AI agent for quality assurance testers
Flaky Test Reproduction Agent
Every flaky test ends with a reproducible cause or a documented set of ruled-out causes.
What it does
Flaky tests pass sometimes and fail other times, and engineers waste hours rerunning them. When CI history flags a test as flaky, this agent reads the failure traces and forms a hypothesis about the cause: random seed, test order, timing or shared state. It runs isolated repetitions that change only that one factor. If the variation reliably triggers the failure, it narrows the setup to a minimal reproduction and writes a candidate test-only patch on a branch. If not, it switches to the next hypothesis within its rerun budget. It never disables tests or changes production code. The QA lead reviews. Edge case: a test fails only when run after another test that leaves a temporary file behind, which the agent confirms by fixing the order.
How it works
Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.
Read the steps as a list
- Test flagged as flaky
- Read CI history and failure traces
- Form a hypothesis: seed, order, timing or shared state
- Run isolated repetitions varying that one factor
- Does the variation reliably trigger the failure?If not: switch to the next hypothesis (within budget). Back to step 3.
- Narrow the fixture to a minimal reproduction
- Write a candidate test-only patch on a branch
- QA lead reviews the patchThe agent waits here for your OK.
- Flaky-test investigation bundle
How it decides
It picks the variation most suggested by the failure trace (timeouts suggest timing, order-dependent failures suggest shared state) and keeps the hypothesis only if controlled reruns confirm it.
- Hypothesis order: guided by the failure signature.
- Confirmed: failure appears consistently under the variation and not without it.
- Stop: rerun budget reached.
Make it yours
Every agent is a starting point. You choose these settings for your own situation.
- Rerun budget per test (default 200 runs)
- Hypothesis order (default guided by the failure signature)
- Flakiness rate that triggers investigation (default 2%)
- Who reviews patches (default QA lead)
What keeps you in control
It always asks you first
- Disabling tests
- Merging fixes
- Changing production behavior
Hard limits
- Never disables tests.
- Test-only patches.
It stops when
- Done: minimal reproduction found.
- Budget reached: report ruled-out causes.
Set it up
We guide you through the set-up, step by step
Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.
- One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
- The agent then walks you through connecting your own data, one source at a time
- A downloadable copy with the flow chart, the rules and the full guide