Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for quality assurance testers

Flaky Test Reproduction Agent

Every flaky test ends with a reproducible cause or a documented set of ruled-out causes.

Flaky Test Reproduction Agent: what goes in, what the agent does and what you get

What it does

Flaky tests pass sometimes and fail other times, and engineers waste hours rerunning them. When CI history flags a test as flaky, this agent reads the failure traces and forms a hypothesis about the cause: random seed, test order, timing or shared state. It runs isolated repetitions that change only that one factor. If the variation reliably triggers the failure, it narrows the setup to a minimal reproduction and writes a candidate test-only patch on a branch. If not, it switches to the next hypothesis within its rerun budget. It never disables tests or changes production code. The QA lead reviews. Edge case: a test fails only when run after another test that leaves a temporary file behind, which the agent confirms by fixing the order.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedNo 1 STARTS WHEN Test flagged as flaky 2 USES A TOOL Read CI history and failure traces 3 DOES Form a hypothesis: seed, order, timing or sharedstate 4 USES A TOOL Run isolated repetitions varying that one factor 5 CHECKS THE RESULT Does the variation reliably trigger the failure? If not: switch to the next hypothesis (within budget).Back to step 3. 6 DOES Narrow the fixture to a minimal reproduction 7 USES A TOOL Write a candidate test-only patch on a branch 8 YOU APPROVE QA lead reviews the patch 9 RESULT Flaky-test investigation bundle
Read the steps as a list
  1. Test flagged as flaky
  2. Read CI history and failure traces
  3. Form a hypothesis: seed, order, timing or shared state
  4. Run isolated repetitions varying that one factor
  5. Does the variation reliably trigger the failure?If not: switch to the next hypothesis (within budget). Back to step 3.
  6. Narrow the fixture to a minimal reproduction
  7. Write a candidate test-only patch on a branch
  8. QA lead reviews the patchThe agent waits here for your OK.
  9. Flaky-test investigation bundle

How it decides

It picks the variation most suggested by the failure trace (timeouts suggest timing, order-dependent failures suggest shared state) and keeps the hypothesis only if controlled reruns confirm it.

  • Hypothesis order: guided by the failure signature.
  • Confirmed: failure appears consistently under the variation and not without it.
  • Stop: rerun budget reached.

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Rerun budget per test (default 200 runs)
  • Hypothesis order (default guided by the failure signature)
  • Flakiness rate that triggers investigation (default 2%)
  • Who reviews patches (default QA lead)

What keeps you in control

It always asks you first

  • Disabling tests
  • Merging fixes
  • Changing production behavior

Hard limits

  • Never disables tests.
  • Test-only patches.

It stops when

  • Done: minimal reproduction found.
  • Budget reached: report ruled-out causes.

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensA checkout test failed in 7% of CI runs. The agent first tested random seeds over 50 runs: no pattern, so the check failed and it moved on. Running the test after the cart suite failed 50 of 50 times, while running it alone passed every time. It narrowed the cause to a shared cache, wrote a cleanup fixture and the QA lead approved the patch.

More agents for quality assurance testers