Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for prompt engineers

Golden Test Set Maintenance Agent

Keep the test set current so real failures become permanent tests

Golden Test Set Maintenance Agent: what goes in, what the agent does and what you get

What it does

The golden test set was written six months ago, the product changed, and real failures keep escaping because they are not in the set. This agent reads production failures, thumbs-down feedback and support escalations. It turns each into a draft test case with the input, the expected behavior and the grading rule. It runs each draft on the current prompt to check that it fails, since a test that already passes adds nothing, and removes duplicates and near-duplicates of existing cases. It adds the surviving cases to a staging copy, reruns the whole suite and reports the new pass rate. The engineer approves the additions. Edge case: a failure came from bad user input that no prompt can fix, so the agent excludes it.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedYes, continueNoNo 1 STARTS WHEN Weekly run 2 USES A TOOL Read new failures, feedback and escalations 3 DOES Draft a test case for each real failure 4 USES A TOOL Run each draft on the current prompt 5 CHECKS THE RESULT Does the draft fail on the current prompt? If not: drop it or rewrite it to capture the realfailure. Back to step 3. 6 USES A TOOL Compare with existing cases and remove duplicates 7 DOES Exclude cases the prompt cannot fix, such asunreadable input 8 YOU APPROVE Engineer approves the additions 9 USES A TOOL Add the cases to staging and rerun the full suite 10 CHECKS THE RESULT Is the grading consistent across two runs? If not: tighten the grading rule or the expected answerand rerun. Back to step 3. 11 RESULT Updated test set and pass rate report
Read the steps as a list
  1. Weekly run
  2. Read new failures, feedback and escalations
  3. Draft a test case for each real failure
  4. Run each draft on the current prompt
  5. Does the draft fail on the current prompt?If not: drop it or rewrite it to capture the real failure. Back to step 3.
  6. Compare with existing cases and remove duplicates
  7. Exclude cases the prompt cannot fix, such as unreadable input
  8. Engineer approves the additionsThe agent waits here for your OK.
  9. Add the cases to staging and rerun the full suite
  10. Is the grading consistent across two runs?If not: tighten the grading rule or the expected answer and rerun. Back to step 3.
  11. Updated test set and pass rate report

How it decides

It keeps a draft case only if it fails on the current prompt, is not a duplicate and tests a behavior the prompt should be able to handle.

  • Add a case only when it fails on the current prompt
  • Treat cases with over 90% text overlap as duplicates
  • Exclude cases caused by missing or unreadable input
  • Cap additions at 25 cases per week

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Failure sources
  • Weekly cap (default 25)
  • Duplicate threshold (default 90%)
  • Grading rubric format
  • Test set location

What keeps you in control

It always asks you first

  • Engineer approves added cases
  • Engineer approves removal of old cases

Hard limits

  • Remove personal data from cases
  • Never delete existing cases without approval

It stops when

  • Done: staging passes grading checks and the set is updated
  • Stop: there are no new real failures this week

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensLast week had 61 thumbs-down logs and 12 escalations. The agent drafted 40 cases. Only 22 failed on the current prompt, and after removing duplicates 16 remained. Three were excluded for blank inputs. After the engineer approved 13, the rerun showed the pass rate fell from 94% to 86%. Grading differed between two runs on 2 cases, so the check failed. The agent tightened the rubric and both runs agreed.

More agents for prompt engineers