Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for prompt engineers

Output Quality Grading Agent

Consistent, rubric-based quality scores for outputs at scale, with reliability checked

Output Quality Grading Agent: what goes in, what the agent does and what you get

What it does

When a prompt runs on hundreds of inputs, grading each output by hand is slow and graders disagree, so quality trends stay unclear. This agent grades a batch of outputs against a written rubric with clear criteria and levels, giving each output a score per criterion and a short reason. To stay reliable, it first grades a set of outputs already graded by a person and measures how often it agrees. If agreement is below your bar, it flags that the rubric or its grading needs tightening before any scores are trusted. Outputs the rubric cannot clearly score go to human review instead of being forced into a level. It summarizes quality across the batch and highlights the weakest criteria. You approve the rubric and review borderline cases.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN Batch of outputs ready to grade 2 USES A TOOL Load the rubric and the human-graded reference set 3 USES A TOOL Grade the reference set against the rubric 4 CHECKS THE RESULT Does it agree closely with the human grades? If not: flag the rubric criteria that disagree and waitfor tightened wording. Back to step 3. 5 USES A TOOL Grade every output in the batch with a reason percriterion 6 CHECKS THE RESULT Can each output be clearly scored on everycriterion? If not: mark the output for human review instead offorcing a level. Back to step 5. 7 DOES Summarize quality and highlight weak criteria 8 YOU APPROVE Engineer approves the rubric and reviews borderlinecases 9 RESULT Quality report with scores
Read the steps as a list
  1. Batch of outputs ready to grade
  2. Load the rubric and the human-graded reference set
  3. Grade the reference set against the rubric
  4. Does it agree closely with the human grades?If not: flag the rubric criteria that disagree and wait for tightened wording. Back to step 3.
  5. Grade every output in the batch with a reason per criterion
  6. Can each output be clearly scored on every criterion?If not: mark the output for human review instead of forcing a level. Back to step 5.
  7. Summarize quality and highlight weak criteria
  8. Engineer approves the rubric and reviews borderline casesThe agent waits here for your OK.
  9. Quality report with scores

How it decides

It scores strictly against the rubric and treats its own grades as trustworthy only when they agree closely with the human reference set.

  • Score strictly against the rubric
  • Trust scores only when agreement is high
  • Send unscoreable outputs to human review

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Rubric and criteria
  • Agreement bar
  • Reference set size
  • Report format

What keeps you in control

It always asks you first

  • Approving the rubric
  • Reviewing borderline cases

Hard limits

  • Does not present scores as reliable below the agreement bar
  • Routes ambiguity to people

It stops when

  • Done: batch graded with reliability shown
  • Stop: no rubric or reference set exists

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensGrading 300 answers at Brookfield Learning in June, the agent first agreed with the human reference on only 79% of cases, below the 85% bar. It flagged a vague level in the citation criterion. The engineer tightened the wording, and the next calibration reached 88%. The agent graded the batch, found citing sources was the weakest criterion and sent 12 ambiguous outputs for human review.

More agents for prompt engineers