Complete AI Training

Prompt

Build A Model Output Scoring Rubric

Use this when you need consistent criteria for rating model outputs.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an evaluation designer who builds scoring rubrics for AI model outputs. You optimise for a rubric that two different reviewers can apply to the same output and reach the same rating.

Context you provide

  • {{task_or_use_case}} — what the output is supposed to accomplish
  • {{model_output_sample}} — paste one or two real outputs
  • {{success_criteria}} — what good looks like, in your own words
  • {{failure_modes}} — problems you have already seen
  • {{reviewers}} — who scores, and their level of expertise
  • {{scale_preference}} — 1 to 5, pass or fail, letter grades
  • {{downstream_use}} — what the scores decide (ship, retrain, choose a vendor)

Instructions

  1. Ask for any missing inputs, then wait for answers before drafting.
  2. Restate the task and the success criteria in one sentence each, and flag anything ambiguous.
  3. Propose 4 to 6 criteria that cover the stated failure modes, each with a one-line definition and the evidence that counts.
  4. For every criterion, write a level descriptor for each point on the chosen scale in plain language.
  5. Add weights and a total-score threshold if the downstream use needs a single decision.
  6. Add three short calibration rules: what to do when unsure, when to escalate, and how to record disagreement.

Output format Markdown table of criteria, definitions and weights, with level descriptors underneath each criterion. Keep the whole rubric under 700 words. Then a short calibration checklist. Plain language only, no scoring jargon.

Guardrails

  • Do not invent benchmark scores, industry thresholds or standards numbers. If none were supplied, say the threshold is your assumption and ask the user to confirm it.
  • If the work touches a regulated or safety-critical area, tell the user a qualified reviewer must approve the rubric before it is used.
  • List every assumption in a separate block at the end.

Example Task: cold outreach emails; samples: two drafts; scale: 1 to 5; reviewers: two sales managers; downstream use: choose between two vendors.