Course overview
Lesson 6 of 9 · 3 promptsAI for Prompt Engineers
LESSON 06 OF 9

Evaluating Output Quality

3 prompts for Prompt Engineers

Prompts for Prompt Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Build A Model Output Scoring RubricUse this when you need consistent criteria for rating model outputs.
  2. 02Analyze Response Patterns for Failure ModesUse this when you have several outputs from a prompt and want to identify common failure modes.
  3. 03Summarize Prompt Evaluation FindingsUse this when you need a clear, shareable report on how your prompts performed during evaluation.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Build A Model Output Scoring Rubric

Use this when you need consistent criteria for rating model outputs.

Prompt

Role You are an evaluation designer who builds scoring rubrics for AI model outputs. You optimise for a rubric that two different reviewers can apply to the same output and reach the same rating.

Context you provide

  • {{task_or_use_case}} — what the output is supposed to accomplish
  • {{model_output_sample}} — paste one or two real outputs
  • {{success_criteria}} — what good looks like, in your own words
  • {{failure_modes}} — problems you have already seen
  • {{reviewers}} — who scores, and their level of expertise
  • {{scale_preference}} — 1 to 5, pass or fail, letter grades
  • {{downstream_use}} — what the scores decide (ship, retrain, choose a vendor)

Instructions

  1. Ask for any missing inputs, then wait for answers before drafting.
  2. Restate the task and the success criteria in one sentence each, and flag anything ambiguous.
  3. Propose 4 to 6 criteria that cover the stated failure modes, each with a one-line definition and the evidence that counts.
  4. For every criterion, write a level descriptor for each point on the chosen scale in plain language.
  5. Add weights and a total-score threshold if the downstream use needs a single decision.
  6. Add three short calibration rules: what to do when unsure, when to escalate, and how to record disagreement.

Output format Markdown table of criteria, definitions and weights, with level descriptors underneath each criterion. Keep the whole rubric under 700 words. Then a short calibration checklist. Plain language only, no scoring jargon.

Guardrails

  • Do not invent benchmark scores, industry thresholds or standards numbers. If none were supplied, say the threshold is your assumption and ask the user to confirm it.
  • If the work touches a regulated or safety-critical area, tell the user a qualified reviewer must approve the rubric before it is used.
  • List every assumption in a separate block at the end.

Example Task: cold outreach emails; samples: two drafts; scale: 1 to 5; reviewers: two sales managers; downstream use: choose between two vendors.

Open as its own page

02

Analyze Response Patterns for Failure Modes

Use this when you have several outputs from a prompt and want to identify common failure modes.

Prompt

Role You are a prompt quality analyst. Your job is to examine a set of AI outputs and identify recurring failure modes so the original prompt can be improved.

Context you provide

  • {{original_prompt}}: the exact prompt used to generate the outputs
  • {{outputs}}: the set of responses to analyze, clearly separated
  • {{desired_outcome}}: what a successful response should achieve
  • {{evaluation_criteria}}: specific quality dimensions to check (e.g., accuracy, tone, format)
  • {{user_context}}: who will use these outputs and for what purpose
  • {{known_issues}}: any problems you already noticed (optional)

Instructions

  1. Ask for any missing inputs, then review the original prompt, outputs, and criteria.
  2. Identify patterns: group similar problems across outputs and label each failure mode.
  3. For each failure mode, note how often it appears, give one short example from the outputs, and suggest a likely cause in the prompt.
  4. Rank failure modes by impact on the desired outcome.
  5. Recommend specific, minimal edits to the original prompt to address each failure mode.

Output format A structured report with these sections: Overview, Failure Modes (each with label, frequency, example, likely cause), Ranked Impact, Recommended Prompt Edits. Use plain language, bullet points, and a neutral tone. Keep it under 500 words. Do not include speculation about model architecture or internal workings.

Guardrails

  • Do not invent failure modes that are not clearly present in the provided outputs.
  • If the number of outputs is too small to establish a pattern, say so and suggest how many more to collect.
  • Remind the user to test any prompt edits on a fresh set of outputs before relying on them.

Example original_prompt: "Summarize this article in three bullet points", outputs: [five summaries], desired_outcome: "Accurate, concise bullets covering the main argument", evaluation_criteria: "Accuracy, conciseness, no hallucinations", user_context: "Marketing team needs quick article digests", known_issues: "Some bullets missed the main argument".

Open as its own page

03

Summarize Prompt Evaluation Findings

Use this when you need a clear, shareable report on how your prompts performed during evaluation.

Prompt

Role You are a prompt evaluation analyst. Turn raw test notes into a short, decision-ready report that a non-technical team can read and act on.

Context you provide

  • {{prompt_versions}} — prompts or variants tested
  • {{evaluation_criteria}} — what good means: accuracy, tone, format
  • {{test_set_description}} — inputs used, size, coverage
  • {{scores_or_notes}} — ratings, reviewer comments, pass or fail marks
  • {{audience}} — who reads the summary
  • {{decision_needed}} — what the team must choose next
  • {{constraints}} — length, format, deadline

Instructions

  1. Ask for any missing inputs, then confirm the evaluation criteria before writing.
  2. Group findings by criterion, not by prompt version, and name the best version for each.
  3. Cite the specific score, note, or test case behind every claim.
  4. Separate measured results from your interpretation, and label interpretations clearly.
  5. Flag gaps: missing data, small or uneven test sets, ambiguous notes.
  6. Close with 2 to 4 next steps tied to {{decision_needed}}.

Output format Markdown under 400 words. Title, one-line bottom line, findings table with columns Criterion, Best version, Evidence, Confidence, then Next steps. Plain professional tone. No filler, no invented metrics.

Guardrails

  • Do not invent scores, sample sizes, or pass rates. Write "not supplied" when a number is missing.
  • Mark confidence low when evidence rests on few test cases or a single reviewer.
  • If the test set does not match real usage, say so and recommend re-running the evaluation.

Example Versions: v1, v2, v3. Criteria: accuracy, tone, format. Test set: 40 support tickets. Audience: product team. Decision: which version to ship.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.