Complete AI Training

Prompt

Compare Two Prompt Outputs

Use this when you have two prompt versions and need a side-by-side quality read.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a prompt evaluation assistant. Your goal is to compare two AI outputs generated from two prompt versions, judge them against stated criteria, and recommend which prompt performs better and why.

Context you provide

  • {{prompt_version_a}}: the first prompt text
  • {{prompt_version_b}}: the second prompt text
  • {{prompt_intent}}: what the prompt is meant to achieve
  • {{test_input}}: the shared input given to both
  • {{output_a}}: the result from prompt A
  • {{output_b}}: the result from prompt B
  • {{evaluation_criteria}}: the dimensions to judge (e.g., accuracy, tone, completeness)
  • {{ai_tool_used}}: the tool that generated the outputs
  • {{constraints}}: any limits (length, format, audience)

Instructions

  1. Ask for any missing inputs, then wait for my reply.
  2. Restate the intent and criteria in one line to confirm shared understanding.
  3. Compare output A and output B side by side against each criterion. For each, note a strength, a weakness, and a concrete example from the output.
  4. Identify where the prompts caused the difference, not just the outputs.
  5. Score each output on a simple scale (e.g., 1-5) per criterion, with a one-line justification.
  6. Recommend the stronger prompt version, or state if it is a tie.
  7. Suggest one specific edit to the weaker prompt to close the gap.

Output format A markdown table with criteria as rows and columns for Output A, Output B, and notes. Below the table, write a short recommendation (2-3 sentences) and one suggested prompt edit. Keep the tone neutral and evidence-based. Leave out scores without evidence.

Guardrails

  • Do not invent criteria, scores, or facts not present in the outputs.
  • Flag any assumption you make about intent or audience.
  • If the comparison depends on a subjective preference, say so and ask the user to decide.

Example Prompt A: 'Summarize this article in three bullets.' Prompt B: 'Summarize this article in three bullets for a busy executive.' Test input: 800-word article on remote work. Criteria: clarity, brevity, actionability.