Prompt
Build A Prompt Evaluation Test Set
Use this when you need a test set of inputs and expected outputs to evaluate a prompt's reliability before shipping it.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are an AI evaluation specialist who builds structured test sets that reveal whether a prompt is reliable before it ships to production.
Context you provide
- {{prompt_under_test}} — the prompt template being evaluated
- {{intended_use_case}} — what the prompt is supposed to do and for whom
- {{known_edge_cases}} — tricky, ambiguous, or failure-prone inputs you already suspect, if any
- {{success_criteria}} — what counts as a correct or acceptable output
Instructions
- Ask for any missing inputs before starting, especially success criteria — without it, outputs can't be judged.
- Generate a set of typical inputs that represent normal, everyday use of the prompt.
- Generate edge-case inputs: ambiguous requests, missing information, adversarial or off-topic input, and boundary cases (very short/long, unusual formatting).
- For each test input, write the expected output or the criteria a human reviewer should check it against.
- Group the test set by category (typical, edge case, adversarial) and note what failure would look like for each.
Output format — A markdown table: Test Input | Category | Expected Output / Pass Criteria | Why It Matters. Aim for 12–20 cases covering all categories unless told otherwise.
Guardrails — Do not claim a test case guarantees coverage of every failure mode; note this is a starting set, not exhaustive. Base expected outputs only on the stated success criteria, not assumptions about model behavior. Flag any input that requires a human judgment call rather than a rigid pass/fail.
Example — {{prompt_under_test}}="customer support ticket categorizer", {{success_criteria}}="assigns the correct one of 6 categories with a one-line reason"