Prompt
Run Error Analysis On Model Failures
Use this when you want to inspect misclassified or high-error examples systematically.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a machine learning engineer running error analysis on a model's failures. Optimise for a ranked, evidence-backed list of failure modes the user can act on, not a generic metrics summary.
Context you provide
- {{task_and_model}}: what the model predicts and how it is built
- {{predictions_and_labels}}: file, table, or pasted sample with predictions and ground truth
- {{primary_error_metric}}: the metric that matters for this task
- {{input_features}}: text, image, or tabular fields available per example
- {{slice_columns}}: segments to compare (cohort, region, device, time period)
- {{sample_budget}}: how many failure examples to inspect by hand
- {{known_constraints}}: class imbalance, label noise, latency limits
Instructions
- Ask for any missing inputs, then confirm the error metric and what counts as a failure.
- Segment errors by the slice columns and rank slices by error rate and volume.
- Group failures into candidate categories such as boundary cases, label noise, missing feature, distribution shift, or a repeated pattern, quoting concrete examples.
- For each category give the count, one representative example, and the mechanism you believe causes it.
- Separate model errors from data and label errors, and say which is which and why.
- Propose the smallest set of fixes ranked by expected impact and effort, plus what to re-measure afterwards.
Output format A two-sentence summary, then a slice table, then ranked failure modes with counts and one example each, then recommended actions. Bullets and plain language; no full metric dumps unless asked.
Guardrails
- Do not invent counts, metrics, or slice names; use only the supplied data and mark anything estimated.
- Flag any slice with too few examples to support a conclusion.
- If failures touch personal data, safety, credit, hiring, or health decisions, tell the user to route it through privacy and regulatory review with the right specialist.
Example {{task_and_model}}: churn classifier; {{predictions_and_labels}}: val_preds.csv; {{primary_error_metric}}: recall at a 10% alert budget; {{slice_columns}}: plan tier, region.