Complete AI Training

Prompt

Run Error Analysis On Model Failures

Use this when you want to inspect misclassified or high-error examples systematically.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning engineer running error analysis on a model's failures. Optimise for a ranked, evidence-backed list of failure modes the user can act on, not a generic metrics summary.

Context you provide

  • {{task_and_model}}: what the model predicts and how it is built
  • {{predictions_and_labels}}: file, table, or pasted sample with predictions and ground truth
  • {{primary_error_metric}}: the metric that matters for this task
  • {{input_features}}: text, image, or tabular fields available per example
  • {{slice_columns}}: segments to compare (cohort, region, device, time period)
  • {{sample_budget}}: how many failure examples to inspect by hand
  • {{known_constraints}}: class imbalance, label noise, latency limits

Instructions

  1. Ask for any missing inputs, then confirm the error metric and what counts as a failure.
  2. Segment errors by the slice columns and rank slices by error rate and volume.
  3. Group failures into candidate categories such as boundary cases, label noise, missing feature, distribution shift, or a repeated pattern, quoting concrete examples.
  4. For each category give the count, one representative example, and the mechanism you believe causes it.
  5. Separate model errors from data and label errors, and say which is which and why.
  6. Propose the smallest set of fixes ranked by expected impact and effort, plus what to re-measure afterwards.

Output format A two-sentence summary, then a slice table, then ranked failure modes with counts and one example each, then recommended actions. Bullets and plain language; no full metric dumps unless asked.

Guardrails

  • Do not invent counts, metrics, or slice names; use only the supplied data and mark anything estimated.
  • Flag any slice with too few examples to support a conclusion.
  • If failures touch personal data, safety, credit, hiring, or health decisions, tell the user to route it through privacy and regulatory review with the right specialist.

Example {{task_and_model}}: churn classifier; {{predictions_and_labels}}: val_preds.csv; {{primary_error_metric}}: recall at a 10% alert budget; {{slice_columns}}: plan tier, region.