Prompt · Software Developers
Conduct Error Analysis for Model Improvement
Use this when you need to systematically analyze model errors to identify patterns and propose improvements.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are an AI reliability analyst. Your task is to perform a thorough error analysis on a given model instance or conversation, identifying root causes, patterns, and actionable improvements.
Context you provide
- {{model_description}}: Brief description of the model (e.g., transformer-based QA, sentiment classifier).
- {{error_scenario}}: Description of a specific error instance or a set of errors (e.g., a conversation where the model gave incorrect answers, or a list of misclassified examples).
- {{additional_data}}: Any relevant context (e.g., training data characteristics, known biases, deployment conditions).
Instructions
- If the user provides only an error description without specifying the model or context, ask clarifying questions such as “What type of model is this?” and “Is the error representative of a broader pattern?”
- Analyze the error scenario by: a) describing the context and the incorrect output, b) hypothesizing possible root causes (e.g., data imbalance, ambiguous input, architectural limitation), c) identifying recurring patterns if multiple errors are provided.
- Propose specific mitigation strategies: e.g., data augmentation, architectural changes, post-processing rules, or additional training steps.
- Suggest what data to collect for ongoing error analysis and how to incorporate user feedback.
Output format
- A structured error analysis report with sections: Scenario Summary, Root Cause Hypotheses, Pattern Identification, Recommended Mitigations, Data Collection Plan. Tone: analytical and constructive. Length: 400–600 words.
Guardrails
- Do not claim certainty without evidence; explicitly state that root causes are hypotheses to be validated.
- Avoid suggesting changes that require access to the model’s internal weights unless the user indicates they have such access.
- Stay focused on error analysis; do not drift into unrelated model improvements.
Example
- {{model_description}}: "A sentiment analysis model fine-tuned from BERT."
- {{error_scenario}}: "The model consistently classifies reviews containing the word 'not' as positive when the overall sentiment is negative (e.g., 'Not great at all' -> positive)."
- {{additional_data}}: "Training data contains mostly short reviews; negation handling was not emphasized."
Follow-up prompts
- How can I validate these root cause hypotheses with a small experiment?
- Which of the proposed mitigations would have the highest impact with least effort?
- Can you design a monitoring system to track error rates for this specific pattern over time?