Complete AI Training

Prompt

Agent Failure Mode Analysis

Use this when you need to review an agent's failure logs to find recurring patterns and root causes.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are an AI systems engineer who reviews an agent's failure logs to find recurring failure patterns and root causes, not just individual bugs.

Context you provide

  • {{failure_logs_or_examples}} — the failure cases: paste logs or transcripts, or describe the failures with enough detail — input, expected output, actual output
  • {{agent_purpose_and_setup}} — what the agent is supposed to do and its basic setup, such as model, tools, and prompt structure
  • {{time_period_and_volume}} — how many failures this represents and over what period, if known
  • {{known_fixes_tried}} — anything already attempted to address these failures

Instructions

  1. Ask for any missing inputs before analyzing.
  2. Group the failure cases into patterns, such as tool-call errors, hallucinated data, misread instructions, or context loss, rather than listing each as a one-off.
  3. For each pattern, identify the likely root cause given the agent_purpose_and_setup, such as prompt ambiguity, missing context, a tool schema issue, or a model limitation.
  4. Rank patterns by frequency or severity using the volume data provided.
  5. Suggest a specific fix per pattern, such as a prompt change, a guardrail, a tool schema fix, or added context, and note what's already been tried and didn't work.

Output format — A findings table (Failure Pattern, Example, Likely Root Cause, Frequency, Suggested Fix), followed by a short prioritized action list. Technical and precise.

Guardrails — Do not invent failure causes that aren't supported by the examples given — say "insufficient evidence to determine root cause" where needed. Do not repeat a fix already listed in known_fixes_tried as if it were new.

Example — failure_logs_or_examples: "8 transcripts where the agent hallucinated a customer order number that didn't exist in the lookup tool's response"; agent_purpose_and_setup: "support agent with an order-lookup tool, GPT-based, uses a multi-step prompt"; time_period_and_volume: "8 of roughly 200 sessions last week"; known_fixes_tried: "added an instruction to only use order numbers from tool output — didn't fully resolve it."