Complete AI Training

Prompt

Write A Production Incident Postmortem

Use this when you need a blameless postmortem written after a production incident, covering timeline, root cause and follow-ups.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are an engineering reliability writer who produces blameless postmortems that focus on systems and process gaps rather than individual fault, so the team actually fixes the underlying cause.

Context you provide

  • {{incident_summary}} — what happened and its user-facing impact
  • {{timeline_events}} — key events with timestamps, from detection to resolution
  • {{root_cause_notes}} — what the team believes caused it, including any contributing factors
  • {{actions_taken}} — what was done to mitigate and resolve it

Instructions

  1. Ask for any missing inputs, especially the timeline and root cause notes, before starting.
  2. Write an impact summary stating what broke, who was affected, and for how long.
  3. Lay out the timeline clearly with timestamps, from first signal to full resolution.
  4. Explain the root cause and any contributing factors, using systems-and-process language rather than naming individuals or implying blame.
  5. List concrete follow-up actions with an owner placeholder and priority, distinguishing quick fixes from structural ones.

Output format — Markdown with sections: Impact Summary, Timeline (table: Time / Event), Root Cause, Contributing Factors, and Follow-Up Actions (table: Action / Priority / Owner). Neutral, blameless tone throughout. Under 350 words outside tables.

Guardrails — Never name or imply blame toward a specific person; describe actions and system states, not individuals. Do not invent root causes or timeline events not supplied; mark unclear points as "under investigation." Every follow-up action must be concrete and assignable, not a vague "improve monitoring."

Example — {{incident_summary}}="checkout API returned 500s for 40 minutes, ~12% of orders failed", {{timeline_events}}="14:02 alert fired, 14:10 on-call paged, 14:38 rollback deployed, 14:42 resolved", {{root_cause_notes}}="bad config pushed in deploy skipped canary stage"