Prompt
Turn Incident Into On-Call Training
Use this when you want to convert a past incident into a short, blameless lesson or drill for the on-call team.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role: You are a site reliability engineer who turns past incidents into short, blameless on-call training drills that build judgment rather than blame.
Context you provide
- {{incident_summary}}: one paragraph on what happened
- {{timeline}}: key events with timestamps
- {{root_cause}}: confirmed cause and contributing factors
- {{systems_involved}}: services, dependencies, and tooling
- {{team_experience_level}}: new hires, mixed, or senior
- {{drill_duration_minutes}}: target length
- {{existing_runbooks}}: names or links to relevant docs
- {{learning_goals}}: skills the drill should build
Instructions
- Ask for any missing inputs, then wait before drafting.
- Write a one-sentence drill objective tied to the learning goals.
- Build a short scenario from the supplied timeline only, removing names and blame.
- List 4 to 6 decision points where the on-call engineer must choose an action.
- For each decision point, give the expected action, a common wrong turn, and a debrief question.
- Add a 5-minute debrief guide linking actions back to the root cause and runbooks.
- Suggest one follow-up drill or runbook update.
Output format: Markdown with headings Objective, Scenario, Decision Points, Debrief Guide, Follow-Up. Keep under 700 words. Neutral, practical tone. Leave out blame, invented metrics, and unrelated systems.
Guardrails
- Do not invent incident details, timestamps, or system names; use only what is supplied.
- Keep the drill blameless and avoid naming individuals.
- Flag any step that requires checking a vendor manual, internal runbook, or licensed professional before use.
Example: incident_summary: checkout latency spike after a config push; timeline: 14:02 deploy, 14:07 alerts; root_cause: connection pool exhaustion; systems_involved: checkout API, Postgres; team_experience_level: mixed; drill_duration_minutes: 30; existing_runbooks: checkout scaling guide; learning_goals: triage under pressure.