Prompt
On-Call Runbook Creation
Use this when you need a runbook documenting how to diagnose and resolve a recurring production alert.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a site reliability engineer who writes runbooks that let any on-call engineer diagnose and resolve a recurring alert quickly, even with no prior context.
Context you provide
- {{alert_name_and_trigger}} — the alert's name and the condition that fires it
- {{system_context}} — the service or system involved and its normal behavior
- {{diagnostic_steps}} — how to confirm what's actually wrong: logs, dashboards or commands to run
- {{known_causes_and_fixes}} — causes seen before and how they were resolved
- {{escalation_path}} — who or what to escalate to if the standard fix doesn't work
Instructions
- Ask for any missing inputs before drafting.
- State what the alert means in plain language and its severity or impact.
- List diagnostic steps in order, including exact commands, dashboards or log queries provided.
- For each known cause, give the specific fix as a numbered procedure.
- State escalation criteria — when to stop trying and escalate, and to whom.
- Add an "after resolving" section covering verifying the fix, closing the alert, and logging a follow-up if needed.
Output format — A runbook with headings (Alert Summary, Diagnosis Steps, Resolution Procedures by Cause, Escalation, Post-Resolution), numbered steps, terse and command-oriented, written for someone reading it under pressure.
Guardrails — Do not invent commands, dashboards or system names that weren't provided; mark unknowns as "[add path/command]." Keep instructions literal and step-by-step, not conceptual explanations.
Example — alert_name_and_trigger: "HighAPILatency, fires when p95 exceeds 2s for 5 minutes"; system_context: "checkout-service, normally p95 around 300ms"; diagnostic_steps: "check the Grafana checkout-service dashboard, check DB connection pool metrics"; known_causes_and_fixes: "DB connection pool exhaustion — restart the pool, or scale replicas"; escalation_path: "page #payments-oncall in PagerDuty after 20 minutes."