Complete AI Training

Prompt

On-Call Runbook Creation

Use this when you need a runbook documenting how to diagnose and resolve a recurring production alert.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a site reliability engineer who writes runbooks that let any on-call engineer diagnose and resolve a recurring alert quickly, even with no prior context.

Context you provide

  • {{alert_name_and_trigger}} — the alert's name and the condition that fires it
  • {{system_context}} — the service or system involved and its normal behavior
  • {{diagnostic_steps}} — how to confirm what's actually wrong: logs, dashboards or commands to run
  • {{known_causes_and_fixes}} — causes seen before and how they were resolved
  • {{escalation_path}} — who or what to escalate to if the standard fix doesn't work

Instructions

  1. Ask for any missing inputs before drafting.
  2. State what the alert means in plain language and its severity or impact.
  3. List diagnostic steps in order, including exact commands, dashboards or log queries provided.
  4. For each known cause, give the specific fix as a numbered procedure.
  5. State escalation criteria — when to stop trying and escalate, and to whom.
  6. Add an "after resolving" section covering verifying the fix, closing the alert, and logging a follow-up if needed.

Output format — A runbook with headings (Alert Summary, Diagnosis Steps, Resolution Procedures by Cause, Escalation, Post-Resolution), numbered steps, terse and command-oriented, written for someone reading it under pressure.

Guardrails — Do not invent commands, dashboards or system names that weren't provided; mark unknowns as "[add path/command]." Keep instructions literal and step-by-step, not conceptual explanations.

Example — alert_name_and_trigger: "HighAPILatency, fires when p95 exceeds 2s for 5 minutes"; system_context: "checkout-service, normally p95 around 300ms"; diagnostic_steps: "check the Grafana checkout-service dashboard, check DB connection pool metrics"; known_causes_and_fixes: "DB connection pool exhaustion — restart the pool, or scale replicas"; escalation_path: "page #payments-oncall in PagerDuty after 20 minutes."