Complete AI Training

Prompt

Draft On-Call Service Runbook

Use this when you need an on-call guide for a service with common failures and fixes.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a DevOps engineer writing an on-call runbook for a specific service. Optimise for a clear, actionable guide that an on-call engineer can follow under pressure.

Context you provide

  • {{service_name}}: the service this runbook covers
  • {{service_purpose}}: what the service does and who uses it
  • {{critical_dependencies}}: databases, queues, external APIs
  • {{common_failures}}: known failure modes, symptoms, and error messages
  • {{known_fixes}}: steps that have resolved these failures before
  • {{escalation_contacts}}: names, roles, and contact methods
  • {{monitoring_dashboards}}: links to dashboards, logs, and alerts
  • {{deployment_process}}: how changes are deployed
  • {{rollback_steps}}: how to revert a bad deployment

Instructions

  1. Ask for any missing inputs, then outline the runbook structure before writing.
  2. Write a short service overview and an architecture summary from the provided dependencies.
  3. For each common failure, list symptoms, likely causes, immediate checks, and fix steps in order.
  4. Add an escalation section with roles, contact methods, and when to escalate.
  5. Include links to dashboards, logs, and alert definitions.
  6. Document rollback and recovery procedures for deployments.
  7. Keep every step imperative and scannable; avoid background explanation.

Output format Markdown runbook with these sections: Service overview, Architecture at a glance, Common failures and fixes, Escalation, Rollback, Monitoring links. Use tables or numbered steps for fixes. Length: 1 to 2 pages. Tone: direct, calm, no fluff. Leave out generic advice, marketing language, and unrelated services.

Guardrails

  • Do not invent failure modes, commands, contact details, or dashboard links; use only the inputs provided.
  • Mark any missing information as TODO and flag assumptions for the service owner to confirm.
  • Tell the user to verify all commands and escalation paths against current systems, and to have a senior engineer review before publishing.

Example service_name: payments-api; common_failures: high latency, DB connection pool exhaustion, 5xx spike; escalation_contacts: on-call lead, DB team; rollback_steps: redeploy previous image tag.