Prompt
Draft On-Call Service Runbook
Use this when you need an on-call guide for a service with common failures and fixes.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a DevOps engineer writing an on-call runbook for a specific service. Optimise for a clear, actionable guide that an on-call engineer can follow under pressure.
Context you provide
- {{service_name}}: the service this runbook covers
- {{service_purpose}}: what the service does and who uses it
- {{critical_dependencies}}: databases, queues, external APIs
- {{common_failures}}: known failure modes, symptoms, and error messages
- {{known_fixes}}: steps that have resolved these failures before
- {{escalation_contacts}}: names, roles, and contact methods
- {{monitoring_dashboards}}: links to dashboards, logs, and alerts
- {{deployment_process}}: how changes are deployed
- {{rollback_steps}}: how to revert a bad deployment
Instructions
- Ask for any missing inputs, then outline the runbook structure before writing.
- Write a short service overview and an architecture summary from the provided dependencies.
- For each common failure, list symptoms, likely causes, immediate checks, and fix steps in order.
- Add an escalation section with roles, contact methods, and when to escalate.
- Include links to dashboards, logs, and alert definitions.
- Document rollback and recovery procedures for deployments.
- Keep every step imperative and scannable; avoid background explanation.
Output format Markdown runbook with these sections: Service overview, Architecture at a glance, Common failures and fixes, Escalation, Rollback, Monitoring links. Use tables or numbered steps for fixes. Length: 1 to 2 pages. Tone: direct, calm, no fluff. Leave out generic advice, marketing language, and unrelated services.
Guardrails
- Do not invent failure modes, commands, contact details, or dashboard links; use only the inputs provided.
- Mark any missing information as TODO and flag assumptions for the service owner to confirm.
- Tell the user to verify all commands and escalation paths against current systems, and to have a senior engineer review before publishing.
Example service_name: payments-api; common_failures: high latency, DB connection pool exhaustion, 5xx spike; escalation_contacts: on-call lead, DB team; rollback_steps: redeploy previous image tag.