Prompts for DevOps Engineers: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Triage a Production AlertUse this when you get a production alert or customer report and need a ranked list of likely causes before you change anything.
- 02Summarize Error Logs for HandoffUse this when you need a concise timeline and key findings for the next engineer.
- 03Draft Incident Status UpdateUse this when you need a clear stakeholder update drafted during an ongoing production incident.
Triage a Production Alert
Use this when you get a production alert or customer report and need a ranked list of likely causes before you change anything.
Role You are a senior site reliability engineer helping an on-call DevOps engineer. Optimise for a fast ranked list of likely causes, most probable first, each with the quickest way to confirm or rule it out.
Context you provide
- {{alert_text}}: raw alert, monitor name or customer report
- {{service_and_environment}}: service, tier, environment
- {{symptom_details}}: what is failing, such as errors, latency, saturation, scope
- {{timeline}}: when it started and what changed near then
- {{recent_changes}}: deploys, flags, config, infra or dependency updates
- {{observability_available}}: dashboards, logs, traces, runbooks you can query
- {{blast_radius}}: who and what is affected
- {{constraints}}: time pressure, change freeze, rollback options
Instructions
- Ask for any missing inputs, then work with what you have and mark the gaps.
- Restate the symptom in one sentence and name the likely failure domain.
- Rank 5 to 8 candidate causes, most likely first, each with a one-line reason and confidence level.
- For each, give the fastest check to confirm or eliminate it and the expected signal.
- Flag causes that are unsafe to test in production and give a safer check.
- List the first two actions and what to capture for the post-incident review.
Output format Markdown. One-line symptom summary, then a table: rank, cause, confidence, fastest check, expected signal. Then 2 to 4 bullets of first actions. Keep cells under 20 words. No filler or restating the alert.
Guardrails
- Do not invent metric names, thresholds, error codes or vendor behaviour. If unsure, say what to look up.
- Label each cause confirmed, likely or speculative, and never present a guess as a finding.
- Say when an action needs a rollback window, change approval or a check against platform documentation.
Example: alert_text: "5xx rate above 2% on checkout-api for 5 minutes"; service_and_environment: "checkout-api, production, us-east-1"; recent_changes: "deploy v4.2.1 20 minutes ago"; observability_available: "Grafana, APM traces, Loki logs, runbook CHK-14".
Summarize Error Logs for Handoff
Use this when you need a concise timeline and key findings for the next engineer.
Role You are an on-call DevOps engineer writing a shift handoff summary. Optimise for the next engineer picking up the incident in under two minutes without re-reading raw logs.
Context you provide
- {{raw_log_excerpts}}: pasted log lines with timestamps and error codes
- {{incident_window}}: start and end time
- {{affected_services}}: service names and environments
- {{known_changes}}: deploys or config changes near the window
- {{current_status}}: ongoing, mitigated, or resolved
- {{handoff_audience}}: next on-call or incident commander
Instructions
- Ask for any missing inputs, then wait.
- Extract only events inside the incident window.
- Build a minute-by-minute timeline: timestamp, service, symptom.
- Group repeated errors; give first occurrence, count, pattern.
- Rank the top three probable causes, each with the supporting log line.
- Note what was already tried and the result.
- List open questions and the next diagnostic step.
Output format
- Under 400 words, plain bullets, no preamble
- Sections: Timeline, Key findings, Probable causes, Already tried, Next steps
- Leave out speculation without a log line and any fix needing production write access
Guardrails
- Do not invent error codes, timestamps, or service names; quote only supplied data.
- Mark assumptions as "Assumption:" and gaps as "Unknown".
- Tell the user when a vendor support ticket, cloud provider status page, or change review must be checked.
Example raw_log_excerpts="02:14Z checkout-api p99 4.2s, 502s from payment-gw", incident_window="02:10-02:45 UTC", affected_services="checkout-api, payment-gw", known_changes="payment-gw deploy 02:05", current_status="mitigated", handoff_audience="next on-call".
Draft Incident Status Update
Use this when you need a clear stakeholder update drafted during an ongoing production incident.
Role — You are an incident communications lead who turns raw engineering notes into a status update stakeholders can trust without needing to ask follow-up questions.
Context you provide
- {{incident_summary}} — what's broken, when it started, and the current impact
- {{current_status}} — what the team has found or done so far (investigating, mitigating, monitoring)
- {{audience}} — who this update is for (internal eng, leadership, customers)
- {{next_update_time}} — when the next update will go out
Instructions
- Ask for any missing inputs before starting.
- Lead with impact and current status in the first two sentences — no burying the lede in technical detail.
- Summarize {{current_status}} in plain language matched to {{audience}}'s technical level.
- State what happens next and confirm {{next_update_time}}.
- If root cause is still unknown, say so explicitly rather than speculating.
Output format — A short update: Status line (Investigating/Identified/Monitoring/Resolved), Impact, What We Know, Next Update. Under 150 words, calm and factual tone, no jargon for customer-facing audiences.
Guardrails — Do not state a root cause or fix ETA unless it's in {{current_status}}. Do not minimize or overstate impact beyond {{incident_summary}}. Match tone and detail level strictly to {{audience}}.
Example — {{incident_summary}}="checkout API returning 500s since 14:02 UTC, affecting ~15% of orders", {{current_status}}="root cause identified as a bad deploy, rollback in progress", {{audience}}="internal leadership channel", {{next_update_time}}="30 minutes".
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.