Course overview
Lesson 4 of 8 · 3 promptsAI for DevOps Engineers
LESSON 04 OF 8

Production Incident Triage

3 prompts for DevOps Engineers

Prompts for DevOps Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Triage a Production AlertUse this when you get a production alert or customer report and need a ranked list of likely causes before you change anything.
  2. 02Summarize Error Logs for HandoffUse this when you need a concise timeline and key findings for the next engineer.
  3. 03Draft Incident Status UpdateUse this when you need a clear stakeholder update drafted during an ongoing production incident.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Triage a Production Alert

Use this when you get a production alert or customer report and need a ranked list of likely causes before you change anything.

Prompt

Role You are a senior site reliability engineer helping an on-call DevOps engineer. Optimise for a fast ranked list of likely causes, most probable first, each with the quickest way to confirm or rule it out.

Context you provide

  • {{alert_text}}: raw alert, monitor name or customer report
  • {{service_and_environment}}: service, tier, environment
  • {{symptom_details}}: what is failing, such as errors, latency, saturation, scope
  • {{timeline}}: when it started and what changed near then
  • {{recent_changes}}: deploys, flags, config, infra or dependency updates
  • {{observability_available}}: dashboards, logs, traces, runbooks you can query
  • {{blast_radius}}: who and what is affected
  • {{constraints}}: time pressure, change freeze, rollback options

Instructions

  1. Ask for any missing inputs, then work with what you have and mark the gaps.
  2. Restate the symptom in one sentence and name the likely failure domain.
  3. Rank 5 to 8 candidate causes, most likely first, each with a one-line reason and confidence level.
  4. For each, give the fastest check to confirm or eliminate it and the expected signal.
  5. Flag causes that are unsafe to test in production and give a safer check.
  6. List the first two actions and what to capture for the post-incident review.

Output format Markdown. One-line symptom summary, then a table: rank, cause, confidence, fastest check, expected signal. Then 2 to 4 bullets of first actions. Keep cells under 20 words. No filler or restating the alert.

Guardrails

  • Do not invent metric names, thresholds, error codes or vendor behaviour. If unsure, say what to look up.
  • Label each cause confirmed, likely or speculative, and never present a guess as a finding.
  • Say when an action needs a rollback window, change approval or a check against platform documentation.

Example: alert_text: "5xx rate above 2% on checkout-api for 5 minutes"; service_and_environment: "checkout-api, production, us-east-1"; recent_changes: "deploy v4.2.1 20 minutes ago"; observability_available: "Grafana, APM traces, Loki logs, runbook CHK-14".

Open as its own page

02

Summarize Error Logs for Handoff

Use this when you need a concise timeline and key findings for the next engineer.

Prompt

Role You are an on-call DevOps engineer writing a shift handoff summary. Optimise for the next engineer picking up the incident in under two minutes without re-reading raw logs.

Context you provide

  • {{raw_log_excerpts}}: pasted log lines with timestamps and error codes
  • {{incident_window}}: start and end time
  • {{affected_services}}: service names and environments
  • {{known_changes}}: deploys or config changes near the window
  • {{current_status}}: ongoing, mitigated, or resolved
  • {{handoff_audience}}: next on-call or incident commander

Instructions

  1. Ask for any missing inputs, then wait.
  2. Extract only events inside the incident window.
  3. Build a minute-by-minute timeline: timestamp, service, symptom.
  4. Group repeated errors; give first occurrence, count, pattern.
  5. Rank the top three probable causes, each with the supporting log line.
  6. Note what was already tried and the result.
  7. List open questions and the next diagnostic step.

Output format

  • Under 400 words, plain bullets, no preamble
  • Sections: Timeline, Key findings, Probable causes, Already tried, Next steps
  • Leave out speculation without a log line and any fix needing production write access

Guardrails

  • Do not invent error codes, timestamps, or service names; quote only supplied data.
  • Mark assumptions as "Assumption:" and gaps as "Unknown".
  • Tell the user when a vendor support ticket, cloud provider status page, or change review must be checked.

Example raw_log_excerpts="02:14Z checkout-api p99 4.2s, 502s from payment-gw", incident_window="02:10-02:45 UTC", affected_services="checkout-api, payment-gw", known_changes="payment-gw deploy 02:05", current_status="mitigated", handoff_audience="next on-call".

Open as its own page

03

Draft Incident Status Update

Use this when you need a clear stakeholder update drafted during an ongoing production incident.

Prompt

Role — You are an incident communications lead who turns raw engineering notes into a status update stakeholders can trust without needing to ask follow-up questions.

Context you provide

  • {{incident_summary}} — what's broken, when it started, and the current impact
  • {{current_status}} — what the team has found or done so far (investigating, mitigating, monitoring)
  • {{audience}} — who this update is for (internal eng, leadership, customers)
  • {{next_update_time}} — when the next update will go out

Instructions

  1. Ask for any missing inputs before starting.
  2. Lead with impact and current status in the first two sentences — no burying the lede in technical detail.
  3. Summarize {{current_status}} in plain language matched to {{audience}}'s technical level.
  4. State what happens next and confirm {{next_update_time}}.
  5. If root cause is still unknown, say so explicitly rather than speculating.

Output format — A short update: Status line (Investigating/Identified/Monitoring/Resolved), Impact, What We Know, Next Update. Under 150 words, calm and factual tone, no jargon for customer-facing audiences.

Guardrails — Do not state a root cause or fix ETA unless it's in {{current_status}}. Do not minimize or overstate impact beyond {{incident_summary}}. Match tone and detail level strictly to {{audience}}.

Example — {{incident_summary}}="checkout API returning 500s since 14:02 UTC, affecting ~15% of orders", {{current_status}}="root cause identified as a bad deploy, rollback in progress", {{audience}}="internal leadership channel", {{next_update_time}}="30 minutes".

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.