Course overview
Lesson 1 of 8 · 3 promptsAI for Site Reliability Engineers
LESSON 01 OF 8

Incident Communication Basics

3 prompts for Site Reliability Engineers

Prompts for Site Reliability Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Draft Incident Status UpdateUse this when you need a clear stakeholder update drafted during an ongoing production incident.
  2. 02Summarize Incident Alert for HandoffUse this when you are handing off an ongoing incident and need a concise summary of what fired and what has been tried.
  3. 03Explain an Outage to StakeholdersUse this when you need to tell non-technical stakeholders what an outage means, what caused it and what happens next.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Draft Incident Status Update

Use this when you need a clear stakeholder update drafted during an ongoing production incident.

Prompt

Role — You are an incident communications lead who turns raw engineering notes into a status update stakeholders can trust without needing to ask follow-up questions.

Context you provide

  • {{incident_summary}} — what's broken, when it started, and the current impact
  • {{current_status}} — what the team has found or done so far (investigating, mitigating, monitoring)
  • {{audience}} — who this update is for (internal eng, leadership, customers)
  • {{next_update_time}} — when the next update will go out

Instructions

  1. Ask for any missing inputs before starting.
  2. Lead with impact and current status in the first two sentences — no burying the lede in technical detail.
  3. Summarize {{current_status}} in plain language matched to {{audience}}'s technical level.
  4. State what happens next and confirm {{next_update_time}}.
  5. If root cause is still unknown, say so explicitly rather than speculating.

Output format — A short update: Status line (Investigating/Identified/Monitoring/Resolved), Impact, What We Know, Next Update. Under 150 words, calm and factual tone, no jargon for customer-facing audiences.

Guardrails — Do not state a root cause or fix ETA unless it's in {{current_status}}. Do not minimize or overstate impact beyond {{incident_summary}}. Match tone and detail level strictly to {{audience}}.

Example — {{incident_summary}}="checkout API returning 500s since 14:02 UTC, affecting ~15% of orders", {{current_status}}="root cause identified as a bad deploy, rollback in progress", {{audience}}="internal leadership channel", {{next_update_time}}="30 minutes".

Open as its own page

02

Summarize Incident Alert for Handoff

Use this when you are handing off an ongoing incident and need a concise summary of what fired and what has been tried.

Prompt

Role You are a site reliability engineer writing a shift handoff for an ongoing incident. You optimise for a next responder who can take over in under two minutes without reading the full alert thread.

Context you provide

  • {{incident_id}} — ticket or incident number
  • {{alert_name}} — alert that fired
  • {{fired_at}} — time and time zone
  • {{affected_service}} — service impacted
  • {{current_status}} — what monitoring shows now
  • {{customer_impact}} — user-facing effect, if known
  • {{actions_taken}} — what was tried, with results
  • {{open_questions}} — unknowns or suspected causes
  • {{next_responder}} — person or team taking over
  • {{handoff_time}} — when the shift changes

Instructions

  1. Ask for any missing inputs, then draft the summary. Do not guess.
  2. Open with one line: alert name, time fired, current status.
  3. List what was tried in order, each with its observed result.
  4. State current impact plainly and whether it is worsening, steady, or recovering.
  5. Name the next action and who owns it.
  6. Flag anything needing a specialist check, such as database, network, or vendor review, before the next person acts.

Output format Markdown, 120 to 200 words: a one-line header, then What Fired, What We Tried, Current State, Next Action. Bullet lists, plain language, active voice. Leave out log dumps, stack traces, credentials, and any cause stated as fact when it is still a guess.

Guardrails

  • Do not invent alert names, timestamps, error codes, or impact numbers. Use only supplied detail; write "unknown" where it is missing.
  • Do not include secrets, tokens, internal hostnames, or customer personal data unless the user provides them.
  • If customer data, security, or a regulatory clock is involved, tell the user to confirm with the incident commander before sending.

Example Incident INC-4821, alert "checkout-api p99 latency above 2s", fired 14:07 UTC, service checkout-api, status degraded, impact some card payments timing out, tried restarting two pods then scaling to six, next responder Priya on payments on-call, handoff 15:00 UTC.

Open as its own page

03

Explain an Outage to Stakeholders

Use this when you need to tell non-technical stakeholders what an outage means, what caused it and what happens next.

Prompt

Role: You are a site reliability engineer who turns incident detail into a plain-English update for non-technical stakeholders, optimising for clarity, accuracy and trust.

Context you provide

  • {{incident_summary}}: what happened, in your words
  • {{affected_services}}: systems or customer journeys impacted
  • {{start_time_and_duration}}: when it started and how long so far
  • {{current_status}}: ongoing, mitigated, monitoring or resolved
  • {{known_cause}}: confirmed cause, or "under investigation"
  • {{customer_impact}}: who is affected and what they notice
  • {{next_steps_and_eta}}: planned actions and committed timing
  • {{audience}}: execs, support, sales or customers
  • {{channel_and_length}}: email, Slack post, status page, word limit
  • {{approved_language}}: terms or claims to use or avoid

Instructions

  1. Ask for any missing inputs, then wait.
  2. Lead with who is affected and what they will notice.
  3. Explain the cause in plain language, with no jargon, acronyms or ticket references.
  4. State current status and next actions, naming an owner only if confirmed.
  5. Keep confirmed facts separate from open questions and label each.
  6. Avoid blame, speculation and unconfirmed root-cause theories.
  7. Close with when the next update will arrive.

Output format Short sections: headline, what is happening, who is affected, what we know, what we are doing, next update. Stay within the word limit given. Tone: calm, factual, plain. Leave out jargon, acronyms, internal codes, unverified causes and blame.

Guardrails

  • Do not invent a cause, timeline, fix date or metric not in the inputs; mark unknowns as under investigation.
  • If the incident touches customer data, contracts, safety or a regulator, say that legal or compliance must review before sending.
  • Keep customer-facing claims inside the approved language provided.

Example audience: sales team; affected_services: login and payments; status: mitigated, monitoring; known_cause: under investigation; impact: some customers could not complete checkout for 40 minutes.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.