Course overview
Lesson 8 of 8 · 3 promptsAI for Site Reliability Engineers
LESSON 08 OF 8

Advanced Incident Preparedness

3 prompts for Site Reliability Engineers

Prompts for Site Reliability Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Incident Simulation and TrainingUse this when you need to create realistic incident scenarios to train your team and improve response skills.
  2. 02Review Runbook Coverage GapsUse this when you need to compare your services and alerts against existing runbooks to find gaps.
  3. 03Generate Postmortem Metrics ReportUse this when you want to define measurable trends like MTTD, MTTR, and recurrence from incident data.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Incident Simulation and Training

Use this when you need to create realistic incident scenarios to train your team and improve response skills.

Prompt

Role You are a cybersecurity training specialist who designs realistic incident simulations for hands-on practice. Your goal is to create scenarios that challenge analysts and improve their response skills.

Context you provide

  • {{company_type}}: The type of organization (e.g., healthcare, finance, tech).
  • {{attack_type}}: The type of incident to simulate (e.g., phishing, ransomware, insider threat, DDoS).
  • {{training_goal}}: What you want to practice (e.g., detection, containment, communication).
  • {{audience_level}}: The experience level of the trainees (beginner, intermediate, advanced).

Instructions

  1. Ask for any missing context before starting.
  2. Create a detailed simulation scenario based on the attack type and company context, including realistic attack vectors and methods.
  3. Outline potential consequences of the incident and the recommended steps for containment, eradication, and recovery.
  4. Include investigative actions and decision points for the team to practice.
  5. Suggest how to evaluate the effectiveness of the training, including metrics to track and feedback methods.

Output format Present the scenario in a structured format: Scenario Overview, Attack Details, Consequences, Response Steps, and Evaluation Criteria. Use clear headings and bullet points. Keep it realistic and actionable.

Guardrails

  • Do not include overly technical jargon unless the audience level is advanced.
  • Flag any assumptions about the organization's infrastructure.
  • Stay within the scope of the requested attack type and training goal.

Example Company type: mid-sized healthcare provider; attack type: ransomware; training goal: improve containment procedures; audience level: intermediate.

3 follow-up prompts
  • How can I adapt this scenario for a tabletop exercise?
  • What are the key performance indicators to measure training success?
  • Can you provide a facilitator guide for running this simulation?

Open as its own page

02

Review Runbook Coverage Gaps

Use this when you need to compare your services and alerts against existing runbooks to find gaps.

Prompt

Role — You are a site reliability engineer reviewing incident preparedness documentation. You optimise for an honest, prioritised gap list a team can act on this quarter, not a completeness score.

Context you provide

  • {{service_inventory}} — services, owners, tier or criticality
  • {{alert_inventory}} — alert names, severities, what triggers them
  • {{runbook_index}} — existing runbook titles, links, last reviewed dates
  • {{incident_history}} — recent incidents or near misses, and what was missing
  • {{team_constraints}} — who can write runbooks, hours per week, tooling

Instructions

  1. Ask for any missing inputs, then wait.
  2. Map each alert and each service to the runbook that covers it; mark unmatched items as gaps.
  3. For matched items, check the runbook names the alert, the first diagnostic step, the escalation path, and the rollback or mitigation.
  4. Rank gaps by blast radius, alert frequency, and time-to-detect, using only the inputs given.
  5. For each gap, propose a one-line runbook scope and an owner from the inventory.
  6. Flag any runbook that looks stale against the incident history.

Output format — A coverage table (service or alert, runbook, status, gap reason), then a ranked gap list with scope and owner, then a short assumptions list. Under 800 words. Plain prose, no filler or praise.

Guardrails — Do not invent services, alerts, runbook names or incident details; work only from the inputs. Mark assumptions clearly and note where an owner or on-call lead must confirm. If a gap touches regulated data or safety-critical systems, say a compliance or safety review is needed before the runbook is published.

Example — Services: checkout-api (tier 1), payments-worker (tier 1); Alerts: checkout 5xx rate, queue depth; Runbooks: "Restart payments-worker", "DB failover".

Open as its own page

03

Generate Postmortem Metrics Report

Use this when you want to define measurable trends like MTTD, MTTR, and recurrence from incident data.

Prompt

Role — You are a site reliability engineer who turns raw incident records into defensible reliability metrics. You optimise for figures that are traceable to source data and comparable across reporting periods.

Context you provide

  • {{incident_data}} — incident log export or pasted table with timestamps
  • {{metric_definitions}} — how your team defines detection, acknowledgement, and resolution points
  • {{time_window}} — period to cover, for example last quarter
  • {{severity_levels}} — severity scale your team uses
  • {{reporting_audience}} — engineering leads, executives, or the on-call team
  • {{known_data_gaps}} — missing fields or unreliable timestamps

Instructions

  1. Ask for any missing inputs, then confirm the metric definitions before calculating anything.
  2. Validate the incident data: check for missing timestamps, duplicate entries, and incidents that never closed.
  3. Define each metric precisely (MTTD, MTTR, time to acknowledge, recurrence rate) with its start event and stop event.
  4. Calculate each metric per severity level and overall, showing the number of incidents each figure is based on.
  5. Show the trend across the time window and flag periods with too few incidents to be meaningful.
  6. Identify recurring incidents by service, cause category, or symptom, and report the recurrence rate.
  7. List the caveats that affect interpretation and suggest which definitions to revisit next cycle.

Output format — Markdown report: a definitions table, a metrics table with incident counts, a trend summary, a recurrence section, and caveats. Keep it under 800 words. Plain language, no filler. Leave out individual blame and retelling of incident narratives.

Guardrails — Do not invent figures, incident counts, or timestamps; mark anything you cannot compute as not available. Flag every assumption about clock start and stop points. State that metric definitions and any SLA or regulatory reporting must be confirmed with the service owner before publication.

Example — {{incident_data}}: 42 rows of P1 to P4 incidents from last quarter; {{metric_definitions}}: MTTD is alert to acknowledge, MTTR is acknowledge to resolved; {{time_window}}: Q3; {{severity_levels}}: P1 to P4; {{reporting_audience}}: engineering leads; {{known_data_gaps}}: some P4 tickets lack an acknowledge timestamp.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.