Course overview
Lesson 2 of 8 · 3 promptsAI for Site Reliability Engineers
LESSON 02 OF 8

Runbooks And Automation

3 prompts for Site Reliability Engineers

Prompts for Site Reliability Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Turn Incident Notes Into RunbookUse this when you have rough notes from a fix or an incident and want them turned into a structured, repeatable runbook.
  2. 02Generate Automation Script SkeletonUse this when you know the automation steps but want a starting script for a language or tool.
  3. 03Explain Infrastructure As Code ChangesUse this when you need to understand or review a Terraform, CloudFormation, or Kubernetes diff before applying it.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Turn Incident Notes Into Runbook

Use this when you have rough notes from a fix or an incident and want them turned into a structured, repeatable runbook.

Prompt

Role: You are a site reliability engineer who converts rough incident notes into a clean, repeatable runbook that another on-call engineer can follow under pressure. Optimise for clarity and safe execution over completeness.

Context you provide

  • {{service_name}}: the system or service this runbook covers
  • {{incident_notes}}: raw notes, chat logs, or the fix steps you actually ran
  • {{trigger_symptoms}}: alerts or symptoms that start this runbook
  • {{environment_or_stack}}: platforms, versions, and tooling involved
  • {{access_and_tools}}: dashboards, CLIs, or consoles the responder needs
  • {{validation_steps}}: how you confirmed the fix worked
  • {{escalation_contacts}}: roles or teams to page if the steps fail
  • {{known_limits}}: what this runbook does not cover

Instructions

  1. Ask for any missing inputs, then confirm the service name and trigger condition before writing.
  2. Reorganise the notes into ordered steps. Where the notes skip a check or a decision point, mark it as unconfirmed rather than filling it in yourself.
  3. For each step, state the action, the command or console path if one was recorded, and the expected result.
  4. Separate diagnosis steps from remediation steps so a responder can stop before changing anything.
  5. Add decision branches for each failure mode mentioned in the notes.
  6. List verification steps and the rollback path.
  7. Close with escalation triggers and the known limits.

Output format Markdown runbook with these sections: title, purpose, triggers, prerequisites, diagnosis, remediation, verification, rollback, escalation. Numbered steps, short sentences, imperative voice. No filler, no marketing language. Flag anything uncertain inline. Keep it under two pages.

Guardrails

  • Do not invent commands, thresholds, contact names, or tool paths. Write TO CONFIRM where the notes are silent.
  • Mark every assumption clearly and keep assumptions separate from facts recorded in the notes.
  • Tell the user to check current vendor or platform documentation and internal change policy before running remediation steps in production.

Example service_name=payments-api, incident_notes="pods OOMKilled after deploy, raised memory limit to 1Gi, restarted, traffic recovered", trigger_symptoms="OOMKilled alerts on payments-api pods".

Open as its own page

02

Generate Automation Script Skeleton

Use this when you know the automation steps but want a starting script for a language or tool.

Prompt

Role You are a site reliability engineer who writes clear, safe automation script skeletons. Optimise for a runnable starting point that the user can extend, not a finished production script.

Context you provide

  • {{automation_goal}}: what the script should achieve in one sentence.
  • {{automation_steps}}: the ordered steps the script must perform.
  • {{language_or_tool}}: the programming language or tool (e.g., Python, Bash, Ansible).
  • {{target_environment}}: where the script will run (e.g., Linux server, Kubernetes cluster, cloud CLI).
  • {{error_handling_requirements}}: how failures should be handled (e.g., retries, alerts, exit codes).
  • {{logging_requirements}}: what to log and at what level.
  • {{dependencies_or_constraints}}: any libraries, permissions, or runtime limits.

Instructions

  1. Ask for any missing inputs, then confirm the automation steps and language or tool before writing.
  2. Produce a script skeleton with clear sections: imports, configuration, functions for each step, error handling, logging, and a main entry point.
  3. Add comments that explain what each section does and where the user must fill in details.
  4. Include placeholder variables for environment-specific values.
  5. Suggest a minimal test or dry-run command if applicable.
  6. Note any assumptions you made.

Output format A single code block with the skeleton, followed by a short bullet list of assumptions and next steps. Keep the skeleton under 100 lines. Use the requested language or tool. Tone: technical and direct. Leave out full implementations, vendor-specific secrets, and long explanations.

Guardrails

  • Do not invent command flags, API endpoints, or library names.
  • Flag any assumption about the environment or steps.
  • Tell the user to test in a non-production environment and check official documentation for the language or tool.

Example automation_goal: restart a failed service; automation_steps: 1. check service status, 2. stop service, 3. start service, 4. verify; language_or_tool: Bash; target_environment: Ubuntu 22.04 server; error_handling_requirements: exit on failure with code 1; logging_requirements: log to syslog; dependencies_or_constraints: systemctl available.

Open as its own page

03

Explain Infrastructure As Code Changes

Use this when you need to understand or review a Terraform, CloudFormation, or Kubernetes diff before applying it.

Prompt

Role You are a site reliability engineer who reviews infrastructure as code changes before they reach production. Optimise for the reviewer knowing exactly what changes, what could break, and how to reverse it.

Context you provide

  • {{iac_tool}} - Terraform, CloudFormation, Kubernetes or similar
  • {{diff_or_plan}} - the plan or diff to review
  • {{environment}} - dev, staging or production, and the account or cluster
  • {{change_context}} - ticket, reason for the change, author
  • {{team_rules}} - naming, tagging, approval or change window rules

Instructions

  1. Ask for any missing inputs above, then work only from what you are given.
  2. Summarise the change in two sentences a non-specialist can follow.
  3. List each resource change: operation (create, update, replace, destroy), the key attributes changing, and whether the resource is stateful.
  4. Flag the risky items first: replacements, destroys, data stores, networking, identity or permission changes, and any new public exposure.
  5. Explain the blast radius and the order the apply will follow.
  6. State what the diff cannot show, such as drift or resources managed outside this code.
  7. Give a short pre-apply checklist and a rollback approach, noting where rollback is not possible. End with the questions the author should answer before approval.

Output format Markdown. Sections: Summary, Resource changes (table: resource, operation, risk), Risk highlights, Blast radius, Pre-apply checklist, Rollback, Questions. Short lines, plain language, gloss any jargon. Leave out generic best-practice advice and anything not visible in the diff.

Guardrails

  • Do not invent resource attributes, values or provider behaviour that are not in the material given. Say so when something is unclear.
  • Mark every assumption with "Assumption:" and note what would confirm it.
  • Tell the user to check provider documentation and their own change process, and to test in a non-production environment before applying anything that replaces or destroys stateful resources.

Example iac_tool: Terraform on AWS; environment: production; change_context: ticket OPS-441, resize a database and add a security group rule.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.