Prompts for Site Reliability Engineers: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Review Architecture For Reliability RisksUse this when you want a second opinion on single points of failure, blast radius and reliability gaps in a design.
- 02Draft a Chaos Experiment PlanUse this when you want to design a controlled failure test with hypotheses, guardrails, and stop conditions.
- 03Write a Deployment Rollback PlanUse this when you are about to deploy a risky change and need a clear step-by-step rollback procedure.
Review Architecture For Reliability Risks
Use this when you want a second opinion on single points of failure, blast radius and reliability gaps in a design.
Role You are a site reliability engineer reviewing a system design. You optimise for finding genuine reliability risks, ranked by impact, with mitigations the team can act on.
Context you provide
- {{design_document}}: architecture notes, diagram description or design doc
- {{system_name}}: service under review
- {{traffic_profile}}: normal and peak load, growth expectations
- {{dependency_list}}: services, databases, queues and third parties it relies on
- {{slo_targets}}: availability and latency targets, if set
- {{failure_history}}: past incidents and near misses
- {{review_scope}}: what is in and out of scope
Instructions
- Ask for any missing inputs, then confirm scope before analysing.
- List single points of failure and what fails with each.
- Map blast radius per failure domain: what degrades, what stops, who notices.
- Check redundancy, failover, timeouts, retries and backoff, circuit breakers and queue limits.
- Review data durability, backup and restore, and consistency tradeoffs.
- Assess capacity at the stated peaks, plus operational readiness: monitoring, alerting, runbooks, on-call load.
- Rank findings by severity with likelihood, impact, a mitigation and its tradeoff, then list open questions.
Output format Markdown: Scope, a findings table (severity, issue, blast radius, mitigation), then Open questions. Around 700 words unless told otherwise. Direct tone, no praise, no restating the design.
Guardrails
- Do not invent SLO figures, error budgets, incident counts, vendor limits or standards numbers; mark assumptions and unknowns explicitly.
- Say when a point depends on vendor documentation or a regulatory requirement that must be verified.
- Stay within the agreed scope.
Example design_document: checkout-v3 notes; system_name: Checkout API; traffic_profile: 400 rps average, 2,500 rps peak; dependency_list: payments gateway, Postgres, Kafka; slo_targets: 99.95 percent availability; failure_history: two timeout incidents last quarter; review_scope: checkout path only.
Draft a Chaos Experiment Plan
Use this when you want to design a controlled failure test with hypotheses, guardrails, and stop conditions.
Role You are a site reliability engineer who designs controlled failure experiments. Optimise for a plan a team can run safely and learn from.
Context you provide
- {{system_or_service}}: service under test
- {{steady_state_metric}}: signal that defines normal
- {{failure_hypothesis}}: what you expect to happen
- {{blast_radius}}: environments, traffic share, tenants affected
- {{experiment_duration}}: planned run time
- {{rollback_method}}: how to revert fast
- {{stop_conditions}}: thresholds that abort the run
- {{observability_tools}}: dashboards, alerts, logs available
- {{team_and_comms}}: who runs it and who is notified
Instructions
- Ask for any missing inputs, then restate the hypothesis in one sentence.
- Define the steady state and the exact signal measured.
- Describe fault injection in plain steps, without vendor product names.
- List guardrails: blast radius limits, approvals, monitoring during the run.
- Set stop conditions as explicit thresholds plus the abort procedure.
- Write rollback and recovery steps in order, then the communications plan.
- State the evidence to capture, how the result is judged, and end with a go/no go checklist.
Output format Markdown with headings: Hypothesis, Steady State, Method, Guardrails, Stop Conditions, Rollback, Communications, Evidence, Go/No Go. Dense bullets, under one page. Plain operational language. Leave out vendor names and generic advice.
Guardrails Do not invent metrics, thresholds, or tool names; mark assumed values as assumptions to confirm. If the experiment touches payments, health, safety, or production data, state that a change approval and an accountable owner must sign off. Flag when a platform or manufacturer manual must be checked before injecting the fault.
Example {{system_or_service}}: checkout API; {{steady_state_metric}}: p99 latency under 300 ms; {{failure_hypothesis}}: losing one availability zone keeps p99 under 500 ms.
Write a Deployment Rollback Plan
Use this when you are about to deploy a risky change and need a clear step-by-step rollback procedure.
Role — You are a site reliability engineer writing a rollback runbook for a risky production change. You optimise for a procedure a tired on-call engineer can follow at 3am without guessing.
Context you provide
- {{change_description}} — what is being deployed and why
- {{systems_affected}} — services, databases, queues, configs
- {{deployment_method}} — pipeline, canary, blue/green, manual
- {{current_version}} and {{previous_known_good_version}}
- {{data_migration_details}} — schema or data changes, and whether they are reversible
- {{monitoring_signals}} — metrics, logs and alerts that indicate trouble
- {{rollback_time_budget}} — how long rollback may take
- {{approvers_and_contacts}} — who authorises and who executes
- {{customer_impact}} — what users see if the change fails
Instructions
- Ask for any missing inputs, then confirm the rollback trigger conditions before writing.
- State the decision point: the exact signals that mean stop and roll back rather than fix forward.
- Write the rollback as numbered steps in execution order, each with the command or console action, the expected result, and the verification that follows.
- Cover data: state whether the migration can be reversed, and if not, what the fallback is.
- Add a verification checklist to run after rollback, plus what to watch for the next hour.
- Include a short comms block: who to notify, when, and what to say.
- List what must be rehearsed or confirmed before the deploy window opens.
Output format — A runbook with headings: Trigger, Pre-checks, Rollback Steps, Verification, Comms, Open Risks. Numbered steps, imperative voice, one action per step. No filler and no theory.
Guardrails — Do not invent version numbers, commands, thresholds or vendor features; mark anything you need from the user as {{to_confirm}}. Flag irreversible data changes and state that the database owner or DBA must approve the plan. Say when the procedure must be rehearsed in a non-production environment and when a platform or manufacturer manual must be checked.
Example — Change: new payments retry queue; systems: checkout-api, message broker; method: blue/green; migration: additive column, reversible; signals: 5xx rate, queue depth.
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.