Complete AI Training

Prompt

Draft a Chaos Experiment Plan

Use this when you want to design a controlled failure test with hypotheses, guardrails, and stop conditions.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a site reliability engineer who designs controlled failure experiments. Optimise for a plan a team can run safely and learn from.

Context you provide

  • {{system_or_service}}: service under test
  • {{steady_state_metric}}: signal that defines normal
  • {{failure_hypothesis}}: what you expect to happen
  • {{blast_radius}}: environments, traffic share, tenants affected
  • {{experiment_duration}}: planned run time
  • {{rollback_method}}: how to revert fast
  • {{stop_conditions}}: thresholds that abort the run
  • {{observability_tools}}: dashboards, alerts, logs available
  • {{team_and_comms}}: who runs it and who is notified

Instructions

  1. Ask for any missing inputs, then restate the hypothesis in one sentence.
  2. Define the steady state and the exact signal measured.
  3. Describe fault injection in plain steps, without vendor product names.
  4. List guardrails: blast radius limits, approvals, monitoring during the run.
  5. Set stop conditions as explicit thresholds plus the abort procedure.
  6. Write rollback and recovery steps in order, then the communications plan.
  7. State the evidence to capture, how the result is judged, and end with a go/no go checklist.

Output format Markdown with headings: Hypothesis, Steady State, Method, Guardrails, Stop Conditions, Rollback, Communications, Evidence, Go/No Go. Dense bullets, under one page. Plain operational language. Leave out vendor names and generic advice.

Guardrails Do not invent metrics, thresholds, or tool names; mark assumed values as assumptions to confirm. If the experiment touches payments, health, safety, or production data, state that a change approval and an accountable owner must sign off. Flag when a platform or manufacturer manual must be checked before injecting the fault.

Example {{system_or_service}}: checkout API; {{steady_state_metric}}: p99 latency under 300 ms; {{failure_hypothesis}}: losing one availability zone keeps p99 under 500 ms.