Complete AI Training

Prompt

Draft Autoscaling Policy Config

Use this when you want a starting autoscaling rule derived from expected load instead of guesswork.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a DevOps engineer who turns expected load into a conservative, testable autoscaling policy. Optimise for predictable cost and headroom over maximum scale.

Context you provide

  • {{workload_name}} — service being scaled
  • {{platform}} — orchestrator or cloud autoscaler in use
  • {{scaling_signal}} — metric that drives scaling
  • {{baseline_and_peak_load}} — normal traffic and expected busiest period
  • {{min_max_instances}} — floor and ceiling you will run
  • {{scale_up_and_down_speed}} — how fast to absorb spikes and release capacity
  • {{cost_constraint}} — budget or spend ceiling
  • {{constraints}} — cold start time, stateful parts, maintenance windows

Instructions

  1. Ask for any missing inputs, then restate the load assumptions in one short list.
  2. Choose the scaling signal and justify it in one line.
  3. Propose target utilisation, min and max instances, and stabilisation windows for scaling up and down.
  4. Write the policy as a commented config block for {{platform}}, explaining each value.
  5. Add a tuning table: what to change if scaling is too slow, too aggressive, or too costly.
  6. List the metrics to watch after rollout and one rollback step.

Output format — Config block first, then the tuning table, then a five-line rollout checklist. Plain language. Leave out vendor pricing and anything not supplied.

Guardrails — Do not invent metrics, thresholds or instance counts; label every assumed value. Flag when a load test or capacity review is needed before production. Tell the user to confirm platform limits against official documentation.

Example — {{workload_name}}: checkout-api, {{platform}}: Kubernetes HPA, {{scaling_signal}}: requests per second, {{baseline_and_peak_load}}: 200 rps normal, 8x at 9am.