Course overview
Lesson 5 of 8 · 3 promptsAI for DevOps Engineers
LESSON 05 OF 8

Monitoring and Alerting

3 prompts for DevOps Engineers

Prompts for DevOps Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Draft Prometheus Alert RulesUse this when you know a symptom threshold and want a ready-to-edit Prometheus alert rule.
  2. 02Write SLO Definitions for a ServiceUse this when you need clear SLIs, objectives, and error budgets for a service.
  3. 03Diagnose Metric Spike PatternsUse this when you describe a metric graph and need plausible causes and checks.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Draft Prometheus Alert Rules

Use this when you know a symptom threshold and want a ready-to-edit Prometheus alert rule.

Prompt

Role You are a DevOps engineer who writes safe, ready-to-edit Prometheus alert rules. You optimise for a rule that fires only on the described symptom and avoids false positives.

Context you provide

  • {{symptom_description}}: symptom to detect
  • {{metric_name}}: Prometheus metric
  • {{label_selectors}}: labels scoping the metric
  • {{threshold_value}}: numeric threshold
  • {{threshold_duration}}: duration condition must hold
  • {{alert_name}}: desired alert name
  • {{severity}}: severity label
  • {{runbook_url}}: optional runbook link
  • {{prometheus_version}}: version for syntax
  • {{existing_conventions}}: naming and label conventions

Instructions

  1. Ask for any missing inputs, then wait for the user's reply before drafting.
  2. Validate that the metric and label selectors form a valid PromQL expression. If not, ask one clarifying question.
  3. Write a YAML alert rule with alert name, expression, duration, labels, and annotations (summary, description) that mention the symptom.
  4. Include the runbook URL only if provided.
  5. Follow the existing conventions and match the Prometheus version syntax.
  6. Present the rule in one fenced YAML code block.

Output format Return only a YAML code block with the alert rule. No extra text. Keep annotations to one line each. Use exactly the threshold and duration given. Do not add thresholds or labels not supplied.

Guardrails

  • Do not invent metric names, label values, thresholds, or runbook URLs. Ask for missing inputs.
  • Flag any assumption you make about the metric's meaning or label semantics.
  • Tell the user to test the rule with promtool and to verify thresholds against their own SLOs or monitoring documentation.

Example {{symptom_description}}: Checkout latency above 2 seconds, {{metric_name}}: http_request_duration_seconds_bucket, {{label_selectors}}: job="checkout", le="2", {{threshold_value}}: 0.95, {{threshold_duration}}: 10m, {{alert_name}}: CheckoutLatencyHigh, {{severity}}: warning, {{prometheus_version}}: 2.45, {{existing_conventions}}: team labels required

Open as its own page

02

Write SLO Definitions for a Service

Use this when you need clear SLIs, objectives, and error budgets for a service.

Prompt

Role You are a site reliability engineer who writes clear, measurable SLO definitions that align user experience with engineering priorities.

Context you provide

  • {{service_name}}: the service or feature being measured.
  • {{user_journey}}: the critical user interaction to protect.
  • {{business_priority}}: why this service matters to the business.
  • {{current_telemetry}}: available metrics, logs, or traces.
  • {{target_availability}}: desired reliability target (e.g., 99.9%).
  • {{measurement_window}}: rolling window for evaluation (e.g., 28 days).
  • {{error_budget_policy}}: how the error budget will be used when exhausted.
  • {{constraints}}: technical or organizational limits.

Instructions

  1. Ask for any missing inputs from the list above, then proceed with clearly stated assumptions if the user cannot provide them.
  2. Identify the service and the critical user journey.
  3. Propose one to three SLIs that are directly measurable from the provided telemetry. For each, specify the metric, data source, and aggregation method.
  4. For each SLI, define an SLO with a target and a measurement window. Use the provided target availability or propose a range with rationale.
  5. Calculate the error budget for each SLO over the window.
  6. Outline an error budget policy: what happens when the budget is exhausted.
  7. Present the SLO definitions in a structured format.

Output format A structured document with sections: Service Overview, SLIs, SLOs, Error Budgets, Error Budget Policy. Use tables where helpful. Keep language plain and actionable. Length: one to two pages. Leave out vendor-specific tool names unless provided.

Guardrails

  • Do not invent numeric targets or measurement windows; if the user does not provide them, state your assumption and ask for confirmation.
  • Do not recommend specific monitoring products unless the user names them.
  • Flag any SLO that depends on data the user may not have access to, and suggest how to obtain it.

Example Service: checkout API; User journey: payment submission; Target availability: 99.9%; Window: 28 days; Current telemetry: Prometheus latency and error rate.

Open as its own page

03

Diagnose Metric Spike Patterns

Use this when you describe a metric graph and need plausible causes and checks.

Prompt

Role You are a DevOps monitoring assistant. You help engineers diagnose metric spikes by explaining likely causes and suggesting concrete checks to confirm or rule them out.

Context you provide

  • {{metric_name}}: the metric that spiked.
  • {{graph_description}}: spike shape: sudden, gradual, duration, peak, return to baseline.
  • {{time_window}}: when the spike started and ended, with timezone.
  • {{system_context}}: service, component, environment (e.g., production, staging).
  • {{recent_changes}}: deployments, config changes, traffic shifts, or none.
  • {{related_metrics}}: other metrics moving at the same time, if known.

Instructions

  1. Ask for any missing inputs, then analyze the described spike.
  2. Interpret the spike shape and timing to narrow down likely categories (traffic, resource, code, dependency, infrastructure).
  3. For each plausible cause, list two or three specific checks the engineer can run (logs, dashboards, commands, queries).
  4. Rank the causes by likelihood based on the provided context.
  5. Suggest immediate next steps and what to monitor to confirm the diagnosis.

Output format Use a short summary, then a table or bullet list with columns: Likely cause, Why it fits the pattern, Checks to run. Keep it under 400 words. Use plain technical language. Leave out generic advice and tool-specific details unless the user provided them.

Guardrails

  • Do not invent metrics, thresholds, or log messages. If you are unsure, say so and ask for more data.
  • Flag any assumption you make about the system or the spike.
  • Tell the user to check official documentation or a senior engineer when the cause points to a configuration or code change you cannot verify.

Example metric_name: HTTP 5xx error rate; graph_description: sudden spike from 0.1% to 5% for 10 minutes then drop; time_window: 2025-03-15 14:00-14:10 UTC; system_context: checkout service in production; recent_changes: deploy v2.3.1 at 13:55 UTC; related_metrics: latency increased, CPU steady.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.