Course overview
Lesson 3 of 8 · 3 promptsAI for Site Reliability Engineers
LESSON 03 OF 8

Monitoring And Alerting

3 prompts for Site Reliability Engineers

Prompts for Site Reliability Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Write A PromQL Query For MetricsUse this when you need a Prometheus query for a metric but do not remember the exact syntax or functions.
  2. 02Reduce Noisy Alert VolumeUse this when an alert fires far more often than it is actionable and you want to refine thresholds, grouping, or inhibition rules.
  3. 03Design an SLO DashboardUse this when you need to turn SLIs and SLOs into a clear dashboard layout for a service.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Write A PromQL Query For Metrics

Use this when you need a Prometheus query for a metric but do not remember the exact syntax or functions.

Prompt

Role You are a site reliability engineer writing PromQL for Prometheus dashboards and alert rules. Optimise for a query that runs, returns the intended signal, and is readable for the team.

Context you provide

  • {{metric_name}}: exact metric or recording rule
  • {{metric_labels}}: labels available, e.g. job, instance, service, status_code
  • {{query_goal}}: what the query must return, e.g. error ratio, saturation
  • {{time_window}}: range or lookback, e.g. 5m
  • {{alert_condition}}: threshold or comparison, if this feeds an alert
  • {{prometheus_version}}: optional, to match available functions

Instructions

  1. Ask for any missing inputs, then restate the query goal in one sentence.
  2. Select the metric and keep only the labels needed for that goal.
  3. Choose the function and range vector that match the goal, such as rate, increase, histogram_quantile, or avg_over_time.
  4. Add aggregation with by or without so the result has one series per intended dimension.
  5. If this feeds an alert, wrap the expression in the comparison and state the for duration separately.

Output format One fenced promql block with the final query. Then a bullet list explaining each part, the expected result shape, and any assumptions. Under 200 words. Skip Prometheus setup and general theory.

Guardrails

  • Do not invent metric names, label names, or function behaviour. If a metric is unknown, say so and ask.
  • Flag any assumption about metric type, counter, gauge, or histogram.
  • Tell the user to compare results against a known dashboard or historical data before using it in a paging alert.

Example metric_name=http_requests_total, metric_labels=job, service, status_code, query_goal=5xx error ratio per service, time_window=5m, alert_condition=ratio > 0.05

Open as its own page

02

Reduce Noisy Alert Volume

Use this when an alert fires far more often than it is actionable and you want to refine thresholds, grouping, or inhibition rules.

Prompt

Role You are a site reliability engineer who tunes production alerting so pages reflect real customer impact. Optimise for fewer actionable alerts without hiding genuine incidents.

Context you provide

  • {{alert_name}} - the alert that fires too often
  • {{alert_rule_or_query}} - current rule, query or expression
  • {{current_threshold_and_duration}} - threshold, evaluation window, "for" duration
  • {{fires_per_week}} - how often it fires and how many are actionable
  • {{monitoring_stack}} - tooling in use, as configured
  • {{service_and_sli}} - service, signal, what users experience
  • {{notification_channels}} - where it pages or posts
  • {{grouping_and_inhibition_rules}} - anything already in place

Instructions

  1. Ask for any missing inputs, then restate the alert's intent in one sentence.
  2. Classify each typical fire: real impact, symptom of a known issue, or non-actionable noise.
  3. For each noise source, propose one or two specific changes: a longer "for" duration, a different aggregation window, a ratio or percentile instead of an absolute value, or a dependency-aware condition.
  4. Propose grouping labels so related alerts arrive as one page, plus inhibition rules that suppress downstream alerts when the upstream cause fires.
  5. Recommend routing by severity: page, ticket, or chat only.
  6. Give a validation plan: silent or shadow mode, expected change in fire count, review period, rollback. List your assumptions and what must be confirmed before any rule is edited.

Output format Sections in this order: Alert Intent, Noise Diagnosis, Threshold Options (table: change, expected effect, trade-off), Grouping and Inhibition, Routing, Rollout and Validation, Open Questions. Under 500 words. Terse, operational tone. Omit vendor config syntax unless the stack is named.

Guardrails

  • Do not invent metric names, thresholds, SLIs or vendor syntax. Use the supplied inputs or ask for them.
  • Flag every assumption, and say when vendor or manufacturer documentation must be checked for exact rule syntax.
  • State that paging, grouping and inhibition changes on customer-facing services need service owner and on-call lead review before rollout.

Example Alert: CheckoutLatencyHigh; query: p95 latency above 800ms for 5m; fires 14 times a week, 3 actionable; stack: Prometheus and Alertmanager.

Open as its own page

03

Design an SLO Dashboard

Use this when you need to turn SLIs and SLOs into a clear dashboard layout for a service.

Prompt

Role You are a site reliability engineer who designs SLO dashboards that let an on-call engineer judge reliability health in under a minute.

Context you provide

  • {{service_name}} — service the dashboard covers
  • {{slis}} — each indicator, how it is measured, its data source
  • {{slo_targets}} — target and rolling window per indicator
  • {{error_budget_policy}} — what happens when burn is high
  • {{audience}} — on-call, product, or leadership
  • {{data_sources}} — metrics backend and query language
  • {{dashboard_tool}} — where the dashboard will live
  • {{review_cadence}} — how often it is reviewed, by whom

Instructions

  1. Ask for any missing inputs, then confirm the SLI list and targets before designing.
  2. Map each SLI to the panel that answers "are we meeting the target right now?"
  3. Group panels into rows: current compliance and error budget remaining first, burn rate over short and long windows next, latency and saturation detail below.
  4. For each panel give a title, metric or query sketch, visualization type, threshold lines, and time window.
  5. Add annotations for deploys, incidents, and maintenance windows.
  6. State the refresh interval and default time range, and note which panels link to runbooks.

Output format Markdown. Two-sentence purpose statement, then a table of rows and panels, then a short annotation list. Panel titles under six words. No filler.

Guardrails

  • Do not invent metric names, thresholds, or SLO targets; use only what the user supplies and mark gaps TBD.
  • Flag assumptions about data availability or window alignment.
  • Tell the user to confirm targets and error budget policy with the service owner before publishing.

Example {{service_name}}: checkout-api; {{slis}}: availability, p99 latency; {{slo_targets}}: 99.9% over 30 days; {{audience}}: on-call engineers.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.