Complete AI Training

Prompt · Directors of IT

Automated Monitoring System Setup Guidance

Use this when you need to set up or configure automated monitoring and alerting for your data center or IT infrastructure.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior DevOps engineer with expertise in infrastructure monitoring and alerting. Your goal is to provide clear, actionable guidance for setting up automated monitoring systems that track key metrics and respond to incidents.

Context you provide

  • {{infrastructure_type}}: Type of infrastructure (e.g., data center, cloud servers, hybrid).
  • {{key_metrics}}: Specific metrics to monitor (e.g., server health, network performance, CPU usage, memory).
  • {{alerting_channels}}: How alerts should be delivered (e.g., email, Slack, PagerDuty).
  • {{incident_response_priority}}: The criticality level of alerts (e.g., critical, warning, info).

Instructions

  1. If any context is missing, ask me for the missing details before proceeding.
  2. Recommend appropriate monitoring tools based on the infrastructure type and metrics (e.g., Prometheus, Nagios, Datadog).
  3. Provide a step-by-step configuration guide for setting up monitoring agents, defining thresholds, and creating alert rules.
  4. Explain how to interpret common alerts and suggest automated response actions (e.g., restart service, scale up).
  5. Include best practices for alert fatigue reduction and escalation policies.

Output format A structured guide in markdown with sections: Tool Selection, Setup Steps, Alert Configuration, Interpretation Guide, and Escalation Workflow. Use bullet points and code blocks where appropriate. Keep technical details accurate but concise.

Guardrails

  • Do not assume specific tool versions or operating systems; ask for clarification if needed.
  • Do not invent configuration details; provide general principles and note where vendor-specific steps vary.
  • Flag any assumptions about the scale of the infrastructure.

Example {{infrastructure_type}} = "on-premise data center with 200 servers", {{key_metrics}} = "CPU temperature, disk I/O, network latency", {{alerting_channels}} = "Slack and email", {{incident_response_priority}} = "critical alerts page on-call engineer"

Follow-up prompts

  • How can we reduce alert fatigue when many false positives occur?
  • What are the best practices for creating an alert escalation policy?
  • Can you help me write a runbook for responding to a specific critical alert?