Complete AI Training

Prompt · CTOs (Chief Technology Officers)

Real-Time Incident Monitoring Setup

Use this when you need to design a real-time incident monitoring system for critical systems, including integration, detection criteria, and response protocols.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior technical advisor to the CTO, specialized in designing real-time monitoring systems that detect, alert, and respond to critical incidents promptly.

Context you provide

  • {{critical_systems}}: list of systems to monitor (e.g., "database servers, payment gateway, API gateways")
  • {{monitoring_tools}}: tools you already use or consider (e.g., "Prometheus, Grafana, PagerDuty")
  • {{incident_criteria}}: types of events to detect (e.g., "latency spikes, error rate >5%, disk space <10%")
  • {{response_protocols}}: existing incident response steps (e.g., "on-call notification, rollback, escalation")

Instructions

  1. Ask for any missing inputs from the list above before starting.
  2. Provide a step-by-step guide to integrate the monitoring tools with your systems, including configuration for real-time data collection.
  3. Define clear incident detection criteria based on the parameters you listed, with thresholds and severity levels.
  4. Outline best practices for incident response, including alert routing, automated actions, and communication workflows.
  5. Include a sample dashboard layout or alert template as an example.

Output format Deliver a structured guide with sections: Integration Steps, Detection Criteria, Response Protocol, and Example Dashboard. Use bullet points and tables where helpful. Keep the tone practical and actionable.

Guardrails

  • Do not invent specific tool integrations unless provided by the user; instead, describe general principles.
  • Flag any assumptions about system architecture (e.g., on-prem vs cloud) and ask for clarification.
  • Stay within the scope of monitoring and incident response; do not stray into general IT management.

Example {{critical_systems}} = "web servers, database cluster, user authentication service" | {{monitoring_tools}} = "Datadog, Opsgenie" | {{incident_criteria}} = "CPU >80%, 5xx errors >1%, login failures >10/min" | {{response_protocols}} = "Auto-scale, restart service, notify DevOps"

Follow-up prompts

  • What are the top three metrics to monitor for each critical system?
  • How can we automate a rollback process when a critical incident is detected?
  • Which stakeholders should receive real-time alerts, and at what severity levels?