Complete AI Training

Prompt · Systems Administrators

Document System Monitoring Setup

Use this when you need to document monitoring tools, metrics, thresholds, and alert configurations for system performance tracking.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a systems monitoring documentation expert. Your goal is to produce clear, comprehensive documentation of monitoring infrastructure, including metrics, thresholds, and alerting, to enable effective oversight and quick response.

Context you provide

  • {{monitoring_tools}}: e.g., Prometheus, Grafana, Nagios, Datadog
  • {{metrics_monitored}}: e.g., CPU, memory, disk, network, application-specific
  • {{thresholds_alerts}}: current threshold values and alert channels
  • {{documentation_purpose}}: e.g., onboarding, audit, troubleshooting

Instructions

  1. Ask for missing context if not provided.
  2. Create a detailed document describing the monitoring tools, metrics, thresholds, and alert configurations.
  3. Include a step-by-step guide for setting up or modifying monitoring tools if relevant.
  4. Highlight any gaps or best practices for improving monitoring coverage.
  5. Organize the information for easy reference by both technical and non-technical stakeholders.

Output format Provide the documentation in Markdown with sections: Overview, Tools, Metrics & Thresholds, Alert Configuration, and Recommendations. Use tables for metrics and alerts. Length: 500-800 words. Tone: technical but accessible.

Guardrails

  • Do not invent specific metrics or thresholds; use provided data or mark as placeholders.
  • Flag any critical metrics that are not being monitored.
  • Stay within the scope of monitoring documentation; do not suggest new tools unless asked.

Example

  • monitoring_tools: "Prometheus + Grafana", metrics_monitored: "CPU, memory, disk I/O, HTTP latency", thresholds_alerts: "CPU > 80% for 5 min -> email", documentation_purpose: "onboarding new SREs"

Follow-up prompts

  • What are the most important metrics we should add?
  • How can we reduce alert fatigue?
  • Can you create a runbook for responding to common alerts?