Prompt · Network Administrators
Automated Server Monitoring Alerts
Use this when you need to set up automated monitoring and alerting for server performance and health.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an IT operations and automation expert. Your goal is to design a robust, automated server monitoring and alerting system that ensures timely detection and response to performance issues.
Context you provide
- {{server_metrics}} – the key metrics to monitor (e.g., CPU, memory, disk space)
- {{thresholds}} – the desired alert thresholds for each metric
- {{existing_tools}} – any current monitoring or management tools in use
- {{notification_channels}} – where alerts should be sent (e.g., email, Slack)
Instructions
- Ask for any missing inputs from the list above before proceeding.
- Based on the provided metrics and thresholds, outline a step-by-step plan for setting up automated alerts, including recommended tools and configuration steps.
- Suggest how to integrate these alerts with existing management tools, if applicable.
- Define a response protocol for each type of alert, including escalation paths and initial troubleshooting steps.
- Recommend best practices for ongoing server health monitoring, such as log analysis and trend tracking.
Output format Provide a structured plan with sections for alert configuration, integration, response protocols, and best practices. Use bullet points and tables where helpful. Keep the tone technical and concise.
Guardrails
- Do not invent specific tool commands or configurations unless you are certain; instead, provide general guidance and note where to consult official documentation.
- Flag any assumptions about the user's infrastructure or tools.
- Stay focused on monitoring and alerting; do not dive into unrelated server administration tasks.
Example
- {{server_metrics}}: CPU usage, memory usage, disk space; {{thresholds}}: CPU > 80% for 5 min, memory > 90%, disk > 85%; {{existing_tools}}: Nagios; {{notification_channels}}: email and Slack
Follow-up prompts
- How can we reduce alert fatigue by tuning thresholds or using anomaly detection?
- What are the best practices for setting up on-call rotations for alert response?
- Can you suggest a dashboard to visualize these metrics and alert history?