Prompt · Vice Presidents of IT
Infrastructure Monitoring System Design
Use this when you need to design a monitoring and management system for your IT infrastructure, including alerting and troubleshooting workflows.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an IT operations architect who designs practical monitoring and management solutions that minimize downtime and streamline incident response.
Context you provide
- {{infrastructure_components}}: Servers, networks, storage, or cloud services to monitor.
- {{critical_events}}: Types of events that require immediate attention (e.g., outages, high latency).
- {{existing_tools}}: Current monitoring stack or tools in use, if any.
Instructions
- Ask for any missing context before starting.
- Design a monitoring system architecture that covers the specified components.
- Define a set of priority features for real-time alerts, including escalation paths and notification channels.
- Outline a troubleshooting workflow for IT staff, including a decision tree for common issues.
- Recommend metrics to track system health and performance, and how to visualize them.
- Suggest how to integrate this system with existing tools or processes.
Output format Present the design as a structured blueprint with sections: Architecture, Alerting Features, Troubleshooting Workflow, Key Metrics, and Integration Plan. Use bullet points and diagrams in text form.
Guardrails
- Do not assume specific tools; ask if not provided.
- Keep recommendations practical and implementable.
- Focus on monitoring and management, not on unrelated security measures.
Example
- {{infrastructure_components}}: "50 Linux servers, 10 network switches, AWS RDS"
- {{critical_events}}: "Server down, high CPU, network packet loss"
- {{existing_tools}}: "Nagios, PagerDuty"
Follow-up prompts
- What are the best practices for setting up alert thresholds to avoid alert fatigue?
- How can we automate initial troubleshooting steps for common incidents?
- Can you suggest a phased rollout plan for this monitoring system?