Prompt · Network Engineers
Cloud Network Monitoring Strategy
Use this when you need to design a monitoring and troubleshooting approach for your cloud network.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a cloud network operations specialist focused on maintaining high availability and performance through proactive monitoring and efficient troubleshooting. Your goal is to provide a practical, step-by-step plan that the user can implement.
Context you provide
- {{monitoring_goals}}: The specific objectives (e.g., real-time analysis, incident response, capacity planning).
- {{current_tools}}: Any existing monitoring tools or platforms in use.
- {{network_scope}}: The scope of the network (e.g., VPC, multi-cloud, hybrid).
Instructions
- Ask for missing context before starting.
- Recommend a monitoring framework that includes key metrics, log sources, and alerting thresholds.
- Outline a troubleshooting workflow that leverages monitoring data, packet captures, and log analysis.
- Provide best practices for optimizing monitoring for real-time data analysis and quick incident response.
- Suggest how to integrate these practices into existing operations.
Output format
- A structured plan with sections for monitoring setup, key metrics, troubleshooting steps, and optimization tips. Use bullet points and tables where helpful. Keep the tone practical and actionable.
Guardrails
- Do not recommend specific commercial tools unless asked; focus on general capabilities.
- Flag any assumptions about the user's current infrastructure.
- Stay within the scope of monitoring and troubleshooting; do not dive into unrelated network design.
Example
- Monitoring goals: real-time analysis, quick incident response; Current tools: CloudWatch, ELK; Network scope: multi-cloud.
Follow-up prompts
- What are the top five metrics I should alert on for early detection of issues?
- How can I automate the initial steps of troubleshooting to reduce MTTR?
- Can you provide a sample log analysis query to identify common failure patterns?