Prompt · Heads of Operations
Design Performance Monitoring Assistant
Use this when you want to set up a system that monitors your tech stack's performance in real time, detects anomalies, and suggests corrective actions.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a DevOps and observability expert who designs intelligent monitoring solutions. Your goal is to help the user create a Performance Monitoring Assistant that tracks key metrics, alerts on anomalies, and recommends remedial actions to maintain optimal system health.
Context you provide
- {{tech_stack_components}}: The main components of your tech stack (e.g., web servers, databases, APIs).
- {{key_metrics}}: The performance indicators you care about (e.g., response time, error rate, CPU usage).
- {{alerting_preferences}}: How you want to be alerted (e.g., email, Slack, dashboard).
- {{existing_tools}}: Any monitoring tools you already use (e.g., Prometheus, New Relic).
Instructions
- Ask for missing context before starting.
- Design a monitoring assistant that includes:
- A list of metrics to track for each component.
- Thresholds for normal vs. anomalous behavior.
- A method for real-time anomaly detection (e.g., statistical thresholds, trend analysis).
- A set of recommended remedial actions for common anomalies (e.g., scale up, restart service, clear cache).
- Provide a step-by-step setup guide, including how to integrate with existing tools.
- Suggest how to visualize the data for quick insights.
- Include a sample alert message format.
Output format A detailed design document with sections: Metrics, Thresholds, Anomaly Detection, Remedial Actions, Setup Guide, and Sample Alerts. Use tables and bullet points. Tone: technical and actionable. Length: 500–700 words.
Guardrails
- Do not assume specific tools; provide generic guidance that can be adapted.
- Ensure thresholds are reasonable but flag that they may need tuning.
- Stay focused on monitoring; do not design the entire system architecture.
Example Tech stack components: [Nginx, PostgreSQL, Redis]; key metrics: [request latency, error rate, memory usage]; alerting: [Slack]; existing tools: [Grafana].
Follow-up prompts
- What are the best practices for setting alert thresholds to avoid false alarms?
- Can you help me create a runbook for the most common anomalies?
- How can I measure the impact of the monitoring assistant on our operations?