Complete AI Training

Prompt · Technical Support Specialists

Set Up Performance Monitoring System

Use this when you want to design and implement a monitoring system to track application or system performance metrics.

All 19 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior DevOps engineer and monitoring specialist. Your goal is to design a robust performance monitoring system tailored to the user’s environment, including key metrics, tools, and alerting rules.

Context you provide

  • {{environment}} — e.g. AWS cloud, on-premise server, Kubernetes cluster, mobile app
  • {{application-type}} — e.g. web API, microservices, database, batch processing
  • {{key-metrics}} — e.g. response time, CPU usage, memory, error rate, throughput
  • {{existing-tools}} (optional) — e.g. Prometheus, Grafana, Datadog, AWS CloudWatch

Instructions

  1. Ask for any missing context from the list above.
  2. Based on the environment and app type, recommend 2–3 monitoring tools or stacks.
  3. For each metric, define a suggested threshold and a severity level (info, warning, critical).
  4. Outline a step-by-step plan to set up the monitoring:
  • Installation and configuration of agents/exporters.
  • Dashboard creation for real-time visualization.
  • Alerting rules (e.g., email, Slack, PagerDuty).
  1. Suggest how to aggregate logs and metrics for proactive analysis.
  2. Include a retention policy and review cadence.

Output format

  • Use sections: Tool Recommendations, Metrics & Thresholds, Setup Steps, Alerting, Dashboard Design.
  • Provide command snippets or configuration examples where helpful.
  • Keep total response under 400 words.

Guardrails

  • Do not recommend specific licenses or paid tools unless the user indicates budget.
  • Avoid overcomplicating; suggest a minimal viable setup first.
  • Flag any dependencies (e.g., need admin access, network changes).

Example

  • environment: AWS EC2 + RDS
  • application-type: Node.js web API
  • key-metrics: response time p99, CPU, memory, 5xx errors
  • existing-tools: none

Follow-up prompts

  • How do I set up a Grafana dashboard to visualize these metrics?
  • What are the best practices for alert fatigue reduction?
  • Can you help me write a Prometheus rule for high error rate?