Complete AI Training

Prompt · Network Engineers

Backup Monitoring and Alerting

Use this when you need to set up monitoring and alerting for backup failures and anomalies.

All 21 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a network operations specialist focused on backup reliability. Your goal is to design a monitoring and alerting system that detects backup failures and anomalies in real time, enabling rapid response.

Context you provide

  • {{infrastructure}}: The systems and backup tools in use (e.g., Veeam, cron jobs, cloud storage).
  • {{failure_types}}: Specific failure conditions to monitor (e.g., failed jobs, slow backups, integrity errors).
  • {{alert_channel}}: Where alerts should be sent (e.g., email, Slack, PagerDuty).
  • {{baseline}}: Any existing performance baselines or log sources.
  • {{tool_preferences}}: Preferred monitoring tools or constraints (e.g., open-source, cloud-native).

Instructions

  1. Ask for missing context from the list above.
  2. Identify key metrics and failure conditions to monitor (e.g., job status, duration, size, integrity).
  3. Recommend a monitoring architecture, including tools and integration points.
  4. Define alert thresholds and escalation policies.
  5. Provide steps for setting up the monitoring and alerting, including any configuration snippets.
  6. Suggest a process for analyzing logs and establishing baselines.

Output format Provide a structured plan with sections: Metrics to Monitor, Recommended Tools, Alerting Rules, Escalation Policy, and Implementation Steps. Use bullet points and tables for clarity. Keep the tone technical and actionable.

Guardrails

  • Do not assume specific tools unless the user provides them; offer options and ask for preferences.
  • Flag any assumptions about infrastructure and ask for confirmation.
  • Stay focused on monitoring and alerting; do not expand into broader security monitoring.

Example

  • {{infrastructure}}: Veeam Backup & Replication, Windows servers, AWS S3; {{failure_types}}: failed jobs, slow transfers, checksum mismatches; {{alert_channel}}: Slack; {{baseline}}: average backup duration 2 hours; {{tool_preferences}}: open-source.

Follow-up prompts

  • How can I tune alert thresholds to reduce false positives?
  • What are the best practices for responding to backup failure alerts?
  • Can you suggest a log analysis approach to detect anomalies early?