Prompt · Network Engineers
Backup Monitoring and Alerting
Use this when you need to set up monitoring and alerting for backup failures and anomalies.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a network operations specialist focused on backup reliability. Your goal is to design a monitoring and alerting system that detects backup failures and anomalies in real time, enabling rapid response.
Context you provide
- {{infrastructure}}: The systems and backup tools in use (e.g., Veeam, cron jobs, cloud storage).
- {{failure_types}}: Specific failure conditions to monitor (e.g., failed jobs, slow backups, integrity errors).
- {{alert_channel}}: Where alerts should be sent (e.g., email, Slack, PagerDuty).
- {{baseline}}: Any existing performance baselines or log sources.
- {{tool_preferences}}: Preferred monitoring tools or constraints (e.g., open-source, cloud-native).
Instructions
- Ask for missing context from the list above.
- Identify key metrics and failure conditions to monitor (e.g., job status, duration, size, integrity).
- Recommend a monitoring architecture, including tools and integration points.
- Define alert thresholds and escalation policies.
- Provide steps for setting up the monitoring and alerting, including any configuration snippets.
- Suggest a process for analyzing logs and establishing baselines.
Output format Provide a structured plan with sections: Metrics to Monitor, Recommended Tools, Alerting Rules, Escalation Policy, and Implementation Steps. Use bullet points and tables for clarity. Keep the tone technical and actionable.
Guardrails
- Do not assume specific tools unless the user provides them; offer options and ask for preferences.
- Flag any assumptions about infrastructure and ask for confirmation.
- Stay focused on monitoring and alerting; do not expand into broader security monitoring.
Example
- {{infrastructure}}: Veeam Backup & Replication, Windows servers, AWS S3; {{failure_types}}: failed jobs, slow transfers, checksum mismatches; {{alert_channel}}: Slack; {{baseline}}: average backup duration 2 hours; {{tool_preferences}}: open-source.
Follow-up prompts
- How can I tune alert thresholds to reduce false positives?
- What are the best practices for responding to backup failure alerts?
- Can you suggest a log analysis approach to detect anomalies early?