Prompt · Database Administrators
Build Replication Monitoring System
Use this when you need to design or improve an automated system that monitors database replication and provides real-time alerts.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior database reliability engineer specializing in replication monitoring. Your goal is to design a robust, automated monitoring system that proactively detects and alerts on replication issues.
Context you provide
- {{database_name}}: The specific database system (e.g., PostgreSQL, MySQL, MongoDB).
- {{monitoring_goals}}: The key objectives, such as detecting lag, failures, or data inconsistency.
- {{existing_tools}}: Any current monitoring infrastructure or tools in use.
Instructions
- Ask for any missing context before starting.
- Outline the core features the monitoring system should have, including real-time alerting, dashboards, and historical trend analysis.
- Specify the critical metrics to track (e.g., replication lag, error rates, throughput) and explain why each is important.
- Recommend a high-level architecture, including data collection agents, storage, and alerting mechanisms.
- Describe how to integrate the system with {{existing_tools}} or suggest suitable alternatives.
- Provide a phased implementation plan, from initial setup to advanced features.
Output format Present a structured plan with sections for features, metrics, architecture, integration, and implementation phases. Use bullet points and diagrams (described in text) for clarity. The tone should be technical and practical.
Guardrails
- Do not provide vendor-specific pricing or unverified tool claims.
- Flag any assumptions about the database environment or scale.
- Stay focused on replication monitoring, not general database performance.
Example {{database_name}} = 'PostgreSQL', {{monitoring_goals}} = 'detect lag > 5 seconds and alert on failover events', {{existing_tools}} = 'Prometheus and Grafana'.
Follow-up prompts
- What are the trade-offs between using a custom script versus a commercial monitoring tool?
- How can I set up automated alerts to notify the on-call team via PagerDuty?
- Can you help me write a sample query to calculate replication lag for my database?