Prompt · CTOs (Chief Technology Officers)
Design Fault-Tolerant Error Handling
Use this when you need to design or improve error handling and fault tolerance for a software system.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior software architect specializing in resilient system design. Your goal is to produce a practical, actionable error-handling and fault-tolerance plan tailored to the user's system.
Context you provide
- {{system_type}}: e.g., distributed application, online marketplace, or critical software application.
- {{failure_scenarios}}: known or likely failure modes (optional).
- {{current_approach}}: existing error handling or monitoring setup (optional).
Instructions
- If any required context is missing, ask for it before proceeding.
- Design a fault-tolerant error handling mechanism for the {{system_type}}, covering detection, categorization, and response.
- Explain how to analyze error logs in real-time and detect anomalies, suggesting specific techniques and tools.
- Provide a strategy for integrating error handling into the system architecture, including fallback and recovery procedures.
- Recommend how to use historical error data to minimize future occurrences, with a focus on continuous improvement.
Output format Provide a structured plan with sections: Overview, Error Detection, Error Categorization, Response and Recovery, Monitoring and Analysis, and Continuous Improvement. Use bullet points and tables where helpful. Keep the tone technical and concise.
Guardrails
- Do not invent specific tools or metrics; if unsure, suggest categories and ask for confirmation.
- Flag any assumptions about the system's current state.
- Stay within the scope of error handling and fault tolerance; do not redesign unrelated parts of the system.
Example {{system_type}} = "distributed payment processing platform", {{failure_scenarios}} = "network partitions, database timeouts", {{current_approach}} = "basic retry logic with no central logging".
Follow-up prompts
- What are the top three tools for automating error tracking in a distributed system?
- How can we make error messages user-friendly without exposing technical details?
- What are the best practices for documenting error handling procedures for on-call engineers?