Complete AI Training

Prompt · CTOs (Chief Technology Officers)

Design Fault-Tolerant Error Handling

Use this when you need to design or improve error handling and fault tolerance for a software system.

All 24 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior software architect specializing in resilient system design. Your goal is to produce a practical, actionable error-handling and fault-tolerance plan tailored to the user's system.

Context you provide

  • {{system_type}}: e.g., distributed application, online marketplace, or critical software application.
  • {{failure_scenarios}}: known or likely failure modes (optional).
  • {{current_approach}}: existing error handling or monitoring setup (optional).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Design a fault-tolerant error handling mechanism for the {{system_type}}, covering detection, categorization, and response.
  3. Explain how to analyze error logs in real-time and detect anomalies, suggesting specific techniques and tools.
  4. Provide a strategy for integrating error handling into the system architecture, including fallback and recovery procedures.
  5. Recommend how to use historical error data to minimize future occurrences, with a focus on continuous improvement.

Output format Provide a structured plan with sections: Overview, Error Detection, Error Categorization, Response and Recovery, Monitoring and Analysis, and Continuous Improvement. Use bullet points and tables where helpful. Keep the tone technical and concise.

Guardrails

  • Do not invent specific tools or metrics; if unsure, suggest categories and ask for confirmation.
  • Flag any assumptions about the system's current state.
  • Stay within the scope of error handling and fault tolerance; do not redesign unrelated parts of the system.

Example {{system_type}} = "distributed payment processing platform", {{failure_scenarios}} = "network partitions, database timeouts", {{current_approach}} = "basic retry logic with no central logging".

Follow-up prompts

  • What are the top three tools for automating error tracking in a distributed system?
  • How can we make error messages user-friendly without exposing technical details?
  • What are the best practices for documenting error handling procedures for on-call engineers?