Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for software architects

Failure Mode Review Agent

A ranked list of failure gaps with tested fixes before they hit production

Failure Mode Review Agent: what goes in, what the agent does and what you get

What it does

Most outages come from failures that someone could have listed beforehand. This agent lists the components in the system, such as databases, queues, third-party services and network links, from the architecture documents. For each it enumerates how it can fail: slow, down, wrong data, full or out of sync. It then checks whether handling exists, like retries, fallbacks or timeouts, and whether monitoring would raise an alert. It ranks the gaps by impact and likelihood. For the top gaps, it proposes a staging test, such as stopping a service or adding delay, and runs it only after approval. After each test it checks the result against what was expected and updates the ranking. The architect approves the final plan. Edge case: a retry setting turns a small slowdown into a flood.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
ApprovedYes, continueYes, continueApprovedNoNo 1 STARTS WHEN Review is scheduled or a major change lands 2 USES A TOOL Read architecture documents and configuration 3 DOES List components and how each can fail 4 USES A TOOL Check existing handling and monitoring for eachfailure 5 DOES Rank the gaps by impact, likelihood anddetectability 6 YOU APPROVE Architect approves the staging tests to run 7 USES A TOOL Run each approved failure test in staging 8 CHECKS THE RESULT Did the system behave as expected and did an alertfire? If not: record the finding, adjust the ranking and add afix proposal. Back to step 4. 9 CHECKS THE RESULT Does each fix remove the gap when retested instaging? If not: revise the fix and retest. Back to step 7. 10 YOU APPROVE Architect approves the remediation plan 11 RESULT Failure mode register with test results
Read the steps as a list
  1. Review is scheduled or a major change lands
  2. Read architecture documents and configuration
  3. List components and how each can fail
  4. Check existing handling and monitoring for each failure
  5. Rank the gaps by impact, likelihood and detectability
  6. Architect approves the staging tests to runThe agent waits here for your OK.
  7. Run each approved failure test in staging
  8. Did the system behave as expected and did an alert fire?If not: record the finding, adjust the ranking and add a fix proposal. Back to step 4.
  9. Does each fix remove the gap when retested in staging?If not: revise the fix and retest. Back to step 7.
  10. Architect approves the remediation planThe agent waits here for your OK.
  11. Failure mode register with test results

How it decides

A gap is ranked by customer impact, likelihood and whether anything would warn the team. Tests start with the highest-ranked gaps.

  • Rank any failure with no alert as high detectability risk
  • Test only in staging
  • Test the highest-impact gaps first
  • Treat retries without a limit as a gap

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Components in scope
  • Test types allowed
  • Review schedule (default every 6 months)
  • Impact scale used for ranking

What keeps you in control

It always asks you first

  • Failure tests in staging
  • Final remediation plan

Hard limits

  • Never runs tests in production
  • Never changes configuration directly

It stops when

  • Done: gaps tested, fixed or accepted by the architect
  • Stop: staging does not match production closely enough

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensThe agent listed 23 components and 61 failure modes, with 14 lacking alerts. Its top test slowed the payments provider by 3 seconds in staging. Checkout retried three times and the queue grew 5x with no alert. It proposed a retry limit and a queue-depth alert. After the fix, retest showed the queue held steady and an alert fired. The architect approved the plan.

More agents for software architects