Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for site reliability engineers

Uptime and Incident Response Agent

Detect and confirm outages quickly, find the likely cause and verify recovery.

Uptime and Incident Response Agent: what goes in, what the agent does and what you get

What it does

A customer emails to say the site is down, and you find out because they told you. This agent watches uptime checks and logs all day. When a check fails, it confirms the outage from several locations to rule out a single network problem. Once confirmed, it collects what changed: recent deploys, error spikes in the logs, certificate dates and the hosting status page. It proposes the most likely cause and a fix, such as rolling back the last deploy. It drafts a status update for customers. After a fix is applied, it checks that the site responds from all locations and that errors have returned to normal. If it has not recovered, it goes back to find the next likely cause. The developer approves any rollback and any public message. Edge case: only one region fails, so the agent says so in the message.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedYes, continueNoNo 1 STARTS WHEN Uptime check fails 2 USES A TOOL Test the site from several locations 3 CHECKS THE RESULT Do at least two locations fail repeatedly? If not: retest in two minutes and log it as a possiblelocal fault. Back to step 2. 4 USES A TOOL Collect recent deploys, error logs and certificatedates 5 DOES Rank the likely causes and propose a fix 6 DOES Draft the internal alert and the public statusmessage 7 YOU APPROVE Developer approves the fix and the public message 8 USES A TOOL Test the site again from all locations 9 CHECKS THE RESULT Has the site recovered and are errors back tonormal? If not: rank the next likely cause and propose the nextfix. Back to step 5. 10 RESULT Incident timeline and recovery update
Read the steps as a list
  1. Uptime check fails
  2. Test the site from several locations
  3. Do at least two locations fail repeatedly?If not: retest in two minutes and log it as a possible local fault. Back to step 2.
  4. Collect recent deploys, error logs and certificate dates
  5. Rank the likely causes and propose a fix
  6. Draft the internal alert and the public status message
  7. Developer approves the fix and the public messageThe agent waits here for your OK.
  8. Test the site again from all locations
  9. Has the site recovered and are errors back to normal?If not: rank the next likely cause and propose the next fix. Back to step 5.
  10. Incident timeline and recovery update

How it decides

An outage is confirmed when at least two of three locations fail twice in a row. The most likely cause is the most recent change that matches the time of the first error.

  • Confirm an outage when two of three locations fail twice in a row
  • Propose a rollback first when a deploy happened in the last 60 minutes
  • Check the certificate and domain expiry before looking at code
  • Draft a customer message after 10 minutes of confirmed downtime

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Uptime check addresses and locations
  • Number of failures before an alert (default 2 of 3)
  • Who is alerted and how
  • Status page and message tone
  • Delay before a public message (default 10 minutes)

What keeps you in control

It always asks you first

  • Developer approves rollbacks and any public message

Hard limits

  • Never roll back or restart a service without approval
  • Never post a public message without approval

It stops when

  • Done: recovery is verified and the timeline is saved
  • Stop: no likely cause is found after three attempts, so wake the developer

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensAt 14:02 the homepage fails from two of three locations. The logs show database connection errors starting at 13:58, four minutes after a deploy. The agent proposes a rollback and drafts a status message. The developer approves both. At 14:11 the homepage works from two locations, but the third still fails because of a stale cache. The agent proposes a cache purge, and after approval all three pass.

More agents for site reliability engineers