Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for database administrators

Maintenance Job Failure Agent

Maintenance jobs that complete on time, with failures fixed or escalated the same day

Maintenance Job Failure Agent: what goes in, what the agent does and what you get

What it does

Nightly maintenance jobs fail quietly, and statistics and backups go stale. This agent reads job histories each morning and finds jobs that failed or ran much longer than normal. It reads the error details and retries with the known safe fix for that kind of error, such as rerunning after a lock clears. It confirms that the next run succeeds. If the same job fails again, it reports the cause and stops retrying. The administrator approves any change to job settings. Edge case: a backup job succeeded but wrote a file of zero bytes, so the agent treats it as a failure.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN Morning job review 2 USES A TOOL Read job histories and durations 3 DOES Find failed, slow and empty-output jobs 4 USES A TOOL Read the error details for each 5 CHECKS THE RESULT Is the error one with a known safe fix? If not: open a ticket with the evidence and tell theadministrator. Back to step 3. 6 DOES Apply the documented fix and rerun the job 7 CHECKS THE RESULT Did the rerun succeed and produce valid output? If not: stop retrying and report the cause. Back to step5. 8 YOU APPROVE Administrator approves any change to job settings 9 DOES Confirm the next scheduled run succeeds 10 RESULT Job health report
Read the steps as a list
  1. Morning job review
  2. Read job histories and durations
  3. Find failed, slow and empty-output jobs
  4. Read the error details for each
  5. Is the error one with a known safe fix?If not: open a ticket with the evidence and tell the administrator. Back to step 3.
  6. Apply the documented fix and rerun the job
  7. Did the rerun succeed and produce valid output?If not: stop retrying and report the cause. Back to step 5.
  8. Administrator approves any change to job settingsThe agent waits here for your OK.
  9. Confirm the next scheduled run succeeds
  10. Job health report

How it decides

A job has failed when it ends in error, runs past its normal time by a set share or produces an output that is empty. A known error gets its documented fix once.

  • Treat zero-byte backups as failures
  • Retry a known fix once
  • Compare runtime with the median of the last 14 runs
  • Stop after one failed retry

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Slow-run threshold (default: 50% above median)
  • Known fix list
  • Jobs in scope
  • Report time

What keeps you in control

It always asks you first

  • Administrator approves any change to job settings

Hard limits

  • Never changes job settings without approval
  • Never deletes backups

It stops when

  • Done: all jobs are healthy and the next run succeeds
  • Stop: the same job fails after the documented fix

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensA nightly backup showed success but wrote a 0 byte file. The agent counted it as failed. The error log showed a full disk. The known fix was clearing the temp folder, which it did. The rerun produced a 38 GB file, and the check passed. The next night's run also succeeded. The administrator approved a disk alert threshold.

More agents for database administrators