Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for database administrators

Replication Lag Investigation Agent

A confirmed cause of replica lag and a tested fix, with the evidence recorded

Replication Lag Investigation Agent: what goes in, what the agent does and what you get

What it does

Replica lag appears and clears with no known cause. When lag passes your threshold, this agent reads long-running transactions, network metrics and replica load, then tests possible causes in a fixed order. It records the likely one with the evidence, for example a bulk update on the primary or a slow disk on the replica. It suggests changes to settings or schedules and rechecks lag after they are applied. The administrator approves any setting change. Edge case: lag spikes every night at 1 am, so the agent compares the time with the scheduled jobs before blaming the network.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedYes, continueNoNo 1 STARTS WHEN Lag passes the threshold 2 USES A TOOL Read replication metrics and recent longtransactions 3 DOES Test for long transactions and bulk jobs at the timeof the lag 4 DOES Test network delay and replica resources 5 CHECKS THE RESULT Does a cause match the timing of the lag? If not: widen the time window and test the next cause.Back to step 3. 6 DOES Record the likely cause and suggest a change 7 YOU APPROVE Administrator approves any setting or schedulechange 8 USES A TOOL Apply the change 9 CHECKS THE RESULT Did lag fall below the threshold afterward? If not: revert the change and test the next cause. Backto step 3. 10 RESULT Incident note with cause and result
Read the steps as a list
  1. Lag passes the threshold
  2. Read replication metrics and recent long transactions
  3. Test for long transactions and bulk jobs at the time of the lag
  4. Test network delay and replica resources
  5. Does a cause match the timing of the lag?If not: widen the time window and test the next cause. Back to step 3.
  6. Record the likely cause and suggest a change
  7. Administrator approves any setting or schedule changeThe agent waits here for your OK.
  8. Apply the change
  9. Did lag fall below the threshold afterward?If not: revert the change and test the next cause. Back to step 3.
  10. Incident note with cause and result

How it decides

Causes are tested in order: long transactions, bulk jobs, network, replica resources. The first one that matches the timing is likely.

  • Test causes in a fixed order
  • Compare lag time with scheduled jobs first
  • Match a cause only on timing and metrics
  • Revert a change that does not help

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Lag threshold (default: 60 seconds)
  • Persistence time
  • Cause test order
  • Job schedule source

What keeps you in control

It always asks you first

  • Administrator approves setting and schedule changes

Hard limits

  • Never changes settings without approval
  • Never restarts replication

It stops when

  • Done: lag stays below the threshold after the fix
  • Stop: replication metrics are missing

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensLag hit 9 minutes at 1 am for the third night. The agent found a nightly bulk update on the primary that ran 40 minutes. Network and replica load were normal. It suggested splitting the update into 5 batches. After approval, the next night's lag peaked at 40 seconds, below the 60 second threshold.

More agents for database administrators