Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for statisticians

A/B Test Validity Check Agent

A test result read only when the data is valid, with a clear recommendation

A/B Test Validity Check Agent: what goes in, what the agent does and what you get

What it does

Test results are often read before the data is valid. This agent checks the split ratio, sample size, run time and tracking. It looks for peeking, where a test was stopped on a good day, and for novelty effects at the start. It computes the result only when the checks pass. If they fail, it suggests extending the test or fixing the setup. It then checks the significance and the size of the effect against the plan. The owner approves the decision. Edge case: the split is 55/45 when 50/50 was planned, so the agent runs a mismatch test and holds the result.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN Test reaches its end date 2 USES A TOOL Read the plan and the assignment and event data 3 DOES Check the split ratio against the plan 4 DOES Check sample size, run time and tracking events 5 CHECKS THE RESULT Do all validity checks pass? If not: suggest extending the test or fixing the setup,and wait. Back to step 2. 6 DOES Look for peeking and novelty effects 7 CHECKS THE RESULT Is the effect stable across the first and secondhalves of the test? If not: extend the test or mark the result as unclear.Back to step 2. 8 DOES Compute the result and effect size 9 YOU APPROVE Owner approves the decision 10 RESULT Test result summary
Read the steps as a list
  1. Test reaches its end date
  2. Read the plan and the assignment and event data
  3. Check the split ratio against the plan
  4. Check sample size, run time and tracking events
  5. Do all validity checks pass?If not: suggest extending the test or fixing the setup, and wait. Back to step 2.
  6. Look for peeking and novelty effects
  7. Is the effect stable across the first and second halves of the test?If not: extend the test or mark the result as unclear. Back to step 2.
  8. Compute the result and effect size
  9. Owner approves the decisionThe agent waits here for your OK.
  10. Test result summary

How it decides

Results are computed only when the split, sample size, duration and tracking checks pass. Otherwise the test is extended or marked invalid.

  • Run a split mismatch test
  • Require the planned sample size and duration
  • Compare early and late periods for novelty
  • Report effect size with the interval

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Required sample size and duration
  • Significance level (default: 95%)
  • Split tolerance
  • Segments to review

What keeps you in control

It always asks you first

  • Owner approves the final decision

Hard limits

  • Never declares a winner on an invalid test
  • Never changes the test setup

It stops when

  • Done: the result is valid and the decision is approved
  • Stop: the test cannot be fixed and is marked invalid

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensA checkout test planned 40,000 visitors over 14 days but was read on day 9. The agent held it. By day 14, the split was 51/49, which passed. The second half showed a lower lift than the first, so the check was run again on days 8 to 14, giving a 2.1% lift. The owner approved a rollout.

More agents for statisticians