Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for meteorologists

Model Run Case Selection and Comparison Agent

A balanced evaluation of a model change with scores by variable and region and an honest statement of confidence

Model Run Case Selection and Comparison Agent: what goes in, what the agent does and what you get

What it does

Model changes are often judged on a few memorable cases, which can hide a worse result elsewhere. This agent picks a balanced set of past events across seasons, regions and weather types, runs or collects both model versions for each, and scores them against observations by variable and region. It compares the overall result to the headline claim. It checks sample size and widens the case set if the difference is too uncertain to call. It flags where results disagree with the headline, for example better precipitation in the east but worse wind in the west. The lead approves conclusions. Edge case: most selected cases are winter storms, so the agent adds summer convection cases.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN New model version ready 2 USES A TOOL Read the event archive and select cases 3 DOES Balance cases across seasons, regions and regimes 4 USES A TOOL Collect or run both model versions for each case 5 USES A TOOL Score forecasts against observations 6 CHECKS THE RESULT Is the case set balanced and large enough for eachregion? If not: add cases for the missing regimes and rescore.Back to step 3. 7 DOES Compare scores by variable and region 8 CHECKS THE RESULT Are differences larger than random variation acrosscases? If not: add cases or mark the result as inconclusive.Back to step 3. 9 DOES Compare findings to the headline claims and flagdisagreements 10 YOU APPROVE Lead approves conclusions 11 RESULT Evaluation report
Read the steps as a list
  1. New model version ready
  2. Read the event archive and select cases
  3. Balance cases across seasons, regions and regimes
  4. Collect or run both model versions for each case
  5. Score forecasts against observations
  6. Is the case set balanced and large enough for each region?If not: add cases for the missing regimes and rescore. Back to step 3.
  7. Compare scores by variable and region
  8. Are differences larger than random variation across cases?If not: add cases or mark the result as inconclusive. Back to step 3.
  9. Compare findings to the headline claims and flag disagreements
  10. Lead approves conclusionsThe agent waits here for your OK.
  11. Evaluation report

How it decides

Cases must cover the main regimes. A difference counts only if it exceeds what random variation across cases would give.

  • Require at least 30 cases per major regime
  • Treat small score differences as inconclusive
  • Report variables where the new version is worse
  • Use the same observations for both versions

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Regimes and regions
  • Minimum cases per regime
  • Scores used
  • Significance rule

What keeps you in control

It always asks you first

  • Conclusions on the model change

Hard limits

  • Never publishes results as final
  • Never drops cases because they look bad

It stops when

  • Done: conclusion supported or marked inconclusive
  • Stop: model output for the cases is unavailable

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensA new version claimed better rainfall forecasts. The first 24 cases included 20 winter events, so the balance check failed. The agent added 14 summer storms. Rainfall improved 6 percent overall, but 10 meter wind errors rose 4 percent in the west, and the rainfall difference was within random variation for summer. The lead approved a conclusion of improved winter rain and unclear elsewhere.

More agents for meteorologists