Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for ai consultants

Model Selection Bake-Off Agent

Choose a model with evidence on quality, cost and speed for the real task

Model Selection Bake-Off Agent: what goes in, what the agent does and what you get

What it does

Teams choose a model because it feels smart in a demo, then discover it is slow or expensive on the real workload. This agent runs a fair contest. It takes the team's own test set and candidate models, runs each, grades the answers with a rubric and a second grader for spot checks, and records cost per task and time per task. It also tries prompt adjustments per model, since a prompt tuned for one model can unfairly handicap another. It then ranks the models by the team's rules, such as quality first, then cost under a limit, then speed. The owner approves the choice. Edge case: the cheapest model wins on average but fails a rare critical case, so the agent applies the must-pass rule.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN Model choice needed 2 USES A TOOL Run the test set on each candidate model 3 USES A TOOL Grade answers with the rubric and spot-check with asecond grader 4 USES A TOOL Record cost and time per task 5 CHECKS THE RESULT Do the two graders agree on at least 90% of cases? If not: clarify the rubric and regrade. Back to step 3. 6 DOES Tune the prompt for each model and rerun the weakestcases 7 DOES Apply must-pass and limit rules to disqualify models 8 DOES Rank the remaining models by the ranking rules 9 CHECKS THE RESULT Is the top model's lead larger than the run-to-runnoise? If not: run the test set again with more samples andrecompute. Back to step 2. 10 YOU APPROVE Owner approves the choice 11 RESULT Bake-off report with the ranking and trade-offs
Read the steps as a list
  1. Model choice needed
  2. Run the test set on each candidate model
  3. Grade answers with the rubric and spot-check with a second grader
  4. Record cost and time per task
  5. Do the two graders agree on at least 90% of cases?If not: clarify the rubric and regrade. Back to step 3.
  6. Tune the prompt for each model and rerun the weakest cases
  7. Apply must-pass and limit rules to disqualify models
  8. Rank the remaining models by the ranking rules
  9. Is the top model's lead larger than the run-to-run noise?If not: run the test set again with more samples and recompute. Back to step 2.
  10. Owner approves the choiceThe agent waits here for your OK.
  11. Bake-off report with the ranking and trade-offs

How it decides

It disqualifies any model that fails the must-pass cases or limits, then ranks the rest by the stated priority of quality, cost and speed.

  • Disqualify any model that fails a must-pass case
  • Disqualify any model over the cost or latency limit
  • Require grader agreement of at least 90%
  • Treat differences under 2 points as a tie and use cost to decide

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Candidate models
  • Must-pass cases
  • Ranking priority
  • Cost and latency limits
  • Runs per case (default 3)

What keeps you in control

It always asks you first

  • Owner approves the final model choice
  • Finance approves a cost increase

Hard limits

  • Never use the test set to tune the prompt for only one model
  • Never send customer data to a model the team has not approved

It stops when

  • Done: a model is chosen with the report saved
  • Stop: no model meets the limits

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensFour models were run on 300 support cases. Model B scored highest at 91.2 but cost 3 times Model C at 89.8. Graders agreed on 84% at first, so the check failed. The rubric was clarified and agreement reached 93%. Model C then failed a must-pass refund-policy case, which disqualified it. Model A scored 90.9 at 40% of B's cost. The lead was inside the noise, so the owner chose A.

More agents for ai consultants