AI agent for ai consultants
Model Selection Bake-Off Agent
Choose a model with evidence on quality, cost and speed for the real task
What it does
Teams choose a model because it feels smart in a demo, then discover it is slow or expensive on the real workload. This agent runs a fair contest. It takes the team's own test set and candidate models, runs each, grades the answers with a rubric and a second grader for spot checks, and records cost per task and time per task. It also tries prompt adjustments per model, since a prompt tuned for one model can unfairly handicap another. It then ranks the models by the team's rules, such as quality first, then cost under a limit, then speed. The owner approves the choice. Edge case: the cheapest model wins on average but fails a rare critical case, so the agent applies the must-pass rule.
How it works
Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.
Read the steps as a list
- Model choice needed
- Run the test set on each candidate model
- Grade answers with the rubric and spot-check with a second grader
- Record cost and time per task
- Do the two graders agree on at least 90% of cases?If not: clarify the rubric and regrade. Back to step 3.
- Tune the prompt for each model and rerun the weakest cases
- Apply must-pass and limit rules to disqualify models
- Rank the remaining models by the ranking rules
- Is the top model's lead larger than the run-to-run noise?If not: run the test set again with more samples and recompute. Back to step 2.
- Owner approves the choiceThe agent waits here for your OK.
- Bake-off report with the ranking and trade-offs
How it decides
It disqualifies any model that fails the must-pass cases or limits, then ranks the rest by the stated priority of quality, cost and speed.
- Disqualify any model that fails a must-pass case
- Disqualify any model over the cost or latency limit
- Require grader agreement of at least 90%
- Treat differences under 2 points as a tie and use cost to decide
Make it yours
Every agent is a starting point. You choose these settings for your own situation.
- Candidate models
- Must-pass cases
- Ranking priority
- Cost and latency limits
- Runs per case (default 3)
What keeps you in control
It always asks you first
- Owner approves the final model choice
- Finance approves a cost increase
Hard limits
- Never use the test set to tune the prompt for only one model
- Never send customer data to a model the team has not approved
It stops when
- Done: a model is chosen with the report saved
- Stop: no model meets the limits
Set it up
We guide you through the set-up, step by step
Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.
- One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
- The agent then walks you through connecting your own data, one source at a time
- A downloadable copy with the flow chart, the rules and the full guide