AI news ·
OpenRouter launches model router benchmarks page to compare quality, speed, and cost
OpenRouter's Model Router Benchmarks let teams compare router quality, speed, and cost against single models, with a 48x per-token price gap between DeepSeek v4 Flash and GPT-6 Astra.

OpenRouter published a new Model Router Benchmarks page on October 2, 2026, giving operations and finance teams a direct way to compare the quality, speed, and cost of AI model routers against using single models. The benchmarks arrive as routers multiply - each promising to cut costs by switching between cheaper and more expensive models mid-task - but often introducing latency and cache rebuilds that can erase those savings.
Brian Thomas announced the benchmarks alongside a breakdown of how routers function in practice. "Every router was run through a set of benchmarks from varied domains to score them on quality, speed, and cost," Thomas said. Individual models representing both best-in-class cost efficiency and intelligence sit alongside routers on the page as a baseline. The page will receive ongoing updates as new routers launch and existing ones improve.
Why model routers exist and where they stumble
Routers address three cost and performance gaps that single-model setups ignore. Cost differences are stark: DeepSeek v4 Flash versus GPT-6 Astra shows a 48x higher average price per token for Astra, and a 21x cost increase for a typical 10-to-49-turn session in Codex. A router can hand cheaper subtasks to a low-cost model and reserve the expensive one for harder steps. Task diversity means the best legal research model is rarely the best coding model - routers can switch between them. Session stage lets routers adjust reasoning effort for complex planning versus routine execution within the same agentic session.
But these techniques carry built-in friction. Switching models forces the input cache to rebuild, making requests more expensive. Routers often lack enough prompt context to judge task complexity accurately. Heuristic signals for session stage sometimes conflict with what the model actually needs next. And every processing layer added for routing decisions introduces latency. The benchmark page is designed to show which routers overcome these trade-offs.
Three types of routers tested
The term "router" covers several distinct mechanisms. Provider routing - what OpenRouter has done since launch - picks an inference provider based on price, speed, uptime, and data policy, then fails over to another if needed. Model routing is what the benchmarks focus on: sending a request to a router like the Auto Router, which then decides which model responds.
Within model routing, three approaches were tested. Routers that execute a blend of models, such as Unbiased's Pareto and Sakana's Fugu, use undisclosed model mixes with standardized per-token billing. Routers that select a model each turn, including OpenRouter's Auto Router and Jev Router, disclose which model was chosen and bill at that model's standard rate. Routers that swap between a pre-defined pair, like NVIDIA's Switchyard, combine a cheaper model with a stronger one as an agent progresses through a task.
Other styles - Pareto Code, the -latest model slugs that alias to a single model, and Fusion approaches that run requests across a council of models - were not benchmarked because they serve specialized purposes needing different evaluation methods.
How the Router Index works
Raw benchmark results span multiple dimensions, making direct comparisons difficult. The Router Index synthesizes quality, speed, and cost into a 0-to-10 score for each benchmark. The default weighting puts quality at 60%, time per task at 20%, and cost at 20%. A slider on the page lets users adjust those weights to match their own priorities.
All routers listed are available on OpenRouter. Teams can swap a router in place of a model in any app, harness, or through the API. The benchmarks represent general tasks rather than any specific company's workload, so the page is a starting point for evaluation rather than a final answer.
Why this matters for customer support, finance, insurance, management, and operations teams
The decision to route or not route directly affects per-request costs and response times - two metrics that show up on operational dashboards and departmental budgets. A router that saves 20% on model costs but adds 300 milliseconds of latency per turn might be acceptable for a batch reporting job and unacceptable for a live chat agent. The adjustable index weighting on the benchmark page lets teams model that trade-off against their actual service-level requirements. Before committing to any router, run it against a sample of your own prompt types and measure both the token bill and the end-to-end response time your users experience.