RouteBench · loading…
Which AI router keeps frontier quality for less?
We ran the same app-shaped tasks through every router we could reach, scored each answer against a known answer or a frontier reference, and counted every cent.
Quality vs. cost
Each point is one router on the 162 held-out tasks. Up is better quality, left is cheaper. The line joins routers nobody beats on both.
Gin v1 → v2
The first run showed Gin below the cost/quality frontier. We changed how production Gin learns, then re-ran the whole benchmark. The simulated Gin calls the same functions the hourly job and the routing edge run.
What changed
- Strict judge for open-ended work. Stories, chat and summaries are now judged pairwise by GPT-5.5 in both orders, with 12 or more samples per candidate. Verifiable tasks keep the cheap Gemini 3.8 Flash judge. In v1, Gemini passed Haiku on 9 of 10 practice stories, and Haiku then won only 17% of held-out stories.
- No switch is an outcome. If no cheaper model holds on an open-ended use case, it stays on its own model and testing pauses. In this run, summaries stay on Opus.
- Wider ladder, ordered per task type. Gin tests 3 to 6 rungs per use case, depending on the test budget. Verifiable tasks add the cheapest models first. Open-ended tasks add the best fits first.
- Calibrated hard-prompt fallback. A "keep on the original" region needs at least 2 failures behind it. Failures where the original also failed don't count. The cosine threshold is chosen per policy on held-out replays: the biggest expected saving that still meets the pass bar, including on the calls the candidate actually serves.
- Choose by expected cost. Candidates within 5 points of the best pass rate count as equal quality. Among them, the lowest expected cost wins, with fallback calls priced at the original model.
What each change was worth
We re-ran v2 with one change undone at a time, on the same 162 held-out tasks.
| Variant | Pass rate | $ / 1k |
|---|---|---|
| Gin v2 | 90.1% | $5.23 |
| − strict judge (v1 judge) | 87.0% | $1.47 |
| − 6-rung ladder (3 rungs) | 87.7% | $6.43 |
| − calibrated fallback (v1 fallback) | 89.5% | $6.46 |
| − ladder (v1's 2 cheapest models) | 90.1% | $6.20 |
The strict judge buys 3.1 points for $3.76 per 1k tasks. Most of that cost is summaries and stories that now stay on Opus or Sonnet. Without it, Gin v2 would sit at 87.0% for $1.47. The other three changes add quality and cut cost at the same time.
Two things didn't work. An uncalibrated v2 fallback let weak candidates "pass" by keeping most calls on Opus: summaries fell to 56%. We now require the routed calls themselves to clear the bar. A ladder that filled extra rungs by task fit skipped the cheap models that hold on verified tasks. We reverted that ordering. v2 was built with this benchmark in view, so its numbers are partly in-sample. The weekly re-run on the same tasks is a check, not new evidence.
By category
Pass rate per task family on held-out tasks. Bold marks the best score in each row; ties share it. Verified families have exact answers; judged families compare against Opus 5.5 with positions swapped.
What each router actually picked
Routers are only as good as the models they pour. Share of held-out calls served by each model.
Method
Tasks
252 synthetic tasks across nine app workloads, 28 each: code review with one seeded bug (or none), security review, JSON extraction with a gold record, support-ticket classification, tool calling with expected calls and arguments, multi-step math with exact answers, meeting/incident summaries, kids' bedtime stories, and policy-grounded support chat. Every task is generated deterministically from a seed; none uses customer data.
Each family is split 10 / 18: Gin learns on the 10 calibration tasks, and every router is scored only on the 18 held-out ones (162 in total).
Scoring
Verified families pass only on the exact answer: right line, right vulnerability class and line, every JSON field, the right label, the right tool with the right arguments, the exact integer.
Judged families must first pass hard checks (required names and words, length limits), then a pairwise judge (GPT-5.5, low reasoning) compares the answer with an Opus 5.5 reference twice, swapping positions. A win or tie passes; a loss fails. Always-Opus is judged against a second Opus sample, so its score shows the noise floor.
Cost and latency
Cost is OpenRouter's own usage.cost per call, scaled to 1,000 tasks. Latency is wall-clock for a non-streamed call from one machine, p50 and p95. Identical (model, task) calls are cached and shared between strategies, so Gin routing a call to Haiku reuses always-Haiku's answer.
How Gin is simulated
Gin learns here the way production learns on a customer's traffic, using production's own code. The app's original model is Opus 5.5. The profiler reads 5 calibration calls per family. The ladder covers 3 to 6 cheaper models, depending on the daily test budget of a workspace holding the default $10 credit. Gin replays calibration calls on each rung. Verifiable families are judged once by Gemini 3.8 Flash. Open-ended families are judged pairwise in both orders by GPT-5.5, on 12 samples: each prompt once, plus a second sample of two prompts.
A use case routes when a candidate reaches ≥90% same-or-better (≥75% pairwise) and is ≥20% cheaper. The cheapest of the equally good candidates wins. Look-alikes of repeated failures stay on Opus, at a threshold calibrated per policy. The routing call is production's own decide(). Gin's one-time calibration cost was $—, not counted in the per-task cost.
Caveats
18 tasks per family is small: treat differences under ~10 points in a single family as noise. Gin's pairwise replay judge is the same model as the scoring judge, with a different prompt and no rubric. It judges calibration tasks only, never the held-out ones. In v1 most verified families are near the ceiling for every model, so the spread comes mostly from the judged families; v2 needs harder verified tasks. The judge is an OpenAI model, and some routers pour OpenAI models. Third-party routers were called with default settings plus the documented cost tiers.
Data and re-runs
The task set and the summary are public: tasks-v1.jsonl (all 252 tasks with answers and rubrics) and data.json (every number on this page). This run cost $— in OpenRouter usage, and a GitHub Action re-runs it every week with a $40 spend cap.
The harness lives in apps/bench in the Gin repo. With an OpenRouter key:
OPENROUTER_API_KEY=sk-or-… bun apps/bench/run.ts
# a cheap smoke run
bun apps/bench/run.ts --routers gin,or-auto --limit 3
# rebuild the task set, publish a run
bun apps/bench/dataset.ts
bun apps/bench/publish.ts
Not yet benchmarked
- Not Diamond, Martian, Requesty, Unify: each needs an account created by a person before an API key exists. We'll add them once we have keys we can use under their terms.
- OpenRouter Pareto Code: a coding-only router, not meant for stories or support chat.
- OpenRouter Auto Beta: an early-access track of Auto; we benchmark the stable one.
Want your router included? It needs a public API we can call with a key.