Gin / LLM router
An LLM router that routes only what passed
Model routing without guesswork. Gin grades cheaper models on your real calls, per use case, and routes only the ones that held quality.
Short answer: an LLM router decides which model answers each request. Most routers guess per prompt with a general-purpose classifier. Gin decides per use case, from replays of your own calls graded against your current model’s answers, and sends anything that errors back to your model.
Three ways to route between models
| Approach | Decides from | Typical failure |
|---|---|---|
| Static rules | You pick a model per feature in code | Picks go stale as models and prices change |
| Per-prompt classifier (OpenRouter Auto, Not Diamond, Martian) | A model trained on public benchmarks scores each prompt | Cheap picks on prompts that only look easy |
| Graded per use case (Gin) | Your own calls, replayed on cheaper models and judged against your answers | Needs real traffic before it routes anything |
How Gin routes
- Wrap your client.
fetch: gin({ useCase: 'review' }). Gin records each call after it finishes; watching adds no latency. - Grade a ladder of cheaper models. For each use case Gin replays recent calls on 3 to 6 cheaper candidates and has a strict judge compare each answer with yours, in both orders.
- Route only what passed. A use case routes when a candidate reaches ≥90% same-or-better and is meaningfully cheaper. The cheapest of the equally good candidates wins.
- Fall back and keep grading. Routed calls go through your own OpenRouter key. Any error falls back to your model, and Gin keeps re-grading as your prompts drift.
Below: the grading in detail, and how this compares with other routers on RouteBench, our open benchmark.
Your daily email
Your AI bill, explained before coffee.
Every morning, to you and your team. Free forever, even if you never route a call.
- One big number. Yesterday’s AI spend and the trend this week.
- Spend by project · use case. Two weeks, stacked, so the feature that spiked stands out.
- What you could save. Counted only from cheaper models that passed grading on your calls.
- One line per use case. Passed and how much it saves, kept on your model, or still grading.
gin.Oct 7on AI yesterday · ▲ 26% this week
api · Support chatapi · Ticket triageweb · Summariesjobs · Code review
You could save
$417/mo
$934 → $517 a month
✓ 64% of spend passed grading on 212 of your calls
Graded on your calls
Acme · Dashboard · Unsubscribe
How grading works
Proven on your calls before anything moves.
Gin doesn’t route on a leaderboard or a hunch. Each use case earns a cheaper model by matching your own answers, or it stays where it is.
- 1 · Watch
Record real calls
One line wraps your AI client. Gin logs each call after it finishes, by project and use case. No latency added.
- 2 · Optimize
Replay and judge
Gin replays recent calls on cheaper models. A strict judge compares each answer with yours, in both orders, so position can’t tip the verdict.
- 3 · Route
Route what passed
A use case routes only when ≥90% of its replays come back the same or better. Any error falls back to your model.
api · Support chaton Claude Opus 5.5
- Claude Sonnet 5.5✓ 9/10−50%
- DeepSeek V4 Pro✗ 6/9—
→ Routes to Claude Sonnet 5.5: −50% on this use case.
web · Summarieson Claude Opus 5.5
- Claude Sonnet 5.5✗ 8/12—
- GLM 5.3 Flash✗ 5/12—
= Stays on your model. Nothing cheaper held up, so nothing changes.
Bars show the share of replays graded the same or better; the tick marks the 90% bar. Example ladders, illustrative.
Benchmarked
Measured against the routers you’d compare us to.
Gin beat Always-Opus quality (93.8% vs 89.5%) for 56% less. OpenRouter Auto (high) came closest: 89.5% for $5.45. OpenRouter Auto is cheaper ($0.42) but passed 74.7%.
| Router | Pass rate | $ / 1k | vs Opus |
|---|---|---|---|
| Gin | 93.8% +4.3 | $4.33 | −56% |
| OpenRouter Auto (high) | 89.5% +0.0 | $5.45 | −45% |
| Always Opus 5.5 | 89.5% | $9.87 | base |
| TypeSafe Jev Router | 86.4% −3.1 | $1.23 | −88% |
| OpenRouter Auto (medium) | 85.2% −4.3 | $2.15 | −78% |
| OpenRouter Auto | 74.7% −14.8 | $0.42 | −96% |
RouteBench 1.3 · 162 held-out tasks across 9 app workloads · run October 8, 2026 · 16 experimental variants on the full page. Gin’s routing was tuned with this benchmark in view, so its numbers are partly in-sample. Full RouteBench results and method →
When you don’t need a router
If one feature makes most of your calls and a small model already handles it, pick that model in code and skip the router. Routing pays when you have several features with different difficulty, a frontier model as the default, and no time to re-test every new model by hand.
Comparing gateways rather than routers? See OpenRouter alternatives, including LiteLLM, Portkey and Helicone.
FAQ
What is an LLM router?
An LLM router sends each request to one of several models, usually to cut cost or latency while keeping quality. Routers differ in what they decide from: fixed rules, a per-prompt classifier, or grades on your own traffic.
How is Gin different from OpenRouter Auto?
OpenRouter Auto picks a model per prompt with a general classifier. Gin picks per use case from replays of your own calls graded against your current model, and routes nothing that hasn’t passed. Gin can route through your OpenRouter key.
Does routing add latency?
Watching adds none. Routing adds one Cloudflare edge hop, usually 10–30 ms, and cheaper models are often faster than the frontier model they replace.
What happens if a routed model fails?
Any error on a routed model falls back to the model you asked for. If Gin itself is unreachable, calls go straight to your provider.