gin

Gin / LLM router

An LLM router that routes only what passed

Model routing without guesswork. Gin grades cheaper models on your real calls, per use case, and routes only the ones that held quality.

Benchmark numbers from the latest RouteBench run

Short answer: an LLM router decides which model answers each request. Most routers guess per prompt with a general-purpose classifier. Gin decides per use case, from replays of your own calls graded against your current model’s answers, and sends anything that errors back to your model.

Three ways to route between models

How LLM routers decide, and how they fail
ApproachDecides fromTypical failure
Static rulesYou pick a model per feature in codePicks go stale as models and prices change
Per-prompt classifier (OpenRouter Auto, Not Diamond, Martian)A model trained on public benchmarks scores each promptCheap picks on prompts that only look easy
Graded per use case (Gin)Your own calls, replayed on cheaper models and judged against your answersNeeds real traffic before it routes anything

How Gin routes

  1. Wrap your client. fetch: gin({ useCase: 'review' }). Gin records each call after it finishes; watching adds no latency.
  2. Grade a ladder of cheaper models. For each use case Gin replays recent calls on 3 to 6 cheaper candidates and has a strict judge compare each answer with yours, in both orders.
  3. Route only what passed. A use case routes when a candidate reaches ≥90% same-or-better and is meaningfully cheaper. The cheapest of the equally good candidates wins.
  4. Fall back and keep grading. Routed calls go through your own OpenRouter key. Any error falls back to your model, and Gin keeps re-grading as your prompts drift.

Below: the grading in detail, and how this compares with other routers on RouteBench, our open benchmark.

Your daily email

Your AI bill, explained before coffee.

Every morning, to you and your team. Free forever, even if you never route a call.

  • One big number. Yesterday’s AI spend and the trend this week.
  • Spend by project · use case. Two weeks, stacked, so the feature that spiked stands out.
  • What you could save. Counted only from cheaper models that passed grading on your calls.
  • One line per use case. Passed and how much it saves, kept on your model, or still grading.
Oct 7
$38.12

on AI yesterday · ▲ 26% this week

api · Support chatapi · Ticket triageweb · Summariesjobs · Code review

You could save

$417/mo

$934 → $517 a month

✓ 64% of spend passed grading on 212 of your calls

api · Support chatClaude Opus 5.5 → Claude Sonnet 5.5 ✓ 9/10
−50%
api · Ticket triageClaude Sonnet 5.5 → GLM 5.3 Flash ✓ 10/10
−93%
web · SummariesClaude Opus 5.5 ✗ cheaper models failed 4/12
keep
jobs · Code reviewClaude Opus 5.5 → DeepSeek V4.1 Flash ✓ 10/10
−96%
web · Kids’ storiesClaude Opus 5.5 grading…
…

Graded on your calls

“Customer asks why invoice #4821 shows two charges for October…”✓ Same answer · Claude Opus 5.5 $0.0214 → Claude Sonnet 5.5 $0.0107
“Label this ticket: ‘App crashes when I upload a PDF over 20 MB’”✓ Same answer · Claude Sonnet 5.5 $0.0031 → GLM 5.3 Flash $0.0002
Route with Gin →

Acme · Dashboard · Unsubscribe

Sample Illustrative numbers, same layout as the real email. Yours is built from your own calls.

How grading works

Proven on your calls before anything moves.

Gin doesn’t route on a leaderboard or a hunch. Each use case earns a cheaper model by matching your own answers, or it stays where it is.

  1. 1 · Watch

    Record real calls

    One line wraps your AI client. Gin logs each call after it finishes, by project and use case. No latency added.

  2. 2 · Optimize

    Replay and judge

    Gin replays recent calls on cheaper models. A strict judge compares each answer with yours, in both orders, so position can’t tip the verdict.

  3. 3 · Route

    Route what passed

    A use case routes only when ≥90% of its replays come back the same or better. Any error falls back to your model.

api · Support chaton Claude Opus 5.5

  • Claude Sonnet 5.5✓ 9/10−50%
  • DeepSeek V4 Pro✗ 6/9—

→ Routes to Claude Sonnet 5.5: −50% on this use case.

web · Summarieson Claude Opus 5.5

  • Claude Sonnet 5.5✗ 8/12—
  • GLM 5.3 Flash✗ 5/12—

= Stays on your model. Nothing cheaper held up, so nothing changes.

Bars show the share of replays graded the same or better; the tick marks the 90% bar. Example ladders, illustrative.

Benchmarked

Measured against the routers you’d compare us to.

Gin beat Always-Opus quality (93.8% vs 89.5%) for 56% less. OpenRouter Auto (high) came closest: 89.5% for $5.45. OpenRouter Auto is cheaper ($0.42) but passed 74.7%.

−56%cost vs Always Opus
93.8%pass rate, highest of 9 tested
$0.38per 1k exact-answer tasks at 99.1% (Opus: $5.20)
Pass rate against cost per 1,000 tasks for 9 routers on RouteBench. Gin: 93.8% at $4.33. Always Opus 5.5: 89.5% at $9.87.70%75%80%85%90%95%$0.3$1$3$10$ per 1k tasks →Always Opus 5.5: 89.5% pass rate, $9.87 per 1k tasksAlways Sonnet 5.5: 86.4% pass rate, $3.63 per 1k tasksAlways Haiku 5.5: 84.0% pass rate, $0.28 per 1k tasksAlways DeepSeek V4.1 Flash: 84.6% pass rate, $0.51 per 1k tasksOpenRouter Auto: 74.7% pass rate, $0.42 per 1k tasksOpenRouter Auto (medium): 85.2% pass rate, $2.15 per 1k tasksOpenRouter Auto (high): 89.5% pass rate, $5.45 per 1k tasksTypeSafe Jev Router: 86.4% pass rate, $1.23 per 1k tasksGin: 93.8% pass rate, $4.33 per 1k tasksGinOpus 5.5OR AutoOR Auto (high)Jev RouterSonnet 5.5Haiku 5.5DeepSeek V4.1 FlashOR Auto (medium)
GinOther routerSingle modelUp and left is better
Pass rate and cost per 1,000 tasks
RouterPass rate$ / 1kvs Opus
Gin93.8% +4.3$4.33−56%
OpenRouter Auto (high)89.5% +0.0$5.45−45%
Always Opus 5.589.5%$9.87base
TypeSafe Jev Router86.4% −3.1$1.23−88%
OpenRouter Auto (medium)85.2% −4.3$2.15−78%
OpenRouter Auto74.7% −14.8$0.42−96%

RouteBench 1.3 · 162 held-out tasks across 9 app workloads · run October 8, 2026 · 16 experimental variants on the full page. Gin’s routing was tuned with this benchmark in view, so its numbers are partly in-sample. Full RouteBench results and method →

When you don’t need a router

If one feature makes most of your calls and a small model already handles it, pick that model in code and skip the router. Routing pays when you have several features with different difficulty, a frontier model as the default, and no time to re-test every new model by hand.

Comparing gateways rather than routers? See OpenRouter alternatives, including LiteLLM, Portkey and Helicone.

FAQ

What is an LLM router?

An LLM router sends each request to one of several models, usually to cut cost or latency while keeping quality. Routers differ in what they decide from: fixed rules, a per-prompt classifier, or grades on your own traffic.

How is Gin different from OpenRouter Auto?

OpenRouter Auto picks a model per prompt with a general classifier. Gin picks per use case from replays of your own calls graded against your current model, and routes nothing that hasn’t passed. Gin can route through your OpenRouter key.

Does routing add latency?

Watching adds none. Routing adds one Cloudflare edge hop, usually 10–30 ms, and cheaper models are often faster than the frontier model they replace.

What happens if a routed model fails?

Any error on a routed model falls back to the model you asked for. If Gin itself is unreachable, calls go straight to your provider.