gin

Gin / Reduce LLM costs

Reduce LLM costs, safely

Cut your LLM bill where nobody can taste the difference. Track cost per feature for free, then move only what a cheaper model passed on your own calls.

Prices from the OpenRouter API · updated October 8, 2026

Short answer: the biggest lever is the model each feature uses: moving an easy feature from a frontier model to a small one cuts that feature’s cost by 80–97%. Caching saves up to 90% of repeated input. Gateway fees are at most about 5.5%.

The catch is quality. Measure cost per feature first, then switch only the features a cheaper model handles as well, on your own prompts.

Step 1: track LLM cost per feature

Provider dashboards group spend by key and model. They can’t tell you that the nightly digest costs $40 a day and chat costs $3, because they never see which feature made the call. Without that split you end up cutting the wrong thing.

Tag every call with the feature that made it. With Gin that is one line per AI client, fetch: gin({ useCase: 'digest' }), and LLM cost tracking is free: a dashboard by project and use case, plus the daily email below.

The levers, ranked by what they usually save

Typical savings on the feature or tokens each lever touches
LeverTypical savingThe risk
Cheaper model for easy features80–97% on that feature (Claude Opus 5.5 $4.00 in → GLM 5.3 Flash $0.15 in, per 1M)Worse answers if you switch without checking
Prompt cachingUp to 90% of repeated inputLow; hit rates are often lower than expected
Lower reasoning effort20–70% on reasoning modelsHard prompts fail more
Cap output, ask for terse JSON10–40%Truncated answers
Batch APIs for offline workAbout 50%Results arrive in hours, not seconds
Trim context and retrieval10–50% of inputMissing context
Switching gateways to avoid feesAt most ~5.5%Migration work

Why “just use a cheaper model” backfires

One global switch saves money on the easy features and quietly breaks the hard ones. Classification, extraction and short follow-up turns usually hold on a small model. Long summaries, stories and open-ended chat often don’t, and a generic benchmark won’t tell you which group your prompts fall into.

So test per feature, on your own calls, against your current answers, and keep a fallback. That is what Gin automates: it replays recent calls on cheaper models, a strict judge compares each answer with yours, and only use cases where ≥90% came back the same or better become eligible to route. Here is what that looks like.

Your daily email

Your AI bill, explained before coffee.

Every morning, to you and your team. Free forever, even if you never route a call.

  • One big number. Yesterday’s AI spend and the trend this week.
  • Spend by project · use case. Two weeks, stacked, so the feature that spiked stands out.
  • What you could save. Counted only from cheaper models that passed grading on your calls.
  • One line per use case. Passed and how much it saves, kept on your model, or still grading.
Oct 7
$38.12

on AI yesterday · ▲ 26% this week

api · Support chatapi · Ticket triageweb · Summariesjobs · Code review

You could save

$417/mo

$934 → $517 a month

✓ 64% of spend passed grading on 212 of your calls

api · Support chatClaude Opus 5.5 → Claude Sonnet 5.5 ✓ 9/10
−50%
api · Ticket triageClaude Sonnet 5.5 → GLM 5.3 Flash ✓ 10/10
−93%
web · SummariesClaude Opus 5.5 ✗ cheaper models failed 4/12
keep
jobs · Code reviewClaude Opus 5.5 → DeepSeek V4.1 Flash ✓ 10/10
−96%
web · Kids’ storiesClaude Opus 5.5 grading…
…

Graded on your calls

“Customer asks why invoice #4821 shows two charges for October…”✓ Same answer · Claude Opus 5.5 $0.0214 → Claude Sonnet 5.5 $0.0107
“Label this ticket: ‘App crashes when I upload a PDF over 20 MB’”✓ Same answer · Claude Sonnet 5.5 $0.0031 → GLM 5.3 Flash $0.0002
Route with Gin →

Acme · Dashboard · Unsubscribe

Sample Illustrative numbers, same layout as the real email. Yours is built from your own calls.

How grading works

Proven on your calls before anything moves.

Gin doesn’t route on a leaderboard or a hunch. Each use case earns a cheaper model by matching your own answers, or it stays where it is.

  1. 1 · Watch

    Record real calls

    One line wraps your AI client. Gin logs each call after it finishes, by project and use case. No latency added.

  2. 2 · Optimize

    Replay and judge

    Gin replays recent calls on cheaper models. A strict judge compares each answer with yours, in both orders, so position can’t tip the verdict.

  3. 3 · Route

    Route what passed

    A use case routes only when ≥90% of its replays come back the same or better. Any error falls back to your model.

api · Support chaton Claude Opus 5.5

  • Claude Sonnet 5.5✓ 9/10−50%
  • DeepSeek V4 Pro✗ 6/9—

→ Routes to Claude Sonnet 5.5: −50% on this use case.

web · Summarieson Claude Opus 5.5

  • Claude Sonnet 5.5✗ 8/12—
  • GLM 5.3 Flash✗ 5/12—

= Stays on your model. Nothing cheaper held up, so nothing changes.

Bars show the share of replays graded the same or better; the tick marks the 90% bar. Example ladders, illustrative.

Benchmarked

Measured against the routers you’d compare us to.

Gin beat Always-Opus quality (93.8% vs 89.5%) for 56% less. OpenRouter Auto (high) came closest: 89.5% for $5.45. OpenRouter Auto is cheaper ($0.42) but passed 74.7%.

−56%cost vs Always Opus
93.8%pass rate, highest of 9 tested
$0.38per 1k exact-answer tasks at 99.1% (Opus: $5.20)
Pass rate against cost per 1,000 tasks for 9 routers on RouteBench. Gin: 93.8% at $4.33. Always Opus 5.5: 89.5% at $9.87.70%75%80%85%90%95%$0.3$1$3$10$ per 1k tasks →Always Opus 5.5: 89.5% pass rate, $9.87 per 1k tasksAlways Sonnet 5.5: 86.4% pass rate, $3.63 per 1k tasksAlways Haiku 5.5: 84.0% pass rate, $0.28 per 1k tasksAlways DeepSeek V4.1 Flash: 84.6% pass rate, $0.51 per 1k tasksOpenRouter Auto: 74.7% pass rate, $0.42 per 1k tasksOpenRouter Auto (medium): 85.2% pass rate, $2.15 per 1k tasksOpenRouter Auto (high): 89.5% pass rate, $5.45 per 1k tasksTypeSafe Jev Router: 86.4% pass rate, $1.23 per 1k tasksGin: 93.8% pass rate, $4.33 per 1k tasksGinOpus 5.5OR AutoOR Auto (high)Jev RouterSonnet 5.5Haiku 5.5DeepSeek V4.1 FlashOR Auto (medium)
GinOther routerSingle modelUp and left is better
Pass rate and cost per 1,000 tasks
RouterPass rate$ / 1kvs Opus
Gin93.8% +4.3$4.33−56%
OpenRouter Auto (high)89.5% +0.0$5.45−45%
Always Opus 5.589.5%$9.87base
TypeSafe Jev Router86.4% −3.1$1.23−88%
OpenRouter Auto (medium)85.2% −4.3$2.15−78%
OpenRouter Auto74.7% −14.8$0.42−96%

RouteBench 1.3 · 162 held-out tasks across 9 app workloads · run October 8, 2026 · 16 experimental variants on the full page. Gin’s routing was tuned with this benchmark in view, so its numbers are partly in-sample. Full RouteBench results and method →

Then cache, cap and batch

  1. Make prompts cacheable. Put static instructions and tools first and the per-request part last. Claude and Gemini need explicit breakpoints. Details: prompt caching costs by provider.
  2. Cap output. Set max_tokens, ask for compact JSON, and turn reasoning down where it doesn’t change the answer.
  3. Batch what can wait. Nightly reports and backfills rarely need a real-time answer.

Want the numbers first? The LLM cost calculator prices any workload on 448 models.

FAQ

What is the fastest way to reduce LLM costs?

Find the most expensive feature that a smaller model can handle, and move only that feature. Model choice moves a feature’s cost 10× or more; fees and infrastructure move it a few percent.

How do I reduce LLM costs without hurting quality?

Test cheaper models per feature on your own recent calls, compare their answers with your current model’s, and switch only features that match. Keep a fallback to the original model for errors.

Is LLM cost tracking free with Gin?

Yes. Watching spend by project and use case, and the daily email, are free forever. You pay only when Gin routes calls: 5% of routed spend after the first $50.