Gin / Reduce LLM costs
Reduce LLM costs, safely
Cut your LLM bill where nobody can taste the difference. Track cost per feature for free, then move only what a cheaper model passed on your own calls.
Short answer: the biggest lever is the model each feature uses: moving an easy feature from a frontier model to a small one cuts that feature’s cost by 80–97%. Caching saves up to 90% of repeated input. Gateway fees are at most about 5.5%.
The catch is quality. Measure cost per feature first, then switch only the features a cheaper model handles as well, on your own prompts.
Step 1: track LLM cost per feature
Provider dashboards group spend by key and model. They can’t tell you that the nightly digest costs $40 a day and chat costs $3, because they never see which feature made the call. Without that split you end up cutting the wrong thing.
Tag every call with the feature that made it. With Gin that is one line per AI client, fetch: gin({ useCase: 'digest' }), and LLM cost tracking is free: a dashboard by project and use case, plus the daily email below.
The levers, ranked by what they usually save
| Lever | Typical saving | The risk |
|---|---|---|
| Cheaper model for easy features | 80–97% on that feature (Claude Opus 5.5 $4.00 in → GLM 5.3 Flash $0.15 in, per 1M) | Worse answers if you switch without checking |
| Prompt caching | Up to 90% of repeated input | Low; hit rates are often lower than expected |
| Lower reasoning effort | 20–70% on reasoning models | Hard prompts fail more |
| Cap output, ask for terse JSON | 10–40% | Truncated answers |
| Batch APIs for offline work | About 50% | Results arrive in hours, not seconds |
| Trim context and retrieval | 10–50% of input | Missing context |
| Switching gateways to avoid fees | At most ~5.5% | Migration work |
Why “just use a cheaper model” backfires
One global switch saves money on the easy features and quietly breaks the hard ones. Classification, extraction and short follow-up turns usually hold on a small model. Long summaries, stories and open-ended chat often don’t, and a generic benchmark won’t tell you which group your prompts fall into.
So test per feature, on your own calls, against your current answers, and keep a fallback. That is what Gin automates: it replays recent calls on cheaper models, a strict judge compares each answer with yours, and only use cases where ≥90% came back the same or better become eligible to route. Here is what that looks like.
Your daily email
Your AI bill, explained before coffee.
Every morning, to you and your team. Free forever, even if you never route a call.
- One big number. Yesterday’s AI spend and the trend this week.
- Spend by project · use case. Two weeks, stacked, so the feature that spiked stands out.
- What you could save. Counted only from cheaper models that passed grading on your calls.
- One line per use case. Passed and how much it saves, kept on your model, or still grading.
gin.Oct 7on AI yesterday · ▲ 26% this week
api · Support chatapi · Ticket triageweb · Summariesjobs · Code review
You could save
$417/mo
$934 → $517 a month
✓ 64% of spend passed grading on 212 of your calls
Graded on your calls
Acme · Dashboard · Unsubscribe
How grading works
Proven on your calls before anything moves.
Gin doesn’t route on a leaderboard or a hunch. Each use case earns a cheaper model by matching your own answers, or it stays where it is.
- 1 · Watch
Record real calls
One line wraps your AI client. Gin logs each call after it finishes, by project and use case. No latency added.
- 2 · Optimize
Replay and judge
Gin replays recent calls on cheaper models. A strict judge compares each answer with yours, in both orders, so position can’t tip the verdict.
- 3 · Route
Route what passed
A use case routes only when ≥90% of its replays come back the same or better. Any error falls back to your model.
api · Support chaton Claude Opus 5.5
- Claude Sonnet 5.5✓ 9/10−50%
- DeepSeek V4 Pro✗ 6/9—
→ Routes to Claude Sonnet 5.5: −50% on this use case.
web · Summarieson Claude Opus 5.5
- Claude Sonnet 5.5✗ 8/12—
- GLM 5.3 Flash✗ 5/12—
= Stays on your model. Nothing cheaper held up, so nothing changes.
Bars show the share of replays graded the same or better; the tick marks the 90% bar. Example ladders, illustrative.
Benchmarked
Measured against the routers you’d compare us to.
Gin beat Always-Opus quality (93.8% vs 89.5%) for 56% less. OpenRouter Auto (high) came closest: 89.5% for $5.45. OpenRouter Auto is cheaper ($0.42) but passed 74.7%.
| Router | Pass rate | $ / 1k | vs Opus |
|---|---|---|---|
| Gin | 93.8% +4.3 | $4.33 | −56% |
| OpenRouter Auto (high) | 89.5% +0.0 | $5.45 | −45% |
| Always Opus 5.5 | 89.5% | $9.87 | base |
| TypeSafe Jev Router | 86.4% −3.1 | $1.23 | −88% |
| OpenRouter Auto (medium) | 85.2% −4.3 | $2.15 | −78% |
| OpenRouter Auto | 74.7% −14.8 | $0.42 | −96% |
RouteBench 1.3 · 162 held-out tasks across 9 app workloads · run October 8, 2026 · 16 experimental variants on the full page. Gin’s routing was tuned with this benchmark in view, so its numbers are partly in-sample. Full RouteBench results and method →
Then cache, cap and batch
- Make prompts cacheable. Put static instructions and tools first and the per-request part last. Claude and Gemini need explicit breakpoints. Details: prompt caching costs by provider.
- Cap output. Set
max_tokens, ask for compact JSON, and turn reasoning down where it doesn’t change the answer. - Batch what can wait. Nightly reports and backfills rarely need a real-time answer.
Want the numbers first? The LLM cost calculator prices any workload on 448 models.
FAQ
What is the fastest way to reduce LLM costs?
Find the most expensive feature that a smaller model can handle, and move only that feature. Model choice moves a feature’s cost 10× or more; fees and infrastructure move it a few percent.
How do I reduce LLM costs without hurting quality?
Test cheaper models per feature on your own recent calls, compare their answers with your current model’s, and switch only features that match. Keep a fallback to the original model for errors.
Is LLM cost tracking free with Gin?
Yes. Watching spend by project and use case, and the daily email, are free forever. You pay only when Gin routes calls: 5% of routed spend after the first $50.