gin

Gin / Resources

Prompt caching costs, provider by provider

Caching is the cheapest win in an AI bill: the same tokens at a tenth of the price. It also silently drops to zero more often than anyone expects.

From provider and OpenRouter docs, checked October 2026 · prices updated October 8, 2026

Short answer: cached input tokens cost 0.1× the input price on Claude and DeepSeek, about 0.25× on Gemini and Grok, and 0.1–0.5× on OpenAI. Claude and Gemini need explicit cache_control breakpoints and charge a 1.25× write premium; OpenAI, DeepSeek and most open models cache automatically.

Who caches what

Prompt caching by model family (multipliers of the normal input price)
FamilyHowMin tokensTTLReadWrite
Anthropic ClaudeExplicit cache_control, up to 4 breakpoints1,024–4,096 by model5 min, or 1 h0.1×1.25× (5 min), 2× (1 h)
Google GeminiImplicit on 2.5+, or explicit1,024 (Flash), 4,096 (Pro)~3–5 min implicit0.25×input + storage (explicit)
OpenAIAutomatic prefix; explicit options on newer models1,024, 128-token steps5–10 min, longer explicit0.1–0.5×free (premium on newest)
DeepSeekAutomatic prefix, on disk—hours~0.1×1×
Z.AI GLMAutomatic, varies by provider——~0.2×free
xAI Grok, Moonshot KimiAutomatic——0.25×free

What that means in dollars

Per million tokens, today, on OpenRouter:

  • Claude Sonnet 5.5: $2.00 input, $0.10 cached.
  • GPT-6 Sol: $2.00 input, $0.20 cached.
  • Gemini 3.8 Flash: $0.75 input, $0.075 cached.
  • DeepSeek V4.1 Flash: $0.30 input, $0.0060 cached.

A feature that sends a 20,000-token system prompt and tool list 10,000 times a day sends 6 billion prefix tokens a month. At a 0.1× read price, caching that prefix cuts its input bill by up to 90%. Try your own numbers with the cache slider in the LLM cost calculator.

When the write premium pays

On Claude, writing a prefix costs 1.25× and every read within 5 minutes costs 0.1×. One read already pays for the write: 1.25 + 0.1 < 2. The 1-hour TTL writes at 2×, so it needs about three reads per prefix per hour to beat the 5-minute default. Don’t add breakpoints to prompts that never repeat; you’d pay the premium for nothing.

Why real hit rates disappoint

On our own production traffic (IonWarp’s code review lanes through OpenRouter), 82% of input tokens already came from cache. Nearly all of the misses had one cause: requests landing on a different provider for the same model. One provider cached 87% of tokens; others serving the same model cached 0–41%, and one listed a cache price but served zero hits.

The usual causes, in order of how often we see them:

  1. Provider hops. On OpenRouter, the sticky key defaults to a hash of the first system and first user message. If the user message varies, requests scatter. Send a stable session_id per conversation or feature.
  2. A volatile line at the top. A timestamp, user name or request ID before the static instructions breaks the prefix at byte 0. Put static content first, per-request content last.
  3. Under the minimum. Prompts shorter than 1,024 tokens (4,096 on some models) never cache.
  4. Parallel fan-out. Ten requests started at once all miss, because none has finished writing the cache. Send one first, then the rest.
  5. provider.order on OpenRouter. It turns off sticky routing, so cache hits become luck.

A listed cache price proves nothing. Send the same long prompt twice and check usage.prompt_tokens_details.cached_tokens on the second response before you count on the discount.

How Gin handles caching

Gin records cached tokens on every call and shows, per use case, how much of the input came from cache and roughly how much a stable session or a breakpoint would save. When routing is on, the edge adds a sticky session_id per use case, avoids providers that don’t actually cache, and adds Claude and Gemini breakpoints only for prefixes it has already seen repeat. Output is unchanged.

FAQ

What is prompt caching?

When the beginning of a prompt is byte-for-byte the same as a recent request, the provider reuses its earlier work on that prefix and bills those tokens at a discounted cache-read price, typically 10 to 50 percent of the normal input price.

How much does Anthropic prompt caching cost?

Claude bills cache reads at 0.1× the input price. Writing to the cache costs 1.25× the input price for the default 5-minute TTL, or 2× for a 1-hour TTL. You mark what to cache with cache_control breakpoints, up to four per request.

Does OpenAI prompt caching cost extra?

OpenAI caches automatically for prompts of 1,024 tokens or more, in 128-token steps, and bills cached tokens at a fraction of the input price. Older models charge nothing extra for writes; newer models with explicit cache options can charge a write premium.

Does prompt caching work through OpenRouter?

Yes. OpenRouter passes through each provider’s cache pricing and keeps you on the same provider after a cached request. Sending a stable session_id makes that stickiness reliable; setting provider.order turns it off.

Why is my cache hit rate low?

Usually because the prefix changes (a timestamp or user name near the top of the prompt), requests land on different providers, the prompt is under the minimum length, or parallel requests start before the first one has written the cache.