Est.

KV Cache and Prompt Caching Across OpenAI, Anthropic, and Google

Different providers require different strategies to actually save money on prompt caching.

Contributing Editor · · 11 min read
Cover illustration for “KV Cache and Prompt Caching Across OpenAI, Anthropic, and Google”
Prompt Optimization · October 2, 2026 · 11 min read · 2,568 words

Prompt caching saves real money on real workloads, but it behaves differently enough on each major provider that treating it as one uniform feature leads straight to missed savings or surprise bills. Most production LLM applications send the same tokens over and over: the same system prompt, the same tool definitions, the same few-shot examples, the same retrieved documents, on every single request. Unless caching is working, every one of those tokens gets recomputed and billed in full repeatedly, for content that never changed.

The underlying mechanism is consistent across OpenAI, Anthropic, and Google. During prefill, the model computes key-value tensors for every token in the prompt. Prompt caching stores those tensors for a given prefix and reuses them the next time a request starts with that same sequence of tokens, so the cached portion skips prefill entirely. Three rules decide whether that reuse happens, and they hold across all three providers. The match has to be exact: one changed byte before the cache boundary breaks the hit for everything downstream. The engines hash fixed-size blocks, so a prefix that ends partway through a block only gets credit for the complete blocks before it. And the cache itself lives somewhere specific, on the GPU worker that ran the original prefill or in a tier of memory behind it, and it disappears on a timer or gets evicted when something else needs the space.

Where the providers split is in who controls that boundary and what it costs to use. OpenAI caches automatically, with no action required from the developer. Anthropic requires an explicit cache breakpoint sent on every request.

How OpenAI's automatic caching works

OpenAI's approach is the easiest to turn on and the hardest to reason about once it's running. Prompt caching activates automatically on prompts past a minimum prefix length, with no setup step, by matching the longest exact prefix it can find from a prior request handled on the same worker.

The pricing has shifted as the model lineup has grown. The cache is held for an extended period on most accounts, and there's no charge for writing to it on models up to GPT-5.5: a miss costs the standard input rate, a hit costs a fraction of that. Starting with GPT-5.6, writes are billed at a premium above the standard input rate. On the newer GPT-5 family models, cached input tokens come at a steep discount off the standard rate, and still carry no write premium. Hannecke found a 30 to 80 percent cut in time-to-first-token on long prompts when the prefix hits, though tokens-per-second during generation doesn't change, so the actual streaming speed once the model starts answering stays the same.

The catch sits in that word "automatic." Because developers can't mark a cache boundary themselves, OpenAI just matches the longest prefix it can find. Change anything near the top of the prompt, a timestamp in the system message, a dynamic instruction injected at the start, and the entire cache invalidates for everything that follows. That makes prompt order the only real lever available on OpenAI. Static content needs to come first. Dynamic content needs to move to the end. And a minimum token threshold means a very short prompt won't qualify for caching no matter how stable it is.

There's also no way to see cache state ahead of time. The usage object returned with the response carries the signal, something like usage.input_tokens_details.cached_tokens, and that field is the only place to confirm whether a prefix actually hit. A team running a workload that should be caching but sees cached_tokens come back at zero has a silent failure on its hands, one that is visible nowhere except in the bill.

Anthropic's cache breakpoints: control and cost

Anthropic takes the opposite stance. Nothing caches unless the developer asks for it: a cache_control breakpoint has to be sent explicitly, on every request, and writing to that cache costs money.

As of 2026, Anthropic runs two TTL tiers. A 5-minute ephemeral cache bills writes at a modest premium above the base input rate, the smaller of the two write premiums. A 1-hour ephemeral cache bills writes at 2 times the base input rate. Reads in both tiers get a steep discount off the base rate. That write cost sets a real break-even point: at the 5-minute tier, it only takes a small number of cache reads to cover the cost of the write, but a workload with a stable prompt that falls below that hit rate ends up paying more for the write premium than it saves on reads.

The TTL choice should track request frequency. A workload firing many requests per minute keeps the 5-minute cache warm on its own, since there's never a long enough gap for it to expire, so the cheaper write tier makes sense there. A workload where requests land every few minutes, with gaps that would blow past a 5-minute window, needs the 1-hour tier instead, even at the higher write cost, because letting the cache expire and rewriting from scratch costs more than the premium.

Anthropic allows up to four breakpoints per request, so tool definitions, the system prompt, retrieved documents, and conversation history can each be treated as their own cacheable block, cached and invalidated independently. The latency payoff for prefix-heavy prompts is substantial: FutureAGI's guide reports a strong drop in time-to-first-token and up to 80 percent cost reduction on cache hits, and cache reads on most Claude models don't count against input-token-per-minute rate limits, which gives high-volume applications real throughput headroom. Anthropic's Batch API compounds with caching, too: the Batch API cuts all token costs in half on its own, and stacking cache reads against the stable part of the prompt on top of that reaches a much deeper discount than either feature alone.

All of this control comes with a corresponding way to lose the savings. A developer who forgets the cache_control header, or sends it wrong, pays full input price for a prefix that never changed. Watching cache_read_input_tokens and cache_creation_input_tokens in the usage object is the only way to confirm the cache is actually active rather than assumed active.

Google Gemini's two-mode caching

Google runs both models at once. Implicit caching works automatically with zero setup, the same way OpenAI's does, and has applied to Gemini 2.5 models since they launched in May 2025. Explicit caching works through named-cache objects, closer in spirit to Anthropic's 1-hour tier, where the developer creates a cache object, gives it a name, and references that name in later requests.

The explicit API surface looks different from Anthropic's breakpoints in practice. A developer calls something like client.caches.create(), passing in the content to store, and gets back a cache object with a name. Later requests then reference that stored content with cached_content=cache.name instead of resending the document. It's a persistent, addressable object rather than a flag set inline on each request.

That persistence changes the economics. Google's implicit mode delivers a solid discount on cached tokens with no write surcharge, matching OpenAI's basic shape, though the discount runs shallower than on OpenAI's newer models, so workload economics differ even when the activation path is the same. The explicit mode delivers a deeper discount but adds per-hour storage billing, charged per million tokens stored and prorated to the minute. Unlike Anthropic, where the cost hits once on write, Google's explicit storage charge accumulates continuously, whether or not any request ever reads from it. A named cache built around a large document corpus that sits idle still racks up storage charges the whole time it exists. That math only favors workloads with high, sustained request volume against the same cached content.

The explicit mode fits one pattern well: large, stable documents such as legal contracts, research papers, or big configuration files, shared across many users or sessions, where the underlying document rarely changes and the same cache object gets referenced by a large number of requests. FutureAGI's 2026 guide names this as the primary use case. A system prompt that changes with every deployment is the wrong candidate for this mode; a 50-page contract referenced by thousands of queries over the course of a day is the right one.

Pricing structures compared side by side

| Provider / mode | Write surcharge | Read discount | TTL | Best for | |---|---|---|---|---| | OpenAI (automatic) | None up to GPT-5.5; 1.25x on GPT-5.6+ | Steep discount off standard input rate | Extended retention, most accounts | Stable prompts, zero setup, passive savings | | Anthropic (5-min tier) | 1.25x base input | Steep discount off base | 5 minutes | High-QPS workloads, cache stays warm naturally | | Anthropic (1-hour tier) | 2x base input | Steep discount off base | 1 hour | Medium-frequency workloads with gaps between requests | | Google implicit | None | Solid discount, shallower than OpenAI's | Automatic | Teams already on Gemini wanting zero-setup savings | | Google explicit | Per-hour storage billing, accumulates over time | 90% discount | Named, persistent | Large, stable documents at high sustained request volume |

All three providers have landed on broadly similar discounts for cache reads, but they diverge sharply on write costs and how long the cache survives, and that divergence, not the read rate, decides which provider is actually cheapest for a given workload. OpenAI asks for nothing on the front end but gives up control over what gets cached. Anthropic charges a modest premium on write in exchange for precise control over the cache boundary, with a low break-even that most prefix-heavy workloads clear easily. Google's explicit caching applies a 90% read discount but bills continuously for storage whether requests arrive or not, making it correct for very large, stable documents shared at high volume rather than for system prompts that change per deployment.

One factor cuts across every provider and every mode. Changing the active model mid-session forces a full re-prefill of the entire context on the next turn. Morph's analysis found that in a 15-turn agent session, that re-prefill can cost as much as many turns of cached reads combined. Staying on the same model and provider once the cache is warm matters just as much as getting caching turned on in the first place. None of these discounts touch output tokens either. Output always gets recomputed and billed in full, so a workload dominated by long generative output, like document drafting, gets proportionally less benefit from caching than one dominated by input, like RAG retrieval or an agent loop carrying a large context window.

Prompt structure decisions that determine whether the cache hits

Diagram: Prompt Order: What to Put First to Protect Your Cache. Visualizes: Illustrate the single most actionable structural rule in the article: order prompt content from most-stable to least-stable so that any change invalidates only the tail.

Pricing structure sets the ceiling on savings; prompt structure decides whether anything underneath that ceiling actually gets captured. The rule holds across every provider: order content from most stable to least stable. Tool definitions first, then the system prompt, then reference documents, then conversation history, then the live user query, because changing any one block invalidates that block and everything placed after it.

The most common way teams lose the cache is by putting dynamic content too early. A timestamp dropped into the system message, a session ID injected at the top of the prompt, retrieval results placed ahead of the static instructions: any one of these resets the match to zero tokens, and the cache never hits regardless of how stable the rest of the prompt actually is.

ProjectDiscovery's production case shows what fixing this looks like in practice. Cache hit rate went from a near-useless single-digit percentage to 84 percent, achieved by moving the agent's evolving working memory out of the system prompt and into a user message appended at the end of the conversation. Overall LLM cost fell substantially, with billions of tokens served from cache afterward, and the fix was purely a matter of prompt ordering, with no change to the model or the provider behind it.

Multi-turn conversations have a natural advantage here. Each turn appends new messages at the end while the earlier turns stay untouched, so the growing prefix forms a stable, cacheable block on its own. That advantage disappears the moment something rewrites earlier messages on every turn, whether that's RAG content injected into the system message, citation blocks inserted into the user message, or memory updates that alter the system prompt after the fact. Open WebUI's documentation catalogs the specific injection patterns that break the cache in agentic setups: file context, where retrieval chunks get injected into the latest message on every turn; citations, which rewrite the system message and the last user message after every tool-calling round; and memory or system context, where stored user memories get injected into the system message when they change. Each one invalidates the cached prefix at a different point in the request's lifecycle, which makes them easy to miss individually even when the aggregate effect on cache hit rate is large.

Anthropic's breakpoint placement deserves specific attention here. The cache_control marker belongs at the boundary between the stable prefix and the first dynamic element, not at the end of the prompt, so the static portion hits the cache every time even when the tail changes completely from request to request. Google's explicit mode carries its own version of this constraint: whatever document or context block gets placed inside a named cache object has to stay genuinely static across every request that references it. Any edit to that content means creating a brand new cache object, and that new object starts fresh storage billing from the moment it's created.

TTL mismatch and regional cache locality as production failure modes

Caching doesn't fail quietly only at the prompt level. It fails at the infrastructure level too, in ways that are harder to spot because the prompt itself can be perfectly well-ordered and the cache can still miss.

TTL mismatch is the most direct version of this. Anthropic's newer models bill cached input tokens at a steep discount off the standard input rate with no write premium, but a workload with real gaps between requests will quietly rewrite the cache on every single request instead of reading from it. That's a silent cost increase, not a crash, which makes it easy to miss until someone checks the bill against the expected hit rate. Choosing the 1-hour tier for that same workload, even at its higher write premium, solves the mismatch directly, which is the reasoning behind matching TTL tier to request cadence covered above.

Locality works the same way as the TTL mismatch: the cache for a given prefix lives on whichever GPU worker ran the original prefill, or in a tier of memory attached to that worker, and it only serves hits to requests that land back on that same worker or within that same locality. A routing layer, a load balancer, or a multi-region deployment that distributes requests across workers without any awareness of where a given prefix was last cached will scatter what should be a single warm cache across several cold ones. Each region or each worker builds its own copy, incurs its own write cost, and serves its own, likely lower, hit rate, even though the prompt itself never changed. Cache-aware routing, keeping a given session or a given stable prefix pinned to the same worker or the same region, makes the savings calculated in the pricing sections appear in a real, distributed deployment instead of in a single-worker benchmark.

Sources

  1. Prompt Caching: How It Works, Provider Pricing, Cache-Aware Routing (2026)
  2. Prompt Caching Explained: What It Is, What It Isn’t, and When to Use It
  3. Prompt Caching (KV Cache) Optimization / Open WebUI
  4. Prompt Caching in 2026: How It Works, Pricing, Wins

More in Prompt Optimization