Prompt Caching / Analysis
Prompt Caching Economics: Break-Even Math for OpenAI vs Anthropic
Anthropic's prompt caching charges an upfront write multiplier that penalizes low hit rates, whereas OpenAI's legacy prefix caching carries zero write surcharge. Here is the mathematical derivation of exact break-even hit rates and cost trade-offs across 1k, 10k, and 100k token prompts.

Abstract technical illustration of prompt cache memory blocks routing LLM requests between cached reads and cache writes
Two Opposing Economic Models: Surcharges vs Free Writes
Key takeaways
- Anthropic rewards high reuse with deep read discounts but penalizes unused writes.
- OpenAI's automatic caching is risk-free but offers a smaller discount per hit.
- Newer OpenAI tiers may blur this split, so check current pricing per model.
Prompt caching skips recomputing the key-value (KV) attention states for a static prefix, such as a system prompt, tool definitions or a long document. A cached prefix cuts time-to-first-token and is billed at a reduced input rate. Providers differ in how they charge for creating that cache, and this difference drives the break-even math in the next section.
Anthropic uses an explicit, write-surcharged model, documented in the Claude platform docs. Writing a prefix to the cache costs 1.25x the base input price for the default 5-minute TTL, or 2.0x for the 1-hour TTL. Reads of that prefix are billed at a steep discount, roughly 90% below base input. As a relative illustration, a prefix that costs 1.0 uncached costs 1.25 to write on the 5-minute TTL, then about 0.1 per read. If the entry expires with no reads, you have paid a 25% penalty (5-minute TTL) or a 100% penalty (1-hour TTL) for nothing.
OpenAI's automatic prefix caching, described in the OpenAI developer docs, has no write step to opt into. On models like GPT-4o, a miss is billed at the normal 1.0x rate, and cached prefix tokens get a flat 50% discount (0.5x). A failed cache therefore costs nothing extra, but the savings per hit are smaller than Anthropic's.
Newer OpenAI frontier tiers such as GPT-5.6 are reported to move toward Anthropic's structure, with a 1.25x write surcharge and a deeper read discount. Confirm the exact multipliers on the official pricing page before modelling them. For background on how input, output and cached rates fit together, see our token pricing explainer. To compare against Anthropic's read-heavy model, see Claude 3.5 Sonnet.

Two Opposing Economic Models: Surcharges vs Free Writes
- Anthropic: 1.25x write (5-minute TTL), 2.0x write (1-hour TTL), about 0.1x read.
- OpenAI GPT-4o: 1.0x on a miss (no write penalty), 0.5x on a hit.
- Risk profile: an Anthropic write that expires unread costs 25% to 100% extra, whereas an OpenAI miss costs no extra.
Mathematical Derivation: Exact Minimum Break-Even Hit Rates
Key takeaways
- General formula: h_min = (W - 1) / (W - R).
- Any model with no write surcharge (W = 1) breaks even at 0% hit rate.
- Anthropic's 1-hour TTL needs more than double the hit rate of the 5-minute TTL to avoid a net surcharge.
Let P be the base input price per token, W the cache-write multiplier, R the cache-read multiplier, and h the hit rate on the cacheable prefix. A miss pays W times P because the prefix is processed and written. A hit pays R times P. The blended cost per prefix token is therefore P * [(1 - h) * W + h * R]. Caching beats uncached input when that bracket is at most 1.0. Rearranging gives W - 1 = h * (W - R), so h_min = (W - 1) / (W - R). The multipliers below come from each vendor's pricing documentation, so confirm them in the Anthropic prompt caching docs before you budget.
Anthropic's two TTLs give very different thresholds. Check the 5-minute case by substitution: at h = 0.2174, 0.7826 * 1.25 + 0.2174 * 0.10 = 1.000.
OpenAI's automatic caching has no write surcharge, so W = 1.00 and the numerator is zero. Any hit rate above zero lowers cost, and the saving is (1 - R) * h. With R = 0.50, as for GPT-4o mini, a 60% hit rate cuts blended prefix cost to 70% of list price. The read multiplier differs across OpenAI models, so check the OpenAI pricing documentation and the model page, for example GPT-5.6 Sol, before plugging in R. You can run your own multipliers through the cost calculators.
Two caveats apply. First, h is measured on cacheable prefix tokens, not total request tokens, and output tokens are unaffected. Second, the 1-hour break-even assumes the cache is hit often enough to stay warm, so bursty traffic with long idle gaps will see lower effective hit rates.
- Anthropic 5-minute TTL (W = 1.25, R = 0.10): h_min = 0.25 / 1.15 = 21.74%. Below that, caching raises your input bill.
- Anthropic 1-hour TTL (W = 2.00, R = 0.10): h_min = 1.00 / 1.90 = 52.63%. Below that, the longer TTL costs more than no caching.
- OpenAI automatic caching (W = 1.00, R = 0.50 for GPT-4o and GPT-4o mini): h_min = 0%. Every hit is a pure saving.
Cost Economics Across Context Lengths: 1k, 10k, and 100k Tokens
The examples below are illustrative: 10 sequential calls that share an identical prefix, with input tokens only (output tokens are excluded). They use Claude 3.5 Sonnet at $3.00/MTok base, with a 5-minute cache write at 1.25x ($3.75) and a cache read at 0.1x ($0.30), and GPT-4o at $2.50/MTok base with cached input at $1.25 and no write surcharge. Confirm current rates on the Anthropic docs and the OpenAI docs before relying on them.
At 1,000 tokens, caching does not apply. The prefix is below the 1,024-token minimum, so both providers bill it at the base rate. Anthropic's minimum is higher on some models, including Haiku-class ones such as Claude 3.5 Haiku, so check the docs for your model.
At 10,000 tokens, 10 uncached calls cost $0.30 on Claude and $0.25 on GPT-4o. With Claude at a 90% hit rate (1 write, 9 reads), the cost is $0.0375 + 9 x $0.003 = $0.0645, a 78.5% saving. At a 10% hit rate, assuming every miss writes a new entry, it is 9 writes plus 1 read, or about $0.3405. That is roughly 13.5% worse than not caching. GPT-4o at 90% costs $0.025 + 9 x $0.0125 = $0.1375, a 45% saving. At 10% it costs $0.2375, a 5% saving, and it can never be worse than uncached.
At 100,000 tokens, the dollar amounts scale by 10x with the same percentages. Ten uncached Claude calls cost $3.00. At a 90% hit rate with 5-minute caching, input costs $0.375 + 9 x $0.03 = $0.645, saving $2.355. The penalty also scales. A 1-hour write at 2x costs $0.60 for a 100k prefix, so each orphaned write that is never reused costs $0.30 more than an uncached call. An orphaned 5-minute write costs $0.075 more. Models with the same base price, such as Claude Sonnet 4, follow the same math. Use the model comparison tool to plug in other rates.

Cost Economics Across Context Lengths: 1k, 10k, and 100k Tokens
- 1k tokens: below the minimum, so no caching and no savings or penalty.
- 10k tokens, 90% hits: Claude $0.0645 vs $0.30 uncached; GPT-4o $0.1375 vs $0.25 uncached.
- 10k tokens, 10% hits: Claude about $0.3405 (a loss); GPT-4o $0.2375 (a small gain).
- 100k tokens, 90% hits on Claude: $0.645 vs $3.00 uncached.
Anthropic TTL Decision Matrix: 5-Minute vs 1-Hour Cache
Key takeaways
- 5-minute TTL: 1.35x vs 2.0x uncached after one read. Choose it when gaps stay under the window.
- 1-hour TTL: 2.20x vs 3.0x uncached after two reads. Choose it when gaps run 10 to 45 minutes.
- Hits refresh the TTL for free, so the real risk is gaps longer than the TTL, not the total session length.
Anthropic prices cache writes by retention window, so TTL choice is a break-even calculation on the gap between requests. Using multipliers on the base input price (confirm current values in the Anthropic prompt caching documentation), a 5-minute write costs 1.25x and a read costs 0.10x. One write plus one read costs 1.35x, against 2.0x for two uncached requests, so the 5-minute cache pays back after exactly one subsequent read. The 1-hour write costs 2.00x. Two reads bring the total to 2.20x, against 3.0x for three uncached requests, so it needs two reads within the hour to pay back.
Every cache hit resets the timer to its full duration (5 minutes or 1 hour) with no new write charge. A chain of requests spaced inside the TTL therefore pays for one write and then only reads. The cost of a wrong choice comes from gaps that exceed the TTL. When the cache expires, the next request pays the write premium again. Illustrative example: six requests spaced 10 minutes apart with a shared prefix cost 7.5x prefix-equivalent on the 5-minute TTL (six rewrites at 1.25x), 6.0x uncached, and 2.5x on the 1-hour TTL (one 2.00x write plus five 0.10x reads).
Use the decision rules below to pick a window. If a pricing change shifts the multipliers, the price tracker is a place to check. For a provider with a different caching structure, the DeepSeek V3 page shows how its pricing compares.
- Use the 5-minute TTL for agent loops, customer support chats and interactive coding assistants where turns arrive less than about 3 minutes apart. This keeps the write premium at 1.25x and leaves margin for slow turns.
- Use the 1-hour TTL for intermittent, batch-burst or human-in-the-loop workloads with 10 to 45 minutes between requests. The 2.00x write is justified because it prevents repeated 1.25x rewrites that would cost more than no caching.
- If traffic is under two requests per hour on the same prefix, skip the 1-hour cache. It will not break even.
Threshold Gates and Granularity: Token Minimums and Alignment
Key takeaways
- Below the minimum, no caching happens on either platform.
- OpenAI rewards left-aligned, deterministic prompts; Anthropic rewards deliberate breakpoint placement.
- Pad short prompts only after running the break-even math with current model rates.
Caching only triggers once a prompt clears a minimum length, and the two platforms differ on how you control the boundary. Check the current figures in the official docs for Anthropic prompt caching and OpenAI prompt caching, because minimums are set per model and change as new models ship.
OpenAI caches automatically. Prompts of 1,024 tokens or more are eligible, and matching is done on the prefix from token index 0, with hits counted in 128-token increments beyond the first 1,024. A prompt that diverges partway through an increment may get no discount on the tokens in that partial block. Anything variable, such as timestamps, user IDs or retrieved snippets, must therefore come after the static content. Anthropic is explicit: you mark boundaries with the cache_control parameter, up to 4 breakpoints per request, so you decide which segments (tools, system prompt, documents, conversation history) are cached. Its minimum is 1,024 tokens for Claude 3.5 Sonnet and Opus, and 2,048 for Claude 3.5 Haiku. A prefix below the minimum is processed normally, with no error and no cache.
Padding a short prompt up to the threshold is a break-even question, not a default. Illustrative case on Anthropic's published 5-minute rates (1.25x write, 0.1x read): padding a 600-token prompt to 1,024 pays off after about three requests within the TTL. On OpenAI, with no write surcharge, padding costs only the extra tokens at the base or cached rate. The exact crossover depends on the model's rates, so test it with the cost calculators and see how token pricing works for the underlying rate structure.
- OpenAI: automatic, 1,024-token minimum, 128-token increments, prefix matched from the start of the prompt.
- Anthropic: explicit cache_control, up to 4 breakpoints, minimum depends on the model (1,024 or 2,048 tokens for the models above).
- Both: keep static content first and variable content last.
Architectural Best Practices: Stabilizing Prefixes in Production
Key takeaways
- Static content first, dynamic content last.
- Any byte change in the prefix is a full-price miss.
- Measure cached versus created tokens on every response.
Cache discounts only apply to an exact prefix match, so the engineering goal is to keep the first N tokens byte-identical across requests. Both providers match from the start of the prompt, and a single changed character early on turns every later token into a full-price miss. Order components from most static to most dynamic: system persona, tool definitions, static documentation, conversation history, then the latest user turn. On Anthropic, the documented cache hierarchy runs tools, then system, then messages, so a change to tool definitions also invalidates the cached system and message blocks behind them. Details are in the Anthropic prompt caching documentation.
Most silent invalidations come from small, innocent-looking changes. The rules below cover the usual causes. For how cached versus uncached input is billed, see our token pricing explainer.
Then verify. Log the usage metadata on every response instead of assuming caching works. If creation tokens keep appearing for a prefix you expected to be warm, something in it is changing. Use the OpenAI developer documentation to confirm current caching behavior and request parameters, since routing and minimum-size rules can change. Once hit rates are measured, you can compare models on effective cost rather than list price.
- Never interpolate timestamps, request IDs, session UUIDs or random seeds into the system prompt or any text ahead of the cached boundary. Put them in the final user turn if the model needs them.
- Serialize tool schemas deterministically: sort functions alphabetically by name and use stable key order within each schema, so JSON serialization cannot reorder between calls.
- Add routing affinity at your gateway so requests sharing a system prompt or project workspace go through the same path. Check whether your provider offers a cache-key or routing hint parameter.
- Append conversation history only; never edit, summarize or reorder earlier turns mid-session, because that rewrites the prefix.
- Track usage.prompt_tokens_details.cached_tokens on OpenAI, and cache_read_input_tokens against cache_creation_input_tokens on Anthropic. Alert when the cache-read share drops for a stable prompt.
Sources & further reading
The details


