The September 2026 Cache Price War: Same List Price, Very Different Bill
Scope: Every cache-related rate below was read from the provider's own pricing or release-notes page and is dated. Third-party figures are labeled as such. The worked bill at the center of this article uses a stated model with every assumption listed, so you can substitute your own numbers.
Two frontier models list at exactly $10 per million input and $50 per million output. Run the same 200-step agent task on both, and one bills about $7.49 while the other bills about $13.46. Nothing about the models' intelligence explains the gap. It is one line item: the cache read.
Between mid-August and mid-September 2026, five labs repriced prompt caching within roughly three weeks of each other. That is not a coincidence, and understanding why it happened will change how you read any pricing page for the next year.
The three weeks, in order
| Date | Vendor | What changed |
|---|---|---|
| Aug 12, 2026 | xAI | Grok 4.6 launches with cache reads at a quarter of input ($0.50 vs $2.00 on the sub-200K tier) |
| Sep 1, 2026 | Anthropic | Fable 5.1 and Mythos 5.1 cut cache reads 75%, from $1.00 to $0.25 per million |
| Sep 2, 2026 | Gemini 3.8 Flash carries cached input at $0.075 per million | |
| Sep 3, 2026 | OpenAI | GPT-6 Astra lists cached input at $1.00, cache writes at $12.50 per million |
| Sep 10, 2026 | DeepSeek | V4.1 Flash serves cache hits at $0.003 off-peak, $0.006 peak, per million |
Source note: the Anthropic figures are from Anthropic's own release notes and pricing documentation; OpenAI, Google, and DeepSeek figures are from their published pricing pages as read on September 19, 2026. The Grok 4.6 row is compiled from xAI's published rates by third-party trackers and is the one line we have not independently re-verified against xAI's own page; treat it accordingly.
Why everyone moved at once
An agent does not read your context once. It re-reads the same system prompt, tool definitions, and conversation history on every step of a task. A 200-step task with a 40,000-token prefix reads that prefix 199 more times than a chat application would. Multiply that across an agent fleet and the cache-read rate — typically a tenth of input or less — quietly becomes the biggest number in the bill.
September is also when agents stopped being a demo. OpenAI put its managed Agents API into public beta on September 10. Salesforce shipped seven named, job-ready Agentforce agents on September 11. GitHub shipped Project HydraFusion on September 4, routing each coding task across multiple models — GitHub's own evaluations report workflow cost down 67 percent on Terminal-Bench 2.1 and 36 percent on DeepSWE versus a single frontier model, with the honest caveat, highlighted by early coverage, that cost fell on every benchmark while quality matched on only one.
When the industry's growth workload is cache-heavy by construction, the cache read becomes the price worth cutting. Five labs reached the same conclusion in the same month.
The verified rate card
Cache-relevant rates per million tokens, as published by each vendor. Where a vendor distinguishes tiers or time windows, both are shown.
| Model | Input (uncached) | Cache read | Cache write | Output |
|---|---|---|---|---|
| Claude Fable 5.1 / Mythos 5.1 | $10.00 | $0.25 | $12.50 (5-min) / $20.00 (1-hr) | $50.00 |
| GPT-6 Astra | $10.00 | $1.00 | $12.50 (1.25x) | $50.00 |
| Gemini 3.8 Flash (intro rate to Dec 31) | $0.75 | $0.075 | Free to write; storage billed hourly | $3.75 |
| DeepSeek V4.1 Flash, off-peak | $0.15 (cache miss) | $0.003 | First write bills at the miss rate | $0.60 |
| DeepSeek V4.1 Flash, peak | $0.30 | $0.006 | First write bills at the miss rate | $1.20 |
| DeepSeek V4 Pro, off-peak | $0.66 | $0.022 | First write bills at the miss rate | $1.98 |
Peak hours on DeepSeek are 01:00–04:00 and 06:00–10:00 UTC on weekdays; everything else, including weekends, is off-peak at half the rate. Anthropic's Fable 5.1 and Mythos 5.1 cache reads are 2.5% of base input — the release notes state other Anthropic models sit at 10%.
One task, five bills
The model task: an agent with a 40,000-token stable prefix — system prompt, tool definitions, working context — runs 200 steps and produces 500 output tokens per step. The prefix is written to cache once, then read on each of the remaining 199 steps. Every step completes within the cache's validity window, so the prefix is never rewritten. Token counts are fixed for cross-vendor comparison; real tokenizers differ, and Claude models from 4.7 onward produce roughly 30% more tokens for the same text, which would raise Anthropic's actual figure.
Total volumes: 7.96 million cache-read tokens, 40,000 cache-write tokens, 100,000 output tokens.
| Model | Cache write | Cache reads | Output | Total |
|---|---|---|---|---|
| Claude Fable 5.1 | $0.50 | $1.99 | $5.00 | $7.49 |
| GPT-6 Astra | $0.50 | $7.96 | $5.00 | $13.46 |
| Claude Fable 5 (previous cache rate) | $0.50 | $7.96 | $5.00 | $13.46 |
| Gemini 3.8 Flash | — | $0.60 | $0.38 | $0.97* |
| DeepSeek V4.1 Flash, off-peak | $0.006 | $0.024 | $0.06 | $0.09 |
*Excludes Google's hourly storage charge for cached content, which Google prices by how long you hold the cache rather than how much you put in; we have not seen a published per-hour figure in the pricing page's summary and are not going to invent one.
Three observations fall out of this table.
The identical-price flagships are not identically priced. Fable 5.1 and GPT-6 Astra share a rate card on input and output, yet the task costs 1.8x more on Astra. Anyone comparing flagships by list price in September 2026 is comparing the wrong numbers.
Anthropic's own estimate checks out. Anthropic states the Fable 5.1 cache cut reduces highly agentic workload cost by up to about 45%. This model, using only published rates, produces 44% — the difference between Fable 5's $13.46 and Fable 5.1's $7.49. A vendor claim matching an independent calculation to within a point is worth recording, because it is not the norm.
Caching is not a discount. It is the product. The same Fable 5.1 task with caching disabled — all input at the uncached rate — bills about $85.00. Caching does not shave 10% off an agent bill; it removes the bill.
The spread across vendors is also wider than the headline rates suggest. On input price alone, the most expensive model here is 67x the cheapest. On this worked task, it is about 150x. Cache pricing amplifies differences rather than compressing them, because the cheapest vendors also discount their cache hardest.
The fine print that moves real bills
Writes are billed, and the majors disagree on how. Anthropic charges 1.25x the uncached input rate for a five-minute cache and 2x for the one-hour tier on Fable 5.1 ($12.50 and $20.00 per million). OpenAI charges 1.25x with no time-tier distinction. Google charges nothing to write but bills storage by the hour. DeepSeek has no separate write price; the first, uncached pass simply bills at the cache-miss input rate. A cache that is written and never read is a net cost on every one of these cards.
Caches expire. The common default window is five minutes from the last request that touches the cache, extendable to an hour on Anthropic. An agent that stays busy keeps its cache warm indefinitely; an agent that idles for six minutes pays a fresh write — at the write premium — on its next step. Bursty and steady workloads see different savings from identical nominal discounts, which is why cache-heavy architectures favor long-running tasks over cron jobs.
The comparison tables are still catching up. As of this week, published AI pricing comparisons still list GPT-5.6 Sol at $5/$30 — the launch-announcement rate — rather than the $4/$20 promotional rate on OpenAI's own model documentation, and most omit cache rows entirely. A comparison without a cache column is a comparison of a workload you probably do not have. Our September pricing update records which provider pages disagree and which one determines your invoice.
What to actually do
- Run your own numbers with your real prefix size and step count; the sensitivity to cache-read rate grows with steps, so long-horizon tasks benefit most.
- If you run agent loops on a $10-input flagship, re-evaluate against Fable 5.1's cache economics before assuming Astra and Fable are interchangeable on cost.
- Measure your cache hit rate before optimizing anything else. The gap between a 90% and a 99% hit rate is larger than the gap between most vendors.
- Check whether your workload is steady or bursty. Bursty workloads lose cache savings to expiry windows and re-write premiums.
- Re-read the vendor pricing page — not an announcement, not a comparison table — before each budget cycle. The September record shows these rates move in weeks, not quarters.
Bottom line
September 2026 moved the real price of agents from the headline row to the cache row, and five vendors repriced within three weeks because that is where the money now is. Two flagships with identical list prices now bill 1.8x apart on the same task. Any model comparison you make after this month without a cache column is incomplete by construction.