Prompt Caching: Design, Economics & Safety
No universal savings claim: caching helps only when an eligible, stable prefix is reused enough times. Measure cache writes, hits, retention, and misses from provider usage data before forecasting savings.
Large prompts often repeat the same system instructions, tool definitions, examples, repository context, or reference documents. Prompt caching—also called context caching by some providers—can reuse work associated with that shared prefix. Provider behavior differs: caching may be automatic or explicit, eligibility and retention vary by model, and billing may separate writes, reads, and storage. Treat caching as an optimization with measurable preconditions, not as a correctness feature.
What gets reused
Caching generally works on a matching prefix. Put stable, broadly shared material first and request-specific content later. Small changes near the beginning can invalidate reuse for everything after them. The exact matching and routing rules are provider-specific, so avoid relying on undocumented behavior.
stable prefix
system policy
tool schemas
reusable examples
shared reference material
changing suffix
user question
current records
request identifiers
Do not move dynamic authorization, current facts, or user-specific secrets into a shared prefix merely to improve hit rate. Correct isolation is more important than cache efficiency.
Automatic and explicit caching
| Mode | Typical behavior | Engineering responsibility |
|---|---|---|
| Automatic or implicit | Provider recognizes eligible repeated prefixes | Keep prefixes stable and inspect cache-hit usage fields |
| Explicit | Application creates or marks cached content and may select retention | Manage lifecycle, identifiers, invalidation, access, and storage cost |
| Application cache | Application reuses a completed response | Define a safe cache key and freshness policy; this is not prompt caching |
Response caching and prompt caching solve different problems. A response cache returns an earlier result and may be unsafe for personalized or time-sensitive questions. Prompt caching still runs the model on the new suffix.
Calculate the break-even point
For a repeated prefix, compare ordinary input cost with the full cache lifecycle:
ordinary cost = uses × prefix tokens × normal input rate
cached cost =
cache-write tokens × write rate
+ cache-read tokens × hit rate
+ missed tokens × normal input rate
+ any retention or storage charge
Some services price writes above normal input and hits below it; others use different rules. A cache used once may cost more than no cache. Use the provider's current rate sheet and model-specific minimums. Include misses caused by prefix changes or retention expiry.
Choose appropriate content
Good candidates are large, stable, and reused:
- versioned system instructions shared across a product;
- unchanged tool definitions and output schemas;
- a document corpus queried repeatedly in a short period;
- long media or transcripts analyzed through multiple questions;
- repository context reused across an engineering session.
Poor candidates include a short prompt, frequently edited context, per-request secrets, current account balances, or content that should not be retained under your privacy requirements.
Version the prefix deliberately
Create a content hash or version identifier for each reusable prefix. A version should change when policy, tool schema, document content, or safety instructions change. Record the version with each request so a regression can be tied to the exact cached context. Never assume a cached policy was invalidated merely because the application deployed new code.
Privacy and tenant isolation
- Review the provider's data controls and retention terms for the chosen endpoint and cache mode.
- Do not share cache identifiers across tenants unless the content is intentionally identical and non-sensitive.
- Exclude credentials, private keys, session tokens, and unnecessary personal information.
- Apply the same data classification to cached context as to the original prompt.
- Verify whether a retention option is compatible with contractual or zero-retention requirements.
A cache discount does not change who is authorized to see the underlying content. Access control remains an application responsibility.
Operational monitoring
Track cache-eligible tokens, write tokens, hit tokens, misses, hit ratio, cost per accepted task, and latency. Segment metrics by prefix version and model. A high request-level hit ratio can still save little if only a small part of each prompt is cached; token-level metrics are more informative.
Alert on a sudden hit-rate drop after prompt or tool-schema releases. Compare quality and safety evaluations before and after prefix changes—the cached material may be stable while its instructions are wrong.
Failure modes to test
- one-character changes near the start of the prefix;
- cache expiry during a long-running workflow;
- requests routed to a different model or region;
- tenant or user identifiers accidentally entering the shared block;
- tool schema changes without a cache-version change;
- provider responses with no cache-usage field;
- fallback requests that use a provider with different cache semantics.
A worked break-even example
Illustrative rates—not a current provider quote: suppose a reusable prefix contains 20,000 tokens. Ordinary input costs $2.50 per million tokens, a cache write costs $3.125 per million, and each cache read costs $0.25 per million. Ignore retention charges for this exercise.
| Prefix uses before expiry | Ordinary input | One write + later reads | Difference |
|---|---|---|---|
| 1 | $0.0500 | $0.0625 | Cache costs $0.0125 more |
| 2 | $0.1000 | $0.0675 | Cache saves $0.0325 |
| 5 | $0.2500 | $0.0825 | Cache saves $0.1675 |
| 10 | $0.5000 | $0.1075 | Cache saves $0.3925 |
Under these assumptions, the second use crosses break-even. Real systems must also multiply reads by the observed hit rate and add misses, storage, retention, model-routing, and expiry behavior. If only 40% of requests hit, the optimistic ten-use row is not a valid forecast.
Build a prefix fingerprint you can audit
A cache version such as support-policy-v7 is readable but does not prove which bytes were sent. Pair a human version with a normalized cryptographic fingerprint. Normalize only transformations that your request builder also performs; otherwise the hash and transmitted prefix can disagree.
async function prefixFingerprint(prefixText) {
const normalized = prefixText.replace(/\r\n/g, "\n").trimEnd();
const bytes = new TextEncoder().encode(normalized);
const digest = await crypto.subtle.digest("SHA-256", bytes);
return [...new Uint8Array(digest)]
.map(byte => byte.toString(16).padStart(2, "0"))
.join("");
}
Log the model ID, prefix version, fingerprint, prefix token count, tenant scope, and provider cache-usage fields. Do not log the prefix itself when it contains sensitive information.
A reproducible cache experiment
This protocol produces evidence without claiming that results transfer between providers or models:
- Select 30 privacy-safe prompts representing short, medium, and long suffixes.
- Freeze one eligible prefix, model version, region, endpoint, tool schema, and generation configuration.
- Run each suffix once with caching disabled or with a deliberately cold prefix.
- Run the same ordered set five times with caching enabled, spacing repetitions across the retention window you intend to use.
- Record billed uncached input, cache-write, cache-read, and output tokens plus latency and task acceptance.
- Repeat with one early-prefix character changed, one late-suffix change, and an expired cache.
- Compare total cost and latency per accepted result; publish the exact date and assumptions with the result.
| Metric | Why request count alone is insufficient |
|---|---|
| Token hit ratio | A request can “hit” while only a small prefix is discounted |
| Write/read/miss cost | A high hit rate can still lose money after expensive writes |
| p50 and p95 latency | Caching may affect long prompts differently from short ones |
| Accepted-task rate | A prefix change can save tokens while degrading instructions |
| Cross-tenant isolation | Cost savings never justify data leakage |
Diagnose misses systematically
| Observed symptom | Likely checks | Evidence to collect |
|---|---|---|
| Hits disappear after deploy | System prompt, tool order, whitespace, serialization | Old/new fingerprint and first differing byte |
| Only some workers hit | Region, model route, endpoint, account scope | Route and provider request identifiers |
| High request hits, low savings | Cached prefix is a small share of total input | Token-level hit ratio |
| Unexpected stale behavior | Application response cache confused with prompt cache | Trace which layer returned the output |
| Quality changes after version bump | Instruction or example content changed | Evaluation delta by prefix version |
Implementation checklist
- Confirm model eligibility, minimum length, pricing, and retention in official docs.
- Measure the repeated prefix in a realistic trace.
- Move stable content before dynamic content without weakening isolation.
- Add a prefix version and tenant-safe cache key.
- Run a no-cache baseline and a cache-enabled test.
- Calculate total write, hit, miss, and retention cost per accepted task.
- Document invalidation and privacy behavior before production use.
Update, September 2026: the cache price war
Between mid-August and mid-September 2026, five vendors repriced cache reads within roughly three weeks of each other: Anthropic cut Fable 5.1 cache reads from $1.00 to $0.25 per million (2.5% of input, versus 10% on its other models), Google listed Gemini 3.8 Flash cached input at $0.075, OpenAI's GPT-6 Astra carries cached input at $1.00 with writes at $12.50, and DeepSeek's V4.1 Flash serves cache hits at $0.003 off-peak. On a worked 200-step agent task, two flagships with identical $10/$50 list prices now bill 1.8x apart, entirely because of this line item. The verified rates and the worked bill are in The September 2026 Cache Price War. The design guidance on this page is unchanged, but the economics it describes moved: re-check your provider's cache row before the next budget cycle.
Primary references
- OpenAI Prompt Caching guide
- Anthropic Prompt Caching guide
- Gemini context caching guide
- OpenAI endpoint data controls
Bottom line
Cache stable, repeated prefixes only after confirming provider rules. Version and isolate the content, measure token-level hits and full lifecycle cost, and keep privacy and invalidation decisions outside the optimization.