Introduction
Large AI applications often send the same instructions, tool definitions, reference material, and conversation context to a model repeatedly.
For example, an AI coding assistant may send:
System instructions
Developer instructions
Tool definitions
Project documentation
Repository context
Previous conversation
New user request
Only the final user request may change between calls.
Without prompt caching, the model still needs to process the reusable input on every request.
OpenAI's prompt caching allows the model to reuse processing for matching prompt prefixes. This can reduce input-token cost and latency when the same context is sent repeatedly. Current OpenAI documentation states that supported models can reuse cached prompt prefixes automatically, while newer models also provide explicit controls for choosing cache breakpoints.
The important point is that prompt caching is not application-level response caching.
It does not return an old answer.
Instead, the model reuses previously processed input context and generates a new response.
What Is Prompt Caching?
A language model processes input tokens before generating a response. During this processing, it creates internal key-value states that help the model work with the input.
Prompt caching allows OpenAI to preserve those states for a reusable prefix.
Consider two requests:
Request 1
Developer instructions
Tool definitions
Project documentation
User question A
and:
Request 2
Developer instructions
Tool definitions
Project documentation
User question B
The first three sections are identical.
The reusable portion is:
Developer instructions
+
Tool definitions
+
Project documentation
The changing part is:
User question A
versus:
User question B
If the common prefix qualifies for caching, the second request can reuse the cached processing.
OpenAI describes this as caching the reusable prompt prefix rather than caching the generated response.
Prompt Caching Is About Prefixes
This is the most important concept to understand.
Prompt caching does not generally mean:
"Cache anything anywhere inside my prompt."
It works around reusable prefixes.
For example:
Stable content
Stable content
Stable content
Stable content
Changing content
Changing content
is a good structure.
But:
Stable content
Changing content
Stable content
Stable content
makes the later stable content less useful for prefix reuse.
This is why prompt structure matters.
OpenAI's current documentation recommends keeping reusable instructions, tools, and reference material stable and placing changing content later in the request.
A Simple Example
Suppose an application sends 8,000 input tokens per request.
The first 6,000 tokens are stable:
6,000 stable tokens
+
2,000 changing tokens
If many requests reuse the same first 6,000 tokens, those tokens can potentially be served from the prompt cache.
Conceptually:
Request 1
6,000 uncached
2,000 uncached
Request 2
6,000 cached
2,000 uncached
Request 3
6,000 cached
2,000 uncached
The model still processes the new 2,000 tokens.
Prompt caching does not mean the entire request becomes free.
How Cache Hits Are Reported
OpenAI API responses expose cache usage through the usage information.
For example:
const usage = response.usage;
console.log(
"Cached tokens:",
usage.input_tokens_details.cached_tokens
);
The current API documentation also exposes cache-write information for models where cache writes have a separate charge.
A production application should capture at least:
input_tokens
cached_tokens
cache_write_tokens
latency
model
This lets you determine whether your prompt structure is actually producing useful cache reuse.
Cache Hit Rate
A simple cache-hit metric can be calculated as:
Cache Hit Rate =
Cached Input Tokens
------------------
Total Input Tokens
× 100
For example:
Total input tokens = 10,000,000
Cached tokens = 7,500,000
Then:
7,500,000 / 10,000,000 × 100
= 75%
This is more useful than simply asking whether prompt caching is enabled.
The important question is:
How much of the application's input traffic is actually being reused?
OpenAI's Prompt Caching documentation recommends monitoring cached-token usage and cache hit rates rather than assuming that caching is working optimally.
Why Cache Hits Can Reduce Cost
Cached input tokens are priced below ordinary input tokens on supported models.
The exact discount depends on the model and current pricing.
For example, OpenAI's current prompt-caching documentation states that GPT-5.6 and later use a reduced cached-input rate, with cached reads priced at 0.1× the standard uncached input rate. Cache writes on those models are priced at 1.25× the standard input rate.
This means developers should not think of caching as:
Cache = always free
A more accurate model is:
First reusable request
↓
Cache write
↓
Later matching requests
↓
Lower-cost cache reads
The savings become more significant when the same prefix is reused repeatedly.
Cache Writes Also Matter
A common mistake is measuring only cached tokens.
For current GPT-5.6-and-later behavior, cache writes themselves can have a higher input-token charge than ordinary uncached input.
Suppose you generate a large unique prompt for every request.
If the prefix changes every time:
Request 1 → Write
Request 2 → Write
Request 3 → Write
Request 4 → Write
you may not receive meaningful cache-read benefits.
A good caching strategy therefore requires:
Stable prefix
+
Repeated requests
+
Useful cache lifetime
=
Potential savings
Keep Stable Instructions at the Beginning
Consider an AI support application.
A poor structure might be:
System instructions
Today's customer information
Tool definitions
Company documentation
User question
If customer information changes frequently, the reusable prefix may be interrupted early.
A better structure is:
System instructions
Company policies
Tool definitions
Reference documentation
Customer information
User question
Now the stable information appears before the dynamic content.
Conceptually:
┌─────────────────────────────┐
│ Stable prefix │
│ │
│ Instructions │
│ Tools │
│ Reference material │
├─────────────────────────────┤
│ Dynamic suffix │
│ │
│ Customer data │
│ Current request │
└─────────────────────────────┘
This structure gives the cache a larger reusable prefix.
Tool Definitions Can Affect Cache Reuse
Tool definitions are part of the model input context.
Suppose your application sends:
Tool A
Tool B
Tool C
Tool D
on every request.
If the definitions remain identical, they can contribute to a reusable prefix.
But if the application dynamically changes tool descriptions or ordering:
Request 1:
Tool A
Tool B
Tool C
Request 2:
Tool B
Tool A
Tool C
the prompt structure is no longer identical.
Even though the same tools exist, their order has changed.
For cache optimization, keep tool definitions stable whenever practical.
OpenAI's prompt-caching guidance specifically recommends keeping tool definitions stable when they are part of reusable context.
Keep Prompt Formatting Stable
Small structural changes can affect prefix matching.
For example:
Project:
Invoice API
Database:
PostgreSQL
versus:
Project: Invoice API
Database: PostgreSQL
may represent the same information semantically, but the token sequence is different.
The cache operates on the processed input sequence, not on the semantic meaning of the text.
Therefore, applications should avoid unnecessary formatting changes in reusable sections.
Keep stable:
wording
ordering
tool definitions
reference material
message structure
serialization format
Put Dynamic Information at the End
A useful pattern is:
[Stable instructions]
[Stable tools]
[Stable reference data]
[Current conversation]
[Current user request]
For example:
You are a support assistant.
Company policy:
...
Available tools:
...
Product documentation:
...
Customer:
Customer ID: 48291
Plan: Enterprise
Current request:
Why did my deployment fail?
The reusable material remains near the beginning.
The dynamic information is appended later.
This pattern is especially useful for:
customer support agents
coding assistants
document analysis
internal enterprise assistants
multi-turn agents
API copilots
Use Explicit Cache Breakpoints When Appropriate
Current OpenAI models such as GPT-5.6 and later support explicit prompt-cache breakpoints.
This provides more control over where reusable context should end.
Conceptually:
Stable instructions
Stable tools
Stable documentation
↓
Cache breakpoint
↓
Current request
OpenAI's API documentation describes prompt_cache_options and explicit prompt_cache_breakpoint controls for supported models.
This is useful when you have a large stable prefix followed by highly variable content that should not be written into the cache.
Implicit vs Explicit Caching
There are two important approaches.
Implicit Caching
The platform determines eligible cache boundaries automatically.
This is convenient when your application naturally sends repeated prefixes.
For example:
Stable conversation history
+
New user message
This can work well for multi-turn interactions.
Explicit Caching
Your application identifies where a reusable section ends.
For example:
System instructions
Tool definitions
Reference documentation
↓
Explicit breakpoint
↓
User-specific information
Explicit caching gives developers more control over what gets written.
OpenAI's current documentation supports both approaches on GPT-5.6 and later.
Example With the Responses API
A simplified request might look like:
const response = await client.responses.create({
model: "gpt-5.6",
input: [
{
role: "developer",
content: [
{
type: "input_text",
text: `
You are an enterprise support assistant.
Follow these company policies:
1. Never expose secrets.
2. Validate customer identity.
3. Use the approved troubleshooting process.
`
}
]
},
{
role: "user",
content: "Why did my deployment fail?"
}
]
});
The exact cache configuration should follow the model and API features supported by the application.
The important architectural principle is to keep the reusable developer context stable.
Monitoring Cache Performance
Do not assume that a well-designed prompt automatically produces a high cache-hit rate.
Measure it.
A simple logging structure might be:
const details = response.usage?.input_tokens_details;
console.log({
inputTokens: response.usage?.input_tokens,
cachedTokens: details?.cached_tokens,
cacheWriteTokens: details?.cache_write_tokens
});
Then aggregate the information.
For example:
Day | Input Tokens | Cached Tokens | Hit Rate |
|---|---|---|---|
Monday | 10M | 7.8M | 78% |
Tuesday | 12M | 9.4M | 78% |
Wednesday | 15M | 8.1M | 54% |
Thursday | 14M | 10.7M | 76% |
The Wednesday drop deserves investigation.
Possible causes include:
changed tool definitions
different model routing
changed system instructions
different prompt structure
conversation compaction
increased one-off requests
insufficient prefix reuse
OpenAI now provides prompt-cache diagnostics that can help identify reasons for cache misses.
Cache Misses Are Diagnostic Signals
A cache miss is not necessarily a problem.
Some requests are supposed to be unique.
For example:
Generate a one-time analysis of this document.
may not have a reusable prefix.
But repeated cache misses in a workload expected to share context can indicate an architecture problem.
Current diagnostics can identify conditions such as model changes and other differences that prevent reuse.
The right question is therefore:
Is this cache miss expected?
rather than:
How do I eliminate every cache miss?
Use a Stable Model
Changing models can break cache reuse because the cached processing is associated with the model.
Consider:
Request 1 → Model A
Request 2 → Model B
Even if the prompts are identical, the requests are not equivalent cache targets.
A production application using multiple models should therefore understand that model routing can influence cache reuse.
OpenAI's prompt-cache diagnostics explicitly identify model changes as one possible reason for cache misses.
Be Careful With Dynamic System Prompts
Consider an application that creates a different system prompt for every customer:
You are an assistant.
Customer name: Alice
Customer plan: Enterprise
Customer region: India
Customer ID: 84921
Then:
You are an assistant.
Customer name: Bob
Customer plan: Standard
Customer region: US
Customer ID: 29174
The prefix changes immediately.
Instead, where appropriate, separate stable instructions from dynamic customer context:
Stable instructions
Company policies
Tool definitions
↓
Customer context
Current request
This increases the amount of reusable input.
Multi-Turn Conversations
Prompt caching is particularly useful for applications where the same conversation context is sent repeatedly.
Consider:
Turn 1
System
Tools
Conversation
User question
Turn 2
System
Tools
Conversation
User question
Assistant answer
Turn 3
System
Tools
Conversation
User question
Assistant answer
New question
Much of the conversation may remain unchanged between requests.
That creates an opportunity for prefix reuse.
However, maintaining a session does not guarantee a cache hit. OpenAI's current Agents API documentation explicitly notes that sessions can preserve context while prompt caching still depends on the requests sharing a matching prompt prefix.
Prewarming a Cache
For applications where a large shared context is known before users arrive, current GPT-5.6-and-later APIs support cache prewarming.
The idea is:
Application startup
↓
Prepare reusable context
↓
Prewarm cache
↓
User request
↓
Reuse prepared prefix
This can be useful for applications with large shared instructions or reference material.
OpenAI documents prompt_cache_options.prewarm for preparing cacheable context without generating a normal response.
Prewarming should still be evaluated economically because cache writes have a cost.
When Prompt Caching Does Not Help Much
Prompt caching is not equally useful for every workload.
It provides less benefit when:
requests are mostly unique
prompts are short
the reusable prefix is below the model's cacheable threshold
system instructions change frequently
tool definitions change frequently
model selection changes frequently
requests have little repeated context
traffic is too sparse for useful reuse
For current GPT-5.6-and-later models, the minimum visible input needed for caching is 1,024 tokens. The exact behavior and thresholds vary by model, so developers should check the model-specific documentation rather than hard-coding assumptions.
Do Not Add Useless Text Just to Reach the Threshold
Suppose an application has:
800 reusable tokens
A developer might think:
Add 224 random tokens so caching becomes available.
That is not a good optimization.
Additional input has a cost and may affect model behavior.
If a prompt naturally contains useful stable information such as:
examples
tool definitions
domain instructions
reference material
then reaching the cacheable threshold may happen naturally.
Do not add meaningless content merely to make a prompt longer.
OpenAI's current documentation also discusses the economics of expanding short prefixes to reach the cacheable threshold and recommends measuring whether the additional tokens are actually worthwhile.
Prompt Caching vs Response Caching
These are different techniques.
Feature | Prompt Caching | Response Caching |
|---|---|---|
Reuses | Processed input prefix | Previous output |
Generates new response | Yes | Usually no |
Useful for changing questions | Yes | Usually not |
Requires identical request | No | Usually |
Reduces input processing | Yes | Not the primary goal |
Suitable for conversational agents | Yes | Limited |
Handles new user questions | Yes | No, unless explicitly designed |
For example:
Same instructions
+
Question A
and:
Same instructions
+
Question B
can still benefit from prompt caching.
A response cache would generally need to recognize whether Question B is equivalent to an earlier request.
Prompt Caching vs RAG
Prompt caching and RAG solve different problems.
RAG answers:
Which information should I retrieve?
Prompt caching answers:
Can previously processed input be reused?
They can work together.
For example:
User request
↓
RAG retrieves relevant documents
↓
Stable instructions + tools + selected context
↓
Model
If the beginning of the resulting prompt remains stable across requests, prompt caching may reduce the processing cost of that repeated prefix.
Security Considerations
Cached context can contain sensitive information.
Do not assume that because caching improves cost efficiency, it should automatically be enabled without considering data boundaries.
Applications handling multiple customers should carefully design cache accounting and isolation.
Current OpenAI documentation supports prompt_cache_key on GPT-5.6 and later for separating cache accounting between customers, users, or workspaces. On those models, the key is not required for optimizing cache routing, but can be useful for accounting and isolation-related considerations.
For earlier models, a stable prompt_cache_key can help route related requests toward reusable cached prefixes.
Common Mistakes
Changing the Prefix on Every Request
If the reusable instructions change unnecessarily, cache reuse decreases.
Putting Dynamic Data at the Beginning
Customer-specific or request-specific content should generally appear after reusable context when the application design permits it.
Assuming Sessions Guarantee Cache Hits
A session preserves conversation state, but it does not guarantee that the resulting request will produce a cache hit.
Measuring Only Latency
A latency improvement is useful, but measure cached tokens and actual input cost as well.
Ignoring Cache Writes
Cache writes can have their own pricing on newer models.
Changing Models Frequently
Model changes can prevent reuse.
Optimizing for 100% Cache Hits
Some requests are naturally unique.
The objective is to maximize useful cache reuse, not eliminate every miss.
Best Practices
For production OpenAI API applications:
Keep reusable instructions at the beginning of the prompt.
Place dynamic user-specific content later.
Keep tool definitions stable when possible.
Avoid unnecessary changes to prompt formatting.
Use explicit cache breakpoints when the model supports them and precise control is useful.
Measure
cached_tokensand cache-write usage.Track cache-hit rates by workload.
Investigate unexpected cache misses.
Keep related requests on compatible model configurations.
Do not add meaningless tokens just to reach a caching threshold.
Consider prewarming for large, frequently reused prefixes.
Review cache behavior when changing prompts or tools.
Separate customer cache accounting where appropriate.
Calculate actual cost savings instead of assuming them.
A Practical Cache Optimization Workflow
Use this process when optimizing an existing application.
Step 1: Measure
Capture:
Input tokens
Cached tokens
Cache writes
Latency
Model
Step 2: Find the Stable Prefix
Identify which parts remain unchanged across requests.
Instructions
Tools
Reference material
Step 3: Move Dynamic Content
Where appropriate, place:
Customer data
Current request
Temporary context
after the stable content.
Step 4: Stabilize Serialization
Keep ordering, formatting, and tool definitions consistent.
Step 5: Test
Run the same workload repeatedly.
Step 6: Compare
Measure:
Before:
Cache hit rate = 32%
After:
Cache hit rate = 81%
Then compare actual input cost and latency.
Step 7: Monitor After Deployment
Prompt changes, tool changes, model upgrades, and routing changes can affect cache behavior.
Treat prompt caching as an observable production feature rather than a one-time optimization.
Production Checklist
Before deploying prompt caching optimization, verify:
Reusable prompt prefixes have been identified.
Stable instructions appear before dynamic content.
Tool definitions remain stable where practical.
Prompt formatting is consistent.
Model selection is understood.
Cached-token usage is logged.
Cache-write usage is monitored where applicable.
Cache-hit rate is measured.
Unexpected cache misses can be investigated.
Cache lifetime is appropriate for the workload.
Sensitive customer data has been reviewed.
Cache accounting is separated where necessary.
Actual input-cost savings have been measured.
Performance improvements have been validated.
Summary
OpenAI Prompt Caching is most useful when an application repeatedly sends the same large input prefix.
The basic optimization is straightforward:
Stable context
↓
Stable tools
↓
Stable reference material
↓
Dynamic request
The more often the stable prefix is reused, the more opportunity there is to reduce repeated input processing, latency, and cached-input cost. Current OpenAI documentation provides automatic prompt caching on supported models and explicit breakpoint controls on GPT-5.6 and later.
The biggest mistake is treating caching as a switch that automatically solves AI costs.
Instead, measure:
Prompt structure
↓
Cache hits
↓
Cache writes
↓
Input-token cost
↓
Latency
A well-designed prompt architecture can turn repeated context from an expensive part of every request into reusable processing.
The practical goal is not to cache everything.
It is to keep the right context stable, reuse it frequently, and verify that the resulting cache hits actually improve cost and performance.

Join the conversation! Your thoughts help the community grow.