Introduction
AI applications often send the same instructions, tools, schemas, documents, and conversation context to a model many times. If that repeated input is processed as new input on every request, the application can spend unnecessary money and increase request latency.
OpenAI Prompt Caching is designed for this type of workload. It allows repeated portions of a prompt to be reused instead of being processed as entirely new input each time. The main opportunity for developers is not simply enabling caching, but structuring prompts so that the reusable prefix stays stable.
This becomes particularly useful for AI agents, customer-support systems, coding assistants, document-processing applications, and other applications where a large system prompt or tool definition is reused across many requests.
In this article, we will look at how prompt caching works, how cache hits are created, common implementation mistakes, and practical patterns for improving cache efficiency.
What Is Prompt Caching?
Prompt caching allows a model provider to reuse previously processed prompt content when a later request contains the same prefix.
Consider an application with a large system instruction:
You are an enterprise customer-support assistant.
Follow these policies:
1. Never expose internal information.
2. Verify account ownership before account-specific actions.
3. Use the available tools when appropriate.
4. Return structured responses.
...
The application may then append a different customer question to that same instruction:
Customer question:
How can I reset my password?
On another request, the question changes:
Customer question:
Why was my payment declined?
The user-specific part changes, but the system instructions remain the same.
A useful prompt structure therefore looks like this:
[Stable instructions]
[Stable tools and schemas]
[Stable reference material]
[Dynamic user request]
[Dynamic conversation state]
The stable prefix provides the opportunity for caching.
Why Cache Hits Matter
There are two practical reasons developers should care about prompt caching:
Lower input-token cost
Lower latency for repeated prompt prefixes
The exact economics depend on the model and pricing configuration. OpenAI's current pricing includes separate cached-input pricing for supported models, so cached input can cost substantially less than normal input processing.
For example, an application might send 10,000 input tokens on every request:
8,000 tokens = system instructions, tools, schemas
2,000 tokens = user-specific information
If the 8,000-token prefix can be reused, the application does not need to treat all 10,000 tokens as new input every time.
The important point is that prompt caching does not reduce the amount of information your application sends conceptually. It reduces the cost and processing associated with repeated input when the request structure allows the provider to reuse that prefix.
How Prompt Caching Works
Prompt caching is primarily a prompt-structure problem.
A simplified request might look like this:
System instructions
↓
Tool definitions
↓
Business rules
↓
Reference context
↓
User question
The first sections are relatively stable, while the final section changes frequently.
For caching to be useful, the beginning of the prompt should remain consistent across requests.
For example:
Request 1:
[System]
[Tools]
[Rules]
[Documentation]
[Question A]
Request 2:
[System]
[Tools]
[Rules]
[Documentation]
[Question B]
The prefix remains the same, while only the dynamic portion changes.
By contrast, this structure makes caching less predictable:
Request 1:
[User data]
[System]
[Tools]
[Question]
Request 2:
[Different user data]
[System]
[Tools]
[Question]
The request starts with different content, so the reusable prefix is reduced.
The Most Important Rule: Keep the Prefix Stable
When designing an AI application, place content in this order:
1. Stable system instructions
2. Stable policies
3. Stable tool definitions
4. Stable schemas
5. Stable reference information
6. Dynamic application context
7. User message
For example:
const messages = [
{
role: "system",
content: SYSTEM_INSTRUCTIONS
},
{
role: "system",
content: BUSINESS_RULES
},
{
role: "system",
content: TOOL_GUIDANCE
},
{
role: "user",
content: userQuestion
}
];
The exact API structure depends on the OpenAI API and SDK version you use, but the architectural principle remains the same: put reusable content before frequently changing content.
A Practical Example
Suppose you are building an internal documentation assistant.
You might have:
const systemPrompt = `
You are an internal documentation assistant.
Follow these rules:
- Answer using the supplied documentation.
- Do not invent APIs.
- Clearly state when information is unavailable.
- Prefer concise technical explanations.
`;
const documentation = loadDocumentation();
const userMessage = getUserQuestion();
The request can then conceptually be organized as:
const request = {
model: "your-model",
input: [
{
role: "system",
content: systemPrompt
},
{
role: "system",
content: documentation
},
{
role: "user",
content: userMessage
}
]
};
The system prompt and documentation may remain unchanged across many requests, while userMessage changes.
This is a better caching candidate than rebuilding the entire prompt in a different order for every request.
Static Versus Dynamic Prompt Content
One of the easiest ways to improve cache efficiency is to explicitly separate static and dynamic content.
Content | Typical behavior | Cache strategy |
|---|---|---|
System instructions | Stable | Keep at the beginning |
Tool definitions | Usually stable | Keep unchanged |
JSON schemas | Usually stable | Avoid unnecessary changes |
Product documentation | Relatively stable | Place before dynamic data |
User profile | Changes occasionally | Keep after stable content |
Current user question | Changes frequently | Put near the end |
Timestamp | Changes frequently | Do not put it at the beginning |
Random request ID | Always changes | Keep out of the stable prefix |
Temporary state | Changes frequently | Append after reusable content |
This separation is useful beyond caching. It also makes prompts easier to test, version, and debug.
Avoid Unnecessary Changes to the Prefix
A common mistake is adding dynamic values near the beginning of a large prompt.
For example:
const prompt = `
Request time: ${new Date().toISOString()}
You are a support assistant.
Follow these rules:
...
`;
The timestamp changes on every request.
If it appears before a large amount of otherwise reusable content, it can interfere with prefix reuse.
A better structure is:
const prompt = `
You are a support assistant.
Follow these rules:
...
`;
const dynamicContext = `
Current request time: ${new Date().toISOString()}
User question: ${question}
`;
Now the stable instructions remain separate from frequently changing information.
What About Tool Definitions?
AI agents frequently use function or tool definitions. These definitions can become a significant part of the request.
For example:
{
"name": "get_customer_order",
"description": "Returns order information for a customer.",
"parameters": {
"type": "object",
"properties": {
"orderId": {
"type": "string"
}
},
"required": ["orderId"]
}
}
If the same tools are supplied on every request, avoid generating slightly different versions of the schema.
For example, unnecessary changes such as:
description = "Returns order information."
and later:
description = "Returns information about a customer's order."
create different prompt content.
Keep tool definitions deterministic and version them intentionally.
How to Measure Cache Performance
Do not assume that caching is working simply because your prompts look similar.
Your application should monitor the token usage returned by the API and distinguish cached input from uncached input where the model/API exposes those usage details.
A useful operational metric is:
Cache Hit Ratio =
Cached Input Tokens / Total Input Tokens
For example:
Total input tokens: 1,000,000
Cached input tokens: 750,000
Cache hit ratio = 75%
This gives engineering teams a more useful measurement than simply counting API requests.
You can also track:
Total input tokens
Cached input tokens
Uncached input tokens
Output tokens
Request latency
Cost per request
Cost per user/session
Cache behavior after deployments
The goal is to understand whether prompt architecture is actually producing reusable input.
Common Mistakes
Putting Dynamic Content at the Beginning
This is one of the most common mistakes.
Avoid:
Current user
Current time
Random request ID
System instructions
Tools
Documentation
Question
Prefer:
System instructions
Tools
Documentation
Current user
Current time
Question
Rebuilding Tool Schemas
Generating tool definitions dynamically can introduce small changes that make otherwise identical requests different.
Keep schemas deterministic whenever possible.
Adding Random Values to Prompts
Request IDs, timestamps, experiment identifiers, and random values should generally stay outside the reusable prefix.
Changing Prompt Formatting Unnecessarily
Even harmless formatting changes can make prompt management harder.
For example, avoid repeatedly changing:
Rules:
1. ...
2. ...
to:
Rules
- ...
- ...
unless there is a real reason to change the prompt.
Assuming Every Large Prompt Will Be Cached
Caching is not a replacement for good prompt design. Applications should verify actual usage rather than assuming that repeated requests automatically produce the desired cache behavior.
Troubleshooting Low Cache Hits
If cache utilization is lower than expected, investigate the request structure systematically.
Step 1: Compare Two Requests
Capture sanitized versions of two consecutive requests.
Check whether their prefixes are actually identical.
Step 2: Find the First Difference
Compare the requests from the beginning rather than only comparing the total prompt.
Ask:
Where does Request A first differ from Request B?
This is often more useful than comparing total token counts.
Step 3: Check Dynamic Metadata
Look for:
Timestamps
Request IDs
User-specific values
Experiment flags
Random values
Dynamic tool descriptions
Frequently changing schemas
Step 4: Check Deployments
A prompt update can change the reusable prefix for subsequent requests.
Treat major prompt changes as application changes and monitor cache behavior after deployment.
Step 5: Check Usage Data
Use the token-usage information returned by the API to determine whether input tokens are actually being served from the cache.
Best Practices for Production Applications
1. Separate Static and Dynamic Content
Keep a clear boundary between reusable instructions and request-specific information.
2. Version Large System Prompts
Instead of constructing large prompts from many unrelated pieces, maintain controlled prompt versions.
support-assistant-v3
coding-assistant-v5
document-review-v2
This makes changes easier to track.
3. Keep Tool Definitions Deterministic
Generate the same schema consistently for the same application version.
4. Minimize Unnecessary Prompt Changes
Do not modify a stable prompt simply to change formatting.
5. Monitor Token Usage
Track cached and uncached input tokens as part of application observability.
6. Test Cache Behavior Before Production
Use representative requests and compare their token usage before releasing a prompt architecture change.
7. Do Not Optimize Caching at the Expense of Correctness
A shorter or more stable prompt is not automatically a better prompt.
Security policies, authorization rules, user-specific context, and required instructions must remain correct. Cost optimization should not remove information required for safe and reliable model behavior.
Advantages and Disadvantages
Advantages | Disadvantages |
|---|---|
Can reduce repeated input-token costs | Requires deliberate prompt organization |
Can improve performance for repeated prefixes | Dynamic content can reduce reuse |
Useful for large system prompts | Prompt changes can affect cache behavior |
Works well for repeated tool definitions and context | Requires usage monitoring |
Encourages cleaner prompt architecture | Not every request has enough repeated content to benefit |
Summary
OpenAI Prompt Caching is most useful when an application repeatedly sends the same large prefix to a model.
The key optimization is straightforward: keep stable content together and place dynamic content later in the request.
System instructions, tool definitions, schemas, policies, and stable reference material are good candidates for the reusable portion of a prompt. User questions, timestamps, request IDs, temporary state, and other changing information should generally come afterward.
For production systems, do not judge caching by intuition. Monitor cached input tokens, total input tokens, latency, and application cost. When cache performance changes, compare requests and identify where their prefixes first diverge.
Prompt caching is therefore not just a pricing feature. It is also a prompt-architecture concern. Designing prompts with a stable, reusable structure can make AI applications easier to operate while reducing unnecessary processing of repeated context.

Join the conversation! Your thoughts help the community grow.