Large language model applications often spend significant time generating tokens, but token generation is only part of the performance problem. When users send repeated or related requests, the model may repeatedly process similar prompt prefixes. That repeated computation can increase latency, GPU usage, and infrastructure costs.
KV cache routing is one approach to reducing this repeated work.
Instead of sending every request to an arbitrary inference worker, a routing layer can direct a request to a worker that already has the relevant Key-Value (KV) cache in GPU or other fast-access memory. The model can then reuse previously computed attention states rather than rebuilding them from scratch.
This article explains how KV caching works, why routing matters, and how the technique can improve LLM inference systems.
What Is the KV Cache?
Transformer models use self-attention to determine how each token relates to other tokens in the input sequence.
During inference, the model calculates Key and Value tensors for tokens that have already been processed. These tensors can be stored and reused when generating subsequent tokens.
This stored data is called the KV cache.
Consider a request such as:
You are a customer-support assistant for Contoso.
Company policy:
1. ...
2. ...
3. ...
What is the refund policy for annual subscriptions?
The system may repeatedly receive requests containing the same company policy prefix.
Without caching, the model repeatedly processes that prefix.
With KV caching, the previously calculated attention states can be reused.
Conceptually:
Request
|
v
Tokenization
|
v
Transformer
|
+----> Compute K/V for prefix
|
v
KV Cache
|
v
Generate new tokens
For long prompts, this can represent a significant amount of computation.
Why KV Cache Routing Matters
A KV cache is generally associated with a particular inference worker or cache location.
Suppose four GPU workers are running the same model:
Load Balancer
|
+--------------+--------------+
| | |
GPU 1 GPU 2 GPU 3
Cache A Cache B Cache C
A request containing a prefix already cached on GPU 2 may benefit from being sent to GPU 2.
If the load balancer sends it to GPU 3 instead, GPU 3 may need to recompute the prefix.
This creates a trade-off:
Traditional routing
Request -> Least busy GPU
KV-aware routing
Request -> GPU with useful KV cache
The second approach is often called KV cache-aware routing or KV cache routing.
How KV Cache Routing Works
A KV-aware inference architecture normally has several components:
Client
Routing layer
Cache metadata or index
Inference workers
KV caches
A simplified architecture looks like this:
Client
|
v
+---------------+
| KV-Aware |
| Router |
+---------------+
/ | \
/ | \
v v v
GPU 1 GPU 2 GPU 3
KV KV KV
Cache Cache Cache
When a request arrives, the router determines which worker has the most useful cached prefix.
Step 1: Identify the Prompt Prefix
The routing system examines the request and determines whether part of the prompt matches a previously processed prefix.
For example:
System instructions
+
Company documentation
+
User question
The first two portions may remain unchanged across many requests.
Step 2: Find Cache Ownership
The router maintains information about which worker contains the relevant cache.
Conceptually:
Prefix Hash Worker
abc123 GPU 1
def456 GPU 3
xyz789 GPU 2
The exact implementation varies by inference stack.
Step 3: Select the Worker
The router can combine cache locality with runtime conditions.
For example:
Candidate Worker Cache Match Load
GPU 1 80% High
GPU 2 70% Low
GPU 3 0% Low
The router does not necessarily have to choose the worker with the highest cache match. It can balance cache reuse against current system load.
Step 4: Reuse the Cache
Once the request reaches the selected worker, the inference engine can reuse the available KV blocks.
Only the uncached portion needs to be processed.
KV Cache Routing vs Traditional Load Balancing
Traditional load balancing generally focuses on distributing requests across available workers.
For example:
Request 1 -> GPU 1
Request 2 -> GPU 2
Request 3 -> GPU 3
Request 4 -> GPU 1
This can keep GPU utilization relatively balanced.
However, it ignores computation already performed for a specific request prefix.
KV-aware routing adds another dimension.
Routing Strategy | Main Goal | Cache Awareness |
|---|---|---|
Round robin | Even distribution | No |
Least connections | Balance active requests | No |
Least latency | Reduce observed latency | Usually no |
KV-aware routing | Reuse cached computation | Yes |
Hybrid routing | Balance cache reuse and load | Yes |
For workloads with long repeated prefixes, cache locality can become an important routing signal.
Why Long Contexts Make This More Important
KV caching becomes particularly relevant as context windows grow.
Imagine an application that sends:
System Prompt: 2,000 tokens
Company Documentation: 10,000 tokens
Conversation History: 4,000 tokens
User Question: 50 tokens
If the system repeatedly processes similar requests, recomputing thousands of tokens can be wasteful.
With prefix reuse, much of the previous computation can potentially remain available.
The benefit depends on several factors:
Prefix length
Prefix reuse frequency
Cache hit rate
GPU memory availability
Model architecture
Request scheduling
Cache eviction policy
Number of inference workers
Therefore, simply enabling a KV cache does not automatically guarantee a large performance improvement.
Prefix Caching and KV Cache Routing
KV cache routing is closely related to prefix caching.
Prefix caching allows an inference system to reuse computation for a repeated prompt prefix.
For example:
Request A:
[System Prompt][Documentation][Question A]
Request B:
[System Prompt][Documentation][Question B]
The shared portion is:
[System Prompt][Documentation]
The inference engine may reuse the KV cache associated with that prefix.
Routing becomes important when the cached prefix exists on only one or a subset of workers.
Shared Prefix
|
v
+----+----+----+
| | |
GPU 1 GPU 2 GPU 3
| | |
Cache None None
Sending the second request to GPU 1 can avoid recomputing the shared prefix.
A Simple Routing Algorithm
A conceptual implementation might look like this:
public Worker SelectWorker(
Request request,
IReadOnlyList<Worker> workers)
{
var prefixHash = ComputePrefixHash(request);
var candidates = workers
.Where(worker => worker.HasCachedPrefix(prefixHash))
.OrderBy(worker => worker.CurrentLoad)
.ToList();
if (candidates.Count > 0)
{
return candidates[0];
}
return workers
.OrderBy(worker => worker.CurrentLoad)
.First();
}
This is only a simplified example.
A production router would usually consider additional information such as:
Cache size
Cache age
Prefix match length
GPU memory pressure
Queue depth
Token generation rate
Network latency
Request priority
The important concept is that cache locality becomes part of the scheduling decision.
Partial Cache Matches
One useful optimization is to consider partial prefix matches.
Suppose the cache contains:
System Prompt
+
Company Documentation
but the incoming request contains:
System Prompt
+
Company Documentation
+
Conversation History
+
Question
The worker can still reuse the cached portion.
A router can therefore assign a score based on the amount of reusable context.
For example:
Worker Cached Tokens Queue Length
GPU 1 8,000 4
GPU 2 2,000 1
GPU 3 0 0
A simplistic load-only router chooses GPU 3.
A cache-aware router may choose GPU 1 if the computational savings outweigh the queueing cost.
This leads to a more realistic routing objective:
Routing Score =
Cache Benefit
-
Queueing Cost
-
Transfer Cost
The exact formula depends on the inference system.
When KV Cache Routing Works Well
KV-aware routing is particularly useful for workloads with repeated context.
Enterprise RAG Applications
An enterprise assistant may repeatedly send the same system instructions and retrieved organizational documentation.
For example:
System Instructions
+
Security Policy
+
Product Documentation
+
User Query
The shared prefix can create cache reuse opportunities.
Customer Support
Customer-support systems often reuse:
Agent instructions
Product documentation
Company policies
Response formatting rules
Only the customer-specific portion changes frequently.
Multi-Tenant Applications
A SaaS platform may have many requests associated with the same tenant.
If tenant-specific instructions or documentation are repeatedly included in prompts, routing requests toward workers containing those caches can improve locality.
Agentic Workloads
AI agents can generate multiple related model calls during a single workflow.
For example:
Agent Instructions
+
Task Context
+
Tool Results
+
New Reasoning Step
Some parts of the context may remain stable across successive calls.
When It May Not Help
KV cache routing is not universally beneficial.
If every request has a completely different prompt, there may be little cache reuse.
For example:
Request A -> Unique 20,000-token prompt
Request B -> Unique 18,000-token prompt
Request C -> Unique 15,000-token prompt
In this scenario, traditional load balancing may be more useful.
There is also a cost associated with maintaining cache metadata and making routing decisions.
Another problem is cache eviction.
GPU memory is limited. If the cache becomes too large, older entries must be removed.
Therefore:
More caching
!=
Unlimited performance improvement
The workload's reuse pattern matters.
KV Cache Routing and Distributed Systems
Distributed inference introduces another important issue: cache locality versus load distribution.
Imagine:
Router
|
+-------------+-------------+
| | |
GPU 1 GPU 2 GPU 3
90% 20% 30%
Cache X Cache Y Cache Z
Suppose a request has a strong cache match on GPU 1.
Sending it there may maximize cache reuse but increase queueing.
Sending it to GPU 2 may provide faster scheduling but require recomputation.
A production router therefore needs a policy that considers both.
One conceptual approach is:
score =
cache_match * cache_weight
- queue_depth * load_weight
- network_cost * network_weight
The weights should be tuned using actual workload measurements rather than arbitrary assumptions.
Metrics to Monitor
If implementing KV-aware routing, do not measure only average request latency.
Track metrics such as:
Metric | Why It Matters |
|---|---|
KV cache hit rate | Measures cache effectiveness |
Prefix reuse ratio | Shows how often context repeats |
Time to first token | Measures prompt-processing impact |
Inter-token latency | Measures generation performance |
GPU memory usage | Shows cache pressure |
Cache eviction rate | Indicates insufficient cache capacity |
Queue depth | Shows worker contention |
Tokens processed | Measures workload volume |
Cache transfer volume | Identifies network overhead |
Time to first token is particularly useful because prompt processing can have a noticeable effect on it.
Common Implementation Mistakes
Treating Cache Hits as the Only Routing Signal
A cache hit does not automatically mean the worker is the right choice.
A heavily overloaded worker may be slower than a lightly loaded worker without the cache.
Ignoring Cache Eviction
Caches are finite.
If the workload constantly replaces cached prefixes, the routing layer may create complexity without achieving meaningful reuse.
Over-Centralizing Cache Metadata
A single routing component can become a bottleneck if every request requires expensive global cache lookups.
Distributed or efficient metadata mechanisms may be necessary at larger scale.
Ignoring Request Affinity
Some workloads naturally benefit from session-level affinity.
For example:
User Session A -> GPU 2
User Session A -> GPU 2
User Session A -> GPU 2
Keeping related requests near their existing cache can be simpler than repeatedly rediscovering cache locality.
Measuring Only Throughput
Higher throughput does not automatically mean a better user experience.
Measure latency, especially:
Time to First Token
and end-to-end response latency.
Best Practices
For production LLM inference systems, consider these practices:
Measure prefix reuse before implementing complex routing.
Track KV cache hit rate explicitly.
Combine cache locality with worker load.
Monitor GPU memory pressure and eviction.
Use stable prefix construction where possible.
Avoid unnecessary changes to system prompts and shared context.
Measure time to first token separately from generation latency.
Account for network cost when cache data must move between workers.
Use workload-specific routing policies instead of assuming one algorithm fits every application.
Test cache-aware routing under realistic concurrent traffic.
Conclusion
KV cache routing changes the way LLM inference systems think about request scheduling.
Traditional routing asks:
Which worker can process this request?
KV-aware routing adds another question:
Which worker has already performed useful computation for this request?
That distinction becomes increasingly important for applications with long prompts, repeated instructions, RAG context, multi-turn conversations, and agent workflows.
The key idea is straightforward:
Repeated context
|
v
Reuse KV cache
|
v
Route toward cache locality
|
v
Reduce repeated computation
The real benefit depends on cache hit rate, prefix length, workload characteristics, GPU memory, queueing, and network overhead. For that reason, KV cache routing should be treated as an inference optimization to measure and tune—not as a universal replacement for conventional load balancing.
Summary
KV cache routing can improve LLM inference by directing requests toward workers that already contain useful attention-state data. When prompts share substantial prefixes, this can reduce repeated prompt computation and improve inference efficiency.
The most important production considerations are cache locality, load balancing, memory pressure, eviction, routing overhead, and measurable latency improvements. A hybrid strategy that considers both cache reuse and worker load is often more practical than relying on either signal alone.
Join the conversation! Your thoughts help the community grow.