Large language model applications often spend significant time generating tokens, but token generation is only part of the performance problem. When users send repeated or related requests, the model may repeatedly process similar prompt prefixes. That repeated computation can increase latency, GPU usage, and infrastructure costs.

KV cache routing is one approach to reducing this repeated work.

Instead of sending every request to an arbitrary inference worker, a routing layer can direct a request to a worker that already has the relevant Key-Value (KV) cache in GPU or other fast-access memory. The model can then reuse previously computed attention states rather than rebuilding them from scratch.

This article explains how KV caching works, why routing matters, and how the technique can improve LLM inference systems.

What Is the KV Cache?

Transformer models use self-attention to determine how each token relates to other tokens in the input sequence.

During inference, the model calculates Key and Value tensors for tokens that have already been processed. These tensors can be stored and reused when generating subsequent tokens.

This stored data is called the KV cache.

Consider a request such as:

You are a customer-support assistant for Contoso.
Company policy:
1. ...
2. ...
3. ...

What is the refund policy for annual subscriptions?

The system may repeatedly receive requests containing the same company policy prefix.

Without caching, the model repeatedly processes that prefix.

With KV caching, the previously calculated attention states can be reused.

Conceptually:

Request
   |
   v
Tokenization
   |
   v
Transformer
   |
   +----> Compute K/V for prefix
   |
   v
KV Cache
   |
   v
Generate new tokens

For long prompts, this can represent a significant amount of computation.

Why KV Cache Routing Matters

A KV cache is generally associated with a particular inference worker or cache location.

Suppose four GPU workers are running the same model:

                 Load Balancer
                      |
       +--------------+--------------+
       |              |              |
     GPU 1          GPU 2          GPU 3
     Cache A        Cache B        Cache C

A request containing a prefix already cached on GPU 2 may benefit from being sent to GPU 2.

If the load balancer sends it to GPU 3 instead, GPU 3 may need to recompute the prefix.

This creates a trade-off:

Traditional routing
Request -> Least busy GPU

KV-aware routing
Request -> GPU with useful KV cache

The second approach is often called KV cache-aware routing or KV cache routing.

How KV Cache Routing Works

A KV-aware inference architecture normally has several components:

  1. Client

  2. Routing layer

  3. Cache metadata or index

  4. Inference workers

  5. KV caches

A simplified architecture looks like this:

                    Client
                      |
                      v
              +---------------+
              | KV-Aware      |
              | Router        |
              +---------------+
                 /     |     \
                /      |      \
               v       v       v
            GPU 1    GPU 2    GPU 3
             KV        KV       KV
           Cache     Cache    Cache

When a request arrives, the router determines which worker has the most useful cached prefix.

Step 1: Identify the Prompt Prefix

The routing system examines the request and determines whether part of the prompt matches a previously processed prefix.

For example:

System instructions
        +
Company documentation
        +
User question

The first two portions may remain unchanged across many requests.

Step 2: Find Cache Ownership

The router maintains information about which worker contains the relevant cache.

Conceptually:

Prefix Hash                  Worker

abc123                       GPU 1
def456                       GPU 3
xyz789                       GPU 2

The exact implementation varies by inference stack.

Step 3: Select the Worker

The router can combine cache locality with runtime conditions.

For example:

Candidate Worker   Cache Match   Load

GPU 1              80%           High
GPU 2              70%           Low
GPU 3               0%           Low

The router does not necessarily have to choose the worker with the highest cache match. It can balance cache reuse against current system load.

Step 4: Reuse the Cache

Once the request reaches the selected worker, the inference engine can reuse the available KV blocks.

Only the uncached portion needs to be processed.

KV Cache Routing vs Traditional Load Balancing

Traditional load balancing generally focuses on distributing requests across available workers.

For example:

Request 1 -> GPU 1
Request 2 -> GPU 2
Request 3 -> GPU 3
Request 4 -> GPU 1

This can keep GPU utilization relatively balanced.

However, it ignores computation already performed for a specific request prefix.

KV-aware routing adds another dimension.

Routing Strategy

Main Goal

Cache Awareness

Round robin

Even distribution

No

Least connections

Balance active requests

No

Least latency

Reduce observed latency

Usually no

KV-aware routing

Reuse cached computation

Yes

Hybrid routing

Balance cache reuse and load

Yes

For workloads with long repeated prefixes, cache locality can become an important routing signal.

Why Long Contexts Make This More Important

KV caching becomes particularly relevant as context windows grow.

Imagine an application that sends:

System Prompt: 2,000 tokens
Company Documentation: 10,000 tokens
Conversation History: 4,000 tokens
User Question: 50 tokens

If the system repeatedly processes similar requests, recomputing thousands of tokens can be wasteful.

With prefix reuse, much of the previous computation can potentially remain available.

The benefit depends on several factors:

  • Prefix length

  • Prefix reuse frequency

  • Cache hit rate

  • GPU memory availability

  • Model architecture

  • Request scheduling

  • Cache eviction policy

  • Number of inference workers

Therefore, simply enabling a KV cache does not automatically guarantee a large performance improvement.

Prefix Caching and KV Cache Routing

KV cache routing is closely related to prefix caching.

Prefix caching allows an inference system to reuse computation for a repeated prompt prefix.

For example:

Request A:
[System Prompt][Documentation][Question A]

Request B:
[System Prompt][Documentation][Question B]

The shared portion is:

[System Prompt][Documentation]

The inference engine may reuse the KV cache associated with that prefix.

Routing becomes important when the cached prefix exists on only one or a subset of workers.

Shared Prefix
     |
     v
+----+----+----+
|         |    |
GPU 1    GPU 2 GPU 3
  |       |     |
Cache    None   None

Sending the second request to GPU 1 can avoid recomputing the shared prefix.

A Simple Routing Algorithm

A conceptual implementation might look like this:

public Worker SelectWorker(
    Request request,
    IReadOnlyList<Worker> workers)
{
    var prefixHash = ComputePrefixHash(request);

    var candidates = workers
        .Where(worker => worker.HasCachedPrefix(prefixHash))
        .OrderBy(worker => worker.CurrentLoad)
        .ToList();

    if (candidates.Count > 0)
    {
        return candidates[0];
    }

    return workers
        .OrderBy(worker => worker.CurrentLoad)
        .First();
}

This is only a simplified example.

A production router would usually consider additional information such as:

  • Cache size

  • Cache age

  • Prefix match length

  • GPU memory pressure

  • Queue depth

  • Token generation rate

  • Network latency

  • Request priority

The important concept is that cache locality becomes part of the scheduling decision.

Partial Cache Matches

One useful optimization is to consider partial prefix matches.

Suppose the cache contains:

System Prompt
+
Company Documentation

but the incoming request contains:

System Prompt
+
Company Documentation
+
Conversation History
+
Question

The worker can still reuse the cached portion.

A router can therefore assign a score based on the amount of reusable context.

For example:

Worker       Cached Tokens     Queue Length

GPU 1        8,000             4
GPU 2        2,000             1
GPU 3            0             0

A simplistic load-only router chooses GPU 3.

A cache-aware router may choose GPU 1 if the computational savings outweigh the queueing cost.

This leads to a more realistic routing objective:

Routing Score =
Cache Benefit
-
Queueing Cost
-
Transfer Cost

The exact formula depends on the inference system.

When KV Cache Routing Works Well

KV-aware routing is particularly useful for workloads with repeated context.

Enterprise RAG Applications

An enterprise assistant may repeatedly send the same system instructions and retrieved organizational documentation.

For example:

System Instructions
+
Security Policy
+
Product Documentation
+
User Query

The shared prefix can create cache reuse opportunities.

Customer Support

Customer-support systems often reuse:

  • Agent instructions

  • Product documentation

  • Company policies

  • Response formatting rules

Only the customer-specific portion changes frequently.

Multi-Tenant Applications

A SaaS platform may have many requests associated with the same tenant.

If tenant-specific instructions or documentation are repeatedly included in prompts, routing requests toward workers containing those caches can improve locality.

Agentic Workloads

AI agents can generate multiple related model calls during a single workflow.

For example:

Agent Instructions
        +
Task Context
        +
Tool Results
        +
New Reasoning Step

Some parts of the context may remain stable across successive calls.

When It May Not Help

KV cache routing is not universally beneficial.

If every request has a completely different prompt, there may be little cache reuse.

For example:

Request A -> Unique 20,000-token prompt
Request B -> Unique 18,000-token prompt
Request C -> Unique 15,000-token prompt

In this scenario, traditional load balancing may be more useful.

There is also a cost associated with maintaining cache metadata and making routing decisions.

Another problem is cache eviction.

GPU memory is limited. If the cache becomes too large, older entries must be removed.

Therefore:

More caching
     !=
Unlimited performance improvement

The workload's reuse pattern matters.

KV Cache Routing and Distributed Systems

Distributed inference introduces another important issue: cache locality versus load distribution.

Imagine:

                    Router
                      |
        +-------------+-------------+
        |             |             |
      GPU 1         GPU 2         GPU 3
       90%           20%            30%
      Cache X        Cache Y        Cache Z

Suppose a request has a strong cache match on GPU 1.

Sending it there may maximize cache reuse but increase queueing.

Sending it to GPU 2 may provide faster scheduling but require recomputation.

A production router therefore needs a policy that considers both.

One conceptual approach is:

score =
    cache_match * cache_weight
    - queue_depth * load_weight
    - network_cost * network_weight

The weights should be tuned using actual workload measurements rather than arbitrary assumptions.

Metrics to Monitor

If implementing KV-aware routing, do not measure only average request latency.

Track metrics such as:

Metric

Why It Matters

KV cache hit rate

Measures cache effectiveness

Prefix reuse ratio

Shows how often context repeats

Time to first token

Measures prompt-processing impact

Inter-token latency

Measures generation performance

GPU memory usage

Shows cache pressure

Cache eviction rate

Indicates insufficient cache capacity

Queue depth

Shows worker contention

Tokens processed

Measures workload volume

Cache transfer volume

Identifies network overhead

Time to first token is particularly useful because prompt processing can have a noticeable effect on it.

Common Implementation Mistakes

Treating Cache Hits as the Only Routing Signal

A cache hit does not automatically mean the worker is the right choice.

A heavily overloaded worker may be slower than a lightly loaded worker without the cache.

Ignoring Cache Eviction

Caches are finite.

If the workload constantly replaces cached prefixes, the routing layer may create complexity without achieving meaningful reuse.

Over-Centralizing Cache Metadata

A single routing component can become a bottleneck if every request requires expensive global cache lookups.

Distributed or efficient metadata mechanisms may be necessary at larger scale.

Ignoring Request Affinity

Some workloads naturally benefit from session-level affinity.

For example:

User Session A -> GPU 2
User Session A -> GPU 2
User Session A -> GPU 2

Keeping related requests near their existing cache can be simpler than repeatedly rediscovering cache locality.

Measuring Only Throughput

Higher throughput does not automatically mean a better user experience.

Measure latency, especially:

Time to First Token

and end-to-end response latency.

Best Practices

For production LLM inference systems, consider these practices:

  1. Measure prefix reuse before implementing complex routing.

  2. Track KV cache hit rate explicitly.

  3. Combine cache locality with worker load.

  4. Monitor GPU memory pressure and eviction.

  5. Use stable prefix construction where possible.

  6. Avoid unnecessary changes to system prompts and shared context.

  7. Measure time to first token separately from generation latency.

  8. Account for network cost when cache data must move between workers.

  9. Use workload-specific routing policies instead of assuming one algorithm fits every application.

  10. Test cache-aware routing under realistic concurrent traffic.

Conclusion

KV cache routing changes the way LLM inference systems think about request scheduling.

Traditional routing asks:

Which worker can process this request?

KV-aware routing adds another question:

Which worker has already performed useful computation for this request?

That distinction becomes increasingly important for applications with long prompts, repeated instructions, RAG context, multi-turn conversations, and agent workflows.

The key idea is straightforward:

Repeated context
      |
      v
Reuse KV cache
      |
      v
Route toward cache locality
      |
      v
Reduce repeated computation

The real benefit depends on cache hit rate, prefix length, workload characteristics, GPU memory, queueing, and network overhead. For that reason, KV cache routing should be treated as an inference optimization to measure and tune—not as a universal replacement for conventional load balancing.

Summary

KV cache routing can improve LLM inference by directing requests toward workers that already contain useful attention-state data. When prompts share substantial prefixes, this can reduce repeated prompt computation and improve inference efficiency.

The most important production considerations are cache locality, load balancing, memory pressure, eviction, routing overhead, and measurable latency improvements. A hybrid strategy that considers both cache reuse and worker load is often more practical than relying on either signal alone.