Introduction
In the enterprise deployment of Large Language Models (LLMs), latency and cost are the twin enemies of scalability. Every token processed by an LLM consumes computational resources and time. For high-traffic applications, such as customer support bots or internal knowledge bases, sending every unique query to a foundational model is financially unsustainable and technically inefficient. Caching strategies offer a powerful solution by storing and reusing previous computations, thereby bypassing the need for redundant processing. However, caching in LLM pipelines is far more complex than traditional key-value storage. It requires handling semantic similarity (where different words mean the same thing), managing stateful conversations, and ensuring that cached data remains fresh relative to dynamic backend systems. A naive cache can lead to stale answers or security leaks.
This article explores a multi-layered caching architecture within an enterprise-grade AI system. We will build a Proof of Concept (POC) for a Technical Documentation Assistant using LangGraph for orchestration, Redis for high-speed semantic caching, Graph RAG for context retrieval, and FastAPI/React for the user interface. This approach demonstrates how to balance speed, cost, and accuracy in a production environment.
Real-Time Use Case: Technical Documentation Assistant
Consider a software company with thousands of developers accessing their API documentation daily. Developers often ask repetitive questions like "How do I authenticate via OAuth?" or "What is the rate limit for the v2 endpoint?"
A standard RAG system would:
Embed the query.
Search the vector database.
Retrieve chunks.
Send everything to the LLM for generation.
This process is slow and expensive for identical queries. Our optimized system will implement:
Semantic Cache: To catch questions that are phrased differently but have the same intent (e.g., "How to login?" vs. "OAuth authentication steps").
Context-Aware Memory: To cache intermediate retrieval results for follow-up questions within a session.
Stale-While-Revalidate: To serve cached answers immediately while updating the cache in the background if the underlying documentation has changed.
Architecture Overview
Semantic Cache (Redis + Vector Index): Stores embeddings of previous queries and their responses. Uses cosine similarity to find "close enough" matches.
LangGraph Orchestrator: Manages the decision logic: Check Cache -> If Miss, Retrieve & Generate -> Update Cache.
Graph RAG (Neo4j): Retrieves structured documentation nodes.
FastAPI Backend: Handles asynchronous requests and cache invalidation hooks.
React Frontend: Displays response times and cache hit status to demonstrate efficiency.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
redis==5.0.0
chromadb==0.4.22
Step 2: Defining the State with Cache Metrics
We track whether a response came from the cache to measure performance gains.
from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field
class CacheState(TypedDict):
user_query: str
session_id: str
cached_response: Optional[str]
retrieved_context: List[dict]
final_answer: str
is_cache_hit: bool
response_time_ms: float
audit_log: List[str]
class DocQuery(BaseModel):
question: str = Field(..., description="Developer's question")
session_id: str = Field(..., description="User session ID")
Step 3: Semantic Caching Engine with Redis
Instead of exact string matching, we use vector similarity to find semantically equivalent queries.
import redis
import json
import numpy as np
from openai import OpenAI
client = OpenAI()
redis_client = redis.Redis(host='localhost', port=6379, db=0)
def get_embedding(text: str) -> list:
response = client.embeddings.create(input=text, model="text-embedding-3-small")
return response.data[0].embedding
def check_semantic_cache(query: str, threshold: float = 0.9) -> Optional[str]:
"""Check Redis for similar queries using vector similarity"""
query_vec = get_embedding(query)
# In production, use Redis Vector Search (RediSearch)
# For this POC, we simulate a lookup by checking a few recent keys
keys = redis_client.keys("cache:*")
for key in keys:
stored_data = json.loads(redis_client.get(key))
stored_vec = stored_data['embedding']
# Calculate cosine similarity
similarity = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec) * np.linalg.norm(stored_vec))
if similarity > threshold:
return stored_data['response']
return None
def update_semantic_cache(query: str, response: str):
"""Store query embedding and response in Redis"""
key = f"cache:{hash(query)}"
data = {
'query': query,
'response': response,
'embedding': get_embedding(query)
}
redis_client.setex(key, 3600, json.dumps(data)) # 1 hour TTL
Step 4: Graph RAG Retrieval Node
If the cache misses, we retrieve fresh context from Neo4j.
from neo4j import GraphDatabase
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
def retrieve_docs(query: str) -> List[dict]:
with driver.session() as session:
result = session.run("""
MATCH (d:DocChunk)-[:PART_OF]->(api:APIEndpoint)
WHERE d.content CONTAINS $q
RETURN d.content LIMIT 3
""", q=query)
return [record['d.content'] for record in result]
Step 5: LangGraph Workflow with Caching Logic
The graph orchestrates the flow between cache checking and generation.
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
import time
llm = ChatOpenAI(model="gpt-4-turbo")
def cache_lookup_node(state: CacheState) -> CacheState:
start_time = time.time()
cached = check_semantic_cache(state['user_query'])
if cached:
state['cached_response'] = cached
state['is_cache_hit'] = True
state['audit_log'].append("Semantic Cache Hit")
else:
state['is_cache_hit'] = False
state['audit_log'].append("Cache Miss - Proceeding to RAG")
state['response_time_ms'] = (time.time() - start_time) * 1000
return state
def rag_generation_node(state: CacheState) -> CacheState:
if state['is_cache_hit']:
state['final_answer'] = state['cached_response']
return state
context = retrieve_docs(state['user_query'])
prompt = f"Context: {context}\nQuestion: {state['user_query']}"
response = llm.invoke(prompt).content
state['final_answer'] = response
state['retrieved_context'] = context
update_semantic_cache(state['user_query'], response)
state['audit_log'].append("Generated and Cached new response")
return state
# Build Graph
workflow = StateGraph(CacheState)
workflow.add_node("lookup", cache_lookup_node)
workflow.add_node("generate", rag_generation_node)
workflow.set_entry_point("lookup")
workflow.add_conditional_edges("lookup",
lambda s: END if s['is_cache_hit'] else "generate")
workflow.add_edge("generate", END)
app = workflow.compile()
Step 6: FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
api_app = FastAPI(title="Cached LLM Pipeline")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])
@api_app.post("/ask-docs")
async def ask_docs(query: DocQuery):
initial_state = CacheState(
user_query=query.question,
session_id=query.session_id,
cached_response=None,
retrieved_context=[],
final_answer="",
is_cache_hit=False,
response_time_ms=0.0,
audit_log=[]
)
result = await app.ainvoke(initial_state)
return {
"answer": result['final_answer'],
"metrics": {
"cache_hit": result['is_cache_hit'],
"latency_ms": result['response_time_ms']
},
"log": result['audit_log']
}
Step 7: React Frontend
import React, { useState } from 'react';
import axios from 'axios';
const DocAssistant = () => {
const [q, setQ] = useState('');
const [res, setRes] = useState(null);
const send = async () => {
const data = await axios.post('http://localhost:8000/ask-docs', {
question: q,
session_id: 'DEV_SESSION_1'
});
setRes(data.data);
};
return (
<div className="p-4 max-w-lg mx-auto">
<h2 className="text-xl font-bold">Tech Docs Assistant</h2>
<input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
<button onClick={send} className="bg-blue-500 text-white p-2 mt-2">Ask</button>
{res && (
<div className="mt-4 p-3 border rounded">
<p>{res.answer}</p>
<div className="text-xs text-gray-500 mt-2">
<span className={`px-2 py-1 rounded ${res.metrics.cache_hit ? 'bg-green-100' : 'bg-yellow-100'}`}>
{res.metrics.cache_hit ? 'Cache Hit' : 'Cache Miss'}
</span>
<span className="ml-2">{res.metrics.latency_ms.toFixed(2)}ms</span>
</div>
</div>
)}
</div>
);
};
export default DocAssistant;
Conclusion
Caching in LLM pipelines is not a one-size-fits-all solution. By implementing a multi-layered strategy that combines semantic vector caching with traditional state management, enterprises can drastically reduce latency and operational costs. The integration of LangGraph allows for flexible orchestration where cache hits bypass expensive retrieval and generation steps entirely, while Graph RAG ensures that when a cache miss occurs, the retrieved context is precise and structured. This architecture provides the scalability needed for high-volume enterprise applications, ensuring that AI remains both responsive and economically viable.

Join the conversation! Your thoughts help the community grow.