Introduction
In the race to deploy Large Language Models (LLMs) at scale, enterprises face a persistent bottleneck: latency and cost. Traditional caching mechanisms, which rely on exact string matching, are ineffective for natural language interactions. A user asking "How do I reset my password?" and another asking "What are the steps to change my login credentials?" are semantically identical but lexically different. A standard cache would treat these as two unique requests, triggering expensive retrieval and generation processes twice. Semantic Caching solves this by leveraging vector embeddings. Instead of comparing raw text, it compares the mathematical representation of the query's meaning. If a new query falls within a certain cosine similarity threshold of a previously answered query, the system retrieves the stored response instantly. This reduces latency from seconds to milliseconds and cuts LLM API costs significantly.
However, implementing semantic caching in an enterprise environment requires more than just a vector store. It demands stateful orchestration to manage cache invalidation, integration with Retrieval-Augmented Generation (RAG) for freshness, and robust memory management. This article demonstrates how to build a production-ready semantic caching layer using LangGraph, Graph RAG, and FastAPI.
Real-Time Use Case: Enterprise HR Policy Assistant
Consider a global corporation with 50,000 employees. The HR department is overwhelmed with repetitive questions about leave policies, benefits, and compliance. An AI assistant is deployed to handle these queries.
Without semantic caching, every variation of a question about "maternity leave" triggers a full RAG pipeline: embedding the query, searching the vector database, retrieving policy documents, and generating a response via GPT-4. With semantic caching, the first employee’s question populates the cache. Subsequent employees asking similar questions in different words receive instant, consistent answers without re-invoking the LLM.
Architecture Overview
Semantic Cache Layer (Redis + Vector Search): Stores query embeddings and responses. Uses cosine similarity to identify hits.
LangGraph Orchestrator: Manages the workflow: Check Cache → If Miss, Retrieve via Graph RAG → Generate → Update Cache.
Graph RAG (Neo4j): Retrieves structured policy relationships (e.g.,
Employee-[:ELIGIBLE_FOR]->Benefit).State Management: Tracks cache hit/miss metrics and conversation history.
FastAPI & React: Provides the interface for employees and administrators.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
redis==5.0.0
numpy==1.26.0
Step 2: Defining the State Model
We define a state that tracks whether the response was served from the cache.
from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field
class HRState(TypedDict):
user_query: str
session_id: str
cached_response: Optional[str]
retrieved_policy_context: List[dict]
final_answer: str
is_cache_hit: bool
similarity_score: float
audit_log: List[str]
class HRQuery(BaseModel):
question: str = Field(..., description="Employee question")
employee_id: str = Field(..., description="Authenticated employee ID")
Step 3: The Semantic Cache Engine
We use Redis to store embeddings. For this POC, we simulate vector search using NumPy for clarity, but in production, you would use RediSearch or a dedicated vector DB.
import redis
import json
import numpy as np
from openai import OpenAI
client = OpenAI()
redis_client = redis.Redis(host='localhost', port=6379, db=0)
def get_embedding(text: str) -> list:
response = client.embeddings.create(input=text, model="text-embedding-3-small")
return response.data[0].embedding
def check_semantic_cache(query: str, threshold: float = 0.92) -> dict:
"""Checks if a semantically similar query exists in Redis"""
query_vec = np.array(get_embedding(query))
# In production, use Redis Vector Search (FT.SEARCH)
# Here we iterate keys for demonstration
keys = redis_client.keys("sem_cache:*")
best_match = None
highest_sim = 0
for key in keys:
data = json.loads(redis_client.get(key))
stored_vec = np.array(data['embedding'])
# Cosine Similarity
sim = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec) * np.linalg.norm(stored_vec))
if sim > highest_sim:
highest_sim = sim
if sim >= threshold:
best_match = data
if best_match:
return {"hit": True, "response": best_match['response'], "score": highest_sim}
return {"hit": False, "score": highest_sim}
def update_semantic_cache(query: str, response: str):
"""Stores the query embedding and response"""
key = f"sem_cache:{hash(query)}"
data = {
'query': query,
'response': response,
'embedding': get_embedding(query)
}
redis_client.setex(key, 86400, json.dumps(data)) # 24h TTL
Step 4: Graph RAG for Policy Retrieval
If the cache misses, we retrieve specific policy nodes from Neo4j.
from neo4j import GraphDatabase
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
def retrieve_policy_context(query: str) -> List[dict]:
with driver.session() as session:
# Graph query to find relevant policies
result = session.run("""
MATCH (p:Policy)-[:APPLIES_TO]->(dept:Department)
WHERE p.keywords CONTAINS $q
RETURN p.title, p.content
LIMIT 2
""", q=query)
return [record.data() for record in result]
Step 5: LangGraph Orchestration Workflow
The graph decides whether to bypass the LLM entirely based on the cache check.
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4-turbo")
def cache_check_node(state: HRState) -> HRState:
result = check_semantic_cache(state['user_query'])
state['is_cache_hit'] = result['hit']
state['similarity_score'] = result['score']
if result['hit']:
state['cached_response'] = result['response']
state['audit_log'].append(f"Cache Hit (Score: {result['score']:.2f})")
else:
state['audit_log'].append("Cache Miss - Initiating RAG")
return state
def rag_generation_node(state: HRState) -> HRState:
if state['is_cache_hit']:
state['final_answer'] = state['cached_response']
return state
context = retrieve_policy_context(state['user_query'])
prompt = f"Policies: {context}\nQuestion: {state['user_query']}"
response = llm.invoke(prompt).content
state['final_answer'] = response
update_semantic_cache(state['user_query'], response)
state['audit_log'].append("Generated new response and updated cache")
return state
# Build Graph
workflow = StateGraph(HRState)
workflow.add_node("check_cache", cache_check_node)
workflow.add_node("generate_rag", rag_generation_node)
workflow.set_entry_point("check_cache")
workflow.add_conditional_edges("check_cache",
lambda s: END if s['is_cache_hit'] else "generate_rag")
workflow.add_edge("generate_rag", END)
app = workflow.compile()
Step 6: FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
api_app = FastAPI(title="HR Semantic Cache API")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])
@api_app.post("/ask-hr")
async def ask_hr(query: HRQuery):
initial_state = HRState(
user_query=query.question,
session_id=query.employee_id,
cached_response=None,
retrieved_policy_context=[],
final_answer="",
is_cache_hit=False,
similarity_score=0.0,
audit_log=[]
)
result = await app.ainvoke(initial_state)
return {
"answer": result['final_answer'],
"source": "Cache" if result['is_cache_hit'] else "RAG",
"similarity_score": result['similarity_score'],
"audit": result['audit_log']
}
Step 7: React Frontend
import React, { useState } from 'react';
import axios from 'axios';
const HRChat = () => {
const [q, setQ] = useState('');
const [res, setRes] = useState(null);
const send = async () => {
const data = await axios.post('http://localhost:8000/ask-hr', {
question: q,
employee_id: 'EMP_123'
});
setRes(data.data);
};
return (
<div className="p-4 max-w-md mx-auto">
<h2 className="text-xl font-bold">HR Policy Assistant</h2>
<input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
<button onClick={send} className="bg-blue-600 text-white p-2 mt-2 rounded">Ask</button>
{res && (
<div className="mt-4 p-3 border rounded bg-gray-50">
<p>{res.answer}</p>
<div className="text-xs text-gray-500 mt-2">
Source: {res.source} | Similarity: {res.similarity_score.toFixed(2)}
</div>
</div>
)}
</div>
);
};
export default HRChat;
Conclusion
Semantic caching transforms LLM pipelines from costly, high-latency systems into responsive, enterprise-grade services. By understanding the intent behind a query rather than just its syntax, organizations can serve instant answers to repetitive questions while reserving expensive generative resources for novel, complex inquiries. Integrating this capability into a LangGraph workflow ensures that caching is not an afterthought but a core, stateful component of the AI architecture, balancing speed, cost, and accuracy effectively.

Join the conversation! Your thoughts help the community grow.