Introduction

In the race to deploy Large Language Models (LLMs) at scale, enterprises face a persistent bottleneck: latency and cost. Traditional caching mechanisms, which rely on exact string matching, are ineffective for natural language interactions. A user asking "How do I reset my password?" and another asking "What are the steps to change my login credentials?" are semantically identical but lexically different. A standard cache would treat these as two unique requests, triggering expensive retrieval and generation processes twice. Semantic Caching solves this by leveraging vector embeddings. Instead of comparing raw text, it compares the mathematical representation of the query's meaning. If a new query falls within a certain cosine similarity threshold of a previously answered query, the system retrieves the stored response instantly. This reduces latency from seconds to milliseconds and cuts LLM API costs significantly.

However, implementing semantic caching in an enterprise environment requires more than just a vector store. It demands stateful orchestration to manage cache invalidation, integration with Retrieval-Augmented Generation (RAG) for freshness, and robust memory management. This article demonstrates how to build a production-ready semantic caching layer using LangGraph, Graph RAG, and FastAPI.

Real-Time Use Case: Enterprise HR Policy Assistant

Consider a global corporation with 50,000 employees. The HR department is overwhelmed with repetitive questions about leave policies, benefits, and compliance. An AI assistant is deployed to handle these queries.

Without semantic caching, every variation of a question about "maternity leave" triggers a full RAG pipeline: embedding the query, searching the vector database, retrieving policy documents, and generating a response via GPT-4. With semantic caching, the first employee’s question populates the cache. Subsequent employees asking similar questions in different words receive instant, consistent answers without re-invoking the LLM.

Architecture Overview

  1. Semantic Cache Layer (Redis + Vector Search): Stores query embeddings and responses. Uses cosine similarity to identify hits.

  2. LangGraph Orchestrator: Manages the workflow: Check Cache → If Miss, Retrieve via Graph RAG → Generate → Update Cache.

  3. Graph RAG (Neo4j): Retrieves structured policy relationships (e.g., Employee-[:ELIGIBLE_FOR]->Benefit).

  4. State Management: Tracks cache hit/miss metrics and conversation history.

  5. FastAPI & React: Provides the interface for employees and administrators.

Step-by-Step Implementation

Step 1: Environment Setup

# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
redis==5.0.0
numpy==1.26.0

Step 2: Defining the State Model

We define a state that tracks whether the response was served from the cache.

from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field

class HRState(TypedDict):
    user_query: str
    session_id: str
    cached_response: Optional[str]
    retrieved_policy_context: List[dict]
    final_answer: str
    is_cache_hit: bool
    similarity_score: float
    audit_log: List[str]

class HRQuery(BaseModel):
    question: str = Field(..., description="Employee question")
    employee_id: str = Field(..., description="Authenticated employee ID")

Step 3: The Semantic Cache Engine

We use Redis to store embeddings. For this POC, we simulate vector search using NumPy for clarity, but in production, you would use RediSearch or a dedicated vector DB.

import redis
import json
import numpy as np
from openai import OpenAI

client = OpenAI()
redis_client = redis.Redis(host='localhost', port=6379, db=0)

def get_embedding(text: str) -> list:
    response = client.embeddings.create(input=text, model="text-embedding-3-small")
    return response.data[0].embedding

def check_semantic_cache(query: str, threshold: float = 0.92) -> dict:
    """Checks if a semantically similar query exists in Redis"""
    query_vec = np.array(get_embedding(query))
    
    # In production, use Redis Vector Search (FT.SEARCH)
    # Here we iterate keys for demonstration
    keys = redis_client.keys("sem_cache:*")
    best_match = None
    highest_sim = 0
    
    for key in keys:
        data = json.loads(redis_client.get(key))
        stored_vec = np.array(data['embedding'])
        
        # Cosine Similarity
        sim = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec) * np.linalg.norm(stored_vec))
        
        if sim > highest_sim:
            highest_sim = sim
            if sim >= threshold:
                best_match = data
                
    if best_match:
        return {"hit": True, "response": best_match['response'], "score": highest_sim}
    return {"hit": False, "score": highest_sim}

def update_semantic_cache(query: str, response: str):
    """Stores the query embedding and response"""
    key = f"sem_cache:{hash(query)}"
    data = {
        'query': query,
        'response': response,
        'embedding': get_embedding(query)
    }
    redis_client.setex(key, 86400, json.dumps(data)) # 24h TTL

Step 4: Graph RAG for Policy Retrieval

If the cache misses, we retrieve specific policy nodes from Neo4j.

from neo4j import GraphDatabase

driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

def retrieve_policy_context(query: str) -> List[dict]:
    with driver.session() as session:
        # Graph query to find relevant policies
        result = session.run("""
            MATCH (p:Policy)-[:APPLIES_TO]->(dept:Department)
            WHERE p.keywords CONTAINS $q
            RETURN p.title, p.content
            LIMIT 2
        """, q=query)
        return [record.data() for record in result]

Step 5: LangGraph Orchestration Workflow

The graph decides whether to bypass the LLM entirely based on the cache check.

from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4-turbo")

def cache_check_node(state: HRState) -> HRState:
    result = check_semantic_cache(state['user_query'])
    state['is_cache_hit'] = result['hit']
    state['similarity_score'] = result['score']
    
    if result['hit']:
        state['cached_response'] = result['response']
        state['audit_log'].append(f"Cache Hit (Score: {result['score']:.2f})")
    else:
        state['audit_log'].append("Cache Miss - Initiating RAG")
        
    return state

def rag_generation_node(state: HRState) -> HRState:
    if state['is_cache_hit']:
        state['final_answer'] = state['cached_response']
        return state
        
    context = retrieve_policy_context(state['user_query'])
    prompt = f"Policies: {context}\nQuestion: {state['user_query']}"
    response = llm.invoke(prompt).content
    
    state['final_answer'] = response
    update_semantic_cache(state['user_query'], response)
    state['audit_log'].append("Generated new response and updated cache")
    return state

# Build Graph
workflow = StateGraph(HRState)
workflow.add_node("check_cache", cache_check_node)
workflow.add_node("generate_rag", rag_generation_node)

workflow.set_entry_point("check_cache")
workflow.add_conditional_edges("check_cache", 
    lambda s: END if s['is_cache_hit'] else "generate_rag")
workflow.add_edge("generate_rag", END)

app = workflow.compile()

Step 6: FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

api_app = FastAPI(title="HR Semantic Cache API")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])

@api_app.post("/ask-hr")
async def ask_hr(query: HRQuery):
    initial_state = HRState(
        user_query=query.question,
        session_id=query.employee_id,
        cached_response=None,
        retrieved_policy_context=[],
        final_answer="",
        is_cache_hit=False,
        similarity_score=0.0,
        audit_log=[]
    )
    
    result = await app.ainvoke(initial_state)
    return {
        "answer": result['final_answer'],
        "source": "Cache" if result['is_cache_hit'] else "RAG",
        "similarity_score": result['similarity_score'],
        "audit": result['audit_log']
    }

Step 7: React Frontend

import React, { useState } from 'react';
import axios from 'axios';

const HRChat = () => {
  const [q, setQ] = useState('');
  const [res, setRes] = useState(null);

  const send = async () => {
    const data = await axios.post('http://localhost:8000/ask-hr', {
      question: q,
      employee_id: 'EMP_123'
    });
    setRes(data.data);
  };

  return (
    <div className="p-4 max-w-md mx-auto">
      <h2 className="text-xl font-bold">HR Policy Assistant</h2>
      <input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
      <button onClick={send} className="bg-blue-600 text-white p-2 mt-2 rounded">Ask</button>
      
      {res && (
        <div className="mt-4 p-3 border rounded bg-gray-50">
          <p>{res.answer}</p>
          <div className="text-xs text-gray-500 mt-2">
            Source: {res.source} | Similarity: {res.similarity_score.toFixed(2)}
          </div>
        </div>
      )}
    </div>
  );
};
export default HRChat;

Conclusion

Semantic caching transforms LLM pipelines from costly, high-latency systems into responsive, enterprise-grade services. By understanding the intent behind a query rather than just its syntax, organizations can serve instant answers to repetitive questions while reserving expensive generative resources for novel, complex inquiries. Integrating this capability into a LangGraph workflow ensures that caching is not an afterthought but a core, stateful component of the AI architecture, balancing speed, cost, and accuracy effectively.