Introduction

In the enterprise deployment of Large Language Models (LLMs), latency and cost are the twin enemies of scalability. Every token processed by an LLM consumes computational resources and time. For high-traffic applications, such as customer support bots or internal knowledge bases, sending every unique query to a foundational model is financially unsustainable and technically inefficient. Caching strategies offer a powerful solution by storing and reusing previous computations, thereby bypassing the need for redundant processing. However, caching in LLM pipelines is far more complex than traditional key-value storage. It requires handling semantic similarity (where different words mean the same thing), managing stateful conversations, and ensuring that cached data remains fresh relative to dynamic backend systems. A naive cache can lead to stale answers or security leaks.

This article explores a multi-layered caching architecture within an enterprise-grade AI system. We will build a Proof of Concept (POC) for a Technical Documentation Assistant using LangGraph for orchestration, Redis for high-speed semantic caching, Graph RAG for context retrieval, and FastAPI/React for the user interface. This approach demonstrates how to balance speed, cost, and accuracy in a production environment.

Real-Time Use Case: Technical Documentation Assistant

Consider a software company with thousands of developers accessing their API documentation daily. Developers often ask repetitive questions like "How do I authenticate via OAuth?" or "What is the rate limit for the v2 endpoint?"

A standard RAG system would:

  1. Embed the query.

  2. Search the vector database.

  3. Retrieve chunks.

  4. Send everything to the LLM for generation.

This process is slow and expensive for identical queries. Our optimized system will implement:

  1. Semantic Cache: To catch questions that are phrased differently but have the same intent (e.g., "How to login?" vs. "OAuth authentication steps").

  2. Context-Aware Memory: To cache intermediate retrieval results for follow-up questions within a session.

  3. Stale-While-Revalidate: To serve cached answers immediately while updating the cache in the background if the underlying documentation has changed.

Architecture Overview

  • Semantic Cache (Redis + Vector Index): Stores embeddings of previous queries and their responses. Uses cosine similarity to find "close enough" matches.

  • LangGraph Orchestrator: Manages the decision logic: Check Cache -> If Miss, Retrieve & Generate -> Update Cache.

  • Graph RAG (Neo4j): Retrieves structured documentation nodes.

  • FastAPI Backend: Handles asynchronous requests and cache invalidation hooks.

  • React Frontend: Displays response times and cache hit status to demonstrate efficiency.

Step-by-Step Implementation

Step 1: Environment Setup

# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
redis==5.0.0
chromadb==0.4.22

Step 2: Defining the State with Cache Metrics

We track whether a response came from the cache to measure performance gains.

from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field

class CacheState(TypedDict):
    user_query: str
    session_id: str
    cached_response: Optional[str]
    retrieved_context: List[dict]
    final_answer: str
    is_cache_hit: bool
    response_time_ms: float
    audit_log: List[str]

class DocQuery(BaseModel):
    question: str = Field(..., description="Developer's question")
    session_id: str = Field(..., description="User session ID")

Step 3: Semantic Caching Engine with Redis

Instead of exact string matching, we use vector similarity to find semantically equivalent queries.

import redis
import json
import numpy as np
from openai import OpenAI

client = OpenAI()
redis_client = redis.Redis(host='localhost', port=6379, db=0)

def get_embedding(text: str) -> list:
    response = client.embeddings.create(input=text, model="text-embedding-3-small")
    return response.data[0].embedding

def check_semantic_cache(query: str, threshold: float = 0.9) -> Optional[str]:
    """Check Redis for similar queries using vector similarity"""
    query_vec = get_embedding(query)
    
    # In production, use Redis Vector Search (RediSearch)
    # For this POC, we simulate a lookup by checking a few recent keys
    keys = redis_client.keys("cache:*")
    for key in keys:
        stored_data = json.loads(redis_client.get(key))
        stored_vec = stored_data['embedding']
        
        # Calculate cosine similarity
        similarity = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec) * np.linalg.norm(stored_vec))
        
        if similarity > threshold:
            return stored_data['response']
    return None

def update_semantic_cache(query: str, response: str):
    """Store query embedding and response in Redis"""
    key = f"cache:{hash(query)}"
    data = {
        'query': query,
        'response': response,
        'embedding': get_embedding(query)
    }
    redis_client.setex(key, 3600, json.dumps(data)) # 1 hour TTL

Step 4: Graph RAG Retrieval Node

If the cache misses, we retrieve fresh context from Neo4j.

from neo4j import GraphDatabase

driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

def retrieve_docs(query: str) -> List[dict]:
    with driver.session() as session:
        result = session.run("""
            MATCH (d:DocChunk)-[:PART_OF]->(api:APIEndpoint)
            WHERE d.content CONTAINS $q
            RETURN d.content LIMIT 3
        """, q=query)
        return [record['d.content'] for record in result]

Step 5: LangGraph Workflow with Caching Logic

The graph orchestrates the flow between cache checking and generation.

from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
import time

llm = ChatOpenAI(model="gpt-4-turbo")

def cache_lookup_node(state: CacheState) -> CacheState:
    start_time = time.time()
    cached = check_semantic_cache(state['user_query'])
    
    if cached:
        state['cached_response'] = cached
        state['is_cache_hit'] = True
        state['audit_log'].append("Semantic Cache Hit")
    else:
        state['is_cache_hit'] = False
        state['audit_log'].append("Cache Miss - Proceeding to RAG")
        
    state['response_time_ms'] = (time.time() - start_time) * 1000
    return state

def rag_generation_node(state: CacheState) -> CacheState:
    if state['is_cache_hit']:
        state['final_answer'] = state['cached_response']
        return state
        
    context = retrieve_docs(state['user_query'])
    prompt = f"Context: {context}\nQuestion: {state['user_query']}"
    response = llm.invoke(prompt).content
    
    state['final_answer'] = response
    state['retrieved_context'] = context
    update_semantic_cache(state['user_query'], response)
    state['audit_log'].append("Generated and Cached new response")
    return state

# Build Graph
workflow = StateGraph(CacheState)
workflow.add_node("lookup", cache_lookup_node)
workflow.add_node("generate", rag_generation_node)

workflow.set_entry_point("lookup")
workflow.add_conditional_edges("lookup", 
    lambda s: END if s['is_cache_hit'] else "generate")
workflow.add_edge("generate", END)

app = workflow.compile()

Step 6: FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

api_app = FastAPI(title="Cached LLM Pipeline")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])

@api_app.post("/ask-docs")
async def ask_docs(query: DocQuery):
    initial_state = CacheState(
        user_query=query.question,
        session_id=query.session_id,
        cached_response=None,
        retrieved_context=[],
        final_answer="",
        is_cache_hit=False,
        response_time_ms=0.0,
        audit_log=[]
    )
    
    result = await app.ainvoke(initial_state)
    return {
        "answer": result['final_answer'],
        "metrics": {
            "cache_hit": result['is_cache_hit'],
            "latency_ms": result['response_time_ms']
        },
        "log": result['audit_log']
    }

Step 7: React Frontend

import React, { useState } from 'react';
import axios from 'axios';

const DocAssistant = () => {
  const [q, setQ] = useState('');
  const [res, setRes] = useState(null);

  const send = async () => {
    const data = await axios.post('http://localhost:8000/ask-docs', {
      question: q,
      session_id: 'DEV_SESSION_1'
    });
    setRes(data.data);
  };

  return (
    <div className="p-4 max-w-lg mx-auto">
      <h2 className="text-xl font-bold">Tech Docs Assistant</h2>
      <input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
      <button onClick={send} className="bg-blue-500 text-white p-2 mt-2">Ask</button>
      
      {res && (
        <div className="mt-4 p-3 border rounded">
          <p>{res.answer}</p>
          <div className="text-xs text-gray-500 mt-2">
            <span className={`px-2 py-1 rounded ${res.metrics.cache_hit ? 'bg-green-100' : 'bg-yellow-100'}`}>
              {res.metrics.cache_hit ? 'Cache Hit' : 'Cache Miss'}
            </span>
            <span className="ml-2">{res.metrics.latency_ms.toFixed(2)}ms</span>
          </div>
        </div>
      )}
    </div>
  );
};
export default DocAssistant;

Conclusion

Caching in LLM pipelines is not a one-size-fits-all solution. By implementing a multi-layered strategy that combines semantic vector caching with traditional state management, enterprises can drastically reduce latency and operational costs. The integration of LangGraph allows for flexible orchestration where cache hits bypass expensive retrieval and generation steps entirely, while Graph RAG ensures that when a cache miss occurs, the retrieved context is precise and structured. This architecture provides the scalability needed for high-volume enterprise applications, ensuring that AI remains both responsive and economically viable.