Introduction

Retrieval-Augmented Generation (RAG) has become the de facto architecture for enterprise LLM applications, yet production deployments routinely underperform expectations. The gap between a working demo and a reliable system is often the silent accumulation of failure modes subtle issues that degrade answer quality without triggering obvious errors. Unlike traditional software bugs, RAG failures are probabilistic, context-dependent, and often invisible until a user loses trust. Understanding the taxonomy of RAG failures is the first step toward building resilient systems. This article presents an enterprise-grade multi-agent LangGraph diagnostician that analyzes query-response pairs, classifies the specific failure mode, traces the root cause through the RAG pipeline, and recommends targeted remediations all while maintaining persistent memory of past failures to detect recurring patterns.

The Taxonomy of RAG Failure Modes

RAG failures can be categorized across four pipeline stages:

Stage

Failure Mode

Symptom

Query

Query misunderstanding

System answers a different question than asked

Query

Poor query decomposition

Multi-part questions answered incompletely

Retrieval

Low recall (missing context)

Relevant documents never retrieved

Retrieval

Low precision (noisy context)

Irrelevant chunks drown out signal

Retrieval

Semantic chunking failure

Meaning split across chunk boundaries

Retrieval

Stale knowledge

Outdated documents retrieved over current ones

Generation

Lost-in-the-middle

Information in middle of context ignored

Generation

Context window overflow

Critical info truncated

Generation

Hallucination despite context

Model fabricates facts not in sources

Generation

Conflicting information

Multiple sources disagree, model picks arbitrarily

Generation

Multi-hop reasoning failure

Requires combining facts across documents

Step-by-Step Implementation

1. Diagnostic State Schema with Memory

from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage
import redis
import json
from datetime import datetime

FailureStage = Literal["query", "retrieval", "generation", "integration"]
FailureType = Literal[
    "query_misunderstanding", "poor_decomposition",
    "low_recall", "low_precision", "chunking_failure", "stale_knowledge",
    "lost_in_middle", "context_overflow", "hallucination",
    "conflicting_info", "multi_hop_failure"
]

class RAGFailureState(TypedDict):
    messages: List
    conversation_id: str
    # Input: the problematic interaction
    original_query: str
    retrieved_chunks: List[Dict]
    generated_answer: str
    ground_truth: Optional[str]
    user_feedback: Optional[str]
    # Diagnostic outputs
    failure_detected: bool
    failure_stage: FailureStage
    failure_type: FailureType
    confidence: float
    root_cause_analysis: str
    pipeline_trace: Dict[str, Any]
    remediations: List[Dict]
    severity: Literal["low", "medium", "high", "critical"]
    # Memory
    historical_failures: List[Dict]
    recurring_patterns: List[str]

2. Retrieval Auditor Agent

from langchain_openai import ChatOpenAI

class RetrievalAuditorAgent:
    """Examines whether retrieval succeeded before generation is blamed"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def audit(self, state: RAGFailureState) -> RAGFailureState:
        chunks_text = "\n---\n".join(
            [f"[Chunk {i+1}] (score={c.get('score', 'N/A')}): {c['content'][:500]}"
             for i, c in enumerate(state["retrieved_chunks"])]
        )
        
        prompt = f"""You are a RAG retrieval auditor. Determine if the retrieved chunks
contain the information needed to answer the query.

Query: {state['original_query']}
{f'Ground Truth: {state["ground_truth"]}' if state.get('ground_truth') else ''}

Retrieved Chunks:
{chunks_text}

Answer JSON:
{{
    "retrieval_adequate": true/false,
    "relevant_chunks": [1, 2, ...],
    "missing_information": "what's missing",
    "noise_ratio": 0.0-1.0,
    "recalls_ground_truth": true/false/null
}}
"""
        response = self.llm.invoke(prompt)
        try:
            import re
            match = re.search(r'\{[\s\S]*\}', response.content)
            audit = json.loads(match.group()) if match else {}
        except Exception:
            audit = {"retrieval_adequate": False}
        
        state["pipeline_trace"] = state.get("pipeline_trace", {})
        state["pipeline_trace"]["retrieval_audit"] = audit
        return state

3. Failure Classifier Agent

class FailureClassifierAgent:
    """Classifies the specific failure mode using retrieved audit + LLM judgment"""
    
    FAILURE_SIGNATURES = {
        "query_misunderstanding": "Answer addresses a different question than asked",
        "poor_decomposition": "Multi-part query only partially answered",
        "low_recall": "Relevant documents exist but weren't retrieved",
        "low_precision": "Retrieved chunks are mostly irrelevant",
        "chunking_failure": "Key information split across chunks, neither complete",
        "stale_knowledge": "Outdated info retrieved, current info missed",
        "lost_in_middle": "Relevant info in middle of context ignored by generator",
        "context_overflow": "Context too long, critical info truncated",
        "hallucination": "Answer contains facts not present in any retrieved chunk",
        "conflicting_info": "Multiple chunks disagree, model picks arbitrarily",
        "multi_hop_failure": "Answer requires combining facts from multiple chunks"
    }
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def classify(self, state: RAGFailureState) -> RAGFailureState:
        retrieval_audit = state["pipeline_trace"].get("retrieval_audit", {})
        
        prompt = f"""Classify the RAG failure mode.

Query: {state['original_query']}
Generated Answer: {state['generated_answer']}
{f'Ground Truth: {state["ground_truth"]}' if state.get('ground_truth') else ''}
{f'User Feedback: {state["user_feedback"]}' if state.get('user_feedback') else ''}

Retrieval Audit:
- Retrieval Adequate: {retrieval_audit.get('retrieval_adequate')}
- Noise Ratio: {retrieval_audit.get('noise_ratio')}
- Missing Info: {retrieval_audit.get('missing_information')}

Failure Mode Signatures:
{chr(10).join([f'- {k}: {v}' for k, v in self.FAILURE_SIGNATURES.items()])}

Respond JSON:
{{
    "failure_detected": true/false,
    "failure_stage": "query|retrieval|generation|integration",
    "failure_type": "<one of the types above>",
    "confidence": 0.0-1.0,
    "severity": "low|medium|high|critical",
    "evidence": "brief explanation"
}}
"""
        response = self.llm.invoke(prompt)
        try:
            import re
            match = re.search(r'\{[\s\S]*\}', response.content)
            classification = json.loads(match.group()) if match else {}
        except Exception:
            classification = {"failure_detected": False}
        
        state["failure_detected"] = classification.get("failure_detected", False)
        state["failure_stage"] = classification.get("failure_stage", "generation")
        state["failure_type"] = classification.get("failure_type", "hallucination")
        state["confidence"] = classification.get("confidence", 0.5)
        state["severity"] = classification.get("severity", "medium")
        state["pipeline_trace"]["classification_evidence"] = classification.get("evidence", "")
        return state

4. Root Cause Analyzer Agent

class RootCauseAnalyzerAgent:
    """Traces failure to specific pipeline component"""
    
    def analyze(self, state: RAGFailureState) -> RAGFailureState:
        ft = state["failure_type"]
        audit = state["pipeline_trace"].get("retrieval_audit", {})
        
        causes = {
            "query_misunderstanding": "Embedding model misaligns query/document semantics; no query rewriting applied",
            "poor_decomposition": "Single-stage retrieval cannot handle compound queries; no query planner",
            "low_recall": f"Top-k too small or embedding model lacks domain specificity; noise ratio {audit.get('noise_ratio', 'N/A')}",
            "low_precision": "Embedding model captures surface similarity, not semantic relevance; no reranker",
            "chunking_failure": "Fixed-size chunking breaks semantic units; no semantic/recursive chunking",
            "stale_knowledge": "No document versioning or recency weighting in retrieval",
            "lost_in_middle": "Context too long for model's attention; no reranking to place key info first/last",
            "context_overflow": "No context budget management; chunks concatenated without priority",
            "hallucination": "Generator not constrained to context; no self-consistency or verification step",
            "conflicting_info": "No conflict detection; no source authority ranking",
            "multi_hop_failure": "Single-hop retrieval; no iterative/adaptive retrieval"
        }
        
        state["root_cause_analysis"] = causes.get(ft, "Unknown failure mode")
        return state

5. Remediation Recommender Agent

class RemediationRecommenderAgent:
    """Recommends concrete fixes based on failure type"""
    
    REMEDIATIONS = {
        "query_misunderstanding": [
            {"action": "Implement HyDE (Hypothetical Document Embeddings)", "effort": "medium"},
            {"action": "Add query rewriting with LLM-based reformulation", "effort": "low"},
            {"action": "Use multi-vector retriever with summary embeddings", "effort": "high"}
        ],
        "poor_decomposition": [
            {"action": "Add query planner agent to decompose compound queries", "effort": "medium"},
            {"action": "Implement multi-query retrieval with result fusion", "effort": "medium"}
        ],
        "low_recall": [
            {"action": "Increase top-k and add reranking", "effort": "low"},
            {"action": "Domain-adapt embedding model on corpus", "effort": "high"},
            {"action": "Add hybrid search (BM25 + vector)", "effort": "low"}
        ],
        "low_precision": [
            {"action": "Add Cohere/BGE reranker post-retrieval", "effort": "low"},
            {"action": "Implement cross-encoder reranking", "effort": "medium"}
        ],
        "chunking_failure": [
            {"action": "Switch to semantic/recursive chunking", "effort": "medium"},
            {"action": "Add parent-child chunk hierarchy", "effort": "medium"},
            {"action": "Use sentence-aware chunking with overlap", "effort": "low"}
        ],
        "stale_knowledge": [
            {"action": "Add document versioning with recency scoring", "effort": "medium"},
            {"action": "Implement TTL-based cache invalidation", "effort": "low"}
        ],
        "lost_in_middle": [
            {"action": "Rerank to place highest-relevance chunks at start/end", "effort": "low"},
            {"action": "Reduce context window to essential chunks only", "effort": "low"}
        ],
        "context_overflow": [
            {"action": "Implement context budget manager with priority scoring", "effort": "medium"},
            {"action": "Use map-reduce or refine patterns for long contexts", "effort": "medium"}
        ],
        "hallucination": [
            {"action": "Add LLM-as-judge verification step", "effort": "medium"},
            {"action": "Implement self-consistency with multiple generations", "effort": "high"},
            {"action": "Strengthen prompt constraints with citation requirements", "effort": "low"}
        ],
        "conflicting_info": [
            {"action": "Add source authority ranking", "effort": "medium"},
            {"action": "Implement conflict detection and user clarification", "effort": "high"}
        ],
        "multi_hop_failure": [
            {"action": "Implement iterative retrieval (Self-RAG / CRAG)", "effort": "high"},
            {"action": "Add knowledge graph for entity relationship traversal", "effort": "high"}
        ]
    }
    
    def recommend(self, state: RAGFailureState) -> RAGFailureState:
        state["remediations"] = self.REMEDIATIONS.get(state["failure_type"], [])
        return state

6. LangGraph Workflow with Memory

class FailureMemory:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def record_failure(self, state: RAGFailureState):
        record = {
            "timestamp": datetime.now().isoformat(),
            "query": state["original_query"][:200],
            "failure_type": state["failure_type"],
            "stage": state["failure_stage"],
            "severity": state["severity"],
            "confidence": state["confidence"]
        }
        key = f"failures:{state['conversation_id']}"
        history = json.loads(self.redis.get(key) or "[]")
        history.append(record)
        self.redis.set(key, json.dumps(history[-100:]))
    
    def load_history(self, conversation_id: str) -> List[Dict]:
        return json.loads(self.redis.get(f"failures:{conversation_id}") or "[]")
    
    def detect_patterns(self, history: List[Dict]) -> List[str]:
        from collections import Counter
        if not history:
            return []
        type_counts = Counter(h["failure_type"] for h in history[-50:])
        patterns = []
        for ft, count in type_counts.most_common(3):
            if count >= 3:
                patterns.append(f"Recurring {ft} ({count} times in last 50 failures)")
        return patterns

def build_failure_diagnostician():
    workflow = StateGraph(RAGFailureState)
    
    auditor = RetrievalAuditorAgent()
    classifier = FailureClassifierAgent()
    analyzer = RootCauseAnalyzerAgent()
    recommender = RemediationRecommenderAgent()
    
    workflow.add_node("audit_retrieval", auditor.audit)
    workflow.add_node("classify_failure", classifier.classify)
    workflow.add_node("analyze_root_cause", analyzer.analyze)
    workflow.add_node("recommend_remediation", recommender.recommend)
    
    workflow.set_entry_point("audit_retrieval")
    workflow.add_edge("audit_retrieval", "classify_failure")
    
    def route_after_classification(state: RAGFailureState):
        if not state["failure_detected"]:
            return "no_failure"
        return "analyze"
    
    workflow.add_conditional_edges(
        "classify_failure",
        route_after_classification,
        {"analyze": "analyze_root_cause", "no_failure": END}
    )
    workflow.add_edge("analyze_root_cause", "recommend_remediation")
    workflow.add_edge("recommend_remediation", END)
    
    return workflow.compile()

7. FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel

app = FastAPI(title="RAG Failure Diagnostician API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])

graph = build_failure_diagnostician()
memory = FailureMemory(redis.Redis())

class DiagnosticRequest(BaseModel):
    conversation_id: str
    original_query: str
    retrieved_chunks: List[Dict]
    generated_answer: str
    ground_truth: Optional[str] = None
    user_feedback: Optional[str] = None

@app.post("/diagnose")
async def diagnose(req: DiagnosticRequest):
    history = memory.load_history(req.conversation_id)
    
    initial_state = RAGFailureState(
        messages=[HumanMessage(content=f"Diagnose: {req.original_query}")],
        conversation_id=req.conversation_id,
        original_query=req.original_query,
        retrieved_chunks=req.retrieved_chunks,
        generated_answer=req.generated_answer,
        ground_truth=req.ground_truth,
        user_feedback=req.user_feedback,
        failure_detected=False,
        failure_stage="generation",
        failure_type="hallucination",
        confidence=0.0,
        root_cause_analysis="",
        pipeline_trace={},
        remediations=[],
        severity="medium",
        historical_failures=history,
        recurring_patterns=[]
    )
    
    result = graph.invoke(initial_state)
    
    if result["failure_detected"]:
        memory.record_failure(result)
        updated_history = memory.load_history(req.conversation_id)
        result["recurring_patterns"] = memory.detect_patterns(updated_history)
    
    return {
        "failure_detected": result["failure_detected"],
        "failure_stage": result["failure_stage"],
        "failure_type": result["failure_type"],
        "confidence": result["confidence"],
        "severity": result["severity"],
        "root_cause": result["root_cause_analysis"],
        "pipeline_trace": result["pipeline_trace"],
        "remediations": result["remediations"],
        "recurring_patterns": result.get("recurring_patterns", [])
    }

8. Frontend: Failure Analysis Dashboard

// components/FailureDiagnostician.tsx
import React, { useState } from 'react';

interface Diagnosis {
  failure_detected: boolean;
  failure_stage: string;
  failure_type: string;
  confidence: number;
  severity: string;
  root_cause: string;
  remediations: { action: string; effort: string }[];
  recurring_patterns: string[];
}

export const FailureDiagnostician: React.FC = () => {
  const [result, setResult] = useState<Diagnosis | null>(null);

  const runDiagnosis = async () => {
    const response = await fetch('http://localhost:8000/diagnose', {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({
        conversation_id: 'finance-compliance-001',
        original_query: "What is the maximum transaction limit for corporate accounts under the new 2024 AML policy?",
        retrieved_chunks: [
          { content: "The 2023 AML policy sets transaction limits at $10,000...", score: 0.82 },
          { content: "Customer onboarding procedures require KYC verification...", score: 0.71 }
        ],
        generated_answer: "The maximum transaction limit is $10,000 per the current AML policy.",
        ground_truth: "The 2024 AML policy increased the limit to $50,000 for corporate accounts.",
        user_feedback: "This is wrong - the policy was updated last quarter"
      })
    });
    setResult(await response.json());
  };

  const severityColor: Record<string, string> = {
    low: 'bg-yellow-100 text-yellow-800',
    medium: 'bg-orange-100 text-orange-800',
    high: 'bg-red-100 text-red-800',
    critical: 'bg-red-200 text-red-900'
  };

  return (
    <div className="p-6 max-w-5xl mx-auto bg-gray-50 min-h-screen">
      <h1 className="text-3xl font-bold mb-6">RAG Failure Diagnostician</h1>
      <button onClick={runDiagnosis} className="bg-blue-600 text-white px-6 py-2 rounded mb-6">
        Run Diagnosis
      </button>
      
      {result && (
        <div className="space-y-4">
          {!result.failure_detected ? (
            <div className="bg-green-100 text-green-800 p-4 rounded">
                No failure detected - RAG pipeline performed correctly
            </div>
          ) : (
            <>
              <div className={`${severityColor[result.severity]} p-6 rounded-lg shadow`}>
                <h2 className="text-2xl font-bold capitalize">
                  {result.failure_type.replace(/_/g, ' ')}
                </h2>
                <p className="mt-2">Stage: <span className="font-semibold">{result.failure_stage}</span></p>
                <p>Confidence: {(result.confidence * 100).toFixed(0)}% | Severity: {result.severity}</p>
              </div>

              <div className="bg-white p-4 rounded-lg shadow">
                <h3 className="font-bold mb-2">Root Cause Analysis</h3>
                <p className="text-gray-700">{result.root_cause}</p>
              </div>

              <div className="bg-white p-4 rounded-lg shadow">
                <h3 className="font-bold mb-2">Recommended Remediations</h3>
                <ul className="space-y-2">
                  {result.remediations.map((r, i) => (
                    <li key={i} className="flex justify-between items-center border-b pb-2">
                      <span>• {r.action}</span>
                      <span className={`text-xs px-2 py-1 rounded ${
                        r.effort === 'low' ? 'bg-green-100' :
                        r.effort === 'medium' ? 'bg-yellow-100' : 'bg-red-100'
                      }`}>{r.effort} effort</span>
                    </li>
                  ))}
                </ul>
              </div>

              {result.recurring_patterns.length > 0 && (
                <div className="bg-purple-50 border-l-4 border-purple-500 p-4 rounded">
                  <h3 className="font-bold mb-2">  Recurring Patterns Detected</h3>
                  <ul className="list-disc list-inside">
                    {result.recurring_patterns.map((p, i) => <li key={i}>{p}</li>)}
                  </ul>
                </div>
              )}
            </>
          )}
        </div>
      )}
    </div>
  );
};

Real-Time Use Case: Financial Services Compliance Q&A

A compliance officer at a major bank asks: "What is the maximum transaction limit for corporate accounts under the new 2024 AML policy?"

The RAG system retrieves two chunks:

  1. The 2023 AML policy (score 0.82) stating $10,000 limit

  2. An unrelated onboarding procedure doc (score 0.71)

The system answers: "The maximum transaction limit is $10,000 per the current AML policy."

The officer flags this as wrong—the 2024 policy increased the limit to $50,000.

The diagnostician analyzes:

  1. Retrieval Auditor detects: retrieval_adequate=false, missing_information="2024 policy document", noise_ratio=0.5

  2. Failure Classifier identifies: stale_knowledge failure at the retrieval stage, severity=critical (regulatory impact), confidence=0.92

  3. Root Cause Analyzer traces: "No document versioning or recency weighting in retrieval; 2023 document scored higher due to lexical overlap"

  4. Remediation Recommender suggests:

    • Add document versioning with recency scoring (medium effort)

    • Implement TTL-based cache invalidation (low effort)

The system also flags a recurring pattern: this is the 4th stale_knowledge failure in the past 50 interactions, prompting the team to prioritize versioning infrastructure.

Conclusion

RAG failures are not monolithic they are a diverse taxonomy of issues spanning query formulation, retrieval quality, and generation behavior. Treating all failures the same leads to generic "improve your prompts" advice that rarely addresses root causes. By building a multi-agent LangGraph diagnostician that audits retrieval, classifies failure modes, traces root causes, and recommends targeted fixes, enterprises gain a systematic approach to RAG reliability. The persistent memory layer transforms isolated incidents into pattern detection, enabling proactive infrastructure improvements rather than reactive firefighting. This diagnostician doesn't just identify what went wrong—it tells you exactly which component to fix and how, turning RAG from a fragile prototype into a production-grade enterprise system.