Introduction
Retrieval-Augmented Generation (RAG) has become the de facto architecture for enterprise LLM applications, yet production deployments routinely underperform expectations. The gap between a working demo and a reliable system is often the silent accumulation of failure modes subtle issues that degrade answer quality without triggering obvious errors. Unlike traditional software bugs, RAG failures are probabilistic, context-dependent, and often invisible until a user loses trust. Understanding the taxonomy of RAG failures is the first step toward building resilient systems. This article presents an enterprise-grade multi-agent LangGraph diagnostician that analyzes query-response pairs, classifies the specific failure mode, traces the root cause through the RAG pipeline, and recommends targeted remediations all while maintaining persistent memory of past failures to detect recurring patterns.
The Taxonomy of RAG Failure Modes
RAG failures can be categorized across four pipeline stages:
Stage | Failure Mode | Symptom |
|---|---|---|
Query | Query misunderstanding | System answers a different question than asked |
Query | Poor query decomposition | Multi-part questions answered incompletely |
Retrieval | Low recall (missing context) | Relevant documents never retrieved |
Retrieval | Low precision (noisy context) | Irrelevant chunks drown out signal |
Retrieval | Semantic chunking failure | Meaning split across chunk boundaries |
Retrieval | Stale knowledge | Outdated documents retrieved over current ones |
Generation | Lost-in-the-middle | Information in middle of context ignored |
Generation | Context window overflow | Critical info truncated |
Generation | Hallucination despite context | Model fabricates facts not in sources |
Generation | Conflicting information | Multiple sources disagree, model picks arbitrarily |
Generation | Multi-hop reasoning failure | Requires combining facts across documents |
Step-by-Step Implementation
1. Diagnostic State Schema with Memory
from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage
import redis
import json
from datetime import datetime
FailureStage = Literal["query", "retrieval", "generation", "integration"]
FailureType = Literal[
"query_misunderstanding", "poor_decomposition",
"low_recall", "low_precision", "chunking_failure", "stale_knowledge",
"lost_in_middle", "context_overflow", "hallucination",
"conflicting_info", "multi_hop_failure"
]
class RAGFailureState(TypedDict):
messages: List
conversation_id: str
# Input: the problematic interaction
original_query: str
retrieved_chunks: List[Dict]
generated_answer: str
ground_truth: Optional[str]
user_feedback: Optional[str]
# Diagnostic outputs
failure_detected: bool
failure_stage: FailureStage
failure_type: FailureType
confidence: float
root_cause_analysis: str
pipeline_trace: Dict[str, Any]
remediations: List[Dict]
severity: Literal["low", "medium", "high", "critical"]
# Memory
historical_failures: List[Dict]
recurring_patterns: List[str]
2. Retrieval Auditor Agent
from langchain_openai import ChatOpenAI
class RetrievalAuditorAgent:
"""Examines whether retrieval succeeded before generation is blamed"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def audit(self, state: RAGFailureState) -> RAGFailureState:
chunks_text = "\n---\n".join(
[f"[Chunk {i+1}] (score={c.get('score', 'N/A')}): {c['content'][:500]}"
for i, c in enumerate(state["retrieved_chunks"])]
)
prompt = f"""You are a RAG retrieval auditor. Determine if the retrieved chunks
contain the information needed to answer the query.
Query: {state['original_query']}
{f'Ground Truth: {state["ground_truth"]}' if state.get('ground_truth') else ''}
Retrieved Chunks:
{chunks_text}
Answer JSON:
{{
"retrieval_adequate": true/false,
"relevant_chunks": [1, 2, ...],
"missing_information": "what's missing",
"noise_ratio": 0.0-1.0,
"recalls_ground_truth": true/false/null
}}
"""
response = self.llm.invoke(prompt)
try:
import re
match = re.search(r'\{[\s\S]*\}', response.content)
audit = json.loads(match.group()) if match else {}
except Exception:
audit = {"retrieval_adequate": False}
state["pipeline_trace"] = state.get("pipeline_trace", {})
state["pipeline_trace"]["retrieval_audit"] = audit
return state
3. Failure Classifier Agent
class FailureClassifierAgent:
"""Classifies the specific failure mode using retrieved audit + LLM judgment"""
FAILURE_SIGNATURES = {
"query_misunderstanding": "Answer addresses a different question than asked",
"poor_decomposition": "Multi-part query only partially answered",
"low_recall": "Relevant documents exist but weren't retrieved",
"low_precision": "Retrieved chunks are mostly irrelevant",
"chunking_failure": "Key information split across chunks, neither complete",
"stale_knowledge": "Outdated info retrieved, current info missed",
"lost_in_middle": "Relevant info in middle of context ignored by generator",
"context_overflow": "Context too long, critical info truncated",
"hallucination": "Answer contains facts not present in any retrieved chunk",
"conflicting_info": "Multiple chunks disagree, model picks arbitrarily",
"multi_hop_failure": "Answer requires combining facts from multiple chunks"
}
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def classify(self, state: RAGFailureState) -> RAGFailureState:
retrieval_audit = state["pipeline_trace"].get("retrieval_audit", {})
prompt = f"""Classify the RAG failure mode.
Query: {state['original_query']}
Generated Answer: {state['generated_answer']}
{f'Ground Truth: {state["ground_truth"]}' if state.get('ground_truth') else ''}
{f'User Feedback: {state["user_feedback"]}' if state.get('user_feedback') else ''}
Retrieval Audit:
- Retrieval Adequate: {retrieval_audit.get('retrieval_adequate')}
- Noise Ratio: {retrieval_audit.get('noise_ratio')}
- Missing Info: {retrieval_audit.get('missing_information')}
Failure Mode Signatures:
{chr(10).join([f'- {k}: {v}' for k, v in self.FAILURE_SIGNATURES.items()])}
Respond JSON:
{{
"failure_detected": true/false,
"failure_stage": "query|retrieval|generation|integration",
"failure_type": "<one of the types above>",
"confidence": 0.0-1.0,
"severity": "low|medium|high|critical",
"evidence": "brief explanation"
}}
"""
response = self.llm.invoke(prompt)
try:
import re
match = re.search(r'\{[\s\S]*\}', response.content)
classification = json.loads(match.group()) if match else {}
except Exception:
classification = {"failure_detected": False}
state["failure_detected"] = classification.get("failure_detected", False)
state["failure_stage"] = classification.get("failure_stage", "generation")
state["failure_type"] = classification.get("failure_type", "hallucination")
state["confidence"] = classification.get("confidence", 0.5)
state["severity"] = classification.get("severity", "medium")
state["pipeline_trace"]["classification_evidence"] = classification.get("evidence", "")
return state
4. Root Cause Analyzer Agent
class RootCauseAnalyzerAgent:
"""Traces failure to specific pipeline component"""
def analyze(self, state: RAGFailureState) -> RAGFailureState:
ft = state["failure_type"]
audit = state["pipeline_trace"].get("retrieval_audit", {})
causes = {
"query_misunderstanding": "Embedding model misaligns query/document semantics; no query rewriting applied",
"poor_decomposition": "Single-stage retrieval cannot handle compound queries; no query planner",
"low_recall": f"Top-k too small or embedding model lacks domain specificity; noise ratio {audit.get('noise_ratio', 'N/A')}",
"low_precision": "Embedding model captures surface similarity, not semantic relevance; no reranker",
"chunking_failure": "Fixed-size chunking breaks semantic units; no semantic/recursive chunking",
"stale_knowledge": "No document versioning or recency weighting in retrieval",
"lost_in_middle": "Context too long for model's attention; no reranking to place key info first/last",
"context_overflow": "No context budget management; chunks concatenated without priority",
"hallucination": "Generator not constrained to context; no self-consistency or verification step",
"conflicting_info": "No conflict detection; no source authority ranking",
"multi_hop_failure": "Single-hop retrieval; no iterative/adaptive retrieval"
}
state["root_cause_analysis"] = causes.get(ft, "Unknown failure mode")
return state
5. Remediation Recommender Agent
class RemediationRecommenderAgent:
"""Recommends concrete fixes based on failure type"""
REMEDIATIONS = {
"query_misunderstanding": [
{"action": "Implement HyDE (Hypothetical Document Embeddings)", "effort": "medium"},
{"action": "Add query rewriting with LLM-based reformulation", "effort": "low"},
{"action": "Use multi-vector retriever with summary embeddings", "effort": "high"}
],
"poor_decomposition": [
{"action": "Add query planner agent to decompose compound queries", "effort": "medium"},
{"action": "Implement multi-query retrieval with result fusion", "effort": "medium"}
],
"low_recall": [
{"action": "Increase top-k and add reranking", "effort": "low"},
{"action": "Domain-adapt embedding model on corpus", "effort": "high"},
{"action": "Add hybrid search (BM25 + vector)", "effort": "low"}
],
"low_precision": [
{"action": "Add Cohere/BGE reranker post-retrieval", "effort": "low"},
{"action": "Implement cross-encoder reranking", "effort": "medium"}
],
"chunking_failure": [
{"action": "Switch to semantic/recursive chunking", "effort": "medium"},
{"action": "Add parent-child chunk hierarchy", "effort": "medium"},
{"action": "Use sentence-aware chunking with overlap", "effort": "low"}
],
"stale_knowledge": [
{"action": "Add document versioning with recency scoring", "effort": "medium"},
{"action": "Implement TTL-based cache invalidation", "effort": "low"}
],
"lost_in_middle": [
{"action": "Rerank to place highest-relevance chunks at start/end", "effort": "low"},
{"action": "Reduce context window to essential chunks only", "effort": "low"}
],
"context_overflow": [
{"action": "Implement context budget manager with priority scoring", "effort": "medium"},
{"action": "Use map-reduce or refine patterns for long contexts", "effort": "medium"}
],
"hallucination": [
{"action": "Add LLM-as-judge verification step", "effort": "medium"},
{"action": "Implement self-consistency with multiple generations", "effort": "high"},
{"action": "Strengthen prompt constraints with citation requirements", "effort": "low"}
],
"conflicting_info": [
{"action": "Add source authority ranking", "effort": "medium"},
{"action": "Implement conflict detection and user clarification", "effort": "high"}
],
"multi_hop_failure": [
{"action": "Implement iterative retrieval (Self-RAG / CRAG)", "effort": "high"},
{"action": "Add knowledge graph for entity relationship traversal", "effort": "high"}
]
}
def recommend(self, state: RAGFailureState) -> RAGFailureState:
state["remediations"] = self.REMEDIATIONS.get(state["failure_type"], [])
return state
6. LangGraph Workflow with Memory
class FailureMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def record_failure(self, state: RAGFailureState):
record = {
"timestamp": datetime.now().isoformat(),
"query": state["original_query"][:200],
"failure_type": state["failure_type"],
"stage": state["failure_stage"],
"severity": state["severity"],
"confidence": state["confidence"]
}
key = f"failures:{state['conversation_id']}"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-100:]))
def load_history(self, conversation_id: str) -> List[Dict]:
return json.loads(self.redis.get(f"failures:{conversation_id}") or "[]")
def detect_patterns(self, history: List[Dict]) -> List[str]:
from collections import Counter
if not history:
return []
type_counts = Counter(h["failure_type"] for h in history[-50:])
patterns = []
for ft, count in type_counts.most_common(3):
if count >= 3:
patterns.append(f"Recurring {ft} ({count} times in last 50 failures)")
return patterns
def build_failure_diagnostician():
workflow = StateGraph(RAGFailureState)
auditor = RetrievalAuditorAgent()
classifier = FailureClassifierAgent()
analyzer = RootCauseAnalyzerAgent()
recommender = RemediationRecommenderAgent()
workflow.add_node("audit_retrieval", auditor.audit)
workflow.add_node("classify_failure", classifier.classify)
workflow.add_node("analyze_root_cause", analyzer.analyze)
workflow.add_node("recommend_remediation", recommender.recommend)
workflow.set_entry_point("audit_retrieval")
workflow.add_edge("audit_retrieval", "classify_failure")
def route_after_classification(state: RAGFailureState):
if not state["failure_detected"]:
return "no_failure"
return "analyze"
workflow.add_conditional_edges(
"classify_failure",
route_after_classification,
{"analyze": "analyze_root_cause", "no_failure": END}
)
workflow.add_edge("analyze_root_cause", "recommend_remediation")
workflow.add_edge("recommend_remediation", END)
return workflow.compile()
7. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="RAG Failure Diagnostician API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_failure_diagnostician()
memory = FailureMemory(redis.Redis())
class DiagnosticRequest(BaseModel):
conversation_id: str
original_query: str
retrieved_chunks: List[Dict]
generated_answer: str
ground_truth: Optional[str] = None
user_feedback: Optional[str] = None
@app.post("/diagnose")
async def diagnose(req: DiagnosticRequest):
history = memory.load_history(req.conversation_id)
initial_state = RAGFailureState(
messages=[HumanMessage(content=f"Diagnose: {req.original_query}")],
conversation_id=req.conversation_id,
original_query=req.original_query,
retrieved_chunks=req.retrieved_chunks,
generated_answer=req.generated_answer,
ground_truth=req.ground_truth,
user_feedback=req.user_feedback,
failure_detected=False,
failure_stage="generation",
failure_type="hallucination",
confidence=0.0,
root_cause_analysis="",
pipeline_trace={},
remediations=[],
severity="medium",
historical_failures=history,
recurring_patterns=[]
)
result = graph.invoke(initial_state)
if result["failure_detected"]:
memory.record_failure(result)
updated_history = memory.load_history(req.conversation_id)
result["recurring_patterns"] = memory.detect_patterns(updated_history)
return {
"failure_detected": result["failure_detected"],
"failure_stage": result["failure_stage"],
"failure_type": result["failure_type"],
"confidence": result["confidence"],
"severity": result["severity"],
"root_cause": result["root_cause_analysis"],
"pipeline_trace": result["pipeline_trace"],
"remediations": result["remediations"],
"recurring_patterns": result.get("recurring_patterns", [])
}
8. Frontend: Failure Analysis Dashboard
// components/FailureDiagnostician.tsx
import React, { useState } from 'react';
interface Diagnosis {
failure_detected: boolean;
failure_stage: string;
failure_type: string;
confidence: number;
severity: string;
root_cause: string;
remediations: { action: string; effort: string }[];
recurring_patterns: string[];
}
export const FailureDiagnostician: React.FC = () => {
const [result, setResult] = useState<Diagnosis | null>(null);
const runDiagnosis = async () => {
const response = await fetch('http://localhost:8000/diagnose', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
conversation_id: 'finance-compliance-001',
original_query: "What is the maximum transaction limit for corporate accounts under the new 2024 AML policy?",
retrieved_chunks: [
{ content: "The 2023 AML policy sets transaction limits at $10,000...", score: 0.82 },
{ content: "Customer onboarding procedures require KYC verification...", score: 0.71 }
],
generated_answer: "The maximum transaction limit is $10,000 per the current AML policy.",
ground_truth: "The 2024 AML policy increased the limit to $50,000 for corporate accounts.",
user_feedback: "This is wrong - the policy was updated last quarter"
})
});
setResult(await response.json());
};
const severityColor: Record<string, string> = {
low: 'bg-yellow-100 text-yellow-800',
medium: 'bg-orange-100 text-orange-800',
high: 'bg-red-100 text-red-800',
critical: 'bg-red-200 text-red-900'
};
return (
<div className="p-6 max-w-5xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-6">RAG Failure Diagnostician</h1>
<button onClick={runDiagnosis} className="bg-blue-600 text-white px-6 py-2 rounded mb-6">
Run Diagnosis
</button>
{result && (
<div className="space-y-4">
{!result.failure_detected ? (
<div className="bg-green-100 text-green-800 p-4 rounded">
No failure detected - RAG pipeline performed correctly
</div>
) : (
<>
<div className={`${severityColor[result.severity]} p-6 rounded-lg shadow`}>
<h2 className="text-2xl font-bold capitalize">
{result.failure_type.replace(/_/g, ' ')}
</h2>
<p className="mt-2">Stage: <span className="font-semibold">{result.failure_stage}</span></p>
<p>Confidence: {(result.confidence * 100).toFixed(0)}% | Severity: {result.severity}</p>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Root Cause Analysis</h3>
<p className="text-gray-700">{result.root_cause}</p>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Recommended Remediations</h3>
<ul className="space-y-2">
{result.remediations.map((r, i) => (
<li key={i} className="flex justify-between items-center border-b pb-2">
<span>• {r.action}</span>
<span className={`text-xs px-2 py-1 rounded ${
r.effort === 'low' ? 'bg-green-100' :
r.effort === 'medium' ? 'bg-yellow-100' : 'bg-red-100'
}`}>{r.effort} effort</span>
</li>
))}
</ul>
</div>
{result.recurring_patterns.length > 0 && (
<div className="bg-purple-50 border-l-4 border-purple-500 p-4 rounded">
<h3 className="font-bold mb-2"> Recurring Patterns Detected</h3>
<ul className="list-disc list-inside">
{result.recurring_patterns.map((p, i) => <li key={i}>{p}</li>)}
</ul>
</div>
)}
</>
)}
</div>
)}
</div>
);
};
Real-Time Use Case: Financial Services Compliance Q&A
A compliance officer at a major bank asks: "What is the maximum transaction limit for corporate accounts under the new 2024 AML policy?"
The RAG system retrieves two chunks:
The 2023 AML policy (score 0.82) stating $10,000 limit
An unrelated onboarding procedure doc (score 0.71)
The system answers: "The maximum transaction limit is $10,000 per the current AML policy."
The officer flags this as wrong—the 2024 policy increased the limit to $50,000.
The diagnostician analyzes:
Retrieval Auditor detects: retrieval_adequate=false, missing_information="2024 policy document", noise_ratio=0.5
Failure Classifier identifies: stale_knowledge failure at the retrieval stage, severity=critical (regulatory impact), confidence=0.92
Root Cause Analyzer traces: "No document versioning or recency weighting in retrieval; 2023 document scored higher due to lexical overlap"
Remediation Recommender suggests:
Add document versioning with recency scoring (medium effort)
Implement TTL-based cache invalidation (low effort)
The system also flags a recurring pattern: this is the 4th stale_knowledge failure in the past 50 interactions, prompting the team to prioritize versioning infrastructure.
Conclusion
RAG failures are not monolithic they are a diverse taxonomy of issues spanning query formulation, retrieval quality, and generation behavior. Treating all failures the same leads to generic "improve your prompts" advice that rarely addresses root causes. By building a multi-agent LangGraph diagnostician that audits retrieval, classifies failure modes, traces root causes, and recommends targeted fixes, enterprises gain a systematic approach to RAG reliability. The persistent memory layer transforms isolated incidents into pattern detection, enabling proactive infrastructure improvements rather than reactive firefighting. This diagnostician doesn't just identify what went wrong—it tells you exactly which component to fix and how, turning RAG from a fragile prototype into a production-grade enterprise system.

Join the conversation! Your thoughts help the community grow.