Introduction
Hallucination the generation of fluent, confident, but factually incorrect content is the single greatest barrier to enterprise LLM adoption. In a 2024 Gartner survey, 78% of enterprises cited hallucination risk as their top concern preventing production deployment. The problem is not merely academic: a hallucinated clause in a legal contract review, a fabricated drug interaction in a clinical decision support system, or an invented financial regulation in a compliance tool can result in lawsuits, patient harm, or regulatory fines.
This article presents a comprehensive, three-layer hallucination defense architecture built as an enterprise multi-agent LangGraph RAG system. Rather than relying on a single mitigation technique, we implement a defense-in-depth strategy: prevention through grounded retrieval and constrained generation, detection through self-consistency voting, natural language inference (NLI) fact-checking, and citation verification, and correction through feedback-driven regeneration. The system maintains persistent memory of hallucination patterns to continuously improve its defenses over time.
Understanding Hallucination: Taxonomy and Root Causes
Hallucinations in LLMs fall into three categories:
Type | Description | Example |
|---|---|---|
Intrinsic | Contradicts the provided source material | Source says "revenue $5M," model outputs "$50M" |
Extrinsic | Introduces information not in any source | Model invents a court case that doesn't exist |
Faithfulness | Logically inconsistent with the context | Source says "Party A may terminate," model says "Party A must terminate" |
Root causes include training data contamination, the model's tendency to maximize fluency over accuracy, attention dilution in long contexts, and the fundamental gap between next-token prediction and factual reasoning.
The Three-Layer Defense Strategy
Layer 1 — Prevention: Constrain the model's generation space by grounding it in retrieved evidence, using strict prompt templates, and limiting temperature.
Layer 2 — Detection: Run multiple independent verification checks in parallel—self-consistency voting (generate N times, check agreement), NLI-based entailment checking (does the source entail the claim?), and citation verification (do the cited sources actually contain the claimed facts?).
Layer 3 — Correction: When hallucination is detected, feed the specific failure feedback back to the generator and regenerate with stronger constraints.
Step-by-Step Implementation
1. Hallucination-Aware State Schema
from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage
from dataclasses import dataclass, field
import redis
import json
from datetime import datetime
class ClaimVerification(TypedDict):
claim: str
source_chunk_id: str
nli_label: Literal["entailment", "contradiction", "neutral"]
nli_confidence: float
citation_valid: bool
class HallucinationState(TypedDict):
messages: List
conversation_id: str
query: str
# Retrieval
retrieved_chunks: List[Dict]
# Generation
primary_answer: str
consistency_answers: List[str]
# Detection
extracted_claims: List[str]
claim_verifications: List[ClaimVerification]
self_consistency_score: float
nli_faithfulness_score: float
citation_accuracy_score: float
overall_hallucination_score: float
hallucination_detected: bool
hallucination_details: str
# Correction
corrected_answer: str
regeneration_count: int
max_regen: int
# Memory
historical_hallucination_rate: float
hallucination_log: List[Dict]
2. Grounded Retrieval Agent (Prevention Layer)
from langchain_community.vectorstores import PGVector
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_openai import ChatOpenAI
class GroundedRetrievalAgent:
"""Prevents hallucination by enforcing strict grounding in retrieved evidence"""
def __init__(self, vector_store: PGVector):
self.vector_store = vector_store
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def retrieve_and_generate(self, state: HallucinationState) -> HallucinationState:
query = state["query"]
docs = self.vector_store.similarity_search_with_score(query, k=6)
chunks = []
for i, (doc, score) in enumerate(docs):
chunks.append({
"id": f"chunk_{i}",
"content": doc.page_content,
"source": doc.metadata.get("source", "unknown"),
"score": score
})
state["retrieved_chunks"] = chunks
# Grounded generation with strict constraints
context = "\n".join([
f"[{c['id']}] ({c['source']}): {c['content']}" for c in chunks
])
prompt = f"""You are a precise factual assistant. Answer ONLY using the provided sources.
RULES:
1. Every factual claim must cite its source chunk ID in brackets, e.g., [chunk_0].
2. If the answer is not fully contained in the sources, explicitly state what is missing.
3. Never fabricate information, statistics, dates, or names not present in the sources.
4. If sources conflict, report both perspectives with their citations.
Sources:
{context}
Question: {query}
Answer:"""
response = self.llm.invoke(prompt)
state["primary_answer"] = response.content
return state
3. Self-Consistency Voting Agent (Detection Layer 1)
class SelfConsistencyAgent:
"""Generates multiple answers and checks agreement to detect hallucination"""
def __init__(self, n_samples: int = 3):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.4)
self.n_samples = n_samples
def check_consistency(self, state: HallucinationState) -> HallucinationState:
context = "\n".join([c["content"] for c in state["retrieved_chunks"]])
query = state["query"]
answers = []
for _ in range(self.n_samples):
prompt = f"Based on this context:\n{context}\n\nAnswer: {query}\n\nBe concise."
resp = self.llm.invoke(prompt)
answers.append(resp.content)
state["consistency_answers"] = answers
# Compute consistency via pairwise similarity
from difflib import SequenceMatcher
scores = []
for i in range(len(answers)):
for j in range(i + 1, len(answers)):
ratio = SequenceMatcher(None, answers[i], answers[j]).ratio()
scores.append(ratio)
state["self_consistency_score"] = sum(scores) / max(len(scores), 1)
return state
4. NLI Fact-Checker Agent (Detection Layer 2)
from transformers import pipeline
class NLIFactCheckerAgent:
"""Uses Natural Language Inference to verify each claim against source chunks"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
# Load a lightweight NLI model for fast inference
self.nli_pipeline = pipeline(
"text-classification",
model="cross-encoder/nli-deberta-v3-small",
top_k=None
)
def extract_claims(self, text: str) -> List[str]:
"""Decompose answer into atomic factual claims"""
prompt = f"""Break this text into individual factual claims.
Each claim should be a single, verifiable statement.
Return as a JSON array of strings.
Text: {text}"""
response = self.llm.invoke(prompt)
try:
import re
match = re.search(r'\[[\s\S]*\]', response.content)
return json.loads(match.group()) if match else [text]
except Exception:
return [text]
def verify_claims(self, state: HallucinationState) -> HallucinationState:
answer = state.get("corrected_answer") or state["primary_answer"]
claims = self.extract_claims(answer)
state["extracted_claims"] = claims
source_text = " ".join([c["content"] for c in state["retrieved_chunks"]])
verifications = []
for claim in claims:
# NLI: Does the source entail this claim?
result = self.nli_pipeline([{"text": source_text, "text_pair": claim}])
# Parse NLI output
labels = {item["label"].lower(): item["score"] for item in result[0]}
entailment_score = labels.get("entailment", 0.0)
contradiction_score = labels.get("contradiction", 0.0)
if entailment_score > 0.6:
nli_label = "entailment"
elif contradiction_score > 0.5:
nli_label = "contradiction"
else:
nli_label = "neutral"
verifications.append(ClaimVerification(
claim=claim,
source_chunk_id="aggregate",
nli_label=nli_label,
nli_confidence=max(entailment_score, contradiction_score),
citation_valid=nli_label == "entailment"
))
state["claim_verifications"] = verifications
# Compute faithfulness score
if verifications:
entailed = sum(1 for v in verifications if v["nli_label"] == "entailment")
state["nli_faithfulness_score"] = entailed / len(verifications)
else:
state["nli_faithfulness_score"] = 1.0
return state
5. Citation Verifier Agent (Detection Layer 3)
class CitationVerifierAgent:
"""Verifies that cited sources actually contain the claimed information"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def verify_citations(self, state: HallucinationState) -> HallucinationState:
answer = state.get("corrected_answer") or state["primary_answer"]
chunks = state["retrieved_chunks"]
import re
citations = re.findall(r'\[(chunk_\d+)\]', answer)
if not citations:
state["citation_accuracy_score"] = 0.5 # No citations to verify
return state
valid_citations = 0
for cite_id in citations:
chunk = next((c for c in chunks if c["id"] == cite_id), None)
if chunk:
prompt = f"""Does this source support the claim made near citation [{cite_id}]?
Source: {chunk['content'][:500]}
Answer excerpt: {answer[max(0, answer.find(cite_id)-100):answer.find(cite_id)+50]}
Respond YES or NO."""
resp = self.llm.invoke(prompt)
if "yes" in resp.content.lower():
valid_citations += 1
state["citation_accuracy_score"] = valid_citations / max(len(citations), 1)
return state
6. Regeneration Agent (Correction Layer)
class RegenerationAgent:
"""Regenerates the answer with specific hallucination feedback"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def regenerate(self, state: HallucinationState) -> HallucinationState:
bad_claims = [
v["claim"] for v in state["claim_verifications"]
if v["nli_label"] in ("contradiction", "neutral")
]
context = "\n".join([
f"[{c['id']}] ({c['source']}): {c['content']}"
for c in state["retrieved_chunks"]
])
prompt = f"""Your previous answer contained hallucinations. Fix them.
ORIGINAL ANSWER:
{state['primary_answer']}
HALLUCINATED CLAIMS TO REMOVE/FIX:
{chr(10).join(['- ' + c for c in bad_claims])}
SOURCES (use ONLY these):
{context}
QUESTION: {state['query']}
REGENERATED ANSWER (remove all unsupported claims, cite sources):"""
response = self.llm.invoke(prompt)
state["corrected_answer"] = response.content
state["regeneration_count"] += 1
return state
7. LangGraph Workflow with Conditional Routing and Memory
class HallucinationMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def log_result(self, state: HallucinationState):
record = {
"timestamp": datetime.now().isoformat(),
"query": state["query"][:200],
"hallucination_detected": state["hallucination_detected"],
"hallucination_score": state["overall_hallucination_score"],
"nli_score": state["nli_faithfulness_score"],
"consistency_score": state["self_consistency_score"],
"regenerations": state["regeneration_count"]
}
key = f"hallucination_log:{state['conversation_id']}"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-200:]))
def get_historical_rate(self, conversation_id: str) -> float:
history = json.loads(
self.redis.get(f"hallucination_log:{conversation_id}") or "[]"
)
if not history:
return 0.0
detected = sum(1 for h in history if h["hallucination_detected"])
return detected / len(history)
def build_hallucination_firewall():
workflow = StateGraph(HallucinationState)
retriever = GroundedRetrievalAgent(None) # Pass real vector store
consistency = SelfConsistencyAgent(n_samples=3)
nli_checker = NLIFactCheckerAgent()
citation_checker = CitationVerifierAgent()
regenerator = RegenerationAgent()
def compute_overall_score(state: HallucinationState) -> HallucinationState:
"""Aggregate all detection signals into a single hallucination score"""
consistency_penalty = 1.0 - state["self_consistency_score"]
nli_penalty = 1.0 - state["nli_faithfulness_score"]
citation_penalty = 1.0 - state["citation_accuracy_score"]
# Weighted combination
state["overall_hallucination_score"] = (
0.25 * consistency_penalty +
0.50 * nli_penalty +
0.25 * citation_penalty
)
state["hallucination_detected"] = state["overall_hallucination_score"] > 0.35
if state["hallucination_detected"]:
bad_claims = [
v["claim"] for v in state["claim_verifications"]
if v["nli_label"] != "entailment"
]
state["hallucination_details"] = (
f"Hallucination score: {state['overall_hallucination_score']:.2f}. "
f"Unsupported claims: {', '.join(bad_claims[:3])}"
)
else:
state["hallucination_details"] = "All checks passed."
return state
workflow.add_node("retrieve_and_generate", retriever.retrieve_and_generate)
workflow.add_node("check_consistency", consistency.check_consistency)
workflow.add_node("nli_fact_check", nli_checker.verify_claims)
workflow.add_node("verify_citations", citation_checker.verify_citations)
workflow.add_node("compute_score", compute_overall_score)
workflow.add_node("regenerate", regenerator.regenerate)
workflow.set_entry_point("retrieve_and_generate")
workflow.add_edge("retrieve_and_generate", "check_consistency")
workflow.add_edge("check_consistency", "nli_fact_check")
workflow.add_edge("nli_fact_check", "verify_citations")
workflow.add_edge("verify_citations", "compute_score")
def route_after_scoring(state: HallucinationState):
if state["hallucination_detected"] and state["regeneration_count"] < state["max_regen"]:
return "regenerate"
return "end"
workflow.add_conditional_edges(
"compute_score",
route_after_scoring,
{"regenerate": "regenerate", "end": END}
)
# After regeneration, re-run detection
workflow.add_edge("regenerate", "nli_fact_check")
return workflow.compile()
8. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="Hallucination Firewall API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_hallucination_firewall()
memory = HallucinationMemory(redis.Redis())
class QueryRequest(BaseModel):
conversation_id: str
query: str
max_regen: int = 2
@app.post("/ask")
async def ask(req: QueryRequest):
initial_state = HallucinationState(
messages=[HumanMessage(content=req.query)],
conversation_id=req.conversation_id,
query=req.query,
retrieved_chunks=[],
primary_answer="",
consistency_answers=[],
extracted_claims=[],
claim_verifications=[],
self_consistency_score=0.0,
nli_faithfulness_score=0.0,
citation_accuracy_score=0.0,
overall_hallucination_score=0.0,
hallucination_detected=False,
hallucination_details="",
corrected_answer="",
regeneration_count=0,
max_regen=req.max_regen,
historical_hallucination_rate=memory.get_historical_rate(req.conversation_id),
hallucination_log=[]
)
result = graph.invoke(initial_state)
memory.log_result(result)
final_answer = result["corrected_answer"] or result["primary_answer"]
return {
"answer": final_answer,
"hallucination_detected": result["hallucination_detected"],
"hallucination_score": round(result["overall_hallucination_score"], 3),
"detection_breakdown": {
"self_consistency": round(result["self_consistency_score"], 3),
"nli_faithfulness": round(result["nli_faithfulness_score"], 3),
"citation_accuracy": round(result["citation_accuracy_score"], 3)
},
"details": result["hallucination_details"],
"regenerations": result["regeneration_count"],
"claims_verified": len(result["claim_verifications"]),
"historical_rate": round(result["historical_hallucination_rate"], 3),
"sources": [
{"id": c["id"], "source": c["source"]}
for c in result["retrieved_chunks"]
]
}
9. Frontend: Hallucination Transparency Dashboard
// components/HallucinationFirewall.tsx
import React, { useState } from 'react';
interface HallucinationResult {
answer: string;
hallucination_detected: boolean;
hallucination_score: number;
detection_breakdown: {
self_consistency: number;
nli_faithfulness: number;
citation_accuracy: number;
};
details: string;
regenerations: number;
claims_verified: number;
historical_rate: number;
sources: { id: string; source: string }[];
}
const ScoreGauge: React.FC<{ label: string; score: number; invert?: boolean }> = ({
label, score, invert
}) => {
const displayScore = invert ? 1 - score : score;
const color = displayScore > 0.7 ? 'text-green-600' :
displayScore > 0.4 ? 'text-yellow-600' : 'text-red-600';
return (
<div className="flex items-center gap-2">
<span className="text-sm w-36">{label}</span>
<div className="flex-1 bg-gray-200 rounded-full h-3">
<div className={`${displayScore > 0.7 ? 'bg-green-500' : displayScore > 0.4 ? 'bg-yellow-500' : 'bg-red-500'} h-3 rounded-full transition-all`}
style={{ width: `${displayScore * 100}%` }} />
</div>
<span className={`text-sm font-bold ${color}`}>{(displayScore * 100).toFixed(0)}%</span>
</div>
);
};
export const HallucinationFirewall: React.FC = () => {
const [query, setQuery] = useState(
"What are the indemnification obligations of Party B under Section 8.3?"
);
const [result, setResult] = useState<HallucinationResult | null>(null);
const [loading, setLoading] = useState(false);
const handleSubmit = async () => {
setLoading(true);
try {
const response = await fetch('http://localhost:8000/ask', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
conversation_id: 'legal-contract-001',
query,
max_regen: 2
})
});
setResult(await response.json());
} finally {
setLoading(false);
}
};
return (
<div className="p-6 max-w-5xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-2"> Hallucination Firewall</h1>
<p className="text-gray-600 mb-6">Multi-layered hallucination detection and correction</p>
<div className="bg-white p-4 rounded-lg shadow mb-6">
<textarea
className="w-full p-3 border rounded mb-3"
rows={2}
value={query}
onChange={e => setQuery(e.target.value)}
placeholder="Ask a question..."
/>
<button onClick={handleSubmit} disabled={loading}
className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
{loading ? 'Analyzing...' : 'Ask with Hallucination Check'}
</button>
</div>
{result && (
<div className="grid grid-cols-3 gap-4">
<div className="col-span-2 space-y-4">
<div className={`p-4 rounded-lg shadow ${
result.hallucination_detected ? 'bg-red-50 border border-red-200' : 'bg-green-50 border border-green-200'
}`}>
<div className="flex items-center gap-2 mb-2">
<span className="text-xl">{result.hallucination_detected ? '⚠️' : '✅'}</span>
<h2 className="font-bold">
{result.hallucination_detected ? 'Hallucination Detected & Corrected' : 'Answer Verified'}
</h2>
</div>
<p className="text-gray-800 whitespace-pre-wrap">{result.answer}</p>
{result.regenerations > 0 && (
<p className="text-sm text-blue-600 mt-2">
Regenerated {result.regenerations} time(s) to fix hallucinations
</p>
)}
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Sources</h3>
{result.sources.map((s, i) => (
<div key={i} className="text-sm text-gray-600">
[{s.id}] {s.source}
</div>
))}
</div>
</div>
<div className="space-y-4">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">Detection Scores</h3>
<div className="space-y-3">
<ScoreGauge label="Self-Consistency" score={result.detection_breakdown.self_consistency} />
<ScoreGauge label="NLI Faithfulness" score={result.detection_breakdown.nli_faithfulness} />
<ScoreGauge label="Citation Accuracy" score={result.detection_breakdown.citation_accuracy} />
<hr />
<ScoreGauge label="Hallucination Risk" score={result.hallucination_score} invert />
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Details</h3>
<p className="text-sm text-gray-700">{result.details}</p>
<p className="text-xs text-gray-500 mt-2">
Claims verified: {result.claims_verified} |
Historical rate: {(result.historical_rate * 100).toFixed(1)}%
</p>
</div>
</div>
</div>
)}
</div>
);
};
Real-Time Use Case: Legal Contract Analysis Assistant
A law firm deploys an LLM assistant to analyze M&A contracts. A junior associate asks: "What are the indemnification obligations of Party B under Section 8.3, and is there a liability cap?"
Layer 1 — Prevention: The retrieval agent fetches 6 chunks from the 120-page contract. The grounded generation prompt forces citation of chunk IDs. The primary answer states: "Party B shall indemnify Party A against all third-party claims [chunk_2]. The liability cap is $10M as stated in Section 8.3 [chunk_4]."
Layer 2 — Detection:
Self-Consistency: 3 generations agree on indemnification scope but disagree on the cap amount ($10M vs $15M vs "uncapped"). Consistency score: 0.52.
NLI Fact-Check: The claim "liability cap is $10M" is checked against chunk_4. The NLI model returns contradiction (chunk_4 actually says "$15M aggregate cap"). Faithfulness score: 0.67.
Citation Verification: The citation [chunk_4] for "$10M" is verified—chunk_4 contains "$15M," not "$10M." Citation accuracy: 0.50.
Overall hallucination score: 0.44 (threshold: 0.35). Hallucination detected.
Layer 3 — Correction: The regeneration agent receives feedback: "The claim 'liability cap is $10M' contradicts chunk_4 which states $15M." It regenerates: "Party B shall indemnify Party A against all third-party claims [chunk_2]. The aggregate liability cap under Section 8.3 is $15M [chunk_4]."
The corrected answer passes all three verification layers on the second pass. The system logs this incident, and the historical hallucination rate for this conversation drops from 12% to 8% over the next 50 queries as the memory-informed prompts improve.
Conclusion
Hallucination in LLMs is not a problem that can be solved with a single technique it requires a defense-in-depth architecture that prevents, detects, and corrects fabricated content at multiple stages of the pipeline. By implementing self-consistency voting, NLI-based fact-checking, and citation verification as independent detection layers within a LangGraph multi-agent workflow, enterprises can quantify hallucination risk with precision rather than relying on intuition. The correction loop ensures that detected hallucinations are fixed before reaching the user, while persistent memory enables continuous improvement over time. This three-layer firewall transforms LLMs from confident but unreliable generators into trustworthy enterprise tools that can be deployed in high-stakes domains like legal, healthcare, and finance with measurable confidence.

Join the conversation! Your thoughts help the community grow.