Introduction

Hallucination the generation of fluent, confident, but factually incorrect content is the single greatest barrier to enterprise LLM adoption. In a 2024 Gartner survey, 78% of enterprises cited hallucination risk as their top concern preventing production deployment. The problem is not merely academic: a hallucinated clause in a legal contract review, a fabricated drug interaction in a clinical decision support system, or an invented financial regulation in a compliance tool can result in lawsuits, patient harm, or regulatory fines.

This article presents a comprehensive, three-layer hallucination defense architecture built as an enterprise multi-agent LangGraph RAG system. Rather than relying on a single mitigation technique, we implement a defense-in-depth strategy: prevention through grounded retrieval and constrained generation, detection through self-consistency voting, natural language inference (NLI) fact-checking, and citation verification, and correction through feedback-driven regeneration. The system maintains persistent memory of hallucination patterns to continuously improve its defenses over time.

Understanding Hallucination: Taxonomy and Root Causes

Hallucinations in LLMs fall into three categories:

Type

Description

Example

Intrinsic

Contradicts the provided source material

Source says "revenue $5M," model outputs "$50M"

Extrinsic

Introduces information not in any source

Model invents a court case that doesn't exist

Faithfulness

Logically inconsistent with the context

Source says "Party A may terminate," model says "Party A must terminate"

Root causes include training data contamination, the model's tendency to maximize fluency over accuracy, attention dilution in long contexts, and the fundamental gap between next-token prediction and factual reasoning.

The Three-Layer Defense Strategy

Layer 1 — Prevention: Constrain the model's generation space by grounding it in retrieved evidence, using strict prompt templates, and limiting temperature.

Layer 2 — Detection: Run multiple independent verification checks in parallel—self-consistency voting (generate N times, check agreement), NLI-based entailment checking (does the source entail the claim?), and citation verification (do the cited sources actually contain the claimed facts?).

Layer 3 — Correction: When hallucination is detected, feed the specific failure feedback back to the generator and regenerate with stronger constraints.

Step-by-Step Implementation

1. Hallucination-Aware State Schema

from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage
from dataclasses import dataclass, field
import redis
import json
from datetime import datetime

class ClaimVerification(TypedDict):
    claim: str
    source_chunk_id: str
    nli_label: Literal["entailment", "contradiction", "neutral"]
    nli_confidence: float
    citation_valid: bool

class HallucinationState(TypedDict):
    messages: List
    conversation_id: str
    query: str
    # Retrieval
    retrieved_chunks: List[Dict]
    # Generation
    primary_answer: str
    consistency_answers: List[str]
    # Detection
    extracted_claims: List[str]
    claim_verifications: List[ClaimVerification]
    self_consistency_score: float
    nli_faithfulness_score: float
    citation_accuracy_score: float
    overall_hallucination_score: float
    hallucination_detected: bool
    hallucination_details: str
    # Correction
    corrected_answer: str
    regeneration_count: int
    max_regen: int
    # Memory
    historical_hallucination_rate: float
    hallucination_log: List[Dict]

2. Grounded Retrieval Agent (Prevention Layer)

from langchain_community.vectorstores import PGVector
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_openai import ChatOpenAI

class GroundedRetrievalAgent:
    """Prevents hallucination by enforcing strict grounding in retrieved evidence"""
    
    def __init__(self, vector_store: PGVector):
        self.vector_store = vector_store
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def retrieve_and_generate(self, state: HallucinationState) -> HallucinationState:
        query = state["query"]
        docs = self.vector_store.similarity_search_with_score(query, k=6)
        
        chunks = []
        for i, (doc, score) in enumerate(docs):
            chunks.append({
                "id": f"chunk_{i}",
                "content": doc.page_content,
                "source": doc.metadata.get("source", "unknown"),
                "score": score
            })
        state["retrieved_chunks"] = chunks
        
        # Grounded generation with strict constraints
        context = "\n".join([
            f"[{c['id']}] ({c['source']}): {c['content']}" for c in chunks
        ])
        
        prompt = f"""You are a precise factual assistant. Answer ONLY using the provided sources.
RULES:
1. Every factual claim must cite its source chunk ID in brackets, e.g., [chunk_0].
2. If the answer is not fully contained in the sources, explicitly state what is missing.
3. Never fabricate information, statistics, dates, or names not present in the sources.
4. If sources conflict, report both perspectives with their citations.

Sources:
{context}

Question: {query}

Answer:"""
        
        response = self.llm.invoke(prompt)
        state["primary_answer"] = response.content
        return state

3. Self-Consistency Voting Agent (Detection Layer 1)

class SelfConsistencyAgent:
    """Generates multiple answers and checks agreement to detect hallucination"""
    
    def __init__(self, n_samples: int = 3):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.4)
        self.n_samples = n_samples
    
    def check_consistency(self, state: HallucinationState) -> HallucinationState:
        context = "\n".join([c["content"] for c in state["retrieved_chunks"]])
        query = state["query"]
        
        answers = []
        for _ in range(self.n_samples):
            prompt = f"Based on this context:\n{context}\n\nAnswer: {query}\n\nBe concise."
            resp = self.llm.invoke(prompt)
            answers.append(resp.content)
        
        state["consistency_answers"] = answers
        
        # Compute consistency via pairwise similarity
        from difflib import SequenceMatcher
        scores = []
        for i in range(len(answers)):
            for j in range(i + 1, len(answers)):
                ratio = SequenceMatcher(None, answers[i], answers[j]).ratio()
                scores.append(ratio)
        
        state["self_consistency_score"] = sum(scores) / max(len(scores), 1)
        return state

4. NLI Fact-Checker Agent (Detection Layer 2)

from transformers import pipeline

class NLIFactCheckerAgent:
    """Uses Natural Language Inference to verify each claim against source chunks"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
        # Load a lightweight NLI model for fast inference
        self.nli_pipeline = pipeline(
            "text-classification",
            model="cross-encoder/nli-deberta-v3-small",
            top_k=None
        )
    
    def extract_claims(self, text: str) -> List[str]:
        """Decompose answer into atomic factual claims"""
        prompt = f"""Break this text into individual factual claims. 
Each claim should be a single, verifiable statement.
Return as a JSON array of strings.

Text: {text}"""
        response = self.llm.invoke(prompt)
        try:
            import re
            match = re.search(r'\[[\s\S]*\]', response.content)
            return json.loads(match.group()) if match else [text]
        except Exception:
            return [text]
    
    def verify_claims(self, state: HallucinationState) -> HallucinationState:
        answer = state.get("corrected_answer") or state["primary_answer"]
        claims = self.extract_claims(answer)
        state["extracted_claims"] = claims
        
        source_text = " ".join([c["content"] for c in state["retrieved_chunks"]])
        verifications = []
        
        for claim in claims:
            # NLI: Does the source entail this claim?
            result = self.nli_pipeline([{"text": source_text, "text_pair": claim}])
            
            # Parse NLI output
            labels = {item["label"].lower(): item["score"] for item in result[0]}
            entailment_score = labels.get("entailment", 0.0)
            contradiction_score = labels.get("contradiction", 0.0)
            
            if entailment_score > 0.6:
                nli_label = "entailment"
            elif contradiction_score > 0.5:
                nli_label = "contradiction"
            else:
                nli_label = "neutral"
            
            verifications.append(ClaimVerification(
                claim=claim,
                source_chunk_id="aggregate",
                nli_label=nli_label,
                nli_confidence=max(entailment_score, contradiction_score),
                citation_valid=nli_label == "entailment"
            ))
        
        state["claim_verifications"] = verifications
        
        # Compute faithfulness score
        if verifications:
            entailed = sum(1 for v in verifications if v["nli_label"] == "entailment")
            state["nli_faithfulness_score"] = entailed / len(verifications)
        else:
            state["nli_faithfulness_score"] = 1.0
        
        return state

5. Citation Verifier Agent (Detection Layer 3)

class CitationVerifierAgent:
    """Verifies that cited sources actually contain the claimed information"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def verify_citations(self, state: HallucinationState) -> HallucinationState:
        answer = state.get("corrected_answer") or state["primary_answer"]
        chunks = state["retrieved_chunks"]
        
        import re
        citations = re.findall(r'\[(chunk_\d+)\]', answer)
        
        if not citations:
            state["citation_accuracy_score"] = 0.5  # No citations to verify
            return state
        
        valid_citations = 0
        for cite_id in citations:
            chunk = next((c for c in chunks if c["id"] == cite_id), None)
            if chunk:
                prompt = f"""Does this source support the claim made near citation [{cite_id}]?
Source: {chunk['content'][:500]}
Answer excerpt: {answer[max(0, answer.find(cite_id)-100):answer.find(cite_id)+50]}
Respond YES or NO."""
                resp = self.llm.invoke(prompt)
                if "yes" in resp.content.lower():
                    valid_citations += 1
        
        state["citation_accuracy_score"] = valid_citations / max(len(citations), 1)
        return state

6. Regeneration Agent (Correction Layer)

class RegenerationAgent:
    """Regenerates the answer with specific hallucination feedback"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def regenerate(self, state: HallucinationState) -> HallucinationState:
        bad_claims = [
            v["claim"] for v in state["claim_verifications"]
            if v["nli_label"] in ("contradiction", "neutral")
        ]
        
        context = "\n".join([
            f"[{c['id']}] ({c['source']}): {c['content']}"
            for c in state["retrieved_chunks"]
        ])
        
        prompt = f"""Your previous answer contained hallucinations. Fix them.

ORIGINAL ANSWER:
{state['primary_answer']}

HALLUCINATED CLAIMS TO REMOVE/FIX:
{chr(10).join(['- ' + c for c in bad_claims])}

SOURCES (use ONLY these):
{context}

QUESTION: {state['query']}

REGENERATED ANSWER (remove all unsupported claims, cite sources):"""
        
        response = self.llm.invoke(prompt)
        state["corrected_answer"] = response.content
        state["regeneration_count"] += 1
        return state

7. LangGraph Workflow with Conditional Routing and Memory

class HallucinationMemory:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def log_result(self, state: HallucinationState):
        record = {
            "timestamp": datetime.now().isoformat(),
            "query": state["query"][:200],
            "hallucination_detected": state["hallucination_detected"],
            "hallucination_score": state["overall_hallucination_score"],
            "nli_score": state["nli_faithfulness_score"],
            "consistency_score": state["self_consistency_score"],
            "regenerations": state["regeneration_count"]
        }
        key = f"hallucination_log:{state['conversation_id']}"
        history = json.loads(self.redis.get(key) or "[]")
        history.append(record)
        self.redis.set(key, json.dumps(history[-200:]))
    
    def get_historical_rate(self, conversation_id: str) -> float:
        history = json.loads(
            self.redis.get(f"hallucination_log:{conversation_id}") or "[]"
        )
        if not history:
            return 0.0
        detected = sum(1 for h in history if h["hallucination_detected"])
        return detected / len(history)

def build_hallucination_firewall():
    workflow = StateGraph(HallucinationState)
    
    retriever = GroundedRetrievalAgent(None)  # Pass real vector store
    consistency = SelfConsistencyAgent(n_samples=3)
    nli_checker = NLIFactCheckerAgent()
    citation_checker = CitationVerifierAgent()
    regenerator = RegenerationAgent()
    
    def compute_overall_score(state: HallucinationState) -> HallucinationState:
        """Aggregate all detection signals into a single hallucination score"""
        consistency_penalty = 1.0 - state["self_consistency_score"]
        nli_penalty = 1.0 - state["nli_faithfulness_score"]
        citation_penalty = 1.0 - state["citation_accuracy_score"]
        
        # Weighted combination
        state["overall_hallucination_score"] = (
            0.25 * consistency_penalty +
            0.50 * nli_penalty +
            0.25 * citation_penalty
        )
        state["hallucination_detected"] = state["overall_hallucination_score"] > 0.35
        
        if state["hallucination_detected"]:
            bad_claims = [
                v["claim"] for v in state["claim_verifications"]
                if v["nli_label"] != "entailment"
            ]
            state["hallucination_details"] = (
                f"Hallucination score: {state['overall_hallucination_score']:.2f}. "
                f"Unsupported claims: {', '.join(bad_claims[:3])}"
            )
        else:
            state["hallucination_details"] = "All checks passed."
        
        return state
    
    workflow.add_node("retrieve_and_generate", retriever.retrieve_and_generate)
    workflow.add_node("check_consistency", consistency.check_consistency)
    workflow.add_node("nli_fact_check", nli_checker.verify_claims)
    workflow.add_node("verify_citations", citation_checker.verify_citations)
    workflow.add_node("compute_score", compute_overall_score)
    workflow.add_node("regenerate", regenerator.regenerate)
    
    workflow.set_entry_point("retrieve_and_generate")
    workflow.add_edge("retrieve_and_generate", "check_consistency")
    workflow.add_edge("check_consistency", "nli_fact_check")
    workflow.add_edge("nli_fact_check", "verify_citations")
    workflow.add_edge("verify_citations", "compute_score")
    
    def route_after_scoring(state: HallucinationState):
        if state["hallucination_detected"] and state["regeneration_count"] < state["max_regen"]:
            return "regenerate"
        return "end"
    
    workflow.add_conditional_edges(
        "compute_score",
        route_after_scoring,
        {"regenerate": "regenerate", "end": END}
    )
    
    # After regeneration, re-run detection
    workflow.add_edge("regenerate", "nli_fact_check")
    
    return workflow.compile()

8. FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel

app = FastAPI(title="Hallucination Firewall API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])

graph = build_hallucination_firewall()
memory = HallucinationMemory(redis.Redis())

class QueryRequest(BaseModel):
    conversation_id: str
    query: str
    max_regen: int = 2

@app.post("/ask")
async def ask(req: QueryRequest):
    initial_state = HallucinationState(
        messages=[HumanMessage(content=req.query)],
        conversation_id=req.conversation_id,
        query=req.query,
        retrieved_chunks=[],
        primary_answer="",
        consistency_answers=[],
        extracted_claims=[],
        claim_verifications=[],
        self_consistency_score=0.0,
        nli_faithfulness_score=0.0,
        citation_accuracy_score=0.0,
        overall_hallucination_score=0.0,
        hallucination_detected=False,
        hallucination_details="",
        corrected_answer="",
        regeneration_count=0,
        max_regen=req.max_regen,
        historical_hallucination_rate=memory.get_historical_rate(req.conversation_id),
        hallucination_log=[]
    )
    
    result = graph.invoke(initial_state)
    memory.log_result(result)
    
    final_answer = result["corrected_answer"] or result["primary_answer"]
    
    return {
        "answer": final_answer,
        "hallucination_detected": result["hallucination_detected"],
        "hallucination_score": round(result["overall_hallucination_score"], 3),
        "detection_breakdown": {
            "self_consistency": round(result["self_consistency_score"], 3),
            "nli_faithfulness": round(result["nli_faithfulness_score"], 3),
            "citation_accuracy": round(result["citation_accuracy_score"], 3)
        },
        "details": result["hallucination_details"],
        "regenerations": result["regeneration_count"],
        "claims_verified": len(result["claim_verifications"]),
        "historical_rate": round(result["historical_hallucination_rate"], 3),
        "sources": [
            {"id": c["id"], "source": c["source"]}
            for c in result["retrieved_chunks"]
        ]
    }

9. Frontend: Hallucination Transparency Dashboard

// components/HallucinationFirewall.tsx
import React, { useState } from 'react';

interface HallucinationResult {
  answer: string;
  hallucination_detected: boolean;
  hallucination_score: number;
  detection_breakdown: {
    self_consistency: number;
    nli_faithfulness: number;
    citation_accuracy: number;
  };
  details: string;
  regenerations: number;
  claims_verified: number;
  historical_rate: number;
  sources: { id: string; source: string }[];
}

const ScoreGauge: React.FC<{ label: string; score: number; invert?: boolean }> = ({
  label, score, invert
}) => {
  const displayScore = invert ? 1 - score : score;
  const color = displayScore > 0.7 ? 'text-green-600' :
                displayScore > 0.4 ? 'text-yellow-600' : 'text-red-600';
  return (
    <div className="flex items-center gap-2">
      <span className="text-sm w-36">{label}</span>
      <div className="flex-1 bg-gray-200 rounded-full h-3">
        <div className={`${displayScore > 0.7 ? 'bg-green-500' : displayScore > 0.4 ? 'bg-yellow-500' : 'bg-red-500'} h-3 rounded-full transition-all`}
          style={{ width: `${displayScore * 100}%` }} />
      </div>
      <span className={`text-sm font-bold ${color}`}>{(displayScore * 100).toFixed(0)}%</span>
    </div>
  );
};

export const HallucinationFirewall: React.FC = () => {
  const [query, setQuery] = useState(
    "What are the indemnification obligations of Party B under Section 8.3?"
  );
  const [result, setResult] = useState<HallucinationResult | null>(null);
  const [loading, setLoading] = useState(false);

  const handleSubmit = async () => {
    setLoading(true);
    try {
      const response = await fetch('http://localhost:8000/ask', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({
          conversation_id: 'legal-contract-001',
          query,
          max_regen: 2
        })
      });
      setResult(await response.json());
    } finally {
      setLoading(false);
    }
  };

  return (
    <div className="p-6 max-w-5xl mx-auto bg-gray-50 min-h-screen">
      <h1 className="text-3xl font-bold mb-2">  Hallucination Firewall</h1>
      <p className="text-gray-600 mb-6">Multi-layered hallucination detection and correction</p>
      
      <div className="bg-white p-4 rounded-lg shadow mb-6">
        <textarea
          className="w-full p-3 border rounded mb-3"
          rows={2}
          value={query}
          onChange={e => setQuery(e.target.value)}
          placeholder="Ask a question..."
        />
        <button onClick={handleSubmit} disabled={loading}
          className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
          {loading ? 'Analyzing...' : 'Ask with Hallucination Check'}
        </button>
      </div>

      {result && (
        <div className="grid grid-cols-3 gap-4">
          <div className="col-span-2 space-y-4">
            <div className={`p-4 rounded-lg shadow ${
              result.hallucination_detected ? 'bg-red-50 border border-red-200' : 'bg-green-50 border border-green-200'
            }`}>
              <div className="flex items-center gap-2 mb-2">
                <span className="text-xl">{result.hallucination_detected ? '⚠️' : '✅'}</span>
                <h2 className="font-bold">
                  {result.hallucination_detected ? 'Hallucination Detected & Corrected' : 'Answer Verified'}
                </h2>
              </div>
              <p className="text-gray-800 whitespace-pre-wrap">{result.answer}</p>
              {result.regenerations > 0 && (
                <p className="text-sm text-blue-600 mt-2">
                    Regenerated {result.regenerations} time(s) to fix hallucinations
                </p>
              )}
            </div>
            
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-2">Sources</h3>
              {result.sources.map((s, i) => (
                <div key={i} className="text-sm text-gray-600">
                  [{s.id}] {s.source}
                </div>
              ))}
            </div>
          </div>

          <div className="space-y-4">
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">Detection Scores</h3>
              <div className="space-y-3">
                <ScoreGauge label="Self-Consistency" score={result.detection_breakdown.self_consistency} />
                <ScoreGauge label="NLI Faithfulness" score={result.detection_breakdown.nli_faithfulness} />
                <ScoreGauge label="Citation Accuracy" score={result.detection_breakdown.citation_accuracy} />
                <hr />
                <ScoreGauge label="Hallucination Risk" score={result.hallucination_score} invert />
              </div>
            </div>
            
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-2">Details</h3>
              <p className="text-sm text-gray-700">{result.details}</p>
              <p className="text-xs text-gray-500 mt-2">
                Claims verified: {result.claims_verified} | 
                Historical rate: {(result.historical_rate * 100).toFixed(1)}%
              </p>
            </div>
          </div>
        </div>
      )}
    </div>
  );
};

Real-Time Use Case: Legal Contract Analysis Assistant

A law firm deploys an LLM assistant to analyze M&A contracts. A junior associate asks: "What are the indemnification obligations of Party B under Section 8.3, and is there a liability cap?"

Layer 1 — Prevention: The retrieval agent fetches 6 chunks from the 120-page contract. The grounded generation prompt forces citation of chunk IDs. The primary answer states: "Party B shall indemnify Party A against all third-party claims [chunk_2]. The liability cap is $10M as stated in Section 8.3 [chunk_4]."

Layer 2 — Detection:

  • Self-Consistency: 3 generations agree on indemnification scope but disagree on the cap amount ($10M vs $15M vs "uncapped"). Consistency score: 0.52.

  • NLI Fact-Check: The claim "liability cap is $10M" is checked against chunk_4. The NLI model returns contradiction (chunk_4 actually says "$15M aggregate cap"). Faithfulness score: 0.67.

  • Citation Verification: The citation [chunk_4] for "$10M" is verified—chunk_4 contains "$15M," not "$10M." Citation accuracy: 0.50.

  • Overall hallucination score: 0.44 (threshold: 0.35). Hallucination detected.

Layer 3 — Correction: The regeneration agent receives feedback: "The claim 'liability cap is $10M' contradicts chunk_4 which states $15M." It regenerates: "Party B shall indemnify Party A against all third-party claims [chunk_2]. The aggregate liability cap under Section 8.3 is $15M [chunk_4]."

The corrected answer passes all three verification layers on the second pass. The system logs this incident, and the historical hallucination rate for this conversation drops from 12% to 8% over the next 50 queries as the memory-informed prompts improve.

Conclusion

Hallucination in LLMs is not a problem that can be solved with a single technique it requires a defense-in-depth architecture that prevents, detects, and corrects fabricated content at multiple stages of the pipeline. By implementing self-consistency voting, NLI-based fact-checking, and citation verification as independent detection layers within a LangGraph multi-agent workflow, enterprises can quantify hallucination risk with precision rather than relying on intuition. The correction loop ensures that detected hallucinations are fixed before reaching the user, while persistent memory enables continuous improvement over time. This three-layer firewall transforms LLMs from confident but unreliable generators into trustworthy enterprise tools that can be deployed in high-stakes domains like legal, healthcare, and finance with measurable confidence.