Table of Contents

  1. Introduction: The Hallucination Challenge in Industrial AI

  2. Understanding Hallucination Types in RAG Systems

  3. Detection Strategies: Self-Correction and External Verification

  4. Solution Architecture: The "Verify-Then-Trust" LangGraph Workflow

  5. Technology Stack Overview

  6. Step-by-Step Implementation: Backend Development

    • Defining the State Schema with Verification Flags

    • Building the Retrieval Agent with Source Tracking

    • Implementing the Hallucination Detector Agent (LLM-as-a-Judge)

    • Creating the Correction/Refinement Agent

    • Constructing the Conditional LangGraph Workflow

  7. Frontend Implementation: Transparency UI with Citation Highlighting

  8. Real-Time Use Case: Pharmaceutical Compliance Assistant

  9. Conclusion: Building Trustworthy Enterprise AI

Introduction

In enterprise environments, particularly in regulated industries like healthcare, finance, and manufacturing, hallucinations—where an AI model generates plausible but factually incorrect information—are not just errors; they are liabilities. While Retrieval-Augmented Generation (RAG) significantly reduces hallucinations by grounding responses in external data, it does not eliminate them entirely. Models may still ignore retrieved context, misinterpret data, or fabricate connections between unrelated documents.

This article presents a robust, enterprise-grade solution using LangGraph to build a multi-agent RAG system that actively detects and fixes hallucinations. We implement a "Verify-Then-Trust" workflow where a dedicated agent acts as a factual auditor, ensuring every claim in the generated response is supported by retrieved evidence before it reaches the user.

Technology Tags

Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, pgvector, Pydantic, Docker, Redis, TypeScript, TailwindCSS, OpenAI API, Sentence Transformers

Step-by-Step Implementation

1. State Schema with Verification Tracking

from typing import List, Dict, Any, TypedDict, Optional
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage

class VerificationState(TypedDict):
    messages: List
    retrieved_documents: List[Dict]  # Contains content and metadata
    initial_draft: str
    verified_answer: str
    hallucination_detected: bool
    verification_feedback: str
    confidence_score: float
    citations: List[str]

2. Retrieval Agent with Source Tracking

from langchain_community.vectorstores import PGVector
from langchain_huggingface import HuggingFaceEmbeddings

class RetrievalAgent:
    def __init__(self, vector_store: PGVector):
        self.vector_store = vector_store
    
    def retrieve(self, state: VerificationState) -> VerificationState:
        """Retrieve relevant documents and store them with IDs"""
        query = state["messages"][-1].content
        docs = self.vector_store.similarity_search_with_score(query, k=5)
        
        retrieved = []
        for doc, score in docs:
            retrieved.append({
                "id": doc.metadata.get("doc_id"),
                "content": doc.page_content,
                "source": doc.metadata.get("source"),
                "relevance_score": score
            })
        
        state["retrieved_documents"] = retrieved
        return state

3. Initial Generation Agent

from langchain_openai import ChatOpenAI

class GeneratorAgent:
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def generate_draft(self, state: VerificationState) -> VerificationState:
        """Generate an initial answer based on retrieved context"""
        context = "\n".join([doc["content"] for doc in state["retrieved_documents"]])
        query = state["messages"][-1].content
        
        prompt = f"""
        Answer the question based ONLY on the following context. 
        If the answer is not in the context, state "Information not found".
        
        Context:
        {context}
        
        Question: {query}
        """
        
        response = self.llm.invoke(prompt)
        state["initial_draft"] = response.content
        return state

4. Hallucination Detector Agent (LLM-as-a-Judge)

class VerifierAgent:
    def __init__(self):
        self.judge_llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def verify_facts(self, state: VerificationState) -> VerificationState:
        """Check if the draft answer is supported by the retrieved documents"""
        draft = state["initial_draft"]
        docs = state["retrieved_documents"]
        context_text = "\n".join([f"[Doc {i+1}]: {d['content']}" for i, d in enumerate(docs)])
        
        prompt = f"""
        You are a factual auditor. Your job is to detect hallucinations.
        
        Retrieved Context:
        {context_text}
        
        Draft Answer:
        {draft}
        
        Instructions:
        1. Identify every claim in the Draft Answer.
        2. Check if each claim is explicitly supported by the Retrieved Context.
        3. If any claim is unsupported or contradicted, mark as HALLUCINATION.
        4. Provide specific feedback on which parts are unsupported.
        
        Output Format (JSON):
        {{
            "hallucination_detected": true/false,
            "confidence_score": 0.0-1.0,
            "feedback": "Specific details on unsupported claims...",
            "supported_citations": ["Doc 1", "Doc 3"]
        }}
        """
        
        response = self.judge_llm.invoke(prompt)
        # In production, parse JSON properly
        result = {
            "hallucination_detected": False, # Placeholder parsing
            "confidence_score": 0.9,
            "feedback": "",
            "supported_citations": ["Doc 1"]
        }
        
        state["hallucination_detected"] = result["hallucination_detected"]
        state["verification_feedback"] = result["feedback"]
        state["confidence_score"] = result["confidence_score"]
        state["citations"] = result["supported_citations"]
        
        return state

5. Correction Agent

class CorrectorAgent:
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def refine_answer(self, state: VerificationState) -> VerificationState:
        """Rewrite the answer to remove hallucinations based on feedback"""
        feedback = state["verification_feedback"]
        context = "\n".join([doc["content"] for doc in state["retrieved_documents"]])
        
        prompt = f"""
        The previous answer contained inaccuracies. 
        Feedback: {feedback}
        
        Please rewrite the answer strictly adhering to the context below.
        Remove any unsupported claims.
        
        Context:
        {context}
        """
        
        response = self.llm.invoke(prompt)
        state["verified_answer"] = response.content
        state["hallucination_detected"] = False  # Reset after correction
        return state

6. LangGraph Workflow Construction

def build_verification_graph():
    workflow = StateGraph(VerificationState)
    
    retriever = RetrievalAgent(None) # Pass actual DB
    generator = GeneratorAgent()
    verifier = VerifierAgent()
    corrector = CorrectorAgent()
    
    workflow.add_node("retrieve", retriever.retrieve)
    workflow.add_node("generate", generator.generate_draft)
    workflow.add_node("verify", verifier.verify_facts)
    workflow.add_node("correct", corrector.refine_answer)
    
    workflow.set_entry_point("retrieve")
    workflow.add_edge("retrieve", "generate")
    workflow.add_edge("generate", "verify")
    
    # Conditional Edge: If hallucination detected, go to correct; else end
    def route_after_verify(state: VerificationState):
        if state["hallucination_detected"]:
            return "correct"
        else:
            state["verified_answer"] = state["initial_draft"]
            return "end"
    
    workflow.add_conditional_edges(
        "verify",
        route_after_verify,
        {"correct": "correct", "end": END}
    )
    
    workflow.add_edge("correct", END)
    
    return workflow.compile()

7. FastAPI Backend

from fastapi import FastAPI

app = FastAPI()
graph = build_verification_graph()

@app.post("/ask")
async def ask_question(query: dict):
    initial_state = VerificationState(
        messages=[HumanMessage(content=query["text"])],
        retrieved_documents=[],
        initial_draft="",
        verified_answer="",
        hallucination_detected=False,
        verification_feedback="",
        confidence_score=0.0,
        citations=[]
    )
    
    result = graph.invoke(initial_state)
    
    return {
        "answer": result["verified_answer"],
        "is_verified": not result["hallucination_detected"],
        "confidence": result["confidence_score"],
        "sources": result["citations"]
    }

8. Frontend: Transparency UI

// components/VerifiedResponse.tsx
import React from 'react';

export const VerifiedResponse: React.FC<{ data: any }> = ({ data }) => {
  return (
    <div className="border-l-4 border-green-500 pl-4">
      <p>{data.answer}</p>
      <div className="mt-2 text-sm text-gray-600">
        <span className="font-bold">Verification Status: </span>
        {data.is_verified ? "  Verified" : "  Unverified"}
      </div>
      <div className="mt-1 text-xs text-blue-600">
        Sources: {data.sources.join(", ")}
      </div>
    </div>
  );
};

Real-Time Use Case: Pharmaceutical Compliance Assistant

A regulatory affairs manager asks: "What is the storage temperature requirement for Drug X?"

  1. Retrieval: System fetches the latest FDA label and internal manufacturing SOPs.

  2. Generation: LLM drafts: "Store at 2-8°C."

  3. Verification: The Judge Agent checks the draft against the retrieved PDFs. It finds that the new SOP updated the requirement to "2-10°C" due to recent stability studies.

  4. Detection: The Judge flags the "2-8°C" claim as outdated/hallucinated relative to the newest document.

  5. Correction: The Corrector Agent rewrites the answer to "Store at 2-10°C as per SOP-2024-09," citing the specific document.

  6. Result: The user receives accurate, compliant information with full traceability.

Conclusion

Hallucinations in RAG systems are mitigated not by hoping for better models, but by designing rigorous verification workflows. By implementing a multi-agent architecture with LangGraph, enterprises can introduce a "fact-checking" layer that acts as a safety gate. This approach transforms RAG from a probabilistic generator into a deterministic, auditable information retrieval system, essential for high-stakes industries.