Introduction
Debugging traditional software is deterministic: if a function returns 5 instead of 6, you trace the logic. Debugging Large Language Model (LLM) outputs, however, is probabilistic and non-deterministic. An LLM might give a perfect answer one time and hallucinate the next, even with the same input. In enterprise environments, where accuracy is paramount, "it works most of the time" is not acceptable. Effective LLM debugging requires observability. It involves tracing the entire lifecycle of a request from the initial prompt and retrieved context to the intermediate reasoning steps and final generation. Without this visibility, fixing issues like hallucinations, bias, or formatting errors is akin to guessing in the dark.
This article demonstrates how to build a debuggable, enterprise-grade AI system using LangGraph for transparent state tracking, Graph RAG for verifiable data retrieval, and a custom Observability Dashboard built with FastAPI and React. We will create a system that doesn't just give an answer but provides a full "audit trail" of how that answer was derived, allowing developers to pinpoint exactly where the pipeline failed.
Real-Time Use Case: Financial Compliance Analyst
Consider a bank using an AI agent to analyze transaction logs for potential money laundering. The agent must:
Retrieve transaction history from a secure database.
Check against regulatory rules stored in a knowledge graph.
Generate a risk assessment report.
If the agent flags a legitimate transaction as high-risk, compliance officers need to know why. Did it retrieve the wrong rule? Did it misinterpret the transaction amount? Or did the LLM simply hallucinate a violation? Our system will expose every step of this process for debugging.
Architecture Overview
LangGraph Orchestrator: Breaks the task into discrete nodes (Retrieve, Analyze, Report), maintaining a strict state object that records inputs and outputs at each stage.
Graph RAG (Neo4j): Provides structured, traceable context. Unlike vector search, graph relationships allow us to see exactly which rule node was connected to the decision.
Observability State: A dedicated part of the LangGraph state that logs timestamps, token usage, and raw data for every step.
FastAPI Backend: Exposes both the chat endpoint and a
/debug-traceendpoint for retrieving execution logs.React Frontend: A developer-focused dashboard that visualizes the execution graph and allows inspection of intermediate states.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
uuid==1.30
Step 2: Defining the Observable State
We extend the standard state to include a detailed execution_trace.
from typing import TypedDict, List, Optional, Any
from pydantic import BaseModel, Field
import time
import uuid
class DebugState(TypedDict):
query: str
session_id: str
retrieved_rules: List[dict]
analysis_reasoning: str
final_report: str
execution_trace: List[dict] # Stores step-by-step debug info
error_log: List[str]
class ComplianceQuery(BaseModel):
transaction_id: str = Field(..., description="ID of the transaction to analyze")
user_id: str = Field(..., description="Analyst ID")
Step 3: Building Traceable Nodes
Each node in our LangGraph workflow will append its activity to the execution_trace.
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
from neo4j import GraphDatabase
llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
def add_trace(state: DebugState, step_name: str, data: Any):
"""Helper to append debug info"""
state['execution_trace'].append({
"step": step_name,
"timestamp": time.time(),
"data_snapshot": str(data)[:500] # Truncate for performance
})
return state
def retrieve_rules_node(state: DebugState) -> DebugState:
"""Retrieves relevant compliance rules from Neo4j"""
try:
with driver.session() as session:
result = session.run("""
MATCH (r:Rule)-[:APPLIES_TO]->(t:TransactionType {name: 'Wire Transfer'})
RETURN r.description, r.threshold
""")
rules = [record.data() for record in result]
state['retrieved_rules'] = rules
state = add_trace(state, "Retrieval", rules)
except Exception as e:
state['error_log'].append(f"Retrieval Error: {str(e)}")
return state
def analyze_transaction_node(state: DebugState) -> DebugState:
"""LLM analyzes transaction against retrieved rules"""
context = str(state['retrieved_rules'])
prompt = f"Rules: {context}\nAnalyze transaction {state['query']} for violations."
response = llm.invoke(prompt)
state['analysis_reasoning'] = response.content
state = add_trace(state, "Analysis", response.content)
return state
def generate_report_node(state: DebugState) -> DebugState:
"""Formats the final compliance report"""
state['final_report'] = f"Compliance Report:\n{state['analysis_reasoning']}"
state = add_trace(state, "Generation", state['final_report'])
return state
Step 4: Orchestrating with LangGraph
workflow = StateGraph(DebugState)
workflow.add_node("retrieve", retrieve_rules_node)
workflow.add_node("analyze", analyze_transaction_node)
workflow.add_node("report", generate_report_node)
workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "analyze")
workflow.add_edge("analyze", "report")
workflow.add_edge("report", END)
app = workflow.compile()
Step 5: FastAPI Backend with Debug Endpoints
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
api_app = FastAPI(title="Debuggable Compliance AI")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])
# In-memory store for traces (use Redis/DB in production)
trace_store = {}
@api_app.post("/analyze")
async def analyze_transaction(query: ComplianceQuery):
initial_state = DebugState(
query=query.transaction_id,
session_id=query.user_id,
retrieved_rules=[],
analysis_reasoning="",
final_report="",
execution_trace=[],
error_log=[]
)
result = await app.ainvoke(initial_state)
trace_id = str(uuid.uuid4())
trace_store[trace_id] = result['execution_trace']
return {
"report": result['final_report'],
"trace_id": trace_id,
"errors": result['error_log']
}
@api_app.get("/debug/{trace_id}")
async def get_debug_trace(trace_id: str):
return trace_store.get(trace_id, {"error": "Trace not found"})
Step 6: React Debug Dashboard
This frontend component allows developers to inspect the internal state of the AI.
import React, { useState } from 'react';
import axios from 'axios';
const DebugDashboard = () => {
const [txnId, setTxnId] = useState('');
const [result, setResult] = useState(null);
const [trace, setTrace] = useState(null);
const runAnalysis = async () => {
const res = await axios.post('http://localhost:8000/analyze', {
transaction_id: txnId,
user_id: 'ANALYST_01'
});
setResult(res.data);
// Fetch detailed trace
const traceRes = await axios.get(`http://localhost:8000/debug/${res.data.trace_id}`);
setTrace(traceRes.data);
};
return (
<div className="p-6 max-w-4xl mx-auto">
<h1 className="text-2xl font-bold mb-4">Compliance AI Debugger</h1>
<input
className="border p-2 mr-2"
placeholder="Transaction ID"
value={txnId}
onChange={e => setTxnId(e.target.value)}
/>
<button onClick={runAnalysis} className="bg-red-600 text-white px-4 py-2 rounded">Analyze & Debug</button>
{result && (
<div className="mt-6 grid grid-cols-2 gap-4">
<div className="p-4 border rounded">
<h3 className="font-bold">Final Report</h3>
<p className="whitespace-pre-wrap">{result.report}</p>
</div>
<div className="p-4 border rounded bg-gray-50">
<h3 className="font-bold">Execution Trace</h3>
{trace && trace.map((step, idx) => (
<div key={idx} className="mb-2 text-sm">
<span className="font-semibold text-blue-600">{step.step}:</span>
<pre className="bg-white p-1 mt-1 border overflow-x-auto">
{step.data_snapshot}
</pre>
</div>
))}
</div>
</div>
)}
</div>
);
};
export default DebugDashboard;
Conclusion
Debugging LLM outputs is no longer about checking a single return value; it is about auditing a complex, multi-step reasoning process. By leveraging LangGraph’s stateful architecture, we can expose the "black box" of AI, providing visibility into retrieval quality, reasoning logic, and generation parameters. This level of observability is essential for enterprise adoption, allowing teams to move from "trusting the AI" to "verifying the AI," ensuring that compliance, financial, and medical applications remain accurate, safe, and reliable.

Join the conversation! Your thoughts help the community grow.