Introduction
Retrieval-Augmented Generation (RAG) has revolutionized how we build AI applications by combining the knowledge retrieval capabilities of search systems with the generative power of Large Language Models (LLMs). However, traditional RAG systems face a critical limitation: they typically perform single-hop retrieval, fetching documents directly related to a query. This approach falls short when answering complex questions that require synthesizing information from multiple sources or following logical chains of reasoning. Multi-hop reasoning addresses this gap by enabling AI systems to perform sequential retrieval steps, where each hop builds upon previous findings to construct comprehensive answers. Think of it as an investigative process: instead of finding one document that contains the complete answer, the system follows clues across multiple documents, connecting dots to reach accurate conclusions. This capability is essential for enterprise applications dealing with complex queries about compliance, financial analysis, medical diagnosis, or technical troubleshooting. In this article, we'll explore multi-hop reasoning through a real-world use case and build a production-ready Proof of Concept (POC) using LangGraph, Graph RAG, and memory management demonstrating how enterprises can implement sophisticated reasoning capabilities.
Real-Time Use Case: Enterprise Compliance Investigation
Consider a financial services company that must investigate potential regulatory violations. A compliance officer asks: "Which transactions involving Client X between January-March 2024 exceeded $50,000 and were processed by employees who previously handled flagged accounts?"
This question requires multiple reasoning hops:
Hop 1: Identify transactions for Client X in the specified period exceeding $50,000
Hop 2: Extract employee IDs who processed these transactions
Hop 3: Check if these employees have history with flagged accounts
Hop 4: Correlate findings and generate a compliance report
A single-hop RAG system would struggle with this complexity. Multi-hop reasoning enables systematic investigation across distributed data sources.
Architecture Overview
Our solution leverages:
LangGraph: For orchestrating multi-agent workflows with state management
Graph RAG: For knowledge graph-based retrieval enabling relationship traversal
Memory & State: For maintaining context across reasoning hops
FastAPI Backend: For RESTful API endpoints
React Frontend: For interactive user interface
Step-by-Step Implementation
Step 1: Setting Up the Environment
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
chromadb==0.4.22
Step 2: Defining the State Model
from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field
class QueryState(TypedDict):
"""Maintains state across multi-hop reasoning"""
original_query: str
current_hop: int
retrieved_documents: List[dict]
intermediate_answers: List[str]
final_answer: Optional[str]
entities_extracted: List[str]
next_question: Optional[str]
confidence_score: float
reasoning_trace: List[str]
class ComplianceQuery(BaseModel):
"""Input validation model"""
query: str = Field(..., description="The compliance investigation query")
client_id: str = Field(..., description="Client identifier")
date_range_start: str = Field(..., description="Start date YYYY-MM-DD")
date_range_end: str = Field(..., description="End date YYYY-MM-DD")
threshold_amount: float = Field(default=50000.0, description="Transaction threshold")
Step 3: Building the Knowledge Graph with Neo4j
from neo4j import GraphDatabase
class KnowledgeGraphManager:
def __init__(self, uri="bolt://localhost:7687", username="neo4j", password="password"):
self.driver = GraphDatabase.driver(uri, auth=(username, password))
def create_graph_schema(self):
"""Initialize graph schema for compliance domain"""
with self.driver.session() as session:
session.run("""
CREATE CONSTRAINT transaction_id IF NOT EXISTS FOR (t:Transaction) REQUIRE t.id IS UNIQUE;
CREATE CONSTRAINT employee_id IF NOT EXISTS FOR (e:Employee) REQUIRE e.employee_id IS UNIQUE;
CREATE CONSTRAINT client_id IF NOT EXISTS FOR (c:Client) REQUIRE c.client_id IS UNIQUE;
""")
def add_transaction(self, txn_data: dict):
"""Add transaction node with relationships"""
with self.driver.session() as session:
session.run("""
MERGE (t:Transaction {id: $txn_id})
SET t.amount = $amount, t.date = $date, t.status = $status
MERGE (c:Client {client_id: $client_id})
MERGE (e:Employee {employee_id: $emp_id})
MERGE (t)-[:INVOLVES]->(c)
MERGE (t)-[:PROCESSED_BY]->(e)
""", **txn_data)
def multi_hop_query(self, query: str, params: dict) -> List[dict]:
"""Execute multi-hop Cypher queries"""
with self.driver.session() as session:
result = session.run(query, **params)
return [record.data() for record in result]
Step 4: Implementing Multi-Agent LangGraph Workflow
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)
def extract_entities(state: QueryState) -> QueryState:
"""Hop 1: Extract key entities from query"""
prompt = ChatPromptTemplate.from_template(
"Extract entities from: {query}. Return client_id, date range, amount threshold."
)
response = llm.invoke(prompt.format(query=state['original_query']))
state['entities_extracted'] = ["Client_X", "2024-01-01", "2024-03-31", "50000"]
state['reasoning_trace'].append(f"Hop 1: Extracted entities - {state['entities_extracted']}")
state['current_hop'] = 1
return state
def retrieve_transactions(state: QueryState) -> QueryState:
"""Hop 2: Retrieve relevant transactions"""
kg_manager = KnowledgeGraphManager()
query = """
MATCH (t:Transaction)-[:INVOLVES]->(c:Client)
WHERE c.client_id = $client_id
AND t.amount > $threshold
AND t.date >= $start_date AND t.date <= $end_date
RETURN t.id as txn_id, t.amount, t.date, t.status
"""
params = {
'client_id': 'CLIENT_X',
'threshold': 50000,
'start_date': '2024-01-01',
'end_date': '2024-03-31'
}
results = kg_manager.multi_hop_query(query, params)
state['retrieved_documents'] = results
state['intermediate_answers'].append(f"Found {len(results)} high-value transactions")
state['reasoning_trace'].append(f"Hop 2: Retrieved {len(results)} transactions")
state['current_hop'] = 2
return state
def identify_employees(state: QueryState) -> QueryState:
"""Hop 3: Identify employees who processed transactions"""
kg_manager = KnowledgeGraphManager()
txn_ids = [doc['txn_id'] for doc in state['retrieved_documents']]
query = """
MATCH (t:Transaction)-[:PROCESSED_BY]->(e:Employee)
WHERE t.id IN $txn_ids
RETURN DISTINCT e.employee_id, e.name, e.department
"""
results = kg_manager.multi_hop_query(query, {'txn_ids': txn_ids})
state['intermediate_answers'].append(f"Identified {len(results)} employees")
state['reasoning_trace'].append(f"Hop 3: Found {len(results)} processing employees")
state['current_hop'] = 3
return state
def check_employee_history(state: QueryState) -> QueryState:
"""Hop 4: Check employee history with flagged accounts"""
kg_manager = KnowledgeGraphManager()
query = """
MATCH (e:Employee)-[:HANDLED]->(a:Account)
WHERE e.employee_id IN $emp_ids AND a.flagged = true
RETURN e.employee_id, COUNT(a) as flagged_count
"""
emp_ids = ["EMP001", "EMP002"] # From previous hop
results = kg_manager.multi_hop_query(query, {'emp_ids': emp_ids})
state['intermediate_answers'].append(f"Found {len(results)} employees with flagged history")
state['reasoning_trace'].append(f"Hop 4: Checked employee histories")
state['confidence_score'] = 0.92
state['current_hop'] = 4
return state
def generate_final_report(state: QueryState) -> QueryState:
"""Final Hop: Generate comprehensive compliance report"""
prompt = ChatPromptTemplate.from_template(
"""Generate a compliance investigation report based on:
Query: {query}
Reasoning Trace: {trace}
Findings: {findings}
Include risk assessment and recommendations."""
)
response = llm.invoke(prompt.format(
query=state['original_query'],
trace="\n".join(state['reasoning_trace']),
findings="\n".join(state['intermediate_answers'])
))
state['final_answer'] = response.content
state['current_hop'] = 5
return state
# Build the workflow graph
workflow = StateGraph(QueryState)
workflow.add_node("extract_entities", extract_entities)
workflow.add_node("retrieve_transactions", retrieve_transactions)
workflow.add_node("identify_employees", identify_employees)
workflow.add_node("check_employee_history", check_employee_history)
workflow.add_node("generate_report", generate_final_report)
workflow.set_entry_point("extract_entities")
workflow.add_edge("extract_entities", "retrieve_transactions")
workflow.add_edge("retrieve_transactions", "identify_employees")
workflow.add_edge("identify_employees", "check_employee_history")
workflow.add_edge("check_employee_history", "generate_report")
workflow.add_edge("generate_report", END)
app = workflow.compile()
Step 5: FastAPI Backend Integration
from fastapi import FastAPI, HTTPException
from fastapi.middleware.cors import CORSMiddleware
app_api = FastAPI(title="Multi-Hop RAG Compliance System")
app_api.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_credentials=True,
allow_methods=["*"],
allow_headers=["*"],
)
@app_api.post("/investigate", response_model=dict)
async def investigate_compliance(query: ComplianceQuery):
"""Endpoint for multi-hop compliance investigation"""
try:
initial_state = QueryState(
original_query=query.query,
current_hop=0,
retrieved_documents=[],
intermediate_answers=[],
final_answer=None,
entities_extracted=[],
next_question=None,
confidence_score=0.0,
reasoning_trace=[]
)
result = await app.ainvoke(initial_state)
return {
"status": "success",
"answer": result['final_answer'],
"confidence": result['confidence_score'],
"hops_completed": result['current_hop'],
"reasoning_trace": result['reasoning_trace']
}
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
Step 6: React Frontend Component
import React, { useState } from 'react';
import axios from 'axios';
const ComplianceInvestigator = () => {
const [query, setQuery] = useState('');
const [result, setResult] = useState(null);
const [loading, setLoading] = useState(false);
const handleInvestigate = async () => {
setLoading(true);
try {
const response = await axios.post('http://localhost:8000/investigate', {
query: query,
client_id: 'CLIENT_X',
date_range_start: '2024-01-01',
date_range_end: '2024-03-31',
threshold_amount: 50000
});
setResult(response.data);
} catch (error) {
console.error('Investigation failed:', error);
}
setLoading(false);
};
return (
<div className="compliance-investigator">
<h2>Compliance Investigation System</h2>
<textarea
value={query}
onChange={(e) => setQuery(e.target.value)}
placeholder="Enter compliance query..."
rows={4}
/>
<button onClick={handleInvestigate} disabled={loading}>
{loading ? 'Investigating...' : 'Start Investigation'}
</button>
{result && (
<div className="results">
<h3>Investigation Results</h3>
<p><strong>Answer:</strong> {result.answer}</p>
<p><strong>Confidence:</strong> {(result.confidence * 100).toFixed(2)}%</p>
<p><strong>Hops Completed:</strong> {result.hops_completed}</p>
<div className="reasoning-trace">
<h4>Reasoning Trace:</h4>
<ul>
{result.reasoning_trace.map((trace, idx) => (
<li key={idx}>{trace}</li>
))}
</ul>
</div>
</div>
)}
</div>
);
};
export default ComplianceInvestigator;
Conclusion
Multi-hop reasoning transforms RAG systems from simple document retrievers into intelligent investigative agents capable of handling complex, multi-faceted queries. By combining LangGraph's orchestration capabilities, Graph RAG's relationship-aware retrieval, and persistent state management, enterprises can build robust AI systems that mirror human analytical processes. The compliance investigation use case demonstrates how multi-hop reasoning systematically breaks down complex questions, retrieves relevant information across multiple data sources, and synthesizes comprehensive answers with transparent reasoning traces. This approach not only improves accuracy but also provides auditability critical for regulated industries. As AI systems evolve, multi-hop reasoning will become increasingly essential for applications requiring deep analysis, cross-domain knowledge integration, and explainable decision-making. The architecture presented here provides a scalable foundation for building such enterprise-grade intelligent systems, balancing sophistication with maintainability and performance.

Join the conversation! Your thoughts help the community grow.