Introduction

In the realm of Retrieval-Augmented Generation (RAG), the granularity of data retrieval is a critical architectural decision that directly impacts accuracy, latency, and cost. Most introductory tutorials focus exclusively on chunk-level retrieval, where documents are split into small, overlapping segments (e.g., 500 tokens) for vector embedding. While this approach maximizes semantic precision for specific queries, it often loses the broader context, leading to hallucinations or incomplete answers when the query requires understanding the document's overall structure or theme. Conversely, document-level retrieval treats entire files as single units. This preserves global context but suffers from the "needle in a haystack" problem, where relevant details are diluted by irrelevant content within large documents, and computational costs for embedding and processing skyrocket. The enterprise solution lies not in choosing one over the other, but in implementing a hybrid multi-agent strategy. By leveraging LangGraph for orchestration and Graph RAG for structural awareness, we can dynamically decide whether to retrieve a whole document for high-level summaries or specific chunks for detailed fact-checking. This article explores this dichotomy through a real-world legal compliance use case, providing a complete, working Proof of Concept (POC) with backend and frontend integration.

Real-Time Use Case: Legal Contract Analysis

Consider a corporate legal department managing thousands of vendor contracts. A lawyer might ask two distinct types of questions:

  1. Chunk-Level: "What is the termination notice period in Contract #123?" (Requires precise extraction from a specific clause).

  2. Document-Level: "Is Contract #123 primarily a service agreement or a licensing deal?" (Requires understanding the entire document’s intent and structure).

A naive RAG system might fail at the second question because no single chunk contains the "overall intent." Our system will use a multi-agent workflow to classify the query intent and route retrieval accordingly.

Architecture Overview

Step-by-Step Implementation

Step 1: Environment Setup and Dependencies

# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
chromadb==0.4.22
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
tiktoken==0.6.0

Step 2: Defining the State and Data Models

We need a state that tracks not just the answer, but how we retrieved it.

from typing import TypedDict, List, Optional, Literal
from pydantic import BaseModel, Field

class RetrievalState(TypedDict):
    query: str
    retrieval_strategy: Optional[Literal["chunk", "document", "hybrid"]]
    chunk_results: List[dict]
    document_metadata: Optional[dict]
    final_answer: Optional[str]
    reasoning_log: List[str]

class QueryRequest(BaseModel):
    query: str = Field(..., description="User's natural language query")
    document_id: str = Field(..., description="Target document identifier")

Step 3: Hybrid Retrieval Engine

This component handles both granular and holistic data access.

import chromadb
from neo4j import GraphDatabase
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4-turbo")

class HybridRetriever:
    def __init__(self):
        self.chroma_client = chromadb.PersistentClient(path="./chroma_db")
        self.collection = self.chroma_client.get_or_create_collection("legal_chunks")
        self.neo4j_driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

    def retrieve_chunks(self, query: str, doc_id: str, k=5) -> List[dict]:
        """Semantic search for specific details"""
        results = self.collection.query(
            query_texts=[query],
            where={"doc_id": doc_id},
            n_results=k
        )
        return [{"content": doc, "metadata": meta} for doc, meta in zip(results['documents'][0], results['metadatas'][0])]

    def retrieve_document_context(self, doc_id: str) -> dict:
        """Graph-based retrieval for holistic context"""
        with self.neo4j_driver.session() as session:
            result = session.run("""
                MATCH (d:Document {id: $doc_id})
                RETURN d.title as title, d.type as type, d.summary as summary, d.page_count as pages
            """, doc_id=doc_id)
            record = result.single()
            if record:
                return dict(record)
            return {}

Step 4: LangGraph Multi-Agent Workflow

The core intelligence lies in the router agent that decides the strategy.

from langgraph.graph import StateGraph, END

def classify_intent(state: RetrievalState) -> RetrievalState:
    """Agent 1: Decide if we need chunks or full doc context"""
    prompt = f"""
    Analyze the query: '{state['query']}'
    If it asks for specific clauses, dates, or numbers, return 'chunk'.
    If it asks about the general nature, summary, or type of document, return 'document'.
    Return only the word 'chunk' or 'document'.
    """
    response = llm.invoke(prompt).content.strip().lower()
    state['retrieval_strategy'] = response if response in ['chunk', 'document'] else 'hybrid'
    state['reasoning_log'].append(f"Intent classified as: {state['retrieval_strategy']}")
    return state

def execute_chunk_retrieval(state: RetrievalState) -> RetrievalState:
    """Agent 2: Fetch specific snippets"""
    retriever = HybridRetriever()
    # In a real app, you'd extract doc_id from state or query
    state['chunk_results'] = retriever.retrieve_chunks(state['query'], "DOC_123")
    state['reasoning_log'].append(f"Retrieved {len(state['chunk_results'])} chunks")
    return state

def execute_doc_retrieval(state: RetrievalState) -> RetrievalState:
    """Agent 3: Fetch holistic metadata"""
    retriever = HybridRetriever()
    state['document_metadata'] = retriever.retrieve_document_context("DOC_123")
    state['reasoning_log'].append("Retrieved document-level metadata")
    return state

def generate_answer(state: RetrievalState) -> RetrievalState:
    """Agent 4: Synthesize final response"""
    context = ""
    if state['retrieval_strategy'] == 'chunk':
        context = "\n".join([c['content'] for c in state['chunk_results']])
    else:
        context = str(state['document_metadata'])
    
    prompt = f"Query: {state['query']}\nContext: {context}\nAnswer:"
    state['final_answer'] = llm.invoke(prompt).content
    return state

# Build Graph
workflow = StateGraph(RetrievalState)
workflow.add_node("classify", classify_intent)
workflow.add_node("get_chunks", execute_chunk_retrieval)
workflow.add_node("get_doc", execute_doc_retrieval)
workflow.add_node("answer", generate_answer)

workflow.set_entry_point("classify")

# Conditional Edges
def route_retrieval(state: RetrievalState):
    if state['retrieval_strategy'] == 'chunk':
        return "get_chunks"
    else:
        return "get_doc"

workflow.add_conditional_edges("classify", route_retrieval, {"get_chunks": "get_chunks", "get_doc": "get_doc"})
workflow.add_edge("get_chunks", "answer")
workflow.add_edge("get_doc", "answer")
workflow.add_edge("answer", END)

app = workflow.compile()

Step 5: FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

api_app = FastAPI(title="Hybrid RAG API")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])

@api_app.post("/query")
async def handle_query(request: QueryRequest):
    initial_state = RetrievalState(
        query=request.query,
        retrieval_strategy=None,
        chunk_results=[],
        document_metadata=None,
        final_answer=None,
        reasoning_log=[]
    )
    result = await app.ainvoke(initial_state)
    return {
        "answer": result['final_answer'],
        "strategy_used": result['retrieval_strategy'],
        "trace": result['reasoning_log']
    }

Step 6: React Frontend Component

import React, { useState } from 'react';
import axios from 'axios';

const RagInterface = () => {
  const [query, setQuery] = useState('');
  const [response, setResponse] = useState(null);

  const submitQuery = async () => {
    const res = await axios.post('http://localhost:8000/query', {
      query,
      document_id: 'DOC_123'
    });
    setResponse(res.data);
  };

  return (
    <div className="p-4 max-w-2xl mx-auto">
      <h1 className="text-2xl font-bold mb-4">Enterprise Legal RAG</h1>
      <input 
        className="border p-2 w-full mb-4" 
        placeholder="Ask about a contract..." 
        value={query}
        onChange={(e) => setQuery(e.target.value)}
      />
      <button onClick={submitQuery} className="bg-blue-500 text-white p-2 rounded">
        Analyze
      </button>
      
      {response && (
        <div className="mt-4 p-4 border rounded bg-gray-50">
          <p><strong>Strategy:</strong> {response.strategy_used}</p>
          <p className="mt-2"><strong>Answer:</strong> {response.answer}</p>
          <div className="mt-2 text-sm text-gray-600">
            <strong>Trace:</strong> {response.trace.join(' -> ')}
          </div>
        </div>
      )}
    </div>
  );
};
export default RagInterface;

Conclusion

Implementing both document-level and chunk-level retrieval is not merely a technical enhancement; it is a necessity for enterprise-grade AI. By using LangGraph to orchestrate these strategies, we ensure that the system is both precise enough to find a specific clause and smart enough to understand the document's overall purpose. This hybrid approach, supported by Graph RAG for structural context and memory for state tracking, provides a robust foundation for building trustworthy, explainable, and highly effective AI applications in complex domains like law, finance, and healthcare.