Introduction
Chunking is the silent architect of RAG performance. A poorly chosen chunking strategy can render even the most sophisticated embedding models and rerankers useless—splitting sentences mid-thought, breaking semantic units across boundaries, or producing chunks so small they lack context. Conversely, the right chunking strategy can dramatically improve retrieval precision, reduce hallucinations, and preserve the logical flow of source material.
Yet most RAG implementations default to naive fixed-size chunking without considering document type, query patterns, or downstream use cases. This article presents an enterprise-grade multi-agent LangGraph system that acts as a "Chunking Strategy Advisor"—analyzing document characteristics, evaluating multiple chunking strategies against quantified trade-offs, and recommending the optimal approach with persistent memory of past decisions.
The Core Chunking Strategies
1. Fixed-Size Chunking: Splits documents into equal-sized character or token windows with optional overlap. Simple, fast, but semantically blind.
2. Recursive Character Chunking: Hierarchically splits by paragraphs → sentences → words using a list of separators. Preserves natural boundaries better than fixed-size.
3. Sentence-Based Chunking: Groups complete sentences into chunks, ensuring no sentence is split mid-way. Good for conversational or narrative content.
4. Semantic Chunking: Uses embedding similarity to detect topic shifts and split at semantic boundaries. Computationally expensive but preserves meaning.
5. Document-Structure Chunking: Respects headings, sections, tables, and lists based on document parsing (Markdown, HTML, PDF structure). Ideal for technical docs.
6. Parent-Child (Small-to-Big) Chunking: Retrieves small chunks for precision but returns their larger parent context for generation. Balances recall and context richness.
7. Late Chunking (Jina-style): Embeds the full document first, then pools token embeddings into chunk-level vectors. Preserves long-range context in embeddings.
Trade-Off Matrix
Strategy | Precision | Context Preservation | Cost | Best For |
|---|---|---|---|---|
Fixed-Size | Low | Low | Very Low | Quick prototypes, homogeneous text |
Recursive | Medium | Medium | Low | General-purpose documents |
Sentence-Based | Medium | Medium-High | Low | Narratives, transcripts |
Semantic | High | High | High | Research papers, mixed-topic docs |
Structure-Aware | High | High | Medium | Technical docs, manuals, reports |
Parent-Child | High | Very High | Medium | QA over long documents |
Late Chunking | High | Very High | High | Long-context retrieval |
Step-by-Step Implementation
1. Chunking Strategy State Schema
from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
from datetime import datetime
ChunkingStrategy = Literal[
"fixed_size", "recursive", "sentence_based",
"semantic", "structure_aware", "parent_child", "late_chunking"
]
class ChunkingState(TypedDict):
messages: List
conversation_id: str
document_id: str
# Input
document_text: str
document_type: str # contract, manual, article, transcript, code
document_length: int
has_structure: bool # headings, sections, tables
query_patterns: List[str] # expected query types
# Analysis outputs
document_characteristics: Dict[str, Any]
strategy_scores: Dict[ChunkingStrategy, float]
trade_offs: Dict[ChunkingStrategy, Dict[str, str]]
# Recommendation
recommended_strategy: ChunkingStrategy
reasoning: str
# Execution
chunks: List[Dict]
chunk_stats: Dict[str, Any]
# Memory
past_decisions: List[Dict]
2. Document Analyzer Agent
from langchain_openai import ChatOpenAI
import re
class DocumentAnalyzerAgent:
"""Analyzes document characteristics to inform strategy selection"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
def analyze(self, state: ChunkingState) -> ChunkingState:
text = state["document_text"]
# Structural analysis
heading_count = len(re.findall(r'^#{1,6}\s|^[A-Z][^.!?\n]{2,50}$', text, re.MULTILINE))
paragraph_count = len([p for p in text.split('\n\n') if len(p.strip()) > 50])
sentence_count = len(re.split(r'[.!?]+', text))
avg_sentence_length = len(text.split()) / max(sentence_count, 1)
# Detect document type signals
has_clauses = bool(re.search(r'Section\s+\d+|Clause\s+\d+|Article\s+\d+', text, re.I))
has_code_blocks = '```' in text or re.search(r'def\s+\w+\(|class\s+\w+:', text)
has_tables = '|' in text and '---' in text
characteristics = {
"length_tokens": len(text.split()) * 1.3,
"heading_count": heading_count,
"paragraph_count": paragraph_count,
"sentence_count": sentence_count,
"avg_sentence_length": avg_sentence_length,
"has_clauses": has_clauses,
"has_code_blocks": has_code_blocks,
"has_tables": has_tables,
"has_structure": state["has_structure"] or heading_count > 3,
"topic_density": sentence_count / max(paragraph_count, 1)
}
state["document_characteristics"] = characteristics
state["document_length"] = len(text)
return state
3. Multiple Chunking Strategy Engines
from langchain_text_splitters import (
RecursiveCharacterTextSplitter,
CharacterTextSplitter
)
from langchain_experimental.text_splitter import SemanticChunker
from langchain_huggingface import HuggingFaceEmbeddings
class ChunkingEngines:
"""Implements all chunking strategies"""
def __init__(self):
self.embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
def fixed_size(self, text: str, chunk_size: int = 500, overlap: int = 50) -> List[str]:
splitter = CharacterTextSplitter(
chunk_size=chunk_size, chunk_overlap=overlap, separator=""
)
return splitter.split_text(text)
def recursive(self, text: str, chunk_size: int = 800, overlap: int = 100) -> List[str]:
splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size, chunk_overlap=overlap,
separators=["\n\n", "\n", ". ", " ", ""]
)
return splitter.split_text(text)
def sentence_based(self, text: str, sentences_per_chunk: int = 5) -> List[str]:
import nltk
nltk.download('punkt', quiet=True)
sentences = nltk.sent_tokenize(text)
chunks = []
for i in range(0, len(sentences), sentences_per_chunk):
chunk = " ".join(sentences[i:i + sentences_per_chunk])
chunks.append(chunk)
return chunks
def semantic(self, text: str, breakpoint_threshold: float = 0.5) -> List[str]:
splitter = SemanticChunker(
embeddings=self.embeddings,
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=breakpoint_threshold * 100
)
return [doc.page_content for doc in splitter.split_text(text)]
def structure_aware(self, text: str) -> List[str]:
"""Split by document headings/sections"""
sections = re.split(r'\n(?=#{1,6}\s|[A-Z][^.!?\n]{2,80}\n)', text)
return [s.strip() for s in sections if len(s.strip()) > 50]
def parent_child(self, text: str, child_size: int = 300, parent_size: int = 1200):
"""Returns (parent_chunks, child_to_parent_mapping)"""
parent_splitter = RecursiveCharacterTextSplitter(
chunk_size=parent_size, chunk_overlap=100
)
child_splitter = RecursiveCharacterTextSplitter(
chunk_size=child_size, chunk_overlap=30
)
parents = parent_splitter.split_text(text)
mapping = []
all_children = []
for i, parent in enumerate(parents):
children = child_splitter.split_text(parent)
for child in children:
all_children.append(child)
mapping.append({"child_idx": len(all_children) - 1, "parent_idx": i})
return parents, all_children, mapping
def execute(self, strategy: ChunkingStrategy, text: str) -> Dict:
if strategy == "fixed_size":
chunks = self.fixed_size(text)
elif strategy == "recursive":
chunks = self.recursive(text)
elif strategy == "sentence_based":
chunks = self.sentence_based(text)
elif strategy == "semantic":
chunks = self.semantic(text)
elif strategy == "structure_aware":
chunks = self.structure_aware(text)
elif strategy == "parent_child":
parents, children, mapping = self.parent_child(text)
return {
"chunks": children,
"parents": parents,
"mapping": mapping,
"strategy": strategy
}
else:
chunks = self.recursive(text) # fallback
return {
"chunks": chunks,
"strategy": strategy,
"parents": None,
"mapping": None
}
4. Strategy Evaluator and Trade-Off Analyzer
class StrategyEvaluatorAgent:
"""Scores each strategy based on document characteristics"""
def evaluate(self, state: ChunkingState) -> ChunkingState:
chars = state["document_characteristics"]
scores = {}
# Fixed-size: good for homogeneous, bad for structured
scores["fixed_size"] = 30 if chars["has_structure"] else 60
# Recursive: general-purpose baseline
scores["recursive"] = 65
# Sentence-based: good for narrative, low topic density
if chars["topic_density"] < 2 and not chars["has_structure"]:
scores["sentence_based"] = 75
else:
scores["sentence_based"] = 50
# Semantic: good for mixed-topic, expensive
if chars["topic_density"] > 3:
scores["semantic"] = 85
else:
scores["semantic"] = 55
# Structure-aware: excellent when structure exists
if chars["has_structure"]:
scores["structure_aware"] = 90
else:
scores["structure_aware"] = 30
# Parent-child: good for long docs with QA patterns
if chars["length_tokens"] > 5000:
scores["parent_child"] = 80
else:
scores["parent_child"] = 50
# Late chunking: good for very long docs
if chars["length_tokens"] > 10000:
scores["late_chunking"] = 85
else:
scores["late_chunking"] = 40
# Domain-specific boosts
if chars["has_clauses"]:
scores["structure_aware"] = min(100, scores["structure_aware"] + 10)
if chars["has_code_blocks"]:
scores["recursive"] = min(100, scores["recursive"] + 15)
state["strategy_scores"] = scores
return state
class TradeOffAnalyzerAgent:
"""Documents pros/cons of each strategy for the given document"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
def analyze(self, state: ChunkingState) -> ChunkingState:
chars = state["document_characteristics"]
trade_offs = {
"fixed_size": {
"pros": "Fast, predictable chunk sizes, simple implementation",
"cons": "Breaks sentences, ignores semantics, poor for structured docs"
},
"recursive": {
"pros": "Respects natural boundaries, good general-purpose choice",
"cons": "May still split semantic units, no topic awareness"
},
"sentence_based": {
"pros": "Never splits sentences, preserves narrative flow",
"cons": "Uneven chunk sizes, poor for dense technical content"
},
"semantic": {
"pros": "Preserves topic coherence, detects meaning shifts",
"cons": "Expensive (embeddings per sentence), slower processing"
},
"structure_aware": {
"pros": "Respects document hierarchy, excellent for technical docs",
"cons": "Requires well-formed structure, fails on plain text"
},
"parent_child": {
"pros": "Best of both worlds—precise retrieval, rich context",
"cons": "Complex indexing, higher storage costs"
},
"late_chunking": {
"pros": "Preserves long-range context in embeddings",
"cons": "Requires specialized models, limited ecosystem support"
}
}
state["trade_offs"] = trade_offs
return state
5. Recommendation and Execution Agents
class RecommendationAgent:
"""Selects optimal strategy and explains reasoning"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
def recommend(self, state: ChunkingState) -> ChunkingState:
scores = state["strategy_scores"]
best_strategy = max(scores, key=scores.get)
state["recommended_strategy"] = best_strategy
prompt = f"""Explain why {best_strategy} chunking is optimal for this document.
Document characteristics:
- Length: {state['document_characteristics']['length_tokens']:.0f} tokens
- Has structure: {state['document_characteristics']['has_structure']}
- Has clauses: {state['document_characteristics']['has_clauses']}
- Topic density: {state['document_characteristics']['topic_density']:.2f}
Strategy scores: {scores}
Provide a 2-3 sentence reasoning."""
response = self.llm.invoke(prompt)
state["reasoning"] = response.content
return state
class ChunkingExecutorAgent:
"""Executes the recommended chunking strategy"""
def __init__(self):
self.engines = ChunkingEngines()
def execute(self, state: ChunkingState) -> ChunkingState:
result = self.engines.execute(
state["recommended_strategy"],
state["document_text"]
)
chunks = result["chunks"]
state["chunks"] = [
{"id": f"chunk_{i}", "content": c, "metadata": {
"strategy": result["strategy"],
"length": len(c),
"parent_id": None
}}
for i, c in enumerate(chunks)
]
# Add parent info if parent_child strategy
if result["parents"] and result["mapping"]:
for m in result["mapping"]:
state["chunks"][m["child_idx"]]["metadata"]["parent_id"] = f"parent_{m['parent_idx']}"
# Compute stats
lengths = [len(c["content"].split()) for c in state["chunks"]]
state["chunk_stats"] = {
"total_chunks": len(state["chunks"]),
"avg_length_words": sum(lengths) / max(len(lengths), 1),
"min_length_words": min(lengths) if lengths else 0,
"max_length_words": max(lengths) if lengths else 0,
"strategy_used": result["strategy"]
}
return state
6. LangGraph Workflow with Memory
class ChunkingMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def save_decision(self, state: ChunkingState):
record = {
"document_type": state["document_type"],
"strategy": state["recommended_strategy"],
"scores": state["strategy_scores"],
"chunk_count": state["chunk_stats"]["total_chunks"],
"timestamp": datetime.now().isoformat()
}
key = f"chunking:{state['document_type']}"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-50:]))
def load_history(self, document_type: str) -> List[Dict]:
return json.loads(self.redis.get(f"chunking:{document_type}") or "[]")
def build_chunking_advisor():
workflow = StateGraph(ChunkingState)
analyzer = DocumentAnalyzerAgent()
evaluator = StrategyEvaluatorAgent()
trade_off = TradeOffAnalyzerAgent()
recommender = RecommendationAgent()
executor = ChunkingExecutorAgent()
workflow.add_node("analyze_document", analyzer.analyze)
workflow.add_node("evaluate_strategies", evaluator.evaluate)
workflow.add_node("analyze_trade_offs", trade_off.analyze)
workflow.add_node("recommend_strategy", recommender.recommend)
workflow.add_node("execute_chunking", executor.execute)
workflow.set_entry_point("analyze_document")
workflow.add_edge("analyze_document", "evaluate_strategies")
workflow.add_edge("evaluate_strategies", "analyze_trade_offs")
workflow.add_edge("analyze_trade_offs", "recommend_strategy")
workflow.add_edge("recommend_strategy", "execute_chunking")
workflow.add_edge("execute_chunking", END)
return workflow.compile()
7. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="Chunking Strategy Advisor API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_chunking_advisor()
memory = ChunkingMemory(redis.Redis())
class ChunkingRequest(BaseModel):
document_id: str
document_text: str
document_type: str = "general"
has_structure: bool = False
query_patterns: List[str] = []
@app.post("/advise_chunking")
async def advise_chunking(req: ChunkingRequest):
history = memory.load_history(req.document_type)
initial_state = ChunkingState(
messages=[HumanMessage(content=f"Analyze chunking for {req.document_type}")],
conversation_id=f"conv_{req.document_id}",
document_id=req.document_id,
document_text=req.document_text,
document_type=req.document_type,
document_length=len(req.document_text),
has_structure=req.has_structure,
query_patterns=req.query_patterns,
document_characteristics={},
strategy_scores={},
trade_offs={},
recommended_strategy="recursive",
reasoning="",
chunks=[],
chunk_stats={},
past_decisions=history
)
result = graph.invoke(initial_state)
memory.save_decision(result)
return {
"recommended_strategy": result["recommended_strategy"],
"reasoning": result["reasoning"],
"strategy_scores": result["strategy_scores"],
"trade_offs": result["trade_offs"],
"chunk_stats": result["chunk_stats"],
"sample_chunks": result["chunks"][:5],
"document_characteristics": result["document_characteristics"]
}
8. Frontend: Interactive Chunking Dashboard
// components/ChunkingAdvisor.tsx
import React, { useState } from 'react';
export const ChunkingAdvisor: React.FC = () => {
const [docText, setDocText] = useState(`# Section 1: Definitions
"Agreement" means this Master Services Agreement.
"Confidential Information" means all non-public data...
# Section 2: Scope of Services
Provider shall deliver the Services as described in Exhibit A.
Any changes to scope require written amendment...
# Section 3: Payment Terms
Client shall pay invoices within 30 days of receipt.
Late payments accrue interest at 1.5% per month...`);
const [docType, setDocType] = useState('contract');
const [result, setResult] = useState<any>(null);
const handleSubmit = async () => {
const response = await fetch('http://localhost:8000/advise_chunking', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
document_id: 'contract-001',
document_text: docText,
document_type: docType,
has_structure: true
})
});
setResult(await response.json());
};
return (
<div className="p-6 max-w-6xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-6">🔪 Chunking Strategy Advisor</h1>
<div className="grid grid-cols-2 gap-6">
<div className="bg-white p-4 rounded-lg shadow">
<h2 className="font-bold mb-3">Document Input</h2>
<label className="text-sm">Document Type:</label>
<select className="w-full p-2 border rounded mb-3"
value={docType} onChange={e => setDocType(e.target.value)}>
<option value="contract">Legal Contract</option>
<option value="manual">Technical Manual</option>
<option value="article">Article/Blog</option>
<option value="transcript">Transcript</option>
<option value="code">Code Repository</option>
</select>
<textarea
className="w-full p-2 border rounded"
rows={12}
value={docText}
onChange={e => setDocText(e.target.value)}
/>
<button onClick={handleSubmit} className="mt-3 bg-blue-600 text-white px-6 py-2 rounded">
Analyze & Recommend
</button>
</div>
{result && (
<div className="space-y-4">
<div className="bg-gradient-to-r from-blue-600 to-purple-600 text-white p-4 rounded-lg shadow">
<h2 className="text-xl font-bold">
Recommended: {result.recommended_strategy.replace(/_/g, ' ').toUpperCase()}
</h2>
<p className="mt-2 text-sm">{result.reasoning}</p>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Strategy Scores</h3>
{Object.entries(result.strategy_scores).map(([k, v]: any) => (
<div key={k} className="flex items-center gap-2 mb-1">
<span className="text-sm w-32">{k.replace(/_/g, ' ')}</span>
<div className="flex-1 bg-gray-200 rounded-full h-2">
<div className={`h-2 rounded-full ${
k === result.recommended_strategy ? 'bg-green-500' : 'bg-blue-400'
}`} style={{width: `${v}%`}} />
</div>
<span className="text-sm font-mono w-10">{v}</span>
</div>
))}
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Chunk Statistics</h3>
<div className="grid grid-cols-2 gap-2 text-sm">
<div>Total chunks: <b>{result.chunk_stats.total_chunks}</b></div>
<div>Avg length: <b>{result.chunk_stats.avg_length_words.toFixed(0)} words</b></div>
<div>Min length: <b>{result.chunk_stats.min_length_words} words</b></div>
<div>Max length: <b>{result.chunk_stats.max_length_words} words</b></div>
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-2">Sample Chunks</h3>
{result.sample_chunks.map((c: any, i: number) => (
<div key={i} className="border-l-4 border-blue-500 pl-3 mb-2 text-sm">
<div className="text-xs text-gray-500">{c.id} ({c.metadata.length} chars)</div>
<div className="text-gray-700">{c.content.slice(0, 150)}...</div>
</div>
))}
</div>
</div>
)}
</div>
</div>
);
};
Real-Time Use Case: Legal Document Intelligence Platform
A law firm ingests a 50-page M&A agreement. The Chunking Advisor analyzes it:
Document Analysis:
Length: ~15,000 tokens
Has structure: Yes (12 sections with headings)
Has clauses: Yes (Section 1, Section 2, etc.)
Topic density: 1.8 (low—each section is self-contained)
Strategy Scores:
fixed_size: 30 (breaks clauses mid-sentence)
recursive: 65 (decent but ignores structure)
sentence_based: 50 (uneven section sizes)
semantic: 55 (overkill for well-structured doc)
structure_aware: 95 ← Winner
parent_child: 80 (good but unnecessary overhead)
late_chunking: 40 (not needed)
Recommendation: Structure-aware chunking, because the contract has clear section headings that define natural semantic boundaries. Each section becomes a chunk, preserving clause integrity.
Result: 12 chunks (one per section), average 1,200 words each. A query about "indemnification obligations" retrieves the exact Section 8 chunk—not a random 500-character window that cuts mid-sentence.
When the same firm later processes a 200-page deposition transcript (no structure, high narrative flow), the advisor recommends sentence-based chunking with 5 sentences per chunk—demonstrating how the system adapts to document characteristics.
Conclusion
Chunking is not a one-size-fits-all preprocessing step it's a strategic decision that directly impacts retrieval quality, context preservation, and downstream generation accuracy. By encoding chunking expertise into a multi-agent LangGraph advisor, enterprises can automatically select the optimal strategy based on document characteristics, quantify trade-offs, and maintain institutional memory of what works for each document type. The system transforms chunking from an afterthought into a first-class architectural decision, ensuring that every document in your RAG pipeline is segmented in the way that best serves your retrieval and generation goals. As document types multiply and use cases diversify, this adaptive approach becomes essential for maintaining RAG performance at enterprise scale.

Join the conversation! Your thoughts help the community grow.