Introduction

Chunking is the silent architect of RAG performance. A poorly chosen chunking strategy can render even the most sophisticated embedding models and rerankers useless—splitting sentences mid-thought, breaking semantic units across boundaries, or producing chunks so small they lack context. Conversely, the right chunking strategy can dramatically improve retrieval precision, reduce hallucinations, and preserve the logical flow of source material.

Yet most RAG implementations default to naive fixed-size chunking without considering document type, query patterns, or downstream use cases. This article presents an enterprise-grade multi-agent LangGraph system that acts as a "Chunking Strategy Advisor"—analyzing document characteristics, evaluating multiple chunking strategies against quantified trade-offs, and recommending the optimal approach with persistent memory of past decisions.

The Core Chunking Strategies

1. Fixed-Size Chunking: Splits documents into equal-sized character or token windows with optional overlap. Simple, fast, but semantically blind.

2. Recursive Character Chunking: Hierarchically splits by paragraphs → sentences → words using a list of separators. Preserves natural boundaries better than fixed-size.

3. Sentence-Based Chunking: Groups complete sentences into chunks, ensuring no sentence is split mid-way. Good for conversational or narrative content.

4. Semantic Chunking: Uses embedding similarity to detect topic shifts and split at semantic boundaries. Computationally expensive but preserves meaning.

5. Document-Structure Chunking: Respects headings, sections, tables, and lists based on document parsing (Markdown, HTML, PDF structure). Ideal for technical docs.

6. Parent-Child (Small-to-Big) Chunking: Retrieves small chunks for precision but returns their larger parent context for generation. Balances recall and context richness.

7. Late Chunking (Jina-style): Embeds the full document first, then pools token embeddings into chunk-level vectors. Preserves long-range context in embeddings.

Trade-Off Matrix

Strategy

Precision

Context Preservation

Cost

Best For

Fixed-Size

Low

Low

Very Low

Quick prototypes, homogeneous text

Recursive

Medium

Medium

Low

General-purpose documents

Sentence-Based

Medium

Medium-High

Low

Narratives, transcripts

Semantic

High

High

High

Research papers, mixed-topic docs

Structure-Aware

High

High

Medium

Technical docs, manuals, reports

Parent-Child

High

Very High

Medium

QA over long documents

Late Chunking

High

Very High

High

Long-context retrieval

Step-by-Step Implementation

1. Chunking Strategy State Schema

from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
from datetime import datetime

ChunkingStrategy = Literal[
    "fixed_size", "recursive", "sentence_based", 
    "semantic", "structure_aware", "parent_child", "late_chunking"
]

class ChunkingState(TypedDict):
    messages: List
    conversation_id: str
    document_id: str
    # Input
    document_text: str
    document_type: str  # contract, manual, article, transcript, code
    document_length: int
    has_structure: bool  # headings, sections, tables
    query_patterns: List[str]  # expected query types
    # Analysis outputs
    document_characteristics: Dict[str, Any]
    strategy_scores: Dict[ChunkingStrategy, float]
    trade_offs: Dict[ChunkingStrategy, Dict[str, str]]
    # Recommendation
    recommended_strategy: ChunkingStrategy
    reasoning: str
    # Execution
    chunks: List[Dict]
    chunk_stats: Dict[str, Any]
    # Memory
    past_decisions: List[Dict]

2. Document Analyzer Agent

from langchain_openai import ChatOpenAI
import re

class DocumentAnalyzerAgent:
    """Analyzes document characteristics to inform strategy selection"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def analyze(self, state: ChunkingState) -> ChunkingState:
        text = state["document_text"]
        
        # Structural analysis
        heading_count = len(re.findall(r'^#{1,6}\s|^[A-Z][^.!?\n]{2,50}$', text, re.MULTILINE))
        paragraph_count = len([p for p in text.split('\n\n') if len(p.strip()) > 50])
        sentence_count = len(re.split(r'[.!?]+', text))
        avg_sentence_length = len(text.split()) / max(sentence_count, 1)
        
        # Detect document type signals
        has_clauses = bool(re.search(r'Section\s+\d+|Clause\s+\d+|Article\s+\d+', text, re.I))
        has_code_blocks = '```' in text or re.search(r'def\s+\w+\(|class\s+\w+:', text)
        has_tables = '|' in text and '---' in text
        
        characteristics = {
            "length_tokens": len(text.split()) * 1.3,
            "heading_count": heading_count,
            "paragraph_count": paragraph_count,
            "sentence_count": sentence_count,
            "avg_sentence_length": avg_sentence_length,
            "has_clauses": has_clauses,
            "has_code_blocks": has_code_blocks,
            "has_tables": has_tables,
            "has_structure": state["has_structure"] or heading_count > 3,
            "topic_density": sentence_count / max(paragraph_count, 1)
        }
        
        state["document_characteristics"] = characteristics
        state["document_length"] = len(text)
        return state

3. Multiple Chunking Strategy Engines

from langchain_text_splitters import (
    RecursiveCharacterTextSplitter,
    CharacterTextSplitter
)
from langchain_experimental.text_splitter import SemanticChunker
from langchain_huggingface import HuggingFaceEmbeddings

class ChunkingEngines:
    """Implements all chunking strategies"""
    
    def __init__(self):
        self.embeddings = HuggingFaceEmbeddings(
            model_name="sentence-transformers/all-MiniLM-L6-v2"
        )
    
    def fixed_size(self, text: str, chunk_size: int = 500, overlap: int = 50) -> List[str]:
        splitter = CharacterTextSplitter(
            chunk_size=chunk_size, chunk_overlap=overlap, separator=""
        )
        return splitter.split_text(text)
    
    def recursive(self, text: str, chunk_size: int = 800, overlap: int = 100) -> List[str]:
        splitter = RecursiveCharacterTextSplitter(
            chunk_size=chunk_size, chunk_overlap=overlap,
            separators=["\n\n", "\n", ". ", " ", ""]
        )
        return splitter.split_text(text)
    
    def sentence_based(self, text: str, sentences_per_chunk: int = 5) -> List[str]:
        import nltk
        nltk.download('punkt', quiet=True)
        sentences = nltk.sent_tokenize(text)
        chunks = []
        for i in range(0, len(sentences), sentences_per_chunk):
            chunk = " ".join(sentences[i:i + sentences_per_chunk])
            chunks.append(chunk)
        return chunks
    
    def semantic(self, text: str, breakpoint_threshold: float = 0.5) -> List[str]:
        splitter = SemanticChunker(
            embeddings=self.embeddings,
            breakpoint_threshold_type="percentile",
            breakpoint_threshold_amount=breakpoint_threshold * 100
        )
        return [doc.page_content for doc in splitter.split_text(text)]
    
    def structure_aware(self, text: str) -> List[str]:
        """Split by document headings/sections"""
        sections = re.split(r'\n(?=#{1,6}\s|[A-Z][^.!?\n]{2,80}\n)', text)
        return [s.strip() for s in sections if len(s.strip()) > 50]
    
    def parent_child(self, text: str, child_size: int = 300, parent_size: int = 1200):
        """Returns (parent_chunks, child_to_parent_mapping)"""
        parent_splitter = RecursiveCharacterTextSplitter(
            chunk_size=parent_size, chunk_overlap=100
        )
        child_splitter = RecursiveCharacterTextSplitter(
            chunk_size=child_size, chunk_overlap=30
        )
        parents = parent_splitter.split_text(text)
        mapping = []
        all_children = []
        for i, parent in enumerate(parents):
            children = child_splitter.split_text(parent)
            for child in children:
                all_children.append(child)
                mapping.append({"child_idx": len(all_children) - 1, "parent_idx": i})
        return parents, all_children, mapping
    
    def execute(self, strategy: ChunkingStrategy, text: str) -> Dict:
        if strategy == "fixed_size":
            chunks = self.fixed_size(text)
        elif strategy == "recursive":
            chunks = self.recursive(text)
        elif strategy == "sentence_based":
            chunks = self.sentence_based(text)
        elif strategy == "semantic":
            chunks = self.semantic(text)
        elif strategy == "structure_aware":
            chunks = self.structure_aware(text)
        elif strategy == "parent_child":
            parents, children, mapping = self.parent_child(text)
            return {
                "chunks": children,
                "parents": parents,
                "mapping": mapping,
                "strategy": strategy
            }
        else:
            chunks = self.recursive(text)  # fallback
        
        return {
            "chunks": chunks,
            "strategy": strategy,
            "parents": None,
            "mapping": None
        }

4. Strategy Evaluator and Trade-Off Analyzer

class StrategyEvaluatorAgent:
    """Scores each strategy based on document characteristics"""
    
    def evaluate(self, state: ChunkingState) -> ChunkingState:
        chars = state["document_characteristics"]
        scores = {}
        
        # Fixed-size: good for homogeneous, bad for structured
        scores["fixed_size"] = 30 if chars["has_structure"] else 60
        
        # Recursive: general-purpose baseline
        scores["recursive"] = 65
        
        # Sentence-based: good for narrative, low topic density
        if chars["topic_density"] < 2 and not chars["has_structure"]:
            scores["sentence_based"] = 75
        else:
            scores["sentence_based"] = 50
        
        # Semantic: good for mixed-topic, expensive
        if chars["topic_density"] > 3:
            scores["semantic"] = 85
        else:
            scores["semantic"] = 55
        
        # Structure-aware: excellent when structure exists
        if chars["has_structure"]:
            scores["structure_aware"] = 90
        else:
            scores["structure_aware"] = 30
        
        # Parent-child: good for long docs with QA patterns
        if chars["length_tokens"] > 5000:
            scores["parent_child"] = 80
        else:
            scores["parent_child"] = 50
        
        # Late chunking: good for very long docs
        if chars["length_tokens"] > 10000:
            scores["late_chunking"] = 85
        else:
            scores["late_chunking"] = 40
        
        # Domain-specific boosts
        if chars["has_clauses"]:
            scores["structure_aware"] = min(100, scores["structure_aware"] + 10)
        if chars["has_code_blocks"]:
            scores["recursive"] = min(100, scores["recursive"] + 15)
        
        state["strategy_scores"] = scores
        return state

class TradeOffAnalyzerAgent:
    """Documents pros/cons of each strategy for the given document"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def analyze(self, state: ChunkingState) -> ChunkingState:
        chars = state["document_characteristics"]
        trade_offs = {
            "fixed_size": {
                "pros": "Fast, predictable chunk sizes, simple implementation",
                "cons": "Breaks sentences, ignores semantics, poor for structured docs"
            },
            "recursive": {
                "pros": "Respects natural boundaries, good general-purpose choice",
                "cons": "May still split semantic units, no topic awareness"
            },
            "sentence_based": {
                "pros": "Never splits sentences, preserves narrative flow",
                "cons": "Uneven chunk sizes, poor for dense technical content"
            },
            "semantic": {
                "pros": "Preserves topic coherence, detects meaning shifts",
                "cons": "Expensive (embeddings per sentence), slower processing"
            },
            "structure_aware": {
                "pros": "Respects document hierarchy, excellent for technical docs",
                "cons": "Requires well-formed structure, fails on plain text"
            },
            "parent_child": {
                "pros": "Best of both worlds—precise retrieval, rich context",
                "cons": "Complex indexing, higher storage costs"
            },
            "late_chunking": {
                "pros": "Preserves long-range context in embeddings",
                "cons": "Requires specialized models, limited ecosystem support"
            }
        }
        state["trade_offs"] = trade_offs
        return state

5. Recommendation and Execution Agents

class RecommendationAgent:
    """Selects optimal strategy and explains reasoning"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def recommend(self, state: ChunkingState) -> ChunkingState:
        scores = state["strategy_scores"]
        best_strategy = max(scores, key=scores.get)
        state["recommended_strategy"] = best_strategy
        
        prompt = f"""Explain why {best_strategy} chunking is optimal for this document.

Document characteristics:
- Length: {state['document_characteristics']['length_tokens']:.0f} tokens
- Has structure: {state['document_characteristics']['has_structure']}
- Has clauses: {state['document_characteristics']['has_clauses']}
- Topic density: {state['document_characteristics']['topic_density']:.2f}

Strategy scores: {scores}

Provide a 2-3 sentence reasoning."""
        response = self.llm.invoke(prompt)
        state["reasoning"] = response.content
        return state

class ChunkingExecutorAgent:
    """Executes the recommended chunking strategy"""
    
    def __init__(self):
        self.engines = ChunkingEngines()
    
    def execute(self, state: ChunkingState) -> ChunkingState:
        result = self.engines.execute(
            state["recommended_strategy"], 
            state["document_text"]
        )
        
        chunks = result["chunks"]
        state["chunks"] = [
            {"id": f"chunk_{i}", "content": c, "metadata": {
                "strategy": result["strategy"],
                "length": len(c),
                "parent_id": None
            }}
            for i, c in enumerate(chunks)
        ]
        
        # Add parent info if parent_child strategy
        if result["parents"] and result["mapping"]:
            for m in result["mapping"]:
                state["chunks"][m["child_idx"]]["metadata"]["parent_id"] = f"parent_{m['parent_idx']}"
        
        # Compute stats
        lengths = [len(c["content"].split()) for c in state["chunks"]]
        state["chunk_stats"] = {
            "total_chunks": len(state["chunks"]),
            "avg_length_words": sum(lengths) / max(len(lengths), 1),
            "min_length_words": min(lengths) if lengths else 0,
            "max_length_words": max(lengths) if lengths else 0,
            "strategy_used": result["strategy"]
        }
        return state

6. LangGraph Workflow with Memory

class ChunkingMemory:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def save_decision(self, state: ChunkingState):
        record = {
            "document_type": state["document_type"],
            "strategy": state["recommended_strategy"],
            "scores": state["strategy_scores"],
            "chunk_count": state["chunk_stats"]["total_chunks"],
            "timestamp": datetime.now().isoformat()
        }
        key = f"chunking:{state['document_type']}"
        history = json.loads(self.redis.get(key) or "[]")
        history.append(record)
        self.redis.set(key, json.dumps(history[-50:]))
    
    def load_history(self, document_type: str) -> List[Dict]:
        return json.loads(self.redis.get(f"chunking:{document_type}") or "[]")

def build_chunking_advisor():
    workflow = StateGraph(ChunkingState)
    
    analyzer = DocumentAnalyzerAgent()
    evaluator = StrategyEvaluatorAgent()
    trade_off = TradeOffAnalyzerAgent()
    recommender = RecommendationAgent()
    executor = ChunkingExecutorAgent()
    
    workflow.add_node("analyze_document", analyzer.analyze)
    workflow.add_node("evaluate_strategies", evaluator.evaluate)
    workflow.add_node("analyze_trade_offs", trade_off.analyze)
    workflow.add_node("recommend_strategy", recommender.recommend)
    workflow.add_node("execute_chunking", executor.execute)
    
    workflow.set_entry_point("analyze_document")
    workflow.add_edge("analyze_document", "evaluate_strategies")
    workflow.add_edge("evaluate_strategies", "analyze_trade_offs")
    workflow.add_edge("analyze_trade_offs", "recommend_strategy")
    workflow.add_edge("recommend_strategy", "execute_chunking")
    workflow.add_edge("execute_chunking", END)
    
    return workflow.compile()

7. FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel

app = FastAPI(title="Chunking Strategy Advisor API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])

graph = build_chunking_advisor()
memory = ChunkingMemory(redis.Redis())

class ChunkingRequest(BaseModel):
    document_id: str
    document_text: str
    document_type: str = "general"
    has_structure: bool = False
    query_patterns: List[str] = []

@app.post("/advise_chunking")
async def advise_chunking(req: ChunkingRequest):
    history = memory.load_history(req.document_type)
    
    initial_state = ChunkingState(
        messages=[HumanMessage(content=f"Analyze chunking for {req.document_type}")],
        conversation_id=f"conv_{req.document_id}",
        document_id=req.document_id,
        document_text=req.document_text,
        document_type=req.document_type,
        document_length=len(req.document_text),
        has_structure=req.has_structure,
        query_patterns=req.query_patterns,
        document_characteristics={},
        strategy_scores={},
        trade_offs={},
        recommended_strategy="recursive",
        reasoning="",
        chunks=[],
        chunk_stats={},
        past_decisions=history
    )
    
    result = graph.invoke(initial_state)
    memory.save_decision(result)
    
    return {
        "recommended_strategy": result["recommended_strategy"],
        "reasoning": result["reasoning"],
        "strategy_scores": result["strategy_scores"],
        "trade_offs": result["trade_offs"],
        "chunk_stats": result["chunk_stats"],
        "sample_chunks": result["chunks"][:5],
        "document_characteristics": result["document_characteristics"]
    }

8. Frontend: Interactive Chunking Dashboard

// components/ChunkingAdvisor.tsx
import React, { useState } from 'react';

export const ChunkingAdvisor: React.FC = () => {
  const [docText, setDocText] = useState(`# Section 1: Definitions
"Agreement" means this Master Services Agreement.
"Confidential Information" means all non-public data...

# Section 2: Scope of Services
Provider shall deliver the Services as described in Exhibit A.
Any changes to scope require written amendment...

# Section 3: Payment Terms
Client shall pay invoices within 30 days of receipt.
Late payments accrue interest at 1.5% per month...`);
  const [docType, setDocType] = useState('contract');
  const [result, setResult] = useState<any>(null);

  const handleSubmit = async () => {
    const response = await fetch('http://localhost:8000/advise_chunking', {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({
        document_id: 'contract-001',
        document_text: docText,
        document_type: docType,
        has_structure: true
      })
    });
    setResult(await response.json());
  };

  return (
    <div className="p-6 max-w-6xl mx-auto bg-gray-50 min-h-screen">
      <h1 className="text-3xl font-bold mb-6">🔪 Chunking Strategy Advisor</h1>
      
      <div className="grid grid-cols-2 gap-6">
        <div className="bg-white p-4 rounded-lg shadow">
          <h2 className="font-bold mb-3">Document Input</h2>
          <label className="text-sm">Document Type:</label>
          <select className="w-full p-2 border rounded mb-3"
            value={docType} onChange={e => setDocType(e.target.value)}>
            <option value="contract">Legal Contract</option>
            <option value="manual">Technical Manual</option>
            <option value="article">Article/Blog</option>
            <option value="transcript">Transcript</option>
            <option value="code">Code Repository</option>
          </select>
          <textarea
            className="w-full p-2 border rounded"
            rows={12}
            value={docText}
            onChange={e => setDocText(e.target.value)}
          />
          <button onClick={handleSubmit} className="mt-3 bg-blue-600 text-white px-6 py-2 rounded">
            Analyze & Recommend
          </button>
        </div>

        {result && (
          <div className="space-y-4">
            <div className="bg-gradient-to-r from-blue-600 to-purple-600 text-white p-4 rounded-lg shadow">
              <h2 className="text-xl font-bold">
                Recommended: {result.recommended_strategy.replace(/_/g, ' ').toUpperCase()}
              </h2>
              <p className="mt-2 text-sm">{result.reasoning}</p>
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-2">Strategy Scores</h3>
              {Object.entries(result.strategy_scores).map(([k, v]: any) => (
                <div key={k} className="flex items-center gap-2 mb-1">
                  <span className="text-sm w-32">{k.replace(/_/g, ' ')}</span>
                  <div className="flex-1 bg-gray-200 rounded-full h-2">
                    <div className={`h-2 rounded-full ${
                      k === result.recommended_strategy ? 'bg-green-500' : 'bg-blue-400'
                    }`} style={{width: `${v}%`}} />
                  </div>
                  <span className="text-sm font-mono w-10">{v}</span>
                </div>
              ))}
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-2">Chunk Statistics</h3>
              <div className="grid grid-cols-2 gap-2 text-sm">
                <div>Total chunks: <b>{result.chunk_stats.total_chunks}</b></div>
                <div>Avg length: <b>{result.chunk_stats.avg_length_words.toFixed(0)} words</b></div>
                <div>Min length: <b>{result.chunk_stats.min_length_words} words</b></div>
                <div>Max length: <b>{result.chunk_stats.max_length_words} words</b></div>
              </div>
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-2">Sample Chunks</h3>
              {result.sample_chunks.map((c: any, i: number) => (
                <div key={i} className="border-l-4 border-blue-500 pl-3 mb-2 text-sm">
                  <div className="text-xs text-gray-500">{c.id} ({c.metadata.length} chars)</div>
                  <div className="text-gray-700">{c.content.slice(0, 150)}...</div>
                </div>
              ))}
            </div>
          </div>
        )}
      </div>
    </div>
  );
};

Real-Time Use Case: Legal Document Intelligence Platform

A law firm ingests a 50-page M&A agreement. The Chunking Advisor analyzes it:

Document Analysis:

Strategy Scores:

Recommendation: Structure-aware chunking, because the contract has clear section headings that define natural semantic boundaries. Each section becomes a chunk, preserving clause integrity.

Result: 12 chunks (one per section), average 1,200 words each. A query about "indemnification obligations" retrieves the exact Section 8 chunk—not a random 500-character window that cuts mid-sentence.

When the same firm later processes a 200-page deposition transcript (no structure, high narrative flow), the advisor recommends sentence-based chunking with 5 sentences per chunk—demonstrating how the system adapts to document characteristics.

Conclusion

Chunking is not a one-size-fits-all preprocessing step it's a strategic decision that directly impacts retrieval quality, context preservation, and downstream generation accuracy. By encoding chunking expertise into a multi-agent LangGraph advisor, enterprises can automatically select the optimal strategy based on document characteristics, quantify trade-offs, and maintain institutional memory of what works for each document type. The system transforms chunking from an afterthought into a first-class architectural decision, ensuring that every document in your RAG pipeline is segmented in the way that best serves your retrieval and generation goals. As document types multiply and use cases diversify, this adaptive approach becomes essential for maintaining RAG performance at enterprise scale.