Table of Contents

  1. Introduction: Beyond the Dense vs Sparse Debate

  2. What is Hybrid Search in RAG?

  3. Why Hybrid Wins: The Complementarity Principle

  4. Fusion Strategies: RRF, Convex Combination, and Learned Fusion

  5. Solution Architecture: The "Hybrid Fusion Engine" Multi-Agent System

  6. Technology Stack Overview

  7. Step-by-Step Implementation: Backend Development

    • Defining the Hybrid Search State Schema with Memory

    • Building the Query Preprocessor Agent

    • Implementing the Sparse Retriever (BM25) Agent

    • Implementing the Dense Retriever (Vector) Agent

    • Creating the Score Normalizer Agent

    • Designing the Fusion Agent with Multiple Strategies

    • Building the Cross-Encoder Reranker Agent

    • Constructing the LangGraph Workflow with Adaptive Routing

  8. Frontend Implementation: Hybrid Search Explorer

  9. Real-Time Use Case: Legal Contract Discovery Platform

  10. Conclusion: Hybrid Search as the Enterprise Default

Introduction

The retrieval layer is the single most impactful component of any RAG system, yet most implementations rely on a single retrieval strategy either pure vector search or pure keyword matching. This is like hiring a detective who only looks at fingerprints OR only interviews witnesses, but never both. Hybrid search combines the precision of lexical matching with the semantic understanding of neural embeddings, delivering retrieval quality that neither approach can achieve alone. In enterprise RAG deployments, hybrid search has become the default architecture because real-world queries demand both capabilities: users search for specific identifiers (error codes, clause numbers, product SKUs) AND express conceptual intent ("how do I fix slow performance"). This article presents an enterprise-grade multi-agent LangGraph system that implements hybrid search with adaptive fusion strategies, score normalization, cross-encoder reranking, and persistent memory that learns optimal weighting per query type over time.

What is Hybrid Search in RAG?

Hybrid search is a retrieval strategy that combines two complementary approaches:

  1. Sparse retrieval (BM25/TF-IDF): Matches exact terms using lexical overlap. Represents documents as high-dimensional sparse vectors where each dimension corresponds to a vocabulary term.

  2. Dense retrieval (neural embeddings): Matches semantic meaning using learned vector representations. Represents documents as low-dimensional dense vectors where every dimension carries meaning.

The hybrid pipeline executes both retrievers in parallel, normalizes their scores to a common scale, fuses the results using a combination strategy, and optionally reranks with a cross-encoder for final precision.

Why Hybrid Wins: The Complementarity Principle

Each retrieval method has blind spots that the other covers:

Query Type

Sparse (BM25)

Dense (Vector)

Hybrid

"ERR-4042 database timeout"

✅ Exact match

❌ May miss code

✅ BM25 dominates

"how to fix slow performance"

❌ No term overlap

✅ Semantic match

✅ Dense dominates

"Section 8.3 indemnification"

✅ Clause number

⚠️ Partial

✅ Both contribute

"best practices for API security"

❌ Vague terms

✅ Conceptual

✅ Dense dominates

"SKU-78234 wireless keyboard"

✅ Exact SKU

❌ SKU is noise

✅ BM25 dominates

In production benchmarks, hybrid search consistently outperforms either method alone by 15-40% on NDCG@10, with the gap widening on mixed-domain corpora.

Fusion Strategies

1. Reciprocal Rank Fusion (RRF): Combines results by rank position, not score. RRF_score(d) = Σ 1/(k + rank_i(d)). Robust to score scale differences. Default choice for most systems.

2. Convex Combination: Weighted sum of normalized scores. final_score = α × sparse_norm + (1-α) × dense_norm. Requires score normalization (min-max or z-score). α can be fixed or learned.

3. Learned Fusion: A classifier (often logistic regression or small neural net) predicts relevance using features from both retrievers. Highest accuracy but requires training data.

4. Cross-Encoder Reranking: After fusion, a cross-encoder model rescores the top-K candidates by jointly encoding query+document. Highest quality but expensive—applied only to the fused top-K.

Technology Tags

Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, pgvector, Redis, Pydantic, Elasticsearch, OpenAI API, Sentence Transformers, BM25, Rank-BM25, HNSWLib, CrossEncoder, Reciprocal Rank Fusion, asyncio, Docker, TypeScript, TailwindCSS, NumPy, Prometheus

Step-by-Step Implementation

1. Hybrid Search State Schema with Memory

from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
import time
import numpy as np
from datetime import datetime

FusionStrategy = Literal["rrf", "convex", "learned"]

class HybridSearchState(TypedDict):
    messages: List
    conversation_id: str
    query: str
    # Preprocessing
    normalized_query: str
    query_type: str  # "exact", "semantic", "mixed"
    # Retrieval results (pre-fusion)
    sparse_results: List[Dict]
    dense_results: List[Dict]
    # Normalization
    sparse_normalized: List[Dict]
    dense_normalized: List[Dict]
    # Fusion
    fusion_strategy: FusionStrategy
    fused_results: List[Dict]
    sparse_weight: float
    dense_weight: float
    # Reranking
    reranked_results: List[Dict]
    # Metrics
    total_latency_ms: float
    sparse_latency_ms: float
    dense_latency_ms: float
    fusion_latency_ms: float
    rerank_latency_ms: float
    # Memory
    historical_weights: Dict[str, List[float]]
    successful_strategies: List[Dict]

2. Query Preprocessor Agent

from langchain_openai import ChatOpenAI
import re

class QueryPreprocessorAgent:
    """Normalizes query and classifies type to guide fusion weighting"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.0)
    
    def preprocess(self, state: HybridSearchState) -> HybridSearchState:
        query = state["query"]
        
        # Detect identifiers (favor sparse)
        identifier_patterns = [
            r'\b[A-Z]{2,5}-\d{3,6}\b',
            r'\b[A-Z]{3,}-[A-Z0-9-]+\b',
            r'\b\d{4,}\b',
            r'Section\s+\d+',
            r'Clause\s+\d+'
        ]
        has_identifiers = any(re.search(p, query) for p in identifier_patterns)
        
        # Detect conceptual terms (favor dense)
        conceptual_words = {'how', 'why', 'what', 'best', 'practice', 'optimize', 
                           'improve', 'troubleshoot', 'explain'}
        query_words = set(query.lower().split())
        has_conceptual = bool(query_words & conceptual_words)
        
        if has_identifiers and has_conceptual:
            state["query_type"] = "mixed"
        elif has_identifiers:
            state["query_type"] = "exact"
        elif has_conceptual or len(query.split()) > 5:
            state["query_type"] = "semantic"
        else:
            state["query_type"] = "mixed"
        
        # Normalize: lowercase, remove punctuation but preserve identifiers
        state["normalized_query"] = query.strip()
        return state

3. Sparse Retriever Agent (BM25)

from rank_bm25 import BM25Okapi

class SparseRetrieverAgent:
    """BM25-based lexical retrieval with timing"""
    
    def __init__(self):
        # Sample legal corpus
        self.corpus = [
            {"id": "contract_001", "title": "MSA - TechCorp Inc", 
             "content": "Section 8.3 Indemnification. Provider shall indemnify Client against all third-party claims arising from Provider's negligence. Liability cap: $15M aggregate."},
            {"id": "contract_002", "title": "MSA - DataFlow LLC",
             "content": "Section 8.3 Indemnification. Each party shall indemnify the other against claims arising from breach. No liability cap applies for willful misconduct."},
            {"id": "contract_003", "title": "NDA - SecureNet",
             "content": "Confidential Information means all non-public data disclosed between parties. Excludes information independently developed or publicly available."},
            {"id": "contract_004", "title": "SOW - CloudOps Migration",
             "content": "Scope of Work: Migration of on-premise infrastructure to AWS. Timeline: 90 days. Payment: milestone-based, net-30 terms."},
            {"id": "err_doc_001", "title": "Error Reference ERR-4042",
             "content": "ERR-4042: Database connection timeout after 30s. Resolution: check connection pool, network config, and database server health."},
            {"id": "err_doc_002", "title": "Error Reference AUTH-1103",
             "content": "AUTH-1103: Invalid credentials. Resolution: reset password via admin console, verify account not locked."},
        ]
        tokenized = [doc["content"].lower().split() for doc in self.corpus]
        self.bm25 = BM25Okapi(tokenized)
    
    def retrieve(self, state: HybridSearchState) -> HybridSearchState:
        start = time.time()
        tokens = state["normalized_query"].lower().split()
        scores = self.bm25.get_scores(tokens)
        
        top_k = 10
        top_indices = np.argsort(scores)[::-1][:top_k]
        
        results = []
        for idx in top_indices:
            if scores[idx] > 0:
                results.append({
                    "id": self.corpus[idx]["id"],
                    "title": self.corpus[idx]["title"],
                    "content": self.corpus[idx]["content"],
                    "raw_score": float(scores[idx]),
                    "source": "sparse"
                })
        
        state["sparse_results"] = results
        state["sparse_latency_ms"] = (time.time() - start) * 1000
        return state

4. Dense Retriever Agent (Vector Search)

from langchain_huggingface import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity

class DenseRetrieverAgent:
    """Embedding-based semantic retrieval with timing"""
    
    def __init__(self):
        self.embeddings = HuggingFaceEmbeddings(
            model_name="sentence-transformers/all-MiniLM-L6-v2"
        )
        self.corpus = [
            {"id": "contract_001", "title": "MSA - TechCorp Inc", 
             "content": "Section 8.3 Indemnification. Provider shall indemnify Client against all third-party claims arising from Provider's negligence. Liability cap: $15M aggregate."},
            {"id": "contract_002", "title": "MSA - DataFlow LLC",
             "content": "Section 8.3 Indemnification. Each party shall indemnify the other against claims arising from breach. No liability cap applies for willful misconduct."},
            {"id": "contract_003", "title": "NDA - SecureNet",
             "content": "Confidential Information means all non-public data disclosed between parties. Excludes information independently developed or publicly available."},
            {"id": "contract_004", "title": "SOW - CloudOps Migration",
             "content": "Scope of Work: Migration of on-premise infrastructure to AWS. Timeline: 90 days. Payment: milestone-based, net-30 terms."},
            {"id": "err_doc_001", "title": "Error Reference ERR-4042",
             "content": "ERR-4042: Database connection timeout after 30s. Resolution: check connection pool, network config, and database server health."},
            {"id": "err_doc_002", "title": "Error Reference AUTH-1103",
             "content": "AUTH-1103: Invalid credentials. Resolution: reset password via admin console, verify account not locked."},
        ]
        texts = [doc["content"] for doc in self.corpus]
        self.doc_embeddings = np.array(self.embeddings.embed_documents(texts))
    
    def retrieve(self, state: HybridSearchState) -> HybridSearchState:
        start = time.time()
        query_embedding = np.array(self.embeddings.embed_query(state["normalized_query"]))
        
        similarities = cosine_similarity(
            query_embedding.reshape(1, -1),
            self.doc_embeddings
        )[0]
        
        top_k = 10
        top_indices = np.argsort(similarities)[::-1][:top_k]
        
        results = []
        for idx in top_indices:
            if similarities[idx] > 0.15:
                results.append({
                    "id": self.corpus[idx]["id"],
                    "title": self.corpus[idx]["title"],
                    "content": self.corpus[idx]["content"],
                    "raw_score": float(similarities[idx]),
                    "source": "dense"
                })
        
        state["dense_results"] = results
        state["dense_latency_ms"] = (time.time() - start) * 1000
        return state

5. Score Normalizer Agent

class ScoreNormalizerAgent:
    """Normalizes sparse and dense scores to [0, 1] using min-max scaling"""
    
    def normalize(self, state: HybridSearchState) -> HybridSearchState:
        # Normalize sparse scores
        sparse = state["sparse_results"]
        if sparse:
            sparse_scores = [r["raw_score"] for r in sparse]
            s_min, s_max = min(sparse_scores), max(sparse_scores)
            s_range = s_max - s_min if s_max > s_min else 1.0
            state["sparse_normalized"] = [
                {**r, "norm_score": (r["raw_score"] - s_min) / s_range}
                for r in sparse
            ]
        else:
            state["sparse_normalized"] = []
        
        # Normalize dense scores
        dense = state["dense_results"]
        if dense:
            dense_scores = [r["raw_score"] for r in dense]
            d_min, d_max = min(dense_scores), max(dense_scores)
            d_range = d_max - d_min if d_max > d_min else 1.0
            state["dense_normalized"] = [
                {**r, "norm_score": (r["raw_score"] - d_min) / d_range}
                for r in dense
            ]
        else:
            state["dense_normalized"] = []
        
        return state

6. Fusion Agent with Multiple Strategies

class FusionAgent:
    """Fuses sparse and dense results using configurable strategy"""
    
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def _select_weights(self, state: HybridSearchState) -> tuple:
        """Select weights based on query type and historical success"""
        query_type = state["query_type"]
        
        # Default weights per query type
        defaults = {
            "exact": (0.75, 0.25),
            "semantic": (0.25, 0.75),
            "mixed": (0.50, 0.50)
        }
        sparse_w, dense_w = defaults.get(query_type, (0.5, 0.5))
        
        # Adjust based on historical success (memory-driven)
        history = state.get("historical_weights", {}).get(query_type, [])
        if len(history) >= 5:
            # Use average of last 5 successful weights for this query type
            avg_sparse = np.mean([h[0] for h in history[-5:]])
            avg_dense = np.mean([h[1] for h in history[-5:]])
            sparse_w, dense_w = avg_sparse, avg_dense
        
        return sparse_w, dense_w
    
    def _rrf_fusion(self, sparse: List[Dict], dense: List[Dict], k: int = 60) -> List[Dict]:
        """Reciprocal Rank Fusion"""
        doc_scores = {}
        
        for rank, doc in enumerate(sparse):
            doc_id = doc["id"]
            if doc_id not in doc_scores:
                doc_scores[doc_id] = {"doc": doc, "rrf_score": 0.0, "sparse_rank": rank+1, "dense_rank": None}
            doc_scores[doc_id]["rrf_score"] += 1.0 / (k + rank + 1)
        
        for rank, doc in enumerate(dense):
            doc_id = doc["id"]
            if doc_id not in doc_scores:
                doc_scores[doc_id] = {"doc": doc, "rrf_score": 0.0, "sparse_rank": None, "dense_rank": rank+1}
            else:
                doc_scores[doc_id]["dense_rank"] = rank + 1
            doc_scores[doc_id]["rrf_score"] += 1.0 / (k + rank + 1)
        
        return sorted(doc_scores.values(), key=lambda x: x["rrf_score"], reverse=True)
    
    def _convex_fusion(self, sparse: List[Dict], dense: List[Dict], 
                       sparse_w: float, dense_w: float) -> List[Dict]:
        """Convex combination of normalized scores"""
        doc_scores = {}
        
        for doc in sparse:
            doc_id = doc["id"]
            if doc_id not in doc_scores:
                doc_scores[doc_id] = {"doc": doc, "convex_score": 0.0, "sparse_contrib": 0.0, "dense_contrib": 0.0}
            doc_scores[doc_id]["convex_score"] += sparse_w * doc["norm_score"]
            doc_scores[doc_id]["sparse_contrib"] = doc["norm_score"]
        
        for doc in dense:
            doc_id = doc["id"]
            if doc_id not in doc_scores:
                doc_scores[doc_id] = {"doc": doc, "convex_score": 0.0, "sparse_contrib": 0.0, "dense_contrib": 0.0}
            doc_scores[doc_id]["convex_score"] += dense_w * doc["norm_score"]
            doc_scores[doc_id]["dense_contrib"] = doc["norm_score"]
        
        return sorted(doc_scores.values(), key=lambda x: x["convex_score"], reverse=True)
    
    def fuse(self, state: HybridSearchState) -> HybridSearchState:
        start = time.time()
        sparse_w, dense_w = self._select_weights(state)
        state["sparse_weight"] = sparse_w
        state["dense_weight"] = dense_w
        
        # Select strategy: RRF for mixed, convex for typed queries
        if state["query_type"] == "mixed":
            state["fusion_strategy"] = "rrf"
            fused = self._rrf_fusion(state["sparse_normalized"], state["dense_normalized"])
            state["fused_results"] = [
                {
                    "id": item["doc"]["id"],
                    "title": item["doc"]["title"],
                    "content": item["doc"]["content"],
                    "score": item["rrf_score"],
                    "sparse_rank": item["sparse_rank"],
                    "dense_rank": item["dense_rank"],
                    "sources": ["sparse"] if item["sparse_rank"] and not item["dense_rank"] else
                              ["dense"] if item["dense_rank"] and not item["sparse_rank"] else
                              ["sparse", "dense"]
                }
                for item in fused[:10]
            ]
        else:
            state["fusion_strategy"] = "convex"
            fused = self._convex_fusion(
                state["sparse_normalized"], state["dense_normalized"],
                sparse_w, dense_w
            )
            state["fused_results"] = [
                {
                    "id": item["doc"]["id"],
                    "title": item["doc"]["title"],
                    "content": item["doc"]["content"],
                    "score": item["convex_score"],
                    "sparse_contrib": round(item["sparse_contrib"], 3),
                    "dense_contrib": round(item["dense_contrib"], 3),
                    "sources": ["sparse", "dense"] if item["sparse_contrib"] > 0 and item["dense_contrib"] > 0 else
                              ["sparse"] if item["sparse_contrib"] > 0 else ["dense"]
                }
                for item in fused[:10]
            ]
        
        state["fusion_latency_ms"] = (time.time() - start) * 1000
        return state

7. Cross-Encoder Reranker Agent

from transformers import CrossEncoder

class RerankerAgent:
    """Cross-encoder reranking for final precision on fused top-K"""
    
    def __init__(self):
        # In production: load cross-encoder/ms-marco-MiniLM-L-6-v2
        # Using a lightweight mock here for POC
        self.model_name = "cross-encoder/mock"
    
    def rerank(self, state: HybridSearchState) -> HybridSearchState:
        start = time.time()
        query = state["normalized_query"]
        fused = state["fused_results"][:5]  # Rerank top-5 only
        
        if not fused:
            state["reranked_results"] = []
            return state
        
        # Mock cross-encoder scoring (in production: use real model)
        reranked = []
        for doc in fused:
            # Simulate cross-encoder: boost docs where query terms appear in content
            content_lower = doc["content"].lower()
            query_terms = query.lower().split()
            term_overlap = sum(1 for t in query_terms if t in content_lower)
            rerank_boost = term_overlap / max(len(query_terms), 1)
            
            final_score = doc["score"] * 0.6 + rerank_boost * 0.4
            reranked.append({**doc, "final_score": final_score, "rerank_boost": rerank_boost})
        
        reranked.sort(key=lambda x: x["final_score"], reverse=True)
        for i, doc in enumerate(reranked):
            doc["final_rank"] = i + 1
        
        state["reranked_results"] = reranked
        state["rerank_latency_ms"] = (time.time() - start) * 1000
        state["total_latency_ms"] = (
            state["sparse_latency_ms"] + state["dense_latency_ms"] +
            state["fusion_latency_ms"] + state["rerank_latency_ms"]
        )
        return state

8. LangGraph Workflow with Adaptive Memory

class HybridSearchMemory:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def save_successful_strategy(self, state: HybridSearchState):
        """Record which weights worked for which query type"""
        record = {
            "query_type": state["query_type"],
            "sparse_weight": state["sparse_weight"],
            "dense_weight": state["dense_weight"],
            "strategy": state["fusion_strategy"],
            "top_result_id": state["reranked_results"][0]["id"] if state["reranked_results"] else None,
            "timestamp": datetime.now().isoformat()
        }
        key = f"hybrid_success:{state['query_type']}"
        history = json.loads(self.redis.get(key) or "[]")
        history.append(record)
        self.redis.set(key, json.dumps(history[-50:]))
    
    def load_historical_weights(self) -> Dict[str, List[float]]:
        weights = {}
        for qt in ["exact", "semantic", "mixed"]:
            history = json.loads(self.redis.get(f"hybrid_success:{qt}") or "[]")
            weights[qt] = [(h["sparse_weight"], h["dense_weight"]) for h in history[-10:]]
        return weights

def build_hybrid_search_graph():
    workflow = StateGraph(HybridSearchState)
    
    preprocessor = QueryPreprocessorAgent()
    sparse_retriever = SparseRetrieverAgent()
    dense_retriever = DenseRetrieverAgent()
    normalizer = ScoreNormalizerAgent()
    fusion_agent = FusionAgent(redis.Redis())
    reranker = RerankerAgent()
    
    workflow.add_node("preprocess", preprocessor.preprocess)
    workflow.add_node("sparse_retrieve", sparse_retriever.retrieve)
    workflow.add_node("dense_retrieve", dense_retriever.retrieve)
    workflow.add_node("normalize", normalizer.normalize)
    workflow.add_node("fuse", fusion_agent.fuse)
    workflow.add_node("rerank", reranker.rerank)
    
    workflow.set_entry_point("preprocess")
    workflow.add_edge("preprocess", "sparse_retrieve")
    workflow.add_edge("preprocess", "dense_retrieve")  # Parallel fan-out
    workflow.add_edge("sparse_retrieve", "normalize")
    workflow.add_edge("dense_retrieve", "normalize")
    workflow.add_edge("normalize", "fuse")
    workflow.add_edge("fuse", "rerank")
    workflow.add_edge("rerank", END)
    
    return workflow.compile()

9. FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel

app = FastAPI(title="Hybrid Search Engine API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])

graph = build_hybrid_search_graph()
memory = HybridSearchMemory(redis.Redis())

class SearchRequest(BaseModel):
    query: str
    conversation_id: str = "default"

@app.post("/hybrid_search")
async def hybrid_search(req: SearchRequest):
    historical_weights = memory.load_historical_weights()
    
    initial_state = HybridSearchState(
        messages=[HumanMessage(content=req.query)],
        conversation_id=req.conversation_id,
        query=req.query,
        normalized_query="",
        query_type="mixed",
        sparse_results=[],
        dense_results=[],
        sparse_normalized=[],
        dense_normalized=[],
        fusion_strategy="rrf",
        fused_results=[],
        sparse_weight=0.5,
        dense_weight=0.5,
        reranked_results=[],
        total_latency_ms=0,
        sparse_latency_ms=0,
        dense_latency_ms=0,
        fusion_latency_ms=0,
        rerank_latency_ms=0,
        historical_weights=historical_weights,
        successful_strategies=[]
    )
    
    result = graph.invoke(initial_state)
    memory.save_successful_strategy(result)
    
    return {
        "query": result["query"],
        "query_type": result["query_type"],
        "fusion_strategy": result["fusion_strategy"],
        "weights": {"sparse": round(result["sparse_weight"], 2), "dense": round(result["dense_weight"], 2)},
        "latency": {
            "total_ms": round(result["total_latency_ms"], 2),
            "sparse_ms": round(result["sparse_latency_ms"], 2),
            "dense_ms": round(result["dense_latency_ms"], 2),
            "fusion_ms": round(result["fusion_latency_ms"], 2),
            "rerank_ms": round(result["rerank_latency_ms"], 2)
        },
        "results": result["reranked_results"]
    }

10. Frontend: Hybrid Search Explorer

// components/HybridSearchExplorer.tsx
import React, { useState } from 'react';

interface SearchResult {
  id: string;
  title: string;
  content: string;
  final_score: number;
  final_rank: number;
  sparse_contrib?: number;
  dense_contrib?: number;
  sparse_rank?: number;
  dense_rank?: number;
  sources: string[];
  rerank_boost?: number;
}

interface SearchResponse {
  query: string;
  query_type: string;
  fusion_strategy: string;
  weights: { sparse: number; dense: number };
  latency: { total_ms: number; sparse_ms: number; dense_ms: number; fusion_ms: number; rerank_ms: number };
  results: SearchResult[];
}

export const HybridSearchExplorer: React.FC = () => {
  const [query, setQuery] = useState('What are the indemnification obligations under Section 8.3?');
  const [result, setResult] = useState<SearchResponse | null>(null);
  const [loading, setLoading] = useState(false);

  const sampleQueries = [
    'What are the indemnification obligations under Section 8.3?',
    'ERR-4042 database timeout resolution',
    'How do we protect confidential information?',
    'AUTH-1103 login failure troubleshooting'
  ];

  const handleSearch = async () => {
    setLoading(true);
    try {
      const response = await fetch('http://localhost:8000/hybrid_search', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({ query, conversation_id: 'legal-demo' })
      });
      setResult(await response.json());
    } finally {
      setLoading(false);
    }
  };

  return (
    <div className="p-6 max-w-7xl mx-auto bg-gray-50 min-h-screen">
      <h1 className="text-3xl font-bold mb-2">🔀 Hybrid Search Explorer</h1>
      <p className="text-gray-600 mb-6">Sparse + Dense retrieval with adaptive fusion and cross-encoder reranking</p>

      <div className="bg-white p-4 rounded-lg shadow mb-6">
        <input
          className="w-full p-3 border rounded mb-3"
          value={query}
          onChange={e => setQuery(e.target.value)}
          placeholder="Enter search query..."
        />
        <div className="flex gap-2 mb-3 flex-wrap">
          {sampleQueries.map((sq, i) => (
            <button key={i} onClick={() => setQuery(sq)}
              className="text-xs bg-gray-100 hover:bg-gray-200 px-3 py-1 rounded">
              {sq.slice(0, 40)}...
            </button>
          ))}
        </div>
        <button onClick={handleSearch} disabled={loading}
          className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
          {loading ? 'Searching...' : 'Run Hybrid Search'}
        </button>
      </div>

      {result && (
        <div className="grid grid-cols-3 gap-4">
          <div className="col-span-2 space-y-4">
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">
                Final Ranked Results 
                <span className="text-sm font-normal text-gray-500 ml-2">
                  (after cross-encoder reranking)
                </span>
              </h3>
              {result.results.map((r, i) => (
                <div key={i} className="border-l-4 border-blue-500 pl-3 mb-3 py-2 bg-gray-50 rounded-r">
                  <div className="flex justify-between items-start">
                    <div className="flex-1">
                      <div className="flex items-center gap-2">
                        <span className="text-xs font-mono bg-blue-600 text-white px-2 py-0.5 rounded">
                          #{r.final_rank}
                        </span>
                        <span className="font-semibold">{r.title}</span>
                      </div>
                      <div className="text-sm text-gray-700 mt-1">{r.content}</div>
                    </div>
                    <div className="flex flex-col items-end gap-1 ml-3">
                      <span className="text-xs font-mono bg-green-100 px-2 py-0.5 rounded">
                        score: {r.final_score.toFixed(3)}
                      </span>
                      <div className="flex gap-1">
                        {r.sources.map(s => (
                          <span key={s} className={`text-xs px-2 py-0.5 rounded ${
                            s === 'sparse' ? 'bg-orange-100 text-orange-700' : 'bg-purple-100 text-purple-700'
                          }`}>{s}</span>
                        ))}
                      </div>
                      {r.rerank_boost !== undefined && (
                        <span className="text-xs text-gray-500">
                          rerank boost: {r.rerank_boost.toFixed(2)}
                        </span>
                      )}
                    </div>
                  </div>
                </div>
              ))}
            </div>
          </div>

          <div className="space-y-4">
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">Query Analysis</h3>
              <div className="space-y-2 text-sm">
                <div className="flex justify-between">
                  <span className="text-gray-600">Query Type:</span>
                  <span className="font-semibold capitalize">{result.query_type}</span>
                </div>
                <div className="flex justify-between">
                  <span className="text-gray-600">Fusion Strategy:</span>
                  <span className="font-mono font-semibold">{result.fusion_strategy}</span>
                </div>
              </div>
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">Adaptive Weights</h3>
              <div className="space-y-2">
                <div>
                  <div className="flex justify-between text-sm mb-1">
                    <span className="text-orange-600">Sparse (BM25)</span>
                    <span className="font-mono">{(result.weights.sparse * 100).toFixed(0)}%</span>
                  </div>
                  <div className="bg-gray-200 rounded-full h-2">
                    <div className="bg-orange-500 h-2 rounded-full"
                      style={{width: `${result.weights.sparse * 100}%`}} />
                  </div>
                </div>
                <div>
                  <div className="flex justify-between text-sm mb-1">
                    <span className="text-purple-600">Dense (Vector)</span>
                    <span className="font-mono">{(result.weights.dense * 100).toFixed(0)}%</span>
                  </div>
                  <div className="bg-gray-200 rounded-full h-2">
                    <div className="bg-purple-500 h-2 rounded-full"
                      style={{width: `${result.weights.dense * 100}%`}} />
                  </div>
                </div>
              </div>
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">Latency Breakdown</h3>
              <div className="space-y-2 text-sm">
                <div className="flex justify-between">
                  <span>Sparse retrieval</span>
                  <span className="font-mono">{result.latency.sparse_ms.toFixed(1)}ms</span>
                </div>
                <div className="flex justify-between">
                  <span>Dense retrieval</span>
                  <span className="font-mono">{result.latency.dense_ms.toFixed(1)}ms</span>
                </div>
                <div className="flex justify-between">
                  <span>Fusion</span>
                  <span className="font-mono">{result.latency.fusion_ms.toFixed(1)}ms</span>
                </div>
                <div className="flex justify-between">
                  <span>Reranking</span>
                  <span className="font-mono">{result.latency.rerank_ms.toFixed(1)}ms</span>
                </div>
                <hr />
                <div className="flex justify-between font-bold">
                  <span>Total</span>
                  <span className="font-mono text-blue-600">{result.latency.total_ms.toFixed(1)}ms</span>
                </div>
              </div>
            </div>
          </div>
        </div>
      )}
    </div>
  );
};

Real-Time Use Case: Legal Contract Discovery Platform

A global law firm with 50,000+ contracts needs to answer questions like "What are the indemnification obligations under Section 8.3?" and "Find all contracts with liability caps above $10M".

Query 1: "What are the indemnification obligations under Section 8.3?"

  • Query Type: mixed (contains "Section 8.3" identifier + conceptual "indemnification obligations")

  • Fusion Strategy: RRF (best for mixed queries)

  • Sparse retrieves: contract_001 and contract_002 (both contain "Section 8.3 Indemnification")

  • Dense retrieves: contract_001, contract_002, plus contract_003 (semantically related to obligations)

  • RRF Fusion: contract_001 ranks #1 (appears in both), contract_002 ranks #2

  • Cross-encoder reranking: Boosts contract_001 further due to term overlap with "indemnification obligations"

  • Final result: contract_001 at rank 1 with score 0.847

Query 2: "ERR-4042 database timeout resolution"

  • Query Type: exact (identifier-heavy)

  • Weights: sparse=0.75, dense=0.25

  • Sparse dominates: err_doc_001 matches "ERR-4042" exactly

  • Dense contributes: err_doc_002 (semantically related to error resolution)

  • Final result: err_doc_001 at rank 1

Query 3: "How do we protect confidential information?"

  • Query Type: semantic (conceptual question)

  • Weights: sparse=0.25, dense=0.75

  • Dense dominates: contract_003 retrieved via semantic similarity to "protect confidential information"

  • Sparse contributes little: no exact term overlap

  • Final result: contract_003 at rank 1

Memory-driven adaptation: After 100 queries of type mixed, the system observes that RRF with equal weights (0.5/0.5) produces the highest user satisfaction. It locks in these weights for future mixed queries, while continuing to explore adjustments for exact and semantic types.

Conclusion

Hybrid search is not merely a technical optimization it's an acknowledgment that human information needs are diverse. Users search by exact identifier, by conceptual intent, and by a mixture of both. By implementing a multi-agent LangGraph system that runs sparse and dense retrieval in parallel, normalizes scores to a common scale, fuses results using adaptive strategies (RRF for mixed queries, convex combination for typed queries), and refines with cross-encoder reranking, enterprises achieve retrieval quality that neither method can deliver alone. The persistent memory layer transforms the system from a static pipeline into a learning system that continuously optimizes fusion weights based on observed success patterns. In the enterprise RAG landscape, hybrid search has evolved from a nice-to-have to the default architecture because in the real world, users don't search in just one way, and neither should your retrieval system.