Table of Contents
Introduction: Beyond the Dense vs Sparse Debate
What is Hybrid Search in RAG?
Why Hybrid Wins: The Complementarity Principle
Fusion Strategies: RRF, Convex Combination, and Learned Fusion
Solution Architecture: The "Hybrid Fusion Engine" Multi-Agent System
Technology Stack Overview
Step-by-Step Implementation: Backend Development
Defining the Hybrid Search State Schema with Memory
Building the Query Preprocessor Agent
Implementing the Sparse Retriever (BM25) Agent
Implementing the Dense Retriever (Vector) Agent
Creating the Score Normalizer Agent
Designing the Fusion Agent with Multiple Strategies
Building the Cross-Encoder Reranker Agent
Constructing the LangGraph Workflow with Adaptive Routing
Frontend Implementation: Hybrid Search Explorer
Real-Time Use Case: Legal Contract Discovery Platform
Conclusion: Hybrid Search as the Enterprise Default
Introduction
The retrieval layer is the single most impactful component of any RAG system, yet most implementations rely on a single retrieval strategy either pure vector search or pure keyword matching. This is like hiring a detective who only looks at fingerprints OR only interviews witnesses, but never both. Hybrid search combines the precision of lexical matching with the semantic understanding of neural embeddings, delivering retrieval quality that neither approach can achieve alone. In enterprise RAG deployments, hybrid search has become the default architecture because real-world queries demand both capabilities: users search for specific identifiers (error codes, clause numbers, product SKUs) AND express conceptual intent ("how do I fix slow performance"). This article presents an enterprise-grade multi-agent LangGraph system that implements hybrid search with adaptive fusion strategies, score normalization, cross-encoder reranking, and persistent memory that learns optimal weighting per query type over time.
What is Hybrid Search in RAG?
Hybrid search is a retrieval strategy that combines two complementary approaches:
Sparse retrieval (BM25/TF-IDF): Matches exact terms using lexical overlap. Represents documents as high-dimensional sparse vectors where each dimension corresponds to a vocabulary term.
Dense retrieval (neural embeddings): Matches semantic meaning using learned vector representations. Represents documents as low-dimensional dense vectors where every dimension carries meaning.
The hybrid pipeline executes both retrievers in parallel, normalizes their scores to a common scale, fuses the results using a combination strategy, and optionally reranks with a cross-encoder for final precision.
Why Hybrid Wins: The Complementarity Principle
Each retrieval method has blind spots that the other covers:
Query Type | Sparse (BM25) | Dense (Vector) | Hybrid |
|---|---|---|---|
"ERR-4042 database timeout" | ✅ Exact match | ❌ May miss code | ✅ BM25 dominates |
"how to fix slow performance" | ❌ No term overlap | ✅ Semantic match | ✅ Dense dominates |
"Section 8.3 indemnification" | ✅ Clause number | ⚠️ Partial | ✅ Both contribute |
"best practices for API security" | ❌ Vague terms | ✅ Conceptual | ✅ Dense dominates |
"SKU-78234 wireless keyboard" | ✅ Exact SKU | ❌ SKU is noise | ✅ BM25 dominates |
In production benchmarks, hybrid search consistently outperforms either method alone by 15-40% on NDCG@10, with the gap widening on mixed-domain corpora.
Fusion Strategies
1. Reciprocal Rank Fusion (RRF): Combines results by rank position, not score. RRF_score(d) = Σ 1/(k + rank_i(d)). Robust to score scale differences. Default choice for most systems.
2. Convex Combination: Weighted sum of normalized scores. final_score = α × sparse_norm + (1-α) × dense_norm. Requires score normalization (min-max or z-score). α can be fixed or learned.
3. Learned Fusion: A classifier (often logistic regression or small neural net) predicts relevance using features from both retrievers. Highest accuracy but requires training data.
4. Cross-Encoder Reranking: After fusion, a cross-encoder model rescores the top-K candidates by jointly encoding query+document. Highest quality but expensive—applied only to the fused top-K.
Technology Tags
Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, pgvector, Redis, Pydantic, Elasticsearch, OpenAI API, Sentence Transformers, BM25, Rank-BM25, HNSWLib, CrossEncoder, Reciprocal Rank Fusion, asyncio, Docker, TypeScript, TailwindCSS, NumPy, Prometheus
Step-by-Step Implementation
1. Hybrid Search State Schema with Memory
from typing import List, Dict, Any, TypedDict, Optional, Literal
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
import time
import numpy as np
from datetime import datetime
FusionStrategy = Literal["rrf", "convex", "learned"]
class HybridSearchState(TypedDict):
messages: List
conversation_id: str
query: str
# Preprocessing
normalized_query: str
query_type: str # "exact", "semantic", "mixed"
# Retrieval results (pre-fusion)
sparse_results: List[Dict]
dense_results: List[Dict]
# Normalization
sparse_normalized: List[Dict]
dense_normalized: List[Dict]
# Fusion
fusion_strategy: FusionStrategy
fused_results: List[Dict]
sparse_weight: float
dense_weight: float
# Reranking
reranked_results: List[Dict]
# Metrics
total_latency_ms: float
sparse_latency_ms: float
dense_latency_ms: float
fusion_latency_ms: float
rerank_latency_ms: float
# Memory
historical_weights: Dict[str, List[float]]
successful_strategies: List[Dict]
2. Query Preprocessor Agent
from langchain_openai import ChatOpenAI
import re
class QueryPreprocessorAgent:
"""Normalizes query and classifies type to guide fusion weighting"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.0)
def preprocess(self, state: HybridSearchState) -> HybridSearchState:
query = state["query"]
# Detect identifiers (favor sparse)
identifier_patterns = [
r'\b[A-Z]{2,5}-\d{3,6}\b',
r'\b[A-Z]{3,}-[A-Z0-9-]+\b',
r'\b\d{4,}\b',
r'Section\s+\d+',
r'Clause\s+\d+'
]
has_identifiers = any(re.search(p, query) for p in identifier_patterns)
# Detect conceptual terms (favor dense)
conceptual_words = {'how', 'why', 'what', 'best', 'practice', 'optimize',
'improve', 'troubleshoot', 'explain'}
query_words = set(query.lower().split())
has_conceptual = bool(query_words & conceptual_words)
if has_identifiers and has_conceptual:
state["query_type"] = "mixed"
elif has_identifiers:
state["query_type"] = "exact"
elif has_conceptual or len(query.split()) > 5:
state["query_type"] = "semantic"
else:
state["query_type"] = "mixed"
# Normalize: lowercase, remove punctuation but preserve identifiers
state["normalized_query"] = query.strip()
return state
3. Sparse Retriever Agent (BM25)
from rank_bm25 import BM25Okapi
class SparseRetrieverAgent:
"""BM25-based lexical retrieval with timing"""
def __init__(self):
# Sample legal corpus
self.corpus = [
{"id": "contract_001", "title": "MSA - TechCorp Inc",
"content": "Section 8.3 Indemnification. Provider shall indemnify Client against all third-party claims arising from Provider's negligence. Liability cap: $15M aggregate."},
{"id": "contract_002", "title": "MSA - DataFlow LLC",
"content": "Section 8.3 Indemnification. Each party shall indemnify the other against claims arising from breach. No liability cap applies for willful misconduct."},
{"id": "contract_003", "title": "NDA - SecureNet",
"content": "Confidential Information means all non-public data disclosed between parties. Excludes information independently developed or publicly available."},
{"id": "contract_004", "title": "SOW - CloudOps Migration",
"content": "Scope of Work: Migration of on-premise infrastructure to AWS. Timeline: 90 days. Payment: milestone-based, net-30 terms."},
{"id": "err_doc_001", "title": "Error Reference ERR-4042",
"content": "ERR-4042: Database connection timeout after 30s. Resolution: check connection pool, network config, and database server health."},
{"id": "err_doc_002", "title": "Error Reference AUTH-1103",
"content": "AUTH-1103: Invalid credentials. Resolution: reset password via admin console, verify account not locked."},
]
tokenized = [doc["content"].lower().split() for doc in self.corpus]
self.bm25 = BM25Okapi(tokenized)
def retrieve(self, state: HybridSearchState) -> HybridSearchState:
start = time.time()
tokens = state["normalized_query"].lower().split()
scores = self.bm25.get_scores(tokens)
top_k = 10
top_indices = np.argsort(scores)[::-1][:top_k]
results = []
for idx in top_indices:
if scores[idx] > 0:
results.append({
"id": self.corpus[idx]["id"],
"title": self.corpus[idx]["title"],
"content": self.corpus[idx]["content"],
"raw_score": float(scores[idx]),
"source": "sparse"
})
state["sparse_results"] = results
state["sparse_latency_ms"] = (time.time() - start) * 1000
return state
4. Dense Retriever Agent (Vector Search)
from langchain_huggingface import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity
class DenseRetrieverAgent:
"""Embedding-based semantic retrieval with timing"""
def __init__(self):
self.embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
self.corpus = [
{"id": "contract_001", "title": "MSA - TechCorp Inc",
"content": "Section 8.3 Indemnification. Provider shall indemnify Client against all third-party claims arising from Provider's negligence. Liability cap: $15M aggregate."},
{"id": "contract_002", "title": "MSA - DataFlow LLC",
"content": "Section 8.3 Indemnification. Each party shall indemnify the other against claims arising from breach. No liability cap applies for willful misconduct."},
{"id": "contract_003", "title": "NDA - SecureNet",
"content": "Confidential Information means all non-public data disclosed between parties. Excludes information independently developed or publicly available."},
{"id": "contract_004", "title": "SOW - CloudOps Migration",
"content": "Scope of Work: Migration of on-premise infrastructure to AWS. Timeline: 90 days. Payment: milestone-based, net-30 terms."},
{"id": "err_doc_001", "title": "Error Reference ERR-4042",
"content": "ERR-4042: Database connection timeout after 30s. Resolution: check connection pool, network config, and database server health."},
{"id": "err_doc_002", "title": "Error Reference AUTH-1103",
"content": "AUTH-1103: Invalid credentials. Resolution: reset password via admin console, verify account not locked."},
]
texts = [doc["content"] for doc in self.corpus]
self.doc_embeddings = np.array(self.embeddings.embed_documents(texts))
def retrieve(self, state: HybridSearchState) -> HybridSearchState:
start = time.time()
query_embedding = np.array(self.embeddings.embed_query(state["normalized_query"]))
similarities = cosine_similarity(
query_embedding.reshape(1, -1),
self.doc_embeddings
)[0]
top_k = 10
top_indices = np.argsort(similarities)[::-1][:top_k]
results = []
for idx in top_indices:
if similarities[idx] > 0.15:
results.append({
"id": self.corpus[idx]["id"],
"title": self.corpus[idx]["title"],
"content": self.corpus[idx]["content"],
"raw_score": float(similarities[idx]),
"source": "dense"
})
state["dense_results"] = results
state["dense_latency_ms"] = (time.time() - start) * 1000
return state
5. Score Normalizer Agent
class ScoreNormalizerAgent:
"""Normalizes sparse and dense scores to [0, 1] using min-max scaling"""
def normalize(self, state: HybridSearchState) -> HybridSearchState:
# Normalize sparse scores
sparse = state["sparse_results"]
if sparse:
sparse_scores = [r["raw_score"] for r in sparse]
s_min, s_max = min(sparse_scores), max(sparse_scores)
s_range = s_max - s_min if s_max > s_min else 1.0
state["sparse_normalized"] = [
{**r, "norm_score": (r["raw_score"] - s_min) / s_range}
for r in sparse
]
else:
state["sparse_normalized"] = []
# Normalize dense scores
dense = state["dense_results"]
if dense:
dense_scores = [r["raw_score"] for r in dense]
d_min, d_max = min(dense_scores), max(dense_scores)
d_range = d_max - d_min if d_max > d_min else 1.0
state["dense_normalized"] = [
{**r, "norm_score": (r["raw_score"] - d_min) / d_range}
for r in dense
]
else:
state["dense_normalized"] = []
return state
6. Fusion Agent with Multiple Strategies
class FusionAgent:
"""Fuses sparse and dense results using configurable strategy"""
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def _select_weights(self, state: HybridSearchState) -> tuple:
"""Select weights based on query type and historical success"""
query_type = state["query_type"]
# Default weights per query type
defaults = {
"exact": (0.75, 0.25),
"semantic": (0.25, 0.75),
"mixed": (0.50, 0.50)
}
sparse_w, dense_w = defaults.get(query_type, (0.5, 0.5))
# Adjust based on historical success (memory-driven)
history = state.get("historical_weights", {}).get(query_type, [])
if len(history) >= 5:
# Use average of last 5 successful weights for this query type
avg_sparse = np.mean([h[0] for h in history[-5:]])
avg_dense = np.mean([h[1] for h in history[-5:]])
sparse_w, dense_w = avg_sparse, avg_dense
return sparse_w, dense_w
def _rrf_fusion(self, sparse: List[Dict], dense: List[Dict], k: int = 60) -> List[Dict]:
"""Reciprocal Rank Fusion"""
doc_scores = {}
for rank, doc in enumerate(sparse):
doc_id = doc["id"]
if doc_id not in doc_scores:
doc_scores[doc_id] = {"doc": doc, "rrf_score": 0.0, "sparse_rank": rank+1, "dense_rank": None}
doc_scores[doc_id]["rrf_score"] += 1.0 / (k + rank + 1)
for rank, doc in enumerate(dense):
doc_id = doc["id"]
if doc_id not in doc_scores:
doc_scores[doc_id] = {"doc": doc, "rrf_score": 0.0, "sparse_rank": None, "dense_rank": rank+1}
else:
doc_scores[doc_id]["dense_rank"] = rank + 1
doc_scores[doc_id]["rrf_score"] += 1.0 / (k + rank + 1)
return sorted(doc_scores.values(), key=lambda x: x["rrf_score"], reverse=True)
def _convex_fusion(self, sparse: List[Dict], dense: List[Dict],
sparse_w: float, dense_w: float) -> List[Dict]:
"""Convex combination of normalized scores"""
doc_scores = {}
for doc in sparse:
doc_id = doc["id"]
if doc_id not in doc_scores:
doc_scores[doc_id] = {"doc": doc, "convex_score": 0.0, "sparse_contrib": 0.0, "dense_contrib": 0.0}
doc_scores[doc_id]["convex_score"] += sparse_w * doc["norm_score"]
doc_scores[doc_id]["sparse_contrib"] = doc["norm_score"]
for doc in dense:
doc_id = doc["id"]
if doc_id not in doc_scores:
doc_scores[doc_id] = {"doc": doc, "convex_score": 0.0, "sparse_contrib": 0.0, "dense_contrib": 0.0}
doc_scores[doc_id]["convex_score"] += dense_w * doc["norm_score"]
doc_scores[doc_id]["dense_contrib"] = doc["norm_score"]
return sorted(doc_scores.values(), key=lambda x: x["convex_score"], reverse=True)
def fuse(self, state: HybridSearchState) -> HybridSearchState:
start = time.time()
sparse_w, dense_w = self._select_weights(state)
state["sparse_weight"] = sparse_w
state["dense_weight"] = dense_w
# Select strategy: RRF for mixed, convex for typed queries
if state["query_type"] == "mixed":
state["fusion_strategy"] = "rrf"
fused = self._rrf_fusion(state["sparse_normalized"], state["dense_normalized"])
state["fused_results"] = [
{
"id": item["doc"]["id"],
"title": item["doc"]["title"],
"content": item["doc"]["content"],
"score": item["rrf_score"],
"sparse_rank": item["sparse_rank"],
"dense_rank": item["dense_rank"],
"sources": ["sparse"] if item["sparse_rank"] and not item["dense_rank"] else
["dense"] if item["dense_rank"] and not item["sparse_rank"] else
["sparse", "dense"]
}
for item in fused[:10]
]
else:
state["fusion_strategy"] = "convex"
fused = self._convex_fusion(
state["sparse_normalized"], state["dense_normalized"],
sparse_w, dense_w
)
state["fused_results"] = [
{
"id": item["doc"]["id"],
"title": item["doc"]["title"],
"content": item["doc"]["content"],
"score": item["convex_score"],
"sparse_contrib": round(item["sparse_contrib"], 3),
"dense_contrib": round(item["dense_contrib"], 3),
"sources": ["sparse", "dense"] if item["sparse_contrib"] > 0 and item["dense_contrib"] > 0 else
["sparse"] if item["sparse_contrib"] > 0 else ["dense"]
}
for item in fused[:10]
]
state["fusion_latency_ms"] = (time.time() - start) * 1000
return state
7. Cross-Encoder Reranker Agent
from transformers import CrossEncoder
class RerankerAgent:
"""Cross-encoder reranking for final precision on fused top-K"""
def __init__(self):
# In production: load cross-encoder/ms-marco-MiniLM-L-6-v2
# Using a lightweight mock here for POC
self.model_name = "cross-encoder/mock"
def rerank(self, state: HybridSearchState) -> HybridSearchState:
start = time.time()
query = state["normalized_query"]
fused = state["fused_results"][:5] # Rerank top-5 only
if not fused:
state["reranked_results"] = []
return state
# Mock cross-encoder scoring (in production: use real model)
reranked = []
for doc in fused:
# Simulate cross-encoder: boost docs where query terms appear in content
content_lower = doc["content"].lower()
query_terms = query.lower().split()
term_overlap = sum(1 for t in query_terms if t in content_lower)
rerank_boost = term_overlap / max(len(query_terms), 1)
final_score = doc["score"] * 0.6 + rerank_boost * 0.4
reranked.append({**doc, "final_score": final_score, "rerank_boost": rerank_boost})
reranked.sort(key=lambda x: x["final_score"], reverse=True)
for i, doc in enumerate(reranked):
doc["final_rank"] = i + 1
state["reranked_results"] = reranked
state["rerank_latency_ms"] = (time.time() - start) * 1000
state["total_latency_ms"] = (
state["sparse_latency_ms"] + state["dense_latency_ms"] +
state["fusion_latency_ms"] + state["rerank_latency_ms"]
)
return state
8. LangGraph Workflow with Adaptive Memory
class HybridSearchMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def save_successful_strategy(self, state: HybridSearchState):
"""Record which weights worked for which query type"""
record = {
"query_type": state["query_type"],
"sparse_weight": state["sparse_weight"],
"dense_weight": state["dense_weight"],
"strategy": state["fusion_strategy"],
"top_result_id": state["reranked_results"][0]["id"] if state["reranked_results"] else None,
"timestamp": datetime.now().isoformat()
}
key = f"hybrid_success:{state['query_type']}"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-50:]))
def load_historical_weights(self) -> Dict[str, List[float]]:
weights = {}
for qt in ["exact", "semantic", "mixed"]:
history = json.loads(self.redis.get(f"hybrid_success:{qt}") or "[]")
weights[qt] = [(h["sparse_weight"], h["dense_weight"]) for h in history[-10:]]
return weights
def build_hybrid_search_graph():
workflow = StateGraph(HybridSearchState)
preprocessor = QueryPreprocessorAgent()
sparse_retriever = SparseRetrieverAgent()
dense_retriever = DenseRetrieverAgent()
normalizer = ScoreNormalizerAgent()
fusion_agent = FusionAgent(redis.Redis())
reranker = RerankerAgent()
workflow.add_node("preprocess", preprocessor.preprocess)
workflow.add_node("sparse_retrieve", sparse_retriever.retrieve)
workflow.add_node("dense_retrieve", dense_retriever.retrieve)
workflow.add_node("normalize", normalizer.normalize)
workflow.add_node("fuse", fusion_agent.fuse)
workflow.add_node("rerank", reranker.rerank)
workflow.set_entry_point("preprocess")
workflow.add_edge("preprocess", "sparse_retrieve")
workflow.add_edge("preprocess", "dense_retrieve") # Parallel fan-out
workflow.add_edge("sparse_retrieve", "normalize")
workflow.add_edge("dense_retrieve", "normalize")
workflow.add_edge("normalize", "fuse")
workflow.add_edge("fuse", "rerank")
workflow.add_edge("rerank", END)
return workflow.compile()
9. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="Hybrid Search Engine API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_hybrid_search_graph()
memory = HybridSearchMemory(redis.Redis())
class SearchRequest(BaseModel):
query: str
conversation_id: str = "default"
@app.post("/hybrid_search")
async def hybrid_search(req: SearchRequest):
historical_weights = memory.load_historical_weights()
initial_state = HybridSearchState(
messages=[HumanMessage(content=req.query)],
conversation_id=req.conversation_id,
query=req.query,
normalized_query="",
query_type="mixed",
sparse_results=[],
dense_results=[],
sparse_normalized=[],
dense_normalized=[],
fusion_strategy="rrf",
fused_results=[],
sparse_weight=0.5,
dense_weight=0.5,
reranked_results=[],
total_latency_ms=0,
sparse_latency_ms=0,
dense_latency_ms=0,
fusion_latency_ms=0,
rerank_latency_ms=0,
historical_weights=historical_weights,
successful_strategies=[]
)
result = graph.invoke(initial_state)
memory.save_successful_strategy(result)
return {
"query": result["query"],
"query_type": result["query_type"],
"fusion_strategy": result["fusion_strategy"],
"weights": {"sparse": round(result["sparse_weight"], 2), "dense": round(result["dense_weight"], 2)},
"latency": {
"total_ms": round(result["total_latency_ms"], 2),
"sparse_ms": round(result["sparse_latency_ms"], 2),
"dense_ms": round(result["dense_latency_ms"], 2),
"fusion_ms": round(result["fusion_latency_ms"], 2),
"rerank_ms": round(result["rerank_latency_ms"], 2)
},
"results": result["reranked_results"]
}
10. Frontend: Hybrid Search Explorer
// components/HybridSearchExplorer.tsx
import React, { useState } from 'react';
interface SearchResult {
id: string;
title: string;
content: string;
final_score: number;
final_rank: number;
sparse_contrib?: number;
dense_contrib?: number;
sparse_rank?: number;
dense_rank?: number;
sources: string[];
rerank_boost?: number;
}
interface SearchResponse {
query: string;
query_type: string;
fusion_strategy: string;
weights: { sparse: number; dense: number };
latency: { total_ms: number; sparse_ms: number; dense_ms: number; fusion_ms: number; rerank_ms: number };
results: SearchResult[];
}
export const HybridSearchExplorer: React.FC = () => {
const [query, setQuery] = useState('What are the indemnification obligations under Section 8.3?');
const [result, setResult] = useState<SearchResponse | null>(null);
const [loading, setLoading] = useState(false);
const sampleQueries = [
'What are the indemnification obligations under Section 8.3?',
'ERR-4042 database timeout resolution',
'How do we protect confidential information?',
'AUTH-1103 login failure troubleshooting'
];
const handleSearch = async () => {
setLoading(true);
try {
const response = await fetch('http://localhost:8000/hybrid_search', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ query, conversation_id: 'legal-demo' })
});
setResult(await response.json());
} finally {
setLoading(false);
}
};
return (
<div className="p-6 max-w-7xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-2">🔀 Hybrid Search Explorer</h1>
<p className="text-gray-600 mb-6">Sparse + Dense retrieval with adaptive fusion and cross-encoder reranking</p>
<div className="bg-white p-4 rounded-lg shadow mb-6">
<input
className="w-full p-3 border rounded mb-3"
value={query}
onChange={e => setQuery(e.target.value)}
placeholder="Enter search query..."
/>
<div className="flex gap-2 mb-3 flex-wrap">
{sampleQueries.map((sq, i) => (
<button key={i} onClick={() => setQuery(sq)}
className="text-xs bg-gray-100 hover:bg-gray-200 px-3 py-1 rounded">
{sq.slice(0, 40)}...
</button>
))}
</div>
<button onClick={handleSearch} disabled={loading}
className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
{loading ? 'Searching...' : 'Run Hybrid Search'}
</button>
</div>
{result && (
<div className="grid grid-cols-3 gap-4">
<div className="col-span-2 space-y-4">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">
Final Ranked Results
<span className="text-sm font-normal text-gray-500 ml-2">
(after cross-encoder reranking)
</span>
</h3>
{result.results.map((r, i) => (
<div key={i} className="border-l-4 border-blue-500 pl-3 mb-3 py-2 bg-gray-50 rounded-r">
<div className="flex justify-between items-start">
<div className="flex-1">
<div className="flex items-center gap-2">
<span className="text-xs font-mono bg-blue-600 text-white px-2 py-0.5 rounded">
#{r.final_rank}
</span>
<span className="font-semibold">{r.title}</span>
</div>
<div className="text-sm text-gray-700 mt-1">{r.content}</div>
</div>
<div className="flex flex-col items-end gap-1 ml-3">
<span className="text-xs font-mono bg-green-100 px-2 py-0.5 rounded">
score: {r.final_score.toFixed(3)}
</span>
<div className="flex gap-1">
{r.sources.map(s => (
<span key={s} className={`text-xs px-2 py-0.5 rounded ${
s === 'sparse' ? 'bg-orange-100 text-orange-700' : 'bg-purple-100 text-purple-700'
}`}>{s}</span>
))}
</div>
{r.rerank_boost !== undefined && (
<span className="text-xs text-gray-500">
rerank boost: {r.rerank_boost.toFixed(2)}
</span>
)}
</div>
</div>
</div>
))}
</div>
</div>
<div className="space-y-4">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">Query Analysis</h3>
<div className="space-y-2 text-sm">
<div className="flex justify-between">
<span className="text-gray-600">Query Type:</span>
<span className="font-semibold capitalize">{result.query_type}</span>
</div>
<div className="flex justify-between">
<span className="text-gray-600">Fusion Strategy:</span>
<span className="font-mono font-semibold">{result.fusion_strategy}</span>
</div>
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">Adaptive Weights</h3>
<div className="space-y-2">
<div>
<div className="flex justify-between text-sm mb-1">
<span className="text-orange-600">Sparse (BM25)</span>
<span className="font-mono">{(result.weights.sparse * 100).toFixed(0)}%</span>
</div>
<div className="bg-gray-200 rounded-full h-2">
<div className="bg-orange-500 h-2 rounded-full"
style={{width: `${result.weights.sparse * 100}%`}} />
</div>
</div>
<div>
<div className="flex justify-between text-sm mb-1">
<span className="text-purple-600">Dense (Vector)</span>
<span className="font-mono">{(result.weights.dense * 100).toFixed(0)}%</span>
</div>
<div className="bg-gray-200 rounded-full h-2">
<div className="bg-purple-500 h-2 rounded-full"
style={{width: `${result.weights.dense * 100}%`}} />
</div>
</div>
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">Latency Breakdown</h3>
<div className="space-y-2 text-sm">
<div className="flex justify-between">
<span>Sparse retrieval</span>
<span className="font-mono">{result.latency.sparse_ms.toFixed(1)}ms</span>
</div>
<div className="flex justify-between">
<span>Dense retrieval</span>
<span className="font-mono">{result.latency.dense_ms.toFixed(1)}ms</span>
</div>
<div className="flex justify-between">
<span>Fusion</span>
<span className="font-mono">{result.latency.fusion_ms.toFixed(1)}ms</span>
</div>
<div className="flex justify-between">
<span>Reranking</span>
<span className="font-mono">{result.latency.rerank_ms.toFixed(1)}ms</span>
</div>
<hr />
<div className="flex justify-between font-bold">
<span>Total</span>
<span className="font-mono text-blue-600">{result.latency.total_ms.toFixed(1)}ms</span>
</div>
</div>
</div>
</div>
</div>
)}
</div>
);
};
Real-Time Use Case: Legal Contract Discovery Platform
A global law firm with 50,000+ contracts needs to answer questions like "What are the indemnification obligations under Section 8.3?" and "Find all contracts with liability caps above $10M".
Query 1: "What are the indemnification obligations under Section 8.3?"
Query Type:
mixed(contains "Section 8.3" identifier + conceptual "indemnification obligations")Fusion Strategy: RRF (best for mixed queries)
Sparse retrieves: contract_001 and contract_002 (both contain "Section 8.3 Indemnification")
Dense retrieves: contract_001, contract_002, plus contract_003 (semantically related to obligations)
RRF Fusion: contract_001 ranks #1 (appears in both), contract_002 ranks #2
Cross-encoder reranking: Boosts contract_001 further due to term overlap with "indemnification obligations"
Final result: contract_001 at rank 1 with score 0.847
Query 2: "ERR-4042 database timeout resolution"
Query Type:
exact(identifier-heavy)Weights: sparse=0.75, dense=0.25
Sparse dominates: err_doc_001 matches "ERR-4042" exactly
Dense contributes: err_doc_002 (semantically related to error resolution)
Final result: err_doc_001 at rank 1
Query 3: "How do we protect confidential information?"
Query Type:
semantic(conceptual question)Weights: sparse=0.25, dense=0.75
Dense dominates: contract_003 retrieved via semantic similarity to "protect confidential information"
Sparse contributes little: no exact term overlap
Final result: contract_003 at rank 1
Memory-driven adaptation: After 100 queries of type mixed, the system observes that RRF with equal weights (0.5/0.5) produces the highest user satisfaction. It locks in these weights for future mixed queries, while continuing to explore adjustments for exact and semantic types.
Conclusion
Hybrid search is not merely a technical optimization it's an acknowledgment that human information needs are diverse. Users search by exact identifier, by conceptual intent, and by a mixture of both. By implementing a multi-agent LangGraph system that runs sparse and dense retrieval in parallel, normalizes scores to a common scale, fuses results using adaptive strategies (RRF for mixed queries, convex combination for typed queries), and refines with cross-encoder reranking, enterprises achieve retrieval quality that neither method can deliver alone. The persistent memory layer transforms the system from a static pipeline into a learning system that continuously optimizes fusion weights based on observed success patterns. In the enterprise RAG landscape, hybrid search has evolved from a nice-to-have to the default architecture because in the real world, users don't search in just one way, and neither should your retrieval system.

Join the conversation! Your thoughts help the community grow.