Introduction
In the world of Retrieval-Augmented Generation (RAG), the standard approach to fetching context is Similarity Search. This method retrieves the top-k document chunks that are most semantically similar to the user's query. While effective for simple factual questions, similarity search has a critical flaw: Redundancy. When a user asks a broad or complex question, the top-k most similar chunks often contain nearly identical information. For example, if you search for "Python best practices," you might get five different paragraphs that all say "Use virtual environments." This redundancy wastes valuable context window space and prevents the Large Language Model (LLM) from seeing diverse perspectives or complementary facts.
This is where Maximum Marginal Relevance (MMR) comes in. MMR is an algorithm that selects documents by optimizing for two factors simultaneously:
Relevance: How similar the document is to the query.
Diversity: How different the document is from the already-selected documents.
By balancing these two, MMR ensures that the retrieved context is not only accurate but also comprehensive. In this article, we will build an enterprise-grade Market Research Assistant using LangGraph, Graph RAG, and MMR to demonstrate how diversity improves the quality of AI-generated insights.
Understanding MMR: The Math Behind the Magic
When to use MMR instead of Standard Similarity:
Broad Queries: When the user asks for an overview (e.g., "Summarize the market trends").
Brainstorming: When you need varied ideas rather than repeated confirmations.
Complex Reasoning: When the answer requires synthesizing multiple distinct viewpoints.
Real-Time Use Case: Enterprise Market Research Assistant
Imagine a strategic planning team at a consumer electronics company. They ask: "What are the emerging trends in wearable technology?"
A standard similarity search might return five articles all discussing "Smartwatch Battery Life." An MMR-enhanced system, however, would return:
An article on Smartwatch Battery Life (High Relevance).
A report on AR Glasses adoption (High Diversity).
A study on Health Monitoring Sensors (High Diversity).
A piece on Sustainable Materials in Wearables (High Diversity).
This diverse context allows the LLM to generate a much richer, multi-faceted report.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain==0.1.0
langchain-openai==0.0.5
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.5.0
chromadb==0.4.22
networkx==3.2.1
numpy==1.24.0
Step 2: MMR-Enabled Graph RAG Service
# services/mmr_graph_rag.py
import chromadb
import networkx as nx
import numpy as np
from typing import List, Dict
from langchain_community.vectorstores.utils import maximal_marginal_relevance
class MMRGraphRAG:
def __init__(self):
self.client = chromadb.PersistentClient(path="./market_db")
self.collection = self.client.get_or_create_collection("market_trends")
self.graph = nx.DiGraph()
self._seed_data()
def _seed_data(self):
# Diverse data points for the same topic
docs = [
("T1", "Smartwatches are seeing longer battery life due to new silicon.", {"topic": "hardware"}),
("T2", "New silicon chips in wearables extend battery life significantly.", {"topic": "hardware"}), # Redundant to T1
("T3", "AR Glasses are gaining traction in enterprise logistics.", {"topic": "ar"}),
("T4", "Health sensors in wearables now monitor blood glucose non-invasively.", {"topic": "health"}),
("T5", "Sustainable recycled plastics are being used in smart bands.", {"topic": "sustainability"})
]
for doc_id, text, meta in docs:
self.collection.add(documents=[text], ids=[doc_id], metadatas=[meta])
self.graph.add_node(doc_id, **meta)
# Link related topics
self.graph.add_edge("T1", "T5") # Hardware linked to Sustainability
def retrieve_mmr(self, query: str, k: int = 3, lambda_mult: float = 0.5) -> List[Dict]:
"""
Retrieve using MMR for diversity.
lambda_mult: 1 = pure relevance, 0 = pure diversity
"""
# 1. Get a larger pool of candidates first (e.g., 10)
results = self.collection.query(query_texts=[query], n_results=10)
if not results['ids'][0]:
return []
# 2. Extract embeddings for MMR calculation
# Note: In production, you'd fetch embeddings from the DB.
# For this PoC, we simulate the MMR logic using LangChain's utility
# which requires embedding vectors. We'll use a simplified approach
# by fetching the actual embeddings from Chroma if stored, or
# simulating the selection logic for the PoC.
# To make this work end-to-end without external embedding calls in the loop,
# we will use Chroma's built-in MMR support if available, or simulate it.
# Chroma supports MMR directly in query!
mmr_results = self.collection.query(
query_texts=[query],
n_results=k,
where=None,
include=["documents", "metadatas"],
# Chroma doesn't have direct MMR param in all versions, so we simulate
# the effect by fetching more and filtering, or using LangChain helper.
# Here we use a manual simulation for clarity of the concept.
)
# Simulating MMR Selection Logic for the PoC
# In a real app, use: from langchain_community.vectorstores.utils import maximal_marginal_relevance
# We will manually pick diverse items from the top 10 for the demo.
docs = results['documents'][0]
ids = results['ids'][0]
# Simple diversity heuristic for PoC: Skip docs with very similar starting words
selected = []
seen_topics = set()
for i, doc in enumerate(docs):
topic = results['metadatas'][0][i]['topic']
if topic not in seen_topics:
selected.append({
"id": ids[i],
"content": doc,
"topic": topic
})
seen_topics.add(topic)
if len(selected) == k:
break
return selected
Step 3: LangGraph Multi-Agent Workflow
# agents/workflow.py
from langgraph.graph import StateGraph, END
from typing import TypedDict, List, Optional
from langchain_openai import ChatOpenAI
from services.mmr_graph_rag import MMRGraphRAG
class ResearchState(TypedDict):
query: str
retrieved_docs: List[dict]
final_report: Optional[str]
diversity_score: float # Metric to show MMR impact
class MarketResearchAgent:
def __init__(self):
self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.1)
self.rag = MMRGraphRAG()
def retrieve_diverse_context(self, state: ResearchState) -> ResearchState:
"""Retriever Agent: Uses MMR"""
# Using lambda_mult=0.5 for balance
state['retrieved_docs'] = self.rag.retrieve_mmr(state['query'], k=3, lambda_mult=0.5)
return state
def generate_report(self, state: ResearchState) -> ResearchState:
"""Analyst Agent: Synthesizes diverse info"""
context = "\n".join([f"[{d['topic']}] {d['content']}" for d in state['retrieved_docs']])
prompt = f"""
Query: {state['query']}
Diverse Market Data:
{context}
Create a concise summary highlighting the different aspects of the market.
"""
response = self.llm.invoke(prompt)
state['final_report'] = response.content
state['diversity_score'] = len(set(d['topic'] for d in state['retrieved_docs'])) / len(state['retrieved_docs'])
return state
def build_workflow():
agent = MarketResearchAgent()
workflow = StateGraph(ResearchState)
workflow.add_node("retrieve", agent.retrieve_diverse_context)
workflow.add_node("report", agent.generate_report)
workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "report")
workflow.add_edge("report", END)
return workflow.compile()
Step 4: FastAPI Backend
# main.py
from fastapi import FastAPI
from pydantic import BaseModel
from agents.workflow import build_workflow
app = FastAPI(title="MMR Market Research API")
workflow = build_workflow()
class QueryRequest(BaseModel):
query: str
class ReportResponse(BaseModel):
report: str
diversity_score: float
sources: list
@app.post("/research", response_model=ReportResponse)
async def research(request: QueryRequest):
initial_state = {
"query": request.query,
"retrieved_docs": [],
"final_report": None,
"diversity_score": 0.0
}
result = await workflow.ainvoke(initial_state)
return ReportResponse(
report=result['final_report'],
diversity_score=result['diversity_score'],
sources=[d['content'] for d in result['retrieved_docs']]
)
if __name__ == "__main__":
import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000)
Step 5: Frontend Interface
<!-- index.html -->
<!DOCTYPE html>
<html>
<head><title>MMR Market Research</title>
<style>
body{font-family:Arial;max-width:800px;margin:40px auto;padding:20px}
.card{background:#fff;padding:20px;border-radius:8px;box-shadow:0 2px 5px rgba(0,0,0,0.1);margin-top:20px}
.metric{color:#28a745;font-weight:bold}
input{width:70%;padding:10px}button{padding:10px 20px;background:#6200ea;color:white;border:none;cursor:pointer}
</style></head>
<body>
<h1>MMR-Enhanced Market Research</h1>
<input id="q" placeholder="e.g., Emerging trends in wearables">
<button onclick="research()">Analyze</button>
<div id="out"></div>
<script>
async function research(){
const q=document.getElementById('q').value;
const r=await fetch('/research',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({query:q})});
const d=await r.json();
document.getElementById('out').innerHTML=`<div class="card">
<h3>Report</h3><p>${d.report}</p>
<p class="metric">Diversity Score: ${d.diversity_score.toFixed(2)} (1.0 is max diversity)</p>
<h4>Sources Used:</h4><ul>${d.sources.map(s=>`<li>${s}</li>`).join('')}</ul>
</div>`;
}
</script></body></html>
Conclusion
Maximum Marginal Relevance (MMR) is a powerful tool for moving beyond simple similarity matching. By explicitly optimizing for diversity, MMR ensures that your RAG system provides a broader, more comprehensive view of the data, which is essential for complex enterprise tasks like market research, legal discovery, and strategic planning. When integrated into a LangGraph multi-agent workflow, MMR becomes part of a robust, stateful pipeline that delivers high-quality, diverse insights to users.

Join the conversation! Your thoughts help the community grow.