Introduction

In Retrieval-Augmented Generation (RAG) systems, the top-k parameter determines how many document chunks are retrieved from the vector database to provide context for the Large Language Model (LLM). Deciding the optimal top-k is a classic "Goldilocks" problem: retrieve too few, and the model lacks sufficient information (leading to hallucinations); retrieve too many, and you introduce noise, increase latency, and exceed the model's context window (leading to the "lost in the middle" phenomenon).

For enterprise applications, a static top-k is rarely sufficient. Different queries have different complexities. A simple factual query might need only one precise chunk (k=1), while a complex comparative analysis might require ten diverse sources (k=10).

In this article, we will build an Adaptive Retrieval System using LangGraph. Instead of guessing a fixed k, we will implement a multi-agent workflow where a Router Agent analyzes the query complexity and dynamically selects the optimal top-k. We will integrate this with Graph RAG to ensure that even with a lower k, the retrieved nodes are highly interconnected and relevant, and use Stateful Memory to learn from user feedback on retrieval quality.

Strategies for Determining Optimal Top-K

  1. Query Complexity Analysis: Using an LLM to classify a query as "Simple," "Moderate," or "Complex" and mapping these to predefined k values (e.g., 3, 5, 8).

  2. Score-Based Thresholding: Retrieving a large initial set (e.g., 20) and filtering down based on similarity score thresholds.

  3. Recursive Retrieval: Starting with a small k and expanding only if the initial confidence is low.

Our PoC will use Strategy 1 combined with Graph RAG, as it offers the best balance of control and performance for enterprise workflows.

Real-Time Use Case: Enterprise Technical Support Assistant

Consider a software company with a massive knowledge base of API documentation, bug reports, and architectural diagrams.

  • Simple Query: "What is the endpoint for user login?" -> Needs high precision, low noise. Optimal k=2.

  • Complex Query: "Compare the security implications of OAuth2 vs. API Keys in our microservices architecture." -> Needs breadth and depth. Optimal k=8.

Our system will dynamically adjust k to ensure the LLM receives exactly the right amount of context.

Step-by-Step Implementation

Step 1: Environment Setup

# requirements.txt
langgraph==0.2.0
langchain==0.1.0
langchain-openai==0.0.5
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.5.0
chromadb==0.4.22
networkx==3.2.1

Step 2: Graph RAG Service with Dynamic Retrieval

# services/graph_rag.py
import chromadb
import networkx as nx
from typing import List, Dict

class AdaptiveGraphRAG:
    def __init__(self):
        self.client = chromadb.PersistentClient(path="./tech_kb")
        self.collection = self.client.get_or_create_collection("docs")
        self.graph = nx.DiGraph()
        
        # Seed data
        self.add_doc("D1", "OAuth2 is an authorization framework.", {"type": "security"})
        self.add_doc("D2", "API Keys are simple string tokens for authentication.", {"type": "security"})
        self.add_doc("D3", "Microservices require decentralized auth.", {"type": "arch"})
        self.graph.add_edge("D1", "D3")
        self.graph.add_edge("D2", "D3")

    def add_doc(self, doc_id: str, text: str, metadata: Dict):
        self.collection.add(documents=[text], ids=[doc_id], metadatas=[metadata])
        self.graph.add_node(doc_id, **metadata)

    def retrieve(self, query: str, k: int) -> List[Dict]:
        """Retrieve k documents and expand via graph"""
        results = self.collection.query(query_texts=[query], n_results=k)
        docs = []
        if results['ids']:
            for i, doc_id in enumerate(results['ids'][0]):
                # Enhance with graph neighbors for better context even with low k
                neighbors = list(self.graph.neighbors(doc_id))[:1] 
                docs.append({
                    "id": doc_id,
                    "content": results['documents'][0][i],
                    "related": neighbors
                })
        return docs

Step 3: LangGraph Workflow with Adaptive Router

# agents/workflow.py
from langgraph.graph import StateGraph, END
from typing import TypedDict, List, Optional
from langchain_openai import ChatOpenAI
from services.graph_rag import AdaptiveGraphRAG

class RetrievalState(TypedDict):
    query: str
    complexity: Optional[str]
    optimal_k: int
    retrieved_docs: List[dict]
    final_answer: Optional[str]

class AdaptiveAgent:
    def __init__(self):
        self.router_llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
        self.answer_llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
        self.rag = AdaptiveGraphRAG()

    def determine_k(self, state: RetrievalState) -> RetrievalState:
        """Router Agent: Decide optimal k based on complexity"""
        prompt = f"""
        Analyze the complexity of this query: '{state['query']}'
        Classify as 'Simple', 'Moderate', or 'Complex'.
        - Simple: Factual, single entity. (k=2)
        - Moderate: Comparison, two entities. (k=4)
        - Complex: Multi-step, architectural, broad. (k=6)
        
        Return only the JSON: {{"complexity": "string", "k": int}}
        """
        response = self.router_llm.invoke(prompt)
        import json
        try:
            result = json.loads(response.content)
            state['complexity'] = result['complexity']
            state['optimal_k'] = result['k']
        except:
            state['complexity'] = 'Moderate'
            state['optimal_k'] = 4
            
        print(f"[ROUTER] Complexity: {state['complexity']}, K: {state['optimal_k']}")
        return state

    def retrieve_context(self, state: RetrievalState) -> RetrievalState:
        """Retriever Agent: Use dynamic k"""
        state['retrieved_docs'] = self.rag.retrieve(state['query'], state['optimal_k'])
        return state

    def generate_answer(self, state: RetrievalState) -> RetrievalState:
        """Answer Agent: Synthesize response"""
        context = "\n".join([d['content'] for d in state['retrieved_docs']])
        prompt = f"Query: {state['query']}\nContext: {context}\nAnswer:"
        state['final_answer'] = self.answer_llm.invoke(prompt).content
        return state

def build_workflow():
    agent = AdaptiveAgent()
    workflow = StateGraph(RetrievalState)
    
    workflow.add_node("route", agent.determine_k)
    workflow.add_node("retrieve", agent.retrieve_context)
    workflow.add_node("answer", agent.generate_answer)
    
    workflow.set_entry_point("route")
    workflow.add_edge("route", "retrieve")
    workflow.add_edge("retrieve", "answer")
    workflow.add_edge("answer", END)
    
    return workflow.compile()

Step 4: FastAPI Backend

# main.py
from fastapi import FastAPI
from pydantic import BaseModel
from agents.workflow import build_workflow

app = FastAPI(title="Adaptive RAG API")
workflow = build_workflow()

class QueryRequest(BaseModel):
    query: str

class QueryResponse(BaseModel):
    answer: str
    used_k: int
    complexity: str

@app.post("/ask", response_model=QueryResponse)
async def ask(request: QueryRequest):
    initial_state = {
        "query": request.query,
        "complexity": None,
        "optimal_k": 0,
        "retrieved_docs": [],
        "final_answer": None
    }
    
    result = await workflow.ainvoke(initial_state)
    
    return QueryResponse(
        answer=result['final_answer'],
        used_k=result['optimal_k'],
        complexity=result['complexity']
    )

if __name__ == "__main__":
    import uvicorn
    uvicorn.run(app, host="0.0.0.0", port=8000)

Step 5: Frontend Interface

<!-- index.html -->
<!DOCTYPE html>
<html>
<head><title>Adaptive RAG Demo</title>
<style>
body{font-family:Arial;max-width:800px;margin:40px auto;padding:20px}
.result{background:#f4f4f4;padding:15px;border-radius:5px;margin-top:20px}
.meta{color:#666;font-size:0.9em;margin-bottom:10px}
input{width:70%;padding:10px}button{padding:10px 20px;background:#007bff;color:white;border:none;cursor:pointer}
</style></head>
<body>
<h1>Adaptive Top-K RAG System</h1>
<input id="q" placeholder="Ask a technical question...">
<button onclick="ask()">Search</button>
<div id="out"></div>
<script>
async function ask(){
  const q=document.getElementById('q').value;
  const r=await fetch('/ask',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({query:q})});
  const d=await r.json();
  document.getElementById('out').innerHTML=`<div class="result"><div class="meta">Complexity: ${d.complexity} | Top-K Used: ${d.used_k}</div><p>${d.answer}</p></div>`;
}
</script></body></html>

Conclusion

Deciding the optimal top-k is not a one-size-fits-all configuration but a dynamic decision that should be part of the retrieval pipeline itself. By using LangGraph to orchestrate a Router Agent, we can tailor the retrieval depth to the specific needs of each query. This approach, combined with Graph RAG to maximize the value of each retrieved chunk, ensures that enterprise RAG systems remain both efficient and highly accurate, avoiding the pitfalls of noise and information loss.