Introduction
In the enterprise adoption of Large Language Models (LLMs), token efficiency is not merely a cost-saving measure; it is a critical performance and scalability lever. Every token processed—whether in the input prompt or the generated output—consumes computational resources, increases latency, and contributes to the overall operational budget. For high-volume applications like customer support, legal document analysis, or real-time code generation, inefficient token usage can lead to prohibitive costs and sluggish user experiences. Token efficiency strategies involve optimizing every stage of the LLM pipeline: from compressing input context and selecting relevant data chunks to guiding the model toward concise outputs. However, implementing these strategies in isolation often leads to fragmented systems. The true power lies in orchestrating them through a multi-agent architecture. This article explores how to build a token-efficient enterprise RAG system using LangGraph for workflow orchestration, Graph RAG for precise data retrieval, and FastAPI with a React frontend. We will demonstrate how to dynamically choose between summarization, chunking, and structured output generation to minimize token waste while maximizing answer quality.
Real-Time Use Case: High-Volume Technical Support Assistant
Consider a SaaS company providing a technical support assistant for its API documentation. Developers frequently ask complex questions that require referencing multiple pages of dense technical manuals. A naive RAG system might retrieve five full pages of documentation (thousands of tokens) and feed them into the LLM, resulting in high costs and slow response times due to the "lost in the middle" phenomenon.
Our token-efficient system will:
Analyze Query Complexity: Determine if a simple lookup or deep reasoning is needed.
Optimize Retrieval: Use Graph RAG to fetch only the most relevant semantic nodes rather than entire documents.
Compress Context: Summarize retrieved chunks before passing them to the final generator.
Enforce Concise Output: Use structured output constraints to prevent verbose, unnecessary explanations.
Architecture Overview
Router Agent (LangGraph): Classifies query complexity to decide the retrieval strategy.
Graph RAG Engine (Neo4j): Retrieves specific conceptual nodes instead of raw text blocks.
Context Compressor Agent: Uses a smaller, cheaper model to summarize retrieved context.
Structured Generator Agent: Produces JSON-formatted answers to ensure brevity and parsability.
State Management: Tracks token counts at each step to monitor efficiency gains.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
tiktoken==0.6.0
Step 2: Defining the Efficiency-Aware State
We track token usage throughout the workflow to provide visibility into cost savings.
from typing import TypedDict, List, Optional, Literal
from pydantic import BaseModel, Field
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
def count_tokens(text: str) -> int:
return len(enc.encode(text))
class EfficiencyState(TypedDict):
user_query: str
strategy: Optional[Literal["simple", "complex"]]
raw_retrieved_context: List[str]
compressed_context: Optional[str]
final_answer: str
input_tokens_used: int
output_tokens_used: int
audit_log: List[str]
class SupportQuery(BaseModel):
question: str = Field(..., description="Developer's technical question")
Step 3: The Router and Graph RAG Retrieval
The router decides the path, and Graph RAG ensures we only pull necessary nodes.
from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, END
from neo4j import GraphDatabase
llm = ChatOpenAI(model="gpt-4-turbo")
compressor_llm = ChatOpenAI(model="gpt-3.5-turbo") # Cheaper model for compression
def route_query(state: EfficiencyState) -> EfficiencyState:
"""Decides if we need deep context or a simple lookup"""
prompt = f"Is '{state['user_query']}' a simple factual lookup or a complex reasoning task? Reply 'simple' or 'complex'."
response = llm.invoke(prompt).content.strip().lower()
state['strategy'] = response if response in ['simple', 'complex'] else 'complex'
state['audit_log'].append(f"Routing strategy: {state['strategy']}")
return state
def retrieve_graph_context(state: EfficiencyState) -> EfficiencyState:
"""Retrieves specific nodes from Neo4j based on query semantics"""
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
with driver.session() as session:
# Semantic search via Graph relationships
result = session.run("""
MATCH (c:Concept)-[:RELATED_TO]->(k:KnowledgeChunk)
WHERE c.name CONTAINS $query
RETURN k.text LIMIT 3
""", query=state['user_query'])
state['raw_retrieved_context'] = [record['k.text'] for record in result]
initial_tokens = count_tokens(state['user_query']) + sum(count_tokens(c) for c in state['raw_retrieved_context'])
state['input_tokens_used'] = initial_tokens
state['audit_log'].append(f"Retrieved {len(state['raw_retrieved_context'])} chunks. Initial tokens: {initial_tokens}")
return state
Step 4: Context Compression and Structured Generation
This is where significant token savings occur. We compress the context and force a concise output format.
from pydantic import BaseModel
class ConciseAnswer(BaseModel):
answer: str = Field(..., description="Direct answer in max 2 sentences")
reference_id: str = Field(..., description="ID of the source concept")
def compress_and_generate(state: EfficiencyState) -> EfficiencyState:
"""Compresses context and generates a structured, token-efficient response"""
# 1. Compress Context using a cheaper model
context_text = "\n".join(state['raw_retrieved_context'])
compress_prompt = f"Summarize the following technical context into key bullet points relevant to: '{state['user_query']}'. Keep it under 100 words.\nContext: {context_text}"
state['compressed_context'] = compressor_llm.invoke(compress_prompt).content
# 2. Generate Structured Output
final_prompt = f"Context: {state['compressed_context']}\nQuestion: {state['user_query']}"
# Using structured output to prevent verbosity
structured_llm = llm.with_structured_output(ConciseAnswer)
result = structured_llm.invoke(final_prompt)
state['final_answer'] = result.answer
state['output_tokens_used'] = count_tokens(result.answer)
state['audit_log'].append(f"Final tokens used: {state['output_tokens_used']}")
return state
Step 5: Orchestrating with LangGraph
workflow = StateGraph(EfficiencyState)
workflow.add_node("route", route_query)
workflow.add_node("retrieve", retrieve_graph_context)
workflow.add_node("process", compress_and_generate)
workflow.set_entry_point("route")
workflow.add_edge("route", "retrieve")
workflow.add_edge("retrieve", "process")
workflow.add_edge("process", END)
app = workflow.compile()
Step 6: FastAPI Backend and React Frontend
# Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
api_app = FastAPI(title="Token Efficient RAG")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])
@api_app.post("/ask-support")
async def ask_support(query: SupportQuery):
initial_state = EfficiencyState(
user_query=query.question,
strategy=None,
raw_retrieved_context=[],
compressed_context=None,
final_answer="",
input_tokens_used=0,
output_tokens_used=0,
audit_log=[]
)
result = await app.ainvoke(initial_state)
return {
"answer": result['final_answer'],
"efficiency_metrics": {
"input_tokens": result['input_tokens_used'],
"output_tokens": result['output_tokens_used'],
"total_tokens": result['input_tokens_used'] + result['output_tokens_used']
},
"audit_trail": result['audit_log']
}
// Frontend
import React, { useState } from 'react';
import axios from 'axios';
const EfficientSupportChat = () => {
const [q, setQ] = useState('');
const [res, setRes] = useState(null);
const send = async () => {
const data = await axios.post('http://localhost:8000/ask-support', { question: q });
setRes(data.data);
};
return (
<div className="p-4 max-w-lg mx-auto">
<h1 className="text-xl font-bold">Efficient Tech Support</h1>
<input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
<button onClick={send} className="bg-blue-500 text-white p-2 mt-2">Ask</button>
{res && (
<div className="mt-4 p-3 border rounded">
<p>{res.answer}</p>
<div className="text-xs text-gray-500 mt-2">
Tokens Used: {res.efficiency_metrics.total_tokens}
(In: {res.efficiency_metrics.input_tokens}, Out: {res.efficiency_metrics.output_tokens})
</div>
</div>
)}
</div>
);
};
export default EfficientSupportChat;
Conclusion
Token efficiency is achieved not by a single trick, but by a holistic architectural strategy. By using LangGraph to orchestrate a pipeline that includes intelligent routing, Graph RAG for precise retrieval, and context compression, enterprises can significantly reduce their LLM operational costs. This multi-agent approach ensures that every token spent contributes directly to the quality of the answer, transforming AI from a costly experiment into a scalable, efficient business asset.

Join the conversation! Your thoughts help the community grow.