Introduction
The transition of Large Language Models (LLMs) from experimental prototypes to enterprise production environments has brought a new challenge to the forefront: cost management. Unlike traditional software where compute costs are relatively predictable, LLM usage is variable, driven by token consumption, model complexity, and request volume. Without strict controls, an unchecked AI agent can quickly escalate operational expenses through redundant retrievals, verbose outputs, and the use of oversized models for simple tasks. Controlling cost in LLM systems requires more than just monitoring bills; it demands an architectural shift toward efficiency-by-design. This involves implementing strategies such as dynamic model routing, context compression, caching, and strict output token limits. In an enterprise setting, these strategies must be orchestrated seamlessly to maintain high-quality user experiences while adhering to budgetary constraints.
This article details how to build a cost-conscious AI system using LangGraph for multi-agent orchestration, Graph RAG for efficient data retrieval, and FastAPI for backend integration. We will demonstrate a real-world Proof of Concept (POC) for a corporate knowledge assistant that dynamically adjusts its resource usage based on query complexity.
Real-Time Use Case: Corporate Knowledge Assistant
Consider a large organization with thousands of internal documents. Employees use an AI assistant to find information about HR policies, technical protocols, and project histories.
A naive implementation might send every query to the most expensive, high-reasoning model (e.g., GPT-4o) and retrieve entire documents for context. This leads to two major cost drivers:
Over-provisioning: Using a "Ferrari" engine for a "grocery run" query like "What is the holiday schedule?"
Context Bloat: Feeding the LLM thousands of tokens of irrelevant document text.
Our cost-controlled system will:
Classify Intent: Use a cheap, fast model to determine if the query is simple or complex.
Route Efficiently: Direct simple queries to a smaller, cheaper model and complex ones to a larger model.
Optimize Retrieval: Use Graph RAG to fetch only specific, relevant nodes rather than whole files.
Enforce Limits: Apply hard caps on output tokens to prevent verbose, costly generations.
Architecture Overview
Cost Monitor Agent (LangGraph): Tracks token usage and enforces budget thresholds per session.
Router Agent: Classifies query complexity to select the appropriate model tier.
Graph RAG Engine (Neo4j): Retrieves precise semantic chunks to minimize input context size.
Tiered LLM Pool: A mix of cost-effective models (e.g., GPT-3.5-turbo) and high-performance models (e.g., GPT-4-turbo).
FastAPI & React: Provides the interface for users and administrators to monitor cost metrics in real-time.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
tiktoken==0.6.0
Step 2: Defining the Cost-Aware State
We maintain a state that tracks cumulative costs and token counts to make real-time routing decisions.
from typing import TypedDict, List, Optional, Literal
from pydantic import BaseModel, Field
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
def estimate_cost(input_tokens: int, output_tokens: int, model_tier: str) -> float:
"""Rough cost estimation based on OpenAI pricing"""
if model_tier == "premium":
return (input_tokens * 0.01 + output_tokens * 0.03) / 1000 # GPT-4 approx
else:
return (input_tokens * 0.0005 + output_tokens * 0.0015) / 1000 # GPT-3.5 approx
class CostState(TypedDict):
user_query: str
model_tier: Optional[Literal["economy", "premium"]]
retrieved_context: List[str]
final_answer: str
total_input_tokens: int
total_output_tokens: int
estimated_cost_usd: float
audit_log: List[str]
class UserQuery(BaseModel):
question: str = Field(..., description="Employee question")
session_id: str = Field(..., description="Unique session identifier")
Step 3: The Router and Cost Monitor Agent
This node decides which model to use based on the query's perceived complexity.
from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, END
# Define two tiers of models
economy_llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
premium_llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)
def classify_and_route(state: CostState) -> CostState:
"""Uses a cheap model to classify intent and route accordingly"""
prompt = f"Is the following query a simple factual lookup or does it require complex reasoning? Reply 'simple' or 'complex'. Query: '{state['user_query']}'"
response = economy_llm.invoke(prompt).content.strip().lower()
if "simple" in response:
state['model_tier'] = "economy"
state['audit_log'].append("Routed to Economy Tier (GPT-3.5)")
else:
state['model_tier'] = "premium"
state['audit_log'].append("Routed to Premium Tier (GPT-4)")
return state
Step 4: Efficient Graph RAG Retrieval
Instead of retrieving full documents, we use Neo4j to find specific conceptual nodes, drastically reducing input tokens.
from neo4j import GraphDatabase
class EfficientRetriever:
def __init__(self):
self.driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
def get_relevant_nodes(self, query: str) -> List[str]:
with self.driver.session() as session:
# Retrieve only specific knowledge chunks linked to concepts
result = session.run("""
MATCH (c:Concept)-[:HAS_DETAIL]->(k:Chunk)
WHERE c.name CONTAINS $q
RETURN k.text LIMIT 2
""", q=query)
return [record['k.text'] for record in result]
def retrieve_context(state: CostState) -> CostState:
retriever = EfficientRetriever()
state['retrieved_context'] = retriever.get_relevant_nodes(state['user_query'])
# Calculate input tokens
context_text = " ".join(state['retrieved_context'])
state['total_input_tokens'] = len(enc.encode(state['user_query'] + context_text))
state['audit_log'].append(f"Retrieved {len(state['retrieved_context'])} nodes. Input tokens: {state['total_input_tokens']}")
return state
Step 5: Generation with Cost Controls
We select the model based on the routing decision and enforce a max token limit.
def generate_response(state: CostState) -> CostState:
context_str = "\n".join(state['retrieved_context'])
prompt = f"Context: {context_str}\nQuestion: {state['user_query']}\nAnswer concisely."
# Select model based on tier
llm = premium_llm if state['model_tier'] == "premium" else economy_llm
# Invoke with max_tokens to control output cost
response = llm.invoke(prompt, max_tokens=150)
state['final_answer'] = response.content
state['total_output_tokens'] = len(enc.encode(response.content))
# Update cost
state['estimated_cost_usd'] = estimate_cost(
state['total_input_tokens'],
state['total_output_tokens'],
state['model_tier']
)
state['audit_log'].append(f"Generated answer. Est. Cost: ${state['estimated_cost_usd']:.6f}")
return state
Step 6: Orchestrating with LangGraph
workflow = StateGraph(CostState)
workflow.add_node("route", classify_and_route)
workflow.add_node("retrieve", retrieve_context)
workflow.add_node("generate", generate_response)
workflow.set_entry_point("route")
workflow.add_edge("route", "retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
app = workflow.compile()
Step 7: FastAPI Backend and React Frontend
# Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
api_app = FastAPI(title="Cost-Controlled AI")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])
@api_app.post("/ask")
async def ask(query: UserQuery):
initial_state = CostState(
user_query=query.question,
model_tier=None,
retrieved_context=[],
final_answer="",
total_input_tokens=0,
total_output_tokens=0,
estimated_cost_usd=0.0,
audit_log=[]
)
result = await app.ainvoke(initial_state)
return {
"answer": result['final_answer'],
"cost_metrics": {
"tier_used": result['model_tier'],
"input_tokens": result['total_input_tokens'],
"output_tokens": result['total_output_tokens'],
"est_cost_usd": result['estimated_cost_usd']
},
"audit": result['audit_log']
}
// Frontend
import React, { useState } from 'react';
import axios from 'axios';
const CostAwareChat = () => {
const [q, setQ] = useState('');
const [res, setRes] = useState(null);
const send = async () => {
const data = await axios.post('http://localhost:8000/ask', {
question: q,
session_id: 'SESSION_001'
});
setRes(data.data);
};
return (
<div className="p-4 max-w-md mx-auto">
<h2 className="text-xl font-bold">Corporate Assistant</h2>
<input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
<button onClick={send} className="bg-green-600 text-white p-2 mt-2 rounded">Ask</button>
{res && (
<div className="mt-4 p-3 bg-gray-100 rounded">
<p>{res.answer}</p>
<div className="text-xs text-gray-600 mt-2">
<p>Tier: {res.cost_metrics.tier_used}</p>
<p>Est. Cost: ${res.cost_metrics.est_cost_usd.toFixed(6)}</p>
</div>
</div>
)}
</div>
);
};
export default CostAwareChat;
Conclusion
Controlling costs in LLM systems is not about restricting capability, but about optimizing resource allocation. By implementing a multi-agent architecture with LangGraph, enterprises can dynamically route queries to the most cost-effective model tier and use Graph RAG to minimize context bloat. This approach ensures that every dollar spent on AI contributes directly to business value, transforming LLMs from a financial liability into a scalable, efficient asset.

Join the conversation! Your thoughts help the community grow.