Introduction

The transition of Large Language Models (LLMs) from experimental prototypes to enterprise production environments has brought a new challenge to the forefront: cost management. Unlike traditional software where compute costs are relatively predictable, LLM usage is variable, driven by token consumption, model complexity, and request volume. Without strict controls, an unchecked AI agent can quickly escalate operational expenses through redundant retrievals, verbose outputs, and the use of oversized models for simple tasks. Controlling cost in LLM systems requires more than just monitoring bills; it demands an architectural shift toward efficiency-by-design. This involves implementing strategies such as dynamic model routing, context compression, caching, and strict output token limits. In an enterprise setting, these strategies must be orchestrated seamlessly to maintain high-quality user experiences while adhering to budgetary constraints.

This article details how to build a cost-conscious AI system using LangGraph for multi-agent orchestration, Graph RAG for efficient data retrieval, and FastAPI for backend integration. We will demonstrate a real-world Proof of Concept (POC) for a corporate knowledge assistant that dynamically adjusts its resource usage based on query complexity.

Real-Time Use Case: Corporate Knowledge Assistant

Consider a large organization with thousands of internal documents. Employees use an AI assistant to find information about HR policies, technical protocols, and project histories.

A naive implementation might send every query to the most expensive, high-reasoning model (e.g., GPT-4o) and retrieve entire documents for context. This leads to two major cost drivers:

  1. Over-provisioning: Using a "Ferrari" engine for a "grocery run" query like "What is the holiday schedule?"

  2. Context Bloat: Feeding the LLM thousands of tokens of irrelevant document text.

Our cost-controlled system will:

  1. Classify Intent: Use a cheap, fast model to determine if the query is simple or complex.

  2. Route Efficiently: Direct simple queries to a smaller, cheaper model and complex ones to a larger model.

  3. Optimize Retrieval: Use Graph RAG to fetch only specific, relevant nodes rather than whole files.

  4. Enforce Limits: Apply hard caps on output tokens to prevent verbose, costly generations.

Architecture Overview

  • Cost Monitor Agent (LangGraph): Tracks token usage and enforces budget thresholds per session.

  • Router Agent: Classifies query complexity to select the appropriate model tier.

  • Graph RAG Engine (Neo4j): Retrieves precise semantic chunks to minimize input context size.

  • Tiered LLM Pool: A mix of cost-effective models (e.g., GPT-3.5-turbo) and high-performance models (e.g., GPT-4-turbo).

  • FastAPI & React: Provides the interface for users and administrators to monitor cost metrics in real-time.

Step-by-Step Implementation

Step 1: Environment Setup

# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
tiktoken==0.6.0

Step 2: Defining the Cost-Aware State

We maintain a state that tracks cumulative costs and token counts to make real-time routing decisions.

from typing import TypedDict, List, Optional, Literal
from pydantic import BaseModel, Field
import tiktoken

enc = tiktoken.get_encoding("cl100k_base")

def estimate_cost(input_tokens: int, output_tokens: int, model_tier: str) -> float:
    """Rough cost estimation based on OpenAI pricing"""
    if model_tier == "premium":
        return (input_tokens * 0.01 + output_tokens * 0.03) / 1000 # GPT-4 approx
    else:
        return (input_tokens * 0.0005 + output_tokens * 0.0015) / 1000 # GPT-3.5 approx

class CostState(TypedDict):
    user_query: str
    model_tier: Optional[Literal["economy", "premium"]]
    retrieved_context: List[str]
    final_answer: str
    total_input_tokens: int
    total_output_tokens: int
    estimated_cost_usd: float
    audit_log: List[str]

class UserQuery(BaseModel):
    question: str = Field(..., description="Employee question")
    session_id: str = Field(..., description="Unique session identifier")

Step 3: The Router and Cost Monitor Agent

This node decides which model to use based on the query's perceived complexity.

from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, END

# Define two tiers of models
economy_llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
premium_llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)

def classify_and_route(state: CostState) -> CostState:
    """Uses a cheap model to classify intent and route accordingly"""
    prompt = f"Is the following query a simple factual lookup or does it require complex reasoning? Reply 'simple' or 'complex'. Query: '{state['user_query']}'"
    response = economy_llm.invoke(prompt).content.strip().lower()
    
    if "simple" in response:
        state['model_tier'] = "economy"
        state['audit_log'].append("Routed to Economy Tier (GPT-3.5)")
    else:
        state['model_tier'] = "premium"
        state['audit_log'].append("Routed to Premium Tier (GPT-4)")
        
    return state

Step 4: Efficient Graph RAG Retrieval

Instead of retrieving full documents, we use Neo4j to find specific conceptual nodes, drastically reducing input tokens.

from neo4j import GraphDatabase

class EfficientRetriever:
    def __init__(self):
        self.driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

    def get_relevant_nodes(self, query: str) -> List[str]:
        with self.driver.session() as session:
            # Retrieve only specific knowledge chunks linked to concepts
            result = session.run("""
                MATCH (c:Concept)-[:HAS_DETAIL]->(k:Chunk)
                WHERE c.name CONTAINS $q
                RETURN k.text LIMIT 2
            """, q=query)
            return [record['k.text'] for record in result]

def retrieve_context(state: CostState) -> CostState:
    retriever = EfficientRetriever()
    state['retrieved_context'] = retriever.get_relevant_nodes(state['user_query'])
    
    # Calculate input tokens
    context_text = " ".join(state['retrieved_context'])
    state['total_input_tokens'] = len(enc.encode(state['user_query'] + context_text))
    state['audit_log'].append(f"Retrieved {len(state['retrieved_context'])} nodes. Input tokens: {state['total_input_tokens']}")
    return state

Step 5: Generation with Cost Controls

We select the model based on the routing decision and enforce a max token limit.

def generate_response(state: CostState) -> CostState:
    context_str = "\n".join(state['retrieved_context'])
    prompt = f"Context: {context_str}\nQuestion: {state['user_query']}\nAnswer concisely."
    
    # Select model based on tier
    llm = premium_llm if state['model_tier'] == "premium" else economy_llm
    
    # Invoke with max_tokens to control output cost
    response = llm.invoke(prompt, max_tokens=150) 
    
    state['final_answer'] = response.content
    state['total_output_tokens'] = len(enc.encode(response.content))
    
    # Update cost
    state['estimated_cost_usd'] = estimate_cost(
        state['total_input_tokens'], 
        state['total_output_tokens'], 
        state['model_tier']
    )
    state['audit_log'].append(f"Generated answer. Est. Cost: ${state['estimated_cost_usd']:.6f}")
    return state

Step 6: Orchestrating with LangGraph

workflow = StateGraph(CostState)
workflow.add_node("route", classify_and_route)
workflow.add_node("retrieve", retrieve_context)
workflow.add_node("generate", generate_response)

workflow.set_entry_point("route")
workflow.add_edge("route", "retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)

app = workflow.compile()

Step 7: FastAPI Backend and React Frontend

# Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

api_app = FastAPI(title="Cost-Controlled AI")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])

@api_app.post("/ask")
async def ask(query: UserQuery):
    initial_state = CostState(
        user_query=query.question,
        model_tier=None,
        retrieved_context=[],
        final_answer="",
        total_input_tokens=0,
        total_output_tokens=0,
        estimated_cost_usd=0.0,
        audit_log=[]
    )
    result = await app.ainvoke(initial_state)
    return {
        "answer": result['final_answer'],
        "cost_metrics": {
            "tier_used": result['model_tier'],
            "input_tokens": result['total_input_tokens'],
            "output_tokens": result['total_output_tokens'],
            "est_cost_usd": result['estimated_cost_usd']
        },
        "audit": result['audit_log']
    }
// Frontend
import React, { useState } from 'react';
import axios from 'axios';

const CostAwareChat = () => {
  const [q, setQ] = useState('');
  const [res, setRes] = useState(null);

  const send = async () => {
    const data = await axios.post('http://localhost:8000/ask', {
      question: q,
      session_id: 'SESSION_001'
    });
    setRes(data.data);
  };

  return (
    <div className="p-4 max-w-md mx-auto">
      <h2 className="text-xl font-bold">Corporate Assistant</h2>
      <input className="border p-2 w-full" value={q} onChange={e => setQ(e.target.value)} />
      <button onClick={send} className="bg-green-600 text-white p-2 mt-2 rounded">Ask</button>
      
      {res && (
        <div className="mt-4 p-3 bg-gray-100 rounded">
          <p>{res.answer}</p>
          <div className="text-xs text-gray-600 mt-2">
            <p>Tier: {res.cost_metrics.tier_used}</p>
            <p>Est. Cost: ${res.cost_metrics.est_cost_usd.toFixed(6)}</p>
          </div>
        </div>
      )}
    </div>
  );
};
export default CostAwareChat;

Conclusion

Controlling costs in LLM systems is not about restricting capability, but about optimizing resource allocation. By implementing a multi-agent architecture with LangGraph, enterprises can dynamically route queries to the most cost-effective model tier and use Graph RAG to minimize context bloat. This approach ensures that every dollar spent on AI contributes directly to business value, transforming LLMs from a financial liability into a scalable, efficient asset.