Table of Contents
Introduction to Resilient AI Architectures
The Problem: Dependency on Fragile Downstream Services
Core Principles of Graceful Degradation
Solution Architecture: Fault-Tolerant LangGraph RAG
Technology Stack Overview
Step-by-Step Implementation: Backend Development
Circuit Breaker Pattern for LLM Calls
Fallback Embedding Strategies
Timeout Handling and Retry Logic with Exponential Backoff
Local Cache as a Safety Net
Building the Resilient LangGraph Workflow
State Management with Error Tracking
Frontend Implementation: User-Friendly Error States
Real-Time Use Case: Manufacturing Quality Control Assistant
Monitoring and Observability
Conclusion and Best Practices
Introduction
In enterprise AI applications, relying on external model services whether proprietary LLM APIs, embedding services, or reranking models introduces significant risk. Network latency, service outages, rate limiting, and unexpected errors can cripple your application if not handled properly. Graceful degradation is the design philosophy that ensures your system continues to provide value, even when components fail, by falling back to simpler but reliable alternatives.
This article demonstrates how to build an enterprise-grade multi-agent RAG system using LangGraph that maintains functionality during downstream service failures. We implement circuit breakers, fallback strategies, local caching, and intelligent error handling to ensure users always receive a response, even if it’s less sophisticated than the ideal path.
Technology Tags
Python, LangGraph, LangChain, FastAPI, React, Redis, PostgreSQL, Tenacity, Pybreaker, Docker, Prometheus, Grafana, TypeScript, TailwindCSS, Sentry
Step-by-Step Implementation
1. Circuit Breaker and Retry Logic
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import pybreaker
import time
from functools import wraps
# Circuit breaker for LLM service
llm_circuit_breaker = pybreaker.CircuitBreaker(
failure_threshold=5,
recovery_timeout=60,
expected_exception=Exception
)
# Circuit breaker for embedding service
embedding_circuit_breaker = pybreaker.CircuitBreaker(
failure_threshold=3,
recovery_timeout=30,
expected_exception=Exception
)
class ServiceUnavailableError(Exception):
pass
def resilient_llm_call(func):
"""Decorator adding circuit breaker and retry logic"""
@wraps(func)
@llm_circuit_breaker
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential(multiplier=1, min=2, max=10),
retry=retry_if_exception_type((TimeoutError, ConnectionError))
)
def wrapper(*args, **kwargs):
try:
return func(*args, **kwargs)
except Exception as e:
raise ServiceUnavailableError(f"LLM service failed: {str(e)}")
return wrapper
def resilient_embedding_call(func):
"""Decorator for embedding service resilience"""
@wraps(func)
@embedding_circuit_breaker
@retry(
stop=stop_after_attempt(2),
wait=wait_exponential(multiplier=0.5, min=1, max=5),
retry=retry_if_exception_type((TimeoutError, ConnectionError))
)
def wrapper(*args, **kwargs):
try:
return func(*args, **kwargs)
except Exception as e:
raise ServiceUnavailableError(f"Embedding service failed: {str(e)}")
return wrapper
2. Fallback Embedding Strategy
from langchain_huggingface import HuggingFaceEmbeddings
from sentence_transformers import SentenceTransformer
import numpy as np
class ResilientEmbeddingService:
def __init__(self):
# Primary: Cloud-based embedding API
self.primary_model = None # e.g., OpenAI embeddings
# Fallback: Local lightweight model
self.fallback_model = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
# Emergency: Simple TF-IDF based approach
self.emergency_mode = False
@resilient_embedding_call
def embed_with_primary(self, texts: list) -> list:
"""Try primary cloud embedding service"""
# Simulate API call
raise TimeoutError("Primary service timeout")
def embed_with_fallback(self, texts: list) -> list:
"""Use local model when primary fails"""
print(" Using fallback embedding model")
return self.fallback_model.embed_documents(texts)
def embed_with_tfidf(self, texts: list) -> list:
"""Emergency: Simple keyword-based similarity"""
print(" Emergency mode: Using TF-IDF approximation")
# Simplified implementation
from sklearn.feature_extraction.text import TfidfVectorizer
vectorizer = TfidfVectorizer()
tfidf_matrix = vectorizer.fit_transform(texts)
return tfidf_matrix.toarray().tolist()
def embed_documents(self, texts: list) -> list:
"""Graceful degradation chain for embeddings"""
try:
return self.embed_with_primary(texts)
except ServiceUnavailableError:
try:
return self.embed_with_fallback(texts)
except Exception:
self.emergency_mode = True
return self.embed_with_tfidf(texts)
3. Local Cache as Safety Net
import redis
import json
import hashlib
class ResponseCache:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
self.ttl = 3600 # 1 hour
def _generate_key(self, query: str) -> str:
query_hash = hashlib.md5(query.encode()).hexdigest()
return f"cache:{query_hash}"
def get_cached_response(self, query: str) -> dict:
"""Retrieve cached response if available"""
key = self._generate_key(query)
cached = self.redis.get(key)
if cached:
print("✅ Serving from cache")
return json.loads(cached)
return None
def cache_response(self, query: str, response: dict):
"""Store response in cache"""
key = self._generate_key(query)
self.redis.setex(key, self.ttl, json.dumps(response))
4. Resilient LangGraph Multi-Agent System
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage, SystemMessage
from typing import List, Dict, Any, TypedDict, Optional
from datetime import datetime
class AgentState(TypedDict):
messages: List
retrieved_context: List[Dict]
final_answer: str
conversation_id: str
error_log: List[str]
degradation_level: int # 0=normal, 1=partial, 2=emergency
use_cache: bool
timestamp: str
class ResilientManufacturingAgent:
def __init__(self, embedding_service: ResilientEmbeddingService,
cache: ResponseCache):
self.embedding_service = embedding_service
self.cache = cache
self.llm = self._initialize_llm()
def _initialize_llm(self):
from langchain_openai import ChatOpenAI
return ChatOpenAI(model="gpt-4", temperature=0.1, timeout=10)
@resilient_llm_call
def call_llm_normal(self, prompt: str) -> str:
"""Normal LLM call with full context"""
response = self.llm.invoke(prompt)
return response.content
def call_llm_fallback(self, prompt: str) -> str:
"""Fallback: Use smaller/faster model or cached responses"""
print(" Using fallback LLM strategy")
# Option 1: Use smaller local model
# Option 2: Return template-based response
return "Based on available manufacturing data, please consult the equipment manual section 4.2 for detailed procedures. Our system is experiencing high load."
def retrieve_node(self, state: AgentState) -> AgentState:
"""Resilient retrieval with degradation tracking"""
query = state["messages"][-1].content
# Check cache first
cached = self.cache.get_cached_response(query)
if cached:
state["use_cache"] = True
state["retrieved_context"] = cached.get("context", [])
state["degradation_level"] = 0
return state
try:
# Normal retrieval path
embeddings = self.embedding_service.embed_documents([query])
# Perform vector search...
state["retrieved_context"] = [{"content": "Sample context", "source": "vector"}]
state["degradation_level"] = 0
except ServiceUnavailableError as e:
state["error_log"].append(f"Retrieval error: {str(e)}")
state["degradation_level"] = 1
state["retrieved_context"] = [{"content": "Limited context available",
"source": "fallback"}]
return state
def generate_answer_node(self, state: AgentState) -> AgentState:
"""Generate answer with graceful degradation"""
context_text = "\n".join([ctx["content"] for ctx in state["retrieved_context"]])
query = state["messages"][-1].content
prompt = f"""Context: {context_text}
Question: {query}
Provide a helpful manufacturing-related answer."""
try:
if state["degradation_level"] == 0:
answer = self.call_llm_normal(prompt)
else:
answer = self.call_llm_fallback(prompt)
state["degradation_level"] = max(state["degradation_level"], 1)
except ServiceUnavailableError as e:
state["error_log"].append(f"Generation error: {str(e)}")
answer = "I'm unable to process your request at the moment due to technical issues. Please try again later or contact support."
state["degradation_level"] = 2
state["final_answer"] = answer
state["messages"].append(AIMessage(content=answer))
state["timestamp"] = str(datetime.now())
# Cache successful responses
if state["degradation_level"] == 0:
self.cache.cache_response(query, {
"context": state["retrieved_context"],
"answer": answer
})
return state
def build_graph(self) -> StateGraph:
"""Build resilient workflow"""
workflow = StateGraph(AgentState)
workflow.add_node("retrieve", self.retrieve_node)
workflow.add_node("generate", self.generate_answer_node)
workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
return workflow.compile()
5. FastAPI Backend with Health Checks
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
app = FastAPI(title="Resilient Manufacturing RAG API")
class QueryRequest(BaseModel):
question: str
conversation_id: str = "default"
class QueryResponse(BaseModel):
answer: str
degradation_level: int
sources: List[Dict]
from_cache: bool
warnings: List[str]
@app.get("/health")
async def health_check():
"""Comprehensive health check"""
status = {
"status": "healthy",
"llm_circuit_breaker": llm_circuit_breaker.current_state,
"embedding_circuit_breaker": embedding_circuit_breaker.current_state,
"timestamp": str(datetime.now())
}
if llm_circuit_breaker.current_state == "open":
status["status"] = "degraded"
return status
@app.post("/query", response_model=QueryResponse)
async def handle_query(request: QueryRequest):
agent = ResilientManufacturingAgent(embedding_service, cache)
graph = agent.build_graph()
initial_state = AgentState(
messages=[HumanMessage(content=request.question)],
retrieved_context=[],
final_answer="",
conversation_id=request.conversation_id,
error_log=[],
degradation_level=0,
use_cache=False,
timestamp=""
)
result = graph.invoke(initial_state)
return QueryResponse(
answer=result["final_answer"],
degradation_level=result["degradation_level"],
sources=result["retrieved_context"],
from_cache=result["use_cache"],
warnings=result["error_log"]
)
6. React Frontend with Degradation Indicators
// components/ResilientChat.tsx
import React, { useState } from 'react';
interface QueryResponse {
answer: string;
degradation_level: number;
from_cache: boolean;
warnings: string[];
}
export const ResilientChat: React.FC = () => {
const [messages, setMessages] = useState<any[]>([]);
const [input, setInput] = useState('');
const getDegradationMessage = (level: number) => {
switch(level) {
case 0: return null;
case 1: return " Running in reduced capability mode";
case 2: return " Limited functionality available";
default: return null;
}
};
const sendMessage = async () => {
const response = await fetch('/api/query', {
method: 'POST',
body: JSON.stringify({ question: input })
});
const data: QueryResponse = await response.json();
setMessages(prev => [...prev, {
role: 'assistant',
content: data.answer,
warning: getDegradationMessage(data.degradation_level),
fromCache: data.from_cache
}]);
};
return (
<div className="chat-interface">
{messages.map((msg, idx) => (
<div key={idx} className="message">
{msg.warning && <div className="warning-banner">{msg.warning}</div>}
{msg.fromCache && <span className="cache-badge">Cached</span>}
<p>{msg.content}</p>
</div>
))}
<input value={input} onChange={e => setInput(e.target.value)} />
<button onClick={sendMessage}>Send</button>
</div>
);
};
Real-Time Use Case: Quality Control Assistant
A quality engineer asks: "What are the acceptance criteria for batch #QC-2024-089?"
Normal Operation: System retrieves specific QC documents, generates detailed answer with references.
Partial Degradation: Embedding service down. System uses local model, returns answer with fewer contextual details but still accurate.
Emergency Mode: Both LLM and embeddings unavailable. System serves cached response or provides template: "Please refer to QC Manual Section 5.3 for batch acceptance criteria. Contact QA lead for specifics."
Conclusion
Graceful degradation transforms fragile AI systems into resilient enterprise solutions. By implementing circuit breakers, fallback strategies, caching, and clear user communication, your multi-agent RAG system maintains value delivery even during service disruptions. The key is designing multiple degradation levels that progressively simplify functionality while preserving core utility. Monitor your circuit breakers, test failure scenarios regularly, and always prioritize user experience over perfect accuracy when systems are under stress. This approach ensures your manufacturing AI assistant remains a reliable tool, not a single point of failure.

Join the conversation! Your thoughts help the community grow.