Table of Contents
Introduction: The Evaluation Imperative in Enterprise RAG
The Four Pillars of RAG Evaluation
Key Metrics: From Retrieval to Generation
Evaluation Frameworks: RAGAS, DeepEval, and LLM-as-Judge
Solution Architecture: The "RAG Evaluation Engine" Multi-Agent System
Technology Stack Overview
Step-by-Step Implementation: Backend Development
Defining the Evaluation State Schema with Memory
Building the Context Precision Agent
Implementing the Context Recall Agent
Creating the Faithfulness Agent
Designing the Answer Relevancy Agent
Building the Answer Correctness Agent
Constructing the Aggregator and Report Agent
Building the LangGraph Workflow with Parallel Execution
Frontend Implementation: Evaluation Dashboard
Real-Time Use Case: Healthcare Clinical Q&A System
Conclusion: Evaluation as Continuous Practice
Introduction
Building a RAG system is only half the battle—evaluating it rigorously is what separates production-grade systems from unreliable demos. Unlike traditional software where tests are deterministic, RAG evaluation must grapple with the probabilistic nature of both retrieval and generation. A system might retrieve the right documents but generate the wrong answer, or generate a plausible-sounding response that isn't grounded in any source. Without systematic evaluation, teams fly blind—unable to detect regressions, compare architectural choices, or prove system reliability to stakeholders.
This article presents an enterprise-grade multi-agent LangGraph evaluation engine that systematically assesses RAG systems across four critical dimensions: retrieval quality, generation faithfulness, answer relevancy, and overall correctness. Each dimension is handled by a specialized agent, with results aggregated into a comprehensive quality report. The system maintains persistent memory of evaluation history, enabling trend analysis and regression detection over time.
The Four Pillars of RAG Evaluation
RAG evaluation decomposes into four independent but complementary dimensions:
Pillar | What It Measures | Why It Matters |
|---|---|---|
Retrieval Quality | Are the right documents retrieved? | Garbage in, garbage out—poor retrieval dooms generation |
Faithfulness | Is the answer grounded in retrieved context? | Prevents hallucinations and fabricated claims |
Answer Relevancy | Does the answer address the actual question? | Ensures the system stays on-topic |
Answer Correctness | Is the answer semantically aligned with ground truth? | Measures factual accuracy against known answers |
Key Metrics Explained
Context Precision@K: Of the top-K retrieved chunks, what fraction are actually relevant to the query? Computed as precision at each relevant chunk position, averaged.
Context Recall@K: Of all the information needed to answer the query (from ground truth), what fraction was retrieved? Requires decomposing ground truth into atomic statements and checking retrieval coverage.
Faithfulness Score: Fraction of claims in the generated answer that can be inferred from the retrieved contexts. Computed by extracting claims from the answer and checking each against context using NLI or LLM-as-judge.
Answer Relevancy: How well the generated answer addresses the original query. Computed by reverse-engineering a query from the answer and measuring similarity to the original.
Answer Correctness: Semantic similarity between the generated answer and ground truth, combining factual accuracy (overlap of claims) and semantic similarity (embedding cosine).
Technology Tags
Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, pgvector, Redis, Pydantic, OpenAI API, Sentence Transformers, RAGAS, DeepEval, scikit-learn, NumPy, NLTK, Docker, TypeScript, TailwindCSS, Prometheus, Pandas
Step-by-Step Implementation
1. Evaluation State Schema with Memory
from typing import List, Dict, Any, TypedDict, Optional
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
from datetime import datetime
class EvaluationItem(TypedDict):
query: str
ground_truth: str
contexts: List[str]
answer: str
class RAGEvaluationState(TypedDict):
messages: List
evaluation_id: str
dataset: List[EvaluationItem]
# Per-item metrics
context_precision_scores: List[float]
context_recall_scores: List[float]
faithfulness_scores: List[float]
answer_relevancy_scores: List[float]
answer_correctness_scores: List[float]
# Aggregate metrics
avg_context_precision: float
avg_context_recall: float
avg_faithfulness: float
avg_answer_relevancy: float
avg_answer_correctness: float
overall_ragas_score: float
# Detailed breakdown
per_item_reports: List[Dict]
failure_analysis: Dict[str, List[int]]
# Memory
evaluation_history: List[Dict]
trend_analysis: Dict[str, List[float]]
2. Context Precision Agent
from langchain_openai import ChatOpenAI
import re
class ContextPrecisionAgent:
"""Measures what fraction of retrieved contexts are actually relevant"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def _judge_relevance(self, query: str, context: str) -> bool:
prompt = f"""Given this query, is the following context useful for answering it?
Respond with only YES or NO.
Query: {query}
Context: {context[:800]}
Answer:"""
resp = self.llm.invoke(prompt)
return "yes" in resp.content.lower()
def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
scores = []
for item in state["dataset"]:
relevance_flags = [
self._judge_relevance(item["query"], ctx)
for ctx in item["contexts"]
]
# Precision@K: average precision across relevant positions
if not any(relevance_flags):
scores.append(0.0)
continue
precision_at_k = []
relevant_count = 0
for k, is_relevant in enumerate(relevance_flags, 1):
if is_relevant:
relevant_count += 1
precision_at_k.append(relevant_count / k)
scores.append(sum(precision_at_k) / len(precision_at_k) if precision_at_k else 0.0)
state["context_precision_scores"] = scores
state["avg_context_precision"] = sum(scores) / max(len(scores), 1)
return state
3. Context Recall Agent
class ContextRecallAgent:
"""Measures what fraction of ground truth information was retrieved"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def _decompose_ground_truth(self, ground_truth: str) -> List[str]:
prompt = f"""Break this answer into individual atomic factual statements.
Return as a JSON array of strings.
Answer: {ground_truth}
JSON array:"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\[[\s\S]*\]', resp.content)
return json.loads(match.group()) if match else [ground_truth]
except Exception:
return [ground_truth]
def _statement_in_contexts(self, statement: str, contexts: List[str]) -> bool:
combined = " ".join(contexts)
prompt = f"""Can this statement be inferred from the given context?
Respond YES or NO.
Statement: {statement}
Context: {combined[:1500]}
Answer:"""
resp = self.llm.invoke(prompt)
return "yes" in resp.content.lower()
def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
scores = []
for item in state["dataset"]:
statements = self._decompose_ground_truth(item["ground_truth"])
if not statements:
scores.append(1.0)
continue
recalled = sum(
1 for stmt in statements
if self._statement_in_contexts(stmt, item["contexts"])
)
scores.append(recalled / len(statements))
state["context_recall_scores"] = scores
state["avg_context_recall"] = sum(scores) / max(len(scores), 1)
return state
4. Faithfulness Agent
class FaithfulnessAgent:
"""Measures what fraction of answer claims are grounded in context"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def _extract_claims(self, answer: str) -> List[str]:
prompt = f"""Extract individual factual claims from this answer.
Return as a JSON array of strings.
Answer: {answer}
JSON array:"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\[[\s\S]*\]', resp.content)
return json.loads(match.group()) if match else [answer]
except Exception:
return [answer]
def _claim_grounded(self, claim: str, contexts: List[str]) -> bool:
combined = " ".join(contexts)
prompt = f"""Is this claim supported by the context?
Respond YES or NO.
Claim: {claim}
Context: {combined[:1500]}
Answer:"""
resp = self.llm.invoke(prompt)
return "yes" in resp.content.lower()
def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
scores = []
for item in state["dataset"]:
claims = self._extract_claims(item["answer"])
if not claims:
scores.append(1.0)
continue
grounded = sum(
1 for claim in claims
if self._claim_grounded(claim, item["contexts"])
)
scores.append(grounded / len(claims))
state["faithfulness_scores"] = scores
state["avg_faithfulness"] = sum(scores) / max(len(scores), 1)
return state
5. Answer Relevancy Agent
from langchain_huggingface import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
class AnswerRelevancyAgent:
"""Measures how well the answer addresses the original query"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
self.embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
def _generate_reverse_query(self, answer: str) -> str:
prompt = f"""Generate a question that this answer would appropriately respond to.
Answer: {answer}
Question:"""
resp = self.llm.invoke(prompt)
return resp.content.strip()
def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
scores = []
for item in state["dataset"]:
reverse_query = self._generate_reverse_query(item["answer"])
orig_emb = np.array(self.embeddings.embed_query(item["query"]))
rev_emb = np.array(self.embeddings.embed_query(reverse_query))
similarity = cosine_similarity(
orig_emb.reshape(1, -1),
rev_emb.reshape(1, -1)
)[0][0]
scores.append(float(max(0, similarity)))
state["answer_relevancy_scores"] = scores
state["avg_answer_relevancy"] = sum(scores) / max(len(scores), 1)
return state
6. Answer Correctness Agent
class AnswerCorrectnessAgent:
"""Measures semantic and factual alignment with ground truth"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
self.embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
def _compute_factual_overlap(self, answer: str, ground_truth: str) -> float:
prompt = f"""Compare these two texts and identify:
1. Claims in the answer that match ground truth (TP)
2. Claims in the answer NOT in ground truth (FP)
3. Claims in ground truth NOT in the answer (FN)
Answer: {answer}
Ground Truth: {ground_truth}
Respond in JSON: {{"tp": N, "fp": N, "fn": N}}"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\{[\s\S]*\}', resp.content)
counts = json.loads(match.group()) if match else {"tp": 0, "fp": 0, "fn": 0}
tp, fp, fn = counts.get("tp", 0), counts.get("fp", 0), counts.get("fn", 0)
precision = tp / (tp + fp) if (tp + fp) > 0 else 0
recall = tp / (tp + fn) if (tp + fn) > 0 else 0
f1 = (2 * precision * recall / (precision + recall)) if (precision + recall) > 0 else 0
return f1
except Exception:
return 0.5
def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
scores = []
for item in state["dataset"]:
# Semantic similarity component
ans_emb = np.array(self.embeddings.embed_query(item["answer"]))
gt_emb = np.array(self.embeddings.embed_query(item["ground_truth"]))
semantic_sim = cosine_similarity(
ans_emb.reshape(1, -1), gt_emb.reshape(1, -1)
)[0][0]
# Factual overlap component
factual_score = self._compute_factual_overlap(item["answer"], item["ground_truth"])
# Weighted combination (70% factual, 30% semantic)
combined = 0.7 * factual_score + 0.3 * float(max(0, semantic_sim))
scores.append(combined)
state["answer_correctness_scores"] = scores
state["avg_answer_correctness"] = sum(scores) / max(len(scores), 1)
return state
7. Aggregator and Report Agent
class AggregatorAgent:
"""Aggregates metrics, identifies failures, generates report"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
def aggregate(self, state: RAGEvaluationState) -> RAGEvaluationState:
# Compute overall RAGAS-style score (weighted average)
weights = {
"context_precision": 0.20,
"context_recall": 0.20,
"faithfulness": 0.25,
"answer_relevancy": 0.15,
"answer_correctness": 0.20
}
state["overall_ragas_score"] = (
weights["context_precision"] * state["avg_context_precision"] +
weights["context_recall"] * state["avg_context_recall"] +
weights["faithfulness"] * state["avg_faithfulness"] +
weights["answer_relevancy"] * state["avg_answer_relevancy"] +
weights["answer_correctness"] * state["avg_answer_correctness"]
)
# Build per-item reports
per_item = []
for i, item in enumerate(state["dataset"]):
per_item.append({
"index": i,
"query": item["query"][:100],
"context_precision": state["context_precision_scores"][i],
"context_recall": state["context_recall_scores"][i],
"faithfulness": state["faithfulness_scores"][i],
"answer_relevancy": state["answer_relevancy_scores"][i],
"answer_correctness": state["answer_correctness_scores"][i]
})
state["per_item_reports"] = per_item
# Failure analysis: identify items below threshold (0.5)
threshold = 0.5
failure_analysis = {
"low_precision": [i for i, s in enumerate(state["context_precision_scores"]) if s < threshold],
"low_recall": [i for i, s in enumerate(state["context_recall_scores"]) if s < threshold],
"unfaithful": [i for i, s in enumerate(state["faithfulness_scores"]) if s < threshold],
"irrelevant": [i for i, s in enumerate(state["answer_relevancy_scores"]) if s < threshold],
"incorrect": [i for i, s in enumerate(state["answer_correctness_scores"]) if s < threshold]
}
state["failure_analysis"] = failure_analysis
return state
def generate_recommendations(self, state: RAGEvaluationState) -> str:
failures = state["failure_analysis"]
recommendations = []
if len(failures["low_precision"]) > len(state["dataset"]) * 0.3:
recommendations.append("High noise in retrieval: add reranker or tighten chunking")
if len(failures["low_recall"]) > len(state["dataset"]) * 0.3:
recommendations.append("Low recall: increase top-k, use hybrid search, or improve embeddings")
if len(failures["unfaithful"]) > len(state["dataset"]) * 0.2:
recommendations.append("Hallucination detected: strengthen grounding prompts or add verification layer")
if len(failures["irrelevant"]) > len(state["dataset"]) * 0.2:
recommendations.append("Answers off-topic: improve prompt constraints or query understanding")
if len(failures["incorrect"]) > len(state["dataset"]) * 0.3:
recommendations.append("Factual errors: review retrieval quality and generation prompts")
if not recommendations:
recommendations.append("System performing well across all dimensions—maintain current configuration")
return "\n".join([f"- {r}" for r in recommendations])
8. LangGraph Workflow with Parallel Execution and Memory
class EvaluationMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def save_evaluation(self, state: RAGEvaluationState):
record = {
"evaluation_id": state["evaluation_id"],
"timestamp": datetime.now().isoformat(),
"overall_score": state["overall_ragas_score"],
"context_precision": state["avg_context_precision"],
"context_recall": state["avg_context_recall"],
"faithfulness": state["avg_faithfulness"],
"answer_relevancy": state["avg_answer_relevancy"],
"answer_correctness": state["avg_answer_correctness"],
"dataset_size": len(state["dataset"])
}
key = "evaluation_history"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-50:]))
def load_history(self) -> List[Dict]:
return json.loads(self.redis.get("evaluation_history") or "[]")
def build_evaluation_graph():
workflow = StateGraph(RAGEvaluationState)
precision_agent = ContextPrecisionAgent()
recall_agent = ContextRecallAgent()
faithfulness_agent = FaithfulnessAgent()
relevancy_agent = AnswerRelevancyAgent()
correctness_agent = AnswerCorrectnessAgent()
aggregator = AggregatorAgent()
workflow.add_node("context_precision", precision_agent.evaluate)
workflow.add_node("context_recall", recall_agent.evaluate)
workflow.add_node("faithfulness", faithfulness_agent.evaluate)
workflow.add_node("answer_relevancy", relevancy_agent.evaluate)
workflow.add_node("answer_correctness", correctness_agent.evaluate)
workflow.add_node("aggregate", aggregator.aggregate)
workflow.set_entry_point("context_precision")
# Fan-out: all 5 evaluation agents run after initial setup
workflow.add_edge("context_precision", "context_recall")
workflow.add_edge("context_precision", "faithfulness")
workflow.add_edge("context_precision", "answer_relevancy")
workflow.add_edge("context_precision", "answer_correctness")
# Fan-in: all converge to aggregation
workflow.add_edge("context_recall", "aggregate")
workflow.add_edge("faithfulness", "aggregate")
workflow.add_edge("answer_relevancy", "aggregate")
workflow.add_edge("answer_correctness", "aggregate")
workflow.add_edge("aggregate", END)
return workflow.compile()
9. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="RAG Evaluation Engine API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_evaluation_graph()
memory = EvaluationMemory(redis.Redis())
class EvaluationDataset(BaseModel):
evaluation_id: str
dataset: List[EvaluationItem]
@app.post("/evaluate")
async def evaluate_rag(req: EvaluationDataset):
history = memory.load_history()
initial_state = RAGEvaluationState(
messages=[HumanMessage(content=f"Evaluate dataset {req.evaluation_id}")],
evaluation_id=req.evaluation_id,
dataset=req.dataset,
context_precision_scores=[],
context_recall_scores=[],
faithfulness_scores=[],
answer_relevancy_scores=[],
answer_correctness_scores=[],
avg_context_precision=0.0,
avg_context_recall=0.0,
avg_faithfulness=0.0,
avg_answer_relevancy=0.0,
avg_answer_correctness=0.0,
overall_ragas_score=0.0,
per_item_reports=[],
failure_analysis={},
evaluation_history=history,
trend_analysis={}
)
result = graph.invoke(initial_state)
memory.save_evaluation(result)
recommendations = AggregatorAgent().generate_recommendations(result)
return {
"evaluation_id": result["evaluation_id"],
"overall_ragas_score": round(result["overall_ragas_score"], 3),
"metrics": {
"context_precision": round(result["avg_context_precision"], 3),
"context_recall": round(result["avg_context_recall"], 3),
"faithfulness": round(result["avg_faithfulness"], 3),
"answer_relevancy": round(result["avg_answer_relevancy"], 3),
"answer_correctness": round(result["avg_answer_correctness"], 3)
},
"per_item_reports": result["per_item_reports"],
"failure_analysis": result["failure_analysis"],
"recommendations": recommendations,
"historical_trend": history[-10:]
}
10. Frontend: Evaluation Dashboard
// components/RAGEvaluationDashboard.tsx
import React, { useState } from 'react';
interface Metrics {
context_precision: number;
context_recall: number;
faithfulness: number;
answer_relevancy: number;
answer_correctness: number;
}
interface EvaluationResult {
evaluation_id: string;
overall_ragas_score: number;
metrics: Metrics;
per_item_reports: any[];
failure_analysis: Record<string, number[]>;
recommendations: string;
historical_trend: any[];
}
const MetricCard: React.FC<{ label: string; score: number; color: string }> = ({ label, score, color }) => (
<div className="bg-white p-4 rounded-lg shadow">
<div className="text-sm text-gray-600">{label}</div>
<div className={`text-3xl font-bold mt-2 ${color}`}>
{(score * 100).toFixed(1)}%
</div>
<div className="bg-gray-200 rounded-full h-2 mt-2">
<div className={`${color.replace('text-', 'bg-')} h-2 rounded-full transition-all`}
style={{ width: `${score * 100}%` }} />
</div>
</div>
);
export const RAGEvaluationDashboard: React.FC = () => {
const [result, setResult] = useState<EvaluationResult | null>(null);
const [loading, setLoading] = useState(false);
const sampleDataset = [
{
query: "What are the symptoms of type 2 diabetes?",
ground_truth: "Common symptoms include increased thirst, frequent urination, fatigue, blurred vision, and slow-healing wounds.",
contexts: [
"Type 2 diabetes symptoms include polydipsia (excessive thirst), polyuria (frequent urination), and unexplained fatigue.",
"Patients may also experience blurred vision and wounds that heal slowly due to impaired circulation."
],
answer: "Type 2 diabetes symptoms include increased thirst, frequent urination, fatigue, blurred vision, and slow-healing wounds."
},
{
query: "What is the first-line treatment for hypertension?",
ground_truth: "First-line treatments include ACE inhibitors, ARBs, calcium channel blockers, or thiazide diuretics.",
contexts: [
"Hypertension management begins with lifestyle modifications including diet and exercise.",
"Pharmacological first-line options include ACE inhibitors and calcium channel blockers."
],
answer: "The first-line treatment for hypertension includes ACE inhibitors, ARBs, calcium channel blockers, or thiazide diuretics, along with lifestyle modifications."
},
{
query: "What causes myocardial infarction?",
ground_truth: "Myocardial infarction is caused by blockage of coronary arteries, typically due to atherosclerotic plaque rupture and thrombosis.",
contexts: [
"Cardiovascular disease is the leading cause of death worldwide.",
"Risk factors for heart disease include smoking, obesity, and sedentary lifestyle."
],
answer: "Myocardial infarction is caused by poor diet and lack of exercise."
}
];
const runEvaluation = async () => {
setLoading(true);
try {
const response = await fetch('http://localhost:8000/evaluate', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
evaluation_id: `eval-${Date.now()}`,
dataset: sampleDataset
})
});
setResult(await response.json());
} finally {
setLoading(false);
}
};
const getScoreColor = (score: number) => {
if (score >= 0.8) return 'text-green-600';
if (score >= 0.6) return 'text-yellow-600';
return 'text-red-600';
};
return (
<div className="p-6 max-w-7xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-2">📊 RAG Evaluation Engine</h1>
<p className="text-gray-600 mb-6">Comprehensive quality assessment across 5 dimensions</p>
<div className="bg-white p-4 rounded-lg shadow mb-6">
<div className="flex items-center justify-between">
<div>
<h2 className="font-bold">Clinical Q&A Evaluation Dataset</h2>
<p className="text-sm text-gray-600">{sampleDataset.length} test cases covering medical knowledge retrieval</p>
</div>
<button onClick={runEvaluation} disabled={loading}
className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
{loading ? 'Evaluating...' : 'Run Full Evaluation'}
</button>
</div>
</div>
{result && (
<div className="space-y-6">
<div className={`p-6 rounded-lg shadow ${
result.overall_ragas_score >= 0.8 ? 'bg-green-50 border border-green-300' :
result.overall_ragas_score >= 0.6 ? 'bg-yellow-50 border border-yellow-300' :
'bg-red-50 border border-red-300'
}`}>
<div className="text-sm opacity-75">OVERALL RAGAS SCORE</div>
<div className={`text-5xl font-bold mt-2 ${getScoreColor(result.overall_ragas_score)}`}>
{(result.overall_ragas_score * 100).toFixed(1)}%
</div>
</div>
<div className="grid grid-cols-5 gap-4">
<MetricCard label="Context Precision" score={result.metrics.context_precision}
color={getScoreColor(result.metrics.context_precision)} />
<MetricCard label="Context Recall" score={result.metrics.context_recall}
color={getScoreColor(result.metrics.context_recall)} />
<MetricCard label="Faithfulness" score={result.metrics.faithfulness}
color={getScoreColor(result.metrics.faithfulness)} />
<MetricCard label="Answer Relevancy" score={result.metrics.answer_relevancy}
color={getScoreColor(result.metrics.answer_relevancy)} />
<MetricCard label="Answer Correctness" score={result.metrics.answer_correctness}
color={getScoreColor(result.metrics.answer_correctness)} />
</div>
<div className="grid grid-cols-2 gap-4">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">🔍 Failure Analysis</h3>
<div className="space-y-2">
{Object.entries(result.failure_analysis).map(([key, indices]) => (
<div key={key} className="flex justify-between items-center border-b pb-2">
<span className="text-sm capitalize">
{key.replace(/_/g, ' ')}
</span>
<span className={`text-sm font-mono ${
(indices as number[]).length > 0 ? 'text-red-600' : 'text-green-600'
}`}>
{(indices as number[]).length} failures
</span>
</div>
))}
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">💡 Recommendations</h3>
<div className="text-sm text-gray-700 whitespace-pre-line">
{result.recommendations}
</div>
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">📋 Per-Item Breakdown</h3>
<div className="overflow-x-auto">
<table className="w-full text-sm">
<thead>
<tr className="border-b">
<th className="text-left py-2">#</th>
<th className="text-left py-2">Query</th>
<th className="text-center py-2">Precision</th>
<th className="text-center py-2">Recall</th>
<th className="text-center py-2">Faithful</th>
<th className="text-center py-2">Relevant</th>
<th className="text-center py-2">Correct</th>
</tr>
</thead>
<tbody>
{result.per_item_reports.map((item, i) => (
<tr key={i} className="border-b hover:bg-gray-50">
<td className="py-2">{i + 1}</td>
<td className="py-2 text-xs">{item.query}...</td>
<td className={`text-center ${getScoreColor(item.context_precision)}`}>
{(item.context_precision * 100).toFixed(0)}%
</td>
<td className={`text-center ${getScoreColor(item.context_recall)}`}>
{(item.context_recall * 100).toFixed(0)}%
</td>
<td className={`text-center ${getScoreColor(item.faithfulness)}`}>
{(item.faithfulness * 100).toFixed(0)}%
</td>
<td className={`text-center ${getScoreColor(item.answer_relevancy)}`}>
{(item.answer_relevancy * 100).toFixed(0)}%
</td>
<td className={`text-center ${getScoreColor(item.answer_correctness)}`}>
{(item.answer_correctness * 100).toFixed(0)}%
</td>
</tr>
))}
</tbody>
</table>
</div>
</div>
{result.historical_trend.length > 1 && (
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">📈 Historical Trend</h3>
<div className="space-y-2">
{result.historical_trend.map((h, i) => (
<div key={i} className="flex items-center gap-3 text-sm">
<span className="text-xs text-gray-500 w-32">
{new Date(h.timestamp).toLocaleDateString()}
</span>
<div className="flex-1 bg-gray-200 rounded-full h-2">
<div className="bg-blue-500 h-2 rounded-full"
style={{ width: `${h.overall_score * 100}%` }} />
</div>
<span className="font-mono w-16 text-right">
{(h.overall_score * 100).toFixed(1)}%
</span>
</div>
))}
</div>
</div>
)}
</div>
)}
</div>
);
};
Real-Time Use Case: Healthcare Clinical Q&A System
A hospital deploys a RAG system to answer clinician questions from medical guidelines, drug databases, and research papers. Before production rollout, the team runs the evaluation engine on a 200-case test set.
Evaluation results reveal:
Context Precision: 0.82 — retrieval is generally relevant
Context Recall: 0.58 — significant information gaps; many queries need docs that aren't retrieved
Faithfulness: 0.91 — strong grounding, low hallucination
Answer Relevancy: 0.87 — answers stay on topic
Answer Correctness: 0.64 — factual gaps due to low recall
Overall RAGAS: 0.75
Failure analysis shows 42 cases with low recall, concentrated in queries about drug interactions and rare conditions. The recommendation engine suggests: "Low recall: increase top-k, use hybrid search, or improve embeddings."
Action taken: The team switches from pure vector search to hybrid search (BM25 + dense), increases top-k from 5 to 10, and adds a domain-adapted medical embedding model.
Re-evaluation after changes:
Context Recall: 0.58 → 0.81 (+39% improvement)
Answer Correctness: 0.64 → 0.79 (+23% improvement)
Overall RAGAS: 0.75 → 0.86
The Redis-backed memory tracks this improvement over time, allowing the team to demonstrate ROI to hospital leadership and detect any future regressions within days rather than months.
Conclusion
RAG evaluation is not a one-time checkpoint it's a continuous practice that must evolve alongside your system. By decomposing evaluation into five independent dimensions (context precision, context recall, faithfulness, answer relevancy, answer correctness) and orchestrating them through a multi-agent LangGraph framework, enterprises gain a systematic, reproducible approach to quality assessment. Each dimension is measured by a specialized agent using LLM-as-judge techniques, with results aggregated into an overall RAGAS-style score and actionable recommendations. The persistent memory layer enables trend analysis, regression detection, and evidence-based architectural decisions. In regulated industries like healthcare, finance, and legal, this evaluation framework isn't optional it's the foundation of responsible AI deployment, providing the auditability and accountability that stakeholders demand. Without rigorous evaluation, you're not building a production system; you're building a liability.

Join the conversation! Your thoughts help the community grow.