Table of Contents

  1. Introduction: The Evaluation Imperative in Enterprise RAG

  2. The Four Pillars of RAG Evaluation

  3. Key Metrics: From Retrieval to Generation

  4. Evaluation Frameworks: RAGAS, DeepEval, and LLM-as-Judge

  5. Solution Architecture: The "RAG Evaluation Engine" Multi-Agent System

  6. Technology Stack Overview

  7. Step-by-Step Implementation: Backend Development

    • Defining the Evaluation State Schema with Memory

    • Building the Context Precision Agent

    • Implementing the Context Recall Agent

    • Creating the Faithfulness Agent

    • Designing the Answer Relevancy Agent

    • Building the Answer Correctness Agent

    • Constructing the Aggregator and Report Agent

    • Building the LangGraph Workflow with Parallel Execution

  8. Frontend Implementation: Evaluation Dashboard

  9. Real-Time Use Case: Healthcare Clinical Q&A System

  10. Conclusion: Evaluation as Continuous Practice

Introduction

Building a RAG system is only half the battle—evaluating it rigorously is what separates production-grade systems from unreliable demos. Unlike traditional software where tests are deterministic, RAG evaluation must grapple with the probabilistic nature of both retrieval and generation. A system might retrieve the right documents but generate the wrong answer, or generate a plausible-sounding response that isn't grounded in any source. Without systematic evaluation, teams fly blind—unable to detect regressions, compare architectural choices, or prove system reliability to stakeholders.

This article presents an enterprise-grade multi-agent LangGraph evaluation engine that systematically assesses RAG systems across four critical dimensions: retrieval quality, generation faithfulness, answer relevancy, and overall correctness. Each dimension is handled by a specialized agent, with results aggregated into a comprehensive quality report. The system maintains persistent memory of evaluation history, enabling trend analysis and regression detection over time.

The Four Pillars of RAG Evaluation

RAG evaluation decomposes into four independent but complementary dimensions:

Pillar

What It Measures

Why It Matters

Retrieval Quality

Are the right documents retrieved?

Garbage in, garbage out—poor retrieval dooms generation

Faithfulness

Is the answer grounded in retrieved context?

Prevents hallucinations and fabricated claims

Answer Relevancy

Does the answer address the actual question?

Ensures the system stays on-topic

Answer Correctness

Is the answer semantically aligned with ground truth?

Measures factual accuracy against known answers

Key Metrics Explained

Context Precision@K: Of the top-K retrieved chunks, what fraction are actually relevant to the query? Computed as precision at each relevant chunk position, averaged.

Context Recall@K: Of all the information needed to answer the query (from ground truth), what fraction was retrieved? Requires decomposing ground truth into atomic statements and checking retrieval coverage.

Faithfulness Score: Fraction of claims in the generated answer that can be inferred from the retrieved contexts. Computed by extracting claims from the answer and checking each against context using NLI or LLM-as-judge.

Answer Relevancy: How well the generated answer addresses the original query. Computed by reverse-engineering a query from the answer and measuring similarity to the original.

Answer Correctness: Semantic similarity between the generated answer and ground truth, combining factual accuracy (overlap of claims) and semantic similarity (embedding cosine).

Technology Tags

Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, pgvector, Redis, Pydantic, OpenAI API, Sentence Transformers, RAGAS, DeepEval, scikit-learn, NumPy, NLTK, Docker, TypeScript, TailwindCSS, Prometheus, Pandas

Step-by-Step Implementation

1. Evaluation State Schema with Memory

from typing import List, Dict, Any, TypedDict, Optional
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
from datetime import datetime

class EvaluationItem(TypedDict):
    query: str
    ground_truth: str
    contexts: List[str]
    answer: str

class RAGEvaluationState(TypedDict):
    messages: List
    evaluation_id: str
    dataset: List[EvaluationItem]
    # Per-item metrics
    context_precision_scores: List[float]
    context_recall_scores: List[float]
    faithfulness_scores: List[float]
    answer_relevancy_scores: List[float]
    answer_correctness_scores: List[float]
    # Aggregate metrics
    avg_context_precision: float
    avg_context_recall: float
    avg_faithfulness: float
    avg_answer_relevancy: float
    avg_answer_correctness: float
    overall_ragas_score: float
    # Detailed breakdown
    per_item_reports: List[Dict]
    failure_analysis: Dict[str, List[int]]
    # Memory
    evaluation_history: List[Dict]
    trend_analysis: Dict[str, List[float]]

2. Context Precision Agent

from langchain_openai import ChatOpenAI
import re

class ContextPrecisionAgent:
    """Measures what fraction of retrieved contexts are actually relevant"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def _judge_relevance(self, query: str, context: str) -> bool:
        prompt = f"""Given this query, is the following context useful for answering it?
Respond with only YES or NO.

Query: {query}
Context: {context[:800]}

Answer:"""
        resp = self.llm.invoke(prompt)
        return "yes" in resp.content.lower()
    
    def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        scores = []
        for item in state["dataset"]:
            relevance_flags = [
                self._judge_relevance(item["query"], ctx)
                for ctx in item["contexts"]
            ]
            
            # Precision@K: average precision across relevant positions
            if not any(relevance_flags):
                scores.append(0.0)
                continue
            
            precision_at_k = []
            relevant_count = 0
            for k, is_relevant in enumerate(relevance_flags, 1):
                if is_relevant:
                    relevant_count += 1
                    precision_at_k.append(relevant_count / k)
            
            scores.append(sum(precision_at_k) / len(precision_at_k) if precision_at_k else 0.0)
        
        state["context_precision_scores"] = scores
        state["avg_context_precision"] = sum(scores) / max(len(scores), 1)
        return state

3. Context Recall Agent

class ContextRecallAgent:
    """Measures what fraction of ground truth information was retrieved"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def _decompose_ground_truth(self, ground_truth: str) -> List[str]:
        prompt = f"""Break this answer into individual atomic factual statements.
Return as a JSON array of strings.

Answer: {ground_truth}

JSON array:"""
        resp = self.llm.invoke(prompt)
        try:
            match = re.search(r'\[[\s\S]*\]', resp.content)
            return json.loads(match.group()) if match else [ground_truth]
        except Exception:
            return [ground_truth]
    
    def _statement_in_contexts(self, statement: str, contexts: List[str]) -> bool:
        combined = " ".join(contexts)
        prompt = f"""Can this statement be inferred from the given context?
Respond YES or NO.

Statement: {statement}
Context: {combined[:1500]}

Answer:"""
        resp = self.llm.invoke(prompt)
        return "yes" in resp.content.lower()
    
    def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        scores = []
        for item in state["dataset"]:
            statements = self._decompose_ground_truth(item["ground_truth"])
            if not statements:
                scores.append(1.0)
                continue
            
            recalled = sum(
                1 for stmt in statements
                if self._statement_in_contexts(stmt, item["contexts"])
            )
            scores.append(recalled / len(statements))
        
        state["context_recall_scores"] = scores
        state["avg_context_recall"] = sum(scores) / max(len(scores), 1)
        return state

4. Faithfulness Agent

class FaithfulnessAgent:
    """Measures what fraction of answer claims are grounded in context"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
    
    def _extract_claims(self, answer: str) -> List[str]:
        prompt = f"""Extract individual factual claims from this answer.
Return as a JSON array of strings.

Answer: {answer}

JSON array:"""
        resp = self.llm.invoke(prompt)
        try:
            match = re.search(r'\[[\s\S]*\]', resp.content)
            return json.loads(match.group()) if match else [answer]
        except Exception:
            return [answer]
    
    def _claim_grounded(self, claim: str, contexts: List[str]) -> bool:
        combined = " ".join(contexts)
        prompt = f"""Is this claim supported by the context?
Respond YES or NO.

Claim: {claim}
Context: {combined[:1500]}

Answer:"""
        resp = self.llm.invoke(prompt)
        return "yes" in resp.content.lower()
    
    def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        scores = []
        for item in state["dataset"]:
            claims = self._extract_claims(item["answer"])
            if not claims:
                scores.append(1.0)
                continue
            
            grounded = sum(
                1 for claim in claims
                if self._claim_grounded(claim, item["contexts"])
            )
            scores.append(grounded / len(claims))
        
        state["faithfulness_scores"] = scores
        state["avg_faithfulness"] = sum(scores) / max(len(scores), 1)
        return state

5. Answer Relevancy Agent

from langchain_huggingface import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

class AnswerRelevancyAgent:
    """Measures how well the answer addresses the original query"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
        self.embeddings = HuggingFaceEmbeddings(
            model_name="sentence-transformers/all-MiniLM-L6-v2"
        )
    
    def _generate_reverse_query(self, answer: str) -> str:
        prompt = f"""Generate a question that this answer would appropriately respond to.
Answer: {answer}
Question:"""
        resp = self.llm.invoke(prompt)
        return resp.content.strip()
    
    def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        scores = []
        for item in state["dataset"]:
            reverse_query = self._generate_reverse_query(item["answer"])
            
            orig_emb = np.array(self.embeddings.embed_query(item["query"]))
            rev_emb = np.array(self.embeddings.embed_query(reverse_query))
            
            similarity = cosine_similarity(
                orig_emb.reshape(1, -1),
                rev_emb.reshape(1, -1)
            )[0][0]
            scores.append(float(max(0, similarity)))
        
        state["answer_relevancy_scores"] = scores
        state["avg_answer_relevancy"] = sum(scores) / max(len(scores), 1)
        return state

6. Answer Correctness Agent

class AnswerCorrectnessAgent:
    """Measures semantic and factual alignment with ground truth"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
        self.embeddings = HuggingFaceEmbeddings(
            model_name="sentence-transformers/all-MiniLM-L6-v2"
        )
    
    def _compute_factual_overlap(self, answer: str, ground_truth: str) -> float:
        prompt = f"""Compare these two texts and identify:
1. Claims in the answer that match ground truth (TP)
2. Claims in the answer NOT in ground truth (FP)
3. Claims in ground truth NOT in the answer (FN)

Answer: {answer}
Ground Truth: {ground_truth}

Respond in JSON: {{"tp": N, "fp": N, "fn": N}}"""
        resp = self.llm.invoke(prompt)
        try:
            match = re.search(r'\{[\s\S]*\}', resp.content)
            counts = json.loads(match.group()) if match else {"tp": 0, "fp": 0, "fn": 0}
            tp, fp, fn = counts.get("tp", 0), counts.get("fp", 0), counts.get("fn", 0)
            precision = tp / (tp + fp) if (tp + fp) > 0 else 0
            recall = tp / (tp + fn) if (tp + fn) > 0 else 0
            f1 = (2 * precision * recall / (precision + recall)) if (precision + recall) > 0 else 0
            return f1
        except Exception:
            return 0.5
    
    def evaluate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        scores = []
        for item in state["dataset"]:
            # Semantic similarity component
            ans_emb = np.array(self.embeddings.embed_query(item["answer"]))
            gt_emb = np.array(self.embeddings.embed_query(item["ground_truth"]))
            semantic_sim = cosine_similarity(
                ans_emb.reshape(1, -1), gt_emb.reshape(1, -1)
            )[0][0]
            
            # Factual overlap component
            factual_score = self._compute_factual_overlap(item["answer"], item["ground_truth"])
            
            # Weighted combination (70% factual, 30% semantic)
            combined = 0.7 * factual_score + 0.3 * float(max(0, semantic_sim))
            scores.append(combined)
        
        state["answer_correctness_scores"] = scores
        state["avg_answer_correctness"] = sum(scores) / max(len(scores), 1)
        return state

7. Aggregator and Report Agent

class AggregatorAgent:
    """Aggregates metrics, identifies failures, generates report"""
    
    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
    
    def aggregate(self, state: RAGEvaluationState) -> RAGEvaluationState:
        # Compute overall RAGAS-style score (weighted average)
        weights = {
            "context_precision": 0.20,
            "context_recall": 0.20,
            "faithfulness": 0.25,
            "answer_relevancy": 0.15,
            "answer_correctness": 0.20
        }
        
        state["overall_ragas_score"] = (
            weights["context_precision"] * state["avg_context_precision"] +
            weights["context_recall"] * state["avg_context_recall"] +
            weights["faithfulness"] * state["avg_faithfulness"] +
            weights["answer_relevancy"] * state["avg_answer_relevancy"] +
            weights["answer_correctness"] * state["avg_answer_correctness"]
        )
        
        # Build per-item reports
        per_item = []
        for i, item in enumerate(state["dataset"]):
            per_item.append({
                "index": i,
                "query": item["query"][:100],
                "context_precision": state["context_precision_scores"][i],
                "context_recall": state["context_recall_scores"][i],
                "faithfulness": state["faithfulness_scores"][i],
                "answer_relevancy": state["answer_relevancy_scores"][i],
                "answer_correctness": state["answer_correctness_scores"][i]
            })
        state["per_item_reports"] = per_item
        
        # Failure analysis: identify items below threshold (0.5)
        threshold = 0.5
        failure_analysis = {
            "low_precision": [i for i, s in enumerate(state["context_precision_scores"]) if s < threshold],
            "low_recall": [i for i, s in enumerate(state["context_recall_scores"]) if s < threshold],
            "unfaithful": [i for i, s in enumerate(state["faithfulness_scores"]) if s < threshold],
            "irrelevant": [i for i, s in enumerate(state["answer_relevancy_scores"]) if s < threshold],
            "incorrect": [i for i, s in enumerate(state["answer_correctness_scores"]) if s < threshold]
        }
        state["failure_analysis"] = failure_analysis
        
        return state
    
    def generate_recommendations(self, state: RAGEvaluationState) -> str:
        failures = state["failure_analysis"]
        recommendations = []
        
        if len(failures["low_precision"]) > len(state["dataset"]) * 0.3:
            recommendations.append("High noise in retrieval: add reranker or tighten chunking")
        if len(failures["low_recall"]) > len(state["dataset"]) * 0.3:
            recommendations.append("Low recall: increase top-k, use hybrid search, or improve embeddings")
        if len(failures["unfaithful"]) > len(state["dataset"]) * 0.2:
            recommendations.append("Hallucination detected: strengthen grounding prompts or add verification layer")
        if len(failures["irrelevant"]) > len(state["dataset"]) * 0.2:
            recommendations.append("Answers off-topic: improve prompt constraints or query understanding")
        if len(failures["incorrect"]) > len(state["dataset"]) * 0.3:
            recommendations.append("Factual errors: review retrieval quality and generation prompts")
        
        if not recommendations:
            recommendations.append("System performing well across all dimensions—maintain current configuration")
        
        return "\n".join([f"- {r}" for r in recommendations])

8. LangGraph Workflow with Parallel Execution and Memory

class EvaluationMemory:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client
    
    def save_evaluation(self, state: RAGEvaluationState):
        record = {
            "evaluation_id": state["evaluation_id"],
            "timestamp": datetime.now().isoformat(),
            "overall_score": state["overall_ragas_score"],
            "context_precision": state["avg_context_precision"],
            "context_recall": state["avg_context_recall"],
            "faithfulness": state["avg_faithfulness"],
            "answer_relevancy": state["avg_answer_relevancy"],
            "answer_correctness": state["avg_answer_correctness"],
            "dataset_size": len(state["dataset"])
        }
        key = "evaluation_history"
        history = json.loads(self.redis.get(key) or "[]")
        history.append(record)
        self.redis.set(key, json.dumps(history[-50:]))
    
    def load_history(self) -> List[Dict]:
        return json.loads(self.redis.get("evaluation_history") or "[]")

def build_evaluation_graph():
    workflow = StateGraph(RAGEvaluationState)
    
    precision_agent = ContextPrecisionAgent()
    recall_agent = ContextRecallAgent()
    faithfulness_agent = FaithfulnessAgent()
    relevancy_agent = AnswerRelevancyAgent()
    correctness_agent = AnswerCorrectnessAgent()
    aggregator = AggregatorAgent()
    
    workflow.add_node("context_precision", precision_agent.evaluate)
    workflow.add_node("context_recall", recall_agent.evaluate)
    workflow.add_node("faithfulness", faithfulness_agent.evaluate)
    workflow.add_node("answer_relevancy", relevancy_agent.evaluate)
    workflow.add_node("answer_correctness", correctness_agent.evaluate)
    workflow.add_node("aggregate", aggregator.aggregate)
    
    workflow.set_entry_point("context_precision")
    
    # Fan-out: all 5 evaluation agents run after initial setup
    workflow.add_edge("context_precision", "context_recall")
    workflow.add_edge("context_precision", "faithfulness")
    workflow.add_edge("context_precision", "answer_relevancy")
    workflow.add_edge("context_precision", "answer_correctness")
    
    # Fan-in: all converge to aggregation
    workflow.add_edge("context_recall", "aggregate")
    workflow.add_edge("faithfulness", "aggregate")
    workflow.add_edge("answer_relevancy", "aggregate")
    workflow.add_edge("answer_correctness", "aggregate")
    workflow.add_edge("aggregate", END)
    
    return workflow.compile()

9. FastAPI Backend

from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel

app = FastAPI(title="RAG Evaluation Engine API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])

graph = build_evaluation_graph()
memory = EvaluationMemory(redis.Redis())

class EvaluationDataset(BaseModel):
    evaluation_id: str
    dataset: List[EvaluationItem]

@app.post("/evaluate")
async def evaluate_rag(req: EvaluationDataset):
    history = memory.load_history()
    
    initial_state = RAGEvaluationState(
        messages=[HumanMessage(content=f"Evaluate dataset {req.evaluation_id}")],
        evaluation_id=req.evaluation_id,
        dataset=req.dataset,
        context_precision_scores=[],
        context_recall_scores=[],
        faithfulness_scores=[],
        answer_relevancy_scores=[],
        answer_correctness_scores=[],
        avg_context_precision=0.0,
        avg_context_recall=0.0,
        avg_faithfulness=0.0,
        avg_answer_relevancy=0.0,
        avg_answer_correctness=0.0,
        overall_ragas_score=0.0,
        per_item_reports=[],
        failure_analysis={},
        evaluation_history=history,
        trend_analysis={}
    )
    
    result = graph.invoke(initial_state)
    memory.save_evaluation(result)
    
    recommendations = AggregatorAgent().generate_recommendations(result)
    
    return {
        "evaluation_id": result["evaluation_id"],
        "overall_ragas_score": round(result["overall_ragas_score"], 3),
        "metrics": {
            "context_precision": round(result["avg_context_precision"], 3),
            "context_recall": round(result["avg_context_recall"], 3),
            "faithfulness": round(result["avg_faithfulness"], 3),
            "answer_relevancy": round(result["avg_answer_relevancy"], 3),
            "answer_correctness": round(result["avg_answer_correctness"], 3)
        },
        "per_item_reports": result["per_item_reports"],
        "failure_analysis": result["failure_analysis"],
        "recommendations": recommendations,
        "historical_trend": history[-10:]
    }

10. Frontend: Evaluation Dashboard

// components/RAGEvaluationDashboard.tsx
import React, { useState } from 'react';

interface Metrics {
  context_precision: number;
  context_recall: number;
  faithfulness: number;
  answer_relevancy: number;
  answer_correctness: number;
}

interface EvaluationResult {
  evaluation_id: string;
  overall_ragas_score: number;
  metrics: Metrics;
  per_item_reports: any[];
  failure_analysis: Record<string, number[]>;
  recommendations: string;
  historical_trend: any[];
}

const MetricCard: React.FC<{ label: string; score: number; color: string }> = ({ label, score, color }) => (
  <div className="bg-white p-4 rounded-lg shadow">
    <div className="text-sm text-gray-600">{label}</div>
    <div className={`text-3xl font-bold mt-2 ${color}`}>
      {(score * 100).toFixed(1)}%
    </div>
    <div className="bg-gray-200 rounded-full h-2 mt-2">
      <div className={`${color.replace('text-', 'bg-')} h-2 rounded-full transition-all`}
        style={{ width: `${score * 100}%` }} />
    </div>
  </div>
);

export const RAGEvaluationDashboard: React.FC = () => {
  const [result, setResult] = useState<EvaluationResult | null>(null);
  const [loading, setLoading] = useState(false);

  const sampleDataset = [
    {
      query: "What are the symptoms of type 2 diabetes?",
      ground_truth: "Common symptoms include increased thirst, frequent urination, fatigue, blurred vision, and slow-healing wounds.",
      contexts: [
        "Type 2 diabetes symptoms include polydipsia (excessive thirst), polyuria (frequent urination), and unexplained fatigue.",
        "Patients may also experience blurred vision and wounds that heal slowly due to impaired circulation."
      ],
      answer: "Type 2 diabetes symptoms include increased thirst, frequent urination, fatigue, blurred vision, and slow-healing wounds."
    },
    {
      query: "What is the first-line treatment for hypertension?",
      ground_truth: "First-line treatments include ACE inhibitors, ARBs, calcium channel blockers, or thiazide diuretics.",
      contexts: [
        "Hypertension management begins with lifestyle modifications including diet and exercise.",
        "Pharmacological first-line options include ACE inhibitors and calcium channel blockers."
      ],
      answer: "The first-line treatment for hypertension includes ACE inhibitors, ARBs, calcium channel blockers, or thiazide diuretics, along with lifestyle modifications."
    },
    {
      query: "What causes myocardial infarction?",
      ground_truth: "Myocardial infarction is caused by blockage of coronary arteries, typically due to atherosclerotic plaque rupture and thrombosis.",
      contexts: [
        "Cardiovascular disease is the leading cause of death worldwide.",
        "Risk factors for heart disease include smoking, obesity, and sedentary lifestyle."
      ],
      answer: "Myocardial infarction is caused by poor diet and lack of exercise."
    }
  ];

  const runEvaluation = async () => {
    setLoading(true);
    try {
      const response = await fetch('http://localhost:8000/evaluate', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({
          evaluation_id: `eval-${Date.now()}`,
          dataset: sampleDataset
        })
      });
      setResult(await response.json());
    } finally {
      setLoading(false);
    }
  };

  const getScoreColor = (score: number) => {
    if (score >= 0.8) return 'text-green-600';
    if (score >= 0.6) return 'text-yellow-600';
    return 'text-red-600';
  };

  return (
    <div className="p-6 max-w-7xl mx-auto bg-gray-50 min-h-screen">
      <h1 className="text-3xl font-bold mb-2">📊 RAG Evaluation Engine</h1>
      <p className="text-gray-600 mb-6">Comprehensive quality assessment across 5 dimensions</p>

      <div className="bg-white p-4 rounded-lg shadow mb-6">
        <div className="flex items-center justify-between">
          <div>
            <h2 className="font-bold">Clinical Q&A Evaluation Dataset</h2>
            <p className="text-sm text-gray-600">{sampleDataset.length} test cases covering medical knowledge retrieval</p>
          </div>
          <button onClick={runEvaluation} disabled={loading}
            className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
            {loading ? 'Evaluating...' : 'Run Full Evaluation'}
          </button>
        </div>
      </div>

      {result && (
        <div className="space-y-6">
          <div className={`p-6 rounded-lg shadow ${
            result.overall_ragas_score >= 0.8 ? 'bg-green-50 border border-green-300' :
            result.overall_ragas_score >= 0.6 ? 'bg-yellow-50 border border-yellow-300' :
            'bg-red-50 border border-red-300'
          }`}>
            <div className="text-sm opacity-75">OVERALL RAGAS SCORE</div>
            <div className={`text-5xl font-bold mt-2 ${getScoreColor(result.overall_ragas_score)}`}>
              {(result.overall_ragas_score * 100).toFixed(1)}%
            </div>
          </div>

          <div className="grid grid-cols-5 gap-4">
            <MetricCard label="Context Precision" score={result.metrics.context_precision}
              color={getScoreColor(result.metrics.context_precision)} />
            <MetricCard label="Context Recall" score={result.metrics.context_recall}
              color={getScoreColor(result.metrics.context_recall)} />
            <MetricCard label="Faithfulness" score={result.metrics.faithfulness}
              color={getScoreColor(result.metrics.faithfulness)} />
            <MetricCard label="Answer Relevancy" score={result.metrics.answer_relevancy}
              color={getScoreColor(result.metrics.answer_relevancy)} />
            <MetricCard label="Answer Correctness" score={result.metrics.answer_correctness}
              color={getScoreColor(result.metrics.answer_correctness)} />
          </div>

          <div className="grid grid-cols-2 gap-4">
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">🔍 Failure Analysis</h3>
              <div className="space-y-2">
                {Object.entries(result.failure_analysis).map(([key, indices]) => (
                  <div key={key} className="flex justify-between items-center border-b pb-2">
                    <span className="text-sm capitalize">
                      {key.replace(/_/g, ' ')}
                    </span>
                    <span className={`text-sm font-mono ${
                      (indices as number[]).length > 0 ? 'text-red-600' : 'text-green-600'
                    }`}>
                      {(indices as number[]).length} failures
                    </span>
                  </div>
                ))}
              </div>
            </div>

            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">💡 Recommendations</h3>
              <div className="text-sm text-gray-700 whitespace-pre-line">
                {result.recommendations}
              </div>
            </div>
          </div>

          <div className="bg-white p-4 rounded-lg shadow">
            <h3 className="font-bold mb-3">📋 Per-Item Breakdown</h3>
            <div className="overflow-x-auto">
              <table className="w-full text-sm">
                <thead>
                  <tr className="border-b">
                    <th className="text-left py-2">#</th>
                    <th className="text-left py-2">Query</th>
                    <th className="text-center py-2">Precision</th>
                    <th className="text-center py-2">Recall</th>
                    <th className="text-center py-2">Faithful</th>
                    <th className="text-center py-2">Relevant</th>
                    <th className="text-center py-2">Correct</th>
                  </tr>
                </thead>
                <tbody>
                  {result.per_item_reports.map((item, i) => (
                    <tr key={i} className="border-b hover:bg-gray-50">
                      <td className="py-2">{i + 1}</td>
                      <td className="py-2 text-xs">{item.query}...</td>
                      <td className={`text-center ${getScoreColor(item.context_precision)}`}>
                        {(item.context_precision * 100).toFixed(0)}%
                      </td>
                      <td className={`text-center ${getScoreColor(item.context_recall)}`}>
                        {(item.context_recall * 100).toFixed(0)}%
                      </td>
                      <td className={`text-center ${getScoreColor(item.faithfulness)}`}>
                        {(item.faithfulness * 100).toFixed(0)}%
                      </td>
                      <td className={`text-center ${getScoreColor(item.answer_relevancy)}`}>
                        {(item.answer_relevancy * 100).toFixed(0)}%
                      </td>
                      <td className={`text-center ${getScoreColor(item.answer_correctness)}`}>
                        {(item.answer_correctness * 100).toFixed(0)}%
                      </td>
                    </tr>
                  ))}
                </tbody>
              </table>
            </div>
          </div>

          {result.historical_trend.length > 1 && (
            <div className="bg-white p-4 rounded-lg shadow">
              <h3 className="font-bold mb-3">📈 Historical Trend</h3>
              <div className="space-y-2">
                {result.historical_trend.map((h, i) => (
                  <div key={i} className="flex items-center gap-3 text-sm">
                    <span className="text-xs text-gray-500 w-32">
                      {new Date(h.timestamp).toLocaleDateString()}
                    </span>
                    <div className="flex-1 bg-gray-200 rounded-full h-2">
                      <div className="bg-blue-500 h-2 rounded-full"
                        style={{ width: `${h.overall_score * 100}%` }} />
                    </div>
                    <span className="font-mono w-16 text-right">
                      {(h.overall_score * 100).toFixed(1)}%
                    </span>
                  </div>
                ))}
              </div>
            </div>
          )}
        </div>
      )}
    </div>
  );
};

Real-Time Use Case: Healthcare Clinical Q&A System

A hospital deploys a RAG system to answer clinician questions from medical guidelines, drug databases, and research papers. Before production rollout, the team runs the evaluation engine on a 200-case test set.

Evaluation results reveal:

  • Context Precision: 0.82 — retrieval is generally relevant

  • Context Recall: 0.58 — significant information gaps; many queries need docs that aren't retrieved

  • Faithfulness: 0.91 — strong grounding, low hallucination

  • Answer Relevancy: 0.87 — answers stay on topic

  • Answer Correctness: 0.64 — factual gaps due to low recall

  • Overall RAGAS: 0.75

Failure analysis shows 42 cases with low recall, concentrated in queries about drug interactions and rare conditions. The recommendation engine suggests: "Low recall: increase top-k, use hybrid search, or improve embeddings."

Action taken: The team switches from pure vector search to hybrid search (BM25 + dense), increases top-k from 5 to 10, and adds a domain-adapted medical embedding model.

Re-evaluation after changes:

  • Context Recall: 0.58 → 0.81 (+39% improvement)

  • Answer Correctness: 0.64 → 0.79 (+23% improvement)

  • Overall RAGAS: 0.75 → 0.86

The Redis-backed memory tracks this improvement over time, allowing the team to demonstrate ROI to hospital leadership and detect any future regressions within days rather than months.

Conclusion

RAG evaluation is not a one-time checkpoint it's a continuous practice that must evolve alongside your system. By decomposing evaluation into five independent dimensions (context precision, context recall, faithfulness, answer relevancy, answer correctness) and orchestrating them through a multi-agent LangGraph framework, enterprises gain a systematic, reproducible approach to quality assessment. Each dimension is measured by a specialized agent using LLM-as-judge techniques, with results aggregated into an overall RAGAS-style score and actionable recommendations. The persistent memory layer enables trend analysis, regression detection, and evidence-based architectural decisions. In regulated industries like healthcare, finance, and legal, this evaluation framework isn't optional it's the foundation of responsible AI deployment, providing the auditability and accountability that stakeholders demand. Without rigorous evaluation, you're not building a production system; you're building a liability.