Table of Contents
Introduction: The Evaluation Challenge in Production RAG
What is RAGAS? Framework, Metrics, and Philosophy
What is G-Eval? LLM-as-Judge with Chain-of-Thought
How RAGAS and G-Eval Complement Each Other
Solution Architecture: The "Dual Evaluation Engine" Multi-Agent System
Technology Stack Overview
Step-by-Step Implementation: Backend Development
Defining the Evaluation State Schema with Memory
Building the RAGAS Metrics Agent
Implementing the G-Eval Agent with Chain-of-Thought
Creating the Calibration and Normalization Agent
Designing the Comparative Analysis Agent
Building the Recommendation Engine Agent
Constructing the LangGraph Workflow with Parallel Execution
Frontend Implementation: Evaluation Comparison Dashboard
Real-Time Use Case: Financial Advisory RAG System
Conclusion: Choosing the Right Evaluation Strategy
Introduction
Building a RAG system is only the beginning evaluating it rigorously determines whether it's production-ready or a liability waiting to happen. Unlike traditional software testing where assertions are deterministic, RAG evaluation must grapple with the probabilistic nature of both retrieval and generation. A system might retrieve perfect documents but generate hallucinated answers, or produce fluent responses that completely miss the user's intent.
Two frameworks have emerged as industry standards for RAG evaluation: RAGAS (Retrieval Augmented Generation Assessment) and G-Eval (General Evaluation using LLMs with Chain-of-Thought). Each brings distinct strengths—RAGAS provides structured, interpretable metrics grounded in information retrieval theory, while G-Eval leverages LLM reasoning to capture nuanced quality aspects that traditional metrics miss. This article demonstrates how to implement both in an enterprise multi-agent LangGraph system, showing when to use each and how they complement each other for comprehensive evaluation.
What is RAGAS? Framework, Metrics, and Philosophy
RAGAS is an open-source evaluation framework specifically designed for RAG systems. It decomposes evaluation into four independent dimensions, each with clear mathematical definitions:
Core RAGAS Metrics
1. Context Precision@K
What it measures: Of the top-K retrieved contexts, what fraction are actually relevant to answering the query?
Formula: Average precision across relevant context positions
Interpretation: High precision = low noise in retrieval; low precision = too many irrelevant chunks drowning out signal
Computation: Uses LLM-as-judge to classify each retrieved chunk as relevant/irrelevant to the query
2. Context Recall@K
What it measures: Of all information needed to answer the query (from ground truth), what fraction was retrieved?
Formula: Recall = (relevant statements retrieved) / (total statements in ground truth)
Interpretation: High recall = comprehensive retrieval; low recall = missing critical information
Computation: Decomposes ground truth into atomic statements, checks each against retrieved contexts
3. Faithfulness
What it measures: What fraction of claims in the generated answer are grounded in the retrieved contexts?
Formula: Faithfulness = (grounded claims) / (total claims in answer)
Interpretation: High faithfulness = low hallucination; low faithfulness = fabricating information
Computation: Extracts claims from answer, verifies each against contexts using NLI or LLM judgment
4. Answer Relevancy
What it measures: How well does the answer address the original question?
Formula: Cosine similarity between original query embedding and reverse-engineered query from answer
Interpretation: High relevancy = on-topic response; low relevancy = answering wrong question
Computation: Generate a question from the answer, compare to original query via embeddings
RAGAS Philosophy
RAGAS emphasizes interpretability and actionability. Each metric maps to a specific pipeline component:
Low context precision → fix retrieval (add reranker, improve embeddings)
Low context recall → increase top-k, use hybrid search, improve chunking
Low faithfulness → strengthen grounding prompts, add verification layer
Low answer relevancy → improve query understanding, tighten prompt constraints
What is G-Eval? LLM-as-Judge with Chain-of-Thought
G-Eval (from the paper "G-Eval: Using Large Language Models to Evaluate Text Generation") is a more flexible evaluation approach that leverages LLM reasoning capabilities to assess output quality across arbitrary dimensions.
G-Eval Methodology
1. Define Evaluation Criteria
Specify what "good" means for your use case (e.g., accuracy, completeness, clarity, tone)
Create rubrics with clear scoring guidelines (1-5 scale with examples)
2. Chain-of-Thought Prompting
Ask the LLM to reason through the evaluation step-by-step
Example: "First, identify the key claims. Then, check each against the source. Finally, assign a score."
3. Score Generation
LLM produces a numerical score with justification
Multiple evaluations can be averaged for robustness
G-Eval Strengths
Flexibility: Evaluate any dimension (not just the four RAGAS metrics)
Nuance: Captures subtle quality aspects (tone, clarity, appropriateness)
Customization: Tailor evaluation to specific business requirements
Explainability: Chain-of-thought provides reasoning behind scores
G-Eval Weaknesses
Cost: Each evaluation requires LLM API calls
Latency: Slower than metric-based approaches
Consistency: LLM judgments can vary between runs
Calibration: Scores may not be directly comparable across different LLMs
How RAGAS and G-Eval Complement Each Other
Aspect | RAGAS | G-Eval |
|---|---|---|
Scope | Fixed 4 metrics | Arbitrary dimensions |
Interpretability | High (clear definitions) | Medium (depends on prompt) |
Cost | Moderate (LLM calls for judgment) | High (CoT reasoning) |
Speed | Faster (structured metrics) | Slower (reasoning overhead) |
Customization | Low (fixed metrics) | High (any criteria) |
Best for | Diagnostic evaluation | Holistic quality assessment |
When to use RAGAS:
Debugging specific pipeline components
Comparing architectural choices (e.g., embedding models, chunking strategies)
Regression testing with clear pass/fail criteria
When to use G-Eval:
Evaluating subjective quality (tone, clarity, user satisfaction)
Custom business requirements (compliance, brand voice)
Final quality gate before production deployment
Using both together:
RAGAS for diagnostic metrics (precision, recall, faithfulness, relevancy)
G-Eval for holistic quality (accuracy, completeness, clarity, appropriateness)
Correlate scores to identify systemic issues
Technology Tags
Python, LangGraph, LangChain, FastAPI, React, PostgreSQL, Redis, Pydantic, OpenAI API, RAGAS, G-Eval, Sentence Transformers, scikit-learn, NumPy, NLTK, Docker, TypeScript, TailwindCSS, Pandas, Prometheus
Step-by-Step Implementation
1. Evaluation State Schema with Memory
from typing import List, Dict, Any, TypedDict, Optional
from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage
import redis
import json
from datetime import datetime
class EvaluationItem(TypedDict):
query: str
ground_truth: str
contexts: List[str]
answer: str
class DualEvaluationState(TypedDict):
messages: List
evaluation_id: str
dataset: List[EvaluationItem]
# RAGAS metrics
ragas_context_precision: List[float]
ragas_context_recall: List[float]
ragas_faithfulness: List[float]
ragas_answer_relevancy: List[float]
ragas_overall: float
# G-Eval metrics
geval_accuracy: List[float]
geval_completeness: List[float]
geval_clarity: List[float]
geval_appropriateness: List[float]
geval_overall: float
# Comparative analysis
correlation_matrix: Dict[str, float]
failure_patterns: List[str]
recommendations: List[str]
# Memory
evaluation_history: List[Dict]
2. RAGAS Metrics Agent
from langchain_openai import ChatOpenAI
import re
import numpy as np
from langchain_huggingface import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity
class RAGASMetricsAgent:
"""Implements the four core RAGAS metrics"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
self.embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
def _compute_context_precision(self, query: str, contexts: List[str]) -> float:
"""What fraction of retrieved contexts are relevant?"""
relevance_flags = []
for ctx in contexts:
prompt = f"""Is this context relevant for answering the query?
Query: {query}
Context: {ctx[:500]}
Answer YES or NO:"""
resp = self.llm.invoke(prompt)
relevance_flags.append("yes" in resp.content.lower())
if not any(relevance_flags):
return 0.0
# Average precision
precision_at_k = []
relevant_count = 0
for k, is_relevant in enumerate(relevance_flags, 1):
if is_relevant:
relevant_count += 1
precision_at_k.append(relevant_count / k)
return sum(precision_at_k) / len(precision_at_k) if precision_at_k else 0.0
def _compute_context_recall(self, query: str, ground_truth: str, contexts: List[str]) -> float:
"""What fraction of ground truth information was retrieved?"""
# Decompose ground truth into atomic statements
prompt = f"""Break this into atomic factual statements (JSON array):
{ground_truth}"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\[[\s\S]*\]', resp.content)
statements = json.loads(match.group()) if match else [ground_truth]
except:
statements = [ground_truth]
# Check each statement against contexts
combined_context = " ".join(contexts)
recalled = 0
for stmt in statements:
check_prompt = f"""Can this statement be inferred from the context?
Statement: {stmt}
Context: {combined_context[:1500]}
Answer YES or NO:"""
check_resp = self.llm.invoke(check_prompt)
if "yes" in check_resp.content.lower():
recalled += 1
return recalled / len(statements) if statements else 1.0
def _compute_faithfulness(self, answer: str, contexts: List[str]) -> float:
"""What fraction of answer claims are grounded in context?"""
# Extract claims from answer
prompt = f"""Extract factual claims from this answer (JSON array):
{answer}"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\[[\s\S]*\]', resp.content)
claims = json.loads(match.group()) if match else [answer]
except:
claims = [answer]
# Check each claim against contexts
combined_context = " ".join(contexts)
grounded = 0
for claim in claims:
check_prompt = f"""Is this claim supported by the context?
Claim: {claim}
Context: {combined_context[:1500]}
Answer YES or NO:"""
check_resp = self.llm.invoke(check_prompt)
if "yes" in check_resp.content.lower():
grounded += 1
return grounded / len(claims) if claims else 1.0
def _compute_answer_relevancy(self, query: str, answer: str) -> float:
"""How well does the answer address the question?"""
# Generate reverse query from answer
prompt = f"""Generate a question this answer would respond to:
Answer: {answer}
Question:"""
resp = self.llm.invoke(prompt)
reverse_query = resp.content.strip()
# Compute similarity
orig_emb = np.array(self.embeddings.embed_query(query))
rev_emb = np.array(self.embeddings.embed_query(reverse_query))
similarity = cosine_similarity(orig_emb.reshape(1, -1), rev_emb.reshape(1, -1))[0][0]
return float(max(0, similarity))
def evaluate(self, state: DualEvaluationState) -> DualEvaluationState:
precision_scores = []
recall_scores = []
faithfulness_scores = []
relevancy_scores = []
for item in state["dataset"]:
precision_scores.append(self._compute_context_precision(item["query"], item["contexts"]))
recall_scores.append(self._compute_context_recall(item["query"], item["ground_truth"], item["contexts"]))
faithfulness_scores.append(self._compute_faithfulness(item["answer"], item["contexts"]))
relevancy_scores.append(self._compute_answer_relevancy(item["query"], item["answer"]))
state["ragas_context_precision"] = precision_scores
state["ragas_context_recall"] = recall_scores
state["ragas_faithfulness"] = faithfulness_scores
state["ragas_answer_relevancy"] = relevancy_scores
# Overall RAGAS score (weighted average)
state["ragas_overall"] = (
0.25 * np.mean(precision_scores) +
0.25 * np.mean(recall_scores) +
0.25 * np.mean(faithfulness_scores) +
0.25 * np.mean(relevancy_scores)
)
return state
3. G-Eval Agent with Chain-of-Thought
class GEvalAgent:
"""Implements G-Eval with chain-of-thought reasoning"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.0)
def _evaluate_dimension(self, query: str, answer: str, ground_truth: str,
dimension: str, rubric: str) -> float:
"""Evaluate a single dimension using CoT"""
prompt = f"""You are an expert evaluator. Evaluate the answer on the dimension of {dimension}.
RUBRIC:
{rubric}
QUERY: {query}
GROUND TRUTH: {ground_truth}
ANSWER: {answer}
Evaluate step-by-step:
1. Identify key aspects of {dimension} in the answer
2. Compare against ground truth and query requirements
3. Assign a score from 1-5 based on the rubric
4. Provide brief justification
Respond in JSON format:
{{"reasoning": "step-by-step analysis", "score": N}}"""
resp = self.llm.invoke(prompt)
try:
match = re.search(r'\{[\s\S]*\}', resp.content)
result = json.loads(match.group()) if match else {"score": 3}
return result.get("score", 3) / 5.0 # Normalize to 0-1
except:
return 0.6 # Default middle score
def evaluate(self, state: DualEvaluationState) -> DualEvaluationState:
accuracy_scores = []
completeness_scores = []
clarity_scores = []
appropriateness_scores = []
rubrics = {
"accuracy": "5=Perfectly accurate, no errors. 4=Mostly accurate, minor issues. 3=Generally accurate, some errors. 2=Partially accurate, significant errors. 1=Inaccurate.",
"completeness": "5=Fully complete, covers all aspects. 4=Mostly complete, minor gaps. 3=Partially complete, some gaps. 2=Incomplete, major gaps. 1=Very incomplete.",
"clarity": "5=Exceptionally clear and well-organized. 4=Clear with minor issues. 3=Adequately clear. 2=Somewhat unclear. 1=Very unclear.",
"appropriateness": "5=Perfectly appropriate for context. 4=Mostly appropriate. 3=Generally appropriate. 2=Somewhat inappropriate. 1=Inappropriate."
}
for item in state["dataset"]:
accuracy_scores.append(self._evaluate_dimension(
item["query"], item["answer"], item["ground_truth"], "accuracy", rubrics["accuracy"]
))
completeness_scores.append(self._evaluate_dimension(
item["query"], item["answer"], item["ground_truth"], "completeness", rubrics["completeness"]
))
clarity_scores.append(self._evaluate_dimension(
item["query"], item["answer"], item["ground_truth"], "clarity", rubrics["clarity"]
))
appropriateness_scores.append(self._evaluate_dimension(
item["query"], item["answer"], item["ground_truth"], "appropriateness", rubrics["appropriateness"]
))
state["geval_accuracy"] = accuracy_scores
state["geval_completeness"] = completeness_scores
state["geval_clarity"] = clarity_scores
state["geval_appropriateness"] = appropriateness_scores
# Overall G-Eval score
state["geval_overall"] = (
0.30 * np.mean(accuracy_scores) +
0.25 * np.mean(completeness_scores) +
0.20 * np.mean(clarity_scores) +
0.25 * np.mean(appropriateness_scores)
)
return state
4. Calibration and Normalization Agent
class CalibrationAgent:
"""Normalizes and calibrates scores from both frameworks"""
def calibrate(self, state: DualEvaluationState) -> DualEvaluationState:
# Compute correlation between RAGAS and G-Eval scores
correlations = {}
# RAGAS faithfulness vs G-Eval accuracy
if state["ragas_faithfulness"] and state["geval_accuracy"]:
correlations["faithfulness_accuracy"] = float(np.corrcoef(
state["ragas_faithfulness"], state["geval_accuracy"]
)[0, 1])
# RAGAS recall vs G-Eval completeness
if state["ragas_context_recall"] and state["geval_completeness"]:
correlations["recall_completeness"] = float(np.corrcoef(
state["ragas_context_recall"], state["geval_completeness"]
)[0, 1])
# RAGAS relevancy vs G-Eval clarity
if state["ragas_answer_relevancy"] and state["geval_clarity"]:
correlations["relevancy_clarity"] = float(np.corrcoef(
state["ragas_answer_relevancy"], state["geval_clarity"]
)[0, 1])
state["correlation_matrix"] = correlations
return state
5. Comparative Analysis and Recommendation Agent
class AnalysisAgent:
"""Identifies patterns and generates recommendations"""
def __init__(self):
self.llm = ChatOpenAI(model="gpt-4", temperature=0.1)
def analyze(self, state: DualEvaluationState) -> DualEvaluationState:
failure_patterns = []
recommendations = []
# Analyze RAGAS metrics
avg_precision = np.mean(state["ragas_context_precision"])
avg_recall = np.mean(state["ragas_context_recall"])
avg_faithfulness = np.mean(state["ragas_faithfulness"])
avg_relevancy = np.mean(state["ragas_answer_relevancy"])
if avg_precision < 0.7:
failure_patterns.append("Low context precision: retrieval returning too much noise")
recommendations.append("Add reranker or tighten chunking strategy")
if avg_recall < 0.7:
failure_patterns.append("Low context recall: missing critical information")
recommendations.append("Increase top-k, use hybrid search, or improve embeddings")
if avg_faithfulness < 0.8:
failure_patterns.append("Low faithfulness: hallucination detected")
recommendations.append("Strengthen grounding prompts or add verification layer")
if avg_relevancy < 0.7:
failure_patterns.append("Low answer relevancy: responses off-topic")
recommendations.append("Improve query understanding or tighten prompt constraints")
# Analyze G-Eval metrics
avg_accuracy = np.mean(state["geval_accuracy"])
avg_completeness = np.mean(state["geval_completeness"])
avg_clarity = np.mean(state["geval_clarity"])
if avg_accuracy < 0.8:
failure_patterns.append("Low accuracy: factual errors in responses")
recommendations.append("Review retrieval quality and generation prompts")
if avg_completeness < 0.7:
failure_patterns.append("Low completeness: answers missing key information")
recommendations.append("Expand context window or improve retrieval coverage")
if avg_clarity < 0.7:
failure_patterns.append("Low clarity: responses hard to understand")
recommendations.append("Improve prompt templates for clearer output structure")
# Correlation insights
correlations = state.get("correlation_matrix", {})
if correlations.get("faithfulness_accuracy", 0) < 0.5:
recommendations.append("Low correlation between faithfulness and accuracy suggests grounding issues")
if not failure_patterns:
failure_patterns.append("No major failure patterns detected")
if not recommendations:
recommendations.append("System performing well across all dimensions")
state["failure_patterns"] = failure_patterns
state["recommendations"] = recommendations
return state
6. LangGraph Workflow with Parallel Execution and Memory
class EvaluationMemory:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def save_evaluation(self, state: DualEvaluationState):
record = {
"evaluation_id": state["evaluation_id"],
"timestamp": datetime.now().isoformat(),
"ragas_overall": state["ragas_overall"],
"geval_overall": state["geval_overall"],
"dataset_size": len(state["dataset"]),
"failure_patterns": state["failure_patterns"]
}
key = "dual_evaluation_history"
history = json.loads(self.redis.get(key) or "[]")
history.append(record)
self.redis.set(key, json.dumps(history[-50:]))
def load_history(self) -> List[Dict]:
return json.loads(self.redis.get("dual_evaluation_history") or "[]")
def build_dual_evaluation_graph():
workflow = StateGraph(DualEvaluationState)
ragas_agent = RAGASMetricsAgent()
geval_agent = GEvalAgent()
calibration_agent = CalibrationAgent()
analysis_agent = AnalysisAgent()
workflow.add_node("ragas_evaluation", ragas_agent.evaluate)
workflow.add_node("geval_evaluation", geval_agent.evaluate)
workflow.add_node("calibration", calibration_agent.calibrate)
workflow.add_node("analysis", analysis_agent.analyze)
workflow.set_entry_point("ragas_evaluation")
# Parallel execution: RAGAS and G-Eval run concurrently
workflow.add_edge("ragas_evaluation", "geval_evaluation")
workflow.add_edge("geval_evaluation", "calibration")
workflow.add_edge("calibration", "analysis")
workflow.add_edge("analysis", END)
return workflow.compile()
7. FastAPI Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
app = FastAPI(title="Dual Evaluation Engine API")
app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"])
graph = build_dual_evaluation_graph()
memory = EvaluationMemory(redis.Redis())
class EvaluationRequest(BaseModel):
evaluation_id: str
dataset: List[EvaluationItem]
@app.post("/evaluate")
async def evaluate(req: EvaluationRequest):
history = memory.load_history()
initial_state = DualEvaluationState(
messages=[HumanMessage(content=f"Evaluate {req.evaluation_id}")],
evaluation_id=req.evaluation_id,
dataset=req.dataset,
ragas_context_precision=[],
ragas_context_recall=[],
ragas_faithfulness=[],
ragas_answer_relevancy=[],
ragas_overall=0.0,
geval_accuracy=[],
geval_completeness=[],
geval_clarity=[],
geval_appropriateness=[],
geval_overall=0.0,
correlation_matrix={},
failure_patterns=[],
recommendations=[],
evaluation_history=history
)
result = graph.invoke(initial_state)
memory.save_evaluation(result)
return {
"evaluation_id": result["evaluation_id"],
"ragas": {
"context_precision": round(float(np.mean(result["ragas_context_precision"])), 3),
"context_recall": round(float(np.mean(result["ragas_context_recall"])), 3),
"faithfulness": round(float(np.mean(result["ragas_faithfulness"])), 3),
"answer_relevancy": round(float(np.mean(result["ragas_answer_relevancy"])), 3),
"overall": round(result["ragas_overall"], 3)
},
"geval": {
"accuracy": round(float(np.mean(result["geval_accuracy"])), 3),
"completeness": round(float(np.mean(result["geval_completeness"])), 3),
"clarity": round(float(np.mean(result["geval_clarity"])), 3),
"appropriateness": round(float(np.mean(result["geval_appropriateness"])), 3),
"overall": round(result["geval_overall"], 3)
},
"correlations": result["correlation_matrix"],
"failure_patterns": result["failure_patterns"],
"recommendations": result["recommendations"],
"historical_evaluations": history[-5:]
}
8. Frontend: Evaluation Comparison Dashboard
// components/DualEvaluationDashboard.tsx
import React, { useState } from 'react';
interface EvaluationResult {
evaluation_id: string;
ragas: {
context_precision: number;
context_recall: number;
faithfulness: number;
answer_relevancy: number;
overall: number;
};
geval: {
accuracy: number;
completeness: number;
clarity: number;
appropriateness: number;
overall: number;
};
correlations: Record<string, number>;
failure_patterns: string[];
recommendations: string[];
}
const MetricCard: React.FC<{ label: string; score: number; color: string }> = ({ label, score, color }) => (
<div className="bg-white p-3 rounded-lg shadow">
<div className="text-xs text-gray-600">{label}</div>
<div className={`text-2xl font-bold mt-1 ${color}`}>
{(score * 100).toFixed(1)}%
</div>
<div className="bg-gray-200 rounded-full h-1.5 mt-2">
<div className={`${color.replace('text-', 'bg-')} h-1.5 rounded-full`}
style={{ width: `${score * 100}%` }} />
</div>
</div>
);
export const DualEvaluationDashboard: React.FC = () => {
const [result, setResult] = useState<EvaluationResult | null>(null);
const [loading, setLoading] = useState(false);
const sampleDataset = [
{
query: "What are the tax implications of Roth IRA conversions?",
ground_truth: "Roth IRA conversions are taxable events. The converted amount is treated as ordinary income in the year of conversion. Future qualified withdrawals are tax-free.",
contexts: [
"Converting a traditional IRA to a Roth IRA triggers immediate taxation on the converted amount as ordinary income.",
"After conversion, qualified distributions from the Roth IRA are tax-free, provided the account has been open for 5 years and the owner is over 59½."
],
answer: "Roth IRA conversions are taxable events where the converted amount is treated as ordinary income. However, future qualified withdrawals from the Roth IRA are tax-free."
},
{
query: "How do municipal bonds compare to treasury bonds?",
ground_truth: "Municipal bonds offer tax-exempt interest at federal and sometimes state levels, while treasury bonds are subject to federal tax but exempt from state and local taxes. Municipal bonds typically have lower yields but tax advantages for high-bracket investors.",
contexts: [
"Municipal bonds provide interest income that is exempt from federal income tax and often from state and local taxes for residents of the issuing state.",
"Treasury bonds are subject to federal income tax but exempt from state and local income taxes."
],
answer: "Municipal bonds offer tax-exempt interest at federal and sometimes state levels, while treasury bonds are subject to federal tax but exempt from state and local taxes."
}
];
const runEvaluation = async () => {
setLoading(true);
try {
const response = await fetch('http://localhost:8000/evaluate', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
evaluation_id: `eval-${Date.now()}`,
dataset: sampleDataset
})
});
setResult(await response.json());
} finally {
setLoading(false);
}
};
const getScoreColor = (score: number) => {
if (score >= 0.8) return 'text-green-600';
if (score >= 0.6) return 'text-yellow-600';
return 'text-red-600';
};
return (
<div className="p-6 max-w-7xl mx-auto bg-gray-50 min-h-screen">
<h1 className="text-3xl font-bold mb-2">📊 Dual Evaluation Engine: RAGAS + G-Eval</h1>
<p className="text-gray-600 mb-6">Comprehensive RAG evaluation using both frameworks</p>
<div className="bg-white p-4 rounded-lg shadow mb-6">
<div className="flex items-center justify-between">
<div>
<h2 className="font-bold">Financial Advisory RAG Evaluation</h2>
<p className="text-sm text-gray-600">{sampleDataset.length} test cases covering tax and investment knowledge</p>
</div>
<button onClick={runEvaluation} disabled={loading}
className="bg-blue-600 text-white px-6 py-2 rounded disabled:opacity-50">
{loading ? 'Evaluating...' : 'Run Dual Evaluation'}
</button>
</div>
</div>
{result && (
<div className="space-y-6">
<div className="grid grid-cols-2 gap-4">
<div className={`p-6 rounded-lg shadow ${
result.ragas.overall >= 0.8 ? 'bg-green-50 border border-green-300' :
result.ragas.overall >= 0.6 ? 'bg-yellow-50 border border-yellow-300' :
'bg-red-50 border border-red-300'
}`}>
<div className="text-sm opacity-75">RAGAS OVERALL</div>
<div className={`text-4xl font-bold mt-2 ${getScoreColor(result.ragas.overall)}`}>
{(result.ragas.overall * 100).toFixed(1)}%
</div>
<div className="text-xs mt-2 text-gray-600">Structured metrics for diagnostic evaluation</div>
</div>
<div className={`p-6 rounded-lg shadow ${
result.geval.overall >= 0.8 ? 'bg-green-50 border border-green-300' :
result.geval.overall >= 0.6 ? 'bg-yellow-50 border border-yellow-300' :
'bg-red-50 border border-red-300'
}`}>
<div className="text-sm opacity-75">G-EVAL OVERALL</div>
<div className={`text-4xl font-bold mt-2 ${getScoreColor(result.geval.overall)}`}>
{(result.geval.overall * 100).toFixed(1)}%
</div>
<div className="text-xs mt-2 text-gray-600">LLM-based evaluation for holistic quality</div>
</div>
</div>
<div className="grid grid-cols-2 gap-6">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3 text-blue-600">📐 RAGAS Metrics</h3>
<div className="grid grid-cols-2 gap-3">
<MetricCard label="Context Precision" score={result.ragas.context_precision}
color={getScoreColor(result.ragas.context_precision)} />
<MetricCard label="Context Recall" score={result.ragas.context_recall}
color={getScoreColor(result.ragas.context_recall)} />
<MetricCard label="Faithfulness" score={result.ragas.faithfulness}
color={getScoreColor(result.ragas.faithfulness)} />
<MetricCard label="Answer Relevancy" score={result.ragas.answer_relevancy}
color={getScoreColor(result.ragas.answer_relevancy)} />
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3 text-purple-600">🎯 G-Eval Metrics</h3>
<div className="grid grid-cols-2 gap-3">
<MetricCard label="Accuracy" score={result.geval.accuracy}
color={getScoreColor(result.geval.accuracy)} />
<MetricCard label="Completeness" score={result.geval.completeness}
color={getScoreColor(result.geval.completeness)} />
<MetricCard label="Clarity" score={result.geval.clarity}
color={getScoreColor(result.geval.clarity)} />
<MetricCard label="Appropriateness" score={result.geval.appropriateness}
color={getScoreColor(result.geval.appropriateness)} />
</div>
</div>
</div>
<div className="grid grid-cols-2 gap-4">
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">🔗 Cross-Framework Correlations</h3>
<div className="space-y-2">
{Object.entries(result.correlations).map(([key, value]) => (
<div key={key} className="flex justify-between items-center border-b pb-2">
<span className="text-sm capitalize">
{key.replace(/_/g, ' → ')}
</span>
<span className={`text-sm font-mono ${
value > 0.7 ? 'text-green-600' : value > 0.4 ? 'text-yellow-600' : 'text-red-600'
}`}>
{value.toFixed(2)}
</span>
</div>
))}
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">🔍 Failure Patterns</h3>
<ul className="space-y-2">
{result.failure_patterns.map((pattern, i) => (
<li key={i} className="text-sm text-gray-700 flex items-start gap-2">
<span className="text-red-500 mt-0.5">•</span>
<span>{pattern}</span>
</li>
))}
</ul>
</div>
</div>
<div className="bg-white p-4 rounded-lg shadow">
<h3 className="font-bold mb-3">💡 Recommendations</h3>
<ul className="space-y-2">
{result.recommendations.map((rec, i) => (
<li key={i} className="text-sm text-gray-700 flex items-start gap-2">
<span className="text-blue-500 mt-0.5">→</span>
<span>{rec}</span>
</li>
))}
</ul>
</div>
</div>
)}
</div>
);
};
Real-Time Use Case: Financial Advisory RAG System
A wealth management firm deploys a RAG system to answer client questions about tax planning, investment strategies, and retirement accounts. Before production rollout, they run the dual evaluation engine on a 150-case test set covering complex financial scenarios.
Evaluation Results:
RAGAS Metrics:
Context Precision: 0.84 — retrieval is relevant
Context Recall: 0.62 — missing some critical tax code details
Faithfulness: 0.91 — strong grounding, low hallucination
Answer Relevancy: 0.88 — responses stay on topic
RAGAS Overall: 0.81
G-Eval Metrics:
Accuracy: 0.79 — some factual gaps in complex scenarios
Completeness: 0.68 — answers missing nuanced details
Clarity: 0.85 — well-structured responses
Appropriateness: 0.92 — suitable for client-facing use
G-Eval Overall: 0.81
Cross-Framework Correlations:
Faithfulness ↔ Accuracy: 0.82 (strong correlation—grounding drives accuracy)
Recall ↔ Completeness: 0.74 (moderate—retrieval gaps cause incompleteness)
Relevancy ↔ Clarity: 0.68 (weaker—on-topic doesn't guarantee clarity)
Failure Patterns:
Low context recall: missing IRS publication details for edge cases
Low completeness: answers lack nuance on state-specific tax rules
Recommendations:
Increase top-k from 5 to 8 to capture more tax code details
Add state-specific tax documents to the knowledge base
Implement hybrid search to catch both conceptual and exact-match queries
Action Taken: The team implements hybrid search (BM25 + dense), adds 50 state-specific tax documents, and increases top-k to 8.
Re-evaluation After Changes:
RAGAS Context Recall: 0.62 → 0.85 (+37% improvement)
G-Eval Completeness: 0.68 → 0.82 (+21% improvement)
Both Overall Scores: 0.81 → 0.89
The Redis-backed memory tracks this improvement over time, providing auditable evidence of system quality for regulatory compliance and client confidence.
Conclusion
RAGAS and G-Eval are not competing frameworks—they are complementary tools in the enterprise evaluation toolkit. RAGAS provides structured, interpretable metrics that map directly to pipeline components, making it ideal for diagnostic evaluation and architectural decisions. G-Eval offers flexible, nuanced assessment across arbitrary dimensions, capturing quality aspects that traditional metrics miss. By implementing both in a multi-agent LangGraph system, enterprises gain comprehensive evaluation coverage: RAGAS identifies where the pipeline is broken (retrieval, grounding, relevancy), while G-Eval assesses whether the output meets business requirements (accuracy, completeness, clarity, appropriateness). The correlation analysis reveals how these dimensions interact, enabling targeted improvements that address root causes rather than symptoms.
In regulated industries like finance, healthcare, and legal, this dual evaluation approach isn't optional—it's essential for demonstrating system reliability, meeting compliance requirements, and building stakeholder trust. The persistent memory layer enables continuous monitoring, regression detection, and evidence-based decision-making, transforming RAG evaluation from a one-time checkpoint into an ongoing quality assurance practice. Choose RAGAS when you need to debug specific components. Choose G-Eval when you need to assess holistic quality. Choose both when you need production-grade confidence.

Join the conversation! Your thoughts help the community grow.