Evaluating RAG
A RAG system can fail in two places: retrieval did not find the right text, or generation misused the text it was given. Evaluate them separately, because they have different fixes. Retrieval metrics (recall@k, MRR, nDCG) are cheap, exact and need no LLM. Generation metrics (faithfulness, relevance, correctness) usually need a judge: a person or, at scale, another LLM that you have checked against people.
Grading a student who answers open-book
First check whether they opened the right pages (retrieval). Then check whether what they wrote is actually on those pages (faithfulness), whether it answers the question asked (relevance), and whether it is right (correctness). A wrong answer from the wrong page and a wrong answer from the right page need very different coaching.
1. What to measure
2. The golden set
Everything starts with a set of questions labelled with what a correct system should do. The one used
throughout this module is in mini_rag.GOLDEN: 32 questions about this site, each with the
lesson(s) that answer it. Half share vocabulary with the lesson, half are paraphrased on purpose.
Where questions come from
Real user questions from logs (the best source), subject-matter experts, and LLM-generated questions from your documents reviewed by a person. Include questions the system should refuse.
How big
30–50 questions already catch big regressions. A few hundred give stable numbers for fine comparisons. Version the set in Git and add every production failure you find.
3. Retrieval metrics, implemented
| Metric | Question it answers | Formula (per question, then averaged) |
|---|---|---|
| recall@k | Did a relevant document make the top k? | 1 if any relevant doc is in the top k (or the fraction of relevant docs found) |
| MRR | How high was the first relevant one? | 1 / rank of the first relevant document |
| nDCG@k | Were relevant docs near the top, with graded credit? | Σ reli / log₂(i+1), divided by the ideal ordering's value |
import math
from mini_rag import GOLDEN, HybridIndex, lessons_in
from site_docs import load_sections
docs = load_sections()
index = HybridIndex(docs)
def metrics(ranked, relevant, ks=(1, 3, 5, 10)):
out = {f"R@{k}": float(bool(relevant & set(ranked[:k]))) for k in ks}
first = next((i for i, d in enumerate(ranked, 1) if d in relevant), None)
out["MRR"] = 1 / first if first else 0.0
dcg = sum(1 / math.log2(i + 1) for i, d in enumerate(ranked[:10], 1) if d in relevant)
ideal = sum(1 / math.log2(i + 1) for i in range(1, min(len(relevant), 10) + 1))
out["nDCG@10"] = dcg / ideal
return out
print(f"{'retriever':9}" + "".join(f"{m:>9}" for m in ("R@1", "R@3", "R@5", "R@10", "MRR", "nDCG@10")))
for mode in ("bm25", "dense", "hybrid"):
per_question = [metrics(lessons_in(index.rank(q, mode), docs, k=10), relevant) for q, relevant in GOLDEN]
avg = {m: sum(r[m] for r in per_question) / len(per_question) for m in per_question[0]}
print(f"{mode:9}" + "".join(f"{v:9.2f}" for v in avg.values()))
Read across a row: every retriever finds nearly every answer somewhere in the top 10, but only about three-quarters at rank 1. That gap is what a reranker (lesson 08) is meant to close. Read down a column to compare retrievers on the same questions. With 32 questions one question is about 0.03, so differences of that size are noise. Grow the set before trusting small gains.
4. Faithfulness: is the answer supported by the sources?
A faithful answer makes only claims that the retrieved context supports. It is the metric that catches hallucination in RAG. The rigorous way is claim-level. Split the answer into claims, then check each claim against the context with an NLI model or an LLM judge. A cheap lexical check makes a useful first alarm:
import re
STOP = set("a an the is are was of to in on for and or it its this that with by as at be can".split())
def content_words(text):
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}
def support(sentence, context):
words = content_words(sentence)
return len(words & content_words(context)) / max(1, len(words))
context = ("A covering index contains every column the query needs, so InnoDB can answer "
"from the index alone without looking up the table rows. EXPLAIN shows 'Using index'.")
answer = ("A covering index holds every column the query needs [1]. "
"EXPLAIN then shows 'Using index' [1]. "
"Covering indexes also compress the table by about 90 percent [1].")
for sentence in re.split(r"(?<=[.!?])\s+", answer):
s = support(sentence, context)
print(f"{s:4.0%} {'OK ' if s >= 0.6 else 'CHECK'} {sentence}")
Word overlap is fooled by negation ("does not hold every column") and misses paraphrase. Use it as a tripwire on every request, and use a proper judge on samples and in CI.
5. LLM-as-judge, and judging the judge
A judge is a prompt plus a schema: give it the question, the context and the answer, and a rubric with a small set of labels. It is itself an LLM feature, so it needs evaluation too. Label 30–100 items by hand, run the judge on them, and measure agreement. Cohen's kappa corrects raw agreement for chance: above about 0.6 is usually considered substantial, and below 0.4 means the judge should not be trusted. The judge's verdicts are scripted here; the agreement maths is real:
import json
from typing import Literal
from pydantic import BaseModel
from fake_openai import fake_client
class Verdict(BaseModel):
unsupported_claims: list[str] # reasoning first: list problems, then decide
label: Literal["faithful", "unfaithful"]
JUDGE = """You grade whether an ANSWER is fully supported by the CONTEXT.
List every claim in the answer that the context does not support, then give the label.
"faithful" only if that list is empty. Reply with JSON: {"unsupported_claims": [...], "label": ...}"""
human = ["faithful", "faithful", "unfaithful", "faithful", "unfaithful", "faithful",
"faithful", "unfaithful", "faithful", "faithful", "unfaithful", "faithful"]
scripted = ["faithful", "faithful", "unfaithful", "faithful", "faithful", "faithful", # judge misses #5
"faithful", "unfaithful", "unfaithful", "faithful", "unfaithful", "faithful"] # and is harsh on #9
client = fake_client(script=[json.dumps({"unsupported_claims": [] if v == "faithful" else ["…"], "label": v})
for v in scripted])
judge = []
for i in range(len(human)):
r = client.responses.create(model="gpt-5.6-luna", instructions=JUDGE,
input=f"<context>…</context><answer>item {i}</answer>")
judge.append(Verdict.model_validate_json(r.output_text).label)
n = len(human)
p_observed = sum(h == j for h, j in zip(human, judge)) / n
p_h, p_j = human.count("faithful") / n, judge.count("faithful") / n
p_chance = p_h * p_j + (1 - p_h) * (1 - p_j)
kappa = (p_observed - p_chance) / (1 - p_chance)
print(f"agreement {p_observed:.0%}, chance agreement {p_chance:.0%}, Cohen's kappa {kappa:.2f}")
print("disagreements:", [i + 1 for i, (h, j) in enumerate(zip(human, judge)) if h != j])
| Judge design rule | Why |
|---|---|
| Few, concrete labels (pass/fail, or 1–3) with a written rubric | 1–10 scores drift and cluster. Binary labels are more consistent. |
| List evidence before the label (see the schema field order) | the label is then conditioned on the reasoning (lesson 04) |
| Use a strong model as the judge, and keep it fixed across comparisons | changing the judge changes the ruler |
| Check it against human labels, and re-check when you change it | an unvalidated judge measures its own biases |
| Watch the known biases: position, verbosity, self-preference | shuffle options, penalise padding, avoid judging a model with itself |
6. Tools that package this
Libraries implement these metrics with tested judge prompts. DeepEval, for example, runs as pytest tests. The snippet uses DeepEval's documented classes; running it calls an LLM judge, so it needs an API key and was not run here:
import pytest
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase
from rag import rag # your pipeline (lesson 09)
@pytest.mark.parametrize("question", ["What is a covering index?", "What does the GIL do?"])
def test_answers_are_faithful_and_relevant(question):
answer = rag.answer(question)
case = LLMTestCase(input=question, actual_output=answer.text,
retrieval_context=[s["text"] for s in answer.sources])
assert_test(case, [FaithfulnessMetric(threshold=0.8), AnswerRelevancyMetric(threshold=0.7)])
Ragas
RAG-specific metrics (faithfulness, context precision and recall) over a dataset.
DeepEval
pytest-style metrics and a CI workflow, as above.
Langfuse
Datasets, experiment runs and scores attached to production traces (lesson 18).
7. Offline, in CI, and online
| When | What runs | Gate |
|---|---|---|
| Every retrieval, prompt or model change (CI) | retrieval metrics on the full golden set; judge metrics on a subset | fail the build if recall@5 or faithfulness drops beyond noise |
| Before a release | full judge run and a human review of the diffs | sign-off |
| In production | cheap tripwires on every request (citation validity, overlap); a judge on a 1–5% sample; user feedback | alerts and dashboards (Module 4) |
Recap
- Evaluate retrieval and generation separately; they have different fixes.
- Retrieval: recall@k, MRR and nDCG against a golden set. Exact and free, so run them on every change.
- Generation: faithfulness, relevance and correctness, judged by an LLM that you have validated against humans (kappa).
- Grow the golden set from production failures, and gate CI on it.