Module 2 · Retrieval-augmented generation

Evaluating RAG

Advanced 18 min read If you cannot measure it, you are guessing

A RAG system can fail in two places: retrieval did not find the right text, or generation misused the text it was given. Evaluate them separately, because they have different fixes. Retrieval metrics (recall@k, MRR, nDCG) are cheap, exact and need no LLM. Generation metrics (faithfulness, relevance, correctness) usually need a judge: a person or, at scale, another LLM that you have checked against people.

🧪

Grading a student who answers open-book

First check whether they opened the right pages (retrieval). Then check whether what they wrote is actually on those pages (faithfulness), whether it answers the question asked (relevance), and whether it is right (correctness). A wrong answer from the wrong page and a wrong answer from the right page need very different coaching.

1. What to measure

question retrieval generation answer + citations retrieval metrics recall@k · MRR · nDCG generation metrics faithfulness (to context) · relevance (to question) · correctness Needs a golden set: questions with the documents that answer them, and ideally reference answers. Retrieval metrics are exact and free. Generation metrics need a judge.

2. The golden set

Everything starts with a set of questions labelled with what a correct system should do. The one used throughout this module is in mini_rag.GOLDEN: 32 questions about this site, each with the lesson(s) that answer it. Half share vocabulary with the lesson, half are paraphrased on purpose.

Where questions come from

Real user questions from logs (the best source), subject-matter experts, and LLM-generated questions from your documents reviewed by a person. Include questions the system should refuse.

How big

30–50 questions already catch big regressions. A few hundred give stable numbers for fine comparisons. Version the set in Git and add every production failure you find.

3. Retrieval metrics, implemented

MetricQuestion it answersFormula (per question, then averaged)
recall@kDid a relevant document make the top k?1 if any relevant doc is in the top k (or the fraction of relevant docs found)
MRRHow high was the first relevant one?1 / rank of the first relevant document
nDCG@kWere relevant docs near the top, with graded credit?Σ reli / log₂(i+1), divided by the ideal ordering's value
import math

from mini_rag import GOLDEN, HybridIndex, lessons_in
from site_docs import load_sections

docs = load_sections()
index = HybridIndex(docs)

def metrics(ranked, relevant, ks=(1, 3, 5, 10)):
    out = {f"R@{k}": float(bool(relevant & set(ranked[:k]))) for k in ks}
    first = next((i for i, d in enumerate(ranked, 1) if d in relevant), None)
    out["MRR"] = 1 / first if first else 0.0
    dcg = sum(1 / math.log2(i + 1) for i, d in enumerate(ranked[:10], 1) if d in relevant)
    ideal = sum(1 / math.log2(i + 1) for i in range(1, min(len(relevant), 10) + 1))
    out["nDCG@10"] = dcg / ideal
    return out

print(f"{'retriever':9}" + "".join(f"{m:>9}" for m in ("R@1", "R@3", "R@5", "R@10", "MRR", "nDCG@10")))
for mode in ("bm25", "dense", "hybrid"):
    per_question = [metrics(lessons_in(index.rank(q, mode), docs, k=10), relevant) for q, relevant in GOLDEN]
    avg = {m: sum(r[m] for r in per_question) / len(per_question) for m in per_question[0]}
    print(f"{mode:9}" + "".join(f"{v:9.2f}" for v in avg.values()))
retriever R@1 R@3 R@5 R@10 MRR nDCG@10 bm25 0.72 0.78 0.81 0.97 0.77 0.82 dense 0.75 0.84 0.88 1.00 0.81 0.85 hybrid 0.72 0.81 0.84 0.97 0.78 0.82

Read across a row: every retriever finds nearly every answer somewhere in the top 10, but only about three-quarters at rank 1. That gap is what a reranker (lesson 08) is meant to close. Read down a column to compare retrievers on the same questions. With 32 questions one question is about 0.03, so differences of that size are noise. Grow the set before trusting small gains.

4. Faithfulness: is the answer supported by the sources?

A faithful answer makes only claims that the retrieved context supports. It is the metric that catches hallucination in RAG. The rigorous way is claim-level. Split the answer into claims, then check each claim against the context with an NLI model or an LLM judge. A cheap lexical check makes a useful first alarm:

import re

STOP = set("a an the is are was of to in on for and or it its this that with by as at be can".split())

def content_words(text):
    return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}

def support(sentence, context):
    words = content_words(sentence)
    return len(words & content_words(context)) / max(1, len(words))

context = ("A covering index contains every column the query needs, so InnoDB can answer "
           "from the index alone without looking up the table rows. EXPLAIN shows 'Using index'.")
answer = ("A covering index holds every column the query needs [1]. "
          "EXPLAIN then shows 'Using index' [1]. "
          "Covering indexes also compress the table by about 90 percent [1].")

for sentence in re.split(r"(?<=[.!?])\s+", answer):
    s = support(sentence, context)
    print(f"{s:4.0%}  {'OK   ' if s >= 0.6 else 'CHECK'}  {sentence}")
75% OK A covering index holds every column the query needs [1]. 67% OK EXPLAIN then shows 'Using index' [1]. 22% CHECK Covering indexes also compress the table by about 90 percent [1].

Word overlap is fooled by negation ("does not hold every column") and misses paraphrase. Use it as a tripwire on every request, and use a proper judge on samples and in CI.

5. LLM-as-judge, and judging the judge

A judge is a prompt plus a schema: give it the question, the context and the answer, and a rubric with a small set of labels. It is itself an LLM feature, so it needs evaluation too. Label 30–100 items by hand, run the judge on them, and measure agreement. Cohen's kappa corrects raw agreement for chance: above about 0.6 is usually considered substantial, and below 0.4 means the judge should not be trusted. The judge's verdicts are scripted here; the agreement maths is real:

import json
from typing import Literal

from pydantic import BaseModel
from fake_openai import fake_client

class Verdict(BaseModel):
    unsupported_claims: list[str]              # reasoning first: list problems, then decide
    label: Literal["faithful", "unfaithful"]

JUDGE = """You grade whether an ANSWER is fully supported by the CONTEXT.
List every claim in the answer that the context does not support, then give the label.
"faithful" only if that list is empty. Reply with JSON: {"unsupported_claims": [...], "label": ...}"""

human = ["faithful", "faithful", "unfaithful", "faithful", "unfaithful", "faithful",
         "faithful", "unfaithful", "faithful", "faithful", "unfaithful", "faithful"]
scripted = ["faithful", "faithful", "unfaithful", "faithful", "faithful", "faithful",       # judge misses #5
            "faithful", "unfaithful", "unfaithful", "faithful", "unfaithful", "faithful"]  # and is harsh on #9
client = fake_client(script=[json.dumps({"unsupported_claims": [] if v == "faithful" else ["…"], "label": v})
                             for v in scripted])

judge = []
for i in range(len(human)):
    r = client.responses.create(model="gpt-5.6-luna", instructions=JUDGE,
                                input=f"<context>…</context><answer>item {i}</answer>")
    judge.append(Verdict.model_validate_json(r.output_text).label)

n = len(human)
p_observed = sum(h == j for h, j in zip(human, judge)) / n
p_h, p_j = human.count("faithful") / n, judge.count("faithful") / n
p_chance = p_h * p_j + (1 - p_h) * (1 - p_j)
kappa = (p_observed - p_chance) / (1 - p_chance)
print(f"agreement {p_observed:.0%}, chance agreement {p_chance:.0%}, Cohen's kappa {kappa:.2f}")
print("disagreements:", [i + 1 for i, (h, j) in enumerate(zip(human, judge)) if h != j])
agreement 83%, chance agreement 56%, Cohen's kappa 0.63 disagreements: [5, 9]
Judge design ruleWhy
Few, concrete labels (pass/fail, or 1–3) with a written rubric1–10 scores drift and cluster. Binary labels are more consistent.
List evidence before the label (see the schema field order)the label is then conditioned on the reasoning (lesson 04)
Use a strong model as the judge, and keep it fixed across comparisonschanging the judge changes the ruler
Check it against human labels, and re-check when you change itan unvalidated judge measures its own biases
Watch the known biases: position, verbosity, self-preferenceshuffle options, penalise padding, avoid judging a model with itself

6. Tools that package this

Libraries implement these metrics with tested judge prompts. DeepEval, for example, runs as pytest tests. The snippet uses DeepEval's documented classes; running it calls an LLM judge, so it needs an API key and was not run here:

import pytest
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase

from rag import rag                                   # your pipeline (lesson 09)

@pytest.mark.parametrize("question", ["What is a covering index?", "What does the GIL do?"])
def test_answers_are_faithful_and_relevant(question):
    answer = rag.answer(question)
    case = LLMTestCase(input=question, actual_output=answer.text,
                       retrieval_context=[s["text"] for s in answer.sources])
    assert_test(case, [FaithfulnessMetric(threshold=0.8), AnswerRelevancyMetric(threshold=0.7)])

Ragas

RAG-specific metrics (faithfulness, context precision and recall) over a dataset.

DeepEval

pytest-style metrics and a CI workflow, as above.

Langfuse

Datasets, experiment runs and scores attached to production traces (lesson 18).

7. Offline, in CI, and online

WhenWhat runsGate
Every retrieval, prompt or model change (CI)retrieval metrics on the full golden set; judge metrics on a subsetfail the build if recall@5 or faithfulness drops beyond noise
Before a releasefull judge run and a human review of the diffssign-off
In productioncheap tripwires on every request (citation validity, overlap); a judge on a 1–5% sample; user feedbackalerts and dashboards (Module 4)

Recap

  • Evaluate retrieval and generation separately; they have different fixes.
  • Retrieval: recall@k, MRR and nDCG against a golden set. Exact and free, so run them on every change.
  • Generation: faithfulness, relevance and correctness, judged by an LLM that you have validated against humans (kappa).
  • Grow the golden set from production failures, and gate CI on it.

Checkpoint

1 · Recall@10 is 0.95 but recall@1 is 0.55. What is the most promising next step?
The document is almost always in the candidate set. The problem is its ordering, which is exactly what reranking fixes.
2 · Your LLM judge agrees with human labels 85% of the time. Is that good?
Raw agreement is inflated by imbalanced labels. Kappa subtracts the agreement you would get by chance.
3 · An answer is correct but not supported by the retrieved context (the model knew it from training). Faithful?
Unsupported-but-correct today is unsupported-and-wrong tomorrow, when the facts change. RAG's promise is answers traceable to your sources.