Module 2 · Retrieval-augmented generation

Why RAG exists

Intermediate 14 min read Open-book answers from your own documents

Retrieval-augmented generation (RAG) means: before asking the model, search your documents for the passages relevant to the question, put those passages in the prompt, and ask the model to answer from them, with citations. It is the most common production LLM pattern, because it fixes the three things a model alone cannot do: know recent facts, know your private facts, and show where an answer came from.

📖

An open-book exam

A closed-book exam tests what you memorised, and under pressure you might invent a plausible answer. In an open-book exam you find the right page first and answer from it, and you can point at the page. RAG turns every question into an open-book exam. The quality of the answer now depends heavily on whether you found the right page, which is why Module 2 spends most of its time on retrieval.

1. What a model alone cannot do

Recent facts

Training stops at the knowledge cut-off. Last week's release notes do not exist for the model.

Private facts

Your wiki, tickets, contracts and code were never in the training data, and should not be.

Changing facts

Prices, policies and inventory change daily. Retraining for each change is absurd. Re-indexing a document takes seconds.

Provenance and access

Users and auditors need to see sources. Different users may see different documents, and retrieval can enforce that per request.

2. RAG in one picture

OFFLINE · when documents change (lesson 07) load documents split into chunks embed each chunk indexvectors + text + metadata ONLINE · every question (lessons 08–09) question"why is my join slow?" retrieve top-ksearch the index build the promptrules + sources + question LLM answersgrounded, cites [1] [2] The model never "learns" your documents. It reads a few relevant ones at question time, like a person with a search engine. Update a document, re-index it, and the very next answer uses the new version.

3. "Why not paste everything into the prompt?"

Context windows are huge now, so it is a fair question. Measure it on a real collection: the lessons of this website.

from statistics import median

from site_docs import load_sections

docs = load_sections()                                  # one record per lesson section
chars = sum(len(d["text"]) for d in docs)
corpus_tokens = chars / 4                               # ~4 characters per token (lesson 02)
rag_tokens = 5 * median(len(d["text"]) for d in docs) / 4 + 600   # 5 sections + instructions

PRICE_IN = 0.20 / 1e6                                   # gpt-5.6-luna, USD per input token
PREFILL_TPS = 5_000                                     # illustrative prefill speed (lesson 03)
lessons = len({d["url"].split("#")[0] for d in docs})
print(f"{len(docs):,} sections from {lessons} lessons, about {corpus_tokens:,.0f} tokens")
for name, t in [("stuff everything", corpus_tokens), ("retrieve 5 sections", rag_tokens)]:
    print(f"{name:20} {t:>9,.0f} input tokens  ${t * PRICE_IN:.4f}/question  "
          f"${t * PRICE_IN * 10_000:>7,.2f} per 10k questions  ~{t / PREFILL_TPS:.1f}s prefill")
print(f"ratio: {corpus_tokens / rag_tokens:.0f}x more tokens to stuff everything")
1,229 sections from 164 lessons, about 256,211 tokens stuff everything 256,211 input tokens $0.0512/question $ 512.42 per 10k questions ~51.2s prefill retrieve 5 sections 1,440 input tokens $0.0003/question $ 2.88 per 10k questions ~0.3s prefill ratio: 178x more tokens to stuff everything

Even this modest collection is a quarter of a million tokens: past many context windows, and nearly 200 times the cost and prefill time of sending five good sections. A company wiki is a hundred times bigger. Long context has its place (one long contract, a whole codebase file), but it complements retrieval rather than replacing it. Recall also degrades in the middle of very long prompts (lesson 02).

4. RAG, fine-tuning or long context?

RAGFine-tuningLong context
Adds new factsyes, and updatable in secondspoorly, and it goes staleyes, for what fits
Changes style and formatsomewhat, via instructionsyes, its real strengthsomewhat
Citationsnatural: you know which chunks were usednopossible
Per-user access controlfilter at retrieval timeno: baked into weightsyou choose what to send
Cost per questionlowlow at inference, high to trainhigh for large inputs
Main failure moderetrieval misses the right textconfidently outdatedlost in the middle, slow
Rule of thumb

Knowledge problem → RAG. Behaviour problem (tone, format, a narrow task done cheaply at scale) → prompt first, then fine-tune. Small, fixed documents → just put them in the prompt. Numbers and aggregates ("revenue by month") → a SQL tool, not RAG (see the OpenAI SDK course project).

5. How RAG fails (and where each lesson fixes it)

SymptomUsual causeFix
"I don't know" when the answer existsretrieval missed it (wording mismatch, bad chunking)hybrid search, better chunks (07, 08)
Answer mixes up two productschunks lost their context ("it supports…" — what does?)structure-aware chunks with titles and metadata (07)
Outdated answerstale index; two versions of a documentincremental re-indexing, version metadata, filters (07, 08)
Correct sources, wrong answerthe model ignored or misread the contextclearer prompt, fewer and better chunks, faithfulness evals (04, 10)
"How many…?" answered wronglyaggregation over many documentsa structured query tool, not top-k retrieval (11)
Leaks a document to the wrong userno access filter at retrievalper-user metadata filters, enforced in code (08, 20)

6. Beyond the basic pipeline

Agentic RAG

Search is a tool. The model decides when to search, reformulates, and searches again if the first results are weak (lesson 11).

Hybrid with tools

Documents for "how and why", SQL for "how many", APIs for "what is the status now".

Graph RAG

Extract entities and relations into a graph for questions that span many documents. Powerful, but costly to build. Reach for it last.

Recap

  • RAG = retrieve, then generate from what was retrieved, with citations.
  • It solves recent, private and changing knowledge, plus provenance and access control.
  • Stuffing everything costs far more (178× on this site) and is slower. Retrieval sends only what matters.
  • Most RAG failures are retrieval failures, so measure retrieval separately (lesson 10).

Checkpoint

1 · A support bot must answer from policy documents that change weekly and differ by country. Best approach?
Weekly changes favour re-indexing over retraining, and per-country filtering keeps answers to the right policy.
2 · "What was our total refund amount last quarter?" What should answer this?
Aggregates need all the rows, computed exactly. Retrieval returns a handful of passages, so the sum would be wrong.
3 · The right section exists, but the bot keeps saying "I don't know". Where do you look first?
If the passage never reaches the prompt, no model can use it. Check retrieval first, with recall@k (lesson 10).