Why RAG exists
Retrieval-augmented generation (RAG) means: before asking the model, search your documents for the passages relevant to the question, put those passages in the prompt, and ask the model to answer from them, with citations. It is the most common production LLM pattern, because it fixes the three things a model alone cannot do: know recent facts, know your private facts, and show where an answer came from.
An open-book exam
A closed-book exam tests what you memorised, and under pressure you might invent a plausible answer. In an open-book exam you find the right page first and answer from it, and you can point at the page. RAG turns every question into an open-book exam. The quality of the answer now depends heavily on whether you found the right page, which is why Module 2 spends most of its time on retrieval.
1. What a model alone cannot do
Recent facts
Training stops at the knowledge cut-off. Last week's release notes do not exist for the model.
Private facts
Your wiki, tickets, contracts and code were never in the training data, and should not be.
Changing facts
Prices, policies and inventory change daily. Retraining for each change is absurd. Re-indexing a document takes seconds.
Provenance and access
Users and auditors need to see sources. Different users may see different documents, and retrieval can enforce that per request.
2. RAG in one picture
3. "Why not paste everything into the prompt?"
Context windows are huge now, so it is a fair question. Measure it on a real collection: the lessons of this website.
from statistics import median
from site_docs import load_sections
docs = load_sections() # one record per lesson section
chars = sum(len(d["text"]) for d in docs)
corpus_tokens = chars / 4 # ~4 characters per token (lesson 02)
rag_tokens = 5 * median(len(d["text"]) for d in docs) / 4 + 600 # 5 sections + instructions
PRICE_IN = 0.20 / 1e6 # gpt-5.6-luna, USD per input token
PREFILL_TPS = 5_000 # illustrative prefill speed (lesson 03)
lessons = len({d["url"].split("#")[0] for d in docs})
print(f"{len(docs):,} sections from {lessons} lessons, about {corpus_tokens:,.0f} tokens")
for name, t in [("stuff everything", corpus_tokens), ("retrieve 5 sections", rag_tokens)]:
print(f"{name:20} {t:>9,.0f} input tokens ${t * PRICE_IN:.4f}/question "
f"${t * PRICE_IN * 10_000:>7,.2f} per 10k questions ~{t / PREFILL_TPS:.1f}s prefill")
print(f"ratio: {corpus_tokens / rag_tokens:.0f}x more tokens to stuff everything")
Even this modest collection is a quarter of a million tokens: past many context windows, and nearly 200 times the cost and prefill time of sending five good sections. A company wiki is a hundred times bigger. Long context has its place (one long contract, a whole codebase file), but it complements retrieval rather than replacing it. Recall also degrades in the middle of very long prompts (lesson 02).
4. RAG, fine-tuning or long context?
| RAG | Fine-tuning | Long context | |
|---|---|---|---|
| Adds new facts | yes, and updatable in seconds | poorly, and it goes stale | yes, for what fits |
| Changes style and format | somewhat, via instructions | yes, its real strength | somewhat |
| Citations | natural: you know which chunks were used | no | possible |
| Per-user access control | filter at retrieval time | no: baked into weights | you choose what to send |
| Cost per question | low | low at inference, high to train | high for large inputs |
| Main failure mode | retrieval misses the right text | confidently outdated | lost in the middle, slow |
Knowledge problem → RAG. Behaviour problem (tone, format, a narrow task done cheaply at scale) → prompt first, then fine-tune. Small, fixed documents → just put them in the prompt. Numbers and aggregates ("revenue by month") → a SQL tool, not RAG (see the OpenAI SDK course project).
5. How RAG fails (and where each lesson fixes it)
| Symptom | Usual cause | Fix |
|---|---|---|
| "I don't know" when the answer exists | retrieval missed it (wording mismatch, bad chunking) | hybrid search, better chunks (07, 08) |
| Answer mixes up two products | chunks lost their context ("it supports…" — what does?) | structure-aware chunks with titles and metadata (07) |
| Outdated answer | stale index; two versions of a document | incremental re-indexing, version metadata, filters (07, 08) |
| Correct sources, wrong answer | the model ignored or misread the context | clearer prompt, fewer and better chunks, faithfulness evals (04, 10) |
| "How many…?" answered wrongly | aggregation over many documents | a structured query tool, not top-k retrieval (11) |
| Leaks a document to the wrong user | no access filter at retrieval | per-user metadata filters, enforced in code (08, 20) |
6. Beyond the basic pipeline
Agentic RAG
Search is a tool. The model decides when to search, reformulates, and searches again if the first results are weak (lesson 11).
Hybrid with tools
Documents for "how and why", SQL for "how many", APIs for "what is the status now".
Graph RAG
Extract entities and relations into a graph for questions that span many documents. Powerful, but costly to build. Reach for it last.
Recap
- RAG = retrieve, then generate from what was retrieved, with citations.
- It solves recent, private and changing knowledge, plus provenance and access control.
- Stuffing everything costs far more (178× on this site) and is slower. Retrieval sends only what matters.
- Most RAG failures are retrieval failures, so measure retrieval separately (lesson 10).