Module 2 · Retrieval-augmented generation

Ingestion & chunking

Intermediate 18 min read Garbage in, confident garbage out

Before anything can be retrieved it must be loaded (turned into clean text), chunked (split into passages small enough to retrieve precisely and send cheaply), given metadata (where it came from, who may see it), and kept in sync as the sources change. This unglamorous pipeline decides more of your answer quality than the choice of LLM does. This lesson builds it on this website's own HTML and measures the effect of chunk size.

✂️

Cutting a textbook into index cards

Cards that are too big (a whole chapter) are hard to find and expensive to carry. Cards that are too small (one sentence) lose the context that made them meaningful. Good cards follow the book's own structure, one idea each, and every card is labelled with the book, chapter and page. That label is the metadata.

1. The ingestion pipeline

sourcesHTML · PDF · docxMarkdown · code loadclean text +headings, tables chunkby structure,capped in size enrichsource, section,access, version embed + indexonly chunks whosehash changed Run it on a schedule or on change events. It must be idempotent: running it twice changes nothing. Log what it did (added / updated / deleted / failed) and alert on failures. A silently stale index is a quality bug.

2. Loading: keep the structure

The loader's job is clean text plus the document's structure: headings, lists, tables, code. Throw away navigation, footers and cookie banners. For this site, site_docs.py keeps each lesson's <main>, drops scripts and quizzes (their wrong answer options are false statements), turns diagrams into their text description, and splits at every <h2>:

from site_docs import parse_lesson, site_root

sections = parse_lesson(site_root() / "sql" / "06-order-limit.html", course="sql")
for s in sections:
    print(f"{len(s['text'].split()):4} words  {s['url']:40} {s['section']}")
print()
print(sections[3]["text"][:260], "…")
41 words sql/06-order-limit.html Introduction 143 words sql/06-order-limit.html#basics 1. Sorting 77 words sql/06-order-limit.html#ties 2. Ties make ORDER BY non-deterministic 51 words sql/06-order-limit.html#limit 3. LIMIT and OFFSET 171 words sql/06-order-limit.html#pagination 4. Why OFFSET pagination breaks 84 words sql/06-order-limit.html#topn 5. Top-N per group is a different problem 122 words sql/06-order-limit.html#perf 6. Sorting and indexes 67 words sql/06-order-limit.html#recap Recap LIMIT 10 -- Postgres, MySQL, SQLite LIMIT 10 OFFSET 20 -- skip 20, take 10 (page 3 at 10/page) LIMIT 20, 10 -- MySQL shorthand: OFFSET 20, LIMIT 10. Confusing; avoid OFFSET 20 ROWS FETCH NEXT 10 ROWS ONLY …
Source typeTooling (Python)Pitfalls
HTMLBeautifulSoup, trafilaturaboilerplate, content injected by JavaScript
PDF with a text layerpypdf, pdfplumber; docling or unstructured for layouttwo-column pages, headers and footers on every page, hyphenation
Scanned PDFs and imagesOCR (Tesseract, cloud OCR), or a vision modelOCR errors; cost; tables become word soup
Word and PowerPointpython-docx, python-pptx, doclingspeaker notes, text boxes out of order
Codeast (split by function or class)splitting in the middle of a function
Tables and spreadsheetspandasa chunk without its header row is meaningless. Often a SQL tool is better than RAG.

3. Two ways to chunk

Fixed-size chunking cuts every N words (or tokens) with some overlap. It is simple, but it ignores meaning and cuts sentences, lists and code in half. Structure-aware chunking splits at the document's own boundaries (sections, then paragraphs, then sentences) and only packs units together up to a size limit. Compare them on the longest section on this site:

import re

from site_docs import load_sections

section = max(load_sections(courses=["sql"]), key=lambda d: len(d["text"]))
text = section["text"]
print(section["url"], "-", len(text.split()), "words")

def fixed_chunks(text, size=200, overlap=40):
    words = text.split()
    return [" ".join(words[i:i + size]) for i in range(0, max(1, len(words) - overlap), size - overlap)]

def units(text):
    """Paragraphs (site_docs writes one block per line); very long ones split into sentences."""
    for para in text.split("\n"):
        yield from ([para] if len(para.split()) <= 200 else re.split(r"(?<=[.!?])\s+", para))

def structured_chunks(text, max_words=200, overlap_units=1):
    chunks, current = [], []
    for u in units(text):
        if current and sum(len(x.split()) for x in current) + len(u.split()) > max_words:
            chunks.append("\n".join(current))
            current = current[-overlap_units:]       # repeat the last paragraph for context
        current.append(u)
    return chunks + ["\n".join(current)] if current else chunks

paragraphs = text.split("\n")
para_starts, pos = set(), 0                       # word offsets where a paragraph (or code block) starts
for p in paragraphs:
    para_starts.add(pos)
    pos += len(p.split())

fixed = fixed_chunks(text)
fixed_cuts = sum(i * 160 not in para_starts for i in range(len(fixed)))          # 160 = size - overlap
structured = structured_chunks(text)
structured_cuts = sum(c.split("\n")[0] not in paragraphs for c in structured)
for name, chunks, cuts in [("fixed 200/40", fixed, fixed_cuts), ("structured ≤200", structured, structured_cuts)]:
    sizes = [len(c.split()) for c in chunks]
    print(f"{name:16} {len(chunks):3} chunks, {min(sizes)}-{max(sizes)} words, "
          f"{cuts} start in the middle of a paragraph or code block")
print("\nfirst fixed chunk ends with:     …", fixed_chunks(text)[0][-70:])
print("first structured chunk ends with: …", structured_chunks(text)[0][-70:])
sql/28-project.html#challenges - 3433 words fixed 200/40 22 chunks, 73-200 words, 19 start in the middle of a paragraph or code block structured ≤200 19 chunks, 94-200 words, 0 start in the middle of a paragraph or code block first fixed chunk ends with: … duct_name, category, unit_price FROM products ORDER BY unit_price DESC first structured chunk ends with: … _name, category, unit_price FROM products ORDER BY unit_price DESC

4. How big should chunks be? Measure it.

Chunk size trades precision against context. The fair comparison fixes the context budget, meaning how many words you are willing to send to the model, and asks how often the lesson that answers each question makes it into that budget. This runs the 32-question golden set from mini_rag.py (lesson 10 explains it) against five chunkings of the same 164 lessons:

import numpy as np

from mini_rag import BM25, GOLDEN, LsaEmbedder, rrf
from site_docs import load_sections

sections = load_sections()
by_lesson = {}
for s in sections:
    by_lesson.setdefault(s["url"].split("#")[0], []).append(s)

def fixed(size, overlap):
    out = []
    for url, secs in by_lesson.items():
        words = " ".join(s["text"] for s in secs).split()
        for i in range(0, max(1, len(words) - overlap), size - overlap):
            out.append({"url": url, "title": secs[0]["title"], "section": "", "text": " ".join(words[i:i + size])})
    return out

def recall_within_budget(chunks, budget_words=800):
    texts = [f"{c['title']} — {c['section']}\n{c['text']}" for c in chunks]     # header + text
    bm25, dense = BM25(texts), LsaEmbedder(texts)
    scores = {}
    for mode in ("bm25", "dense", "hybrid"):
        hits = 0
        for question, answers in GOLDEN:
            b = list(np.argsort(-bm25.scores(question))[:50])
            d = list(np.argsort(-dense.scores(question))[:50])
            ranking = {"bm25": b, "dense": d, "hybrid": rrf(b, d)}[mode]
            sent, used = set(), 0
            for i in ranking:                                  # fill the budget in rank order
                n = len(chunks[i]["text"].split())
                if used + n > budget_words and sent:
                    break
                used += n
                sent.add(chunks[i]["url"].split("#")[0])
            hits += bool(answers & sent)
        scores[mode] = hits / len(GOLDEN)
    return scores

print(f"{'strategy':18}{'chunks':>7}{'median words':>14}   recall: bm25  dense  hybrid")
for name, chunks in [("fixed 60 / 10", fixed(60, 10)), ("fixed 100 / 20", fixed(100, 20)),
                     ("fixed 300 / 50", fixed(300, 50)), ("fixed 1000 / 100", fixed(1000, 100)),
                     ("h2 sections", sections)]:
    r = recall_within_budget(chunks)
    med = int(np.median([len(c["text"].split()) for c in chunks]))
    print(f"{name:18}{len(chunks):7}{med:14}          {r['bm25']:.2f}   {r['dense']:.2f}    {r['hybrid']:.2f}")
strategy chunks median words recall: bm25 dense hybrid fixed 60 / 10 2943 60 0.88 0.94 0.91 fixed 100 / 20 1856 100 0.91 0.91 0.88 fixed 300 / 50 628 300 0.75 0.81 0.78 fixed 1000 / 100 212 738 0.72 0.81 0.72 h2 sections 1229 97 0.78 0.84 0.81

On this data, big chunks lose. With 1,000-word chunks the budget holds one chunk, so one bad ranking decision costs the whole answer. Small chunks let several candidates into the prompt. The h2 sections sit in between: their median is under 100 words, but the few very long ones fill the budget on their own. A size cap, like the structured chunker above, fixes that. But small chunks have their own cost that this metric cannot see: a 60-word window often holds only part of an answer. That tension has a well-known resolution.

Small-to-big ("parent document") retrieval

Index small chunks for precise matching, but send the model the parent they came from (the whole section). Store parent_id in each chunk's metadata, retrieve on children, and de-duplicate parents before building the prompt.

Contextual headers

Prefix every chunk with where it lives ("SQL › Pagination › Keyset") before embedding it. A chunk that says "it is faster because…" then still carries what "it" is. mini_rag.doc_text() does this.

Starting points, then measure

Structure-aware chunks of a few hundred tokens at most, with a small overlap and a contextual header, are a sensible default for prose. On this site, chunks of about 100 words did best. Then run your golden set (lesson 10) and let the numbers decide. The best size depends on your documents, your questions and your context budget, which is exactly what the table above shows.

5. Metadata: the part people skip

FieldUsed for
source_id, url + anchorcitations that link to the exact place; deleting a document's chunks
title, sectioncontextual headers; showing sources to users
updated_at, versionpreferring the newest document; excluding superseded ones
tenant, allowed_groupsaccess control at retrieval time, enforced by your code, never by the prompt
language, doc_type, productfilters ("only the API reference", "only German")
content_hashincremental re-indexing (next section)

6. Keep the index in sync, cheaply

Re-embedding everything nightly works for small collections but wastes money and time at scale. Give every chunk a stable id (for example, source URL + anchor) and a content hash. Each run then upserts what changed and deletes what vanished:

import hashlib

from site_docs import load_sections

def content_hash(text):
    return hashlib.sha256(text.encode()).hexdigest()[:16]

def sync_plan(index, docs):
    """index: {chunk_id: hash} as stored in the vector DB. Returns what to upsert and delete."""
    current = {d["url"]: content_hash(d["text"]) for d in docs}
    upsert = [cid for cid, h in current.items() if index.get(cid) != h]   # new or changed
    delete = [cid for cid in index if cid not in current]                 # gone from the source
    return upsert, delete

v1 = load_sections(courses=["sql"])
index = {d["url"]: content_hash(d["text"]) for d in v1}                   # first full run

v2 = [dict(d) for d in v1]                                                # the docs change:
v2[10]["text"] += " (updated for version 2)"                             # one edited
v2[20]["text"] = v2[20]["text"].replace("SELECT", "select")              # one reformatted
del v2[30]                                                               # one removed
v2.append({"url": "sql/29-new-lesson.html", "text": "A brand new lesson."})   # one added

upsert, delete = sync_plan(index, v2)
print(f"{len(v2)} chunks in source: re-embed {len(upsert)}, delete {len(delete)}, "
      f"skip {len(v2) - len(upsert)} unchanged")
print("upsert:", upsert)
print("delete:", delete)
198 chunks in source: re-embed 2, delete 1, skip 196 unchanged upsert: ['sql/02-execution-order.html#alias', 'sql/29-new-lesson.html'] delete: ['sql/04-select.html#where']

Deletes matter as much as upserts. An index that still contains last year's refund policy will retrieve it, and the model will quote it confidently.

Recap

  • Load clean text with its structure; drop boilerplate and anything false (quiz distractors, outdated copies).
  • Chunk by structure with a size cap; add contextual headers; consider small-to-big retrieval.
  • Measure chunk size against a fixed context budget. On this site, huge chunks performed worst.
  • Metadata powers citations, filters and access control. Hashes make re-indexing incremental.

Checkpoint

1 · A chunk reads "It is 3× faster than the previous method." Retrieval never finds it for relevant questions. Best fix?
The chunk lost its subject. A contextual header ("MySQL › Covering indexes") puts the missing words back into what is indexed.
2 · Where should "user X may only see HR documents from their own country" be enforced?
Anything that reaches the prompt can leak. Access control belongs in code, before retrieval results exist.
3 · Why store a content hash with every chunk?
Compare stored and current hashes: equal means skip, different means re-embed, missing means delete.