Ingestion & chunking
Before anything can be retrieved it must be loaded (turned into clean text), chunked (split into passages small enough to retrieve precisely and send cheaply), given metadata (where it came from, who may see it), and kept in sync as the sources change. This unglamorous pipeline decides more of your answer quality than the choice of LLM does. This lesson builds it on this website's own HTML and measures the effect of chunk size.
Cutting a textbook into index cards
Cards that are too big (a whole chapter) are hard to find and expensive to carry. Cards that are too small (one sentence) lose the context that made them meaningful. Good cards follow the book's own structure, one idea each, and every card is labelled with the book, chapter and page. That label is the metadata.
1. The ingestion pipeline
2. Loading: keep the structure
The loader's job is clean text plus the document's structure: headings, lists, tables, code. Throw
away navigation, footers and cookie banners. For this site, site_docs.py keeps each lesson's
<main>, drops scripts and quizzes (their wrong answer options are false statements), turns
diagrams into their text description, and splits at every <h2>:
from site_docs import parse_lesson, site_root
sections = parse_lesson(site_root() / "sql" / "06-order-limit.html", course="sql")
for s in sections:
print(f"{len(s['text'].split()):4} words {s['url']:40} {s['section']}")
print()
print(sections[3]["text"][:260], "…")
| Source type | Tooling (Python) | Pitfalls |
|---|---|---|
| HTML | BeautifulSoup, trafilatura | boilerplate, content injected by JavaScript |
| PDF with a text layer | pypdf, pdfplumber; docling or unstructured for layout | two-column pages, headers and footers on every page, hyphenation |
| Scanned PDFs and images | OCR (Tesseract, cloud OCR), or a vision model | OCR errors; cost; tables become word soup |
| Word and PowerPoint | python-docx, python-pptx, docling | speaker notes, text boxes out of order |
| Code | ast (split by function or class) | splitting in the middle of a function |
| Tables and spreadsheets | pandas | a chunk without its header row is meaningless. Often a SQL tool is better than RAG. |
3. Two ways to chunk
Fixed-size chunking cuts every N words (or tokens) with some overlap. It is simple, but it ignores meaning and cuts sentences, lists and code in half. Structure-aware chunking splits at the document's own boundaries (sections, then paragraphs, then sentences) and only packs units together up to a size limit. Compare them on the longest section on this site:
import re
from site_docs import load_sections
section = max(load_sections(courses=["sql"]), key=lambda d: len(d["text"]))
text = section["text"]
print(section["url"], "-", len(text.split()), "words")
def fixed_chunks(text, size=200, overlap=40):
words = text.split()
return [" ".join(words[i:i + size]) for i in range(0, max(1, len(words) - overlap), size - overlap)]
def units(text):
"""Paragraphs (site_docs writes one block per line); very long ones split into sentences."""
for para in text.split("\n"):
yield from ([para] if len(para.split()) <= 200 else re.split(r"(?<=[.!?])\s+", para))
def structured_chunks(text, max_words=200, overlap_units=1):
chunks, current = [], []
for u in units(text):
if current and sum(len(x.split()) for x in current) + len(u.split()) > max_words:
chunks.append("\n".join(current))
current = current[-overlap_units:] # repeat the last paragraph for context
current.append(u)
return chunks + ["\n".join(current)] if current else chunks
paragraphs = text.split("\n")
para_starts, pos = set(), 0 # word offsets where a paragraph (or code block) starts
for p in paragraphs:
para_starts.add(pos)
pos += len(p.split())
fixed = fixed_chunks(text)
fixed_cuts = sum(i * 160 not in para_starts for i in range(len(fixed))) # 160 = size - overlap
structured = structured_chunks(text)
structured_cuts = sum(c.split("\n")[0] not in paragraphs for c in structured)
for name, chunks, cuts in [("fixed 200/40", fixed, fixed_cuts), ("structured ≤200", structured, structured_cuts)]:
sizes = [len(c.split()) for c in chunks]
print(f"{name:16} {len(chunks):3} chunks, {min(sizes)}-{max(sizes)} words, "
f"{cuts} start in the middle of a paragraph or code block")
print("\nfirst fixed chunk ends with: …", fixed_chunks(text)[0][-70:])
print("first structured chunk ends with: …", structured_chunks(text)[0][-70:])
4. How big should chunks be? Measure it.
Chunk size trades precision against context. The fair comparison fixes the context
budget, meaning how many words you are willing to send to the model, and asks how often the lesson that
answers each question makes it into that budget. This runs the 32-question golden set from
mini_rag.py (lesson 10 explains it) against five chunkings of the same 164 lessons:
import numpy as np
from mini_rag import BM25, GOLDEN, LsaEmbedder, rrf
from site_docs import load_sections
sections = load_sections()
by_lesson = {}
for s in sections:
by_lesson.setdefault(s["url"].split("#")[0], []).append(s)
def fixed(size, overlap):
out = []
for url, secs in by_lesson.items():
words = " ".join(s["text"] for s in secs).split()
for i in range(0, max(1, len(words) - overlap), size - overlap):
out.append({"url": url, "title": secs[0]["title"], "section": "", "text": " ".join(words[i:i + size])})
return out
def recall_within_budget(chunks, budget_words=800):
texts = [f"{c['title']} — {c['section']}\n{c['text']}" for c in chunks] # header + text
bm25, dense = BM25(texts), LsaEmbedder(texts)
scores = {}
for mode in ("bm25", "dense", "hybrid"):
hits = 0
for question, answers in GOLDEN:
b = list(np.argsort(-bm25.scores(question))[:50])
d = list(np.argsort(-dense.scores(question))[:50])
ranking = {"bm25": b, "dense": d, "hybrid": rrf(b, d)}[mode]
sent, used = set(), 0
for i in ranking: # fill the budget in rank order
n = len(chunks[i]["text"].split())
if used + n > budget_words and sent:
break
used += n
sent.add(chunks[i]["url"].split("#")[0])
hits += bool(answers & sent)
scores[mode] = hits / len(GOLDEN)
return scores
print(f"{'strategy':18}{'chunks':>7}{'median words':>14} recall: bm25 dense hybrid")
for name, chunks in [("fixed 60 / 10", fixed(60, 10)), ("fixed 100 / 20", fixed(100, 20)),
("fixed 300 / 50", fixed(300, 50)), ("fixed 1000 / 100", fixed(1000, 100)),
("h2 sections", sections)]:
r = recall_within_budget(chunks)
med = int(np.median([len(c["text"].split()) for c in chunks]))
print(f"{name:18}{len(chunks):7}{med:14} {r['bm25']:.2f} {r['dense']:.2f} {r['hybrid']:.2f}")
On this data, big chunks lose. With 1,000-word chunks the budget holds one chunk, so one bad
ranking decision costs the whole answer. Small chunks let several candidates into the prompt. The
h2 sections sit in between: their median is under 100 words, but the few very long ones fill
the budget on their own. A size cap, like the structured chunker above, fixes that.
But small chunks have their own cost that this metric cannot see: a 60-word window often holds only
part of an answer. That tension has a well-known resolution.
Small-to-big ("parent document") retrieval
Index small chunks for precise matching, but send the model the parent they came from
(the whole section). Store parent_id in each chunk's metadata, retrieve on children, and
de-duplicate parents before building the prompt.
Contextual headers
Prefix every chunk with where it lives ("SQL › Pagination › Keyset") before embedding it. A chunk
that says "it is faster because…" then still carries what "it" is. mini_rag.doc_text()
does this.
Structure-aware chunks of a few hundred tokens at most, with a small overlap and a contextual header, are a sensible default for prose. On this site, chunks of about 100 words did best. Then run your golden set (lesson 10) and let the numbers decide. The best size depends on your documents, your questions and your context budget, which is exactly what the table above shows.
5. Metadata: the part people skip
| Field | Used for |
|---|---|
source_id, url + anchor | citations that link to the exact place; deleting a document's chunks |
title, section | contextual headers; showing sources to users |
updated_at, version | preferring the newest document; excluding superseded ones |
tenant, allowed_groups | access control at retrieval time, enforced by your code, never by the prompt |
language, doc_type, product | filters ("only the API reference", "only German") |
content_hash | incremental re-indexing (next section) |
6. Keep the index in sync, cheaply
Re-embedding everything nightly works for small collections but wastes money and time at scale. Give every chunk a stable id (for example, source URL + anchor) and a content hash. Each run then upserts what changed and deletes what vanished:
import hashlib
from site_docs import load_sections
def content_hash(text):
return hashlib.sha256(text.encode()).hexdigest()[:16]
def sync_plan(index, docs):
"""index: {chunk_id: hash} as stored in the vector DB. Returns what to upsert and delete."""
current = {d["url"]: content_hash(d["text"]) for d in docs}
upsert = [cid for cid, h in current.items() if index.get(cid) != h] # new or changed
delete = [cid for cid in index if cid not in current] # gone from the source
return upsert, delete
v1 = load_sections(courses=["sql"])
index = {d["url"]: content_hash(d["text"]) for d in v1} # first full run
v2 = [dict(d) for d in v1] # the docs change:
v2[10]["text"] += " (updated for version 2)" # one edited
v2[20]["text"] = v2[20]["text"].replace("SELECT", "select") # one reformatted
del v2[30] # one removed
v2.append({"url": "sql/29-new-lesson.html", "text": "A brand new lesson."}) # one added
upsert, delete = sync_plan(index, v2)
print(f"{len(v2)} chunks in source: re-embed {len(upsert)}, delete {len(delete)}, "
f"skip {len(v2) - len(upsert)} unchanged")
print("upsert:", upsert)
print("delete:", delete)
Deletes matter as much as upserts. An index that still contains last year's refund policy will retrieve it, and the model will quote it confidently.
Recap
- Load clean text with its structure; drop boilerplate and anything false (quiz distractors, outdated copies).
- Chunk by structure with a size cap; add contextual headers; consider small-to-big retrieval.
- Measure chunk size against a fixed context budget. On this site, huge chunks performed worst.
- Metadata powers citations, filters and access control. Hashes make re-indexing incremental.