Build a RAG pipeline in Python
Time to connect the pieces: retrieve, build a prompt with numbered sources, call the model, and return an answer whose citations are checked against what was actually retrieved. Building it without a framework once is the fastest way to understand what LangChain and LlamaIndex do for you, and to debug them when they misbehave. It is about 60 lines.
Cook it from scratch once, then buy the sauce
A jar of pasta sauce is convenient, but if you have never made one you cannot tell why this jar tastes burnt. Build the pipeline by hand, and framework abstractions become shortcuts you understand instead of magic you fear.
1. The request path
2. The pipeline, no framework
The retrieval, prompt building and citation checking below are real. The model is the course's offline
stand-in (a small function that answers with the first prose sentence of source 1), plugged into the real
OpenAI SDK. With OpenAI() in its place nothing else changes.
import re
from dataclasses import dataclass, field
from html import escape
from fake_openai import fake_client
from mini_rag import HybridIndex
from site_docs import COURSE_NAMES, load_sections
INSTRUCTIONS = """You answer questions about the course lessons in the <source> tags.
- Use only the sources. Cite every claim inline as [1], [2].
- If the sources do not answer the question, reply exactly: I don't know.
- At most 4 sentences."""
PROMPT_VERSION = "answer-v3"
@dataclass
class Answer:
text: str
sources: list = field(default_factory=list)
usage: dict = field(default_factory=dict)
grounded: bool = True
class RagPipeline:
def __init__(self, client, index, model="gpt-5.6-luna", budget_words=700):
self.client, self.index, self.model, self.budget = client, index, model, budget_words
def retrieve(self, question):
chosen, seen, used = [], set(), 0
for hit in self.index.search(question, k=20):
lesson = hit.doc["url"].split("#")[0]
words = len(hit.doc["text"].split())
if lesson in seen or (chosen and used + words > self.budget):
continue # one chunk per lesson, within budget
chosen.append(hit.doc)
seen.add(lesson)
used += words
return chosen
@staticmethod
def build_input(question, docs):
blocks = [f'<source id="{i}" title="{escape(COURSE_NAMES[d["course"]])} › {escape(d["title"])}">\n'
f'{escape(d["text"], quote=False)}\n</source>' for i, d in enumerate(docs, 1)]
return "\n".join(blocks) + f"\n<question>{escape(question, quote=False)}</question>"
def answer(self, question):
docs = self.retrieve(question)
if not docs:
return Answer("I don't know.", grounded=False)
r = self.client.responses.create(model=self.model, instructions=INSTRUCTIONS,
input=self.build_input(question, docs),
metadata={"prompt_version": PROMPT_VERSION})
cited = sorted({int(n) for n in re.findall(r"\[(\d+)\]", r.output_text)})
valid = [n for n in cited if 1 <= n <= len(docs)]
usage = {"input": r.usage.input_tokens, "output": r.usage.output_tokens}
if not valid or len(valid) != len(cited): # no citation, or a made-up one
return Answer("I don't know.", usage=usage, grounded=False)
return Answer(r.output_text, [docs[n - 1] for n in valid], usage)
def stand_in_model(body):
"""Offline stand-in for the LLM: quotes the first prose sentence of source 1, cited."""
m = re.search(r'<source id="1"[^>]*>\n(.*?)\n</source>', body["input"], re.S)
for line in (m.group(1).splitlines() if m else []):
sentence = re.match(r"([A-Z][^.!?]{20,}[.!?])(\s|$)", line) # skips code and diagrams
if sentence:
return f"{sentence.group(1)} [1]"
return "I don't know."
rag = RagPipeline(fake_client(responder=stand_in_model), HybridIndex(load_sections()))
for q in ["What is a covering index?", "What does the GIL do?", "When should I use a pandas UDF?",
"What does the KV cache store?"]:
a = rag.answer(q)
print(f"Q: {q}\nA: {a.text}")
print(" sources:", [s["url"] for s in a.sources], "| tokens:", a.usage)
The last question is a trap. The KV cache is explained in this course, which is not in the indexed collection, yet retrieval still returned its best guess: a section on caching DataFrames in PySpark. Retrieval always returns something. Here the stand-in found no sentence to quote and answered "I don't know", so the pipeline returned no sources. A real model has to make that call by judging relevance, and sometimes it will happily explain Spark caching instead. That is the failure faithfulness and relevance evaluation catch (lesson 10), and why a reranker's relevance threshold is a useful second guard.
3. Never trust a citation you did not check
Models occasionally cite a source number that does not exist, or cite nothing at all. The pipeline treats both as "not grounded" rather than passing them to the user. Here the model's reply is scripted to cite a source that was never provided:
import re
from fake_openai import fake_client
docs = [{"url": "sql/26-indexes.html#covering", "text": "A covering index contains every column the query needs."}]
client = fake_client(script=["Covering indexes avoid table lookups [1]. They also compress data by 90% [4]."])
reply = client.responses.create(model="gpt-5.6-luna", input="...").output_text
cited = sorted({int(n) for n in re.findall(r"\[(\d+)\]", reply)})
invalid = [n for n in cited if not 1 <= n <= len(docs)]
print("reply: ", reply)
print("cited: ", cited, "| invalid:", invalid)
print("verdict:", "reject, not grounded" if invalid else "ok")
Some teams drop only the invalid sentence, some retry once with a stricter prompt, some show the answer with a warning. Whatever you choose, count it. The rate of ungrounded answers is one of the most useful RAG health metrics (Module 4).
4. What the 60 lines leave out
| Concern | Production answer | Lesson |
|---|---|---|
| Latency | stream the answer; run retrieval and any rewrite concurrently where possible; cache embeddings of frequent queries | 03, OpenAI SDK 05 |
| Follow-up questions | rewrite "and in MySQL?" into a standalone question using the chat history before retrieving | 08 |
| Access control | a metadata filter built from the user's identity on every search | 07, 20 |
| Failures | timeouts and retries on the model call; a clear message when the index is unavailable | OpenAI SDK 12 |
| Observability | a trace per request: retrieval hits, prompt version, tokens, cost, grounded flag | 16–18 |
| Quality | a golden set, run on every change | 10 |
5. The same pipeline with a framework
Frameworks give you loaders, splitters, vector-store adapters and model wrappers behind common interfaces. These snippets follow the current documented APIs but were not run here, because the packages are not installed on this site's machine.
from langchain.chat_models import init_chat_model
from langchain_core.documents import Document
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_openai import OpenAIEmbeddings
from langchain_text_splitters import RecursiveCharacterTextSplitter
documents = [Document(page_content=d["text"], metadata={"url": d["url"]}) for d in load_sections()]
chunks = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=100).split_documents(documents)
store = InMemoryVectorStore(OpenAIEmbeddings(model="text-embedding-3-small"))
store.add_documents(chunks)
llm = init_chat_model("openai:gpt-5.6-luna")
def answer(question):
hits = store.similarity_search(question, k=4)
sources = "\n".join(f'<source id="{i}">{h.page_content}</source>' for i, h in enumerate(hits, 1))
reply = llm.invoke([("system", INSTRUCTIONS), ("user", f"{sources}\n<question>{question}</question>")])
return reply.content, [h.metadata["url"] for h in hits]
from llama_index.core import Document, Settings, VectorStoreIndex
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI
Settings.llm = OpenAI(model="gpt-5.6-luna") # a new model name may need an up-to-date integration package
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
index = VectorStoreIndex.from_documents(
[Document(text=d["text"], metadata={"url": d["url"]}) for d in load_sections()])
engine = index.as_query_engine(similarity_top_k=4)
response = engine.query("Why should I avoid OFFSET for deep pagination?")
print(response)
print([node.metadata["url"] for node in response.source_nodes])
| Choose a framework when… | Stay framework-free when… |
|---|---|
| you need many loaders (PDF, Notion, Confluence…) and store adapters quickly | the pipeline is small and you want every token of the prompt under your control |
| you are prototyping several retrieval strategies | you must debug latency and cost precisely |
| the team already knows it | you want few dependencies and a stable upgrade path |
A common path in practice: prototype with a framework, then replace the parts that matter (retrieval, prompt, citation checks) with your own code as the product matures, keeping the framework's loaders. The course project (lesson 21) is framework-free for the same reason as this lesson.
Recap
- Retrieve wide, select within a budget, number the sources, escape them.
- Verify citations against what was retrieved. Count ungrounded answers.
- Retrieval always returns something, so relevance judgement needs the model, a reranker threshold, or both.
- Frameworks save plumbing time. Knowing the 60-line version lets you debug them.