Module 1 · How LLMs work

What an LLM actually is

Foundations 18 min read The one idea everything else builds on

A large language model is a function with one job: given some text, predict the next token (a word or a piece of a word). It returns a probability for every token in its vocabulary. The program that calls it picks one, appends it to the text and asks again. Repeat that a few hundred times and you get an essay, a SQL query or a tool call. Everything in this course (RAG, agents, MCP, observability) is engineering around that loop.

📱

Your phone's autocomplete, after reading a library

Your phone suggests the next word from what you usually type. Now imagine an autocomplete that has read a good fraction of the public internet, books and code, and has billions of adjustable dials to capture patterns far longer than "the next word after good". It still only suggests what comes next. It has no database of facts to look things up in. It has patterns that often, but not always, line up with facts.

1. The generation loop

text so far (the context) "An index speeds up" split into tokens first (lesson 02) the model billions of weights one forward pass same input → same output next-token probabilities " lookups" " queries" " reads" " bananas" pick one token (sampling, lesson 03), append it, run again It stops when it emits a special end token or hits the max_output_tokens limit you set.

Three consequences follow directly from this picture, and they explain most LLM behaviour you will see:

Output costs time

Each output token needs its own pass through the model, so a 500-token answer takes roughly 500 steps. Output is slower and pricier than input.

No memory between calls

The model is a pure function of its input. A "conversation" is your code resending the whole history each time.

Plausible, not true

Training rewards the likely next token. Likely and correct usually coincide, but nothing in the loop checks facts.

2. Build a tiny language model

The smallest possible language model counts which word follows which (a bigram model). Real LLMs are far more capable, but the interface is identical: context in, next-token probabilities out. Here it is, trained on the text of this site's SQL course:

import random
import re
from collections import Counter, defaultdict

from site_docs import load_sections          # the lessons of this website as plain text

text = " ".join(d["text"] for d in load_sections(courses=["sql"])).lower()
words = re.findall(r"[a-z_']+|[.,]", text)

follows = defaultdict(Counter)              # word -> Counter of the words that came next
for current, nxt in zip(words, words[1:]):
    follows[current][nxt] += 1

def next_word_probs(word, k=4):
    counts = follows[word]
    total = sum(counts.values())
    return [(w, round(n / total, 2)) for w, n in counts.most_common(k)]

print(f"trained on {len(words):,} words, vocabulary {len(follows):,}")
print("after 'primary':", next_word_probs("primary"))
print("after 'group':  ", next_word_probs("group"))
print("after 'left':   ", next_word_probs("left"))

random.seed(3)
def generate(word, n=14):
    out = [word]
    for _ in range(n):
        counts = follows[out[-1]]
        out.append(random.choices(list(counts), weights=list(counts.values()))[0])
    return " ".join(out)

for start in ["the", "an", "indexes"]:
    print(">", generate(start))
trained on 26,993 words, vocabulary 2,869 after 'primary': [('key', 0.89), ('keys', 0.11)] after 'group': [('by', 0.83), (',', 0.02), ('totals', 0.01), ('at', 0.01)] after 'left': [('join', 0.74), ('rows', 0.07), ('aligned', 0.04), (',', 0.04)] > the result rows . unit_price from orders in memory . store results you stand on > an anti join , instr email primary key , s . . the product has > indexes avoid replace into order_items oi . not mysql default and then else changed histories

The probabilities are real statistics: "primary" is almost always followed by "key", and "group" by "by". Every adjacent pair in the generated text really occurs in the course. Yet the sentences as a whole say nothing true, and some say nothing at all. That is hallucination in miniature. The model continues text plausibly and has no notion of whether the result is correct. A real LLM conditions on thousands of previous tokens instead of one, which is why its output is coherent over pages. Coherent is still not the same as verified.

3. How a model is made: three training stages

1 · Pretraining predict the next token on trillions of tokens of web, books and code months · thousands of GPUs → a base model 2 · Instruction tuning (SFT) same objective, on curated prompt → ideal answer pairs (chat format, tool calls) days · far less data → follows instructions 3 · Preference & RL RLHF / DPO: prefer answers people rate higher; RL on checkable tasks (maths, code) shapes tone, safety, and "reasoning" behaviour Knowledge comes almost entirely from stage 1, which ends at the knowledge cut-off. Stages 2 and 3 change behaviour: format, helpfulness, refusals, step-by-step thinking.
TermPlain meaningWhy an engineer cares
Parameters (weights)The numbers learned in training; the model's entire "knowledge"More parameters usually means more capability, and more cost and latency per token
Base modelOutput of pretraining; continues any textRarely used directly. You call instruction-tuned models.
Reasoning modelTrained to produce hidden "thinking" tokens before answeringBetter at multi-step problems. You pay for the thinking tokens, and it is slower (lesson 03).
Knowledge cut-offThe date the training data endsAnything newer, or private to your company, must be supplied in the prompt (RAG, lesson 05)
Fine-tuningExtra training on your examplesTeaches format and style well, and facts poorly. Usually try prompting and RAG first.

4. Hallucination: why, and what you do about it

A hallucination is fluent output that is false or unsupported: an invented function name, a citation to a paper that does not exist, a confident wrong number. It is not a bug that will be patched away. It follows from the objective. The model is trained to produce likely text, and when it lacks the facts, a likely-sounding answer is still likely text.

Makes it worse

Questions about recent or private facts; long exact values (IDs, prices); a false premise in the question ("why did Postgres drop JOIN support?"); asking for citations it cannot see.

Makes it better

Grounding: put the source text in the prompt (RAG). Tools: let it query the real system. Ask for citations to the supplied text, allow "I don't know", and measure with evals (lesson 10).

Engineering rule

Treat model output like user input: untrusted until checked. Validate structure with a schema, check facts against a source, run generated SQL through a read-only gateway (see the OpenAI SDK course project), and never execute generated code with more permissions than the least trusted user has.

5. What LLMs are good and bad at

StrongWeak (without help)The help
Summarising, rewriting, translatingFacts after the cut-off, private dataRAG, tools
Extracting structure from messy textExact arithmetic on big numbersA calculator or code tool
Writing and explaining codeCounting letters or wordsCode: len() (tokens hide letters, lesson 02)
Classifying intent, routingRemembering past conversationsYour app stores and resends history
Planning steps with toolsKnowing when it is wrongEvals, validators, human review

6. Where models come from

Hosted APIs

OpenAI (the GPT-5.6 family used in the OpenAI SDK course), Anthropic, Google and others. The best quality with nothing to run, and you pay per token. Your data goes to the provider under their terms.

Open-weight models

Families such as Llama, Mistral, Qwen, DeepSeek and gpt-oss. You download the weights and run them with Ollama or llama.cpp on a laptop, or vLLM on a GPU server. The data stays with you, but you operate it.

The skills in this course do not depend on the vendor. Retrieval, tool use, MCP, tracing and evaluation work the same way whichever model sits in the middle. The Python examples use the OpenAI SDK because most providers and local servers (Ollama and vLLM included) offer an OpenAI-compatible endpoint. Change base_url and the code keeps working.

Non-technical summary, for explaining this to a manager

An LLM is a very well-read autocomplete. It writes convincing text about almost anything, but it does not look things up and cannot tell you when it is guessing. We make it reliable the way we make any clever-but-unreliable component reliable. We give it the right documents (retrieval), let it use real systems through controlled tools, check its output automatically, and monitor it in production. Those four things are what "AI engineering" means, and they are the four modules that follow.

Recap

  • An LLM predicts the next token. Generation is a loop that appends one token at a time.
  • Knowledge comes from pretraining and stops at the cut-off. Later stages shape behaviour.
  • It is stateless. Memory is whatever your code puts back in the prompt.
  • Hallucination is built in, so ground the model with sources and tools, validate its output, and measure it.

Checkpoint

1 · A user says "you told me yesterday that…". What actually happened?
The model is a pure function of its input. Chat memory is a feature of the application: it saves history and includes it in the next prompt.
2 · Which is the best first step to make a model answer questions about your company's internal wiki?
Fine-tuning is poor at adding facts and goes stale as the wiki changes. Retrieval supplies current text the model can quote and cite.
3 · Why does a 1,000-token answer take much longer than reading a 1,000-token prompt?
Prefill handles the whole prompt at once. Decoding is sequential because each new token depends on the previous one (more in lesson 03).