What an LLM actually is
A large language model is a function with one job: given some text, predict the next token (a word or a piece of a word). It returns a probability for every token in its vocabulary. The program that calls it picks one, appends it to the text and asks again. Repeat that a few hundred times and you get an essay, a SQL query or a tool call. Everything in this course (RAG, agents, MCP, observability) is engineering around that loop.
Your phone's autocomplete, after reading a library
Your phone suggests the next word from what you usually type. Now imagine an autocomplete that has read a good fraction of the public internet, books and code, and has billions of adjustable dials to capture patterns far longer than "the next word after good". It still only suggests what comes next. It has no database of facts to look things up in. It has patterns that often, but not always, line up with facts.
1. The generation loop
Three consequences follow directly from this picture, and they explain most LLM behaviour you will see:
Output costs time
Each output token needs its own pass through the model, so a 500-token answer takes roughly 500 steps. Output is slower and pricier than input.
No memory between calls
The model is a pure function of its input. A "conversation" is your code resending the whole history each time.
Plausible, not true
Training rewards the likely next token. Likely and correct usually coincide, but nothing in the loop checks facts.
2. Build a tiny language model
The smallest possible language model counts which word follows which (a bigram model). Real LLMs are far more capable, but the interface is identical: context in, next-token probabilities out. Here it is, trained on the text of this site's SQL course:
import random
import re
from collections import Counter, defaultdict
from site_docs import load_sections # the lessons of this website as plain text
text = " ".join(d["text"] for d in load_sections(courses=["sql"])).lower()
words = re.findall(r"[a-z_']+|[.,]", text)
follows = defaultdict(Counter) # word -> Counter of the words that came next
for current, nxt in zip(words, words[1:]):
follows[current][nxt] += 1
def next_word_probs(word, k=4):
counts = follows[word]
total = sum(counts.values())
return [(w, round(n / total, 2)) for w, n in counts.most_common(k)]
print(f"trained on {len(words):,} words, vocabulary {len(follows):,}")
print("after 'primary':", next_word_probs("primary"))
print("after 'group': ", next_word_probs("group"))
print("after 'left': ", next_word_probs("left"))
random.seed(3)
def generate(word, n=14):
out = [word]
for _ in range(n):
counts = follows[out[-1]]
out.append(random.choices(list(counts), weights=list(counts.values()))[0])
return " ".join(out)
for start in ["the", "an", "indexes"]:
print(">", generate(start))
The probabilities are real statistics: "primary" is almost always followed by "key", and "group" by "by". Every adjacent pair in the generated text really occurs in the course. Yet the sentences as a whole say nothing true, and some say nothing at all. That is hallucination in miniature. The model continues text plausibly and has no notion of whether the result is correct. A real LLM conditions on thousands of previous tokens instead of one, which is why its output is coherent over pages. Coherent is still not the same as verified.
3. How a model is made: three training stages
| Term | Plain meaning | Why an engineer cares |
|---|---|---|
| Parameters (weights) | The numbers learned in training; the model's entire "knowledge" | More parameters usually means more capability, and more cost and latency per token |
| Base model | Output of pretraining; continues any text | Rarely used directly. You call instruction-tuned models. |
| Reasoning model | Trained to produce hidden "thinking" tokens before answering | Better at multi-step problems. You pay for the thinking tokens, and it is slower (lesson 03). |
| Knowledge cut-off | The date the training data ends | Anything newer, or private to your company, must be supplied in the prompt (RAG, lesson 05) |
| Fine-tuning | Extra training on your examples | Teaches format and style well, and facts poorly. Usually try prompting and RAG first. |
4. Hallucination: why, and what you do about it
A hallucination is fluent output that is false or unsupported: an invented function name, a citation to a paper that does not exist, a confident wrong number. It is not a bug that will be patched away. It follows from the objective. The model is trained to produce likely text, and when it lacks the facts, a likely-sounding answer is still likely text.
Makes it worse
Questions about recent or private facts; long exact values (IDs, prices); a false premise in the question ("why did Postgres drop JOIN support?"); asking for citations it cannot see.
Makes it better
Grounding: put the source text in the prompt (RAG). Tools: let it query the real system. Ask for citations to the supplied text, allow "I don't know", and measure with evals (lesson 10).
Treat model output like user input: untrusted until checked. Validate structure with a schema, check facts against a source, run generated SQL through a read-only gateway (see the OpenAI SDK course project), and never execute generated code with more permissions than the least trusted user has.
5. What LLMs are good and bad at
| Strong | Weak (without help) | The help |
|---|---|---|
| Summarising, rewriting, translating | Facts after the cut-off, private data | RAG, tools |
| Extracting structure from messy text | Exact arithmetic on big numbers | A calculator or code tool |
| Writing and explaining code | Counting letters or words | Code: len() (tokens hide letters, lesson 02) |
| Classifying intent, routing | Remembering past conversations | Your app stores and resends history |
| Planning steps with tools | Knowing when it is wrong | Evals, validators, human review |
6. Where models come from
Hosted APIs
OpenAI (the GPT-5.6 family used in the OpenAI SDK course), Anthropic, Google and others. The best quality with nothing to run, and you pay per token. Your data goes to the provider under their terms.
Open-weight models
Families such as Llama, Mistral, Qwen, DeepSeek and gpt-oss. You download the weights and run them with Ollama or llama.cpp on a laptop, or vLLM on a GPU server. The data stays with you, but you operate it.
The skills in this course do not depend on the vendor. Retrieval, tool use, MCP, tracing and evaluation
work the same way whichever model sits in the middle. The Python examples use the OpenAI SDK because most
providers and local servers (Ollama and vLLM included) offer an OpenAI-compatible endpoint. Change
base_url and the code keeps working.
Non-technical summary, for explaining this to a manager
An LLM is a very well-read autocomplete. It writes convincing text about almost anything, but it does not look things up and cannot tell you when it is guessing. We make it reliable the way we make any clever-but-unreliable component reliable. We give it the right documents (retrieval), let it use real systems through controlled tools, check its output automatically, and monitor it in production. Those four things are what "AI engineering" means, and they are the four modules that follow.
Recap
- An LLM predicts the next token. Generation is a loop that appends one token at a time.
- Knowledge comes from pretraining and stops at the cut-off. Later stages shape behaviour.
- It is stateless. Memory is whatever your code puts back in the prompt.
- Hallucination is built in, so ground the model with sources and tools, validate its output, and measure it.