Tokens, embeddings & the transformer
Inside the model, text goes through four steps. It is split into tokens (integers), each token becomes a vector, a stack of transformer layers lets every token gather information from the tokens before it (attention), and the last layer scores every possible next token. You do not need to train models to be an AI engineer. You do need this picture, because it explains token counts, context limits, prompt caching, and why models cannot count the r's in "strawberry".
A meeting where every word asks "who here is relevant to me?"
In the sentence "The index on orders made it fast", the word it needs to know what it refers to. In each attention layer every token sends out a question (a query). Each earlier token holds up a label (a key), and the tokens whose labels match best pass along their information (their values). After dozens of rounds, each token's vector carries the context it needs.
1. The whole pipeline in one picture
2. Tokens: the unit of everything
Models do not see characters or words. They see tokens from a fixed vocabulary, built by byte-pair encoding (BPE). Start from single characters, repeatedly merge the most frequent adjacent pair into a new symbol, and stop at the vocabulary size you want. Common words end up as one token, and rare words split into pieces. Here is the real algorithm, trained on the text of the Advanced Python course:
import re
from collections import Counter
from site_docs import load_sections
text = " ".join(d["text"] for d in load_sections(courses=["python"])).lower()
word_freq = Counter(re.findall(r"[a-z]+", text))
# every word starts as a tuple of characters, with an end-of-word marker
vocab = {tuple(w) + ("_",): n for w, n in word_freq.items()}
def best_pair(vocab):
pairs = Counter()
for symbols, n in vocab.items():
for pair in zip(symbols, symbols[1:]):
pairs[pair] += n
return max(pairs, key=pairs.get)
def apply_merge(symbols, pair):
out, i = [], 0
while i < len(symbols):
if i + 1 < len(symbols) and (symbols[i], symbols[i + 1]) == pair:
out.append(symbols[i] + symbols[i + 1])
i += 2
else:
out.append(symbols[i])
i += 1
return tuple(out)
merges = []
for _ in range(250): # real tokenizers: ~100k-200k merges
pair = best_pair(vocab)
merges.append(pair)
vocab = {apply_merge(s, pair): n for s, n in vocab.items()}
print("first merges:", ["".join(p) for p in merges[:10]])
def tokenize(word):
symbols = tuple(word) + ("_",)
for pair in merges:
symbols = apply_merge(symbols, pair)
return list(symbols)
for w in ["the", "function", "decorators", "unhashable", "strawberry"]:
print(f"{w:12} -> {tokenize(w)}")
total_chars = sum(len(w) * n for w, n in word_freq.items())
total_tokens = sum(len(tokenize(w)) * n for w, n in word_freq.items())
print(f"{total_chars / total_tokens:.1f} characters per token with only {len(merges)} merges")
Frequent words in this corpus ("the", "function") become single tokens after only 250 merges. A rarer word like "strawberry" is still in pieces. A production tokenizer has a vocabulary a thousand times larger, so English averages about 4 characters (¾ of a word) per token. Code, numbers and most non-English languages use more tokens per character.
Billing
Prices are per token, input and output separately. Count tokens, not words, when you estimate cost (lesson 03).
Limits
Context windows and max_output_tokens are measured in tokens.
Blind spots
The model sees str|aw|berry, not letters. Counting letters, reversing strings and exact arithmetic are hard for it. Give it a code tool.
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # the encoding used by recent OpenAI models
ids = enc.encode("Indexes make lookups fast.")
print(len(ids), ids, [enc.decode([i]) for i in ids])
Or skip the library: every API response reports exact usage (response.usage.input_tokens,
output_tokens), and that is what you log and bill from (Module 4).
3. Embeddings: tokens become vectors
Each token id selects a row in a learned embedding table, a vector of a few thousand numbers. During training these vectors arrange themselves so that tokens used in similar ways end up near each other. Position is added too (modern models rotate the vectors by an amount that depends on position, called RoPE), because attention on its own has no idea of word order.
These token embeddings are internal to the LLM. The text embeddings used for search (lesson 06) come from a separate embedding model that turns a whole passage into one vector. Same idea, different job.
4. Attention, computed
For each token, attention computes a query q, a key k and a value v by multiplying its vector by three learned matrices. The score between token i and token j is qi·kj, scaled by √d. A softmax turns each row of scores into weights, and token i's output is the weighted sum of the values. In a language model a causal mask hides future tokens, because you cannot attend to words that have not been generated yet.
import numpy as np
rng = np.random.default_rng(0)
tokens = ["The", "index", "made", "it", "fast"]
n, d = len(tokens), 16
X = rng.normal(size=(n, d)) # token vectors (learned in a real model)
Wq, Wk, Wv = (rng.normal(size=(d, d)) / np.sqrt(d) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv
scores = Q @ K.T / np.sqrt(d) # n x n: how relevant is token j to token i?
scores[np.triu_indices(n, k=1)] = -np.inf # causal mask: no peeking at later tokens
weights = np.exp(scores - scores.max(axis=1, keepdims=True))
weights /= weights.sum(axis=1, keepdims=True) # softmax per row: each row sums to 1
print("attention weights (row = token doing the looking):")
print(" " * 7 + "".join(f"{t:>7}" for t in tokens))
for t, row in zip(tokens, weights):
print(f"{t:>6} " + "".join(f"{w:7.2f}" for w in row))
out = weights @ V # each token: a weighted mix of earlier values
print("output shape:", out.shape, "(same as the input: one updated vector per token)")
The weights here are random because the matrices are random. In a trained model the row for "it" would put high weight on "index". Real models run many such heads in parallel (each learns a different kind of relationship) and stack dozens of layers. The shape of the computation is exactly this.
5. The context window and the KV cache
The context window is the most tokens the model can attend over at once: prompt plus output. Anything outside it does not exist for the model. Windows are now hundreds of thousands of tokens and more, but every token in the window costs money and time on every call.
When generating, each new token needs the keys and values of all earlier tokens. Recomputing them at every step would be wasteful, so inference servers keep them in memory: the KV cache.
prompt_tokens, output_tokens = 2_000, 500
# K/V projections computed while generating the answer:
without_cache = sum(prompt_tokens + step for step in range(1, output_tokens + 1)) # redo everything each step
with_cache = prompt_tokens + output_tokens # each token once
print(f"without KV cache: {without_cache:,} token-projections")
print(f"with KV cache: {with_cache:,} token-projections ({without_cache / with_cache:.0f}x fewer)")
Prompt caching: the KV cache across requests
Providers keep the KV cache of a prompt prefix for a few minutes. If your next request starts with the same tokens (same system prompt, same tool list, same documents), those tokens are billed at a steep discount and processed faster. Put stable content first and variable content last.
Long context is not perfect recall
Models retrieve facts from the start and end of a long prompt more reliably than from the middle ("lost in the middle"). A 200k-token window does not make retrieval obsolete: sending the 5 relevant passages is cheaper, faster and often more accurate than sending all 500.
6. From logits to a token
The last layer produces one score (a logit) per vocabulary entry for the next position. A softmax turns logits into probabilities, and a sampling rule picks one token. Temperature and top-p, the knobs you set in the API, act exactly here. That is the next lesson.
Recap
- Text → tokens (BPE) → vectors → N × (attention + MLP) → logits.
- Tokens drive cost and limits: about 4 characters per token for English, more for code and other languages.
- Attention lets each token pull in information from earlier tokens through query, key and value vectors.
- The KV cache makes generation affordable, and prompt caching reuses it across requests, so put stable text first.