Module 1 · How LLMs work

Sampling, latency & cost

Foundations 16 min read The three numbers every LLM feature is judged on

The model hands back a probability for every possible next token. Sampling decides which one is used, and it is why the same prompt can give different answers. Latency is how long the user waits: the time to the first token, then one step per output token. Cost is tokens in plus tokens out, each at its own price. Engineers who can estimate all three before writing code build features that survive contact with production.

🎲

A weighted die, re-carved for every word

For each next word the model carves a die where likely words get big faces and unlikely words get tiny ones. Temperature sharpens or flattens the faces. Top-p shaves off the tiny faces so absurd words can never come up. Temperature 0 means "always take the biggest face".

1. Temperature: sharpen or flatten

The model outputs logits (raw scores). Softmax turns them into probabilities. Dividing the logits by a temperature T before the softmax changes how peaked the distribution is:

import numpy as np

tokens = ["key", "keys", "index", "column", "table", "banana"]
logits = np.array([6.0, 4.2, 3.9, 3.1, 2.8, -1.0])      # "The table needs a primary ___"

def softmax(z):
    e = np.exp(z - z.max())                  # subtract the max for numerical stability
    return e / e.sum()

for T in (0.2, 0.7, 1.0, 1.5):
    p = softmax(logits / T)
    print(f"T={T:<4}" + "".join(f"{t:>8}:{x:.3f}" for t, x in zip(tokens, p)))
T=0.2 key:1.000 keys:0.000 index:0.000 column:0.000 table:0.000 banana:0.000 T=0.7 key:0.868 keys:0.066 index:0.043 column:0.014 table:0.009 banana:0.000 T=1.0 key:0.722 keys:0.119 index:0.088 column:0.040 table:0.029 banana:0.001 T=1.5 key:0.549 keys:0.165 index:0.135 column:0.079 table:0.065 banana:0.005

Low (0–0.3)

Nearly always the top token. Use it for extraction, classification, SQL and code, anything with one right answer.

Medium (0.7–1.0)

The default. Natural-sounding text with some variety.

High (>1.2)

Creative, then incoherent. Unlikely tokens gain real probability. Rarely what you want.

2. Top-p and top-k: cut off the tail

Top-k keeps only the k most likely tokens. Top-p (nucleus sampling) keeps the smallest set of tokens whose probabilities add up to p, which adapts to how confident the model is. Everything else gets probability zero, so a one-in-a-thousand "banana" can never appear.

import numpy as np
from collections import Counter

tokens = np.array(["key", "keys", "index", "column", "table", "banana"])
logits = np.array([6.0, 4.2, 3.9, 3.1, 2.8, -1.0])
softmax = lambda z: np.exp(z - z.max()) / np.exp(z - z.max()).sum()

def top_p_filter(p, top_p):
    order = np.argsort(-p)                               # most likely first
    cumulative = np.cumsum(p[order])
    keep = order[: np.searchsorted(cumulative, top_p) + 1]
    q = np.zeros_like(p)
    q[keep] = p[keep]
    return q / q.sum()                                   # renormalise the survivors

rng = np.random.default_rng(42)
p = softmax(logits / 1.5)                                # a high temperature makes the tail visible
for label, dist in [("T=1.5", p), ("T=1.5, top_p=0.9", top_p_filter(p, 0.9))]:
    draws = Counter(str(t) for t in rng.choice(tokens, size=1000, p=dist))
    print(f"{label:17}", dict(sorted(draws.items(), key=lambda kv: -kv[1])))

print("greedy (T=0):     always", tokens[np.argmax(logits)])
T=1.5 {'key': 545, 'keys': 163, 'index': 142, 'column': 84, 'table': 61, 'banana': 5} T=1.5, top_p=0.9 {'key': 587, 'keys': 166, 'index': 151, 'column': 96} greedy (T=0): always key
Determinism is not guaranteed

Temperature 0 is close to deterministic, but hosted models can still vary slightly between calls (batching and floating-point order on GPUs). Many reasoning models also ignore or reject temperature and top_p; you steer them with a reasoning-effort setting instead. Never build a test that expects identical text twice. Test properties: the JSON validates, the answer cites a source, the number matches the database (lesson 10).

3. Where the time goes

network+queue prefill whole prompt, in parallel … one step per output token (decode) time to first token (TTFT) grows with prompt length total ≈ TTFT + output tokens ÷ tokens-per-second Streaming (OpenAI SDK lesson 05) shows tokens as they arrive: perceived wait ≈ TTFT, not the total.
def latency(input_tokens, output_tokens, prefill_tps=5_000, decode_tps=90, overhead_s=0.35):
    ttft = overhead_s + input_tokens / prefill_tps
    return ttft, ttft + output_tokens / decode_tps

cases = [("classify a review", 300, 3), ("RAG answer", 4_000, 300),
         ("long report", 2_000, 2_500), ("stuff a 100k-token context", 100_000, 300)]
print(f"{'feature':28}{'TTFT':>8}{'total':>9}")
for name, i, o in cases:
    ttft, total = latency(i, o)
    print(f"{name:28}{ttft:7.2f}s{total:8.2f}s")
feature TTFT total classify a review 0.41s 0.44s RAG answer 1.15s 4.48s long report 0.75s 28.53s stuff a 100k-token context 20.35s 23.68s

Two lessons hide in that table. Output length dominates total time: the report is slow because of its 2,500 output tokens, not its prompt. Huge prompts hurt TTFT: stuffing 100k tokens delays the first token by seconds on every request. Retrieval that sends 4k relevant tokens instead is the cure (Module 2). Reasoning models add a third factor: their hidden thinking tokens are output tokens too, so a high reasoning effort can add many seconds before the visible answer starts.

4. Where the money goes

cost = input_tokens × input_price + output_tokens × output_price, with prices per million tokens. Output is typically several times more expensive than input, and cached input is discounted.

# USD per 1M tokens. gpt-5.6-luna is the real price when this was written (OpenAI SDK lesson 04);
# the second model is a HYPOTHETICAL 10x-priced flagship, to show the shape of the trade-off.
PRICES = {"gpt-5.6-luna": (0.20, 1.20), "flagship (hypothetical 10x)": (2.00, 12.00)}
CACHED_DISCOUNT = 0.90        # assumption: cached input billed at 10% of the input price

def request_cost(model, input_tokens, output_tokens, cached_tokens=0):
    inp, out = PRICES[model]
    billed_input = (input_tokens - cached_tokens) + cached_tokens * (1 - CACHED_DISCOUNT)
    return (billed_input * inp + output_tokens * out) / 1_000_000

per_day = 50_000
for model in PRICES:
    plain = request_cost(model, 4_000, 300)
    cached = request_cost(model, 4_000, 300, cached_tokens=1_500)   # stable system prompt + tools
    print(f"{model:28} ${plain:.5f}/req  ${plain * per_day * 30:>9,.0f}/month"
          f"  with caching ${cached * per_day * 30:>9,.0f}/month")
gpt-5.6-luna $0.00116/req $ 1,740/month with caching $ 1,335/month flagship (hypothetical 10x) $0.01160/req $ 17,400/month with caching $ 13,350/month
LeverSavesCost of pulling it
Send less context (better retrieval, trimmed history)input tokens and TTFTengineering time, and recall must stay high
Stable prefix for prompt cachinginput cost and TTFTnone: just order your prompt well
Cap output (max_output_tokens, "answer in 3 bullet points")output cost and total latencyrisk of truncated answers: check the finish status
Smaller model, or route easy requests to it (lesson 20)often 5–20×quality must be proven by evals first
Lower reasoning effortthinking tokens and latencyharder tasks may get worse
Batch API for offline jobsabout 50% (per OpenAI's batch pricing)results within hours, not seconds
Cache whole answers for repeated questions (lesson 20)the entire callstaleness, and cache-key design

5. Throughput and rate limits

Providers limit requests per minute (RPM) and tokens per minute (TPM) per model and account tier. A batch job that fires 200 concurrent requests will hit 429 errors. Handle them with bounded concurrency (a semaphore) and retries with exponential backoff, both covered in the OpenAI SDK course (lessons 11–12). For capacity planning: peak requests per minute × average tokens per request must stay under your TPM limit, with headroom.

Recap

  • Temperature reshapes the distribution; top-p/top-k cut its tail. Use low temperature when there is one right answer.
  • Latency = TTFT + output tokens ÷ speed. Long outputs and huge prompts are the usual culprits. Streaming hides the rest.
  • Cost = tokens × price, with output pricier and cached input cheaper. Estimate before you build.
  • Never assert exact text in tests. Sampling and server nondeterminism make it flaky.

Checkpoint

1 · A text-to-SQL feature sometimes returns different queries for the same question. First fix?
Low temperature concentrates probability on the most likely tokens. Validation catches what randomness is left, and model mistakes too.
2 · Users complain a summary feature "hangs" for 20 seconds. The prompt is 1,500 tokens and the answer 1,800. Most effective change?
Decoding 1,800 tokens dominates. Streaming makes the wait feel like the TTFT, and a shorter output cuts the real total.
3 · Top-p = 0.9 means…
Nucleus sampling adapts: when the model is confident the nucleus may be one token; when unsure, many.