Sampling, latency & cost
The model hands back a probability for every possible next token. Sampling decides which one is used, and it is why the same prompt can give different answers. Latency is how long the user waits: the time to the first token, then one step per output token. Cost is tokens in plus tokens out, each at its own price. Engineers who can estimate all three before writing code build features that survive contact with production.
A weighted die, re-carved for every word
For each next word the model carves a die where likely words get big faces and unlikely words get tiny ones. Temperature sharpens or flattens the faces. Top-p shaves off the tiny faces so absurd words can never come up. Temperature 0 means "always take the biggest face".
1. Temperature: sharpen or flatten
The model outputs logits (raw scores). Softmax turns them into probabilities. Dividing the logits by a temperature T before the softmax changes how peaked the distribution is:
import numpy as np
tokens = ["key", "keys", "index", "column", "table", "banana"]
logits = np.array([6.0, 4.2, 3.9, 3.1, 2.8, -1.0]) # "The table needs a primary ___"
def softmax(z):
e = np.exp(z - z.max()) # subtract the max for numerical stability
return e / e.sum()
for T in (0.2, 0.7, 1.0, 1.5):
p = softmax(logits / T)
print(f"T={T:<4}" + "".join(f"{t:>8}:{x:.3f}" for t, x in zip(tokens, p)))
Low (0–0.3)
Nearly always the top token. Use it for extraction, classification, SQL and code, anything with one right answer.
Medium (0.7–1.0)
The default. Natural-sounding text with some variety.
High (>1.2)
Creative, then incoherent. Unlikely tokens gain real probability. Rarely what you want.
2. Top-p and top-k: cut off the tail
Top-k keeps only the k most likely tokens. Top-p (nucleus sampling) keeps the smallest set of tokens whose probabilities add up to p, which adapts to how confident the model is. Everything else gets probability zero, so a one-in-a-thousand "banana" can never appear.
import numpy as np
from collections import Counter
tokens = np.array(["key", "keys", "index", "column", "table", "banana"])
logits = np.array([6.0, 4.2, 3.9, 3.1, 2.8, -1.0])
softmax = lambda z: np.exp(z - z.max()) / np.exp(z - z.max()).sum()
def top_p_filter(p, top_p):
order = np.argsort(-p) # most likely first
cumulative = np.cumsum(p[order])
keep = order[: np.searchsorted(cumulative, top_p) + 1]
q = np.zeros_like(p)
q[keep] = p[keep]
return q / q.sum() # renormalise the survivors
rng = np.random.default_rng(42)
p = softmax(logits / 1.5) # a high temperature makes the tail visible
for label, dist in [("T=1.5", p), ("T=1.5, top_p=0.9", top_p_filter(p, 0.9))]:
draws = Counter(str(t) for t in rng.choice(tokens, size=1000, p=dist))
print(f"{label:17}", dict(sorted(draws.items(), key=lambda kv: -kv[1])))
print("greedy (T=0): always", tokens[np.argmax(logits)])
Temperature 0 is close to deterministic, but hosted models can still vary slightly between calls
(batching and floating-point order on GPUs). Many reasoning models also ignore or reject
temperature and top_p; you steer them with a reasoning-effort setting instead.
Never build a test that expects identical text twice. Test properties: the JSON validates, the answer
cites a source, the number matches the database (lesson 10).
3. Where the time goes
def latency(input_tokens, output_tokens, prefill_tps=5_000, decode_tps=90, overhead_s=0.35):
ttft = overhead_s + input_tokens / prefill_tps
return ttft, ttft + output_tokens / decode_tps
cases = [("classify a review", 300, 3), ("RAG answer", 4_000, 300),
("long report", 2_000, 2_500), ("stuff a 100k-token context", 100_000, 300)]
print(f"{'feature':28}{'TTFT':>8}{'total':>9}")
for name, i, o in cases:
ttft, total = latency(i, o)
print(f"{name:28}{ttft:7.2f}s{total:8.2f}s")
Two lessons hide in that table. Output length dominates total time: the report is slow because of its 2,500 output tokens, not its prompt. Huge prompts hurt TTFT: stuffing 100k tokens delays the first token by seconds on every request. Retrieval that sends 4k relevant tokens instead is the cure (Module 2). Reasoning models add a third factor: their hidden thinking tokens are output tokens too, so a high reasoning effort can add many seconds before the visible answer starts.
4. Where the money goes
cost = input_tokens × input_price + output_tokens × output_price, with prices per million
tokens. Output is typically several times more expensive than input, and cached input is discounted.
# USD per 1M tokens. gpt-5.6-luna is the real price when this was written (OpenAI SDK lesson 04);
# the second model is a HYPOTHETICAL 10x-priced flagship, to show the shape of the trade-off.
PRICES = {"gpt-5.6-luna": (0.20, 1.20), "flagship (hypothetical 10x)": (2.00, 12.00)}
CACHED_DISCOUNT = 0.90 # assumption: cached input billed at 10% of the input price
def request_cost(model, input_tokens, output_tokens, cached_tokens=0):
inp, out = PRICES[model]
billed_input = (input_tokens - cached_tokens) + cached_tokens * (1 - CACHED_DISCOUNT)
return (billed_input * inp + output_tokens * out) / 1_000_000
per_day = 50_000
for model in PRICES:
plain = request_cost(model, 4_000, 300)
cached = request_cost(model, 4_000, 300, cached_tokens=1_500) # stable system prompt + tools
print(f"{model:28} ${plain:.5f}/req ${plain * per_day * 30:>9,.0f}/month"
f" with caching ${cached * per_day * 30:>9,.0f}/month")
| Lever | Saves | Cost of pulling it |
|---|---|---|
| Send less context (better retrieval, trimmed history) | input tokens and TTFT | engineering time, and recall must stay high |
| Stable prefix for prompt caching | input cost and TTFT | none: just order your prompt well |
Cap output (max_output_tokens, "answer in 3 bullet points") | output cost and total latency | risk of truncated answers: check the finish status |
| Smaller model, or route easy requests to it (lesson 20) | often 5–20× | quality must be proven by evals first |
| Lower reasoning effort | thinking tokens and latency | harder tasks may get worse |
| Batch API for offline jobs | about 50% (per OpenAI's batch pricing) | results within hours, not seconds |
| Cache whole answers for repeated questions (lesson 20) | the entire call | staleness, and cache-key design |
5. Throughput and rate limits
Providers limit requests per minute (RPM) and tokens per minute (TPM) per model and
account tier. A batch job that fires 200 concurrent requests will hit 429 errors. Handle them with bounded
concurrency (a semaphore) and retries with exponential backoff, both covered in the OpenAI SDK course
(lessons 11–12). For capacity planning:
peak requests per minute × average tokens per request must stay under your TPM limit, with
headroom.
Recap
- Temperature reshapes the distribution; top-p/top-k cut its tail. Use low temperature when there is one right answer.
- Latency = TTFT + output tokens ÷ speed. Long outputs and huge prompts are the usual culprits. Streaming hides the rest.
- Cost = tokens × price, with output pricier and cached input cheaper. Estimate before you build.
- Never assert exact text in tests. Sampling and server nondeterminism make it flaky.