Latency, throughput & queueing
Two numbers describe performance: latency (how long one request takes) and throughput (how many requests per second the system completes). They are linked by queues, and queues behave in a way that surprises everyone once: a server that is 50% busy feels fast, 80% busy feels a bit slow, and 95% busy feels broken, even though it is "only" 15% busier. This lesson simulates that, so you never plan a system to run hot again.
A supermarket checkout
If a cashier takes a minute per customer and customers arrive every two minutes, there is rarely a queue. At one customer every 70 seconds, a queue forms whenever two arrive close together, and it takes a while to clear. At one every 62 seconds the queue almost never clears. The cashier did not get slower; the waiting exploded. Opening a second till with one shared queue fixes it far more than you'd expect.
1. Averages hide the pain: use percentiles
Suppose 98% of Linkly's redirects take 20 to 40 ms, and 2% hit a slow path (a cold cache, a lock, a garbage collection pause) and take 1 to 3 seconds:
import random
from simkit import summarize
rng = random.Random(7)
latencies = [rng.uniform(1_000, 3_000) if rng.random() < 0.02 else rng.uniform(20, 40)
for _ in range(10_000)]
for name, value in summarize(latencies).items():
print(f"{name:>4}: {value:7.0f} ms")
# A page that makes 30 requests (HTML, API calls, redirects...) is only fast if ALL of them are.
for calls in (1, 10, 30, 100):
print(f"page with {calls:3d} requests: {1 - 0.99 ** calls:5.1%} chance at least one is slower than p99")
The mean lies
The mean is dragged up by the slow 2% yet still describes nobody's experience: most requests are far faster, and the slow ones far slower.
p99 is common
"1 in 100" sounds rare. With 100 requests behind one page, most page loads include a p99 request.
What to track
p50 (typical), p95/p99 (the tail) and max over a window, per endpoint. Alert on the tail, not the mean (lesson 18).
2. Little's law: the one formula to memorise
L = λ × W: the average number of requests in the system (L) equals the arrival rate (λ) times the average time each spends in it (W). It holds for any stable system, whatever the distributions. It tells you how many connections, threads or workers you need:
| Question | Little's law |
|---|---|
| 1,200 req/s, each takes 50 ms: how many in flight? | 1,200 × 0.05 = 60 concurrent requests |
| A database call takes 5 ms at 2,000 calls/s: pool size? | 2,000 × 0.005 = 10 connections busy on average (size the pool with headroom) |
| Latency doubles because a dependency slows down: what happens? | in-flight requests double at the same traffic: threads, memory and connections run out |
The last row is how one slow dependency takes down a healthy service. Lesson 15 is about stopping that.
3. Why 90% utilisation feels broken
This simulation sends Poisson arrivals (random, like real users) to a server whose service time averages 10 ms, at increasing utilisation, first with one server and then with four servers sharing one queue:
import heapq, random
from simkit import summarize, table
def simulate(utilisation, servers=1, service_ms=10.0, n=100_000, seed=1):
"""FIFO queue with `servers` workers; returns each request's total time (wait + service) in ms."""
rng = random.Random(seed)
arrival_rate = utilisation * servers / service_ms # requests per ms
free_at = [0.0] * servers # when each worker is next free
now, times = 0.0, []
for _ in range(n):
now += rng.expovariate(arrival_rate) # next arrival
start = max(now, heapq.heappop(free_at)) # wait for the first free worker
finish = start + rng.expovariate(1 / service_ms)
heapq.heappush(free_at, finish)
times.append(finish - now)
return times
rows = []
for u in (0.5, 0.7, 0.8, 0.9, 0.95, 0.99):
one, four = summarize(simulate(u)), summarize(simulate(u, servers=4))
rows.append([f"{u:.0%}", f"{10 / (1 - u):.0f} ms", f"{one['mean']:.0f} ms", f"{one['p99']:.0f} ms",
f"{four['mean']:.0f} ms", f"{four['p99']:.0f} ms"])
table(rows, ["utilisation", "theory (1 server)", "1 server mean", "1 server p99", "4 servers mean", "4 servers p99"])
The hockey stick
Response time is service time divided by (1 − utilisation). Going from 80% to 90% busy doubles latency; 90% to 95% doubles it again. The p99 grows even faster. That is why capacity plans target 50–70% (lesson 22).
Pooling wins
Four workers sharing one queue at the same utilisation have a fraction of the latency of one worker, because a slow request only blocks one of four. It is also why one shared queue beats four separate queues (and why lesson 07's load-balancing algorithm matters).
At 99% utilisation the queue swings so wildly that 100,000 simulated requests are not enough for the mean to settle near the theoretical 1,000 ms. That instability is the real-world lesson too: a system run that close to capacity has no predictable latency at all.
4. Throughput has a ceiling; latency tells you when you're near it
| Signal | What it means | Where you'll use it |
|---|---|---|
| Throughput flattens while latency climbs | you have found the bottleneck's capacity | load tests (lessons 04, 24) |
| Queue length (in-flight requests) grows steadily | arrivals exceed capacity: the queue will grow until something times out | queues and backpressure (lesson 11) |
| p99 rises long before the mean | you are entering the knee of the curve | alerts and autoscaling (lessons 18, 22) |
| Latency high at low utilisation | not a capacity problem: slow dependency, lock, cold cache | tracing (lesson 18) |
Recap
- Report percentiles (p50, p95, p99), never just the mean; tail latency compounds across many requests.
- Little's law: in-flight = arrival rate × time in system. Use it to size pools and to predict what a slow dependency does.
- Latency ∝ 1 / (1 − utilisation): plan for 50–70% busy, not 90%.
- Shared queues and more workers tame the tail far more than their raw capacity suggests.