Module 1 · Thinking in scale

Latency, throughput & queueing

Core 20 min read The maths behind "it was fine yesterday"

Two numbers describe performance: latency (how long one request takes) and throughput (how many requests per second the system completes). They are linked by queues, and queues behave in a way that surprises everyone once: a server that is 50% busy feels fast, 80% busy feels a bit slow, and 95% busy feels broken, even though it is "only" 15% busier. This lesson simulates that, so you never plan a system to run hot again.

🛒

A supermarket checkout

If a cashier takes a minute per customer and customers arrive every two minutes, there is rarely a queue. At one customer every 70 seconds, a queue forms whenever two arrive close together, and it takes a while to clear. At one every 62 seconds the queue almost never clears. The cashier did not get slower; the waiting exploded. Opening a second till with one shared queue fixes it far more than you'd expect.

1. Averages hide the pain: use percentiles

Suppose 98% of Linkly's redirects take 20 to 40 ms, and 2% hit a slow path (a cold cache, a lock, a garbage collection pause) and take 1 to 3 seconds:

import random
from simkit import summarize

rng = random.Random(7)
latencies = [rng.uniform(1_000, 3_000) if rng.random() < 0.02 else rng.uniform(20, 40)
             for _ in range(10_000)]
for name, value in summarize(latencies).items():
    print(f"{name:>4}: {value:7.0f} ms")

# A page that makes 30 requests (HTML, API calls, redirects...) is only fast if ALL of them are.
for calls in (1, 10, 30, 100):
    print(f"page with {calls:3d} requests: {1 - 0.99 ** calls:5.1%} chance at least one is slower than p99")
mean: 72 ms p50: 30 ms p95: 39 ms p99: 1984 ms max: 2982 ms page with 1 requests: 1.0% chance at least one is slower than p99 page with 10 requests: 9.6% chance at least one is slower than p99 page with 30 requests: 26.0% chance at least one is slower than p99 page with 100 requests: 63.4% chance at least one is slower than p99

The mean lies

The mean is dragged up by the slow 2% yet still describes nobody's experience: most requests are far faster, and the slow ones far slower.

p99 is common

"1 in 100" sounds rare. With 100 requests behind one page, most page loads include a p99 request.

What to track

p50 (typical), p95/p99 (the tail) and max over a window, per endpoint. Alert on the tail, not the mean (lesson 18).

2. Little's law: the one formula to memorise

L = λ × W: the average number of requests in the system (L) equals the arrival rate (λ) times the average time each spends in it (W). It holds for any stable system, whatever the distributions. It tells you how many connections, threads or workers you need:

QuestionLittle's law
1,200 req/s, each takes 50 ms: how many in flight?1,200 × 0.05 = 60 concurrent requests
A database call takes 5 ms at 2,000 calls/s: pool size?2,000 × 0.005 = 10 connections busy on average (size the pool with headroom)
Latency doubles because a dependency slows down: what happens?in-flight requests double at the same traffic: threads, memory and connections run out

The last row is how one slow dependency takes down a healthy service. Lesson 15 is about stopping that.

3. Why 90% utilisation feels broken

This simulation sends Poisson arrivals (random, like real users) to a server whose service time averages 10 ms, at increasing utilisation, first with one server and then with four servers sharing one queue:

import heapq, random
from simkit import summarize, table

def simulate(utilisation, servers=1, service_ms=10.0, n=100_000, seed=1):
    """FIFO queue with `servers` workers; returns each request's total time (wait + service) in ms."""
    rng = random.Random(seed)
    arrival_rate = utilisation * servers / service_ms      # requests per ms
    free_at = [0.0] * servers                               # when each worker is next free
    now, times = 0.0, []
    for _ in range(n):
        now += rng.expovariate(arrival_rate)                # next arrival
        start = max(now, heapq.heappop(free_at))            # wait for the first free worker
        finish = start + rng.expovariate(1 / service_ms)
        heapq.heappush(free_at, finish)
        times.append(finish - now)
    return times

rows = []
for u in (0.5, 0.7, 0.8, 0.9, 0.95, 0.99):
    one, four = summarize(simulate(u)), summarize(simulate(u, servers=4))
    rows.append([f"{u:.0%}", f"{10 / (1 - u):.0f} ms", f"{one['mean']:.0f} ms", f"{one['p99']:.0f} ms",
                 f"{four['mean']:.0f} ms", f"{four['p99']:.0f} ms"])
table(rows, ["utilisation", "theory (1 server)", "1 server mean", "1 server p99", "4 servers mean", "4 servers p99"])
utilisation theory (1 server) 1 server mean 1 server p99 4 servers mean 4 servers p99 ----------- ----------------- ------------- ------------ -------------- ------------- 50% 20 ms 20 ms 90 ms 11 ms 47 ms 70% 33 ms 33 ms 143 ms 13 ms 54 ms 80% 50 ms 49 ms 221 ms 17 ms 68 ms 90% 100 ms 96 ms 482 ms 29 ms 125 ms 95% 200 ms 214 ms 1059 ms 58 ms 270 ms 99% 1000 ms 793 ms 2510 ms 203 ms 632 ms
0%25% 50%75% 100% 0100 ms 200 ms300 ms utilisation danger zone 80%: 50 ms (5×) 90%: 100 ms (10×) 50%: 20 ms mean time = service time ÷ (1 − utilisation) (one server, random arrivals, 10 ms average service time)

The hockey stick

Response time is service time divided by (1 − utilisation). Going from 80% to 90% busy doubles latency; 90% to 95% doubles it again. The p99 grows even faster. That is why capacity plans target 50–70% (lesson 22).

Pooling wins

Four workers sharing one queue at the same utilisation have a fraction of the latency of one worker, because a slow request only blocks one of four. It is also why one shared queue beats four separate queues (and why lesson 07's load-balancing algorithm matters).

About the 99% row

At 99% utilisation the queue swings so wildly that 100,000 simulated requests are not enough for the mean to settle near the theoretical 1,000 ms. That instability is the real-world lesson too: a system run that close to capacity has no predictable latency at all.

4. Throughput has a ceiling; latency tells you when you're near it

SignalWhat it meansWhere you'll use it
Throughput flattens while latency climbsyou have found the bottleneck's capacityload tests (lessons 04, 24)
Queue length (in-flight requests) grows steadilyarrivals exceed capacity: the queue will grow until something times outqueues and backpressure (lesson 11)
p99 rises long before the meanyou are entering the knee of the curvealerts and autoscaling (lessons 18, 22)
Latency high at low utilisationnot a capacity problem: slow dependency, lock, cold cachetracing (lesson 18)

Recap

  • Report percentiles (p50, p95, p99), never just the mean; tail latency compounds across many requests.
  • Little's law: in-flight = arrival rate × time in system. Use it to size pools and to predict what a slow dependency does.
  • Latency ∝ 1 / (1 − utilisation): plan for 50–70% busy, not 90%.
  • Shared queues and more workers tame the tail far more than their raw capacity suggests.

Checkpoint

1 · A service handles 500 req/s with 200 ms average latency. How many requests are in flight on average?
Little's law: 500 × 0.2 s = 100 requests in the system at any moment.
2 · A server runs at 90% utilisation. Traffic rises 5%. What happens to latency?
Near saturation, small increases in load cause huge increases in waiting time.
3 · Why is p99 more useful than the mean for user-facing latency?
A page with 30 requests has about a 26% chance of including at least one request slower than p99.