The journey: 1 user to 1 million
Meet Linkly, a URL shortener. Its founder writes the first version over a weekend: one small server, one database, one happy user. This course follows Linkly all the way to a million users. At every stage something specific breaks, we measure it, and we add exactly the piece of architecture that fixes it. By the end you will know not just what a load balancer, a cache or a shard is, but when each one earns its place and what it costs.
From a food stall to a restaurant chain
A food stall does not start with a central kitchen, a delivery fleet and a franchise manual. It starts with one cook. When the queue gets long, it hires a second cook (more servers). When the same dish is ordered all day, it pre-cooks it (a cache). When people come from the other side of town, it opens a branch (a CDN, then a region). Each change answers a problem that actually happened. Build the chain on day one and you go bankrupt before the first customer arrives.
1. What "scale" actually means
"Can it scale?" hides three different questions. A system can be fine on one axis and failing on another, so always ask which one you mean.
| Axis | Measured in | What breaks first | Typical fixes |
|---|---|---|---|
| Load | requests per second, concurrent connections | CPU of the app tier, database connections, one hot key | more servers, caches, queues, rate limits |
| Data | gigabytes, rows, write rate | disk, backups taking hours, queries that scan, index size vs RAM | indexes, archiving, partitioning, sharding, object storage |
| Organisation | engineers, teams, deploys per day | merge conflicts, slow releases, "who owns this?" | modules, clear ownership, sometimes services |
There is also a quality you get for free at stage 0 and lose as you grow: simplicity. Every box you add is something that can fail, must be monitored, and must be understood by the next engineer.
2. From users to requests
"A million users" sounds enormous. What does it mean for the servers? We need a few assumptions about how people use Linkly. Every one of them is a guess you would check against real analytics later. The point is to write them down so they can be argued with.
from simkit import table, human, human_bytes
STAGES = [1, 100, 1_000, 10_000, 100_000, 1_000_000] # registered users
DAU_SHARE = 0.20 # 20% of users are active on a given day
LINKS_PER_DAU = 2 # each active user creates two short links a day (writes)
CLICKS_PER_DAU = 100 # their audiences click 100 times a day in total (reads: redirects)
PEAK_FACTOR = 5 # the busiest hour runs at 5x the daily average
LINK_BYTES = 500 # code + long URL + owner + timestamps + index overhead
CLICK_EVENT_BYTES = 100 # one analytics record per click: time, code, country, referrer
rows = []
for users in STAGES:
dau = users * DAU_SHARE
writes, reads = dau * LINKS_PER_DAU, dau * CLICKS_PER_DAU
avg_rps = (reads + writes) / 86_400 # seconds in a day
rows.append([human(users, 0), human(dau), human(writes), human(reads), f"{avg_rps:.2f}",
f"{avg_rps * PEAK_FACTOR:.1f}", human_bytes(writes * 365 * LINK_BYTES),
human_bytes(reads * 365 * CLICK_EVENT_BYTES)])
table(rows, ["users", "DAU", "links/day", "clicks/day", "avg req/s", "peak req/s", "links/yr", "click log/yr"])
The surprise
A million users is about 240 requests per second on average and roughly 1,200 at peak. A single well-tuned server can handle that. The reasons to go beyond one server are elsewhere.
Reads dominate
With 50 clicks for every new link, this is a read-heavy system. That points at caches and replicas long before sharding.
Data grows quietly
The links themselves stay small (73 GB a year). The click log does not: 730 GB a year, and it never shrinks. Analytics becomes the data problem.
3. Averages lie: spikes, hot keys and availability
If average load is small, why do real systems at this size need load balancers, caches and replicas? Because traffic is not average. Here is one link going viral, and what three years of the click log look like:
from simkit import human, human_bytes
# A celebrity posts a Linkly link to 30M followers; 4% click, almost all in the first ten minutes.
followers, click_rate, window_s = 30_000_000, 0.04, 600
viral_rps = followers * click_rate / window_s
normal_peak = 1_180 # from the table above, at 1M users
print(f"viral link: {human(followers * click_rate)} clicks in 10 minutes = {viral_rps:,.0f} req/s "
f"on ONE key ({viral_rps / normal_peak:.1f}x the normal peak for the whole site)")
# The click log keeps everything unless someone decides otherwise.
per_day = 20_000_000 * 100 # clicks/day at 1M users x bytes per event
for year in (1, 2, 3):
print(f"click log after {year} year{'s' if year > 1 else ''}: {human_bytes(per_day * 365 * year)}")
# Availability: what a single server's downtime means in a year
for nines, label in [(0.99, "99%"), (0.999, "99.9%"), (0.9999, "99.99%")]:
minutes = (1 - nines) * 365 * 24 * 60
print(f"{label:>7} available = {minutes:,.0f} minutes of downtime a year")
Spikes (the viral link is more than the rest of the site put together, and it all hits one row), data volume (terabytes of history), availability (one server means every reboot and every deploy is an outage, and 99.9% still allows almost nine hours of downtime a year), and latency for distant users. Average requests per second is the least interesting number on the page.
4. The stages, and where this course covers them
| Stage | Users | What hurts | What we add | Lessons |
|---|---|---|---|---|
| 0 | 1–100 | nothing yet; speed of building matters most | one server, one process, backups | 04 |
| 1 | 100–1K | app and database fight for memory; one crash loses both | a separate, managed database; indexes | 05 |
| 2 | 1K–10K | deploys cause downtime; peaks overload one server | stateless apps behind a load balancer | 06, 07 |
| 3 | 10K–100K | database reads saturate; slow pages abroad; slow requests doing heavy work | cache, CDN, read replicas, queues | 08–11 |
| 4 | 100K–1M | writes and data outgrow one database; abuse; partial failures cascade | sharding, rate limits, timeouts and breakers, SLOs | 12–18 |
| 5 | 1M+ | fan-out, region outages, global users, the cloud bill | feeds, multiple regions, security at the edge, capacity planning | 19–22 |
5. Decide what is cheap now and expensive later
"Don't build for a million users on day one" is not "don't think". Some decisions are two-way doors (easy to change later) and some are one-way doors. Spend your early care on the one-way doors:
| Decision | Cheap on day one | Painful at a million |
|---|---|---|
| Identifiers | random or time-ordered IDs (UUIDv7, Snowflake-style) | auto-increment IDs leak volume, collide across shards (lesson 12, 21) |
| State | sessions and uploads outside the app process | rewriting login and storage while adding servers (lesson 06) |
| Time | store UTC everywhere | migrating billions of rows of local times |
| Configuration | environment variables, not hard-coded hosts | every new environment needs a code change |
| Data model | a clear owner for each table; no cross-module joins by habit | untangling a shared database to split anything (lesson 17) |
| Observability | request IDs and structured logs from the start | debugging production blind (lesson 18) |
6. How every lesson works
Symptom → measurement → fix → new bottleneck
We never add a component because diagrams usually have one. Each lesson starts from something that hurts, measures it with runnable code, applies the fix, and names the next bottleneck.
Simulations you can run
Networks of hundreds of servers don't fit in a lesson, so many examples are small seeded simulations (see simkit.py in this course folder). Where a real server fits on a laptop, the course runs a real one and measures it.
Recap
- Scale has three axes: load, data and organisation. Name the one you mean.
- Turn users into numbers with written-down assumptions: DAU, actions per user, read/write ratio, peak factor, bytes per record.
- Averages are misleading: spikes, hot keys, data growth and availability drive most architecture.
- Add components when something forces you to, but get the one-way doors (IDs, state, time, config) right early.