Interactive course · 24 lessons · ~8 hours

Scaling from 1 user to 1 million: the real architecture journey.

Most system design material shows you the finished diagram: load balancers, caches, shards, queues, regions. This course shows you why each box appears, and when. It follows one product, Linkly (a URL shortener), from a weekend project on one server to a million users. At every stage something specific breaks, we measure it with code that runs, add exactly the piece of architecture that fixes it, and find the next bottleneck.

0 of 24 lessons complete0%

The journey this course is built on

Each stage keeps everything before it, and nothing is added until something forces it. The modules follow the stages.

users → (each step is ten times the last) 1 userone serverapp + databasemodule 2 100 – 1Kown database,stateless appsmodule 2 1K – 100Kload balancer, cache,CDN, replicas,queuesmodule 3 100K – 1Msharding, rate limits,retries & breakers,consistency, servicesmodule 4 operating at 1M+SLOs, feeds, regions,security, capacity, costmodule 5 · then the interviewand the project (module 6)
What "runnable" means in a system design course

You can't run a thousand servers in a browser tab, so most lessons use small seeded simulations (queues, load balancers, caches, retry storms, quorums, autoscalers, twenty years of failures) written in plain Python with a shared helper, simkit.py. Where the real thing fits on a laptop, the course runs it: a real server pushed to its ceiling, real SQLite indexes, real connections, and a project that starts real processes behind a real load balancer. Every output on every page was produced by running the code. Numbers that are assumptions (cloud prices, network distances) are labelled as such.

Pick a path

~3 hours

New to system design

Lessons 01–11 in order: estimation, queueing, databases, stateless servers, load balancers, caches, CDNs, replication, queues.

~3 hours

Interview in two weeks

Lessons 02, 08, 12, 13, 14, 16, 19, then 23 (the framework and a worked chat design).

~2.5 hours

Keeping production up

Lessons 03, 07, 15, 18, 20, 22: queueing, health checks, retry storms, SLOs, regions, capacity.

~2 hours

Architect / tech lead

Lessons 01, 12, 16, 17, 21, 22: one-way doors, sharding, consistency, service boundaries, security, cost.

Ships with a hands-on project

🛠 Scale Lab: take Linkly from 1 to 1M users

A real URL shortener pushed through six architecture stages under real load on your own machine: no index, an index, a click queue, a cache, three servers behind a least-connections load balancer with health checks, and four shards on a consistent-hash ring. A raw-socket load generator measures throughput, p50/p99, database load and cache hit ratio at each stage; a failover experiment kills a server mid-run; and the measurements become a capacity plan for a million users. 30 tests, no services to install.

Open the project →

The curriculum