Learn System Design
How big systems stay fast and stay up.
Basic
Start here. The ideas everything else is built on, from the first command.
- Load Balancing: One address, many servers: why one server falls over, how a balancer chooses between several, and what it does when one of them is slow, broken or dead. (about 30 minutes)
- Caching: Keep answers close so the database does not have to give them again: cache-aside, hit ratios, TTLs, invalidation, eviction, and the stampede when a hot key expires. (about 32 minutes)
- CDNs: Serve bytes from the edge, close to every user: what distance costs, what an edge saves the origin, how Cache-Control and versioned file names decide what users see after a deploy, and what happens when an edge or the origin goes down. (about 30 minutes)
- Scaling Up vs Scaling Out: Bigger machines or more machines: what each buys, what each costs per month and per million requests, where each stops working, and the state and timing problems that come with more machines. (about 30 minutes)
- SQL vs NoSQL: One shop stored four ways (Postgres, MongoDB, Redis Cluster, Cassandra): which questions each answers cheaply, which it refuses, and what each does with a schema change, a half-finished order, a write flood and a dead machine. (about 35 minutes)
- Back-of-the-Envelope Estimates: Turn daily users into requests a second, bytes a day, cache size and servers, round like an engineer, know the latency numbers by heart, and then prove the estimate against a running system. (about 30 minutes)
Intermediate
What you need to run it for real: the moving parts and the ways they fail.
- Consistent Hashing: Add a server and move only 1/N of the keys: why hash mod N reshuffles everything, how a hash ring limits a change to one arc, how virtual nodes even out the load, and what no hashing scheme can do about a hot key. (about 30 minutes)
- Database Replication: Leaders, followers, failover and the lag that bites: what copies of a database buy you, and what each kind of copy costs. (about 32 minutes)
- Sharding: Split a database that no longer fits on one machine: shard keys, routers, modulo versus range versus consistent hashing, and the queries and hot keys that sharding cannot fix. (about 30 minutes)
- Rate Limiting: Token buckets, leaky buckets and sliding windows: how an API says 429 to one noisy client so that everyone else still gets served, and where that decision lives. (about 30 minutes)
- Message Queues: Decouple services, absorb spikes and retry safely: acknowledgements, visibility timeouts, duplicates and idempotency, dead-letter queues, backpressure, ordering and pub/sub. (about 32 minutes)
- Indexes & Query Plans: Read EXPLAIN ANALYZE line by line, turn a 2-second sequential scan into a sub-millisecond index lookup, and learn why the planner sometimes ignores your index, or trusts statistics that lie. (about 35 minutes)
- API Design: REST, gRPC & Pagination: An API is a contract that outlives its first client: status codes and error bodies, conditional requests, cursors instead of offsets, REST against gRPC on the wire, and changing a schema without breaking the apps already installed on people's phones. (about 32 minutes)
- Idempotency & Retries: A payment whose answer is lost: why a blind retry charges twice, how an Idempotency-Key makes the retry safe, which failures are worth retrying, how retries at every layer multiply into 27 requests, and how jitter and retry budgets keep a recovering service alive. (about 32 minutes)
- Search & Inverted Indexes: Why LIKE '%word%' scans every row while a search engine answers in milliseconds: analyzers, postings lists, BM25 scoring, near-real-time refresh, and scatter-gather across shards that can be slow or gone. (about 35 minutes)
Advanced
Production depth: hardening, recovery and the trade-offs behind the design.
- CAP & Consistency: Three copies of a shopping cart in two data centres: read and write quorums, what a network partition forces you to choose, and how diverged copies are put back together. (about 32 minutes)
- Consensus & Raft: Five machines keep one history of a shop's stock: elections and terms, commit on a majority, what crashes and partitions do, and why clusters come in odd sizes. (about 32 minutes)
- Distributed Transactions: 2PC & Sagas: One order touches four services with four databases. Two-phase commit makes it atomic and blocks when the coordinator dies; a saga never blocks and pays with compensations, idempotent steps and the outbox. Watch both fail, and learn which to choose. (about 34 minutes)
- Event Sourcing & CQRS: Store what happened instead of what is: an append-only log of events as the source of truth, state rebuilt by replaying it, optimistic concurrency on stream versions, snapshots, questions about the past, read models that lag and can be rebuilt, schemas that change, and personal data you must forget in a log that never forgets. (about 32 minutes)
- Observability: Metrics, Logs & Traces: Finding the slow hop in a request that crossed eight services: traces and their critical path, RED metrics and the percentiles averages hide, alerts on symptoms rather than causes, sampling, cardinality, and logs that join up by trace id. (about 34 minutes)
- Case Study: A URL Shortener: The classic interview question done properly and then run in production: requirements, numbers, API, data model, five ways to make a short code, a 301 that hides your clicks, a hot link, a bot and a dead shard. (about 35 minutes)
- Case Study: A Chat System: WhatsApp in an interview and then in production: requirements, numbers, a WebSocket protocol, per-conversation sequence numbers, receipts, offline inboxes and push, retries without duplicates, presence, a gateway crash and the 5,000-member group. (about 35 minutes)
System design you can watch under load
System design interviews and real architecture reviews ask the same questions: what happens when traffic doubles, when a server dies, when the cache is cold, when the network splits? Diagrams on a whiteboard cannot answer them. These lessons run the system: requests flow through load balancers, caches, queues, databases and replicas in your browser, and every latency, hit ratio and error rate comes from the simulation.
You learn the building blocks one at a time: load balancing, caching and CDNs, consistent hashing, replication, sharding, rate limiting, message queues, the CAP theorem, Raft consensus, SQL versus NoSQL, indexes and query plans, API design, idempotency and retries, search, distributed transactions with two-phase commit and sagas, event sourcing and CQRS, and observability with metrics, logs and traces. Then you put them together in case studies such as a URL shortener and a chat system, estimating the load first and testing the design against it.
Every concept is something you break and fix: crash a replica and watch failover, partition the network and watch the trade-off CAP describes, let retries storm a recovering service and add jitter.
After these lessons you can
- Estimate traffic, storage and server counts with back-of-the-envelope maths
- Choose between caching, replication, sharding and queues for a given bottleneck
- Explain consistency trade-offs, CAP and Raft with concrete failure scenarios
- Design safe retries, idempotent APIs and distributed transactions
- Read a query plan, pick an index, and choose between SQL and NoSQL stores
- Walk through a complete design, such as a URL shortener or a chat system, in an interview
Who it is for
Engineers preparing for system design interviews, developers moving into senior or staff roles, and anyone who wants to understand why large systems are built the way they are.
Common questions
Is this useful for system design interviews?
Yes. The lessons cover the topics interviews ask about and the two case studies follow the interview structure: requirements, estimates, API, data model, design, deep dives, failure modes and scaling.
Are the numbers realistic?
They come from the simulation, which models latency, capacity, cache behaviour and failures. Defaults follow real systems such as Postgres, Redis, Kafka-style queues and nginx, and each lesson says where a number comes from.
Do I need to know distributed systems already?
No. The basic lessons start with a single server and add one idea at a time.