Scaling Up vs Scaling Out
Bigger machines or more machines: what each buys, what each costs per month and per million requests, where each stops working, and the state and timing problems that come with more machines.
An interactive System Design lesson: 21 steps, about 30 minutes, on a live simulation in your browser.
The shop runs on one machine, web-1, an AWS m6i.xlarge: 4 vCPU, 16 GiB, $0.192 an hour. Every request costs it about 21 ms of CPU, then a query on db. Four vCPU at 21.22 ms each is about 189 CPU-bound requests a second: the most this box could do if CPU were all that mattered.
Send 100 requests a second for five seconds. Every one succeeds in 50.2 ms. The capacity lens reads web-1 at about 53% CPU, the database at 12%, and the bottleneck none.
What you will learn
One box, one bill
- One server, measured: Capacity is the scarcest resource divided by what one request uses of it. Cost per million requests is the bill divided by the work: it falls as a machine fills up.
- Full before the CPU is: A server is full when its slots are full, and a slot waiting on a database is still taken. CPU can read 75% on a server that is turning customers away.
Scaling up
- Break it: resize the only server: Scaling up a single machine is an outage: the box must stop to change size, and with one box nothing serves while it does.
- Twice the vCPU, not twice the work: Double the cores and you get less than double the throughput: cores contend for shared resources, and that contention grows with the size of the box.
- The biggest box there is: Scaling up has two ceilings: the largest machine anyone sells, and the point on the curve where more cores stop adding throughput. Both arrive sooner than the price list suggests.
- Drill: what one box can do
Scaling out
- Five small boxes: Scaling out adds capacity in small, cheap steps with no ceiling on the server tier, and losing one box costs a fraction of capacity instead of all of it. Every server still shares whatever sits behind them.
- The bottleneck moves: Scaling one tier moves the bottleneck to the next shared one. Find the component every request touches and compute its ceiling: that is the system's ceiling.
- Three more servers: Servers added in front of a saturated database do not add throughput. They move the queue, and latency grows instead of errors.
- Scale the database's reads: Each tier scales out by its own technique: servers by copies behind a balancer, database reads by replicas, writes by sharding.
- Drill: cost per million requests
State gets in the way
- Break it: sessions in memory: A server that keeps state is no longer interchangeable. Round-robin only works when any server can answer any request.
- Sticky sessions: Sticky routing keeps a user on the server that holds their state. It works until that server leaves, and it spreads load only as evenly as the users hash.
- Scale in, with sticky sessions: Connection draining protects requests in flight, not state in memory. Every scale-in, deploy and crash logs out the users that server held.
- Move the state out: Stateless servers are what make scaling out work: put sessions, carts and uploads in a shared store, and every server becomes disposable.
- Drill: a session in Redis
Autoscaling and its lag
- Let a machine decide: An autoscaler is a control loop: measure, compare with a target, adjust. The target utilisation is the headroom you keep for the time it takes to react.
- A spike meets the boot time: Autoscaling is late by design: detection interval plus boot time plus health checks. Whatever arrives faster than that must be absorbed by headroom you already pay for.
- Shrink the lag: The fastest autoscaler is the one with the least to do: shorten boot time, keep headroom for the gap, and expect it to overshoot before it settles.
Recap & playground
- Cheat sheet
- Playground