Load Balancers: L4 vs L7
One address in front of many servers: spreading connections at layer 4, understanding requests at layer 7, and what each one does when a server dies.
An interactive Networking lesson: 24 steps, about 35 minutes, on a live simulation in your browser.
The shop outgrew one machine, so it now runs on three: web-1, web-2 and web-3. Customers still type one name, shop.example, and that name resolves to one address: 203.0.113.80.
That address belongs to none of the three. It sits on the load balancer in front of them and is called the virtual IP, or VIP. Run curl from the client and watch: the packets go to the VIP, and a server you never named sends the answer. Read the Server: header in the output.
What you will learn
One address, many servers
- Three servers, one address: A VIP is an address that belongs to the service and not to any server, so the servers behind it can change without clients noticing.
- Ask the same address again
Layer 4: one choice per connection
- What the balancer does to a packet: A layer-4 balancer picks a backend when the SYN arrives and rewrites addresses on every packet after that. It never reads the request.
- The balancer's whole memory
- Drill: list the backends
- Four requests, one connection: Layer 4 balances connections, not requests. One long-lived connection is one decision, however many requests travel down it.
Who gets the next connection
- Least connections: Round-robin counts turns. Least-connections counts what is still open. They only differ when connections live for different lengths of time.
- Hashing the client address: ip-hash is a pure function of the client address and the backend list: no state, the same answer every time, and no regard for how busy the backend is.
When a backend dies
- Health checks: A health check decides who gets new connections. It does nothing for connections that already exist.
- The app crashes between two checks
- Break it: the whole machine dies: Refused means a live machine answered no. Timeout means nothing answered. Through a layer-4 balancer the client sees exactly what the chosen backend did.
- No backend left: A layer-4 balancer can only fail in TCP: with no backend it resets the connection, and the client reads that as connection refused.
Layer 7: reading the request
- Layer 7: the balancer answers itself: A layer-7 balancer is a server to the client and a client to the backends. It balances requests because requests are what it can see.
- The backend no longer sees the client: At layer 7 there are two TCP connections. The backend's peer is the balancer, so the client's address has to travel inside the request, in X-Forwarded-For.
- TLS ends at the balancer
- Routing by what the request asks for: Layer 7 routes first and balances second: the request picks the group of backends, the algorithm picks one inside the group.
- The API's only backend goes down
- Break it: 502 Bad Gateway: A layer-7 balancer fails in HTTP. 503 means nobody to ask, 502 means the one it asked failed, 504 means the one it asked was too slow.
- Drill: who wrote this answer?
Costs and stickiness
- What each layer sees, and what it costs: Layer 4 is cheap because it understands nothing above TCP. Layer 7 is capable because it understands the request. You pay for the second with CPU, memory and a second connection.
- Sticky sessions, and the office behind NAT: ip-hash balances addresses, not people. Everyone behind one NAT address is one client, and lands on one backend.
- An idle backend leaves the pool: Stickiness by hash lasts only as long as the pool does. Change the number of backends and clients move, including ones whose backend was fine.
Recap & playground
- Cheat sheet
- Playground