learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Back-of-the-Envelope Estimates

Turn daily users into requests a second, bytes a day, cache size and servers, round like an engineer, know the latency numbers by heart, and then prove the estimate against a running system.

An interactive System Design lesson: 20 steps, about 30 minutes, on a live simulation in your browser.

An interviewer says: "Design the backend for a social feed with 10 million daily active users." Before any box is drawn, you need to know whether this is one server or a thousand. That takes five minutes of arithmetic, not a load test.

Start by saying your assumptions out loud, because every number after this one is built on them: each user makes 10 requests a day; reads outnumber writes 10 to 1; a post is about 1 KB; we keep data 5 years in 3 copies; the busiest hour runs at twice the daily average.

What you will learn

  1. Assumptions in, requests out

    • The question behind the question: An estimate is assumptions times arithmetic. State the assumptions first; then anyone can check the arithmetic, and a wrong assumption costs one line, not the whole design.
    • Users a day to requests a second: A day is about 100,000 seconds. So a million requests a day is about 10 a second, and 100 million a day is about 1,000 a second.
    • Round to one figure: Round inputs to one figure so the arithmetic fits in your head, then round the answer up. A napkin estimate is good to a factor of two; that is all a design decision needs.
    • Drill: requests a second
  2. Storage, bandwidth, cache

    • Bytes a day, bytes in five years: Storage = writes a day × object size × days kept × copies. Then divide by users and ask whether the answer per user is believable.
    • What if posts are photos?: Object size multiplies storage, bandwidth and cache, and leaves requests a second untouched. Small records are a request-rate problem; media is a bytes problem.
    • How big a cache?: Cache size = a day's reads × object size × the hot fraction (20% by the 80/20 rule). If it fits in one machine's RAM, start with one node.
    • Drill: a year of uploads
  3. Numbers every engineer should know

    • The latency ladder: Latency comes in rungs a factor of 10 to 1000 apart: memory, SSD, network in the data centre, disk, the other side of the world. Count the expensive rungs a request climbs, not the cheap ones.
    • Drill: read a gigabyte
  4. From numbers to machines

    • Requests a second to servers: Servers needed = peak requests a second ÷ what one server can do. The second number must come from a measurement or a model of your own server, not from a blog post.
    • Break it: run the estimated peak: Check the estimate against every shared component, not only the one you sized: peak requests × queries per request against what the database can do.
    • Add the cache you estimated: A capacity number measured at 100% is a ceiling, not a plan. Real servers wait on dependencies and need headroom, so plan for a fraction of the ceiling.
    • Plan at 70%: Size for the peak at a target utilisation (60 to 70% is common), not at 100%. The gap is what absorbs a bad minute, a lost server and the autoscaler's lag.
    • Break it: launch day: The peak factor is the assumption most likely to be wrong. When traffic doubles, every tier doubles with it, and the shared one gives first.
  5. What if?

    • Double the users: Most estimate lines are linear in their inputs: double the users and you double requests, storage, bandwidth, cache and servers. Knowing that lets you answer "what if we grow 10×?" in one breath.
    • Double the object size: Users and actions drive requests and everything else; object size drives only bytes. Ask which input a question changes before you redo the whole sheet.
    • The five-minute interview version: Assumptions, traffic, bytes, machines, sanity check: the same five beats every time, rounded to one figure and said out loud.
  6. Recap & playground

    • Cheat sheet
    • Playground