learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

The KV Cache

Why every generated token needs the keys and values of every token before it, what that costs per token (MHA, GQA, MQA; bf16 and fp8), how many sequences fit after the weights, paged blocks, shared prefixes, what happens when the cache is full, offloading and disaggregation, and the two vLLM metrics that tell you.

An interactive AI Infrastructure lesson: 21 steps, about 35 minutes, on a live simulation in your browser.

The developer-tools team serves a coding assistant: Llama 3.1 8B in bf16 on one H100 of dgx-1, behind vLLM. Before a single request arrives, nvidia-smi shows the GPU almost full: about 67 GiB of 80 used.

That is on purpose. vLLM takes 90% of the memory (--gpu-memory-utilization 0.9), loads 16 GB of weights, keeps 1.5 GB of workspace and 1 GB for the CUDA context, and hands everything else to the KV cache: room for 407,714 tokens, allocated at start-up so nothing else can take it.

What you will learn

  1. Why decode needs a cache

    • A coding assistant on one H100: An LLM server pre-allocates its KV cache at start-up: GPU memory looks full before the first request, and that is the healthy state.
    • What each new token reads: Each new token attends to the keys and values of every earlier token in every layer; because those never change, they are computed once and kept: that store is the KV cache.
    • Recompute instead of remember: Without a KV cache, every output token repeats the prefill of the whole sequence so far; the cache trades memory for that compute, and memory is the price you pay for the rest of this lesson.
  2. Bytes per token

    • The formula on the whiteboard: KV bytes per token = 2 × layers × KV heads × head dim × bytes per value; read layers, num_key_value_heads and head_dim from the model's config.json.
    • MHA, GQA and MQA: MHA keeps one K and V per query head, MQA one for all, GQA one per group; the KV cache scales with KV heads, not with parameters.
    • Why GQA made 70B practical: GQA cut the 70B KV cache by 8×; at tp=4 that is the difference between six 8K-token conversations at once and forty-eight.
    • Break it: a context the cache can't hold: max-model-len is a per-request limit; the KV cache is the shared pool. A server will not start unless one max-length request fits the pool.
    • Drill: 70B with an fp8 cache: Llama 3.1 70B costs 320 KiB per token in bf16 and 160 KiB in fp8: a 32K-token conversation is 10.7 GB or 5.4 GB of HBM.
  3. Capacity and context length

    • From bytes to sequences: Concurrent sequences = (0.9 × GPU memory − weights − workspace) ÷ (KV bytes per token × tokens per sequence); context length divides it directly.
    • Break it: whole files as context: Sequences you need in flight = requests per second × seconds each one lives; when that exceeds what the KV cache holds, the rest wait, however idle the arithmetic units look.
    • Drill: how many fit?: Size concurrency from the longest requests you accept, not the average: the long tail fills the cache.
  4. Contiguous versus paged

    • Break it: contiguous slabs at 32K: Contiguous allocation reserves max-model-len per request; raising the context limit for a few users divides everyone's concurrency.
    • PagedAttention: blocks, not slabs: Paged KV allocates fixed 16-token blocks on demand through a block table: waste falls to under one block per sequence, so concurrency is set by tokens actually used.
    • The flags that shape the cache: Every KV decision is a start-up flag: memory share, max length, block size, cache dtype, prefix sharing; the start-up log prints the resulting capacity in tokens.
  5. Sharing and running out

    • One repository, many questions: Prefix caching stores identical leading blocks once and shares them by reference: a shared system prompt or document costs its KV once per replica, not once per request.
    • When the cache is full: A full KV cache means new requests queue, and an optimistic scheduler also preempts running ones and recomputes them later; steady preemptions mean you need more cache, not more patience.
    • Reading KV usage and the queue: Diagnose a slow LLM server from KV usage and waiting requests together: full cache plus a queue is a memory problem; a queue with a half-empty cache is a scheduling or compute problem.
  6. Smaller, and beyond one GPU

    • Half the bytes per token: An fp8 KV cache halves bytes per token, independently of the weights, so the same memory holds twice the sequences and every decode step reads half the KV; it is the first knob to turn when the cache binds, and the next bottleneck shows right after.
    • Beyond HBM: offload and move the KV: Offloading and disaggregation both treat KV as data you move; they pay when moving the bytes is faster than recomputing them, which for long prompts it usually is.
  7. Recap & playground

    • Cheat sheet
    • Playground