learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Batching for LLM Inference

Why one request at a time wastes a memory-bound GPU, static and dynamic batching and where each belongs, continuous iteration-level batching, the throughput-latency curve and its knee, prefill against decode and the token budget, max-num-seqs, KV-bound batches, and settings for a chat SLO versus an offline job.

An interactive AI Infrastructure lesson: 20 steps, about 35 minutes, on a live simulation in your browser.

A helpdesk company answers customers with Llama 3.1 8B in bf16 on one H100 of dgx-1. A request carries about 512 tokens (history plus question) and gets about 256 back. The first version of the server takes one request at a time.

One request: prefill 14 ms, then 6.02 ms per token, 166 tokens/s. Each decode step reads all 16 GB of weights to produce one token: 16 GB in 6.02 ms is 2.67 TB/s, 80% of the H100's 3.35 TB/s. The maths is 2 FLOPs per parameter, 16 GFLOP per token, 2.7 TFLOPS at this pace: 0.3% of the 989 TFLOPS the tensor cores can do.

What you will learn

  1. One request at a time

    • A support bot, one request at a time: One decode step reads every weight to make one token per sequence; with one sequence the GPU is bandwidth-bound and its tensor cores sit 99% idle.
    • Drill: the single-stream ceiling: Single-stream decode speed ≈ memory bandwidth ÷ bytes of weights; real servers reach about 80% of that.
  2. Batching that waits

    • Break it: static batches at peak: Static batching fixes the batch at the start and runs it to the slowest member: short answers leave padded slots, and newcomers wait a whole batch.
    • Dynamic batching for classic models: Dynamic batching trades a bounded wait for a much bigger batch; it is right when every request costs the same, as in classifiers, embedders and vision models.
  3. Iteration-level batching

    • Same 64 slots, refilled every step: Continuous batching refills a slot the moment its sequence finishes, so no slot is padding and no request waits for someone else's answer to end.
    • Inside one engine iteration: One iteration = one token for each decoding sequence + prefill chunks for newcomers, within a token budget, a sequence cap and the free KV blocks.
    • Throughput versus latency: As the batch grows, throughput rises much faster than per-token latency, until compute, the token budget or the KV cache binds; below that knee, batching is nearly free.
    • Break it: past the knee: Past saturation, more load adds no throughput, only queueing: TTFT explodes while TPOT and tokens/s stay flat.
    • Drill: sequences in flight: Sequences in flight = requests per second × seconds per request; it tells you the batch size your traffic will build, before you run anything.
  4. Prefill against decode

    • Break it: long prompts stall the stream: A prefill in the same iteration as decodes delays every decoding sequence by its compute time; long prompts show up as TPOT spikes for other users.
    • A smaller token budget: The token budget trades TTFT against TPOT: small budgets protect the decode stream, large ones finish prompts sooner and run bigger, more efficient passes.
    • The scheduler's three dials: max-num-batched-tokens shapes each iteration, max-num-seqs caps the batch, and the KV cache caps both: tune them in that order, one at a time, against measured TTFT and TPOT.
  5. What caps the batch

    • Capping the batch for speed: max-num-seqs below the batch your traffic builds (rate × time per request) does not make users faster; it moves them from the decode stream into the queue.
    • Break it: the cache caps the batch: With long sequences the KV cache, not max-num-seqs, sets the batch: batch ≤ cache tokens ÷ tokens per sequence.
    • Drill: the batch the cache allows: Before tuning max-num-seqs, divide cache tokens by tokens per sequence: if that is smaller, the cap you set will never be reached.
  6. Choosing settings

    • Settings for a chat SLO: For a latency SLO, find the per-replica rate where the curve still meets TTFT and TPOT, divide the peak by it, and add headroom; do not buy latency by shrinking the batch.
    • Settings for an offline job: Offline jobs want the largest batch the KV cache allows: per-token latency is irrelevant, so push past the knee that a chat service must stay below.
    • Reading TTFT and TPOT percentiles: TTFT p99 is the early warning for queueing; TPOT p99 is the warning for batch size and prefill interference; read them per pool, at p50 and p99.
  7. Recap & playground

    • Cheat sheet
    • Playground