learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Speculative Decoding

Why decode leaves the tensor cores idle, how a draft-and-verify loop with rejection sampling buys several tokens per weight read without changing the output, the acceptance rate and the expected-tokens formula, draft models, n-gram lookup, Medusa and EAGLE, when speculation pays and when it hurts under load, what it costs in memory, how to read its metrics, and how it combines with quantization and batching.

An interactive AI Infrastructure lesson: 20 steps, about 35 minutes, on a live simulation in your browser.

The developer-tools team serves complete, the code assistant inside the company IDE: Llama 3.1 70B in fp8 with an fp8 KV cache on two H100s of dgx-1, the two-GPU replica from "Quantization". A request carries 1,024 tokens of the file around the cursor and asks for about 256 tokens of code.

One request: the first token after 61.5 ms, then 14.49 ms per token, 69 tokens a second. A 40-line function takes almost four seconds to appear, and developers stop waiting for suggestions that slow.

What you will learn

  1. Why decode is slow

    • An autocomplete that types slowly: For a latency-bound service the metric is time per output token for one stream, and that is set by how long one decode step takes.
    • One weight read per token: A decode step costs one full read of the weights whether it processes one token or five; the extra tokens ride on compute that was idle anyway.
  2. Draft and verify

    • Guess cheaply, check in one pass: Draft k tokens cheaply, verify all k + 1 positions in one target pass, keep the agreeing prefix plus one target token: each step yields 1 to k + 1 tokens.
    • Rejection sampling keeps it exact: With rejection sampling, the draft can only change how fast tokens arrive, never which tokens the target would have produced.
    • Turn it on for code: The speed-up is tokens per verify step divided by the cost of a verify step; at high acceptance and low batch that is close to the tokens per step.
  3. Acceptance and k

    • The acceptance rate α: Expected tokens per verify step = (1 − α^(k+1)) ÷ (1 − α). Acceptance compounds: a token counts only if all the guesses before it were right.
    • Drill: tokens per verify step: At α = 0.8 and k = 4 a verify step is worth 3.36 tokens: a 3× speed-up at most, before paying for the draft.
    • Choosing k: Raise k only where acceptance is high: the gain from the k-th draft token is α^k, its cost is a full draft step and one more verified position.
  4. Where drafts come from

    • A draft model, and what it costs: A draft model costs its weights plus its own KV cache for every token in flight, paid out of the target's KV cache, which is your concurrency.
    • n-gram lookup, Medusa and EAGLE: Pick the draft by your traffic: n-gram when outputs copy inputs, EAGLE or Medusa heads trained for your exact target otherwise, a small sibling model when neither exists.
  5. When it pays, when it hurts

    • A few developers typing: At low batch the decode step is memory-bound with idle tensor cores, so speculation's single-stream gain carries over to real traffic almost intact.
    • The busy hour: Speculation's gain shrinks as the batch grows: the spare compute it spends is the same spare compute batching spends on more users.
    • Break it: speculation in a full batch: In a full batch the tensor cores are no longer idle: verifying k + 1 positions per sequence costs real compute, rejected guesses become waste, and every longer step delays the prefills behind it.
  6. With quantization and batching

    • Stacked: NVFP4 plus speculation: Quantization makes each step cheaper, which keeps batches small and steps memory-bound; that is exactly where speculation pays. The two multiply.
    • One config for day and night: Quantize to make steps cheap, speculate while batches are small, and cap speculation by batch size so the busy hour belongs to batching.
  7. Measuring it

    • Reading the acceptance counters: Mean acceptance length (1 + accepted ÷ drafts) is the number to watch; per-position rates show where guesses fail; accepted ÷ drafted is not α.
    • Drill: mean acceptance length: Tokens per verify step = 1 + accepted ÷ drafts; divide the plain step time by it (and add the draft cost) to estimate TPOT.
    • Break it: acceptance collapses: Acceptance is a property of the draft-target pair and the traffic. Change any of the three and re-measure; a silent drop shows up only as slower tokens.
  8. Recap & playground

    • Cheat sheet
    • Playground