Quantization
Number formats from FP32 to NVFP4, what quantization shrinks and why that speeds decode, weight-only versus weight-and-activation schemes, scales, GPTQ, AWQ, SmoothQuant and FP8, measuring accuracy, which GPUs support which format, and the tools: the same 70B model served in several precisions, on fewer GPUs per replica.
An interactive AI Infrastructure lesson: 22 steps, about 35 minutes, on a live simulation in your browser.
The support team's copilot runs Llama 3.1 70B in bf16 on dgx-1: tensor parallel 4, so each replica takes four H100s and each GPU holds 35 GB of weights. One request streams at 14.49 ms per token.
The copilot is going to three more regions. At four GPUs a replica, that is twelve more H100s, and finance has asked whether it can be done with fewer.
What you will learn
The bill
- Four GPUs per replica: A 70B model in bf16 is 141 GB of weights. Every serving decision starts from that number divided by the GPUs you split it over.
- Decode is paid in bytes: Decode time per step ≈ weight bytes per GPU ÷ memory bandwidth. Quantization attacks the numerator.
Number formats
- Bits, range and precision: Exponent bits buy range, mantissa bits buy precision, and below 8 bits nothing works without a scale shared by a block of values.
- Drill: a 4-bit 70B: Weights in GB = parameters in billions × bytes per parameter: 70.6 × 2 = 141, × 1 = 71, × 0.5 = 35.
Fewer GPUs per replica
- fp8 on half the GPUs: Quantization buys either speed or GPUs: half the bytes per weight is half the step time on the same GPUs, or the same step time on half the GPUs.
- The same service on two GPUs: A W8A8 fp8 replica on N GPUs does the work of a bf16 replica on 2N: half the bytes for decode, double the FLOPS for prefill.
- Two replicas on the same four GPUs: Cost per token follows GPUs per replica. Quantization that halves the GPUs a replica needs halves the bill for the same traffic and latency.
- Break it: fp8 on one GPU: A model is servable on a GPU only if weights plus workspace leave room for the KV cache of real requests, not merely if the weights fit.
Weights, activations, KV cache
- int8 weights, and the KV cache: Weights, activations and the KV cache are three separate quantization choices. Weight bits set step time; KV bits set how many tokens fit.
- Weight-only: 70B on one H100: Weight-only quantization (W4A16, W8A16) shrinks the bytes decode reads and nothing else: the maths runs at the activation precision.
- Break it: W4A16 under daytime load: Weight-only quantization helps memory-bound work. Under load, prefill and big decode batches are compute-bound, and W4A16 does that maths in bf16 on however few GPUs you left it.
- W8A8: quantize the activations too: Weight-only quantization makes the reads cheaper; W8A8 also makes the maths cheaper. Prefill-heavy and high-batch services need the second.
Methods and accuracy
- Scales, groups and outliers: A scale is shared by everything it covers, so its largest member decides everyone's precision. Smaller groups and per-channel scales keep one outlier from erasing its neighbours.
- Drill: a group scale: Symmetric scale = max |value| ÷ largest level (7 for INT4, 127 for INT8, 448 for FP8 E4M3): the biggest value maps to the top level.
- GPTQ, AWQ, SmoothQuant, FP8: GPTQ corrects rounding error as it goes, AWQ protects the weights that matter, SmoothQuant moves activation outliers into the weights; all three learn from calibration data that must look like production.
- Accuracy: measure, do not assume: Quantization is accepted on evals, not on perplexity: same harness, baseline versus quantized, on tasks that look like your product, and maths and code first.
Hardware decides
- Break it: the fp8 checkpoint on an A100: A format is only as fast as the tensor cores that execute it. FP8 maths needs Hopper, Ada, Blackwell or MI300; on an A100 an fp8 checkpoint is weight-only at best.
- Four bits on one Blackwell GPU: Bits only turn into speed when the tensor cores compute in those bits. Native FP4 on Blackwell makes 4-bit fast for prefill and decode; on Hopper it only shortens decode.
- AMD: one MI300X, then fp8: Capacity is a quantization substitute: a 192 GB GPU holds a bf16 70B whole. Quantize there for speed and concurrency, not to fit.
- Drill: fp8 at load time: --quantization sets how the weights (and activations) are stored and computed; --kv-cache-dtype sets the cache. A pre-quantized checkpoint declares its own scheme and needs neither flag for the weights.
Recap & playground
- Cheat sheet
- Playground