LLM Inference at Scale
What limits an LLM service at scale and the tools that lift each limit: the KV cache and paged attention, continuous batching under load, quantization, speculative decoding, prefix caching, long contexts, autoscaling on the right signal, cost per million tokens and a p99 incident.
An interactive AI Infrastructure lesson: 22 steps, about 35 minutes, on a live simulation in your browser.
The platform team runs the company assistant: Llama 3.1 70B in bf16, tensor parallel 4, on half of dgx-1. Prompts average 1,024 tokens, answers 256. "Serving Models" covered how one request runs; this lesson is about running thousands.
Open the memory view. Each GPU holds 35 GB of weights. vLLM takes 90% of the 80 GB, subtracts weights, workspace and overhead, and turns the rest into KV cache: 399,169 tokens for the replica.
What you will learn
The KV cache sets concurrency
- A 70B assistant on four GPUs: After the weights, what is left of GPU memory is KV cache, and KV cache is concurrency: every token of every request in flight lives there until the request ends.
- KV cache arithmetic: KV bytes per token = 2 × layers × KV heads × head dim × bytes. Multiply by tokens in flight and compare with what is left after the weights: that ratio is your concurrency.
- Drill: KV bytes per token: An 8B model with GQA costs 128 KiB per token; a 4,000-token conversation is half a gigabyte of HBM for as long as it is in flight.
- Break it: reserving the maximum: Contiguous KV allocation reserves the worst case for every request. Memory is full of promises, and concurrency is capped at capacity ÷ max length.
- Paged attention: Paged KV allocates cache in small blocks as sequences grow. Waste drops from most of the cache to under one block per sequence, and concurrency rises until compute is the limit.
Fewer bytes per token
- Twice the traffic: At saturation, a decode step's cost is weights plus the batch's KV cache, read from HBM, plus any prefill work scheduled into it. Cut the bytes and the step gets cheaper for everyone.
- Quantize the weights to fp8: Quantization is a memory-bandwidth optimisation first: half the bytes per weight is nearly half the time per decode step, and the memory it frees becomes KV cache.
- int8, 4-bit and the accuracy trade: Fewer bits per weight means fewer bytes per decode step and more room for KV cache, paid for in accuracy that you must measure on your own tasks.
- Speculative decoding: Speculative decoding turns spare compute into fewer memory-bound steps. The gain is set by the acceptance rate, and the output distribution is unchanged.
- Break it: speculation at saturation: Speculation spends idle compute. When the batch is big enough to keep the tensor cores busy there is no idle compute left, and rejected guesses become pure waste.
Shared prefixes and long contexts
- Break it: a 2,500-token system prompt: A shared prompt prefix is identical work repeated for every request. Without caching you pay its full prefill each time.
- Prefix caching: Prefix caching reuses the KV blocks of identical leading tokens. Stable content first, variable content last: the order of your prompt decides your hit rate.
- What a long context costs: KV cache grows linearly with context and prefill faster than linearly (attention is quadratic): one 100K-token request occupies the memory of about eighty 1,280-token conversations.
Autoscaling on the right signal
- Break it: scaling on GPU-Util: GPU-Util says the GPU is doing something, not how much more it could do. An LLM server is 'busy' at any load, so GPU-Util can only scale it up.
- Scale on the queue: Scale an LLM service on queue depth or KV cache use, and remember the scale-up delay is pod start plus weights ÷ bandwidth. Keep enough warm headroom to cover that window.
Cost and the frontier
- Cost per million tokens: Cost per million tokens = hourly GPU cost ÷ (tokens/s × 3,600) × 10⁶. The same model on the same GPUs is ten times cheaper per token at ten times the batch.
- Drill: price a replica: Price a deployment from two measured numbers: GPU-hours and sustained tokens per second at your latency target.
- Disaggregation and multi-node: Separate what has different bottlenecks: prefill from decode across pools, tensor parallel inside a server, pipeline parallel across servers.
Incident: p99 TTFT
- Incident: TTFT p99 at 36 seconds: Diagnose LLM latency from the serving metrics: queue depth, KV cache use and tokens per request. Request rate and GPU-Util can both look normal while the service drowns.
- The fix: Separate workloads whose shapes differ (short chat, long documents, batch jobs) into pools with their own limits and scaling. One pool tuned for all of them serves none of them well.
Recap & playground
- Cheat sheet
- Playground