Serving Models
A model behind an API, answering in milliseconds all day: prefill and decode, the latency metrics users feel, batching as the lever for throughput, queues, replicas, tensor parallelism, cold starts and the serving stacks that do it.
An interactive AI Infrastructure lesson: 23 steps, about 32 minutes, on a live simulation in your browser.
Training produced a file of weights. Serving is everything that turns that file into an API: a process that holds the model in GPU memory, takes requests over HTTP, and answers in milliseconds, all day, for every user at once.
The support team wants a chatbot on Llama 3.1 8B. The first version is the simplest thing that works: one server process on one GPU of dgx-1, answering one request at a time. ("How an LLM Writes Text" shows what the model does with each one.)
What you will learn
One request, two phases
- A model behind an API: A served model is a long-running process that keeps the weights resident in GPU memory. Starting it costs seconds to minutes; answering must cost milliseconds.
- Follow one request: An LLM request is a short prefill that processes the whole prompt in one pass, then a long decode that emits one token per pass. Most of the wall time is decode.
- A prompt eight times longer: Prompt length drives time to first token; answer length drives total time. They are different costs and you tune them separately.
- Why decode is slow: bytes, not FLOPs: Decode time per step ≈ bytes read per step ÷ memory bandwidth. For one sequence that is the size of the weights, so a GPU with faster memory decodes faster, not one with more FLOPs.
The numbers users feel
- TTFT, TPOT, tokens/s and the p99: Averages hide queueing. Quote TTFT and TPOT as p50 and p99: the tail is where random arrivals meet a busy server.
- Drill: capacity without batching: Capacity = concurrency ÷ time per request. With one request at a time and 1.55 s per answer, one GPU serves 0.65 requests a second.
- Break it: one request a second: Without batching, a GPU serving an LLM is 100% busy and nearly idle at the same time: every decode step reads all the weights to serve a single user.
Batching is the lever
- Eight sequences per decode step: In a memory-bound decode, a batch of 8 costs about the same as a batch of 1. Batching multiplies throughput almost for free until compute or memory runs out.
- Static and dynamic: run to completion: Static and dynamic batching fix the batch when it starts. For LLMs, whose answers vary in length, that means waiting for the slowest sequence while slots sit empty.
- Continuous batching: Continuous batching schedules at every decode step: finished sequences leave, new ones join. It removes the wait for the slowest sequence, which is where static batching lost its latency.
- Latency versus throughput: Throughput grows with batch size while latency per token grows too. Pick the batch limit from the latency target, then buy throughput with more replicas.
Past capacity: queues and replicas
- Break it: 100 requests a second: Past capacity, TTFT climbs until the queue is full, then requests are refused. A queue absorbs a burst; it cannot absorb a sustained excess.
- More replicas behind one address: When a model fits on one GPU, scale out with data-parallel replicas: capacity grows linearly and each replica is an independent failure domain.
A model bigger than one GPU
- Break it: 70B on one GPU: Weights in GB ≈ parameters in billions × bytes per parameter. If that does not fit one GPU with room to spare, the model must be split (tensor parallel) or quantised before anything else matters.
- Split it over two GPUs: Fitting the weights is not enough: a serving GPU needs room for KV cache, and KV cache is what sets how many users it can hold at once.
- tp=4: room to work: Tensor parallelism divides the weights, and so the bytes each GPU reads per token, by tp. It buys memory and per-token speed with NVLink traffic, and it belongs inside one NVLink domain.
Loading and cold starts
- Break it: scaled to zero: A model server's cold start is pod start plus weights ÷ storage bandwidth. Scale-to-zero trades idle GPU cost for that delay on the first request.
- What a cold start costs: Cold start = pod scheduling + image pull + weights ÷ read bandwidth + engine warm-up. For big models the weights term dominates, so storage bandwidth sets how fast you can scale.
- Drill: weights over the wire: Load time = weight bytes ÷ read bandwidth. Halve the bytes (fp8) or double the bandwidth (local NVMe) and scale-up is twice as fast.
Serving stacks and the API
- The serving stacks, by role: An inference engine (vLLM, SGLang, TensorRT-LLM, TGI) runs the model fast; a model server or platform (Triton, NIM, KServe) deploys, routes and scales it. Job descriptions name both layers.
- An OpenAI-compatible API: The OpenAI-compatible API is the interface contract of LLM serving: messages in, streamed tokens out, token counts in usage. Keep it, and the engine behind it becomes replaceable.
Recap & playground
- Cheat sheet
- Playground