What AI Workloads Need
A model is billions of numbers and a lot of matrix multiplication. Training learns them over weeks; inference uses them in milliseconds. Three numbers decide what hardware either needs: FLOPs, memory capacity and memory bandwidth.
An interactive AI Infrastructure lesson: 21 steps, about 30 minutes, on a live simulation in your browser.
Your company wants its own chat assistant, and you have been handed dgx-1: one server with eight NVIDIA H100 GPUs. The data science team says they will use Llama 3.1 8B. To an infrastructure engineer, that model is a file of 8.03 billion numbers called parameters (or weights), arranged as large matrices.
Each parameter is stored in bf16, a 2-byte number format. So the model is 8.03 billion × 2 bytes = 16 GB, and it has to sit in GPU memory to be used. Ask the engine for a plan on one GPU: 16 GB of weights plus a little working space comes to 19 GB of 80 GB. It fits.
What you will learn
A model is a pile of numbers
- A model is a pile of numbers: A model is its parameters: billions of numbers in matrices. Its size in bytes is parameters × bytes per number, and using it costs about 2 × parameters FLOPs per token.
- AI, machine learning, deep learning: AI is the goal, machine learning is learning from examples, deep learning is doing it with many-layered neural networks. To the hardware, all of it is matrix multiplication.
Learning the numbers
- Training needs far more memory: Training with Adam costs 16 bytes per parameter before activations, eight times the 2 bytes that serving needs. A model you can serve on one GPU usually cannot be trained on one.
- Break it: train it anyway: Out of memory is decided by arithmetic before the job starts. If the plan does not fit, no amount of retrying will make it fit.
- A training run you can watch: Training is a batch job: 6 × parameters × tokens of maths, run flat out for hours to months, judged by tokens per second rather than by the latency of any one step.
- Training at real scale: Large-model training needs many GPUs for two separate reasons: its state does not fit in one GPU's memory, and its FLOPs would take one GPU centuries.
- Drill: bytes per parameter to train
Using the numbers
- Serve the model: Inference is a service, not a batch job: load the weights once, keep them in GPU memory, and answer requests as they come.
- Follow one request: A request is one prefill (all prompt tokens at once, maths-heavy) and then one decode step per output token (each reads every weight). Long answers cost far more than long questions.
- Training and inference side by side: Training is throughput-bound and scheduled; inference is latency-bound and on demand. The same GPU can do either, but it is sized, scheduled and monitored differently for each.
Why GPUs
- A CPU and a GPU, compared: A CPU is a few fast, clever cores for varied work; a GPU is thousands of simple lanes doing the same maths on different data, fed by far more memory bandwidth. Neural networks are the second kind of work.
- The same prompt on the CPUs: For maths-heavy work the question is FLOPs ÷ FLOPS. A GPU has tens of times a server CPU's peak, so the same prompt, step or batch takes tens of times less time.
The three numbers
- Number one: compute: Time for maths-heavy work = FLOPs needed ÷ (peak FLOPS × efficiency). MFU is that efficiency, measured.
- Break it: number two, capacity: Capacity is a gate: if the weights plus working memory exceed the GPU's memory, nothing runs, however fast the chip is.
- Number three: memory bandwidth: When each byte read is used for only a little maths, time = bytes read ÷ memory bandwidth. Single-user decode is the classic case: one token costs one full read of the weights.
- Maths or memory, never both: Every workload is limited by compute or by memory bandwidth, whichever takes longer. FLOPs per byte read decides which, and batching is how inference moves from the memory side toward the compute side.
- Drill: will it fit?
Where AI runs, and who runs it
- Where AI runs: Inference runs wherever users are, from phones to data centres; large-scale training runs only where thousands of GPUs share one fast network. Owning pays when the GPUs stay busy.
- Who runs it, and what comes next: AI infrastructure work is answering five questions with numbers: does it fit, what limits its speed, how many GPUs, why did it fail, and what does it cost.
Recap & playground
- Cheat sheet
- Playground