Inside a GPU
Open one accelerator: streaming multiprocessors and compute units, tensor and matrix cores, HBM and the memory hierarchy, precisions from FP32 to FP4, power and clocks, and how to read nvidia-smi and amd-smi without being fooled.
An interactive AI Infrastructure lesson: 23 steps, about 35 minutes, on a live simulation in your browser.
Your team has a mixed fleet: H100 and B200 servers from NVIDIA, an AMD MI300X server and two L40S boxes. A bake-off is already running the same chat model on four of them. Before you pick hardware for the next model, open one GPU. Start a GPT-2 1.5B training job on dgx-1/gpu0 and look inside.
The grid is the H100's 132 streaming multiprocessors (SMs). An SM is the GPU's unit of work: it has its own scheduler, 128 FP32 lanes, 4 tensor cores, a 256 KB register file and 256 KB of on-chip memory it splits between L1 cache and shared memory. A kernel (one GPU function) is cut into thread blocks and the blocks are spread over the SMs.
What you will learn
Thousands of small workers
- One GPU, up close: A GPU is many identical streaming multiprocessors. Work is cut into blocks and dealt out to them, so a job is fast only if there is enough parallel work to keep all of them busy.
- The same idea on AMD: compute units: SM (NVIDIA) and CU (AMD) name the same thing: the repeated processor inside a GPU. Count them, multiply by what each can do per clock, and you have the chip's peak.
- Where the TFLOPS are: The TFLOPS on a datasheet are tensor-core TFLOPS. Work that does not run as matrix multiplies in a tensor-core format gets a small fraction of them.
HBM and the memory ladder
- HBM: 80 GB beside the chip: GPU memory is a ladder: registers, shared memory, L2, HBM. Each rung down is bigger and slower. HBM holds the model; the upper rungs decide how often you have to go back to it.
- One token, one read of HBM: Generating one token reads every weight from HBM once, so per-token time is roughly model bytes ÷ memory bandwidth.
Precision: bytes versus accuracy
- Bytes per number: A precision is a trade: half the bytes buys roughly double the tensor-core speed and half the memory traffic, and costs some accuracy that you must measure, not assume.
- Break it: 70B on one H100: Fitting the weights is not enough for serving: the KV cache needs room too, and a model that barely fits serves nobody.
- Quantise to go faster: For memory-bound serving, halving the bytes per weight nearly halves the time per token. Speed is measured by a stopwatch; quality has to be measured by evaluations.
Reading nvidia-smi and amd-smi
- nvidia-smi, column by column: nvidia-smi answers four questions per GPU at a glance: who is using it, how much memory, how much power against its limit, and how hot. It does not tell you how much maths is being done.
- amd-smi on the MI300X: Each vendor ships its own management tool with the same core readings: memory, utilisation, power, temperature, clocks. The tool that works tells you which vendor's driver is loaded.
- Drill: nvidia-smi for a script
Power, clocks and heat
- Boost clocks and the power limit: A GPU's clock is whatever its power and temperature limits allow. Clock event reasons tell you which limit is holding it back.
- Cap the power: Power caps cost speed less than proportionally: the last watts buy the fewest megahertz. Capping is how you fit more GPUs into a fixed power budget.
- Break it: a hot GPU: Heat does not break a GPU first; it slows it. A thermally throttled GPU shows as lower clocks and a clock event reason, not as an error.
Busy is not the same as working
- 100% GPU-Util, mostly idle: GPU-Util measures whether a kernel was running, not how much of the chip it used. 100% GPU-Util with low SM activity means memory-bound or underfilled, not full.
- What to watch instead: Judge a GPU by SM activity, tensor activity and memory activity together. GPU-Util alone only says whether something was running.
Comparing and choosing GPUs
- Which GPU holds 70B in BF16?: Compare GPUs on the whole working set per GPU: weights plus KV cache plus runtime. A GPU whose memory equals the weights cannot serve the model.
- The same request on four GPUs: For LLM decode, rank GPUs by memory bandwidth; for training and prefill, by tensor-core TFLOPS at the precision you use. The datasheet has both, and the workload tells you which to read.
- Five accelerators, one table: Each generation moves three numbers at different rates: TFLOPS, memory capacity and memory bandwidth, with power rising alongside. Choosing a GPU means knowing which of them your workload is limited by.
- Picking a GPU for a job: Pick an accelerator by asking, in order: does it fit, is the job compute- or bandwidth-bound, does it span GPUs, does the software stack support it, and can you power, cool and buy it.
- Drill: tokens per second from bandwidth
Recap & playground
- Cheat sheet
- Playground