learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Inside a GPU

Open one accelerator: streaming multiprocessors and compute units, tensor and matrix cores, HBM and the memory hierarchy, precisions from FP32 to FP4, power and clocks, and how to read nvidia-smi and amd-smi without being fooled.

An interactive AI Infrastructure lesson: 23 steps, about 35 minutes, on a live simulation in your browser.

Your team has a mixed fleet: H100 and B200 servers from NVIDIA, an AMD MI300X server and two L40S boxes. A bake-off is already running the same chat model on four of them. Before you pick hardware for the next model, open one GPU. Start a GPT-2 1.5B training job on dgx-1/gpu0 and look inside.

The grid is the H100's 132 streaming multiprocessors (SMs). An SM is the GPU's unit of work: it has its own scheduler, 128 FP32 lanes, 4 tensor cores, a 256 KB register file and 256 KB of on-chip memory it splits between L1 cache and shared memory. A kernel (one GPU function) is cut into thread blocks and the blocks are spread over the SMs.

What you will learn

  1. Thousands of small workers

    • One GPU, up close: A GPU is many identical streaming multiprocessors. Work is cut into blocks and dealt out to them, so a job is fast only if there is enough parallel work to keep all of them busy.
    • The same idea on AMD: compute units: SM (NVIDIA) and CU (AMD) name the same thing: the repeated processor inside a GPU. Count them, multiply by what each can do per clock, and you have the chip's peak.
    • Where the TFLOPS are: The TFLOPS on a datasheet are tensor-core TFLOPS. Work that does not run as matrix multiplies in a tensor-core format gets a small fraction of them.
  2. HBM and the memory ladder

    • HBM: 80 GB beside the chip: GPU memory is a ladder: registers, shared memory, L2, HBM. Each rung down is bigger and slower. HBM holds the model; the upper rungs decide how often you have to go back to it.
    • One token, one read of HBM: Generating one token reads every weight from HBM once, so per-token time is roughly model bytes ÷ memory bandwidth.
  3. Precision: bytes versus accuracy

    • Bytes per number: A precision is a trade: half the bytes buys roughly double the tensor-core speed and half the memory traffic, and costs some accuracy that you must measure, not assume.
    • Break it: 70B on one H100: Fitting the weights is not enough for serving: the KV cache needs room too, and a model that barely fits serves nobody.
    • Quantise to go faster: For memory-bound serving, halving the bytes per weight nearly halves the time per token. Speed is measured by a stopwatch; quality has to be measured by evaluations.
  4. Reading nvidia-smi and amd-smi

    • nvidia-smi, column by column: nvidia-smi answers four questions per GPU at a glance: who is using it, how much memory, how much power against its limit, and how hot. It does not tell you how much maths is being done.
    • amd-smi on the MI300X: Each vendor ships its own management tool with the same core readings: memory, utilisation, power, temperature, clocks. The tool that works tells you which vendor's driver is loaded.
    • Drill: nvidia-smi for a script
  5. Power, clocks and heat

    • Boost clocks and the power limit: A GPU's clock is whatever its power and temperature limits allow. Clock event reasons tell you which limit is holding it back.
    • Cap the power: Power caps cost speed less than proportionally: the last watts buy the fewest megahertz. Capping is how you fit more GPUs into a fixed power budget.
    • Break it: a hot GPU: Heat does not break a GPU first; it slows it. A thermally throttled GPU shows as lower clocks and a clock event reason, not as an error.
  6. Busy is not the same as working

    • 100% GPU-Util, mostly idle: GPU-Util measures whether a kernel was running, not how much of the chip it used. 100% GPU-Util with low SM activity means memory-bound or underfilled, not full.
    • What to watch instead: Judge a GPU by SM activity, tensor activity and memory activity together. GPU-Util alone only says whether something was running.
  7. Comparing and choosing GPUs

    • Which GPU holds 70B in BF16?: Compare GPUs on the whole working set per GPU: weights plus KV cache plus runtime. A GPU whose memory equals the weights cannot serve the model.
    • The same request on four GPUs: For LLM decode, rank GPUs by memory bandwidth; for training and prefill, by tensor-core TFLOPS at the precision you use. The datasheet has both, and the workload tells you which to read.
    • Five accelerators, one table: Each generation moves three numbers at different rates: TFLOPS, memory capacity and memory bandwidth, with power rising alongside. Choosing a GPU means knowing which of them your workload is limited by.
    • Picking a GPU for a job: Pick an accelerator by asking, in order: does it fit, is the job compute- or bandwidth-bound, does it span GPUs, does the software stack support it, and can you power, cool and buy it.
    • Drill: tokens per second from bandwidth
  8. Recap & playground

    • Cheat sheet
    • Playground