GPU Monitoring & Observability
Why nvidia-smi's GPU-Util says 100% for a GPU doing almost nothing, the DCGM profiling metrics that tell the truth, MFU, health signals, exporters and alerts, and finding the GPUs a cluster is wasting.
An interactive AI Infrastructure lesson: 23 steps, about 32 minutes, on a live simulation in your browser.
The VP of research asks one question on Monday: are the 32 H100s we pay for busy? The platform team opens the dashboard. It plots GPU-Util, the number nvidia-smi prints, and three of the four servers sit at 100%.
Three teams share the Kubernetes cluster. llm runs pretrain on dgx-1, vision runs vision on dgx-2, and research runs finetune, a Llama 3.1 70B fine-tune across dgx-3 and dgx-4. A new AMD server, mi300x-1, joined last week.
What you will learn
The number on the dashboard
- Is the cluster busy?: Allocated is not busy, and busy by GPU-Util is not busy doing maths. A GPU dashboard has to answer 'how much useful work', not 'is anything running'.
- Two jobs at 100%: GPU-Util is a time fraction: 'was any kernel running?'. SM activity is a space fraction: 'how much of the chip was working?'. Only the second one moves when a job gets better or worse.
- What GPU-Util actually counts: GPU-Util at 0% proves a GPU is idle; GPU-Util at 100% proves nothing.
Metrics that do not lie
- The DCGM profiling fields: Four numbers describe a GPU's work: SM activity (how much of the chip), occupancy (how full each SM is), tensor activity (how much is matrix maths) and DRAM activity (how hard memory is working).
- Reading the pattern: Everything low means the GPU is starved (input, CPU, network); util high with SM low means it is busy waiting (communication, sync); SM high with tensor low means it is working inefficiently.
- 100% busy, doing nothing: A hung collective is the purest GPU-Util lie: 100% busy, 2% SM activity, falling power. Alert on SM activity and power, never on GPU-Util.
- What ten minutes of nothing cost: Every failure costs detection time + restart time + work since the last checkpoint. Monitoring shrinks the first term; checkpoints and fast restarts shrink the others.
Job metrics and MFU
- The job's own numbers: Tokens per second (or images, or samples) and step time are the job's heartbeat. Plot them per job, and alert when step time rises or the step counter stops.
- Work out MFU: MFU = 6 × parameters × tokens per second ÷ (GPUs × peak FLOPS). 40–55% is good for large dense models; under 30% means something is wasting the cluster.
- SM activity is not MFU: SM activity says the chip was occupied; MFU says the occupation trained the model. Communication, recomputation and padding raise the first and not the second.
Health: power, heat, clocks, errors
- Power, temperature, clocks: A GPU running below its boost clock always has a reason, and the driver tells you which one. Watch the clock event reasons, not the temperature alone.
- One capped GPU, eight slow ones: In a synchronous job the slowest GPU sets the pace for all of them. One throttled GPU shows up as a slower step on every rank, and only its clocks and event reasons name it.
- ECC and Xid: the error counters: Correctable errors are the warning; uncorrectable errors and Xids are the event. Track the rate of the first so you can drain the GPU before the second.
- A health line per GPU: nvidia-smi --query-gpu with -l is a poor man's exporter: one CSV line per GPU per interval, easy to diff and graph.
Exporters, Prometheus and alerts
- The same questions on AMD: Every vendor has a local CLI (nvidia-smi / amd-smi / hl-smi), a daemon that reads deeper counters (DCGM on NVIDIA), and a Prometheus exporter. Build dashboards on the exporter's metrics, with a vendor label, and the questions stay the same.
- DCGM exporter into Prometheus: DCGM exporter turns GPU counters into Prometheus series, and the Kubernetes pod mapping turns them into per-team, per-job series. Without that label you can see busy GPUs, but not whose.
- Alert rules: page versus ticket: Page when work is being lost right now and a human can stop it; ticket when hardware is degrading; put everything else on a dashboard.
- Xid 92 at 03:40: An Xid that kills work pages; an Xid that predicts failure opens a ticket. Know which is which before 3 a.m.
Capacity: idle and wasted GPUs
- Effective GPUs per team: Capacity has three layers: owned, allocated, and effectively used. Report all three per team; the gaps between them are where the money goes.
- Feeding the starved job: When every counter is low, the fix is upstream of the GPU. More loader workers, faster storage or preprocessed data lift all the counters together.
- Effective GPUs in one query: sum by (namespace) of SM activity counts effective GPUs; count by (namespace) of the same series counts allocated GPUs. Their ratio is the team's efficiency.
Recap & playground
- Cheat sheet
- Playground