learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

GPU Monitoring & Observability

Why nvidia-smi's GPU-Util says 100% for a GPU doing almost nothing, the DCGM profiling metrics that tell the truth, MFU, health signals, exporters and alerts, and finding the GPUs a cluster is wasting.

An interactive AI Infrastructure lesson: 23 steps, about 32 minutes, on a live simulation in your browser.

The VP of research asks one question on Monday: are the 32 H100s we pay for busy? The platform team opens the dashboard. It plots GPU-Util, the number nvidia-smi prints, and three of the four servers sit at 100%.

Three teams share the Kubernetes cluster. llm runs pretrain on dgx-1, vision runs vision on dgx-2, and research runs finetune, a Llama 3.1 70B fine-tune across dgx-3 and dgx-4. A new AMD server, mi300x-1, joined last week.

What you will learn

  1. The number on the dashboard

    • Is the cluster busy?: Allocated is not busy, and busy by GPU-Util is not busy doing maths. A GPU dashboard has to answer 'how much useful work', not 'is anything running'.
    • Two jobs at 100%: GPU-Util is a time fraction: 'was any kernel running?'. SM activity is a space fraction: 'how much of the chip was working?'. Only the second one moves when a job gets better or worse.
    • What GPU-Util actually counts: GPU-Util at 0% proves a GPU is idle; GPU-Util at 100% proves nothing.
  2. Metrics that do not lie

    • The DCGM profiling fields: Four numbers describe a GPU's work: SM activity (how much of the chip), occupancy (how full each SM is), tensor activity (how much is matrix maths) and DRAM activity (how hard memory is working).
    • Reading the pattern: Everything low means the GPU is starved (input, CPU, network); util high with SM low means it is busy waiting (communication, sync); SM high with tensor low means it is working inefficiently.
    • 100% busy, doing nothing: A hung collective is the purest GPU-Util lie: 100% busy, 2% SM activity, falling power. Alert on SM activity and power, never on GPU-Util.
    • What ten minutes of nothing cost: Every failure costs detection time + restart time + work since the last checkpoint. Monitoring shrinks the first term; checkpoints and fast restarts shrink the others.
  3. Job metrics and MFU

    • The job's own numbers: Tokens per second (or images, or samples) and step time are the job's heartbeat. Plot them per job, and alert when step time rises or the step counter stops.
    • Work out MFU: MFU = 6 × parameters × tokens per second ÷ (GPUs × peak FLOPS). 40–55% is good for large dense models; under 30% means something is wasting the cluster.
    • SM activity is not MFU: SM activity says the chip was occupied; MFU says the occupation trained the model. Communication, recomputation and padding raise the first and not the second.
  4. Health: power, heat, clocks, errors

    • Power, temperature, clocks: A GPU running below its boost clock always has a reason, and the driver tells you which one. Watch the clock event reasons, not the temperature alone.
    • One capped GPU, eight slow ones: In a synchronous job the slowest GPU sets the pace for all of them. One throttled GPU shows up as a slower step on every rank, and only its clocks and event reasons name it.
    • ECC and Xid: the error counters: Correctable errors are the warning; uncorrectable errors and Xids are the event. Track the rate of the first so you can drain the GPU before the second.
    • A health line per GPU: nvidia-smi --query-gpu with -l is a poor man's exporter: one CSV line per GPU per interval, easy to diff and graph.
  5. Exporters, Prometheus and alerts

    • The same questions on AMD: Every vendor has a local CLI (nvidia-smi / amd-smi / hl-smi), a daemon that reads deeper counters (DCGM on NVIDIA), and a Prometheus exporter. Build dashboards on the exporter's metrics, with a vendor label, and the questions stay the same.
    • DCGM exporter into Prometheus: DCGM exporter turns GPU counters into Prometheus series, and the Kubernetes pod mapping turns them into per-team, per-job series. Without that label you can see busy GPUs, but not whose.
    • Alert rules: page versus ticket: Page when work is being lost right now and a human can stop it; ticket when hardware is degrading; put everything else on a dashboard.
    • Xid 92 at 03:40: An Xid that kills work pages; an Xid that predicts failure opens a ticket. Know which is which before 3 a.m.
  6. Capacity: idle and wasted GPUs

    • Effective GPUs per team: Capacity has three layers: owned, allocated, and effectively used. Report all three per team; the gaps between them are where the money goes.
    • Feeding the starved job: When every counter is low, the fix is upstream of the GPU. More loader workers, faster storage or preprocessed data lift all the counters together.
    • Effective GPUs in one query: sum by (namespace) of SM activity counts effective GPUs; count by (namespace) of the same series counts allocated GPUs. Their ratio is the team's efficiency.
  7. Recap & playground

    • Cheat sheet
    • Playground