Troubleshooting GPU Clusters
A triage order for GPU clusters, reading Xid codes, and six real incidents diagnosed from their symptoms: a GPU off the bus, an NCCL hang, a thermal straggler, a flapping link, an uncorrectable ECC error and a half-finished driver upgrade.
An interactive AI Infrastructure lesson: 22 steps, about 35 minutes, on a live simulation in your browser.
The cluster runs one big job: llm, Llama 3.1 70B on all 32 H100s of four DGX servers, tensor parallel 8 inside each server and data parallel 4 across them. 34.25 s per step, 30,618 tokens/s, MFU 41%, a checkpoint every 20 steps.
When it breaks at 3 a.m., the temptation is to start with the most interesting theory. Resist it. Ask four questions in order, cheapest to check first: Is the hardware healthy? (GPUs, memory, heat) Is the fabric healthy? (NVLink, InfiniBand) Is the software stack consistent? (driver, fabric manager, CUDA, NCCL, the same everywhere) Is the job configured sanely? (parallelism, memory, environment).
What you will learn
A method before the incidents
- Four questions, in order: Hardware, fabric, software, job: check them in that order, because failures propagate upwards and never downwards.
- The triage script: A triage script is a checklist that runs itself: same commands, same order, every incident, so nobody skips the boring check that would have found it.
- Reading Xid codes: An Xid number sorts the fault: 13/31/43 point at the application, 48/63/64/92/94/95 at memory, 74 at NVLink, 79 at the GPU's connection to the host.
A GPU falls off the bus
- Xid 79: Xid 79 means the host lost the GPU on PCIe. It is always an incident, it always kills the job, and software cannot fix it.
- Confirming it on the node: Match the Xid's PCI address to a GPU index before doing anything else; every later step (drain reason, RMA, spare part) refers to that one GPU.
- Drain before you touch it: Drain first, repair second, resume last. A node under repair that the scheduler can still use will eat the next job.
- Power cycle, test, resume: Recover with the cheapest action that reaches the fault, prove it with a diagnostic, and count recurrences: the second identical failure on the same part is an RMA.
The job that stopped without an error
- A peer vanishes: A collective waits for every rank. When one rank disappears without an error, the others do not fail: they wait, and a hang looks like a busy job.
- What the log says: Silence plus 100% GPU-Util plus a frozen step counter is a hang. The watchdog timeout that ends it is the symptom's last line, not its cause.
- Finding the rank that left: In a hang, the missing rank is the one whose progress stopped first. Per-rank logs and NCCL_DEBUG=INFO turn a rank number into a host, a GPU and a NIC.
One hot GPU, 32 slow ones
- The job got slower: A synchronous job runs at the speed of its slowest GPU. A 34% slowdown on 32 GPUs can come from one GPU running 30% slow.
- Finding the straggler: Stragglers hide in averages. Compare each GPU's clock, temperature and event reasons with its peers', and the odd one out is the answer.
- Slow is better than stopped: Without spares, draining a straggler's node stops the whole job. Weigh 'slower' against 'stopped' before you drain.
A flapping link
- Symbol errors on one cable: A marginal link halves bandwidth long before it breaks. Test the fabric on a schedule; job step time will not tell you until it is too late.
- Then it goes down: A link that flaps kills whatever was using it and heals without healing the job. State says up; error counters say healthy. Check the second.
Bad memory, mismatched software
- An uncorrectable ECC error: An uncorrectable ECC error kills the job and schedules a row remap. The remap takes effect only at the next GPU reset, so a GPU with a pending remap is not fixed yet.
- Reset, reboot or RMA?: Reset for a pending remap, power-cycle for a GPU that is gone, RMA for a remap that failed or keeps recurring.
- A half-finished upgrade: On NVSwitch servers the driver and the fabric manager are one unit: same version, upgraded together. A mismatch fails every CUDA job on the node with error 802.
- Finding version drift: Version drift is found by comparing nodes, never by looking at one. Pin the whole stack as one release and check it before every job.
Reset, reboot or replace
- Reset, reboot or replace: Reset fixes state, reboot fixes connections, replacement fixes parts. Pick the cheapest one that reaches the fault, and escalate on the second occurrence.
- Cheat sheet
- Playground