Bring-up & Validation
From delivered racks to a cluster accepted for production: out-of-band access and firmware, the software stack in the right order, DCGM diagnostics and HPL burn-in, per-rail and NCCL tests, written acceptance criteria, and the first real job.
An interactive AI Infrastructure lesson: 23 steps, about 32 minutes, on a live simulation in your browser.
Four DGX H100 servers, dgx-1 to dgx-4, are racked, powered and cabled to eight InfiniBand leaves. The research team wants to start a Llama 3.1 70B run on all 32 GPUs on Monday. Your job between now and then is bring-up: turning delivered hardware into a cluster you are willing to sign for.
Right now the servers have an operating system and nothing else. nvidia-smi on dgx-1 cannot talk to a driver, because there is none. Slurm already knows the nodes, and all four sit in drain with the reason bring-up, so no user job can land on hardware nobody has tested.
What you will learn
Racks on the floor
- Four servers and a deadline: A GPU node is not in production when it boots. It is in production when it has passed tests you wrote down in advance, and until then the scheduler keeps it drained.
- Out of band before anything else: The BMC is how you recover a node that cannot recover itself. Check it on day one, while the node is healthy, not on the night you need it.
- Same firmware on every node: Pick one version of everything and make every node match it. A cluster where each node is slightly different turns every bug into a question about which node you were on.
The stack, in order
- Install the stack bottom-up: Install bottom-up: driver, fabric manager, CUDA, container toolkit, DCGM, then the scheduler. Each layer checks the one below it, so a mistake low down shows up as a confusing error high up.
- One node, one wrong package: nvidia-smi proves the driver loaded, nothing more. Only a real CUDA program proves the whole stack works together.
- Audit versions across every node: Version drift is found by comparing every node to a written baseline, not by looking at one node and assuming the others match.
One node at a time
- nvidia-smi: count, driver, memory, ECC: The first test is the count. Eight GPUs expected, eight listed, same memory, same driver, zero uncorrected errors; anything else stops the bring-up of that node.
- nvidia-smi topo -m: is it wired right?: topo -m is the server's wiring diagram as the driver sees it. Every node of the same model must print the same matrix; a difference is a hardware fault or a misplaced card.
- dcgmi diag: quick, medium, long: dcgmi diag -r 1 is a smoke test, -r 2 a health check, -r 3 a stress test. Only the levels that load the GPU can find parts that fail under load.
- Burn-in: HPL at full power for hours: Burn-in turns early-life failures into bring-up findings. HPL gives you two things at once: a stress load, and a performance number every identical node must match.
Burn-in finds the weak part
- A GPU that fails only under load: A GPU can pass every configuration and health check and still be broken. Faults that appear only under sustained load are found only by sustained load.
- What does the burn-in report?: In a synchronous computation the slowest GPU sets the pace for all of them. Compare every node's HPL to the reference; a node at half speed has one bad part, not eight mediocre ones.
- Drain, replace, burn in again: A node that failed burn-in is drained with a reason, repaired, and burned in again from the start. Repairs are a new source of faults, not the end of them.
Many nodes: rails and NCCL
- ib_write_bw, one rail at a time: Test the fabric one rail at a time with a raw RDMA tool before running collectives. A bad link is obvious alone and invisible inside an eight-rail average.
- A cable that works, badly: A link with errors stays up and runs slowly. Link state tells you it is cabled; only bandwidth tests and error counters tell you it is good.
- NCCL across two nodes: what to expect: Before any NCCL test, write down the busbw the hardware should give: the slower of NVLink inside the server and rails × NIC speed across servers. Then the test is a comparison, not a number to admire.
- All four nodes, one bad cable: In a ring, one slow link is a slow cluster. Multi-node NCCL tests exist to find the one link in hundreds that is not like the others.
Acceptance and hand-over
- Acceptance criteria, written down: Acceptance criteria are numbers agreed before the tests run. Without them every result is "looks fine", and a half-speed node gets signed off.
- Drill: what should the fabric give?: Cross-node busbw per server is rails × NIC bytes per second × efficiency: 8 × 50 GB/s × 0.92 = 368 GB/s on a DGX H100 with NDR.
- Drill: the long diagnostic: dcgmi diag -r 3 is the acceptance and RMA test: it is the level that loads the GPU with power and stress, and the log vendors ask for.
- Hand-over: the first real job: Hand-over ends with a real workload meeting its predicted throughput. Acceptance tests prove the parts; the first job proves the system.
Recap & playground
- Cheat sheet
- Playground