learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

AI Networking: InfiniBand & RoCE

Why a training cluster needs a different network: RDMA, InfiniBand and RoCE, GPUDirect, rail-optimised fat trees, the bandwidth maths, and how one bad cable or one dead link stalls every GPU in the job.

An interactive AI Infrastructure lesson: 22 steps, about 32 minutes, on a live simulation in your browser.

Your team runs four DGX H100 servers, dgx-1 to dgx-4, and wants to train a 70B model across all 32 GPUs. Before you trust them with that, you measure the network. There are two.

The scale-up network lives inside one server: NVLink through NVSwitch (AMD's is Infinity Fabric). An NCCL all-reduce across dgx-1's 8 GPUs reaches a bus bandwidth of 360 GB/s per GPU: 450 GB/s each way, at the 80 % a real collective gets.

What you will learn

  1. Why AI traffic is different

    • Two networks in one cluster: Scale-up links GPUs inside a server (NVLink, Infinity Fabric, hundreds of GB/s); scale-out links servers (InfiniBand or Ethernet NICs). Collectives use both, and the slower one sets the pace.
    • Everyone talks at once: Training traffic is a few huge synchronised bursts. A collective finishes when its slowest link finishes, so the network is judged by its worst path, not its average.
    • Leaving the server: A DGX-class server has one 400 Gb/s NIC per GPU so that the scale-out network keeps up with NVLink for collectives. Take NICs away and the network becomes the bottleneck.
  2. RDMA and GPUDirect

    • RDMA: memory to memory: RDMA lets the NIC move data from one machine's memory into another's with no kernel and no CPU copies. That is how 400 Gb/s per NIC is reachable at all.
    • Drill: a server's network in GB/s: Divide network bits by 8 before comparing them with GPU bytes: 400 Gb/s is 50 GB/s, and eight of them are 400 GB/s per server.
    • GPUDirect RDMA: GPUDirect RDMA lets the NIC read and write GPU memory directly over PCIe, so the GPU's data never makes a trip through the CPU's memory.
    • Break it: GPUDirect switched off: Without GPUDirect RDMA nothing fails: NCCL stages through host memory and each rail roughly halves. Check the NCCL log for GDRDMA on every channel.
  3. InfiniBand and RoCE

    • InfiniBand: lossless by credit: InfiniBand is managed centrally by a subnet manager and never drops packets: a sender needs a credit, a promise of buffer space, before it transmits.
    • RoCE: RDMA over Ethernet: RoCE v2 is InfiniBand's transport over UDP/IP. It is only fast if the Ethernet under it is made lossless with PFC and ECN, end to end.
    • Which one, and when: InfiniBand buys a lossless, centrally managed fabric that works for collectives out of the box; RoCE buys Ethernet's vendors, tools and people at the price of careful lossless tuning.
  4. Rail-optimised fat trees

    • One leaf per rail: In a rail-optimised fabric NIC i of every server shares leaf i, so same-rank GPUs are one switch hop apart. The spines carry only what crosses rails.
    • Lose half a leaf's uplinks: Rail alignment is what protects collectives from the spine: as long as GPU i talks to GPU i, a spine problem is invisible to the all-reduce.
    • Oversubscription and growing up: Oversubscription is down-bandwidth over up-bandwidth. Web fabrics tolerate 3:1; a training fabric must be 1:1 anywhere a collective can cross.
  5. Bad cables and dead links

    • A link goes dark mid-step: A dead link under a running collective is a hang, not an error: every rank waits for the missing peer while the GPUs report 100 % busy.
    • Find it with ibstat: For a hung multi-node job, check the ports before the code: ibstat on each server, looking for any port that is not Active at full rate.
    • Break it: a cable that lies: A cable with errors does not fail; it slows every collective that crosses it, and the whole ring runs at its speed. Jobs that hide communication will not show it.
    • Find it with the tests: Find a slow link by narrowing: nccl-tests across servers says something is slow, ib_write_bw per rail says which NIC, the port counters say why.
    • Drill: test one rail: ib_write_bw takes a device and a peer: one NIC against one NIC, which is what isolates a single cable.
  6. DPUs and the road ahead

    • DPUs and SmartNICs: A DPU is a computer on the NIC: it runs the infrastructure (networking, storage, security) so the host's CPUs belong to the workload, and the tenant cannot touch it.
    • Ultra Ethernet and what comes next: Ultra Ethernet aims to give Ethernet InfiniBand-class behaviour for collectives by changing the transport (packet spraying, fast loss recovery) instead of forcing the network to be lossless.
  7. Recap & playground

    • Cheat sheet
    • Playground