learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Anatomy of a GPU Server

What is inside an 8-GPU server from NVIDIA or AMD and a PCIe L40S box: CPUs, memory, PCIe switches, the scale-up fabric, one NIC per GPU, the BMC and the power, and how reading the topology tells you which workloads a server will run well.

An interactive AI Infrastructure lesson: 21 steps, about 30 minutes, on a live simulation in your browser.

Your team is buying capacity and three quotes are on the table. The cluster already has one of each, so open them up, starting with dgx-1, a DGX H100 (NVIDIA's own build of the HGX H100 board that other vendors also sell).

Inside: 2 CPUs (Intel Xeon, 56 cores each, 224 threads in all) with 2 TB of system memory; 8 H100 SXM GPUs with 80 GB of HBM3 each, mounted on one baseboard with 4 NVSwitch chips; PCIe Gen5 switches that join each GPU to its CPU, its NIC and its NVMe drives; 8 ConnectX-7 NICs at 400 Gb/s, one per GPU, for the compute network, plus separate ports for storage and management; about 30 TB of local NVMe; and a BMC for remote management.

What you will learn

  1. What is in the box

    • An 8-GPU server, opened: An AI server is a GPU baseboard (GPUs plus their scale-up fabric) bolted to an ordinary two-socket host that feeds it: CPUs, memory, PCIe, NICs, disks and a BMC.
    • Three servers, three designs: Servers with the same GPU count can differ completely in how the GPUs connect to each other and to the network; that wiring, not the GPU count, sets what they are good for.
  2. The scale-up fabric

    • NVLink and NVSwitch: all to all: NVSwitch makes the scale-up fabric all to all: every GPU pair gets the full 450 GB/s per direction, at the same time.
    • Infinity Fabric: point to point: A point-to-point mesh has the same total bandwidth as a switch but gives any single pair only one link; it shines when all GPUs talk at once.
    • No fabric: the PCIe server: Without a scale-up fabric, GPUs talk over PCIe and through the CPUs: fine for work that stays on one GPU, slow for anything that splits a model across GPUs.
  3. Reading the topology

    • Reading nvidia-smi topo -m: topo -m ranks every device pair by what the traffic crosses, NV# best, then PIX, PXB, PHB, NODE, and SYS worst; each step down costs bandwidth and latency.
    • One NIC per GPU: One NIC per GPU, on the same PCIe switch, gives every GPU its own full-speed path to the network; the NIC is the rail that GPU rides.
    • NUMA: pin the process to the GPU's socket: Run each GPU's process on the CPU socket and memory that topo -m lists for that GPU; the wrong socket costs speed silently.
  4. The bandwidth ladder

    • Every level, a number: Per H100 the ladder is HBM 3,350, NVLink 450, PCIe 64, NIC ~50 GB/s: roughly sevenfold drops going outward. Put the chattiest traffic on the fastest level.
    • What the ladder means for a job: Communication only costs step time when it cannot hide behind compute; the faster the link, the more of it hides.
  5. BMC and power

    • The BMC: out-of-band management: The BMC is a separate computer on its own network that can power-cycle, monitor and reinstall the server when the OS is unreachable.
    • A GPU falls off the bus: Xid 79 means the GPU is gone from PCIe: no in-band tool can reach it, so recovery needs a power cycle, and a repeat means an RMA.
    • Power-cycle through the BMC: Out-of-band power control is the recovery path of last resort; a server whose BMC is down has lost it.
    • Power: what a whole server draws: An 8-GPU H100 server draws about 8 kW flat out and is provisioned at 10.2 kW; rack power, not floor space, limits how many you can install.
  6. Design decides the workload

    • One all-reduce, three servers: An all-reduce runs at the speed of its slowest hop: NVSwitch ~360, Infinity Fabric ~313, PCIe across sockets ~15 GB/s bus bandwidth.
    • Tensor parallel over PCIe: Tensor parallelism puts an all-reduce inside every layer, so it belongs inside a scale-up fabric; over PCIe it spends most of the step waiting.
    • Match the workload to the server: Choose the server by the workload's communication pattern and memory need: split models need a scale-up fabric, big single-GPU models need big HBM, independent replicas need neither.
  7. Drills, cheat sheet, playground

    • Drill: print the wiring matrix: nvidia-smi topo -m is the first command on an unfamiliar GPU server: it shows the fabric, the NIC rails and the NUMA layout in one table.
    • Drill: pin GPU 2's process: Pair CUDA_VISIBLE_DEVICES with numactl: the first picks the GPU, the second puts the process on that GPU's socket.
    • Cheat sheet
    • Playground