learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

The Accelerator Software Stack

Firmware, driver, CUDA or ROCm, the math and collective libraries, the framework and the container: what each layer does, how they depend on each other, and the version mismatches that stop a GPU job before it computes anything.

An interactive AI Infrastructure lesson: 22 steps, about 30 minutes, on a live simulation in your browser.

Your team fine-tunes Llama 3.1 8B with a PyTorch script, train.py. Nothing in it names a GPU vendor. Submit it to dgx-1, an 8 × H100 server, and it starts: 13.07 s per step, about 80,000 tokens a second.

Between that Python file and the silicon sit about seven layers of software, each written by a different team and each with its own version number. When all of them agree, nobody notices they exist. When two of them disagree, the job dies before its first step, and the error message names the layer that noticed, not the layer that is wrong.

What you will learn

  1. One script, three machines

    • One training script, one server: A GPU job runs on a stack of layers, and the job starts only when every layer agrees with the ones next to it.
    • Seven layers between Python and silicon: Firmware and driver belong to the machine; everything above them belongs to the job and can travel inside a container.
  2. Firmware and the driver

    • Firmware and the kernel driver: The driver is a kernel module owned by the host; nvidia-smi's "CUDA Version" is the newest CUDA the driver supports, not what is installed.
    • Upgrading the driver under a job: A driver change is a node operation: drain, upgrade, reboot or reload, validate, resume. Never under a running job.
  3. Compute platform, libraries, framework

    • The compute platform and its kernels: A kernel is a function the CPU launches onto the GPU; the compute platform (CUDA, ROCm/HIP, SynapseAI) is the compiler and runtime that builds and launches kernels.
    • Math libraries: the fast kernels: The speed of a training step lives mostly in the vendor's math libraries; the framework is the dispatcher that picks which library kernel to call.
    • The collective library: The collective library turns the server's wiring into bandwidth; busbw from nccl-tests is the number to compare against the link speed.
    • The framework hides the vendor: The framework keeps your code vendor-neutral; the vendor is chosen by which build of the framework you install, not by the script.
  4. Containers

    • Why AI runs in containers: The image brings the user space (CUDA runtime, cuDNN, NCCL, framework); the host brings the driver; the container toolkit connects the two at start-up.
    • The container toolkit goes missing: "could not select device driver … [[gpu]]" means Docker has no GPU hook: install or configure the container toolkit, not the driver.
    • Wiring GPUs into containers: NVIDIA containers get GPUs through a runtime hook (--gpus); AMD containers get them as plain device files (/dev/kfd, /dev/dri).
    • An NGC image on the AMD server: An image is built for one vendor's platform; the framework hides the vendor from your code, not from the image.
  5. The version chain

    • The compatibility chain: Compatibility points downward: a newer driver runs older CUDA, never the reverse; and on NVSwitch systems the fabric manager must equal the driver exactly.
    • A CUDA newer than the driver: "CUDA driver version is insufficient for CUDA runtime version" means the user-space CUDA is newer than the host driver: lower the CUDA or raise the driver.
    • Diagnose: read both version numbers: Diagnose a start-up failure by reading the versions of adjacent layers and checking them against the compatibility table, before touching anything.
    • CUDA error 802: the fabric manager: On NVSwitch systems, CUDA error 802 almost always means nvidia-fabricmanager is stopped or does not match the driver version.
    • Match the fabric manager, validate: Every driver change ends with a validation run on the node: nvidia-smi, an NCCL test, a short DCGM diagnostic, then back into service.
    • nvidia-smi on the AMD node: Each vendor's stack has its own tools; "command not found" for nvidia-smi on an AMD or Gaudi node is a question about the node, not a broken GPU.
  6. Drills, cheat sheet, playground

    • Drill: the driver version, script-ready: nvidia-smi --query-gpu=<fields> --format=csv,noheader turns nvidia-smi into a data source for scripts and health checks.
    • Drill: prove a container sees the GPUs: docker run --rm --gpus all <cuda image> nvidia-smi is the one-line smoke test of driver, toolkit and runtime together.
    • Cheat sheet
    • Playground