learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

GPUs on Kubernetes

Why Kubernetes cannot see a GPU until a device plugin advertises it, what the NVIDIA and AMD GPU operators install, how pods ask for accelerators, how to steer them to the right model, and how to share one GPU with time-slicing, MPS and MIG.

An interactive AI Infrastructure lesson: 23 steps, about 32 minutes, on a live simulation in your browser.

Your platform team has joined four DGX H100 servers, dgx-1 to dgx-4, to a Kubernetes 1.35 cluster. Log in to dgx-1 and nvidia-smi lists eight H100s with 81,560 MiB each: the driver is loaded and the hardware is fine.

Now ask Kubernetes. kubectl describe node dgx-1 lists 224 CPUs, 2 TiB of memory and 110 pods under Capacity. Not one GPU. As far as the scheduler knows, this is a very large CPU server.

What you will learn

  1. Kubernetes cannot see a GPU

    • Eight H100s the cluster cannot see: The kubelet only counts CPU and memory by itself. An accelerator exists for the scheduler only once something advertises it as a resource.
    • A one-GPU pod on a GPU cluster: "Insufficient nvidia.com/gpu" on a cluster full of GPUs means nothing is advertising them: check the device plugin before the hardware.
    • The device plugin and extended resources: A GPU becomes schedulable through a vendor's device plugin, which advertises an extended resource such as nvidia.com/gpu; the pod asks for that name and a whole number.
  2. The GPU operator

    • The GPU operator rolls out: The GPU operator turns a GPU node into a Kubernetes GPU node: driver, container toolkit, device plugin, labels, metrics and MIG manager, all as DaemonSets.
    • Capacity, Allocatable and the labels: kubectl describe node answers three questions for GPUs: does the node advertise them (Capacity), how many are free (Allocatable minus Allocated), and what are they (the feature-discovery labels).
    • Break it: scheduled, but no GPU hook: Pending means the scheduler found no GPU; CreateContainerError on a GPU pod means the node advertised a GPU its runtime cannot hand over.
  3. Asking for GPUs

    • Whole GPUs, limits equal requests: A GPU request is a whole number, set as a limit, never overcommitted, and satisfied on a single node.
    • Drill: two GPUs for one pod
  4. The right GPU, and only GPU work

    • An AMD server joins: The resource name carries the vendor: nvidia.com/gpu, amd.com/gpu and habana.ai/gaudi are different resources, so a pod's request already decides whose hardware it can use.
    • 148 GB on an 80 GB GPU: The scheduler matches a count and labels, never your job's memory: making a job land on a GPU big enough is your job, done with the resource name and selectors.
    • Target the GPU model with labels: Feature-discovery labels make GPU models selectable: nodeSelector for one model, node affinity for a list or a preference.
    • Drill: keep CPU work off GPU nodes: Taints keep everything else off the expensive nodes; selectors and affinity steer GPU pods onto the right ones. You need both.
  5. Sharing a GPU: time-slicing and MPS

    • Time-slicing: one GPU, four tickets: Time-slicing advertises each GPU N times and lets the pods take turns: more pods fit, each runs slower, and nothing separates their memory.
    • Break it: the third notebook: Time-sliced pods share one pool of GPU memory and one fault domain: one tenant's allocation or crash is every tenant's problem.
    • MPS: concurrent, still one fault domain: MPS runs several processes' kernels concurrently with per-client memory and SM limits, but they still share one fault domain.
  6. MIG: hardware partitions

    • MIG: seven GPUs inside one: MIG cuts a GPU into hardware-isolated instances with their own SMs and memory; each profile is advertised as its own Kubernetes resource.
    • Two 4g.40gb on one H100?: A MIG layout must fit two budgets at once, 7 compute slices and 8 memory slices, plus each profile's own instance limit.
    • Break it: repartition a busy GPU: Changing a MIG layout needs an idle GPU: repartitioning is a maintenance operation, not a live resize.
    • Pods on MIG instances: A MIG instance runs at its share of the GPU all the time: isolation and predictability, paid for by giving up idle capacity next door.
    • AMD: compute and memory partitions: AMD Instinct partitions a GPU by dies: SPX/DPX/QPX/CPX for compute and NPS1/NPS4 for memory, CPX + NPS4 giving eight 24 GB GPUs on an MI300X.
    • Choosing, and asking for, a slice: Pick sharing by trust and predictability: time-slicing for trusted and bursty, MPS for cooperative concurrency, MIG for isolation, whole GPUs for training.
  7. Recap & playground

    • Cheat sheet
    • Playground