learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Scheduling GPU Clusters

Why GPU jobs need all-or-nothing scheduling: Slurm's partitions, backfill, draining and priorities; the deadlock Kubernetes' default scheduler walks into and the gang scheduling of Kueue and Volcano that prevents it; quotas, borrowing and preemption; Ray on Kubernetes; and when to pick Slurm or Kubernetes.

An interactive AI Infrastructure lesson: 21 steps, about 32 minutes, on a live simulation in your browser.

Two teams share this cluster: research and product, 32 H100s in four DGX servers under Kubernetes. Research starts pretrain, a Llama 3.1 70B run: tensor parallel 8 inside each server, data parallel 4 across them, four pods of eight GPUs.

Each step takes 34.28 s at 41% MFU, and every one of those seconds needs all 32 GPUs: every rank computes a slice, then waits in an all-reduce for the other 31. A web service with 10 of its 12 replicas is degraded. A training job with 31 of its 32 GPUs does not run at all.

What you will learn

  1. Why GPU scheduling is different

    • Thirty-two GPUs or nothing: A distributed training job is one unit of work: it needs every GPU it asked for at the same time, or it does nothing.
    • Take one server away: Evicting one pod of a training job kills the job: on GPU clusters, maintenance is scheduled with the scheduler, not around it.
  2. Slurm: the HPC way

    • Slurm: partitions, sbatch and srun: Slurm hands a job its whole allocation at once or not at all; partitions are the queues and GRES counts the GPUs.
    • An idle server and a waiting job: Strict FIFO protects the big job at the head of the queue by leaving everything behind it idle.
    • Backfill: fill the hole, delay nobody: Backfill lets a job jump the queue only if its time limit ends before the head job could start anyway, so accurate --time values turn idle nodes into work.
    • Break it: a GPU falls off the bus: A Slurm drain lets running jobs finish and takes no new ones; the node leaves service only when it is empty.
    • Drill: a whole server, interactively: sbatch queues a script for later; srun runs now, inside an allocation it creates or one it is given.
    • Fair share, priority and preemption: Priority is a weighted sum (age, fair share, QOS, size); preemption turns priority into eviction, and checkpoints decide how much a preempted job loses.
  3. Kubernetes' default scheduler

    • Back to Kubernetes: two jobs wait: The default kube-scheduler places pods one by one; it has no notion of a job whose pods are worthless apart.
    • Break it: two jobs, sixteen GPUs each: Without gang scheduling, two jobs that together need more than the cluster can each grab part of it and wait forever for the rest.
    • Thirty-two GPUs doing nothing: A deadlocked cluster looks healthy to Kubernetes: Running pods, Pending pods, busy nodes, and no work.
  4. Gang scheduling: Kueue and Volcano

    • Gang scheduling: all pods or none: Gang scheduling places a job's pods all together or not at all, so a waiting job never holds GPUs.
    • Kueue, Volcano and Run:ai: Kueue adds queues and all-or-nothing admission in front of the default scheduler; Volcano and KAI replace the scheduler for batch pods; all three speak gangs, quotas and preemption.
  5. Quotas, borrowing and preemption

    • Idle GPUs and a waiting job: A quota is checked before capacity: a job over its queue's quota waits on an empty cluster unless its queue may borrow.
    • Borrowing idle quota: Borrowing lends a queue's idle quota to its neighbours; the loan is only safe because it can be taken back.
    • The owner wants its GPUs back: Preemption is how a borrowed or low-priority GPU is returned; what the victim loses is exactly the work since its last checkpoint.
  6. Ray, and choosing a scheduler

    • Ray on Kubernetes: KubeRay: KubeRay runs a Ray cluster as pods: Kubernetes and Kueue give it GPUs, and Ray schedules tasks and actors onto them.
    • Slurm, Kubernetes, or both: Choose Slurm for large, long, tightly coupled training on dedicated hardware, Kubernetes with a batch scheduler for mixed services and training; the concepts are the same in both.
    • Drill: send a Job to a Kueue queue: A Job joins a Kueue queue with one label; Kueue suspends it until the queue's quota admits the whole Job.
  7. Recap & playground

    • Cheat sheet
    • Playground