Scheduling GPU Clusters
Why GPU jobs need all-or-nothing scheduling: Slurm's partitions, backfill, draining and priorities; the deadlock Kubernetes' default scheduler walks into and the gang scheduling of Kueue and Volcano that prevents it; quotas, borrowing and preemption; Ray on Kubernetes; and when to pick Slurm or Kubernetes.
An interactive AI Infrastructure lesson: 21 steps, about 32 minutes, on a live simulation in your browser.
Two teams share this cluster: research and product, 32 H100s in four DGX servers under Kubernetes. Research starts pretrain, a Llama 3.1 70B run: tensor parallel 8 inside each server, data parallel 4 across them, four pods of eight GPUs.
Each step takes 34.28 s at 41% MFU, and every one of those seconds needs all 32 GPUs: every rank computes a slice, then waits in an all-reduce for the other 31. A web service with 10 of its 12 replicas is degraded. A training job with 31 of its 32 GPUs does not run at all.
What you will learn
Why GPU scheduling is different
- Thirty-two GPUs or nothing: A distributed training job is one unit of work: it needs every GPU it asked for at the same time, or it does nothing.
- Take one server away: Evicting one pod of a training job kills the job: on GPU clusters, maintenance is scheduled with the scheduler, not around it.
Slurm: the HPC way
- Slurm: partitions, sbatch and srun: Slurm hands a job its whole allocation at once or not at all; partitions are the queues and GRES counts the GPUs.
- An idle server and a waiting job: Strict FIFO protects the big job at the head of the queue by leaving everything behind it idle.
- Backfill: fill the hole, delay nobody: Backfill lets a job jump the queue only if its time limit ends before the head job could start anyway, so accurate --time values turn idle nodes into work.
- Break it: a GPU falls off the bus: A Slurm drain lets running jobs finish and takes no new ones; the node leaves service only when it is empty.
- Drill: a whole server, interactively: sbatch queues a script for later; srun runs now, inside an allocation it creates or one it is given.
- Fair share, priority and preemption: Priority is a weighted sum (age, fair share, QOS, size); preemption turns priority into eviction, and checkpoints decide how much a preempted job loses.
Kubernetes' default scheduler
- Back to Kubernetes: two jobs wait: The default kube-scheduler places pods one by one; it has no notion of a job whose pods are worthless apart.
- Break it: two jobs, sixteen GPUs each: Without gang scheduling, two jobs that together need more than the cluster can each grab part of it and wait forever for the rest.
- Thirty-two GPUs doing nothing: A deadlocked cluster looks healthy to Kubernetes: Running pods, Pending pods, busy nodes, and no work.
Gang scheduling: Kueue and Volcano
- Gang scheduling: all pods or none: Gang scheduling places a job's pods all together or not at all, so a waiting job never holds GPUs.
- Kueue, Volcano and Run:ai: Kueue adds queues and all-or-nothing admission in front of the default scheduler; Volcano and KAI replace the scheduler for batch pods; all three speak gangs, quotas and preemption.
Quotas, borrowing and preemption
- Idle GPUs and a waiting job: A quota is checked before capacity: a job over its queue's quota waits on an empty cluster unless its queue may borrow.
- Borrowing idle quota: Borrowing lends a queue's idle quota to its neighbours; the loan is only safe because it can be taken back.
- The owner wants its GPUs back: Preemption is how a borrowed or low-priority GPU is returned; what the victim loses is exactly the work since its last checkpoint.
Ray, and choosing a scheduler
- Ray on Kubernetes: KubeRay: KubeRay runs a Ray cluster as pods: Kubernetes and Kueue give it GPUs, and Ray schedules tasks and actors onto them.
- Slurm, Kubernetes, or both: Choose Slurm for large, long, tightly coupled training on dedicated hardware, Kubernetes with a batch scheduler for mixed services and training; the concepts are the same in both.
- Drill: send a Job to a Kueue queue: A Job joins a Kueue queue with one label; Kueue suspends it until the queue's quota admits the whole Job.
Recap & playground
- Cheat sheet
- Playground