learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Reliability & Cost at Scale

Why a job on hundreds of GPUs fails every few hours, what each failure costs, how often to checkpoint, goodput versus uptime, and how GPU-hours, utilisation and capacity plans turn into money.

An interactive AI Infrastructure lesson: 21 steps, about 35 minutes, on a live simulation in your browser.

Alice's team is pre-training a Llama 3.1 70B-shaped model on dgx-1, dgx-2 and dgx-3: 24 H100s as tensor parallel 8 inside each server and data parallel 3 across them, with ZeRO-3. Each step takes 45.66 s and moves 22,964 tokens a second at 41% MFU (model FLOPs utilisation: the share of the GPUs' peak maths that goes into the model).

Every 25 steps the job writes a checkpoint: 988 GB of weights and optimizer state to the parallel filesystem at 20 GB/s, 49.42 s during which every GPU waits.

What you will learn

  1. Failure is a rate

    • A 70B run on 24 GPUs: A synchronous training job is one machine made of many: it runs at the pace of its slowest part and stops when any part stops.
    • Failures add up across servers: Failure rates add: a job on N servers fails N times as often as one server. At a few thousand GPUs, a failure every few hours is normal operation, not bad luck.
  2. What one failure costs

    • Break it: a GPU falls off the bus: A failure throws away everything since the last checkpoint, on every GPU in the job, not only the one that broke.
    • Restart on the hot spare: Each failure costs the work since the checkpoint plus the restart overhead, times every GPU in the job. A hot spare removes the wait for a repair; nothing removes the rest.
    • The failure that does not report itself: A dead peer does not raise an error; it makes everyone else wait. Hangs cost the watchdog timeout on top of the lost work, so detection time is part of failure cost.
    • No spare left: You need as many spares as there are servers waiting for repair at once: failures per day times days to repair.
  3. How often to checkpoint

    • Checkpoint every step?: Checkpointing is insurance paid in GPU time. Too often and the premium eats the run; too rarely and each failure throws away hours.
    • The square-root rule: Checkpoint every √(2 × write time × MTBF). The interval shrinks with the square root of the failure rate, so a cluster 4 times bigger checkpoints 2 times as often.
    • Drill: the checkpoint interval
    • Make each checkpoint cheaper: The cheapest checkpoint is the one the GPUs do not wait for: sharded across every rank, staged in host memory, written in the background.
  4. Goodput, not uptime

    • Goodput versus uptime: Goodput is progress kept divided by time paid for. Uptime counts machines that are on; goodput counts steps you did not have to redo.
    • Drill: goodput from the ledger
  5. GPU-hours into money

    • The bill is GPU-hours: You pay for GPU-hours held; you get GPU-hours used, times MFU. The gap between them is the number FinOps is about.
    • Cloud or your own cluster?: Owned and reserved GPUs are cheap per hour and expensive per idle hour. Buy or commit for the load you can keep busy; rent the peaks.
    • Break it: spot capacity is reclaimed: Spot capacity is a discount for accepting failures on a schedule you do not control. Take it when a checkpoint fits inside the notice period.
  6. Capacity planning

    • Plan training from the FLOPs: GPU-hours = 6 × parameters × tokens / (peak FLOPS × MFU) / goodput. Measure MFU on a short real run before you quote the budget.
    • Drill: GPU-hours for a training run
    • Plan inference from tokens per day: Inference GPUs = peak tokens per second / tokens per second one replica sustains within the latency target, rounded up, plus spare for failures.
  7. FinOps for GPUs

    • FinOps: find the idle GPUs: An allocated GPU doing nothing costs the same as a busy one. Measure SM activity per allocation and attribute every GPU-hour to an owner.
  8. Recap & playground

    • Cheat sheet
    • Playground