learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Power, Cooling & the AI Data Center

Why AI broke the old data-centre assumptions: 10 kW servers and 40 kW racks, power caps that cost less speed than power, air against liquid cooling, what a cooling failure does to a training run, PUE, reference architectures, and planning a room.

An interactive AI Infrastructure lesson: 20 steps, about 30 minutes, on a live simulation in your browser.

Four DGX H100 servers sit in two racks, dgx-1 and dgx-2 in rack-1, dgx-3 and dgx-4 in rack-2. Idle, each H100 draws 70 W and sits at 37 °C; a whole server draws about 2.2 kW, mostly CPUs, fans and NICs.

Start the team's Llama 3.1 70B training run on all 32 GPUs. Within a step every GPU is at 555 W and 62 °C, and each server draws about 6.7 kW. The rack view shows two racks going from under 5 kW to over 13 kW in a minute, and staying there for weeks.

What you will learn

  1. Why AI changed the room

    • Seventy watts, then five hundred: A training cluster is a constant maximum load: every GPU near its power limit, around the clock. Size power and cooling for that, not for an average.
    • Drill: one rack of DGX H100s: Rack power is servers × nameplate: four DGX H100s are 40.8 kW, four to eight times what an ordinary rack was built for.
  2. Air, and when it runs out

    • Every watt becomes heat: A GPU's temperature is its inlet temperature plus a rise set by its power. Anything that warms the inlet warms every GPU by the same amount.
    • The chillers stop: Heat first costs speed, silently: clocks drop to 70% at 87 °C and the synchronous job slows with them. Only temperature alerts turn it into a page.
    • Ninety seconds later: A thermal shutdown looks like a hang, not a crash: the dead nodes vanish, the survivors wait in NCCL, and the scheduler still says R.
    • What the outage cost the job: A facility failure costs a training job the time since its last checkpoint, the hang until the watchdog fires, and the restart. Minutes of outage become an hour of lost GPU time.
  3. Liquid cooling

    • Cold plates instead of fans: Liquid cooling buys thermal headroom and density, not speed: the same GPU at the same power runs about 12 °C cooler, and a rack can carry about three times the heat.
    • The same failure, on liquid: Headroom is time: the 12 °C that liquid cooling saves at the chip, and a loop that warms more slowly than air, are minutes of warning when the facility fails — not immunity.
    • Where each one stops: Air tops out near 40 kW a rack; liquid goes to about 120 kW. The GPU generation you buy decides which one you need, and fewer servers per rack is the air-cooled workaround.
  4. Rack power budgets

    • A rack has a power budget: A rack's real power budget is what it can draw with one feed lost. With A and B feeds, size so that either one alone carries the worst case.
    • Feed B trips on rack-1: Power capping turns an overload into a slowdown. It is what lets a rack lose a feed, or be oversubscribed on purpose, without going dark.
    • How much slower is the job?: Performance falls much more slowly than power: an H100 capped to 63% of its power keeps about 83% of its clock. In a synchronous job the capped GPUs set everyone's pace.
    • Reading a power cap in nvidia-smi: When a GPU is slow with no errors, read its clock event reasons: SW Power Cap, HW Slowdown and thermal slowdown each point to a different cause.
    • Drill: the cap that fits: Per-GPU cap = (rack budget − servers × non-GPU power) ÷ GPUs: (12,000 − 4,960) ÷ 16 = 440 W.
  5. Planning the room

    • PUE: the power that is not compute: PUE = facility power ÷ IT power. The GPU servers set the IT load; PUE decides how much more the site must buy on top.
    • Reference architectures and scalable units: A scalable unit is a pre-designed block of servers, switches, power and cooling. Clusters grow by whole units, so capacity planning is counting units, not redesigning.
    • Planning checklist for a deployment: Plan a GPU deployment in the order constraints bite: power with a feed lost, heat rejection, floor loading, then cabling. The smallest of these sets the cluster's size.
    • Drill: how many GPUs fit the site?: Facility power caps GPU count: 1 MW ÷ 10.2 kW is 98 DGX H100s, 784 GPUs, before a single switch or storage server is counted.
  6. Recap & playground

    • Cheat sheet
    • Playground