learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

The Accelerator Landscape

NVIDIA Ampere to Blackwell, AMD Instinct, Intel Gaudi, Google TPU, AWS Trainium and Inferentia: the four numbers that separate them, the software that decides whether your code runs, and how to choose on memory, tokens per second, price and power.

An interactive AI Infrastructure lesson: 20 steps, about 30 minutes, on a live simulation in your browser.

Your platform runs Kubernetes over a mixed fleet: two DGX H100 servers, one DGX B200, an 8-GPU AMD Instinct MI300X server, two L40S PCIe boxes, and an Intel Gaudi 3 server on evaluation. Next year's budget is due, and you have to say what to buy.

To make the comparison fair, the same model runs on each: Llama 3.1 70B in bf16, four accelerators per replica. Each vendor's device plugin advertises its own resource: nvidia.com/gpu, amd.com/gpu, habana.ai/gaudi. A pod asks for the one it was built for.

What you will learn

  1. Four numbers per accelerator

    • A fleet with three vendors in it: In Kubernetes an accelerator is an extended resource named by its vendor. A workload is built for one of them; the scheduler will not translate.
    • The four numbers on a datasheet: Capacity decides what fits, bandwidth decides decode speed, compute decides prefill and training, the scale-up link decides how many chips can act as one.
  2. NVIDIA: Ampere to Blackwell

    • Three generations in four numbers: Each NVIDIA generation raised all four numbers, but not evenly: H200 is H100 with more memory, B200 roughly doubles compute and bandwidth again.
    • B200 against H100, one user: For one sequence, decode time per token ≈ bytes of weights per GPU / memory bandwidth. A new generation speeds decode by its bandwidth ratio, not its TFLOPS ratio.
    • The scale-up link doubled too: The scale-up link sets how many chips can share one model efficiently. Measure it with an all-reduce test; busbw near 80 % of the per-direction link speed is healthy.
  3. AMD Instinct

    • MI300X: more memory per chip: AMD Instinct competes on memory: more HBM and more bandwidth per chip than the NVIDIA part of the same generation, with its own stack (ROCm, HIP, RCCL, amd-smi).
    • 70B on one GPU: More HBM per chip means fewer chips per replica: 70B in bf16 on one MI300X, 405B in bf16 on one 8-GPU server. That is where AMD's capacity pays off.
    • Break it: fewest chips under load: Fitting a model on the fewest chips maximises tokens per chip, not latency. A replica's KV capacity and bandwidth decide how many users it can hold.
    • Break it: an NVIDIA habit on AMD: Porting the model is the easy part; porting images, kernels, monitoring and runbooks is the real cost of a second accelerator vendor.
  4. Gaudi, TPU, Trainium and the rest

    • Gaudi 3: Ethernet on the chip: Gaudi 3 puts RoCE Ethernet on the accelerator itself, so scale-up and scale-out both use ordinary Ethernet switches.
    • TPU, Trainium, Inferentia: cloud-only chips: TPU and Trainium trade portability for price-performance: they live in one cloud and run code that goes through one compiler (XLA or Neuron).
  5. Software is the moat

    • The layers that keep code portable: Portability comes from the layer you write against: code written to PyTorch, Triton or vLLM moves between vendors; code written to CUDA directly does not.
    • Which move costs the most?: The cost of a move grows with the distance between software stacks: same vendor < another GPU vendor < another kind of accelerator in another operations model.
    • The link is the other lock-in: Whoever owns the scale-up link owns the size of the model-parallel group. Open standards (UALink for scale-up, Ultra Ethernet for scale-out) are the industry's attempt to separate the chip from the fabric.
  6. How to choose

    • Price per token, not price per GPU: Compare accelerators by cost per token at your latency target: tokens per second per chip, divided into price per chip-hour and watts per chip.
    • Names on the price list: Choose in order: workload, software maturity, availability, price per token, power. The fastest chip you cannot get, run or cool is not an option.
    • Drill: fewest H200s for 405B: The minimum chip count is total weight bytes / HBM per chip, rounded up; the real count is the next layout that also leaves room for the KV cache (for 405B on H200: 8).
    • Drill: the decode speed limit: Decode floor per token = weight bytes / memory bandwidth. Real servers reach about 80 % of it; batching shares the read across many sequences.
  7. Recap & playground

    • Cheat sheet
    • Playground