ROCm & HIP on AMD
The same GPU ideas on AMD Instinct: the ROCm stack beside CUDA's, HIP and hipify for porting, wavefronts of 64 and the bugs they expose, LDS and 304 compute units, rocprof and rocprof-compute, and an honest view of where the stack still lags.
An interactive AI Infrastructure lesson: 20 steps, about 30 minutes, on a live simulation in your browser.
The platform team has taken delivery of mi300x-1: eight AMD Instinct MI300X GPUs, 192 GB of HBM3 each, joined by Infinity Fabric. The ML team's CUDA code has to run on it. amd-smi takes the place of nvidia-smi.
Each layer of NVIDIA's stack has an AMD counterpart, and most of them are open source. The kernel driver is amdgpu (here 6.12.12); ROCm (6.4.1) is the platform; HIP is the runtime API and C++ dialect, compiled by HIP-Clang (hipcc, LLVM's AMDGPU backend); the libraries are rocBLAS and hipBLASLt for GEMMs, MIOpen for convolutions and normalisation, RCCL (2.22.3) for collectives; the profilers are rocprofv3 and rocprof-compute.
What you will learn
The ROCm stack
- An MI300X server and its stack: ROCm is AMD's CUDA: a driver, a runtime (HIP), a compiler, libraries and tools, layer for layer. The names change; the jobs of the layers do not.
- Break it: run the CUDA binary: GPU binaries are vendor-specific. Portability lives in source code and in frameworks, never in a compiled CUDA binary.
HIP and hipify
- The same vector add, in HIP: HIP is CUDA with a different prefix: the kernel language is the same, the host API is renamed, and one source builds for AMD and NVIDIA.
- Porting with hipify: hipify translates API names, not assumptions. Its warnings about warp size are the part of the port a human has to do.
- Drill: build for the MI300X
Wavefronts of 64
- A wavefront is 64 lanes: AMD's wavefront is 64 lanes. Block sizes, lane arithmetic, masks and occupancy all have to be re-derived for 64, not copied from code written for 32.
- Break it: the reduction returns half: Warp-size bugs on AMD are silent: the port compiles, runs at full speed and returns wrong numbers. Test ported kernels for correctness against a reference, not only for speed.
- 128 registers, two vendors: Occupancy is per-architecture arithmetic: registers per lane, lanes per wavefront, slots per CU. Recompute it for each target instead of carrying tuning across vendors.
LDS and 304 compute units
- Break it: a 96 KB tile on 64 KB of LDS: LDS is 64 KB per CU on CDNA 3, a third of an H100 SM's shared memory. Tiles sized for NVIDIA either fail to launch or crush occupancy on AMD.
- Halve the tile: While a kernel is latency-bound, its speed scales with resident wavefronts. Every halving of LDS per workgroup buys a doubling, until bandwidth becomes the roof.
- 1.32× the TFLOPS, how much faster?: Datasheet TFLOPS are a ceiling: delivered GEMM speed is peak × the share the vendor's libraries reach × how well the shape fills the GPU. 304 CUs need 304 × (workgroups per CU) tiles per wave.
- A grid sized for 132 SMs: Never hard-code a GPU's size. Query the CU or SM count at run time and size grids from it; the next GPU will have a different number.
rocprof and rocprof-compute
- rocprofv3: trace and counters: A hardware counter proves an event happened, not that it is the bottleneck. Read the speed-of-light numbers before chasing a counter.
- rocprof-compute: the speed of light: rocprof-compute is to MI300X what Nsight Compute is to H100: Speed-of-Light first, then the wavefront, LDS and cache sections it points at.
- Drill: trace every kernel of a program
Portability, honestly
- PyTorch on ROCm: still torch.cuda: For most teams, AMD support means the ROCm build of their framework and serving engine. The porting work is in the custom kernels and images around them.
- Triton: one source, both vendors: Portability strategy, in order: the framework's ROCm build, Triton for custom kernels, HIP (which also compiles for NVIDIA) where C++ control is needed.
- Where the AMD stack still lags: AMD hardware is competitive on memory and FLOPs; the cost is in software breadth, tuning and people. Budget for those, and benchmark your own workload rather than the datasheet.
Recap & playground
- Cheat sheet
- Playground