learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Learn AI Infrastructure

The accelerators, networks, storage and software that train and serve AI models, from one GPU to a cluster, across NVIDIA, AMD and the cloud.

Basic

Start here. The ideas everything else is built on, from the first command.

  • What AI Workloads Need: A model is billions of numbers and a lot of matrix multiplication. Training learns them over weeks; inference uses them in milliseconds. Three numbers decide what hardware either needs: FLOPs, memory capacity and memory bandwidth. (about 30 minutes)
  • How an LLM Writes Text: What really happens between a prompt and an answer: tokens and the tokenizer, one score per vocabulary entry, softmax, the generation loop that runs the whole model once per token, greedy decoding, temperature, top-k and top-p, seeds, and the three ways generation stops. (about 30 minutes)
  • Inside the Transformer: What sits between token ids and next-token scores: embeddings, attention as queries, keys and values, the √d scaling and softmax, position information, the causal mask and why it makes a KV cache possible, many heads, the MLP, the output head, and where a real model's parameters, FLOPs and memory go, including mixture of experts. (about 30 minutes)
  • Inside a GPU: Open one accelerator: streaming multiprocessors and compute units, tensor and matrix cores, HBM and the memory hierarchy, precisions from FP32 to FP4, power and clocks, and how to read nvidia-smi and amd-smi without being fooled. (about 35 minutes)
  • The Accelerator Software Stack: Firmware, driver, CUDA or ROCm, the math and collective libraries, the framework and the container: what each layer does, how they depend on each other, and the version mismatches that stop a GPU job before it computes anything. (about 30 minutes)
  • GPU Memory Math: Will it fit, and on how many GPUs? Weights by precision, the KV cache, the 16 bytes per parameter of training with Adam, and how ZeRO, FSDP, tensor parallelism and activation checkpointing divide the bill. (about 30 minutes)
  • Anatomy of a GPU Server: What is inside an 8-GPU server from NVIDIA or AMD and a PCIe L40S box: CPUs, memory, PCIe switches, the scale-up fabric, one NIC per GPU, the BMC and the power, and how reading the topology tells you which workloads a server will run well. (about 30 minutes)

Intermediate

What you need to run it for real: the moving parts and the ways they fail.

  • Prefill, Decode and the Cost of a Token: Why reading a prompt is compute-bound and writing an answer is memory-bound: FLOPs per byte, the roofline and the ridge point, time to first token and time per output token, batching, what long contexts do to it, fp8 and int4 weights, when a model needs more GPUs, and how training differs. (about 30 minutes)
  • Distributed Training: Data Parallel: Many GPUs, one model: copy it to every GPU, split the batch, average the gradients with a ring all-reduce, hide that behind the backward pass, and find out why adding GPUs eventually stops helping. (about 32 minutes)
  • Model Parallelism: When a model does not fit one GPU even for one step: shard its optimizer state, gradients and weights with ZeRO and FSDP, split its matrices with tensor parallelism, its layers with pipeline parallelism, its experts with expert parallelism, and combine them into a 3D plan for Llama 3.1 70B on 32 H100s. (about 34 minutes)
  • AI Networking: InfiniBand & RoCE: Why a training cluster needs a different network: RDMA, InfiniBand and RoCE, GPUDirect, rail-optimised fat trees, the bandwidth maths, and how one bad cable or one dead link stalls every GPU in the job. (about 32 minutes)
  • Storage & Data Pipelines: Keeping GPUs fed and their work safe: what training reads and writes, the data loader that starves a node, storage from NVMe to parallel filesystems and S3, checkpoint maths, GPUDirect Storage and model cold starts. (about 30 minutes)
  • GPUs on Kubernetes: Why Kubernetes cannot see a GPU until a device plugin advertises it, what the NVIDIA and AMD GPU operators install, how pods ask for accelerators, how to steer them to the right model, and how to share one GPU with time-slicing, MPS and MIG. (about 32 minutes)
  • Scheduling GPU Clusters: Why GPU jobs need all-or-nothing scheduling: Slurm's partitions, backfill, draining and priorities; the deadlock Kubernetes' default scheduler walks into and the gang scheduling of Kueue and Volcano that prevents it; quotas, borrowing and preemption; Ray on Kubernetes; and when to pick Slurm or Kubernetes. (about 32 minutes)
  • Serving Models: A model behind an API, answering in milliseconds all day: prefill and decode, the latency metrics users feel, batching as the lever for throughput, queues, replicas, tensor parallelism, cold starts and the serving stacks that do it. (about 32 minutes)
  • The Accelerator Landscape: NVIDIA Ampere to Blackwell, AMD Instinct, Intel Gaudi, Google TPU, AWS Trainium and Inferentia: the four numbers that separate them, the software that decides whether your code runs, and how to choose on memory, tokens per second, price and power. (about 30 minutes)
  • Ray & Distributed AI Frameworks: The layer between the cluster and the model code: PyTorch distributed and NCCL, DeepSpeed, Megatron-LM and FSDP, Ray and KubeRay, JAX on TPUs, who owns what, and how a distributed launch fails. (about 30 minutes)
  • The KV Cache: Why every generated token needs the keys and values of every token before it, what that costs per token (MHA, GQA, MQA; bf16 and fp8), how many sequences fit after the weights, paged blocks, shared prefixes, what happens when the cache is full, offloading and disaggregation, and the two vLLM metrics that tell you. (about 35 minutes)
  • Batching for LLM Inference: Why one request at a time wastes a memory-bound GPU, static and dynamic batching and where each belongs, continuous iteration-level batching, the throughput-latency curve and its knee, prefill against decode and the token budget, max-num-seqs, KV-bound batches, and settings for a chat SLO versus an offline job. (about 35 minutes)
  • Quantization: Number formats from FP32 to NVFP4, what quantization shrinks and why that speeds decode, weight-only versus weight-and-activation schemes, scales, GPTQ, AWQ, SmoothQuant and FP8, measuring accuracy, which GPUs support which format, and the tools: the same 70B model served in several precisions, on fewer GPUs per replica. (about 35 minutes)

Advanced

Production depth: hardening, recovery and the trade-offs behind the design.

  • LLM Inference at Scale: What limits an LLM service at scale and the tools that lift each limit: the KV cache and paged attention, continuous batching under load, quantization, speculative decoding, prefix caching, long contexts, autoscaling on the right signal, cost per million tokens and a p99 incident. (about 35 minutes)
  • GPU Monitoring & Observability: Why nvidia-smi's GPU-Util says 100% for a GPU doing almost nothing, the DCGM profiling metrics that tell the truth, MFU, health signals, exporters and alerts, and finding the GPUs a cluster is wasting. (about 32 minutes)
  • Bring-up & Validation: From delivered racks to a cluster accepted for production: out-of-band access and firmware, the software stack in the right order, DCGM diagnostics and HPL burn-in, per-rail and NCCL tests, written acceptance criteria, and the first real job. (about 32 minutes)
  • Troubleshooting GPU Clusters: A triage order for GPU clusters, reading Xid codes, and six real incidents diagnosed from their symptoms: a GPU off the bus, an NCCL hang, a thermal straggler, a flapping link, an uncorrectable ECC error and a half-finished driver upgrade. (about 35 minutes)
  • Power, Cooling & the AI Data Center: Why AI broke the old data-centre assumptions: 10 kW servers and 40 kW racks, power caps that cost less speed than power, air against liquid cooling, what a cooling failure does to a training run, PUE, reference architectures, and planning a room. (about 30 minutes)
  • Reliability & Cost at Scale: Why a job on hundreds of GPUs fails every few hours, what each failure costs, how often to checkpoint, goodput versus uptime, and how GPU-hours, utilisation and capacity plans turn into money. (about 35 minutes)
  • Speculative Decoding: Why decode leaves the tensor cores idle, how a draft-and-verify loop with rejection sampling buys several tokens per weight read without changing the output, the acceptance rate and the expected-tokens formula, draft models, n-gram lookup, Medusa and EAGLE, when speculation pays and when it hurts under load, what it costs in memory, how to read its metrics, and how it combines with quantization and batching. (about 35 minutes)
  • GPU Programming with CUDA: How GPU code really runs: grids of blocks of threads, warps of 32 and waves over SMs; why most kernels are memory-bound; coalescing, occupancy, shared memory and tiling; fusion and launch overhead; and how to read Nsight Compute and Nsight Systems. (about 35 minutes)
  • ROCm & HIP on AMD: The same GPU ideas on AMD Instinct: the ROCm stack beside CUDA's, HIP and hipify for porting, wavefronts of 64 and the bugs they expose, LDS and 304 compute units, rocprof and rocprof-compute, and an honest view of where the stack still lags. (about 30 minutes)
  • Triton Kernels: Writing GPU kernels in Python with Triton: block programs, fusion, tiled matmul and flash attention, autotuning, and one kernel for NVIDIA and AMD. (about 35 minutes)

The infrastructure behind AI, from one GPU to a cluster

Training and serving AI models is now an infrastructure problem: GPUs that sit idle waiting for data, collectives that stall on one bad cable, inference servers that run out of KV cache, clusters that cost thousands of dollars an hour. These lessons simulate that infrastructure in your browser, with GPUs, NVLink, InfiniBand and RoCE networks, storage, schedulers and inference servers, and real numbers for memory, bandwidth, throughput and latency.

You start with how a large language model actually works, from first principles: tokens, the generation loop, sampling, attention and the transformer, and why reading a prompt and writing an answer cost a GPU so differently. Then you go inside a single GPU and learn memory maths, the software stack from drivers to CUDA and NCCL, and how a GPU server is built. Then you scale out: data, tensor and pipeline parallelism, collectives and all-reduce, AI networking, storage pipelines, GPUs on Kubernetes and Slurm scheduling. On the serving side you learn LLM inference: the KV cache, continuous batching, quantization, speculative decoding and vLLM. The advanced lessons cover GPU programming with CUDA, ROCm and HIP on AMD, and Triton kernels, plus monitoring, bring-up, troubleshooting, data-centre power and cooling, and cost. NVIDIA and AMD hardware are both covered.

The goal is to be useful on an AI platform, MLOps or GPU infrastructure team: to size a cluster, read nvidia-smi and DCGM, find the straggler, and explain why a model is memory-bound.

After these lessons you can

  • Explain how an LLM generates text: tokens, logits, softmax, sampling and attention
  • Calculate how much GPU memory a model needs to train or serve
  • Choose data, tensor or pipeline parallelism and explain the network traffic each creates
  • Run and read NCCL tests, nvidia-smi and DCGM, and find a failing GPU or link
  • Serve LLMs efficiently with batching, KV cache management, quantization and speculative decoding
  • Schedule GPU jobs on Kubernetes and Slurm without wasting accelerators
  • Read a GPU kernel profile and tell whether it is limited by memory or by compute

Who it is for

Infrastructure, platform and DevOps engineers moving into AI, MLOps engineers, and ML engineers who want to understand the hardware and systems their models run on.

Common questions

Do I need a GPU to take these lessons?

No. The GPUs, networks and clusters are simulated in your browser, with realistic numbers taken from published specifications and benchmarks.

Is this only about NVIDIA?

No. The lessons cover NVIDIA and AMD GPUs, compare other accelerators, and teach CUDA, ROCm and HIP, and Triton, which runs on both vendors.

Do I need machine learning knowledge?

Only the basics. The lessons explain the parts of training and inference that matter for infrastructure, such as parameters, activations, gradients and tokens, when they first appear.