Learn infrastructure by running it
Linux, networking, Kubernetes, system design and AI infrastructure, taught on live simulations in your browser. Predict what happens, break it, fix it. Free, no sign-up.
How every lesson works
- Watch a real system run: requests move, queues fill, pods get scheduled.
- Make a call: predict what happens next and say how sure you are.
- Break it: crash a server, cut a link, fail a GPU, fill a disk.
- Fix it with the real command and watch it recover.
Linux
The operating system under every container and every node.
- Files & the Filesystem: One tree, rooted at /. Learn to move around it, read a listing, and see what a file name really points at.
- Permissions & Ownership: An owner, a group and nine bits decide every "Permission denied". Learn to read them, change them, and work out which rule said no.
- Users, Groups & sudo: A user is a number, a group is a list of numbers, and sudo is a rule file. Read all three and you can say exactly who may do what.
- Processes & Signals: A process is a program running, with a number, a parent, an owner and a state. Signals are how you talk to one, and the process table is where you read what happened.
- The Shell, Pipes & Redirection: What the shell does to a line before any program runs, the three streams every process gets, and how small tools are joined into pipelines.
- Text Tools: grep, sed, awk: Find lines, cut columns, count and rank, rewrite text and add up numbers: a web server's log and a config file, investigated from the command line.
- systemd & Services: A server is a set of programs that must start in order, stay up and leave a log. systemd is the program that does all three.
- Packages & Repositories: Software on a server comes from a signed catalogue, not from a download page. Learn to read what apt does, and what it leaves behind.
- Disks, Partitions & LVM: From raw blocks to a mounted filesystem, a full disk at 3 a.m., and a volume you can grow without downtime.
- Host Networking: One machine's view of the network: its addresses, routes, names and sockets, and how to tell refused from timed out from unresolved.
- SSH & Keys: Log in to other machines without passwords, read every refusal ssh can give you, and harden sshd without locking yourself out.
- Namespaces & cgroups: Build a container by hand from ordinary processes: namespaces for what it can see, cgroups for what it may use, and the Kubernetes name for each piece.
- Syscalls, Capabilities & seccomp: A program can do nothing on its own. Watch it ask the kernel, then see how root's power is split up and how a filter decides which requests get through.
- Logs & Troubleshooting: A method for a machine that misbehaves, then six incidents to practise it on: CPU, memory, disk, a service, the network and a permission.
- Archives & Compression: Bundle a directory with tar, squeeze it with gzip, xz or zstd, and get it back intact: what each tool really saves, what it costs, and the ways an archive bites when you extract it.
- Environment & Dotfiles: Shell variables versus the environment, how PATH finds a program, which startup files each kind of shell reads, and why a command works for you but not under sudo, in a script or for a colleague.
- Bash Scripting: Turn a release you type by hand into a script you can trust: arguments and options, conditions and loops, functions, exit codes, set -euo pipefail with traps, tracing, quoting and the portability traps.
- Scheduling: cron & Timers: Run things later and on a schedule with cron and systemd timers, and find out why a job did not run: lost output, a different environment, a stray %, overlapping runs and a machine that was off.
- Performance: CPU, Memory & I/O: Load, saturation, iowait, page cache, swap, steal and queue depth: make each kind of slowness on purpose, learn its fingerprint, then triage a real one in sixty seconds.
- Boot, Kernel & Modules: Firmware, GRUB, kernel, initramfs, systemd: what each stage reads, what survives a reboot, how to rescue a machine that will not boot, and how kernel modules load.
- SELinux & AppArmor: A page with mode 644 still returns 403, and root is still told no. Learn the second lock on every door: read the denial, fix the label or the profile, and never switch the lock off.
Networking
What actually happens between two computers.
- What Happens When You Hit Enter: One URL, four conversations: DNS, TCP, TLS and HTTP, what each one costs, and which one failed when the page does not load.
- The OSI Model: Build a packet from the bottom up, one header per problem, then use the same ladder to find which layer is broken.
- IP, Subnets & CIDR: An address is a network number and a host number. The mask decides which is which, and every host uses it to choose between a neighbour and a router.
- Routing & NAT: Routers pass a packet along one decision at a time, and NAT lets a whole private network borrow one public address.
- TCP Handshakes & Reliability: The network loses packets and tells nobody. TCP is how two machines get every byte across anyway.
- DNS Resolution: One question from your laptop, a walk from the root down by the resolver, and caches at every level that decide what you see.
- TLS & HTTPS: How two strangers agree on a secret while everyone listens, and how the client knows who it is talking to.
- HTTP/1.1 vs 2 vs 3: One page is many requests. Three versions of HTTP, three answers to where those requests queue and what one lost packet costs.
- Load Balancers: L4 vs L7: One address in front of many servers: spreading connections at layer 4, understanding requests at layer 7, and what each one does when a server dies.
- Firewalls & Packet Filtering: An ordered list of rules and a default: first match wins, a drop is a timeout, a reject is a refusal, and state is why replies get back in.
- WebSockets: HTTP can only answer. A WebSocket turns one TCP connection into a line both ends may speak on, and you have to keep that line alive.
Kubernetes
From your first pod to a hardened production cluster. Covers every CKAD, CKA and CKS exam competency.
- Pods: The smallest thing Kubernetes runs: what a pod is, how it starts, what restarts it, and what does not bring it back.
- Container Images: From Dockerfile to running container: layers, tags, registries, pull policy and the pull failures you will meet.
- Multi-container Pods: Init containers, sidecars and the volumes they share: what containers in one pod have in common, and what they do not.
- Choosing a Workload: Deployment, StatefulSet, DaemonSet, Job or CronJob: what each one promises about your pods, and how to pick.
- Deployments & Rollouts: Ship a new version with no downtime, and get back to the old one when it goes wrong.
- Blue/Green & Canary: Two Deployments, one Service, and a label selector that decides who gets the customers.
- Helm & Kustomize: Install, upgrade and roll back packaged apps with Helm; adapt plain YAML per environment with Kustomize.
- Probes & Health Checks: Three questions the kubelet asks your container, and what it does with each answer.
- Logs & Debugging: Five broken pods, one routine: get, describe, logs, events, exec. Learn which one answers which question.
- ConfigMaps & Secrets: Keep settings out of the image: hand them to pods as variables or files, and know what base64 does not protect.
- Requests, Limits & Quotas: Requests reserve room on a node. Limits are enforced by the kernel. Quotas cap what a whole namespace may ask for.
- SecurityContext & ServiceAccounts: A pod has two identities: a Linux user on the node and a ServiceAccount to the API server. Shrink both.
- CRDs, Operators & API Versions: Teach the API a new noun, let an operator act on it, and keep your manifests alive when old API versions are removed.
- Services & DNS: Pods die and change IP. A Service gives them one name that never moves.
- Ingress: One public entry point for many Services: host and path rules, TLS, and the controller that makes the rules real.
- NetworkPolicies: Every pod can reach every pod until you say otherwise. Select pods, isolate them, and allow back only the paths the shop needs.
- Cluster Architecture: Follow one kubectl apply through every component until a container is running.
- kubeadm: Install & Upgrade: Bootstrap a cluster, join nodes, and upgrade it one minor version at a time.
- HA Control Plane & etcd: Quorum, leader election, and backing up the one database that holds everything.
- RBAC: Roles, bindings and subjects: who may do what, and how to prove it.
- Scheduling: How the scheduler picks a node, and every way you can steer it.
- Autoscaling & Self-healing: Let the cluster add replicas under load and replace what breaks.
- Pod Networking & CoreDNS: How a packet gets from one pod to another: CNI, kube-proxy and cluster DNS.
- Gateway API: The successor to Ingress: a GatewayClass, a Gateway and HTTPRoutes, owned by different people, with traffic splitting built in.
- Storage: Volumes that outlive pods: PVs, PVCs, StorageClasses and reclaim policies.
- Troubleshooting Workloads: Eight broken workloads, one habit: read the STATUS column as a diagnosis, then run the one command that holds the reason.
- Troubleshooting the Cluster: Nodes go NotReady and control-plane components die. Check the API, the nodes, the system pods and the workload, in that order.
- Troubleshooting Networking: A request fails somewhere between the name and the pod. Read the error, test one hop at a time, and find the hop.
- CIS Benchmarks & API Hardening: Audit an inherited control plane with kube-bench, close what it finds, and check what you run before you run it.
- Network Lockdown: Default-deny everywhere, a guarded metadata endpoint, and TLS at the edge.
- Least Privilege: Shrink every identity to exactly what it needs, starting with ServiceAccounts.
- Kernel & Host Hardening: seccomp, AppArmor and capabilities: what stands between a container and the kernel.
- Pod Security Standards: Privileged, baseline, restricted: enforce a floor for every pod in a namespace.
- Secrets, Sandboxes & mTLS: Encrypt Secrets at rest, sandbox untrusted pods, encrypt traffic between them.
- Image Footprint & Scanning: Smaller images, scanned images, and manifests checked before they ship.
- Signed Images & Trusted Registries: Only run what you built: allowed registries and signatures verified at admission.
- Runtime Security with Falco: Watch syscalls for behaviour that should never happen, and make containers immutable.
- Audit Logs & Investigation: Record who did what, then reconstruct an attack from the evidence.
System Design
How big systems stay fast and stay up.
- Load Balancing: One address, many servers: why one server falls over, how a balancer chooses between several, and what it does when one of them is slow, broken or dead.
- Caching: Keep answers close so the database does not have to give them again: cache-aside, hit ratios, TTLs, invalidation, eviction, and the stampede when a hot key expires.
- CDNs: Serve bytes from the edge, close to every user: what distance costs, what an edge saves the origin, how Cache-Control and versioned file names decide what users see after a deploy, and what happens when an edge or the origin goes down.
- Consistent Hashing: Add a server and move only 1/N of the keys: why hash mod N reshuffles everything, how a hash ring limits a change to one arc, how virtual nodes even out the load, and what no hashing scheme can do about a hot key.
- Database Replication: Leaders, followers, failover and the lag that bites: what copies of a database buy you, and what each kind of copy costs.
- Sharding: Split a database that no longer fits on one machine: shard keys, routers, modulo versus range versus consistent hashing, and the queries and hot keys that sharding cannot fix.
- Rate Limiting: Token buckets, leaky buckets and sliding windows: how an API says 429 to one noisy client so that everyone else still gets served, and where that decision lives.
- Message Queues: Decouple services, absorb spikes and retry safely: acknowledgements, visibility timeouts, duplicates and idempotency, dead-letter queues, backpressure, ordering and pub/sub.
- CAP & Consistency: Three copies of a shopping cart in two data centres: read and write quorums, what a network partition forces you to choose, and how diverged copies are put back together.
- Consensus & Raft: Five machines keep one history of a shop's stock: elections and terms, commit on a majority, what crashes and partitions do, and why clusters come in odd sizes.
- Scaling Up vs Scaling Out: Bigger machines or more machines: what each buys, what each costs per month and per million requests, where each stops working, and the state and timing problems that come with more machines.
- SQL vs NoSQL: One shop stored four ways (Postgres, MongoDB, Redis Cluster, Cassandra): which questions each answers cheaply, which it refuses, and what each does with a schema change, a half-finished order, a write flood and a dead machine.
- Back-of-the-Envelope Estimates: Turn daily users into requests a second, bytes a day, cache size and servers, round like an engineer, know the latency numbers by heart, and then prove the estimate against a running system.
- Indexes & Query Plans: Read EXPLAIN ANALYZE line by line, turn a 2-second sequential scan into a sub-millisecond index lookup, and learn why the planner sometimes ignores your index, or trusts statistics that lie.
- API Design: REST, gRPC & Pagination: An API is a contract that outlives its first client: status codes and error bodies, conditional requests, cursors instead of offsets, REST against gRPC on the wire, and changing a schema without breaking the apps already installed on people's phones.
- Idempotency & Retries: A payment whose answer is lost: why a blind retry charges twice, how an Idempotency-Key makes the retry safe, which failures are worth retrying, how retries at every layer multiply into 27 requests, and how jitter and retry budgets keep a recovering service alive.
- Search & Inverted Indexes: Why LIKE '%word%' scans every row while a search engine answers in milliseconds: analyzers, postings lists, BM25 scoring, near-real-time refresh, and scatter-gather across shards that can be slow or gone.
- Distributed Transactions: 2PC & Sagas: One order touches four services with four databases. Two-phase commit makes it atomic and blocks when the coordinator dies; a saga never blocks and pays with compensations, idempotent steps and the outbox. Watch both fail, and learn which to choose.
- Event Sourcing & CQRS: Store what happened instead of what is: an append-only log of events as the source of truth, state rebuilt by replaying it, optimistic concurrency on stream versions, snapshots, questions about the past, read models that lag and can be rebuilt, schemas that change, and personal data you must forget in a log that never forgets.
- Observability: Metrics, Logs & Traces: Finding the slow hop in a request that crossed eight services: traces and their critical path, RED metrics and the percentiles averages hide, alerts on symptoms rather than causes, sampling, cardinality, and logs that join up by trace id.
- Case Study: A URL Shortener: The classic interview question done properly and then run in production: requirements, numbers, API, data model, five ways to make a short code, a 301 that hides your clicks, a hot link, a bot and a dead shard.
- Case Study: A Chat System: WhatsApp in an interview and then in production: requirements, numbers, a WebSocket protocol, per-conversation sequence numbers, receipts, offline inboxes and push, retries without duplicates, presence, a gateway crash and the 5,000-member group.
AI Infrastructure
The accelerators, networks, storage and software that train and serve AI models, from one GPU to a cluster, across NVIDIA, AMD and the cloud.
- What AI Workloads Need: A model is billions of numbers and a lot of matrix multiplication. Training learns them over weeks; inference uses them in milliseconds. Three numbers decide what hardware either needs: FLOPs, memory capacity and memory bandwidth.
- How an LLM Writes Text: What really happens between a prompt and an answer: tokens and the tokenizer, one score per vocabulary entry, softmax, the generation loop that runs the whole model once per token, greedy decoding, temperature, top-k and top-p, seeds, and the three ways generation stops.
- Inside the Transformer: What sits between token ids and next-token scores: embeddings, attention as queries, keys and values, the √d scaling and softmax, position information, the causal mask and why it makes a KV cache possible, many heads, the MLP, the output head, and where a real model's parameters, FLOPs and memory go, including mixture of experts.
- Inside a GPU: Open one accelerator: streaming multiprocessors and compute units, tensor and matrix cores, HBM and the memory hierarchy, precisions from FP32 to FP4, power and clocks, and how to read nvidia-smi and amd-smi without being fooled.
- The Accelerator Software Stack: Firmware, driver, CUDA or ROCm, the math and collective libraries, the framework and the container: what each layer does, how they depend on each other, and the version mismatches that stop a GPU job before it computes anything.
- GPU Memory Math: Will it fit, and on how many GPUs? Weights by precision, the KV cache, the 16 bytes per parameter of training with Adam, and how ZeRO, FSDP, tensor parallelism and activation checkpointing divide the bill.
- Anatomy of a GPU Server: What is inside an 8-GPU server from NVIDIA or AMD and a PCIe L40S box: CPUs, memory, PCIe switches, the scale-up fabric, one NIC per GPU, the BMC and the power, and how reading the topology tells you which workloads a server will run well.
- Prefill, Decode and the Cost of a Token: Why reading a prompt is compute-bound and writing an answer is memory-bound: FLOPs per byte, the roofline and the ridge point, time to first token and time per output token, batching, what long contexts do to it, fp8 and int4 weights, when a model needs more GPUs, and how training differs.
- Distributed Training: Data Parallel: Many GPUs, one model: copy it to every GPU, split the batch, average the gradients with a ring all-reduce, hide that behind the backward pass, and find out why adding GPUs eventually stops helping.
- Model Parallelism: When a model does not fit one GPU even for one step: shard its optimizer state, gradients and weights with ZeRO and FSDP, split its matrices with tensor parallelism, its layers with pipeline parallelism, its experts with expert parallelism, and combine them into a 3D plan for Llama 3.1 70B on 32 H100s.
- AI Networking: InfiniBand & RoCE: Why a training cluster needs a different network: RDMA, InfiniBand and RoCE, GPUDirect, rail-optimised fat trees, the bandwidth maths, and how one bad cable or one dead link stalls every GPU in the job.
- Storage & Data Pipelines: Keeping GPUs fed and their work safe: what training reads and writes, the data loader that starves a node, storage from NVMe to parallel filesystems and S3, checkpoint maths, GPUDirect Storage and model cold starts.
- GPUs on Kubernetes: Why Kubernetes cannot see a GPU until a device plugin advertises it, what the NVIDIA and AMD GPU operators install, how pods ask for accelerators, how to steer them to the right model, and how to share one GPU with time-slicing, MPS and MIG.
- Scheduling GPU Clusters: Why GPU jobs need all-or-nothing scheduling: Slurm's partitions, backfill, draining and priorities; the deadlock Kubernetes' default scheduler walks into and the gang scheduling of Kueue and Volcano that prevents it; quotas, borrowing and preemption; Ray on Kubernetes; and when to pick Slurm or Kubernetes.
- Serving Models: A model behind an API, answering in milliseconds all day: prefill and decode, the latency metrics users feel, batching as the lever for throughput, queues, replicas, tensor parallelism, cold starts and the serving stacks that do it.
- LLM Inference at Scale: What limits an LLM service at scale and the tools that lift each limit: the KV cache and paged attention, continuous batching under load, quantization, speculative decoding, prefix caching, long contexts, autoscaling on the right signal, cost per million tokens and a p99 incident.
- GPU Monitoring & Observability: Why nvidia-smi's GPU-Util says 100% for a GPU doing almost nothing, the DCGM profiling metrics that tell the truth, MFU, health signals, exporters and alerts, and finding the GPUs a cluster is wasting.
- Bring-up & Validation: From delivered racks to a cluster accepted for production: out-of-band access and firmware, the software stack in the right order, DCGM diagnostics and HPL burn-in, per-rail and NCCL tests, written acceptance criteria, and the first real job.
- Troubleshooting GPU Clusters: A triage order for GPU clusters, reading Xid codes, and six real incidents diagnosed from their symptoms: a GPU off the bus, an NCCL hang, a thermal straggler, a flapping link, an uncorrectable ECC error and a half-finished driver upgrade.
- Power, Cooling & the AI Data Center: Why AI broke the old data-centre assumptions: 10 kW servers and 40 kW racks, power caps that cost less speed than power, air against liquid cooling, what a cooling failure does to a training run, PUE, reference architectures, and planning a room.
- Reliability & Cost at Scale: Why a job on hundreds of GPUs fails every few hours, what each failure costs, how often to checkpoint, goodput versus uptime, and how GPU-hours, utilisation and capacity plans turn into money.
- The Accelerator Landscape: NVIDIA Ampere to Blackwell, AMD Instinct, Intel Gaudi, Google TPU, AWS Trainium and Inferentia: the four numbers that separate them, the software that decides whether your code runs, and how to choose on memory, tokens per second, price and power.
- Ray & Distributed AI Frameworks: The layer between the cluster and the model code: PyTorch distributed and NCCL, DeepSpeed, Megatron-LM and FSDP, Ray and KubeRay, JAX on TPUs, who owns what, and how a distributed launch fails.
- The KV Cache: Why every generated token needs the keys and values of every token before it, what that costs per token (MHA, GQA, MQA; bf16 and fp8), how many sequences fit after the weights, paged blocks, shared prefixes, what happens when the cache is full, offloading and disaggregation, and the two vLLM metrics that tell you.
- Batching for LLM Inference: Why one request at a time wastes a memory-bound GPU, static and dynamic batching and where each belongs, continuous iteration-level batching, the throughput-latency curve and its knee, prefill against decode and the token budget, max-num-seqs, KV-bound batches, and settings for a chat SLO versus an offline job.
- Quantization: Number formats from FP32 to NVFP4, what quantization shrinks and why that speeds decode, weight-only versus weight-and-activation schemes, scales, GPTQ, AWQ, SmoothQuant and FP8, measuring accuracy, which GPUs support which format, and the tools: the same 70B model served in several precisions, on fewer GPUs per replica.
- Speculative Decoding: Why decode leaves the tensor cores idle, how a draft-and-verify loop with rejection sampling buys several tokens per weight read without changing the output, the acceptance rate and the expected-tokens formula, draft models, n-gram lookup, Medusa and EAGLE, when speculation pays and when it hurts under load, what it costs in memory, how to read its metrics, and how it combines with quantization and batching.
- GPU Programming with CUDA: How GPU code really runs: grids of blocks of threads, warps of 32 and waves over SMs; why most kernels are memory-bound; coalescing, occupancy, shared memory and tiling; fusion and launch overhead; and how to read Nsight Compute and Nsight Systems.
- ROCm & HIP on AMD: The same GPU ideas on AMD Instinct: the ROCm stack beside CUDA's, HIP and hipify for porting, wavefronts of 64 and the bugs they expose, LDS and 304 compute units, rocprof and rocprof-compute, and an honest view of where the stack still lags.
- Triton Kernels: Writing GPU kernels in Python with Triton: block programs, fusion, tiled matmul and flash attention, autotuning, and one kernel for NVIDIA and AMD.