learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Namespaces & cgroups

Build a container by hand from ordinary processes: namespaces for what it can see, cgroups for what it may use, and the Kubernetes name for each piece.

An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.

People call a container "a lightweight virtual machine". It is not one. The kernel has no object called a container: no system call creates one and no table lists them. It has processes, like the eleven in this ps -ef.

What the kernel does have is two properties on every process. Its namespaces decide what it can see: which process list, which hostname, which network interfaces, which mounted filesystems. Its cgroup decides what it may use: how much memory, how much CPU.

What you will learn

  1. There is no container

    • The kernel has no containers: There is no container object in the kernel. A container is a normal process with a restricted view (namespaces) and a budget (cgroups).
    • Every process already has namespaces: Every process is always in exactly one namespace of each kind. On a plain host they all share the same ones, created at boot.
    • lsns: the machine's namespaces: lsns lists namespaces, not processes. More than one row of a type means something on this machine has its own view.
  2. One namespace at a time

    • A private hostname: unshare gives a process a private copy of one kind of kernel state. Changes made inside stay inside.
    • A new PID namespace: what does ps see?: One process, two pids: the number the host uses and the number it has inside its own PID namespace. NSpid in /proc/PID/status lists both.
    • A /proc of its own: PID 1: A PID namespace is a private numbering that starts at 1. From inside you cannot see, name or signal anything outside it.
    • When PID 1 exits: A PID namespace lives exactly as long as its PID 1. When that process exits, the kernel kills everything else inside.
    • A private mount table: A mount namespace is a private mount table. What is mounted inside is invisible outside, which is how a container gets its own root filesystem.
    • A network with nothing in it: A new network namespace contains one loopback interface, down. No eth0, no routes, no open ports: a network has to be plugged in from outside.
  3. Stepping inside

    • A namespace needs an occupant: A namespace lives as long as one process is in it. A process that does nothing but sit there is enough; that is what a pod's pause container is.
    • nsenter: join someone else's namespaces: nsenter -t PID runs a command inside the namespaces of an existing process. docker exec and kubectl exec are this, with the pid looked up for you.
    • Drill: get a shell inside
  4. cgroups: the budget

    • Which cgroup is a process in?: A cgroup is a directory under /sys/fs/cgroup. Its members are listed in cgroup.procs; its limits are the files next to it.
    • Create a cgroup, set a limit: mkdir makes a cgroup, writing memory.max sets its limit, writing a pid into cgroup.procs moves a process in. The filesystem is the API.
    • Break it: spend more than the budget: A cgroup that cannot stay under memory.max gets one of its processes killed with SIGKILL. The program never sees an error it could handle.
    • Read the kill in the kernel log: "Memory cgroup out of memory" in dmesg, oom_kill in memory.events, exit status 137: three views of the same kill.
    • Now limit the CPU: Memory is incompressible: over the limit means a kill. CPU is compressible: over the limit means waiting. Same cgroup, two very different failures.
    • The same limits, set by systemd: MemoryMax= and CPUQuota= in systemd, --memory and --cpus in a runtime, limits in a pod spec: three spellings of writing memory.max and cpu.max.
    • Drill: a budget for batch jobs
  5. Putting the pieces together

    • A container, built by hand: Container = an ordinary process + its own namespaces + a cgroup with limits. A container runtime is a careful way of setting those up.
    • Break it: two small processes, one limit
    • The missing piece: a root filesystem: An image is a directory tree; the runtime makes it / inside a new mount namespace. Namespaces + cgroup + root filesystem = container.
    • What Kubernetes calls each piece: Pod sandbox = namespaces held open by pause. limits = memory.max and cpu.max. OOMKilled = the cgroup OOM kill, exit code 137. kubectl exec = nsenter.
  6. Recap & playground

    • Cheat sheet
    • Playground