Namespaces & cgroups
Build a container by hand from ordinary processes: namespaces for what it can see, cgroups for what it may use, and the Kubernetes name for each piece.
An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.
People call a container "a lightweight virtual machine". It is not one. The kernel has no object called a container: no system call creates one and no table lists them. It has processes, like the eleven in this ps -ef.
What the kernel does have is two properties on every process. Its namespaces decide what it can see: which process list, which hostname, which network interfaces, which mounted filesystems. Its cgroup decides what it may use: how much memory, how much CPU.
What you will learn
There is no container
- The kernel has no containers: There is no container object in the kernel. A container is a normal process with a restricted view (namespaces) and a budget (cgroups).
- Every process already has namespaces: Every process is always in exactly one namespace of each kind. On a plain host they all share the same ones, created at boot.
- lsns: the machine's namespaces: lsns lists namespaces, not processes. More than one row of a type means something on this machine has its own view.
One namespace at a time
- A private hostname: unshare gives a process a private copy of one kind of kernel state. Changes made inside stay inside.
- A new PID namespace: what does ps see?: One process, two pids: the number the host uses and the number it has inside its own PID namespace. NSpid in /proc/PID/status lists both.
- A /proc of its own: PID 1: A PID namespace is a private numbering that starts at 1. From inside you cannot see, name or signal anything outside it.
- When PID 1 exits: A PID namespace lives exactly as long as its PID 1. When that process exits, the kernel kills everything else inside.
- A private mount table: A mount namespace is a private mount table. What is mounted inside is invisible outside, which is how a container gets its own root filesystem.
- A network with nothing in it: A new network namespace contains one loopback interface, down. No eth0, no routes, no open ports: a network has to be plugged in from outside.
Stepping inside
- A namespace needs an occupant: A namespace lives as long as one process is in it. A process that does nothing but sit there is enough; that is what a pod's pause container is.
- nsenter: join someone else's namespaces: nsenter -t PID runs a command inside the namespaces of an existing process. docker exec and kubectl exec are this, with the pid looked up for you.
- Drill: get a shell inside
cgroups: the budget
- Which cgroup is a process in?: A cgroup is a directory under /sys/fs/cgroup. Its members are listed in cgroup.procs; its limits are the files next to it.
- Create a cgroup, set a limit: mkdir makes a cgroup, writing memory.max sets its limit, writing a pid into cgroup.procs moves a process in. The filesystem is the API.
- Break it: spend more than the budget: A cgroup that cannot stay under memory.max gets one of its processes killed with SIGKILL. The program never sees an error it could handle.
- Read the kill in the kernel log: "Memory cgroup out of memory" in dmesg, oom_kill in memory.events, exit status 137: three views of the same kill.
- Now limit the CPU: Memory is incompressible: over the limit means a kill. CPU is compressible: over the limit means waiting. Same cgroup, two very different failures.
- The same limits, set by systemd: MemoryMax= and CPUQuota= in systemd, --memory and --cpus in a runtime, limits in a pod spec: three spellings of writing memory.max and cpu.max.
- Drill: a budget for batch jobs
Putting the pieces together
- A container, built by hand: Container = an ordinary process + its own namespaces + a cgroup with limits. A container runtime is a careful way of setting those up.
- Break it: two small processes, one limit
- The missing piece: a root filesystem: An image is a directory tree; the runtime makes it / inside a new mount namespace. Namespaces + cgroup + root filesystem = container.
- What Kubernetes calls each piece: Pod sandbox = namespaces held open by pause. limits = memory.max and cpu.max. OOMKilled = the cgroup OOM kill, exit code 137. kubectl exec = nsenter.
Recap & playground
- Cheat sheet
- Playground