learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Kernel & Host Hardening (CKS)

seccomp, AppArmor and capabilities: what stands between a container and the kernel.

An interactive Kubernetes lesson: 20 steps, about 32 minutes, on a live simulation in your browser.

Look inside node-a. The pod web serves the shop's pages on port 80. It looks like its own machine, but it is an ordinary Linux process, fenced in by namespaces and cgroups, and it talks to the node's kernel. There is no other kernel.

An attacker has found a bug in web and can run commands in the container. The pod has no securityContext, so look at what the defaults allow. They send a raw network packet. They add a line to /etc/passwd. Both work.

What you will learn

  1. One kernel, one host

    • A container is a process on the node: Every container on a node shares that node's one kernel. A kernel bug reached from one container is a bug in all of them.
    • What is this node running?: A node should run the kubelet, a container runtime, kube-proxy and SSH. Every other service and package is attack surface with no benefit.
    • Stop is not disable: stop is for now, disable is for the next boot, purge is for good. A package that is not installed cannot be restarted by anyone.
    • Who can log in to a node: Anyone with a root shell on a node owns every pod on it, and RBAC never sees them. Keep that list of people very short.
    • Drill: stop it for good
  2. Capabilities

    • Root in a container: Root's power is split into about 40 capabilities. A container's root gets 14 of them by default, and most apps need none.
    • Drop ALL, add back one: capabilities: drop ["ALL"], then add the one or two the app needs. An allow-list you wrote is smaller than a default you inherited.
    • Break it: privileged: true: privileged: true is not one extra permission. It switches off every gate at once and hands over the node's devices.
  3. seccomp

    • The syscall filter nobody turned on: seccomp filters which syscalls a process may make. In Kubernetes it is off unless you ask for it.
    • RuntimeDefault: RuntimeDefault is the seccomp profile to put on every pod: it blocks the dangerous syscalls and breaks almost nothing.
    • A profile of your own: A Localhost seccomp profile is a file on the node under /var/lib/kubelet/seccomp. The pod names it by relative path.
    • Break it: the profile is on one node
    • Changing the default itself: The kubelet's seccompDefault turns "no profile" from Unconfined into RuntimeDefault for every new container on that node.
  4. AppArmor

    • AppArmor: rules about paths: An AppArmor profile is loaded into a node's kernel with apparmor_parser. Until a container asks for it by name, it confines nothing.
    • Root meets AppArmor: File permissions ask who you are. AppArmor asks which program you are, and root gets the same answer as everyone else.
    • Break it: not loaded on this node: A pod carries only the name of a seccomp or AppArmor profile. The profile itself must exist on every node the pod can land on.
    • Drill: load the profile
  5. Recap & playground

    • All four gates, one manifest: Capabilities limit root, seccomp limits syscalls, AppArmor limits paths. Each stops attacks the other two let through.
    • Cheat sheet
    • Playground