Syscalls, Capabilities & seccomp
A program can do nothing on its own. Watch it ask the kernel, then see how root's power is split up and how a filter decides which requests get through.
An interactive Linux lesson: 24 steps, about 35 minutes, on a live simulation in your browser.
cat /etc/hostname prints web01. It looks as if cat read a file and wrote to your screen. It did neither. cat has no way to touch a disk, a terminal, a network card or another process.
A Linux machine runs code in two modes. The kernel runs in kernel space, with full access to hardware and to every process's memory. Everything else, from cat to a database, runs in user space, where the CPU refuses those operations outright.
What you will learn
Asking the kernel
- A program cannot open a file: User space cannot touch files, the network or other processes. It can only ask the kernel, one system call at a time.
- Watch cat ask: A running program is a sequence of system calls. strace prints that sequence, so you can watch what it does instead of guessing.
- Read one line, field by field: name(arguments) = return value. For open the return is a file descriptor; for read and write it is a byte count.
- When the kernel says no: A failed syscall returns -1 and an errno. Every error message a program prints started life as one of those codes.
strace as a debugger
- Break it: an error that says nothing: When the message is useless, trace the program and read the last failing call: it names the file and the reason.
- From errno to fix
- A different errno, a different fix: The errno names the kind of fault: EACCES is a permission check, ENOENT is a path that does not exist. Read it before you reach for chmod.
- strace -c: the summary
- Drill: keep the evidence
Root is not one switch
- Root's power comes in pieces: Root is a set of about forty separate capabilities. A process can hold all of them, some of them, or none.
- CapEff: what a process may do now: CapEff in /proc/<pid>/status is the truth about a process's privileges. The uid is only a hint.
- Bind port 80 as alice
- Drill: one privilege, not root
- The same command, now allowed: A file capability grants one named privilege to whoever runs that program. Setuid root grants all of them.
- Why ping is no longer setuid
What a container keeps
- Root in a container is a smaller root: A container's root starts with about a third of root's capabilities. Dropping is the default; the dangerous ones are already gone.
- What --privileged hands back: privileged: true means host root with a different view of the filesystem. Treat it as giving the workload the node.
- Break it: drop everything: uid 0 with no capabilities is an ordinary user with a special name. The kernel checks capabilities, not the uid.
A filter on the boundary
- seccomp: a filter on the door: Capabilities decide which privileged calls succeed. seccomp decides which calls can be made at all, and it is asked first.
- Break it: a call the filter refuses
- The strict action: kill: A seccomp denial is either an errno the program sees or a kill the kernel logs. When a process dies with no message, read dmesg.
The same knobs in Kubernetes
- Four fields in a securityContext: A securityContext is a list of kernel settings: capability sets, the seccomp filter, no_new_privs. Verify it in /proc/<pid>/status on the node.
Recap & playground
- Cheat sheet
- Playground