learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Syscalls, Capabilities & seccomp

A program can do nothing on its own. Watch it ask the kernel, then see how root's power is split up and how a filter decides which requests get through.

An interactive Linux lesson: 24 steps, about 35 minutes, on a live simulation in your browser.

cat /etc/hostname prints web01. It looks as if cat read a file and wrote to your screen. It did neither. cat has no way to touch a disk, a terminal, a network card or another process.

A Linux machine runs code in two modes. The kernel runs in kernel space, with full access to hardware and to every process's memory. Everything else, from cat to a database, runs in user space, where the CPU refuses those operations outright.

What you will learn

  1. Asking the kernel

    • A program cannot open a file: User space cannot touch files, the network or other processes. It can only ask the kernel, one system call at a time.
    • Watch cat ask: A running program is a sequence of system calls. strace prints that sequence, so you can watch what it does instead of guessing.
    • Read one line, field by field: name(arguments) = return value. For open the return is a file descriptor; for read and write it is a byte count.
    • When the kernel says no: A failed syscall returns -1 and an errno. Every error message a program prints started life as one of those codes.
  2. strace as a debugger

    • Break it: an error that says nothing: When the message is useless, trace the program and read the last failing call: it names the file and the reason.
    • From errno to fix
    • A different errno, a different fix: The errno names the kind of fault: EACCES is a permission check, ENOENT is a path that does not exist. Read it before you reach for chmod.
    • strace -c: the summary
    • Drill: keep the evidence
  3. Root is not one switch

    • Root's power comes in pieces: Root is a set of about forty separate capabilities. A process can hold all of them, some of them, or none.
    • CapEff: what a process may do now: CapEff in /proc/<pid>/status is the truth about a process's privileges. The uid is only a hint.
    • Bind port 80 as alice
    • Drill: one privilege, not root
    • The same command, now allowed: A file capability grants one named privilege to whoever runs that program. Setuid root grants all of them.
    • Why ping is no longer setuid
  4. What a container keeps

    • Root in a container is a smaller root: A container's root starts with about a third of root's capabilities. Dropping is the default; the dangerous ones are already gone.
    • What --privileged hands back: privileged: true means host root with a different view of the filesystem. Treat it as giving the workload the node.
    • Break it: drop everything: uid 0 with no capabilities is an ordinary user with a special name. The kernel checks capabilities, not the uid.
  5. A filter on the boundary

    • seccomp: a filter on the door: Capabilities decide which privileged calls succeed. seccomp decides which calls can be made at all, and it is asked first.
    • Break it: a call the filter refuses
    • The strict action: kill: A seccomp denial is either an errno the program sees or a kill the kernel logs. When a process dies with no message, read dmesg.
  6. The same knobs in Kubernetes

    • Four fields in a securityContext: A securityContext is a list of kernel settings: capability sets, the seccomp filter, no_new_privs. Verify it in /proc/<pid>/status on the node.
  7. Recap & playground

    • Cheat sheet
    • Playground