Boot, Kernel & Modules
Firmware, GRUB, kernel, initramfs, systemd: what each stage reads, what survives a reboot, how to rescue a machine that will not boot, and how kernel modules load.
An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.
A boot is a relay. The firmware (UEFI on this VM) finds a disk and starts the boot loader. GRUB reads its menu, loads a kernel and an initramfs (a small compressed filesystem with just enough drivers to find the real root disk) and passes the kernel a command line. The kernel sets up memory and devices, mounts the root filesystem and starts pid 1. systemd, as pid 1, starts units until it reaches the default target, and you can log in.
Each stage leaves a trace. uname -r is the kernel GRUB chose: 6.8.0-45-generic. /proc/cmdline is the line it was given: the image, root=UUID=…, ro, quiet splash. systemd-analyze times the stages: 4.218s (firmware) + 1.746s (loader) + 2.911s (kernel) + 14.302s (userspace). Ubuntu's initramfs reports no time of its own, so it is inside "kernel".
What you will learn
From power-on to login
- Five handovers: Firmware finds GRUB, GRUB loads a kernel and initramfs with a command line, the kernel mounts root and starts systemd, systemd starts the default target. A boot fails at exactly one of those handovers.
- What the boot waited for: blame lists durations; critical-chain shows what the boot waited for. Only a unit on the chain can make the boot faster.
What survives a reboot
- Four changes, one reboot: A reboot keeps only what is on disk. Every runtime change you want to keep needs a file that the boot reads.
- The journal remembers the last boot: journalctl -b is this boot, -b -1 the previous one. After an unexpected reboot, read the end of -b -1.
When the boot stops
- One wrong line in /etc/fstab: Every /etc/fstab line without nofail is required for the boot. One missing device means 90 seconds of waiting, then emergency mode with no ssh.
- Read why it stopped: In emergency mode read journalctl -b -p err from the bottom up, then check the fstab line against blkid.
- Drill: repair /etc/fstab
- Fixed. Now what?: Fix the cause in emergency mode, verify it, then exit: the same boot continues. Reboot only to prove the fix survives one.
GRUB and the kernel command line
- Edit GRUB's settings, reboot
- update-grub writes what GRUB reads: /etc/default/grub is the source, /boot/grub/grub.cfg is what boots. Nothing changes until update-grub regenerates grub.cfg.
- One boot with different parameters: Pressing e at the GRUB menu changes the kernel command line for one boot. systemd.unit=rescue.target is the gentlest way into a machine whose services break the boot.
- Locked out of root
- init=/bin/bash: a root shell as pid 1: init=/bin/bash replaces systemd with a root shell: no password, no services, and / still read-only until you remount it.
- Remount, repair, reboot -f: In an init=/bin/bash shell: mount -o remount,rw /, fix the one thing, sync, reboot -f.
Targets
- Targets, and switching to one: set-default changes the next boot; isolate changes the machine now. Isolating rescue.target over ssh cuts your own session.
Kernels and modules
- Install a kernel, check uname: Installing a kernel changes /boot, not the running system. The fix you installed protects nothing until you reboot into it.
- Boot the old kernel, once: grub-reboot ENTRY boots that entry once and then falls back to the default: the safe way to test or roll back a kernel.
- Modules and their dependencies: modprobe loads a module and everything it depends on, in order. lsmod's Used by column shows who holds each module.
- A module that will not leave: rmmod removes one module and refuses while it is held. modprobe -r removes a module and the dependencies that it alone was holding.
- Drill: br_netfilter, now and at boot
- Module options: Module parameters are set at load time: on the modprobe line or in /etc/modprobe.d. /sys/module/NAME/parameters shows the values the module is really using.
- A blacklist is not a ban: blacklist stops a module loading by itself; install NAME /bin/false stops it loading at all.
Recovery, cheat sheet & playground
- When a server will not boot: Recover from the top: console, read the error, previous kernel, rescue target, init=/bin/bash, then the disk from another machine. Each step needs less of the system to work.
- Cheat sheet
- Playground: make the next boot boring