Scheduling: cron & Timers
Run things later and on a schedule with cron and systemd timers, and find out why a job did not run: lost output, a different environment, a stray %, overlapping runs and a machine that was off.
An interactive Linux lesson: 24 steps, about 34 minutes, on a live simulation in your browser.
The shop on web01 needs work done when nobody is logged in: a health check every few minutes, a backup every night. Linux has done this for fifty years with cron, a daemon that wakes once a minute, reads its tables, and starts whatever is due.
It is already running: pid 640, started at boot by systemd ("systemd & Services" covers how). Note the -P in its command line; it matters later. Each user has their own table, a crontab, and alice has none yet.
What you will learn
A clock that runs commands
- cron: one daemon, awake every minute: cron is one daemon, separate from your shell. Once a minute it reads every table and starts whatever is due, as the table's owner.
- Five fields and a command: minute hour day-of-month month day-of-week, then a command. The job runs in each minute where all five fields match.
- Let an hour pass
Where the output goes
- Break it: the shop goes down, nobody hears
- Send the output somewhere yourself: Whatever a cron job prints is mailed, and with no mail server it is thrown away. Redirect stdout and stderr to a file in the crontab line itself.
Works in my terminal
- Break it: backup: not found
- The environment a cron job really gets: A cron job gets a tiny environment: systemd's PATH, /bin/sh, your home as the working directory, and nothing from your dotfiles. Write every job as if no one had ever logged in.
- A date in a file name: In a crontab line, % means newline. Escape it as \% or move the command into a script, where % is ordinary.
- Drill: make the backup run from cron
System jobs
- /etc/cron.d: system jobs with a user field: System jobs live in /etc/cron.d with a user field after the schedule. A file name with a dot in it is silently ignored.
- /etc/cron.daily and run-parts: /etc/cron.hourly, daily, weekly and monthly are directories of scripts that run-parts runs in name order. No dots in the names, and the scripts must be executable.
Slow jobs
- A one-minute job that takes 90 seconds: cron starts a job on schedule whether or not the last run has finished. A job slower than its interval overlaps itself.
- flock -n: one run at a time: flock -n LOCKFILE CMD: if the last run still holds the lock, this run exits at once. One line, and overlaps are impossible.
systemd timers
- The health check as a systemd timer: A timer is two units: NAME.service says what to run, NAME.timer says when. Enable the timer, not the service.
- An hour of a 15-minute timer: A timer's job logs to the journal under its own unit name: journalctl -u NAME.service is its whole history, output included.
- Test a schedule before you trust it: Never install a schedule you have not tested: systemd-analyze calendar shows how systemd read it and when it fires next.
- Drill: a nightly timer
- cron or a timer?: cron is one line and no guarantees. A timer is two files and gets logging, status, no overlap and catch-up for free.
Missed runs and time zones
- The first night
- The machine is off at 02:00: cron never catches up: a minute missed is gone. Persistent=true timers and anacron remember the last run and catch up after boot.
- 02:00 in which time zone?: The clock counts UTC. cron and OnCalendar read their times in the machine's local zone, so changing the zone moves every schedule.
Recap & playground
- "My cron job didn't run": the checklist: Debug a scheduled job in order: did it start (CMD line), what did it print, how did it exit, does it work in cron's environment.
- Cheat sheet
- Playground