learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Logs & Troubleshooting

A method for a machine that misbehaves, then six incidents to practise it on: CPU, memory, disk, a service, the network and a permission.

An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.

That is the whole ticket. web01 runs the shop: nginx on port 80, the backend shop-api on 8080, and a database on another host. You have a shell and no idea where to look.

People who are good at this follow a loop. What changed? Read the error. Look at the logs. Form one hypothesis. Test it. Fix the cause, not the symptom. Before any of that they spend one minute collecting the same eight facts, every time, so that nothing obvious is missed. This chapter is that minute, on a healthy machine, because you cannot spot abnormal until you know normal.

What you will learn

  1. The first sixty seconds

    • "web01 is slow": Load average is a queue length. Divide it by the number of CPUs: under 1 there is spare capacity, over 1 work is waiting.
    • Read top from the top: In top, read the summary before the table: load, task states, the CPU split, memory. The table only tells you who; the summary tells you what kind of problem.
    • free and df: the two that fill up: Free memory is wasted memory. Read "available", not "free", to know whether the machine is short.
    • Ask what is already complaining: Sixty seconds, eight commands: uptime, top, free -h, df -h, dmesg, journalctl -p err -b, systemctl --failed, ss -tlnp. Then start thinking.
  2. Where the evidence is

    • Where logs live
    • Filter by unit, time and priority: Narrow a log three ways: which unit, which minutes, how severe. An empty result is evidence as well.
    • Drill: attach the log
  3. Incidents: CPU and memory

    • Incident 1: pages are slow: Observe before you act. The next command should be the one that splits the possibilities, and it should be read-only.
    • Who started it, and when
    • Drill: stop the runaway jobs
    • Incident 2: memory that never comes back: One reading is a number, two readings are a trend. Shrinking "available" with shrinking cache means real memory pressure.
    • "It only said Killed": A process that vanishes with no error was killed from outside. When memory runs out, the kernel's OOM killer picks a victim and says so in dmesg.
  4. Incident: the disk is full

    • Incident 3: every write fails: When several unrelated things break at once, look for the one resource they share: disk, memory, DNS, the clock.
    • df says where, du says what: df finds the full filesystem. du, one level at a time, finds the directory. It is nearly always a log, a cache or a core dump.
    • Delete it and check again: A file's space is freed when its last name is removed and its last open descriptor is closed. rm only does the first half.
    • The file du cannot see: df and du disagree when a deleted file is still open. lsof +L1 names the process; closing the file, not deleting it, frees the space.
  5. Incident: it will not start

    • Incident 4: nginx will not start: systemctl tells you that a start failed. The journal for that unit tells you why, in the daemon's own words.
    • Who has port 80?: One hypothesis, one test. Pick the command whose output would be different if you were wrong.
    • A config that will not load: Test a config before you load it, and prefer reload to restart: a failed reload leaves the old process serving.
  6. Incidents: unreachable, denied

    • Incident 5: "cannot reach db01": Test a connection one layer at a time, nearest first: interface, route, gateway, name, host, port. The first link that fails is the diagnosis.
    • The host is back, the app still fails: Unreachable means no host or no route. Refused means the host answered and nothing is listening on that port. A timeout means something is silently dropping packets.
    • Incident 6: the deploy fails for bob
    • Fix the cause, not the symptom: A fix is finished when you can say what caused the problem and why it will not come back. Anything less is a workaround.
  7. Recap & playground

    • Cheat sheet
    • Playground: you are on call