learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Troubleshooting the Cluster (CKA)

Nodes go NotReady and control-plane components die. Check the API, the nodes, the system pods and the workload, in that order.

An interactive Kubernetes lesson: 22 steps, about 34 minutes, on a live simulation in your browser.

Game day. A colleague has root on every machine and a list of ways to break a cluster. You have kubectl, SSH and a shop that must keep serving: web calling three api pods through a Service.

In the last lesson one workload was broken and the cluster was fine. Today it is the other way round, and a broken cluster makes every workload look guilty. So you never start at the workload. You ask four questions, from the outside in.

What you will learn

  1. Outside in

    • Four questions, outside in: API, nodes, system pods, workloads: check them in that order, because each one only works if the one before it does.
  2. A node goes NotReady

    • A node goes quiet: Ready=Unknown: the kubelet stopped talking. Ready=False: the kubelet is talking and says what is wrong. Both print as NotReady.
    • Go to the node: The kubelet is a systemd service, not a pod: systemctl status says whether it runs, journalctl -u kubelet says why not.
    • Five minutes of silence: The control plane can only ask. A pod is not gone until the kubelet on its node confirms it.
    • Fix it, reload, restart: Edit, daemon-reload, restart. And a recovered node stays empty until something creates new pods.
    • Break it: every node NotReady at once
    • What a node needs to be Ready: A node is Ready when four things hold: kubelet running, runtime answering, CNI configured, no resource pressure.
  3. Control-plane components

    • The control plane is pods on disk: Control-plane components are files in /etc/kubernetes/manifests, run by the kubelet. Edit the file to change one. Move the file out to stop one.
    • Pending, and no events: Pending with a FailedScheduling event: the scheduler said no. Pending with no events: there is no scheduler.
    • Debug the scheduler like any pod
    • Break it: nothing reconciles: Objects change and nothing acts on them: the controller manager is down. It is the only thing that turns a desired count into pods.
  4. When kubectl stops working

    • kubectl itself stops working: When kubectl is refused, drop one layer: crictl on the control-plane node talks to the runtime and needs no API.
    • No container to read: Container died: crictl logs. Container never created: journalctl -u kubelet. Container already removed: /var/log/pods.
    • Break it: the API server loses etcd: An API server failure is often etcd's failure one step removed. Read the API server's log and follow the address it cannot reach.
  5. Outages that are not outages

    • The outage that is only yours: 'Connection refused' names the address kubectl dialled. Check it against your kubeconfig before you touch the cluster.
    • Certificates run out: Refused: nothing is listening. Timeout: something is dropping. x509: you reached it, and the certificate is expired or not trusted.
  6. Symptom to component

    • From symptom to component
    • Drill: ask the kubelet
    • Drill: start it and keep it started
    • Drill: when do the certificates expire
  7. Recap & playground

    • Cheat sheet
    • Playground