Troubleshooting the Cluster (CKA)
Nodes go NotReady and control-plane components die. Check the API, the nodes, the system pods and the workload, in that order.
An interactive Kubernetes lesson: 22 steps, about 34 minutes, on a live simulation in your browser.
Game day. A colleague has root on every machine and a list of ways to break a cluster. You have kubectl, SSH and a shop that must keep serving: web calling three api pods through a Service.
In the last lesson one workload was broken and the cluster was fine. Today it is the other way round, and a broken cluster makes every workload look guilty. So you never start at the workload. You ask four questions, from the outside in.
What you will learn
Outside in
- Four questions, outside in: API, nodes, system pods, workloads: check them in that order, because each one only works if the one before it does.
A node goes NotReady
- A node goes quiet: Ready=Unknown: the kubelet stopped talking. Ready=False: the kubelet is talking and says what is wrong. Both print as NotReady.
- Go to the node: The kubelet is a systemd service, not a pod: systemctl status says whether it runs, journalctl -u kubelet says why not.
- Five minutes of silence: The control plane can only ask. A pod is not gone until the kubelet on its node confirms it.
- Fix it, reload, restart: Edit, daemon-reload, restart. And a recovered node stays empty until something creates new pods.
- Break it: every node NotReady at once
- What a node needs to be Ready: A node is Ready when four things hold: kubelet running, runtime answering, CNI configured, no resource pressure.
Control-plane components
- The control plane is pods on disk: Control-plane components are files in /etc/kubernetes/manifests, run by the kubelet. Edit the file to change one. Move the file out to stop one.
- Pending, and no events: Pending with a FailedScheduling event: the scheduler said no. Pending with no events: there is no scheduler.
- Debug the scheduler like any pod
- Break it: nothing reconciles: Objects change and nothing acts on them: the controller manager is down. It is the only thing that turns a desired count into pods.
When kubectl stops working
- kubectl itself stops working: When kubectl is refused, drop one layer: crictl on the control-plane node talks to the runtime and needs no API.
- No container to read: Container died: crictl logs. Container never created: journalctl -u kubelet. Container already removed: /var/log/pods.
- Break it: the API server loses etcd: An API server failure is often etcd's failure one step removed. Read the API server's log and follow the address it cannot reach.
Outages that are not outages
- The outage that is only yours: 'Connection refused' names the address kubectl dialled. Check it against your kubeconfig before you touch the cluster.
- Certificates run out: Refused: nothing is listening. Timeout: something is dropping. x509: you reached it, and the certificate is expired or not trusted.
Symptom to component
- From symptom to component
- Drill: ask the kubelet
- Drill: start it and keep it started
- Drill: when do the certificates expire
Recap & playground
- Cheat sheet
- Playground