Troubleshooting Workloads (CKA)
Eight broken workloads, one habit: read the STATUS column as a diagnosis, then run the one command that holds the reason.
An interactive Kubernetes lesson: 23 steps, about 32 minutes, on a live simulation in your browser.
It is Monday, 09:00, and you are on call for the shop. web calls the api Service and every request comes back 200. Friday evening's release touched eight other workloads, and the tickets are already waiting.
Every ticket starts with kubectl get pods. A pod goes through the same stages every time: the scheduler picks a node, the kubelet pulls the image, builds the container from its configuration, runs any init containers, starts the app, then probes it.
What you will learn
The board
- STATUS is a diagnosis: STATUS names the stage where a pod is stuck: scheduling, pulling, configuring, initialising, running or probing. The stage tells you who to ask.
It never started
- Pending: nobody took the pod: Pending with no node is the scheduler's problem, and its FailedScheduling event counts every node and says why each one was rejected.
- Break it: a tag that does not exist
- Same status, different cause: ImagePullBackOff has two causes. Not found: the name or tag is wrong. Denied or unauthorized: the node has no credentials. The event says which.
- CreateContainerConfigError: CreateContainerConfigError: the image is on the node, but the spec points at a ConfigMap, Secret or key that does not exist in this namespace.
- Init:0/1: stuck before the app: Init:N/M means init container N+1 has not finished. describe gives its name; kubectl logs -c <that name> gives its reason.
It started, then died
- CrashLoopBackOff: it started and died: A crashing pod of a Deployment is a symptom of its template. Deleting the pod buys a fresh copy of the same bug.
- Read the dead container: CrashLoopBackOff: describe shows how it died (Last State, Exit Code), logs --previous shows why (its last words).
- A crash with no last words: OOMKilled leaves no last words. Reason OOMKilled and exit code 137 in describe are the only record; the fix is the memory limit or the leak.
Nothing crashed, still broken
- Break it: Running, and serving nobody: Running only means the process exists. READY says whether it gets traffic. Read both columns, every time.
- A Job that fails: A failed Job leaves its evidence behind: one Error pod per attempt. describe job for the verdict, logs on a pod for the reason.
Output streams
- Where logs come from: kubectl logs reads a file the container runtime writes on the node: stdout and stderr, merged, kept only as long as the pod.
- An app that logs to a file: If it is not written to stdout or stderr, kubectl logs cannot see it.
- A sidecar turns the file into a stream: One output stream per container. -c picks a container, --all-containers takes them all, and a tailing sidecar gives a file a stream of its own.
- Cut a stream down to size
Resource usage
- Usage: install the meter: kubectl top is live usage, served by metrics-server. No metrics-server: no top, and no autoscaling on CPU or memory either.
- Idle and full at the same time: top shows what is used. The scheduler counts what is requested. A node can be idle and full at the same time.
The decision tree
- The decision tree: Did it get a node? Was the container created? Did it stay up? Is it Ready? The first no picks the command.
- Drill: the container that died
- Drill: who is using the memory
- Drill: every replica at once
Recap & playground
- Cheat sheet
- Playground