Probes & Health Checks (CKAD)
Three questions the kubelet asks your container, and what it does with each answer.
An interactive Kubernetes lesson: 20 steps, about 30 minutes, on a live simulation in your browser.
The shop's backend is a Deployment, api, with three replicas behind a Service. Customers' requests are spread across the three pods and every one succeeds.
The pod template has an image and a port, and nothing else. So ask what the kubelet, the agent on each node, actually knows about these containers. It knows one thing: whether the process is still there. If the process exits, the kubelet restarts it. While the process exists, the pod is Running, and with no other information it is also counted as Ready.
What you will learn
What Kubernetes cannot see
- Three pods, no probes: With no probes, Running and Ready both mean one thing: the process has not exited.
- The process hangs: Kubernetes only knows what you teach it to ask. A probe is a question the kubelet puts to your container on a timer.
- Three probes, three questions: Startup: are you up yet? Liveness: are you stuck? Readiness: can you take work now? Same mechanism, three different consequences.
Readiness: should it get traffic?
- Add a readiness probe
- One pod loses its database: A failed readiness probe takes the pod out of the Service. Nothing is killed, and the other pods carry the load.
- Readiness lasts the pod's whole life: Readiness is a switch for traffic that flips both ways, for the whole life of the pod. It never restarts anything.
Liveness: should it be restarted?
- Add a liveness probe
- The deadlock comes back: A failed liveness probe means: kill this container and start it again. It is the fix for a process that cannot recover by itself.
- Break it: a liveness probe too eager: Liveness should only fail for problems a restart can fix. Put a dependency in it and one slow database restarts your whole fleet.
- Shallow liveness, deep readiness
How a probe asks, and how often
- Four ways to ask
- Five numbers that set the pace: Reaction time ≈ periodSeconds × failureThreshold. With the defaults that is about 30 seconds before anything happens.
Startup: has it finished booting?
- Break it: a slow starter gets killed
- A startup probe holds the others back: A startup probe buys boot time: failureThreshold × periodSeconds. After its first success it never runs again and liveness takes over.
Probes guard your rollouts
- Ship a version that cannot serve: Readiness is the gate of a rolling update: a version that never reports Ready never receives traffic and never replaces a working pod.
- Find the cause, then undo
At exam speed
- Drill: get a manifest to edit
- Drill: why is it not Ready?
Recap & playground
- Cheat sheet
- Playground