Autoscaling & Self-healing (CKA)
Let the cluster add replicas under load and replace what breaks.
An interactive Kubernetes lesson: 20 steps, about 32 minutes, on a live simulation in your browser.
It is the day before the shop's big sale. The backend, api, runs as a Deployment with replicas: 3, and shoppers reach it through a load balancer.
That number is not a one-time instruction. The Deployment owns a ReplicaSet, and the ReplicaSet runs a loop that never stops: count the pods matching my selector, compare with replicas, create or delete pods to close the gap.
What you will learn
A loop that keeps the count
- Desired state, enforced: A controller is a loop: observe, compare with desired, act. Self-healing and autoscaling are the same loop with different inputs.
- The kubelet restarts containers: The kubelet heals a container by restarting it inside the same pod. The pod, its IP and its node do not change.
- The ReplicaSet replaces pods: A ReplicaSet does not repair pods. It counts them, and creates or deletes until the count matches.
When a node dies
- Break it: a node goes dark: A silent node is first taken out of traffic, not replaced. Endpoints react in seconds; rescheduling waits.
- The five-minute wait: Node failure recovery = detect (under a minute) + tolerationSeconds (300 by default) + start a new pod. Plan for about six minutes on one replica fewer.
- The node comes back: Deleting a pod is a request; the kubelet's confirmation completes it. An unreachable node leaves its pods Terminating until it returns or is removed.
Disruption on purpose
- A budget for planned disruption: A PDB is a rule for the eviction API: allow this eviction only if enough matching pods stay Ready. It turns a drain into a paced move.
- Break it: a budget that blocks everything
- Drill: create a PodDisruptionBudget
Scaling on CPU
- Sale day: the load arrives: Self-healing defends a number. Autoscaling chooses the number. They are separate controllers.
- Break it: an autoscaler with no eyes
- The formula: desired = ceil(current × currentUtil ÷ targetUtil). The HPA changes the replica count by the same ratio the metric is off by.
- Utilisation of what?: HPA utilisation is a percentage of the pod's request, not of the node and not of the limit. No request, no percentage, no scaling.
Limits and coming back down
- minReplicas and maxReplicas: min and max are hard walls around the formula. max protects your nodes and your bill; min protects you from scaling down to too few.
- Scaling down is deliberately slow: Scale up fast, scale down slow. The HPA acts on the highest recommendation of the last five minutes, so brief dips do not remove capacity.
- Drill: autoscale a Deployment
Nodes and pod size
- When the nodes run out: HPA adds pods when CPU is high. Cluster Autoscaler adds nodes when pods are Pending. Neither does the other's job.
- The other axis: bigger pods: Horizontal = more pods (HPA). Vertical = bigger pods (VPA, in-place resize). More nodes = Cluster Autoscaler. Three controllers, three axes.
Recap & playground
- Cheat sheet
- Playground