learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Autoscaling & Self-healing (CKA)

Let the cluster add replicas under load and replace what breaks.

An interactive Kubernetes lesson: 20 steps, about 32 minutes, on a live simulation in your browser.

It is the day before the shop's big sale. The backend, api, runs as a Deployment with replicas: 3, and shoppers reach it through a load balancer.

That number is not a one-time instruction. The Deployment owns a ReplicaSet, and the ReplicaSet runs a loop that never stops: count the pods matching my selector, compare with replicas, create or delete pods to close the gap.

What you will learn

  1. A loop that keeps the count

    • Desired state, enforced: A controller is a loop: observe, compare with desired, act. Self-healing and autoscaling are the same loop with different inputs.
    • The kubelet restarts containers: The kubelet heals a container by restarting it inside the same pod. The pod, its IP and its node do not change.
    • The ReplicaSet replaces pods: A ReplicaSet does not repair pods. It counts them, and creates or deletes until the count matches.
  2. When a node dies

    • Break it: a node goes dark: A silent node is first taken out of traffic, not replaced. Endpoints react in seconds; rescheduling waits.
    • The five-minute wait: Node failure recovery = detect (under a minute) + tolerationSeconds (300 by default) + start a new pod. Plan for about six minutes on one replica fewer.
    • The node comes back: Deleting a pod is a request; the kubelet's confirmation completes it. An unreachable node leaves its pods Terminating until it returns or is removed.
  3. Disruption on purpose

    • A budget for planned disruption: A PDB is a rule for the eviction API: allow this eviction only if enough matching pods stay Ready. It turns a drain into a paced move.
    • Break it: a budget that blocks everything
    • Drill: create a PodDisruptionBudget
  4. Scaling on CPU

    • Sale day: the load arrives: Self-healing defends a number. Autoscaling chooses the number. They are separate controllers.
    • Break it: an autoscaler with no eyes
    • The formula: desired = ceil(current × currentUtil ÷ targetUtil). The HPA changes the replica count by the same ratio the metric is off by.
    • Utilisation of what?: HPA utilisation is a percentage of the pod's request, not of the node and not of the limit. No request, no percentage, no scaling.
  5. Limits and coming back down

    • minReplicas and maxReplicas: min and max are hard walls around the formula. max protects your nodes and your bill; min protects you from scaling down to too few.
    • Scaling down is deliberately slow: Scale up fast, scale down slow. The HPA acts on the highest recommendation of the last five minutes, so brief dips do not remove capacity.
    • Drill: autoscale a Deployment
  6. Nodes and pod size

    • When the nodes run out: HPA adds pods when CPU is high. Cluster Autoscaler adds nodes when pods are Pending. Neither does the other's job.
    • The other axis: bigger pods: Horizontal = more pods (HPA). Vertical = bigger pods (VPA, in-place resize). More nodes = Cluster Autoscaler. Three controllers, three axes.
  7. Recap & playground

    • Cheat sheet
    • Playground