learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

HA Control Plane & etcd (CKA)

Quorum, leader election, and backing up the one database that holds everything.

An interactive Kubernetes lesson: 19 steps, about 30 minutes, on a live simulation in your browser.

The shop's first cluster had one control-plane node. When that machine rebooted for a kernel patch, the pods kept serving, but for six minutes nobody could deploy, nothing was rescheduled and no crashed pod was replaced. If its disk had died instead, the cluster's entire memory would have gone with it.

So the new cluster has three control-plane nodes: cp-1, cp-2 and cp-3. Each one runs its own API server, etcd member, scheduler and controller manager.

What you will learn

  1. One control plane is one failure

    • Three of everything
    • Lose a control-plane node: A highly available control plane is three complete copies behind one address. Any one of them can disappear.
  2. How three copies cooperate

    • One address in front of three API servers: API servers are stateless and all active. Clients talk to one load-balanced address, never to a particular node.
    • One scheduler works, two wait: API servers are all active. The scheduler and controller manager are one active, the rest on standby, and a Lease decides who.
    • Adding control-plane nodes
    • Stacked or external etcd: Stacked: etcd lives on the control-plane nodes and fails with them. External: etcd has its own machines and its own failures.
  3. etcd and the majority rule

    • A write needs a majority: An etcd write needs a majority of all members, counted against the full membership, not against whoever is still alive.
    • One etcd member down
    • Two etcd members down: Without a majority, etcd stops writing rather than risk two versions of the truth. The cluster freezes. The pods do not.
    • Break it: frozen, not dead
    • Would a fourth member help?: Use an odd number of etcd members: 3 tolerates one failure, 5 tolerates two. An even number buys nothing.
  4. Backup and restore

    • Take a snapshot: A Kubernetes backup is an etcd snapshot. Every object in the cluster is in that one file.
    • Break it: a delete, replicated three times
    • Restore the snapshot: A restore is a rewind, not a merge. Everything written after the snapshot is forgotten.
    • The restore, step by step
    • Drill: save a snapshot
    • Drill: restore to a new directory
  5. Recap & playground

    • Cheat sheet
    • Playground