Scheduling (CKA)
How the scheduler picks a node, and every way you can steer it.
An interactive Kubernetes lesson: 24 steps, about 35 minutes, on a live simulation in your browser.
Your shop's frontend, web, starts with three replicas. A freshly created pod is only a record in etcd with an empty spec.nodeName. It runs nowhere, so it sits in the Pending tray.
kube-scheduler watches for exactly that: pods with no node. For each one it works in two phases. Filter removes every node that cannot run the pod. Score ranks the nodes that are left. The winner's name is written into spec.nodeName, and the kubelet on that node starts the containers.
What you will learn
Filter, then score
- A pod with no node: The scheduler does one job: fill in spec.nodeName. Filter decides where a pod can run; score decides where it should.
- nodeName skips the scheduler: nodeName is an order, not a request. It bypasses filter and score entirely, taints included.
- A pod too big for any node: The scheduler counts requests, never usage. A node is full when its requests add up to allocatable, even if it is idle.
- Reading FailedScheduling: FailedScheduling is a tally of nodes by rejection reason. The counts add up to the node total; fix any one group and the pod can land.
Pulling pods toward nodes
- nodeSelector: only nodes with this label: nodeSelector is a hard filter on node labels: every listed key=value must be present, or the node is not considered.
- Break it: a label no node has
- Node affinity: required and preferred: Required affinity is a filter. Preferred affinity is a bonus in scoring. Only a filter can leave a pod Pending.
Keeping pods away
- A taint pushes pods away: A label on a node attracts pods that select it. A taint repels every pod that does not tolerate it. They solve opposite problems.
- The cache gets restarted: Every hard rule removes nodes, and a pod needs one node that survives all of them. Rules never cancel each other out.
- A toleration opens the door: Taint on the node is the lock; toleration on the pod is the key. Key, value and effect must line up.
- A toleration does not attract: A toleration is permission, not direction. A dedicated node needs both halves: a taint to keep others out and a label plus selector to bring its own pods in.
- Break it: NoExecute evicts: NoSchedule guards the door. NoExecute also clears the room: running pods that do not tolerate it are evicted.
- PreferNoSchedule: a soft no: Three taint effects: NoSchedule filters new pods, PreferNoSchedule only lowers the score, NoExecute filters and evicts.
Pods relative to pods
- Anti-affinity: never two on one node: Pod anti-affinity: not in the same topology domain as pods matching this selector. With topologyKey hostname, a domain is one node.
- Pod affinity and the topologyKey: topologyKey names a node label. Nodes with the same value for it are one domain: hostname means per node, zone means per zone.
- Spread evenly across zones: Topology spread balances pod counts between domains. maxSkew is the largest allowed gap between the fullest and the emptiest domain.
Full clusters and maintenance
- Priority and preemption: Priority orders the scheduling queue. Preemption goes further: a pod that fits nowhere may evict lower-priority pods to make room.
- Cordon: no new pods here: Cordon closes a node to new pods and leaves the running ones alone. It is NoSchedule with a friendlier command.
- Drain: evict everything, politely
Pending, decoded
- Four reasons in one line: Pending means the filter left zero nodes. describe the pod, split the message by reason, and remove one reason.
- Drill: empty a node for maintenance
- Drill: taint a node
Recap & playground
- Cheat sheet
- Playground