Requests, Limits & Quotas (CKAD)
Requests reserve room on a node. Limits are enforced by the kernel. Quotas cap what a whole namespace may ask for.
An interactive Kubernetes lesson: 20 steps, about 30 minutes, on a live simulation in your browser.
The shop shares two worker nodes, each with 2 CPUs and 4Gi of memory. api answers web, and a batch workload crunches numbers beside it. Until now the api pods declared nothing about what they need.
To the scheduler, a pod that declares nothing needs nothing. It can be squeezed onto any node, next to anything, with no promise that CPU or memory will be there when it wants them.
What you will learn
Requests: what the scheduler reserves
- A request is a reservation: A request is a reservation on a node. The scheduler adds up requests to decide whether a pod fits.
- Full on paper, idle in practice: A node is full when the requests on it add up to its capacity, even if every pod on it is idle.
- The scheduler's ledger: describe node is the scheduler's ledger: requests per pod, and their total under Allocated resources.
Limits: what the kernel enforces
- A limit is a ceiling: Requests are for the scheduler: what is reserved. Limits are for the kernel: what is enforced.
- More CPU than the limit: CPU is compressible. Over the limit a container is slowed down, never killed.
- Break it: more memory than the limit: Memory is not compressible. Over the limit the kernel kills the container: OOMKilled, exit code 137, restart.
- Choosing the numbers
Who goes first when a node is full
- Three classes, never set by hand: QoS is derived, not declared. Requests equal to limits everywhere: Guaranteed. Nothing set: BestEffort. Anything else: Burstable.
- The node runs out of memory: Under node pressure, pods using more than they requested are evicted first. BestEffort requested nothing, so it is always first in line.
Guardrails for a namespace
- LimitRange: defaults at the door: A LimitRange edits pods on the way in: containers that say nothing get the namespace's default request and limit.
- Break it: ask for more than the maximum
- ResourceQuota: a namespace budget: A LimitRange bounds each container. A ResourceQuota bounds the sum of the whole namespace.
- A workload meets the quota: A quota refusal lands on whoever creates the pod: a StatefulSet, a Job, the ReplicaSet behind a Deployment. Look for FailedCreate in its events.
- Break it: a quota with no defaults: A quota on CPU or memory makes requests mandatory. A LimitRange default is what keeps that from rejecting every plain pod.
Reading the numbers
- Promised versus used: describe node: what was promised. top: what is used. Right-sizing is closing the gap between them.
- Units: m, Mi and M: CPU: 1000m is one core. Memory: Mi is 1024-based, M is 1000-based, and m is never what you want.
Exam speed
- Drill: resources on a Deployment
- Drill: a quota in one line
Recap & playground
- Cheat sheet
- Playground