Kubernetes Node Allocatable: Reservations, Eviction Headroom, and Real Capacity
Understand how Kubernetes computes node allocatable capacity, sizes kube and system reservations, budgets eviction headroom, and verifies enforcement per node pool.
Node.status.capacity and Node.status.allocatable answer different operational questions. Capacity is what the node reports as installed resource capacity; allocatable is the amount Kubernetes exposes for Pods after configured reservations and eviction headroom. The scheduler accounts Pod requests against allocatable, not against the advertised size of the VM and not against a live CPU or memory sample. If an operator sizes from capacity alone, system daemons and kubelet recovery work compete with application Pods for the same finite host.
Node Allocatable is a capacity contract between kubelet and scheduler, not a physical partition by itself. It can prevent the scheduler from placing more requested workload than the node is budgeted to host, and kubelet can enforce some of those boundaries. Whether host daemons are actually confined to their reservation depends on cgroup setup and kubelet enforcement settings. A reservation that only changes reported allocatable is not a magically fenced block of RAM or CPU.
Start with the node’s advertised budget
For the resources Kubernetes uses for Pod placement, the useful mental model is:
Pod allocatable ~= capacity - kubeReserved - systemReserved - eviction reservation
The exact values are node-specific and should be read from the API. CPU, memory, and ephemeral storage are exposed as allocatable resources. Memory and disk eviction thresholds contribute recovery headroom; CPU does not use the same threshold mechanism. A scheduler considers requests against allocatable, including requests from already-bound Pods, whereas kubectl top reports sampled usage. A quiet CPU graph cannot make an over-requested node schedulable, and low requests can let placement consume headroom that live usage has not yet consumed.
Kubernetes documents an illustrative 16-CPU, 32-GiB-memory node with 1 CPU and 2 GiB in kubeReserved, 500 millicores and 1 GiB in systemReserved, and a 500-MiB memory eviction threshold. Its Pod allocatable values are 14.5 CPUs and 28.5 GiB of memory. Local-storage allocatable also reflects supported root-filesystem reservation and nodefs eviction settings. This arithmetic explains why an instance size is not the same as safe application capacity.
Reserve the right owners, not a generic percentage
kubeReserved is intended for Kubernetes system daemons such as kubelet and the container runtime. Its sizing tends to rise with Pod density, image churn, runtime work, and the number of node-level agents. It does not reserve resources for system components that run as Pods; those Pods need their own requests and are accounted for as workload.
systemReserved is for operating-system daemons and host overhead such as SSH services, udev, user sessions, and kernel memory not accounted to Pods. Do not make this a guess copied between instance families. Profile representative nodes under startup storms, image pulls, logging bursts, upgrades, and the maximum supported Pod density. Include a recovery margin: a node that becomes unable to run kubelet, the runtime, networking, or monitoring reliably is not actually at full usable capacity.
Resource keys supported by the reservation maps are not identical. The current KubeletConfiguration v1beta1 reference documents CPU and memory for systemReserved, while kubeReserved also supports local storage on the root filesystem. For disk headroom, separately review the kubelet’s nodefs and imagefs eviction signals; do not copy an ephemeral-storage value into systemReserved without confirming that the target release accepts it.
An example Linux kubelet configuration follows. The figures are placeholders, not universal production defaults. This sample explicitly retains the other Linux hard-eviction defaults while customizing memory headroom; inspect the kubelet config reference and your Kubernetes release before distributing it.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
kubeReserved:
cpu: "500m"
memory: "1Gi"
ephemeral-storage: "1Gi"
systemReserved:
cpu: "300m"
memory: "1Gi"
evictionHard:
memory.available: "500Mi"
nodefs.available: "10%"
nodefs.inodesFree: "5%"
imagefs.available: "15%"
imagefs.inodesFree: "5%"
mergeDefaultEvictionSettings: true
enforceNodeAllocatable:
- pods
The sample enforces the aggregate Pod allocatable boundary; it does not claim the OS daemons are cgroup-limited to the two reservation maps. To enforce kubeReserved or systemReserved on the host processes, set the corresponding kubeReservedCgroup or systemReservedCgroup, add kube-reserved or system-reserved to enforceNodeAllocatable, and ensure those cgroups already exist with the intended process tree. Kubelet does not create them for you and can fail to start for an invalid path. Under systemd, names must follow the slice naming rules.
Understand the enforcement boundary
With pods enforcement, kubelet can evict Pods when aggregate Pod usage exceeds allocatable. This is a node recovery response, not a synchronous per-container quota. The scheduler’s request accounting still does not physically reserve each request as a dedicated block. If node daemons exceed their reservation and their cgroups are not enforcing it, they may consume resources the allocatable calculation assumed would remain available. Memory exhaustion can then lead to node-level pressure or system OOM before the capacity plan behaves as operators expected.
CPU, memory, and storage also have different mechanics. CPU requests influence placement and CPU shares; CPU limits, when set, are enforced separately. Memory cannot be throttled like CPU, and memory pressure can require reclaim or eviction. Local storage is measured through kubelet-supported filesystem layouts and has its own pressure and accounting limits. Do not treat a single aggregate reserved percentage as proof that each resource is isolated. Verify each resource, each filesystem, the cgroup driver, and the exact node image.
Enforcing systemReserved deserves particular caution. An undersized OS cgroup can starve or OOM critical host services, making the node harder to recover. Kubernetes recommends beginning with Pod allocatable enforcement, adding measured compressible-resource enforcement where appropriate, and only then considering non-compressible host reservations with a tested recovery route. CPU is compressible; memory and disk are not equivalent controls.
Eviction headroom is part of capacity planning
Hard eviction thresholds have no grace period once crossed. Soft thresholds require a configured grace period and can use a capped Pod termination grace period. The kubelet evaluates threshold signals periodically, so a configured threshold is not a guarantee that an application receives a particular amount of warning or termination time. Thresholds protect node operability; they do not replace realistic Pod requests or workload-level limits.
Memory.available is a kubelet signal derived from node statistics and cgroup accounting, not simply the output of free -m. Filesystem signals such as nodefs.available, imagefs.available, and inode availability refer to kubelet-observed filesystems. Several identifiers can map to one mount. If you override any hard threshold, understand mergeDefaultEvictionSettings: without merging or explicitly specifying the rest, changed settings can cause the other defaults to become zero rather than silently inheriting the values you expected.
An allocatable target should leave space for the eviction threshold and any expected workload burst. If reservations are too small, host daemons can destabilize the node. If too large, Pods remain Pending while idle resources appear on dashboards. Measure both outcomes instead of optimizing only for bin-packing density.
Audit a node pool with API evidence
Start by comparing capacity and allocatable on every node class, not only a sample from one pool:
NODE=worker-01
kubectl get node "$NODE" -o jsonpath='{.status.capacity.cpu}{" CPU / "}{.status.allocatable.cpu}{" allocatable CPU; "}{.status.capacity.memory}{" memory / "}{.status.allocatable.memory}{" allocatable memory\n"}'
kubectl describe node "$NODE"
kubectl get events --all-namespaces --field-selector involvedObject.kind=Node,involvedObject.name="$NODE"
The node description includes allocated requests and limits, which helps explain why a Pod cannot schedule even if the host looks quiet. Compare that table with Pod specifications after admission, because defaulting and injected sidecars can alter the scheduler’s request total. Inspect node conditions and eviction events alongside allocatable; repeated MemoryPressure, DiskPressure, or evictions indicate that the nominal reserve is not protecting the workload envelope.
For each pool, record the node image and Kubernetes version, cgroup mode and driver, kubelet configuration source, reservations, eviction thresholds, observed daemon peaks, maximum Pod density, and actual allocatable. Managed services may generate kubelet configuration from provider settings, so edit the supported source of truth rather than a file that will be overwritten. If node pools intentionally differ, document the reason and test workloads against each pool’s true allocatable value.
Acceptance-test capacity, not just configuration
Use a disposable node pool to test the contract. First, create Pods whose aggregate requests approach the intended allocatable limit and verify that additional requests remain Pending with a clear scheduling reason. Then test controlled CPU and memory load while observing host daemons, cgroups, kubelet events, node conditions, and application latency. Verify that the configured eviction threshold triggers before the node loses its ability to report health or run recovery agents. Never induce genuine memory or filesystem exhaustion on a production node to prove the configuration.
Also test node replacement and scale-out. A cluster autoscaler can only help if its template advertises compatible allocatable resources and its node shape satisfies placement constraints. If it predicts capacity from instance capacity instead of the final registered Node allocatable, scaling can repeatedly produce nodes that still cannot fit the pending workload. Record the measured difference between instance capacity and usable Pod capacity so application teams can size requests and replicas honestly.
The operational target is not zero unused memory. It is a node that has enough declared Pod capacity to meet density goals while preserving the kubelet, runtime, OS, and eviction margin needed to remain schedulable and recoverable under real load.
Related:
- How the Kubernetes Scheduler Actually Places Workloads
- Kubernetes ResourceQuotas in Production: Admission Budgets and Failure Modes
Sources: