Kubernetes Linux Node Swap: Kubelet Policy, Limits, and Pressure
Operate Linux swap on Kubernetes nodes safely: configure NoSwap or LimitedSwap, isolate scheduling, inspect runtime limits, and plan for memory pressure.
Linux swap can give a Kubernetes node a reclaimable backing store for cold memory pages, but it is not extra RAM and Kubernetes does not treat it as schedulable memory. Whether it helps depends on the workload’s access pattern, the swap device’s latency and I/O budget, kubelet policy, container-runtime support, and the node’s eviction configuration. A node that remains responsive during a short memory spike can still suffer severe tail-latency regressions or sustained paging under load.
This guide focuses on current upstream Linux node behavior: what failSwapOn controls, how NoSwap differs from LimitedSwap, what the scheduler cannot see, how to observe actual use, and how to roll the change out safely. Kubernetes distributions and releases can impose different constraints; verify their supported kubelet configuration and runtime before applying any upstream example.
Start with the policy boundary
Swap must first be provisioned and enabled by the node operating system. Separately, the kubelet decides whether it may start when it detects swap and what swap settings it asks the Container Runtime Interface (CRI) runtime to apply to Kubernetes workloads. These are distinct controls:
| Kubelet setting | Effect | What it does not mean |
|---|---|---|
failSwapOn: true (default) |
Kubelet refuses to start if swap is detected | It does not disable swap in Linux |
failSwapOn: false |
Kubelet tolerates a swap-enabled node | It does not grant Pods swap access |
memorySwap.swapBehavior: NoSwap (default) |
Kubernetes-managed workloads receive no swap | Host processes and system services may still use swap |
memorySwap.swapBehavior: LimitedSwap |
Eligible Kubernetes workloads can use a kubelet/runtime-enforced amount | It does not make swap capacity visible to the scheduler |
With failSwapOn: false but the default NoSwap behavior, swap can still help processes outside Kubernetes, including system services or kubelet, during a host-level memory spike. That is not a workaround for setting limits incorrectly: if the workload itself needs swap, LimitedSwap and compatible runtime behavior must be configured deliberately. Do not assume that a healthy kubelet proves that containers have the expected swap limits.
What LimitedSwap allows
Upstream currently describes LimitedSwap as allowing swap for Burstable Pods subject to automatic per-container limits. BestEffort and Guaranteed Pods are excluded; high-priority Pods are also excluded so their memory remains resident. A Burstable container whose memory request equals its limit can opt out of swap. Check the documentation for the cluster’s Kubernetes minor release because eligibility rules and reporting capabilities can change.
The documented limit is proportional to a container’s memory request:
container swap allowance =
(container memory request / node physical memory)
* swap capacity available to Pods
For example, if a node has 64 GiB of physical memory and 32 GiB of swap is available to Pods, a Burstable container requesting 2 GiB receives a proportional allowance of 1 GiB. This is an example of the documented calculation, not a recommendation to provision those capacities. Some node swap may be reserved for system use, and not all workloads qualify; therefore total configured swap is not the same as aggregate Pod-usable swap.
The kubelet passes configuration through CRI. The runtime is responsible for applying it to the container cgroup, for example through memory.swap.max on cgroup v2. Verify the actual kernel, cgroup mode, kubelet release, CRI implementation, and runtime version as a tested combination. A YAML fragment copied into a managed-service node bootstrap may be ignored, rejected, or superseded by provider configuration.
Decide whether swap belongs on this node pool
Swap is most defensible for workloads with a large memory footprint but a significant cold or infrequently accessed portion, where latency variation is acceptable and eviction during a transient spike is more damaging than slower access to cold pages. It is a poor fit for latency-critical hot paths, tightly bounded real-time tasks, and systems where predictable memory-resident behavior is part of the service objective. Measure workload behavior instead of enabling swap as a generic cure for OOMKilled containers.
The principal trade-offs are operational, not just capacity-related:
- Swap-in and swap-out are storage I/O. Under IOPS throttling or on slow media, paging can add large latency and contend with kubelet, the container runtime, logs, and application data.
- A memory request remains the scheduler’s accounting signal; Pod specifications do not request swap. Current Kubernetes scheduling does not account for swap usage, so a node may be overpacked from the perspective of backing-store I/O even when requested RAM fits.
- During sustained pressure, paging can become thrashing: the node spends time moving pages rather than completing useful work. It can hide an undersized capacity plan and delay the symptom without removing the cause.
- Host-wide swap can affect neighboring Pods. A container-level allowance limits an individual container but does not guarantee isolated storage bandwidth or a responsive node.
The Kubernetes project recommends keeping control-plane nodes free of swap. For worker nodes, use a separate, explicitly identified pool when only a subset of workloads is suitable. A taint and a matching label can keep ordinary workloads away while permitting opted-in workloads to target the pool:
kubectl label node worker-swap-01 node.example.com/swap=enabled
kubectl taint node worker-swap-01 node.example.com/swap=enabled:NoSchedule
Only workloads that intentionally tolerate the taint and select the label should land there. Replace the example domain with an organization-controlled label prefix, and manage the label and taint through the node-pool source of truth rather than a one-off command. A taint is not a substitute for capacity planning or workload-level memory requests.
Provision swap as an operating-system change
Provisioning is distribution-specific. The upstream kubeadm tutorial demonstrates swap files, permissions, encryption, activation, and kubelet restart, but managed Kubernetes providers may require a launch-template, machine configuration, or supported node image instead. Do not edit an active node’s /etc/fstab, systemd unit, or kubelet files from a one-off shell session and assume the change will survive replacement.
If the security policy permits swap, encrypt it. Swap can contain pages originating from application memory, so encryption at the OS or storage layer is important, especially for memory-backed volumes and sensitive data. Linux supports the noswap tmpfs mount option upstream from kernel 6.3, and distributions may backport it. Kubernetes uses this protection for supported memory-backed volumes, but an older or unsupported kernel/runtime combination can have different behavior. Confirm the kernel’s support and kubelet logs; do not infer that a volume is protected solely because its source is a Secret or emptyDir with medium: Memory.
The swap device shares the node’s I/O budget unless isolated. Prefer a fast, appropriately provisioned device and test contention against the actual VM or storage service limits. If system services degrade when swapping is active, upstream guidance discusses preventing swap for the system slice and prioritizing system-critical I/O. Those are host-level cgroup and systemd decisions: apply the OS vendor’s supported mechanism, verify it after reboot, and avoid hand-editing cgroup controls without accounting for hierarchy and service lifecycle.
Configure and roll out the kubelet behavior
For kubeadm-managed nodes, the upstream example uses a kubelet configuration fragment like this:
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
This is a configuration fragment, not a complete KubeletConfiguration document and not a universal managed-cluster procedure. Kubelet usually must be restarted after swap is provisioned and configuration is changed. Follow the provider’s node-image or configuration-management workflow, and coordinate swap activation so it is available before kubelet starts. Otherwise kubelet may fail during boot because the expected swap device is not yet active.
Use a canary node pool, not a cluster-wide simultaneous restart. First record the current node image, kubelet and runtime versions, cgroup mode, physical memory, swap size and backing device, eviction thresholds, and workload baseline. Cordon and drain a canary using the platform’s supported procedure; configure swap and kubelet; restart and confirm Ready; then run a representative workload and verify its cgroup limits and observed swap usage. Exercise node replacement and reboot before broad rollout. Keep enough unaffected capacity to absorb drained nodes and rollback if the canary exhibits paging storms, failed kubelet startup, or new evictions.
Set memory requests and limits based on measured working sets, not on swap capacity. Review kubelet eviction thresholds with the operating system’s vm.min_free_kbytes and reclaim behavior. Kubernetes upstream notes that eviction can happen too early to permit useful reclaim if configured above the kernel’s aggressive-reclaim threshold, while thresholds set too high risk an OOM kill. Its guidance is to set eviction thresholds slightly below vm.min_free_kbytes; this is a relationship to validate with units and the actual kernel, not a copy-paste numeric value for every distribution. Test it under controlled pressure and verify that kubelet evicts before the node becomes unstable.
Observe capacity, limits, and pressure together
Begin with node status and human-readable metrics. Current upstream kubectl supports swap display on top commands, and node status can report swap capacity:
kubectl get nodes -o custom-columns='NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,SWAP-BYTES:.status.nodeInfo.swap.capacity'
kubectl top nodes --show-swap
kubectl top pods -A --show-swap
The swap capacity field may be absent or unknown when the kubelet cannot determine it or no swap is provisioned. kubectl top requires a compatible client and metrics pipeline; it is not a substitute for inspecting node-level pressure and container limits. For machine-readable collection, upstream documents node_swap_usage_bytes, container_swap_usage_bytes, and container_swap_limit_bytes through kubelet resource metrics, along with /stats/summary. Access to kubelet endpoints should remain authenticated and restricted to trusted monitoring components.
Check the host and workload evidence side by side:
swapon --show
free -h
cat /proc/sys/vm/min_free_kbytes
kubectl describe node worker-swap-01
kubectl get events -A --sort-by=.lastTimestamp
Investigate sustained swap-in/out, memory PSI, disk latency and throttling, node MemoryPressure, kubelet eviction events, container restarts, OOM kills, and service latency. Swap usage alone is not a failure signal: cold pages may remain swapped after a spike. A rising trend combined with reclaim stalls, I/O saturation, or elevated latency is more actionable. Compare per-container usage and allowance with the Pod’s QoS class and priority; if a workload expected to swap reports zero allowance, validate policy eligibility and runtime-applied cgroup settings before changing limits.
Troubleshoot common failure modes
| Symptom | Likely checks | Safer next step |
|---|---|---|
| Kubelet will not start after enabling OS swap | failSwapOn, whether swap activates before kubelet, node bootstrap overrides, kubelet logs |
Restore the known-good node configuration; correct the declarative startup order on a canary |
| Kubelet is healthy but Pods show no swap | swapBehavior, Pod QoS and priority, equal request/limit, runtime and cgroup support |
Inspect effective cgroup limits and the release-specific eligibility rules |
| Pod swap usage grows while latency spikes | Swap device I/O latency, VM IOPS throttling, memory PSI, workload working set | Roll back or isolate the workload; increase real capacity or reduce memory demand before widening swap |
| Evictions persist despite free swap | Kubelet available-memory signal, vm.min_free_kbytes, eviction thresholds, node-level versus container-level pressure |
Reconcile reclaim and eviction thresholds in staging; swap does not change the scheduler’s request accounting |
| Sensitive data may have reached swap | Encryption state, key lifecycle, tmpfs noswap support and kubelet warnings |
Treat it as a host data-handling incident and follow the organization’s response process; do not assume deleting a Pod erases disk pages |
Avoid using swapoff -a as an incident reflex on a pressured node: forcing swapped pages back into RAM can itself trigger memory pressure and OOM behavior. If swap must be disabled, first drain or otherwise protect the node, confirm sufficient physical headroom, then follow the OS and provider procedure. Preserve metrics and kernel/kubelet logs before the node is recycled.
Production acceptance checklist
- The target workloads have a measured use case for swap and a latency/error budget that tolerates its worst-case I/O behavior.
- The node OS, kernel, kubelet, CRI runtime, cgroup mode, and provider-supported configuration have been validated together.
failSwapOnandmemorySwap.swapBehaviorare understood as separate controls; no Pod is assumed to use swap merely because kubelet starts.- Swap is encrypted, memory-backed volume protection is verified, and control-plane nodes follow upstream/provider guidance.
- Swap-enabled workers are identified and isolated deliberately; scheduling and memory requests still reflect RAM rather than swap capacity.
- Canary tests cover startup order, Pod eligibility, cgroup limits, reboot/replacement, eviction behavior, and rollback.
- Monitoring joins swap use with memory pressure, I/O latency, eviction, OOM, and application SLO signals.
Kubernetes swap is a node-level resilience and memory-management option, not a replacement for accurate requests, limits, capacity, or load testing. A production rollout succeeds only when the kubelet policy, kernel reclaim, CRI enforcement, scheduler blind spots, storage performance, and incident response are all understood as one system.
Related:
- Kubernetes Node-Pressure Eviction: Thresholds, Ranking, and Recovery
- Fixing OOMKilled Containers with Correct Resource Limits
Sources: