Skip to content
SRE & DevOpsDeep Dive Published Updated 11 min readViews unavailable

Kubernetes Linux Node Swap: Kubelet Policy, Limits, and Pressure

Operate Linux swap on Kubernetes nodes safely: configure NoSwap or LimitedSwap, isolate scheduling, inspect runtime limits, and plan for memory pressure.

Linux swap can give a Kubernetes node a reclaimable backing store for cold memory pages, but it is not extra RAM and Kubernetes does not treat it as schedulable memory. Whether it helps depends on the workload’s access pattern, the swap device’s latency and I/O budget, kubelet policy, container-runtime support, and the node’s eviction configuration. A node that remains responsive during a short memory spike can still suffer severe tail-latency regressions or sustained paging under load.

This guide focuses on current upstream Linux node behavior: what failSwapOn controls, how NoSwap differs from LimitedSwap, what the scheduler cannot see, how to observe actual use, and how to roll the change out safely. Kubernetes distributions and releases can impose different constraints; verify their supported kubelet configuration and runtime before applying any upstream example.

Start with the policy boundary

Swap must first be provisioned and enabled by the node operating system. Separately, the kubelet decides whether it may start when it detects swap and what swap settings it asks the Container Runtime Interface (CRI) runtime to apply to Kubernetes workloads. These are distinct controls:

Kubelet setting Effect What it does not mean
failSwapOn: true (default) Kubelet refuses to start if swap is detected It does not disable swap in Linux
failSwapOn: false Kubelet tolerates a swap-enabled node It does not grant Pods swap access
memorySwap.swapBehavior: NoSwap (default) Kubernetes-managed workloads receive no swap Host processes and system services may still use swap
memorySwap.swapBehavior: LimitedSwap Eligible Kubernetes workloads can use a kubelet/runtime-enforced amount It does not make swap capacity visible to the scheduler

With failSwapOn: false but the default NoSwap behavior, swap can still help processes outside Kubernetes, including system services or kubelet, during a host-level memory spike. That is not a workaround for setting limits incorrectly: if the workload itself needs swap, LimitedSwap and compatible runtime behavior must be configured deliberately. Do not assume that a healthy kubelet proves that containers have the expected swap limits.

What LimitedSwap allows

Upstream currently describes LimitedSwap as allowing swap for Burstable Pods subject to automatic per-container limits. BestEffort and Guaranteed Pods are excluded; high-priority Pods are also excluded so their memory remains resident. A Burstable container whose memory request equals its limit can opt out of swap. Check the documentation for the cluster’s Kubernetes minor release because eligibility rules and reporting capabilities can change.

The documented limit is proportional to a container’s memory request:

container swap allowance =
  (container memory request / node physical memory)
  * swap capacity available to Pods

For example, if a node has 64 GiB of physical memory and 32 GiB of swap is available to Pods, a Burstable container requesting 2 GiB receives a proportional allowance of 1 GiB. This is an example of the documented calculation, not a recommendation to provision those capacities. Some node swap may be reserved for system use, and not all workloads qualify; therefore total configured swap is not the same as aggregate Pod-usable swap.

The kubelet passes configuration through CRI. The runtime is responsible for applying it to the container cgroup, for example through memory.swap.max on cgroup v2. Verify the actual kernel, cgroup mode, kubelet release, CRI implementation, and runtime version as a tested combination. A YAML fragment copied into a managed-service node bootstrap may be ignored, rejected, or superseded by provider configuration.

Decide whether swap belongs on this node pool

Swap is most defensible for workloads with a large memory footprint but a significant cold or infrequently accessed portion, where latency variation is acceptable and eviction during a transient spike is more damaging than slower access to cold pages. It is a poor fit for latency-critical hot paths, tightly bounded real-time tasks, and systems where predictable memory-resident behavior is part of the service objective. Measure workload behavior instead of enabling swap as a generic cure for OOMKilled containers.

The principal trade-offs are operational, not just capacity-related:

  • Swap-in and swap-out are storage I/O. Under IOPS throttling or on slow media, paging can add large latency and contend with kubelet, the container runtime, logs, and application data.
  • A memory request remains the scheduler’s accounting signal; Pod specifications do not request swap. Current Kubernetes scheduling does not account for swap usage, so a node may be overpacked from the perspective of backing-store I/O even when requested RAM fits.
  • During sustained pressure, paging can become thrashing: the node spends time moving pages rather than completing useful work. It can hide an undersized capacity plan and delay the symptom without removing the cause.
  • Host-wide swap can affect neighboring Pods. A container-level allowance limits an individual container but does not guarantee isolated storage bandwidth or a responsive node.

The Kubernetes project recommends keeping control-plane nodes free of swap. For worker nodes, use a separate, explicitly identified pool when only a subset of workloads is suitable. A taint and a matching label can keep ordinary workloads away while permitting opted-in workloads to target the pool:

kubectl label node worker-swap-01 node.example.com/swap=enabled
kubectl taint node worker-swap-01 node.example.com/swap=enabled:NoSchedule

Only workloads that intentionally tolerate the taint and select the label should land there. Replace the example domain with an organization-controlled label prefix, and manage the label and taint through the node-pool source of truth rather than a one-off command. A taint is not a substitute for capacity planning or workload-level memory requests.

Provision swap as an operating-system change

Provisioning is distribution-specific. The upstream kubeadm tutorial demonstrates swap files, permissions, encryption, activation, and kubelet restart, but managed Kubernetes providers may require a launch-template, machine configuration, or supported node image instead. Do not edit an active node’s /etc/fstab, systemd unit, or kubelet files from a one-off shell session and assume the change will survive replacement.

If the security policy permits swap, encrypt it. Swap can contain pages originating from application memory, so encryption at the OS or storage layer is important, especially for memory-backed volumes and sensitive data. Linux supports the noswap tmpfs mount option upstream from kernel 6.3, and distributions may backport it. Kubernetes uses this protection for supported memory-backed volumes, but an older or unsupported kernel/runtime combination can have different behavior. Confirm the kernel’s support and kubelet logs; do not infer that a volume is protected solely because its source is a Secret or emptyDir with medium: Memory.

The swap device shares the node’s I/O budget unless isolated. Prefer a fast, appropriately provisioned device and test contention against the actual VM or storage service limits. If system services degrade when swapping is active, upstream guidance discusses preventing swap for the system slice and prioritizing system-critical I/O. Those are host-level cgroup and systemd decisions: apply the OS vendor’s supported mechanism, verify it after reboot, and avoid hand-editing cgroup controls without accounting for hierarchy and service lifecycle.

Configure and roll out the kubelet behavior

For kubeadm-managed nodes, the upstream example uses a kubelet configuration fragment like this:

failSwapOn: false
memorySwap:
  swapBehavior: LimitedSwap

This is a configuration fragment, not a complete KubeletConfiguration document and not a universal managed-cluster procedure. Kubelet usually must be restarted after swap is provisioned and configuration is changed. Follow the provider’s node-image or configuration-management workflow, and coordinate swap activation so it is available before kubelet starts. Otherwise kubelet may fail during boot because the expected swap device is not yet active.

Use a canary node pool, not a cluster-wide simultaneous restart. First record the current node image, kubelet and runtime versions, cgroup mode, physical memory, swap size and backing device, eviction thresholds, and workload baseline. Cordon and drain a canary using the platform’s supported procedure; configure swap and kubelet; restart and confirm Ready; then run a representative workload and verify its cgroup limits and observed swap usage. Exercise node replacement and reboot before broad rollout. Keep enough unaffected capacity to absorb drained nodes and rollback if the canary exhibits paging storms, failed kubelet startup, or new evictions.

Set memory requests and limits based on measured working sets, not on swap capacity. Review kubelet eviction thresholds with the operating system’s vm.min_free_kbytes and reclaim behavior. Kubernetes upstream notes that eviction can happen too early to permit useful reclaim if configured above the kernel’s aggressive-reclaim threshold, while thresholds set too high risk an OOM kill. Its guidance is to set eviction thresholds slightly below vm.min_free_kbytes; this is a relationship to validate with units and the actual kernel, not a copy-paste numeric value for every distribution. Test it under controlled pressure and verify that kubelet evicts before the node becomes unstable.

Observe capacity, limits, and pressure together

Begin with node status and human-readable metrics. Current upstream kubectl supports swap display on top commands, and node status can report swap capacity:

kubectl get nodes -o custom-columns='NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,SWAP-BYTES:.status.nodeInfo.swap.capacity'
kubectl top nodes --show-swap
kubectl top pods -A --show-swap

The swap capacity field may be absent or unknown when the kubelet cannot determine it or no swap is provisioned. kubectl top requires a compatible client and metrics pipeline; it is not a substitute for inspecting node-level pressure and container limits. For machine-readable collection, upstream documents node_swap_usage_bytes, container_swap_usage_bytes, and container_swap_limit_bytes through kubelet resource metrics, along with /stats/summary. Access to kubelet endpoints should remain authenticated and restricted to trusted monitoring components.

Check the host and workload evidence side by side:

swapon --show
free -h
cat /proc/sys/vm/min_free_kbytes
kubectl describe node worker-swap-01
kubectl get events -A --sort-by=.lastTimestamp

Investigate sustained swap-in/out, memory PSI, disk latency and throttling, node MemoryPressure, kubelet eviction events, container restarts, OOM kills, and service latency. Swap usage alone is not a failure signal: cold pages may remain swapped after a spike. A rising trend combined with reclaim stalls, I/O saturation, or elevated latency is more actionable. Compare per-container usage and allowance with the Pod’s QoS class and priority; if a workload expected to swap reports zero allowance, validate policy eligibility and runtime-applied cgroup settings before changing limits.

Troubleshoot common failure modes

Symptom Likely checks Safer next step
Kubelet will not start after enabling OS swap failSwapOn, whether swap activates before kubelet, node bootstrap overrides, kubelet logs Restore the known-good node configuration; correct the declarative startup order on a canary
Kubelet is healthy but Pods show no swap swapBehavior, Pod QoS and priority, equal request/limit, runtime and cgroup support Inspect effective cgroup limits and the release-specific eligibility rules
Pod swap usage grows while latency spikes Swap device I/O latency, VM IOPS throttling, memory PSI, workload working set Roll back or isolate the workload; increase real capacity or reduce memory demand before widening swap
Evictions persist despite free swap Kubelet available-memory signal, vm.min_free_kbytes, eviction thresholds, node-level versus container-level pressure Reconcile reclaim and eviction thresholds in staging; swap does not change the scheduler’s request accounting
Sensitive data may have reached swap Encryption state, key lifecycle, tmpfs noswap support and kubelet warnings Treat it as a host data-handling incident and follow the organization’s response process; do not assume deleting a Pod erases disk pages

Avoid using swapoff -a as an incident reflex on a pressured node: forcing swapped pages back into RAM can itself trigger memory pressure and OOM behavior. If swap must be disabled, first drain or otherwise protect the node, confirm sufficient physical headroom, then follow the OS and provider procedure. Preserve metrics and kernel/kubelet logs before the node is recycled.

Production acceptance checklist

  • The target workloads have a measured use case for swap and a latency/error budget that tolerates its worst-case I/O behavior.
  • The node OS, kernel, kubelet, CRI runtime, cgroup mode, and provider-supported configuration have been validated together.
  • failSwapOn and memorySwap.swapBehavior are understood as separate controls; no Pod is assumed to use swap merely because kubelet starts.
  • Swap is encrypted, memory-backed volume protection is verified, and control-plane nodes follow upstream/provider guidance.
  • Swap-enabled workers are identified and isolated deliberately; scheduling and memory requests still reflect RAM rather than swap capacity.
  • Canary tests cover startup order, Pod eligibility, cgroup limits, reboot/replacement, eviction behavior, and rollback.
  • Monitoring joins swap use with memory pressure, I/O latency, eviction, OOM, and application SLO signals.

Kubernetes swap is a node-level resilience and memory-management option, not a replacement for accurate requests, limits, capacity, or load testing. A production rollout succeeds only when the kubelet policy, kernel reclaim, CRI enforcement, scheduler blind spots, storage performance, and incident response are all understood as one system.

Related:

Sources:

Comments