Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux zswap: Operate the Compressed Swap Cache Under Memory Pressure

Diagnose Linux zswap pool growth, writeback, compressor behavior, and cgroup limits without confusing its compressed cache with a zram swap device.

Linux zswap is a compressed RAM cache for pages that the swap subsystem is trying to move out of memory. It can reduce writes to a backing swap device and may make a later swap-in faster when decompression in memory costs less than reading the page from storage. The tradeoff is not free: compression uses CPU, the compressed pool itself consumes RAM, and pages that do not fit or are evicted from the pool can still reach the backing swap area.

Zswap is often confused with zram because both compress data in memory. Their placement in the swap path is different. Zram exposes a compressed block device that can itself be enabled as a swap area. Zswap is an optional cache in front of one or more ordinary swap areas. The distinction matters when interpreting swap accounting, configuring capacity, or reasoning about where a page will be written when the cache is full.

Follow a page through the zswap path

When the virtual-memory subsystem swaps out an eligible page, zswap can attempt to compress it into a dynamically allocated pool in RAM. It associates the compressed object with the swap entry so a later page fault can find and decompress it. The pool is not a reservation of the configured maximum: it grows as entries are stored and shrinks as entries are invalidated or faulted back in.

If a store cannot be accepted, or the pool reaches its configured limit, the system can fall back to the backing swap device according to the active policy. When zswap needs space, it can evict older entries and write them to that device. Consequently, “swap used” does not tell an operator whether the bytes are currently held compressed in RAM or occupy the backing area. Inspect zswap’s counters alongside /proc/swaps, memory pressure, cgroup accounting, CPU use, and device I/O.

The compressed representation is workload-dependent. Zero-filled and other same-value pages may be represented specially; encrypted or already compressed contents often yield little benefit. A benchmark over repeated zeros is not representative of a database, browser, virtual machine, or mixed production workload. Measure both the reduction in swap I/O and the CPU time, tail latency, memory footprint, and reclaim behavior introduced by compression.

Establish what the running kernel actually enabled

Whether zswap starts enabled depends on the kernel build default and boot command line. Runtime configuration is exposed under /sys/module/zswap/parameters/ when the feature is available. Begin with read-only inspection and record the kernel, active swap areas, and current settings:

uname -r
cat /proc/cmdline
cat /proc/swaps
for key in enabled compressor max_pool_percent accept_threshold_percent shrinker_enabled; do
    path="/sys/module/zswap/parameters/$key"
    if [ -r "$path" ]; then
        printf '%s=' "$key"
        cat "$path"
    fi
done

An absent file is not proof that a distribution never supports zswap. The feature may be disabled in the kernel configuration, unavailable in that kernel build, or the expected sysfs interface may not be present. Confirm the target kernel’s configuration and matching documentation before automating a change. Do not assume a setting from one vendor kernel applies unchanged to another release.

The enabled parameter controls whether new swap-out pages are accepted into zswap. Disabling it at runtime does not immediately flush the entries already stored in the compressed pool. Those entries can remain until they are faulted back in or invalidated. The kernel documentation describes swapoff as a way to fault swapped pages back into memory; it is therefore not a harmless cache-clear command. On a pressured host, swapping an area off can demand substantial RAM and fail or worsen an incident. Establish headroom and an alternate swap/recovery plan before changing active swap policy.

Choose pool and compressor settings from evidence

max_pool_percent controls the maximum zswap pool size as a percentage of system memory. Raising it may keep more pages in compressed RAM, but that also allows the pool to compete with application memory and kernel reclaim. Lowering it can increase writes to the backing swap device. Neither direction is universally safer; correlate the pool limit with available memory, workload latency, writeback capacity, and the host’s swap policy.

The compressor is selected by the kernel’s build default or boot-time configuration and can be changed at runtime on supported kernels. A live compressor change does not recompress every existing page. Entries already stored remain associated with their original compressor until they are removed, so more than one pool may temporarily exist. Avoid treating the setting change as an immediate conversion or comparing a short transient window as though all pages used the new algorithm.

For a controlled maintenance test, record current values, change one variable, run the same workload, and keep a rollback plan. The following writes are examples of root-only runtime configuration, not a universal recommendation:

# Review first; do not copy these values into production without a test plan.
cat /sys/module/zswap/parameters/max_pool_percent
cat /sys/module/zswap/parameters/compressor

# Example only: a bounded test value on a kernel exposing this parameter.
echo 20 > /sys/module/zswap/parameters/max_pool_percent

Select a compressor based on representative pages and CPU capacity, not only its name or a published ratio. More compression can consume additional CPU and increase latency; faster compression can leave a larger pool. Include incompressible pages, memory pressure, and simultaneous swap-in workloads in tests. A kernel may expose different algorithms or tunables than another build, so confirm the available choices before writing to the interface.

Read counters as a set, not a success score

When debugfs is enabled and mounted, current kernel builds expose zswap statistics under /sys/kernel/debug/zswap/. Depending on the kernel version, these include total pool size, stored page counts, incompressible pages, pool-limit events, store rejection reasons, decompression failures, and pages written back. Availability and exact files depend on the target kernel configuration and implementation; inspect the directory rather than assuming every counter exists.

if [ -d /sys/kernel/debug/zswap ]; then
    for counter in /sys/kernel/debug/zswap/*; do
        [ -r "$counter" ] || continue
        printf '%s=' "${counter##*/}"
        cat "$counter"
    done
else
    echo "zswap debugfs statistics are unavailable or debugfs is not mounted"
fi

Counters such as pool_limit_hit, reject_alloc_fail, reject_compress_fail, reject_compress_poor, and written_back_pages help distinguish capacity pressure from allocation failure, compression rejection, and backing-device writeback. Their presence is kernel-dependent. Many are cumulative counters, so take timestamped samples and compare deltas rather than interpreting a nonzero value as an active failure. A counter can remain nonzero after the original event has ended.

Correlate these values with /proc/pressure/memory, swap-in and swap-out activity, device latency, CPU utilization, and application request metrics. A growing pool with low writeback may indicate retained compressed pages, not a leak. Rising rejection and writeback counts together with memory stalls can indicate the pool is not absorbing pressure as expected. If a cgroup is involved, host-wide totals can hide the boundary that reached its limit.

Use cgroup v2 controls deliberately

On cgroup v2 systems with the memory controller active, non-root cgroups can expose memory.zswap.current, memory.zswap.max, and memory.zswap.writeback. The current value measures memory consumed by the zswap compression backend for that cgroup. The maximum limits zswap usage at that boundary. These controls complement, rather than replace, memory.max, memory.high, and memory.swap.max.

The writeback control is especially easy to misinterpret. Setting memory.zswap.writeback to 0 disables swapping attempts to the backing device, including writeback and swap caused by zswap store failures, while still permitting pages to be stored in zswap. If pages repeatedly fail to store, reclaim can become inefficient because the same pages may be rejected repeatedly. This differs from setting memory.swap.max to zero, so validate the exact policy needed before changing either knob.

cg=/sys/fs/cgroup/workload
for key in memory.zswap.current memory.zswap.max memory.zswap.writeback \
           memory.current memory.high memory.max memory.swap.current \
           memory.swap.max memory.events memory.pressure; do
    path="$cg/$key"
    if [ -r "$path" ]; then
        printf '\n[%s]\n' "$key"
        cat "$path"
    fi
done

The cgroup example is observational. Writing memory limits or disabling writeback can affect descendants and change reclaim behavior. Check the hierarchy, parent policy, ownership, and current usage first. If memory.zswap.max is lowered below usage, the kernel can take time to reclaim entries; do not expect every counter to drop immediately. Use memory pressure and application latency as guardrails during a controlled rollout.

Distinguish zswap from zram in an incident

Use swapon --show and /proc/swaps to identify active swap areas. A /dev/zramN entry is a zram block device; zswap is not represented as a separate swap device. Check the zswap enabled state and counters separately. This distinction matters when a host has both mechanisms: a page may be compressed by zswap before being written to a swap area, and that area could itself be zram if the host deliberately configured that stack. Such layering is not automatically beneficial. It can add CPU work and confusing accounting, so benchmark the exact combination before operating it.

For a clean diagnosis, capture a baseline before changing settings. Record active swap devices and priorities, kernel command line, cgroup membership and limits, zswap parameters, available debugfs counters, memory PSI, CPU cost, and device I/O. Run a bounded and representative workload, take a second sample, and report which values are current gauges versus cumulative counters. Make one change per experiment and retain an immediate rollback path.

Zswap can be useful when compressed RAM meaningfully reduces costly swap I/O without unacceptable CPU or memory overhead. It cannot create physical memory, guarantee a compression ratio, or eliminate the need for a backing swap policy. Operate it as one measured part of reclaim, respect cgroup boundaries, and never use swapoff or runtime disabling as an unplanned cache flush.

Related:

Sources:

Comments