Linux Memory Hotplug: Online, Offline, and Capacity-Change Failure Modes
Operate Linux memory hotplug safely by inspecting block granularity, online policy, NUMA placement, movable-page limits, and rollback conditions first.
Memory hotplug lets supported systems add or remove physical memory while Linux is running. In virtual machines, a hypervisor may expose capacity dynamically; in hardware platforms, firmware and the memory controller must also support the operation. Linux then has to create metadata, present memory blocks to the page allocator, and, for removal, migrate movable contents away. A successful request at one layer does not prove every layer completed: capacity changes need explicit observation at the device, kernel, zone, NUMA, and workload levels.
Support is architecture- and configuration-dependent. The generic kernel documentation describes selected 64-bit architectures and SPARSEMEM; the fact that a sysfs path exists is not a guarantee that a platform can physically add or remove memory. Coordinate the change with the hypervisor or hardware control plane, the guest kernel configuration, memory policy, and applications that depend on a specific NUMA shape.
Two phases to add or remove memory
Adding memory has a physical or platform phase and a logical Linux phase. The kernel establishes metadata such as memory maps and page tables and creates sysfs memory-block objects. Those blocks must then be onlined before they become available to the page allocator and ordinary memory statistics. Some configurations automatically online new blocks, but operators should inspect actual state rather than infer it from successful hypervisor API completion.
Removal reverses the order. Linux first offlines a memory block, migrating movable pages and removing its free pages from allocation. Only after that succeeds can the block be removed from the kernel’s memory map. The platform then retracts the capacity. This ordering prevents memory from disappearing while Linux still uses it.
Block granularity depends on architecture and kernel configuration. Read the exported block-size information and inspect the discovered block IDs; do not hard-code memory0 or a familiar section size. A DIMM or hypervisor range may map to several blocks, and a requested capacity decrease may not align with the kernel’s block boundary.
# Read-only discovery: capture these values before planning a change.
cat /sys/devices/system/memory/block_size_bytes
for state in /sys/devices/system/memory/memory*/state; do
printf '%s: ' "$state"
cat "$state"
done
The glob can be empty on a system without the expected sysfs interface. Run this discovery from a privileged diagnostic shell only as needed; it is read-only. Record the target block list and NUMA association before initiating an online/offline operation.
Onlining and zone policy
When automatic online is disabled, user space can write an accepted state such as online to a memory block’s state file. The kernel chooses an online zone based on policy. Depending on the platform, explicit choices such as online_movable may be available. The zone choice affects reclaim, allocation, and future removability; it is not a cosmetic setting.
ZONE_MOVABLE can improve the chance of later offlining because memory placed in a movable zone is intended for pages that can migrate. It does not mean every page on the machine becomes movable. Kernel structures, pinned pages, long-term DMA mappings, huge pages, or driver allocations may prevent migration. Determine the workload and device constraints before choosing an online policy.
After onlining, verify the block’s state, its zone, NUMA node, and changes in memory accounting such as MemTotal. In a virtualized system, compare guest observations with the host’s assigned capacity and balloon/hotplug counters. free output alone is insufficient because it reports aggregate capacity and cannot show whether the intended NUMA node or block was added.
Offlining and why it fails
Linux must migrate movable pages off the target block and ensure no remaining references prevent removal. Most kernel allocations, including page tables, are not generally movable, and some page types have constraints that prevent migration. Offlining may therefore fail with a busy or invalid-argument error even when the block appears mostly free. The failure is a safeguard; do not retry blindly or force the platform to retract memory anyway.
Before a planned change, collect a baseline of per-node memory, zone information, active workloads, device assignments, and kernel logs. Drain workloads where practical. Check for pinned memory, huge pages, device passthrough, memory-tiering or CXL policy, and guest-agent constraints. The exact diagnostics vary by kernel and platform; use the kernel hotplug documentation and vendor control-plane runbook rather than treating one sysfs error as a universal recipe.
The operation may take a long time as pages migrate. Cancellation and timeout behavior should be understood before a production window. If the state remains going-offline, determine whether migration is progressing before initiating another capacity change. Preserve host-side telemetry and the guest kernel log around the attempt.
Memory accounting can differ between control planes. A cloud console may display provisioned memory, the guest firmware may expose a changed address range, and Linux may still show the block offline. Record a timestamped state transition at every layer rather than collapsing those reports into one “memory added” flag. In NUMA deployments, confirm that CPU and memory locality still match the workload’s placement policy. An aggregate increase can mask an unintended node imbalance that reduces performance or exhausts one node first.
The capacity budget should include migration headroom. Offlining requires room to move pages elsewhere, so a request to remove too much memory at once can fail even if the final system would have adequate free memory after a successful migration. Stage large reductions into supported blocks and re-evaluate after each completed phase. This is a planning technique, not a universal workaround: the kernel may still reject a block because it contains pinned or otherwise non-migratable pages.
# Observe state and logs. Do not issue "offline" until the change is approved.
cat /sys/devices/system/memory/memoryXXX/state
journalctl -k --since "15 minutes ago" | tail -200
Replace the illustrative block name with a discovered block ID. The kernel log can contain unrelated messages; correlate timestamps and memory-block identifiers instead of assuming the last line is causal. A site-specific automation should fail closed if the target ID is missing, has changed state, or is outside the approved set.
Operational sequence and rollback
An operational runbook should coordinate four actors: hardware or hypervisor control plane, guest kernel, orchestration agent, and application workload. First confirm support and reserve a maintenance window. Discover granularity and state. Drain or migrate workloads and validate remaining capacity on every NUMA node. Request the change from the platform control plane. Online or offline the documented blocks in the prescribed order. Then verify both layers before returning the node to service.
Rollback differs for add and remove. If newly added blocks fail to online, the extra capacity may remain present but unusable; decide whether to retry online with an approved zone policy or reverse the platform add. If offlining fails, keep the physical memory allocated and restore workloads; do not remove it from the hypervisor. If Linux offlines successfully but host removal fails, determine whether the block can be safely returned online according to the platform’s contract and kernel state.
Do not encode a generic script that writes online or offline to every discovered block. The script must understand capacity intent, block state, ordering, NUMA placement, platform API result, and rollback. Repeating writes after partial success can produce a topology different from the planned one.
Monitoring and verification
Compare per-node and per-zone counters before and after the change. Confirm new capacity is usable by the intended workload, not merely counted by the kernel. For removal, check that the block is offline, that it disappeared only after the platform phase, and that system memory accounting matches the expected remaining capacity. Review application memory pressure and allocation failures during the observation window.
Test the runbook in a representative VM or lab platform first. Exercise a successful add, automatic online, manual online, a block that cannot offline, a timeout, partial multi-block success, control-plane failure after guest preparation, and guest reboot during orchestration. Define idempotency and recovery behavior for each stage.
The test must also prove alerting and service behavior. Verify that automation detects a partially completed change, that operators can identify the exact block and node involved, and that capacity dashboards do not double-count memory still offline. Rehearse aborting before the platform retracts memory and recovering from a control-plane timeout whose guest operation succeeded but whose API response was lost. Reconcile actual guest state before retrying such a request.
Memory hotplug is a coordinated topology change, not a single sysfs write. Reliable operation starts with platform support and discovered block geometry, preserves the kernel’s online/offline phases, treats migration failure as a hard stop, and verifies NUMA and workload behavior before declaring completion.
Related:
- Inside the OOM Killer: How Linux Decides What to Kill When Memory Runs Out
- Linux process_madvise(): Advising the Kernel About Another Process’s Memory
Sources: