Linux Memory Failure Handling: Hardware Poison and Process Recovery
Trace Linux memory-failure handling from ECC or machine-check reports through poisoned pages, SIGBUS delivery, containment, and evidence preservation.
When hardware reports that a physical memory location may contain corrupted data, Linux cannot treat the affected page as ordinary RAM. The kernel’s memory-failure machinery attempts to identify users of the page, isolate or unmap it where possible, and report an error to affected processes or higher-level subsystems. The outcome depends on whether the error was corrected, whether the page was consumed, what kind of page it is, and which architecture and recovery features are available.
The visible event may be a machine-check record, a RAS trace, a kernel log, a process signal, or a later I/O error. A missing panic does not prove that memory was healthy; a panic does not prove the physical DIMM alone is at fault. Preserve the complete evidence and distinguish detection, containment, and data integrity.
What a poisoned page means
Hardware ECC can detect and sometimes correct a memory error. If an uncorrectable error is associated with a page, the kernel may mark or treat that page as poisoned so future consumers do not unknowingly use its contents. The poison indication is a containment signal, not a repair of the underlying memory or a guarantee that all copies of the data were correct.
Linux can encounter memory errors through architecture-specific machine checks or other RAS reporting paths. Some events are consumed synchronously by the current instruction; others are reported asynchronously after the memory access. This distinction affects whether the kernel can identify a process and deliver a synchronous error. The exact architecture behavior and recovery options must be checked for the deployed CPU and kernel.
File-backed cache pages, anonymous memory, slab objects, huge pages, and device-mapped memory have different recovery constraints. The kernel may be able to invalidate a page-cache entry and cause a later reread, but an anonymous page with no recoverable copy may be lost. A corrupted kernel object can affect the entire system. Do not assume every memory error results in a clean per-process SIGBUS.
Interpret process-facing signals carefully
When a recoverable user mapping is affected, Linux can deliver SIGBUS with machine-check-specific information such as BUS_MCEERR_AR or BUS_MCEERR_AO on supported systems. These codes indicate aspects of the memory error and delivery model; they do not identify a replaceable DIMM or guarantee that the process’s data structure is safe to continue using.
Applications that need resilience must define how to handle the signal and what state can be discarded or reconstructed. Returning from a signal handler and continuing to read the same corrupted page is not a recovery strategy. Databases, virtual machines, and scientific applications may need to abort a transaction, terminate a worker, restore from a validated copy, or fail the service.
Signal address information and page offsets may be limited or architecture-dependent. Collect process mappings, core-dump policy, kernel logs, and RAS events while protecting sensitive memory contents. A core dump can preserve corrupted or confidential data and should follow the organization’s retention policy.
Diagnosis and containment workflow
Begin with the boot ID, kernel version, CPU model, firmware versions, machine-check logs, RAS daemon state, EDAC counters, and the affected process. Record whether the event was corrected or uncorrected, whether a page was isolated, which process received a signal, and whether the machine continued operating. Do not reboot before preserving volatile logs unless the system is unstable or vendor guidance requires immediate shutdown.
journalctl -k -b --no-pager
journalctl -b -u rasdaemon --no-pager
find /sys/devices/system/edac -maxdepth 5 -type f -print
Service names and sysfs attributes vary. The commands are examples for evidence collection; an absent rasdaemon service or EDAC path means that particular reporting path may not be configured. Correlate with BMC event logs, firmware memory inventory, thermal and power events, workload traces, and the platform’s documented physical slot map.
Do not clear RAS counters while collecting a time series. Record before-and-after values over a defined period, but avoid exposing raw physical addresses or customer data in shared reports. If a page is reported as poisoned, preserve the report and follow the platform vendor’s service procedure. Moving memory modules or injecting errors requires a maintenance window and controlled handling.
Huge pages, filesystems, and other hard cases
Huge pages aggregate many base pages and can have compound-page state, so an error may affect a larger mapping or require a subsystem-specific recovery path. A memory error in a file-backed page may be recoverable by invalidation and reread only if the storage source remains healthy and the data has not already been corrupted on disk. Filesystem checksums, redundancy, and application checksums can provide evidence or recovery, but none should be assumed.
Pages pinned for DMA or mapped into a device can complicate isolation. The device may still access the affected memory until its queue or DMA is stopped. A driver or subsystem needs to coordinate invalidation and report whether an operation may have used bad data. Containers do not isolate hardware memory faults; a page used by one process can trigger kernel-wide consequences.
Virtual machines add another layer. The host may detect a physical page error and notify or poison a guest mapping, depending on virtualization support and policy. A guest can also report a virtual memory machine check that the host must map back to backing memory. Preserve host and guest logs with synchronized timestamps and machine identities; a guest-only signal is not enough to localize the DIMM.
Recovery boundaries and operational response
If an error is isolated and the process receives a signal, the kernel may continue, but the service still needs a defined failure policy. Restarting the process can be safe only if its persistent state is validated and external operations are idempotent. A single hardware memory error can affect in-flight buffers, filesystem writeback, or another process through shared pages.
If kernel memory or a page table is involved, panic may be safer than continuing with uncertain state. Kdump or pstore can preserve evidence when configured, but crash capture itself requires reserved memory and a working capture path. Verify those facilities before the incident; do not configure a new crash kernel on a failing host as a live experiment.
Escalate repeated uncorrected events, rising corrected counts, or machine-checks to the hardware vendor. Provide model, serial or service tag through a protected channel, memory map, firmware, kernel, event rate, workload, and exact event logs. Replace or offline components only using the platform’s supported procedure.
Validate the response path safely
Error injection is dangerous and driver-specific. Some kernels expose test hooks for memory failure, but an injected event can panic the host or corrupt data. Use a disposable lab machine, local console, test data, and a recovery image. Read the exact kernel documentation and configuration before enabling any injection path.
Prefer validating alert and incident logic with synthetic telemetry or vendor simulator data. Test that an EDAC or RAS event is retained, that the alert includes severity and machine identity, and that runbooks tell operators how to preserve evidence. Do not test the alert by injecting an ECC fault into a production memory controller.
Acceptance criteria define the reporting path for corrected and uncorrected errors, process signal handling, page isolation evidence, crash capture availability, vendor escalation, and data-integrity checks after recovery. Include which failure classes are expected to continue and which require host reset.
Memory-failure handling is a containment protocol, not an automatic repair. A sound response correlates the hardware report with page and process state, stops consumers from reusing poisoned data, preserves evidence, and applies a recovery policy that is safe for the service.
Related:
- Linux Kdump: Reserved Memory, Capture Kernels, and vmcore Validation
- Inside the OOM Killer: How Linux Decides What to Kill When Memory Runs Out
Sources: