Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux EDAC and RAS: Correlate ECC Errors Before Replacing Memory

Interpret Linux EDAC and RAS evidence by separating corrected and uncorrected errors, mapping controllers to DIMMs, and preserving counters and logs.

Linux RAS facilities collect reliability events from CPUs, memory controllers, interconnects, and other hardware. EDAC focuses substantially on memory-controller error reporting, while architecture-specific machine-check and RAS paths can report additional fault classes. A corrected ECC count is evidence that hardware detected and corrected a condition; it is not a complete root cause or a guarantee that the system remains safe indefinitely.

Memory replacement decisions should not be made from one unexplained counter. The counter may be cumulative since boot, a driver may expose only part of the controller topology, and a logical rank or channel label may not match a chassis slot without platform documentation. Preserve the original report and correlate it with firmware, machine-check, kernel, and service evidence.

Locate the available EDAC and RAS interfaces

EDAC sysfs devices are commonly exposed below /sys/devices/system/edac, but names and attributes vary by driver. Begin with inventory and kernel logs:

find /sys/devices/system/edac -maxdepth 5 -type f -print
journalctl -k -b --no-pager
lsmod

The last command shows loaded modules but does not prove a particular module owns a memory controller. Some EDAC drivers are built into the kernel, and platform firmware can reserve or hide controller registers. Use the driver name and kernel configuration to establish which reporting path is active.

An EDAC memory controller may expose corrected-error and uncorrected-error counts, labels for channels or DIMMs, and reset or injection interfaces on supported platforms. Read-only status is appropriate for initial diagnosis. Do not write to injection nodes or reset counters on a production host while collecting evidence.

Corrected and uncorrected errors are different evidence

A corrected error means the memory controller detected an error and delivered corrected data according to its ECC capability. A rising corrected-error count can indicate a degrading DIMM, a channel or slot issue, a controller or firmware fault, temperature or power conditions, or a transient event. The count’s location and rate matter more than a single nonzero value.

An uncorrected error means the controller could not correct the affected data or could not guarantee its integrity. Depending on the hardware and architecture, the event can trigger a machine check, poison a page, terminate a process, panic, or be contained by a recovery mechanism. The exact result depends on whether the error was consumed, the page type, kernel support, and firmware behavior.

Corrected and uncorrected labels do not necessarily map one-to-one to replaceable modules. A memory controller can report channel, rank, DIMM label, row, bank, or address information with varying precision. The platform’s SMBIOS, service manual, and firmware inventory are needed to map a logical label to a physical part. A broken label can misdirect replacement work.

Preserve time series and event context

Capture the initial counters, kernel boot ID, system model, BIOS/UEFI and memory firmware, kernel release, and recent hardware events. Then collect deltas over a defined interval. A count that increases only under a repeatable memory workload differs from a static historical value or a counter reset after reboot.

Where configured, rasdaemon and its reporting tools can consume kernel RAS trace events and persist records. Check whether the daemon is enabled and whether its database spans the event. An absent database row does not prove no error occurred if the event was emitted before service startup, the tracepoint was unavailable, or the daemon was not running.

Correlate EDAC records with machine-check or corrected-error logs, BMC hardware event logs, memory training errors, thermal/power events, ECC scrubbing, and workload symptoms. Store timestamps and machine identity. Avoid sharing raw physical addresses publicly; they can reveal layout and are meaningful only with the same boot and memory map.

Distinguish a memory fault from the reporting path

A controller driver may fail to initialize because the platform exposes registers differently, firmware owns the reporting interface, or a kernel update changes driver binding. Missing EDAC counters can be a visibility gap rather than evidence of healthy RAM. Check the kernel log for probe errors and compare with vendor hardware telemetry.

Conversely, a noisy corrected-error counter does not automatically justify replacing an entire memory set. Reproduction after reseating or moving a DIMM should be done only under vendor maintenance procedures, with power removed and electrostatic precautions. On systems with mirrored or memory sparing features, the controller may have already remapped a region; firmware records are needed to understand the policy.

EDAC labeling can be incomplete or abstract. A driver may report a memory controller and channel but not a DIMM label, or it may expose a label inherited from firmware that does not match the chassis. The DIMM topology can include ranks, subchannels, chip selects, and interleaving. A physical address may be hashed across controllers or ranks, so a simple address-range calculation is not a reliable slot locator unless the platform documentation defines that mapping.

Some systems use firmware-first RAS handling, where BIOS or a platform management controller collects an event and may suppress direct kernel access to registers. Others use native kernel handling. Compare OS logs with BMC event records and vendor tools, but keep their timestamps and event identifiers. Two records can describe the same hardware error from different collection layers; deduplicate by evidence rather than counting every message as a separate fault.

When corrected errors recur on one channel, review whether they cluster in one physical location, rise with temperature or workload, or follow a firmware update. If they move with a DIMM after a supported service action, that is stronger evidence than a logical channel label alone. Do not reseat parts in a live system or use software to offline memory without an approved maintenance procedure.

Do not run destructive memory tests during production workload or treat a one-pass test as proof that an intermittent ECC error is resolved. Use vendor-approved diagnostics, controlled windows, and a copy of critical data. If uncorrected errors repeat, preserve evidence and follow the platform’s service process rather than suppressing the log.

Error injection belongs in a lab

Some EDAC drivers and kernel configurations expose error injection for testing the reporting pipeline. Injection support and semantics are hardware-specific; a test can corrupt memory or trigger a fatal machine check. Read the driver’s documentation and kernel RAS guide before enabling any injection path.

A safe lab test has a disposable machine, local console, known-good firmware, recovery media, test-only workload, and an explicit restoration plan. Verify whether injection targets an address, rank, channel, or a synthetic event. Confirm the reported event reaches the expected sysfs, trace, and daemon path. Do not use an injection interface to validate production alarms on the live memory controller.

Turn telemetry into an operational decision

Define thresholds using platform guidance, error rate, and service criticality. An alert should state whether it is corrected or uncorrected, which controller and label reported it, how quickly the counter rose, and whether a page or process was affected. A daily nonzero count may be acceptable on one platform and urgent on another.

Acceptance criteria should include complete detection coverage, stable device labeling, persistent event retention, counter-delta alerting, vendor escalation rules, and verified recovery for uncorrected faults. Test the alert pipeline with synthetic monitoring data or a lab injection, not with a real memory fault on a production server.

EDAC and RAS turn hardware error signals into useful service evidence only when the reporting path is understood. Keep corrected and uncorrected events separate, map labels to the physical platform, and preserve correlated logs before changing hardware.

Related:

Sources:

Comments