Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux PCIe AER: Diagnose and Recover from Link Errors

Interpret Linux PCIe Advanced Error Reporting, identify the affected hierarchy, and distinguish corrected events from recovery that requires driver coordination.

PCI Express Advanced Error Reporting (AER) messages are evidence about a link or transaction, not a diagnosis by themselves. A corrected error can be logged without taking a device offline; an uncorrectable error may require the PCIe hierarchy and participating drivers to stop I/O, reset or isolate a component, and restore service. The same message can arise from a marginal link, power transition, device firmware, platform firmware, or a real endpoint defect.

Linux’s AER service collects error status and coordinates recovery through PCI error-recovery callbacks for drivers that participate in the framework. It does not make every device recoverable, and a successful recovery callback does not prove that application data remained valid. Diagnose the affected hierarchy, severity, recovery trace, and workload impact together.

Start by capturing the complete kernel log around the event and identify each BDF address mentioned. AER logs commonly name a Root Port or Downstream Port as well as an endpoint. The reporting port is not necessarily the faulty component: it may be where an error was observed or contained. Map every BDF to its parent bridge, slot, and driver before replacing hardware.

Read-only inventory examples:

lspci -D -nn
lspci -D -t
lspci -D -vv -s 0000:03:00.0
journalctl -k -b --no-pager

Replace the example address with a BDF from the log. The topology view helps map devices to bridges; verbose configuration can expose link capability and negotiated state when the device and platform report it. Do not infer the physical slot solely from bus numbering. Correlate with firmware inventory, chassis labels, device paths in sysfs, and the platform’s slot mapping.

Record kernel release, BIOS/UEFI and device firmware revisions, link width and speed, workload, temperature, recent power transitions, and whether errors repeat on the same hierarchy. A single corrected event after boot has a different meaning from a rapidly increasing corrected-error counter under sustained traffic.

Understand severity and containment

PCIe defines correctable and uncorrectable errors, with uncorrectable errors further classified by whether they are fatal. Correctable means the link or protocol machinery recovered the transaction according to the specification; it does not mean the underlying physical condition is harmless. Repeated corrected errors can be an early signal of a degrading cable, connector, slot, signal integrity, power, or device.

AER registers record status and masks for different error classes. The Linux service may log an error and continue, or it may invoke recovery when an uncorrectable error affects a hierarchy. Kernel options, platform firmware ownership, native AER support, and device capabilities influence what Linux can observe. If the platform firmware handles AER, Linux may have little or no control over the path.

Error containment is important. An endpoint error can be isolated to a function, link, or subordinate hierarchy depending on the error and topology. Do not assume that one endpoint log line proves only that function was affected. Conversely, a fatal event at an upstream port does not prove every child device is defective.

Recovery is a driver protocol

The PCI error-recovery framework lets a participating driver respond to stages such as detecting an error, deciding whether it can recover, resetting a channel or slot, and resuming normal operation. Drivers must stop new work, quiesce DMA and interrupts where needed, preserve or invalidate device state appropriately, and reinitialize after the platform reports recovery. A driver that does not implement the required callbacks may leave the device unusable even if the PCIe link itself is restored.

The framework’s callback return values have defined meanings; they are not generic success booleans. A driver may report that it recovered, needs a slot reset, or cannot recover. The PCI core coordinates the sequence across the affected hierarchy. For multi-function devices, dependent functions may also need to participate. A reset can destroy volatile device state, interrupt queues, or outstanding operations, so higher layers must be prepared for I/O errors and reconnection.

Never trigger a surprise removal, hot reset, or AER injection against a production device simply to see whether recovery works. Use the kernel’s documented test facility only on a dedicated lab system with a recoverable device and a prepared out-of-band management route. Review the exact test interface for the running kernel and understand that a reset can disrupt all functions below the selected port.

Separate an error burst from service impact

Count corrected and uncorrected events over time, but avoid converting a raw count into a universal alert threshold. Devices differ in link speed, reporting, firmware, and workload. Correlate AER with device resets, driver messages, NVMe/SCSI timeouts, network carrier changes, filesystem errors, application latency, and hardware-management logs.

AER status is often latched until software or hardware clears it, so a displayed bit can describe an earlier event rather than a current active fault. Kernel logs may summarize multiple status bits, suppress repeated messages, or report an error at the Root Port that observed it. Preserve the raw log and compare it with the reporting device’s link state and the endpoint’s own counters where available. Do not clear a status register simply to see whether the message returns; clearing it can destroy evidence and some registers have write-one-to-clear semantics.

PCIe link retraining or a downstream device reset can change negotiated width or speed. Read the capability and current status fields for the relevant port and endpoint, and compare before and after a controlled load. A reduced link width can be a platform policy or an effect of a failed lane, not necessarily a bandwidth bottleneck. AER severity and link state should be interpreted with the device’s workload and vendor guidance.

For storage endpoints, capture device health and transport counters before changing queue depth or retry settings. For a network adapter, correlate PCIe errors with link state, queue resets, and packet drops. AER may be one contributor in a broader failure chain; it cannot tell whether an application committed a transaction before the device reset.

Avoid immediately disabling AER or masking errors to quiet a log. Masking changes visibility and may remove evidence needed for diagnosis. If an operational exception is required, scope it to a documented platform issue, preserve counters through another channel, and define an expiration and rollback.

Evidence collection without changing the machine

Capture logs and topology before rebooting, because counters and transient state may be lost:

lspci -D -t
lspci -D -vv -s 0000:00:1c.0
journalctl -k -b --no-pager
find /sys/bus/pci/devices -maxdepth 1 -type l -print

The BDF is an example. Some sysfs files require elevated read access, and a container may not expose the host topology. Keep output with time stamps and the exact machine identity. Logs can contain serial numbers and device paths; redact them before sending outside the organization.

If the device remains usable, compare AER counters before and after a controlled workload. Do not clear status registers as an exploratory step: some counters are write-one-to-clear or have device-specific effects, and clearing destroys historical evidence. Prefer vendor-supported diagnostics and a maintenance window before changing firmware, link speed, slot, cable, or power settings.

Acceptance and escalation

A useful incident report identifies the reporting port, subordinate device, severity, status bits, frequency, driver callback sequence, recovery outcome, and application symptoms. State whether Linux or firmware owned the AER service, whether recovery callbacks were available, and which device state was lost. Include the physical topology so a replacement decision is evidence-based.

Escalate to the platform or device vendor when errors persist after supported firmware and physical-link checks, when firmware owns AER without exposing useful diagnostics, or when the driver cannot recover repeatedly. Preserve unmodified logs and avoid drawing a root-cause conclusion from one status bit. Reproduction on a supported kernel and a known-good slot or system helps distinguish endpoint defects from platform signal integrity.

PCIe AER is most useful when treated as a protocol event with hierarchy and recovery context. The corrective action follows from the trace: stabilize a physical link, correct firmware, repair a driver recovery path, or replace a failing component. Quieting the message without understanding its origin only hides the next failure.

Related:

Sources:

Comments