Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux pstore and ramoops: Recover Crash Evidence After Reboot

Preserve Linux oops and panic evidence across reboot with pstore backends, ramoops memory regions, boot-time sizing, and disciplined log recovery.

A kernel crash can destroy the ordinary journal path before the final diagnostics reach persistent storage. The Linux pstore framework provides a way for supported backends to preserve selected records across reboot. ramoops stores oops and panic records in a reserved RAM region that survives a system restart, while pstore/blk can use a block-device path designed for crash-time logging. These mechanisms are not full memory dumps, not universal across hardware, and not automatically enabled just because /sys/fs/pstore exists.

This guide explains the evidence pipeline and a careful validation plan. Backend choice depends on architecture, firmware, kernel configuration, persistent-memory availability, and storage-driver behavior at panic time. Use the kernel documentation and platform runbook for the exact target; do not paste a sample physical address or block device into a production boot configuration.

pstore front ends and storage backends

The pstore filesystem exposes records produced by front ends such as kernel message dumps. A backend stores those records in a way that can be recovered after reboot. ramoops reserves a region of persistent RAM and writes records through a circular buffer. It requires memory that retains contents across the reset or reboot path used by the platform. Ordinary DRAM that loses power will not preserve a record through power loss.

pstore/blk is a distinct backend that writes crash records to a configured block or non-block device using panic-appropriate operations. A block device that works during normal operation may not be safe at panic time if the controller, driver, locks, or interrupt path are unavailable. The kernel documentation requires the panic I/O path to avoid allocation, scheduling, sleeping, and ordinary locking. That makes backend design a hardware and driver contract rather than a generic file write from panic context.

The backend capacity is limited. Ramoops divides reserved space among record types and rotates records as space fills. Block-backed storage similarly has configured sizes and replacement behavior. Decide how many records and which front ends matter for diagnosis, and ensure the chosen configuration cannot overwrite a larger incident record before recovery.

Discover support before configuring

Inspect the running kernel’s configuration, module availability, boot messages, and the platform’s reserved-memory or crash-storage layout. For ramoops, the physical range, size, memory type, and region definition must match the hardware and firmware map. Device Tree bindings or platform-specific reservation mechanisms can own these values. A guessed address can collide with usable memory or another firmware reservation.

For pstore/blk, identify the device by a stable identifier supported in the target boot configuration. A /dev name can change as storage discovery order changes. When the backend is built into the kernel, the device or partition may need to be specified in a form available before ordinary device-node creation. Verify the built-in versus module case because accepted configuration forms differ.

Before changing a production boot line, record the present kernel command line, memory map, and pstore-related kernel messages. Confirm the selected backend claims the intended region or device and that no other subsystem uses it. Do not make crash logging compete with swap, a filesystem journal, or firmware-reserved memory.

Configure record sizing intentionally

Ramoops can store distinct record classes such as console, ftrace, pmsg, and kernel message dumps according to its configuration and build options. The capacity allocated to each class affects what survives. A large continuous console buffer can crowd out discrete oops records; a very small kmsg area can truncate the trace needed to locate a panic. Measure realistic incident output in a controlled test and size accordingly.

For pstore/blk, user parameters can include chunk sizes such as kmsg_size and pmsg_size, while the backend driver owns device-specific read and write behavior. Read the target kernel’s module parameter and Kconfig documentation before specifying values. The documentation states that kmsg chunks have alignment requirements, so validate units and constraints carefully rather than copying a value from a different version.

Do not use persistent crash storage as an uncontrolled log archive. It is a bounded recovery channel with a defined overwrite policy. A userspace service should collect and rotate records after boot, preserve relevant timestamps and boot identity, and then delete or archive the pstore files according to incident-retention policy.

Recover records after reboot

Mounting the pstore filesystem exposes records when a backend registered data. Some distributions mount it automatically or collect records early in boot. If early boot cleanup removes records before your collection service runs, you may lose evidence. Inspect service ordering and the actual files immediately after a controlled reboot.

findmnt /sys/fs/pstore || true
find /sys/fs/pstore -maxdepth 1 -type f -print

This is a read-only discovery example. A production collector should copy records to persistent storage before unlinking them, associate them with the boot ID and timestamp, and handle an empty directory as a normal case. The kernel documentation notes that deleting a stored pstore record can be performed by unlinking the respective file; do that only after a durable copy exists and your retention policy allows it.

Treat filenames and record ordering as backend-specific. A record name can indicate a dump sequence or front end, but the filename alone is not a complete incident identifier. Preserve the original bytes, record metadata and any decoding steps, and correlate the record with the boot’s journal, systemd-coredump data, watchdog reports, and firmware logs.

Controlled failure testing

Crash testing is destructive and should occur only on a disposable machine, VM, or lab system with a recovery path. Do not trigger a kernel panic on a production host or a remote machine whose storage and access cannot be recovered. The test must verify the complete chain: backend registered, a controlled oops or panic produced a record, system restarted, a collector found it, the copy was durable, and the next boot did not silently discard earlier evidence.

For ramoops, test whether the region survives the actual reset path. A warm reboot may preserve RAM while a power cycle does not; document which incident classes the platform can capture. For pstore/blk, validate the panic path and controller behavior under the same storage topology used in production. A successful normal write is not proof that panic-time polling I/O works.

Use the kernel’s supported test mechanism for the relevant backend and verify platform-specific prerequisites. Keep console access and a tested boot recovery route. Repeat across kernel upgrades because Kconfig, command-line syntax, backend implementations, and boot collection order can change.

Operational failure modes

No files can mean the backend is unsupported, not configured, failed to reserve storage, or has no records. It can also mean an early service already collected or deleted the previous boot’s evidence. Check dmesg, boot parameters, kernel configuration, service logs, and the reserved resource as separate facts.

Truncated data can result from undersized buffers or several front ends sharing limited capacity. Missing latest records can reflect a circular overwrite policy. Corrupted or stale data can indicate the region was not actually persistent or was used by another component. Do not treat one missing dmesg-ramoops-* file as proof that the system did not crash.

Persistent kernel logs may contain information about processes, paths, devices, or workload configuration. Restrict access to the collector and exported evidence, use appropriate retention, and redact before external sharing. Keep diagnostics useful without turning a crash channel into an unbounded store of machine data.

Acceptance checklist

Document the kernel version and configuration, selected backend, exact reserved region or device identity, size breakdown, panic-time I/O constraints, record rotation behavior, early-boot collector, deletion policy, and recovery process. Test reboot and power-loss assumptions separately. Confirm an operator can retrieve evidence when the root filesystem is unavailable or damaged.

pstore is a recovery interface, not a guarantee that every crash becomes a complete dump. Choose a backend supported by the platform, respect the restricted execution context at panic time, budget the bounded storage, and prove record survival with a controlled test before relying on it during an incident.

Related:

Sources:

Comments