Linux Kdump: Reserved Memory, Capture Kernels, and vmcore Validation
Design and verify Linux kdump around crashkernel reservation, capture-kernel dependencies, vmcore storage, and controlled panic testing.
Kdump preserves a kernel memory image after a panic by using kexec to boot a separate capture kernel. The first kernel reserves memory for that capture environment before the failure; when a panic occurs, the capture kernel starts in the reserved region and can expose the crashed kernel’s memory as /proc/vmcore. A dump service then copies or filters that image to storage.
Kdump is a diagnostic pipeline, not a guarantee that every panic produces a complete dump. Reservation can fail, the capture kernel can lack storage or network drivers, the dump target can be unavailable, or a second failure can interrupt collection. Validate every stage on the actual hardware and kernel build. A boot parameter appearing in /proc/cmdline proves only that a parameter was passed, not that the capture kernel is loaded or that a useful dump can be written.
Separate the first kernel from the capture kernel
The production kernel runs the workload. A memory region is reserved for a second kernel, which is loaded in advance. During a panic, the first kernel transfers control to the capture kernel using the kexec crash path. The reserved-memory boundary is essential: the capture kernel must not overwrite the memory image it is supposed to inspect.
The capture environment can be a small initramfs or another platform-supported root. It needs enough CPU, memory, drivers, and userspace tooling to locate and write the dump target. If the destination is local storage, include the relevant controller and filesystem support. If it is remote, include the network device, link configuration, route, credentials, and transport tools needed to reach the collection server during early boot.
Do not assume the ordinary system’s mounted filesystem, network manager, or secrets are automatically available in the capture environment. The kdump service and initramfs builder for the distribution assemble that environment, and their configuration differs. Inspect the generated capture initramfs and service logs instead of copying one distribution’s commands to another.
Reserve crash memory at boot
The crashkernel= boot parameter requests memory for the capture kernel. The supported syntax and sizing behavior are documented by the kernel and may depend on architecture and bootloader. A reservation that is too small may prevent the capture kernel from loading or leave insufficient room for dump processing; a reservation that is unnecessarily large reduces memory available to the first kernel.
cat /proc/cmdline
grep -i 'Crash kernel' /proc/iomem
cat /sys/kernel/kexec_crash_loaded
These checks answer different questions. /proc/cmdline shows the boot arguments, /proc/iomem can show a reserved crash region when the kernel exposes it, and kexec_crash_loaded reports whether a crash kernel is currently loaded. Availability and exact presentation vary with kernel configuration and architecture; also inspect the distribution’s kdump service status and logs.
Avoid selecting a fixed address or memory size from an unrelated machine. Physical-memory layout, firmware reservations, device mappings, kernel configuration, architecture, and capture-initramfs contents affect the required reservation. Use the distribution’s supported sizing policy, review bootloader output after regeneration, and retain a known-good boot entry while testing changes.
Load and verify the capture environment
The normal operating system prepares the crash kernel before an incident. A kdump service usually coordinates loading it and building the capture environment; the user-space commands and service names are distribution-specific. Verify that the service reports success after each kernel update, because a new kernel can change the capture image, module dependencies, or initramfs composition.
Check that the crash kernel is loaded after boot and after service restarts. Review logs for missing modules, invalid memory reservations, unsupported kexec operations, or initramfs generation errors. If the platform offers an explicit kdump status command, include it in routine health checks. Do not treat an enabled service unit as proof that its capture kernel loaded successfully.
The capture kernel may need a command line different from the first kernel. It must avoid recursively taking another dump and must find the reserved memory and intended destination. Use the distribution’s documented configuration rather than hand-assembling a command line from examples. Confirm that drivers and firmware required for the destination are present in the capture image; they may not be loaded in time if the capture environment expects to mount its root from that same destination.
Choose a dump destination and capacity plan
An unfiltered kernel memory image can be large. Estimate capacity using the system’s memory size, the kernel’s dump format, the filtering policy, and operational retention. A compressed or filtered dump may reduce storage but can still be substantial, and the crash kernel requires its own working memory. Avoid placing the only copy on a filesystem that is likely to be full during the incident.
For a local destination, verify that the block device is identified reliably and that the capture environment can mount or write it. Device names can change; use the platform’s stable identifiers and test the actual capture path. For a remote destination, test network bring-up from the capture initramfs, including DNS or static addressing as configured, firewall policy, routing, and authentication. A successful upload from the normal system does not prove that the early capture environment can reach the same server.
Treat dump files as sensitive. A vmcore can contain process memory, credentials, keys, customer data, and fragments of prior activity. Restrict access and transport, protect storage at rest under the organization’s policy, limit retention, and define who can analyze or export the artifact. These controls should not prevent the capture kernel from writing the dump, so test them as part of the end-to-end procedure.
Validate vmcore generation without risking a production host
The kernel exposes the crashed memory image to the capture environment through /proc/vmcore. The kdump service or its configured helper copies that image to the destination, optionally applying a filtering tool such as makedumpfile. The final artifact should be identifiable with the originating host, kernel release, boot identifier, incident time, and collection status; otherwise a technically successful dump can be difficult to interpret later.
Test using the supported kdump test mechanism for the distribution on a disposable host or a maintenance window with a recovery plan. A forced panic is intentionally disruptive: it terminates the running kernel and workload. Never trigger it on a production system merely to see whether a command works. Verify that the test produces a readable artifact, that the dump analyzer recognizes the kernel and architecture, and that the capture service returns or powers off according to the configured policy.
Exercise failure cases too: a full destination, unreachable network, missing module, and interrupted collection should produce a clear status and leave the host in a known state. Confirm the next normal boot does not repeatedly enter the capture path, and that the kdump service reloads the expected kernel. Test after kernel, initramfs, storage, firmware, or network-driver changes that could affect capture.
Diagnose a missing or unusable dump
If no dump was produced, determine how far the pipeline progressed. Was the reservation present? Was a crash kernel loaded? Did the panic transfer to it? Did its initramfs start? Could it read /proc/vmcore? Did it locate the target and complete the write? Service logs, serial or console output, persistent kernel records, and the target’s filesystem state can distinguish those failures.
An absent /proc/vmcore in the capture environment points to a different problem than a failed network upload. A partial file can indicate capacity exhaustion, I/O failure, a second crash, or a filter/write error. Preserve the relevant logs before retrying or cleaning the target. Avoid repeatedly forcing a panic while changing several variables at once; isolate one failed stage and verify the correction in a controlled test.
Kdump complements pstore rather than replacing it. Pstore can preserve selected panic or oops records through a supported backend, while kdump aims to capture a broader memory image. Use both only if the platform supports the required persistence mechanisms and the diagnostic value justifies their operational cost. A normal user-space core dump is a separate mechanism for one process and does not substitute for kernel crash capture.
Kdump is dependable only when the reservation, capture kernel, early userspace, destination, and analysis workflow are all tested together. Verify each layer after system changes, protect collected memory images, and keep a safe recovery path for deliberate test failures.
Related:
- Linux pstore and ramoops: Recover Crash Evidence After Reboot
- Linux Core-Dump Triage with systemd-coredump, coredumpctl, and GDB
Sources: