Linux Core-Dump Triage with systemd-coredump, coredumpctl, and GDB
A systemd/Linux core-dump workflow for limits, kernel routing, coredumpctl, GDB symbols, storage budgets, retention, and post-crash verification.
A process that terminates after a fatal signal may leave a core dump: a snapshot of some or all of its userspace memory that a debugger can inspect after the process is gone. On a system using systemd, the kernel often pipes that dump to systemd-coredump, which records metadata in the journal and may store the core separately. The exact result depends on the kernel’s core_pattern, the process’s resource limits, service configuration, systemd-coredump policy, and available disk space.
Core dumps are different from kernel crash dumps. A userspace core is for one terminated process and can contain application state; a kernel panic or machine crash uses a separate crash-dump mechanism. It is also different from strace: a core captures state at termination, while a syscall trace records interactions as they happen. Choose the evidence that fits the failure instead of treating every crash as a kernel problem.
Confirm how this host routes core dumps
Start by checking the kernel’s configured destination and the process’s applicable limit:
sysctl kernel.core_pattern
ulimit -c
If core_pattern begins with a pipe (|), the kernel sends the dump to a userspace handler rather than creating a file named core in the crashing process’s current directory. A systemd installation commonly uses a systemd-coredump handler, but do not assume this from the presence of systemd alone; inspect the actual pattern and installed units. If the pattern is a filename template instead, check the process’s working directory, directory permissions, file-size limits, mount namespace, free bytes, inodes, and filesystem quota.
ulimit -c reports the current shell’s soft core-file size limit. It applies to the shell and child processes it starts; it does not change a daemon’s already-running resource limits. For a process started as a systemd service, inspect the unit and effective property instead:
systemctl cat myapp.service
systemctl show myapp.service -p LimitCORE
Replace myapp.service with the actual unit. A service inherits limits when it starts, so a changed unit configuration does not alter a process that is already running. Other launchers, containers, schedulers, and PAM sessions may apply different limits.
Use coredumpctl to find the event and inspect its metadata
First establish whether an event was recorded and whether the corresponding core is still available:
coredumpctl --no-pager list --since today
coredumpctl --no-pager info 12345
Replace the sample PID 12345 with the ID shown for the crash, and match the timestamp and executable too. PIDs are reused, and more than one crash can share an executable name; do not assume that the newest record is the failure you are investigating. coredumpctl list distinguishes a recorded event from an accessible core file, and coredumpctl info can expose the signal, timestamp, command line, executable, unit, storage path, and any stack trace systemd collected.
If the core is available and GDB is installed, invoke it through the matching record:
coredumpctl debug 12345
Or extract it for a controlled analysis environment:
coredumpctl --output=./core.12345 dump 12345
file ./core.12345
gdb /path/to/the/matching/executable ./core.12345
Replace sample PID 12345 and the executable path, preserve the exact executable build, and use a destination with enough space and appropriate access controls. An extracted dump can be as large as the process address space permitted by policy. The core is not self-describing source code: install or retrieve debug symbols matching the exact executable and shared-library build IDs before trusting source lines or variable names. A backtrace without matching symbols can still identify a crash path, but it is not proof that the displayed source matches the running binary.
Diagnose why no usable core exists
When the application crashed but there is no dump, separate “no event was generated,” “metadata was recorded but core storage was skipped,” and “the file existed but was later removed.” Check, in order:
- Was termination by a signal whose default action includes a core dump, rather than a clean exit,
SIGKILL, or an application-managed restart? - What did
kernel.core_patterncontain at the time, and did a handler run? If the pattern is piped, the usualRLIMIT_COREfile-size rule is not enforced by the kernel in the same way as direct file output. - What resource limit applied to the actual crashing process? For a shell-launched reproduction, inspect
ulimit -c; for a service, inspectLimitCORE=and the running service’s start time. - Does
coredumpctl listshow a record markednone,missing,truncated, orpresent? These states point to different storage or retention paths. - Was the handler’s processing or storage limited by systemd-coredump configuration, filesystem capacity, quota, or cleanup policy?
- Could the process’s memory mappings have been excluded from the dump, or could a security policy, namespace, or executable property prevent normal dumping?
The Linux core(5) documentation lists additional reasons a dump may not be created, including an unwritable destination, full filesystem, zero limit in direct-file mode, or an empty core pattern. Do not change kernel.core_pattern to a hard-coded filename just because no file appeared in the working directory: that can bypass the existing collector, change retention and permissions, and cause an uncontrolled disk-growth problem.
Configure a service limit only for the service that needs it
If the service is expected to produce a core and its effective limit is too low, use a narrow systemd drop-in for that unit rather than changing every login shell or the global service-manager policy. For a temporary diagnostic window, a drop-in can set a chosen limit:
# Output of: systemctl edit myapp.service
[Service]
LimitCORE=infinity
infinity is an example for a controlled capture window, not a universal production recommendation. After saving the drop-in, reload the manager and restart the service in a planned window so the new process inherits it:
sudo systemctl daemon-reload
sudo systemctl restart myapp.service
systemctl show myapp.service -p LimitCORE
A larger process limit does not guarantee that systemd-coredump will process or store an equally large dump. Its own size and storage policy is independent, and direct-file behavior differs from a piped kernel handler. Decide the required limit from the diagnostic need, memory footprint, filesystem capacity, and incident-retention policy; remove or narrow a temporary override when the capture window ends.
Bound storage and retention before enabling broad capture
systemd-coredump has configurable controls in /etc/systemd/coredump.conf and drop-ins under /etc/systemd/coredump.conf.d/. Options include whether dumps are stored externally or in the journal, maximum sizes processed or stored, aggregate disk-use limits, free-space reserve, and compression behavior. Read the installed coredump.conf(5) because names, defaults, and behavior can vary by systemd version. The systemd-tmpfiles policy also participates in cleanup; a journal entry can remain after an external core file has been deleted.
Before increasing a limit, measure the target filesystem and decide what should happen during a burst of repeated crashes. Set bounded processing/storage policy appropriate to the workload, alert on space pressure, and keep core files outside source-control directories and general-purpose backups unless policy explicitly requires them. Test the configured maximum with a controlled test process on a nonproduction machine rather than inferring the effective cap from a configuration file alone.
Treat a core as a memory image, not a harmless stack trace. It can contain credentials, request data, cryptographic material, and user information that happened to be in process memory. Limit who can retrieve it, keep it only as long as the investigation requires, and use an approved transfer channel for debugging. This is data-handling for a diagnostic artifact, not a substitute for the host’s broader access-control program.
Turn the dump into a reproducible bug report
Record the exact UTC/local timestamp and time zone, executable path and build ID, package version, signal, service unit, kernel and systemd versions, command line with secrets removed, and whether the issue reproduces under the same workload. Keep the core, coredumpctl info, journal excerpt, symbols, and matching executable together as a labeled evidence set. If systemd’s automatic backtrace points into a library, verify that the installed library matches the one loaded by the process; recent package updates can otherwise make the investigation misleading.
For a repeatable application crash, collect one representative dump and compare multiple crashes only when it helps identify a stable signature. For an intermittent race, a core is a single endpoint snapshot and may not reveal the prior event sequence; targeted logs, tracing, or a controlled reproducer may be more informative. If the core is truncated or storage says missing, report that fact rather than presenting a backtrace as complete.
Acceptance criteria for a useful core-dump path
The workflow is ready when a controlled test on the same service proves that the configured signal is captured, coredumpctl lists the matching event, the expected core is present and retrievable, the analysis uses matching symbols, and the storage policy remains within an explicit budget. Verify that normal service startup, restart, and cleanup still work after the diagnostic override is removed. A core file existing somewhere on disk is not enough; it must be attributable to the correct binary, safe to handle, and useful to a debugger.
Related:
Sources: