Linux Hardware Watchdogs: Device Ownership, Timeouts, and Safe Recovery
Operate Linux hardware watchdogs by understanding open/close ownership, keepalives, timeouts, nowayout behavior, pretimeout, and reboot recovery.
A hardware watchdog is an independent timer that can reset a system if software stops servicing it. Linux exposes watchdog devices to userspace, but the interface is intentionally hardware-dependent: timeouts, pretimeouts, stop behavior, magic close, and reboot handling vary by driver and device. A watchdog that is accidentally enabled, serviced by the wrong process, or stopped during a crash can fail precisely when recovery is needed.
Do not confuse a hardware watchdog with a service watchdog implemented through systemd notifications. A service watchdog detects missing application progress and can restart one unit; a hardware watchdog can reset the whole machine after the kernel or userspace stops servicing the device. They can be layered, but their owners and failure domains must be explicit.
Identify the watchdog and its owner
Linux may expose one or more watchdog character devices, commonly named watchdog0 and watchdog1. The device node is not a stable mapping to a physical controller across all machines. Read the sysfs identity and kernel log, then check the service manager or watchdog daemon that owns it.
ls -l /dev/watchdog*
ls -l /sys/class/watchdog/
cat /sys/class/watchdog/watchdog0/identity
journalctl -k -b --no-pager
Some sysfs attributes are optional; a missing timeout, state, or identity file can reflect driver capability rather than a broken device. Access to the character device may be privileged. Do not open it as an exploratory check: many drivers start the timer when the device is opened.
Determine whether systemd, a dedicated watchdog daemon, a vendor agent, or a custom application owns keepalives. Two independent processes must not race to open and service one watchdog unless their coordination contract is deliberate. A process that keeps pinging while the host’s critical services are dead can hide the very failure the watchdog should recover.
Opening and closing can change hardware state
The traditional watchdog API commonly starts servicing behavior when userspace opens the device. If keepalives stop, the timer expires and the hardware may reset the system. Closing the file can disable the timer on some drivers, while other devices support a magic-close protocol or a nowayout mode that prevents stopping after activation.
Magic close is not a universal safety feature. Where supported, a special character before close can indicate intentional shutdown; an unexpected close can leave the watchdog running until timeout. A driver configured as nowayout may ignore attempts to stop it entirely. Always confirm the driver documentation and current kernel configuration before testing process termination.
The file descriptor is therefore an operational resource, not a passive status handle. A monitoring script that opens and closes the device to inspect it can arm a reboot timer. Use sysfs and the owning service’s documented status interface for observation. Never run a generic write or ioctl test against the production watchdog to see whether it responds.
Timeouts, keepalives, and pretimeout
The requested timeout is a goal constrained by hardware and driver capabilities. A driver may round it to supported values, reject an out-of-range value, or expose a maximum and minimum. Read back the effective value after configuration. The moment of last keepalive and the hardware’s countdown may not align exactly with wall-clock samples.
Choose a timeout based on the failure-detection objective and the slowest expected healthy boot or maintenance interval. A very short interval can reset a machine during normal storage recovery, firmware update, suspend transition, or scheduler stall. A very long interval may violate recovery objectives. Add margin for scheduling and hardware tolerance, but do not inflate the timeout without documenting the resulting maximum outage.
Pretimeout support can generate an earlier signal or interrupt before final reset. It is useful only when the platform can preserve evidence and the handler has enough time to act. Pretimeout is not guaranteed to provide a full crash dump; it may merely log an event or produce a nonmaskable interrupt. Verify exactly what the platform implements.
Keepalive writes or the documented ioctl refresh the countdown. The application should service the hardware only while its health policy passes. A separate timer thread that continues pings while the request path, storage, or essential worker threads are frozen defeats end-to-end monitoring. Conversely, tying the watchdog to one optional worker can create false resets during expected idleness.
Understand the API and capability negotiation
The watchdog API defines ioctls to query support, set or get timeout, send keepalive, configure pretimeout, and inspect boot status where implemented. Drivers can support only a subset. A command returning an unsupported error is not proof that the watchdog itself is defective.
A robust manager should query capabilities and effective values, record the chosen device identity, and have a single clear owner. It must handle device disappearance, write errors, service restart, and intentional shutdown. If the watchdog is configured to remain active across close or reboot, the system’s shutdown and boot sequence must continue to service or intentionally hand off the device.
Avoid writing directly to the watchdog from a one-off shell command. A successful ping only proves the device accepted that operation; it does not validate the timeout, owner, reboot action, or system recovery. Any destructive test belongs on a lab system with console access and a recovery plan.
Integrate with host health and systemd
Systemd can manage hardware watchdog activity through manager configuration on systems that support it. A service-level sd_notify watchdog is a different protocol. The former attempts to reset a host that is no longer servicing the hardware; the latter asks the service manager to react when an application stops reporting progress.
If both are used, define the health chain. The application reports genuine progress to systemd. The manager’s own watchdog is serviced only while its main control loop is responsive. Hardware reset is the final recovery layer. Avoid a design where an independent process feeds the hardware watchdog forever while PID 1 or the storage stack is dead.
Observe restart counters, boot status, kernel logs, watchdog manager logs, and the hardware’s reset reason after a controlled failure. A reset can lose volatile data and interrupt writes. Services must recover idempotently, storage must have an appropriate durability contract, and the system should avoid reboot loops through bounded retries or maintenance mode.
Safe validation procedure
Test in a disposable machine or lab device with serial or out-of-band console access. First document the current watchdog identity, timeout, kernel configuration, service owner, magic-close and nowayout settings, boot behavior, and recovery route. Then use the vendor or kernel-documented test procedure to trigger a missed keepalive; do not improvise a process kill on production.
Capture the last keepalive time, pretimeout behavior if supported, reset reason, service restart count, boot journal, filesystem check status, and application recovery. Verify the host returns to a healthy state and does not simply restart into the same failure. Test a planned reboot and shutdown separately, because stop semantics can differ from an unexpected process crash.
Acceptance criteria should state who owns each watchdog, the effective timeout and tolerance, how service health gates keepalives, how an intentional close or handoff works, and what evidence is available after reset. Include a negative test where progress stops and the intended recovery occurs, plus a false-positive test under the longest expected healthy pause.
The hardware watchdog is a last-resort reset mechanism, not a substitute for monitoring or safe application recovery. Treat open, keepalive, close, reboot, and pretimeout as explicit lifecycle transitions. That makes the device useful under real failures instead of an unexplained source of periodic reboots.
Related:
- Linux Kdump: Reserved Memory, Capture Kernels, and vmcore Validation
- systemd Readiness and Watchdogs with sd_notify
Sources: