Operating FreeBSD Watchdogs Safely: Detection, Recovery, and Failure Modes
Deploy FreeBSD watchdogd with verified hardware support, bounded health checks, safe timeouts, shutdown behavior, and measurable recovery tests.
A watchdog is a recovery mechanism for a host that has stopped making progress, not an availability system and not a substitute for service monitoring. FreeBSD’s watchdogd(8) periodically interacts with the kernel watchdog facility. If the daemon or the system stops servicing the timer for long enough, the selected watchdog implementation can take a configured action. That action may help recover an unattended machine, but it can also cause a needless reset, destroy volatile incident evidence, or create an endless reboot loop when the underlying fault persists.
The first production decision is therefore not the timeout value. It is the failure you intend to detect, the system boundary that can still observe it, and the recovery action that is safe for that failure. A hardware timer can detect a kernel or scheduling stall that a userland process cannot recover from. It cannot prove that an application is serving correct responses, that a database has a healthy replica, or that a remote dependency is reachable.
Know which watchdog is actually armed
FreeBSD exposes hardware and software watchdog implementations through the kernel watchdog facility. A request can arm a supported hardware watchdog when one is available; if no hardware implementation can process the request, the default software watchdog may be used. A software watchdog has different failure coverage: if the kernel itself is unable to run the code that services that timer, it may not provide the same independent reset path as a motherboard timer. Hardware presence, driver attachment, accepted timeout granularity, BIOS configuration, and the action after expiry are host-specific.
Inspect the actual host before enabling the service:
freebsd-version -kru
dmesg | grep -i watchdog
kldstat
sysrc watchdogd_enable watchdogd_flags watchdogd_timeout watchdogd_shutdown_timeout
The dmesg output and loaded-module list are clues, not proof that the desired hardware timer is armed. Some devices are compiled into the kernel rather than loaded as a module. Check the exact device driver’s manual and platform firmware settings, then verify behavior in a disposable or redundant test host. Do not infer hardware reset capability merely because watchdogd starts without an error.
watchdog(4) documents the kernel interface and watchdog(8) is a direct control utility. Directly setting a short timeout with watchdog -t is capable of making the machine reset if the watchdog is not serviced. Do not use it as an exploratory command on a production host. In addition, watchdog disarming can itself fail; the manual warns that an error while disabling a timer does not prove that it is no longer armed.
Decide what “healthy” means before writing a check
Without a custom command, watchdogd performs a trivial filesystem check. With -e, it invokes the supplied command through system(3) and resets the watchdog only when the command exits successfully. This is not a general service supervisor. The return code is the signal; a check that only prints “OK,” backgrounds its work, or always returns zero gives no useful health guarantee.
A custom check should be deterministic, fast, bounded, and local enough that it does not convert a remote outage into a host reboot. A script that waits indefinitely for DNS, a database, or an external endpoint may stop feeding the watchdog during a dependency failure even though the kernel and host are otherwise healthy. Conversely, an over-simple check can keep the timer refreshed while the service users care about is broken. Define explicitly whether the check proves operating-system progress, storage availability, or a particular local service, and use application health monitoring for deeper semantics.
For a controlled dry run, watchdogd -n runs the check without arming the system watchdog. A harmless foreground invocation can validate the command path and exit-status handling:
watchdogd --debug -n -e /usr/bin/true
--debug prevents daemonization; -n is the documented dry-run mode. Replace /usr/bin/true only after the intended check is ready, and test both a zero exit and a deliberate non-zero exit. Dry-run verifies command execution, not hardware support, timer servicing under load, or the eventual reset path. Those require a planned test on a recoverable system.
Configure the service as explicit policy
The rc configuration controls whether watchdogd starts at boot, its flags, and shutdown behavior. Confirm the installed release’s rc.conf(5) because the timeout and shutdown settings can override corresponding daemon flags.
# /etc/rc.conf, illustrative values only
watchdogd_enable="YES"
watchdogd_flags="-t 120 -s 10"
watchdogd_shutdown_timeout="180"
These values are examples, not universal defaults. -t requests the watchdog timeout and -s sets the interval between checks. The hardware may round, reject, or not support a requested timeout. The manual page’s defaults are not a capacity plan: measure the longest legitimate check duration, worst expected scheduler delay, boot behavior, and maintenance window for the target host. Leave margin between the check interval and timeout, and ensure a slow check cannot overlap in a way the daemon handles unexpectedly.
Only add -e /path/to/check after reviewing script ownership, executable permissions, dependencies, and logging. If the script needs credentials, store them with narrowly restricted access and do not echo them. Keep it independent of network mounts if the host must recover from a network stall. A non-zero result intentionally stops watchdog refresh, so return failure only for a condition for which host-level reset is an acceptable remedy.
Before making the setting persistent, validate the rc values and examine the service state during a maintenance window. Do not edit /etc/defaults/rc.conf; use the local configuration and sysrc so the effective settings can be reviewed. Stopping or restarting the watchdog service changes kernel timer state, so do not test shutdown semantics through a remote-only session without console or out-of-band access.
Understand timeout actions and reset loops
The watchdog facility may be hardware-backed, software-backed, or a combination mediated by the kernel. watchdogd(8) supports soft timeouts and pretimeouts with actions such as logging, printing, entering the debugger, or panicking, depending on flags and kernel configuration. A pretimeout is an opportunity to record evidence before the main watchdog deadline; it is not guaranteed to run if the whole machine or hardware path has failed. Debugger actions are unsuitable for an unattended server unless a human or an explicit recovery design can handle the stop.
Choose the policy based on recovery objective:
- Log or print before reset: useful when the system can still write a diagnostic message, but not proof that persistent logs will survive a hard reset.
- Panic before reset: can produce a crash dump when dump capture is correctly configured, but requires storage capacity, matching symbols, and a tested
savecorepath. - Immediate hardware reset: can restore service after a hard lockup, but loses volatile state and may repeat indefinitely if boot reaches the same failure.
- No reset: may preserve a debuggable console for a human, but leaves a stuck host unavailable until intervention.
A nonzero shutdown timeout (-x or the corresponding rc setting) can leave the watchdog armed during reboot to reset a machine whose software shutdown hangs. That behavior must be tested with the actual init path and hardware; a normal reboot that exits watchdogd is not enough. The rc manual distinguishes ordinary service stops and return to single-user mode from the system shutdown path. Do not assume the shutdown timeout is applied whenever an administrator runs service watchdogd stop.
Plan for repeated-failure control. A watchdog reset is not a repair. If a kernel module, storage device, filesystem, or configuration consistently breaks during startup, the watchdog can produce a boot loop that makes remote diagnosis impossible. Keep a boot environment or alternate kernel where applicable, preserve console access, document the firmware reset policy, and confirm how the hosting platform reports repeated resets. In a clustered system, fencing and quorum policy must be designed separately; a watchdog reboot is not proof that the peer has safely taken over.
Validate the recovery path without endangering production
Testing should prove each layer separately:
- Configuration: the intended
watchdogd_enable, flags, timeout override, and shutdown timeout are active after reboot. - Check behavior: a known-success case refreshes the timer, a controlled failure is logged, and execution cannot hang beyond the chosen threshold.
- Implementation: logs or hardware management telemetry identify whether the host is using the intended timer, and the requested timeout is supported.
- Recovery: in a lab or redundant maintenance window, the configured event leads to the expected reset or diagnostic action, the machine boots, filesystems recover, services return, and monitoring observes the event.
- Negative control: disabling or stopping the service during a safe test produces the documented shutdown behavior rather than an unexpected reset.
Correlate watchdog messages with kernel logs, the BMC event log, boot timestamps, crash dumps, and service-health probes. A reset counter or a new uptime proves only that a restart occurred. It does not establish that applications recovered, data remained consistent, or failover completed. Alert on repeated boots, missing heartbeat after restart, crash-dump extraction failure, and a watchdog that is enabled in configuration but absent from observed runtime state.
The reliable use of a watchdog is deliberately conservative: make the check match a recoverable host-level failure, make the deadline tolerant of expected pauses, and exercise the end-to-end reset path where loss of the machine is acceptable. Keep service-level probes, backups, replication, and human escalation independent. That is what turns a timer into a controlled recovery mechanism rather than an unexplained source of reboots.
Related:
- Diagnosing a FreeBSD Kernel Panic from a Crash Dump
- FreeBSD System Logging in Production: syslogd, newsyslog, and Rotation
Sources: