Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux timerfd: Deadline Clocks, Missed Periods, and Event-Loop Timers

Choose the right timerfd clock, consume expiration counts correctly, avoid wall-clock surprises, and integrate deadlines with poll or epoll.

Linux timerfd exposes timer expirations through a file descriptor. An event loop can monitor that descriptor with poll, select, or epoll, alongside sockets and other work sources. Unlike a signal-based timer, a timerfd reports the number of expirations since the last successful read, which makes missed periods observable instead of silently compressing them into one handler invocation.

The descriptor does not decide what missed work means. A service still has to choose whether to run every overdue task, coalesce them into one reconciliation pass, skip stale intervals, or mark itself unhealthy. Correct timer design starts by choosing a clock and defining how the application interprets expiration counts.

Choose a clock that matches the deadline

CLOCK_MONOTONIC is a non-settable clock that advances while the system is running and is not stepped when an administrator corrects wall time. Use it for elapsed-time deadlines, timeouts, retries, and periodic maintenance that should not jump when the calendar clock changes. The clock does not count time while the system is suspended.

CLOCK_BOOTTIME is also monotonic, but includes time spent suspended. It is appropriate when a deadline should become overdue while a laptop or embedded device sleeps. It does not itself wake the suspended system. The CLOCK_BOOTTIME_ALARM and CLOCK_REALTIME_ALARM clock variants can wake a suspended system, but setting timers against them requires CAP_WAKE_ALARM.

CLOCK_REALTIME tracks the system wall clock. Use it for calendar deadlines such as a specified absolute date and time, with a policy for clock corrections. A realtime clock can jump forward or backward due to time administration or synchronization. A long duration timeout generally should not be tied to wall time unless that behavior is explicitly wanted.

Arm a timer and consume the expiration count

Create the descriptor with TFD_NONBLOCK and TFD_CLOEXEC when it belongs to an event loop and should not leak across execve(). A periodic timer can be configured with an initial delay and an interval:

#define _GNU_SOURCE
#include <errno.h>
#include <stdint.h>
#include <sys/timerfd.h>
#include <unistd.h>

int arm_periodic_timer(void) {
    int timer_fd = timerfd_create(CLOCK_MONOTONIC,
                                  TFD_NONBLOCK | TFD_CLOEXEC);
    if (timer_fd == -1)
        return -1;

    struct itimerspec spec = {
        .it_value = { .tv_sec = 5, .tv_nsec = 0 },
        .it_interval = { .tv_sec = 5, .tv_nsec = 0 }
    };
    if (timerfd_settime(timer_fd, 0, &spec, NULL) == -1) {
        int saved_errno = errno;
        close(timer_fd);
        errno = saved_errno;
        return -1;
    }
    return timer_fd;
}

int consume_timer_expirations(int timer_fd, uint64_t *expirations) {
    ssize_t n;
    do {
        n = read(timer_fd, expirations, sizeof(*expirations));
    } while (n == -1 && errno == EINTR);

    if (n == (ssize_t)sizeof(*expirations))
        return 0;
    if (n == -1 && errno == EAGAIN)
        return 1; /* Readiness was stale or another reader consumed it. */
    if (n >= 0)
        errno = EIO; /* Unexpected short or zero-byte read in this clock model. */
    return -1;
}

The example returns ownership of the new descriptor to its caller and makes cleanup the caller’s responsibility. A read returns an unsigned 64-bit expiration count in host byte order and requires an eight-byte buffer. The count may be greater than one if the event loop was busy or the machine was descheduled across several timer periods.

Do not interpret the count as a list of timestamps or as proof that each corresponding job ran. If the timer represents “refresh current state,” one read with count 7 may reasonably trigger one refresh and record six coalesced periods. If it represents seven billing intervals that must each be processed, the application needs durable period identities and a policy for replay, not merely a timerfd counter.

Account for delayed and periodic work

A periodic timer continues on its configured expiration cadence while the process is busy. After a long pause, the expiration count exposes how many periods became due. For a fixed-rate scheduler, use the count to advance the logical deadline by the number of intervals and decide whether to catch up or skip. Re-arming a relative timer only after work completes instead produces a “delay after completion” schedule whose phase drifts with task duration.

For a stable fixed-rate schedule, maintain an absolute next deadline on a monotonic clock. After processing a wake, advance the deadline by an integral number of periods according to the chosen catch-up rule, then arm with TFD_TIMER_ABSTIME. If work takes longer than one period, decide whether to run each missed unit, coalesce, or move to the first future deadline. Avoid computing a new target as “now plus period” when the requirement is tied to the original schedule; that converts scheduling delay into permanent drift.

If the event loop uses edge-triggered epoll, keep the timer nonblocking and read the expiration counter when signaled. Timerfd has a single counter-like read result rather than an unbounded sequence of individual records; after a successful read the accumulated count is consumed. With multiple readers, one may consume the count before another runs, so assign a single timer owner or synchronize the work contract.

Absolute timers and wall-clock changes

TFD_TIMER_ABSTIME interprets it_value as an absolute time on the descriptor’s selected clock. This is useful for durable deadlines computed against the same clock domain. Do not pass a wall-clock timestamp to a monotonic timer or compare values from different clocks; the units are both timespecs, but their epochs and adjustment behavior differ.

For an absolute CLOCK_REALTIME or CLOCK_REALTIME_ALARM timer, TFD_TIMER_CANCEL_ON_SET asks the kernel to report a discontinuous wall-clock change as ECANCELED on a read. Treat that result as “re-evaluate the calendar deadline,” not as an ordinary timer expiration. The application should recompute the next wall-clock target and rearm according to its calendar policy.

Without cancellation, a negative discontinuous realtime change can create an unusual case where a read is awakened but returns zero bytes. Never assume that every readable timerfd necessarily yields an eight-byte count in every realtime clock-change scenario. Handle ECANCELED, EAGAIN, EINTR, short or zero reads, and terminal descriptor errors as distinct states.

An absolute timer set in the past expires immediately. Validate conversion from civil time to timespec, including timezone policy, daylight-saving transitions, leap-second behavior provided by the system clock, and overflow. If a job is “at the next local 09:00,” recompute from the timezone database after clock changes instead of adding a fixed 24 hours to a prior epoch value.

Integrate the descriptor lifecycle with the event loop

Register the timerfd with the same ownership and close discipline as sockets. Under level-triggered readiness, an expired count keeps the descriptor readable until a read consumes it. Under edge-triggered operation, read after each reported expiration and return to the loop only when the counter is consumed or a nonblocking read yields EAGAIN.

Do not run expensive maintenance while holding a global event-loop lock. Convert an expiration into a state transition or queued work item, and let the appropriate worker handle it. A timer read and a task completion are separate events: record the deadline, start time, completion, and outcome if operational correctness depends on punctuality.

Closing the last descriptor reference disarms and releases the timer object. A descriptor inherited through fork() refers to the same underlying timer; reads in parent and child consume shared expiration state. Set TFD_CLOEXEC when inheritance is not part of the protocol, and close duplicate references when ownership ends. A leaked timer descriptor can keep the object active and make shutdown behavior surprising.

Compare timerfd with other scheduling primitives

Use timerfd when timer readiness belongs in a descriptor-based event loop. Use clock_nanosleep() when one thread can block until a deadline and no multiplexing is needed. Use POSIX timers or a signal-wait design when their notification and threading semantics are a better fit. Use a higher-level event library when it provides the required clock and recurrence semantics without hiding a critical deadline policy.

Timerfd does not make a process real-time, reserve CPU, or guarantee that work begins exactly at the expiry instant. Scheduler latency, CPU pressure, virtualization, power states, and event-loop load all affect when userspace handles readiness. For strict latency requirements, measure actual deadline miss distributions and configure the system under an explicit real-time policy; an accurate timer alone is not a real-time guarantee.

Test deadline policy as well as wakeup

Test one-shot and periodic timers, multiple expirations while the loop is paused, disarm and rearm, absolute deadlines in the past, wall-clock jumps for realtime cancellation, suspend/resume behavior for the chosen clock, descriptor inheritance, and resource exhaustion. Verify what the application does with a count greater than one, not merely that a read succeeded.

Log the clock ID, scheduled deadline, expiration count, handling delay, missed-period policy, and resulting work state. Do not collapse “timer fired,” “job started,” and “job completed” into one metric. When a task misses a deadline, the distinction is essential for determining whether the clock, event loop, scheduler, or workload caused the delay.

Timerfd makes Linux deadlines composable with ordinary I/O, but correctness rests on a deliberate clock domain and a defined catch-up policy. Use monotonic time for elapsed durations, realtime for calendar deadlines, read and interpret the expiration count, and keep timer lifecycle tied to the owning event loop.

Related:

Sources:

Comments