Skip to content
LinuxDeep Dive Published Updated 9 min readViews unavailable

Linux eventfd: Counter Semantics, Pollability, and Safe Notification

Use eventfd as a pollable kernel counter for Linux event loops, understand semaphore reads and nonblocking writes, and keep notification separate from payload.

Linux eventfd provides a file descriptor around a kernel-maintained 64-bit counter. A process can increment the counter with a write, consume it with a read, and monitor readiness through poll or epoll. This makes eventfd useful when a worker needs to wake an event loop, a kernel subsystem needs to notify userspace, or a program needs a lightweight signal that integrates with other file descriptors.

An eventfd is a notification mechanism, not a message queue. It has no per-event payload, and multiple writes can accumulate into one counter value. If the application needs to associate each wakeup with distinct work, store that work in a synchronized queue and use eventfd only to make the consumer wake. Treating the counter as both the work queue and the wake signal creates ambiguous ownership and can lose the meaning of individual events.

Counter behavior and reads

Create the descriptor with eventfd, passing an initial unsigned counter value and flags such as EFD_NONBLOCK or EFD_CLOEXEC. A successful read must transfer exactly eight bytes. With ordinary counter semantics, the read returns the current nonzero value and resets the counter to zero. With EFD_SEMAPHORE, each successful read returns one and decrements the counter by one instead.

These behaviors have different work-distribution meaning. Ordinary mode allows one consumer to observe an accumulated count and handle a batch. Semaphore mode lets consumers take one unit at a time, but the counter still contains no identity for the work represented by each unit. Neither mode makes an application queue thread-safe. If several consumers read concurrently, decide which one owns queued items and synchronize that ownership independently.

The counter can hold values up to UINT64_MAX minus one. A write must also be exactly eight bytes, and a value of UINT64_MAX is invalid. A write that would overflow the counter blocks when the descriptor is blocking; with EFD_NONBLOCK, it fails with EAGAIN. In normal event notification designs, increments are small and the consumer drains the descriptor before the counter approaches its maximum.

Readiness and event-loop integration

The eventfd is readable when its counter is nonzero. It is writable when another value can be added without blocking. Since the returned descriptor behaves like a pollable file descriptor, an event loop can watch it alongside sockets, timers, and other sources. EFD_CLOEXEC avoids leaking it across exec, and EFD_NONBLOCK allows the loop to drain or attempt a write without unexpectedly sleeping.

For level-triggered epoll, a nonzero counter remains readable until consumed. For edge-triggered epoll, drain reads until EAGAIN so the loop does not leave unread counter state behind while waiting for a new edge. If the application uses ordinary mode, one successful read clears the accumulated value at that instant. If other writers can increment immediately afterward, the descriptor becomes readable again.

Do not treat one readiness notification as one work item. Readiness means the counter can be read; it does not describe how many separate tasks are queued, whether a task is still valid, or whether a producer has already removed it. The consumer should drain the eventfd and then inspect the synchronized queue or state that defines actual work.

Producer and consumer pattern

A common design combines a mutex-protected queue with an eventfd. Producers enqueue work, then write one to the descriptor. The event loop wakes, reads the counter, and drains the queue under the queue’s synchronization rules. The counter may be greater or less than the queue length because notifications can coalesce, failures can occur, or one wake can cover multiple enqueued items.

#define _GNU_SOURCE
#include <errno.h>
#include <stdint.h>
#include <sys/eventfd.h>
#include <unistd.h>

int notify_event_loop(int event_fd) {
    const uint64_t increment = 1;

    for (;;) {
        ssize_t written = write(event_fd, &increment, sizeof(increment));
        if (written == (ssize_t)sizeof(increment))
            return 0;
        if (written == -1 && errno == EINTR)
            continue;
        if (written == -1 && errno == EAGAIN)
            return 0; /* A notification is already pending; queue state is authoritative. */
        return -1;
    }
}

This helper assumes the application has already published work to a queue and that queue state is the source of truth. Treating EAGAIN as an acceptable coalesced notification is only correct if the consumer will inspect the queue after every readable wake and no required transition depends on counting each eventfd write. If the counter is itself the only state, EAGAIN is not equivalent to success and must be handled according to the application’s protocol.

The consumer must validate read size and errors, handle EINTR, and drain nonblocking descriptors until EAGAIN when using edge-triggered readiness:

int drain_event_fd(int event_fd, uint64_t *total) {
    *total = 0;

    for (;;) {
        uint64_t value = 0;
        ssize_t n = read(event_fd, &value, sizeof(value));
        if (n == (ssize_t)sizeof(value)) {
            *total += value;
            continue;
        }
        if (n == -1 && errno == EINTR)
            continue;
        if (n == -1 && errno == EAGAIN)
            return 0;
        return -1;
    }
}

This drain function is for ordinary counter mode. In semaphore mode, each read returns one, so an implementation may stop after a single token or continue draining while it has work capacity. Do not sum counters without considering overflow in the application’s own total; eventfd’s counter bound does not protect an accumulator with a smaller or signed type.

Nonblocking operation and backpressure

Use EFD_NONBLOCK when the descriptor belongs to an event loop or when a producer must never stall while the counter is saturated. A nonblocking read with a zero counter returns EAGAIN. A nonblocking write that would overflow the counter also returns EAGAIN. The application must decide whether EAGAIN means an existing notification is sufficient, whether it should retry later, or whether the work submission itself failed.

A common safe pattern makes the queue authoritative: enqueue work under a mutex, then attempt to signal. If the signal counter is already nonzero, the loop is already readable and will eventually inspect the queue. If the write fails for an error other than a documented coalescing condition, the producer must not silently claim that the consumer will wake. It can roll back the queue item, use an alternate wakeup path, or mark the component as failed.

A blocking eventfd write can sleep if adding the value would overflow. That may be useful in a narrowly defined producer-consumer protocol but is usually the wrong behavior on an event-loop thread or inside a lock. A blocked producer can hold a mutex needed by the consumer, preventing the read that would free counter capacity. Prefer nonblocking I/O and explicit backpressure.

Descriptor ownership and process boundaries

The descriptor refers to a kernel eventfd object. Duplicating or inheriting the descriptor gives another descriptor reference to the same underlying counter, so all holders can affect the shared state. Coordinate who may read, who may write, and which component is responsible for closing each descriptor. A forked child can inherit descriptors unless close-on-exec and explicit descriptor hygiene are used.

Closing one descriptor does not destroy the eventfd while other references remain open. A child inherits a descriptor across fork(); EFD_CLOEXEC closes it only on a later exec, so explicitly close unwanted references in each process. Closing the last reference removes the object, and a later write through a stale or unrelated descriptor is not a valid notification. Treat descriptor lifetime as part of the event-loop shutdown protocol: stop producers, drain or discard pending work according to policy, wake blocked consumers if needed, and close after no component can use it.

An eventfd can be shared across processes if the descriptor is passed or inherited, but it does not authenticate the meaning of the counter or carry task data. Use a Unix-domain socket, pipe, shared-memory protocol, or another mechanism when the producer must send structured data or establish peer identity. Choose the IPC primitive for the contract, not only for the smallest syscall count.

Eventfd is not a memory-ownership protocol

The eventfd counter provides kernel-managed readiness and atomic counter operations, but it does not define who owns application memory, whether a queue item is visible to a consumer, or how a multi-step state update is synchronized. Protect shared data with a mutex or an explicitly correct atomic protocol. Publish data before signaling, and make the consumer acquire the same synchronization before reading it.

Do not infer queue length from the eventfd count unless the protocol rigorously guarantees one outstanding counter unit per item and handles saturation, consumer batching, cancellation, and shutdown. In most systems, count is a wakeup hint; the queue or state machine is authoritative. This distinction makes coalescing safe and prevents a stale counter from being mistaken for a complete work inventory.

Choosing eventfd, a pipe, or a condition variable

Use eventfd when the process needs a small counter-like notification that integrates with poll or epoll. Use a pipe or socket when the wakeup should carry bytes or when backpressure should be expressed through a byte stream. Use a condition variable when all waiters are coordinated through in-process shared state and no file-descriptor integration is required.

Eventfd can reduce descriptor count compared with a pipe used only for signaling, but that is not a universal performance guarantee. Measure the real workload, include queue contention and event-loop wakeup costs, and avoid introducing a kernel interface where a simpler in-process primitive is sufficient. Document whether one wake means “inspect state,” “consume one token,” or “process this exact count.”

Operational verification

Test the event loop with producers that signal before the consumer begins waiting, multiple producers, bursts that coalesce, EAGAIN, EINTR, descriptor shutdown, and a consumer that is deliberately delayed. Verify that every work item is either completed or explicitly canceled even when one counter read represents many writes. Test both ordinary and semaphore mode if both are supported.

Log queue depth separately from eventfd counter values. Track producer signal failures, wake-to-drain latency, time spent blocked, and work items that remain queued after a wake. If a process appears idle while work is pending, inspect the actual descriptor state and queue protocol rather than assuming eventfd lost an event. A persistent unread nonzero counter should remain pollable; a missed wake is often an ownership, descriptor, or synchronization bug around it.

Eventfd is a compact building block for Linux event loops, but its simplicity can invite an underspecified protocol. Define whether it is a counter or only a wake hint, keep payload and ownership elsewhere, use nonblocking operations where stalls would be dangerous, and make descriptor lifetime explicit.

Related:

Sources:

Comments