Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux epoll Readiness Semantics: Edge-Triggered I/O, Lifetimes, and Fairness

Build correct Linux epoll loops by handling readiness, edge-triggered drains, duplicated descriptors, one-shot rearming, and backpressure explicitly.

epoll is Linux’s readiness notification interface for file descriptors. An application registers an interest list, waits for events, and performs ordinary nonblocking reads or writes when an operation is reported ready. It is not a completion queue: an event does not mean that a full request finished, only that an operation may make progress without blocking at the time the readiness condition is observed.

This distinction governs the whole design. Correctness depends on the trigger mode, draining behavior, descriptor identity, and queue ownership. A loop that works under a small test can stall under edge-triggered traffic, spin on a permanently writable socket, or keep receiving events after a descriptor was closed because a duplicate still refers to the same open file description.

The interest list and ready list

An epoll instance has an interest list describing watched file descriptors and an internally maintained ready list of descriptors that currently satisfy readiness conditions. epoll_ctl() adds, modifies, or removes interest; epoll_wait() retrieves a batch of ready events. This avoids repeatedly submitting a large unchanged descriptor set as with select() or poll(), but it does not make application work free or guarantee constant-time behavior for every workload.

Create instances with epoll_create1(EPOLL_CLOEXEC) so they are not unintentionally inherited across exec. Set watched sockets nonblocking before registering them. Use one owner for each connection’s input buffer, output queue, and close path; the kernel reports readiness, but the application owns protocol framing and lifecycle.

Regular files are usually not useful epoll targets because they are considered ready in ways that do not provide asynchronous disk completion; unsupported descriptor types can make epoll_ctl() fail with EPERM. Use an API designed for the operation instead of interpreting every descriptor as a socket-like event source.

Level-triggered mode is the default

Without EPOLLET, epoll is level-triggered (LT). If unread data remains, the descriptor can continue to be reported ready. That makes LT straightforward: process a bounded amount, return to the loop, and expect another event while the condition remains true. LT is often the better initial choice because it makes backpressure and fairness easier to implement.

An LT server still needs nonblocking operations, bounded per-connection work, and a policy for EPOLLOUT. A TCP socket is commonly writable most of the time. Registering write interest permanently can produce a busy loop. Enable EPOLLOUT when an application has queued bytes, write until the queue is empty or the call returns EAGAIN, then remove write interest. If only part of the buffer was accepted, retain the remainder and wait for another writable notification.

Readiness can change between epoll_wait() and the subsequent I/O operation, especially when multiple threads share state. Nonblocking I/O makes that race safe: EAGAIN means there is no progress now, not that the connection is broken. Serialize ownership or use a deliberate multi-worker scheme so two handlers do not corrupt one protocol buffer.

Edge-triggered mode requires draining

EPOLLET asks for edge-triggered notifications. If an application reads only part of the available data and then waits for another event, the remaining bytes may stay buffered without producing a new edge. The documented rule is to use nonblocking descriptors and wait again only after read() or write() returns EAGAIN or EWOULDBLOCK.

enum drain_result {
	DRAINED_TO_EAGAIN,
	PEER_CLOSED,
	DRAIN_FAILED,
};

static enum drain_result drain_socket(int fd)
{
	char buffer[16384];

	for (;;) {
		ssize_t count = recv(fd, buffer, sizeof(buffer), 0);
		if (count > 0) {
			consume_protocol_bytes(fd, buffer, (size_t)count);
			continue;
		}
		if (count == 0)
			return PEER_CLOSED;
		if (errno == EINTR)
			continue;
		if (errno == EAGAIN || errno == EWOULDBLOCK)
			return DRAINED_TO_EAGAIN;
		return DRAIN_FAILED;
	}
}

This is a Linux/POSIX C method fragment that assumes the usual socket, errno, and integer headers plus an application-defined protocol consumer. Call it when an edge-triggered read event arrives; close or report an error according to the returned status. Do not return to an indefinite epoll_wait() merely because a convenient byte or request quota was reached while data remains readable. If fairness requires bounded work, put that connection on an application-ready queue and continue it without waiting for a fresh edge.

Write handling follows the same principle: keep writing queued output until the queue is empty or the nonblocking send() returns EAGAIN. A short successful write is progress, not completion. Monitor write readiness only while bytes remain. Do not spin on EAGAIN; wait for the appropriate readiness event.

Treat event flags as hints with protocol meaning

EPOLLERR and EPOLLHUP can be reported even if they were not requested in the event mask. They do not mean the application may skip buffered input: a stream may have readable bytes before EOF, so drain it and then process closure. EPOLLRDHUP can be requested to notice a peer’s write-side shutdown on a stream socket. Inspect socket errors and the result of the actual I/O operation instead of treating all flags as interchangeable “disconnect” events.

EPOLLONESHOT disables the registered descriptor after an event is delivered, until the application rearms it with EPOLL_CTL_MOD. This can help transfer one connection to one worker at a time, but it is not a complete locking protocol. Rearm only after shared connection state is ready for another handler, and handle the case where the descriptor becomes ready while it is disabled. If several threads wait on one epoll fd, event delivery and application ownership must still be designed explicitly.

Understand descriptor identity and teardown

The registration key is the combination of the file descriptor number and its open file description. dup(), fork(), or fcntl(F_DUPFD) can create another descriptor referring to the same open file description. Closing one descriptor does not necessarily remove the epoll registration while another duplicate remains open; events can still be reported through that registration.

If a descriptor is duplicated or shared across workers, establish who removes the interest and who closes every reference. Use EPOLL_CTL_DEL before the final owner hands a descriptor to an unrelated lifecycle when stale events would be dangerous. Store a generation or stable connection object in event.data.ptr/u64 and validate it before acting, rather than trusting a recycled integer fd as a permanent identity. Keep the pointed-to event data alive until no returned batch or worker can use it.

Shutdown is a protocol: stop accepting new work, remove or disable interests as appropriate, prevent queued handlers from using freed connection state, drain or cancel outstanding application work, then close descriptors and the epoll fd. Handle errors from epoll_ctl() such as ENOENT during idempotent cleanup without masking unrelated failures.

Bound queues and keep the loop fair

Epoll reports readiness, not application capacity. If a handler continues reading while downstream work is full, it can move unbounded data from the kernel socket buffer into process memory. Couple read interest to bounded application queues or stop reading until consumers recover. If a queue is paused in edge-triggered mode, arrange an explicit recheck or local reschedule when capacity returns; waiting for a new network edge can leave already-buffered bytes stranded.

Likewise, bound the number of bytes, messages, or CPU time one connection consumes per scheduling turn. In LT mode, remaining readiness will be reported again. In ET mode, preserve a user-space ready queue for connections that have not reached EAGAIN. This prevents a hot connection from starving quiet ones without violating the edge-triggered drain contract.

epoll_wait() can return fewer events than requested; process the returned count only. A signal can interrupt the wait with EINTR, which should be retried or handled according to shutdown state. Its millisecond timeout is rounded up to clock granularity and actual sleep can overrun because the thread still needs CPU scheduling. Do not treat a timeout as an exact deadline; use a monotonic clock for application deadlines and recalculate remaining time.

The per-user watch limit is a resource boundary, not a tuning target to raise blindly. Monitor descriptor counts, registrations, memory use, and connection churn. Reuse a long-lived epoll instance where appropriate, remove closed connections, and size admission control to the service’s actual memory budget.

Validate the loop under races and load

Test partial reads and writes, multiple frames in one read, EOF with unread data, resets, EINTR, EAGAIN, half-close, duplicate descriptors, concurrent close, one-shot rearming, and shutdown while an event batch is being processed. Add a test where the peer sends a burst larger than one application fairness quota; verify ET mode continues draining from a local queue without waiting for a nonexistent new edge.

Measure accepted connections, active watches, events returned, bytes per readiness event, EAGAIN rates, queue depth, event-loop CPU, context switches, allocation pressure, and tail latency. Compare LT and ET with the same workload and bounded queues. Include connection churn and overload, not just many idle sockets. A higher connection count at startup does not demonstrate that the service can sustain useful work or shut down cleanly.

Use epoll when readiness-based event handling matches the workload and the team can preserve the descriptor and queue invariants. A sound LT design is better than an ET loop that occasionally sleeps with unread data. More advanced mechanisms such as io_uring have different completion semantics; they are not interchangeable merely because both can participate in an event loop.

Related:

Sources:

Comments