Skip to content
LinuxHow-To Published Updated 8 min readViews unavailable

Linux inotify Overflow Recovery: Reconcile Events with Filesystem State

Build resilient inotify watchers that parse variable-length records, recover from queue overflow, reconcile renames, and rebuild directory-tree watches.

Linux inotify reports filesystem events through a file descriptor that can be monitored by poll or epoll. It is useful for caches, file indexers, development tools, and daemons that want to react promptly when names or metadata change. It is not a durable event log and cannot be treated as an exact count of every filesystem operation.

Events may be coalesced, the bounded queue can overflow, names can change before the consumer processes an event, and the API does not recursively watch directory trees. A reliable watcher uses events as hints to reconcile an authoritative filesystem view, not as the only record of state.

Create one instance and read complete event records

inotify_init1() can create a nonblocking descriptor with close-on-exec behavior. Watches are added to that instance, and reads return a packed sequence of struct inotify_event records. Each record has a fixed header plus an optional, padded name field whose length is given by len.

#define _GNU_SOURCE
#include <errno.h>
#include <stddef.h>
#include <sys/inotify.h>
#include <unistd.h>

int drain_inotify(int fd) {
    union {
        struct inotify_event align;
        char bytes[64 * 1024];
    } buffer;

    for (;;) {
        ssize_t n = read(fd, buffer.bytes, sizeof(buffer.bytes));
        if (n == -1 && errno == EINTR)
            continue;
        if (n == -1 && errno == EAGAIN)
            return 0;
        if (n == 0) {
            errno = EIO;
            return -1;
        }
        if (n == -1)
            return -1;

        size_t offset = 0;
        while (offset < (size_t)n) {
            if ((size_t)n - offset < sizeof(struct inotify_event)) {
                errno = EIO;
                return -1;
            }
            const struct inotify_event *event =
                (const struct inotify_event *)(buffer.bytes + offset);
            size_t available = (size_t)n - offset;
            if ((size_t)event->len > available - sizeof(*event)) {
                errno = EIO;
                return -1;
            }
            size_t record_size = sizeof(*event) + (size_t)event->len;
            if (event->mask & IN_Q_OVERFLOW) {
                /* Invalidate derived state and schedule reconciliation. */
            } else {
                /* Resolve wd and process the event as a state-change hint. */
            }
            offset += record_size;
        }
    }
}

The union provides alignment suitable for the event header. A production parser should also handle IN_IGNORED, IN_UNMOUNT, unknown mask bits, and watch-descriptor reuse, and should copy a name if it must retain it after the read buffer is reused. The event’s len includes space for the name and padding; do not use strlen() on untrusted bytes without first checking the record bounds.

Set IN_NONBLOCK at descriptor creation when the reader drains from an event loop. Under edge-triggered epoll, consume records until EAGAIN so the queue is empty before waiting for another edge. A read can contain multiple records, and one logical operation can produce several events.

Treat queue overflow as loss of synchronization

IN_Q_OVERFLOW has watch descriptor -1 and indicates that events beyond the queue limit were dropped. The kernel generates the overflow record, but it cannot tell the application which dropped events matter to its model. The correct response is not to retry the previous read or guess the missing filenames; the application’s cache or index has become potentially inconsistent.

Mark affected state dirty, stop making decisions that assume the event history is complete, and schedule a reconciliation scan from the authoritative filesystem. Depending on the application, it may be sufficient to rescan one directory or necessary to rebuild the entire watch tree and cache. Do not clear the dirty state until the scan has completed and new watches are installed according to the application’s race policy.

The queue size is controlled through /proc/sys/fs/inotify/max_queued_events when an instance is created. Per-user instance and watch limits are controlled separately. Raising limits can absorb larger bursts, but it cannot eliminate overload or guarantee lossless history. Size the queue for expected burst and drain latency, measure overflow frequency, and retain a recovery path even when the limit is generous.

Event coalescing also means that repeated identical events can collapse before the application reads them. Inotify is therefore unsuitable as an exact event counter. If every state transition must be durable and ordered, use an application-owned journal or a filesystem protocol that records those transitions; use inotify only to prompt reconciliation.

Pair rename events cautiously

Renames within watched directories commonly produce IN_MOVED_FROM and IN_MOVED_TO records with the same cookie. Events from one inotify descriptor form an ordered queue, but the matching pair is not guaranteed to be adjacent or inserted atomically. Other events may appear between them, and a move out of the watched tree may have no corresponding IN_MOVED_TO event inside it.

Keep a bounded table of pending move cookies with an expiry policy rather than assuming the next event completes the pair. If the pair does not arrive, interpret the object as removed from the observed tree or trigger a reconciliation scan. A watch descriptor can be removed and later reused, so do not retain a stale mapping from a numeric watch ID to a path indefinitely.

The event identifies a watch descriptor and, for directory events, a name relative to that watched directory. It does not guarantee the named object still exists when the event is processed. Another process can rename or delete it immediately. Reopen relative to a trusted directory descriptor and verify the current object if the application needs to act on it.

Atomic file replacement often uses a temporary file and rename, which can appear as move events rather than a simple sequence of writes to the final pathname. A watcher should inspect the final directory state rather than assume that every update emits IN_MODIFY on the name it originally watched.

Build recursive coverage as a reconciliation problem

Inotify watches are not recursive. Each directory in a tree requires its own watch. When a new directory appears or an existing directory is moved into the tree, the watcher must add the watch and scan its current contents. Files may be created inside before that watch is installed, so the event stream alone cannot guarantee that the watcher saw the entire subtree.

A practical strategy is to maintain a versioned inventory: observe a directory event, add a watch to the new directory, scan it, and reconcile the scan result against events already queued. Another approach is to periodically rescan and use inotify only to trigger earlier scans. For high-churn trees, designing the data model to tolerate duplicate discoveries and idempotent updates is simpler than attempting to infer a perfect operation history.

Use IN_ONLYDIR when a watch must target a directory, and IN_MASK_CREATE when a new watch must not silently change an existing watch on the same inode. Multiple pathnames can refer to the same filesystem object, and adding a watch without this flag can replace the mask associated with that object. Keep a mapping from watch descriptor to current object identity and path context, and invalidate it on IN_IGNORED.

Know what inotify cannot observe

Inotify reports events caused through the filesystem API on the local kernel. It does not report remote changes performed on a network filesystem server, and pseudo-filesystems such as /proc, /sys, and /dev/pts are not monitorable in the ordinary way. Changes made through mmap() and related synchronization or unmapping operations are not reported as ordinary file modifications by this interface.

The API does not tell the watcher which user or process caused an event. A service cannot use an inotify event alone as an authorization decision or as proof that a particular client made a change. If remote state matters, combine inotify with polling, server-side notifications, or an application protocol that conveys identity and version information.

Mounting another filesystem over a watched directory can hide the contents below the mount without generating the event a naive subtree cache expects. Monitor mount topology separately where that affects correctness, and reconcile after mount or unmount operations. A pathname cache may also become stale after ancestor renames; a watch descriptor names the object, while your cached path is application-maintained state.

Keep the event loop and recovery path bounded

The event descriptor integrates naturally with epoll, but event processing should not perform unbounded scans inside the read callback. Copy or summarize the event, mark a region dirty, and schedule a bounded reconciliation worker. Limit pending work and coalesce redundant scan requests so a burst does not create an unbounded userspace queue on top of the kernel queue.

Track queue overflow count, event-to-reconciliation latency, number of active watches, watch-add failures, pending rename-cookie age, and scan duration. If inotify_add_watch() fails because the per-user watch limit is exhausted, do not report recursive coverage as complete. Either reduce scope, request an operational limit change, use polling, or fail the feature explicitly.

Use a single owner for the read stream when possible. Multiple readers on one descriptor complicate ordering and parser state. If several components need notifications, have one reader fan out normalized state-change hints through an internal queue, and keep filesystem reconciliation as the shared source of truth.

Test lost events and path races deliberately

Test burst loads large enough to overflow the configured queue, repeated identical operations that may coalesce, renames between watched and unwatched directories, watch removal and reuse, directory creation followed by immediate population, mount overlays, and a slow reconciliation worker. Verify that after overflow the application repairs its inventory rather than merely logging a warning.

Test under concurrent writers and include mmap-based updates, remote filesystems if used, and pseudo-filesystem paths that are known not to be supported. Compare the watcher’s final model with a fresh filesystem scan. The pass condition is eventual correct state after supported events and explicit recovery after loss, not a one-to-one event count.

Inotify is an efficient wakeup and change-hint mechanism, but it is intentionally not a durable journal. Parse records defensively, understand rename ambiguity and recursive-watch gaps, detect queue overflow, and make a rescan-based recovery path part of the original design.

Related:

Sources:

Comments