Linux FUSE Architecture: Request Lifetimes, Cache Modes, and Daemon Backpressure
Trace Linux FUSE requests between VFS and a userspace daemon, then design cache semantics, cancellation, backpressure, and shutdown around their real lifetimes.
Filesystem in Userspace (FUSE) lets an ordinary userspace process provide filesystem operations through the Linux VFS. The kernel module receives operations such as lookup, getattr, open, read, and write, then exchanges requests and replies with a filesystem daemon through a FUSE connection. Applications continue to use normal path and file APIs; they do not call a private FUSE library for every read.
That convenience creates a real service boundary in the middle of filesystem I/O. A blocked daemon can stall processes that touch the mount, cache choices change which operations reach userspace, and an unmount does not necessarily end a connection whose references remain alive. Production reliability depends on understanding request identity, negotiated protocol features, cache ownership, daemon queues, and teardown.
The kernel is a client and the daemon is a server
The low-level FUSE device interface is a request/reply protocol. A userspace daemon obtains a file descriptor for /dev/fuse; that descriptor is associated with a particular filesystem connection. The kernel queues operations on the connection, and the daemon reads requests and writes matching replies. Libraries such as libfuse provide high- and low-level APIs over this interface, but the request ownership and kernel semantics remain important even when the library hides the wire structures.
Each request has an input header containing a length, opcode, unique request ID, node ID, and caller credentials. The operation payload depends on the opcode. A reply starts with an output header that includes the matching unique ID and an error value. The daemon must validate that it has received a complete request, return a reply with the proper identity and length, and avoid treating an error reply as if it also had a success payload.
Node IDs identify filesystem objects in the connection, while file handles returned by open operations are opaque state for later operations. They are not interchangeable: the kernel can ask for attributes by node ID, while reads and releases often refer to an opened handle and an offset. At the low level, a filesystem must preserve the node and generation rules for its connection and release per-open state only after the protocol says the handle is no longer in use. A high-level API may manage some of this bookkeeping, but it cannot repair inconsistent backend semantics.
Negotiate the protocol instead of assuming a version
The first FUSE_INIT request negotiates the protocol major and minor versions and capabilities such as read-ahead and maximum write sizes. A daemon must remember what was negotiated and only use fields or behavior supported by that connection. The protocol evolves over time, and the low-level manual page explicitly warns that its operation list is incomplete and version-sensitive. A kernel header on the build machine is not proof that a deployed kernel supports every flag.
Treat capabilities as a negotiated contract. Use the library’s supported-feature checks where available, probe behavior on the minimum supported kernel, and keep a fallback for unsupported optional operations. Never parse version-specific structures using hard-coded offsets copied from a different release. Build tests that cover the oldest and newest supported kernel and library pair.
Decide who owns cached truth
FUSE has distinct I/O modes. Direct I/O bypasses the page cache for reads and writes, with read-ahead disabled and shared mmap disabled by default. Cached operation can use page-cache reads and read-ahead; the default cached write-through mode sends writes to userspace while maintaining cache state. Writeback cache can let write() complete after data reaches cache, with dirty pages sent later by background writeback, reclaim, fsync(), close, or final unmap-related activity.
Writeback cache is not a free performance switch. It assumes changes to a file go through the FUSE kernel module so the kernel can maintain size and timestamp attributes. It is generally unsuitable when a remote or external writer can modify the same backing data without cache invalidation. Even an O_WRONLY partial-page write can cause a read request so the kernel can preserve the untouched part of a cached page. Validate write ordering and durability with the actual backing store, not just the return time of write().
Metadata and directory-entry timeouts are also consistency policy. Longer cache lifetimes reduce daemon round trips but can make external changes less visible; short lifetimes improve freshness at the cost of more lookup and getattr work. If the backend is mutable outside the mount, define how notifications invalidate cached entries and attributes, what happens when invalidation is delayed, and whether stale reads are acceptable. Keep cache settings tied to a documented consistency model.
Bound daemon concurrency and make congestion observable
FUSE background requests have connection-level limits. When the configured background request limit is reached, later operations can block. A congestion threshold lets the kernel alter behavior such as asynchronous read-ahead and non-synchronous writeback while the daemon is overloaded. These controls protect the system from unlimited work in flight, but a poorly sized daemon queue can turn a throughput increase into higher tail latency and memory use.
When the FUSE control filesystem is mounted, each connection exposes operational state such as waiting, max_background, congestion_threshold, and abort. waiting counts requests waiting to reach userspace or being processed by the daemon. A nonzero waiting count with no filesystem activity is a strong sign of a hung or deadlocked connection. The abort control can terminate a stuck connection and cause waiting and new operations to fail; it is a recovery action, not an ordinary health check.
Size the daemon’s workers and queues from measured backend latency and concurrency. Record queue depth, time from kernel request receipt to daemon dispatch, backend duration, reply duration, error count, and canceled work. Bound both queued requests and any per-request buffers. When the backend is slow, return an explicit filesystem error or apply a bounded wait policy; do not allow unbounded worker creation to conceal overload.
Cancellation does not erase a request
If a caller is interrupted after its request has reached userspace, the kernel can queue a FUSE_INTERRUPT referencing the original unique request ID. Receiving the interrupt message does not itself cancel the operation. The daemon may ignore it, or honor it by completing the original request with an interruption result. Races are possible: the original operation might already have completed, or the interrupt might arrive before the worker has found the request.
Maintain a request table keyed by the protocol request identity for operations that can be canceled. Make cancellation idempotent, coordinate it with the backend, and ensure exactly one valid completion path owns the reply. Do not free the request state merely because an interrupt arrived; the kernel still expects the original request to finish according to protocol. Test cancellation during queue wait, backend I/O, response serialization, and shutdown.
Unmount and daemon shutdown are separate events
A FUSE connection exists until the daemon dies or the filesystem is unmounted, but a lazy detach does not immediately destroy the connection while references to the filesystem remain. Processes can retain open files or working directories after the mount point has disappeared from the namespace. Therefore, “the mount is gone” is not a sufficient proof that every request has drained or that the daemon can discard all state.
Use an orderly shutdown sequence: stop accepting new application work, unmount through the supported helper or service manager, stop new backend scheduling, complete or fail in-flight requests, and wait for worker ownership to end before destroying connection state. Make daemon death visible to supervising services and applications. Recovery tooling should distinguish a detached mount, an unresponsive daemon, a disconnected FUSE connection, and a backend outage.
Validate behavior through both interfaces
On a Linux test host, inspect the mount with findmnt -T /mount/point, and inspect /sys/fs/fuse/connections only when the FUSE control filesystem is mounted. Compare syscall traces from the application with request counters or daemon logs. This distinguishes a cache hit from a request that actually crossed into the daemon. A missing daemon log entry does not by itself mean the application performed no filesystem operation; the kernel may have served cached metadata or data.
Acceptance tests should cover lookup storms, metadata changes made through and outside the mount, direct and cached reads, partial writes, fsync(), delayed backend errors, full worker queues, interrupted requests, daemon restart, forced abort, lazy unmount with an open file, and shutdown while reads are active. Verify data checksums and ordering, then measure p50/p99 latency, throughput, CPU, memory, queue depth, and recovery time.
FUSE is most useful when a userspace implementation boundary is worth the kernel-to-daemon round trip and its additional failure mode. It can provide a clean filesystem interface without putting the full filesystem implementation in kernel space. It does not remove filesystem semantics; it moves responsibility for cache validity, request completion, backend failure, and lifecycle into a process that must be operated like any other critical service.
Related:
- The Linux Virtual File System: One Interface, Many Filesystems
- Setting Up Bind Mounts and Overlay Filesystems
Sources: