Linux userfaultfd: Delegating Page-Fault Resolution to a Userspace Manager
Understand Linux userfaultfd registration, missing and write-protect faults, feature negotiation, manager synchronization, and failure handling.
Linux normally resolves virtual-memory faults inside the kernel. userfaultfd lets a userspace process register selected virtual-memory ranges and receive notifications when accesses to those ranges fault. A manager can then provide page contents, map an existing page, or apply another supported resolution before the faulting thread continues. This is the foundation for mechanisms such as demand paging from a remote store and QEMU/KVM post-copy migration, but it is not a general replacement for mmap, mprotect, or signal handlers.
The design separates the thread that touches a page from the thread that decides how to resolve it. That separation is powerful and hazardous: a faulting thread may block indefinitely if the manager crashes, stalls, loses its backing data, or mishandles a concurrent address-space change.
Create, negotiate, then register
The API has three conceptual stages. Create a userfault file descriptor, initialize its protocol with UFFDIO_API, then register one or more ranges with UFFDIO_REGISTER. A successful API negotiation returns feature and ioctl bitmasks. A correct program checks those returned capabilities instead of assuming that a kernel supports a feature just because its headers define the constant.
The userfaultfd(2) system call can create descriptors that handle user-mode faults with UFFD_USER_MODE_ONLY. Kernel-mode fault handling has additional permission controls. The kernel also documents /dev/userfaultfd plus USERFAULTFD_IOC_NEW, whose access is controlled by device-node permissions. These creation routes have different security properties; granting a process CAP_SYS_PTRACE solely to obtain userfaultfd access may give it unrelated authority.
Registrations describe both a virtual-memory range and the fault modes to intercept. Missing-page mode lets a manager populate absent pages. Write-protect mode reports writes to protected ranges. Minor-fault mode applies to supported backing types where a page exists in page cache but has not yet been mapped into the process. Support for memory types and modes is kernel-dependent and must be discovered during negotiation.
The manager is part of the application’s liveness path
For a missing page, the manager can use UFFDIO_COPY to copy bytes into the faulting range or UFFDIO_ZEROPAGE to provide a zero-filled page. For supported minor faults, UFFDIO_CONTINUE can continue with an existing page. These operations coordinate visibility so a reader does not observe a partially populated page. By default, successful resolution wakes blocked faulting threads; a DONTWAKE mode allows a manager to coordinate wake-up separately.
That means the manager needs bounded queues, explicit timeouts for backing-store operations, and a policy for unrecoverable faults. An unbounded remote fetch can turn one slow network into many blocked application threads. Returning a zero page when data is missing is not a neutral fallback: it changes program state. Fail closed, fail open, or synthesize data only when the application semantics explicitly permit it.
Address-space changes require synchronization
When a descriptor is used to monitor a process that does not cooperate with its manager, the registered virtual range can change while work is in flight. The kernel can report events for fork, mremap, madvise removal, and munmap when the relevant features are requested and supported. The manager must process these events and synchronize them with resolution ioctls. For example, attempting to populate a range that has been unmapped or changed concurrently must be treated as a normal race to handle, not as a reason to retry blindly.
Keep the address-space owner and manager lifecycle explicit. If the manager is in another process, passing the descriptor over a Unix-domain socket is supported. Define what happens when either process exits, how outstanding faults are drained, and whether children inherit tracking. A descriptor alone does not guarantee that the manager’s view of mappings remains current.
Minimal implementation sequence
A production implementation should follow this outline and check every return value:
create userfaultfd with the least privilege mode required
issue UFFDIO_API and inspect returned feature/ioctl masks
register only the intended mapping range and fault mode
poll/read fault messages on a dedicated manager path
validate the message and backing-store bounds
resolve with the ioctl appropriate to that fault mode
handle exit, unmap, remap, and manager-shutdown races
unregister ranges and close the descriptor during teardown
This is intentionally not a copy-and-paste C program. The kernel UAPI structures and constants are architecture headers, and usable features vary by running kernel. Start from the current kernel documentation and man page example, compile against the target system’s headers, and test the exact backing type and fault mode that the application will use.
Security and operational testing
Userfaultfd exists partly because user-controlled fault handling can be relevant to kernel attack surfaces. Keep permissions narrow, validate every address and length, and never accept an untrusted manager’s page contents without an integrity and authorization design. A handler that processes attacker-controlled descriptors should also bound message parsing and ensure that it cannot resolve a fault into an unrelated mapping.
Test both correctness and failure: a normal missing-page read, manager termination, backing-store timeout, concurrent unmap, fork, range removal, and unsupported-feature negotiation. Assert that the faulting workload either makes progress with correct bytes or terminates through a documented error path. If the only success test is that the userfaultfd descriptor was created, the important failure modes remain untested.
Related:
- Understanding the Linux Page Cache and How Writeback Actually Works
- Reading /proc and /sys: The Kernel’s Window into Userspace
Sources: