Linux memfd_secret: Isolating Sensitive Userspace Memory
Learn what Linux memfd_secret protects, how its file-descriptor and mapping lifecycle works, and where its security guarantees stop.
Linux’s memfd_secret() system call creates an anonymous RAM-backed file that can be mapped into a process. The kernel removes those pages from its normal direct mapping, so ordinary kernel code cannot access the secret bytes through the usual direct-map path. The interface is intended to make selected secrets harder to expose through some classes of kernel compromise; it is not a hardware enclave, a complete process sandbox, or a promise that no privileged attacker can ever recover the data.
This design is useful only when the application already has a clear secret-lifetime model. The caller must create the descriptor, size the file, map it, keep the descriptor and mapping under control, and release both on every exit path. Current Linux man-pages documentation notes that glibc does not provide a wrapper, so programs generally invoke the system call through syscall() and handle unsupported kernels explicitly.
Descriptor, size, mapping
The returned file descriptor starts with a zero-length object. Use ftruncate() to set a bounded size before calling mmap(). The mapping is locked in memory, cannot be swapped in the ordinary way, and is constrained by the process’s locked-memory resource limit. Pages are faulted in as they are touched rather than necessarily being allocated all at once by mmap().
This bounded C sketch preserves errno during cleanup and returns the descriptor and mapping only after both have been acquired successfully. It deliberately leaves the application-specific secret handling and final zeroization policy to the caller.
#define _GNU_SOURCE
#include <errno.h>
#include <stddef.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#define SECRET_BYTES 4096
static int create_secret_region(int *fd_out, void **mapping_out)
{
if (fd_out == NULL || mapping_out == NULL) {
errno = EINVAL;
return -1;
}
*fd_out = -1;
*mapping_out = MAP_FAILED;
#ifndef SYS_memfd_secret
errno = ENOSYS;
return -1;
#else
long result = syscall(SYS_memfd_secret, 0);
if (result == -1)
return -1;
int fd = (int)result;
if (ftruncate(fd, SECRET_BYTES) == -1) {
int saved_errno = errno;
(void)close(fd);
errno = saved_errno;
return -1;
}
void *mapping = mmap(NULL, SECRET_BYTES, PROT_READ | PROT_WRITE,
MAP_SHARED, fd, 0);
if (mapping == MAP_FAILED) {
int saved_errno = errno;
(void)close(fd);
errno = saved_errno;
return -1;
}
*fd_out = fd;
*mapping_out = mapping;
return 0;
#endif
}
/* On success, use the region, then munmap(mapping, SECRET_BYTES) and close(fd). */
Compile-time headers and syscall numbers vary by architecture and libc version. Check that target headers define the syscall number, and test the actual running kernel rather than inferring support from the build machine. Do not hard-code a syscall number copied from another architecture. In production, wrap this platform-specific setup behind a small interface that can return a clear “secure memory unavailable” result.
The file descriptor is a capability to the backing object. Passing it to another process changes who may map the region; inheriting or duplicating it therefore changes the security boundary. Closing the descriptor is not enough if a mapping remains active, and unmapping is not enough if another process still has a descriptor. Teardown should account for every mapping and reference, then close the descriptor.
What the protection does and does not mean
Removing the secret pages from the kernel’s direct map makes it harder for an exploited kernel to read them through the common direct-map route. The kernel also prevents ordinary mechanisms such as get_user_pages() from pinning the region for arbitrary kernel access. Active secret-memory users inhibit hibernation so the contents are not copied into a hibernation image.
These are meaningful mitigations, but their scope is narrow. The owning process can read and modify its mapping; memory disclosure within that process remains a serious threat. A sufficiently privileged attacker may use other mechanisms, including controlling the process or manipulating page tables. Physical attacks, firmware compromise, side channels, and implementation defects are separate threat classes. Do not describe the feature as “memory the kernel can never read” without qualifying the documented threat model.
The API also does not encrypt bytes at rest in application-managed files, secure data while it is copied into ordinary buffers, or guarantee crash-safe cleanup. Minimize copies, keep secret buffers short-lived, avoid logging their contents, and consider clearing the mapping before release when the data type and compiler behavior make that meaningful. Zeroization complements the mapping boundary; it does not turn it into a complete defense.
Limits and failure policy
Because the mapping consumes locked memory, large allocations can fail even when the machine has abundant free RAM. Inspect and deliberately configure RLIMIT_MEMLOCK for the service instead of raising it broadly. Reserve only the amount the application needs, and test allocation failures under the same service manager and container limits used in production.
Kernel support and configuration are also relevant. The system call was introduced in Linux 5.14; older kernels or kernels where the feature is disabled can return ENOSYS. Unknown flag bits fail with EINVAL, while ordinary descriptor or memory exhaustion can fail creation or sizing. Treat all such outcomes as explicit security-policy decisions. Falling back silently to a normal heap buffer may keep an application running while invalidating its security claim.
A defensible deployment records the kernel version, architecture, syscall result, locked-memory limit, mapping size, and fallback outcome. Test normal creation, resource-limit exhaustion, unsupported-kernel behavior, descriptor transfer, process exit, and cleanup. The relevant result is not merely “the syscall returned an fd”; it is whether the complete secret lifecycle behaves as intended under failure.
Related:
- Linux userfaultfd: Delegating Page-Fault Resolution to a Userspace Manager
- Reading /proc and /sys: The Kernel’s Window into Userspace
Sources: