Linux close_range: File-Descriptor Hygiene Before exec
Use close_range to close or mark descriptor ranges atomically, avoid procfs enumeration races, and preserve only the descriptors a child process needs.
Every file descriptor that survives into a child process can become an accidental capability. A leaked listening socket may keep a service port open, a pipe end may prevent another process from observing end-of-file, and a directory descriptor may expose a filesystem location the child was not meant to access. Closing descriptors before execve() is therefore part of process isolation, not just cleanup.
Linux close_range() makes this boundary easier to express. It closes every open descriptor in an inclusive numeric range, or can mark that range close-on-exec. The call avoids enumerating /proc/self/fd and iterating through a list that can change while other threads open or close descriptors. Its flags also make it useful in multithreaded launchers that share a descriptor table.
Decide which descriptors are part of the child contract
Start by listing the descriptors the new program is supposed to inherit: usually standard input, output, and error, plus a deliberately passed socket, pipe, or pre-opened file. Everything else should be closed or marked close-on-exec before the new image runs.
Descriptor numbers are process-local indexes, not object identities. The same number can be reused after a close, and duplicated descriptors may refer to the same open file description. A launch plan should name the intended descriptors by role, then make sure each one has a single owner and documented inheritance behavior.
The simplest common policy preserves descriptors 0 through 2 and closes everything from 3 upward:
#define _GNU_SOURCE
#include <limits.h>
#include <linux/close_range.h>
#include <unistd.h>
int close_child_extras(void) {
return close_range(3, UINT_MAX, 0);
}
This is a policy example, not a universal rule. If the child needs a descriptor numbered 5, either move it into a reserved slot and close only the surrounding ranges, or mark unwanted descriptors close-on-exec while preserving the allowlist. Do not assume that descriptors 0, 1, and 2 are open: launchers sometimes close or repurpose them, so explicitly establish their intended values as part of the child contract.
close_range(first, last, 0) includes both bounds. A high bound of UINT_MAX expresses “through the end of the descriptor-number space” without first asking how many descriptors happen to be open. Individual close errors are ignored by the interface; callers should handle a failure of the range operation itself and should not use this syscall as proof that an unrelated I/O operation completed successfully.
Close now or defer closure until exec
Immediate closure is appropriate when the current process will never use the selected descriptors again. CLOSE_RANGE_CLOEXEC instead sets close-on-exec for every descriptor in the range. The descriptors remain available to the current image and disappear when a later execve() successfully replaces it.
#define _GNU_SOURCE
#include <limits.h>
#include <linux/close_range.h>
#include <unistd.h>
int mark_extras_for_exec(void) {
return close_range(3, UINT_MAX, CLOSE_RANGE_CLOEXEC);
}
This deferred mode is useful when child setup needs temporary access to descriptors before execve(), or when a process must establish a seccomp policy before closing descriptors that the policy setup itself needs. It also reduces the number of timing-sensitive close operations between “prepare child” and “replace image.”
Close-on-exec does not close the descriptor when the function returns, and it does not affect a fork() by itself. The flag is applied to descriptor-table entries, so code that continues running in the same image can still use those descriptors. A failed execve() leaves the original process image and descriptors present, with their close-on-exec bits still set. A launcher that falls back after an exec failure must know whether that state is acceptable.
Use atomic close-on-exec creation flags such as O_CLOEXEC, SOCK_CLOEXEC, and EFD_CLOEXEC whenever possible. Setting the flag later with fcntl() can race with another thread that forks and execs between descriptor creation and the flag update. A final close_range() sweep is a useful defense-in-depth boundary, but it should not excuse APIs that create inheritable descriptors by default.
Unshare a shared descriptor table before changing it
Threads commonly share a file-descriptor table. In Linux terms this can result from normal thread creation or from processes created with CLONE_FILES. If one thread closes a descriptor while another is setting up an exec, a loop of separate close() calls can observe a changing table and produce a hard-to-audit inheritance boundary.
CLOSE_RANGE_UNSHARE asks the kernel to first give the caller a separate descriptor table, then apply the requested range operation to that table. This prevents the close operation from changing the shared table used by the other threads. It is particularly useful in a child-launch path that cannot safely serialize all descriptor creation and closure across the entire process.
#define _GNU_SOURCE
#include <limits.h>
#include <linux/close_range.h>
#include <unistd.h>
int isolate_child_descriptor_table(void) {
return close_range(3, UINT_MAX,
CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC);
}
After this call, the caller has a private table for the range operation; other threads retain their own references. Combining unshare with close-on-exec is often convenient in a launcher because it prevents both concurrent table mutation and accidental inheritance by the replacement program. The operation can fail while the kernel constructs the new table, including resource-related failures. Treat failure as a failed isolation step and do not continue as if descriptors had been sanitized.
CLOSE_RANGE_UNSHARE is not a general lock for application state. It does not stop other threads from changing files, writing to shared sockets, or modifying objects that both tables reference. It isolates the descriptor-table entries for this operation; duplicated descriptors in either table may still refer to the same underlying open file description and therefore share file status and offset state.
Why procfs enumeration is more fragile
Older launchers often open /proc/self/fd, enumerate its entries, and call close() on each descriptor. That approach has several edge cases: the enumeration directory itself has a descriptor, another thread can allocate new descriptors after the scan begins, and numeric entries can be reused between listing and closing. A fixed upper bound loop avoids procfs but may perform a very large number of syscalls and still races with concurrent descriptor creation.
close_range() expresses the range operation in one syscall and does not depend on procfs being mounted. This is useful in containers, early boot, restricted sandboxes, and small pre-exec helpers. The performance benefit is workload- and kernel-dependent, but the semantic benefit is a smaller and clearer race surface.
There are still cases where an allowlist is clearer than a blanket range. If a launcher must retain several noncontiguous descriptors, a careful strategy can mark the whole range close-on-exec and then clear FD_CLOEXEC only on the intended descriptors, or close the ranges around explicitly retained slots. Coordinate that work with the descriptor-table sharing model. A retained descriptor is authority: verify its peer, path, access mode, and intended use before handing it to a less-trusted child.
Make compatibility and fallback deliberate
close_range() was added in Linux 5.9; the glibc wrapper is documented since glibc 2.34, and CLOSE_RANGE_CLOEXEC is available since Linux 5.11. Kernel support, libc support, and header support are separate questions. Build against the target sysroot, and do not infer runtime support merely because a constant exists at compile time.
If an older target must be supported, create a fallback with an explicit contract. A /proc/self/fd scan is not always available and needs careful handling of the scan descriptor. A loop to a configured descriptor limit is simpler but can be expensive. A launcher can also set close-on-exec atomically on each descriptor as it creates it, then pass an explicit descriptor allowlist to the child. Treat ENOSYS, EINVAL for unsupported flags, permission restrictions, and resource errors differently; do not retry with a weaker path after every failure without recording the downgrade.
When using a raw syscall instead of a libc wrapper, keep architecture-specific syscall-number concerns out of application logic. Prefer the libc API where it exists, or centralize a compatibility wrapper and test it on each supported ABI. Header constants and the running kernel can diverge on older systems.
Verify the boundary, not just the syscall result
Test with a child that inspects /proc/self/fd and reports only descriptor numbers and expected roles. Include descriptors created by sockets, pipes, eventfd, signalfd, epoll, and temporary files. Test concurrent descriptor creation, a failed exec path, a successful exec path, a child that receives one explicitly preserved descriptor, and a parent with a non-default descriptor limit.
In production, log the launch policy and whether the range call succeeded, but do not log secret file contents or credentials. A child can verify its own descriptor set as an additional invariant. If a service unexpectedly keeps a pipe or socket alive, inspect every process that inherited a writer or duplicate; the original creator may not be the remaining reference.
close_range() is a precise tool for descriptor lifecycle, but the security property comes from the allowlist and launch contract around it. Decide what the child may inherit, create descriptors close-on-exec where possible, isolate shared tables when required, and fail closed when the final sanitation step cannot be confirmed.
Related:
- Linux pidfds: Race-Free Process Handles Beyond Numeric PIDs
- Linux signalfd: Consume Signals in an Event Loop Without Async Handlers
Sources: