Linux Open File Description Locks: Byte-Range Coordination Without PID Ownership
Use F_OFD_SETLK byte-range locks with deliberate descriptor ownership, nonblocking conflict handling, safe thread semantics, and explicit portability checks.
Linux open file description (OFD) locks are advisory byte-range locks acquired with fcntl() operations such as F_OFD_SETLK and F_OFD_SETLKW. Their defining property is ownership: a lock belongs to the open file description, the kernel object shared by duplicated descriptors, rather than to a process ID. That model avoids several surprising process-associated lock behaviors and makes OFD locks useful for cooperating processes and carefully designed multithreaded file access.
An OFD lock is not a transaction, a mandatory access-control rule, or proof that the file cannot change. It coordinates only with participants that use compatible locking operations on the same file object. A production design still needs an ownership protocol, a policy for timeouts and process failure, and a strategy for filesystems whose lock semantics differ from a local filesystem.
Choose the ownership model before choosing the command
Traditional F_SETLK record locks are associated with a process. OFD locks instead belong to one open file description. Two descriptors created by dup() refer to the same open file description, as do descriptors inherited across fork(). Locks made through those references therefore have the same lock owner. A second independent open() of the same file creates a different open file description and can conflict with the first.
This distinction matters especially in a multithreaded program. Threads commonly share a process and its descriptor table, so traditional process-associated record locks do not provide a per-thread exclusion boundary. OFD locks can provide one if each participant opens the file independently and retains its own descriptor. Calling dup() for each thread does not create independent lock ownership; it duplicates a reference to the same open file description.
Linux has supported OFD lock operations since kernel 3.15. The current Linux man-pages list the operations under POSIX.1-2024, but deployed kernel, libc headers, and non-Linux systems can differ. Build and runtime compatibility must be checked for the actual target rather than inferred from the presence of a constant in one development environment.
Acquire a bounded lock deliberately
struct flock describes a lock type and a byte range. With SEEK_SET, l_start is an absolute file offset. A positive l_len locks that many bytes; zero means from l_start through end-of-file, including bytes added later. The kernel requires l_pid to be zero for OFD lock operations. A read lock permits compatible readers, while a write lock conflicts with both read and write locks from a different owner.
#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <sys/types.h>
int try_ofd_write_lock(int fd, off_t start, off_t length) {
if (start < 0 || length < 0) {
errno = EINVAL;
return -1;
}
struct flock lock = {
.l_type = F_WRLCK,
.l_whence = SEEK_SET,
.l_start = start,
.l_len = length,
.l_pid = 0,
};
return fcntl(fd, F_OFD_SETLK, &lock);
}
The helper performs one nonblocking attempt. A return of zero means the lock was acquired. -1 with EAGAIN means an incompatible lock currently prevents acquisition; it is a normal conflict result, not necessarily a system failure. Other errors need separate treatment. In particular, EINVAL can indicate an unsupported operation or an invalid request, so validate the structure and test kernel support using a valid request on a controlled file before classifying it as a compatibility failure.
The descriptor must be open with access suitable for the requested lock: a read lock requires a readable descriptor, and a write lock requires a writable descriptor. Open the intended object first, then lock through that descriptor. If a pathname can be replaced or redirected by an untrusted actor, pathname checks performed before open() do not bind the eventual lock to the inspected object; use a trusted directory, descriptor-relative operations, and an explicit ownership policy.
Make contention, cancellation, and deadlines explicit
F_OFD_SETLK returns immediately on a conflict. A service can use it to implement a bounded retry policy with monotonic deadlines, jitter, cancellation, and observability. Do not busy-loop on EAGAIN: a tight retry can consume CPU while making no progress. A waiter should have a maximum wait, a cancellation path, and a useful record of which resource range it wanted.
F_OFD_SETLKW waits for the conflicting lock to be released. It may return EINTR when a signal interrupts the wait, so callers need to decide whether that means retry, cancellation, or a higher-level state transition. The Linux kernel does not perform deadlock detection for OFD locks. If two participants acquire ranges in inconsistent orders, both can wait indefinitely. Establish a global ordering for multi-range or multi-file locks, or use nonblocking acquisition with rollback and a bounded retry protocol.
Do not use F_OFD_GETLK as a reservation. It reports whether a conflicting lock exists at the moment of the query, but another participant can change lock state before a later acquisition. The query is useful for diagnostics or status displays; correctness must depend on the result of the actual F_OFD_SETLK or F_OFD_SETLKW operation.
Release locks through the same ownership boundary
An explicit unlock uses the same range description with l_type = F_UNLCK and an OFD set-lock command. Closing one duplicate descriptor does not necessarily release the lock if another reference still keeps the same open file description alive. The lock is automatically released when the last reference to that open file description is closed, including references inherited by a child process. A child that unintentionally inherits a descriptor can therefore extend a lock’s lifetime after the parent thinks it has shut down.
Create descriptors with close-on-exec behavior when a new program should not inherit the file or its lock lifetime. Audit fork(), descriptor duplication, descriptor passing, and every error path so the component that owns a lock also owns its release. A leaked descriptor can look like a stuck lock even after the original worker has exited.
OFD locks and traditional process-associated record locks do interact: conflicting lock types can block each other even when issued by the same process. Avoid mixing the two models on one file unless a documented protocol requires it. Linux flock() locks use another locking interface; do not assume they coordinate with fcntl() record locks on Linux merely because both are described as file locks.
Treat advisory locks as a cooperation protocol
Advisory locking does not stop a process from issuing read(), write(), or memory-mapped access without taking a compatible lock. Every writer that can violate the protected invariant must follow the same protocol. If a tool, legacy component, or administrator edits the file outside that protocol, the lock alone cannot preserve consistency.
Keep the lock scope aligned with the data invariant. Locking one record can protect an independently updated record only if all code agrees on record offsets and boundaries. A global metadata update may require a separate lock range or a higher-level transaction. Document whether a reader needs a shared lock while consuming the bytes or can safely read an immutable version after releasing the lock.
Locks are also not durability barriers. Successfully acquiring and releasing a lock says nothing about whether file contents have reached stable storage. If the application publishes durable state, combine its lock protocol with a separately designed write, synchronization, and recovery sequence. Likewise, a lock does not validate input, authenticate the other participant, or protect data from a privileged process.
Account for network and unusual filesystems
On a local filesystem, OFD locking is commonly used as a host-level coordination mechanism. On NFS and other network filesystems, lock state depends on client/server protocol, recovery, and the filesystem implementation. A network partition or administrative action can cause a previously acquired lock to be lost; subsequent I/O may fail, and there may be no asynchronous notification that lets the application safely assume its old ownership still holds.
Test the exact mount type, kernel, server, and failure behavior used in production. Do not infer distributed consensus from a successful fcntl() call. If correctness must survive a client partition or server failover, use an application protocol with fencing tokens or another mechanism that can reject stale owners at the resource being modified.
Keep file identity stable during the critical section. Replacing a pathname atomically can leave one participant locking an old inode while another opens the new inode under the same name. A lock coordinates on the opened file object, not on the spelling of its pathname. Include inode replacement, rename, hard-link, and temporary-file publication behavior in the design review.
Diagnose lock behavior without mistaking a snapshot for truth
Log the lock mode, file identity, byte range, caller, attempt duration, and final result. Avoid logging sensitive file contents or assuming that a process ID uniquely identifies an OFD lock owner. The kernel reports l_pid as -1 for a conflicting OFD lock in the F_OFD_GETLK result because the lock belongs to an open file description rather than a process.
A diagnostic query can become stale immediately. Pair it with syscall tracing or application-level ownership logs when investigating contention. Check descriptor lifetime, inherited references, lock ranges, traditional-versus-OFD interactions, and remote filesystem behavior before concluding that the kernel has left a lock orphaned.
Production checklist
- Confirm that every participant opens the same underlying file object and uses the same range convention.
- Use independent
open()calls when participants need independent OFD ownership;dup()does not create a new owner. - Set
l_pidto zero and use an access mode compatible with the requested lock. - Treat
EAGAINas expected contention, and give blocking waits a cancellation and deadlock-prevention strategy. - Close inherited and duplicated descriptors deliberately; the last reference controls automatic release.
- Do not mix lock families casually, and require every writer to follow the advisory protocol.
- Test the target kernel, libc, local or network filesystem, replacement behavior, and recovery paths.
OFD locks improve the ownership model for byte-range coordination, but they do not remove the need for a protocol. Make open-file-description lifetime, lock ordering, cancellation, filesystem behavior, and durability separate design decisions, then test each one under the same concurrency and failure conditions expected in production.
Related:
- Linux fallocate: Reserve, Punch, Zero, and Reshape File Space
- Fixing ‘Disk Full’ on Linux When df Shows Space Available
Sources: