flock in Shell: Lock Descriptors, Contention, and Safe Release
Coordinate Linux shell jobs with flock file descriptors, bounded waits, explicit exit codes, stable lock inodes, and tested cleanup behavior.
Two copies of a maintenance script can corrupt shared state even when each copy is individually correct. They may rotate the same logs, refresh the same cache, publish the same generated file, or run an expensive job twice. On Linux, flock provides advisory locks that can coordinate processes through an open file description. It is a small primitive, but its behavior is tied to file descriptors, filesystem semantics, process inheritance, and the exact scope of the protected operation.
This article covers the command-line flock utility commonly shipped with util-linux and the Linux flock(2) semantics it uses. It is not a portable POSIX shell facility, and filesystem behavior can vary on NFS, CIFS, containers, and network storage. Check the local flock(1) and flock(2) manuals, then test on the same filesystem and process model used in production.
A lock is held through an open file, not by a magic filename
The lock pathname is a rendezvous point: cooperating processes open the same underlying file and request compatible or conflicting locks. An exclusive lock allows one holder; shared locks can coexist with other shared locks while excluding an exclusive holder. By default, flock waits for an incompatible lock. -n requests nonblocking acquisition, and -w seconds bounds how long it waits.
The existence of the lock file does not mean that a job is running. The file can remain after every process exits. Conversely, deleting a lock file while another process still has it open can create a second inode at the same pathname. One process then holds a lock on the old inode while a new process locks the replacement, and the two processes can overlap. Keep the inode stable: create it once if needed, but do not unlink it as a normal cleanup step.
Choose a directory with the correct ownership and lifecycle. A per-user job may use a private runtime directory if one is guaranteed on that platform. A system job may use a root-owned lock directory. Avoid a world-writable directory with a predictable filename unless the directory is protected against symlink and replacement attacks. The path must resolve identically for every process that is supposed to coordinate; relative paths, chroots, mount namespaces, and container boundaries can silently create different lock domains.
Wrap a command when the lock covers one process lifetime
The simplest form asks flock to acquire a file lock and run a command while the lock is held:
flock --wait 15 /run/lock/cache-refresh.lock \
/usr/local/sbin/refresh-cache --source production
If the lock cannot be acquired before the wait expires, flock returns a conflict status (1 by default, or the value set by -E). If it starts the child successfully, the wrapper returns the child command’s status. Those outcomes have different meanings, so a caller should decide whether contention means “skip this duplicate run,” “retry,” or “fail the service.”
Use nonblocking mode for jobs where an already-running copy should simply suppress this invocation:
if flock --nonblock --conflict-exit-code 75 \
/run/lock/cache-refresh.lock \
/usr/local/sbin/refresh-cache --source production; then
printf '%s\n' 'refresh completed'
else
status=$?
if [ "$status" -eq 75 ]; then
printf '%s\n' 'refresh already in progress; skipping this run'
else
printf 'refresh failed (status %s)\n' "$status" >&2
exit "$status"
fi
fi
This assumes the chosen command’s own failure statuses do not collide with the custom conflict code. If they can, wrap the child in a helper that maps its status into a separately defined result contract. A lock collision should not be confused with a failed refresh, and a failed refresh should not be reported as a harmless collision.
The command form starts a child process and holds the lock around that command’s lifecycle. --no-fork (-F) replaces the flock process with the command while keeping the locked descriptor open; it can simplify supervision but is incompatible with options that require the wrapper to retain the descriptor. Understand whether your service manager tracks the wrapper or final executable before choosing it.
Lock inside a Bash script with a dedicated descriptor
When the critical section spans multiple shell commands, open a stable lock file on a dedicated file descriptor and ask flock to lock that descriptor:
#!/usr/bin/env bash
set -u
lock_file=/run/lock/report-build.lock
mkdir -p -- "${lock_file%/*}"
exec {lock_fd}>"$lock_file"
if ! flock --nonblock "$lock_fd"; then
printf '%s\n' 'another report build owns the lock' >&2
exit 75
fi
build_report
validate_report
publish_report
The descriptor remains open in the shell, so the lock stays held across the commands. The script should use set -u only if every referenced variable is initialized; it does not enable strict error propagation by itself. Critical commands still need explicit checks. If a validation step fails, do not publish a partial report just because the lock worked correctly.
The {lock_fd} dynamic descriptor allocation is a Bash feature, not POSIX syntax. If supporting an older Bash release, verify the minimum version. A fixed descriptor such as 9 can be simpler, but confirm it is not already in use by the caller or an included script. Use exec to keep the descriptor open in the current shell; an ordinary command-scoped redirection may close it when that command finishes.
Opening a file with > truncates it, but that does not invalidate the lock; flock coordinates the open file description and does not use the file’s contents. Some teams prefer >> to make it obvious that the path is a persistent lock object. Neither convention substitutes for correct permissions and stable path resolution.
Descriptor inheritance defines how long the lock lives
On Linux, locks are associated with an open file description. Duplicated descriptors refer to the same lock, and locks are released when explicitly unlocked or when all descriptors referring to that open file description are closed. A descriptor can survive fork and exec, so a child process that inherits the locked descriptor may unintentionally keep the lock after the parent shell exits.
This often surprises scripts that start background work:
exec {lock_fd}>/run/lock/nightly.lock
flock -n "$lock_fd" || exit 75
long_job &
exit 0
If long_job inherits lock_fd, the lock can remain held until that child exits. That may be desired if the child is the actual protected job, or it may block later runs unexpectedly. A process that daemonizes and leaves a descriptor open can make the lock appear stuck long after the initiating shell is gone.
If a child should not extend the critical section, close its copy explicitly or use the command wrapper’s --close (-o) option where the semantics fit. But closing the descriptor before executing a child also means the child no longer holds the lock; the parent must remain alive and keep its own descriptor open for the duration of the protected work. Test the actual descriptor inheritance behavior with your shell and tool versions rather than guessing from source indentation.
Prefer a simple ownership model: either the shell holds the lock for a synchronous critical section, or a foreground command owns the lock for its entire operation. Avoid launching detached work while relying on a lock held only by the launcher. If a worker needs the lock, let the worker acquire it itself or pass a deliberately inherited descriptor with clear documentation.
Separate serialization from atomic publication
A lock prevents cooperating writers from entering the protected region simultaneously. It does not make file writes atomic for readers, validate the data, or prevent a process that ignores the lock from modifying the same path. For generated state, combine the lock with a staging file and rename:
directory=/var/lib/report
temporary=$(mktemp "$directory/.report.XXXXXX") || exit 1
cleanup() {
status=$?
trap - EXIT
rm -f -- "$temporary" || printf '%s\n' 'warning: could not remove temporary report' >&2
exit "$status"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
if generate_report >"$temporary" && validate_report "$temporary"; then
chmod 0644 "$temporary" || exit 1
mv -f -- "$temporary" "$directory/report.json" || exit 1
else
printf '%s\n' 'report generation or validation failed' >&2
exit 1
fi
The temporary file is in the destination directory so the rename normally remains on one filesystem. Lock acquisition must happen before reading state that can be changed by another run, not just before the final write. Otherwise, two jobs can compute from the same stale input and serialize only their final publication. Readers that need a consistent snapshot should open the published file once and continue reading that descriptor.
Advisory means a non-cooperating process can ignore the lock. Use filesystem permissions, service ownership, or a stronger transactional storage mechanism if the data must be protected against arbitrary writers. If a lock controls a distributed job across hosts, verify that the shared filesystem implements the lock operation consistently under failure and failover. A local file lock is not a distributed consensus protocol.
Filesystem and portability limitations
The lock behavior of flock(2) is filesystem-dependent. The Linux man page documents limitations on some NFS and CIFS configurations, where locks may fail or behave differently according to mount and server options. Containers can also have distinct mount namespaces: two containers that use the same text path but different filesystems do not coordinate, while a shared bind mount can. Test contention across the actual nodes or containers that will run the job.
flock is not required by POSIX and is not identical across all Unix variants. The util-linux CLI supports options such as --close, --no-fork, and --conflict-exit-code; another implementation may not. If portability is a product requirement, document the dependency, feature-detect it, or choose a portable locking design whose semantics you can support. Do not claim a shell script is portable merely because its syntax is POSIX while its synchronization primitive is Linux-specific.
The kernel lock is advisory and generally tied to an open file description, which is different from older process-associated record locks. Do not mix different locking APIs or assume that applications using another API will necessarily conflict with flock; on Linux, flock and POSIX fcntl locks are distinct in important cases. Select one coordination protocol and make every participant use it.
Diagnose contention without deleting the lock file
On Linux, lslocks can show held locks, and process inspection can identify descriptors and command lines. Use these tools as diagnostics, not as an invitation to kill an arbitrary process. A lock file’s timestamp is not a reliable “last heartbeat”; it can be old while a valid lock is held or recent after a process crashed. The lock itself is authoritative for cooperating processes.
Log acquisition outcomes with the job identifier, lock path, wait duration, and operation phase. Do not record secrets or untrusted payloads. If a job waits too long, investigate whether a legitimate process is still working, whether a descriptor escaped into a detached child, whether the path points to a different inode, or whether the filesystem lock service is unhealthy. Deleting the path will not release a lock held on an open inode and can split the coordination domain.
For automated tests, launch two copies against the same temporary lock path. Have the first hold the lock for a bounded interval and assert that the second either waits or receives the documented conflict status. Then terminate the first process and verify that a later invocation acquires the lock. Repeat with a child process that inherits the descriptor, a path with spaces, and the production mount type. A single successful run proves only that the command executed; it does not prove mutual exclusion.
Finally, define what duplicate invocation means to the scheduler. Some cron jobs should skip a run if a previous one is active; others should queue, wait, or raise an alert. flock supplies a lock and acquisition result, not that business policy. When the timeout, descriptor lifetime, filesystem, and exit-status contract are explicit, it becomes a reliable building block for shell orchestration.
Related:
- Shell File Descriptors and Redirection: Ordering, Duplication, and Lifetime
- How to Run Parallel Shell Jobs and Collect Every Exit Status
Sources: