Skip to content
Shell & TerminalDeep Dive Published Updated 7 min readViews unavailable

Bash wait -n and -p: Reap One Child Without Losing Its Exit Status

Use Bash wait -n and wait -p to supervise asynchronous children while preserving PID-to-result mapping, signal outcomes, and bounded job concurrency.

When a Bash script launches several background commands, wait is the boundary between starting work and knowing how it ended. Calling wait on every process in launch order is simple, but it can leave the script blocked on a slow first child while later children have already failed. Bash’s wait -n waits for one eligible child or job to complete, and wait -p records which one supplied the returned status. Together they support a completion-driven supervisor rather than a polling loop.

These options are Bash extensions, not POSIX sh syntax. wait -n arrived in Bash 4.3; wait -p was added in Bash 5.1. Scripts that may run on macOS’s system Bash 3.2, BusyBox ash, or a POSIX shell need another strategy or an explicit interpreter/version requirement. Check BASH_VERSION at runtime and fail with a clear diagnostic before launching expensive work.

Understand what wait actually reports

wait PID waits for a child process known to the current shell and returns that child’s status. A job specification can name a shell job when job control is available; a jobspec may represent a pipeline with multiple processes. A shell can wait only for children it owns, not arbitrary PIDs found by scanning the process table.

With no IDs, plain wait waits for all running background jobs, but it does not give the caller an individual status record for each one. Its return status is zero under the documented no-argument behavior. If the script needs to associate each failure with a particular task, it must retain its own IDs and wait for them deliberately.

wait -n [IDs...] returns when one eligible child or job completes and returns that completion’s status. With an ID list, it restricts the candidates to the named children/jobs. Bash 5.1’s -p variable stores the identifier corresponding to the result and unsets that variable before assignment; it is intended for use with -n. If the wait is interrupted by a signal, the variable remains unset and the return status is greater than 128. These outcomes make the identifier variable part of the control flow, not optional decoration.

Maintain an explicit set of children

Capture $! immediately after launching each asynchronous command. Store the value alongside a label or task description, so the final report can identify which unit failed. $! refers to the most recently backgrounded process or process substitution; a later asynchronous launch overwrites it. Do not attempt to reconstruct the child list later from jobs output, which is intended for shell job control and is not a durable supervisor database.

An illustrative bounded batch is:

declare -A task_for_pid=()

build_one alpha &
task_for_pid[$!]=alpha
build_one beta &
task_for_pid[$!]=beta
build_one gamma &
task_for_pid[$!]=gamma

while ((${#task_for_pid[@]})); do
    finished_pid=
    if wait -n -p finished_pid "${!task_for_pid[@]}"; then
        status=0
    else
        status=$?
    fi

    if [[ -z $finished_pid ]]; then
        printf 'wait was interrupted or no child was reaped\n' >&2
        break
    fi

    printf '%s completed with status %d\n' \
        "${task_for_pid[$finished_pid]}" "$status"
    unset 'task_for_pid[$finished_pid]'
done

This example requires Bash 5.1 or newer for associative arrays and wait -p; a production supervisor should also decide whether an interrupted wait should retry, initiate shutdown, or return a distinct failure. The function names are placeholders. The important invariant is that the ID returned by -p maps to exactly one outstanding child and is removed exactly once.

Bound concurrency instead of launching everything

Completion-driven waiting makes it natural to keep at most N tasks in flight. Start children until the active set reaches the configured limit, wait for any one to finish, record its status, then start the next input. This provides backpressure: input volume does not become an unbounded number of processes, open files, or simultaneous remote requests.

Treat task launch failure separately from child failure. A shell error while preparing a command may prevent a process from being started at all, while a launched child later returns a nonzero status. Record a child only after a successful asynchronous launch has assigned a valid $!; keep the task label available so failure to launch can be reported without inventing a PID.

Do not use wait -n as a generic “sleep until something changes” primitive. It only concerns child processes/jobs known to the shell. If no eligible child exists, Bash reports an error status (commonly 127 for no unwaited child); a loop that ignores that condition can spin or misreport success.

Distinguish a nonzero child from an interrupted wait

The child exit code is the return value from wait -n, but shells conventionally use values greater than 128 to represent signal termination. A process can also deliberately exit with such a numeric status. Therefore, a status alone may not always distinguish “the command returned 143” from “the child died from SIGTERM” in a portable, universally reliable way. If this distinction matters, use a wrapper protocol or platform-specific process metadata and state the assumptions.

An unset -p variable is a critical clue: it means the call did not provide a child identifier, such as when the wait was interrupted. Never index an associative array using that unset value under set -u. Initialize a scratch variable before every wait, verify it is set, then process the returned status. If interruption is expected, install an explicit signal policy and make retry/termination behavior deterministic.

With job control enabled, wait -f changes behavior by requiring termination rather than returning when a child merely changes state. Most noninteractive scripts leave job control disabled, but do not assume that inherited shell options make this universal. If stopped jobs are possible, decide whether the supervisor cares about termination or any state change and select the option accordingly.

Preserve pipeline and wrapper semantics

If a background task is a pipeline, its job result and its individual command statuses are not automatically the same thing. Without set -o pipefail, a pipeline usually reports the final command’s status; with pipefail, Bash reports a nonzero status according to pipeline rules. If diagnostics need every component, capture PIPESTATUS immediately in the process that executed the pipeline or have a wrapper write a structured result. Reading PIPESTATUS later is unsafe because another command overwrites it.

Subshells and command substitutions also establish different process and variable lifetimes. A child cannot mutate the parent’s associative array or ordinary shell variable directly. Send results through its exit code, a dedicated file/pipe, or a structured message with clearly defined ownership and cleanup. Do not try to pass complex data through a single exit status.

Test the supervisor with controlled outcomes

Use small children that sleep for different durations and exit with known statuses. Confirm that the completion log follows finish order rather than launch order, every PID is reported once, and a failed child does not prevent the remaining jobs from being reaped. Add tests for a signal during wait -n, no eligible children, one child exiting before the first wait, a wrapper pipeline, and a maximum concurrency limit.

Run the tests under the minimum supported Bash release as well as the newest target. This host’s macOS system Bash may be 3.2, which cannot validate wait -n or wait -p; use a separately installed supported Bash and verify its resolved path before running the feature tests. Static parsing with bash -n on an older binary can itself fail because it does not recognize newer syntax, so it is not a substitute for a modern interpreter.

The GNU Bash Reference Manual defines the builtin’s accepted operands and edge cases. Use the Bash 4.3 and 5.1 release notes to establish feature floors, and state those floors in the script’s documentation. A reliable supervisor is not just a loop around wait -n: it is an explicit policy for which children are owned, how results map to tasks, what interruption means, how concurrency is bounded, and how every child is cleaned up.

Related:

Sources:

Comments