Skip to content
FreeBSDDeep Dive Published Updated 8 min readViews unavailable

FreeBSD Process Descriptors: Track Child Lifecycles Without PID Races

Use FreeBSD process descriptors to signal, wait for, and monitor child processes by file descriptor, with explicit close and Capsicum semantics.

Traditional process control stores a child PID and later passes it to kill or wait. That model is familiar, but a numeric PID is a name in a reusable namespace: code that retains a stale PID after the child has been reaped risks referring to a different process. FreeBSD process descriptors provide a file-descriptor-oriented reference to a process. The pdfork family combines process creation with a descriptor that can be used for signaling, waiting, and lifecycle observation.

The interface is useful when a program manages its own child processes, especially when it needs to hand a constrained process reference into Capsicum capability mode. It is not a general service supervisor and does not implement restart policy, readiness, durable queues, logging, or resource limits. Choose it for the process-reference and event-notification semantics, then use the appropriate service or resource-management layer for the rest.

Understand the reference and version boundary

Include sys/procdesc.h and use the functions documented for the target release. On FreeBSD 15.1, pdfork creates a child and returns a PID in the fork-style way; the parent also receives a process descriptor through the fdp pointer. pdkill takes that descriptor rather than a PID, pdgetpid can retrieve the conventional PID, and pdwait can retrieve status for the process described by the descriptor. The newer pdrfork and pdwait interfaces first appeared in 15.1, so software targeting earlier releases must feature-detect or use an older compatible wait design.

The descriptor remains a kernel reference to a particular process. While it remains open, the PID of the referenced process is not reused according to the manual. This avoids a common stale-PID race in the controlling program. It does not eliminate all races in the application: the program must still close descriptors on every path, handle failed creation, manage child-side behavior, and avoid accidentally sharing control descriptors with unrelated processes.

Here is a minimal FreeBSD 15.1 example that starts a short-lived child and waits for it through the descriptor:

#include <sys/procdesc.h>
#include <sys/wait.h>

#include <err.h>
#include <stdio.h>
#include <unistd.h>

int
main(void)
{
    int process_fd;
    pid_t child = pdfork(&process_fd, PD_CLOEXEC);

    if (child == -1)
        err(1, "pdfork");
    if (child == 0) {
        execl("/bin/sh", "sh", "-c", "exit 7", (char *)NULL);
        _exit(127);
    }

    int status;
    if (pdwait(process_fd, &status, 0, NULL, NULL) == -1)
        err(1, "pdwait");
    if (close(process_fd) == -1)
        err(1, "close process descriptor");

    if (WIFEXITED(status))
        printf("child exited with status %d\n", WEXITSTATUS(status));
    return 0;
}

Compile and run this only on FreeBSD 15.1 or a later release whose installed header and manual expose pdwait with this signature. The host used to author this article is not FreeBSD, so the example has not been compiled here. The code intentionally uses an absolute program path and handles exec failure with _exit; a production launcher should validate its own arguments and child error-reporting contract.

Treat descriptor close as a process-control operation

The default process-descriptor behavior is significant: if the referenced process is still alive and the last process-descriptor reference is closed, the kernel terminates it with SIGKILL. That can make cleanup deterministic, but it also means a careless close in an error path can kill a process the application intended to keep running.

The PD_DAEMON flag changes this close behavior so the process can survive explicit descriptor closure until killed through the ordinary process interface. The manual says PD_DAEMON is not permitted in Capsicum capability mode. Do not use it as a casual workaround for descriptor ownership bugs; it changes the lifetime contract and can leave an untracked process behind.

PD_CLOEXEC requests close-on-exec behavior. This is usually important when a parent launches another program: an unintended descriptor inherited across exec can keep the reference alive, change whether a last-close cleanup occurs, or give the new process a control handle it was not meant to receive. Make descriptor inheritance explicit, including in child code that forks again.

When a program needs the process to continue after losing the descriptor, it should make that decision deliberately and arrange an alternate lifecycle authority before closing the last reference. Conversely, when descriptor close is the cancellation mechanism, test it under normal exit, exec failure, parent cancellation, and duplicate-descriptor paths. An unclosed descriptor can prevent the expected last-close behavior.

Signal and wait through the descriptor

pdkill has signaling semantics similar to kill, but uses a process descriptor. It returns an error when the descriptor is invalid or lacks the necessary rights. A successful signal request is not proof that the child has completed cleanup; the controller still needs to wait for termination and inspect the final status. Treat ESRCH-like or capability-right errors as lifecycle evidence rather than retrying blindly.

The pdwait call is intended to wait for the process referenced by the descriptor and return status information. On older releases, the available interface may instead expose pdwait4 or rely on conventional wait calls. Do not infer API availability from the host kernel version alone; check the target system’s header, libc, and manual together. A binary compiled against a newer interface may fail to link or run on a release that lacks it.

You can also ask for the conventional PID when an existing API requires it:

pid_t child_pid;

if (pdgetpid(process_fd, &child_pid) == -1)
    err(1, "pdgetpid");

Once the PID is obtained, it is still only a numeric identifier. Keep the process descriptor open while using it as a correlation value and prefer pdkill or pdwait for operations the descriptor API supports. Do not pass the PID to an unrelated asynchronous subsystem and later assume the number still identifies the same process after the descriptor has been closed.

Integrate lifecycle monitoring with kqueue

FreeBSD supports monitoring a process descriptor with kqueue using the EVFILT_PROCDESC filter. The documented NOTE_EXIT event indicates process exit. This lets an event loop wait for child state transitions alongside sockets, timers, and other descriptors rather than polling process tables or installing one signal handler per child.

The event is not a readiness notification and does not report application health. A NOTE_EXIT event means the process exited; it does not mean the work succeeded. Retrieve and interpret the child status with the appropriate wait call. Use the exact event flags and behavior documented by the target release because the manual marks currently supported notifications and fields explicitly.

When adding a process descriptor to an event loop, design for cancellation, descriptor closure, and event deregistration. A descriptor that is closed while registered may result in event-loop behavior that differs from a still-live child. Remove or update the registration according to the kqueue manual and preserve one owner for the process descriptor. Test fast-exiting children as well as long-running ones so an exit before event registration does not become a missed notification.

Use capability rights and process descriptors for different questions

Process descriptors complement Capsicum because a descriptor can act as a capability-like reference to one process instead of a global PID lookup. Capability mode further restricts which operations the process can perform. The API can report insufficient rights, such as when CAP_PDKILL is unavailable. Grant only the rights the worker needs, and keep the control descriptor out of code that does not own the child lifecycle.

Do not confuse a process descriptor with a jail, a process group, a reaper, or a service manager. A process descriptor refers to one process. It does not automatically contain grandchildren, limit CPU or memory, isolate filesystem access, restart failed programs, or preserve work across reboot. Use procctl reaper features, RCTL, jails, rc.d, or an external orchestrator where those contracts are required.

Failure handling and validation

For each child, record the descriptor owner, creation result, PID for logs, event registration, termination request, wait result, exit status, and close outcome. If pdfork fails, no child should be assumed to exist. If descriptor copyout reports EFAULT, the manual notes the child may already have been created and both parent and child continue; code that must handle this rare case needs a documented way to determine which branch is running.

Test at least these cases in a FreeBSD test environment: successful exit, nonzero exit, exec failure, signal termination, parent cancellation while the child is active, and descriptor close with and without PD_DAEMON where allowed. Confirm there is no orphaned worker and no unintended SIGKILL. Test Capsicum rights separately from ordinary process-descriptor behavior.

Do not publish code that silently discards pdwait status or closes the descriptor before confirming the intended child outcome. The descriptor solves identity and lifecycle-reference problems; it does not make error handling optional.

Acceptance criteria

The design is ready when every child has one clear control reference, exec inheritance is explicit, the parent can receive exit status without a PID reuse window, and cancellation behavior is tested. If kqueue is used, exit events must be correlated with a successful wait. If capability mode is used, the exact descriptor rights must be documented and exercised.

Keep the system boundary explicit: procdesc handles a process reference, while service supervision, isolation, resource control, and application readiness remain separate responsibilities. That boundary prevents a low-level kernel handle from being mistaken for an end-to-end reliability mechanism.

Related:

Sources:

Comments