Skip to content
FreeBSDDeep Dive Published Updated 8 min readViews unavailable

FreeBSD Asynchronous I/O: Bound Requests and Reap Completions Correctly

Understand FreeBSD POSIX AIO request lifetimes, safe descriptor types, completion errors, queue limits, and operational checks before tuning.

Asynchronous I/O (AIO) lets a process submit an I/O request and collect its result later, rather than blocking the submitting thread until the operation finishes. That distinction can help an application overlap work, but it does not guarantee higher throughput, eliminate storage latency, or make every file descriptor safe for asynchronous operations. FreeBSD’s AIO facility has kernel-managed requests, completion APIs, resource limits, and explicit caveats for descriptor types that can block an AIO daemon.

An AIO request is not “a thread that will finish eventually.” The caller owns an aiocb control block and its referenced buffer until the operation completes and the result has been collected. Reusing that memory too early, submitting duplicate work through the same control block, or treating an accepted request as successful data transfer creates correctness bugs. Start with one bounded read on a regular file, measure it, and expand concurrency only when the workload and application design justify it.

Identify the supported operation before changing tunables

Confirm the release, kernel, and AIO-related settings before comparing behavior across hosts:

freebsd-version -kru
sysctl vfs.aio

Some vfs.aio nodes may not be present or readable on every supported release and kernel configuration. Capture the exact error instead of copying a tuning recipe from another operating system. FreeBSD documents asynchronous reads, writes, vector operations, list I/O, cancellation, result inspection, and waiting through separate system calls. The portable POSIX interfaces are preferable when source portability matters; FreeBSD-specific extensions such as aio_read2(2) and aio_readv(2) should be selected only when their platform dependence is intentional.

The AIO facility is integrated into supported FreeBSD kernels, but the installed aio(4) manual remains authoritative for the actual system. Do not assume that loading an aio module is required merely because an older example includes a kldload command. Verify the release-specific manual and kldstat only if the kernel documentation for that release says a module is applicable.

Keep the request and buffer alive until reaping

The following complete example performs one read from a regular file. It waits for completion, reports an asynchronous operation error separately from a submission error, calls aio_return() once, and only then closes the descriptor and exits. It deliberately uses no signal handler or queue of outstanding requests:

#include <aio.h>
#include <errno.h>
#include <fcntl.h>
#include <signal.h>
#include <stdio.h>
#include <string.h>
#include <time.h>
#include <unistd.h>

int
main(int argc, char **argv)
{
    struct aiocb cb;
    char buffer[4096];
    const struct aiocb *pending[1];
    int fd, error;
    ssize_t bytes;

    if (argc != 2) {
        fprintf(stderr, "usage: %s file\n", argv[0]);
        return 2;
    }
    fd = open(argv[1], O_RDONLY);
    if (fd == -1) {
        perror("open");
        return 1;
    }
    memset(&cb, 0, sizeof(cb));
    cb.aio_fildes = fd;
    cb.aio_buf = buffer;
    cb.aio_nbytes = sizeof(buffer);
    cb.aio_offset = 0;
    cb.aio_sigevent.sigev_notify = SIGEV_NONE;
    if (aio_read(&cb) == -1) {
        perror("aio_read");
        close(fd);
        return 1;
    }
    pending[0] = &cb;
    for (;;) {
        error = aio_error(&cb);
        if (error != EINPROGRESS)
            break;
        if (aio_suspend(pending, 1, NULL) == -1) {
            if (errno == EINTR)
                continue;
            perror("aio_suspend");

            /* Keep the request and buffer alive until it is terminal. */
            do {
                error = aio_error(&cb);
                if (error != EINPROGRESS)
                    break;
                (void)nanosleep(&(struct timespec){
                    .tv_sec = 0, .tv_nsec = 100000000
                }, NULL);
            } while (error == EINPROGRESS);
            break;
        }
    }
    if (error == -1) {
        perror("aio_error");
        close(fd);
        return 1;
    }
    if (error != 0) {
        fprintf(stderr, "asynchronous read failed: %s\n", strerror(error));
        (void)aio_return(&cb);
        close(fd);
        return 1;
    }
    bytes = aio_return(&cb);
    if (bytes == -1) {
        perror("aio_return");
        close(fd);
        return 1;
    }
    printf("read %zd bytes\n", bytes);
    close(fd);
    return 0;
}

If aio_suspend() fails for a reason other than signal interruption, the fallback keeps the control block and buffer alive and polls aio_error() until it no longer reports EINPROGRESS. A production service should expose this degraded wait path in telemetry and choose a polling interval appropriate to its latency and CPU budget.

Compile this on FreeBSD with:

cc -Wall -Wextra -O2 aio-read-one.c -o aio-read-one

The example’s fixed buffer is large enough for one bounded sample, not a general file-copy strategy. A short read is valid at end of file, and a successful completion with zero bytes means end-of-file at the chosen offset. A production reader must define how it handles partial results, retries, file growth, and whether file contents can change while work is in flight.

Distinguish submission from completion

aio_read() returning zero means the request was queued; it does not mean the data arrived. aio_error() returns EINPROGRESS while the operation is incomplete, zero after successful completion, or an error number describing a failed or canceled operation. aio_return() retrieves the result of the completed operation and must be called only once per request. Keep the aiocb, buffer, file descriptor, and any notification state valid until that lifecycle ends.

aio_suspend() can wait for one or more requests to complete. It can be interrupted by a signal, so a loop must handle EINTR without discarding request state. aio_waitcomplete() is another FreeBSD interface that waits and returns a completed control block; its semantics differ from a polling loop and it is not a portable POSIX API. aio_cancel() requests cancellation but does not justify freeing the buffer until the operation is known to be canceled or complete. Review the specific cancellation result and then reap the request.

For multiple operations, maintain an explicit table of request ownership and state. A minimal state machine is “prepared, submitted, in progress, completed, reaped.” Never submit two operations through one aiocb while the first is still outstanding. Do not use a stack buffer whose scope ends before completion. A request must also use a file offset compatible with the other I/O paths: concurrent writes to the same range can race even when each individual call is asynchronous.

Know which descriptors can stall the AIO worker pool

FreeBSD’s aio(4) manual warns that some descriptor types can block an AIO daemon indefinitely. Operations on those unsafe descriptor types are disabled by default because they can hang a process or the system. The manual identifies sockets, raw disk devices, and regular files on local filesystems as descriptor classes that do not block indefinitely in this implementation, while some other file types require special caution. The exact supported and unsafe sets belong to the release’s aio(4) page, not an assumption based on Linux or a library’s API surface.

Do not enable vfs.aio.enable_unsafe merely to make an error disappear. First determine the descriptor type, the application’s requirement, and the documented failure mode. A daemon blocked on a FIFO, device, or remote filesystem can consume a worker without producing completion. If an operation needs readiness-based I/O on a socket, use kqueue(2) or an event framework designed for that descriptor instead of enabling unsafe AIO globally.

Bound concurrency and backpressure

Every outstanding request consumes kernel and application resources. FreeBSD exposes AIO tunables under vfs.aio; inspect the nodes documented by the installed aio(4) before raising limits. Raising a maximum can permit more queued work but cannot make a device complete requests faster. A huge application queue can increase latency, memory use, and recovery time while concealing a saturated disk.

Set a per-process in-flight limit based on the target’s tested service rate and memory budget. When the limit is reached, stop submitting and process completions before accepting more work. Track queue depth, oldest request age, error counts, bytes completed, and time from submission to completion. A growing queue with stable throughput indicates the producer is outrunning the consumer or the storage path is saturated. Do not respond by continually increasing queue limits.

Measure with a repeatable workload and control cache effects. Compare synchronous and asynchronous implementations using the same file, block sizes, concurrency, storage, and durability requirements. Avoid using a raw device against production media to “benchmark” AIO. A benchmark that omits fsync, O_SYNC, or the application’s actual durability path can report an impressive number that does not represent safe writes.

Diagnose errors by phase

If aio_read() or aio_write() fails immediately, investigate request validation, descriptor access mode, invalid offset, exhausted per-process or system resources, or missing AIO support. An immediate failure is different from a queued operation that later returns an error through aio_error(). Preserve errno immediately after the system call and record the asynchronous error returned by aio_error() as its own field.

If requests remain EINPROGRESS for an unexpectedly long time, collect the descriptor type, storage state, mount type, queue depth, outstanding request age, and application logs. Compare a bounded test against a regular file on a local filesystem and use iostat or gstat to check whether the device is already saturated. Do not conclude that a hung request proves an AIO kernel defect; blocked I/O, remote storage, device firmware, or an application never reaping results can look similar.

If memory corruption or intermittent wrong data appears, audit lifetime and reuse first. The buffer cannot be read or written by the application while the kernel may still be using it. An aiocb cannot be reused until its prior result has been collected. Validate alignment only when the descriptor or API requires it; do not invent direct-I/O alignment rules for ordinary buffered file I/O.

Acceptance criteria

An AIO implementation is operationally acceptable when its supported descriptor types are documented, every submitted request has exactly one completion path, request storage remains valid through reaping, queue depth is bounded, and partial/error results are handled. Demonstrate behavior under normal load, a full queue, cancellation, signal interruption, and a file ending before the requested buffer length. Capture the installed release and applicable manual page alongside performance measurements.

The safe default is to use AIO only when concurrency solves a measured application problem. Where simple blocking I/O is already sufficient, asynchronous machinery can add state and error paths without improving service. Where it is needed, correctness comes from explicit ownership and completion discipline, not from the word “asynchronous.”

Related:

Sources:

Comments