Skip to content
WindowsDeep Dive Published Updated 7 min readViews unavailable

Windows APCs and Alertable Waits: Thread-Affine Completion Callbacks

Understand when user-mode APCs run, how alertable waits deliver completions, and why APC callbacks require strict thread-affine lifetime and reentrancy rules.

An asynchronous procedure call (APC) is queued to a particular Windows thread and, for a user-mode APC, executes on that thread only when it enters an alertable wait. This is different from a generic task queue: a worker cannot pick up the callback merely because it is idle, and the caller cannot assume that queueing the APC immediately runs it. The target thread’s control flow, wait behavior, and continued lifetime are part of the callback contract.

APCs are useful in APIs that define completion routines, such as ReadFileEx, WriteFileEx, and waitable-timer completion callbacks. They are also a low-level way to request that a thread run a small piece of work in its own context. The same properties make them easy to misuse. A target that never calls an alertable wait can retain queued work indefinitely; a callback can run reentrantly inside a wait operation; and thread-pool worker lifetimes make APC-based signaling a poor substitute for pool-native callbacks.

Per-thread queues and alertable waits

Each thread has its own APC queue. QueueUserAPC requests execution on the selected thread; success means the request was queued, not that the callback has completed. The target thread becomes alertable when it calls APIs such as SleepEx, WaitForSingleObjectEx, WaitForMultipleObjectsEx, SignalObjectAndWait, or MsgWaitForMultipleObjectsEx with alertable behavior enabled.

If an APC is already queued when the target enters an alertable wait, the callback runs instead of the thread remaining asleep. The wait commonly returns WAIT_IO_COMPLETION, which tells the caller that one or more APC routines ran. Code should treat this as a distinct control-flow result and re-evaluate its stop state and predicates before continuing. Do not treat it as if the waited-on event became signaled.

VOID CALLBACK NotifyWorker(ULONG_PTR value)
{
    auto* notification = reinterpret_cast<Notification*>(value);
    // Keep this callback short and use only state that remains valid
    // until the target thread has processed the APC.
    notification->MarkObserved();
}

DWORD WINAPI WorkerMain(void*)
{
    while (!ShouldStop()) {
        const DWORD result = SleepEx(1000, TRUE);
        if (result == WAIT_IO_COMPLETION) {
            // An APC or I/O completion routine ran on this worker.
            // Re-check the worker's state before starting more work.
            continue;
        }
        if (result == 0) {
            PerformTimedMaintenance();
        }
    }
    return 0;
}

The example omits ownership and synchronization around Notification and ShouldStop; production code must define both. A thread handle used by another thread to queue user APCs needs the appropriate access rights, and the notification storage must remain alive until the callback finishes. The target should also be designed to reach alertable waits regularly, or the sender needs a different mechanism with delivery and cancellation semantics suited to its contract.

APC routines introduce reentrancy

An APC callback executes in the target thread, not on a separate callback stack owned by a scheduler. The thread may be paused inside a wait when the callback runs, so code that resumes after the wait must tolerate the callback having modified shared state. Avoid acquiring locks that the target thread already holds across the alertable wait. Avoid blocking I/O, lengthy computation, or calling arbitrary application code from the callback; that delays the waiter’s return and can create hard-to-debug lock cycles.

Keep APC callbacks small: set a flag, append a small item to a thread-owned queue, or complete the exact operation for which the API defines an APC. If the callback must transfer a heap object, use an explicit ownership protocol. A pointer captured by an APC is not protected merely because it was valid when QueueUserAPC returned. The sender and receiver need a rule for cancellation, target exit, and exactly-once cleanup.

Do not use an APC as a general cross-thread interrupt that is expected to safely abort arbitrary code. User-mode APCs are cooperative in the sense that the target must enter an alertable wait. They are not a replacement for cancellation tokens, events, I/O cancellation, or a work queue. A target that is running CPU-bound code without alertable waits will not dispatch the APC until it reaches a suitable point.

APC-based I/O completion

ReadFileEx and WriteFileEx use completion routines that are delivered through the issuing thread’s APC queue. The issuing thread must enter alertable waits to run them. The buffer, OVERLAPPED structure, and callback context must remain valid until the completion routine executes. If the thread blocks forever in a non-alertable wait, the operation may complete but the application callback is not dispatched through that path.

This model can be appropriate when one dedicated thread owns the I/O and already has a well-designed alertable loop. It is generally not the best model for an I/O-heavy server with many concurrent operations. I/O completion ports and thread-pool I/O objects allow completion work to be scheduled through a shared completion mechanism without requiring a specific target worker to enter an alertable wait. Choose one completion model for each operation and do not accidentally wait for both a completion packet and a thread APC that will never run.

Cancellation does not erase the obligation to observe completion and release state safely. An I/O cancellation request can race with normal completion, and the callback may still report success if the request finished first. Keep operation objects alive until the terminal result has been consumed. A timeout in the caller is a policy decision, not proof that the kernel stopped using the buffer.

Waitable timer completion routines

SetWaitableTimer can associate an APC completion routine with a timer. When the timer expires, the routine is queued to the thread that set the timer and runs only when that thread enters an alertable wait. For a periodic timer, Windows queues the completion APC only if no outstanding APC is already queued. If the target stays non-alertable across multiple expirations, those expirations do not create an unbounded APC backlog; the timer is reactivated, while callback delivery is coalesced rather than guaranteed once per elapsed period. Do not use completion-callback count to reconstruct every missed timer interval.

If the only goal is to run periodic background work, a thread-pool timer is usually easier to manage: the pool schedules callbacks and provides operations for disarming the timer and waiting for callbacks to complete. Microsoft’s guidance specifically cautions that APC delivery does not fit thread-pool worker lifetime management, because a pool thread can be retired before a queued APC is delivered. Use APC-based timers only when the target thread and its alertable-wait loop are explicit parts of the design.

Shutdown and error handling

Shutdown must coordinate the producer, the target thread, and any queued callback context. Stop queueing new APCs before freeing the shared data. If the APC itself is intended to wake the target for shutdown, ensure the target will enter an alertable wait and that the callback recognizes the shutdown request. Join the target thread after its callback work has finished, then release the thread handle and associated context. A thread that exits before dispatch is not a successful delivery path for application state.

Check the return value from QueueUserAPC. Failure means the request was not queued, so the sender must choose a fallback or report the failure. Success does not provide an acknowledgement. If acknowledgement matters, have the callback signal an event or update an owned completion object that the sender can observe. Do not use a global boolean without atomic or lock-based synchronization; the callback and sender run concurrently.

For message-pump threads, do not replace the ordinary message loop with a simple alertable sleep. Use an API designed to wait for both GUI messages and kernel objects while enabling APC delivery, and process all return values according to its contract. A thread that owns windows has message-queue liveness requirements in addition to APC delivery requirements.

When APCs are the wrong abstraction

Prefer a work queue or thread-pool callback when the work can run on any worker. Prefer an event when the target needs a durable state change rather than a callback. Prefer I/O completion ports or thread-pool I/O for scalable completion processing. Prefer a cooperative cancellation flag or a documented cancel API over trying to interrupt arbitrary execution. APCs are precise but thread-affine: their strongest use cases are the ones where that affinity is intentional and the target’s alertable waits are part of the architecture.

Test with the callback queued just before and just after the target enters its wait, with the target busy, with shutdown racing delivery, and with repeated queueing. Instrument queue failures, callback start/end, and time spent between queue and dispatch. Those measurements expose the main operational truth of APCs: enqueueing is only a request; alertable control flow is what makes the callback run.

Related:

Sources:

Comments