Skip to content
WindowsDeep Dive Published Updated 7 min readViews unavailable

Windows Thread-Pool Waits: Re-Registration and Safe Handle Teardown

Register a Windows thread-pool wait on one event, re-arm it deliberately, and cancel queued callbacks before closing handles or freeing callback state.

The Windows Vista-era thread-pool API can monitor a waitable handle and invoke a callback when that handle is signaled or a timeout expires. CreateThreadpoolWait creates the callback object; SetThreadpoolWait registers a single handle and optional timeout. This lets an application react to an event without dedicating one sleeping thread per handle, but it does not transfer ownership of the handle or callback context to the pool.

Thread-pool waits are one-shot registrations. After a handle signals and the callback runs, the application must register the wait again before a later signal can queue another callback. That explicit re-registration is easy to miss when porting code from the older registered-wait API. It also makes the lifecycle visible: the owner can stop future callbacks, cancel queued work, wait for running callbacks, and only then release the handle and context.

Register one wait at a time

A PTP_WAIT can watch one waitable object, not a set of handles. A mutex is not supported as the handle for this API. If the component must watch multiple events, use multiple wait objects or a different wait design. A finite timeout is supplied as a FILETIME; NULL means no timeout. A relative timeout is encoded as a negative 100-nanosecond value, while a positive value denotes an absolute time since January 1, 1601 UTC.

The wait callback receives the wait result. It should distinguish a signaled object from a timeout and from other result values. Do not infer success merely because the callback ran. The callback may also be queued just as shutdown begins, so it must validate the owning component’s state before acting on the event.

#include <mutex>

struct WaitState {
    PTP_WAIT wait = nullptr;
    HANDLE event = nullptr; // Auto-reset event, owned by the component.
    SRWLOCK lifecycleLock = SRWLOCK_INIT;
    std::mutex callbackLock;
    bool stopping = false; // Read and written under lifecycleLock.
};

VOID CALLBACK EventWaitCallback(
    PTP_CALLBACK_INSTANCE,
    PVOID raw,
    PTP_WAIT wait,
    TP_WAIT_RESULT result)
{
    auto* state = static_cast<WaitState*>(raw);

    // Re-registration and a rapid subsequent signal can overlap callbacks.
    std::lock_guard<std::mutex> workGuard(state->callbackLock);
    AcquireSRWLockExclusive(&state->lifecycleLock);
    const bool stopping = state->stopping;
    ReleaseSRWLockExclusive(&state->lifecycleLock);

    if (!stopping) {
        if (result == WAIT_OBJECT_0) {
            ProcessNotification();
        } else if (result == WAIT_TIMEOUT) {
            HandleDeadline();
        } else {
            ReportUnexpectedWaitResult(result);
        }
    }

    AcquireSRWLockExclusive(&state->lifecycleLock);
    if (!state->stopping) {
        // Every subsequent notification requires a new registration.
        SetThreadpoolWait(wait, state->event, nullptr);
    }
    ReleaseSRWLockExclusive(&state->lifecycleLock);
}

The sample assumes ProcessNotification and HandleDeadline are safe to call on a thread-pool worker, and that WaitState remains alive until all callbacks have drained. In production, check and record every registration failure where the API provides a result; SetThreadpoolWait is void, so validate the handle before calling it and surface failures through the callback or component state. Do not perform UI-thread-affine work directly in this callback.

The lifecycle lock serializes re-registration against shutdown. Without that rule, shutdown could unregister the wait and a callback could race afterward to register it again. It is held only while checking or changing component state, not while doing work or waiting for callbacks. The separate work lock prevents callback bodies from overlapping if a rapid subsequent signal queues another callback. A callback already in progress may finish after shutdown sets stopping; the owner retains its context until the drain completes.

Signal semantics and re-arming

SetThreadpoolWait replaces any previous wait registration for that object. Passing a null handle stops it from queueing new callbacks, but callbacks already queued can still occur. A signaled event does not mean the callback has completed, and a timeout callback does not imply that the event is unsignaled. The event’s own manual-reset or auto-reset behavior determines how a signal is consumed.

For an auto-reset event, a successful wait normally consumes one signal. For a manual-reset event, the signal remains set until the application resets it; re-registering while it remains signaled can cause another callback immediately. Choose event semantics with the producer and callback together. If every notification is a distinct item, use a protected queue and treat the event as a wake hint; a manual-reset event can be paired with a “queue nonempty” predicate. If notifications can be coalesced into one current-state update, an auto-reset event may be sufficient.

A timeout can also be used to implement periodic checks, but do not turn the callback into a polling loop by re-registering an immediate timeout. Define whether each timeout is relative to the time of registration or tied to an absolute deadline. A long callback delays the next registration because this design re-arms after work completes; if fixed-rate behavior is required, track the intended deadline separately and compute the next timeout.

Handle lifetime is part of the registration

Do not close the event while it remains registered. Microsoft’s documentation states that closing a handle while the wait is pending produces undefined behavior. The wait object does not duplicate ownership in a way that makes it safe for the application to discard the original handle. Keep the event open until the registration is canceled and any callbacks that can use it are complete.

Do not free the WaitState passed as the callback context while a callback is queued or running. The thread pool stores and later passes that pointer to the callback; it does not copy the pointed-to object. A callback already queued before unregistration still has the right to execute unless the pending callback is canceled and drained.

A shutdown sequence that closes the race

Use a lifecycle lock or another serialized state transition to prevent a callback from re-registering after shutdown begins:

  1. Acquire the lifecycle lock and set stopping = true.
  2. Call SetThreadpoolWait(wait, nullptr, nullptr) to stop new registrations, then release the lock.
  3. Call WaitForThreadpoolWaitCallbacks(wait, TRUE) to wait for outstanding callbacks and cancel queued callbacks that have not begun.
  4. Call CloseThreadpoolWait(wait) after the callback boundary is quiescent.
  5. Close the event handle and release WaitState only after no callback can use them.

The callbacks that have already started are not forcibly terminated by the wait API. They must be able to finish, and the component must not hold a lock they need while waiting for them. Never perform this blocking sequence from a callback associated with the same wait object; it can wait for itself. A cleanup group can centralize the close of related wait, work, and timer objects, but the producer still needs to stop creating new work before the group is drained.

CloseThreadpoolWait alone is not a sufficient “no more callbacks can run” guarantee in every case. Microsoft documents that a callback may run after close; to prevent that behavior, unregister with a null handle, call WaitForThreadpoolWaitCallbacks with cancellation of pending callbacks, and then close the object. This ordering is what lets the owner release the callback context safely.

Worker-thread assumptions

A thread-pool callback can run on different worker threads at different times. Do not use thread identity as the lifetime of the wait object, and do not leave thread-local state, impersonation, COM registration, or priority changes behind after the callback returns. A wait callback must be thread-pool safe: it cannot assume a dedicated persistent thread or an STA message pump.

If the callback starts an asynchronous operation whose completion requires the same thread to remain alive or enter an alertable wait, use a dedicated thread or a thread-pool primitive designed for that operation. Do not make a callback wait synchronously for work whose completion is queued to the same constrained pool without accounting for pool starvation. Keep callback work bounded and use the wait only as an event-to-work bridge.

Verification under shutdown races

Stress the sequence with repeated event signals while shutdown toggles the stopping state. Count callback starts and finishes; after the drain returns, assert that the counts match and that no callback observes freed context. Race timeout expiry with explicit event signaling, test event reset policy, and make a callback deliberately block on a controllable event to prove the owner does not hold its lifecycle lock while draining.

Log registration changes, callback result codes, stop requests, and drain duration. Avoid logging each high-frequency signal if that materially changes scheduling. If teardown stalls, identify which callback remains active and inspect its lock, I/O, and nested wait dependencies. Thread-pool waits are an efficient notification mechanism when their one-shot re-registration and callback lifetime are explicit. The safe design treats registration as state, signals as hints, and callback drain as the boundary before ownership ends.

Related:

Sources:

Comments