Windows I/O Rings: Capability Checks, Queue Backpressure, and Completion Lifetime
Build Windows I/O ring code around runtime capability checks, bounded submission queues, completion correlation, cancellation races, and orderly teardown.
Windows I/O rings expose a submission queue and a completion queue for supported asynchronous file operations. An application builds queue entries, submits a batch to the operating system, and later consumes completion entries that carry both a result and application-supplied correlation data. The design can reduce per-operation transition and dispatch overhead for workloads that have many concurrent operations, but it does not eliminate the need to manage file handles, buffers, queue capacity, cancellation, errors, or shutdown.
Do not treat an I/O ring as a faster synonym for an I/O Completion Port (IOCP). IOCP is a mature completion notification mechanism for overlapped operations across supported handle types. An I/O ring is a separate API with its own operation support, queue-builder calls, API versions, and per-system limits. Choose it only when the target operations are supported on the systems you deploy and measurements show the queue model is beneficial. Maintain a tested fallback where the application must also run on systems that do not expose the required capability.
The I/O ring API surface is versioned and capability-dependent. Microsoft documents QueryIoRingCapabilities, CreateIoRing, GetIoRingInfo, IsIoRingOpSupported, operation builders such as BuildIoRingReadFile, SubmitIoRing, PopIoRingCompletion, and CloseIoRing. A successful process launch or supported Windows version alone does not prove a particular opcode is usable. Ask the running system, then select a path the program can actually operate.
Probe the running system instead of hard-coding a release number
Call QueryIoRingCapabilities at process startup or when initializing the I/O subsystem. Its result includes the maximum supported API version, maximum submission and completion queue sizes, and feature flags. Microsoft notes that capability results should not be persisted beyond the current process or assumed to be identical between process runs. Even on the same OS, deployed hardware, virtualization, policy, or servicing state can affect observed capability.
Create a ring with a version no greater than the reported maximum and queue sizes within the documented limits. The system can round requested queue sizes to powers of two. Completion capacity is rounded to a value at least twice the actual submission queue size, so do not assume the sizes you requested are the actual sizes. GetIoRingInfo provides the created ring’s version, creation flags, and actual queue sizes. Record these values in diagnostics so an unsupported-operation report can be tied to what the process actually received.
The following small C++ program exercises capability discovery, creates a bounded ring, queries the effective queue information, and releases the ring with the correct API. It intentionally submits no I/O, making it a deployment smoke test rather than a throughput benchmark.
#include <windows.h>
#include <ioringapi.h>
#pragma comment(lib, "Kernel32.lib")
int wmain()
{
IORING_CAPABILITIES capabilities{};
HRESULT hr = QueryIoRingCapabilities(&capabilities);
if (FAILED(hr) || capabilities.MaxVersion == IORING_VERSION_INVALID)
return 1;
IORING_CREATE_FLAGS flags{};
flags.Required = IORING_CREATE_REQUIRED_FLAGS_NONE;
flags.Advisory = IORING_CREATE_ADVISORY_FLAGS_NONE;
HIORING ring{};
hr = CreateIoRing(capabilities.MaxVersion, flags, 8, 16, &ring);
if (FAILED(hr))
return 2;
IORING_INFO info{};
hr = GetIoRingInfo(ring, &info);
const HRESULT closeHr = CloseIoRing(ring);
if (FAILED(hr) || FAILED(closeHr))
return 3;
return info.SubmissionQueueSize >= 8 ? 0 : 4;
}
Build against a Windows SDK that declares the I/O ring interfaces and run on each supported target build. This program does not test every opcode, file system, storage device, or deployment policy. Add a runtime operation probe after creation for each operation the product uses; use the API’s documented result convention and test it on the exact SDK and Windows versions in the support matrix. On Windows targets where the SDK or runtime is unavailable, that native compile/runtime validation cannot be replaced by a Linux or macOS build.
Model the queue as bounded work, not an infinite buffer
Builder calls add entries to the submission queue. The queue is finite. If a builder returns IORING_E_SUBMISSION_QUEUE_FULL, the application must submit already built entries and allow enough completions to be consumed before building more. Treat this as normal backpressure, not as a reason to increase queue sizes without limit. Oversized queues consume memory and can increase tail latency when an application accumulates work faster than the storage layer can complete it.
Define a maximum number of in-flight requests per ring and per file or service. Track queued, submitted, and completed counts separately. If producers can build work concurrently, serialize the ring operations or put them behind a bounded producer queue whose capacity and cancellation behavior are explicit. The ring’s presence does not automatically make concurrent calls from arbitrary threads safe under every wrapper or application ownership model.
Submission is a distinct state transition. SubmitIoRing sends constructed entries to the kernel and can optionally wait for a specified number of completion entries. A successful submission does not mean every file operation succeeded; each operation reports its own HRESULT in a completion. Conversely, an error from SubmitIoRing other than the documented wait-timeout case means the queue was not fully processed and constructed entries remain queued. Handle the queue state deliberately before retrying so that the same operation is not built twice.
Use a unique user-data value for each operation when the application supports cancellation or needs correlation. Maintain a map from user-data ID to operation metadata, including the file, byte range, buffer ownership, deadline, and application request. Remove the entry only when the original operation’s completion has been consumed and processed. Reusing an identifier while an earlier request is active can make a cancellation target the wrong operation or make telemetry ambiguous.
Keep every referenced resource alive through completion
BuildIoRingReadFile accepts a file reference, a buffer reference, the requested length and offset, a user-data value, and submission flags. The supplied buffer must be large enough for the requested byte count. A buffer and file reference are not copies of their underlying resources. Keep the file handle open and the memory stable until the operation completes, including when cancellation was requested. Moving or freeing an in-flight buffer can turn an ordinary cancellation race into a memory corruption defect.
Open files with access, sharing, and flags appropriate to the workload. The operation’s offset and length must be valid for the file and device. Handle short reads and end-of-file explicitly; the completion’s Information field is operation metadata, not a promise that the entire requested size was transferred. Validate each operation result, not just the ring-level SubmitIoRing result. If an application registers file handles or buffers for repeated reference by index, model the registration itself as an asynchronous operation and keep the registration data valid until its completion.
The required order for teardown follows directly from this lifetime rule: stop admitting new requests, decide whether to drain or cancel submitted work, consume the completion for every request whose memory or handle will be released, then close or deregister the resources in the correct order. CloseIoRing releases ring resources and abandons entries that were built but never submitted; it does not cancel operations already in flight. Microsoft explicitly warns that reads or writes to buffers may continue after CloseIoRing returns unless outstanding work has been completed or canceled and observed.
Consume completions without losing identity
PopIoRingCompletion removes one completion entry when one is available. S_OK means the caller received a completion; S_FALSE means the completion queue is empty. Drain the queue until it is empty before sleeping for a new event. If the application registers a completion event, the kernel signals it when the queue transitions from empty to non-empty. An event is a wakeup mechanism, not a count of completed operations: one signal can correspond to multiple entries, so always drain all available completions after waking.
Completion processing should be idempotent. A completion can report a file error even when submission was successful. Convert HRESULT and any operation-specific information into a structured result, release the associated buffer only after completion, and fulfill the client request exactly once. Keep user-data correlation independent of a raw pointer to a request object that may already have been freed. Use a generation counter or stable request ID if identifiers can be reused after the old request is fully retired.
Keep the completion thread responsive. Do not perform expensive parsing, UI work, synchronous network calls, or arbitrary callbacks while holding the ring’s central lock. Copy the completion metadata, release the queue lock, then hand the result to a bounded worker or event-loop task. If completions arrive faster than workers can process them, apply backpressure and expose queue depth and age metrics rather than allowing unbounded memory growth.
Treat cancellation as a request with a race
BuildIoRingCancelRequest attempts to cancel a previously submitted operation. It is not a synchronous guarantee that the original request stopped. The I/O can finish before cancellation is processed, and the cancellation’s own completion can arrive after the original operation’s completion. Microsoft advises checking the original operation’s completion queue entry for its final status. A successful cancel-request submission does not prove that the original operation completed as canceled.
Give the original operation and cancellation request distinct user-data IDs. Keep their metadata separate. When the caller times out or closes its logical request, mark that request as no longer deliverable, but keep the memory and file reference alive until the I/O completion is observed. If a read completes successfully after the caller timed out, discard or route the data according to the API contract; do not write it into a buffer that has been reassigned to another request.
Cancellation should be tested under races: cancellation before submission, cancellation while in flight, completion immediately before cancellation, ring shutdown, file-handle closure, and repeated cancel attempts. Define what the API does when the operation is already complete and ensure the caller receives one terminal outcome. Avoid treating a timeout from SubmitIoRing as cancellation; it means the wait expired after submission and the original work can remain active.
Handle errors and unsupported operations at the boundary
Check HRESULTs from capability discovery, ring creation, builders, submission, and close. Check each completion independently. If a required opcode is unsupported, use a fallback rather than repeatedly submitting a known-incompatible operation. If the queue is full, submit and drain; if the file handle is invalid, fix ownership and lifetime; if the storage path is slow, measure device and file-system latency before attributing the problem to queue semantics.
Keep diagnostics bounded and safe. Log ring version, queue sizes, operation code, unique request ID, requested bytes and offset, elapsed time, result HRESULT, bytes reported, cancellation state, and queue depth. Avoid logging file contents, credentials, or sensitive path segments unless approved. Add counters for builder-full events, submit failures, empty polls, wait timeouts, canceled requests, and operations that complete after caller timeout.
Benchmark against an equivalent IOCP or other supported path using the same storage, data, concurrency, and application-level work. Measure throughput, CPU, memory, latency percentiles, queue depth, short I/O, errors, and cancellation responsiveness. An I/O-ring microbenchmark that measures only the queue entry path but skips buffer validation, completion processing, and application logic is not evidence that the production service improved.
I/O ring implementation checklist
- Query capabilities at runtime and create a ring with a supported version and bounded queue sizes.
- Probe every required opcode; preserve a fallback for targets that do not support it.
- Keep file handles, buffers, IDs, and request metadata alive until the original completion is consumed.
- Treat queue-full and wait-timeout outcomes as backpressure and pending work, not as operation success or cancellation.
- Drain completion entries until the queue is empty; process per-operation HRESULTs and byte counts.
- Model cancellation races and do not free resources until all outstanding I/O has reached a terminal state.
- Stop producers, drain or cancel in-flight work, then call CloseIoRing; do not assume ring close cancels submitted I/O.
- Compare the full production path with the existing IOCP implementation on Windows targets.
The core engineering challenge is lifecycle discipline. An I/O ring can batch operations, but the application still owns the correctness contract around each request, buffer, file, completion, and cancellation. Treat those contracts as explicit state transitions and the queue becomes measurable rather than mysterious.
Related:
- Windows I/O Completion Ports: Building Correct Overlapped-I/O Workers
- CancelIoEx: Correct Cancellation and Completion for Windows Overlapped I/O
Sources:
- I/O ring APIs - Microsoft Learn
- QueryIoRingCapabilities - Microsoft Learn
- CreateIoRing - Microsoft Learn
- BuildIoRingReadFile - Microsoft Learn
- SubmitIoRing - Microsoft Learn
- PopIoRingCompletion - Microsoft Learn
- BuildIoRingCancelRequest - Microsoft Learn
- CloseIoRing - Microsoft Learn
- IORING_CREATE_FLAGS - Microsoft Learn