Windows CPU Sets and Processor Groups: Affinity Without Topology Assumptions
Discover Windows processor groups and CPU Sets safely, then choose soft scheduling preferences or hard affinity only when measurements justify it.
Processor affinity code often starts with a correct observation and an incorrect conclusion: a machine has many logical processors, therefore a thread should be pinned to one. Windows exposes several related mechanisms, but they express different constraints. Processor groups model systems with more than 64 logical processors. CPU Sets provide scheduler-aware placement preferences. Legacy affinity masks can impose a harder restriction. Treating these as interchangeable can silently strand threads, conflict with power management, or make an application perform worse on a newer Windows release.
For most applications, the best first decision is to leave placement to the scheduler and measure the workload. Use the topology APIs to describe and observe the machine, and add placement policy only when a repeatable benchmark or latency requirement justifies the complexity.
Processor groups are topology, not processor numbers
Windows partitions logical processors into processor groups, each containing at most 64 logical processors. A processor is therefore identified by a group number and a processor number within that group. A plain DWORD_PTR affinity mask can represent only one group’s processors; it is not a machine-wide CPU bitmap on a large server.
The defaults have changed. On Windows versions before Windows 11 and Windows Server 2022, a process was generally assigned to one processor group by default. An application that wanted to use processors in other groups needed to opt in with group-aware APIs. Starting with Windows 11 and Windows Server 2022, processes and threads can span groups by default on large systems. Older affinity APIs retain primary-group semantics for compatibility. Code that equates processor number zero with one globally unique CPU can therefore select the wrong processor or report misleading telemetry.
Query the actual topology instead of inferring it from a count returned by an unrelated API. The following small C++ program reports the active group count and the number of active logical processors in each group:
#include <windows.h>
#include <stdio.h>
int main() {
const WORD groups = GetActiveProcessorGroupCount();
if (groups == 0) {
fprintf(stderr, "No active processor groups were reported.\n");
return 1;
}
for (WORD group = 0; group < groups; ++group) {
const DWORD processors = GetActiveProcessorCount(group);
printf("Group %u: %lu active logical processors\n",
static_cast<unsigned>(group),
static_cast<unsigned long>(processors));
}
return 0;
}
This is inventory, not proof that a particular process can use every processor. Process and thread affinity, job-object limits, CPU Set reservations, processor state, and the OS version all affect effective placement. Record the Windows build and the process’s own policy alongside topology output when investigating a discrepancy.
Hard affinity and CPU Sets answer different questions
SetThreadGroupAffinity and SetThreadAffinityMask constrain execution to a selected processor or group-relative mask. Such a constraint can be appropriate for a carefully measured device-interrupt relationship, a real-time-like workload with a controlled deployment envelope, or compatibility with a component that has a documented processor-local assumption. It can also make a thread ineligible to run on otherwise idle CPUs. Hard pinning is not a generic performance optimization.
CPU Sets are a more scheduler-friendly mechanism for expressing placement preferences. An application can query the system’s SYSTEM_CPU_SET_INFORMATION records, select CPU Set IDs, and apply them to a thread with SetThreadSelectedCpuSets or as process defaults with SetProcessDefaultCpuSets. CPU Sets work across groups and are intended to coexist with operating-system placement and power-management decisions. They are not a promise that a thread is permanently pinned to one physical core: the scheduler may choose among allowed Sets, and restrictive affinity masks still take precedence where they conflict.
CPU Set records also carry topology and allocation information, including the group and logical processor relationship and flags that indicate whether a Set is allocated for exclusive use. Do not assume ID values are dense, stable across boots, or equal to a processor index. Enumerate fresh information at runtime and treat a CPU Set ID as an opaque identifier rather than deriving it arithmetically.
The common buffer-query pattern for GetSystemCpuSetInformation is a size probe followed by an allocation and a second call. The API returns variable-length records; a production enumerator must advance by each record’s documented Size, validate every boundary against the returned byte count, and reject malformed or truncated data before dereferencing a record. Do not parse the buffer as a fixed array of structures or assume one record per active CPU without checking its Type.
Choose the least restrictive policy that meets the objective
Start with no explicit placement, then compare a baseline against the proposed policy under a representative load. Keep the process affinity mask, job-object CPU limits, and any CPU Set assignments in the test notes. A per-thread CPU Set selection can be useful when a worker pool has distinct latency and throughput roles, but assigning every worker to a tiny Set can create a queue behind one logical processor. On SMT systems, two logical processors may share execution resources; a count of CPUs is not equivalent to a count of independent cores.
If the application has a hard affinity requirement, use the group-aware APIs and keep the group number attached to every processor identifier. When creating a thread, Windows supports specifying group affinity through the extended startup attribute PROC_THREAD_ATTRIBUTE_GROUP_AFFINITY; after creation, SetThreadGroupAffinity can change it. Validate the return value and preserve the previous affinity if the operation must be rolled back. For processes with threads in multiple groups, do not assume SetProcessAffinityMask can express a machine-wide policy.
Job objects add another policy layer. A job can define processor-group information or CPU-rate controls that constrain processes independently of their own affinity calls. If placement seems ignored, inspect the containing job and the process’s effective limits before changing application code. Containers, virtual machines, and hosted services may also expose a processor topology different from the physical host.
Measure outcomes, not affinity settings
Collect throughput, tail latency, CPU time, context switches, ready time, and power state for the same workload before and after a change. Include cold and steady-state runs where the application has both. A faster microbenchmark on an idle developer workstation is not enough to justify pinning a production service whose CPU count, firmware, core layout, virtualization layer, and Windows version differ.
Test process startup, thread creation, shutdown, and topology changes as well as the hot loop. A worker pool created before policy is applied may inherit a different group or CPU Set state than one created afterward. Rebuild or deliberately update the pool when changing policy; do not rely on undocumented inheritance behavior. Include a graceful fallback when the requested CPU Set is unavailable or reserved, and emit a diagnostic that names the policy and OS version without dumping unrelated process data.
Affinity can also obscure a bottleneck. If a pinned worker consumes a full processor while other CPUs are idle, first check lock contention, serialized queues, blocking I/O, garbage collection, and thread-pool starvation. Moving the same serialized work to another core cannot make it parallel. Keep placement decisions reversible through configuration, and turn them off if a workload or deployment topology changes.
Operational checklist
- Record Windows edition, build, active processor groups, and group-relative CPU counts.
- Inspect process and thread affinity plus any job-object CPU controls.
- Establish an unpinned baseline under representative load.
- Prefer CPU Sets for soft placement preferences; use hard affinity only for a demonstrated requirement.
- Handle API failure and unavailable or reserved Sets without assuming the request was honored.
- Re-test after OS, firmware, VM-size, or application worker-pool changes.
The important design rule is to keep the scheduler informed rather than fighting it. Processor groups make large topologies addressable; CPU Sets make placement intent expressible; neither is a substitute for a measured bottleneck analysis.
Related:
- Windows Process and Thread Internals: Handles, Tokens, and Objects
- Windows Performance Recorder and Analyzer: A Defensible Trace Workflow
Sources: