FreeBSD CPU Sets and NUMA: Affinity Without False Isolation
Use FreeBSD cpuset masks and NUMA domain policies deliberately, verify effective placement, and measure whether affinity improves real workloads.
CPU pinning is easy to request and surprisingly easy to misinterpret. On FreeBSD, cpuset(1) can constrain processes, threads, jails, interrupt sources, and other targets to CPU masks. On NUMA systems it can also select memory domains and an allocation policy. These controls are useful for repeatable benchmarks, cache-sensitive services, and locality experiments, but a mask is not automatically an exclusive reservation, a CPU quota, or proof that memory is local.
The reliable approach is to understand the layers, inspect the machine’s actual topology, change the smallest scope possible, and compare measurements against an unpinned baseline. This guide uses the current cpuset(1) and NUMA interfaces without assuming CPU IDs, domain numbering, or performance gains that vary by hardware.
Three masks, not one magic affinity value
FreeBSD’s cpuset(1) manual describes two sets applicable to each process plus a private CPU mask for each thread. A process belongs to a numbered cpuset. Its threads also have masks that must be subsets of the set’s allowed CPUs. The immutable root set, numbered zero, represents all possible CPUs in the system. By default, ordinary processes start in set 1.
This hierarchy explains why two commands that look similar have different scope. cpuset -l ... -p PID adjusts the target’s private CPU mask. Adding -c applies the operation to the cpuset associated with the target, which can affect other processes in that set. The distinction is operationally important: a per-process experiment is usually safer than changing a shared parent set or the default set. A mask limits where the selected thread may run; it does not make those processors unavailable to other workloads.
The scheduler still chooses when runnable threads execute. ULE documents thread CPU affinity, topology awareness, per-CPU run queues, and interactivity heuristics. Pinning narrows scheduler choices; it does not disable scheduler behavior or guarantee a fixed percentage of CPU time. For actual resource ceilings and enforcement, use the appropriate resource-control mechanism separately. Affinity is placement eligibility, not a budget.
Inventory before changing placement
CPU numbering and NUMA topology differ across firmware, platforms, virtual machines, and kernel configurations. Start with the root CPU mask and the current process mask, then inspect the memory-domain data that the host exposes:
cpuset -g -r
cpuset -g -p $$
sysctl vm.ndomains vm.phys_locality vm.phys_segs
cpuset -g -r queries the root set; cpuset -g -p displays the current shell’s allowed CPU mask. vm.ndomains reports the number of detected VM memory domains, while vm.phys_locality and vm.phys_segs expose locality and physical-memory grouping information when available. A one-domain system still benefits from CPU affinity experiments, but changing NUMA policy there is unlikely to demonstrate cross-domain locality effects.
Use the domain query supported by cpuset(1) to inspect CPU visibility for an actual domain:
cpuset -g -d 0
Domain 0 is only an example; use IDs reported by the host. FreeBSD’s numa(4) manual describes a VM-focused domain numbering model. Those identifiers are not necessarily identical to platform-specific hardware numbering. Map CPU IDs to domains using the target system’s output and SMP(4) guidance, not by assuming that CPU 0 and memory domain 0 are always the closest pair.
Capture the current process and set identifiers before making a reversible change:
cpuset -g -p "$pid"
cpuset -g -i -p "$pid"
The first command queries the target process mask; the second reports its set ID. Record both, along with the service manager and restart procedure. Commands that change an existing process may require appropriate credentials and may fail if the requested mask is invalid or outside the process’s permitted set.
Apply the narrowest useful CPU mask
For a new benchmark process, the manual’s launch form creates a new group and starts the command in it:
cpuset -c -l 4-7 ./worker --config ./benchmark.conf
The example assumes CPUs 4 through 7 exist and are appropriate; discover the mask first. A comma-separated CPU list is also accepted. This method keeps the experiment’s placement context attached to the launched workload instead of changing the machine-wide default set.
For an existing PID, a private mask can be changed with:
cpuset -l 4-7 -p "$pid"
cpuset -g -p "$pid"
The second command is an immediate verification, not a performance test. When a target has several threads, inspect the relevant thread IDs and effective masks if the application creates threads dynamically. A multithreaded benchmark can behave differently if only one thread or one shared set was changed. Prefer testing through the service’s normal launch path so child threads and restarts inherit the intended process context.
Changing the containing set has a wider blast radius:
cpuset -l 4-7 -c -p "$pid"
The manual uses this form to modify the cpuset to which the process belongs. Before doing this in production, determine which other processes share that set and what happens to them if the mask excludes CPUs they need. Avoid modifying set 1 or an interrupt target during an application experiment. Although cpuset(1) can target IRQs, driver and hardware behavior vary; interrupt placement is a separate change that should be measured and have a rollback plan.
CPU locality and memory locality are separate decisions
On a NUMA machine, a CPU mask can keep threads near a selected processor group while a domain policy governs where eligible memory is allocated. cpuset(1) accepts a -n policy:domain-list option. Policies include first-touch, round-robin, prefer, and interleave. First-touch allocates on the local domain when memory is available; prefer prioritizes one selected domain and consults the parent set if it is unavailable; interleave distributes allocations with an implementation-defined stripe width. Their exact behavior and trade-offs are described in domainset(9).
A launch-time example can constrain both CPU eligibility and memory-domain policy:
cpuset -c -l 4-7 -n first-touch:0 ./worker --config ./benchmark.conf
This is not a universal recommendation to use domain zero. It is valid only if the selected IDs exist and the workload’s CPU and memory access pattern match that policy. First-touch can help when the thread that initializes a page is also the thread or domain that will use it. It can be counterproductive when one thread initializes a large shared heap and many other CPUs access it, or when work migrates across domains. Interleaving can improve aggregate bandwidth for some access patterns while increasing remote access for a latency-sensitive single-thread workload. Benchmark both policies with representative data and concurrency.
Changing a future allocation policy should not be confused with relocating every page that the process already touched. Existing mappings, shared file-backed pages, allocator behavior, and pages first accessed by another thread can all affect observed locality. Plan experiments from process start, control initialization order, and report cold-start and steady-state results separately. numa(4) notes that FreeBSD does not keep statistics indicating how often NUMA policies succeed or fail, so a configured policy alone is not evidence that allocations landed as intended.
Measure the effect instead of trusting the mask
Before pinning, define a workload and collect a baseline with the same binary, data, load, power state, and duration. Record throughput, latency distribution, CPU utilization by processor, context switching where available, and memory footprint. Then run one placement change at a time. Repeat enough runs to distinguish improvement from noise, and include tail latency rather than only an average.
Check for common failure modes: a mask that omits a CPU the application expected, a narrowed set shared with unrelated processes, a hot thread pinned onto an oversubscribed SMT sibling, worker threads spread across memory domains despite a locality goal, or interrupts competing with a busy polling thread. A process can report the intended affinity and still fail to improve because it was I/O-bound, lock-bound, or memory-latency-bound. Likewise, scheduler topology awareness already influences normal placement, so manual pinning can remove a useful scheduler choice.
Compare an unconstrained run, a CPU-only mask, and only then a CPU-plus-memory policy. Keep system load and CPU frequency behavior in the record; power management can change the result even though CPU affinity did not. Do not compare a cold first-touch run against a warmed baseline or infer locality from CPU utilization alone. Use application timing and, where available, hardware performance counters that are supported on the actual processor.
Rollback and acceptance criteria
For a test launched under cpuset, stopping it and relaunching it through the original service path is usually the cleanest rollback. For an existing process, restore the recorded mask and set ID, and verify with cpuset -g before declaring the change reverted. If the containing set was changed, restore its prior CPU mask as well; restoring only the process’s private mask may leave its parent set narrowed. Avoid speculative all changes because “all” refers to the root set and can expand a workload beyond the allocation policy intended by an administrator.
Accept the change only if it is reproducible on the target hardware, improves the preselected service metric without violating latency or throughput objectives, survives a restart, and has an understood rollback. Verify the effective CPU mask after launch and again after the service manager replaces the process. For NUMA changes, test the application’s real allocation and access pattern; the policy string is configuration evidence, not proof of placement.
FreeBSD cpusets are precise controls over scheduler eligibility and memory-domain policy. They are not exclusive CPUs, a CPU quota, a cache partition, or an automatic NUMA optimizer. Treat the masks as one part of a measured workload design, preserve the default scheduling behavior unless evidence says otherwise, and make every wider-scope change explicit.
Related:
- FreeBSD CPU Power Management: Driver-Aware Frequency Tuning in Production
- FreeBSD RCTL: Resource Limits for Users, Jails, and Processes
Sources: