Linux NUMA Memory Placement: CPU Affinity, Memory Policy, and Locality
Diagnose NUMA locality by separating where threads run from where pages are allocated, then measure policy, cpuset constraints, and remote access.
On a NUMA machine, a memory access has a topology: a CPU reaches some memory through a local controller and reaches other memory through an inter-socket fabric. The cost difference depends on the platform and workload; “remote is always twice as slow” is not a portable rule. The useful operational question is whether a process’s runnable CPUs, allocated pages, and actual access pattern agree well enough for its latency and throughput goals.
CPU affinity and memory policy solve different problems. Pinning a thread constrains where it may execute. A memory policy guides where eligible page allocations are made. Neither setting alone guarantees that every byte is local, and neither measures the performance improvement. Treat placement as an observed relationship among scheduler decisions, page faults, policy scopes, cpuset restrictions, and the application’s sharing pattern.
NUMA nodes are allocation domains, not CPU labels
The kernel describes CPUs and memory in NUMA nodes, but node numbering is an identifier rather than a promise about physical distance or ordering. A machine may have nodes without local memory, asymmetric distances, memory-only nodes, or firmware-described proximity domains that do not correspond to a simple two-socket diagram. Inspect the live topology instead of assuming that node 0 is “the first socket” or that successive IDs are equally distant.
Useful read-only starting points include lscpu -e=CPU,NODE,SOCKET,CORE, numactl --hardware when the numactl package is installed, and /sys/devices/system/node/node*/distance. The distance matrix is a relative topology hint, not a calibrated nanosecond table. Also inspect the online CPU and memory lists: hotplug, virtualization, and boot parameters can make the runtime topology differ from a product brochure.
lscpu -e=CPU,NODE,SOCKET,CORE
for node in /sys/devices/system/node/node[0-9]*; do
printf '%s cpulist=' "$node"
cat "$node/cpulist"
printf 'distance='
cat "$node/distance"
done
The loop is observational. On systems without lscpu, use /proc/cpuinfo together with sysfs, but remember that a processor’s package ID is not itself a complete NUMA distance map. Container namespaces and CPU/memory cgroups may expose a restricted view; compare host and service views before concluding that a node is missing.
First touch and policy scope
Most process memory is virtual until pages are faulted in. For anonymous mappings, the thread that first causes a page allocation commonly influences local placement under the default policy, subject to allocator fallback and allowed-node constraints. This is the source of the “first-touch” technique: initialize a large array in parallel using the same partitioning and CPU placement expected during steady-state work. It is not a universal law for all memory. File-backed page-cache allocations, shared mappings, huge pages, kernel allocations, and memory reclaimed or migrated later have different rules.
Linux supports memory policies with different scopes. A task policy provides a process-level default; a policy attached to a virtual memory area can be more specific; shared memory can have a shared policy. A more specific policy takes precedence for the allocations it governs. Policy selection does not retroactively move all existing pages unless the relevant API requests migration and migration is possible. A program that installs a policy after startup can therefore retain a large old allocation on its original node.
The common modes express intent, not magic capacity. Local allocation prefers the node associated with the allocating CPU and can fall back. Preferred mode expresses a preferred node while allowing fallback. Bind restricts the eligible nodes to a set, but cpuset restrictions intersect with policy and can make the requested set empty or unusable. Interleave distributes eligible allocations across a node set, which can help bandwidth for evenly accessed data but can hurt a latency-sensitive working set. The correct mode follows the access pattern, not a slogan about locality.
CPU placement is an independent control plane
sched_setaffinity() and tools such as taskset constrain where threads may run. NUMA memory policy controls a different decision. A process bound to CPUs on node 1 can still use pages allocated on node 0, especially if those pages were faulted in before affinity changed or are shared with other workers. Conversely, a process can allocate locally and later migrate its threads to another node. Always measure both sides.
For services managed by systemd, CPU placement may come from AllowedCPUs= or a cpuset controller, while allowed memory nodes are separately constrained. A container runtime can impose both. The kernel’s NUMA policy operates only within the memory nodes permitted by the active cpuset; it does not override that administrative boundary. If a requested node is outside the allowed set, inspect the cgroup and effective masks before changing the application’s policy.
taskset -pc "$PID"
cat "/proc/$PID/status" | rg 'Cpus_allowed_list|Mems_allowed_list'
cat "/proc/$PID/numa_maps" | head -n 20
These commands help form a hypothesis, not prove that a hot data structure is local. numa_maps summarizes virtual memory areas and page counts by node; correlate it with the actual allocation and access phase. taskset -pc reports the process affinity mask, but a multithreaded service may need per-thread inspection under /proc/$PID/task/*/status.
Measure placement and remote traffic
Start with a before-and-after baseline under the same workload and duration. Record throughput, tail latency, CPU utilization, context switches, per-node memory use, and the placement of the workload’s major mappings. numastat -p PID is useful when available. Per-node kernel counters under /sys/devices/system/node/node*/numastat report page counts, while /proc/PID/numa_maps exposes mapping-level residency and policy hints. Counter names such as numa_miss and numa_foreign describe allocation outcomes relative to preferences; they are not direct measurements of cache misses or all remote reads.
Memory-residency counters are not a substitute for a performance counter or a workload-level profile. A page can be resident on a remote node while rarely accessed, or local in aggregate while a small hot region dominates tail latency. Where platform PMUs expose remote-memory events, verify their meaning in the processor vendor documentation and check whether virtualization permits them. Use application-level latency distributions to determine whether a placement change matters.
Avoid comparing runs that differ in warm-up, page cache, frequency state, memory pressure, or worker count. A clean experiment has a known initial allocation phase, repeatable input, fixed CPU masks, and enough repetitions to see noise. If only one node’s memory is pressured, a policy can trigger fallback or allocation failure that does not appear in a small synthetic run.
Safe policy experiments
Run policy changes on a dedicated test process first. numactl --cpunodebind=0 --membind=0 command is a convenient experiment when numactl is installed, but it is deliberately restrictive: if node 0 is not available to the task or cannot satisfy allocations, the command may fail or experience pressure. --preferred=0 and --interleave=all describe different tradeoffs; do not substitute one for another without deciding whether latency, bandwidth, or capacity is the objective.
For application code, check every memory-policy system call’s return value and log the requested mask plus the effective allowed-node mask. For a mbind() experiment, account for alignment, range length, mapping type, and whether existing pages are to be moved. Migration can fail because pages are pinned, shared, unevictable, or otherwise not migratable. A successful policy call is not proof that every requested page moved.
Use numactl --show in the launched process to confirm inherited task policy when the utility is available. Then inspect node residency after the workload has touched the relevant mappings. Keep a rollback path: terminate the test process and remove the test service’s affinity or policy override rather than trying to repair a live production allocation by repeated policy changes.
A practical acceptance matrix
Validate at least four cases: default local allocation; worker threads pinned to their intended nodes before allocation; policy installed after allocation; and the production cpuset restrictions. Include a shared-data case if workers communicate through common mappings. For each, capture effective masks, per-node residency, CPU placement, throughput, and p50/p95/p99 latency. Test a memory-pressure case to observe fallback behavior, and verify what happens when one target node lacks capacity.
An acceptance threshold should state the workload and metric, for example, “the 99th-percentile request latency improves by at least the agreed amount without reducing throughput or exhausting node 0 memory.” Do not accept a tuning change because a single NUMA counter becomes zero. Also test restart and deployment behavior: service managers, container runtimes, and orchestration layers can replace manual affinity settings at the next rollout.
The most robust optimization is often to align data partitioning and worker ownership, not to bind the entire process to a node. That design limits cross-node sharing while preserving scheduler flexibility. Keep explicit policy only where measurement justifies it, document its assumptions, and revisit the result when hardware, kernel, or workload topology changes.
Related:
- Control Groups (cgroup v2) Explained: Limiting and Accounting for Resources
- Linux Memory Hotplug: Online, Offline, and Capacity-Change Failure Modes
Sources: