Linux cgroup v2 I/O Control: Weights, Limits, Latency, and Verification
Operate Linux cgroup v2 I/O controls with io.weight, io.max, and io.latency while validating device mapping, writeback accounting, and workload impact.
The cgroup v2 I/O controller can distribute or constrain block-I/O work among workload groups. It is not a universal “make this process faster” switch: weights influence competition, maximums cap consumption, and latency targets can throttle peers when a protected group misses its target. These controls operate in a hierarchy, depend on device accounting, and can change application latency immediately. Establish which block device handles the workload before tuning any file.
Confirm the hierarchy and controller
Start by verifying that cgroup v2 is mounted and that the io controller is available in the relevant parent. Controller availability is not the same as enablement: the parent enables a controller for its children through cgroup.subtree_control, and controller interface files then appear in the child cgroups. A non-root domain cgroup that distributes resources generally must not contain processes of its own when enabling domain controllers. Follow the cgroup v2 no-internal-process rule rather than writing a control file in an arbitrary directory.
findmnt -t cgroup2
cat /sys/fs/cgroup/cgroup.controllers
cat /sys/fs/cgroup/my.slice/cgroup.controllers
Those paths are examples; on a managed host, systemd usually owns the hierarchy. Do not create a competing tree or move production processes by hand. Use the service manager’s supported resource properties where applicable, and inspect the resulting cgroup files to verify effective state. If configuring a delegated subtree, operate only within the delegation and never change ancestor controls you do not own.
Choose the control model that matches the objective
io.weight is a relative weight for distributing I/O among active sibling cgroups. The kernel documentation describes weights from 1 through 10,000 with a default of 100. A weight is not a fixed throughput reservation: when siblings are not competing, a group can use otherwise idle capacity. When one sibling is busy, relative weights influence the distribution implemented by the I/O scheduler and block layer.
io.max is a per-device limit for bytes per second and/or I/O operations per second. A limit is a ceiling, not a minimum share. The io.latency control instead protects a target latency for a group by throttling lower-priority peers when the protected group misses its target. It is work-conserving while targets are met, and its behavior depends on the device and hierarchy. Do not set a latency target below the device’s achievable baseline and expect the controller to manufacture performance.
Weights are useful when the goal is preference under contention. io.max is useful for bounding noisy-neighbor consumption. io.latency is an experiment for workload protection when the workload and device can be measured meaningfully. Combining all three without an explicit policy makes diagnosis difficult; change one control at a time.
Apply a narrow io.max limit
The io.max file uses a block-device major:minor key followed by optional rbps, wbps, riops, and wiops fields. max removes a limit for that field. The block device key must correspond to the layer where the kernel accounts the I/O, not simply the pathname shown by df.
cat /sys/fs/cgroup/workload-a/io.max
printf '%s\n' '259:0 rbps=10485760 wbps=max riops=max wiops=max' \
| sudo tee /sys/fs/cgroup/workload-a/io.max
cat /sys/fs/cgroup/workload-a/io.max
Assume workload-a already exists as a non-root cgroup, the io controller is enabled for it by its parent, and the manager or delegation permits this write. The 259:0 device and 10 MiB/s read cap are illustrative, not a recommended production value. Find the actual major:minor with tools such as lsblk -o NAME,MAJ:MIN,TYPE,MOUNTPOINTS, and trace stacked devices such as device mapper, RAID, or virtual disks before choosing a key. A path may traverse several block devices, and accounting at the wrong layer can make a control appear ineffective or constrain unrelated workloads.
Use exact io.max syntax supported by the running kernel, write one device entry at a time, and read the file back. A successful write confirms configuration acceptance, not the measured rate. Rate limiting can add queueing delay and interact with the device’s scheduler. Start with a reversible canary cgroup and a limit derived from a measured workload objective.
Interpret io.weight as a share, not a cap
The parent cgroup distributes the configured weight among active children according to the controller’s model. If only one child has runnable I/O, it may consume spare bandwidth. If several siblings are continuously active, a higher relative weight favors one group; it does not promise a constant number of IOPS or a particular latency. Device firmware, request size, queue depth, scheduler, and other cgroups all influence observed results.
Set the weight on the child that should receive a different share and compare siblings under simultaneous load. Testing a weight with no competing I/O says little. Record the exact hierarchy and parent weight because an ancestor constraint remains effective and a nested child cannot override it. Avoid moving a process repeatedly between groups as an ad hoc throttle; the kernel documentation recommends organizing workload membership once and changing controller settings rather than frequent migrations.
Use io.latency as a peer-protection mechanism
The io.latency file takes a device major:minor and a target in microseconds. Its protection is organized among peer groups: a group missing its target can cause the controller to throttle peers with less stringent targets. The kernel documentation recommends starting above expected device latency and observing io.stat; rotational devices expose average-latency information under debug-stat conditions, while non-rotational devices expose missed/total-style counters when available.
This control is not a per-request deadline and does not guarantee that every operation completes below the target. A target should be based on the workload’s measured service latency on the same device and under representative load. Protecting a latency-sensitive group can intentionally reduce throughput for a bulk peer. Make that tradeoff explicit to service owners before enabling it.
Read back the target and capture io.stat for the target group and its siblings before and after the change. A controller that is not seeing the expected device may show no relevant changes. Debug fields may require a kernel configuration or module parameter and can add observability overhead; do not assume they are present on every distribution kernel.
Understand accounting before blaming the controller
Buffered writes may be issued later by writeback, not at the application syscall that dirtied the page. The kernel’s cgroup writeback support attempts to associate writeback with cgroups, but filesystem and workload behavior affects attribution. Direct I/O and metadata paths can behave differently from buffered file writes. Swap, journal, and filesystem metadata traffic also complicate simple per-process accounting.
io.stat reports device-level I/O counters for cgroups; it does not necessarily answer “which request from my application caused this physical command.” Compare deltas over a controlled interval and correlate them with application operations, block-device statistics, and workload placement. A counter snapshot is not a latency histogram. Use a suitable tracing or benchmarking tool when you need per-request latency distributions.
Check whether the cgroup controller is enabled at the ancestor that owns the resource distribution, whether the target device is visible in the file, and whether the process is actually inside the intended cgroup. cat /proc/$PID/cgroup and service-manager status are better first checks than repeatedly writing io.max.
Validate with a safe experiment
Use a disposable cgroup, a dedicated test device or scratch file, and a bounded benchmark. Record baseline throughput, IOPS, p50/p95/p99 latency, queue depth, device utilization, and application-level errors. Apply one control, repeat under both idle and competing load, then remove the limit or restore the old value. Never benchmark by writing to an important raw device.
Test reads and writes separately; a read cap does not constrain writes. Test multiple request sizes and synchronous versus queued workloads because an IOPS limit and a byte-rate limit affect them differently. Include a workload with no siblings to demonstrate that a weight is not a hard cap, and a workload with busy siblings to demonstrate that weight only matters under competition.
If an application becomes slow, remove the experimental control first and compare with the baseline. Keep a record of the original file contents, device identifier, cgroup path, timestamp, kernel version, filesystem, and benchmark command. Re-check after reboot or service-manager reload because the persistent configuration may not match the temporary cgroupfs change.
The cgroup v2 I/O controller is a set of distinct policies, not one generic throttle. Identify the actual block layer, choose a weight, hard maximum, or peer-latency target according to the intended outcome, and validate it with controlled contention and application-level latency measurements.
Related:
- Linux blk-mq Internals: Software Queues, Hardware Dispatch, and I/O Scheduling
- Limiting a Service’s CPU and Memory with cgroups v2
Sources: