Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux blk-mq Internals: Software Queues, Hardware Dispatch, and I/O Scheduling

Understand Linux blk-mq request flow, CPU-to-hardware queue maps, tags, scheduler tradeoffs, queue limits, and practical storage diagnostics.

The Linux multi-queue block layer, usually called blk-mq, is the request path that lets storage drivers use device and CPU parallelism without routing every operation through one global submission lock. It sits below filesystems and above block drivers. Filesystem I/O can arrive as bios, be merged or represented as requests, pass through optional scheduling, and ultimately be dispatched to one of the driver’s hardware contexts.

The word “queue” refers to several different objects. A per-CPU software context, a hardware dispatch context, a device’s actual submission queue, a scheduler queue, and a hardware completion queue are related but not synonymous. Diagnosing a slow disk by changing one queue count without knowing which layer is saturated often increases outstanding work without fixing the bottleneck.

From a bio to a driver request

Filesystems and other block clients submit bios describing data operations against a block device. The block layer can combine compatible adjacent bios into a request, enforce device limits, account for the operation, and pass it through the queueing path. A request is the unit the driver eventually submits to the device; a bio is not automatically equivalent to one hardware command.

blk-mq uses software staging contexts and hardware contexts. Software contexts provide local entry points associated with CPUs or NUMA placement. A queue map selects which hardware context receives work from a given software context. Hardware contexts represent driver submission paths and connect the block layer to device-specific queues. A driver can expose more than one queue type, such as default, read, or polled I/O, and map them differently.

The fast path can dispatch directly when the queue has resources and no merge or scheduler work is needed. Otherwise, the software staging path can plug or merge requests, apply scheduling policy, then submit toward a hardware context. If hardware tags or device resources are unavailable, requests wait or are kept for later dispatch. The details depend on the driver, device, selected scheduler, and request flags.

This multiqueue layout reduces contention and lets capable devices accept independent commands in parallel. It does not guarantee a particular CPU-to-device-queue mapping, completion order, or application latency. Neither the block layer nor device protocol guarantees requests complete in submission order; higher layers must preserve any ordering required by their semantics.

Tags are bounded in-flight request slots

Tags let the block layer and driver correlate a completed command with its request without searching all outstanding I/O. A hardware tag is assigned as a request is dispatched to the hardware queue; a scheduler may use a separate scheduler tag when one is attached. Queue depth and reserved tags are driver and device resources, not a safe proxy for application concurrency.

When tags or device command slots run out, callers may wait even if a CPU appears idle. Increasing queue depth can improve utilization for a device that has unused parallel capacity, but it can also increase queueing delay, memory use, and the amount of work that must drain during failure. A high device queue depth is not inherently good: latency-sensitive workloads often need a bounded queue rather than maximum saturation.

Completion can occur on a CPU different from the one that submitted the request. CPU topology, IRQ affinity, driver completion handling, NUMA locality, and the queue map affect cache traffic and the cost of processing completions. Treat CPU affinity and hardware queues as one topology problem, not as isolated per-device knobs.

Understand when an I/O scheduler participates

An I/O scheduler operates at the block layer before dispatch to hardware. It can reorder or group requests and apply a policy such as deadline-oriented dispatch or fairness. The none scheduler removes scheduler-specific ordering; it does not disable the block layer, remove hardware queues, or prevent firmware from internally scheduling commands. Scheduler availability and behavior depend on the kernel configuration, device stack, and distribution.

Schedulers can trade throughput for fairness or latency. A scheduler that helps a rotational device or an interactive desktop may add overhead or harm throughput on a high-parallelism NVMe workload. Stacked devices complicate the decision: a logical device can sit above device-mapper, RAID, multipath, or a virtual disk, and the scheduler visible at one layer may not be the one that controls the physical queue.

Inspect the active scheduler before changing it:

DEV=nvme0n1
cat "/sys/block/$DEV/queue/scheduler"

The selected scheduler is usually shown in brackets, and the other names are the schedulers available for that queue. Changing it through sysfs is a live behavioral change, not a universal performance fix. Test on a representative device, capture latency and throughput first, and record how the distribution persists the selection across reboot. Do not assume that a scheduler listed for one host exists on another.

Read queue attributes as capability and policy

The block queue’s sysfs attributes report device limits and tunables such as logical block size, maximum request size, rotational hint, read-ahead, polling, scheduler, and queue accounting. These values are often inherited or transformed on stacked devices. A zero, absent file, or rejected write can mean the underlying driver does not support the feature or the attribute is read-only; it is not a request for the kernel to guess a value.

DEV=nvme0n1
cat "/sys/block/$DEV/queue/scheduler"
cat "/sys/block/$DEV/queue/logical_block_size"
cat "/sys/block/$DEV/queue/rotational"
cat "/sys/block/$DEV/queue/nr_requests"
cat "/sys/block/$DEV/stat"

The list is for observation. It intentionally does not write tunables. Device names vary, and not every attribute exists or means the same thing on every kernel and driver. For a multipath or device-mapper stack, compare the logical device with each lower block device rather than attributing all latency to the leaf disk.

/sys/block/<device>/stat provides a consistent snapshot containing cumulative block statistics (completed reads and writes, merges, sectors, and accumulated timing counters) plus an in-flight request gauge. Compute deltas only for cumulative counters; interpret the in-flight field as a point-in-time count. Handle counter resets or device recreation. The data does not identify the process that caused each request, nor does device activity alone tell whether a userspace request was cache-served by the filesystem.

Diagnose queueing rather than guessing from utilization

Correlate block statistics with application latency, filesystem writeback, device latency, queue depth, CPU time, and scheduler behavior. If request latency grows while the device is saturated, reduce concurrency or add actual device capacity rather than blindly increasing queue depth. If the device is underused but one CPU is hot, inspect submission/completion placement, IRQ affinity, locks, driver queue mappings, and whether the workload exposes enough independent I/O.

Read and write paths also differ. Buffered filesystem writes may first dirty page-cache pages and reach the block layer later; a fast write() is not necessarily a completed device write. Direct I/O bypasses some page-cache behavior but has its own alignment and application semantics. Flushes, barriers, discard operations, and error recovery can follow different paths. This article focuses on blk-mq after block requests are submitted, not on the complete filesystem durability contract.

Use one workload definition across comparisons: block device and stack, block size, read/write mix, random/sequential distribution, queue depth, concurrency, cache state, durability flags, and CPU placement. Measure p50/p99/p999 completion latency, throughput, errors, CPU cycles, context switches, and time in queue. Compare a single queue and multiple independent flows only when the device supports the underlying parallelism.

Test failure and recovery paths

Stress queue saturation and I/O errors on disposable devices or test virtual machines. Verify that request timeouts, device reset, hot removal, suspend/resume, and driver recovery do not strand application work. Test how a stacked device propagates errors and flushes. A benchmark that only measures successful steady-state I/O misses the conditions that turn a queue backlog into a hung filesystem or database.

Record the kernel release and configuration, driver and firmware, block topology, scheduler, queue limits, CPU/NUMA map, workload generator version, and all tunables. Re-run after kernel or firmware changes; queue defaults, supported scheduler modules, and driver mappings can change independently of the userspace workload.

blk-mq provides a scalable request framework, not an automatic performance setting. Correct diagnosis starts by identifying whether requests are waiting in a software context, blocked by scheduling or tags, delayed in a driver queue, or completing slowly on the device. Only then can a queue or scheduler change be tested against a real service objective.

Related:

Sources:

Comments