Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux DMA API: Mapping, Coherency, and Device Ownership

Use the Linux DMA API correctly by separating CPU and device addresses, mapping lifetimes, cache coherency, scatter-gather paths, and teardown.

Direct Memory Access lets a device read or write memory without the CPU copying every byte. It does not mean a device can safely use a CPU virtual address. Linux’s DMA API translates and manages device-visible addresses, enforces platform constraints, and provides synchronization operations for memory that may be cached differently by the CPU and device.

The key distinction is that a CPU pointer, a physical address, and a DMA address are different types of value even when they happen to have the same numeric representation on one machine. An IOMMU, bus offset, bounce buffer, DMA mask, or architecture cache rule can make them diverge. Drivers should use the DMA API and preserve the ownership protocol; casting a pointer to an integer is not a portable mapping strategy.

Select the right address and mask

Before allocating or mapping memory, the driver sets the device’s DMA capability through the supported mask APIs. Streaming and coherent DMA masks can differ. A device with a restricted address width may need a bounce path or may be unable to access a high-memory allocation. If mask negotiation fails, do not continue using an address that the hardware cannot represent.

Treat the struct device as the authority for DMA operations. Map and unmap through that device so the architecture and IOMMU can apply the right domain. For PCI devices, the mapping may be translated through an IOMMU; for other buses, it may go through a different DMA domain. A platform may also require specific alignment, boundary, or maximum segment constraints.

The DMA API HOWTO explains coherent and streaming mappings, masks, and errors. A successful mapping returns a DMA address that the device can use for the mapping lifetime. It does not make the memory safe to access from userspace or guarantee a physically contiguous buffer. Check the mapping error with the API helper and propagate failure before programming device registers.

Coherent memory and streaming memory serve different jobs

Coherent allocations are useful for long-lived control structures that CPU and device access concurrently, such as descriptor rings, when supported by the device contract. Coherent means the API handles visibility requirements; it does not mean ordering is automatic. The CPU may still need memory barriers before notifying the device that a descriptor is ready, and before reading data the device has published.

Streaming mappings are intended for buffers with explicit ownership phases. The CPU prepares data, synchronizes or maps it for the device, then the device owns it. After completion, the CPU synchronizes for CPU access and processes the result. For a reused buffer, map once and alternate sync-for-device and sync-for-CPU where the API and platform permit it. Do not let both sides modify the same cache lines concurrently unless the API and protocol explicitly support that.

On cache-coherent systems, missing synchronization may appear to work during simple tests. On noncoherent systems, stale cache lines can produce corruption that depends on CPU architecture, alignment, or timing. Test on the weakest supported coherency model or use the API contract regardless of the developer workstation.

Direction and mapping lifetime are correctness properties

The DMA direction describes which side reads or writes the buffer. A transmit buffer that the device reads is mapped to-device; a receive buffer that the device writes is mapped from-device; a bidirectional mapping is for cases where both directions are genuinely needed. An incorrect direction can cause cache maintenance to discard or expose the wrong data on some platforms.

The mapping must remain active until the device can no longer access it. On completion, timeout, cancellation, or reset, stop DMA first, confirm the device has quiesced, then unmap or free the buffer. A software timeout alone does not stop hardware. Freeing a buffer while a bus master can still write into it can corrupt unrelated memory or trigger an IOMMU fault.

Every error path matters. If submitting a command fails after mapping, unmap the buffer. If mapping succeeds but descriptor setup fails, unwind in reverse order. If device removal occurs, disable interrupts and DMA using the device’s documented sequence before releasing mappings. Keep the exact owner transition in the driver’s state machine.

Scatter-gather lists and segment constraints

Large buffers may be represented by scatter-gather entries. Mapping a scatterlist can merge entries according to device and platform constraints, so the number of mapped DMA segments can differ from the original number of CPU-side entries. Program hardware from the mapped segment count and DMA addresses, not the pre-map list count.

A device can impose maximum segment size, address boundaries, alignment, and segment-count limits. These limits should be configured or enforced through the appropriate DMA and block/network subsystem APIs. When a mapping cannot be represented, the driver needs a fallback, segmentation strategy, or clear failure. It must not truncate a high address or silently issue fewer descriptors than the payload requires.

Use an unwind table in code review: original entries allocated, mapping call and direction, mapped count, descriptor count, completion or reset, unmap call, and memory free. Validate zero-length, one-segment, maximum-segment, and mapping-failure cases. Include IOMMU enabled and disabled configurations in testing when the platform supports both.

Ordering is separate from cache visibility

DMA synchronization and memory ordering solve related but different problems. Sync operations manage visibility of streaming buffers between CPU and device. Barriers order descriptor writes relative to a doorbell or ensure a device-written completion status is observed before dependent fields. The correct barrier depends on the device protocol and the architecture’s memory model.

Do not replace DMA API calls with generic CPU barriers. A barrier cannot make a noncoherent cache line visible to a device. Conversely, coherent memory does not guarantee a doorbell cannot overtake preceding descriptor stores. Follow the relevant subsystem and architecture helpers, then read the device specification for the required ownership and notification order.

Ring-buffer designs commonly use producer and consumer indices, phase bits, or ownership flags. Publish an entry only after all its fields are initialized. On receive, observe the device’s completion marker before reading payload fields. Serialize concurrent CPU producers if the hardware ring is not multi-producer safe.

Diagnose with the platform evidence

When DMA fails, capture the device identity, driver, kernel, IOMMU state, DMA mask, mapping direction, size, alignment, and device completion state. IOMMU fault logs can identify an unmapped or forbidden device address, but they do not by themselves prove which driver bug produced it. Check descriptor dumps and verify addresses are DMA addresses, not physical or virtual addresses.

Use kernel tracing, dynamic debug, or driver-specific statistics to count map failures, resets, completions, and timeouts. Avoid printing every DMA address in production logs; addresses can reveal layout information and overwhelm logging. For a reproducer, use a small deterministic transfer, guard pages or DMA API debugging where supported, and test repeated teardown under load.

Acceptance criteria

A reliable DMA path has a negotiated address mask, correct device and direction, checked map return values, proper sync and barrier operations, and an explicit lifetime from map through device quiescence to unmap. Test normal completion, device timeout, reset, removal, suspend/resume, and mapping failure. Confirm that payload bytes and descriptor ownership are correct across the supported architectures and IOMMU modes.

The debugging question is not “what physical address should I cast to?” It is “which DMA mapping does this device own now, what constraints were applied, and what transition returns that memory to the CPU?” Keeping those distinctions explicit prevents architecture-specific corruption and use-after-free DMA.

Related:

Sources:

Comments