Linux dm-integrity: Journal, Bitmap, and Authenticated Storage Trade-offs
Use Linux dm-integrity deliberately: choose journal or bitmap semantics, distinguish corruption detection from authentication, and plan recovery safely.
Linux Device Mapper’s dm-integrity target adds per-sector integrity tags to a block device. Depending on its mode, it can detect accidental corruption, authenticate data with a keyed integrity algorithm, or supply integrity metadata to another device-mapper target such as dm-crypt. It is not simply a filesystem checksum switch: it changes on-disk layout, usable device capacity, write ordering, recovery behavior, and the consequences of formatting or changing parameters.
The important design choice is the integrity contract. A checksum such as CRC can detect many accidental changes, but an attacker with raw write access to the backing device may be able to modify both data and an unkeyed tag. A keyed HMAC can authenticate data without encrypting it. When dm-integrity is combined correctly with dm-crypt, the stack can provide authenticated disk encryption. These are distinct properties; do not describe corruption detection as confidentiality or cryptographic authenticity.
Understand why the journal exists
Every protected data block has associated integrity metadata. A crash between writing the data and its tag can leave a new block paired with an old tag, or the reverse. Journal mode handles this by writing data and tags to the journal, committing the journal, and later copying the data and tags to their regular locations. It provides the atomic relationship needed for recovery, at the cost of extra writes and journal work.
The kernel documentation describes journaled writes as requiring the data to be written twice. That is a useful warning, not a promise that every workload will measure exactly half the throughput. Queue depth, cache behavior, block size, storage firmware, tag size, flush behavior, and the workload’s write pattern all affect observed performance. Benchmark the real stack and include flush-heavy and crash-recovery scenarios rather than extrapolating from sequential throughput alone.
dm-integrity also has a bitmap mode. The bitmap records regions where data and tags may not be synchronized; after a crash, those regions are recalculated. It avoids the journal’s second data write and can be faster, but the kernel documentation warns that corruption occurring around a crash may go undetected. Bitmap mode is available for internal-hash operation, not as a universally interchangeable journal setting. Select it only if its failure semantics fit the data and recovery objective.
Direct-write mode does not journal data and tags. Because they are written separately, a crash can leave them inconsistent. Recovery mode is more restricted still: it does not replay the journal, does not check tags, and does not permit writes. It exists as a recovery path when normal activation is not possible, not as a performance setting or routine read-write fallback.
Some hardware exposes inline integrity metadata in sector protection-information fields. This mode has device-profile and tag-size requirements, and may require device-specific provisioning. It does not use the journal or bitmap. Do not infer support from an NVMe or SCSI device label alone; confirm the kernel integrity profile and the storage stack end to end before designing around inline metadata.
Decide what a tag can and cannot prove
In standalone mode, dm-integrity calculates and verifies its own tags. CRC algorithms such as CRC32C are useful for detecting accidental corruption but are not keyed against a hostile writer who can alter the raw device. A non-cryptographic hash has a similar limitation for adversarial tampering. A keyed HMAC provides cryptographic authenticity of the protected blocks if the key remains secret and the selected format and algorithms are configured correctly; it does not encrypt the underlying data.
When used beneath or with dm-crypt, integrity metadata can be supplied by the upper target. In the authenticated-encryption arrangement documented by the kernel, dm-crypt creates integrity information and passes it to dm-integrity, so a modified encrypted device can return an I/O error instead of unauthenticated random plaintext. The exact configuration is managed through supported cryptsetup/LUKS2 interfaces, not by assembling arbitrary table parameters from memory. Validate the chosen LUKS2 mode, kernel support, cryptsetup version, key lifecycle, and recovery process for the distribution you operate.
Think through the attacker and failure models separately:
- A device that fails or returns bad bits is an accidental-corruption problem.
- A device exposed to untrusted raw writes requires a keyed authentication property, not just CRC.
- A lost integrity key can make authenticated data unavailable even when the sectors are intact.
- An unencrypted HMAC-protected device can reveal file content to anyone who can read raw sectors.
- Encryption without authentication may not detect malicious modification in the same way as authenticated encryption.
- A journal may reveal recent write locations; its contents and metadata need an explicit confidentiality and integrity assessment.
Do not treat one layer as an automatic substitute for another. Filesystem checksums, block integrity tags, full-disk encryption, application-level signatures, and backups answer different questions. dm-integrity cannot recover data after a device failure; it can report integrity errors or maintain metadata needed by its chosen mode.
Treat formatting and on-disk parameters as lifecycle decisions
integritysetup format initializes the device metadata and, by default, wipes the device. Formatting an existing volume is a destructive operation. The tool will not safely guess how to reinterpret an arbitrary existing layout, and a nonzero invalid superblock prevents normal target loading. Before any format or migration, inventory the device identity, size, current mappings, filesystem signatures, backup, and recovery keys. Use a disposable loop-backed test device for experiments, never an ambiguous /dev/sdX copied from an example.
These commands inspect metadata and active mapping status without formatting a device:
integritysetup dump /dev/mapper/integrity-volume
integritysetup status integrity-volume
dump reports parameters from the on-disk superblock; status reports the active mapping. Neither replaces a backup or a complete storage inventory. Preserve the command line, kernel version, cryptsetup version, integrity algorithm, tag size, key identifier, and recovery instructions in the volume’s operational record. Never store the actual HMAC key alongside an unencrypted copy of the device it protects.
The volume’s on-disk geometry depends on format-time parameters such as tag size, block size, interleave, metadata placement, journal size, and layout-related feature options. Some runtime parameters can be changed by reloading a device-mapper table, while other arguments must not be changed because the on-disk layout depends on them. Use the exact tool and kernel documentation to distinguish mutable controls from format-time choices. Do not experiment by changing options on a production mapping.
The --no-wipe option is not a generic way to preserve valid integrity-protected data during first-time formatting. The manual warns that a device not initially wiped contains invalid checksums; a subsequent recalculation pass must complete before the device is fully integrity protected. Capacity must also be planned around metadata and padding: the exposed data sector count is smaller than the entire backing device. Use the reported provided-data size rather than assuming every physical sector belongs to the filesystem.
Use trim and discard options with care
Discard forwarding leaks information about allocation and use patterns on encrypted storage, and has separate tag semantics for integrity-protected devices. The kernel documentation distinguishes ordinary allow_discards from allow_discards_keyed: the former uses a constant filler tag for discarded blocks that a raw writer can forge, while the keyed option uses a keyed checksum of the sector number. The keyed form is not compatible with volumes whose discarded blocks were marked using the earlier method and is intended for freshly formatted volumes.
This is a format and threat-model decision, not an option to toggle during incident response. If trim support is required, document the metadata leakage, verify that the selected feature is supported throughout the storage stack, and test discard, recovery, and migration behavior on a disposable volume. Do not turn on discard forwarding merely because the underlying SSD supports TRIM.
Plan the initial setup as a controlled change
For a new standalone test volume, integritysetup format establishes the target’s on-disk structure; integritysetup open creates an active mapping. The examples below are intentionally not a copy-and-paste setup: TEST_DEVICE must refer only to a disposable loop device that contains no required data, and an HMAC key file must be provisioned and protected separately. Formatting the wrong device destroys data.
# Inspect the test device identity and size before any write operation.
lsblk -o NAME,SIZE,TYPE,FSTYPE,MOUNTPOINTS "$TEST_DEVICE"
# DESTRUCTIVE: use only a disposable loop device prepared for this test.
integritysetup format "$TEST_DEVICE" --integrity crc32c
integritysetup open "$TEST_DEVICE" integrity-test --integrity crc32c
integritysetup status integrity-test
This example demonstrates standalone CRC32C integrity checking. It does not provide confidentiality or strong authentication against an attacker who can rewrite the device. For keyed integrity, integritysetup supports an HMAC algorithm with a key file, explicit key size, and tag size; key generation, storage, rotation, access, and recovery must be designed before formatting. For authenticated disk encryption, prefer a supported cryptsetup/LUKS2 configuration and validate its exact semantics instead of layering hand-written dmsetup commands.
After activation, format a filesystem only on the mapped device that the integrity layer exposes, not on the raw backing device. Test clean shutdown, forced power-loss simulation in a VM, full-device conditions, detected tag mismatch, recovery-mode access, and restoration from backup. Do not trigger corruption on a real production volume to prove the alerting path. Ensure monitoring distinguishes an integrity mismatch from ordinary media errors and that responders know when to stop writes and image the backing device.
Handle recalculation and recovery conservatively
Automatic recalculation can populate tags for pre-existing data when the documented workflow uses --no-wipe, but the volume is only fully integrity protected after recalculation finishes. The superblock records a progress position that integritysetup dump can report. Make that state visible to monitoring and do not claim complete coverage while background recalculation is still in progress. If interrupted, resume only with the intended algorithm, key, and device identity; a new calculation from the wrong data cannot prove that the data was originally correct.
Recovery mode disables tag checking and journal replay and disallows writes. It can help expose data for recovery when ordinary activation fails, but it is not evidence that the content is valid. Capture the original backing device before attempting repair, record the status and logs, and avoid writing through a recovery mapping. If tags mismatch, distinguish a stale or torn data/tag pair from a corrupted journal, wrong key, wrong algorithm, wrong device, or damaged media. Escalate to the storage owner with the preserved image and metadata rather than reformatting or recalculating blindly.
When resizing, follow the supported integritysetup and cryptsetup order for the exact stack. Growing an active mapping can trigger recalculation over newly exposed space; verify the new data-sector count and wait for the status to reach completion before representing the whole expanded range as protected. Shrinking or changing a backing layout is a separate migration with filesystem and device-mapper constraints, not an inverse of growth.
Benchmark the complete stack
Compare journal and bitmap only when the latter is valid for the selected internal-hash mode and its weaker crash semantics are acceptable. Use representative block sizes, sequential and random writes, flush-heavy operations, queue depths, and the actual upper layer (filesystem, LUKS2, RAID, multipath, or virtualization). Measure throughput, tail latency, CPU usage, metadata I/O, usable capacity, recovery time, and behavior after an interrupted write. Report the integrity mode and kernel, cryptsetup, and storage versions with each result.
Do not compare a journaled integrity volume against a plain device and attribute every difference to tag calculation. Journaling adds writes; HMAC adds computation; data encryption changes CPU and I/O paths; filesystem barriers and device caches affect durability. Use fresh disposable volumes with the same backing device and test data, and include a baseline for each stack layer. Performance is one axis of the decision; crash consistency and adversarial integrity are separate requirements.
Production review checklist
- Does the selected tag algorithm meet the accidental-corruption or adversarial-tampering threat model?
- Is confidentiality provided separately or through a validated authenticated-encryption stack?
- Are journal, bitmap, direct-write, recovery, or inline semantics documented and accepted?
- Is the integrity key protected, backed up, recoverable, and excluded from the data device?
- Are format-time geometry and metadata capacity recorded before provisioning?
- Is
formatrestricted to verified disposable or newly provisioned devices? - Are recalculation progress and integrity mismatches monitored and understood?
- Are discard leakage and keyed-discard compatibility requirements explicit?
- Have crash recovery, mismatch handling, resizing, and backup restoration been tested without risking production data?
dm-integrity is useful when the data path needs sector-level integrity semantics that ordinary block storage does not provide. Choose its journal and tag model based on the failure and threat model, preserve keys and metadata as first-class recovery assets, and test the exact block stack. A faster mode is not an improvement if its crash behavior violates the service’s durability contract.
Related:
- How to Set Up Full-Disk Encryption on Linux with LUKS
- Linux NVMe Multipath: Path Selection, ANA State, and Failover Evidence
Sources: