Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Operating Linux Software RAID with mdadm: Assembly, Recovery, and Monitoring

An operational guide to Linux md RAID: disk identity, safe array creation, assembly, rebuild monitoring, replacement, alerts, and recovery boundaries.

Linux MD (multiple devices) RAID combines block devices into a virtual block device managed by the kernel. mdadm creates, assembles, monitors, and reshapes those arrays. A redundant level can keep some storage available after certain component failures, but it does not replace backups, independent failure domains, filesystem checks, or a tested restore. This guide focuses on the operational lifecycle of a conventional RAID1 mirror and calls out where other levels differ.

The most important boundary is ownership: MD protects block-level redundancy; the filesystem and application still own their own consistency. A mirror can faithfully replicate an accidental deletion, bad application write, or filesystem corruption. Use RAID to meet an availability or device-failure requirement, and use a separate, independently protected backup to recover from logical damage or loss of the host.

Inventory stable device identity before changing anything

Never identify a production component by assuming /dev/sdb will remain the same after reboot. Compare model, serial, size, transport, filesystem signatures, current mountpoints, and existing MD metadata:

lsblk -o NAME,PATH,TYPE,SIZE,FSTYPE,MOUNTPOINTS,MODEL,SERIAL
ls -l /dev/disk/by-id/
cat /proc/mdstat
sudo mdadm --detail --scan
sudo mdadm --examine /dev/disk/by-id/REPLACE_WITH_DEVICE_ID

Replace the placeholder with the actual stable path. Do not run a create, zero-superblock, or format command until every candidate has been positively mapped to the intended physical device and any data that must survive has a tested backup. Confirm no filesystem, swap area, LVM PV, existing array, or active workload depends on the target. A single incorrect component path can overwrite valuable metadata or make a valid array harder to assemble.

For an existing array, inspect the array and each member independently. mdadm --detail /dev/md0 reports the assembled array’s state; mdadm --examine reads metadata on a component. These answer different questions. If metadata suggests two arrays, a stale clone, or an unexpected UUID, stop and resolve that identity conflict before assembling or forcing anything.

Choose a level from failure and capacity requirements

RAID1 mirrors data across members and can tolerate a member failure as long as a complete usable copy remains. RAID0 stripes data without redundancy, so loss of a member loses the array. RAID5 and RAID6 distribute parity and have more complex degraded-write and recovery behavior; RAID10 combines mirrors and striping with layout-dependent failure tolerance. A RAID level is not a promise against every combination of failures: read the supported level’s semantics and model which specific member-loss sets preserve data.

Capacity is constrained by the array geometry. A two-member RAID1 generally provides usable data capacity no larger than its smallest member; extra mirror members add copies, not additive capacity. Mixing disk models or sizes may leave space unused. Component devices that are nominally equal in size can expose slightly different usable sectors, so plan for the smallest actual component and verify what the array reports.

The failure domain matters as much as the level. Two members in one enclosure, on one controller, sharing one power supply, or in the same physical host can disappear together. Mirroring across two partitions on the same physical disk does not provide disk-failure tolerance. A spare helps shorten degraded time only if it is available, compatible, monitored, and not exposed to the same correlated failure.

Create only a new array on verified empty devices

The following is a syntax template for a new two-member RAID1. The component paths are placeholders; substitute only devices that have been verified empty and intentionally dedicated to this array. mdadm --create writes array metadata and the initial synchronization may overwrite prior assumptions about their contents. Do not paste it unchanged:

sudo mdadm --create /dev/md0 \
  --metadata=1.2 --level=1 --raid-devices=2 \
  /dev/disk/by-id/REPLACE_WITH_MEMBER_A \
  /dev/disk/by-id/REPLACE_WITH_MEMBER_B

After creation, wait for the initial synchronization to complete or reach the approved operational state. Observe it from both interfaces:

cat /proc/mdstat
sudo mdadm --detail /dev/md0
cat /sys/block/md0/md/sync_action
cat /sys/block/md0/md/degraded

Do not create a filesystem merely because /dev/md0 exists. Confirm the expected level, UUID, member count, component paths, and synchronization state first. On a deliberately new array only, filesystem creation (for example, mkfs.ext4 /dev/md0) is itself destructive and belongs after that verification. Prefer mounting the filesystem by its filesystem UUID rather than a volatile /dev/md0 name.

Persist assembly and mount configuration deliberately

Linux distributions differ in the location and boot integration of mdadm.conf; common locations include /etc/mdadm/mdadm.conf and /etc/mdadm.conf. Ask the installed package which file its initramfs tooling consumes. Use mdadm --detail --scan to produce candidate ARRAY records, inspect and add the intended record without blindly duplicating the whole output, then rebuild the initramfs with the distribution’s supported tool when required. A correct config file that is not included in the boot image may still leave a root or data array unavailable at startup.

Test the boot path during a maintenance window: stop workloads that use the mount, perform a controlled restart, and verify that the expected UUID assembled and mounted automatically. Check journalctl -b for assembly and mount errors. On systems where a separate service starts arrays, confirm that service’s installed unit and initramfs hooks rather than assuming a service name is universal across distributions.

Monitor health and rebuilds, not just device presence

The kernel exposes a live array summary in /proc/mdstat; mdadm --detail provides the member and array state; sysfs provides machine-readable state and synchronization attributes. During a rebuild, record the action (recover or resync), progress, degraded count, member state, read/write errors, and completion. A device appearing in lsblk is not evidence that it is a healthy in-sync member.

watch -n 2 cat /proc/mdstat
sudo mdadm --detail /dev/md0
for f in /sys/block/md0/md/degraded /sys/block/md0/md/sync_action \
         /sys/block/md0/md/mismatch_cnt; do
  printf '%s: ' "$f"
  cat "$f"
done

Replace watch and the md0 path for the local operating system and device. Arrange notifications through the distribution’s supported mdadm monitor service or monitoring agent, and verify that a simulated alert reaches the on-call destination in a test environment. Polling /proc/mdstat from a terminal is a useful diagnosis, not a durable alerting system. Alert on degraded arrays, failed members, rebuilds that stop progressing, and hardware or transport errors before a second fault removes the remaining redundancy.

For redundant levels that support it, a scheduled consistency check reads the array’s data and redundancy information. The kernel exposes sync_action=check and a mismatch count; repair is a distinct write operation. A mismatch is evidence to investigate, not proof that a particular copy is authoritative. Schedule checks according to workload and maintenance capacity, monitor their impact, and do not automatically launch a repair without understanding which data source the level and kernel will use.

Replace a failed member with an evidence-based sequence

First confirm the array is actually degraded and identify the exact failed component from mdadm --detail, sysfs, kernel logs, and device serials. Do not infer the failed physical disk from a familiar /dev/sdX letter after reboot. Check whether the drive is fully failed, merely disconnected, or has a recoverable path/cable/controller issue; an intermittent path problem and media failure need not have the same remedy.

For a confirmed failed member, the general management sequence is to mark that component failed if MD has not already done so, remove it from the array, install and verify a replacement device, then add the replacement. These commands are templates, and the selected member path must be checked again immediately before use:

sudo mdadm --manage /dev/md0 --fail /dev/disk/by-id/OLD_MEMBER
sudo mdadm --manage /dev/md0 --remove /dev/disk/by-id/OLD_MEMBER
sudo mdadm --manage /dev/md0 --add /dev/disk/by-id/NEW_MEMBER

Adding a device to a degraded redundant array usually starts recovery automatically, but confirm its sync_action, progress, and final in-sync state instead of assuming it completed. Keep a tested backup, avoid unnecessary reboots while a rebuild is in progress, and ensure the replacement provides enough usable sectors. If the array contains a boot/root filesystem, follow the distribution’s bootloader and initramfs recovery instructions too; the data array’s recovery alone does not restore bootability.

Never use --force, --assume-clean, --zero-superblock, or an assembly command that overrides consistency checks as generic recovery advice. In particular, the kernel may refuse to start a dirty degraded parity array because data cannot be reconstructed with confidence. Overriding that refusal can expose silent corruption. Preserve member metadata and take a block-level image when forensic recovery is required; seek specialist review rather than experimenting on the only surviving copies.

Define acceptance before declaring the array healthy

An operational array is not “green” only because mdadm --detail says active. Confirm all expected members are in_sync, the degraded count is zero, no rebuild or check is stuck, boot assembly works, the mounted filesystem UUID is correct, and alerts reach the responsible operator. Exercise a replacement procedure on a disposable lab or documented maintenance plan; do not deliberately fail a production disk to prove a dashboard works.

Most importantly, test restore from an independent backup. A mirror preserves availability during some component failures; it does not preserve a second point-in-time copy, protect against host-level damage, or undo a mistaken write. Document the exact array UUID, member identities, boot configuration, owner, alert path, and tested restore procedure so another operator can recover it without guessing.

Related:

Sources:

Comments