Skip to content
Shell & TerminalDeep Dive Published Updated 8 min readViews unavailable

du in Production Shell Scripts: Disk Usage, Hard Links, and Mount Boundaries

Interpret du results with explicit apparent-size, hard-link, symlink, mount, and unit policies before using them for capacity decisions.

du estimates space used by files and directories. Its result is not simply the sum of file sizes: allocated blocks, sparse files, hard links, symlinks, filesystem compression, reflinks, snapshots, mount traversal, and reporting units all affect interpretation. A cleanup job that deletes data based on a misunderstood total can remove the wrong files while appearing to follow a numeric threshold.

Define the question first. Apparent size asks how many bytes a file reports as content length. Disk usage asks how much filesystem allocation is attributed to entries under a traversal. Quotas, snapshots, deduplicated storage, and remote filesystems can report different capacity views. du is an estimate for a path walk, not a complete accounting system.

Apparent size and allocated blocks differ

A sparse file can have a large logical length but few allocated blocks. Conversely, small files can consume more allocation than their byte length because filesystems allocate blocks. GNU du’s apparent-size option reports logical size; default block usage reports allocation according to implementation and filesystem data. Human-readable units can be rounded and are unsuitable for exact threshold comparisons.

Use explicit machine units for automation and avoid comparing rounded output. GNU -B1 and BSD equivalents have different interfaces, so test the target utility. If policy is based on a quota or filesystem availability, query that quota or filesystem capacity API rather than summing du output and assuming it equals free space.

Hard links complicate totals because multiple names can refer to one inode. Implementations generally avoid counting a multiply linked file more than once within a traversal unless an option changes that policy. Traversal root and ordering can determine which name is counted. This matters when reporting per-directory usage across overlapping roots.

GNU du documents this behavior explicitly: it normally counts a multiply linked file once, and the order of file operands can determine which pathname receives that count. Adding overlapping directory operands can therefore make the sum of their individual rows differ from a unique-inode total for the combined tree. Keep the traversal scope and operand list stable when comparing reports over time. GNU du --inodes answers a different operational question by reporting inode consumption, which is useful when a volume runs out of inodes despite having free bytes. It is not a byte-capacity measurement, and GNU documents that block-size and apparent-size options do not apply in inode mode.

GNU du documents the hard-link rule explicitly: by default it counts a multiply linked file once, and operand order can determine which pathname receives that count. Adding two directory operands that overlap can therefore produce surprising per-root figures. Do not assume that summing independently collected subtree totals yields a unique-inode total for the parent tree. Keep the traversal scope and operand list stable when comparing reports over time.

GNU du --inodes answers a different operational question by reporting inode consumption rather than block usage. This is useful when a volume runs out of inodes despite having free bytes, which can happen with vast numbers of tiny files. It is not a byte-capacity measurement, and the manual notes that block-size and apparent-size options do not apply to this mode. Report inode counts separately from byte or block totals rather than combining them into one “disk usage” number.

By default, du commonly accounts for a symlink itself rather than following its target; options can change that policy. Following links may count content outside the intended tree or revisit targets. Decide whether storage belongs to the link entry or referenced object and use an explicit mode.

A filesystem-boundary option such as GNU -x can keep traversal on one device, but device identity is not a universal security boundary. Bind mounts, containers, snapshots, and network filesystems complicate path topology. For a strict inventory, record which filesystems were traversed and whether permission errors prevented a complete scan.

Treat -x as a same-device traversal rule, not as a guarantee that the path tree matches a container, mount namespace, or administrative ownership boundary. POSIX defines the option in terms of device identity; a mount arrangement that presents the same device number can differ from a policy that means “do not cross any mount point.” If mount topology itself is the policy, enumerate and validate mounts using an operating-system-specific interface before walking the tree.

Treat -x as a same-device traversal rule, not as a guarantee that the path tree matches a container, mount namespace, or administrative ownership boundary. POSIX defines the option in terms of device identity; a mount arrangement that presents the same device number can therefore differ from a policy that says “do not cross any mount point.” If mount topology itself is the policy, enumerate and validate mounts using an operating-system-specific interface before walking the tree.

The root path must be explicit and quoted. A missing or unreadable subtree can make totals incomplete. A partial report should not drive automatic deletion. Capture stderr and process status; if a pipeline consumes du output, preserve producer failures with pipefail or stage results.

Threshold checks must avoid display formats

Do not parse human-readable K/M/G output with shell arithmetic unless units and rounding are deliberately defined. Use byte counts and a safe integer parser, then compare against a threshold in the same unit. Shell arithmetic may have bounded integer width and can overflow for large volumes; use a tool with suitable range or compare normalized values safely.

Sampled outputs are observations, not locks. Files can grow or disappear after du traverses them. Before deleting or rotating data, recheck each candidate and apply retention rules based on ownership, age, and service state. Do not delete the largest path simply because it dominates a snapshot; it may be an active database file.

Filesystem reporting can also differ from physical device use. GNU’s manual notes that the operating system may not report duplicate or backup blocks; on copy-on-write filesystems, du can count the space that would be used if shared non-hole data were rewritten rather than the space physically consumed at that instant. Compression may be reported as uncompressed size, and network filesystems may not expose precise server-side allocation to the client. Snapshots or storage-pool reservations outside the traversed path tree are not counted merely because every visible pathname was scanned. For these reasons, interpret du as a path-tree estimate, not a device accounting API.

Use df for filesystem-reported used and available capacity, and use the storage platform’s quota, snapshot, or pool-management interface when a decision concerns those counters. The values can legitimately disagree because they answer different questions. Capture the filesystem type, mount options, privilege context, traversal root, and time of collection if a report will be compared during incident response. A denominator such as “percentage full” should come from the same filesystem-level source as its numerator; do not combine a path-tree total from du with unrelated capacity figures and label the ratio authoritative.

Machine-readable output needs an explicit pathname framing policy too. GNU du -0 terminates output records with NUL so a filename containing a newline cannot masquerade as multiple records. A consumer still has to parse the command’s size-and-path format correctly; do not split on arbitrary whitespace or store NUL-delimited output in a shell variable. Quote the operand and use -- before a path that may begin with a hyphen. If only the aggregate number matters, request a summary and keep the pathname separately in report metadata.

Filesystem implementation details also limit the meaning of allocated-space totals. GNU’s manual cautions that operating systems may not report duplicate or backup blocks, and that copy-on-write filesystems can make du count the space that would be used if shared non-hole data were rewritten rather than the space physically consumed at that instant. Compression may be reported as uncompressed size, while network filesystems may not expose precise server-side usage. A snapshot, deduplicated extent, or storage-pool reservation that is not represented by ordinary files under the traversed root will not become visible merely because du visits every pathname.

For these reasons, interpret du as a path-tree estimate, not as a device accounting API. Use df to inspect filesystem-reported capacity and free space, and use the storage platform’s quota, snapshot, or pool-management interface when the decision concerns those counters. The values can legitimately disagree because they answer different questions. Capture the filesystem type, mount options, privilege context, and time of collection if a report is used for incident analysis.

Machine-readable output needs an explicit pathname framing policy too. GNU du -0 terminates output records with NUL so a filename containing a newline cannot look like multiple records. A consumer still has to parse the command’s size-and-path format correctly; do not split on arbitrary whitespace or store NUL-delimited output in a shell variable. Quote the operand and use -- before a path that may begin with a hyphen. If only the aggregate number matters, ask for a summary and keep the pathname separate in the report metadata.

Reliable capacity workflow

A capacity report should include exact root, unit, apparent versus allocated-size mode, link policy, mount policy, timestamp, implementation, and whether warnings occurred. For high-stakes cleanup, produce a dry-run manifest first, review it, and then run deletion using a separately validated command. Keep a recovery window and never let shell glob expansion define the deletion set implicitly.

Test sparse files, hard links, symlinks to internal and external targets, nested mount points, unreadable directories, files changing during traversal, and unusual names. Compare against filesystem-level reports and known fixtures. Ensure labels state whether numbers are bytes, blocks, or estimates.

du is useful for path-scoped diagnosis and trend reports, but no single flag turns it into a quota oracle. Make allocation, traversal, and failure assumptions explicit before automating capacity decisions.

Related:

Sources:

Comments