OverlayFS Metadata-Only Copy-Up: metacopy, Layer Integrity, and Safe Operations
Understand OverlayFS metacopy behavior, delayed data copy-up, xattr trust boundaries, incompatible options, and safe lower-layer lifecycle practices.
OverlayFS normally copies a lower-layer file into the writable upper layer before applying a write or metadata change. That copy-up preserves the lower layer, but it can move far more data than a metadata operation requires. On supported kernels, the metacopy feature can copy metadata first and defer copying file data until a process opens the file for writing. This can reduce immediate copy-up work for workloads that adjust ownership or mode on large lower-layer files but do not modify their contents.
metacopy is not a transparent free optimization. It changes the relationship between upper and lower layer content, adds an extended-attribute trust boundary, has mount-option conflicts, and depends on the lower layer remaining consistent. Treat it as a deliberate filesystem policy that must be validated against the actual kernel, backing filesystems, runtime, and data lifecycle.
Copy-up normally moves metadata and content together
An OverlayFS mount presents a merged namespace. The upper layer wins when a name exists in both upper and lower trees. When a file exists only in a lower tree and a write-like operation needs a writable upper object, OverlayFS creates the corresponding upper object and copies up the lower object’s metadata and, for regular files, data. Future access then uses the upper file. This is the ordinary copy-up path and is the behavior most administrators expect when they first reason about a container’s writable layer.
A metadata-only operation such as chmod or chown may not need the file’s content to change. Without metadata-only copy-up, that operation can still trigger a full copy. For a large file that receives an ownership adjustment during initialization and is never written afterward, the data movement can dominate startup time and upper-layer storage. The benefit of metacopy is to defer that content copy until there is an actual write-open that requires it.
With metacopy enabled, the upper object initially contains updated metadata but no file data. OverlayFS marks it with a trusted.overlayfs.metacopy extended attribute. Reads can continue to obtain the file’s data from the lower layer. When the file is opened for writing, the data is copied up, and the metacopy marker is removed. This is delayed copy-up, not a permanent overlay that writes data back to the lower layer.
The optimization helps only when the workload does metadata work without soon modifying file content. If every file is immediately rewritten, the data copy has not been avoided; it has merely happened later. In that case the feature may shift latency from initialization to the first write and make the copy-up point less obvious in an application trace.
Evaluate the filesystem and mount prerequisites
OverlayFS uses upper, lower, and work directory trees. The lower tree may be read-only. A writable upper filesystem must support the extended attributes OverlayFS needs and return valid directory-entry type information; not every Linux-mountable filesystem is suitable for a writable upper. The workdir must be on the same filesystem as upperdir. Validate these properties before tuning copy-up behavior, and do not infer suitability merely because a mount command succeeded once on a different host.
The following is an isolated lab example, not a command to run against a container runtime’s live storage directories. It assumes that the directories are controlled, the kernel includes OverlayFS with metacopy support, and the backing filesystems meet OverlayFS requirements. upper and work must reside on the same filesystem.
mkdir -p /srv/ovl-lab/upper /srv/ovl-lab/work /srv/ovl-lab/merged
findmnt -T /srv/ovl-lab/upper
findmnt -T /srv/ovl-lab/work
mount -t overlay overlay \
-o lowerdir=/srv/immutable/base,upperdir=/srv/ovl-lab/upper,workdir=/srv/ovl-lab/work,metacopy=on \
/srv/ovl-lab/merged
Run this in a disposable Linux environment with the privileges required to mount filesystems. Use test content, not production data. A managed container runtime may create its own overlay mounts and storage metadata; do not change mount flags behind the runtime or edit its upper/lower trees while containers are active. Configure supported runtime options through the runtime’s documented interface, if it exposes them, and verify the actual mount rather than assuming the host mount configuration was inherited.
Test the delayed copy-up behavior
Create a lower-layer file large enough to make unnecessary copying measurable. In a disposable mount, apply only a metadata change through the merged path, inspect the upper tree and its OverlayFS xattrs with sufficient privilege, and verify that the upper object is marked as metadata-only. Then read the file and confirm that the content remains accessible. Finally, open it for writing, change a byte, and inspect again to confirm that content copy-up occurred and the metacopy marker is gone.
The xattr namespace depends on mount configuration. The default trusted xattr namespace requires the relevant privilege to inspect; a mount using userxattr uses the user.overlay. namespace instead. The marker can also be absent because the file has already been fully copied up. An xattr listing is one piece of evidence, not proof of byte-for-byte integrity or of which lower tree supplied the data. Compare the merged content against the intended immutable lower content and record the mount’s actual options.
Measure both sides of the tradeoff. Record the time for metadata-heavy initialization, bytes allocated in the upper layer, latency of the first subsequent write, and the steady-state write path. Repeat with the same file and filesystem options with metacopy disabled. Sparse files, compression, reflinks, filesystem accounting, and runtime-specific behavior can change apparent allocation, so use the measurements from the target stack rather than relying on du alone. Use application-level tests that detect content changes, not just a faster startup timestamp.
Useful checks on an already mounted test filesystem include:
findmnt -no SOURCE,FSTYPE,OPTIONS --target /srv/ovl-lab/merged
getfattr -d -m 'trusted.overlay.*' -- /srv/ovl-lab/upper/path/to/test-file
If the mount uses userxattr, inspect the corresponding user.overlay.* namespace instead. These commands are diagnostic examples; the exact xattr visibility depends on privileges and kernel configuration. Do not modify OverlayFS control xattrs by hand as a repair technique. A crafted redirect or metacopy marker can change which lower data file is exposed.
Treat layer trust as a correctness and security boundary
The kernel documentation explicitly warns against enabling metacopy=on when the upper or lower directories are untrusted. An attacker able to create appropriately crafted OverlayFS redirect and metacopy extended attributes could direct access to content in a lower layer. The kernel protects trusted xattrs from ordinary unprivileged writes in standard configurations, but imported or externally prepared layers may cross trust boundaries in ways that are not equivalent to an ordinary local process setting an xattr.
Only use metacopy with upper and lower trees whose ownership, provenance, and mutation permissions are controlled. Verify downloaded or imported layers before mounting them. Do not let untrusted users populate the upper tree, and do not assume that a removable drive, shared build cache, or extracted archive is safe merely because it is mounted read-only. If the data or xattr provenance is uncertain, keep the optimization disabled until the layer can be validated and isolated.
There is no mount-time verification mechanism that makes reusing a metacopy upper layer with a different set of lower layers safe. The kernel documentation warns that such a mount may succeed but produce unexpected behavior later. Keep the lower-layer identity stable for the lifetime of an upper layer. If the base image changes, create a new overlay instance or follow the container/storage runtime’s supported upgrade and cleanup procedure rather than reusing a writable upper tree with a different base.
Respect option conflicts and lower-layer immutability
The kernel documentation lists conflicts between metacopy=on and nfs_export=on or several redirect_dir modes. Incompatible options cause mount failure. Do not respond by removing options blindly: they may exist to support a required NFS export or directory-rename behavior. Determine which feature the workload needs, consult the kernel documentation for the exact kernel in use, and test the complete option combination in a staging environment.
More importantly, do not modify an underlying tree while it is part of a mounted OverlayFS. The kernel documents behavior as undefined if a lower or upper filesystem is changed while mounted. Offline lower-layer changes are also restricted once features including metacopy, index, xino, or redirect handling are in use. This makes immutable, content-addressed base layers a natural operational fit: build a new lower-layer version, create a new overlay, validate it, switch traffic or mount ownership, and retire the old instance only after it is unused.
Do not conflate metacopy with data-only lower layers. Data-only lower layers are an advanced arrangement in which content can come from layers that do not contribute visible path names or metadata. The kernel documents special double-colon syntax for such lower layers and a newer file-descriptor mount API for configuring them. This is a distinct feature with its own redirect and privilege requirements; it should not be added to a routine container mount as a casual extension of the metadata-copy optimization.
Diagnose unexpected space, latency, or content behavior
When the upper layer grows unexpectedly, check what operation triggered copy-up before concluding that metacopy is broken. A file opened for writing will need data copy-up even if the first operation is a small change. A workload that performs chown -R and then rewrites every file may simply move the copy cost into a later phase. Compare the sequence of metadata calls, open modes, actual writes, and upper-layer allocations.
If a mount fails, capture the exact error, kernel version, filesystem types, mount options, and relevant kernel log messages. Check for workdir/upperdir placement and unsupported option combinations before changing the lower data. If a mount succeeds but presents unexpected content, stop writing to that overlay, preserve the upper and lower trees, and reconstruct the layer identity and xattrs. Do not edit redirect or metacopy xattrs to force a mount to appear correct; that can expose the wrong lower content or make the state harder to recover.
For a container host, collect the runtime version, storage driver configuration, kernel version, mount namespace, and findmnt output from the namespace that owns the overlay. A host-level mount listing may not show the same view as the container process. Reproduce with an isolated test image and fresh upper/work directories before changing a shared storage configuration.
Operational checklist
- Is the performance profile actually metadata-heavy, and do measurements show that data copy-up is avoidable?
- Are upper, lower, and work directories controlled and supported by the active kernel/filesystems?
- Are
upperdirandworkdiron the same filesystem? - Are imported layers and xattrs trusted before
metacopy=onis enabled? - Is the lower-layer identity immutable for the lifetime of the upper layer?
- Are mount options compatible with the required redirect and NFS-export behavior?
- Have metadata-only changes, reads, first writes, and recovery paths been tested separately?
- Are operators inspecting the mount namespace and exact runtime configuration rather than changing live runtime storage by hand?
- Is the fallback a fresh, verified overlay instead of an unreviewed manual xattr edit?
metacopy is a targeted way to delay file-data copy-up when the workload changes metadata more often than content. Its value comes from a measured reduction in unnecessary work; its risk comes from coupling upper metadata to lower-layer data. Keep layers immutable, trust every xattr-bearing tree, test the later write path, and make the mount configuration reproducible before enabling it outside a lab.
Related:
- Setting Up Bind Mounts and Overlay Filesystems
- Linux’s File-Descriptor Mount API: fsopen, fsmount, and move_mount
Sources: