Linux F2FS Garbage Collection: Foreground Stalls, Background Cleaning, and Space
Diagnose F2FS segment cleaning by separating foreground and background GC, checking free-space pressure, discard behavior, and workload latency.
F2FS is a filesystem designed for flash-based storage. Like other log-structured designs, it writes new data and metadata into segments and later reclaims segments that contain invalidated blocks. Garbage collection (GC) selects a victim segment, moves still-valid data elsewhere, then makes the segment reusable. This can create bursts of I/O and latency, especially when free space is low or the workload keeps invalidating data in ways that leave few inexpensive victims.
The common operational mistake is to interpret every latency spike as an SSD failure or to disable background GC without understanding the resulting allocation path. Diagnose the filesystem’s free-space pressure, GC mode, discard stack, block-device behavior, and workload write pattern as a single system.
Segment cleaning and write amplification
F2FS maintains multiple logs for data and metadata classes. Updates are generally written to new locations rather than overwriting the old block in place. The previous copy becomes invalid, and the filesystem later cleans segments to recover capacity. Cleaning cost depends on how many valid blocks must be copied from a candidate segment and how much metadata must be updated.
The flash translation layer beneath the filesystem also performs its own reclamation. Filesystem GC and device-level flash management are separate layers, but their work interacts. Discard/TRIM can inform the lower device that logical blocks are no longer needed; it does not guarantee that the device will immediately erase cells or that a future write will have a fixed latency. If discard is disabled or unsupported through a layer, the device may retain mappings for filesystem-free blocks until another mechanism communicates them.
F2FS supports background cleaning when the I/O subsystem is idle and can perform foreground GC when an allocation needs space. Foreground work can be visible directly in the writer’s latency. Background GC aims to do useful cleaning before urgent allocation pressure, but it competes with workload I/O and is not a guarantee that foreground GC never occurs.
Establish filesystem and device context
Before changing mount options, capture the filesystem type, mount options, kernel release, device stack, and capacity:
findmnt -no SOURCE,FSTYPE,OPTIONS /mount/point
df -hT /mount/point
df -i /mount/point
lsblk -o NAME,TYPE,SIZE,FSTYPE,MOUNTPOINTS
journalctl -k -b --no-pager | grep -i -E 'f2fs|discard|I/O error|timeout'
Replace the mount point with the actual path. df reports filesystem allocation, not the SSD’s physical wear or flash translation layer state. lsblk shows a block-device topology but does not reveal every controller cache or internal FTL behavior. Record whether the filesystem resides on raw flash, eMMC, a virtual disk, a device mapper layer, or a loopback image.
F2FS-specific sysfs interfaces and statistics depend on kernel version and configuration. Enumerate available attributes under the mounted device’s F2FS entry rather than assuming one set of counter names. Read-only status files are useful for trends, but do not treat a single count as a direct measure of GC health. Correlate deltas with write rate, free segments, foreground GC events, I/O latency, and device errors.
Inspect mount policy and GC mode
The F2FS documentation describes background_gc choices for controlling cleaning, including background operation and synchronous behavior. The exact behavior should be confirmed against the documentation for the running kernel. Disabling background GC can postpone cleaning until foreground allocation pressure, which may move work onto application writes and create sharper latency spikes. Selecting synchronous background behavior can also alter timing and contention. Neither option is a universal performance fix.
The gc_merge option, when supported and when background GC is enabled, allows the background GC thread to handle foreground GC requests. It may reduce sluggish behavior in some circumstances, but it is still a policy choice that should be tested under the actual CPU and I/O constraints. Other F2FS options affect discard, age-threshold victim selection, checkpointing, and zoned devices. Do not stack several options in one change and attribute the result to only one.
discard behavior deserves particular care. Mount-time and periodic discard paths may differ, and the lower device must support discard propagation. A discard request is not secure deletion and does not prove physical erasure. In stacked storage, inspect each layer’s advertised support and the filesystem documentation before enabling continuous discard. A periodic fstrim workflow can have a different performance profile from issuing discard with each free operation.
Diagnose latency spikes
Capture application write latency and block-layer latency at the same time. A brief pause aligned with a foreground GC event suggests segment cleaning is in the writer path, but correlation is not proof that the SSD is healthy or that F2FS is the only cause. Compare the block device’s queue activity, controller error logs, writeback pressure, and device temperature when available.
Check available space and workload churn. A nearly full filesystem can have fewer low-cost segments and less room to relocate valid data. Large amounts of small random overwrites, snapshots at an upper storage layer, or a workload that retains many recently written blocks can change victim costs. Deleting files may not immediately create a large contiguous free segment; filesystem and flash-layer reclamation are distinct operations.
Check whether the observed latency follows a workload phase: mass file deletion, database compaction, package update, checkpointing, or a sudden increase in random writes. A lower average latency with worse high percentiles may indicate cleaning moved to quieter periods but still contends with the application. Record p50, p95, p99, throughput, and tail-event duration instead of relying on a single average.
If the system reports read/write errors, timeouts, or controller resets, investigate hardware and block-layer health before changing GC policy. Filesystem cleaning cannot correct media errors, power loss, or a broken device firmware path. Likewise, a high filesystem used percentage is not sufficient evidence of an F2FS bug.
Safe test plan
Use a disposable test filesystem or a representative canary. Keep the existing mount options and baseline measurements. Replay a bounded workload that includes sequential writes, small random overwrites, file deletion, and the production application’s update pattern. Record free space, write amplification proxies, GC activity, block I/O latency, discard counters, and application tail latency.
Change one mount option at a time and keep the same kernel, device, data set, and write pattern. A remount may not change every option, and some options are only read at mount time; verify the effective state rather than assuming a command succeeded. Never run filesystem repair or formatting commands on the production volume as part of performance testing.
If using a benchmark, ensure the test is not limited by a page cache or an artificial device that lacks the target SSD’s FTL behavior. Do not issue global cache-drop commands on a live host to create a cold-cache test. Use a disposable image or lab device and maintain enough free space to avoid testing only the emergency allocation path.
Common interventions to avoid
Turning off GC because it appears in a latency trace. This can shift cleaning into foreground allocation and make stalls worse. First determine whether background activity is competing with the workload or preventing urgent cleaning.
Adding discard options without checking the storage stack. The request may not propagate, may be unsupported, or may change latency. Validate every layer and compare periodic versus continuous behavior when appropriate.
Assuming more free bytes automatically fixes segment pressure. Filesystem block availability and clean segment availability are related but not identical. Observe F2FS state and workload churn.
Running repair tools during a performance incident. fsck.f2fs and repair modes are for specific offline recovery procedures, not runtime tuning. Follow vendor and filesystem recovery guidance and preserve data first.
Comparing different devices or kernels. Flash controller firmware, overprovisioning, write cache, discard behavior, and kernel changes all affect results. Record them as part of the experiment.
Acceptance and operations
Keep a change record with device topology, kernel and F2FS versions, mount options, filesystem utilization, discard capability, workload profile, and latency distributions. Accept a GC-policy change only if it reduces the target tail latency or throughput cost over repeated tests without increased errors, unexpected space exhaustion, or degraded durability. Test restart, unmount/remount, and recovery behavior in the same environment.
For production, alert on filesystem capacity, I/O errors, device timeouts, and application latency. Retain periodic measurements so a gradual change in free-space behavior can be distinguished from a sudden hardware fault. Revisit the policy after kernel, firmware, or workload changes.
F2FS garbage collection is normal maintenance for a log-structured filesystem. The engineering problem is where that work occurs and how much data a victim segment requires. Measure those effects across the filesystem and flash stack before changing the policy that schedules cleaning.
Related:
- Btrfs Scrub: Verify Checksums and Repair from Redundant Copies
- XFS Online Scrub: Metadata Validation, Repair Boundaries, and Recovery
Sources: