Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux NFS Client Caching, Delegations, and State Recovery

Troubleshoot Linux NFS client cache behavior and recovery by separating attribute caching, close-to-open visibility, protocol state, and mount evidence.

An NFS-mounted path looks like a local filesystem path to many applications, but its caching and failure behavior span a client, a network, and a server. Linux NFS clients cache data and metadata to reduce round trips; the server remains authoritative for shared filesystem state. Applications that assume every local stat() or directory lookup immediately reflects a remote change can therefore behave unexpectedly. Protocol state, delegations, leases, client identity, and network recovery add a second class of behavior beyond ordinary cache freshness.

The right investigation separates several questions: is the application reading cached file data, cached attributes, or a cached directory entry; has the server acknowledged a write; is the client holding protocol state that must be recovered; and does the server/export provide the semantics the application expects? Mount options change tradeoffs but do not turn NFS into a globally coherent local filesystem. Validate the workload’s actual consistency requirements before tuning caches.

Inventory the mount and client state

Start with the effective mount options and kernel view:

findmnt -t nfs,nfs4 -o TARGET,SOURCE,FSTYPE,OPTIONS
nfsstat -m
cat /proc/self/mountstats

findmnt and nfsstat output can differ in formatting and available counters. Save it before changing anything. Record NFS version, transport, server address, mount options, client kernel release, and whether the mount is hard or soft. A hostname in a mount source may resolve to multiple addresses or change over time; compare the actual endpoint with DNS and server-side logs.

Use nfsstat -c and the relevant mount’s statistics to identify retransmissions, RPC timeouts, operation mix, and throughput over an interval. Counters are cumulative and need before/after deltas. They do not directly identify a particular application’s operation. Correlate with process-level I/O, server metrics, network captures where permitted, and application logs.

Distinguish data, attribute, and directory caching

NFS clients cache file data and metadata. Attribute cache timers influence how often attributes are revalidated, while directory cache behavior affects lookup and readdir visibility. Options such as acregmin, acregmax, acdirmin, acdirmax, and actimeo adjust cache intervals. Shorter intervals can improve freshness in some workloads but increase metadata RPC load and still do not establish a universal distributed transaction boundary.

The noac option disables attribute caching and also forces application writes to be synchronous, which can impose a significant performance penalty. It still does not eliminate data caching or every application-level consistency issue, so it should not be applied as a reflexive fix for a stale listing. Diagnose which observation is stale: an already open file descriptor, a new stat, a directory listing, a name lookup, or content read through an application cache. These use different paths and can be affected differently.

Many NFS deployments rely on close-to-open behavior for common file-sharing patterns: closing a file and opening it again causes the client to revalidate as appropriate. This is not equivalent to strict cache coherence for concurrent writers that keep a file open. Applications requiring concurrent transactional updates need application-level coordination, locking, or a storage design with explicit semantics. Compare behavior using two clients and controlled writes before changing production cache options.

Understand write visibility and durability

An application write returning does not automatically mean that every client has observed the new bytes. NFS write stability and COMMIT operations help define server-side durability behavior, while client writeback and cache revalidation govern visibility. The exact guarantees depend on protocol version, mount options, server implementation, and operation type. Consult the NFS protocol/server documentation and application durability requirements before making claims about when data is safely on persistent media.

When investigating “write succeeded but peer sees old content,” record the file descriptor lifetime, whether each application closed and reopened the file, file size and mtime observations, NFS RPC errors, and server-side file state. Check whether an application or runtime has its own caching layer. A client-side cache timer will not explain data served from a process cache, and invalidating an OS cache may not affect that process.

Avoid drop_caches as an NFS consistency tool. It affects local caches broadly, is not a guarantee that all remote state is refreshed, and can distort performance. Prefer a controlled test with new processes, explicit close/reopen boundaries, documented mount settings, and observation on both client and server.

NFSv4 state, delegations, and recovery

NFSv4 adds stateful operations such as opens, locks, delegations, and sessions. A delegation allows a client to cache certain state until it is recalled or otherwise invalidated under the protocol. Server reboot, network partition, client restart, or identity changes can require recovery and state reclaim. During recovery, applications may see delays or errors even when the mount eventually becomes usable.

Client identity must remain stable and unique according to the kernel client implementation and deployment. The Linux client documentation describes nfs4_unique_id for distinguishing clients in certain scenarios. Do not set a unique ID casually on every boot or clone a configured identity across multiple live clients; changing identity can affect recovery, and collisions can cause state confusion. Verify the option’s semantics against the running kernel documentation and coordinate changes with server administrators.

Delegations and state recovery should be diagnosed with kernel logs and server-side NFS events. Look for lease expiration, reclaim, grace-period, callback, session, and transport messages. Message strings differ by kernel and server versions, so capture the entire time window rather than searching only one literal phrase. A mount that responds to ls does not prove every open file’s lock state was recovered correctly.

Choose failure semantics deliberately

Hard mounts generally continue retrying operations when the server is unavailable, so processes can block during outages. Soft-style behavior can return errors to applications after retries, which may be unsafe for workloads that do not correctly handle partial or ambiguous I/O failures. Understand the documented behavior of the exact mount options and NFS version. Do not switch from hard to soft solely to make a hung command return faster; that changes application error semantics.

Timeout and retransmission settings influence retry timing and server load. Tune them only after characterizing network latency, failover duration, server recovery, and application deadlines. Aggressive retry settings can amplify an outage, while long retries can exceed an application’s useful request window. Monitor server availability, client RPC statistics, TCP behavior, and process blocked tasks together.

For an incident, avoid force-unmounting an active filesystem before identifying processes and data operations. A forced unmount can hide which operations were outstanding and complicate application recovery. Capture findmnt, mountstats, kernel logs, process wait channels, and server-side state first when feasible. Follow site runbooks for unmount and remount procedures.

Build a consistency test and acceptance criteria

Use two clients and a disposable export to test the exact application pattern: client A creates and writes a file, closes it, client B looks it up and reads it, and both clients repeat with concurrent writers and open descriptors if the workload uses them. Record timestamps, mount options, NFS version, server logs, and RPC counters. Include directory rename/unlink and lock behavior only if the real workload depends on them.

Test server interruption and recovery in a staging environment. Confirm how long client calls block, whether operations return errors, whether open/lock state is reclaimed, and whether application retries can duplicate side effects. Restore the server, then validate application data. Do not equate “mount is responsive again” with “application state is correct.”

An acceptance record should identify the required visibility model, the tested cache options, failure behavior, application-level locking or coordination, and server durability assumptions. If a workload needs stronger coordination than NFS caching provides, fix it at the application or architecture layer rather than indefinitely shortening attribute timers.

Linux NFS behavior is a protocol and cache contract, not merely a mount command. Separate data from metadata visibility, inventory real client options, observe state recovery, and test the application’s concurrency pattern. Doing so turns vague stale-file reports into evidence about which cache or protocol boundary needs attention.

Related:

Sources:

Comments