Windows Cloud Files API: Build a Sync Root with Correct Hydration Semantics
Design a Windows Cloud Files provider around sync-root registration, placeholder state, hydration callbacks, cancellation, and recoverable transfer operations.
The Windows Cloud Files API (CfAPI) is the operating-system contract for presenting remote files as ordinary entries in a local sync root while fetching file content only when needed. It is not just a way to tag a file as “in the cloud.” The platform coordinates placeholder metadata, hydration policy, file operations, shell state, and calls between the filter driver and a desktop sync provider. A provider still owns remote authentication, versioning, upload, conflict resolution, integrity, and offline behavior.
This distinction matters because an implementation can appear correct in a happy-path demo and fail under a real workload: an application seeks to the end of a file, a user pins it during a download, antivirus opens it concurrently, the network request is canceled, or the provider restarts while a callback is pending. Production design begins with explicit state transitions and durable operation identity, not with a file icon.
Define the provider and sync-root boundary
A sync root is the namespace boundary registered with Windows for a provider and a user. Registration associates the root path with provider identity and policies and allows the shell to expose a branded location and hydration controls. Register only a directory the product owns. Before changing registration, confirm the current user, package or provider identity, path, volume, and existing registration state. A second provider must not silently adopt a root whose database or placeholders belong to another provider.
CfAPI uses the Cloud Filter minifilter (cldflt.sys) to mediate file-system operations and communicate with a user-mode provider. The documented design currently depends on NTFS features, so validate the target volume and Windows version before creating a root. Do not infer support from the fact that an ordinary directory can be created on the volume. Product packaging, desktop-app lifecycle, user sign-in, provider startup, and upgrade/uninstall behavior all need to agree with the registration model.
Registration is a stateful operation, not a per-launch toggle. Define how the application handles first registration, update of existing provider metadata, temporary disconnect, clean sign-out, uninstall, and recovery after a crash. A service or desktop process must call the supported connection API and remain available for callbacks that require remote data. When the provider is unavailable, Windows cannot invent remote bytes; describe that limitation to users and return bounded, diagnosable errors instead of leaving file operations indefinitely blocked.
Model placeholder state separately from synchronization state
A placeholder contains file-system metadata and enough provider identity for Windows to represent an item without storing its full data stream locally. Hydration materializes requested content. Pinning is a user’s availability preference, not proof that the remote copy is current. “In sync” is a provider assertion about a particular state, not an automatic consequence of a successful download. Keep these concepts separate in the provider database and UI.
Use immutable remote object IDs and revision IDs rather than a pathname alone. Paths can be renamed or case-changed, and a cloud object can be replaced while retaining the same name. Store a versioned, bounded identity blob in placeholder metadata so the provider can map a callback back to the intended remote revision after an application restart. If the remote version changes, reconcile metadata using the supported placeholder-update operations; do not serve new remote bytes under stale size, timestamp, or content identity.
Create placeholders in batches where appropriate, but inspect per-entry results. CfCreatePlaceholders can process several entries and return an HRESULT for the call as well as a result for each placeholder. A batch-level success is not a substitute for checking every entry. Treat name collisions, invalid parent paths, access errors, and a lost sync-root registration as distinct outcomes. Persist the remote listing cursor or generation so a retry can be idempotent rather than creating duplicate local state.
Directory population also needs a policy. A provider can populate a directory eagerly or on demand depending on the product’s experience and documented population policy. Large trees should not require a full remote crawl every time Explorer opens a folder. Bound enumeration work, preserve remote consistency across pages, and make cancellation cheap. A partially returned directory listing must not be mistaken for an authoritative deletion set.
Choose hydration behavior as an application contract
CfAPI exposes primary hydration policies with different tradeoffs, and Windows computes effective behavior from provider and application preferences when the file is opened. A provider cannot assume that every read arrives as a single sequential request. Progressive hydration lets applications consume ranges as they are fetched; full hydration downloads the whole file before satisfying the open; always-full restrictions can prevent placeholder operations that would leave a file partial. Select policy based on what the service can deliver and what the application workload expects.
Progressive access requires correct random-range semantics. A media player, archive tool, indexer, antivirus engine, or office application may read a header, seek elsewhere, and return. Implementing a stream as “download the next byte after the last one” is insufficient. Validate the requested offset and length against the remote object’s current size, request exactly the needed range, and transfer only bytes whose integrity and revision are known. Respect the API’s documented alignment and end-of-file rules for hydration ranges; do not round a request in a way that writes bytes from a different object version.
Hydration policy is fixed for the relevant open. Changing a provider setting during an incident does not retroactively change an already opened handle’s policy. When a workload unexpectedly downloads whole files, compare the application preference, provider policy, placeholder flags, and which code path opened the file. A shell pin action and an application’s open policy can produce different requests even for the same path.
Treat callbacks as concurrent, cancelable work
After connecting the sync root, the provider receives callbacks for requests such as fetching file data, validating data, and reacting to placeholder operations. A callback is an operation with an identifier and parameters; it is not a durable queue record by itself. If the provider completes work asynchronously, keep the operation identifier and connection key valid until the matching completion call, handle cancellation, and never complete the same operation twice.
Keep callback dispatch separate from network and disk I/O. Use a bounded worker queue, per-request deadline, byte limit, and cancellation token. A single global lock held across a remote request can serialize unrelated file opens and make Explorer appear frozen. Conversely, unlimited workers can overwhelm a network endpoint, exhaust sockets, or cause the provider to compete with the foreground application. Use per-file or per-object coalescing only when the operation identity and requested ranges remain correct.
Cancellation is a normal state transition. Windows may cancel a fetch because the requesting handle closes or the user changes the operation. The remote request may already be in flight; cancellation should stop work where possible and discard late data safely. If the remote server returns bytes after cancellation, do not write them into a placeholder without a still-valid transfer context. During shutdown, stop admitting new work, drain or cancel outstanding callbacks, disconnect using the documented API, and only then free callback state.
Build an integrity and conflict boundary
For each data transfer, bind together the sync-root identity, placeholder identity, remote object revision, requested byte range, callback operation, and expected size. A successful HTTP response alone does not prove that the content corresponds to the placeholder. Verify the remote ETag or immutable version, expected range, response length, and cryptographic digest where the service protocol provides one. If an object changes during transfer, restart or report a conflict according to policy rather than merging bytes from two versions.
Local writes, renames, deletes, and pin changes can race remote synchronization. Decide whether the product supports read-only placeholders, local edits with upload, or application-managed conflict copies. Track dirty state and upload acknowledgments outside an ephemeral callback. If an upload fails, preserve the local file and present the retry state. Never dehydrate or overwrite unsynchronized local data just because the provider’s remote catalog says the file exists.
Coordinate with file-system filters such as antivirus, backup, and indexing software. Their reads can trigger hydration and add pressure to the provider. Set product policies and diagnostic logging to distinguish user opens from background scans where the platform provides that evidence; do not claim that every access can be reliably attributed. Test simultaneous opens, memory mapping, partial reads, file rename, delete, pin/unpin, and a provider restart while an operation is outstanding.
Triage a sync-root incident without forcing downloads
Start with non-invasive evidence. Record Windows build, provider version, user identity, sync-root path, filesystem, registration result, provider process state, and time. Inspect a small, known set of placeholder paths with metadata-only operations. Operations such as hashing, reading file content, opening an application, or recursively scanning a tree can hydrate data and change the system under investigation. Explain this before running a test that reads content.
For provider-side diagnostics, log callback type, operation ID, connection key, relative path or privacy-safe object ID, range, remote revision, result code, elapsed time, cancellation, and bytes transferred. Avoid logging access tokens, file contents, or sensitive names by default. Correlate provider logs with Windows event records and application activity using UTC timestamps. A missing callback can be a registration/connection problem; a callback that returns failure can be remote auth, revision, quota, path, or provider logic. Use different counters for these conditions.
Before repair, capture the registered root and a representative placeholder’s metadata. Re-registering or deleting a root can affect shell state and user data. Do not remove cldflt.sys, delete provider-owned reparse data, or run an indiscriminate cleanup script. Follow the documented CfAPI lifecycle and the provider’s recovery procedure, beginning in a test profile or copy of the root.
Validate lifecycle and failure paths
The test matrix should include cold sign-in, offline startup, reconnect, token expiry, large directory pagination, random-range reads, file growth or replacement, unpinning, full-disk conditions, duplicate names, a provider crash, application cancellation, antivirus access, and OS restart. Verify that retries are bounded and idempotent and that a local unsynchronized write remains recoverable. Include upgrade and uninstall because stale registrations can leave a shell root that no longer has a working provider.
Measure first-directory latency, time to first byte, full hydration duration, request count, bytes read versus bytes requested, and the age of the metadata shown to the user. A service can be functionally correct but unusable if a shell operation blocks behind an unbounded remote request. Define which operations are allowed to wait, which can return a retryable error, and what progress the user sees.
Provider release checklist
- Confirm supported Windows versions, NTFS requirement, provider identity, and sync-root ownership.
- Select hydration and population policies from actual workload behavior, not assumptions about sequential reads.
- Bind each placeholder to stable remote object and revision identifiers.
- Process per-placeholder results and model callback cancellation, deadlines, and shutdown.
- Protect dirty local content and define remote-change conflict behavior.
- Test pinning, offline access, random reads, filters, crashes, upgrades, and uninstall.
- Keep telemetry useful while excluding tokens and file contents.
CfAPI supplies the integration layer that makes remote content behave like a local namespace. Reliability still depends on an explicit state model, correct range semantics, a provider that can recover from interrupted work, and a user experience that does not confuse “present,” “available offline,” and “synchronized.”
Related:
- Windows ProjFS: Building a Filesystem Projection Provider That Hydrates on Demand
- Windows Reparse Points: Safe Traversal, Inspection, and Cleanup
Sources: