Skip to content
macOSDeep Dive Published Updated 7 min readViews unavailable

AVAssetReader on macOS: Sample Providers, Timing, and Cancellation

Read AVFoundation media incrementally with typed outputs, async sample providers, precise timestamps, bounded processing, and explicit failure handling.

AVAssetReader is AVFoundation’s lower-level API for extracting timed media from an asset. It can vend samples from a file-backed asset or a composition, but it is not a general-purpose decoder callback, an editor, or a file exporter. The application chooses a track or composition output, chooses an output representation, consumes samples, and decides how much work and memory to retain. A robust reader pipeline therefore treats format descriptions, timestamps, completion, and cancellation as part of the data contract.

Current AVFoundation documentation exposes output providers and asynchronous next() calls for consuming samples. Older code often uses add(_:), startReading(), and copyNextSampleBuffer(); these APIs are deprecated in the current documentation. If an application must support older deployment targets, isolate that compatibility path and test it separately rather than mixing both lifecycle models in one loop.

Decide which representation the consumer needs

An AVAssetReaderTrackOutput reads one track. Audio settings can request linear PCM, while video settings can request uncompressed pixel buffers. A nil output-settings value requests the track’s stored sample representation rather than a decoded working format. That can be useful for packet inspection, passthrough, or handing compressed media to another compatible stage; it is not automatically the right choice for image processing. In this mode the output skips decoding and returns samples in decode order, which may not match display order for inter-frame video.

Output format affects CPU cost, memory bandwidth, and fidelity. Decoding video into a convenient RGB format may introduce a conversion and a larger frame footprint. Apple documents choosing a decoder-supported pixel format to avoid unnecessary conversion. When the next stage accepts the source representation, preserving it can be cheaper. When code needs pixel access, explicitly request a supported uncompressed format and account for plane layout, row stride, range, bit depth, and color metadata. A pixel buffer is not necessarily tightly packed RGB bytes.

The reader’s output type also determines what transformations happen. A track output reads one track. Audio-mix and video-composition outputs represent different processing paths. Reading raw video frames does not apply a custom edit just because an AVComposition exists elsewhere in the app. Select the output that matches the intended timeline and inspect its format description instead of inferring type from a filename.

Configure before starting and consume incrementally

Load the asset properties the operation requires before building the reader. A remote or protected asset may fail while loading its tracks. Select a track intentionally: taking .first is acceptable for a small example, not a production policy for files with commentary, alternate languages, multiple camera angles, or disabled tracks. Preserve a selected track identifier or the user’s language choice when the feature depends on a specific stream.

The provider attaches an output to the reader and delivers ready sample buffers asynchronously. try await provider.next() returns one sample at a time and returns nil when there is no further sample. The application should process a sample and release references before requesting or retaining large additional batches. Avoid collecting an entire movie’s sample buffers into an array: a long 4K stream can turn an otherwise incremental reader into an avoidable memory spike.

import AVFoundation

func scanVideoTimestamps(at url: URL) async throws -> [CMTime] {
    let asset = AVURLAsset(url: url)
    let tracks = try await asset.loadTracks(withMediaType: .video)
    guard let track = tracks.first else {
        throw ReaderFailure.noVideoTrack
    }

    let reader = try AVAssetReader(asset: asset)
    let output = AVAssetReaderTrackOutput(track: track, outputSettings: nil)
    let provider = reader.outputProvider(for: output)
    try reader.start()

    var presentationTimes: [CMTime] = []
    while let sample = try await provider.next() {
        presentationTimes.append(sample.presentationTimeStamp)
    }

    guard reader.status == .completed else {
        throw reader.error ?? ReaderFailure.incompleteRead
    }
    return presentationTimes
}

enum ReaderFailure: Error {
    case noVideoTrack
    case incompleteRead
}

The example deliberately returns only timestamps, but even that array grows with duration. For real processing, replace the array with a bounded consumer, accumulator, or streaming writer. If a consumer is slow, backpressure should slow the read loop; launching one detached task per sample merely moves an unbounded queue elsewhere. Attach explicit limits to any concurrency window and make the consumer’s ordering requirements clear.

Preserve media time, not loop position

Sample buffers carry timing and format information. Use presentation timestamps for display order, decode timestamps where decode dependencies require them, and durations only when the source provides meaningful duration information. Do not derive time from a frame counter multiplied by an assumed frame interval. Variable-frame-rate footage, edits, gaps, and audio packetization make that arithmetic unreliable.

For synchronization, compare values as CMTime and convert into a chosen time scale only at boundaries that require it. Floating-point seconds are useful for UI labels but lose exact rational timing if repeatedly used as a transport format. When writing samples downstream, preserve source presentation and decode timing unless the transformation explicitly changes the timeline. If trimming, scaling, or reordering content, define the mapping from source time to output time and test boundary samples.

A sample buffer’s format description can change at a format transition. A pipeline that caches dimensions, channel layout, or codec parameters once at startup can process later samples incorrectly. Inspect descriptions where transitions are possible and rebuild downstream converters when the format contract actually changes. Attachments can carry information that ordinary timestamps do not capture, including per-sample flags and color-related metadata.

Separate end of stream from failure

An exhausted provider and a successful reader status are not the same assertion. Check the reader’s final status and surface its error if the operation failed. Cancellation is a separate result, not an empty successful asset. A caller that intentionally stops after a selected time range should report that policy distinctly from a source that ended unexpectedly.

Cancellation should be connected to the owning task or user operation. When a task is cancelled, stop the reader and let the provider unwind; do not leave a decode operation running after the user closes the document. Ensure cleanup is idempotent because cancellation can race with natural end-of-stream or a read error. A useful state model records whether the pipeline reached completed, failed, or cancelled, plus the asset URL, selected track, output settings, and sanitized error domain/code.

Time ranges and random access

Set a reader time range before starting when the job needs a segment rather than the whole asset. Define whether endpoints are inclusive or exclusive in the surrounding product contract, especially when concatenating adjacent ranges. For random-access workflows, use the documented random-access output-provider API and its controller instead of assuming a linear provider can be rewound. Seekable sample extraction still has decoder preroll and format constraints; request the frames the decoder needs and discard output outside the final presentation interval when appropriate.

If a user scrubs repeatedly, cancel obsolete reads and attach a generation number to results. A late sample from an earlier seek must not replace a newer preview. For a background analysis job, make the asset and range part of the job identity so a retry cannot accidentally reuse stale output from another source revision.

Operational failure modes

  • Unsupported output settings: Validate that the requested representation is supported by the track and current system. Fall back only when the product can accept the alternate format; do not silently change color or precision.
  • Track selection errors: Log track IDs, media types, and relevant language metadata, not user content. Test multi-track assets rather than only single-track phone recordings.
  • Memory growth: Bound sample retention, pixel-buffer pools, and downstream work. Measure peak resident memory on long assets and high-resolution frames.
  • Timing drift: Compare emitted sample timestamps against expected ranges and the downstream writer or renderer timeline. Do not equate array index with presentation time.
  • False success on nil: Verify terminal status after the provider ends and distinguish .completed, .failed, and .cancelled outcomes.
  • Stale results after cancellation: Tag asynchronous work with a request generation and discard outputs that no longer belong to the visible operation.

Acceptance checks

Test audio-only, video-only, and multiplexed assets; a file with multiple candidate tracks; variable-frame-rate footage; a truncated or unreadable file; a protected or remotely backed asset; a zero-length range; and cancellation during a slow consumer. Record duration, track selection, first and last presentation timestamps, sample count, status, and error code. Assert that the consumer never exceeds its configured in-flight limit and that every successful run ends in a completed reader state.

AVAssetReader gives an application explicit access to media samples, not a promise that every source can be decoded into every requested format. Treat selection, timing, memory, cancellation, and terminal status as a single pipeline design. That discipline makes sample extraction reproducible and keeps downstream editing or analysis from inheriting hidden assumptions.

Related:

Sources:

Comments