OperationQueue on macOS: Dependencies, Cancellation, and Bounded Work
Design Foundation OperationQueue pipelines with truthful dependencies, cooperative cancellation, explicit concurrency bounds, and safe completion ownership.
OperationQueue schedules Operation objects according to readiness, priority, dependencies, and its concurrency policy. It is useful when work needs a visible dependency graph, cancellation state, progress integration, or observable completion. It is not a transactional workflow engine: dependencies say that a predecessor finished, not that it succeeded, and cancellation is a request for work to stop rather than a forced thread kill.
The most common production defect is treating a queue as a list that can be edited freely after submission. Apple documents that an operation remains in the queue until it finishes and cannot be directly removed after enqueueing. Queues retain operations, and a suspended queue with unfinished operations can retain memory. Model each unit of work with a clear owner, completion rule, cancellation policy, and durable side-effect boundary before adding it to a queue.
Build the dependency graph before enqueueing
Dependencies describe ordering. An operation becomes ready only after its dependencies finish; they do not carry a success result. Add dependencies before submitting the dependent operation so the queue never schedules an incomplete graph. Use operation results or a shared typed state to decide whether the next stage should proceed.
import Foundation
func makePipeline() -> OperationQueue {
let queue = OperationQueue()
queue.name = "com.example.import-pipeline"
queue.maxConcurrentOperationCount = 3
let download = BlockOperation {
// Store a durable, validated input or publish a result object.
}
let parse = BlockOperation {
// Check the recorded download result before parsing.
}
let index = BlockOperation {
// Check the parse result before making the index visible.
}
parse.addDependency(download)
index.addDependency(parse)
queue.addOperations([download, parse, index], waitUntilFinished: false)
return queue
}
The sample shows only scheduling. The comments mark places where a real pipeline needs result state and error propagation. A dependency is satisfied when its operation is finished, including when it was cancelled. If parsing must not run after a failed download, the parser must inspect a result that distinguishes success, cancellation, and error. Do not encode that distinction by hoping queue order or cancellation behavior will prevent it from running.
Dependencies are most useful for a directed acyclic graph. Cycles prevent readiness and can leave the queue with work that can never start. Build the graph from product stages and validate it in tests; do not create a global “everything depends on everything” queue to coordinate unrelated work. Separate independent pipeline instances when they have different resource limits or cancellation ownership.
Cancellation is cooperative
Calling cancel() sets cancellation state and informs the operation. It does not forcibly interrupt arbitrary code that is already running. A long-running operation must check isCancelled at safe points, stop work, release resources, and return. For a block operation, break large tasks into bounded chunks or move the work into a custom operation that can check cancellation between chunks.
Cancellation also affects dependencies. Apple documents that a cancelled operation may ignore unfinished dependencies so the queue can start it and move it to a finished state sooner. A dependent operation can therefore become ready even though a predecessor’s work did not complete successfully. Model the result, not just the dependency relation. If cancellation should stop an entire import, propagate a shared cancellation token or cancel the relevant operations and ensure each stage refuses to publish partial state.
Do not cancel by abandoning a reference to the queue. The queue may retain operations while they are executing. Keep a coordinator that can cancel the active request, observe completion, and retire its state. A generation identifier is useful when users start a newer request while an older one is winding down; late completion from the old generation must not replace current UI or persistent state.
Concurrency limits and resource budgets
maxConcurrentOperationCount applies to one queue only. Other operation queues can run their own operations simultaneously, so setting a queue to one does not globally serialize the process. Lowering the property does not stop work that is already executing. Apple recommends leaving the default concurrency value when possible; the system can choose a count based on current conditions. Set a small explicit cap when each operation consumes a scarce resource such as file descriptors, memory, remote API quota, or a non-thread-safe device.
A concurrency count is not a memory limit. Three image decode operations may still allocate gigabytes if each accepts an unbounded source. Bound both parallelism and per-operation input/output. Use backpressure at enqueue boundaries so a producer cannot create an arbitrarily large backlog of pending operations. For network or disk tasks, distinguish active operations from queued intent and consider coalescing obsolete work by key.
Quality of service expresses the importance of the work, not a guaranteed priority or deadline. Avoid labeling background maintenance as user-interactive to make it seem faster; this can compete with foreground rendering. Pick a QoS that matches whether the work is directly required to respond to a user action, and measure observed latency on representative machines rather than interpreting the setting as a real-time guarantee.
Avoid queue-level deadlocks and starvation
Do not synchronously wait for an operation from work that occupies the only slot needed to execute it. For example, a parent operation running on a single-concurrency queue that waits for a child submitted to the same queue can deadlock. Prefer dependencies or asynchronous continuation. Similarly, avoid waitUntilAllOperationsAreFinished() on the main thread for work that may need UI progress or user input; it blocks the caller and can freeze the interface.
Suspension prevents the queue from scheduling new operations but does not pause a task that is already executing. It is not a safe way to freeze arbitrary user work. If the app suspends a queue while operations retain large inputs, the queue may keep those objects alive. Use a short, explicit suspension window for controlled intake, and always have a path that resumes the queue even when setup fails.
Operation completion blocks and KVO callbacks can arrive away from the main thread. Marshal UI changes to the main actor or main queue. Keep the operation’s result immutable after publication, and do not make completion handlers retain an owner indefinitely. KVO is useful for observation but is not a reason to bind queue internals directly to UI controls.
Custom asynchronous operations require a real state machine
A synchronous main() implementation is the simplest operation type. A custom asynchronous operation is more demanding: it must report executing and finished state correctly, make those properties thread-safe, and send KVO changes when they transition. Finishing is essential because the queue cannot retire the operation or satisfy dependents until isFinished becomes true. Cancellation before the underlying callback arrives must still transition the operation to finished exactly once.
Prefer a modern structured-concurrency task when the work is naturally async/await and no operation-graph integration is required. If an operation must wrap asynchronous work, protect the completion state against races among cancel, network completion, timeout, and owner teardown. Implement one idempotent finish(result:) transition that releases resources and changes observable state once. Test every order in which those events can arrive.
Side effects and retries
An operation queue does not make file writes, database transactions, or remote mutations atomic. Write to a temporary destination, validate, then publish with an appropriate atomic replacement strategy. For server changes, include an idempotency key or reconcile an ambiguous timeout before retry. If a stage can be retried, its outputs must either be idempotent or isolated until validation completes.
Attach a stable operation identifier and stage name to logs. Record queued, started, cancelled, failed, and completed states with durations and bounded error categories. Avoid logging user documents or credential-bearing request details. Use operation name and queue name for diagnostics, but do not expose internal object descriptions as a support contract.
Acceptance tests
Test a successful dependency chain, a cancelled predecessor, a failed predecessor, cancellation before start, cancellation during a long task, a slow task that completes after a newer generation, a queue concurrency cap, a second queue executing concurrently, and queue suspension/resume. Assert that downstream work inspects result status, partial outputs are not published, every custom asynchronous operation reaches finished, and UI state updates on the correct executor.
Measure backlog length, operation wait time, execution time, memory per active stage, cancellation latency, and final error rate. A queue that reports zero active operations may still have pending work; distinguish queue depth from work in flight. Add overload behavior before production rather than discovering that the queue can accumulate more work than the machine can finish.
OperationQueue gives an app a scheduler and observable operation lifecycle. The application still owns dependency success semantics, cancellation cooperation, memory budgets, side-effect safety, and user-facing completion. Use the queue to make those rules explicit, not to hide them inside callback ordering.
Related:
- Grand Central Dispatch on macOS: Queues, QoS, Barriers, and Deadlocks
- NSProgress on macOS: Composable Progress, Cancellation, and Task Ownership
Sources: