Metal on macOS: Command Buffers, Resource Hazards, and GPU Completion
Build reliable macOS Metal work submission with command-buffer lifecycles, hazard scopes, storage synchronization, bounded frames, and GPU diagnostics.
Metal code is asynchronous by design. The CPU creates a command queue, encodes work into command buffers, and commits those buffers; the GPU schedules and executes the encoded commands later. A successful commit() means the work was submitted, not that a frame was drawn or a result is already visible to the CPU. Reliable rendering and compute therefore require an explicit lifecycle, correctly scoped resource synchronization, and a way to observe completion without blocking the interface.
This article focuses on the long-established MTLCommandQueue and MTLCommandBuffer model used by macOS applications. Metal capabilities and storage behavior vary by device and queue API, so query the actual device and follow the documentation for the specific queue type you create. Do not infer support from a Mac model name or assume every GPU uses the same memory topology.
Choose a device and retain a queue
MTLCreateSystemDefaultDevice() can return no device in an environment where Metal is unavailable. Creating an MTLCommandQueue can also fail. Treat both as ordinary initialization failures and select a fallback, disable the GPU-backed feature, or present an actionable error. Create queues during renderer or compute-service setup and retain them; do not allocate a new queue for every draw call.
A command queue belongs to one MTLDevice and maintains an ordered submission list. Enqueued command buffers are scheduled in enqueue order. If a buffer is not explicitly enqueued, committing it enqueues it implicitly. This ordering is valuable for sequential passes that depend on previous output, but it does not make CPU mutations to a resource safe while the GPU is using that resource. Queue ordering and memory visibility are related but separate contracts.
A command buffer is a one-shot submission
The usual lifecycle is: create a command buffer, create an encoder, bind the pipeline and resources, encode commands, end the encoder, register any handlers, and commit. A command buffer cannot be committed twice and cannot accept more encoded commands after commit. Register scheduled and completion handlers before committing. A completion handler reports that the GPU finished executing the command buffer; a scheduled handler reports that the buffer has been scheduled, which is an earlier milestone.
import Dispatch
import Metal
enum ComputeFailure: Error {
case commandBufferUnavailable
case encoderUnavailable
case emptyWorkload
}
func submit(
queue: MTLCommandQueue,
pipeline: MTLComputePipelineState,
input: MTLBuffer,
elementCount: Int,
completion: @escaping (Result<Void, Error>) -> Void
) throws {
guard elementCount > 0 else { throw ComputeFailure.emptyWorkload }
guard let commandBuffer = queue.makeCommandBuffer() else {
throw ComputeFailure.commandBufferUnavailable
}
guard let encoder = commandBuffer.makeComputeCommandEncoder() else {
throw ComputeFailure.encoderUnavailable
}
encoder.setComputePipelineState(pipeline)
encoder.setBuffer(input, offset: 0, index: 0)
let width = pipeline.threadExecutionWidth
let groups = (elementCount + width - 1) / width
encoder.dispatchThreadgroups(
MTLSize(width: groups, height: 1, depth: 1),
threadsPerThreadgroup: MTLSize(width: width, height: 1, depth: 1)
)
encoder.endEncoding()
commandBuffer.addCompletedHandler { finishedBuffer in
if let error = finishedBuffer.error {
DispatchQueue.main.async { completion(.failure(error)) }
} else {
DispatchQueue.main.async { completion(.success(())) }
}
}
commandBuffer.commit()
}
The example assumes a compute kernel that reads the buffer at index zero and bounds-checks its thread index against the logical element count. The pipeline width is chosen for that pipeline; applications should check device limits and ensure the kernel’s dispatch geometry matches its data layout. The completion closure is marshalled to the main queue only because this example models a UI caller. A non-UI consumer should receive completion on an executor that owns its state, not touch AppKit objects from a Metal callback.
Never call waitUntilCompleted() on the main thread to make a result appear synchronous. Waiting blocks CPU progress and can freeze input and drawing while the GPU is busy. Use completion handlers for dependent CPU work, and keep multiple frames or batches in flight only up to a deliberate bound. An unbounded stream of command buffers can increase memory usage and latency even when the GPU eventually catches up.
Encoders describe passes, not arbitrary interleaving
A command encoder records one kind of pass, such as render, compute, or blit work. Finish the encoder with endEncoding() before creating the next encoder or committing the buffer. Keep render and compute responsibilities explicit. A compute dispatch writes outputs; a later render pass consumes those outputs. Putting both passes into a single command buffer can establish useful ordering, but resource access rules still determine whether a hazard is tracked automatically or needs an explicit dependency.
A command buffer’s status and error are diagnostic inputs. A completion callback should distinguish a successful completion from an error and record the error domain/code without assuming that every GPU fault has one repair path. Do not reuse buffers or mutable CPU-side state merely because the submission call returned. Design resource reuse around the GPU completion boundary.
Resource hazards depend on scope and tracking mode
Two commands conflict when one writes a resource that another reads or writes before the dependency is satisfied. Metal can automatically track some conflicts for commands submitted to an MTLCommandQueue, but that behavior depends on the resource’s hazard-tracking mode and how the resource is bound. Apple’s synchronization guide distinguishes directly bound tracked resources from untracked resources, including resources created from heaps. Do not rely on automatic tracking for every resource or across independent queues.
Use the smallest synchronization primitive that matches the dependency. A fence coordinates relevant encoder work within its supported scope. An event can coordinate work across queues on a device; a shared event provides broader CPU/GPU coordination. Modern Metal queue APIs have their own barrier model and must be evaluated against their documentation rather than copied from an older command-queue example. A synchronization primitive does not fix an incorrect producer/consumer graph: write down which pass produces each resource, which later pass consumes it, and what signal or ordering edge connects them.
For independent command queues, queue order alone cannot order work on the other queue. Use the appropriate supported event or dependency mechanism. Avoid broad waits that stall unrelated GPU work. A GPU access hazard often appears as flickering, stale output, nondeterministic results, or behavior that changes under load; those symptoms are not proof that adding a CPU sleep will fix the race.
CPU/GPU visibility and storage modes
Storage modes describe how CPU and GPU access resource contents. Shared storage lets both sides access shared backing memory on supported devices. Managed storage, available in relevant macOS device configurations, maintains CPU and GPU copies that need explicit synchronization. When the CPU modifies a managed resource, report the modified range as documented. When GPU work updates one, encode a synchronization operation for the resource and wait for that synchronization command buffer to complete before reading the CPU-visible copy.
The order matters: encode the producer GPU pass, encode the resource synchronization in the appropriate command buffer, commit, and wait asynchronously for completion before the CPU consumes the result. Calling a synchronization method is not itself proof that the GPU has finished. Conversely, issuing a CPU-side update while a GPU pass is reading the same bytes can race even when the command queue itself is correctly ordered.
Do not hard-code a storage mode from “Apple Silicon” or “Intel” alone. Check the device, resource usage, and API requirements. Keep synchronization boundaries narrow, synchronize only the ranges or resources that need CPU visibility, and avoid a CPU readback when the result can remain on the GPU for another pass.
Limit work in flight and retain resource lifetime
Interactive rendering is a pipeline: while the GPU renders one frame, the CPU may encode the next. A small in-flight limit prevents the CPU from getting too far ahead, which bounds transient resources and input-to-display latency. Use a semaphore or another explicit frame-slot mechanism to cap in-flight work, signal it from the completion handler, and make shutdown release outstanding state safely. Keep the slot count tied to the rendering design, not to a guessed universal constant.
Command buffers normally maintain the resources they reference for execution. APIs that create command buffers with unretained references change that lifetime contract; only use them when the application can prove that every resource survives until GPU completion. Per-frame uniform buffers, argument data, and temporary textures should be recycled only after their owning command buffer finishes. A CPU reference going out of scope is not a safe resource-reuse signal when work remains in flight.
Diagnose with GPU-aware evidence
Enable Metal API Validation during development to catch invalid resource usage and encoding errors. Capture a representative frame with Xcode’s GPU tools and inspect encoder boundaries, resource bindings, and dependencies. Use Metal System Trace to distinguish CPU encoding delay, queue backlog, GPU execution time, and synchronization stalls. Put signposts around frame acquisition, encoding, commit, and completion; correlate them with a frame identifier rather than logging every resource’s private contents.
Test under GPU load, with the window resized, after app suspension/resumption where applicable, and on each supported device family. Include low-memory and device-unavailable fallback behavior. For a compute pipeline, compare output with a CPU reference for boundary sizes such as zero, one, a partial threadgroup, and multiple groups. For a renderer, validate frame pacing and input latency, not merely whether a triangle appears once.
An acceptance review should establish that each buffer is committed once, encoders end before submission, handlers are installed before commit, CPU work does not synchronously wait on the UI thread, every resource dependency has a valid scope, managed-resource visibility is explicit where required, and in-flight work is bounded. Metal becomes dependable when submission, execution, and visibility are treated as separate stages with observable completion.
Related:
- AUv3 on macOS: Extension Discovery and Real-Time Audio Rendering
- Apple Virtualization.framework: Build, Validate, and Run macOS and Linux VMs
Sources: