Core Audio Process Taps on macOS: Capture Scope, Consent, and Teardown
Design macOS process-audio capture with Core Audio taps, explicit scope, user consent, aggregate-device ownership, and reliable teardown.
Core Audio process taps provide a system-supported way for a macOS application to capture outgoing audio from a process or a selected group of processes. They are not microphone inputs, ordinary output-device discovery, or a license to record everything a user hears. A tap is an audio object with a declared capture scope and mixdown behavior. To consume it as an input stream, an application normally places it into a Core Audio aggregate device and owns the complete lifetime of both objects.
This distinction matters for privacy and reliability. Recording a single meeting application for a user-requested transcript has a different scope from capturing all system output. Make the capture target visible, explain what is recorded, retain only what the feature needs, and provide a clear stop action. A technically successful tap can still be an unacceptable product design if the user cannot understand or control it.
Establish the supported platform and consent model
Apple’s process-tap sample requires macOS 14.2 or later. A release plan should therefore define a deployment target that supports the API or provide an explicit fallback that does not silently capture through an unrelated mechanism. Check the SDK availability annotations during compilation and test the exact OS versions in the support matrix; do not infer availability from a machine on which the app happens to build.
The app must include the NSAudioCaptureUsageDescription information-property-list key. Its text should say what audio is captured and why, in terms a user can evaluate. The first time the app starts recording from an aggregate device that contains a tap, macOS presents the system-audio recording permission prompt. Treat denial as a normal result: keep the rest of the application usable, explain how the feature behaves without access, and never retry in a loop or attempt to circumvent the system decision.
Use separate product states for permission not yet requested, allowed, denied, capture configured, recording, and stopped. The permission dialog is not a substitute for an in-product recording indicator. When capture is active, show a persistent indicator and identify the target processes or scope. If the product has export or upload behavior, explain that separately from local capture. Keep the recorded data out of diagnostic logs and crash annotations.
Define the smallest useful capture scope
CATapDescription describes the tap input stream. Apple provides initializers for a global tap that excludes selected processes, a mixdown of selected processes, and other configurations. Prefer a process list that matches the user-selected operation. A global tap has a much broader privacy impact and should be an intentional, separately explained mode rather than the default for a feature that only needs one app.
The description also includes properties such as a name, privacy visibility, mixdown, exclusivity, output device UID, stream index, and mute behavior. The process selectors in CATapDescription are Core Audio AudioObjectID values, not Unix pid_t values; resolve them through the appropriate Core Audio process-object API rather than passing a PID directly. Set only properties needed by the feature and document their effects. A tap that mutes the original process output changes what the user hears; it must not be enabled as an incidental implementation detail. Mono or stereo mixdown changes the sample stream and should be tested against the product’s actual analysis or recording requirements.
Private visibility controls whether a tap is visible outside the creating process; it does not grant permission or make captured data harmless. A private tap is not a security boundary around the resulting audio file. Similarly, excluding a process from a global tap is not equivalent to positively selecting a specific source. Test whether target application restarts or process replacement changes the set of process identifiers your capture setup relies on.
Treat tap creation and the aggregate device as one owned session
The lower-level workflow creates the tap with AudioHardwareCreateProcessTap and receives an AudioObjectID. The tap’s UID can then be used when configuring the aggregate device’s tap list. The aggregate device is a separate Core Audio object with its own creation and destruction calls. Keep the tap ID, aggregate-device ID, and any I/O procedure or stream resources in one session object, and make teardown idempotent.
There are failure points at every stage: the target process can exit, the device UID can disappear, an aggregate device may fail to instantiate, a stream may be unavailable, a device can be unplugged, or the user can revoke authorization. Return these as explicit session errors. If aggregate-device configuration fails after tap creation, destroy the tap before returning. If the app is shutting down, stop I/O first, remove listeners and callbacks, then destroy the aggregate device and tap in the order required by resources that reference them.
The following minimal Swift example creates a private stereo tap for explicitly selected process object IDs and destroys it on scope exit. It demonstrates tap-object ownership only; it does not create an aggregate device or read audio samples. Production capture must configure and run the aggregate-device input path, and must also implement permission, error reporting, and user-visible recording state.
import CoreAudio
import Foundation
@available(macOS 14.2, *)
func withProcessTap<T>(
processObjectIDs: [AudioObjectID],
body: (AudioObjectID) throws -> T
) throws -> T {
guard !processObjectIDs.isEmpty else {
throw NSError(
domain: "ProcessTap",
code: 1,
userInfo: [NSLocalizedDescriptionKey: "Select at least one process."]
)
}
// These are Core Audio process-object IDs, not Unix pid_t process IDs.
let description = CATapDescription(stereoMixdownOfProcesses: processObjectIDs)
description.name = "User-requested application capture"
description.isPrivate = true
var tapID = AudioObjectID(kAudioObjectUnknown)
let status = AudioHardwareCreateProcessTap(description, &tapID)
guard status == noErr else {
throw NSError(domain: NSOSStatusErrorDomain, code: Int(status))
}
defer {
AudioHardwareDestroyProcessTap(tapID)
}
return try body(tapID)
}
The closure makes cleanup happen on both normal return and a thrown Swift error. It is not a complete recorder: the caller must not retain the tap ID after the closure returns, because the tap has already been destroyed. A real session object should instead retain resources for exactly as long as capture is active and provide an explicit asynchronous stop operation.
Configure and remove aggregate-device membership carefully
An aggregate device combines the tap with the audio input stream a capture client consumes. Apple documents obtaining the tap UID through the tap object’s property selector and adding the UID to the aggregate device’s tap list. Property reads and writes are Core Audio calls with buffer sizes, selectors, and status codes; check every return value. Do not assume a property list contains only the entries your app created, and do not overwrite another component’s tap membership.
Generate a stable, collision-resistant aggregate-device UID for a session and give the device a name that makes its purpose understandable in system diagnostics. Keep the aggregate device private if no other process should use it. When changing the tap list, preserve existing entries unless the app owns the entire device configuration. After removal, destroy the aggregate device and tap only after the I/O path has stopped using them.
Use device and stream identifiers rather than display names as identity. Names can be localized, duplicated, or changed by the user. Device removal and route changes are normal runtime events, not proof that the audio framework is broken. Re-enumerate supported objects and either reconfigure under a documented policy or stop with a clear message. Avoid retry storms that repeatedly create new taps and aggregate devices.
Design the audio callback for bounded real-time work
Capture callbacks run under strict latency constraints. Do not block on disk access, network I/O, locks with unbounded wait time, UI work, or synchronous logging. Move sample buffers into a bounded queue and let a worker encode or write them. If the queue fills, define an explicit policy such as dropping newest frames with a counter, pausing capture, or stopping and notifying the user. Unbounded queues convert a brief disk stall into memory growth.
Record timestamps, sample rate, channel count, frame count, and discontinuity markers alongside the audio data when they are required to interpret it. Validate the actual stream format instead of assuming a fixed sample rate or channel layout. A device switch or aggregate-device rebuild can change the format. Reject unexpected buffer layouts safely and update the recording metadata when a supported format transition occurs.
Do not perform process enumeration or Core Audio property queries in the render callback. Resolve target processes and construct the tap before starting I/O. Use stable session state and atomics or a bounded single-producer/single-consumer queue where appropriate. Document ownership of every buffer and callback context, because a callback that outlives a destroyed session can access freed memory.
Diagnose permission and signal-path failures separately
A useful diagnostic records the failure stage and OSStatus code, not the captured signal. Separate permission denial, tap creation failure, aggregate-device failure, stream initialization, callback starvation, and output-file finalization. Include OS version, hardware family, selected device UID, negotiated format, and whether the target process was still present, while minimizing identifying information.
Test with no selected process, a process that exits during setup, a process that restarts, an output-device change, a sleep and wake cycle, denied permission, permission revocation in System Settings, a slow storage destination, and an app termination while capture is active. Confirm each case ends in a stopped state, destroys resources, and leaves no orphan aggregate device. Test actual audio continuity and channel mapping with a known signal; compiling the API calls is not evidence that the stream is correct.
When a capture session stops, flush only within a bounded shutdown budget. Finalize the file format, close descriptors, release buffers, and update user-visible state even if finalization reports an error. A partially written file should be marked incomplete rather than presented as a successful recording.
Keep retention and access proportional to the feature
Audio can contain credentials, private conversations, health information, and information about people who did not initiate the capture. Store recordings in an appropriately protected location, set a documented retention period, and delete temporary data on cancellation or failed sessions. If the product transcribes audio, preserve whether the transcription is local or remote and provide controls for deletion and export.
Limit process selection to the user’s intended task. Do not silently fall back to a global tap when a selected-process tap fails. Do not infer consent from the fact that the operating system displayed a permission dialog once. Revisit product messaging when capture scope or data flows change, and ensure testing scripts do not leave real user audio in build artifacts.
Release checklist
Before shipping, confirm the 14.2-or-later availability strategy, usage-description text, first-run permission behavior, process-selection UI, aggregate-device cleanup, format negotiation, backpressure policy, and a visible recording indicator. Exercise denial and revocation on a clean test account, and test both Intel and Apple silicon systems if both are supported. Verify the final artifact’s privacy declarations and do not claim that a private tap is equivalent to a sandbox boundary.
Core Audio taps are a powerful, narrow capture primitive. Production reliability comes from combining correct object lifetimes with honest consent, bounded callback work, and a clear fail-closed policy. If any of those properties is missing, successful audio capture is not a successful product feature.
Related:
- Core Audio Device Discovery: Enumerating and Tracking macOS Audio Hardware
- ScreenCaptureKit on macOS: Stream Ownership, Frame Delivery, and Recovery
Sources: