Speech Recognition on macOS: SFSpeechRecognizer Requests and Results
Build resilient macOS transcription with explicit authorization, locale and availability checks, partial-result handling, cancellation, and privacy boundaries.
Apple’s Speech framework provides SFSpeechRecognizer for recognizing audio from a file or an audio-buffer stream. Recognition involves asynchronous service availability, language selection, request configuration, task completion, and a privacy boundary. A valid recognizer object does not mean the selected language is available at that moment, and a partial transcript is not a final result.
This article focuses on the SFSpeechRecognizer request/task API. Treat the captured or imported audio as sensitive user data. Apple’s current documentation says this recognition path can send audio to Apple servers, requires authorization, and requires an NSSpeechRecognitionUsageDescription entry in the app’s Info.plist. Do not enable it invisibly as a side effect of opening a document.
Request permission when the feature is invoked
Ask for authorization when the user enters the transcription workflow, not at unrelated app startup. Explain the purpose in the usage description, handle authorized, denied, restricted, and not-determined states, and provide a useful app path when permission is unavailable. The Speech usage key is required for this API; Apple’s documentation warns that a missing key causes the app to crash when it requests authorization or uses Speech APIs.
import Speech
func requestSpeechAuthorization(_ completion: @escaping (Bool) -> Void) {
SFSpeechRecognizer.requestAuthorization { status in
let allowed = status == .authorized
DispatchQueue.main.async {
completion(allowed)
}
}
}
Authorization is a system permission, not a guarantee about network or language availability. Re-check the recognizer and handle service unavailability even after the user grants access. Do not repeatedly prompt after a denial; direct the user to the system settings path when appropriate and preserve a non-speech workflow.
For live microphone capture, microphone permission and audio capture lifecycle are separate from speech-recognition authorization. The recognition API’s permission does not itself configure an AVAudioEngine input tap or authorize audio capture. Explain both data flows and request each permission only when the feature needs it. This example processes an existing file and therefore does not implement microphone capture.
Select a locale and check service availability
Each SFSpeechRecognizer supports a language locale. Construct it with the intended locale, and show the chosen language to the user if the product supports multiple languages. Falling back silently to a different locale can create convincing but incorrect text.
Check isAvailable immediately before starting work. Availability can change due to service state or network conditions, and some languages may need an internet connection. supportsOnDeviceRecognition indicates whether the recognizer supports on-device operation for that locale; it does not by itself configure a request to stay on device.
When the product requires local processing, set requiresOnDeviceRecognition on the request and handle an unsupported or unavailable result as a normal product state. Do not silently fall back to server processing when local-only is a privacy requirement. If server processing is acceptable, explain that behavior and use the required authorization flow.
Choose a request type for the source
Use SFSpeechURLRecognitionRequest for an existing audio file. Use SFSpeechAudioBufferRecognitionRequest when the app supplies live or pre-collected audio buffers. The abstract base request is not instantiated directly. Configure options such as partial results, contextual strings, task hint, punctuation, and on-device requirements before starting the task.
import Foundation
import Speech
func transcribeFile(
at url: URL,
recognizer: SFSpeechRecognizer,
completion: @escaping (String?, Error?) -> Void
) -> SFSpeechRecognitionTask {
let request = SFSpeechURLRecognitionRequest(url: url)
request.shouldReportPartialResults = false
request.taskHint = .dictation
return recognizer.recognitionTask(with: request) { result, error in
if let error {
completion(nil, error)
return
}
guard let result, result.isFinal else { return }
completion(result.bestTranscription.formattedString, nil)
}
}
The caller must check permission, recognizer creation, locale, and isAvailable before calling this helper. Retain the returned SFSpeechRecognitionTask for the owning workflow so the app can cancel it. In production, marshal completion into the actor or queue that owns the UI and guard against the result belonging to a closed document or superseded request.
For buffer streaming, append audio in the expected format and call endAudio() when no more samples will arrive. Without that terminal signal, the recognizer can continue waiting for additional input. A live pipeline must also stop or detach its audio capture source, release its tap, and cancel the recognition task together when the user ends the session.
Partial and final results
Partial results can improve perceived responsiveness, but they are provisional. A later callback may revise the transcript. Keep partial text in transient UI state and only commit it as a final result after the task reports completion. If the user edits or accepts an intermediate transcript, store provenance so the app does not later overwrite the user’s changes with a delayed final callback.
Each task can produce multiple callbacks. Treat the error and final-result paths as terminal state transitions, cancel or release resources once, and ignore duplicate completion attempts. A result’s isFinal field distinguishes a finalized transcription from a partial hypothesis. Do not report success just because a non-final string is non-empty.
Contextual strings can help recognition with specialized terms, but they do not guarantee exact spelling and should not be used to smuggle sensitive data into a service request. Set a task hint that matches the actual use, such as dictation or search, rather than assuming all speech is generic dictation. Measure accuracy on representative accents, microphones, noise levels, and domain vocabulary.
Audio quality and file ownership
Validate that the input URL points to an audio file the app can read and that the file’s format is supported. Keep file access valid for the full task lifetime and coordinate with other writers if the audio can change while recognition is running. A file that is replaced after the request begins can produce inconsistent results or an access failure.
For live capture, consider input device changes, sample-rate negotiation, interruptions, route changes, buffer duration, and task teardown. Speech recognition is downstream from the audio engine; a healthy recognition task cannot repair a missing input tap or silent samples. Inspect audio level and format separately from transcript status when diagnosing no-result behavior.
Do not place recognition on the main thread or synchronously wait for it. The callback is asynchronous; update UI through the owning actor, throttle partial-result rendering, and avoid relayout for every word when a coarse refresh is sufficient. Cancel obsolete tasks promptly when the user starts a replacement request.
Transcript editing and confidence
Recognition output is a hypothesis generated from audio, not verified ground truth. If the transcript triggers commands, changes a record, or enters a regulated workflow, require a confirmation or use a downstream validation step. Confidence values and alternatives can help prioritize review, but they do not prove that a proper name, number, or negation was recognized correctly.
Store the transcript separately from the source audio and track whether it is partial, final, user-edited, or machine-generated. Once the user edits a final transcript, a late callback must not replace that text. Attach a task generation and a document revision to callbacks, and compare both before applying results. When the underlying audio is replaced, invalidate the old task and any cached transcript associated with it.
For long recordings, split work only when the audio workflow and API contract support a clean boundary. Arbitrary byte slicing can cut through encoded frames or utterances. Preserve timestamps and segment order if the product needs searchable alignment, and record the segment range associated with each transcript fragment. A plain concatenated string loses timing provenance.
Error taxonomy and user recovery
Distinguish authorization denial, recognizer construction failure, current unavailability, invalid audio, request failure, cancellation, and a successful task that produced no useful speech. Each outcome has a different recovery: request permission only after user action, choose another supported locale, retry after service availability returns, ask the user to select a valid file, or let them enter text manually.
Retry only transient failures and avoid an immediate tight loop. If on-device processing is required and unavailable, explain the limitation rather than silently changing the data-processing boundary. Record the task’s terminal state and clear any spinner or capture indicator in every completion path, including cancellation and app teardown.
Privacy, retention, and observability
Tell users whether audio or transcript data is stored, where it is processed, and how they can delete it. Apple’s documented SFSpeechRecognizer server path processes audio as sensitive data and requires permission. For a distinct API such as SpeechAnalyzer, consult its own current processing and authorization documentation rather than assuming it has identical behavior or limitations.
Avoid logging raw audio, complete transcripts, contextual phrases, or user names in ordinary diagnostics. Prefer request identifier, locale, input duration, whether on-device recognition was required, availability result, task duration, and terminal status. If transcript telemetry is needed for quality review, make it explicit, consented, minimized, and governed by the product’s retention policy.
Validation matrix
Test authorization states, missing usage description in a development fixture, unsupported locale, temporarily unavailable service, local-only requirement on an unsupported recognizer, empty file, corrupt file, partial then final results, task cancellation, file replacement, app backgrounding, and a second task superseding the first. For live audio, test microphone denial and input-device removal separately from speech authorization.
Verify the final transcript against known audio fixtures and track error classes instead of treating every failure as “no speech.” Confirm the UI never commits an obsolete partial result and that all tasks and capture resources are stopped after cancellation. Test any usage-description change in the signed app bundle users actually install.
SFSpeechRecognizer supplies recognition requests and results, while the application owns permission timing, audio lifetime, locale policy, transcript finality, and data retention. Keep each boundary explicit so recognition can fail gracefully without surprising the user or corrupting document state.
Related:
- AVSpeechSynthesizer on macOS: Utterance Queues, Voice Selection, and Cancellation
- AVCaptureSession on macOS: Device Setup, Frame Delivery, and Interruption Recovery
Sources: