Skip to content
macOSDeep Dive Published Updated 7 min readViews unavailable

AVSpeechSynthesizer on macOS: Utterance Queues, Voice Selection, and Cancellation

Integrate AVSpeechSynthesizer with retained ownership, language-aware voices, UTF-16 speech ranges, queue policy, interruption handling, and cleanup.

AVSpeechSynthesizer turns text utterances into audible speech and reports lifecycle events through properties and an optional delegate. The synthesizer keeps its own utterance queue, but the application must retain the synthesizer object until speech concludes. A read-aloud feature should model the queue, current content revision, selected voice, pause/stop behavior, and user-visible progress independently from the UI view that started it.

Speech synthesis is useful for read-aloud, guided workflows, and spoken feedback. It is not a replacement for macOS VoiceOver or semantic accessibility. A custom spoken presentation should respect assistive-technology settings and avoid competing with system speech unnecessarily.

Retain the synthesizer and configure utterances before enqueueing

An AVSpeechUtterance represents one unit of speech and carries properties such as voice, rate, pitch, volume, and pre/post delays. Set these parameters before calling speak(_:) and treat queued utterances as immutable; if the user changes voice or rate, apply the new settings to newly created utterances instead of relying on mutation of work already in the synthesizer’s queue. Split text into meaningful units if the product needs per-section progress, different pronunciation, or pauses between paragraphs.

import AVFAudio

final class SpeechQueue {
    private let synthesizer = AVSpeechSynthesizer()

    func speak(_ text: String, language: String = "en-US") {
        let utterance = AVSpeechUtterance(string: text)
        utterance.voice = AVSpeechSynthesisVoice(language: language)
        utterance.rate = AVSpeechUtteranceDefaultSpeechRate
        synthesizer.speak(utterance)
    }

    func stopImmediately() {
        synthesizer.stopSpeaking(at: .immediate)
    }
}

This small owner keeps the synthesizer alive for as long as the SpeechQueue is retained. A production controller should hold that owner while any utterance is queued or active. The synthesizer is not automatically retained by the system, so making it a temporary local variable can stop speech unexpectedly when the function returns.

Voice choice and language

Select a voice by language or by a product preference that the system currently exposes. Voice availability and quality can vary by OS installation and installed voice resources. Ask the framework for available speech voices rather than persisting an assumption that one identifier exists on every Mac. If a preferred voice is unavailable, fall back to a voice matching the utterance’s language and tell the user which preference was selected.

Do not assign one language to a document that contains multiple scripts unless the product explicitly supports that behavior. Segment content by language when necessary, but avoid splitting inside a word or grapheme cluster. Keep language metadata as part of the document or selection model instead of guessing from one short sentence.

Rate values are bounded by the framework’s speech-rate constants. A product slider should map user preference into that documented range and be tested for intelligibility, not merely maximum speed. Pitch and volume changes can also affect comprehension. Provide a preview and a reset-to-default option when voice settings are user-configurable.

For names or domain-specific vocabulary, use the supported pronunciation mechanisms deliberately. An utterance can use attributed text and the framework’s speech attributes; SSML is another input form when the product owns a carefully validated markup string. Do not construct markup by interpolating raw document text into tags. Keep pronunciation dictionaries tied to a language and test them against actual voices because a pronunciation that works for one locale may be wrong for another.

Queue policy and content revisions

When the synthesizer is already speaking, new utterances are queued in order. That is convenient for a document reader, but it can also create a stale speech queue if the user selects new content. Decide whether a new request appends, replaces, pauses, or cancels existing speech. For “read selection,” cancel the old queue and start the new selection; for a playlist of chapters, appending may be expected.

Associate each utterance with a content ID and revision in the app’s own queue model. Delegate callbacks report the utterance but not necessarily every product concept the UI needs. If a document changes while speech is queued, either continue reading the captured immutable text or stop and restart from the new revision. Do not let an old completion move the highlight or reading cursor in a newer document.

Treat queue replacement as a transaction in the coordinator: record the requested revision, cancel or drain the old queue according to policy, then enqueue the new utterances and publish that revision to the UI. Delegate events should be matched to the utterance identity before they advance progress. This avoids a brief stale callback making the interface claim that a canceled paragraph is still being read.

Use pauseSpeaking(at:) for a pause that should complete at a word or sentence boundary and stopSpeaking(at:) for cancellation at the chosen boundary. An immediate stop is useful for an explicit stop control but can cut off the current word. After cancellation, clear pending app-side queue entries and update the UI only when the synthesizer reports the corresponding terminal callback or state transition.

Delegate callbacks and text highlighting

The delegate can report that an utterance started, paused, continued, finished, or was cancelled. It can also report a range that is about to be spoken. Use these events to update a read-along highlight, but remember that NSRange offsets are UTF-16 units. Swift String.count is not interchangeable with those offsets. Convert ranges at a well-defined boundary and test emoji, combining marks, and non-Latin scripts.

Delegate callbacks can outlive a view that initiated speech. The speech coordinator should own them and publish value updates to the active UI. Do not capture a view strongly in the synthesizer delegate or callback path. If the reading window closes but speech should continue, keep the model owner alive at an app-level scope; otherwise cancel when the window’s task ends.

Audio routing and interruptions

AVSpeechSynthesizer controls the route used for synthesized audio. On macOS, the app should not assume that a particular output device remains selected if the user changes audio devices. Coordinate spoken output with other app audio and expose a clear mute/stop control. If the product has its own AVAudioEngine or playback system, test how speech and that audio are expected to coexist.

Do not assume every OS interruption maps to one delegate callback. Monitor the synthesizer’s speaking and paused state and observe system audio conditions only where the product needs them. Keep the app responsive if output is unavailable, and avoid busy-looping attempts to restart speech after a device or system change.

User trust and privacy

Speech output can be audible to people nearby. Do not read sensitive content automatically when a document opens. Require a user action or a clear preference, and pause or stop on the product’s defined lock or account-switch event. Keep spoken text out of diagnostic logs. If speech synthesis uses a text string that came from OCR or another model, make it clear when the content may be inaccurate.

Keep the text-to-speech path separate from analytics and do not upload full document text merely to report a read-aloud interaction. If the app sends text to any remote service for another feature, make that transfer explicit and do not infer consent from the fact that a user pressed Play.

Acceptance tests

Test empty text, long text split into utterances, unavailable preferred voice, language switch, pause/resume, stop immediate, stop at word or sentence boundary, enqueue while speaking, replace current selection, document edit during speech, view closure, output-device change, and a delegate callback after the original controller disappears. Verify that speech ends when the owner releases its task and that app-side queue state matches synthesizer state.

Measure time to first spoken output, callback ordering, cancellation latency, and memory for long documents. Confirm highlight ranges map correctly to visible text in multiple scripts. Include a manual listening review for voice quality and intelligibility; API type checks cannot establish that spoken content sounds correct.

A robust speech feature retains the synthesizer, establishes a deliberate utterance queue, selects voices with fallbacks, and reconciles delegate callbacks against the current content revision. AVFAudio generates speech; the app defines what is spoken, when it stops, and how the user remains in control.

Related:

Sources:

Comments