The Libretro Audio Callback Contract: Frames, Buffering, and Timing
A practical engineering guide to libretro audio callbacks, interleaved frames, partial batch writes, emulator clocks, threading, and measurable audio tests.
An emulator can produce a correct waveform and still sound broken in a frontend. A callback may receive the wrong number of samples, stereo channels may be swapped, audio can drift against video, or an asynchronous producer can call a function that was assumed to be single-threaded. The libretro audio interface is small, but its details are a contract between independently built code.
The first useful distinction is between the emulated sound clock and the host audio device. A core advances its sound hardware according to emulated time. It then submits a sequence of samples to the frontend, which resamples, mixes, queues, and eventually gives them to an operating-system audio backend. The callback is a delivery boundary, not a guarantee that a sample will become audible immediately.
A frame means a left/right pair
The standard signed 16-bit API has two alternatives: retro_audio_sample_t, which submits one stereo frame at a time, and retro_audio_sample_batch_t, which submits several frames in one call. A frame is a pair of samples, left followed by right. It is not one scalar sample and it is not necessarily one video frame. For a batch of N frames, the interleaved buffer contains 2 * N int16_t values:
static retro_audio_sample_batch_t audio_batch;
void retro_set_audio_sample_batch(retro_audio_sample_batch_t callback)
{
audio_batch = callback;
}
static size_t submit_stereo_frames(const int16_t *samples, size_t frames)
{
if (audio_batch == NULL || samples == NULL)
return 0;
/* samples[0] is L0, samples[1] is R0, then L1, R1, and so on. */
return audio_batch(samples, frames);
}
The callback’s return value is the number of frames processed. A core should not silently reinterpret it as a count of scalar samples. If the callback accepts fewer frames than offered, the core must have an intentional policy for the remainder, such as retaining it in a bounded queue; ignoring the return value can create discontinuities or unbounded buffering. The canonical header says only one of the ordinary audio callbacks should be used. Choose the per-sample form when the emulator naturally produces one frame at a time; choose batch form when it already has a block. The documented core guide discourages batches smaller than 32 frames because call overhead can erase the benefit.
Interleaving is part of the API, not a frontend preference. A non-interleaved layout such as left[N] followed by right[N] is not a valid standard stereo batch. Validate the first few callback buffers with a known channel-identification signal before debugging a mixer. Also keep the value representation clear: ordinary audio samples are signed 16-bit native-endian integers. A separately negotiated float extension exists in newer headers as an experimental environment command; older frontends can reject it, so it must not replace the compatible int16_t fallback without an explicit successful negotiation.
Convert emulated time into sample counts without accumulating drift
The nominal average audio frames needed during one video interval is sample_rate / frames_per_second. At 44,100 Hz and 60 Hz that quotient is 735 exactly. At 44,100 Hz and 59.94 Hz it is not an integer, so rounding every frame to 736 emits audio too quickly, while rounding to 735 emits it too slowly. The correct long-run total comes from preserving the fractional remainder.
A core can keep a fixed-point phase accumulator: add the audio rate for each video frame, divide by the video rate to obtain the number of samples due, and retain the remainder. For a variable-refresh emulated system, derive output from emulated master-clock progress rather than a host wall-clock timer. For systems with an established non-integer refresh rate, represent the rate as a rational value instead of repeatedly converting a rounded decimal. The exact synthesis path differs by emulator, but the invariant is measurable: cumulative submitted frames should track elapsed emulated time within the expected rounding error, not walk steadily ahead or behind.
Do not assume that one retro_run() call corresponds to a fixed integer number of audio samples. The API describes retro_run() as one video frame of core work and recommends submitting the matching audio for that interval, but the sample count can vary by one as the fractional phase is carried forward. A frontend can fast-forward, slow down, or frame-step. Those modes are another reason to avoid sleeping against the host clock inside an emulator core; the frontend controls playback scheduling.
Silence is data too. If a game produces no sound during an interval, submit the appropriate silent output or follow the core’s established audio policy. Skipping callbacks accidentally can cause a frontend queue to underrun. Conversely, duplicating a previous block to conceal a missed deadline changes the waveform and can conceal the actual timing bug. Make underruns observable in a test build rather than treating them as a normal audio strategy.
Keep callback ownership and threading explicit
In the ordinary model, a frontend installs callback function pointers through the core’s retro_set_* entry points. The core stores those pointers and calls the audio callback while processing a frame. The standard API does not promise general thread safety; core code should not move ordinary libretro calls to worker threads unless a specific contract permits it.
There is an optional asynchronous-audio mechanism for software whose audio is genuinely produced asynchronously. It is not a faster replacement for the batch callback. The frontend notification tells the core that audio can be accepted; the core still supplies samples through the normal audio functions. The header warns that the notification may run on any thread, requires thread-safe handling, recommends the frame-time callback alongside it, and says the core must remain capable of rendering through the normal interface if the frontend disables the feature. This is a poor fit for a conventional emulator whose audio clock is advanced deterministically during retro_run().
If a core has a mixer thread for internal reasons, establish a bounded producer-consumer queue with clear ownership. Do not let one thread mutate a buffer while another invokes the frontend callback on it. Define what happens at unload, state restore, fast-forward, and audio-device reconfiguration. Stop or join workers before releasing their queue memory. Save states should capture emulated sound-chip and resampler state that affects future output, not the host frontend’s transient playback queue.
Diagnose audible failures at the API boundary
Useful telemetry separates emulated production from frontend acceptance. Record frames generated per retro_run(), frames offered, frames accepted, queue depth, underrun/overrun counts, and the emulated time represented by each block. Do not log every sample in a production loop; counters and periodic summaries expose trends without distorting scheduling.
Symptoms have different signatures:
- Pitch is wrong but tempo is steady: inspect the core’s reported sample rate, resampler ratio, chip clock, and any accidental conversion from Hz to kHz.
- Stereo is reversed or collapsed: inspect interleaving and channel order at the callback boundary before changing the sound-chip model.
- A click occurs at frame boundaries: check whether the callback is receiving short batches, whether a fractional sample remainder is discarded, and whether each channel’s phase is continuous.
- Audio gradually drifts against video: compare cumulative sample frames with emulated master-clock time; search for per-frame rounding and host-timer sleeps.
- Audio breaks only in fast-forward or frame-step: inspect assumptions that one callback always arrives at real-time cadence and confirm the core still advances from frontend-driven frame calls.
- A frontend-specific crash occurs: verify that the chosen optional callback was negotiated, that ordinary callbacks are not called from an unapproved worker thread, and that callback-owned storage remains alive for the call.
Acceptance tests that make the contract reproducible
Use a deterministic test core or a test ROM that emits separate left/right tones, a channel impulse, silence, and a known sample-count sequence. Load it on at least one frontend with a debug audio backend and one ordinary user-facing frontend. Capture the callback input before frontend mixing where possible, and include the frontend version, core commit, reported AV information, audio driver, output rate, and platform in the result.
An acceptance run should establish that the batch buffer has exactly two values per frame, channel order is stable, callback return counts are honored, no steady queue growth occurs, and cumulative audio duration matches the expected emulated duration. Repeat under normal speed, fast-forward, pause/resume, frame-step, and save-state restore. A test does not prove that every host audio driver has identical latency; it proves the core respects the API at its boundary. That distinction keeps frontend buffering problems from being misdiagnosed as an emulated sound-chip defect.
The durable rule is simple: generate audio from emulated time, transport complete stereo frames using one negotiated callback model, and measure what crossed the boundary. Once those facts are known, resampling quality and host latency can be tuned without guessing about buffer shape or callback ownership.
Related:
- Audio Resampling in Emulators: Reconciling Console Clocks with Modern Sound Hardware
- Inside Libretro: The Core/Frontend Architecture Behind RetroArch
Sources: