Nintendo 64 Audio Interface: DMA FIFO, Interrupt Edges, and DAC Timing
Model the N64 Audio Interface as a clocked PCM DMA endpoint, from aligned RDRAM buffers and a two-entry queue to status transitions and sample-rate divisors.
The Nintendo 64 Audio Interface (AI) is easy to mis-model because it sits at the end of a larger sound pipeline. The CPU and Reality Signal Processor (RSP) cooperate to produce waveform data in RDRAM; the AI then reads that output at a programmed rate and feeds the audio DAC. The AI is therefore not the N64’s synthesizer, voice mixer, or sample decoder. It is a clocked PCM transfer endpoint with its own DMA registers, a small queue, status, and interrupt behavior.
That distinction is useful when diagnosing silence, clicks, or drift. A correct RSP audio task can still produce silence if the AI address, length, enable state, or DAC divisor is wrong. Conversely, a correctly running AI cannot repair malformed PCM produced upstream. Keep the software/RSP synthesis stage, the AI DMA stage, and the analog output stage visible as separate links in the trace.
The register pair describes queued work
The AI register block begins at 0x04500000. AI_DRAM_ADDR_REG stages a source address in RDRAM; AI_LEN_REG supplies its transfer length. The N64 register header documents these as double-buffered registers: software can program an address/length pair twice before the interface is full, and the address must be written before its matching length. The function reference describes AI_STATUS_FIFO_FULL as set when both transfer slots have been programmed, while AI_STATUS_DMA_BUSY indicates active DMA.
This is a two-entry transfer queue, not a general-purpose ring buffer managed by arbitrary host code. One transfer can be playing while the next waits. The producer’s job is to prepare complete PCM blocks in safe RDRAM, enqueue address then length, and replenish a slot when the device has freed one. The exact producer schedule depends on whether software waits on the AI event, polls remaining length, or coordinates its audio updates with another clock.
The address must be 64-bit aligned, and the transfer length must be a multiple of eight bytes. The hardware manual notes that the bottom three length bits are ignored. The final hardware’s length register is 18 bits and supports up to 256 KiB per transfer; version 1.0 exposes only 15 bits and supports at most 32 KiB. An emulator should model the relevant hardware revision rather than silently accepting any host buffer size.
Before writing a register, validate the emulated RDRAM range without allowing integer wraparound. Address translation from a CPU virtual or cached pointer into the N64-visible RDRAM address is a separate step:
#include <stdbool.h>
#include <stdint.h>
static bool ai_range_is_valid(uint32_t dram_address,
uint32_t byte_length,
uint32_t rdram_size,
uint32_t hardware_length_limit) {
if (byte_length == 0 || (dram_address & 7u) != 0 ||
(byte_length & 7u) != 0 || byte_length > hardware_length_limit)
return false;
/* Subtract only after the address is known to be in range. */
if (dram_address > rdram_size)
return false;
return byte_length <= rdram_size - dram_address;
}
Choose hardware_length_limit from the modeled AI revision (32 KiB for version 1.0, 256 KiB for the final 18-bit implementation). A real driver also needs the platform’s correct address mapping and cache-visibility rules; this helper only validates alignment, length, and range.
The AI interrupt follows FIFO state, not a vague “audio finished” event
One subtle detail is the relationship between the full flag and the interrupt. The N64 register header states that a transition of ai_full from 1 to 0 sets the AI interrupt. The status register is also writable to clear the audio interrupt. This describes a queue-capacity edge: when a previously full two-entry queue frees a slot, software is notified that another buffer can be submitted.
Do not replace that edge with a generic interrupt on every sample, an interrupt on every buffer’s final sample regardless of queue state, or an interrupt simply because a write to the length register occurred. A common producer keeps one transfer playing and one queued; when the front transfer completes and the queue ceases to be full, the edge gives software time to enqueue the next block. AI_LEN_REG separately reports the number of bytes remaining in the current transfer through the OS API, so code can inspect progress without confusing it with queue capacity.
The distinction matters in edge cases. If software never fills both slots, the full-to-not-full transition may not occur as it does in the steady-state pipeline. If software writes after the device is already full, osAiSetNextBuffer() reports failure rather than magically growing the queue. A deterministic emulator must track queued address/length pairs, the active transfer, the full and busy bits, and the interrupt’s pending/acknowledged state. Derive the interrupt at the documented state transition, not from a host audio callback’s timing.
Use a small state table in tests:
| Event | Queue state | Expected observation |
|---|---|---|
| First valid address/length pair | One entry | A transfer can start; queue is not full |
| Second valid pair while first is active | Two entries | Full is asserted |
| Active transfer completes with a queued successor | One entry | Full-to-not-full edge raises AI interrupt |
| Software acknowledges AI status | Queue unchanged | Pending AI interrupt is cleared |
| Attempt to enqueue while both slots remain occupied | Still full | OS helper reports failure; no third slot appears |
Exact timing of bus activity and interrupt delivery should follow the selected hardware model. The table captures queue and observable event semantics rather than claiming that an interrupt arrives at the exact host callback or wall-clock instant that consumes the last sample.
DACRATE is a divider, not the sample rate itself
AI_DACRATE_REG programs a sample-period divisor derived from the system’s video clock. The N64 register documentation gives the relationship as sample_rate = video_clock / (dperiod + 1). For a target rate, software chooses the closest representable divisor and should use the resulting actual rate when calculating audio buffer duration. Nintendo’s osAiSetFrequency() returns the closest rate generated by the internal divisors, rather than guaranteeing that every requested frequency is exact.
The register AI_BITRATE_REG is a separate half-period setting for the serial audio clock. It does not mean the interface can accept arbitrary PCM word formats or channel counts. The hardware header constrains the DAC period relative to this clock divider; invalid combinations can prevent output or corrupt sound. For ordinary emulator work, model both divisors and their relation, but do not reinterpret the bitrate field as “bits per sample” in a WAV-style format header.
Region matters because NTSC, PAL, and MPAL units have different video clock frequencies. The programming manual lists 48,681,812 Hz for NTSC, 49,656,530 Hz for PAL, and 48,628,316 Hz for MPAL. The same divider therefore produces slightly different sample rates on different system regions. A frontend that hard-codes one clock can slowly drift against the emulated machine even when its output buffer size looks sensible.
Keep requested rate, selected divisor, actual rational clock rate, and host device rate as separate values. If the host runs at 48 kHz while the N64 output is, for example, a nearby rate determined by its clock and divider, resample at the frontend boundary. Do not change emulated AI timing to match the host device; that makes software-visible remaining-length and interrupt behavior depend on the user’s sound card.
Buffer duration couples audio to frame scheduling
For signed 16-bit stereo PCM, one sample frame contains four bytes: two channels times two bytes per channel. A buffer of N sample frames occupies 4N bytes and plays for N / actual_sample_rate seconds. Because the hardware length is eight-byte granular, such a stereo buffer contains an even number of sample frames. When choosing its size, balance interrupt/queue servicing cost against latency: smaller blocks need more frequent replenishment, while larger blocks increase the time represented by each submitted transfer.
Nintendo’s audio-memory chapter makes a related distinction between audio frame rate and video frame rate. Audio can be updated at a different cadence from graphics. Its examples also show that output-buffer count depends on whether the application synchronizes to vertical retrace or to audio completion. Therefore, an emulator should not assume exactly one AI buffer per rendered video frame or tie sound generation solely to a graphics present call.
For deterministic tests, calculate a buffer’s duration from the selected AI divisor and its byte length, then compare the number of consumed sample frames against emulated time. Test both a producer that keeps two queue slots full and one that deliberately starves the AI. The latter should reveal a reproducible underrun or silence policy, not silently replay whichever host samples happen to remain in memory.
Cache visibility is upstream of audio quality
The AI reads RDRAM independently of the CPU’s cached view. Before starting DMA, the data the AI is meant to consume must be visible through the correct hardware address and cache-coherency path. A CPU can inspect a freshly generated buffer and hear old samples if the backing RDRAM was not updated or if the submitted address names a different alias. Likewise, a save state that preserves only the host audio callback queue but not AI address/length and pending state cannot resume the same transfer.
For an emulator, record both the AI-visible physical RDRAM address and the bytes actually read at each sample boundary. For a native N64 development setup, follow libultra or the chosen SDK’s cache-management requirements and use the documented aligned DMA buffers. Keep the producer-owned buffer separate from a buffer still owned by the AI; never overwrite the active source region simply because the host has finished mixing the next block.
A stage-by-stage validation plan
Test the interface independently of the game’s synthesizer first. Fill RDRAM with a short known stereo pattern, program an aligned buffer, set a legal DAC divisor, enable DMA, and verify the output sample sequence and duration. Then add queueing and the interrupt. Only after those pass should the test produce PCM through audio microcode or game sound code.
Useful cases include:
- An eight-byte transfer at the start and end of valid RDRAM, plus deliberately misaligned address and size.
- Lengths at the selected hardware revision’s maximum, one alignment unit below it, and one unit above it.
- One queued transfer, a full two-transfer queue, front-transfer completion, full-flag falling edge, and status acknowledgement.
- A write attempted while both entries are occupied, checking that no third transfer appears.
- The minimum and maximum supported DAC divisors and representative NTSC, PAL, and MPAL clocks.
- PCM buffers whose size maps to an integral duration and buffers whose duration crosses the game’s video update boundary.
- Save/restore while idle, while one buffer is active, and with two queue entries, then compare address, remaining length, status, pending interrupt, and sample phase.
Log enqueue and dequeue times in emulated ticks, buffer start address, requested and masked length, remaining bytes, AI_STATUS transitions, interrupt assertion/acknowledgement, active DAC divisor, and each sample fetch. A click then becomes attributable to an address change, an unexpected queue edge, a bad clock ratio, or stale memory instead of an undifferentiated “audio bug.”
Keep synthesis, transfer, and presentation separate
The N64 audio path has at least three useful debugging layers: software and RSP code create PCM; the AI transfers that PCM from RDRAM on its own schedule; and the host audio backend converts or resamples the emulated output for physical playback. The AI’s FIFO and interrupt are part of machine behavior. The host’s output queue is not a substitute for that FIFO.
Model the address-before-length programming sequence, two-entry capacity, alignment limits, hardware-version-specific length width, full/busy status, full-to-not-full interrupt edge, and divider-based sample timing. Once those contracts are explicit, crackles and clock drift can be investigated with reproducible register traces instead of guessed buffer sizes. The AI is a small interface, but it is a real asynchronous device boundary - and treating it that way is what makes audio emulation reliable.
Related:
- The Libretro Audio Callback Contract: Frames, Buffering, and Timing
- Nintendo 64 RDP: From Microcode Display Lists to Filtered Pixels
Sources:
- Nintendo 64 Programming Manual, audio memory usage and buffer sizing (Chapter 22)
- Nintendo 64 hardware register definitions for the Audio Interface
- Nintendo OS Function Reference:
osAiSetNextBuffer, status, alignment, and hardware-version limits - Nintendo 64 introduction: audio data path from RDRAM through AI to the DAC
- libdragon low-level AI register reference