NES APU Timing: DMC DMA, Frame Sequencing, and CPU Bus Side Effects
Trace NES DMC sample fetches, APU frame events, OAM-DMA contention, and controller-read glitches with cycle-aware emulator models and tests.
The NES audio processing unit is not just five waveform generators attached to a sample mixer. Its delta modulation channel reads sample bytes through the CPU-visible memory system, and its frame sequencer clocks envelopes, sweep units, and length counters on a schedule that is related to, but independent from, picture processing. Those interactions are observable by software. A game can use DMC playback while polling a controller, and a cycle-insensitive emulator may occasionally lose a button press even though audio sounds plausible.
This article concentrates on two timing interfaces: the APU frame counter and the DMC memory reader. It does not attempt to describe every pulse, triangle, or noise-channel quirk. The goal is to make clear which state belongs to the APU, which bus cycles belong to the CPU and DMA units, and which tests can distinguish a waveform approximation from a useful hardware model.
The frame sequencer is not the video frame
The frame counter is programmed through $4017. In four-step mode, it clocks envelope and triangle linear-counter work at quarter-frame events, and clocks length counters and pulse sweeps at half-frame events. The sequence also produces a frame interrupt unless inhibited. Five-step mode has a different event pattern and does not generate that four-step frame IRQ. The familiar names are convenient, but they should not be implemented as a callback that fires once per rendered video frame: the APU sequencer maintains its own CPU-cycle schedule.
On NTSC hardware, the four-step sequence is approximately 240 Hz at its quarter-frame cadence, but its exact event positions matter more than the rounded frequency. A write to $4017 resets/restarts the sequence with parity-sensitive timing. Depending on the phase of the write, the first sequencer action follows after a small delay rather than happening at the same abstract instant as the CPU store. Five-step mode also has a startup clocking behavior that differs from four-step mode. PAL uses different timing constants. Do not use the host display refresh rate as the APU clock.
An emulator should therefore schedule frame-counter transitions against emulated CPU time. If audio is generated in blocks, the APU’s register and sequencer state must still advance at the original event positions inside each block. Otherwise, length expiration or envelope changes move relative to CPU writes, and software that synchronizes to the frame counter can diverge even when the final average pitch seems right.
DMC playback separates a reader from an output unit
The DMC plays one-bit delta-encoded data. Its memory reader obtains bytes from CPU address space, while a sample buffer feeds the output unit. The output unit consumes bits at a rate selected through $4010, adjusting a seven-bit output level; the reader fetches another byte when needed. That separation explains why a DMC channel can continue producing the current buffered bits while a later memory fetch is pending.
$4012 and $4013 specify the sample start and length using hardware-defined units and address formation, not arbitrary byte pointers. DMC sample fetches are confined to the upper part of the CPU address map. $4010 controls rate, looping, and optional IRQ behavior; $4011 directly changes the output level; $4015 enables the channel and exposes status. The byte-count status is not identical to “audible output is nonzero”: the output unit can still be shifting buffered bits after the reader has no further bytes to fetch.
This distinction is valuable while debugging silence at the end of a sample. Log the remaining-byte counter, sample-buffer occupancy, bits remaining in the shift register, output level, loop/IRQ flags, and the next scheduled reader request. A single dmc_playing boolean cannot represent all of those states.
DMC DMA owns real CPU bus cycles
The DMC reader does not magically index a ROM array in parallel with the CPU. A fetch is performed by the console’s DMA logic and temporarily halts CPU progress. DMC DMA uses a sequence of halt/dummy/alignment/read cycles; the exact alignment depends on which APU bus phase is available and whether the CPU is currently performing a read or write. For the common documented sequence, a fetch consumes three or four CPU cycles. A write cycle cannot be treated as interchangeable with a read cycle when deciding when a halt succeeds.
OAM DMA, started by writing $4014, also suspends normal CPU execution while copying a page into PPU OAM. The two DMA engines can overlap parts of their schedules. A DMC read takes priority over an OAM-DMA read, delaying that transfer and potentially changing its alignment cost. Modeling these as two independent timers that each add a fixed number of cycles produces the wrong total duration and can reorder side effects.
The bus effects are observable. A DMC fetch that lands on a CPU-visible register read can cause an extra read of an address with side effects. Controller ports are the best-known example: repeated reads can shift the serial button stream, making a button appear released. PPUDATA reads and APU register accesses also require careful handling. The visible symptom may look like a controller defect, but the root cause is DMA bus arbitration.
Keep bus arbitration explicit
A useful model has separate CPU, APU, OAM-DMA, and DMC-DMA state, but one shared cycle scheduler. Represent whether each CPU cycle is eligible for a DMA get or put, which bus address is driven, and whether the CPU operation is a read or write. The following is an architectural sketch, not cycle-exact drop-in code:
on_cpu_cycle(cycle):
advance_apu_timers_and_frame_sequencer(cycle)
request_dmc_fetch_if_sample_buffer_needs_data()
if dmc_fetch_can_halt_this_cpu_cycle():
perform_dmc_dma_state_transition()
return
if oam_dma_active and oam_dma_can_use_bus(cycle):
perform_one_oam_dma_bus_phase()
return
perform_cpu_bus_operation()
The ordering inside the real scheduler must be chosen from the target’s timing rules and verified with test ROMs. This pseudocode emphasizes one invariant: DMA, CPU reads, and APU events share a timeline. It is unsafe to complete all CPU instructions first and then subtract a frame-level DMA duration afterward.
Record the emulated cycle, APU phase, CPU operation type, address, DMA owner, OAM index, DMC sample address, and any register side effect. If a frame counter edge or timer advances on a particular cycle, record it too. A trace that says only “DMC DMA took four cycles” cannot explain a mismatch if the disputed event is whether a controller read occurred before or after the fetch.
Sample playback is also an interrupt contract
When the DMC IRQ-enable bit is set, reaching the end of a non-looping sample can request an IRQ. Looping changes the byte reader’s restart behavior; it is not equivalent to repeatedly issuing a new $4015 write. Writes to $4015 clear the DMC interrupt flag, while disabling the DMC stops future byte fetches and clears the remaining-byte state. The output unit’s buffered bits may take time to drain, so the channel’s audible tail and status bit can differ during shutdown.
Software that uses DMC IRQs must coexist with other interrupts and with music code. On the emulator side, do not conflate the DMC interrupt flag with the CPU’s global interrupt-enable state. The APU requests an interrupt; the CPU decides when to recognize it according to its own instruction-boundary rules. This distinction becomes important when a DMA halt overlaps an instruction whose final bus cycle changes interrupt state.
A focused verification matrix
Use one test for each contract instead of judging accuracy from a game boot alone:
- Run APU frame-counter tests in four- and five-step modes, both write parities, with IRQ inhibit toggled. Compare event positions, not only the audible envelope.
- Run DMC output/rate and IRQ tests for several sample lengths, including the shortest supported samples, looped playback, and termination while a byte is buffered.
- Trigger DMC fetches while OAM DMA is active. Record total CPU stall cycles and verify the final OAM contents.
- Read controller ports with DMC playback at several rates. Confirm that the serial stream matches the documented hardware behavior for the tested system revision.
- Repeat timing tests under NTSC and PAL configurations. Do not silently reuse NTSC event constants for PAL.
- Save and restore a state while a DMC byte fetch is pending. The scheduler phase, remaining bytes, sample buffer, and interrupt flags should resume deterministically.
Keep the test ROM, console/revision, region, emulator build identifier, and trace hash with the result. Hardware observations and reverse-engineered documentation can be revision-sensitive; an emulator should label what it models rather than claim universal behavior from one test board.
Engineering acceptance criteria
For a production emulator, the APU should advance independently of host audio callback size, and each DMA request should be placed on the emulated bus rather than charged as a coarse fixed penalty. The test suite should assert cycle totals, bus ownership, register side effects, sample-byte sequence, and IRQ timing. If a faster execution mode relaxes those details, expose it as a deliberate accuracy tradeoff and keep the cycle-aware mode as the comparison oracle.
This model also helps avoid a common debugging detour: do not “fix” a lost controller press by changing input polling frequency until you have checked whether DMC DMA is causing the extra register read. Similarly, do not move APU frame ticks to match the renderer’s VBlank callback. The NES’s CPU bus, PPU beam, and APU sequencer are related clocks, not one interchangeable frame counter.
Related:
- NES PPU VBlank and NMI Timing: Model the Event, Not Just the Frame
- Retro Sound Chips: PSG, FM, Wavetable, and Sample Playback Architectures
Sources: