PlayStation SPU: ADPCM Voices, Envelopes, Sound RAM, and Reverb
Trace PlayStation SPU voices from ADPCM blocks and key-on state through ADSR, pitch, sound RAM, CD mixing, IRQs, and reverb memory.
The original PlayStation sound system combines an SPU with a CD-ROM decoder. The SPU provides 24 voices, local sound RAM, per-voice pitch and envelope controls, noise and modulation options, and a digital reverb path. The CD-ROM decoder can also produce CD audio or XA audio that enters the mixer. A game that streams music from disc therefore uses more than one audio path, and a sound emulator that models everything as “24 samples mixed at 44.1 kHz” misses important distinctions.
Sony’s PlayStation Hardware manual is the primary architectural source for the licensed development model. It describes the SPU’s voices, sound buffer, pitch, envelope, reverb, and mixing with CD decoder output. PSX-SPX supplies a much more detailed community-assembled register reference; its reverse-engineered observations should be labeled as such, especially where it explicitly marks uncertain bus or timing behavior.
Twenty-four voices share one SPU memory space
Each SPU voice can play ADPCM sample data from the SPU’s own sound RAM and has programmable playback properties such as pitch, left/right volume, envelope, and loop point. The voice hardware is not a CPU-side software mixer. Software configures voice registers and issues key-on/key-off operations; the SPU then advances playback state and contributes samples to the output path.
The SPU sound RAM is a separate address space. The CPU does not read it as ordinary main RAM; data is moved through the SPU’s I/O registers and DMA path. This matters for both sample upload and save-state handling. A host audio buffer may contain the decoded waveform, but it is not a substitute for the emulated sound RAM if software can update or reuse that memory while a voice is active.
A voice’s control state includes an address into sound RAM, pitch step, volume/envelope state, playback/loop state, and flags. A trace that logs only the output amplitude cannot reveal whether the error began at a wrong sample address, block header, pitch accumulator, key transition, or envelope phase. Capture per-voice state when diagnosing a missing note.
ADPCM blocks are decoded with history
The SPU’s sample voices consume compressed ADPCM data. Decoding depends on block metadata and predictor history; it is not equivalent to independently expanding each nibble into an unrelated PCM sample. The reconstructed values then feed interpolation and voice processing. Loop points are associated with sample data and affect where playback resumes after reaching a marked region.
PSX-SPX documents the block structure and register semantics, while Sony’s manual confirms ADPCM voices at the architectural level. Keep claims about undocumented filter behavior tied to the chosen reference. Do not apply the PlayStation SPU block format to CD-XA sectors: XA decoding is handled by the CD-ROM decoder and has its own sector-level format and route into the mixer.
Pitch control changes how the decoded stream is consumed over time. A voice may read at a rate different from the nominal source sample rate, and adjacent-voice modulation can vary pitch. An emulator should preserve the integer/fixed-point state used by its chosen hardware model; calculating each host output sample from a fresh floating-point resampling ratio can accumulate drift or alter loop boundaries.
ADSR and key transitions define note shape
The SPU supports attack, decay, sustain, and release envelope phases with programmable rates and curve behavior. Key-on starts a voice according to its current address and envelope state; key-off initiates release rather than necessarily silencing output instantly. Some games alter pitch or volume while a voice is playing. Therefore, register writes during playback must update the appropriate running state at the correct SPU event, not wait for the next video frame.
Volume can be controlled independently for left and right channels. Some voice modes use noise or pitch modulation instead of a straightforward sample stream. The SPU’s voice flags and masks are part of the emulated machine state. A correct mixer can still be wrong if its key-on mask, end flag, loop behavior, or envelope timer is stale.
Implement each voice as a state machine advanced on the SPU clock. Keep the following explicit: ADPCM block address, decoded sample history, current interpolation phase, pitch step, left/right volumes, ADSR phase and timer, key state, loop/end flags, and any modulation/noise source. Then let a separate mixer combine the per-voice outputs and external audio inputs.
Reverb is a memory algorithm
The PlayStation reverb unit is not merely a host effect toggle. It uses a work area in SPU RAM and a set of delay/filter address and coefficient registers. Voice sends select which voice outputs feed the effect. The reverb pipeline combines reflections, comb-like delays, and all-pass stages, then applies output levels. Its working memory can be observed and modified through the SPU memory access path.
This architecture explains why reducing reverb to a generic platform reverb plug-in can alter timing, stereo image, clipping, and interaction with memory writes. Games may configure reverb parameters per scene or use the work area for changing presets. An emulator can optimize the filter math, but it should preserve the register/state transition and deterministic output contract.
The reverb area is part of save-state data. If the emulator saves voice registers but not the SPU RAM region and current reverb address state, resuming a game can produce a discontinuity or a different tail. Validate a state capture while a sound and reverb tail are active, not only at a silent menu.
CD audio and XA take different routes
Sony’s hardware manual describes a CD-ROM decoder as a separate block from the SPU. The decoder supports CD-DA and CD-ROM XA audio, and its output is mixed with SPU output before final playback. CD-DA samples do not need to be represented as one of the 24 sample voices. XA sectors are decoded by the CD path rather than by the SPU voice ADPCM reader.
This difference is useful in diagnosis. If a Red Book audio track is silent but sample voices work, inspect disc track metadata, CD command state, and external mixer routing. If XA audio fails in a video sequence, inspect sector coding information and CD decoder behavior. If all 24 voices fail but CD music plays, inspect SPU RAM, voice masks, key state, or SPU DMA. Treating both as “ADPCM” obscures the actual decoder responsible.
The final mix has to align timestamps. CD decoder output, SPU voice samples, and host audio callback blocks may have different internal buffering. Use a common emulated audio clock and explicit resampling between source and host rate. Do not adjust per-game latency to compensate for a missing CD/SPU synchronization event.
Timing and I/O caveats
PSX-SPX records edge cases and measurements around SPU register access, DMA, and reverb. Some details are reverse-engineered and marked uncertain. When implementing them, tag evidence in tests as official-manual, measured-hardware, or cross-emulator. A single implementation quirk should not be presented as a guaranteed silicon rule unless a primary source supports it.
The CPU interacts with SPU state through memory-mapped registers and DMA transfers. A write is not necessarily audible immediately at the same host sample. The sound engine has its own processing cadence, so a register write can take effect on a later SPU update. When exact same-cycle ordering is unknown, preserve a documented convention and test for stable behavior rather than introducing nondeterministic host thread races.
Verification strategy
Build a test suite around individual voices and routes:
- Decode known ADPCM blocks with loop flags off and on; verify block address, predictor history, and end/loop transitions.
- Trigger key-on and key-off at known emulated times and record ADSR phase/output samples.
- Sweep pitch and verify source-address progression without host-rate drift.
- Exercise left/right volume independently, noise, and voice-to-voice pitch modulation.
- Upload samples through SPU DMA, read them back using the documented interface, and compare sound RAM bytes.
- Enable reverb with a controlled impulse, capture work-area changes, and restore from a save state mid-tail.
- Route CD-DA and XA separately from SPU voices and verify that each decoder reaches the mixer through the expected path.
- Change a voice register while playback is active and verify the state transition at the emulated sound event, not just the next rendered video frame.
Store the disc image hash, region, SPU register trace, sound RAM snapshot, emulator build, and raw output samples. Human listening remains valuable, but objective waveforms reveal a one-sample loop discontinuity or a short underrun more reliably than a compressed screen recording.
Acceptance criteria
A dependable SPU implementation keeps voices, SPU RAM, DMA, envelopes, pitch, key/loop flags, reverb state, CD decoder, and final mixer as separate but timestamped components. Tests cover each boundary and distinguish official Sony documentation from reverse-engineered assumptions. Save states include all local memory and pending audio events. Output is deterministic given the same disc, firmware, register trace, and host-independent sample rate.
The SPU’s engineering value lies in that division of work: compressed voice playback, modulation, envelopes, memory-backed effects, and external CD audio converge in one device but retain distinct inputs. Emulators reproduce it more faithfully when they preserve those boundaries instead of flattening every sound source into one pre-mixed stream.
Related:
- Retro Sound Chips: PSG, FM, Wavetable, and Sample Playback Architectures
- Audio Resampling in Emulators: Reconciling Console Clocks with Modern Sound Hardware
Sources: