Skip to content
RetrogamingDeep Dive Published Updated 8 min readViews unavailable

Sega Saturn SCU DSP: Program RAM, Data Banks, and Parallel Operations

Inspect the Saturn SCU DSP's small Harvard-style memories, instruction buses, multiply-accumulate path, program loading, and cycle-aware emulator tests.

The Sega Saturn’s System Control Unit contains its own programmable digital signal processor. This SCU DSP is a geometry and general arithmetic coprocessor, not the Saturn sound processor. Audio belongs to the separate SCSP subsystem; the SCU DSP has a compact program store, banked data RAM, arithmetic units, and paths to exchange data with other system memory. Confusing these processors leads to the wrong memory map, instruction model, and timing assumptions.

The SCU manual describes a DSP with a 32-bit by 32-bit multiplication path feeding a 48-bit result, 256 words of program RAM, and four banks of 64 32-bit data words. This is a highly constrained execution environment. Programs must fit a small instruction budget or move data and code through explicit host interfaces. The DSP’s value is not general-purpose capacity but the ability to perform a narrow sequence of arithmetic operations near the system’s memory and graphics coordination logic.

A separate execution and memory domain

The SCU DSP has its own program counter, instruction store, working registers, flags, stacks, and data RAM banks. The main SH-2 processors do not execute these instructions as ordinary code. They initialize and control the DSP through SCU register interfaces, while the DSP runs its own compact instruction stream. Keep that ownership boundary visible in an emulator debugger.

The small program RAM affects software architecture. A graphics or transform routine may be copied into program memory before execution, while parameters and intermediate vectors occupy data RAM or external memory. A DMA instruction can move data between the DSP and system buses, but the transfer count and address space are part of the instruction semantics. Do not model it as an unbounded host memcpy that completes outside the DSP schedule.

The manual’s instruction set exposes operations through multiple internal buses. Arithmetic/logic work, multiplier input, X/Y data paths, and D1-bus movement can be coordinated within an instruction. This is why a DSP operation should not be flattened into a sequence of unrelated host arithmetic calls without preserving the documented register-transfer timing and architectural side effects.

Multiply-accumulate and fixed-width behavior

The DSP’s 32-bit operands and 48-bit result path make fixed-width arithmetic central. A host language with arbitrary-precision integers will not automatically reproduce truncation, sign extension, carry, or the placement of high and low product bits. Model each register width explicitly, then define where arithmetic results are masked and when the condition flags are updated.

A multiply-accumulate routine often works on fixed-point geometry values. Scale is a software convention layered over the hardware’s integer datapath; the processor does not know that a particular bit means a fractional position. Documentation should state the chosen scaling and rounding at the program level. The emulator should implement the underlying bit-accurate multiply, add, shift, and saturation behavior rather than special-case one observed geometry library.

Instruction issue and completion can be pipelined or overlap across internal resources. The Sega manual’s instruction format and simulator material should drive the emulator schedule. Avoid inventing superscalar behavior from a modern DSP analogy, and avoid forcing all operations into a single “one instruction equals one cycle” assumption unless the documentation and hardware evidence support it.

Host setup and data transfer

The main CPU must place a valid program and its input data in the correct DSP-visible locations, configure any DMA or bus access, establish entry state, and enable execution. A missing result can be caused by a wrong upload order, a stale program counter, a data-bank selection error, a transfer direction mistake, or a result that has not yet been copied back. The same final CPU value can hide several different broken intermediate states.

Use a control-plane trace for DSP reset/enable, program RAM writes, program counter changes, execution start, stop conditions, and interrupt requests. A data-plane trace should record data RAM bank and address, source/destination bus, count, and each transfer’s start/completion cycles. Keep the two traces correlated through an operation identifier rather than dumping every DSP register at every cycle.

The SCU itself participates in multiple system functions, including DMA and interrupts. A DSP request can therefore contend with other SCU activity or be restricted by documented bus rules. Do not merge SCU DSP DMA with the three CPU-usable SCU DMA levels. Give each master separate status and arbitration identity, and consult the SCU manual and applicable technical bulletins for restrictions.

Instruction-level verification

Begin with isolated instructions whose outputs are easy to calculate: zero and nonzero logical operations, signed and unsigned additions if both apply, shifts at boundary counts, multiplication producing nonzero high bits, and moves among each register and data RAM bank. For a 48-bit result, test values that exercise sign, upper-word placement, and low-word carry. Check that the program counter and branch conditions change at the right point.

Then test short programs that exercise instruction packing, loops, calls/returns, stack state, interrupts, and DMA. Include a program that fills the entire program store boundary and one that attempts to fetch beyond valid program memory. The documented hardware response to invalid or out-of-range accesses should be represented deliberately, not left to host array behavior.

The Sega manual’s simulator instructions are useful as a debugging reference: a machine-level simulator can inspect program area and data area values, which mirrors the observability an emulator debugger should provide. A concise trace should show the instruction word, decoded instruction fields, pre-state, resource requests, post-state, and cycle count. This provides evidence when an arithmetic result disagrees with the source program.

A fixed-width arithmetic invariant

This Python reference illustrates a 48-bit masking boundary for tests; it does not define the Saturn’s exact multiply instruction encoding:

MASK48 = (1 << 48) - 1


def product_32x32_to_48(left, right):
    left &= 0xFFFFFFFF
    right &= 0xFFFFFFFF
    return (left * right) & MASK48


assert product_32x32_to_48(0xFFFFFFFF, 1) == 0xFFFFFFFF
assert product_32x32_to_48(0x10000, 0x10000) == 0x100000000

If a particular instruction interprets operands as signed, the conversion to signed values and the observable result-register layout must be tested from the SCU specification. Do not use this helper as a replacement for sign and flag semantics.

Save states and integration tests

A save state must include the active program image or a reliable reference to it, program counter, working registers, data RAM banks, stacks, flags, pending instruction or pipeline state, DMA progress, and pending SCU interrupt state. If the emulator schedules DSP steps on a shared event queue, also serialize the next execution deadline. Restoring only the SH-2 state can corrupt a game that was midway through a DSP transform.

Integration tests should compare DSP output buffers before they are submitted to VDP1 or VDP2. That isolates arithmetic from rendering. Test a deterministic program with known input matrices, then capture SCU bus activity and final command or table data. A visual mismatch alone cannot tell whether the fault came from the DSP, DMA, cache/bus view, or a downstream video processor.

When physical Saturn access is unavailable, label results accurately: manual conformance, source comparison, or agreement between independent emulators is not the same as silicon validation. The manual is a primary source, but any supplemental correction or emulator-derived quirk should be cited and versioned separately.

Acceptance criteria

A production-quality SCU DSP model exposes its private memories and state, implements fixed-width arithmetic and instruction transfers, respects separate DMA ownership, and advances in emulated time. It can show the exact instruction and data movement behind a result and restore that state across a save state.

The SCU DSP is best treated as a compact coprocessor with a strict instruction and memory contract. Do not call it a Saturn audio DSP, and do not equate an apparently correct frame with a verified arithmetic pipeline. Preserving processor identity and resource boundaries is the path to reliable emulation.

Program loading and instruction budget

With only 256 32-bit words of program RAM, code size is itself a runtime constraint. A routine that performs a matrix transform, geometry setup, or other repeated operation may be copied in compact form and reused with different data rather than treating the DSP as a large resident library. The host-side upload sequence should be testable independently from the instruction interpreter: verify the programmed address, word order, write completion, entry point, and enable/start transition before evaluating arithmetic output.

The four data banks give software a small workspace for operands and intermediate values. Make bank selection explicit in debugger views; displaying the same 8-bit address without its bank invites false conclusions about aliasing. Data RAM load/store instructions and host-programmed data-port accesses should use the same underlying storage model but remain distinguishable in traces because they have different initiators and timing.

Parallel paths need ordered side effects

The DSP instruction format can request multiple internal operations in one encoded word. An interpreter that decodes one field and discards the other parallel fields can appear correct on trivial programs while corrupting loops that combine arithmetic with data movement. Decode an instruction into a typed set of resource actions first, then apply register and memory side effects in the order required by the manual.

This is also why a JIT or recompiler should share the same semantic helpers as the interpreter. Differentially execute a compact test program in both engines and compare every architectural state boundary: program counter, accumulators, flags, data-bank writes, DMA state, and cycle count. A final transformed vertex array is not enough evidence if both engines accidentally share an incorrect output packing shortcut.

Related:

Sources:

Comments