Skip to content
RetrogamingDeep Dive Published Updated 9 min readViews unavailable

SNES SA-1 Coprocessor: Dual CPUs, Shared Memory, and DMA

Trace the SNES SA-1's two CPU views, shared work RAM, interrupt flags, arithmetic hardware, DMA paths, and deterministic bus-visible synchronization.

The SA-1 is not a graphics command processor that receives a finished job from the SNES CPU. It is a cartridge chip containing a second 65C816-family execution engine and a collection of memory-mapping, interrupt, timer, arithmetic, and transfer functions. A game can let the two processors cooperate through flags and shared memory, use the SA-1 for arithmetic or data movement, and expose the resulting timing to the console CPU. Correct emulation therefore requires two independently resumable processor states plus explicit rules for when each side can observe a change.

The SNES development documentation describes the programmer-visible registers and operating model. Snes9x’s upstream SA-1 implementation is useful for mapping those registers to execution and memory behavior, but implementation details such as instruction batching are not automatically statements about chip timing. Keep the documented interface separate from emulator scheduling policy.

Two processors, one cartridge

The console’s 65C816 and the SA-1 execute separate instruction streams. Each has its own program counter, status, stack, data bank, and execution state. The cartridge exposes register windows that let software configure SA-1 execution, control interrupts, select memory behavior, and inspect results. The two CPUs are not independent computers: they meet at cartridge ROM, battery-backed work RAM, internal RAM, and the SA-1’s mapped control registers.

Represent both processors with separate architectural contexts. A single shared PC or status register makes it impossible to resume one CPU while the other is waiting or executing. Serialize both contexts, pending interrupt state, timer state, mapping registers, and any in-flight transfer state. A save taken during a handoff must preserve which processor last changed each shared byte and which interrupt or DMA completion is still pending.

The work is cooperative rather than automatically race-free. A common pattern is for one processor to publish a command or buffer-ready flag, then allow the other to process it and set a completion flag. The exact meaning of a flag is game-defined; the hardware does not turn it into a queue with ownership or atomic transactions. An emulator must preserve byte-write visibility and the ordering of register effects rather than inventing a higher-level mailbox abstraction.

Model the address spaces independently

The SA-1 and the SNES CPU can have different views of the cartridge address map. Mapping controls select ROM banks and work-RAM behavior, including bitmap-oriented access paths used by software. It is not sufficient to create one host pointer table and let both processors use it unchanged. A read at one CPU-visible address may resolve differently depending on the active mapping and which processor issued it.

Keep address translation in distinct, testable functions for the console CPU and SA-1. The functions should consume the current bank and mode registers, then resolve a logical address to ROM, internal RAM, work RAM, or an I/O register. Do not cache a translation across a write to a mapping register unless the cache is invalidated at that exact transition. Snes9x’s SA-1 map setup illustrates why the map is stateful: bank selection and bitmap mode update the SA-1 map separately from the console view.

Shared work RAM is both storage and a synchronization surface. If the emulator stores it twice for convenience, every write must update the correct visible copy at the correct point. Prefer one canonical byte store with processor-specific access rules. Add tests in which one processor writes a sentinel, changes banks, and the other reads at the documented window. Cover ordinary byte accesses, boundary addresses, and each supported mapping mode.

Internal RAM, work RAM, and bitmap access

The chip includes internal RAM for SA-1-local work and can access cartridge battery-backed RAM. Software can also select bitmap-oriented work-RAM modes that expose packed pixel data through a different access interpretation. These modes are not the same as a general-purpose framebuffer API: the mapping and packing rules determine how CPU reads and writes are interpreted.

Treat memory mode as part of the address decoder, not as a postprocessing step applied to a completed image. For packed bitmap writes, test adjacent pixel fields that share a byte, reads after partial writes, and addresses at mode boundaries. Verify whether the console CPU can observe the same backing bytes or a different mapped window under the selected mode. Keep the raw work-RAM bytes available in debugging tools so a display renderer cannot conceal a mapping defect.

Do not assume all RAM is initialized to zero at every reset or content load. Distinguish power-on defaults, reset effects, cartridge-provided battery data, and emulator policy. Snes9x source initializes register fields and mapping state, then reconstructs derived maps after state restoration; that is a useful reminder that pointers are derived state while memory contents and register values are persistent state.

Interrupt flags and processor handoff

The SA-1 has control and status registers for interrupt enable and pending conditions. Software may use CPU-to-CPU flags as a lightweight handoff: one side raises a condition and the other observes or acknowledges it. Timer and DMA activity can also contribute interrupt conditions. Keep enabled, pending, and acknowledged state distinct. A pending event should not disappear merely because the receiving CPU is currently masked or waiting.

An interrupt is a visible boundary, not just a callback. Record the event’s emulated time, source, target processor, mask state, and acknowledge transition. If the target is waiting for an interrupt, the wait state must end at the correct point and normal instruction entry behavior must follow. Test both event-before-mask and mask-before-event sequences, repeated events while a flag is already set, and explicit software clearing.

Register names and bit meanings should come from the documented SA-1 interface. Avoid copying a bit mask from an emulator and assuming it is a complete hardware definition: the implementation may combine pending flags, compatibility behavior, and derived state. Keep register reads and writes centralized so read side effects, write-one-to-clear behavior, and open-bus behavior can be tested independently.

Timers and arithmetic hardware

The SA-1 has timer functions that can be used to generate an interrupt at programmed video positions, as well as arithmetic support that relieves the CPUs from some repeated calculations. A timer is a scheduled emulated event, not a host timer. Advance it from the console’s emulated clock and preserve its phase across scanlines, CPU yields, and state serialization. A frame-rate callback cannot reproduce a horizontal and vertical counter match that software reads or uses to coordinate rendering.

Arithmetic registers are also state machines. Software writes operands and an operation selection, then reads result and overflow state according to the device’s completion rules. Represent operand capture, result width, signedness, and availability timing explicitly when documented. Do not immediately compute a host integer and expose it earlier than the hardware would. Use boundary vectors: zero, maximum unsigned values, signed extremes, overflow, negative results, and a second operation started before the first result is consumed where the manual defines behavior.

When timing details are incomplete, make the assumption explicit in tests and isolate it from the register interface. Compare multiple implementations and game traces, but do not label an emulator’s cycle count as the chip’s exact count without a primary timing reference.

DMA is several behaviors, not one copy loop

The SA-1 exposes data-transfer facilities used for moving bytes between mapped sources and destinations, and it has character-conversion DMA paths that transform source data into bitmap-oriented output. These functions have different control registers and effects. Calling all of them one generic memory copy can silently skip alignment, address stepping, completion flags, source restrictions, and timing visible to either processor.

Create separate transfer engines or explicitly modeled modes. Each transfer needs source and destination address state, length or progress, mode, completion status, and an emulated-time budget. If the console CPU and SA-1 can access a resource during the transfer, define the observable ordering. Do not let a host memcpy finish the entire operation instantaneously unless the hardware contract says no intermediate observation is possible.

For character conversion, validate the produced bytes against known source patterns and the selected format. Test one tile, a source boundary, and a transfer that crosses a bank or RAM window. For ordinary DMA, verify byte order and which address increments. Check status before, during, and after a transfer, and test reset or abort behavior only where documentation or a known test establishes it. The implementation should not infer a universal cancellation rule from one game’s usage.

Deterministic scheduling and save states

The SA-1 and console CPU should advance against a shared emulated timeline. A simple cooperative scheduler can run a bounded number of cycles from one processor, stop at observable events, then advance the other. The exact scheduler may vary, but it must not run a whole frame of SA-1 instructions before processing all console CPU writes. That reverses causality when software changes a bank, signals a flag, or starts a transfer mid-frame.

Define synchronization points at control-register writes, shared-memory accesses, interrupt assertion and acknowledgment, transfer progress, and scanline events. Use cycle budgets that carry remainders rather than rounding every run to a line or frame. Keep host thread scheduling out of emulated ordering: two runs of the same deterministic test should yield the same register and memory trace.

For state files, store architectural registers, internal and battery-backed RAM as appropriate, flags, timer counters and phase, mapping modes, DMA progress, and pending results. Rebuild host pointers and lookup tables after loading from saved register state. A round-trip test should save during active SA-1 execution and during a transfer, then compare the subsequent instruction-visible writes, interrupt order, and final memory against an uninterrupted run.

A practical validation plan

Start with a minimal synthetic cartridge or a test harness that can control both processor contexts. Log emulated timestamp, processor identity, PC, register access, address translation result, memory byte, interrupt source, and transfer progress. Then test one contract at a time:

  • Boot with the SA-1 held or disabled according to documented control behavior; verify its reset vector and independent register state.
  • Let each CPU write a different byte to shared RAM and confirm visibility, bank mapping, and read order.
  • Raise and acknowledge each supported CPU-to-CPU flag in both directions, including masked and waiting states.
  • Program timer matches at a known scanline position and verify status and interrupt timing separately from frame IRQ behavior.
  • Exercise arithmetic result width and overflow at edge values.
  • Run ordinary and character-conversion DMA with short known buffers, checking every output byte and completion transition.
  • Save and restore with a running CPU, a pending interrupt, and an in-flight transfer; compare traces to uninterrupted execution.

Commercial-game validation should record the exact dump revision, region, emulator revision, and scene. A game booting or displaying a correct title screen is not proof of dual-CPU synchronization. Prefer small traces and memory assertions before judging screenshots or performance.

The central rule is to preserve independent CPU state while treating shared RAM, mapping, interrupts, and DMA as explicitly timed interfaces. Once each boundary is observable and testable, the SA-1 stops looking like a mysterious speed chip and becomes a set of concrete processor and bus contracts.

Related:

Sources:

Comments