Skip to content
RetrogamingDeep Dive Published Updated 8 min readViews unavailable

SNES Super FX GSU: Register Pipeline, Cache, and Bus Coordination

Model the SNES Super FX as an executing cartridge processor, tracing its R0-R15 state, program and data banks, cache, pixel operations, and CPU-visible timing.

The Super FX is not a polygon-rendering command that the SNES CPU hands to a modern graphics API. It is a cartridge-resident processor, commonly called the GSU, with its own instruction stream, register file, memory-bank controls, cache RAM, and pixel operations. The SNES CPU communicates through mapped cartridge registers and shared memory, while the GSU executes work that can continue independently. Accurate emulation therefore needs a second processor state machine and a rule for when its effects become visible to the console CPU.

The SNES Development Manual documents the GSU programming model, and Snes9x’s upstream source is a useful implementation reference for register access and instruction behavior. Source code is not the physical chip specification: emulator execution batching and compatibility workarounds can differ from silicon timing. Use the manual for programmer-visible behavior and use implementations as comparative evidence, especially when a timing detail is not fully specified.

The register file and execution state

The GSU exposes sixteen 16-bit general registers, R0 through R15. R15 is the program counter. The other registers can serve as operands, addresses, counters, or results depending on the instruction. Special registers hold status flags and bank or rendering configuration, including the program bank, ROM bank, RAM bank, cache-base, color, and screen-mode state.

The status register is not just an arithmetic result. It carries flags used for conditional flow and control state, including the Go state that indicates whether GSU execution is active. Instruction prefixes and register-selection state affect how following operations interpret source and destination registers. A disassembler or debugger that prints only the opcode and R15 can miss the status that makes the same opcode behave differently.

Serialize the program counter, visible registers, status and prefix state, bank selections, ROM data buffer, cache contents and mode, screen parameters, pixel state, and any pending execution timing. A save state taken while the GSU is running must resume at the same next instruction boundary and produce the same write ordering. Do not represent an active GSU as a single host callback with no intermediate state.

Program, ROM, RAM, and cache are distinct spaces

The GSU has separate bank controls for instruction fetches and data access. PBR identifies the program bank used by the instruction stream; ROMBR and RAMBR select banks for data paths. Treat a register value as a bank selection input, not a host pointer. Address translation should be centralized and account for the cartridge mapping and hardware-visible bank limits.

The 512-byte cache is a small internal execution resource, not another range of cartridge ROM or game RAM. Programs can arrange code for cached execution and use the cache controls to select its base behavior. Cache state can change which memory traffic occurs and when the SNES CPU can observe access to cartridge resources. If an emulator treats the cache as always-valid host instruction memory, code that depends on cache fill or invalidation behavior may diverge.

The GSU’s ROM data path also has buffering and timing behavior. Loads from cartridge data are not equivalent to arbitrary reads from an unconstrained host array. Keep the ROM bank, data address, fetch buffer, and visible instruction pipeline state in the model. When a program changes a bank or reads across a boundary, validate both the returned data and the subsequent execution state.

The instruction pipeline matters

The GSU fetches and executes instructions through a pipeline. A simple interpreter can still be accurate, but it must model which instruction is in the current pipe, how the next opcode is fetched, and the effects of branch and register-control operations on subsequent fetches. Treating the instruction pointer as a conventional CPU PC that advances only after a complete instruction can produce incorrect behavior around branches, cache execution, and writes to R15.

Instructions do not all have identical cost. Memory operations, multiplies, pixel operations, and cache activity have different hardware costs. The console and GSU share cartridge resources, so CPU-visible availability depends on the access type and execution state. An emulator can batch work for performance, but it needs synchronization points at observable shared-memory or register events. A batch must not run past a point where the other processor should see a result or change the GSU’s control state.

Use one clock domain or an explicit conversion between master clocks, GSU cycles, and SNES CPU cycles. Avoid host wall-clock throttling as a substitute. If the GSU executes until a scanline boundary in a fast mode, preserve the exact budget and pending remainder rather than assuming every line ends at the same instruction boundary.

Pixel instructions are device operations

The GSU includes operations for graphics-oriented work, such as plotting and reading a pixel, alongside ordinary arithmetic, branches, loads, and stores. Pixel instructions write to the cartridge’s framebuffer-related memory according to current color, plot options, screen mode, and coordinates. They are not interchangeable with drawing a host triangle: color-depth packing, address translation, transparency and screen mode belong to the emulated device.

Create test programs that set a known screen mode and color, plot a small set of pixels at boundaries, and then have the SNES CPU inspect the result through the documented mapping. Test each supported bit depth, clipping or wrapping behavior where specified, and both cached and uncached code paths if relevant. Compare raw emulated memory before scaling or host shader processing.

The rendering path may be part of game logic. A title can read the framebuffer or wait on a GSU-published flag before using generated graphics. A host renderer that draws pixels to a texture without updating emulated memory can display an image while leaving CPU-visible state wrong.

Starting, stopping, and sharing work

The SNES CPU configures GSU-visible registers and memory before allowing the processor to run. GSU programs can stop themselves, and the CPU can observe or alter control state through mapped registers. Implement reset and start as device transitions, not as side effects of content loading. Reset must initialize only the state the hardware reset affects; cartridge RAM contents and save data are separate.

The two processors may access common cartridge resources. Model bank changes and memory writes in deterministic order. If the 65816 writes a control register while the GSU is executing, define whether that change takes effect immediately, at an instruction boundary, or at a documented synchronization point. Avoid running the entire GSU frame before processing all SNES CPU writes; that reverses causality.

In a multithreaded emulator, host thread safety is only an implementation concern. The hardware model still requires deterministic arbitration and visibility. Use an event timeline, instruction stepping, or another documented scheduler approach, then verify that changing host thread scheduling does not change the emulated trace.

Diagnostics and conformance

For every instruction trace, record emulated cycle, current opcode, pipeline byte, R15, source and destination register selection, status/prefix flags, program bank, data bank, memory access, and output side effect. At shared boundaries, record which processor issued an access and whether another processor was active. A screenshot alone cannot prove correct bank mapping, register state, or interrupt timing.

Build focused tests:

  • Execute arithmetic and conditional branches with each relevant status flag state.
  • Branch across a bank boundary and verify the next fetch address.
  • Start and stop the processor through each documented path.
  • Execute loads and stores from ROM and RAM banks with known patterns.
  • Fill and execute from cache, then modify the corresponding backing memory.
  • Plot/read pixels at normal and edge coordinates across screen modes.
  • Save and restore while the GSU is active, then compare the exact future write sequence.
  • Run the same test with different host scheduling and verify identical emulated output.

For commercial-game testing, record the exact ROM revision, region, emulator build, and test scene. Keep reference captures unmodified and distinguish expected hardware output from shader, scaling, or color-management changes.

Common shortcuts to reject

Calling a host graphics routine to draw the GSU’s intended result bypasses program execution and will fail games that inspect memory, rely on timing, or change the scene based on processor state. Treating R15 as the only serialized register discards cache, bank, and prefix state. Treating all memory operations as a constant cycle cost hides shared-bus effects. Running the GSU only once per video frame introduces large timing steps and may let the SNES CPU read data before real hardware would publish it.

Snes9x source includes both performance strategies and compatibility corrections. Do not copy a convenient batching decision as if it were a silicon guarantee. Cross-check the manual, compare at least one independent implementation, and keep uncertain cycle-level behavior isolated so test evidence can refine it.

Acceptance criteria

A reliable GSU implementation preserves a resumable processor state, distinct program and data bank mapping, observable cache state, instruction and memory timing, pixel memory effects, and deterministic synchronization with the SNES CPU. Regression tests cover both computation and the moment another processor sees each result.

Thinking of Super FX as a second cartridge processor rather than a graphics shortcut clarifies the engineering boundary. The GSU is a stateful device with a program, a pipeline, and shared memory. Rendering correctness follows from reproducing those contracts.

Related:

Sources:

Comments