Nintendo 64 RDP: From Microcode Display Lists to Filtered Pixels
Follow Nintendo 64 graphics from RSP microcode and RDP command streams through TMEM, rasterization, blending, framebuffer writes, and VI scanout.
The Nintendo 64 graphics path is not one “GPU” that receives triangles and immediately displays them. The Reality Signal Processor (RSP) executes microcode that interprets game display lists and prepares work. The Reality Display Processor (RDP) consumes a command stream, rasterizes primitives, samples texture data, performs depth and color operations, and writes images into memory. The Video Interface (VI) then reads an image for display and applies output-format and filtering behavior. Each stage has separate state and synchronization.
This pipeline division explains why an emulator can render the right geometry and still look wrong. A missing texture may originate in an RSP microcode path, an incorrectly loaded TMEM tile, a combiner input, coverage/blending state, framebuffer feedback, or VI scanout. Treating every mismatch as “a shader bug” skips the hardware boundary where the error entered.
The Nintendo 64 Programming Manual documents the architecture and graphics programming interface. Low-level software implementations such as Angrylion RDP Plus are useful executable references for behavior. Those projects are emulation code, not replacement hardware specifications; where the manual is silent, label an implementation-derived rule as such.
RSP output is not the same thing as RDP work
Game code commonly builds a graphics display list from Graphics Binary Interface (GBI) macros. The RSP runs a microcode program that interprets those commands and emits work for the RDP. Different microcode families can implement different high-level command sets even though the downstream RDP remains the same rasterization device. A high-level emulator plugin that recognizes a familiar microcode can translate its intent into host graphics operations, while a low-level path must reproduce the generated RDP command behavior more directly.
That distinction matters for compatibility. A command decoder that assumes one GBI can misread another microcode’s opcodes. Conversely, an RDP implementation should not infer application-level “draw triangle” semantics from an arbitrary byte sequence without first establishing which processor generated it. Preserve diagnostic traces at both layers: decoded RSP/GBI operations and the resulting RDP command stream.
The RDP receives state-setting commands as well as primitive commands. State includes color/texture image addresses, tile descriptors, texture filtering and wrapping, combine modes, render modes, scissor bounds, and depth behavior. A triangle command consumes the state active when it executes. If the emulator defers rendering, it must snapshot that state or preserve the original command order; reading the final register values after a whole frame can apply later state to earlier triangles.
TMEM is a small working store
The RDP includes 4 KiB of texture memory, TMEM. Texture data is loaded into that local store from system memory through RDP commands. Tile descriptors define how subsequent texture operations interpret an area of TMEM: format, size, line stride, bounds, palette selection, and clamp/mirror/mask behavior. They do not make the source texture a permanently resident, arbitrary-size host image.
The small store is a key part of the architecture. Games reuse TMEM strategically, load texture blocks or tile rows, and arrange image layouts around the available space. Texture load commands and tile setup are distinct operations. An emulator that simply samples the original RDRAM texture pointer can miss swizzles, tile boundaries, masks, palette lookups, and overwrites caused by later loads.
Texture formats include different texel widths and indexed modes, with lookup tables stored in the RDP’s texture memory. The maximum texture region per tile depends on format and stride; “TMEM holds 4 KiB” does not imply that every game can address one contiguous 4 KiB texture in the same way. Validate addressing with format-specific cases, odd row widths, tile origins, masks, and transfers near the end of TMEM.
Rasterization includes coverage, depth, and color stages
The RDP rasterizes triangles and rectangles from command parameters that include edge equations and interpolated attributes. It determines which subpixel samples are covered, derives texture/shade/depth values, tests depth where configured, runs the color combiner, and passes results through the blender and memory interface. Coverage is not interchangeable with a simple integer pixel-center test. It participates in antialiasing and in how partially covered pixels interact with stored framebuffer values.
The combiner selects and combines inputs such as texture, shade, primitive/environment colors, and constants. The blender controls how the result is mixed with memory or other inputs, subject to the render mode. A “texture looks too dark” report may be caused by incorrect alpha, cycle mode, blender flags, or a framebuffer read rather than a texture decode. Log command state in the same order the RDP consumes it.
Z-buffer behavior is also coupled to primitive order and coverage. A renderer that uses the host GPU’s default depth convention must translate near/far ranges, compare modes, update masks, and precision deliberately. Host APIs can differ in depth range, coordinate conventions, filtering, and edge rules. Validate the emulated result before adding host post-processing.
Framebuffer memory is observable state
The RDP writes color and, when enabled, depth information into memory. Those images can be read later by the CPU or reused by another graphics operation. Some games use framebuffer effects, motion blur, copy operations, or render-to-texture-like techniques. A pure host display list that never materializes emulated memory can be fast but may not preserve every CPU-visible result.
The RDP memory interface and the RSP’s command-generation path should expose synchronization points. A command processor may be busy while the CPU edits related memory or submits more work. The emulator must model whichever ordering guarantees software can observe, including cache visibility where relevant. Do not use an unconditional “finish GPU after every primitive” workaround as a correctness architecture; it can obscure the dependency and destroy throughput.
Coverage and framebuffer state are a particularly useful forensic clue. If a pixel differs only when rendering over an existing color image, capture the prior pixel, coverage value, alpha, depth result, combiner inputs, blender mode, and final memory write. A screenshot shows the symptom but not the command state that produced it.
VI scanout is a separate graphics stage
The VI reads the selected framebuffer using its own origin, width, size, scale, and timing registers. It converts the stored image to a display signal and can apply filtering behavior. Thus a correct RDP framebuffer may still be presented with wrong crop, aspect, gamma, or filtering if VI state is ignored. Conversely, changing a host output filter may make the screen resemble a reference while concealing an incorrect VI model.
Capture both the RDP color image and the final VI output. The first isolates command processing and memory; the second validates scanout. Use a documented VI register snapshot when comparing two runs. Account for interlaced or field-based modes and framebuffer origin changes before concluding that the RDP drew the wrong rows.
Emulator design and validation
An accurate implementation can be divided into command fetch/parse, state setup, TMEM loads, primitive rasterization, memory access, synchronization, and VI scanout. Keep each layer testable. A deterministic trace should include RSP microcode identity, command addresses, RDP state changes, tile descriptors, TMEM writes, primitive bounds, framebuffer addresses, and VI registers. Such a trace is substantially more useful than a single global “graphics accuracy” flag.
Use tests that isolate interfaces:
- Decode representative GBI operations under each supported microcode, then compare the emitted RDP commands.
- Load textures in multiple formats into TMEM and test tile stride, palette lookup, wrap, mirror, clamp, and masks.
- Render triangles with shared edges, subpixel movement, partial coverage, and reversed winding.
- Exercise depth compare/update combinations and overlapping primitives.
- Use blender and combiner combinations with alpha, shade, and memory inputs.
- Read back or reuse framebuffer data, rather than validating only the final screen.
- Change VI origin, width, scaling, filtering, and interlace state while keeping the RDP image fixed.
- Save and restore while command processing or VI scanout has pending work.
For regression captures, preserve native framebuffer dumps and final display output separately. Compare the image after controlling for crop and scale. If a software RDP reference and a host-accelerated path disagree, first compare their input command stream and state snapshots; only then investigate raster rules or host API differences.
Practical accuracy boundaries
High-level graphics translation can be a reasonable performance strategy when the emulator recognizes the microcode and the title does not depend on low-level RDP behavior. Low-level rendering is preferable for unusual commands, framebuffer effects, or pixel-sensitive coverage. These are implementation tradeoffs, not universal labels such as “HLE is inaccurate” or “LLE is perfect.” A low-level renderer can still have bugs, and a well-tested high-level path may reproduce a game’s visible contract.
The crucial requirement is observability. Make it possible to identify which stage produced a pixel and which source data it consumed. Version reference dumps and test ROMs, record the target console mode, and distinguish manual-supported behavior from reverse-engineered assumptions. This allows a compatibility workaround to remain narrow instead of silently changing the RDP model for every title.
The N64 pipeline rewards disciplined boundaries. The RSP interprets a program, the RDP evaluates a command stream, memory stores the result, and the VI turns memory into a display. Preserve those boundaries and a blurry or missing pixel can be traced to a concrete fetch, interpolant, blend, or scanout parameter rather than guessed at from the final screenshot.
Related:
- PlayStation GPU Texture Pages and CLUTs: Decode the VRAM Addressing Model
- Inspecting MAME Vector Displays with Lua Device Notifiers
Sources: