Nintendo 64 RSP Vector Arithmetic: Fixed Point, Accumulators, and Saturation
Understand Nintendo 64 RSP vector multiply, fixed-point rounding, the 48-bit accumulator, saturation, and emulator tests that expose hidden-state bugs.
The Nintendo 64’s Reality Signal Processor (RSP) is not a generic GPU shader engine. It runs microcode, pairing a scalar unit with a vector unit whose operations are part of the console’s observable behavior. The vector unit has 32 128-bit registers, each divided into eight 16-bit lanes. Many arithmetic instructions operate lane by lane, but their visible 16-bit results are only part of the state: multiply instructions also use an eight-lane, 48-bit accumulator.
That hidden precision is where seemingly plausible emulators go wrong. If a core calculates a product, immediately truncates it to 16 bits, and uses that truncated value in a later multiply-accumulate, the first output may look correct while geometry, lighting, audio, or microcode-specific results drift over a sequence. Accurate RSP arithmetic therefore means emulating both the value written to a vector register and the accumulator state left behind.
The data model: lanes and fixed-point values
An RSP vector register contains eight 16-bit elements, not four 32-bit floats. For the fractional multiply instructions, the bit patterns are interpreted as signed 1.15 fixed-point values: the most significant bit is the sign, and the binary point lies before the 15 fractional bits. The representable range is -1.0 through 0.999969482421875. Consequently, 0x8000 represents -1.0, while 0x7FFF is the largest positive value below 1.0.
The instruction can select or broadcast an element from one source vector while processing the destination lanes. This is useful for dot products and matrix operations, but it means a test that always uses the same source lane will not exercise the element-selection behavior. Keep vector lane order and selection semantics explicit in the decoder and test fixtures; do not infer them from host SIMD conventions, which may use a different lane numbering or byte order.
Why the accumulator is architectural state
Each of the eight lanes has a 48-bit accumulator. The RSP guide describes the accumulator as high, middle, and low 16-bit portions for the VSAR instruction, which exposes a selected slice to a vector register. Those slices are a readout mechanism, not three unrelated accumulators: carries and signs cross the boundaries. An implementation that stores three int16_t values and updates them independently will lose information unless it explicitly propagates carries and sign correctly.
For ordinary debugging, it is useful to think of each lane as one signed 48-bit number. The high and middle portions contribute to the 16-bit value returned by common fractional multiply instructions; the low portion preserves precision for later operations and can be inspected with VSAR. The accumulator is not simply a temporary copy of the destination vector. It survives operations that return a clamped value, which makes subsequent instructions depend on information that software cannot see in vd alone.
This distinction matters to save states and deterministic replay too. If a state is captured between two multiply-accumulate instructions, saving only the vector registers is insufficient. The eight accumulator lanes must be restored along with the other RSP state, or the next instruction can produce a different result even though every visible input register matches.
VMULF: fractional multiply with rounding
VMULF multiplies signed 1.15 inputs. Per lane, the mathematical product is doubled to form the fractional 1.31 result, and 0x8000 is added before the value is placed in the accumulator. The returned vector element is derived from the accumulator shifted right by 16 and saturated to signed 16-bit range.
The following is explanatory pseudocode, not a drop-in host-language implementation. In particular, a C or C++ implementation must define its signed-width and arithmetic-shift behavior rather than relying on undefined signed overflow:
for each lane i:
product = signed16(vs[i]) * signed16(vt[selected_element])
accumulator[i] = sign_extend_to_48(2 * product + 0x8000)
vd[i] = saturate_signed16(arithmetic_shift_right(accumulator[i], 16))
The classic boundary case is 0x8000 × 0x8000, or -1.0 multiplied by -1.0. The exact result is +1.0, which cannot be represented in signed 1.15. A correct signed saturation therefore returns 0x7FFF, not a wrapped negative value. This single case distinguishes saturation from truncation and catches a common fixed-point edge error.
VMACF: keep the precision between operations
VMACF performs a related signed fractional multiply, but adds the unrounded product into the current accumulator. It then returns a saturated signed 16-bit view of the accumulator’s upper result bits. Unlike VMULF, it does not add the round-to-nearest constant as part of that accumulation step. Crucially, returning a 16-bit vd value does not discard the full-width accumulator.
for each lane i:
product = signed16(vs[i]) * signed16(vt[selected_element])
accumulator[i] = add_in_48_bits(accumulator[i], sign_extend_to_48(2 * product))
vd[i] = saturate_signed16(arithmetic_shift_right(accumulator[i], 16))
Consider several small products added into one lane. Each individual product may be too small to change the returned 16-bit value, but their accumulated low bits can eventually carry into the middle portion. An emulator that rounds or truncates after every step changes when that carry becomes visible. Test the whole instruction sequence and inspect VSAR output; comparing only the final vd for one operation cannot prove the accumulator is right.
The VMULU naming trap
The U in VMULU does not mean that both input vectors are unsigned. The instruction multiplies signed fractional inputs like VMULF, then applies unsigned saturation when it returns the result. Negative values clip to zero, and positive overflow saturates to 0xFFFF. Its output interpretation is therefore different even though its multiplication inputs are signed.
This is a useful warning against implementing instructions from their mnemonics alone. Similar-looking opcodes can share a product calculation yet differ in rounding, accumulator update, extraction, or clamp behavior. Keep those stages separate in an emulator: decode the operation, calculate the precisely defined product, update the accumulator according to that opcode, select the output bits, and then apply the correct signed or unsigned saturation rule.
A regression plan that checks hidden state
Build small deterministic tests around instruction semantics before using complete games as an oracle. For every arithmetic test, initialize all eight lanes with distinct values; repeated lanes can hide a swapped index or an incorrect broadcast. Include positive and negative operands, zero, the minimum signed value, values immediately around the saturation boundary, and multiple element selectors.
For VMULF, test the zero product, positive and negative fractional products, and 0x8000 × 0x8000. Assert both the vector result and the accumulator slices returned by VSAR. For VMACF, begin from a known accumulator state, issue a sequence of products whose low bits carry into the next slice, and check after each step. Include positive and negative accumulation near saturation so that the visible result clamps without destroying the underlying accumulator.
Then test persistence boundaries. Save and restore an RSP state between two multiply-accumulate instructions and verify that the resumed trace matches an uninterrupted run. Exercise microcode that uses different element selectors and compare results against a trusted hardware test or a well-documented reference implementation. When hardware is unavailable, record the exact reference revision and test vectors; a game booting or rendering one scene is not an arithmetic conformance test.
Finally, keep performance and accuracy claims separate. Host SIMD can accelerate the eight lanes, but only if its signed widths, saturation, byte order, lane mapping, and accumulator precision reproduce the RSP rules. A fast approximate path may be a deliberate compatibility choice, but it should not silently masquerade as instruction-accurate emulation.
The practical rule is simple: treat each multiply opcode as a state transition over eight 48-bit accumulator lanes plus its vector output. Once tests verify both sides of that transition, errors in graphics or audio microcode become much easier to localize than when every discrepancy is blamed on the renderer.
Related:
- Nintendo 64 RDP: From Microcode Display Lists to Filtered Pixels
- Nintendo 64 DMA Cache Coherency: Keeping R4300, RSP, and RDRAM in Agreement
Sources: