PlayStation GTE: Fixed-Point Transforms, Saturation, and FLAG Semantics
Trace PS1 GTE inputs, fixed-point transforms, command timing, saturation, and FLAG summaries without replacing guest arithmetic with host floats.
The PlayStation Geometry Transformation Engine is a fixed-function coprocessor attached to the R3000A execution model. It is not the GPU, and it is not a general floating-point unit. Game code loads vectors, matrices, translation values, and color parameters into COP2 registers, starts a named command, then reads results and status. An emulator that turns those operations into a host graphics API call can produce a similar picture while returning different coordinates, flags, or values to software.
The observable contract includes more than the final screen point. It includes the register file, signed widths, fixed-point scaling, command-specific accumulator behavior, saturation, the FLAG register, and the time at which results become available. Those details matter to games that chain GTE commands, poll busy state, use results in collision or lighting code, or depend on overflow behavior. The best debugging unit is therefore a command with captured input registers, output registers, flags, and guest-cycle timestamps.
COP2 is a programmed arithmetic device
The CPU accesses GTE data and control registers through coprocessor transfer instructions and launches operations through COP2 commands. The engine has no ordinary memory-mapped register window. This changes how an emulator should connect it to the CPU: loads and stores participate in the processor’s instruction pipeline, while a command has its own execution latency and result-visibility rules.
The data-register bank carries values such as input vectors, projected screen coordinates, depth values, color inputs, and intermediate results. The control-register bank carries matrices, translation vectors, projection parameters, and the command FLAG state. Some logical registers expose overlapping views of packed coordinates or color channels. Treating the bank as an untyped array of 32-bit integers loses the signed 16-bit halves and sign-extension rules that software sees.
The rotation, light, and light-color matrix elements are signed fixed-point values. The hardware reference documents their packed representation, while the translation vectors and several intermediate values have different widths and scaling. Decode each register according to its own role. A useful trace prints both the raw word and the interpreted fields, for example a packed pair of signed matrix entries, rather than displaying only a decimal integer.
Fixed-point arithmetic is part of the API
GTE commands combine matrix multiplication, translation, color multiplication, and shifts using fixed-width arithmetic. The command word includes options that affect the shift and limiting behavior. The shift option commonly selects whether the result is retained at a wider intermediate scale or shifted to the architectural output scale. The limiting option changes how negative intermediate values are clamped for selected IR outputs. These are not renderer quality settings; they are guest-visible command semantics.
The intermediate MAC registers are wider than the ordinary 32-bit data registers. A correct implementation should compute a command’s products and sums at the specified width, record overflow conditions, then place the architectural result in the destination registers using the command’s exact saturation and shift rules. If host signed overflow happens first, or a generic 64-bit expression is shifted at the wrong stage, the final value can be off by one or wrap in a way no software instruction could observe on the console.
Use explicit integer types for each stage. Keep register packing separate from arithmetic, and write unit tests around positive and negative extremes, zero, half-unit fixed-point values, and values just on either side of a saturation boundary. Comparing only random mid-range inputs is weak: most implementations agree when no intermediate range is stressed.
Saturation and flags explain divergence
The FLAG register records command conditions such as component saturation, depth-cue or screen-coordinate limits, and arithmetic overflow. It also contains a summary indication derived from selected error bits. Software can inspect this register and use it to diagnose or branch around a result. A renderer that clamps a coordinate for safety but never updates FLAG has changed the program’s behavior even if the rendered triangle appears acceptable.
Saturation must occur at the documented point. For example, a transformed component may first be accumulated at an intermediate width, then shifted and clamped into an IR register. A screen coordinate may be calculated from projection parameters and a depth quotient before being limited to its output range. Clamping the original vector, clamping every multiply, or applying a final host-coordinate clamp is not equivalent. Preserve intermediate values in traces so a discrepancy can be localized to multiplication, accumulator overflow, shift, or output clamp.
The perspective divide has its own defined range and approximation behavior. When the denominator is too small relative to the projection scale, the quotient saturates and sets a flag rather than producing an arbitrary host infinity. Do not delegate the operation to float division: floating-point rounding, infinities, and NaNs do not reproduce the GTE’s integer result or overflow status. The PSX-SPX reference describes the divide range and its hardware-oriented reciprocal approximation; DuckStation’s implementation is useful as an independent code-level cross-check.
Command families share state, not one formula
Projection commands use rotation, translation, projection distance, screen offsets, and depth cue parameters. Lighting commands combine normal vectors with light and color matrices, then pass values through IR and RGB output limits. General-purpose commands can add vectors, interpolate color, or transform matrices. A common implementation mistake is to create one convenient matrix helper and use it for every command while overlooking command-specific accumulator updates, FIFOs, and flag rules.
The screen-coordinate and depth FIFOs are also architectural state. A command that emits several vertices can shift prior values through the queues and expose the newest results in specific register positions. Preserve that sequence even if the renderer consumes only the newest screen point. Software can read the queue registers, and subsequent commands can use them. A saved state taken between two commands must serialize the queue contents rather than regenerate them from the current frame.
Use command-specific functions over a shared, explicit register model. Each function should identify source registers, output registers, affected MAC and IR values, queue pushes, FLAG bits, and documented latency. This makes a command reviewable against a reference table and allows differential testing against an established emulator without turning that emulator into ground truth.
Timing and CPU hazards are observable
The GTE does not make every result available immediately at the instruction that starts a command. Command execution has documented latency, and CPU instructions that read a GTE register can encounter load delays or a busy condition. Games may schedule useful CPU work between command issue and result consumption. An emulator that completes every operation synchronously can pass a screenshot test but fail code that polls or reads at a specific instruction boundary.
Model latency in the same guest-time scheduler used by the CPU. When a command starts, snapshot the required inputs and schedule its completion. If hardware permits overlapping or pipelined operations for the relevant command sequence, follow that rule explicitly; do not infer it from a single game’s instruction stream. At completion, publish all related results and flags consistently. If a save state is captured while the command is active, store the command identity, remaining time, captured inputs, and pending outputs.
Instruction hazards need tests distinct from command arithmetic. Test the COP2 enable check, register transfer delays, command busy polling, reads around completion, and interruption or state restoration while work is pending. For every test, preserve the instruction sequence and cycle count alongside the expected register snapshots. This prevents an apparent arithmetic bug from actually being a missing CPU/GTE synchronization rule.
A small arithmetic boundary test
This helper is not a GTE emulator. It demonstrates how to make signed register decoding explicit before adding command-specific fixed-point rules. Tests should compare the returned value with the raw register and exercise the sign bit in both halves.
def sign_extend(value, bits):
if bits <= 0:
raise ValueError("bit width must be positive")
mask = (1 << bits) - 1
value &= mask
sign = 1 << (bits - 1)
return value - (1 << bits) if value & sign else value
def unpack_matrix_pair(word):
low = sign_extend(word, 16)
high = sign_extend(word >> 16, 16)
return low, high
assert unpack_matrix_pair(0x0002FFFE) == (-2, 2)
assert sign_extend(0x7FFF, 16) == 32767
assert sign_extend(0x8000, 16) == -32768
The real command path must then apply its documented fractional scale, accumulator width, shift, and saturation. Keeping decoding separate ensures that an endian or packing mistake is not misdiagnosed as matrix math.
Build a trace that names the first wrong value
For a failing scene, record the raw COP2 command word, data and control registers before issue, command start and completion cycles, MAC outputs, IR values, screen/depth queue contents, FLAG, and the first CPU instruction that consumes each result. Record game revision and region because the test program may take a different path. A screenshot can tell you where a polygon ended up; it cannot tell you whether the cause was a packed register, a missing saturation flag, a divide boundary, or an instruction delay.
Differential tests should use the same guest register inputs and command sequence in the target implementation and a reference emulator. Compare each architecturally visible result before comparing final geometry. If outputs differ, reduce the case to one command with a hand-calculated ordinary-range example and one boundary case. Mark implementation-derived behavior separately from values documented by a hardware reference, particularly for obscure timing and silicon-specific quirks.
Acceptance criteria
A dependable GTE implementation keeps data and control register semantics distinct, performs arithmetic at documented widths, applies shifts and saturation in command-specific order, updates FLAG consistently, maintains result queues, and publishes results on the right guest-time boundary. It can replay a captured command from raw inputs and explain every output word. Only after that model is stable should the GPU consume projected coordinates and colors.
Related:
- PlayStation GPU Drawing Environment: Clip Areas, Offsets, and Mask Bits
- Nintendo 64 RSP Vector Arithmetic: Fixed Point, Accumulators, and Saturation
Sources: