Game Boy Advance WAITCNT: Game Pak Timing, Sequential Access, and Prefetch
Decode GBA WAITCNT timing for Game Pak regions, distinguish sequential from nonsequential bus cycles, and test the prefetch pipeline safely.
The Game Boy Advance CPU can execute code from cartridge ROM, but the cartridge bus is slower than internal work RAM. WAITCNT controls access timing for three Game Pak regions and also enables a prefetch buffer. Correctly emulating this behavior requires distinguishing first/nonsequential accesses from subsequent sequential accesses, tracking the 16-bit external data path, and treating prefetch as a pipeline with invalidation conditions rather than as a universal cache.
This register has performance consequences and is also part of guest-visible timing. A game may set shorter waits for a compatible cartridge or use prefetch for code-heavy loops. If an emulator models every ROM read as one constant delay, CPU timing, DMA overlap, and code execution measurements can drift.
WAITCNT contains three independent timing regions
WAITCNT is at 0x04000204. Its fields configure SRAM/FRAM access timing, first and second access timing for wait-state regions 0, 1, and 2, the PHI output clock, and the Game Pak prefetch enable. The three Game Pak regions are mapped at 0x08000000, 0x0A000000, and 0x0C000000, allowing different timings for devices behind those ranges. The settings describe wait cycles; total bus time also includes the base access cycle.
“First” and “second” are often described as nonsequential and sequential accesses. Sequential timing applies only when the bus protocol and address progression qualify. A branch, unrelated bus access, or region/boundary transition can break a sequential run. The GBA’s 16-bit Game Pak bus means a 32-bit CPU read is assembled from two halfword transfers; the second half follows a sequential timing rule even though the overall CPU operation began as one word load. The model should represent these bus beats, not charge one scalar wait to the instruction.
The bus forces a nonsequential access at each 128 KiB Game Pak ROM boundary. That boundary rule can make a long block transfer have a timing discontinuity even when addresses are contiguous. Tests that only read from the middle of a ROM image will miss it. Include accesses straddling the boundary and record each halfword’s sequential classification.
Prefetch helps instruction fetches, not arbitrary reads
The external Game Pak prefetch buffer can queue upcoming instruction data during appropriate idle bus time. It is distinct from the ARM CPU’s own instruction pipeline. It is not a generic data cache: a load from ROM can consume the external bus and interfere with the stream of prefetched opcodes. Branches and changes in access pattern also alter what can be reused. Emulation should track the buffer’s contents/validity and the bus events that fill or invalidate it.
The exact queue depth and prefetch rules should come from a trusted hardware reference and be encoded with named constants. Tests should verify sequential code fetch with prefetch disabled and enabled, branch target fetch, a data load from ROM while code executes there, DMA activity, and crossing a 128 KiB boundary. Compare total cycle counts over a known instruction sequence; do not infer prefetch success just because the game runs faster.
Cartridge wait-state choices are a compatibility/performance contract. Short timings are not safe for every flash cartridge or external device. A game’s selected value is guest software state and must be honored unless the emulation target deliberately exposes a cartridge profile. Do not silently rewrite WAITCNT to a favorite fast preset as an emulator optimization.
A typical setting is an example, not a universal default
The reference notes that some manufactured cartridges commonly used a setting represented as 0x4317, with faster ROM waits and prefetch enabled. This is historical context, not a safe value for every cartridge or a recommendation to write it blindly. SRAM/FRAM and EEPROM/flash devices can require different timing from ROM, and third-party devices may not tolerate the same settings. A game or test program must select values based on its actual hardware.
The following code shows how to isolate the register write in a GBA program. It deliberately leaves the chosen value as a caller-supplied policy so hardware-specific timing is not smuggled into a universal helper:
#include <stdint.h>
#define REG_WAITCNT (*(volatile uint16_t *)0x04000204u)
static void set_gamepak_waitcnt(uint16_t value)
{
REG_WAITCNT = value;
}
The snippet is complete C syntax for a bare-metal GBA toolchain. It does not validate the value or measure timing; the caller must use the documented bit layout and cartridge requirements. On an emulator, the register write updates the emulated control bits and invalidates/reclassifies any in-flight bus timing as required by the hardware model.
Measure the bus, not just wall-clock speed
Build a small ROM fixture that performs a known sequence of nonsequential and sequential reads, word reads, instruction fetches, and branch refills. Record internal cycle counts and compare with the reference timing model. Repeat with each wait-state region and prefetch state. Use DMA separately because DMA bus ownership can affect CPU progress and prefetch fill windows.
At the emulator boundary, trace the initiating CPU instruction, address, width, region, first/second beat, sequential flag, WAITCNT snapshot, prefetch hit or miss, DMA owner, and charged cycles. This instrumentation reveals whether an apparent timing mismatch comes from address classification, prefetch, region decoding, or DMA arbitration. Keep ROM-data reads distinct from SRAM and serial EEPROM access; the cartridge backup protocols are not ordinary ROM wait states.
Save states should retain WAITCNT, prefetch queue and valid count, next fetch address, any partially completed halfword transfer, and DMA/bus ownership. A load into the middle of a pipeline must not discard a prefetched opcode or complete a half-transfer twice.
Regression matrix
Test reset defaults; each region’s slow and fast encodings; sequential chains; breaks caused by branch and data access; unaligned word access; the 128 KiB edge; prefetch enabled/disabled; and contention with DMA. Test actual game code with a verified trace after synthetic tests pass. If only wall-clock frame rate changes while emulated instruction time does not, the test is measuring host scheduling rather than WAITCNT.
WAITCNT is small, but it couples software configuration to external-bus scheduling. Explicit bus beats, sequential rules, prefetch state, and cartridge-specific limits turn it from a magic timing penalty into a testable hardware model.
One practical emulator test should deliberately defeat prefetch. Execute a ROM loop, branch to a new page, perform a data read from the cartridge, and then return to sequential instruction fetch. Instrument which halfwords are already buffered before each access and which bus cycles are used to refill them. Repeat the exact sequence with one timing bit changed and compare the cycle ledger. This test is more revealing than a benchmark that runs a single linear loop, because it distinguishes a buffer hit from the cheaper sequential wait-state setting. Keep the cartridge image immutable and record its hash so performance results can be compared across emulator builds. If a flash cart or unusual cartridge profile is modeled, put its electrical timing into a separate profile selected by explicit metadata; do not infer it from a filename or silently override a game’s register writes. These controls make WAITCNT tests reproducible while protecting compatibility with slower devices.
Instruction timing should be attributed carefully when interpreting those measurements. A load instruction includes core execution work in addition to external-bus wait, and a word read may require more than one bus phase. Measure the cycle delta around a controlled sequence and subtract a baseline sequence that runs from internal work RAM, while preserving the instruction mix. Do not conclude that a prefetch change saved a particular number of cycles from elapsed host time; host scheduling, JIT block size, and audio/video pacing all contaminate wall-clock measurements. In a dynamic recompiler, the WAITCNT model must still update emulated cycle accounting at the guest bus operation even if the host emits the whole translated block at once. Keep the timing result attached to guest instructions and device events, not the duration of the host function call.
Related:
- Game Boy Advance DMA: Channel Priority, Trigger Timing, and Sound FIFOs
- Game Boy Advance Windows and Blending: Layer Masks, Priority, and Pixel Rules
Sources: