Skip to content
RetrogamingDeep Dive Published Updated 10 min readViews unavailable

Nintendo 64 DMA Cache Coherency: Keeping R4300, RSP, and RDRAM in Agreement

A production-minded guide to Nintendo 64 DMA cache maintenance, writeback and invalidation order, cache-line alignment, and emulator regression tests.

The Nintendo 64 lets the R4300 CPU, Reality Co-Processor (RCP), and I/O engines use the same RDRAM, but shared physical memory does not mean coherent CPU caches. The R4300 has a write-back data cache: a CPU store may update a cached line without immediately updating RDRAM. RSP tasks and peripheral DMA read or write RDRAM outside that cached CPU path. If software does not explicitly transfer ownership at the right time, a device can consume stale input, the CPU can read stale output, or a later cache writeback can overwrite newer DMA data.

This is not merely a hardware-programming concern. An emulator that models the R4300 cache, memory aliases, RSP/RDP access, or DMA-visible state must preserve the same ordering. A renderer may appear correct until it reads a freshly written display list, a PI transfer completes into a buffer, or an older dirty cache line is evicted over the device’s result.

One physical address space, separate visibility paths

The Nintendo 64 Programming Manual describes the main memory as shared by the R4300, RSP microcode, RDP, and I/O interfaces. The CPU commonly accesses data through cached KSEG0 addresses, while the RSP and RDP use physical addresses and DMA mechanisms. The R4300’s 8 KiB data cache uses a write-back policy, so a CPU write can remain newer than the corresponding bytes in RDRAM.

The important boundary is not “CPU versus graphics” in the abstract. It is whether the next consumer reads the CPU’s cached copy or the backing memory, and whether the producer of new bytes bypasses the cache. Before a device reads a CPU-produced buffer, software must make the dirty CPU data visible in RDRAM. Before the CPU consumes device-produced bytes, it must prevent an old cached copy from being used or written back over the transfer.

The KSEG0/KSEG1 aliases do not remove this ownership problem. An uncached CPU access can avoid allocating a new data-cache line for that particular access, but it does not automatically write back an older dirty KSEG0 line that still refers to the same physical memory. Mixing cached and uncached aliases while a line is dirty is therefore not a substitute for an explicit, ordered handoff.

CPU-produced data: write back before the device reads

For an outbound transfer, the CPU first finishes populating the buffer and any descriptors the device will inspect. It then writes back the relevant cache lines before starting the RSP task, RDP operation, or I/O DMA that reads them. The Nintendo manual specifically calls out flushing the dynamic portion of an RSP task list before starting the RSP, and gives the same rule for buffers read by DMA.

The sequence is conceptually:

CPU writes payload and descriptors
write back the full cache-line range the device will read
submit or start the device operation
wait for the device's documented completion point before reusing the buffer

The writeback must cover every device-visible byte, including fields filled by a producer thread or allocator after an earlier flush. A flush performed before the final store is not sufficient. Keep task headers, command lists, descriptors, and payloads in well-defined regions so their ownership is easy to audit. Where software has several producers, establish CPU-side synchronization first; cache maintenance does not make concurrent writes safe.

In libdragon, data_cache_hit_writeback() provides a range writeback operation. The exact API and alignment contract depend on the SDK, so use the target SDK’s documentation rather than copying cache helper names from another toolchain. Do not flush the entire cache by habit when a smaller, correct range is known; broad operations can hide ownership bugs and add avoidable work.

Device-produced data: invalidate before the DMA writes

For an inbound transfer, a CPU cache line may contain an older copy of the destination, possibly with dirty bytes. The safe handoff is to invalidate the destination’s cache lines before the device starts writing RDRAM, then avoid CPU access to that range until the transfer has completed. After completion, the CPU reads the new bytes and refills its cache from memory.

The order matters. If stale or dirty destination lines survive while DMA writes new data, a later CPU writeback can restore old bytes over the device’s result. Invalidating only after completion can be too late if the line remains dirty and is written back between the DMA and that invalidate. The Nintendo manual explicitly requires invalidation before RCP or I/O DMA produces the data and warns that a later writeback can clobber the new contents.

Treat the completion notification and cache operation as separate parts of the protocol. A PI-manager message or another device-specific completion event establishes that the transfer finished; it does not itself invalidate the R4300 cache. Conversely, invalidating a buffer does not mean an asynchronous DMA has completed. The consumer should wait for the documented completion mechanism and must not read the region while the device owns it.

Cache-line boundaries are an ownership boundary

The N64 Programming Manual identifies 16-byte data-cache lines and recommends aligning I/O buffers to avoid cache-line tearing. The reason is subtle: a cache operation acts on lines, not on an abstract C object. If a DMA destination shares a line with an unrelated dirty variable, invalidating the line can discard that variable’s CPU-only update. If an unrelated CPU store later writes back a line containing the DMA destination, it can overwrite device-produced bytes.

Use dedicated, cache-line-aligned buffers for DMA, and round their allocated size up so the first and last line contain no unrelated live data. Keep control structures separate from payload ranges when their ownership changes at different times. If a partial-line transfer is unavoidable, follow the SDK’s documented writeback/invalidate procedure and ensure the CPU does not touch the whole affected line until the operation finishes. A helper that preserves neighboring dirty bytes before invalidation is not permission to share a line concurrently.

Libdragon documents that data_cache_hit_invalidate() requires a 16-byte-aligned address and a length that is an exact multiple of 16. It also warns that a non-aligned range may expand to a larger region and discard unrelated CPU-written data. Its data_cache_hit_writeback_invalidate() operation first writes back and then invalidates affected lines, which can preserve dirty neighboring data in cases where that behavior is intended. These details are library-specific; inspect the actual SDK contract before using equivalent-looking functions elsewhere.

A small ownership table prevents ordering mistakes

Before implementing a transfer, write down the producer, consumer, direction, and point at which ownership changes. This catches the common mistake of treating “DMA finished” and “cache is coherent” as synonyms.

Transfer Before device starts While device owns the buffer Before CPU consumes or reuses it
CPU to RSP/RDP/I/O Finish CPU writes, then write back the device-readable lines Do not mutate the submitted bytes Wait for completion before repurposing data still in use
RSP/RDP/I/O to CPU Invalidate the destination lines before the device writes; do not retain CPU-only data in shared lines Do not read or write the destination from the CPU Wait for completion, then read the updated bytes
Bidirectional or reused buffer Define a separate handoff for each direction and phase Never let two owners modify the same line concurrently Complete the previous phase’s cache operation before the next owner begins

This table is deliberately conservative. A device may have its own queueing and ordering rules, and a CPU-side message queue can establish a completion point without changing cache state. Record both mechanisms in the code review and test plan.

Alignment requirements differ by device and operation

Cache-line alignment and DMA alignment are not interchangeable. The manual describes 16-byte cache lines and separately lists 8-byte alignment for most DMA operations, with special constraints for cartridge/PI transfers and specific graphics buffers. Satisfying a DMA engine’s address and length restrictions does not automatically make cache maintenance safe. Conversely, aligning a cache range does not guarantee a valid device transfer.

For each engine, verify the source-address, destination-address, length, and cache-operation constraints independently. Use the OS or SDK transfer API when it owns device arbitration or provides a documented completion queue. Do not infer a peripheral’s behavior from another channel just because both are called DMA. A PI transfer, RSP DMEM transfer, and RDP framebuffer operation have different interfaces and may have different alignment and lifecycle requirements.

Emulator consequences and regression strategy

An emulator does not have to imitate host cache instructions literally, but it must reproduce their guest-visible results. If its RDRAM array is always immediately updated by CPU stores, while the emulated N64 exposes a write-back cache, a guest that forgot a writeback may work accidentally. Conversely, an emulator that models memory as separate device and CPU copies must implement explicit flush and invalidate effects at the correct line granularity. Both approaches should be tested against the observable contract, not against assumptions about host cache hardware.

Build focused regression cases around ownership transitions:

  • Write a known pattern into a CPU buffer, submit it to an RSP or I/O consumer, and verify the device sees the latest bytes only after the guest performs the required writeback.
  • Let a device write a buffer after the CPU has cached its previous contents; test that correct pre-transfer invalidation exposes the new pattern and that a stale dirty line cannot overwrite it later.
  • Place an unrelated dirty word adjacent to a DMA buffer in the same 16-byte line to expose false sharing and range-rounding mistakes.
  • Test an aligned full-line transfer, a partial-line transfer, zero length if the API permits it, and the largest supported transfer size.
  • Exercise asynchronous completion: verify the CPU cannot consume the destination before the device completion event, and that a completion message alone does not stand in for cache maintenance.
  • Repeat through cached and uncached aliases, and verify that aliasing does not silently repair an omitted writeback.
  • Save and restore around an in-flight operation only if the emulator supports that state; include the buffer bytes, pending transfer, and ownership phase in the snapshot model.

When debugging a visual artifact, capture the physical RDRAM bytes, the CPU’s cached line state, the transfer address and length, the cache operation range, and the device’s completion event. A screenshot cannot show whether a display list was stale in memory or whether the RDP read an old descriptor. Logging these boundaries makes a rare timing-dependent defect reproducible.

Practical review checklist

For each shared-memory object, answer five questions: who writes it first, who reads it next, whether that consumer uses the CPU cache, which cache lines contain the object, and what event ends the current owner’s access. Then verify cache maintenance against the exact SDK version, align and size the allocation accordingly, and test both normal completion and error paths. Reuse of a buffer is a new ownership transition, not an automatic consequence of one earlier flush.

The durable rule is simple: before a device reads CPU-written data, make it visible in RDRAM; before a device writes data that the CPU will read, invalidate the relevant cache lines before the write and wait for completion. Respect full cache-line ownership, not just C-level object boundaries. This discipline prevents stale command lists, discarded neighboring variables, and DMA results that mysteriously revert after a later cache eviction.

Related:

Sources:

Comments