Skip to content
RetrogamingDeep Dive Published Updated 8 min readViews unavailable

Amiga Copper and Bitplane DMA: Beam-Synchronized Display Lists

Trace Amiga Copper MOVE, WAIT, and SKIP instructions alongside bitplane DMA, Chip RAM fetches, bus slots, and safe display-list validation.

Classic Amiga display effects are often credited to the Copper, but the Copper is only one participant in a shared custom-chip memory system. Bitplane DMA fetches the image data that Denise or Lisa turns into pixels. The Copper fetches and executes a compact instruction list, waiting for beam positions and writing selected chipset registers. The Blitter, audio, sprites, disk, and CPU also compete for memory slots. A raster effect that is mathematically correct can still appear late if the Copper cannot fetch its next instruction when expected.

The Commodore-Amiga Hardware Reference Manual describes the Copper as a DMA coprocessor with MOVE, WAIT, and SKIP instruction forms. It also explains Chip RAM requirements and the interaction of display DMA with bus slots. WinUAE is a useful implementation reference for behavior across the OCS/ECS/AGA families. Keep the manual as the hardware contract and the emulator source as an executable interpretation, particularly where later chipsets or compatibility choices differ.

The three-instruction machine

Every Copper instruction occupies two 16-bit words in memory. MOVE writes a value to an allowed custom-chip register. WAIT suspends instruction execution until the beam comparator reaches a selected horizontal/vertical position, subject to masks and enable bits. SKIP conditionally skips the next instruction when the beam has already reached a comparison point. Copper lists use these operations to load colors, change scroll values, set bitplane pointers, trigger blitter work, or establish split-screen state.

This is a state machine, not a script that the emulator should execute all at once at frame start. The Copper fetches its next words through DMA, compares against the evolving beam position, and may be delayed by occupied memory slots. A WAIT does not consume an unbounded sequence of host time while the emulated machine pauses; it yields bus slots until its condition becomes true. Then instruction-fetch timing determines when the following MOVE takes effect.

Because a Copper instruction is two words, a list pointer that is misaligned or points outside accessible Chip RAM is not a valid list. The Copper’s program pointer and instruction state are separate from the CPU program counter. A host implementation should represent the current pointer, first word, decoded operation, pending WAIT condition, DMA availability, and beam comparator state.

Bitplanes are fetched by DMA, not read as a framebuffer

The Amiga graphics chipset constructs pixels from one or more bitplanes. For each display region, bitplane DMA fetches words from Chip RAM using the configured bitplane pointers and row layout. The fetched planes feed pixel-value generation and color-register selection. In common low-resolution modes, each pixel’s bit from each plane contributes to a color index; additional planes increase the number of available indexes. The display is therefore assembled from parallel one-bit streams rather than a conventional packed-color framebuffer.

At the end of a displayed row, modulo values adjust the bitplane pointers so the next fetch can continue with a wider source row, a cropped region, or another layout. The Copper can change pointer or scroll registers at beam positions, enabling split screens and scrolling changes. The effective image depends on when the pointer update arrives relative to bitplane fetch, not only on the final pointer value saved at end of line.

Chip memory bandwidth is finite. Display modes that request more bitplane data consume more DMA opportunities, leaving fewer slots for the CPU, Copper, or Blitter. If the Copper waits for a beam position while a dense bitplane display occupies memory slots, the write may occur later than the nominal comparator coordinate. This is why a precise raster log needs both beam position and actual bus/fetch slot.

WAIT compares the beam; it does not schedule a host callback

A WAIT instruction encodes a beam comparison plus mask/enable bits. The comparator tests the current raster position against the instruction. The Copper remains halted until the condition becomes true, then continues its instruction stream when it can fetch again. Horizontal resolution and counter encodings have chipset-specific details; software should construct waits through the correct assembler macros rather than assume that a nominal pixel coordinate maps one-to-one to a bitfield.

Use symbolic assembler definitions for production lists. A short conceptual sequence might look like:

        dc.w    $0180,$0F00      ; MOVE: COLOR00 = white
        ; WAIT for a later beam position using the assembler's Copper macro
        ; MOVE: change color or scroll state for the next raster region
        dc.w    $FFFF,$FFFE      ; conventional end-of-list marker

The example demonstrates list structure, not a complete video mode setup or portable beam wait. Register addresses, permissible destinations, wait encodings, display fetch windows, and chipset revision must be checked in the Hardware Reference Manual and assembler include files. Never transplant raw WAIT constants between PAL/NTSC or custom screen modes without verifying the beam-counter assumptions.

SKIP is useful for one-conditional-instruction branching within a list, but it is not a general-purpose CPU branch. It skips the next instruction only when its comparison succeeds. If code inserts a variable number of Copper words between the SKIP and intended target, the list does something different. Validate list offsets after macro expansion, not from source indentation.

Bus priority and display deadlines

The Copper and bitplanes share DMA scheduling. Bitplane fetches have priority needed to keep the display fed, and the manual describes how other DMA channels and the CPU use remaining slots. A Copper list that attempts many MOVEs in a line may not execute all of them before the next line. The blitter-nasty setting (BLTPRI) changes the blitter’s competition with the 68000 for otherwise available Chip-RAM cycles; it does not promote the blitter ahead of the Copper or display DMA. Attribute a late Copper operation to its actual fetch/slot constraints, not to BLTPRI alone.

This makes Copper throughput a bus-budget problem. Count instruction fetch slots as well as visible color writes. A WAIT itself consumes instruction and comparison resources; a MOVE needs fetch cycles and a register write. Add display DMA load, sprites, and blitter activity. If a list has a visible tear only when a large blit runs, compare the bus schedule before changing the wait position.

The operating system’s ownership matters too. Applications that take direct control of custom-chip registers must coordinate with system display state, memory allocation, Copper pointers, and interrupt lifecycle. The hardware manual’s examples assume the programmer has control of the hardware; a normal OS application should use documented graphics APIs unless it intentionally owns the display. Emulator behavior should not turn an illegal or conflicting list into an implicitly safe one.

Emulator architecture

Model bitplane fetches and Copper execution on a shared beam/bus timeline. Bitplane DMA should request slots at the required fetch positions and load shifter state. The Copper should request instruction words from Chip RAM, evaluate WAIT/SKIP against the current beam counter, and write chipset registers only when its operation reaches execution. CPU and Blitter accesses should compete according to chipset rules. This may be implemented with an event queue, slot scheduler, or cycle-level loop; the key invariant is that one deterministic ordering governs all masters.

Avoid a renderer that samples all registers at the beginning of a frame and emits a complete bitmap. That loses the precise mid-frame register writes which the Copper exists to create. A line renderer can be sufficient if it processes all relevant events and fetch boundaries within each line. If it snapshots state at line start, it must support partial-line changes rather than silently postponing them.

Keep OCS/ECS/AGA variations explicit. The Copper instruction model has compatibility nuances, and chipset display modes alter fetch requirements and memory bandwidth. Do not cite a timing value for one chipset family as universal. Store the selected chipset and video standard in traces and test metadata.

Validation strategy

Create small diagnostic lists that isolate each operation:

  1. MOVE a color register before the visible display and verify the resulting color index.
  2. WAIT to a known line/horizontal point, then change one register and capture where the pixel stream changes.
  3. SKIP exactly one following MOVE under true and false beam comparisons.
  4. Change bitplane pointers or modulo at a line boundary and verify word fetch addresses.
  5. Increase plane count and add sprites to measure which Copper instructions are delayed by DMA usage.
  6. Run the same list with Blitter activity and its priority modes, observing instruction-fetch stalls.
  7. Compare PAL/NTSC and selected chipset configurations rather than changing only the host window size.
  8. Save/load while Copper is halted in WAIT and while a two-word instruction fetch is partial.

Trace the beam counter, current Copper pointer, decoded instruction, Chip RAM bus owner, bitplane fetch address/plane, custom register writes, and any Blitter busy state. Pair the trace with a native-resolution output capture. When testing with WinUAE, record version and chipset configuration so that the emulator is a specified reference, not a vague authority.

Acceptance criteria

A robust Copper/bitplane implementation has beam-synchronized register writes, explicit Chip RAM access, shared DMA scheduling, and per-chipset timing configuration. Tests prove MOVE/WAIT/SKIP behavior independently, and raster effects remain stable when CPU or Blitter traffic changes only to the extent allowed by the documented bus model. Final output should be derived from the same event sequence used by the Copper and bitplane fetchers.

The Copper demonstrates that old display hardware could expose powerful composition through a tiny instruction set, but the effect was never “free.” Every color change was paid for in instruction fetches, bus slots, and beam time. Emulation becomes both more accurate and more understandable when those costs remain visible.

Related:

Sources:

Comments