Skip to content
RetrogamingDeep Dive Published Updated 9 min readViews unavailable

Amiga Blitter Minterms and Overlapping Copies: DMA Semantics That Matter

Model Amiga blitter DMA precisely: derive raw minterms, preserve overlapping copies with descending mode, honor word masks and modulos, and test completion safely.

The Amiga blitter is not merely a fast memory-copy instruction. In block mode it reads up to three source streams, applies a programmable Boolean function to each data bit, and writes a fourth stream. Shifts, first- and last-word masks, row modulos, ascending or descending traversal, and bus arbitration all affect the result. An emulator that copies rectangles with a host memcpy can render simple screens while failing masked sprites, scroll operations, fill mode, or code that observes the blitter’s busy state.

The reliable way to model it is as a DMA state machine whose inputs and pointer updates are visible in hardware order. Start with the raw register-level contract, then separately implement the higher-level graphics.library helpers that translate bitmap operations into that contract.

Raw channel model and minterm truth table

In a raw block blit, channels A, B, and C are inputs to the minterm logic; channel D receives its output. For each bit position in the fetched words, A, B, and C form a three-bit truth-table index. The eight-bit logic field in BLTCON0 supplies the output for all eight possible input combinations. Thus the field describes a complete Boolean function, not a transparency mode or a preselected “copy” opcode.

Input bits Output selector
A B C = 000 LF bit 0
A B C = 001 LF bit 1
A B C = 010 LF bit 2
A B C = 011 LF bit 3
A B C = 100 LF bit 4
A B C = 101 LF bit 5
A B C = 110 LF bit 6
A B C = 111 LF bit 7

For example, LF = 0xCC selects the combinations where B is 1, so D copies B. LF = 0xCA implements D = (A AND B) OR (NOT A AND C): use A as a mask, B as source image, and C as the prior destination/background. Each source bit either chooses the sprite bit or preserves the background bit.

index = (A_bit << 2) | (B_bit << 1) | C_bit
D_bit = (LF >> index) & 1

That formula is a compact conformance oracle for the raw minterm unit. Apply it independently to every bit position in the words after the channel’s shift and mask stages. In a planar display, a bit position can correspond to a pixel in one bitplane; the minterm itself does not know about RGB colors or sprite alpha.

There is a frequent API trap here. The AmigaOS graphics.library documentation describes minterms in the context of library-level channel roles and rectangle masking. Values used by a graphics.library call should not automatically be copied into the raw hardware BLTCON0 field. The library can configure A as an “inside the rectangle” mask and assign B/C differently from a hand-written raw blit. Keep the library API’s truth table and the chip register’s A/B/C truth table explicitly separate in code and documentation.

Shifts, masks, and row modulos are part of the operation

The blitter processes words, not arbitrary host pixels. Channel A and B can use barrel shifts to align source data to destination word boundaries. Channel A also has first- and last-word masks so an operation can preserve bits outside a partial-width rectangle at its edges. A shift changes which source bits are aligned with the word being processed; masks determine which edge bits participate. Applying the minterm to unshifted words and clipping the final rectangle on the CPU is not equivalent.

For a bitmap whose rows are wider than the blit rectangle, each channel has a modulo: the signed byte adjustment made after the row’s words have been traversed to reach the next row. Source A/B/C and destination D can have different strides. Do not assume every channel uses the same pitch, and do not calculate the next row solely as base + width_words * 2. An emulator should maintain separate current pointers and modulo values for each enabled channel.

An OCS/ECS blit begins with its pointer, control, mask, modulo, data, and size state. BLTSIZE is written last because that write starts the operation. The familiar OCS/ECS size field has a limited word width and line-height encoding; ECS and AGA provide extended size registers, so the size model must be selected by chipset rather than stretched to match a modern bitmap dimension.

Choose direction before moving overlapping memory

When source and destination rectangles overlap, processing order determines whether the blitter overwrites source data that has not yet been read. The manual’s example of moving an image toward higher memory addresses uses descending mode. In descending mode, word pointers decrement rather than increment, row modulos are subtracted, and the pointers must begin at the last word of the block. For a standard copy with no shift or masking, changing the starting pointers and traversal direction is the essential difference; the data should be consumed before a later write can destroy it.

Use ascending traversal when the destination lies below the source in address order, and descending when it lies above, after taking row layout and actual overlap into account. Do not choose direction from the visual idea of “up” or “down” alone: planar buffers, signed modulos, and rectangular strides determine the actual address ranges. If the source and destination do not overlap, either direction may produce the same bytes, but still test pointer initialization and modulo handling separately.

Descending mode also changes shift direction and which word mask is applied first versus last. A simplistic implementation that reverses only the pointer increment can pass a same-alignment copy and fail as soon as the source shift or partial-word masks are enabled. Test a two-row buffer where destination overlaps source by one word, then repeat with a nonzero shift and masked edge words.

Fill mode and line mode are not ordinary copies

The BLTCON1 control register selects behaviors that modify the normal data path. Area-fill mode uses a fill carry as set bits are encountered across a row; the manual requires descending traversal for the described fill operation. Inclusive and exclusive fill differ at boundary pixels. This means “apply minterm, then independently fill the destination” is not a sufficient model unless it preserves the hardware’s word traversal and carry state.

Line mode is another specialized operation with its own control fields and size convention. In the hardware reference, BLTSIZE starts line drawing with width set to the line-mode value and height controlling the line length. Do not decode every BLTSIZE as a rectangle width and height without checking BLTCON1 mode. The same register write starts the operation, but its fields have a different interpretation.

Busy state, register ownership, and bus contention

Writing BLTSIZE launches asynchronous work. The CPU may continue executing while blitter DMA is pending or underway, and other custom-chip DMA users compete for memory slots. Code that reads the destination or reprograms the blitter before the prior operation has completed can observe partial output or corrupt the operation’s register state.

The Amiga Hardware Reference Manual documents a blitter-done/busy status flag and warns about chipset differences in when that flag becomes visible relative to the start write. For OS-level software, the ROM kernel’s WaitBlit() is the appropriate synchronization interface when the operation must be complete; direct hardware code needs a revision-aware wait strategy and exclusive or arbitrated ownership of the hardware. A CPU that is fast enough to reach the next operation before the blit finishes can expose races that were hidden on slower configurations.

Timing is also observable indirectly. Blitter, Copper, display, sprite, audio, and disk DMA share system resources under priorities and time-slot rules. A functional renderer can calculate the correct final pixels but still be inaccurate if software depends on CPU stalls, blitter completion timing, or Copper/Blitter interaction. Keep final-data correctness and DMA scheduling as separate milestones in an emulator; only claim cycle accuracy after both have tests.

Emulator test matrix

Start by validating the minterm lookup for all eight A/B/C combinations, then apply it to words containing alternating patterns such as 0xAAAA and 0x5555. Compare 0xCC source-copy and 0xCA mask-composite operations against a simple software reference. Test each channel’s enable bit and confirm disabled source channels use the documented data-register behavior for the selected mode rather than reading arbitrary host memory.

Next vary shifts, A first/last masks, independent modulos, and odd rectangle widths. Use guard words around each source and destination and assert that untouched edge bits and neighboring rows remain unchanged. For overlap, create a small known buffer, copy a rectangle one row toward higher addresses in descending mode, and compare the result with memmove; then copy toward lower addresses in ascending mode. Add a shifted case where the destination crosses a word boundary.

Test fill and line operations as distinct state machines. For fill mode, include rows with an even and odd number of set bits and verify the documented inclusive/exclusive boundary behavior. For line mode, test each octant and error-sign transition against reference vectors. Verify that writing BLTSIZE starts work only after setup state is complete, that busy clears after completion, and that CPU reads are synchronized before the result is consumed.

Finally, run the same trace on the chipset revisions the emulator claims to support or compare against a validated hardware test suite. Record chipset class, DMA enable state, blitter priority, display mode, pointers, modulos, controls, and memory before and after each operation. These inputs turn a graphical mismatch from a subjective screenshot into a reproducible DMA and logic trace.

The Amiga blitter is best understood as a four-channel, word-oriented DMA engine with a programmable Boolean datapath and direction-sensitive traversal. Preserve the chip-level minterm truth table, treat each channel’s alignment and stride independently, and model fill, line, and completion behavior as explicit state. That discipline prevents the common mistake of turning a precise hardware contract into a rectangle copy that only works for the easiest case.

Related:

Sources:

Comments