Skip to content
RetrogamingDeep Dive Published Updated 10 min readViews unavailable

Atari 8-Bit ANTIC Display Lists: LMS, Scrolling, DLI, and DMA Steals

Understand ANTIC as a display-list processor with independent list and screen pointers, LMS fetches, scroll modifiers, display-list interrupts, and cycle stealing.

Atari 8-bit graphics are not generated by a single framebuffer engine that scans a fixed rectangle of memory. ANTIC fetches and interprets a display-list program, requests screen and character data with DMA, and coordinates its output with GTIA. The display list can mix blank lines, text modes, bitmap modes, scroll regions, address reloads, and display-list interrupts within one frame.

That flexibility comes with a concrete cost: ANTIC is a bus master. The CPU loses memory cycles when the video system fetches display-list instructions, playfield data, player/missile graphics, and DRAM refresh. An emulator that renders the correct pixels but advances the 6502 as if all cycles were available will fail software that synchronizes to the beam, relies on a DLI, or writes display data at carefully chosen times.

Two address streams, one display program

The display-list program counter and the memory-scan address are related but not interchangeable. ANTIC fetches an instruction byte from the display list. Depending on the instruction, it may fetch an additional 16-bit address and load the memory-scan pointer, which selects where playfield bytes come from. As a mode line consumes graphics data, the scan pointer advances; the display-list pointer advances through commands. A jump changes the display-list pointer, while LMS changes the playfield scan address.

This distinction makes mixed-mode screens possible. A display list can start with blank lines, present a text mode, load a different screen address for a bitmap band, enable fine scrolling for a region, and end with a wait-for-vertical-blank jump back to the top. The screen data is not required to be a single contiguous image. The instruction stream is a compact program interpreted at video rate.

Instructions use the low nibble to identify blank, jump, or a playfield mode. Modifier bits can request a display-list interrupt, load-memory-scan address, vertical scroll, or horizontal scroll. A blank-line instruction encodes a count of blank scanlines. A jump includes a 16-bit destination; the jump-and-wait variant suspends the current list until vertical blank timing permits the next frame. Mode instructions may also be followed by an LMS address when their modifier requests it.

The bytes following an instruction are conditional. A parser that always consumes three bytes for every opcode will interpret the next instruction as an address and desynchronize the whole list. Conversely, one that consumes only one byte after a mode with LMS will begin reading graphics bytes as opcodes. Decode each instruction into an explicit length and update the display-list PC only after all conditional operands have been fetched.

Line modes determine fetch demand

ANTIC’s playfield modes differ in the number of bytes fetched per mode line, the number of TV scanlines represented, and how memory bytes become character or bitmap data. Character modes fetch character codes and use a character set; later scanlines in a multi-scanline character row can reuse the character code while fetching the next glyph row. Bitmap modes fetch pixel-pattern data as the mode line is drawn. These modes therefore have different DMA schedules even when their visible widths appear similar.

Screen width settings change how much data is needed. Horizontal fine scroll can extend the fetch span beyond the nominal width to make room for pixels that are shifted into view. ANTIC’s fetch requests are not a single burst at the beginning of a line: depending on mode, screen and glyph bytes are requested at defined positions during the scan. The video shifter consumes the fetched bytes later, so fetch timing and visible-pixel timing should be represented separately.

Vertical fine scroll changes which scanlines of a display-list mode are presented and can alter how many lines ANTIC spends on a mode instruction. The VSCROL state is shared across instruction processing and must be changed at the intended line boundary, usually with a DLI routine. Software often places a carefully calculated sequence of blank instructions and mode lines around the scroll region. A renderer that applies the new vertical offset to the entire frame after the fact can produce smooth output while completely missing CPU-visible timing.

Player/missile DMA is a separate source of memory requests from playfield DMA. Its resolution and enabled objects affect additional stolen cycles and fetched object data. Memory refresh also occupies cycles independently of the visible graphics workload. Do not attribute every missing CPU cycle to the current playfield width: compute the requests made by each ANTIC subsystem and arbitrate them against the CPU bus at the correct cycle.

LMS and scroll modifiers are stateful

Load Memory Scan is an operand-bearing modifier. When present on a playfield mode instruction, ANTIC consumes a low and high address byte and reloads the screen-data pointer. This is useful not only at the first visible line but whenever a program wants to redirect the source. If the address is omitted, the instruction stream has a different interpretation and the screen pointer continues according to its current state.

Horizontal and vertical scroll bits mark the mode lines that participate in a fine-scroll region. The corresponding scroll registers supply pixel or line offset state. A scroll modifier can change fetch demand and timing, so model it where the line’s requests are scheduled, not only at the final pixel-composition stage. Changing HSCROL or VSCROL from a DLI affects later work; the already-issued DMA requests cannot be retroactively moved.

An LMS address near the edge of the 16-bit address space is a valuable test. Confirm whether the scan pointer wraps according to the hardware’s address width and memory map, and test the memory decoder independently. Do not treat a wrap as permission to read outside emulated RAM. The CPU and ANTIC share memory but can have different access arbitration and open-bus consequences; keep the address generator separate from the host array bounds check.

DLIs and WSYNC are part of the same schedule

A display-list interrupt is requested at the mode-line boundary encoded by the instruction. It is an NMI source, distinct from the operating system’s vertical-blank routine. A DLI handler commonly changes color registers, scroll values, or player/missile state for the next part of the screen. If the emulator raises the interrupt at instruction-fetch time instead of at the specified display boundary, the handler may run early and modify state used by the wrong scanline.

CPU interrupt entry takes time, and the CPU itself can be held by ANTIC DMA. Therefore a DLI’s visible effect is determined by the DLI request, CPU instruction boundary, interrupt latency, bus steals, and the handler’s write cycle. An accurate trace should record all of these, not just the scanline number at which the NMI was requested. A border-color raster bar is a useful test because it exposes a one-cycle shift that a coarse frame screenshot hides.

WSYNC provides another synchronization mechanism. Writing it causes the CPU to wait for horizontal synchronization, so code can schedule a register update at a repeatable line boundary. WSYNC is not a generic delay for a fixed host duration; it is tied to ANTIC’s horizontal timing. A regression should verify CPU cycle position before and after the wait, including when the write occurs late in a scanline.

Cycle stealing and CPU-visible behavior

Display-list instruction fetches, playfield bytes, player/missile data, and refresh cycles all consume bus time. The CPU’s effective progress is consequently content-dependent: a narrow screen, wide screen, character row’s first line, and bitmap mode can leave different numbers of cycles available. The Atari hardware reference describes deterministic cycle stealing, while the Altirra hardware reference presents mode-specific timing. Use a per-cycle request schedule or a proven batched scheduler that has equivalent observable boundaries.

The first line in a character row can demand more memory traffic than the following glyph lines because ANTIC fetches character codes and pattern data. A simplistic “one fixed CPU penalty per scanline” misses that distinction. Horizontal scrolling can increase fetch demand, and object DMA can add it again. These are exactly the workloads that make software timing loops and raster effects behave differently across modes.

CPU writes to screen memory, display-list bytes, or ANTIC control registers can race a fetch. The emulator should define whether the write occurs before or after a same-cycle ANTIC request and test that boundary. Dynamic display-list modifications are valid software behavior; caching a decoded list for an entire frame is safe only if the cache invalidates at the correct bus-visible point. A display-list page can also have address and alignment constraints that should be validated from the hardware reference rather than inferred from a single program.

A safe instruction-stream decoder fixture

This Python example parses the variable-length display-list structure without trying to render graphics or emulate cycle stealing. It validates bounds before consuming operands and exposes the difference between a jump destination and an LMS screen address. Production code should retain the byte address and fetch cycle for each instruction so writes and DMA races can be traced.

def decode_display_list(data, start=0):
    pc = start
    decoded = []
    while pc < len(data):
        instruction = data[pc]
        pc += 1
        opcode = instruction & 0x0F
        item = {"instruction": instruction, "opcode": opcode}

        if opcode == 0:  # blank lines
            item["kind"] = "blank"
            item["dli"] = bool(instruction & 0x80)
            item["lines"] = ((instruction >> 4) & 0x07) + 1
        elif opcode == 1:  # jump or jump-and-wait
            if pc + 2 > len(data):
                raise ValueError("truncated ANTIC jump operand")
            item["kind"] = "jvb" if instruction & 0x40 else "jump"
            item["target"] = data[pc] | (data[pc + 1] << 8)
            pc += 2
        else:  # playfield mode 2..15
            item["kind"] = "mode"
            item["mode"] = opcode
            item["dli"] = bool(instruction & 0x80)
            item["lms"] = bool(instruction & 0x40)
            item["vscroll"] = bool(instruction & 0x20)
            item["hscroll"] = bool(instruction & 0x10)
            if item["lms"]:
                if pc + 2 > len(data):
                    raise ValueError("truncated ANTIC LMS address")
                item["screen_address"] = data[pc] | (data[pc + 1] << 8)
                pc += 2

        item["next_pc"] = pc
        decoded.append(item)
        if item["kind"] == "jvb":
            break
    return decoded


program = bytes((0x70, 0x42, 0x00, 0x20, 0x41, 0x00, 0x10))
ops = decode_display_list(program)
assert ops[0]["kind"] == "blank" and ops[0]["lines"] == 8
assert ops[1]["lms"] and ops[1]["screen_address"] == 0x2000
assert ops[2]["kind"] == "jvb" and ops[2]["target"] == 0x1000

This structural pass does not know how many scanlines a playfield mode occupies or how many bytes it consumes. That depends on the mode, display width, fine-scroll state, and hardware timing. Keep structural decoding and line scheduling as separate modules; it makes malformed operands, mode-length mistakes, and DMA accounting independently testable.

Validation strategy

Start with a minimal display list containing a blank, one text mode, one bitmap mode with LMS, a DLI, and a JVB. Verify the decoded PC and screen pointer after every operand. Then change one modifier at a time and compare the number and timing of ANTIC memory requests. Use a memory trace to establish the exact fetch address, cycle, and bus owner; check a first text row separately from subsequent character rows.

For raster tests, schedule a DLI at a known mode-line boundary and write a visible color register in the handler. Repeat with WSYNC, HSCROL, VSCROL, narrow/normal/wide width, and player/missile DMA. Compare CPU instruction progress and write timestamps, not just screenshots. Test NTSC and PAL line geometry separately, and do not apply one scanline count or cycle-to-frame ratio to both.

Fuzz the parser with truncated jumps, truncated LMS operands, blank instructions, all mode opcodes, and jump loops. A malformed list should not read beyond mapped memory or spin the host thread forever. Any hardware-accurate wrap or open-bus behavior should be covered by a distinct fixture. Include state restoration mid-line, where the display-list PC, scan address, shift state, pending DMA requests, and DLI state must all resume consistently.

Acceptance criteria

A dependable ANTIC model keeps the display-list PC separate from the memory-scan address, decodes conditional operands exactly, schedules line-specific memory fetches, accounts for display and refresh DMA, and requests DLIs at their observable boundaries. It can explain CPU delay and pixel output from the same cycle trace. That makes Atari mixed-mode effects accurate without game-specific scanline patches.

Related:

Sources:

Comments