Skip to content
RetrogamingDeep Dive Published Updated 8 min readViews unavailable

ZX Spectrum 48K ULA Contention and Floating Bus: Fetch Phases and CPU Waits

Trace the Spectrum 48K ULA's shared display RAM fetches, CPU contention delays, floating-bus reads, frame phase, and model-specific timing evidence.

The ZX Spectrum 48K ULA and Z80 share access to the RAM that holds the display. The ULA must fetch bitmap and attribute bytes at specific points in the television scan, so it receives priority over the CPU during those fetch windows. The Z80 can consequently be held for additional T-states when it accesses contended memory. The same ULA bus activity can also leak into reads from otherwise unconnected I/O ports, creating the floating-bus behavior used by some software.

Sinclair’s service manual describes the memory arbitration principle: the ULA needs access to the display area at set intervals and has priority over CPU accesses. The precise timing pattern and floating-bus sequence are documented by later Spectrum timing research and implemented by mature emulators such as Fuse. Keep those sources distinguished: the service manual establishes the hardware arrangement, while model-specific edge timing comes from measured and cross-checked technical work rather than a single short block diagram.

Shared memory and display fetch ownership

On the 48K model, the display bitmap and its color attributes are in the lower 16 KiB of RAM, conventionally addressed from $4000 through $7FFF. The ULA fetches these bytes while generating active picture lines. The Z80 and ULA are not independent agents with private VRAM: they contend for a shared RAM resource. When the CPU makes a contended access during an interval needed by the ULA, the ULA wins and the CPU’s clock is inhibited for a model-specific delay.

This arbitration stretches machine time. An instruction that normally occupies a fixed count of Z80 T-states can take longer because one or more of its memory cycles were delayed. The exact delay depends on the current frame phase and the timing of the individual bus cycle, not just the instruction’s starting timestamp. A cycle-accurate core should therefore apply contention to bus accesses as they occur, rather than adding a flat penalty to every instruction that references $4000-$7FFF.

The 48K contention pattern commonly used for active display phases is 6, 5, 4, 3, 2, 1, 0, 0 T-states over an eight-T-state repeating window. That sequence is not a universal penalty for the full frame. Outside the display fetch region, the same address may not be delayed. The phase origin depends on the selected machine timing, ULA behavior, frame geometry, and when the emulated event occurs. A 128K Spectrum, Amstrad +2/+3, or clone can have a different timing profile and must not inherit the 48K constants by filename alone.

Why the bitmap and attribute reads alternate

The ULA’s bus schedule is tied to the Spectrum’s display organization. During a fetch group it reads bitmap data and the corresponding attribute byte, then the next bitmap and attribute, while arranging pixel data for the display shift register. A floating-bus timing guide documents an eight-T-state pattern for the 48K: four early transfers expose alternating pixel and attribute data, followed by bus-idle intervals. Repeated groups fill the active line before the ULA returns to border work.

The bitmap address order is not a simple linear row-major byte sequence. The Spectrum’s screen layout interleaves character rows and scanlines. An emulator should calculate the memory address for the ULA’s current display fetch and use that address both to retrieve the pixel/attribute value and to drive the bus-observation state. A separately hard-coded floating-bus pattern can appear correct on one screen while disagreeing with the actual bytes when software writes a new display pattern during a frame.

The implementation should preserve the distinction between “ULA is fetching a byte” and “the bus is idle.” Technical research identifies an idle value during border/idle phases, while fetch phases expose the byte currently being read. A CPU port read at the wrong T-state can therefore return an attribute instead of a pixel, or a bus-idle value instead. The value is tied to the ULA’s real fetch timeline; it is not a random open-bus byte and not necessarily the previous CPU memory read.

Floating bus is a bus observation, not a port device

The floating bus occurs because a read from an unconnected I/O address can observe data currently present on the ULA/memory bus. A program can issue an IN at a carefully timed point and see the pixel or attribute byte that the ULA is fetching. If the ULA is not reading display RAM at that time, the observable value follows the bus’s idle behavior documented by the relevant machine profile.

The CPU’s input transaction has its own phase. A floating-bus test must account for the Z80 I/O read sampling point, instruction duration, any contention delay on the I/O address, interrupt-entry timing, and the ULA’s fetch phase. Comparing only the T-state at which an IN instruction begins with the ULA’s memory-fetch table can produce an off-by-several-T-state error because the data is sampled later in the I/O cycle.

Refresh activity creates another special case. On some Spectrum hardware, the ULA’s contention logic can react to Z80 refresh/address lines in a way that is not captured by a simplistic rule that “only CPU reads and writes of display RAM contend.” This can produce snow effects or alter floating-bus observations when the Z80’s I register places refresh addresses in a particular range. Treat that behavior as a separate model-specific bus phenomenon; do not conflate it with ordinary screen-memory contention.

Timing profiles instead of one global constant

A stable emulator design defines a machine timing profile containing frame length, line length, interrupt position, active display start, contention window, fetch sequence, and any ULA-revision-specific behavior. CPU accesses query the profile at their exact T-state. The display renderer and floating-bus path use the same fetch scheduler, so all three consumers share one source of truth.

This avoids a common inconsistency: a renderer draws a line using one fetch phase, a CPU wait-state table uses a slightly different frame origin, and a port-read handler estimates the bus byte from yet another counter. The output may look visually correct while timing-sensitive games fail or border effects shift. An integrated event schedule should drive memory arbitration, pixel output, floating-bus state, and interrupt timing.

Do not port the 48K delay values to 128K or later models by changing only the line length. Memory banks, contention policy, ULA or gate-array revisions, and frame timing all matter. A useful machine profile is explicit about what evidence it represents: original 48K board issue and ULA where known, later ULA revision, 128K machine, Amstrad model, or a clone. If the hardware model is unknown, report that uncertainty rather than labeling it “48K accurate.”

A contention lookup fixture

The following Python function illustrates one 48K active-display delay table. The caller must supply the correct T-state origin and active interval from its machine profile. It does not model Z80 bus-cycle phase, I/O contention, refresh behavior, or any later Spectrum variant.

CONTENTION_48K = (6, 5, 4, 3, 2, 1, 0, 0)


def ula_wait_tstates(tstate, active_start, active_end):
    if active_end < active_start:
        raise ValueError("active interval must be ordered")
    if not active_start <= tstate < active_end:
        return 0
    return CONTENTION_48K[(tstate - active_start) & 7]


assert [ula_wait_tstates(t, 0, 8) for t in range(8)] == list(CONTENTION_48K)
assert ula_wait_tstates(8, 0, 8) == 0

For a real CPU core, a bus cycle may be delayed before the memory transfer starts, and the following phases should be recalculated from the new machine time. Merely summing the lookup result once per instruction is not a sufficient contention model.

Verification matrix

Begin with a 48K timing fixture that records every ULA display fetch and every Z80 memory transaction by T-state. Place CPU reads and writes at each phase of the eight-state contention window, on both contended and uncontended addresses, before active display, during active display, and in the border. Confirm the six nonzero waits and two zero-wait phases where the selected 48K profile specifies them. Include read-modify-write instructions and opcode fetches so the test covers the actual bus-cycle sequence rather than only data loads.

Next build a floating-bus pattern with deliberately distinct bitmap and attribute bytes. Execute I/O reads at the documented sampling phase of each fetch slot and verify the returned byte order, then sample idle phases and border. Change the screen RAM during active display to ensure the fetch path sees the correct value at the correct time. Repeat with an interrupt routine whose entry timing is known, and account for I/O access contention if the selected ULA profile requires it.

Test refresh/snow behavior separately by varying the Z80 refresh address and relevant machine issue. Do not let a workaround for this effect alter ordinary contended RAM accesses. Then run a suite across 48K ULA timing profiles and at least one 128K/Amstrad profile; the tests should fail if the wrong machine profile is selected.

Save-state tests should restore at an active fetch, a CPU wait state, a floating-bus read boundary, and just before the next display line. The restored ULA fetch address, pixel shift state, memory arbitration phase, pending CPU bus operation, and interrupt timing must agree. A video screenshot is not enough: save a machine-cycle trace and compare it against a trusted Fuse configuration or physical hardware capture where available.

Acceptance criteria

A trustworthy Spectrum model uses one ULA fetch timeline to drive video, shared-memory arbitration, and floating-bus reads. It applies CPU wait states at bus-cycle boundaries and selects a model-specific timing profile rather than treating every Spectrum as a 48K. It can explain which byte the ULA was fetching when a port read occurred, and how much the CPU was held while accessing contended RAM. That precision is necessary for raster tricks, display timing, and software that observes the otherwise invisible ULA bus.

Related:

Sources:

Comments