Skip to content
Written and edited in-house.Every figure, date and quote is taken from a named primary source — never from another site’s summary.Editorial policySpotted an error?
Emulation

Emulation accuracy versus speed, explained without the arguments

Retro··5 min read

A single misplaced cycle can turn a classic platformer into a slideshow. The difference between a jump that connects and a fall that feels unfair is often fewer than a thousand CPU cycles, and those cycles must arrive in the right order with respect to every component on the bus—the video chip, the sound processor, the memory controller, the input latch. Emulation that respects that order is called cycle-accurate, and it is enormously expensive, not because the original hardware was fast, but because every piece of it must be kept in lockstep with every other.

A single-board computer with video and audio connectors
A single-board machine. Accuracy costs host power, which is why the smallest hardware runs the oldest systems. Damianvila · CC BY-SA 4.0 · Wikimedia Commons

State and time: two things an emulator must get right

The emuStudio project’s documentation separates the problem into two obligations that are often conflated. An emulator must first be faithful to the internal state of the silicon—the register values, the flags, the exact contents of memory at every instant. That is the data fidelity most casual players notice. The second obligation, far more costly, is timing fidelity: preserving the instruction-cycle time of the real machine so that state transitions occur when the hardware would have produced them.

To do this, the emulator uses a scheduler that assigns each emulated component a fixed window of simulated time, which the project calls a timeSlice. The scheduler then computes how many machine cycles can fit inside that interval and lets each component run its allotment. Every chip, every bus transfer, every interrupt ticks forward inside its own timeSlice, and when the interval expires the whole ensemble is synchronised. Without that boundary, one component would race ahead of another and the delicate cooperation that the original engineers baked into the circuit board would dissolve.

Why a slow original does not mean a simple emulation

A common misunderstanding is that a machine that ran at a few megahertz ought to be trivial to reproduce. The speed of the target is almost irrelevant; what matters is how tightly the original designers coupled the components. A Carnegie Mellon storage-emulation paper describes a timing loop that runs alongside the simulation engine, deliberately holding emulated time close to the real-world clock so that a storage operation reports completion only after the determined physical duration has passed.

The same principle appears in a conference paper on high-performance microprocessor emulation. There, emulators post events to a processor’s timed event queue, interleaving instructions with disk accesses, interrupt servicing, and bus transactions. Even when the simulation is split across multiple threads, the paper notes that it must use fixed quanta to keep the parallel timelines roughly synchronised. A console from the early home-computing era can therefore demand tens of thousands of timed checkpoints per frame, every one of them a scheduling decision that stops the main loop dead. The slowness of the original is a comfort; the rigidity of its timing is the burden.

Cycle-approximate models and what they leave behind

Not every emulation project aims for perfect cycle fidelity, and the documentation that ships with commercial tools makes that plain. AMD’s hardware emulation documentation explains that the technology balances accuracy and execution speed by blending SystemC modelling with register-transfer-level co-simulation. It warns that the result does not reproduce hardware behaviour with one hundred per cent accuracy, and that differences in functionality can be observed. Some models are explicitly labelled cycle-approximate rather than fully cycle-accurate.

A separate design reference for AMD’s adaptive computing systems reiterates the point: hardware emulation gets very close to the real thing but stops short of a complete cycle-level replica. The gap between “very close” and “fully accurate” is where the characteristic flaws of an approximate core live—sprites that flicker one frame late, audio that drifts, timing-sensitive copy-protection checks that fail. Those glitches are not bugs in the emulator; they are predictable consequences of a deliberate trade-off.

Why full-cycle reproduction burns so much compute

The computational bill for removing that gap is steep. AMD’s hardware emulation guide states that the process runs orders of magnitude slower than the real hardware it is copying, precisely because register-transfer-level co-simulation forces the host to model signal transitions at the gate level rather than at the functional block level. Every wire, every latch, every propagation delay becomes a line item in the schedule.

The timing loop from the Carnegie Mellon paper and the timed event queue from the aerospace conference describe the same drain from the software side. The emulator cannot execute a single instruction without first consulting a schedule, checking whether a peripheral has finished its last operation, and advancing time by a thin slice. A game that on real hardware would chew through a million cycles per second may require the emulator to process tens of millions of scheduled intervals, each of which halts the simulation loop, recomputes the state of multiple components, and only then permits the next instruction. The machine is not being emulated quickly; it is being reconstructed, instant by instant, at a resolution the original engineers never needed to see.

What the published record does not yet show

The available evidence on emulation timing is strong on mechanism and lighter on casework. Which specific consoles from the eighties and nineties most clearly expose timing errors—and which titles on those platforms break in reproducible ways under cycle-approximate cores—is not settled in the vendor documentation or conference proceedings gathered here. The claim that certain bugs appear only under cycle-accurate emulation, and the broader argument that cycle-accurate cores are essential for software preservation rather than merely for playability, rest on what preservation foundations and platform holders may have documented internally. The primary sources do not yet put names and frame numbers to those failures. Similarly, the role of reconfigurable-logic reproductions as a preservation path, often invoked in discussions of field-programmable gate array cores, needs a firmer chain of primary evidence than the present documents supply.

The emulation community has spent decades learning that the last ten per cent of timing fidelity costs more than the first ninety. Whether the target is a games console from the early home-computing era or a signal processor on a modern aerospace board, the pattern holds: the slower and more deliberate the emulator, the less it falsifies the original’s behaviour. The unresolved question is not whether a machine can be made to boot and render a screen—that problem is largely solved. It is which specific moments in that machine’s library of software deserve the expense, and whether the institutions that preserve those works can afford to record them in flawless detail.