SOURCE ARCHIVE
EXTRACTED CONTENT
68,330 chars0 White Rose eprints@whiterose.ac.uk (¥) Research Online. https://eprints.whiterose.ac.uk Universities of Leeds, Sheffield and York
Deposited via The University of York.
White Rose Research Online URL for this paper:
https://eprints.whiterose.ac.uk/id/eprint/244135/
Version: Accepted Version
Proceedings Paper:
Sun, Jialin, Hu, Yuchen, You, Dean et al. (2026) Prelude: Priming-Guided State
Reconstruction for Efficient FPGA Processor Debugging. In: Design Automation
Conference 2026.
Reuse This article is distributed under the terms of the Creative Commons Attribution (CC BY) licence. This licence allows you to distribute, remix, tweak, and build upon the work, even commercially, as long as you credit the authors for the original work. More information and the full terms of the licence here: https://creativecommons.org/licenses/
Takedown If you consider content in White Rose Research Online to be in breach of UK law, please notify us by emailing eprints@whiterose.ac.uk including the URL of the record and the reason for the withdrawal request.
A University. . of er 53
LEEDS| ¢’ Sheffield A J Flt/
UNIVERSITY OF 07
Prelude: Priming-Guided State Reconstruction for Efficient FPGA
Processor Debugging
1,2
Jialin Sun1,2, Yuchen Hu1,2, Dean You1,2, Hui Wang , Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†
1School of Integrated Circuits, Southeast University, China 2National Center of Technology Innovation for EDA, China
3Department of Computer Science, University of York, UK
Abstract Prelude
Debugging complex FPGA prototypes of modern processors in em- FPGA Tradeoff
bedded systems is challenging due to limited signal visibility and System DESSERT[30] > \
significant tracing overhead. Existing approaches struggle to bal- Existing Method SW-like
ance execution efficiency with debugging capability, often requiring || Well Balanced Method NO) ENCORE[32] System
either expensive continuous tracing or heavyweight snapshots. We Extreme Method O
propose Prelude, a lightweight snapshot-based debugging frame- Ideal Method Trends StateMover[31]
work that records only essential architectural states and memory lower Debugging Capability higher
footprint of the processor on FPGA. During replay, a short visibility Figure 1. Tradeoff between execution efficiency and debugging capability.
warm-up reconstructs internal micro-architectural states, enabling typically requires observing thousands of interacting signals. As
cycle-accurate analysis. Implemented on BOOM and Rocket, Pre- design complexity increases, the demand for observable signals
lude provides comparable visibility to prior work while significantly grows accordingly. Moreover, because bugs often occur under un-
improving debugging efficiency: 32.88× / 2191.2× speedup over predictable conditions, it is difficult to determine in advance which
DESSERT / ENCORE on BOOM, and 18.09× / 896.4× on Rocket. signals are essential for diagnosis. Therefore, higher signal visi-
1 Introduction bility greatly improves the likelihood of capturing the root cause
Modern processor architectures have become increasingly com- when failures arise. Studies indicate that debugging can consume
plex [1ś3], incorporating deep pipelines, out-of-order execution, 60% to 70% of the total verification effort [26ś28], and the limited
and complex cache hierarchies [4ś7]. Verifying these designs re- waveform visibility on FPGA prototypes further magnifies this bot-
quires executing billions of cycles with real-world workloads such ~~ tleneck, motivating extensive research on improving observability
as operating system boots and application suites [8ś10]. Traditional and controllability during FPGA-based debugging.
RTL simulation (e.g., VCS [11], Verilator [12]), while providing full Existing work. Full-state capture methods have been proposed to
signal visibility, operates at extremely low speeds (∼1kHz), making achieve complete signal observability for FPGA debugging. One
extensive verification impractical. An IBM study illustrates this ~~ class of approaches relies on scan-chain-based designs (e.g., STATE- limitation, showing that booting a Linux system on an out-of-order ACCESS [29], DESSERT [30]), which insert dedicated observation superscalar processor through such software-based RTL simulation channels for every register and storage element. While this method would take nearly five years to complete [13]. Consequently, FPGA- can, in principle, provide full visibility, it scales poorly in practice. based prototyping has become essential for pre-silicon verification, ~~ Modern CPUs may contain millions of registers and other state offering execution speeds 1000 to 10000 times faster than software elements. Adding scan logic for each of them introduces over 100% simulation and enabling realistic workload testing. additional area overhead [31]. This massive instrumentation quickly However, FPGA-based verification suffers from fundamentally overwhelms FPGA resources, often preventing placement and rout- limited signal visibility, as only a restricted subset of internal states ing from completing. Even when implementation succeeds, the can be observed compared with the full transparency of software resulting system suffers from severe execution slowdown, making simulation. This limitation primarily stems from the reliance on normal program execution nearly impossible and limiting this ap- on-chip logic analyzers such as ILA [14], SignalTap [15], or Chip- proach to very small or simplified designs. Another class of methods Scope [16], which monitor internal signals by inserting additional ~~ (e.g., StateMover [31], ENCORE [32]) employs state readback [33ś probing logic either manually or through synthesis tools [17ś22]. 36], leveraging FPGA hardware features to extract the complete Such instrumentation is highly resource-intensive, consuming sig- internal state through JTAG [37]. JTAG provides only a single low- nificant block RAM, logic elements, and routing resources, and often speed serial channel that typically operates at several hundred kHz disrupts critical timing paths, which may lead to implementation to a few MHz, which makes transferring gigabytes of state data failures or prevent the design from meeting frequency targets [23ś extremely time-consuming. During readback, the running design 25]. As a result, only a minute fraction of internal signals (typically on the FPGA must be halted, and gigabytes of state data are trans- less than 0.01% of the total) can be traced within a short observa- ferred to the host, introducing substantial latency. Even when using tion window, offering only partial insight into system behavior. a high-speed PCIe interface, reading back the complete state re- This limited visibility is far from sufficient for debugging complex ~~ mains costly (e.g., it takes around 10s to read back a crossbar-big designs such as CPUs, where understanding erroneous behaviors ~~ design, which is orders of magnitude smaller than an out-of-order superscalar processor [31].). In addition, debugging requires repeat- †Corresponding author: Zhe Jiang. Email: zhejiang.uk@gmail.com. edly capturing system snapshots throughout execution to monitor Dac’26, Long Beach, CA, USA state evolution, leading to frequent interruptions and excessive data 2026. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM transfer overhead. This prevents continuous execution and makes https://doi.org/10.1145/3770743.3804273 it infeasible to debug long-running programs or workloads.
Execution Efficiency lower higher
1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA Jialin Sun , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†
In summary, existing FPGA debugging methods face a funda-
mental dilemma (Figure 1), as full-state approaches are costly while 90.0%
partial-state approaches offer only limited insight, preventing both 85.0%
high execution efficiency and comprehensive visibility.
Contributions. To address this challenge, we propose Prelude, a 80.0%
lightweight snapshotśbased debugging framework that achieves 75.0%
near-full signal visibility through replay. Instead of continuously 70.0%
tracing signals or capturing the entire design state, our method 10 100 1000 10000 100000 1000000
records only a compact snapshot containing essential architectural Instructions
states and partial signal information on FPGA. During replay on a Figure 2. Comparison of similarity between normal execution and partial-
software simulator, a short visibility warm-up sequence reconstru- state reconstruction across executed instructions.
cts most internal signals from these partial states before execut- summarized in Table 1, reproducing a segment correctly requires
ing the error segment, enabling cycle-accurate debugging with less than 0.001% of the full-state snapshot, yet achieves identical
minimal performance interference. We implement and evaluate functional replay fidelity.
this framework on BOOM [38], an open-source out-of-order super- Key Insight. Functionally correct replay can be achieved us-
scalar RISC-V processor, and Rocket [39], a five-stage in-order core, ing only architectural registers and accessed memory, without
achieving comparable debugging visibility with 32.88× / 2191.2× preserving the entire hardware state.
speedup over DESSERT [30]/ ENCORE [32] on BOOM and 18.09× /
896.4× on Rocket, demonstrating both efficiency and portability. 2.2 Error Trigger Capability.
2 Design Philosophy When debugging real-world processor bugs, engineers rarely need
Full-state capture provides a compelling demonstration that software- to observe all internal signals. Typically, only a subset of a few
level replay can enable detailed debugging if the hardware state hundred to a few thousand relevant signals is monitored, rather
is correctly reproduced. However, directly applying this approach than the complete set of registers and logic structures. Similarly,
to large out-of-order processors is impractical due to the prohib- triggering the same bug during replay does not require the entire
itive cost of capturing and transferring the complete state. This micro-architectural state, just like on the FPGA where the bug
observation highlights two essential requirements for practical occurs; for example, an ALU-related bug is independent of the state
replay-oriented debugging: the captured state must support (i) of unrelated units such as the FPU.
functionally correct replay and the system must retain sufficient Building on this insight and leveraging the transient nature
information to maintain (ii) error trigger capability. of micro-architectural states (Section 2.1), our method executes a
2.1 Functionally Correct Replay. short warm-up sequence of instructions prior to the segment of
Table 1. Comparison of full snapshot contents and our method (✓ = required interest. This warm-up allows the micro-architectural structures to
by our method, ✗ = not collected) converge to their original states, effectively reconstructing over 90%
of the necessary state for error triggering (Figure 2). By combining
Category Full Snapshot [31] Our Method selective signal capture with targeted warm-up, we achieve high
Architectural registers (PC, GPRs, Architectural registers (PC, debug fidelity without capturing the entire hardware state1.
CSRs, FPRs) GPRs, CSRs, FPRs) ✓
Pipeline registers ✗ Key Insight. Error triggering depends only on a limited subset
Micro-arch. Instruction queues ✗ of the micro-architectural state, which can be naturally recon-
State Reorder buffers ✗ structed through a short warm-up execution.
Load/store queues ✗
Branch predictors ✗ 3
Other internal buffers ✗ Framework
Size Based on the analysis in Section 2, we design our framework, Pre-
(micro-arch.) 23.4MB (full) 1KB (0.004%) lude, to meet these requirements. We employ lightweight architec-
Memory L1/L2/L3 caches ✗ tural snapshots combined with an online error detection mechanism
System Full main memory ✗ based on differential testing (Section 3.1) to identify bugs during
Size (memory) Accessed memory Accessed memory ✓ FPGA execution with minimal overhead. To enable deterministic
Size (total) 4GB (full) 17.59KB (0.0004%) replay from these snapshots, we augment the captured state with
4.023GB (full) 18.6KB (0.00046%) memory system reconstruction (Section 3.2), allowing a software
From the perspective of software execution, a program can simulator to replay the execution with full micro-architectural visi-
only observe and operate on architectural registers, benefiting bility (Section 3.3). Section 3.4 presents the complete workflow.
from the architectural abstraction that hides micro-architectural 3.1 Differential Testing
details. These architectural registers remain stable and fully de- Differential testing [4, 27, 32, 40ś43] has proven efficient for detect-
termine the instruction behavior within a segment, while other ing functional errors in complex processor designs by executing
micro-architectural structures are transient and evolve from them. the design under test (DUT) alongside a golden reference model,
Therefore, accurate segment replay only requires the architectural usually using the instruction set architecture (ISA) simulator [44],
portion rather than complete micro-architectural snapshot. and comparing their architectural states at synchronization points.
Similarly, for the memory system, only the memory locations ac- Any mismatch indicates a potential bug and triggers detailed state
cessed within the segment are relevant to its replay. Other memory
contents remain unused and thus have no effect on execution. As 1
The setup for the experiment is described in Section 4.
Reproduction Rate
Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA
FPGA Host PC Debugging
DUT Processor Differential Testing a Software Simulator c Segment e
Memory Arch. Log ISA Simulator Snapshot ...
Footprint Register ... Snapshot GPRs DUT Processor bez x4,74
FPRs add x3,x5 PC
Snapshot PCIe ... Compare Memory CSRs DEU AMU MRU ...
Rebuild Memory Data Data Data
Memory Footprint Memory Rebuild b FPGA SW Sim. Replay Hardware
Load Addr1 Data1 1
Read after Write Store Addr2 Data2 Initial Memory Arch. State d Software
Overlapping Load Load Addr2 Data2 Analyze Addr1 Data1 2 Replay 2 Warm‐up Log Message
Load Addr3 Data3 Addr3 Data3 FPGA Thrd.
Load Addr4 Data4 Addr5 Data5 3 3 Micro‐Arch.
Inclusive Load Load Addr5 Data5 ... ... Error State SW Sim. Thrd.
... ... ... Detect! Add Module
Figure 3. An overview of Prelude. (GPRs: General Purpose Registers; FPRs: Floating Point Registers; CSRs: Control and Status Registers; DEU: Data Extract Unit; AMU: Architectural Register Management Unit; MRU: Memory Rebuild Unit). a Differential testing identifies errors on the FPGA by comparing snapshots from the DUT and an ISA simulator. b Upon detecting an error, the memory footprint is analyzed to reconstruct the initial memory state. c The architectural registers and reconstructed memory are then used for software replay. d The DUT starts from the captured architectural state and replays instructions forward to regenerate the micro-architectural state. e The error-containing segment is replayed with nearly full signal visibility for debugging. replay. Previous work [32, 41] transmit complete architectural state Algorithm 1: Memory Footprint Analysis at every comparison, incurring substantial communication over- Input: address range 𝐼 = [𝐴 head for processors with large register files. Prelude addresses this set W 𝑠, 𝐴𝑒 ] with data 𝐷; read set R; write by employing lightweight snapshots (Figure 3 a ) that are sufficient Output: Updated read set R′ 1 foreach ( [𝑅 to reflect the execution of the DUT processor, while more detailed 𝑠 Inclusive Load Ranges 2 // 𝑖 , 𝑅𝑒𝑖 ], 𝐷𝑖 ) ∈ R do micro-architectural states are reconstructed from these snapshots. 3 if 𝐼 ⊇ [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ] then Snapshot. In Prelude, a snapshot captures only the architectural 4 R ← R \ { ( [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ], 𝐷𝑖 ) } ; // remove old load state and the memory locations accessed within the segment (Sec- 5 end 𝐷 [𝑅𝑠𝑖 ,𝑅𝑒𝑖 ] ← 𝐷𝑖 ; // copy data into current load tion 2.1), following the design philosophy that the software-visible 6 // Overlapping Load Addresses state fully determines execution behavior. 7 𝑖 , 𝑅𝑒𝑖 ] ∩ 𝐼 ≠ ∅ then 8 else if [𝑅 • Architectural Registers. All general-purpose registers 𝐼 𝐼 𝑠 9 ← \ ( [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ] ∩ 𝐼 ) ; // remove overlap (GPRs), program counter (PC), floating-point registers (FPRs) 10 end 𝐷 ← 𝐷 |𝐼 ; // keep non-overlapping data and relevant control and status registers (CSRs). These are 11 end sufficient to reproduce instruction execution within the seg- 12 // Read-after-write Dependencies 13 ment, as other micro-architectural structures, e.g., pipeline 14 foreach [𝑊𝑠𝑗,𝑊𝑒𝑗 ] ∈ W do buffers, evolve in a predictable manner from them. 15 if 𝐼 ⊆ [𝑊𝑠𝑗,𝑊𝑒𝑗 ] then • Accessed Memory Footprint. Only memory addresses 16 end 𝐼 ← ∅ ; // discard load accessed by instructions within the segment are included. 17 end 18 This ensures that memory contents, which do not affect 19 if 𝐼 ≠ ∅ then execution, are excluded, minimizing snapshot size. 20 R′ ← R ∪ { (𝐼, 𝐷 ) }; // add remaining load to read set Snapshots are transmitted from FPGA to host PC. During differ- 21 end ential testing, the reference model executes the same program to is analyzed (Figure 3 b ), with particular attention to the addresses the corresponding synchronization points, where its architectural state is compared against the DUT. Any mismatch triggers further and data of all load operations. As illustrated in Algorithm 1, special analysis or additional snapshot collection. By restricting snapshots handling is required for several cases that may affect correctness: to the minimal set of software-visible state, Prelude significantly Read-after-write Dependencies. If a load reads from an address reduces storage and communication overheads to less than 0.001% that has already been written within the same segment, both the of a full-state snapshot, while preserving the ability to detect errors. address and the data of this load are discarded, since the value is 3.2 Memory Rebuild internally defined and does not depend on the initial memory state. Overlapping Load Addresses. When two load operations partially During checking, all memory accesses within a segment are recorded. overlap in their accessed address ranges, only the portion of the For replay, however, only the memory state at the segment’s be- later access that covers previously unread bytes is retained. The ginning is required. The memory rebuild process bridges this gap overlap region is removed, and the effective access range of the by analyzing the recorded memory footprint to reconstruct the preserved load is narrowed accordingly. necessary initial memory state for the segment’s execution. Inclusive Load Ranges. If a later load covers a larger address range To identify the required initial memory state, we focus on the first that includes one or more earlier accesses, the earlier and smaller read of each memory address in the memory footprint, provided loads are removed. For the overlapping region, the data in the later that the address has not been written before the read occurs. In load are replaced with the corresponding bytes from the earlier this process, the complete memory footprint obtained from FPGA accesses to ensure consistency with the original access order.
1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA Jialin Sun , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†
Data Extract Unit (DEU) the Load-Store Queue (LSQ) to clear any in-flight operations, then
DF start
!end State⁰ˣ⁰¹ + | Mem_start!Mem_end 0x03 + restores the memory footprint by injecting data through the DCache
Arch. Memory 0x01 interface according to the reconstructed snapshot. After both the
0x1F Calculation Mem_size Calculation 0x02 a AMU and MRU finish register and memory initialization, the DUT
resumes execution from the replay segment’s starting PC. A brief
Arch. Reg Management Unit (AMU)
5 64bits Clr.(MRU)lsq FIFO 64bits 2bits 64bits warm-up phase then allows internal queues, buffers, and predictors
Start Condition bits | Memory Rebuild Unit
flush Clr. rob A_Addr Data Addr Size Data to settle, ensuring that replay proceeds under a consistent and
0x01 ... Addr1 Size1 Data1 reproducible micro-architectural context.
Ctrl.Init 0x02 ... Addr2 Size2 Data2
... b ... c 3.4 Workflow
Redirect_PC In this subsection, we present the overall workflow of our FPGA
Instruction Fetch Unit (IFU) Re‐Order Buffer (ROB) debugging framework (Figure 3). The process consists of captur-
Next_PC |
Clr. rob Files (PRFs) Clr. lsq ing lightweight snapshots during normal execution, performing
Ctrl.Init PRFs Controllers Physical Register
5bits 7bits Addr 6bits 64bits Micro differential testing to detect errors, and reconstructing both archi-
A_Addr P_Addr Write_Ctrl[0] Data Addr Data ‐ops Load
0x01 0x04 Arch. andPhy. 0x00 ... Store §Addr tectural and memory states for replay (Figure 5). If a replay fails
0x02 0x07 RegisterAddr. 0x01 ... wb Unit Data to reproduce a bug, a full-state snapshot is captured as a fallback.
... Write_Ctrl[2] 0x02 ... Micro This workflow balances visibility and efficiency, minimizing per-
0x1d 0x24 0x03 ... ‐ops
0x1e 0x2f ... ... wb Execute formance overhead while ensuring accurate debugging.
0x1f 0x01 Write_Ctrl[3] 0x7f ... = Units Compare
Map Table ISA Sim. 1 Wait 2 Wait 3 Wait Error Detect REF Snapshot
FPGA Snapshot Warm up, DUT SW Sim.
DCache 1 2 3 reproduce micro‐arch.
Replay Arch. Micro‐arch. Bug Found
SW Sim. 2 3
Figure 4. The micro-architecture of the DUT processor in software simula- time
tor (DF: Data Fetch; blue: Architectural Register Initialization; red: Memory 5 10 15 20 25 30 35 40 45 50 55 60 65 70
State Initialization; green: Control Signal). a DEU fetches data from log file Figure 5. FPGA debugging workflow timeline, showing snapshot capture,
and transmits to AMU and MRU. b AMU completes architectural register differential testing, replay, warm-up and debugging.
initialization and c MRU completes memory state initialization. Error Detection. The DUT on the FPGA is executed in parallel with
3.3 Micro-architecture Add-on the ISA simulator on the host PC (Figure 3 a ). Their architectural
A general, non-intrusive mechanism reconstructs memory and ar- states are periodically compared at predefined synchronization
chitectural states from lightweight logs to enable deterministic re- points. Any mismatch between the two indicates a potential func-
play and consistent micro-architectural execution. This mechanism tional error, which triggers the subsequent debugging workflow.
is demonstrated on BOOM, an out-of-order superscalar RISC-V Snapshot Processing. After an error is detected, the software snap-
processor, as a representative case study (Figure 4). shot is analyzed to reconstruct only the minimal memory state
Data Extract Unit (DEU, Figure 4 a ). The DEU serves as the needed for replay (Figure 3 b ), which is then loaded into the simu-
bridge between snapshot processing and hardware replay. It re- ~~ lator to accurately reproduce the failing segment (Figure 3 c ).
trieves architectural register values and reconstructed memory con- Visibility Warm-up. In this stage, the DUT executes a short in-
tents from the snapshot log and transfers them to the simulation struction sequence in the simulator to propagate architectural state
environment via the standard DPI-C interface. After parsing, the into micro-architectural state (Figure 3 d ), ensuring all internal
DEU routes register data to the Architectural Register Management signals are properly initialized for consistent replay and debugging.
Unit (AMU) and memory data to the Memory Rebuild Unit (MRU), Debugging. Starting from the prepared micro-architectural state,
enabling precise initialization of both subsystems. By decoupling the erroneous segment is fully replayed in the software simulator,
data extraction from replay control, the DEU provides a clean and where all signals and architectural states are visible for fine-grained
efficient path for propagating snapshot information, ensuring that inspection and accurate root-cause identification (Figure 3 e ).
deterministic replay begins from the correct architectural state. Fallback. If the bug cannot be successfully reproduced during
Architectural Register Management Unit (AMU, Figure 4 b ). replay, the FPGA must be re-run to capture a full-state snapshot.
The AMU reconstructs the processor’s architectural state for replay ~~ In this case, execution only needs to proceed up to the start of the after the DEU provides the snapshot data. It first resets the pipeline erroneous segment, at which point all internal signals are taken. by flushing the ROB and returning all physical registers to the free Intermediate differential testing and non-essential snapshots are list, ensuring a clean architectural configuration without specula- ~~ skipped, avoiding unnecessary overhead while still providing the tive or in-flight state. Using BOOM’s map table, the AMU writes complete state required for subsequent replay and debugging. each architectural register value to its assigned physical register and 4 Evaluation redirects the Instruction Fetch Unit (IFU) to the replay segment’s We evaluate Prelude on the BOOM, an open-source out-of-order starting program counter, reconstructing the processor state ac- superscalar RISC-V processor. The design is implemented on an cording to the snapshot. After this reset-and-restore procedure, the AMD Virtex UltraScale+ VU19P FPGA [45] using Xilinx Vivado AMU waits for the MRU to complete memory initialization so that 2024.2. For functional reference, we use Spike as the ISA simulator, replay can begin from a fully consistent architectural state. while software-level replay and verification are performed with Memory Rebuild Unit (MRU, Figure 4 c ). The MRU reconstructs Verilator. We use CoreMark [46] and Embench [47] benchmark the initial memory state required for replay. It begins by flushing suites to measure debugging performance metrics.
Snapshot
... ......
Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA
Verilator[12] ENCORE[32] DESSERT[30] Prelude(Ours)
104
103
102
10
1
Figure 6. Debugging performance results for Prelude, Verilator [12], ENCORE [32] and DESSERT [30], running CoreMark and Embench.
90.0% 100 1,000 10,000 100,000 1,000,000 100.0% 100 1,000 10,000 100,000 1,000,000
85.0% 95.0%
80.0% 90.0%
75.0% 85.0%
70.0% 80.0%
65.0% Fetch Decode Rename 75.0% Register-read Dispatch Issue
(a) Front-end Reproduction Rate. (b) Mid-end Reproduction Rate.
100.0% 100 1,000 10,000 100,000 1,000,000 90.0% 100 1,000 10,000 100,000 1,000,000
95.0% 85.0%
90.0% 80.0%
85.0% 75.0%
80.0% Execute LSU ROB 70.0% DCache L2Cache Bus
(c) Back-end Reproduction Rate. (d) Memory System Reproduction Rate.
Figure 7. Signal reproduction rates of different hardware module groups at various pipeline stages, under varying instruction counts.
4.1 Evaluation Metrics efficiency. We compare Prelude against three representative base-
To quantify the debugging efficiency of our framework, we define lines: (1)Verilator [12], a software simulator which is capable of full
the effective debugging time 𝑇eff as: waveform tracing and provides the highest visibility but extremely
bugging via readback of internal configuration and state data; and
𝑇eff = 1 𝑁 (𝑇exe,𝑖 + (1 − 𝑃err,𝑖 ) · 𝑇Penalty,𝑖 ) (1) low execution speed; (2) ENCORE [32], which performs FPGA de-
𝑁 𝑖=1 (3) DESSERT [30], which uses scan-chain based state extraction
Here, 𝑇exe,𝑖 denotes the FPGA execution time required to reach a for post-mortem analysis. All baseline approaches rely on either
debuggable state, 𝑃err,𝑖 is the probability that a bug is successfully full-state or near-complete signal retrieval, leading to substantial
triggered and diagnosed within the segment, and 𝑇Penalty,𝑖 repre- runtime and data transfer overhead during debugging.
sents the extra time consumed if the bug is not detected. 𝑁 is the Across all workloads, Prelude achieves substantial improvements
number of checkpoints. This metric reflects both the execution in debugging efficiency, with a geomean speedup of 2191.2× over
efficiency and the likelihood of effective error localization. ENCORE and 32.88× over DESSERT. The maximum speedups reach
We further report the normalized slowdown as: 5714.6× and 73.94×, respectively. Even for very short workloads (e.g.,
slowdown = 𝑇𝑇others (2) nbody, st), Prelude still achieves 66.67× and 2.11× improvements.
ours The smaller gains on short workloads occur because the warm-up
where 𝑇others and 𝑇ours represent the effective debugging time of phase runs on a software simulator, which is slower than the FPGA.
baseline methods and our system, respectively. The warm-up cost is fixed, so when the FPGA execution segment
We additionally define the reproduction rate as: becomes short, the warm-up occupies a larger fraction of the total
𝑅𝑒𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑅𝑎𝑡𝑒 = 𝑆match (3) runtime. For longer workloads (e.g., nettle-aes and nettle-sha256),
𝑆total the FPGA segment takes most of the total time, the relative impact
of the warm-up decreases, and the performance improves.
where 𝑆match is the number of micro-architectural signals that Table 2. Time Breakdown Ratios
match the ground-truth execution after replay, and 𝑆total is the total
number of compared signals. A higher value indicates that the Execution Snapshot Warm-up Fallback
replayed micro-architectural state agrees with the actual hardware Ratio (%) 65.25 12.20 13.48 9.07
state and supports reliable error analysis.
4.2 Results and Analysis Performance Breakdown.To better analysis Prelude’s internal
Result#1: Outperforming prior tools in debugging perfor- performance characteristics, we further break down its end-to-end
mance. Figure 6 shows the normalized slowdown when running debugging time into four components: FPGA execution, snapshot
CoreMark and the Embench benchmark suite under different debug- transfer, warm-up reconstruction, and replay fallback. As shown
ging frameworks. A smaller slowdown indicates higher runtime in Table 2, this breakdown shows that the dominant cost lies in
Reproduction Rate Reproduction Rate Slowdown
Reproduction Rate Reproduction Rate
sglib-combined slre st ud wikisort Geo.Mean
coremarkaha-mont64 crc32 cubic edn huffbenchmatmult-int minver nbody nettle-aesnettle-sha256 nsichneu picojpeg qrduino
1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA Jialin Sun , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†
ENCORE[32] DESSERT[30] Ours Table 3. Area overhead of BOOM and Rocket on FPGA.
103 BOOM Rocket Core Resource Pure With Absolute Relative
Core Prelude Overhead Overhead
100 Logic LUTs 853,622 879,230 25,608 3.00%
10 BOOM LUTRAMs 26,749 27,328 579 2.16%
Flip-Flops 477,416 491,818 14,402 3.02%
1 Logic LUTs 201,402 211,136 9,734 4.83%
Figure 8.CoreMark Embench CoreMark Embench Rocket LUTRAMs 22,816 23,306 490 2.15%
Debugging performance evaluation on BOOM and Rocket Flip-Flops 271,477 280,255 8,778 3.23%
the fast FPGA-execution portion, while snapshot and warm-up Result#4: Efficient integration into diverse RISC-V cores with
overheads remain small and stable, validating the efficiency of very low area cost. The area overhead of integrating Prelude is
Prelude’s snapshot-plus-warm-up design. evaluated on BOOM and Rocket cores using FPGA implementations.
Result#2: Effectiveness of reproduction across all pipeline Table 3 reports the utilization of logic LUTs, LUTRAMs, and flip-
stages. Figure 7 reports the signal reproduction rate of each hard- flops for each core. Compared to the baseline designs, Prelude
ware module in pipeline stage under varying instruction counts. introduces only minor resource increases. These results indicate
Overall, the geomean rises from 86.7% at 102 instructions to 91.5% that Prelude can be efficiently incorporated into different RISC-V
at 106 (Figure 2), showing that longer workloads provide more cores without significant hardware cost.
complete context for reconstructing micro-architectural state.
Front-end (Figure 7(a)). The Fetch module has a lower repro- 4.3 Case Study
duction rate due to dynamic structures like the branch predictor During the evaluation of Prelude, we identified a previously un-
requiring long execution history, while Decode and Rename achieve known bug on BOOM’s floating-point (F) extension. Certain in-
higher rates as their states mainly depend on current instructions, structions produced incorrect results due to incomplete implemen-
which can be reconstructed from short segments. tation. While running our framework, a mismatch was detected at
Mid-end (Figure 7(b)). Similar to Decode and Rename, the Register- a checkpoint in the program counter (PC). Using the state saved at
read, Dispatch, and Issue modules all achieve high reproduction the previous checkpoint, we performed a warm-up to restore the
rates, as their states mainly depend on current instructions and can micro-architectural state before executing the erroneous segment.
be reconstructed from short segments. 1 0x80002390: sd s4 ,0( gp )
Back-end (Figure 7(c)). The ROB module reaches very high re- 2 0x80002394: fcvt.l.s a2 , fs11
production rates because most of its signals store instruction in- 3 0x80002398: auipc tp ,0x0
formation, with only a small portion representing dynamic control Listing 1. Instructions from the erroneous segment
signals. The LSU module maintains relatively stable reproduction Full-waveform debugging of the segment revealed that when
rates, as its state reflects memory access patterns that are partially execution reached address 0x80002394 (Listing 1), the reorder buffer
dynamic but largely deterministic over short instruction segments. (ROB) finished the instruction, but the corresponding exception was
Memory System (Figure 7(d)). The L2Cache module has relatively not committed. Step-by-step analysis of the waveform showed that
low reproduction rates because it relies on complex caching policies the decode unit did not have an implementation for this instruction,
and dynamic replacement decisions, which require long execution causing it to be flagged as an exception. After identifying the root
history to reconstruct accurately. This low reproduction rate does cause, we confirmed that executing the corrected instruction in
not affect bug triggering, because errors depend primarily on archi- BOOM produced results consistent with the expected model. Using
tectural memory operations rather than detailed L2Cache control Prelude, the entire bug discovery and debugging process, including
states. The DCache module improves moderately with longer in- state restoration and warm-up, was completed in under three min-
struction segments, similar to other cache-related modules. The Bus utes. In comparison, performing the same analysis with Verilator
module remains stable across all instruction counts, as its signals would have taken roughly 50 hours, roughly 1, 000 times longer,
mainly represent deterministic data transfers and arbitration states. demonstrating the significant efficiency advantage of Prelude.
Result#3: Effortlessly portable between BOOM and Rocket. 5
Figure 8 shows the normalized slowdown of Prelude when ported Conclusion
from BOOM to Rocket [39] using the same CoreMark and Embench We have presented Prelude, a lightweight and portable FPGA de-
workloads. Despite differences in pipeline depth, execution model, bugging framework that achieves near-full visibility without full-
and micro-architectural state, Prelude requires no design-specific state capture. By combining compact snapshots with a short replay
tuning beyond connecting replay units. Rocket lacks a reorder warm-up, Prelude reconstructs the required micro-architectural
buffer and physical register files, so initialization only flushes the state for cycle-accurate debugging while imposing minimal over-
pipeline and writes values directly into logical registers, making head. Overall, Prelude offers a practical and efficient approach for
migration straightforward while preserving framework generality. debugging modern FPGA-based processors.
Across both benchmarks, Prelude maintains a clear debugging 6
advantage over ENCORE and DESSERT on Rocket. The gap narrows Acknowledgement
compared with BOOM due to Rocket’s lower IPC increasing execu- We’d like to thank the reviewers for the helpful feedback. This
tion time, yet Prelude still achieves 27.85× and 18.10× speedups on work is supported by the National Natural Science Foundation of
CoreMark and Embench, demonstrating that lightweight snapshots, China (No. 62472086, 92464204), the Science and Technology Major
deterministic replay, and warm-upśguided reconstruction provide Special Program of Jiangsu (No. BG2024010), and the Fundamental
a portable, efficient debugging solution across RISC-V cores. Research Funds for the Central Universities (No. 2242025K20013).
Slowdown
Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA
References llms. Proceedings of the IEEE/ACM Design Automation Conference (DAC) (2025). [1] Bryan H Fletcher. 2005. FPGA embedded processors. In Embedded Systems [27] Jialin Sun, Yuchen Hu, Dean You, Yushu Du, Hui Wang, Xinwei Fang, Weiwei Conference. 18. Shan, Nan Guan, and Zhe Jiang. 2025. ISAAC: Intelligent, Scalable, Agile, and [2] Fuming Sun, Xiaoying Li, Qin Wang, and Chunlin Tang. 2008. FPGA-based Accelerated CPU Verification via LLM-aided FPGA Parallelism. arXiv preprint embedded system design. In APCCAS 2008-2008 IEEE Asia Pacific Conference on arXiv:2510.10225 (2025). Circuits and Systems. IEEE, 733ś736. [28] Ke Xu, Jialin Sun, Yuchen Hu, Xinwei Fang, Weiwei Shan, Xi Wang, and Zhe [3] Michael Gschwind, Valentina Salapura, and Dietmar Maurer. 2002. FPGA proto- Jiang. 2024. Meic: Re-thinking rtl debug automation using llms. In Proceedings of typing of a RISC processor core for embedded applications. IEEE Transactions on the IEEE/ACM International Conference on Computer-Aided Design (ICCAD). 1ś9. Very Large Scale Integration (VLSI) Systems 9, 2 (2002), 241ś250. [29] Dirk Koch, Christian Haubelt, and Jürgen Teich. 2007. Efficient hardware check- [4] Nursultan Kabylkas, Tommy Thorn, Shreesha Srinath, Polychronis Xekalakis, pointing: concepts, overhead analysis, and implementation. In Proceedings of and Jose Renau. 2021. Effective processor verification with logic fuzzer enhanced the 2007 ACM/SIGDA 15th international symposium on Field programmable gate co-simulation. In MICRO-54: 54th Annual IEEE/ACM International Symposium on arrays. 188ś196. Microarchitecture. 667ś678. [30] Donggyu Kim, Christopher Celio, Sagar Karandikar, David Biancolin, Jonathan [5] Chloe Tain, Savita Patil, and Hussain Al-Asaad. 2025. Survey of Verification of Bachrach, and Krste Asanović. 2018. DESSERT: Debugging RTL effectively RISC-V Processors. Journal of Electronic Testing (2025), 1ś28. with state snapshotting for error replays across trillions of cycles. In 2018 28th [6] Eric Sprangle and Doug Carmean. 2002. Increasing processor performance by International Conference on Field Programmable Logic and Applications (FPL). IEEE, implementing deeper pipelines. ACM SIGARCH Computer Architecture News 30, 76ś764. 2 (2002), 25ś34. [31] Sameh Attia and Vaughn Betz. 2020. StateMover: Combining simulation and [7] Jin Li, Kristin Tufte, Vladislav Shkapenyuk, Vassilis Papadimos, Theodore John- hardware execution for efficient FPGA debugging. In Proceedings of the 2020 son, and David Maier. 2008. Out-of-order processing: a new architecture for ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. 175ś high-performance stream systems. Proceedings of the VLDB Endowment 1, 1 185. (2008), 274ś288. [32] Kan Shi, Shuoxiang Xu, Yuhan Diao, David Boland, and Yungang Bao. 2023. [8] Google. 2019. RISC-V DV. https://github.com/google/riscv-dv. ENCORE: Efficient architecture verification framework with FPGA accelera- [9] Sadullah Canakci, Chathura Rajapaksha, Leila Delshadtehrani, Anoop Nataraja, tion. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field Michael Bedford Taylor, Manuel Egele, and Ajay Joshi. 2023. Processorfuzz: Programmable Gate Arrays. 209ś219. Processor fuzzing with control and status registers guidance. In 2023 IEEE In- [33] Hari Angepat, Gage Eads, Christopher Craik, and Derek Chiou. 2010. NIFD: ternational Symposium on Hardware Oriented Security and Trust (HOST). IEEE, Non-intrusive FPGA DebuggerśDebugging FPGA’Threads’ for Rapid HW/SW 1ś12. Systems Prototyping. In 2010 International Conference on Field Programmable [10] Chen Chen, Rahul Kande, Nathan Nguyen, Flemming Andersen, Aakash Tyagi, Logic and Applications. IEEE, 356ś359. Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. 2023. HyPFuzz:Formal- [34] Ashfaquzzaman Khan, Richard Neil Pittman, and Alessandro Forin. 2010. gNOSIS: Assisted processor fuzzing. In 32nd USENIX Security Symposium (USENIX Security A board-level debugging and verification tool. In 2010 International Conference 23). 1361ś1378. on Reconfigurable Computing and FPGAs. IEEE, 43ś48. [11] Synopsys. 2024. VCS: Synopsys Verification Continuum. https://www.synopsys. [35] Changgong Li, Alexander Schwarz, and Christian Hochberger. 2016. A readback com/verification/simulation/vcs.html based general debugging framework for soft-core processors. In 2016 IEEE 34th [12] Wilson Snyder. 2024. Verilator. https://www.veripool.org/wiki/verilator International Conference on Computer Design (ICCD). IEEE, 568ś575. [13] Sameh Asaad, Ralph Bellofatto, Bernard Brezzo, Chuck Haymes, Mohit Kapur, [36] Georgios Tzimpragos, Da Cheng, Stephanie Tapp, Balakrishna Jayadev, and Benjamin Parker, Thomas Roewer, Proshanta Saha, Todd Takken, and José Tierno. Amitava Majumdar. 2016. Application debug in FPGAs in the presence of multiple 2012. A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating asynchronous clocks. In 2016 International Conference on Field-Programmable multi-core processor simulation. In Proceedings of the ACM/SIGDA international Technology (FPT). IEEE, 189ś192. symposium on Field Programmable Gate Arrays. 153ś162. [37] Stephanie Tapp. 2015. XAPP1230: Configuration Readback Capture in UltraScale [14] Xilinx. 2016. Integrated Logic Analyzer v6.2. https://www.xilinx.com/support/ FPGAs. documentation/ip_documentation/ila/v6_2/pg172-ila.pdf [38] Jerry Zhao, Ben Korpan, Abraham Gonzalez, and Krste Asanovic. 2020. Sonic- [15] Intel. 2020. Intel Quartus Prime Pro Edition User Guide: Debug boom: The 3rd generation berkeley out-of-order machine. In Fourth Workshop on Tools. https://www.intel.com/content/dam/www/programmable/us/en/pdfs/ Computer Architecture Research with RISC-V, Vol. 5. International Symposium on literature/ug/ug-qpp-debug.pdf Computer Architecture Valencia, Spain, 1ś7. [16] Xilinx. 2012. Chipscope Pro Software and Cores: User Guide. https://docs.amd. [39] Krste Asanovic, Rimas Avizienis, Jonathan Bachrach, Scott Beamer, David Bian- com/v/u/en-US/chipscope_pro_sw_cores_ug029 colin, Christopher Celio, Henry Cook, Daniel Dabbelt, John Hauser, Adam Izraele- [17] Eddie Hung and Steven JE Wilton. 2012. Scalable signal selection for post-silicon vitz, Sagar Karandikar, Ben Keller, Donggyu Kim, and John Koenig. 2016. The debug. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 21, 6 rocket chip generator. EECS Department, University of California, Berkeley, Tech. (2012), 1103ś1115. Rep. UCB/EECS-2016-17 4 (2016), 6ś2. [18] Jeffrey Goeders and Steven JE Wilton. 2016. Signal-tracing techniques for in- [40] Kunlin You, Yinan Xu, Kehan Feng, Luoshan Cai, Yaoyang Zhou, and Yungang system FPGA debugging of high-level synthesis circuits. IEEE Transactions on Bao. 2025. DiffTest-H: Toward Semantic-Aware Communication in Hardware- Computer-Aided Design of Integrated Circuits and Systems 36, 1 (2016), 83ś96. Accelerated Processor Verification. In Proceedings of the 58th IEEE/ACM Interna- [19] Daniel Holanda Noronha, Ruizhe Zhao, Jeff Goeders, Wayne Luk, and Steven JE tional Symposium on Microarchitecture®. 1462ś1476. Wilton. 2019. On-chip FPGA debug instrumentation for machine learning ap- [41] Yinan Xu, Zihao Yu, Dan Tang, Guokai Chen, Lu Chen, Lingrui Gou, Yue Jin, plications. In Proceedings of the 2019 ACM/SIGDA International Symposium on Qianruo Li, Xin Li, Zuojun Li, Jiawei Lin, Tong Liu, Zhigang Liu, Jiazhan Tan, Field-Programmable Gate Arrays. 110ś115. Huaqiang Wang, Huizhe Wang, Kaifan Wang, Chuanqi Zhang, Fawang Zhang, [20] David Sidler and Ken Eguro. 2016. Debugging framework for FPGA-based soft Linjuan Zhang, Zifei Zhang, Yangyang Zhao, Yaoyang Zhou, Yike Zhou, Jiangrui processors. In 2016 International Conference on Field-Programmable Technology Zou, Ye Cai, Dandan Huan, Zusong Li, Jiye Zhao, Zihao Chen, Wei He, Qiyuan (FPT). IEEE, 165ś168. Quan, Xingwu Liu, Sa Wang, Kan Shi, Ninghui Sun, and Yungang Bao. 2022. To- [21] Fatemeh Eslami and Steven JE Wilton. 2014. Incremental distributed trigger wards developing high performance RISC-V processors using agile methodology. insertion for efficient FPGA debug. In 2014 24th International Conference on Field In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). Programmable Logic and Applications (FPL). IEEE, 1ś4. IEEE, 1178ś1199. [22] Gefei Zuo, Jiacheng Ma, Andrew Quinn, and Baris Kasikci. 2023. Vidi: Record [42] Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, and replay for reconfigurable hardware. In Proceedings of the 28th ACM International Byoungyoung Lee. 2021. Difuzzrtl: Differential fuzz testing to find cpu bugs. In Conference on Architectural Support for Programming Languages and Operating 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 1286ś1303. Systems, Volume 3. 806ś820. [43] Fabian Thomas, Lorenz Hetterich, Ruiyi Zhang, Daniel Weber, Lukas Gerlach, and [23] Brad L Hutchings and Jared Keeley. 2014. Rapid post-map insertion of embed- Michael Schwarz. 2024. RISCVuzz: Discovering architectural CPU vulnerabilities ded logic analyzers for Xilinx FPGAs. In 2014 IEEE 22nd annual international via differential hardware fuzzing. https://ghostwriteattack. com/ (2024). symposium on field-programmable custom computing machines. IEEE, 72ś79. [44] RISC-V. 2024. Spike RISC-V ISA Simulator. https://github.com/riscv-software- [24] Pavan Kumar Bussa, Jeffrey Goeders, and Steven JE Wilton. 2017. Accelerating src/riscv-isa-sim in-system FPGA debug of high-level synthesis circuits using incremental compi- [45] AMD. 2025. Virtex UltraScale+ VU19P. https://www.amd.com/en/products/ lation techniques. In 2017 27th International Conference on Field Programmable adaptive-socs-and-fpgas/fpga.html. Logic and Applications (FPL). IEEE, 1ś4. [46] EEMBC. 2018. CoreMark. https://github.com/eembc/coremark. [25] Eddie Hung and Steven JE Wilton. 2013. Incremental trace-buffer insertion for [47] David Patterson, Jeremy Bennett, Mary Bennett, Hélène Chelin, David Harris, FPGA debug. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 22, Jennifer Hellar, William Jones, Konrad Moron, Paolo Savini, Roger Shepherd, 4 (2013), 850ś863. Ray Simar, Zachary Susskind, and Stefan Wallentowitz. 2025. Embench IOT 2.0 [26] Yuchen Hu, Junhao Ye, Ke Xu, Jialin Sun, Shiyue Zhang, Xinyao Jiao, Dingrong and DSP 1.0: Modern Embedded Computing Benchmarks. Computer 58, 5 (2025), Pan, Jie Zhou, Ning Wang, Weiwei Shan, Xinwei Fang, Xi Wang, Nan Guan, and 37ś47. doi:10.1109/MC.2024.3511352 Zhe Jiang. 2025. Uvllm: An automated universal rtl verification framework using