Skip to content
STIMSMITH

SOURCE ARCHIVE

SHA256: defc2e808dfab516d7cac20ed559643c613e0af0a64c90b62a00c86912d21be1
TYPE: application/pdf
SIZE: 739.7 KB
FETCHED: 8/16/2026, 10:15:42 AM
EXTRACTOR: liteparse
CHARS: 68,330

EXTRACTED CONTENT

68,330 chars

0 White Rose eprints@whiterose.ac.uk (¥) Research Online. https://eprints.whiterose.ac.uk Universities of Leeds, Sheffield and York

Deposited via The University of York.

White Rose Research Online URL for this paper:
https://eprints.whiterose.ac.uk/id/eprint/244135/

Version: Accepted Version

Proceedings Paper:
Sun, Jialin, Hu, Yuchen, You, Dean et al. (2026) Prelude: Priming-Guided State
Reconstruction for Efficient FPGA Processor Debugging. In: Design Automation
Conference 2026.

Reuse This article is distributed under the terms of the Creative Commons Attribution (CC BY) licence. This licence allows you to distribute, remix, tweak, and build upon the work, even commercially, as long as you credit the authors for the original work. More information and the full terms of the licence here: https://creativecommons.org/licenses/

Takedown If you consider content in White Rose Research Online to be in breach of UK law, please notify us by emailing eprints@whiterose.ac.uk including the URL of the record and the reason for the withdrawal request.

    A University. . of er 53
    LEEDS| ¢’ Sheffield A J Flt/

UNIVERSITY OF 07

Prelude: Priming-Guided State Reconstruction for Efficient FPGA
    Processor Debugging

    1,2
    Jialin Sun1,2, Yuchen Hu1,2, Dean You1,2, Hui Wang , Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†
    1School of Integrated Circuits, Southeast University, China 2National Center of Technology Innovation for EDA, China
    3Department of Computer Science, University of York, UK

Abstract                                                                              Prelude
Debugging complex FPGA prototypes of modern processors in em-                  FPGA      Tradeoff
bedded systems is challenging due to limited signal visibility and             System  DESSERT[30]      > \
significant tracing overhead. Existing approaches struggle to bal-                Existing Method                                SW-like
ance execution efficiency with debugging capability, often requiring          ||  Well Balanced Method NO)             ENCORE[32] System
either expensive continuous tracing or heavyweight snapshots. We                   Extreme Method                             O
propose Prelude, a lightweight snapshot-based debugging frame-                    Ideal Method Trends      StateMover[31]
work that records only essential architectural states and memory               lower   Debugging Capability                   higher
footprint of the processor on FPGA. During replay, a short visibility         Figure 1. Tradeoff between execution efficiency and debugging capability.
warm-up reconstructs internal micro-architectural states, enabling            typically requires observing thousands of interacting signals. As
cycle-accurate analysis. Implemented on BOOM and Rocket, Pre-                 design complexity increases, the demand for observable signals
lude provides comparable visibility to prior work while significantly         grows accordingly. Moreover, because bugs often occur under un-
improving debugging efficiency: 32.88× / 2191.2× speedup over                 predictable conditions, it is difficult to determine in advance which
DESSERT / ENCORE on BOOM, and 18.09× / 896.4× on Rocket.                      signals are essential for diagnosis. Therefore, higher signal visi-
1 Introduction                                                                bility greatly improves the likelihood of capturing the root cause
Modern processor architectures have become increasingly com-                  when failures arise. Studies indicate that debugging can consume
plex [1ś3], incorporating deep pipelines, out-of-order execution,             60% to 70% of the total verification effort [26ś28], and the limited
and complex cache hierarchies [4ś7]. Verifying these designs re-              waveform visibility on FPGA prototypes further magnifies this bot-
quires executing billions of cycles with real-world workloads such           ~~ tleneck, motivating extensive research on improving observability
as operating system boots and application suites [8ś10]. Traditional          and controllability during FPGA-based debugging.
RTL simulation (e.g., VCS [11], Verilator [12]), while providing full         Existing work. Full-state capture methods have been proposed to
signal visibility, operates at extremely low speeds (∼1kHz), making           achieve complete signal observability for FPGA debugging. One

extensive verification impractical. An IBM study illustrates this ~~ class of approaches relies on scan-chain-based designs (e.g., STATE- limitation, showing that booting a Linux system on an out-of-order ACCESS [29], DESSERT [30]), which insert dedicated observation superscalar processor through such software-based RTL simulation channels for every register and storage element. While this method would take nearly five years to complete [13]. Consequently, FPGA- can, in principle, provide full visibility, it scales poorly in practice. based prototyping has become essential for pre-silicon verification, ~~ Modern CPUs may contain millions of registers and other state offering execution speeds 1000 to 10000 times faster than software elements. Adding scan logic for each of them introduces over 100% simulation and enabling realistic workload testing. additional area overhead [31]. This massive instrumentation quickly However, FPGA-based verification suffers from fundamentally overwhelms FPGA resources, often preventing placement and rout- limited signal visibility, as only a restricted subset of internal states ing from completing. Even when implementation succeeds, the can be observed compared with the full transparency of software resulting system suffers from severe execution slowdown, making simulation. This limitation primarily stems from the reliance on normal program execution nearly impossible and limiting this ap- on-chip logic analyzers such as ILA [14], SignalTap [15], or Chip- proach to very small or simplified designs. Another class of methods Scope [16], which monitor internal signals by inserting additional ~~ (e.g., StateMover [31], ENCORE [32]) employs state readback [33ś probing logic either manually or through synthesis tools [17ś22]. 36], leveraging FPGA hardware features to extract the complete Such instrumentation is highly resource-intensive, consuming sig- internal state through JTAG [37]. JTAG provides only a single low- nificant block RAM, logic elements, and routing resources, and often speed serial channel that typically operates at several hundred kHz disrupts critical timing paths, which may lead to implementation to a few MHz, which makes transferring gigabytes of state data failures or prevent the design from meeting frequency targets [23ś extremely time-consuming. During readback, the running design 25]. As a result, only a minute fraction of internal signals (typically on the FPGA must be halted, and gigabytes of state data are trans- less than 0.01% of the total) can be traced within a short observa- ferred to the host, introducing substantial latency. Even when using tion window, offering only partial insight into system behavior. a high-speed PCIe interface, reading back the complete state re- This limited visibility is far from sufficient for debugging complex ~~ mains costly (e.g., it takes around 10s to read back a crossbar-big designs such as CPUs, where understanding erroneous behaviors ~~ design, which is orders of magnitude smaller than an out-of-order superscalar processor [31].). In addition, debugging requires repeat- †Corresponding author: Zhe Jiang. Email: zhejiang.uk@gmail.com. edly capturing system snapshots throughout execution to monitor Dac’26, Long Beach, CA, USA state evolution, leading to frequent interruptions and excessive data 2026. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM transfer overhead. This prevents continuous execution and makes https://doi.org/10.1145/3770743.3804273 it infeasible to debug long-running programs or workloads.

Execution Efficiency lower higher

      1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA        Jialin Sun         , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†

             In summary, existing FPGA debugging methods face a funda-
mental dilemma (Figure 1), as full-state approaches are costly while             90.0%
partial-state approaches offer only limited insight, preventing both             85.0%
high execution efficiency and comprehensive visibility.
Contributions. To address this challenge, we propose Prelude, a                  80.0%
lightweight snapshotśbased debugging framework that achieves                     75.0%
near-full signal visibility through replay. Instead of continuously              70.0%
tracing signals or capturing the entire design state, our method                     10 100             1000     10000     100000    1000000
records only a compact snapshot containing essential architectural                                      Instructions
states and partial signal information on FPGA. During replay on a            Figure 2. Comparison of similarity between normal execution and partial-
software simulator, a short visibility warm-up sequence reconstru-           state reconstruction across executed instructions.
cts most internal signals from these partial states before execut-           summarized in Table 1, reproducing a segment correctly requires
ing the error segment, enabling cycle-accurate debugging with                less than 0.001% of the full-state snapshot, yet achieves identical
minimal performance interference. We implement and evaluate                  functional replay fidelity.
this framework on BOOM [38], an open-source out-of-order super-               Key Insight. Functionally correct replay can be achieved us-
scalar RISC-V processor, and Rocket [39], a five-stage in-order core,         ing only architectural registers and accessed memory, without
achieving comparable debugging visibility with 32.88× / 2191.2×               preserving the entire hardware state.
speedup over DESSERT [30]/ ENCORE [32] on BOOM and 18.09× /
896.4× on Rocket, demonstrating both efficiency and portability.             2.2 Error Trigger Capability.
2   Design Philosophy                                                        When debugging real-world processor bugs, engineers rarely need
Full-state capture provides a compelling demonstration that software-        to observe all internal signals. Typically, only a subset of a few
level replay can enable detailed debugging if the hardware state             hundred to a few thousand relevant signals is monitored, rather
is correctly reproduced. However, directly applying this approach            than the complete set of registers and logic structures. Similarly,
to large out-of-order processors is impractical due to the prohib-           triggering the same bug during replay does not require the entire
itive cost of capturing and transferring the complete state. This            micro-architectural state, just like on the FPGA where the bug
observation highlights two essential requirements for practical              occurs; for example, an ALU-related bug is independent of the state
replay-oriented debugging: the captured state must support (i)               of unrelated units such as the FPU.
functionally correct replay and the system must retain sufficient             Building on this insight and leveraging the transient nature
information to maintain (ii) error trigger capability.                       of micro-architectural states (Section 2.1), our method executes a
2.1 Functionally Correct Replay.                                             short warm-up sequence of instructions prior to the segment of
Table 1. Comparison of full snapshot contents and our method (✓ = required   interest. This warm-up allows the micro-architectural structures to
by our method, ✗ = not collected)                                            converge to their original states, effectively reconstructing over 90%
                                                                             of the necessary state for error triggering (Figure 2). By combining
     Category   Full Snapshot [31]                   Our Method              selective signal capture with targeted warm-up, we achieve high
        Architectural registers (PC, GPRs,  Architectural registers (PC,     debug fidelity without capturing the entire hardware state1.
                   CSRs, FPRs)                  GPRs, CSRs, FPRs) ✓
                Pipeline registers                       ✗                    Key Insight. Error triggering depends only on a limited subset
   Micro-arch.  Instruction queues                       ✗                    of the micro-architectural state, which can be naturally recon-
      State      Reorder buffers                         ✗                    structed through a short warm-up execution.
                Load/store queues                        ✗
                Branch predictors                        ✗                   3
              Other internal buffers                     ✗                       Framework
       Size                                                                  Based on the analysis in Section 2, we design our framework, Pre-
  (micro-arch.)   23.4MB (full)                     1KB (0.004%)             lude, to meet these requirements. We employ lightweight architec-
      Memory     L1/L2/L3 caches                         ✗                   tural snapshots combined with an online error detection mechanism
      System     Full main memory                        ✗                   based on differential testing (Section 3.1) to identify bugs during
  Size (memory)  Accessed memory                 Accessed memory ✓           FPGA execution with minimal overhead. To enable deterministic
   Size (total)     4GB (full)                   17.59KB (0.0004%)           replay from these snapshots, we augment the captured state with
                  4.023GB (full)                 18.6KB (0.00046%)           memory system reconstruction (Section 3.2), allowing a software
  From the perspective of software execution, a program can                  simulator to replay the execution with full micro-architectural visi-
only observe and operate on architectural registers, benefiting              bility (Section 3.3). Section 3.4 presents the complete workflow.
from the architectural abstraction that hides micro-architectural            3.1 Differential Testing
details. These architectural registers remain stable and fully de-           Differential testing [4, 27, 32, 40ś43] has proven efficient for detect-
termine the instruction behavior within a segment, while other               ing functional errors in complex processor designs by executing
micro-architectural structures are transient and evolve from them.           the design under test (DUT) alongside a golden reference model,
    Therefore, accurate segment replay only requires the architectural       usually using the instruction set architecture (ISA) simulator [44],
portion rather than complete micro-architectural snapshot.                   and comparing their architectural states at synchronization points.
  Similarly, for the memory system, only the memory locations ac-            Any mismatch indicates a potential bug and triggers detailed state
cessed within the segment are relevant to its replay. Other memory
contents remain unused and thus have no effect on execution. As              1
                                                                              The setup for the experiment is described in Section 4.

Reproduction Rate

Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging                                                    Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA

         FPGA                                          Host PC                                                                                           Debugging
        DUT Processor     Differential Testing     a       Software Simulator                                                      c   Segment                            e
        Memory      Arch.       Log ISA Simulator                      Snapshot                                                                ...
        Footprint  Register     ...   Snapshot                           GPRs                       DUT Processor                      bez x4,74
                                                                         FPRs                                                          add x3,x5            PC
         Snapshot  PCIe         ...   Compare          Memory            CSRs                      DEU      AMU        MRU                     ...
                                                       Rebuild          Memory                                                                              Data Data Data

                                 Memory Footprint                             Memory Rebuild         b        FPGA                   SW Sim.   Replay                 Hardware
                                Load  Addr1   Data1                                                               1
        Read after Write       Store  Addr2   Data2                           Initial Memory                                           Arch. State       d            Software
        Overlapping Load        Load  Addr2   Data2    Analyze Addr1                   Data1                      2  Replay             2      Warm‐up                Log Message
                                Load  Addr3   Data3            Addr3                   Data3                                                                          FPGA Thrd.
                                Load  Addr4   Data4            Addr5                   Data5                      3                     3     Micro‐Arch.
         Inclusive Load         Load  Addr5   Data5             ...                    ...                                Error                State                  SW Sim. Thrd.
                                ...    ...    ...                                                                      Detect!                                        Add Module

Figure 3. An overview of Prelude. (GPRs: General Purpose Registers; FPRs: Floating Point Registers; CSRs: Control and Status Registers; DEU: Data Extract Unit; AMU: Architectural Register Management Unit; MRU: Memory Rebuild Unit). a Differential testing identifies errors on the FPGA by comparing snapshots from the DUT and an ISA simulator. b Upon detecting an error, the memory footprint is analyzed to reconstruct the initial memory state. c The architectural registers and reconstructed memory are then used for software replay. d The DUT starts from the captured architectural state and replays instructions forward to regenerate the micro-architectural state. e The error-containing segment is replayed with nearly full signal visibility for debugging. replay. Previous work [32, 41] transmit complete architectural state Algorithm 1: Memory Footprint Analysis at every comparison, incurring substantial communication over- Input: address range 𝐼 = [𝐴 head for processors with large register files. Prelude addresses this set W 𝑠, 𝐴𝑒 ] with data 𝐷; read set R; write by employing lightweight snapshots (Figure 3 a ) that are sufficient Output: Updated read set R′ 1 foreach ( [𝑅 to reflect the execution of the DUT processor, while more detailed 𝑠 Inclusive Load Ranges 2 // 𝑖 , 𝑅𝑒𝑖 ], 𝐷𝑖 ) ∈ R do micro-architectural states are reconstructed from these snapshots. 3 if 𝐼 ⊇ [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ] then Snapshot. In Prelude, a snapshot captures only the architectural 4 R ← R \ { ( [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ], 𝐷𝑖 ) } ; // remove old load state and the memory locations accessed within the segment (Sec- 5 end 𝐷 [𝑅𝑠𝑖 ,𝑅𝑒𝑖 ] ← 𝐷𝑖 ; // copy data into current load tion 2.1), following the design philosophy that the software-visible 6 // Overlapping Load Addresses state fully determines execution behavior. 7 𝑖 , 𝑅𝑒𝑖 ] ∩ 𝐼 ≠ ∅ then 8 else if [𝑅 • Architectural Registers. All general-purpose registers 𝐼 𝐼 𝑠 9 ← \ ( [𝑅𝑠𝑖 , 𝑅𝑒𝑖 ] ∩ 𝐼 ) ; // remove overlap (GPRs), program counter (PC), floating-point registers (FPRs) 10 end 𝐷 ← 𝐷 |𝐼 ; // keep non-overlapping data and relevant control and status registers (CSRs). These are 11 end sufficient to reproduce instruction execution within the seg- 12 // Read-after-write Dependencies 13 ment, as other micro-architectural structures, e.g., pipeline 14 foreach [𝑊𝑠𝑗,𝑊𝑒𝑗 ] ∈ W do buffers, evolve in a predictable manner from them. 15 if 𝐼 ⊆ [𝑊𝑠𝑗,𝑊𝑒𝑗 ] then • Accessed Memory Footprint. Only memory addresses 16 end 𝐼 ← ∅ ; // discard load accessed by instructions within the segment are included. 17 end 18 This ensures that memory contents, which do not affect 19 if 𝐼 ≠ ∅ then execution, are excluded, minimizing snapshot size. 20 R′ ← R ∪ { (𝐼, 𝐷 ) }; // add remaining load to read set Snapshots are transmitted from FPGA to host PC. During differ- 21 end ential testing, the reference model executes the same program to is analyzed (Figure 3 b ), with particular attention to the addresses the corresponding synchronization points, where its architectural state is compared against the DUT. Any mismatch triggers further and data of all load operations. As illustrated in Algorithm 1, special analysis or additional snapshot collection. By restricting snapshots handling is required for several cases that may affect correctness: to the minimal set of software-visible state, Prelude significantly Read-after-write Dependencies. If a load reads from an address reduces storage and communication overheads to less than 0.001% that has already been written within the same segment, both the of a full-state snapshot, while preserving the ability to detect errors. address and the data of this load are discarded, since the value is 3.2 Memory Rebuild internally defined and does not depend on the initial memory state. Overlapping Load Addresses. When two load operations partially During checking, all memory accesses within a segment are recorded. overlap in their accessed address ranges, only the portion of the For replay, however, only the memory state at the segment’s be- later access that covers previously unread bytes is retained. The ginning is required. The memory rebuild process bridges this gap overlap region is removed, and the effective access range of the by analyzing the recorded memory footprint to reconstruct the preserved load is narrowed accordingly. necessary initial memory state for the segment’s execution. Inclusive Load Ranges. If a later load covers a larger address range To identify the required initial memory state, we focus on the first that includes one or more earlier accesses, the earlier and smaller read of each memory address in the memory footprint, provided loads are removed. For the overlapping region, the data in the later that the address has not been written before the read occurs. In load are replaced with the corresponding bytes from the earlier this process, the complete memory footprint obtained from FPGA accesses to ensure consistency with the original access order.

                                                                                                                                      1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA                                                                             Jialin Sun   , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†

                                                                                              Data Extract Unit (DEU)                       the Load-Store Queue (LSQ) to clear any in-flight operations, then
       DF           start
                     !end       State⁰ˣ⁰¹   +       |       Mem_start!Mem_end     0x03          +                                           restores the memory footprint by injecting data through the DCache
                           Arch.                                       Memory                         0x01                                  interface according to the reconstructed snapshot. After both the
                  0x1F     Calculation                      Mem_size   Calculation            0x02                                 a        AMU and MRU finish register and memory initialization, the DUT
                                                                                                                                            resumes execution from the replay segment’s starting PC. A brief
      Arch. Reg Management Unit (AMU)
                            5             64bits                       Clr.(MRU)lsq   FIFO    64bits           2bits       64bits           warm-up phase then allows internal queues, buffers, and predictors
      Start Condition           bits                   |        Memory Rebuild Unit
      flush Clr. rob        A_Addr         Data                                               Addr             Size        Data             to settle, ensuring that replay proceeds under a consistent and
                            0x01           ...                                                Addr1            Size1       Data1            reproducible micro-architectural context.
       Ctrl.Init            0x02           ...                                                Addr2            Size2       Data2
                                      ...              b                                                       ...                 c        3.4 Workflow
           Redirect_PC                                                                                                                      In this subsection, we present the overall workflow of our FPGA
                            Instruction Fetch Unit (IFU)                                     Re‐Order Buffer (ROB)                          debugging framework (Figure 3). The process consists of captur-
      Next_PC                                                                                 |
                                                                                      Clr. rob Files (PRFs)    Clr. lsq                     ing lightweight snapshots during normal execution, performing
       Ctrl.Init           PRFs Controllers                            Physical Register

                  5bits    7bits                            Addr             6bits    64bits          Micro                                 differential testing to detect errors, and reconstructing both archi-
                  A_Addr   P_Addr       Write_Ctrl[0]       Data              Addr             Data   ‐ops             Load
                  0x01     0x04       Arch. andPhy.                           0x00             ...                    Store    §Addr        tectural and memory states for replay (Figure 5). If a replay fails
                  0x02     0x07       RegisterAddr.                           0x01             ...    wb               Unit     Data        to reproduce a bug, a full-state snapshot is captured as a fallback.
                     ...                Write_Ctrl[2]                         0x02             ...    Micro                                 This workflow balances visibility and efficiency, minimizing per-
                  0x1d     0x24                                               0x03             ...    ‐ops
                  0x1e     0x2f                                               ...              ...    wb          Execute                   formance overhead while ensuring accurate debugging.
                  0x1f     0x01         Write_Ctrl[3]                         0x7f             ...             =      Units                                     Compare
                  Map Table                                                                                                                   ISA Sim.   1    Wait 2 Wait  3   Wait  Error Detect            REF    Snapshot
                                                                                                                                              FPGA             Snapshot                  Warm up,            DUT    SW Sim.
                                                            DCache                                                                                         1        2        3       reproduce micro‐arch.
                                                                                                                                                               Replay                Arch.      Micro‐arch.         Bug Found
                                                                                                                                              SW Sim.                                    2                   3
Figure 4. The micro-architecture of the DUT processor in software simula-                                                                                                                                               time
tor (DF: Data Fetch; blue: Architectural Register Initialization; red: Memory                                                                     5        10 15   20  25    30  35  40     45     50 55  60  65  70
State Initialization; green: Control Signal).                          a DEU fetches data from log file                                     Figure 5. FPGA debugging workflow timeline, showing snapshot capture,
and transmits to AMU and MRU.                               b     AMU completes architectural register                                      differential testing, replay, warm-up and debugging.
initialization and            c                                                  MRU completes memory state initialization.                 Error Detection. The DUT on the FPGA is executed in parallel with
3.3   Micro-architecture Add-on                                                                                                             the ISA simulator on the host PC (Figure 3 a ). Their architectural
A general, non-intrusive mechanism reconstructs memory and ar-                                                                              states are periodically compared at predefined synchronization
chitectural states from lightweight logs to enable deterministic re-                                                                        points. Any mismatch between the two indicates a potential func-
play and consistent micro-architectural execution. This mechanism                                                                           tional error, which triggers the subsequent debugging workflow.
is demonstrated on BOOM, an out-of-order superscalar RISC-V                                                                                 Snapshot Processing. After an error is detected, the software snap-
processor, as a representative case study (Figure 4).                                                                                       shot is analyzed to reconstruct only the minimal memory state
Data Extract Unit (DEU, Figure 4 a ). The DEU serves as the                                                                                 needed for replay (Figure 3 b ), which is then loaded into the simu-
bridge between snapshot processing and hardware replay. It re-                                                                                           ~~ lator to accurately reproduce the failing segment (Figure 3 c ).
trieves architectural register values and reconstructed memory con-                                                                         Visibility Warm-up. In this stage, the DUT executes a short in-
tents from the snapshot log and transfers them to the simulation                                                                            struction sequence in the simulator to propagate architectural state
environment via the standard DPI-C interface. After parsing, the                                                                            into micro-architectural state (Figure 3 d ), ensuring all internal
DEU routes register data to the Architectural Register Management                                                                           signals are properly initialized for consistent replay and debugging.
Unit (AMU) and memory data to the Memory Rebuild Unit (MRU),                                                                                Debugging. Starting from the prepared micro-architectural state,
enabling precise initialization of both subsystems. By decoupling                                                                           the erroneous segment is fully replayed in the software simulator,
data extraction from replay control, the DEU provides a clean and                                                                           where all signals and architectural states are visible for fine-grained
efficient path for propagating snapshot information, ensuring that                                                                          inspection and accurate root-cause identification (Figure 3 e ).
deterministic replay begins from the correct architectural state.                                                                           Fallback. If the bug cannot be successfully reproduced during
Architectural Register Management Unit (AMU, Figure 4 b ).                                                                                  replay, the FPGA must be re-run to capture a full-state snapshot.

The AMU reconstructs the processor’s architectural state for replay ~~ In this case, execution only needs to proceed up to the start of the after the DEU provides the snapshot data. It first resets the pipeline erroneous segment, at which point all internal signals are taken. by flushing the ROB and returning all physical registers to the free Intermediate differential testing and non-essential snapshots are list, ensuring a clean architectural configuration without specula- ~~ skipped, avoiding unnecessary overhead while still providing the tive or in-flight state. Using BOOM’s map table, the AMU writes complete state required for subsequent replay and debugging. each architectural register value to its assigned physical register and 4 Evaluation redirects the Instruction Fetch Unit (IFU) to the replay segment’s We evaluate Prelude on the BOOM, an open-source out-of-order starting program counter, reconstructing the processor state ac- superscalar RISC-V processor. The design is implemented on an cording to the snapshot. After this reset-and-restore procedure, the AMD Virtex UltraScale+ VU19P FPGA [45] using Xilinx Vivado AMU waits for the MRU to complete memory initialization so that 2024.2. For functional reference, we use Spike as the ISA simulator, replay can begin from a fully consistent architectural state. while software-level replay and verification are performed with Memory Rebuild Unit (MRU, Figure 4 c ). The MRU reconstructs Verilator. We use CoreMark [46] and Embench [47] benchmark the initial memory state required for replay. It begins by flushing suites to measure debugging performance metrics.

Snapshot

                                                                                              ...
    
    
    
    
    
    
    
    
    
    
                                                                                              ......
    

Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA

Verilator[12]                                                        ENCORE[32]    DESSERT[30]   Prelude(Ours)

104

103

102

10

1

          Figure 6. Debugging performance results for Prelude, Verilator [12], ENCORE [32] and DESSERT [30], running CoreMark and Embench.
    90.0%           100       1,000      10,000      100,000        1,000,000      100.0%   100  1,000      10,000    100,000    1,000,000
    85.0%                                                                           95.0%
    80.0%                                                                           90.0%
    75.0%                                                                           85.0%
    70.0%                                                                           80.0%
    65.0%               Fetch      Decode      Rename                               75.0%   Register-read   Dispatch             Issue
                         (a) Front-end Reproduction Rate.                                        (b) Mid-end Reproduction Rate.
    100.0%          100      1,000      10,000      100,000      1,000,000         90.0%    100     1,000   10,000    100,000    1,000,000
           95.0%                                                                   85.0%
           90.0%                                                                   80.0%
           85.0%                                                                   75.0%
           80.0%        Execute           LSU      ROB                             70.0%     DCache          L2Cache             Bus
                         (c) Back-end Reproduction Rate.                                     (d) Memory System Reproduction Rate.
    Figure 7. Signal reproduction rates of different hardware module groups at various pipeline stages, under varying instruction counts.
4.1 Evaluation Metrics                                                            efficiency. We compare Prelude against three representative base-
To quantify the debugging efficiency of our framework, we define                  lines: (1)Verilator [12], a software simulator which is capable of full
the effective debugging time 𝑇eff as:                                            waveform tracing and provides the highest visibility but extremely

                                                                                  bugging via readback of internal configuration and state data; and
            𝑇eff   = 1  𝑁 (𝑇exe,𝑖 + (1 − 𝑃err,𝑖 ) · 𝑇Penalty,𝑖 )  (1)     low execution speed; (2) ENCORE [32], which performs FPGA de-
                    𝑁 𝑖=1                                                       (3) DESSERT [30], which uses scan-chain based state extraction
 Here, 𝑇exe,𝑖           denotes the FPGA execution time required to reach a     for post-mortem analysis. All baseline approaches rely on either
debuggable state, 𝑃err,𝑖      is the probability that a bug is successfully     full-state or near-complete signal retrieval, leading to substantial
triggered and diagnosed within the segment, and 𝑇Penalty,𝑖           repre-     runtime and data transfer overhead during debugging.
sents the extra time consumed if the bug is not detected. 𝑁 is the                 Across all workloads, Prelude achieves substantial improvements
number of checkpoints. This metric reflects both the execution                    in debugging efficiency, with a geomean speedup of 2191.2× over
efficiency and the likelihood of effective error localization.                    ENCORE and 32.88× over DESSERT. The maximum speedups reach
 We further report the normalized slowdown as:                                    5714.6× and 73.94×, respectively. Even for very short workloads (e.g.,
                           slowdown = 𝑇𝑇others                          (2)     nbody, st), Prelude still achieves 66.67× and 2.11× improvements.
                                          ours                                    The smaller gains on short workloads occur because the warm-up
            where 𝑇others and 𝑇ours represent the effective debugging time of   phase runs on a software simulator, which is slower than the FPGA.
baseline methods and our system, respectively.                                    The warm-up cost is fixed, so when the FPGA execution segment
 We additionally define the reproduction rate as:                                 becomes short, the warm-up occupies a larger fraction of the total
                    𝑅𝑒𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑅𝑎𝑡𝑒 = 𝑆match           (3)     runtime. For longer workloads (e.g., nettle-aes and nettle-sha256),
                                              𝑆total                             the FPGA segment takes most of the total time, the relative impact
                                                                                  of the warm-up decreases, and the performance improves.
               where 𝑆match is the number of micro-architectural signals that                   Table 2. Time Breakdown Ratios
match the ground-truth execution after replay, and 𝑆total is the total
number of compared signals. A higher value indicates that the                                  Execution    Snapshot  Warm-up        Fallback
replayed micro-architectural state agrees with the actual hardware                  Ratio (%)    65.25         12.20  13.48          9.07
state and supports reliable error analysis.
4.2 Results and Analysis                                                          Performance Breakdown.To better analysis Prelude’s internal
Result#1: Outperforming prior tools in debugging perfor-                          performance characteristics, we further break down its end-to-end
mance. Figure 6 shows the normalized slowdown when running                        debugging time into four components: FPGA execution, snapshot
CoreMark and the Embench benchmark suite under different debug-                   transfer, warm-up reconstruction, and replay fallback. As shown
ging frameworks. A smaller slowdown indicates higher runtime                      in Table 2, this breakdown shows that the dominant cost lies in

Reproduction Rate Reproduction Rate Slowdown

Reproduction Rate Reproduction Rate

                                                                                             sglib-combined slre      st  ud     wikisort Geo.Mean
    coremarkaha-mont64   crc32 cubic  edn huffbenchmatmult-int minver nbody nettle-aesnettle-sha256 nsichneu picojpeg qrduino

               1,2
Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA        Jialin Sun       , Yuchen Hu1,2, Dean You1,2, Hui Wang1,2, Yushu Du2, Xinwei Fang3, Zhe Jiang1,2†

           ENCORE[32]                      DESSERT[30]  Ours                               Table 3. Area overhead of BOOM and Rocket on FPGA.
    103    BOOM        Rocket                                                 Core     Resource      Pure     With   Absolute  Relative
                                                                                                     Core   Prelude  Overhead  Overhead
    100                                                                                Logic LUTs  853,622  879,230   25,608    3.00%
     10                                                                       BOOM     LUTRAMs      26,749   27,328    579      2.16%
                                                                                       Flip-Flops  477,416  491,818   14,402    3.02%
      1                                                                                Logic LUTs  201,402  211,136   9,734     4.83%
Figure 8.CoreMark Embench                     CoreMark  Embench               Rocket   LUTRAMs      22,816   23,306    490      2.15%
    Debugging performance evaluation on BOOM and Rocket                                Flip-Flops  271,477  280,255   8,778     3.23%
the fast FPGA-execution portion, while snapshot and warm-up                Result#4: Efficient integration into diverse RISC-V cores with
overheads remain small and stable, validating the efficiency of            very low area cost. The area overhead of integrating Prelude is
Prelude’s snapshot-plus-warm-up design.                                    evaluated on BOOM and Rocket cores using FPGA implementations.
Result#2: Effectiveness of reproduction across all pipeline                Table 3 reports the utilization of logic LUTs, LUTRAMs, and flip-
stages. Figure 7 reports the signal reproduction rate of each hard-        flops for each core. Compared to the baseline designs, Prelude
ware module in pipeline stage under varying instruction counts.            introduces only minor resource increases. These results indicate
Overall, the geomean rises from 86.7% at 102 instructions to 91.5%         that Prelude can be efficiently incorporated into different RISC-V
at 106            (Figure 2), showing that longer workloads provide more   cores without significant hardware cost.
complete context for reconstructing micro-architectural state.
Front-end (Figure 7(a)). The Fetch module has a lower repro-               4.3 Case Study
duction rate due to dynamic structures like the branch predictor           During the evaluation of Prelude, we identified a previously un-
requiring long execution history, while Decode and Rename achieve          known bug on BOOM’s floating-point (F) extension. Certain in-
higher rates as their states mainly depend on current instructions,        structions produced incorrect results due to incomplete implemen-
which can be reconstructed from short segments.                            tation. While running our framework, a mismatch was detected at
Mid-end (Figure 7(b)). Similar to Decode and Rename, the Register-         a checkpoint in the program counter (PC). Using the state saved at
read, Dispatch, and Issue modules all achieve high reproduction            the previous checkpoint, we performed a warm-up to restore the
rates, as their states mainly depend on current instructions and can       micro-architectural state before executing the erroneous segment.
be reconstructed from short segments.                                      1  0x80002390:  sd     s4 ,0( gp )
Back-end (Figure 7(c)). The ROB module reaches very high re-               2  0x80002394:  fcvt.l.s a2 , fs11
production rates because most of its signals store instruction in-         3  0x80002398:  auipc   tp ,0x0
formation, with only a small portion representing dynamic control                          Listing 1. Instructions from the erroneous segment
signals. The LSU module maintains relatively stable reproduction              Full-waveform debugging of the segment revealed that when
rates, as its state reflects memory access patterns that are partially     execution reached address 0x80002394 (Listing 1), the reorder buffer
dynamic but largely deterministic over short instruction segments.         (ROB) finished the instruction, but the corresponding exception was
Memory System (Figure 7(d)). The L2Cache module has relatively             not committed. Step-by-step analysis of the waveform showed that
low reproduction rates because it relies on complex caching policies       the decode unit did not have an implementation for this instruction,
and dynamic replacement decisions, which require long execution            causing it to be flagged as an exception. After identifying the root
history to reconstruct accurately. This low reproduction rate does         cause, we confirmed that executing the corrected instruction in
not affect bug triggering, because errors depend primarily on archi-       BOOM produced results consistent with the expected model. Using
tectural memory operations rather than detailed L2Cache control            Prelude, the entire bug discovery and debugging process, including
states. The DCache module improves moderately with longer in-              state restoration and warm-up, was completed in under three min-
struction segments, similar to other cache-related modules. The Bus        utes. In comparison, performing the same analysis with Verilator
module remains stable across all instruction counts, as its signals        would have taken roughly 50 hours, roughly 1, 000 times longer,
mainly represent deterministic data transfers and arbitration states.      demonstrating the significant efficiency advantage of Prelude.
Result#3: Effortlessly portable between BOOM and Rocket.                   5
Figure 8 shows the normalized slowdown of Prelude when ported                  Conclusion
from BOOM to Rocket [39] using the same CoreMark and Embench               We have presented Prelude, a lightweight and portable FPGA de-
      workloads. Despite differences in pipeline depth, execution model,   bugging framework that achieves near-full visibility without full-
and micro-architectural state, Prelude requires no design-specific         state capture. By combining compact snapshots with a short replay
tuning beyond connecting replay units. Rocket lacks a reorder              warm-up, Prelude reconstructs the required micro-architectural
buffer and physical register files, so initialization only flushes the     state for cycle-accurate debugging while imposing minimal over-
pipeline and writes values directly into logical registers, making         head. Overall, Prelude offers a practical and efficient approach for
migration straightforward while preserving framework generality.           debugging modern FPGA-based processors.
             Across both benchmarks, Prelude maintains a clear debugging   6
advantage over ENCORE and DESSERT on Rocket. The gap narrows                   Acknowledgement
compared with BOOM due to Rocket’s lower IPC increasing execu-             We’d like to thank the reviewers for the helpful feedback. This
tion time, yet Prelude still achieves 27.85× and 18.10× speedups on        work is supported by the National Natural Science Foundation of
CoreMark and Embench, demonstrating that lightweight snapshots,            China (No. 62472086, 92464204), the Science and Technology Major
deterministic replay, and warm-upśguided reconstruction provide            Special Program of Jiangsu (No. BG2024010), and the Fundamental
a portable, efficient debugging solution across RISC-V cores.              Research Funds for the Central Universities (No. 2242025K20013).

Slowdown

Prelude: Priming-Guided State Reconstruction for Efficient FPGA Processor Debugging Dac’26, July 26śJuly 29, 2026, Long Beach, CA, USA

References llms. Proceedings of the IEEE/ACM Design Automation Conference (DAC) (2025). [1] Bryan H Fletcher. 2005. FPGA embedded processors. In Embedded Systems [27] Jialin Sun, Yuchen Hu, Dean You, Yushu Du, Hui Wang, Xinwei Fang, Weiwei Conference. 18. Shan, Nan Guan, and Zhe Jiang. 2025. ISAAC: Intelligent, Scalable, Agile, and [2] Fuming Sun, Xiaoying Li, Qin Wang, and Chunlin Tang. 2008. FPGA-based Accelerated CPU Verification via LLM-aided FPGA Parallelism. arXiv preprint embedded system design. In APCCAS 2008-2008 IEEE Asia Pacific Conference on arXiv:2510.10225 (2025). Circuits and Systems. IEEE, 733ś736. [28] Ke Xu, Jialin Sun, Yuchen Hu, Xinwei Fang, Weiwei Shan, Xi Wang, and Zhe [3] Michael Gschwind, Valentina Salapura, and Dietmar Maurer. 2002. FPGA proto- Jiang. 2024. Meic: Re-thinking rtl debug automation using llms. In Proceedings of typing of a RISC processor core for embedded applications. IEEE Transactions on the IEEE/ACM International Conference on Computer-Aided Design (ICCAD). 1ś9. Very Large Scale Integration (VLSI) Systems 9, 2 (2002), 241ś250. [29] Dirk Koch, Christian Haubelt, and Jürgen Teich. 2007. Efficient hardware check- [4] Nursultan Kabylkas, Tommy Thorn, Shreesha Srinath, Polychronis Xekalakis, pointing: concepts, overhead analysis, and implementation. In Proceedings of and Jose Renau. 2021. Effective processor verification with logic fuzzer enhanced the 2007 ACM/SIGDA 15th international symposium on Field programmable gate co-simulation. In MICRO-54: 54th Annual IEEE/ACM International Symposium on arrays. 188ś196. Microarchitecture. 667ś678. [30] Donggyu Kim, Christopher Celio, Sagar Karandikar, David Biancolin, Jonathan [5] Chloe Tain, Savita Patil, and Hussain Al-Asaad. 2025. Survey of Verification of Bachrach, and Krste Asanović. 2018. DESSERT: Debugging RTL effectively RISC-V Processors. Journal of Electronic Testing (2025), 1ś28. with state snapshotting for error replays across trillions of cycles. In 2018 28th [6] Eric Sprangle and Doug Carmean. 2002. Increasing processor performance by International Conference on Field Programmable Logic and Applications (FPL). IEEE, implementing deeper pipelines. ACM SIGARCH Computer Architecture News 30, 76ś764. 2 (2002), 25ś34. [31] Sameh Attia and Vaughn Betz. 2020. StateMover: Combining simulation and [7] Jin Li, Kristin Tufte, Vladislav Shkapenyuk, Vassilis Papadimos, Theodore John- hardware execution for efficient FPGA debugging. In Proceedings of the 2020 son, and David Maier. 2008. Out-of-order processing: a new architecture for ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. 175ś high-performance stream systems. Proceedings of the VLDB Endowment 1, 1 185. (2008), 274ś288. [32] Kan Shi, Shuoxiang Xu, Yuhan Diao, David Boland, and Yungang Bao. 2023. [8] Google. 2019. RISC-V DV. https://github.com/google/riscv-dv. ENCORE: Efficient architecture verification framework with FPGA accelera- [9] Sadullah Canakci, Chathura Rajapaksha, Leila Delshadtehrani, Anoop Nataraja, tion. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field Michael Bedford Taylor, Manuel Egele, and Ajay Joshi. 2023. Processorfuzz: Programmable Gate Arrays. 209ś219. Processor fuzzing with control and status registers guidance. In 2023 IEEE In- [33] Hari Angepat, Gage Eads, Christopher Craik, and Derek Chiou. 2010. NIFD: ternational Symposium on Hardware Oriented Security and Trust (HOST). IEEE, Non-intrusive FPGA DebuggerśDebugging FPGA’Threads’ for Rapid HW/SW 1ś12. Systems Prototyping. In 2010 International Conference on Field Programmable [10] Chen Chen, Rahul Kande, Nathan Nguyen, Flemming Andersen, Aakash Tyagi, Logic and Applications. IEEE, 356ś359. Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. 2023. HyPFuzz:Formal- [34] Ashfaquzzaman Khan, Richard Neil Pittman, and Alessandro Forin. 2010. gNOSIS: Assisted processor fuzzing. In 32nd USENIX Security Symposium (USENIX Security A board-level debugging and verification tool. In 2010 International Conference 23). 1361ś1378. on Reconfigurable Computing and FPGAs. IEEE, 43ś48. [11] Synopsys. 2024. VCS: Synopsys Verification Continuum. https://www.synopsys. [35] Changgong Li, Alexander Schwarz, and Christian Hochberger. 2016. A readback com/verification/simulation/vcs.html based general debugging framework for soft-core processors. In 2016 IEEE 34th [12] Wilson Snyder. 2024. Verilator. https://www.veripool.org/wiki/verilator International Conference on Computer Design (ICCD). IEEE, 568ś575. [13] Sameh Asaad, Ralph Bellofatto, Bernard Brezzo, Chuck Haymes, Mohit Kapur, [36] Georgios Tzimpragos, Da Cheng, Stephanie Tapp, Balakrishna Jayadev, and Benjamin Parker, Thomas Roewer, Proshanta Saha, Todd Takken, and José Tierno. Amitava Majumdar. 2016. Application debug in FPGAs in the presence of multiple 2012. A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating asynchronous clocks. In 2016 International Conference on Field-Programmable multi-core processor simulation. In Proceedings of the ACM/SIGDA international Technology (FPT). IEEE, 189ś192. symposium on Field Programmable Gate Arrays. 153ś162. [37] Stephanie Tapp. 2015. XAPP1230: Configuration Readback Capture in UltraScale [14] Xilinx. 2016. Integrated Logic Analyzer v6.2. https://www.xilinx.com/support/ FPGAs. documentation/ip_documentation/ila/v6_2/pg172-ila.pdf [38] Jerry Zhao, Ben Korpan, Abraham Gonzalez, and Krste Asanovic. 2020. Sonic- [15] Intel. 2020. Intel Quartus Prime Pro Edition User Guide: Debug boom: The 3rd generation berkeley out-of-order machine. In Fourth Workshop on Tools. https://www.intel.com/content/dam/www/programmable/us/en/pdfs/ Computer Architecture Research with RISC-V, Vol. 5. International Symposium on literature/ug/ug-qpp-debug.pdf Computer Architecture Valencia, Spain, 1ś7. [16] Xilinx. 2012. Chipscope Pro Software and Cores: User Guide. https://docs.amd. [39] Krste Asanovic, Rimas Avizienis, Jonathan Bachrach, Scott Beamer, David Bian- com/v/u/en-US/chipscope_pro_sw_cores_ug029 colin, Christopher Celio, Henry Cook, Daniel Dabbelt, John Hauser, Adam Izraele- [17] Eddie Hung and Steven JE Wilton. 2012. Scalable signal selection for post-silicon vitz, Sagar Karandikar, Ben Keller, Donggyu Kim, and John Koenig. 2016. The debug. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 21, 6 rocket chip generator. EECS Department, University of California, Berkeley, Tech. (2012), 1103ś1115. Rep. UCB/EECS-2016-17 4 (2016), 6ś2. [18] Jeffrey Goeders and Steven JE Wilton. 2016. Signal-tracing techniques for in- [40] Kunlin You, Yinan Xu, Kehan Feng, Luoshan Cai, Yaoyang Zhou, and Yungang system FPGA debugging of high-level synthesis circuits. IEEE Transactions on Bao. 2025. DiffTest-H: Toward Semantic-Aware Communication in Hardware- Computer-Aided Design of Integrated Circuits and Systems 36, 1 (2016), 83ś96. Accelerated Processor Verification. In Proceedings of the 58th IEEE/ACM Interna- [19] Daniel Holanda Noronha, Ruizhe Zhao, Jeff Goeders, Wayne Luk, and Steven JE tional Symposium on Microarchitecture®. 1462ś1476. Wilton. 2019. On-chip FPGA debug instrumentation for machine learning ap- [41] Yinan Xu, Zihao Yu, Dan Tang, Guokai Chen, Lu Chen, Lingrui Gou, Yue Jin, plications. In Proceedings of the 2019 ACM/SIGDA International Symposium on Qianruo Li, Xin Li, Zuojun Li, Jiawei Lin, Tong Liu, Zhigang Liu, Jiazhan Tan, Field-Programmable Gate Arrays. 110ś115. Huaqiang Wang, Huizhe Wang, Kaifan Wang, Chuanqi Zhang, Fawang Zhang, [20] David Sidler and Ken Eguro. 2016. Debugging framework for FPGA-based soft Linjuan Zhang, Zifei Zhang, Yangyang Zhao, Yaoyang Zhou, Yike Zhou, Jiangrui processors. In 2016 International Conference on Field-Programmable Technology Zou, Ye Cai, Dandan Huan, Zusong Li, Jiye Zhao, Zihao Chen, Wei He, Qiyuan (FPT). IEEE, 165ś168. Quan, Xingwu Liu, Sa Wang, Kan Shi, Ninghui Sun, and Yungang Bao. 2022. To- [21] Fatemeh Eslami and Steven JE Wilton. 2014. Incremental distributed trigger wards developing high performance RISC-V processors using agile methodology. insertion for efficient FPGA debug. In 2014 24th International Conference on Field In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). Programmable Logic and Applications (FPL). IEEE, 1ś4. IEEE, 1178ś1199. [22] Gefei Zuo, Jiacheng Ma, Andrew Quinn, and Baris Kasikci. 2023. Vidi: Record [42] Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, and replay for reconfigurable hardware. In Proceedings of the 28th ACM International Byoungyoung Lee. 2021. Difuzzrtl: Differential fuzz testing to find cpu bugs. In Conference on Architectural Support for Programming Languages and Operating 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 1286ś1303. Systems, Volume 3. 806ś820. [43] Fabian Thomas, Lorenz Hetterich, Ruiyi Zhang, Daniel Weber, Lukas Gerlach, and [23] Brad L Hutchings and Jared Keeley. 2014. Rapid post-map insertion of embed- Michael Schwarz. 2024. RISCVuzz: Discovering architectural CPU vulnerabilities ded logic analyzers for Xilinx FPGAs. In 2014 IEEE 22nd annual international via differential hardware fuzzing. https://ghostwriteattack. com/ (2024). symposium on field-programmable custom computing machines. IEEE, 72ś79. [44] RISC-V. 2024. Spike RISC-V ISA Simulator. https://github.com/riscv-software- [24] Pavan Kumar Bussa, Jeffrey Goeders, and Steven JE Wilton. 2017. Accelerating src/riscv-isa-sim in-system FPGA debug of high-level synthesis circuits using incremental compi- [45] AMD. 2025. Virtex UltraScale+ VU19P. https://www.amd.com/en/products/ lation techniques. In 2017 27th International Conference on Field Programmable adaptive-socs-and-fpgas/fpga.html. Logic and Applications (FPL). IEEE, 1ś4. [46] EEMBC. 2018. CoreMark. https://github.com/eembc/coremark. [25] Eddie Hung and Steven JE Wilton. 2013. Incremental trace-buffer insertion for [47] David Patterson, Jeremy Bennett, Mary Bennett, Hélène Chelin, David Harris, FPGA debug. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 22, Jennifer Hellar, William Jones, Konrad Moron, Paolo Savini, Roger Shepherd, 4 (2013), 850ś863. Ray Simar, Zachary Susskind, and Stefan Wallentowitz. 2025. Embench IOT 2.0 [26] Yuchen Hu, Junhao Ye, Ke Xu, Jialin Sun, Shiyue Zhang, Xinyao Jiao, Dingrong and DSP 1.0: Modern Embedded Computing Benchmarks. Computer 58, 5 (2025), Pan, Jie Zhou, Ning Wang, Weiwei Shan, Xinwei Fang, Xi Wang, Nan Guan, and 37ś47. doi:10.1109/MC.2024.3511352 Zhe Jiang. 2025. Uvllm: An automated universal rtl verification framework using