Skip to content
STIMSMITH

SOURCE ARCHIVE

SHA256: 328809f564fc0855c254d228a0bc4534fa5c4bdd3c37518782067a4377540d1b
TYPE: application/pdf
SIZE: 627.5 KB
FETCHED: 8/5/2026, 10:06:00 AM
EXTRACTOR: liteparse
CHARS: 75,359

EXTRACTED CONTENT

75,359 chars

0 White Rose eprints@whiterose.ac.uk (¥) Research Online. https://eprints.whiterose.ac.uk Universities of Leeds, Sheffield and York

Deposited via The University of York.

White Rose Research Online URL for this paper:
https://eprints.whiterose.ac.uk/id/eprint/244137/

Version: Accepted Version

Proceedings Paper:
Hu, Yuchen, Sun, Jialin, Du, Yushu et al. (2026) Beyond Fuzzer Islands: CPU Fuzzing via
Smart Coordination. In: Design Automation Conference 2026.

Reuse This article is distributed under the terms of the Creative Commons Attribution (CC BY) licence. This licence allows you to distribute, remix, tweak, and build upon the work, even commercially, as long as you credit the authors for the original work. More information and the full terms of the licence here: https://creativecommons.org/licenses/

Takedown If you consider content in White Rose Research Online to be in breach of UK law, please notify us by emailing eprints@whiterose.ac.uk including the URL of the record and the reason for the withdrawal request.

                       University.  .  of er
|                      Sheffield      53    Flt/
UNIVERSITY OF LEEDS        07

Beyond Fuzzer Islands: CPU Fuzzing via Smart Coordination

2
Yuchen Hu1,2, Jialin Sun1,2, Yushu Du                                   , Renshuang Jiang3, Ning Wang4,
                                                                                             Weiwei Shan1,2, Xinwei Fang5, Xi Wang1,2, Nan Guan4, Zhe Jiang1,2†
                                                                                   1Southeast University, China 2National Center of Technology Innovation for EDA, China
                                                                     3National University of Defense Technology, China 4City University of Hong Kong, Hong Kong 5University of York, UK
                                                                                                        †Corresponding author: zhejiang.uk@gmail.com

Abstract Independently One Fuzzing Thread Legends C1 Fuzzer1 DUT Seed Despite recent progress, every CPU fuzzer explores only a narrow C1 DUT C* slice of the vast micro-architectural state space due to its fixed C2 Fuzzer2 DUT Database SO C* Corpus mutation and feedback biases. While different fuzzers thus excel C2 DUT Mergequick TA Final Target in disjoint regions, naïve combination fails because of conflicting C3 Fuzzer3 Int. C3 DUT Result Trans. DUT Lbut great strategies, seed pollution, and early saturation. LiFU introduces overlap possible growth micro-architecture-aware orchestration that dynamically profiles Serially L but slow Result C1 Fuzzer1 DUT C2 Fuzzer2 DUT C3 Fuzzer3 DUT heterogeneous fuzzers, detects complementary strengths, decreases C1 DUT C2 DUT C3 DUT harmful interactions, and steers each fuzzer in real time using Interactively *Beneficial *Posionous coverage and bug feedback, augmented by semantic seed triage. Fuzzer1 J Ensemble Linherent conflicts Ensemble Evaluated on the BOOM core, LiFU achieves 93.4% line, 43.6% FSM, Fuzzer1 DUT DUT CC DUT and 95.7% condition coverage (90.1%, 75.0%, 78.3% on Rocket) with CC DUT Starting Point SOTA SOTA around 40% fewer tests than the best standalone fuzzer, consistently Fuzzer2 Merged Fuzzer3 Merged closing long-standing verification gaps in modern CPU. Figure 1. Independent, serial, and interactive parallel execution of multiple 1 Introduction fuzzers. While interactive execution enables real-time collaboration, certain fuzzer pairs exhibit severe negative interference. Modern hardware systems continue to scale in complexity, inte- grating heterogeneous intellectual properties (IPs) and advanced requires a carefully ordered sequence such as ADD → STORE → LOAD, processors, ranging from in-order designs (e.g., Rocket [1]) to out-of- where the store and load access the same address within a few order superscalar cores (e.g., BOOM [2]). This scaling dramatically cycles. 1 Longer hazards, such as forwarding through multiple expands the architectural state space and micro-architectural inter- pipeline stages or dependent branches, reduce the likelihood expo- actions that require validation, encompassing pipeline speculation, nentially, making them virtually unreachable under random mu- memory coherency and privilege transitions. With post-silicon bug tation. Learning-based generators [18, 27, 29ś32] mitigate early fixes often incurring billion-dollar costs and pre-silicon verification randomness by producing semantically coherent programs, but already accounting for over 70% of the total design effort [3ś9], en- rare multi-cycle interactions remain unreachable. hancing verification efficiency and depth is paramount for ensuring Biasing mutations [20, 27, 33] toward dependency-heavy pat- functional correctness and security of the CPU cores. terns improve depth yet sacrifice diversity, systematically missing In recent years, hardware fuzzing has emerged as a scalable alter- uncorrelated bug classes. In contrast, more stochastic fuzzers [22, native to formal verification [10, 11] and constrained random test- 24] maintain higher diversity and occasionally uncover unexpected ing [12, 13] for RTL designs [14ś21]. Inspired by software fuzzing, behaviours, but they struggle to penetrate the structured semantic hardware fuzzing techniques generate instruction streams and re- dependencies. This diversity, however, has led to a natural intuition: fine them via coverage feedback from DUT simulation. Early black- combine multiple fuzzers to leverage their strengths [34]. However, box tools (RFUZZ [22], DirectFuzz [23]) treated streams as raw naïve composition rarely works [35ś37]. As shown in Figure 1, run- bit-sequences, scaling well but converging quickly due to poor ning fuzzers independently and merging corpora post-hoc results architectural visibility. Later hardware-aware fuzzers introduced in the loss of cross-guided exploration advantages; running them register coverage (DifuzzRTL [24]), explicit dependency chains sequentially is slow and ordering-sensitive. Interactive parallel exe- (Cascade [25]), formal counterexample seeding (HyPFuzz [26]), and cution enables runtime seed sharing and can yield more balanced assembly synthesis guided by language model (LM) (GenHuzz [27]). performance. However, we observe empirically that some fuzzer Despite continuous progress, no existing fuzzer simultaneously combinations accelerate coverage, while others are poisonous, de- delivers high instruction diversity (e.g., rapidly hit basic arithmetic grading performance, even below a single fuzzer, due to incompati- operations) and deep multi-cycle dependency chains (e.g., trigger ble heuristics or mismatched abstraction levels. These behaviours pipeline corner cases). On Rocket, even state-of-the-art (SOTA) indicate that the challenge is no longer merely łdesign a better tools reach only 80% coverage, leaving substantial portions of fuzzer,ž but how to manage multiple fuzzers. the pipeline untested [27]. This phenomenon occurs because cur- Contributions. We introduce LiFU, a verification framework or- rent fuzzers suffer from weaknesses that leave distinct portions of chestrating fuzzing engines and formal methods via feedback-driven the micro-architectural state space unexplored. Closer inspection coordination. LiFU detects coverage growth conditions and selects shows that these tools often converge early because certain micro- abstraction levels using lightweight heuristics and semantic-aware architectural behaviours are statistically or structurally hard to trig- LLM-powered analysis. On Rocket, LiFU achieves 90.1% line, 75.0% ger [28]. For example, exposing a data-forwarding hazard typically FSM, 78.3% condition coverage (versus individual Cascade [25] (the

                                                                        1

DAC 2026, Long Beach, USA Assuming uniform instruction sampling over 1146 opcodes and 32 registers, the 2026. ACM ISBN 978-X-XXXX-XXXX-X/XX/XX probability of generating that exacts three-instruction dependency chain by chance is https://doi.org/10.1145/3770743.3804351 roughly (1/1146) 3 × (1/32) 2 ≈ 6 × 10−12. 1

                  Fetch                                Mutate                                                                                     Execute            Analyse                                            Update
                                                                                                                                                                                                                        Save Bugs
                                                           Random                                                         Time     Invalid                                                                              and Add to
                                                           Random                                                         ISS
                                                           Random
                                                           Random
                                                           Mutator
                     Seeds for                             Mutator
                     Seeds for                                                                                                             Testcase1       …                                                            Knowledge
        W$                                                                                                                DUT#1                                  Testcase5        Human                                      Final
        W$            Mutation                             Mutator
                                                           Mutator                                                        DUT#2                     Testcase2
                      Mutation                             Directed                                    Mutate | Execute   DUT#3                                 Testcase4      Execute | Analyse  Expert   Interesting      Output
                                                                                                                          Filter           ISS                                                                          Seeds for
                                                                                                                                                                                                           Stimulus      Mutation
        Seed                                                                                                                                                                                                             Mutation
        Seed             S$                                Mutator                                                                                  Trace_Log   DUT#1     Checker        Mis.                           Seeds for
       Selector          S$
       Selector                      Cfg.
                                     Cfg.              User-defined                                                                        debug                                     Log

Legends Configure 0 DUT#2 Parser W$ Data Runtime Gen. 1 Discard Property Analyser W$ Corpus $$ Cache Property Mutator Interest Constant DUT#3 Coverpoint Idle Run Knwl. Coverpoint Property State Knwl. Property Generator State Knowledge for LLMs Generator Figure 2. The LiFU Framework Overview: The five-stage pipeline begins with the seed fetch. Phase I (Fetch) selects seeds from a hybrid corpus (initially, prior instruction/program sets). Phase II (Mutate) applies a suite of hybrid mutators to produce candidate testcases. Phase III (Execute) runs the ISS ahead as a fast predictive oracle, generating a golden trace/log; DUTs follow asynchronously, streaming partial traces to the Checker for live differential comparison (execution timeline inset). Phase IV (Analyse) parses logs, evaluates uncovered points, and identifies interesting stimuli and mismatches with human-in-the-loop triage. Finally, Phase V (Update) saves confirmed bugs, updates the knowledge base, and feeds high-value seeds back for the next iteration. Cycle0 FFF MMM EEE AAA UUU Serial Table 1. Idealised throughput on Rocket core (unit time is 20 minutes) Cycle1 FF MM EE AA UU Config. Testcases / Unit Time Coverage Gain Cycle2 F -Fetch E -Execute A -Analyse FF MM EE Cycle3 M - Mutate U -Update FF MM EE AA UU Random baseline 20 5.78% Cycle0 FFF MMM EEE AAA UUU Invalid Pipelined Serial monolithic 3 6.71% Cycle1 FFF MMM EEE AAA UUU Testcases LiFU pipelined 8 10.28% Cycle2 F MFF MM EEE Cycle3 F MFF MM EEE AAA UUU mutation rarely preserves the register dependencies WhisperFuzz Figure 3. The Fuzzing Pipeline. relies on, collapsing its gradient signal. On CVA6 [39], this uncoordi- SOTA fuzzer): 89.2%, 60.0%, 77.5%). On BOOM, it attains 93.4% line, nated ensemble drops timing-path coverage from 39.6% to 29ś33%, 43.6% FSM, 95.7% condition (versus individual Cascade: 83.2%, 11.6%, misses 4/12 timing bugs, and increases runtime by 1.5ś2.2×. 90.8%, with 4ś5× faster convergence. When compared to the merged Requirement for effectiveness. Effective ensembles demand coverage database obtained by combining the best results from cur- continuous assessment of which seeds benefit which engine, plus rent the SOTA fuzzers Cascade [25] and ProcessorFuzz [16], LiFU dynamic resource allocation and selective seed transformation. achieves relative gains of 10.9%, 221.5%, 4.9% in line, FSM, condition coverage on BOOM. It cuts time-to-90% coverage by more than 50%, Efficiency is defined as the capacity of the fuzzing system to exe- boosts final coverage by around 15% and finds six bugs. cute the maximum number of test cases and reach coverage goals as quickly as possible, by fully utilising available computational 2 Background and Motivation resources. The challenge is that any added coordination (as needed CPU fuzzers exhibit complementary but asymmetric strengths. Ran- for effectiveness) will introduce overhead. dom engines (e.g., RISCV-Torture [38], RFUZZ [22]) provide broad Classic fuzzers use a serial loop, so any coordination barrier testcase diversity, whereas directed fuzzers (e.g., HyPFuzz [26], would block the entire pipeline and directly compound the already- WhisperFuzz [20]) are designed to drive long dependency chains severe load imbalance of slow RTL simulation: fast cores sit idle and explore deeper in micro-architectural behaviours. Although waiting for the slowest simulator and for coordination decisions. A such complementary strengths suggest that ensembles should out- central scheduler naïvely onto the serial fuzzing loop would make perform individual fuzzers, the performance of existing ensemble coordination overhead directly worsen the imbalance. fuzzing is insufficient due to lack of effectiveness and efficiency. We reconstruct the pipeline into an asynchronous five-stage Effectiveness refers to the ability that an ensemble consistently one (Fetch → Mutate → Execute → Analyse → Update; Figures generate high-value test cases without interference between its 2 & 3). All coordination (seed ownership, heuristic priority, etc.) fuzzes. Achieving this is non-trivial: when multiple fuzzzers share is handled in lightweight, lock-free stages that execute in parallel seed, their interplay can inadvertently hinder progress. with the slow Execute stage, completely hidden in its slack. High- HyPFuzz couples a formal verifier with TheHuzz [14], sharing value seeds and partial traces flow instantly without stalling any their seeds. On Rocket, its adaptive proof-time budgeting policy fast stage. Result on Rocket (Table 1) shows, the pipelined design yields 41× faster architectural-state coverage than TheHuzz alone. processes 2.7× more testcases per unit time and achieves roughly However, switching to fixed time slices (common practice for multi- 77% greater coverage gain than a traditional serial loop. component system) lets the formal engine to dominate trivial proofs, Requirement for efficiency. Coordination must never sit on the hindering progress; the speed-up decreases by 75% and 107 unnec- critical path of the no-idle-stage pipeline, requiring analysis and essary context switches waste 18% of the total budgets [26]. scheduling to guide the next few cycles. Similar observations are shown when combining a targeted fuzzer with random mutators. WhisperFuzz targets data-dependent LLM-aided fuzzing: a step forward. The two requirements above timing channels. For example, DIVUW shows a 14ś56-cycle gap be- show that effective and efficient ensembles demand semantic aware- tween divisors 0/1, which activates WhisperFuzz’s leakage analyser ness to avoid poisonous interactions and predictive, non-blocking and yields highly focused seeds. When these seeds are passed di- scheduling that existing heuristics fail to provide. LLMs offer a rectly to a random fuzzer that mutates raw 32-bit streams, two promising path forward. Because they can interpret RTL behaviour, issues arise: many seeds induce division-by-zero traps or unstable infer dependency patterns, and anticipate how different fuzzing branches, producing many unexecutable mutants, and random byte engines will react to a given seed, they have the potential to enable 2

                                                                                                         Syntax
    Lᴿᵉᵈᵘⁿᵈᵃⁿᵗ      LLes likely to                                                        Binary         Syntax
                                                                                          Binary                      Directed    Generative
                                                                                                         -aware                   Generative
               Tests                     Trigger Unknown                                                 -aware       Directed
                                         Pattern                                   Validity                            Scalability
                                                                                                            Novelty                     Speed

 J  More                                                                          Breadth                              Cost
    Exploration                             JGet to Target More
    Possiblities                              Straightforward                                         Determinism                    Feedback
                                                                                                                                  Sensitivity
 Architectural level                     Micro-Architectural level                        Exploration                  Efficiency
    (Grey-box)                                  (White-box)                 Figure 5. Qualitative trade-offs of mutation strategies, grouped into explo-

Figure 4. Exploration dynamics in processor fuzzing. Grey-box fuzzing (left) ration and efficiency axes: novelty (new micro-architectural behaviours), receives only architectural feedback (registers, memory, ISA-level coverage). breadth (state-space volume), validity (runnable tests), determinism (pre- White-box fuzzing (right) has direct visibility into the RTL source and dictability and reproducibility); feedback sensitivity (response to precise internal pipeline/functional-unit states. coverage gaps), speed/cost/scalability (throughput and resource demands). LLM-based generative mutation, by reasoning over the actual RTL, achieves compatibility-aware seed exchange and eliminate the poisonous in- superior novelty and feedback sensitivity at the cost of some determinism. teractions that plague existing ensembles. Meanwhile, their ability to reason about micro-architectural structures and forecast a seed’s effect seen in coverage-guided fuzzers, where a few high-yield seeds possible coverage payoff opens the door to predictive, pipeline- dominate mutation cycles and stall global exploration. friendly scheduling that keeps every stage saturated without in- Seed Selector. Guided by W$, the Seed Selector constructs an troducing stalls. Together, these abilities allow LLMs to act as the ordered queue that balances exploitation and exploration. Each seed coordinating layer that current ensemble fuzzers lack. receives a composite score reflecting both its cumulative coverage potential and its structural novelty relative to the current corpus. 3 LiFU: The Framework Pipeline The Selector consumes the ranked list, then performs adaptive LiFU introduces a staged workflow that generates, filters, executes, mutator dispatching: high-value seeds with complex instruction and analyses testcases. The framework combines simulation ac- dependencies are routed to LLM-guided mutators that preserve curacy with selective LLM assistance. LLMs contribute semantic semantic consistency, while low-performing or stagnant seeds are cues where helpful, while structured mutators, fast filtering, and handed to binary-level mutators for disruptive exploration. differential checking ensure correctness and stability. The following 3.2 Mutate: Hybrid Test Case Generation subsections describe each stage of the workflow. The Fetch stage Mutation lies at the core of every fuzzer, yet existing hardware selects high-value seeds using novelty and historical performance fuzzers almost exclusively operate at the architectural surface: ran- (Section 3.1). Section 3.2 explains how LiFU applies hybrid-level dom bit flips, instruction splicing, or grammar-guided changes. mutation strategies. Section 3.3 outlines the Execute stage, includ- These quickly saturate visible architectural states but leave micro- ing fast ISS pre-runs, filtering for non-progressing tests, and RTL architectural corner cases untouched, because they cannot see the simulation with mismatch and coverage logging. The Analyse stage, data and timing dependencies hidden inside the RTL (Figure 4, left). where mismatches are classified and unreachable coverpoints are Semantic-Aware Generative Mutation. LiFU breaks this barrier pruned, is introduced in Section 3.4. Finally, the data from the with a micro-architecture-aware generative mutator powered by Analyse stage will be inserted into the data corpus (Section 3.5). LLMs that have ingested the processor’s full design hierarchy, cov- Together, these stages form a feedback-driven loop that expands erage gaps and gap-related RTL codes. Instead of blind edits, LLMs coverage efficiently and exposes bugs missed by traditional fuzzers. interpret uncovered functional blocks (e.g., an unexercised carry 3.1 Fetch: Better Seeds to Start chain or shifter bypass) and synthesises short, semantically precise The Fetch stage governs how LiFU selects and schedules input instruction sequences that activate the missing logic (typically 3ś30 programs for mutation and execution. Its goal is to transform a static instructions; see Listing 1). As Figure 5 illustrates, this white-box queue into an adaptive feedback-driven process that continuously strategy trades some speed and determinism for dramatically higher reallocates effort toward the most promising test directions. novelty and feedback sensitivity compared with traditional binary, Seed Corpus. The pipeline begins with a curated seed corpus de- syntax-aware, or directed grey-box mutations. rived from standard RISC-V test suites [18], e.g., riscv-dv [40] 1 li t0 , 0x F0F0F0F0 # Alternating nibbles and riscv-tests [41]. These corpora provide structurally valid 2 and semantically diverse instruction sequences that cover com- 3 bset t0 , t0 , zero , 31 # Set sign bit bext t1 , t0 , zero , 4 # Extract low nibble mon ISA behaviours, serving as a reliable bootstrap for the early 4 bins t0 , t1 , t0 , 8, 12 # Insert into upper half fuzzing stages. This avoids the cold-start inefficiency observed in Listing 1. LLM-generated testcase targeting Rocket ALU paths random initialisation, where initial inputs often fail trivial decode or privilege checks and waste simulation cycles. The generated fragments are never run standalone. We insert Weight Cache (W$). Each executed seed updates a runtime Weight them into valid base programs drawn from the regular corpus, pre- Cache (W$) , which records historical metadata including coverage serving liveness while injecting targeted behaviour. Conventional contribution, mutation lineage, and observed novelty. Rather than low-level and syntactic mutators continue to run in parallel, supply- serving as a static priority queue, W$ acts as a dynamic scoring ing raw diversity and preventing entrapment in local optima. The oracle that evolves alongside the campaign (see Section 3.5). Seeds resulting hierarchy (random and syntactic mutations for breadth, that repeatedly unlock new coverage are reinforced, while those directed mutations for validity, LLM guidance for depth) keeps cov- exhibiting diminishing returns are decayed in priority. This persis- erage growing long after pure grey-box techniques have plateaued, tent weighting mechanism mitigates the common łearly winnerž without sacrificing their raw throughput advantage. 3

Algorithm 1: Filter for Non-Progressing Testcases Algorithm 2: Reachability Analysis for Uncovered Point Input: Testcase 𝑡, ISS simulator, timeout 𝑇max, min progress 𝛿, max Input: RTL coverpoint 𝑐 , Trace set T, signal driver map 𝐷 repair attempts 𝑅 Output: ReachClass(𝑐 ) ∈ {unreachable, unlikely, reachable} Output: (Valid, 𝑡′) where Valid ∈ {⊤, ⊥} and 𝑡′ is the (possibly 1 𝑠𝑖𝑔 ← GetSignal(𝑐 ) ; // signal driving coverpoint repaired) testcase 2 if 𝑠𝑖𝑔 ∈ 𝐷 and 𝐷 (𝑠𝑖𝑔) = constant then 1 𝑠𝑡𝑎𝑡𝑒0 ← InitState() 𝑝𝑐hist ← ∅, 𝑐𝑦𝑐𝑙𝑒 ← 0 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐 ← ⊥, 3 end return unreachable ; // hard-wired 0/1 𝑠𝑡𝑢𝑐𝑘_𝑠𝑡𝑎𝑡𝑒 ← ⊥ 4 2 while 𝑐𝑦𝑐𝑙𝑒 < 𝑇max and ¬Timeout do 5 𝑣𝑎𝑙𝑠 ← {Read (𝑠𝑖𝑔, 𝑡𝑟 ) | 𝑡𝑟 ∈ T } 3 𝑠𝑡𝑎𝑡𝑒 ← Step𝐼 𝑆𝑆 (𝑡, 𝑠𝑡𝑎𝑡𝑒 ) ; // single-cycle step 6 if |𝑣𝑎𝑙𝑠 | = 1 and T spans ≥ 𝑁 diverse seeds then 4 𝑝𝑐 ← GetPC(𝑠𝑡𝑎𝑡𝑒) 7 return unlikely ; // empirically constant 5 if 𝑝𝑐 ∈ 𝑝𝑐hist and ArchDelta (𝑠𝑡𝑎𝑡𝑒0, 𝑠𝑡𝑎𝑡𝑒) < 𝛿 then 8 end 6 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐 ← 𝑝𝑐, 𝑠𝑡𝑢𝑐𝑘_𝑠𝑡𝑎𝑡𝑒 ← 𝑠𝑡𝑎𝑡𝑒 ; 9 return reachable 7 break ; // loop without architectural progress 8 end 9 𝑝𝑐hist ← 𝑝𝑐hist ∪ {𝑝𝑐 } ; Algorithm 3: LLM-Guided Property Generation 10 𝑐𝑦𝑐𝑙𝑒 ← 𝑐𝑦𝑐𝑙𝑒 + 1 ; Input: Uncoverpoint 𝑢, fused coverage 𝐶, knowledge base 𝐾 , 11 end prompt template 𝑃 12 if 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐 = ⊥ then Output: Set of synthesised properties P 13 return (⊤, 𝑡 ) ; // normal progressing testcase 1 𝑐𝑜𝑛𝑡𝑒𝑥𝑡 ← ExtractSnippet(𝐶, 𝑢) ; // uncoverpoint log 14 end 2 𝑝𝑟𝑜𝑚𝑝𝑡 ← 𝑃 .𝑓 𝑜𝑟𝑚𝑎𝑡 (𝑐𝑜𝑛𝑡𝑒𝑥𝑡, 𝐾 ) ; // grounded prompt 15 𝑡′ ← 𝑡 for 𝑎𝑡𝑡𝑒𝑚𝑝𝑡 = 1 to 𝑅 do 16 𝑡′′ ← 𝑡′ ; // Repair at 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐 construction 17 if 𝑎𝑡𝑡𝑒𝑚𝑝𝑡 ≤ 𝑅/2 then 3 P ←ˆ LLMQuery(𝑝𝑟𝑜𝑚𝑝𝑡 ) ; // initial generation // Insert NOPs / privilege-neutral instructions 4 P ← ∅ foreach 𝑝ˆ ∈ Pˆ do 18 𝑡′′ ← InsertNopsAfter (𝑡′, 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐, Rand(1, 8) ) 5 if ValidateSyntax (𝑝ˆ) and SimCheck (𝑝, 𝐶ˆ ) ≠ ⊥ then 19 else 6 P ← P ∪ 𝑝ˆ ; // add if syntactically valid and // Inject reg.-forwarding breakers or CSR writes 7 end consistent 20 𝑟𝑒𝑔 ← RandLiveOutReg(𝑠𝑡𝑢𝑐𝑘_𝑠𝑡𝑎𝑡𝑒) ; 8 end 21 𝑖𝑛𝑠𝑡𝑟 ← GenForwardingBreaker(𝑟𝑒𝑔 ); 9 if | P | = 0 then 22 𝑡′′ ← ReplaceOrInsertAfter(𝑡′, 𝑠𝑡𝑢𝑐𝑘_𝑝𝑐, 𝑖𝑛𝑠𝑡𝑟 ) 10 P ← RefinePrompt (𝑝𝑟𝑜𝑚𝑝𝑡 ) ; // refinement 23 end 11 24 if QuickSim(𝑡′′, 𝑠𝑡𝑢𝑐𝑘_𝑠𝑡𝑎𝑡𝑒, 64) shows ≥ 𝛿 progress then end 25 return (⊤, 𝑡′′ ) ; // repair succeeded 12 return P 26 end 27 𝑡′ ← 𝑡′′ ; // keep best attempt so far differential check between DUT outputs and the ISS golden trace. 28 end Comparisons span register states and memory transactions, with 29 return (⊥, 𝑡 ) ; // failed to repair within budget mismatches (e.g., register value mismatch) captured in a structured 3.3 Execute: Asymmetric Simulation with ISS Pre-Run mismatch log. This log, enriched with context such as triggering instructions and affected modules, is forwarded to the Analyse It is necessary to get feedback for the fuzzing process in the Exe- stage for classification and guided repair. For cases with no direct cute stage. RTL simulation provides the highest micro-architectural mismatches, the Checker merges coverage data across runs and per- fidelity but is prohibitively slow, often several orders of magnitude forms micro-architectural reachability refinement (Algorithm 2) slower than instruction-set simulation. To balance precision and to exclude constant or empirically unreachable RTL coverpoints throughput, LiFU adopts an asymmetric simulation strategy that de- (e.g., tied-off debug ports or inactive configuration lines). 2 couples reference generation from detailed RTL execution. Drawing 3.4 Analyse: Fusion with LLM-Guided Interpretation from techniques in prior work [42], a fast Instruction Set Simula- Log Parsing and Coverage Fusion. The Analyse stage begins tor (ISS) (e.g., Spike) runs ahead as a golden oracle to produce with a log parser ingesting Checker artefacts and simulator cover- architectural traces, while multiple RTL DUT instances execute age databases (e.g., URG reports): ISS/DUT traces, mismatch logs, asynchronously using these traces as reference checkpoints. and metrics (uncovered points and architectural events). For each Early Filtering for Efficiency. Although each testcase undergoes testcase 𝑡, it computes Δcov(𝑡) (new-bit fraction), extracts cycle- standard syntactic compilation, correctness at the ISA level does not accurate mismatches with severity, and ranks uncoverpoints by guarantee meaningful progress at runtime. Random or mutation- criticality, proximity, and difficulty. Metrics are normalised across based generators often produce non-progressing testsÐprograms DUTs for consistent scoring, quantifying completeness and expos- that loop indefinitely, execute very few static instructions, or repeat- ing under-tested regions (e.g., rare instructions or pipeline stalls) edly stall due to dependency or privilege traps. Such deadlock-like to enable targeted reseeding and LLM-guided synthesis. behaviour wastes valuable RTL cycles and distorts coverage feed- LLM-Guided Property Generation and Interpretation. This back, since most instructions remain unexecuted. To mitigate this, stage fuses coverage data with LLM-driven analysis to generate the Execute stage incorporates a lightweight pre-filter (Algorithm 1) temporal properties for hard-to-reach coverpoints flagged as un- that evaluates each candidate through bounded static and dynamic covered by simulator, e.g., redundant FSM transitions or condi- checks, including loop-depth analysis, control-flow graph traversal, tion sequences. The LLM examines fused logs and prioritised un- and minimal progress estimation via a short ISS pre-run. Instead of coverpoints via targeted prompts (e.g., łSynthesise an SVA prop- complete rejection, the filter can also trigger repair actions, such as erty requiring FSM state S5 to S2 after a 3-cycle stall and branch inserting bounded-loop exits or safe termination sequences, when a high-value seed (as indicated by its Weight Cache score) is blocked 2 This prevents such artefacts from misleading the whole fuzzing process into futile only by trivial control-flow issues. search directions. Points labelled as łunlikely reachablež are withheld from feedback, Checker and Reachability Refinement. Upon completion of while those remaining uncertain may later escalate to formal reachability checks. This RTL simulation, the Checker module performs instruction-granular layered processing ensures that coverage metrics reflect genuine design behaviour rather than simulation noise or unreachable artefacts. 4

      90                                                      70                                            70          Large amount of debug-related
      89                                                      60                                                        conditions (e.g., assert)
                                      Cascade[25]             50                         Cascade[25]                                             Cascade[25]
      88                              This work               40                         This work          60                                   This work
           0   10000 20000 30000      40000   50000                  0   10000 20000
                                                                           (a) Rocket30000    40000  50000        0     10000 20000         30000 40000   50000
               Number of test cases                                        Number of test cases                            Number of test cases

                                                              40                     Proof of generated                   Randomness of
      80                                                                             cover property         80          Generated Test Cases
                                      Cascade[25]             20                         Cascade[25]                                             Cascade[25]
      60                              This work                0                         This work          60                                   This work
           0   10000 20000 30000      40000   50000                  0   10000 20000 30000    40000  50000        0     10000 20000         30000 40000   50000
               Number of test cases                                      Number of test cases                              Number of test cases
                                                                             (b) BOOM
Figure 6. Coverage benchmark on the CPUs. LiFU continuously outperforms Cascade [25] in all coverage metrics, with BOOM achieving large FSM coverage
gains from the formal proof of LLM-generated properties. The largest gap is in BOOM FSM coverage, where LiFU has LLMs turn uncoverpoints into formal
properties to prove rare controller transitions. On Rocket, condition coverage plateaus early for both tools because many remaining conditions are unreachable
debug-related logic. On BOOM, Cascade occasionally edges ahead in condition coverage simply because of the randomness of generated testcases.

               Table 2. Weight Cache (W$) Scoring Factors                            Table 3. Final Coverage on BOOM. łMergedž means merging the database of
    Factor     Definition                  Computation                               Cascade [25] and ProcessorFuzz [16] after independent fuzzing; Percentages
                                                                                     report the relative improvement of LiFU over the merged solution.
   Δcov(𝑡 )   New coverage contribution   Í|cov|𝐶(𝑡)\𝐶ᵖʳᵉᵛ|
                                             total|                                            Method             Line          FSM                     Condition
  bugsc(𝑡 )   Bug detection impact            severity(𝑚);                                Cascade [25]          83.2          11.6                       90.8
                                           ∀𝑚 ∈ mismatches(𝑡 )                         ProcessorFuzz [16]       80.3         13.55                      86.79
  nov(𝑡 )     Structural uniqueness           1 − sim(𝑡, Sseen)                              Merged             84.2         13.56                       91.2
                                           (e.g., normalised edit distance)                  This work      93.4 (10.9%↑)  43.6 (221.5%↑)              95.7 (4.9%↑)
  eff(𝑡 )     Insight per cycle           Δcov(𝑡)+bugsc(𝑡)
                                               cycles(𝑡)                            VCS (O-2018.09-SP2) [43], while formal verification was conducted
mispredictž), producing assertions that explain coverage gaps and                    with VC Formal (S-2021.09-SP2) [44]. Stimulus generation in LiFU
guide stimulus mutation. Grounded in empirical traces and a con-                     leveraged direct calls to the OpenAI API, using GPT-4-turbo as the
strained knowledge base, the LLM translates raw metrics into                         default model to synthesise assembly sequences. We also adopted
natural-language insights (e.g., łThis bin is blocked by missing                     Cascade [25] and ProcessorFuzz [16] for seed mutation.
hazard resolutionž) while reducing hallucination. Valid properties                   All experiments were conducted on an AMD EPYC 7763 2.45GHz
and high-impact stimuli are prioritised for reseeding. Algorithm 3                   CPU. We compare LiFU against five SOTA hardware fuzzers: Dif-
formalises this iterative, validation-gated workflow.                                fuzzRTL [24], ProcessorFuzz, ChatFuzz [45], Cascade, and Gen-
Mismatch Detection and Triage. Human experts may intervene                           Huzz [27]. Coverage (line, FSM, condition) and bug detection data
to analyse the errors, identify the causes, and check for false posi-                were collected via VCS coverage reports and differential check-
tives. Analysed results will be annotated , providing ground-truth                   ing against Spike. Considering the limited simulation resources,
labels that refine automated thresholds.                                             the time budget for one testcase is 1.5 hours. The thread shall be
                                                                                     cancelled exceeding the budget, with executed information logged.
3.5 Update: Knowledge Integration
Following analysis, the Update stage archives confirmed bugs and                     4.1   Results and Discussions
high-priority stimuli in the knowledge base. New interesting seeds
are added to the seed corpus, along with weights obtained from                       Comparison with Individual Cascade. We compare LiFU against
the priority formula that emphasises coverage deltas and novelty,                    Cascade, currently the strongest published open-source hardware
thereby closing the feedback loop with minimal overhead.                             fuzzer. Across both Rocket and BOOM, our approach delivers higher
W$ Update. For each testcase 𝑡, W$ computes a priority weight                       final coverage and significantly faster convergence in every metric
𝑤 (𝑡) using a multi-factor scoring function:                                       (Figure 6). On Rocket, LiFU achieves 90.1% line, 75.0% FSM, and
    𝑤 (𝑡) = 𝛼 · Δcov(𝑡) + 𝛽 · bugsc(𝑡) + 𝛾 · nov(𝑡) + 𝛿 · eff(𝑡)     (1)   78.3% condition coverage, beating Cascade’s 89.2%, 60.0%, and 77.5%.
where 𝛼 +𝛽 +𝛾 +𝛿 = 1. Table 2 details each factor and its calculation.           We further reach 89.0% line coverage using fewer than 1,000 tests,
Coefficients are user-configurable or adapted via online feedback                    whereas Cascade needs around 5,000, and our FSM/condition curves
(e.g., multi-armed bandit). High-𝑤 (𝑡) testcases are promoted to the               rise far more steeply from the start. The gap grows larger on BOOM,
runtime cache and prioritised for reseeding.                                         where we obtain 93.4% line (vs. 83.2%), 43.6% FSM (vs. 11.6%), and
4   Evaluation                                                                       95.7% condition coverage (vs. 90.8%). The FSM gap is especially
                                                                                     revealing: with the combination of formal tools, LiFU discovers
This section presents our experimental setup, evaluation metrics,                    3.76× more micro-architectural controller states. Although Cascade
and discussions to the results.                                                      temporarily leads in BOOM condition coverage during some inter-
Setup. We evaluated LiFU on two widely-used open-source RISC-V                       vals (roughly 1,000ś8,000 tests), these brief crossings stem from
processor cores: the in-order RocketChip [1] and the out-of-order                    random variation and Cascade’s targeting in deep dependencies.
BOOM [2]. All RTL simulations were performed using Synopsys                          Overall, our method consistently ends with greater results across
                                                                             5

Line (%) Line (%)

FSM (%) FSM (%)

Cond. (%) Cond. (%)

 80                                                                                      Table 4. Reported bugs with LiFU in BOOM and Rocket. B* denotes bugs
                       in BOOM, and R* denotes bugs in Rocket.
 75                                                                                        ID   Core     Bug Description                       New
 70                                                                                        B1   BOOM     Misdecoding when processing single-    ✓
                                                                                                         precision operations like fcvt.l(u).s
 65                                                                                        B2   BOOM     Improper floating-point state          ×
                                                                                                         after mstatus.FS state transition
 60                                                                                        B3   BOOM     Incorrect Restoration of MPP field in  ✓
                   Cascade[25]                                                                           mstatus during trap return
 55                DifuzzRTL[24]                                                           B4   BOOM     Uninitialised CSR registers return     ✓
                   ProcessorFuzz[16]                                                                     non-deterministic values after reset
 50                ChatFuzz[45]                                                            R1   Rocket   Misdecoding when processing single-    ✓
                   GenHuzz[27]                                                                           precision operations like fcvt.l(u).s
 45                This Work                                                               R2   Rocket   Uninitialised CSR registers return     ✓
                                                                                                         non-deterministic values after reset
     0 10000 20000 30000     40000     50000
     Figure 7. Number of test cases                                                      instructions that are deliberately replaced in the BOOM, plus par-
     Condition coverage benchmark on Rocket.                                             tially instantiated parametrised modules. Some of these items are
all metrics (line, FSM, condition) on both designs, confirming clear                     waived during fuzzing execution yet still counted in final metrics,
and stable improvement over the SOTA individual fuzzer.                                  and a few remaining holes are simply unreasonable artefacts of the
LiFU vs. Naïve Merging. Different fuzzers show complementary                             coverage tool rather than genuine verification gaps.
strengths. A simple merge of Cascade and ProcessorFuzz coverage                          4.2   Bugs Reported
databases on BOOM yield 84.2% line, 13.56% FSM, and 91.2% condi-                                 Table 4 summarises the bugs identified in the BOOM and Rocket
tion coverage (Table 3). These values exceed the individual fuzzer                       during our verification. Both cores (B1 in the BOOM, R1 in the
(+1.0% line, +0.4% condition). Our method LiFU finally achieves                          Rocket) misdecode the single-precision floating-point-to-integer
93.4% line, 43.6% FSM, and 95.7% condition coverage under the                            conversion instructions such as fcvt.l(u).s, resulting in incor-
same time budget. The large gap (relative increases of 10.9% for                         rect integer outputs. A previously known defect in the BOOM (B2)
lines, 221.5% for FSM states, 4.9% for conditions) proves that LiFU                      causes improper handling of the floating-point register file when the
succeeds not by loosely combining separate tools, but by tightly                         mstatus.FS field transitions (e.g., from Off to Initial), potentially
integrating sequence generation, micro-architectural state tracking,                     leaving stale or corrupted data in the FP registers. Another bug in
and coverage feedback. Naïve database merging fails to drive the                         the BOOM (B3) involves incorrect restoration of the mstatus.MPP
deep, systematic exploration that our design consistently attains.                       field during trap return; the processor may update or restore this
Benchmark against Published SOTAs. We benchmark LiFU                                     field erroneously, causing mret to return to an unintended privi-
against these published results, with Cascade included for refer-                        lege mode (e.g., from machine mode directly to user mode), which
ence3. As shown in Figure 7, our approach significantly outperforms                      triggers subsequent illegal instruction faults on operations like
all prior methods. While DiffuzzRTL exhibits rapid saturation after                      sret. Also, both cores suffer from a post-reset behaviour (B4 in the
modest initial gains, indicating constrained exploration capacity,                       BOOM, R2 in the Rocket) where certain CSR registers, including
our method sustains strong upward progress even as the test-case                         performance counters such as scounteren and sscratch, remain
count scales to 50,000 and beyond. Most strikingly, we achieve cov-                      no-zero-initialised and return non-deterministic values, resulting
erage comparable to or exceeding the results of all prior tools using                    in mismatches between CPU and Spike.
only a small fraction, roughly 1/4 to 1/3, of their test cases. This
superior efficiency in exploring diverse hardware states translates                      5   Conclusion
directly into accelerated bug and vulnerability detection, establish-                    We present LiFU, a CPU testing framework that coordinates fuzzers
ing our approach as the most effective tool for Rocket.                                  and formal tools via runtime analysis and seed selection. Evaluated
Coverage Gap Analysis. Although we achieve the highest cov-                              on BOOM, LiFU outperforms Cascade with 93.4% line, 43.6% FSM,
erage on both Rocket and BOOM, a small residual gap remains.                             and 95.7% condition coverage (vs. 83.2%, 11.6%, 90.8%), achieving 4ś
Manual analysis of VCS reports shows that nearly all uncovered                           5
elements are untestable or irrelevant logic belonging to three main                        × faster convergence and 89% line coverage in less than 1,000 tests
categories: (1) debug and verification constructs (asserts, assumes,                       (vs. around 5,000). It sustains gains beyond 50,000 tests, matching
$fatal) that only trigger on illegal states never generated by the core                  SOTA using 25ś33% of the budget; overall, it cuts time-to-90% cov-
itself and should therefore be excluded from functional coverage; (2)                    erage by more than 50% and uncovers six bugs. No individual fuzzer
severe VCS over-reporting that inflates trivial logic (e.g., BOOM’s                      can achieve comprehensive performance; superior coordination is
BranchMaskGenerationLogic expands simple 2ś4-bit mux/priority-                           essential to advance CPU verification. LiFU establishes adaptive,
encoder circuitry into 420 FSM states, and less than 10% states                          semantically informed orchestration as a scalable paradigm.
are triggered during massive simulation processes without formal                         6
proof); and (3) unreachable redundant code introduced by design                              Acknowledgement
reuseÐmost notably Rocket ALU paths for rotate and miscellaneous                         We greatly appreciate the reviewers for the helpful feedback. This
                                                                                         work is supported by the National Natural Science Foundation of
                                                                                         China (No. 62472086, 92464204), the Science and Technology Major
3Since most recent RTL fuzzers such as GenHuzz and TheHuzz remain closed-source,         Special Program of Jiangsu (No. BG2024010), and the Fundamental
direct evaluation is not feasible. We therefore utilise the data reported in the [27].   Research Funds for the Central Universities (No. 2242025K20013).
                       6

Cond. Coverage (%)

References ACM/IEEE Design Automation Conference (DAC). IEEE, 1ś7. [1] Krste Asanovic, Rimas Avizienis, Jonathan Bachrach, Scott Beamer, David Bian- [22] Kevin Laeufer, Jack Koenig, Donggyu Kim, Jonathan Bachrach, and Koushik colin, Christopher Celio, Henry Cook, Daniel Dabbelt, John Hauser, Adam Izraele- Sen. 2018. RFUZZ: Coverage-directed fuzz testing of RTL on FPGAs. In 2018 vitz, et al. 2016. The rocket chip generator. EECS Department, University of IEEE/ACM International Conference on Computer-Aided Design (ICCAD). IEEE, California, Berkeley, Tech. Rep. UCB/EECS-2016-17 4 (2016), 6ś2. 1ś8. [2] Christopher Celio, David A Patterson, and Krste Asanovic. 2015. The berkeley out- [23] Sadullah Canakci, Leila Delshadtehrani, Furkan Eris, Michael Bedford Taylor, of-order machine (boom): An industry-competitive, synthesizable, parameterized Manuel Egele, and Ajay Joshi. 2021. Directfuzz: Automated test generation risc-v processor. EECS Department, University of California, Berkeley, Tech. Rep. for rtl designs using directed graybox fuzzing. In 2021 58th ACM/IEEE Design UCB/EECS-2015-167 (2015). Automation Conference (DAC). IEEE, 529ś534. [3] Sakari Lahti, Panu Sjövall, Jarno Vanne, and Timo D Hämäläinen. 2018. Are we [24] Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, and there yet? A study on the state of high-level synthesis. IEEE Transactions on Byoungyoung Lee. 2021. Difuzzrtl: Differential fuzz testing to find cpu bugs. In Computer-Aided Design of Integrated Circuits and Systems 38, 5 (2018), 898ś911. 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 1286ś1303. [4] Kan Shi, Shuoxiang Xu, Yuhan Diao, David Boland, and Yungang Bao. 2023. [25] Flavien Solt, Katharina Ceesay-Seitz, and Kaveh Razavi. 2024. Cascade:{CPU} ENCORE: Efficient architecture verification framework with FPGA accelera- fuzzing via intricate program generation. In 33rd USENIX Security Symposium tion. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field (USENIX Security 24). 5341ś5358. Programmable Gate Arrays. 209ś219. [26] Chen Chen, Rahul Kande, Nathan Nguyen, Flemming Andersen, Aakash Tyagi, [5] Ke Xu, Jialin Sun, Yuchen Hu, Xinwei Fang, Weiwei Shan, Xi Wang, and Zhe Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. 2023. HyPFuzz:Formal- Jiang. 2024. MEIC: Re-thinking RTL Debug Automation using LLMs. arXiv Assisted processor fuzzing. In 32nd USENIX Security Symposium (USENIX Security preprint arXiv:2405.06840 (2024). 23). 1361ś1378. [6] Ziqing Zhang, Weijie Weng, Yaning Li, Lijia Cai, Haoyu Wang, David Boland, [27] Lichao Wu, Mohamadreza Rostami, Huimin Li, Jeyavijayan Rajendran, and Yungang Bao, and Kan Shi. 2024. Hassert: Hardware Assertion-Based Verification Ahmad-Reza Sadeghi. 2025. GenHuzz: An Efficient Generative Hardware Fuzzer. Framework with FPGA Acceleration. In Proceedings of the 29th ACM International In 34th USENIX Security Symposium (USENIX Security 25). 1787ś1805. Conference on Architectural Support for Programming Languages and Operating [28] Flavien Solt, Patrick Jattke, and Kaveh Razavi. 2022. Rememberr: Leveraging Systems, Volume 4. 142ś154. microprocessor errata for design testing and validation. In 2022 55th IEEE/ACM [7] Yuchen Hu, Junhao Ye, Ke Xu, Jialin Sun, Shiyue Zhang, Xinyao Jiao, Dingrong International Symposium on Microarchitecture (MICRO). IEEE, 1126ś1143. Pan, Jie Zhou, Ning Wang, Weiwei Shan, et al. 2024. Uvllm: An automated [29] Lisa Zhang, Gregory Rosenblatt, Ethan Fetaya, Renjie Liao, William Byrd, universal rtl verification framework using llms. arXiv preprint arXiv:2411.16238 Matthew Might, Raquel Urtasun, and Richard Zemel. 2018. Neural guided con- (2024). straint logic programming for program synthesis. Advances in Neural Information [8] Jie Zhou, Youshu Ji, Ning Wang, Yuchen Hu, Xinyao Jiao, Bingkun Yao, Xinwei Processing Systems 31 (2018). Fang, Shuai Zhao, Nan Guan, and Zhe Jiang. 2025. Insights from rights and [30] Kyriakos Ispoglou, Daniel Austin, Vishwath Mohan, and Mathias Payer. 2020. wrongs: A large language model for solving assertion failures in rtl design. arXiv FuzzGen: Automatic fuzzer generation. In 29th USENIX Security Symposium preprint arXiv:2503.04057 (2025). (USENIX Security 20). 2271ś2287. [9] Junhao Ye, Yuchen Hu, Ke Xu, Dingrong Pan, Qichun Chen, Jie Zhou, Shuai [31] Sameer Reddy, Caroline Lemieux, Rohan Padhye, and Koushik Sen. 2020. Quickly Zhao, Xinwei Fang, Xi Wang, Nan Guan, et al. 2025. From Concept to Practice: generating diverse valid test inputs with reinforcement learning. In Proceedings of an Automated LLM-aided UVM Machine for RTL Verification. arXiv preprint the ACM/IEEE 42nd International Conference on Software Engineering. 1410ś1421. arXiv:2504.19959 (2025). [32] Ethan Brooks, Janarthanan Rajendran, Richard L Lewis, and Satinder Singh. [10] Czea Sie Chuah, Christian Appold, and Tim Leinmueller. 2023. Formal verification 2021. Reinforcement learning of implicit and explicit control flow instructions. of security properties on risc-v processors. In Proceedings of the 21st ACM-IEEE In International Conference on Machine Learning. PMLR, 1082ś1091. International Conference on Formal Methods and Models for System Design. 159ś [33] Muhammad Monir Hossain, Arash Vafaei, Kimia Zamiri Azar, Fahim Rahman, 168. Farimah Farahmandi, and Mark Tehranipoor. 2023. Socfuzzer: Soc vulnerability [11] Lennart Weingarten, Kamalika Datta, Abhoy Kole, and Rolf Drechsler. 2024. Com- detection using cost function enabled fuzz testing. In 2023 Design, Automation & plete and efficient verification for a RISC-V processor using formal verification. Test in Europe Conference & Exhibition (DATE). IEEE, 1ś6. In 2024 Design, Automation & Test in Europe Conference & Exhibition (DATE). [34] Matej Bölcskei, Flavien Solt, Katharina Ceesay-Seitz, and Kaveh Razavi. 2025. IEEE, 1ś6. Encarsia: Evaluating cpu fuzzers via automatic bug injection. In 34th USENIX [12] Sallar Ahmadi-Pour, Vladimir Herdt, and Rolf Drechsler. 2021. Constrained Security. random verification for RISC-V: overview, evaluation and discussion. In MBMV [35] Shuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu, Kang Li, Zhongyu Pei, and 2021; 24th Workshop. VDE, 1ś8. Zuoning Chen. 2018. Collafl: Path sensitive fuzzing. In 2018 IEEE Symposium on [13] Muhammad Kashif Minhas, Haroon Waris, Yasir Farooq, Nasir Mohyuddin, and Security and Privacy (SP). IEEE, 679ś696. Sajid Baloch. 2023. Coverage-Driven and Constrained-Randomized Sub-System [36] Chenyang Lyu, Shouling Ji, Chao Zhang, Yuwei Li, Wei-Han Lee, Yu Song, and Level Verification Methodology for RISC-V Based SoCs. In 2023 20th International Raheem Beyah. 2019. MOPT: Optimized mutation scheduling for fuzzers. In 28th Bhurban Conference on Applied Sciences and Technology (IBCAST). IEEE, 80ś85. USENIX security symposium (USENIX security 19). 1949ś1966. [14] Rahul Kande, Addison Crump, Garrett Persyn, Patrick Jauernig, Ahmad-Reza [37] Yuanliang Chen, Yu Jiang, Fuchen Ma, Jie Liang, Mingzhe Wang, Chijin Zhou, Sadeghi, Aakash Tyagi, and Jeyavijayan Rajendran. 2022. TheHuzz: Instruc- Xun Jiao, and Zhuo Su. 2019. EnFuzz: Ensemble fuzzing with seed synchronization tion fuzzing of processors using Golden-Reference models for finding Software- among diverse fuzzers. In 28th USENIX Security Symposium (USENIX Security 19). Exploitable vulnerabilities. In 31st USENIX Security Symposium (USENIX Security 1967ś1983. 22). 3219ś3236. [38] RISC-V International and contributors. 2023. RISC-V Torture. https://github. [15] Chathura Rajapaksha, Leila Delshadtehrani, Manuel Egele, and Ajay Joshi. 2023. com/riscv/riscv-torture. SIGFuzz: A framework for discovering microarchitectural timing side channels. [39] Florian Zaruba and Luca Benini. 2019. The cost of application-class processing: In 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE). Energy and performance analysis of a Linux-ready 1.7-GHz 64-bit RISC-V core IEEE, 1ś6. in 22-nm FDSOI technology. IEEE Transactions on Very Large Scale Integration [16] Sadullah Canakci, Chathura Rajapaksha, Leila Delshadtehrani, Anoop Nataraja, (VLSI) Systems 27, 11 (2019), 2629ś2640. Michael Bedford Taylor, Manuel Egele, and Ajay Joshi. 2023. Processorfuzz: [40] Google. 2019. RISC-V DV. https://github.com/google/riscv-dv. Processor fuzzing with control and status registers guidance. In 2023 IEEE In- [41] RISC-V Software contributors. 2025. riscv-tests. https://github.com/riscv- ternational Symposium on Hardware Oriented Security and Trust (HOST). IEEE, software-src/riscv-tests. [Online; accessed]. 1ś12. [42] Jialin Sun, Yuchen Hu, Dean You, Yushu Du, Hui Wang, Xinwei Fang, Weiwei [17] Deepak Narayan Gadde, Aman Kumar, Djones Lettnin, and Sebastian Simon. Shan, Nan Guan, and Zhe Jiang. 2025. ISAAC: Intelligent, Scalable, Agile, and 2024. FuzzWiz-Fuzzing Framework for Efficient Hardware Coverage. In 2024 Accelerated CPU Verification via LLM-aided FPGA Parallelism. arXiv preprint International Symposium on Electronics and Telecommunications (ISETC). IEEE, arXiv:2510.10225 (2025). 1ś5. [43] Synopsys, Inc. 2023. VCS: Synopsys Verification Compiler System. https://www. [18] Yinan Xu, Sa Wang, Dan Tang, Ninghui Sun, and Yungang Bao. 2024. PathFuzz: synopsys.com/verification/simulation/vcs.html Broadening Fuzzing Horizons with Footprint Memory for CPUs. In Proceedings [44] Synopsys, Inc. 2023. VC Formal: Synopsys Verification Compiler Formal Verification of the 61st ACM/IEEE Design Automation Conference. 1ś6. Solution. https://www.synopsys.com/verification/static-and-formal-verification/ [19] Fabian Thomas, Lorenz Hetterich, Ruiyi Zhang, Daniel Weber, Lukas Gerlach, and vc-formal.html Michael Schwarz. 2024. RISCVuzz: Discovering architectural CPU vulnerabilities [45] Mohamadreza Rostami, Marco Chilese, Shaza Zeitouni, Rahul Kande, Jeyavijayan via differential hardware fuzzing. https://ghostwriteattack. com/ (2024). Rajendran, and Ahmad-Reza Sadeghi. 2024. Beyond random inputs: A novel ML- [20] Pallavi Borkar, Chen Chen, Mohamadreza Rostami, Nikhilesh Singh, Rahul Kande, based hardware fuzzing. In 2024 Design, Automation & Test in Europe Conference Ahmad-Reza Sadeghi, Chester Rebeiro, and Jeyavijayan Rajendran. 2024. Whis- & Exhibition (DATE). IEEE, 1ś6. perFuzz: White-Box Fuzzing for Detecting and Locating Timing Vulnerabilities in Processors. In 33rd USENIX Security Symposium (USENIX Security 24). 5377ś5394. [21] Rihui Sun, Jin Wu, Hanyin Liu, Zikang Tao, Gang Qu, Dongsheng Wang, Yongqiang Lyu, and Jian Dong. 2025. BPUFuzzer: Effective Fuzz Testing for Branching Transient Execution Vulnerabilities of RISC-V CPU. In 2025 62nd

 7