Full text
Phantom Trails: Practical Pre-Silicon Discovery of Transient Data Leaks Alvise de Faveri Tron Vrije Universiteit Amsterdam Raphael Isemann Vrije Universiteit Amsterdam Hany Ragab Vrije Universiteit Amsterdam Cristiano Giuffrida Vrije Universiteit Amsterdam Klaus von Gleissenthall Vrije Universiteit Amsterdam Herbert Bos Vrije Universiteit Amsterdam Abstract Transient execution vulnerabilities have affected CPUs for the better part of the decade, yet, we are still missing methods to efficiently uncover them at the design stage. Existing approaches try to find programs that leak explicitly defined secrets, sometimes including the transmission over a sidechannel, which severely restricts the space of programs that can trigger detection. As a result, current fuzzers are forced to constrain the search space using templates of known vulnerabilities, which risks overfitting. What is missing is a general detection mechanism that (1) makes it easy for the fuzzer to trigger a violation and (2) catches vulnerabilities at their root cause — similarly to sanitizers in software. In this paper, we propose Phantom Trails, an efficient yet generic method for discovering transient execution vulnerabilities. Phantom Trails relies on a fuzzer-friendly detection model that can be applied without the need for templating. Our detector builds on two key design choices. First, it concentrates on finding microarchitectural data leaks independently of the covert channel, thereby focusing on the core of the attack. Second, it automatically infers all secret locations from the architectural behavior of a program, making it easier for the detector to find leaks. We evaluate Phantom Trails by fuzzing the BOOM RISC-V CPU, where it finds all known speculative vulnerabilities in 24-hours, starting from an empty seed and without pre-defined templates, as well as a new Spectre variant specific to BOOM — Spectre-LoopPredictor. 1 Introduction Transient execution attacks are a critical security threat that has plagued CPUs for the best part of the last decade. After the initial discoveries of Spectre [39] and Meltdown [42], recent years have brought on a variety of new attacks [6,40,44,46, 52,65,66,70]. Once discovered, these issues are unfortunately not easy to fix: post-silicon mitigations have often proven to be either incomplete [6,13,45,69,71], opening the door for new attacks, or so detrimental to performance as to render them impractical [28]. Ideally, such vulnerabilities should be found, and fixed, at the pre-silicon stage, i.e., during the design phase of the CPU. However, automatically detecting them in hardware designs is challenging. Pre-Silicon Fuzzing. While exhaustive approaches such as formal verification [19,21,26,67] are difficult to scale to real-world CPUs, a promising approach for finding hardware bugs in RTL designs is fuzzing, which has been applied to both architectural [11,34,38,59,64,72] and microarchitectural [25,33] bugs. For CPUs, pre-silicon fuzzing generally requires to iteratively generate random inputs (i.e., programs) and verify their behavior on a cycle-accurate simulation of the Design Under Test (DUT). This approach poses some unique challenges when compared to traditional software fuzzing – especially when looking for transient execution vulnerabilities. First, the size of the input space – the space of all possible programs, initial memory states, and CPU configurations – paired with the complexity of the designs and the slow speeds of cycle-accurate simulations make efficient fuzzing hard. Second, hardware does not inherently “crash”, which raises the problem of how to detect violations during fuzzing. Transient execution vulnerabilities represent a further challenge: while architectural bugs can be detected through HDL assertions or golden reference models, modelling transient vulnerabilities at the RTL level is still an open problem. Problem Statement. Current state-of-the-art fuzzers for transient vulnerabilities address the problem of navigating the huge search space by employing some form of templating, i.e., by either (1) breaking up known end-to-end attacks into individual stages that serve as a blueprint for creating new variants [33], or (2) providing the fuzzer with program snippets such as “try to access a secret” or “slow down an instruction” to mimic the behavior of known PoCs [25]. While these restrictions help make the search practical, they bias the fuzzer towards known issues, which risks overfitting. Our main insight is that current fuzzers need restrictions like templates because their underlying detection models overly constrain the space of programs that can trigger a violation.
In particular, current detection models rely on explicitly defined secrets, i.e., values that should not be leaked by the microarchitecture, which are defined as all data residing in specific memory regions protected by specific hardware flags [25,33,60]. On top of this, state-of-the-art fuzzers [33] only detect violations after a secret is transmitted through a covert channel, thereby requiring full end-to-end attacks. Both choices make life for the fuzzer unnecessarily hard. Phantom Trails. With Phantom Trails, we present a new approach to efficiently finding transient vulnerabilities without templates or smart seeds. Phantom Trails builds on a fuzzer-friendly detector which imposes fewer constraints on the programs and detects violations at their root cause. Instead of end-to-end exploits, Phantom Trails concentrates on finding transient data leaks — ways in which secrets can enter the microarchitecture through transient execution — independently of the side-channel transmission, providing early detection of vulnerabilities. Unlike previous methods, our detection model implicitly defines secrets specific to a program by deriving which memory locations are never accessed architecturally. Taking a key insight from software sanitizers such as ASAN [57] and MSAN [61], whose implicit “tainting” of most of a program’s memory greatly increases the probability of detecting memory errors, Phantom Trails’s tainting of all memory not accessed architecturally maximizes the likelihood of finding violations. This allows to model a wide variety of vulnerabilities including Spectre-v1, Spectre-v2, Spectre-RSB, Spectre-SSB, and Meltdown variants. Evaluation. To demonstrate the practical benefits of our approach, we run a fuzzing campaign on BOOM [5], a popular open-source RISC-V core equipped with an out-of-order pipeline and speculation. On BOOM, Phantom Trails is able to reliably detect all Spectre and Meltdown variants known on this core within 24 hours without the need for templating, unlike the state of the art [25,33]. Phantom Trails also uncovers Spectre-LoopPredictor – a new Spectre variant, specific to BOOM, through which an attacker can cause mispredictions on an uncontrolled branch by training a nearby control-flow instruction. We disclosed Spectre-LP to the maintainers of BOOM, who acknowledged the issue. Contributions. We make the following contributions: 1. We describe a new, fuzzer-friendly detection model offering sanitizer-like functionality for transient execution vulnerabilities in CPU designs; 2. We build an extensible, software-only detector based on LLVM and Verilator to enforce our model on CPU simulations, with minimal knowledge of the DUT and no hardware modifications; 3. We integrate the detector into an open-source fuzzer that can find Spectre and Meltdown samples within 24 hours on BOOM without templating or smart seeds; 4. We uncover a new speculation primitive on BOOM (Spectre-LP) that can be used to mispredict an uncontrolled branch towards a disclosure gadget. Open Sourcing. All the source code of our detector, including our LLVM instrumentation for Verilator (BFSan) and a taint-tracking wrapper for BOOM, is available at https://zenodo.org/records/14726711 , along with all transient leak experiments and fuzzing infrastructure. 2 Background In this section, we briefly recap the nature of transient execution attacks and existing detection models. 2.1 Transient Execution Attacks Phases. End-to-end transient execution attacks consist of three main phases: 1 apriming step, in which an attacker massages the microarchitecture to a vulnerable state, 2 a secret access step, in which the CPU transiently accesses some secret data as a result of the attacker’s priming, and 3 atransmission step, in which the secret is first encoded into a non-transient microarchitectural state (e.g., the cache) and then recovered into an architectural value by the attacker. Classification. In Meltdown-like attacks, the attacker accesses data belonging to a different security domain through a faulting instruction. In particular, in Meltdown a faulty attacker load brings victim data from the L1 cache into the pipeline, while the faulting load in an MDS attack accesses inflight data belonging to the victim. In contrast, in Spectre-like attacks the access occurs in the same security domain, by a victim acting as a confused deputy, while the transient window is generated through speculation. Figure 1shows the phases of a typical Spectre attack employing FLUSH+RELOAD [73]. Covert Channels. Step 3 transmits the secret via a timing covert channel. Many such covert channels have been discovered over the years, including those based on caches [73], translation structures [29,62], prefetchers [30], predictors [3, 20], contention on execution units [8], etc. 2.2 Existing Detection Models Previous work has proposed different models for detecting transient execution vulnerabilities at the pre-silicon stage. Templating. IntroSpectre [25] and SpecDoctor [33] are presilicon fuzzers that are aimed at transient execution vulnerabilities. Due to the complexity of the DUT, they both employ strategies to restrict the search space. IntroSpectre defines a set of “gadgets”, i.e., snippets of code taken from known vulnerabilities (such as “M5 – Generate store and load instructions with overlapping addresses.”) and combines them.
arch. µ-arch. 1 flush 3b reload 2 access 3a encode Attacker Victim D-Cache direct data flow indirect flow Figure 1: Phases of a Spectre attack employing Flush+Reload. The attacker first primes the microarchitecture by flushing the cache 1 , then forces the victim to transiently access a secret 2 , the value of which gets encoded in the microarchitectural state (here the cache) which leaks to the attacker by means of subsequent timed loads 3 . SpecDoctor uses multi-phased fuzzing starting from a predefined template that mimics the different phases of known attacks and tries to fill them until an end-to-end leakage is found, including the transmission and recovery through a covert channel. Secret Tracking. Detection mechanisms for transient vulnerabilities generally involve tracking a secret through the microarchitecture. In particular, IntroSpectre uses a secret value generator to populate secret memory with specific values, and triggers detection if such values are found in a microarchitectural buffer (e.g. the Line-Fill Buffer). SpecDoctor uses differential testing by changing the values of secret memory between two different runs of the same program and checking if the hash of the microarchitectural state differs. STT [74] and CellIFT [60] propose a different approach for detection (but not in the context of fuzzing) that uses hardware Information Flow Tracking, or taint tracking, to precisely follow the flow of secret data during its manipulation. Secret Definition. All existing detection approaches rely on defining secrets. STT [74] is a microarchitectural defense that considers all speculatively-accessed data as secret, until the corresponding instruction is past a Point-of-No-Return in the RoB, by which time the data is considered architectural. This approach is not suitable for fuzzing, as any speculative window would trigger a violation, even those where the speculation turns out to be correct. All other approaches use explicitly defined secrets. CellIFT and SpecDoctor start from a predefined secret memory region, which is isolated using hardware primitives (PMP or page flags). IntroSpectre also uses page flags to identify secrets, but allows them to evolve based on the permissions changes operated by the gadgets. 3 Challenges and Observations for Fuzzing In this section, we highlight some of the obstacles that existing pre-silicon detectors and fuzzers for transient vulnerabilities face, as well as key insights to overcome them. Sources of Entropy. To generate a program that uncovers a transient vulnerability, fuzzers need to beat a variety of entropy sources. First, the fuzzer needs to generate a set of valid instructions, and a program that exhibits some non-trivial control and data flow. Next, the program needs to open a speculative window. On top of this, the program needs to access a memory location containing a secret during speculation. Finally, in the case of SpecDoctor, the program also needs to encode the secret into the microarchitecture, and the fuzzer needs to generate the receiver code that extracts the secret. Creating programs that follow all of these steps is a considerable effort for a fuzzer, and makes efficient fuzzing impractical. Existing fuzzers tackle this complexity through templating, which aims at reducing the entropy of program generation. While this approach speeds up fuzzing, it risks overfitting on known vulnerabilities. We observe that, by concentrating on other sources of entropy, we might be able to significantly speed up fuzzing without the need for templates. Observation #1: To generate samples of transient execution attacks, fuzzers must beat a variety of entropy sources. By focusing on sources other than program generation, we can eliminate the need for templates. Indirect Flows. Our second observation stems from analyzing the different steps of transient execution attacks in Figure 1. We observe that in the priming step (step 1 ) the attacker massages the microarchitecture indirectly, i.e., performs actions that modify the content of prediction structures, without directly accessing their content (which is not available architecturally). Similarly, in step 3 , the victim modifies the microarchitectural state in a secret-dependent way, but there is no direct flow of information between victim and attacker. In contrast, in step 2 (secret access) there is a direct data flow between secret data and some microarchitectural buffer, e.g., the Register File or the Line-Fill Buffer. A key observation is that, while indirect flows are a known issue for taint tracking frameworks and can often lead to overtainting, direct data flows can be precisely tracked—making the secret access step an ideal place to catch speculative attacks. Moreover, as the secret access happens independently of the transmission and recovery step, it is orthogonal to the side-channel being used. Focusing on the secret access step therefore targets the core of the attack, reducing fuzzer entropy. Observation #2: We can remove the entropy of the side-channel by focusing on the secret access phase, where we have a direct data flow of the secret.
Secret Model. A core challenge of defining transient execution attacks at the RTL level is modelling secrets. Existing techniques rely on explicitly defined secrets—for instance, by marking some pages as secret [33,60]. Explicit secret models restrict the number of speculative accesses that trigger detection, making it harder for the fuzzer to find a violation. Moreover, they require the detector to commit to specific threat models. For example, attacks can leak data across hardware-defined boundaries (e.g., user code reading supervisor memory) or software-only boundaries (e.g., JavaScript programs breaking website isolation). Similarly, the secret may be read directly from within the attacker context (Meltdown), or through the victim (via a gadget in the victim code), and then exfiltrated by the attacker (Spectre). Approaches based on explicit secrets protected with PMP/Page Flags must explicitly pick an attacker model before fuzzing [33], and cannot account for leakage across software-only boundaries. Observation #3: By avoiding explicit secrets we can greatly reduce the entropy of the secret address (and possibly catch same-domain leaks). 4 Phantom Trails We now discuss how Phantom Trails addresses these fuzzing challenges, based on our observations. 4.1 Transient Data Leaks While detecting end-to-end leaks is a challenging task and represents a considerable obstacle for fuzzing, detecting secret accesses, which happen before and independently of the side-channel transmission, can be achieved with precise tainttracking, and catches vulnerabilities at their root. In particular, given a program to test, we track all data flows through the DUT by accessing the RTL-level design and applying taint tracking to the cycle-accurate simulation of the CPU. This includes speculative data flows, e.g., speculative loads, which are visibile in the microarchitecture for a restricted period of time (until the speculation is squashed) but not from the architectural execution. Such data flows can move secret data from the memory subsystem to an exposed buffer inside the CPU, i.e., a taint sink, such as the Register File. Once the secret has entered the RF, it can be leaked in a variety of ways, for example, by a subsequent load or a variable-time instruction. We call direct leaks from secret memory to exposed buffers transient data leaks, and focus our detection method on them. 4.2 Implicit Secrets Instead of relying on explicit secrets, which make it hard for the fuzzer to find vulnerabilities and risk missing attacks, we introduce the concept implicit secrets, depicted in Figure 2. Given a stream of instructions executed by the CPU, we can DRAM x = *A0 if (x < 10) y = array[x] A0 A1 arch. spec. secret (a) Explicit Secrets \ x = *A0 if (x < 10) y = array[x] DRAM A0 A1 arch. spec. (b) Implicit Secrets Figure 2: A visualization of the difference between explicit secret models a and implicit secret models b (red is secret). ISA Simulator *A0 = 100 x = *A0 if (x < 10) y = array[x] A0 arch. not taken (a) ISA Simulation Cycle-Accurate Simulator DRAM A0 *A0 = 100 x = *A0 if (x < 10) y = array[x] (b) Taint Initialization DRAM A0 A1 busy busy spec . RoB x = *A 0 if (x < 10) y = array[x] Cycle-Accurate Simulator (c) Taint Detected DRAM A0 A1 done done flush RoB x = *A 0 if (x < 10) y = array[x] Cycle-Accurate Simulator (d) Pipeline Flush Figure 3: Different phases of our detection model. derive the set of all memory locations that should be accessed architecturally. We use this intuition to define our notion of implicit secrets: all data that is not accessed architecturally by a program is considered a secret. During fuzzing, we generate programs that start from the same initial state and can run for a maximum number of cycles proportional to the size of the binary (see Section 5.2). We infer secrets by first executing a generated program on an ISA simulator, such as Spike [2], which recovers the list of architectural accesses for a single run, and then taint every other memory location in the simulated DRAM before running the same program from the same initial state on the microarchitectural simulator. Example. Figure 3represents an example of how a Speculative Bounds Check Bypass (Spectre-v1) can be detected by our model. First, we run a sequence of instructions with an ISA simulator (Phase a ) and infer the set of architecturallyaccessed locations ( {A0} ). Then, we taint every other location in the simulated DRAM as “secret” (Phase b ). Finally, on the cycle-accurate simulation of the CPU, we observe taint coming from A1 , which was never loaded architecturally, inside of the Register File (Phase c ), which triggers detection.
4.3 Flush-Based Classification The transient nature of the vulnerabilities we are looking for implies that the instruction accessing the secret is speculative, and therefore has to be squashed when the speculation is revealed to be incorrect. Microarchitectures typically have at least three ways to signal that an instruction has to be squashed: (1) pipeline exceptions, generated by faulty instructions (e.g., loads that cause a page fault), (2) mispredictions that indicate incorrect control-flow speculation, (3) and rollbacks, which might happen on value speculation, e.g., with store-to-load forwarding. We can use these signals for classification: whenever the microarchitecture brings a secret into a sink, instead of immediately crashing the execution, we wait until one of such signals is detected (Phase d in Figure 3) and perform a preliminary classification of the leak based on it. If no pipeline flush is detected before the end of the program, we report an unidentified leak. This might indicate either the presence of an additional, unidentified flush signal, or an architectural bug that leaks transiently-accessed data. 4.4 Tainting Software Simulations The implementation of our detector has two major requirements: (1) we need a taint-tracking engine to track secrets in the microarchitecture (2) we need easy access to the microarchitectural state during simulation, in particular the simulated DRAM, the Physical Register File (i.e., our sink) and any relevant component for classification (Re-Order Buffer, and signals indicating a pipeline flush). Additionally, for fuzzing, we need to instrument the simulation to gather feedback. To tackle these requirements, we adopt an approach similar to Trippel et al. [64] and instrument the software cycleaccurate simulation of the CPU generated by Verilator [58], an open-source cycle-accurate simulator. With such approach, we can benefit from the power and maturity of existing software such as LLVM [41] and AFL [22,23,75]. More specifically, a software-only approach has the following advantages: 1. Reusing mature infrastructure widely adopted in academia and industry (LLVM, AFL) makes this approach compatible with past and future research/tooling on software fuzzing and vulnerability discovery 2. Using LLVM instrumentation for both taint tracking and fuzzing means that the same taint infrastructure can be used for both detection and fuzzing feedback 3. A software-only infrastructure makes it easier to prototype new detection mechanisms and taint policies, and guarantees easier scalability (does not require FPGAs) ISA Simulator ICache ReOrder Buffer Exec Units input program record accessed memory run program CPU Core Cycle-Accurate Simulator DRAM L2 Cache DCache Load Ports Register File Taint Engine taint everything else detection + classification 1 3 4 6 2 5 Figure 4: Structure of Phantom Trails’ detector. Given an input binary 1 , the detector runs an ISA Simulator to obtain a list of architectural accesses 2 . Every address that is not in this list is tainted in the simulated DRAM 3 . Then, the program is run through the cycle-accurate simulator 4 , where taint is allowed to propagate until it reaches a sink 5 . On the next pipeline flush, the detector will abort the execution with an error code and produce a detection report 6 that marks the input as problematic. 5 Detector Design We now present our implementation of Phantom Trails’ detection component, and evaluate its ability to correctly identify and classify PoCs of known vulnerabilities. 5.1 Components Figure 4represents an overview of the structure of Phantom Trails’ detector. Similar to previous work, we use the BOOM RISC-V core [5] as the design-under-test for our prototype. ISA simulator. To infer secret locations,we use a modified version of the RISC-V ISA simulator Spike [2] to architecturally simulate an input program. We modified Spike to log all memory locations (and instructions) accessed during simulation as well as the number of executed instructions. Additionally, we added the possibility of discarding test cases that hinder correct classification, such as self-modifying code (see Section 6.2). It is worth noting that, while our system is currently implemented for RISC-V architectures, ISA simulators exist also for other instruction sets.
BOOM source (Chisel) Verilog "Verilated" C Binary FIRRTL Compiler Verilator Clang Add simulation wrapper Compile with data-flow sanitizer Runtime taint propagation Figure 5: Compilation pipeline that transforms the BOOM design into an instrumented binary that can be used for fuzzing. Taint tracking engine. Our prototype uses a custom taint tracking engine called BFSan, which supports bit-precise tracking of taint throughout a program. We only propagate taint through direct data flows. BFSan is built on top of the MemorySanitizer (MSan) error detector from the LLVM compiler infrastructure. Similar to MSan, it divides the memory space into two parts: normal program memory, and a shadow map, used to track the taint of each bit. BFSan follows the flow of taint by instrumenting a program during compilation. The instrumented code propagates the taint through the program by updating the contents of the shadow map on each executed instruction. Cycle-accurate simulator. The cycle-accurate simulator used in our prototype is generated by the Chipyard [1] build system. The source code of BOOM is first lowered from Chisel [17] to Verilog, then translated by Verilator into a compilable C ++ object that contains all simulation logic and whose members represent hardware registers and wires. Figure 5shows an overview of the compilation process. Once the C ++ object is generated, we identify the relevant components to monitor and execute the simulator through a software wrapper. The software wrapper applies the initial taint to the DRAM, advances the simulation clock, and monitors taint sinks for classification. The resulting C ++ program is then compiled with BFSan, which adds logic for taint tracking. 5.2 Challenges Termination. Since hardware is reactive, that is, it does not terminate as long as a clock signal is provided, we face the problem of deciding when to stop the simulation for a given program. We employ the following strategy: 1. During architectural simulation (Spike), we terminate on any faulty instruction, thus ensuring we only execute the loaded program. In particular, we initialize all DRAM locations outside the loaded program to 0 – an illegal instruction in RISC-V – and ensure that our trap handler also contains an illegal instruction. This means that when the program reaches its end, the CPU will encounter an illegal instruction, which we use as a termination signal. Since our program may contain unbounded loops, we further put a bound on the total number of execution steps. If no illegal instruction is found, the simulation terminates after the maximum number of steps, which is calculated depending on the size of the input program. 2. During cycle-accurate simulation, we monitor the ReOrder Buffer (RoB) to count the number of retired instructions. Whenever this number matches the number of instructions reported by Spike, we end the simulation. Taint sources. Since all the input program’s code and data are loaded in memory at the start of each execution, we use DRAM as our initial taint source. In particular, we leverage BOOM’s option to provide a black-box implementation for the simulated DRAM. We use the list of memory accesses generated by the ISA simulator to initialize taint. In particular, we apply taint to all DRAM locations that have not been accessed architecturally. This includes locations whose initial value is overwritten by a subsequent store before being read. Taint sink(s). The simulation wrapper monitors the presence of taint in a predefined sink after each clock cycle. For our prototype, we chose the Physical Register File (PRF) as sink for two main reasons: (1) non-architectural data reaching the PRF can be leaked through a variety of side-channels, e.g., port contention, cache, TLB; (2) if tainted data reaches a physical register, we can easily infer which instruction is responsible for it by inspecting entries in the RoB. Taint washing. If a program speculatively jumps to a tainted value rather than loading it, taint might end up in the Register File. While this correctly implies that speculative code is being executed, we only care about speculative code that brings new data into the Register File, like Spectre gadgets for example. To avoid marking speculative code that does not directly leak values as a vulnerability, we make sure that taint is washed for instructions passing through the instruction cache, which prevents taint from spreading to the RF. Self-modifying code. Differently from x86, RISC-V architectures do not guarantee that the instruction cache is invalidated if code is modified during execution, and instead require explicit synchronization from software through FENCE.I. Programs that modify cached instructions without flushing the I-Cache are expected to produce a different behavior than the ISA simulation. For our use-case, this means that any program that modifies a load (e.g., by turning it into a nop ), will still observe the microarchitectural effects of that load, while the ISA simulation will not. To avoid reporting such cases, we detect programs that contain self-modifying code during the ISA simulation, and immediately discard the program without wasting time on the slow cycle-accurate simulation.
5.3 Classification To aid the analysis of the reported leaks, we perform an initial classification of the bug using the pipeline flush signal. In particular, instead of aborting the simulation immediately when taint reaches an exposed sink, the simulation continues executing the program and records: 1. The Taint Event, i.e., when taint is first observed in the Register File. We refer to the instruction responsible for this event as the tainting instruction, which can be derived by observing the Re-Order Buffer. 2. The Flush Event, i.e., a flush signal that squashes the taint instruction. We refer to the instruction that triggers the flush event as the flushing instruction. The simulator finally crashes whenever it detects that a pipeline flush is about to “remove” (squash) the tainting instruction from the pipeline. If taint is found in a sink but the corresponding instruction is never squashed, the crash is generated at the end of the test-case execution, i.e., when all the architecturally-executed instructions have retired from the pipeline, and the test case is marked as Unknown Flush. We use this information to perform a preliminary classification of the violation found. In particular, by observing the flush signal we can distinguish between Spectre violations (mispredictions), Meltdown violations (pipeline exceptions), and memory ordering faults (Spectre-v4). For Spectre variants other than Spectre-v4, we observe the flush instruction to determine if the misprediction was caused by a branch (Spectre-v1), indirect jump (Spectre-v2) or return (Spectre-RSB). For pipeline exceptions, we check if taint was introduced by the flush instruction itself (Meltdown) or by a younger instruction in the pipeline (OOO - Out-Of-Order). Finally, for branch mispredictions, we further report if the branch was predicted taken or not-taken, and if the tainting instruction was architecturally executed at least once before the taint detection. This allows us to differentiate between Spectre-v1-static (predicted not-taken, new instruction), Spectre-v1-training (predicted taken, previously executed instruction), and Spectre-v1-new (predicted taken, new instruction). 5.4 Extensions MDS Detection. MDS attacks, such as RIDL [66] and Fallout [13] and derivative attacks such as LVI [65] and CrossTalk [53], showed that, on Intel microarchitectures, an attacker can leak in-flight data from Line-Fill Buffers, Load Ports, and Store Buffers. Differently from traditional Spectre and Meltdown attacks, these vulnerabilities incorrectly access values in internal CPU buffers, as opposed to secrets in memory. These vulnerabilities can be modeled in Phantom Trails by adding such internal buffers as taint sources. In particular, we extended our prototype with a simple userspace initialization snippet (a set of loads and stores) that runs right before the start of the program under test, without any fence. Once the last instruction of the initialization snippet retires, Phantom Trails taints the initial values of all internal buffers, while the user program is ready to start. If the program is able to leak such stale values, e.g., through a faulty load, their taint will be observed in the Register File. Note that the user program is not allowed to directly access such values, so, whenever they are leaked, we are sure that there is a violation. As BOOM is not vulnerable to MDS, to test this setup we added a simple MDS-StoreBuffer vulnerability, as described by the Fallout [13] paper, to the BOOM core design, and verified that the leakage is detected. For the benefit of future research, we open-source the patch for adding the vulnerability to BOOM. Secure Speculation. Phantom Trails can be extended to incorporate knowledge of both software and hardware defenses. For instance, the instruction generator can be constrained to always emit an LFENCE [37] after each branch, to mimic cases where this mitigation is deployed. For secure speculation defenses such as STT [74], finer-grained detection policies can be added for taint sinks, e.g., discarding tainted entries that are read by instructions deemed "safe" by STT. Other Taint Sources. Similarly to MDS, other data sampling attacks such as Gather Data Sampling [46], AEPIC Leak [10] and ZenBleed [51] have been shown to be possible on x86 cores. In particular, Downfall [46] shows that the gather instruction can transiently leak stale data from a temporal buffer called the SIMD register buffer, confirmed by Intel. AEPIC Leak and ZenBleed instead can read stale data architecturally from the superqueue (buffer between L2 and LLC) and XMM registers, respectively. Similarly to the MDS case, Phantom Trails can be extended to handle more taint sources by making sure such internal buffers are initialized and tainted right before the start of the program. For initial taint residing in the Register File, more fine-grained taint sink policies can be applied to ignore specific initially-tainted locations until they are overwritten by another operation. 6 Fuzzing This section describes how we integrated Phantom Trails’ detector into a pre-silicon fuzzer, and highlights the benefits of our detection model to the fuzzing use-case. 6.1 Overview Phantom Trails resembles a traditional greybox fuzzer that exercises a software representation of the hardware as the DUT [64]. Figure 6presents a high-level overview of its components. In particular we used the setup described in Section 7.1 as the DUT in our fuzzing campaigns, which includes
ISA simulator mutated program program + metadata Verilated BOOM core BFSAN + AFL instrumentation detection 3 45 coverage Mutator pick random mutate fuzzer queue DETECTOR FUZZER Software wrapper 6 1 2 Instr. Generator Figure 6: Phantom Trails’s fuzzing cycle. A sequence of instructions is picked randomly from the fuzzer queue 1 and modified by the mutator 2 . The resulting sequence is then translated into a RISCV program 3 and executed by the ISA simulator. If the program is not discarded, the metadata is passed to a cycle-accurate simulator 4 where, in case of detection, the program will be saved 5 . If the program execution produced new coverage in the simulator, it is also added to the fuzzer queue 6 . a minimal setup for the BOOM core in its MEDIUMBOOM configuration and a black-box DRAM module. Fuzzing infrastructure. We build our fuzzer on top of the state-of-the-art libafl [23] fuzzing framework and run it in fork mode—forking after the completion of a hardware reset to avoid the cost of restarting the simulation on each input. To adapt the libafl software-based infrastructure to hardware designs, we developed a set of components suitable for generalized hardware fuzzing. Our entire infrastructure is open-source and available at https://github.com/ vusec/phantom-trails. DUT warmup. Before any code is run, the hardware simulation is reset through the default Verilator wrapper for BOOM by asserting the reset signal for 100 cycles. Boot phase. Upon reset, execution starts from the content of the boot ROM, which we modify to simply jump to the first DRAM address. At the beginning of DRAM we place our initialization code, which is responsible for: 1. Setting up the trap handler, which in our case simply contains an illegal instruction to terminate simulation; 2. Configuring the Physical Memory Protection (PMP) unit to permit access to all memory; 3. Setting up page tables and enabling virtual memory management; in particular, we map a contiguous set of pages starting from the beginning of DRAM with different page flags; 4. Optionally, initializing register values (optimization D2); 5. Jumping to U-mode (code with user privileges), where the input program is located; Predictors initialization. By default, the BOOM processor initializes all entries of the Bi-Modal Table ( BIM ) to 2 on reset, which corresponds to the “weakly taken” state. While this does not prevent detection, it can cause the classifier to mistake cases of static branch prediction for cases where the branch was trained, and incorrectly label Spectre-v1 samples. To ensure a correct fine-grained classification, we modify the BIM initialization procedure to instead set entries to the “not-taken” state. Taint initialization. While in supervisor-mode, the initialization code reads from the supervisor data region, which fills the D-Cache with tainted (supervisor) data. As discussed in Section 5.4, and optional user-mode initialization can be performed for MDS to fill internal CPU buffers, such as the Load Queue or the Store Buffer, with tainted data as well, right before the start of the generated program. 6.2 Program Generation As stated in section 3, one of the entropy sources that the fuzzer has to beat is program generation. While sophisticated approaches [59] can be added on top of our fuzzer, in this paper we want to demonstrate that our detector already helps even with a minimal program generator. In particular, in our prototype we make sure to generate and mutate syntactically valid RISC-V instructions, and we adopt a set of minimal optimizations to increase the chance of generating complex control and data flow. Such optimizations differ from templates, as they are aimed at maximizing the odds of generated well-formed, complex programs rather than following the blueprint of a specific vulnerability. 6.2.1 Instructions Mutator. Phantom Trails’ custom program mutator is aware of what constitutes syntactically valid RISC-V instructions, but possesses no further (semantic) information about them. In its basic form, it generates instructions by choosing a random RISC-V instruction type and applying a random mutation operation. In particular, the current prototype supports inserting a new instruction, replacing an instruction with a new one, replacing the argument of an existing instruction, repeating an existing instruction, deleting an existing instruction, replacing
an instruction with a nop , and swapping two instructions. Optionally, the mutator can be biased towards emitting jalr and ret instructions, and towards reusing previous values when choosing arguments, as we will discuss in Section 6.2. Program generator backend. Since applying random mutations like bit-flips at the assembly level has a high chance of generating invalid programs which would waste precious simulation time, the fuzzer instead uses a structured internal representation to apply mutations. Programs stored in this internal representation are then translated into valid RISC-V programs by the instruction generator, before entering the detector component. 6.2.2 Optimizations To avoid wasting simulation cycles on uninteresting inputs and to maximize the likelihood of finding bugs quickly, we develop a set of optimizations that bias our program generation towards valid programs. Unlike templates [33], the optimizations are general so as to avoid overfitting. We group our optimizations into: Basic (B), biasing the generator towards reusing arguments, Control-Flow (C), maximizing the probability of generating well-formed function calls, and Data-Flow (D) optimizations that increase the likelihood of using valid pointers (code and data). B1 - Register reuse. With this optimization, when deciding on the argument of an instruction, the mutator has a bias towards selecting the registers used by previous instructions (e.g., a probability of 50% in the current prototype). The idea is to improve the chance of generating data flow between instruction sequences, as well as that of creating race conditions in the microarchitecture through aliasing. B2 - Power-of-two constants. When picking immediate values, this optimization adds a bias towards powers of 2, which reduces the amount of entropy for constants and helps with alignment. C1 - Indirect calls. To help the fuzzer reduce the entropy for indirect calls, this optimization adds the possibility of inserting one of two code snippets shown in Listing 1. Snippet 1 performs a well-formed indirect jump to a nearby address (i.e., a small offset from the current program counter); Snippet 2 emits a valid ret instruction, which, in RISC-V, is a pseudonym for an indirect jump to ra . While this does not guarantee the generation of a valid indirect call, it improves the chances of executing valid calls and returns. Note that more sophisticated approaches [59,72] to program generation can build on top of such a basic optimization—albeit at the cost of additional overhead. C2 - Discard invalid jumps. To further reduce the likelihood of wasting simulation cycles on programs with invalid controlflow, the ISA simulator terminates the execution immediately whenever a program jumps outside the range of valid mem- # Load current PC auipc x2,0 # Jump to PC + offset jalr ra,rand_offset(x2) (a) Snippet 1: Indirect call # Return jalr zero,0(ra) (b) Snippet 2: Return Listing 1: Control-flow snippets inserted by the mutator. ory locations, and discards the corresponding input program. This guarantees that, at least for jump and call instructions, Phantom Trails runs the expensive cycle-accurate simulation only if the program jumps to valid targets. D1 - Map address 0. Since most of the memory is filled with 0s at startup, and it is not uncommon for predictors to default to address 0 [36] on empty prediction structures, this optimization makes sure virtual address 0 is mapped to a valid memory page before the input program starts execution. D2 - Initialize registers. This optimization ensures that some registers are filled with valid pointers before jumping to the input program’s code. We fill half of the logical Register File with addresses of both code and data pages that have an associated page table entry. This means that, whenever an instruction uses a register for the first time there is a 50% chance that it will use one of the initialized pointers. 6.3 Feedback Currently, there is no consensus on the best feedback metric for hardware fuzzing [59], nor is there is a “standard” strategy for transient execution fuzzers. As we want to show the advantages of our detection model on a simple fuzzer, we use as baseline feedback the standard coverage metric provided by the AFL ++ software fuzzer, which we call ‘SW Feedback’. To evaluate additionally Phantom Trails’s sensitivity to feedback metrics, we additionally implemented an alternative, taint-based feedback mechanism which is inserted into the cycle-accurate simulation via an LLVM pass, to explore the possibility of using taint as feedback. SW Feedback. In this case, the metric is an approximation of the edge coverage of the system-under-test as described in previous work [64]. We adapted this metric by only counting whether an edge in the simulator has been executed at all, and not how often it was executed. Doing so avoids labelling mutations that merely traverse the same edges as interesting, while adding little relevance to the program. Taint Feedback. Instead of tracking the edge coverage of the simulator during the input program’s execution, this metric tries to measure how much taint has spread through the design—across all the CPU’s wires. Since most Verilogspecific information, including the list of wires, is lost during Verilator’s translation process, we identify the code for each
[14] Boru Chen, Yingchen Wang, Pradyumna Shome, Christopher W Fletcher, David Kohlbrenner, Riccardo Paccagnella, and Daniel Genkin. Gofetch: Breaking constant-time cryptographic implementations using data memory-dependent prefetchers. In USENIX Security, 2024. [15] Chen Chen, Vasudev Gohil, Rahul Kande, Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. Psofuzz: Fuzzing processors with particle swarm optimization. In ICCAD, 2023. [16] Chen Chen, Rahul Kande, Nathan Nguyen, Flemming Andersen, Aakash Tyagi, Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. { HyPFuzz } : { Formal-Assisted } processor fuzzing. In USENIX Security, 2023. [17] ChipsAlliance. Chisel. https://www.chisel-lang. org/. [18] Yaakov Cohen, Kevin Sam Tharayil, Arie Haenel, Daniel Genkin, Angelos D Keromytis, Yossi Oren, and Yuval Yarom. Hammerscope: observing DRAM power consumption using Rowhammer. In CCS, 2022. [19] S. Dinesh, M. Parthasarathy, and C. Fletcher. Conjunct: Learning inductive invariants to prove unbounded instruction safety against microarchitectural timing attacks. In IEEE S&P, 2024. [20] Dmitry Evtyushkin, Ryan Riley, Nael CSE AbuGhazaleh, ECE, and Dmitry Ponomarev. Branchscope: A new side-channel attack on directional branch predictor. ACM SIGPLAN Notices, 2018. [21] Mohammad Rahmani Fadiheh, Alex Wezel, Johannes Müller, Jörg Bormann, Sayak Ray, Jason M Fung, Subhasish Mitra, Dominik Stoffel, and Wolfgang Kunz. An exhaustive approach to detecting transient execution side channels in rtl designs of processors. IEEE Transactions on Computers, 2022. [22] Andrea Fioraldi, Dominik Maier, Heiko Eißfeldt, and Marc Heuse. AFL++: Combining incremental steps of fuzzing research. In WOOT, 2020. [23] Andrea Fioraldi, Dominik Christian Maier, Dongjia Zhang, and Davide Balzarotti. Libafl: A framework to build modular and reusable fuzzers. In CCS, 2022. [24] Jacob Fustos, Michael Bechtel, and Heechul Yun. SpectreRewind: Leaking secrets to past instructions. In ASHES, 2020. [25] Moein Ghaniyoun, Kristin Barber, Yinqian Zhang, and Radu Teodorescu. Introspectre: A pre-silicon framework for discovery and analysis of transient execution vulnerabilities. In ISCA, 2021. [26] Klaus v Gleissenthall, Rami Gökhan Kıcı, Deian Stefan, and Ranjit Jhala. { IODINE } : Verifying { ConstantTime } execution of hardware. In USENIX Security, 2019. [27] Enes Göktas, Kaveh Razavi, Georgios Portokalidis, Herbert Bos, and Cristiano Giuffrida. Speculative probing: Hacking blind in the Spectre era. In CCS, 2020. [28] Google. Retpoline: a software construct for preventing branch-targetinjection. https://support.google. com/faqs/answer/7625886. [29] Ben Gras, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. TLBleed: When Protecting Your CPU Caches is not Enough. In Black Hat USA, 2018. [30] Daniel Gruss, Clémentine Maurice, Anders Fogh, Moritz Lipp, and Stefan Mangard. Prefetch side-channel attacks: Bypassing smap and kernel aslr. In CCS, 2016. [31] Mathé Hertogh, Sander Wiebing, and Cristiano Giuffrida. Leaky Address Masking: Exploiting Unmasked Spectre Gadgets with Noncanonical Address Translation. In IEEE S&P, 2024. [32] Muhammad Monir Hossain, Nusrat Farzana Dipu, Kimia Zamiri Azar, Fahim Rahman, Farimah Farahmandi, and Mark Tehranipoor. Taintfuzzer: Soc security verification using taint inference-enabled fuzzing. In ICCAD, 2023. [33] Jaewon Hur, Suhwan Song, Sunwoo Kim, and Byoungyoung Lee. Specdoctor: Differential fuzz testing to find transient execution vulnerabilities. In CCS, 2022. [34] Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, and Byoungyoung Lee. Difuzzrtl: Differential fuzz testing to find cpu bugs. In IEEE S&P, 2021. [35] Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, and Byoungyoung Lee. Difuzzrtl: Differential fuzz testing to find cpu bugs. In IEEE S&P, 2021. [36] Intel. Bhi disclosure documentation. https://www.intel.com/content/ www/us/en/developer/articles/ technical/software-security-guidance/ technical-documentation/ branch-history-injection.html. [37] Intel. Intel analysis of speculative execution side channels. https://www.intel.com/ content/www/us/en/developer/articles/ technical/software-security-guidance/ technical-documentation/
analysis-speculative-execution-side-channels. html. [38] Rahul Kande, Addison Crump, Garrett Persyn, Patrick Jauernig, Ahmad-Reza Sadeghi, Aakash Tyagi, and Jeyavijayan Rajendran. TheHuzz: Instruction fuzzing of processors using Golden-Reference models for finding Software-Exploitable vulnerabilities. In USENIX Security, 2022. [39] Paul Kocher, Jann Horn, Anders Fogh, , Daniel Genkin, Daniel Gruss, Werner Haas, Mike Hamburg, Moritz Lipp, Stefan Mangard, Thomas Prescher, Michael Schwarz, and Yuval Yarom. Spectre attacks: Exploiting speculative execution. In IEEE S&P, 2019. [40] Esmaeil Mohammadian Koruyeh, Khaled N. Khasawneh, Chengyu Song, and Nael Abu-Ghazaleh. Spectre returns! speculation attacks using the return stack buffer. In WOOT, 2018. [41] Chris Lattner and Vikram Adve. Llvm: A compilation framework for lifelong program analysis & transformation. In International symposium on code generation and optimization, 2004. CGO 2004. IEEE, 2004. [42] Moritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher, Werner Haas, Anders Fogh, Jann Horn, Stefan Mangard, Paul Kocher, Daniel Genkin, Yuval Yarom, and Mike Hamburg. Meltdown: Reading kernel memory from user space. In USENIX Security, 2018. [43] Kevin Loughlin, Ian Neal, and Jiacheng Ma. DOLMA: Securing speculation with the principle of transient nonobservability. In USENIX Security, 2021. [44] Giorgi Maisuradze and Christian Rossow. ret2spec: Speculative execution using return stack buffers. In CCS, 2018. [45] Alyssa Milburn, Ke Sun, and Henrique Kawakami. You cannot always win the race: Analyzing mitigations for branch target prediction attacks. In EuroS&P, 2023. [46] Daniel Moghimi. Downfall: Exploiting speculative data gathering. In USENIX Security, 2023. [47] Daniel Moghimi, Moritz Lipp, Berk Sunar, and Michael Schwarz. Medusa: Microarchitectural data leakage via automated attack synthesis. In USENIX Security, 2020. [48] Hamed Nemati, Pablo Buiras, Andreas Lindner, Roberto Guanciale, and Swen Jacobs. Validation of abstract sidechannel models for computer architectures. In CAV, 2020. [49] Oleksii Oleksenko, Christof Fetzer, Boris Köpf, and Mark Silberstein. Revizor: Testing black-box cpus against speculation contracts. In ASPLOS, 2022. [50] Oleksii Oleksenko, Marco Guarnieri, Boris Köpf, and Mark Silberstein. Hide and seek with spectres: Efficient discovery of speculative information leaks with random testing. In IEEE S&P, 2023. [51] Tavis Ormandy. Zenbleed. https://lock. cmpxchg8b.com/zenbleed.html. [52] Hany Ragab, Enrico Barberis, Herbert Bos, and Cristiano Giuffrida. Rage against the machine clear: A systematic analysis of machine clears and their implications for transient execution attacks. In USENIX Security, 2021. [53] Hany Ragab, Alyssa Milburn, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. Crosstalk: Speculative data leaks across cores are real. In IEEE S&P, 2021. [54] Chathura Rajapaksha, Leila Delshadtehrani, Manuel Egele, and Ajay Joshi. Sigfuzz: A framework for discovering microarchitectural timing side channels. In DATE, 2023. [55] Xida Ren, Logan Moody, Mohammadkazem Taram, Matthew Jordan, Dean M Tullsen, and Ashish Venkat. I see dead micro-ops: Leaking secrets via Intel/AMD micro-op caches. In ISCA, 2021. [56] Michael Schwarz, Martin Schwarzl, Moritz Lipp, Jon Masters, and Daniel Gruss. NetSpectre: Read arbitrary memory over network. In ESORICS, 2019. [57] Konstantin Serebryany, Derek Bruening, Alexander Potapenko, and Dmitriy Vyukov. AddressSanitizer: A Fast Address Sanity Checker. In USENIX ATC, 2012. [58] Wilson Snyder. Verilator. https://www.veripool. org/verilator/. [59] Flavien Solt, Katharina Ceesay-Seitz, and Kaveh Razavi. Cascade: Cpu fuzzing via intricate program generation. In USENIX Security, 2024. [60] Flavien Solt, Ben Gras, and Kaveh Razavi. CellIFT: Leveraging cells for scalable and precise dynamic information flow tracking in RTL. In USENIX Security, 2022. [61] Evgeniy Stepanov and Konstantin Serebryany. MemorySanitizer: fast detector of uninitialized memory use in C++. In CGO, 2015. [62] Andrei Tatar, Daniël Trujillo, Cristiano Giuffrida, and Herbert Bos. { TLB; DR } : Enhancing { TLB-based } attacks with { TLB } desynchronized reverse engineering. In USENIX Security, 2022.
[63] Youssef Tobah, Andrew Kwong, Ingab Kang, Daniel Genkin, and Kang G Shin. SpecHammer: Combining Spectre and Rowhammer for new speculative attacks. In IEEE S&P, 2022. [64] Timothy Trippel, Kang G. Shin, Alex Chernyakhovsky, Garret Kelly, Dominic Rizzo, and Matthew Hicks. Fuzzing hardware like software. In USENIX Security, 2022. [65] Jo Van Bulck, Daniel Moghimi, Michael Schwarz, Moritz Lipp, Marina Minkin, Daniel Genkin, Yarom Yuval, Berk Sunar, Daniel Gruss, and Frank Piessens. LVI: Hijacking Transient Execution through Microarchitectural Load Value Injection. In IEEE S&P, 2020. [66] Stephan van Schaik, Alyssa Milburn, Sebastian Österlund, Pietro Frigo, Giorgi Maisuradze, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. Ridl: Rogue inflight data load. In IEEE S&P, 2019. [67] Zilong Wang, Gideon Mohr, Klaus von Gleissenthall, Jan Reineke, and Marco Guarnieri. Specification and verification of side-channel security for open-source processors via leakage contracts. In CCS, 2023. [68] Daniel Weber, Ahmad Ibrahim, Hamed Nemati, Michael Schwarz, and Christian Rossow. Osiris: Automated discovery of microarchitectural side channels. In USENIX Security, 2021. [69] Sander Wiebing, Alvise de Faveri Tron, Herbert Bos, and Cristiano Giuffrida. InSpectre gadget: Inspecting the residual attack surface of cross-privilege spectre v2. In USENIX Security, 2024. [70] Johannes Wikner and Kaveh Razavi. RETBLEED: Arbitrary speculative code execution with return instructions. In USENIX Security, 2022. [71] Johannes Wikner and Kaveh Razavi. Breaking the Barrier: Post-Barrier Spectre Attacks. In IEEE S&P, 2025. [72] Jinyan Xu, Yiyuan Liu, Sirui He, Haoran Lin, Yajin Zhou, and Cong Wang. { MorFuzz } : Fuzzing processor via runtime instruction morphing enhanced synchronizable co-simulation. In USENIC Security, 2023. [73] Yuval Yarom and Katrina Falkner. FLUSH+RELOAD: A high resolution, low noise, l3 cache Side-Channel attack. In USENIX Security, 2014. [74] Jiyong Yu, Mengjia Yan, Artem Khyzha, Adam Morrison, Josep Torrellas, and Christopher W. Fletcher. Speculative taint tracking (stt): A comprehensive protection for speculatively accessed data. In MICRO, 2019. [75] Michal Zalewski. American fuzzy lop. https:// github.com/google/AFL.