Accurate Analysis of Silent Data Corruptions in Programmable AI Accelerator Microarchitectures
Full text
Accurate Analysis of Silent Data Corruptions in Programmable AI Accelerator Microarchitectures Odysseas Chatzopoulos Maria Trakosa Dimitris Gizopoulos University of Athens, Greece {od.chatzopoulos |mariatrak |dgizop}@di.uoa.gr Abstract—Programmable AI accelerators become increasingly important to modern computing infrastructure, thus, their reliability is critical for the integrity of the produced results. Silent Data Corruptions (SDCs)—incorrect program outputs that occur without any warning or notification—have been reported by hyperscalers such as Meta, Google, and Alibaba, affecting both CPUs and AI chips in production environments. SDCs originate from a range of low-level causes including manufacturing defects, aging-induced degradation, process variation, particle strikes, and electromagnetic interference. In this work, we revisit the modeling debate between software-level and microarchitecturelevel fault injection for estimating SDC vulnerability, in the context of programmable AI accelerators. While software-level (hardware agnostic) techniques are fast and easy to deploy, studies on CPUs and GPUs have shown they produce misleading results due to their lack of the hardware notion which determines faults propagation or filtering. We show that these issues also persist dramatically in AI accelerators. Using detailed microarchitectural modeling, we demonstrate that even so-called hardware-aware software-level approaches can misestimate FIT rates by more than 4×across realistic accelerator configurations. Our findings support microarchitecture-level simulation as the most effective tradeoff point between accuracy and scalability for early-stage reliability analysis of programmable AI hardware. Index Terms—AI accelerators, reliability, silent data corruptions, silicon defects, fault injection, microarchitectural modeling I. INTRODUCTION Artificial Intelligence (AI) is reshaping modern computing, enabling breakthroughs across domains (healthcare, autonomous systems, natural language processing, scientific research, to name a few). The increasing complexity and scale of AI workloads—particularly deep learning—have driven the need for specialized hardware accelerators [1]–[5], which provide higher performance and energy efficiency than CPUs and GPUs. While performance has been the main design driver, reliability has emerged as an equally critical concern—especially at scale. Silent Data Corruptions (SDCs)—incorrect program outputs that occur without any warning or notification—pose a growing threat. They can originate from manufacturing defects, aging, process variation, or external particle strikes, and often escape conventional detection mechanisms. Hyperscalers (Meta, Google, Alibaba) have reported alarmingly high SDC occurrences across CPUs and AI accelerators. Meta estimates that approximately one CPU chip per thousand in their production fleets experiences SDCs [6]–[8]. Google similarly reports SDCs in their machine learning chips [9], while large-scale training disruptions have been attributed to undetected hardware faults in accelerator deployments [10], [11]. While significant research has focused on reliability analysis in CPUs [12]–[21], AI accelerators present new challenges. Their dataflow-oriented microarchitectures, specialized compute paths, and deterministic control flow introduce unique vulnerabilities. The fault impact is often entangled with the AI model structure, the memory hierarchy, and data reuse pattern—demanding tailored reliability methodologies. Early-stage modeling (and the broad design space exploration it offers) is key not only for performance evaluation but also for reliability assessment. Traditional cycleaccurate microarchitectural simulators [22] support detailed performance analysis, but modeling reliability requires the ability to inject faults and propagate their effects through the full stack. Fault injection can be performed at several abstraction levels. Software-level injection—targeting highlevel variables or model parameters—is easy to implement and fast. However, it completely ignores the hardware behavior, and has been shown to lead to misleading conclusions on CPUs [23] and GPUs [24]. On the other end, RTL and gatelevel fault injection provide high accuracy but at the cost of prohibitively low simulation throughput. Microarchitectural simulation is a practical balance point: it offers sufficient accuracy for modeling fault propagation in hardware structures while maintaining very high throughput for large-scale studies [23], [25]. For AI accelerators in particular, we argue that this abstraction level is essential to capture timing and the data movements that affect fault impact. In this work, we revisit the debate between softwareand microarchitecture-level fault modeling in the context of AI accelerators. We show that many recently proposed softwarebased modeling techniques suffer from severe inaccuracies, and we present measurements using a detailed modeling infrastructure capable of full-system simulation and hardwareaware Statistical Fault Injection (SFI). Our goal is to highlight the modeling choices that matter—and to motivate the use of microarchitectural-level approaches for reliability evaluation of future AI hardware (analysis of SDCs and other fault effects) and design of countermeasures. II. SOFTWARE VS MICROARCHITECTURE-LEVEL FAULT INJECTION IN CPUS AND GPUS The debate between abstraction levels for fault injection—software-level versus microarchitecture-level—has979-8-3315-3334-2/25/$31.00 ©2025 IEEE
long been active in the reliability research community, particularly for CPUs and GPUs. While software-level methods offer speed and ease of deployment, they abstract away hardware details, leading to misleading conclusions about fault behavior and Silent Data Corruptions (SDCs). For CPUs, Papadimitriou and Gizopoulos [23] perform a detailed analysis using SFI at software and microarchitecture levels. They show that software-based injection—typically done via LLVM-level perturbations—dramatically underestimates SDC rates and misrepresents fault propagation. Their results highlight that fault masking, propagation latency, and architectural dependencies are only accurately captured at the microarchitecture level. In some CPU workloads, softwarelevel techniques underestimated SDC rates by up to 4× compared to microarchitecture-level injection, risking false confidence in resilience analysis and driving wrong mitigation. Yang et al. [24] extend this analysis to the dataparallel architecture of modern GPUs and present a crosslayer study comparing microarchitecture-level injection (using gpuFI-4) against software-level injection (via NVBitFI). Their study demonstrates that software-level vulnerability estimation (measured via Software Vulnerability Factor, SVF) not only misestimates the absolute vulnerability, but also yields inconsistent application ranking compared to the accurate metric of Architectural Vulnerability Factor (AVF) derived from microarchitecture-level injection. In over 40% of application pairs, SVF and AVF disagreed on which workload was more vulnerable. Even under strong protections like Triple Modular Redundancy (TMR), software-level methods falsely indicated complete SDC mitigation, while AVF-based analysis revealed lingering corruption and increased Detected Unrecoverable Errors (DUEs). These discrepancies were traced back to software-level methods ignoring faults in inactive hardware bits and undervaluing control/data path interactions. These studies collectively highlight that software-level fault modeling and vulnerability analysis which lack the notion of hardware, can lead to dramatically misleading conclusions about system resilience - and misguide mitigation plans. As AI accelerators inherit many characteristics from CPUs and GPUs—such as pipeline depth, hardware-managed parallelism, and sensitive control/data flows—there is reason to suspect similar pitfalls in using high-level abstractions for reliability assessment. III. REVISITING THE SOFTWARE VS MICROARCHITECTURE LEVEL FAULT INJECTION DEBATE FOR AI ACCELERATORS AI accelerators while similar in some ways to CPUs and GPUs have unique characteristics that set them apart from their more general-purpose counterparts. Their dataflowcentric compute model, tightly coupled memory hierarchies, and highly parallel execution pipelines make them especially vulnerable to Silent Data Corruptions (SDCs) that arise from low-level hardware faults. These properties demand fault modeling techniques that are accurate, microarchitecture-aware, but also fast. Evaluation of AI accelerator reliability can be performed at the same abstraction levels. At the lowest level, gate-level or RTL injection offers the highest accuracy but suffers from prohibitively low throughput. At the other extreme, software-level fault injection—such as those perturbing inputs, weights, or intermediate tensors [26], [27]—scale well but completely abstract away the underlying hardware, making them fundamentally incapable of modeling datapath or memory hierarchy effects which significantly affect timing. Hardware-aware (but still software-level) methods attempt to bridge this gap by mapping hardware faults to software-visible artifacts. However, they still rely on assumptions that may not hold across system configurations. To assess the accuracy of these approaches and explore the design space of AI accelerators under realistic fault conditions, we employ a comprehensive microarchitectural simulation and fault injection infrastructure. This infrastructure supports fullsystem modeling of programmable, systolic-array-based AI accelerators within the state-of-the-art gem5 simulation environment. The modeled compute stack can be seen in Fig. 1. Our infrastructure enables detailed performance and reliability exploration across a wide range of accelerator configurations, dataflows, datatypes, and memory hierarchies. Critically, it includes a statistical fault injection engine targeting all major architectural elements: SRAMs, register files, and functional units and injects faults at the bit level and the gate level. Fig. 1. The compute stack modeled in our microarchitecture-level framework. Due to the well-documented shortcomings of softwarelevel fault injection (for CPUs and GPUs), several efforts have focused on incorporating hardware-aware “proxies” to improve modeling fidelity without incurring the overhead of full hardware simulation. Among these, FIdelity [28] stands out as a representative example. FIdelity introduces the Reuse Factor (RF), a metric designed to approximate the likelihood that a transient fault in a flip-flop will manifest as an error in software-visible outputs. This is achieved by statically analyzing which registers are reused in the computation of model parameters such as weights and activations. In theory, registers with higher reuse are more likely to propagate errors if perturbed. While this approach is a step that aims to fix the pitfalls of na¨ ıve software-level injection, it is still built on highly optimistic assumptions. The authors claim that the RF can be simply derived from high-level architectural block diagrams, bypassing the need for RTL or low-level microarchitectural
details. However, in practice, understanding whether and how a transient fault will propagate through the hardware pipeline requires accurate modeling of timing, control logic, and pipeline stalls—all of which are typically abstracted away in high-level architectural views. Fault propagation behavior is tightly coupled with these dynamics, and simplistically assigning reuse values risks overlooking critical masking or error amplification paths. Moreover, FIdelity attempts to refine fault impact estimates through an Activity Analysis step that accounts for whether the flip-flops under consideration are “temporally inactive.” In this context, temporal inactivity refers to flip-flops that are stalled or unused while waiting on long-latency operations such as memory fetches. In their NVDLA case study, FIdelity estimates this activity profile using an NVIDIA-provided performance analysis tool. This dependence on such tools or detailed RTL models introduces a serious limitation: it hinders portability and prevents the methodology from being applied to arbitrary accelerators or memory hierarchies. As SoC designs grow more heterogeneous and complex, these simplifying assumptions become increasingly untenable. Fig. 2. Architecture of the modeled System-on-Chip. Arrows show the different locations in the memory hierarchy where the accelerator can be connected. To investigate how sensitive such simplistic activity-based modeling is to microarchitectural parameters, we used our infrastructure to measure the actual register activity of processing elements in an AI accelerator while varying the parameters of the L2 cache (similar results hold for other hardware units). The high-level block diagram of the accelerator can be seen in Fig. 2. We changed both the cache size (four sizes) and the replacement policy (LRU vs. MRU) and recorded the PE register utilization while executing the same inference workload. The results, shown in Figure 3, demonstrate that even such small system-level changes can significantly alter the observed activity levels. In our experiments, switching from LRU to MRU in a 1MB cache caused a 3.82×reduction in register activity. Overall, we observed nearly a 4×variation in activity between configurations. This variability is especially problematic for resilience analysis because FIdelity’s FIT rate estimation depends directly on these activity levels. In other words, if the activity estimation is off by a factor of four, so will the predicted reliability degradation. This illustrates the brittleness of proxy-based modeling: what appears accurate under one configuration may be wildly off under another, even for the same workload. Fig. 3. Register activity measured in our infrastructure. The accelerator is attached to the L2 cache and we vary its size and replacement policy. This shows the potential accuracy loss of FIdelity [28]. In contrast, our methodology—built on cycle-level, fullsystem simulation—avoids such approximations entirely. By simulating the entire workload execution across compute stack layers having realistic traffic through memory hierarchies, we observe flip-flop usage, error propagation, and masking in context. There is no reliance on approximate tools, architectural guesswork, or static reuse analysis. The simulator natively tracks architectural state and temporal behavior, ensuring accurate fault modeling regardless of configuration. Frameworks that inherit the FIdelity methodology, such as Thales [29], also suffer from these weaknesses. Any conclusions drawn about system vulnerability or the effectiveness of mitigation strategies are fundamentally limited by the fidelity of the underlying system model. Without capturing memory stalls, instruction-level concurrency, and register lifetimes in situ, these models risk misrepresenting the actual SDC risk. IV. CONCLUSION As AI accelerators become foundational components of hyperscale infrastructure for ML workloads training and inference, the need for accurate Silent Data Corruption (SDC) and other fault effects modeling is more pressing than ever. While software-level fault injection remains popular due to its ease of deployment, our analysis shows that even hardwareaware proxy methods like FIdelity fall short of capturing the complex, timing-sensitive behavior of real systems. These methods rely on unverifiable assumptions and lack the configurational sensitivity needed to reflect realistic system-level effects. Our case study demonstrates that minor changes in memory hierarchy—such as cache policy—can lead to over 4×differences in register activity, directly impacting SDC rate estimation. In contrast, our microarchitecture-level infrastructure enables detailed, in-context fault modeling that reflects true system behavior across diverse accelerator configurations. We conclude that only microarchitectural simulation provides the fidelity required for trustworthy resilience analysis and should be the preferred methodology in future design-time evaluations of AI hardware.
ACKNOWLEDGMENTS This work is supported through research gifts by Meta, AMD, and the Open Compute Project (OCP). It is also supported by the European Union’s Horizon Europe research and innovation programme under grant agreement No 101093062 (Vitamin-V), the Chips JU grant No 101097224 (REBECCA), and the EuroHPC JU grant No 101202459 (DARE SGA1). Views and opinions expressed are however, those of the authors only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them. REFERENCES [1] S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” in Proceedings of the 43rd International Symposium on Computer Architecture (ISCA ’16), 2016, p. 243–254. [Online]. Available: https://doi.org/10.1109/ISCA.2016.30 [2] Y. Turakhia, G. Bejerano, and W. J. Dally, “Darwin: A genomics coprocessor provides up to 15,000x acceleration on long read assembly,” in Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’18), 2018, p. 199–213. [Online]. Available: https://doi.org/10.1145/3173162.3173193 [3] W. Qadeer, R. Hameed, O. Shacham, P. Venkatesan, C. Kozyrakis, and M. A. Horowitz, “Convolution engine: Balancing efficiency & flexibility in specialized computing,” in Proceedings of the 40th Annual International Symposium on Computer Architecture (ISCA ’13), 2013, p. 24–35. [Online]. Available: https://doi.org/10.1145/2485922.2485925 [4] M. Maddury, P. Kansal, and O. Wu, “Next gen mtia-recommendation inference accelerator,” in 2024 IEEE Hot Chips 36 Symposium (HCS). IEEE, 2024, pp. 1–27. [5] N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. A. Patterson, “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3579371.3589350 [6] H. D. Dixit, S. Pendharkar, M. Beadon, C. Mason, T. Chakravarthy, B. Muthiah, and S. Sankar, “Silent Data Corruptions at Scale,” 2021. [Online]. Available: https://arxiv.org/abs/2102.11245 [7] P. H. Hochschild, P. Turner, J. C. Mogul, R. Govindaraju, P. Ranganathan, D. E. Culler, and A. Vahdat, “Cores That Don’t Count,” in Proceedings of the Workshop on Hot Topics in Operating Systems, ser. HotOS ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 9–16. [Online]. Available: https://doi.org/10.1145/3458336.3465297 [8] S. Wang, G. Zhang, J. Wei, Y. Wang, J. Wu, and Q. Luo, “Understanding silent data corruptions in a large production cpu population,” in Proceedings of the 29th Symposium on Operating Systems Principles, ser. SOSP ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 216–230. [Online]. Available: https://doi.org/10.1145/3600006.3613149 [9] Y. He, M. Hutton, S. Chan, R. De Gruijl, R. Govindaraju, N. Patil, and Y. Li, “Understanding and mitigating hardware failures in deep learning training systems,” in Proceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3579371.3589105 [10] R. Bonderson, “Training in turmoil: Silent data corruption in systems at scale,” in International Test Conference Silicon Lifecycle Management Workshop, 2021, keynote Presentation. [Online]. Available: https://marcello.altervista.org/SLM.tttcevents.org/program.html#Keynote1 [11] J. Dean and A. Vahdat, “Exciting directions for ml models and the implications for computing hardware,” in Keynote 1, Hot Chips 2023, 2023. [12] N. Karystinos, O. Chatzopoulos, G. Fragkoulis, G. Papadimitriou, G. Gizopoulos, and S. Gurumurthi, “Harpocrates: Breaking the silence of cpu faults through hardware-in-the-loop program generation,” in Proceedings 51st ACM/IEEE International Symposium on Computer Architecture (ISCA), 2024. [13] O. Chatzopoulos, N. Karystinos, G. Papadimitriou, D. Gizopoulos, H. D. Dixit, and S. Sankar, “Veritas – demystifying silent data corruptions: µarch-level modeling and fleet data of modern x86 cpus,” in 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA), March 2025. [14] R. Dutta, H. D. Dixit, R. Van Riel, G. Vunnam, and S. Sankar, “Hardware sentinel: Protecting software applications from hardware silent data corruptions,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 482–497. [Online]. Available: https://doi.org/10.1145/3676641.3716258 [15] D. Gizopoulos, G. Papadimitriou, O. Chatzopoulos, N. Karystinos, H. D. Dixit, and S. Sankar, “Silent data corruptions in computing systems: Early predictions and large-scale measurements,” in 2024 IEEE European Test Symposium (ETS), 2024, pp. 1–10. [16] H. D. Dixit, L. Boyle, G. Vunnam, S. Pendharkar, M. Beadon, and S. Sankar, “Detecting silent data corruptions in the wild,” 2022. [17] G. Papadimitriou, D. Gizopoulos, H. D. Dixit, and S. Sankar, “Silent data corruptions: The stealthy saboteurs of digital integrity,” in 2023 IEEE 29th International Symposium on On-Line Testing and Robust System Design (IOLTS), 2023, pp. 1–7. [18] A. Singh, S. Chakravarty, G. Papadimitriou, and D. Gizopoulos, “Silent data errors: Sources, detection, and modeling,” in 2023 IEEE 41st VLSI Test Symposium (VTS), 2023. [19] T. Macieira, S. Gurumurthy, S. Gurumurthi, A. Haggag, G. Papadimitriou, and D. Gizopoulos, “Silent data corruptions in computing: Understand and quantify,” in 2024 IEEE 30th International Symposium on On-Line Testing and Robust System Design (IOLTS), 2024, pp. 1–7. [20] D. Gizopoulos, “Sdcs: A b c,” Sep 2024. [Online]. Available: https://www.sigarch.org/sdcs-a-b-c/ [21] ——, “The Dark Side of Computing: Silent Data Corruptions,” Computer, vol. 58, no. 6, pp. 101–106, Jun. 2025. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MC.2025.3554306 [22] A. Samajdar, Y. Zhu, P. Whatmough, M. Mattina, and T. Krishna, “Scale-sim: Systolic cnn accelerator simulator,” arXiv preprint arXiv:1811.02883, 2018. [23] G. Papadimitriou and D. Gizopoulos, “Demystifying the system vulnerability stack: Transient fault effects across the layers,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 902–915. [24] L. Yang, G. Papadimitriou, D. Sartzetakis, A. Jog, E. Smirni, and D. Gizopoulos, “Gpu reliability assessment: Insights across the abstraction layers,” in 2024 IEEE International Conference on Cluster Computing (CLUSTER), 2024, pp. 1–13. [25] A. Chatzidimitriou, P. Bodmann, G. Papadimitriou, D. Gizopoulos, and P. Rech, “Demystifying Soft Error Assessment Strategies on ARM CPUs: Microarchitectural Fault Injection vs. Neutron Beam Experiments,” in 2019 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN), June 2019, pp. 26–38. [Online]. Available: https://doi.org/10.1109/DSN.2019.00018 [26] Z. Chen, N. Narayanan, B. Fang, G. Li, K. Pattabiraman, and N. DeBardeleben, “Tensorfi: A flexible fault injection framework for tensorflow applications,” in 2020 IEEE 31st International Symposium on Software Reliability Engineering (ISSRE). IEEE, 2020, pp. 426–435. [27] A. Mahmoud, N. Aggarwal, A. Nobbe, J. R. S. Vicarte, S. V. Adve, C. W. Fletcher, I. Frosio, and S. K. S. Hari, “Pytorchfi: A runtime perturbation tool for dnns,” in 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W). IEEE, 2020, pp. 25–31. [28] Y. He, P. Balaprakash, and Y. Li, “Fidelity: Efficient resilience analysis framework for deep learning accelerators,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2020, pp. 270–281. [29] A. Tyagi, Y. Gan, S. Liu, B. Yu, P. Whatmough, and Y. Zhu, “Thales: Formulating and estimating architectural vulnerability factors for dnn accelerators,” 2024. [Online]. Available: https://arxiv.org/abs/2212.02649