scieee AI-readable full text Open interactive document viewer

Mixed-Precision For Energy Efficient Computations

Gedik, Gülçin; Schöne, Robert; Iakymchuk, Roman

Abstract

As simulations become more realistic, the pursuit of higher accuracy results in extended computation times and substantial energy consumption. This study explores mixed-precision computing as a promising strategy to address these challenges, leveraging computer arithmetic tools to optimize performance. To do so, we used Reactor Simulator and LULESH benchmarks as case studies to evaluate the potential of mixed-precision strategies to reduce both time-to-solution and energy-to-solution. For Reactor Simulator, we achieved more then 30 % reduction in both metrics without compromising accuracy. Similarly, results for LULESH demonstrated improvements of up to 31.5 % in time-to-solution and 25.6 % savings in energy-to-solution.

Full text

Mixed-Precision For Energy Efficient Computations G¨ ulc¸in Gedik∗†‡ †Universit´ e Paris-Saclay, UVSQ Dresden, Germany [email protected] Robert Sch¨ one‡ ‡ZIH, CIDS, TU Dresden Dresden, Germany [email protected] Roman Iakymchuk∗ ∗Ume˚ a University Ume˚ a, Sweden [email protected] Abstract—As simulations become more realistic, the pursuit of higher accuracy results in extended computation times and substantial energy consumption. This study explores mixed-precision computing as a promising strategy to address these challenges, leveraging computer arithmetic tools to optimize performance. To do so, we used Reactor Simulator and LULESH benchmarks as case studies to evaluate the potential of mixed-precision strategies to reduce both time-to-solution and energy-to-solution. For Reactor Simulator, we achieved more then 30 % reduction in both metrics without compromising accuracy. Similarly, results for LULESH demonstrated improvements of up to 31.5 % in time-to-solution and 25.6 % savings in energy-to-solution. Index Terms—Mixed-precision, Time-to-solution, Energy-tosolution, LULESH, Verificarlo. I. INTRODUCTION AND BACKGROUND Recent efforts have focused on designing applications that not only deliver high accuracy, but also aim to minimize two metrics: time-to-solution and energy-to-solution [1]–[3]. Although mixed-precision computing offers potential gains in both, achieving those gains without compromising scientific accuracy is non-trivial [4], [5] as it represents a multi-objective optimization challenge. Our work addresses this challenge by exploring the optimization space defined by hardware capabilities, algorithmic choices, and the selective application of mixed-precision techniques. In parallel with theoretical advances in floating point arithmetic error mitigation and estimation techniques [6], [7], practical tools for exploring precision and error propagation have emerged. One such tool is Verificarlo, which operates at the intermediate representation level of the compiler to instrument and analyze floating point operations without modifying the source code. Verificarlo leverages two complementary backends: the Monte Carlo Arithmetic (MCA) backend [8] to assess error sensitivity in code regions by evaluating numerical stability under randomized perturbations, and the Variable Precision (VPREC) backend [9] to explore mixed-precision strategies, quantify rounding-error effects, and determine the minimal precision to preserve accuracy or convergence. Exascale computing demands an understanding of how hardware limits, algorithm design, and precision choices interact to impact performance. When these factors are balanced well, applications can be both energy-efficient and reliable. To demonstrate this, we present our methodology in the following sections and apply it to two case studies, both of which employ explicit solvers and exhibit distinct computational characteristics. II. METHODOLOGY Changing an application from double precision to mixed (including lower) precision requires balancing accuracy and performance. Hence, mechanisms that apply these changes have to be adaptable since each application has unique features. A key challenge is to identify routines that can use reduced precision safely – without compromising accuracy or degrading performance [10]. This involves weighing the performance gains of lower precision against the overhead of copying and casting between variable types. Achieving this balance demands a deep understanding of the application’s architecture, including data structures, computations, libraries, and communication patterns. To guide this, we follow the methodology proposed in [11]. We start with profiling performance hotspots to identify regions with high potential speedup. Then, we analyze numerical hotspots by tracking variables and quantifying the error growth with the help of Verificarlo’s VPREC and MCA backends. We then implement mixed precision for the most time-consuming regions that contribute minimal error, which maximizes speedup with minimal accuracy loss. This functionlevel analysis illustrates the workflow and performance-energy trade-offs for sequential versions, but also directly extends to highly parallel scientific codes, where lower-precision storage shrinks message sizes and reduces communication overhead in addition to higher computational throughput. While finergrained tuning exists [12], [13], function-level granularity offers a practical balance by significantly reducing the search space. Throughout, we eliminate regions unlikely to benefit from reduced precision. We adopt a staged strategy: converting the most time-critical inner routines to single precision first, then extending to outer bottlenecks where each stage includes the previous one. Finally, we implement the mixed-precision code and monitor accuracy, time-to-solution, and energy-to-solution on the LUMI system using SLURM’s energy accounting plugin and HPE Cray PM Counters [14], [15], guided by our CEEC Best Practice Guide [16]. III. REACTOR SIMULATOR BENCHMARK The Reactor Simulator [17], [18] implements a probabilistic Monte Carlo application that models interactions between neutron-source particles and a slab. It emits particles iteratively, updates their states, and accumulates total-energy in double precision while tracking outcomes with integer counters, making it a robust benchmark for mixed-precision Type Stage 1 Stage 2 Stage 3 Stage 4 Stage 5 Double Energy Median (J) 1060 1220 1200 1250 1230 1700 Energy Savings (%) 37.6% 28.2% 29.4% 26.4% 27.6% - Time Median (s) 3.89 4.15 4.12 4.22 4.12 5.63 Time Savings (%) 30.7% 26.2% 26.6% 25% 26.7% - Error 0 0 0 10−810−7TABLE I: Energy (in Joules) and time-to-solution (in seconds) with Reactor Simulator with 10 million elements. Type Stage 1 Stage 2 Stage 3 Double Energy Median (J) 1060 803 965 1080 Energy Savings (%) 1.8% 25.6% 10.6% - Time Median (s) 3.53 2.71 3.20 3.96 Time Savings (%) 10.8% 31.5% 19.0% - Error 10−910−910−9TABLE II: Energy (in Joules) and time-to-solution (in seconds) with LULESH with 203elements. techniques. Applying a higher precision to the entire application reduces the forward error in total-energy nearly linearly, as the mantissa varies from 3to 52, regardless of the number of particles. This linear relationship is not trivial, as highlighted in [11]. Profiling reveals that shared math library calls (e.g., sin_ fma(),exp_fma()) dominate execution. To evaluate precision impacts, we instrumented these functions using the VPREC backend. We observed that the error plateaus after 17 bit, indicating that FP32 accuracy is sufficient as shown in Fig. 1. We proceed to instrument each function individually at single precision with the VPREC backend, while leaving the others in double precision. By following this approach we explore the extent to which we can reduce precision. We notice that only two functions out of ten have the potential to contribute to forward error under this approach. These insights guided our implementation of a five-stage mixed-precision workflow. In Stage 5, most functions run in single precision, with only accuracy-critical routines in double precision. Table I summarizes time and energy savings on the LUMI system. IV. LULESH LULESH (Livermore Unstructured Lagrangian Explicit Shock Hydrodynamics) is a proxy application developed by Lawrence Livermore National Lab [19]–[21] that replicates the computational behavior of a hydrodynamics code [22]. The application represents real-world simulations with its computational and memory access patterns. With a mesh of 203elements, our experiments reveal that error decreases nearly linearly with increased precision. Notably, at 23 bits of mantissa (corresponding to FP32), we observe a stagnation at the timestep selection with Verificarlo, indicating a precision limit. We also observe that 10 out of 40 functions do not contribute to the final error of the energy variable, when emulated in single precision separately while the rest of the application is kept in double precision. Profiling revealed that a small number of routines accounted for a significant share of the runtime. Functions requiring high numerical accuracy, such as TimeIncrement() and Calc PositionForNodes(), were retained in double precision. In contrast, performance-critical computations were selected for reduced precision to improve efficiency. We applied mixedprecision optimization in three stages: In the first stage, only the core routine CalcElemFBHourglassForce() was converted to single precision; however, per-element casts and copies imposed on its caller introduced overhead that yielded negligible performance gains as seen in Table II. In 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 ... 52 Size of Mantissa [bit] 0.0000 0.0002 0.0004 0.0006 0.0008 0.0010 0.0012 Mean Relative Error [%] cross/sin cross/exp Fig. 1: Mean absolute relative forward error for varying mantissa bits [3,52] for sin() and exp() calls in the cross() function of the Reactor Simulator case study. the second stage, we extended this conversion to CalcFB HourglassForceForElems() by modifying its signature to accept single-precision inputs and relocating all cast/copy operations into its caller, CalcHourglassControlFor Elems(). This eliminated indirect-copy overhead and produced the largest performance gain. In the final stage, we implemented higher-level routines CalcHourglass ControlForElems(),CollectDomainNodestoElem Nodes(), which showed zero error when instrumented alone in single precision with the help of VPREC backend, and VoluDer(), which significantly impacts runtime, its contribution to error is comparatively limited, thereby completing the full optimization. V. CONCLUSION AND FUTURE WORK In this work, we applied a well-established methodology to two explicit solver use cases and showed that mixedprecision tuning is inherently application-specific. Our results demonstrate that selectively reducing precision in key computational kernels can improve performance and energy up to 31.5 % and 37.6 %, respectively, while acceptable numerical accuracy for both the Reactor Simulator and LULESH benchmarks is preserved. These results underscore the potential of mixed-precision optimizations as an effective approach to optimize scientific simulations for both performance and energy efficiency. Although challenges remain in narrowing the mixed-precision search space and automating its implementation, future work will address these and evaluate our approach across a wider range of scientific applications. REFERENCES [1] M. Malms, L. Cargemel, E. Suarez, N. Mittenzwey, M. Duranton, S. Sezer, C. Prunty, P. Ross´ e-Laurent, M. P´ erez-Harnandez, M. Marazakis, G. Lonsdale, P. Carpenter, G. Antoniu, S. Narasimharmurthy, A. Brinkman, D. Pleiter, U.-U. Haus, J. Krueger, H.-C. Hoppe, E. Laure, A. Wierse, V. Bartsch, K. Michielsen, C. Allouche, T. Becker, and R. Haas, “ETP4HPC’s SRA 5 - Strategic Research Agenda for HighPerformance Computing in Europe - 2022,” Nov. 2022. [2] M. Govett, B. Bah, P. Bauer, D. Berod, V. Bouchet, S. Corti, C. Davis, Y. Duan, T. Graham, Y. Honda, A. Hines, M. Jean, J. Ishida, B. Lawrence, J. Li, J. Luterbacher, C. Muroi, K. Rowe, M. Schultz, M. Visbeck, and K. Williams, “Exascale computing and data handling: Challenges and opportunities for weather and climate prediction,” Bulletin of the American Meteorological Society, vol. 105, no. 12, pp. E2385 – E2404, 2024. [3] DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, and et. al., “Deepseek-v3 technical report,” 2025. [4] Y. Chen, P. de Oliveira Castro, P. Bientinesi, N. Jansson, and R. Iakymchuk, “Enabling mixed-precision in spectral element codes,” Future Generation Computer Systems, vol. 174, p. 107990, 2026. [5] A. Kashi, H. Lu, W. Brewer, D. Rogers, M. Matheson, M. Shankar, and F. Wang, “Mixed-precision numerics in scientific applications: survey and perspectives,” 2025. [6] N. J. Higham and T. Mary, “A new approach to probabilistic rounding error analysis,” SIAM Journal on Scientific Computing, vol. 41, no. 5, pp. A2815–A2835, 2019. [7] E.-M. El Arar, D. Sohier, P. de Oliveira Castro, and E. Petit, “Stochastic rounding variance and probabilistic bounds: A new approach,” SIAM Journal on Scientific Computing, vol. 45, no. 5, pp. C255–C275, 2023. [8] C. Denis, P. De Oliveira Castro, and E. Petit, “Verificarlo: Checking floating point accuracy through monte carlo arithmetic,” in 2016 IEEE 23nd Symposium on Computer Arithmetic (ARITH), pp. 55–62, 2016. [9] Y. Chatelain, E. Petit, P. de Oliveira Castro, G. Lartigue, and D. Defour, “Automatic exploration of reduced floating-point representations in iterative methods,” in Euro-Par 2019: Parallel Processing (R. Yahyapour, ed.), (Cham), pp. 481–494, Springer International Publishing, 2019. [10] P. de Oliveira Castro, High Performance Computing Code Optimizations: Tuning Performance and Accuracy. PhD thesis, Universit´ e ParisSaclay, 2022. [11] Y. Chen, P. d. O. Castro, P. Bientinesi, and R. Iakymchuk, “Enabling mixed-precision with the help of tools: A nekbone case study,” in Parallel Processing and Applied Mathematics (R. Wyrzykowski, J. Dongarra, E. Deelman, and K. Karczewski, eds.), (Cham), pp. 34–50, Springer Nature Switzerland, 2025. [12] C. Rubio-Gonz´ alez, C. Nguyen, H. D. Nguyen, J. Demmel, W. Kahan, K. Sen, D. H. Bailey, C. Iancu, and D. Hough, “Precimonious: Tuning assistant for floating-point precision,” in SC ’13: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pp. 1–12, 2013. [13] M. O. Lam, T. Vanderbruggen, H. Menon, and M. Schordan, “Tool integration for source-level mixed precision,” in 2019 IEEE/ACM 3rd International Workshop on Software Correctness for HPC Applications (Correctness), pp. 27–35, 2019. [14] A. Hart, H. Richardson, J. Doleschal, T. Ilsche, M. Bielert, and M. Kappel, “User-level power monitoring and application performance on cray xc30 supercomputers,” Cray User Group, 2014. [15] S. J. Martin and M. Kappel, “Cray xc30 power monitoring and management,” Cray User Group, 2014. [16] R. Iakymchuk, G. Gedik, K. Kulkarni, Y. Chen, D. Kempf, S. Kemmler, D. Papageorgiou, D. Konioris, S. Kiebdaj, J. Corbalan, and H. K¨ ostler, “Best Practice Guide – Harvesting energy consumption on European HPC systems: Sharing Experience from the CEEC project,” Aug. 2024. [17] D. Kahaner, C. Moler, S. Nash, and J. Burkardt, “Reactor simulator,” 1989. [18] D. Kahaner, C. Moler, and S. Nash, Numerical Methods and Software. Englewood Cliffs, NJ: Prentice Hall, 1989. LC: TA345.K34. [19] I. Karlin, A. Bhatele, B. L. Chamberlain, J. Cohen, Z. Devito, M. Gokhale, R. Haque, R. Hornung, J. Keasler, D. Laney, E. Luke, S. Lloyd, J. McGraw, R. Neely, D. Richards, M. Schulz, C. H. Still, F. Wang, and D. Wong, “Lulesh programming model and performance ports overview,” Tech. Rep. LLNL-TR-608824, Livermore CA, December 2012. [20] I. Karlin, J. Keasler, and R. Neely, “Lulesh 2.0 updates and changes,” Tech. Rep. LLNL-TR-641973, Livermore, CA, August 2013. [21] M. B. G. R. D. Hornung, J. A. Keasler, “Hydrodynamics challenge problem,” 2011. [22] L. L. N. L. (LLNL), “Lulesh – livermore unstructured lagrangian explicit shock hydrodynamics.” https://github.com/LLNL/LULESH/tree/ 46c2a1d6db9171f9637d79f407212e0f176e8194, n.d. Accessed: 202507-31. ACKNOWLEDGMENTS The author gratefully acknowledges the financial support provided by the Center of Excellence in Exascale Computational Fluid Dynamics (CEEC) under Grant No. 101093393, funded by the European Union through the EuroHPC Joint Undertaking and Sweden, Germany, Spain, Greece and Denmark. The author would like to thank the NHR-Verein e.V.(www.nhrverein.de) for supporting this work/project within NHR Graduate School of National High Performance Computing (NHR). The author would like to thank Pablo de Oliveira Castro for their valuable guidance and contributions. APPENDIX A MEASUREMENT METHODOLOGY DETAILS All experiments were performed on the CPU partition of the LUMI system, where we compiled each benchmark with g++ version 12.2.0 and the -O2 optimization flag. To obtain statistically significant measurements, every application was run 32 times under the performance-scaling governor, and both runtime (in seconds) and energy consumption (in joules) were recorded. Energy data were gathered via SLURM’s energy-accounting plugin, which reads HPE Cray PM Counters through the Baseboard Management Controller (BMC) as detailed in our CEEC Best Practice Guide; nodes were dedicated exclusively to each job to eliminate interference from co-scheduled workloads. CPU time was measured with the GNU time command, capturing only user and system time to reflect the precise CPU resources devoted to each process. APPENDIX B ERROR ANALYSIS DETAILS To verify the numerical accuracy of our results, we recompiled the code using Verificarlo (version 1.0.0) with the -O2 optimization flag to avoid unintended alterations from compiler optimizations. During standard runs, Verificarlo’s VPREC backend was employed with its default settings, preserving IEEE-754 binary64 precision and range. For mixed-precision experiments, we invoked VPREC with the -precision-binary64=<desired_mantissa_bits> option to emulate double precision at the specified mantissa width. This approach ensures that only the targeted precision changes, rather than other compiler transformations, impact our accuracy measurements. We quantify the error introduced by reduced precision using the mean absolute relative error.