BBView: A View-Aware Burst-Buffer Mechanism for MPI-IO
Full text
BBView: A View-Aware Burst-Buffer Mechanism for MPI-IO Sohei Koyama, Osamu Tatebe University of Tsukuba September 2, 2025 1 / 18
Table of Contents •Background •Research Goal •Related Work •Positioning of BBView in Related Work •Proposed Method •Evaluation of Proposed Method •Conclusion 2 / 18
Background: HPC Simulation and Output Phase ref: OpenMHD •Consists of computation phase and output phase (alternating between computation and output) •In the output phase, simulation results are written (at each timestep) If the output phase is slower than the computation phase, execution time becomes longer, or the interval between timesteps must be extended. 3 / 18
Background: Issues in the Output Phase •In the computation phase, the multidimensional space is divided into grids and distributed to processes •Example: divided into 2 ×2 for 4 processes (to minimize communication) •Each process needs to write out many small chunks •I/O in small units (tens of KiB) cannot exploit the performance of parallel file systems 4 / 18
Research Goal Without modifying the application or the output file format, shorten or hide the output phase time, thereby reducing the execution time of HPC simulations (or enabling more frequent outputs). Usage of BBView Just add arguments at runtime, no recompilation required, no change in output file mpirun –np 4 –mca io bbview –mca fcoll individual ./a.out 5 / 18
Related Work: Two Phase I/O [Dickens and Thakur, 1998] •Each process calls MPI File write all() simultaneously •Buffers of each process are aggregated in the MPI runtime mpirun –np 4 –mca fcoll dynamic gen2 ./a.out The write unit becomes larger, but still only a few hundred KiB. 6 / 18
Related Work: Node-Local Burst Buffers (e.g. UnifyFS [Brim et al., 2023]) •Construct a file system by aggregating local storage on compute nodes •Temporarily write to the burst buffer quickly, then flush later to the parallel file system •Hide writes to the parallel file system by moving them off the critical path However, write unit size to local storage does not become larger. 7 / 18
Positioning of BBView in Related Work Advantages of related work: •Larger write units to the parallel file system reduce write time •Writing to local storage hides write time to the parallel file system BBView combines both advantages: •Writing to local storage in large units reduces write time •Writing to local storage hides write time to the parallel file system 8 / 18
Proposed Method Overview •Write MPI File Views sequentially (in large units)into a single file •Reconstruct views later and asynchronously write to the parallel file system Process •When a View is set, each process opens /tmp/{filename}-{rank}-{idx} •All writes are appended sequentially into that file (internally set to an identity View) •When MPI File close is called, a background daemon asynchronously reconstructs the view and writes to the parallel file system 9 / 18
Conclusion BBView shortens or hides the output phase time without modifying the application or output file format, thereby reducing execution time of HPC simulations. Released at: https://github.com/tsukuba-hpcs/bbview 16 / 18
Acknowledgments This work was partially supported by •JSPS KAKENHI JP22H00509 •NEDO “Post-5G Infrastructure Enhancement R&D Program” (JPNP20017) •JST CREST JPMJCR24R4 •Interdisciplinary Joint Usage Program at the Center for Computational Sciences, University of Tsukuba •Special joint research with Fujitsu 17 / 18
References Brim, M. J., Moody, A. T., Lim, S.-H., Miller, R., Boehm, S., Stanavige, C., Mohror, K. M., and Oral, S. (2023). UnifyFS: A user-level shared file system for unified access to distributed local storage. In 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pages 290–300. IEEE. Dickens, P. M. and Thakur, R. (1998). A performance study of two-phase i/o. In Euro-Par’98 Parallel Processing: 4th International Euro-Par Conference Southampton, UK, September 1–4, 1998 Proceedings 4, pages 959–965. Springer. Hiraga, K. (2024). 2d reaction-diffusion system benchmark using mpi and mpi-io. 18 / 18