D5.3 - Report on testing provided software
Abstract
Report on the monitoring and testing of the central shared software stack across the range of supported platforms. Dashboard to present current support status of central software stack on current system architectures.
Full text
Report on testing provided software MultiXscale Deliverable 5.3 Deliverable Type: Report Delivered in June, 2025 MultiXscale EuroHPC Centre of Excellence for Multiscale Modelling Acknowledgement Funded by the European Union. This work has received funding from the European High Performance Computing Joint Undertaking (JU) under grant agreement No 101093169. Disclaimer Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European High Performance Computing Joint Undertaking (JU). Neither the European Union nor the granting authority can be held responsible for them.
MultiXscale Deliverable 5.3 Page ii Project and Deliverable Information Project Title MultiXscale: EuroHPC Centre of Excellence for Multiscale Modelling Project Ref. Grant Agreement 101093169 Project Website https://www.multixscale.eu EuroHPC Project Officer Dr. Matteo Mascagni Deliverable ID D5.3 Deliverable Nature Report Dissemination Level Public Contractual Date of Delivery Project Month 30 (30th June, 2025) Actual Date of Delivery 27th June, 2025 Description of Deliverable Report on the monitoring and testing of the central shared software stack across the range of supported platforms. Dashboard to present current support status of central software stack on current system architectures. Document Control Information Document Title: Report on testing provided software ID: D5.3 Version: As of June, 2025 Status: Accepted by Steering Committee Available at: https://www.multixscale.eu/deliverables Document history: Internal Project Management Link Review Review Status: Reviewed Authorship Written by: Satish Kamath (SURF) and Maksim Masterov (SURF) Contributors: Caspar van Leeuwen (SURF), Casper van Leeuwen (SURF), Paul Melis (SURF), Kenneth Hoste (UGent), Lara Peeters (UGent), Alan Ó Cais (UB), Thomas Röblitz (UiB) Reviewed by: Caspar van Leeuwen (SURF), Kenneth Hoste (UGent) Approved by: Alan O’Cais (UB) Document Keywords Keywords: MultiXscale, High Performance Computing (HPC), software, applications, infrastructure 27th June, 2025 Disclaimer: This deliverable has been prepared by the responsible Work Package of the Project in accordance with the Consortium Agreement and the Grant Agreement. It solely reflects the opinion of the parties to such agreements on a collective basis in the context of the Project and to the extent foreseen in such agreements. Copyright notices: This deliverable was co-ordinated by Satish Kamath1(SURF) and Maksim Masterov2(SURF) on behalf of the MultiXscale consortium with contributions from Caspar van Leeuwen (SURF), Casper van Leeuwen (SURF), Paul Melis (SURF), Kenneth Hoste (UGent), Lara Peeters (UGent), Alan Ó Cais (UB), Thomas Röblitz (UiB) . This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this license, visit: http://creativecommons.org/licenses/by/4.0 cb 1[email protected] 2maksim.mastero[email protected]
MultiXscale Deliverable 5.3 Page iii Contents Executive Summary 1 1 Introduction 2 1.1 Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Target audience for this deliverable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.3 Deliverable outline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.4 Partner contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2 Deployments 3 2.1 Running the test suite as part of the EESSI deployment pipeline . . . . . . . . . . . . . . . . . . . . . . . . 3 2.2 Periodic runs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2.2.1 Testing setup .................................................. 4 3 Dashboard 5 3.1 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3.2 Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3.3 Data storage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.4 Security .......................................................... 7 3.5 Dashboards . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3.5.1 Performance as time series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3.5.2 Performance as beeswarm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.5.3 Identity matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 3.6 Connected systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 4 Analysis 11 4.1 Functional verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 4.2 Time series analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.2.1 Performance baseline and variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.2.2 Performance patterns and implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.3 Hardware based comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 5 Conclusion and outlook 17 References 18 List of Figures 1 Dashboard ecosystem. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2 Dashboard network and security restrictions. The green line indicates full (unrestricted) access, the yellow lines indicate restricted access granted to systems based on their IP addresses, the red line indicates restricted access for the generic public. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3 An example view on the dashboard depicting performance time series for the TensorFlow test executed on Snellius “genoa” partition. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 4 Filtering panel with highlighted sections (left) and detailed information about a test point (right). . . . . 8 5 An example of a beeswarm plot. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 6 Identity matrix. An overview of test names over systems. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 7 Identity matrix tooltip with detailed information on a test. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 8 Identity matrix. Detailed view on the module names over systems. . . . . . . . . . . . . . . . . . . . . . . . 10 9 Performance (time steps per second, higher is better) vs time for LJ test for Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) application on Vega CPU partition (Rome 7H12) executed on a scale of 2 nodes. Time spans from 01-01-2024 till 01-05-2025. The blue dashed line indicates the mean performance. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 10 Performance (ns per day, higher is better) vs time for GROMACS application benchmark on Snellius CPU partition (Genoa 9654) executed on a scale of 1 node. Time spans from 01-07-2024 till 09-05-2025. . . . . 12 11 Bandwidth (MB per second, higher is better) vs time for point to point test for OSU Microbenchmarks application on Snellius CPU partition (Genoa 9654) executed within a node. Time spans from 01-012024 till 01-03-2025. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 12 Bandwidth (MB per second, higher is better) vs time for point to point test for OSU Microbenchmarks application on Snellius CPU partition (Genoa 9654) executed across 2 nodes. Time spans from 01-012024 till 01-03-2025. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
MultiXscale Deliverable 5.3 Page iv 13 Performance (seconds per step, lower is better) vs time for the P3M test for ESPResSo application on Vega CPU partition (Rome 7H12) executed on 1 node. Time spans from 01-01-2024 till 04-01-2025. . . . 15 14 Performance (seconds per step, lower is better) vs time for the LJ test for ESPResSo application on Vega CPU partition (Rome 7H12) executed on 1 node. Time spans from 01-01-2024 till 04-01-2025. . . . . . . 15 15 Performance (images per second, higher is better) vs time for TensorFlow application on various CPU partitions on the cloud (AWS), EuroHPC systems (Vega and Karolina), Snellius (Dutch Tier-1 system) executed on 1 node. Time spans from 01-01-2025 till 19-05-2025. . . . . . . . . . . . . . . . . . . . . . . . 16 16 Performance (images per second, higher is better) vs time for TensorFlow application on various CPU partitions on the cloud (AWS), EuroHPC systems (Vega and Karolina), Snellius (Dutch Tier-1 system) executed on 2 nodes. Time spans from 01-01-2025 till 19-05-2025. . . . . . . . . . . . . . . . . . . . . . . . 16 List of Tables 1 List of HPC sites and clouds presented in the dashboard. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
MultiXscale Deliverable 5.3 Page 1 Executive Summary The European Environment for Scientific Software Installations (EESSI) test suite described in Deliverable 1.5 is deployed on various platforms in two phases, namely (i) the pre-deployment phase and (ii) the post-deployment phase. These platforms include the cloud such as Amazon Web Services (AWS), EuroHPC systems such as Vega and Karolina, National Tier 1 systems such as Snellius (The Netherlands), Hortense (Belgium) and Betzy (Norway) and National Tier 2 systems such as Fram (Norway). Pre-deployment testing is performed before the application is ingested into EESSI software stack and post-deployment testing is performed in a periodic manner on various systems listed in this deliverable. The purpose of the pre-deployment test step is not only to test the software itself but also to test that the underlying libraries that are used by the software are also working in a sane manner on a given hardware architecture. A mapping is implemented between the software to be ingested and the tests to be executed. Of course, this testing is designed to consume minimal resources while testing the software, some of its components (if applicable), and another application from the stack to check if the installation affects them. The post-deployment tests are periodically run on various systems listed in this deliverable. The results of these tests are automatically ingested into a database and can be visualized using a dashboard. The technical details of the database ingestion pipeline, the design of the dashboard, its views and other aspects such as storage, security etc. are discussed in detail within this deliverable. The results of the test suite show the indicative performance of a particular software that can be achieved on a given hardware. The dashboard provides a platform to analyze this indicative performance on various hardware as a time series and perform a comparative analysis across different systems. This is not only useful for the end-users, but also for the system maintainers and administrators. To summarize: the developed test suite is deployed across various systems to collect indicative performance data and a dashboard is developed for the analysis and continuous monitoring of the EESSI software stack.
MultiXscale Deliverable 5.3 Page 2 1 Introduction 1.1 Scope This deliverable aims to describe the periodic tests that are running via the developed portable test suite described in Deliverable 1.5, running on several European High Performance Computing Joint Undertaking (EuroHPC) systems as well as in our Continuous Integration (CI) bot. The primary purpose of running periodic tests is to check the performance of the EESSI software stack in a time series and provide the results to the users so that performance expectations for various software can be set on different hardware for the users on different systems. To do this in an efficient manner, a test suite dashboard (also described in this document) is developed where the performance of the EESSI software stack can be checked on various systems and compared. 1.2 Target audience for this deliverable This deliverable is targeted at readers with experience in High Performance Computing (HPC) system support and HPC end-users, particularly in the context of maintaining and testing software stacks (or specific applications) for HPC infrastructure. 1.3 Deliverable outline Section 2describes how the EESSI test suite is deployed as a test step in the software deployment pipeline and as periodic runs where the installed software are tested on various hardware platforms. The data produced by these periodic runs are collected in a database and can be visualized using a dashboard whose development and technical details are described in Section 3. Furthermore, Section 4describes the potential usage of the deployed test suite and the data collected from periodic runs, through the dashboard, by end users, system maintainers, and administrators. 1.4 Partner contributions SURF and UGent contributed as planned to the work in this deliverable. Note that Task 5.3 had other partners (RIJKSUNIGRON, Uib, Ub), who contributed to the development of the pre-deployment pipeline and to the deployment of the test suite (as periodic runs) in various systems both at the national and EuroHPC level.
MultiXscale Deliverable 5.3 Page 3 2 Deployments Due to the separation of system-specific and test-specific information, the EESSI test suite (described in deliverable 1.5) can be deployed on any system, while only requiring that one write a ReFrame configuration file that describes the specifics of that system. There are two ways in which the test suite is used within the MultiXscale project: 1. In the test step of the deployment pipeline to add new software to EESSI. 2. Regular runs on a large variety of HPC systems providing EESSI. 2.1 Running the test suite as part of the EESSI deployment pipeline All EESSI software is built by the EESSI build bots (see [1]). Bot instances are running on a variety of systems: from Magic Castle clusters in AWS and Azure, to HPC clusters from MultiXscale partners (e.g. at SURF and UGent), to EuroHPC clusters (e.g. on Deucalion, and in preparation for JUPITER). The bot configuration describes which partitions a bot instance can submit to. For example, the EESSI build bot running on the AWS Magic Castle cluster is configured with: 1arch_target_map = { 2"linux /x86_64/generic" : "−−partition x86−64−generic−node" , 3"linux /x86_64/ intel /haswell" : "−−partition x86−64−intel −haswell−node" , 4"linux /x86_64/ intel /sapphirerapids" : "−−partition x86−64−intel −srapids−node" , 5"linux /x86_64/ intel /skylake_avx512" : "−−partition x86−64−intel −skylake−node" , 6"linux /x86_64/ intel /cascadelake ": "−−partition x86−64−intel −caslake −node" , 7"linux /x86_64/ intel /icelake ": "−−partition x86−64−intel −icelake −node" , 8" linux /x86_64/amd/zen2 " : "−− partition x86−64−amd−zen2−node" , 9"linux /x86_64/amd/zen3" : "−−partition x86−64−amd−zen3−node" , 10 "linux /aarch64/generic" : "−−partition aarch64−generic−node" , 11 "linux /aarch64/neoverse_n1" : "−−partition aarch64−neoverse−n1−node" , 12 "linux /aarch64/neoverse_v1" : "−−partition aarch64−neoverse−v1−node" } The build pipeline that the bot runs has three stages: a build stage, a test stage, and a deploy stage. To be able to run the test step, a ReFrame configuration file is required that matches the partitions for which the bot is configured. For example, the section for the Intel Skylake nodes on this cluster would look like this in the ReFrame configuration file: 1{ 2’name’: ’x86_64_intel_skylake_avx512 ’ , 3’scheduler ’ : ’ local ’ , 4’launcher ’ : ’mpirun ’ , 5’ access ’ : [’−−nodes=1 ’ , ’−−ntasks−per−node=16 ’ , ’−− partition x86−64−intel −skylake −node’ ] , 6’ environs ’ : [ ’ default ’ ] , 7’ features ’ : [ 8FEATURES.CPU 9] + l i s t (SCALES. keys () ) , 10 ’ resources ’ : [ 11 { 12 ’name’ : ’memory’ , 13 ’options ’ : [’−−mem={ size } ’] , 14 } 15 ] , 16 ’ extras ’ : { 17 # Make sure to round down, otherwise a job might ask for more mem than is available 18 # per node 19 EXTRAS.MEM_PER_NODE: 31342, 20 } , 21 ’max_jobs ’ : 1 22 } , We don’t run the full test suite in the test step: if there are tests in the EESSI test suite for the software that is being added, we run those. In addition, we run a small number of tests to make sure the installation does not inadvertently break other components of the software stack. A mapping is done between which software is being installed, and the corresponding set of tests from the EESSI test suite that should be run for this software. 2.2 Periodic runs Periodic testing of applications within the EESSI software stack is setup on various EuroHPC systems as well as local Tier 1 systems, and also in the cloud. In any cluster, the performance of a certain software depends on performance
MultiXscale Deliverable 5.3 Page 4 of various system level components such as the OS-level stack, the compilers/toolchains, the intermediate math libraries and finally application level software and its dependencies. The primary purpose of the testing is to check the performance of the software provided by the EESSI stack which mainly covers everything above the OS-level stack. Furthermore, since the software stack is being tested on various different systems, observations regarding the systems themselves can be drawn from the data which is collected and displayed in the dashboard. These will be covered in section 4. 2.2.1 Testing setup These tests are set up using the cron jobs on the login nodes of the systems and then the test results are pushed to the dashboard storage via an ingestion script. Based on the resource availability, the frequency and scale of the tests are defined. The daily tests are limited to a maximum scale of 2 nodes and the weekly tests can scale up to 16 nodes which is the current maximum chosen within the test suite. For more information regarding possible scales, please refer to Deliverable 1.5. For each of the periodic runs, a ReFrame configuration file describing the system is also required. Currently, periodic runs are performed on the following clusters: • Vega (IZUM) • Karolina (IT4innovations) • Snellius (SURF) • Doduo, Donphan, Gallade, Hortense, Shinx, Skitty (UGent) • BETZY, FRAM, SAGA (Sigma2) •AWS Magic Castle All of these clusters run the EESSI test suite on the EESSI software environment, and push their data to the dashboard (discussed more extensively in section 3).
MultiXscale Deliverable 5.3 Page 5 3 Dashboard 3.1 Overview The developed online dashboard offers an intuitive and straightforward way to visualize test results generated by the EESSI ReFrame test suite. It presents data either as a series of performance metrics (e.g. time series) or as an identity matrix showing test pass rates. The former can be accessed by visiting https://dashboard.eessi.io, whereas the latter is integrated into the documentation page on https://eessi.io/docs/test suite/dashboard/. These visual representations of the test results help system administrators and EESSI developers to quickly identify and address inconsistencies in the performance of scientific applications integrated into the EESSI shared software stack. 3.2 Structure Figure 1: Dashboard ecosystem. Figure 1illustrates the generic structure of the dashboard ecosystem. Each vertical layer represents a separate hosting site dedicated to a specific task: •HPC site – an HPC system where the ReFrame tests are executed and test reports are generated. • Database Virtual Machine (Database VM) – a cloud virtual machine that hosts the Elasticsearch database where the test data is stored. • Dashboard Virtual Machine (Dashboard VM) – a cloud virtual machine that runs the front-end application with a dashboard. At the HPC site, test reports from the EESSI ReFrame suite are created and saved in JavaScript Object Notation (JSON) format. These JSON files are then parsed by a Python ingestion script, which pushes the parsed data into the Elasticsearch (ES)database on the Database VM. On the Dashboard VM, an Elastic proxy handles queries to the database, retrieving relevant data and making it accessible to the front-end for visualization. On all participating HPC sites, the ingestion script is scheduled to run daily via a cron job. Upon execution, the script scans the designated directory where ReFrame stores its report files, parses each file to extract test metadata, and computes a unique hash value for every test entry. This hash value is generated using the following fields from each test report: • Timestamp of execution • Test name • Hostname • System name • ReFrame hash value (unique per run) • Scheduler assigned job ID (e.g., from SLURM) • Test elapsed time • Command line used to run the test A combination of these fields is hashed using MD5 to create a truly unique identifier for each test that is used by the ingestion script for de-duplication. The script then queries the Elasticsearch (ES)database to check for the existence
MultiXscale Deliverable 5.3 Page 12 4.2 Time series analysis 4.2.1 Performance baseline and variance Establishing a baseline performance for a given test is generally difficult. For (very) basic synthetic tests (e.g. a bandwidth test), the baseline is typically based on the hardware specifications and firmware settings. However, for application tests, it is near-impossible to determine a theoretical baseline performance. A practical result of the periodic testing is that it provides a clear baseline, which then allows detection of any changes compared to that baseline. Figure 9shows the LAMMPS performance over time on Vega’s Rome CPU partition. Apart from the mean, which is the baseline performance achieved (in time steps per second), the variance can provide information on the stability of performance of a system. Some tests inherently show a higher variability than others - this is no reason for concern. However, if a test shows substantially higher variability on one system than on others, this may be a reason for the system administrators to investigate, as it may indicate issues with a subset of the nodes, or an overload on shared resources (e.g. parallel filesystem, network congestion etc.). 4.2.2 Performance patterns and implications Figure 10: Performance (ns per day, higher is better) vs time for GROMACS application benchmark on Snellius CPU partition (Genoa 9654) executed on a scale of 1 node. Time spans from 01-07-2024 till 09-05-2025. A concrete example of where the dashboard clearly indicated a systematic problem on a system (in this case: Snellius) can be seen in Figure 10. The single node and two node GROMACS performance had dropped during the time period August to September 2024 on the Snellius’ genoa CPU partition. This drop in performance was caused by: • A security mitigation that was applied at the firmware level and which included an upgrade of the operating system from RHEL 8 to RHEL 9. The security mitigation was only applicable on Zen4 hardware, and thus only implemented there. • Firmware related issues which required a full system reboot. These problems were identified in the order they are mentioned above. The security mitigation related problem can also be seen in Figure 11 where the performance drop and the recovery after the fix in January 2025 can be clearly identified. It can also be seen from Figure 12 that the drop inperformance was not network related since the internode performance didn’t show any degradation. It is to be noted that some data of the OSU tests is missing here, most likely because the tests were not running. The firmware related problem was fixed in February 2025 and once the system was rebooted, the GROMACS performance returned to normal. A factor that also has an effect on performance is the binding of Message Passing Interface (MPI) ranks and OpenMP threads. In the release of the test suite during July 2024, proper binding was enforced within the test suite for the
MultiXscale Deliverable 5.3 Page 13 Figure 11: Bandwidth (MB per second, higher is better) vs time for point to point test for OSU Microbenchmarks application on Snellius CPU partition (Genoa 9654) executed within a node. Time spans from 01-01-2024 till 01-032025. systems by distinguishing systems where hyper-threading is enabled and where it is not. If hyper-threading is enabled, then the tests can be written such that one can pin a task on each hardware thread or on each physical core based on the option that is chosen in the test. This change was adopted in the Vega system later around August 2024 and a stark improvement in Extensible Simulation Package for Research on Soft Matter Systems (ESPResSo) performance (lower is better) can be seen in Figures 13 and 14. In the ESPResSo test, one task per physical core is chosen due to which the number of tasks per node was halved. The pinning also resulted in a more consistent performance (less variance), which can be clearly seen in the figures. The reduction in variance can also be attributed to less OS jitter and cache misses compared to the situation where all hyper-threads were occupied by the migrating MPI processes. 4.3 Hardware based comparison An important utility of this system is to compare the performance across various systems, including the cloud. It is important to note that the test suite is not fine-tuned to achieve maximum performance on a given system (it is not a benchmark suite), but the tests are designed to show indicative performance on all systems since the same tests are run on each system. The reasons we call it indicative are the following: • The software installed within the EESSI software stack is optimized for that hardware. • The pinning of the MPI tasks can be controlled within individual tests via the functions present in the Mixin class which in turn controls the scheduler options using environment variables. This can be at the thread level (assuming hyper-threading is enabled), physical CPU level or the socket level. This rich overview gives the user an indication as to which hardware their application performs the best. Several tier 0, 1 and 2 systems running on similar hardware should also get an indication if the applications within their systems perform optimally on a given set of hardware. If not, then they can approach the respective system maintainers or administrators to exchange information regarding the settings that each of them applies to achieve this, which promotes further collaboration. As an example, the Tensorflow application from the test suite executed on various CPU partitions is shown in Figures 15 and 16. Here we also compare performance on various ARM architectures available on AWS cluster. Even with Ethernet based interconnect, the performance almost doubles from one to two nodes, which shows that the test itself is not network dependent but is more memory dependent with highest performance reported on Snellius-genoa partition which has a AMD Zen4 based architecture. Another interesting point is that the CPU partitions Karolina-qcpu, Vega and Snellius-rome have the same CPU hardware, namely AMD 7H12 but interestingly the average performance is different, Vega performing 15 percent better some times on a single node but on two nodes the Snellius-rome per-
MultiXscale Deliverable 5.3 Page 14 Figure 12: Bandwidth (MB per second, higher is better) vs time for point to point test for OSU Microbenchmarks application on Snellius CPU partition (Genoa 9654) executed across 2 nodes. Time spans from 01-01-2024 till 01-032025. formance seems to be the best (apart from a few blips due to system related issues). This could be due to differences in clock frequencies set by the system admins and may also be due to differences in network hardware employed by the systems.
MultiXscale Deliverable 5.3 Page 15 Figure 13: Performance (seconds per step, lower is better) vs time for the P3M test for ESPResSo application on Vega CPU partition (Rome 7H12) executed on 1 node. Time spans from 01-01-2024 till 04-01-2025. Figure 14: Performance (seconds per step, lower is better) vs time for the LJ test for ESPResSo application on Vega CPU partition (Rome 7H12) executed on 1 node. Time spans from 01-01-2024 till 04-01-2025.
MultiXscale Deliverable 5.3 Page 16 Figure 15: Performance (images per second, higher is better) vs time for TensorFlow application on various CPU partitions on the cloud (AWS), EuroHPC systems (Vega and Karolina), Snellius (Dutch Tier-1 system) executed on 1 node. Time spans from 01-01-2025 till 19-05-2025. Figure 16: Performance (images per second, higher is better) vs time for TensorFlow application on various CPU partitions on the cloud (AWS), EuroHPC systems (Vega and Karolina), Snellius (Dutch Tier-1 system) executed on 2 nodes. Time spans from 01-01-2025 till 19-05-2025.
MultiXscale Deliverable 5.3 Page 17 5 Conclusion and outlook A functional test setup and pipeline is developed, deploying the test suite developed within Task 1.3 in a periodic manner, across various Tier 0, Tier 1 and Tier 2 systems. The pipeline involves pushing the results of the performed test into a public dashboard. The test results in this dashboard provide value not only to maintainers of the EESSI shared software stack, but also to system administrators and end-users. As illustrated by the examples in this deliverable, the dashboard can be used to monitor problems within the EESSI software stack, but also identify issues in the underlying systems themselves, which can then be resolved by the respective system administrators.
MultiXscale Deliverable 5.3 Page 18 References Acronyms used AWS Amazon Web Services CI Continuous Integration EESSI European Environment for Scientific Software Installations HPC High Performance Computing MPI Message Passing Interface JSON JavaScript Object Notation EuroHPC European High Performance Computing Joint Undertaking VM Virtual Machine SIMD Single Instruction Multiple Data Database VM Database Virtual Machine Dashboard VM Dashboard Virtual Machine ES Elasticsearch SRC SURF Research Cloud API Application Programming Interface Software mentioned ESPResSo Extensible Simulation Package for Research on Soft Matter Systems GROMACS GROningen MAChine for Chemical Simulation LAMMPS Large-scale Atomic/Molecular Massively Parallel Simulator URLs referenced Page ii https://www.multixscale.eu ... https://www.multixscale.eu https://www.multixscale.eu/deliverables ... https://www.multixscale.eu/deliverables Internal Project Management Link ... https://github.com/multixscale/planning/issues/40 [email protected] ... mailto:[email protected] [email protected] ... mailto:[email protected] http://creativecommons.org/licenses/by/4.0 ... http://creativecommons.org/licenses/by/4.0 Page 5 https://dashboard.eessi.io ... https://dashboard.eessi.io https://eessi.io/docs/test suite/dashboard/ ... https://eessi.io/docs/testsuite/dashboard ingestion script ... https://github.com/EESSI/dashboard-ingestion Page 6 Vue ... https://vuejs.org Nginx ... https://nginx.org/en/docs/beginners_guide.html JavaScript ... https://developer.mozilla.org/en-US/docs/Web/JavaScript TypeScript ... https://www.typescriptlang.org D3 ... https://d3js.org Elastic ... https://www.elastic.co/licensing/elastic-license distribution ... https://github.com/EESSI/dashboard-ingestion/blob/main/docs/setup_es_on_cloud. md Page 7 Nginx ... https://nginx.org/en/docs/beginners_guide.html Page 11 GROMACS issue ... https://gitlab.com/gromacs/gromacs/-/issues/5057 Citations [1] T. Röblitz and K. Hoste, “D5.1 - community contribution policy and github app,” Jan. 2024. [Online]. Available: https://doi.org/10.5281/zenodo.10451793