scieee AI-readable full text Open interactive document viewer

Teaching a semester course in GPU-centric, scalable scientific computing for mechanistic modeling

Pariksheet, Nanda

Abstract

Abstract Graphical Processing Units (GPUs) are now integrated into the worlds fastest supercomputers to solve the most computationally difficult, mechanistic scientific problems in physics, chemistry, biology, material sciences, energy, earth and space sciences, national security, data analytics, and optimization[1-4] an 8 year, $1.8 billion effort called the Exascale Computing Project involving 2,800 scientists and engineers recently finished modernizing the underlying scientific numerical software4 to efficiently use this new generation of machines at more than a billion, billion floating point calculations per second referred to as exaflops. While NSF ACCESS[5,6] and the computing facilities themselves directly grant U.S. institutional and industry researchers access to these GPU-accelerated machines at no cost, research software engineers and scientists must first show that their computations scale to efficiently use multiple GPUs across many computer servers / nodes. Therefore, it is imperative to teach the data parallel software development skills[7-9], calibration techniques[10,11], and sustainable software practices[12] now required to enable large scale, GPU-powered computational research. This poster discusses the unique challenges of semester-long training and course infrastructure to cross-train domain scientists as research software engineers. The course is architected around solving a real-world research problem with a faculty mentor as a capstone project and is tailored to undergraduate students, graduates students, faculty and staff. Undergraduate students and staff members are paired with a faculty mentor ideally before the course and within the first 2 weeks of the course before the drop date; this helps scope the capstone project and group learners in complementary research areas together. This course does not strongly focus on AI tools—there are plenty of courses covering those—but instead focuses on high performance computing (HPC) tools and HPC numerical libraries that can be instrumented for hardware performance analysis. The capstone project is intended to support summer NSF research experiences for undergraduates (REU)[13], collaborations between faculty and core-facility staff, and preliminary grant work of faculty, postdoctoral trainees, and graduate students; this intent is similar to introductory grant writing classes. The course is organized into three phases. The first phase is a crash-course covering efficient within-node, multi-node, and data parallel GPU programming; this provides the initial foundations for learners to begin writing their capstone project code. These topics are covered using C and C++ because C more easily allows inspecting vectorized instructions and C++ is central to vendor-agnostic, GPU and numerical HPC libraries. The second phase is basic knowledge and skills in distributed methods for parameter fitting, performance tuning, automated unit testing and performance testing, and stress-free peer code review along with writing various types of documentation. The third phase covers advanced topics with less immediate applicability to the capstone project, doing deeper dives into topics of interest to the cohort, and for progress reports and final presentations of individual capstone projects. As this course is still being developed, invaluable feedback from the USRSE community about their teaching, learning, and pedagogical experiences would help make it a enjoyable experience for the future learners. The course assumes experience with at least one programming language, therefore, assigned reading material with a short, low-stakes quiz before each class would help level the knowledge and increase confidence of learners for the classroom instruction. Providing classroom laptops or portable PCs are being investigated to deliver instruction because they would allow group activities such as connecting machines to form a cluster and allow students to watch local resource usage, understand latencies, topologies, and provide memorable, active learning experiences[14] complementing terminal work on remote HPCs. It would be of great interest to learn from USRSE instructors about creating engaging opportunities for classroom learning with such portable GPU machines. Lastly, the University of Pittsburgh libraries support developing open educational resources[15-17] for which USRSE attendees can weight in on gradable, pedagogically useful exercises often missing from documentation and training resources to help teach RSE skills[18]. Hopefully this poster discussion and feedback from the USRSE community can help make this course a success and support similar teaching and learning endeavors[19]. References Lax PD. Report of the Panel on Large Scale Computing in Science and Engineering. Tech. rep. 1982 Dec. Available from: https://science.osti.gov/-/media/ascr/pdf/program-documents/archive/Lax_report.pdf Messina P. The Exascale Computing Project. Computing in Science & Engineering 2017 May; 19:63–7. DOI: 10.1109/MCSE.2017.57 Heroux MA. Scalable Delivery of Scalable Libraries and Tools: How ECP Delivered a Software Ecosystem for Exascale and Beyond. Computing in Science & Engineering 2024 Jan; 26:9–18. DOI: 10.1109/MCSE.2024.3384937 Exascale Computing Project (ECP). Available from: https://www.exascaleproject.org/ Boerner TJ, Deems S, Furlani TR, Knuth SL, and Towns J. ACCESS: Advancing Innovation: NSF’s Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support. en. Practice and Experience in Advanced Research Computing. Portland OR USA: ACM, 2023 Jul :173–6. DOI: 10.1145/3569951.3597559 National Science Foundation (NSF) Advanced Cyberinfrastructure Coordination Ecosystem: Services and Support (ACCESS). Available from: https://access-ci.org Carter Edwards H, Trott CR, and Sunderland D. Kokkos: Enabling manycore performance portability through polymorphic memory access patterns. en. Journal of Parallel and Distributed Computing 2014 Dec; 74:3202–16. DOI: 10.1016/j.jpdc.2014.07.003 Beckingsale DA, Scogland TR, Burmark J, Hornung R, Jones H, Killian W, Kunen AJ, Pearce O, Robinson P, and Ryujin BS. RAJA: Portable Performance for Large-Scale Scientific Applications. 2019 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). Denver, CO, USA: IEEE, 2019 Nov :71–81. DOI: 10 . 1109/P3HPC49587.2019.00012. Available from: https://conferences.computer.org/sc19w/2019/#!/toc/14 Reinders J, Ashbaugh B, Brodman J, Kinsner M, Pennycook J, and Tian X. Data Parallel C++: Programming Accelerated Systems Using C++ and SYCL. en. Berkeley, CA: Apress, 2023 Nanda P and Kirschner DE. Calibration methods to fit parameters within complex biological models. Frontiers in Applied Mathematics and Statistics 2023 Oct; 9:1256443. DOI: 10.3389/fams.2023.1256443 Nanda P, Budak M, Michael CT, Krupinsky K, and Kirschner DE. Development and Analysis of Multiscale Models for Tuberculosis: From Molecules to Populations. en. Predicting Pandemics in a Globally Connected World, Volume 2. Ed. by Aguiar M, Bellomo N, and Chaplain M. Series Title: Modeling and Simulation in Science, Engineering and Technology. Cham: Springer Nature Switzerland, 2024 :11–43. DOI: 10.1007/978-3-031-56794-0_2 Katz DS, McInnes LC, Bernholdt DE, Mayes AC, Hong NPC, Duckles J, Gesing S, Heroux MA, Hettrick S, Jimenez RC, Pierce M, Weaver B, and Wilkins-Diehr N. Community Organizations: Changing the Culture in Which Research Software Is Developed and Sustained. Computing in Science Engineering 2019 Mar; 21. Conference Name: Computing in Science Engineering:8–24. DOI: 10.1109/MCSE.2018.2883051 National Science Foundation (NSF) Research Experiences for Undergraduates (REU). Available from: https://www.nsf.gov/funding/initiatives/reu Freeman S, Eddy SL, McDonough M, Smith MK, Okoroafor N, Jordt H, and Wenderoth MP. Active learning increases student performance in science, engineering, and mathematics. en. Proceedings of the National Academy of Sciences 2014 Jun; 111:8410–5. DOI: 10.1073/pnas.1319030111 Open Educational Resources | University of Pittsburgh Library System. Available from: https://library.pitt.edu/OER Scholarly Publishing and Academic Resources Coalition (SPARC). Available from: https://sparcopen.org Open Syllabus Project. Available from: https://opensyllabus.org/ Miller MC. Formal Course Resources for Learning about HPC. en. 2024 Sep. Available from: https://bssw.io/items/formal-course-resources-for-learning-about-hpc Measures of Effective Teaching Project Releases Final Research Report. 2013 Jan. Available from: http://www.gatesfoundation.org/media-center/press-releases/2013/01/measures-of-effective-teaching-project-releases-final-research-report

Full text

Teaching a semester course in GPU-centric, scalable scientific computing for mechanistic modeling Pariksheet Nanda University of Pittsburgh, Department of Chemical Engineering Course goals •Support graduate students developing preliminary GPU-centric scientific applications with sufficient performance and scaling to independently apply for NSF ACCESS allocations. •Teach domain scientists relevant research software engineering skills for industry, academia, and government careers. Course requirements •Familiarity with at least one procedural computer programming language; C++, python, bash shell used by not required (open access resources available for self-guided study). •Basic statistical knowledge of mean, median, standard deviation, L2-norm, etc. Background Graphical Processing Units (GPUs) are integrated into the core architecture of the world’s fastest supercomputers that solve the most computationally difficult, mechanistic scientific problems. An 8 year, $1.8 billion effort called the Exascale Computing Project (ECP) involving 2,800 scientists and engineers recently finished modernizing the underlying scientific numerical software to efficiently use this new generation of machines. Computing facilities grant access at no cost to these GPUaccelerated1–3 machines only to research software engineers and scientists who demonstrate that their computations scale to efficiently use multiple GPUs across many computer servers / nodes. Therefore, it is imperative to teach the data parallel software development skills, appropriate domain science methods, and sustainable software practices to enable large scale, GPU-centric computational research. Modern supercomputers use GPU architectures. Photo of the Aurora supercomputer at Argonne National Laboratory that is ranked 3rd in the Top 500 List of June 2025 and one of three US exascale supercomputers. Germany recently launched their first exascale supercomputer, JUPITER. China may have as many as 10 exascale supercomputers by the end of this year4but stopped participating in the Top 500 List after advanced GPU US embargoes. Course objectives and topics 1. Write vectorized code; analyze efficiency: assembly instructions, memory hierarchy, FLOP-memory bandwidth, compare to algorithmic complexity. 2. Architect multi-node distributed software; analyze bottlenecks using SLURM statistics, Darshan, and HPCToolKit. 3. Fit models to heterogeneous datasets using distributed, parallelizable samplers. 4. Develop unit tests and performance tests. Describe and document code at appropriate levels. 5. Peer constructive code criticism using issue trackers, git merge/pull requests, and review comments. 6. Apply the above skills to a capstone portfolio project. Write an NSF ACCESS report using a scaling study to justify a request for computational resources. course topics capstone write abstract write proposal present, submit code perf. optim. within node multi node GPU accel. tuning software dev. test, CI, benchmark debug, checkpoint peer review, package domain methods mech. multiscale models scientific param. fitting. coupling AI, cloud, data sci. cluster admin. launch, profile E n d S t a r t The course topics are organized into categories of the ■capstone project, ■performance optimization, ■software development, ■domain science methods, and ■cluster administration. Each node corresponds to a single week of semester classroom sessions. Exascale training coverage Topic comparison. This course was inspired by the 2024 Argonne Training Program on Extreme-Scale Computing (ATPESC’24) and the 2025 Lawrence Livermore National Laboratory High Performance Computing Innovation Center (HPC-IC’25) tutorial series. Category Topic This Course ATPESC’24 HPC-IC’25 Performance optimization Within node ✓ ✓ × Multi-node ✓ ✓ ✓ GPU performance portability ✓ ✓ ✓ Performance tuning ✓ ✓ ✓ Software development Unit testing ✓ ✓ × Performance testing ✓ ✓ ✓ Continuous integration ✓×✓ Checkpointing ✓×✓ Benchmarking ✓×✓ Parallel debugging ✓ ✓ × Code peer review ✓× × Documentation ✓× × Packaging ✓ ✓ ✓ Domain science methods Mechanistic multiscale models ✓× × Scientific parameter fitting ✓× × Coupling AI, cloud, and . . . . . . data science ✓×✓ Sysadmin Cluster launch ✓×✓ Cluster profiling ✓ ✓ ✓ Future directions •Extend course to undergraduates by preparing a list of projects with sufficient scope with other faculty members; instead of ACCESS allocations they would apply for REUs. •Upstream class assignments with automated grading to HPC library documentation and NSF/IEEE-TCPP Peachy assignments. •Compartmentalize topics into a 2-day workshop format to be taught at the Pittsburgh Supercomputing Center (PSC). Acknowledgments Mentorship was provided by Elijah MacCarthy from the Oak Ridge National Laboratory as part of the FacultyHack program of SGX3 funded by the National Science Foundation under award number 2231406. Critical feedback on improving the structure of this course was provided by Tze Meng Low from Carnegie Mellon University and John Urbanic from the Pittsburgh Supercomputing Center. Topic training was sponsored by the in-perso 2024 Argonne Training Program on Extreme-Scale Computing (ATPESC) by the Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program. Additional training was provided by the virtual 2025 Lawrence Livermore National Laboratory High Performance Computing Innovation Center (HPC-IC’25) tutorial series. The agent-based, mechanistic, multiscale modeling and associated uncertainty quantification methodologies were supported by NIH grant R01 AI50684. References 1. Carter Edwards H et al. Kokkos: Enabling manycore performance portability through polymorphic memory access patterns. en. Journal of Parallel and Distributed Computing 2014 Dec; 74:3202– 16. doi:10 . 1016 / j . jpdc . 2014.07.003 2. Beckingsale DA et al. RAJA: Portable Performance for LargeScale Scientific Applications. 2019 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). Denver, CO, USA: IEEE, 2019 Nov :71–81. doi: 10 . 1109 / P3HPC49587 . 2019 . 00012. Available from: https:// conferences.computer.org/ sc19w/2019/#!/toc/14 3. Reinders J et al. Data Parallel C++: Programming Accelerated Systems Using C++ and SYCL. en. Berkeley, CA: Apress, 2023 4. Dongarra J et al. Can the United States Maintain Its Leadership in High-Performance Computing? - A report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR Office. Tech. rep. USDOE Office of Science (SC) (United States), 2023 Jun. doi:10 . 2172/1989107 Write abstract Begin a capstone project after completing the performance optimization topics Assignment 0: Write a half page proposal for a new or existing computeintensive mechanistic model to be accelerated using GPU resources. Within-node optimization Strategies for within-node vectorization, caching, memory bandwidth, and I/O Assignment 1: Exercises vectorizing Python code, inspecting C assembler CPU vectorized instructions, creating roofline plots to inspect CPU-memory, hwloc to inspect memory CPU hierarchy. Multi-node scaling Strategies for task distribution and multi-node cluster scaling Assignment 2: Exercises with C, TaskWorks, Spack, MPI, and SLURM. GPU acceleration Introduction to data parallel programming for GPUs Assignment 3: Exercises with C++ Kokkos to fill in the blanks of partially completed programs. Assignment 3capstone: Outline an informal list of goals and objectives to create a minimum viable product of your capstone code through the remainder of the semester. Distributed performance tuning Introduction to distributed performance tuning Assignment 4: Exercises with Darashan, Drishti, and HPCtoolkit Assignment 4capstone: Literature review of software similar to yours and how you intend to make yours different. Testing, CI, benchmarking Automated unit testing, performance regression testing, test coverage, and continuous integration Hands-On 5: Exercises with pytest, Benchpark, and GitLab CI. Assignment 5-capstone: Start adding tests, coverage, and CI to your capstone project. Peer review, usability, packaging Practicing stress-free code peer review, documentation, and packaging Hands-On 7: Exercise with professor and pairing. Assignment 7-capstone: Document, package with Spack or Apptainer. Parallel debugging, checkpointing Introduction to parallel debugging, code correctness, and application checkpoints Hands-On 6: MPIGDB debugging session, resuming from checkpoint to understand error, correctness with sanitizers. Assignment 6-capstone: Start adding checkpointing to your capstone project. Mechanistic, multiscale modeling Simulations using an agent-based, mechanistic, multiscale model Hands-On 8: Analyze how model layers and linking are implemented in DOI: 10.1007/s12195-014-0363-6. Parallel parameter fitting Parallel parameter fitting using gradientfree, pseudo-likelihood based sampling Hands-On 9: Calibration our model from last time with pyabc using adaptive sampling; discuss checkpointing, debugging. Assignment 9-capstone: Start coupling with scientific parameter sampling. Coupling AI, cloud, data science Coupling HPC paradigms with data science, AI, and cloud environments Assignment 10: Summarize chapters 4 and 8 of ISBN 978-3-031-78698-3; in 1 page. Assignment 10-capstone: Assess possible extensions to your capstone project using components of AI, big data, or cloud applications and specialized hardware needs by connecting the concepts from each of the two articles, your domain science, and to what extent we discussed these concepts so far in the class. Cluster sysadmin and profiling Cluster launch, system administration, and profiling Hands-On 11: Setup VMs with network, shared disk, and SLURM to launch and profile an ad-hoc, suboptimal cluster. Final presentations and code Final presentations of capstone projects Hands-On 12-capstone: Walk though your project with ACCESS-relevant details and future directions. Assignment 12-capstone: Submit code for grading. NSF ACCESS proposal NSF ACCESS allocation request Assignment 13-capstone: Write a 3-page ACCESS CI Accelerate proposal with scaling studies and justification for compute resources requested; trim actual Discover proposal to 1-page. Course topics with assignments and hands-on exercises Start End