scieee AI-readable full text Open interactive document viewer

Teaching Reproducible Science Through Software Engineering and Generative AI: A Training Paradigm for Emerging Researchers

Charlie, Dey; Powell, Jeaime; Hayden, Linda; Stites, Nole

Abstract

Reproducibility in scientific research hinges on a deep understanding of software engineering, data workflows, and transparent analysis pipelines. In this work, we present a novel training paradigm developed for undergraduate students in computational and data sciences. Student skills were developed using a framework that integrated core Python programming, data analysis with Pandas, web development with Flask, job queuing with Redis, and containerization with Docker.to promote self-aware thinking and independent problem-solving, we also incorporate narrative reasoning and prompt engineering alongside use of generative AI tools.Grounded in real-world data—specifically, traffic incident reports from the City of Austin—our students build and deploy reproducible data workflows and application programming interfaces (APIs). We discuss the pedagogical strategy, curriculum design, tools used, and alignment with goals from the Science Gateways Community. Results from student feedback and learning assessments suggest increased engagement, skill acquisition, and awareness of reproducibility in scientific computation.

Full text

Teaching Reproducible Science Through Software Engineering and Generative AI: A Training Paradigm for Emerging Researchers S. Charlie Dey Texas Advanced Computing Center Austin, TX [email protected] Jeaime H. Powell Omnibond Systems Orlando, FL [email protected] Dr. Linda Hayden Elizabeth City State University Elizabeth City, NC [email protected] Nole Stites Southern Oregon University Ashland, OR [email protected] Abstract: Reproducibility in scientific research hinges on a deep understanding of software engineering, data workflows, and transparent analysis pipelines. In this work, we present a novel training paradigm developed for undergraduate students in computational and data sciences. Student skills were developed using a framework that integrated core Python programming, data analysis with Pandas, web development with Flask, job queuing with Redis, and containerization with Docker.to promote self-aware thinking and independent problem-solving, we also incorporate narrative reasoning and prompt engineering alongside use of generative AI tools.Grounded in real-world data—specifically, traffic incident reports from the City of Austin—our students build and deploy reproducible data workflows and application programming interfaces (APIs). We discuss the pedagogical strategy, curriculum design, tools used, and alignment with goals from the Science Gateways Community. Results from student feedback and learning assessments suggest increased engagement, skill acquisition, and awareness of reproducibility in scientific computation. Keywords—Reproducible Science; Science Gateways; High-Performance Computing Education; Jupyter HPC Integration; Docker Containerization; RESTful APIs; Pandas Data Workflows; Flask Web Services; Redis Job Queuing; Generative AI; Prompt Engineering; Undergraduate Workforce Development; Hackathon; HackHPC Model; projectEUREKA! I. INTRODUCTION Reproducibility is a foundational principle of scientific inquiry and a growing priority in high-performance computing (HPC) and data-driven research. As computational experiments become more complex, involving intricate software stacks, vast datasets, and rapidly evolving AI tools, ensuring that results can be independently verified is both technically and pedagogically challenging. Reproducible research enables transparency, accountability, and—critically—the reusability of scientific artifacts such as datasets, code, workflows, and execution environments. This reusability accelerates discovery by eliminating redundant replication efforts and supports cumulative knowledge—researchers can extend or adapt existing work without re-implementing entire experiments. In computer systems research, papers that share artifacts receive, on average, 75 % more citations than those without, even after controlling for confounding factors—demonstrating their broader scientific impact and integration potential [1]. Moreover, journal and funder policies guided by the FAIR principles (Findable, Accessible, Interoperable, Reusable) underscore that reusability requires not just availability, but well-defined and machine-actionable metadata and environment specifications [2], [3]. This becomes indispensable in high-performance computing (HPC) and data science, where workflows often rely on intricate software stacks, parallel execution, and heterogeneous hardware. Without reusable artifacts—complete with containerized environments, version-controlled code, and precise documentation—reproduction efforts frequently fail due to missing dependencies or unrecorded configurations [4], [5]. Integrating reproducible artifacts into undergraduate education offers profound pedagogical benefits. First, students gain a deeper technical understanding by exploring existing codebases, data schemas, and performance architecture—skills that go beyond textbook exercises. Second, they develop environmental literacy through practical exposure to containerization and dependency resolution, learning to navigate platform-specific challenges using tools like Docker and Singularity [6]. As students troubleshoot mismatches—such as library versions or hardware configurations—they sharpen their debugging, their inquiry and problem-solving skills. Meanwhile, consistent exposure to professional standards in documentation and modular code fosters best practices often absent in traditional coursework. Most importantly, as students reproduce or extend authentic research artifacts, they evolve from passive consumers to active contributors; this shift cultivates a research-oriented mindset, reinforcing the principles of scientific rigor and sharing [7]. Educational programs such as NSF-supported REU sites and recent reproducibility training initiatives have validated XXX-X-XXXX-XXXX-X/XX/$XX.00 ©20XX IEEE this approach. Undergraduates trained to reproduce and extend computations demonstrate stronger data provenance awareness and computational robustness than their peers [8], [9]. Therefore, embedding reusable artifacts within undergraduate curricula serves both as a pedagogical scaffold and a mechanism to instill reproducibility as an operational ethic—preparing students to engage responsibly and effectively in modern computational science. However, training emerging researchers to practice reproducible science remains a non-trivial task. Undergraduate students, particularly in computational fields, often lack exposure to formal software engineering practices and reproducibility-aware workflows [6]. Moreover, the integration of generative AI into research introduces new opportunities—and potential pitfalls—for maintaining reproducibility. This paper introduces a novel undergraduate training paradigm that merges foundational software development skills with modern AI tools, situated within a real-world, project-based environment. The curriculum is designed around the core principles of reproducibility and aligns with community-driven efforts like the Science Gateways Community Institute (SGCI) to make scientific computation more transparent and accessible [12]. Students participating in the institute gain hands-on experience in reproducible workflows using Python, Pandas, Flask, Redis, and Docker, while simultaneously engaging with generative AI tools to enhance their reasoning, debugging, and documentation practices. The primary learning outcome is not just technical fluency, but an operational understanding of how reproducibility functions as both a scientific ethic and an engineering discipline. II. BACKGROUND AND RELATED WORK The call for reproducible research in computational science has been persistent for over a decade. Donoho [10] emphasized that the scientific method depends not just on publishing conclusions, but also on publishing the complete computational environment—code, data, and configuration—that led to those conclusions. Peng [11] and Stodden [13] further expanded this conversation by examining the cultural and technical barriers to reproducibility in data-intensive science. Despite these efforts, surveys indicate that most published computational research lacks sufficient detail for replication [14]. To address this gap, the community has proposed several frameworks and tools. Sandve et al. [6] offered ten practical rules for reproducible computational research, many of which underscore the importance of version control, automated workflows, and comprehensive documentation. Boettiger [5] demonstrated how containerization technologies like Docker can preserve complex software environments, providing a practical path to reproducibility. Similarly, initiatives like Whole Tale [15] and ReproZip [16] have aimed to simplify the sharing of computational experiments by packaging data, code, and execution context. Education-focused responses have included Software Carpentry [17] and Data Carpentry, which teach core practices such as Unix command-line skills, Git usage, and Python programming. However, these are typically short-form workshops and may not offer sustained project-based learning opportunities. Our approach extends this landscape by embedding reproducibility within a multi-week coding institute, where students build web-accessible APIs using real-world datasets and modern development stacks. To ensure all students in our reproducible science program have consistent access to computational resources, we deployed Omnibond’s projectEUREKA!, which provides both browser-based JupyterLab environments and remote Linux shell access hosted on HPC infrastructure [18,19]. Students launch interactive Jupyter notebooks via a centralized web interface that transparently submits batch jobs to scheduled compute nodes, ensuring notebooks run within a reproducible, container-backed software environment on shared hardware. In parallel, secure Linux terminals are made available for command-line workflows and scripting, also scheduled to maintain consistent execution conditions. By combining notebook-based interfaces with remote terminal access, this approach levels the computational playing field—independent of personal devices—while embedding reproducibility into the learning environment through uniform environments, resource provisioning, and data management. In high-performance computing (HPC) and data science education, reproducible workflows are critical for validating results and sharing knowledge. Accordingly, the training program emphasizes a modern toolchain comprising Python, Pandas, Flask, Redis, and Docker, each selected for its role in supporting reproducible practices. Python is the core programming language due to its widespread use in scientific computing and robust ecosystem of libraries that facilitate reproducible analysis [20]. Pandas provides high-level data structures for data manipulation, enabling consistent and repeatable data analysis pipelines [21]. Flask, a lightweight web framework, introduces the development of simple web interfaces for sharing results and deploying reproducible computational tools [22]. Redis serves as an in-memory data store and job queueing system, supporting scalable task scheduling in distributed HPC environments and ensuring deterministic execution of workflows [23]. Finally, Docker containerization is employed to encapsulate software environments, guaranteeing that analyses can be portably reproduced across different platforms and HPC systems [24]. Recently, generative AI tools such as Gemini, Manus, Preplexity, and ChatGPT have emerged as assistants for code generation and reasoning. While these tools can accelerate development, their integration into scientific training requires careful scaffolding to ensure that AI-generated outputs are understood, verified, and properly documented. Prior studies have begun exploring this space [25], but few have investigated its implications for reproducibility in undergraduate education. This work addresses that gap by combining AI-assisted development with reproducibility-focused pedagogy. III. DESIGN AND IMPLEMENTATION Our pedagogical framework was grounded in high-performance computing (HPC) workforce development principles. It was structured around three core pillars: Software Engineering for Science, Narrative Reasoning, and Generative AI & Prompt Engineering. We guided students to write modular Python code and design RESTful APIs, fostering clean and maintainable development practices. They were encouraged to embed narrative through docstrings and commit messages, effectively treating code as technical storytelling. By crafting effective prompts, students learned to use AI tools as collaborative partners—for debugging, design iteration, and refining code. Throughout the four-week program, these pillars shaped both instruction and applied learning, preparing students for real-world reproducible science in HPC. The curriculum unfolded as a cohesive four-week immersive experience, with Weeks 1–3 framed by the SGX3 Coding Institute [26], focusing on foundational HPC and gateway development, followed by Week 4’s HackHPC@ADMI25 Hackathon [27], during which students applied and extended their work toward targeted learning goals. The schedule was as follows: ● Week 1 – Onboarding & Foundations: Students built proficiency in Linux shell operations, Python scripting, Git version control, and used the Bandit wargame to reinforce security and command-line skills. They also practiced documenting and communicating their work through tools like Canva. ● Week 2 – HPC & Data Analytics: Learners advanced to create HTML pages, use Pandas and Matplotlib for data workflows, and incorporated generative AI to support coding and conceptual reasoning, bridging foundational knowledge to higher‑order HPC tasks. ● Week 3 – Gateway Frameworks & Reproducible Pipelines: Students constructed Flask-based microservices, managed asynchronous jobs with Redis, implemented Docker containerization, and developed basic machine learning workflows. They used prompt engineering to structure modular, reproducible pipelines. ● Week 4 – Hackathon Integration via HackHPC Model: Informed by the HackHPC model ([28]), teams of up to four students engaged in a reproducible-science hackathon. Deliverables included code repositories with documentation and comments, a project poster, a final presentation, and a web portal displaying their reproducibility scorecard and team profile. Awards given to the teams targeted three areas of reproducibility in research. The Scorecard Analytics Award prompted teams to justify their reproducibility scores through transparent documentation of code, data, and runtime environments. The Data Portal Award challenged teams to craft a user-friendly web portal—built with Flask and dynamic visualizations—to communicate their scorecard findings clearly. Finally, the Gateways Impact Award evaluated the team's integration of scorecard rigor, portal usability, and final presentation, emphasizing contributions to the Science Gateways vision of accessible HPC tools. IV. STUDENT EXPERIENCE AND OUTCOMES Students reported increased confidence in using development tools and understanding reproducibility. Narrative reasoning helped them internalize why they wrote code a certain way. Prompt engineering allowed them to move beyond "copy/paste" AI usage and instead engage in meaningful conversations with generative tools. Anecdotally, students began personifying their AI assistant as a teammate in the development process. Following the structured training, students participated in a hackathon designed to apply their newly acquired skills in a real-world challenge. The task was to build a reproducible science scoring system—a tool to assess the reproducibility of published scientific research papers. Students were asked to define and implement their own scoring metrics based on criteria such as the presence of GitHub or DockerHub links, whether the code was open source, if the datasets were accessible, and whether the article was paywalled or published in an open access journal. Teams built APIs that parsed metadata, scraped article references, and generated reproducibility scores. Several students integrated Redis queues and containerization to ensure their analysis pipeline was scalable and portable. The hackathon highlighted students' ability to synthesize technical skills with critical thinking about scientific transparency. Post-hackathon survey results highlighted that students valued the learning experience most, citing exposure to new tools such as Python, GitHub, Linux, and generative AI systems. They appreciated the open-ended nature of the challenge, which encouraged diverse problem-solving approaches, and noted the flexible scheduling as a benefit. However, some participants expressed a desire for clearer instructions and earlier definition of project goals, with a few reporting initial confusion about expectations. Additional feedback included suggestions for improving website navigation and offering an in-person format. Overall, students found the hackathon engaging and educational, with one respondent describing it as "perfect" and "very beneficial for learning new skills." V. CONCLUSIONS AND FUTURE WORK Science gateways aim to lower barriers to computational research by providing user-friendly access to advanced cyberinfrastructure. Our training model directly supports this mission by preparing students to build their own micro-gateways—small, functional services that reflect the principles of modularity, reproducibility, and openness. These student projects not only mirror real-world scientific workflows but also provide a scalable path for contributing to larger gateway ecosystems. By blending containerization, RESTful APIs, and job queuing with narrative reasoning and prompt engineering, we created a learning experience that is both technically rigorous and accessible. Students developed critical software engineering skills while engaging with reproducibility as a foundational scientific value. The integration of generative AI and prompt-based development also fostered self-aware problem-solving and collaborative thinking. Future iterations of this training program may include modules on automated testing, continuous integration/continuous deployment (CI/CD) pipelines, and direct integration with gateway platforms such as Apache Airavata or Tapis. We also plan to develop and share open-source assessment rubrics, curricular materials, and reproducibility checklists. By grounding these efforts in community-driven goals, we hope to help scale reproducible science education and further align training pipelines with the needs of the science gateway community. REFERENCES [1] E. Frachtenberg, “Citation analysis of computer systems papers,” PeerJ Comput. Sci., vol. 9, p. e1389, 2023, doi: 10.7717/peerj‑cs.1389. [2] M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Sci. Data, vol. 3, 2016, doi: 10.1038/sdata.2016.18. [3] F. Markowetz, “Five selfish reasons to work reproducibly,” Genome Biol., vol. 16, article 274, Dec. 2015, doi: 10.1186/s13059‑015‑0850‑7. [4] B. A. Antunes et al., “Reproducibility, replicability, and repeatability: A survey of reproducible research with a focus on high performance computing,” Comput. Sci. Rev., vol. 53, p. 100655, 2024, doi: 10.1016/j.cosrev.2024.100655. [5] C. Boettiger, “An introduction to Docker for reproducible research,” ACM SIGOPS Oper. Syst. Rev., vol. 49, no. 1, pp. 71–79, 2015. [6] G. K. Sandve et al., “Ten simple rules for reproducible computational research,” PLoS Comput. Biol., vol. 9, no. 10, p. e1003285, 2013. [7] G. Wilson et al., “Best practices for scientific computing,” PLoS Biol., vol. 12, no. 1, p. e1001745, 2014. [8] L. Vilhuber, H. H. Son, M. Welch, D. N. Wasser, and M. Darisse, “Teaching for large-scale reproducibility verification,” J. Stat. Data Sci. Educ., vol. 30, no. 3, pp. 1–8, Jun. 2022, doi: 10.1080/26939169.2022.2074582. [9] NSF, “Research Experiences for Undergraduates (REU),” National Science Foundation, 2024. [10] D. Donoho, “An invitation to reproducible computational research,” Biostatistics, vol. 11, no. 3, pp. 385–388, 2010. [11] R. D. Peng, “Reproducible research in computational science,” Science, vol. 334, no. 6060, pp. 1226–1227, 2011, doi: 10.1126/science.1213847. [12] N. Wilkins‑Diehr et al., “Science gateways: Leveraging community‑focused cyberinfrastructure,” IEEE Comput., vol. 52, no. 1, pp. 32–41, 2019. [13] V. Stodden, “The scientific method in practice: Reproducibility in the computational sciences,” MIT Sloan Research Paper no. 4773‑10, 2010. [14] M. Baker, “1,500 scientists lift the lid on reproducibility,” Nature, vol. 533, no. 7604, pp. 452–454, 2016, doi: 10.1038/533452a. [15] K. Chard et al., “Enabling reproducible, interactive, and scalable science with Whole Tale,” Concurr. Comput. Pract. Exp., vol. 33, no. 6, p. e5847, 2021. [16] F. Chirigati, R. Rampin, D. Shasha, and J. Freire, “ReproZip: Computational reproducibility with ease,” in Proc. 2016 ACM SIGMOD Int. Conf. Manag. Data, pp. 2085–2088, 2016. [17] G. Wilson, “Software Carpentry: Lessons learned,” F1000Research, vol. 3, p. 62, 2014. [18] B. Wilson, C. Milligan, and D. Landry, “projectEureka: A Gateway for Cloud and K8s HPC & AI,” Zenodo, Gateways 2024 conference, Jul. 2024, doi: 10.5281/zenodo.13864077 [19] Omnibond Systems LLC, “ProjectEureka™: HPC, AI and Interactive Computation with Integrated Storage,” Omnibond.com, 2024. [Online]. Available: https://omnibond.com/project-eureka. [Accessed: Jul. 9, 2025]. [20] Python Software Foundation, Python Language Reference, version 3.10, Python.org, 2021. [Online]. Available: https://docs.python.org/3.10/reference/. [Accessed: Jul. 9, 2025]. [21] W. McKinney, “Data structures for statistical computing in Python,” in Proc. 9th Python in Science Conf. (SciPy 2010), Austin, TX, 2010, pp. 56–61. [22] Pallets, Flask Documentation, Release 2.0, Palletsprojects.com, 2021. [Online]. Available: https://flask.palletsprojects.com/. [Accessed: Jul. 9, 2025]. [23] Redis Labs, Redis Documentation, Version 6.0, Redis.io, 2020. [Online]. Available: https://redis.io/docs/. [Accessed: Jul. 9, 2025]. [24] D. Merkel, “Docker: lightweight Linux containers for consistent development and deployment,” Linux Journal, no. 239, p. 2, Mar. 2014. [25] R. Beale, “Computer science education in the age of generative AI,” arXiv preprint arXiv:2507.02183, 2025. [26] HackHPC, “SGX3 Coding Institute 2025,” HackHPC, 2025. [Online]. Available: https://hackhpc.github.io/sgx3codinginstitute25. [Accessed: Jul. 9, 2025]. [27] HackHPC, “ADMI 2025 Hackathon,” HackHPC, 2025. [Online]. Available: https://hackhpc.github.io/admi25. [Accessed: Jul. 9, 2025]. [28] J. Holmen et al., “The HackHPC Model: Fostering Workforce Development in High‑Performance Computing,” NSF-funded report, 2023. [Online]. Available: https://par.nsf.gov/servlets/purl/10450534.