scieee AI-readable full text Open interactive document viewer

Golden rules for building sustainable bioinformatics capacity using Nextflow and nf-core with a focus on early-mid career researchers: The Kids Research Institute Australia, case report.

Agudelo-Romero, Patricia; Conradie, Talya; Caparros-Martin, Jose A; Martino, David J; Kicic, Anthony; Stick, Stephen M.; Hakkaart, Christopher; Sharma, Abhinav; Theme Collaboration Group

Abstract

The increasing adoption of high-throughput "omics" technologies has heightened the demand for standardized, scalable, and reproducible bioinformatics workflows. Nextflow and nf-core provide a robust framework for researchers, particularly early- and mid-career researchers (EMCRs), to navigate complex data analysis. At The Kids Research Institute Australia, we implemented a structured approach to bioinformatics capacity building using these tools. This perspective presents nine practical “golden” rules that facilitated the successful adoption of Nextflow and nf-core, addressing implementation, knowledge gaps, resource allocation, and community support. Our experience serves as a guide for institutions

Full text

Golden rules for building sustainable bioinformatics capacity using 1 Nextflow and nf-core with a focus on early-mid career researchers: The 2 Kids Research Institute Australia, case report. 3 Patricia Agudelo-Romero1,2,3,*,#, Talya Conradie1, Jose A. Caparros-Martin1,4,5, David J. 4 Martino1, Anthony Kicic1,6,7, Stephen M. Stick7,8, Christopher Hakkaart9, Abhinav Sharma10,#, 5 and the Theme Collaboration Group. 6 7 1Wal-Yan Respiratory Research Centre, The Kids Research Institute Australia, Perth, Western 8 Australia, Australia. 9 2Australian Research Council Centre of Excellence in Plant Energy Biology, School of Molecular 10 Sciences, The University of Western Australia, Perth, Western Australia, Australia. 11 3European Virus Bioinformatics Center, Friedrich-Schiller-Universitat Jena, Thuringia, Germany. 12 4Curtin Health Innovation Research Institute (CHIRI), Curtin University, Perth, Western Australia, 13 Australia. 14 5School of Medicine, The University of Western Australia, Perth, Western Australia, Australia. 15 6School of Population Health, Curtin University, Perth, Western Australia, Australia 16 7Centre for Cell Therapy and Regenerative Medicine, Medical School, The University of Western 17 Australia, Perth, Western Australia, Australia. 18 8Department of Respiratory and Sleep Medicine, Perth Children’s Hospital, Perth, Western Australia, 19 Australia. 20 9Seqera Labs, Barcelona, Catalonia, Spain. 21 10DSI-NRF Centre of Excellence for Biomedical Tuberculosis Research; SAMRC Centre for 22 Tuberculosis Research; Division of Molecular Biology and Human Genetics, Faculty of Medicine and 23 Health Sciences, Stellenbosch University, Cape Town, Western Cape, South Africa. 24 * Correspondence: 25 Corresponding Author: [email protected] 26 # Co-senior authorship 27 Keywords: Bioinformatics, Capacity Building, Nextflow, Early-Mid Career Researcher, Omics, 28 Pipelines 29 Running Title: Building sustainable bioinformatics capacity using Nextflow and nf-core 30 31 32 33 34 35 THIS PREPRINT WAS IMPROVED FURTHER AND THE FINAL FORM IS NOW PUBLISHED IN FRONTIERS IN BIOINFORMATICS 2 Abstract 36 The increasing adoption of high-throughput "omics" technologies has heightened the demand 37 for standardized, scalable, and reproducible bioinformatics workflows. Nextflow and nf-core provide 38 a robust framework for researchers, particularly earlyand mid-career researchers (EMCRs), to 39 navigate complex data analysis. At The Kids Research Institute Australia, we implemented a structured 40 approach to bioinformatics capacity building using these tools. This perspective presents nine practical 41 “golden” rules that facilitated the successful adoption of Nextflow and nf-core, addressing 42 implementation, knowledge gaps, resource allocation, and community support. Our experience serves 43 as a guide for institutions aiming to establish sustainable bioinformatics capabilities and empower 44 EMCRs. 45 46 Introduction 47 The demand for researchers skilled in bioinformatics has surged due to the increasing adoption 48 of high-throughput “omics” technologies, presenting both, opportunities and challenges. Researchers 49 without formal training in programming or workflow engines must now analyze high-dimensional 50 datasets using various tools, while ensuring analysis are reproducible across computational 51 environments. This has highlighted a need for omics data workflows that are standardized, scalable, 52 and reproducible. In this context, we aimed to provide earlyand mid-career researchers (EMCRs) with 53 opportunities to gain knowledge and establish a bioinformatics support system (Woelmer et al., 2021) 54 using workflow managers like Nextflow and the nf-core community. 55 Nextflow is a workflow management system designed to streamline bioinformatics research by 56 composing multiple tools into a single pipeline (Di Tommaso et al., 2017). Its flexibility, portability 57 and scalability ranges from local servers, high-performance clusters, and cloud environments with 58 minimal reconfiguration. Nextflows native task parallelization, makes it ideal for handling high-59 volume of datasets originating from modern "omics" technologies (Dai and Shen, 2022). In addition, 60 the nf-core community enhances the Nextflow ecosystem by connecting users through hackathons, 61 seminars, training, and collaborative platforms like Slack (Ewels et al., 2020; Langer et al., 2024). For 62 newcomers, this support is invaluable for troubleshooting and professional growth. Additionally, nf-63 core community develops and maintains pipelines used across different fields of life sciences. Each 64 pipeline follows strict guidelines to ensure robustness, reproducibility, and ease of use, which are key 65 features for EMCRs entering bioinformatics. 66 THIS PREPRINT WAS IMPROVED FURTHER AND THE FINAL FORM IS NOW PUBLISHED IN FRONTIERS IN BIOINFORMATICS 3 For institutions and researchers in "omics" studies, adopting Nextflow and nf-core is strategic. 67 These tools enable management of high-throughput data, effective collaboration, and reproducible 68 results, ensuring continued scientific advancement. This is especially crucial for EMCRs driving 69 innovation. By integrating these tools into training, institutions can better prepare researchers to meet 70 modern bioinformatics challenges (Stephens et al., 2015). 71 Here, we share our experience at The Kids Research Institute Australia (The Kids), a research 72 institution in Perth (Australia) with a diverse workforce and a strong commitment to building EMCR 73 expertise in computational biology. The Theme Collaboration Group, a multidisciplinary team across 74 The Kids, played a key role in shaping this initiative (Supplemental Data S1). We outline practical 75 “golden” rules essential for adopting and implementing Nextflow (Table 1). These guidelines aim to 76 help other institutions establish sustainable and robust bioinformatics capabilities, enabling researchers 77 to fully leverage Nextflow and nf-core for scientific discovery. Figure 1 summarizes the key rules and 78 the timeline of routine events critical to Nextflow adoption. 79 80 Rule 1: Develop a strategic plan aligned to the institutional priorities 81 A well-designed capacity development program within any institute should (i) reflect the needs 82 of the institution and (ii) adhere to a well-scoped curriculum that benefits the entire organization. 83 Identifying core stakeholders and aligning the program with the institution's current and future data 84 analysis needs is essential for ensuring its long-term relevance and impact (Aron et al., 2021). 85 At The Kids, our vision is "Happy Healthy KIDS," and our objective is to enhance the health, 86 development, and lives of children and adolescents through exemplary research. Our work 87 encompasses a wide array of disciplines, from fundamental to applied science. In light of this diversity, 88 we have concentrated on research streams dependent on omics data processing, providing training for 89 students and EMCRs engaged in these data-intensive fields. We hosted an internal town hall meeting, 90 attended by research teams engaged in bioinformatics, as well as members of the Pawsey 91 Supercomputing Research Centre, an Australian government-supported high-performance computing 92 (HPC) facility. During this meeting, we documented common roadblocks and challenges experienced 93 by bioinformatic users, which laid the foundation for developing a strategic plan and the formation of 94 a special interest group (SIG) with focus on bioinformatics. Building on these insights, we established 95 a three-pillar strategy aimed at improving the efficiency, repeatability, and scalability of our 96 bioinformatics investigations. This strategy focuses on transitioning from ad-hoc shell scripts to the 97 4 integration of Nextflow workflow manager into our computational research (Sztuka et al., 2024) while 98 making the best use of shared resources. 99 1. Pillar one of our strategies involved identifying available computational resources, 100 understanding their allocation across the different research groups, and assessing existing 101 knowledge gaps and bioinformatics needs. 102 2. Pillar two focused on promoting a community of “power users” by integrating support and 103 mentorship. A mentorship program was established to guide EMCRs through their learning 104 journey, while promoting collaboration, a cornerstone of data-driven research. Researchers 105 were encouraged to share their experiences, common challenges, and best practices, 106 ultimately strengthening knowledge exchange, and sparking innovation with a strong, 107 supportive community. 108 3. Pillar three emphasized adherence to established standards and best practices. We 109 promoted the use of bioinformatics standards and provided templates and guidelines for 110 creating reproducible and well-documented Nextflow pipelines, thereby enhancing research 111 quality and credibility. 112 These three pillars informed the design of engagement activities, including seminars and online 113 tutorials (Rule 5), Hacky-Hours (Rule 6), one-on-one sessions and bootcamps (Rule 7). Additionally, 114 continuous evaluation and feedback (Rule 8 and 9) were pivotal in refining and adapting these activities 115 to meet the evolving needs of researchers effectively. 116 Rule 2: Promote awareness of reproducible, robust and open science. 117 Promoting awareness of reproducible data analysis and bioinformatics best practices is crucial 118 in “omics” fields like transcriptomics, genomics, proteomics, metabolomics, and microbiomics. As 119 researchers increasingly adopt “multi-omics” approaches to gain deeper insights into biological 120 phenomenon (Reel et al., 2021; Vahabi and Michailidis, 2022), the importance of reproducibility, 121 robustness and open science cannot be overstated. These practices ensure reliable and actionable 122 research findings while promoting transparency and allowing for validation, which is vital for scientific 123 progress (Kraus, 2014; Stewart et al., 2022). A successful implementation of the capacity-building 124 strategy (Rule 1) should ensure clear communication of the best bioinformatic practices. Participants 125 must understand the importance of these practices in programming and computational experiments, 126 such as testing simulations, algorithms, and models. Additionally, they should recognize the clear 127 5 benefits of applying these practices within the context of their own work and research groups. (Carey 128 and Papin, 2018; Heise et al., 2023). 129 At The Kids, we accomplished this through (i) direct approaches, using seminars and Hacky-130 Hours, as well as (ii) indirectly through mailing lists and a dedicated Microsoft Teams channel. These 131 platforms allowed us to share relevant resources and promote awareness effectively. 132 Rule 3: Identify the needs and skill gaps of different research groups within the institute 133 The key to successful capacity building effort lies in understanding existing skills and 134 identifying future needs of research groups (Woelmer et al., 2021). We began by conducting an initial 135 survey aimed at comprehensively evaluating key requirements of the EMCRs (Supplemental Data S2). 136 The survey gathered information on (i) basic computational skills (e.g. Linux command line 137 proficiency or scripting using Python and/or R languages), (ii) familiarity with bioinformatics tools 138 and workflows; (iii) experience with Nextflow and nf-core; (iv) and specific challenges faced in their 139 research programs. 140 We encouraged participation by emphasizing the importance of the surveys in identifying 141 bioinformatic pipeline needs at The Kids, with weekly follow-ups over four weeks. This approach 142 ensured detailed and honest responses. The results provided an accurate overview of the current 143 bioinformatics capabilities and needs, and the insights shaped our implementation strategy, helping us 144 identify priority topics for training and skill development. 145 Using insights from the survey, we categorized researchers’ skills into three levels (beginner, 146 intermediate, and advanced). This allowed us to design targeted training programs to address the 147 identified gaps by building upon the sensitize, train, hack and collaborate model (Karega et al., 2023). 148 1. Beginner users focused on basic skills and attended introductory courses on Linux, 149 bioinformatics fundamentals (Brandies and Hogg, 2021), and monitoring pipelines using 150 the Seqera Platform (https://seqera.io/platform/). 151 2. Intermediate users focused on workflow development with Nextflow and nf-core pipeline 152 templates (Roach et al., 2022). 153 3. Advanced users specialized in customizing existing nf-core pipelines and integrating new 154 tools for cluster systems, such as those provided within Australia by Australian Pawsey 155 Supercomputing Research Centre (Pawsey Supercomputing Research Centre Perth, 2023d, 156 6 2023b, 2023c, 2023a) and the Australian BioCommons Leadership Share (ABLeS) 157 program (Gustafsson et al., 2023). 158 This targeted approach ensured that researchers received the precise support they needed to advance 159 their bioinformatics skills and progress efficiently within their research projects. 160 Rule 4: Engage with the IT team for infrastructure automation and sustainability through 161 documentation (Nextflow-Biowiki). 162 Building a sustainable bioinformatics environment requires close collaboration with the 163 Information Technology (IT) team. Their expertise in infrastructure management, data storage (short 164 and long-term), and data sharing and protection is invaluable for adhering to responsible big data 165 research practices (Zook et al., 2017). At The Kids, we partnered with the IT team to standardize and 166 streamline frequent computational requests for “omics” data analysis and to establish long-term storage 167 solutions for the Nextflow-Biowiki documentation created during the program (Agudelo-Romero et 168 al., 2023). 169 To optimize infrastructure, the IT team developed a baseline virtual machine (VM) template. 170 This template standardized configurations for efficient installation of the core dependencies, including 171 Java (LTS version) (Arnold et al., 2005), Nextflow (Di Tommaso et al., 2017), and nf-core (Ewels et 172 al., 2020). Package managers like Bioconda (Dale et al., 2018) and Mamba (mamba-Org, n.d.), and 173 containerization tools such as Docker (da Veiga Leprevost et al., 2017) and Singularity (Kurtzer et al., 174 2017), were also preconfigured. Shared access to critical resources (e.g. genome reference files) was 175 enabled, and computing resources, such as number of CPU core, memory and disk space (scratch 176 workspaces), were allocated based on project-specific data and analysis needs. To ensure robustness, 177 we validated the setup by running various nf-core pipelines, addressing any issues, ensuring a seamless 178 performance, and securing enough storage to the outcome’s files on the VMs. 179 At the same time, to ensure long-term sustainability, we created comprehensive documentation 180 that could be used for future reference for EMCR and researchers’ community. We hosted the 181 Nextflow-BioWiki project (Agudelo-Romero et al., 2023) on The Kids’ institutional GitHub repository 182 (https://github.com/TelethonKids). This resource includes detailed guides on bioinformatics tools, 183 workflows execution, troubleshooting in our infrastructure, online manuals, video tutorials, and 184 Frequently Asked Questions (FAQs); catering to diverse learning styles. The documentation was 185 designed to be user-friendly and accessible to all researchers. Regular updates based on users’ feedback 186 THIS PREPRINT WAS IMPROVED FURTHER AND THE FINAL FORM IS NOW PUBLISHED IN FRONTIERS IN BIOINFORMATICS 7 have ensured that resources remain relevant and aligned with technological advancements and 187 community needs. 188 Rule 5: Conduct regular seminars using relevant nf-core pipelines. 189 Regular seminars aimed to introduce participants to conducting computational experiments 190 under optimal conditions (“the happy path”), focusing on data analysis and interpretation. For the 191 implementation of monthly seminars, we followed principles outlined by Fadlelmola et al 2019. 192 Seminars were designed to address real-world data analysis challenges, such as parameter optimization 193 and chaining multiple pipelines for comprehensive analyses. 194 At The Kids, regular seminars covered the breadth of available resources, as key references, 195 without overwhelming participants. These regular seminars also offered a unique opportunity to 196 increase awareness about the resources available from the Institute and the Nextflow and nf-core 197 communities, such as pipeline-specific “byte size” talks. To facilitate focused learning, we organized 198 dedicated seminars for “omics” pipelines under cohesive themes, including: 199 1. Transcriptomics: nf-core/rnaseq (Patel et al., 2024), nf-core/differentialabundance 200 (WackerO et al., 2023), nf-core/smrnaseq (Peltzer et al., 2024b). 201 2. Methylation: nf-core/methylseq (Ewels et al., 2024) , nf-core/atacseq (Patel et al., 2023). 202 3. Microbiome: nf-core/ampliseq (Straub et al., 2024) , nf-core/mag (Yates et al., 2024), 203 agudeloromero/everest_nf (Agudelo-Romero et al., 2025). 204 4. Single-cell: nf-core/scrnaseq (Peltzer et al., 2024a). 205 This thematic grouping allowed participants to dive deeper into specific areas, establishing a 206 better understanding of key biological concepts related to the “omics” strategies, while also 207 encouraging collaborations among students and EMCRs working on similar topics. 208 Ultimately, during our seminars, we encouraged participants to explore and embrace cloud 209 platforms as an integral part of their continued growth. Offering auxiliary opportunities for talent 210 enhancement and performing their data analysis in the cloud. Consequently, certain participants in this 211 program obtained research credits for performing their analysis and deploying Nextflow on cloud 212 infrastructure. Supplemental Table S1 delineates cloud computing companies that extend credits for 213 academic purposes. 214 215 8 Rule 6: Organize practical training sessions (Hacky-Hours) 216 While regular seminars introduce theoretical concepts, Hacky-Hours translate them into 217 practical skills through coordinated sessions tailored to the specific needs of EMCRs (Heise et al., 218 2023). These sessions provide hands-on opportunities to address challenges such as accessing different 219 infrastructures or optimizing computing sources for diverse projects, empowering EMCRs to build 220 their analytical capacity effectively. 221 The involvement of the IT team and researchers in designing Hacky-Hours provided valuable 222 insight in the sessions. The IT team contributed by offering guidance on structured use of resources, 223 ensuring efficient resource allocation. Meanwhile, engaging with researchers helped identify 224 opportunities to refine the program’s content. For instance, when a specific nf-core pipeline was 225 routinely used, an interactive Hacky-Hour session was organized on that pipeline. This session 226 included step-by-step instructions and troubleshooting, delivering practical and targeted training 227 (Boudabous and Tekaia, 2020). 228 At The Kids, bi-weekly Hacky-Hours (30-90 minutes) were structured around topics identified 229 via participant surveys. Initially, these sessions focused on fundamental skills such as downloading 230 and configuring nf-core pipelines. As Nextflow and nf-core adoption grew, surveys revealed a growing 231 need for advanced topics like troubleshooting and handling complex configurations (Supplemental 232 Data S2). This iterative feedback process allowed us to adapt and prioritize future session, ensuring 233 relevance to the evolving needs of the community. Hacky-Hours were scheduled flexibly and only held 234 when there was clear interest, maximizing participation and avoiding unnecessary strain on 235 communication channels. 236 Rule 7: Structured and Personalized Training 237 To accommodate diverse skill levels and learning styles, a successful capacity-building 238 program should combine structured group training with personalized mentoring, such as bootcamps 239 and mentoring session, respectively. Bootcamps are an efficient way to immerse participants in the 240 theoretical and practical aspects, exposing them to the tools and techniques essential for modern 241 bioinformatics. While mentoring sessions is personalized approach empowered individuals to navigate 242 the complexities of bioinformatics. 243 At The Kids, we implemented an integrated approach that included immersive week-long 244 bootcamps alongside tailored one-to-one mentorship sessions. Bootcamps effectively introduced 245 9 participants, especially students and EMCRs, to the theoretical and practical aspects of Nextflow and 246 nf-core. These events balanced conceptual learning with hands-on pipeline development and included 247 dedicated "Bring Your Own Data" (BYO-D) sessions to ensure relevance to individual projects. 248 Including IT team members in these sessions fostered mutual understanding between researchers and 249 technical support, aligning infrastructure solutions with research needs. 250 To complement group training, we offered one-to-one mentoring tailored to participants with 251 limited bioinformatics experience or specific project challenges. Mentors from the organizing team 252 provided individualized support, helping participants gain confidence in using workflow tools. These 253 sessions also generated valuable feedback, which shaped and expanded the BioWiki documentation 254 (https://github.com/TelethonKids/Nextflow-BioWiki), ensuring it remained practical and user-255 informed (McGrath et al., 2019). For broader reach, we recommend starting with foundational seminars 256 (Rule 5) and Hacky-Hours (Rule 6) before launching advanced bootcamps focused on troubleshooting 257 and complex configurations (Hagan et al., 2020). This layered approach promotes continuous learning 258 and maximizes long-term impact. 259 Rule 8: Facilitate continuous learning and engagement 260 To ensure the long-term success of the capacity development program and build a community 261 of “power-users,” it's crucial to establish continual learning and engagement opportunities. Leveraging 262 internal platforms, such as messaging apps or newsletters, is an effective strategy for informing 263 participants about current activities and upcoming events. A monthly newsletter is also an effective 264 medium for sharing updates on Nextflow developments, like module binaries and wave containers, and 265 nf-core community projects, including new pipelines, modules, and configurations. 266 At The Kids, we established a Microsoft Teams channel as a space for discussion and promoted 267 a local volunteer community. The long-term success of the program depends on users willing to assist 268 each other in areas such as creating new software/pipelines (Brack et al., 2022), customizing existing 269 pipelines, addressing infrastructure questions, and designing bioinformatics experiments within the 270 infrastructure. 271 Beyond the internal communication, engagement with the global Nextflow and nf-core 272 communities is also encouraged and should be a major goal of the capacity-building program for long-273 term success. Communities such as online forums offer valuable resources, this is demonstrated in the 274 nf-core Slack group (https://nf-co.re/join) and the organized community events, including online 275 training sessions (https://nf-co.re/events/hackathon). By participating in these community events, 276 16 Table 1. Summary of the motivation for the rules’ development. 454 Motivation Rule Implementation Rule 1 and 2 Knowledge gap Rule 3 Resources Rule 4 Support Rule 5, 6, 7 and 8 Engagement and evaluation Rule 9 455 456 17 Author Contributions 457 458 PAR: Conceptualization, Funding acquisition, Resources, Supervision, Writing (original draft, review 459 & editing). TC: Resources, Writing (review & editing). JAC: Conceptualization, Funding acquisition, 460 Writing (original draft, review & editing). DJM: Funding acquisition, Writing (review & editing). AK: 461 Funding acquisition, Writing (review & editing). SMS: Funding acquisition, Writing (review & 462 editing). CH: Resources, Writing (review & editing). AS: Conceptualization, Resources, Writing 463 (original draft, review & editing). 464 Conflict of Interest 465 The authors declare that the research was conducted in the absence of any commercial or financial 466 relationships that could be construed as a potential conflict of interest. 467 468 Funding 469 This project was funded by The Kids Research Institute Australia's Theme Collaboration Award and a 470 Google Cloud Education Program grant. 471 Acknowledgments 472 We would like to thank the key contributions of The Kids IT team that participated in the development 473 of this project, Steven Figliomeni, Kieran Gee and Chris Hadden. 474