scieee AI-readable full text Open interactive document viewer

FAIR Software and FAIR Workflow

COZZINI, STEFANO

Full text

Software for FAIR data Stefano Cozzini , PRP@CERIC coordinator Area Science Park 10.11.2025 Agenda •Data Workflow examples: •The basic: Data from SEM to repo •The luxury: Data from Oxford Nanophore to ORFEO •FAIR approach to scientific software •What/Why software in research •Software sustainability •FAIR4RS: Fair principles for research software •Recommendation to make software FAIR •How to make a modern data workflow FAIR ? •How to make a modern computational workflow FAIR ? Software in Research: A pillar of Open Science •Different roles: •a tool •a research outcome or result •the object of research Where is software in your daily research ? Where is software in your daily research ? •software is a component in our scientific instruments Where is software in your daily research ? •software is the instrument ! Where is software in your daily research ? •software analyses your research data Where is software in your daily research ? •software analyses your research data Where is software in your daily research ? •software presents your research results Updated version: taken on 20.10.2024 Updated version: taken on 11.11.2025 Lots of wheel reinvention? it’s fun and was funded! Sustainable?? We do need more than one kind of wheel … … technical issues, types of data, analysis, users, communities …. Analysis Code one-off me research Prototype Tools Research Software Infrastructure professionalised product Concept: Thanks to Tom Honeyman, ARDC Slide taken from Research Software Sustainability takes a Village (zenodo.org) to have a sustainable future you must create and sustain a future workforce that can develop, support and/or want to use your software. Image: It Takes a Village: Open Source Software Sustainability A Guidebook for Programs Serving Cultural and Scientific Heritage February 2018 More on https://future-of-research-software.org/ Software patchwork: Act Local, Think Global A web of dependencies, a spectrum of visibility. https://xkcd.com/2347/ User facing shiny thing Applications, tools, scripts Domain specific reusability Visible - Rocket Underware Platforms, infrastructure, libraries Big codes and little codes Cross-domain generic reusability Overly familiar Invisible – Rocket launcher Community <-> software closeness Direct, indirect, unconnected to the software use or rely on in supply chain Biologist Bioinformatician Specialised software developer Platform developer Library maintainer Infrastructure provider Informs the need for funding Six different software layer for reproducibility 1. Project specific software Written by scientists for a specific research project 2. Domain specific research software. These are tools and libraries thati mplement models and methods which are developed and used by communities ranging in size from a single research lab to thousands of researchers” 3. Scientific infrastructure created specifically for scientific computing, but not any particular domain.” 4. Infrastructure software not specific toscientific computing....obtain[ed]from the wider non-scientific software market” 5. operating system 6. hardware Research software in a snapshot Software developed and used for the purpose of research: to generate, process, analyze results within the scholarly process Fundamental in the research process •But •Software will collapse if not maintained •Software bugs are found, new features are needed, new platforms arise •Software development and maintenance is human intensive •Much software developed specifically for research, by researchers •Researchers know their disciplines, but often not software best practices •Researchers are not rewarded for software development and maintenance in academia • Developers don’t match the diversity of user community FAIR4RS: •released in 2022 by the FAIR for Research Software (FAIR4RS) Working Group (WG), •jointly convened by •ReSA · Research Software Alliance •FORCE11 – The Future of Research Communications and e-Scholarship •Research data alliance: rd-alliance.org •This milestone reflects the maturation of the research community in understanding the benefits of having FAIR research software, and coming together as the FAIR4RS WG to achieve this. • The FAIR4RS WG is a global and interdisciplinary community whose members share an interest in the application of FAIR principles to research software, such as researchers, software users, developers and maintainers, policy makers, infrastructure support staff, and funders. •The FAIR4RS Principles are relevant to increase transparency, reproducibility, and reusability of research. More comments •The Interoperable and Reusable FAIR4RS Principles are somewhat different than the equivalent FAIR data principles •This because of differences between software and data where the FAIR4RS group had to choose how to define these terms in the context of software, reaching the definitions shown above in the principles. •They define interoperability as how information (data, metadata, application programming interfaces (APIs)) is exchanged. •They then define reusability as both usability (the software can be executed) and reusability (it can be understood, modified, built upon, or incorporated into other software). Main reference: •Data Workflow examples: •The basic: Data from SEM to repo •The luxury: Data from Oxford Nanophore to ORFEO Ten quick tips for building FAIR workflows | PLOS Computational Biology Tip1: register the workflow •Registering the workflow to any public record, preferably one that is also indexed by popular search engines, will increase findability. •we recommend registries that enable systematic scientific annotations and are catering for workflows written in different languages. •Examples of these registries are •WorkflowHub •Docksto re. Tip 2: Describe the workflow with rich metadata •The metadata should cover information on all data entities that are present in the workflow •workflow language files, •scripts, configuration files, •example input data, etc.. •research data can be packaged along with the associated metadata using the ROCrate (Research Object Crate) specification. Research Object Crate (RO-Crate) - Research Object Crate (RO-Crate) (stain.github.io) • A workflow RO-Crate should follow the community curated Bioschemas [30] specification for a computational workflow, which defines the workflow properties that are mandatory or recommended to be described [31]. •The metadata is captured in a JSON-LD file, using the Linked Data principles [32]. •Following these principles, the metadata file describes all data and contextual entities (researchers, organizations, etc.) of the workflow with uniform resource identifiers (URIs). This ensures that all entities in the RO-Crate are described unambiguously and can be easily searched for. Moreover, workflow RO-Crate objects can be directly uploaded to WorkflowHub to register the workflow. Altogether, the RO-Crate method offers a good trade-off between usability (human readable formats) and richness (sufficient metadata). Tip 3: Make source code available in a public code repository •Multiple conventional repository services for software development are available such as GitHub, GitLab, and Bitbucket •Source code should be written following widely used style conventions, e.g., PEP 8 for Python and the Google Style Guide for a variety of programming languages •Code analysis tools that can assist workflow developers in adhering to these style conventions are available •These tools should be integrated in the workflow development routine, for example, through automatic testing protocols. From: https://google.github.io/styleguide/shellguide.html Tip 4: Provide example input data and results along with the workflow •Accessibility of the workflow’s input data and associated results will help the end-user to understand how the workflow should function and improves reproducibility. Example data can be provided along with the workflow, for example, when using RO-Crate to package the workflow •Alternatively, the workflow documentation should give guidance on how to retrieve the data, preferably from a FAIR data repository. •example data can be used to verify the users’ configuration. Running a pipeline in another computational environment can require adjustments to the configuration file (see Tip 8). The example data with results can be used to verify that the workflow runs correctly with this new configuration profile. •test functions can be incorporated in the workflow to guarantee a proper workflow execution and, if not executed correctly, reveal quickly where the execution halts •Unit tests are small tests that can be implemented in a workflow to test the execution of single scripts or even functions within a script. For popular programming languages, there are libraries available that are designed to implement unit tests, Tip 5: The tools integrated in a workflow should adhere to file format standards •Adopting standardized file formats increases interoperability. Not only workflow inand output files, but also intermediate files that are exchanged by processes within the workflow should be written in standardized formats where possible. •It is important to realize that current data standards might not be persistent over time. Using data standards does not mean being blind to emerging data standards that possibly offer more advantages. • In the long run, it is a community effort to determine which domain-specific standards should be retained or replaced by better alternatives. Therefore, we recommend closely keeping track of the latest developments in the respective field that a researcher is working in. A reference for reproducible computational research Ten Simple Rules for Reproducible Computational Research | PLOS Computational Biology A reference for software tool workflowready ? Yet another example A final note… This Pilot training activity has been funded by the European Union –NextGenerationEU within the PNRR projects funded pursuant to Article 11, paragraph1, of Notice 594/2024: •“NFFA-DI cod. IR0000015, Missione 4, “Istruzione e Ricerca” – Componente 2, “Dalla ricerca all'impresa” – Linea di investimento 3.1, “Fondo per la realizzazione di un sistema integrato di infrastrutture di ricerca e innovazione” – Azione 3.1.1, “Creazione di nuove IR o potenziamento di quelle esistenti che concorrono agli obiettivi di Eccellenza Scientifica di Horizon Europe e costituzione di reti” (CUP B53C22004310006). •“EFC cod. SSU2024-00002, Missione 4 "Istruzione e ricerca" - Componente 1, “Potenziamento dell'offerta dei servizi all'istruzione: dagli asili nido all'universita'” - Investimento 3.4 “Didattica e competenze universitarie avanzate” - Sub-Investimento “Rafforzamento delle scuole universitarie superiori” (CUP: G97G24000100007). E-ARGO Thanks for the attention