scieee AI-readable full text Open interactive document viewer

Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science

Wilkinson, Sean; Widener, Patrick; Oral, Sarp; Ferreira da Silva, Rafael

Abstract

High Performance Computing (HPC) centers provide advanced infrastructure that enables scientific research at extreme scale. These centers operate with hardware configurations, software environments, and security requirements that differ substantially from most users' local systems. As a result, users often develop customized digital artifacts that are tightly coupled to a given HPC center. This practice can lead to significant duplication of effort as multiple users independently create similar solutions to common problems. The FAIR Principles offer a framework to address these challenges. Initially designed to improve data stewardship, the FAIR approach has since been extended to encompass software, workflows, models, and infrastructure. By encouraging the use of rich metadata and community standards, FAIR practices aim to make digital artifacts easier to share and reuse, both within and across scientific domains. Many FAIR initiatives have emerged within individual research communities, often aligned by discipline (e.g. bioinformatics, earth sciences). These communities have made progress in adopting FAIR practices, but their domain-specific nature can lead to silos that limit broader collaboration. Thus, we propose that HPC centers play a more active role in fostering FAIR ecosystems that support research across multiple disciplines. This requires designing infrastructure that enables researchers to discover, share, and reuse computational components more effectively. Here, we build on the architecture of the European Open Science Cloud (EOSC) EOSC-Life FAIR Workflows Collaboratory to propose a model tailored to the needs of HPC. Rather than focusing on entire workflows, we emphasize the importance of making individual workflow components FAIR. This component-based approach better supports the diverse and evolving needs of HPC users while maximizing the long-term value of their work.

Full text

Designing FAIR Workflows at OLCF ORNL/TM-2025/4241 Building Scalable and Reusable Ecosystems for HPC Science Sean R. Wilkinson Patrick Widener Sarp Oral Rafael Ferreira da Silva DESIGNING FAIR WORKFLOWS AT OLCF Disclaimer. This research used resources of the Oak Ridge Leadership Computing Facility at ORNL, which is supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC05-00OR22725. ORNL is managed by UT-Battelle LLC on behalf of the U. S. Department of Energy. This report was prepared as an account of work sponsored by agencies of the United States Government. Neither the United States Government nor any agency thereof, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. License. This report is made available under a Creative Commons Attribution 4.0 International Public license (https: //creativecommons.org/licenses/by/4.0). Preferred citation S. R. Wilkinson, P. Widener, S. Oral, R. Ferreira da Silva, “Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science”, Technical Report, ORNL/TM-2025/4241, November 2025, DOI: 10.5281/zenodo.17290392. @techreport{fair−olcf−2025, author = {Wilkinson, Sean R. and Widener, Patrick and Oral, Sarp and Ferreira da Silva, Rafael}, title = {{Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science}}, year = {2025}, publisher = {Zenodo}, number = {ORNL/TM−2025/4241}, doi = {10.5281/zenodo.17290392}, url = {https://doi.org/10.5281/zenodo.17290392}, institution = {Oak Ridge National Laboratory} } 2 DESIGNING FAIR WORKFLOWS AT OLCF Contents 1 INTRODUCTION AND MOTIVATION 4 2 BACKGROUND 4 2.1 THE FAIR PRINCIPLES .................................... 4 2.2 FAIR COMMUNITIES ..................................... 5 2.3 CHALLENGES AT HPC CENTERS .............................. 6 2.4 EXISTING US DOE EFFORTS ................................ 6 2.5 EOSC-LIFE FAIR WORKFLOWS COLLABORATORY ................... 6 3 FAIR ECOSYSTEM ARCHITECTURES 7 3.1 REPOSITORIES ........................................ 8 3.2 REGISTRIES .......................................... 9 3.3 COMPUTING INFRASTRUCTURE ............................. 9 3.4 AUTHENTICATION AND AUTHORIZATION ........................ 10 3.5 METADATA STANDARDS AND INTERCHANGE FORMATS ............... 10 4 PROPOSED ARCHITECTURE FOR OLCF 10 4.1 REPOSITORIES ........................................ 11 4.2 REGISTRIES .......................................... 13 4.3 COMPUTING INFRASTRUCTURE ............................. 14 4.4 AUTHENTICATION AND AUTHORIZATION ........................ 14 4.5 METADATA STANDARDS AND INTERCHANGE FORMATS ............... 15 5 ENCOURAGING ADOPTION OF FAIR AT OLCF 15 6 FUTURE DIRECTIONS 18 REFERENCES 20 3 DESIGNING FAIR WORKFLOWS AT OLCF 1 INTRODUCTION AND MOTIVATION High Performance Computing (HPC) centers, such as the Oak Ridge Leadership Computing Facility (OLCF), provide advanced infrastructure that enables scientific research at extreme scale. These centers operate with unique hardware configurations, specialized software environments, and elevated security requirements that differ substantially from what most users encounter on their local systems. As a result, users often develop customized digital artifacts that are tightly coupled to the specific configuration of a given HPC center. Although necessary, this practice can lead to significant duplication of effort as multiple users independently create similar solutions to common problems. The FAIR Principles, which stand for Findable, Accessible, Interoperable, and Reusable, offer a framework to address these challenges (Section 2.1). Initially designed to improve data stewardship, the FAIR approach has since been extended to encompass software, workflows, models, and infrastructure. By encouraging the use of rich metadata and community standards, FAIR practices aim to make digital artifacts easier to share and reuse, both within and across scientific domains. Many FAIR initiatives have emerged within individual research communities, often aligned by discipline, such as bioinformatics or earth sciences. These communities have made progress in adopting FAIR practices, but their domain-specific nature can lead to silos that limit broader collaboration. To overcome this, we propose that HPC centers play a more active role in fostering FAIR ecosystems that support research across multiple disciplines. This requires designing infrastructure that enables researchers to discover, share, and reuse computational components more effectively. In this report, we build on the architecture of the European Open Science Cloud (EOSC) EOSCLife FAIR Workflows Collaboratory [1] to propose a model that is tailored to the needs of HPC environments. Rather than focusing on entire workflows, which are often difficult to reuse due to rapid changes in HPC infrastructure, we emphasize the importance of making individual workflow components FAIR. This component-based approach better supports the diverse and evolving needs of HPC users while maximizing the long-term value of their work. The sections that follow present a conceptual architecture for implementing a FAIR ecosystem at OLCF, describe existing capabilities and gaps, and outline a strategy to encourage adoption among the user community. 2 BACKGROUND To design effective FAIR ecosystems for HPC environments, it is important to understand the foundational principles and existing efforts that inform this work. This section introduces the FAIR Principles, explores how various research communities have implemented them, and outlines specific challenges that arise in the context of HPC centers like OLCF. We also review related initiatives within the US Department of Energy and examine lessons learned from the EOSC-Life FAIR Workflows Collaboratory, which serves as a model for our proposed approach. 2.1 THE FAIR PRINCIPLES The “Findable, Accessible, Interoperable, Reusable” (FAIR) Principles for making digital objects Findable, Accessible, Interoperable, and Reusable were first established to improve the management and stewardship of data [2]. There have been a number of other efforts in this space, some of which even sound 4 DESIGNING FAIR WORKFLOWS AT OLCF similar or related to the word “fair”, such as the CARE Principles [3], the Fair-code software model1, the TRUST Principles [4], and several distinct efforts which use the name “SHARE Principles” [5,6]. What makes FAIR unique and distinguishes it from other peer initiatives is that, rather than making data easy for just humans to use, it places specific emphasis on enhancing the ability of machines to automatically find and use data, in ways that also happen to support reuse by humans. To accomplish this, FAIR focuses on metadata in order to make data • Findable, so that metadata and data are findable for both humans and computers (e.g., persistent identifiers, rich metadata, indexing/registering); • Accessible, so that users know how to access the data (e.g., clear access protocols, being able to access metadata even if the data have disappeared); • Interoperable, so that data can be used by applications and workflows for analysis, storage, and processing (e.g., machine processable metadata using standards); and • Reusable, so that data reuse is optimized via comprehensive, well-described metadata (e.g., metadata standards, data usage licenses, provenance). Since that original publication, FAIR’s focus on metadata has found broad application beyond data to research software [7], computational workflows [8], open hardware [9], Artificial Intelligence (AI) and Machine Learning (ML) models [10], and even facilities and instruments [11]. In fact, in 2021, Oak Ridge National Laboratory (ORNL) performed some of the earliest research on applying FAIR to workflows [12], and OLCF hosted a lab-wide workshop on the topic [13]. OLCF was also instrumental in recent efforts to enumerate the FAIR Principles for computational workflows formally [14], following in the same format and even the same journal that previously published the principles for data [2] and research software [15]. Following the FAIR Principles helps make research artifacts more easily discoverable, shareable, and reusable. These principles improve collaboration, transparency, and efficiency in scientific research by ensuring that artifacts like code and data are well-documented, allowing them to be reused effectively across different studies and disciplines both by humans and machines. There may be challenges in implementing the FAIR Principles [16], but they do play an important role in facilitating reproducibility, accelerating discoveries, and supporting long-term stewardship of research artifacts. Their adoption helps researchers maximize the impact of their artifacts while fostering open science and innovation. 2.2 FAIR COMMUNITIES The FAIR Principles focus on metadata practices to be implemented by vaguely defined “communities”; the various publications have all deliberately left this up to interpretation. In practice, these communities tend to gather by domain (e.g. bioinformatics, geosciences, and agriculture). Such efforts have improved sharing and reusing artifacts within their respective disciplines, but these domain-based communities can end up functioning as silos. This hampers sharing and reusing artifacts as well as the adoption of the FAIR Principles themselves across computational research areas. For example, the bioinformatics community has made significant strides in developing FAIR repositories such as the European Nucleotide Archive [17]. Similarly, the geoscience community has developed metadata-rich repositories such as Earth Observing System Data and Information System (EOSDIS) [18], 1https://faircode.io/ 5 DESIGNING FAIR WORKFLOWS AT OLCF which provides comprehensive access to satellite and remote sensing data while supporting research in areas such as climate change, natural disasters, and ecosystem monitoring. However, interoperability between these repositories remains limited, reducing the potential for cross-domain collaboration. 2.3 CHALLENGES AT HPC CENTERS HPC centers present additional challenges not present in commodity computing environments that must be addressed when implementing FAIR Principles [19]. In order to provide resources to users who require greater scale to “get science done”, HPC centers frequently deploy infrastructure with exotic hardware architectures, cutting-edge software environments, and stricter security measures as compared with users’ own resources [20]. Variability in system architectures, ranging from CPU-based clusters to GPUand TPU-accelerated platforms, complicates software portability and reproducibility. Furthermore, interdisciplinary and inter-project collaboration within and across HPC centers remains challenging due to security constraints, data governance policies, and technical barriers. As a result, users often create, configure, and customize digital artifacts in ways that are specialized for the unique infrastructure at a given HPC center. This means that the research artifacts produced at HPC centers have limited potential for reuse by others unless they are shared in ways that are discoverable by other users of those HPC centers. 2.4 EXISTING US DOE EFFORTS Several efforts in the United States (US) Department of Energy (DOE) share the common goal of advancing scientific research by providing cutting-edge computational, data, and networking resources to support interdisciplinary collaboration: Integrated Research Infrastructure (IRI) [21], Framework for Accelerating Science and Scientific Training (FASST) Framework [22], and National Artificial Intelligence Research Resource (NAIRR) [23]. These initiatives focus on improving access to HPC, data management, and security frameworks to foster cross-disciplinary research and enable more efficient, data-driven scientific discoveries. They emphasize the importance of open science principles and FAIR, ensuring that research data and tools are accessible and reusable. These efforts are more concerned with enabling the autonomous laboratories of the future, however, rather than the human users of HPC facilities. 2.5 EOSC-LIFE FAIR WORKFLOWS COLLABORATORY The EOSC-Life FAIR Workflows Collaboratory is a European initiative aimed at enhancing the FAIRness of scientific workflows in life sciences [1,24]. It provides a collaborative platform, including WorkflowHub [25], that enables researchers to share, develop, and execute workflows across various life science domains, ensuring that data, tools, and computational resources are well-documented and standardized for broader use. It leverages the EOSC to foster collaboration between institutions and projects, promote transparency, and improve the efficiency of research through the integration of diverse data sources and computational tools. The initiative is crucial for supporting reproducibility, advancing data-driven discoveries, and accelerating innovation in European life sciences. In this report, we argue that users could increase impactful scientific output if HPC centers enhanced their unique environments with a FAIR-based strategy similar to the one pursued by EOSC-Life. We observe that a focus on FAIR workflow components may be more impactful than focusing on FAIR for entire workflows. We also propose a general architecture and discuss how to encourage adoption by users. 6 DESIGNING FAIR WORKFLOWS AT OLCF 3 FAIR ECOSYSTEM ARCHITECTURES Constructing a FAIR ecosystem is an ambitious undertaking because such an ecosystem has many moving parts, including data, software, hardware, services, and most importantly, users. Luckily, there is an existing ecosystem for us to analyze to abstract some of the general design patterns to inform our own designs, which are detailed in Section 4. The EOSC-Life FAIR Workflows Collaboratory provides an example for constructing an ecosystem that incorporates FAIR “all the way down” to aid domain scientists to produce impactful scientific artifacts that are shareable, discoverable, and reusable. A diagram depicting many of the moving parts is shown in Figure 1. Figure 1: The EOSC-Life FAIR Workflows Collaboratory [24]. Although the EOSC-Life Collaboratory was constructed specifically to meet the needs of European life scientists, it does accurately portray many of the complications shared by domain scientists in general and, we argue, those faced by the users of HPC facilities. There are some key differences, however. The EOSC-Life Collaboratory focuses mainly on the practice of open science, and life scientists try to maintain infrastructure for as long as possible to avoid breaking the pipelines they have constructed. This results in a stronger focus on applying FAIR to workflows as a whole. We argue that FAIR ecosystems for HPC centers are better served by focusing on the sharing, discovery, and reuse of workflow components rather than entire workflows. This is partly because HPC workflows are often short-lived due to being customized for HPC resources with short design lifetimes (due to the rapid pace of hardware and software innovation). For example, the flagship machines at centers like OLCF are designed for a lifespan of 5 years; after that, executing the same workflows on a machine’s successor is almost guaranteed to fail. Having witnessed this process repeatedly, we have noticed that workflow 7 DESIGNING FAIR WORKFLOWS AT OLCF components themselves are less likely to break, but they will have to be re-plumbed to port the workflow to the new machine. Another reason for focusing on components is that, in our experience, it is extremely rare to locate an existing off-the-shelf workflow that exactly matches your needs. Users at HPC centers come from a variety of scientific disciplines, so it is unlikely that any particular workflow will match the needs of other users exactly. Because these users are all solving challenges in trying to scale using the same HPC resources, however, it is much more likely that solutions to separable pieces of problems will be reusable across disciplines, supporting a component-based approach. For these reasons we propose to focus on applying FAIR to workflow components, making sure they are well-described by rich metadata. This strategy requires a combination of repositories, registries, computing infrastructure, authentication/authorization, and metadata standards. This also allows us to abstract many of the details from Figure 1into the simpler Figure 2, which depicts an abstract FAIR ecosystem at an HPC center. Registry Services Component repository Data repository Container repository Software repository Supercomputer HPC center Figure 2: Illustrated examples of a FAIR ecosystem in action at an HPC center. Dashed lines (––) show the registry’s records tracking components wherever they are. Dotted lines (· · ·) show services in an onprem cloud monitoring the registry and component repository for updates and providing support such as databases to supercomputer jobs. Solid lines (—) show that the supercomputer can consult the registry to locate components and request the service infrastructure to retrieve external artifacts if needed [19]. 3.1 REPOSITORIES Repositories, roughly speaking, are places where things are stored; this is in contrast to registries (see below), which are where records of things are stored. Registries and repositories provide curation and best practices for recording workflows [26]. Workflow components are what workflows are made of, which can include various forms of data (e.g. simulation output, trained AI/ML models, and provenance) and software (e.g. scripts, programs, and computational workflows). Having a robust repository system aligns with IRI-style distributed, collaborative science and its aims for creating a seamless, interoperable research infrastructure across DOE labs. The FAIR Principles do not imply openness, but we also recommend following open science princi8 DESIGNING FAIR WORKFLOWS AT OLCF ples [27] when possible by enabling integration with externally hosted public repositories (e.g. Conda2, Spack3, Docker Hub4, TensorFlow Hub5). Security or scale concerns may not allow for open dissemination or integration, however. HPC centers typically handle their own user identity management, and those operated by US Government entities must conform to well-specified security policies. Scale of workflow components tends to be more a technical issue; centers typically address these through data transport middleware designed for large scale (e.g. Globus [28]) and specialized repositories (e.g. DataFed [29]) which can encapsulate both data transport and storage capabilities. These concerns can also take different forms for commercial cloud compute providers than at leadership-class HPC centers where the ability to scale in compute and data is a goal in itself. 3.2 REGISTRIES Registries, as mentioned previously, are places where records of artifacts are stored, and they can be distinct from the repositories that contain the artifacts described by the records. A well-designed registry such as WorkflowHub [25] is critically important for enabling a FAIR ecosystem anywhere, but particularly as illustrated in 2, they are crucial. Registries complement repositories by maintaining metadata records for workflow components, ensuring the discoverability and interoperability of research artifacts via metadata and Persistent Identifiers (PIDs) such as Digital Object Identifiers (DOIs). In this case, the registry stores metadata about the artifacts as well as where they are stored, allowing users to search for artifacts relevant to their needs across repositories both internal and external to the HPC center. 3.3 COMPUTING INFRASTRUCTURE Computing infrastructure, although not digital, must still be described by rich metadata. Reproducibility and reusability rely on preservation and discovery of the computing contexts in which workflows are executed. It is therefore important that metadata about the resources and infrastructure provided by an HPC center be made available and exportable in useful forms. Tools such as FlowCept [30] help to capture provenance about execution and environment that should be stored and linked to the workflow components. Reproducing execution context is a critical enabler for the seamless cross-center data exchange and workflow portability that IRI needs. Following the model of the EOSC-Life Collaboratory [1] requires infrastructure for service execution, separate from but adjacent to HPC supercomputing resources. Infrastructure provided by Kubernetes6and OpenStack7-based platforms, for example, enables persistent services for automatically testing workflow component viability as well as continuously indexing and updating registries in response to artifact updates in repositories. Commonly available services running on instances of these platforms also reduce the need for workflows to deploy their own versions (e.g., a key-value store). Validation and benchmarking services help align with NAIRR, ensuring that AI/ML models trained on HPC resources remain FAIR-compliant and reusable even as the FAIR Principles are updated in the future. 2https://anaconda.org/anaconda/conda 3https://spack.io/ 4https://hub.docker.com/ 5https://www.tensorflow.org/hub 6https://kubernetes.io/ 7https://www.openstack.org/ 9 DESIGNING FAIR WORKFLOWS AT OLCF Despite the clear benefits of FAIR practices, their widespread implementation within the OLCF user community remains inconsistent. This is due in part to deeply rooted research norms, varied domain-specific data practices, and the perceived cost or complexity of adopting new standards. Driving cultural change in such an environment requires not only policy guidance and technical support, but also a systematic approach to shifting community behaviors and incentives. Figure 4: Strategy for Culture Change. Image credit: https://www.cos.io/blog/ strategy-for-culture-change. In this context, the “Strategy for Culture Change” proposed by Brian Nosek in a 2019 blog post25 provides a useful framework. Nosek suggests that sustainable cultural transformation can be achieved by altering the default behaviors in a system through three coordinated levers: infrastructure, incentives, and norms. Applying this framework to the OLCF community offers a pathway to embed the FAIR Principles into standard research workflows gradually, making the FAIR choice not only possible, but easy and rewarding. The diagram shown in Figure 4illustrates an approach that works from the bottom up, showing that for us to encourage adoption of FAIR among OLCF users, we must first make adoption possible, especially at the infrastructure level. As we showed in Section 4, OLCF already possesses most of the pieces to construct a FAIR ecosystem. The most glaring hole that is missing is a registry for tracking components. OLCF needs a registry to make a FAIR ecosystem possible. After making adoption possible, Nosek recommends that OLCF make adoption easy. More often than not, OLCF users are highly technically skilled in addition to being leaders in their scientific fields. It is a lot to ask these users to become FAIR experts, too! We can “make it easy” by providing simple user interfaces to walk them through what would otherwise seem like very complicated processes for registering and sharing their artifacts, or for finding relevant existing artifacts they can reuse. Some efforts are already underway, such as the FlowCept tool [30] for capturing provenance about a workflow’s execution and environment; we should also provide interfaces for storing and linking to the workflow components themselves. The API for S3M is a user interface that is designed to simplify some very complicated systems on behalf of the user. 25https://www.cos.io/blog/strategy-for-culture-change 16 DESIGNING FAIR WORKFLOWS AT OLCF In both cases, adopting FAIR practices is already possible, but OLCF is providing tools and services that make adoption significantly easier. Next, Figure 4shows that OLCF should make FAIR adoption normative, which requires psychology beyond the scope of this report to fully explain. The example given by Nosek for shifting norms is to make the desired behaviors visible, such as when journals adopt badges to acknowledge authors who preregister their studies and share their data and materials. Those awarded badges become signals to other researchers as they read the articles both that their colleagues do these behaviors (descriptive norm setting) and that the journal values these behaviors (prescriptive norm setting). At OLCF, we might spotlight certain projects in newsletters or at the monthly user group meetings to increase visibility for the desired behaviors. Visibility is critical for accelerating adoption among those who are willing but who have not yet adopted. For those users who are not quite as willing, the remaining two steps of the strategy attempt to make adoption rewarding before ultimately making adoption required. Colloquially, this is also known as the “carrot and stick” metaphor26, which evokes imagery of leading a horse either towards a positive incentive (eating a tasty carrot) or away from a negative consequence (being forcefully thrashed with a stick). Positive incentives for users include the creation of reward systems that recognize and incentivize researchers for sharing their components. This can be achieved through citation and attribution mechanisms which formally acknowledge the contributions of those who make their computational resources, data, or code available for public use. At an HPC center like OLCF, however, there are other currencies available, too, such as increased allocations and job priority. Allocations (measured in node-hours) and job priority (measured in wait time) both represent ways to reward time back to users who actively work to save time for other users. It will be important to balance such rewards against potential resulting disincentives, such as users’ equating reuse of shared components with reduction in resources available to themselves. Rather than utilizing “negative incentives”, however, OLCF would most likely enact and enforce policies to make adoption required among the final users who have still held out. Receiving allocations of resources at OLCF could be conditioned on FAIR compliance and participation in the workflow component and research artifact-sharing ecosystem. Currently, allocations are frequently awarded at no cost to researchers or are governed by proposal processes funded by US taxpayers. Artifacts resulting from that usage should (perhaps after some embargo period to protect intellectual investment pending publication) be freely available and reusable. Researchers will want assurances that they will be appropriately credited for their efforts, and so implementation of such mandates will require careful consideration. The incentives for OLCF to construct a FAIR ecosystem and encourage its adoption are important, too. Not least among these incentives is to uphold the mission of OLCF itself, which is to accelerate scientific discovery and engineering progress by providing world-leading computational performance and advanced data infrastructure. The FAIR Principles maximize the value and impact of research artifacts [14], which is critically important for artifacts that are the products of hundreds of thousands of node-hours on a Top50027 machine like Frontier [45]. It is prohibitively expensive to reproduce such artifacts, when it is possible to reproduce them at all without access to a leadership-class supercomputer. Frontier is a time-sensitive scientific instrument with a finite lifespan; avoiding recomputing results not only speeds time to discovery for subsequent studies, but it also frees the machine to continue crunching numbers on behalf of new artifacts that could lead to additional discoveries. In short, the benefits of adopting FAIR allow Frontier to produce more results that can lead to more scientific discovery, thereby increasing the overall scientific impact of Frontier and of OLCF. 26https://en.wikipedia.org/wiki/Carrot_and_stick 27https://top500.org 17 DESIGNING FAIR WORKFLOWS AT OLCF 6 FUTURE DIRECTIONS Having discussed in previous sections why and how FAIR relates to HPC centers in general and OLCF in particular, we now similarly connect and relate FAIR with the recently announced American Science Cloud (AmSC). At the time of this writing, not much is publicly known about the AmSC beyond what has been revealed by the One Big Beautiful Bill [46], which defines it as “a system of United States government, academic, and private sector programs and infrastructures utilizing cloud computing technologies to facilitate and support scientific research, data sharing, and computational analysis across various disciplines while ensuring compliance with applicable legal, regulatory, and privacy standards.” A bundled summary of its enclosing section provides a little more information, however: This section provides funding for partnerships between the National Laboratories and U.S. industry to organize DOE data for use in artificial intelligence and machine learning models. DOE must also initiate seed efforts for self-improving artificial intelligence models for science and engineering using this data. These models must be provided to the scientific community through a system of programs and infrastructure using cloud computing. This section also allows this data to be used to develop next-generation microelectronics. From this information, we can deduce that the future AmSC may have quite a lot in common with the EOSC, the website28 for which states that “[t]he ambition of the European Open Science Cloud, known as EOSC, is to develop a ‘Web of FAIR Data and Services’ for science in Europe.” This sentence is intended to evoke the “Web of Linked Data” [47] concept introduced in 2006, as updated by the FAIR Principles published 10 years later [2]. Interestingly, the European Commission began pursuing EOSC in 2015, which now provides 10 years’ worth of potentially relevant insights for constructing the AmSC. In Sections 3and 4, we repeatedly referenced insights and patterns from the EOSC-Life FAIR Workflows Collaboratory when designing a FAIR ecosystem for workflows for HPC centers and for OLCF. Without loss of generality, these same arguments and recommendations would apply to the broader AmSC, from what is currently known. The AmSC does seem to place greater focus on AI / ML models than the EOSCLife ecosystem shown in Figure 1, however, and so for that, we recommend following the work of the DOE-funded HPC-FAIR project [48]. HPC-FAIR was a multi-institutional project that aimed to develop a generic HPC data management framework to make both training data and AI models of scientific applications FAIR [10]. The framework was designed to centralize HPC-related datasets and AI models within a unified hub and to ensure interoperability through a standardized representation and vocabulary (ontology) for both data and models. The project also implemented automated workflows for streamlined data processing, model access, and benchmarking, and it focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics [48]. In a nutshell, the application of the FAIR Principles at every level of an ecosystem – be it inside OLCF or a much more geographically distributed system like the AmSC – is crucial for enabling autonomous, self-assembling workflows. Although Large Language Model (LLM)s have proven capable of generating performant code on demand [49], the value of building from vetted components has been a cornerstone of software engineering for quite some time. For any humanor machine-based system to build from components, those components will inevitably need to be Findable, Accessible, Interoperable, and Reusable, 28https://eosc.eu/eosc-about/ 18 DESIGNING FAIR WORKFLOWS AT OLCF which immediately implies that many of us who lay the groundwork now must learn FAIR or waste time reinventing it. FAIR is certain to be a critical enabler and accelerator of the construction of the AmSC. 19 DESIGNING FAIR WORKFLOWS AT OLCF References [1] C. Goble, S. Soiland-Reyes, F. Bacall, S. Owen, A. Williams, I. Eguinoa, B. Droesbeke, S. Leo, L. Pireddu, L. Rodr´ ıguez-Navas, J. M. Fern´ andez, S. Capella-Gutierrez, H. M´ enager, B. Gr¨ uning, B. Serrano-Solano, P. Ewels, and F. Coppens, “Implementing FAIR Digital Objects in the EOSC-Life Workflow Collaboratory,” Mar. 2021. doi: 10.5281/ZENODO.4605654. [Online]. Available: https://doi.org/10.5281/ZENODO.4605654 [2] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, J. Bouwman, A. J. Brookes, T. Clark, M. Crosas, I. Dillo, O. Dumon, S. Edmunds, C. T. Evelo, R. Finkers, A. Gonzalez-Beltran, A. J. Gray, P. Groth, C. Goble, J. S. Grethe, J. Heringa, P. A. ’t Hoen, R. Hooft, T. Kuhn, R. Kok, J. Kok, S. J. Lusher, M. E. Martone, A. Mons, A. L. Packer, B. Persson, P. Rocca-Serra, M. Roos, R. van Schaik, S.-A. Sansone, E. Schultes, T. Sengstag, T. Slater, G. Strawn, M. A. Swertz, M. Thompson, J. van der Lei, E. van Mulligen, J. Velterop, A. Waagmeester, P. Wittenburg, K. Wolstencroft, J. Zhao, and B. Mons, “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data, vol. 3, no. 1, p. 160018, Mar. 2016. doi: 10.1038/sdata.2016.18. [Online]. Available: https://doi.org/10.1038/sdata.2016.18 [3] S. R. Carroll, I. Garba, O. L. Figueroa-Rodr´ ıguez, J. Holbrook, R. Lovett, S. Materechera, M. Parsons, K. Raseroka, D. Rodriguez-Lonebear, R. Rowe, R. Sara, J. D. Walker, J. Anderson, and M. Hudson, “The care principles for indigenous data governance,” Data Science Journal, vol. 19, 2020. doi: 10.5334/dsj-2020-043. [Online]. Available: http://doi.org/10.5334/dsj-2020-043 [4] D. Lin, J. Crabtree, I. Dillo, R. R. Downs, R. Edmunds, D. Giaretta, M. De Giusti, H. L’Hours, W. Hugo, R. Jenkyns, V. Khodiyar, M. E. Martone, M. Mokrane, V. Navale, J. Petters, B. Sierman, D. V. Sokolova, M. Stockhause, and J. Westbrook, “The TRUST Principles for digital repositories,” Scientific Data, vol. 7, no. 1, May 2020. doi: 10.1038/s41597-020-0486-7. [Online]. Available: http://doi.org/10.1038/s41597-020-0486-7 [5] “Principles for conducting research in the arctic,” Interagency Arctic Research Policy Committee, Washington D.C, Tech. Rep., 2018. [Online]. Available: https://www.iarpccollaborations.org/uploads/ cms/documents/principles for conducting research in the arctic final 2018.pdf [6] S. S. Jordan, H. M. Matusovich, J. M. Case, L. Benson, D. A. Delaine, R. L. Kajfez, S. M. Lord, M. C. Paretti, E. T. Young, and Y. V. Zastavker, “Share: A framework for secondary qualitative data analysis,” Studies in Engineering Education, vol. 5, no. 1, p. 125–133, 2024. doi: 10.21061/see.175. [Online]. Available: http://doi.org/10.21061/see.175 [7] W. Hasselbring, L. Carr, S. Hettrick, H. S. Packer, and T. Tiropanis, “FAIR and Open Computer Science Research Software,” CoRR, vol. abs/1908.05986, 2019. [Online]. Available: http://arxiv.org/abs/1908.05986 [8] C. Goble, S. Cohen-Boulakia, S. Soiland-Reyes, D. Garijo, Y. Gil, M. R. Crusoe, K. Peters, and D. Schober, “FAIR Computational Workflows,” Data Intelligence, vol. 2, no. 1-2, pp. 108–121, 2020. doi: 10.1162/dint a 00033. [Online]. Available: https://doi.org/10.1162/dint a 00033 20 DESIGNING FAIR WORKFLOWS AT OLCF [9] N. Miljkovi´ c, A. Trisovic, and L. Peer, “Towards FAIR Principles for open hardware,” Feb. 2022. [Online]. Available: https://doi.org/10.5281/zenodo.6506428 [10] E. A. Huerta, B. Blaiszik, L. C. Brinson, K. E. Bouchard, D. Diaz, C. Doglioni, J. M. Duarte, M. Emani, I. Foster, G. Fox, P. Harris, L. Heinrich, S. Jha, D. S. Katz, V. Kindratenko, C. R. Kirkpatrick, K. Lassila-Perini, R. K. Madduri, M. S. Neubauer, F. E. Psomopoulos, A. Roy, O. R¨ ubel, Z. Zhao, and R. Zhu, “FAIR for AI: An interdisciplinary and international community building perspective,” Sci. Data, vol. 10, no. 1, p. 487, Jul. 2023. doi: 10.1038/s41597-023-02298-6. [Online]. Available: https://doi.org/10.1038/s41597-023-02298-6 [11] A. Johnson, R. Julian, M. Mayernik, C. Mundoma, M. Murray, A. Ranganath, and G. Stossmeister, “FAIR Facilities and Instruments Workshop #1 Report: Exploring Persistent Identifier Needs, Barriers, and Incentives,” National Center for Atmospheric Research, Tech. Rep., 2024. [Online]. Available: https://opensky.ucar.edu/islandora/object/technotes:601 [12] M. Wolf, J. Logan, K. Mehta, D. A. Jacobson, M. C. McDevitt, A. Walker, G. Eisenhauer, P. Widener, and A. Cliff, “Reusability First: Toward FAIR Workflows.” Oak Ridge National Lab. (ORNL), Oak Ridge, TN (United States), 09 2021. [Online]. Available: https://www.osti.gov/biblio/1827005 [13] S. Wilkinson, K. Knight, O. Kuchar, K. Mehta, M. Shankar, and M. Wolf, “Official report on the 2021 computational and autonomous workflows workshop (CAW 2021),” Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States), Tech. Rep., 03 2022. [Online]. Available: https://www.osti.gov/biblio/1862119 [14] S. R. Wilkinson, M. Aloqalaa, K. Belhajjame, M. R. Crusoe, B. de Paula Kinoshita, L. Gadelha, D. Garijo, O. J. R. Gustafsson, N. Juty, S. Kanwal, F. Z. Khan, J. K¨ oster, K. Peters-von Gehlen, L. Pouchard, R. K. Rannow, S. Soiland-Reyes, N. Soranzo, S. Sufi, Z. Sun, B. Vilne, M. A. Wouters, D. Yuen, and C. Goble, “Applying the FAIR Principles to computational workflows,” Scientific Data, vol. 12, no. 1, Feb. 2025. doi: 10.1038/s41597-025-04451-9. [Online]. Available: https://doi.org/10.1038/s41597-025-04451-9 [15] M. Barker, N. P. Chue Hong, D. S. Katz, A.-L. Lamprecht, C. Martinez-Ortiz, F. Psomopoulos, J. Harrow, L. J. Castro, M. Gruenpeter, P. A. Martinez, and T. Honeyman, “Introducing the FAIR Principles for research software,” Scientific Data, vol. 9, no. 1, p. 622, Oct. 2022. doi: 10.1038/s41597-022-01710-x. [Online]. Available: https://doi.org/10.1038/s41597-022-01710-x [16] S. R. Wilkinson, G. Eisenhauer, A. J. Kapadia, K. Knight, J. Logan, P. Widener, and M. Wolf, “F*** workflows: when parts of FAIR are missing,” in 2022 IEEE 18th International Conference on e-Science (e-Science). Salt Lake City, UT, USA: IEEE, Oct. 2022. doi: 10.1109/eScience55777.2022.00090. ISBN 9781665461245 pp. 507–512. [Online]. Available: https://doi.org/10.1109/eScience55777.2022.00090 [17] D. Yuan, A. Ahamed, J. Burgin, C. Cummins, R. Devraj, K. Gueye, D. Gupta, V. Gupta, M. Haseeb, M. Ihsan, E. Ivanov, S. Jayathilaka, V. B. Kadhirvelu, M. Kumar, A. Lathi, R. Leinonen, J. McKinnon, L. Meszaros, C. O’Cathail, D. Ouma, J. Paup´ erio, S. Pesant, N. Rahman, G. Rinck, S. Selvakumar, S. Suman, Y. Sunthornyotin, M. Ventouratou, S. Vijayaraja, Z. Waheed, P. Woollard, A. Zyoud, T. Burdett, and G. Cochrane, “The european nucleotide archive in 2023,” Nucleic Acids 21 DESIGNING FAIR WORKFLOWS AT OLCF Research, vol. 52, no. D1, pp. D92–D97, 11 2023. doi: 10.1093/nar/gkad1067. [Online]. Available: https://doi.org/10.1093/nar/gkad1067 [18] H. Ramapriyan and J. Behnke, “NASA’s Earth Observing System Data and Information System (EOSDIS) and FAIR - A Self-Assessment,” in AGU Fall Meeting Abstracts, vol. 2020, Dec. 2020, pp. IN044–08. [Online]. Available: https://ui.adsabs.harvard.edu/abs/2020AGUFMIN044..08R [19] S. R. Wilkinson and P. Widener, “FAIR Ecosystems for Science at Scale,” in Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration, ser. PEARC ’25. New York, NY, USA: Association for Computing Machinery, 2025. doi: 10.1145/3708035.3736097. ISBN 9798400713989. [Online]. Available: https://doi.org/10.1145/3708035.3736097 [20] K. B. Antypas, D. J. Bard, J. P. Blaschke, R. Shane Canon, B. Enders, M. A. Shankar, S. Somnath, D. Stansberry, T. D. Uram, and S. R. Wilkinson, “Enabling discovery data science through cross-facility workflows,” in 2021 IEEE International Conference on Big Data (Big Data), 2021. doi: 10.1109/BigData52589.2021.9671421 pp. 3671–3680. [Online]. Available: https://doi.org/10.1109/BigData52589.2021.9671421 [21] W. Miller, D. Bard, A. Boehnlein, K. Fagnan, C. Guok, E. Lanc¸on, S. J. Ramprakash, M. Shankar, N. Schwarz, and B. Brown, “Integrated research infrastructure architecture blueprint activity (final report 2023),” US Department of Energy (USDOE), Washington, DC (United States). Office of Science; Lawrence Berkeley National Laboratory (LBNL), Berkeley, CA (United States), Tech. Rep., 07 2023. [Online]. Available: https://www.osti.gov/biblio/1984466 [22] U.S. Department of Energy, “Frontiers in artificial intelligence for science, security, and technology initiative,” 2025, accessed: 2025-03-27. [Online]. Available: https://www.energy.gov/ science-innovation/energy-science-research/frontiers-artificial-intelligence-science-security-and [23] ——, “National artificial intelligence research resource,” 2025, accessed: 2025-03-27. [Online]. Available: https://www.energy.gov/science-innovation/national-artificial-intelligence-research-resource [24] S. R. Wilkinson, J. Gustafsson, F. Bacall, K. Belhajjame, S. Capella, J. M. F. Gonzalez, J. F. Tande, L. Gadelha, D. Garijo, P. Grubel, B. Gr¨ uning, F. Z. Khan, S. Kanwal, S. Leo, S. Owen, L. Pireddu, L. Pouchard, L. Rodr´ ıguez-Navas, B. Serrano-Solano, S. Soiland-Reyes, B. Vilne, A. Williams, M. A. Wouters, F. Coppens, and C. Goble, “An Ecosystem of Services for FAIR Computational Workflows,” 2025. [Online]. Available: https://arxiv.org/abs/2505.15988 [25] O. J. R. Gustafsson, S. R. Wilkinson, F. Bacall, S. Soiland-Reyes, S. Leo, L. Pireddu, S. Owen, N. Juty, J. Fern´ andez, T. Brown, H. M´ enager, B. Gr¨ uning, S. Capella-Gutierrez, F. Coppens, and C. Goble, “Workflowhub: a registry for computational workflows,” Scientific Data, vol. 12, no. 1, May 2025. doi: 10.1038/s41597-025-04786-3. [Online]. Available: https://doi.org/10.1038/s41597-025-04786-3 [26] R. Ferreira da Silva, R. M. Badia, V. Bala, D. Bard, T. Bremer, I. Buckley, S. Caino-Lores, K. Chard, C. Goble, S. Jha, D. S. Katz, D. Laney, M. Parashar, F. Suter, N. Tyler, T. Uram, I. Altintas, S. Andersson, W. Arndt, J. Aznar, J. Bader, B. Balis, C. Blanton, K. R. Braghetto, A. Brodutch, P. Brunk, H. Casanova, A. C. Lierta, J. Chigu, T. Coleman, N. Collier, I. Colonnelli, F. Coppens, M. Crusoe, W. Cunningham, B. De Paula Kinoshita, P. Di Tommaso, C. Doutriaux, M. Downton, 22 DESIGNING FAIR WORKFLOWS AT OLCF W. Elwasif, B. Enders, C. Erdmann, T. Fahringer, L. Figueiredo, R. Filgueira, M. Foltin, A. Fouilloux, L. Gadelha, A. Gallo, A. G. Saez, D. Garijo, R. Gerlach, R. Grant, S. Grayson, P. Grubel, J. Gustafsson, V. Hayot-Sasson, O. Hernandez, M. Hilbrich, AnnMary Justine, I. Laflotte, F. Lehmann, A. Luckow, J. Luettgau, K. Maheshwari, M. Matsuda, D. Medic, P. Mendygral, M. Michalewicz, Jorji Nonaka, M. Pawlik, L. Pottier, L. Pouchard, M. Putz, S. K. Radha, L. Ramakrishnan, S. Ristov, P. Romano, D. Rosendo, M. Ruefenacht, K. Rycerz, Nishant Saurabh, V. Savchenko, M. Schulz, C. Simpson, R. Sirvent, T. Skluzacek, S. Soiland-Reyes, R. Souza, S. R. Sukumar, Z. Sun, A. Sussman, D. Thain, M. Titov, B. Tovar, A. Tripathy, M. Turilli, B. Tuznik, H. Van Dam, A. Vivas, L. Ward, P. Widener, S. Wilkinson, J. Zawalska, and Mahnoor Zulfiqar, “Workflows Community Summit 2022: A Roadmap Revolution,” Oak Ridge National Laboratory, Tech. Rep. ORNL/TM-2023/2885, Mar. 2023. [Online]. Available: https://doi.org/10.5281/zenodo.7750670 [27] National Academies of Sciences, Engineering, and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. Washington, DC: The National Academies Press, 2018. ISBN 978-0-309-47624-9. [Online]. Available: https://nap.nationalacademies.org/catalog/25116/ open-science-by-design-realizing-a-vision-for-21st-century [28] I. Foster and C. Kesselman, “The globus project: a status report,” in Proceedings Seventh Heterogeneous Computing Workshop (HCW’98), 1998. doi: 10.1109/HCW.1998.666541 pp. 4–18. [Online]. Available: https://doi.org/10.1109/HCW.1998.666541 [29] D. Stansberry, S. Somnath, J. Breet, G. Shutt, and M. Shankar, “Datafed: Towards reproducible research via federated data management,” in 2019 International Conference on Computational Science and Computational Intelligence (CSCI), 2019. doi: 10.1109/CSCI49370.2019.00245 pp. 1312–1317. [Online]. Available: https://doi.org/10.1109/CSCI49370.2019.00245 [30] R. Souza, T. J. Skluzacek, S. R. Wilkinson, M. Ziatdinov, and R. F. da Silva, “Towards lightweight data integration using multi-workflow provenance and data observability,” in 2023 IEEE 19th International Conference on e-Science (e-Science), 2023. doi: 10.1109/e-Science58273.2023.10254822 pp. 1–10. [Online]. Available: https://doi.org/10.1109/e-Science58273.2023.10254822 [31] G. von Laszewski, W. Brewer, S. R. Wilkinson, A. Shao, J. P. Fleischer, H. Pitkar, C. R. Kirkpatrick, and G. C. Fox, “Towards experiment execution in support of community benchmark workflows for HPC,” 2025. [Online]. Available: https://arxiv.org/abs/2507.22294 [32] A. Gray, C. Goble, and R. Jimenez, “Bioschemas: From potato salad to protein annotation,” in ISWC 2017 Posters & Demonstrations and Industry Tracks, ser. CEUR workshop proceedings, N. Nikitina, D. Song, A. Fokoue, and P. Haase, Eds. Germany: RWTH Aachen University, Oct. 2017, the 16th International Semantic Web Conference 2017, ISWC 2017 ; Conference date: 21-10-2018 Through 25-10-2018. [Online]. Available: https://iswc2017.semanticweb.org/paper-579/ [33] F. Suter, T. Coleman, ˙ Ilkay Altintas¸, R. M. Badia, B. Balis, K. Chard, I. Colonnelli, E. Deelman, P. Di Tommaso, T. Fahringer, C. Goble, S. Jha, D. S. Katz, J. K¨ oster, U. Leser, K. Mehta, H. Oliver, J.-L. Peterson, G. Pizzi, L. Pottier, R. Sirvent, E. Suchyta, D. Thain, S. R. Wilkinson, J. M. Wozniak, and R. Ferreira da Silva, “A terminology for scientific workflow systems,” Future Generation Computer Systems, vol. 174, p. 107974, 2026. doi: https://doi.org/10.1016/j.future.2025.107974. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167739X25002699 23 DESIGNING FAIR WORKFLOWS AT OLCF [34] M. R. Crusoe, S. Abeln, A. Iosup, P. Amstutz, J. Chilton, N. Tijani´ c, H. M´ enager, S. Soiland-Reyes, B. Gavrilovi´ c, C. Goble, and T. C. Community, “Methods included: standardizing computational reuse and portability with the Common Workflow Language,” Communications of the ACM, vol. 65, no. 6, pp. 54–63, Jun. 2022. doi: 10.1145/3486897. [Online]. Available: https://doi.org/10.1145/3486897 [35] S. Leo, M. R. Crusoe, L. Rodr´ ıguez-Navas, R. Sirvent, A. Kanitz, P. De Geest, R. Wittner, L. Pireddu, D. Garijo, J. M. Fern´ andez, I. Colonnelli, M. Gallo, T. Ohta, H. Suetake, S. Capella-Gutierrez, R. de Wit, B. P. Kinoshita, and S. Soiland-Reyes, “Recording provenance of workflow runs with ro-crate,” PLoS one, vol. 19, no. 9, p. e0309210, 2024. [Online]. Available: https://doi.org/10.1371/journal.pone.0309210 [36] R. Ferreira da Silva, D. Bard, K. Chard, d. W. Shaun, I. T. Foster, T. Gibbs, C. Goble, W. Godoy, J. Gustafsson, U.-U. Haus, S. Hudson, S. Jha, L. Los, D. Paine, F. Suter, L. Ward, S. Wilkinson, M. Amaris, Y. Babuji, J. Bader, R. Balin, D. Balouek, S. Beecroft, K. Belhajjame, R. Bhattarai, W. Brewer, P. Brunk, S. Caino-Lores, H. Casanova, D. Cassol, J. Coleman, T. Coleman, I. Colonnelli, A. A. Da Silva, D. de Oliveira, P. Elahi, N. Elfaramawy, W. Elwasif, B. Etz, T. Fahringer, W. Ferreira, R. Filgueira, J. Fosso Tande, L. Gadelha, A. Gallo, D. Garijo, Y. Georgiou, P. Gritsch, P. Grubel, A. Gueroudji, Q. Guilloteau, C. Hamalainen, R. Hong Enriquez, L. Huet, K. Hunter Kesling, P. Iborra, S. Jahangiri, J. Janssen, J. Jordan, S. Kanwal, L. Kunstmann, F. Lehmann, U. Leser, C. Li, P. Liu, J. Luettgau, R. Lupat, J. M. Fernandez, K. Maheshwari, T. Malik, J. Marquez, M. Matsuda, D. Medic, S. Mohammadi, A. Mulone, J.-L. Navarro, K. W. Ng, K. Noelp, B. P. Kinoshita, R. Prout, M. R. Crusoe, S. Ristov, S. Robila, D. Rosendo, B. Rowell, J. Rybicki, H. Sanchez, N. Saurabh, S. K. Saurav, T. Scogland, D. Senanayake, W. Shin, R. Sirvent, T. Skluzacek, B. Sly-Delgado, S. Soiland-Reyes, A. Souza, R. Souza, D. Talia, N. Tallent, L. Thamsen, M. Titov, B. Tovar, K. Vahi, E. Vardar-Irrgang, E. Vartina, Y. Wang, M. Wouters, Q. Yu, Z. Al Bkhetan, and M. Zulfiqar, “Workflows Community Summit 2024: Future Trends and Challenges in Scientific Workflows,” Oak Ridge National Laboratory, Tech. Rep. ORNL/TM-2024/3573, 2024. [Online]. Available: https://doi.org/10.5281/zenodo.13844759 [37] S. S. Vazhkudai, J. Harney, R. Gunasekaran, D. Stansberry, S.-H. Lim, T. Barron, A. Nash, and A. Ramanathan, “Constellation: A science graph network for scalable data and knowledge discovery in extreme-scale scientific collaborations,” in 2016 IEEE International Conference on Big Data (Big Data), 2016. doi: 10.1109/BigData.2016.7840959 pp. 3052–3061. [Online]. Available: https://doi.org/10.1109/BigData.2016.7840959 [38] J. Ison, K. Rapacki, H. M´ enager, M. Kalaˇ s, E. Rydza, P. Chmura, C. Anthon, N. Beard, K. Berka, D. Bolser, T. Booth, A. Bretaudeau, J. Brezovsky, R. Casadio, G. Cesareni, F. Coppens, M. Cornell, G. Cuccuru, K. Davidsen, G. D. Vedova, T. Dogan, O. Doppelt-Azeroual, L. Emery, E. Gasteiger, T. Gatter, T. Goldberg, M. Grosjean, B. Gr¨ uning, M. Helmer-Citterich, H. Ienasescu, V. Ioannidis, M. C. Jespersen, R. Jimenez, N. Juty, P. Juvan, M. Koch, C. Laibe, J.-W. Li, L. Licata, F. Mareuil, I. Miˇ ceti´ c, R. M. Friborg, S. Moretti, C. Morris, S. M¨ oller, A. Nenadic, H. Peterson, G. Profiti, P. Rice, P. Romano, P. Roncaglia, R. Saidi, A. Schafferhans, V. Schw¨ ammle, C. Smith, M. M. Sperotto, H. Stockinger, R. S. Vaˇ rekov´ a, S. C. Tosatto, V. de la Torre, P. Uva, A. Via, G. Yachdav, F. Zambelli, G. Vriend, B. Rost, H. Parkinson, P. Løngreen, and S. Brunak, “Tools and data services registry: a community effort to document bioinformatics resources,” Nucleic Acids 24 DESIGNING FAIR WORKFLOWS AT OLCF Research, vol. 44, no. D1, pp. D38–D47, 11 2015. doi: 10.1093/nar/gkv1116. [Online]. Available: https://doi.org/10.1093/nar/gkv1116 [39] J. Ison, H. Ienasescu, P. Chmura, E. Rydza, H. M´ enager, M. Kalaˇ s, V. Schw¨ ammle, B. Gr¨ uning, N. Beard, R. Lopez, S. Duvaud, H. Stockinger, B. Persson, R. S. Vaˇ rekov´ a, T. Raˇ cek, J. Vondr´ aˇ sek, H. Peterson, A. Salumets, I. Jonassen, R. Hooft, T. Nyr¨ onen, A. Valencia, S. Capella, J. Gelp´ ı, F. Zambelli, B. Savakis, B. Leskoˇ sek, K. Rapacki, C. Blanchet, R. Jimenez, A. Oliveira, G. Vriend, O. Collin, J. van Helden, P. Løngreen, and S. Brunak, “The bio.tools registry of software tools and data resources for the life sciences,” Genome Biology, vol. 20, no. 1, Aug. 2019. doi: 10.1186/s13059-019-1772-6. [Online]. Available: https://doi.org/10.1186/s13059-019-1772-6 [40] F. da Veiga Leprevost, B. A. Gr¨ uning, S. Alves Aflitos, H. L. R¨ ost, J. Uszkoreit, H. Barsnes, M. Vaudel, P. Moreno, L. Gatto, J. Weber, M. Bai, R. C. Jimenez, T. Sachsenberg, J. Pfeuffer, R. Vera Alvarez, J. Griss, A. I. Nesvizhskii, and Y. Perez-Riverol, “Biocontainers: an opensource and community-driven framework for software standardization,” Bioinformatics, vol. 33, no. 16, p. 2580–2582, Mar. 2017. doi: 10.1093/bioinformatics/btx192. [Online]. Available: https://doi.org/10.1093/bioinformatics/btx192 [41] R. Julian, A. Johnson, M. Mayernik, C. Mundoma, M. Murray, and A. Ranganath, “FAIR Facilities and Instruments Workshop #2 Report: Recent Progress, Remaining Challenges, and Emerging PID Strategies,” NSF National Center for Atmospheric Research, Tech. Rep., 2024. [Online]. Available: https://opensky.ucar.edu/islandora/object/technotes:42004 [42] T. J. Skluzacek, P. Bryant, A. Ruckman, D. Rosendo, S. Prentice, M. J. Brim, R. Adamson, S. Oral, M. Shankar, and R. Ferreira da Silva, “Secure api-driven research automation to accelerate scientific discovery,” in Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration, ser. PEARC ’25. New York, NY, USA: Association for Computing Machinery, 2025. doi: 10.1145/3708035.3736072. ISBN 9798400713989. [Online]. Available: https://doi.org/10.1145/3708035.3736072 [43] S. Soiland-Reyes, P. Sefton, M. Crosas, L. J. Castro, F. Coppens, J. M. Fern´ andez et al., “Packaging research artefacts with RO-Crate,” Data Science, vol. 5, no. 2, pp. 97–138, Jul. 2022. doi: 10.3233/DS-210053. [Online]. Available: https://doi.org/10.3233/DS-210053 [44] H. L. Rehm, A. J. Page, L. Smith, J. B. Adams, G. Alterovitz, L. J. Babb, M. P. Barkley, M. Baudis, M. J. Beauvais, T. Beck, J. S. Beckmann, S. Beltran, D. Bernick, A. Bernier, J. K. Bonfield, T. F. Boughtwood, G. Bourque, S. R. Bowers, A. J. Brookes, M. Brudno, M. H. Brush, D. Bujold, T. Burdett, O. J. Buske, M. N. Cabili, D. L. Cameron, R. J. Carroll, E. Casas-Silva, D. Chakravarty, B. P. Chaudhari, S. H. Chen, J. M. Cherry, J. Chung, M. Cline, H. L. Clissold, R. M. Cook-Deegan, M. Courtot, F. Cunningham, M. Cupak, R. M. Davies, D. Denisko, M. J. Doerr, L. I. Dolman, E. S. Dove, L. J. Dursi, S. O. Dyke, J. A. Eddy, K. Eilbeck, K. P. Ellrott, S. Fairley, K. A. Fakhro, H. V. Firth, M. S. Fitzsimons, M. Fiume, P. Flicek, I. M. Fore, M. A. Freeberg, R. R. Freimuth, L. A. Fromont, J. Fuerth, C. L. Gaff, W. Gan, E. M. Ghanaim, D. Glazer, R. C. Green, M. Griffith, O. L. Griffith, R. L. Grossman, T. Groza, J. M. Guidry Auvil, R. Guig´ o, D. Gupta, M. A. Haendel, A. Hamosh, D. P. Hansen, R. K. Hart, D. M. Hartley, D. Haussler, R. M. Hendricks-Sturrup, C. W. Ho, A. E. Hobb, M. M. Hoffman, O. M. Hofmann, P. Holub, J. S. Hsu, J.-P. Hubaux, S. E. Hunt, 25