scieee AI-readable full text Open interactive document viewer

ETP4HPC SRA 6 White Paper - Federation of Computing Infrastructure/Framework

Magugliani, Fabrizio; Hoppe, Hans-Christian; Nortamo, Henrik

Abstract

This is a white paper released as part of the ETP4HPC’s Strategic Research Agenda 6. A Federated Computing Infrastructure/Framework is the concrete implementation and operation term for a suite of technologies and functionalities aiming at the coordinated management and interoperability of data, resources, and processes across different systems, locations, and organizational boundaries. The framework envisions a decentralized architecture for creating and managing an interconnected network of resources where each participant (which could be a data centre, a system, an application, or any other device) may be used to achieve the intended result, breaking away from the traditional monolithic approach. This framework allows organizations to collaborate and share resources without fully giving up local control and thus to maintain a selectable level of autonomy. Federated Computing Infrastructures are designed from the beginning to be modular and adaptable, allowing resource-providing organizations to choose their level of involvement in the federation to ensure seamless integration with their existing platforms and regulatory requirements. The broader applicability of Federated Computing Infrastructure is pivotal for utilizing distributed data assets among and across organizations and within an organization. Federated Computing Infrastructures enable organizations and users to leverage distributed assets and drive innovation, collaboration, and value creation. The key characteristics of a Federated Computing Infrastructure/Framework are: 1. Decentralization: Each organisation participating in the Federated Computing Infrastructure makes available its resources while retaining a selectable level of autonomy and independence, 2. Interoperability: The Federated Computing Infrastructure is designed to enable the interaction among different systems built on different technologies, protocols, or standards. 3. Openness: The Federated Computing Infrastructure is open and accessible by any entitled organisation aiming at using its resources, provided that the authorization and data security criteria are met. 4. Scalability: The Federated Computing Infrastructure enables seamless scaling because new organizations, devices and resource-providing institutions can be added to the federation without interfering with the existing Infrastructure. 5. Data security and sovereignty: The concerns of data protection, data security and privacy-preserving computation are ubiquitous, and the Federated Computing Infrastructures provide tools and safeguards 6. Fault tolerance, redundancy and resiliency: Because resources are distributed across systems, locations, and organizational, the Federated Computing Infrastructure offers natural redundancy, enhancing the overall resiliency of the architecture. When one unit experiences a fault or is withdrawn from the pool of available resources, the problem is generally isolated to that particular unit, minimizing the impact on the entire system. The concept of Federated Computing Infrastructure/Framework provides a strategic framework that allows for the coordinated management and interoperability of data, resources, and processes. It’s particularly valuable in the EU environment for the exploitation of the computational resources deployed e.g. within the EuroHPC JU programs and initiatives, the Worldwide LHC Computing Grid (WLCG) and the likes, where it is of paramount importance for systems to be autonomous yet able to share compute power, data and resources effectively. However, the benefits come with their own set of challenges, such as increased complexity, inter-system security concerns and potential governance conflicts. Nonetheless, when designed and managed within a proper governance, adequate operational and policy guidelines, Federated Computing Infrastructures provide a powerful way to build flexible, scalable, and collaborative systems. This white paper explores the conceptual models of Federated Computing Infrastructure/Framework, the current deployments, the challenges facing the widespread exploitation of the Federated Computing Infrastructure and proposes a number of key R&I recommendations for a smooth and effective exploitation of the current and future deployments. Looking forward, the white paper outlines a Post Exascale Vision providing guidelines for designing future-proof Federated Computing Infrastructures.

Full text

WHITE PAPER Federation of Computing Infrastructure / Framework November 2025 Federation of Computing Infrastructure/Framework 2 Leading authors Fabrizio Magugliani (E4 Computer Engineering) Hans-Christian Hoppe (ParTec AG) Henrik Nortamo (CSC - IT Center for Science) Table of contents Executive Summary ..................................... 3 Glossary of acronyms .................................. 5 Introduction ............................................... 6 State of the Art ........................................... 7 Current HPC Federation Infrastructure/ Framework Landscape in the EU ................. 7 Enabling Technologies ................................. 7 EuroHPC Federation Platform ..................... 9 Current Federation Platforms in Development or Production ...................... 10 The open source aspect ............................ 13 Technologies for interconnecting FCI components .............................................. 13 Upcoming Challenges (2025 – 2028) .......... 15 Security ..................................................... 15 Management............................................. 15 Automatic Resource Discovery .............. 15 Resource Management ......................... 16 Adding Artificial Intelligence ................. 17 Adding Quantum Computing ................ 17 HPC-as-a-Service ................................... 17 Post Exascale Vision .................................. 18 Novel Approaches for Energy Consumption .................................................................. 18 Leverage AI to Manage Complexity ........... 18 Guarantee Interoperability ........................ 19 Integrate Quantum Computers in the Picture ....................................................... 19 Key R&I Recommendations ....................... 20 Conclusions .............................................. 21 Contributing Authors ................................ 22 References ................................................ 23 Federation of Computing Infrastructure/Framework 3 Executive Summary This is a white paper released as part of the ETP4HPC’s Strategic Research Agenda 6. A Federated Computing Infrastructure/Framework is the concrete implementation and operation term for a suite of technologies and functionalities aiming at the coordinated management and interoperability of data, resources, and processes across different systems, locations, and organizational boundaries. The framework envisions a decentralized architecture for creating and managing an interconnected network of resources where each participant (which could be a data centre, a system, an application, or any other device) may be used to achieve the intended result, breaking away from the traditional monolithic approach. This framework allows organizations to collaborate and share resources without fully giving up local control and thus to maintain a selectable level of autonomy. Federated Computing Infrastructures are designed from the beginning to be modular and adaptable, allowing resource-providing organizations to choose their level of involvement in the federation to ensure seamless integration with their existing platforms and regulatory requirements. The broader applicability of Federated Computing Infrastructure is pivotal for utilizing distributed data assets among and across organizations and within an organization. Federated Computing Infrastructures enable organizations and users to leverage distributed assets and drive innovation, collaboration, and value creation. The key characteristics of a Federated Computing Infrastructure/Framework are: 1. Decentralization: Each organisation participating in the Federated Computing Infrastructure makes available its resources while retaining a selectable level of autonomy and independence, 2. Interoperability: The Federated Computing Infrastructure is designed to enable the interaction among different systems built on different technologies, protocols, or standards. 3. Openness: The Federated Computing Infrastructure is open and accessible by any entitled organisation aiming at using its resources, provided that the authorization and data security criteria are met. 4. Scalability: The Federated Computing Infrastructure enables seamless scaling because new organizations, devices and resource-providing institutions can be added to the federation without interfering with the existing Infrastructure. 5. Data security and sovereignty: The concerns of data protection, data security and privacy-preserving computation are ubiquitous, and the Federated Computing Infrastructures provide tools and safeguards 6. Fault tolerance, redundancy and resiliency: Because resources are distributed across systems, locations, and organizational, the Federated Computing Infrastructure offers natural redundancy, enhancing the overall resiliency of the architecture. When one unit experiences a fault or is withdrawn from the pool of available resources, the problem is generally isolated to that particular unit, minimizing the impact on the entire system. The concept of Federated Computing Infrastructure/Framework provides a strategic framework that allows for the coordinated management and interoperability of data, resources, and processes. It’s particularly valuable in the EU environment for the exploitation of the computational resources deployed e.g. within the EuroHPC JU programs and initiatives, the Worldwide LHC Computing Grid (WLCG) and the likes, where it is of paramount importance for systems to be autonomous yet able to share compute power, data and Federation of Computing Infrastructure/Framework 4 resources effectively. However, the benefits come with their own set of challenges, such as increased complexity, inter-system security concerns and potential governance conflicts. Nonetheless, when designed and managed within a proper governance, adequate operational and policy guidelines, Federated Computing Infrastructures provide a powerful way to build flexible, scalable, and collaborative systems. This white paper explores the conceptual models of Federated Computing Infrastructure/Framework, the current deployments, the challenges facing the widespread exploitation of the Federated Computing Infrastructure and proposes a number of key R&I recommendations for a smooth and effective exploitation of the current and future deployments. Looking forward, the white paper outlines a Post Exascale Vision providing guidelines for designing future-proof Federated Computing Infrastructures. Federation of Computing Infrastructure/Framework 5 Glossary of acronyms AAI: Authentication and Authorization Infrastructure(s) API: Application Programming Interface CERN: Conseil Européen pour la Recherche Nucléaire EFP: EuroHPC Federation Platform EOSC: European Open Science Cloud FCI: Federated Computing Infrastructure GÉANT: Gigabit European Academic Network GUI: Graphical User Interface HPC: High Performance Computing LDAP: Lightweight Directory Access Protocol LHC: Large Hadron Collider OIDC: OpenID Connect RBAC: Role-Based Access Control REST: Representational State Transfer interface SAML: Security Assertion Markup Language SKA: Square Kilometre Array SIMPL: Synchronous Interprocess Messaging Project for LINUX SLURM: Simple Linux Utility for Resource Management WLCG: World-wide LHC Compute Grid Federation of Computing Infrastructure/Framework 6 Introduction This white paper describes various aspects and challenges related to the federation of computing and storage resources across Europe. As the term 'federation' is used quite often in this document, it is worth discussing its meaning in the context of complex computing systems. One of the core aspects of federation as a union of different computing centres is that each member of such a union keeps its autonomy while adhering to a common set of policies, which can sometimes mean that one entity must make a compromise on its side for the benefit of the whole union. The main benefits of such a federation are to increase usability and foster efficient access to similar computing resources, which in turn can increase the research output and foster innovation across various scientific domains that leverage supercomputing resources. Computing centres across Europe usually operate on national or even institutional level, adhering to local policies and operating principles. Some centres have decades-long experience operating supercomputers, while others are quite young and have yet to receive their first procured clusters. This is in stark contrast with commercial (cloud) computing resource providers, where each one exposes its services in its own unified way, thus creating a continuum of (virtually) infinite resources. Some aspects of this approach create a good user experience by hiding away parts of the technical complexities, which, at the same time, may be a deterrent to some of the more advanced power users of computing resources. In this work we mainly focus on the modern High Performance Computing (HPC) infrastructure, however there were similar efforts to unify computing resources in the past, often referred to as “grid”. The grid approach relies on massively distributed computing resources, often leveraging idle computing power across multiple institutions to form a powerful system. It shares some of the desired properties of the federation like decentralized control, resource sharing and open standards. However, at the same time, the grids often lacked the user experience (GUIs, APIs vs. terminal tools) which is common in the modern cloud environment and is being adopted by the HPC ecosystem as well. One of the well-known grid systems is the Worldwide LHC Computing Grid (WLCG) [1] tailored specifically for particle physics applications, or the SETI@Home [2] project, which was used to analyse radio signals from radio telescopes in a screensaver running on home desktop computers. The latter has been generalised as BOINC [3], a general-purpose grid system still widely used by several scientific fields. In this white paper we discuss current state and challenges related to establishing a federation of diverse computing centres around Europe by seeking balance between the user friendliness, efficient access and maintaining autonomy of each resource provider. Increasing accessibility and user friendliness can help flatten the learning curve needed to start using the supercomputing infrastructures, which can in turn attract more users from both research and industry areas as well as shorten the time needed to develop new innovations and valuable research outcomes. Clean and well-defined APIs and GUIs for the federated access to the computing resources are first step, while the set of actual set of the unified services spans over aspects such as user Authentication and Authorization Infrastructure(s) (AAI), compute time allocation management, interactive access, declarative workflows and distributed data management and transfers. These aspects are further elaborated in the sections below. The next stage involves the consolidation of pre-processed data in data lakes in the cloud as well as locally. These centralized storage solutions enable extensive data aggregation and management, providing scalable storage and Federation of Computing Infrastructure/Framework 7 State of the Art Current HPC Federation Infrastructure/ Framework Landscape in the EU The current HPC landscape in the European Union is characterised by the co-existence of Tier0, Tier-1 and Tier-2 HPC systems across multiple performance ranges (from single PetaFlop/s to ExaFlop/s) operated in dedicated HPC centres which are funded at the European, national or regional level. Similarly, the HPC centres also operate high-performance tiered storage systems sized to feed their supercomputers to enable efficient execution of the day-to-day workloads. Federation of resources is usually provided within the centres themselves, yet federation across HPC centres and/or with data-oriented infrastructures is still not supported today in any general way. The Federated Computing Infrastructure (FCI) concept is of course not new, and it is supported by specific infrastructures for data-intensive science domains. First attempts to develop and deploy a FCI date back to 1997, with the Unicore0F 1 and Globus1F 2 activities in Europe and the US, leading to deployments and of course the development of “Grid Computing” systems and the standardisation efforts of the Global (or Open) Grid Forum2F 3. From this, a number of both multidisciplinary and domainspecific federated infrastructures did emerge, including EGI3F 4 ,WLCG4F 5 and FENIX, the latter two targeting high-energy physics and brain sciences. For the HPC and AI systems co-funded by the EuroHPC Joint Undertaking5F 6, a federation platform named the “EuroHPC Federation Platform” or EFP currently under development will be generally deployed in 2026. In addition, the 1 See https://www.unicore.eu/; the system is still in use today in certain niches 2 See https://toolkit.globus.org/; the Globus toolkit has been retired from active use 3 See https://en.wikipedia.org/wiki/Open_Grid_Forum 4 See https://www.egi.eu/ SKA project will stand up several regional centres6F 7 in the near future which will support federation of radio-astronomy data and compute resources. Enabling Technologies Building and operating an FCI will require a number of key functionalities to be provided, including local user administration, accounting of system use, policies and mechanisms to grant access quotas, secure compute and data access protocols, scheduling and resource management/orchestration services, and abstraction of usage/programming environments. To get the most out of a FCI, end users will also require methods to define cross-resource or cross-site dynamic workflows, and services which schedule, orchestrate and execute the workflow steps (which can be compute or data analysis/transfer tasks) on the best suited resources in the FCI. At the OS level, all HPC centres rely on standard Linux user ids; sometimes (and in centre-specific ways), additional information (such as project ids) is added, or several users are subsumed under a larger project for accounting use of resources and setting quotas. File access control is usually performed on the base of the local user ids. The traditional, local authentication used in HPC (such as usernames and associated secrets like passwords or private keys) is not sufficient for access to a federation, since the identification used and/or the means to verify it differ between organisations. Authentication and Authorization Infrastructure(s) (AAI) systems like 5 See https://home.cern/science/computing/grid 6 See https://eurohpc-ju.europa.eu/supercomputers/oursupercomputers_en and https://eurohpc-ju.europa.eu/ai-factories_en 7 See https://www.skao.int/en/science-users/119/ska-regional-centres Federation of Computing Infrastructure/Framework 8 Shibboleth7F 8, MyAccessID8F 9 and Keycloak9F 10 enable single-sign by end users based on a single identity and way of verification; access to the constituent organisations and resources is then provided based on transactions between the AAI and the local access control systems. Customarily, the single-sign-on identity is customarily mapped onto local credentials, and AAI systems use standardised protocols and APIs like OpenID, OAuth and SAML and interface with local systems such as LDAP and Active Directory. Policies for granting compute or data quotas and general access do differ across centres, with two basic mechanisms being prevalent: (i) granting privileges based on an evaluation of a written access request, which describes the intended use of resources, the scientific results to be obtained and the ways to ensure efficient use of resources, and (ii) granting access based on membership of an end-user in a larger project, to which a centre has pledged a certain use of resources to. Access is usually limited to 1 - 2 years, and can be extended. Encrypted connections with ssh generally provide access to supercomputers (or more precisely, their login nodes); depending on the centre setup, intermediate “springboard” systems are used for added security. Centres pose different requirements as to how to secure the private keys – password-protection is now required by most, and some centres have started to require multi-factor authentication. An alternative approach, for instance implemented by CSCS in Switzerland, enables users to create tokens after initial authentication (using ssh 8 Shibboleth (https://www.shibboleth.net/) was originally developed by the late 1990s IUnternet2 project in the US; it is available as OPen Source and widely used in federation infrastructures. 9 MyAccessID is a European solution supported by GÉANT – see https://wiki.geant.org/display/MyAccessID; it used, for instance, by the EFP and FENIX federations 10 See https://www.keycloak.org/ – the system is used widely by OSS Cloud solutions 11 See https://heappe.eu/ 12 See https://www.redhat.com/en/topics/api/what-is-arest-api 13 See https://openondemand.org/ and https://jupyter.org/ methods) that can be used for a prescribed period of time (24 hours to access resources and services without providing private key passwords or a second authentication factor. Besides access to a Linux shell, some centres, such as CSCS (FirecREST) and IT4I (HEAppE)10F 11, provide REST11F 12 interfaces and other centres, such as VEGA, provide GUI access with Open On Demand or Jupyter Notebooks12F 13 For data access, the fallback solution provided by the HPC centres is scp; for higher bandwidth transfers, multi-stream solutions such as Globus Connect13F 14 or FTS14F 15 can be used, again with differences across the centres. Systems for data federation such as SIMPL or Rucio (see the section below) can be layered on top. For batch-oriented scheduling and resource management on HPC systems, SLURM15F 16 is the incumbent solution, with the newer Flux16F 17 scheduler gaining ground, such as on the world’s largest system at the time of writing (El Capitan). Many data analytics and AI workloads use scheduling and resource management systems such as Kubernetes17F 18 which support interactive and dynamically changing workloads, and how to combine the capabilities of both classes of systems without compromising efficiency on large HPC systems is an active research and development topic. HPC use (certainly by scientists) traditionally relies on issuing Linux shell commands and interacting with bespoke application interfaces. Since HPC systems are becoming increasingly different (different CPU architectures/features, accelerators, networks, SW stacks and versions), it is a challenge for users to be 14 See https://docs.globus.org/guides/tutorials/managefiles/transfer-files/ 15 See https://fts.web.cern.ch/fts/ 16 Slurm is open source controlled by a commercial company (SchedMD) – see https://slurm.schedmd.com/ 17 Flux (https://computing.llnl.gov/projects/flux-buildingframework-resource-management) is an open source scheduler & resource manager originally developed by the US DoE Lawrebnce Livermore lab, to be transitioned developed under the auspices of the HPC SW foundation, which is part of the Linux Software Foundation 18 Kubernetes is a very widely used solution for the orchestration of containerised workloads – see https://kubernetes.io/ Federation of Computing Infrastructure/Framework 9 productive in using a variety of systems and achieve near-optimal efficiency and performance across these. Grid Computing research in the late 1990s and early 2000s has proposed approaches to define and deploy higher-level interfaces which hide the differences from application developers and end-users. Notable systems which provided such an abstraction for general-purpose HPC end-uses include NICE EnginFrame18F 19 and Unicore19F 20; domain-specific solutions include WLCG, which provides highenergy physics scientists with high-level services for data analytics and simulation which hide the myriad system differences. Recent SW solutions include Eviden’s Nimbix20F 21 offering. EuroHPC Federation Platform21F 22 In October 2023, the EuroHPC JU launched a call for tender targeting the federation of supercomputers, quantum computers, and data management resources. A consortium led by CSC-IT Centre for Science was awarded the contract to build the EuroHPC JU Federation Platform (EFP) in December 2024. The EFP addresses diverse user requirements by providing unified access to extensive EuroHPC computing and storage resources. The platform targets users from both private and public sectors, including small and mediumsized enterprises. It will provide a single point of access to EuroHPC supercomputing and quantum resources, including AI Factories. The platform aims to facilitate federated access to data lakes and data spaces across Europe by seamlessly integrating both private and public solutions, including established platforms like SIMPL, EOSC, and FENIX. The platform implements several features important for establishing a successful federation of compute and storage resources while keeping their autonomy and upholding specific requirements of their local environments. A key 19 NICE was acquired by Amazon - a description of the EnginFrame Web portal is available at https://docs.aws.amazon.com/enginframe/latest/ag/about.html. 20 The European Unicore system was in active use until the 2020s, with large activities like the Human Brain project enabling feature, the unified user Authentication and Authorization Infrastructure(s) (AAI), uses the MyAccessID service operated by GÉANT and leverage modern protocols such as OIDC in combination with certificate based SSH authentication. This will allow the enforcement of modern multi-factor authentication both for all web-based access methods as well as access via direct SSH access. The EFP provides provisioning, management and tracking of projects and allocations granted by the EuroHPC JU across all the federated systems. For this the EFP utilizes Waldur. The hosting entities keep their own resource management systems, an integration between them and the EFP is created and tailored in cooperation with each hosting entity. The source for the granted resources is the separate EuroHPC peer-review platform (https://access.eurohpcju.europa.eu/). Traditionally, HPC resources have typically been accessed through SSH and command line interfaces. While still providing direct access to all federated resources, the EFP also provides web-based access to the federated resources, either via offering direct command line access via a browser or through various powerful and high-level graphical interfaces for e.g. workflows and data management. adopting the technology to provide high-level interfaces to domain scientists – see https://www.unicore.eu/. 21 See https://eviden.com/solutions/high-performancecomputing/hpc-as-a-service/nimbix-federated-supercomputing/. 22 https://my-eurohpc.eu/ Federation of Computing Infrastructure/Framework 16 systems, like the case of federation, provide advantages for the end users since available resources can scale more than what is feasible within the single infrastructure (horizontal scale). However, coordinating the resource on the single infrastructures requires to keep and periodically distribute status information. ARD aims at providing a set of methods to automate the process of determining which resources are available within the DS in terms of their type, number, capabilities and constraints, –ontologies [9], [13]; it can be extended to a more flexible way of managing large datasets as well [14]. As such, DSs use a unified abstraction of the resources that eases the implementation of discovery protocols. This latter calls for an adequate data sharing mechanism which allows to: i) make sure that the information is consistent across the various infrastructures, ii) maintain a reliable communication upon an unreliable communication network(s), and iii) dealing with failures and unavailability of resources without impacting on the overall federation. To this purpose, distributed key-values datastores and distributed databases can be used. However, one of the challenges that still remains in the implementation of ARD for federated HPC centres is the lack of a uniform abstraction method. This is particularly true if the federation extends to the computing continuum to the Cloud, edge and IoT resources [9]. Resource Management Resource management (RM) is a critical point in an HPC centre, and becomes even more critical in the context of a federation of compute resources. RM touches different dimensions in the management spectrum, where different metrics (sometimes also contrasting each other) are used to define the ultimate goal. On the one hand, (energy) efficiency46F 47 is a growing point of attention from the building blocks to the system level [26], since being able to reduce the energy consumption is of primary importance to reduce the overall costs for running the infrastructures. An effective RM system should select the most energy efficient set of 47 https://top500.org/lists/green500/2025/06/ 48 https://nephele-project.eu/ compute resources to complete the user tasks; however, challenges come from special cases, like urgent computing applications, that impose strong constraints to the RM systems. On the other hand, failures and resource unavailability in a large-scale, distributed environment introduce disturbances on the RM’s planning for which a quick and dynamic counteraction is required. To this end, well known approaches, such as checkpointing and restarts [16] should be revised to take into account the vast diversity of resources and software components. The continuous growing in size of HPC and AI infrastructures [17], [18], as well their federation just emphasizes the need for better understanding the impact of failures on the typical workloads execution, and for proposing novel methods to enhance resiliency of large-scale distributed systems. Another challenge relates to the typical batch modality of allocating resources in an HPC centre: being able to distribute the users’ tasks among different infrastructures may require a clever resource allocation plan to avoid introducing large delays in the execution of tasks due to long queueing times. This situation is exacerbated in a federated system since different RM systems (e.g., SLURM, PBS, etc.) at the various infrastructures may enforce different configurations and resource allocation policies. To this purpose, federated systems will have to implement an additional coordination layer (meta-orchestrator) that, taking advantage from distributed databases and datastores and from an (homogenous) abstraction of the resources (already used to support ARD), can quickly generate distributed resource allocation plans and use underlying RM systems as executor entities. As such, Cloud and network domains offer a large base of solutions from which to source [19], as well as in examples already funded EU projects47F 48, 48F 49, 49F 50. Furthermore, this top layer is expected to provide the user interface and thus it represents the main entry point for the users to the federated system. 49 https://aeros-project.eu/ 50 https://www.8ra.com/ Federation of Computing Infrastructure/Framework 17 Adding Artificial Intelligence Artificial Intelligence (AI) is becoming more prominent even in many scientific and industrial applications, with foundational models (FMs) and large language models (LLMs) driving a large industrial sector in recent years. Despite the ‘demands for high-performance’ commonality with traditional HPC applications, large differences exist in the software stacks and in the way users may use compute resources [20]. Software stacks are profoundly different, as well as the hardware needs; while HPC applications are bound to high-precision arithmetic operations (FP64), AI models demand less arithmetic precision for their operations, but still need strong scaling. This motivated the emergence of specialized data formats (BF16, FP16, FP8, FP4, etc.) and the customization of the hardware, which ultimately culminated in the design of dedicated large infrastructures (AI factories and upcoming Giga-AI factories). Indeed, modern supercomputers tend to provide a compromise between supporting traditional HPC applications and new AI-tailored ones. Also, RM tends to be quite different from that of traditional HPC systems, since model design needs a more interactive process which contrasts with the typical batch scheduling of HPC systems. So, the integration of AI tailored infrastructures within a federated system appears as a good approach to, on the one hand, keep the infrastructure specialization and, on the other side, extend the range of resources (both in terms of number and type) that users can explore within their workflows. Again, most of the challenges are moved up to the meta-orchestration layer, which needs to be aware of all these infrastructural peculiarities. Interestingly, the tendency of AI to become more ubiquitous, brings new opportunities to tackle most of the aforementioned challenges. In this sense, the emergence of agent-based AI models can ease the management of such a huge number and type of resources, towards a more energy efficient and sustainable future [21]. Adding Quantum Computing Quantum computing (and with less emphasis quantum communication) revolution is not just a dream anymore, but it has become a concrete reality. Europe already made large investments in quantum computing technologies, and a growing number of smaller scale quantum computers (QC) started to be installed in research centres and private entities. Despite the fast and continuous progress in building ever larger quantum computers strong limitations still exist in their scaling up, which is a key aspect for approaching the quantum advantage. The challenge is further emphasised by the need of being able to map as much as possible logical qubits on top of noisy physical ones. This said, it is quite easy to see how quantum computers are and still will be a scarce resource in current and future HPC infrastructures. From this standpoint, the integration of QC within existing HPC machines is a critical, albeit challenging aspect. Indeed, the specificities of the QC that derive from the chosen technological basis (e.g., neutral atoms, photonics, superconducting, etc.) must be considered within the RM policies. This translates into a challenging problem of appropriately and efficiently abstracting such special computing resources even at the level of the meta-orchestrator. This is a challenging problem because, the involved timescales for running ‘quantum’ tasks are quite different among the different technological implementations of QC, and are also quite different from the involved timescales of the classical domain (queueing time, job execution, etc.). HPC-as-a-Service In the context of “HPC-as-a-Service” (HaaS), opening the supercomputing resources to external end-users needs to provide remote access APIs but also to guarantee the integrity of the overall platform federation in order to avoid security issues. For that purpose, there is a need to define access modalities describing how end-users can access the Federation of computing infrastructure from a technical point of view but also from a legal point of view. A contract between end-users and the Federation authority needs to be defined to ensure a good and safe operation of the overall system. Federation of Computing Infrastructure/Framework 18 Post Exascale Vision Exascale computers are a reality, but the race for even more performance is still running. There is no doubt that scientific and engineering applications are and will remain eager for getting access to highest performance machines; however, their construction and operational costs are growing correspondingly, making just few entities capable of sustaining them. The situation for the AI factories and future Giga-AI factories50F 51 will not be different. These are envisioned as very large-scale facilities whose primary purpose will be the development, training and deployment of next generation AI models at the scale of trillions of parameters, by running on an adequate infrastructure built on top of hundreds of thousands of AI-tailored processors. Although their ultimate goal, software stack and resource delivery models are different from those of traditional (post-exascale) HPC systems, the need for performance will not stop; on the contrary, it will be well aligned to that of the future forefront supercomputers. Novel Approaches for Energy Consumption As a consequence of the need for more and more performance, energy consumption (and consequently carbon footprint) of such large infrastructures is growing year by year [25], a situation that demands for novel approaches to be sustainable in the coming future. In this scenario, computer resource federation is a promising approach to systemizing a multitude of resources covering diverse scales. Indeed, a federated approach offers a way of harmonising the different software layers of the various infrastructures without the need of being very invasive or forcing large restructuring of their architecture. Federation will also empower end users to get access to very specialized infrastructures at large scale without any compromise, thus helping to progress in science and 51 EuroHPC-JU AI GigaFactories - Workshop Presentation - 22 May 2025 engineering domains, as well as to support new discoveries. As previously stated, there are a number of challenges to address in order to effectively federate compute resources; however, being able to pursue this goal will allow us to benefit from a more efficient and sustainable use of distributed heterogeneous resources. Cloud computing and networking represent a good starting point, since it is in their nature to provide on-demand access to distributed compute resources. Leverage AI to Manage Complexity To implement in the near future a federation of heterogeneous compute infrastructures, AI can be of primary importance to provide new and more intelligent tools for managing such complex systems [21], as well as to enhance their energy efficiency and security. As such, we can envision a federation of infrastructures, where the meta-orchestration layer is built upon a set of distributed agents which are in charge of allocating local compute resources, as well as delegate execution of specific tasks to other federated entities. Such a decentralized (peer-to- Federation of Computing Infrastructure/Framework 19 peer) approach will help to keep open the door for computing resources that span from traditional HPC to AI-factories, Cloud resources, network of edge/IoT devices and Quantum Computers. Cloudification of the HPC resource management is also supported by modern HPC clusters that in many cases include dedicated Cloud partitions, and that can be exploited to run application components as well as orchestration services that need to last for a long time. They can be also used to facilitate the integration with large deployments of IoT and Edge devices, through more neutral and standardized interfaces. Guarantee Interoperability Interoperability should be guaranteed by the creation and adoption of abstraction models that are agnostic with respect to the underlying implementations. For instance, workflows can be described at a high-level focusing more on expressing task dependency rather than explicitly imposing a set of resources where to execute. Examples of such infrastructure-agnostic intermediate workflow representations already exist [22], but a more general consensus and standardization is still required. A decentralized approach should provide more entry points for the users without impacting on the security aspects of the infrastructures' management; also, it comes with the beneficial effect of scaling more easily with respect to more centralized approaches. Integrate Quantum Computers in the Picture Looking more closely at the integration of Quantum Computers, as previously stated, there are a quite large set of open challenges. One of the critical aspects to deal with is their scarcity; indeed, quantum computers are and will remain in the near future an unbalanced resource compared to the number of compute nodes available on modern supercomputers. Depending on the composition of quantum tasks and classical ones in a workflow, and depending on the targeted quantum machine (i.e., there can be different technologies at the basis of a QC that imply different time scales for executing a quantum task), different approaches can be considered. Malleability can make more flexible and dynamic the acquisition and release of classical and quantum resources, resulting in a less constrained execution environment that may easily fit with the requirements of multiple concurrent users. Similarly, quantum resource virtualization can help to deal with the execution of tasks on a QC that last longer than their classical counterparts [23], [24]. Federation of Computing Infrastructure/Framework 20 Key R&I Recommendations Federated Computing Infrastructures/Frameworks are a decentralized architecture for creating and managing an interconnected network of resources where each participant may be used to achieve the result, breaking away from the traditional monolithic approach, This framework allows organizations to collaborate and share resources without fully giving up local control and thus to maintain a selectable level of autonomy. Federated Computing Infrastructures/Frameworks represent a key component of the European strategy for providing a user-friendly, harmonized and heterogeneous set of compute, storage and networking components and attracting a larger number of users (Academic, SMEs, ROs, Industries) to the use of HPC and AI resources. To reap the full benefits of Federated Computing Infrastructures/Frameworks, the following R&I recommendations are suggested: 1. Continuously explore how new technologies (e.g. 6G, IoT/edge/Cloud continuum, Quantum Computing, Neuromorphic Computing, Nano Technologies, Green Energy Technologies, Sustainable Technologies) could be seamless and economically integrated into the Federated Computing Infrastructures/Frameworks 2. Co-design and implement programs and projects of future Federated Computing Infrastructures/Frameworks based on these new technologies involving the end users and develop proper objectives, KPI, assessment metrics based on a real scenario 3. Enable the interconnection and interoperability with different networks (e.g. GEANT) and with the next generation of mobile communication standard (e.g. 6G) 4. Enable the dynamic integration in EuroHPC’s future infrastructures (e.g. Gigafactories), private infrastructures as well as cloud providers 5. Enable common and user-friendly remote access modalities for any end-user aiming at the widest cross-compatibility and crossoperation among different Federated Computing Infrastructures/Frameworks 6. Design from the beginning any Federated Computing Infrastructures/Frameworks to implement multi-level, tight secure access and controls with respect to data protection, data security and privacy-preserving computation to a) Enable seamless access for open science access, b) Enforce stringent criteria and full compliance with current regulations for privacy-sensitive data (e.g. medical) 7. Define a financial model for dealing with licensed software, like ANSYS, MATLAB, and any copyrighted or licensed applications/environments 8. Inform of the advantages of large-scale compute facilities new user/groups/communities, which are currently not familiar with the classical use of supercomputing resources or are used to a different set of tools or paradigms Please note that the actual implementation of the above recommendations in the current infrastructures does not require any technologi-cal breakthroughs or discoveries. Federation of Computing Infrastructure/Framework 21 Conclusions Federated Computing Infrastructure/Framework aims at providing seamless access to High Performance Computing-based infrastructures and services to a wide range of users from the research and scientific community, as well as the industry including SMEs, and the public sector, fostering science and engineering applications by providing uninterrupted and secure access to supercomputing resources for the new and emerging data and compute-intensive applications and services. Federated Computing Infrastructure/Framework offers several key benefits, primarily centred around service continuity, data and application security, compliance with privacy regulation, scalability, access to diverse data, and efficiency in data analysis. By enabling computations without centralizing them, Federated Computing Infrastructure/Framework minimize the risk of system downtime and resource oversubscription. While the key tools for creating and managing a Federated Computing Infrastructure/Framework are available or are in an advanced stage of development, harvesting the full benefits of Federated Computing Infrastructure/Framework requires tight cross-organizational collaboration and standardized set of APIs to enable the system interoperability among diverse, heterogeneous, geographically distributed platforms. Federation of Computing Infrastructure/Framework 22 Contributing Authors Pierre-Yves Danet, 6G Industry Association Venkatesh Kannan, ICHEC (Irish Centre for High-End Computing) Alberto Scionti, LINKS Foundation Martin Golasowski, IT4Innovations, VSB - Technical University of Ostrava Jan Martinovič, IT4Innovations, VSB - Technical University of Ostrava Federation of Computing Infrastructure/Framework 23 References [1] WLCG: https://home.cern/science/computing/grid [2] SETI@home: Data Analysis and Findings: https://arxiv.org/abs/2506.14737 [3] BOINC https://boinc.berkeley.edu/pubs.php [4] Recordon, David, and Drummond Reed. "OpenID 2.0: a platform for user-centric identity management." Proceedings of the second ACM workshop on Digital identity management. 2006. [5] Baker, Matthew. "OAuth2." Secure Web Application Development: A Hands-On Guide with Python and Django. Berkeley, CA: Apress, 2022. 351-397. [6] Ferry, Eugene, John O Raw, and Kevin Curran. "Security evaluation of the OAuth 2.0 framework." Information & Computer Security 23.1 (2015): 73-101. [7] Konvička, Jakub, Václav Svatoň, and Jan Křenek. "HEAppE Middleware: From Desktop to HPC." European Conference on Parallel Processing. Cham: Springer Nature Switzerland, 2023. [8] Hachinger, Stephan, et al. "Leveraging high-performance computing and cloud computing with unified big-data workflows: the LEXIS project." Technologies and Applications for Big Data Value. Cham: Springer International Publishing, 2021. 159-180. [9] Zhou, Aolong, et al. "Semantic-based discovery method for high-performance computing resources in cyber-physical systems." Microprocessors and Microsystems 80 (2021): 103328. [10] Lazaro, Daniel, et al. "Decentralized resource discovery mechanisms for distributed computing in peer-to-peer environments." ACM Computing Surveys (CSUR) 45.4 (2013): 1-40. [11] Datta, Soumya Kanti, Rui Pedro Ferreira Da Costa, and Christian Bonnet. "Resource discovery in Internet of Things: Current trends and future standardization aspects." 2015 IEEE 2nd world forum on internet of things (WF-IoT). IEEE, 2015. [12] Lazaro, Daniel, et al. "Decentralized resource discovery mechanisms for distributed computing in peer-to-peer environments." ACM Computing Surveys (CSUR) 45.4 (2013): 1-40. [13] Castañé, Gabriel G., et al. "An ontology for heterogeneous resources management interoperability and HPC in the cloud." Future Generation Computer Systems 88 (2018): 373-384. [14] Liao, Chunhua, et al. "Hpc ontology: Towards a unified ontology for managing training datasets and ai models for high-performance computing." 2021 IEEE/ACM Workshop on Machine Learning in High Performance Computing Environments (MLHPC). IEEE, 2021 [15] Alam, Sadaf R., et al. "Federated Single Sign-On and Zero Trust Co-design for AI and HPC Digital Research Infrastructures." SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2024. [16] Timalsina, Madan, et al. "Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC." arXiv preprint arXiv:2407.19117 (2024). [17] Kokolis, Apostolos, et al. "Revisiting reliability in large-scale machine learning research clusters." 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2025. [18] Dixit, Harish Dattatraya, et al. "Silent data corruptions at scale." arXiv preprint arXiv:2102.11245 (2021). Federation of Computing Infrastructure/Framework 24 [19] Vaño, Rafael, Ignacio Lacalle, and Carlos E. Palau. "Federation of distributed domains in the CloudEdge-IoT Continuum." Proceedings of the 2nd International Workshop on MetaOS for the Cloud-EdgeIoT Continuum. 2025 [20] Dube, Nicolas, et al. "Future of HPC: The Internet of workflows." IEEE Internet Computing 25.5 (2021): 26-34. [21] Pauloski, J. Gregory, et al. "Empowering Scientific Workflows with Federated Agents." arXiv preprint arXiv:2505.05428(2025). [22] Colonnelli, Iacopo, et al. "Introducing SWIRL: An intermediate representation language for scientific workflows." International Symposium on Formal Methods. Cham: Springer Nature Switzerland, 2024. [23] Viviani, Paolo, et al. "Assessing the Elephant in the Room in Scheduling for Current Hybrid HPC-QC Clusters." arXiv preprint arXiv:2504.10520 (2025) [24] Döbler, Philip, and Manpreet Singh Jattana. "A Survey on Integrating Quantum Computers into High Performance Computing Systems." arXiv preprint arXiv:2507.03540 (2025). [25] Wu, Carole-Jean, et al. "Beyond efficiency: Scaling AI sustainably." IEEE Micro 44.5 (2024): 37-46. [26] Alvarez, Lluc, et al. "eProcessor: european, extendable, energy-efficient, extreme-scale, extensible, processor ecosystem." Proceedings of the 20th ACM International Conference on Computing Frontiers. 2023. Federation of Computing Infrastructure/Framework 25 Cite as: Magugliani F. et al., ETP4HPC SRA6 White Paper – Federation of Computing Infrastructure/Framework, 2025, ETP4HPC. https://doi.org/10.5281/zenodo.17544940 DOI: 10.5281/zenodo.17544940 © ETP4HPC 2025