scieee AI-readable full text Open interactive document viewer

Deliverable D6.3 Initial Design of 6G Smart Network Management Framework

Landi, Giada

Abstract

This document describes the initial design and implementation of the enablers for 6G Smart Network Management, contributing to the 2nd iteration of the Hexa-X-II end-to-end system blueprint. The deliverable provides an initial overview of the Smart Network Management framework defined in the project. It analyses its contributions to the Hexa-X-II architecture design principles and lists the technical enablers for the management and orchestration functionalities, highlighting the main enhancements and innovations with respect to the Hexa-X architecture.The categorization of the enablers has been updated with respect to D6.2, to provide an improved and unified picture of their role and contribution in the Hexa-X-II Smart Network Management system. The enablers cover the functionalities and objectives of a 6G network Management & Orchestration (M&O) system, providing the technical foundations for programmable network configuration and monitoring, network capability exposure, synergetic resource orchestration across the continuum, and zero-touch network automation. All these functionalities are supported by advanced techniques for security and trustworthiness, AI/ML algorithms, and network digital twins.For each enabler, the document illustrates the design and presents the ongoing implementation, providing early validation results and discussing the enablers’ contributions to KPIs and KVIs. The role of the enablers in Hexa-X-II PoCs is presented as well, with initial integration and evaluation results.Finally, as key contribution towards the design of the end-to-end Hexa-X-II system, the document analyses the alignment of the M&O enablers with the end-to-end system blueprint, identifying the mapping with the current components and potential gaps to be filled with new elements or interfaces. This will feed the next iteration of the end-to-end Hexa-X-II system design, under development in WP2, following the overall methodology defined in the project.

Full text

A holistic flagship towards the 6G network platform and system, to inspire digital transformation, for the world to act together in meeting needs in society and ecosystems with novel 6G services Deliverable D6.3 Initial Design of 6G Smart Network Management Framework Hexa-X-II project has received funding from the Smart Networks and Services Joint Undertaking (SNS JU) under the European Union’s Horizon Europe research and innovation programme under Grant Agreement No 101095759. Date of delivery: 07/11/2025 Version: 1.1 Project reference: 101095759 Call: HORIZON-JU-SNS-2022 Start date of project: 01/01/2023 Duration: 30 months Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 2 / 207 Document properties: Document Number: D6.3 Document Title: Initial Design of 6G Smart Network Management Framework Editor(s): G. Landi (NXW) Authors: R. Vilalta, R. Muñoz, P. Alemany, Ll. Gifre, C. Manso, B. Ojaghi (CTT), C. Ayimba, A. Calvillo (UC3), G. Landi, P.G. Giardina, M. De Angelis, P. Piscione (NXW), W. Tavernier, J. Miserez (IMEC), S. Rodríguez, X.R. Sousa, F. Lamela (OPT), T. Dimitrovski, N. Toumi (TNO), V. Lamprousi, S. Barmpounakis, P. Demestichas, A. Karaolanis, V. Tsekenis, C. Karousatou (WIN), E. Lluesma Marti, L.F. González (ATO), I. Labrador Pavón (ASA), R. Nicolicchia (TID), I. Tzanettis, Grigorios Kakkavas, and A. Zafeiropoulos (ICC). Contractual Date of Delivery: 30/06/2024 Dissemination level: PU Status: Public version Version: 1.1 File Name: Hexa-X-II_D6.3_v1.1 Revision History Revision Date Issued by Description 1.0 28.06.2024 Hexa-X-II WP6 First public version 1.1 07.11.2025 Hexa-X-II WP6 Broken links fixed in Section 8.2.3 (Interfaces of the Management Capabilities Exposure framework). Abstract This document describes the initial design and implementation of the enablers for 6G Smart Network Management, contributing to the 2nd iteration of the Hexa-X-II end-to-end system blueprint. The deliverable provides an initial overview of the Smart Network Management framework defined in the project. It analyses its contributions to the Hexa-X-II architecture design principles and lists the technical enablers for the management and orchestration functionalities, highlighting the main enhancements and innovations with respect to the Hexa-X architecture. The categorization of the enablers has been updated with respect to D6.2, to provide an improved and unified picture of their role and contribution in the Hexa-X-II Smart Network Management system. The enablers cover the functionalities and objectives of a 6G network Management & Orchestration (M&O) system, providing the technical foundations for programmable network configuration and monitoring, network capability exposure, synergetic resource orchestration across the continuum, and zero-touch network automation. All these functionalities are supported by advanced techniques for security and trustworthiness, AI/ML algorithms, and network digital twins. For each enabler, the document illustrates the design and presents the ongoing implementation, providing early validation results and discussing the enablers’ contributions to KPIs and KVIs. The role of the enablers in Hexa-X-II PoCs is presented as well, with initial integration and evaluation results. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 3 / 207 Finally, as key contribution towards the design of the end-to-end Hexa-X-II system, the document analyses the alignment of the M&O enablers with the end-to-end system blueprint, identifying the mapping with the current components and potential gaps to be filled with new elements or interfaces. This will feed the next iteration of the end-to-end Hexa-X-II system design, under development in WP2, following the overall methodology defined in the project. Keywords Smart Network Management, Management and Orchestration, Network Programmability, Monitoring, Telemetry, Network Exposure, Security, Continuum Orchestration, Artificial Intelligence and Machine Learning, Trustworthy AI, Network Digital Twins, Zero-Touch Closed Control Loops Disclaimer Funded by the European Union. The views and opinions expressed are however those of the author(s) only and do not necessarily reflect the views of Hexa-X-II Consortium nor those of the European Union or Horizon Europe SNS JU. Neither the European Union nor the granting authority can be held responsible for them. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 4 / 207 Executive Summary The Smart Network and Services Joint Undertaking (SNS JU) 6G Flagship project Hexa-X-II leads the way to next generation 6G end-to-end (E2E) system design, based on integrated and interacting technology enablers. The project will continue the track of the previous 5G-PPP Hexa-X project [HEX24], which has laid the foundation for the global communication network of the 2030s by developing the 6G vision and basic concepts. As a continuation of Hexa-X, the interaction between the cyber-physical world and human world will be further evolved with the advancement of information, communication, and computation technologies towards a pervasive human centred cyber-physical world in 2030. To reach this 6G vision, Hexa-X-II has set up the goals to design the system blueprint of a sustainable, inclusive, and trustworthy 6G platform, which will require novel enablers regarding Smart Management & Orchestration (M&O) functionalities. In this direction, this report is the second public deliverable produced by the Hexa-X-II “Smart Network Management” Work Package, documenting the evolution in the design of the envisioned enablers and their first implementation, with preliminary evaluation results. The deliverable has the primary objective of contributing to the 2nd iteration of the Hexa-X-II end-to-end system blueprint, providing a well-structured and consolidated analysis of the M&O technical enablers defined in the project. An initial overview of the Smart Network Management framework highlights the contributions to the Hexa-X-II architecture design principles and the mapping with 6G stakeholders, explaining the main enhancements and innovations introduced in the Hexa-X-II M&O framework, especially with respect to the Hexa-X architecture. The categorization of the enablers has been updated with respect to the previous deliverable D6.2, to provide an improved and unified picture of their role and complementary contribution in the Hexa-X-II Smart Network Management system. The whole set of enablers covers, jointly, the functionalities and objectives of a 6G network Management & Orchestration (M&O) system. They provide the technical foundations for programmable network configuration and monitoring, network capability exposure, synergetic resource orchestration across a multi-domain continuum, and zero-touch network automation. All these functionalities are supported, in a transversal manner, by advanced techniques for security and trustworthiness, Artificial Intelligence (AI) and Machine Learning (ML) - AI/ML algorithms, and network digital twins. The restructuring of the M&O enablers classification has identified 8 main enablers, where some of them are organized in additional sub-enablers for specific aspects or functionalities. Moreover, their naming and terminology has been updated to highlight their applicability to concrete elements or systems of the future 6G Smart Network Management system. The updated list of proposed enablers and sub-enablers is the following: • Enabler 1: Network Programmability Framework • Enabler 2: Monitoring and Telemetry Framework • Enabler 3: Management Capabilities Exposure Framework • Enabler 4: Security and trustworthiness Framework o Sub-enabler 4.1: Resource controllability for 3rd parties o Sub-enabler 4.2: User-centric service provisioning o Sub-enabler 4.3: Trust Management System • Enabler 5: Synergetic Orchestration Mechanisms for the Computing Continuum o Sub-enabler 5.1: Multi-agent systems for multi-cluster orchestration o Sub-enabler 5.2: Decentralised orchestration system o Sub-enabler 5.3: Federated orchestration system • Enabler 6: AI/ML Algorithms o Sub-enabler 6.1: AI/ML-based control algorithms for sustainability o Sub-enabler 6.2: Trustworthy AI/ML-based control algorithms • Enabler 7: Network Digital Twins Creation Mechanisms • Enabler 8: Real-time Zero-touch Control Loops Automation and Coordination System For each enabler, the document presents: (i) the technical approach; (ii) the ongoing implementation; (iii) the early validation results, and, (iv) their impact to relevant Key Performance Indicators (KPI) and Key Value Indicators (KVI). At this stage of the activities, most of the M&O enablers have been contextualized and Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 5 / 207 integrated in software prototypes within two Hexa-X-II Proof of Concepts (PoC): PoC#A.1 and PoC#B.1. These allow to validate and demonstrate the feasibility of the conceptual ideas with concrete scenarios, welldefined workflows and testing environments. Finally, as key contribution towards the design of the end-to-end Hexa-X-II system, the document analyses the alignment of the M&O enablers with the end-to-end system blueprint, identifying the mapping with the current components and potential gaps to be filled with new elements or interfaces. This will feed the next iteration of the end-to-end Hexa-X-II system design, under development in WP2, following the overall methodology defined in the project. In this direction, the document constitutes an intermediate output towards the design and implementation of Hexa-X-II Smart Network Management system, with preliminary software prototypes, early validation results and initial integration in Hexa-X-II PoCs. The next deliverable D6.4 will further consolidate these assets, providing the final and fully validated version of the M&O system enablers. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 6 / 207 Table of Contents 1 Introduction ......................................................................................................................... 21 1.1 Objective of the document ........................................................................................... 21 1.2 Structure of the document ............................................................................................ 21 2 Smart Network Management ............................................................................................. 23 2.1 M&O technical enablers .............................................................................................. 26 2.2 Mapping with Hexa-X-II architecture design principles .............................................. 26 2.3 Mapping with the 6G stakeholders .............................................................................. 31 2.4 Innovations in Hexa-X-II M&O framework ................................................................ 34 2.4.1 Enhancements to the Hexa-X architecture .............................................................. 38 3 Enablers for Management and Orchestration ................................................................. 41 3.1 Enabler 1: Network programmability framework ........................................................ 41 3.1.1 Enabler design ......................................................................................................... 41 3.1.1.1 Internal Architecture of the system components ............................................... 42 3.1.1.2 Workflows ......................................................................................................... 43 3.1.2 Preliminary implementation and early validation results ........................................ 44 3.1.2.1 TeraFlowSDN using data-plane in-a-box .......................................................... 44 3.1.2.2 SmartNIC Transceiver support using OpenConfig extensions .......................... 45 3.1.2.3 Time Sensitive Networking and Deterministic Networking SDN controller .... 47 3.1.2.4 Automated transport network re-configuration ................................................. 49 3.1.2.5 MEC Bandwidth Management service integration with SDN controller .......... 51 3.1.2.6 Integration of TM Forum APIs with ETSI TFS NBI ......................................... 53 3.1.3 Impacted KPIs and KVIs ........................................................................................ 55 3.2 Enabler 2: Monitoring and telemetry framework ......................................................... 55 3.2.1 Enabler design ......................................................................................................... 56 3.2.1.1 System components ........................................................................................... 56 3.2.1.2 Workflows ......................................................................................................... 57 3.2.2 Preliminary implementation and early validation results ........................................ 58 3.2.2.1 TeraFlowSDN event-driven monitoring ............................................................ 58 3.2.2.2 Energy monitoring ............................................................................................. 60 3.2.2.3 Monitoring platform for integration in closed loop ........................................... 62 3.2.3 Conceptual solutions ............................................................................................... 64 3.2.3.1 Passive/in-band and active telemetry for TSN/DetNet networks ...................... 64 3.2.3.2 Data fusion for signals correlation and remediation actions .............................. 65 3.2.4 Impacted KPIs and KVIs ........................................................................................ 66 3.3 Enabler 3: Management capabilities exposure framework .......................................... 67 3.3.1 Enabler design ......................................................................................................... 67 3.3.1.1 System components ........................................................................................... 68 3.3.1.2 Internal Architecture of the system components ............................................... 69 3.3.1.3 Workflows ......................................................................................................... 72 3.3.2 Preliminary implementation and early validation results ........................................ 73 3.3.2.1 Apache Kafka cluster structure .......................................................................... 73 3.3.3 Impacted KPIs and KVIs ........................................................................................ 74 3.4 Enabler 4: Security and trustworthiness framework .................................................... 75 3.4.1 Sub-enabler 4.1: 3rd party resource control separation enabler ............................... 76 3.4.1.1 Sub-enabler design ............................................................................................. 76 3.4.1.2 Impacted KPIs and KVIs ................................................................................... 82 3.4.2 Sub-enabler 4.2: User-centric service provisioning system .................................... 82 3.4.2.1 Sub-enabler design ............................................................................................. 82 3.4.2.2 Impacted KPIs and KVIs ................................................................................... 86 3.4.3 Sub-enabler 4.3: Trust Management System .......................................................... 87 3.4.3.1 Sub-enabler design ............................................................................................. 87 Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 7 / 207 3.4.3.2 Preliminary implementation and early validation results .................................. 88 3.4.3.3 Impacted KPIs and KVIs ................................................................................... 89 3.5 Enabler 5: Synergetic orchestration mechanisms for the computing continuum ......... 89 3.5.1 Sub-enabler 5.1: Multi-agent systems for multi-cluster orchestration .................... 90 3.5.1.1 Sub-enablers design ........................................................................................... 90 3.5.1.2 Preliminary implementation and early validation results .................................. 97 3.5.1.3 Impacted KPIs and KVIs ................................................................................. 102 3.5.2 Sub-enabler 5.2: Decentralised orchestration system ........................................... 102 3.5.2.1 Sub-enabler design ........................................................................................... 103 3.5.2.2 Preliminary Implementation and early validation results ................................ 110 3.5.2.3 Impacted KPIs and KVIs ................................................................................. 112 3.5.3 Sub-enabler 5.3: Federated orchestration system .................................................. 113 3.5.3.1 Enabler Design ................................................................................................. 113 3.5.3.2 Preliminary implementation and early validation results ................................ 115 3.5.3.3 Impacted KPIs and KVIs ................................................................................. 116 3.6 Enabler 6: AI/ML algorithms ..................................................................................... 116 3.6.1 Sub-enabler 6.1: AI/ML-based control algorithms for sustainability ................... 117 3.6.1.1 Sub-enabler design ........................................................................................... 117 3.6.1.2 Preliminary implementation and early validation results ................................ 118 3.6.1.3 Impacted KPIs and KVIs ................................................................................. 133 3.6.2 Sub-enabler 6.2: Trustworthy AI/ML-based control algorithms .......................... 133 3.6.2.1 Sub-enabler design ........................................................................................... 133 3.6.2.2 Preliminary implementation and early validation results ................................ 134 3.6.2.3 Conceptual solutions ........................................................................................ 137 3.6.2.4 Impacted KPIs and KVIs ................................................................................. 140 3.7 Enabler 7: Network digital twins creation mechanisms ............................................. 140 3.7.1 Enabler design ....................................................................................................... 140 3.7.1.1 System Components ........................................................................................ 140 3.7.2 Preliminary implementation and early validation results ...................................... 141 3.7.2.1 Applicable datasets / environment ................................................................... 142 3.7.2.2 Modelling virtualization effects ....................................................................... 142 3.7.3 Impacted KPIs and KVIs ...................................................................................... 143 3.8 Enabler 8: Real-time Zero-touch control loops automation and coordination system144 3.8.1 Enabler design ....................................................................................................... 144 3.8.1.1 Internal Architecture of the system components ............................................. 145 3.8.1.2 Workflows ....................................................................................................... 150 3.8.2 Preliminary implementation and early validation results ...................................... 153 3.8.2.1 Closed Loop Governance and Coordination Functions ................................... 154 3.8.2.2 Specialized CLs: Automation of Transport Network Slices ............................ 156 3.8.2.3 Penalty-based management of concurrent service CLs ................................... 158 3.8.3 Conceptual solutions ............................................................................................. 161 3.8.3.1 Specialized CLs: TSN/DetNet control ............................................................. 161 3.8.3.2 Specialized CLs: Dependencies in AI/ML models and NDTs ........................ 162 3.8.4 Relation with other enablers ................................................................................. 163 3.8.4.1 Specialized CLs for the decentralized orchestration system ............................ 163 3.8.4.2 Specialized CLs: Service autoscaling in the continuum .................................. 165 3.8.5 Impacted KPIs and KVIs ...................................................................................... 166 4 Alignment of M&O enablers with the E2E system blueprint ....................................... 167 4.1 Enabler 1: Network programmability framework ...................................................... 167 4.2 Enabler 2: Monitoring and telemetry framework ....................................................... 168 4.3 Enabler 3: Management capabilities exposure framework ........................................ 169 4.4 Enabler 4: Security and trustworthiness framework .................................................. 170 4.5 Enabler 5: Synergetic orchestration mechanisms for the computing continuum ....... 171 4.5.1 Sub-enabler 5.1: Multi-agent systems for multi-cluster orchestration .................. 171 Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 8 / 207 4.5.2 Sub-enabler 5.2: Decentralised Orchestration System .......................................... 172 4.5.3 Sub-enabler 5.3: Federated Orchestration System ................................................ 173 4.6 Enabler 6: AI/ML Algorithms .................................................................................... 173 4.6.1 Sub-enabler 6.1: AI/ML-based control algorithms for sustainability ................... 173 4.6.2 Sub-enabler 6.2: Trustworthy AI/ML-based control algorithms .......................... 174 4.7 Enabler 7: Network digital twins creation mechanisms ............................................. 175 4.8 Enabler 8: Real-time zero-touch control loops automation and coordination system.175 5 Contributions to dissemination activities ........................................................................ 177 6 Conclusions ........................................................................................................................ 179 7 References .......................................................................................................................... 180 8 Appendix ............................................................................................................................ 188 8.1 Information models .................................................................................................... 188 8.1.1 Information model for Trustworthy 3P ................................................................. 188 8.1.2 Information models for Resource Orchestration in the continuum ....................... 190 8.1.3 Information model for Closed Loop Descriptor.................................................... 194 8.2 Interfaces .................................................................................................................... 197 8.2.1 Interfaces of SDN Orchestrator and SDN controllers ........................................... 197 8.2.1.1 E2E SDN Orchestrator ..................................................................................... 197 8.2.1.2 Technological-Domain SDN controller ........................................................... 198 8.2.2 Interfaces of the Monitoring and Telemetry framework ....................................... 198 8.2.3 Interfaces of the Management Capabilities Exposure framework ........................ 199 8.2.4 Interfaces of CL Governance and CL Coordination functions ............................. 200 8.3 Overview of TM Forum, IETF, and ETSI Framework Alignments .......................... 202 8.3.1 NBI TMForum ...................................................................................................... 202 8.3.2 Perspective on TM Forum APIs and IETF Standards ........................................... 202 8.3.2.1 TM Forum - Interoperability and Standardization ........................................... 202 8.3.2.2 TM Forum alignment with IETF Standards ..................................................... 203 8.3.2.3 Intent-Based Management and Autonomous Networks .................................. 203 8.4 Related technologies .................................................................................................. 206 8.4.1 The compute continuum and the convergence with the edge ............................... 206 Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 9 / 207 List of Tables Table 2-1: WP6 Enablers vs. Hexa-X-II Design Principles. ........................................................................... 28 Table 3-1: Management Capabilities Exposure Framework components. ...................................................... 68 Table 3-2: Management capabilities exposure framework interface mapping. ............................................... 71 Table 3-3: AuthN/AuthZ – protocols comparison. .......................................................................................... 80 Table 3-4: MSE values of the original model after applying adversarial attacks. ......................................... 137 Table 3-5: Comparison of the MSE values of the original and robust model after applying FGSM attack. . 137 Table 3-6: Comparison of the MSE values of the original and robust model after applying BIM attack. .... 137 Table 3-7: CL components. ........................................................................................................................... 145 Table 3-8: CL Coordination internal components. ........................................................................................ 149 Table 3-9: Mapping between software components and CL functions ......................................................... 159 Table 4-1: Alignment of the enablers with E2E system blueprint. ................................................................ 167 Table 5-1: Contributions to papers. ............................................................................................................... 177 Table 5-2: Contributions to demonstrations. ................................................................................................. 178 Table 8-1: “Identity” class attributes. ............................................................................................................ 188 Table 8-2: “Role” class attributes. ................................................................................................................. 188 Table 8-3: “AccessRule” class attributes. ...................................................................................................... 189 Table 8-4: “Permission” dataType attributes. ................................................................................................ 189 Table 8-5: E2E SDN Orchestrator exposed interfaces. ................................................................................. 197 Table 8-6: Technological-Domain SDN controller exposed interfaces. ........................................................ 198 Table 8-7: Subscription API. ......................................................................................................................... 199 Table 8-8: Listing API. .................................................................................................................................. 200 Table 8-9: Security API. ................................................................................................................................ 200 Table 8-10: CL Governance and Coordination interfaces. ............................................................................ 201 Table 8-11: Intent Standards Classification. .................................................................................................. 206 List of Figures Figure 2-1: WP6 and its relationship with other WPs. .................................................................................... 23 Figure 2-2: Extension of E2E System Blueprint provided from WP2. ........................................................... 24 Figure 2-3: WP6 enablers re-definition. .......................................................................................................... 26 Figure 2-4: Roles in 5G provisioning systems [5GP21]. ................................................................................. 32 Figure 2-5: Roles in 6G ecosystem.................................................................................................................. 33 Figure 3-1 Design of Enabler 1: Network programmability framework. ........................................................ 41 Figure 3-2: ETSI TeraFlowSDN internal architecture. ................................................................................... 43 Figure 3-3 Network programmability framework end-to-end workflow for provisioning IETF slice. ........... 44 Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 16 / 207 IdP Identity Provider IETF Internet Engineering Task Force ILE Infrastructure Layer Emulator IMF Intent Management Function IMS IP Multimedia Subsystem IoT Internet of Things IP Internet Protocol IRN Infrastructure Registry Node IRS Infrastructure Registry Service IRTF Internet Research Task Force ISG Industry Specification Groups ISPS Infrastructure Status Prediction Service ISPN Infrastructure Status Prediction Node ITU International Telecommunication Union JSON JavaScript Object Notation K8s Kubernetes KPI Key Performance Indicator KVI Key Value Indicator L2VPN Layer 2 Virtual Private Network LCM Life Cycle Management LDAP Lightweight Directory Access Protocol LoTAF Level of Trust Assessment Function LSTM Long Short-Term Memory LTS Long Term Support M&O Management and Orchestration MARL Multi-Agent Reinforcement Learning MDAF Management Data Analytics Function MEC Multi-access Edge Computing MIB Management Information Base MIP Mixed Integer Programming ML Machine Learning MNO Mobile Network Operator MQTT Message Queuing Telemetry Transport MSE Mean Squared Error Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 17 / 207 NBI NorthBound Interface NB-IoT NarrowBand IoT NDT Network Digital Twin NEF Network Exposure Function NF Network Function NFV Network Functions Virtualization NFV-NS NFV Network Service NFVI Network Function Virtualization Infrastructure NFVO Network Function Virtualization Orchestrator NIC Network Interface Card NM Network Modelling NMI Network Management Interface NOS Network Operating System NPN Non-Public Network NRM Network Resource Model NS Network Service NSaaS Network Slice as a Service NSM Network Service Mesh NSMF Network Slice Management Function NSSMF Network Slice Subnet Management Function NSSAI Network Slice Selection Assistance Information NTN Non Terrestrial Network NWDAF Network Data Analytics Function OAM Operations, Administration, and Maintenance OIDC OpenID Connect ONF Open Networking Foundation O-RAN Open Radio Access Network OS Operating System OSM Open Source MANO OSS Operations Support System OTLP OpenTelemetry Protocol OTN Optical Transport Network OTT Over the Top PCF Policy Control Function Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 18 / 207 PDL Policy Description Language PDU Packet Data Unit PM Performance Management PMF Privacy Management Function PoC Proof of Concept POF Privacy Operation Function QoE Quality of Experience QoS Quality of Service QT Quantitative Target RAM Radio Access Memory RAN Radio Access Network RAW Reliable and Available Wireless RBAC Role-Based Access Control REC-EXEC REsource orchestrator for Continuum across EXtreme-edge, Edge and Cloud REST Representational state transfer RF Radio Frequency RFC Request for Comments RFS Resource Facing Service RL Reinforcement Learning RNN Recurrent Neural Network ROADM Reconfigurable Optical Add-Drop Multiplexers RPC Remote Procedure Call SAI Securing Artificial Intelligence SBA Service Based Architecture SBI SouthBound Interface SBMA Service Based Management Architecture SCI Security Control Interface SDN Software Defined Network SDO Standard Development Organizations SLA Service Level Agreement SLO Service Level Objective SMF Session Management Function SMO Service Management and Orchestration SNMP Simple Network Management Protocol Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 19 / 207 SNS-JU Smart Networks and Services Joint Undertaking SON Self-Organizing Network SoTA State of the Art SP Service Provider SQL Structured Query Language SRN Service Registry Node SRS Services Registry Service SW Software SWV Software Vendor TAPI Transport API TAS Time-Aware Shaping TE Traffic Engineering TEAS Traffic Engineering Architecture and Signalling TFS TeraFlowSDN TMF TM Forum TLA Trust Level Agreement TN Transport Network TLS Transport Layer Security TSN Time-Sensitive Networking UAV Unmanned Aerial Vehicle UDM Unified Data Management UE User Equipment UI User Interface UL Uplink UML Unified Modelling Language UPF User Plane Function URLLC Ultra-Reliable Low Latency Communication URSP UE Route Selection Policy VIM Virtual Infrastructure Manager VISP Virtual Infrastructure Service Provider VM Virtual Machine VMAF Virtualized MEC application function VNF Virtual Network Function VPN Virtual Private Network Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 20 / 207 WAN Wide Area Network WP Work Package XR eXtended Reality XRL eXplainable Reinforcement Learning ZSM Zero touch network and Service Management Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 21 / 207 1 Introduction 1.1 Objective of the document This document aims to describe the design, initial implementation and early validation results of the enablers for the Hexa-X-II Smart Network Management framework. The global objective is to provide feedback for the design of the Hexa-X-II End-to-End (E2E) 6G System Blueprint, in its 2nd iteration. The enablers presented in this document constitute an evolution of the initial set identified in D6.2 [HEX223-D62], proposing a restructuring of their classification and introducing their implementation and contributions to Hexa-X-II Proof of Concepts (PoC) implementation. This deliverable can be considered as an intermediate milestone in the overall process of defining the Management and Orchestration (M&O) system for the future 6G networks, providing a preliminary version of M&O enablers’ design and implementation. However, it plays a crucial role in the methodology planned in the project for the definition of the Hexa-X-II E2E 6G System Blueprint. This methodology is based on an inter-work-package collaboration structured in multiple cycles of (i) initial bottom-up contributions from the technical Work Packages (WP) focused on specific aspects of the 6G system architecture, which are then (ii) consolidated in progressive iterations of the end-to-end blueprint. This is in turn (iii) analysed by the technical WPs in order to provide feedback and contribute to the refinement of the next iteration of the E2E system blueprint. In this process, this deliverable will provide feedback to an intermediate iteration of the Hexa-X-II E2E 6G System Blueprint contributing to its 2nd official iteration (see section 2 for further details). The perspective provided in this deliverable is particularly relevant because it captures lessons learnt from early implementation and validation activities, going well beyond the initial conceptual design provided in D6.2 [HEX223-D62]. The initial integration of M&O enablers in Hexa-X-II Proof-of-Concepts (PoC) allows to consolidate the design of the enablers in concrete frameworks and systems, offering an opportunity to evaluate and compare different architectural choices in real scenarios. In this direction and following an evolutionary approach, the enablers initially identified in D6.2 [HEX223D62] have been re-organized in a short list of 8 enablers, with some of them further structured in sub-enablers. For each of them, the document presents: (i) the technical approach; (ii) the ongoing implementation; (iii) the early validation results, and (iv) their impact to relevant Key Performance Indicators (KPI) and Key Value Indicators (KVI), showing as well how they have been integrated in the project PoCs. Feedback to the E2E system blueprint design is presented at the end of the document following on a per-enabler approach and collecting the results of the latest interactions. 1.2 Structure of the document The document is structured as follows: • Section 1 introduces the document with an overview of its content and structure, defining the main objectives of the deliverable together with its position and main contribution to the Hexa-X-II project. • Section 2 provides a global view of the Smart Network Management framework presenting the refined list of M&O technical enablers and their classification. This section analyses the alignment and the contribution of the framework to the design principles of Hexa-X-II architecture, as well as the mapping of the proposed system with the 6G stakeholders. Finally, it highlights the main innovations introduced in the Hexa-X-II M&O framework, with particular reference to the enhancements with regards to the original Hexa-X architecture. • Section 3 is focused on the M&O enablers and sub-enablers, constituting the core of the document. For each enabler, this section reports the current status of design, implementation and evaluation results, discussing the contributions to KPIs and KVIs. It should be noted that each enabler may have several implementations, e.g., focusing on particular sub-components, examples of algorithms or customization of workflows and configurations for a given target scenario or objective. For some enablers, conceptual solutions not yet at the implementation level are also reported and, where applicable, the relation with other technologies or other enablers are documented. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 22 / 207 • Section 4 discusses the alignment of M&O enablers with the ongoing version of the E2E system blueprint used as reference for feedback on smart network management aspect. For each enabler, this section analyses the mapping with layers, components and interfaces of the E2E system blueprints, identifying possible gaps and requirements for architectural updates, refinements or even new components and interfaces. This is the main outcome of the document in what regards its contribution to the end-to-end Hexa-X-II system design. • Section 5 reports the contributions to dissemination and demonstrations. • Section 6 provides the conclusions and a roadmap towards the finalization of the Smart Network Management framework. • The Appendixes provide additional details on information models and interfaces. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 23 / 207 2 Smart Network Management The Hexa-X-II project approaches the overall design of the 6G E2E System in the context of its WP2, which is intended to harmonize the E2E design principles towards producing the 6G SNS platform blueprint. In this context, WP6 (to which this deliverable belongs) receives inputs from WP2 regarding the overall platform design and, combining it with the information from other work packages (WP1 describing the use cases, and WP3 addressing the 6G Architecture design), provides outputs to WP2 regarding the specific WP6 smart network management topics to be integrated as part of the end-to-end system. Figure 2-1 outlines this information flow between WP6 and the other related WPs mentioned here (interactions between WP2 and WP3 are not shown in the picture for simplicity). Figure 2-1: WP6 and its relationship with other WPs. To address this flow of information towards WP2, WP6 has defined a framework considering: (i) the contributions to the Hexa-X-II architecture design principles, (ii) the mapping with the envisaged 6G stakeholders, (iii) the definition and the alignment of the WP6 M&O technical enablers (already introduced in the previous Deliverable D6.2 [HEX223-D62]) with the initial blueprint provided from WP2 (through the flow ① in Figure 2-1), and (iv), the envisaged contributions of those M&O enablers towards the future 6G smart networks (flow ② in the figure). This Section contextualizes the M&O framework in the E2E system blueprint, initially introducing the M&O technical enablers (Subsection 2.1) and discussing their contribution to the architecture design principles (Subsection 2.2). The 6G stakeholders, as defined in WP2, are briefly reported in Subsection 2.3 highlighting their interaction with the M&O framework. Besides, Subsection 2.4 introduces the main innovations envisaged in the Hexa-X-II M&O framework considering the Hexa-X M&O architecture [HEX22-D62] as baseline. Figure 2-2 depicts the current version (at the time this document is being edited) of the E2E System blueprint provided by WP2. It represents the main building blocks for each Mobile Network Operator (MNO) that would aim to align its network design with the future envisaged 6G E2E System. The design is basically a four-layer stack (left) consisting of an Infrastructure layer (bottom), a Network Functions layer (middle), a Networkcentric Application Layer (top), and an Application Layer (on top of the previous one). Besides these four layers there is also a so-called Pervasive Functionalities block (right). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 24 / 207 Figure 2-2: Extension of E2E System Blueprint provided from WP2. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 25 / 207 - Networks Functions Layer: consists of User Equipment (UE) and subnetwork functions (concerning the UEs protocol stack and their associated subnetworks), Radio Access Network (RAN) functions, Core Network (CN) Functions, Transport Network functions, Beyond communication functions (concerning functions beyond the communications topic itself, e.g., novel Radio Frequency (RF)- based sensing, Artificial Intelligence (AI)-related, 6G positioning and mapping, etc.), and the Data network (as defined in the 3GPP architecture, targeting those functionalities related to services provided by the operator, the Internet access, or 3rd party services). - Network-centric Application layer: consists of network-centric applications, i.e., applications that are more centred on the network capabilities, possibly depending on network features and that can interact with the control plane of the Network Functions Layer to provide enriched services to vertical applications. Moreover, these network-centric applications can provide certain actions/updates to the network functions and thereby provide quick ways to modify the network functionality, enabling additional operation and optimization of the network (e.g., rApp in O-RAN). - Application layer: represents the applications that may have certain service expectations from the network itself. These expectations may depend on the application requirements, such as bandwidth, latency, etc. The Network functions layer and the Network-centric applications layer expose network information and network control Application Programming Interfaces (API) through the interface/exposure towards this application layer. This enables application developers to specify their requirements to the network as well as to adjust the expectations of the applications based on network conditions. Operations APIs for the service M&O can also be exposed to the application ecosystem actors (e.g., to vertical providers) through developer-friendly intent-based abstractions and humanoriented intent-based APIs. - Pervasive Functionalities block: refers to those functionalities and frameworks that operate at the various layers of the 6G E2E system blueprint, i.e., at the infrastructure, network function, and application levels, and can interact with all the different layers, as well. The M&O functionalities are part of the Pervasive Functionalities block, represented by the “M&O” block on the right-hand side of the diagram. Besides, there are also other additional “Management Functions” within the “Networks Functions Layer” (middle). For the pervasive M&O block [HEX223-D21] provides the following definition: “The management and orchestration functionality will cope with novel and more complex services combining both network capabilities, and beyond communications capabilities such as sensing and computing. Smart network management will provide a uniform orchestration across a continuum of resources from extreme edge to edge to central clouds. Automation functionality will provide more and more degrees towards autonomy with fully automated closed-loop control, supported by intent-based management and AI/ML techniques. It will also comprise the Continuous Integration/Delivery (CI/CD) intrinsically associated to the development of software components”. Beyond this, in the most updated version of the blueprint (the one in Figure 2-2), this pervasive M&O block has been enriched with the following internal elements: - A “Service Layer M&O” block, intended to provide specific M&O functionalities for the so-called Network-centric Application Layer through a set of Service APIs (“Service APIs” blue arrow). - A “Network Layer M&O” block, with a functionality similar to the previous block, but with respect to the Network Functions Layer. - An “Infrastructure Layer M&O” block, with a similar role to the two previous blocks, but focusing on the Infrastructure Layer. - A “3rd Party Trust Management” block (top), intended to offer multi-tenancy support in resource sharing environments, in terms of (i) resource controllability separation (providing tenants with segregated yet customized management spaces), (ii) user-centric network management (defining policies for tenant subscribers in relation to service delivery and consumption), and (iii) Service Level Agreement (SLA) enforcement, assurance and verifiability. - An “Intent Digital Service Manager (DSM)” block, associated to the Service Layer M&O block, intended to allow tenants (with or without technical knowledge) to interact with the M&O system and to request a desired service without specifying how it needs to be deployed, leaving the system to find Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 32 / 207 Figure 2-4: Roles in 5G provisioning systems [5GP21]. c) A new role belonging to the Hardware and Software Suppliers category (13) in the previous Figure 2-4, designated as AI/ML Software Provider [HEX22-D62]. d) The new 6G specific roles mentioned in [HEX223-D22], namely 2 : a. Capability Operator (COP), as a decomposition of the already existing Network Operator role that would oversee the E2E service management within the operator administrative domain, integrating the different available capabilities (e.g., Network, Cloud, AI, applications, etc.). COPs are envisaged to operate those entities/elements in charge of managing the available domain resources, as well as handling the control and management elements, such as a Network Slice Management Function (NSMF), or other domain specific managers. b. Resource Provider, another decomposition of the same Network Operator role that would offer different domain resources, such as applications, AI, network (including RAN, transport, and core), and cloud resources (extreme-edge, edge, and cloud). c. TechCos [STL23], considered a step forward with respect to regular Tier-1 telco operators (referred as just ‘Telcos’ in [HEX223-D22]), and having a wider scope on service offerings, as well as further flexibility on the composition and operation of managed resources. TechCos are expected to articulate their systems with much more granularity, breaking functions into its fundamental components (microservices), each representing stand-alone capabilities that can be individually programmed and chained on a per service/use case needs. Also, their service offering is envisaged not to be limited to communication services, but to other digital services (e.g., Web3, big data, or security services) and beyond communication services (e.g., AI, or applications) offered to new customers (e.g., aggregators) through APIs, in line with the DSP (Digital Service Provider) stakeholder definition already in [5GP21]. However, in the new 6G context, the DSP is also envisaged as the stakeholder in charge to receive intent-based requests from the Service Customers and manage them to achieve the right translation into a set of specific requests for a selected set of the available Capability Operators. d. In line with the TechCos approach, Service Customers already in [5GP21] are renamed as Digital Service Customers (i.e., tenants), which in turn are divided into three categories: i. Aggregator Tenants, which would include actors such as Hyperscalers, Marketplaces, Telco Consortium, etc., following the Business-to-Business-to-X (B2B2X) model and offering intent-based services to other tenants. 2 [HEX223-D22] also refers some of the roles already in Figure 2-5, namely: Network Operator (NOP), Communication Service Provider (CSP), and Digital Service Provider (DSP). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 33 / 207 ii. Business-to-Business (B2B) Verticals. This type of tenants would group either vertical industries following a B2B model (e.g., Extended Reality (XR) providers) or Application Service Providers following a Business to Business to Consumer (B2B2C) model. iii. Other Verticals having direct access to the offered intent-based services. Considering all this, the roles in the 5G system presented in Figure 2-4 would be updated to include the new roles envisaged towards 6G, as shown in Figure 2-5. Figure 2-5: Roles in 6G ecosystem. Besides aggregating the different roles mentioned above, the main update in this figure compared to Figure 2-4 is the deletion of the “boundary” line separating the Service Customers in the Vertical Ecosystem scope from the Provisioning Ecosystem. This is intentional, trying to illustrate the “continuum” orchestration concept towards 6G. The network operator (designated as TechCo now) is seen now as part of the network continuum, just like the other stakeholders in the ecosystem. The idea behind this is those vertical industries, or even endusers, will no longer be mere network service consumers, but also able to perform some of the roles that in previous generations were exclusive to the network operator, even including the deployment and the operation of the network services. Besides, those network services can be also decomposed to be deployed on heterogeneous network resources made available by different infrastructure providers, spanning on different technical and administrative domains across the network continuum (core, edge, and extreme-edge), and using cloud-native techniques for this purpose. Below some examples considering possible interactions among different stakeholders based on this approach: a) A Software Supplier could automatically deploy (using cloud-native Continuous Integration and Continuous Deployment (CI/CD) techniques) a network service (or a subset of network service components) on the infrastructure resources (e.g., in a specific staging environment) provided by a TechCo and/or other Infrastructure Providers. The Software Supplier could also monitor certain parameters once the service is deployed, to automatically trigger its re-deployment (e.g., based on certain monitoring parameters) or to perform other service life cycle management actions (e.g., the service re-configuration). b) A Software Supplier could request data to an AI/ML Data Supplier and to a TechCo, in order to train an AI/ML model that would be part of a network service to be deployed on the infrastructure belonging to that TechCo. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 34 / 207 c) A TechCo could deploy a network service consisting of multiple microservices on infrastructure resources belonging to different network domains: on its own domain (e.g., on its core network), on certain infrastructure resources provided by a vertical industry (acting as both, Infrastructure Provider and Digital Service Customer), and other additional infrastructure resources at the extreme-edge belonging to endusers (also acting as Infrastructure Providers). d) A Vertical Industry could act not only as a Digital Service Consumer, but also as Software Provider: certain microservices deployed on its own domain could be part of a network service that, as a whole, would be deployed through the network using the resources of different Infrastructure Providers, and that would be operated by the TechCo, or by a third-party Operation Support Provider. This use case is considered especially relevant to integrate the vertical industries with the operators, without this meaning the usage of specific software provided by the operator. e) A hyperscaler could act as both Service Provider and Infrastructure Provider, and could request the deployment of RAN service components on a TechCo that would act as Infrastructure Provider as well. f) Complementary to the previous, a TechCo could request the deployment of certain AI/ML service components on the infrastructure provided by a hyperscaler (e.g., to enrich a network service with certain image recognition or natural language processing capabilities). The examples list could continue. However, the key idea here is that, as can be appreciated, contrary to what occurred in previous generations, the role-stakeholder relationship is not so strictly defined. Taking advantage of the cloud-native technology, network services can be decomposed and deployed throughout the entire network continuum using agile DevOps techniques, enabling each stakeholder (including the operator) to perform different roles, both in terms of the network services provisioning and their execution. 2.4 Innovations in Hexa-X-II M&O framework The main innovations envisaged for the Hexa-X-II M&O framework are those derived from the WP6 enablers presented in the previous Section 2.1. Below are the main innovations associated with each of these enablers. Enabler #1 Network programmability framework. Broadly speaking, this enabler is about including Software Defined Networking (SDN) technology in the 6G architecture. Although the SDN technology itself is not something new (it is a key enabler already for the 5G technology) it is considered that SDN should continue being part of the next generation of the mobile technology as well, in order to make it possible to interact with the network infrastructure in a programmable and flexible manner, and also based on well-defined standards. It is in this regards that R17 3GPP documents consider Transport Network as part of xHaul network and a dedicated sub-slice manager is introduced for the Transport Network. Thus, some innovations associated with this SDN technology are also considered towards 6G. The main ones are: • Integration of Transport Networks in 6G through Transport Network - Network Slice Subnet Management Functions (TN-NSSMFs). • Its alignment with the cloud-native approach, in line with the overall architectural design approach already posed in the previous Hexa-X project. • The design of interfaces for new devices, in order to support network disaggregation, where network elements such as routers are divided between whiteboxes and network operating systems (NOS). New support for control and management of multiple NOS through OpenConfig or gNMI (gRPC Network Management Interface) protocols. • To align the SDN architecture with the cloud continuum concept, starting with ease of integration of transport network in Multi-Access Edge Computing (MEC), considered as an example of edge architecture. Enabler #2: Monitoring and telemetry framework. The goal of this enabler is to develop an innovative Monitoring and telemetry framework architecture for future 6G networks that is both scalable and driven by data and events. This architecture focuses on advanced monitoring and telemetry systems and also includes features for tracking energy usage. Additionally, it supports the integration of diverse data sources, facilitating data fusion. This approach aims to enhance the Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 35 / 207 efficiency and effectiveness of network operations, enabling more sophisticated management and optimization of resources in 6G environments. While Network Data Analytics Function (NWDAF) in 5G focuses on network-specific data analytics, the proposed 6G enabler expands this concept by introducing the following innovations: • Enhancing scalability • Diverse data types integration (including energy usage data and extreme edge monitoring and telemetry). • Data and event driven architecture to provide instant access to alarms and event alerts. • Ease of integration with closed loop mechanisms for network automation. This evolution reflects a shift towards more comprehensive network management tools that can adapt to the increasing complexity and demands of future telecommunications networks. The move towards an eventdriven architecture in 6G suggests a response to real-time changes and conditions, which is critical for managing the dynamic environments expected in next-generation networks. Enabler #3: Management capabilities exposure framework. This enabler proposes an implementation of the Integration Fabric inspired by the ETSI ZSM specification [ZSM-002]. It basically provides two main innovations: • Dynamic registration and discovery of network elements or services, following a plug-and-play concept 3 . This innovation streamlines the operational aspects by automating the process of registering and discovering new network elements or services within the system. By adopting a plug-and-play approach, it reduces the need for manual intervention, thus minimizing operational overhead and allowing a quick integration of new services and functionalities. This not only enhances operational efficiency but also facilitates rapid service innovation, a crucial aspect in the dynamic landscape of 6G deployments. • The potential of being implemented in cloud-native environments, especially in scenarios involving several stakeholders. This feature considers the changing operational requirements of 6G networks, particularly in situations where a variety of stakeholders operate together. The enabler provides interoperability with cloud-native architectures, which guarantees resource efficiency, scalability, and flexibility in operations. It supports dynamic workload orchestration and resource allocation, giving operators the ability to manage resources across distributed settings with efficiency. This feature is necessary to maximize resource utilization and optimize operating expenses, which will improve the overall operational performance. This enabler was present in the Hexa-X architecture [HEX22-D62]. Its tasks included handling exposure blocks and API administration. Hexa-X-II comprises a prototype implementation and extends the use of the previously proven concepts to multi-site settings. Enabler #4: Security and trustworthiness framework A security and trustworthiness framework is essential for the successful deployment and operation of 6G technology. It will ensure the protection of data, maintain the integrity of communications, support regulatory compliance, and foster trust among users and stakeholders. This enabler is split into 3 sub-enablers, namely: • Sub-enabler #4.1: Third-party resource control separation enabler. This sub-enabler is intended to define the scope and impact of the “tenancy” concept in an M&O system. The proposed innovation goes beyond the static, manually configured Role-Based Access (RBAC) and assisted with Lightweight Directory Access Protocol (LDAP) solutions that are being used. This sub-enabler deals with the definition of a granular access control solution that targets authentication, authorization and auditability, ensuring that tenants can be provisioned with tailored management spaces where the permissions that characterize the readable/writable attributes are built in the model of the managed resources and services, avoiding conflict between tenants on resource sharing environments. 3 The term 'plug and play' is used to describe an architectural approach where software components or modules can be added, removed, or replaced within a system without requiring significant modifications to the existing codebase [WAC08]. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 36 / 207 • Sub-enabler #4.2: User-centric service provisioning system. This sub-enabler covers the tenant subscribers provision based on a personalized service definition according to their preferences. This sub-enabler will mean innovations in service fulfilment and service activation stages (with the definition of customized policies based on UE Route Selection for accessing subscribed services and keeping protection of user’s data), and also in service assurance stage (enriching Service Level Agreement (SLA) with trustworthiness KVIs and exploring how closed loop automation solutions can be reused and adapted to fulfil SLA assurance activities). • Sub-enabler #4.3: Trust management system. This sub-enabler proposes a trustworthiness estimation procedure of compute nodes and infrastructure components through trust evaluation function. It also aims to estimate the trustworthiness level of communication paths that hold sensitive data workloads as they transit the network. Cloud orchestration engines use the output of the trust evaluation functions, among other factors, to allocate workloads into compute nodes, targeting maximum trustworthiness and efficiency. Enabler #5: Synergetic orchestration mechanisms for the computing continuum The term computing continuum refers to resources that span different parts of the programmable infrastructure from Internet of Things (IoT) devices to extreme edge devices to edge and cloud computing infrastructure [DPD23]. The term orchestration mechanisms for the computing continuum refers to solutions for management of the deployment and operation of service graphs across various resources in the continuum [DPD23] [PDM23]. Solutions can be centralized or federated, depending on the applied orchestration scheme. Various types of synergies are considered, either between agents managing various parts of the continuum, or between different stakeholders within a 6G ecosystem (e.g., network providers, cloud providers, OTT players). This enabler is split into 3 sub-enablers, each of them targeting the orchestration mechanisms for the computing continuum from a different perspective. Below the innovations associated to each one: • Sub-enabler #5.1: Multi-agent systems for multi-cluster orchestration. This sub-enabler focuses on the proper abstraction and management of multi-cluster resources that may span the computing continuum. With the term multi-cluster resources, we refer to compute resources that can host virtualized network and application functions in the various parts of the continuum. In a 6G E2E system, the deployment of distributed services and the provision of distributed applications in multiple clusters is envisaged to better satisfy QoS, privacy and security needs. Integration of IoT and edge computing technologies is examined, where computational resources may be offered by IoT/edge devices, while end-to-end network management has to be supported. The sub-enabler also introduces mechanisms that follow a “system of systems” approach in orchestration, considering hierarchical decision-making principles. With the term “system of systems” we refer to an approach where the overall responsibility is assigned to a centralized entity, while the control is distributed across various entities [NLF15]. The objective is to introduce distributed intelligence and autonomy characteristics and enable the optimal management of virtualized compute resources in cases where the deployments are made over multi-cluster environments. The combination of techniques coming from multi-agent systems and ML is envisaged for the development of synergetic orchestration mechanisms. For instance, multiple agents may collaborate for managing autoscaling of functions for an overall application/service graph, taking advantage of Reinforcement Learning techniques. Synergies may be also applied between different stakeholders in a 6G ecosystem. For instance, interaction between network providers and over the top (OTT) players such as Service Consumers and Digital Service Providers. • Sub-enabler #5.2: Decentralised orchestration system. This sub-enabler considers the M&O mechanisms for the computing continuum, focusing mainly on integrating the extreme-edge domain. The targeting of this extreme-edge domain was already introduced in the previous Hexa-X [HEX22-D62] project, meaning that domain beyond the MNO own domain, and including even the end-user devices. The objective is to integrate that domain as part of the network continuum, deploying network service components on it, which is considered quite challenging, since the infrastructure resources in that domain can be highly heterogeneous, they can be asynchronously connected or disconnected (they are not in well controlled premises), they can be mobile devices, or devices with limited computing capabilities. Besides, the extreme-edge domain can be huge, with millions of devices, on a cloud native scale. This sub-enabler is associated with the work regarding virtualisation and the cloud transformation studies, already introduced in [HEX223-D22], [HEX223-D32], and [HEX224-D33]. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 37 / 207 • Sub-enabler #5.3: Federated orchestration system. This sub-enabler provides enhanced flexibility in the provisioning of services across domains by automating the establishment of dynamic SLAs between them. Services can therefore be scaled up beyond the installed capacity of a provider leveraging the resources of allied providers communicating over a secure blockchain network. The corresponding negotiation and service extension or migration procedures would be handled by aptly designed smart contracts. Enabler #6: AI/ML algorithms. This enabler builds on top of the AI/ML framework already provided in the previous Hexa-X project [HEX22D62] but provides specific algorithms targeting sustainability and trustworthiness, which are considered two of key KVIs. The enabler is split into two sub-enablers, targeting each of these aspects. Below are the innovations related to each of these sub-enablers. • Regarding sub-enabler 6.1 (AI/ML-based control algorithms for sustainability), this enabler brings several innovations over the 5G technology, namely: - Energy-efficient configuration recommendations based on high-level requirements related to energy efficiency that are provided as input, then translated into decisions and configurations to be implemented by the 6G system in order to satisfy the expressed requirements. This helps further automate the M&O of the 6G system compared to 5G by adding an abstraction layer for the consumer that only has to provide high-level requirements instead of detailed configurations. - Sustainability targets in optimization processes of network and service M&O, where AI/ML algorithms are trained to optimize energy consumption and carbon footprint while satisfying the performance requirements. This helps realize a performant but also sustainable and energy-efficient 6G system, when AI/ML optimization in 5G systems was mostly focused on performance. - Adding energy spent in ML training to the overall energy consumption and optimization process. Indeed, AI/ML model training consumes energy, particularly when the model is complex and computation intensive. Thus, tracking and optimizing energy consumption during all phases of the MLOps pipeline can help to further reduce the overall energy usage for orchestration and management operations in 6G. - Energy footprint driven federation of orchestration domains, by optimizing the clustering of federated edges to reduce the communication load and energy cost, which minimizes the overall energy footprint. • Regarding sub-enabler 6.2 (Trustworthy AI/ML-based control algorithms), the main innovation regarding this enabler is to improve the robustness, privacy levels and explainability of AI/ML models in network and service M&O against adversarial attacks meant to influence model decisions for M&O, or privacy attacks for accessing sensitive data. This results in a more trustworthy, robust and resilient AI-based 6G M&O system with assured performance. Enabler #7: Network digital twins creation mechanisms This enabler is intended to develop a general framework for the use of NDTs in the network and the service M&O. In 5G systems, AI/ML models for M&O are generally trained using input from datasets, or simulators. However, real data from operational networks is difficult to obtain due to data privacy regulations, while networks simulators often fall short when mimicking real network behaviour, which results in sub-optimal models. This enabler helps bridge that gap by providing mechanisms for generating Network Digital Twins that provide an accurate representation of the network state in real time while taking into account the virtualization aspects, which allows for more efficient AI/ML models for M&O in 6G systems. Enabler #8: Real-time zero-touch control loops automation and coordination system. This enabler provides mechanisms for the automated provisioning and the lifecycle management of closed loops (CL) for zero-touch network automation. CLs can be specialized to operate in different layers (infrastructure, network, service) and domains (edge/cloud, access, core, transport) and to control particular aspects of dynamic elements like network slices and services. Closed loops are delivered as a set of cloud native functions, fully programmable and possibly assisted by AI functions. They are managed through the CL Governance as an integrated step of service and network provisioning lifecycle, so that their deployment and configuration is fully associated and jointly managed with the entities they automatically control. Moreover, the same orchestration mechanisms adopted to optimize provisioning of network functions and services can Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 38 / 207 be applied to CL instances, with resource allocation algorithms, placement strategies and scaling or migration procedures customized for each CL. This approach is aligned with the closed loop governance defined by ETSI ZSM [ZSM-009-1] and an example prototype will be implemented in the project. As a new feature with respect to CL management in 5G networks, the main innovation proposed for 6G is related to the coordination of multiple and concurrent closed loops, that may operate in different time scales and work at different layers, but with interdependencies that impact the same resources or services. CL coordination allows to resolve issues related to potential conflicts among their decisions, to schedule the execution of their actions, to identify possible issues in the combination of actions proposed by different closed loops and to find solutions to arbitrate and negotiate among contrasting decisions. Moreover, coordination techniques among hierarchical closed loops allow more scalable network automation, among different domains, through delegation and escalation procedures. Finally, collaboration among peer closed loops allows to share knowledge and insight on different network domains or services, allowing for more extensive network data analysis and more effective actions via multi-objective decisions. 2.4.1 Enhancements to the Hexa-X architecture Hexa-X-II is the continuation of previous 5G-PPP Hexa-X project [HEX24], whose main results are also considered towards designing the E2E 6G system design. The main novel capabilities for the future 6G M&O systems identified in that previous project were the following ([HEX22-D62] – Section 5.3): a) Unified orchestration across the “extreme-edge, edge, core” continuum, considering that the end-user devices can contribute to provide additional infrastructure resources which can leverage the deployment of innovative 6G services. b) Unified management and orchestration across multiple domains, owned and administered by different stakeholders, characterised by heterogeneous technologies, platforms and management systems with their own specific interfaces. This feature involves the definition of converging interfaces, as well as the mechanisms to dynamically register and expose the resources and capabilities offered from each domain, including also federation strategies. c) Increasing levels of automation in the functionalities of service and network planning, design, provisioning, optimisation, operation, and control, leveraging on closed-loop and zero-touch solutions that may strongly reduce the required manual interventions. This innovation also involves the continuous monitoring of different aspects of the network and the services performance, so that the M&O system can be able to automatically identify, detect or predict potential issues, failures, bottlenecks, or inefficient configurations, to trigger and coordinate dynamic reactions. This would also be enabled through the extensive programmability of networks and their computing resources. d) Adoption of data-driven and AI/ML techniques in the M&O system, supporting frameworks for distributed and collaborative AI, AI/MLOps, pervasive monitoring of service and network KPIs, with support for scalable data and trained models sharing along the “extreme-edge, edge and core” continuum in multi-stakeholder environments. The scope of AI/ML techniques would also cover several optimisation aspects, as well as the lifecycle actions regarding the services M&O, including the resource allocation and the slices sharing at provisioning time, services composition, scaling, migration, re-configuration, and reoptimisation of the network functions and their related resources. e) Intent-based approaches for service planning and definition. The M&O system would implement automated mechanisms for translating service specifications and commands based on intents, which could be expressed even in natural language. f) Adoption of the cloud-native principles in the telco-grade environment, involving three aspects: (i) the usage of micro-services for implementing the network functions, (ii) the implementation of service meshes, targeting to optimise the communication between applications and to reduce downtimes by using a specific built-in infrastructure layer, and (iii), the enabling mechanisms for the network services to be deployed and updated using DevOps practices, by implementing CI/CD pipelines with a very high automation degree. This third aspect was considered innovative in the telco-grade environment, involving the fact of putting together development and operational teams, which could be challenging in this scope, where services development is typically a collaborative effort involving multiple external software vendors. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 39 / 207 In the following we describe how the Hexa-X-II WP6 enablers are related to these innovative aspects identified in the previous Hexa-X project. Enabler 1: Network programmability framework. This enabler is related to innovation (c) stated in Hexa-X, and specifically, in what regards the mention about the extensive programmability of networks and their computing resources. However, the work done in HexaX regarding this is more specific, targeting the implementation of an SDN controller towards 6G, based on TeraFlowSDN [TER24] controller. So, beyond the theoretical approach outlined in Hexa-X, the work in HexaX-II is more specific and tangible, contributing to the project with a specific PoC, and also, to the Open-Source community, since TeraFlowSDN is an Open-Source ETSI hosted project [TFS24]. Enabler 2: Monitoring and telemetry framework. This enabler is directly related to innovations (c) and (d) stated in Hexa-X, although it can be also related to innovations (a) and (b), regarding the M&O over different technical and administrative domains. In fact, monitoring can be considered a kind of pervasive functionality, in the sense that for the upcoming 6G networks the collection of monitoring data from multiple and heterogeneous sources is considered necessary. Anyway, all these innovative aspects considered in Hexa-X are more oriented to the data gathering and integration problem, while the approach in Hexa-X-II goes beyond data collection, also addressing the immediate reception and the forwarding of network related events, as well as its fusion and processing. Also, a specific scalable architecture is proposed to provide a cloud-scale solution. Enabler 3: Management capabilities exposure framework. This enabler is directly related to innovation (b) stated in Hexa-X, and specifically, in what regards the definition of converging interfaces and the mechanisms to dynamically register and expose the resources and capabilities in different network domains. However, although the work performed in Hexa-X already introduced a component called “API Management Exposure” to perform this functionality, the approach there was basically theoretical. Beyond that, the approach in Hexa-X-II is more specific, targeting to implement such functionality by means of a prototype, inspired by the Integration Fabric concept defined in the ETSI ZSM specification [ZSM-002]. This will also convey contributions to ETSI. Enabler 4: Security and trustworthiness framework. Although certain security related components were included in the Hexa-X architectural design, as can be seen in the listing at the beginning of this section, that was not considered a main novel capability. In Hexa-X-II the approach consists of introducing three components that are considered key enablers: the third-party resource control separation enabler, a user-centric service provisioning system, and a specific trust management system. Enabler 5: Synergetic orchestration mechanisms for the computing continuum. This enabler is mainly related with novel capability (a) from Hexa-X, but also with (b). In this case, the stepforward in Hexa-X-II is different considering the three sub-enablers in this Enabler 5: - Regarding sub-enabler 5.1 (multi-agent systems for multi-cluster orchestration), the main added value is the development of specific data models and compute resource management solutions targeted to multicluster environments. Also, the way that the resources will be modelled will be in accordance with the developed data management schemes defined in WP2. In this way, orchestration mechanisms will be able to consider in a unified way the available resources to schedule the necessary deployment/optimization/reconfiguration mechanisms. Synergetic orchestration mechanisms based on the adoption of multi-agent systems and ML techniques will enable the better collaboration among distributed orchestration entities that collaborate towards joint objectives (e.g., assurance of a KPI or SLA for an end-to-end service). This approach has been already introduced in [HEX223-D62]. - Regarding sub-enabler 5.2 (decentralised orchestration system), although it still relies on applying the cloud-native principles (as those mentioned in (f.i) and (f.ii) above), it represents a novel approach to orchestrate resources and services towards 6G, not considered in Hexa-X. It targets the integration of the compute-continuum infrastructure resources for the network services M&O, and with focus on integrating the extreme-edge domain considering its key challenging features (the aggregation of devices beyond the MNO own premises, the diversity of stakeholders in this domain, the high heterogeneity of devices, the Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 40 / 207 high volatility of those devices, and the size of this domain, that can be huge). In short, the approach consists in (i) delegating the network services provisioning on a new set of distributed network elements, and (ii), to embed as part of the network services themselves the specific set of M&O resources they may need once they are deployed. This approach has already been introduced in [HEX223-D22], [HEX223D32] and [HEX224-D33], associated to the work regarding virtualisation and the cloud transformation studies. - Regarding sub-enabler 5.3 (federated orchestration system), this is an innovative approach, not considered in Hexa-X, to the static and slow process of establishing SLAs between service providers is proposed. This allows service continuity in roaming scenarios that may not have been envisaged by providers and as such lack pre-established agreements that in legacy M&O systems led to service disruption. Enabler 6: AI/ML algorithms. This enabler is clearly related to innovation (d) in the previous Hexa-X project (Adoption of data-driven and AI/ML techniques in the M&O system). However, while in Hexa-X the approach was in the need of defining a general AI/ML framework to support the M&O functionalities, Hexa-X-II focuses on more specific solutions to address two of the aspects that are considered relevant regarding the usage of AI/ML algorithms: trustworthiness and sustainability. Regarding trustworthiness, this enabler provides specific mechanisms intended to improve robustness, privacy levels, and the explainability of the M&O actions that may be triggered by the AI/ML models. Regarding sustainability, the enabler provides mechanisms to improve certain M&O optimization processes such as resource allocation, function placement, migration and scaling, by adding consideration of energy-related metrics to drive the AI/ML models training processes, among other performance related KPIs. Enabler 7: Network digital twins creation mechanisms. This enabler is completely new with respect to Hexa-X, where network digital twins (NDT) were not considered regarding the M&O processes optimization. However, it could be somehow related with innovation (d), i.e., as part of the general AI/ML framework, since those NDTs, as they are addressed in Hexa-X-II, will be created based on AI/ML techniques. The generated NDTs will also be used for training more efficient AI/ML models for optimizing M&O processes. Anyway, all the work regarding this in Hexa-X-II certainly adds value to what was initially stated in Hexa-X. Enabler 8: Real-time zero-touch control loops automation and coordination system. This enabler is related to innovation (c) from Hexa-X in what regards to increase the level of automation based on closed-loop zero-touch solutions. However, while in Hexa-X closed loops were considered just as static functionalities, in Hexa-X-II the introduction of the control-loops governance mechanisms allows to deploy the closed-loops in a more customizable and automatic manner over the edge/cloud continuum, as well as adjusting their placement, scaling, and configuration to the dynamicity of the network. On the other hand, Hexa-X did not explicitly investigate the cooperation among multiple closed loops. In summary, the smart network management enablers in Hexa-X-II are compliant with the main novel capabilities for the future 6G M&O systems identified in the previous Hexa-X project 4 , while they go a step further in providing more specific and tangible implementations in some cases, and more elaborated and complete concepts in other cases. Besides, new aspects are also included, such as the usage of network digital twins as part of the M&O processes, the programmability of transport networks through SDN, decentralized orchestration through the network continuum, or more specific and complete security and sustainability related enablers. 4 Intent management is not reported in this deliverable, since in Hexa-X-II it is addressed in WP2 (ref. [HEX223-D22]). CI/CD is considered out of scope for the project. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 41 / 207 3 Enablers for Management and Orchestration 3.1 Enabler 1: Network programmability framework The solution/enabler for "Network programmability framework" is designed to provide service providers with a flexible, programmable, and scalable approach to network management. This solution is based on several key technologies and approaches, including software-defined networking (SDN) [NG13], application programming interfaces (APIs), and cloud-based network management platforms [VMC+21]. D6.2 [HEX223-D62] presented the SoTA and expected beyond SoTA of this enabler. To this end, work on multiple components was proposed, as well as possible external interfaces. The following subsections present the internal architecture that could be used to implement this enabler, as well as preliminary implementation details, and early validation results. 3.1.1 Enabler design The proposed high-level architecture for Enabler 1 assumes a single administrative domain and it involves a hierarchical approach with an End-to-End (E2E) controller as the parent controller and technological domain controllers as child controllers. Specifically, it mentions an E2E SDN (Software-Defined Networking) orchestrator and technological domain SDN controllers in the IP (Internet Protocol), Optical, Time-Sensitive Networks (TSN) / Deterministic Networks (DetNet) and other domains. The Network programmability framework, which acts as End-to-End SDN Orchestrator, is the parent controller responsible for managing and orchestrating the entire network infrastructure using Software-Defined Networking. The E2E SDN orchestrator oversees the coordination and control of the network across different technological domains. The E2E SDN orchestrator can rely on existing solutions, such as ETSI TeraFlowSDN (TFS) [TFS24]. Figure 3-1 shows the hierarchical SDN orchestration architecture proposed and also the internal design of Enabler 1: Network programmability framework. TeraFlowSDN and its components has been described at [VMC+21]. Figure 3-1 Design of Enabler 1: Network programmability framework. As the parent controller, the E2E SDN Orchestrator assumes the role of overseeing the coordination and control of the network across different technological domains. It acts as a unifying entity that harmonizes the Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 48 / 207 of high time synchronization accuracy such as strict priority queueing and tagged Cyclic Queueing and Forwarding (tCQF). Still, a unified network model that incorporates all these forwarding techniques, while also allowing to reserve resources on the end-to-end path, is very complex and difficult to maintain. This implementation proposes an East-Westbound (EW) control architecture that allows a modular DetNet network using a divide-and-conquer approach. As shown in Figure 3-8, the main idea is that multiple centralized network controllers (CNCs) are employed, where each controller is only concerned about its specific network segment. This segment is a collection of devices that employ the same packet scheduling techniques, e.g., a TSN network segment with TAS-based switches or an L3 network segment with routers employing strict priority queueing. The controllers use the EW protocol to exchange various types of information. Figure 3-8: Multi-Segment DetNet Network Architecture. The EW protocol has the following main features: - It provides discovery of attached neighbouring network segments in a BGP-like fashion. - It coordinates the routing and resource reservation problem of an end-to-end flow over the different segments. Importantly, it needs to translate a given traffic specification into a representation that is useful at the next segment. Network calculus can be used to derive arrival curves, which can be used as a traffic specification across multiple segments. - It spreads global configuration along the different segments. Figure 3-9: Local actions and EW interactions between CNCs. Given the network topology in Figure 3-8, Figure 3-9 describes the different local and EW interactions in order to setup an end-to-end path. It works as follows: first, each controller discovers the topology of its own network segment. Then, it transforms this detailed topology into an abstracted version. This abstract topology consists Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 49 / 207 of neighbouring segments, local end stations, and additional statistics such as network diameter. The EW protocol is then used to flood this information to the other controllers. By gathering all abstracted topologies, a multi-segment topology can be built at every controller. Then, as a flow request arrives at TSN CNC1, the multi-segment routing algorithm computes a path of segments for an end-to-end flow. After computing this path, the end-to-end delay budget division algorithm divides the end-to-end delay among the different segments. Then, the EW protocol coordinates the end-to-end path setup by computing and installing an internal every involved segment given the assigned local delay budget and translating the traffic specification across segments. Finally, a confirmation is signalled back to the source. In this initial implementation, the multi-segment routing algorithm computes a multi-segment path that crosses the least number of segments to the destination. The end-to-end delay is divided among the segments proportionally based on the network diameter. Early Validation Results The proposed implementation was demonstrated on Linux-based routers/switchers in an emulated network setup. Five different networks of increasing size and complexity were used. As shown in Figure 3-10, operations that are contained to a local segment only, retain constant complexity in all networks. Interactions that require coordination over multiple segments, such as multi-segment topology discovery and path signalling upon a flow request, have increasing complexity, which is also reflected in the increase in control overhead. Additional details on the architecture, the used algorithms, the different interactions, and the experimental setup of this implementation can be found in [MCP+24]. Moreover, in the context of PoC#B.1, this implementation can help to demonstrate the seamless deployment of a latency-sensitive application/flow over a complex/heterogeneous TSN/DetNet network. Figure 3-10: TSN/DetNet SDN controller early validation results. 3.1.2.4 Automated transport network re-configuration This implementation is expected to be demonstrated in ADRENALINE Testbed at CTTC premises [MNC+17]. To this end, preliminary implementation details are provided, but full feature will be available for next D6.4. This new workflow will be part of TeraFlowSDN R4. The new ETSI ZSM-aligned Monitoring-AnalyticsAutomation loop architecture is being described in Enabler 8 and it is also under development in ETSI TeraFlowSDN, as depicted in Figure 3-11. This architecture may be used as an example to implement closedloops (see enabler 8 in section 3.8) for the automation of transport networks. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 50 / 207 Figure 3-11: TFS ZSM-aligned Monitoring-Analytics-Automation Loop. The TeraFlowSDN components that are involved in this implementation are: 1. The Telemetry data ingestion component is based on plugins (to provide extensibility). It supports polling-based and streaming-based raw telemetry data ingestion, e.g., raw samples, events and state data, from the underlying entities, i.e., network devices and intermediate controllers. The collected data is forwarded to the Monitoring component for further processing. 2. The Monitoring component is responsible for managing the KPI descriptors defining the type of data to be collected, the observation point used for interrogating the devices, e.g., the path in the data model used to retrieve or subscribe to receive the data, and the TeraFlowSDN resources associated with the KPI descriptor, e.g., the device, port, link, service, etc. Monitoring also checks and adjusts the timestamps of the collected data, performs the mapping with existing KPI Descriptors, and produces usable KPI Values that are later stored in a TimeSeries DataBase. 3. The Analytics component implements a number of pluggable algorithms including simple thresholdbased detectors, forecasters, and AI/ML-based algorithms. It subscribes to Monitoring to receive new KPI Values produced and analyses them according to the configured analyses. As a result, it produces Events, Alarms, and Insight that can be consumed by other components. 4. The Policy component implements a set of pluggable Event-Condition-Action—based policies. It consumes outcomes from Analytics and, whenever they fulfil a configured event (e.g., type of notification) and condition (e.g., for a specific resource) it triggers the associated policy (e.g., a set of actions to be sequentially executed). Among others, the actions supported include performing Create/Read/Update/Delete operations and/or configuring explicit configuration rules over TeraFlowSDN-managed transport network slices, services, or devices. 5. The Slice, Service, and SBI components are the existing TeraFlowSDN components used to manage transport network slices, connectivity services supporting the transport network slices, and devices where the configuration rules are installed to implement the connectivity services. 6. Finally, on top of all the components, the Automation component is responsible for managing the overall set of components involved in the Monitoring-Analytics-Automation loop. It takes as input the Service Level Agreement constraints defined in newly-configured connectivity services and transport network slices, and extrapolates the required KPIs to be monitored, the analytics algorithms required to process them, and the policies to be triggered when specific events and conditions are met. Given the wide range of scenarios that might arise, the Automation component is based on plugins, shaped Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 51 / 207 as workflows, to provide the appropriate levels of flexibility. Interestingly, the Automation component can also consume events, alarms, and insight from the Analytics component to self-reconfigure its internal workflows. Figure 3-12: Sequence diagram: operation of the Monitoring-Analytics-Automation Control Loop. The sequence diagram illustrating the operation of the Monitoring-Analytics-Automation Control Loop is depicted in Figure 3-12. The workflow assumes the transport network slices and services are created in the TeraFlowSDN controller and, if the slice/service carries some SLA to be enforced through the utilization of the Monitoring-Analytics-Automation Control Loop, it issues a request to the Automation component. The Automation component first infers the required KPIs, Alarms, and Policies to be used through the pluggable workflows it features. Next, it uses the outcome from the selected workflow to configure: 1. the Monitoring component to perform the telemetry data collection, 2. the Analytics component to perform the analysis over the data, 3. the Policy component to execute the appropriate sequence of actions encoded in the policies, and 4. it self-subscribes to alarms from Analytics to self-reconfigure as needed. Note that each of Analytics, Policy, and Automation creates an internal self-running Control Loop in order to provide the appropriate levels of scalability and that Control Loop is the one responsible for subscribing to receive data, execute the analysis algorithms, policies, and workflows, respectively. In the bottom part of the figure, two cases are illustrated: a) the execution of a ZSM loop triggering a reconfiguration of a slice/service/device, and b) the self-reconfiguration of Automation component, which might trigger changes in the configuration of Monitoring, Analytics, and Policy, based on the alarms, events, and insight produced by the Analytics component. 3.1.2.5 MEC Bandwidth Management service integration with SDN controller This section explores the synergy between MEC BandWidth Management (BWM) service and TeraFlowSDN (TFS) [TFS24] in dedicating resources for optimal network resource allocation in the gaming domain. This implementation is part of presented ETSI MEC PoC 14: Network resource allocation [MECP014] for Application specific requests using MEC BandWidth Management service and TeraFlowSDN. The source Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 52 / 207 code has been introduced in Release 3. BWM empowers applications to earmark specific bandwidth quotas (among other quality of service constraints, such as latency) for gaming applications, while TFS orchestrates the management and control of traffic flows. The resulting prioritization of gaming traffic over other data holds the potential to significantly enhance the overall gaming experience for users. The benefits that MEC BWM and TFS offer in the context of application networking improved application performance, reduced congestion, and increased scalability stand out as key advantages. The prioritization of application traffic, facilitated by BWM and TFS, contributes to latency reduction, thereby elevating the overall quality of the experience. The allocation of designated bandwidth to applications mitigates network congestion, fostering an improved application environment for all users. Crucially, MEC BWM and TFS showcase scalability, effortlessly accommodating the evolving landscape of applications and the surge in user traffic. This scalability is particularly pertinent as applications might continue to gain popularity, necessitating robust solutions that can seamlessly adapt to increasing demands. In the ever-evolving landscape of network resource management for applications, the integration of MEC BWM service and TFS stands as a groundbreaking innovation. This section delves into the unique aspects of this approach, detailing how it enhances the experience for users. • Application: MEC BWM and TFS allow application-triggered requests to the network. This allows ensuring a targeted and responsive allocation of network resources for the most bandwidth-demanding applications. The application has been modified so it is able to request the necessary network resources. • End-to-End Dynamic and Adaptive Bandwidth Allocation: The core innovation lies in the dynamic and adaptive nature of end-to-end (E2E) bandwidth allocation. The integration with TFS enables real-time responsiveness to the changing demands for the packet-optical network. This allocation mechanism ensures that resources are optimally distributed in a multi-layer network. • Application-network interface in TeraFlowSDN: A pivotal innovation lies in the novel API facilitated by TFS to directly allow applications to request bandwidth allocation resources. This dynamic allocation ensures that applications receive dedicated and exclusive network services. TFS can monitor in real-time a departure from static allocation methods, offering a proactive approach to safeguard against potential congestion and maintain an uninterrupted SLA request. • User-Centric Focus: Ultimately, the innovation in MEC BWM and TFS is driven by a user-centric philosophy. By prioritizing applications and tailoring resource allocation to their specific needs, this approach elevates the gaming experience to new heights. Reduced latency, improved performance, and minimized congestion collectively contribute to a network environment that aligns with the expectations and demands of modern applications. Figure 3-13 shows the implemented demonstration using a gaming server and client as example applications. It can be observed that ETSI TFS acts as an E2E SDN controller/orchestrator of a multi-layer multi-domain packet optical network. To this end, it controls hierarchically two domain-specific SDN controllers (IP and optical SDN controllers). Also, edge and cloud resources are shown. Figure 3-13: TeraFlowSDN and MEC integration proof-of-concept. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 53 / 207 Moreover, the figure depicts two applications: Moonlight and Sunshine. Moonlight is a game streaming application client. It is run on the network edge or extreme edge. Moonlight allows to play PC games on almost any device, through game streaming, as it is an open-source implementation of NVIDIA's GameStream protocol [MOO24]. On the other hand, Sunshine is a game streaming application server. It is a self-hosted game stream host for Moonlight, offering low latency, cloud gaming server capabilities with support for AMD, Intel, and Nvidia GPUs for hardware encoding. Sunshine is run in the cloud resources, typically located in a Data Centre (DC). Upon client application request through MEC 015 BWM allocation request [MEC015], a certain degree of bandwidth is granted from end-to-end (E2E) perspective by ETSI TFS. To this end, TFS interacts with underlying SDN controllers and Network Elements, such as whiteboxes (IP routers) and Reconfigurable Optical Add-Drop Multiplexers (ROADMs) in order to deploy constrained connectivity service. Figure 3-14 presents the implemented sequence diagram for a complex network architecture with E2E communication involving various components such as E2E SDN controllers, IP and optical SDN controllers, whiteboxes, ROADMs, and cloud resources. Figure 3-14 Sequence diagram of integration of BWM services and ETSI TeraFlowSDN. 3.1.2.6 Integration of TM Forum APIs with ETSI TFS NBI The evaluation and detailed exposition on the intersection and integration of TM Forum, IETF and ETSI, focused on enhancing network management through API standards is analysed in Appendix 8.3, while this section focuses on the concrete example of TM Forum APIs integration in ETSI TFS NBI. The TM Forum experience in proving guidelines, best practices and Open APIs from service providers and their suppliers, as well the vision that has in autonomous networks and aligned with the focus on increasing automation and flexibility in network management from ETSI’s perspective, is an opportunity to integrate and test the alignment between ETSI’s framework concept of intents and TM Forum’s approach where network operations are driven by high-level customer requirements rather than detailed technical configurations. With a common technology-agnostic approach from both TM Forum and ETSI we can purse a generic solution between all the parties, allowing independence of the underlying network technology and allowing the specification of customers need without any concern to the specifics of the technology. To achieve a holistic solution, it is needed to integrate the TM Forum APIs [TMF-OA] into the existing NBI interface of ETSI’s TeraFlow SDN Controller, offering to network service providers the best of both worlds: the robust, telecom-focused business processes and service models from TM Forum, coupled with the technical and architectural strengths of ETSI's network management frameworks. This can lead to more dynamic, efficient, and customer-centric network operations. This integration involves several aspects and it is implemented under OPTARE’s testbed (two cloud instances, Linux based, with 4 CPU each, one having 16GB RAM and the other 7,5 GB RAM and 60GB and 250GB for storage). The integration involves the following aspects: Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 54 / 207 • Service Abstraction: TM Forum APIs, such as TMF640 (Service Activation and Configuration) [TMF-640] and TMF664 (Resource Function Activation and Configuration) [TMF-664], are used to abstract network functions and services in a way that they can be consumed by business applications. • Standardization: TM Forum APIs are standardized and widely accepted in the telecom industry, which means they can help ensure that the integration into the ETSI TFS NBI will be compliant with industry practices, allowing for easier interoperability and communication among systems. • Functionality Mapping: Functions that the TM Forum APIs perform, such as activating a service or allocating a resource, are being mapped to the corresponding functionalities within the ETSI TFS NBI. • API Extension or Adaptation: There is the need to extend the TM Forum APIs or adapt them to fit into the specific requirements and capabilities of the ETSI TFS NBI. • Protocol and Data Model Alignment: TM Forum APIs need to work with the protocols and data models supported by the ETSI TFS NBI. The TM Forum APIs translate their service requests into the YANG models understood by the network controller, so a data transformation is being implemented that can translate the protocols and data models into YANG models and the ones defined by ETSI TFS NBI. • Workflow and Orchestration: The service orchestration and workflow capabilities provided by TM Forum APIs complement the orchestration logic of the ETSI TFS NBI, allowing a seamless lifecycle management of services and resources and fully integrated within ETSI architecture. Figure 3-15 represents a sequence diagram that outlines the process of integrating TM Forum APIs into the ETSI TFS NBI to facilitate dynamic network management. It features several key elements: TMF640 and TMF664, representing Service Activation & Configuration API and Resource Function Activation & Configuration API, respectively. These interfaces interact through the ETSI TFS NBI. The diagram illustrates the flow of service and resource activation requests from the TM Forum side, which are then translated and forwarded by the ETSI TFS NBI to manage network resources. It also shows feedback loops where the network resources provide status updates and metrics back through the ETSI TFS system, which aggregates and translates these updates. Figure 3-15: TMF API integration within ETSI TFS NBI interface. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 55 / 207 3.1.3 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this Enabler 1 are the following: • Scalability. As it considers a cloud-native architecture, the proposed Enabler 1 is intrinsically highly scalable, which is considered quite relevant regarding the provisioning of intense volume of connectivity services. The distributed orchestration network elements would make possible to handle an increased amount of network services and workloads of different shapes and sizes, and without requiring complex centralized systems that could become bottlenecks or single points of failure, and that would need to be upgraded to be able to manage more services and resources. • Latency: The proposed approach relies on configuring network elements on the whole network continuum. The optimal allocation of services, through network programmability in the cloud continuum can lead to a large reduction in latency. • Flexibility: Scalable and cloud-native network control and management introduces support multiple network hierarchies and technologies. • Services Creation Time is a KPI that would be of course affected by the Enabler, which shall focus on minimizing it. • Reliability. The enabler has been designed to specifically target the necessary redundance by design, so multiple instances can be easily managed. • Programmability. The proposed system is highly programmable by itself, relying on the cloud-native principles as a whole (e.g., all the network elements, such as routers, communicate using exposed interfaces, which enable programmability, and could be provided relying on highly automated DevOps practices). 3.2 Enabler 2: Monitoring and telemetry framework Programmable network monitoring and telemetry are revolutionizing the way network data is collected and analysed, by harnessing the power of Software-Defined Networking (SDN) and a suite of automation technologies. While monitoring ensures ongoing oversight of network performance metrics and status indicators, telemetry facilitates the automated collection and transmission of real-time data from diverse network sources. Network Data Analytics Function (NWDAF) is a functionality already available in 5G core. The purpose of this enabler is to extend the current SoA and consider monitoring a kind of pervasive functionality, in the sense that for the upcoming 6G networks the collection of monitoring data from multiple and heterogeneous sources is considered necessary. Enabler 2 approach also goes beyond data collection, as it also addresses the immediate reception and the forwarding of network related events, as well as its fusion and processing. Also, a specific scalable architecture is proposed to provide a cloud-scale solution. This methodical approach allows for the real-time gathering of intricate data from both virtual and physical network components as well as applications. The essence of this framework is its multi-layered data collection capability, which spans across various services and applications, ensuring a comprehensive view of the network's operational dynamics. By implementing a multi-vendor strategy, it achieves higher scalability and flexibility. This real-time data provides a robust foundation for decision-making, allowing network administrators and automated frameworks to dissect network traffic flows and performance metrics meticulously. Through this analysis, insights into network utilization patterns emerge, pinpointing areas of resource underutilization. With this knowledge, network configurations can be refined and optimized, fostering enhanced performance and cost-efficiency. Moreover, this sophisticated monitoring extends to the measurement of energy consumption across network elements, including computational resources. The collation of these data is pivotal for the creation of algorithms aimed at deploying networks that balance operational demands with energy efficiency. The breadth of technologies and protocols that the proposed enabler will leverage is crucial for its efficacy. And they play a vital role in enabling real-time data harvesting and analysis. These protocols not only facilitate the automation of network management tasks but also aid in the orchestration of network functions and policies. Moreover, the system is designed to incorporate robust authentication and privacy measures, addressing the ever-important concerns of security, privacy, and systemic resilience. To manage cloud-scale operations and circumvent any potential monitoring bottlenecks, the system architecture contemplates a Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 56 / 207 sequential processing approach for metrics and alerts. This includes a series of independent yet interconnected steps—acquisition, normalization, visualization, evaluation, and publication of metrics and alerts—ensuring that each aspect of network monitoring is handled with precision and efficiency, thus maintaining the smooth operation and scalability of the system. D6.2 [HEX223-D62] presented the SoTA and expected beyond State of The Art of this enabler. To this end, work on the multiple components was proposed, as well as possible external interfaces. The following subsections present the internal architecture, as well as preliminary implementation details, and early validations results. 3.2.1 Enabler design Monitoring and telemetry are crucial components that enable system reliability, performance, and security. The framework depicted in Figure 3-16 provides a clear pathway from data collection to actionable insights, utilizing a set of tools and processes. The first stage involves the Collector, which is responsible for gathering raw data. Tools like gnmic [GNMIC24], OpenTelemetry [OTL24], and Prometheus [PRO24] are often employed at this stage. Once data is collected, it is passed to the Processors. The processors, which may use OpenTelemetry and tools like Apache Airflow, are responsible for parsing, aggregating, and transforming the data into a coherent format that can be analysed. After processing, the data is sent to the Exporters, which may also utilize OpenTelemetry for exporting the processed data to various destinations. This step is essential for delivering the information to the appropriate tools used for analysis and visualization. Finally, the data arrives at tools like Prometheus for storage and Grafana for visualization. Prometheus can store the processed timeseries data, making it queryable, while Grafana specializes in turning this data into actionable insights through its powerful dashboards. These insights enable operations teams to detect and respond to issues in real-time, making informed decisions based on comprehensive data. This architecture has the challenge to be adapted and extended to support all specific requirements for upcoming 6G networks. Figure 3-16: High-level architecture of monitoring and telemetry framework. 3.2.1.1 System components Figure 3-17 depicts a schematic of a data processing system, arranged horizontally to illustrate the flow and interaction of data between various components. It consists of a cloud-native design with microservices connected through a common bus. At the forefront of the system is the "Subscription Manager," which likely oversees the configurations for data collection across the system. Next in line are "Exporter #1" and "Exporter #2," denoting two distinct pathways through which data is extracted and possibly disseminated to different destinations or for various uses. An "In-band telemetry exporter" is also featured, implying a mechanism tailored to handle telemetry data within the operational bandwidth, ensuring efficient monitoring. As we move further right, there's an "Event Processing" unit, indicative of a subsystem dedicated to handling discrete events that occur within the system—these could be irregular, significant, or require special processing separate from the main data flow. Adjacent to this is a "Time Series #1" database, which suggests storage and retrieval functionality for data points collected sequentially over time, a common requirement for monitoring trends and patterns. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 57 / 207 The "Non-SQL DB #1" component indicates the use of a non-relational database, which is designed to store and manage large volumes of unstructured or semi-structured data, signifying flexibility and scalability in data handling. Finally, the "Visualization #1" module posits an endpoint for the system's data, where it is likely transformed into graphical representations to aid users in interpreting and analysing the data. Supporting these main components are several "Collector" blocks, labelled "Collector #1" through "Collector #N." These are depicted beneath the main data pathway, suggesting a hierarchical relationship where these collectors feed into the upper-level components. The term "Collector #N" alludes to a scalable array of data collection points, with "N" representing an indefinite number that can be expanded as needed. Among these, an "Active Collector (in-band telemetry)" is specified, highlighting its active role in gathering in-band telemetry data, which is crucial for real-time monitoring and analysis. Below the description of each of the components in the Figure 3-17. Figure 3-17: Monitoring and telemetry framework internal architecture. • Distributed event streaming platform (depicted as Data Bus). • Communication Protocols: Choose protocols for data transfer between devices and the central system. Common ones would include MQTT, HTTP, or CoAP for IoT. • Real-time Processing: Implements real-time processing for immediate analysis and response. Tools like Apache Kafka [KAF24] or Apache Flink [FLI24] can help with streaming data. • Collector: Collects data from various sources. This could include IoT devices, or other data points. Sensors and probes are also included in this category. • Exporter: Set up alerts for abnormal conditions or thresholds. This ensures the delivery of notifications when something requires attention. • Subscription Manager: is the responsible for handling the specific topics in the data bus. • Time series: Plan for long-term storage and analysis of historical data. This is crucial for trend analysis and making informed decisions. Services like Prometheus [PRO24] or Nagios [NAG24] could be used. • Database: robust database system to store the collected data efficiently. Options include SQL databases or NoSQL databases, depending on the data structure. • Visualization: Tools like Grafana [GRA24] can help in visualizing data and creating dashboards. • Security Measures: Implement robust security practices to protect the platform from unauthorized access. This includes encryption, authentication, and access controls. • Scalability: Design the platform to scale easily as the number of monitored devices or data points increases. The solution can be deployed in the entire network continuum. 3.2.1.2 Workflows Figure 3-18 shows the sequence diagram of the flow of events through a system that collects, processes, and visualizes data. In the diagram, there are several entities or components involved: Collector1, Data Bus, Exporter1, Timeseries Component, External App, Database, and Visualization. A data-driven process begins with Collector1, which is responsible for generating an event. This event is then passed to the Data Bus. The role of the Data Bus is to distribute the event to various components within the system. In this case, it distributes the event to both Exporter1 and Timeseries Component. Exporter1 seems to have a role in further distributing the event, perhaps to different systems or for different uses. Meanwhile, the Timeseries Component relates with Database and Visualization. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 64 / 207 Figure 3-24: Monitoring platform dashboard showing battery data collected from cobots. 3.2.3 Conceptual solutions 3.2.3.1 Passive/in-band and active telemetry for TSN/DetNet networks Given the critical nature of TSN and DetNet, monitoring is a very important component to verify that the network can meet its requirements regarding end-to-end latency and packet loss. There exist different monitoring strategies. With passive network measurements, data is gathered by passively listening to network traffic, i.e., without interfering with data traffic. Passive monitoring allows to collect simple statistics at switches and routers, such as bytes sent, lost packets, and other similar statistics. Conversely, active probing is an active monitoring strategy where artificial data packets (probes) are sent into the network, using the same links as the data traffic and collecting statistics along the way. Active probing can be used to e.g., measure the end-to-end delay along a specified path. Since active measurements generate additional network traffic, they interfere with the normal traffic flow, and as such they must be carefully planned. In-band network telemetry offers low-overhead monitoring possibilities, enabling an end-to-end performance view of the network (including end devices), while not injecting any new packets in the network. The idea is to embed monitoring information in the actual data packets, and as such information can be collected per-hop, per-packet and perflow. A monitoring algorithm for a TSN/DetNet network segment computes a suitable monitoring strategy based on the monitoring configuration and the current traffic flows in the network, as shown in Figure 3-25. This algorithm should carefully analyse the trade-offs between different monitoring strategies, and make sure that the monitoring overhead is kept between certain limits such that it does not interfere with the data traffic, which is one of the main challenges/requirements outlined in the framework of Operations, Administration, and Maintenance (OAM) for DetNet (RFC9551). In the case of active monitoring via probes, the optimal placement and deployment frequency of these probes are crucial challenges that need to be addressed. The resulting monitoring strategy might be a combination of passive, active, and in-band telemetry. After collecting the monitoring data, this data will be stored in a local database and can be consumed by external parties, or by a closed control loop (cf. Section 3.8.3.1). The coordination of monitoring strategies in a multi-segment TSN/DetNet setup is another challenge, as a single traffic flow might cross multiple segments. This allows for the virtualization of probes across segments, and more importantly end-to-end statistics for traffic flows. The EW protocol, as described in Section 3.1.2.3, can be extended with end-to-end monitoring coordination capabilities for a multi-segment TSN/DetNet network. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 65 / 207 Figure 3-25: Monitoring for TSN/DetNet network segment. 3.2.3.2 Data fusion for signals correlation and remediation actions The idea behind signal data fusion [TAZ+22] is concentrated on active and passive telemetry of different entities and sources of the system, aiming at facilitating agents to reactively mitigate system and network failures, as well as proactively preventing them altogether by analysing real-time information. A clear methodology of collecting the necessary data from the different entities contributing to a heterogeneous and distributed system is crucial to its efficient maintenance and orchestration. OpenTelemetry [OTL24] provides such a methodology for accessing data signals at real-time in distributed architectures. The OTLP (Open Telemetry Protocol) Collector provided by the OpenTelemetry framework comes in two different flavours supporting a variety of distributed data collection designs. The OTLP Gateway can be used for centralized deployments where data collection sources are directly accessible, while the OTLP Agent is available for exporters in distributed locations which forward the collected signals to the OTLP Gateways residing in centralized locations. Data signals refer to performance information such as metrics, traces and logs collected by both infrastructure and network nodes. Application and system-related signals are collected by using OpenTelemetry instrumentation libraries, while TeraFlow SDN Controller’s NBI provides accessibility to QoS metrics related to the performance of the network. As shown in Figure 3-26, this solution leverages the implementations of the OTLP Collector to deploy across the heterogeneous infrastructure of the computing continuum to collect local observability signals using the OTLP agent version of the collector, preprocess the data close to the source to identify early identification of events and finally aggregate this information to be fused with other signals at more ‘central’ locations using the OTLP Gateway version of the collector. Through the collection data bus, the collected information is processed for identifying anomalies or specific events at runtime and is then exported to the data fusion observability tool for further analysis and aggregation of the information, to assess higher-level orchestration components. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 66 / 207 Figure 3-26: Observability signals data collection and fusion. 3.2.4 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this Enabler 2 are the following: - Scalability: this enabler ensures that network monitoring and telemetry can scale to meet increasing demands without degradation in performance. The system's architecture, designed to handle cloud-scale operations, effectively manages large volumes of data and network activities, preventing bottlenecks that can impede scalability. - Latency: The real-time data gathering capabilities of this system minimize latency in network monitoring and decision-making processes. By leveraging protocols like NETCONF, gRPC, and gNMI, the system can quickly collect, process, and react to data from various network components, ensuring timely responses to dynamic network conditions. - Flexibility: The suite of automation technologies and the use of diverse protocols (NETCONF, REST, YANG, gRPC, gNMI, SNMP) enable this enabler to be highly adaptable to different network configurations and requirements. This flexibility allows network administrators to tailor the monitoring and management solutions to specific needs, enhancing the overall efficiency of network operations. - Reliability: The enablers data acquisition mechanism underpins dependable network performance monitoring and fault detection, thereby increasing the overall reliability of network operations. - Automation: Automation is a core component of this enabler, with protocols designed to facilitate the automation of network management tasks and the orchestration of network functions and policies. This allows for more efficient use of resources and reduces the need for manual intervention, thereby enhancing operational efficiencies and reducing the potential for human error. Regarding the KVIs, it is considered that the most evident to be affected by this sub-enabler is Sustainability in what regards the following aspects: - One of the distinctive features of this enabler is its focus on monitoring energy consumption across network elements, including computational resources. This capability is pivotal for developing algorithms aimed at optimizing energy use, thus supporting sustainable network operations. By enabling more energy-efficient network configurations and operations, this enabler helps reduce the overall environmental impact of network infrastructures. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 67 / 207 3.3 Enabler 3: Management capabilities exposure framework Hexa-X-II is intended to provide highly configurable solutions which may be tuned to the requirements of different 6G stakeholders, as introduced in section 2.3. Semantics (referring to network, cloud, zero-touch…), operational domains, and infrastructure domains are used to categorize these capabilities. Programmable compositional patterns [IS21] are considered required for the integration of these capabilities, with an emphasis on plug-and-play methods and cloud-native strategies. In order to accomplish this, a management capabilities exposure framework — which offers a service bus for smooth Hexa-X-II capability interoperability — is considered. This fabric enables integration between administrative domains and capacities by implementing features such as connection, dependability, security, and observability. By enabling these features, this component accomplishes two primary objectives: it first enables the modularization and statelessness of management services, allowing them to be deployed as scalable, containerized microservices. Second, it facilitates the system's transition to APIfication, in which traditional interfaces, i.e. older methods of system interaction and communication, like tightly coupled integrations, and legacy protocols like RPC (Remote Procedure Call), SOAP (Simple Object Access Protocol) etc., are replaced with HTTP-based RESTful APIs, which handle producer-consumer interactions. Deliverable D6.2 [HEX223-D62] presented the SoTA and expected beyond SoTA of this enabler. A draft architecture definition was also pointed out, along with the relevant internal and external interfaces. This deliverable presents the work carried out to realise the architecture planned in deliverable D6.2, the improvements to the first design, and the first deployment results. 3.3.1 Enabler design Figure 3-27 shows the high-level architectural view of the suggested solution. Thanks to the use of a message broker, the design is event driven. The central component of the architecture, the broker, enables communication between the external services serving as a shared message bus and the framework internal components. The establishment of communication channels among the different components facilitates the connection and global orchestration of the operations of the various enablers. Figure 3-27: High level architecture of Management Capabilities Exposure Framework. The implementation draws inspiration from the ETSI ZSM Integration Fabric [ZSM-002]. This component is a key element within the wider ETSI Zero-touch Network and Service Management (ZSM) framework and is Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 68 / 207 responsible for facilitating seamless integration and interoperability between a variety of management and orchestration entities in a network environment. It acts as a centralised middleware layer which enables communication and data exchange between different ZSM components, such as network controllers, orchestrators, and management systems. The ETSI ZSM Integration Fabric plays a pivotal role in the realisation of the vision of zero-touch automation. It provides a set of standardised interfaces, protocols, and mechanisms which facilitate the orchestration of network resources, the automation of service delivery, and the optimisation of network operations. As the Management Capabilities Exposure Framework would be set to serve as the central component to facilitate interoperability in the E2E system, as stated in section 2, the design principle of the ZSM Integration Fabric would align perfectly with the design view. Further details of the internal architectural definition, in addition with early implementation details and validation findings will be presented in the following sections. 3.3.1.1 System components The components of the Management Capabilities Exposure Framework are described in Table 3-1. Table 3-1: Management Capabilities Exposure Framework components. Component Description Services Consist of network functions, management systems, and external applications, which need to communicate with each other. Each of them is isolated from each other, unaware of the topology. They can communicate only by interfacing with the integration fabric. To have a standardized way to connect and deploy them, each of these services would be implemented in the form of containers. Each of them alongside the container scope is equipped with a client that makes it possible to communicate with the message broker. Service registry It would have two main functions: it is the creation/deletion services subscriptions and manage inter-service relations. It is the insertion and removal endpoint for new services. This single-entry point makes it possible to better regulate the inter-service relations updates and address the horizontal scalability also in a multi-tenancy scenario. Inserting a new service means subscribing it to the base functionalities’ topic (alarm, coordination…). After this registration step the service will not be aware of the current structure around it, but it will be aware of the relevant events happening in the architecture. One service that requires to take advantage of integration fabric services must start a one-time onboarding procedure, that may be built as a volatile process, exploiting virtualization technologies like serverless computing. The second task shipped out by this module embraces all the operations to keep the list of services managed by the integration fabric updated. It shows the amount and the type of services managed, and it keeps the current policies used, relations within the system in terms of topic and subscriptions in the scope of the integration fabric up to date. This enables service discovery features, crucial in scope of the integration fabric. Security module Secures access to the resources managed by the integration fabric, as well as the communication within the modules. It ensures AuthN/Z mechanisms to guarantee that the resources are accessed only by accredited stakeholders. It enables to manage token/key to unlock different QoS or certain services. Its last feature is to manage the encryption within modules (Transport Layer Security - TLS) to avoid leaking information in the inter-service communication. Connectivity manager It offers advanced management in terms of subscription, retention policies and all the aspects related to the management of topic and related queue. The connectivity manager routes and distributes traffic between services based on defined policies and subscriptions. This module is connected, directly, to the message broker to modify the subscription. A change in the subscription implies a change in the involvement of a particular service in the system. In fact, the interaction of the involved services depends on the consumed and published messages, and for this the service must subscribe to a certain communication channel, i.e. a topic. The modification of a subscription therefore leads to a modification of the communication patterns within the systems. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 69 / 207 Tracing module The tracing module oversees recording and following the requests flows through various services. Analysis, monitoring, and troubleshooting of the system is made possible by tracing Admin endpoint Observability is an important point of the Management Capabilities Exposure Framework, enabling the continuous check of the health and key metrics of the system. This is possible because this endpoint embeds a tracing and metric generator module (that refer to the enabler performances) that are connected to the topic of interest. An admin endpoint is also important to have a clear view and control at a high level on the general services relations, and routing rules between services. That block would be directly connected to the management’s blocks of the enabler architecture, i.e., the connectivity manager and service registry, and the admin endpoint is secured through the security module, to ensure only authorized subjects can access this integration fabric control panel. Message broker It is the central component of the architecture. It represents the central node in which all packets are routed according to the subscription, the available topics, and the retention policies. Each service will be asynchronously connected to this module. From the broker each of them would receive all the information of the other components and will react to certain event or request by directly publishing on the broker itself (event-driven architecture). 3.3.1.2 Internal Architecture of the system components The enabler at this stage is composed by two main macro components: • The Kafka [KAF24] cluster and all the related services, more details are presented in Section 3.3.2.1. • The services that manage onboarding/offboarding, listing and security aspects of the framework. Figure 3-28 shows the general structure and how the internal components of the framework communicate each other. The core of the framework is the Apache Kafka cluster, which implements the event-driven architecture at the heart of this framework. It makes use of the Apache Kafka protocol [KAFP24], a high-performance binary messaging protocol that enables the efficient and reliable exchange of streamed data over TCP between producers and consumers in a distributed, fault-tolerant and scalable manner. The management of the data stream is topic-oriented, using the concept of queues. A Kafka queue is a logical stream of records belonging to a topic. Queues are not a single stream, but a more granular structure divided into partitions. Partitions are ordered, immutable sequences of records within the queue, identified by unique offsets. This allows data to be replicated over multiple brokers, and allows multiple consumers to read data simultaneously, providing parallelism, horizontal scaling, and fault tolerance. This structure makes Apache Kafka a highly scalable, distributed, and fault-tolerant streaming platform that enables real-time processing of continuous data streams. As a result, with Apache Kafka: - It is possible to handle massive amounts of data with high throughput and low latency. - The distributed architecture provides data redundancy and fault tolerance, allowing applications to continue to operate in the event of machine failures. - Provides strong data durability and flexible retention policies, allowing a high degree of freedom in managing message persistence. Users can either store and access data for an extended period of time or persist for a limited period of time. - This feature fits perfectly with the design vision of the enabler, which needs to be flexible, consistent, responsive and easily integrated with a discrete number of actors. Another key benefit of using Apache Kafka is its Access Control Lists (ACLs) feature. ACLs provide finegrained control over user access to Kafka resources such as topics, consumer groups and brokers. This feature is critical for organisations that require strict data security and regulatory compliance, as it allows them to effectively manage and monitor user access to prevent unauthorised access and data breaches. The added value is that it is fully compliant with an mTLS approach (chosen as method to secure communications within the system), using the certificate as an identity to encrypt communications between the actor and the Kafka cluster. Finally, it is important to note that Apache Kafka is a mature open-source project with a large community. The solid background of the project makes it able to offer a lot of integration with a plethora of frameworks and Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 70 / 207 programming languages, which is crucial in a multi-domain scenario, as the one have to deal with this framework. Figure 3-28: Management capabilities exposure framework main components and interactions overview. Figure 3-28 highlights that an enabler willing to exploit the framework can interact in two ways: using REST APIs with HTTP protocol or using Apache Kafka protocol. In the former case, the new entities do not interact directly with the services REST API interfaces, but all the requests are received by a centralized API gateway in charge to forward them the requests. This makes possible to have an additional control on the received requests, in term for instance of traffic control, and, in addition, with the support of an Identity provider to secure a route. For instance, the route connecting the enabler with the Security Manager is OAuth2.0 secured, to avoid leaking security information. Figure 3-29: Management capabilities exposure framework interfaces. The combination of the API Gateway and the IdP basically represent the AuthZ/AuthN interfaces, shown in Figure 3-29, of Security modules (displayed both in Figure 3-27 and Figure 3-29). The latter case involves the actual communication with the system, i.e. using Apache Kafka protocol. An enabler that is already onboarded Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 71 / 207 to the functionalities of the Management capabilities exposure framework would be able to publish to and/or consume from a set of topics. This kind of communication is secured through the usage of mTLS (mutual TLS), that makes possible to jointly encrypt the communication and give a method to identify the requesting party. Each entity interacting with this enabler would be identified by the certificate (distributed by the usage Security Manager REST endpoint shown in Figure 3-28, which API is described in Appendix 8.2.3) that is used for the TLS connection with Apache Kafka cluster. As consequence, the authorization would be proportioned by mean of ACL (Access Control List) offered by Apache Kafka, using the CN (Common Name) declared in the certificate. A more detailed description of this interaction will be described in the next Section 3.3.1.3. Figure 3-29 reports the entry points of the management capabilities exposure framework, for its main components. Each interface would cover a different aspect of the communication, such as on boarding of new entities, service listing of the entities that are plugged to the management capabilities exposure framework, security and monitoring. The interfaces are marked with different colours to distinguish the different stage of development of the system. The ones highlighted in green are related to core functionalities, fundamental to deliver a basilar service. Given that, these functionalities are prioritized. The ones in yellow will be delivered at a later stage, providing not core functionalities but added values to complete the solution. The red ones will be developed in the final stage. Table 3-2 describes the interfaces for the core functionalities of the management capabilities exposure framework and identifying protocols and technologies that could be used as baseline for their implementation. Table 3-2: Management capabilities exposure framework interface mapping. Interface Description Registration management It is the system interface that guarantees the onboarding of new enablers in the integration fabric scope. It basically creates, after a REST query bringing all the key information, a new topic in the message broker related to the registered service. In addition will create a consumer group and a consumer instance, to the message broker, to link the new added service. Finally, this interface will make it possible to change the settings of the registration, with ad hoc REST queries. Discovery management This interface querying the message broker forward the information to communicate with another enablers. In detail, after a REST call, it forwards all the useful information of the topic related to the requested enabler. Service listing management It has the role of forwarding the information about all the enablers registered to the integration fabric. Querying this interface with a REST call it forwards a list of all the topic registered in the message broker. Subscription management This interface will permit to manage the subscription of all the enablers by the admin, and to manage internal creation/revocation of the subscription of a certain topic. Error exposure management This interface will expose the information published to the topic that collects the errors of all enablers. This can be queried with a REST API in an asynchronous way. Inter-service management It is essentially represented by the message broker. Each Kafka topic can represent (i) a communication channel that a producer, i.e. an enabler, can use to deliver information to an authorized consumer, i.e. another enabler, (ii) a broadcast topic that deliver global info that are of interest to all the enablers, for instance critical errors. Authentication management Authentication will be offered using a 3rd party IdP (Identity Provider), to manage the identities in a more secure way. Authorization management Authorization features are offered as a combination offered by the ACL of the message broker and the OAuth of the IdP. Communication encryption management This feature is guaranteed by the functionality natively offered by the message broker, that leverages TLS. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 72 / 207 3.3.1.3 Workflows Figure 3-30 shows the sequence of operations for a single requesting actor for subscribing to the Management capabilities exposure framework services. The entity to onboard must interface with the subscription manager via a RESTful API call. This API is built using OpenAPI 3.0 specification. That makes it easier to integrate, speeds up development, and improves collaboration and use among developers, consumers, and other stakeholders. API description using OpenAPI 3.0 provides a standard, machine-readable format for defining and documenting RESTful APIs. This promotes consistency, interoperability and reusability across platforms and programming languages. In the message body the requesting actor should include all the requested information for the onboarding as well as the necessary info to authenticate that it is a trusted party. After the subscription manager has proven the identity of the requesting actor, it will directly communicate to the message broker. Using the info passed by the requesting actor, this will create a custom topic for the requesting actor, a consumer group, and a consumer instance. After all this roll-out is completed, the subscription manager will send all this information to the requesting actor in order to make it able to communicate with the message broker correctly. Figure 3-30: Registration to Management capabilities exposure framework. The flows presented in Figure 3-31 and Figure 3-32 represent the functioning of the listing calls. The first is the global topic listing. When a requesting actor asks for the list of all topics, after being authenticated, the listing service will forward a list of all topics of interest to it. This means, that in addition to the global topic used for errors and general information broadcasting, the response body contains the list of all topics to which the requesting actor has access. Each of them specifies the name and the related access rights (consume/produce, consume only). When a requesting actor asks for information of a single topic, in addition to the above-mentioned information, it receives additional information on this topic. These details include the number of replicas, partitions, how the replicas and partitions are distributed among the brokers, last offsets, and finally the list of entities that have access to the information and the related access rights. Figure 3-31: Management capabilities exposure framework topic listing. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 73 / 207 Figure 3-32: Management capabilities exposure framework single topic information retrieval. The last workflow represented in Figure 3-33 shows the operations performed by the requesting actor when it has to produce or consume. First, out of the loop, the message broker is constantly updating the permissions that each requesting actor has, by using the native Apache Kafka features offered by ACL. All the information about identity is stored by the security module that oversees triggering the updates to the ACL of the message broker when an identity is added or changed. Based on these access rights, the client, i.e. the requesting actor, will be able to produce or consume to/from a specific topic. Figure 3-33: Enabler communication with Management capabilities exposure framework. The interfaces of the management capabilities exposure framework are reported in Appendix 8.2.3. 3.3.2 Preliminary implementation and early validation results An initial prototype of this enabler is being developed and tested at one of the partners' local testbed (Telefonica), which will be made available to other partners in the consortium. The integration with the other partners is planned in the PoC B.1. The framework developed will serve as the central communication bus and exposition of interface in the E2E architecture. In addition, once the aforementioned integration has been completed and the results achieved, the implementation will be proposed to the ETSI ZSM community as a reference implementation of the ETSI ZSM Integration Fabric [ZSM-002]. The following section presents additional information regarding the implementation works performed. 3.3.2.1 Apache Kafka cluster structure The Apache Kafka cluster is the core component for providing the communication medium offered by Management capabilities exposure framework. The Figure 3-34 shows the internal structure of the cluster. At first glance it is possible to see that the cluster is deployed using KRaft methods. This mean that no ZooKeeper Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 80 / 207 Figure 3-41: 3P onboarding – workflow. From that moment on, the operation time starts, whereby the tenant can access capabilities offered by COP MFs via granular access control. The actual solution for this access control is highly specific to the protocol used for data model specification, either REST/Open API or Netconf/YANG. Table 3-3 summarizes the main differences between both options. Table 3-3: AuthN/AuthZ – protocols comparison. Aspect Solutions REST/OpenAPI Netconf/YANG AuthN and AuthZ functionalities Co-located: both functions are integrated into the same server (AuthN/Z server) Separated: Authentication takes place on the transport layer, and the authorization on the Netconf server AuthN protocol Client credentials assertion based authentication (with optional OIDC) Mutual TLS (mTLS) AuthZ protocol Token based authorization framework with various grant modes, as specified in RFC 6749 [RFC6749] OAuth2.0. Static authorization, as specified in RFC 8341 [RFC8341] Network Configuration Access Control Model (NACM). Figure 3-42 illustrates the workflow for the operation time using REST/OpenAPI protocol. In this option, the tenant first asks for authentication (step 1), issuing the “tenant-uuid” attribute. The COP AuthN/Z server identifies which “identity” instance is populated with this “tenant-uuid” attribute (step 2). With this information, the server generates an identity token (step 3), which is sent back to the tenant (step 4). The tenant then asks for authorization (step 5), issuing the received token. The server proceeds with token decryption (steps 6-7), in order to identify which roles and permissions have been allocated for that tenant (steps 7-8). Upon this identification, the server is informed about the tenant’s management space, and proceeds with the generation of an access token (step 9), which is issued to the tenant (step 10). From this moment on, the tenant can gain access to the MF and invoke necessary capabilities (steps 13-14). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 81 / 207 Figure 3-42: Operation time – REST/OpenAPI workflow. Figure 3-43 illustrates the workflow for the operation time using Netconf/YANG protocol. The process is rather similar to the REST/OpenAPI solution, with three main differences: i) authentication occurs at the transport layer, using mTLS; ii) once authenticated, the tenant has direct visibility on COP MFs; and iii) the COP MFs delegates authorization in NACM server. Figure 3-43: Operation time – Netconf/YANG workflow. The interfaces used for this solution are APIs conveying YAML or YANG models, as illustrated in the workflows. To ensure clarity regarding this enabler, it is essential to understand its inherent relation with enabler 3 and the specific role it plays in defining the information model that identifies a third party within a system. The identity onboarding process described here is complementary to the definition provided in the enabler 3 definition. The identity, roles, and access rules determine whether and how a third party can access the M&O resources that are exposed by the enabler 3. The security described in enabler 3 is solely related to internal security. It is Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 82 / 207 certain that the identity of the third party defined in this sub-enabler 4.1 must be translated into an entity that can interact with enabler 3 (thus enabling the establishment of mTLS with the Apache Kafka cluster). In conclusion, this enabler can be regarded as a layer that separates the enabler 3 and the resources it exposes from the external environment. 3.4.1.2 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this sub-enabler 4.1 are the following: • Availability: by separating resource control for third-party entities, the system can prevent one tenant from monopolizing resources or causing disruptions that affect the availability of services for other tenants. This ensures that each tenant has fair access to resources and reduces the likelihood of downtime due to resource contention or malicious activities. • Reliability: resource control separation helps maintain the reliability of the system by isolating the impact of failures or errors within individual tenant environments. If a particular tenant's application or activities lead to a failure, it can be contained within their allocated resources without affecting the reliability of other tenants' services. This isolation minimizes the propagation of failures and enhances the overall reliability of the system. • Security: separating resource control for third-party entities enhances security by enforcing strict access controls and permissions. It allows administrators to define and enforce policies governing the interactions between tenants and system resources, mitigating the risk of unauthorized access, data breaches, or malicious activities. Additionally, by isolating each tenant's environment, the impact of security breaches or vulnerabilities can be contained, limiting their scope and protecting the integrity and confidentiality of other tenants' data and services. • Maintainability: resource control separation simplifies system maintenance and management by providing clear boundaries between tenants' environments. This allows to apply updates, patches, and configurations more efficiently, without affecting the operations of other tenants. It also facilitates troubleshooting and debugging processes by isolating issues to specific tenant environments, making it easier to identify and resolve problems without disrupting overall system maintainability. Regarding the KVIs, the most affected by this sub-enabler is Trustworthiness. Third-party resource control separation significantly impacts the trustworthiness of 6G networks by reinforcing confidentiality, integrity, availability, data privacy, operation resilience and security. It serves as a foundational element in building trust among stakeholders providing a resilient, secure, and trustworthy digital ecosystem. 3.4.2 Sub-enabler 4.2: User-centric service provisioning system 3.4.2.1 Sub-enabler design In the context of 6G networks, the relevance of the user-centric service provisioning lies in its ability to enhance user experience, optimize network resources and support the diverse and dynamic requirements of future applications and services. This sub-enabler aims at provisioning tenant subscribers (e.g., enterprise users, endusers) with optimal and personalized Quality of Experience (QoE), according to their preferences, SLAs of subscribed services, and network context (e.g., context scenarios). This sub-enabler would provide the following capabilities: • Definition of customized service policies for individual tenant users. These policies would be based on the UE Route Selection Policy (URSP) concept defined in 3GPP [23.501] and introduced in D6.2 [HEX223-D62]. These rules are interpreted and executed by mobile modems of user devices. The range of eligible devices not only includes OS-type devices (e.g., smartphones), but also other novel devices that Hexa-X-II is considering [HEX223-D52]. This capability is relevant for service fulfilment. • Assisting users to gain access to subscribed services, while keeping their data safely stored, preventing any unauthorized entity to read/modify them. This would be done by injecting URSP rules into user devices. This capability is relevant for service activation. • Ensuring the behaviour of tenant services comply with the KPIs (section 3.4.2.2) defined in the SLA, throughout the entire service lifetime. In case of violation, corrective actions will be taken to set Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 83 / 207 service back to the desired stated, leveraging closed control loop automation features. This capability is relevant for service assurance. This enabler is strongly related to the concepts of intents, SLAs and Closed Loops (CL), which are correlated through a hierarchical relationship. This relationship can be split into two parts: intents to SLA and SLA to CL. With regards to the former, intents are translated into SLAs: the high-level goals defined by intents are broken down into specific, measurable requirements in SLAs. This translation involves understanding what performance metrics are necessary to achieve the high-level intent. Conversely, SLAs provide the benchmarks that CLs strive to maintain. CLs continuously monitor performance metrics and make adjustments in real time to ensure that the service meets the SLA criteria. For each UE, the relevant SLAs and associated CLs are stored and utilised to keep the KPIs within the defined parameters. It can be argued that the aforementioned flow from intents to SLAs to CLs allows for the systematic management of resources and service quality. This enables the rapid identification and rectification of any deviations from the defined performance thresholds, thus maintaining operational stability and performance consistency. The contents presented in this and the following sections pertain to a conceptual design, with no provision for subsequent implementation. 3.4.2.1.1 System components Figure 3-35 illustrated the framework eligible for enabler 4. In this section, we will be focusing on those components that are inherent to enabler 4.2, highlighted in red. These components are the following: • Intra-tenant space: it specifies the management space provisioned to each registered tenant. This management space is the result of mapping information from the 3P profiling component (within the Intent-based Digital Service Manager, see [HEX223-D22]) into properties that the M&O layer can interpret and act upon. For user-centric service provisioning, the intra-tenant space captures the URSP rules eligible to individual tenant users. These URSP rules are injected into mobile user devices (service fulfilment), so that devices can grant access to user subscribed services (service activation). • Imported models: this is a repository that stores the models for SLA (defined in Intent-based Digital Service Manager) and Closed Control Loop (defined in enabler 8). The mission of these models is to help provide zero-touch solutions for service assurance, using closed loop automation to supervise the compliance of tenant services to the signed SLAs. To that end, SLA attributes must be allowed as input for the goal of Closed Control Loop, so both models can work together. Figure 3-44: User-centric service provisioning – internal architecture. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 84 / 207 Figure 3-44 shows the internal architecture of impacted components. As seen, the DSP sends the relevant information from “3P profiling” to the COP, which allocates it within the trustworthy 3P service provisioning framework. The relevant information for this sub-enabler consists of: i) contracted services, ii) SLAs, and iii) subscriber data. All contribute to the definition of URSP rules, which are stored within the “intra-tenant space”. The SLA attributes are also captured according to the SLA model definition registered in the “imported model”. The key point in this solution for this sub-enabler is the use of URSP. URSP rules can be pre-configured on the device or be provisioned by the network to the device. In addition, a device may store a local configuration about the association of an application to a e.g. a PDU session (e.g., an operator provided S-NSSAI and DNN or application-specific parameters to set up a PDU Session). The format and contents of Local Configuration is left for specific device implementation. For further details on URSP structure and working, please refer to 3GPP TS 23.503 [23.503] and 3GPP TS 24.526 [24.526]. For a comparative analysis of different URSP rule options, see [HEX223-D62]. 3.4.2.1.2 Workflows As introduced in Section 3.4.2.1, this enabler provides three capabilities, each corresponding to a different stage: service fulfilment, service activation and service assurance. This section reports on the workflows for these stages. Figure 3-45 shows the workflow for the service fulfilment. Firstly, the DSP sends the relevant information from “3P profiling component” (see Figure 3-44) to the COP, including the following data: • Contracted services, among the service offerings available in the service portfolio. In case the contracted service includes application servers that are not managed by the DSP (i.e., 3rd party application servers deployed in Internet or in a local data network), the tenant shall specify the IP address ranges where these application servers are reachable. The tenant should also ensure these IP addresses are resolvable via existing DNS mechanisms. • Subscriber data, needed for tenant users to become subscribers of contracted services. This information, received from the tenant, includes i) the identifiers of tenant users, e.g. MSISDN or IPv6; ii) which contracted service(s) are eligible for each device to be subscribed on; and iii) other related information as to privacy and consent management, as required by EU regulation. Figure 3-45: Service fulfilment workflow. The Management Capabilities Exposure Framework receives data from the DSP and queries the mobile network to identify devices for tenant users. Based on the attributes received, it defines URSP rules, filling out Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 85 / 207 the rule attributes as defined in 3GPP TS 24.526 [24.526], and sends the information to the intra-tenant space for rule creation. Once this process is complete, it notifies the user and formats SLA values according to the SLA model. These values are then registered for later use, involving message exchanges between "imported models" and the Management Capabilities Exposure Framework. Figure 3-46 shows the service activation stage, which requires the provisioning of URSP rules to the devices of tenant users, so that the client applications can issue service requests to the application server(s), so the connectivity between service endpoints can get established according to the registered SLAs. First, the mobile network must know which URSP rules to inject on which devices. To that end, the integration fabric gets this information from the intra-tenant space (steps 1-2), and forwards it to the subscriber database (step 3), which in turn makes it available to the rest of control plane (CP) network functions from the core network and RAN (step 4). These network functions are responsible to inject applicable URSP rules into the individual devices, using the signalling interfaces defined in the mobile network for that purpose (step 5). From then on, the application client is enabled to issue service requests (step 6). Upon receiving the application request, the modem (or other device component, e.g. the Operating System) evaluates the URSP rules to use in order of priority, and proceeds to select a rule as follows (step 7): • If an URSP rule other than the default URSP rule matches the request, then the device selects this URSP rule as matching rule. • If no matching URSP rule is found, then the device selects the default URSP rule as matching rule. Finally, a matching URSP rule instructs the modem to start a PDU session establishment procedure (step 8). The modem uses the Route Selection Descriptor in the matching URSP rule to determine PDU Session connectivity parameters such as S-NSSAI (slice identifier), DNN (data network through which the application server is reachable), etc. These parameters enable the device to determine if the user traffic data can be routed through an already established PDU Session or if there is a need to trigger the establishment of a new PDU Session. Upon completion of step 8, the service gets activated, and traffic between service endpoints (application client and server) starts flowing. Figure 3-46: Service activation – workflow. In the service assurance stage (Figure 3-47), the mission is to supervise that the state of the service instance is conformant to the SLA, and take corrective actions otherwise. To that end, closed loop control (CL) instances will be used. In this regard, it is needed to insert SLA attributes into the CL model, such that the SLA becomes the goal that the CL instance must fulfil and assure. The “imported models” component oversees this (step 2), and informs the management capabilities exposure framework accordingly, specifying the CL class attributes associated to each service (step 3). With this information, the integration fabric reaches out to the CL Governance (step 4), requesting the creation of an CL class instance for every contracted service. The CL Governance completes the provisioning (step 5-6). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 86 / 207 Figure 3-47: Service assurance – workflow. 3.4.2.2 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this sub-enabler 4.2 are the following: • Availability: it measures the proportion of time a service is fully operational and accessible to users. High availability ensures that the network services are always accessible, minimizing downtime and ensuring that subscribers can rely on continuous service performance. • Reliability: it assesses the ability of the network to perform its required functions under stated conditions for a specified period of time. Reliability includes the network's capacity to handle errors, recover from failures, and maintain service continuity in the face of challenges such as hardware malfunctions or high traffic. • Maintainability: this involves the ease with which a system can be maintained in order to correct defects or their causes, improve performance or other attributes, or adapt to a changed environment. In network services, maintainability can refer to the ease of updating the system, fixing issues, and scaling services to meet changing demands. Additionally, user-centric service provisioning may include other KPIs similar to those found in existing service templates such as the Generic Network Slice Template (GST) [NG116]. These could include: • Guaranteed device throughput: this KPI measures the minimum data transfer rate that must be maintained for each device connected to the network, ensuring efficient and consistent performance. • Maximum number of PDU sessions: it represents the upper limit of simultaneous packet data unit sessions that can be handled by the network, ensuring that the network can support a high volume of data traffic without degradation in service quality. • Maximum number of registered subscribers: this is the maximum number of subscribers that can be registered and supported by the network at any given time, indicating the network's capacity. • Maximum latency: this KPI specifies the maximum allowable delay in data transmission, ensuring timely communication which is critical for applications requiring real-time interaction, such as VoIP or gaming. • Error rates: measures the frequency of errors during data transmission, which can affect the quality and reliability of network services. These KPIs are essential for monitoring the performance and health of user-centric services, allowing providers to guarantee compliance with the terms specified in SLAs and ensuring a high-quality user experience. Regarding the KVIs, the most affected by this sub-enabler is Trustworthiness. User-centric service provisioning significantly enhances the trustworthiness of future 6G networks by focusing on the specific needs and expectations of users regarding security, privacy, and reliability. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 87 / 207 3.4.3 Sub-enabler 4.3: Trust Management System 3.4.3.1 Sub-enabler design 6G networks are expected to support massive connectivity, ultra-reliability, and high mobility for IoT devices, which are typically situated within the extreme-edge domain. Within this domain, a myriad of devices with different characteristics continuously communicate and exchange information. However, they often do so over networks that are unreliable and unstable. To address this challenge, there is a pressing need for a flexible, lightweight, and adaptive access control mechanism. Such a mechanism is crucial for ensuring secure communications among trusted devices, particularly in the context of 6G’s architectural framework. Trust management system sub-enabler proposes Trust Evaluation Functions (TEFs) for assessing the trustworthiness level of entities like infrastructure components, compute nodes as well as services, applications, and 3rd party consumers. Also, this sub-enabler is planned to establish trustworthy communication paths holding sensitive data workloads as they transit a network, for data transferred between subnets, or connecting different elements that belong to different trust domains. TEFs estimate the trustworthiness level of the system of interest depending on the use case. In the case of workload placement/orchestration, a TEF focusing on the infrastructure layer is used for assessing the trustworthiness level to the infrastructure and compute nodes. In the case of flexible topologies, a TEF of flexible networks and flexible nodes is utilised. In the case of service-centric orchestration, a TEF focusing on the service layer is utilised, and so on. The TEF of the infrastructure layer is used by cloud orchestration engines to allocate workloads to compute nodes with the goal of achieving maximum trustworthiness. The TEF takes input from the monitoring telemetry data (enabler 2) of the various components of the system and triggers one of the TEFs depending on the intent request. The TEF for infrastructure layer analyses the data coming from the compute nodes and the computational workloads placed on them and outputs the trust indexes of the compute nodes of the infrastructure of interest based on the novel mathematical formula described in the following subsection. The output is fed to the functionality allocation component which along with other metrics and data suggests a close-to optimal placement of the computational workloads to the available compute nodes. This suggestion is analysed by the orchestrator which would be the responsible for enforcing the needed decision. A schematical representation of the high-level architecture of TEFs is show in Figure 3-48. Figure 3-48: High-level architecture of trust evaluation functions. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 88 / 207 3.4.3.2 Preliminary implementation and early validation results The early stage implementation outlined in this section includes a detailed description of how the trust indexes of the compute and infrastructure nodes were calculated. These indexes are used by cloud orchestration engines to allocate workloads to compute nodes targeting maximum trustworthiness. In particular, the engine used in this implementation is the functionality allocation mechanism described in section 3.6.1.2.3. The functionality allocation mechanism is a metaheuristic optimisation algorithm based on the genetic algorithm paradigm. It optimally places the computational workloads to the available compute nodes, towards maximum energy efficiency and trustworthiness by minimising the objective function consisting of an energy consumption term, an E2E latency term and a trustworthiness related term. The trust indexes are estimated by the TEF python algorithm which is described below. The described TEF is integrated to the PoC#A and PoC#B of the project, but these results are described in the respective deliverables of E2E system PoCs’ evaluations [HEX223-D21, HEX223-D22]. Here, some preliminary simulation results are presented considering fifty compute nodes with various capabilities and increasing number of workloads with various requirements. As described in Hexa-X [HEX23-D14], trustworthiness has a wide scope and is related to security, privacy [HEX23-D13] as well as availability and reliability which fall under the dependability framework [HEX21D71]. Based on this, the trust evaluation function envisioned for infrastructure layer, for now, considers the following aspects: • Availability, 𝐴𝑗, which is initially approached as the percentage of a time window during which the compute node was not fully loaded or unavailable. • Reliability, 𝑅𝑗, which is approached as the times that the compute node succeeded to execute a task/workload within a time threshold. • Security, 𝑆𝑗, which is related to secure communication and trusted computing but for this preliminary implementation it is envisioned as a binary variable indicating if Transport Layer Security (TLS) is available or not. • Multi-connectivity capabilities, 𝑀𝑗, which spans between 0 to 1 depending on which capabilities (e.g., 4G/5G, Wi-Fi, NB-IoT, BT, interfaces - device/UE-based adaptive RAT selection) are supported. • Battery level, 𝐵𝑗, of the battery powered devices/nodes (e.g., robotic units). Hence, the trust index 𝑡𝑟𝑢𝑠𝑡𝑁𝑗 of the compute node 𝑗 is calculated based on the following formula: 𝑡𝑟𝑢𝑠𝑡𝑁𝑗=𝑤1𝐴𝑗+𝑤2𝑅𝑗+𝑤3𝑆𝑗+𝑤4𝑀𝑗+𝑤5𝐵𝑗 The weights, 𝑤1,𝑤2,𝑤3,𝑤4,𝑤5 vary depending on the use case and the intent request. Figure 3-49 shows the preliminary results obtained related to trustworthiness which were calculated with the trustworthiness formula described before. In particular, the graph on the figure shows the percentage increase of trustworthiness obtained by using the metaheuristic functionality allocation mechanism (see section 3.6.1.2.3) compared to the feasible round-robin placement algorithm (baseline). The functionality allocation optimisation mechanism estimates the close to optimal placement of the computational workloads to the available compute nodes by minimising the weighted (𝑎1,𝑎2,𝑎3) objective function, min 𝑥,𝑧 (𝑎1𝐸+𝑎2𝐿−𝑎3𝑇), which consists of the energy consumption term 𝐸, the E2E latency term 𝐿, and the trustworthiness term 𝑇. The trustworthiness term is the sum of the trust indexes of the compute nodes utilised for the workload placement. These measurements were taken for different trust weight levels 𝑎3, low, medium and high. These levels correspond to an almost zero contribution, a moderate contribution and full contribution, respectively, compared to the other terms of the objective function. Fifty fixed virtualised compute nodes were used, with various capabilities and an increasing number of compute workloads with varied requirements. In particular, the compute nodes used had {1000, 1200, 1500, 2000, 2500, 2600, 8800} MIPs levels of available CPU, {2048, 4096, 8192} MB levels of available memory, {160, 360, 460} W levels of power consumption when fully loaded and {70, 100, 170} W when idle. The links between nodes have capacity 3.3 − 20 Mbps. The workloads have {250, 300, 500} MIPS levels of required CPU, {256, 512} MB levels of required memory, {10, 20} MB levels of data transferred. The weights of the trust index formula were all set to 0.2 (𝑤1,𝑤2,𝑤3,𝑤4,𝑤5). In these measurements/experiments the initial population of genetic algorithm's chromosomes was 100, the crossover rate was 0.8 and the mutation rate was 0.008-0.07. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 89 / 207 As shown in the graph, the higher the weight 𝑎3 of the trust term of the functionality allocation mechanism used, the higher the percentage of increase of the trustworthiness was obtained which was as expected. Also, the more workloads we need to place, the percentage increase of trustworthiness decreases. This is because when many workloads need to be placed, the number of feasible placement solutions decreases and the potential gains in the percentage increase of trustworthiness become limited. In general, in this scenario, the metaheuristic functionality allocation algorithm, compared to the baseline (round-robin placement), can gain up to 43% increase of trustworthiness. However, beyond these measurements, a significant benefit of this enabler is its applicability to integrate the extreme-edge domain, where devices are highly volatile. Therefore, employing a TEF to assess the trustworthiness of devices is crucial for deploying network service components within this domain. Of course, one important challenge that is under investigation is the actual utilisation of this enabler to the possible huge extreme-edge domain (cloud native) where thousands or millions of devices are to be assessed. Figure 3-49: Trustworthiness measurements of the metaheuristic functionality allocation mechanism. 3.4.3.3 Impacted KPIs and KVIs The main KPIs influenced by the adoption of this sub-enabler 4.3 are the following: • Reliability: This KPI is directly impacted by sub-enabler 4.3 as it plays a critical role in the trust evaluation functions, which are essential for accessing trust indexes. • Availability: The implementation of sub-enabler 4.3 has a direct effect on availability, again due to its involvement in the trust evaluation functions. • Scalability: Sub-enabler 4.3 impacts scalability through its close communication with enablers 5 and 6, which collectively contribute to the optimisation of resource management and functionality allocation. Regarding the KVIs, trustworthiness is the most significantly impacted due to the influence of sub-enabler 4.3. 3.5 Enabler 5: Synergetic orchestration mechanisms for the computing continuum The enabler 5 focuses on the development of synergetic orchestration mechanisms for the management of network services over programmable resources in the computing continuum. With the term computing continuum, we refer to resources that span from the IoT to the edge to the cloud part of the infrastructure. The development of distributed services that are deployed over resources in the continuum, while having strict QoS requirements in the edge (or extreme edge) part is arising in the 6G ecosystem, taking advantage of emerging technologies such as integrated communication and sensing. The enabler aims to develop different orchestration techniques that can be applicable for distributed, de-centralized and/or federated management of the available resources. Depending on the orchestration needs and the involved stakeholders, the selection of the most appropriate technique can take place. Peculiarities based on the need to manage resources that span Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 96 / 207 Figure 3-54 depicts the high-level steps performed by the REC-EXEC to start monitoring a new platform, enabling the dynamic discovery and inventory of its underlying nodes and embedded resources, along with the ones already available. When a system administrator onboards a new platform, the REC-EXEC detects the virtualization platform type and the underlying device types that are involved, and spins-up dedicated drivers that collect the platform-specific and device-specific information. For example, in case of nodes with computing resources managed by Kubernetes-based [K8S] platforms, the REC-EXEC spawns a platform driver that leverages the Kubernetes API server to collect the computing capabilities and other relevant information of the Kubernetes nodes that compose the cluster. In the case of nodes corresponding to specific types of devices with additional capabilities (e.g., drones, cobots, IoT gateways, etc.), the REC-EXEC deploys dedicated drivers to collect additional device-specific characteristics and information -beyond computing resource capabilitiesfollowing the way those devices expose them. Devices are always considered as part of a cluster, which may be limited to a single element. The driver-based approach is used to enable the dynamic discovery, continuous monitoring and inventory of the computing resources available in the nodes within the clusters, controlled through computing platforms like, e.g., Kubernetes or K3S [K3S]. Each platform, in turn, can expose one or more clusters. Moreover, extreme-edge nodes often consist of devices with additional capabilities, beyond computing resources, which could be jointly controlled (e.g., movement of a cobot, commands to an IoT actuator, etc.). This per-device information is collected with additional drivers, customized for each type of device. The implementation allows to plug and unplug specific drivers, even dynamically at runtime, depending on the scenario and on the types of involved devices. In parallel, it introduces a unified approach to handle computing resources in the continuum for multi-cluster management. The same driver-based approach is used for the execution of orchestration operations towards different types of virtualization platforms, e.g., for service provisioning, scaling, migration, etc. For example, the Kubernetes specific deployer (i.e., driver) leverages the Kubernetes API server to execute the deployment of new virtualized applications. Figure 3-55: REC-EXEC workflow for service deployment. The REC-EXEC exposes a unified and platform-agnostic set of APIs, leveraging the platform-agnostic service deployment request information model detailed in Appendix 8.1.2, to enable the deployment of applications. Thus, each deployer, when invoked for the deployment of a service component, has to translate the agnostic high-level requirements specified in the application component deployment request in the platform-specific deployment payload making also use of the orchestrator template (e.g., Helm Chart, plain Kubernetes descriptors, Heat Template, etc.) specified for that particular application component. Figure 3-55 depicts the full workflow executed by the REC-EXEC to deploy an application upon receiving a deployment request from a Service Orchestrator that has already performed the allocation of the components using the available Continuum resources retrieved from the REC-EXEC itself (i.e., the continuum resource that has been discovered from REC-EXEC drivers). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 97 / 207 3.5.1.2 Preliminary implementation and early validation results 3.5.1.2.1 REC-EXEC This section describes the preliminary implementation of the REC-EXEC multi-technology resource orchestration platform. It constitutes an example of multi-cluster resource manager, which has been developed as an evolution of the extreme-edge resource orchestrator initially implemented in the Hexa-X project [HEX23-D63]. The software architecture is depicted in Figure 3-56: the prototype is a microservices-based platform where each module, developed as a REST Java Spring Boot Application, contributes to implement the functionalities of the orchestrator. The Platform Manager software component enables the dynamic discovery, continuous monitoring of Extreme-Edge, Edge and Cloud Continuum resources, organized in multiple clusters, and the virtualized applications orchestration operations, leveraging a driver-based approach where platform-specific and devicespecific drivers operate over the underlying virtualization infrastructures and devices. These drivers are embedded in an agnostic skeleton and plugged or unplugged depending on the specific scenario. The Platform Manager includes a PostgreSQL database to store the information about the onboarded platforms (i.e., the platforms whose resources and characteristics, also in terms of devices, are discovered and monitored by the Platform Manager drivers). The Resource Manager software component retrieves the platform-specific and device specific information collected by the Platform Manager drivers from an internal Kafka message broker and exposes the discovered information (i.e., the Extreme-Edge, Edge and Cloud Continuum resources) through a REST interface. The Resource Manager includes an H2 in-memory database to store the information of the platforms discovered and monitored by the Platform Manager. Both the Platform Manager and the Resource Manager implement the information models outlined in Appendix 8.1.2. Figure 3-56: REC-EXEC implementation. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 98 / 207 The Service Deployer module enables the platform-agnostic applications orchestration operations exposing a REST interface that leverages the deployment information models outlined in Appendix 8.1.2. The Service Deployer takes care of translating the application orchestration requests in the exact formats that can be handled by the platform-specific deployer driver in the Platform Manager that can manage the target virtualization platform types. The Service Template Catalogue exposes a REST interface consumed by the Service Deployer to retrieve the platform-specific orchestration templates to be used to fulfil specific orchestration operations (e.g., a Helm Chart [HEL24] template to perform the deployment of a service component in a Kubernetes cluster managed by the REC-EXEC). The Service Template Catalogue leverages an H2 in-memory database to maintain the metadata of the orchestration templates stored in a MinIO Object Storage managed by the Service Template Catalogue software component itself. In the context of the PoC-B, the REC-EXEC has been used to discover the resources and the device characteristics of a Kubernetes cluster made of cobots (provided by one of the consortium partners, VTT) and to perform the deployment and migration of a cobot application. The deployment of the application is requested by a Service Orchestrator to the REC-EXEC targeting a specific node of the cluster chosen by the Service Orchestrator upon the resources discovered and monitored (i.e., the cobot cluster resources and battery level information of the cobots themselves) by the REC-EXEC. Additionally, a migration of the application is triggered (and executed by the REC-EXEC) by the decision and execution functions of a closed loop operating at service and infrastructure layers. This automates the migration of the application in the Extreme-Edge domain when the battery level of the cobot where the application has been originally deployed decreases beyond a certain threshold (for further details on the closed loop for service migration see section 3.8.2.1). In order to leverage the REC-EXEC in the PoC-B, the information models for resources in the continuum have been extended to represent the characteristics of the target cobots. In particular, a VTT Cobot class has been introduced as concrete extension of the abstract class Cobot (sub-class of Device Info) to maintain the peculiar information of the VTT cobot, including the battery level of the devices following the format exposed by the cobots themselves. This information has been collected by a new device driver introduced for the PoC. Early Validation Results In PoC-B, the Resource Orchestrator REC-EXEC has been deployed in the VTT testbed alongside the Monitoring Platform (section 3.2.2.3), the Closed-Loop Governance (section 3.8.2.1) and a Service Orchestrator. The testbed dedicated to the PoC-B is also providing a Kubernetes cluster made out of six nodes where two of them are cobot devices working as targets for the deployment of a network monitoring application; such application will be deployed in one of the two cobots and then moved to the other to carry on the monitoring job when the battery level of the first target goes below a certain threshold. In the described scenario, the roles of the above-mentioned services are the following: • Resource Orchestrator REC-EXEC: discovers and monitors the computing resources of the PoC-dedicated Kubernetes Cluster and the device characteristics of the cobots that are part of the cluster itself, in particular their battery information collecting the latter directly from the Monitoring Platform. Deploys and migrates the network monitoring application among the cobot Kubernetes workers. • Monitoring Platform: collects the battery level information of the cobots in order to make the latter available for historical and real-time management. The information is collected from an MQTT message broker available in the Kubernetes cluster where the cobot devices push the information regarding their batteries. • Closed-Loop Governance: responsible for the management of the lifecycle of closed-loop instances associated with a service application. In case of the PoC-B, it is responsible for the closed-loop associated with the network monitoring application, having as goal the migration of the application when the battery level of the target cobot goes below a predetermined threshold. The closed-loop functions (analysis, decision and execution) that are part of the closed-loop involved in this scenario are being deployed by the Resource Orchestrator in the PoC-B dedicated Kubernetes cluster. Their objective is to continuously monitor the battery level of the cobot chosen as target for the deployment of the network monitoring application by fetching the relevant real-time data from the Monitoring Platform and then request the Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 99 / 207 migration once the threshold is reached. The communication (i.e., event-based messaging) between the closed-loop functions is implemented leveraging a Kafka message broker. • Service Orchestrator: works as a higher-level orchestrator managing the full lifecycle of service application instances and their associated closed-loops when its instantiation is requested through its APIs. The Service Orchestrator makes use of a Resource Allocation service to determine the placements of the components of the service application and the closed-loop functions making use of the resources discovered, monitored and collected by the Resource Orchestrator. In the context of the PoC it will chose one of the two cobots for the network monitoring application and other worker nodes of the Kubernetes cluster for the three closed-loop functions. The Service Orchestrator leverages the APIs of the Resource Orchestrator for the deployment operations of the service applications and the APIs of the Closed-Loop Governance for the deployment operations of the closed-loops (i.e., that in turn makes use of the Resource Orchestrator APIs). The Service Orchestrator also exposes APIs for starting the migration of a service application instance; such APIs are used by the execution function of the closed-loop to request the migration of the network monitoring application when the threshold is reached. The deployment of the network monitoring application alongside the dedicated closed-loop (i.e., the closedloop analysis, decision and execution functions) in the PoC-B Kubernetes cluster has been tested and validated successfully, migrating the application on a threshold base from one cobot to another. The instantiation operation of the service application instance, comprehensive of the network monitoring application and the closed-loop functions, (i.e., from Service Orchestrator to Resource Orchestrator and Closed-Loop Governance) requires an average of ~36 seconds: ~20 seconds for the deployment of the network monitoring application and ~16 seconds for the deployment of the three closed-loop functions. The above-mentioned timings consider also the time needed by the network monitoring application and the closed-loop functions to be up and running in the PoC-B Kubernetes cluster (i.e., RUNNING pods’ status). The average time needed for the decision closed-loop function to check the cobot battery levels received through the analysis function (i.e., collected by the Monitoring Platform; in this scenario the analysis function works as an intermediary) and sends an event to the execution closed-loop function if the threshold is triggered, inclusive of the time needed by the execution function to receive such an event, is less than a second. Finally, the average time needed to execute the migration operation, from the reception of the event by the execution closed-loop function to the network monitoring application being up and running in the other cobot worker, is ~5 seconds. 3.5.1.2.2 RL-driven autoscaling An initial implementation of the RL-driven autoscaling mechanisms towards their integration with Component PoC#B.1 is illustrated in Figure 3-57. A setup of two clusters has been deployed as an initial environment for the application graph of the PoC. A preliminary loop instance of the RL autoscaling workflow is the main object of this description and of the corresponding experiments later on. As shown in Figure 3-57, the application designed for PoC#B.1 is deployed in a two-cluster multi-cluster setup in a lab deployment emulating the edge-cloud environment. For validating the RL agent’s functionality, a set of initial experiments were executed based on a low-high workload scenario. In specific, a low-level workload rate (frames/sec) is considered for normal operation of the application, while a high-level workload is applied when there is potential risk, so a faster detection is required. For the specific experiments, the state of the agent is constituted by the replicas deployed for the service (1-3 replicas), their placement (edge=0 or cloud=1) and the current workload. The agent is responsible for deciding the replica number suitable for the deployment and whether these replicas should be deployed at the cloud or the edge. The reward considers end-to-end latency and is calculated as -100 if the SLA is not respected or else using the following formula: R=100∗𝑆𝐿𝐴−𝑙𝑎𝑡 𝑆𝐿𝐴 The complexity of the specific problem comes from the fact that the edge server has by design low resources to support high numbers of replicas and increased parallelization may cause performance to deteriorate. Endto-end latency of a service includes communication as well as computation latency. When executing a service at the cloud introduces higher communication time, but the average computation time when more replicas (which the cloud server may be able to support) are available is decreased. Thus, the agent needs to identify Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 100 / 207 the optimal point to which to increase the replica number of the service without decreasing performance, while considering the difference in communication time in each case. Figure 3-57: RL-driven autoscaling implementation. Figure 3-58 shows the initial results from the deployment described above. The described workload is given as input in the form of a pulse, which is what is expected based on the specific scenario. The agent learns how to optimally combine the placement and scaling decisions and identifies that for low workloads, placement with 1 replica at the edge (placement=0) provides lower latency, while high workloads demonstrate additional stress and need to be deployed at the cloud (placement=1) with 3 replicas. The deployed version of the algorithm is reactive, meaning that it does not consider forecasted values of the workload and thus, demonstrates a delay in optimal decision making. This is obvious in Figure 3-58 where we see that latency is below the SLA for all cases instead of the workload level shift points where unanticipated workload increase causes latency increase, which is regulated in the next steps. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 101 / 207 Figure 3-58: RL autoscaling early results. 3.5.1.2.3 Development and communications of functionality allocation mechanism The computational workload placement functionality allocation algorithm developed is a metaheuristic algorithm based on the genetic algorithm paradigm [Yan14] as described in Section 3.6.1.2.3. The algorithm accepts computational workloads with their computation and physical demands, the volume of data, and the data’s origination points corresponding to each workload. It also considers the available compute nodes with their capabilities, trust indexes, along with the network topology graph. The algorithm’s output is a nearly optimal assignment of these computational workloads to the available compute nodes towards energy efficiency and trustworthiness. It communicates with the observability/KPI monitoring component for collecting the capabilities of the available compute nodes and the network topology graph. It also communicates with the service registry for the computation and physical requirements of the computational workloads, the data’s origination points and the size of the data corresponding to each workload. The functionality allocation algorithm additionally considers the various trust indexes of the compute and infrastructure nodes coming from the trust evaluation function (sub-enabler 4.3). The output of the algorithm is sent to the multi cluster manager which then manages the system’s resources. All these communications are controlled by the northbound interface (NBI) component named API Server. This component triggers the algorithm in the case of a possible intent coming from an end-user or when there is a need for reallocation (e.g., increased latency, malfunction in an extreme-edge component) and accepts requests from the functionality allocation component for data as well as it ensures the feasibility of the algorithm’s output before sending it to the multi cluster manager (see Figure 3-59). The data model used for this procedure is shown in Figure 8-7. In this schema, the three types of compute nodes are shown with their capabilities, some example application components are given with their requirements and the location of the generation of the utilised data. Finally, some main metrics are quoted for physical and virtual resources, network, and application layer. The described development could be extended for multi domain orchestration where the management capabilities exposure framework which collects the various available resources (application, AI, cloud, network domain) could communicate with the observability/ KPI monitoring component and the API Server and then trigger the functionality allocation mechanism if there is a need. SLA Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 102 / 207 Figure 3-59: High level view of the communications for the functionality allocation mechanism. 3.5.1.3 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this sub-enabler are the following: - Latency: the provided mechanisms for optimal placement and runtime management of network/application service graphs can lead to improvements in the end-to-end latency for the provision of a service/application. - Reliability: the continuous monitoring/observability of the performance metrics for the various deployments can lead to proactive and reactive decision making for the management of events and failures in the service provision and the infrastructure management, - Scalability: the provided mechanisms for autoscaling of the compute resources across the continuum lead to increased efficiency of the scaling decisions, considering performance, energy and cost metrics. - Programmability: the provided mechanisms manage programmable compute and network resources across the computing continuum, while the provided orchestration interfaces are fully programmable and configurable. - Automation: automation and decentralized intelligence characteristics are injected within the various orchestration mechanisms, leading to reduction of the administration overhead by network/system administrators and optimal services provision. Regarding the KVIs, it is considered that the most evident to be affected by this sub-enabler is Sustainability. The developed orchestration mechanisms can support optimal placement and lifecycle management of distributed network services and applications from an energy efficiency perspective. 3.5.2 Sub-enabler 5.2: Decentralised orchestration system This sub-enabler is based on the work regarding virtualisation and the cloud transformation studies in WP3 (6G Architecture design), and initially described in [HEX223-D32] and [HEX224-D33]. It proposes a decentralised M&O approach, targeting to integrate the broad heterogeneity of stakeholders envisaged for 6G in the M&O processes, as well as the consequent diversity of administrative and technical network domains, including the extreme-edge. The approach basically relies on two main ideas: - The deployment of multiple instances of a reduced set of network elements intended for the network resources management and the network services provisioning, which would be distributed through the entire network continuum. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 103 / 207 - The embedding of service-specific, tailor-made, and diverse M&O components as part of the network services themselves, intended for the network services assurance, which would also be distributed through the network continuum. Aligned with the cloud-native concept, this enabler tries to address one of the main challenges regarding the M&O of the network resources and the network services in the future 6G networks: the extension of the M&O scope beyond the MNO’s own domain, integrating also the extreme-edge domain, in line with the continuum orchestration concept [HEX22-D62] [KLM+22] [RCV+22]. In this regard, the integration of the extreme-edge is considered specially challenging, due, on the one hand, to its size and complexity. The extreme-edge domain can include high number of domains and devices, with a cloud native scale, and with devices integrated in different networks belonging to multiple stakeholders that could be deployed, e.g., in homes, vehicles, industrial environments, robots, satellites, smart cities, ground transport routes, agricultural fields, and more. On the other hand, the network resources in that domain can be highly heterogeneous, with many kinds of devices that could be asynchronously connected or disconnected (they won’t be necessarily in well-controlled premises), move, or with limited and/or unexpectedly changing computing or storage capabilities. However, putting this concept into practice involves more than simply offering connectivity to a variety of devices outside the MNO access network (which is already the case today, e.g., by integrating IoT devices to gather data from them), but also, to make it possible to deploy and orchestrate NS components on those “beyond the edge” infrastructure resources in a cloud-native way. The reason is that, unlike what happened in previous generations, the computing power and storage capacity of the extreme-edge devices can be significant, especially if they are considered as a whole, which could provide an additional valuable set of infrastructure resources that in many cases could be very close to the end-users, which may help reducing latency and distributing workloads, and reducing data communication needs. As its name states, this sub-enabler 5.2 tackles this challenge relying on a decentralised M&O approach. The rationale behind it is to make the system highly flexible and scalable. It is considered that it would be impractical to deal with the large number and diversity of devices on the extreme-edge domain, as well as with the large number of microservices that could be deployed on them, in a centralized manner (e.g., just the gathering and the processing of the monitoring and diagnostics data from a huge amount and diversity of resources in a centralised way could be a relevant challenge by itself). Besides, there are other problems envisaged for a possible centralised approach, e.g., the capacity planning with non-owned resources (the extreme-edge includes resources beyond the operator own premises), the well-know “single-point-of-failure” issue associated with the centralised approaches, or the increased operational costs (managing a network with many elements and mechanisms can be very costly for a single stakeholder). On the other hand, a decentralised approach offers better resources optimization, since network services may be very different in scale and complexity, so not always requiring the same kind of orchestration needs (a common MNO-centric orchestration solution would bring the same orchestration framework for all the network services, while the decentralised approach relies on tailor-made “adapted to each service” approach). Besides, as described in the following sections, the decentralised approach is “multi-domain by design”: service chaining would be performed through multiple domains relying on the service components exposed interfaces, in a cloud-native way, relying on the microservices federation concept [GKV+19] [FSP+20], which contrasts with the centralised approach that typically relies on complex business and technological agreements among different MNOs to communicate different centralized orchestration frameworks. 3.5.2.1 Sub-enabler design From a design perspective, this sub-enabler 5.2 proposes to include two main conceptual updates in the highlevel view of the well-know ETSI NFV MANO framework [NFV13] represented in Figure 3-60, as highlighted in Figure 3-61, i.e.: − To add the explicit declaration of the application scope of the ETSI NFV MANO framework which, as can be appreciated, extends through the entire network continuum, represented in the figure by the outer frame with the cloud on top, and not only in the domain of a single MNO, as in Figure 3-60. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 104 / 207 − To introduce the concept of cardinality (represented using the crow's foot convention [Eve76]) to explicitly denote that there would be multiple MANO functional blocks (which would be distributed) associated to the VNFs and the NFVI layers. Figure 3-60: ETSI NFV MANO Framework. Figure 3-61: Proposed update aligned with the ETSI NFV MANO Framework. Although at first glance it may appear to be only two minor changes, they may have relevant implications in the M&O architecture towards 6G (though still in good alignment with the abstractions of the ETSI NFV framework): on one hand the expansion of the ETSI NFV MANO framework through the whole network continuum shows that the other architectural blocks (VNFs on top and infrastructure resources at the bottom) do not necessarily have to be associated to a single stakeholder (e.g. the MNO) but distributed through different domains of the network continuum, and associated to different stakeholders. I.e., VNFs in the Virtualized Network Functions block on top would be grouped together to compose Network Services as the ETSI NFV MANO framework defines, but with the decentralised approach, those VNFs could be provided and executed by different stakeholders distributed through the whole continuum (instead of being deployed on a single MNO domain), so that both, industry and MNOs can benefit in tandem. Besides, the NFVI layer at the bottom would still represent the physical and virtualized infrastructure (also as the ETSI NFV MANO framework defines), but in this case, such infrastructure represents all the computing/storage/networking resources distributed along the entire network continuum, including the regular (5G) core and edge domains, but also, extreme-edge resources beyond the MNO own domain. On the other hand, the inclusion of a cardinality “>1” regarding the MANO block, and its implementation as part of the network continuum, indicates the possible implementation of multiple MANO blocks, which could be also distributed throughout the network continuum. However, to make this view possible in practice, in addition to these two conceptual changes, this sub-enabler also proposes other design updates in terms of implementation, namely: • The multiple MANO blocks resulting from considering that cardinality ">1" mentioned above would be implemented in two ways: − There would be one “common” decentralised (i.e., distributed through the network continuum) functionality specifically devoted to the network resources orchestration (those at the NFVI) and the network services provisioning. This functionality would be provided by the so-called Common Infrastructure Management and Services Provisioning System, which will be explained in more detail in Section 3.5.2.1.1 (System components) below. − A set of tailor-made, distributed, and diverse MANO resources embedded in the network services themselves, specially oriented to implement the necessary services assurance mechanisms for those services to which they are attached. • Beyond to what was originally proposed for the ETSI NFV MANO framework, VNFs would be used to implement not just common network infrastructure devices (e.g., routers, switches, firewalls…), but other network functions as well (e.g., AI/ML functions, monitoring functions, billing functions, etc.), either to implement the network specific functions (e.g., management functions, RAN functions…) or functions oriented to implement the service logic of those network services deployed on the network (e.g., application oriented network functions). Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 105 / 207 • The implementation would be based on modern cloud-native principles, i.e., − Instead of relying on VMs to implement VNFs (as suggested in [NFV13]), they would be primarily implemented through lightweight containerised micro-services (or Containerised Network Functions -CNFsas they are sometimes referred). However, it is also not excluded that other virtualization technologies could be used to implement these VNFs as well (e.g., legacy technologies such as VMs, or new virtualization technologies that may appear in the future). − Instead of based on strictly defined reference-points (like in [NFV13]), communication among VNFs would be also cloud-native, i.e., based on the micro-services exposed interfaces. This would make this approach “multi-domain by design”, enabling services federation through multiple network domains by relying on the micro-services exposed interfaces. The following sections provide more detailed information about the system components’ design, as well as the main envisaged workflows and interfaces. Information is also provided on a preliminary implementation of this concept (which is being integrated as part of the PoC B of the project), as well as on the most relevant KPIs and KVIs that could result positively impacted by this sub-enabler. 3.5.2.1.1 System components The application of the ETSI NFV MANO framework on the entire network continuum considers the deployment of multiple MANO blocks, with one of these blocks considered common (the so-called Common Infrastructure Management and Services Provisioning System – CIM&SPS), while the others would be tailormade, and attached to each NS to provide the specific service assurance mechanisms each service may require. The idea behind is to decouple service onboarding (which is considered a main “common” feature of the network) from service assurance (which can be addressed specifically for each service, in a decentralized way, and with an approach that in many cases may be more lightweight than in the regular MNO-centric approach). The components considered for each of type of MANO blocks are described below. Common Infrastructure Management and Services Provisioning System components The CIM&SPS is envisaged to be composed of four so-called distributed network stakeholder support services, which were already introduced in [HEX223-D32] and [HEX224-D33]). They are the following: • The Deployment Service (DS). • The Infrastructure Registry Service (IRS). • The Services Registry Service (SRS). • The Infrastructure Status Prediction Service (ISPS). Figure 3-62: Common Infrastructure Management and Services Provisioning System. Collectively, these four stakeholder support services facilitate the provisioning of new network services in the network, being implemented by four specific components associated with each service: Deployment Nodes (DN), Infrastructure Registry Nodes (IRN), Service Registry Nodes (SRN), and Infrastructure Status Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 112 / 207 a new microservice instance at runtime when an instance that was running on an extreme-edge node becomes unavailable due to the unexpected disconnection of that node. However, the work continues to improve this scenario targeting the demonstration with a more realistic network service (the low latency network service developed for PoC#B.1), and also, to trigger the orchestration action (i.e., the relocation of the network service components) based also on energy-related measurements (certain network nodes will be configured as “battery-powered”, and the VNF migration will be triggered based on the measured battery levels). Regarding (b) the focus is on the implementation of an Infrastructure Status Prediction Node (ISPN) based on AI/ML techniques, in order to provide predictions about the extreme-edge nodes availability status (those deployed in the ILE). This ISPN will provide forecasts to the service referred in (a) to trigger proactive orchestration actions, i.e., to make K3S to be able to migrate certain service components “before” the specific node on which they would be deployed is actually unavailable. While writing this document an initial design for this has been already provided, though the prototype is still under development. 3.5.2.3 Impacted KPIs and KVIs The main KPIs considered to be impacted by the adoption of this sub-enabler 5.2 are the following: • Scalability. As a distributed system, the approach presented here is intrinsically highly scalable, which is considered quite relevant regarding the integration of the extreme-edge domain. The distributed orchestration network elements make it possible to handle an increased amount of network services and workloads of different shapes and sizes, and without requiring complex centralized systems that could become bottlenecks or single points of failure, and that would need to be upgraded to be able to manage more services and resources. • Latency: The proposed approach relies on deploying service components on the whole network continuum, beyond the MNO own domain, and relying on network resources from multiple datacentres and in different geographic regions, so making it possible to move the necessary service components in close proximity to where they are actually requested by end users. This can lead to a large reduction in latency for certain time-sensitive applications, beyond to what can be done in 5G with regular edge nodes. • Flexibility, mainly in what regards the integration of vertical parties that, instead of having to adapt to an external MNO-centric orchestrator could just integrate their own service components exposing their interfaces in a cloud native way. Also, in what regards relying on an extensive diversity of infrastructure resources, which would be abstracted by the DN. • Processing Capacity, which would be considerably expanded by integrating the resources at the extremeedge domain. • Automation, which appears in different aspects of the model: the devices discovery processes and the registry of the infrastructure resources performed by the IRS would be a highly automated process. Also, the migration of the NS components through the volatile extreme-edge infrastructure nodes (which can be also performed proactively, based on the ISPS predictions). Also, the deployment of the network services from the DN: finding the most suitable infrastructure resources to deploy the network services (according to the requirements in the deployment descriptors) would be fully automated as well. • Services Creation Time is a KPI that would be of course affected by this approach, depending on the facilities and mechanisms provided by the DS. • Integrated intelligence. AI/ML algorithms could be embedded as part of the ISPS to enable proactive M&O actions. Also, as part of the service specific tailor-made MANO resources for the services assurance processes. • Reliability. The sub-enabler has been designed to specifically target the high-volatility of the resources extreme-edge domain, considering that those resources could unexpectedly vary their capabilities, move or even fully disconnect, and providing a M&O framework to specifically manage such situations. This should therefore affect the reliability of the system as a whole, and of the deployed network services. • Programmability. The proposed system is highly programmable by itself, relying on the cloud-native principles as a whole (e.g., all the network service components communicate using exposed interfaces, which enable programmability, and could be provided relying on highly automated DevOps practices). • Maintainability. The system is designed to be easily maintainable, in what regards the on-boarding/offboarding of the infrastructure resources, which would be done in a highly automated manner supported by the infrastructure discovery mechanisms associated to the IRS. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 113 / 207 • Intent expressiveness. As anticipated in [HEX223-D22], network services definition could be performed intent-based. Intents, that could be expressed even in natural language, would be translated into the lowlevel regular language used to define the service components deployment descriptors. • OPEX would be reduced for MNOs: instead of a large and complex M&O system within their premises to manage a big amount of network services and infrastructure resources, they could delegate on a wide set of external distributed resources and M&O mechanisms. Regarding KVIs, it is considered that the most evident to be affected by this sub-enabler is Sustainability in what regards the following aspects: - Extreme-edge nodes that would be already connected (so consuming Energy) but not hosting any service could be used to execute network service components, which would avoid having to put into operation new infrastructure nodes which would consume an additional amount of energy. Although the amount of energy consumed by the execution of the service components could be considered roughly the same in both cases, the baseline energy required to keep the nodes up and running would be saved. Considering the large size of the extreme-edge domain this could be a great energy saver. - Relying on the extreme-edge devices to deploy network service components would also contribute to reducing the hardware to be deployed by MNOs (or other stakeholders) in their datacentres. If certain network service components would be deployed on extreme-edge devices that are “already there” (e.g., vehicles, industrial robots, home appliances, etc.), that could contribute to reduce the amount of specific infrastructure devices in datacentres, which would lead to direct energy savings, as well as to reduce the carbon footprint associated with the manufacturing and the installation of those devices. - Energy used in data transmission would be also reduced, since certain workloads could be directly executed on edge and extreme-edge resources, without needing to transmit certain data to central datacentres (e.g., inference or certain training AI/ML algorithms could be performed right on the extremeedge nodes). 3.5.3 Sub-enabler 5.3: Federated orchestration system As the operating landscape of cloud networking becomes more disaggregated and more resource providers enter the market, it becomes more complex to establish SLAs to provide service continuity among them should the need arise. Moreover, given that it makes more business sense for the providers to modify the pricing for the use of their resources depending on demand, SLAs need to be dynamic to capture this characteristic. Using technologies like NFV and Network Slicing enables the creation of vertical-specific services tailored to meet the unique requirements of each industry sector. Vertical-specific services are transformed into network services (NFV-NS), which are then deployed and orchestrated across both local and external domains. The orchestration of services across multiple administrative domains known as federation, presents a promising opportunity for future 6G networks. Specifically, federation of services occurs when a consumer domain instantiates NFV-NSs (or parts thereof) in external provider domains, thereby orchestrating their life cycle. A promising approach to tackle these issues is leveraging distributed ledger technologies or blockchains to define and execute smart contracts. This system ensures security, transparency as well as verifiability without the need of costly and slow third-party intermediaries. 3.5.3.1 Enabler Design The key idea involves employing a permissioned blockchain where each administrative domain operates a single node within the blockchain. A unified Federation Smart Contract (SC) is deployed on the blockchain to serve as a decentralized authority, ensuring security and trust throughout the federation process. Each administrative domain running a single node operates the same instance of the Federation SC, ensuring synchronous execution of code across all nodes in the permissioned blockchain network, bolstering security and trust among participants. The design of the Federation SC is pivotal in safeguarding the privacy of sensitive information for each administrative domain while overseeing federation procedures involving all domains. To join the blockchain network, a new administrative domain must register with the Federation SC, providing its unique blockchain address and administrative information along with its service footprint. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 114 / 207 For seamless interoperability, administrative domains access the Federation SC via a REST API. Once registered, domains can participate in federation processes as consumers or providers. When a consumer domain initiates an announcement or federation offer, it is recorded as a new auction on the blockchain by the Federation SC, which then broadcasts the auction to all registered domains. To protect privacy, the consumer domain's address is concealed in the broadcast announcement, preventing passive data collection by other domains. As such, a single-blinded reverse auction mechanism is employed, where consumer domains anonymously create offer announcements, and potential provider domains bid for them. The Federation SC serves as a facilitator rather than an authority, allowing the bidding process to be closed only by the consumer domain. This empowers the consumer domain to apply selection policies freely. Periodically polling the Federation SC, the consumer domain caches bidding offers, selects a provider domain, and closes the auction. The chosen provider is recorded as the winner by the Federation SC, which notifies the selected provider and broadcasts the completion of the auction to all domains. Following this, the negotiation and acceptance phases conclude. Direct communication and information sharing occur between the consumer and selected provider domains. The federated service is deployed and integrated into the E2E service by the consumer domain, following legacy provisioning procedures. Upon deployment completion, the provider domain initiates charging for the federated service. The same permissioned blockchain network can be utilized, incorporating micropayment channels to facilitate unbiased charging records, immutable for both consumer and provider domains. 3.5.3.1.1 System components This enabler leverages the monitoring and telemetry capabilities of Enabler 2 to obtain the data that serves as part of the input to trigger the smart contract execution. It is also closely aligned with the objectives of usercentric service provisioning given that it allows for dynamic SLA creation and enhances service continuity. By eliminating the need for a human in the loop for the business process of SLA creation and policing, this enabler contributes to the real time zero touch control of Enabler 8. Figure 3-67 illustrates the design of the Digital Ledger Technology (DLT) federation system. Each cloud domain includes a node that implements a blockchain protocol that allows it to participate in the DLT network. Figure 3-67: S/W Design of DLT based federation. The resources, required to provision a service, of each cloud domain are managed by an orchestrator which determines where in the domain’s network they will be deployed. An East-West Bound interface ensures that the orchestration mechanism is governed by the DLT by enabling it to connect to its corresponding domain’s node that participates in the blockchain network. The smart contract consists of a set of conditions (based on the ad-hoc SLAs between the consumer and provider domains) that when met, trigger the federation process highlighted in Figure 3-68. 3.5.3.1.2 Workflows In order that this proposed system works as desired, the available providers have to be first registered on the blockchain. When the need arises to exploit resources of a provider distinct from the consumer domain currently offering a service, it publishes an advertisement on order to discover the candidate provider domains. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 115 / 207 This is followed by a negotiation of terms using such techniques as reverse auctioning. Once an agreement is reached, it is announced to all the bidding provider domains and the chosen one(s) proceed to provision the resources to deploy the service. The service is then deployed on the provider domains and the corresponding life-cycle management including the records for subsequent billing carried out by the consumer domain in a transparent manner over the blockchain. The sequence of the federation mechanism is depicted in Figure 3-68. Figure 3-68: Sequence diagram of service federation using DLT. 3.5.3.2 Preliminary implementation and early validation results A deployment of a preliminary version of the federation system using KVM as the VIM and Kubernetes as the orchestrator is depicted in Figure 3-69. Ethereum [ETH] has been used as the base DLT network allowing the use of smart contracts for triggering and managing the federation process. This implementation is aligned with PoC B.1 and facilitates the autonomous provisioning of the object detection service in third party domains. Figure 3-69: Initial implementation of DLT federation. The consumer initiates a transaction in the smart contract to request the deployment of a desired service with specific requirements. The provider monitors federation events and responds by creating a bid-offer transaction that includes the service price... The consumer then goes on to select the winning provider which proceeds to deploy the requested federated service. Once the deployment is successfully completed, the consumer is informed by the provider and it furnishes the latter with pertinent service information, such as the connection Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 116 / 207 parameters. A load balancer is employed to allow the deployed service to communicate to the consumer domain. These data were collected by recording the precise timestamps of transaction times as perceived by the consumer domain. Figure 3-70 depicts the average duration of the federation workflow steps. The transactions between the consumer and provider domain take longer than the service deployment which is executed only on the provider domain. The foregoing is as a result of the complex cryptographic mechanisms associated with blockchain transactions that underpin the authenticity, integrity and confidentiality of this technology. Figure 3-70: Average duration of service federation steps. 3.5.3.3 Impacted KPIs and KVIs The main KPIs considered impacted by this contribution are: • Reliability: The autonomous provisioning of a service in third-party domains ensures service continuity in the event of unforeseen roaming scenarios which in legacy systems would necessitate services to be dropped. • OPEX reduction: Given that this system provides a trustworthy mechanism through which service level agreements can be reached, executed and policed without a costly intermediary legal framework, the operational costs of network operators is greatly reduced. • Scalability: This mechanism is highly scalable given that all that is required for operators to participate in this scheme is a representative node in the blockchain, they need not establish prior agreements with every possible third-party provider as in traditional systems. • Automation: This scheme fully automates the legacy BSS based SLA establishment process and integrates it into the OSS processes related to service provisioning. 3.6 Enabler 6: AI/ML algorithms This enabler focuses on providing AI/ML-based mechanisms for the M&O to be included in the E2E system blueprint design. This enabler is split into two sub-enablers, the first one is the AI/ML-based control algorithms for sustainability which provides AI-based solutions that fulfil the objectives of improving the overall performance in terms of QoE metrics, energy efficiency, and increased zero-touch automation for reducing OPEX. The second sub-enabler is aims to improve the trustworthiness of the AI/ML based control system by protecting the AI/ML models against privacy and adversarial attacks, and by providing explainable decision outputs. The solutions developed within the enabler are part of the 6G architecture blueprint as a set of AI/MLbased solutions within the AI/ML framework for automated, sustainable and performant network control, together with the security, privacy and explainability mechanisms to ensure the trustworthiness of the system Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 117 / 207 and its decisions. The solutions can be deployed in a distributed manner, having different AI/ML models running and inferring close to or even as part of specific M&O functions or procedures. These procedures can also be part of a single layer seen in the blueprint Figure 2-2, from the network-centric application down to the actual computing infrastructure, or any combination of layers with resulting actions being performed in multiple layers. For example, a Reinforcement Learning policy for function placement and scaling can be part of an orchestrator in the network layer deploying and managing those functions. 3.6.1 Sub-enabler 6.1: AI/ML-based control algorithms for sustainability In recent years, an increasing interest has been witnessed for environmental sustainability and energy efficiency, which have become one of the main goals for future network technologies. This is particularly relevant considering increasingly complex network environments that include edge resources and virtualization overlays that make it difficult to select the appropriate management action, in particular when energy saving management actions may conflict with network and service performance targets, or with other energy saving actions. In this context, this sub-enabler aims to provide AI/ML-based control solutions that help achieve 6G networks requirements in terms of environmental sustainability and energy-saving while satisfying performance targets. The solutions developed within this sub-enabler provide AI/ML-based network control solutions with the objective of optimizing energy efficiency, to be integrated into automated network and service management procedures. Namely, energy-efficient solutions are provided to perform dynamic resource allocation for service chains and multi-domain system federation, where the decision process optimizes network and service performance metrics such as latency as well as energy related indicators such as energy consumption, carbon footprint, and energy sources. To facilitate network management automation, AI/ML is employed to translate high-level energy-related requirements into decision recommendations. On the other hand, while using powerful ML models can improve efficiency, the training process of those models consumes significant amounts of energy and computing resources. Thus, sustainable AI/ML solutions are developed where the energy consumption of AI/ML model training is monitored, and a trade-off is made between performance and energy consumption when implementing AI/ML based solutions. 3.6.1.1 Sub-enabler design This sub-enabler aims to provide AI/ML-based solutions for automating and optimizing M&O control in terms of sustainability and energy efficiency, while maintaining the service performance levels. Thus, this subenabler supports the different M&O lifecycle workflows such as service deployment, task scheduling, migration and scaling by providing the AI/ML models for making automated (zero-touch) decisions. The MLOps workflows for training, deploying and managing the AI/ML models are also supported and optimized to reduce energy consumption. Additionally, this sub-enabler interfaces with the monitoring functionalities from Enabler 2 to collect relevant data on energy consumption and service performance metrics such as latency and throughput. Figure 3-71: High-level architecture for AI/ML-based M&O. The high-level architecture for this sub-enabler is illustrated in Figure 3-71 and designed as follows: - Monitoring and data collection: The monitoring functionality interfaces with different layers and components of the architecture and collects data on the state of the network infrastructure and services. The collected metrics considered relate to service KPIs such as latency and throughput, as well as Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 118 / 207 energy consumption values. The data is provided as input to the AI/ML model to make decisions, as well as for evaluating the results of the applied decisions. - Model training: This functionality is responsible for the training of AI/ML models for performing automated M&O actions such as service deployment, migration, or scaling while optimizing energy efficiency as well as different performance metrics such as latency and throughput. The model training may be centralized on a single instance, or distributed over multiple instances or domains such as in Federated Learning. - Model inference: This component is responsible for performing inference of the trained models, and providing control actions to be performed by the M&O system. - MLOps Management Function: This functionality is responsible for orchestrating the lifecycle of the AI/ML model, and may trigger model training, data collection, and deployment on the inference instances. This functionality also aims to monitor and minimize the energy consumption of the steps in the MLOps control loop of the models. 3.6.1.2 Preliminary implementation and early validation results Multiple implementations for the sub-enabler are described below together with early validation results in the following sub-sections, where the implemented solutions aim to minimize energy and resource consumption while satisfying performance metrics. Six implementations are detailed: 1. AI/ML-based recommendation system for access node configuration. 2. Resource-efficient network function deployment in shared Edge environments. 3. Energy efficient computational workload placement. 4. Energy consumption assessment and optimization for MLOps workflows. 5. Federated Learning for decentralized resource allocation in multi-domain collaborative architectures. 6. Multi-agent Reinforcement Learning for dynamic scaling of resources. All of these except 4 and 5 are different AI/ML models that fit in the blue elements of Figure 3-71, implementing (part of) the intelligence for different M&O procedures. The other two detail improvements to the MLOps Management Function with regards to energy and other costs related to model training and model aggregation across domains. 3.6.1.2.1 ML based configuration recommendation for energy saving Despite the numerous energy-efficiency solutions already integrated into mobile networks, energy consumption continues to escalate due to the rapid expansion of both network traffic and data volumes. Research indicates [CKS+21] that further enhancements in energy efficiency can be realized by leveraging ML techniques, facilitating higher levels of automation. Through the recommendation of configuration settings applicable to base stations and other equipment, ML-based techniques enable the reduction of energy consumption in network elements without adversely affecting Quality of Experience (QoE). Through the recommendation of configuration settings applicable to base stations and other equipment, ML-based techniques enable the reduction of energy consumption in network elements without adversely affecting Quality of Experience (QoE). The focus on improving energy efficiency in networks revolves around access nodes. An access node, often referred to simply as a node, represents the relationship between connected user devices and the network elements to which those devices are linked. Configuration settings for access nodes significantly impact node energy consumption [PLD+22] and potentially influence numerous observable network performance Quality of Service (QoS) metrics. Recognizing that certain configuration settings may have varying timeframes for implementation, the necessity arises for accurate predictive models. These models should forecast when changes can be applied in advance, aiming to minimize potential disruptions to the network's operation. An optimal solution would identify numerous potential avenues to reduce energy consumption within the current functionality of network elements. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 119 / 207 Figure 3-72: End-to-end energy optimization from power system to node to network [VHI+21]. Considering these aspects, a concept for E2E energy optimization has been formulated, spanning from the power system to the nodes and the network level. Figure 3-72 depicts the concept, emphasizing the energy recommendation engine that serves as its core [VHI+21]. Implementation All data was derived from a live network The primary sources included performance management (PM) and configuration management (CM) data obtained from the base station. Energy measurements are inherently part of the PM-collected dataset, eliminating the need for deploying new hardware to gather the required information. The radio network performance counter data sets offer insights into cell performance, encompassing metrics like activity count in downlink (DL) and uplink (UL) directions, cell utilization, and units. Example of data collected includes transmit power of the base station, total energy consumption, important key KPIs such as throughput and latency. The environment is a real network with different types of traffics. As depicted in Figure 3-72, the collected data is first pre-processed to be ready for feature selection. Not all features are important for energy and KPI prediction and hence some of them are eliminated with feature engineering. Then, the solution described in the next Methodology section is applied to the final data and features where it recommends new network configurations, such as new transmit power level, and it also provides the prediction of this new recommended configurations on the network KPI, such as throughput or latency. Domain expert can also play a role in feature selection and in final network configuration to be applied to real network. To ensure no adverse impact on QoE, the models were constrained by KPIs. While additional or alternative KPIs could technically be included based on operator preferences, the five of them are selected considering their influence on energy consumption: Number of connection attempts to a cell, Average number of users in a cell, Throughput, Latency, Interference. Telecom networks are initially configured with parameters such as the number of cells and hardware unit types, but reconfigurations may occur due to issues like software glitches or hardware failures. Subtle changes and tuning over time can lead to varying energy consumption levels, either positive (less energy consumption) or negative (more energy consumption) for the same traffic volume. The CM data set that is utilized comprises numerous configuration attributes for the sector, including radio cell settings (e.g., frequency in DL and UL directions) and installed hardware types. This information enables the recommendation of multiple configuration changes simultaneously, as opposed to focusing on individual adjustments. The output from the energy recommendation engine provides a set of configuration attribute changes for a corresponding node, capturing the interplay between different nodes and configurations rather than isolated fine-tuning at a per-node level, which could impact other nodes. Methodology The high-level definition of a generative model involves deducing the generalized distribution of observation features within the constraints specified. The recommendation engine, based on a Conditional Variational Autoencoder (CVAE) generative model [SLY15], comprises multiple components, including an encoder, decoder, prediction model for target KPIs, and a prediction model for the energy consumption target (see Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 120 / 207 Figure 3-73). The encoder model compresses the representation of the raw CM dataset into a low-dimensional matrix known as the latent space. This space consists of points, where each point in a 2D latent space represents a complete CM configuration setting constrained by energy consumption and a KPI value. Given the nature of the CVAE, CM configurations with the same energy and KPI constraint categories are situated closely in the latent space. The decoder reconstructs the complete configuration file, potentially containing numerous CM attributes, from the embedded representation in the latent space. Latent variables representing the targeted KPI and energy-consumption levels are randomly selected from those representing the same category of targeted KPI and energy levels. The decoder utilizes these inputs to generate new configuration attributes. During deployment, constraints are provided as input to the CVAE model's decoder, confined by the specified energy-efficient network parameters. Figure 3-73: CVAE Structure. Experimental results By applying CVAE, a KPI versus energy trade-off curve is obtained as shown in Figure 3-74. In general, these curves tend to show that higher energy savings yield a higher negative impact on the KPI. For example, the operator can select a configuration from this map, depending on its priority. As the impact of different configurations to different KPIs is measured, the operator can decide that the KPIs coming with more valuable services are favoured, and the configuration reflecting this priority is selected. Figure 3-74: Energy savings and KPI impact according to the CVAE model. An all-encompassing strategy for energy optimization is the most effective means to attain comprehensive energy savings. This approach guarantees that advancements achieved at one level do not counteract increased energy consumption at another level. The concept developed here for E2E energy optimization relies on an energy recommendation engine fuelled by artificial intelligence. This solution holds significant potential for automation, facilitated by specific interfaces capable of directly adjusting nodes without human intervention. Moreover, it can be entirely software-based, eliminating the need for additional hardware. From the blueprint architecture point of view, the idea can sit in Pervasive Functionalities block where it can be placed at AI function module. It can communicate RAN NF through Network Layer APIs located in M & O. For example, Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 121 / 207 it can be deployed as rApp in O-RAN architecture and can communicate with other rApps or different NFs (NWDAF) to cooperate or to obtain the required information or data. 3.6.1.2.2 Resource efficient network function deployment Edge Computing enables the deployment of services closer to the end users, which reduces the communication latency, offloads the core, and allows more efficient resource management. Thus, services that are traditionally deployed on centralized servers can now be distributed over multiple Edge servers and their allocated resources can be adapted on-demand to adjust to local requirements. In this context, the same compute infrastructure can be used to host network functions and services whenever possible. To further optimize physical infrastructure usage, the physical resources of a node can be shared between the services and functions that the node hosts based on the current demand without a strict allocation scheme. However, this leads to the hosted functions and services competing for the available resources, and possible SLA violations in case of increased user demand. Therefore, an efficient resource allocation process should find a trade-off between user admission and SLA satisfaction. The solution described below allows to train a Deep Reinforcement Learning (DRL) [ADB+17] model for dynamic and intelligent network orchestration with resource allocation, user demand admission and routing in a resource-constrained edge environment with shared resources, where the objective is to maximize user demand admission and SLA satisfaction per user, while minimizing resource and energy consumption by performing service consolidation. The solution developed in this work allows to train and provide a model for resource efficient network function deployment in shared Edge environments, the model is assumed to be trained by the AI/ML framework, and used in a centralized manner by the M&O control functionality of the architecture. This model would then be deployed for inference as illustrated with the blue elements from Figure 3-71, providing actions to the placement and resource allocation functionalities. Implementation As illustrated in Figure 3-75, the developed solution consists of an architecture comprising a DRL model which has been trained using the Deep Deterministic Policy Gradient (DDPG) algorithm [LHP+16], an actor-critic algorithm that comprises an actor for providing actions, and a critic for evaluating the decisions provided by the actor. The environment for training the model is developed using a network simulator (OMNeT++) [OMN24] for deploying services on a network infrastructure comprising 5 interconnected Edge clusters of 10 nodes each with varying user demands, and observing the packet delay and drop for each of the user flows. Multiple scripts have been developed for integrating the model and the network simulator by generating the state input to be fed into the DRL agent, translating the DRL agent actions into simulation configuration and deployment files to be run by the network simulator, collecting the flow metrics from the simulator and calculating the overall reward to provide to the agent. Figure 3-75: DRL model training environment. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 128 / 207 From one hand, the granularity of the information retrieved (resources level, application level, process level…) needs to be defined, first, considering the overall measurements corresponding to the different computing resources where the MLOps workflows are running, as well as those measurements from the network, considering the communications needs among computing nodes, and the data sources for modelling, training, and executing. With these measurements, different metrics related to the overall sustainability could be computed (illustrated in Figure 3-81) in order to take a clear and detailed picture of the context where the MLOps workflow is running, allowing the decision making based on that information. Second, considering the specific measurements corresponding, for example, with: • AI/ML applications and models. • Data Management Frameworks. • Pipeline scripts (e.g. data normalization scripts). That would be used in the different phases and steps of the MLOps workflows, to obtain a detailed contribution of each process to the energy consumption, being capable to aggregate the information by stage and then, monitor and take decisions based on that information. Third, considering the specific measurements that can be made inside the ML artifacts, if a more fine-grained information to take decisions were needed. The objective is to have more fine-grained information, to take decisions, regarding the MLOps lifecycle energy consumption adding the capability to compare the performance of different models in both, training and execution (Figure 3-83). Figure 3-80 illustrates a high-level diagram with the network elements that could be used to retrieve the sustainability related information, considering the different stages of the MLOps workflows implemented by SW vendors and operators.The diagram considers the needed measurements level (related with each domain and step of the MLOps life cycle) and the relevance of having enough data for a clear view of the contributions in terms of sustainability (e.g. OS processes and common frameworks), power consumption at ML application level (MLApp represents the machine learning application/s in execution over the different computing nodes - e.g. a forecasting app that implements a model trained for execution -), and the equivalent in CO2 emissions per region (depending on electricity production scenario). In the diagram is being considered too the operators IT platform where a centralized view and management functionalities related with MLOps is conceceptually placed. To perform those measurements, a variety of energy sensors need to be deployed on each node where the different stages of the MLOps workflow are running, considering the different architectures capabilities, limitations and the additional requirement of exposing metrics calculated using normalized ontologies for each granularity layer. As well as the multi technology energy sensors, different instances of the eqCO2 agent need to be deployed by region, to add information or estimations about the distribution of energy sources. In addition to this, global modules need to be considered for the management of the sustainability metrics as well as to implement the management of the work loads, SLAs and policies. With the information retrieved from the different levels of granularity, even in a distributed scenario as the one in Figure 3-80, energy-saving decisions can be taken based on information retrieved from nodes and policies/slas defined, on the operator platform side (MLOps Sla Management & Sustainable Smart Decisions), such as moving part of the workloads to different nodes (inside or outside the same region, depending on the policies defined), or workflows re-design decisions, by comparing the efficiency of the models. Besides, it would be also possible to rely on temporal data analysis to timely start/stop different process in a more optimal way. An additional feature would be also the possibility to perform MLOps efficiency grade certifications, since all the measurements and metrics could be used to certify computing nodes, applications, or models, among others. This could be useful to compare different results and decide where/when to execute the MLOps workflows. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 129 / 207 Icons left to right: Python & Measurement libs, Scaphandre, Kepler, Grafana and Promethus Figure 3-80: High level architecture for sustainability measurements Preliminary implementation and early validation results To test some of the capabilities defined in the above paragraphs a couple of AI/ML artifacts (ML applications for base classification and clustering models) were implemented and deployed over a testbed node. Based on that, the following plots were obtained, illustrating the measurements retrieved from the sensors deployed using different levels of granularity, targeting the AI/ML models and the nodes on which they were deployed, and considering different activities related with the implemented MLOps workflow (considering training, evaluation, deploying, execution and swapping for a simple classification use case). This demonstrates the capability of the system to retrieve sustainability information with different granularity related with the different stages of a MLOps workflow. Specifically, Figure 3-81 shows an example with carbon production measurements for the computing nodes used during a particular MLOps workflow. Figure 3-82, on the other hand, shows power consumption measurements in mW, while Figure 3-83 shows the power consumption for the ML Artifacts, also during a particular MLOps workflow. Figure 3-81: Example of carbon production equivalent measurements over a node, in gCO2. Figure 3-82: Example of power consumption measurements over a node in mW. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 130 / 207 Figure 3-83: Example of power consumption measurements over different ML Artifacts. Finally, and regarding technology support, some limitations were detected while performing the testing regarding some hardware and software restrictions that impacted the capacity to perform measurements without assuming adaptations and developing new sensors. These limitations ware mainly related with the CPU architectures (x86 or ARM), GPU drivers’ availability and, in case of virtualization of the computing nodes, in the virtualization engine itself. 3.6.1.2.5 Multi-domain federated learning The overwhelming resource pre-requisites of training AI models at scale has meant that only expensive, high performance compute nodes would be self-sufficient in these kinds of applications. Networked nodes, working cooperatively as a distributed compute fabric, therefore provide a more apt solution to this problem. Network management therefore comes to the fore as an important concept in ensuring that such demanding envisioned 6G services as AIaaS can be satisfactorily handled with little impact on co-existing services. Moreover, in the contemporary landscape of artificial intelligence, the distributed nature of data poses challenges for model training. Privacy concerns emerge when sensitive data is shared, and the resource-intensive process of transporting training data over networks requires energy [AMK+18]. Decentralized Learning (DL) emerges as a solution by training models across multiple entities, collecting data locally, and aggregating models from each entity. However, DL faces the delicate challenge of accurately identifying and integrating pertinent models given that errors in this process can result in inaccurate aggregated models. Recent advancements propose an E-TREE DL architecture [YLC+21] that organizes nodes into clusters, facilitating the hierarchical aggregation of model weights. In collaborative scenarios involving multiple administrative domains collaborating on a training task, the objective is to define a DL logical training topology that determines interactions among domains. Decisions regarding optimal resource allocation (computing nodes, data sources, and computational/network resources) are made with the overarching goal of minimizing the overall energy consumption. Preliminary implementation and early validation results This optimized decentralised learning technique enhances the performance of the ML training component of the ML-Ops management function depicted in Figure 3-71. We define an administrative domain as a set of nodes (each consisting of a distinct data source and compute resources) that are individually managed and can interact with each other sharing model weights to produce a local model. The domains could be implemented as virtual machines and the compute resources as multiple containers within each. In the proposed system, multiple administrative domains collaborate within a Federated Domain Set (FDS) to collectively optimize the utilization of shared resources. The interconnection between these domains is via an End-to-End Federation Layer, administered by a Federating Orchestrator (FO) operating within each domain, as depicted in Figure 3-84. Each administrative domain maintains one dataset for training and another for testing. The latter is dynamically updated with input data collected within the domain, ensuring that it consistently reflects the current distribution of input data and accommodates any potential drift that might arise over time. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 131 / 207 Figure 3-84: Schematic of decentralised learning. To proceed with the problem and solution, a reliable way of selecting significant data for training the model is needed. To this end, we measure differences between datasets to identify unexplored samples. These measures are represented in Figure 3-85, for datasets with 500 samples (a) and 980 samples (b). Green lines are the distance between the raw datasets, treated as a matrix to compute the Frobenius distance; blue lines are the L2 distance between the class distributions of the datasets. Remark that the distances are normalized to fit in the plot. (a) (b) Figure 3-85: (a) 500 samples and (b) 980 samples. The decentralised learning system proposed here is based on [YLC+21], which is enhanced by optimizing the selection of the participant nodes in the model aggregation. In this way only nodes that contain sufficiently uncorrelated data are enjoined in model aggregation thereby avoiding the unnecessary synchronization between relatively similar nodes. 3.6.1.2.6 Multi-agent Reinforcement Learning for adaptive scaling The increased complexity of network services and application graphs deployed across the continuum does not fit well with centralised approaches such as optimization techniques or single-agent-based AI algorithms whose decision space grows exponentially with the addition of services or links. Handling resources or tasks in a decentralised manner becomes more and more crucial for the satisfaction of the service level objectives (SLOs) defined in a set of available deployment options. For implementing such a scenario, a Directed Acyclic Graph (DAG) -formed application graph with individual services communicating with each other is deployed in a multi-cluster setup. The traffic created by user requests is distributed across the graph according to the application’s structure and so, each service needs to handle different loads at each time point. The developed End to End Federation layer Local Federated Learner Training Agent Local Domain 2 Local Federated Learner Training Agent Local Domain 1 Local Federated Learner Training Agent Local Domain N … Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 132 / 207 autoscaling mechanisms make the application adaptive to the variable incoming workload by spawning the necessary number of replicas for each service to handle the corresponding load, while considering the defined optimization objectives, i.e. latency and energy requirements, and the respective resource constraints. Implementation An initial implementation of the QMIX algorithm [RSS+18] is considered for service autoscaling towards latency minimization based on a Kubernetes setup. In specific, a three-service application (as shown in Figure 3-86) is considered for deployment on a single node infrastructure with sufficient resources (16 CPUs, 32GiB RAM). For each service, the corresponding agent monitors the environment and collects all relevant information for the service. During training, the agents aggregate their outputs and provide the input for an extra layer that provides an output for the training actions. This process is referred to as centralized training. For the inference part, the local models of the agents are used for action selection (decentralized execution). The environment of each agent is defined as follows: • State space: o Workload: Request arrival rate (reqs/s) o Predicted workload: Predicted request arrival rate (reqs/s) o Service computation latency: The latency currently measured (ms) o Service CPU/memory usage (MiB) o Node CPU/memory usage (mcores) o Service pods number deployed (#) • Action space: The number of replicas spawned for the service. • Reward: The reward is calculated as a function of the consumed resources (CPU) and latency measured: 𝑅𝑒𝑤𝑎𝑟𝑑=𝑤𝑙𝑎𝑡∗𝑅𝑙𝑎𝑡(𝑙𝑎𝑡)+𝑤𝑟𝑒𝑠 ∗𝑅𝑟𝑒𝑠(𝐶𝑃𝑈) The design of the formula guarantees minimization of end-to-end latency, while considering the resource consumption constraints that may be existent. Figure 3-86: QMIX application on service auto-scaling. The algorithm showcases how different agents in the continuum controlling different interacting resources can work together towards a common goal. The idea leverages on collaborative multi-agent systems theory as well as adaptive ML to identify features of the environment that influence individual decision making in coordination with similar agents that coexist ad influence each other. This work will provide a real-world Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 133 / 207 implementation of the QMIX algorithm applied in a real testbed and service to validate its applicability in unstable and evolving network environments. The goal is for a global reward (centralised training) to guide local decision making (decentralised execution) and enabling collaboration between independent entities with separate (but dependent) objectives in the compute continuum. 3.6.1.3 Impacted KPIs and KVIs The main KPIs impacted by this sub-enabler are the following: - Energy efficiency: AI-based resource efficient network function deployment minimizes resource use through consolidation by reducing the number of used nodes, which translates into energy saving, while optimizing end-to-end latency and packet loss metrics. Energy efficient resource allocation and physical tasks scheduling also leads to improvements in the total energy consumption of the system by collecting sensing, localization, traffic and mobility pattern data as input for the AI/ML algorithm. AI-based solutions can also be used to automate and optimize post-deployment orchestration operations such as service migration or scaling to minimize energy usage. On the other hand, the energy consumption of the MLOps workflow for the different phases of model training and maintenance can be assessed and optimized by retrieving metrics of different granularities from the workflow. When using Federated Learning in a multi-domain edge architecture context, each step to convergence bears a data communication load and energy cost as the FL members exchange data. Thus, optimizing the local clustering of edge nodes reduces data communication and energy cost. - Service KPIs: Service KPIs such as the end-to-end latency and throughput are optimized through the AI-based M&O control procedures for resource allocation, task scheduling and post-deployment lifecycle management. - Automation: This enabler contributes to automating the M&O control system and reducing OPEX through full zero-touch M&O decision mechanisms using AI-based solutions which provide more efficient decisions compared to manual or default decision mechanisms. Additionally, the developed configuration recommendation system allows for increased automation of the network and service M&O by translating high-level requirements into configurations that best match the performance and energy related requirements. The main KVI impacted by this sub-enabler is sustainability, as the energy and resource efficiency targets in the different AI-based solutions lead to reduced energy consumption by the system for computing, data transmission, MLOps workflows, as well as the energy consumed for using hardware edge nodes. 3.6.2 Sub-enabler 6.2: Trustworthy AI/ML-based control algorithms The purpose of trustworthy AI/ML is to ensure that AI/ML systems are created and used in a transparent, accountable, dependable, safe, and ethical manner. There are several fundamental characteristics connected with trustworthy AI/ML [S21], including security, privacy, and explainability. The security element of trustworthy AI/ML examines the AI/ML system's design to be resilient, dependable, and safe for both individuals and society. The privacy component of trusted AI/ML strives to construct the AI/ML system in compliance with data protection laws and regulations, respecting individual privacy and protecting personal data disclosure. The third important factor connected with trustworthiness is explainability. ML models that handle difficult computational tasks with almost no human intervention are naturally complex black boxes. This has heightened the need for accountability while also raising worries about AI/ML systems' decisionmaking processes. Explainability strives to increase the transparency of black-box ML models, allowing humans to understand the decisions made by AI/ML systems. The trustworthy AI/ML-based control enabler aims to provide conceptual solutions and implementations for more robust models by protecting against adversarial attacks, reducing leakage of personal and sensitive data, and providing clear and concise explanations and justifications for the AI's reasoning or decision-making process. 3.6.2.1 Sub-enabler design This sub-enabler focuses on the development of mechanisms for improving the trustworthiness of the developed AI/ML-based control solutions by protecting the AI/ML models against adversarial attacks, data leaks and by providing explainability. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 134 / 207 Figure 3-87: High-level architecture for trustworthy AI/ML-based control. The envisioned high-level architecture for this sub-enabler is shown in Figure 3-87 and includes the following components: - Robust model training: This function is responsible for performing AI/ML model training by including adversarial attack samples in the training dataset to create AI/ML models that are more resistant against adversarial attacks. - Privacy functions: In distributed training scenarios such as Federated Learning, training data is exchanged between the components, which makes the data vulnerable to privacy attacks. The privacy functions are deployed with each distributed instance, and used to exchange data in a secure and privacy-preserving manner. - Explainability function: This function applies XRL techniques to provide human-interpretable explanations of AI/ML-based decisions within the M&O framework. The AI/ML model-level explanations also serve as fundamental components for aggregation and abstraction, potentially contributing to achieving system-level explainability. While system-level explainability is not the main objective of this function at this stage, it represents a significant area for future investigation. 3.6.2.2 Preliminary implementation and early validation results This section describes the preliminary implementation of an example of trustworthy AI/ML-based control, with particular reference to a secure AI/ML-based control for an Intent-based Management System. 3.6.2.2.1 Secure AI/ML-based control for Intent-based Management System The advent of intent-based management (IbM) in overseeing telecommunication systems within the 6G landscape underscores the importance of recognizing potential threats and risks accompanying its adoption and execution. IbM relies on automation and AI/ML to administer networks according to overarching goals and requirements outlined through intents, in a high-level and abstract manner. While AI/ML models empower IbM systems to dynamically adjust network parameters and resources to fulfil intents, these models possess inherent vulnerabilities to adversarial attacks. Such attacks have the potential to degrade the IbM system's performance and prompt erroneous decision-making, consequently impacting network performance and incurring financial costs for the operator. Thus, it is important to be aware of the vulnerability of IbM system against adversarial attacks originated through malicious component and the mitigations to decrease the chance of such attacks. In 6G networks, due to their complexity, the networks should be self-organizing and self-adaptive to maintain high performance under different circumstances. Thus, it is important to provide management solution with the knowledge about the expectations, requirements and constraints in order to enable these capabilities in the network. This knowledge can be defined by an intent for an autonomous system. The goals and the expected state can be given to the technical system through the intents in a high-level format. Thus, the network needs to take required actions to satisfy the defined expectations by the intents. According to the TM forum [TMF- Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 135 / 207 IG1253], the Intent Management Function (IMF) is responsible for handling intents and executing the intentbased operations. Information including the data gathered from the managed environment and network state, intents and their current state and expectation, as well as domain models and expert knowledge and rules are stored in knowledge base. It seems that the knowledge base is a key component in intent-based management, as it is a place to store the most important data, so it must be privacy-preserving and protected against adversarial attacks. The source of adversarial attacks in IbM systems can stem from vulnerabilities in the underlying ML models which are foreseen to be used by agents to facilitate the predictions and the decisions that are taken by the IbM system. They may also originate from other sources such as an adversary can generate various randomly constructed inputs and inject them in an intent and then observe what happens to the network. In another case, the adversary can collect system information and gain knowledge about the managed environment, as well as disclosing the ML models by inferencing input/output APIs. Adversaries may create carefully crafted inputs to cause models to take incorrect decisions. In IbM system, adversaries may exploit the vulnerabilities in the models that are used by agents to propose actions, made predictions, and evaluate actions. The adversaries may craft adversarial examples by modifying the submitted inputs to the AI model in the agents and may deceive the agents to propose wrong actions, predict wrong impacts, or make wrong evaluations. In addition, adversaries can cause the system to take action, while there is no need for such action and vice versa. Implementation To observe the effect of adversarial attacks on the generated model which are used by the agents in the IbM system, experiments are conducted using an internal network emulator that simulates end-to-end network connectivity. The Intent Management Function (IMF) prototype developed in [BJZ+22] has been used. It incorporates a knowledge base, reasoner mechanism, and a set of agents. To facilitate connectivity and execute pertinent functionalities for data and service exposure, a set of user plane and control plane network functions are established. IMF uses these exposures to keep an eye on the network and adjust as necessary. A clouddeployed conversational video service instance is set up as part of the trials. The QoE reported by the application itself based on a formulation is the primary KPI for this sort of service. As a result, the intention to target this kind of service is predicated on a need that establishes a minimum QoE threshold. In the experiments, a topology was generated in the network emulator with 2 UEs consuming the video service and downlink video traffic flowing from application to the UEs via UPF and gNB as it is shown in Figure 3-88. Data were collected under different network configurations. To achieve this, the network was configured on the run by taking arbitrary actions such as modifying the priority of the video traffic flows or maximum bit rate~(MBR) set for each UE. The data set for training a QoE prediction model consists of measurements and recent configurations of the environment that are collected from both network and the application. The data set consists of 15 different features (e.g., number of received frames, throughput). Figure 3-88:Topology in our in-house network emulator It should be noted that even though there is a QoE formulation integrated into the conversational video service application, some of the input features are found irrelevant and redundant due to the modelling of the network emulator. In practice, the QoE is estimated as QoE = f(x) + noise, where f(x) is the QoE formulation. Since the traffic from application towards UEs flow in real-time, the queue management might create noise in the data. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 136 / 207 The aim of the experiment is to show how adversarial attacks can degrade the performance of the model and impact the decisions which are taken by IbM system as well as the actions taken by the reasoner. A Neural Network was used to generate a regression model for QoE using measurements collected from both the network and the application. There are several parameters in the Neural Network to be configured which impact the performance of the model. In the experiments, the number of layers is set as 3, and the activation function is selected as Rectified Linear Unit (ReLU), the batch size is 100, and number of epochs is 15. The choice of loss function for the regression problem is also important: Mean Squared Error (MSE), which is a commonly used loss function for regression tasks, is used in the experiments. The training and test losses for a regression problem using MSE are numerical values that show how well a prediction is made by regression model. The dataset for the experiments contains 13437 samples where 80% of the data is used for training and the rest for the test. A surrogate model was also created using the same distribution of the main dataset with a different neural network architecture. Experimental results In order to show how adversarial attacks can degrade the performance of the model, a set of experiments have been conducted, testing the performance of the model against both types of attacks, white-box and black-box attacks with two different attack methods, Fast Gradient Signed Method (FGSM) and Basic Iterative Method (BIM). FGSM experiment for both types of white-box and black-box attacks. In FGSM method, the magnitude of the perturbations added to the input data is controlled by the epsilon parameter. Larger epsilon values produce more pronounced alterations in the input data, whereas smaller epsilon values provide less obvious disturbances. The selection of epsilon is important because it establishes the trade-off between the adversarial example's imperceptibility and its capacity to trick the model. In the experiments, as shown in Figure 3-89, the epsilon value is selected as 0.1 and 0.2. a - FGSM EPS=0.1 b - FGSM EPS=0.1 Figure 3-89: Adversarial attack using FGM method. BIM experiment for both types of white-box and black-box attacks. In the BIM method, the magnitude of the perturbations added to the input data is again controlled by the epsilon parameter. Alpha is the step size and commonly is set to a fraction of epsilon. The number of iterations determines how many times the perturbation is applied to the input. Two other parameters, alpha and number of iterations are respectively set as epsilon*0.2 and 15 (Figure 3-90). a - BIM EPS=0.1 b - BIM EPS=0.2 Figure 3-90: Adversarial attack using BIM method. Table 3-4 shows the MSE value obtained when adversarial attacks are applied. The value of the MSE before the attacks are applied is 0.01379. Adversarial attacks can lead to an increase in MSE if the perturbations Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 137 / 207 introduced by the attack significantly distort the input data in a way that the model’s predictions are further from the true values. It is important to note that the effect of adversarial attacks on MSE can vary based on several factors such as robustness of the model, the nature of the data, type and strength of the attack. Table 3-4: MSE values of the original model after applying adversarial attacks. Mean square error increases may be less for models that have been specifically trained or altered to be resistant to adversarial attack types. Also, the effect of adversarial assaults on MSE may also be influenced by the qualities of the input data and the model’s sensitivity to adjustments in that data. The impact of adversarial attacks on the mean square error of a machine learning model is contingent upon both their type and strength. Stronger adversarial attacks, characterized by larger distortions to the input data, often lead to higher MSE values. Furthermore, the kind of attack used—such as targeted, untargeted, black-box, or white-box—can have varying effects on the model's performance and, in turn, the MSE. In the experiment, to improve the performance and make the models resistant against adversarial attacks, adversarial training methods are employed. In the experiments, the BIM technique is used to create adversarial examples and include these samples while training the machine learning model. FGSM and BIM adversarial attack methods are applied to evaluate how much the performance of model is improved against these attacks. As expected, the loss value of the model generated using augmented training dataset is decreased, which shows that the model is more resistant to adversarial attacks. The comparison of MSE values for the trained model before and after augmenting adversarial examples and applying FGSM and BIM adversarial attacks are demonstrated in Table 3-5 and Table 3-6. Table 3-5: Comparison of the MSE values of the original and robust model after applying FGSM attack. Table 3-6: Comparison of the MSE values of the original and robust model after applying BIM attack. As it is shown in Table 3-4, the MSE of the original model on clean data should be relatively low, assuming the model has been trained well on the given task and data. Applying FGSM or BIM attacks to the original model led to an increase in MSE. In general, BIM tends to generate adversarial examples that are closer to the decision boundary of the model compared to FGSM. This iterative approach of BIM often results in more potent perturbations, which can lead to higher distortion in the input data. Consequently, we would expect BIM to result in a higher MSE compared to FGSM. The robust model trained using adversarial examples should also have a relatively low MSE on clean data similar to the original model. As shown in Table 3-5 and Table 3-6, the robust model exhibits greater resilience against FGSM and BIM attacks compared to the original model. While the MSE still increases after applying the attacks, the expected increase is smaller compared to the original model. 3.6.2.3 Conceptual solutions 3.6.2.3.1 Privacy Protection Framework for data analytics in M&O ML is expected to be used in 5G and 6G where one of the centralized frameworks responsible for ML-based analytics is Management Data Analytic Service (MDAS)/Management Data Analytic Function (MDAF) in the management and orchestration (M&O) layer. MDAS/MDAF functionality enables a service consumer to obtain management data analytics. The primary role of the MDAF in a 5G network is to perform data analytics Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 144 / 207 - Service KPIs: This enabler provides the mechanisms for creating NDTs that best reflect the state of networks, which means that advanced algorithms for network and service management can be trained in a more efficient manner, which improves their performance and decision quality. This results in improvements for the use case KPIs such as energy efficiency, latency, or bit rate. - Automation: The use of advanced models that are true copies of the infrastructure hosting NFs complements the work on closed management loops. It allows for autonomous prediction of the impact certain management actions will have, thus removing complex policy decision trees made by network operational teams. The KVIs impacted by this enabler are trustworthiness by providing a more accurate representation of the network state and thus improving the performance and accuracy of the decision mechanisms, and sustainability as the NDTAI/ML models for optimizing energy consumption. 3.8 Enabler 8: Real-time Zero-touch control loops automation and coordination system Zero-Touch Closed Loops (CL) are expected to play a primary role in the management of 6G networks. They provide the foundations to bring self-configuration, self-adaptation, and self-optimization capabilities in the network logic. This makes complex control and orchestration tasks fully autonomous in highly scalable, dynamic and multi-technology environments. The automation mechanisms introduced with CLs limit the complexity of the network operation. CLs are able to handle heterogeneous and multi-layer objectives that can reduce the infrastructure OPEX (e.g., maximizing resource utilization, reducing global energy consumption, etc.) while guaranteeing the performance and QoS/QoE level required for the services. This could be achieved through pervasive M&O functions that cooperate following a data-driven approach, often supported by AI/ML techniques and other complementary technologies like Digital Twins and Intent Management. These CL functions contribute to the global M&O procedures and they can be specialized for several scopes. They can automatically update the resource allocation at the infrastructure level, the network configuration and/or the service and application settings to guarantee service continuity, compliance with user intents and continuous re-optimization based on monitored data and target objectives. A set of concrete CL implementation examples will be presented in this section, together with a system for CL governance, implementing dynamic provisioning and configuration of several CL instances. CLs can be applied to different domains and layers, even with different scopes, e.g., with CLs specialized for a given service, a given network slice or, more broadly, for a given tenant. Multiple CLs can run in parallel, each of them with its own objectives and time scaling, but often operating over the same set of resources. CLs can be characterized by interdependencies where the actions triggered by one of them may impact the decisions of others, or actions from multiple CLs should be executed in a proper order, or more CLs can cooperate together to achieve a common objective in more efficient manner. Moreover, the decisions of concurrent CLs may be in contrast and lead to conflicts that need to be promptly identified and mitigated or resolved through suitable arbitration strategies before actuating their actions. This enabler also addresses these aspects proposing a system for the coordination of multiple CLs to guarantee a consistent end-to-end, cross-layer and crossdomain network automation. The design of the enabler (see section 3.8.1) is strongly aligned with the current work on ETSI ZSM ISG ([ZSM-009-1], [ZSM-009-2], [ZSM-009-3]) and the preliminary implementations are intended to provide concrete examples of ETSI ZSM CLs’ models and functions. 3.8.1 Enabler design The closed loop (CL) approach introduced in D6.2 [HEX223-D62] is based on four stages (Monitoring, Analysis, Decision, and Execution), which are chained together with the support of a transversal Knowledge function for the sharing of common information among the CL stages. The CL concept would be applicable at all the three layers of the Hexa-X-II E2E System Blueprint (ref. Figure 2-2 in Section 2), i.e., at the Infrastructure, Network and Application layers, with CLs that may consume monitoring data from various layers and operate with a cross-layer scope. Each CL starts from the collection and pre-processing of detailed information on target metrics, from service, network or underlying infrastructure layers. These data are then analysed to build advanced insights or predictions on relevant KPIs and/or detect anomalous conditions. In cognitive closed loops, ML techniques can adopt continuous learning mechanisms and adjust the CL behaviour Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 145 / 207 according to the results of previous actions. Processing the analysis output, the decision stage identifies the actions needed to continuously meet the CL objectives and triggers their execution (e.g., resource reconfiguration or network/service lifecycle management commands) on the entities managed by the CL. Enabling a seamless integration with the management system of the mobile network, these stages can be implemented through a composition of cloud-native functions, provisioned and orchestrated dynamically over the cloud/edge continuum of the infrastructure. The management of CL functions is handled through the CL Governance service, delivered through one or more functions in the M&O framework. The following sections provide further details regarding the architecture, functional blocks and workflows supporting this enabler. 3.8.1.1 Internal Architecture of the system components The internal architecture of a CL includes the CL functions (Monitoring, Analysis, Decision, Execution and Knowledge) and the function providing the mechanisms for its governance, i.e., the CL Governance system. In scenarios with multiple CLs, additional functions for CL coordination are also required. The functional representation of a CL can be implemented with several deployment models, e.g., with CL functions split in multiple components or, vice versa, aggregated in single elements. CLs can be specialized to guarantee different objectives or to operate over different managed entities (e.g., infrastructure resources, network slices, service instances, etc.). As such, several CL instances of different types can be instantiated dynamically to deal with the whole set of objectives defined by the operator or derived from the users’ intents. CL functions can be deployed and configured on-demand or initially pre-provisioned and re-configured at runtime. Moreover, they can be shared among CL instances as it happens for VNFs in NFV network services. The CL provisioning, lifecycle management and runtime is handled by the CL Governance function, which includes a mix of generalized and CL-specific logic and procedures. To efficiently handle the CLs diversity, multiple CL Governance functions can exist, each of them specialized for one or more CL categories. Moreover, more instances of the same CL Governance function type can run in parallel for scalability reasons, under the coordination of an upper layer entity where needed. The list of components with possible technologies for their implementation is provided in Table 3-7. The table refers to generic CLs. A survey of the particular CLs under implementation in the project is provided in section 3.8.2. Table 3-7: CL components. Component Description Possible technologies CL – Monitoring Represents the first stage of the CL and collects data required for the system automation from system itself (e.g., MDAF or an equivalent NF in 6G) or external sources. Prometheus, ELK, Grafana, Monitoring Platforms developed for enabler #2. CL – Analysis Analyses the data collected by the Monitoring stage and derives information on what is happening and or happened in the monitored systems. Can make use of information resulting from external AI/ML process e.g., belonging to the 6G system TensorFlow, KubeFlow, Seldon. CL – Decision Takes decisions based on the information derived by the Analysis service. Such decisions aim to produce possible corrective actions that would project the system towards a desired state. Can make use of information resulting from external AI/ML process e.g., belonging to the 6G system CL – Execution Enforces on the system the decision(s) taken at the decision state, if any. This can include interactions with different management functions (e.g., CSMF, NSMF), NEF, etc. Orchestrators, SDN controllers, RAN controllers, any other entity that can accept and apply a configuration in the target domain. Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 146 / 207 CL – Knowledge Stores data (e.g., configuration) used by the other stages of the CL or by other elements such as 6G system’s AI/ML functions. Information stored can also include AI/ML models to be employed at Analysis/Decision stage. For ML models: TF Serving, Model Registry in MLFlow, Seldon. For monitoring data and events: InfluxDB, MongoDB, MINIO. For configuration and status information: SQL databases. CL Governance Service Allows the management of a CL by external entities such as Orchestrators and CL Coordinators. In this regard, it exposes a number of interfaces for CL lifecycle management, configuration, operations, status, and information retrieval. OSM, Service Orchestration Platforms [HEX23-D63] CL Coordination Service Allow the overall coordination logic for groups of registered CLs. It interacts with the various CL Governance to receive notifications related to status, evolution and decisions of single CLs and to send commands for their configuration or activation. Internally, it invokes specialized coordination functions (see Table 3-8) implementing the logic associated to a group of interdependent CLs. Not applicable. CL Governance Figure 3-97 shows the high-level view of the interactions between CL Governance and CL functions (represented in the green boxes) and other potential components of a 6G architecture (represented in the blue boxes) during the provisioning of a CL instance. The picture assumes a single CL Governance Service and a single CL instance for simplicity, but it can be easily extended to scalable scenarios without changing the global principles. Figure 3-97: CL functions and CL Governance during CL provisioning. CLs associated to specific service or network slice instances are deployed during the provisioning of their “parent” element, under the coordination of their orchestrator (step 1). The creation of the CL instance is triggered through a request to the CL Governance function, which identifies the CL functions to be provisioned (or re-configured in case of pre-existing ones) and their optimal placement in the target infrastructure (step 2). The actual allocation of the CL functions could be performed through a Resource Orchestrator (see also enabler Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 147 / 207 5.1) operating at the Infrastructure layer of the Hexa-X-II Architecture Blueprint (step 3), triggered by invoking the Cloud APIs exposed by this layer. It should be noted that the CL functions, depending on their scope, managed entities and required input metrics, shall be able to interact with some entities or functions running at the Infrastructure layer, Network function layer, or Network-centric applications layer of the Hexa-X-II Architecture Blueprint. This could be enabled through the consumption of the Cloud APIs, Network APIs and Service APIs, respectively, for both monitoring and configuration purposes. The communication is supposed to be mediated through an Integration Fabric, i.e., a possible implementation of the Management Capabilities Exposure Framework (see enabler #3 in section 3.3), which provides the required mechanisms for access control, security, message brokering, tracing, etc. In case of cognitive CLs based on AI/ML techniques, the CL Governance may interact with the AI Framework to request the deployment and/or configuration of the required AI functions (step 4). These AI functions can be considered as logical components of the CL functions they assist, mostly CL Analysis or CL Decision (even if ML techniques, in principle, can be applied also to monitoring and execution processes). The ML models are assumed pre-trained and they can be selected from the ML models’ catalogue exposed by the AI Framework. However, during the CL runtime, a new training can be requested in case of poor performance of the active models. This would allow the cognitive CLs to autonomously adapt to the dynamic environment and to the new conditions generated by applying the CL actions. The target placement for the deployment of the AI functions exposing the ML model (step 5) is CL-dependent. They can either be instantiated within the AI framework in a centralized manner, or be co-located with the associated CL functions or with the Knowledge function, which may have a component dedicated to the exposure of ML models to be shared across CL functions. Figure 3-98: CL functions and CL Governance during CL runtime. Figure 3-98 shows the interactions during the service runtime, when the CL is active. The cycle starts with the collection of monitoring data (step 1) at the CL Monitoring function. The example in the picture assumes to retrieve data from monitoring functions operating at the network layer (e.g., from the NWDAF or the MDAF, or equivalent functions in 6G), but additional sources, from different layers or even external services, can be considered. The incoming data are unified, aggregated and pre-processed at the CL Monitoring function and made available for the following analysis stage (step 2). The CL Analysis function produces an output (statistics, predictions, events, etc.) that feeds the decision stage (step 3), where a set of actions is elaborated depending on the target objective and the configured policies. In case of fully autonomous CLs, the decided actions are sent to the execution stage (step 4). Here, the CL Execution function would coordinate their actuation interacting with the managed entities at the various architecture layers using the respective APIs. This is represented in the picture with step 5a for an Infrastructure-Layer CL, step 5b for a Network-Layer CL and step 5c for a Service-Layer CL. However, as further discussed in the CL Coordination section below, CLs are usually working in concurrency with other Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 148 / 207 ones and their decisions may conflict. For this reason, some CLs need to be coordinated together, evaluating the impact of the decisions taken by each of them before permitting their actuation. In this case, the CL is configured to operate in a controlled mode and the decision is not immediately transferred to the CL Execution function but notified to the CL Governance (step 3a) and, from here, to the CL Coordination service (step 3b). The CL Coordination is in charge of guaranteeing the consistency of the actions of multiple CLs, in compliance with the policies established by the operator. In this case it validates the actions proposed by the CL to identify and mitigate potential conflicts with concurrent CLs and it grants the permission to proceed with the action execution, which is triggered from the CL Governance to the CL Execution function (step 4’). It should be noted that in this stage the CL Coordinator may need to interact with other CLs and the conflict detection procedure may result in denying the execution of the proposed actions or modifying them. CL Coordination The CL coordination approach has been introduced in D6.2 [HEX223-D62], considering two main different models for the interaction between CLs. Peer-to-peer interaction allows to coordinate CLs operating at the same layer, e.g., in different domains with visibility on their own set of resources, or applied to different services to guarantee service continuity, performance or SLAs. In this case the objective is to guarantee the consistency of the end-to-end resource allocation, without incurring conflicting configurations. The hierarchical coordination model, e.g., based on vertical delegation and escalation actions, is particularly suitable for scenarios where CLs are applied to different layers. For example, a “parent” CL in charge of service automation can be assisted by multiple lower-layer CLs, where each of them receives a delegation to focus on a specific aspect of the service or a subset of its resources (e.g., the distribution of service components on the edge resources, the transport network connectivity, etc.). The interoperability and the coordination among CLs whose components are provided by different vendors can be considered as a challenge that can be partially mitigated with (i) unified and standard information models for the definition and description of CLs, CL functions and groups of CLs to be coordinated together, and (ii) standard APIs to be exposed at the level of single CL functions, CL Governance and CL Coordination. Standardization work in these aspects is still in progress (e.g., in 3GPP SA5 [28.536] and in ETSI ZSM ISG [ZSM-009-1]) and the information models and interfaces (see section 0) defined in the project and validated through PoCs can contribute to the effort. Figure 3-99: CL coordination functions. Figure 3-99 shows the system components for the coordination of multiple CLs, with a main function responsible for the overall CL Coordination Service and exposing the related external APIs. Internally, it interacts with a set of functions specialized for the various aspects of the CL coordination, e.g., the definition of the order of the actions to be executed by concurrent CLs, the coordination of CLs that can collaborate towards a joint multi-objective target, as well as the management of conflicts. The scope of the single functions is described in Table 3-8. The interaction between the CL Coordination and the single CLs is mediated through their CL Governance function. It should be noted that CLs may belong to different administrative domains. In Hexa-X-II Deliverable D6.3 Dissemination level: Public Page 149 / 207 this case, the interaction between CL Coordination and CL Governance would be based on an interface between different stakeholders, with the CL Governance function responsible to regulate the proper level of visibility on the internal status of the controlled CLs. Table 3-8: CL Coordination internal components. Component Description Concurrency Coordination Specialized coordination function responsible for deciding and regulating the execution order of the actions decided by concurrent CLs, in order to avoid temporary inconsistencies of the overall configuration. Pre-/PostExecution Coordination Specialized coordination function responsible for deciding and triggering commands (e.g., to temporarily pause one or more CL functions, to reconfigure them, etc.) before or after the execution of a CL action. This can help to avoid unexpected and undesired re-configuration decisions due e.g., long convergence times after CL actions execution. Multiobjective Coordination & Arbitration Specialized coordination function responsible for coordinating the decisions of CLs with different objectives, possibly combining them by adopting joint decision strategies. In case of contrasting objectives, this function can arbitrate among multiple CLs depending on global policies. Global Impact Assessment Specialized coordination function responsible to elaborate the global impact of the decisions coming from multiple CLs before their actual execution. This function may be assisted by network digital twins, for a preliminary evaluation of the consequences of several commands over a virtual copy of the network. Conflict Mitigation / Resolution Specialized coordination function that takes as input a detected conflict and tries to find alternative decisions to mitigate or resolve the conflict. Conflict Detection Specialized coordination function responsible for the detection of conflicting decisions taken by concurrent CLs and, possibly, the identification of the origin of the conflict. This may help in the mitigation or resolution step. 3.8.1.1.1 Closed Loop common information model The information model of a CL Descriptor shall contain the elements to: • Describe the main objectives of the CL. This shall enable the specification of a human-readable description of the CL goal and the definition of a set of statements to be either targeted or enforced by the CL. • Describe the scope of the CL in terms of the entities (e.g. infrastructure resources, services, network slices, etc.) a specific CL is able to manage. Moreover, a CL shall state the type of management operations it can perform on top of these resources. • Detail the functions which compose the CL in terms of type (e.g. analysis, decision, execution, monitoring, knowledge), description, configuration parameters required and the templates to be used for the provisioning and orchestration of the software components in a virtualized environment. • Define runtime policies which can govern the behavioural pattern that a closed loop follows. For instance, these policies may be used to determine the communication mechanisms between the functions of the closed loop, or the decision which shall be tracked in the knowledge base. • Define the reporting to be associated to the CL, for the generation of periodical or event-based reports with information on the history and the status of the given CL. A preliminary proposal for the information model of a generalized CL Descriptor is reported in Appendix 8.1.3. Starting from this template, CL-specific descriptors can be derived, e.g., specializing the CL goals depending on the particular type of CL. For example, a CL for SLA Assurance may specify the SLA objectives on the Goal Condition, while a CL for KPI Assurance may list the target KPIs, how to calculate them from a set of measurable metrics and the range of acceptable values. [Document text truncated for crawler view.]