Full text
Reinforced Fairness-Aware Multi-Agent Self-Organization for 6G Radio Access Network Orchestration Elham Hashemi Nezhad, Antonio Di Maio, Torsten Braun University of Bern, Bern, Switzerland {elham.hasheminezhad, antonio.dimaio, torsten.braun}@unibe.ch Abstract—The orchestrators’ deployment problem presents numerous challenges in 6G Network Radio Access Networks due to their large-scale, dynamic conditions, and variable user demands. Most works propose singleor hierarchical-orchestrator solutions, which offer poor resiliency, high signaling overhead, and slow adaptation to variable network dynamics. To tackle these challenges, in this work, we propose an online, datadriven, fully decentralized, Multi-Agent Reinforcement Learning (MARL)-based, self-organization orchestrator deployment system for 6G networks , which jointly optimizes the tradeoff between user throughput and fairness, based on time-varying system conditions. In the proposed approach, a flexibly variable number of decentralized, cooperative, peer self-organization agents autonomously adapt their associated orchestrator’s deployment location and activity to optimize network operation, without requiring centralized coordination. Simulations show improvements of up to 77% in user throughput and over 200% in fairness compared to Hierarchical and Single Orchestrator baselines in a broad range of realistic scenarios. Index Terms—6G Networks, Self-Organization (SO), Decentralized Orchestration, Multi-Agent Reinforcement Learning (MARL), Open-Radio Access Network (O-RAN) I. INTRODUCTION The Open-Radio Access Network (O-RAN) architecture enables the decoupling of physical and control layers, improving the scalability and adaptability of the network to different use cases, such as 5G and beyond. The O-RAN Alliance has introduced RAN Intelligent Controller (RIC) as a key architectural component that provides a centralized network abstraction, enabling operators to implement customized control functions in Radio Access Network (RAN). The RIC exists in two forms: the Non-Real-Time (Non-RT) RIC, which integrates with the network orchestrator and operates on a time scale longer than 1 s, and the Near-Real-Time (Near-RT) RIC, which manages control loops with RAN nodes on a time scale between 10 ms and 1 s [1]–[5]. Various studies [6]–[8], have explored the Controller Placement Problem (CPP) in different architectures, including Software-Defined Network (SDN) and O-RAN. However, the problem of orchestrator organization to manage controllers remains NP-hard and has received limited attention, particularly within the context of O-RAN. Some studies [9]–[11] present an additional management layer for orchestrator deployment, which brings a single point of failure and overhead communication to the centralized management. In contrast, we propose a decentralized Self-Organization (SO) approach, where orchestrators autonomously manage their placement, eliminating dependence on centralized control and enhancing adaptability through localized, real-time decisionmaking. Our design aligns with the 6G Self-Organizing Networks (SONs) vision, which seeks to overcome scalability, stability, security, and cost challenges via autonomous network restructuring. A SON is a network that autonomously configures, optimizes, and heals itself using automated mechanisms to improve performance, reduce manual intervention, and adapt to changing conditions. [12]. Building on our previous work, we proposed a decentralized orchestration model utilizing Multi-agent Reinforcement Learning (MARL) systems to tackle CPP. Our work highlighted the significance of decentralized orchestration (i.e., orchestrator domains) for deploying distributed Near-Real-Time RICs (i.e., controller domains) within the O-RAN architecture, especially when compared to centralized RAN Intelligent controller orchestration under similar network conditions to maximize user Packet Delivery Ratio (PDR). We adopt the 6G network architecture (Figure 1) which is composed of extreme edge (i.e., at User Equipments (UEs)), RAN, edge, core, and central cloud, and is built on three key components: Management and Orchestration Framework (MOF), Cloud Continuum Framework (CCF), and Artificial Intelligence and Machine Learning Framework (AIMLF). A single Master Service Orchestrator (MSO) manages multiple Distributed Service Orchestrators (DSOs), which control Network Services (NSs) consisting of Virtual Network Functions (VNFs), including controller software modules. In alignment with the O-RAN architecture, we interpret VNFs in the edge domain as RICs that manage and optimize RAN operations, and are deployed by DSOs. In this paper, we enhance our previous 6G network architecture by replacing the hierarchical orchestration model with a decentralized self-organization (SO) approach. Using a MARL system, orchestrators autonomously manage their placement, improving system resilience while maximizing user throughput and enhancing fairness. This goal is achieved through adaptive workload distribution within the core network, without relying on a centralized management layer. This paper addresses two research questions: 1) How can we optimize the problem of orchestrators’ deployment to reduce centralized overhead and maximize user
throughput even with highly dense UEs population? 2) How can 6G network bandwidth be orchestrated to ensure fair allocation among users under dynamic network conditions? To address these questions, we propose DEcentralized Reinforced RAN Intelligent Controller orchestration (DERRIC) Self-Organization orchestration through MARL in O-RAN (DERRIC-SO) to optimize the problem of orchestrator placement. This paper provides the following contributions. 1) By enabling orchestrators to autonomously manage their placement and actions using MARL, the network can adapt in real time to dynamic conditions, reduce dependence on centralized control, and enhance scalability, resilience, and responsiveness of orchestration. 2) Decentralized self-organization allows orchestrators to balance controller loads and user assignments more effectively, reducing performance disparities and enhancing fairness in throughput distribution, particularly under varying network loads and user densities. II. RELATED WORK We present a comprehensive analysis of related works, emphasizing how they address the orchestrators’ deployment problem. A. Decentralized Orchestration Several studies [13]–[15] have explored centralized orchestration approaches to enhance core network performance and coordinate management across O-RAN and SDN architectures. Centralized orchestration introduces vulnerability to a single point of failure, creates scalability bottlenecks as network demands grow, and increases control-plane latency due to centralized decision-making. In contrast, decentralized orchestration enhances network resilience and scalability by distributing control functions, thereby mitigating congestion and accommodating more connected users efficiently. B. Orchestration Organization A number of works [16]–[18] employ a hierarchical orchestration model, where a top-level orchestrator coordinates lower-level orchestrators to address the adaptability of network management and orchestration. While such external management systems can enhance security and trust, they also introduce coordination overhead, increase communication latency, and create upper-layer dependency, resulting in a single point of failure. These limitations ultimately reduce system resilience, hinder scalability, and slow real-time adaptability in dynamic network environments. C. Self-Organization Moura et al. [19] and Lyu et al. [20] propose selfoptimization strategies for orchestrator placement based on workload and multi-timescale tuning. However, their methods rely on fixed parameters and lack adaptability in highly dynamic networks. In contrast, DERRIC-SO empowers orchestrators to autonomously manage their lifecycle through Fig. 1. Proposed 6G Network Architecture TABLE I RELATED WORKS Works Orchestration Architecture Method Centralized Overhead Adaptive Intelligence Fairness Index [13]– [15] Centralized Data-driven ✓ ✓ × [9] Decentralized Blockchain ✓× × [16], [17] Decentralized Hierarchical ✓ ✓ × [19], [20] Decentralized Selfoptimization × × × Proposed Decentralized Selforganization ×✓ ✓ duplication, relocation, and termination using Reinforcement Learning (RL), enabling continuous learning and real-time decision-making. Duplicating a single orchestrator, rather than deploying three or more, allows the agent to act more efficiently and reach convergence faster due to the smaller action space. By intelligently relocating orchestrators closer to areas of demand, our method improves adaptability, balances controller workloads, and enhances fairness in user throughput. Table I categorizes the relevant works according to the proposed method, architecture, controller, and orchestration paradigms. III. SYSTEM MODEL AND PROBLEM FORMULATION A. System Model We assume that the system operates within a RAN deployment, adhering to the O-RAN specifications. The network topology is modeled as an undirected graph G= (V, E), where V={v1, . . . , v|V|}represents a set of infrastructure devices containing physical communication and processing devices such as fixed Base Stations (BSs) (e.g., gNodeBs), Multi-access Edge Computing (MEC) servers, and large cloud data centers. Moreover, we consider a set of edges E= {e1...,e|E|}representing a set of physical links between two
nodes in V, including wireless and wired links. We assume time is discretized into time steps t∈ T , where T ⊂ Ndenotes the set of all time steps during which the system operates. Each link in the link set Eis characterized by link latency Lij (t), which is determined by the physical length of the link, the device transmission rate and the congestion levels experienced during packet transmission at time t. For any pair of nodes i, j ∈V, we define fij ⊆Eas the set of edges forming the multi-hop path between nodes iand j. The end-to-end latency L(fij)(t) = P(k,l)∈fij Lkl(t)is defined as the sum of link latencies along the path fij at time step t∈ T . In this architecture, orchestrators and controllers are stateful, virtualized, and migratable software modules that can be deployed on a set of 6G network devices (i.e., cloud servers and gNodeBs). Both controller and orchestrator agents can operate on the same network topology graph G. The system model has a set O(t)of orchestrators at time step tthat partition the network into a Voronoi-like [21] set of contiguous orchestrator domains that are migrated on a M⊆Vnumber of orchestrator hosts, denoted as {m1, . . . , mM}, with the highest available infrastructure resources for orchestrator operations. These specialized hardware devices are capable of running orchestration software modules efficiently. Each orchestrator o∈ O(t)has a logically time-varying deployment location po(t)∈Mon the network topology at time step tthat optimizes the time-varying location and number through a self-organization process. Based on the current deployment location of the orchestrator, an orchestrator domain Dois created by clustering the RAN nodes in Vwith the lowest latency to the orchestrator node. Within its domain Do⊆V, each orchestrator o∈ O(t)is also responsible for deploying a time-varying set of controllers Co(t)to regulate user metrics such as transmission power Pu(t)for a set Uc(t)of UEs which are connected to BSs at time tin the controller domain Dc. We assume the UEs can move in the scenario and connect to the closest BS. In our model, there are a Knumber of BS, each with finite capacity to connect UEs to the network, where users are assumed to be mobile and may change their association with base stations over time. This capacity is governed by hardware constraints and regulatory policies that influence the allocation of physical parameters such as transmission power and bandwidth. To quantify the network performance from the user’s perspective, we model link capacity (i.e., userBS link) based on the Shannon-Hartley theorem [22], and each BS allocates the maximum achievable data rate in a communication channel subject to noise and interference to its connected users. We consider a scenario where users transmit data to the base station at the maximum capacity of their channels. Accordingly, the user data throughput Tu(t)[bit/s] at time step tis modeled using Shannon’s capacity theorem, as shown in Equation 1a. Tu(t) = Bu(t)·log21 + ρu(t) Iu(t) + N(1a) Iu(t) = K X k=1 Pk(t)·c 4πνduk(t)α (1b) Bu(t)[Hz] represents the bandwidth allocated to user u by BS, and Signal to Interference and Noise Ratio (SINR) experienced by user uis calculated by the fraction of the quality of the received signal ρu(t)by user uin the presence of the interference power Iu(t)(Equation 1b) at user ubased on the Free-Space Path Loss [23] from neighboring transmissions Pk(t)of Knumber of BSs and a function of cthe speed of light, νthe frequency of a radio wave, duk(t)the distance between BS kand user uwith the path loss exponent α > 0, and ambient thermal noise N. We also define the average total throughput as T=1 |T | Pt∈T 1 |U| Pu∈U Tu(t). To quantify the equity of resource (i.e., throughput) distribution within a subset U′⊆Uof users in the system at time t, we use the Jain Fairness Index J(U′, t)as in Equation 2. J(U′, t) = (Pu∈U′Tu(t))2 |U′|Pu∈U′Tu(t)2(2) The higher the value of the Jain Fairness Index, the more uniform the throughput allocation across the users in the considered set. We define the Global Fairness J(U, t)and the Local Fairness J(Uc, t)as the throughput fairness among all users in the system and among all users managed by controller c, respectively, at time t. We also define the Average Global Fairness as J(U) = 1 |T | Pt∈T J(U, t), the Average Local Fairness for controller cas J(Uc) = 1 |T | Pt∈T J(Uc(t), t). B. Problem Formulation Let us define the fairness sensitivity coefficient β∈[0,1] as a weighting coefficient that represents the importance of user throughput over fairness for the network policymaker. Let us assume the optimizer should exclude orchestrator placement solutions that induce a global user fairness below a minimum threshold Jmin. Let us also assume that there exists a maximum number mmax of orchestrators a network node can host due to physical limitations. We now formulate the orchestrator placement problem as a constrained optimization problem to maximize a utility function (Equation 3a), defined as a convex combination of cumulative user throughput and fairness over time, subject to system constraints (Equations 3b and 3c). maximize O(t), po(t)∈M, ∀o∈O(t),∀t∈T X t∈T βX u∈U Tu(t) + (1 −β)J(U, t)(3a) subject to J(U, t)≥Jmin,∀t∈ T (3b) X o∈O(t) 1[po(t)=m]≤mmax,∀m∈M, ∀t∈ T (3c) , where 1[q]represents the indicator function that returns 1 if predicate qis true. This ”offline” problem requires complete knowledge of the network history to be solved, requires such knowledge to be collected at a centralized location, and
requires the solver to explore an immense solution space as a combination of all possible orchestrator sets and related deployment locations, making the solution of such a problem impractical. Therefore, we convert such a problem into an online version that does not require complete network information but only historical and current state estimation to make optimal orchestrator deployment decisions with limited information. Simultaneously, we reformulate the presented problem into a decentralized version whose solution can be approximated through the collaboration among multiple agents, leveraging the computational capabilities offered by the set of network devices. IV. METHODOLOGY The main objective of DERRIC-SO is to continuously adapt the network orchestrators’ deployment based on the observed system state, so that their managed controllers’ placement maximizes a tradeoff between user throughput and fairness. We assume each orchestrator executes a decentralized selforganization RL agent, which solves a sequential decisionmaking problem (so, ao, ro)(t)by selecting a local action ao(t)∈ Ao(t)on the environment at each time step t, based on the observation so(t)∈ So(t)of the system at time t, and an associated reward function ro(t)∈ Ro(t), defined as follows. 1) State: Each orchestrator node builds a local estimation so(t)of the network state at time t(Equation 4) by gathering environment observations such as the current orchestrators’ deployment locations {po(t)}o∈O(t), the end-to-end orchestratoruser latency matrix L(fou, t)∈R|O(t)|×|Uc(t)|, the number of users managed by all controllers in the orchestrator domain {Uc(t)}c∈Co(t), and the number of managed controllers Co(t) at time t. We assume self-organization agents can periodically exchange part of the state information so(t)among themselves, such as their position po(t)∈Mand the number Co(t)of their currently managed controllers, to improve the true network state estimation. so(t)=({po}, L(fou),{Uc},Co)(t)(4) 2) Action: Each orchestrator adopts a shared policy π to optimize its placement through an action ao(t)∈ {0,...,2|M|} =Ao(t)of one of four different types: (1) Relocation or Stay, (2) Termination, and (3) Duplication (Figure 2), depending on the local state estimation so(t)and detailed here after. All agents use a shared policy for centralized coordination during training while enabling decentralized decision-making during execution, ensuring that the agents can effectively collaborate in a multi-agent environment. The following section outlines the possible actions that each agent can take. Relocation or Stay: When the orchestrator selects an action ao(t)∈ {1,...,|M|} it migrates to the target host location ao(t)∈M. If ao(t) = po(t), i.e. relocation to the current location, the orchestrator ”stays” on its current host. Service interruption is minimized through hot migration, i.e., starting the orchestrator at the target host location before terminating the instance at the previous location. This action focuses on reducing end-to-end orchestrator-user latency and maximizing user throughput and fairness in the orchestrator domain. Duplication: When the orchestrator selects an action ao(t)∈ {|M|+1,...,2|M|}, it stays on the current host location and creates a new orchestrator instance on the target host location ao(t)−|M| ∈ M. Once a new orchestrator instance is created as part of this process, and controllers are redistributed between the original and new orchestrators to balance their workload. This action improves user throughput and fairness by distributing controller management responsibilities through load balancing, leading to more responsive control plane operations and reducing resource block allocation delays. Termination: When the orchestrator selects an action ao(t)=0it terminates its operation and frees the resources. An orchestrator opts for termination when it detects that its workload is significantly below capacity, indicating inefficient resource utilization. Before terminating, it coordinates with neighboring orchestrators to ensure its controllers can be redistributed without overloading other domains or significantly increasing latencies. 3) Reward: The reward function guides each orchestrator’s decision-making process by evaluating the effectiveness of its chosen action. We define the reward function ro(t)∈ Ro(t) = R+for the self-organization agent on orchestrator o∈ O(t) as a convex combination of the user throughput and Local Fairness among the users managed by all its controllers c∈ Co(t)at time t(Equation 5). ro(t) = 1 |Co(t)|X c∈Co(t) βX u∈Uc(t) Tu(t) + (1 −β)J(Uc(t), t) (5) This reward structure encourages the orchestrator agents to improve fair user throughput in the network system. 4) RL Algorithm: We employ the Proximal Policy Optimization (PPO) algorithm under the Centralized Training, Decentralized Execution (CTDE) paradigm, using a single shared policy among all orchestrator agents to facilitate coordinated yet independent decision-making. All agents interact with the environment and collect experiences. The experience for each orchestrator agent oat tcan be denoted as Xo(t) = so(t), ao(t), ro(t), so(t+ 1),ˆ Ao(t), ωo(t), where ˆ Ao(t)is the advantage estimate for oat t, and ωo(t)is the probability ratio for agent oat t. Each agent computes the advantage estimate ˆ Ao(t) = P∞ j=0(γλ)jδt+jusing Generalized Advantage Estimator (GAE) and during centralized training, all agents’ experiences are aggregated to compute the total loss. The step index jrepresents how far into the future the advantage estimator looks from the current time step t. Here, δt= ro(t) + γF(so(t+ 1)) −F(so(t)) is the Temporal Difference (TD) residual, where F(so(t)) denotes the value function that estimates the expected return from state so(t). The parameters γ, λ ∈[0,1] are the discount rate and the GAE parameter,
Fig. 2. An example of a self-organization approach in which a single orchestrator agent can take one of four possible actions. respectively. The term ωo(t) = π(ao(t)|so(t)) πold(ao(t)|so(t)) represents the probability ratio between the current policy πand previous policy πold. The policy update uses the aggregated advantage estimates from all agents’ experiences. For the shared policy π, the clipped objective for the policy loss is calculated as Lπ= EtPo∈|O(t)|min{ωo(t)ˆ Ao(t),clip(ωo(t),1−ς, 1 + ς)ˆ Ao(t)} , where ς > 0is the clipping parameter. The loss is computed for each agent, but since the policy is shared, the gradients from all agents are aggregated to compute the final gradient update. Algorithm 1 describes how each orchestrator agent organizes itself using the RL algorithm. In the first section (lines 1 to 2), the state of the orchestrator agent ois initialized by random O(t)orchestrator nodes which are placed on M number of orchestrator hosts. The second section (lines 3 to 14) describes how the orchestrator gathers the state information from the previous time step to optimize its number and placement. It also covers reward evaluation and policy updates using the GAE [24] method. The third section (lines 15 to 22) outlines the determination of state parameters such as orchestrator deployment location, orchestrator-user latency, user-count of controllers managed by the orchestrator node, and the number of managed controllers by the orchestrator, which are measured at each time step. Lastly, the fourth section (lines 23 to 26) details the calculation of the orchestrator’s reward function. V. EXPERIMENTAL EVALUATION A. Experiment Setup We perform simulations to verify the performance of our method in terms of user throughput, Global and Local Fairness. We implemented our method and two baselines, namely a Hierarchical Orchestrator and a Single Orchestrator, in a simulated NetworkX Python environment. In the hierarchical environment, there is a master orchestrator agent to deploy local orchestrators by the single-agent RL system, and local orchestrators utilize the MARL system to deploy controllers. Moreover, the Single Orchestrator environment presents only one single agent to manage the entire network, such as the deployment of controllers by the single agent RL system. Table II gives more details of the simulation parameters. We considered the Gigabit European Advanced Network (GEANT) [25] topology with 34 infrastructure nodes that can host orchestrators and controllers, and 20 BSs. UEs are randomly placed within normalized unit square scenario, [0,1]2, and users move according to a Truncated Levy Walk model [26]. A fixed number of BSs are randomly placed on the hosts in V with the same bandwidth, coverage area, and noise. Figure 3 shows the simulation environments for DERRIC-SO and the state-of-the-art baselines. In these setups, UEs and BSs are fixed to enable a consistent performance comparison across the baselines. B. Result Performance Figure 4 shows the performance of average user throughput in increasing the number of users, in which DERRIC-SO
Fig. 3. Example of simulation environments for DERRIC-SO, Hierarchical Orchestrator and Single Orchestrator with the same number of randomly placed UEs and BSs Fig. 4. Impact of the number of users |U|on the Average User Throughput T Fig. 5. Impact of the number of users |U|on the Average Global Fairness J(U). D-SO denotes DERRIC-SO, HO denotes Hierarchical Orchestrator, and SO denotes Single Orchestrator. Fig. 6. Impact of the number of users |U|on the Average Local Fairness J(Uc) consistently outperforms both Hierarchical and Single orchestrator up to 35% and 77%, respectively. This outcome arises from optimal decentralized orchestrator domains in which orchestrators and controllers are placed optimally to distribute the workload on the nodes and efficiently manage UEs while accounting for UE mobility and density patterns. Moreover, the MARL system accelerates decision-making processes over time by leveraging a set of agents equal to the number of orchestrators. These agents operate in parallel, each independently taking actions at every time step based on local network conditions. In contrast, a single orchestrator with a single agent is limited to sequential decision-making, substantially reducing the responsiveness and effectiveness of the placement optimization process in complex network environments. Figure 5 presents the Global Fairness achieved by our proposed learning algorithm, as well as by the MARL and SO orchestrator approaches. A higher value of Jain’s index indicates a more balanced throughput allocation across increasing the number of users for three bandwidth allocation levels at 200,400, and 800 MHz dedicated to the BSs. Even at the lowest bandwidth B= 200, DERRIC-SO demonstrates superior performance with fairness values 0.50 −0.35, while Hierarchical Orchestrator achieves 0.38 −0.27 and Single Orchestrator only manages 0.21−0.07. These numerical differences remain consistent across all user scales, with DERRICSO maintaining higher absolute fairness values, demonstrating better scalability and resource allocation effectiveness compared to the state-of-the-art approaches. Figure 6 illustrates the average Local Fairness among users connected to BSs in each controller domain. Although a slight decline in fairness is observed as the number of users increases, our method maintains a relatively stable performance. In contrast, both baseline methods exhibit substantial degradation in fairness metrics as the user count escalates. This enhanced fairness can be primarily attributed to optimizing orchestrator domain boundaries and the strategic distribution
Algorithm 1: DERRIC-SO Operation // All orchestrators execute this process in parallel Data: Orchestrator set O, orchestrator hosts M, Controller set Co, User set Uc, discount rate γ, learning rate η // Reward and Policy Initialization 1R←0,π←InitializePolicy() // Initialize first orchestrator location uniformly at random 2O(t)← {o},po(1) sample ←−−−− U(M); // each episode 3for ϵ∈ E do // each time step 4for t∈ T do // Update state of orchestrator o 5so(t)← GetOrchestratorState(t, {po}, L(fou),Co,{Uc}) // Select action according to policy π 6ao(t)sample ←−−−− π(ao(t)|so(t)) // Optimize location according to action 7O(t)←SelfOrganization(ao(t)) // Collect reward 8ro(t)←GetOrchestratorReward(Uc) // Update returns 9R←ro(t) + γR // Compute TD residual 10 δt←ro(t) + γF (so(t+ 1)) −F(so(t)) // Compute advantage estimates using GAE 11 ˆ Ao(t) = P∞ j=0(γλ)jδt+j // Compute probability ratio between current and previous policy 12 ωo(t) = π(ao(t)|so(t)) πold(ao(t)|so(t)) // Update policy with PPO clipped objective 13 Lπ← min ωo(t)ˆ Ao(t),clip(ωo(t),1−ϵ, 1 + ϵ)ˆ Ao(t) // Update policy using gradient ascent with learning rate 14 π←π+η∇πLπ 15 Function GetOrchestratorState(t, {po}, L(fou),Co,{Uc}): // Collect orchestrator-user latency in the orchestrator domain 16 L(fou, t)←MeasureUserLatency(Uc) // Collect user-count of controllers in the orchestrator domain 17 Uc(t)←MeasureUserCount(Co) // Collect deployment location of orchestrator 18 po(t)←MeasureOrchestratorLocation(O) // Collect the number of controllers for each orchestrator 19 Co(t)←MeasureOrchestratorLoad(O) // Broadcast current position and controller-count to all other orchestrators 20 for x∈ O(t)in parallel do 21 sx(t)transmit ←−−−− (po(t),Co(t)) 22 return (sx, L(fou),{Uc})t 23 Function GetOrchestratorReward(t, Uc): 24 Tu(t)←MeasureUserThroughput(Uc) 25 J(Uc, t)←MeasureIndexFairness(Uc) 26 return 1 |Co(t)|Pc∈Co(t)βPu∈Uc(t)Tu(t) + (1 −β)J(Uc(t), t) TABLE II EXPERIMENT PARAMETERS Parameter Value Number of infrastructure nodes V34 Number of orchestrator hosts Mand BS K6,20 Number of users |U| {8,16,...,512,1024} Number of episodes |E| and time steps |T | 250,15000 Bandwidth of BS Band noise N400MHz,7dB Learning rate η, Discount rate γ0.0001,0.9 Batch size 64 of users across controllers, which collectively regulate usercentric performance metrics such as throughput. Figure 7 illustrates the progression of the reward training of the orchestrator agents in terms of user performance for DERRIC-SO compared to the two baseline approaches. The results demonstrate that DERRIC-SO achieves up to 18% and 102% higher rewards and faster convergence compared to state-of-the-art Hierarchical and Single Orchestrator baselines. This superior performance stems from the MARL system architecture, where multiple agents, equal in number to the orchestrator nodes, simultaneously observe the environment and optimize decisions in parallel for more efficient network management. In contrast, the Hierarchical Orchestrator relies on a sequential process in which the upper-level orchestrator, functioning as a single agent RL system, must first observe the environment and communicate with the lowerlevel orchestrators to optimize their deployment, after which the lower-level orchestrators employ an MARL approach to deploy controllers and manage the network. This additional communication overhead results in slower convergence and lower cumulative rewards for the orchestration layer. Similarly, the Single Orchestrator model employs just one agent to manage the entire network and determine the controller deployment, requiring more computational resources to learn an optimal policy for system orchestration, resulting in poorer performance and slower convergence. Figure 8 shows the mean policy loss in training episodes for three different orchestration models evaluated using the PPO algorithm with the same hyperparameters. DERRICSO consistently achieves 55% and 89% lower policy loss, indicating more stable and effective learning dynamics than the Hierarchical and Single Orchestrator models. In contrast, the Single Orchestrator shows higher and less stable policy loss, reflecting limitations in scalability and adaptability. VI. CONCLUSION This paper addresses the problem of orchestrator organization, where each orchestrator agent autonomously makes decisions, such as relocation or remaining stationary, duplication, and termination, using an MARL algorithm. This improvement stems from DERRIC-SO’s fully decentralized architecture, which eliminates the need for centralized communication at higher management layers. As a result, it enables faster
Fig. 7. Orchestrator training reward ro(t)over number of episodes containing 100 UEs and a set of controllers and orchestrators deployed over the GEANT network topology Fig. 8. Policy loss over training episodes using PPO, with 95% confidence band around the average decision-making and mitigates the risk of a single point of failure inherent in centralized approaches. Moreover, DERRICSO achieves a significantly higher and more equitable throughput distribution among users, even under varying user densities and mobility patterns. This performance aligns closely with the scalability, adaptability, and fairness requirements envisioned for next-generation 6G networks. ACKNOWLEDGMENT This work was funded by the SNS-JU 6G Cloud project under the European Union’s Horizon Europe Research and Innovation Programme under Grant Agreement No. 101139073. REFERENCES [1] M. Polese, L. Bonati, S. D’oro, S. Basagni, and T. Melodia, “Understanding O-RAN: Architecture, Interfaces, Algorithms, Security, and Research Challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 2, pp. 1376–1411, 2023. [2] X. Lin, L. Kundu, C. Dick, and S. Velayutham, “Embracing AI in 5GAdvanced toward 6G: A joint 3GPP and O-RAN Perspective,” IEEE Communications Standards Magazine, vol. 7, no. 4, pp. 76–83, 2023. [3] S. Niknam, A. Roy, H. S. Dhillon, S. Singh, R. Banerji, J. H. Reed, N. Saxena, and S. Yoon, “Intelligent O-RAN for beyond 5G and 6G Wireless Networks,” in 2022 IEEE Globecom Workshops (GC Wkshps). IEEE, 2022, pp. 215–220. [4] J. A. Ayala-Romero, A. Garcia-Saavedra, X. Costa-Perez, and G. Iosifidis, “EdgeBOL: A Bayesian Learning Approach for the Joint Orchestration of vRANs and Mobile Edge AI,” IEEE/ACM Transactions on Networking, vol. 31, no. 6, pp. 2978–2993, 2023. [5] G. M. Almeida, G. Z. Bruno, A. Huff, M. Hiltunen, E. P. Duarte, C. B. Both, and K. V. Cardoso, “RIC-O: Efficient Placement of a Disaggregated and Distributed RAN Intelligent Controller with Dynamic Clustering of Radio Nodes,” IEEE Journal on Selected Areas in Communications, 2023. [6] E. H. Bouzidi, A. Outtagarts, R. Langar, and R. Boutaba, “Dynamic clustering of software defined network switches and controller placement using deep reinforcement learning,” Computer networks, vol. 207, p. 108852, 2022. [7] G. M. Almeida, G. Z. Bruno, A. Huff, M. Hiltunen, E. P. Duarte, C. B. Both, and K. V. Cardoso, “RIC-O: Efficient Placement of a Disaggregated and Distributed RAN Intelligent Controller With Dynamic Clustering of Radio Nodes,” IEEE Journal on Selected Areas in Communications, vol. 42, no. 2, pp. 446–459, 2024. [8] A. Narwaria, K. Soni, and A. P. Mazumdar, “A Position and Energy Aware Multi-Objective Controller Placement and Re-placement Scheme in Distributed SDWSN,” The Journal of Supercomputing, pp. 1–29, 2024. [9] C. N´ u˜ nez-G´ omez, C. Carri´ on, B. Caminero, and F. M. Delicado, “SHIDRA: A Blockchain and SDN domain-based Architecture to Orchestrate fog Computing Environments,” Computer Networks, vol. 221, p. 109512, 2023. [10] B. Li, X. Deng, and Y. Deng, “Mobile-edge Computing-based Delay Minimization Controller Placement in SDN-IoV,” Computer Networks, vol. 193, p. 108049, 2021. [11] J. Baranda and J. Mangues-Bafalluy, “End-to-End Network Service Orchestration in Heterogeneous Domains for Next-Generation Mobile Networks,” in NOMS 2022-2022 IEEE/IFIP Network Operations and Management Symposium. IEEE, 2022, pp. 1–6. [12] A. Chaoub, A. M¨ ammel¨ a, P. Martinez-Julia, R. Chaparadza, M. Elkotob, L. Ong, D. Krishnaswamy, A. Anttonen, and A. Dutta, “Hybrid SelfOrganizing Networks: Evolution, Standardization Trends, and a 6G architecture vision,” IEEE Communications Standards Magazine, vol. 7, no. 1, pp. 14–22, 2023. [13] C. Valente, P. Valente, P. Rito, D. Raposo, M. Lu´ Is, and S. Sargento, “5G RAN and Core Orchestration with ML-Driven QoS Profiling,” in IEEE INFOCOM 2024-IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2024, pp. 1–6. [14] G. Z. Bruno, V. K. Radhakrishnan, G. M. Almeida, A. Huff, A. P. da Silva, K. V. Cardoso, L. A. DaSilva, and C. B. Both, “RIC-O: An Orchestrator for the Dynamic Placement of a Disaggregated RAN Intelligent Controller,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2023, pp. 1–2. [15] S. D’Oro, L. Bonati, M. Polese, and T. Melodia, “OrchestRAN: Network Automation through Orchestrated Intelligence in the Open RAN,” in IEEE INFOCOM 2022-IEEE Conference on Computer Communications. IEEE, 2022, pp. 270–279. [16] S. Kukli´ nski, R. Kołakowski, L. Tomaszewski, L. Sanabria-Russo, C. Verikoukis, C.-T. Phan, L. Zanzi, F. Devoti, A. Ksentini, C. Tselios et al., “Monb5g: Ai/ml-capable Distributed Orchestration and Management Framework for Network Slices,” in 2021 IEEE International Mediterranean Conference on Communications and Networking (MeditCom). IEEE, 2021, pp. 29–34. [17] M. A. Habib, H. Zhou, P. E. Iturria-Rivera, M. Elsayed, M. Bavand, R. Gaigalas, Y. Ozcan, and M. Erol-Kantarci, “Intent-driven Intelligent Control and Orchestration in O-RAN via Hierarchical Reinforcement Learning,” in 2023 IEEE 20th International Conference on Mobile Ad Hoc and Smart Systems (MASS). IEEE, 2023, pp. 55–61. [18] J. F. Santos, W. Liu, X. Jiao, N. V. Neto, S. Pollin, J. M. Marquez-Barja, I. Moerman, and L. A. DaSilva, “Breaking down network slicing: Hier-
archical orchestration of end-to-end networks,” IEEE Communications Magazine, vol. 58, no. 10, pp. 16–22, 2020. [19] J. Moura, “Decentralized Control Orchestration for Dynamic Edge Programmable Systems,” in 2023 3rd International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME). IEEE, 2023, pp. 1–6. [20] X. Lyu, C. Ren, W. Ni, H. Tian, R. P. Liu, and Y. J. Guo, “MultiTimescale Decentralized Online Orchestration of Software-Defined Networks,” IEEE Journal on Selected Areas in Communications, vol. 36, no. 12, pp. 2716–2730, 2018. [21] W. Qi, Y. Xia, T. Ma, L. Zhu, and J. Zhu, “ELBCFVD: An Efficient LowEnergy Balanced Clustering Algorithm Based on Fast Voronoi Division for Mobile Sensor Networks,” IEEE Sensors Journal, 2024. [22] M. E. Ekpenyong and P. J. Udoh, “Modeling the Effect of Bandwidth Allocation on Network Performance,” Science World Journal, vol. 9, no. 4, pp. 12–22, 2014. [23] M. Gao, S. Raman, Z. Sipus, and A. K. Skrivervik, “Analytic Approximation of Free-Space Path Loss for Implanted Antennas,” IEEE Open Journal of Antennas and Propagation, 2024. [24] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “Highdimensional Continuous Control using Generalized Advantage Estimation,” arXiv preprint arXiv:1506.02438, 2015. [25] J.-I. Castillo-Velazquez, I. Mu˜ noz-Mart´ ınez, J.-A. D´ ıaz-Ram´ ırez, and E. F. Ordo˜ nez-Morales, “Management Emulation for GEANT Advanced Network: 2020 Topology under IPv6,” in 2020 IEEE ANDESCON. IEEE, 2020, pp. 1–6. [26] L. Cao and M. Grabchak, “Smoothly truncated levy walks: Toward a realistic mobility model,” in 2014 IEEE 33rd International Performance Computing and Communications Conference (IPCCC). IEEE, 2014, pp. 1–8.