Full text
DERRIC: Decentralized Reinforced RAN Intelligent Controller Orchestration for 6G Networks Elham HashemiNezhad, Antonio Di Maio, Torsten Braun University of Bern, Bern, Switzerland {elham.hasheminezhad, antonio.dimaio, torsten.braun}@unibe.ch Abstract—Open-Radio Access Network (O-RAN) facilitates the scalability of cellular networks by introducing a RAN Intelligent Controller (RIC) component whose functions can be flexibly distributed over large-scale 6G networks. Artificial Intelligence (AI) is effective in optimizing RIC placement in 6G O-RAN, mitigating the limited adaptability of non-data-driven methods in complex time-varying network conditions. However, the centralized orchestration of current approaches for RIC placement hinders scalability. This work introduces a data-driven DEcentralized Reinforced RAN Intelligent Controller orchestration (DERRIC) method for 6G networks, leveraging the online learning capabilities of decentralized multi-agent Reinforcement Learning (RL) orchestration to solve the RAN Intelligent Controller Placement Problem (CPP). DERRIC is a two-layer network management scheme with decentralized orchestrators that adapt to network conditions, deploy controllers, and allocate resources. These orchestrators manage distributed controllers to optimize RAN parameters, such as user transmission power. DERRIC’s main goal is to increase the system’s overall user Packet Delivery Ratio (PDR) by optimal controller deployment and operation. Optimal controller deployment reduces controller-user latency and accelerates user-transmission-power control decisions, leading to further enhancement to user PDR. We show that DERRIC reduces the controller-user latency and power consumption by up to 66% and 29% and increases user PDR by up to 14% compared to state-of-the-art baselines in a broad range of simulated scenarios. Index Terms—6G, O-RAN, Decentralized Controller Placement, Multi-Agent RL I. INTRODUCTION The evolution toward 6G networks requires architectural changes to support diverse services and multi-connectivity coordination. O-RAN enables this transformation through virtualization and intelligence [1]. One of the main objectives of 6G is to provide AI-native management of dense, highly distributed, and potentially cell-free multi-domain networks based on O-RAN [2]. O-RAN architecture includes two RAN Intelligent Controllers (RICs) performing control and management of the network at Near-Real-Time (Near-RT) RIC between 10 ms and 1 sand Non-Real-Time (Non-RT) RIC >1s time scales [3]. Distributed controllers are necessary to enhance network performance and ensure the sustainability of user connections because a single controller poses a single point of failure. The problem of distributing several controllers at optimal locations is known as the Controller Placement Problem (CPP) [4]. Some works [5], [6] used static optimization approaches for controller placement, which cannot handle dynamic environments with time-varying network conditions due to static problem parametrization. To manage such environments, Wu et al. [7] applied Deep Q-Network (DQN) for controller placement based on deep RL to optimize a single controller, which introduces a single point of failure. Other works [8], [9] propose a centralized RIC Orchestrator (RICO) to deploy distributed controllers to enhance network performance. Centralized orchestration is a single point of failure and hinders latency and scalability. However, decentralized orchestration reduces the size and complexity of orchestration domains, improving latency, management, and network resiliency and scalability. To address the problems of centralized orchestration, Bruno et al. [10] introduced a distributed orchestration framework using a distributed cloud infrastructure to deploy disaggregated Near-RT RIC components by non-datadriven optimization, which lacks the intelligence to quickly adapt to complex and dynamic network conditions. Conversely, AI-based optimization provides flexibility, scalability, and low latency, which enables modern demanding applications in beyond-5G networks [11], [12]. Bouzidi et al. [13] proposed a decentralized approach to deploy distributed controllers using Deep Q-Network-based Dynamic Clustering and Placement (DDCP) in Software-defined Networks (SDNs), which significantly improves network response time and resource utilization. However, they apply the DQN method based on a single-agent RL to learn the whole environment, which leads to slower convergence and challenges in dividing tasks effectively. In multi-agent RL, agents autonomously learn and adapt to distinct regions of the environment and share their experiences to handle various environmental conditions more effectively [14]. In this paper, we extend our previous work [15], in which we introduced a data-driven Orchestration for Distributed RAN Intelligent Controller Placement by studying real-world computational requirements and learning convergence of multi-agent RL and perform a fine-grained parameter study on the impact of the number of User Equipments (UEs) in the system on the global transmission power. This paper addresses two research questions: 1) How can a multi-agent RL approach optimize transmission power allocation by controllers for UEs? 2) How can a decentralized orchestration framework be designed to efficiently deploy and manage distributed controllers? To address these questions, we propose DERRIC, a decentralized orchestration method that deploys distributed controllers to minimize controller-user and orchestrator-controller latencies and improve user PDR, ensuring faster and more optimal transmission power decisions by controller agents. The main contributions of this work are: 1) We use multi-agent RL approach in the RIC layer to optimize
TABLE I RELATED WORKS Works Method Architecture Controller Orchestration RL [7] DQN SDN C C ✓ [8] Dynamic Clustering O-RAN D C × [9] RIC-O O-RAN D C × [10] Dynamic Optimization O-RAN D D × [13] DDCP SDN D D ✓ DERRIC RL multiple agents O-RAN D D ✓ D: Decentralized, C: Centralized Fig. 1. Example of deployment of controllers and orchestrators over the modeled physical network infrastructure transmission power allocation by collecting user metrics such as controller-user latency and Signal-to-Noise Ratio (SNR), thereby maximizing user PDR across the network. 2) This work is the first decentralized orchestration for controller deployment in O-RAN using multi-agent RL to reduce orchestrator-controller and controller-user latencies. II. SYSTEM MODEL The system operates within a Radio Access Network (RAN) deployment adhering to O-RAN specifications, as represented in Figure 1. The network topology is modeled as an undirected graph G= (V, E), where V={v1, . . . , v|V|}represents a set of network devices. The node set Vcan contain physical communication and processing devices such as personal UEs, fixed base stations (gNodeBs), Multi-access Edge Computing (MEC) servers, and large cloud data centers. We consider a set of edges E={e1...,e|E|}representing a set of physical links between two nodes in Vthat the link set Ecan contain wireless and wired links. In this architecture, orchestrators and controllers are deployed on 6G network devices like cloud servers and gNodeBs, depending on the RL agent decision. Both controller and orchestrator agents can operate on the same network topology graph G. Each link in Eis characterized by its total link latency Lij between two devices vi, vj∈V. We assume that the link latency is defined as a one-way control latency, representing the time for control signals to travel between nodes, which varies with the transmission medium (wired or wireless). This latency is measured for control operations, such as those between orchestrators and controllers or controllers and users. In our system, we assume that the UEs in Vare connected to the 6G core network through a set of fixed base stations (e.g., gNodeBs) deployed across a geographical area, with each UE associated with the closest base station through a wireless link in E. We define PDR as Qt ifor the i-th UE at time step tas the ratio between the number of packets correctly received by the associated base station and the total number of packets transmitted by iduring the time step t. In particular, each base station in Vmust adopt an optimal transmission power Pt i[W] toward its i-th connected UEs, to maximize the Signal to Interference and Noise Ratio (SINR) at each receiver with the final goal of maximizing the average number of correctly received packets in the system. Our system contains a set C ⊆ V of controllers, which are stateful, virtualized, and migratable software modules that regulate power transmissions for UEs and can be installed on any physical devices in Vsuch as base stations and cloud servers, depending on computational and communication capabilities. Each controller c∈ C is in charge of power allocation for NcUEs associated with all base stations managed by the controller c(i.e., the controller domain). Let us define the SNR vector ρct = (ρ1c, . . . , ρNcc)t∈RNc +and the latency vector Lct = (L1c, . . . , LNcc)t∈RNc +as the vectors respectively containing the SNR of transmissions from the base station associated with each user managed by controller c, and the latency of communications from controller cto each UE managed by it at time step t. We define the average transmission power Pc tselected by the controller cfrom its managed base stations to all its managed Ncusers in time step tas Pc t=1 NcPi∈[Nc]Pt i. The DERRIC system model has a set O ⊆ Vof orchestrators that partition the network into a Voronoi-like set of contiguous orchestrator domains. Each orchestrator domain is created by clustering the RAN nodes in Vwith the lowest latency to the orchestrator o∈ O. Each orchestrator o∈ O is responsible for deploying a time-varying set of controllers Coon a set Ko⊆Vof possible deployment locations that are determined based on the orchestrator agent’s decision in its domain. We define the orchestrator user-count vector Not = (N1o, . . . , N|Co|o)tas the vector containing the number of users Nco managed by the controller c∈ Coin the set Coof controllers orchestrated by orchestrator o. We define the controller-user latency vector Lot = (L1o, . . . , L|Co|o)tas the vector containing the average latencies Lco between UEs and the controller c∈ Coin the set Coof controllers orchestrated by orchestrator o, where Lco =1 Nco Pi∈[Nco]Lic. Finally, we define the orchestratorcontroller latency vector Lot = (L1o, . . . , L|Co|o)tas the vector containing the latencies between the orchestrator oand the controller c∈ Coin the set Coof controllers orchestrated by
Fig. 2. Logical DERRIC Model orchestrator o. Figure 2 presents the logical architecture of our system based on the O-RAN framework. DERRIC focuses on the strategic placement of Near-RT RIC components representing as c∈ C. The Service Management and Orchestration (SMO) oversees the entire O-RAN architecture, utilizing the Non-RT RIC for advanced RAN optimization that we consider decentralized Orchestration represented o∈ O. Figure 2 shows that O-RAN adopts a disaggregated approach to the gNodeB, dividing it into a Central Unit with control and user plane functions (O-CU), a Distributed Unit (O-DU), and a Radio Unit (O-RU) [16]. III. DERRIC DERRIC aims at enhancing controller-user latency vector and user PDR in the transmission between the controller node and the assigned user in O-RAN. We discuss the operation of each controller and propose a multi-agent strategy to deploy distributed controllers through decentralized orchestrators. A. Controller Operation Each controller allocates the transmission power to each user in its domain by leveraging a local RL agent that observes the controller state and selects transmission power for every user managed by the controller to maximize the expected value of a reward function that considers the managed users PDR. This power allocation process can be modeled as a sequential decision-making problem (sc, ac, rc)(t), where each controller adjusts its action ac(t)∈ Ac(t)on the environment at each time step t, based on the current system’s state sc(t)∈ Sc(t) and a reward function rc(t)∈ Rc(t). We define the controller’s state, action, and reward as follows. 1) Controller State: At each time step t, every controller builds the local state sc(t)by collecting system metrics such as the latency vector Lc(t−1) and the SNR vector ρc(t−1), which contains information about the communication latency between the controller and all managed UEs and the SNR received by all managed UEs at time step t−1. The controller also collects an average transmission power vector Pt= (P1 t, . . . , P|C| t) through inter-controller connections at each time step t, which contains the latest average transmission power selected by all controllers (itself and all others) at time step t. This coordination between controllers reduces interference, where controllers allocate power to users simultaneously over the same frequency, and reduces interference between domains in the power allocation process. As the number of base stations and users Ncmanaged by the generic controller cvary over time, the dimension of the state space Sc(t)that contains the state sc(t)∈ Sc(t)is also time-varying and Sc(t)⊆R2Nc+|C|. We design the state to include user-to-controller latency because the actions that modify transmission power are adopted by the base station with a time-varying delay, which should be considered by the RL agent to select delay-predictive transmission power decisions. Equation 1 formally characterizes the c-th controller’s state sc(t)at every time step t∈N. sc(t) = Lc(t−1), ρc(t−1),Pt−1(1) 2) Controller Action: Each controller cmust determine the transmission power for each user in its control domain by executing a local controller policy πc. We define the controller agent’s action ac(t)∈ Ac(t) = [0, Pmax)Nc, where Pmax represents the maximum allowed transmission power, which is determined by each orchestrator for the controllers in its domain. Initially, all UEs are assigned equal transmission power. Subsequently, the controller agents adjust their transmission power levels based on the controller state information. Each orchestrator shares the average transmission power of its controllers Cowith other orchestrators through inter-orchestrator connections to avoid interference among orchestrator domains. 3) Controller Reward: The objective function of RL for allocating transmission power to users is to maximize PDR in users’ transmission. We define the reward rc(t)∈ Rc(t) = R+ for the generic controller cat time step tas the average of all Qt iof all UEs in the c-th controller’s domain at time step t: rc(t) = 1 NcX i∈Nc Qt i(2) Each controller agent in the multi-agent power allocation system follows a policy πcthat maps the observed state sc(t)to a transmission power action Pifor the user. Controllers exchange information on their power levels to reduce interference based on Pt−1in the state space. Policies incorporate this data to adjust actions and minimize network interference. 4) Controller Agent Algorithm: Each controller executes its local RL agent in our proposed scheme according to Algorithm 1. The first section of the algorithm (lines 1 to 2) describes initialization parameters such as random transition power allocation at time step t= 0. The second section (lines 3 to 9) explains how controller cadapts power for Ncusers and receives a reward, with the controller policy updates via the Generalized Advantage Estimator (GAE) at each time step. The third section (lines 10 to lines 14) describes how user latency and SNR are measured in the domain controller, along with the average transmission power from all controllers. The last section (lines 15 to 17) details the reward function for the Ncnumber of users calculated based on PDR. B. Orchestrator Operation We propose decentralized orchestration for controller placement and network management, where orchestrators organize their operations using RL. The orchestrator agents aim to deploy controllers while reducing user latency to affect quicker decisions on user management, such as optimized power allocation. The controller placement process can be modeled
Algorithm 1: Controller Operation // All controllers execute this process in parallel Data: Controller set C, discount rate γ // Reward, Power, and Policy Initialization 1(R, P 0 1,...,P0 Nc)←(0,...,0) 2πc←InitializePolicy() // each time step 3for t∈Ndo // Update state of controller c 4sc(t)←GetControllerState(t−1, Nc,C) // Select action according to policy πc 5ac(t)sample ←−−−−− πc(a|sc(t)) // Adapt TX power for Ncusers according to action 6(Pt 1,...,Pt Nc)←AdaptPower(ac(t)) // Collect reward 7rc(t)←GetControllerReward(Nc) // Update returns 8R←rc(t) + γR // Update policy using GAE 9πc←PolicyUpdatePPO(πc, R) 10 Function GetControllerState(t, Nc,C): // Collect user latency 11 Lct ←MeasureLatency(Nc) // Collect user SNR 12 ρct ←MeasureSNR(Nc) // Collect average power allocation from other controllers in the previous time step 13 Pt←MeasurePowerLevel(C) 14 return (Lct, ρct,Pt) 15 Function GetControllerReward(t, Nc): 16 (Qt 1,...,Qt Nc)←MeasurePacketDeliveryRatio(NC) 17 return 1 NcPi∈NcQt i as a sequential decision-making problem, where the generic orchestrator o∈ O decides the deployment of controller nodes among a set Ko⊆Vof possible deployment locations at each time step, based on the outcomes of its decisions at previous steps. The process can be represented using the tuple (so, ao, ro)(t), where each orchestrator performs ao(t)∈ Ao(t) on the environment to select controller nodes at each time step t based on the current system’s state so(t)∈ So(t)and achieve a reward function ro(t)∈ Ro(t). The orchestrator’s state, action, and reward are denoted as follows. 1) Orchestrator State: Each orchestrator builds a state so(t) by gathering the controller-user latency vector Lo(t−1) and the orchestrator user-count vector No(t−1) at time step t−1. This observation provides the orchestrator with more information to deploy controllers, adapting to the real-time demands of users within its domain. Each agent observes the latency between the orchestrator and all managed controllers in the previous time step as orchestrator-controller latency vector Lo(t−1) to deploy controllers in the possible lowest orchestrator-controller latency at time step t. Each orchestrator also collects the previous-timestep number of controllers managed by any other orchestrators o∈ O at time step t−1through inter-orchestrator connections to build the vector Ct−1= (|C1|,...,|C|O||)t−1∈N|O|. This exchanged information distributes the controller-management workload among orchestrators. Equation 3 describes the o-th orchestrator’s state so(t)at every time step t∈N. so(t) = (Lo(t−1), No(t−1), Lo(t−1),Ct−1)(3) 2) Orchestrator Action: The action ao(t)for a single orchestrator agent is to select controller nodes in the orchestrator domain as a binary decision by executing a local orchestrator policy πo. The agent selects a node with less average user latency, more number of users, and less latency between the node and the orchestrator as a controller and gets 1; otherwise, it gets 0 for non-controllers. The orchestrator agent’s action is defined as ao(t)∈ {0,1}|Ko|, which is a logical value representing controller and non-controller nodes. 3) Orchestrator Reward: The RL agent for each orchestrator in the process of controller placement aims to minimize the latency between users and their controllers managed by the orchestrator node and the latency between the controllers and their orchestrator. Therefore, we define the reward function ro(t)for each orchestrator (Equation 4) as the negative sum of the norm of the average controller-user latency vector and the norm of the orchestrator-controller latency vector. ro(t) = −∥Lot∥−∥Lot∥(4) Each orchestrator agent in the multi-agent controller placement system follows a policy πothat maps the observed state so(t) to select a controller as an action in the orchestrator domain. Orchestrators as agents exchange information on their workload to make workload balanced among themselves based on Ct−1in the state space. Policies incorporate this data to adjust actions and minimize network interference. 4) Orchestrator Agent Algorithm: Algorithm 2 describes how each orchestrator agent selects the controller nodes using the RL algorithm. In the first section (lines 1 to 2), the state of the orchestrator agent ois initialized by random Cocontroller deployment. The second section (lines 3 to 9) describes how the orchestrator gathers the state information from the previous time step t−1to deploy controllers at Kopossible locations. It also covers reward evaluation and policy updates using the GAE method. The third section (lines 10 to 15) outlines the determination of state parameters such as average user latency, user count, and latency between the Cocontrollers and the orchestrator, and how the total controller count is measured at each time step. Lastly, the fourth section (lines 16 to 19) details the calculation of the orchestrator’s reward function. IV. EXPERIMENTAL EVALUATION A. Simulation Setup We perform simulations to verify the performance of our method regarding latency, transmission power, and PDR. We compare DERRIC’s performance against two baselines introduced by us, namely the Single Orchestrator - Single Controller (SOSC) and the Single Orchestrator - Distributed Controllers (SODC), as no other works in the literature tackle CPP in O-RAN with RL. We model PDR as Qt i=e−αdiNc, where di, the distance between i-th UE and its associated base station, and the controller node load Nc, which depends on the number of users managed by controller c. The coefficient α∈(0,+∞)
Algorithm 2: Orchestrator Operation // All orchestrators execute this process in parallel Data: Controller domain Ko, orchestrator set O, discount rate γ // Reward and Policy Initialization 1R←0,πo←InitializePolicy() // Random controller deployment initialization 2Co sample ←−−−− {0,1}|Ko| // each time step 3for t∈Ndo // Update state of orchestrator o 4so(t)←GetOrchestratorState(t−1,O,Co) // Select action according to policy πo 5ao(t)sample ←−−−− πo(a|so(t)) // Deploy controllers on Koaccording to action 6Co←DeployControllers(ao(t)) // Collect reward 7ro(t)←GetOrchestratorReward(Co) // Update returns 8R←ro(t) + γR // Update policy using GAE 9πo←PolicyUpdatePPO(πo, R) 10 Function GetOrchestratorState(t, O,Co): // Collect average controller-user latency 11 Lot ←MeasureAverageUserLatency(Co) // Collect user-count of controllers in the orchestrator domain 12 Not ←MeasureUserCount(Co) // Collect the latency between controllers and the orchestrator 13 Lot ←MeasureOrchestratorLatency(Co) // Collect the number of controllers for each orchestrator 14 Ct←MeasureOrchestratorLoad(O) 15 return (Lot, Not, Lot,Ct) 16 Function GetOrchestratorReward(t, Co): 17 Lot ←MeasureAverageUserLatency(Co) 18 Lot ←MeasureOrchestratorLatency(Co) 19 return −∥Lot∥−∥Lot∥ TABLE II EXPERIMENT PARAMETERS Parameter Value Number of RAN nodes V10,20,...,100 Number of users N50,100,...,500 Number of time steps 1000 Maximum transmission power Pmax 1W Coefficient of user PDR α0.01 Learning rate, Discount rate γ0.0001,0.9 Distance-load tradeoff coefficient α0.01 jointly controls the impact of distance and load on packet delivery. Figure 3 shows our simulation environment in which users and RAN nodes are uniformly distributed within a normalized unit square [0,1]2, assuming a free-space model without obstacles in the network. We compare the performance of the selected baselines across increasing numbers of users in the system. We implement our RL-based methods and environments using Python and Ray RLlib to train the Proximal Policy Optimization (PPO) algorithm that optimizes each agent’s policy. Table II summarizes the used simulation parameters. B. Results Figure 4 shows controller-user latency vector between N number of UEs and the controllers in which DERRIC consisFig. 3. Example of simulation environment with two orchestrators managing a total of four controllers tently outperforms both SOSC and SODC as the number of UEs increases. Notably, the average controller-user latency for DERRIC is almost 42% lower than SODC and around 66% lower than SOSC. The accomplished average latency LNis calculated for different locations and numbers of UEs, reflecting user mobility within the network. The lowest average latency in DERRIC is gained by deploying controllers at the lowest latency from its users in the controller domain. Figure 5 shows the average Pt iacross different systems containing a varying number of UEs for the three considered baselines. DERRIC’s lower controller-user latency allows agents to make decisions and allocate optimal power to users more quickly and effectively, which induces an observed overall lower power consumption than baselines. Furthermore, lower latency allows agents to allocate the required power to users without the need for excessive power on transmission paths. The experimental results show that DERRIC consumes about 17% and 29% lower power than SOSC and SODC, respectively. Figure 6 shows the average Qt ifor different numbers of UEs that DERRIC consistently achieves approximately 9% higher than SODC and 14% higher than SOSC. This outcome arises from using distributed controllers and decentralized orchestrators, which effectively balance the workload among controllers. By strategically deploying controllers with the lowest latency to their users and allocating appropriate transmission power to UEs, leading to quicker delivery of packets. We compare the worst-case cumulative wall-clock execution time over the slowest agent’s time steps tin a scenario containing 1000 UEs and 500 nodes hosting orchestrators and controllers between DERRIC and SODC. In DERRIC, the multiple agents (i.e., orchestrators and controllers), collect the state, make decisions, and take actions in parallel, whereas, in SODC, a single orchestrator agent manages all controllers. Figure 7 shows that the multi-agent system is faster than the single-agent in decision-making and action-taking based on observations. Therefore, the workload is distributed among multiple agents; in DERRIC agents correspond to the number of controllers and orchestrators. The UEs are distributed
Fig. 4. Impact of the number of UEs on the average user latency LN Fig. 5. Average power allocation Pt iacross the N number of User Equipment Fig. 6. Average user packet delivery ratio Qt i across the Nnumber of User Equipment Fig. 7. Cumulative execution time per time step Fig. 8. Controller reward rc(t)over time across controllers, reducing the burden on any single controller, while orchestrators effectively balance the workload among controllers, enhancing network performance. Figure 8 reports the controller reward in a system containing N= 25 UEs connected to a network containing 10 RAN nodes. These results show that DERRIC significantly outperforms SODC in terms of the packet delivery ratio, i.e., the controller’s reward, of up to 13% compared to SODC at convergence. Furthermore, DERRIC converges to a higher controller reward faster than the SODC baseline due to its collaborative nature among controllers. The faster convergence of DERRIC means that agents compute less to learn the optimal policy for system orchestration and user power control. V. CONCLUSION This paper addresses Near-RT RIC placement in O-RAN using a multi-agent system where orchestrators and controllers collaborate to minimize controller-user and orchestratorcontroller latencies, improve user PDR, and optimize transmission power allocation to UEs. Orchestrators determine the optimal number and location of the controllers, while controllers adjust transmission power based on user metrics. Simulations show that DERRIC improves controller-user latency and PDR of different numbers of UEs, outperforming existing methods. REFERENCES [1] S. Niknam, A. Roy, H. S. Dhillon, S. Singh, R. Banerji, J. H. Reed, N. Saxena, and S. Yoon, “Intelligent O-RAN for Beyond 5G and 6G Wireless Networks,” in 2022 IEEE Globecom Workshops (GC Wkshps). IEEE, 2022, pp. 215–220. [2] S. Faye, M. Camelo, J.-S. Sottet, C. Sommer, M. Franke, J. Baudouin, G. Castellanos, R. Decorme, M. P. Fanti, R. Fuladi et al., “Integrating Network Digital Twinning into Future AI-based 6G Systems: The 6GTWIN Vision,” in 2024 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit). IEEE, 2024, pp. 883–888. [3] L. Bonati, S. D’Oro, M. Polese, S. Basagni, and T. Melodia, “Intelligence and Learning in O-RAN for Data-Driven NextG Cellular Networks,” IEEE Communications Magazine, vol. 59, no. 10, pp. 21–27, 2021. [4] M. Abdel-Rahman, E. Mazied, F. Hassan, K. Teague, A. AL-Shaggah, A. Mackenzie, S. Midkiff, and K. V. Cardoso, “A Stochastic Optimization Framework for Joint RAN Intelligent Controller Placement and RAN Nodes Assignment in O-RAN Networks,” Authorea Preprints, 2023. [5] M. J. Abdel-Rahman, E. A. Mazied, K. Teague, A. B. MacKenzie, and S. F. Midkiff, “Robust Controller Placement and Assignment in SoftwareDefined Cellular Networks,” in 2017 26th International Conference on Computer Communication and Networks (ICCCN). IEEE, 2017, pp. 1–9. [6] A. Narwaria, K. Soni, and A. P. Mazumdar, “A position and energy aware multi-objective controller placement and re-placement scheme in distributed SDWSN,” The Journal of Supercomputing, pp. 1–29, 2024. [7] Y. Wu, S. Zhou, Y. Wei, and S. Leng, “Deep reinforcement learning for controller placement in software defined network,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2020, pp. 1254–1259. [8] G. M. Almeida, G. Z. Bruno, A. Huff, M. Hiltunen, E. P. Duarte, C. B. Both, and K. V. Cardoso, “RIC-O: Efficient Placement of a Disaggregated and Distributed RAN Intelligent Controller With Dynamic Clustering of Radio Nodes,” IEEE Journal on Selected Areas in Communications, vol. 42, no. 2, pp. 446–459, 2024. [9] G. Z. Bruno, V. K. Radhakrishnan, G. M. Almeida, A. Huff, A. P. da Silva, K. V. Cardoso, L. A. DaSilva, and C. B. Both, “RIC-O: An Orchestrator for the Dynamic Placement of a Disaggregated RAN Intelligent Controller,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2023, pp. 1–2. [10] G. Z. Bruno, G. M. Almeida, A. Sathish, A. P. da Silva, L. A. D. A. Huff, K. V. Cardoso, and C. B. Both, “Evaluating the deployment of a disaggregated open ran controller on a distributed cloud infrastructure,” IEEE Transactions on Network and Service Management, 2024. [11] C.-X. Wang, M. Di Renzo, S. Stanczak, S. Wang, and E. G. Larsson, “Artificial Intelligence Enabled Wireless Networking for 5G and Beyond: Recent Advances and Future Challenges,” IEEE Wireless Communications, vol. 27, no. 1, pp. 16–23, 2020. [12] S. Niknam, H. S. Dhillon, and J. H. Reed, “Federated Learning for Wireless Communications: Motivation, Opportunities, and Challenges,” IEEE Communications Magazine, vol. 58, no. 6, pp. 46–51, 2020. [13] E. H. Bouzidi, A. Outtagarts, R. Langar, and R. Boutaba, “Dynamic Clustering of Software-defined Network Switches and Controller Placement Using Deep Reinforcement Learning,” Computer networks, vol. 207, p. 108852, 2022. [14] J. Hao, T. Yang, H. Tang, C. Bai, J. Liu, Z. Meng, P. Liu, and Z. Wang, “Exploration in deep reinforcement learning: From singleagent to multiagent domain,” IEEE Transactions on Neural Networks and Learning Systems, 2023. [15] E. Hashemi Nezhad, A. Di Maio, and T. Braun, “Data-Driven Orchestration for Distributed RAN Intelligent Controller Placement in 6G Networks,” 2024. [Online]. Available: https://boris-portal.unibe.ch/ handle/20.500.12422/194060 [16] L. Bonati, M. Polese, S. D’Oro, S. Basagni, and T. Melodia, “Open, Programmable, and Virtualized 5G Networks: State-of-the-art and the Road ahead,” Computer Networks, vol. 182, p. 107516, 2020.