Data-Driven Orchestration for Distributed RAN Intelligent Controller Placement in 6G Networks
Abstract
Open-Radio Access Network (O-RAN) brings an innovative ap- proach to address the issues of controller placement in large-scale 6G networks. In this work, we introduce a Reinforcement Learn- ing (RL) algorithm for decentralized RAN Intelligent Controller Orchestration in 6G Networks, which leverages the online learning capabilities of a multi-agent RL system. Our method achieves up to 66% lower user latency and up to 14% higher user packet delivery ratio compared to state-of-the-art baselines in a broad range of simulated scenarios.
Full text
Data-Driven Orchestration for Distributed RAN Intelligent Controller Placement in 6G Networks Elham HashemiNezhad University of Bern Bern, Switzerland [email protected] Antonio Di Maio University of Bern Bern, Switzerland [email protected] Torsten Braun University of Bern Bern, Switzerland [email protected] ABSTRACT Open-Radio Access Network ( O-RAN ) brings an innovative approach to address the issues of controller placement in large-scale 6G networks. In this work, we introduce a Reinforcement Learning ( RL ) algorithm for decentralized RAN Intelligent Controller Orchestration in 6G Networks, which leverages the online learning capabilities of a multi-agent RL system. Our method achieves around 42-66% lower user latency and 9-14% higher user packet delivery ratio compared to state-of-the-art baselines in a broad range of simulated scenarios. ACM Reference Format: Elham HashemiNezhad, Antonio Di Maio, and Torsten Braun. 2025. DataDriven Orchestration for Distributed RAN Intelligent Controller Placement in 6G Networks. In The 40th ACM/SIGAPP Symposium on Applied Computing (SAC ’25), March 31-April 4, 2025, Catania, Italy. ACM, New York, NY, USA, 3 pages. https://doi.org/10.1145/3672608.3707972 1 INTRODUCTION The evolution beyond 5G and the development of sixth-generation (6G) wireless networks call for an architectural transformation that supports service heterogeneity, coordinates multi-connectivity technologies, and enables on-demand service deployment. The O-RAN architecture includes two RAN Intelligent Controllers (RICs) that perform management and control of the network at Near-RealTime (Near-RT) RIC between 10 [ ms ] and 1 [s] and Non-Real-Time (Non-RT) RIC more than 1 [s] time scales [ 3 ]. Optimal placement of controllers has a prominent effect on minimizing this response time. Distributing a minimum number of controllers at optimal locations to complete control functions promptly is known as the Controller Placement Problem (CPP) [ 1 ]. Distributed controllers are necessary to enhance network performance and ensure the sustainability of user connections. A single controller poses a single point of failure, compromising network reliability and availability. Several works have tackled the problem of distributed controller placement in different types of networks. Almeida et al. [ 2 ] proposed a RIC Orchestrator (RIC-O) to optimize the deployment of the Near-RT RIC components. The orchestrator seeks to place Near-RT RIC components to the closest edge computing nodes based on the least cost in an O-RAN, so it reduces latency and improves cost efficiency using a greedy strategy. The approach also employs a monitoring system Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). SAC ’25, March 31-April 4, 2025, Catania, Italy ©2025 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-0629-5/25/03 https://doi.org/10.1145/3672608.3707972 for the control loop. However, convergence may occur when all edge nodes host Near-RT RIC components, potentially increasing latency, particularly in large, dynamic 6G networks. Wu et al. [ 6 ] proposed a Deep Q-Network (DQN) approach based on deep RL for controller placement in Software-Defined Networks (SDN). This method optimizes latency and load-balancing metrics by optimally placing the controller and adjusting the switch-controller mapping according to the flow fluctuations. However, this method does not consider the number of controllers. Lyu et al. [ 5 ] proposed a fully distributed orchestration to minimize SDN’s time-average cost by stochastically optimizing controllers’ on-demand activation, an adaptive association of controllers and switches, and real-time request processing and dispatching. However, RL for intelligent and decentralized decision-making would be more efficient in meeting the dynamic, complex, and high-demand environment of 6G networks. The main contributions of this research to tackle the mentioned issues are summarized as follows. • Develop a multi-agent RL approach for distributed controllers to optimize transmission power allocation, thereby reducing average latency and maximizing packet delivery ratio across the network. • Design a decentralized orchestration framework for efficient controller deployment and orchestration management using a multi-agent RL system. 2 METHODOLOGY We assume the system operates within a Radio Access Network (RAN) deployment adhering to O-RAN specifications, as represented in Figure 1. The network topology is modeled as an undirected graph 𝐺=(𝑉, 𝐸) , where 𝑉={𝑣1, . . . , 𝑣|𝑉|} represents a set of network devices. Moreover, we consider a set of edges 𝐸={𝑒1. . . , 𝑒|𝐸|} representing a set of physical links between two nodes in 𝑉 . Each link in 𝐸 is characterized by its link latency 𝐿𝑖 𝑗 between devices 𝑣𝑖∈𝑉 and 𝑣𝑗∈𝑉 . We assume that each base station in 𝑉 must adopt the optimal transmission power 𝑃𝑡 𝑖[ W ] towards its 𝑖 -th connected User Equipments ( UE s). Our system model contains a set C ⊆ 𝑉 of controllers that each is in charge of power allocation for 𝑁𝑐UE s associated with all base stations managed by controller 𝑐 (i.e., the controller domain). The average transmission power 𝑃𝑐 𝑡 selected by controller 𝑐 from its managed base stations to all its managed 𝑁𝑐 users at the time step 𝑡 as 𝑃𝑐 𝑡=1 𝑁𝑐Í𝑖∈[𝑁𝑐]𝑃𝑡 𝑖 . A set O ⊆ 𝑉 of orchestrators partitioning the network into contiguous domains using a Voronoi-tessellation [ 4 ] approach. Each domain 𝐾𝑜 groups RAN nodes 𝑉 with the lowest latency to the orchestrator 𝑜∈ O . Orchestrators are tasked with dynamically deploying controllers C𝑜on a set 𝐾𝑜⊆𝑉of possible deployment locations.
SAC ’25, March 31-April 4, 2025, Catania, Italy E. HashemiNezhad et al. Figure 1: Example of deployment of controllers and orchestrators over the modeled physical network infrastructure 2.1 Controller Operation Each controller allocates the transmission power to each user in its domain by leveraging a local RL agent. This power allocation process can be modeled as a sequential decision-making problem (𝑠𝑐, 𝑎𝑐,𝑟𝑐)(𝑡) , where each controller adjusts its action 𝑎𝑐(𝑡) ∈ A𝑐(𝑡) in the environment, i.e., the transmission power allocated to each user at each time step 𝑡 , based on the current system’s state 𝑠𝑐(𝑡) ∈ S𝑐(𝑡) and a reward function 𝑟𝑐(𝑡) ∈ R𝑐(𝑡) . We define the controller’s state, action, and reward as follows. 2.1.1 Controller State. At each time step 𝑡 , every controller builds the local state 𝑠𝑐(𝑡) by collecting system metrics such as the latency vector 𝐿𝑐(𝑡−1) and the SNR vector 𝜌𝑐(𝑡−1) , which contain information about the communication latency between the controller and all managed UE s and the SNR received by all managed UE s at time step 𝑡− 1. Moreover, the controller also collects an average transmission power vector P𝑡=(𝑃1 𝑡, . . . , 𝑃 | C| 𝑡) through Inter-controller connections at every time step 𝑡 , which contains the latest average transmission power selected by all controllers (itself and all others) at time step 𝑡 . As the number of base stations and users 𝑁𝑐 managed by the generic controller 𝑐 vary over time, the dimension of the state space S𝑐(𝑡) that contains the state 𝑠𝑐(𝑡) ∈ S𝑐(𝑡) is also timevarying and S𝑐(𝑡) ⊆ R2𝑁𝑐+| C | . Equation 1 formally characterizes the 𝑐-th controller’s state 𝑠𝑐(𝑡)at every time step 𝑡∈N. 𝑠𝑐(𝑡)=𝐿𝑐(𝑡−1), 𝜌𝑐(𝑡−1),P𝑡−1(1) 2.1.2 Controller Action. Each controller 𝑐 must determine the transmission power for each user in its control domain by executing a local controller policy 𝜋𝑐 . We define the controller agent’s action 𝑎𝑐(𝑡) ∈ A𝑐(𝑡)=[ 0 , 𝑃max)𝑁𝑐 , where 𝑃max represents hardware or regulatory limitations of base stations on the maximum transmission power. 2.1.3 Controller Reward. The reward function 𝑟𝑐(𝑡) ∈ R𝑐(𝑡)=R+ for the generic controller 𝑐 at time step 𝑡 is to maximize the average of all packet delivery ratios 𝑄𝑡 𝑖 of all UE s in the 𝑐 -th controller’s domain at time step 𝑡(Equation 2). 𝑟𝑐(𝑡)= 1 𝑁𝑐∑︁ 𝑖∈𝑁𝑐 𝑄𝑡 𝑖(2) 2.2 Orchestrator Operation We propose a decentralized orchestration framework where orchestrators employ a multi-agent RL algorithm to deploy controllers; the decisions at each time step rely on the outcomes of previous steps. The process can be represented using the tuple (𝑠𝑜, 𝑎𝑜,𝑟𝑜)(𝑡) , where each orchestrator performs 𝑎𝑜(𝑡) ∈ A𝑜(𝑡) on the environment to select controller nodes at each time step 𝑡 based on the current system’s state 𝑠𝑜(𝑡) ∈ S𝑜(𝑡) and achieve a reward function 𝑟𝑜(𝑡) ∈ R𝑜(𝑡) . The orchestrator’s state, action, and reward are denoted as follows. 2.2.1 Orchestrator State. Each orchestrator has a state 𝑠𝑜(𝑡) by gathering the controller-user latency vector 𝐿𝑜(𝑡−1) and the orchestrator user-count vector 𝑁𝑜(𝑡−1) at time step 𝑡− 1. Each agent observes the latency between the orchestrator and all managed controllers in the previous time step as orchestrator-controller latency vector 𝐿𝑜(𝑡−1) to deploy controllers in the possible lowest orchestrator-controller latency at time step 𝑡 . Each orchestrator also collects number of controllers vector, C𝑡−1=(|C1|, . . . , |C| O| |)𝑡−1∈ N| O| through Inter-orchestrator connections which represents the latest number of controllers managed by orchestrator 𝑜 at time step 𝑡− 1. Equation 3 describes the 𝑜 -th orchestrator’s state 𝑠𝑜(𝑡) at every time step 𝑡∈N. 𝑠𝑜(𝑡)=(𝐿𝑜(𝑡−1), 𝑁𝑜(𝑡−1), 𝐿𝑜(𝑡−1),C𝑡−1)(3) 2.2.2 Orchestrator Action. The action 𝑎𝑜(𝑡) for a single orchestrator agent is to select controller nodes in the orchestrator domain as a binary decision by executing a local orchestrator policy 𝜋𝑜 . The agent selects a node with less average user latency, more number of users, and less latency between the node and the orchestrator as a controller and gets 1; otherwise, it gets 0 for non-controllers. The orchestrator agent’s action is defined as 𝑎𝑜(𝑡) ∈ { 0 , 1 }|𝐾𝑜| , which is a logical value representing controller and non-controller nodes. 2.2.3 Orchestrator Reward. The reward function 𝑟𝑜(𝑡) in the RL framework for each orchestrator is designed to minimize user latency within the controller’s domain and the latency between the controller and its orchestrator node (Equation 4). 𝑟𝑜(𝑡)=−∥𝐿𝑜𝑡 ∥−∥𝐿𝑜𝑡 ∥(4) 3 EXPERIMENTAL EVALUATION In our evaluation, we conduct simulations to verify the performance of our method in aspects of latency, transmission power, and packet delivery ratio. We compare three environments, including Single Orchestrator - Single Controller (SOSC), Single Orchestrator - Distributed Controllers (SODC) that are implemented by a single-agent system, and the proposed method with a multi-agent system in the same network conditions and topology in terms of latency, packet delivery ratio, and transmission power.
Data-Driven Orchestration for Distributed RAN Intelligent Controller Placement in 6G Networks SAC ’25, March 31-April 4, 2025, Catania, Italy Figure 2: The average user latency 𝐿𝑁 across the number 𝑁 of User Equipment In our experiment, we assume that each node 𝑖∈𝑉 is located at a position (𝑥𝑖,𝑦𝑖) on a 2D plane randomly, and we define 𝑑𝑖 𝑗 as the Euclidean distance between the positions of nodes 𝑖 and 𝑗 . The latency between two nodes 𝑖 and 𝑗 is defined as 𝐿𝑖 𝑗 =𝐿p 𝑖 𝑗 +𝐿t 𝑖 𝑗 +𝐿q 𝑖 𝑗 + 𝐿c 𝑖 𝑗 . The propagation latency 𝐿p 𝑖 𝑗 = 𝑑𝑖 𝑗 𝜂𝑖 𝑗 between two nodes is the ratio between their distance 𝑑𝑖 𝑗 and the speed of electromagnetic waves 𝜂𝑖 𝑗 in the transmission medium between nodes 𝑖 and 𝑗 . The transmission latency 𝐿t 𝑖 𝑗 =𝑆 𝑅𝑖 𝑗 , where 𝑆[bit] is the packet size and 𝑅𝑖 𝑗 [bit/ s ] is the transmission rate. 𝐿q 𝑖 𝑗 represents the queuing latency as a packet’s time in a queue before it can be transmitted. Finally, the processing latency 𝐿c 𝑖 𝑗 reflects the time required for the devices to analyze and route packets. We also define the packet delivery ratio in each user transmission as 𝑄𝑡 𝑖=𝑒−𝛼𝑑𝑖𝑁𝑐 , which is used to simulate the packet delivery ratio 𝑄𝑡 𝑖 based on 𝑑𝑖 , the distance between the 𝑖 -th UE and its associated base station, and the controller node load 𝑁𝑐 at time step 𝑡 . The coefficient 𝛼∈ ( 0 ,+∞) jointly controls the impact of distance and load on packet delivery. A set of 𝑉 is randomly deployed in a normalized unit square [ 0 , 1 ]2 in our simulation environment. We consider the number of users 𝑁={ 50 , 100 , . . . , 500 } to compare the performance of the selected baselines under different numbers of users in the system. We implement our RL -based methods with Python and use the Ray RLlib to train the Proximal Policy Optimization (PPO) algorithm, which optimizes policy performance for orchestrators and controllers. Figure 2 demonstrates that the proposed method consistently outperforms both SOSC and SODC as the number of UE s increases in terms of lower user latency. The proposed method’s average user latency 𝐿𝑁 is almost 42% lower than SODC and around 66% lower than SOSC at the UE level. The gained average latency in the proposed method is the lowest because each controller is placed at the lowest latency from its users in the controller domain, minimizing delays in their user’s transitions. Figure 3 shows the packet delivery ratio 𝑄𝑡 𝑖 for different numbers of UE s that the proposed method consistently achieves approximately 9% higher than SODC and 14% higher than SOSC. This outcome arises from using distributed controllers and decentralized orchestration, which effectively balance the workload among controllers. By deploying controllers with the lowest latency from Figure 3: Average user packet delivery ratio 𝑄𝑡 𝑖 across the number 𝑁of User Equipment their users, the communication and decision-making between the controllers and the users are minimized, leading to quicker delivery of packets. 4 CONCLUSION This paper addresses the Near-RT RIC placement problem, crucial for managing UE s in the O-RAN architecture. We propose a multi-agent approach where decentralized orchestrators place controllers to minimize latency, sharing data to optimize controller placement within each domain. Controllers act as agents, adjusting user transmission power based on latency and Signal-to-Noise Ratio ( SNR ) observations. Extensive experiments show that the proposed method reduces latency and improves packet delivery ratios, outperforming state-of-the-art baselines. ACKNOWLEDGMENTS This work was funded by the SNS-JU 6G Cloud project under the European Union’s Horizon Europe Research and Innovation Programme under Grant Agreement No. 101139073. REFERENCES [1] Mohammad Abdel-Rahman, EMADELDIN MAZIED, FAHID HASSAN, Kory Teague, ATHEER AL-SHAGGAH, ALLEN MACKENZIE, Scott Midkiff, and Kleber V Cardoso. 2023. A Stochastic Optimization Framework for Joint RAN Intelligent Controller Placement and RAN Nodes Assignment in O-RAN Networks. Authorea Preprints (2023). [2] Gabriel Matheus Almeida, Gustavo Zanatta Bruno, Alexandre Huff, Matti Hiltunen, Elias Procopio Duarte, Cristiano Bonato Both, and Kleber Vieira Cardoso. 2024. RIC-O: Efficient Placement of a Disaggregated and Distributed RAN Intelligent Controller With Dynamic Clustering of Radio Nodes. IEEE Journal on Selected Areas in Communications 42, 2 (2024), 446–459. https://doi.org/10.1109/JSAC.2023. 3336159 [3] Leonardo Bonati, Salvatore D’Oro, Michele Polese, Stefano Basagni, and Tommaso Melodia. 2021. Intelligence and Learning in O-RAN for Data-Driven NextG Cellular Networks. IEEE Communications Magazine 59, 10 (2021), 21–27. [4] Zakhar Kabluchko and Christoph Thäle. 2021. The Typical Cell of a Voronoi Tessellation on the Sphere. Discrete & Computational Geometry 66, 4 (2021), 1330– 1350. [5] Xinchen Lyu, Chenshan Ren, Wei Ni, Hui Tian, Ren Ping Liu, and Y Jay Guo. 2018. Multi-Timescale Decentralized Online Orchestration of Software-Defined Networks. IEEE Journal on Selected Areas in Communications 36, 12 (2018), 2716– 2730. [6] Yiwen Wu, Sipei Zhou, Yunkai Wei, and Supeng Leng. 2020. Deep reinforcement learning for controller placement in software defined network. In IEEE INFOCOM 2020-IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 1254–1259.