Full text
MARC-6G: Multi-Agent Reinforcement Learning for Distributed Context-Aware SFC Deployment and Migration in 6G Networks Solomon Fikadie Wassie, Eric Samikwa, Antonio Di Maio, and Torsten Braun Institute of Computer Science, University of Bern, Switzerland Email: {solomon.wassie, eric.samikwa, antonio.dimaio, torsten.braun}@unibe.ch Abstract—The Cloud Continuum Framework (CCF) extends computing capabilities across near-edge, far-edge, and extremeedge nodes beyond the traditional edge to meet the diverse performance demands of emerging 6G applications. While Deep Reinforcement Learning (DRL) has demonstrated potential in automating Virtual Network Function (VNF) migration by learning optimal policies, centralized DRL-based orchestration faces challenges related to scalability and limited visibility in distributed, heterogeneous network environments. To address these limitations, we introduce MARC-6G (Multi-Agent Reinforcement Learning for Distributed Context-Aware Service Function Chain (SFC) Deployment and Migration in 6G Networks), a novel framework that leverages decentralized agents for distributed, dynamic, and service-aware SFC placement and migration. MARC-6G allows agents to monitor different portions of the network, collaboratively optimize network control policies via experience sharing, and make local decisions that collectively enhance global orchestration under time-varying traffic conditions. We show through simulations that MARC-6G improves SFC deployment efficiency, reduces migration costs by 34%, and lowers energy consumption by 12.5% compared to the state-of-the-art centralized DRL baseline. Index Terms—Multi-Agent Reinforcement Learning, Distributed Service Orchestration, Distributed Intelligence, Service Function Chain I. INTRODUCTION The Sixth Generation (6G) mobile communication network is expected to leverage the concept of the Cloud Continuum Framework (CCF), which provides more flexible computational resources closer to end users beyond traditional edge computing, thereby meeting diverse application requirements [1]. However, realizing the full potential of this continuum requires scalable and intelligent orchestration of distributed resources across a heterogeneous, multi-tier infrastructure. Network services are provisioned as sequences of heterogeneous, predefined, and ordered Virtual Network Function (VNF) in the form of SFC on standardized, general-purpose servers enabled by Network Function Virtualization (NFV) technology [2]. VNFs are software-based implementations of network services such as Network Address Translation (NAT), Firewalls (FW), Intrusion Detection and Prevention Systems (IDPS), WAN optimizers (WO), Video Optimization Controllers (VOC), Traffic monitors (TM), and encoding/decoding functionalities. Those network functions provide a wide range of emerging applications, including video streaming, virtual/augmented/mixed reality, Industry 4.0, holographic communication, smart factories, autonomous vehicles, tactile industrial internet [3]. Optimal VNF deployment is one of the design requirements of modern mobile networks to ensure sustainable long-term performance and minimize operational costs. This, in turn, enables fast, reliable, and cost-effective delivery of network services. Several studies have proposed machine learning based approaches for the centralized orchestration of network functions [4], [5]. Deep Reinforcement Learning (DRL) is employed for network state awareness by leveraging Deep Learning (DL) to extract complex, high-dimensional network patterns and using Reinforcement Learning (RL) to optimize decisionmaking through interactions with dynamic network states. A Service Orchestrator (SO) is a network management system designed to automate the provisioning, scaling, and lifecycle management of network services [6]. Despite being effective in small-scale networks, centralized service orchestrators exhibit several limitations in large-scale environments due to their limited visibility of the global network state. These limitations include a single point of failure, high signaling overhead for network-wide data collection, and reduced responsiveness in real-time decision-making, which degrade the performance of latency-sensitive applications, and often lack the flexibility and scalability required to efficiently manage dynamic and heterogeneous workload demands [7]. The main research question addressed in this work is: How to optimally deploy and dynamically reconfigure multiple SFC requests in large-scale 6G networks, while adapting to timevarying traffic demands and heterogeneous infrastructure resources, to satisfy end-to-end performance requirements? To address this challenge, we introduce MARC-6G (MultiAgent Reinforcement Learning for Distributed Context-Aware SFC Deployment and Migration in 6G Networks). In MARC6G, agents monitor portions of the CCF and collaboratively learn VNF placement and migration policies from real-time network metrics, enabling scalable and dynamic orchestration across heterogeneous 6G infrastructures. The key contributions of this paper are summarized as follows: •We model the problem of scalable and distributed serviceaware orchestration of VNF deployment and migration for multiple SFC requests in large-scale 6G networks. •We design multi-agent RL–based distributed orchestrators
that use local state to jointly learn deployment and migration policies, minimizing delay, energy consumption, and VNF migration cost under dynamic traffic and resource conditions, while concurrently provisioning multiple SFCs.ovisioning multiple SFCs. •We evaluate MARC-6G against baseline methods and demonstrate higher request acceptance, improved energy efficiency, and reduced migration costs, validating its effectiveness for dynamic SFC management in 6G networks. The remainder of the paper is organized as follows: Section II describes the related works. Section III presents the system model and problem formulation. Section IV outlines the proposed methods. Section V presents the performance evaluation. Finally, Section VI draws the conclusions. II. RELATED WORKS Existing adaptive centralized provisioning techniques address VNF deployment as an elastic resource provisioning problem, aiming for flexible and on-demand resource allocation to meet network service requirements and service level agreements [4], [8], [9]. However, they often overlook the fact that multiple service requests may arrive at the orchestrator simultaneously, with varying performance requirements. Tang et al. [10] employ a digital twin powered by an attention model to guide an RL agent by predicting resource requirements a priori. However, because prediction is decoupled from action selection, the system’s ability to adapt in real time is compromised. Onsu et al. [8] apply DRL for VNF placement using fixed data center priorities based on residual capacity, but static scoring overlooks context and traffic, resulting in suboptimal placement. Dynamic priority assignment enables more adaptive and scalable orchestration under resource variability. Tanuboddi et al.[11] addressed VNF migration by leveraging softwarebased network functions to enable dynamic scaling, facilitating seamless migration in response to user mobility, load variations, and hardware failures. Chen et al. [12] and J.Chen et al. [13] address cost-efficient and fault-tolerant SFC migration using DRL and optimization techniques, respectively, but both approaches overlook key aspects such as fairness, realistic service lifetimes. Table I provides a comparison of the parameters considered in this study with those reported in the literature. III. SYSTEM MODEL AND PROBLEM FORMULATION A. Distributed Service Orchestration in 6G Networks We envision the 6G cloud network architecture that consists of three frameworks: the CCF, the Management and Orchestration Framework (MOF), and the Artificial Intelligence and Machine Learning Framework (AIMLF), as shown in Fig. 1 [15]. The CCF provides logically unified resource management across cloud-to-edge environments by dynamically integrating resources into Cloud, Near-edge, Far-edge, and Extreme-edge. A portion of CCF is highlighted in a different color to illustrate that VNFs of a single network service can be deployed across heterogeneous CCF nodes. TABLE I: Comparison of Related Works. Reference Concurrent VNF Migration Multiple SFC Requests Stateful VNF Migration Migration Cost Tang et al. [10]✓×✓× J.Chen et al.[13]✓×✓ ✓ Onsu et al. [8]×✓× × Tanub. et al. [11]×✓×✓ Zhang et al [14]×✓ ✓ × S.Long et al [4]×✓× × Chen et al. [12]✓×✓ ✓ Liu et al. [9]✓ ✓ × × MARC-6G ✓ ✓ ✓ ✓ DSO#M AIMLF Orchestrator Portal Model Serving and CI/CD AIMLF Storage Training Data Inference Data Performance Data AIMLF Functions Fuctions, Algorithms, Libraries AIMLF Learning Supervised Unsupervised Reinforcement NDT AIMLF Collaborative Functions FL, Transfer, MARL AIMLF Performance Monitoring AIMLF Models Database MMMM M R W Data Manager MOF NS1 NS2 NS3 AIMLF MSO VNF1 VNF4 VNF5 VNF6 VNF10 VNF11 VNF9 VNF12 VNF2 SO #1 SO #1 SO #N SO #N VNF6 VNF7 Cloud EdgeFar-Edge Extreme-edge VNF1 AI AI DSO#1 CCF AI Fig. 1: AI-native 6G network architecture with distributed service orchestrators for scalable network state adaptive VNF deployment [15]. The MOF provides distributed orchestration capabilities to enable scalable orchestration, interfacing with the CCF for infrastructure control and with the AIMLF for learning-based decision support. It comprises two components: the Master Service Orchestrator (MSO) and the Distributed Service Orchestrators (DSOs). The MSO is responsible for the initial deployment of Network Services (NSs) and DSOs across the CCF, while the DSOs manages runtime operations and the network service lifecycle management. Each DSOs contains multiple SO to address the dynamic workloads and scalability challenges posed by the heterogeneous 6G network infrastructure. The AIMLF is the intelligent control framework, supporting real-time monitoring and continuous learning in dynamic network environments. It ensures autonomous orchestration and adaptive service management by cooperating with the CCF and MOF. DSOs comprise intelligent agents that leverage the AIMLF to manage the CCF in real time, enabling contextaware decisions. DSOs monitor resource availability to support autonomous VNF allocation, predictive scaling, and proactive SFC migration. This ensures low latency, energy efficiency, and enhanced resilience and fault tolerance.
B. System Model We model distributed orchestration over the CCF at a high level (Fig. 2), where service orchestrators manage batches of SFC requests in a queue. We consider CCF network infrastructure that comprises heterogeneous physical nodes vi, each with CPU/GPU capacity Ci[cycles/s], distributed across four tiers: (i) centralized cloud data centers, (ii) near-edge nodes, (iii) far-edge nodes, and (iv) end-user devices. We represent the network infrastructure as a weighted undirected graph G= (V, E, W ), where Vis the set of physical nodes, E⊆V×Vdenotes the set of links connecting them, and the weight function W:E→R+represents each link’s available bandwidth. Each node v∈Vrepresents a physical network entity, such as an extreme-edge device (e.g., smartphone, electric vehicle, or drone), an edge or near-edge server, or a centralized cloud data center within the CCF. Each link e∈Erepresents a high-speed communication path, typically implemented via fiber connections. We denote the bandwidth capacity between nodes vi, vj∈Vas Bij [bit/s]. We consider a scenario involving multiple SFC requests arriving at the orchestrator. Each SFC request is modeled as a Directed Acyclic Graph (DAG) fi= (Ki, Li, δi, ζi,Λi, Bmin i, Dmax i, σi), where Kidenotes the set of VNFs for the i-th SFC; Lidenotes the set of logical links between VNFs; δiand ζidenote the source and destination endpoints, respectively; Λisignifies the traffic arrival time; Bmin i [bit/s] is the minimum bandwidth requirement; Dmax i[s] is the maximum tolerable end-to-end delay; and σi[cycle/s] is the total computational demand across all VNFs. Each VNF k∈Kirepresents a softwarized network function that can process incoming packets. The logical links (ki, kj)∈Li represent the connections between successive VNFs kiand kj, which represent a sequential dependency between VNFs. To support concurrent deployment, we consider a batch of Nactive SFC requests simultaneously for deployment, denoted by F= (f1, f2, . . . , fN). These requests, predefined according to application-specific requirements and submitted by tenants, are orchestrated in parallel over the physical infrastructure. The topology fiof an SFC is determined by the application it serves and is assumed to be specified by the tenant and forwarded to the network management plane for processing and deployment. Each VNF k∈Kimaintains an internal state, making stateful migration essential to preserve session continuity and avoid service disruption during reallocation. Given user mobility and fluctuating link quality, proactive resource management and adaptive state transfer are essential. The selection of a target node for migrating a stateful VNF can be modeled as a tuple Sk= (Mi, Dc, Qv, Ps, Tm)where Miis the size of the context to be migrated, Dcrepresents the deployment cost of the service on the target node, Qvdenotes the SLA violation impact during migration, Psindicates the congestion level along the selected migration path, and Tmis the total migration time. C. Problem Formulation We formulate the problem of simultaneous, elastic selfscaling placement of VNFs for a batch of Nactive SFC requests Cloud Continuum Framework Far edge Edge CloudExtreme Edge VK7 VK14 VK10 VK11 VK8 VK5 VK1 VK2 VK3 VK4 VK9 VK13 VK6 VK6 VK7 VK8 VK9 VK10 VK11 VK12 VK13 VK15 VK1 VK2 VK3 VK4 VK14 SO #N RL Agent SO#1 RL Agent SO#2 RL Agent f2 f1 VK5 fN SFC requests waiting in a queue Deployment Policy Experience sharing f1 f2 f3 fM Wired link Virtual linkDeployment action State Information Traffic flow Fig. 2: Several SFC requests arrive at the orchestrator in queues, while multiple intelligent agents concurrently deploy VNFs over CCF. over a shared physical network, with the goal of determining the optimal placement that minimizes end-to-end SFC latency. The total end-to-end delay of an SFC typically comprises propagation, communication, queuing, VNF computing, and virtualization delays. For tractability, we simplify our formulation by considering communication delay, VNF computing delay, and queuing delay. Our objective is to determine an optimal allocation that minimizes overall system latency, energy consumption, and migration cost. For each SFC request fi, we define the allocation vector as αi= (αi,1, αi,2, . . . , αi,|Ki|)∈V|Ki|, where each element αi,j ∈Vrepresents the physical node on which the j-th VNF for the SFC request fiis deployed. The complete set of allocation vectors for a batch of NSFC requests is denoted by A={α1, α2, . . . , αN} ∈ Ω, where Ω = QN i=1 V|Ki|is the generalized cartesian product of allocation vectors. Specifically, Ω = (α1, . . . , αN)|αi∈V|Ki|,∀i∈ {1, . . . , N}defines the feasible solution space comprising tuples of allocation vectors over heterogeneous domains V|K1|, V |K2|, . . . , V |KN|, corresponding to the variable number of VNFs across the N SFC requests. The communication latency for an SFC request fi, given the allocation αi, is modeled as Γ(αi) = P(km,kn)∈Lil(αi,m, αi,n), where l(αi,m, αi,n)denotes the shortest-path transmission delay between VNF deployment locations αi,m and αi,n. We define the processing latency for an SFC request fias P(αi) = PKi j=1(Pc(αi,j) + Pq(αi,j)),
where Pc(αi,j)and Pq(αi,j )represent the computing delay and queuing delay waiting for processing of the j-th VNF at its assigned node, respectively. The total SFC delay for a request fiis then T(αi) = P(αi) + Γ(αi), comprising the communication, processing, and queuing delays. We therefore define the total delay for a batch of Nrequests as T(A) = Pαi∈AT(αi), representing the aggregate delay across all SFCs in the batch. The energy consumption at time tgiven an allocation αiis modeled as Pαi(t) = X v∈V ωv(αi)·PPr v(t) + X v∈V (1 −ωv(αi)) ·PN v(t)(1) where ωv(αi)∈ {0,1}indicates whether node vis active under allocation αi. The active processing energy consumption is defined as PPr v(t) = PN v(t) + βv(t)·Pmax v(t)−PN v(t), with PN v(t)representing baseline energy usage and Pmax v(t)the energy drawn under full utilization [16]. The utilization factor βv(t)according to [13] is given by βv(t) = wp·xk vf(t)·νk f(t)·σp v Cp v +wu·xk vf(t)·νk f(t)·σu v Cu v (2) where xk vf(t)indicates whether flow fkis processed on node v,νk f(t)is the flow’s processing rate, σp vand σu vare the resource demands per unit of computing and storage capacity, respectively, and Cp v,Cu vare the corresponding available capacities. The weights wpand wureflect the relative importance of CPU and storage utilization. Let us define Pτ v(t)as the transmission energy required to migrate VNF kfrom one node to another. We therefore define the total energy consumption for a batch of Nrequests as P(A) = Pαi∈APαi(t), representing the cumulative energy usage induced by the current allocation across all requests in the batch. The VNF migration cost comprises both time and energy components. The total migration time for a VNF k∈Kiis defined as tk=P(i,j)∈lk Mk Bij , which represents the sum of the transmission times required to transfer the VNFs’ state size Mkover each physical link (i, j)with bandwidth Bij along the shortest path lkfrom the VNFs’ current deployment location to its new candidate node. The associated energy consumption to migrate VNF kalong the shortest path lk is ek=P(i,j)∈lkPτ v(t). We then define the SFC migration cost M(αi) = Pk∈Kiλ·tk+ (1 −λ)·ekas the sum of all VNF migration cost, where λ∈[0,1] is a tunable parameter used to balance the trade-off between time and energy costs, expressing different units as percentages. We therefore define the total migration cost for a batch of Nrequests as M(A) = Pαi∈AM(αi), representing the aggregate migration costs across all SFCs in the batch. Let us define nk vas the CPU cycles per second required by a VNF kwhen deployed on a physical node v∈V. Similarly, let bkl ij denote the bandwidth consumed by the logical link (k, l) of an SFC when mapped onto the physical link (i, j)∈E. We assume a total of MSFC requests arrive over time. To support scalable orchestration, we divide these into mbatches, each consisting of Nconcurrent requests, such that m=M/N. In each batch, VNFs are allocated jointly for the NSFCs. Given that the optimization problem is multi-objective, we define β= (β1, β2, β3)as the weight vector balancing the trade-offs between aggregated delay, energy consumption, and migration cost. Recall that T(A),P(A), and M(A)denote the total delay, energy consumption, and migration cost, respectively, aggregated over the batch of NSFC requests under allocation A. The objective is to minimize the scalarized cost function β⊤C(A), where C(A) = (T, P, M)(A)captures the three cost components. The goal of the optimization problem is to determine the optimal allocation A∗∈Ωfor a batch of NSFC requests such that system-wide resource utilization is efficient and constraint-satisfying, as shown in Equation 3. minimize A∈Ω Total SFC delay z }| { β1·T(A) + Energy consumption z }| { β2·P(A) + Migration cost z }| { β3·M(A)(3) subject to m X r=1 X i∈[N]X k∈Ki nk v≤Cv,∀v∈V(3a) m X r=1 X i∈[N]X (k,l)∈Li bkl ij ≤Bij,∀(i, j)∈E(3b) T(αi)≤Dmax i,∀i∈ {1, . . . , N}(3c) Constraint 3a ensures that for each node in the network, the total processing demand of all VNFs assigned to that node does not exceed its available computational capacity. Constraint 3b ensures that for each logical link, the combined bandwidth requirements of all SFC requests do not exceed the link’s available bandwidth capacity. Constraint 3c ensures that for each SFC request, the aggregate processing and communication delay does not exceed its E2E delay tolerance required for successful service completion. IV. MULTI-AGENT DEEP REINFORCEMENT LEARNING FOR DISTRIBUTED SFC DEPLOYMENT AND MIGRATION We present MARC-6G, a distributed cooperative MARL framework that orchestrates VNF deployment and migration in dynamic networks. We redefine the optimization equation 1 in the context of the Multi Agent Reinforcement learning (MARL) problem, where each agent operates under a Partially Observable Markov Decision Process (POMDP), observing a local portion of the CCF, acting independently, and exchanging state, action, and reward to learn a joint policy that anticipates future conditions. The policy selects physical nodes for VNFs by accounting for inter-VNF delay, deployment cost, SLA constraints, congestion, energy, and migration cost. Agents refine strategies from experience and peer sharing, enabling scalable management of heterogeneous, time-varying workloads. The broader system-wide objective is to jointly learn a set of optimal policies, denoted as π∗={π∗ 1, . . . , π∗ N}, which collectively maximize the performance of all agents in a cooperative manner.
A. Modeling VNF Deployment as Markov Decision Process We model the distributed SFC deployment and migration task as a cooperative MARL problem, where the detailed formulation is expressed in terms of the joint state space St, joint action space At, and joint reward function Rtat a given time t, as described below. 1) Joint State Space Stdescribes the current situation of each agent in the environment. The global state at time tis the collection of local states observed by each agent. The individual state of agent iis Si t= VKi i,t , V K2 i,t , ...L|Ki| Vi,h, Gtopo, M, fiwhere VKi i,t indicates that VNF Kiis deployed on physical node Viat time t, and L|Ki| Vi,h represents the link delay between node Vi(hosting Ki) and node Vh(hosting |Ki|). The state information also includes the observed network topology (Gtopo), the number of SFC requests in the queue M, and the characteristics of the SFC request fi. 2) Joint Action Space Atexplores optimal physical-node placements for hosting VNFs.Each action corresponds to selecting a sequence of physical nodes that meet the performance requirements of incoming SFC requests. At each time step t, each agent iselects a sequence of physical nodes to deploy the VNFs of its SFC. The action of agent i is defined as: αi t= (αi,1, αi,2, . . . , αi,|Ki|)∈V|Ki|where each element αi,Ki∈Vrepresents the selected physical node on which the i-th VNF for the SFC request fiis deployed. 3) Joint Reward Function Rtassigns a numerical score to each agent’s decision on the network performance. The agent evaluates each physical node’s placement by interpreting patterns it learns from the environment. At time t, the reward for agent iaccording to Equation 3 is given as follows. Ri t=−PT t=0 γt·(β1·T(A) + β2· P(A) + β3·M(A)) where T is the maximum time-step at which an agent learn optimal deployment policies. The collective reward is then defined as: R=PN i=1 Ri t,where atis the joint action across all agents. B. Multi-Agent Proximal Policy Optimization for Autonomous VNF deployment In Multi Agent Proximal Policy Optimization (MAPPO), distributed learning via experience sharing employs separate policy networks, each parameterized by πθi, and separate value networks, each parameterized by Vϕi. The RL agents exchange experiences through their value networks and policy networks to learn coordinated deployment policies. Each agent iobserves a local state si t(e.g., a subset of the network topology, current VNF placements, resource utilization, and KPI requirements for incoming traffic) and selects a deployment action ai t∼πθi(a|si t). After agent iexecutes its action, it receives a reward ri t+1, which is weighted by βiand exchanged with its neighboring agents. Based on this reward, each agent iupdates its value function Vϕi(si t)and the policy network parameters θi. In each training epoch, agents also share their latest policy πθi(· | si t)with neighbors, enabling them to anticipate others’ behaviors and coordinate VNF placements across the network. Thus, by exchanging weighted reward signals βiri t+1 (i.e, experience of an agent), MAPPO ensures that each agent’s value network gains a more global perspective on performance. Algorithm 1: MARC-6G Operation Input: F= (f1, f2,...,fN),Tmax,G,l(i, j),M,N // Initialize policy, value, replay buffer 1Initialization: ϕi 0,πi 0, Dt // Randomly place Initially VNF 2Si t=VKi i,t , V K2 i,t , ...L|Ki| Vi,h // Initialize the reward to zero 3Ri t←0 4for t∈Tmax do 5for i∈Ndo // Each agent selects one SFC request 6fi←Sample(M, i) 7c←MeasureCPUrequirment(fi) 8b←MeasureBandwidth(fi) 9l←MeasureLatency(fi) 10 q←MonitorLinkquality(Gtopo) // Translate fi& number of SFC requests in a queue as state 11 Si t←(l, b, c, q, Gtopo, M) // Execute Ai t∼πi θbased on the current policy 12 Si t Ai t,πθ(Ai t|Si t,A−i t) −−−−−−−−−−−−→ Si t+1, Ri t+1 // Deploy Kinew location 13 Si t+1 ←VKi t // Store each si t, ai t, Ri tin replay buffer 14 Dt← {si 0, ai 0, ri 0,...,si t, ai t, ri t} 15 if l(fi)< l(i, j)then 16 Ai t∼πθ(Ai t|Si t, A−i t)Execute −−−−→ Ri t+1, Si t+1 // Each agent calculates advantage estimate 17 Aπi θt(st, at)←Qi(s, a)−Vi ϕ(s) // Update θiand ϕiparameters 18 θi t+1 = arg maxθi1 |Dt|TPτ∈DtPT t=0 min πθ(at|st) πθt(at|st)Aπθt(st, at), g(ϵ, Aπθt(st, at)) 19 ϕi t+1 = arg minϕ1 |Dt|Pτ∈DtPT t=0(Vϕ(st)−Rt)2 20 else // Relocate VNF 21 VKi t+1 ←VKi t // Assign VNF Kioptimally 22 Ki Deploy −−−−→ VKi t 23 return ϕi(st), πθi(st) C. MARC-6G Algorithm The step-by-step workflow of Algorithm 1 is described below in detail. The inputs to the MARC-6G algorithm are the physical network topology Gtopo, the corresponding resource capacities (e.g., link bandwidth, link delay, node CPU), and a set of SFC requests fiwith their characteristics. Each agent selects a single SFC request from the pool M, based on its traffic
arrival time Λiand E2E delay requirement Dmax i(lines 5-7). Each agent constructs the state stby measuring the SFC request KPIs fi(bandwidth, CPU, latency) and the current physical resource availability (lines 7–12). Next each agent execute the deployment action at∼πθ(st)by considering the deployment action of others, following current policy πθ, update the system to state st+1, and set the new VNF placement St+1 =Vk istore si t, ai t, ri ttrajectory in replay buffer Dt(lines 13–14). Then the requirements of l(fi)are compared with available infrastructure resources: l(fi)< l(i, j)for the given allocation vector. The reward Rtand the advantage Aπi θt(st, at)are then computed (lines 11-15). Then update θi t+1(the policy) and ϕi t+1 (value parameters), and repeat this iteration until the optimal policy is developed (lines 14-21). The decision about the SFC request deployment or migration to another node (lines 21-22). Finally, return the value parameter ϕi(st)and the policy parameter πθi(st)(line 23). V. PERFORMANCE EVALUATION A. Experimental Setup We performed the simulation experiments using a custom simulator developed in Python with the NetworkX [17] to generate USA NET [18] network topologies that represent the underlying network infrastructure. For the MARC-6G implementation, we employ the open-source libraries Gymnasium v1.0 and RLlib v2.4.0 to train and evaluate RL agents within a custom environment. Given the variability of incoming traffic, each SFC request fiis managed in a way that preserves better link quality. Multiple PPO agents operate in parallel, continuously monitoring network conditions and adjusting VNF placements in response to traffic patterns and system dynamics. We model both the state space Sand the action space A as multi-discrete, making PPO a suitable choice due to its robustness, stability, and convergence properties, and update its policy online in discrete control tasks. B. Baselines and Evaluation Metrics The performance of MARC-6G, compared with centralized orchestration from the previous work [19] and baseline greedybased VNF allocation. The Greedy allocator, deploying VNFs immediately without considering long-term impacts, often leads to suboptimal resource utilization. Centralized orchestrator Single Agent Proximal Policy Optimization (SAPPO), suffers from scalability limitations due to its reliance on a single global controller, preventing real-time adaptability in largescale, dynamic network environments. MARC-6G is evaluated using: (i) Reward, the weighted objective function that combines total E2E delay, energy consumption, and migration cost (Equation 3); (ii) Number of Accepted Requests, the fraction of VNF requests successfully deployed with the required performance relative to all incoming requests; (iii) Energy Consumption, highlighting energy-efficient VNF placement that minimizes usage by aggregating workloads from underutilized nodes onto a minimal set of active servers in real time; and (iv) Migration Cost, include the computational, energy, and bandwidth overhead incurred when migrating VNF instances to optimize resource utilization and service quality. C. Discussion and Simulation Results Figure 3a compares learning curves for MARC-6G, SAPPO, and a greedy (non-learning) allocator: rewards fluctuate during exploration and SAPPO leads early (no coordination overhead), but once a shared state emerges, MARC-6G’s multiagent cooperation accelerates learning past SAPPO, ultimately outperforming both SAPPO and the greedy baseline with more stable, efficient VNF placements. Figure 3b shows that across the SFC types in [3] (ID4.0, MIoT, CG, AR, VS), MARC-6G with five agents consistently outperforms SAPPO and the greedy allocator in terms of the number of accepted requests. For ID4.0, SAPPO and MARC6G are comparable because ID4.0 has few VNFs, reducing contention. MARC-6G deploys SFCs concurrently with five agents, whereas SAPPO and the greedy baseline place one per step; the greedy is further hindered by random VNF placement. Figure 3c shows energy consumption rising with the number of devices because VNFs are spread across many servers without accounting for overor underprovisioning. MARC6G lowers energy consumption by up to 12.5% and 39.2% compared to SAPPO and greedy across device scales by learning utilization-aware, energy-efficient VNF placements. Figure 4a shows that as the number of physical devices grows, migration cost decreases: a larger pool of nodes enables optimal nearby node selection, and MARC-6G’s multi-agent monitoring further minimizes cost by up to 34% and 41.25% compared to SAPPO and greedy approaches, respectively. Figure 4b reports E2E delay for 40/70/100-node networks under concurrent SFC loads of 3, 6, 9, and 12; delay drops across all loads as node count grows, indicating that MARC-6G learns scalable, resource-aware deployments that minimize latency even under heavier traffic. Figure 4c illustrates the scalability of MARC-6G: as the number of SFC requests and hence agents increases proportionally, the end-to-end delay decreases, since each agent, managing only a segment of the network, and deploy multiple requests concurrently. VI. CONCLUSION This paper addresses distributed, context-aware VNF placement and migration under time-varying traffic for 6G networks, with the objective of minimizing end-to-end delay, energy consumption, and migration cost. We present MARC-6G, a MARL framework for concurrent SFC deployment that adapts online to network dynamics, placing and migrating VNFs to meet the strict delay requirements of fluctuating applications. Unlike centralized orchestration, characterized by slow, global state collection as request volume and topology size grow, MARC6G monitors the local portion of the CCF, shares experience among agents, and updates policies in real time. Experimental results show that MARC-6G has higher request acceptance, better energy efficiency, lower migration cost, and improved scalability compared to baseline approaches.
(a) Reward convergence (b) Number of accepted requests for different specialized network services (c) Energy-consumption comparison across different physical devices Fig. 3: Learning performance of MARC-6G: accepted SFC requests and energy consumption across the number of devices, compared with the baseline methods. (a) VNF migration costs for different physical devices (b) Variable number of physical devices and SFC requests (c) Variable number of agents and SFC requests Fig. 4: Analysis of migration cost and scalability under varying SFC request loads and numbers of agents, and their impact on E2E delay. ACKNOWLEDGMENT This work is funded by the SNS-JU 6G Cloud project under the EU Horizon Europe programme (Grant Agreement No. 101139073). REFERENCES [1] C. Campolo, A. Iera, and A. Molinaro, “Network for distributed intelligence: A survey and future perspectives,” IEEE Access, vol. 11, pp. 52 840–52 861, 2023. [2] I. Avgouleas, D. Yuan, N. Pappas, and V. Angelakis, “Virtual network functions scheduling under delay-weighted pricing,” IEEE Networking Letters, vol. 1, no. 4, pp. 160–163, 2019. [3] J. M. Ziazet, B. Jaumard, H. Duong, P. Khoshabi, and E. Janulewicz, “A dynamic traffic generator for elastic 5g network slicing,” in 2022 IEEE international symposium on measurements & networking (M&N). IEEE, 2022, pp. 1–6. [4] S. Long, B. Liu, H. Gao, X. Su, and X. Xu, “Deep reinforcement learning-based sfc deployment scheme for 6g iot scenario,” in 2023 IEEE Symposium on Computers and Communications (ISCC). IEEE, 2023, pp. 1189–1192. [5] N. Toumi, M. Bagaa, and A. Ksentini, “Machine learning for service migration: a survey,” IEEE Communications Surveys & Tutorials, vol. 25, no. 3, pp. 1991–2020, 2023. [6] T. Haga, “Orchestration of networking processes,” 2007. [7] T. Mai, H. Yao, N. Zhang, W. He, D. Guo, and M. Guizani, “Transfer reinforcement learning aided distributed network slicing optimization in industrial iot,” IEEE Transactions on Industrial Informatics, vol. 18, no. 6, pp. 4308–4316, 2022. [8] M. A. Onsu, P. Lohan, B. Kantarci, E. Janulewicz, and S. Slobodrian, “Unlocking reconfigurability for deep reinforcement learning in sfc provisioning,” IEEE Networking Letters, 2024. [9] Q. Liu, L. Tang, T. Wu, and Q. Chen, “Deep reinforcement learning for resource demand prediction and virtual function network migration in digital twin network,” IEEE Internet of Things Journal, vol. 10, no. 21, pp. 19 102–19 116, 2023. [10] L. Tang, Z. Li, J. Li, D. Fang, L. Li, and Q. Chen, “Dt-assisted vnf migration in sdn/nvf-enabled iot networks via multiagent deep reinforcement learning,” IEEE Internet of Things Journal, vol. 11, no. 14, pp. 25 294– 25 315, 2024. [11] B. R. Tanuboddi, G. Gad, Z. M. Fadlullah, and M. M. Fouda, “Optimizing vnf migration in b5g core networks: A machine learning approach,” in 2024 International Conference on Smart Applications, Communications and Networking (SmartNets). IEEE, 2024, pp. 1–5. [12] R. Chen, H. Lu, Y. Lu, and J. Liu, “Msdf: A deep reinforcement learning framework for service function chain migration,” in 2020 IEEE Wireless communications and networking conference (WCNC). IEEE, 2020, pp. 1–6. [13] J. Chen, J. Chen, K. Guo, R. Hu, T. Zou, J. Zhu, H. Zhang, and J. Liu, “Fault tolerance oriented sfc optimization in sdn/nfv-enabled cloud environment based on deep reinforcement learning,” IEEE Transactions on Cloud Computing, vol. 12, no. 1, pp. 200–218, 2024. [14] Y. Zhang, R. Wang, J. Hao, Q. Wu, Y. Teng, P. Wang, and D. Niyato, “Service function chain deployment with vnf-dependent software migration in multi-domain networks,” IEEE Transactions on Mobile Computing, pp. 1–18, 2024. [15] 6G-Cloud, “D2.2 - Initial Results on Architecture, Service Interfaces and AI/ML,” 6G-Cloud Consortium, Deliverable D2.2, 2025. [16] J. A. Aroca, A. Chatzipapas, A. F. Anta, and V. Mancuso, “A measurement-based characterization of the energy consumption in data center servers,” IEEE Journal on selected areas in communications, vol. 33, no. 12, pp. 2863–2877, 2015. [17] NetworkX Developers, “Networkx,” https://networkx.org/, 2025, accessed: 31 May 2025. [18] N. Spring, R. Mahajan, D. Wetherall, and T. Anderson, “Measuring isp topologies with rocketfuel,” IEEE/ACM Transactions on networking, vol. 12, no. 1, pp. 2–16, 2004. [19] S. F. Wassie, A. Di Maio, and T. Braun, “Deep reinforcement learning for context-aware online service function chain deployment and migration over 6g networks,” in Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing, 2025, pp. 1361–1370.