Full text
Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks Solomon Fikadie Wassie University of Bern Switzerland [email protected] Antonio Di Maio University of Bern Switzerland [email protected] Torsten Braun University of Bern Switzerland [email protected] ABSTRACT The Cloud Continuum Framework ( CCF ) logically integrates distributed extreme edge, far edge, near edge, and cloud data centers in 6G networks. Deploying VNFs over the CCF can enhance network performance and Quality of Service ( QoS ) for modern delay-sensitive applications and use cases in 6G networks. Deep Reinforcement Learning ( DRL ) has shown potential to automate Virtual Network Function ( VNF ) migrations by learning optimal policies through continuous monitoring of the network environment. In this work, we leverage Deep Reinforcement Learning to optimize network control policies that continuously update VNF placement for optimal Service Function Chain ( SFC ) deployment in time-varying user traffic scenarios. By leveraging dynamic VNF relocation, this approach seeks to improve network performance in terms of latency, operational costs, scalability, and flexibility. This study addresses the gap in existing solutions by jointly considering network performance requirements and migration costs, providing a more comprehensive strategy for efficient VNF deployment and management. We show that our proposed DRL-based VNF deployment method achieves a 28.8% lower delay and a 34% lower migration overhead compared to state-of-the-art baselines in a broad range of large-scale simulated scenarios, showing the proposed method’s scalability features. CCS CONCEPTS •Networks → Network architectures;Network management; Network services. KEYWORDS 6G Network Architecture,Cloud Continuum Framework,Service Orchestrator, Deep reinforcement learning ACM Reference Format: Solomon Fikadie Wassie, Antonio Di Maio, and Torsten Braun. 2025. Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks. In The 40th ACM/SIGAPP Symposium on Applied Computing (SAC ’25), March 31-April 4, 2025, Catania, Italy. ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/3672608. 3707975 Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. SAC ’25, March 31-April 4, 2025, Catania, Italy ©2025 ACM. ACM ISBN 979-8-4007-0629-5/25/03 https://doi.org/10.1145/3672608.3707975 1 INTRODUCTION Software Defined Network ( SDN ) and Network Function Virtualization ( NFV ) are key technologies that enable elastically resource provisioning for Virtual Network Function (VNF) using virtualization technology. This approach leverages the potential of SDN technology replacing traditional hardware-based network functions with software programs. Service functions are typically deployed as Service Function Chains (SFCs), which consist of multiple VNF s in a predefined sequence to deliver end-to-end services. These VNF s can be hosted in a virtualized environment on standard Commercial Off-The-Shelf ( COTS ) servers, reducing both Capital Expenditure ( CAPEX ) and Operational Expense ( OPEX ) for network operators,[1],[2]. Network Address Translation ( NAT ), Intrusion Detection and Prevention System ( IDPS ), Firewall (FW), Load Balancer ( LB ), Video Optimization controller ( VOC ), Traffic Monitoring ( TM ), WAN Optimizer ( WO ), Deep Packet Inspection (DPI), and more VNF s can be interconnected in specific predefined sequences to create SFC requests, enabling the provision of specialized network services such as Video Streaming ( VS ), Augmented Reality ( AR ), Virtual Reality, Industry 4.0 (Ind 4.0), Holographic-Type Communications, Smart Factory, Autonomous driving, Cloud gaming and tactile industrial Internet [3, 4]. The main challenge for an Internet Service Provider ( ISP ) in enhancing Quality of Service (QoS) and Quality of Experience ( QoE ) is determining the optimal VNF deployment locations to meet stringent, variable service requests. Optimal VNF placement on physical servers is crucial for network performance, OPEX , and reliability [ 5 ],[ 6 ]. Machine learning models, particularly Deep Learning ( DL ) and Reinforcement Learning ( RL ), make VNF deployment dynamic and adaptive, enabling real-time adjustments. DL handles complex high-dimensional features, while RL optimizes strategies through interaction with network states, improving performance, reliability, and service continuity, while reducing operational costs. Few studies have explored VNF deployment and migration in time-varying traffic, typically involving traffic prediction and a migration index to represent node load trends. However, this approach is complex, needs both traffic prediction and node scheduling based on load. Many works address VNF deployment, migration, and SFC reconfiguration in three stages: VNF Resource Prediction,SFC Deployment Optimization, and Destination Node Scheduling [ 7 ],[ 8 ],[ 9 ]. A DRL-based approach enables intelligent agents to monitor real-time network performance, adapt to traffic variations, and continuously improve decision-making by tracking user traffic and node status through periodic interactions and feedback from the environment. Fluctuating VNF resource requirements due to time-varying user
SAC ’25, March 31-April 4, 2025, Catania, Italy S. Wassie et al. traffic require service relocation to maintain reliability and performance. The key research question is how to optimally deploy VNFs while considering both time-varying user traffic and the variable underlying network infrastructure. We propose DRL-based approaches for automatic VNF deployment and network-state adaptive network reconfiguration. To the best of our knowledge, this paper is the first to address VNF deployment and migration from the extreme edge, through the edge, to cloud data centers, aiming for long-term operational cost benefits and network-state adaptive VNF deployment within 6G mobile network architecture. The three key contributions of this paper are as follows: • Develop a data-driven service orchestrator, which is part of the 6G network architecture management plane, to manage network-state-aware VNF deployment for delay-sensitive applications. • Develop an intelligent network function deployment management entity to place VNFs, aimed at predicting the effect of long-term operational costs, network performance, and reliability by continuously monitoring time-varying user traffic demand and network infrastructure. • Propose a novel DRL-based connectivity continuum with VNF migration from the extreme edge through the edge to the cloud over the CCF for 6G Network architecture. The remainder of the paper is organized as follows: Section 2 describes the related works and considered scenario. Section 3 presents the system model. Section 4 outlines the proposed methods. Section 4 explains the experimental setup and simulation results. Finally, Section 6 draws the conclusions. 2 RELATED WORKS Recent research has approached the problems of VNF deployment, self-scaling, and elastic resource allocation from various perspectives. We have reviewed studies from recent years that attempt to solve the VNF deployment problem across three categories: resource provisioning, VNF migration, and resource prediction and scheduling, often leveraging predictions of time-varying user traffic. While many studies focus on QoS-aware VNF deployment and migration across distributed data centers, a few approach have explored DRL as a potential solution. Significant efforts have optimized VNF placement to enhance network performance. Despite these advances, automatic and network-adaptive operational cost and long-term effect considerations, as well as efficient, reliable, and scalable VNF deployment in large-scale networks, remain challenging. 2.1 Resource provisioning Researchers have addressed the VNF placement and chaining problem as a resource provisioning issue, focusing on flexible resource allocation to meet service requirements and service level agreements. They primarily consider how much resource allocation is needed to satisfy QoS and network performance, but often overlook the impact of external traffic fluctuations [ 10 ],[ 11 ]. Furthermore, many studies tackled VNF placement in MEC-NFV networks, formulating optimization models to enhance resource utilization through deep learning techniques that intelligently select nodes and place VNFs for SFC requests. Service deployment typically involves allocating a Virtual Network Function - Forwarding Graph to meet the QoS requirements of VNFs [12],[13]. 2.2 VNF migration NFV technologies enable VNFs as software-based network services. However, frequent user mobility necessitates re-scaling and re-provisioning of VNFs. Akrem et al. [ 14 ] address this with an AI-Based Network-Aware Service Function Chain Migration for 5G, enabling low-latency slice transfers between service areas. Like Virtual Machine ( VM ) migration and serverless computing, stateful VNFs can migrate within telecom data centers, facilitating context transfer across geographically distributed setups. Li et al. [ 15 ] proposed a joint resource optimization and delayaware VNF migration method focusing on resource availability and delay constraints. He et al. [ 16 ] introduced an SLA-aware approach for multiple migration planning in SDN-NFV clouds, optimizing sequence and timing to minimize migration time and prevent QoS degradation. However, these heuristics overlook dynamic variations in link quality and computational resources over time. 2.3 SFC resource requirement predictions and scheduling This approach predicts the time-varying resource and QoS requirements of SFC requests, proactively allocating resources on available nodes based on fluctuating user traffic demands. This proactive allocation is essential for efficiently addressing the VNF deployment and resource provisioning problem. Efficient scheduling aims to reduce total deployment cost, communication cost, and enhance QoS by dynamically allocating resources based on time-varying demand. Gu et al. [ 17 ] proposed a mixed-integer linear programming solution for VNF deployment and flow scheduling in distributed data centers, considering network topology, VNF instances, and deployment. Tang et al. [ 18 ] developed a method predicting future resource needs based on time-varying user traffic and deep belief networks, addressing dynamic VNF resource requirements. The parameters considered in this method, compared with other literature, are shown in Table 1. Table 1: Comparison of Related Works Parameters End to end delay Concurrent VNF Migration Node Resource Variable Traffic Stateful VNF Migration Migration Cost [10]-2020 ✓×✓××× [16]-2020 ✓ ✓ × × ✓× [5]-2021 ×✓ ✓ × × ✓ [7]-2021 ✓ ✓ × × ✓ ✓ [14]-2022 ✓ ✓ ✓ × × ✓ [8]-2023 ✓×✓ ✓ × × [6]-2023 ✓×✓ ✓ ×✓ [13]-2024 ✓×✓××× Proposed Method ✓✓ ✓✓✓✓
Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks SAC ’25, March 31-April 4, 2025, Catania, Italy NDT Automation Optimization Orchestration Operations/Business Support Systems NFV-MANOService orchestration (SO) AI/ML Framework OSS/BSS Application layer S VNF1 VNF3 VNF5 VNF6 Extreme edge Edge Cloud Business interface Cloud continuum framework(CCF) Management & orchestration Frame work(MOF) AI models VNF2 Resource Orchestration(RO) MOF Bus AI/ML Bus Network function layer VNF4End devices smartphone drone AR/VR headset Control policy and data flow VNF Deployment action Message bus Figure 1: Illustration of Proposed Native AI 6G Network architecture with network state adaptive VNF deployment 3 SYSTEM MODEL AND PROBLEM FORMULATION 3.1 6G Network Architecture We envision a 6G network architecture depicted in Figure 1 composed of three main components: the Cloud Continuum Framework ( CCF ), the Management and Orchestration Framework ( MOF ), and the Artificial Intelligence and Machine Learning Framework ( AIMLF ).Each framework can use message buses both internally and for inter-framework communication. Specifically, the MOF message bus is used for MOF and AIMLF communication, while the AIMLF message bus facilitates communication between CCF and AIMLF. 3.1.1 Cloud Continuum Framework ( CCF ). The CCF offers a unified resource pool that orchestrates resources across multiple clouds and dynamically composes network and cloud resources from the extreme edge to central clouds based on availability and service requirements. The nodes in a CCF can be classified based on their geographical proximity to the end user and their computational capacity into the categories of Cloud,Near-edge,Far-edge, and Extremeedge. It integrates AI-driven resource management to optimize utilization and energy efficiency by predicting demand and dynamically adjusting allocations in real time. Additionally, the framework provides business interfaces for cloud providers to enhance Service level Agreements ( SLA s) and ensure security, reliability, trust, and energy efficiency. 3.1.2 Management and Orchestration Framework ( MOF ). The MOF orchestrates network services across the cloud continuum in the 6G service-oriented network, integrating various technological domains and supporting AI/ML frameworks for real-time monitoring and updates of AI-driven functions. Its distributed management approach separates concerns, federates functional domains, and delivers end-to-end network services (E2E NS). It also allows tenants to request deployment, modification, or termination of networks or applications. MEC MEC MEC MEC RAN RAN RAN RAN St+1 Sn At t+1 An VNF2 VNF 1 VNF3 Cloud DC SDN controller Wired link Migration Wireless link States fi=(Bi, Di,𝜎i ) VNF2 VNF3 St Actions Environmental information 1 A VNF deployment action 3 VNF1 2 SFC request traffic generation Figure 2: Example of physical network containing a Cloud Data Center, several interconnected MEC servers serving one or more RAN domains serving a diverse set of User Equipments (UEs), AR headsets, and IoT devices In the 6G network architecture, each Service Orchestrator (SO) is tightly integrated with the Operations Support System (OSS) and Business Support System (BSS), which handle the orchestration and lifecycle management of Network Services (NS) as a set of VNFs. The OSS manages network operations like monitoring, fault management, and performance, while the BSS oversees business tasks like billing and customer management. Together, they ensure efficient service deployment, scaling, and resource optimization. SO automate FCAPS functions—Fault, Configuration, Accounting, Performance, and Security—ensuring network health and security. At the business layer, OSS/BSS handles Lifecycle Management (LCM) requests for NS, distributing them across orchestration domains to ensure programmability and service integration across diverse cloud environments. This coordination of technical and business layers enables the architecture to adapt to service demands efficiently while maintaining seamless business operations. 3.1.3 Artificial Intelligence and Machine Learning Framework ( AIMLF ). Designed to provide unified AI/ML management and orchestration across various segments and frameworks of the 6G network architecture, the AIMLF supports the development, training, and distribution of AI/ML models. It incorporates mechanisms for continuous integration and continuous development (CI/CD) of AI/ML deployment within a service-oriented architecture. The framework is capable of evaluating and updating AI-driven functions during system runtime and utilizes resources provided by the CCF. Additionally, AIMLF employs a network digital twin to improve AI/ML training, enhance simulations behavior, and optimize AI-driven functions, increase efficiency and adaptability. 3.2 System Model End-user devices generate dynamic traffic patterns and initiate network-access requests through base stations, as shown in Figure 1. The traffic traverses through the application layer in the proposed 6G network architecture, requesting an orchestrator to
SAC ’25, March 31-April 4, 2025, Catania, Italy S. Wassie et al. deploy VNFs optimally while meeting strict performance requirements.The service orchestrator, part of the MOF, receives this traffic and determines the optimal deployment on physical servers within the CCF, over the Cloud,Near-edge,Far-edge, and Extreme-edge. We model a physical network at high level as shown in Figure 2, where Mobile Edge Computing ( MEC ) nodes are connected to the cloud via high-speed fiber. All computational and communication resources of the edge nodes are controlled by an overlay controller(i.e RL agent), which orchestrates SDN entities at the cloud DC. The network infrastructure consists of: (i) a cloud data center (DC) capable of hosting the proposed DRL algorithm and potentially deploying multiple VNFs; (ii) several MEC data centers 𝑣𝑖 , each with a CPU capacity 𝐶𝑖 [ cycle/s ] ,positioned near end-users to minimize latency; and (iii) a set of end-user devices that access network services through base stations as decipted in figure 2. We consider heterogeneous resources and varying link capacities among distributed edge servers, where delays between VNFs on different servers, dictated by dynamic user traffic, affect the optimal VNF deployment locations. In such a system, end-user devices such as smartphones, AR headsets, and IoT devices request dynamic communication and computational resources through base stations, initiating network-access requests (e.g.registration, attach requests, and radio resource requests for channel allocation). These devices produce dynamic traffic as shown in figure 2 in step 1 with strict Key Performance Indicator ( KPI ) requirements, demanding flexible resource allocation for optimal computing and communication performance.The RL agent operates as an intelligent SDN controller, processing complex state information to optimize network performance and service delivery. The RL agent continuously adapts to network state information, which serves as its input (Figure 2, step 2 ).The environmental state information includes traffic loads, resource availability (e.g., MEC CPU, memory, bandwidth), network performance metrics (e.g., delay, link utilization), and infrastructure status (e.g., edge node status). Based on this analysis, the RL agent predicts optimal VNF deployment actions and dynamically allocates resources for efficient operation, as shown in Figure 2, step 3 . To model the above described physical network, We consider a scenario involving a physical network infrastructure modeled as an undirected graph 𝐺=(𝑉, 𝐸) , where 𝑉 and 𝐸 represent the sets of physical network nodes and links between nodes, respectively. Each node 𝑣∈𝑉 represents a physical network entity, such as an extreme-edge node (i.e., a UE end device such as a smartphone, electric vehicle, or drone), and edge server, or a central cloud data center within the CCF , as depicted in Figure 1. Each link 𝑒∈𝐸 corresponds to a physical network connection, which represents high-speed fiber links between nodes. The bandwidth (BW) capacity of the physical link between nodes 𝑣𝑖, 𝑣𝑗∈𝑉 is denoted as 𝐵𝑖 𝑗 [bit/s]. We model a generic SFC as a Directed Acyclic Graph ( DAG ) 𝐻=(𝐾, 𝐿) , where 𝐾 represents the set of VNFs within the SFC and 𝐿 denotes the set of logical links between VNFs. Each VNF 𝑘∈𝐾 represents a softwarized network function that can process incoming packets. The logical links (𝑘𝑖,𝑘𝑗) ∈ 𝐿 represent the connections between successive VNFs 𝑘𝑖 and 𝑘𝑗 . The topology 𝐻 of an SFC depends on the application that the SFC aims to support, and we assume it is already determined by the tenant and submitted to the network management plane for acceptance and deployment. The position and logical order of VNFs in an SFC are critical parameter for network performance.The detailed granularity of traffic arrival at the orchestrator is depicted in Figure 3. End-to-end user traffic from the application layer generates multiple SFC requests ( 𝑓1 , 𝑓2 , 𝑓3 , ..., 𝑓𝑛 ), which are transmitted to the service orchestrator. These requests follow a standardized structure within a defined queuing model. The communication patterns encompass various interaction types, including human-to-human (H2H), machine-tomachine (M2M), and machine-to-human (M2H) communication. The service orchestrator processes the incoming traffic, identifies the relevant VNFs and their inter dependencies, and maps them to appropriate physical nodes. This process ensures optimal operation while maintaining logical connections between VNFs across the network infrastructure within the Cloud Continuum Framework. We define an SFC request as a tuple 𝑓𝑖=𝐻𝑖, 𝐵min 𝑖, 𝐷max 𝑖, 𝜎𝑖 where 𝐻𝑖=(𝐾𝑖, 𝐿𝑖) is the SFC’s topology, 𝐵min 𝑖[bit/ s ] represents the minimal end-to-end bandwidth requirement, 𝐷max 𝑖[ s ] is the maximum allowable end-to-end delay, 𝜎𝑖[cycle/ s ] denotes the overall SFC’s computational capacity requirement. As the incoming traffic pattern from user traffic fluctuates over time, by taking into account the network design, packet arriving at the orchestrator follows a Poisson distribution with mean arrival intensity rate 𝜆𝑖 . We denote the sequence of SFC requests arriving to the network’s resource orchestrator for being deployed onto the physical network as 𝐹=(𝑓1, 𝑓2, . . . , 𝑓𝑛) , where 𝑓𝑖 indicates the 𝑖 -th SFC request in the queue. Each VNF 𝑘∈𝐾 is associated with a specific state, making stateful migration essential for maintaining service continuity and optimizing network performance. The target node selection for each state can be modeled as a tuple: S𝑘=(𝑀𝑖, 𝐷𝑐, 𝑄𝑣, 𝑃𝑠,𝑇𝑚) , where 𝑀𝑖[ B ] is the size of the context information to be migrated, 𝐷𝑐 represents service deployment costs, 𝑄𝑣 is the impact of SLA violations during migration, 𝑃𝑠 is selected path congestion status, and 𝑇𝑚[ s ] is the total migration time. Stateful migration preserves active sessions and data processing, minimizing disruptions. The inability to relocate services may result in failures, leading to interruptions and increased delays. Our proposed method can also be applied to systems where SFC are modeled as generic Directed Acyclic Graphs (DAGs), supporting emerging 6G mission-critical applications with stringent KPI requirements. The ingress traffic from the application layer first enters the NAT as inbound traffic for many network services, while the outbound traffic from the last VNF often exits through the IDPS. Several studies [ 19 ] show that egress traffic depends on the network services and is not unique to specific VNFs (e.g Ind 4.0 traffic exit through FW). To illustrate how traffic traverses through VNFs for various applications, VNFs are logically arranged in sequence. For example, VNFs are organized as 𝐾=(NAT,FW,TM,VOC,IDPS) for video streaming, and as 𝐾=(NAT,TM,encryption,decompression,decryption) for autonomous vehicles. An augmented reality (AR) application typically employs a linear SFC request 𝑓𝑖 , with VNFs ordered as 𝐾=(NAT,FW,VOC,TM,WO,IDPS) . Incoming traffic first enters the NAT for address translation, then passes through the firewall (FW) to filter unauthorized traffic. Next, it goes through the VOC
Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks SAC ’25, March 31-April 4, 2025, Catania, Italy VK1 S d Cloud Continuum Network Infrastructure f2 f3 . . . . fn Physical network Virtual link SFC logical link Embedding Extreme edge Edge Central cloud VK4 VK5 VK6 VK7 VK8 VK9 d S S VK2 VK3 VK4 VK5 VK6 VK7 d VK1 VK2 VK3 f1 VK3 VK5 VK1 VK4 VK6 VK7 Service orchestrator VK9 VK2 VK1 L1 L2 L3 L4 L5 Figure 3: SFC deployment with generic structure for video quality and bandwidth optimization. The traffic manager (TM) analyzes patterns, while the WAN Optimizer (WO) enhances performance by optimizing data flow and reducing latency. Finally, the IDPS scans for malicious activity. 3.3 Problem Formulation We formulate the problem of elastically auto-scaling VNF deployment over a physical network, aiming to determine the optimal physical location of VNFs by minimizing the latency of SFC requests across the physical network. This latency comprises propagation delay, transmission delay, queuing delay, and processing delay. However, we consider the transmission delay and VNF computing delay. The optimization problem is formulated as follows. We define the SFC allocation vector 𝛼=(𝛼1, . . . , 𝛼|𝐾|) ∈ 𝑉|𝐾| , where each component 𝛼𝑖 represents the physical node in 𝑉 on which the 𝑖 -th VNF in the SFC is deployed on. Let us define the VNF processing latency 𝑃(𝛼𝑖) on node 𝛼𝑖 as the time needed by the 𝑖 -th VNF to perform its computing task when deployed on node 𝛼𝑖 , and we define the worst-case SFC processing latency as the sum 𝑃(𝛼)=Í|𝐾| 𝑖=1𝑃(𝛼𝑖) of all SFC’s VNFs’ processing delays. We also define the VNF communication latency 𝑙(𝛼𝑖, 𝛼𝑗 ) as the shortest-path latency between VNF deployment locations 𝛼𝑖 and 𝛼𝑗 over the physical links 𝐸 , and we define the worst-case SFC communication latency as the sum Γ(𝛼)=Í(𝑖,𝑗 ) ∈𝐿𝑙(𝛼𝑖, 𝛼𝑗) of all VNF communication latency’s over all SFC’s links. Finally, we define the total SFC delay as the sum of the SFC processing and communication latancies, i.e., 𝑃(𝛼) + Γ(𝛼) . Even though the functions 𝑃 and Γ depend on the physical network topology 𝐺 and the SFC topology 𝐻 , and their time-variant properties, we drop such dependency in the notation for conciseness. Another parameter we considered in formulating objective function to select the optimal location for a stateful VNF is its migration time, which depends on the size of the VNF and the throughput of all links constituting the shortest path from the VNF’s source node to a candidate target relocation node. We define the SFC’s 𝑘 -th VNF migration time 𝑡𝑘=Í(𝑖,𝑗 ) ∈𝑝𝑘 𝑀𝑘 𝐵𝑖 𝑗 as the sum of all the times needed to transmit the VNF’s state of size 𝑀𝑘 over all physical links (𝑖, 𝑗) ∈ 𝑝𝑘 that constitute the shortest path 𝑝𝑘 from the VNF’s previous deployment location to the new candidate deployment node 𝛼𝑘 , which depends on time-varying VNF state size and network conditions. We define the worst-case SFC migration time 𝑇(𝛼)=Í𝑘∈𝐾𝑡𝑘 as the sum of all VNF migration times, from their deployment locations to their respective candidate target nodes represented by 𝛼 . It is worth noting that the migration time might not be the only cost operator incur for migrating SFCs, for example adding economical or energy expenses for infrastructure activation, SLA violations, bandwidth quota excess, the size of migrated VNF context information, energy consumed for migration task, et cetera. Therefore, we consider 𝑇(𝛼) as a more generic definition of SFC migration cost that may not necessarily be expressed in latency but includes other financial, energy, and resource aspects. Given that the optimization problem is multi-objective, we define 𝛽=(𝛽1, 𝛽2) as the weight factors that balance the trade-off between delay and migration cost when selecting physical nodes, and aim to minimize 𝛽⊤𝐶(𝛼) , where 𝐶=(𝑃+Γ,𝑇 )(𝛼) . Let us define 𝑛𝑘 𝑣 as the cycles/ sused by a VNF 𝑘 when deployed on node 𝑣∈𝑉 , and 𝑏𝑘𝑙 𝑖 𝑗 as the bandwidth in bit/ sused by an SFC’s logical link (𝑘, 𝑙) if deployed over the physical link (𝑖, 𝑗) ∈ 𝐸. The optimization problem’s goal is to find the optimal SFC allocation vector 𝛼∗ for each SFC request 𝑓𝑖 , which minimizes the objective function 𝛽⊤𝐶(𝛼) under a set of infrastructure-induced constraints, as in Equation 1. minimize 𝛼∈𝑉|𝐾| Total SFC delay z }| { 𝛽1· (𝑃(𝛼) + Γ(𝛼)) + SFC migration cost z }| { 𝛽2·𝑇(𝛼)(1) subject to ∑︁ 𝐾𝑖:𝑖∈[𝑛]∑︁ 𝑘∈𝐾𝑖 𝑛𝑘 𝑣≤𝐶𝑣,∀𝑣∈𝑉(1a) ∑︁ 𝐿𝑖:𝑖∈[𝑛]∑︁ (𝑘,𝑙 ) ∈𝐿𝑖 𝑏𝑘𝑙 𝑖𝑗 ≤𝐵𝑖 𝑗,∀(𝑖, 𝑗) ∈ 𝐸(1b) 𝑃(𝛼) + Γ(𝛼) ≤ 𝐷max (1c) Constraint 1a imposes that the sum of the processing requirements of all VNFs deployed on each node in the system should be less than the locally available computational capacity. Constraint 1b indicates that the bandwidth usage of all SFC requests must remain within the available bandwidth capacity of the network logical links. Finally, Constraint 1c implies that the total communication and processing delay for an SFC request does not exceed the E2E delay tolerance limit required for the successful completion of the given service.
SAC ’25, March 31-April 4, 2025, Catania, Italy S. Wassie et al. 4 METHODOLOGY 4.1 Deep Reinforcement Learning Approach In this section, we redefine the optimization problem in Equation 1 with the context of RL, where the RL agent performs VNF deployment actions by continuously monitoring the physical network infrastructure and receiving feedback in the form of rewards. The goal is to derive an optimal control policy that enables the agent to select VNF deployment actions optimally, considering the future network state. To achieve this, we employ DRL with deep neural networks, enabling the agent to manage complex environments and high-dimensional state spaces. The proposed DRL-based Proximal Policy Optimization ( PPO ) agent identifies optimal physical nodes to meet application delay requirements by calculating delays between VNF-hosting nodes, taking into account session information size, service deployment costs, SLA violations, network congestion, energy use for active transfers, and total migration time. It refines its strategies by leveraging rewards from actions within a Markov Decision Process (MDP). The problem is formulated in terms of a state space 𝑆 , an action space 𝐴 , and a reward function 𝑅 . We demonstrate how PPO solves the problem using both a value network and a policy network. The discount factor 𝛾∈ [ 0 , 1 ) is used to mathematically represent continuing tasks. The main objective of the RL agent is to discover a policy that maximizes the expected sum of discounted future rewards, known as the return, given by 𝑅𝑡=Í∞ 𝑖=0𝛾𝑖+𝑡𝑟𝑡+𝑖 . The optimal policy, given by 𝜋∗=arg max𝜋E𝜋{𝑟0|𝑠0=𝑠} , is the one that maximizes the expected return from any given state 𝑠 . The detailed description of the state space, action space, and reward function is as follows. 1) State space 𝑆𝑡 :Describes the current situation of the agent in the environment for our VNF deployment problem, designed as a vector initially randomly deployed with 𝐷= {𝑉𝐾𝑗 𝑖,𝑉𝐾𝑗+1 𝑖+1, . . . ,𝑉 𝐾𝑛 𝑛} in the physical network which represents VNFs 𝐾𝑗, 𝐾𝑗+1, . . . , 𝐾𝑛 that are deployed on physical nodes 𝑉𝑖,𝑉𝑖+1, . . . ,𝑉𝑛 at time 𝑡0 , as well as the link delay between the source and destination nodes of the physical nodes hosting the VNFs in an SFC. The state S is defined as 𝑆=𝑉𝐾𝑗 𝑖,𝑡 , 𝐿𝐾𝑗 𝑉𝑖,ℎ ∀𝑖∈𝑉𝑛,∀𝑗∈𝐾𝑛,(2) , where 𝑉𝐾𝑗 𝑖,𝑡 means that VNF 𝐾𝑗 is deployed on physical server 𝑉𝑖 at time 𝑡 , and 𝐿𝐾𝑗 𝑉𝑖,ℎ denotes the link delay 𝐿𝑉𝑖 ℎ,𝑡 between the source node hosting VNF 𝐾𝑗 at 𝑉𝑖 and the destination node where VNF 𝐾𝑗+1 is hosted on physical server 𝑉ℎ . Where Both (𝑣𝑖, 𝑣ℎ) ∈ 𝑉2. 2) Action space 𝐴𝑡 :The agent explores where are the optimal locations of physical nodes to host VNFs that meet the QoS requirements of incoming traffic. The action space provides boundaries the agent how to search the physical nodes regarding VNF deployment, with each action representing the selection of a sequence of nodes that satisfies the QoS requirements. The agent will select a number of nodes equal to the length of the VNFs in the SFCs at each time step 𝑡. 𝐴=(𝑉𝐾𝑖 1, . . . ,𝑉 𝐾𝑛 𝑛),(3) L= 3ms B=100 Mbps CPU: 3GHz RAM:9GB memo:36GB CPU: 3GHz RAM:9GB Memo:36GB CPU: 4GHz RAM:12GB memo:42GB CPU:6GHz RAM:15GB memo:256GB VNF1 VNF2 VNF3 L=4ms B=120Mbps L=5ms B=50Mbps L=4ms B=40Mbps L=12ms B=60 Mbps Enviroment Reward Rt Policy network State st Action At Vϕ (st) Value network State st Figure 4: Value and policy network for VNF deployment. , where 𝑉1, . . . ,𝑉𝑛 represent the number of nodes selected at each time step 𝑡1, . . . , 𝑡𝑛. 3) Reward Function 𝑅𝑡 :The reward function quantitatively measures the impact of decisions on network operations. The agent needs to quantify the quality of the deployment location of physical node by learning from the pattern extracted from environment. The QoS utility function is given by 𝑄𝑗=𝛼𝜇𝑗·𝑆𝑗(𝐷𝑗) + 𝛼𝑓 𝜇𝑗·𝑇𝑗(𝑊𝑗,𝐶𝑗) where 𝛼𝜇𝑗 and 𝛼𝑓 𝜇𝑗 , are weighting parameters used to prioritize user satisfaction 𝑆𝑗(𝐷𝑗) ,and the cost savings score 𝑇𝑗(𝑊𝑗,C𝑗) for the given traffic request. Assume 𝑉={𝑉1, . . . ,𝑉𝑛} as the set of all possible locations where VNFs can be deployed and 𝐾={𝐾1, . . . , 𝐾𝑛} are possible VNFs to be deployed on available locations of physical server. Equation 2 represents the delay between VNF 𝐾𝑗 running at the source node 𝑉𝑖 and VNF 𝐾𝑗+1 running on the target node 𝑉ℎat time 𝑡. The reward function is given by 𝑅𝑡=−𝜔1∑︁ (𝑣𝑖,ℎ) ∈𝑉2 𝐿𝐾𝑗 𝑉𝑖,ℎ −𝜔2∑︁ 𝑓∈𝐹 𝑄𝑖,(4) , where 𝜔1 and 𝜔2 prioritize latency reduction and maximize QoS respectively, 4.2 Proximal Policy Optimization for autonomous VNF deployment PPO is policy-based DRL algorithm, directly updates its policy based on observed rewards, allowing for rapid adaptation in dynamic environments. Its use of policy gradients makes it highly efficient and adaptable, making it a valuable choice for optimization problems in complex and changing network environment. As shown in Figure 4, the state space consists of the network topology, current VNF placements, resource utilization, and KPI requirements for incoming traffic. The RL agent makes VNF placement decisions by considering SFC routing choices and resource
Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks SAC ’25, March 31-April 4, 2025, Catania, Italy allocation. The reward is based on improvements in network performance and the efficiency of resource utilization. The functionalities of each component in this reinforcement learning framework are detailed as follows. Algorithm 1: Autonomous VNF deployment Input: 𝑓𝑖=(𝐻𝑖, 𝐵min 𝑖, 𝐷max 𝑖, 𝜎𝑖),𝛼,𝑇max,𝐺topo,𝑙(𝑖, 𝑗) 1 Initialization: 𝜙0,𝜋0// Initialize 𝜙,𝜋parameters 2𝑆𝑡← {𝑉𝐾𝑗 𝑖,𝑡 , 𝐿𝑉𝑖,𝑗 𝑡}// Initial VNF location 3𝑅𝑡←0// Initial the reward to zero 4 for 𝑡∈𝑇max do 5𝑐←MeasureCPUrequirment(𝑓𝑖) 6𝑏←MeasureBandwidth(𝑓𝑖) 7𝑙←MeasureLatency(𝑓𝑖) 8𝑞←MonitorLinkquility(𝐺topo) 9𝑆𝑡← (𝑙,𝑏,𝑐,𝑞)// Assign 𝑓𝑖KPI requirement 10 𝑆𝑡 𝐴𝑡,𝜋𝜃(𝐴𝑡|𝑆𝑡) −−−−−−−−−−−→ 𝑆𝑡+1// Select 𝐴𝑡∼𝜋𝜃 11 𝑆𝑡+1←𝑉𝐾𝑖 𝑡// assign VNF 𝐾𝑖new location 12 𝐷𝑡← {𝑠0, 𝑎0,𝑟0, . . . ,𝑠𝑡, 𝑎𝑡, 𝑟𝑡}// collect state action interaction 13 if 𝑙(𝑓𝑖)<𝑙(𝑖, 𝑗)then 14 𝐴𝑡∼𝜋𝜃(𝐴𝑡|𝑆𝑡)Execute 𝐴𝑡 −−−−−−−−−→ 𝑅𝑡+1, 𝑆𝑡+1 15 𝐴𝜋𝜃𝑡(𝑠𝑡, 𝑎𝑡) ← 𝑄(𝑠, 𝑎) − 𝑉𝜙(𝑠)// Advantage estimate // Update 𝜃and 𝜙parameters 16 𝜃𝑡+1=arg max𝜃1 | D𝑡|𝑇Í𝜏∈ D𝑡Í𝑇 𝑡=0 17 min 𝜋𝜃(𝑎𝑡|𝑠𝑡) 𝜋𝜃𝑡(𝑎𝑡|𝑠𝑡)𝐴𝜋𝜃𝑡(𝑠𝑡, 𝑎𝑡),𝑔(𝜖, 𝐴𝜋𝜃𝑡(𝑠𝑡, 𝑎𝑡)) 18 𝜙𝑡+1=arg min𝜙1 | D𝑡|Í𝜏∈ D𝑡Í𝑇 𝑡=0(𝑉𝜙(𝑠𝑡) − 𝑅𝑡)2 19 else 20 SFC_req ←rejected 21 return 𝜙𝑖(𝑠𝑡), 𝜋𝜃(𝑠𝑡) PPO develops an optimal policy by utilizing the collaboration between value and policy networks for decision-making optimization. The detailed working principle of how the value and policy networks monitor the environment state and discover the optimal policy is illustrated in (Algorithm 1). The value network estimates the long-term effectiveness of network states and reconfiguration actions, factoring in VNF placements, SFC configurations, resource utilization, and traffic demands. This estimation, represented by 𝑉𝜙(𝑠)=EÍ∞ 𝑡=0𝛾𝑡𝑅𝑡+1|𝑆𝑡=𝑠 guides the policy network. Here, 𝑉𝜙(𝑠) is parameterized by 𝜙 , 𝛾 is the discount factor, 𝑅𝑡+1 is the reward at time 𝑡, and 𝑆𝑡=𝑠specifies the state at time 𝑡. In RL, a policy 𝜋 defines how an agent acts based on the observed state. It guides the agent actions in response to environmental states. The policy network maps states to actions, aiming to maximize cumulative rewards by increasing the probability of high-reward actions and decreasing that of less effective ones. PPO refines this approach using the advantage function 𝐴(𝑠, 𝑎)=𝑄(𝑠, 𝑎) −𝑉𝜙(𝑠) to evaluate actions and constrain updates to stay close to the current policy. Large updates to the policy can lead to significant changes in the behavior of the agent and an unstable training process. 𝐴(𝑠, 𝑎) quantifies how much better or worse an action 𝑎 in state 𝑠 is compared to the expected outcome under the current policy. A positive 𝐴(𝑠, 𝑎) suggests the action is beneficial; a negative value suggests a suboptimal choice that does not maximize the expected return, indicating the need to lower the probability of selection. L(𝑠, 𝑎, 𝜃𝑘)=min 𝜋𝜃(𝑎|𝑠) 𝜋𝜃𝑘(𝑎|𝑠)𝐴(𝑠, 𝑎), clip 𝜋𝜃(𝑎|𝑠) 𝜋𝜃𝑘(𝑎|𝑠),1−𝜖, 1+𝜖𝐴(𝑠, 𝑎)(5) The loss function L(𝑠, 𝑎, 𝜃𝑘) evaluates the policy 𝜋𝜃 at state 𝑠 for action 𝑎 , where 𝜃𝑘 are the old policy parameters. The terms 𝜋𝜃(𝑎|𝑠) and 𝜋𝜃𝑘(𝑎|𝑠) represent the probabilities of taking action 𝑎 under the current and old policies, respectively. The advantage function 𝐴(𝑠, 𝑎) estimates the improvement in reward for action 𝑎 at 𝑠 relative to the average action under 𝜋𝜃𝑘 . To limit large updates, clip 𝜋𝜃(𝑎|𝑠) 𝜋𝜃𝑘(𝑎|𝑠),1−𝜖, 1+𝜖 constrains the policy ratio within [ 1 −𝜖, 1 +𝜖] , where 𝜖 ensures stability by preventing excessive deviations. 4.3 Autonomous VNF deployment algorithm Step 1: The input for the PPO algorithm includes physical network communication, computational resources, and SFC requests with specific requirements for bandwidth, delay, computing capacity, and memory. Step 2: The PPO model, with both value and policy networks, initializes the parameters 𝜙0 and policy 𝜋0 as decipted in (Algorithm line 1). Step 3: Start with a random VNF deployment 𝑆𝑡={𝑉𝑘 𝑖, 𝐿𝑘 𝑖} , which represents both the physical location and the link quality and initialize the reward 𝑅𝑡= 0(Algorithm 1 line 2-3). Step 4: Measure KPI Requirements of SFC request 𝑓𝑖 such as bandwidth,CPU,latency and observe the available resources in the physical network. Those are translated as state and taken as input to the value network and policy network (Algorithm 1 line 4-9). Step 5: Select deployment action based on the current policy 𝜋𝜃 . Update the system state and select the new VNF location 𝑆𝑡+1=𝑉𝑘 𝑖 (Algorithm 1 line 10-11). Step 6: Compare the incoming traffic requirements with available infrastructure resources: 𝑙(𝑓𝑖)<𝑙(𝑖, 𝑗) for the allocation vector. Compute the reward 𝑅𝑡 , the advantage 𝐴𝜋𝜃𝑡(𝑠𝑡, 𝑎𝑡) , and assign the new VNF location 𝑆𝑡+1=𝑉𝑘 𝑖 (Algorithm 1 line 11-15). Step 7: Update the policy and value function 𝜃𝑡+1 and 𝜙𝑡+1 , and repeat this iteration until the optimal policy is developed (Algorithm 1 line 15-18). Step 8: Decision about the SFC acceptance and SFC request rejection (Algorithm 1 line 13-20). 5 EXPERIMENTAL EVALUATION AND RESULTS 5.1 Simulation Setup The simulation experiments are conducted using a NetworkX-based Python simulator to generate the network topology and infrastructure. For the DRL implementation, we employ the open-source tools Open AI Gymnasium and Stable Baselines3 for training and testing the DRL-based PPO agent in a customized RL environment inspired by the USA NET network topology [ 20 ]. The performance of the proposed DRL-based solution is evaluated across various
SAC ’25, March 31-April 4, 2025, Catania, Italy S. Wassie et al. Table 2: Simulation Parameters Parameter Value Number of nodes (𝑁){10,20, . . . , 60} SFC length (|𝐾|){3,5,7,9, . . . , 13} Maximum steps (𝑇max)106 Discount factor (𝛾) 0.99 Clip ratio (𝜖) 0.2 Number of epochs (𝑁epoch)50 Batch sizes (𝑁batch) {64,128,256} Learning rate (𝛼) {0.0005,0.003,0.001} Coefficient (𝛽) 0.01 Clip range 0.1 Generalized Advantage Estimation (GAE) 0.95 network settings, with topologies ranging from 10, 20 to 60 nodes. Given the dynamic nature of incoming traffic, each SFC request 𝑓𝑖 is managed to maintain the quality of the link, while the PPO agent monitors the available communication resources in the background and adapts VNF deployment when necessary. We model both the state space 𝑆 and action space 𝐴 as discrete space. Given PPO suitability for discrete spaces and its stability and convergence properties, we select PPO for this scenario. PPO employs a stochastic policy that refines over time to favor rewarding actions. Proper hyper parameter tuning balances exploration and exploitation, avoiding local optima and improving solution quality. Large policy updates can cause performance collapse, so PPO uses a surrogate loss to keep updates within a safe range. Simulation & hyper parameters limiting policy updates are shown in Table 2. The clip ratio ( 𝜖 ) in PPO prevents large policy updates, while the regularization coefficient ( 𝛽 ) controls entropy and value function effects. The clip range (0.1) ensures stable updates, and Generalized Advantage Estimation (GAE) balances bias and variance. 5.2 Baseline and reference for proposed method To evaluate the performance of the proposed DRL-based VNF deployment, we compare it to a baseline greedy algorithm . The greedy algorithm, known for its effectiveness in optimization tasks, makes locally optimal choices at each step to maximize or minimize an objective. This serves as a valuable benchmark for assessing against distributed data-driven approach, as it consistently selects the best available option without necessarily considering the global optimum with short-sighted decisions. To address the VNF deployment challenge, the greedy algorithm selects random source and destination nodes, deploying VNFs and choosing the lowest-latency paths between them based on physical link delays. Though focused on immediate gains, it recognizes that locally optimal choices may not ensure globally optimal VNF placement and network performance. 5.3 Evaluation Metrics To investigate the performance of the proposed DRL model, we conducted simulations with varying parameters to measure the agent’s performance within the environment. Furthermore, we utilized several evaluation metrics: average reward, loss function, acceptance ratio, VNF request violation ratio, and migration overhead. Algorithm 2: Baseline greedy VNF allocation Input: 𝐺topo,𝑇max,𝑓𝑖,𝑉 𝑁𝐹_𝑖𝑛𝑓 𝑜 Output: 𝑉𝑘 1Initialization: 𝛼𝑖=(𝑉1, . . . ,𝑉𝑛),𝑃∗ path =0,𝐿∗ path =0 2for 𝑡∈𝑇max do 3src,dst ←Sample (𝐺topo,|𝐾|) 4while src,dst ∉𝑉𝑘do // Randomly sample nodes 5src,dst ←Sample (𝐺topo,|𝐾|) // Select the shortest path 6𝑃path ← Pshort [src][dst] 7𝐿P=Ílen( P) 𝑣𝑖,𝑣ℎ∈𝑣dst 𝐿𝐾𝑗 𝑉𝑖,ℎ 8for 𝑉∈ [src,dst]do 9if 𝑎𝑙𝑙𝑜𝑐_𝑠𝑟𝑐 ≠𝑉then 10 𝐿∗ path =Í𝑣𝑖∈ Pshortest 𝑣src→𝑣dst 𝐿𝐾𝑗 𝑉𝑖,ℎ // Sum the selected path delay 11 𝐿total =Í𝑣𝑖∈ Pselected_path 𝑣src→𝑣dst 𝐿(𝑣𝑖, 𝑣𝑗) 12 if 𝐿P<𝐿∗ path then 13 𝑃∗ path ← (src,dst) 14 allocate 𝑉[𝑡] ← 𝑉∗ 15 return 𝑉𝑘,𝑉∗ • Reward: Defined as the end-to-end delay between SFC-deployed physical servers, as described in Equation 4. • Loss Function: The PPO model convergence performance for VNF deployment is described in Equation 5. A smaller loss indicates better performance. • Acceptance Ratio: The proportion of accepted VNF requests relative to the total incoming requests. It indicates the percentage of accepted requests. • QoS violation Ratio: It is the fraction of VNF requests that fail to meet QoS constraints, calculated relative to the total number of received requests. • Migration Overhead: It measures the computational, energy, and bandwidth costs of migrating VNF instances to optimize resource use and service quality. 5.4 Simulation Results Figure 5a shows the cumulative reward per time step during the training of the PPO algorithm for VNF deployment and migration, compared to the analytical, greedy-based VNF allocation method. This results highlights the superior learning capability of intermediate learning rates for the PPO algorithm against the myopic approach of the greedy algorithm that is unable to learn effective strategies. The cumulative reward reflects the PPO agent’s decisionmaking, with initial variability due to extensive exploration. These fluctuations highlight the dynamic learning process as the PPO model optimizes VNF allocation. In contrast, the greedy approach,
Deep Reinforcement Learning for Context-Aware Online Service Function Chain Deployment and Migration over 6G Networks SAC ’25, March 31-April 4, 2025, Catania, Italy 0 200 400 600 800 1000 Time Step 0.08 0.06 0.04 0.02 0.00 Reward ×103 Greedy VNF allocation PPO lr=0.0001 PPO lr=0.003 PPO lr=0.0005 (a) Comparison of reward convergence in terms of negative latency between the DRL-based VNF deployment agent and the analytical greedy method. 0.00 0.25 0.50 0.75 1.00 Time Step ×105 0.00 0.25 0.50 0.75 1.00 Loss Function ×105 Learning rate LR 0.001 LR 0.003 LR 0.0005 (b) Convergence of the loss function for varying learning rates, highlighting superior performance of intermediate values of learning rate 30 40 50 60 Number of Nodes N 0 200 400 600 800 Service Relocation Cost SFC Length 5 SFC Length 7 SFC Length 9 SFC Length 11 SFC Length 13 (c) Impact of network size and SFC length on the migration overhead in terms of number of VNF migrations Figure 5: Learning capability of the DRL-based VNF deployment agent against greedy-based VNF allocation for various learning rates and an SFC composed of 13 VNFs in a network containing 60 nodes, and a parameter study of network scalability on migration overhead which ignores future states, also shows high variability. The comparison shows that PPO consistently outperforms the greedy method, offering more stable and efficient results. Figure 5b illustrates the effect of learning rates 𝛼= 0 . 001,0 . 003, and 0 . 0005 on loss convergence during training. Initially, the DRL PPO agent explores randomly, leading to higher loss. As training progresses, the agent refines its strategy, reducing the loss toward 0, indicating improved decision-making and convergence of the VNF deployment policy. Figure 5c describes the relationship between the number of nodes and the service relocation cost for various SFC lengths. As the number of nodes increases from 30,40 to 60, the service relocation cost shows a rising trend across all SFC lengths. Notably, longer SFCs, such as those with lengths 11 and 13, exhibit significantly higher relocation costs compared to shorter SFCs, like those with lengths 5 and 7. This indicates that as the network grows and the complexity of the SFC increases, the cost associated with relocating services also increase, with larger SFCs incurring disproportionately higher costs. Figure 7 shows the impact of communication overhead across various network configurations as the number of VNFs per SFC increases from 5,7 to 13. The results indicate that the number of migration (i.e migration overhead ) rises with the number of VNFs. Notably, the network setting with 60-node consistently experiences the highest communication overhead, suggesting larger networks have greater capacity for VNF migration. In contrast, the network with 30-node configuration exhibits small number of migration which means indirectly lower communication overhead, highlighting the importance of managing network size and VNF allocation to minimize overhead. Figure 6a shows the VNF acceptance and link violation ratios over time as the PPO algorithm learns to meet QoS constraints. Initially, VNF requests are rejected, and the link violation ratio is high. As the agent refines its policy, VNF placement improves, reducing violations. Eventually, the PPO agent converges to a policy that maximizes VNF acceptance and minimizes link violations, demonstrating improved system performance and efficient QoS management. The scalability of the DRL-based PPO system is evaluated by analyzing its response to increasing workloads. Figure 6b shows that as the number of physical servers in edge nodes increases, E2E traffic delay decreases for different SFC lengths. More servers reduce latency, allowing the RL agent to identify optimal deployment locations. Additionally, shorter SFCs result in lower latency, with E2E delay rising as SFC length increases. Overall, more physical servers enhance network performance by reducing delay and providing more deployment options for variable latency links. Figure 6c illustrates scalability across SFC lengths and network configurations. As SFC length increases, latency also rises, with longer SFCs leading to higher delays and reduced efficiency. Configurations with fewer VNFs per SFC perform better. Notably, larger networks (e.g., 45 nodes) show lower latency than smaller ones (e.g., 10 nodes), emphasizing the benefits of larger networks in reducing delays. This underscores the importance of strategic VNF placement to optimize latency and service delivery in varying network sizes. 6 CONCLUSION This paper addresses the problem of network state-adaptive(i.e network context aware) optimal VNF deployment and migration while minimizing E2E delay. VNFs are organized in a predefined sequence to satisfy the strict delay requirements of SFC requests and accommodate fluctuating communication resource demands. Unlike analytical algorithm for network load and failure detection techniques, the DRL model adapts to real-time network changes by continuously updating its decision-making policies. The proposed approach demonstrates superior performance compared to baseline methods by reducing delay link violations, improving VNF request acceptance rates, and minimizing E2E latency through optimal VNF deployment decisions. Furthermore, this work can be extended using a federated reinforcement learning to support the new "network of networks" concept for 6G network architecture