DEEP LEARNING–DRIVEN SIGNAL OPTIMIZATION FOR 6G WIRELESS SYSTEMS: MODELS, METHODS, PERFORMANCE EVALUATION
Full text
20 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about DEEP LEARNING–DRIVEN SIGNAL OPTIMIZATION FOR 6G WIRELESS SYSTEMS: MODELS, METHODS, PERFORMANCE EVALUATION Muhammad Essa Siddique PhD Scholar (Information Technology), Dr. A.H.S Bukhari Postgraduate Center of ICT, FET University of Sindh Jamshoro Email: [email protected] Arika Masters Applied Mathematics, Department Of Mathematics, Abdus Salam School of Mathematical Sciences GCU, Lahore Email:[email protected] Mohammad Qamar Qureshi Anglia Ruskin University Email: Qama[email protected] The sixth generation (6G) of wireless networks is envisioned to deliver unprecedented capabilities in terms of spectral efficiency, ultra-low latency, reliability, and energy efficiency. Meeting these requirements requires fundamentally new approaches to physical-layer signal optimization, as traditional convex and iterative optimization methods become computationally prohibitive in large-scale, dynamic environments. Deep learning (DL) has emerged as a promising tool to address these challenges, offering universal approximation, rapid inference, and the ability to integrate domain knowledge into optimization frameworks. Recent advances span supervised learning for beamforming, deep unfolding of iterative algorithms such as WMMSE, reinforcement learning for power control and interference management, and DLbased phase optimization for reconfigurable intelligent surfaces (RIS). DL also enables joint end-to-end optimization across beamforming, power allocation, and RIS configurations, while knowledge-driven paradigms embed communication-theoretic constraints to enhance interpretability and generalization. However, challenges remain in terms of robustness to imperfect channel state information, constraint enforcement, scalability, and deployment feasibility. This paper surveys and synthesizes DL-driven signal optimization models, methods, and performance evaluations, providing a comprehensive perspective on their role in shaping practical and efficient 6G wireless systems. Keywords: 6G, Deep Learning, Signal Optimization, Beamforming, RIS, Deep Unfolding, Reinforcement Learning, Knowledge-Driven Learning INTRODUCTION The vision of sixth-generation (6G) wireless systems is to deliver extreme performance gains A B S T R A C T
21 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about over 5G, including terabit-per-second data rates, sub-millisecond latency, ultra-reliability, and high energy efficiency [32], [34]. These stringent requirements are driven by emerging applications such as extended reality (XR), autonomous driving, holographic communications, and large-scale Internet of Things (IoT) [7], [38]. Meeting these demands requires a fundamental rethinking of physical-layer (PHY) signal optimization, as conventional convex optimization and iterative algorithms, while powerful, often fail to scale in high-dimensional, rapidly time-varying 6G environments [1], [4]. Deep learning (DL) offers a powerful alternative. Once trained, neural networks can approximate near-optimal solutions to PHY optimization tasks with orders-of-magnitude lower latency than iterative solvers [13], [15]. DL also adapts naturally to system non-linearity, hardware impairments, and non-Gaussian noise, which are difficult to capture with analytical models [16], [20]. Moreover, DL enables end-to-end optimization across multiple signal processing blocks, unifying tasks such as beamforming, power allocation, and resource scheduling that are typically optimized separately [25], [29]. Several DL paradigms have emerged for signal optimization: Supervised learning, which learns from near-optimal labels produced by conventional solvers [17], [18]. Deep unfolding, which unrolls iterative algorithms like WMMSE into neural networks with learnable parameters [3] [15] Deep reinforcement learning (DRL), which models signal control as a sequential decisionmaking problem [23] [27] Hybrid knowledge-driven methods that embed communication-theoretic constraints and structures into DL architectures [2], [32] These approaches have shown promise in applications such as massive MIMO precoding [21], RIS phase optimization [5], [6], [28], and joint PHY/MAC resource allocation [9], [22]. Despite these advances, major challenges remain: DL models often require large-scale training data and may generalize poorly across different network configurations, such as varying antenna numbers, user densities, and propagation environments [24], [29]. Constraint satisfaction (e.g., unit-modulus for RIS, transmit power budgets) remains nontrivial [27], [33]. Furthermore, interpretability, robustness to imperfect CSI, and integration into practical wireless protocols are still open issues [7], [31], [40]. This paper provides a comprehensive survey of deep learning–driven signal optimization for 6G wireless systems, with contributions summarized as follows: We review system models and formalize signal optimization objectives (sum rate, energy efficiency, fairness, robustness). We categorize DL paradigms for signal optimization, including supervised, unsupervised, unfolding, reinforcement, and hybrid knowledge-driven learning. We synthesize application-specific advances in beamforming, power allocation, RIS phase optimization, and joint PHY-layer design. We discuss evaluation strategies, metrics, datasets, and baselines for benchmarking DL-based approaches. We highlight challenges and future research directions for deploying DL-driven optimization in real 6G networks. Through this survey, we aim to consolidate fragmented research and provide a roadmap for
22 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about leveraging DL in 6G signal optimization. Scope and Contribution The scope of this paper is limited to deep learning–driven signal optimization for 6G wireless systems at the physical (PHY) layer. We focus on optimization problems such as beamforming and precoding, power allocation, reconfigurable intelligent surface (RIS) phase design, waveform adaptation, and joint multi-objective optimization under 6G constraints [5], [6], [12], [22]. Unlike broader surveys that cover higher-layer protocol design or general AI applications in 6G [1], [32], we emphasize PHY-layer signal optimization, where DL can directly replace or augment traditional mathematical solvers. Our main contributions are as follows: Comprehensive review of DL paradigms for signal optimization We categorize DL approaches into supervised, unsupervised/self-supervised, deep unfolding, reinforcement learning, and knowledge-driven hybrids [2], [3], [15], [23]. For each, we describe architectures, optimization formulations, and constraint-handling strategies. Application-focused synthesis We survey DL-driven methods for key PHY tasks, including massive MIMO beamforming [17], [18], uplink/downlink power allocation [21], [22], RIS/STAR-RIS phase optimization [6], [12], [27], and OFDM signal recovery [11], [14]. Performance evaluation framework We review metrics (spectral efficiency, energy efficiency, latency, and robustness), dataset generation strategies, and simulation benchmarks. We also compare DL-based schemes against classical optimization algorithms such as WMMSE, alternating optimization, and semidefinite relaxation [3], [26]. Discussion of open challenges We highlight unresolved issues including generalization across scenarios, constraint satisfaction, robustness to imperfect CSI, interpretability, scalability, and integration into standardized 6G architectures [7], [29], [34]. Future research roadmap We propose directions such as federated and online learning for real-time adaptation [39], hybrid model-data-driven methods [2], [15], and AI-native air interfaces [29]. Through this structure, the paper provides both a state-of-the-art survey and a forward-looking perspective on how DL can enable practical and efficient PHY-layer optimization in 6G systems. LITERATURE REVIEW Research on the use of deep learning (DL) in wireless signal optimization has accelerated in recent years, bridging traditional communication theory with data-driven methods. This section reviews relevant studies, grouped by application area and methodology.
23 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Deep Learning for PHY-layer Optimization Initial work demonstrated that neural networks could approximate PHY-layer operations such as channel estimation, detection, and beamforming with near-optimal performance [13], [14], [16]. He et al. [15] and Qin et al. [16] highlighted the advantages of model-driven DL, which embeds domain knowledge into neural architectures. These approaches outperform purely data-driven black-box models in terms of interpretability and training efficiency. Jagannath et al. [3] introduced deep unfolding, showing that iterative algorithms like weighted minimum mean square error (WMMSE) can be unrolled into neural networks, enabling fast inference while retaining algorithmic structure. Similar concepts have been applied to beamforming, power allocation, and signal detection [15], [37]. Beamforming and Precoding Beamforming and precoding have received significant attention. Huttunen et al. [17] proposed DeepTx, a CNN-based approach for downlink beamforming, which demonstrated robustness under channel aging. Chen et al. [18] applied DL for hybrid beamforming in mmWave massive MIMO, while Ma et al. [19] used CNN and LSTM to reduce beam training overhead in mobile scenarios. Vahapoglu et al. [21] introduced NNBF, an unsupervised neural beamforming scheme that directly maximizes sum rate during training, avoiding dependence on costly optimization labels. For more general adaptability, Nguyen et al. [20] applied deep reinforcement learning (DRL) to mmWave beamforming, showing its ability to adapt to time-varying environments. Collectively, these works show that DL-based beamforming can outperform conventional methods such as zero-forcing (ZF) and minimum mean square error (MMSE) while enabling faster real-time decisions. Power Control and Resource Allocation Power control has also benefited from DL approaches. Song et al. [22] proposed an energyefficient transmission framework for STAR-RIS-assisted cell-free massive MIMO, using DL to jointly optimize transmission power and RIS configurations. Mismar et al. [23] formulated joint beamforming, power control, and interference coordination as a DRL problem, achieving performance close to optimal with far lower computational overhead. Khan et al. [24] surveyed ML and DL-based resource allocation strategies, highlighting scalability and generalization challenges. Camana et al. [25] combined DL with beamforming optimization in SWIPT-RSMA systems, predicting rate components and designing beam formers with improved efficiency. Reconfigurable Intelligent Surfaces (RIS) RIS optimization is a particularly challenging task due to unit-modulus phase constraints. Huang et al. [27] applied DRL to jointly optimize RIS phase shifts and BS beamforming, while Zhong et al. [28] proposed AI-assisted RIS for non-orthogonal multiple access (NOMA), jointly predicting RIS phases and power allocation. Zhou et al. [5] and Alexandropoulos et al. [6] provided comprehensive surveys, comparing model-based, heuristic, and ML-driven optimization approaches.
24 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Recent work has extended DL methods to STAR-RIS, which can both reflect and transmit signals. Megahed et al. [12] applied DL to optimize STAR-RIS in 6G, demonstrating improvements in both throughput and energy efficiency. Joint Optimization and End-to-End Design End-to-end optimization, where DL handles multiple coupled PHY tasks jointly, has been explored in several studies. O’Shea and Hoydis [13] introduced the auto encoder paradigm for communication systems, inspiring end-to-end neural designs. Hoydis et al. [29] advocated for AI-native air interfaces where DL architectures jointly optimize modulation, coding, and beamforming. Sun et al. [2] advanced knowledge-driven DL, embedding communication-theoretic constraints into neural architectures for more generalizable and interpretable optimization. Elsayed and Erol-Kantarci [31] discussed AI-enabled wireless networks, noting hybrid model-data-driven approaches as key to practical 6G deployments. Challenges in Existing Literature Despite these advances, several limitations remain. Most studies assume perfect channel state information (CSI), whereas real-world systems suffer from estimation errors and feedback delays [30], [33]. Constraint handling, such as enforcing power budgets or RIS phase modulus, is often addressed by heuristic projection layers without guarantees [27], [33]. Scalability is another concern: many DL models are trained for fixed antenna or user configurations and generalize poorly to new scenarios [24], [29]. Moreover, interpretability and robustness are open research problems, especially in safety-critical applications such as vehicular or medical communications [34], [40]. Finally, the literature suffers from a lack of standardized datasets and benchmarks, making it difficult to compare approaches fairly [7], [38]. Federated learning [39] and continual learning are emerging directions to improve adaptability and reduce reliance on centralized data. SYSTEM MODELS AND PROBLEM FORMULATIONS Signal optimization in 6G systems is typically defined within the context of multi-antenna, multi-user wireless networks operating across sub-6 GHz, millimeter-wave, and terahertz bands [7], [32]. This section introduces the generic channel and system models, followed by common optimization objectives and their reformulations for deep learning (DL). Generic 6G Channel and Signal Models Consider a downlink multi-user MIMO system with a base station (BS) equipped with 𝑁𝑡 antennas serving 𝐾 users. The received signal is expressed as 𝒚 = 𝑯𝒙+𝒏 where 𝑯𝝐ℂ𝑲∗𝑵𝒕 is the channel matrix, 𝑿 = 𝑾𝒔 is the transmitted signal vector, 𝑾𝝐ℂ𝑵𝟏∗𝑲 is the precoding matrix, 𝒔𝝐ℂ𝑲∗𝟏 is the data symbol vector with normalized power, and 𝒏~𝑪𝑵(𝟎,𝝈𝟐𝑰) is additive Gaussian noise. In reconfigurable intelligent surface (RIS)-assisted systems, the signal model is extended to incorporate the reflection matrix Θ: 𝒚 = 𝑯𝒅𝒙+𝑯𝒓𝚯𝑮𝑿+𝒏,
25 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Where 𝑯𝒅 is the direct BS–user channel, 𝑮 is the BS–RIS channel, 𝑯𝒓 is the RIS–user channel, and 𝚯 = 𝒅𝒊𝒂𝒈(𝒆𝒋𝜽𝟏),…,𝒆𝒋𝜽𝑴 represents the RIS phase shifts [5], [6], [27]. Extensions of this model to STAR-RIS (simultaneous transmit and reflect) have been studied for 6G [12]. At higher frequencies (mmWave, THz), hybrid analog-digital architectures are often employed, where only a limited number of RF chains exist compared to antenna elements, adding further constraints to beamforming optimization [18], [30]. Optimization Objectives Key optimization goals in 6G PHY include: Sum rate maximization: 𝒎𝒂𝒙 𝑾,𝜣 ∑𝒍𝒐𝒈(𝟏+ |𝒉𝑯 𝒌𝒘𝒌|𝟐 ∑|𝒉𝑯 𝒌𝒘𝒋|𝟐+𝝈𝟐 𝒋≠𝒌 ) 𝑲 𝒌=𝟏 , Subject to transmit power constraints [3], [22] Energy efficiency maximization: 𝒎𝒂𝒙 𝑹𝒔𝒖𝒎 𝑷𝒕𝒙 +𝑷𝒄, Where 𝑷𝒄 captures circuit power consumption, relevant in STAR-RIS and cell-free MIMO networks [12], [22] Fairness (max–min rate): Ensuring that the lowest-rate user achieves acceptable performance, critical in dense IoT or URLLC scenarios [24], [34] Transmit power minimization: Minimize ∥ 𝑾 ∥𝑭 𝟐 subject to QoS rate constraints, widely studied for green communication objectives [26]. Error performance minimization: Directly optimize for metrics like BER or SER, particularly in auto encoder-style end-to-end learning frameworks [13], [16]. Multi-objective tradeoffs: Jointly balance rate, energy, latency, and fairness, which is increasingly relevant in 6G heterogeneous deployments [9], [40]. Reformulation for Deep Learning Traditional optimization of these objectives involves convex relaxation, alternating optimization, or semidefinite programming. However, these methods are computationally intensive, particularly in large-scale RIS or massive MIMO scenarios [4], [33]. DL-based
26 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about approaches reformulate the problems as follows: Supervised learning: Use near-optimal solutions from solvers (e.g., WMMSE, SDR) as labels to train neural networks for direct mapping from CSI to pre coders, power vectors, or RIS phases [17], [18]. Unsupervised/self-supervised learning: Define the loss as the negative of the objective (e.g.,−𝑅𝑠𝑢𝑚), enabling networks to optimize directly without ground-truth labels [21]. Deep unfolding: Map iterations of algorithms such as WMMSE into neural layers with learnable step sizes and parameters, accelerating convergence [3], [15]. Reinforcement learning: Model resource optimization as a Markov decision process (MDP), where actions (beam selection, phase configuration) maximize long-term rewards such as throughput or energy efficiency [20], [23], [27]. Knowledge-driven learning: Embed communication constraints into network structures (e.g., normalization layers for power, unit-modulus projection layers for RIS), yielding more interpretable and generalizable models [2], [29]. These reformulations allow DL models to approximate or even outperform traditional solvers under real-time constraints, especially in scenarios with imperfect CSI, non-linear hardware effects, or massive numbers of antennas and RIS elements [6], [28], [31]. DEEP LEARNING ARCHITECTURES AND PARADIGMS Deep learning (DL) models for 6G signal optimization can be categorized into several broad paradigms depending on how they are trained, how they incorporate domain knowledge, and how they address optimization constraints. This section outlines the main approaches and their relevance to physical-layer tasks. Supervised Learning Supervised DL methods train neural networks on datasets generated by traditional solvers. Given channel state information (CSI) or user positions as inputs, the network learns to approximate outputs such as beamforming vectors, power allocations, or RIS phase shifts. Feedforward deep neural networks (DNNs): Multi-layer perceptron’s (MLPs) have been widely used to regress precoding vectors from CSI [17], [18]. Power constraints are often enforced by normalization layers. Convolutional neural networks (CNNs): CNNs exploit spatial correlation in channel matrices, treating them as images. Huttunen et al. [17] used CNNs to predict downlink pre
27 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about coders under channel aging, while Ma et al. [19] applied CNN/LSTM models to mmWave beam training. Graph neural networks (GNNs): In multi-cell or RIS-enabled systems, GNNs capture interactions between BSs, RIS elements, and users, supporting scalable optimization across topologies [24]. Limitations: Supervised methods depend heavily on labeled datasets from computationally expensive solvers such as WMMSE or SDR [3], [26]. Their generalization across scenarios (e.g., varying antenna/user counts) is limited [29]. Unsupervised and Self-Supervised Learning Unsupervised DL avoids reliance on solver-generated labels. Instead, the loss function directly encodes communication objectives. Sum-rate maximization: Vahapoglu et al. [21] designed NNBF, where the loss is negative sum rate, enabling direct optimization of uplink beamforming without labels. Error probability minimization: Auto encoder-based communication systems [13], [16] train end-to-end without labels by minimizing symbol error rates. Advantage: More scalable than supervised learning in large-scale networks. Challenge: Requires careful loss function design to ensure convergence and constraint satisfaction. Deep Unfolding (Model-Driven DL) Deep unfolding (also called unrolling) bridges iterative algorithms with DL. Each layer corresponds to one iteration of an algorithm, but with trainable parameters such as step sizes or thresholds. Jagannath et al. [3] showed how WMMSE for sum-rate maximization can be unfolded into a neural network. He et al. [15] emphasized that model-driven DL improves interpretability and reduces parameter counts compared to black-box models. Applications include beamforming, power control, and MIMO detection [37]. Limitation: Performance depends on the number of layers (iterations) unfolded; too few layers reduce accuracy, while too many increase complexity. Reinforcement Learning (RL) and Deep RL (DRL) RL methods treat optimization as a sequential decision-making process. Markov decision process (MDP): States include CSI, user demand, and interference; actions include beamforming vectors, RIS phases, or power levels; rewards are functions such as sum rate or energy efficiency [20], [23]. Applications:
28 http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about RIS phase optimization [27] Joint beamforming and interference management [23] Dynamic mmWave beam selection [20] Advantages: Adaptation to time-varying and partially observable environments. Limitations: Training requires large numbers of episodes; convergence may be slow in highly dynamic networks. Knowledge-Driven and Hybrid Learning Knowledge-driven DL integrates wireless communication theory into learning architectures. Knowledge-assisted: Classical optimization outputs are used as auxiliary features for training networks [2]. Knowledge-fused: DL provides coarse solutions refined by optimization solvers in a hybrid loop [12], [25]. Knowledge-embedded: Constraints (e.g., power normalization, unit-modulus RIS phases) are implemented as deterministic layers within the network [2], [15]. Sun et al. [2] argued that such paradigms improve generalization and interpretability, while Hoydis et al. [29] highlighted their role in AI-native air interfaces. Comparative Insights Supervised vs. unsupervised: Supervised methods achieve high accuracy under known conditions but lack robustness; unsupervised methods scale better but may converge more slowly Deep unfolding: Provides interpretability and faster convergence with fewer parameters. RL/DRL: Well-suited for dynamic or partially observed environments but require careful design of reward structures. Hybrid knowledge-driven learning: Likely to dominate 6G deployments due to balance between domain knowledge and DL adaptability [2], [31], [40]. APPLICATION AREAS: DL-BASED SIGNAL OPTIMIZATION FOR 6G Deep learning (DL) has been applied to a range of physical-layer optimization problems in 6G. This section reviews key application domains, highlighting models, methods, and insights from the literature. Beamforming and Precoding Beamforming optimization is a cornerstone of massive MIMO and mmWave/THz systems. Supervised approaches: CNNs and fully connected DNNs have been trained to map CSI to precoding vectors. Huttunen et al. [17] developed DeepTx, which outputs downlink precoders robust to channel aging. Chen et al. [18] used supervised DL to design hybrid analog–digital beamformers in mmWave systems.
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 34 Surveys [5], [10], [33] highlight that DL-based RIS optimization often outperforms heuristic or alternating optimization baselines, especially as RIS element counts scale into the thousands. OFDM and Multicarrier Systems Orthogonal frequency-division multiplexing (OFDM) remains essential for wideband 6G. Ye et al. [14] used DL for channel estimation and detection in OFDM, showing BER improvements under non-linear distortion. Kumar et al. [11] proposed a fused deep convolutional network for OFDM signal recovery, reducing error rates compared to conventional equalizers. Auto encoder-based frameworks jointly optimize modulation and detection as a learned system [13]. DL-based OFDM optimization can handle non-linear hardware impairments and outperform least-squares or MMSE estimators under realistic noise models [16]. Joint Optimization and End-to-End Design DL enables joint or end-to-end optimization across multiple layers of the communication stack. Beamforming + RIS + power control: Joint networks can output beamforming vectors, RIS phases, and power allocations simultaneously, optimizing the overall sum rate or energy efficiency [22], [25]. Auto encoder frameworks: O’Shea and Hoydis [13] pioneered end-to-end learned communications, later extended to multi-user and MIMO systems. Knowledge-driven designs: Sun et al. [2] and Hoydis et al. [29] argue that hybrid learning embedding PHY constraints and structure into networks will be central to AI-native 6G interfaces. These methods move beyond optimizing isolated components, targeting global system-level objectives such as latency–throughput trade-offs and fairness [9], [40]. PERFORMANCE EVALUATION AND BENCHMARKING Evaluating deep learning (DL)-based signal optimization methods for 6G requires rigorous metrics, realistic datasets, and strong baselines. This section reviews the methodologies used in the literature. Key Performance Metrics Spectral efficiency (SE) / sum rate: The most common metric, measuring throughput in bits/s/Hz [3], [17]. Energy efficiency (EE): Bits per Joule, especially relevant for RISand STAR-RIS-assisted systems [12], [22].
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 35 Bit error rate (BER) / symbol error rate (SER): Captures error resilience in auto encoderstyle frameworks and OFDM recovery [11], [14]. Fairness: Often measured via max–min user rate or 5th percentile user throughput, critical for dense IoT and URLLC use cases [24], [34]. Inference latency / complexity: Measured in runtime (ms) or FLOPs, important for real-time 6G applications [7], [29]. Robustness to CSI error: Evaluated by adding channel estimation noise or feedback delay [30], [33]. Constraint satisfaction rate: Percentage of outputs that respect power or unit-modulus constraints, critical in RIS settings [27]. Dataset Generation DL methods require diverse and representative training data: Synthetic CSI generation: Rayleigh, Rician, and 3GPP channel models are widely used [15], [18]. Solver-generated labels: Supervised training often relies on WMMSE, semidefinite relaxation (SDR), or alternating optimization solutions [3], [26]. Unsupervised training: Avoids labels by embedding the optimization objective into the loss function [21]. Reinforcement learning (RL): Generates data through simulated interaction episodes, where agents learn from accumulated rewards [20], [23], [27]. Few public benchmarks exist, leading to reproducibility concerns [7], [38]. Calls for standardized datasets are growing in the 6G research community. Baselines for Comparison Classical optimization methods: WMMSE for beamforming/power allocation [3] Alternating optimization for RIS [5], [6] Water-filling for power allocation [26] Heuristic approaches: Random beamforming, zero-forcing (ZF), maximum ratio transmission (MRT), or random RIS phases [19], [28]. Other DL architectures: Comparing supervised vs. unsupervised vs. unfolding vs. RL within the same task [15], [21], [23]. Ablation studies: Removing components such as attention blocks, projection layers, or knowledge-driven modules to quantify their contribution [2], [29]. Simulation Setup Considerations
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 36 Rigorous simulation setups are critical for fair benchmarking: Channel models: Use of standardized models (3GPP, QuaDRiGa) for mmWave/THz ensures realism [18], [30]. System configurations: Results should span different numbers of antennas, RIS elements, and user densities [12], [28]. Hardware impairments: Non-linear amplifiers, quantization noise, and phase noise are increasingly included in 6G evaluations [14], [16]. Statistical evaluation: Average and percentile results (e.g., 5th percentile rate) are needed to capture both average and worst-case performance [24], [34]. Reported Performance Insights Beamforming: DL-based pre coders achieve near-optimal sum rate compared to WMMSE but with orders-of-magnitude faster inference [17], [18]. Power control: DRL-based power allocation adapts better to non-stationary environments than supervised approaches [23]. RIS optimization: DL and DRL approaches outperform alternating optimization when the RIS element count grows large, with STAR-RIS showing notable EE gains [12], [27]. OFDM recovery: CNN/auto encoder methods outperform MMSE equalization under channel distortion and non-linearity [11], [14]. Joint designs: End-to-end learning reduces overall latency by integrating beamforming, power allocation, and RIS optimization into a single inference step [2], [29]. CHALLENGES AND FUTURE DIRECTIONS Despite the strong progress in applying deep learning (DL) to signal optimization, several open challenges must be addressed before widespread deployment in 6G networks. Generalization across Scenarios Most DL models are trained for fixed system settings such as antenna numbers, user densities, or channel statistics. When deployed in mismatched conditions, performance often degrades sharply [24], [29]. Future work should explore domain adaptation, meta-learning, and transfer learning to enable cross-scenario generalization. Robustness to Imperfect CSI Many studies assume perfect channel state information (CSI), but in practice, CSI is corrupted by estimation error, quantization, or feedback delay [30], [33]. Robust DL models must incorporate uncertainty modeling and adversarial training to remain reliable under CSI imperfections. Hybrid model-driven approaches that embed error statistics into network architectures show promise [2], [15].
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 37 Constraint Satisfaction and Interpretability DL-based methods often require ad hoc projection layers to enforce constraints such as unitmodulus RIS phases or transmit power budgets [27], [33]. This may not guarantee strict feasibility. Moreover, the black-box nature of many DL architectures reduces interpretability, which is critical in safety-critical applications such as vehicular or medical communications [34]. Future research should integrate explainable AI (XAI) and optimization-theoretic guarantees into PHY learning systems. Scalability in Large-Scale Networks As antenna counts, RIS elements, and user densities scale up in 6G, the dimensionality of optimization problems increases dramatically. While DL can accelerate inference, the training cost and memory requirements also scale [7], [12] Lightweight models such as graph neural networks (GNNs) and sparse DL architectures may improve scalability [24]. Data Availability and Standardized Benchmarks A lack of publicly available datasets hinders fair comparison of DL-based methods. Most studies rely on proprietary or synthetic CSI generation, preventing reproducibility [7], [38]. Establishing benchmark datasets and open-source platforms for 6G signal optimization would enable systematic evaluation. Online, Distributed, and Federated Learning Centralized training may be infeasible in 6G, where distributed base stations and edge devices continuously generate local data. Federated learning has emerged as a way to train models collaboratively without sharing raw data [39]. Future research should extend federated and online learning to PHY optimization tasks, enabling continual adaptation to dynamic environments. Hardware Constraints and Deployment Feasibility Practical deployment faces challenges from hardware impairments such as low-resolution ADCs, amplifier non-linearity’s, and phase noise [14], [16]. DL-based solutions must be codesigned with hardware-aware models to ensure robustness. Furthermore, inference latency must remain below 1 ms for ultra-reliable low-latency communication (URLLC) [32], [40]. Toward AI-Native 6G Architectures The ultimate vision is an AI-native air interface, where DL is embedded into the design of modulation, coding, beamforming, and resource allocation [29]. This requires integrating DL with 6G standards, ensuring compatibility, reliability, and security [31], [40]. Future systems may use hybrid model-data-driven designs that exploit both mathematical structure and datadriven adaptability [2].
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 38 HARDWARE AND IMPLEMENTATION CONSIDERATIONS While deep learning (DL) offers significant gains for signal optimization in 6G, practical deployment depends heavily on hardware capabilities and system integration. This section highlights key considerations regarding computational platforms, inference acceleration, and implementation trade-offs. Centralized vs. Distributed Inference In 6G networks, DL models may be deployed in centralized cloud data centers, at the base station (BS), or on user equipment (UE). Centralized inference allows access to powerful GPUs/TPUs, enabling large-scale models and batch optimization [7], [29]. However, latency constraints and fronthaul overhead limit its applicability in ultra-reliable low-latency communication (URLLC) scenarios [32]. Distributed or edge inference, where models are deployed at BSs, RIS controllers, or even UEs, reduces latency and bandwidth consumption but requires lightweight, energy-efficient architectures [14], [16]. Hardware Accelerators Specialized hardware accelerators play a crucial role in meeting real-time constraints: GPUs and TPUs: Effective for training and inference in cloud or edge servers, though energy consumption remains high [7]. Field-programmable gate arrays (FPGAs): Offer customizable low-latency inference pipelines, suitable for PHY-layer DL models embedded in BSs [15]. Application-specific integrated circuits (ASICs): Provide energy-efficient deployment for repetitive tasks such as beamforming or RIS phase prediction [29]. Emerging edge AI chips designed for low-power inference are expected to support deployment in mobile devices and RIS controllers [16]. Memory and Latency Constraints Large DL models often exceed the memory capacity of BS and UE hardware. This motivates research into: Model compression: Pruning, quantization, and knowledge distillation reduce parameter size without large accuracy losses [34]. Low-complexity architectures: Graph neural networks (GNNs) and deep unfolding reduce memory and latency compared to black-box DNNs [24], [37]. Pipeline parallelism: Splitting computation across multiple processors can lower latency for wideband and multi-antenna optimization [7]. Meeting the sub-millisecond latency requirement for URLLC remains a critical challenge [32]. Integration with Reconfigurable Intelligent Surfaces (RIS) RIS and STAR-RIS hardware impose additional constraints. Phase adjustments must be executed within microseconds to track fast channel variations [6], [12]. Embedding DL inference directly into RIS controllers could support this, but requires ultra-lightweight,
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 39 hardware-friendly networks. Mixed-signal designs, combining analog RIS hardware with digital DL controllers, are being explored [28]. Online and Federated Training Feasibility Training large DL models in real time is impractical at the BS or UE. However, federated learning enables collaborative training across distributed devices without sharing raw CSI, mitigating privacy risks [39]. Edge devices can update local models, while aggregation occurs in the cloud. Still, communication overhead and hardware heterogeneity remain barriers [7]. Implementation Trade-Offs A key trade-off exists between accuracy and deploy ability. Large networks may achieve higher spectral or energy efficiency but cannot meet the strict latency and energy budgets of 6G hardware. Lightweight model-driven DL approaches offer a practical balance, providing both interpretability and hardware efficiency [2], [15], [29]. SECURITY AND PRIVACY IN DL-BASED 6G OPTIMIZATION The integration of deep learning (DL) into 6G signal optimization introduces new security and privacy challenges. Since DL models control critical physical-layer parameters such as beamforming vectors, power allocations, and RIS phase shifts, adversarial manipulation or data leakage could compromise network reliability, user confidentiality, and even safety-critical applications. Adversarial Attacks on DL Models DL models are vulnerable to adversarial examples, where small perturbations in input data (e.g., channel state information) cause significant degradation in performance [34]. An attacker could exploit this to force misaligned beamforming, increase interference, or reduce throughput. Reinforcement learning (RL)-based methods are particularly susceptible during training, as malicious agents may manipulate the reward function [23]. Defense strategies include adversarial training, robust optimization, and model certification techniques. Model Poisoning and Data Integrity In federated or distributed learning setups, adversaries may attempt model poisoning by uploading corrupted updates to the global model [39]. This could bias beamforming or resource allocation toward inefficient or malicious outcomes. Detecting poisoned updates and applying secure aggregation protocols are essential for maintaining integrity. Blockchain-based authentication mechanisms have been proposed to enhance trust in distributed training [31]. Privacy Concerns in CSI and User Data Channel state information (CSI) inherently contains spatial and mobility patterns of users. If exposed, this could reveal sensitive user location data [7]. Privacy-preserving techniques such as differential privacy, secure multiparty computation, and homomorphic encryption can be
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 40 used in CSI collection and DL training [39]. However, these methods often introduce latency and computational overhead, requiring careful trade-offs in URLLC scenarios [32]. Robustness to Eavesdropping and Jamming DL-driven signal optimization introduces novel attack surfaces: Eavesdropping: If adversaries can predict DL outputs (e.g., RIS configurations), they may align receivers to intercept communications. Jamming: Attackers may target DL-based beam management by injecting false CSI or interfering during training [6], [27]. Resilient architectures require anti-jamming strategies integrated into DL training and inference pipelines, such as training under worst-case interference assumptions [30]. Secure Deployment and Standardization For large-scale adoption, DL-driven PHY optimization must align with 6G security standards [31], [40]. This includes: Secure model updates in distributed/federated settings. Authentication of DL inference outputs at BSs and RIS controllers. Built-in monitoring systems to detect adversarial anomalies in real time Standardization bodies (e.g., 3GPP, ITU) are beginning to incorporate AI/ML security guidelines into 6G frameworks [31], [40]. CROSS-LAYER INTEGRATION Most existing studies on DL-based signal optimization focus on the physical (PHY) layer, but the true potential of deep learning (DL) in 6G arises when PHY optimization is integrated with medium access control (MAC), transport, and application layers. Cross-layer design can unlock system-level improvements in latency, throughput, energy efficiency, and fairness. PHY–MAC Coordination DL-driven beamforming and power control directly affect scheduling, user association, and interference coordination at the MAC layer [23], [24]. For example, a DL-based beamformer that maximizes sum rate may inadvertently disadvantage edge users unless scheduling policies incorporate fairness constraints [34]. Joint PHY–MAC optimization, where DL models consider both channel states and traffic loads, has been shown to improve spectral efficiency and fairness simultaneously [25]. Cross-Layer Resource Allocation In ultra-dense 6G networks, cross-layer DL frameworks can jointly allocate spectrum, power, and computing resources. Reinforcement learning (RL) agents have been designed to optimize end-to-end throughput by coordinating PHY-layer power control with MAC-layer scheduling [23]. Similarly, RIS-assisted networks require coordination between PHY phase-shift design and MAC-layer user association strategies [6], [12].
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 41 Transport-Layer and QoS Integration Future applications such as XR/VR, telemedicine, and autonomous driving impose stringent QoS requirements, including ultra-low latency and jitter [32]. DL-based PHY optimization must therefore be QoS-aware, ensuring that transport-layer requirements are met. For instance, deep unfolding frameworks for beamforming can integrate latency penalties into the objective, aligning PHY-layer inference with higher-layer delay constraints [15]. Edge Intelligence and Computing Offloading Cross-layer integration also extends to computing. In mobile edge computing (MEC) scenarios, DL models can jointly optimize communication and computation resources, balancing task offloading decisions with beamforming and bandwidth allocation [29]. Federated learning frameworks at the edge require joint PHY optimization (for communication) and learning resource scheduling (for computation) [39]. End-to-End Learning Architectures End-to-end neural communication systems represent the most direct form of cross-layer integration [13], [16], [29]. In these architectures, modulation, coding, beamforming, and even application-level QoS requirements are jointly optimized within a single DL model. While promising, challenges remain in scalability, interpretability, and standardization [7], [40]. STANDARDIZATION AND INDUSTRY PERSPECTIVES The successful adoption of deep learning (DL)-based signal optimization in 6G networks requires alignment with ongoing standardization efforts and industry initiatives. This section reviews progress in standardization bodies and examines industry perspectives on DL deployment in future wireless systems. Standardization in 3GPP and ITU The 3rd Generation Partnership Project (3GPP) has already initiated studies on AI/ML for 5G Advanced (Release 18–19), including model management, data collection, and signaling support [31]. These initiatives lay the groundwork for incorporating DL-based optimization into 6G. The International Telecommunication Union (ITU), through its IMT-2030 framework, has also identified AI-native air interfaces as a key enabler for 6G [40]. Key areas under discussion include: ML model representation and exchange formats. Standardized signaling for CSI and feature feedback between BSs and UEs Guidelines for model training, validation, and security [31] Industry Initiatives Global alliances and industry consortia, such as the Next G Alliance (U.S.), Hexa-X (Europe),
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 42 and 6G Flagship (Finland), are actively exploring AI-native solutions for wireless systems. Companies including Nokia, Ericsson, Huawei, and Samsung have demonstrated prototypes of DL-driven beamforming, RIS control, and intelligent scheduling for beyond-5G testbeds [29], [32]. These prototypes highlight the feasibility of deploying DL inference in real-time base stations, though challenges such as hardware constraints, standard compliance, and security remain. Use Cases Driving Industry Adoption Industry adoption of DL-based optimization is being accelerated by use cases requiring low latency, high reliability, and adaptability, including: Autonomous vehicles and UAVs: Requiring ultra-reliable low-latency communication (URLLC). Extended reality (XR/VR): Demanding high throughput and low jitter [32]. Tactile internet and industrial IoT: Dependent on precise beamforming and interference control [34]. Satellite-terrestrial integration: Where DL assists in beam alignment and power management [40]. Challenges in Standardization Despite strong interest, several barriers remain: Lack of consensus on model training and distribution mechanisms in standard architectures Uncertainty over benchmark datasets for reproducibility and evaluation [7], [38] Concerns about security and robustness of DL models in standardized networks [31] Nevertheless, momentum is building toward formalizing AI/ML frameworks in 6G standards, particularly for PHY-layer tasks where DL has shown consistent performance gains [29]. CASE STUDIES AND SIMULATION RESULTS To demonstrate the potential of deep learning (DL)-driven signal optimization for 6G, this section presents representative case studies. The results are drawn from recent literature and supported with simulation-style datasets (see Figs. 3–7). These case studies highlight performance trends in beamforming, power control, and reconfigurable intelligent surfaces (RIS). A. Case Study 1: DL-Based Beamforming DL methods for beamforming were compared against classical zero-forcing (ZF), MMSE, and weighted MMSE (WMMSE) baselines. Sum Rate Performance: As shown in Fig. 3, DL-based approaches such as supervised, unsupervised, deep unfolding, and deep reinforcement learning (DRL) achieve spectral
http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 10 (2025) Online ISSN Print ISSN . . 3007-3197 3007-3189 http://amresearchreview.com/index.php/Journal/about Page 43 efficiency close to WMMSE while significantly outperforming ZF and MMSE [17], [18], [21], [23]. Inference Latency: Fig. 4 shows latency results, where WMMSE incurs high computational costs (~50 ms), whereas DL models reduce inference to under 5 ms. Rate–Latency Trade-off: The trade-off is visualized in Fig. 5, confirming that deep unfolding achieves a strong balance between sum rate and latency [3], [15]. These results support the claim that DL can achieve near-optimal beamforming performance with orders-of-magnitude lower latency. Case Study 2: Power Control and Resource Allocation Energy efficiency and robustness were compared across methods such as water-filling, WMMSE, supervised DL, unsupervised DL, and DRL. Energy Efficiency: As illustrated in Fig. 6, DL-based approaches provide up to a 15% improvement in bits/Joule compared to classical methods [22], [23]. Robustness to CSI Error: DL frameworks also achieve higher robustness levels under CSI uncertainty (see Fig. 6 inset), with DRL performing best. This aligns with prior studies showing that unsupervised and reinforcement learning generalize better to non-ideal channels [21], [23]. These findings confirm that DL-based power control not only improves energy efficiency but also enhances robustness in practical conditions. Case Study 3: RIS Optimization RIS and STAR-RIS pose unique challenges due to high-dimensional unit-modulus constraints. DL-based methods were evaluated against alternating optimization (AO) baselines. Spectral Efficiency vs. RIS Size: As shown in Fig. 7, DL consistently outperforms AO, with the performance gap increasing as RIS size grows [6], [12], [27]. Scalability: While AO converges slowly in large RIS configurations (>512 elements), DL models scale more efficiently, producing rapid inference while maintaining constraint satisfaction. These results demonstrate that DL provides a scalable alternative to iterative AO methods, particularly as RIS element counts approach thousands in 6G systems. Discussion The simulation results confirm several key insights: DL-based beamforming achieves near-WMMSE performance at a fraction of the computational cost.