Full text
Network Data Based Transfer Learning Failure Prediction Agent Pre-Trained using Digital Twin Mashboob Cheruvakkadu Mohamed Politecnico di Torino, Italy mashboob.cheruv[email protected] Muhammad Umar Masood Politecnico di Torino, Italy [email protected] Imran Chowdhury Dipto Politecnico di Torino, Italy [email protected] Renato Ambrosone Politecnico di Torino, Italy [email protected] Gulmina Malik Politecnico di Torino, Italy [email protected] Stefano Straullu Links Foundation, Torino, Italy [email protected] Sai Kishore Bhyri Nokia, India [email protected] Gabriele Maria Galimberti Nokia, USA [email protected] Jo˜ ao Pedro Nokia, Portugal [email protected] Antonio Napoli Nokia, Germany [email protected] Walid Wakim Nokia, USA walid.w[email protected] Vittorio Curri Politecnico di Torino, Italy [email protected] Abstract—This paper describes the use of Transfer Learning (TL) using experimental data and a Machine Learning (ML) model pre-trained with a Digital Twin (DT) for the prediction of amplifier failures in optical networks. Using GNPy, an opensource framework, amplifier failure conditions are simulated, creating the required training dataset for the ML model. Later, by implementing TL using optical transmission testbed network data, the model is able to capture realtime network fluctuations, thereby enabling it to distinguish network parameters variations due to any incoming failures from the regular network dynamics, thus enhancing the practical applicability of the model. The model is based on the Long Short-Term Memory (LSTM) ML Technique and is shown to achieve a TL accuracy of 99%, demonstrating the ability of the model to effectively predict failures. This method facilitates early identification and intervention, reduces service interruptions, and improves network reliability. Leveraging TL from network data provides a scalable and datadriven solution to enhance the resilience, efficiency, and ongoing operation of contemporary optical communication systems. Index Terms—Machine learning, digital twin, optical amplifiers, fault detection, proactive failure management. I. INTRODUCTION In today’s increasingly data-driven and connectivitydependent environments, it is crucial to maintain network availability in optical communication networks, as they are the backbone of modern communication systems. Although protection and restoration mechanisms are usually utilized in these networks, these strategies are reactive in nature. Hence, to reduce operational expenditures and ensure high-quality service, network operators are interested in completing them with proactive maintenance and failure detection in optical networks. By identifying and addressing potential problems before they lead to service interruptions, operators can maintain availability, reliability, and performance while minimizing downtime. This approach optimizes network resources, reduces repair costs, and improves overall network stability. The application of machine learning for the detection and localization of failures in optical networks has been gaining attention since earlier works [1], [2]. However, practical implementation remains a significant challenge and has not yet been fully realized. The application of ML models in practical use cases is heavily dependent on the coverage of failure scenarios that can be handled by these ML models. The irregular incidence of network failures makes it difficult to collect the necessary data required for the development of reliable ML models. Alternatively, the required training dataset can be obtained by modeling failure scenarios using digital twins [3]. However, the limitation of the synthetic dataset is that it may lack real-time network dynamics. In order to combine the advantages of each of these approaches, a possible solution for implementing the ML models in the practical use case is to pre-train the ML model with a larger training dataset generated using DT and then apply TL using the available real network data. Previous studies [4], [5] emphasize the scarcity of real network datasets to train ML models for various types of failures. These works exploit strategies, such as data augmentation, to compensate for data scarcity and dataset imbalance. The use of DT is a promising alternative for generating synthetic training datasets. In [6], soft failure localization using telemetry twin is demonstrated where failures are localized on streaming telemetry based on predefined thresholds with the ML approach relying on instantaneous values without addressing time-series properties. This results in reduced usability, considering the fact that how much real-time synchronization of the telemetry twin with the real network is possible. Additionally, using this approach, localization can take seconds, whereas in real networks, downtime exceeding 50 milliseconds is considered critical. The idea of prediction of future changes has been proposed in [7] based on historical data in DT. However, the lack of sufficient past data for fault scenarios and the prolonged time to collect all types of failures that happen unevenly in real networks makes this solution less viable today. In this article, a novel concept of applying TL 978-3-903176-67-6 © 2025 IFIP 2025 International Conference on Optical Network Design and Modeling (ONDM)
Optical Field Network Network Data Digital Twin (DT) Transfer Learning Model Trained on Network Data Failure Prediction Machine Learning Model Pre-trained on DT Fig. 1: Schematic representation of failure prediction using TL model utilizing network data and a pre-trained ML model based on DT. using real network data on a pre-trained ML model trained with synthetic dataset from DT is introduced. The advantage of transfer learning is that it requires a smaller amount of data from the target domain, as it leverages the knowledge gained from the source domain [8]. The general concept of applying real network data based TL to the ML model pretrained using DT is represented in Fig. 1. The goal is to use this approach with field network data, but for the purpose of obtaining results for this work, TL is implemented with experimental data obtained from the optical transmission test bed. A ML model based on the time-series data from normal conditions of an optical link is demonstrated in an existing research work [9], where failures are detected and localized based on deviations in linear thresholds calculated from normal conditions data. This article demonstrates the application of Long Short-Term Memory (LSTM) ML technique to capture the time-series characteristics of failure scenarios, using a synthetic failure condition dataset created by simulating the progression of failures with GNPy and then applying TL using the time-series failures dataset generated using the experimental setup of the laboratory optical network. The remainder of this paper is organized as follows. Section II provides an explanation of the generation of a synthetic dataset of amplifier failure conditions using DT. Section III describes the experimental setup and the data collection process. In Section IV, an overview of the ML framework and TL approach is presented. Section V discusses the results and provides an analysis. Finally, Section VI concludes the paper and outlines directions for future work. II. OPTICAL MULTIPLEXED SECTION DIGITAL TWIN USING GNPY FOR PRE-TRAINING AMPLIFIER FAILURE FORECAST The optical amplifiers can exhibit performance degradation due to failure of components such as pump lasers and gain media (Erbium-doped fiber, EDF) due to saturation, aging, etc. This can result in a decline in the gain achieved by the amplifiers, which can affect the overall Optical Signalto-Noise Ratio (OSNR) of the optical link, thereby causing a low Quality-of-Transmission (QoT). The persistence of these issues can lead to critical failures that disrupt services. It is crucial to monitor and identify potential forthcoming failures in amplifiers, as they play a critical role in sustaining the proper OSNR to maintain the desired QoT over the network. In an Optical Line System (OLS) consisting of multiple amplifiers, the slow degradation in the target gain of a given amplifier, if still below the threshold set for alarm reporting, will be left unchecked. Also, this incidental decline of gain in a faulty-to-be amplifier will be compensated by the subsequent amplifiers in the optical link and may go unnoticed due to no changes detected in the power levels at the power monitors at the end of the lightpath. This can result in a hidden fault for a longer period of time. This scenario of amplifier gain degradation can be simulated by failure modeling in GNPy. GNPy is a widely used software tool for designing optical networks by simulating the configurations of the different key components of a network [10]. The network topology consisting of transponders, Re-configurable Optical Add/Drop Multiplexers (ROADMs), amplifiers, and fibers can be defined in GNPy and the tool simulates the physical layer and computes the Generalized Signal-to-Noise Ratio (GSNR) of the optical 2025 International Conference on Optical Network Design and Modeling (ONDM)
TX/RXTransceiver WSSWavelength Selective Switch AmpOptical Amplifier ILAIn-Line Amplifier SMFSingle Mode Fiber OSAOptical Spectrum Analyzer OSNROptical Signal to Noise Ratio ILA 1 ILA 2 ILA 3 ILA 4 ILA 5 ILA 6 ILA 7 Failure generation (Gain degradation) on randomly selected Amps OSNR Monitoring Measurement of baseline values of amplifier gain/OSNR Amplifier failure generation Monitoring of current values of amplifier gain/OSNR due to degradation - Data collection Transfer Learning of the DT pre-trained ML model Fig. 2: Experimental optical transmission setup, amplifier failure generation and data collection process - an illustration. signal [11]. In GNPy a DT of the experimental setup illustrated in Fig. 2 is created consisting of transceivers, ROADMs, optical amplifiers and fiber spans. The span includes DT of single mode fiber of total 520 km, consisting of eight spans of 65 km. The network topology includes DT of commercially deployed amplifiers constituting of seven In-Line Amplifiers (ILAs), pre-amplifier and booster amplifier in each direction. In the above DT, slow gain degradation is created in the amplifier by varying the gain in a range of smaller step sizes and simultaneously compensating the power loss by adjusting the Variable Optical Attenuator (VOA) at the output of the amplifier. This results in no change of the output power resembling the scenario of non-detection of the effect of gain degradation of an individual amplifier by the power monitors at the end of line in the real network, as explained above. Step sizes are randomly selected such that overall degradation is added slowly to replicate field defects. The step sizes are held for a different period of time starting from shorter duration for smaller steps and extending it to increasing values of degradation simulating network dynamics. The time-series dataset consists of changes in GSNR corresponding to the degradation values of individual amplifiers that have different gain and tilt levels. The distribution of each ILA in the dataset varies between 10% to 20% which is significant for ML to classify the faulty amplifier. III. EXPERIMENTAL SETUP AND DATA ACQUISITION This section provides a detailed description of the experimental setup and the data collection process, highlighting the procedures, equipment used, and the methodology followed to collect data from the failure condition of the optical amplifiers. A. Experimental Details In this study, the optical testbed illustrated in Fig. 2 is employed as the physical layer (PHY) system. The amplified optical path comprises 9 commercial Erbium-Doped Fiber Amplifiers (EDFAs), including 7 ILAs, one booster amplifier (BST), and one pre-amplifier (PRE), linked by 8 spans of standard Single-Mode Fiber (SMF), each with a nominal length of 65 km. At the input of the BST, a Wavelength Division Multiplexing (WDM) comb spanning the C-band is generated, consisting of 64 channels spaced 75 GHz apart, each carrying a signal modulated at 64 GBd. The two channels under test (CUTs) are generated using Galileo Phoenix devices. These CUTs are then propagated into the WDM comb through Adtran ROADM. The Phoenix devices support configuration and real-time monitoring through NETCONF interfaces. At the output of the OLS, the two CUTs are extracted by the ROADM and directed to the receiver-side coherent modules. Given the partially disaggregated nature of the optical infrastructure, the Optical Multiplex Section (OMS) controller plays a critical role in the software control framework. The controller interfaces with amplifiers via proprietary (i.e. vendor-specific) protocols, allowing it to monitor input/output power and dynamically adjust gain and tilt settings [12]. B. Data Collection Process The dataset collection in the optical transmission testbed is achieved in two steps: the baseline values of OSNR at 2025 International Conference on Optical Network Design and Modeling (ONDM)
Simulated Dataset Lab Dataset Pre-Trained LSTM Model LSTM Model Failure Forecast/ Alerts TL Model Fig. 3: Process of training the LSTM Model for Transfer Learning. the end of the line, along with various parameters of the optical amplifiers, such as gain, tilt, VOA etc., are recorded by telemetry. Subsequently, failure conditions are introduced in optical amplifiers by gradually degrading the gain in steps of 0.1 dB for randomly selected amplifiers. During failure generation, amplifier parameter values are recorded and OSNR values are monitored through the Optical Spectrum Analyzer (OSA) during each iteration. The VOA in amplifiers are adjusted in parallel to offset power loss, resembling hidden gain degradation failures causing change in noise figures which are not directly detectable by power monitors in the network. Based on the degree of degradation, three fault levels are assigned: no fault is considered below 2 dB, a minor fault between 2 dB and 3 dB, and a major fault above 3 dB. To prevent service interruptions, the amplifier’s maximum fault is restricted within the VOA limit, as this helps predict soft failures. Surpassing this threshold could result in hard failures or loss of the optical signal. IV. TRANSFER LEARNING FRAMEWORK This section outlines the background of the LSTM technique, highlighting its relevance in the scope of this work, and the TL approach employed in this study. A. Long Short-Term Memory (LSTM) ML Technique LSTM networks were first introduced by Hochreiter and Schmidhuber [13]. LSTM is a specialized form of Recurrent Neural Networks (RNNs) created to address the vanishing gradient issue. This problem arises during training when the gradients of the loss function with respect to the model parameters shrink during back-propagation. Such small gradients can make it difficult to update model weights effectively, a challenge that LSTM networks are designed to solve [14], [15]. Using LSTM in machine learning, the accuracy of forecasting is greatly improved, as it effectively integrates recent and historical time-series data [16]. This capability allows the model to capture complex temporal dependencies and patterns over time, making it well-suited for tasks such as predicting future trends or behaviors. LSTM is a widely used and highly effective model to handle input features that exhibit sequential dependencies [14]. By learning from both shortterm and long-term dependencies in sequential data, LSTM networks can model complex patterns that evolve over time (e.g., when order and timing of data points are crucial for making accurate predictions). This makes it a better-suited machine learning approach for predicting amplifier failures in optical networks. B. Transfer Learning Transfer learning enhances the learning process in a target task by leveraging knowledge from a related source task. TL improves learning in three key areas: First, it increases initial performance in the target task by utilizing the knowledge transferred, compared to starting with no prior knowledge. Second, it reduces the time required to fully learn the target task when using transferred knowledge, compared to learning it from scratch. Lastly, it leads to better performance on the target task compared to the performance achieved without any transfer. In summary, TL accelerates and improves the learning of a machine learning model by transferring knowledge from a task that has already been learned [17], [18]. In this work, TL is implemented by training an LSTM in two different stages. The process flow of TL is depicted in Fig. 3. Initially, the model is trained on a source dataset generated from GNPy simulations, as described in Section II. In this stage the model is trained on normal conditions and relationships in the data such as correlation between amplifier failure conditions and the end-of-line GSNR values. These learned features form the foundation of the transfer learning approach. The last stage involves training of the pre-trained LSTM model on a dataset containing the real data from the experimental set up as described in Section III, which contains OSNR values that represent amplifier failures due to signal degradation. To effectively train the model, the dataset includes the differences between the baseline GSNR/OSNR values and the failure condition values along with the duration of the faults. By retaining the knowledge learned by the model in the previous training and adapting it to the target domain, the LSTM model effectively identifies patterns of gain degradation, enabling robust fault detection in optical amplifiers. 2025 International Conference on Optical Network Design and Modeling (ONDM)
V. RESULTS AND DISCUSSION As outlined, this study utilizes LSTM to effectively detect faulty amplifiers. Furthermore, the model categorizes the gradual degradation of gain into different fault levels, as shown in Table I, allowing for a more precise and systematic assessment of amplifier performance over time. This approach identifies the presence of faults and tracks their progression, providing a detailed analysis of the severity of the issues. TABLE I: Amplifier Fault Levels Fault level Gain degradation(dB) Severity 0<2No fault 1 2-3 Minor fault 2>3Major fault The performance of the model is assessed using a set of standard classification metrics, such as accuracy, precision, recall, and F1 score. In classification tasks, the classifier assigns each input to one of two classes, typically labeled ”true” or ”false.” This process generates four key outcomes: true positives, true negatives, false positives, and false negatives. Accuracy measures the proportion of correct predictions made by the classifier, reflecting how often the model correctly classifies both true and false instances. Precision refers to the percentage of truly positive predictions. It highlights how reliable the model is when it predicts a positive class, ensuring that false positives are minimized. Recall, also known as the sensitivity or true positive rate, indicates the percentage of actual positive cases that are correctly identified by the classifier. It emphasizes the model’s ability to detect all relevant positive instances. F1-Score is the harmonic mean of precision and recall, offering a balanced metric that considers both false positives and false negatives [14]. A comparison of Accuracy, Precision, Recall, and F1-Scores for the pre-trained model and the TL model is shown in Table II. The pre-trained model achieved an accuracy of 98% with a dataset size of 14177 data points, whereas the TL model achieved 99% accuracy from a dataset size of 4174 data points showing a 70% reduction in the required data while maintaining improved performance. The train:test data ratio is 80:20. TABLE II: Accuracy, Precision, Recall and F1 score - a comparison of pre-trained and TL model Model Accuracy Precision (weighted average) Recall (weighted average) F1-Score (weighted average) Pre-trained model 98% 0.98 0.98 0.98 TL model 99% 0.99 0.99 0.99 The confusion matrix is constructed by comparing the predicted amplifiers and their corresponding fault levels with the actual amplifiers and their associated fault levels. The confusion matrix for the pre-trained model and the TL model Predicted Actual (a) Confusion Matrix of pre-trained model Predicted Actual (b) Confusion Matrix of TL model Fig. 4: ML confusion matrices - a comparison is shown in Fig. 4a and Fig. 4b respectively. This matrix shows correctly classified instances (true positives and true negatives) and misclassified instances (false positives and false negatives). Each amplifier is represented by a distinct identifier name, ilaxy, where xshows the name of the amplifier and y shows the fault level. As is evident from the matrices, the misclassifications of faulty amplifiers is significantly reduced after TL. For example, the pre-trained model misclassified ila4 0 instead of ila6 0 for 16 instances; all these misclassifications were cleared after TL. These metrics provide a comprehensive evaluation of the model’s ability to correctly classify faulty amplifiers and categorize the varying levels of gain degradation. The training time for the TL model was 497 seconds, whereas the pre-trained model took 1222 seconds on the same 2025 International Conference on Optical Network Design and Modeling (ONDM)
machine. This represents a time reduction of 40%, highlighting the efficiency gained through TL. Taking into account multiple performance measures, the evaluation provides evidence of the effectiveness of the proposed framework that leverages a DT for initial ML model training and TL using field data to further improve the ML model. VI. CONCLUSIONS AND FUTURE WORK This paper presented a novel approach to predicting optical network failures using digital twins and transfer learning. DT can be used to generate a large amount of training dataset to pre-train a ML model. TL helps reduce the computational costs associated with developing machine learning models for new problems. In this experiment, we applied TL to reduce: training time, the amount of field training data, and other computational expenses, leading to a faster and more efficient model training process. Additionally, TL supports model optimization and improves generalizability, as demonstrated by the higher accuracy achieved by the TL model. This improvement occurs because TL involves retraining an existing model using a new dataset, allowing the model to retain knowledge from the previous dataset it was trained on. As a result, the TL model is better equipped to prevent overfitting when trained on the new data. Future work will address a wider range of failures and challenges encountered in optical networks. For example, the study can be extended to explore issues such as filter shifts and filter tightening in ROADMs, which can significantly impact network performance and signal integrity. Additionally, research can explore the effects of nonlinear impairments in optical fibers, such as four-wave mixing, self-phase modulation, and cross-phase modulation, which can degrade signal quality and hinder data transmission efficiency. By investigating these complex phenomena, future work can contribute to the development of more robust and efficient optical network systems, enhancing their reliability and performance under diverse operational conditions. Furthermore, exploring advanced mitigation techniques and optimization strategies for these challenges could pave the way for the design of next-generation optical networks capable of supporting highcapacity, low-latency communication. ACKNOWLEDGMENTS This publication is part of the PNRR-NGEU projects which have received funding from MUR–DM117/2023 and have been supported by the EU projects ALLEGRO, GA No. 101092766 and DN NESTOR GA No. 101119983. REFERENCES [1] Francesco Musumeci, Cristina Rottondi, Giorgio Corani, Shahin Shahkarami, Filippo Cugini, and Massimo Tornatore. A tutorial on machine learning for failure management in optical networks. J. Lightwave Technol., 37(16):4125–4139, Aug 2019. [2] Josh W. Nevin, Sam Nallaperuma, Nikita A. Shevchenko, Xiang Li, Md. Saifuddin Faruk, and Seb J. Savory. Machine learning for optical fiber communication systems: An introduction and overview. APL Photonics, 6(12):121101, 12 2021. [3] Mashboob Cheruvakkadu Mohamed, Muhammad Umar Masood, Renato Ambrosone, Gulmina Malik, Rocco D’Ingillo, Stefano Straullu, Sai Kishore Bhyri, Gabriele Maria Galimberti, Joao Pedro, Antonio Napoli, Walid Wakim, and Vittorio Curri. Machine learning agents leveraging digital twins for failure prediction in optical networks. In 2025 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN) (2025 ICMLCN), page 5.63, Barcelona, Spain, May 2025. [4] Lareb Zar Khan, Jo ao Pedro, Nelson Costa, Andrea Sgambelluri, Antonio Napoli, and Nicola Sambo. Model and data-centric machine learning algorithms to address data scarcity for failure identification. J. Opt. Commun. Netw., 16(3):369–381, Mar 2024. [5] Cheng Xing, Chunyu Zhang, Bing Ye, Danshi Wang, Yinqiu Jia, Jin Li, and Min Zhang. Failure data augmentation for optical network equipment using time-series generative adversarial networks. In Optical Fiber Communication Conference (OFC) 2023, page M3G.4. Optica Publishing Group, 2023. [6] Kayol S. Mayer, Rossano P. Pinto, Jonathan A. Soares, Dalton S. Arantes, Christian E. Rothenberg, Vinicius Cavalcante, Leonardo L. Santos, Filipe D. Moraes, and Darli A. A. Mello. Demonstration of ml-assisted soft-failure localization based on network digital twins. J. Lightwave Technol., 40(14):4514–4520, Jul 2022. [7] Qunbi Zhuge, Xiaomin Liu, Yihao Zhang, Meng Cai, Yichen Liu, Qizhi Qiu, Xueying Zhong, Jiaping Wu, Ruoxuan Gao, Lilin Yi, and Weisheng Hu. Building a digital twin for intelligent optical networks. J. Opt. Commun. Netw., 15(8):C242–C262, Aug 2023. [8] Francesco Musumeci, Virajit G. Venkata, Yusuke Hirota, Yoshinari Awaji, Sugang Xu, Masaki Shiraiwa, Biswanath Mukherjee, and Massimo Tornatore. Transfer learning across different lightpaths for failurecause identification in optical networks. In 2020 European Conference on Optical Communications (ECOC), pages 1–4, 2020. [9] Mois´ es Felipe Silva, Alessandro Pacini, Andrea Sgambelluri, and Luca Valcarenghi. Learning longand short-term temporal patterns for mldriven fault management in optical communication networks. IEEE Transactions on Network and Service Management, 19(3):2195–2206, 2022. [10] Vittorio Curri. Gnpy model of the physical layer for open and disaggregated optical networking [invited]. Journal of Optical Communications and Networking, 14(6):C92–C104, 2022. [11] Alessio Ferrari, Mark Filer, Karthikeyan Balasubramanian, Yawei Yin, Esther Le Rouzic, Jan Kundrat, Gert Grammel, Gabriele Galimberti, and Vittorio Curri. Gnpy: an open source application for physical layer aware open optical networks. Journal of Optical Communications and Networking, 12(6):C31–C40, 2020. [12] Renato Ambrosone, Giacomo Borraccini, Andrea D’Amico, Stefano Straullu, Francesco Aquilino, Dirk Breuer, Rainer Schatzmayr, Gert Grammel, and Vittorio Curri. Open line controller architecture in partially disaggregated optical networks. In 2023 International Conference on Photonics in Switching and Computing (PSC), pages 1–3. IEEE, 2023. [13] Sepp Hochreiter and J¨ urgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 11 1997. [14] Lars E. Kruse, Sebastian K¨ uhl, Annika Dochhan, and Stephan Pachnicke. Experimental validation of machine learning-based joint failure management and quality of transmission estimation. IEEE Photonics Journal, 15(6):1–9, 2023. [15] Qing Wang, Rong-Qun Peng, Jia-Qiang Wang, Zhi Li, and Han-Bing Qu. Newlstm: An optimized long short-term memory language model for sequence prediction. IEEE Access, 8:65395–65401, 2020. [16] Stefanos Giaremis, Noujoud Nader, Clint Dawson, Hartmut Kaiser, Carola Kaiser, and Efstratios Nikidis. Storm surge modeling in the ai era: Using lstm-based machine learning for enhancing forecasting accuracy, 2024. [17] Lisa Torrey and Jude Shavlik. Transfer learning. In Handbook of research on machine learning applications and trends: algorithms, methods, and techniques, pages 242–264. IGI global, 2010. [18] Emilio Soria Olivas, Jos David Mart Guerrero, Marcelino MartinezSober, Jose Rafael Magdalena-Benedito, L Serrano, et al. Handbook of research on machine learning applications and trends: Algorithms, methods, and techniques: Algorithms, methods, and techniques. IGI global, 2009. 2025 International Conference on Optical Network Design and Modeling (ONDM)