Landscape of Machine Learning Methods and Data Representations for Antimicrobial Resistance: Toward a Benchmarking Framework in HPC Environments
Full text
Landscape of Machine Learning Methods and Data Representations for Antimicrobial Resistance: Toward a Benchmarking Framework in HPC Environments Camilo Cerda Sarabia1 , Esteban Gómez Terán1 , Fernanda Bravo Cornejo1 , Belén Díaz Díaz1 , Fausto Cabezas-Mera2, Raúl Caulier-Cisterna3 , Jorge Vergara-Quezada3, Ana Moya-Beltrán3 . 1Escuela de Informática, Facultad de Ingeniería, Universidad Tecnológica Metropolitana, Santiago, Chile. 2Programa de Doctorado en Informática Aplicada a Salud y Medio Ambiente, Escuela de Postgrado, Universidad Tecnológica Metropolitana, Santiago, Chile 3Departamento de Informática y Computación, Facultad de Ingeniería, Universidad Tecnológica Metropolitana, Santiago, Chile . To analyze and compare the performance of various coding strategies and ML models for predicting antimicrobial resistance, identifying optimal, reproducible, and scalable configurations in HPC environments. Models available in the literature Conclusions References Antimicrobial resistance (AMR) is one of the most urgent threats to global health, demanding computationally robust and scalable solutions. In recent years, machine learning (ML) has emerged as a powerful strategy for analyzing large-scale genomic data to predict resistance profiles and uncover genetic patterns linked to resistance mechanisms. However, existing tools vary in terms of input features, encoding strategies, model architectures, and execution environments. AMR prediction studies exhibit considerable heterogeneity: data sets vary in species, sample size, and data type. Computational environments differ across laboratories; feature extraction and encoding methods are inconsistent; and model architectures range from convolutional neural networks (CNNs) to ensemble methods. This diversity limits reproducibility, complicates performance comparisons, and increases the time and resources required to reliably evaluate tools. MLP And CNN Acknowledgments: Laboratorio de Investigación Aplicada, Departamento de Informática y Computación, UTEM; Escuela de Informática, UTEM;. This work was supported in part by Project supported by the “Competition for Research Regular Projects”, year 2023, code LPR23-09 and “Competition for Research Assistant Funding UTEM”, year 2023, code AI23-06, Universidad Tecnológica Metropolitana (AM-B) Contact: [email protected] [email protected] Conclusions Nguyen, M., Olson, R., Shukla, M., VanOeffelen, M., & Davis, J. J. (2020). Predicting antimicrobial resistance using conserved genes. PLOS Computational Biology, 16(8), e1008319. https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008319 Wang, S.-C. (2024). E-CLEAP: An ensemble learning model for efficient and accurate identification of antimicrobial peptides. PLOS ONE, 19(3), e0300125. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0300125 Ren, Y., Chakraborty, T., Doijad, S., Falgenhauer, L., Falgenhauer, J., Goesmann, A., Hauschild, A.-C., Schwengers, O., & Heider, D. (2022). Prediction of antimicrobial resistance based on whole-genome sequencing and machine learning. Bioinformatics, 38(2), 325–334. https://academic.oup.com/bioinformatics/article/38/2/325/6382301 The CNN.model.h5 model was evaluated with over 2,400 Escherichia coli genomes paired with resistance phenotypes, and the E-CLEAP model was evaluated with 3,500 peptide samples. These models were tested in Ubuntu and Windows environments, respectively, with one-hot encoding and PseAAC (Pseudo Amino Acid Composition), with a runtime of 2,589.87 seconds and a size of 87 MB for CNN and 100 seconds and 0.84 MB for MLP. Goal A state-of-the-art analysis was carried out using three bibliographic repositories, Scopus, WOS, and PubMed. Method RF-LR-SVM-training E-CLEAP TrainXGBoost CNN.model.h5 Models evaluated Models Optimized SVM And XGBoost The Logistic Regression and SVM models were optimized using strategic Grid Search, evaluating 16 and 12 combinations of hyperparameters, respectively. Logistic Regression implemented L1/L2/ElasticNet regularization with StandardScaler, improving ROC-AUC from 85% to 95.42%. SVM used linear and RBF kernels with RobustScaler, increasing ROC-AUC from 92.4% to 96.53%. Both models achieved >95% discriminatory power with robust cross-validation. Random Forest was optimized using strategic Grid Search evaluated eight combinations of key hyperparameters: n_estimators (150-300 trees versus the original fixed 200), max_depth (limited to 15 levels for regularization, versus no limit to capture complex patterns), and class_weight (equal treatment versus automatic balancing by class frequency), achieving significant improvements with ROC-AUC of 97.14% and especially MCC of 86.74%, indicating a better balance between sensitivity and specificity critical for clinical diagnosis. The RF-LR-SVM model was evaluated with its dataset of over 2,400 Escherichia coli genomes paired with resistance phenotypes, and the TrainXGBoost model was evaluated with its dataset of Staphylococcus aureus genomic data (1,274 genomes paired with seven antibiotic resistance profiles), each with different environments and encoding methods. RF LR SVM MLP (Multilayer Perceptron) DOI:10.1093/bioinfo rmatics/btab681 . MLP AUC = 97.33% Resistant phenotypes with F1 = 98% for methicillin. Convolutional Neural Network (CNN) CNN AUC = 93% DOI:10.1093/bioi nformatics/btab6 81 XGBoost study Random Forest AUC = 96% DOI:10.1371/journ al.pone.0300125 DOI:10.1371/jour nal.pcbi.1008319 The second model mentioned had execution time metrics of 1.42 hours and a weight of 186.36 MB, differing from the first with 25 MB and 11 seconds. This aspect shows the heterogeneity of the data present in antimicrobial resistance. With the aim of seeing what is currently available in terms of antimicrobial resistance and machine learning models. Modifications were made to the Python codes due to the removal of incompatibilities libraries (Python 3 to Python 2.7). Hyperparameters optimization was performed on three methods o (RF,LR,SVM), training model. The models were downloaded and evaluated with their respective environment, datasets and encoding . Hyperparameter optimization using Grid Search achieved improvements in the methodological models, although critical trade-offs between accuracy and computational efficiency persist, with some models requiring 60 times more resources while maintaining similar performance. Incompatibilities between Python versions and obsolete libraries underscore the urgent need for standardized computational frameworks that facilitate knowledge transfer between laboratories and support the effective clinical implementation of AMR predictive tools. By systematically evaluating different data types, encoding methods, and feature sets across diverse computational environments, our study highlights that heterogeneity remains a major barrier to reproducibility in ML-based AMR prediction. Identifying optimal encoding-model combinations on unified datasets provides a foundation for reliable, scalable, and reproducible AMR prediction pipelines, supporting equitable access to computational tools in the fight against antimicrobial resistance. Models Optimized Models evaluated Virulence factors and antibiotic resistance of Streptococcus pyogenes. Methods Performance Reference Problem