Comparative Performance Analysis of DNA Sequence Encoding Methods for Machine Learning-Based Bacterial Classification
Full text
Comparative Performance Analysis of DNA Sequence Encoding Methods for Machine Learning-Based Bacterial Classification Authors: Diego Santibáñez Oyarce1, Esteban Gómez Terán1, Jorge Vergara-Quezada2, Ana Moya-Beltrán2. Affiliations: 1Escuela de Informática, Facultad de Ingeniería, Universidad Tecnológica Metropolitana, Santiago, Chile. 2Departamento de Informática y Computación, Facultad de Ingeniería, Universidad Tecnológica Metropolitana, Santiago,Chile. Efficient encoding of DNA sequences remains a major challenge for applying machine learning (ML) in genomics. The choice of encoding method significantly impacts model performance, computational cost, and memory efficiency, particularly in taxonomic classification, antimicrobial resistance prediction, and metagenomic analysis. We systematically evaluate multiple DNA encoding strategies using the universal 16S rRNA marker gene for bacterial genus classification in ML. We compare traditional approaches (One-Hot, K-mers) against signal transformation techniques (Fast Fourier Transform, Wavelet) and hybrid combinations, tested with SVM, Random Forest, and XGBoost classifiers. All methods were evaluated using both aligned sequences (AS) from multiple sequence alignment and padded sequences (PS) with N-padding to uniform length. Transform-based encodings demonstrated notable efficiency: Fourier methods achieved 20.1 seconds execution time compared to 48.2 seconds for One-Hot encoding, while using 1.9GB versus 19.1GB memory—representing 2.4x speed improvement and 10× memory reduction. Wavelet transform showed similar efficiency at 21.9 seconds and 3.8GB memory usage. Peak classification accuracy reached 99.6% with hybrid approaches, while efficient methods like AS-K-mers achieved 98.7% accuracy. Across 30 bacterial genera and 5,256 sequences, aligned sequences consistently outperformed padded sequences, with SVM showing superior performance across most encoding strategies. Our results demonstrate that encoding choice is pivotal for scaling ML models in genomics, with implications beyond taxonomic classification. The evaluated methods can be extended to other sequence-based prediction tasks, enabling more efficient and scalable pipelines. This work contributes to standardizing sequence representation strategies, supporting broader ML adoption in computational biology. Keywords: 16S rRNA marker, Wavelet transformation, Fast Fourier Transform, Machine learning Acknowledgement: Departamento de Informática y Computación, UTEM; Escuela de Informática, UTEM; Laboratorio de Investigación Aplicada, Departamento de Informática y Computación, UTEM. This work was supported in part by Project supported by the “Competition for Research Regular Projects”, year 2023, code LPR23-09 and in part by the “Scientific and Technological Equipment Projects Competition, year 2024, code LE24-03”.