scieee AI-readable full text Open interactive document viewer

BMD-SRA: A Boosting Model for Differentiating Sequence Read Archive sequences

Bole, Martin

Abstract

The number of sequence files deposited in the Sequence Read Archive (NCBI-SRA) has been growing exponentially through the years, and with it, the number of incorrectly annotated types of sequences. The submitted sequences are then used for genomic, metagenomic, and taxonomic studies. This presents a need in the research community for a model that facilitates the collection of correctly annotated data. This study aimed to develop a boosting classification model called BMDSRA that classifies input sequences into four sequence types: 1)Metagenomes, 2)Amplicons, 3)Single-Amplified Genomes (SAGs), 4)Isolated-Genomes. For developing the Machine Learning (ML) algorithm, we gathered 3000 test samples for each sequence type respectively. Test samples were used for supervised ML. Metagenomes were collected from various metagenome databases (DBs) (Kasmanas et al., Nucleic Acids Research, 2020) (750 samples from each), manually curated, and created by our team. Amplicon samples were gathered from the Joint Genome Institute portal based on their library strategy. The SAG samples were collected by manually inspecting published research papers, proving they were sequenced from a single cell. The Isolated-Genomes were gathered from SRA, searching for bacteria-type strain Genomes from different taxonomies. The BDMSRA reads a small portion of the sequence file using a sub-sampling approach (SRA Toolkit Development Team, https://trace.ncbi.nlm.nih.gov/Traces/sra/sra.cgi?view=software) and extracts statistical features generated based on Shannon entropy, Tsallis entropy, and Fourier z-curve. The extracted features were evaluated using the QPFS method (Soheili et al., Scientific Programming, 2020), and the reliability of training data was tested with an outlier analysis. From the 119 generated features, we chose 38 with the highest importance for developing the model. The outlier analysis showed that the SAG and Amplicon data sets were the most reliable, with few outliers. The outliers from Metagenomes and Isolated-Genomes were subjected to further manual investigation. The model was created and evaluated by using 5-fold cross-validation. The confusion matrix showed an overall accuracy of 92% (96% for SAGs, 95% for Amplicons, 92% for Metagenomes, and 85% for Isolated-Genomes). The false negatives from Isolated-Genomes classified as Metagenome (7.6%) and SAGs (5.9 %) are likely due to the wrong classification in the SRA. The false negatives from Metagenomes classified as Isolated-Genomes (6.7%) are potentially due to downloading process from our Dbs. BMDSRA can help researchers verify that the sequences they submit or collect from public repositories are correctly annotated. Further, our tool could also select samples for metastudies and determine if sequence projects are well performed.

Full text

BMD-SRA: A Boosting Model for Differentiating Sequence Read Archive sequences Ulisses Nunes da Rocha1*, Martin Bole1,2, Jonas Coelho Kasmanas1,2,3, Majid Soheili1 1 Department of Environmental Microbiology, Helmholtz Centre for Environmental Research – UFZ GmbH, Leipzig, Saxony, Germany, 2 Department of Computer Science and Interdisciplinary Centre of Bioinformatics, University of Leipzig, Leipzig, Saxony, Germany 3 Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, Brazil Contacts: ulisses.roc[email protected],martin.bol[email protected],[email protected],[email protected], @ulisses_rocha REFERENCES  The Sequence Read Archive (NCBI-SRA) has seen exponential growth in the number of deposited sequence files over the years.  Incorrectly annotated types of sequences have also increased due to the expanding volume of submissions.  The research community requires a model to aid in collecting accurately annotated data to support genomic, metagenomic, and taxonomic studies. This study aimed to develop a boosting classification model called BMDSRA that classifies input sequences into four sequence types: 1) Metagenomes, 2) Amplicons, 3) Single-Amplified Genomes (SAGs), 4) Isolated-Genomes. Importance Feature Name Figure 1: The importance of features measured with QPFS Figure 2: Outliers of the training data based on their category Figure 4: Confusion Matrix of BMDSRA after 5-fold cross validation Type Source Distribution Meta-Genomes Published Datasets Marine (750), Human (750), Animal (750), and Terrestrial (750) Amplicons JGI Soil (1216), Marine (547), Water (116), Human (322), 16s (742), 18s (484), ITS (291), V1 (94), V2 (11), V3 (76), V4 (63), V5 (12), V6 (12) SAGs Published Papers 3000 various bacterial species Isolated-Genomes NCBI-SRA 3000 various bacterial species Table 1: Structure of the training dataset STEP 1 – Random subsampling and downloading sequences from SRA (Fastqdump), extracting Shannon & Tsalis entropy (MathFeature), Z-curve Fourier transformation (MathFeature), Feature importance measurement (QPFS) STEP 2 – Systematic detection of outliers to increase training set reliability (DBScan) STEP 3 – Feature Forward Addition Curve (XGBoost, 3-fold crossvalidation) to reduce overfitting, selection of 38 optimal features Number of Selected Features Error Rate Figure 3: Feature Forward Addition Curve for the error rate of the classification model STEP 4 – Final XGBoost model trained using 5-fold crossvalidation RESULT – BMDSRA has 92 % overall accuracy in classifying SRA sequences True Label Predicted Label (a) METAGENOMIC outliers (c) SAG outliers (b) AMPLICON outliers (d) ISOLATED-GENOME outliers  Selected 38 most important features out of 119 generated.  Model developed and evaluated using 5-fold cross-validation.  Confusion matrix showed 92% overall accuracy.  False negatives in Isolated-Genomes: 7.6% as Metagenome, 5.9% as SAGs (SRA misclassification).  False negatives in Metagenomes: 6.7% as Isolated-Genomes (download issues from DBs). Alneberg, J., Karlsson, C. M. G., Divne, A.-M., Bergin, C., Homa, F., Lindh, M. V., Hugerth, L. W., Ettema, T. J. G., Bertilsson, S., Andersson, A. F., and Pinhassi, J. (2018). Genomes from uncultivated prokaryotes: a comparison of metagenome-assembled and single-amplified genomes. Microbiome, 6(1):173. For Biotechnology Information (U.S.), N. C. SRA hand-book. National Center for Biotechnology Information, Bethesda. Hosokawa, M., Endoh, T., Kamata, K., Arikawa, K., Nishikawa, Y., Kogawa, M., Saeki, T., Yoda, T., and Takeyama, H. (2022). Strain-level profiling of viable microbial community by selective single-cell genome sequencing. Sci Rep, 12(1):4443. Soheili, M., Moghadam, A.-M. E., and Dehghan, M. (2020). Statistical analysis of the performance of rank fusion methods applied to a homogeneous ensemble feature ranking. Scientific Programming, 2020:8860044. Torres, P. J., Edwards, R. A., and McNair, K. A. (2017). PARTIE: a partition engine to separate metagenomic and amplicon projects in the sequence read archive. Bioinformatics, 33(15):2389--2391.  BMDSRA helps verify correct annotation of submitted or collected sequences.  Improve and refine the model for better accuracy and performance.  Address SRA misclassifications and downloading issues from databases.  Continuously evaluate and validate the model with diverse datasets. DISCUSSION NEXT STEPS AND FUTURE PROSPECTS METHODOLOGY AND RESULTS BACKGROUND AND AIMS