scieee AI-readable full text Open interactive document viewer

Dataset for: High-performance machine learning for peptide classification from nanopore translocation events, leveraging event kinetics and duration filtering

Krantz, Bryan

Abstract

Abstract Understanding single-molecule translocation dynamics through biological nanopores is fundamental to advancing next-generation biosensing and sequencing technologies. Here, using the anthrax toxin protective antigen nanopore, we describe a high-performance machine learning (ML) framework for classifying a diverse series of guest-host peptides based on individual translocation events. The approach leverages carefully engineered, event-level biophysical features extracted from either scaled current and conductance state sequences. Through systematic UMAP analysis of this feature space, we reveal that filtering away the shortest events effectively enriches the dataset with more discriminative longer events, leading to improved classification. Various deep learning (DL) and traditional ML architectures, including convolutional neural networks (CNN), temporal convolutional networks (TCN), and eXtreme Gradient Boosting (XGBoost), were investigated. The dual-input CNN-Dense model, which utilized current sequences and features, achieved strong classification performance (accuracy ~0.80). However, the most robust classification was achieved with XGBoost acting solely on the engineered feature set, demonstrating superior performance (accuracy ~0.90). This ML approach provided a significant computational advantage in both training and inference over DL models. Notably, these models consistently discriminated between peptides differing only in backbone stereochemistry, highlighting the exquisite sensitivity of the nanopore to subtle conformational dynamics. These findings underscore that carefully engineered event-level features, particularly from longer translocations, combined with efficient tree-based models, offer a highly effective and computationally favorable strategy for high-fidelity peptide classification for biosensing applications.

Full text

High-performance peptide classification with a nanopore biosensor enabled by filtering for longer translocation events Jennifer M. Colby1 and Bryan A. Krantz2* 1Present address: Premier Biotech Labs, 723 Kasota Ave SE, Minneapolis, MN 55414, U.S.A. 2Department of Microbial Pathogenesis, School of Dentistry, University of Maryland, Baltimore, 650 W. Baltimore Street, Baltimore, MD 21201, U.S.A. * Corresponding author Bryan A. Krantz Department of Microbial Pathogenesis School of Dentistry, University of Maryland, Baltimore 650 W. Baltimore Street Baltimore, MD 21201, U.S.A (410) 706-1656 [email protected] Running title: Machine learning for peptide biosensing Preprint server: deposited at bioRxiv under doi: https://doi.org/10.1101/2025.08.21.671556 with a CC-BY 4.0 International license. Classification: Major: Biological Sciences, Minor: Biophysics and Computational Biology Keywords: peptide biosensor, machine learning, deep learning, nanopore, anthrax toxin, protective antigen, electrophysiology, translocation Abstract Understanding single-molecule translocation dynamics through biological nanopores is fundamental to advancing next-generation biosensing and sequencing technologies. Here, using the anthrax toxin protective antigen nanopore, we describe a high-performance machine learning (ML) framework for classifying a diverse series of guest-host peptides based on individual translocation events. The approach leverages carefully engineered, event-level biophysical features extracted from either scaled current and conductance state sequences. Through systematic Uniform Manifold Approximation and Projection (UMAP) analysis of this feature space, we reveal that filtering away the shortest events effectively enriches the dataset with more discriminative longer events, leading to improved classification. Various deep learning (DL) and traditional ML architectures, including convolutional neural networks (CNN) and eXtreme Gradient Boosting (XGBoost), were investigated. The dual-input CNN-Dense model, which utilized current sequences and features, achieved strong classification performance (accuracy ~0.80). However, the most robust classification was achieved with XGBoost acting solely on the engineered feature set, demonstrating superior performance (accuracy ~0.90). This ML approach provided a significant computational advantage in both training and inference over DL models. Notably, these models consistently discriminated between peptides differing only in backbone stereochemistry, highlighting the exquisite sensitivity of the nanopore to subtle conformational dynamics. These findings underscore that carefully engineered event-level features, particularly from longer translocations, combined with efficient tree-based models, offer a highly effective and computationally favorable strategy for high-fidelity peptide classification for biosensing applications. Significance Statement Nanopore biosensors ensure detailed single-molecule peptide analysis, but deciphering their complex electrical signals is a major bottleneck to their application. We developed a machine learning framework to classify diverse peptides using the anthrax toxin nanopore. Classification fidelity is shown to be critically dependent on translocation event duration; filtering out short translocations enhances for information-rich events and dramatically improves accuracy across all models. Surprisingly, a computationally efficient tree-based model, which uses engineered biophysical features, achieved ~91% accuracy, outperforming more complex deep learning architectures. This work establishes that leveraging event duration and expert-derived features provides a robust and computationally tractable strategy for developing high-performance nanopore biosensors. Introduction The rapid and accurate detection and characterization of biomolecules, particularly peptides and proteins, represent a critical and often unmet challenge across diagnostics, drug discovery, and fundamental biological research. Peptide biomarkers, indicative of various disease states including heart disease, infectious diseases, and cancer (1-3), offer immense potential for timely and precise diagnosis. However, analyzing complex peptide mixtures in situ, often at low concentrations, demands analytical technologies with unparalleled sensitivity and specificity—capabilities that remain challenging for many conventional methods. Singlemolecule nanopore biosensing provides a transformative approach (4), enabling the detection and characterization of individual molecules as they traverse a nanometer-scale pore by measuring picoamp-scale modulations in ionic current. Beyond simple presence/absence detection, nanopore technology holds the promise to revolutionize biopolymer analysis, including the ambitious goal of direct, high-throughput peptide and protein sequencing, a major frontier where robust, widely applicable methods are still lacking (5-7). Nanopore biosensors consist of a membrane-embedded pore separating two electrolytefilled compartments. Under an applied driving force (either a voltage or proton gradient), biomolecules are directed through the pore, generating unique current signatures dependent on their size, shape, and chemical properties. Biological protein nanopores, such as those formed by transmembrane proteins inserted into lipid bilayers, offer exquisite control over pore geometry and molecular interactions (8-14). Single-channel recordings capture the dynamic, stochastic interactions of individual molecules with the nanopore in high-resolution ionic current traces. Extracting the maximum analytical information from these complex, high-dimensional translocation event streams is the central bottleneck preventing the full realization of nanopore technology's potential for analyzing complex biological samples. Among the most promising biological nanopores for polypeptide analysis is the anthrax toxin protective antigen (PA) channel from Bacillus anthracis (15) (Fig. 1A). As a natural protein translocase, PA possesses unique structural/functional features, which are highly advantageous for peptide biosensing and sequencing. It is remarkably robust, facilitates highly processive polypeptide translocation driven by an applied voltage (16) or proton gradient (17) with distinct kinetic substeps, and utilizes multiple internal 'peptide-clamp' sites (9, 12-14, 18, 19) to engage heterologous polypeptides (20) without the need for cumbersome tags like DNA handles (18, 19, 21, 22). Importantly, polypeptide translocation through the PA pore generates rich, multistate current signatures (18), often involving multiple discrete partially blocked sub-conductance intermediates in addition to the fully blocked and open states. These distinct, peptide-dependent kinetic and conductance characteristics are information-rich readouts that can serve as powerful discriminatory features for peptide identification and classification in complex mixtures. Furthermore, the fine structure and temporal dynamics of these multi-state transitions encode information potentially sufficient to resolve amino acid sequence, presenting a unique opportunity for developing label-free, peptide sequencing capabilities. Analyzing the massive, complex, and often noisy datasets generated by multi-state nanopore systems requires advanced computational approaches. Traditional manual or simple threshold-based analysis methods are fundamentally inadequate for extracting the full information content from nuanced multi-state kinetics or dissecting complex mixtures of analytes. Machine learning (ML) and deep learning (DL) offer powerful, modular, and adaptable tools uniquely suited to identify subtle patterns in translocation state sequences and correlate them with computed biophysical features (23-27). While ML/DL has been applied to nanopore data, its application to analyzing real-world, multi-state peptide translocation data, either for precise peptide classification in mixed samples or for determination of sequence information, remains largely an underdeveloped area. Here, leveraging anthrax toxin’s PA nanopore, we describe a high-performance ML framework for classifying a diverse series of guest-host peptides based on individual translocation events. Results Diverse guest-host peptide translocation events via PA nanopores. Previous investigations extensively characterized the broad ensemble properties and single-channel dynamics of the guest-host peptide series translocating through the anthrax toxin protective antigen (PA) nanopore (18, 21). Building upon this foundation, we sought to explore whether the intrinsic dynamic/kinetic properties of these peptides, as observed during single-channel translocation events, could serve as information-rich signatures for classification. Specifically, we aimed to assess if the PA nanopore, when coupled with powerful computational ML/DL methods, could reliably distinguish peptides based on subtle sequence differences that are challenging to discern by conventional analysis. This study thus assesses the PA nanopore's potential as a sophisticated peptide biosensor. Relative to other protein nanopores commonly utilized in biosensing and nucleic acid sequencing, the PA nanopore (Fig. 1A) possesses a significantly longer lumen. This architecture, however, features numerous active site clamps (e.g., α-clamp, ϕ-clamp, and charge clamp) and loop regions (e.g., 397-loop), specifically evolved to processively translocate large (~100 kDa) proteins, and, at nanomolar concentrations, shorter peptides. The 10-residue guest-host peptide series, with a general sequence of KKKKKXXSXX, was initially designed to systematically probe differences in binding, translocation, and dynamics based on the guest residue (X) (Fig. 1B). Our guest-host panel included peptides with standard natural ʟstereochemistry for guest residues Ala, Leu, Phe, Thr, Trp, and Tyr. To investigate the sensitivity of the nanopore to stereochemical variations, a seventh peptide, called guest-host TrpDL, was included; it shared the same amino acid sequence as guest-host Trp but featured an alternating pattern of ᴅand ʟ-stereoisomers along its backbone. For this assessment of the PA nanopore as a biosensor platform, we focused on collecting extensive single-channel translocation event streams via planar lipid bilayer electrophysiology. The data acquisition and processing workflow is shown in Fig. S1A. All analyses presented herein utilized data acquired under a consistent 70 mV driving force (cis positive). This potential strongly favors complete translocation events, which is critical for consistent feature extraction as it ensures the peptide fully interacts with the entire nanopore rather than more superficially engaging the entrance, as might occur at lower potentials. Furthermore, the signal-to-noise ratio is inherently higher at larger potentials, providing additional support for this selection. A comprehensive breakdown of the total recording times per peptide in our complete dataset is provided in Table S1. Across samples of translocation event streams for the seven guest-host peptide classes, four discrete conductance states are consistently observed, which are enumerated as states 0-3 (Fig. 1C). However, the translocation event dynamics, as characterized by current blockade patterns and durations, cover a broad range. Qualitatively, events for guest-host Trp and guesthost TrpDL exhibit noticeably longer durations compared to the other five peptides. In contrast, peptides with smaller guest residues, such as guest-host Ala and guest-host Thr, display rapid dynamics that are often difficult to distinguish reliably by visual inspection alone. For peptides like guest-host Leu and guest-host Phe, which are similar on hydrophobicity scales, subtle visual differences in their translocation dynamics exist but are similarly challenging to resolve. The aromatic guest residues Phe and Tyr present event lengths and flickering dynamics that are visually analogous and difficult to discriminate. Therefore, these qualitative observations, particularly the subtle or complex nature of visual distinctions between peptide classes, directly suggest that advanced ML/DL methods can be effectively exploited to robustly classify these peptides, even from individual single-channel translocation events. Discriminatory potential of engineered features is event length dependent. Achieving robust peptide classification from individual translocation events necessitated a comprehensive and generalized feature engineering approach. To ensure high-fidelity state assignments critical for feature extraction, raw current records were state-labeled using the 'Single-Channel Search' routine in CLAMPFIT, providing expert-validated assignments. For our specific four-state event streams (18), a consistent enumeration scheme was used: state 0 for fully blocked, states 1 and 2 for intermediate blockades (closest to state 0 and state 3, respectively), and state 3 for the fully open state (Fig. 1C). From these labeled records, both conductance state sequences and scaled current sequences were segmented for each translocation event, from which a rich feature set (including scalar, vector, and matrix features) was calculated (see Materials and Methods in SI Appendix) (Fig. S1B). While the segmentation process offers filtering and baseline correction capabilities, these were not employed for the current datasets. However, we critically explored the impact of minimum event duration as a preprocessing filter. This filter initially aimed to remove spurious, very rapid blockade dynamics, which are sometimes attributed to water structure formation ('wetting' and 'dewetting') around the ϕ clamp (28). Beyond this initial rationale, it became evident through preliminary analyses on simulated data that shorter events inherently possessed less discriminatory information, analogous to attempting to classify images with significantly fewer pixels. This observation motivated a systematic investigation into the effect of minimum event duration on feature discriminability. The main practical downside to using a minimum event duration filter is that translocation event data are exponentially distributed, and higher filtering values for this parameter could remove large numbers of translocation events from the dataset (Table S2). To qualitatively and quantitatively assess the impact of different minimum event duration thresholds on the discriminative power of the extracted feature sets, Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction was employed to generate 2D cluster representations. For events filtered at a minimum duration of 5 ms, UMAP analysis (Fig. 2A) resulted in poor clustering performance. Visually, peptide classes remained largely intermixed, with few well-structured, distinct clusters and numerous 'stray' points dispersed throughout the 2D projection. In contrast, increasing the minimum event-length filter to 20 ms visibly improved clustering in the UMAP analysis (Fig. 2B). While some intermixing persisted and a smaller fraction of stray points remained, distinct clusters became apparent for several peptides, notably Phe, Thr, Trp, and Tyr. Interestingly, despite its strong classification performance in subsequent ML/DL models (as shown later), the events for guest-host TrpDL in the 20 ms UMAP embedding presented as several smaller, somewhat dispersed clusters, suggesting inherent sub-populations or more complex relationships not fully captured by the 2D projection. To quantify these visual observations, clustering metrics Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI) were computed by applying K-Means clustering (K=7, for seven peptide classes) to the UMAP embeddings and comparing the resulting clusters to the known peptide labels. For data filtered at 5 ms, the ARI was 0.0356 and NMI was 0.0402. However, for data filtered at 20 ms, the ARI increased to 0.0887 and the NMI to 0.1282. The substantially higher ARI and NMI values for the 20 ms filtered data strongly indicate a greater alignment between UMAP-derived clusters and true peptide identities at longer event durations. These results collectively demonstrate that the engineered feature sets gain significant discriminative power by excluding very short events, albeit with the inherent trade-off of reducing the total number of analyzed events (Tabel S2). Performance of DL classification models. The core design approach for DL-based peptide classification from individual translocation events involved a branched, dual-input neural network architecture (Fig. 3A, S1C) (29). In this configuration, either the conductance state sequences (S) or the raw current sequences (C) from translocation events served as input to a multi-layered CNN or TCN branch. Simultaneously, the corresponding event-level features (F), extracted during preprocessing, were fed into a separate fully connected (Dense) network. The outputs of these two parallel branches were then concatenated before leading to a final classification layer. Based on this design pattern, three distinct configurations were assessed: TCN-Dense (S+F), CNN-Dense (S+F), and CNN-Dense (C+F). Given that the discriminative power of the engineered features was shown to be dependent on the minimum event duration parameter during preprocessing (Fig. 2), these models were subjected to an initial, singlereplicate performance scan across a range of minimum event duration values (5, 7.5, 10, 12.5, Acknowledgments. We thank members of the department as well as Tobin Sosnick for useful feedback and discussions. J.M.C. and B.A.K. conceived of the experiments. J.M.C. collected the data. B.A.K performed the analysis. B.A.K. and J.M.C. wrote the manuscript. Portions of this document, including some of the Python code and language refinement, were generated with the assistance of AI-powered tools. All content was reviewed and approved by the authors, who take full responsibility for its accuracy. References 1. V. Castiglione et al., Biomarkers for the diagnosis and management of heart failure. Heart Fail Rev 27, 625–643 (2022). 2. P. Povoa et al., How to use biomarkers of infection or sepsis at the bedside: guide to clinicians. Intensive Care Med 49, 142–153 (2023). 3. R. Mastrantonio, H. You, L. Tamagnone, Semaphorins as emerging clinical biomarkers and therapeutic targets in cancer. Theranostics 11, 3262–3277 (2021). 4. L. Ratinho, N. Meyer, S. Greive, B. Cressiot, J. Pelta, Nanopore sensing of protein and peptide conformation for point-of-care applications. Nat Commun 16, 3211 (2025). 5. B. Lin, J. Hui, H. Mao, Nanopore Technology and Its Applications in Gene Sequencing. Biosensors (Basel) 11 (2021). 6. Y. Goto, R. Akahori, I. Yanagi, Challenges of Single-Molecule DNA Sequencing with Solid-State Nanopores. Adv Exp Med Biol 1129, 131–142 (2019). 7. X. Wei et al., Engineering Biological Nanopore Approaches toward Protein Sequencing. ACS Nano 17, 16369–16395 (2023). 8. J. Jiang, B. L. Pentelute, R. J. Collier, Z. H. Zhou, Atomic structure of anthrax protective antigen pore elucidates toxin translocation. Nature 521, 545–549 (2015). 9. N. J. Hardenbrook et al., Atomic structures of anthrax toxin protective antigen channels bound to partially unfolded lethal and edema factors. Nat Commun 11, 840 (2020). 10. K. Zhou et al., Atomic Structures of Anthrax Prechannel Bound with Full-Length Lethal and Edema Factors. Structure 28, 879–887 e873 (2020). 11. A. J. Machen, M. T. Fisher, B. D. Freudenthal, Anthrax toxin translocation complex reveals insight into the lethal factor unfolding and refolding mechanism. Sci Rep 11, 13038 (2021). 12. G. K. Feld et al., Structural basis for the unfolding of anthrax lethal factor by protective antigen oligomers. Nature Struct. Mol. Biol. 17, 1383–1390 (2010). 13. S. L. Wynia-Smith, M. J. Brown, G. Chirichella, G. Kemalyan, B. A. Krantz, Electrostatic ratchet in the protective antigen channel promotes anthrax toxin translocation. J Biol Chem 287, 43753–43764 (2012). 14. B. A. Krantz et al., A phenylalanine clamp catalyzes protein translocation through the anthrax toxin pore. Science 309, 777–781 (2005). 15. B. A. Krantz, Anthrax Toxin: Model System for Studying Protein Translocation. J Mol Biol 436, 168521 (2024). 16. S. Zhang, E. Udho, Z. Wu, R. J. Collier, A. Finkelstein, Protein translocation through anthrax toxin channels formed in planar lipid bilayers. Biophys. J. 87, 3842–3849 (2004). 17. B. A. Krantz, A. Finkelstein, R. J. Collier, Protein translocation through the anthrax toxin transmembrane pore is driven by a proton gradient. J. Mol. Biol. 355, 968–979 (2006). 18. K. Ghosal et al., Dynamic Phenylalanine Clamp Interactions Define Single-Channel Polypeptide Translocation through the Anthrax Toxin Protective Antigen Channel. J Mol Biol 429, 900–910 (2017). 19. D. Das, B. A. Krantz, Peptideand proton-driven allosteric clamps catalyze anthrax toxin translocation across membranes. Proc Natl Acad Sci U S A 113, 9611–9616 (2016). 20. S. R. Blanke, J. C. Milne, E. L. Benson, R. J. Collier, Fused polycationic peptide mediates delivery of diphtheria toxin A chain to the cytosol in the presence of anthrax protective antigen. Proc. Natl Acad. Sci. U.S.A. 93, 8437–8442 (1996). 21. J. M. Colby, B. A. Krantz, Peptide Probes Reveal a Hydrophobic Steric Ratchet in the Anthrax Toxin Protective Antigen Translocase. J Mol Biol 427, 3598–3606 (2015). 22. D. Das, B. A. Krantz, Secondary Structure Preferences of the Anthrax Toxin Protective Antigen Translocase. J Mol Biol 429, 753–762 (2017). 23. N. Celik et al., Deep-Channel uses deep neural networks to detect single-molecule events from patch-clamp data. Commun Biol 3, 3 (2020). 24. S. Yang et al., Deep Learning-Based Ion Channel Kinetics Analysis for Automated Patch Clamp Recording. Adv Sci (Weinh) 12, e2404166 (2025). 25. D. Dematties, C. Wen, M. D. Perez, D. Zhou, S. L. Zhang, Deep Learning of Nanopore Sensing Signals Using a Bi-Path Network. ACS Nano 15, 14419–14429 (2021). 26. C. Cao et al., Aerolysin nanopores decode digital information stored in tailored macromolecular analytes. Sci Adv 6 (2020). 27. D. Rodriguez-Larrea, Single-aminoacid discrimination in proteins with homogeneous nanopore sensors and neural networks. Biosens Bioelectron 180, 113108 (2021). 28. G. Yamini et al., Hydrophobic Gating and 1/f Noise of the Anthrax Toxin Channel. J Phys Chem B 125, 5466–5478 (2021). 29. B. A. Krantz, Deep learning-based classification of peptide analytes from single-channel nanopore translocation events. PLoS One 20, e0324777 (2025). 30. T. Chen, C. Guestrin (2016) XGBoost: A Scalable Tree Boosting System. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (ACM), pp 785–794. 31. J. M. Colby, B. A. Krantz, Comparative study of nanopore phenylalanine clamp variants reveals unique peptide biosensing and classification properties. bioRxiv 10.1101/2025.08.22.671566 (2025). 32. E. F. Pettersen et al., UCSF Chimera—a visualization system for exploratory research and analysis. J. Comput. Chem. 25, 1605–1612 (2004). Table 1. Model performance metrics at two minimum event duration extremes.1 Model2 Metric3 Minimum event duration 5 ms 20 ms XGBoost (F) Accuracy 0.5419 (±0.0066) 0.9112 (±0.0069) XGBoost (S+F) 0.5334 (±0.005) 0.9047 (±0.0064) CNN-Dense (C+F) 0.5511 (±0.0141) 0.7857 (±0.0116) CNN-Dense (S+F) 0.4201 (±0.0069) 0.7155 (±0.0095) TCN-Dense (S+F) 0.3821 (±0.0144) 0.6287 (±0.0157) XGBoost (F) F1-score 0.5371 (±0.0056) 0.9109 (±0.0068) XGBoost (S+F) 0.5255 (±0.0069) 0.9044 (±0.0063) CNN-Dense (C+F) 0.5517 (±0.0146) 0.7845 (±0.0108) CNN-Dense (S+F) 0.4265 (±0.0071) 0.7112 (±0.0071) TCN-Dense (S+F) 0.3948 (±0.0157) 0.6296 (±0.0166) 1Performance based on test set evaluation metrics of different ML/DL classification models at different minimum event durations. Values are means and std. dev. (N=5). 2Models are named as defined in the text. 3Metric names are abbreviated and refer to overall accuracy and macro-averaged F1-score. Figures Fig. 1. PA nanopore peptide biosensor. (A) Sagittal section of the anthrax toxin PA7 nanopore cryo-EM structure (PDB: 3J9C) (8) rendered in Chimera (32) as a molecular surface. Peptide clamps and loop active sites are colored and labeled: α clamp (magenta), 397-loop (cyan), ϕ clamp (green), and charge clamp (red). Overall scale of the upper vestibule and elongated lower β barrel are indicated. Narrowest point, at the ϕ clamp, has a luminal diameter of 6 Å. Direction of translocation from cis to trans is indicated by an arrow. Membrane bilayer position is indicated with a solid gray rectangle. (B) Guest-host peptide design schematic for the 10-residue guesthost peptide with general sequence, KKKKKXXSXX, with guest residue (X). (C) Representative current versus time records of guest-host peptide translocations carried out at 70 mV (cis positive) in symmetric succinate buffer, 100 mM KCl, pH 5.6. To the left are guest-host peptide names. In general, this nanopore-peptide system populates multiple discrete conductance state intermediates (indicated by dashed lines), which are enumerated by state on the far right: fully blocked (state 0), partially blocked intermediates (state 1 and state 2), and fully open (state 3). Scalebar at upper right for guest-host Ala, Leu, Phe, Thr, and Tyr peptides represents 2 pA by 100 ms. For guest-host Trp and TrpDL peptides, the scalebar represents 2 pA by 500 ms due to their characteristically longer events. Fig. 2. Discriminative potential of extracted features at different minimum event durations. UMAP clustering analysis of event-level features. Data points represent individual translocation events, colored by their corresponding guest-host peptide identity: Ala (black), Leu (red), Phe (green), Thr (blue), Trp (yellow), TrpDL (magenta), and Tyr (cyan). (A) UMAP embedding of events filtered at a minimum duration of 5 ms (n_neighbors = 15, min_dist = 0.5). Peptide classes are more intermixed with less clear separation, indicating more limited discriminative power of features for very short events (ARI: 0.0356; NMI: 0.0402). (B) UMAP embedding of events filtered at a minimum duration of 20 ms (n_neighbors = 15, min_dist = 0.5). In contrast, events filtered at 20 ms show visibly improved clustering for several guest-host peptides (e.g., Phe, Thr, Trp, Tyr), albeit TrpDL is more spread out and peripheral. While some scattered points remain, suggesting inherent event variability or limitations of the 2D projection, the overall separation is markedly enhanced (ARI: 0.0887; NMI: 0.1282). Fig. 3. DL-based classification of guest-host peptide translocation events. (A) General dual-input neural network architecture as a block diagram illustrating the common design pattern for the DL classifiers. Sequence data (either conductance states (S) or raw current (C)) are processed by a multi-layered CNN or TCN branch. Concurrently, event-level features (F) are processed by a fully connected (Dense) branch. The outputs of these two branches are concatenated and fed into a final Dense layer for classification. This architecture was the basis for the TCN-Dense (S+F), CNN-Dense (S+F), and CNN-Dense (C+F) models. (B) Normalized confusion matrix displaying the per-class prediction accuracy of the CNN-Dense (C+F) model on the test set for a minimum event duration of 5 ms. Rows represent true peptide labels, and columns represent predicted labels. Values indicate the proportion of events from a given true class that were predicted as each class. (C) Normalized confusion matrix for CNN-Dense (C+F) at 20 ms minimum event duration, which is plotted as in panel B. Note the improved diagonal elements (correct predictions) and reduced off-diagonal elements (misclassifications) compared to the shorter minimum event duration predictions in panel B. Fig. 4. High-performance tree-based classification of translocation events. Normalized confusion matrix for XGBoost classification using the event-level feature set. This matrix represents the best-performing XGBoost (F) model at a minimum event duration of 20 ms. Rows represent true peptide labels, and columns represent predicted labels. Values indicate the proportion of events from a given true class that were predicted as each class. Despite the model's overall high performance, a very moderate level of misclassification is observed between the chemically similar guest-host Ala and guest-host Thr classes. 1 Supporting Information for High-performance peptide classification with a nanopore biosensor enabled by filtering for longer translocation events Jennifer M. Colby1 and Bryan A. Krantz2* 1Present address: Premier Biotech Labs, 723 Kasota Ave SE, Minneapolis, MN 55414, U.S.A. 2Department of Microbial Pathogenesis, School of Dentistry, University of Maryland, Baltimore, 650 W. Baltimore Street, Baltimore, MD 21201, U.S.A. * Corresponding author Bryan A. Krantz Department of Microbial Pathogenesis School of Dentistry, University of Maryland, Baltimore 650 W. Baltimore Street Baltimore, MD 21201, U.S.A (410) 706-1656 [email protected] This PDF file includes: Supporting text Figure S1 Tables S1 to S3 SI References 7 features, were loaded from a local database. To ensure balanced class representation, all peptide classes were downsampled to match the class with the minimum number of events (Table S2). The comprehensive dataset was then split into an 80% training set and a 20% testing set for model development and evaluation, respectively. For clarity, throughout this features-only model is referred to as XGBoost (F). The XGBoost classifier was configured with the following key parameters: a multiclass classification objective, the number of target classes set to the total number of peptides, 1000 boosting rounds (trees), and a learning rate of 0.05 to control the step size shrinkage. Regularization was applied with max depth of 5 to limit tree complexity, minimum child weight was 1 to control minimum sum of instance weight (hessian) needed in a child, gamma was set to 0 for minimum loss reduction required to make a further partition on a leaf node, subsample was 0.8 (fraction of samples used per tree), and the fraction of features used per tree was 0.8. L1 and L2 regularization were used. For reproducibility, a random state was fixed, and computation was distributed across all available CPU cores. The model's performance during training was monitored using the multiclass classification error. The trained model's performance was evaluated on the unseen testing dataset. Classification metrics including accuracy, precision, recall, and F1-score were summarized in a standard classification report, and a confusion matrix was generated to visualize per-class prediction accuracy. Hybrid DL/ML-based peptide classification. A hybrid classification approach, XGBoost (S+F), was developed to leverage the representational power of deep learning for sequence data alongside the robust performance of tree-based ensemble methods. This model combined sequence-derived embeddings from a supervised CNN with the previously defined event-level features, which were then merged and fed into the XGBoost classifier. 8 CNN-based sequence embedding generation. To generate sequence embeddings, a dedicated CNN model was trained to perform multiclass peptide classification directly on conductance state sequences. Raw state sequences were first remapped to integer IDs, where original states (e.g., 0, 1, 2, 3) were shifted by one (e.g., 1, 2, 3, 4) to reserve '0' as a dedicated padding token. These remapped sequences were then post padded to a fixed sequence length of 1300 time points (which contained 99% of the events), thus ensuring a uniform input dimension for the CNN. During data loading for CNN training, all peptide classes were downsampled to the size of the smallest class to maintain class balance. The dataset was subsequently split into an 80% training set and a 20% validation set. The CNN embedding model architecture consisted of an embedding layer (input dimension vocab size was number of unique remapped states + 1 (to include padding token) and output dimension was 128) that also masked the padding token (value 0). This was followed by a stack of two 1D CNN layers, each with 128 filters and a kernel size of 3, utilizing rectified linear unit activation and 'same' padding. Each layer's output was subjected to batch normalization and max pooling with a pool size of 2, followed by a dropout layer of 0.4. A global max pooling layer then summarized the processed sequence into a fixed-size representation. This was fed into a Dense layer with an output embedding dimension of 128, producing the final sequence embedding. A separate Dense classification head (with softmax activation) was attached to this embedding layer, allowing the entire CNN to be trained in a supervised manner for peptide classification. The model was compiled using the Adam optimizer with sparse categorical cross-entropy for the classification head (the embedding output had loss of none as it was not directly optimized during this phase). Training was performed for 100 epochs with a batch size of 32, incorporating early stopping (restoring best weights) and reduce learning rate on plateau callbacks to prevent overfitting and optimize 9 learning. The training history (loss and accuracy) were monitored on the validation set. After training, the final sequence embeddings were extracted from the trained model's embedding layer for both the training and testing sets. Hybrid model training and evaluation. For the hybrid model, event data (including sequences and features) for all peptides were loaded from a local database, applying the same minimum event duration filter range as used for the individual DL and ML models (e.g., ≥ 20 ms was most optimal empirically). The dataset was then split 80-20 into training and testing sets. The extracted CNN sequence embeddings (from the previously trained encoder) were then horizontally concatenated with their corresponding event-level features for both the training and testing sets. This combined feature vector served as the input for the final XGBoost classifier. The XGBoost classifier was configured with identical parameters to those used when trained solely on event-level features (as detailed in the previous section). Early stopping was also applied during XGBoost training. The performance of the hybrid XGBoost (S+F) model was assessed on the independent test set using a comprehensive classification report, providing overall accuracy, precision, recall, and F1-score. A confusion matrix was also generated to visualize per-class prediction accuracy. 10 Figures Fig. S1. Processing and analysis schemes. (A) Data acquisition and state labeling. (B) Event segmentation and feature calculation. (C) ML/DL classification workflow. 11 Tables Table S1. Recording time1 per peptide class in the dataset. Guest-Host Peptide Time (seconds) Time (hours) Ala 1717.8125 0.4772 Leu 513.3825 0.1426 Phe 514.57 0.1429 Thr 5141.785 1.4283 Trp 1929.47 0.536 TrpDL 1975.4625 0.5487 Tyr 2559.91 0.7111 TOTAL Dataset 14352.3925 3.9868 1These recording times include all data at the 70 mV voltage condition and peptide concentration range (5-20 nM). 12 Table S2. Translocation event counts at different minimum event duration filters. Guest-Host Peptide 5 ms Minimum Event Duration 20 ms Minimum Event Duration Events Before Downsampling Events After Downsampling1 Events Before Downsampling Events After Downsampling2 Ala 83,546 2,435 6,164 1,340 Leu 15,983 2,435 2,231 1,340 Phe 5,307 2,435 2,342 1,340 Thr 368,562 2,435 31,546 1,340 Trp 10,483 2,435 5,109 1,340 TrpDL 2,435 2,435 1,340 1,340 Tyr 77,338 2,435 37,978 1,340 Total 563,654 17,045 86,710 9,380 1For the 5 ms minimum event duration dataset, all peptide classes were downsampled to 2,435 events, based on the class with the fewest events (TrpDL), to create a balanced dataset for model training and evaluation. 2For the 20 ms minimum event duration dataset, all peptide classes were downsampled to 1,340 events, based on the class with the fewest events (TrpDL), to create a balanced dataset for model training and evaluation. 13 Table S3. Model performance metrics at varying minimum event duration.1 Model2 Metric3 Minimum event duration 5 ms 7.5 ms 10 ms 12.5 ms 15 ms 20 ms XGBoost (F) Accuracy 0.5483 0.7934 0.8235 0.8697 0.8737 0.9083 XGBoost (S+F) 0.5345 0.7905 0.8084 0.8755 0.8536 0.9078 CNN-Dense (C+F) 0.5456 0.6949 0.7312 0.7564 0.769 0.8006 CNN-Dense (S+F) 0.4303 0.6397 0.6471 0.686 0.6848 0.7281 TCN-Dense (S+F) 0.3711 0.5407 0.568 0.6242 0.6182 0.6578 XGBoost (F) Precision 0.5526 0.797 0.8262 0.8692 0.8739 0.9091 XGBoost (S+F) 0.5531 0.7913 0.8121 0.8746 0.8548 0.9089 CNN-Dense (C+F) 0.5602 0.7118 0.7457 0.7635 0.776 0.802 CNN-Dense (S+F) 0.4982 0.6455 0.6615 0.6968 0.6897 0.7412 TCN-Dense (S+F) 0.5103 0.6088 0.6056 0.6576 0.6627 0.692 XGBoost (F) Recall 0.5483 0.7934 0.8234 0.8697 0.8737 0.9083 XGBoost (S+F) 0.5345 0.7905 0.8083 0.8755 0.8536 0.9078 CNN-Dense (C+F) 0.5456 0.6949 0.731 0.7563 0.7689 0.8006 CNN-Dense (S+F) 0.4303 0.6396 0.647 0.6858 0.6847 0.7281 TCN-Dense (S+F) 0.3711 0.5407 0.568 0.6241 0.6181 0.6578 XGBoost (F) F1-score 0.5436 0.7941 0.8232 0.8687 0.8731 0.9085 XGBoost (S+F) 0.5159 0.7899 0.8083 0.8742 0.8533 0.9082 CNN-Dense (C+F) 0.5457 0.6983 0.7364 0.7583 0.7687 0.7981 CNN-Dense (S+F) 0.4308 0.6353 0.6483 0.6857 0.6825 0.7233 TCN-Dense (S+F) 0.3875 0.5596 0.5765 0.6318 0.6264 0.6597 1Performance based on test set evaluation metrics of different ML/DL classification models at different minimum event durations (N=1). Finalized replicated performance metrics (N=5) for these models at either minimum event duration extreme are presented in Table 1. 2Models are named as defined in the text. 3Metric names are abbreviated and refer to overall accuracy, macro-averaged precision, macro-averaged recall, and macro-averaged F1-score. 14 SI References 1. A. F. Kintzer et al., The protective antigen component of anthrax toxin forms functional octameric complexes. J. Mol. Biol. 392, 614–629 (2009). 2. A. Bernard, M. Payton, Fermentation and growth of Escherichia coli for optimal protein production. J. E. Coligan, B. M. Dunn, H. L. Plough, D. W. Speicher, P. T. Wingfield, Eds., Current Protocols in Protein Science (John Wiley & Sons, Inc., New York, 1995), vol. 5.3, pp. 1–18. 3. J. M. Colby, B. A. Krantz, Peptide Probes Reveal a Hydrophobic Steric Ratchet in the Anthrax Toxin Protective Antigen Translocase. J Mol Biol 427, 3598–3606 (2015). 4. K. Ghosal et al., Dynamic Phenylalanine Clamp Interactions Define SingleChannel Polypeptide Translocation through the Anthrax Toxin Protective Antigen Channel. J Mol Biol 429, 900–910 (2017). 5. K. L. Thoren, E. J. Worden, J. M. Yassif, B. A. Krantz, Lethal factor unfolding is the most force-dependent step of anthrax toxin translocation. Proc. Natl Acad. Sci. U.S.A. 106, 21555–21560 (2009). 6. M. Abadi et al. (2015) TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. 7. P. Remy, Temporal convolutional networks for keras. GitHub repository (2020). 8. T. Chen, C. Guestrin (2016) XGBoost: A Scalable Tree Boosting System. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (ACM), pp 785–794. 9. B. A. Krantz, Deep learning-based classification of peptide analytes from singlechannel nanopore translocation events. PLoS One 20, e0324777 (2025).