scieee AI-readable full text Open interactive document viewer

[WIP] Complete genome and QbD-guided reverse vaccinology for Streptococcus iniae strain SIKU01

ANDRES, Quentin Ludovic Stephane; Uchuwittayakul, Anurak; Srisapoome, Prapansak

Abstract

🧬 Whole Genome Assembly of 5 Streptococcus iniae bacteria isolated from diseased Asian Seabass, Thailand, comparative genomics, validation of WGS and in-sillico identification of antigens for vaccine biomanufacturing via Quality by Design (QbD). Organism: Streptococcus iniae isolates from diseased farmed Asian seabass (Lates calcarifer) Technologies: Illumina PE (short-reads) Methodologies: Single reference mapping De novo assembly Reference-guided de novo assembly Multi-reference mapping onto pangenome graphs Literature review and functional annotation of S. iniae proteome Identification of protein subset of candidate antigens based on functional annotations Pre-filtering using a QbD approach with a scoring matrix based on physico-chemical properties of Ags Second-filtering using a QbD approach with a scoring matrix based on E. coli expression system Identification of shared epitopes versus IEDB database of B- Cell epitopes in other animals Scoring and final selection of sets of best-scoring antigens for vaccine biomanufacturing using a range of downstream separation methods. NCBI Submission: Bioproject PRJNA933632 GenBank Sequence of SIKU01 Streptococcus iniae GenomeResults 98 proteins suitable for anion-exchange purification, 20 for cation-exchange, 100 for cellulose-affinity, 49 for silica-affinity, and 57-65 for plasmid DNA platforms. Importantly, this approach recovered well-validated antigens including enolase and GAPDH, which showed minimal sequence variation across our global dataset and have demonstrated 62-80% relative percent survival in previous trials.

Full text

Complete genome and QbD-guided reverse vaccinology for Streptococcus iniae 1 strain SIKU01 2 Quentin Ludovic Stephane Andres1,2, Worapong Singchat1,3, Kornsorn Srikulnath1,3,5 and 3 Prapansak Srisapoome1,4,* 4 5 1Animal Genomics and Bioresource Research Unit (AGB Research Unit), Faculty of Science, 6 Kasetsart University, Bangkok 10900, Thailand 7 2Doctor of Philosophy Program in Fishery Science and Technology (International Program), 8 Faculty of Fisheries, Kasetsart University, Bangkok, Thailand 9 3Special Research Unit for Wildlife Genomics (SRUWG), Department of Forest Biology, Faculty 10 4Department of Aquaculture, Faculty of Fisheries, Kasetsart University, Bangkok, Thailand 11 5Biodiversity Center Kasetsart University (BDCKU), Bangkok 10900, Thailand 12 13 *Corresponding author: Dr. Prapansak Srisapoome 14 4Department of Aquaculture, Faculty of Fisheries, Kasetsart University, Bangkok, Thailand 15 16 17 18 19 20 21 22 23 24 Abstract 25 Streptococcus iniae causes major losses in aquaculture. We assembled a complete, circular genome 26 for strain SIKU01 (2.09 Mb; 1,855 proteins) and applied reverse vaccinology to identify 27 conserved, epitope-rich antigens. We then integrated a Quality-by-Design framework to connect 28 antigen properties with manufacturability across recombinant protein and plasmid DNA platforms 29 across the proteome of S. iniae. The pipeline produced ranked, platform-aligned shortlists and 30 recovered known protective antigens (e.g., enolase, GAPDH). By linking genome-scale antigen 31 discovery to QbD downstream production constraints, our pipeline identified 98 AEX, 20 CEX, 32 100 cellulose, 49 silica, and 57–65 pDNA antigens for immediate downstream development. This 33 work provides a practical route to cost-effective S. iniae vaccine development and a general 34 template for selecting proteins for vaccine development in aquatic pathogens. 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 Background & Summary 50 Streptococcus iniae is a major bacterial pathogen in global aquaculture. It was initially identified 51 in the 1970s from an Amazon River dolphin (Inia geoffrensis) (GB PIER et al. 1976). It is 52 estimated that the pathogen is responsible for annual losses exceeding US$100 million worldwide 53 (CA Shoemaker et al. 2001)2 and causes mortality rates reaching 30-80% during outbreaks in 54 aquaculture ponds (M Chen et al. 2012)3 and a cumulative mortality at 3+ months to up to 70% in 55 some species (FP Mmanda et al. 2014)7. S. iniae has been identified across 27+ fish species (W 56 Agnew et al. 2007)4 and detected on all continents (Baiano & Barnes 2009 ; A Mishra et al. 2018)5,6 57 with particularly severe effects in intensive aquaculture systems where high stocking densities and 58 environmental stressors create optimal conditions for rapid disease transmission (M Chen et al. 59 2012)3 causing severe infections in a wide range of economically important fish species including 60 barramundi Asian seabass (Lates calcarifer), tilapia (Oreochromis spp.), and catfish (Siluriformes 61 spp.) (MNA Amal et al. 2011 ; RA Nawawi et al. 2008)8,9. In addition, it is also zoonotic and 62 human infections, while rare, have been previously reported in individuals handling infected fish. 63 64 Conventional biosecurity control of S. iniae relies on broad-spectrum antibiotics and 65 prevention prophylaxis with whole-cell or polysaccharide-based vaccines, which can provide 66 short-term protection in experimental trials. These vaccines, however, typically lose efficacy 67 within 4–6 months and are prone to long-term failure due to antigenic variation. Capsular 68 polysaccharides have long been considered key protective antigens in Streptococcus, but 69 especially unstable targets because S. iniae can rapidly alter its surface structures under immune 70 pressure and the genes responsible for capsule production are not always present in all strains of 71 S. iniae and are constituents of the accessory pangenome. As a result, existing vaccine strategies 72 relying on whole-cells often lack durability and fail to ensure broad protection across diverse S. 73 iniae strains. 74 Developing new vaccines for S. iniae requires addressing both biological and 75 manufacturing challenges especially for cost-efficiency because aquaculture fish affected by S. 76 iniae such as tilapia and seabass are considered low to medium value fish and hence require 77 cheaper vaccines. Therefore, production constrains look for candidate antigens that must be 78 conserved, immunogenic, and compatible with scalable production processes. The Quality by 79 Design (QbD) framework, widely used in pharmaceutical manufacturing, provides a systematic 80 way to resolve these challenges and connect antigen properties with downstream 81 manufacturability. Yet its application to the earliest stage of vaccine design—antigen discovery 82 through reverse vaccinology—remains underexplored. 83 In this study, we used new-generation sequencing to sequence pathogenic S. iniae bacteria 84 isolated from Asian seabass in Thailand in aquaculture pond and applied a combined reverse 85 vaccinology and QbD approach. One isolate, SIKU01, was selected for complete genome 86 assembly and subsequent in silico antigen screening. By integrating immunoinformatics with 87 manufacturability criteria, we created a ranked shortlist of epitope-rich antigens aligned with 88 multiple production platforms that is relevant to low-cost vaccinology. This “omics-to89 manufacturing” pipeline provides a practical framework for developing cost-effective and scalable 90 vaccines against S. iniae and other pathogens of economic importance in aquaculture. 91 92 Results 93 Genome Assembly and Functional Annotation of Streptococcus iniae SIKU01 94 We assembled a complete, circular genome of S. iniae strain SIKU01 (2.09 Mb, 37% GC) with 95 high completeness and no plasmids detected (Fig. 1a-b; Supplementary Fig. 1). Whole-genome 96 comparisons with reference strains QMA0248 (AS Alsheikh-Hussain et al. 2022), 89353 (HY 97 Gong et al. 2017), SF1 (Zhang et al. 2014) and LSSM211007Si which confirmed strong 98 macrosyntenic conservation (Fig. 1c-d). We also studied the synteny with the original Amazon 99 dolphin isolate QMA0141 from 1976 (GB PIER et al. 1976) (Supplementary Fig. 2). 100 Annotation of strain SIKU01 identified 2,004 genes, including 1,855 protein-coding 101 sequences and 77 RNAs. Functional classification revealed a broad repertoire of metabolic 102 enzymes, transporters, DNA repair proteins, and cell-surface factors, with additional subsets of 103 stress-response proteins, toxins, and antimicrobial resistance determinants (Fig. 1e). Domain 104 analysis of proteins highlighted extensive transmembrane features and conserved protein families 105 essential for growth, adaptation, and host interaction, while a subset of proteins had signal peptides 106 for secretion. 107 Together, these results confirm that SIKU01 retains a complete set of housekeeping 108 functions while harboring specialized virulence systems commonly found in other Streptococcus 109 iniae genomes. This well-annotated genome provided the basis for proteome-wide reverse 110 vaccinology and subsequent Quality by Design (QbD) manufacturability screening. 111 112 Figure 1. Complete genome assembly and functional annotation of Streptococcus iniae strain 113 SIKU01. (a) Genome assembly and reference-guided workflow from raw reads to validated 114 circular chromosome. (b) Assembly and annotation statistics, including genome size, GC content, 115 gene counts, and RNA features. (c) Comparative macrosynteny of SIKU01 against public S. iniae 116 reference strains, showing conserved genomic architecture. (d) Metadata of strain SIKU01 and 117 related isolates, including collection date, country, and host species, used in hybrid reference118 guided assembly. (e) Functional annotation of the SIKU01 proteome, including KEGG Mapper 119 categories, Gene Ontology subcellular localization, and InterProScan domain assignments. 120 121 Reverse Vaccinology and Identification of Antigenic Epitopes 122 Using the annotated SIKU01 proteome as input, we applied reverse vaccinology to identify 123 antigens with extracellular/surface signals and known B-/T-cell epitopes by mapping DIAMOND 124 hits to immunogenic epitopes found in humans and mice from the database IEDB (E ≤ 1×10⁻³) (K 125 Rawal et al. 2021 ; L Liu et al. 2009 ; DG Moriel et al. 2010)51–53. This yielded 17 epitope-positive 126 proteins spanning classic streptococcal targets that has been validated experimentally from the 127 literature. (Table 1). These epitope containing epitopes were carried forward, along with the rest 128 of the proteome of SIKU01, into the QbD manufacturability funnel (Fig. 2). 129 130 Table 1. Predicted antigenic epitopes from Streptococcus iniae SIKU01. Seventeen proteins 131 were identified by DIAMOND searches against the IEDB database (E ≤ 1×10⁻³) and are shown 132 with their locus tags in SIKU01, UniProtKB annotations for gene name, and matched amino acid 133 epitope sequences with cross-homology to IEDB database. 134 Locus Tag Subject ID (Gene Name in UniProtKB) Epitope Sequence (Amino Acid) SIKU01_001181 Enolase RAAADYLEVPLYNYLG SIKU01_001378 metal binding protein of ABC transporter (lipoprotein) EINTEEEGTPDQISSLIEK SIKU01_001395 cell envelope proteinase A LYIKAIEDAVALGADAINLS SIKU01_001681 pneumococcal histidine triad protein D YVTSHGDHYHYYNGKVPYDA SIKU01_001709 immunogenic secreted protein DASGTGKRRAEVMEKLDQWIDRHGGTP SIKU01_001765 Glyceraldehyde 3 phosphate dehydrogenase MVVKVGINGFGRIGRLAFRRIQ SIKU01_001769 surface exclusion protein VMLTAVGLTGFKLKKDIK SIKU01_001890 M protein QLPSTGDSYNPFFTASAMAI SIKU01_001902 GroEL LPTLVLNKIRGTFNVVAVKAPGFGDRRKAM SIKU01_001991 Inosine 5' monophosphate dehydrogenase [Streptococcus agalactiae] VVKVGIGPGSIC SIKU01_001995 conserved hypothetical protein AQNLRNILHGSDSIFYTFTS SIKU01_000047 Amidase EEKQAANQEAINTVA SIKU01_000165 DNA directed RNA polymerase subunit beta' [Streptococcus pyogenes] IDNGRRGRPITGP SIKU01_000216 pneumococcal histidine triad protein D YVTSHGDHFHFYNGKVPYDA SIKU01_000259 YSIRK signal domain/LPXTG anchor domain surface protein GVNQFIPFELFGGDGMLTRL SIKU01_000404 trigger factor ELDDELAKDIDEEV SIKU01_000973 hypothetical protein SPy1154 DSVLEAQMASQQLPVIGGIA 135 CQAs and CPPs in the QbD Framework 136 In Quality by Design (QbD), Critical Quality Attributes (CQAs) describe the antigen or plasmid 137 properties that determine safety, efficacy, and manufacturability, while Critical Process Parameters 138 (CPPs) are the thresholds, ranges, or decision rules imposed on those attributes to ensure the 139 product meets quality for cost-efficient purification of these antigens through various common 140 downstream manufacturability routes. 141 Figure 2 presents the QbD lifecycle (Fig. 2a) along with the QbD framework applied to 142 the proteome of S. iniae (Fig. 2b-d) and is structured as a multi-level funnel in which the number 143 of candidate antigens is progressively reduced from M0 (the full proteome of SIKU01), through 144 pre-M1 (core genome filters), M1 general (for M1 (antigenic variation, size, literature, epitope 145 evidence, and route-specific CAI for E. coli in protein expression or zebrafish/tilapia in pDNA 146 expression), and finally M2 (route-specific manufacturability matrices for silica and cellulose 147 affinity tags, ion-exchange purification of unmodified proteins, or pDNA design), yielding two 148 major downstream strategies: pDNA in two target hosts or recombinant protein in E. coli as 149 biofactory where affinity-purified proteins are modified with silica-binding peptides or CBMs, 150 while unmodified proteins are recovered by ion-exchange. 151 152 153 Figure 2. Quality by Design (QbD) lifecycle workflow for antigen selection and downstream 154 purification strategy. (a) Target Product Profile (TPP) definition (antigen type, delivery route, 155 expression system) and evaluation of Critical Material Attributes (CMAs). Antigens are scored in 156 silico for Critical Quality Attributes (CQAs) such as size, charge, conservation, codon adaptation, 157 and predicted purification behavior. (b) Upstream funnel showing stepwise reduction from M0 158 (protein-coding CDS, N=1,855) → pre-M1 (core genome, N=1,374) → M1-general survivors 159 (N=1,101). (c) From this shared pool, candidates branch into route-specific funnels: protein route 160 (E. coli biofactory, N=141) and plasmid DNA routes (zebrafish N=212; tilapia N=207). (d) M2 161 downstream manufacturability options, including silica and cellulose affinity fusions, ion162 exchange chromatography, and plasmid DNA optimization. Detailed thresholds and scoring rules 163 are given in the Design Space (DS) section, Table 2, and Supplementary Table N. 164 165 Definition of Design Spaces (DS) 166 We integrated the immunoinformatic dataset into a Quality by Design (QbD) framework that links 167 immunological relevance with manufacturability. For S. iniae SIKU01 antigens, the Critical 168 Quality Attributes (CQAs) included expression and annotation features (the CDS must encode a 169 protein, the open reading frame must be complete, and the gene must be carried as part of the core 170 pangenome), sequence and structural properties (protein length, molecular weight, hydrophobicity 171 by GRAVY, instability index, isoelectric point, and net charge at pH 7), conservation (normalized 172 Shannon entropy, H.norm, from multiple nucleotide sequence alignments), translational potential 173 (codon adaptation index, CAI, in E. coli as well as in zebrafish and Nile tilapia), gene 174 characteristics (GC3 fraction, gene length in nucleotides, and the presence or absence of internal 175 Type IIS restriction sites), and external evidence (literature support in PubMed or WHO vaccine 176 candidate databases, and curated epitope presence in IEDB as shown in Table 1). Route-specific 177 descriptors, such as silica-binding peptide length (~20 aa) and cellulose-binding module length 178 (≈120–300 aa), although not scored, were also included as CQAs. 179 M0 (mandatory inclusion). At this stage, candidates were retained only if they encoded a 180 valid protein-coding CDS, had a full-length ORF. Any sequence failing these requirements was 181 excluded. 182 Pre-M1 (mandatory inclusion). At this stage, candidates were retained only if they were 183 present in the core genome fraction of the pangenome. Any sequence failing these requirements 184 was excluded. 185 M1 (general filters). Candidates that passed M0 were then assessed for general 186 immunological relevance. Acceptable protein length was defined as 100–699 amino acids; 187 proteins shorter than 100 aa or longer than 699 aa were excluded. Molecular weight between 20– 188 Platform-specific gates (M1) Results 276 Filters were adapted to expression platforms in a host-dependent manner (Fig. 5). For the protein 277 route in E. coli (141 proteins), feasible antigens required CAI (E. coli) ≥ Q3, moderate 278 hydrophobicity (GRAVY −0.5 to +0.5), and stability (instability index ≤ 40) (Fig. 5a). For the 279 pDNA routes, the only M1 gate was host-specific CAI ≥ Q3, yielding zebrafish (212 proteins) and 280 tilapia (207 proteins) (Fig. 5b–c). GC3 and nucleotide length constraints were not applied until 281 M2. In both hosts, density contours showed clustering of survivors within narrow codon usage and 282 sequence length ranges, reflecting platform-specific optimization for translation efficiency and 283 manufacturability. 284 285 Figure 5. Expression system and purification design spaces for downstream 286 manufacturability filtering. (a) M1 Protein (E. coli) candidates projected by MW, pI, and codon 287 adaptation index (CAI), highlighting physicochemical compatibility with bacterial recombinant 288 expression. (b) pDNA platform (Danio rerio, zebrafish) candidates filtered by nucleotide length, 289 GC3 composition, and CAI. (c) pDNA platform (Oreochromis niloticus, Nile tilapia) candidates 290 filtered by nucleotide length, GC3 composition, and CAI. Density contours indicate clustering of 291 feasible antigens under each platform’s codon usage constraints. Final survivors represent 292 platform-ready antigen pools compatible with host translation and manufacturing requirements. 293 294 Manufacturability design spaces (M2) Results 295 To evaluate manufacturability, each downstream platform was subjected to stepwise Quality by 296 Design (QbD) filtering (M2 criteria) (Figure 6). Mapping antigens into ion-exchange charge space 297 (Fig. 6a) showed that most proteins were acidic at pH 7 and thus compatible with anion exchange 298 (AEX), whereas only a minor fraction were sufficiently basic to survive cation exchange (CEX). 299 This was reflected in the survivor pools (Fig. 6b), with 98 proteins retained by AEX and only 20 300 by CEX. Buffer-specific ranges (Figs. 6c–f) confirmed that AEX survivors were stable across 301 nearly all chemistries (82–98 proteins), while CEX buffers consistently yielded the same restricted 302 set (~20 proteins). The full antigen lists for AEX and CEX are provided in Supplementary Table 303 S1. Affinity routes imposed distinct orthogonal constraints: cellulose excluded proteins longer than 304 400 aa (Fig. 6g), because the CBM fusion tag is itself large and places a metabolic burden on 305 recombinant expression; minimizing the size of the fused antigen reduces energy demand and 306 steric hindrance, thereby improving folding and yield. In contrast, silica affinity (Fig. 6j) was 307 governed by electrostatic interactions with surface silanol groups: proteins with pI between 7–9 308 acquire a net positive charge at physiological pH, enabling stable adsorption to negatively charged 309 silica. These filters reduced the pools to 100 cellulose-compatible and 49 silica-compatible 310 proteins, detailed in Supplementary Table S2 and Supplementary Table S3. 311 For plasmid DNA (pDNA) platforms, Critical Quality Attribute (CQA) filters were applied 312 sequentially. In O. niloticus (Figs. 6h–i), proteins were first assessed for CAI ≥ Q3 (~0.60), 313 ensuring codon usage was well adapted to the host translation machinery; this retained 188 314 sequences. Applying GC3 ≥ Q3 (~0.32) halved the pool to 95, as high G/C content at the third 315 codon position improves mRNA stability and reduces metabolic stress. A further antigen length 316 cutoff (≤ 2200 nt) yielded 57 candidates, reflecting the reduced transcriptional and translational 317 burden of shorter constructs. Finally, proteins containing Type IIS restriction sites were excluded, 318 because these enzymes (e.g., BsaI, SapI) cleave outside their recognition sequence and disrupt 319 modular DNA assembly workflows. The final Nile tilapia pDNA pool remained at 57 survivors. 320 In D. rerio (Figs. 6k–l), the thresholds were slightly higher (CAI ≥ Q3 ~0.70, GC3 ≥ Q3 ~0.31). 321 Here, 186 proteins passed the CAI gate, 82 remained after GC3 filtering, 66 after the length filter, 322 and 65 after Type IIS removal. Comprehensive candidate sets for both pDNA platforms are 323 available in Supplementary Table S4 and Supplementary Table S5. 324 Overall, the QbD funnel outputs platform-ready shortlists: 98 AEX, 20 CEX, 100 325 cellulose, 49 silica, and 57–65 pDNA antigens for immediate downstream development, which 326 are available as lists in the Supplementary Tables. 327 328 Figure 6. Manufacturability design spaces across purification platforms. (a) Ion-exchange 329 charge space for protein candidates: pI vs. net charge at pH 7. Shaded bands mark the AEX and 330 CEX inclusion regions; vaccine-referenced antigens are labeled. (b) Survivors after M2 filtering 331 by platform; bars show unique genes per route with counts and percentages. Only candidates fitting 332 the biophysiochemical criteria for their respective purification pathway are displayed in color bars. 333 This step reflects manufacturability constraints and downstream process optimization in the QbD 334 framework. (c-d) Buffer operating pH ranges for AEX (C) and CEX (D). Short tick marks indicate 335 the working pH (range midpoint). (e-f) Protein-route survivors per buffer evaluated at the 336 midpoint: AEX rule = base gate passed, z@7 ≤ 0, and pI ≤ pH_mid; CEX rule = base gate passed, 337 z@7 ≥ 0, and pI ≥ pH_mid. (g) Cellulose: antigen length distribution for M2 survivors; dashed 338 line at 400 aa. (j) Silicon binding peptides favors proteins with moderate pI (7-9) and positive net 339 charge at pH >7. (h, k) pDNA (M2) scatterplots by host—tilapia (h) and zebrafish (k)—showing 340 CAI vs. GC3; dashed lines mark the dataset thresholds. (i, l) pDNA-route survivors passing each 341 manufacturability gate for tilapia (i) and zebrafish (l): CAI ≥ threshold, GC3 ≥ threshold, length ≤ 342 median, and no Type IIS sites. 343 344 Cross-validation with literature 345 We validated our shortlisted antigens against previously tested S. iniae vaccine targets reported in 346 multiple hosts, including Channel catfish (Channa striata) (E Wang et al. 2016)73, Asian seabass 347 (Lates calcarifer) (P Kayansamruaj et al. 2017 ; J Wang et al. 2014)74, 36, Olive flounder 348 (Paralichthys olivaceus) (X Sheng et al. 2018)39, Zebrafish (Danio rerio) (JD Membrebe et al. 349 2016)75, Mouse (Mus musculus) (J Wang et al. 2015)38, Turbot fish (Scophthalmus maximus) 350 (Zhang et al. 2014)27. 351 352 Among the top-ranked candidates consistently found across purification routes, well353 studied vaccines such as enolase, GAPDH, and GroEL were recovered, overlapping with antigens 354 that have already conferred in vivo protection and supporting the predictive accuracy of the QbD 355 pipeline (Supplementary Table N). 356 To further assess their suitability, we modelled two representative antigens, GAPDH and 357 enolase, both well established in the literature. Structural predictions generated with AlphaFold2 358 (Z. Yang et al., 2023)⁷⁶ for S. iniae SIKU01 confirmed conserved folds and surface-exposed loops 359 (L. Wang et al., 2017)⁷⁷. Several epitope-containing regions identified through reverse vaccinology 360 were surface-accessible (Fig. 7a–b, yellow), consistent with prior reports (V. Gent et al., 2024)⁷⁸. 361 These epitopes (Table 1) can be directly incorporated into subunit vaccines (X Sheng et al. 2018)39 362 or used in chimeric multi-epitope constructs (A Pumchan et al. 2020)79, further validating their 363 suitability as vaccine targets. Both enolase and GAPDH displayed epitope-rich regions spatially 364 separated from catalytic or hypervariable sites, reinforcing their accessibility and stability as 365 broad-spectrum candidates. 366 367 Figure 7. Structural mapping of variability, epitopes, and active sites in Streptococcus iniae 368 vaccine candidates. (a) Shannon entropy profiles show that GAPDH and enolase are largely 369 conserved with limited variable regions. (b) GAPDH (336 aa, UniProt Q7BB80) carries a 370 predicted N-terminal epitope (MVVKVGINGFGRIGRLAFRRIQ) positioned near, but not 371 overlapping with, the active site and adjacent to a hypervariable patch (1–89% variation). (c) 372 Enolase (435 aa, UniProt T1TFA0) contains a predicted epitope (RAAADYLEVPLYNYLG) 373 located opposite the active site and spatially separated from five hypervariable surface regions 374 (HR1–HR5), indicating stability and accessibility as a vaccine target. 375 376 Discussion 377 This study presents the first integrated reverse vaccinology and Quality by Design (QbD) 378 framework for developing cost-effective vaccines against Streptococcus iniae, addressing a critical 379 need in global aquaculture where annual losses exceed US$1 billion (JM Amillano-Cisneros et al. 380 2025 ; M Chen et al. 2012)3. By combining complete genome sequencing of strain SIKU01 with 381 systematic manufacturability screening, we identified platform-specific antigen shortlists that 382 balance immunological relevance with production feasibility—a crucial consideration for vaccines 383 targeting low-to-medium value fish species. 384 385 Integration of QbD with Reverse Vaccinology 386 Our QbD-guided approach represents a paradigm shift from traditional vaccine discovery 387 pipelines. Rather than identifying antigens solely based on immunological properties, we 388 systematically evaluated 1,855 proteins through staged manufacturability filters that connect 389 upstream discovery with downstream production constraints. This "omics-to-manufacturing" 390 pipeline reduced the candidate pool to 98 proteins suitable for anion-exchange purification, 20 for 391 cation-exchange, 100 for cellulose-affinity (G Carrard et al. 2000), 49 for silica-affinity (AI Freitas 392 et al. 2022), and 57-65 for plasmid DNA platforms. Importantly, this approach recovered well393 validated antigens including enolase and GAPDH, which showed minimal sequence variation 394 across our global dataset and have demonstrated 62-80% relative percent survival in previous 395 trials. 396 The biophysical landscape analysis (Figure 3) revealed that the S. iniae proteome exhibits 397 distinct clustering patterns that inform manufacturability decisions. The bimodal distribution of 398 isoelectric points (mean 7.03, SD 2.21) naturally segregates proteins into acidic and basic 399 populations, explaining why anion-exchange chromatography captured nearly five times more 400 candidates than cation-exchange (98 vs 20 proteins). This finding has immediate practical 401 implications: vaccine manufacturers can prioritize AEX-compatible antigens to maximize yields 402 while minimizing purification complexity. 403 404 Platform-Specific Considerations and Vaccine Efficacy 405 Our analysis of published vaccine trials reveals clear platform-dependent efficacy patterns that 406 validate our multi-route approach. DNA vaccines consistently achieve the highest protection (80407 95% RPS), likely due to sustained antigen expression and proper post-translational modifications 408 in the host. This aligns with our identification of 57-65 pDNA candidates optimized for zebrafish 409 and tilapia codon usage. The slightly higher CAI threshold for zebrafish (≥0.70) versus tilapia 410 (≥0.60) reflects species-specific translation machinery differences that must be considered during 411 vaccine deployment. 412 For recombinant protein production, our filtering for E. coli expression (141 candidates) 413 prioritized proteins with moderate hydrophobicity (GRAVY -0.5 to +0.5) and stability indices ≤40, 414 addressing the common challenge of inclusion body formation in bacterial systems (DM Francis 415 et al. 2010). The integration of affinity tags (either cellulose-binding modules or silica-binding 416 peptides) offers a cost-effective alternative to traditional His-tag purification (EA Woestenenk et 417 561 Functional Annotations and Epitope Detection in S. iniae Proteome 562 - Homology-Based Functional Annotations 563 Following genome annotation, functional characterization of the predicted proteome 564 was performed using InterProScan v5.48-83.0 (E Quevillon et al. 2005)49 to identify 565 InterPro protein families (M Blum et al. 2021) and their protein domains, motifs, and 566 assign Gene Ontology (GO) terms. Functional pathway assignments and hierarchical 567 classifications were conducted with KEGG Mapper and KEGG Brite (M Kanehisa et al. 568 2020)50, while GO terms followed the Gene Ontology Consortium framework (M 569 Ashburner et al. 2000) 51; The Gene Ontology Consortium, 2021). Predicted protein 570 sequences were further analyzed for structural and localization features. TMHMM v2.0 (A 571 Krogh et al. 2001) was used to predict transmembrane helices, and SignalP v5.0 (JJA 572 Armenteros et al. 2019) identified N-terminal signal peptides, applying Gram-positive 573 bacterial models. Additional domain annotations were obtained from PFAM (J Mistry et 574 al. 2021), SMART (I Letunic et al. 2021), TIGRFAMs (W Li et al. 2021), and 575 SUPERFAMILY (AP Pandurangan et al. 2018). Manual curation integrated UniProtKB 576 keyword annotations, KEGG pathway mappings, and GO terms reported in prior literature 577 to classify virulence factors (VFs) into literature-supported functional categories. 578 579 - Epitope Detection (IEDB) 580 Epitope detection was carried out by retrieving known antigenic peptides from 581 Streptococcus species in the Immune Epitope Database (IEDB) (R Vita et al. 2024)70 and 582 aligning them against the S. iniae proteome using DIAMOND BlastP (B Buchfink et al. 583 2021)71. Proteins containing identical or homologous regions to known B-cell epitopes 584 were flagged as epitope-containing candidates. 585 586 - Physiochemical Characterization of Proteins 587 Homology-based functional annotation was refined through alignment of the annotated 588 SIKU01 proteome against the S. iniae reference proteome in UniProtKB. Physicochemical 589 properties—including isoelectric point (pI) and molecular weight (MW)—were calculated 590 from translated protein sequences (.faa) using the SeqinR package in R(D Charif et al. 591 2007). Coding sequences were derived from Prokka and NCBI PGAP annotations. Finally, 592 a targeted literature review, supported by a custom text-mining script, was conducted to 593 scrap PUBMED literature and integrate experimental data and contextualize antigen 594 candidates with known virulence or immunogenic roles. 595 596 - Codon Adaptation Index (CAI) Calculation 597 Codon usage frequencies for target organisms were obtained from the Kazusa Codon 598 Usage Database available online at (https://www.kazusa.or.jp/codon/cgi599 bin/showcodon.cgi?species=8128&aa=1&style=GCG). Species-specific codon frequency 600 tables were downloaded providing frequencies per thousand codons for each organism's 601 coding sequences. The Codon Adaptation Index (CAI) quantifies the degree of preference 602 for synonymous codons in a given organism, where values range from 0 (least preferred) 603 to 1 (most preferred). For each amino acid family, the relative adaptiveness of codon i was 604 calculated as wi = fi / fmax, where fi represents the frequency per thousand of codon i for 605 its amino acid and fmax represents the frequency per thousand of the most frequently used 606 codon for that amino acid. For example, proline codons in Oreochromis niloticus were 607 calculated as follows: CCG had wi = 7.39 / 16.53 = 0.447, CCA had wi = 14.59 / 16.53 = 608 0.883, CCT had wi = 16.53 / 16.53 = 1.000 (optimal), and CCC had wi = 14.91 / 16.53 = 609 0.902. The value 16.53 represents the highest frequency among all proline codons, making 610 CCT the optimal reference codon. Next, for complete gene sequences, CAI was computed 611 as the geometric mean of relative adaptiveness values using the formula CAI = (∏Lk=1 612 wk)1/L, where the product encompasses all L sense codons in the gene sequence and L 613 represents the total number of sense codons excluding stop codons. Stop codons (UAA, 614 UAG, UGA) and the single methionine codon (AUG) were excluded from CAI 615 calculations as they lack synonymous alternatives for optimization. CAI values were 616 implemented by first parsing codon frequency tables from the Kazusa database, identifying 617 the optimal codon for each amino acid, calculating relative adaptiveness for all 61 sense 618 codons, and finally computing the geometric mean of constituent codon values for gene 619 sequences. 620 621 Data Records 622 All genome assemblies and raw sequencing data generated in this study have been deposited in 623 public repositories. The project has been registered under NCBI BioProject PRJNA933632, with 624 the BioSample accession number SAMN33244440. Raw sequencing reads are hosted in the NCBI 625 Sequence Read Archive (SRA): Illumina (150PE), 10 FASTQ files, 5 raw sequencing datasets; 626 SRR23406918 for SIKU01, while other accessions were SRR23406921, SRR23406920, 627 SRR23406919, and SRR23406922 for SIKU02–SIKU05, respectively. The annotated genome for 628 SIKU01 has been submitted to the NCBI GenBank database and is available at CP121692.1. Due 629 to time constraints, only the SIKU01 isolate was submitted to NCBI GenBank. 630 631 Code Availability 632 All bioinformatics tools used in this study are publicly available, and command-line parameters 633 have been specified in the Methods section. Custom scripts for data processing and analysis are 634 available from the corresponding author upon reasonable request. 635 636 Acknowledgements 637 To the GenBank team for handling the genome annotation after 2 years of delay. 638 639 Author Contributions Statement 640 Quentin Ludovic Stephane Andres: Conceptualized the study, developed the methodology, 641 extracted DNA, performed bioinformatic analysis, curated data, generated visualizations, wrote 642 the original draft. Anurak Uchuwittayakul: Collected samples, extracted DNA, and performed 643 DNA sequencing. Kornsorn Srikulnath and Worapong Singchat: Collected samples. Prapansak 644 Srisapoome: Collected samples, supervised the project, and managed project administration. 645 646 Competing Interests 647 The authors declare no competing interests. 648 649 References 650 1. Pier, G. B., & Madin, S. H. (1976). Streptococcus iniae sp. nov., a beta-hemolytic 651 streptococcus isolated from an Amazon freshwater dolphin, Inia geoffrensis. 652 International Journal of Systematic Bacteriology, 26(4), 545–553. 653 https://doi.org/10.1099/00207713-26-4-545 654 2. Shoemaker, C. A., Klesius, P. H., & Evans, J. J. (2001). Prevalence of Streptococcus 655 iniae in tilapia, hybrid striped bass, and channel catfish on commercial fish farms in 656 the United States. American journal of veterinary research, 62(2), 174–177. 657 https://doi.org/10.2460/ajvr.2001.62.174 658 3. Chen, M., Li, L. P., Wang, R., Liang, W. W., Huang, Y., Li, J., Lei, A. Y., Huang, W. Y., 659 & Gan, X. (2012). PCR detection and PFGE genotype analyses of streptococcal clinical 660 isolates from tilapia in China. Veterinary microbiology, 159(3-4), 526–530. 661 https://doi.org/10.1016/j.vetmic.2012.04.035 662 4. Agnew, W., & Barnes, A. C. (2007). Streptococcus iniae: an aquatic pathogen of global 663 veterinary significance and a challenging candidate for reliable vaccination. Veterinary 664 microbiology, 122(1-2), 1–15. https://doi.org/10.1016/j.vetmic.2007.03.002 665 5. Baiano, J. C., & Barnes, A. C. (2009). Towards control of Streptococcus iniae. 666 Emerging infectious diseases, 15(12), 1891–1896. 667 https://doi.org/10.3201/eid1512.090232 668 6. Mishra, A., Nam, G. H., Gim, J. A., Lee, H. E., Jo, A., & Kim, H. S. (2018). Current 669 Challenges of Streptococcus Infection and Effective Molecular, Cellular, and 670 Environmental Control Methods in Aquaculture. Molecules and cells, 41(6), 495–505. 671 https://doi.org/10.14348/molcells.2018.2154 672 7. Mmanda, F. P., Zhou, S., Zhang, J., Zheng, X., An, S., & Wang, G. (2014). Massive 673 mortality associated with Streptococcus iniae infection in cage-cultured red drum 674 (Sciaenops ocellatus) in Eastern China. African Journal of Microbiology Research, 675 8(16), 1722-1729. https://doi.org/10.5897/AJMR2014.6659 676 8. Azmai, M. N. A. & Saad, M. (2011). Streptococcosis in tilapia (Oreochromis niloticus): 677 A review. PERTANIKA J. Trop. Agric. Sci. 34, 195–206. 678 9. Nawawi, R. A., Baiano, J. & Barnes, A. C. (2008). Genetic variability amongst 679 Streptococcus iniae isolates from Australia. J. Fish Dis. 31, 305–309. 680 https://doi.org/10.1111/j.1365-2761.2007.00880.x 681 10. Facklam, R. (2002). What happened to the Streptococci: Overview of taxonomic and 682 nomenclature changes. Clin. Microbiol. Rev. 15, 613–630. 683 https://doi.org/10.1128/cmr.15.4.613-630.2002 684 11. Irion, S. et al. (2021). Molecular investigation of recurrent Streptococcus iniae 685 epizootics affecting coral reef fish on an oceanic island suggests at least two distinct 686 emergence events. Front. Microbiol. 12. https://doi.org/10.3389/fmicb.2021.749734 687 12. Schar, D., Klein, E. Y., Laxminarayan, R., Gilbert, M. & Boeckel, T. P. V. (2020). 688 Global trends in antimicrobial use in aquaculture. Sci. Reports 10. 689 https://doi.org/10.1038/s41598-020-78849-3 690 13. Schar, D. et al. (2021). Twenty-year trends in antimicrobial resistance from aquaculture 691 and fisheries in Asia. Nat. Commun. 12. https://doi.org/10.1038/s41467-021-25655-8 692 14. Bromage, E. & Owens, L. (2007). Environmental factors affecting the susceptibility of 693 barramundi to Streptococcus iniae. Aquaculture 290, 224–228. 694 https://doi.org/10.1016/j.aquaculture.2009.02.038 695 15. Al-fattah, H. A. A., Mahmoud, N., Al-razik, M. A., Al-moghny, F. A. & Ibrahim, M. S. 696 (2020). Characterization and Pathogenicity of Streptococcus Iniae Isolated from 697 Oreochromis Niloticus Fish Farms in Kafr-Elshiekh Governorate, Egypt. Alexandria 698 Journal of Veterinary Sciences, 64 (2), 123-128. 699 http://dx.doi.org/10.5455/ajvs.4984617 700 16. Figueiredo, H. C., Netto, L. N., Leal, C. A., Pereira, U. P., & Mian, G. F. (2012). 701 Streptococcus iniae outbreaks in Brazilian Nile tilapia (Oreochromis niloticus L.) 702 farms. Brazilian journal of microbiology, 43(2), 576–580. 703 https://doi.org/10.1590/s1517-83822012000200019 704 17. Facklam, R., Elliott, J., Shewmaker, L. & Reingold, A. (2005). Identification and 705 characterization of sporadic isolates of Streptococcus iniae isolated from humans. J. 706 Clin. Microbiol. 43, 933–937. https://doi.org/10.1128/jcm.43.2.933-937.2005 707 18. Defoirdt, T., Sorgeloos, P. & Bossier, P. (2011). Alternatives to antibiotics for the 708 control of bacterial disease in aquaculture. Curr. Opin. Microbiol. 14, 251–258. 709 https://doi.org/10.1016/j.mib.2011.03.004 710 19. Jeong, Y. U., Subramanian, D., Jang, Y. H., Kim, D. H., Park, S. H., Park, K. I., ... & 711 Heo, M. S. (2016). Protective efficiency of an inactivated vaccine against 712 Streptococcus iniae in olive flounder, Paralichthys olivaceus. Fisheries & Aquatic Life, 713 24(1), 23-32. https://doi.org/10.1515/aopf-2016-0003 714 20. Eldar, A., Horovitcz, A., & Bercovier, H. (1997). Development and efficacy of a 715 vaccine against Streptococcus iniae infection in farmed rainbow trout. Veterinary 716 immunology and immunopathology, 56(1-2), 175–183. https://doi.org/10.1016/s0165717 2427(96)05738-8 718 21. Tanpichai, P., Chaweepack, S., Senapin, S., Piamsomboon, P., & Wongtavatchai, J. 719 (2023). Immune Activation Following Vaccination of Streptococcus iniae Bacterin in 720 Asian Seabass (Lates calcarifer, Bloch 1790). Vaccines, 11(2), 351. 721 https://doi.org/10.3390/vaccines11020351 722 22. Eyngor, M. et al. (2008). Emergence of novel Streptococcus iniae exopolysaccharide723 producing strains following vaccination with nonproducing strains. Appl. Environ. 724 Microbiol. 74, 6892–6897. https://doi.org/10.1128/aem.00853-08 725 23. Millard, C. M. et al. (2012). Evolution of the capsular operon of Streptococcus iniae in 726 response to vaccination. Appl. Environ. Microbiol. 78, 8219–8226. 727 https://doi.org/10.1128/aem.02216-12 728 24. Lawrence, X. Y. et al. (2014). Understanding pharmaceutical quality by design. The 729 AAPS journal, 16(4), 771–783. https://doi.org/10.1208/s12248-014-9598-3 730 25. Alsheikh-Hussain, A. S., Ben Zakour, N. L., Forde, B. M., Rudenko, O., Barnes, A. C., 731 & Beatson, S. A. (2022). A high-quality reference genome for the fish pathogen 732 Streptococcus iniae. Microbial genomics, 8(3), 000777. 733 https://doi.org/10.1099/mgen.0.000777 734 26. Gong, H. Y., Wu, S. H., Chen, C. Y., Huang, C. W., Lu, J. K., & Chou, H. Y. (2017). 735 Complete Genome Sequence of Streptococcus iniae 89353, a Virulent Strain Isolated 736 from Diseased Tilapia in Taiwan. Genome announcements, 5(4), e01524-16. 737 https://doi.org/10.1128/genomeA.01524-16 738 27. Zhang B-c, Zhang J, Sun L (2014). Streptococcus iniae SF1: Complete Genome 739 Sequence, Proteomic Profile, and Immunoprotective Antigens. PLoS ONE 9(3): 740 e91324. https://doi.org/10.1371/journal.pone.0091324 741 28. Didelot, X. et al. (2022). A scalable analytical approach from bacterial genomes to 742 epidemiology Phil. Trans. R. Soc. B377: 20210246 743 http://doi.org/10.1098/rstb.2021.0246 744 29. Page, A. J. et al. (2015). Roary: rapid large-scale prokaryote pan genome analysis. 745 Bioinformatics 31(22), 3691-3693. https://doi.org/10.1093/bioinformatics/btv421 746 30. Tonkin-Hill, G., Lees, J. A., Bentley, S. D., Frost, S. D. W. & Corander, J. (2019). Fast 747 hierarchical Bayesian analysis of population structure. Nucleic Acids Res. 47, 5539– 748 5549. https://doi.org/10.1093/nar/gkz361 749 31. Croucher, N. J. et al. (2014). Rapid phylogenetic analysis of large samples of 750 recombinant bacterial whole genome sequences using Gubbins. Nucleic Acids Res. 751 43(3), e15. https://doi.org/10.1093/nar/gku1196 752 32. Yadav, A. K., Espaillat, A., & Cava, F. (2018). Bacterial Strategies to Preserve Cell 753 Wall Integrity Against Environmental Threats. Frontiers in microbiology, 9, 2064. 754 https://doi.org/10.3389/fmicb.2018.02064 755 33. Locke JB Colvin KM Datta AK, Patel SK, Naidu NN, Neely MN, Nizet V, Buchanan 756 JT (2007). Streptococcus iniae Capsule Impairs Phagocytic Clearance and Contributes 757 to Virulence in Fish. J Bacteriol 189: https://doi.org/10.1128/jb.01175-06 758 34. Eyngor, M. et al. (2007). Transcytosis of Streptococcus iniae through skin epithelial 759 barriers: an in vitro study. FEMS Microbiology Letters, Volume 277, Issue 2, 238–248. 760 https://doi.org/10.1111/j.1574-6968.2007.00973.x 761 35. Milani, C. J. E., Aziz, R. K., Locke, J. B., Dahesh, S., Nizet, V., & Buchanan, J. T. 762 (2010). The novel polysaccharide deacetylase homologue Pdi contributes to virulence 763 of the aquatic pathogen Streptococcus iniae. Microbiology, 156(Pt 2), 543–554. 764 https://doi.org/10.1099/mic.0.028365-0 765 36. Wang, J., Zou, L. L., & Li, A. X. (2014). Construction of a Streptococcus iniae sortase 766 A mutant and evaluation of its potential as an attenuated modified live vaccine in Nile 767 tilapia (Oreochromis niloticus). Fish & shellfish immunology, 40(2), 392–398. 768 https://doi.org/10.1016/j.fsi.2014.07.028 769 37. Baiano, J. C., Tumbol, R. A., Umapathy, A., & Barnes, A. C. (2008). Identification and 770 molecular characterisation of a fibrinogen binding protein from Streptococcus iniae. 771 BMC microbiology, 8, 67. https://doi.org/10.1186/1471-2180-8-67 772 38. Wang, J., Wang, K., Chen, D., Geng, Y., Huang, X., He, Y., Ji, L., Liu, T., Wang, E., 773 Yang, Q., & Lai, W. (2015). Cloning and Characterization of Surface-Localized α774 Enolase of Streptococcus iniae, an Effective Protective Antigen in Mice. International 775 journal of molecular sciences, 16(7), 14490–14510. 776 https://doi.org/10.3390/ijms160714490 777 39. Sheng, X., Gao, J., Liu, H., Tang, X., Xing, J., Zhan, W. (2018). Recombinant 778 phosphoglucomutase and CAMP factor as potential subunit vaccine antigens induced 779 high protection against Streptococcus iniae infection in flounder (Paralichthys 780 olivaceus). Journal of Applied Microbiology, Volume 125, Issue 4, 997–1007. 781 https://doi.org/10.1111/jam.13948 782 40. Bergmann, S., Rohde, M., & Hammerschmidt, S. (2004). Glyceraldehyde-3-phosphate 783 dehydrogenase of Streptococcus pneumoniae is a surface-displayed plasminogen784 binding protein. Infection and immunity, 72(4), 2416–2419. 785 https://doi.org/10.1128/iai.72.4.2416-2419.2004 786 41. Zou, L., Wang, J., Huang, B., Xie, M., & Li, A. (2010). A solute-binding protein for 787 iron transport in Streptococcus iniae. BMC microbiology, 10, 309. 788 https://doi.org/10.1186/1471-2180-10-309 789 42. Wang, J., Zou, L. L., & Li, A. X. (2013). A novel iron transporter in Streptococcus 790 iniae. Journal of fish diseases, 36(12), 1007–1015. https://doi.org/10.1111/j.1365791 2761.2012.01439.x 792 43. Molloy, E., Cotter, P., Hill, C. et al. (2011). Streptolysin S-like virulence factors: the 793 continuing sagA. Nat Rev Microbiol 9, 670–681. https://doi.org/10.1038/nrmicro2624 794 44. Lee, S.W., Mitchell, D.A., Markley, A.L., Hensler, M.E., Gonzalez, D., Wohlrab, A., 795 Dorrestein, P.C., Nizet, V., & Dixon, J.E. (2008). Discovery of a widely distributed 796 89. Tonkin-Hill, G. et al. (2020). Producing polished prokaryotic pangenomes with the 935 Panaroo pipeline. Genome Biol. 21. https://doi.org/10.1186/s13059-020-02090-4 936 90. Lischer, H.E.L., Shimizu, K.K. (2017). Reference-guided de novo assembly approach 937 improves genome reconstruction for related species. BMC Bioinformatics 18, 474. 938 https://doi.org/10.1186/s12859-017-1911-6 939 91. Carrard, G., Koivula, A., Söderlund, H., & Béguin, P. (2000). Cellulose-binding 940 domains promote hydrolysis of different sites on crystalline cellulose. Proceedings of 941 the National Academy of Sciences of the United States of America, 97(19), 10342– 942 10347. https://doi.org/10.1073/pnas.160216697 943 92. Francis, D. M., & Page, R. (2010). Strategies to optimize protein expression in E. coli. 944 Current protocols in protein science, Chapter 5(1), 5.24.1–5.24.29. 945 https://doi.org/10.1002/0471140864.ps0524s61 946 93. Woestenenk, E. A., Hammarström, M., van den Berg, S., Härd, T., & Berglund, H. 947 (2004). His tag effect on solubility of human proteins produced in Escherichia coli: a 948 comparison between four expression vectors. Journal of structural and functional 949 genomics, 5(3), 217–229. https://doi.org/10.1023/b:jsfg.0000031965.37625.0e 950 94. Andrews, S. et al. (2010). FastQC: A quality control tool for high throughput sequence 951 data. 952 95. Sturm, M., Schroeder, C. & Bauer, P. (2016). SeqPurge: highly-sensitive adapter 953 trimming for paired-end NGS data. BMC Bioinformatics. 17, 208. 954 https://doi.org/10.1186/s12859-016-1069-7 955 96. Wick, R. R., Judd, L. M., Gorrie, C. L. & Holt, K. E. (2017). Unicycler: Resolving 956 bacterial genome assemblies from short and long sequencing reads. PLOS Comput. 957 Biol. 13, e1005595. https://doi.org/10.1371/journal.pcbi.1005595 958 97. Bankevich, A. et al. (2012). SPAdes: A new genome assembly algorithm and its 959 applications to single-cell sequencing. J. Comput. Biol. 19, 455–477. 960 https://doi.org/10.1089/cmb.2012.0021 961 98. Wick, R. R., Schultz, M. B., Zobel, J. & Holt, K. E. (2015). Bandage: interactive 962 visualization of de novo genome assemblies. Bioinformatics 31, 3350–3352. 963 https://doi.org/10.1093/bioinformatics/btv383 964 99. Quast, C. et al. (2013). The SILVA ribosomal RNA gene database project: improved 965 data processing and web-based tools. Nucleic Acids Res. 41, D590–6. 966 https://doi.org/10.1093/nar/gks1219 967 100. Langmead, B. & Salzberg, S. L. (2012). Fast gapped-read alignment with bowtie 968 2. Nat. Methods 9, 357–359. https://doi.org/10.1038/nmeth.1923 969 101. Darling, A. E., Mau, B. & Perna, N. T. (2010). progressiveMauve: Multiple genome 970 alignment with gene gain, loss and rearrangement. PLoS ONE 5, e11147. 971 https://doi.org/10.1371/journal.pone.0011147 972 102. Walker, B. J. et al. (2014). Pilon: An integrated tool for comprehensive microbial 973 variant detection and genome assembly improvement. PLoS ONE 9, e112963. 974 https://doi.org/10.1371/journal.pone.0112963 975 103. Md, V., Misra, S., Li, H. & Aluru, S. (2019). Efficient architecture-aware 976 acceleration of bwa-mem for multicore systems. BioarXiv. 977 https://doi.org/10.1109/IPDPS.2019.00041. 978 104. Robinson, J. T. et al. (2011). Integrative genomics viewer. Nat. Biotechnol. 29, 24– 979 26. https://doi.org/10.1038/nbt.1754 980 105. Tatusova, T. et al. (2016). NCBI prokaryotic genome annotation pipeline. Nucleic 981 Acids Res. 44, 6614–6624. https://doi.org/10.1093/nar/gkw569 982 106. Minh, B. Q., Schmidt, H. A., Chernomor, O., Schrempf, D., Woodhams, M. D., von 983 Haeseler, A., Lanfear, R. (2020). IQ-TREE 2: New Models and Efficient Methods for 984 Phylogenetic Inference in the Genomic Era. Molecular Biology and Evolution, Volume 985 37, Issue 5, 1530–1534. https://doi.org/10.1093/molbev/msaa015 986 107. Kalyaanamoorthy, S., Minh, B., Wong, T. et al. (2017). ModelFinder: fast model 987 selection for accurate phylogenetic estimates. Nat Methods 14, 587–589. 988 https://doi.org/10.1038/nmeth.4285 989 108. Charif, D., Lobry, J.R. (2007). SeqinR 1.0-2: A Contributed Package to the R 990 Project for Statistical Computing Devoted to Biological Sequences Retrieval and Analysis. 991 In: Bastolla, U., Porto, M., Roman, H.E., Vendruscolo, M. (eds) Structural Approaches to 992 Sequence Evolution. Biological and Medical Physics, Biomedical Engineering. Springer, 993 Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-35306-5_10 994 109. Taouk, M. L., Featherstone, L. A., Taiaroa, G., Seemann, T., Ingle, D. J., Stinear, T. 995 P., & Wick, R. R. (2025). Exploring SNP filtering strategies: the influence of strict vs soft 996 core. Microbial genomics, 11(1), 001346. https://doi.org/10.1099/mgen.0.001346 997 110. Tonkin-Hill, G., Gladstone, R. A., Pöntinen, A. K., Arredondo-Alonso, S., Bentley, 998 S. D., & Corander, J. (2023). Robust analysis of prokaryotic pangenome gene gain and loss 999 rates with Panstripe. Genome research, 33(1), 129–140. 1000 https://doi.org/10.1101/gr.277340.122 1001 111. Martí I Líndez, A. A., & Reith, W. (2021). Arginine-dependent immune responses. 1002 Cellular and molecular life sciences : CMLS, 78(13), 5303–5324. 1003 https://doi.org/10.1007/s00018-021-03828-4 1004 112. Tettelin, H., Riley, D., Cattuto, C., & Medini, D. (2008). Comparative genomics: 1005 the bacterial pan-genome. Current opinion in microbiology, 11(5), 472–477. 1006 https://doi.org/10.1016/j.mib.2008.09.006 1007 113. Sun, Y., Hu, Y. H., Liu, C. S., & Sun, L. (2010). Construction and analysis of an 1008 experimental Streptococcus iniae DNA vaccine. Vaccine, 28(23), 3905–3912. 1009 https://doi.org/10.1016/j.vaccine.2010.03.071 1010 114. Sun, Y., Zhang, M., Liu, C. S., Qiu, R., & Sun, L. (2012). A divalent DNA vaccine 1011 based on Sia10 and OmpU induces cross protection against Streptococcus iniae and Vibrio 1012 anguillarum in Japanese flounder. Fish & shellfish immunology, 32(6), 1216–1222. 1013 https://doi.org/10.1016/j.fsi.2012.03.024 1014 115. Sun, Y., Sun, L., Xing, M. Q., Liu, C. S., & Hu, Y. H. (2013). SagE induces highly 1015 effective protective immunity against Streptococcus iniae mainly through an immunogenic 1016 domain in the extracellular region. Acta veterinaria Scandinavica, 55(1), 78. 1017 https://doi.org/10.1186/1751-0147-55-78 1018 116. Sun, Y., Hu, Y. H., Liu, C. S., & Sun, L. (2012). Construction and comparative 1019 study of monovalent and multivalent DNA vaccines against Streptococcus iniae. Fish & 1020 shellfish immunology, 33(6), 1303–1310. https://doi.org/10.1016/j.fsi.2012.10.004 1021 117. Liu, C., Hu, X., Cao, Z., Sun, Y., Chen, X., & Zhang, Z. (2019). Construction and 1022 characterization of a DNA vaccine encoding the SagH against Streptococcus iniae. Fish & 1023 shellfish immunology, 89, 71–75. https://doi.org/10.1016/j.fsi.2019.03.045 1024 118. Wang, E., Long, B., Wang, K., Wang, J., He, Y., Wang, X., Yang, Q., Liu, T., Chen, 1025 D., Geng, Y., Huang, X., Ouyang, P., & Lai, W. (2016). Interleukin-8 holds promise to 1026 serve as a molecular adjuvant in DNA vaccination model against Streptococcus iniae 1027 infection in fish. Oncotarget, 7(51), 83938–83950. 1028 https://doi.org/10.18632/oncotarget.13728 1029 119. Sheng, X., Zhang, H., Liu, M., Tang, X., Xing, J., Chi, H., & Zhan, W. (2023). 1030 Development and Evaluation of Recombinant B-Cell Multi-Epitopes of PDHA1 and 1031 GAPDH as Subunit Vaccines against Streptococcus iniae Infection in Flounder 1032 (Paralichthys olivaceus). Vaccines, 11(3), 624. https://doi.org/10.3390/vaccines11030624 1033 120. Cao, Y., Liu, J., Liu, G., Du, H., Liu, T., Wang, G., Wang, Q., Zhou, Y., & Wang, E. 1034 (2023). Exploring the Immunoprotective Potential of a Nanocarrier Immersion Vaccine 1035 Encoding Sip against Streptococcus Infection in Tilapia (Oreochromis niloticus). Vaccines, 1036 11(7), 1262. https://doi.org/10.3390/vaccines11071262 1037 121. Jia Liu, Ye Cao, Haixiang Ma, Hui Du, Tianqiang Liu, Gaoxue Wang, Mingzhu 1038 Liu, Qing Wang, Pengfei Li, Erlong Wang, Enolase-based nanovaccine immersion 1039 immunization induces robust immunity and protection against Streptococcus infection in 1040 tilapia, Aquaculture, Volume 576, (2023), 1041 https://doi.org/10.1016/j.aquaculture.2023.739849 1042 122. Sun, Y., Hu, Y. H., Liu, C. S., & Sun, L. (2012). A Streptococcus iniae DNA vaccine 1043 delivered by a live attenuated Edwardsiella tarda via natural infection induces cross-genus 1044 protection. Letters in applied microbiology, 55(6), 420–426. 1045 https://doi.org/10.1111/j.1472-765X.2012.03307.x 1046 123. Buchanan, J. T., Stannard, J. A., Lauth, X., Ostland, V. E., Powell, H. C., 1047 Westerman, M. E., & Nizet, V. (2005). Streptococcus iniae phosphoglucomutase is a 1048 virulence factor and a target for vaccine development. Infection and immunity, 73(10), 1049 6935–6944. https://doi.org/10.1128/IAI.73.10.6935-6944.2005 1050 124. Pridgeon, J. W., & Klesius, P. H. (2011). Development and efficacy of a 1051 novobiocin-resistant Streptococcus iniae as a novel vaccine in Nile tilapia (Oreochromis 1052 niloticus). Vaccine, 29(35), 5986–5993. https://doi.org/10.1016/j.vaccine.2011.06.036 1053 125. Pridgeon, J. W., Zhang, D., & Zhang, L. (2014). Complete Genome Sequence of 1054 the Attenuated Novobiocin-Resistant Streptococcus iniae Vaccine Strain ISNO. Genome 1055 announcements, 2(3), e00510-14. https://doi.org/10.1128/genomeA.00510-14 1056 126. Wang, J., Zou, L. L., & Li, A. X. (2014). Construction of a Streptococcus iniae 1057 sortase A mutant and evaluation of its potential as an attenuated modified live vaccine in 1058 Nile tilapia (Oreochromis niloticus). Fish & shellfish immunology, 40(2), 392–398. 1059 https://doi.org/10.1016/j.fsi.2014.07.028 1060 127. Liu, Y., Li, L., Yu, F., Luo, Y., Liang, W., Yang, Q., Wang, R., Li, M., Tang, J., Gu, 1061 Q., Luo, Z., & Chen, M. (2020). Genome-wide analysis revealed the virulence attenuation 1062 mechanism of the fish-derived oral attenuated Streptococcus iniae vaccine strain YM011. 1063 Fish & shellfish immunology, 106, 546–554. https://doi.org/10.1016/j.fsi.2020.07.046 1064 128. Heckman, T. I., Shahin, K., Henderson, E. E., Griffin, M. J., & Soto, E. (2022). 1065 Development and efficacy of Streptococcus iniae live-attenuated vaccines in Nile tilapia, 1066 Oreochromis niloticus. Fish & shellfish immunology, 121, 152–162. 1067 https://doi.org/10.1016/j.fsi.2021.12.043 1068 129. Xiong, X., Peng, Y., Chen, R., Liu, X., & Jiang, F. (2023). Efficacy and 1069 transcriptome analysis of golden pompano (Trachinotus ovatus) immunized with a 1070 formalin-inactived vaccine against Streptococcus iniae. Fish & shellfish immunology, 134, 1071 108489. https://doi.org/10.1016/j.fsi.2022.108489 1072 130. Blum, M., Chang, H. Y., Chuguransky, S., Grego, T., Kandasaamy, S., Mitchell, A., 1073 Nuka, G., Paysan-Lafosse, T., Qureshi, M., Raj, S., Richardson, L., Salazar, G. A., 1074 Williams, L., Bork, P., Bridge, A., Gough, J., Haft, D. H., Letunic, I., Marchler-Bauer, A., 1075 Mi, H., … Finn, R. D. (2021). The InterPro protein families and domains database: 20 1076 years on. Nucleic acids research, 49(D1), D344–D354. 1077 https://doi.org/10.1093/nar/gkaa977 1078 131. Krogh, A., Larsson, B., von Heijne, G., & Sonnhammer, E. L. (2001). Predicting 1079 transmembrane protein topology with a hidden Markov model: application to complete 1080 genomes. Journal of molecular biology, 305(3), 567–580. 1081 https://doi.org/10.1006/jmbi.2000.4315 1082 132. Mistry, J., Chuguransky, S., Williams, L., Qureshi, M., Salazar, G. A., 1083 Sonnhammer, E. L. L., Tosatto, S. C. E., Paladin, L., Raj, S., Richardson, L. J., Finn, R. D., 1084 & Bateman, A. (2021). Pfam: The protein families database in 2021. Nucleic acids 1085 research, 49(D1), D412–D419. https://doi.org/10.1093/nar/gkaa913 1086 133. Letunic, I., Khedkar, S., & Bork, P. (2021). SMART: recent updates, new 1087 developments and status in 2020. Nucleic acids research, 49(D1), D458–D460. 1088 https://doi.org/10.1093/nar/gkaa937 1089 134. Li, W., O'Neill, K. R., Haft, D. H., DiCuccio, M., Chetvernin, V., Badretdin, A., 1090 Coulouris, G., Chitsaz, F., Derbyshire, M. K., Durkin, A. S., Gonzales, N. R., Gwadz, M., 1091 Lanczycki, C. J., Song, J. S., Thanki, N., Wang, J., Yamashita, R. A., Yang, M., Zheng, C., 1092 Marchler-Bauer, A., … Thibaud-Nissen, F. (2021). RefSeq: expanding the Prokaryotic 1093 Genome Annotation Pipeline reach with protein family model curation. Nucleic acids 1094 research, 49(D1), D1020–D1028. https://doi.org/10.1093/nar/gkaa1105 1095 135. Pandurangan, A. P., Stahlhacke, J., Oates, M. E., Smithers, B., & Gough, J. (2019). 1096 The SUPERFAMILY 2.0 database: a significant proteome update and a new webserver. 1097 Nucleic acids research, 47(D1), D490–D494. https://doi.org/10.1093/nar/gky1130 1098 136. Green, E. R., & Mecsas, J. (2016). Bacterial Secretion Systems: An Overview. 1099 Microbiology spectrum, 4(1), 10.1128/microbiolspec.VMBF-0012-2015. 1100 https://doi.org/10.1128/microbiolspec.VMBF-0012-2015 1101 137. Gene Ontology Consortium (2021). The Gene Ontology resource: enriching a GOld 1102 mine. Nucleic acids research, 49(D1), D325–D334. https://doi.org/10.1093/nar/gkaa1113 1103 138. Almagro Armenteros, J. J., Tsirigos, K. D., Sønderby, C. K., Petersen, T. N., 1104 Winther, O., Brunak, S., von Heijne, G., & Nielsen, H. (2019). SignalP 5.0 improves signal 1105 peptide predictions using deep neural networks. Nature biotechnology, 37(4), 420–423. 1106 https://doi.org/10.1038/s41587-019-0036-z 1107 139. Strait, B. J., & Dewey, T. G. (1996). The Shannon information entropy of protein 1108 sequences. Biophysical journal, 71(1), 148–155. https://doi.org/10.1016/S00061109 3495(96)79210-X 1110 140. Freitas, A. I., Domingues, L., & Aguiar, T. Q. (2022). Bare silica as an alternative 1111 matrix for affinity purification/immobilization of His-tagged proteins. Separation and 1112 Purification Technology, 286, 120448. https://doi.org/10.1016/j.seppur.2022.120448 1113 141. Ferretti J, Köhler W. History of Streptococcal Research. 2016 Feb 10. In: Ferretti 1114 JJ, Stevens DL, Fischetti VA, editors. Streptococcus pyogenes : Basic Biology to Clinical 1115 Manifestations [Internet]. Oklahoma City (OK): University of Oklahoma Health Sciences 1116 Center; 2016-. Available from: https://www.ncbi.nlm.nih.gov/books/NBK333430/ 1117 142. Muio, K. D.-R. (2012). Identification of novel up-regulated virulence-associated 1118 proteins in Streptocossus Iniae using quantitative proteomics (1–) [Master thesis, 1119 University of Prince Edward Island]. https://islandscholar.ca/islandora/object/17713 1120 143. Abd, Ashraf & Tawab, El & Hofy, Fatma & Ali, Nadia & Saad, Walaa & El-Mougy, 1121 Emad & Mohammed, Amira. (2022). Antibiotic resistance genes in Streptococcus iniae 1122 isolated from diseased Oreochromis niloticus. Egyptian Journal of Aquatic Biology and 1123 Fisheries. 26. 413-428. https://doi.org/10.21608/EJABF.2022.230602 1124 144. Grant, B. J., Skjaerven, L., & Yao, X. Q. (2021). The Bio3D packages for structural 1125 bioinformatics. Protein science : a publication of the Protein Society, 30(1), 20–30. 1126 https://doi.org/10.1002/pro.3923 1127 145. Holcomb, D. D., Alexaki, A., Katneni, U., & Kimchi-Sarfaty, C. (2019). The 1128 Kazusa codon usage database, CoCoPUTs, and the value of up-to-date codon usage 1129 statistics. Infection, genetics and evolution : journal of molecular epidemiology and 1130 evolutionary genetics in infectious diseases, 73, 266–268. 1131 https://doi.org/10.1016/j.meegid.2019.05.010 1132 1133 Figure legends 1134 Figure 1. Complete genome assembly and functional annotation of Streptococcus iniae strain 1135 SIKU01. (a) Genome assembly and reference-guided workflow from raw reads to validated 1136 circular chromosome. (b) Assembly and annotation statistics, including genome size, GC content, 1137 gene counts, and RNA features. (c) Comparative macrosynteny of SIKU01 against public S. iniae 1138 reference strains, showing conserved genomic architecture. (d) Metadata of SIKU01 and related 1139 isolates used in hybrid reference-guided de novo assembly, including collection date, country, and 1140 host species. (e) Functional annotation of the SIKU01 proteome, including KEGG Mapper 1141 categories, Gene Ontology subcellular localization, and InterProScan domain assignments. 1142 Figure 2. Quality by Design (QbD) lifecycle workflow for antigen selection and downstream 1143 purification strategy. The pipeline starts with defining the Target Product Profile (TPP), including 1144 expression system, delivery route, and antigen type. Candidate purification strategies are then 1145 chosen by evaluating Critical Material Attributes (CMAs), such as tag mechanism, cost, and host 1146 compatibility. Antigens are scored in silico against Critical Quality Attributes (CQAs), including 1147 molecular weight, isoelectric point, codon adaptation index (CAI), hydrophobicity, and predicted 1148 purification behavior. The circular layout highlights the iterative refinement of design spaces, 1149 linking antigen selection with downstream manufacturability. 1150 Figure 3. Biophysical design space filtering of Streptococcus iniae proteins across QbD pre1151 selection, M0, and M1 criteria and expression system comparison. (a) Full proteome (1,855 1152 proteins) before Quality-by-Design (QbD) filtering, plotted in three dimensions by molecular 1153 weight (x-axis), hydrophobicity (y-axis), and isoelectric point (pI, z-axis). Symbols indicate prior 1154 annotation: PubMed only (triangle), epitope only (square), both (dark dot), or neither (light grey). 1155 (b) Pre-M1 subset showing only proteins that pass M0 mandatory inclusion criteria: presence of a 1156 coding sequence with protein product, core gene carriage (≥ 99% of isolates), and full-length open 1157 reading frame. (c) M1 General matrix applying physicochemical scoring gates: protein length 100– 1158 699 aa, Shannon entropy > 0.98, molecular weight 20–60 kDa (+2), PubMed support (+1, 1159 cumulative), and predicted Band T-cell epitopes from IEDB (+3). Points are colored by 1160 supporting evidence as in (a). This stepwise filtering enriches for conserved, structurally stable, 1161 and immunologically supported candidates suitable for downstream antigen selection. 1162 Figure 4. Manufacturability design spaces across purification platforms. (a) Ion-exchange 1163 charge space for protein candidates: pI vs. net charge at pH 7. Shaded bands mark the AEX and 1164 CEX inclusion regions; vaccine-referenced antigens are labeled. (b) Survivors after M2 filtering 1165 by platform; bars show unique genes per route with counts and percentages. (c-d) Buffer operating 1166 pH ranges for AEX (C) and CEX (D). Short tick marks indicate the working pH (range midpoint). 1167 (e-f) Protein-route survivors per buffer evaluated at the midpoint: AEX rule = base gate passed, 1168 z@7 ≤ 0, and pI ≤ pH_mid; CEX rule = base gate passed, z@7 ≥ 0, and pI ≥ pH_mid. (g) Cellulose: 1169 antigen length distribution for M2 survivors; dashed line at 400 aa. (h, k) pDNA (M2) scatterplots 1170 penalty/bonus. Reports score_m1_protein alongside MW, pI, hydrophobicity, CAI_ec, and 1312 instability index. 1313 Table S6. M1 pDNA (zebrafish). Adds the zebrafish-specific CAI contribution (sc_cai_dr) to 1314 score_m1_general, reporting score_m1_pDNA_dr with supporting fields (MW, pI, 1315 hydrophobicity, CAI_dr). 1316 Table S7. M1 pDNA (tilapia). Analogous to S6 but for tilapia CAI (sc_cai_on), reporting 1317 score_m1_pDNA_on with MW, pI, hydrophobicity, CAI_on. 1318 Table S8. M2 Silica (additive route score). Silica feasibility subscores based on pI (7–9) and 1319 charge at pH 7 (≤0), with score_m2_silica_add and total score_m2_silica_total (= M1 + silica 1320 add). 1321 Table S9. M2 Cellulose (additive route score). Cellulose feasibility via low pI (2–4) and shorter 1322 length (<400 aa), with score_m2_cellulose_add and score_m2_cellulose_total. 1323 Table S10. M2 Anion exchange (AEX). AEX feasibility subscores using low pI (≤6) and non1324 positive charge at pH 7, with score_m2_anion_add and score_m2_anion_total. 1325 Table S11. M2 Cation exchange (CEX). CEX feasibility subscores using high pI (≥8) and non1326 negative charge at pH 7, with score_m2_cation_add and score_m2_cation_total. 1327 Table S12. M2 pDNA (zebrafish). DNA-vaccine cloning feasibility subscores for zebrafish: 1328 CAI_dr, GC3 quartiles, zero Type IIS sites, and shorter nucleotide length (≤ median), reporting 1329 score_m2_pDNA_add and score_m2_pDNA_total. 1330 Table S13. M2 pDNA (tilapia). As in S12 but using CAI_on for tilapia; reports 1331 score_m2_pDNA_add and score_m2_pDNA_total. 1332 Table S14. M2 summary (route passes and composite score). For every M1-passing candidate, 1333 lists per-route pass flags (silica, cellulose, anion, cation, size, salt, cloning), the matrix2_count, 1334 and the overall score_composite (= M1 score + number of M2 passes), alongside reference 1335 properties (CAI, arginine content, Type IIS sites, predicted charge at multiple pH values). 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354