scieee AI-readable full text Open interactive document viewer

Complete genome and QbD-guided reverse vaccinology for Streptococcus iniae strain SIKU01

ANDRES, Quentin Ludovic Stephane; Uchuwittayakul, Anurak; Srisapoome, Prapansak

Abstract

🧬 Whole Genome Assembly of 5 Streptococcus iniae bacteria isolated from diseased Asian Seabass, Thailand, comparative genomics, validation of WGS and in-sillico identification of antigens for vaccine biomanufacturing via Quality by Design (QbD). Organism: Streptococcus iniae isolates from diseased farmed Asian seabass (Lates calcarifer) Technologies: Illumina PE (short-reads) Methodologies: Single reference mapping De novo assembly Reference-guided de novo assembly Multi-reference mapping onto pangenome graphs Literature review and functional annotation of S. iniae proteome Identification of protein subset of candidate antigens based on functional annotations Pre-filtering using a QbD approach with a scoring matrix based on physico-chemical properties of Ags Second-filtering using a QbD approach with a scoring matrix based on E. coli expression system Identification of shared epitopes versus IEDB database of B- Cell epitopes in other animals Scoring and final selection of sets of best-scoring antigens for vaccine biomanufacturing using a range of downstream separation methods. NCBI Submission: Bioproject PRJNA933632 GenBank Sequence of SIKU01 Streptococcus iniae GenomeResults 98 proteins suitable for anion-exchange purification, 20 for cation-exchange, 100 for cellulose-affinity, 49 for silica-affinity, and 57-65 for plasmid DNA platforms. Importantly, this approach recovered well-validated antigens including enolase and GAPDH, which showed minimal sequence variation across our global dataset and have demonstrated 62-80% relative percent survival in previous trials. MS_Reverse_Vaccinology_QbD_Streptococcus_iniae_29-10-2025_RV.pdf / .docx — Final peer-reviewed version of the manuscript describing the integration of reverse vaccinology and Quality-by-Design (QbD) for Streptococcus iniae vaccine antigen discovery. MS_Reverse_Vaccinology_Identification_Vaccine_antigens_Streptococcus_iniae_29-10-2025_RV.docx — Supporting version emphasizing antigen discovery pipeline and candidate selection. Supplementary_Informations_MS_Reverse_Vaccinology_QbD_Streptococcus_iniae_29-10-2025.docx / .pdf — Full Supplementary Information including methods, figures, and QbD scoring matrix explanations. Tables_1-3.xlsx — Summary tables of genome statistics, annotation metrics, and antigen scoring results. Supplementary Data Workbooks Supplementary_Data_1_Metadata_and_Proteome.xlsx — Genome metadata, proteome annotation (PGAP), InterProScan results, and initial antigen preselection datasets (S01–S18). Supplementary_Data_2_Pangenomics_and_MSAs.xlsx — Pan-genome presence/absence matrices, multiple sequence alignments, and sequence entropy metrics across 90 S. iniae isolates (S19–S22). Supplementary_Data_3_QbD_Manufacturability.xlsx — Quality-by-Design manufacturability matrices (M0–M2), codon usage, route-specific subscores, and composite vaccine candidate rankings (S23–S37). Figures Figures 1–4: Genome assembly overview, QbD workflow, manufacturability design spaces, and structural epitope mapping. Figures S1–S12: Supplementary visualizations — assembly QC (Circos plots), synteny, antigenic variation, physico-chemical landscapes, filtering stages, and literature-based comparisons of RPS vaccine systems. Analysis Scripts S00–S14 — Custom Python, R, and Bash scripts for genome annotation parsing, UniProt and IEDB mapping, pangenome generation, MSA and entropy computation, and QbD scoring. Examples: S00_gbk_to_table.py — Converts GenBank annotations into tabular format. S03_IEDB_Epitope_Mapping_DIAMOND_SIKU01.py — Performs epitope homology search against the IEDB dataset. S11_Conservation_Shannon_SIKU01.R — Calculates Shannon entropy for conserved core gene alignments. S14_QbD_Ranking_SIKU01.R — Implements QbD-based multi-criteria scoring for antigen manufacturability. Includes auxiliary scripts for Panaroo integration, conservation visualization (ChimeraX), and core genome concatenation. General Data General_Data.zip — Consolidated auxiliary data (reference sequences, KEGG mappings, and intermediate outputs) supporting the analysis pipeline. 🧬 Summary This dataset supports the publication:“Reverse Vaccinology and Quality-by-Design (QbD) Framework for Vaccine Antigen Discovery in Streptococcus iniae”It includes the complete genome and proteome analysis, pangenomic context, Quality-by-Design scoring matrices, and reproducible scripts for candidate antigen prioritization.

Full text

1/15 Supplementary Information Complete genome and QbD-guided reverse vaccinology for Streptococcus iniae strain SIKU01 Andres et al. 2/15 Figure S1. Circular representation of the S. iniae strain SIKU01 compared with four additional Thai isolates (SIKU02-SIKU05). From outer to inner rings: coding sequences on forward strand (green), coding sequences on reverse strand (red), RNA genes (blue), GC content (black), and GC skew (purple/green). Connecting lines in the center represent regions of synteny between genomes. 3/15 Figure S2. Genome alignment and synteny analysis between Streptococcus iniae strain SIKU01 and Streptococcus iniae strain QMA0141 (9117) isolated from Inia geoffrensis Amazon River dolphin (1976). The top and bottom tracks represent the genomes of strains QMA0141 (9117) and SIKU01, respectively. Colored blocks indicate Locally Collinear Blocks (LCBs) conserved between the two genomes, while grey shading links homologous regions. Multiple large-scale inversions and rearrangements are visible, suggesting genome structural divergence over time. The conserved core genome is largely preserved, but structural plasticity and potential horizontal gene transfer events contribute to the observed variation. Scale bar indicates 0.5 Mb. 4/15 Figure S3. Antigenic variation (gene carriage and conservation) across 90 S. iniae genomes for 17 epitope-containing proteins (IEDB) with percent amino acid identity. Amino acid identity was calculated for each candidate antigen using reciprocal DIAMOND BlastP and multiple sequence alignment via MAFFT. The heatmap displays pairwise amino acid identity (%) of each antigen across 90 S. iniae genomes, with darker shades indicating higher conservation. Gene carriage analysis was performed to assess the presence or absence of each target protein, revealing varying levels of sequence conservation across isolates. Several targets, including GAPDH, Enolase, and 60 kDa chaperonin, showed >98% conservation across nearly all genomes, while others such as LPXTG-anchored and metal substrate-binding proteins exhibited greater variability. Proteins marked as "SIKU01_XXXXXX" represent strain-specific open reading frames predicted in SIKU01. Sortase A and Trigger Factor (TF) also showed low conservation in some isolates, suggesting variable antigenicity within the population. 5/15 M0 biophysical landscape Across the full S. iniae SIKU01 proteome (M0; N = 1,855), the distributions of Quality Attributes (CQAs) for CDS coding sequences (hydrophobicity (GRAVY), aliphatic index, instability index (II), isoelectric point (pI), molecular weight (MW), length (nt or/ aa), net charge (z) at pH 7, GC3 fraction, arginine content, and binding potential) and their correlations were well resolved (Supplementary Fig. S4). Hydrophobicity centered near neutrality (GRAVY mean −0.13 (SD 0.46) and median −0.22), displaying a biphasic pattern consistent with a mixed population of soluble and membrane proteins. Isoelectric points were similarly bimodal (pI mean 7.03 (SD 2.21) and median 6.51), reflecting acidic and basic proteome subpopulations (Supplementary Fig. S4a). Protein length spanned broadly (mean 320.9 aa, SD 207.0; median 281 aa), while molecular weight was similarly variable (MW mean 35.4 kDa, SD 20.5; median 31.4 kDa). Instability index was generally low-to-moderate (II mean 34.6 (SD 10.2); median 33.5). Composition-linked variables were narrow: GC3 fraction mean 0.28 (SD 0.05); median 0.28, arginine content mean 0.04 (SD 0.02); median 0.04. Net charge at pH 7 clustered near zero (z mean −2.58 (SD 15.4); median −1.78), with the expected tight coupling to pI (Supplementary Fig. S4a). Pairwise structure recapitulated known trends and revealed coherent clusters (Supplementary Fig. S4b-g). Hydrophobicity vs. aliphatic index showed a strong positive association (ρ ≈ 0.82, p < 10⁻¹⁵; Supplementary Fig. S4d), whereas hydrophobicity vs. binding potential was strongly negative (Supplementary Fig. S4b). Aliphatic index vs. binding potential was negatively sloped (Supplementary Fig. S4e), and arginine content increased with binding potential but decreased with hydrophobicity (Supplementary Fig. S4f-g). As expected, length and molecular weight (MW) were nearly collinear (ρ ≈ 1.00), and pI vs. net charge (z) at pH 7 was tightly positive (ρ ≈ 0.93). Amino-acid usage frequencies for the M0 proteome are provided in 6/15 Supplementary Fig. S5, and a full Spearman correlation matrix summarizing these relationships is shown in Supplementary Fig. S6. Figure S4. Biophysical landscape of the proteome of S. iniae strain SIKU01 (M0). (a) Distributions of intrinsic features across all 1,855 proteins, including aliphatic index, arginine content, binding potential, GC3 fraction, hydrophobicity (GRAVY), instability index (II), isoelectric point (pI), protein length (aa), molecular weight (MW), and net charge (z) at pH 7. Mean, median, and variance are shown in each facet. (b–g) Selected pairwise relationships highlighting strong or structured correlations: hydrophobicity vs. binding potential (b), pI vs. net charge at pH 7 (c), hydrophobicity vs. aliphatic index (d), binding potential vs. aliphatic index (e), binding potential vs. arginine content (f), and hydrophobicity vs. arginine content (g). Points are overlaid with density contours. 7/15 Supplementary Note 1 — Virulence-associated protein subset of S. iniae SIKU01 Based on InterPro domain architectures, Gene Ontology (GO) functional annotations, and gene name-based inference, a total of 283 proteins were identified as belonging to Critical Quality Attribute (CQA) categories relevant to Streptococcus iniae virulence (Supplementary Table S14). Functional annotation was derived from curated bioinformatics databases (InterPro, UniProt, GOA, and VFDB) and complemented by expert-driven literature validation using PubMed, ensuring both computational and knowledge-based evidence. The functional risk space indicated that nutrient acquisition systems represented the most prevalent CQA group (> 100 proteins), followed by secretion systems (~ 25 proteins), immune evasion factors (~ 30 proteins), and adhesins (~ 25 proteins). Additional groups included toxins and hemolysins (e.g., TOMMassociated or pore-forming domains; ~ 6 proteins), capsular biosynthesis–related enzymes (~ 50 proteins), and infection-associated hydrolases such as nucleases, S8 family proteases, and other peptidases (~ 50 proteins). Subcellular localization analyses of these 283 VFs further identified 31 proteins with Nterminal signal peptides (Supplementary Table S05), 164 proteins containing predicted transmembrane domains (Supplementary Table S06), and 112 predicted cytoplasmic antigen candidates (Supplementary Table S07). Together, these results confirm that strain SIKU01 retains a complete set of housekeeping functions while harboring specialized virulence systems commonly found in other Streptococcus iniae genomes. 8/15 Figure S5. Summary of functional and physicochemical properties of virulence factors (VFs) identified in Streptococcus iniae strain SIKU01. Top panels show distributions of VF types, protein lengths, molecular weights, and isoelectric points. Middle panels illustrate relationships between physicochemical parameters. Bottom panel shows AA composition of predicted VFs. 9/15 Figure S6. Spearman correlation matrix of biophysical features in the Streptococcus iniae SIKU01 proteome (M0). Pairwise ρ values, scatterplots, and density overlays illustrate relationships among intrinsic properties, including aliphatic index, binding potential, GC₃ fraction, hydrophobicity, instability index (II), isoelectric point (pI), arginine content, molecular weight (MW), and net charge (z). Hydrophobicity and aliphatic index show a strong positive correlation (ρ = 0.82), while hydrophobicity is inversely correlated with binding potential (ρ = −0.94) and II (ρ = −0.32). Other associations, such as GC₃ fraction with pI, are weaker but significant.