scieee AI-readable full text Open interactive document viewer

Genomic studies of piscine streptococci and genome assembly of Clariid catfishes (Siluriformes) for aquaculture enhancement

Andres, Quentin Ludovic Stephane

Abstract

Abstract This thesis presents integrated genomic research aimed at enhancing aquaculture productivity through two main contributions: (1) the generation of high-quality, haplotype-resolved genome assemblies of Clarias macrocephalus (bighead catfish) and its F1 hybrid with Clarias gariepinus, and (2) the genomic analysis of Streptococcus iniae, a bacterial pathogen causing severe losses in farmed fish. Using long-read (PacBio HiFi, Oxford Nanopore), short-read (Illumina), and proximity-ligation (Hi-C) sequencing, we assembled chromosome-scale genomes for both catfish taxa, enabling improved trait analysis, hybrid assessment, and comparative transposable element studies in Siluriformes. In parallel (3), five S. iniae isolates were sequenced and analyzed using pangenomics, variant calling, recombination detection, and reverse vaccinology pipelines. This work contributes essential genomic resources for selective breeding and disease control, promoting more sustainable and genetically informed aquaculture practices.

Full text

THESIS PROPOSAL GENOMIC STUDIES OF PISCINE STREPTOCOCCI AND GENOME ASSEMBLY OF CLARIIDAE CATFISHES (SILURIFORMES) FOR AQUACULTURE ENHANCEMENT QUENTIN LUDOVIC STEPHANE ANDRES GRADUATE SCHOOL KASETSART UNIVERSITY 2025 THESIS PROPOSAL APPROVAL GRADUATE SCHOOL KASETSART UNIVERSITY DEGREE Doctor of Philosophy (Fishery Science and Technology) MAJOR FIELD Fishery Science and Technology FACULTY Fisheries TITLE Genomic studies of piscine streptococci and genome assembly of Clariidae catfishes (Siluriformes) for aquaculture enhancement NAME MR. QUENTIN LUDOVIC STEPHANE ANDRES THIS THESIS HAS BEEN ACCEPTED BY (Associate Professor Prapansak Srisapoome, Ph.D.) THESIS ADVISOR (Professor Kornsorn Srikulnath, Ph.D.) THESIS CO-ADVISOR (Mr. Worapong Singchat, Ph.D.) THESIS CO-ADVISOR (Assistant Professor Methee Kaewnern, D.Tech.Sc.) GRADUATE COMMITTEE CHAIRMAN (Associate Professor Weeraphart Khunrattanasiri, Dr.rer.nat.) DEAN THESIS PROPOSAL Genomic studies of piscine streptococci and genome assembly of Clariidae catfishes (Siluriformes) for aquaculture enhancement QUENTIN LUDOVIC STEPHANE ANDRES A Thesis Proposal Submitted in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy (Fishery Science and Technology) Graduate School, Kasetsart University Academic Year 2025 1 Contents Page LIST OF TABLES 3 LIST OF FIGURES 5 LIST OF ACRONYMS AND GLOSSARY 5 INTRODUCTION 1 OBJECTIVES AND RATIONALE 3 LITERATURE REVIEW 4 Genomic Technologies for Aquaculture 4 Clarias Catfish Aquaculture in Southeast Asia 10 Case Study: Genome Assembly Applications in Tilapia 11 GENOME ASSEMBLY – BIGHEAD CATFISH 12 Overview of Genome Assembly Strategy 12 Haplotype Resolution, Hi-C Scaffolding, and Phasing 18 GENOME ASSEMBLY – F1 HYBRID CATFISH 32 Overview of Genome Assembly Strategy 32 VALIDATION OF CATFISH ASSEMBLIES 44 Benchmarking Methodology 44 Results – Bighead Catfish 47 Results – F1 Hybrid Catfish 61 GENOME ASSEMBLY, REVERSE VACCINOLOGY, AND QUALITY BY DESIGN — Streptococcus iniae 77 Introduction 77 Methods 80 Results 96 DATA AVAILABILITY AND NCBI SUBMISSIONS 110 Bighead Catfish 110 F1 Hybrid Catfish 110 2 Streptococcus iniae 111 Associated Publications 111 DISCUSSION AND CONCLUSIONS 112 Genomic Insights and Technical Achievements 112 Streptococcus iniae Vaccine Development 113 Overall Conclusions 114 Recommendations 115 Appendix 138 Personal Information 161 3 List of Tables Table Page 1 Bighead Catfish – Assembly Stats. 48 2 Bighead Catfish – Scaffold Metrics – Haplotype 1. 55 3 Bighead Catfish – Scaffold Metrics – Haplotype 2. 56 4 Bighead Catfish – TE Content. 58 5 F1 Hybrid Catfish – North African Subgenome – Scaffold Metrics. 72 6 F1 Hybrid Catfish – Bighead Subgenome – Scaffold Metrics. 73 7 F1 Hybrid Catfish – North African Subgenome – Structural Metrics. 75 8 F1 Hybrid Catfish – Bighead Subgenome – Structural Metrics. 76 9S. iniae RV – Predicted antigenic epitopes from strain SIKU01. 99 10 S. iniae QbD – QTPP and CQAs definition workflow. 153 11 S. iniae – Compiled vaccine studies. 154 12 S. iniae – Supplementary Code for analysis. 157 13 S. iniae – Supplementary Data 1: Metadata and Proteome. 158 14 S. iniae – Supplementary Data 2: Pangenomics and MSAs. 159 15 S. iniae – Supplementary Data 3: QbD Manufacturability. 160 LIST OF FIGURES Figure Page 1 Long-reads vs Short-reads 4 2 ONT sequencing 5 3 HiFi Sequencing 6 4 Illumina PE Sequencing 7 5 Hi-C Sequencing 8 6 Clarias catfish species 10 7 Bighead Catfish – Specimen Photograph. 16 8 Bighead Catfish – Genome Assembly Workflow. 17 9 Bighead Catfish – Genome Assembly Graph – Haplotype 1 and 2. 20 10 Bighead Catfish – Genome Assembly 1-Year Progress. 27 11 Bighead Catfish – mtDNA Alignment Against 208 Catfish Species. 29 12 F1 Hybrid Catfish – Genome Assembly Workflow. 37 13 F1 Hybrid Catfish – Assembly validation using IGV. 42 14 F1 Hybrid Catfish – Genome Assembly 1-Year Progress. 43 15 Bighead Catfish – DNA Sequencing. 47 16 Bighead Catfish – Hi-C Heatmap – Haplotype 1. 50 17 Bighead Catfish – Hi-C Heatmap – Haplotype 2. 51 18 Bighead Catfish – Hi-C Heatmaps – Haplotype 1 and 2. 52 19 Bighead Catfish – Assembly Completeness Assessment. 54 20 Bighead Catfish – TE Divergence Profile. 57 21 Four Species Macrosynteny – Catfishes Chromosomes. 60 22 F1 Hybrid Catfish – Sequencing and GenomeScope2.0 Survey. 61 23 F1 Hybrid Catfish – North African Subgenome – Hi-C Heatmap. 67 24 F1 Hybrid Catfish – Bighead Subgenome – Hi-C Heatmap. 68 25 F1 Hybrid Catfish – Genome Assembly Validation. 70 26 Streptococcus iniae - Phase contrast micrograph 77 27 S. iniae – Genome assembly and annotation. 97 4 5 28 S. iniae QbD – Workflow QTPP and manufacturability. 102 28 S. iniae QbD – Manufacturability design spaces. 107 29 S. iniae RV – 3D Structure of Enolase and GAPDH Immunogens. 109 30 Four Catfish Genome Survey – GenomeScope2.0 Profiles. 139 31 C. gariepinus – Synteny vs F1 hybrid genome. 140 32 C. macrocephalus – Synteny vs F1 hybrid genome. 141 33 S. iniae – Circos synteny & QC (SIKU01–05). 142 34 S. iniae – Synteny vs dolphin isolate. 143 35 S. iniae – Antigenic variation (17 epitopes). 144 36 S. iniae – Proteome landscape (M0). 145 37 S. iniae – Proteome annotation. 146 38 S. iniae – QbD CQAs correlation matrix. 147 39 S. iniae – QbD filtering workflow (M0–PreM1–M1-general). 148 40 S. iniae – Expression and pDNA platform filters. 149 41 S. iniae – Vaccine type and RPS distribution in literature. 150 42 S. iniae – RPS distribution by fish host species. 151 43 S. iniae – Comparison of pDNA and protein vaccine RPS. 152 1 GENOMIC STUDIES OF PISCINE STREPTOCOCCI AND GENOME ASSEMBLY OF CLARIID CATFISHES (SILURIFORMES) FOR AQUACULTURE ENHANCEMENT INTRODUCTION Global aquaculture is the process of farming aquatic organisms, both marine and freshwater, for human consumption and other uses, such as producing feed, pharmaceuticals, and related products. It represents the fastest-growing food production sector, providing over half of the seafood consumed globally, with Asia dominating production. In Thailand, inland aquaculture is critical for global food security, where freshwater species such as tilapia (Oreochromis spp.), catfish (Clarias spp., Silurus spp., and Pangasius spp.), and Asian seabass (Lates calcarifer) contribute significantly to domestic consumption and seafood exports (FAO,2020;Naylor et al.,2021). In 2022, aquaculture surpassed capture fisheries as the primary source of aquatic animals, producing 94.4 million tonnes compared to 92.3 million tonnes from wildcapture fisheries, according to the Global Seafood Alliance. Globally, aquaculture production (including seaweeds) reached 130.9 million tonnes in 2022, valued at USD 313 billion, according to the Food and Agriculture Organization. Major gains in aquaculture feed efficiency and fish nutrition have lowered the fish-in–fish-out ratio for many fed species, although dependence on marine ingredients persists and reliance on terrestrial ingredients has increased. To ensure sustainable growth, genetic improvement through selective breeding has become a key strategy, where optimizing traits such as growth, feed efficiency, sex determination, robustness, and disease resistance requires identifying and deploying relevant genetic variants. Modern aquaculture breeding programs do not rely solely on conserving “pure” strains or extreme inbreeding. Instead, they are typically structured around families or lines that serve as a defined starting point and are improved cumulatively over gener- 8 2. Hi-C for Genome Scaffolding Hi-C captures chromosome three-dimensional organization through proximityligation, where physically close DNA fragments become joined. Paired-end sequencing of these junctions reveals long-range linkage: contigs from the same chromosome show many Hi-C contacts while those from different chromosomes show few (Putnam et al., 2016). This enables clustering, ordering, and orienting contigs into chromosome-scale scaffolds for de novo assembly—construction without a reference genome—rather than reference-guided assembly which requires an existing genome. Hi-C has become essential for achieving chromosome-level assemblies in aquaculture species, enabling precise localization of genes and QTLs for breeding traits. In polyploid fish genomes, Hi-C correctly separates duplicated loci and provides insights into 3D genome structure, chromatin conformation, and regulatory interactions (Dudchenko et al.,2017). Figure 5 Ultra-long-range sequencing for genome scaffolding Credits: Ultra-long-range Sequencing for Genome Scaffolding 3. Haplogenome Assembly of Catfish Genomes Haplogenome assembly reconstructs each parental haplotype independently instead of producing a collapsed consensus sequence. This captures the full spectrum of 9 allelic and structural variation, valuable for outbred species and highly heterozygous organisms such as interspecific catfish hybrids. By resolving phase-specific sequences, haplogenome assemblies reveal allele-specific expression patterns critical for understanding gene function and optimizing selective breeding. The approach typically uses long-read sequencing (PacBio HiFi or ONT) with bioinformatic phasing to separate maternal and paternal haplotypes. A well-known example includes the assembly of haplogenomes in interspecific Eucalyptus hybrids, analogous to the genome-wide structural differences observed between Silurus asotus and S. meridionalis that are linked to hybrid traits Chen et al. (2021a). Similarly, in catfishes, haplogenome assemblies of Clarias macrocephalus,C. gariepinus, and their F1 hybrids enable precise characterization of chromosomal rearrangements, gene duplications, and introgression events—laying the foundation for genomic selection and improved aquaculture productivity Duong and Scribner (2018). 4. In Summary These sequencing platforms capture genome structure and identify variants— SNPs1, indels2, and structural variants (SVs)3—associated with economically important traits (Chaivichoo et al.,2020). Short reads excel at accurate SNP/indel detection, long reads resolve structural variants and complex regions, while Hi-C provides chromosomal-level scaffolding. As of May 2025, NCBI hosts over 3.01 million genomes including 51,820 eukaryotic genomes, with fewer than 500 being haplotype-resolved, complete vertebrate genomes (https://www.ncbi.nlm.nih.gov/datasets/genome/). 1Single-base changes at specific genomic positions, used as genetic markers. 2Short insertions or deletions of 1–50 bp. 3Large genomic alterations >50 bp including deletions, insertions, duplications, inversions, and translocations. 10 Clarias Catfish Aquaculture in Southeast Asia 1. Evolutionary History and Distribution Catfish (Siluriformes) are widely cultivated freshwater species crucial for food production across Africa, the Americas, and Asia (Lisachov et al.,2023). This diverse order diverged during the early Cretaceous ( 145 Ma) and comprises over 36 families and 3,000 species (Ferraris,2007). While distributed globally in freshwater environments with some marine extensions, catfish are most diverse in tropical South America, Asia, and Africa. Key representatives like Clarias and Ictalurus serve as both research models and economically valuable aquaculture species. 2. Classification and Biology Bighead catfish (Clarias macrocephalus) and North African catfish (C. gariepinus) belong to family Clariidae, order Siluriformes. These air-breathing catfish are economically important in Southeast Asia and Africa, respectively. C. macrocephalus, native to Southeast Asia and widely farmed in Thailand, has a broad flattened head, short barbels, and reaches 40–50 cm. Its body is dark brown to black with pale blotches. C. gariepinus, introduced globally for its fast growth and environmental tolerance, grows to 1.7 m with an elongated body and long dorsal fins. It survives oxygen-depleted waters using its suprabranchial organ for atmospheric breathing. Figure 6 Left: Female C. macrocephalus (Credits: Matillano, J.D.). Right: C. gariepinus (Credits: Larsen, J.H.) 11 F1 hybrids from C. gariepinus males × C. macrocephalus females combine parental advantages: body form from C. macrocephalus with rapid growth from C. gariepinus. Both species thrive at 26–32°C. 3. Role in Food Production and Economy Clarias species are extensively farmed across Asia and Africa. C. macrocephalus is primarily used for hybrid production with C. gariepinus (Lisachov et al.,2023;NaNakorn et al.,2004), yielding sterile F1 hybrids with hybrid vigor (heterosis) that combine rapid growth and desirable body characteristics. However, hybrid sterility necessitates separate parental production systems. The lack of high-quality reference genomes limits implementation of modern breeding tools like ddRAD sequencing, low-coverage genome sequencing, and CRISPR/Cas9. Studies indicate low introgression risk in C. macrocephalus, supporting potential genetic improvement efforts. Case Study: Genome Assembly Applications in Tilapia Tilapia (Oreochromis spp.) exemplifies successful genomic applications in aquaculture (Yu et al.,2022). Researchers identified adaptive responses4to salinity stress through genome-wide SNP analysis linked to osmoregulation and survival (Gu et al., 2018;Jiang et al.,2019;Rengmark et al.,2007). Key tolerance genes include prolactin (PRL) (Breves et al.,2013), growth hormone (GH1) (Deane and Woo,2008), insulinlike growth factor 1 (IGF1) (Yan et al.,2020), and plasma-membrane Ca2+-ATPases (PMCAs) (Rengmark et al.,2007). Similar studies in rainbow trout (Le Bras et al., 2011), Atlantic salmon (Norman et al.,2012), and stickleback (Kusakabe et al.,2016) collectively established foundations for marker-assisted selection (MAS) in aquaculture. Non-synonymous mutations in salt regulation genes enhance tolerance and alter phenotypes. These achievements required high-quality reference genome assemblies. 4Biological adjustments to environmental stress through gene regulation, metabolic changes, or immune modulation. 12 GENOME ASSEMBLY – BIGHEAD CATFISH Overview of Genome Assembly Strategy The haplotype-resolved genome of Clarias macrocephalus was assembled using an hybrid sequencing strategy comprising four DNA sequencing platforms generating various types of DNA reads (i.e., in DNA sequencing, a read is an inferred sequence of base pairs (or base pair probabilities) corresponding to all or part of a single DNA fragment). The four DNA sequencing technologies used to read the gDNA of bighead catfish into DNA reads (short-reads and long-reads) were: PacBio HiFi CSS (Wenger et al., 2019) (Circular consensus sequencing5, long-reads of base accuracy 99.9%), used for maximizing genome assembly quality. Oxford Nanopore (ONT) noisy 1D long-reads (20%.err)6(Jain et al.,2018b) for increasing assembly continuity and to span complex repeat regions of the genome. Proximity-ligation (Hi-C) (Dudchenko et al.,2017; Putnam et al.,2016) data for long-range linking information to phase (i.e., the specific arrangement of alleles on the same chromosome to distinguish maternal and paternal haplotypes) and scaffold contigs (i.e., separating haplotypes). These complementary data types facilitated haplotype resolution and scaffold anchoring and allow to generate dual assemblies7from a single diploid tissue. Among the tested assemblers (Hifiasm (Cheng et al.,2021), Flye (Kolmogorov et al.,2020), wtdbg2 (Ruan and Li,2019)), I selected Hifiasm for the haplotype-resolved assembly because of its ability to separate parental haplotypes (phasing) on low genome coverage (< 13X) HiFi data. In contrast I found that Flye and wtdbg2 produced collapsed consensus assemblies. Below, I outline the main steps for the complete genome assembly of bighead catfish Haplotype 1 and Haplotype 2. 5Circular consensus sequencing (CCS) is a high-accuracy long-read sequencing method, commonly used with PacBio HiFi reads, where multiple passes of the same DNA molecule are combined to generate a consensus sequence. 6Oxford Nanopore 1D reads are single-strand long reads that historically exhibit high error rates, often around 10–20%, due to insertions, deletions, and base-calling inaccuracies. These reads are valuable for spanning long genomic regions but require polishing for accuracy. 7Dual assemblies refer to separate genome assemblies generated for each haplotype in a diploid organism, allowing resolution of heterozygous regions and structural differences between homologous chromosomes. 13 1. Genomic DNA Sequencing, Genome Survey, and k-mer Meryl Databases Genomic DNA was extracted from the liver tissue of a single male bighead catfish individual using a custom protocol (Supikamolseni et al.,2015). Sex was determined based on histology and external gonadal examination (Kitano et al.,2007; Wyneken et al.,2007). Sequencing was performed using four complementary platforms generating four synergistic raw sequencing data types: (i) Pacific Biosciences (PacBio) High-Fidelity (HiFi) sequencing (Wenger et al.,2019) for generating highly accurate contigs, (ii) Oxford Nanopore Technology (ONT) longread sequencing (Jain et al.,2018b) for scaffolding and resolving repetitive regions, and (iii and iv) Illumina short-read 150 base paired-end sequencing (Hi-C, and standard non Hi-C) (Bentley et al.,2008) for chromosomal phasing and error correction, respectively. Genome survey aims in assessing genome characteristics (genome size, heterozygosity, repeat content) from raw sequencing data (e.g., Illumina short-reads) using k-mer analysis in Jellyfish (Marçais and Kingsford,2011) and GenomeScope2.0 (Ranallo-Benavidez et al.,2020). Meryl databases (N=2) for 21-mer and 31-mer (bighead.hifi.illuminaPCRfree.gt1.db.meryl) hybrid k-mer databases, made from combined Illumina and HiFi reads using Meryl (Rhie et al.,2020), as explained in the T2T-polish GitHub repository (https:// github.com/arangrhie/T2T-Polish) (Rhie et al.,2022). K-mer Meryl databases use 21-mer and 31-mer DNA strings to support quality evaluation and error correction in genome assemblies. The shorter 21-mer spectra provide high sensitivity for detecting potential errors, while the longer 31-mer spectra offer higher specificity, validating true variants and minimizing false positives. During genome polishing, these k-mer databases enhance effective coverage in low-depth HiFi regions by leveraging k-mer spectra to guide targeted corrections and improve assembly accuracy. 2. Initial Assembly Generation (Contigguing and Phasing): Contigging is the process of assembling overlapping DNA reads into continuous sequences without gaps (i.e., contigs). These contigs are then separated and assigned to a phase group using proximity-ligation data (Hi-C) based on their parental origin and spacial physical distance in the cell nucleus. Therefore, a contig that comes from a 14 haplotype is referred to as a haplotig8, and in a diploid phased assembly there are two groups of homologous haplotigs (i.e., unitigs, one per haploid ”homotype” haplotype). Here, Hifiasm was used in ”Hi-C UL” mode (Cheng et al.,2021) to generate a haplotype-resolved assembly using Hi-C + ONT + HiFi reads (Figure 8D, left), followed by GreenHill (Ouchi et al.,2023) using all data types: Hi-C + Illumina + ONT + HiFi reads for further scaffolding and phasing of the haplotigs (Figure 8E), and finally wtdbg2 (Hu et al.,2024;Ruan and Li,2019) was employed with HiFi and Illumina short-reads correcting small errors at the nucleotide level (i.e., SNPs) while preventing switch errors (i.e., allele switching between homologous haplotypes). This produced two phased contig sets representing homologous haplotypes (Figure 9). 3. Hi-C Scaffolding and Their Validation: Scaffolds are computationally ordered and oriented arrays of contigs that have sequence gaps along their length. The process of scaffolding of contigs is about arranging and orienting contigs into larger structures, sometimes with gaps, using additional data like proximity ligation (Hi-C) in order to reconstruct an in-silico equivalent to the karyotype (referred to as ”Scaffotype”) which consists of Hi-C scaffolds of ”pseudochromosomes” (i.e., chromosome pseudomolecules), visualized as Hi-C maps. Prior to Hi-C scaffolding, it is possible to break the contigs at erroneous sites, for that, misassembly corrections are done using CRAQ (Li et al.,2023), then Hi-C scaffolds (i.e., contigs and scaffolds) are flipped, reoriented, and refined to improve structural accuracy, this step is knows as the ”manual review step9” step and is performed using JBAT (JuiceBox Assembly Tools) (Dudchenko et al.,2017), next, a ”post-review step10 ” step seals contigs (i.e., by creating gaps between newly adjacent contigs and inserting 500 ’N’ characters to represent unknown 8In an unphased assembly, a contig may join alleles from different parental haplotypes in a diploid or polyploid genome (see (Lewin et al.,2019) and https://lh3.github.io/2021/04/17/ concepts-in-phased-assemblies). The process of separating sequences based on their parental haplotype in diploid genomes is referred to as ”haplotype phasing”. 9In Hi-C scaffolding, manual review involves visually inspecting contact maps (e.g., using Juicebox or HiGlass) to detect misassemblies, orientation errors, or misplaced contigs that automated tools may miss. 10Post-review refers to the stage after Hi-C contact map inspection, where validated corrections (e.g., flips, cuts, or joins) are applied to the genome assembly. 15 sequences) to form Hi-C scaffolds to ultimately reconstruct pseudochromosomes followed by 3D-DNA ”post-review” validation (Dudchenko et al.,2017) (Figure 8F). Finally more gaps were closed with TGS-GapCloser (Xu et al.,2020) (Figure 8G). These steps — including CRAQ break detection, Hi-C read alignment, manual review in JBAT, and post-review in the 3D-DNA pipeline — were iteratively performed three times. As illustrated in Figure 8F, this cycle can be repeated if needed after contig polishing (Figure 8H), and typically results in well-defined Hi-C scaffolded pseudochromosomes (Figure 8C). 4. Assembly Polishing of Contigs and Benchmarking: Polishing a genome assembly consist in correcting sequencing errors using additional data (Meryl k-mer databases, long-reads and short-reads). To further improve assembly accuracy, reads were aligned to the polished assembly, the three primary long-read mappers employed were Winnowmap (Jain et al.,2022), Minimap2 (Li,2018), and VerityMap (Mikheenko et al.,2020) (previously known as TandemTools), these were used in assembly validation only (i.e., not for variant calls and genome polishing). Secondary alignments (i.e., reads flagged with bit 0x10011 ), low-quality regions (i.e., MAPQ12 = 0), and hard-clipped reads (i.e., alignments with hard clipping operations indicated by ‘H’ in the CIGAR string13 ) were excluded from analysis using samtools (Li and Durbin,2009). This read realignment step was crucial for validating assembly completeness and reducing errors. Non-haplotypeaware tools14 (e.g., Racon (Vaser et al.,2017), Clair3 (Zheng et al.,2022)) were applied cautiously to minimize errors from parental haplotype switches, while Merfin (Formenti et al.,2022) filtered edits for polishing and BCFtools consensus produced the final polished sequence. For large variants and closing gaps using read alignments I used Sniffles2 (Smolka et al.,2024) for large structural 11In SAM/BAM flags, bit 0x100 indicates a secondary alignment. Tools like Picard use this flag to mark reads that are not the primary alignment of a read with multiple mappings. 12MAPQ (Mapping Quality) is a Phred-scaled score that reflects the confidence in the read’s alignment position. Higher values indicate more reliable mappings. 13The CIGAR (Compact Idiosyncratic Gapped Alignment Report) string encodes how a read aligns to a reference, using operations such as matches (M), insertions (I), deletions (D), soft clips (S), and others. For example, 100M indicates 100 matching bases. 14Non-haplotype-aware tools do not distinguish between maternal and paternal sequences in diploid or polyploid genomes, often collapsing allelic variants into a single consensus sequence and potentially masking structural differences between haplotypes. 16 variant (SV) calling, specifically for insertions (INS) and deletions (DEL). I evaluate genome quality in term of errors found in the assembly that are not found in raw reads, for that I calculate the Quality Value, usually refered to as QV15 of a genome assembly where QV refers to the Phred quality score (or quality (Q) score), an integer representing the estimated probability of an error (i.e., that a base is incorrect). The polishing workflow implemented here for correcting bighead catfish ensured high nucleotide accuracy Quality Values (QV) (from initial QV33 to around QV46) after Benchmarking against recommended assembly metrics, from the VGP paper (Rhie et al.,2021). Figure 7 Representative specimens of the bighead catfish (Clarias macrocephalus), collected for high-quality haplotype-resolved genome assembly. The individuals were reared under controlled aquaculture conditions prior to tissue sampling. C. macrocephalus is an economically important freshwater species native to Southeast Asia, widely cultured in Thailand for selective breeding and genomic improvement programs. Methodological details are presented in the following sections and in Figure 8. 15Quality Value (QV) is a Phred-scaled score used to represent base-level or assembly-level accuracy. If Pis the probability of error, then Q=−10log10(P), or equivalently, P=10−Q/10. Higher QV indicates greater accuracy. 17 Figure 8 Comprehensive haplotype-resolved genome assembly and scaffolding workflow for bighead catfish (Clarias macrocephalus). (A) Hi-C contact map (Juicebox v2.16.0). (B) Genome assembly using PacBio HiFi, ONT, and Hi-C reads with Hifiasm (Hi-C UL mode) and Flye for consensus. (C) Hi-C scaffolding via GreenHill and 3D-DNA. (D) Manual post-review with JBAT. (E) Gap filling using TGS-GapCloser and QuarTeT GapFiller. (F) Genome polishing with NextPolish2 and CRAQ. (G) Assembly QV validation with Merqury, Pilon, BCFtools, and VerityMap. (H) Mitochondrial genome assembly with MitoHiFi and Minimap2. The pipeline yields two high-quality phased haplotype assemblies representing the complete bighead catfish genome. 24 3. Running pilon and calling the consensus: Pilon was employed on assemblies for correcting SNPs, INDELs (i.e., 1-2 base insertions and deletions), and other base-level errors. In individual haplotypes, each scaffold was processed by Pilon with parameters (’--genome,--frags,--bam,--targets,--fix all, --vcf, --diploid,--minmq 30,--minqual 30’) to use the alignments from all data types (HiFi, ONT, and short reads). The output Variant Calling Files (VCF) of containing the detected variants were sorted by position using (’bcftools sort’), compressed using bgzip, and indexed using (bcftools index)28 , and final consensus sequences were generated for each scaffold using (’bcftools consensus -f $genome.fa -H 1’). The overall median quality value metric increased from 41 to approximately 45-47 after haplotype-aware targeted assembly polishing. 5. Hi-C Scaffolding with HapHiC To reintegrate unplaced scaffolds in each haplotype in the pseudochromosomes, I used Quartet and HapHiC. First, I aligned the unplaced contigs to reference genome C. fuscus (GCA_030347435.1), and reference genome C. gariepinus (GCA_024256425.2) using Quartet AssemblyMapper (Lin et al.,2023) v1.2.1 with the following parameters (’-r $reference -q $contigs -c 50000 -l 2000 -i 90 -a minimap2’). I identified 53 MB and 33 MB of unplaced scaffold sequences in bighead catfish Haplotype 1 and Haplotype 2, respectively, with strong homology to C. fuscus pseudochromosomes, next I filtered bighead catfish Haplotype 1 and Haplotype 2 separately with SeqKit (Shen et al.,2016) v0.8.1 to retain pseudochromosomes and concatenated them to unplaced scaffolds. For the preparation of Hi-C scaffolding step, Hi-C reads were mapped to separate haplotypes using BWA-MEM (’-5SP’) after making a BWT index29 28bcftools index creates a compressed index (.csi or .tbi) for VCF or BCF files, allowing rapid random access to variants by genomic position during downstream processing or visualization. 29The BWT index refers to a data structure derived from the Burrows-Wheeler Transform (BWT), which allows fast and memory-efficient alignment of sequencing reads to a reference genome. It compresses the genome while retaining the ability to search for exact or approximate matches, enabling tools like Bowtie2 and BWA to align short reads with high performance. 25 for Haplotype 1 and Haplotype 2 (’bwa index $genome.fasta’) and Hi-C read alignments filtered for PCR duplicates30 and secondary alignments (’samblaster $BAM | samtools view - -@ $threads -S -h -b -F 3340’) using the software Samblaster (Faust and Hall,2014) v0.1.26. Hi-C scaffolding was performed on bighead catfish haplotypes using HapHic (Zeng et al.,2024) v1.0. The resulting Hi-C maps were visualized in JBAT and using (’haphic plot’) for separate scaffolds from each haplotype (Figure 16 and Figure 17) but also with all scaffolds represented in a wide-map (Figure 18), and a post-review of the Hi-C scaffold was performed as described in the previous section. Finally, three rounds of TGS-GapCloser followed by one round of targeted Pilon polishing specifying the new targets resulted in a genome of global higher quality (i.e., both in term of QV quality values and CRAQ’s metrics of CRE/CSE structural accuracy). I addressed the duplication errors31 visible in the heterozygous peak of the k-mer coverage (i.e., from the Merqury spectra CN) with more or less success by using haplotypespecific k-mer databases. This involved applying (’meryl difference’) command to identify erroneous k-mers and using (’meryl-lookup’) command to extract reads associated with these k-mers. Finally, I used BCFtools to call a consensus32 using the opposite haplotype (’-H 1’) within bcftools consensus command, effectively reversing most haplotype switch errors. 6. SV and SNP Consensus Polishing Consensus polishing33 was automated by following instructions from the T2Tpolish GitHub repository (https://github.com/arangrhie/T2T-Polish). A HiFi mapping file of repetitive k-mers (k=15) was generated with Meryl and used in Winnowmap for 30PCR duplicates refer to read pairs that originate from the same DNA fragment and are amplified multiple times during library preparation, potentially leading to biased coverage if not removed. 31Duplication errors refer to the erroneous presence of redundant sequences in genome assemblies, often caused by uncollapsed haplotypes, repetitive regions, or misassemblies. These can inflate genome size and complicate downstream analyses. 32Calling a consensus in bcftools refers to generating a modified reference sequence by applying variant calls (e.g., from a .vcf file) to a reference genome, resulting in a consensus sequence that reflects the sample’s specific genotype. 33Consensus polishing is the process of correcting errors in a draft genome assembly using aligned reads, typically by averaging observed base calls at each position in a read pileup. Most polishing tools are not haplotype-aware and must be run on both haplotypes together, meaning reads are distributed across haplotigs without distinguishing parental origin. This can lead to miscorrections in heterozygous regions. 26 HiFi read alignment (’-MD -W .repetitive_k15.txt -ax map-pb’), followed by alignment filtering with samtools (’-Sb’). pb-falconc filtered hard-clipped reads and for bit 0x104 in Picard34 (’bam-filter-clipped -t -F 0x104 --output-count-fn’) (https://github.com/bio-nim/pb-falconc) v1.15.0 was then used to filter clipped reads. Genome polishing was performed using the liftover branch of the Racon GitHub repository, using Racon with (’-L -S’) options (Vaser et al.,2017) v1.5.0. After polishing, the k-mers present in the genome (seqmers) were counted using Meryl count (k=21). Merfin was then applied using (’-readmers Illumina.HiFi.gt1.PCRfree.hybrid.meryl -seqmers’) and the hybrid-kmer db of reads to evaluate the results by comparing the distributions in the reads and in the polished genome (Formenti et al.,2022;Rhie et al., 2022). For consensus generation (i.e., to apply polishing edits to the genome assembly), I used BCFtools (’-H 1’) for two rounds of genome correction. Assembly quality metrics were measured for QV, completeness and BUSCOs scores, I found a large increase in QV in most chromosomes (min. increase > +1-5 QV points) after ONT Racon and Merfin, the median QV was 50, the progress over months is presented in Figure 10. 34The SAM flag 0x104 is a combination of 0x004 (read unmapped) and 0x100 (not primary alignment). It indicates that the read is both unmapped and not the primary alignment, typically seen in secondary alignments of unaligned reads or ambiguous multi-mappings. 27 Figure 10 Bighead Catfish Assembly Status January 2024 - November 2025. (A) Haplotype 1 (blue) and Haplotype 2 (purple) ideograms showing gaps on pseudochromosomes (orange rectangles) and telomeres (blue triangles) after manual-review in Juicebox (JBAT). (B) Same assemblies a few months later. 28 7. Mitochondrial Genome Assembly The mitochondrial genome (mtDNA) was assembled by mapping to reference (NC_046749.1) bighead catfish mtDNA, available at NCBI Nucleotide35 . Minimap2 (’-ax map-ont --secondary=no’) mapped nanopore reads to the reference and minimap2 (’-ax sr’) mapped Illumina reads. PCR duplicates in short reads were removed from the alignments using samtools markdup and the results were visualized in the Integrative Genome Viewer (IGV) (Thorvaldsdottir et al.,2012). Pilon (’--fix all --diploid --changes --vcf --tracks --minmq 10’) was used to correct mtDNA (NC_046749.1) and call SNPs, SVs, gaps and local variants and obtain the consensus mtDNA sequence. Reads were re-aligned to the consensus mtDNA sequence and no more SNPs were visible in the IGV. Next, mapped reads were filtered with samtools view (’-F4 -q 20’) and re-assembled de novo with Unicycler (Wick et al.,2017) for comparison. The results were visualized in Bandage-ng (Wick et al.,2015) v2022.09. The assembly was polished with Pilon and the two homologous mitochondria were compared using minimap2 (’--eqx -x asm5’). Gene annotations were generated using MitoFinder (Allio et al.,2020) v1.4.2. To ensure that the mitochondrial genome was correct and not dissimilar to other mtDNAs in Siluriformes, I downloaded all reference mtDNA sequences for (N = 209) species of Siluriformes catfish from NCBI Nucleotide, last accessed September 2024. All sequences were taken from NCBI RefSeq36 and not from NCBI GenBank37. All mtDNA nucleotide sequences (N = 210) were renamed to the Pan-SN naming scheme (https://github.com/pangenome/PanSN-spec), concatenated in a single multi-FASTA including the bighead catfish mtDNA generated here after Pilon, and all-versus-all alignments were performed with wfmash, then ODGI was used to filter the alignment graph, and visualizations were made with Bandage and multiQC (Ewels et al.,2016). All tools of the pipeline were part of the larger PGGB (Pan Genome Graph Builder) (Guarracino et al.,2022). The results of the alignments are shown in Figure 11, the mtDNA aligns well with other sequences in the phylogenetic order and is not an outlier. 35NCBI Nucleotide is a public database providing access to sequences from GenBank and RefSeq. 36NCBI RefSeq is a curated, non-redundant database of genomic, transcript, and protein sequences. 37NCBI GenBank is a comprehensive archive of nucleotide sequences submitted by researchers. 29 Figure 11 All-versus-all 2D representations of mitochondrial DNA sequences alignments in Siluriformes catfishes. The mtDNA assembled for bighead catfish is on the last row (blue color). 30 8. Transposable Element Annotation Two complementary approaches were used to identify transposable elements (TEs)38 in the bighead catfish genome. (1) A species-specific TE library was generated de novo for the bighead catfish. (2) A curated fish TE library was retrieved from the FishTEDB database (https://www.fishtedb.com/) and used as a reference for crossspecies TE annotation. 8.1 De novo transposable element library for bighead catfish De novo TE annotation was performed using EDTA (Extensive de novo TE Annotator) (Ou et al.,2019) v2.2.0. The pipeline integrates several tools to detect different element types. LTR harvest (Ellinghaus et al.,2008) v1.5.10 identifies LTR retrotransposons by efficiently locating structural features such as border positions, LTR lengths, and motifs in large genomic datasets. LTR FINDER (Xu and Wang,2007) and the parallel version of LTR FINDER (Ou and Jiang,2019) v1.1.0 accelerate LTR detection using parallel computing for large genomes. LTR retriever (Ou and Jiang, 2017) v2.9.0 refines LTR annotations to improve accuracy and remove false positives. TEsorter is an accurate and fast way to classify LTR retrotransposons (Zhang et al., 2022) v1.4.7, GRF (Generic Repeat Finder) (Shi and Liang,2019) v1.1 for genomewide de novo repeat detection. TIR-Learner (Su et al.,2019) v1.1.2 uses machine learning to detect terminal inverted repeat (TIR) transposons, including miniature inverted TEs (MITEs). HelitronScanner (Xiong et al.,2014) v1.0 identifies Helitrons by recognizing their characteristic structural motifs. RepeatModeler (Flynn et al.,2020) v2.0.3 performs de novo repeat discovery and builds a comprehensive repeat library. MAFFT (Katoh and Standley,2013) v7.526 is used for multiple sequence alignments, and HMMER (Finn et al.,2011) v3.4 is used for domain-based searching and TE classification. 38Transposable elements (TEs), also known as “jumping genes,” are mobile DNA sequences that can move or replicate within the genome. They play key roles in genome evolution, structural variation, and gene regulation. 31 8.2 FishTEDB general fish-specific TE library collection Next, libraries of TE consensus sequences39 for transposons were retrieved from the FishTEDB database (Shao et al.,2018) version 1, which contains different species of fish and their associated transposable element libraries in .fasta file format, and these sequences were merged with EDTA’s de novo bighead catfish initial TE annotations. To reduce the library complexity and to create a non-redundant TE library (i.e., to lower the number of sequences and redundancy of the combined TE library dataset), I used CD-HIT-EST (’-d 20 -aS 0.95 -c 0.95 -G 0 -g 1 -b 500’) (Li and Godzik,2006) for an 80% identity threshold and 80% alignment coverage of TE sequences in all-against-all sequence clustering following the commonly used 80%-80% rule (Goubert et al.,2022). The reduced TE library was used to mask (i.e., to annotate) the bighead catfish genome for TEs with RepeatMasker (Smit, AFA, Hubley, R. & Green, P RepeatMasker at http://www.repeatmasker.org). 39TE (transposable element) consensus sequences are artificial, representative sequences constructed by aligning and averaging multiple genomic copies of a TE family. They do not correspond to any specific real insertion, but serve as idealized references for annotation and classification. 32 GENOME ASSEMBLY – F1 HYBRID CATFISH Overview of Genome Assembly Strategy Starting from an F1 hybrid catfish (i.e., a first filial generation catfish made from the cross of a pure bighead catfish and a pure North African catfish), I reconstructed two complete haploid, non-homotypic parental genomes, comprising 27 pseudochromosomes for C. macrocephalus and 28 for C. gariepinus (Fig. 25). In most F1 hybrids derived from divergent parental species, the observed heterozygosity rate—typically around 10%—reflects the inter-specific genomic divergence between the two haploid sub-genomes, encompassing both single nucleotide polymorphisms (SNPs) and structural variants (SVs). The complete workflow for generating chromosome-scale assemblies of the Thai F1 hybrid catfish is presented in Figure 12. The workflow consists of six sequential steps (from Step 1A to Step 6) that integrate multiple sequencing technologies and analysis tools and pipelines in specific order used to achieve optimal phasing and Hi-C scaffolding at a high-contiguity, low structural error rate, high solve rates, and high base-accuracy genome assemblies in highly heterozygous F1 genomes. Below, I outline each step in detail. •Step 1: Genomic DNA Sequencing, Genome Survey, and Meryl Databases Genomic DNA was extracted from the liver tissue of a single F1 hybrid individual using a custom protocol (Supikamolseni et al.,2015). Sex was determined to be female based on histology and external gonadal examination (Kitano et al.,2007;Wyneken et al.,2007). Sequencing was performed using four complementary platforms generating four synergistic raw sequencing data types: (i) Pacific Biosciences (PacBio) High-Fidelity (HiFi) sequencing (Wenger et al., 2019) for generating highly accurate contigs, (ii) Oxford Nanopore Technology (ONT) long-read sequencing (Jain et al.,2018b) for scaffolding and resolving repetitive regions, and (iii and iv) Illumina short-read 150 base paired-end sequencing (Hi-C, and standard non Hi-C) (Bentley et al.,2008) for chromoso- 33 mal phasing and error correction, respectively. Genome survey aims in assessing genome characteristics (genome size, heterozygosity, repeat content) from raw sequencing data (e.g., Illumina short-reads) using k-mer analysis in Jellyfish (Marçais and Kingsford,2011) and GenomeScope2.0 (Ranallo-Benavidez et al.,2020). Raw sequencing reads quality was assessed FastQC (http://www. bioinformatics.babraham.ac.uk/projects/fastqc/) v0.1.12 for short reads and using NanoPlot v1.42.0 for long reads. Meryl databases (N = 2) consisted in 21mer and 31-mer (hybrid.hifi.illuminaPCRfree.gt1.db.meryl) hybrid kmer databases, made from combined Illumina and HiFi reads using Meryl (Rhie et al.,2020), as explained in the T2T-polish GitHub repository (https://github. com/arangrhie/T2T-Polish) (Rhie et al.,2022). k-mer Meryl databases use 21mer and 31-mer DNA strings to support quality evaluation and error correction in genome assemblies. The shorter 21-mer spectra provide high sensitivity for detecting potential errors, while the longer 31-mer spectra offer higher specificity, validating true variants and minimizing false positives. During genome polishing, these k-mer databases enhance effective coverage in low-depth HiFi regions by leveraging k-mer spectra to guide targeted corrections and improve assembly accuracy. •Step 2: Contig Assembly and Their Haplotype Phasing Contig assembly was performed using Hifiasm (Cheng et al.,2021), a haplotype-resolved assembler integrating data from HiFi reads, Hi-C reads, and ONT reads. This resulted in two haplotype-separated sub-genomes40 (hap1 and hap2). However, I observed that the Hifiasm algorithm artificially de-duplicated the primary phased haplotype assembly, producing two redundant copies, namely .hic.hap1.p_ctg.fa) and .hic.hap2.p_ctg.fa) which were nearly identical (combined genome size of 3.6 Gb versus expected 1.8 Gb, GenomeScope2.0). The best strategy was therefore to cluster parental haplotigs (i.e., the haplotype-specific contigs also referred to as ”unitigs”, contained in the haplotype-resolved polished unitig graph without 40In the context of an F1 hybrid, a subgenome refers to the haploid genome inherited from one parent species. Since diploid hybrids carry one genome copy from each parent, they contain two sub-genomes— each representing a distinct parental lineage. 40 1,602 •C. gariepinus chromosome 18: left telomeric repeat count increased from 268 to 1,411 •C. macrocephalus chromosome 7: left telomeric repeat count increased from 199 to 924 These refinements increased the number of pseudochromosomes with bilateral telomeres from 43 to 48, while reducing those with unilateral telomeres from 8 to 7. Telomeric additions contributed approximately 400 kb to the total assembly length, representing substantial progress toward telomere-to-telomere (T2T) completeness for multiple chromosomes. Subgenomic analysis revealed additional differences between parental haplotypes, including heterozygosity rate variations and assembly complexity differences attributable to limited nanopore coverage. While most genomic gaps have been resolved, centromeric and satellite regions remain challenging, particularly where HiFi read coverage is absent or where only MAPQ0 Illumina reads provide support, necessitating reliance on lower-accuracy ONT data. 1.4 Quality Value Enhancement Through Manual and Automated Polishing The polishing strategy combined automated QV correction with manual curation. Initial variant calling with Clair3 identified over 60,000 variants, subsequently resolved using Merfin. Despite achieving QV60 at k=21, a final manual curation round introduced 8,410 edits, including 551 large structural variants, resulting in QV improvement of 1–5 points. The complete manual curation process comprised four distinct rounds: Round 1: Sniffles2 identified 755 target regions across all alignment files for manual review. 41 Round 2: Cross-assembly curation using Unimap examined scaffoldsfrom Flye, GreenHill, and Hifiasm P-contigs, incorporating seqmer error files. This resulted in 2,221 manual edits and application of 5,951 variants. Round 3: Combined Winnowmap and Unimap alignments enabled identification of 8,710 variants from 5,740 alignments, with 1,932 applied to the assembly. Additionally, 300 gaps were resolved following Sniffles2 and BCFtools consensusbased corrections. Round 4: Correction of 97 large structural variants, including 10 visible only through VerityMap alignments. To ensure genome reliability, 65 gaps were introduced to mask sequences lacking proper read support or exhibiting high error profiles. Post-masking, 247 structural variants (196 insertions, 73 deletions) identified using Sniffles2 were applied to the genome. These variants, derived from ONT pb-falconc-filtered alignments (https://github.com/bio-nim/pb-falconc) realigned at Hifiasm contig boundaries with 50 nt margins, were approximately 90% homozygous and 10% heterozygous. Genomic error k-mers showed non-uniform chromosomal distribution, clustering in regions with ONT-only or single-depth coverage (511 loci total). Further correction using Clair3 (Zheng et al.,2022) addressed 22,625 ONT-based consensus errors and 7,558 HiFi Winnowmap variants, achieving median QV31. Assembly validation using IGV demonstrated quality improvements and structural resolution. Chromosome 03 showed clear reduction in k-mer error rates and increased QV post-polishing. Chromosome 02 analysis revealed a false duplication present in Hifiasm but absent in Flye, confirmed by doubled homozygous coverage and lack of supporting HiFi alignments. A single spanning ONT read and clipped Illumina read validated boundary precision. Consistent single-nucleotide markers across all data types (HiFi, ONT, Flye) suggest potential for further refinement using high-confidence variants (AF > 0.9) with Clair3 (Zheng et al.,2022) and WhatsHap (Martin et al.,2016). 42 Figure 13 Quality value and coverage validation using Integrative Genomics Viewer. (A) Chromosome 03 demonstrating QV improvement following polishing procedures. (B) Chromosome 02 showing false duplication and gap analysis through comparative assembly assessment. Following eight months of genome curation, the estimated QV improved from Q40 to Q60, demonstrating that F1 hybrid genomes can achieve superior nucleotidelevel accuracy compared to single-species assemblies. For comparison, the pure-breed bighead catfish assembly achieved median QV40–50 (k=21, k=31, pileup), while the current F1 hybrid genome demonstrates ten-fold higher accuracy using comparable sequencing coverage, analytical tools, and curation effort. 43 Figure 14 F1 Hybrid Catfish Assembly status January 2024 - November 2024. 44 VALIDATION OF CATFISH ASSEMBLIES Benchmarking Methodology This chapter presents the results and technical validations performed on the genome assemblies of the two catfish species. First, I present the most common metrics to measure the quality and completeness of the assemblies, next I present each individual assembly with results, finally comparative genomic analysis are provided and illustrate how these two genomes can be used for comparative studies across species. Metrics of continuity, structural accuracy, base accuracy, and functional completeness were used for benchmarking as described in the Vertebrate Genome Project (VGP) paper (Rhie et al.,2021). 1. Metrics Assessed 1.1 Continuity and summary statistics To assess continuity and summary statistics of the assembly, I compute the following measures for scaffold/contig (N50, N90, NG50, LG50, and LG90)50 with RagTag (’ragtag.py asmstats -g’) (Alonge et al.,2022) v2.1.0. 1.2 Repeat completeness and continuity of repeats To measure repeat completeness and continuity of repeats for assessing assembly quality, I estimated the percentage of fully assembled LTR retroelements (LTRRTs)51 and computed the long terminal repeat (LTR) Assembly Index (LAI) using two 50N50 and N90 represent the contig or scaffold length such that 50% or 90% of the total assembly length is contained in contigs/scaffolds of that length or longer. NG50 is similar to N50 but calculated relative to an expected reference genome size. LG50 and LG90 indicate the minimum number of contigs or scaffolds whose combined length makes up 50% or 90% of the assembly, respectively. 51LTR-RTs (Long Terminal Repeat Retrotransposons) are a class of transposable elements characterized by direct long terminal repeats at both ends. They replicate via an RNA intermediate and reverse transcription, and are major contributors to genome size and structure in many eukaryotic organisms. 45 reference-free programs, LTR Assembly Index (LAI) (Ou et al.,2018) and LTR_retriever (Ou and Jiang,2017) v2.9.00. To assess the assembly quality of complex repeats (centromeres), I used TandemTools and TandemQUAST (Mikheenko et al.,2020) v1. Telomere prediction for the presence / absence and orientation of telomeres was done with TIDK, a Telomere Identification Toolkit Brown et al. (2023), implemented in TeloExplorer (’-m 50 -c animal’), a module of QuarTeT Lin et al. (2023). To estimate gaps, I used the script detgaps, available on GitHub (https://github.com/dfguan/asset). 1.3 Structural accuracy (regional and structural errors and reliable blocks) To assess structural accuracy I used the CRAQ software (’sms_coverage=5 ngs_coverage=20 -B T --minimap2-sensitve’) (Li et al.,2023) v1.0.9, more specifically (’-q 20 -m 2 -f 0.4 -h 0.6 -r 0.75 -a 20’). CRAQ is a method for assembly structural validation relying on the study of mapped reads, clipped reads and coverage support by two or more simultaneous sequencing platforms to find supporting regions or reliable blocks (Rhie et al.,2021).and visualized results using the Integrative Genome Viewer (IGV) (Thorvaldsdottir et al.,2012). CRAQ evaluates genome assembly quality based on clipped-read evidence from both short and long reads. It reports two main types of errors: 1. Clip-based Regional Errors (CREs), which represent small-scale local misassemblies identified by short-read clipping without structural disruption. 2. Clip-based Structural Errors (CSEs), which represent large-scale misassemblies supported by long-read clipping near breakpoint regions. These errors are quantified into two sub-scores: R-AQI (based on CREs) and S-AQI (based on CSEs). Both scores contribute to the global Assembly Quality Index (AQI), ranging from 0–100, where higher values indicate better assembly integrity. AQI is computed based on error density relative to genome size. 46 1.4 Cross-species structural correctness To assess cross-species structural correctness, I used the same references from related catfish species for one-to-one nucleotide-level alignments of orthologous segments with MashMap252 (’-s 2000000 --pi 90 -c 100000’) (Jain et al.,2018a) v3.1.3. 1.5 Base accuracy and assembly completeness To assess base accuracy and completeness, specifically Mercury’s quality values (QV) and 21-mer genome completeness (%), I used Mercury (Rhie et al.,2020) v1.3. Merqury was run three times using different k-mer databases: Illumina, HiFi, and a hybrid 21-mer database combining Illumina and HiFi reads, as explained above, in the ”Consensus polishing” section. To assess functional completeness and evaluate the completeness of the gene set, I used the lineage of ray-finned fishes and BUSCO (’-l actinopterygii_odb10’) (Simão et al.,2015) v5.6.1. Finally, nucleotide accuracy was verified by read-to-assembly mapping, performed using minimap2 (’-ax map-hifi --secondary=no’) and Winnowmap (’-W repetitive.15.txt’) at MAPQ >10 for HiFi reads, while for ONT reads I used the '-ax map-ont' preset for minimap2 and the '-ax map-pb' preset in Winnowmap MAPQ >10. For Visualisations, I used IGV (’igv -g $genome $BAM(s) $merqury_only_bed_wig_kmer_files’). 52MashMap2 is a fast and approximate algorithm for computing whole-genome homology maps. It uses minimizer-based locality-sensitive hashing to rapidly identify high-confidence regions of sequence similarity, making it suitable for genome-to-genome alignments, structural variation detection, and referenceguided scaffolding (Jain et al.,2018a). 47 Results – Bighead Catfish The haplotype-resolved de novo assembly resulted in 27 Hi-C scaffolds, and the chromosome number was 2n = 2x = 54 pseudochromosomes (Maneechot et al.,2016). 1. Raw Read Quality Figure 15 summarizes raw data quality for the bighead catfish genome. Figure 15 Sequencing summary and genome survey of bighead catfish. (A) Sequencing protocols and coverages. (B) GenomeScope k-mer profile showing genome size, repeat content, and 17.6% heterozygosity. (C) HiFi reads: high quality and 15 kbp peak length. (D) ONT reads: broader size range with N50 30 kbp and moderate quality. 48 2. Global Assembly Metrics Sequence variation analysis53 was performed with PlotSR suite (Goel and Schneeberger,2022), revealing 1,968,666 heterozygous SNPs across both haplotypes after minimap2 ('-ax asm5') cross-haplotype mapping, resulting in a heterozygosity rate54 of 0.594%. Structural variant analysis identified 392,973 insertions totaling 7.67 Mb in Haplotype 2, while Haplotype 1 had 393,127 deletions spanning 7.75 Mb. Copy number variations55 included 114 copy gains in Haplotype 2 (184 Kb) and 123 copy losses in Haplotype 1 (416 Kb). Highly divergent regions56 were identified, spanning 57.9 Mb in Haplotype 1 and 56.0 Mb in Haplotype 2. Additionally, 47 tandem repeat clusters57 were detected, covering 10 Kb in Haplotype 1 and 7 Kb in Haplotype 2. Table 1 Summary statistics of the haplotype-resolved Clarias macrocephalus genome assembly. Features Haplotype 1 Haplotype 2 Overall quality (x.y.P.Q.C) 6.30.P7.53.Q48.C94.25 6.30.P7.53.Q48.C94.25 Genome size (Mb) 875 Mb 880 Mb pseudochromosomes 27 27 Mitochondrion length (bp) 16,510 bp N/A NG50 of contigs (Mb) 3 Mb 3 Mb N50 of scaffolds (Mb) 34 Mb 34 Mb LG50/LG90 of scaffolds 11 / 24 (Hap1) 11 / 24 (Hap2) Number of gaps 752 653 GC content (%) 39.32% 39.32% Complete BUSCOs N (%) 93.0% (Hap1) 93.5% (Hap2) Median QV (Merqury k21) 42.99 (Hap1) 43.33 (Hap2) 48.22 (All) Completeness (Merqury k21) 89.11 (Hap1) 88.88 (Hap2) 95.25 (All) 53Comparison of Haplotype 1 and 2 chromosomes to detect heterozygous variants such as SNPs, indels, and structural changes. 54Proportion of sites where the two haplotypes differ, reflecting allelic diversity. 55Duplications or deletions causing variable copy numbers of genomic regions. 56Genomic segments showing strong sequence divergence between haplotypes or species. 57Short motifs repeated head-to-tail, often in telomeres or centromeres. 49 56 Table 3 Summary of scaffold metrics in the haplotype-resolved genome assembly of Clarias macrocephalus,Haplotype 2. Scaffold Name Size Errors Quality Values Gaps Left Telomere Right Telomere S-AQI Haplotype_Linkage (Mb) 21-mer hyb.k21.meryl (N) (5’) AACCCT) (AGGGTT 3’) (%) fClaMac_2_LG_01 51,51 2649 55.5298 17 - - 89.28 fClaMac_2_LG_02 44,94 6893 51.161 25 Left 255 (+) - 71.62 fClaMac_2_LG_03 39,44 8894 49.5912 16 - - 83.57 fClaMac_2_LG_04 38,01 6912 50.5588 18 Left 293 (+) Right 526 (-) 88.17 fClaMac_2_LG_05 38,16 3424 53.4601 23 - Right 125 (-) 78.26 fClaMac_2_LG_06 38,66 2675 54.3325 17 - - 80.86 fClaMac_2_LG_07 33,93 5651 50.8217 17 Left 169 (+) - 89.91 fClaMac_2_LG_08 47,19 10458 49.5856 29 Left 174 (+) Right 190 (-) 89.44 fClaMac_2_LG_09 28,8 10239 47.5736 14 Left 641 (+) - 91.91 fClaMac_2_LG_10 29,07 5001 50.8593 14 Left 268 (+) - 95.24 fClaMac_2_LG_11 25,76 6627 48.9338 21 - - 89.75 fClaMac_2_LG_12 30,93 6915 48.7951 19 - - 82.58 fClaMac_2_LG_13 23,66 3213 51.9117 15 - Right 100 (+) 86.65 fClaMac_2_LG_14 21,06 3487 50.9324 7 - - 100.00 fClaMac_2_LG_15 24,93 7673 48.2988 17 Left 134 (+) Right 137 (-) 85.58 fClaMac_2_LG_16 24,83 3011 52.2777 10 - - 81.74 fClaMac_2_LG_17 22,34 2012 53.6414 15 Left 175 (+) - 85.99 fClaMac_2_LG_18 24,44 1986 54.0521 11 - - 100.00 fClaMac_2_LG_19 30,86 3566 52.521 16 - Right 362 (-) 90.62 fClaMac_2_LG_20 34,51 2384 53.5357 13 - - 100.00 fClaMac_2_LG_21 27,99 2550 53.5741 13 - - 100.00 fClaMac_2_LG_22 33,35 4588 51.8225 16 Left 128 (+) Right 908 (-) 88.64 fClaMac_2_LG_23 31,45 9701 48.2459 6 Left 214 (+) Right 793 (-) 82.68 fClaMac_2_LG_24 27,39 2840 53.0349 6 Left 114 (+) - 93.87 fClaMac_2_LG_25 37,88 7891 49.9914 15 - - 87.01 fClaMac_2_LG_26 40,17 11271 48.6822 13 - Right 607 (-) 88.08 fClaMac_2_LG_27 29,56 1236 56.9751 12 Left 155 (+) - 90.78 MT_from_NC_046749 0,02 - - - - - - 57 6. TE Metrics 6.1 Dating of TE Replicative Bursts To estimate the expansion history of major transposable element (TE) superfamilies in Clarias macrocephalus, I calculated the nucleotide divergence of each TE copy relative to its consensus sequence and converted this divergence into approximate insertion ages using a neutral mutation rate58 ( µ =6×10−9per site per generation) previously determined in electric catfish (Liu et al.,2023b). Assuming a one-year generation time, the estimated expansion peaks indicate that LTR/Gypsy and TIR/Mutator families expanded around 30–33 Mya, LTR/Copia and TIR/hAT around 28–30 Mya, and TIR/CACTA, PIF/Harbinger, and Tc1/Mariner families between 16 and 25 Mya. These results (Figure 20) reveal multiple waves of TE proliferation, likely reflecting episodes of genome instability or evolutionary transition in the species’ history. Figure 20 TE divergence profiles in bighead catfish genome. 58The neutral mutation rate is the rate at which mutations accumulate in non-functional genomic regions, providing a baseline for molecular-clock dating. 58 6.2 Genome composition in TE Transposable element (TE) annotation indicated that 35.25% of the bighead catfish genome consisted of repetitive sequences. Among these, TIR DNA transposons comprised 19.12%, Helitrons 4.47%, LTR retrotransposons 8.30%, LINE elements 0.46%, and the remaining 2.82% consisted of uncharacterized repeats. Table 4 Transposable element content in Bighead catfish Haplotype 1. Class Superfamily Haplotype 1 Haplotype 2 Features # of elements LINE I128 L1 547 L2 1,129 Rex 556 LINE/Unknown 7,412 LTR Bel_Pao 151 Copia 53 Gypsy 68,215 LTR/Unknown 13,466 TIR CACTA 357,388 Mutator 172,578 PIF_Harbinger 36,538 Tc1_Mariner 41,500 hAT 80,189 PiggyBac 758 Polinton 281 NonLTR DIRS_YR 678 Penelope 417 NonTIR Helitron 166,251 Other Repeat Fragment 78,900 Total Interspersed 1,148,329 59 LINEs (Long Interspersed Nuclear Elements) are non-LTR retrotransposons that transpose via a copy-and-paste mechanism using reverse transcriptase; although less abundant in fish genomes than in mammals, they contribute to structural variation. LTR retrotransposons replicate through an RNA intermediate using reverse transcription and are flanked by long terminal repeats, playing key roles in genome expansion and gene regulation. TIR DNA transposons (Terminal Inverted Repeat transposons) are cut-and-paste elements flanked by inverted repeats recognized by a transposase enzyme, contributing to genome rearrangement and regulatory diversification. Non-LTR retrotransposons such as DIRS and Penelope elements add further diversity to the repetitive landscape through distinct mobilization mechanisms. Helitrons, belonging to the NonTIR class, are rolling-circle DNA transposons that lack terminal inverted repeats and transpose through a mechanism similar to rolling-circle replication, often capturing and mobilizing gene fragments that promote genomic innovation. Together, these TE classes reflect a dynamic evolutionary history for C. macrocephalus, consistent with high DNA transposon activity typical of teleost genomes. 7. Macrosynteny Analysis To examine orthologous relationships and identify syntenic regions between catfish and related teleost species, I obtained proteomes (.pep.fa files) from Ensembl v2024 (Harrison et al.,2023), ensuring the latest assembly versions for consistency. Species are Oncorhynchus mykiss (USDA_OmykA_1.1), Oryzias latipes (ASM223467v1), Lates calcarifer (ASB_HGAPassembly_v1), Cyprinus carpio (Cypcar_WagV4.0), Danio rerio (GRCz11), Oreochromis niloticus (O_niloticus_UMD_NMBU), and Lepisosteus oculatus (LepOcu1). Orthologous gene clusters (Orthogroups) were first identified with OrthoFinder (’orthofinder -f /proteomes -M msa -a 126 -t 126’) (Emms and Kelly,2019) v.2.5.5. Subsequently, I utilized xrhbexpress (https://github.com/SamiLhll/ rbhXpress) and macrosyntR (El Hilali and Copley,2023) v0.2.19 to analyze the resulting orthogroups and visualize their syntenic relationships across species. This combined approach allowed for an in-depth view of conserved and divergent genomic regions within the catfish lineage and related teleosts (Figure 21). 60 Figure 21 Synteny analysis for various catfish samples. (A) Whole-genome 4-way synteny analysis from 1413 shared single-copy orthogroups across Clarias gariepinus (GCA_024256425.2), Clarias macrocephalus,Danio rerio (GRCz11)Danio rerio and Oreochromis niloticus (O_niloticus_UMD_NMBU), reveals conserved chromosome structures and rearrangements. (B) Simple phylogeny and geological timescales (TimeTree5, https://timetree. org/). 61 Results – F1 Hybrid Catfish The haplotype-resolved de novo assembly of the F1 hybrid catfish (C. macrocephalus ×C. gariepinus) yielded a total of 55 Hi-C scaffolds, representing 27 + 28 chromosomes from the parental species (Lewin et al.,2019;Maneechot et al.,2016). 1. Raw Read Quality Figure 22 summarizes the sequencing data quality and genome survey statistics. Figure 22 Sequencing summary and genome survey of the F1 hybrid catfish (C. macrocephalus ×C. gariepinus). (A) GenomeScope2.0 k-mer profile (k=21) estimating genome size (903 Mb), heterozygosity (10.1%), and sequencing error rate (0.28%). (B) Sequencing coverage by data type: HiFi (30×), ONT (36×), Hi-C (37×), and Illumina (55×). (C) HiFi reads show high quality with 15 kbp modal length. (D) ONT reads display broader length distribution (N50 ≈ 30 kbp). 62 1.1 Illumina paired-end short-reads Standard Illumina paired-end short-reads sequencing Bentley et al. (2008) was performed on the Illumina NextSeq2000 platform (by NovoGene), resulting in more than 50 Gb of raw read data (read-length=151 nucleotides), which is equivalent to a genome coverage of 36X. Raw reads had an average quality of Q25, and 92% of the dataset had a median quality of Q30. 1.2 Proximity-ligation (Hi-C) paired-end short-reads Proximity-ligation (Hi-C) data was obtained by following a standard in situ Hi-C protocol from 2009 Dudchenko et al. (2017). Briefly, a nuclear ligation was performed by cross-linking chromosomes, then a restriction digestion was carried out with DpnII endonuclease. The chromatin conformation capture library was prepared using Phase Genomics (https://info.phasegenomics.com/protocols), and the genomic DNA was sequenced on the Illumina NextSeq2000 platform, resulting in paired-end short-reads (151 nucleotides in length) for 37.19 Gb of raw data (approximately 100M pairs of reads per billion bases in genome length) corresponding to a genome coverage of 39.77X. 97.18% of the dataset had a quality score of Q30 or higher (Supplementary Table 3). For Haplotype 1, 123.96M read pairs were sequenced, yielding 74.77M unique Hi-C contacts, with 59.40M valid contacts (22.71M inter-chromosomal, 36.69M intrachromosomal), including 17.47M short-range (<20Kb) and 19.72M long-range (>20Kb) contacts. Haplotype 2 showed similar metrics. 1.3 PacBio HiFi long-reads PacBio’s HiFi long-read sequencing was performed using a SMRT cell on the Sequel II system, resulting in 10.62 Gb of raw data corresponding to a genome coverage of approximately 12.9X or 6-7X per haplotype. The read length N50 value was 10,379 bases, the raw reads’ mean quality score was Q28.7, and the median quality score was Q36. Over 72% were above Q30, and 86% above Q25. There were 191,790 63 reads with a cumulative length of 5,066,253,114 bases (5 Gb) in the high-quality reads subset (> Q10). 1.4 Oxford nanopore technologies (ONT) 1D-long noisy reads ONT long noisy reads were prepared from 1D-fragmented and size-selected pooled libraries of end-polished, nick-repaired, and A-tailed DNA fragments, produced by ligating sequencing adapters onto double-stranded DNA fragments purified with AMPure XP beads. DNA quality was assessed with QuBit before sequencing on the Nanopore MiniION system (NovoGene). The standard ONT libraries comprised 32.62 Gb of raw read data, corresponding to a genome coverage of 36X, or 18X per haplotype. The read N50 value was 4 kb. We estimate an ONT data median read quality of Q11, with 97.0% of reads having quality > Q7 (4,023,582 reads), corresponding to 20% base-errors. Unfortunately the quality of the data was extremely low resulting in only < 4X of genome coverage of reads usable from 36X sequenced. 2. Global Assembly Metrics Our F1 hybrid catfish genome released into GenBank consists of 55 scaffolds (all exceeding 10,000 bp), totaling 1,858,695,542 bp, one mtDNA and 0 unplaced sequences. Notably, every contig in scaffolds exceeds 1,000,000 bp, with the longest spanning 51,692,840 bp and the second longest being 51,114,572 bp which correspond to a Telomere-to-Telomere (T2T) chromosome with 0 gaps. Our assembly reached an N50 of 33,844,116 bp which is an indicator of high contiguity. Read-to-contig HiFi minimap2 alignment performed in the context of Inspector confirms a 99.96% mapping rate, a 0.52% split-read rate, and an average mapped depth of 10.65×, which remains consistent for contigs > 1 Mbp. Structural validation identified 22 structural errors (12 expansions, 7 collapses, 2 to 5 haplotype switches, and 1 inversion), while small-scale errors averaged 4.79 per Mbp (8,897 total), comprising 6,226 base substitutions, 1,189 expansions, and 1,482 collapses. Further quality control was performed with Merqury (k=21) to evaluate k-mer-based metrics (Figure 25, Table 5and Table 6). Consensus Quality 64 Values (QVs) for the hybrid genome and each sub-genome were 50.1, 48.3, and 49.5, respectively, with k-mer completeness scores of 96.7%, 95.8%, and 97.4% for hybrid (2 haps, combined), bighead, and North African catfish sub-genomes. Merqury’s spectrum analysis confirmed the high quality and completeness of the haplotype-resolved assembly with minimal assembly errors. Overall, the polished assembly achieved a QV of 50.28 (pileup-based, Inspector v1.3), indicating a robust base-level accuracy and structural integrity suitable for various types of downstream analyses. 3. Comparative synteny with previous C. macrocephalus and C. gariepinus Comparative synteny analyses revealed extensive genomic conservation between the hybrid catfish genome and its parental species (Appendix Figure 31; Appendix Figure 32). In the North African catfish (C. gariepinus) comparison, macrosynteny across 28 pseudo-chromosomes indicated strong structural collinearity with the hybrid genome, interrupted only by localized inversions, translocations, and small duplicated regions. The overall sequence variation landscape was dominated by highly divergent regions and SNPs, while structural variants represented a minor fraction of total differences. Similarly, in the Bighead catfish (C. macrocephalus) comparison, 27 pseudo-chromosomes showed high collinearity and preserved chromosomal organization, confirming the structural stability of both parental lineages. Variation analyses identified modest insertions, deletions, and inversion tracts, consistent with evolutionary divergence preceding hybridization. Together, these results indicate that both parental genomes contributed largely conserved chromosomal architectures to the hybrid lineage, supporting a high-quality, structurally stable hybrid genome assembly. 3.1 Expected Genome Sizes, k-mer Profiles and Hybrid Genome Assessment (Appendix Figure 30) presents GenomeScope2.0 k-mer profiles to evaluate genome size, heterozygosity, and repeat content for the F1 hybrid genome and its parental species. Separate k-mer distributions are shown for individual parental lineages from distinct bloodlines to assess intra-species variation. 65 •Parental individuals (same species, different bloodlines): The C. macrocephalus individual exhibits low heterozygosity (~0.056%), while C. gariepinus shows moderate heterozygosity (~1.56%), consistent with expected intraspecific levels. Broader peaks at k=21 likely reflect lower Illumina coverage (<20×), but specieslevel divergence remains evident. •F1 hybrid genome: At k=21, the estimated genome size is approximately 1.8 Gb, aligning with the combined haploid contributions of both species. The apparent heterozygosity rate of 10% reflects typical divergence in inter-specific F1 hybrids with distinct sub-genomes. This divergence results in a bimodal distribution due to haplotype differences. At k=31, heterozygosity estimates in the F1 hybrid decrease as k-mer size increases, reflecting partial subgenome separation and reducing false overlaps. 4. Hi-C Contact Validation Separate Hi-C contact matrices were generated for each chromosome-scale scaffold in the F1 hybrid genome to confirm subgenome identity and structural integrity. As shown in Figures 23 and 24, the two sub-genomes display distinct, internally consistent chromosomal contact patterns, supporting successful haplotype separation and chromosome-scale phasing. •African catfish subgenome (28 chromosomes): The majority of scaffolds show strong diagonal contact patterns without major off-diagonal disruptions, indicating high contiguity and few mis-joins. A few chromosomes exhibit minor local noise or contact gaps, which were further curated using Juicebox manual correction and IGV inspection. •Bighead catfish subgenome (27 chromosomes): Hi-C interaction maps demonstrate similar quality, with well-defined diagonals consistent with chromosomelength assemblies. Chromosomes such as Chr07 and Chr23 displayed additional 72 Table 5 North African subgenome scaffold metrics with Quality Values from Merqury (k=21 and k=31), telomere orientation, and gap counts. Scaffold_name Length QV_k21 Errors_k21 QV_k31 Errors_k31 Left_telomere Gaps Right_telomere Sub-Genome_Chromosome Mb. 21-mer 21-mer 31-mer 31-mer (5’ AACCCT) (N) (AGGGTT 3’) fClaHyb_african_1_Chr_01 51,1 64.5442 377 59.4881 1783 119 (+) T2T 0 fClaHyb_african_1_Chr_02 51,7 56.0219 2713 56.0911 3942 174 (+) 10 fClaHyb_african_1_Chr_03 47,6 59.3028 1174 54.5941 5125 1701 (+) 31603 (-) fClaHyb_african_1_Chr_04 44,3 49.4714 10490 56.9825 2716 185 (+) 8 160 (-) fClaHyb_african_1_Chr_05 41,5 62.5139 488 61.5864 892 1472 (+) T2T 1883 (-) fClaHyb_african_1_Chr_06 41,3 54.5824 3021 56.2081 3068 278 (+) 11447 (-) fClaHyb_african_1_Chr_07 42,6 58.0864 1389 54.3111 4889 384 (+) 3382 (-) fClaHyb_african_1_Chr_08 40,2 54.3182 3124 59.8602 1287 0 T2T 130 (-) fClaHyb_african_1_Chr_09 37,3 55.9044 2013 53.8378 4784 1449 (+) 1315 (-) fClaHyb_african_1_Chr_10 34,9 67.6478 126 60.3736 993 0 2127 (-) fClaHyb_african_1_Chr_11 33,5 51.4996 4986 58.2486 1555 1614 (+) T2T 1997 (-) fClaHyb_african_1_Chr_12 35,3 54.4649 2653 58.1993 1657 0 6 0 fClaHyb_african_1_Chr_13 33,6 66.4695 159 63.3235 483 1525 (+) 21557 (-) fClaHyb_african_1_Chr_14 32,4 47.2905 12717 61.6029 690 105 (+) T2T 226 (-) fClaHyb_african_1_Chr_15 35 64.2407 275 58.2795 1602 328 (+) 2118 (-) fClaHyb_african_1_Chr_16 33,8 60.5363 627 57.9679 1675 1602 (+) 4137 (-) fClaHyb_african_1_Chr_17 30,3 55.2666 1892 56.604 2052 335 (-) T2T 276 (-) fClaHyb_african_1_Chr_18 30,5 56.641 1382 57.9088 1522 1411 (+) 32289 (-) fClaHyb_african_1_Chr_19 30,7 64.5469 226 58.6026 1311 784 (+) 2178 (+) fClaHyb_african_1_Chr_20 27,1 57.0276 1128 60.0252 835 171 (+) 2406 (-) fClaHyb_african_1_Chr_21 30,6 60.6832 549 60.1377 919 1936 (+) 20 fClaHyb_african_1_Chr_22 30 62.328 369 60.2163 886 0 2129 (-) fClaHyb_african_1_Chr_23 26,2 52.9697 2773 58.0814 1260 279 (+) 11739 (-) fClaHyb_african_1_Chr_24 25,9 60.3587 501 60.6659 689 185 (+) 2114 (-) fClaHyb_african_1_Chr_25 25,6 58.5637 748 59.0404 989 232 (+) T2T 129 (-) fClaHyb_african_1_Chr_26 26,3 56.7889 1163 57.358 1506 1343 (+) 30 fClaHyb_african_1_Chr_27 24 68.2243 76 59.5233 832 0 T2T 264 (-) fClaHyb_african_1_Chr_28 20,3 59.8335 443 56.9713 1261 0 20 73 Table 6 Bighead subgenome scaffold metrics with Quality Values from Merqury (k=21 and k=31), telomere orientation, and gap counts. Scaffold_name Length QV_k21 Errors_k21 QV_k31 Errors_k31 Left_telomere Gaps Right_telomere Sub-Genome_Chromosome Mb. 21-mer 21-mer 31-mer 31-mer (5’ AACCCT) (N) (AGGGTT 3’) fClaHyb_bighead_1_Chr_01 50,1 55.7992 2767 50.7048 13207 998 (+) 14 740 (-) fClaHyb_bighead_1_Chr_02 44,3 51.8276 6102 54.4421 4934 815 (+) 50 fClaHyb_bighead_1_Chr_03 40,4 44.33 31251 55.4715 3531 327 (+) 12 279 (-) fClaHyb_bighead_1_Chr_04 38,4 52.1302 4933 52.1322 7281 222 (+) 13 819 (-) fClaHyb_bighead_1_Chr_05 38,1 47.0334 15835 53.7219 4994 321 (+) 14 301 (-) fClaHyb_bighead_1_Chr_06 37,4 52.5469 4370 50.6785 9921 282 (+) 14 0 fClaHyb_bighead_1_Chr_07 33,8 50.998 5644 48.4795 14888 924 (+) 15 629 (-) fClaHyb_bighead_1_Chr_08 45,6 56.2767 2252 51.0889 10984 161 (+) 11 1237 (-) fClaHyb_bighead_1_Chr_09 30 48.1331 9677 48.5674 12936 184 (+) 13 751 (-) fClaHyb_bighead_1_Chr_10 29,8 54.8059 2065 49.5164 10314 316 (+) 12 664 (-) fClaHyb_bighead_1_Chr_11 25 51.0866 4085 53.4888 3469 0 21 267 (-) fClaHyb_bighead_1_Chr_12 31,6 48.5768 9241 47.4853 17546 559 (+) 13 0 fClaHyb_bighead_1_Chr_13 24,9 44.8933 16909 48.6904 10418 740 (+) 8220 (-) fClaHyb_bighead_1_Chr_14 23,3 44.5272 17249 51.8503 4715 0 8190 (-) fClaHyb_bighead_1_Chr_15 27,6 39.2488 68816 52.3266 4848 200 (+) 12 107 (-) fClaHyb_bighead_1_Chr_16 25,2 51.936 3389 53.2655 3686 656 (+) 9818 (-) fClaHyb_bighead_1_Chr_17 24 40.0722 49568 46.7745 15616 326 (+) 15 222 (-) fClaHyb_bighead_1_Chr_18 25,2 52.6726 2856 52.3131 4582 632 (+) 11 973 (-) fClaHyb_bighead_1_Chr_19 31,1 53.8536 2692 52.9809 4841 674 (+) 8751 (-) fClaHyb_bighead_1_Chr_20 32,7 51.6697 4681 51.0227 7929 410 (+) 18 422 (-) fClaHyb_bighead_1_Chr_21 29,2 43.0312 30479 51.9822 5662 805 (+) 16 253 (-) fClaHyb_bighead_1_Chr_22 34,9 51.0828 5705 50.9355 8715 568 (+) 13 121 (-) fClaHyb_bighead_1_Chr_23 31,7 42.7592 35295 48.8582 12797 931 (+) 81022 (-) fClaHyb_bighead_1_Chr_24 29,6 50.1103 6080 52.8055 4829 851 (+) 10 159 (-) fClaHyb_bighead_1_Chr_25 40,3 52.391 4875 52.1653 7582 0 9798 (-) fClaHyb_bighead_1_Chr_26 40,1 53.1914 4035 51.6319 8529 1038 (+) 11 0 fClaHyb_bighead_1_Chr_27 30,5 48.7036 8641 48.5145 13324 1858 (+) 10 219 (-) 74 6.3 Structural Accuracy Assessment Table 7and Table 8present structural accuracy metrics from CRAQ for both sub-genomes, including Structural Accuracy Index (S-AQI), Regional Accuracy (R-AQI), and heterozygous loci for structural and regional metrics (Avg.CSH/Avg.CRH). Note that CRH and CSH values are expected to be close to 0 due to the high genomic diversity between the two parental species (divergence > 10%, mash-based (Ondov et al., 2019), not shown here). North African catfish scaffolds displayed near-perfect S-AQI and R-AQI values, both ranging approximately from 97% to 100%, indicating minimal misassemblies at both large-scale and regional-scale levels. In contrast, bighead catfish scaffolds exhibited slightly lower accuracy, with S-AQI values around 90–97% and R-AQI between 90–96%. This reduction in quality may be attributed to factors such as lower species heterozygosity, increased local assembly complexity due to transposable element content, or potential sequencing quality issues, possibly arising from degraded genomic DNA used in ONT sequencing. 75 Table 7 North African subgenome structural accuracy metrics from CRAQ: S-AQI, R-AQI, and heterozygous regions (Avg.CSH/Avg.CRH). Scaffold_name Mapping.rate Avg.CRH Avg.CRE Avg.CSE Regional-AQI Avg.CSH Structural-AQI Sub-Genome_Chromosome ONT/PE (%) Count Count Count Accuracy (%) Count Accuracy (%) Genome ( 55 Chromosomes) >97% 0.046 0.270 0.021 97.33 0.001 97.92 fClaHyb_african_1_Chr_01 0.992 0.020 0.078 0.000 99.22 0.000 100.00 fClaHyb_african_1_Chr_02 0.981 0.000 0.098 0.000 99.02 0.000 100.00 fClaHyb_african_1_Chr_03 0.955 0.033 0.154 0.000 98.47 0.000 100.00 fClaHyb_african_1_Chr_04 0.980 0.046 0.207 0.046 97.95 0.000 95.50 fClaHyb_african_1_Chr_05 0.986 0.048 0.097 0.000 99.04 0.000 100.00 fClaHyb_african_1_Chr_06 0.998 0.024 0.085 0.000 99.15 0.024 100.00 fClaHyb_african_1_Chr_07 0.978 0.000 0.120 0.000 98.81 0.000 100.00 fClaHyb_african_1_Chr_08 0.992 0.025 0.100 0.000 99.00 0.000 100.00 fClaHyb_african_1_Chr_09 0.992 0.076 0.083 0.027 99.17 0.000 97.34 fClaHyb_african_1_Chr_10 0.998 0.066 0.172 0.000 98.29 0.000 100.00 fClaHyb_african_1_Chr_11 0.905 0.030 0.188 0.000 98.13 0.000 100.00 fClaHyb_african_1_Chr_12 0.873 0.086 0.230 0.000 97.72 0.000 100.00 fClaHyb_african_1_Chr_13 0.917 0.075 0.119 0.000 98.81 0.000 100.00 fClaHyb_african_1_Chr_14 0.990 0.031 0.062 0.062 99.38 0.000 93.96 fClaHyb_african_1_Chr_15 0.922 0.076 0.152 0.000 98.49 0.000 100.00 fClaHyb_african_1_Chr_16 0.952 0.000 0.248 0.000 97.55 0.000 100.00 fClaHyb_african_1_Chr_17 0.997 0.033 0.083 0.000 99.18 0.000 100.00 fClaHyb_african_1_Chr_18 0.954 0.175 0.315 0.000 96.90 0.000 100.00 fClaHyb_african_1_Chr_19 0.998 0.000 0.294 0.033 97.10 0.000 96.79 fClaHyb_african_1_Chr_20 0.995 0.056 0.148 0.000 98.53 0.000 100.00 fClaHyb_african_1_Chr_21 0.986 0.033 0.131 0.098 98.70 0.000 90.65 fClaHyb_african_1_Chr_22 0.992 0.067 0.134 0.034 98.67 0.000 96.70 fClaHyb_african_1_Chr_23 0.998 0.058 0.186 0.038 98.16 0.000 96.23 fClaHyb_african_1_Chr_24 0.999 0.077 0.232 0.000 97.71 0.000 100.00 fClaHyb_african_1_Chr_25 0.989 0.039 0.118 0.000 98.83 0.000 100.00 fClaHyb_african_1_Chr_26 0.994 0.057 0.324 0.000 96.81 0.000 100.00 fClaHyb_african_1_Chr_27 0.980 0.042 0.166 0.000 98.35 0.000 100.00 fClaHyb_african_1_Chr_28 0.995 0.000 0.297 0.000 97.07 0.000 100.00 76 Table 8 Bighead subgenome structural accuracy metrics from CRAQ: S-AQI, R-AQI, and heterozygous regions (Avg.CSH/Avg.CRH). Scaffold_name Mapping.rate Avg.CRH Avg.CRE Avg.CSE Regional-AQI Avg.CSH Structural-AQI Sub-Genome_Chromosome ONT/PE (%) Count Count Count Accuracy (%) Count Accuracy (%) Genome ( 55 Chromosomes) >97% 0.046 0.270 0.021 97.33 0.001 97.92 fClaHyb_bighead_1_Chr_01 0.981 0.000 0.329 0.000 96.76 0.000 100.00 fClaHyb_bighead_1_Chr_02 0.992 0.046 0.148 0.023 98.53 0.000 97.75 fClaHyb_bighead_1_Chr_03 0.980 0.050 0.276 0.000 97.28 0.000 100.00 fClaHyb_bighead_1_Chr_04 0.990 0.000 0.370 0.026 96.36 0.000 97.40 fClaHyb_bighead_1_Chr_05 0.972 0.053 0.360 0.160 96.46 0.000 85.22 fClaHyb_bighead_1_Chr_06 0.934 0.000 0.497 0.027 95.15 0.000 97.33 fClaHyb_bighead_1_Chr_07 0.990 0.030 0.403 0.030 96.05 0.000 97.06 fClaHyb_bighead_1_Chr_08 0.987 0.000 0.429 0.044 95.80 0.000 95.66 fClaHyb_bighead_1_Chr_09 0.986 0.051 0.422 0.067 95.87 0.000 93.47 fClaHyb_bighead_1_Chr_10 0.982 0.034 0.460 0.034 95.50 0.000 96.65 fClaHyb_bighead_1_Chr_11 0.961 0.042 0.382 0.042 96.26 0.000 95.92 fClaHyb_bighead_1_Chr_12 0.946 0.194 0.606 0.069 94.12 0.000 93.38 fClaHyb_bighead_1_Chr_13 0.989 0.041 0.509 0.081 95.04 0.000 92.19 fClaHyb_bighead_1_Chr_14 0.982 0.080 0.544 0.000 94.71 0.044 100.00 fClaHyb_bighead_1_Chr_15 0.922 0.136 0.506 0.039 95.07 0.000 96.19 fClaHyb_bighead_1_Chr_16 0.993 0.040 0.240 0.000 97.63 0.000 100.00 fClaHyb_bighead_1_Chr_17 0.985 0.042 0.793 0.042 92.38 0.000 95.87 fClaHyb_bighead_1_Chr_18 0.977 0.000 0.474 0.000 95.37 0.000 100.00 fClaHyb_bighead_1_Chr_19 0.992 0.000 0.341 0.000 96.65 0.000 100.00 fClaHyb_bighead_1_Chr_20 0.971 0.000 0.236 0.062 97.66 0.000 93.96 fClaHyb_bighead_1_Chr_21 0.977 0.064 0.591 0.000 94.26 0.000 100.00 fClaHyb_bighead_1_Chr_22 0.935 0.102 0.406 0.044 96.02 0.000 95.74 fClaHyb_bighead_1_Chr_23 0.990 0.000 0.303 0.032 97.02 0.000 96.87 fClaHyb_bighead_1_Chr_24 0.992 0.303 0.608 0.000 94.10 0.000 100.00 fClaHyb_bighead_1_Chr_25 0.990 0.062 0.350 0.025 96.56 0.000 97.53 fClaHyb_bighead_1_Chr_26 0.991 0.050 0.378 0.025 96.29 0.000 97.51 fClaHyb_bighead_1_Chr_27 0.979 0.033 0.447 0.000 95.63 0.000 100.00 77 GENOME ASSEMBLY, REVERSE VACCINOLOGY, AND QUALITY BY DESIGN — Streptococcus iniae Introduction 1. Streptococcus iniae as a Major Aquaculture Pathogen Streptococcus iniae, first identified in an Amazon River dolphin (Inia geoffrensis) in the 1970s (Pier and Madin,1976), has emerged as a major bacterial pathogen in global aquaculture. The pathogen causes annual losses exceeding USD 100 million worldwide (Shoemaker et al.,2001), with mortality rates of 30–80% during outbreaks (Chen et al.,2012) and cumulative mortality up to 70% after three months in some species (Mmanda et al.,2014). Detected across all continents (Baiano and Barnes, 2009;Mishra et al.,2018), S. iniae affects over 27 fish species (Agnew and Barnes, 2007), particularly in intensive aquaculture systems where high stocking densities and environmental stressors facilitate rapid disease transmission (Chen et al.,2012). Economically important species including Asian seabass (Lates calcarifer), tilapia (Oreochromis spp.), and catfish (Siluriformes spp.) are particularly susceptible (Azmai and Saad,2011;Nawawi et al.,2008). While rare, zoonotic transmission to humans handling infected fish has been documented (Facklam et al.,2005). Figure 26 Phase contrast micrograph of S. iniae strain QMA0076, showing characteristic spherical morphology. Credit: Baiano et al. (2008). 78 2. Taxonomic Classification and Molecular Characteristics S. iniae belongs to the phylum Bacillota (formerly Firmicutes), comprising Grampositive bacteria with thick peptidoglycan cell walls. Within the genus Streptococcus, S. iniae shares molecular mechanisms with human pathogens S. pyogenes and S. pneumoniae, particularly the sortase A pathway for cell wall protein anchoring. Taxonomic hierarchy: Domain: Bacteria Phylum: Bacillota Class: Bacilli Order: Lactobacillales Family: Streptococcaceae Genus: Streptococcus Species: S. iniae Key virulence factors identified include M-like protein (SimA), hyaluronidase (Hyl), enolase (eno), and glyceraldehyde-3-phosphate dehydrogenase (GAPDH). Notably, GAPDH localizes to the outer membrane despite its primary glycolytic function, enhancing host cell adherence and immune evasion. Sortase A (SrtA) anchors these surface proteins to the cell wall through a conserved mechanism shared with other pathogenic streptococci. Despite these molecular insights, knowledge gaps remain regarding genomic variability, horizontal gene transfer, and antimicrobial resistance mechanisms. 3. Current Disease Management Challenges Disease control in aquaculture relies primarily on broad-spectrum antibiotics (Schar et al.,2020,2021), with vaccines serving as complementary prophylactic measures. However, vaccine adoption remains limited, particularly for low-value fresh- 79 water species. In Thailand, no commercial vaccines exist for S. iniae in barramundi (Kayansamruaj et al.,2020). The economic barrier is substantial: vaccines cost approximately 100× more than antibiotics (Hoelzer et al.,2018;Pham et al.,2015), hindering adoption despite their potential to reduce antibiotic dependence and environmental impact. Existing whole-cell or polysaccharide-based vaccines provide only short-term protection (4–6 months) before losing efficacy due to antigenic variation (Eldar et al., 1997;Eyngor et al.,2008;Tanpichai et al.,2023). Capsular polysaccharides, traditionally considered key protective antigens, prove unstable targets as S. iniae rapidly alters surface structures under immune pressure. Moreover, genes responsible for capsule production belong to the accessory pangenome and are inconsistently present across strains (Millard et al.,2012), limiting broad protection. Traditional diagnostic methods (bacterial culture, ELISA, histopathology) lack specificity, while molecular approaches (PCR, MLST) can differentiate S. iniae from other streptococci (Glazunova et al.,2009) but require whole-genome sequencing for accurate genotyping. 4. Study Objectives Developing effective vaccines for S. iniae requires addressing both biological and manufacturing challenges. Candidate antigens must be conserved across strains, immunogenic, and compatible with scalable production processes—particularly crucial for lowto medium-value aquaculture species where vaccines must be cost-competitive with antimicrobial treatments. The Quality by Design (QbD) framework (Yu et al., 2014), established in pharmaceutical manufacturing, provides systematic methodology to connect antigen properties with downstream manufacturability, yet its application to reverse vaccinology remains largely unexplored. This study aimed to: (1) sequence and assemble complete genomes of pathogenic 80 S. iniae isolated from diseased Asian seabass in Thai aquaculture systems; (2) apply integrated reverse vaccinology and QbD approaches to identify conserved, immunogenic antigens; and (3) classify candidates by manufacturability criteria aligned with multiple production platforms. Using isolate SIKU01 for complete genome assembly and in silico screening, we developed an ”omics-to-manufacturing” pipeline that provides a practical framework for developing cost-effective vaccines against S. iniae and other economically important aquaculture pathogens. Methods 1. Bacterial Isolation and Cultivation Five Streptococcus iniae pathogens (designated SIKU01–SIKU05) were isolated from diseased Asian seabass (Lates calcarifer) exhibiting clinical signs of streptococcosis disease. All isolates originated from a single farm located in Chachoengsao Province, Thailand. Pure bacterial cultures were maintained in Tryptic Soy Broth (TSB) containing 25% glycerol at -80°C for long-term storage and subsequent molecular confirmation and genomic DNA extraction. 2. Genomic DNA Extraction and Sequencing Genomic DNA extraction from S. iniae cells involved cultivating bacteria in TSB to logarithmic phase, followed by cell harvesting and lysis. DNA purification was performed using a column-based method with silica membranes (QIAamp DNA Mini Kit, Qiagen, Germany) according to the manufacturer’s protocol with minor modifications for Gram-positive bacteria. The bacterial cell pellets were subjected to enzymatic lysis using hen lysozyme (20 mg/mL) and overnight incubation at 37°C to digest the peptidoglycan layer, which is particularly thick in Gram-positive bacteria. Following enzymatic treatment, a detergent-based lysis buffer containing proteinase-K was added to complete cell disruption and protein degradation. 81 The quality and quantity of extracted DNA were assessed using spectrophotometry (NanoDrop2000, Thermo Fisher Scientific, USA). High Molecular Weight (HMW) DNA integrity was verified by agarose gel electrophoresis. DNA library preparation and quality control were performed according to standard protocols for Illumina sequencing. Sequencing was performed on the Illumina HiSeq 2000 platform (Illumina, USA) using paired-end 151 bp reads, generating approximately 3 million paired-end reads per sample. 3. Genome Assembly and Quality Control 3.1 De Novo Assembly Initial quality assessment of raw Illumina paired-end reads was performed using FastQC v0.11.9 (Andrews,2010) to evaluate read quality, GC content, and adapter contamination. SeqPurge v2019 (Sturm et al.,2016) was used to trim adapters and lowquality bases using a sliding window approach, removing bases with quality scores below Q20. De novo assembly was conducted with Unicycler v0.4.9 (Wick et al.,2017), which employs SPAdes v3.15.2 (Bankevich et al.,2012) as its assembly algorithm. Resulting assembly graphs were visualized using Bandage v0.8.1 (Wick et al.,2015) to assess connectivity and identify potential misassemblies. Assembly quality metrics were evaluated using QUAST v5.0.2 (Quast et al.,2013). 3.2 Reference-Guided Assembly Multiple reference strains (QMA0248, 89353, SF1) were obtained from NCBI RefSeq (Supplementary Data 13). Reference-guided assembly was performed by mapping trimmed reads to each reference strain genome using Bowtie2 v2.4.4 (Langmead and Salzberg,2012) with the (’--fast-local’) option. To assess synteny and structural variations across strains, assembly contigs were aligned against all reference genomes using ProgressiveMauve v2.4.0 (Darling et al.,2010). 88 of sequence variability, conservation, and structure-based mapping. All MSAs were analyzed in R (Supplementary Code 12) using a custom script, which calculated normalized Shannon entropy (H.norm) at both codon and amino acid levels via the Bio3D R package (Grant et al.,2006). For each orthogroup, H.norm quantified sequence conservation on a continuous scale from 0 (fully conserved) to 1 (maximally variable). Antigens with median H.norm ≤0.05 were classified as highly conserved, while those above 0.90 were excluded as excessively variable. These conservation metrics, together with gene-carriage data from Panaroo, defined the gene-carriage and sequence-variability components of the Quality-by-Design (QbD) framework (PreM1 and M1 matrices) (Supplementary Data 14). 11. Protein Structure Prediction and Visualization Three-dimensional structures of prioritized S. iniae antigens (GAPDH, enolase, GroEL, Sortase A) were predicted using AlphaFold2 (Jumper et al.,2021) and crossvalidated against available homologous templates in the Protein Data Bank (PDB) (Berman, 2000). Predicted PDB files were examined in UCSF ChimeraX v1.6 (Meng et al.,2023) for fold integrity, residue geometry, and solvent accessibility. To quantify site-specific structural conservation, we implemented a per-residue identity analysis (Supplementary Code 12) that scanned each aligned FASTA previously produced by MAFFT. For every alignment, the script compared all sequences to the reference (first entry) and recorded, for each residue position, the number and percentage of identical residues across all genomes. These per-site identity profiles were mapped onto AlphaFold2 models in ChimeraX using custom color scripts to generate gradient-based visualizations where blue indicated highly conserved residues and red denoted variable positions. This approach provided a direct link between MSA entropy (H.norm) values and 3D spatial conservation, highlighting structurally stable and immunologically relevant regions. 89 Epitope localization was performed using (Supplementary Code 12), a custom Python script which scanned PDB chains for exact matches to epitope peptide sequences identified by IEDB mapping. For each hit, the script reported the model ID, chain ID, and PDB residue range corresponding to the matched peptide, enabling precise overlay of Band T-cell epitopes onto predicted antigen structures. Surface visualization and electrostatic potential mapping were carried out in ChimeraX using ”surface color” and ”cartoon” representations to assess accessibility and topological context. 12. Codon Adaptation Index (CAI) and Expression Compatibility Codon usage frequencies for target organisms were obtained from the Kazusa Codon Usage Database (Holcomb et al.,2019), available at https://www.kazusa.or.jp/codon/. Species-specific codon frequency tables were downloaded (Supplementary Data 15), providing frequencies per thousand codons for each organism’s coding sequences. The Codon Adaptation Index (CAI) quantifies the degree of preference for synonymous codons in a given organism, where values range from 0 (least preferred) to 1 (most preferred). For each amino acid family, the relative adaptiveness of codon iwas calculated as wi=fi/fmax, where firepresents the frequency per thousand of codon ifor its amino acid and fmax represents the frequency per thousand of the most frequently used codon for that amino acid. For example, proline codons in Oreochromis niloticus were calculated as follows: CCG had wi=7.39/16.53 =0.447, CCA had wi=14.59/16.53 =0.883, CCT had wi=16.53/16.53 =1.000 (optimal), and CCC had wi=14.91/16.53 =0.902. The value 16.53 represents the highest frequency among all proline codons, making CCT the optimal reference codon. For complete gene sequences, CAI was computed as the geometric mean of rel- 90 ative adaptiveness values using the formula: CAI =(L ∏ k=1 wk)1/L where the product encompasses all Lsense codons in the gene sequence. Stop codons (UAA, UAG, UGA) and the single methionine codon (AUG) were excluded from CAI calculations as they lack synonymous alternatives for optimization. 13. Quality by Design (QbD) Framework To implement QbD principles effectively within the reverse vaccinology framework, specific criteria were established to guide systematic antigen selection. These criteria served as the foundation for developing a quantitative scoring matrix computed in a custom script (Supplementary Code 12) that ensures both immunological efficacy and manufacturing feasibility. 13.1 QbD Design Space: Gene Carriage (Matrix Pre-M1) The gene carriage design space defines the level of genomic conservation of an antigen across the S. iniae pangenome. This Critical Quality Attribute (CQA) prioritizes stable and widely distributed antigens to prevent loss of vaccine efficacy in heterologous strains. Based on the pan-genome presence/absence matrix, the core genome (present in ≥99% of isolates) was retained whereas soft-core genome (95–98%), shell (15–94%), and cloud (<15%) genes were excluded from further evaluation. This filtering yielded 1,538 core and 110 soft-core genes (Supplementary Data 15), forming the foundation for downstream antigen variability and QbD scoring due to their broad distribution and genetic stability. 91 13.2 QbD Design Space: Amino Acid and Nucleotide Variability (Matrix 1) The normalized Shannon entropy (H.norm)defined the antigen variability design space, quantified from per-cluster MSAs. Each residue received two complementary measures: (1) H.norm from sequence alignments (0 = fully conserved, 1 = maximally variable) and (2) percent identity (%ID) relative to the amino acid reference sequence (SIKU01). These values were averaged per gene to yield gene-wise conservation indices subsequently integrated into the QbD variability matrix (M1) (Supplementary Data 15). Genes with median H.norm ≤0.05 and mean %ID ≥95% were prioritized as structurally and functionally stable antigens. Residues with low entropy and high surface exposure (as confirmed by ChimeraX mapping) defined the structural conservation-driven QbD design space, used to constrain antigen selection and ensure manufacturing consistency across S. iniae lineages. 13.3 QbD Design Space: General Physicochemical and Expression Characteristics (Matrix 1) The protein length design space favors candidates with 100 or more amino acids to ensure sufficient epitope diversity while avoiding overly short sequences that lack functional relevance. Proteins containing 50–99 amino acids are classified as too short and excluded due to instability concerns, while those with fewer than 50 amino acids face exclusion for insufficient immunogenic potential. Molecular weight constraints target the 20–80 kDa range to balance immune system interaction with proper folding capabilities, with the optimal 20–50 kDa range receiving highest priority for E. coli expression systems. The isoelectric point (pI)design space recognizes multiple optimal ranges depending on purification strategy. Proteins with pI values between 4.0–7.5 receive priority due to superior solubility characteristics and broad purification platform compatibility. The pI range of 7–9 provides negative charge under physiological conditions, 92 facilitating purification via histidine tag chromatography or silicon dioxide-based filters as specified in the foundational selection criteria. Additionally, proteins with pI values between 2–4 offer specialized advantages for silica-based purification systems through enhanced electrostatic interactions. Hydrophobicity design space assessment relies on the GRAVY index, which prioritizes proteins within the -0.5 to +0.5 range to minimize aggregation tendencies and support aqueous solubility throughout the manufacturing process. Protein stability design space relies on the instability index (II) calculated from primary sequence data, with values of 40 or lower indicating favorable in vivo stability characteristics. Literature support provides empirical validation through PubMed literature and PMID, with proteins having prior experimental mentions, particularly as vaccine candidates or virulence factors, receiving graduated positive scoring based on evidence strength. 13.4 QbD Design Space: Immunogenicity (Matrix 1) Epitope prediction analysis identifies the presence of known sequences that stimulate T and B lymphocytes, considering the 30% similarity between fish and human immunoglobulins in the scoring framework. Proteins predicted to contain both Tand B-cell epitopes (inferred from IEDB) receive maximum immunological relevance scoring, while those lacking predicted epitopes face penalty assessment. 13.5 QbD Design Space: Host-specific Expression (Matrix 1 – CAI) In vitro expression efficiency design space prediction utilizes the Codon Adaptation Index (CAI) to assess translational compatibility with E. coli systems, following the principle of optimal codon usage for maximizing protein expression. Top- 93 quartile CAI values receive positive weighting, while bottom-quartile scores indicate potential expression difficulties (Supplementary Data 15–15). 13.6 QbD Design Space: Matrix 2 – Secondary Selection In Matrix 2, candidate antigens from Matrix 1 were screened multiple times for compatibility with specific downstream purification modalities by aligning their intrinsic molecular descriptors with platform-specific Critical Process Parameters (CPPs). Five distinct sets of design spaces were evaluated: silica affinity, cellulose affinity, ion exchange chromatography (anion exchange (AEX) and cation exchange (CEX)), and plasmid DNA (pDNA) expression. This pre-downstream design space cross-references in silico Critical Quality Attributes (CQAs) with platform-specific CPPs to ensure process– antigen compatibility before physical development for maximum cost-efficiency postbiomanufacturing (Supplementary Data 15–15). Silica Affinity Purification: Silica-based purification technologies offer ubiquity, low cost, and adaptability across laboratory and industrial bioprocessing applications. Critical material attributes (CMAs) include silanol density, surface topology, and functionalization chemistry, while key CPPs encompass pH, ionic strength, and buffer composition (Supplementary Data 15). Affinity tags with high silica specificity, including Si-tag, SB7, Car9, R5, and the synthetic octapeptide (RH)4, composed of four repeating Arginine-Histidine units, interact through electrostatic and hydrogen-bonding mechanisms with silanol-rich matrices. Proteins presenting net positive charge at pH 7.4 demonstrate enhanced adsorption due to electrostatic attraction to partially deprotonated silanol groups (≡Si–OH → ≡Si–O-). QbD selection criteria required pI > 4 and a surface-exposed cationic patch at loading pH (offered by the fusion of a silica binding peptide short tag). Elution is achieved by increasing pH from 7.4 to 8.5, expanding silanol deprotonation and reducing hydrogen-bond donor capacity. Short tag lengths ( 7–20 aa) were favored to avoid steric hindrance and folding interference (Supplementary Data 15). 94 Cellulose Affinity Purification: Cellulose-based purification uses cellulosebinding modules (CBMs) such as CBM3, CBM9, or CelD, which recognize repeating β-(1→4)-D-glucopyranose units through hydrogen bonding with hydroxyl groups and hydrophobic stacking with planar glucan rings. Key CPPs include buffer pH (6.0–8.5), salt concentration, cellulose type, and contact time. Since binding is mediated by the CBM domain rather than antigen charge properties, no pI requirement was applied. However, CBMs add substantial polypeptide segments (30–200 aa), necessitating length constraints. Antigens were preferentially kept ≤400 aa to ensure total fusion constructs remained ≤430–630 aa, limiting metabolic burden, misfolding risk, and aggregation while preserving tag accessibility (Supplementary Data 15). Ion Exchange Chromatography (IEX): Ion exchange utilizes the net charge of a protein to separate it from other proteins. Based on the type of resins and the charge of the protein, anion exchange chromatography or cation exchange chromatography techniques may be used. Anion exchange chromatography was modeled on quaternary amine functional groups (e.g., –N+(CH3)3) binding negatively charged proteins through electrostatic interaction with deprotonated carboxyl groups. The QbD filter retained proteins with pI ≤ 6.5 and net charge ≤ -1 at pH 7.4 (Supplementary Data 15). Cation exchange chromatography employed sulfonate ligands (–SO3-) that bind protonated amines from Lys, Arg, and His side chains. Selection criteria favored proteins with pI ≥ 8.0 and net charge ≥ +1 at pH 7.4, ensuring positive charge under working conditions for effective resin binding (Supplementary Data 15). Plasmid DNA Expression: Plasmid DNA expression was treated as a distinct design space, optimizing antigens for high-yield, high-fidelity eukaryotic expression in DNA vaccine formats. Sequence-level CQAs included GC3 enrichment to improve mRNA stability and translation efficiency, ORF length threshold (set at the median coding sequence length of core genes ≈2.2 kb), and complete absence of internal Type IIS restriction sites (BsaI, BsmBI, Eco31I) to ensure seamless Golden Gate Assembly cloning compatibility (Supplementary Data 15–15). 95 14. Statistical Analysis All statistical analyses were performed in R version 4.2.0. Spearman’s rank correlation coefficients ( ρ ) were calculated using the cor() function along with option method="spearman". Statistical significance of correlations was assessed using the rcorr() function from the Hmisc package, with significance set at p< 0.05. For the correlation between hydrophobicity and aliphatic index ( ρ =0.82,p<10−15), significance was calculated using cor.test(). While no formal corrections for multiple testing were applied in this exploratory analysis, we note that key correlations (e.g., hydrophobicity-aliphatic index, ρ =0.82) would remain significant even after Bonferroni correction. Kernel density estimation was performed using the density() function with 512 interpolation points. Density contours in scatter plots were generated using ggplot2’s stat_density_2d() with default bandwidth selection. All histograms used 50 bins for consistency across distributions. Normalized Shannon entropy for sequence conservation was calculated for proteins present in at least 95% of genomes as H=−Σ(pi×log2(pi)), where piis the frequency of amino acid iat each position. Normalized entropy (H.norm) was calculated as H/log2(20)to scale values between 0 (fully conserved) and 1 (maximum variation), using the Bio3D R package (Grant et al.,2006). Quartile thresholds for codon adaptation index (CAI) and GC3 content were calculated using quantile() at probabilities 0.25 and 0.75. All thresholds were calculated on the Pre-M1 filtered dataset (n=1,374 proteins) to ensure consistency across platforms. For multimodal distribution detection, hierarchical density-based clustering (HDBSCAN) was applied with minPts=40 using the dbscan package. Statistical summaries (mean, median, standard deviation) were calculated for all biophysical properties. The dataset M0 comprising the full SIKU01 proteome (N= 1,855 proteins) (Supplementary 96 Data 15) and the Pre-M1 filtered subset (n= 1,374 proteins) (Supplementary Data 15) provided adequate statistical power for correlation analyses. Results 1. Genome Assembly and Annotation Results 1.1 Genome Assembly of Streptococcus iniae strain SIKU01 We assembled a complete, circular genome of S. iniae strain SIKU01 (2.09 Mb, 37% GC) with high completeness and no plasmids detected (Figs. 27a–b); Appendix Figure 33). Whole-genome comparisons with reference S. iniae strains QMA0248 (Alsheikh-Hussain et al.,2022), 89353 (Gong et al.,2017), SF1 (Zhang et al.,2014b), and LSSM211007Si confirmed strong macrosyntenic conservation (Fig. 27c–d). We also analyzed synteny relative to the original Amazon dolphin isolate QMA0141 from 1976 (Pier and Madin,1976) (Appendix Figure 34). 1.2 Functional Annotation of Streptococcus iniae strain SIKU01 Annotation of strain SIKU01 identified 2,004 genes, including 1,855 proteincoding sequences and 77 RNAs (Supplementary Data 13). Functional classification with InterProScan (Quevillon et al.,2005), KEGG Mapper (Kanehisa and Sato,2020), and Gene Ontology (GO) (Ashburner et al.,2000) (Supplementary Data 13–13) revealed a broad repertoire of metabolic enzymes, transporters, DNA-repair proteins, and cellsurface factors, along with subsets of stress-response proteins, toxins, and antimicrobial resistance determinants (Figure 27e). This comprehensive genome annotation provided the foundation for proteome-wide reverse vaccinology and subsequent Qualityby-Design (QbD) manufacturability screening of candidate vaccines (Supplementary Data 15–15). 97 Figure 27 Complete genome assembly and functional annotation of Streptococcus iniae strain SIKU01. (a) Genome assembly and reference-guided workflow from raw reads to validated circular chromosome. (b) Assembly and annotation statistics, including genome size, GC content, gene counts, and RNA features. (c) Comparative macrosynteny of SIKU01 against public S. iniae reference strains, showing conserved genomic architecture. (d) Metadata of strain SIKU01 and related isolates, including collection date, country, and host species, used in hybrid reference-guided assembly. (e) Functional annotation of the SIKU01 proteome, including KEGG Mapper categories, Gene Ontology subcellular localization, and InterProScan domain assignments. 104 tion were sufficiently basic to survive cation exchange (CEX). This was reflected in the survivor pools (Fig. 3.5), with 98 proteins retained by AEX and only 20 by CEX. Bufferspecific ranges (Figs. 3.5–3.5) confirmed that AEX survivors were stable across nearly all chemistries (n = 82–98 proteins), while CEX buffers consistently yielded the same restricted set (n = 20 proteins). The full antigen lists for AEX and CEX are provided in (Supplementary Data 15-15). Affinity purification routes imposed distinct orthogonal constraints: cellulose excluded proteins longer than 400 aa (Fig. 3.5), because the CBM fusion tag is itself large and places a metabolic burden on recombinant expression; minimizing the size of the fused antigen reduces energy demand and steric hindrance, thereby improving folding and yield. In contrast, silica affinity (Fig. 3.5) was governed by electrostatic interactions with surface silanol groups: proteins with pI between 7–9 acquire a net positive charge at physiological pH, enabling stable adsorption to negatively charged silica. These filters reduced the pools to 100 cellulose-compatible and 49 silica-compatible proteins, detailed in (Supplementary Data 15-15). For plasmid DNA (pDNA) platforms, Critical Quality Attribute (CQA) filters were applied separately on a per-host basis. In O. niloticus (Figs. 3.5–3.5), genes and their product were first assessed for CAI ≥ Q3 ( 0.60), ensuring codon usage was well adapted to the host translation machinery; this retained 188 sequences. Applying GC3 ≥ Q3 ( 0.32) halved the pool to 95, as high G/C content at the third codon position improves mRNA stability and reduces metabolic stress. A further antigen length cutoff (≤ 2,200 nt) yielded 57 candidates, favoring shorter constructs with reduced transcriptional burden. Finally, genes containing internal Type IIS restriction sites (e.g., BsaI, SapI) were excluded because these enzymes cut outside their recognition sites, disrupting modular DNA assembly workflows by preventing a one-step cloning in a Golden Gate Assembly. The final Nile tilapia pDNA pool remained at 57 survivors. In D. rerio (Figs. 3.5–3.5), the thresholds were slightly higher (CAI ≥ Q3 0.70, GC3 ≥ Q3 0.31). Here, 186 proteins passed the CAI filter, 82 remained after the 105 GC3 filter, 66 after the length filter, and 65 after the Type IIS removal filter. Comprehensive candidate sets for both pDNA platforms (Nile tilapia and zebrafish) are available in (Supplementary Data 15-15). Overall, the reverse vaccinology–QbD funnel outputs platform-ready shortlists: 98 AEX, 20 CEX, 100 cellulose, 49 silica, and 57–65 pDNA antigens for immediate downstream development. 106 107 Figure 28 Manufacturability design spaces across purification platforms. (a) Ion-exchange charge space for protein candidates: pI vs. net charge (z) at pH 7. Shaded bands mark the AEX and CEX inclusion regions; vaccine-referenced antigens are labeled. (b) Survivors after M2 filtering by platform; bars show unique genes per route with counts and percentages. Only candidates fitting the biophysiochemical criteria for their respective purification pathway are displayed in color bars. This step reflects manufacturability constraints and downstream process optimization in the QbD framework. (c–d) Buffer operating pH ranges for AEX (c) and CEX (d). Short tick marks indicate the working pH (range midpoint). (e–f) Protein-route survivors per buffer evaluated at the midpoint: AEX rule = base gate passed, net charge (z) at pH 7 ≤0, and pI ≤pHmid; CEX rule = base gate passed, net charge (z) at pH 7 ≥0, and pI ≥pHmid.(g) Cellulose: antigen length distribution for M2 survivors; dashed line at 400 aa. (j) Silica binding peptides favor proteins with moderate pI (7–9) and positive net charge at pH > 7. (h, k) pDNA (M2) scatterplots by host: tilapia (h) and zebrafish (k), showing CAI vs. GC3; dashed lines mark dataset thresholds. (i, l) pDNA-route survivors passing each manufacturability gate for tilapia (i) and zebrafish (l): CAI ≥threshold, GC3 ≥threshold, length ≤median, and no Type IIS sites (count). 108 4. Cross-Validation with Literature We validated our shortlisted antigens against previously tested S. iniae vaccine targets reported in multiple hosts, including Channel catfish (Ictalurus punctatus) (Wang et al.,2016b), Nile tilapia (Oreochromis niloticus) (Kayansamruaj et al.,2017), Olive flounder (Paralichthys olivaceus) (Sheng et al.,2018a,2023), Zebrafish (Danio rerio) (Membrebe et al.,2016), Mouse (Mus musculus) (Wang et al.,2015a), and Turbot fish (Scophthalmus maximus) (Zhang et al.,2014a) (Table 11). Among the top-ranked candidates consistently found across purification routes, antigens with demonstrated in vivo protection (enolase, GAPDH, and GroEL) were recovered, supporting the predictive accuracy of the QbD framework. To further assess their suitability, we modeled GAPDH and enolase, two representative antigens, both well-established in the literature. Structural predictions generated using AlphaFold2 (Yang et al.,2023) for S. iniae SIKU01 confirmed highly conserved folds and surface-exposed loops (Wang et al.,2017). Several epitope-containing regions identified through reverse vaccinology were surface-accessible ((Figures 29a– c), highlighted in yellow), consistent with prior reports (Gent et al.,2024). These epitopes (Table 9) are suitable for direct incorporation into subunit vaccines or for the design of chimeric multi-epitope constructs (Pumchan et al.,2020), further validating their suitability as vaccine targets. Both enolase (eno) and GAPDH (gap) displayed epitoperich regions spatially separated from catalytic or hypervariable sites, reinforcing their accessibility and stability as broad-spectrum candidates. 109 Figure 29 Structural mapping of variability, epitopes, and active sites in two model Streptococcus iniae vaccine candidates. (a) Shannon entropy profiles show that two model and well-known antigens of the scientific literature, GAPDH (gap) and enolase (eno), are largely conserved with limited variable regions. (b) GAPDH (336 aa, UniProt Q7BB80) carries a predicted N-terminal epitope (MVVKVGINGFGRIGRLAFRRIQ) positioned near, but not overlapping with, the active site and adjacent to a region of high nucleotide-level (codon-derived) variability (1–89 % variation). (c) Enolase (435 aa, UniProt T1TFA0) contains a predicted epitope (RAAADYLEVPLYNYLG) located opposite the active site and spatially separated from five nucleotide-driven hypervariable surface regions (HR1–HR5), indicating stability and accessibility as a vaccine target. Variability values are based on codon-level polymorphisms within the core-genome alignment and projected onto corresponding amino acid positions in the 3-D models for visualization. 110 DATA AVAILABILITY AND NCBI SUBMISSIONS Bighead Catfish The final diploid genome assembly was screened for contamination using the NCBI Foreign Contamination Screen (FCS) (Astashyn et al.,2024) and submitted following the Vertebrate Genome Project (VGP) naming conventions (https://github.com/VGP/ vgp-assembly) (Rhie et al.,2021). The project is registered under NCBI BioProject PRJNA1132508, with BioSample accession SAMN41769988 corresponding to a Thai (male, adult) bighead catfish (isolate: CMAM; TaxID: 35657). Raw sequencing reads are deposited in the NCBI Sequence Read Archive (SRA): Nanopore (20% error) — SRR29723575, HiFi — SRR29723576, Hi-C (150PE) — SRR29723577, and Illumina (150PE) — SRR29723578. The final diploid assembly is available in GenBank under accession numbers JBLWMO000000000 (Haplotype 1) and JBLWMP000000000 (Haplotype 2). A complete dataset, including genome assemblies and supporting files, is permanently archived at Zenodo (10.5281/zenodo.14826875). F1 Hybrid Catfish The hybrid genome assembly (Clarias macrocephalus ×C. gariepinus) was processed and submitted using the same quality control and standardization workflow as described above. GenBank accessions are JBLWFY000000000.1 and JBLWFZ000000000.1, corresponding to the C. macrocephalus and C. gariepinus sub-genomes, respectively. Associated records are hosted under NCBI BioProject PRJNA1153495 and BioSample SAMN43395848 (TaxID: 1334085). Raw sequencing reads are Nanopore (SRR30599638), HiFi (SRR30599641), Illumina (SRR30599640), and Hi-C (SRR30599639). Additional Illumina datasets correspond to female C. macrocephalus (SAMN42503781) and male C. gariepinus (SAMN43548335). The complete F1 hybrid dataset, including assemblies and metadata, is archived at Zenodo (10.5281/zenodo.15269601). 111 Streptococcus iniae The Streptococcus iniae project is registered under NCBI BioProject PRJNA933632, with BioSample accession SAMN33244440. Raw sequencing reads are available in the NCBI Sequence Read Archive (SRA): Illumina (150PE), 10 FASTQ files across five datasets—SRR23406918 (SIKU01), and SRR23406921,SRR23406920,SRR23406919, and SRR23406922 for SIKU02–SIKU05. The annotated genome of SIKU01 is deposited in NCBI GenBank under accession CP121692.1. Supplementary annotation files, intermediate datasets, and related analysis materials are permanently archived at Zenodo (https://zenodo.org/records/17476104. All computational tools used in this study are publicly available, and all command-line parameters are specified in the Methods section. Custom scripts used for data processing, proteome annotation, and Qualityby-Design (QbD) manufacturability analysis are provided as Supplementary Code, available at Zenodo (see above) and in the Appendix chapter of this document. Associated Publications 1. Andres, Q. L. S., Singchat, W., & Srikulnath, K. (2025). Haplotype-Resolved Chromosome-scale Assembly of the Bighead Catfish (Clarias macrocephalus) Genome. https://doi.org/10.5281/zenodo.14826876 2. Andres, Q. L. S., Singchat, W., & Srikulnath, K. (2025). Dual Reference Genomes from F1 Hybrids: Phased Assembly of North African Catfish and Bighead Catfish with Hi-C Data. https://doi.org/10.5281/zenodo.15269601 3. ANDRES, Q. L. S., Uchuwittayakul, A., & Srisapoome, P. (2025). Complete genome and QbD-guided reverse vaccinology for Streptococcus iniae strain SIKU01. https://doi.org/10.5281/zenodo.15264953 112 DISCUSSION AND CONCLUSIONS Genomic Insights and Technical Achievements 1. Catfish Genome Assembly This study presents the first haplotype-resolved genome assemblies of both bighead catfish (Clarias macrocephalus) and North African catfish (C. gariepinus), achieved through sequencing a single F1 hybrid offspring. Despite using identical pipelines and data sources, the bighead catfish subgenome exhibited lower quality metrics (higher base error rates, lower QV) compared to the African subgenome—likely reflecting differences in transposable element content and the inherent assembly challenges of repeatrich teleost genomes. Technical validation followed VGP benchmarks using k-mer sizes of 21 (QV analysis), 28 (alignment applications), and 31 (GenomeScope 2.0). While contamination was detected in raw reads by Mash, it was successfully excluded from final assemblies (confirmed by FCS-NCBI and Mash screening). Primary limitations included low sequencing depth (PacBio HiFi <13X, ONT <36X), suboptimal ONT library quality (longest read <134kb), and software-related challenges including tool maintenance and quality control gaps. The F1 hybrid genome assembly recovered complete genomic sets from both parental species at exceptional quality (QV50-QV60): 27 pseudochromosomes from bighead catfish and 28 from North African catfish, totaling 1.8 Gb with median QV of 55. The assembly achieved 99.35% completeness (21-mer analysis) with over 40 pseudochromosomes assembled to near telomere-to-telomere continuity. All data are available under NCBI BioProject PRJNA115349. 113 Streptococcus iniae Vaccine Development This study presents an integrated reverse vaccinology–Quality by Design (QbD) framework for S. iniae vaccines, addressing critical needs in aquaculture where annual streptococcal losses exceed USD 1 billion (Amillano-Cisneros et al.,2025;Defoirdt et al.,2011). By evaluating all 1,855 proteins from strain SIKU01 through manufacturability filters, we reduced the candidate pool by 74–95% before wet-lab validation while retaining previously validated protective antigens. The pipeline yielded platformspecific candidates: 98 for anion-exchange (Duong-Ly and Gabelli,2014), 100 for cellulose-affinity (Carrard et al.,2000), 49 for silica-affinity purification (Freitas et al., 2022), and 57–65 for plasmid DNA vaccines (Marillonnet and Grutzner,2020), including validated antigens enolase and GAPDH (62–80% RPS in previous trials). The bimodal pI distribution of the S. iniae proteome (peaks at 5.5 and 9.0) explains the five-fold higher capture rate of anion-exchange chromatography, providing practical guidance for purification strategy selection. DNA vaccines showed highest efficacy (80–95% RPS), while our E. coli expression filters addressed inclusion-body formation through hydrophobicity and stability constraints (Francis and Page,2010;Rosano and Ceccarelli,2014). To address S. iniae’s rapid antigenic variation, we restricted candidates to the conserved core genome (n=1,374) with high conservation thresholds (H.norm >0.98). The framework explicitly links antigen discovery to production feasibility, ensuring candidates remain cost-competitive with current antibiotic treatments while providing manufacturing flexibility through alternative affinity tags (Bachmann and Jennings,2010b;Woestenenk et al.,2004) and Golden Gate Assembly compatibility (Bird et al.,2022b;Marillonnet and Grutzner,2020). Key limitations include incomplete capture of teleost immune complexity through in silico prediction, uneven geographic representation in our 90-isolate pangenome, and the need for experimental validation of manufacturing yields. 120 I. Machol, E. S. Lander, A. P. Aiden and E. L. Aiden. 2017. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science. 356 (6333): 92–95. Duong, T.-Y. and K. T. Scribner. 2018. Regional variation in genetic diversity between wild and cultured populations of bighead catfish (Clarias macrocephalus) in the Mekong Delta. Fisheries Research. 207: 118–125. Duong-Ly, K. C. and S. B. Gabelli. 2014. Using ion exchange chromatography to purify a recombinantly expressed protein. Methods in Enzymology. 541: 95–103. Durand, N. C., M. S. Shamim, I. Machol, S. S. Rao, M. H. Huntley, E. S. Lander and E. L. Aiden. 2016. Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. Cell Systems. 3 (1): 95–98. Eisenstein, M. 2017. An ace in the hole for DNA sequencing. Nature. 550 (7675): 285–288. El Hilali, S. and R. R. Copley. 2023. macrosyntR: Drawing automatically ordered Oxford Grids from standard genomic files in R. Archive Ouverte HAL. Eldar, A., A. Horovitcz and H. Bercovier. 1997. Development and efficacy of a vaccine against Streptococcus iniae infection in farmed rainbow trout. Veterinary Immunology and Immunopathology. 56 (1-2): 175–183. Elfaitouri, A., B. Herrmann, A. Bolin-Wiener, Y. Wang, C. Gottfries, O. Zachrisson, R. Pipkorn, L. Ronnblom and J. Blomberg. 2013. Epitopes of microbial and human heat shock protein 60 and their recognition in myalgic encephalomyelitis. PLoS ONE. 8 (11): e81155. Ellinghaus, D., S. Kurtz and U. Willhoeft. 2008. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics. 9 (1). Emms, D. M. and S. Kelly. 2019. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biology. 20 (1). Ewels, P., M. Magnusson, S. Lundin and M. Käller. 2016. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics. 32 (19): 3047. Eyngor, M. and others. 2008. Emergence of novel Streptococcus iniae exopolysaccharideproducing strains following vaccination with nonproducing strains. Applied and Environmental Microbiology. 74 (22): 6892–6897. Facklam, R., J. Elliott, L. Shewmaker and A. Reingold. 2005. Identification and characterization of sporadic isolates of Streptococcus iniae isolated from humans. Journal of Clinical Microbiology. 43 (2): 933–937. 121 FAO, . 2020. The State of World Fisheries and Aquaculture 2020. FAO. Faust, G. G. and I. M. Hall. 2014. SAMBLASTER: fast duplicate marking and structural variant read extraction. Bioinformatics. 30 (17): 2503–2505. Ferraris, C. J. 2007. Checklist of catfishes, recent and fossil (Osteichthyes: Siluriformes), and catalogue of siluriform primary types. Zootaxa. 1418 (1): 1–628. Finn, R. D., J. Clements and S. R. Eddy. 2011. HMMER web server: interactive sequence similarity searching. Nucleic Acids Research. 39 (suppl): W29–W37. Flynn, J. M., R. Hubley, C. Goubert, J. Rosen, A. G. Clark, C. Feschotte and A. F. Smit. 2020. RepeatModeler2 for automated genomic discovery of transposable element families. Proceedings of the National Academy of Sciences. 117 (17): 9451–9457. Formenti, G., A. Rhie, B. P. Walenz, F. Thibaud-Nissen, K. Shafin, S. Koren, E. W. Myers, E. D. Jarvis and A. M. Phillippy. 2022. Merfin: improved variant filtering, assembly evaluation and polishing via k-mer validation. Nature Methods. 19 (6): 696–704. Fourie, K. and H. Wilson. 2020. Understanding GroEL and DnaK Stress Response Proteins as Antigens for Bacterial Diseases. Vaccines. 8 (4): 773. Francis, D. M. and R. Page. 2010. Strategies to optimize protein expression in E. coli. Current Protocols in Protein Science. (1): 5.24.1–5.24.29. Freitas, A. I., L. Domingues and T. Q. Aguiar. 2022. Bare silica as an alternative matrix for affinity purification/immobilization of his-tagged proteins. Separation and Purification Technology. 286: 120448. Gent, V., Y.-J. Lu, S. Lukhele and others. 2024. Surface protein distribution in Group B Streptococcus isolates from South Africa and identifying vaccine targets through in silico analysis. Scientific Reports. 14: 22665. Giovanni, A., Y.-Z. Shi, P.-C. Wang, M.-A. Tsai and S.-C. Chen. 2025. Recombinant C5a Peptidase and Formalin-Killed Cell: A Synergistic Vaccine Against Streptococcus iniae in Four-Finger Threadfin Fish (Eleutheronema tetradactylum). Journal of Fish Diseases. 48: e14154. Glazunova, O. O., D. Raoult and V. Roux. 2009. Partial sequence comparison of the rpoB, sodA, groEL and gyrB genes within the genus Streptococcus. INTERNATIONAL JOURNAL OF SYSTEMATIC AND EVOLUTIONARY MICROBIOLOGY. 59 (9): 2317–2322. Goel, M. and K. Schneeberger. 2022. plotsr: visualizing structural similarities and rearrangements between multiple genomes. Bioinformatics. 38 (10): 2922–2926. 122 Gong, H. and others. 2017. Complete Genome Sequence of Streptococcus iniae 89353, a Virulent Strain Isolated from Diseased Tilapia in Taiwan. Genome Announcements. 5 (8): e01524–16. Goubert, C., R. J. Craig, A. F. Bilat, V. Peona, A. A. Vogan and A. V. Protasio. 2022. A beginner’s guide to manual curation of transposable elements. Mobile DNA. 13 (1). Grant, B. J., L. Skjaerven and X. Q. Yao. 2021. The Bio3D packages for structural bioinformatics. Protein Science. 30 (1): 20–30. Grant, B. J. and others. 2006. Bio3D: an R package for the comparative analysis of protein structures. Bioinformatics. 22 (21): 2695–2696. Gu, X. H., D. L. Jiang, Y. Huang, B. J. Li, C. H. Chen, H. R. Lin and J. H. Xia. 2018. Identifying a Major QTL Associated with Salinity Tolerance in Nile Tilapia Using QTL-Seq. Marine Biotechnology. 20 (1): 98–107. Guan, D., S. A. McCarthy, J. Wood, K. Howe, Y. Wang and R. Durbin. 2020. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 36 (9): 2896–2898. Guarracino, A., S. Heumos, S. Nahnsen, P. Prins and E. Garrison. 2022. ODGI: understanding pangenome graphs. Bioinformatics. 38 (13): 3319–3326. Guruprasad, K., B. V. Reddy and M. W. Pandit. 1990. Correlation between stability of a protein and its dipeptide composition: a novel approach for predicting in vivo stability of a protein from its primary sequence. Protein Engineering. 4 (2): 155–161. Harrison, P. W., M. R. Amode, O. Austine-Orimoloye, A. G. Azov, M. Barba, I. Barnes, A. Becker, R. Bennett, A. Berry, J. Bhai, S. K. Bhurji, S. Boddu, P. R. Branco Lins, L. Brooks, S. B. Ramaraju, L. I. Campbell, M. C. Martinez, M. Charkhchi, K. Chougule, A. Cockburn, C. Davidson, N. H. De Silva, K. Dodiya, S. Donaldson, B. El Houdaigui, T. E. Naboulsi, R. Fatima, C. G. Giron, T. Genez, D. Grigoriadis, G. S. Ghattaoraya, J. G. Martinez, T. A. Gurbich, M. Hardy, Z. Hollis, T. Hourlier, T. Hunt, M. Kay, V. Kaykala, T. Le, D. Lemos, D. Lodha, D. Marques-Coelho, G. Maslen, G. A. Merino, L. P. Mirabueno, A. Mushtaq, S. N. Hossain, D. N. Ogeh, M. P. Sakthivel, A. Parker, M. Perry, I. Piližota, D. Poppleton, I. Prosovetskaia, S. Raj, J. G. Pérez-Silva, A. I. A. Salam, S. Saraf, N. Saraiva-Agostinho, D. Sheppard, S. Sinha, B. Sipos, V. Sitnik, W. Stark, E. Steed, M.-M. Suner, L. Surapaneni, K. Sutinen, F. F. Tricomi, D. UrbinaGómez, A. Veidenberg, T. A. Walsh, D. Ware, E. Wass, N. L. Willhoft, J. Allen, J. Alvarez-Jarreta, M. Chakiachvili, B. Flint, S. Giorgetti, L. Haggerty, G. R. Ilsley, 123 J. Keatley, J. E. Loveland, B. Moore, J. M. Mudge, G. Naamati, J. Tate, S. J. Trevanion, A. Winterbottom, A. Frankish, S. E. Hunt, F. Cunningham, S. Dyer, R. D. Finn, F. J. Martin and A. D. Yates. 2023. Ensembl 2024. Nucleic Acids Research. 52 (D1): D891–D899. Heckman, T. I., K. Shahin, E. E. Henderson, M. J. Griffin and E. Soto. 2022. Development and efficacy of Streptococcus iniae live-attenuated vaccines in Nile tilapia (Oreochromis niloticus). Fish & Shellfish Immunology. 121: 152–162. Hoelzer, K., L. Bielke, D. P. Blake, E. Cox, S. M. Cutting, B. Devriendt, E. Erlacher-Vindel, E. Goossens, K. Karaca, S. Lemiere, M. Metzner, M. Raicek, M. C. Suriñach, N. M. Wong, C. Gay and F. V. Immerseel. 2018. Vaccines as alternatives to antibiotics for food producing animals. Part 1: challenges and needs. Veterinary Research. 49 (1). Holcomb, D. D., A. Alexaki, U. Katneni and C. Kimchi-Sarfaty. 2019. The Kazusa codon usage database, CoCoPUTs, and the value of up-to-date codon usage statistics. Infection, Genetics and Evolution. 73: 266–268. Hon, T., K. Mars, G. Young, Y.-C. Tsai, J. W. Karalius, J. M. Landolin, N. Maurer, D. Kudrna, M. A. Hardigan, C. C. Steiner, S. J. Knapp, D. Ware, B. Shapiro, P. Peluso and D. R. Rank. 2020. Highly accurate long-read HiFi sequencing data for five complex genomes. Scientific Data. 7 (1). Hu, J., Z. Wang, F. Liang, S.-L. Liu, K. Ye and D.-P. Wang. 2024. NextPolish2: A Repeataware Polishing Tool for Genomes Assembled Using HiFi Long Reads. Genomics, Proteomics and Bioinformatics. 22 (1). Jain, C., S. Koren, A. Dilthey, A. M. Phillippy and S. Aluru. 2018a. A fast adaptive algorithm for computing whole-genome homology maps. Bioinformatics. 34 (17): i748–i756. Jain, C., A. Rhie, H. Zhang, C. Chu, B. P. Walenz, S. Koren and A. M. Phillippy. 2020. Weighted minimizer sampling improves long read mapping. Bioinformatics. 36 (Supplement1): 111–118. Jain, C., A. Rhie, N. F. Hansen, S. Koren and A. M. Phillippy. 2022. Long-read mapping to repetitive reference sequences using Winnowmap2. Nature Methods. 19 (6): 705–710. Jain, M., S. Koren, K. H. Miga, J. Quick, A. C. Rand, T. A. Sasani, J. R. Tyson, A. D. Beggs, A. T. Dilthey, I. T. Fiddes, S. Malla, H. Marriott, T. Nieto, J. O’Grady, H. E. Olsen, B. S. Pedersen, A. Rhie, H. Richardson, A. R. Quinlan, T. P. Snutch, L. Tee, B. Paten, A. M. Phillippy, J. T. Simpson, N. J. Loman and M. Loose. 2018b. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nature Biotechnology. 36 (4): 124 338–345. Jiang, D. L., X. H. Gu, B. J. Li, Z. X. Zhu, H. Qin, Z. n. Meng, H. R. Lin and J. H. Xia. 2019. Identifying a Long QTL Cluster Across chrLG18 Associated with Salt Tolerance in Tilapia Using GWAS and QTL-seq. Marine Biotechnology. 21 (2): 250–261. Jumper, J. and others. 2021. Highly accurate protein structure prediction with AlphaFold. Nature. 596 (7873): 583–589. Kanehisa, M. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research. 28 (1): 27–30. Kanehisa, M. and Y. Sato. 2020. KEGG Mapper for inferring cellular functions from protein sequences. Protein Science. 29 (1): 28–35. Katoh, K. and D. M. Standley. 2013. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Molecular Biology and Evolution. 30 (4): 772–780. Katoh, K., K. Misawa, K. Kuma and T. Miyata. 2002. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Research. 30 (14): 3059–3066. Kayansamruaj, P., H. T. Dong, N. Pirarat, D. Nilubol and C. Rodkhum. 2017. Efficacy of α -enolase-based DNA vaccine against pathogenic Streptococcus iniae in Nile tilapia (Oreochromis niloticus). Aquaculture. 468: 102–106. Kayansamruaj, P., N. Areechon and S. Unajak. 2020. Development of fish vaccine in Southeast Asia: A challenge for the sustainability of SE Asia aquaculture. Fish and Shellfish Immunology. 103: 73–87. Kitano, J., S. Mori and C. L. Peichel. 2007. Sexual Dimorphism in the External Morphology of the Threespine Stickleback (Gasterosteus Aculeatus). Copeia. 2007 (2): 336–349. Kolberg, J., A. Aase, S. Bergmann, T. Herstad, G. Rodal, R. Frank, M. Rohde and S. Hammerschmidt. 2006. Streptococcus pneumoniae enolase is important for plasminogen binding despite low abundance of enolase protein on the bacterial cell surface. Microbiology (Reading, England). 152 (Pt 5): 1307–1317. Kolmogorov, M., D. M. Bickhart, B. Behsaz, A. Gurevich, M. Rayko, S. B. Shin, K. Kuhn, J. Yuan, E. Polevikov, T. P. L. Smith and P. A. Pevzner. 2020. metaFlye: scalable longread metagenome assembly using repeat graphs. Nature Methods. 17 (11): 1103–1110. Krogh, A., B. Larsson, G. vonHeijne and E. L. Sonnhammer. 2001. Predicting transmembrane protein topology with a hidden markov model: application to complete 125 genomes11Edited by F. Cohen. Journal of Molecular Biology. 305 (3): 567–580. Kusakabe, M., A. Ishikawa, M. Ravinet, K. Yoshida, T. Makino, A. Toyoda, A. Fujiyama and J. Kitano. 2016. Genetic basis for variation in salinity tolerance between stickleback ecotypes. Molecular Ecology. 26 (1): 304–319. Kutzler, M. and D. Weiner. 2008. DNA vaccines: ready for prime time? Nature Reviews Genetics. 9: 776–788. Kyte, J. and R. F. Doolittle. 1982. A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology. 157 (1): 105–132. Langmead, B. and S. L. Salzberg. 2012. Fast gapped-read alignment with Bowtie 2. Nature Methods. 9 (4): 357–359. Le Bras, Y., N. Dechamp, F. Krieg, O. Filangi, R. Guyomard, M. Boussaha, H. Bovenhuis, T. G. Pottinger, P. Prunet, P. Le Roy and E. Quillet. 2011. Detection of QTL with effects on osmoregulation capacities in the rainbow trout (Oncorhynchus mykiss). BMC Genetics. 12 (1). Letunic, I., S. Khedkar and P. Bork. 2021. SMART: recent updates, new developments and status in 2020. Nucleic Acids Research. 49 (D1): D458–D460. Lewin, H. A., J. A. M. Graves, O. A. Ryder, A. S. Graphodatsky and S. J. O’Brien. 2019. Precision nomenclature for the new genomics. GigaScience. 8 (8). Li, H. 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 34 (18): 3094–3100. Li, H. and R. Durbin. 2009. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics. 25 (14): 1754–1760. Li, H. and R. Durbin. 2010. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics. 26 (5): 589–595. Li, H., B. Handsaker, A. Wysoker, T. Fennell, J. Ruan, N. Homer, G. Marth, G. Abecasis and R. Durbin. 2009. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 25 (16): 2078–2079. Li, K., P. Xu, J. Wang, X. Yi and Y. Jiao. 2023. Identification of errors in draft genome assemblies at single-nucleotide resolution for quality assessment and improvement. Nature Communications. 14 (1). Li, W., K. R. O’Neill, D. H. Haft and others. 2021. RefSeq: expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation. Nucleic Acids Research. 49 (D1): D1020–D1028. 126 Li, W. and A. Godzik. 2006. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 22 (13): 1658–1659. Lin, X., J. Tan, Y. Shen, B. Yang, Y. Zhang, Y. Liao, P. Wang, D. Zhou, G. Li and C. Tian. 2022. A high-density genetic linkage map and QTL mapping for sex in Clarias fuscus. Aquaculture. 561: 738723. Lin, Y., C. Ye, X. Li, Q. Chen, Y. Wu, F. Zhang, R. Pan, S. Zhang, S. Chen, X. Wang, S. Cao, Y. Wang, Y. Yue, Y. Liu and J. Yue. 2023. quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification. Horticulture Research. 10 (8). Lisachov, A., D. H. M. Nguyen, T. Panthum, S. F. Ahmad, W. Singchat, J. Ponjarat, K. Jaisamut, P. Srisapoome, P. Duengkae, S. Hatachote, K. Sriphairoj, N. Muangmai, S. Unajak, K. Han, U. Na-Nakorn and K. Srikulnath. 2023. Emerging importance of bighead catfish (Clarias macrocephalus) and north African catfish (C. gariepinus) as a bioresource and their genomic perspective. Aquaculture. 573: 739585. Lischer, H. E. L. and K. K. Shimizu. 2017. Reference-guided de novo assembly approach improves genome reconstruction for related species. BMC Bioinformatics. 18 (1). Liu, C., X. Hu, Z. Cao, Y. Sun, X. Chen and Z. Zhang. 2019. Construction and characterization of a DNA vaccine encoding the SagH against Streptococcus iniae.Fish & Shellfish Immunology. 89: 71–75. Liu, J., Y. Cao, H. Ma, H. Du, T. Liu, G. Wang, M. Liu, Q. Wang, P. Li and E. Wang. 2023a. Enolase-based nanovaccine immersion immunization induces robust immunity and protection against Streptococcus infection in tilapia. Aquaculture. 576. Liu, L. and others. 2009. Identification and experimental verification of protective antigens against Streptococcus suis serotype 2 based on genome sequence analysis. Current Microbiology. 58: 11–17. Liu, M., Y. Song, S. Zhang, L. Yu, Z. Yuan, H. Yang, M. Zhang, Z. Zhou, I. Seim, S. Liu, G. Fan and H. Yang. 2023b. A chromosome-level genome of electric catfish (Malapterurus electricus) provided new insights into order Siluriformes evolution. Marine Life Sciencec and Technology. 6 (1): 1–14. Liu, Y., L. Li, F. Yu and others. 2020. Genome-wide analysis revealed the virulence attenuation mechanism of the fish-derived oral attenuated Streptococcus iniae vaccine strain YM011. Fish & Shellfish Immunology. 106: 546–554. Mahmoud, M., Y. Huang, K. Garimella, P. A. Audano, W. Wan, N. Prasad, R. E. Handsaker, 127 S. Hall, A. Pionzio, M. C. Schatz, M. E. Talkowski, E. E. Eichler, S. E. Levy and F. J. Sedlazeck. 2023. Utility of long-read sequencing for All of Us. Cold Spring. Maneechot, N., C. F. Yano, L. A. C. Bertollo, N. Getlekha, W. F. Molina, S. Ditcharoen, B. Tengjaroenkul, W. Supiwong, A. Tanomtong and M. deBello Cioffi. 2016. Genomic organization of repetitive DNAs highlights chromosomal evolution in the genus Clarias (Clariidae, Siluriformes). Molecular Cytogenetics. 9 (1). Marçais, G. and C. Kingsford. 2011. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 27 (6): 764–770. Marillonnet, S. and R. Grutzner. 2020. Synthetic DNA Assembly Using Golden Gate Cloning and the Hierarchical Modular Cloning Pipeline. Current Protocols in Molecular Biology. 130 (1): e115. Martin, M., M. Patterson, S. Garg, S. O Fischer, N. Pisanti, G. W. Klau, A. Schöenhuth and T. Marschall. 2016. WhatsHap: fast and accurate read-based phasing. OXford Academics - Bioinformatics. Mc Cartney, A. M., K. Shafin, M. Alonge, A. V. Bzikadze, G. Formenti, A. Fungtammasan, K. Howe, C. Jain, S. Koren, G. A. Logsdon, K. H. Miga, A. Mikheenko, B. Paten, A. Shumate, D. C. Soto, I. Sović, J. M. D. Wood, J. M. Zook, A. M. Phillippy and A. Rhie. 2022. Chasing perfection: validation and polishing strategies for telomere-totelomere genome assemblies. Nature Methods. 19 (6): 687–695. Md, V., S. Misra, H. Li and S. Aluru. 2019. Efficient architecture-aware acceleration of bwamem for multicore systems. In IEEE International Parallel and Distributed Processing Symposium (IPDPS). Membrebe, J. D., N. K. Yoon, M. Hong, J. Lee, H. Lee, K. Park, S. H. Seo, I. Yoon, S. Yoo, Y. C. Kim and J. Ahn. 2016. Protective efficacy of Streptococcus iniae derived enolase against streptococcal infection in a zebrafish model. Veterinary Immunology and Immunopathology. 170: 25–29. Meng, E. C., T. D. Goddard, E. F. Pettersen and others. 2023. UCSF ChimeraX: Tools for structure building and analysis. Protein Science. 32 (11): e4792. Mikheenko, A., A. V. Bzikadze, A. Gurevich, K. H. Miga and P. A. Pevzner. 2020. TandemTools: mapping long reads and assessing/improving assembly quality in extra-long tandem repeats. Bioinformatics. 36 (Supplement_1): i75–i83. Millard, C. and others. 2012. Evolution of the capsular operon of Streptococcus iniae in response to vaccination. Applied and Environmental Microbiology. 78 (23): 8219–8226. 128 Mishra, A. and others. 2018. Current Challenges of Streptococcus Infection and Effective Molecular, Cellular, and Environmental Control Methods in Aquaculture. Molecules and Cells. 41 (6): 495–505. Mistry, J. and others. 2021. Pfam: The protein families database in 2021. Nucleic Acids Research. 49 (D1): D412–D419. Mmanda, F. and others. 2014. Massive mortality associated with Streptococcus iniae infection in cage-cultured red drum (Sciaenops ocellatus) in Eastern China. African Journal of Microbiology Research. 8: 1722–1729. Moriel, D. and others. 2010. Identification of protective and broadly conserved vaccine antigens from the genome of extraintestinal pathogenic Escherichia coli. Proceedings of the National Academy of Sciences. 107 (20): 9072–9077. Na-Nakorn, U., W. Kamonrat and T. Ngamsiri. 2004. Genetic diversity of walking catfish, Clarias macrocephalus, in Thailand and evidence of genetic introgression from introduced farmed C. gariepinus. Aquaculture. 240 (1-4): 145–163. Nawawi, R., J. Baiano and A. Barnes. 2008. Genetic variability amongst Streptococcus iniae isolates from Australia. Journal of Fish Diseases. 31 (4): 305–309. Naylor, R. L., R. W. Hardy, A. H. Buschmann, S. R. Bush, L. Cao, D. H. Klinger, D. C. Little, J. Lubchenco, S. E. Shumway and M. Troell. 2021. Publisher Correction: A 20-year retrospective review of global aquaculture. Nature. 595 (7868): E36–E36. Norman, J. D., M. Robinson, B. Glebe, M. M. Ferguson and R. G. Danzmann. 2012. Genomic arrangement of salinity tolerance QTLs in salmonids: A comparative analysis of Atlantic salmon (Salmo salar) with Arctic charr (Salvelinus alpinus) and rainbow trout (Oncorhynchus mykiss). BMC Genomics. 13 (1): 420. Ondov, B. D., G. J. Starrett, A. Sappington, A. Kostic, S. Koren, C. B. Buck and A. M. Phillippy. 2019. Mash Screen: high-throughput sequence containment estimation for genome discovery. Genome Biology. 20 (1). Osorio, D., P. Rondon-Villarreal and R. Torres. 2015. Peptides: A package for data mining of antimicrobial peptides. The R Journal. 7 (1): 4–14. Ou, S. and N. Jiang. 2017. LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons. Plant Physiology. 176 (2): 1410–1422. Ou, S. and N. Jiang. 2019. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mobile DNA. 10 (1). 129 Ou, S., J. Chen and N. Jiang. 2018. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic Acids Research. Ou, S., W. Su, Y. Liao, K. Chougule, J. R. A. Agda, A. J. Hellinga, C. S. B. Lugo, T. A. Elliott, D. Ware, T. Peterson, N. Jiang, C. N. Hirsch and M. B. Hufford. 2019. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biology. 20 (1). Ouchi, S., R. Kajitani and T. Itoh. 2023. GreenHill: a de novo chromosome-level scaffolding and phasing tool using Hi-C. Genome Biology. 24 (1). Page, A. J., C. A. Cummins, M. Hunt, V. K. Wong, S. Reuter, M. T. G. Holden, M. Fookes, D. Falush, J. A. Keane and J. Parkhill. 2015a. Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics. 31 (22): 3691–3693. Page, A. J., C. A. Cummins, M. Hunt, V. K. Wong, S. Reuter, M. T. Holden, M. Fookes, D. Falush, J. A. Keane and J. Parkhill. 2015b. Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics. 31 (22): 3691–3693. Pandurangan, A. P., J. Stahlhacke, M. E. Oates, B. Smithers and J. Gough. 2019. The SUPERFAMILY 2.0 database: a significant proteome update and a new webserver. Nucleic Acids Research. 47 (D1): D490–D494. Pertea, G. and M. Pertea. 2020. GFF Utilities: GffRead and GffCompare. F1000Research. 9: 304. Pham, D. K., J. Chu, N. T. Do, F. Brose, G. Degand, P. Delahaut, E. D. Pauw, C. Douny, K. V. Nguyen, T. D. Vu, M.-L. Scippo and H. F. L. Wertheim. 2015. Monitoring Antibiotic Use and Residue in Freshwater Aquaculture for Domestic Use in Vietnam. EcoHealth. 12 (3): 480–489. Pier, G. and S. Madin. 1976. Streptococcus iniae sp. nov., a beta-hemolytic streptococcus isolated from an Amazon freshwater dolphin, Inia geoffrensis. International Journal of Systematic Bacteriology. 26 (4): 545–553. Pridgeon, J. W. and P. H. Klesius. 2011. Development and efficacy of a novobiocin-resistant Streptococcus iniae as a novel vaccine in Nile tilapia (Oreochromis niloticus). Vaccine. 29 (35): 5986–5993. Pumchan, A., S. Krobthong, S. Roytrakul and others. 2020. Novel chimeric multiepitope vaccine for streptococcosis disease in Nile tilapia (Oreochromis niloticus Linn.). Scientific Reports. 10: 603. Putnam, N. H., B. L. O’Connell, J. C. Stites, B. J. Rice, M. Blanchette, R. Calef, C. J. Troll, 136 analysis. PLoS ONE. 5 (1): e10546. Woestenenk, E. A., M. Hammarström, S. van denBerg, T. Härd and H. Berglund. 2004. His tag effect on solubility of human proteins produced in Escherichia coli: a comparison between four expression vectors. Journal of Structural and Functional Genomics. 5 (3): 217–229. Wyneken, J., S. P. Epperly, L. B. Crowder, J. Vaughan and K. Blair Esper. 2007. DETERMINING SEX IN POSTHATCHLING LOGGERHEAD SEA TURTLES USING MULTIPLE GONADAL AND ACCESSORY DUCT CHARACTERISTICS. Herpetologica. 63 (1): 19–30. Xiong, W., L. He, J. Lai, H. K. Dooner and C. Du. 2014. HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes. Proceedings of the National Academy of Sciences. 111 (28): 10263–10268. Xiong, X., Y. Peng, R. Chen, X. Liu and F. Jiang. 2023. Efficacy and transcriptome analysis of golden pompano (Trachinotus ovatus) immunized with a formalin-inactivated vaccine against Streptococcus iniae.Fish & Shellfish Immunology. 134: 108489. Xu, M., L. Guo, S. Gu, O. Wang, R. Zhang, B. A. Peters, G. Fan, X. Liu, X. Xu, L. Deng and Y. Zhang. 2020. TGS-GapCloser: A fast and accurate gap closer for large genomes with low coverage of error-prone long reads. GigaScience. 9 (9). Xu, Z. and H. Wang. 2007. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Research. 35 (Web Server): W265–W268. Yan, J.-J., Y.-C. Lee, Y.-L. Tsou, Y.-C. Tseng and P.-P. Hwang. 2020. Insulin-like growth factor 1 triggers salt secretion machinery in fish under acute salinity stress. Journal of Endocrinology. 246 (3): 277–288. Yang, Z., X. Zeng, Y. Zhao and others. 2023. AlphaFold2 and its applications in the fields of biology and medicine. Signal Transduction and Targeted Therapy. 8: 115. Yu, L. X. and others. 2014. Understanding pharmaceutical quality by design. AAPS Journal. 16 (4): 771–783. Yu, X., P. Setyawan, J. W. Bastiaansen, L. Liu, I. Imron, M. A. Groenen, H. Komen and H.-J. Megens. 2022. Genomic analysis of a Nile tilapia strain selected for salinity tolerance shows signatures of selection and hybridization with blue tilapia (Oreochromis aureus). Aquaculture. 560: 738527. Yue, G. H. 2013. Recent advances of genome mapping and marker�assisted selection in aquaculture. Fish and Fisheries. 15 (3): 376–396. 137 Zayas, J. F. 1997. Solubility of Proteins. Functionality of Proteins in Food. pp. 6–75. Zeng, X., Z. Yi, X. Zhang, Y. Du, Y. Li, Z. Zhou, S. Chen, H. Zhao, S. Yang, Y. Wang and G. Chen. 2024. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. Nature Plants. 10 (8): 1184–1200. Zhang, B.-C., J. Zhang and L. Sun. 2014a. Streptococcus iniae SF1: Complete genome sequence, proteomic profile, and immunoprotective antigens. PLoS ONE. 9 (3): e91324. Zhang, B.-c., J. Zhang and L. Sun. 2014b. Streptococcus iniae SF1: Complete Genome Sequence, Proteomic Profile, and Immunoprotective Antigens. PLoS ONE. 9 (3): e91324. Zhang, R.-G., G.-Y. Li, X.-L. Wang, J. Dainat, Z.-X. Wang, S. Ou and Y. Ma. 2022. TEsorter: An accurate and fast method to classify LTR-retrotransposons in plant genomes. Horticulture Research. 9. Zheng, Z., S. Li, J. Su, A. W.-S. Leung, T.-W. Lam and R. Luo. 2022. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nature Computational Science. 2 (12): 797–803. Zhou, Z., Y. Dang, M. Zhou, L. Li, C. Yu, J. Fu, S. Chen and Y. Liu. 2016. Codon usage is an important determinant of gene expression levels largely through its effects on transcription. Proceedings of the National Academy of Sciences. 113 (41): E6117–E6125. 138 Appendix 139 Figure 30 GenomeScope2.0 profiles for (a) male C. gariepinus, (b) female C. macrocephalus, and the F1 hybrid catfish at (c) k=21 and (d) k=31. The hybrid genome shows an intermediate heterozygosity level ( 1%) and genome size of 1.8 Gb, consistent with contributions from both parental subgenomes. Parental species exhibit 0.056% (C. macrocephalus) and 1.56% (C. gariepinus) heterozygosity, respectively. BUSCO analyses confirmed assembly completeness and the presence of two subgenomes. Low Illumina coverage (<20×) in parental datasets resulted in broad peaks in panels (a–b). 140 Figure 31 Comparative synteny and structural variation between the North African catfish (C. gariepinus) reference genome (GCA_024256425.2) and the hybrid catfish genome (fClaHyb_Gar, this study). (A) Genome-wide macrosynteny across 28 pseudochromosomes showing conserved collinearity and localized inversions, translocations, and duplications. (B) Sequence variation lengths and percentages of genome size, highlighting highly divergent regions. (C) Structural variation composition including duplications, translocations, and inversions. (D) Variant feature counts (SNPs, indels, CNVs, tandem repeats) showing genome-wide heterogeneity. 141 Figure 32 Comparative synteny and structural variation between the Bighead catfish (C. macrocephalus) Haplotype 1 (GCA_048544425.1) and the hybrid genome (fClaHyb_Mac, this study). (A) Genome-wide macrosynteny across 27 pseudochromosomes showing extensive collinearity with limited rearrangements. (B) Sequence variation profiles showing divergence and insertion–deletion patterns. (C) Structural variation composition between parental and hybrid genomes. (D) Variant feature counts across all categories, emphasizing SNP dominance and minor structural variants. 142 Figure 33 Circular genome synteny and quality control overview of the five sequenced Streptococcus iniae strains (SIKU01–SIKU05). The Circos plot shows inter-strain genomic alignments, GC content variation, GC skew, and contig connectivity metrics. Outer rings represent genome coordinates, while inner links highlight conserved syntenic regions among isolates, confirming overall structural stability across strains. 143 Figure 34 Comparative genome synteny of Streptococcus iniae strain SIKU01 relative to the original Amazon River dolphin isolate QMA0141 (collected in 1976) (Pier and Madin,1976). Homologous regions show strong macrosyntenic conservation with limited structural rearrangements, confirming longterm genomic stability across lineages. 144 Figure 35 Antigenic variation and gene carriage across the 17 proteins carrying predicted epitopes in Streptococcus iniae SIKU01. Bar plots show the distribution of antigenic determinants, sequence variability, and cross-strain presence of corresponding loci, indicating conserved immunogenic targets for vaccine development. 145 Figure 36 Biophysical landscape of the unfiltered SIKU01 proteome (M0), showing distributions of molecular weight, hydrophobicity (GRAVY), isoelectric point (pI), and instability index (II). These parameters establish the baseline design space for subsequent QbD filtering.