Full text
THESIS PROPOSAL GENOMIC STUDIES OF PISCINE STREPTOCOCCI AND GENOME ASSEMBLY OF CLARIIDAE CATFISHES (SILURIFORMES) FOR AQUACULTURE ENHANCEMENT QUENTIN LUDOVIC STEPHANE ANDRES GRADUATE SCHOOL KASETSART UNIVERSITY 2025
THESIS PROPOSAL APPROVAL GRADUATE SCHOOL KASETSART UNIVERSITY DEGREE Doctor of Philosophy (Fishery Science and Technology) MAJOR FIELD Fishery Science and Technology FACULTY Fisheries TITLE Genomic studies of piscine streptococci and genome assembly of Clariidae catfishes (Siluriformes) for aquaculture enhancement NAME MR. QUENTIN LUDOVIC STEPHANE ANDRES THIS THESIS HAS BEEN ACCEPTED BY (Associate Professor Prapansak Srisapoome, Ph.D.) THESIS ADVISOR (Professor Kornsorn Srikulnath, Ph.D.) THESIS CO-ADVISOR (Mr. Worapong Singchat, Ph.D.) THESIS CO-ADVISOR (Assistant Professor Methee Kaewnern, D.Tech.Sc.) GRADUATE COMMITTEE CHAIRMAN (Associate Professor Weeraphart Khunrattanasiri, Dr.rer.nat.) DEAN
THESIS PROPOSAL Genomic studies of piscine streptococci and genome assembly of Clariidae catfishes (Siluriformes) for aquaculture enhancement QUENTIN LUDOVIC STEPHANE ANDRES A Thesis Proposal Submitted in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy (Fishery Science and Technology) Graduate School, Kasetsart University Academic Year 2025
1 Contents Page LIST OF TABLES 3 LIST OF FIGURES 5 LIST OF ACRONYMS AND GLOSSARY 5 INTRODUCTION 1 OBJECTIVES AND RATIONALE 3 LITERATURE REVIEW 4 Genomic Technologies for Aquaculture 4 Clarias Catfish Aquaculture in Southeast Asia 10 Case Study: Genome Assembly Applications in Tilapia 11 GENOME ASSEMBLY – BIGHEAD CATFISH 12 Overview of Genome Assembly Strategy 12 Haplotype Resolution, Hi-C Scaffolding, and Phasing 18 GENOME ASSEMBLY – F1 HYBRID CATFISH 33 Overview of Genome Assembly Strategy 33 VALIDATION OF CATFISH ASSEMBLIES 45 Benchmarking Methodology 45 Results – Bighead Catfish 48 Results – F1 Hybrid Catfish 63 GENOME ASSEMBLY, REVERSE VACCINOLOGY, AND QUALITY BY DESIGN — Streptococcus iniae 80 Introduction 80 Methods 83 Results 99 DATA AVAILABILITY AND NCBI SUBMISSIONS 113 Bighead Catfish 113 F1 Hybrid Catfish 113
2 Streptococcus iniae 114 Associated Publications 114 DISCUSSION AND CONCLUSIONS 115 Genomic Insights and Technical Achievements 115 Streptococcus iniae Vaccine Development 116 Overall Conclusions 117 Recommendations 118 Appendix 141 Personal Information 164
3 List of Tables Table Page 1 Bighead Catfish – Assembly Stats. 49 2 Bighead Catfish – Scaffold Metrics – Haplotype 1. 57 3 Bighead Catfish – Scaffold Metrics – Haplotype 2. 58 4 Bighead Catfish – TE Content. 60 5 F1 Hybrid Catfish – North African Subgenome – Scaffold Metrics. 75 6 F1 Hybrid Catfish – Bighead Subgenome – Scaffold Metrics. 76 7 F1 Hybrid Catfish – North African Subgenome – Structural Metrics. 78 8 F1 Hybrid Catfish – Bighead Subgenome – Structural Metrics. 79 9S. iniae RV – Predicted antigenic epitopes from strain SIKU01. 102 10 S. iniae QbD – QTPP and CQAs definition workflow. 156 11 S. iniae – Compiled vaccine studies. 157 12 S. iniae – Supplementary Code for analysis. 160 13 S. iniae – Supplementary Data 1: Metadata and Proteome. 161 14 S. iniae – Supplementary Data 2: Pangenomics and MSAs. 162 15 S. iniae – Supplementary Data 3: QbD Manufacturability. 163
LIST OF FIGURES Figure Page 1 Long-reads vs Short-reads 4 2 ONT sequencing 5 3 HiFi Sequencing 6 4 Illumina PE Sequencing 7 5 Hi-C Sequencing 8 6 Clarias catfish species 10 7 Bighead Catfish – Specimen Photograph. 16 8 Bighead Catfish – Genome Assembly Workflow. 17 9 Bighead Catfish – Genome Assembly Graph – Haplotype 1 and 2. 20 10 Bighead Catfish – Genome Assembly 1-Year Progress. 27 11 Bighead Catfish – mtDNA Alignment Against 208 Catfish Species. 30 12 F1 Hybrid Catfish – Genome Assembly Workflow. 38 13 F1 Hybrid Catfish – Assembly validation using IGV. 43 14 F1 Hybrid Catfish – Genome Assembly 1-Year Progress. 44 15 Bighead Catfish – DNA Sequencing. 48 16 Bighead Catfish – Hi-C Heatmap – Haplotype 1. 51 17 Bighead Catfish – Hi-C Heatmap – Haplotype 2. 52 18 Bighead Catfish – Hi-C Heatmaps – Haplotype 1 and 2. 53 19 Bighead Catfish – Assembly Completeness Assessment. 55 20 Bighead Catfish – TE Divergence Profile. 59 21 Four Species Macrosynteny – Catfishes Chromosomes. 62 22 F1 Hybrid Catfish – Sequencing and GenomeScope2.0 survey. 64 23 F1 Hybrid Catfish – North African Subgenome – Hi-C Heatmap. 70 24 F1 Hybrid Catfish – Bighead Subgenome – Hi-C Heatmap. 71 25 F1 Hybrid Catfish – Genome Assembly Validation. 73 26 Streptococcus iniae - Phase contrast micrograph 80 27 S. iniae – Genome assembly and annotation. 100 4
5 28 S. iniae QbD – Workflow QTPP and manufacturability. 105 28 S. iniae QbD – Manufacturability design spaces. 110 29 S. iniae RV – 3D Structure of Enolase and GAPDH Immunogens. 112 30 Four Catfish Genome Survey – GenomeScope2.0 Profiles. 142 31 C. gariepinus – Synteny vs F1 hybrid genome. 143 32 C. macrocephalus – Synteny vs F1 hybrid genome. 144 33 S. iniae – Circos synteny & QC (SIKU01–05). 145 34 S. iniae – Synteny vs dolphin isolate. 146 35 S. iniae – Antigenic variation (17 epitopes). 147 36 S. iniae – Proteome landscape (M0). 148 37 S. iniae – Proteome annotation. 149 38 S. iniae – QbD CQAs correlation matrix. 150 39 S. iniae – QbD filtering workflow (M0–PreM1–M1-general). 151 40 S. iniae – Expression and pDNA platform filters. 152 41 S. iniae – Vaccine type and RPS distribution in literature. 153 42 S. iniae – RPS distribution by fish host species. 154 43 S. iniae – Comparison of pDNA and protein vaccine RPS. 155
1 GENOMIC STUDIES OF PISCINE STREPTOCOCCI AND GENOME ASSEMBLY OF CLARIID CATFISHES (SILURIFORMES) FOR AQUACULTURE ENHANCEMENT INTRODUCTION Global aquaculture is the process of farming aquatic organisms, both marine and freshwater, for human consumption and other uses, such as producing feed, pharmaceuticals, and related products. It represents the fastest-growing food production sector, providing over half of the seafood consumed globally, with Asia dominating production. In Thailand, inland aquaculture is critical for global food security, where freshwater species such as tilapia (Oreochromis spp.), catfish (Clarias spp., Silurus spp., and Pangasius spp.), and Asian seabass (Lates calcarifer) contribute significantly to domestic consumption and seafood exports (FAO,2020;Naylor et al.,2021). In 2022, aquaculture surpassed capture fisheries as the primary source of aquatic animals, producing 94.4 million tonnes compared to 92.3 million tonnes from wildcapture fisheries, according to the Global Seafood Alliance. Globally, aquaculture production (including seaweeds) reached 130.9 million tonnes in 2022, valued at USD 313 billion, according to the Food and Agriculture Organization. Major gains in aquaculture feed efficiency and fish nutrition have lowered the fish-in–fish-out ratio for many fed species, although dependence on marine ingredients persists and reliance on terrestrial ingredients has increased. To ensure sustainable growth, genetic improvement through selective breeding has become a key strategy, where optimizing traits such as growth, feed efficiency, sex determination, robustness, and disease resistance requires identifying and deploying relevant genetic variants. Modern aquaculture breeding programs do not rely solely on conserving “pure” strains or extreme inbreeding. Instead, they are typically structured around families or lines that serve as a defined starting point and are improved cumulatively over gener-
8 2. Hi-C for Genome Scaffolding Hi-C captures chromosome three-dimensional organization through proximityligation, where physically close DNA fragments become joined. Paired-end sequencing of these junctions reveals long-range linkage: contigs from the same chromosome show many Hi-C contacts while those from different chromosomes show few (Putnam et al., 2016). This enables clustering, ordering, and orienting contigs into chromosome-scale scaffolds for de novo assembly—construction without a reference genome—rather than reference-guided assembly which requires an existing genome. Hi-C has become essential for achieving chromosome-level assemblies in aquaculture species, enabling precise localization of genes and QTLs for breeding traits. In polyploid fish genomes, Hi-C correctly separates duplicated loci and provides insights into 3D genome structure, chromatin conformation, and regulatory interactions (Dudchenko et al.,2017). Figure 5 Ultra-long-range sequencing for genome scaffolding Credits: Ultra-long-range Sequencing for Genome Scaffolding 3. Haplogenome Assembly of Catfish Genomes Haplogenome assembly reconstructs each parental haplotype independently instead of producing a collapsed consensus sequence. This captures the full spectrum of
9 allelic and structural variation, valuable for outbred species and highly heterozygous organisms such as interspecific catfish hybrids. By resolving phase-specific sequences, haplogenome assemblies reveal allele-specific expression patterns critical for understanding gene function and optimizing selective breeding. The approach typically uses long-read sequencing (PacBio HiFi or ONT) with bioinformatic phasing to separate maternal and paternal haplotypes. A well-known example includes the assembly of haplogenomes in interspecific Eucalyptus hybrids, analogous to the genome-wide structural differences observed between Silurus asotus and S. meridionalis that are linked to hybrid traits Chen et al. (2021a). Similarly, in catfishes, haplogenome assemblies of Clarias macrocephalus,C. gariepinus, and their F1 hybrids enable precise characterization of chromosomal rearrangements, gene duplications, and introgression events—laying the foundation for genomic selection and improved aquaculture productivity Duong and Scribner (2018). 4. In Summary These sequencing platforms capture genome structure and identify variants— SNPs1, indels2, and structural variants (SVs)3—associated with economically important traits (Chaivichoo et al.,2020). Short reads excel at accurate SNP/indel detection, long reads resolve structural variants and complex regions, while Hi-C provides chromosomal-level scaffolding. As of May 2025, NCBI hosts over 3.01 million genomes including 51,820 eukaryotic genomes, with fewer than 500 being haplotype-resolved, complete vertebrate genomes (https://www.ncbi.nlm.nih.gov/datasets/genome/). 1Single-base changes at specific genomic positions, used as genetic markers. 2Short insertions or deletions of 1–50 bp. 3Large genomic alterations >50 bp including deletions, insertions, duplications, inversions, and translocations.
10 Clarias Catfish Aquaculture in Southeast Asia 1. Evolutionary History and Distribution Catfish (Siluriformes) are widely cultivated freshwater species crucial for food production across Africa, the Americas, and Asia (Lisachov et al.,2023). This diverse order diverged during the early Cretaceous ( 145 Ma) and comprises over 36 families and 3,000 species (Ferraris,2007). While distributed globally in freshwater environments with some marine extensions, catfish are most diverse in tropical South America, Asia, and Africa. Key representatives like Clarias and Ictalurus serve as both research models and economically valuable aquaculture species. 2. Classification and Biology Bighead catfish (Clarias macrocephalus) and North African catfish (C. gariepinus) belong to family Clariidae, order Siluriformes. These air-breathing catfish are economically important in Southeast Asia and Africa, respectively. C. macrocephalus, native to Southeast Asia and widely farmed in Thailand, has a broad flattened head, short barbels, and reaches 40–50 cm. Its body is dark brown to black with pale blotches. C. gariepinus, introduced globally for its fast growth and environmental tolerance, grows to 1.7 m with an elongated body and long dorsal fins. It survives oxygen-depleted waters using its suprabranchial organ for atmospheric breathing. Figure 6 Left: Female C. macrocephalus (Credits: Matillano, J.D.). Right: C. gariepinus (Credits: Larsen, J.H.)
11 F1 hybrids from C. gariepinus males × C. macrocephalus females combine parental advantages: body form from C. macrocephalus with rapid growth from C. gariepinus. Both species thrive at 26–32°C. 3. Role in Food Production and Economy Clarias species are extensively farmed across Asia and Africa. C. macrocephalus is primarily used for hybrid production with C. gariepinus (Lisachov et al.,2023;NaNakorn et al.,2004), yielding sterile F1 hybrids with hybrid vigor (heterosis) that combine rapid growth and desirable body characteristics. However, hybrid sterility necessitates separate parental production systems. The lack of high-quality reference genomes limits implementation of modern breeding tools like ddRAD sequencing, low-coverage genome sequencing, and CRISPR/Cas9. Studies indicate low introgression risk in C. macrocephalus, supporting potential genetic improvement efforts. Case Study: Genome Assembly Applications in Tilapia Tilapia (Oreochromis spp.) exemplifies successful genomic applications in aquaculture (Yu et al.,2022). Researchers identified adaptive responses4to salinity stress through genome-wide SNP analysis linked to osmoregulation and survival (Gu et al., 2018;Jiang et al.,2019;Rengmark et al.,2007). Key tolerance genes include prolactin (PRL) (Breves et al.,2013), growth hormone (GH1) (Deane and Woo,2008), insulinlike growth factor 1 (IGF1) (Yan et al.,2020), and plasma-membrane Ca2+-ATPases (PMCAs) (Rengmark et al.,2007). Similar studies in rainbow trout (Le Bras et al., 2011), Atlantic salmon (Norman et al.,2012), and stickleback (Kusakabe et al.,2016) collectively established foundations for marker-assisted selection (MAS) in aquaculture. Non-synonymous mutations in salt regulation genes enhance tolerance and alter phenotypes. These achievements required high-quality reference genome assemblies. 4Biological adjustments to environmental stress through gene regulation, metabolic changes, or immune modulation.
12 GENOME ASSEMBLY – BIGHEAD CATFISH Overview of Genome Assembly Strategy The haplotype-resolved genome of Clarias macrocephalus was assembled using an hybrid sequencing strategy comprising four DNA sequencing platforms generating various types of DNA reads (i.e., in DNA sequencing, a read is an inferred sequence of base pairs (or base pair probabilities) corresponding to all or part of a single DNA fragment). The four DNA sequencing technologies used to read the gDNA of bighead catfish into DNA reads (short-reads and long-reads) were: PacBio HiFi CSS (Wenger et al., 2019) (Circular consensus sequencing5, long-reads of base accuracy 99.9%), used for maximizing genome assembly quality. Oxford Nanopore (ONT) noisy 1D long-reads (20%.err)6(Jain et al.,2018b) for increasing assembly continuity and to span complex repeat regions of the genome. Proximity-ligation (Hi-C) (Dudchenko et al.,2017; Putnam et al.,2016) data for long-range linking information to phase (i.e., the specific arrangement of alleles on the same chromosome to distinguish maternal and paternal haplotypes) and scaffold contigs (i.e., separating haplotypes). These complementary data types facilitated haplotype resolution and scaffold anchoring and allow to generate dual assemblies7from a single diploid tissue. Among the tested assemblers (Hifiasm (Cheng et al.,2021), Flye (Kolmogorov et al.,2020), wtdbg2 (Ruan and Li,2019)), I selected Hifiasm for the haplotype-resolved assembly because of its ability to separate parental haplotypes (phasing) on low genome coverage (< 13X) HiFi data. In contrast I found that Flye and wtdbg2 produced collapsed consensus assemblies. Below, I outline the main steps for the complete genome assembly of bighead catfish Haplotype 1 and Haplotype 2. 5Circular consensus sequencing (CCS) is a high-accuracy long-read sequencing method, commonly used with PacBio HiFi reads, where multiple passes of the same DNA molecule are combined to generate a consensus sequence. 6Oxford Nanopore 1D reads are single-strand long reads that historically exhibit high error rates, often around 10–20%, due to insertions, deletions, and base-calling inaccuracies. These reads are valuable for spanning long genomic regions but require polishing for accuracy. 7Dual assemblies refer to separate genome assemblies generated for each haplotype in a diploid organism, allowing resolution of heterozygous regions and structural differences between homologous chromosomes.
13 1. Genomic DNA Sequencing, Genome Survey, and k-mer Meryl Databases Genomic DNA was extracted from the liver tissue of a single male bighead catfish individual using a custom protocol (Supikamolseni et al.,2015). Sex was determined based on histology and external gonadal examination (Kitano et al.,2007; Wyneken et al.,2007). Sequencing was performed using four complementary platforms generating four synergistic raw sequencing data types: (i) Pacific Biosciences (PacBio) High-Fidelity (HiFi) sequencing (Wenger et al.,2019) for generating highly accurate contigs, (ii) Oxford Nanopore Technology (ONT) longread sequencing (Jain et al.,2018b) for scaffolding and resolving repetitive regions, and (iii and iv) Illumina short-read 150 base paired-end sequencing (Hi-C, and standard non Hi-C) (Bentley et al.,2008) for chromosomal phasing and error correction, respectively. Genome survey aims in assessing genome characteristics (genome size, heterozygosity, repeat content) from raw sequencing data (e.g., Illumina short-reads) using k-mer analysis in Jellyfish (Marçais and Kingsford,2011) and GenomeScope2.0 (Ranallo-Benavidez et al.,2020). Meryl databases (N=2) for 21-mer and 31-mer (bighead.hifi.illuminaPCRfree.gt1.db.meryl) hybrid k-mer databases, made from combined Illumina and HiFi reads using Meryl (Rhie et al.,2020), as explained in the T2T-polish GitHub repository (https:// github.com/arangrhie/T2T-Polish) (Rhie et al.,2022). K-mer Meryl databases use 21-mer and 31-mer DNA strings to support quality evaluation and error correction in genome assemblies. The shorter 21-mer spectra provide high sensitivity for detecting potential errors, while the longer 31-mer spectra offer higher specificity, validating true variants and minimizing false positives. During genome polishing, these k-mer databases enhance effective coverage in low-depth HiFi regions by leveraging k-mer spectra to guide targeted corrections and improve assembly accuracy. 2. Initial Assembly Generation (Contigguing and Phasing): Contigging is the process of assembling overlapping DNA reads into continuous sequences without gaps (i.e., contigs). These contigs are then separated and assigned to a phase group using proximity-ligation data (Hi-C) based on their parental origin and spacial physical distance in the cell nucleus. Therefore, a contig that comes from a
14 haplotype is referred to as a haplotig8, and in a diploid phased assembly there are two groups of homologous haplotigs (i.e., unitigs, one per haploid ”homotype” haplotype). Here, Hifiasm was used in ”Hi-C UL” mode (Cheng et al.,2021) to generate a haplotype-resolved assembly using Hi-C + ONT + HiFi reads (Figure 8D, left), followed by GreenHill (Ouchi et al.,2023) using all data types: Hi-C + Illumina + ONT + HiFi reads for further scaffolding and phasing of the haplotigs (Figure 8E), and finally wtdbg2 (Hu et al.,2024;Ruan and Li,2019) was employed with HiFi and Illumina short-reads correcting small errors at the nucleotide level (i.e., SNPs) while preventing switch errors (i.e., allele switching between homologous haplotypes). This produced two phased contig sets representing homologous haplotypes (Figure 9). 3. Hi-C Scaffolding and Their Validation: Scaffolds are computationally ordered and oriented arrays of contigs that have sequence gaps along their length. The process of scaffolding of contigs is about arranging and orienting contigs into larger structures, sometimes with gaps, using additional data like proximity ligation (Hi-C) in order to reconstruct an in-silico equivalent to the karyotype (referred to as ”Scaffotype”) which consists of Hi-C scaffolds of ”pseudochromosomes” (i.e., chromosome pseudomolecules), visualized as Hi-C maps. Prior to Hi-C scaffolding, it is possible to break the contigs at erroneous sites, for that, misassembly corrections are done using CRAQ (Li et al.,2023), then Hi-C scaffolds (i.e., contigs and scaffolds) are flipped, reoriented, and refined to improve structural accuracy, this step is knows as the ”manual review step9” step and is performed using JBAT (JuiceBox Assembly Tools) (Dudchenko et al.,2017), next, a ”post-review step10 ” step seals contigs (i.e., by creating gaps between newly adjacent contigs and inserting 500 ’N’ characters to represent unknown 8In an unphased assembly, a contig may join alleles from different parental haplotypes in a diploid or polyploid genome (see (Lewin et al.,2019) and https://lh3.github.io/2021/04/17/ concepts-in-phased-assemblies). The process of separating sequences based on their parental haplotype in diploid genomes is referred to as ”haplotype phasing”. 9In Hi-C scaffolding, manual review involves visually inspecting contact maps (e.g., using Juicebox or HiGlass) to detect misassemblies, orientation errors, or misplaced contigs that automated tools may miss. 10Post-review refers to the stage after Hi-C contact map inspection, where validated corrections (e.g., flips, cuts, or joins) are applied to the genome assembly.
15 sequences) to form Hi-C scaffolds to ultimately reconstruct pseudochromosomes followed by 3D-DNA ”post-review” validation (Dudchenko et al.,2017) (Figure 8F). Finally more gaps were closed with TGS-GapCloser (Xu et al.,2020) (Figure 8G). These steps — including CRAQ break detection, Hi-C read alignment, manual review in JBAT, and post-review in the 3D-DNA pipeline — were iteratively performed three times. As illustrated in Figure 8F, this cycle can be repeated if needed after contig polishing (Figure 8H), and typically results in well-defined Hi-C scaffolded pseudochromosomes (Figure 8C). 4. Assembly Polishing of Contigs and Benchmarking: Polishing a genome assembly consist in correcting sequencing errors using additional data (Meryl k-mer databases, long-reads and short-reads). To further improve assembly accuracy, reads were aligned to the polished assembly, the three primary long-read mappers employed were Winnowmap (Jain et al.,2022), Minimap2 (Li,2018), and VerityMap (Mikheenko et al.,2020) (previously known as TandemTools), these were used in assembly validation only (i.e., not for variant calls and genome polishing). Secondary alignments (i.e., reads flagged with bit 0x10011 ), low-quality regions (i.e., MAPQ12 = 0), and hard-clipped reads (i.e., alignments with hard clipping operations indicated by ‘H’ in the CIGAR string13 ) were excluded from analysis using samtools (Li and Durbin,2009). This read realignment step was crucial for validating assembly completeness and reducing errors. Non-haplotypeaware tools14 (e.g., Racon (Vaser et al.,2017), Clair3 (Zheng et al.,2022)) were applied cautiously to minimize errors from parental haplotype switches, while Merfin (Formenti et al.,2022) filtered edits for polishing and BCFtools consensus produced the final polished sequence. For large variants and closing gaps using read alignments I used Sniffles2 (Smolka et al.,2024) for large structural 11In SAM/BAM flags, bit 0x100 indicates a secondary alignment. Tools like Picard use this flag to mark reads that are not the primary alignment of a read with multiple mappings. 12MAPQ (Mapping Quality) is a Phred-scaled score that reflects the confidence in the read’s alignment position. Higher values indicate more reliable mappings. 13The CIGAR (Compact Idiosyncratic Gapped Alignment Report) string encodes how a read aligns to a reference, using operations such as matches (M), insertions (I), deletions (D), soft clips (S), and others. For example, 100M indicates 100 matching bases. 14Non-haplotype-aware tools do not distinguish between maternal and paternal sequences in diploid or polyploid genomes, often collapsing allelic variants into a single consensus sequence and potentially masking structural differences between haplotypes.
16 variant (SV) calling, specifically for insertions (INS) and deletions (DEL). I evaluate genome quality in term of errors found in the assembly that are not found in raw reads, for that I calculate the Quality Value, usually refered to as QV15 of a genome assembly where QV refers to the Phred quality score (or quality (Q) score), an integer representing the estimated probability of an error (i.e., that a base is incorrect). The polishing workflow implemented here for correcting bighead catfish ensured high nucleotide accuracy Quality Values (QV) (from initial QV33 to around QV46) after Benchmarking against recommended assembly metrics, from the VGP paper (Rhie et al.,2021). Figure 7 Representative specimens of the bighead catfish (Clarias macrocephalus), collected for high-quality haplotype-resolved genome assembly. The individuals were reared under controlled aquaculture conditions prior to tissue sampling. C. macrocephalus is an economically important freshwater species native to Southeast Asia, widely cultured in Thailand for selective breeding and genomic improvement programs. Methodological details are presented in the following sections and in Figure 8. 15Quality Value (QV) is a Phred-scaled score used to represent base-level or assembly-level accuracy. If Pis the probability of error, then Q=−10log10(P), or equivalently, P=10−Q/10. Higher QV indicates greater accuracy.
17 Figure 8 Comprehensive haplotype-resolved genome assembly and scaffolding workflow for bighead catfish (Clarias macrocephalus). (A) Hi-C contact map (Juicebox v2.16.0). (B) Genome assembly using PacBio HiFi, ONT, and Hi-C reads with Hifiasm (Hi-C UL mode) and Flye for consensus. (C) Hi-C scaffolding via GreenHill and 3D-DNA. (D) Manual post-review with JBAT. (E) Gap filling using TGS-GapCloser and QuarTeT GapFiller. (F) Genome polishing with NextPolish2 and CRAQ. (G) Assembly QV validation with Merqury, Pilon, BCFtools, and VerityMap. (H) Mitochondrial genome assembly with MitoHiFi and Minimap2. The pipeline yields two high-quality phased haplotype assemblies representing the complete bighead catfish genome.
24 3. Running pilon and calling the consensus: Pilon was employed on assemblies for correcting SNPs, INDELs (i.e., 1-2 base insertions and deletions), and other base-level errors. In individual haplotypes, each scaffold was processed by Pilon with parameters (’--genome,--frags,--bam,--targets,--fix all, --vcf, --diploid,--minmq 30,--minqual 30’) to use the alignments from all data types (HiFi, ONT, and short reads). The output Variant Calling Files (VCF) of containing the detected variants were sorted by position using (’bcftools sort’), compressed using bgzip, and indexed using (bcftools index)28 , and final consensus sequences were generated for each scaffold using (’bcftools consensus -f $genome.fa -H 1’). The overall median quality value metric increased from 41 to approximately 45-47 after haplotype-aware targeted assembly polishing. 5. Hi-C Scaffolding with HapHiC To reintegrate unplaced scaffolds in each haplotype in the pseudochromosomes, I used Quartet and HapHiC. First, I aligned the unplaced contigs to reference genome C. fuscus (GCA_030347435.1), and reference genome C. gariepinus (GCA_024256425.2) using Quartet AssemblyMapper (Lin et al.,2023) v1.2.1 with the following parameters (’-r $reference -q $contigs -c 50000 -l 2000 -i 90 -a minimap2’). I identified 53 MB and 33 MB of unplaced scaffold sequences in bighead catfish Haplotype 1 and Haplotype 2, respectively, with strong homology to C. fuscus pseudochromosomes, next I filtered bighead catfish Haplotype 1 and Haplotype 2 separately with SeqKit (Shen et al.,2016) v0.8.1 to retain pseudochromosomes and concatenated them to unplaced scaffolds. For the preparation of Hi-C scaffolding step, Hi-C reads were mapped to separate haplotypes using BWA-MEM (’-5SP’) after making a BWT index29 28bcftools index creates a compressed index (.csi or .tbi) for VCF or BCF files, allowing rapid random access to variants by genomic position during downstream processing or visualization. 29The BWT index refers to a data structure derived from the Burrows-Wheeler Transform (BWT), which allows fast and memory-efficient alignment of sequencing reads to a reference genome. It compresses the genome while retaining the ability to search for exact or approximate matches, enabling tools like Bowtie2 and BWA to align short reads with high performance.
25 for Haplotype 1 and Haplotype 2 (’bwa index $genome.fasta’) and Hi-C read alignments filtered for PCR duplicates30 and secondary alignments (’samblaster $BAM | samtools view - -@ $threads -S -h -b -F 3340’) using the software Samblaster (Faust and Hall,2014) v0.1.26. Hi-C scaffolding was performed on bighead catfish haplotypes using HapHic (Zeng et al.,2024) v1.0. The resulting Hi-C maps were visualized in JBAT and using (’haphic plot’) for separate scaffolds from each haplotype (Figure 16 and Figure 17) but also with all scaffolds represented in a wide-map (Figure 18), and a post-review of the Hi-C scaffold was performed as described in the previous section. Finally, three rounds of TGS-GapCloser followed by one round of targeted Pilon polishing specifying the new targets resulted in a genome of global higher quality (i.e., both in term of QV quality values and CRAQ’s metrics of CRE/CSE structural accuracy). I addressed the duplication errors31 visible in the heterozygous peak of the k-mer coverage (i.e., from the Merqury spectra CN) with more or less success by using haplotypespecific k-mer databases. This involved applying (’meryl difference’) command to identify erroneous k-mers and using (’meryl-lookup’) command to extract reads associated with these k-mers. Finally, I used BCFtools to call a consensus32 using the opposite haplotype (’-H 1’) within bcftools consensus command, effectively reversing most haplotype switch errors. 6. SV and SNP Consensus Polishing Consensus polishing33 was automated by following instructions from the T2Tpolish GitHub repository (https://github.com/arangrhie/T2T-Polish). A HiFi mapping file of repetitive k-mers (k=15) was generated with Meryl and used in Winnowmap for 30PCR duplicates refer to read pairs that originate from the same DNA fragment and are amplified multiple times during library preparation, potentially leading to biased coverage if not removed. 31Duplication errors refer to the erroneous presence of redundant sequences in genome assemblies, often caused by uncollapsed haplotypes, repetitive regions, or misassemblies. These can inflate genome size and complicate downstream analyses. 32Calling a consensus in bcftools refers to generating a modified reference sequence by applying variant calls (e.g., from a .vcf file) to a reference genome, resulting in a consensus sequence that reflects the sample’s specific genotype. 33Consensus polishing is the process of correcting errors in a draft genome assembly using aligned reads, typically by averaging observed base calls at each position in a read pileup. Most polishing tools are not haplotype-aware and must be run on both haplotypes together, meaning reads are distributed across haplotigs without distinguishing parental origin. This can lead to miscorrections in heterozygous regions.
26 HiFi read alignment (’-MD -W .repetitive_k15.txt -ax map-pb’), followed by alignment filtering with samtools (’-Sb’). pb-falconc filtered hard-clipped reads and for bit 0x104 in Picard34 (’bam-filter-clipped -t -F 0x104 --output-count-fn’) (https://github.com/bio-nim/pb-falconc) v1.15.0 was then used to filter clipped reads. Genome polishing was performed using the liftover branch of the Racon GitHub repository, using Racon with (’-L -S’) options (Vaser et al.,2017) v1.5.0. After polishing, the k-mers present in the genome (seqmers) were counted using Meryl count (k=21). Merfin was then applied using (’-readmers Illumina.HiFi.gt1.PCRfree.hybrid.meryl -seqmers’) and the hybrid-kmer db of reads to evaluate the results by comparing the distributions in the reads and in the polished genome (Formenti et al.,2022;Rhie et al., 2022). For consensus generation (i.e., to apply polishing edits to the genome assembly), I used BCFtools (’-H 1’) for two rounds of genome correction. Assembly quality metrics were measured for QV, completeness and BUSCOs scores, I found a large increase in QV in most chromosomes (min. increase > +1-5 QV points) after ONT Racon and Merfin, the median QV was 50, the progress over months is presented in Figure 10. 34The SAM flag 0x104 is a combination of 0x004 (read unmapped) and 0x100 (not primary alignment). It indicates that the read is both unmapped and not the primary alignment, typically seen in secondary alignments of unaligned reads or ambiguous multi-mappings.
27 Figure 10 Bighead Catfish Assembly Status January 2024 - November 2025. (A) Haplotype 1 (blue) and Haplotype 2 (purple) ideograms showing gaps on pseudochromosomes (orange rectangles) and telomeres (blue triangles) after manual-review in Juicebox (JBAT). (B) Same assemblies a few months later.
28 7. Mitochondrial Genome Assembly The mitochondrial genome (mtDNA) was assembled by mapping to reference (NC_046749.1) bighead catfish mtDNA, available at NCBI Nucleotide35 . Minimap2 (’-ax map-ont --secondary=no’) mapped nanopore reads to the reference and minimap2 (’-ax sr’) mapped Illumina reads. PCR duplicates in short reads were removed from the alignments using samtools markdup and the results were visualized in the Integrative Genome Viewer (IGV) (Thorvaldsdottir et al.,2012). Pilon (’--fix all --diploid --changes --vcf --tracks --minmq 10’) was used to correct mtDNA (NC_046749.1) and call SNPs, SVs, gaps and local variants and obtain the consensus mtDNA sequence. Reads were re-aligned to the consensus mtDNA sequence and no more SNPs were visible in the IGV. Next, mapped reads were filtered with samtools view (’-F4 -q 20’) and re-assembled de novo with Unicycler (Wick et al.,2017) for comparison. The results were visualized in Bandage-ng (Wick et al.,2015) v2022.09. The assembly was polished with Pilon and the two homologous mitochondria were compared using minimap2 (’--eqx -x asm5’). Gene annotations were generated using MitoFinder (Allio et al.,2020) v1.4.2. To ensure that the mitochondrial genome was correct and not dissimilar to other mtDNAs in Siluriformes, I downloaded all reference mtDNA sequences for (N = 209) species of Siluriformes catfish from NCBI Nucleotide, last accessed September 2024. All sequences were taken from in NCBI RefSeq36 and not from NCBI GenBankNCBI GenBank37. All mtDNA nucleotide sequences (N = 210) were renamed to the Pan-SN naming scheme (https://github.com/ pangenome/PanSN-spec), concatenated in a single multi-FASTA including the bighead catfish mtDNA generated here after Pilon, and all-versus-all alignments were performed with wfmash, then ODGI was used to filter the alignment graph, and visualizations were 35NCBI Nucleotide is a public database maintained by the National Center for Biotechnology Information that provides access to nucleotide sequences from GenBank, RefSeq, and other sources for a wide range of organisms. 36NCBI RefSeq (Reference Sequence Database) is a curated, non-redundant collection of genomic, transcript, and protein sequences provided by the National Center for Biotechnology Information, serving as a standard reference for genome annotation and comparative analysis. 37NCBI GenBank is a comprehensive public database of annotated nucleotide sequences submitted by researchers worldwide. It includes raw and curated data and serves as a primary archive for sequence data in molecular biology.
29 made with Bandage and multiQC (Ewels et al.,2016). All tools of the pipeline were part of the larger PGGB (Pan Genome Graph Builder) (Guarracino et al.,2022). The results of the alignments are shown in Figure 11, the mtDNA aligns well with other sequences in the phylogenetic order and is not an outlier.
30 Figure 11 All-versus-all 2D representations of mitochondrial DNA sequences alignments in Siluriformes catfishes. The mtDNA assembled for bighead catfish is on the last row (blue color) of the linear 2D representation.
31 8. Transposable Element Annotation Two complementary approaches were used to identify transposable elements (TEs)38 in the bighead catfish genome. (1) A species-specific TE library was generated de novo for the bighead catfish. (2) A curated fish TE library was retrieved from the FishTEDB database (https://www.fishtedb.com/) and used as a reference for crossspecies TE annotation. 8.1 De novo transposable element library for bighead catfish De novo TE annotation was performed using EDTA (Extensive de novo TE Annotator) (Ou et al.,2019) v2.2.0. The pipeline integrates several tools to detect different element types. LTR harvest (Ellinghaus et al.,2008) v1.5.10 identifies LTR retrotransposons by efficiently locating structural features such as border positions, LTR lengths, and motifs in large genomic datasets. LTR FINDER (Xu and Wang,2007) and the parallel version of LTR FINDER (Ou and Jiang,2019) v1.1.0 accelerate LTR detection using parallel computing for large genomes. LTR retriever (Ou and Jiang, 2017) v2.9.0 refines LTR annotations to improve accuracy and remove false positives. TEsorter is an accurate and fast way to classify LTR retrotransposons (Zhang et al., 2022) v1.4.7, GRF (Generic Repeat Finder) (Shi and Liang,2019) v1.1 for genomewide de novo repeat detection. TIR-Learner (Su et al.,2019) v1.1.2 uses machine learning to detect terminal inverted repeat (TIR) transposons, including miniature inverted TEs (MITEs). HelitronScanner (Xiong et al.,2014) v1.0 identifies Helitrons by recognizing their characteristic structural motifs. RepeatModeler (Flynn et al.,2020) v2.0.3 performs de novo repeat discovery and builds a comprehensive repeat library. MAFFT (Katoh and Standley,2013) v7.526 is used for multiple sequence alignments, and HMMER (Finn et al.,2011) v3.4 is used for domain-based searching and TE classification. 38Transposable elements (TEs), also known as “jumping genes,” are mobile DNA sequences that can move or replicate within the genome. They play key roles in genome evolution, structural variation, and gene regulation.
32 8.2 FishTEDB general fish-specific TE library collection Next, libraries of TE consensus sequences39 for transposons were retrieved from the FishTEDB database (Shao et al.,2018) version 1, which contains different species of fish and their associated transposable element libraries in .fasta file format, and these sequences were merged with EDTA’s de novo bighead catfish initial TE annotations. To reduce the library complexity and to create a non-redundant TE library (i.e., to lower the number of sequences and redundancy of the combined TE library dataset), I used CD-HIT-EST (’-d 20 -aS 0.95 -c 0.95 -G 0 -g 1 -b 500’) (Li and Godzik,2006) for an 80% identity threshold and 80% alignment coverage of TE sequences in all-against-all sequence clustering following the commonly used 80%-80% rule (Goubert et al.,2022). The reduced TE library was used to mask (i.e., to annotate) the bighead catfish genome for TEs with RepeatMasker (Smit, AFA, Hubley, R. & Green, P RepeatMasker at http://www.repeatmasker.org). 39TE (transposable element) consensus sequences are artificial, representative sequences constructed by aligning and averaging multiple genomic copies of a TE family. They do not correspond to any specific real insertion, but serve as idealized references for annotation and classification.
33 GENOME ASSEMBLY – F1 HYBRID CATFISH Overview of Genome Assembly Strategy Starting from an F1 hybrid catfish (i.e., a first filial generation catfish made from the cross of a pure bighead catfish and a pure North African catfish), I reconstructed two complete haploid, non-homotypic parental genomes, comprising 27 pseudochromosomes for C. macrocephalus and 28 for C. gariepinus (Fig. 25). In most F1 hybrids derived from divergent parental species, the observed heterozygosity rate—typically around 10%—reflects the inter-specific genomic divergence between the two haploid sub-genomes, encompassing both single nucleotide polymorphisms (SNPs) and structural variants (SVs). The complete workflow for generating chromosome-scale assemblies of the Thai F1 hybrid catfish is presented in Figure 12. The workflow consists of six sequential steps (from Step 1A to Step 6) that integrate multiple sequencing technologies and analysis tools and pipelines in specific order used to achieve optimal phasing and Hi-C scaffolding at a high-contiguity, low structural error rate, high solve rates, and high base-accuracy genome assemblies in highly heterozygous F1 genomes. Below, I outline each step in detail. •Step 1: Genomic DNA Sequencing, Genome Survey, and Meryl Databases Genomic DNA was extracted from the liver tissue of a single F1 hybrid individual using a custom protocol (Supikamolseni et al.,2015). Sex was determined to be female based on histology and external gonadal examination (Kitano et al.,2007;Wyneken et al.,2007). Sequencing was performed using four complementary platforms generating four synergistic raw sequencing data types: (i) Pacific Biosciences (PacBio) High-Fidelity (HiFi) sequencing (Wenger et al., 2019) for generating highly accurate contigs, (ii) Oxford Nanopore Technology (ONT) long-read sequencing (Jain et al.,2018b) for scaffolding and resolving repetitive regions, and (iii and iv) Illumina short-read 150 base paired-end sequencing (Hi-C, and standard non Hi-C) (Bentley et al.,2008) for chromoso-
40 that C. macrocephalus exhibits lower heterozygosity (0.62%) compared to C. gariepinus (1.0%) (Figure ??). The final assembly achieved high quality values (QV) at both k=21 and k=31, with robust completeness metrics including CRAQ structural accuracy. The combination of manual curation, iterative polishing, and uniform read coverage proved critical for overcoming inherent limitations in both software algorithms and sequencing depth. The final assembly comprises 55 pseudochromosomes, accurately representing both parental haplotypes from a single F1 individual. During initial assembly, optimal results were achieved by generating multiple independent assemblies from the same dataset followed by post-hoc merging and scaffolding. Specifically, combining outputs from Flye and GreenHill with Hifiasm improved resolution of complex genomic regions. Under conditions of limited read depth, Hifiasm occasionally exhibited problematic behavior, extending reads and switching to longer overlapping reads despite significant nucleotide differences (>20 SNPs)—a level of variation inconsistent with sequencing error alone. Additionally, using related reference genomes to fill assembly gaps requires careful validation, as incorrect placements frequently outnumber accurate ones without manual verification. Hi-C scaffolding underwent three rounds of manual review using Juicebox, involving contig splitting at offdiagonal signals and identification of scaffold gaps, ultimately reducing contig count from >2,000 to <500. 1.3 Telomeric Refinement and Structural Improvements Telomeric sequence refinement represented a significant technical achievement. Manual extension of clipped sequences at chromosome termini successfully incorporated five additional telomeric regions containing canonical [TTAGGG]nmotifs. While telomeres may not be essential for most genomic analyses, their accurate representation enables precise chromosome boundary delineation. Notable improvements included: •C. gariepinus chromosome 16: left telomeric repeat count increased from 138 to
41 1,602 •C. gariepinus chromosome 18: left telomeric repeat count increased from 268 to 1,411 •C. macrocephalus chromosome 7: left telomeric repeat count increased from 199 to 924 These refinements increased the number of pseudochromosomes with bilateral telomeres from 43 to 48, while reducing those with unilateral telomeres from 8 to 7. Telomeric additions contributed approximately 400 kb to the total assembly length, representing substantial progress toward telomere-to-telomere (T2T) completeness for multiple chromosomes. Subgenomic analysis revealed additional differences between parental haplotypes, including heterozygosity rate variations and assembly complexity differences attributable to limited nanopore coverage. While most genomic gaps have been resolved, centromeric and satellite regions remain challenging, particularly where HiFi read coverage is absent or where only MAPQ0 Illumina reads provide support, necessitating reliance on lower-accuracy ONT data. 1.4 Quality Value Enhancement Through Manual and Automated Polishing The polishing strategy combined automated QV correction with manual curation. Initial variant calling with Clair3 identified over 60,000 variants, subsequently resolved using Merfin. Despite achieving QV60 at k=21, a final manual curation round introduced 8,410 edits, including 551 large structural variants, resulting in QV improvement of 1–5 points. The complete manual curation process comprised four distinct rounds: Round 1: Sniffles2 identified 755 target regions across all alignment files for manual review.
42 Round 2: Cross-assembly curation using Unimap examined scaffoldsfrom Flye, GreenHill, and Hifiasm P-contigs, incorporating seqmer error files. This resulted in 2,221 manual edits and application of 5,951 variants. Round 3: Combined Winnowmap and Unimap alignments enabled identification of 8,710 variants from 5,740 alignments, with 1,932 applied to the assembly. Additionally, 300 gaps were resolved following Sniffles2 and BCFtools consensusbased corrections. Round 4: Correction of 97 large structural variants, including 10 visible only through VerityMap alignments. To ensure genome reliability, 65 gaps were introduced to mask sequences lacking proper read support or exhibiting high error profiles. Post-masking, 247 structural variants (196 insertions, 73 deletions) identified using Sniffles2 were applied to the genome. These variants, derived from ONT pb-falconc-filtered alignments (https://github.com/bio-nim/pb-falconc) realigned at Hifiasm contig boundaries with 50 nt margins, were approximately 90% homozygous and 10% heterozygous. Genomic error k-mers showed non-uniform chromosomal distribution, clustering in regions with ONT-only or single-depth coverage (511 loci total). Further correction using Clair3 (Zheng et al.,2022) addressed 22,625 ONT-based consensus errors and 7,558 HiFi Winnowmap variants, achieving median QV31. Assembly validation using IGV demonstrated quality improvements and structural resolution. Chromosome 03 showed clear reduction in k-mer error rates and increased QV post-polishing. Chromosome 02 analysis revealed a false duplication present in Hifiasm but absent in Flye, confirmed by doubled homozygous coverage and lack of supporting HiFi alignments. A single spanning ONT read and clipped Illumina read validated boundary precision. Consistent single-nucleotide markers across all data types (HiFi, ONT, Flye) suggest potential for further refinement using high-confidence variants (AF > 0.9) with Clair3 (Zheng et al.,2022) and WhatsHap (Martin et al.,2016).
43 Figure 13 Quality value and coverage validation using Integrative Genomics Viewer. (A) Chromosome 03 demonstrating QV improvement following polishing procedures. (B) Chromosome 02 showing false duplication and gap analysis through comparative assembly assessment. Following eight months of genome curation, the estimated QV improved from Q40 to Q60, demonstrating that F1 hybrid genomes can achieve superior nucleotidelevel accuracy compared to single-species assemblies. For comparison, the pure-breed bighead catfish assembly achieved median QV40–50 (k=21, k=31, pileup), while the current F1 hybrid genome demonstrates ten-fold higher accuracy using comparable sequencing coverage, analytical tools, and curation effort.
44 Figure 14 F1 Hybrid Catfish Assembly status January 2024 - November 2024.
45 VALIDATION OF CATFISH ASSEMBLIES Benchmarking Methodology This chapter presents the results and technical validations performed on the genome assemblies of the two catfish species. First, I present the most common metrics to measure the quality and completeness of the assemblies, next I present each individual assembly with results, finally comparative genomic analysis are provided and illustrate how these two genomes can be used for comparative studies across species. Metrics of continuity, structural accuracy, base accuracy, and functional completeness were used for benchmarking as described in the Vertebrate Genome Project (VGP) paper (Rhie et al.,2021). 1. Metrics Assessed 1.1 Continuity and summary statistics To assess continuity and summary statistics of the assembly, I compute the following measures for scaffold/contig (N50, N90, NG50, LG50, and LG90)50 with RagTag (’ragtag.py asmstats -g’) (Alonge et al.,2022) v2.1.0. 1.2 Repeat completeness and continuity of repeats To measure repeat completeness and continuity of repeats for assessing assembly quality, I estimated the percentage of fully assembled LTR retroelements (LTRRTs)51 and computed the long terminal repeat (LTR) Assembly Index (LAI) using two 50N50 and N90 represent the contig or scaffold length such that 50% or 90% of the total assembly length is contained in contigs/scaffolds of that length or longer. NG50 is similar to N50 but calculated relative to an expected reference genome size. LG50 and LG90 indicate the minimum number of contigs or scaffolds whose combined length makes up 50% or 90% of the assembly, respectively. 51LTR-RTs (Long Terminal Repeat Retrotransposons) are a class of transposable elements characterized by direct long terminal repeats at both ends. They replicate via an RNA intermediate and reverse transcription, and are major contributors to genome size and structure in many eukaryotic organisms.
46 reference-free programs, LTR Assembly Index (LAI) (Ou et al.,2018) and LTR_retriever (Ou and Jiang,2017) v2.9.00. To assess the assembly quality of complex repeats (centromeres), I used TandemTools and TandemQUAST (Mikheenko et al.,2020) v1. Telomere prediction for the presence / absence and orientation of telomeres was done with TIDK, a Telomere Identification Toolkit Brown et al. (2023), implemented in TeloExplorer (’-m 50 -c animal’), a module of QuarTeT Lin et al. (2023). To estimate gaps, I used the script detgaps, available on GitHub (https://github.com/dfguan/asset). 1.3 Structural accuracy (regional and structural errors and reliable blocks) To assess structural accuracy I used the CRAQ software (’sms_coverage=5 ngs_coverage=20 -B T --minimap2-sensitve’) (Li et al.,2023) v1.0.9, more specifically (’-q 20 -m 2 -f 0.4 -h 0.6 -r 0.75 -a 20’). CRAQ is a method for assembly structural validation relying on the study of mapped reads, clipped reads and coverage support by two or more simultaneous sequencing platforms to find supporting regions or reliable blocks (Rhie et al.,2021).and visualized results using the Integrative Genome Viewer (IGV) (Thorvaldsdottir et al.,2012). CRAQ evaluates genome assembly quality based on clipped-read evidence from both short and long reads. It reports two main types of errors: 1. Clip-based Regional Errors (CREs), which represent small-scale local misassemblies identified by short-read clipping without structural disruption. 2. Clip-based Structural Errors (CSEs), which represent large-scale misassemblies supported by long-read clipping near breakpoint regions. These errors are quantified into two sub-scores: R-AQI (based on CREs) and S-AQI (based on CSEs). Both scores contribute to the global Assembly Quality Index (AQI), ranging from 0–100, where higher values indicate better assembly integrity. AQI is computed based on error density relative to genome size.
47 1.4 Cross-species structural correctness To assess cross-species structural correctness, I used the same references from related catfish species for one-to-one nucleotide-level alignments of orthologous segments with MashMap252 (’-s 2000000 --pi 90 -c 100000’) (Jain et al.,2018a) v3.1.3. 1.5 Base accuracy and assembly completeness To assess base accuracy and completeness, specifically Mercury’s quality values (QV) and 21-mer genome completeness (%), I used Mercury (Rhie et al.,2020) v1.3. Merqury was run three times using different k-mer databases: Illumina, HiFi, and a hybrid 21-mer database combining Illumina and HiFi reads, as explained above, in the ”Consensus polishing” section. To assess functional completeness and evaluate the completeness of the gene set, I used the lineage of ray-finned fishes and BUSCO (’-l actinopterygii_odb10’) (Simão et al.,2015) v5.6.1. Finally, nucleotide accuracy was verified by read-to-assembly mapping, performed using minimap2 (’-ax map-hifi --secondary=no’) and Winnowmap (’-W repetitive.15.txt’) at MAPQ >10 for HiFi reads, while for ONT reads I used the '-ax map-ont' preset for minimap2 and the '-ax map-pb' preset in Winnowmap MAPQ >10. For Visualisations, I used IGV (’igv -g $genome $BAM(s) $merqury_only_bed_wig_kmer_files’). 52MashMap2 is a fast and approximate algorithm for computing whole-genome homology maps. It uses minimizer-based locality-sensitive hashing to rapidly identify high-confidence regions of sequence similarity, making it suitable for genome-to-genome alignments, structural variation detection, and referenceguided scaffolding (Jain et al.,2018a).
48 Results – Bighead Catfish The haplotype-resolved de novo assembly resulted in 27 Hi-C scaffolds, and the chromosome number was 2n = 2x = 54 pseudochromosomes (Maneechot et al.,2016). 1. Raw Read Quality Figure 15 summarizes raw data quality for the bighead catfish genome. Figure 15 Sequencing summary and genome survey of bighead catfish. (A) Sequencing protocols and coverages. (B) GenomeScope k-mer profile showing genome size, repeat content, and 17.6% heterozygosity. (C) HiFi reads: high quality and 15 kbp peak length. (D) ONT reads: broader size range with N50 30 kbp and moderate quality.
49 2. Global Assembly Metrics Sequence variation analysis53 was performed with PlotSR suite (Goel and Schneeberger,2022), revealing 1,968,666 heterozygous SNPs across both haplotypes after minimap2 ('-ax asm5') cross-haplotype mapping, resulting in a heterozygosity rate54 of 0.594%. Structural variant analysis identified 392,973 insertions totaling 7.67 Mb in Haplotype 2, while Haplotype 1 had 393,127 deletions spanning 7.75 Mb. Copy number variations55 included 114 copy gains in Haplotype 2 (184 Kb) and 123 copy losses in Haplotype 1 (416 Kb). Highly divergent regions56 were identified, spanning 57.9 Mb in Haplotype 1 and 56.0 Mb in Haplotype 2. Additionally, 47 tandem repeat clusters57 were detected, covering 10 Kb in Haplotype 1 and 7 Kb in Haplotype 2. Table 1 Summary statistics of the haplotype-resolved Clarias macrocephalus genome assembly. Features Haplotype 1 Haplotype 2 Overall quality (x.y.P.Q.C) 6.30.P7.53.Q48.C94.25 6.30.P7.53.Q48.C94.25 Genome size (Mb) 875 Mb 880 Mb pseudochromosomes 27 27 Mitochondrion length (bp) 16,510 bp N/A NG50 of contigs (Mb) 3 Mb 3 Mb N50 of scaffolds (Mb) 34 Mb 34 Mb LG50/LG90 of scaffolds 11 / 24 (Hap1) 11 / 24 (Hap2) Number of gaps 752 653 GC content (%) 39.32% 39.32% Complete BUSCOs N (%) 93.0% (Hap1) 93.5% (Hap2) Median QV (Merqury k21) 42.99 (Hap1) 43.33 (Hap2) 48.22 (All) Completeness (Merqury k21) 89.11 (Hap1) 88.88 (Hap2) 95.25 (All) 53Comparison of Haplotype 1 and 2 chromosomes to detect heterozygous variants such as SNPs, indels, and structural changes. 54Proportion of sites where the two haplotypes differ, reflecting allelic diversity. 55Duplications or deletions causing variable copy numbers of genomic regions. 56Genomic segments showing strong sequence divergence between haplotypes or species. 57Short motifs repeated head-to-tail, often in telomeres or centromeres.
56 5. Scaffold Metrics Per-scaffold analyses confirmed uniform base and structural accuracy across chromosomes. Key quality metrics included base accuracy (median QV ≈ 50), gap content, and telomere orientation. Structural integrity was assessed using the Structural Assembly Quality Index (S-AQI %), base-level accuracy with Merqury (k=21), structural precision with CRAQ, telomere polarity with QuarTeT (+ inward, – outward, > 100 repeats), and gap statistics with Detgaps.
57 Table 2 Summary of scaffold metrics in the haplotype-resolved genome assembly of Clarias macrocephalus,Haplotype 1. Scaffold Name Size Errors Quality Values Gaps Left Telomere Right Telomere S-AQI Haplotype_Linkage (Mb) 21-mer hyb.k21.meryl (N) (5’) AACCCT) (AGGGTT 3’) (%) fClaMac_1_LG_01 48,4 3242 54.8396 23 Left 507 (-) -93.24 fClaMac_1_LG_02 46,18 6400 51.574 24 Left 607 (+) Right 1016 (-) 73.94 fClaMac_1_LG_03 40,54 6784 50.6951 20 - Right 215 (-) 91.18 fClaMac_1_LG_04 37,64 4160 52.7678 23 Left 147 (+) Right 122 (-) 100.00 fClaMac_1_LG_05 37,19 9007 49.2436 23 Left 579 (+) Right 190 (-) 93.23 fClaMac_1_LG_06 37,01 5548 51.4013 25 - - 84.09 fClaMac_1_LG_07 34,64 7089 49.8928 25 Left 214 (+) - 93.24 fClaMac_1_LG_08 47,03 12668 48.7208 32 Left 170 (+) Right 190 (-) 89.37 fClaMac_1_LG_09 28,28 7205 49.2196 13 Left 153 (+) Right 1630 (-) 81.31 fClaMac_1_LG_10 29,87 12042 47.0424 24 Left 282 (+) - 95.23 fClaMac_1_LG_11 25,51 6087 49.3463 19 Left 1008 (+) - 80.92 fClaMac_1_LG_12 28,04 8367 48.2834 21 Left 124 (+) - 95.73 fClaMac_1_LG_13 24,08 3715 51.2825 18 Left 377 (+) - 86.62 fClaMac_1_LG_14 21,22 4707 49.7286 13 - Right 136 (-) 94.05 fClaMac_1_LG_15 25,46 3974 51.1515 20 Left 127 (+) Right 115 (-) 95.00 fClaMac_1_LG_16 24,8 4054 50.9451 15 - Right 109 (-) 73.65 fClaMac_1_LG_17 22,98 67523 38.5013 15 Left 120 (+) - 77.54 fClaMac_1_LG_18 24,8 2060 53.9205 13 - - 90.26 fClaMac_1_LG_19 31,01 3329 52.8263 22 - Right 363 (-) 92.45 fClaMac_1_LG_20 30,49 4074 51.8251 15 - Right 406 (-) 83.79 fClaMac_1_LG_21 28,95 4609 51.2 28 - - 100.00 fClaMac_1_LG_22 33,48 2784 54.0062 23 - Right 904 (-) 92.37 fClaMac_1_LG_23 31,45 8435 48.8646 16 - Right 790 (-) 86.84 fClaMac_1_LG_24 28,07 4984 50.6636 12 Left 423 (+) - 88.33 fClaMac_1_LG_25 38,33 6711 50.6434 24 - - 96.55 fClaMac_1_LG_26 40,06 4967 52.1555 21 Left 556 (+) Right 604 (-) 87.85 fClaMac_1_LG_27 29,17 2146 54.5417 13 - - 74.96
58 Table 3 Summary of scaffold metrics in the haplotype-resolved genome assembly of Clarias macrocephalus,Haplotype 2. Scaffold Name Size Errors Quality Values Gaps Left Telomere Right Telomere S-AQI Haplotype_Linkage (Mb) 21-mer hyb.k21.meryl (N) (5’) AACCCT) (AGGGTT 3’) (%) fClaMac_2_LG_01 51,51 2649 55.5298 17 - - 89.28 fClaMac_2_LG_02 44,94 6893 51.161 25 Left 255 (+) - 71.62 fClaMac_2_LG_03 39,44 8894 49.5912 16 - - 83.57 fClaMac_2_LG_04 38,01 6912 50.5588 18 Left 293 (+) Right 526 (-) 88.17 fClaMac_2_LG_05 38,16 3424 53.4601 23 - Right 125 (-) 78.26 fClaMac_2_LG_06 38,66 2675 54.3325 17 - - 80.86 fClaMac_2_LG_07 33,93 5651 50.8217 17 Left 169 (+) - 89.91 fClaMac_2_LG_08 47,19 10458 49.5856 29 Left 174 (+) Right 190 (-) 89.44 fClaMac_2_LG_09 28,8 10239 47.5736 14 Left 641 (+) - 91.91 fClaMac_2_LG_10 29,07 5001 50.8593 14 Left 268 (+) - 95.24 fClaMac_2_LG_11 25,76 6627 48.9338 21 - - 89.75 fClaMac_2_LG_12 30,93 6915 48.7951 19 - - 82.58 fClaMac_2_LG_13 23,66 3213 51.9117 15 - Right 100 (+) 86.65 fClaMac_2_LG_14 21,06 3487 50.9324 7 - - 100.00 fClaMac_2_LG_15 24,93 7673 48.2988 17 Left 134 (+) Right 137 (-) 85.58 fClaMac_2_LG_16 24,83 3011 52.2777 10 - - 81.74 fClaMac_2_LG_17 22,34 2012 53.6414 15 Left 175 (+) - 85.99 fClaMac_2_LG_18 24,44 1986 54.0521 11 - - 100.00 fClaMac_2_LG_19 30,86 3566 52.521 16 - Right 362 (-) 90.62 fClaMac_2_LG_20 34,51 2384 53.5357 13 - - 100.00 fClaMac_2_LG_21 27,99 2550 53.5741 13 - - 100.00 fClaMac_2_LG_22 33,35 4588 51.8225 16 Left 128 (+) Right 908 (-) 88.64 fClaMac_2_LG_23 31,45 9701 48.2459 6 Left 214 (+) Right 793 (-) 82.68 fClaMac_2_LG_24 27,39 2840 53.0349 6 Left 114 (+) - 93.87 fClaMac_2_LG_25 37,88 7891 49.9914 15 - - 87.01 fClaMac_2_LG_26 40,17 11271 48.6822 13 - Right 607 (-) 88.08 fClaMac_2_LG_27 29,56 1236 56.9751 12 Left 155 (+) - 90.78 MT_from_NC_046749 0,02 - - - - - -
59 6. TE Metrics 6.1 Dating of TE Replicative Bursts To estimate the expansion history of major transposable element (TE) superfamilies in Clarias macrocephalus, I calculated the nucleotide divergence of each TE copy relative to its consensus sequence and converted this divergence into approximate insertion ages using a neutral mutation rate58 ( µ =6×10−9per site per generation) previously determined in electric catfish (Liu et al.,2023b). Assuming a one-year generation time, the estimated expansion peaks indicate that LTR/Gypsy and TIR/Mutator families expanded around 30–33 Mya, LTR/Copia and TIR/hAT around 28–30 Mya, and TIR/CACTA, PIF/Harbinger, and Tc1/Mariner families between 16 and 25 Mya. These results (Figure 20) reveal multiple waves of TE proliferation, likely reflecting episodes of genome instability or evolutionary transition in the species’ history. Figure 20 TE divergence profiles in bighead catfish genome. 58The neutral mutation rate is the rate at which mutations accumulate in non-functional genomic regions, providing a baseline for molecular-clock dating.
60 6.2 Genome composition in TE Transposable element (TE) annotation indicated that 35.25% of the bighead catfish genome consisted of repetitive sequences. Among these, TIR DNA transposons comprised 19.12%, Helitrons 4.47%, LTR retrotransposons 8.30%, LINE elements 0.46%, and the remaining 2.82% consisted of uncharacterized repeats. Table 4 Transposable element content in Bighead catfish Haplotype 1. Class Superfamily Haplotype 1 Haplotype 2 Features # of elements LINE I128 L1 547 L2 1,129 Rex 556 LINE/Unknown 7,412 LTR Bel_Pao 151 Copia 53 Gypsy 68,215 LTR/Unknown 13,466 TIR CACTA 357,388 Mutator 172,578 PIF_Harbinger 36,538 Tc1_Mariner 41,500 hAT 80,189 PiggyBac 758 Polinton 281 NonLTR DIRS_YR 678 Penelope 417 NonTIR Helitron 166,251 Other Repeat Fragment 78,900 Total Interspersed 1,148,329
61 LINEs (Long Interspersed Nuclear Elements) are non-LTR retrotransposons that transpose via a copy-and-paste mechanism using reverse transcriptase; although less abundant in fish genomes than in mammals, they contribute to structural variation. LTR retrotransposons replicate through an RNA intermediate using reverse transcription and are flanked by long terminal repeats, playing key roles in genome expansion and gene regulation. TIR DNA transposons (Terminal Inverted Repeat transposons) are cut-and-paste elements flanked by inverted repeats recognized by a transposase enzyme, contributing to genome rearrangement and regulatory diversification. Non-LTR retrotransposons such as DIRS and Penelope elements add further diversity to the repetitive landscape through distinct mobilization mechanisms. Helitrons, belonging to the NonTIR class, are rolling-circle DNA transposons that lack terminal inverted repeats and transpose through a mechanism similar to rolling-circle replication, often capturing and mobilizing gene fragments that promote genomic innovation. Together, these TE classes reflect a dynamic evolutionary history for C. macrocephalus, consistent with high DNA transposon activity typical of teleost genomes. 7. Macrosynteny Analysis To examine orthologous relationships and identify syntenic regions between catfish and related teleost species, I obtained proteomes (.pep.fa files) from Ensembl v2024 (Harrison et al.,2023), ensuring the latest assembly versions for consistency. Species are Oncorhynchus mykiss (USDA_OmykA_1.1), Oryzias latipes (ASM223467v1), Lates calcarifer (ASB_HGAPassembly_v1), Cyprinus carpio (Cypcar_WagV4.0), Danio rerio (GRCz11), Oreochromis niloticus (O_niloticus_UMD_NMBU), and Lepisosteus oculatus (LepOcu1). Orthologous gene clusters (Orthogroups) were first identified with OrthoFinder (’orthofinder -f /proteomes -M msa -a 126 -t 126’) (Emms and Kelly,2019) v.2.5.5. Subsequently, I utilized xrhbexpress (https://github.com/SamiLhll/ rbhXpress) and macrosyntR (El Hilali and Copley,2023) v0.2.19 to analyze the resulting orthogroups and visualize their syntenic relationships across species. This combined approach allowed for an in-depth view of conserved and divergent genomic regions within the catfish lineage and related teleosts (Figure 21).
62 Figure 21 Synteny analysis for various catfish samples. (A) Whole-genome 4-way synteny analysis from 1413 shared single-copy orthogroups across Clarias gariepinus (GCA_024256425.2), Clarias macrocephalus,Danio rerio (GRCz11)Danio rerio and Oreochromis niloticus (O_niloticus_UMD_NMBU), reveals conserved chromosome structures and rearrangements. (B) Simple phylogeny and geological timescales (TimeTree5, https://timetree. org/).
63 Results – F1 Hybrid Catfish The haplotype-resolved de novo assembly of the F1 hybrid catfish (C. macrocephalus ×C. gariepinus) yielded a total of 55 Hi-C scaffolds, corresponding to the combined 27 + 28 pseudochromosomes inherited from the parental species (Lewin et al., 2019;Maneechot et al.,2016). The resulting assembly spans both subgenomes, representing the first complete haplotype-resolved hybrid genome for aquaculture catfishes. 1. Raw Read Quality Figure 22 summarizes the sequencing data quality and genome survey statistics for the hybrid genome.
64 Figure 22 Sequencing summary and genome survey of the F1 hybrid catfish (C. macrocephalus ×C. gariepinus). (A) GenomeScope2.0 k-mer profile (k=21) showing estimated haploid genome size (903 Mb), 10.1 % heterozygosity, and low sequencing error rate (0.28 %). (B) Sequencing protocols and mapped coverages across data types, including HiFi (30×), ONT (36×), Hi-C (37×), and Illumina WGS (55×). (C) HiFi reads: high quality with a modal length around 15 kbp. (D) ONT reads: broader length distribution with N50 ≈ 30 kbp and moderate quality.
65 1.1 Illumina paired-end short-reads Standard Illumina paired-end short-reads sequencing Bentley et al. (2008) was performed on the Illumina NextSeq2000 platform (by NovoGene), resulting in more than 50 Gb of raw read data (read-length=151 nucleotides), which is equivalent to a genome coverage of 36X. Raw reads had an average quality of Q25, and 92% of the dataset had a median quality of Q30. 1.2 Proximity-ligation (Hi-C) paired-end short-reads Proximity-ligation (Hi-C) data was obtained by following a standard in situ Hi-C protocol from 2009 Dudchenko et al. (2017). Briefly, a nuclear ligation was performed by cross-linking chromosomes, then a restriction digestion was carried out with DpnII endonuclease. The chromatin conformation capture library was prepared using Phase Genomics (https://info.phasegenomics.com/protocols), and the genomic DNA was sequenced on the Illumina NextSeq2000 platform, resulting in paired-end short-reads (151 nucleotides in length) for 37.19 Gb of raw data (approximately 100M pairs of reads per billion bases in genome length) corresponding to a genome coverage of 39.77X. 97.18% of the dataset had a quality score of Q30 or higher (Supplementary Table 3). For Haplotype 1, 123.96M read pairs were sequenced, yielding 74.77M unique Hi-C contacts, with 59.40M valid contacts (22.71M inter-chromosomal, 36.69M intrachromosomal), including 17.47M short-range (<20Kb) and 19.72M long-range (>20Kb) contacts. Haplotype 2 showed similar metrics. 1.3 PacBio HiFi long-reads PacBio’s HiFi long-read sequencing was performed using a SMRT cell on the Sequel II system, resulting in 10.62 Gb of raw data corresponding to a genome coverage of approximately 12.9X or 6-7X per haplotype. The read length N50 value was 10,379 bases, the raw reads’ mean quality score was Q28.7, and the median quality score was Q36. Over 72% were above Q30, and 86% above Q25. There were 191,790
72 5. Genome Completeness Evaluation To further validate gene content and completeness, BUSCO analysis (Figure 25) was performed, yielding high completeness scores. For the North African catfish sub-genome, 97.1% of core genes were complete [95.7% single-copy, 1.4% duplicated], with 2.2% missing. The bighead catfish sub-genome showed similar completeness, with 96.6% of core genes [95.3% single-copy, 1.3% duplicated] and 2.7% missing. This analysis underscores the high integrity and completeness of gene content within the assembly. Notably, the F1 genome exceeded 97% BUSCO completeness (Fig. 25B) and achieved a median QV of 55 (Merqury Figs. 25D, 25E), with no multiplicity peaks over 2×—indicating near-reference-grade assemblies. Together with k-mer analysis and BUSCO metrics, these profiles confirm the F1 hybrid assembly’s completeness, chromosome number consistency (27 + 28), and dual subgenome contribution, while BUSCO-derived gene synteny in Clarias confirms the F1 hybrid nature of the genome two subspecies (Fig. 25F). One mitochondrial sequence, confirmed through IGV polymorphism analysis, belongs to C. macrocephalus. Other results were a final assembly k-mer completeness estimated at 99.4% (k=21, Merqury) to 99.38% (k=31, Merqury),
73 Figure 25 Overview of F1 Hybrid Catfish (Clarias gariepinus ×Clarias macrocephalus) genome assembly validation. Hi-C phasing confirms 55 pseudochromosomes, with BUSCO scores indicating 97% gene completeness. Synteny analysis and k-mer distributions support the F1 hybrid nature, demonstrating chromosome separation into sub-genomes and their evolutionary affiliations.
74 6. Scaffold Metrics 6.1 Assembly Statistics and Quality Assessment Assembly metrics are summarized in Table 5and Table 6, including scaffold count, length in Mb, QV, error rates, and scaffold length in bases, with a gradient indicating QV quality: red for QV < 50, yellow for QV between 50-60, and green for QV > 60. The leftmost column displays the count of error seqmers (k-mers present in the assembly but absent in reads), indicating potential assembly errors. Alongside are given the relative telomeres orientation (5’ +inward -> / <- inward -3’), and gaps count in chromosomes. Structural validation, performed using 3D-DNA and JBAT, confirmed scaffold accuracy, with additional polishing steps through NextPolish2, TGS GapCloser, Pilon, and automated consensus polishing using ONT 1D reads, Racon, Merfin, and BCFtools, resulting in a final quality score of QV50, a 99.999% base precision, 99% mapping rate, and 99.4% k-mer completeness. Despite the presence of complex repeats, the assembly achieved CRAQ structural accuracy (S-AQI) above 97%, indicating reference-grade quality (Li et al.,2023). 6.2 Comparative Analysis of Subgenomes North African catfish scaffolds exhibit consistently higher QVs (up to ~65.79) than bighead catfish scaffolds (ranging ~50–54), indicating fewer sequence errors and higher assembly quality overall. Several chromosomes exhibit Telomere-to-Telomere (T2T) continuity, validating successful resolution of full-length chromosomes. North African catfish scaffolds contain fewer unclosed gaps than bighead catfish, reflecting superior completeness and assembly contiguity. Note that this may result from local genome complexity or differences in transposable element content and diversity, aside from variations in raw sequencing data quality.
75 Table 5 North African subgenome scaffold metrics with Quality Values from Merqury (k=21 and k=31), telomere orientation, and gap counts. Scaffold_name Length QV_k21 Errors_k21 QV_k31 Errors_k31 Left_telomere Gaps Right_telomere Sub-Genome_Chromosome Mb. 21-mer 21-mer 31-mer 31-mer (5’ AACCCT) (N) (AGGGTT 3’) fClaHyb_african_1_Chr_01 51,1 64.5442 377 59.4881 1783 119 (+) T2T 0 fClaHyb_african_1_Chr_02 51,7 56.0219 2713 56.0911 3942 174 (+) 10 fClaHyb_african_1_Chr_03 47,6 59.3028 1174 54.5941 5125 1701 (+) 31603 (-) fClaHyb_african_1_Chr_04 44,3 49.4714 10490 56.9825 2716 185 (+) 8 160 (-) fClaHyb_african_1_Chr_05 41,5 62.5139 488 61.5864 892 1472 (+) T2T 1883 (-) fClaHyb_african_1_Chr_06 41,3 54.5824 3021 56.2081 3068 278 (+) 11447 (-) fClaHyb_african_1_Chr_07 42,6 58.0864 1389 54.3111 4889 384 (+) 3382 (-) fClaHyb_african_1_Chr_08 40,2 54.3182 3124 59.8602 1287 0 T2T 130 (-) fClaHyb_african_1_Chr_09 37,3 55.9044 2013 53.8378 4784 1449 (+) 1315 (-) fClaHyb_african_1_Chr_10 34,9 67.6478 126 60.3736 993 0 2127 (-) fClaHyb_african_1_Chr_11 33,5 51.4996 4986 58.2486 1555 1614 (+) T2T 1997 (-) fClaHyb_african_1_Chr_12 35,3 54.4649 2653 58.1993 1657 0 6 0 fClaHyb_african_1_Chr_13 33,6 66.4695 159 63.3235 483 1525 (+) 21557 (-) fClaHyb_african_1_Chr_14 32,4 47.2905 12717 61.6029 690 105 (+) T2T 226 (-) fClaHyb_african_1_Chr_15 35 64.2407 275 58.2795 1602 328 (+) 2118 (-) fClaHyb_african_1_Chr_16 33,8 60.5363 627 57.9679 1675 1602 (+) 4137 (-) fClaHyb_african_1_Chr_17 30,3 55.2666 1892 56.604 2052 335 (-) T2T 276 (-) fClaHyb_african_1_Chr_18 30,5 56.641 1382 57.9088 1522 1411 (+) 32289 (-) fClaHyb_african_1_Chr_19 30,7 64.5469 226 58.6026 1311 784 (+) 2178 (+) fClaHyb_african_1_Chr_20 27,1 57.0276 1128 60.0252 835 171 (+) 2406 (-) fClaHyb_african_1_Chr_21 30,6 60.6832 549 60.1377 919 1936 (+) 20 fClaHyb_african_1_Chr_22 30 62.328 369 60.2163 886 0 2129 (-) fClaHyb_african_1_Chr_23 26,2 52.9697 2773 58.0814 1260 279 (+) 11739 (-) fClaHyb_african_1_Chr_24 25,9 60.3587 501 60.6659 689 185 (+) 2114 (-) fClaHyb_african_1_Chr_25 25,6 58.5637 748 59.0404 989 232 (+) T2T 129 (-) fClaHyb_african_1_Chr_26 26,3 56.7889 1163 57.358 1506 1343 (+) 30 fClaHyb_african_1_Chr_27 24 68.2243 76 59.5233 832 0 T2T 264 (-) fClaHyb_african_1_Chr_28 20,3 59.8335 443 56.9713 1261 0 20
76 Table 6 Bighead subgenome scaffold metrics with Quality Values from Merqury (k=21 and k=31), telomere orientation, and gap counts. Scaffold_name Length QV_k21 Errors_k21 QV_k31 Errors_k31 Left_telomere Gaps Right_telomere Sub-Genome_Chromosome Mb. 21-mer 21-mer 31-mer 31-mer (5’ AACCCT) (N) (AGGGTT 3’) fClaHyb_bighead_1_Chr_01 50,1 55.7992 2767 50.7048 13207 998 (+) 14 740 (-) fClaHyb_bighead_1_Chr_02 44,3 51.8276 6102 54.4421 4934 815 (+) 50 fClaHyb_bighead_1_Chr_03 40,4 44.33 31251 55.4715 3531 327 (+) 12 279 (-) fClaHyb_bighead_1_Chr_04 38,4 52.1302 4933 52.1322 7281 222 (+) 13 819 (-) fClaHyb_bighead_1_Chr_05 38,1 47.0334 15835 53.7219 4994 321 (+) 14 301 (-) fClaHyb_bighead_1_Chr_06 37,4 52.5469 4370 50.6785 9921 282 (+) 14 0 fClaHyb_bighead_1_Chr_07 33,8 50.998 5644 48.4795 14888 924 (+) 15 629 (-) fClaHyb_bighead_1_Chr_08 45,6 56.2767 2252 51.0889 10984 161 (+) 11 1237 (-) fClaHyb_bighead_1_Chr_09 30 48.1331 9677 48.5674 12936 184 (+) 13 751 (-) fClaHyb_bighead_1_Chr_10 29,8 54.8059 2065 49.5164 10314 316 (+) 12 664 (-) fClaHyb_bighead_1_Chr_11 25 51.0866 4085 53.4888 3469 0 21 267 (-) fClaHyb_bighead_1_Chr_12 31,6 48.5768 9241 47.4853 17546 559 (+) 13 0 fClaHyb_bighead_1_Chr_13 24,9 44.8933 16909 48.6904 10418 740 (+) 8220 (-) fClaHyb_bighead_1_Chr_14 23,3 44.5272 17249 51.8503 4715 0 8190 (-) fClaHyb_bighead_1_Chr_15 27,6 39.2488 68816 52.3266 4848 200 (+) 12 107 (-) fClaHyb_bighead_1_Chr_16 25,2 51.936 3389 53.2655 3686 656 (+) 9818 (-) fClaHyb_bighead_1_Chr_17 24 40.0722 49568 46.7745 15616 326 (+) 15 222 (-) fClaHyb_bighead_1_Chr_18 25,2 52.6726 2856 52.3131 4582 632 (+) 11 973 (-) fClaHyb_bighead_1_Chr_19 31,1 53.8536 2692 52.9809 4841 674 (+) 8751 (-) fClaHyb_bighead_1_Chr_20 32,7 51.6697 4681 51.0227 7929 410 (+) 18 422 (-) fClaHyb_bighead_1_Chr_21 29,2 43.0312 30479 51.9822 5662 805 (+) 16 253 (-) fClaHyb_bighead_1_Chr_22 34,9 51.0828 5705 50.9355 8715 568 (+) 13 121 (-) fClaHyb_bighead_1_Chr_23 31,7 42.7592 35295 48.8582 12797 931 (+) 81022 (-) fClaHyb_bighead_1_Chr_24 29,6 50.1103 6080 52.8055 4829 851 (+) 10 159 (-) fClaHyb_bighead_1_Chr_25 40,3 52.391 4875 52.1653 7582 0 9798 (-) fClaHyb_bighead_1_Chr_26 40,1 53.1914 4035 51.6319 8529 1038 (+) 11 0 fClaHyb_bighead_1_Chr_27 30,5 48.7036 8641 48.5145 13324 1858 (+) 10 219 (-)
77 6.3 Structural Accuracy Assessment Table 7and Table 8present structural accuracy metrics from CRAQ for both sub-genomes, including Structural Accuracy Index (S-AQI), Regional Accuracy (R-AQI), and heterozygous loci for structural and regional metrics (Avg.CSH/Avg.CRH). Note that CRH and CSH values are expected to be close to 0 due to the high genomic diversity between the two parental species (divergence > 10%, mash-based (Ondov et al., 2019), not shown here). North African catfish scaffolds displayed near-perfect S-AQI and R-AQI values, both ranging approximately from 97% to 100%, indicating minimal misassemblies at both large-scale and regional-scale levels. In contrast, bighead catfish scaffolds exhibited slightly lower accuracy, with S-AQI values around 90–97% and R-AQI between 90–96%. This reduction in quality may be attributed to factors such as lower species heterozygosity, increased local assembly complexity due to transposable element content, or potential sequencing quality issues, possibly arising from degraded genomic DNA used in ONT sequencing.
78 Table 7 North African subgenome structural accuracy metrics from CRAQ: S-AQI, R-AQI, and heterozygous regions (Avg.CSH/Avg.CRH). Scaffold_name Mapping.rate Avg.CRH Avg.CRE Avg.CSE Regional-AQI Avg.CSH Structural-AQI Sub-Genome_Chromosome ONT/PE (%) Count Count Count Accuracy (%) Count Accuracy (%) Genome ( 55 Chromosomes) >97% 0.046 0.270 0.021 97.33 0.001 97.92 fClaHyb_african_1_Chr_01 0.992 0.020 0.078 0.000 99.22 0.000 100.00 fClaHyb_african_1_Chr_02 0.981 0.000 0.098 0.000 99.02 0.000 100.00 fClaHyb_african_1_Chr_03 0.955 0.033 0.154 0.000 98.47 0.000 100.00 fClaHyb_african_1_Chr_04 0.980 0.046 0.207 0.046 97.95 0.000 95.50 fClaHyb_african_1_Chr_05 0.986 0.048 0.097 0.000 99.04 0.000 100.00 fClaHyb_african_1_Chr_06 0.998 0.024 0.085 0.000 99.15 0.024 100.00 fClaHyb_african_1_Chr_07 0.978 0.000 0.120 0.000 98.81 0.000 100.00 fClaHyb_african_1_Chr_08 0.992 0.025 0.100 0.000 99.00 0.000 100.00 fClaHyb_african_1_Chr_09 0.992 0.076 0.083 0.027 99.17 0.000 97.34 fClaHyb_african_1_Chr_10 0.998 0.066 0.172 0.000 98.29 0.000 100.00 fClaHyb_african_1_Chr_11 0.905 0.030 0.188 0.000 98.13 0.000 100.00 fClaHyb_african_1_Chr_12 0.873 0.086 0.230 0.000 97.72 0.000 100.00 fClaHyb_african_1_Chr_13 0.917 0.075 0.119 0.000 98.81 0.000 100.00 fClaHyb_african_1_Chr_14 0.990 0.031 0.062 0.062 99.38 0.000 93.96 fClaHyb_african_1_Chr_15 0.922 0.076 0.152 0.000 98.49 0.000 100.00 fClaHyb_african_1_Chr_16 0.952 0.000 0.248 0.000 97.55 0.000 100.00 fClaHyb_african_1_Chr_17 0.997 0.033 0.083 0.000 99.18 0.000 100.00 fClaHyb_african_1_Chr_18 0.954 0.175 0.315 0.000 96.90 0.000 100.00 fClaHyb_african_1_Chr_19 0.998 0.000 0.294 0.033 97.10 0.000 96.79 fClaHyb_african_1_Chr_20 0.995 0.056 0.148 0.000 98.53 0.000 100.00 fClaHyb_african_1_Chr_21 0.986 0.033 0.131 0.098 98.70 0.000 90.65 fClaHyb_african_1_Chr_22 0.992 0.067 0.134 0.034 98.67 0.000 96.70 fClaHyb_african_1_Chr_23 0.998 0.058 0.186 0.038 98.16 0.000 96.23 fClaHyb_african_1_Chr_24 0.999 0.077 0.232 0.000 97.71 0.000 100.00 fClaHyb_african_1_Chr_25 0.989 0.039 0.118 0.000 98.83 0.000 100.00 fClaHyb_african_1_Chr_26 0.994 0.057 0.324 0.000 96.81 0.000 100.00 fClaHyb_african_1_Chr_27 0.980 0.042 0.166 0.000 98.35 0.000 100.00 fClaHyb_african_1_Chr_28 0.995 0.000 0.297 0.000 97.07 0.000 100.00
79 Table 8 Bighead subgenome structural accuracy metrics from CRAQ: S-AQI, R-AQI, and heterozygous regions (Avg.CSH/Avg.CRH). Scaffold_name Mapping.rate Avg.CRH Avg.CRE Avg.CSE Regional-AQI Avg.CSH Structural-AQI Sub-Genome_Chromosome ONT/PE (%) Count Count Count Accuracy (%) Count Accuracy (%) Genome ( 55 Chromosomes) >97% 0.046 0.270 0.021 97.33 0.001 97.92 fClaHyb_bighead_1_Chr_01 0.981 0.000 0.329 0.000 96.76 0.000 100.00 fClaHyb_bighead_1_Chr_02 0.992 0.046 0.148 0.023 98.53 0.000 97.75 fClaHyb_bighead_1_Chr_03 0.980 0.050 0.276 0.000 97.28 0.000 100.00 fClaHyb_bighead_1_Chr_04 0.990 0.000 0.370 0.026 96.36 0.000 97.40 fClaHyb_bighead_1_Chr_05 0.972 0.053 0.360 0.160 96.46 0.000 85.22 fClaHyb_bighead_1_Chr_06 0.934 0.000 0.497 0.027 95.15 0.000 97.33 fClaHyb_bighead_1_Chr_07 0.990 0.030 0.403 0.030 96.05 0.000 97.06 fClaHyb_bighead_1_Chr_08 0.987 0.000 0.429 0.044 95.80 0.000 95.66 fClaHyb_bighead_1_Chr_09 0.986 0.051 0.422 0.067 95.87 0.000 93.47 fClaHyb_bighead_1_Chr_10 0.982 0.034 0.460 0.034 95.50 0.000 96.65 fClaHyb_bighead_1_Chr_11 0.961 0.042 0.382 0.042 96.26 0.000 95.92 fClaHyb_bighead_1_Chr_12 0.946 0.194 0.606 0.069 94.12 0.000 93.38 fClaHyb_bighead_1_Chr_13 0.989 0.041 0.509 0.081 95.04 0.000 92.19 fClaHyb_bighead_1_Chr_14 0.982 0.080 0.544 0.000 94.71 0.044 100.00 fClaHyb_bighead_1_Chr_15 0.922 0.136 0.506 0.039 95.07 0.000 96.19 fClaHyb_bighead_1_Chr_16 0.993 0.040 0.240 0.000 97.63 0.000 100.00 fClaHyb_bighead_1_Chr_17 0.985 0.042 0.793 0.042 92.38 0.000 95.87 fClaHyb_bighead_1_Chr_18 0.977 0.000 0.474 0.000 95.37 0.000 100.00 fClaHyb_bighead_1_Chr_19 0.992 0.000 0.341 0.000 96.65 0.000 100.00 fClaHyb_bighead_1_Chr_20 0.971 0.000 0.236 0.062 97.66 0.000 93.96 fClaHyb_bighead_1_Chr_21 0.977 0.064 0.591 0.000 94.26 0.000 100.00 fClaHyb_bighead_1_Chr_22 0.935 0.102 0.406 0.044 96.02 0.000 95.74 fClaHyb_bighead_1_Chr_23 0.990 0.000 0.303 0.032 97.02 0.000 96.87 fClaHyb_bighead_1_Chr_24 0.992 0.303 0.608 0.000 94.10 0.000 100.00 fClaHyb_bighead_1_Chr_25 0.990 0.062 0.350 0.025 96.56 0.000 97.53 fClaHyb_bighead_1_Chr_26 0.991 0.050 0.378 0.025 96.29 0.000 97.51 fClaHyb_bighead_1_Chr_27 0.979 0.033 0.447 0.000 95.63 0.000 100.00
80 GENOME ASSEMBLY, REVERSE VACCINOLOGY, AND QUALITY BY DESIGN — Streptococcus iniae Introduction 1. Streptococcus iniae as a Major Aquaculture Pathogen Streptococcus iniae, first identified in an Amazon River dolphin (Inia geoffrensis) in the 1970s (Pier and Madin,1976), has emerged as a major bacterial pathogen in global aquaculture. The pathogen causes annual losses exceeding USD 100 million worldwide (Shoemaker et al.,2001), with mortality rates of 30–80% during outbreaks (Chen et al.,2012) and cumulative mortality up to 70% after three months in some species (Mmanda et al.,2014). Detected across all continents (Baiano and Barnes, 2009;Mishra et al.,2018), S. iniae affects over 27 fish species (Agnew and Barnes, 2007), particularly in intensive aquaculture systems where high stocking densities and environmental stressors facilitate rapid disease transmission (Chen et al.,2012). Economically important species including Asian seabass (Lates calcarifer), tilapia (Oreochromis spp.), and catfish (Siluriformes spp.) are particularly susceptible (Azmai and Saad,2011;Nawawi et al.,2008). While rare, zoonotic transmission to humans handling infected fish has been documented (Facklam et al.,2005). Figure 26 Phase contrast micrograph of S. iniae strain QMA0076, showing characteristic spherical morphology. Credit: Baiano et al. (2008).
81 2. Taxonomic Classification and Molecular Characteristics S. iniae belongs to the phylum Bacillota (formerly Firmicutes), comprising Grampositive bacteria with thick peptidoglycan cell walls. Within the genus Streptococcus, S. iniae shares molecular mechanisms with human pathogens S. pyogenes and S. pneumoniae, particularly the sortase A pathway for cell wall protein anchoring. Taxonomic hierarchy: Domain: Bacteria Phylum: Bacillota Class: Bacilli Order: Lactobacillales Family: Streptococcaceae Genus: Streptococcus Species: S. iniae Key virulence factors identified include M-like protein (SimA), hyaluronidase (Hyl), enolase (eno), and glyceraldehyde-3-phosphate dehydrogenase (GAPDH). Notably, GAPDH localizes to the outer membrane despite its primary glycolytic function, enhancing host cell adherence and immune evasion. Sortase A (SrtA) anchors these surface proteins to the cell wall through a conserved mechanism shared with other pathogenic streptococci. Despite these molecular insights, knowledge gaps remain regarding genomic variability, horizontal gene transfer, and antimicrobial resistance mechanisms. 3. Current Disease Management Challenges Disease control in aquaculture relies primarily on broad-spectrum antibiotics (Schar et al.,2020,2021), with vaccines serving as complementary prophylactic measures. However, vaccine adoption remains limited, particularly for low-value fresh-
88 Comments, Annotation, and Features. These were mapped to SIKU01 via Entry_name derived from DIAMOND subjects after one-to-one matching between isolate SF1 from UniProtKB reference proteome and isolate SIKU01 (Supplementary Data 13). Pathway and hierarchy assignments utilized KEGG Mapper/BRITE (Kanehisa, 2000), and GO terms followed the Gene Ontology framework (The Gene Ontology Consortium,2021) (Supplementary Data 13–13). Transmembrane helices and bacterial Gram-positive N-terminal signal peptides were predicted with TMHMM v2.0 (Krogh et al.,2001) and SignalP v5.0 (Almagro Armenteros et al.,2019), respectively (Supplementary Data 13 and 13). Additional family/domain calls from InterPro’s member databases (Blum et al.,2021) were considered where present. An epitope evidence layer (Supplementary Data 13–13) was merged on the variable Locus_tag, carrying evalue_epitope_prediction,pct_identity_epitope_prediction, and bit_score_epitope_prediction for matches to curated Streptococcus spp. epitopes retrieved from the Immune Epitope Database (IEDB) (Vita et al.,2019) and aligned to the S. iniae proteome with DIAMOND BlastP (Buchfink et al.,2014) (Supplementary Code 12). Epitope hits were retained at E ≤ 1×10e-3 (observed range 10e-3 to 10e-16) (Supplementary Data 13). 7. Protein-Level Biophysical Descriptors Protein-level biophysical descriptors were computed from amino acid sequences using the Peptides R package (Osorio et al.,2015) and a custom script to annotate SIKU01 proteins of the proteome (Supplementary Code 12) and the SeqinR package in R (Charif and Lobry,2007) from translated protein sequences in Supplementary Data 13. For each protein, we recorded theoretical isoelectric point (pI), molecular weight (MW), net charge at pH 6.8, 7.0, and 7.4, GRAVY-like hydrophobicity, aliphatic index, instability index, Boman binding potential, predicted membrane propensity class, and sequence length (aa). These descriptors fed Matrix-1 filters (solubility, stability, manufacturability biases) and Matrix-2 platform compatibility rules. Physicochemical descriptors
89 were appended as Charge_pH_6_8,Charge_pH_7,Charge_pH_7_4,aliphatic_Index, hydrophobicity,instability_index,binding_potential,length_AA,mw, and pI to a subset of UniProtKB TSV data (Supplementary Data 13), summarizing SIKU01 physicochemical properties of individual proteins (Supplementary Data 13). 8. PubMed Literature Search PubMed literature and PMID metadata were retrieved using a custom R script (Supplementary Code 12). Proteins with prior experimental mentions, particularly as vaccine candidates or virulence factors, received graduated positive scoring based on evidence strength (Supplementary Data 13). 9. Annotation Merging and Biophysical Analysis All previous annotation layers were merged using (Supplementary Code 12) a custom script made for that purpose while for manual curation, cross-checking UniProtKB keywords/feature texts with InterPro domain calls, KEGG/GO assignments, and primary PGAP/Prokka annotations to produce the combined initial table used in downstream QbD scoring (Supplementary Data 15). Plots and distributions were generated using a custom R script (Supplementary Code 12). 10. Pangenomics and Sequence Conservation Analysis Ninety representative S. iniae genomes obtained from NCBI Genomes (listed in Supplementary Data 14) were annotated using Prokka (Seemann,2014) to homogenize coding sequence (CDS) annotations prior to graph-based pan-genome clustering with Panaroo (Tonkin-Hill et al.,2020b), which implements an improved orthology graph algorithm originally inspired by Roary (Page et al.,2015b). Clustering was performed at high amino acid identity threshold (≥85%), with a core gene inclusion cutoff of 95%, and paralogs excluded. This graph-based ortholog clustering produced Supplementary Data 14, including a gene presence/absence (P/A) table and summary links for each gene
90 sequence and annotation used in parsing the final pangenome graph. To annotate each genome with consistent pan-genome metadata, we used a custom Python script (Supplementary Code 12), adapted from post_run_gff_output.py, a Panaroo-based utility. This script integrated Panaroo’s graph and table outputs to reconstruct isolate-specific GFF3 annotation files containing standardized pangenome attributes. The Rtab file served as quantitative input to determine gene carriage frequency across isolates and to assign each orthogroup to a pangenome class (core, soft-core, shell, or cloud) based on prevalence thresholds: ≥99% for core, 95–98% for soft-core, 15–94% for shell, and <15% for cloud. These classifications generated augmented presence/absence tables linking each Panaroo cluster ID to its class, representative gene, and genome distribution. To derive isolate-specific orthogroups, we extracted S. iniae SIKU01 clusters from Panaroo tables using a custom script (Supplementary Code 12), which filtered the pangenome Rtab file to produce Supplementary Data 14. This dataset was enriched with GFF-derived attributes using (Supplementary Code 12), another custom script which cross-referenced pangenome_id entries against the Panaroo-generated GFF file to create a unified annotation table. Although biased toward the SIKU01 reference isolate, this approach provided one-to-one mapping between pangenome clusters and genomespecific locus tags, enabling precise tracking of orthogroups during downstream conservation analysis. For each genome, CDS and translated protein sequences were extracted from annotated GFF3 files using gffread (Pertea and Pertea,2020), and all orthologous sequences were concatenated into two tagged FASTA files with SeqKit (Shen et al.,2016). These combined FASTA files were processed by a custom Python pipeline (Supplementary Code 12), which automatically retrieved sequences by cluster ID and performed per-cluster multiple sequence alignments (MSAs) with MAFFT (Katoh et al.,2002) or Clustal Omega (Sievers and Higgins,2018) depending on cluster size. The resulting nucleotideand amino acid-level MSAs formed the foundation for subsequent analyses
91 of sequence variability, conservation, and structure-based mapping. All MSAs were analyzed in R (Supplementary Code 12) using a custom script, which calculated normalized Shannon entropy (H.norm) at both codon and amino acid levels via the Bio3D R package (Grant et al.,2006). For each orthogroup, H.norm quantified sequence conservation on a continuous scale from 0 (fully conserved) to 1 (maximally variable). Antigens with median H.norm ≤0.05 were classified as highly conserved, while those above 0.90 were excluded as excessively variable. These conservation metrics, together with gene-carriage data from Panaroo, defined the gene-carriage and sequence-variability components of the Quality-by-Design (QbD) framework (PreM1 and M1 matrices) (Supplementary Data 14). 11. Protein Structure Prediction and Visualization Three-dimensional structures of prioritized S. iniae antigens (GAPDH, enolase, GroEL, Sortase A) were predicted using AlphaFold2 (Jumper et al.,2021) and crossvalidated against available homologous templates in the Protein Data Bank (PDB) (Berman, 2000). Predicted PDB files were examined in UCSF ChimeraX v1.6 (Meng et al.,2023) for fold integrity, residue geometry, and solvent accessibility. To quantify site-specific structural conservation, we implemented a per-residue identity analysis (Supplementary Code 12) that scanned each aligned FASTA previously produced by MAFFT. For every alignment, the script compared all sequences to the reference (first entry) and recorded, for each residue position, the number and percentage of identical residues across all genomes. These per-site identity profiles were mapped onto AlphaFold2 models in ChimeraX using custom color scripts to generate gradient-based visualizations where blue indicated highly conserved residues and red denoted variable positions. This approach provided a direct link between MSA entropy (H.norm) values and 3D spatial conservation, highlighting structurally stable and immunologically relevant regions.
92 Epitope localization was performed using (Supplementary Code 12), a custom Python script which scanned PDB chains for exact matches to epitope peptide sequences identified by IEDB mapping. For each hit, the script reported the model ID, chain ID, and PDB residue range corresponding to the matched peptide, enabling precise overlay of Band T-cell epitopes onto predicted antigen structures. Surface visualization and electrostatic potential mapping were carried out in ChimeraX using ”surface color” and ”cartoon” representations to assess accessibility and topological context. 12. Codon Adaptation Index (CAI) and Expression Compatibility Codon usage frequencies for target organisms were obtained from the Kazusa Codon Usage Database (Holcomb et al.,2019), available at https://www.kazusa.or.jp/codon/. Species-specific codon frequency tables were downloaded (Supplementary Data 15), providing frequencies per thousand codons for each organism’s coding sequences. The Codon Adaptation Index (CAI) quantifies the degree of preference for synonymous codons in a given organism, where values range from 0 (least preferred) to 1 (most preferred). For each amino acid family, the relative adaptiveness of codon iwas calculated as wi=fi/fmax, where firepresents the frequency per thousand of codon ifor its amino acid and fmax represents the frequency per thousand of the most frequently used codon for that amino acid. For example, proline codons in Oreochromis niloticus were calculated as follows: CCG had wi=7.39/16.53 =0.447, CCA had wi=14.59/16.53 =0.883, CCT had wi=16.53/16.53 =1.000 (optimal), and CCC had wi=14.91/16.53 =0.902. The value 16.53 represents the highest frequency among all proline codons, making CCT the optimal reference codon. For complete gene sequences, CAI was computed as the geometric mean of rel-
93 ative adaptiveness values using the formula: CAI =(L ∏ k=1 wk)1/L where the product encompasses all Lsense codons in the gene sequence. Stop codons (UAA, UAG, UGA) and the single methionine codon (AUG) were excluded from CAI calculations as they lack synonymous alternatives for optimization. 13. Quality by Design (QbD) Framework To implement QbD principles effectively within the reverse vaccinology framework, specific criteria were established to guide systematic antigen selection. These criteria served as the foundation for developing a quantitative scoring matrix computed in a custom script (Supplementary Code 12) that ensures both immunological efficacy and manufacturing feasibility. 13.1 QbD Design Space: Gene Carriage (Matrix Pre-M1) The gene carriage design space defines the level of genomic conservation of an antigen across the S. iniae pangenome. This Critical Quality Attribute (CQA) prioritizes stable and widely distributed antigens to prevent loss of vaccine efficacy in heterologous strains. Based on the pan-genome presence/absence matrix, the core genome (present in ≥99% of isolates) was retained whereas soft-core genome (95–98%), shell (15–94%), and cloud (<15%) genes were excluded from further evaluation. This filtering yielded 1,538 core and 110 soft-core genes (Supplementary Data 15), forming the foundation for downstream antigen variability and QbD scoring due to their broad distribution and genetic stability.
94 13.2 QbD Design Space: Amino Acid and Nucleotide Variability (Matrix 1) The normalized Shannon entropy (H.norm)defined the antigen variability design space, quantified from per-cluster MSAs. Each residue received two complementary measures: (1) H.norm from sequence alignments (0 = fully conserved, 1 = maximally variable) and (2) percent identity (%ID) relative to the amino acid reference sequence (SIKU01). These values were averaged per gene to yield gene-wise conservation indices subsequently integrated into the QbD variability matrix (M1) (Supplementary Data 15). Genes with median H.norm ≤0.05 and mean %ID ≥95% were prioritized as structurally and functionally stable antigens. Residues with low entropy and high surface exposure (as confirmed by ChimeraX mapping) defined the structural conservation-driven QbD design space, used to constrain antigen selection and ensure manufacturing consistency across S. iniae lineages. 13.3 QbD Design Space: General Physicochemical and Expression Characteristics (Matrix 1) The protein length design space favors candidates with 100 or more amino acids to ensure sufficient epitope diversity while avoiding overly short sequences that lack functional relevance. Proteins containing 50–99 amino acids are classified as too short and excluded due to instability concerns, while those with fewer than 50 amino acids face exclusion for insufficient immunogenic potential. Molecular weight constraints target the 20–80 kDa range to balance immune system interaction with proper folding capabilities, with the optimal 20–50 kDa range receiving highest priority for E. coli expression systems. The isoelectric point (pI)design space recognizes multiple optimal ranges depending on purification strategy. Proteins with pI values between 4.0–7.5 receive priority due to superior solubility characteristics and broad purification platform compatibility. The pI range of 7–9 provides negative charge under physiological conditions,
95 facilitating purification via histidine tag chromatography or silicon dioxide-based filters as specified in the foundational selection criteria. Additionally, proteins with pI values between 2–4 offer specialized advantages for silica-based purification systems through enhanced electrostatic interactions. Hydrophobicity design space assessment relies on the GRAVY index, which prioritizes proteins within the -0.5 to +0.5 range to minimize aggregation tendencies and support aqueous solubility throughout the manufacturing process. Protein stability design space relies on the instability index (II) calculated from primary sequence data, with values of 40 or lower indicating favorable in vivo stability characteristics. Literature support provides empirical validation through PubMed literature and PMID, with proteins having prior experimental mentions, particularly as vaccine candidates or virulence factors, receiving graduated positive scoring based on evidence strength. 13.4 QbD Design Space: Immunogenicity (Matrix 1) Epitope prediction analysis identifies the presence of known sequences that stimulate T and B lymphocytes, considering the 30% similarity between fish and human immunoglobulins in the scoring framework. Proteins predicted to contain both Tand B-cell epitopes (inferred from IEDB) receive maximum immunological relevance scoring, while those lacking predicted epitopes face penalty assessment. 13.5 QbD Design Space: Host-specific Expression (Matrix 1 – CAI) In vitro expression efficiency design space prediction utilizes the Codon Adaptation Index (CAI) to assess translational compatibility with E. coli systems, following the principle of optimal codon usage for maximizing protein expression. Top-
96 quartile CAI values receive positive weighting, while bottom-quartile scores indicate potential expression difficulties (Supplementary Data 15–15). 13.6 QbD Design Space: Matrix 2 – Secondary Selection In Matrix 2, candidate antigens from Matrix 1 were screened multiple times for compatibility with specific downstream purification modalities by aligning their intrinsic molecular descriptors with platform-specific Critical Process Parameters (CPPs). Five distinct sets of design spaces were evaluated: silica affinity, cellulose affinity, ion exchange chromatography (anion exchange (AEX) and cation exchange (CEX)), and plasmid DNA (pDNA) expression. This pre-downstream design space cross-references in silico Critical Quality Attributes (CQAs) with platform-specific CPPs to ensure process– antigen compatibility before physical development for maximum cost-efficiency postbiomanufacturing (Supplementary Data 15–15). Silica Affinity Purification: Silica-based purification technologies offer ubiquity, low cost, and adaptability across laboratory and industrial bioprocessing applications. Critical material attributes (CMAs) include silanol density, surface topology, and functionalization chemistry, while key CPPs encompass pH, ionic strength, and buffer composition (Supplementary Data 15). Affinity tags with high silica specificity, including Si-tag, SB7, Car9, R5, and the synthetic octapeptide (RH)4, composed of four repeating Arginine-Histidine units, interact through electrostatic and hydrogen-bonding mechanisms with silanol-rich matrices. Proteins presenting net positive charge at pH 7.4 demonstrate enhanced adsorption due to electrostatic attraction to partially deprotonated silanol groups (≡Si–OH → ≡Si–O-). QbD selection criteria required pI > 4 and a surface-exposed cationic patch at loading pH (offered by the fusion of a silica binding peptide short tag). Elution is achieved by increasing pH from 7.4 to 8.5, expanding silanol deprotonation and reducing hydrogen-bond donor capacity. Short tag lengths ( 7–20 aa) were favored to avoid steric hindrance and folding interference (Supplementary Data 15).
97 Cellulose Affinity Purification: Cellulose-based purification uses cellulosebinding modules (CBMs) such as CBM3, CBM9, or CelD, which recognize repeating β-(1→4)-D-glucopyranose units through hydrogen bonding with hydroxyl groups and hydrophobic stacking with planar glucan rings. Key CPPs include buffer pH (6.0–8.5), salt concentration, cellulose type, and contact time. Since binding is mediated by the CBM domain rather than antigen charge properties, no pI requirement was applied. However, CBMs add substantial polypeptide segments (30–200 aa), necessitating length constraints. Antigens were preferentially kept ≤400 aa to ensure total fusion constructs remained ≤430–630 aa, limiting metabolic burden, misfolding risk, and aggregation while preserving tag accessibility (Supplementary Data 15). Ion Exchange Chromatography (IEX): Ion exchange utilizes the net charge of a protein to separate it from other proteins. Based on the type of resins and the charge of the protein, anion exchange chromatography or cation exchange chromatography techniques may be used. Anion exchange chromatography was modeled on quaternary amine functional groups (e.g., –N+(CH3)3) binding negatively charged proteins through electrostatic interaction with deprotonated carboxyl groups. The QbD filter retained proteins with pI ≤ 6.5 and net charge ≤ -1 at pH 7.4 (Supplementary Data 15). Cation exchange chromatography employed sulfonate ligands (–SO3-) that bind protonated amines from Lys, Arg, and His side chains. Selection criteria favored proteins with pI ≥ 8.0 and net charge ≥ +1 at pH 7.4, ensuring positive charge under working conditions for effective resin binding (Supplementary Data 15). Plasmid DNA Expression: Plasmid DNA expression was treated as a distinct design space, optimizing antigens for high-yield, high-fidelity eukaryotic expression in DNA vaccine formats. Sequence-level CQAs included GC3 enrichment to improve mRNA stability and translation efficiency, ORF length threshold (set at the median coding sequence length of core genes ≈2.2 kb), and complete absence of internal Type IIS restriction sites (BsaI, BsmBI, Eco31I) to ensure seamless Golden Gate Assembly cloning compatibility (Supplementary Data 15–15).
104 •M1 (protein route). Host-specific CAI in E. coli, CAIec (+2 top quartile / −2 bottom quartile / 0 otherwise); hydrophobicity (GRAVY −0.5 to +0.5) +2; instability index (II) ≤ 40 +1 (Guruprasad et al.,1990). •M1 (pDNA routes). Host-specific CAI in zebrafish (CAIdr) or tilapia (CAIon) scored identically (+2 top quartile / −2 bottom quartile / 0 otherwise). •M2 (downstream manufacturability, platform specific). In the final stage, manufacturability descriptors were applied. For silica affinity, net charge (z) at pH 7 was used: z ≤ −15 led to exclusion, −15 < z ≤ −5 scored −1, −5 < z ≤ +5 scored +2, and z > +5 scored +3 (capped at +3). A soft preference of +1 was also added for proteins with isoelectric points (pI) between 7 and 9. Silica-binding peptides were constrained to 20 aa to minimize metabolic cost. For cellulose affinity, candidate antigens were required to be <400 aa to offset the 30-200 aa carbohydrate-binding module fusion; longer proteins were excluded. For ionexchange chromatography, anion exchange (AEX) required pI ≤ 6.0; within this constraint, proteins with z at pH 7 were scored as 0 to −5 (+1), −5 to −15 (+2), and ≤ −15 (+3). Cation exchange (CEX) required pI ≥ 8.0; within this constraint, proteins with z at pH 7 were scored as 0 to +5 (+1), +5 to +15 (+2), and ≥ +15 (+3). For plasmid DNA (pDNA) manufacturability, GC3 fraction [0–1] in the top quartile scored +2, the bottom quartile scored −2, and the middle quartiles scored 0. Gene length (nt) shorter than the cohort median scored +1, while longer genes scored 0. Absence of internal Type IIS sites (count = 0) scored +2, while presence scored 0. Figure 28 summarizes the Quality-by-Design (QbD) lifecycle applied to the S. iniae proteome, illustrating how successive filtering stages (M0–M2) progressively refine the antigen candidate space. The workflow integrates antigen design and manufacturability routes through iterative filtering of Critical Quality Attributes (CQAs) into Critical Process Parameters (CPPs) across multiple production systems. This staged CQA→CPP mapping and definition of design spaces ensure that retained candidates are both immunologically relevant and manufacturable at scale, accommodating mul-
105 tiple expression systems (plasmid DNA in vivo and recombinant protein vaccines in E. coli) as well as diverse downstream purification routes (silica and cellulose affinity, ion-exchange chromatography, and pDNA expression in zebrafish and Nile tilapia). Full thresholds and scoring rules are provided in (Table 10). Complete proteome-level distributions of CQA attributes and their pairwise correlations (Spearman’s ρ) are shown in (Appendix Figures 36–38). Figure 28 Quality-by-Design (QbD) workflow integrating antigen design and manufacturability routes. (a) The conceptual QbD lifecycle links the Target Product Profile (TPP), in silico antigen scoring, and affinity-tag or plasmid purification strategy through iterative design feedback. (b) The M2 manufacturability stage defines platform-specific routes—silica and cellulose affinity, ion-exchange chromatography (AEX/CEX), and plasmid DNA optimization—each translating Critical Quality Attributes (CQAs) into measurable Critical Process Parameters (CPPs) for scalable vaccine production. (c) The workflow illustrates the progressive narrowing of the vaccineantigen design space through successive QbD stages (M0 → Pre-M1 → M1 → M2).
106 3.3 Initial QbD Proteome Reduction (M0 → Pre-M1 → M1-general) Results At the unfiltered M0 stage (Supplementary Data 15), the proteome (N = 1,855 proteins) spanned a wide biophysical space, with hydrophobicity and isoelectric point distributions showing clear bimodality (Appendix Figure 39a). Applying the Pre-M1 inclusion gate (core genome carriage) reduced the set to 1,374 proteins (Appendix Figure 39b). (Supplementary Data 15). M1-general filters then retained 1,101 proteins by enforcing protein length 100–699 aa (else excluded), conservation (H.norm > 0.98, else excluded), molecular weight (MW) 20–60 kDa (+2 points), literature support (+1 point), and IEDB epitope presence (+3 points) (Appendix Figure 39c). (Supplementary Data 15). 3.4 Platform-specific Gates (M1) Results Filters were adapted to expression platforms in a host-dependent manner (Appendix Figure 40a). For the M1-protein route in E. coli (141 proteins), feasible antigens required CAIec ≥ Q3, moderate hydrophobicity (GRAVY −0.5 to +0.5), and stability (instability index (II) ≤ 40) (Supplementary Fig. S08a), detailed in (Supplementary Data 15). For the M1-pDNA routes, the only gate was host-specific CAI ≥ Q3, yielding tilapia (207 proteins) and zebrafish (212 proteins) (Appendix Figure 40b-c), detailed in (Supplementary Data 15-15). GC3 and nucleotide length constraints were not applied until M2. In both hosts, density contours showed clustering of survivors within narrow codon usage and sequence length ranges, reflecting platform-specific optimization for translation efficiency and manufacturability. 3.5 Downstream Manufacturability Design Spaces (M2) Results To evaluate manufacturability, each downstream platform was subjected to stepwise Quality by Design (QbD) filtering (M2 criteria) (Figure 3.5). Mapping antigens into ion-exchange charge space (Figure 3.5) showed that most proteins were acidic at pH 7 and thus compatible with anion exchange (AEX), whereas only a minor frac-
107 tion were sufficiently basic to survive cation exchange (CEX). This was reflected in the survivor pools (Fig. 3.5), with 98 proteins retained by AEX and only 20 by CEX. Bufferspecific ranges (Figs. 3.5–3.5) confirmed that AEX survivors were stable across nearly all chemistries (n = 82–98 proteins), while CEX buffers consistently yielded the same restricted set (n = 20 proteins). The full antigen lists for AEX and CEX are provided in (Supplementary Data 15-15). Affinity purification routes imposed distinct orthogonal constraints: cellulose excluded proteins longer than 400 aa (Fig. 3.5), because the CBM fusion tag is itself large and places a metabolic burden on recombinant expression; minimizing the size of the fused antigen reduces energy demand and steric hindrance, thereby improving folding and yield. In contrast, silica affinity (Fig. 3.5) was governed by electrostatic interactions with surface silanol groups: proteins with pI between 7–9 acquire a net positive charge at physiological pH, enabling stable adsorption to negatively charged silica. These filters reduced the pools to 100 cellulose-compatible and 49 silica-compatible proteins, detailed in (Supplementary Data 15-15). For plasmid DNA (pDNA) platforms, Critical Quality Attribute (CQA) filters were applied separately on a per-host basis. In O. niloticus (Figs. 3.5–3.5), genes and their product were first assessed for CAI ≥ Q3 ( 0.60), ensuring codon usage was well adapted to the host translation machinery; this retained 188 sequences. Applying GC3 ≥ Q3 ( 0.32) halved the pool to 95, as high G/C content at the third codon position improves mRNA stability and reduces metabolic stress. A further antigen length cutoff (≤ 2,200 nt) yielded 57 candidates, favoring shorter constructs with reduced transcriptional burden. Finally, genes containing internal Type IIS restriction sites (e.g., BsaI, SapI) were excluded because these enzymes cut outside their recognition sites, disrupting modular DNA assembly workflows by preventing a one-step cloning in a Golden Gate Assembly. The final Nile tilapia pDNA pool remained at 57 survivors. In D. rerio (Figs. 3.5–3.5), the thresholds were slightly higher (CAI ≥ Q3 0.70, GC3 ≥ Q3 0.31). Here, 186 proteins passed the CAI filter, 82 remained after the
108 GC3 filter, 66 after the length filter, and 65 after the Type IIS removal filter. Comprehensive candidate sets for both pDNA platforms (Nile tilapia and zebrafish) are available in (Supplementary Data 15-15). Overall, the reverse vaccinology–QbD funnel outputs platform-ready shortlists: 98 AEX, 20 CEX, 100 cellulose, 49 silica, and 57–65 pDNA antigens for immediate downstream development.
109
110 Figure 28 Manufacturability design spaces across purification platforms. (a) Ion-exchange charge space for protein candidates: pI vs. net charge (z) at pH 7. Shaded bands mark the AEX and CEX inclusion regions; vaccine-referenced antigens are labeled. (b) Survivors after M2 filtering by platform; bars show unique genes per route with counts and percentages. Only candidates fitting the biophysiochemical criteria for their respective purification pathway are displayed in color bars. This step reflects manufacturability constraints and downstream process optimization in the QbD framework. (c–d) Buffer operating pH ranges for AEX (c) and CEX (d). Short tick marks indicate the working pH (range midpoint). (e–f) Protein-route survivors per buffer evaluated at the midpoint: AEX rule = base gate passed, net charge (z) at pH 7 ≤0, and pI ≤pHmid; CEX rule = base gate passed, net charge (z) at pH 7 ≥0, and pI ≥pHmid.(g) Cellulose: antigen length distribution for M2 survivors; dashed line at 400 aa. (j) Silica binding peptides favor proteins with moderate pI (7–9) and positive net charge at pH > 7. (h, k) pDNA (M2) scatterplots by host: tilapia (h) and zebrafish (k), showing CAI vs. GC3; dashed lines mark dataset thresholds. (i, l) pDNA-route survivors passing each manufacturability gate for tilapia (i) and zebrafish (l): CAI ≥threshold, GC3 ≥threshold, length ≤median, and no Type IIS sites (count).
111 4. Cross-Validation with Literature We validated our shortlisted antigens against previously tested S. iniae vaccine targets reported in multiple hosts, including Channel catfish (Ictalurus punctatus) (Wang et al.,2016b), Nile tilapia (Oreochromis niloticus) (Kayansamruaj et al.,2017), Olive flounder (Paralichthys olivaceus) (Sheng et al.,2018a,2023), Zebrafish (Danio rerio) (Membrebe et al.,2016), Mouse (Mus musculus) (Wang et al.,2015a), and Turbot fish (Scophthalmus maximus) (Zhang et al.,2014a) (Table 11). Among the top-ranked candidates consistently found across purification routes, antigens with demonstrated in vivo protection (enolase, GAPDH, and GroEL) were recovered, supporting the predictive accuracy of the QbD framework. To further assess their suitability, we modeled GAPDH and enolase, two representative antigens, both well-established in the literature. Structural predictions generated using AlphaFold2 (Yang et al.,2023) for S. iniae SIKU01 confirmed highly conserved folds and surface-exposed loops (Wang et al.,2017). Several epitope-containing regions identified through reverse vaccinology were surface-accessible ((Figures 29a– c), highlighted in yellow), consistent with prior reports (Gent et al.,2024). These epitopes (Table 9) are suitable for direct incorporation into subunit vaccines or for the design of chimeric multi-epitope constructs (Pumchan et al.,2020), further validating their suitability as vaccine targets. Both enolase (eno) and GAPDH (gap) displayed epitoperich regions spatially separated from catalytic or hypervariable sites, reinforcing their accessibility and stability as broad-spectrum candidates.
112 Figure 29 Structural mapping of variability, epitopes, and active sites in two model Streptococcus iniae vaccine candidates. (a) Shannon entropy profiles show that two model and well-known antigens of the scientific literature, GAPDH (gap) and enolase (eno), are largely conserved with limited variable regions. (b) GAPDH (336 aa, UniProt Q7BB80) carries a predicted N-terminal epitope (MVVKVGINGFGRIGRLAFRRIQ) positioned near, but not overlapping with, the active site and adjacent to a region of high nucleotide-level (codon-derived) variability (1–89 % variation). (c) Enolase (435 aa, UniProt T1TFA0) contains a predicted epitope (RAAADYLEVPLYNYLG) located opposite the active site and spatially separated from five nucleotide-driven hypervariable surface regions (HR1–HR5), indicating stability and accessibility as a vaccine target. Variability values are based on codon-level polymorphisms within the core-genome alignment and projected onto corresponding amino acid positions in the 3-D models for visualization.
113 DATA AVAILABILITY AND NCBI SUBMISSIONS Bighead Catfish The final diploid genome assembly was screened for contamination using the NCBI Foreign Contamination Screen (FCS) (Astashyn et al.,2024) and submitted following the Vertebrate Genome Project (VGP) naming conventions (https://github.com/VGP/ vgp-assembly) (Rhie et al.,2021). The project is registered under NCBI BioProject PRJNA1132508, with BioSample accession SAMN41769988 corresponding to a Thai (male, adult) bighead catfish (isolate: CMAM; TaxID: 35657). Raw sequencing reads are deposited in the NCBI Sequence Read Archive (SRA): Nanopore (20% error) — SRR29723575, HiFi — SRR29723576, Hi-C (150PE) — SRR29723577, and Illumina (150PE) — SRR29723578. The final diploid assembly is available in GenBank under accession numbers JBLWMO000000000 (Haplotype 1) and JBLWMP000000000 (Haplotype 2). A complete dataset, including genome assemblies and supporting files, is permanently archived at Zenodo (10.5281/zenodo.14826875). F1 Hybrid Catfish The hybrid genome assembly (Clarias macrocephalus ×C. gariepinus) was processed and submitted using the same quality control and standardization workflow as described above. GenBank accessions are JBLWFY000000000.1 and JBLWFZ000000000.1, corresponding to the C. macrocephalus and C. gariepinus sub-genomes, respectively. Associated records are hosted under NCBI BioProject PRJNA1153495 and BioSample SAMN43395848 (TaxID: 1334085). Raw sequencing reads are Nanopore (SRR30599638), HiFi (SRR30599641), Illumina (SRR30599640), and Hi-C (SRR30599639). Additional Illumina datasets correspond to female C. macrocephalus (SAMN42503781) and male C. gariepinus (SAMN43548335). The complete F1 hybrid dataset, including assemblies and metadata, is archived at Zenodo (10.5281/zenodo.15269601).
120 A. A. Brown, D. H. Buermann, A. A. Bundu, J. C. Burrows, N. P. Carter, N. Castillo, M. Chiara E. Catenazzi, S. Chang, R. Neil Cooley, N. R. Crake, O. O. Dada, K. D. Diakoumakos, B. Dominguez-Fernandez, D. J. Earnshaw, U. C. Egbujor, D. W. Elmore, S. S. Etchin, M. R. Ewan, M. Fedurco, L. J. Fraser, K. V. Fuentes Fajardo, W. Scott Furey, D. George, K. J. Gietzen, C. P. Goddard, G. S. Golda, P. A. Granieri, D. E. Green, D. L. Gustafson, N. F. Hansen, K. Harnish, C. D. Haudenschild, N. I. Heyer, M. M. Hims, J. T. Ho, A. M. Horgan, K. Hoschler, S. Hurwitz, D. V. Ivanov, M. Q. Johnson, T. James, T. A. Huw Jones, G.-D. Kang, T. H. Kerelska, A. D. Kersey, I. Khrebtukova, A. P. Kindwall, Z. Kingsbury, P. I. Kokko-Gonzales, A. Kumar, M. A. Laurent, C. T. Lawley, S. E. Lee, X. Lee, A. K. Liao, J. A. Loch, M. Lok, S. Luo, R. M. Mammen, J. W. Martin, P. G. McCauley, P. McNitt, P. Mehta, K. W. Moon, J. W. Mullens, T. Newington, Z. Ning, B. Ling Ng, S. M. Novo, M. J. O’Neill, M. A. Osborne, A. Osnowski, O. Ostadan, L. L. Paraschos, L. Pickering, A. C. Pike, A. C. Pike, D. Chris Pinkard, D. P. Pliskin, J. Podhasky, V. J. Quijano, C. Raczy, V. H. Rae, S. R. Rawlings, A. Chiva Rodriguez, P. M. Roe, J. Rogers, M. C. Rogert Bacigalupo, N. Romanov, A. Romieu, R. K. Roth, N. J. Rourke, S. T. Ruediger, E. Rusman, R. M. Sanches-Kuiper, M. R. Schenker, J. M. Seoane, R. J. Shaw, M. K. Shiver, S. W. Short, N. L. Sizto, J. P. Sluis, M. A. Smith, J. Ernest Sohna Sohna, E. J. Spence, K. Stevens, N. Sutton, L. Szajkowski, C. L. Tregidgo, G. Turcatti, S. vandeVondele, Y. Verhovsky, S. M. Virk, S. Wakelin, G. C. Walcott, J. Wang, G. J. Worsley, J. Yan, L. Yau, M. Zuerlein, J. Rogers, J. C. Mullikin, M. E. Hurles, N. J. McCooke, J. S. West, F. L. Oaks, P. L. Lundberg, D. Klenerman, R. Durbin and A. J. Smith. 2008. Accurate whole human genome sequencing using reversible terminator chemistry. Nature. 456 (7218): 53–59. Berman, H. M. 2000. The Protein Data Bank. Nucleic Acids Research. 28 (1): 235–242. Bird, J. E., J. Marles-Wright and A. Giachino. 2022a. A user’s guide to golden gate cloning methods and standards. ACS Synthetic Biology. 11 (12): 3551–3563. Bird, J. E., J. Marles-Wright and A. Giachino. 2022b. A User’s Guide to Golden Gate Cloning Methods and Standards. ACS Synthetic Biology. 11 (11): 3551–3563. Blum, M. and others. 2021. The InterPro protein families and domains database: 20 years on. Nucleic Acids Research. 49 (D1): D344–D354. Breves, J. P., S. B. Serizier, V. Goffin, S. D. McCormick and R. O. Karlstrom. 2013. Prolactin regulates transcription of the ion uptake Na+/Cl− cotransporter (ncc) gene in zebrafish gill. Molecular and Cellular Endocrinology. 369 (1–2): 98–106.
121 Brown, M., P. M. González De laRosa and B. Mark. A Telomere Identification Toolkit, 2023. URL https://doi.org/10.5281/zenodo.10091385. Buchanan, J. T., J. A. Stannard, X. Lauth and others. 2005. Streptococcus iniae phosphoglucomutase is a virulence factor and a target for vaccine development. Infection and Immunity. 73 (10): 6935–6944. Buchfink, B., K. Reuter and H. Drost. 2021. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nature Methods. 18 (4): 366–368. Buchfink, B., C. Xie and D. H. Huson. 2014. Fast and sensitive protein alignment using DIAMOND. Nature Methods. 12 (1): 59–60. Cao, Y., J. Liu, G. Liu, H. Du, T. Liu, G. Wang, Q. Wang, Y. Zhou and E. Wang. 2023. Exploring the Immunoprotective Potential of a Nanocarrier Immersion Vaccine Encoding Sip against Streptococcus Infection in Tilapia (Oreochromis niloticus). Vaccines. 11 (7): 1262. Carrard, G., A. Koivula, H. Söderlund and P. Béguin. 2000. Cellulose-binding domains promote hydrolysis of different sites on crystalline cellulose. Proceedings of the National Academy of Sciences of the United States of America. 97 (19): 10342–10347. Chaivichoo, P., S. Koonawootrittriron, S. Chatchaiphan, W. Srimai and U. Na-Nakorn. 2020. Genetic components of growth traits of the hybrid between ♂ North African catfish (Clarias gariepinus Burchell, 1822) and ♀ bighead catfish (C. macrocephalus Gunther, 1864). Aquaculture. 521: 735082. Chaivichoo, P., S. Sukhavachana, R. Khumthong, P. Srisapoome, S. Chatchaiphan and U. NaNakorn. 2023. Genome–wide association study and genomic prediction of growth traits in bighead catfish (Clarias macrocephalus Günther, 1864). Aquaculture. 562: 738748. Charif, D. and J. R. Lobry. SeqinR 1.0-2: A Contributed Package to the R Project for Statistical Computing Devoted to Biological Sequences Retrieval and Analysis. In Structural Approaches to Sequence Evolution, pp. 207–232. Springer Berlin Heidelberg, 2007. doi: 10.1007/978-3-540-35306-5_10. URL https://doi.org/10.1007/978-3-540-35306-5_ 10. Chen, M. and others. 2012. PCR detection and PFGE genotype analyses of streptococcal clinical isolates from tilapia in China. Veterinary Microbiology. 159 (3-4): 526–530. Chen, S., Y. Zhou, Y. Chen and J. Gu. 2018. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 34 (17): i884–i890. Chen, W., M. Zou, Y. Li, S. Zhu, X. Li and J. Li. 2021a. Sequencing an F1 hybrid of Silurus
122 asotus and S. meridionalis enabled the assembly of high-quality parental genomes. Scientific Reports. 11 (1). Chen, Y., Y. Zhang, A. Y. Wang, M. Gao and Z. Chong. 2021b. Accurate long-read de novo assembly evaluation with Inspector. Genome Biology. 22 (1). Cheng, H., G. T. Concepcion, X. Feng, H. Zhang and H. Li. 2021. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nature Methods. 18 (2): 170– 175. Cheng, H., E. D. Jarvis, O. Fedrigo, K.-P. Koepfli, L. Urban, N. J. Gemmell and H. Li. 2022. Haplotype-resolved assembly of diploid genomes without parental data. Nature Biotechnology. 40 (9): 1332–1335. Dale, J., P. Smeesters, H. Courtney, T. Penfound, C. Hohn, J. Smith and J. Baudry. 2017. Structure-based design of broadly protective group A streptococcal M protein-based vaccines. Vaccine. 35 (1): 19–26. Danecek, P., J. K. Bonfield, J. Liddle, J. Marshall, V. Ohan, M. O. Pollard, A. Whitwham, T. Keane, S. A. McCarthy, R. M. Davies and H. Li. 2021. Twelve years of SAMtools and BCFtools. GigaScience. 10 (2). Darling, A. E., B. Mau and N. T. Perna. 2010. progressiveMauve: Multiple Genome Alignment with Gene Gain, Loss and Rearrangement. PLoS ONE. 5 (6): e11147. De Coster, W. and R. Rademakers. 2023. NanoPack2: population-scale evaluation of long-read sequencing data. Bioinformatics. 39 (5). De Coster, W., S. D’Hert, D. T. Schultz, M. Cruts and C. Van Broeckhoven. 2018. NanoPack: visualizing and processing long-read sequencing data. Bioinformatics. 34 (15): 2666– 2669. Deane, E. E. and N. Y. S. Woo. 2008. Modulation of fish growth hormone levels by salinity, temperature, pollutants and aquaculture related stress: a review. Reviews in Fish Biology and Fisheries. 19 (1): 97–120. Defoirdt, T., P. Sorgeloos and P. Bossier. 2011. Alternatives to antibiotics for the control of bacterial disease in aquaculture. Current Opinion in Microbiology. 14 (3): 251–258. Dobrut, A., E. Brzozowska, S. Gorska, M. Pyclik, A. Gamian, M. Bulanda, E. Majewska and M. Brzychczy-Wloch. 2018. Epitopes of Immunoreactive Proteins of Streptococcus Agalactiae: Enolase, Inosine 5’-Monophosphate Dehydrogenase and Molecular Chaperone GroEL. Frontiers in Cellular and Infection Microbiology. 8: 349. Dudchenko, O., S. S. Batra, A. D. Omer, S. K. Nyquist, M. Hoeger, N. C. Durand, M. S. Shamim,
123 I. Machol, E. S. Lander, A. P. Aiden and E. L. Aiden. 2017. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science. 356 (6333): 92–95. Duong, T.-Y. and K. T. Scribner. 2018. Regional variation in genetic diversity between wild and cultured populations of bighead catfish (Clarias macrocephalus) in the Mekong Delta. Fisheries Research. 207: 118–125. Duong-Ly, K. C. and S. B. Gabelli. 2014. Using ion exchange chromatography to purify a recombinantly expressed protein. Methods in Enzymology. 541: 95–103. Durand, N. C., M. S. Shamim, I. Machol, S. S. Rao, M. H. Huntley, E. S. Lander and E. L. Aiden. 2016. Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. Cell Systems. 3 (1): 95–98. Eisenstein, M. 2017. An ace in the hole for DNA sequencing. Nature. 550 (7675): 285–288. El Hilali, S. and R. R. Copley. 2023. macrosyntR: Drawing automatically ordered Oxford Grids from standard genomic files in R. Archive Ouverte HAL. Eldar, A., A. Horovitcz and H. Bercovier. 1997. Development and efficacy of a vaccine against Streptococcus iniae infection in farmed rainbow trout. Veterinary Immunology and Immunopathology. 56 (1-2): 175–183. Elfaitouri, A., B. Herrmann, A. Bolin-Wiener, Y. Wang, C. Gottfries, O. Zachrisson, R. Pipkorn, L. Ronnblom and J. Blomberg. 2013. Epitopes of microbial and human heat shock protein 60 and their recognition in myalgic encephalomyelitis. PLoS ONE. 8 (11): e81155. Ellinghaus, D., S. Kurtz and U. Willhoeft. 2008. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics. 9 (1). Emms, D. M. and S. Kelly. 2019. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biology. 20 (1). Ewels, P., M. Magnusson, S. Lundin and M. Käller. 2016. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics. 32 (19): 3047. Eyngor, M. and others. 2008. Emergence of novel Streptococcus iniae exopolysaccharideproducing strains following vaccination with nonproducing strains. Applied and Environmental Microbiology. 74 (22): 6892–6897. Facklam, R., J. Elliott, L. Shewmaker and A. Reingold. 2005. Identification and characterization of sporadic isolates of Streptococcus iniae isolated from humans. Journal of Clinical Microbiology. 43 (2): 933–937.
124 FAO, . 2020. The State of World Fisheries and Aquaculture 2020. FAO. Faust, G. G. and I. M. Hall. 2014. SAMBLASTER: fast duplicate marking and structural variant read extraction. Bioinformatics. 30 (17): 2503–2505. Ferraris, C. J. 2007. Checklist of catfishes, recent and fossil (Osteichthyes: Siluriformes), and catalogue of siluriform primary types. Zootaxa. 1418 (1): 1–628. Finn, R. D., J. Clements and S. R. Eddy. 2011. HMMER web server: interactive sequence similarity searching. Nucleic Acids Research. 39 (suppl): W29–W37. Flynn, J. M., R. Hubley, C. Goubert, J. Rosen, A. G. Clark, C. Feschotte and A. F. Smit. 2020. RepeatModeler2 for automated genomic discovery of transposable element families. Proceedings of the National Academy of Sciences. 117 (17): 9451–9457. Formenti, G., A. Rhie, B. P. Walenz, F. Thibaud-Nissen, K. Shafin, S. Koren, E. W. Myers, E. D. Jarvis and A. M. Phillippy. 2022. Merfin: improved variant filtering, assembly evaluation and polishing via k-mer validation. Nature Methods. 19 (6): 696–704. Fourie, K. and H. Wilson. 2020. Understanding GroEL and DnaK Stress Response Proteins as Antigens for Bacterial Diseases. Vaccines. 8 (4): 773. Francis, D. M. and R. Page. 2010. Strategies to optimize protein expression in E. coli. Current Protocols in Protein Science. (1): 5.24.1–5.24.29. Freitas, A. I., L. Domingues and T. Q. Aguiar. 2022. Bare silica as an alternative matrix for affinity purification/immobilization of his-tagged proteins. Separation and Purification Technology. 286: 120448. Gent, V., Y.-J. Lu, S. Lukhele and others. 2024. Surface protein distribution in Group B Streptococcus isolates from South Africa and identifying vaccine targets through in silico analysis. Scientific Reports. 14: 22665. Giovanni, A., Y.-Z. Shi, P.-C. Wang, M.-A. Tsai and S.-C. Chen. 2025. Recombinant C5a Peptidase and Formalin-Killed Cell: A Synergistic Vaccine Against Streptococcus iniae in Four-Finger Threadfin Fish (Eleutheronema tetradactylum). Journal of Fish Diseases. 48: e14154. Glazunova, O. O., D. Raoult and V. Roux. 2009. Partial sequence comparison of the rpoB, sodA, groEL and gyrB genes within the genus Streptococcus. INTERNATIONAL JOURNAL OF SYSTEMATIC AND EVOLUTIONARY MICROBIOLOGY. 59 (9): 2317–2322. Goel, M. and K. Schneeberger. 2022. plotsr: visualizing structural similarities and rearrangements between multiple genomes. Bioinformatics. 38 (10): 2922–2926.
125 Gong, H. and others. 2017. Complete Genome Sequence of Streptococcus iniae 89353, a Virulent Strain Isolated from Diseased Tilapia in Taiwan. Genome Announcements. 5 (8): e01524–16. Goubert, C., R. J. Craig, A. F. Bilat, V. Peona, A. A. Vogan and A. V. Protasio. 2022. A beginner’s guide to manual curation of transposable elements. Mobile DNA. 13 (1). Grant, B. J., L. Skjaerven and X. Q. Yao. 2021. The Bio3D packages for structural bioinformatics. Protein Science. 30 (1): 20–30. Grant, B. J. and others. 2006. Bio3D: an R package for the comparative analysis of protein structures. Bioinformatics. 22 (21): 2695–2696. Gu, X. H., D. L. Jiang, Y. Huang, B. J. Li, C. H. Chen, H. R. Lin and J. H. Xia. 2018. Identifying a Major QTL Associated with Salinity Tolerance in Nile Tilapia Using QTL-Seq. Marine Biotechnology. 20 (1): 98–107. Guan, D., S. A. McCarthy, J. Wood, K. Howe, Y. Wang and R. Durbin. 2020. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 36 (9): 2896–2898. Guarracino, A., S. Heumos, S. Nahnsen, P. Prins and E. Garrison. 2022. ODGI: understanding pangenome graphs. Bioinformatics. 38 (13): 3319–3326. Guruprasad, K., B. V. Reddy and M. W. Pandit. 1990. Correlation between stability of a protein and its dipeptide composition: a novel approach for predicting in vivo stability of a protein from its primary sequence. Protein Engineering. 4 (2): 155–161. Harrison, P. W., M. R. Amode, O. Austine-Orimoloye, A. G. Azov, M. Barba, I. Barnes, A. Becker, R. Bennett, A. Berry, J. Bhai, S. K. Bhurji, S. Boddu, P. R. Branco Lins, L. Brooks, S. B. Ramaraju, L. I. Campbell, M. C. Martinez, M. Charkhchi, K. Chougule, A. Cockburn, C. Davidson, N. H. De Silva, K. Dodiya, S. Donaldson, B. El Houdaigui, T. E. Naboulsi, R. Fatima, C. G. Giron, T. Genez, D. Grigoriadis, G. S. Ghattaoraya, J. G. Martinez, T. A. Gurbich, M. Hardy, Z. Hollis, T. Hourlier, T. Hunt, M. Kay, V. Kaykala, T. Le, D. Lemos, D. Lodha, D. Marques-Coelho, G. Maslen, G. A. Merino, L. P. Mirabueno, A. Mushtaq, S. N. Hossain, D. N. Ogeh, M. P. Sakthivel, A. Parker, M. Perry, I. Piližota, D. Poppleton, I. Prosovetskaia, S. Raj, J. G. Pérez-Silva, A. I. A. Salam, S. Saraf, N. Saraiva-Agostinho, D. Sheppard, S. Sinha, B. Sipos, V. Sitnik, W. Stark, E. Steed, M.-M. Suner, L. Surapaneni, K. Sutinen, F. F. Tricomi, D. UrbinaGómez, A. Veidenberg, T. A. Walsh, D. Ware, E. Wass, N. L. Willhoft, J. Allen, J. Alvarez-Jarreta, M. Chakiachvili, B. Flint, S. Giorgetti, L. Haggerty, G. R. Ilsley,
126 J. Keatley, J. E. Loveland, B. Moore, J. M. Mudge, G. Naamati, J. Tate, S. J. Trevanion, A. Winterbottom, A. Frankish, S. E. Hunt, F. Cunningham, S. Dyer, R. D. Finn, F. J. Martin and A. D. Yates. 2023. Ensembl 2024. Nucleic Acids Research. 52 (D1): D891–D899. Heckman, T. I., K. Shahin, E. E. Henderson, M. J. Griffin and E. Soto. 2022. Development and efficacy of Streptococcus iniae live-attenuated vaccines in Nile tilapia (Oreochromis niloticus). Fish & Shellfish Immunology. 121: 152–162. Hoelzer, K., L. Bielke, D. P. Blake, E. Cox, S. M. Cutting, B. Devriendt, E. Erlacher-Vindel, E. Goossens, K. Karaca, S. Lemiere, M. Metzner, M. Raicek, M. C. Suriñach, N. M. Wong, C. Gay and F. V. Immerseel. 2018. Vaccines as alternatives to antibiotics for food producing animals. Part 1: challenges and needs. Veterinary Research. 49 (1). Holcomb, D. D., A. Alexaki, U. Katneni and C. Kimchi-Sarfaty. 2019. The Kazusa codon usage database, CoCoPUTs, and the value of up-to-date codon usage statistics. Infection, Genetics and Evolution. 73: 266–268. Hon, T., K. Mars, G. Young, Y.-C. Tsai, J. W. Karalius, J. M. Landolin, N. Maurer, D. Kudrna, M. A. Hardigan, C. C. Steiner, S. J. Knapp, D. Ware, B. Shapiro, P. Peluso and D. R. Rank. 2020. Highly accurate long-read HiFi sequencing data for five complex genomes. Scientific Data. 7 (1). Hu, J., Z. Wang, F. Liang, S.-L. Liu, K. Ye and D.-P. Wang. 2024. NextPolish2: A Repeataware Polishing Tool for Genomes Assembled Using HiFi Long Reads. Genomics, Proteomics and Bioinformatics. 22 (1). Jain, C., S. Koren, A. Dilthey, A. M. Phillippy and S. Aluru. 2018a. A fast adaptive algorithm for computing whole-genome homology maps. Bioinformatics. 34 (17): i748–i756. Jain, C., A. Rhie, H. Zhang, C. Chu, B. P. Walenz, S. Koren and A. M. Phillippy. 2020. Weighted minimizer sampling improves long read mapping. Bioinformatics. 36 (Supplement1): 111–118. Jain, C., A. Rhie, N. F. Hansen, S. Koren and A. M. Phillippy. 2022. Long-read mapping to repetitive reference sequences using Winnowmap2. Nature Methods. 19 (6): 705–710. Jain, M., S. Koren, K. H. Miga, J. Quick, A. C. Rand, T. A. Sasani, J. R. Tyson, A. D. Beggs, A. T. Dilthey, I. T. Fiddes, S. Malla, H. Marriott, T. Nieto, J. O’Grady, H. E. Olsen, B. S. Pedersen, A. Rhie, H. Richardson, A. R. Quinlan, T. P. Snutch, L. Tee, B. Paten, A. M. Phillippy, J. T. Simpson, N. J. Loman and M. Loose. 2018b. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nature Biotechnology. 36 (4):
127 338–345. Jiang, D. L., X. H. Gu, B. J. Li, Z. X. Zhu, H. Qin, Z. n. Meng, H. R. Lin and J. H. Xia. 2019. Identifying a Long QTL Cluster Across chrLG18 Associated with Salt Tolerance in Tilapia Using GWAS and QTL-seq. Marine Biotechnology. 21 (2): 250–261. Jumper, J. and others. 2021. Highly accurate protein structure prediction with AlphaFold. Nature. 596 (7873): 583–589. Kanehisa, M. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research. 28 (1): 27–30. Kanehisa, M. and Y. Sato. 2020. KEGG Mapper for inferring cellular functions from protein sequences. Protein Science. 29 (1): 28–35. Katoh, K. and D. M. Standley. 2013. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Molecular Biology and Evolution. 30 (4): 772–780. Katoh, K., K. Misawa, K. Kuma and T. Miyata. 2002. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Research. 30 (14): 3059–3066. Kayansamruaj, P., H. T. Dong, N. Pirarat, D. Nilubol and C. Rodkhum. 2017. Efficacy of α -enolase-based DNA vaccine against pathogenic Streptococcus iniae in Nile tilapia (Oreochromis niloticus). Aquaculture. 468: 102–106. Kayansamruaj, P., N. Areechon and S. Unajak. 2020. Development of fish vaccine in Southeast Asia: A challenge for the sustainability of SE Asia aquaculture. Fish and Shellfish Immunology. 103: 73–87. Kitano, J., S. Mori and C. L. Peichel. 2007. Sexual Dimorphism in the External Morphology of the Threespine Stickleback (Gasterosteus Aculeatus). Copeia. 2007 (2): 336–349. Kolberg, J., A. Aase, S. Bergmann, T. Herstad, G. Rodal, R. Frank, M. Rohde and S. Hammerschmidt. 2006. Streptococcus pneumoniae enolase is important for plasminogen binding despite low abundance of enolase protein on the bacterial cell surface. Microbiology (Reading, England). 152 (Pt 5): 1307–1317. Kolmogorov, M., D. M. Bickhart, B. Behsaz, A. Gurevich, M. Rayko, S. B. Shin, K. Kuhn, J. Yuan, E. Polevikov, T. P. L. Smith and P. A. Pevzner. 2020. metaFlye: scalable longread metagenome assembly using repeat graphs. Nature Methods. 17 (11): 1103–1110. Krogh, A., B. Larsson, G. vonHeijne and E. L. Sonnhammer. 2001. Predicting transmembrane protein topology with a hidden markov model: application to complete
128 genomes11Edited by F. Cohen. Journal of Molecular Biology. 305 (3): 567–580. Kusakabe, M., A. Ishikawa, M. Ravinet, K. Yoshida, T. Makino, A. Toyoda, A. Fujiyama and J. Kitano. 2016. Genetic basis for variation in salinity tolerance between stickleback ecotypes. Molecular Ecology. 26 (1): 304–319. Kutzler, M. and D. Weiner. 2008. DNA vaccines: ready for prime time? Nature Reviews Genetics. 9: 776–788. Kyte, J. and R. F. Doolittle. 1982. A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology. 157 (1): 105–132. Langmead, B. and S. L. Salzberg. 2012. Fast gapped-read alignment with Bowtie 2. Nature Methods. 9 (4): 357–359. Le Bras, Y., N. Dechamp, F. Krieg, O. Filangi, R. Guyomard, M. Boussaha, H. Bovenhuis, T. G. Pottinger, P. Prunet, P. Le Roy and E. Quillet. 2011. Detection of QTL with effects on osmoregulation capacities in the rainbow trout (Oncorhynchus mykiss). BMC Genetics. 12 (1). Letunic, I., S. Khedkar and P. Bork. 2021. SMART: recent updates, new developments and status in 2020. Nucleic Acids Research. 49 (D1): D458–D460. Lewin, H. A., J. A. M. Graves, O. A. Ryder, A. S. Graphodatsky and S. J. O’Brien. 2019. Precision nomenclature for the new genomics. GigaScience. 8 (8). Li, H. 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 34 (18): 3094–3100. Li, H. and R. Durbin. 2009. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics. 25 (14): 1754–1760. Li, H. and R. Durbin. 2010. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics. 26 (5): 589–595. Li, H., B. Handsaker, A. Wysoker, T. Fennell, J. Ruan, N. Homer, G. Marth, G. Abecasis and R. Durbin. 2009. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 25 (16): 2078–2079. Li, K., P. Xu, J. Wang, X. Yi and Y. Jiao. 2023. Identification of errors in draft genome assemblies at single-nucleotide resolution for quality assessment and improvement. Nature Communications. 14 (1). Li, W., K. R. O’Neill, D. H. Haft and others. 2021. RefSeq: expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation. Nucleic Acids Research. 49 (D1): D1020–D1028.
129 Li, W. and A. Godzik. 2006. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 22 (13): 1658–1659. Lin, X., J. Tan, Y. Shen, B. Yang, Y. Zhang, Y. Liao, P. Wang, D. Zhou, G. Li and C. Tian. 2022. A high-density genetic linkage map and QTL mapping for sex in Clarias fuscus. Aquaculture. 561: 738723. Lin, Y., C. Ye, X. Li, Q. Chen, Y. Wu, F. Zhang, R. Pan, S. Zhang, S. Chen, X. Wang, S. Cao, Y. Wang, Y. Yue, Y. Liu and J. Yue. 2023. quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification. Horticulture Research. 10 (8). Lisachov, A., D. H. M. Nguyen, T. Panthum, S. F. Ahmad, W. Singchat, J. Ponjarat, K. Jaisamut, P. Srisapoome, P. Duengkae, S. Hatachote, K. Sriphairoj, N. Muangmai, S. Unajak, K. Han, U. Na-Nakorn and K. Srikulnath. 2023. Emerging importance of bighead catfish (Clarias macrocephalus) and north African catfish (C. gariepinus) as a bioresource and their genomic perspective. Aquaculture. 573: 739585. Lischer, H. E. L. and K. K. Shimizu. 2017. Reference-guided de novo assembly approach improves genome reconstruction for related species. BMC Bioinformatics. 18 (1). Liu, C., X. Hu, Z. Cao, Y. Sun, X. Chen and Z. Zhang. 2019. Construction and characterization of a DNA vaccine encoding the SagH against Streptococcus iniae.Fish & Shellfish Immunology. 89: 71–75. Liu, J., Y. Cao, H. Ma, H. Du, T. Liu, G. Wang, M. Liu, Q. Wang, P. Li and E. Wang. 2023a. Enolase-based nanovaccine immersion immunization induces robust immunity and protection against Streptococcus infection in tilapia. Aquaculture. 576. Liu, L. and others. 2009. Identification and experimental verification of protective antigens against Streptococcus suis serotype 2 based on genome sequence analysis. Current Microbiology. 58: 11–17. Liu, M., Y. Song, S. Zhang, L. Yu, Z. Yuan, H. Yang, M. Zhang, Z. Zhou, I. Seim, S. Liu, G. Fan and H. Yang. 2023b. A chromosome-level genome of electric catfish (Malapterurus electricus) provided new insights into order Siluriformes evolution. Marine Life Sciencec and Technology. 6 (1): 1–14. Liu, Y., L. Li, F. Yu and others. 2020. Genome-wide analysis revealed the virulence attenuation mechanism of the fish-derived oral attenuated Streptococcus iniae vaccine strain YM011. Fish & Shellfish Immunology. 106: 546–554. Mahmoud, M., Y. Huang, K. Garimella, P. A. Audano, W. Wan, N. Prasad, R. E. Handsaker,
136 Simão, F. A., R. M. Waterhouse, P. Ioannidis, E. V. Kriventseva and E. M. Zdobnov. 2015. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 31 (19): 3210–3212. Smolka, M., L. F. Paulin, C. M. Grochowski, D. W. Horner, M. Mahmoud, S. Behera, E. KalefEzra, M. Gandhi, K. Hong, D. Pehlivan, S. W. Scholz, C. M. B. Carvalho, C. Proukakis and F. J. Sedlazeck. 2024. Detection of mosaic and population-level structural variants with Sniffles2. Nature Biotechnology. 42 (10): 1571–1580. Spriestersbach, A., J. Kubicek, F. Schäfer, H. Block and B. Maertens. 2015. Purification of his-tagged proteins. Laboratory Methods in Enzymology: Protein Part D. pp. 1–15. Strait, B. J. and T. G. Dewey. 1996. The Shannon information entropy of protein sequences. Biophysical Journal. 71 (1): 148–155. Sturm, M., C. Schroeder and P. Bauer. 2016. SeqPurge: highly-sensitive adapter trimming for paired-end NGS data. BMC Bioinformatics. 17 (1). Su, W., X. Gu and T. Peterson. 2019. TIR-Learner, a New Ensemble Method for TIR Transposable Element Annotation, Provides Evidence for Abundant New Transposable Elements in the Maize Genome. Molecular Plant. 12 (3): 447–460. Sun, Y., Y. H. Hu, C. S. Liu and L. Sun. 2010. Construction and analysis of an experimental Streptococcus iniae DNA vaccine. Vaccine. 28 (23): 3905–3912. Sun, Y., Y. H. Hu, C. S. Liu and L. Sun. 2012. Streptococcus iniae DNA vaccine delivered by a live attenuated Edwardsiella tarda via natural infection induces cross-genus protection. Letters in Applied Microbiology. 55 (6): 420–426. Sun, Y., L. Sun, M. Q. Xing, C. S. Liu and Y. H. Hu. 2013. SagE induces highly effective protective immunity against Streptococcus iniae mainly through an immunogenic domain in the extracellular region. Acta Veterinaria Scandinavica. 55 (1): 78. Supikamolseni, A., N. Ngaoburanawit, M. Sumontha, L. Chanhome, S. Suntrarachun, S. Peyachoknagul and K. Srikulnath. 2015. Molecular barcoding of venomous snakes and species-specific multiplex PCR assay to identify snake groups for which antivenom is available in Thailand. Genetics and Molecular Research. 14 (4): 13981–13997. Sémon, M., D. Mouchiroud and L. Duret. 2005. Relationship between gene expression and GCcontent in mammals: statistical significance and biological relevance. Human Molecular Genetics. 14 (3): 421–427. Tanpichai, P. and others. 2023. Immune Activation Following Vaccination of Streptococcus iniae Bacterin in Asian Seabass (Lates calcarifer, Bloch 1790). Vaccines. 11 (2): 351.
137 Tatusova, T., M. DiCuccio, A. Badretdin, V. Chetvernin, E. P. Nawrocki, L. Zaslavsky, A. Lomsadze, K. D. Pruitt, M. Borodovsky and J. Ostell. 2016. NCBI prokaryotic genome annotation pipeline. Nucleic Acids Research. 44 (14): 6614–6624. Tettelin, H., D. Riley, C. Cattuto and D. Medini. 2008. Comparative genomics: the bacterial pan-genome. Current Opinion in Microbiology. 11 (5): 472–477. The Gene Ontology Consortium, . 2021. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Research. 49 (D1): D325–D334. Thorvaldsdottir, H., J. T. Robinson and J. P. Mesirov. 2012. Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration. Briefings in Bioinformatics. 14 (2): 178–192. Tonkin-Hill, G., N. MacAlasdair, C. Ruis, A. Weimann, G. Horesh, J. A. Lees, R. A. Gladstone, S. Lo, C. Beaudoin, R. A. Floto and others. 2020a. Producing polished prokaryotic pangenomes with the Panaroo pipeline. Genome Biology. 21: 180. Tonkin-Hill, G., N. MacAlasdair, C. Ruis, A. Weimann, G. Horesh, J. A. Lees, R. A. Gladstone, S. Lo, C. Beaudoin, R. A. Floto, S. D. Frost, J. Corander, S. D. Bentley and J. Parkhill. 2020b. Producing polished prokaryotic pangenomes with the Panaroo pipeline. Genome Biology. 21 (1). Treepong, P., C. Guyeux, A. Meunier, C. Couchoud, D. Hocquet and B. Valot. 2018. panISa: ab initio detection of insertion sequences in bacterial genomes from short read sequence data. Bioinformatics. 34 (22): 3795–3800. Vaser, R., I. Sović, N. Nagarajan and M. Šikić. 2017. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Research. 27 (5): 737–746. Vinogradov, A. E. 2005. Dualism of gene GC content and CpG pattern in regard to expression in the human genome: magnitude versus breadth. Trends in Genetics. 21 (12): 639–643. Vita, R., N. Blazeska, D. Marrama, I. C. T. Members, S. Duesing, J. Bennett, J. Greenbaum, M. Mendes, J. Mahita, D. Wheeler, J. Cantrell, J. Overton, D. Natale, A. Sette and B. Peters. 2025. The Immune Epitope Database (IEDB): 2024 update. Nucleic Acids Research. 53 (D1): D436–D443. Vita, R. and others. 2019. The Immune Epitope Database (IEDB): 2019 update. Nucleic Acids Research. 47 (D1): D339–D343. Walker, B. J., T. Abeel, T. Shea, M. Priest, A. Abouelliel, S. Sakthikumar, C. A. Cuomo, Q. Zeng, J. Wortman, S. K. Young and A. M. Earl. 2014. Pilon: An Integrated Tool for Comprehensive Microbial Variant Detection and Genome Assembly Improvement. PLoS
138 ONE. 9 (11): e112963. Wang, E., B. Long, K. Wang and others. 2016a. Interleukin-8 holds promise to serve as a molecular adjuvant in DNA vaccination model against Streptococcus iniae infection in fish. Oncotarget. 7 (51): 83938–83950. Wang, E., J. Wang, B. Long and others. 2016b. Molecular cloning, expression and the adjuvant effects of interleukin-8 of channel catfish (Ictalurus punctatus) against Streptococcus iniae.Scientific Reports. 6: 29310. Wang, J., L. L. Zou and A. X. Li. 2014. Construction of a Streptococcus iniae sortase A mutant and evaluation of its potential as an attenuated modified live vaccine in Nile tilapia (Oreochromis niloticus). Fish & Shellfish Immunology. 40 (2): 392–398. Wang, J., K. Wang, D. Chen, Y. Geng, X. Huang, Y. He, L. Ji, T. Liu, E. Wang, Q. Yang and W. Lai. 2015a. Cloning and characterization of surface-localized α -enolase of Streptococcus iniae, an effective protective antigen in mice. International Journal of Molecular Sciences. 16 (7): 14490–14510. Wang, L., D. Xing, A. Le Van, A. E. Jerse and S. Wang. 2017. Structure-based design of ferritin nanoparticle immunogens displaying antigenic loops of Neisseria gonorrhoeae.FEBS Open Bio. 7 (2): 262–272. Wang, L., Z. Y. Wan, B. Bai, S. Q. Huang, E. Chua, M. Lee, H. Y. Pang, Y. F. Wen, P. Liu, F. Liu, F. Sun, G. Lin, B. Q. Ye and G. H. Yue. 2015b. Construction of a high-density linkage map and fine mapping of QTL for growth in Asian seabass. Scientific Reports. 5 (1). Wenger, A. M., P. Peluso, W. J. Rowell, P.-C. Chang, R. J. Hall, G. T. Concepcion, J. Ebler, A. Fungtammasan, A. Kolesnikov, N. D. Olson, A. Töpfer, M. Alonge, M. Mahmoud, Y. Qian, C.-S. Chin, A. M. Phillippy, M. C. Schatz, G. Myers, M. A. DePristo, J. Ruan, T. Marschall, F. J. Sedlazeck, J. M. Zook, H. Li, S. Koren, A. Carroll, D. R. Rank and M. W. Hunkapiller. 2019. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nature Biotechnology. 37 (10): 1155–1162. Wick, R. R., M. B. Schultz, J. Zobel and K. E. Holt. 2015. Bandage: interactive visualization of de novo genome assemblies. Bioinformatics. 31 (20): 3350–3352. Wick, R. R., L. M. Judd, C. L. Gorrie and K. E. Holt. 2017. Unicycler: Resolving bacterial genome assemblies from short and long sequencing reads. PLOS Computational Biology. 13 (6): e1005595. Widmann, M., P. Trodler and J. Pleiss. 2010. The isoelectric region of proteins: A systematic
139 analysis. PLoS ONE. 5 (1): e10546. Woestenenk, E. A., M. Hammarström, S. van denBerg, T. Härd and H. Berglund. 2004. His tag effect on solubility of human proteins produced in Escherichia coli: a comparison between four expression vectors. Journal of Structural and Functional Genomics. 5 (3): 217–229. Wyneken, J., S. P. Epperly, L. B. Crowder, J. Vaughan and K. Blair Esper. 2007. DETERMINING SEX IN POSTHATCHLING LOGGERHEAD SEA TURTLES USING MULTIPLE GONADAL AND ACCESSORY DUCT CHARACTERISTICS. Herpetologica. 63 (1): 19–30. Xiong, W., L. He, J. Lai, H. K. Dooner and C. Du. 2014. HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes. Proceedings of the National Academy of Sciences. 111 (28): 10263–10268. Xiong, X., Y. Peng, R. Chen, X. Liu and F. Jiang. 2023. Efficacy and transcriptome analysis of golden pompano (Trachinotus ovatus) immunized with a formalin-inactivated vaccine against Streptococcus iniae.Fish & Shellfish Immunology. 134: 108489. Xu, M., L. Guo, S. Gu, O. Wang, R. Zhang, B. A. Peters, G. Fan, X. Liu, X. Xu, L. Deng and Y. Zhang. 2020. TGS-GapCloser: A fast and accurate gap closer for large genomes with low coverage of error-prone long reads. GigaScience. 9 (9). Xu, Z. and H. Wang. 2007. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Research. 35 (Web Server): W265–W268. Yan, J.-J., Y.-C. Lee, Y.-L. Tsou, Y.-C. Tseng and P.-P. Hwang. 2020. Insulin-like growth factor 1 triggers salt secretion machinery in fish under acute salinity stress. Journal of Endocrinology. 246 (3): 277–288. Yang, Z., X. Zeng, Y. Zhao and others. 2023. AlphaFold2 and its applications in the fields of biology and medicine. Signal Transduction and Targeted Therapy. 8: 115. Yu, L. X. and others. 2014. Understanding pharmaceutical quality by design. AAPS Journal. 16 (4): 771–783. Yu, X., P. Setyawan, J. W. Bastiaansen, L. Liu, I. Imron, M. A. Groenen, H. Komen and H.-J. Megens. 2022. Genomic analysis of a Nile tilapia strain selected for salinity tolerance shows signatures of selection and hybridization with blue tilapia (Oreochromis aureus). Aquaculture. 560: 738527. Yue, G. H. 2013. Recent advances of genome mapping and marker�assisted selection in aquaculture. Fish and Fisheries. 15 (3): 376–396.
140 Zayas, J. F. 1997. Solubility of Proteins. Functionality of Proteins in Food. pp. 6–75. Zeng, X., Z. Yi, X. Zhang, Y. Du, Y. Li, Z. Zhou, S. Chen, H. Zhao, S. Yang, Y. Wang and G. Chen. 2024. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. Nature Plants. 10 (8): 1184–1200. Zhang, B.-C., J. Zhang and L. Sun. 2014a. Streptococcus iniae SF1: Complete genome sequence, proteomic profile, and immunoprotective antigens. PLoS ONE. 9 (3): e91324. Zhang, B.-c., J. Zhang and L. Sun. 2014b. Streptococcus iniae SF1: Complete Genome Sequence, Proteomic Profile, and Immunoprotective Antigens. PLoS ONE. 9 (3): e91324. Zhang, R.-G., G.-Y. Li, X.-L. Wang, J. Dainat, Z.-X. Wang, S. Ou and Y. Ma. 2022. TEsorter: An accurate and fast method to classify LTR-retrotransposons in plant genomes. Horticulture Research. 9. Zheng, Z., S. Li, J. Su, A. W.-S. Leung, T.-W. Lam and R. Luo. 2022. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nature Computational Science. 2 (12): 797–803. Zhou, Z., Y. Dang, M. Zhou, L. Li, C. Yu, J. Fu, S. Chen and Y. Liu. 2016. Codon usage is an important determinant of gene expression levels largely through its effects on transcription. Proceedings of the National Academy of Sciences. 113 (41): E6117–E6125.
141 Appendix
142 Figure 30 GenomeScope2.0 profiles for (a) male C. gariepinus, (b) female C. macrocephalus, and the F1 hybrid catfish at (c) k=21 and (d) k=31. The hybrid genome shows an intermediate heterozygosity level ( 1%) and genome size of 1.8 Gb, consistent with contributions from both parental subgenomes. Parental species exhibit 0.056% (C. macrocephalus) and 1.56% (C. gariepinus) heterozygosity, respectively. BUSCO analyses confirmed assembly completeness and the presence of two subgenomes. Low Illumina coverage (<20×) in parental datasets resulted in broad peaks in panels (a–b).
143 Figure 31 Comparative synteny and structural variation between the North African catfish (C. gariepinus) reference genome (GCA_024256425.2) and the hybrid catfish genome (fClaHyb_Gar, this study). (A) Genome-wide macrosynteny across 28 pseudochromosomes showing conserved collinearity and localized inversions, translocations, and duplications. (B) Sequence variation lengths and percentages of genome size, highlighting highly divergent regions. (C) Structural variation composition including duplications, translocations, and inversions. (D) Variant feature counts (SNPs, indels, CNVs, tandem repeats) showing genome-wide heterogeneity.
144 Figure 32 Comparative synteny and structural variation between the Bighead catfish (C. macrocephalus) Haplotype 1 (GCA_048544425.1) and the hybrid genome (fClaHyb_Mac, this study). (A) Genome-wide macrosynteny across 27 pseudochromosomes showing extensive collinearity with limited rearrangements. (B) Sequence variation profiles showing divergence and insertion–deletion patterns. (C) Structural variation composition between parental and hybrid genomes. (D) Variant feature counts across all categories, emphasizing SNP dominance and minor structural variants.
145 Figure 33 Circular genome synteny and quality control overview of the five sequenced Streptococcus iniae strains (SIKU01–SIKU05). The Circos plot shows inter-strain genomic alignments, GC content variation, GC skew, and contig connectivity metrics. Outer rings represent genome coordinates, while inner links highlight conserved syntenic regions among isolates, confirming overall structural stability across strains.