scieee AI-readable full text Open interactive document viewer

Comparative Genomics Reveals 13 Different Isoforms of Mytimycins (A-M) in Mytilus galloprovincialis

Rey-Campos, Magalí,Novoa, Beatriz,Pallavicini, Alberto,Gerdol, Marco,Figueras Huerta, Antonio

Abstract

15 pages, 9 figures, 1 table.--This is an open access article distributed under the Creative Commons Attribution License

Full text

International Journal of Molecular Sciences Article Comparative Genomics Reveals 13 Different Isoforms of Mytimycins (A-M) in Mytilus galloprovincialis MagalíRey-Campos 1, Beatriz Novoa 1, Alberto Pallavicini 2,3, Marco Gerdol 2,* and Antonio Figueras 1,*   Citation: Rey-Campos, M.; Novoa, B.; Pallavicini, A.; Gerdol, M.; Figueras, A. Comparative Genomics Reveals 13 Different Isoforms of Mytimycins (A-M) in Mytilus galloprovincialis.Int. J. Mol. Sci. 2021, 22, 3235. https://doi.org/10.3390/ ijms22063235 Academic Editor: Igor Rogozin Received: 1 February 2021 Accepted: 19 March 2021 Published: 22 March 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1Institute of Marine Research (IIM), CSIC, Eduardo Cabello 6, 36208 Vigo, Spain; [email protected] (M.R.-C.); [email protected] (B.N.) 2Department of Life Sciences, University of Trieste, Via Giorgieri 5, 34127 Trieste, Italy; [email protected] 3National Institute of Oceanography and Applied Geophysics–OGS, via Auguste Piccard, 54, 34151 Trieste, Italy *Correspondence: [email protected] (M.G.); [email protected] (A.F.); Tel.: +34-986214462 (A.F.) Abstract: Mytimycins are cysteine-rich antimicrobial peptides that show antifungal properties. These peptides are part of the immune network that constitutes the defense system of the Mediterranean mussel (Mytilus galloprovincialis). The immune system of mussels has been increasingly studied in the last decade due to its great efficiency, since these molluscs, particularly resistant to adverse conditions and pathogens, are present all over the world, being considered as an invasive species. The recent sequencing of the mussel genome has greatly simplified the genetic study of some of its immune genes. In the present work, we describe a total of 106 different mytimycin variants in 16 individual mussel genomes. The 13 highly supported mytimycin clusters (A–M) identified with phylogenetic inference were found to be subject to the presence/absence variation, a widespread phenomenon in mussels. We also identified a block of conserved residues evolving under purifying selection, which may indicate the “functional core” of the mature peptide, and a conserved set of 10 invariable plus 6 accessory cysteines which constitute a plastic disulfide array. Finally, we extended the taxonomic range of distribution of mytimycins among Mytilida, identifying novel sequences in M. coruscus,M. californianus,P. viridis,L. fortunei,M. philippinarum,M. modiolus, and P. purpuratus. Keywords: Mytilus galloprovincialis; mytimycins; mussel genome; RNA-seq; isoelectric point; positive and negative selection; promoter 1. Introduction Mytimycins are cysteine-rich antimicrobial peptides isolated for the first time from Mytilus edulis in 1996 [ 1 ]. Like several other antimicrobial peptides (AMPs), mytimycins contain several conserved cysteines in the mature peptide, whose connectivity is still unclear [ 1 , 2 ]. Since these molecules show antifungal properties [ 3 ], they have been notably less studied than other mussel antibacterial and antiviral peptides. Although multiple mytimycin variants have been previously reported in Mytilus galloprovincialis so far [ 4 ], only a single variant, described in 2011, has been the subject of detailed studies [ 4 ]. This AMP is produced as a precursor peptide, which includes a signal peptide of 23 amino acids, a mature peptide of 54 amino acids (containing 12 cysteines), and a C-terminal extension of 75 amino acids, which contains an EF hand-motif [4]. The mytimycin gene is mainly expressed in circulating hemocytes, the central immune cells of mussels. These AMPs are strongly up-regulated after a stimulation with the fungus Fusarium oxysporum [ 5 ]. However, the relevant inter-individual differences in response to specific stimuli in these animals [ 6 ] have hampered the definition of consistent patterns of regulation for the mytimycin gene [ 7 ]. Mytimycins are part of the complex immune network that constitutes the defense system of the Mediterranean mussel, along with other AMPs such as defensins, myticins, or mytilins [ 8 – 10 ]. These, in association with other recognition and effector molecules, create a highly efficient immune system, which likely Int. J. Mol. Sci. 2021,22, 3235. https://doi.org/10.3390/ijms22063235 https://www.mdpi.com/journal/ijms Int. J. Mol. Sci. 2021,22, 3235 2 of 15 prevents the massive mortality events that often occur in the natural environment for other bivalve species [ 11 , 12 ]. The advance of massive sequencing technologies has provided a strong contribution in unveiling this scenario, in particular thanks to the release of the Mytilus galloprovincialis genome sequence [ 13 ], which revealed the presence of widespread gene presence/absence variation for the first time in a metazoan. This dispensable nature of 25% of the mussel coding genes undoubtedly offers a source of enormous genetic variability. A striking example of this molecular diversity is represented by another class of cysteine-rich AMPs, i.e., myticins, which show a great level of inter-individual variability, with very little overlap among individuals (120 different isoforms were described in only 16 individuals) [14]. The main objective of this work was to continue studying the molecular diversity of the immune genes of mussel Mytilus galloprovincialis, extending investigations to the case of mytimycins. We explored the sequence variants belonging to this gene family, exploiting the availability of a reference genome and resequencing data of 16 different individuals [ 13 ]. We also investigated several previously unexplored aspects that might help to further characterize the biological roles of these molecules, such as their isoelectric point, the regulatory elements present in the promoter region of the gene, and their expression in different animals and tissues. 2. Results 2.1. Searching, Screening, and Identifying M. galloprovincialis Mytimycins A total of 106 different nucleotide sequences encoding mytimycins were found in the 16 mussel genome assemblies, i.e., all the mytimycin variants that were present in the analyzed genomes. These 106 sequences code 94 different peptides, 76 of which are mytimycins with an uninterrupted CDS and 18 of which are pseudogenes (sequences that incorporate a STOP codon which interrupts the open reading frame) (File S1). 2.2. Phylogenetic Analysis The Bayesian phylogenetic tree of all the 106 mytimycin variants identified in this study (Figure 1) displayed a remarkable subdivision of mytimycins in 13 groups (A–M), with well-supported branches posterior probabilities. All the sequences grouped within the same cluster showed a pairwise identity threshold higher than 95% (File S1) and displayed a variable number of cysteine residues, as will be reported in detail below along with the report of mytimycins in other Mytilida. An alignment of the peptide sequence of all the variants is shown in the Figure S1. 2.3. Presence/Absence Variation An evaluation of the presence/absence of all the 106 sequences was performed in the 16 mussel genomes (File S2). On average, each mussel genome showed nine different mytimycin sequences, with each of the 16 individuals analyzed displaying an average of five unique variants (i.e., variants not present in any of the other 15 genomes). This fact highlights a scenario of enormous diversity, which results in a unique collection of mytimycin variants in each mussel. The sequences which displayed the highest frequency of occurrence are F1, F2, E1, and J1. Even so, none of them were present in all the genomes analyzed. As several of the 106 variants identified only displayed minor differences in pairwise comparisons, we cannot exclude the possibility that they represent polymorphic alleles of the same gene. To investigate this possibility, we generated a cladogram showing which variants were present in each genome (Figure 2). Int. J. Mol. Sci. 2021,22, 3235 3 of 15 Int. J. Mol. Sci. 2021, 22, 3235 3 of 18 Figure 1. Phylogenetic analysis. The full coding region of 106 mytimycin sequences obtained from 16 genomes were subjected to multiple sequence alignment and analyzed with Bayesian phylogenetic inference. The employed evolution model was the Hasegawa, Kishino, and Yano 1985 model (HKY + G). Thirteen clusters of sequences are defined with high posterior probability support. * indicates previously described isoforms. 2.3. Presence/Absence Variation An evaluation of the presence/absence of all the 106 sequences was performed in the 16 mussel genomes (File S2). On average, each mussel genome showed nine different mytimycin sequences, with each of the 16 individuals analyzed displaying an average of five unique variants (i.e., variants not present in any of the other 15 genomes). This fact highlights a scenario of enormous diversity, which results in a unique collection of mytimycin variants in each mussel. The sequences which displayed the highest frequency of occurrence are F1, F2, E1, and J1. Even so, none of them were present in all the genomes analyzed. As several of the 106 variants identified only displayed minor differences in pairwise comparisons, we cannot exclude the possibility that they represent polymorphic alleles of the same gene. To investigate this possibility, we generated a cladogram showing which variants were present in each genome (Figure 2). Figure 1. Phylogenetic analysis. The full coding region of 106 mytimycin sequences obtained from 16 genomes were subjected to multiple sequence alignment and analyzed with Bayesian phylogenetic inference. The employed evolution model was the Hasegawa, Kishino, and Yano 1985 model (HKY + G). Thirteen clusters of sequences are defined with high posterior probability support. * indicates previously described isoforms. Int. J. Mol. Sci. 2021, 22, 3235 4 of 18 Figure 2. Presence/absence variation cladogram. The graph shows the presence/absence variation of each mytimycin nucleotide variant in each mussel genome, along with the arbitrarily rooted phylogenetic tree shown in Figure 1. Numbers close to each node are posterior probabilities. To verify to which extent mytimycins were subjected to presence/absence variation, we used a maximum parsimony approach, tentatively assigning all the 106 sequence variants to the 13 clusters identified by the phylogenetic analysis, under the assumption that each cluster may include different allelic variants belonging to the same gene. This approach confirmed that PAV was the most likely explanation for the observed patterns of distribution, as different mytimycin clusters displayed largely different frequencies of occurrence (Figure 3). Notably, the clusters belonging to a monophyletic branch of the tree (MKJAI) were generally absent in genomes, whereas cluster D was present in all genomes, and B and C were found in 15 out of the 16 studied genomes. Figure 2. Presence/absence variation cladogram. The graph shows the presence/absence variation of each mytimycin nucleotide variant in each mussel genome, along with the arbitrarily rooted phylogenetic tree shown in Figure 1. Numbers close to each node are posterior probabilities. Int. J. Mol. Sci. 2021,22, 3235 4 of 15 To verify to which extent mytimycins were subjected to presence/absence variation, we used a maximum parsimony approach, tentatively assigning all the 106 sequence variants to the 13 clusters identified by the phylogenetic analysis, under the assumption that each cluster may include different allelic variants belonging to the same gene. This approach confirmed that PAV was the most likely explanation for the observed patterns of distribution, as different mytimycin clusters displayed largely different frequencies of occurrence (Figure 3). Notably, the clusters belonging to a monophyletic branch of the tree (MKJAI) were generally absent in genomes, whereas cluster D was present in all genomes, and B and C were found in 15 out of the 16 studied genomes. Int. J. Mol. Sci. 2021, 22, 3235 5 of 18 Figure 3. Clusters presence/absence variation. Matrix shows presence/absence of the 13 mytimycin clusters in all genomes, along with a Bayesian phylogenetic tree. Numbers close to each node are posterior probabilities. 2.4. Isoelectric Point The isoelectric point and predicted charge at cytoplasmic pH (i.e., 7.4) of all the mytimycins with an uninterrupted CDS were obtained (mature peptide), and they are reported in File S3. Figure 4A shows that the mature peptides of almost all the mytimycins display very narrow variations in terms of isoelectric point (pI), which varies between 7 and 9. However, the C isoform has a slightly higher pI (between 9 and 10). In terms of charge, the mature peptides are usually cationic, with a positive net charge varying from 0 to 10 (Figure 4B), with the C isoform once again showing the highest values. Figure 3. Clusters presence/absence variation. Matrix shows presence/absence of the 13 mytimycin clusters in all genomes, along with a Bayesian phylogenetic tree. Numbers close to each node are posterior probabilities. 2.4. Isoelectric Point The isoelectric point and predicted charge at cytoplasmic pH (i.e., 7.4) of all the mytimycins with an uninterrupted CDS were obtained (mature peptide), and they are reported in File S3. Figure 4A shows that the mature peptides of almost all the mytimycins display very narrow variations in terms of isoelectric point (pI), which varies between 7 and 9. However, the C isoform has a slightly higher pI (between 9 and 10). In terms of charge, the mature peptides are usually cationic, with a positive net charge varying from 0 to 10 (Figure 4B), with the C isoform once again showing the highest values. The sliding-window analysis of pI along the whole peptide, carried out on the 13 representative precursor peptides of mytimycins (Figure 4C), showed very similar profiles, with a notable decrease at the C-terminal region, characterized by the presence of several acidic residues. 2.5. Positive and Negative Selection Analysis The selection analysis mostly identified sites subject to negative selection in the mature peptide region. Despite some minor differences, the various tests were concordant in recognizing a major block of residues under pervasive purifying selection between positions 30 and 43 in the multiple sequence alignment (Figure 5). These sites match with four of the cysteines that form the characteristic disulfide array of mytimycins and could constitute the “core functional region” of these peptides. Some other cases of negatively selected codons were also identified along the mature peptide regions, mostly corresponding to conserved cysteines. Moreover, the 3 0 end of the mature peptide region displays signatures of negative selection in conjunction with a dibasic site, which may be recognized as a furin-like proteolytic cleavage site of the precursor protein. On the other hand, the few (i.e., 8 out of 67) codons under positive selection were spread along Int. J. Mol. Sci. 2021,22, 3235 5 of 15 the mature peptide region and they were often accompanied by low statistical support (p-value < 0.09). Int. J. Mol. Sci. 2021, 22, 3235 6 of 18 Figure 4. Isoelectric Point. (A) Isoelectric point (X axis) and molecular weight (Y axis) of the mature peptide of mytimycins with an uninterrupted CDS. (B) Isoelectric point (X axis) and charge at pH = 7.4 (Y axis) of the mature peptide of mytimycins with an uninterrupted CDS. (C) Isoelectric point of the whole sequence of the 13 mytimycin isoforms (consensus sequence of each isoform). The isoelectric point distribution was analyzed through the calculation of the average isoelectric point based on a sliding window of 15 amino acids. The alignment shows the consensus sequence of the 13 mytimycins clusters obtained from 16 genomes. X represents polymorphic amino acid residues found in each cluster. * represents STOP codons. The sliding-window analysis of pI along the whole peptide, carried out on the 13 representative precursor peptides of mytimycins (Figure 4C), showed very similar profiles, with a notable decrease at the C-terminal region, characterized by the presence of several acidic residues. 2.5. Positive and Negative Selection Analysis The selection analysis mostly identified sites subject to negative selection in the mature peptide region. Despite some minor differences, the various tests were concordant in recognizing a major block of residues under pervasive purifying selection between positions 30 and 43 in the multiple sequence alignment (Figure 5). These sites match with four of the cysteines that form the characteristic disulfide array of mytimycins and could constitute the “core functional region” of these peptides. Some other cases of negatively Figure 4. Isoelectric Point. ( A ) Isoelectric point (X axis) and molecular weight (Y axis) of the mature peptide of mytimycins with an uninterrupted CDS. ( B ) Isoelectric point (X axis) and charge at pH = 7.4 (Y axis) of the mature peptide of mytimycins with an uninterrupted CDS. ( C ) Isoelectric point of the whole sequence of the 13 mytimycin isoforms (consensus sequence of each isoform). The isoelectric point distribution was analyzed through the calculation of the average isoelectric point based on a sliding window of 15 amino acids. The alignment shows the consensus sequence of the 13 mytimycins clusters obtained from 16 genomes. X represents polymorphic amino acid residues found in each cluster. * represents STOP codons. Int. J. Mol. Sci. 2021,22, 3235 6 of 15 Int. J. Mol. Sci. 2021, 22, 3235 7 of 18 selected codons were also identified along the mature peptide regions, mostly corresponding to conserved cysteines. Moreover, the 3′ end of the mature peptide region displays signatures of negative selection in conjunction with a dibasic site, which may be recognized as a furin-like proteolytic cleavage site of the precursor protein. On the other hand, the few (i.e., 8 out of 67) codons under positive selection were spread along the mature peptide region and they were often accompanied by low statistical support (pvalue < 0.09). Figure 5. Positive and negative selection analysis. The mature peptide of the mytimycins with an uninterrupted CDS were selected to perform the analysis. Four prediction models have been used (MEME, FEL, FUBAR, and SLAC). Positive selection sites and negative selection sites are collected on the graph, using +/− symbols respectively. Orange arrows point to the beginning and end site of the mature peptide, marking the signal peptide and propeptide cleavage sites, respectively. 2.6. Promoter Analysis The promoter analysis identified 10 different motifs located in the 1000 pb upstream of the gene transcription start site (Figure 6). The most conserved motif shared by almost all the available sequences among the 13 clusters (i.e., the promoter regions could not be retrieved in some cases due to the excessive fragmentation of the genome assemblies) was the motif 1 (TGCTTGTTTAYTTWTAACAAYAATTTTAAG). This motif was only missing in the H cluster, which included several sequences with in-frame stop codons, and therefore is likely pseudogenic (Figure S1). The two phylogenetically close clusters J and K (Figure 1) shared several other motifs that could not be found in any of the other clusters. In general, mytimycin genes do not appear to show a robust promoter, as evidenced by the lack of conserved motifs in fixed relative positions. Figure 5. Positive and negative selection analysis. The mature peptide of the mytimycins with an uninterrupted CDS were selected to perform the analysis. Four prediction models have been used (MEME, FEL, FUBAR, and SLAC). Positive selection sites and negative selection sites are collected on the graph, using +/ − symbols respectively. Orange arrows point to the beginning and end site of the mature peptide, marking the signal peptide and propeptide cleavage sites, respectively. 2.6. Promoter Analysis The promoter analysis identified 10 different motifs located in the 1000 pb upstream of the gene transcription start site (Figure 6). The most conserved motif shared by almost all the available sequences among the 13 clusters (i.e., the promoter regions could not be retrieved in some cases due to the excessive fragmentation of the genome assemblies) was the motif 1 (TGCTTGTTTAYTTWTAACAAYAATTTTAAG). This motif was only missing in the H cluster, which included several sequences with in-frame stop codons, and therefore is likely pseudogenic (Figure S1). The two phylogenetically close clusters J and K (Figure 1) shared several other motifs that could not be found in any of the other clusters. In general, mytimycin genes do not appear to show a robust promoter, as evidenced by the lack of conserved motifs in fixed relative positions. Int. J. Mol. Sci. 2021, 22, 3235 8 of 18 Figure 6. Promoter analysis. The promoter (1 kb) of 7 complete available isoforms was analyzed. Colored boxes show the 10 different significant motifs found. The p-value of the combined position of all motifs present in each sequence is also shown. 2.7. Genomic vs. Transcriptomic Data Having the genome and transcriptome of the same individual offers a great opportunity to investigate RNA editing processes which may occur after the transcription [13]. Six different mytimycin sequence variants were found in the reference genome (LOLA) (File S2) and only four of these were present in the transcriptome assembly, indicating that two of the genomic forms were not expressed. The comparison between the genomic sequences and mRNA sequences of the four expressed mytimycin variants highlighted that there were no discrepancies, allowing us to disregard mRNA editing as a plausible phenomenon responsible for the molecular diversity observed in this family of AMPs. 2.8. Expression Analysis Five different transcriptomes from different mussel tissues and geographical locations were analyzed to investigate whether sequences belonging to the 13 previously defined clusters were broadly expressed (Figure 7). Some of these variants, such as those from clusters D and J, were broadly expressed across the different samples analyzed (e.g., in 4 out of the 5 transcriptomes analyzed). Others were scarcely expressed or, as in the case of the M, H, and E clusters, not expressed at all in any of the studied transcriptomes. Figure 6. Promoter analysis. The promoter (1 kb) of 7 complete available isoforms was analyzed. Colored boxes show the 10 different significant motifs found. The p-value of the combined position of all motifs present in each sequence is also shown. Int. J. Mol. Sci. 2021,22, 3235 7 of 15 2.7. Genomic vs. Transcriptomic Data Having the genome and transcriptome of the same individual offers a great opportunity to investigate RNA editing processes which may occur after the transcription [ 13 ]. Six different mytimycin sequence variants were found in the reference genome (LOLA) (File S2) and only four of these were present in the transcriptome assembly, indicating that two of the genomic forms were not expressed. The comparison between the genomic sequences and mRNA sequences of the four expressed mytimycin variants highlighted that there were no discrepancies, allowing us to disregard mRNA editing as a plausible phenomenon responsible for the molecular diversity observed in this family of AMPs. 2.8. Expression Analysis Five different transcriptomes from different mussel tissues and geographical locations were analyzed to investigate whether sequences belonging to the 13 previously defined clusters were broadly expressed (Figure 7). Some of these variants, such as those from clusters D and J, were broadly expressed across the different samples analyzed (e.g., in 4 out of the 5 transcriptomes analyzed). Others were scarcely expressed or, as in the case of the M, H, and E clusters, not expressed at all in any of the studied transcriptomes. Int. J. Mol. Sci. 2021, 22, 3235 9 of 18 Figure 7. Expression analysis. A total of five transcriptomes of different tissues and mussel locations were analyzed to find evidence of expression of sequences belonging to each cluster. The isoforms expressed in each transcriptome are highlighted in yellow. Red color indicates no evidence of expression. 2.9. Taxonomical Distribution of Mytimycins Although the -omic information available for marine mussels is still incomplete, the screening of available genomes and transcriptomes allowed us to extend their range of distribution beyond M. galloprovincialis and M. edulis [1,3]. We showed that mytimycins are found in M. coruscus and M. californianus, belonging to the same genus, but not interfertile with the other species included in the M. edulis species complex. Mytimycins were also present in Perna viridis (subfamily Mytilinae) [15], in Limnoperna fortunei (subfamily Arcuatulinae) [16], in Modiolus philippinarum and Modiolus modiolus (subfamily Modiolinae) [17], and in Perumytilus purpuratus (subfamily Brachidontinae) [18], but not in Bathymodiolinae [17]. The mytimycin sequences detected in the different mytilid subfamilies widely differed due to the number of cysteine residues, which varied from 10 to 14, and due to the type of disulfide array present (Table 1). Mytimycins adopted five possible types of disulfide arrays, with only three of them (i.e., type I, type II, and type III) found in M. galloprovincialis. Type II array, found in the A, I, J, and K groups within M. galloprovincialis, denoted the more widespread architecture, found in all Mytilida, except L. fortunei. Curiously, the type I array, characterizing groups B, D, E, F, G, L, and M in the Mediterranean mussel, was exclusively present in Mytilus spp. Type III array was only shared by Mytilus spp. and L. fortunei, whereas the type IV array, found in three different mytilid subfamilies, was not present in M. galloprovincialis and its congeneric species. Finally, a particular disulfide array with 14 cysteine residues was only found in sequences of the golden mussel. Figure 7. Expression analysis. A total of five transcriptomes of different tissues and mussel locations were analyzed to find evidence of expression of sequences belonging to each cluster. The isoforms expressed in each transcriptome are highlighted in yellow. Red color indicates no evidence of expression. 2.9. Taxonomical Distribution of Mytimycins Although the -omic information available for marine mussels is still incomplete, the screening of available genomes and transcriptomes allowed us to extend their range of distribution beyond M. galloprovincialis and M. edulis [ 1 , 3 ]. We showed that mytimycins are found in M. coruscus and M. californianus, belonging to the same genus, but not inter-fertile with the other species included in the M. edulis species complex. Mytimycins were also present in Perna viridis (subfamily Mytilinae) [ 15 ], in Limnoperna fortunei (subfamily Arcuatulinae) [ 16 ], in Modiolus philippinarum and Modiolus modiolus (subfamily Modiolinae) [ 17 ], and in Perumytilus purpuratus (subfamily Brachidontinae) [ 18 ], but not in Bathymodiolinae [17]. The mytimycin sequences detected in the different mytilid subfamilies widely differed due to the number of cysteine residues, which varied from 10 to 14, and due to the type of disulfide array present (Table 1). Mytimycins adopted five possible types of disulfide Int. J. Mol. Sci. 2021,22, 3235 8 of 15 arrays, with only three of them (i.e., type I, type II, and type III) found in M. galloprovincialis. Type II array, found in the A, I, J, and K groups within M. galloprovincialis, denoted the more widespread architecture, found in all Mytilida, except L. fortunei. Curiously, the type I array, characterizing groups B, D, E, F, G, L, and M in the Mediterranean mussel, was exclusively present in Mytilus spp. Type III array was only shared by Mytilus spp. and L. fortunei, whereas the type IV array, found in three different mytilid subfamilies, was not present in M. galloprovincialis and its congeneric species. Finally, a particular disulfide array with 14 cysteine residues was only found in sequences of the golden mussel. Table 1. Taxonomic distribution of the mytimycin cysteine arrays observed in Mytilida. For M. galloprovincialis, the type Cys Array Details Cys Residues M. galloprovincialis Mytilus spp P. viridis Modiolus spp L. fortunei P. purpuratus B. platifrons type I CCC-C-CC-CCC-C-C-C-C-C 14 √(B, D, E, F, G, L, M) √× × × × × type II CCC-CC-CC-C-C-C-C-C 12 √(A, I, J, K) √ √ √ ×√× type III CCC-C-CC-CCC-C-C-C 12 √(C, H) √× × √× × type IV CCC-CC-CC-C-C-C 10 × × √ √ ×√× type V CC-C-C-C-C-CCC-CC-C-C-C 14 × × × × √× × Based on the multiple alignment of the sequences, and on the presence/absence of paired cysteine residues in the five aforementioned types of disulfide arrays, it was possible to tentatively predict the connectivity of a few cysteine residues, even though the threedimensional structure of mytimycins currently remains unsolved. These considerations are limited to the three optional disulfide bonds, whereas the connectivity among the 10 cysteine residues which characterize the backbone of mytimycins are yet to be investigated (Figure 8). Int. J. Mol. Sci. 2021, 22, 3235 10 of 18 Table 1. Taxonomic distribution of the mytimycin cysteine arrays observed in Mytilida. For M. galloprovincialis, the type Cys Array Details Cys Residues M. galloprovincialis Mytilu s spp P. viridis Modio lus spp L. fortun ei P. purpur atus B. platifr ons type I CCC-C-CC-C-CC-C-C-C-C-C 14 √ (B, D, E, F, G, L, M) √ × × × × × type II CCC-CC-CC-C-C-C-C-C 12 √ (A, I, J, K) √ √ √ × √ × type III CCC-C-CC-C-CC-C-C-C 12 √ (C, H) √ × × √ × × type IV CCC-CC-CC-C-C-C 10 × × √ √ × √ × type V CC-C-C-C-C-CC-C-CC-C-C-C 14 × × × × √ × × Based on the multiple alignment of the sequences, and on the presence/absence of paired cysteine residues in the five aforementioned types of disulfide arrays, it was possible to tentatively predict the connectivity of a few cysteine residues, even though the three-dimensional structure of mytimycins currently remains unsolved. These considerations are limited to the three optional disulfide bonds, whereas the connectivity among the 10 cysteine residues which characterize the backbone of mytimycins are yet to be investigated (Figure 8). Figure 8. Predicted disulfide connectivity based on in silico analyses. The 10 cysteine residues which compose the backbone of mytimycins are highlighted with a yellow background. The three optional bonds are as follows: (i) bond 1 connects two consecutive cysteine residues found towards the N-terminal portion of the mature peptide of type V mytimycins, between the 2nd and 3rd Cys of the backbone array; (ii) bond 2 is found in type I, III, and V mytimycins; this disulfide bond connects the two optional cysteine residues found in tandem with the 5th and 7th Cys residues of the backbone array, and (iii) bond 3 is only found in type I and type II mytimycins and connects the two cysteine residues found at the C-terminal end of the mature peptide. From an evolutionary perspective, the phylogenetic clustering of mytimycins was based on taxonomy rather than on the type of disulfide array (Figure 9). The observed topology of the tree strongly hints that different disulfide arrays have been independently acquired in the three main clades. Figure 8. Predicted disulfide connectivity based on in silico analyses. The 10 cysteine residues which compose the backbone of mytimycins are highlighted with a yellow background. The three optional bonds are as follows: (i) bond 1 connects two consecutive cysteine residues found towards the N-terminal portion of the mature peptide of type V mytimycins, between the 2nd and 3rd Cys of the backbone array; (ii) bond 2 is found in type I, III, and V mytimycins; this disulfide bond connects the two optional cysteine residues found in tandem with the 5th and 7th Cys residues of the backbone array, and (iii) bond 3 is only found in type I and type II mytimycins and connects the two cysteine residues found at the C-terminal end of the mature peptide. From an evolutionary perspective, the phylogenetic clustering of mytimycins was based on taxonomy rather than on the type of disulfide array (Figure 9). The observed topology of the tree strongly hints that different disulfide arrays have been independently acquired in the three main clades. Int. J. Mol. Sci. 2021,22, 3235 9 of 15 Int. J. Mol. Sci. 2021, 22, 3235 11 of 18 Figure 9. Bayesian phylogeny (run under a WAG+G model of molecular evolution, with a MCMC analysis with 500,000 generations) of the mytimycin sequences identified in different Mytilida species. Posterior probability support values are shown next to each node. Colored dots indicate the type of disulfide array observed in each sequence. Figure 9. Bayesian phylogeny (run under a WAG+G model of molecular evolution, with a MCMC analysis with 500,000 generations) of the mytimycin sequences identified in different Mytilida species. Posterior probability support values are shown next to each node. Colored dots indicate the type of disulfide array observed in each sequence.