Analysis of CACTA transposases reveals intron loss as major factor influencing their exon/intron structure in monocotyledonous and eudicotyledonous hosts
Full text
RESEARCH Open Access Analysis of CACTA transposases reveals intron loss as major factor influencing their exon/intron structure in monocotyledonous and eudicotyledonous hosts Jan P Buchmann 1,4* , Ari Löytynoja 1 , Thomas Wicker 2 and Alan H Schulman 1,3 Abstract Background: CACTA elements are DNA transposons and are found in numerous organisms. Despite their low activity, several thousand copies can be identified in many genomes. CACTA elements transpose using a ‘cut-and-paste’ mechanism, which is facilitated by a DDE transposase. DDE transposases from CACTA elements contain, despite their conserved function, different exon numbers among various CACTA families. While earlier studies analyzed the ancestral history of the DDE transposases, no studies have examined exon loss and gain with a view of mechanisms that could drive the changes. Results: We analyzed 64 transposases from different CACTA families among monocotyledonous and eudicotyledonous host species. The annotation of the exon/intron boundaries showed a range from one to six exons. A robust multiple sequence alignment of the 64 transposases based on their protein sequences was created and used for phylogenetic analysis, which revealed eight different clades. We observed that the exon numbers in CACTA transposases are not specific for a host genome. We found that ancient CACTA lineages diverged before the divergence of monocotyledons and eudicotyledons. Most exon/intron boundaries were found in three distinct regions among all the transposases, grouping 63 conserved intron/exon boundaries. Conclusions: We propose a model for the ancestral CACTA transposase gene, which consists of four exons, that predates the divergence of the monocotyledons and eudicotyledons. Based on this model, we propose pathways of intron loss or gain to explain the observed variation in exon numbers. While intron loss appears to have prevailed, a putative case of intron gain was nevertheless observed. Keywords: Transposases, Intron loss, Molecular evolution, DNA transposons, Plants Background CACTA elements are DNA transposons found in genomes across the phylogenetic spectrum, from algae [1] to vascular plants [2-6] to animals [7,8]. The first CACTA element described at the molecular level was En-1 in Zea mays [2]; since then, they have been well documented in the grasses. Although CACTA elements usually do not account for the largegenomesizesfoundingrasses,CACTA families nevertheless can be highly abundant. In a few cases, however, including Tpo1 in Lolium perenne (ryegrass) and Caspar in the Triticeae, CACTA elements are known to have contributed considerably to the expansion of the genome size of their host [9-12]. Moreover, CACTAs can influence the evolution of the host genome in other ways [12]. In Glycine max (soybean), CACTA elements can affect flower color and capture host genes [13-16]. CACTA elements are sometimes associated with regulatory elements of genes, therefore possibly influencing gene expression [10,17]. Despite their prevalence and impact, evolutionary studies about * Correspondence: [email protected] 1 Institute of Biotechnology, Viikki Biocenter, University of Helsinki, PO Box 65, FIN-00014 Helsinki, Finland 4 Present address: Marie Bashir Institute for Infectious Diseases and Biosecurity, Charles Perkins Center, University of Sydney, Sydney NSW 2006, Australia Full list of author information is available at the end of the article © 2014 Buchmann et al.; licensee BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly credited. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated. Buchmann et al. Mobile DNA 2014, 5:24 http://www.mobilednajournal.com/content/5/1/24
CACTA elements, or DNA transposons in general, are scarce. The CACTA superfamily belongs to the Class II of transposable elements, proliferating by a ‘cut and paste’ mechanism. In contrast to Class I elements, which transpose via an RNA intermediate and therefore copy the original element, CACTAs transpose the original element itself. CACTA elements constitute approximately 2 to 5% of a grass genome [16,18]. However, only few active CACTA elements have been identified in plants [2-6,19]. In addition, only seven putative transcribed transposases have been identified in the Triticeae [10]. A full-length CACTA element consists of two terminal inverted repeats (TIRs) bordering two open reading frames(ORFs), one encoding a transposase and the other, called ORF2, a protein of unknown function. The first and last 5 bp of the TIRs consist of the highly conserved CACTA and TAGTG motifs, respectively, hence the name of the element. The function of the ORF2 protein has been determined in specific CACTA families to support excision and transposition [20]. However, the transposase is the key transposition enzyme. It binds to the TIR during excision, creating a 3-bp target site duplication (TSD) [21]. The catalytic center of the transposase is the acidic triad known as the ‘DDD/E’motif, which is highly conserved [22]. The presence of CACTA elements across the phylogenetic spectrum and the highly conserved catalytic core of their transposases indicate an ancient presence. Interestingly, the number of exons in transposases among CACTA transposons differs even among the grasses. Transposases in rice were found that have four exons [23], while studies in maize reported up to eleven exons for CACTA transposases [2,24]. In the recently sequenced grass Brachypodium distachyon, the exon number for transposases among CACTA superfamilies ranges from one to three. Therefore, the analysis of the exon/ intron configuration of CACTA transposases offers an excellent opportunity to study the evolutionary mechanisms of intron gain and loss in DNA transposons. In addition, analyzing exon number variations in such a highly conserved and ancient gene as the CACTA transposase can offer a perspective on the ‘intron-early’and ‘intron-late’models [25,26]. The goal of this study was to analyze the differences in exon numbers in CACTA transposases in monocotyledonous and eudicotyledonous plants and to identify an evolutionary mechanism to explain those differences. This was accomplished using phylogenetic and comparative analyses, which required a solid and robust multiple sequence alignment (MSA). We constructed such an MSA based on protein consensus sequences of 64 transposases from CACTA families annotated in ten monocotyledonous and eudicotyledonous species. Our phylogenetic analysis revealed that ancient CACTA lineages diverged before the divergence of the monocotyledons and eudicotyledons, supporting an intron-early model for CACTA transposases. The analysis of the MSA identified conserved exon/intron boundaries and putative intron gain among the transposases examined. Combining these analyses lead to a model for a putative ancient CACTA transposase, in which intron loss was the main mechanism shaping the exon/intron configurations of current transposases found in monocotyledonous and eudicotyledonous plants. Results We analyzed 64 autonomous CACTA transposases from ten different monocotyledonous and eudicotyledonous species. All analyzed transposases are derived from consensus sequences from distinctive CACTA families. Because families of transposable elements (TEs) differ from each other based on the 80-80-80 rule, they were considered orthologous [27]. Therefore, the name of the family, for example, Calvin, will indicate the consensus sequence of the transposase and not the consensus of the whole element. We refer to the plant in which a CACTA family and its transposase were annotated as its host. Except for transposases identified in B. distachyon,we searched the PTREP [28] and Repbase [29] databases for CACTA families with annotated transposases (see Materials and Methods). The selection was based on two criteria: i) the annotation had to clearly state ‘transposase’, that is annotations without ORFs described as transposases were omitted because CACTA elements have two ORFs, the transposase and ORF2; ii) the presence of two ORFs was expected, thereby avoiding selection of annotations having a predicted transposase that spans most of a consensus sequence, such as ATENSPM10 in Repbase, where the consensus is 8,272 bp and the predicted transposase covers positions 1,201 to 7,766. We selected nine transposases from Sorghum bicolor,eight transposases from Z. mays,fivetransposasesfrom Triticum aestivum,13fromOryza sativa,and11from B.distachyon (Additional file 1). This resulted in a total of 46 transposases from monocotyledonous hosts. For the eudicotyledonous dataset, we selected all transposases from eudicotyledonous hosts in Repbase fitting our criteria, totaling in eighteen elements: seven transposases from elements annotated in Arabidopsis thaliana,fivefromFragaria vesca,threefromVitis vinifera, and one each from Petunia hybrida, Malus domestica, and G. max (Additional file 1). Annotation of exon/intron boundaries on CACTA transposases For simplicity, the term ‘boundary’will indicate exon/ intron boundaries in this study. Except for transposases Buchmann et al. Mobile DNA 2014, 5:24 Page 2 of 15 http://www.mobilednajournal.com/content/5/1/24
in B. distachyon, boundaries were extracted from the respective PTREP and Repbase entries (Table 1, Material and Methods). The eleven Brachypodium distachyon transposases were derived from consensus sequences of the autonomous families in this genome [18]. We manually annotated the transposases and boundaries by aligning to the most similar BLASTX hit within the PTREP database. Additional alignments against transcription databases from rice and B. distachyon did not increase the quality of the boundary predictions, because transcriptome data is scarce for CACTA transposases. De novo gene prediction did not return significant results. Our final dataset consisted of 64 transposases with 86 annotated boundaries on the 40 transposases that contained more than one exon (Table 1). Out of the 64 annotated transposases, 24 contained only one exon and therefore no boundaries. On the remaining 40 transposases, we annotated between two and six exons (Additional file 1). The length of the transposases ranged from 552 amino acids (amino acids; PSL, 1 exon) to 4,785 amino acids (EnSpm4_Fves, 4 exons), and averaged 1,163 amino acids. The six transposases Isidor, Rufus, Sandro, Radon, Ivan, and Isaac were annotated on the 3’end of the corresponding CACTA consensus sequence (Additional file 1). Generation of a robust multiple sequence alignment using confidence scores Our phylogenetic and comparative analyses were based on an MSA derived from the selected 64 consensus transposase protein sequences. Due to the possibly ancient origin of certain CACTA transposases and their generally low activity, we assumed that some parts of sequences might be more evolutionarily diverged than others. In addition, the formation of consensus sequences can introduce weak regions into an MSA. A robust MSA is therefore crucial because errors or uncertainties can influence the downstream analysis. In addition, identifying weakly aligned regions or positionsinanMSAandthenremovingthemmayimprove downstream phylogenetic analysis [30]. GUIDANCE is a method to infer unreliable regions in an MSA and remove the potentially erroneous signal from subsequent analyses ([31]; Materials and Methods). The final MSA was 2,516 residues long and contained five unstable regions placed between positions 120 to 186, 196 to 251, 381 to 416, 728 to 766, and in the 3’ end, starting from position 1,665 (Additional file 2). GUIDANCE scores range from 0 (low confidence) to 1 (high confidence) and are calculated for single residues as well as for whole columns. Because there is no recommended confidence score for residues and columns in an MSA, a trade-off between sensitivity and specificity is required. High sensitivity (low cutoff value) retains as many columns as possible while high specificity (high cutoff value) keeps only columns of very high confidence. The default GUIDANCE cutoff of 0.93 removed 638 columns (approximately 25%) from the alignment, including the badly aligned regions and 34 annotated boundaries. However, GUIDANCE kept columns with only one residue, for example, most of the badly aligned 3’end. To retain as many boundaries as possible for the analysis we applied our own trimming: we removed columns containing only residues with scores below 0.804 (keeping boundaries) and columns with only one residue (not comparable and/or bad aligned). This approach removed 1,398 columns (approximately 44%): the badly aligned regions but only 13 annotated boundaries. This final MSA was 1,118 residues long and contained 73 annotated boundaries in 64 transposases (Figure 1). Because the first boundary is also the beginning of the first intron, introns were named in the 5’to 3’direction and designated as subscripts to the name of the transposase, for example, the first intron and boundary of transposase Baron is described as Baron 1 . We mapped conserved DDE motifs [22] onto the MSA, which were all in positions with high confidence values (Figure 1). This MSA was used for all further analysis. Exon numbers in CACTA transposases are not specific to a host genome RAxML [32] was used to calculate the phylogenetic tree (Figure 2). A maximum likelihood (ML) tree was generated based on 200 distinct, randomized, maximum parsimony trees and its robustness assessed by using 1,000 bootstrap replicates and by testing the influence of several outgroups (Additional file 3, Material and Methods). The resulting tree shows the relation between individual transposases but not their evolution over time; that is the branch lengths do not indicate the time when transposases diverged from each other but how close they are on the molecular level (Figure 2). We identified eight clades, designated αto θ(Figure 2). Crucially, the transposases grouped primarily by their exon numbers rather than by their hosts and the analysis of the clusters found no host-specific exon numbers for CACTA transposases (Figure 2). Ancient CACTA lineages diverged before the divergence of monocotyledons and eudicotyledons We identified three clades in which monocotyledonous and eudicotyledonous transposases clustered together. EnSpm2_Gmax from soybean grouped in Clade αwith transposases from several monocotyledonous hosts, analogous to EnSpm3_Fves and EnSpm4_Fves from strawberry in Clade ζ. Clade δgrouped transposases from strawberry, apple, and several grasses. The other clades contained only transposases from either eudicotyledonous or Buchmann et al. Mobile DNA 2014, 5:24 Page 3 of 15 http://www.mobilednajournal.com/content/5/1/24
Table 1 Exon/intron boundaries of the 34 analyzed CACTA transposases with more than one exon. 12 3 4 5 EnSpm12_Fves 462 | 564 G C 718 | 771 I EnSpm10_Fves 826 | 846 Joey 842 | 893 II Janus 837 | 894 II F 846 | 894 II G 847 | 894 II Norman 879 | 921 III En1 879 | 925 III Alfred 885 | 925 III H 838 | 972 EnSpm3_Vvin 821 | 894 II 856 | 925 III EnSpm8_Sbic 754 | 783 I 976 | 0 Storm 827 | 782 I 951 | 0 Sherman 831 | 782 I 954 | 0 J 495 | 521 750 | 885 EnSpm2_Mdom 755 | 782 I 886 | 893 II Baldur 731 | 782 I 837 | 895 II I 834 | 782 I 954 | 910 Isidor 857 | 894 II 892 | 920 III Radon 841 | 894 II 877 | 921 III Rufus 851 | 894 II 887 | 921 III EnSpm13_Vvin 821 | 894 II 856 | 925 III EnSpm5_Vvin 824 | 894 II 859 | 925 III Isaac 861 | 894 II 900 | 925 III Sandro 744 | 782 I 851 | 928 III Balduin 850 | 895 II 890 | 930 III DOPPIA 843 | 894 II 890 | 936 III K 744 | 7,82I 850 | 936 III Horace 712 | 711 981 | 1,054 EnSpm4_Fves 812 | 0 992 | 0 1,244 | 0 EnSpm3_Fves 681 | 0 770 | 781 I 919 | 0 Seamus 730 | 782 I 833 | 892 II 878 | 925 III Dario 726 | 711 842 | 839 895 | 890 III Aron 851 | 833 899 | 879 II 1,013 | 1,060 Korbin 510 | 567 G 718 | 782 I 814 | 894 II 853 | 0 Chester 520 | 563 G 728 | 777 I 823 | 889 II 858 | 920 III Baron 522 | 568 G 730 | 781 I 825 | 893 II 861 | 925 III EnSpm8_Fves 158 | 163 830 | 893 II 975 | 0 1,219 | 0 1,500 | 0 ATENSPM6_Athal 802 | 809 918 | 922 III 978 | 981 1,011 | 1,012 1,141 | 0 The positions are relative to the beginning of the transcription start and given as follows: On the protein sequence | on the trimmed multiple sequence alignment (MSA). 0 and numbers in italic indicate boundaries with GUIDANCE scores below 0.804 and removed in the final MSA. Superscripts indicate Regions I to III and G cluster, respectively (Figure 1). Buchmann et al. Mobile DNA 2014, 5:24 Page 4 of 15 http://www.mobilednajournal.com/content/5/1/24
monocotyledonous hosts (Figure 2). Despite the long evolutionary time separating monocotyledonous and eudicotyledonous hosts, the presence of mixed clades and the close relation of clades with only monocotyledonous or eudicotyledonous hosts suggests that the CACTA transposase phylogeny rather than the host phylogeny is primary, that is that the main transposase branches diverged already before the divergence of monocotyledons and eudicotyledons. Indeed, a closer look at the phylogenetic tree revealed that transposases within clades tend to have thesamenumberofexons(Figure2). The majority of CACTA transposase boundaries are found in three regions on the MSA To analyze the evolution of exon/intron arrangements in CACTA transposases, we compared the boundaries from the 33 transposases containing 73 introns that were not removed in the trimming process (Table 1, Figure 1). We identified 3 regions, labeled I to III, in the MSA, which contain 63 out of the 73 boundaries (Figure 1). Outside those regions, we identified eight boundaries inside the DDE motif, four boundaries between Regions I and II, one boundary between Regions II and III and five boundaries downstream of Region III. Most boundaries are close to each other but not in the same position on the alignment. This can be due to small errors introduced by calculating the MSA or consensus sequences. Therefore, we analyzed the distances between boundaries to identify which were shared among transposases. We analyzed the boundaries by clustering them based on their positions on the MSA. We set the maximal distance between boundaries still considered to be in the Figure 1 Multiple sequence alignment based on protein sequences of the 64 analyzed CACTA transposases. Colored boxes indicate amino acids, gray boxes indicate residues with a GUIDANCE score below 0.804, and white boxes indicate gaps in the multiple sequence alignment (MSA). The plot below the MSA shows GUIDANCE scores for the corresponding position in the MSA. Columns with a score below 0.804 are indicated in light blue while columns with a score of 0.804 and above in dark blue. Positions relative to the MSA and corresponding GUIDANCE score are shown between the MSA and the plot. Highly conserved DDE transposase motifs as described in [22] are depicted on top. In the phylogenetic tree, colors indicate the host as shown in the legend. Major clades are depicted αto θ. Exon/intron boundaries are depicted as blue circles if their GUIDANCE score was above 0.804 and red otherwise. The number in the boundary indicates the boundary number on the corresponding transposase. Regions I to III are indicated by dashed lines and corresponding roman capitals. Positions of putative intron gain are depicted as described in the legend. Buchmann et al. Mobile DNA 2014, 5:24 Page 5 of 15 http://www.mobilednajournal.com/content/5/1/24
same region to 16 residues, which is half the length of the shortest intron annotated (33 amino acids in ATENSPM_Athal 3 ). Boundaries that were closer than 16 residues to each other were grouped together. No boundaries within a region were further than 16 residues apart (Tables 2, 3, Additional files 4, 5, 6). The distances between the closest boundaries of Regions I and II is 98 residues (Additional file 7), but 30 residues between Region II and III (Additional file 7). The closest boundary upstream of Region I is 60 residues away, whereas the closest boundary downstream of Region III is 36 residues away. This clustering confirmed the previously identified regions as clearly distinct. The four boundaries EnSpm10_Fves 1 , Dario 2 ,Aron 1 , and ATENSPM6_Athal 1 between Region I Figure 2 Majority-rule based phylogram of the 64 analyzed CACTA transposases. The phylogenetic tree is the same as in Figure 1. Bootstrap values represent the percentage out of 1,000 bootstrap replicates. Only bootstraps below 100% are indicated. Transposase hosts are colored as indicated in the legend. Numbers in parentheses indicate the number of exons. Clades are indicated by dashed lines and labeled αto θ. Table 2 Distances between exon/intron boundaries within Region I Baldur 1 Baron 2 1 Baron 2 C 1 11 10 C 1 Chester 2 5 4 6 Chester 2 EnSpm2_Mdom 1 0 1 11 5 EnSpm2_Mdom 1 EnSpm3_Fves 2 1 0 10 4 1 EnSpm3_Fves 2 I 1 01115 0 1 I 1 K 1 0 1 11 5 0 1 0 K 1 Korbin 2 0 1 11 5 0 1 0 0 Korbin 2 Sandro 1 0 1 11 5 0 1 0 0 0 Sandro 1 Seamus 1 0 1 11 5 0 1 0 0 0 0 Seamus 1 Sherman 1 0 1 11 5 0 1 0 0 0 0 0 Sherman 1 Storm 1 0 1 11 5 0 1 0 0 0 0 0 0 Distances between exon/intron boundaries in the MSA within Region I (depicted in Figure 1). The distances are given in residues in the alignment. Buchmann et al. Mobile DNA 2014, 5:24 Page 6 of 15 http://www.mobilednajournal.com/content/5/1/24
and II, as well as I 2 between Region II and III could not be clustered in those Regions. We identified only one additional cluster containing four boundaries outside Regions I to III. It groups the first introns from all members of Clade γand was therefore named Region G. Based on these analyses of distances between all boundaries, we established that Regions I to III and G in the MSA were clearly separated from each other as well as from all other boundaries. Given the distinctness of the four boundary regions, we examined if the boundaries themselves were conserved among the analyzed transposases. Boundaries in Regions I to III are conserved among most transposases while Region G represents putative intron gain Due to the proximity of boundaries in Regions I to III and their clear separation from other boundaries, we established that boundaries within a region are shared between the different transposases. The clustering of boundaries within Regions I to III indicates that the boundaries are conserved among the analyzed transposases. This is supported by the phylogenetic tree, in which purely monocotyledonous or eudicotyledonous clades share boundaries (Figure 1). Boundaries in Region I are on, or close to, the position of the conserved E from the DDE motif, supporting the claim that Region I represents conserved boundaries among the transposases (Figure 1). Therefore, we considered the 63 boundaries in Regions I to III as conserved within each region. All transposases in Clade γ share their first introns with a maximum distance of five residues (Figure 1, Table 4). This is a unique cluster in the whole tree, indicating intron gain since all members of Clade γshare this intron but none of its ancestor nodes and transposases in other clades. Only two boundaries from a monocotyledonous host are found outside Regions I to III We identified 17 boundaries outside Regions I to III (Figure 1). Only J 1 and H1 are from a monocotyledonous host, whereas the remaining 15 boundaries were annotated in transposases from eudicotyledonous hosts. Boundaries I 1 and ATENSPM6 1,2,3 cannot be clustered and therefore were not further characterized. The transposases Horace, Dario, and Aron have three separate boundaries which are not farther apart than six residues: Horace 1 and Dario 1 , Daron 2 and Aron 1 , Horace 2 and Aron 3 . While this appears as another case of intron gain, their relation in the phylogenetic tree is not properly resolved and does not support this interpretation. Our analysis of the boundaries identified 63 conserved boundaries and 4 cases of putative intron gain in Region G. Most conserved introns were identified in transposases from monocotyledonous hosts. In contrast, all unique boundaries except two were identified in eudicotyledonous hosts. We decided to combine the results of the phylogenetic and boundary analyses to develop a model to understand how the observed exon/intron configuration evolved. Defining consensus exon numbers for each phylogenetic clade A comparison of the phylogenetic tree and the conserved boundaries revealed a high consistency between clades and boundary positions. Based on the majority of exons per clade, we constructed a loose consensus to represent the exon number for transposases in the corresponding clade. For example, Clade ζgroups together seven transposases of which four, the majority, have two exons. Therefore, a representative transposase from Clade ζhas two exons and one consensus boundary. We used this approach for each clade (Figure 3). Our approach resulted in following exon numbers for representative transposases: one exon for Clade α;Cladesβ,δ,andθthree exons each; Clade ηfour exons; Clade γfive exons. Designating consensus exon numbers for each clade simplified further the analysis to develop a model for the loss and gain of boundaries in CACTA transposases. A model for loss and gain of exon/intron boundaries in CACTA transposases Because it had the largest number of confirmed exons, we compared all consensus boundaries to Clade γ(Figure 3). Clade αhas no annotated introns. The second, third, and fourth intron of Clade γcan be found throughout the phylogenetic tree, whereby the third intron of Clade γis the most conserved, followed by its fourth and second intron. The fourth intron of Clade γis found among Clades β,θ,ι,and in Isaac. The third intron is missing in the Clades EnSpm8, δ,andθ, but otherwise is found in all clades containing introns. The second intron of Clade γis present in Clades δ, EnSpm8, and η. This comparison indicates that CACTA transposases were as a whole losing rather than gaining introns. However, Clades γand ζhave introns that are not found in other clades (Figure 3), the first intron in Clade γ representing an intron gain. The unique introns in Clade ζ cannot be classified as losses or gains because the phylogenetic tree does not allow a definitive classification. We propose that the consensus transposase in Clade γ represents the most likely exon/intron configuration of an ancient transposase, containing at least four exons and three introns (Figure 3). The three boundaries correspond to those identified in Regions I to III in the MSA (Figures 1, 3). Using the putative ancestor model transposase, we can infer the emergence of the known transposases through intron loss and gain (Figure 3). Discussion In sum, we analyzed 64 CACTA transposases from 11 monocotyledonous and eudicotyledonous hosts. Our Buchmann et al. Mobile DNA 2014, 5:24 Page 7 of 15 http://www.mobilednajournal.com/content/5/1/24
Table 3 Distances between exon/intron boundaries within Region III Alfred 1 Balduin 2 5 Balduin 2 Baron 4 0 5 Baron 4 Chester 4 5 10 5 Chester 4 En1 1 0 5 0 5 En1 1 EnSpm13_Vvin 2 0 5 0 5 0 EnSpm13_Vvin 2 EnSpm3_Vvin 2 0 5 0 5 0 0 EnSpm3_Vvin 2 EnSpm5_Vvin 2 0 5 0 5 0 0 0 EnSpm5_Vvin 2 Isaac 2 0 5 0 5 0 0 0 0 Isaac 2 Isidor 2 5 10 5 0 5 5 5 5 5 Isidor 2 Norman 1 4 9 4 1 4 4 4 4 4 1 Norman 1 Radon 2 4 9 4 1 4 4 4 4 4 1 0 Radon 2 Rufus 2 4 9 4 1 4 4 4 4 4 1 0 0 Rufus 2 Sandro 2 3 2 3 8 3 3 3 3 3 8 7 7 7 Sandro 2 Seamus 3 05 05 00 0 0 054 443 Distances between exon/intron boundaries in the MSA within Region III (depicted in Figure 1). The distances are given in residues in the alignment. Buchmann et al. Mobile DNA 2014, 5:24 Page 8 of 15 http://www.mobilednajournal.com/content/5/1/24
phylogenetic analysis indicates divergence of ancient CACTA lineages already before the divergence of the monocotyledons and eudicotyledons. The analysis of 73 boundaries across 33 transposases with more than one exon identified 55 conserved exon/intron boundaries and allowed us to reconstruct the exon/intron configuration of a CACTA transposase representing the ancestral state before the divergence of monocotyledonous and eudicotyledonous plants. The model consists of at least four exons. We propose a mechanism for the evolution of the extant CACTA transposases in which they were shaped mainly by intron loss, although one case of putative intron gain was found. Potential for greater regulation of CACTA elements in eudicotyledons Studies of the PElement in Drosophila and Ac/Ds in maize have shown that alternative splicing can regulate tissue-specific transposition of elements. For example, the Pelement retains its third intron in somatic cells, inhibiting transposition [33,34]. Should this occur with CATCA transposases as well, our data suggests that elements in dicotyledonous hosts have more possibilities for regulation. Interestingly, most non-clustered boundaries and the putative intron gain cluster were found in transposases from dicotyledonous hosts, whereas the majority of boundaries in Regions I to III were found in transposases from monocotyledonous hosts. The number of transposable elements in eudicotyledonous genomes is generally lower than in monocotyledonous genomes, consistent with a tighter control of transposable elements in eudicotyledonous hosts. Therefore, the large number of unique boundaries found outside Regions I to III could be associated with more control of expression of CACTA elements in eudicotyledons than in monocotyledons. Differences in intron gain and loss among TE transposases Previously, intron gain and loss in transposases of DNA transposable elements was studied for Mariner-like elements in flowering plants [35]. In that study, degenerate primers were used to extract fragments of DDE transposases from 54 plant species for phylogenetic analysis. The results were consistent with vertical transmission Table 4 Distances between exon/intron boundaries within Cluster G Baron 1 Chester 1 5 Chester 1 EnSpm12_Fves 1 4 1 EnSpm12_Fves 1 Korbin 1 14 3 Distances between exon/intron boundaries in the MSA within Cluster G (depicted in Figure 1). The distances are given in residues in the alignment. Figure 3 Model for the loss and gain of introns in CACTA transposases. Simplified phylogenetic tree based on the consensus exon numbers per clade as described in the text. Below the tree the putative ancestor transposase with four exons is depicted. Exons are depicted as gray rectangles with introns as colored lines. Blue, red and green depict introns conserved in Regions I to III, G indicates cluster G with the putative intron gain. Conserved introns share the same color band. Intron loss is depicted by its corresponding color and circled −, intron gain by an encircled +. Gray balloons indicate how the observed configuration arose from the putative ancestor. Buchmann et al. Mobile DNA 2014, 5:24 Page 9 of 15 http://www.mobilednajournal.com/content/5/1/24