Full text
Evolution of the Metazoan Mitochondrial Replicase Marcos T. Oliveira 1,2 , Jani Haukka 1 , and Laurie S. Kaguni 1,3, * 1 Institute of Biosciences and Medical Technology, University of Tampere, Finland 2 Departamento de Tecnologia, Faculdade de Cie ˆncias Agra ´rias e Veterina ´rias, Universidade Estadual Paulista “Ju ´lio de Mesquita Filho,” Jaboticabal, SP, Brazil 3 Department of Biochemistry and Molecular Biology and Center for Mitochondrial Science and Medicine, Michigan State University *Corresponding author: E-mail: [email protected] Accepted: February 26, 2015 Abstract The large number of complete mitochondrial DNA (mtDNA) sequences available for metazoan species makes it a good system for studying genome diversity, although little is known about the mechanisms that promote and/or are correlated with the evolution of this organellar genome. By investigating the molecular evolutionary history of the catalytic and accessory subunits of the mtDNA polymerase, pol g, we sought to develop mechanistic insight into its function that might impact genome structure by exploring the relationships between DNA replication and animal mitochondrial genome diversity. We identified three evolutionary patterns among metazoan pol gs. First, a trend toward stabilization of both sequence and structure occurred in vertebrates, with both subunits evolving distinctly from those of other animal groups, and acquiring at least four novel structural elements, the most important of which is the HLH-3b(helix-loop-helix, 3 b-sheets) domain that allows the accessory subunit to homodimerize. Second, both subunits of arthropods and tunicates have becomeshorter and evolved approximately twice asrapidly as their vertebrate homologs. And third, nematodes have lost the gene for the accessory subunit, which was accompanied by the loss of its interacting domain in the catalytic subunit of pol g, and they show the highest rate of molecular evolution among all animal taxa. These findings correlate well with the mtDNA genomic features of each group describedabove, and with their modes ofDNA replication, although a substantive amount of biochemical work is needed to draw conclusive links regarding the latter. Describing the parallels between evolution of pol gand metazoan mtDNA architecture may also help in understanding the processes that lead to mitochondrial dysfunction and to human disease-related phenotypes. Key words: mitochondria, mitochondrial DNA replication, structural evolution, mitochondrial replicase, pol g. Introduction Mitochondrial DNA (mtDNA) replication is accomplished by the sole DNA polymerase found in animal mitochondria, DNA polymerase g(pol g) (reviewed in Kaguni 2004), which functions as part of a larger replication machinery called the mtDNA replisome. The identity of all of the components of the mtDNA replisome is still unknown, but biochemical studies (Korhonen et al. 2003,2004;Oliveira and Kaguni 2010, 2011) have shown that a group of mitochondrial proteins can interact functionally to promote DNA synthesis in vitro, forming the minimal mtDNA replisome: The mtDNA helicase, also known as Twinkle in humans, unwinds the duplex DNA at the replication fork; pol gcatalyzes nascent DNA synthesis on both the leading and lagging DNA strands; and the mitochondrial single-stranded DNA-binding protein (mtSSB) coordinates their functions while binding and stabilizing the single-stranded DNA template (Korhonen et al. 2004;Oliveira and Kaguni 2011). This scenario, as in most DNA replication events, requires the presence of short RNA molecules for priming of DNA synthesis by the DNA polymerase, which in animal mitochondria might be achieved by the action of the mitochondrial RNA polymerase (Wanrooij et al. 2008;Fuste et al. 2010;Reyes et al. 2013) and/or the newly identified primase PrimPol (Garcia-Gomez et al. 2013). Animal pol gis a hetero-oligomeric enzyme, in most known cases: The catalytic core, pol g-a(also known as POLG or PolGA), contains both the 50–30DNA polymerase and 30–50 exonuclease activities of the holoenzyme, whereas the accessory subunit, pol g-b(or POLG2, PolGB), serves as a processivity factor, enhancing the interactions between the holoenzyme and the DNA substrate (Lewis et al. 1996; Wang et al. 1997;Carrodeguas et al. 1999;Lim et al. 1999). The catalytic core is a member of the family A DNA polymerase group, to which bacterial DNA polymerase I and the catalytic core of bacteriophage T7 DNA polymerase (gp5) also belong (reviewed in Kaguni 2004); pol g-ashares GBE ßThe Author(s) 2015. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited. Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 943 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
substantial structural and functional properties with both Pol I and T7 Pol. Notably, Pol I is not the replicative DNA polymerase in bacteria. It is a single-subunit enzyme that lacks the high fidelity and processivity of the holoenzyme form of T7 Pol, which acquires these properties as a result of the association of the catalytic core with the bacterial protein thioredoxin that serves as its accessory subunit. Although the presence of pol g as the mitochondrial replicase appears to be conserved among the metazoans (and other eukaryotes), important variations in its structure have been described that may impact the mode in which mtDNA is replicated. It has long been known that pol g of Drosophila melanogaster, one of the best studied insect model organisms, is a heterodimer comprising one catalytic subunit and a single accessory subunit (Wernette and Kaguni 1986;Olson et al. 1995;Wang and Kaguni 1999). On the other hand, the human and mouse holoenzymes have a heterotrimeric conformation, consisting of one pol g-aand a dimeric pol g-b(Carrodeguas et al. 2001;Yakubovskaya et al. 2006;Lee et al. 2009). Interestingly, the mtDNA of the nematode Caenorhabditis elegans appears to be replicated by a machinery containing only a single subunit pol g(only pol g-a) (Bratic et al. 2009;Addo et al. 2010), which resembles the catalytic core in other eukaryotes, such as that of the yeast Saccharomyces cerevisiae (Foury 1989). The accessory subunit of pol g, which indeed appears to be present only in Metazoa, has a remarkable evolutionary origin, because amino acid sequence alignments, phylogenetic inferences, and general protein structure demonstrate its homology to class II aminoacyl-tRNA synthetases (Fan et al. 1999,2006;Carrodeguas et al. 2001;Wolf and Koonin 2001). We sought to explore the sequence and structural diversity of pol gin the animal kingdom, taking advantage of the current increase in nuclear genomes and transcriptomes for which complete sequences are available in public databases. We retrieved as many animal pol g-aand -bgene sequences as are available and performed in silico analyses to infer their molecular evolutionary history, taking into account the substantial biochemical and structural data reported by our group and others. Here we report the oligomeric plasticity of animal pol g, the finding of new structural elements, and the distinct rates of molecular evolution for different taxa, which may reflect differences in the fundamental mechanisms of mtDNA replication. We discuss our findings in the context of mitochondrial genome diversity, structure, replication, and evolution. Materials and Methods Searches for Animal pol g-aand -bHomologs and Multiple Sequence Alignments TBLASTN searches (Altschul et al. 1990) in the NCBI (National Center for Biotechnology Information) nonredundant sequence database were performed using the translated mRNA reference sequences from Homo sapiens (pol g-a, NM_001126131.1; pol g-b, NM_007215.3) and D. melanogaster (pol g-a, NM_057473.3; pol g-b, FJ635829.1) as queries. To retrieve sequences from Porifera, Placozoa, Cnidaria, Mollusca, and Hemichordata species, complementary HMMR3 BLAST (Basic Local Alignment Search Tool) searches (Eddy 2011) were performed (http://toolkit.lmb.unimuenchen.de/hmmer3, last accessed August 2014), followed by BLAST searches against the Ensembl Metazoa database (http://metazoa.ensembl.org, last accessed August 2014), allowing the inclusion of missing exons. Most sequences from Nematoda species were retrieved from genomic scaffolds with no gene models deposited in the 959 Nematode Genomes databank (http://www.nematodes.org/nematodegenomes, last accessed August 2014), followed by manual processing of the exons. A Python script utilizing functions of the Biopython library (Cock et al. 2009) was developed to extract the coding sequences and other specific information from the original files retrieved from the diverse databanks. The script can be provided by the authors upon request. Highly divergent sequences were tested for true orthology by using them as queries for BLAST searches against the human genome database and by predicting possible mitochondrial localization using the TargetP service (Emanuelsson et al. 2007). The complete set of sequences retrieved is shown in supplementary table S1,Supplementary Material online. Because of the overrepresentation by mammalian and Drosophila sequences, this data set was reduced before the alignments were performed to diminish bias in interpretation of the results. The multiple amino acid sequence alignments were performed with the software MAFFT (Katoh and Toh 2008), using the G-INSi and E-INSi algorithms for the pol g-a and -bsequences, respectively, and are shown in supplementary figures S1 and S2,Supplementary Material online. Phylogenetic Inferences To create the input files for the phylogenetic inferences, the software PAL2NAL (Suyama et al. 2006) was used to convert the amino acid sequence alignments into codon-based nucleotide alignments, which were then converted to NEXUS format. The phylogenetic trees were inferred using the Bayesian algorithm built in the software MrBayes, version 3.2.2 (Ronquist et al. 2012). The consensus tree, run with 200,000 cycles for pol g-aand 1 million cycles for pol g-b, was set to the 50-majority rule, gamma variation was expected among the sites and Generalized Time Reversible model was used; other parameters were kept as default. Modeling of Protein Structure The structure of the disordered regions in the human pol g-a crystallography data (PDB accession number 3IKM, chain A) was predicted with the software I-TASSER (Bazzoli et al. 2011), using default parameters. I-TASSER was also used Oliveira et al. GBE 944 Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
with default parameters to model the whole structure of pol g-afrom C. elegans (using the human pol gcrystal structure 3IKM:A), and of pol g-bfrom Strongylocentrotus purpuratus, Ciona intestinalis, and Trichoplax adhaerens (using the human pol g-bcrystal structure 2G4C:A). Drosophila melanogaster pol g-bstructure was modeled with the software MODELLER, version9.12(Sali and Blundell 1993), using the 2G4C:A file as template and the multiple sequence alignment (MSA) shown in supplementary figure S2,Supplementary Material online, as a parameter. The selected models were evaluated by the Zscores, the discrete optimized protein energy values, and the residue error plot (SwissProt website, http://swissmodel. expasy.org/workspace/index.php?func=tools_structureassessment1, last accessed August 2014). Structures and models were analyzed and figures were produced using Pymol (www.pymol.org, last accessed August 2014). Results Our searches of public genomic sequence databases found nonredundant sequences for pol g-aand -bof 62 and 52 animal species, respectively (supplementary table S1, Supplementary Material online). The resulting data set is overrepresented by sequences from insect and vertebrate (especially mammalian) species due to the bias in the databases. Animal groups in the Protostomia clade, other than Arthropoda and Nematoda, were most often absent, except for one pol g-asequence from Mollusca. Nonetheless, we obtained sequences from key species of basal groups of Metazoa, such as Porifera, Placozoa and Cnidaria, and basal and sister groups of Chordata, such as Tunicata, Cephalochordata, Hemichordata and Echinodermata, which are crucial for the analyses described below. We excluded several sequences of mammalian and Drosophila species from our pol g-aand -bdata sets to balance the taxa representation, and also a few of the BLAST results because they contained only partial gene sequences (see supplementary table S1,Supplementary Material online). Supplementary figure S1,Supplementary Material online, shows the alignment of 43 pol g-aaminoacidsequences plus the outgroup (the sequence from S. cerevisiae pol g-a, accession number NM_001183750.1). The alignment of 35 pol g-bamino acid sequences plus the outgroup (the glycyl-tRNA synthetase sequence from the bacterium Thermus thermophilus, accession number AJ222643.1) is shown in supplementary figure S2,Supplementary Material online. The number of pol g-bsequences retrieved was lower than that for pol g-amainly because of the absence of the pol g-b gene in the genome of nematode species, as discussed below. In addition, we were unable to find complete pol g-bgene sequences for the poriferan Amphimedon queenslandica,the mollusk Crassostrea gigas, and the crustacean Daphnia pulex. Considering that the pol g-asequences from these species do have a potential to form an accessory-interacting determinant (AID) structure (see below), our failure to find the corresponding pol g-bmight be because of the current low coverage of their genome/transcriptome sequences. A summary of our most interesting findings is presented in figure 1, along with schematics of the catalytic and accessory subunit polypeptides. We identified specific features for all sequences retrieved to indicate that our data set most likely consists of true orthologs of the catalytic and accessory subunits of the mitochondrial replicase, and not random genes coding for other family A DNA polymerases or aminoacyl-tRNA synthetases, respectively: 1) Conserved active site motifs in the exonuclease (Exo I–III) and polymerase (Pol A–C) domains of pol ga(reviewed in Kaguni 2004)(supplementary fig. S1, Supplementary Material online) that are shared among family A DNA polymerases; 2) conserved pol g-a-specific sequences in the spacer region and polymerase domains of pol g-a(Lewis et al. 1996;Lecrenier et al. 1997;Kaguni 2004); 3) conserved sequence for the AID subdomain of pol g-a(except in nematodes), which is a pol g-a-specific feature of metazoans (reviewed in Kaguni 2004) (see below); 4) conserved hydrophobic residues in the C-terminal region of pol g-bthat are most relevant for the interactions with the pol g-aAID, and are found almost exclusively in pol g-bs(Fan et al. 1999; Carrodeguas et al. 2001;Kaguni 2004;Lee et al. 2009); and 5) absence in pol g-bof the residues required for dimerization of class II aminoacyl-tRNA synthetase (Logan et al. 1995;Arnez et al. 1999). Distinct Rates of Evolutionary Changes in pol g-aand -b Sequences Phylogenetic inferences with the pol g-anucleotide sequences using Bayesian analysis (fig. 2) reproduced moderately well the currently accepted relationships among animal taxa (Philippe et al. 2009), with few exceptions, and provided important findings regarding the evolution of the gene. First, it is clear that the sequences from nematodes have evolved at a higher rate than those from insect species, which in turn have accumulated more substitutions than the sequences from vertebrates. Second, Tunicata, which is represented only by the sequence from Ci. intestinalis (Ascidiacea: Cionidae), was grouped within the pol g-asequences from Arthropoda. Finally, the branches for the pol g-asequences from Deuterostomia (vertebrates and sister groups), excluding Ci. intestinalis, and from the mollusk Cr. gigas were as short as those for the sequences from basal animal groups, such as Porifera (A. queenslandica), Placozoa (T. adhaerens), and Cnidaria (Nematostella vectensis). In combination, this may indicate that most deuterostome pol g-as retained more ancestral characters, whereas pol g-asequences from nematodes, arthropods (especially insects), and tunicates have diverged considerably from the original enzyme. Bayesian phylogenetic analysis using pol g-bnucleotide sequences (fig. 3) also shows that the insect proteins have Evolution of Animal pol gGBE Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 945 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
evolved at a much higher rate than vertebrates. The rate of substitutions observed for the sequences from most Deuterostomia species again matches that for the sequences from basal animal groups, suggesting that pol g-balso retained more ancestral characters for these groups. Moreover, the sequence from the tunicate Ci. intestinalis again grouped with those from Arthropoda, indicating a remarkable resemblance and their significant divergence from FIG.1.—Schematics of animal (vertebrate) pol g-aand -bsequence and structure. (A) Representation of the amino acid sequence of the pol g-aand -b polypeptides, showing the protein domains, subdomains, conserved motifs, and new motifs proposed in this work. For pol g-a: NTD, the N-terminal domain for which no functional data are available; Exo, the exonuclease domain responsible for the 30–50exonuclease activity that edits misincorporated nucleotides and increases the fidelity of DNA synthesis several-hundred fold; AID, the accessory-interacting determinant subdomain that provides the primary contacts between the catalytic core and the accessory subunit; IP, the intrinsic processivity subdomain that contributes to the ability of the catalytic core to polymerize multiple nucleotides in a single enzyme binding cycle; Pol, the DNA polymerase domain responsible for the 50–30DNA polymerase activity; T, the bipartite thumb subdomain (according to Lee et al. 2009) that contributes to template–primer DNA binding; boxes I–III, the conserved motifs Exo I–III that form the 30–50exonuclease active site; boxes A–C, the conserved motifs Pol A–C that comprise the 50–30polymerase active site (for more detailed descriptions of each of these features, see Kaguni [2004] and Euro et al. [2011]; blue box, the new vertebrate Exo motif; red box, the region absent in the nematode species; and orange box, the new vertebrate IP motif [see the text for details]. For pol g-b, the domain designations follow that described by [Fan and Kaguni 2001;Fan et al. 2006]. Green box, the vertebrate dimerization [HLH-3b] interface; cyan box, the vertebrate M loop [see text for details]; and brown boxes, the conserved hydrophobic regions for pol g-ainteraction [according to Lee et al. 2009]). All structural elements are represented to scale, except motifs Exo I–III and Pol A–C. (B) Surface representation of the crystal structure of the human pol gapo-holoenzyme (3IKM; Lee et al. 2009), highlighting the interactions among the subunits and pol g-afunctional domains. Structures are colored as shown in (A); dark gray, distal pol g-b.(C) Surface representation of the crystal structure of the human pol gapo-holoenzyme (3IKM; Lee et al. 2009), highlighting the motifs and domains identified in this study. The structures are colored asshownin(A)and(B). Oliveira et al. GBE 946 Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
the ancestral accessory subunit polypeptide. We were unable to find any pol g-bcoding sequence from species of nematodes, as discussed below. The fast evolutionary rates observed for the pol g-aand -b genes in insects, tunicates, and nematodes appear to follow a general tendency described for the majority of nuclearencoded genes from these groups (Li 1997;Mitreva et al. 2005;Drosophila 12 Genomes Consortium et al. 2007; Denoeud et al. 2010). Interestingly, our preliminary analyses of the genes encoding mtSSB and mtDNA helicase show substitution rates that differ from those of the pol g-aand -b genes among animal groups (Oliveira MT, Haukka J, Kaguni LS, unpublished data). Thus, it appears that the mtSSB and mtDNA helicase genes may have atypical evolutionary constraints, but this requires further validation. New Motif in the Exonuclease Domain of Vertebrate pol g-a: Implications for DNA Binding The MSA identified an insertion of approximately 17 amino acid residues between motifs Exo II and III of the pol g-a exonuclease domain (H320–A336 in humans; A308–A309 in D. melanogaster) that is present in all species of Vertebrata, and possibly other Deuterostomia species (except Ci. intestinalis;fig. 4A). Similarly long regions in the same position are present in Da. pulex (Arthropoda: Crustacea), Oscheius tipulae (Nematoda: Rhabditida), Cr. gigas (Mollusca: Bivalvia), and A. queenslandica (Porifera: Demospongiae), but these have little sequence conservation and may represent independent insertion events into the pol g-agenes of these species. The insertion is several residues downstream of the orienter, a structural module that is highly conserved within eukaryotes, and which has been reported to coordinate the balance between the polymerase and exonuclease functions (Szczepanowska and Foury 2010). Locating the new vertebrate Exo motif in the crystal structure of the human pol g-a(PDB: 3IKM; Lee et al. 2009) revealed that the insertion is part of a disordered region (K319–S344) for which we have no structural information, but which may indicate a flexible domain involved in transient interactions. Modeling the missing residues in the human pol g-astructure based upon the secondary structure prediction algorithm built into the I-TASSER server resulted in two short alpha helices connected by a short loop, which orient this element toward the DNA binding cleft of the enzyme (fig. 4B). In particular, residues H320, K327, K331, and K335 are in close proximity to the minor groove of the primed DNA template modeled onto the putative DNA binding cleft of pol g-a(Euro et al. 2011). We postulate that the FIG.2.—Bayesian phylogenetic inference for animal pol g-anucleotide sequences. The outgroup sequence used was the pol g-a(mip-1) from the yeast S. cerevisiae (accession number NM_001183750.1). The 50% majority-rule consensus tree was inferred using MrBayes 3.2, as described under Materials and Methods. Bayesian posterior probability values are indicated for almost all nodes. The scale bar indicates substitutions per site. Evolution of Animal pol gGBE Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 947 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
new Exo motif may represent a module that serves to enhance DNA binding in vertebrate pol gs. New Motif in the IP Subdomain of Vertebrate pol g-a: Implications for pol g-bInteractions The spacer region of pol g-a, which connects the N-terminal exonuclease domain to the C-terminal polymerase domain, contains intrinsic processivity (IP) and AID subdomains (fig. 1). As their names suggest, these subdomains provide structural platforms for supporting both the intrinsic processivity of pol g-aaloneandtheenhancedprocessivityofthe holoenzyme, respectively (Lee et al. 2009). Our MSA also identified an insertion of 30 amino acid residues on average in the IP subdomain (E692–R722 in humans; L640–S641 in D. melanogaster) that is present consistently in all species of Vertebrata (fig. 5A). Other deuterostome species, such as St. purpuratus (Echinodermata) and Branchiostoma floridae (Cephalochordata), and some other noninsect animal species also have insertions of varying sizes in this position. However, the low sequence similarity does not provide enough evidence for homology among the new vertebrate IP motif and the regions to which it aligns in these other animals. In fact, the N-terminal region of this motif in vertebrates is highly conserved, but its C-terminus shows high variability. Again, the sequence from Ci. intestinalis resembles significantly those of insect species due to the lack of any amino acid residues in this region. In summary, this element is a distinct and conserved, derived feature in vertebrate pol g-a; for other metazoan groups, there is no clear indication of its evolutionary constraints. The new IP motif was localized in a region of the human pol g-acrystal structure for which most of the residues were also disordered (G674–R709) (Lee et al. 2009). Although the C-terminus of the motif contains a residue mutated in some cases of human patients with Alpers disease (R722H) and FIG.3.—Bayesian phylogenetic inference for animal pol g-bnucleotide sequences. The outgroup sequence used was the glycyl-tRNA synthetase from the bacterial species Thermus thermophilus (accession number AJ222643.1). The 50% majority-rule consensus tree was inferred using MrBayes 3.2, as described under Materials and Methods. Bayesian posterior probability values are indicated for almost all nodes. The scale bar indicates substitutions per site. Oliveira et al. GBE 948 Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
possibly implicated in DNA binding (Euro et al. 2011), modeling the missing 36 amino acids resulted in a helix-turn-helix structure located in close proximity to a loop at the end of the Middle domain of the proximal pol g-bprotomer (fig. 5B). Part of this loop in the proximal pol g-b, which to our knowledge has not been described previously and is hereafter called the “M loop” (T357–K364 in humans; D238–H239 in D. melanogaster), is again almost exclusive to vertebrates (fig. 5A), FIG.4.—Identification of a new Exo motif in vertebrate pol g-a, potentially implicated in primer–template DNA binding. (A) Amino acid sequence alignment indicates the presence of the extra residues (boxed) between pol g-amotifs Exo II and III in all species of Vertebrata, and a few other animal groups, including other deuterostome species. These residues are disordered in the crystal structure of the human pol gapo-holoenzyme (3IKM; Lee et al. 2009). (B) Structural model of the human residues indicated in (A) showing their proximity to the primer–template DNA binding cleft. The left panel shows the DNA binding cleft structure without the new vertebrate Exo motif, as it appears in the PDB data file 3IKM. Primer–template DNA binding to pol g-awas modeled by Euro et al. (2011) and the orienter module (see text for details) is shown as described by Szczepanowska and Foury (2010). Colors are as in figure 1. Evolution of Animal pol gGBE Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 949 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
although Echinodermata (St. purpuratus), Cnidaria (N. vectensis), and Placozoa (T. adhaerens) do possess several residues in this same region. A firm correlation between the conserved new IP motif in pol g-aand the M loop of pol g-boccurs only for vertebrate species; the absence of both structural elements is firmly correlated only for insects and tunicates. Unfortunately, without a better taxa representation, we are unable to conclude whether both elements have been lost in insects and tunicates or gained in vertebrates. Dimerization of pol g-b Structural and biochemical data have documented that the human and mouse pol g-bs form homodimers in solution and that the homodimer form associates with pol g-ato constitute a functional heterotrimeric holoenzyme (Carrodeguas et al. 2001;Fan et al. 2006;Yakubovskaya et al. 2006;Lee et al. 2009). The major pol g-bdimerization interface is provided by the HLH-b3 [helix-loop-helix, 3 b-sheets] domain (H133–R182 in humans; N63–Q65 in D. melanogaster), which is present consistently across all vertebrate species (fig. 6A). Interestingly, the echinoderm St. purpuratus and the placozoan T. adhaerens also have amino acid residues that could potentially fold into a partial HLH-b3 structure. We tested this hypothesis by modeling the structure of pol g-bfrom these two species and that from D. melanogaster and Ci. intestinalis (control species that lack completely the HLH-b3 element). Neither the helixloop-helix structure, responsible for the formation of the fourhelix bundle with the adjacent pol g-b, nor the three bsheets found at the base of the four-helix bundle are clearly observed for St. purpuratus and T. adhaerens (fig. 6B). As expected, the pol g-bmodels for D. melanogaster and Ci. intestinalis also have none of the structural elements necessary for the HLH-b3 folding. All these proteins are, therefore, most likely unable to homodimerize using the same structural features adopted by FIG.5.—Identification of a new, putative interacting surface between the catalytic and accessory subunits of vertebrate pol g.(A) Amino acid sequence alignments indicate the presence of new motifs in the pol g-aIP subdomain (boxed, left panel) and in the pol g-bmiddle domain (boxed, right panel) of all species of Vertebrata. Only the species for which both pol g-aand -bsequences were retrieved are shown. The indicated residues in pol g-aare disordered in the crystal structure of the human pol gholoenzyme (3IKM; Lee et al. 2009); likewise those in pol g-bare disordered in the crystal structure of the human pol g-bdimer (2G4C; Fan et al. 2006). (B) Model of the human residues indicated in (A), suggesting that the predicted structural elements are in close proximity to each other. The right panel shows the possible interacting region in the absence of the newly identified motifs, as they appear in the PDB files 3IKM and 2G4C. Primer–template DNA binding to pol gwas modeled by Euro et al. (2011). Colors are as in figure 1. Oliveira et al. GBE 950 Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
FIG.6.—Dimerization of vertebrate pol g-bthrough the formation of the four-helix bundle structure. (A) Amino acid sequence alignment indicates the presence of the HLH-3bdomain (boxed) in all species of Vertebrata and possibly in few other animal groups. (B) Comparison of the crystal structure of the human pol g-bdimer and structural models for pol g-bof Trichoplax adhaerens,Strongylocentrotus purpuratus,Drosophila melanogaster, and Ciona intestinalis, showing that only vertebrate pol g-bcanfoldintoaHLH-3bstructure and therefore form the four-helix bundle dimerization interface. The inset shows the three short b-sheets at the base of the HLH-3bstructure. For more information about the structural features of the four-helix bundle fold, see Kamtekar and Hecht (1995) and Carrodeguas et al. (2001). Evolution of Animal pol gGBE Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 951 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
Fan L, Sanschagrin PC, Kaguni LS, Kuhn LA. 1999. The accessory subunit of mtDNA polymerase shares structural homology with aminoacyltRNA synthetases: implications for a dual role as a primer recognition factor and processivity clamp. Proc Natl Acad Sci U S A. 96: 9527–9532. Farnum GA, Nurminen A, Kaguni LS. 2014. Mapping 136 pathogenic mutations into functional modules in human DNA polymerase gamma establishes predictive genotype-phenotype correlations for the complete spectrum of POLG syndromes. Biochim Biophys Acta. 1837:1113–1121. Farr CL, Wang Y, Kaguni LS. 1999. Functional interactions of mitochondrial DNA polymerase and single-stranded DNA-binding protein. Template–primer DNA binding and initiation and elongation of DNA strand synthesis. J Biol Chem. 274:14779–14785. Foury F. 1989. Cloning and sequencing of the nuclear gene MIP1 encoding the catalytic subunit of the yeast mitochondrial DNA polymerase. J Biol Chem. 264:20552–20560. Fukuoh A, et al. 2014. Screen for mitochondrial DNA copy number maintenance genes reveals essential role for ATP synthase. Mol Syst Biol. 10:734. Fuste JM, et al. 2010. Mitochondrial RNA polymerase is needed for activation of the origin of light-strand DNA replication. Mol Cell. 37: 67–78. Garcia-Gomez S, et al. 2013. PrimPol, an archaic primase/polymerase operating in human cells. Mol Cell. 52:541–553. Gissi C, Iannelli F, Pesole G. 2008. Evolution of the mitochondrial genome of Metazoa as exemplified by comparison of congeneric species. Heredity (Edinb). 101:301–320. Hance N, Ekstrand MI, Trifunovic A. 2005. Mitochondrial DNA polymerase gamma is essential for mammalian embryogenesis. Hum Mol Genet. 14:1775–1783. Holt IJ, Lorimer HE, Jacobs HT. 2000. Coupled leadingand lagging-strand synthesis of mammalian mitochondrial DNA. Cell 100:515–524. Hyvarinen AK, et al. 2007. The mitochondrial transcription termination factor mTERF modulates replication pausing in human mitochondrial DNA. Nucleic Acids Res. 35:6458–6474. Hyvarinen AK, Pohjoismaki JL, Holt IJ, Jacobs HT. 2011. Overexpression of MTERFD1 or MTERFD3 impairs the completion of mitochondrial DNA replication. Mol Biol Rep. 38:1321–1328. Iyengar B, Luo N, Farr CL, Kaguni LS, Campos AR. 2002. The accessory subunit of DNA polymerase gamma is essential for mitochondrial DNA maintenance and development in Drosophila melanogaster.ProcNatl Acad Sci U S A. 99:4483–4488. Iyengar B, Roote J, Campos AR. 1999. The tamas gene, identified as a mutation that disrupts larval behavior in Drosophila melanogaster, codes for the mitochondrial DNA polymerase catalytic subunit (DNApol-gamma125). Genetics 153:1809–1824. Jiang ZJ, et al. 2007. Comparative mitochondrial genomics of snakes: extraordinary substitution rate dynamics and functionality of the duplicate control region. BMC Evol Biol. 7:123. Joers P, et al. 2013. Mitochondrial transcription terminator family members mTTF and mTerf5 have opposing roles in coordination of mtDNA synthesis. PLoS Genet. 9:e1003800. Joers P, Jacobs HT. 2013. Analysis of replication intermediates indicates that Drosophila melanogaster mitochondrial DNA replicates by a strand-coupled theta mechanism. PLoS One 8:e53249. Kaguni LS. 2004. DNA polymerase gamma, the mitochondrial replicase. Annu Rev Biochem. 73:293–320. Kamtekar S, Hecht MH. 1995. Protein Motifs. 7. The four-helix bundle: what determines a fold? FASEB J. 9:1013–1022. Katoh K, Toh H. 2008. Recent developments in the MAFFT multiple sequence alignment program. Brief Bioinform. 9:286–298. Korhonen JA, Gaspari M, Falkenberg M. 2003. TWINKLE Has 50!30 DNA helicase activity and is specifically stimulated by mitochondrial single-stranded DNA-binding protein. J Biol Chem. 278: 48627–48632. Korhonen JA, Pham XH, Pellegrini M, Falkenberg M. 2004. Reconstitution of a minimal mtDNA replisome in vitro. EMBO J. 23:2423–2429. Lavrov DV, Boore JL, Brown WM. 2000. The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus. Mol Biol Evol. 17:813–824. Lecrenier N, Van Der Bruggen P, Foury F. 1997. Mitochondrial DNA polymerases from yeast to man: a new family of polymerases. Gene 185: 147–152. Lee YS, et al. 2010. Each monomer of the dimeric accessory protein for human mitochondrial DNA polymerase has a distinct role in conferring processivity. J Biol Chem. 285:1490–1499. Lee YS, Kennedy WD, Yin YW. 2009. Structural insight into processive human mitochondrial DNA synthesis and disease-related polymerase mutations. Cell 139:312–324. Lewis DL, Farr CL, Kaguni LS. 1995. Drosophila melanogaster mitochondrial DNA: completion of the nucleotide sequence and evolutionary comparisons. Insect Mol Biol. 4:263–278. Lewis DL, Farr CL, Wang Y, Lagina AT 3rd, Kaguni LS. 1996. Catalytic subunit of mitochondrial DNA polymerase from Drosophila embryos. Cloning, bacterial overexpression, and biochemical characterization. J Biol Chem. 271:23389–23394. Lewis SC, Joers P, Wilcox S, Griffith JD, Jacobs HT, Hyman BC. 2015. A rolling cirlce replication mechanism produces multimeric lariats of mitochondiral DNA in Caenorhabditis elegans. PLoS Genet. 11: e1004985. Li W-H. 1997. Molecular evolution. Sunderland (MA): Sinauer Associates. Lim SE, Longley MJ, Copeland WC. 1999. The mitochondrial p55 accessory subunit of human DNA polymerase gamma enhances DNA binding, promotes processive DNA synthesis, and confers N-ethylmaleimide resistance. J Biol Chem. 274:38197–38203. Logan DT, Mazauric MH, Kern D, Moras D. 1995. Crystal structure of glycyl-tRNA synthetase from Thermus thermophilus.EMBOJ.14: 4156–4167. Longley MJ, Prasad R, Srivastava DK, Wilson SH, Copeland WC. 1998. Identification of 50-deoxyribose phosphate lyase activity in human DNA polymerase gamma and its role in mitochondrial base excision repair in vitro. Proc Natl Acad Sci U S A. 95: 12244–12248. Mitreva M, Blaxter ML, Bird DM, McCarter JP. 2005. Comparative genomics of nematodes. Trends Genet. 21:573–581. Moraes CT, Bacman SR, Williams SL. 2014. Manipulating mitochondrial genomes in the clinic: playing by different rules. Trends Cell Biol. 24: 209–211. Oliveira MT, Garesse R, Kaguni LS. 2010. Animal models of mitochondrial DNA transactions in disease and ageing. Exp Gerontol. 45:489–502. Oliveira MT, Kaguni LS. 2010. Functional roles of the Nand C-terminal regions of the human mitochondrial single-stranded DNA-binding protein. PLoS One 5:e15379. Oliveira MT, Kaguni LS. 2011. Reduced stimulation of recombinant DNA polymerase gamma and mitochondrial DNA (mtDNA) helicase by variants of mitochondrial single-stranded DNA-binding protein (mtSSB) correlates with defects in mtDNA replication in animal cells. J Biol Chem. 286:40649–40658. Olson MW, Wang Y, Elder RH, Kaguni LS. 1995. Subunit structure of mitochondrial DNA polymerase from Drosophila embryos. Physical and immunological studies. J Biol Chem. 270:28932–28937. Philippe H, et al. 2009. Phylogenomics revives traditional views on deep animal relationships. Curr Biol. 19:706–712. Pohjoismaki JL, et al. 2010. Mammalian mitochondrial DNA replication intermediates are essentially duplex but contain extensive tracts of RNA/DNA hybrid. J Mol Biol. 397:1144–1155. Oliveira et al. GBE 958 Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from
Reyes A, et al. 2013. Mitochondrial DNA replication proceeds via a “bootlace” mechanism involving the incorporation of processed transcripts. Nucleic Acids Res. 41:5837–5850. Reyes A, Yang MY, Bowmaker M, Holt IJ. 2005. Bidirectional replication initiates at sites throughout the mitochondrial genome of birds. J Biol Chem. 280:3242–3250. Ronquist F, et al. 2012. MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space. Syst Biol. 61: 539–542. Sali A, Blundell TL. 1993. Comparative protein modelling by satisfaction of spatial restraints. J Mol Biol. 234:779–815. Spelbrink JN, et al. 2000. In vivo functional analysis of the human mitochondrial DNA polymerase POLG expressed in cultured human cells. J Biol Chem. 275:24818–24828. Stiban J, Farnum GA, Hovde SL, Kaguni LS. 2014. The N-terminal domain of the Drosophila mitochondrial replicative DNA helicase contains an iron-sulfur cluster and binds DNA. J Biol Chem. 289: 24032–24042. Suyama M, Torrents D, Bork P. 2006. PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res. 34:W609–W612. Szczepanowska K, Foury F. 2010. A cluster of pathogenic mutations in the 30-50exonuclease domain of DNA polymerase gamma defines a novel module coupling DNA synthesis and degradation. Hum Mol Genet. 19:3516–3529. Wang Y, Farr CL, Kaguni LS. 1997. Accessory subunit of mitochondrial DNA polymerase from Drosophila embryos. Cloning, molecular analysis, and association in the native enzyme. J Biol Chem. 272: 13640–13646. Wang Y, Kaguni LS. 1999. Baculovirus expression reconstitutes Drosophila mitochondrial DNA polymerase. J Biol Chem. 274: 28972–28977. Wanrooij S, et al. 2008. Human mitochondrial RNA polymerase primes lagging-strand DNA synthesis in vitro. Proc Natl Acad Sci U S A. 105: 11122–11127. Wernette CM, Kaguni LS. 1986. A mitochondrial DNA polymerase from embryos of Drosophila melanogaster. Purification, subunit structure, and partial characterization. J Biol Chem. 261:14764–14770. Wolf YI, Koonin EV. 2001. Origin of an animal mitochondrial DNA polymerase subunit via lineage-specific acquisition of a glycyl-tRNA synthetase from bacteria of the Thermus-Deinococcus group. Trends Genet. 17:431–433. Yakubovskaya E, Chen Z, Carrodeguas JA, Kisker C, Bogenhagen DF. 2006. Functional human mitochondrial DNA polymerase gamma forms a heterotrimer. J Biol Chem. 281:374–382. Yang MY, et al. 2002. Biased incorporation of ribonucleotides on the mitochondrial L-strand accounts for apparent strand-asymmetric DNA replication. Cell 111:495–505. Yasukawa T, et al. 2006. Replication of vertebrate mitochondrial DNA entails transient ribonucleotide incorporation throughout the lagging strand. EMBO J. 25:5358–5371. Ye F, Samuels DC, Clark T, Guo Y. 2014. High-throughput sequencing in mitochondrial DNA research. Mitochondrion 17C:157–163. Young MJ, et al. 2011. Biochemical analysis of human POLG2 variants associated with mitochondrial disease. Hum Mol Genet. 20: 3052–3066. Associate editor: Sarah Schaack Evolution of Animal pol gGBE Genome Biol. Evol. 7(4):943–959. doi:10.1093/gbe/evv042 Advance Access publication March 3, 2015 959 at Tampere University Library. Department of Health Sciences on November 15, 2016http://gbe.oxfordjournals.org/Downloaded from