scieee AI-readable full text Open interactive document viewer

Study of Protein-nucleic acid Complexes

Luís Filipe de Castro Fernandes

Full text

Study of Proteinnucleic acid Complexes Luis Filipe de Castro Fernandes Mestrado em Química Departamento de Química e Bioquímica 2013 Orientadora Maria João Ramos, Professora catedrática, Faculdade de Ciências da Universidade do Porto Co-orientadora Irina Sousa Moreira, Investigadora externa, Faculdade de Ciências da Universidade do Porto 2 FCUP Study of Protein-nucleic acid Complexes FCUP Study of Protein-nucleic acid Complexes 3 Todas as correções determinadas pelo júri, e só essas, foram efetuadas. O Presidente do Júri, Porto, ______/______/_________ 4 FCUP Study of Protein-nucleic acid Complexes Acknowledges O meu primeiro agradecimento terá que ir para as minhas orientadoras, em especial para a professora João, que me deu a oportunidade de trabalhar no grupo mesmo estando fora do país e de férias quando a abordei. Só mesmo esta oportunidade para compensar a falha do estágio na GALP. Depois queria agradecer todo o apoio e ajuda dos meus colegas de sala; O Rui, o João Martins e a Elisabete pelo ambiente descontraído e pela ajuda crucial que me permitiu avançar; os colegas da sala do lado; o Zé, o Eduardo, a Sílvia, o Rui e o João Ribeiro, que estavam sempre disponíveis para ajudar independentemente da quantidade de trabalho que tinham, sem eles não teria conseguido aprender tanto; os colegas do outro lado, o Óscar, a Natércia e o João Coimbra, que também estavam sempre lá! Sempre dispostos a partilhar o que sabiam para eu poder avançar; e também o Sérgio e o Nuno que com o seu conhecimento teórico me permitiram perceber onde estava a errar. Sem a ajuda de todos com certeza não teria aprendido tanto e com certeza não teria conseguido montar esta investigação e a sua consequência, esta tese. Para acabar não posso deixar passar as qualidades deste grupo/família, conseguem passar o espírito de entre ajuda e o orgulho de pertencer ao grupo com a maior das naturalidades. Não admira que continuem a crescer desta forma! Gostava também de agradecer ao núcleo duro composto pela Filipe, a Cláudia, a Ana e a Tânia que foram o pilar do primeiro ano de mestrado e o suporte do ano de tese, espero que esta amizade continue para a vida. Um agradecimento também ao Jumbo por me ajudar a pagar as propinas desde a licenciatura. E o grande obrigado à Carolina que tem sido a minha alegria de viver nos últimos 8 anos, por todo o suporte, motivação e capacidade para me obrigar a levantar nos momentos difíceis e a manter-me lá em cima quando estou motivado. Sem ti não tinha acabado a tese. FCUP Study of Protein-nucleic acid Complexes 5 Abstract The study of biological systems, their structure and function is of great importance for science and can lead to advances in the scientific knowledge and contributions to possible discoveries from a pharmaceutical point of view. In this work we studied 4 Protein-DNA complexes and 2 Protein-RNA complexes using several computational techniques such as Molecular Dynamics simulations (MD) and the application of the Alanine Scanning Mutagenesis (ASM) methodology for the identification of hot-spots (HS) and null-spots (NS). We also made Root Mean Square Deviation (RMSD) profiles as well as Radial Distribution Function (RDF) and solvent accessible surface area (SASA) analysis. The MD simulations were carried out for 10ns in explicit solvent, using the ff99SB force field for DNA-based complexes and rnaff99 force field for RNA-based complexes. A total of 30 residues from the 6 complexes (23 from DNA and 7 from RNA), were mutated to alanine and their binding free energy calculated and compared to experimental values. In the end we were able to get a good correlation with the experimental values with an average error of 3.08 kcal/mol. These values are as valuable as they can be because there are very few studies in this field for DNA-based complexes and no-one for RNA-based complexes. Therefore, it leaves a good base for future studies, future optimizations and future generalization and implementation of this kind of studies. We also complemented this study with RDF and SASA analysis to support Oring theory that states that HS would be surrounded by regions with higher packing density, more deeply buried. This leads to solvent exclusion around them and results in a lower local dielectric constant environment and enhancement of specific electrostatic and hydrogen bond interactions. This region would be surrounded by another one formed by NS, whose role would be to shelter the HS from bulk solvent. 6 FCUP Study of Protein-nucleic acid Complexes Keywords Molecular Dynamics (MD), Alanine Scanning Mutagenesis (ASM), Radial Distribution Function (RDF), Solvent Accessible Solvent Area (SASA), Protein-DNA, Protein-RNA, Hot-Spots (HS), Null-Spots (NS). FCUP Study of Protein-nucleic acid Complexes 7 Index Acknowledges............................................................................................................... 4 Abstract ........................................................................................................................ 5 Keywords ...................................................................................................................... 6 Index ............................................................................................................................. 7 Table index ................................................................................................................... 9 Figure index ................................................................................................................ 10 Abbreviation Index ...................................................................................................... 12 1 - Introduction ............................................................................................................ 14 1.1 - Context............................................................................................................ 14 1.2 - Hot spots and the O-ring theory ...................................................................... 15 1.3 - Protein – Protein interfaces ............................................................................. 16 1.4 - Protein – nucleic acid complexes .................................................................... 17 1.5 - Protein-based complexes ................................................................................ 18 1.5.1 - Escherichia Coli replication terminator protein .......................................... 19 1.5.2 - DNA and protein NHP6A .......................................................................... 19 1.5.3 - DNA subunit RPA70 and Human replication protein ................................. 20 1.5.4 - DNA and integrase protein TN916 ............................................................ 20 1.5.5 - U1A mutant and RNA complex ................................................................. 21 1.5.6 - RNA binding domain of Human fox-1 in complex UGCAUGU ................... 22 1.6 - Methodology ................................................................................................... 23 1.6.1 - Computational chemistry/Biochemistry ..................................................... 23 1.7 - Molecular Mechanics and force field ............................................................... 24 1.7.1 - Energy minimization ................................................................................. 28 1.7.2 - Molecular dynamic .................................................................................... 28 1.7.3 – Ensembles ............................................................................................... 28 1.7.4 Integration time ........................................................................................... 29 1.7.5 – Periodic boundary conditions ................................................................... 30 8 FCUP Study of Protein-nucleic acid Complexes 1.8 - MMPBSA ...................................................................................................... 31 1.8.1 - Solvation ................................................................................................... 32 1.9 - Alanine Scanning Mutagenesis ....................................................................... 33 2 - Methodology .......................................................................................................... 35 2.1 - Systems preparation ....................................................................................... 35 2.2 - Molecular Dynamics ........................................................................................ 36 2.3 - Alanine scanning mutagenesis ........................................................................ 36 2.4 - System analysis .............................................................................................. 38 2.4.1 – Root Mean Square Deviation ................................................................... 38 2.4.2 – Radial Distribution Function ..................................................................... 39 2.4.3 – Solvent Accessible Solvent Area .............................................................. 39 3 - Results and Discussion.......................................................................................... 40 3.1 - RMSD ............................................................................................................. 40 3.2 - RDF ................................................................................................................ 42 3.2.1 HS ≥ 2.0 kcal/mol cut-off ............................................................................. 43 3.2.2 HS ≥ 1.0 kcal/mol cut-off ............................................................................. 45 3.3 - SASA .............................................................................................................. 46 3.4 - Mutagenesis in Protein acid nucleic interfaces ................................................ 53 Conclusion .................................................................................................................. 62 References ................................................................................................................. 65 Annexes...................................................................................................................... 70 RDF ........................................................................................................................ 70 MM-PBSA ............................................................................................................... 73 Paper – Extending the applicability of the O-ring theory to protein-DNA complexes .... 76 FCUP Study of Protein-nucleic acid Complexes 9 Table index Table 1 - Differences between DNA and RNA ............................................................ 14 Table 2 – Characteristics of ab-initio methodologies ................................................... 23 Table 3 – Characteristics of Semi-empiric methodologies........................................... 23 Table 4 – Characteristics of Density Functional theories methodologies ..................... 24 Table 5 – Characteristics of Molecular Mechanics methodologies .............................. 24 Table 6 - Composition of the 6 systems subjected to MD simulations ......................... 36 Table 7 - Description of the 30 residues that constitute the dataset, evidencing the respective system, PDB numeration, amino acid type and experimentally ∆∆Gbinding. .. 38 Table 8 - Average number of water molecules around NS and HS for each complex and global considering HS ≥ 2.0 kcal/mol. .................................................................. 43 Table 9 - Average number of water molecules around NS and HS for each complex and global considering HS ≥ 1.0 kcal/mol. .................................................................. 45 Table 10 – Results for ∆SASA and relSASA for all residues with known ∆∆Gbinding for the studied complexes. ............................................................................................... 48 Table 11 - MM-PBSA results for ɛ1 to ɛ4 .................................................................... 54 Table 12 - MM-PBSA results for ɛ5 to ɛ9 .................................................................... 55 Table 13 - Results of average errors obtained with the NLPB equation. ..................... 56 Table 14 - Results of Statistical tests for the 2.0 kcal/mol cut-off ................................ 57 Table 15 - Results of Statistical tests for the 1.0 kcal/mol cut-off ................................ 59 16 FCUP Study of Protein-nucleic acid Complexes protein-RNA interfaces, have as much biological interest as the protein-protein. However, the information regarding experimentally detected HS in these complexes or the application of the alanine scanning mutagenesis method to this type of interface is still scarce. It probably occurs due to the difficulties in energetic characterizing this type of system as it possesses a highly charged character. Regardless, it was observed the same organization of HS in the central region of the interface but with a different composition. For protein-DNA interfaces there is a higher occurrence of positively charged residues (Arginine and Lysine), as well as, a lower occurrence of hydrophobic and negatively charged residues [22]. The identification of HS can be made in laboratory (in vivo or in vitro). Among them are Chromatin immunoprecipitation (ChIP), where living cellules are treated with formaldehyde to stabilize Protein-DNA interactions allowing their purification and detection, DNA electrophoretic mobility shift essays, used to test the affinity and specificity levels of the interactions and microplate capture and detection assay, just to name a few. The computational techniques for the study of the free energy differences upon alanine mutation of acid nucleic systems are not fully understood. Therefore, we will study how to implement in silico detection of HS in these systems as well as their characteristic accessibility to solvent. 1.3 - Protein – Protein interfaces Protein - protein interactions (PPIs) are involved in a wide variety of cellular processes and are critical events in most biological pathways and their function or malfunction results in a variety of diseases, turning these interfaces compelling targets for drug discovery [23, 24] . There have been several attempts to understand them in terms of physical features of the associating surfaces and energetic contributions made by each residue [25]. This identification of the key residues that are important for the interaction is very difficult, due in part to an incomplete understanding of the sources of affinity and specificity of interfaces [24]. An accurate understanding of the factors that make certain residues more important than others in facilitating these interactions is also going to be of enormous importance at redesigning the affinity and specificity of natural occurring interactions as well as for new protein-protein designs [25]. The MMPBSA (molecular mechanics-Poisson Boltzmann surface area) is widely used to investigate PPIs and other interactions for around a decade combining the speed of a continuum approach to modelling solvent interactions with theoretical accuracy of an MM-based approach to atomistically modelling protein-protein interactions. In order to FCUP Study of Protein-nucleic acid Complexes 17 predict the location of HS at interfaces, they also have being used in various alaninescanning mutagenesis protocols, calculating the relative free energy change (∆∆Gbinding) between the wild-type and mutant complex upon alanine mutation.[23, 26]. 1.4 - Protein – nucleic acid complexes Protein-DNA interactions play an essential role in many cellular functions such as transcription, replication, recombination and DNA packaging. Protein-RNA interactions are also essential in biological process and some of their functions are transcription termination, mRNA splicing, mRNA export to the nucleus and cytoplasm, intracellular localization of transcripts, mRNA translation, mRNA stability and processing tRNA and rRNA [27]. Since the development of computational methods powerful and reliable enough there has been a continuous growth in the numbers of Protein-DNA structures available to study. (Protein Data Bank[28] - PDB) and they are created using X-ray crystallography. All the research work done on this matter relies on the quality (resolution) of the structures. With a higher resolution, better results are possible regarding interaction between atoms and between molecules, organization and reorganization of base sequences and structure modification after bond breaking/formation. There are other useful databases that can complete PDB’s information like ProNIT(Thermodynamic database for protein-nucleic acid interactions) [29-31] where can be found experimental data for several thermodynamic and energy parameters for Protein-DNA and ProteinRNA complexes and AANT (amino acid-nucleotide interaction database) [32] where can be found statistical information regarding aminoacid-nucleotid interactions. DNA structures can be classified according to their function and can be divided in three classes: (i) Enzyme – if its primary function is the modification of DNA; (ii) Transcription factor – if its function id the regulation of the expression and transcription of genes (iii) Support protein – if its function is solely to provide DNA support. These classes can be further divided into types considering their function and structure: (a) for enzymes we have 6 sub-categories ( oxidoreductases, transferases, hydrolases, lyases, isomerases and ligases); (b) for transcription factors we have 7 sub-categories ( Alpha Helix, Alpha/Beta, Beta Sheet, Helix turn Helix, Ribbon/Helix/Helix, Zinc Coordinating and Zipper type); (c) for supporting Proteins we have 8 sub-categories[33, 34]. 18 FCUP Study of Protein-nucleic acid Complexes RNA structures can be classified according to their function and can be divided in two major classes: (i) ribosomal RNA that ensures the correct protein sequence correcting missing codons and (ii) RNA polymerase that recognizes the correct sequence and synthetizes it. 1.5 - Protein-based complexes In the following sub sections we will give a very light description of each complex under study. Figure 1 - Representation of the 4 protein-DNA complexes and 2 protein-RNA complexes studied in this work. Protein and DNA are in cartoon and stick representation respectively 1ECR 1J5N 1JMC 1TN9 1URN 2ERR FCUP Study of Protein-nucleic acid Complexes 19 1.5.1 - Escherichia Coli replication terminator protein The Escherichia Coli replication terminator protein (PDBid: 1ECR [35]), occurs at discrete Ter sites. These sites block replication fork progression in vivo, when the replication fork approaches from on direction, the non-permissive direction, there is a function creating a trap that restricts the meeting of the convergent forks of a certain chromosome to a certain region.[36] The replication of DNA in many prokaryotes and in certain regions of eukaryotic chromosomes is specifically terminated at specialized sequences called replication termini, Ter, that cause orientation-dependent forks arrest, which performs important physiological functions[37]. The DNA replication termination protein, TUS, blocks the progress of the replisome in the final stages of the chromosomal replication in Escherichia Coli and related bacterial species [38]. In vitro analyses have shown that the replication terminator protein of Escherichia Coli is a polar contrahelicase, meaning, the protein causes a unidirectional arrest of the replicative helicase DnaB upon binding to the ter sequence.[37] The crystal structure of the Tus-Ter complex indicates that the core DNA-binding domain of the protein consisting in two pairs of antiparallel beta-strands that lie in the major groove of the DNA. [38] 1.5.2 - DNA and protein NHP6A The complex between DNA and the protein NHP6A (PDBid: 1J5N [39]), is a HMG box protein that can be found in Saccharomyces cerevisiae. HMG is the acronym of High mobility group, and it’s a conserved domain of ~80 amino acids witch mediates DNA binding of many proteins.[40]The first class (HMG1) is generally transcription factors that bind to DNA in a sequence specific fashion and are expressed only in a few cell types, containing only one HMG box, while the second class(HMG2) is more abundant and often contains two or more HMG boxes, binding to DNA with little or no sequence specificity[40]. HMG proteins are small chromatin associated eukaryotic proteins that alter the physical properties of DNA in vitro and in vivo[41]. These proteins are members of a class of small proteins that are abundant in eukaryotic cells and are sequence-nospecifically bind to DNA.[42] There are two groups of HMGs, the A group and the B group and they differ in terms of shape and orientation of its first alfa-helix and the identity of potential intercalating residues. The A box domains is known to bend 20 FCUP Study of Protein-nucleic acid Complexes less DNA[43]. Each homologous motif contains amino acids that form three alfa helices to bind DNA as an L-shaped structure. [41] 1.5.3 - DNA subunit RPA70 and Human replication protein The complex (PDBid: 1JMC [44]) is the representation of the human replication protein (RPA) which is a key factor in DNA metabolism including DNA replication, DNA repair and recombination [45]. It’s a modular multi domain protein that functions in a wide range of DNA pathways required to maintain and propagate the genome of all living organisms[46], its constituted by a stable single stranded DNA binding protein composed by three subunits (70kDa, 32kDa and 14kDa; RPA70, RPA32 and RPA14 respectively)[45, 47, 48]. RPA function by interfacing with dynamic multi-protein machinery and acts as a central hub that links many DNA transactions, it also provides the primary single-stranded DNA (ssDNA) binding activity in eukaryotes and even servers as a scaffold and coordinator of DNA processing machinery[46]. RPA is highly conserved throughout evolution, and homologous, heterotrimeric single stranded DNAbinding proteins have being identified in all eukaryotes examined [45]. The primary interaction of RPA is with ssDNA, however , RPA function requires interactions with other forms of DNA, it binds to damaged DNA and double stranded DNA (dsDNA) and can cause dsDNA helices destabilization, this destabilization is a manifestation of ssDNA activity[47]. 1.5.4 - DNA and integrase protein TN916 The crystal structure of the DNA binding domain of Tn916 integrase (PDBid: 1TN9 [49]) is essential for excision and reintegration of bacterial Tn916 conjugative transposon and the latter spreads antibiotic resistance among pathogenic bacteria[50]. Tn916 is a conjugative transposon (also called Integrative conjugative elements, ICEs [51]), and like most transposons is extremely promiscuous genetic element that disseminates antibiotic resistance among gram positive and gram negative bacteria[52], serving as a major contributor to bacterial evolution by passing the antibiotic resistance, virulence genes and metabolic genes across species and genus lines [51]. Tn916 is also of the most extensively studied transposon. FCUP Study of Protein-nucleic acid Complexes 21 Unlike most DNA-binding domains reported that bind to major groove using - helix [53], the Tn916 N-terminal domain (INT-DBD) recognizes the major groove using the face of a three-stranded beta-sheet. The major protein-DNA contacts occur at the largely hydrophobic interface formed by turn T1 and strands Beta2 and beta3[50]. This N-terminal domain, INT-DBD recognizes DNA by a rare structural motif, the three stranded beta sheet[54]. 1.5.5 - U1A mutant and RNA complex The crystal structure of an RNA recognition motif (RRM) is also known as ribonucleoprotein (RNP) consensus domain or RNA binding domain (RBD). It is characterized by highly conserved regions located centrally on a beta sheet, which forms the RNA binding surface [55-57], this domain is the third most common in human proteins [58] (PDBid: 1URN [59]). It is present in one or more copies in hundreds of RNA binding domains and proteins that carry RRM domains play critical roles in a wide variety of cellular processes, including RNA processing and packaging, mRNA export, translation, RNA degradation and gene regulation [27, 55, 56, 58]. These domains are about 90 amino acids long and fold into a globular structure consisting of a four-stranded antiparallel beta-sheet (the RNA binding surface) backed by two alfa-helices and are characterized by the presence of two highly conserved stretches of 8 and 6 amino acids, known as RNP1 and RNP2 consensus sequences, which lie strategically in the center of the beta sheet surface and domain conserved aromatic residues critical for RNA binding [55, 56, 58], contrasting to most DNA-binding proteins, which are presented with a double-stranded b-form helix of uniform structure. RNA-binding structures must be able to bind targets with widely differing structures and must be able to bind to its correct RNA target with appropriate kinetics, affinities that correspond to the function of the complex, ranging from relatively nonspecific, transient binding (such as the binding involved in general RNA processing), to highly specific and stable interactions (such as those involved in the formation of intracellular machinery) [27, 56].This recognition is done by both sequence and structure displaying a considerable variety in the binding affinities [57]. Because the steep and narrow groove of double stranded RNA does not provide proteins easy access to the bases for sequence-specific recognition, most RNA-binding protein recognize single-stranded regions to distort double-stranded regions in which the major groove has been widened by bulges, hairpins or loops [56]. 22 FCUP Study of Protein-nucleic acid Complexes 1.5.6 - RNA binding domain of Human fox-1 in complex UGCAUGU The RNA element UGCAUGU (represented with PDBID:2ERR [60]) has long known to strongly influence splicing of a variety of alternative exons in mammalian genes, including the c-src N1 exon, the calcitonin/CGRP exon4, the fibronectin exon IIIB, the fibroblast growth factor receptor 2 exon and the nonmuscle myosin II heavy chain B exon N30 [60, 61]. RNA splicing plays a critical role in the programming of neuronal differentiation and has a consequence in human neurodevelopment [62] so genes targeted by neuronal FOX-1 are much more likely to be involved in neuronal cytoskeletal rearrangements and neuronal vesicular and protein transport functions, as an example, analysis of RNA recognition sites characterized for brain specific Fox-1 showed that these sequences are highly represented in alternatively spliced transcripts preferentially expressed in neurons [63]. The fox-1 gen was originally identified in Caenorhabditis elegans, where it acts as a numerator element in counting the number of X chromosomes relative to ploidity, and determining male or hermaphrodite development. It is thought to posttranscriptionally repress the expression of Xol-1 (the main switch controlling sex determination). But since several alternatively spliced isoforms of Xol-1 exist while only one of these splice variants is necessary and sufficient as a sex determinant, it was speculated that Fox-1 might led to unproductive splicing of the Xol-1 gene [60, 61]. The Fox-1 family of RNA binding proteins are regulated by alternative splicing in neurons, so Fox-1 and alike proteins are expressed predominantly in brain, skeletal muscle and cardiac muscle [63, 64]. The Fox-1, in addition to the numerous hydrophobic and electrostatic interactions that provide affinity, also has a dense network of hydrogen bonds that provide sequence specificity to the first six nucleotides 5’-UGCAU-3’, being the most important interactions at the protein-RNA interface [60]. FCUP Study of Protein-nucleic acid Complexes 23 1.6 - Methodology 1.6.1 - Computational chemistry/Biochemistry Since the exponential development of computers over the last decades, it was possible to merge the traditional experimental Chemistry/Biochemistry with the processing ability of machines giving birth to a new way of create science. It was the creation of Computational Chemistry and Computational Biochemistry. This new way of doing science allowed scientists among other things to study and simulate a great deal of protein-based systems and to describe its properties such as molecular structure, interactions between its components, geometry and a vast array of thermodynamic properties. One of the main goals has been the creation and discovery of new drugs through hit/lead methodologies and their optimization. The various methodologies used can be split in four major groups with its advantages, disadvantages and specific targets as shown in the Tables 2, 3, 4 and 5. ab-initio Description - Uses quantic physics - Does not include empiric parameters - Mathematical rigorous - Based on wave functions Advantages - Can be used in all kind of systems - Does not depend on experimental data - Allows calculation of transition and excited states Disadvantages - Very demanding from the computational view Systems studied - Not over a few hundred atoms Table 2 – Characteristics of ab-initio methodologies Semi-empiric Description - Uses quantic physics - Uses experimental parameters and empiric simplifications - Includes several approximations - Based on wave functions Advantages - Less demanding from the computational view when compared with ab-initio and DFT methods - Allows calculations of transition and excited states Disadvantages - Require experimental data or ab-initio calculations to the parameter derivation - Less rigorous than the ab-initio methods Systems studied - Not above the thousands of atoms Table 3 – Characteristics of Semi-empiric methodologies 24 FCUP Study of Protein-nucleic acid Complexes DFT (Density Functional Theory) Description - Based on density functionals - Uses parameters on the functionals Advantages - Include the electronic correlation term - Less demanding from the computational view than abinitio methods, using the same quality of calculations - The results achieved are better in systems of open layer - Allows calculation of transition states Disadvantages - Less rigorous than the ab-initio methods when the functionals don’t adapt to the parameters to calculate - Not possible to calculate excited states - Not possible to improve systematically the results - Not very accurate to describe dispersive interactions Systems studied - Not above the hundred atoms Table 4 – Characteristics of Density Functional theories methodologies Molecular mechanics Description - Based on the laws of the classic mechanics - Use of force fields based on empirical parameters Advantages - The calculations are very fast and useful specially when the computational resources are scarce and the parameters used are adequate to the system - Can be used to study large systems like enzymes Disadvantages - Does not calculate electronic properties - Require experimental data or ab-initio calculations for the derivation of parameters - Force field use is limited to a certain type of system Systems studied - Systems may have hundreds of thousand atoms Table 5 – Characteristics of Molecular Mechanics methodologies The method used must be adequate to the system studied and in some cases more than one method can be used. One example of that is using a hybrid method like QM/MM (Quantum Mechanics/Molecular Mechanics), in which we can have the advantages of the quantum physics like the precision of the quantic physics and the speed of the molecular mechanics in the study of chemical processes in solution or in proteins. The quantum mechanics is used to study the smaller parts of the system like the nucleus and molecular mechanics to study the rest of the system. 1.7 - Molecular Mechanics and force field This is the most suitable method to study protein systems, since it requires a lower level of computational power than quantic methods. Like said in Table 5, it uses the laws of classic mechanics and Newton Laws to describe the particles movement. FCUP Study of Protein-nucleic acid Complexes 25 Molecular mechanics calculations, also known as force field calculations, can be of considerable use in the qualitative descriptions of systems. In these cases, we concentrate on the structural aspect and not on the electronic and/or spectroscopic properties. In essence, we describe the potential energy surface without invoking any quantum mechanical calculations or descriptions. The Born-Oppenheimer approximation, fundamental to our molecular description, states that the Schrodinger Equation for a molecule can be separated into a part describing the motions of the electrons and a part describing the motions of the nuclei and that these two motions can be studied independently. This can be interpreted in one of two manners, one of which allows the study of the electronic structure, one of which allows the study of the molecular mechanics structure. But sine this method only considers the most important nuclear movements to describe the molecule and does not consider its electrons, it is not capable of acknowledge the formation and break of bonds and electronic excited states. In molecular mechanics, the smallest particle of the system is the atom, so the nuclei and the electrons are treated using parameterization with a force field. In molecular mechanics a force field is a mathematic expression of physical variables to describe the potential energy of a system. There is a common expression to all force fields to calculate the energy: The bond stretching (energy required to stretch or compress a bond between two atoms), bending (energy required to bend a bond from its equilibrium angle) and torsional (torsional energy to dihedral angles) terms are called bonded interactions because the atoms involves must be directly bonded or bonded to a common atom. The Van der Waals (energy responsible for the liquefaction of non-polar gases like O2 and N2, also govern the energy of interaction of non-bonded atoms within a molecule. These interactions contribute to the steric interactions in molecules and are often the most important factors in determining the overall molecular conformation (shape), being the most important to determine the three dimensional structure of many biomolecules, especially proteins) and electrostatic ( when bonds in the molecule are polar, partial electrostatic charges will reside on the atoms. These interactions are represented with a Columbic potential function) terms are between non-bonded atoms. The last term correlate the previous ones, but is often omitted because it increases greatly the computational time. 32 FCUP Study of Protein-nucleic acid Complexes 1.8.1 - Solvation The free energy of each residue was estimated as the sum of the molecular mechanical free energy, the solvation free energy and the contributions from the vibrational, rotational and translation entropy. This energy can be divided into polar (Gpolar) and nonpolar (Gnp) contributions (equation 12). [12] The nonpolar solvation term includes the energetic cost of the cavity formation, solvent re-arrangement and interactions between solvent-solute, so, this term represents the free binding energy of the molecule when its removed from all charged (partial charges are taken as zero) as seen in equation 13). ∆Gdispersion is the energy of the Van der Waals interactions between solventsolute and the ∆Gcavity term includes the entropic penalization due to the rearrangement of the solvent molecules around the solute and the work realized to create the cavity needed to make the solute emerge. Both terms are proportional to solutes SASA making the nonpolar term able to be estimated with equation 14, Where A is the SASA value estimated by Molsurf software included in the AMBER package and σ and β are empiric constants with values of 0.00542 kcal Å-2 mol-1 and 0.92 kcal mol-1 respectively. The polar solvation term can be calculated by solving the Generalized Born (GB)[82, 83] equation or by the approximation Poisson Boltzmann (PB)[84]. Since the PB model is considered the most precise it was used as reference in GB models which are considered to be more efficient in a computational point of view. Since this method was designed for Protein-Protein interfaces, its results on Protein-Nucleic acid are not as accurate, and to solve this problem, it was needed to use a different method to calculate the polar solvation term. The program used was DelPhi [85, 86] which uses a different approach where the protein is modelled as a dielectric continuum of low polarizability embedded in a dielectric medium of high polarizability[87]. The Gpolar solvation was calculated by solving the Linear Poisson−Boltzmann (LPB) equation, the traditional method, and the Nonlinear Poisson−Boltzmann (NLPB) equation, which accounts for the importance of salt concentration in the medium. This factor is FCUP Study of Protein-nucleic acid Complexes 33 particularly important in protein−DNA interfaces due their highly charged and polar character. With this in mind, we used a value of 2.5 grids/Å for scale; a value of 0.001 kT/c for the convergence criterion; a 90% for the fill of the grid box; and the Coulombic method to set the potentials at the boundaries of the finite-difference grid. The dielectric boundary was taken as the molecular surface defined by a 1.4 Å probe sphere and by spheres centered on each atom with radii taken from the Parse33 vdW radii parameter set. The salt concentrations used were 0.010 M and 0.145 M, which are in the physiological range.34 We have also calculated the electrostatic solvation energy term using PB solver implemented in the pbsa module from the AMBER package. We tested a set of nine different dielectric constants, from 1 to 9, to mimic the expected rearrangement upon alanine mutation and to assess the importance of each dielectric constant in the determination of ∆∆Gbinding. 1.9 - Alanine Scanning Mutagenesis Alanine-Scanning Mutagenesis is an extension of the MM-PBSA and can be used to identify mutations that can enhance the binding affinities of the complex due to its ability to estimate the contribution of each residue to the protein-protein, proteinDNA/RNA or protein-ligand binding[78]. It’s also one of the most used methods to analyse and detect HS and to study the functional groups of the lateral chains of the amino acids in specific points. It works by replacing the original residue with an alanine and calculating its free binding energy to compare with the original residue’s free binding energy. In this work were considered two ways of recognizing HS since there still isn’t a consensual value. So, first HS were considered to have a free binding energy of 1.0 kcal/mol upon alanine mutation and then HS were considered to have a free binding energy of 2.0 kcal/mol. 34 FCUP Study of Protein-nucleic acid Complexes Figure 3 - Schematic representation of the initial method formulation for the determination of ∆∆Gbinding. Wild-type structure MD simulation with explicit solvent representation ASM calculation and analysis using electrostatic and polar solvation terms ɛ1 to ɛ9 ∆∆Gbinding FCUP Study of Protein-nucleic acid Complexes 35 2 - Methodology 2.1 - Systems preparation The first step was to find Protein-DNA systems with known experimental data for free binding energies upon alanine mutation of the interfacial residues. To this end we used the ProNIT database. [28-30] Six different complexes, 4 Protein-DNA and 2 Protein-RNA were studied (Figure 1) for a total of 30 residues: (i)The protein - replication-termination-protein and DNA (PDBid: 1ECR[35]); (ii) the nonhistone chromosomal protein and DNA (PDBid: 1J5N [39]) (iii) the human replication protein A and DNA (PDBid: 1JMC [44]); (iv) the N-terminal domain of the Tn916 integrase protein bound to its DNA-binding site (PDBid: 1TN9 [49]); (v) the protein U1A and RNA (PDBid:1URN[49]) and (vi) Ataxin-2-binding protein 1 and RNA (PDBid: 2ERR[60]). Then, we retrieved the 3D structures from the PDB [28] and process them. We first protonate the amino acids since the crystallographic structures in PDB do not possess enough resolution to have the hydrogen atoms. To access the protonation state of residues we used the Propka [88-90] software within the PDB2PQR [91, 92] server. Then the leap program included in the AMBER[69] package was used to create the necessary input files to run the MD simulations using the ff99SB force field for protein-DNA complexes and rnaff99 for protein-RNA complexes. Leap was used to solvate each system with a 10Å TIP3P [93, 94] water box. An appropriate number of Na+ ions were added to properly neutralize the system and the input files for the simulation were saved: the topology one (.top) and the coordinates one (.crd). The composition of the system in study is summarized in Table 6. 36 FCUP Study of Protein-nucleic acid Complexes Complex Residues Atoms AA DNA Waters Ions Total 1ECR 305 30 13052 20Na+ 13407 45164 1J5N 93 30 9791 22Na+ 9936 31886 1JMC 238 8 11408 10 Na+ 11664 38207 1TN9 69 26 7963 20Na+ 8078 25905 1URN 96 21 7716 11Na+ 7833 25377 2ERR 109 7 6388 3Na+ 6504 20805 Table 6 - Composition of the 6 systems subjected to MD simulations 2.2 - Molecular Dynamics The MD simulations were executed in three steps: (i) Minimization step; (ii) Heating run and (iii) Production run. The minimization step is required to eliminate bad contacts in the crystallographic structures and the interaction of the proteins with the solvent. We used the SANDER module in AMBER09[69] package. The systems were subject to 2 ns of heating where the temperature was gradually increased since 0 to 300K with an ensemble NVT, followed with 8 ns of production with an ensemble NPT. The Langevin algorithm was used to regulate the temperature of the system. The Particle Mesh Ewald (PME) was used to treat electrostatic interactions of long range, being the non-ligand interactions blocked to a 10 Aº radius. In every simulation, the SHAKE[76] algorithm was used to constrain all covalent bonds involving hydrogen atoms. The integration step was 2 fs. For all systems were executed 10 ns simulations using explicit solvent and the ff99SB force field for DNA-based complexes and rnaff99 force field for RNA-based complexes. 2.3 - Alanine scanning mutagenesis MM-PBSA method was used to calculate the bond free energies after alanine mutation of the interfacial residues. Equation 4 was used to calculate this free energy. The entropic term of this equation (the last one), can be neglected since its partial contributions tend to be negligible. The first three terms were introduced the way they FCUP Study of Protein-nucleic acid Complexes 37 are given by the method, but the free energy of polar solvation had to be calculated using the resolution of the linear and non-linear equation of Poisson-Boltzmann recurring to DelPhi software. In this continuum method, the protein is modelled as a dielectric continuum of low polarizability embedded in a dielectric medium of high polarizability. Because the following parameters have been shown in earlier works to constitute a good compromise between accuracy and computing time, they were set as: (i) a scale of 2.5 grids/Aº; (ii) a convergence criterion of 0.001 kT/c and (iii) the molecule filled 90% of the grid box. The dielectric boundary was taken as the molecular surface defined by a 1.4Aº probe sphere and by spheres centred in each atom. To verify the most correct dielectric constant that should be used to simulate the rearrangement of the protein upon alanine we used values from 1 to 9. To ease the systematic work required, we used a VMD plugin called CompASM [95] which provides an easy way to prepare input files and analyse results through the use of a graphical interface. The mutated residues are listed in Table 7. Protein #AA PDB #AA Mutated Gbinding / kcal mol-1 references 1ECR 198 R 1.19 [38] 1ECR 250 Q 0.11 1J5N 22 K 0.43 [40] 1J5N 23 R 0.63 1J5N 28 Y 0.90 1J5N 29 M 0.46 1J5N 33 N 0.74 1J5N 36 R 0.84 1J5N 40 R 0.72 1J5N 48 F 0.40 1J5N 53 K 0.49 1J5N 54 K 0.00 1J5N 58 K 0.00 1J5N 60 K 0.40 1J5N 67 K 0.25 38 FCUP Study of Protein-nucleic acid Complexes 1J5N 78 K 0.43 1J5N 81 Y 0.21 1J5N 88 Y 0.36 1JMC 238 F 0.20 [45] 1JMC 361 W 1.09 1JMC 234 R 2.15 [47] 1JMC 277 E 1.36 1JMC 382 R 1.91 1TN9 15 T 0.08 [53] 1URN 51 M 0.54 [55] 1URN 54 Q 4.85 [58] 1URN 56 F 3.23 2ERR 120 H 2.98 [60] 2ERR 126 F 4.31 2ERR 158 F 3.87 2ERR 160 F 6.08 Table 7 - Description of the 30 residues that constitute the dataset, evidencing the respective system, PDB numeration, amino acid type and experimentally ∆∆Gbinding. 2.4 - System analysis 2.4.1 – Root Mean Square Deviation The first step of the analysis was the determination of the RMSD (Root Mean Square Deviation) which to infer the stability of the complexes and monomers along the MD simulation. To perform this calculation, it was used the PTRAJ program which is included in AMBER9 [69] package. FCUP Study of Protein-nucleic acid Complexes 39 2.4.2 – Radial Distribution Function The measurement of the Radial Distribution Function allows the characterization of the interaction between the solute and the solvent molecules. The water molecules around the protein complexes can be divided in three categories: (i) Water molecules involving the protein structure, which are free to move themselves and assist in the protein diffusion comparing to the other molecules by moving casually in the solution; (ii) hydration water molecules on the protein surface; and (iii) individual water molecules connected to each other and forming hydrogen bridges with charged and polar residues, stabilizing the protein structure. With this, is possible to get the density of the solvent particles that are at a distance r from the solute particles. 2.4.3 – Solvent Accessible Surface Area SASA is the acronym to Solvent accessible Surface area and is a way of quantifying hydrophobic burial, by others words it describes the area around the protein on which is possible to occur interactions with the solvent. Therefore, SASA is the solvent-accessible surface area that was estimated using the MSMS algorithm with probe radius of 1.4 Å. In house scripts were used in the VMD to perform this calculation. For all residues SASA calculations were done with the objective of getting to know the importance of water molecules around HS and NS. These calculations were done in the last 2ns of the MD in explicit solvent. In each case we have calculated SASA for the complex (SASAcpx) and the monomer (SASAmon). ∆SASA and relSASA were also calculated as relSASA allows the differentiation of residues with equal ∆SASA but different solvent exposure; this was done according to equations 15 and 16. 40 FCUP Study of Protein-nucleic acid Complexes 3 - Results and Discussion In this work were studied 30 mutations in 6 systems, statistically the different amino acids can be distributed: Arg (17%), Lys (24%), Glu (3%), Tyr (10%), Asn (3%), Thr (3%), Met (7%), Phe (20%), Trp (3%), Gln (7%) and His (3%). In this group 47% are charged, 23% polar and 30% non-polar. Two HS considerations were made: (i) when the minimum free binding energy upon alanine mutation was considered 2.0 kcal/mol: we had 23% of HS and 77% of NS and (ii) when the minimum free binding energy value upon alanine mutation was considered 1.0 kcal/mol we had 37% as HS and 63% os NS. All residue choices are limited to the existence of experimental free binding energy values upon alanine mutation in the protein-based interfaces studied. 3.1 - RMSD RMSD profiles were calculated for each of the systems, considering separately the protein and the nucleic acid contribution, to assure their equilibration throughout the MD simulation. All six complexes were stable throughout the MD simulation with variations lower than 2 Å in the DNA-based complexes and 3 Å in the RNA-based complexes. FCUP Study of Protein-nucleic acid Complexes 41 48 FCUP Study of Protein-nucleic acid Complexes 1URN M 0.5400 -42.47 ± 4.29 -0.68 ± 0.07 Q 4.8500 -30.24 ± 4.88 -0.87 ± 0.04 F 3.2300 -63.25 ± 5.71 -0.92 ± 0.03 2ERR H 2.9800 -31.58 ± 5.58 -0.70 ± 0.12 F 4.3100 -111.90 ± 6.38 -0.75 ± 0.07 F 3.8700 -53.51 ± 7.09 -0.85 ± 0.08 F 6.0800 -51.06 ± 7.74 -0.96 ± 0.04 Table 10 – Results for ∆SASA and relSASA for all residues with known ∆∆Gbinding for the studied complexes. Since the type of amino acid plays a crucial role in the definition of the interface, they were grouped according to their chemical character: charged (Glu, His, Lys and Arg); Polar (Thr, Asn, Gln and Tyr) and nonpolar (Met, Phe and Trp). Figure 12 - Graphical Representation of average ∆SASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. Figure 13 - Graphical Representation of average relSASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. 0,00 20,00 40,00 60,00 80,00 All Charged Polar Non Polar ∆ SASA / A2 Hot-Spots Null-Spots 0 0,2 0,4 0,6 0,8 1 All Charged Polar Non Polar rel SASA Hot-Spots Null-Spots FCUP Study of Protein-nucleic acid Complexes 49 For an easier analysis of the results presented in Table 10, they were plotted in 2 graphics and shown in Figures 11 and 12. The SASA analysis by itself is insufficient to make a clear distinction between HS and NS as we can see in Figure 11, (both HS and NS have high values of ∆SASA). The average value for HS is 56.39 ± 6.40 Å2 (∆SASA) and 0.81 ± 0.07 (relSASA) and 52.07 ± 7.15 Å2 (∆SASA) and 0.48 ± 0.06 (relSASA) for NS. For this 2.0 kcal/mol cut-off all three categories of NS have higher values of ∆SASA than the respective ones of HS despite the total average being higher for the HS. The higher difference is in the charged group as NS have a ∆SASA value of 52.10 ± 7.49 Å2, 10 Å2 more than the HS average, while in the polar and non-polar group the differences are 2 Å2 and 5 Å2 respectively. RelSASA values don’t follow this tendency and are higher for all the considered HS groups and reach their maximum difference in the polar group where HS have 0.87 ± 0.04 and NS only 0.41 ± 0.07. Figure 14 - Graphical Representation of average ∆SASA values for each of the amino acid groups for the 1.0 kcal/mol cut-off. Figure 15 - Graphical Representation of average relSASA values for each of the amino acid groups for the 1.0 kcal/mol cut-off. 0,00 20,00 40,00 60,00 80,00 100,00 All Charged Polar Non Polar ∆ SASA / A2 Hot-Spots Null-Spots 0 0,2 0,4 0,6 0,8 1 All Charged Polar Non Polar rel SASA Hot-Spots Null-Spots 50 FCUP Study of Protein-nucleic acid Complexes The results for the 1.0 kcal/mol cut-off are plotted in Figures 13 and 14. In this case the HS of the charged group has a much lower value of ∆SASA (60.02 ± 7.38 Å2) while the NS average is only 45.53 ± 7.33 Å2 making a difference of roughly 15 Å2. The polar group remains very close with a difference of only 2 Å2 and the non-polar group is the one with the bigger difference in ∆SASA values, 91.54 ± 7.63 Å2 for NS and 58.49 ± 6.15 Å2 for HS. For the relSASA values is notorious a big difference (double) between HS and NS where HS have a higher value but when comparing the non-polar group the value is the same, 0.81 ± 0.08 for HS and 0.81 ± 0.04 for NS. To try to understand the possible differences between DNA and RNA-based complexes, it was made a separately analysis of ∆SASA and relSASA as well as for the 2.0 kcal/mol and 1.0 kcal/mol cut-offs. Figure 16 - Graphical Representation of average ∆SASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. Figure 17 - Graphical Representation of average relSASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. 0,00 20,00 40,00 60,00 80,00 100,00 All Charged Polar Non Polar ∆ SASA / A2 Hot-Spots Null-Spots 0 0,2 0,4 0,6 0,8 1 All Charged Polar Non Polar rel SASA Hot-Spots Null-Spots FCUP Study of Protein-nucleic acid Complexes 51 The results for the DNA-based complexes 2.0 kcal/mol cut-off are plotted in Figures 15 and 16. This separate analysis is poorer than the global one due to the lack of HS in the polar and non-polar group remaining only the charged group. In this group we can observe that HS have a slightly higher value of ∆SASA than NS and a bigger difference for relSASA 53.19 ± 7.39 Å2; 0.66 ± 0.09 (HS) and 52.10 ± 7.49 Å2; 0.39 ± 0.06 (NS). Figure 18 - Graphical Representation of average ∆SASA values for each of the amino acid groups for the 1.0 kcal/moll cut-off. Figure 19 - Graphical Representation of average relSASA values for each of the amino acid groups for the 1.0 kcal/mol cut-off. The results for the 1.0 kcal/mol cut-off are plotted in Figures 17 and 18. Once again there are no HS in the polar group so the comparison can only be made for the charged and non-polar groups. In the Charged group HS have a higher value of ∆SASA and relSASA, 67.14 ± 7.84 Å2; 0.65 ± 0.08 (HS) and 45.53 ± 7.33 Å2; 0.30 ± 0.05 (NS) but in the non-polar group is the other way around and the difference is massive in favour of NS, 12.72 ± 3.8 Å2; 0.57 ± 0.16 (HS) and 107.9 ± 8.74 Å2; 0.85 ± 0,00 20,00 40,00 60,00 80,00 100,00 120,00 All Charged Polar Non Polar ∆ SASA / A2 Hot-Spots Null-Spots 0 0,2 0,4 0,6 0,8 1 All Charged Polar Non Polar rel SASA Hot-Spots Null-Spots 52 FCUP Study of Protein-nucleic acid Complexes 0.03 (NS). This huge result is not normal, but can be explained with the small number of non-polar residues at these interfaces and the size of these complexes, making them vulnerable to the solvent action. Figure 20 - Graphical Representation of average ∆SASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. Figure 21 - Graphical Representation of average rel SASA values for each of the amino acid groups for the 2.0 kcal/mol cut-off. 0,00 20,00 40,00 60,00 80,00 All Charged Polar Non Polar ∆ SASA / A2 Hot-Spots Null-Spots 0 0,2 0,4 0,6 0,8 1 All Charged Polar Non Polar rel SASA Hot-Spots Null-Spots FCUP Study of Protein-nucleic acid Complexes 53 The results for ∆SASA and relSASA in the RNA-based complexes for the 2.0 kcal/mol cut-off are plotted in Figures 19 and 20. If in the DNA-based complexes we lacked HS, in RNA-based complexes we lack NS, and so in this cut-off the only possible comparison is the non-polar group which has an opposite result when compared with the previous analysis. In this non-polar group HS have a higher value of ∆SASA and relSASA, 69.93 ± 6.73 Å2; 0.87 ± 0.06 (HS) and 42.47 ± 4.29 Å2; 0.68 ± 0.07 (NS). For the 1.0 kcal/mol cut-off is an analysis is not required because there is no change in the ∆SASA and relSASA values. 3.4 - Mutagenesis in Protein acid nucleic interfaces We have applied the ASM methodology first in its original formulation, using the LPB equation to calculate the Gpolar solvation term. For that we calculated the ∆∆Gbinding value for a set of 25 structures generated from the last 2 ns of MD simulations in explicit solvent. Throughout this work we performed the various calculations with dielectric constants ranged from ɛ1 to ɛ9 to access the importance of the dielectric constant in the determination. As the results were not accurate enough we will not present them in this thesis. Then we used the NLPB equation instead of the LPB implemented in the DelPhi program for the calculation of the binding free energy upon alanine mutation. This was done because it was proved that using the LPB equation is not the most appropriate method for dealing with highly charged systems such as the Protein-nucleic acid systems in study. Having this in mind only the results after the use of the NLPB were treated and are shown in Table 11(from ɛ1 to ɛ4) and Table 12(from ɛ5 to ɛ9). We also have to stress out that instead of presenting ∆∆Gbinding, we will present ∆∆Gpolar solv + ∆∆ɛelectric as it show to correlate more accurately with the experimental values. 54 FCUP Study of Protein-nucleic acid Complexes AA Muted ∆∆Gexp / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ɛ1 ɛ2 ɛ3 ɛ4 1ECR R 1,2000 -8,08 0,31 3,02 4,32 Q 0,1100 -3,84 -1,89 -1,24 -0,91 1J5N K 0,4300 3,33 5,11 5,69 5,98 R 0,6300 -8,25 1,51 4,58 6,03 Y 0,9000 -1,16 -0,56 -0,30 -0,16 M 0,4600 -1,77 -0,88 -0,57 -0,42 N 0,7400 -3,25 -1,32 -0,68 -0,37 R 0,8400 2,23 4,86 5,72 6,14 R 0,7200 -5,43 1,09 3,21 4,24 F 0,4000 -4,26 -2,10 -1,37 -1,01 K 0,4900 -6,96 0,08 2,42 3,58 K 0,0000 -0,11 2,21 3,00 3,40 K 0,0000 -0,32 1,93 2,70 3,09 K 0,4000 -3,01 1,62 3,15 3,92 K 0,2500 0,43 3,57 4,59 5,09 K 0,4300 -2,36 2,63 4,25 5,06 Y 0,2100 -0,86 -0,32 -0,16 -0,08 Y 0,3600 -0,99 -0,40 -0,21 -0,12 1JMC R 2,1500 7,14 6,51 5,95 5,56 F 0,1900 -2,70 -1,31 -0,84 -0,61 E 1,3600 1,34 -0,26 -0,86 -1,19 W 1,0900 -0,50 -0,10 0,05 0,11 R 1,9100 6,84 6,63 6,46 6,29 1TN9 T 0,0800 -3,96 -1,95 -1,28 -0,94 1URN M 0,5400 -2,78 -1,19 -0,70 -0,47 Q 4,8500 5,91 2,85 1,82 1,31 F 3,2300 -2,68 -1,16 -0,71 -0,49 2ERR H 2,9800 -4,88 -2,30 -1,46 -1,05 F 4,3100 -5,40 -2,60 -1,68 -1,22 F 3,8700 -5,77 -2,69 -1,68 -1,19 F 6,0800 -2,11 -1,01 -0,65 -0,47 Table 11 - MM-PBSA results for ɛ1 to ɛ4 FCUP Study of Protein-nucleic acid Complexes 55 AA Muted ∆∆Gexp / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ɛ5 ɛ6 ɛ7 ɛ8 ɛ9 1ECR R 1,2000 5,08 5,56 5,90 6,14 6,31 Q 0,1100 -0,71 -0,58 -0,49 -0,42 -0,37 1J5N K 0,4300 6,16 6,27 37,06 6,40 6,45 R 0,6300 6,85 7,35 7,69 7,93 8,10 Y 0,9000 -0,08 -0,03 0,01 0,03 0,05 M 0,4600 -0,33 -0,27 -0,23 -0,20 -0,17 N 0,7400 -0,19 -0,07 0,02 0,08 0,12 R 0,8400 6,38 6,54 6,65 6,73 6,79 R 0,7200 4,84 5,22 5,48 5,67 5,81 F 0,4000 -0,79 -0,64 -0,54 -0,46 -0,40 K 0,0000 3,65 3,82 3,95 4,05 4,13 K 0,0000 3,34 3,51 3,63 3,73 3,81 K 0,4000 4,37 4,68 4,89 5,05 5,17 K 0,2500 5,38 5,56 5,69 5,78 5,84 K 0,4300 5,53 5,85 6,06 6,22 6,35 Y 0,2100 -0,03 0,00 0,02 0,03 0,04 Y 0,3600 -0,07 -0,03 -0,01 0,01 0,02 1JMC R 2,1500 5,27 5,03 4,82 4,64 4,49 F 0,1900 -0,48 -0,38 -0,32 -0,27 -0,23 E 1,3600 -1,39 -1,54 -1,64 -1,72 -1,78 W 1,0900 0,14 0,16 0,17 0,18 0,18 R 1,9100 6,13 5,99 5,86 5,74 5,64 1TN9 T 0,0800 -0,57 -0,47 -0,39 -0,34 -0,29 1URN M 0,5400 -0,34 -0,26 -0,21 -0,17 -0,14 Q 4,8500 1,01 0,81 0,67 0,56 0,48 F 3,2300 -0,37 -0,29 -0,24 -0,20 -0,17 2ERR H 2,9800 -0,81 -0,66 -0,55 -0,47 -0,41 F 4,3100 -0,95 -0,77 -0,65 -0,55 -0,48 F 3,8700 -0,86 -0,71 -0,58 -0,49 -0,41 F 6,0800 -0,36 -0,29 -0,24 -0,20 -0,18 Table 12 - MM-PBSA results for ɛ5 to ɛ9 The average errors of the calculated values for ∆∆Gpolar solv + ∆∆ɛele energy are shown in Table 13, where they were separated in the three: charged, polar and nonpolar. This alone is not enough to infer about the applicability of the method and its accuracy, so a statistical analysis is also necessary and will be analysed by a set of tests: (i) F1 score (equation 17) defined as a function of Precision (P, equation 18), which indicates the reliability of the predictions and the outcome of alanine mutations; and (ii) Recall which is related with the number of HS correctly predicted and therefore is crucial in these studies (R, equation 19). TP stands for true positive (predicted HS 56 FCUP Study of Protein-nucleic acid Complexes that are actual HS) and FP stands for false positive (predicted HS that are not an actual HS), TN stands for true negative (predicted NS that are actual NS) and FN stands for false negative (predicted NS that are not actual NS). Specificity (equation 20) is another measure of performance, especially for NS. F1 and Accuracy (equation 21) give the overall performance of the methods, so the ideal method would have these values as close to 100% as possible. For a better display and discussion, this statistical analysis will be presented separately for the 2.0 kcal/moll cut-off and 1.0 kcal/mol cut-off. |∆∆GMM-PBSA - ∆∆Gexp| kcal/mol 1 2 3 4 5 6 7 8 9 All 3,96 2,72 2,92 3,04 3,10 3,14 4,19 3,18 3,19 Charged 3,45 2,49 3,42 3,92 4,21 4,39 6,70 4,59 4,65 Polar 3,17 2,02 1,72 1,57 1,46 1,41 1,37 1,34 1,32 Non-polar 5,38 3,72 3,18 2,91 2,75 2,66 2,59 2,53 2,49 Table 13 - Results of average errors obtained with the NLPB equation. In Table 13 are displayed the average errors obtained for each dielectric constant, globally and for each group of residues. The average error value goes from 1.32 kcal/mol with ɛ9 for the polar group to 5.38 kcal/mol with ɛ1 for the non-polar group (the total average of error average is 3.08 kcal/mol). Having these error values in consideration, the best dielectric constant overall is ɛ2 with an average error of 2.72 kcal/mol; for the charged groups, ɛ2 is also the best with an average error of 2.49 kcal/mol; for the polar and non-polar groups the best dielectric constant is also ɛ2 with average errors of 2.02 kcal/mol and 3.72 kcal/mol respectively. Although ɛ9 values FCUP Study of Protein-nucleic acid Complexes 57 were better (1.32 kcal/mol and 2.49 kcal/mol respectively), they were missing a physical explanation and therefore were excluded. Statistical tests/all ɛ1 ɛ 2 ɛ 3 ɛ 4 ɛ 5 ɛ 6 ɛ 7 ɛ 8 ɛ 9 P 33.3 41.7 8.3 8.3 8.3 8.3 8.3 8.3 8.3 R 100.0 71.4 14.3 14.3 14.3 14.3 14.3 14.3 14.3 F1 50.0 52.6 10.5 10.5 10.5 10.5 10.5 10.5 10.5 Accuracy 53.3 70.0 43.3 43.3 43.3 43.3 43.3 43.3 43.3 Specificity 39.1 69.6 52.2 52.2 52.2 52.2 52.2 52.2 52.2 Statistical tests/charged ɛ1 ɛ 2 ɛ 3 ɛ 4 ɛ 5 ɛ 6 ɛ 7 ɛ 8 ɛ 9 P 20.0 25.0 8.3 8.3 8.3 8.3 8.3 8.3 8.3 R 100.0 100.0 50.0 50.0 50.0 50.0 50.0 50.0 50.0 F1 33.3 40.0 14.3 14.3 14.3 14.3 14.3 14.3 14.3 Accuracy 42.9 57.1 14.3 14.3 14.3 14.3 14.3 14.3 14.3 Specificity 33.3 50.0 8.3 8.3 8.3 8.3 8.3 8.3 8.3 Statistical tests/polar ɛ1 ɛ 2 ɛ 3 ɛ 4 ɛ 5 ɛ 6 ɛ 7 ɛ 8 ɛ 9 P 25.0 100.0 - - - - - - - R 100.0 100.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 F1 40.0 100.0 - - - - - - - Accuracy 57.1 100.0 85.7 85.7 85.7 85.7 85.7 85.7 85.7 Specificity 50.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 Statistical tests/non-polar ɛ1 ɛ 2 ɛ 3 ɛ 4 ɛ 5 ɛ 6 ɛ 7 ɛ 8 ɛ 9 P 57.1 66.7 - - - - - - - R 100.0 50.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 F1 72.7 57.1 - - - - - - - Accuracy 66.7 66.7 55.6 55.6 55.6 55.6 55.6 55.6 55.6 Specificity 40.0 80.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 Table 14 - Results of Statistical tests for the 2.0 kcal/mol cut-off In Table 14 are presented the statistical results for the 2.0 kcal/mol cut-off. For an easier reading, the discussion will be made in terms of HS prediction accuracy overall and for each of the groups considered (charged, polar and non-polar), for each dielectric constant. Overall, there are 7HS for a total of 30 residues, where 1HS belongs to DNA-based complexes and the other 6 to RNA-based complexes. The charged group has 14 residues, the polar group 7 and the non-polar has 9. 64 FCUP Study of Protein-nucleic acid Complexes predicted for all dielectric constants was a tryptophan (Trp) which is also the only one of the data set. ɛ1 showed the bigger number of FP again and the ɛ2 the lower but ɛ2 failed to predict 4/11 HS, where 3 of them belong to DNA-based complexes. In this cutoff HS from RNA-based complexes were correctly predicted with an accuracy of 100% for ɛ1, ɛ2 (83%), ɛ3 (67%), ɛ4 (67%), ɛ5 (16%) and falling to 0% from ɛ6 to ɛ9. Considering the ɛ2, it was once again the one with the least FP, only 6 (again half from ɛ1). HS prediction was better in RNA-based complexes with 5/6 but worse in DNAbased complexes with 2/4 correctly predicted HS. Considering the ɛ9, which is statistically identical to ɛ6, ɛ7 and ɛ8 with the same exact number of TP, FN, TP and FN, all failing to identify a single HS in the RNA-based complexes despite the 80% (4/5) HS prediction in DNA-based complexes. The major problem for both cut-offs was the charged group of amino acids which has 100% of the FP off all data set when considering ɛ3 to ɛ9. In the other groups there is only 1 FP in ɛ2 in the non-polar group and 6 in ɛ1 distributed 3 each in the polar and non-polar group. Figure 22 - Schematic representation of the final method formulation for the determination of ∆∆Gbinding Wild-type structure MD simulation with explicit solvent representation ASM calculation and analysis using electrostatic and polar solvation terms Application of the NLPB equation ɛ2 for all amino acids ∆∆Gbinding FCUP Study of Protein-nucleic acid Complexes 65 References 1. Janin, J., ELUSIVE AFFINITIES. Proteins-Structure Function and Genetics, 1995. 21(1): p. 30-39. 2. Jones, S. and J.M. Thornton, Principles of protein-protein interactions. Proceedings of the National Academy of Sciences of the United States of America, 1996. 93(1): p. 13-20. 3. Chothia, C. and J. Janin, PRINCIPLES OF PROTEIN-PROTEIN RECOGNITION. Nature, 1975. 256(5520): p. 705-708. 4. Clackson, T., et al., Structural and functional analysis of the 1 : 1 growth hormone : receptor complex reveals the molecular basis for receptor affinity. Journal of Molecular Biology, 1998. 277(5): p. 1111-1128. 5. DeLano, W.L., et al., Convergent solutions to binding at a protein-protein interface. Science, 2000. 287(5456): p. 1279-1283. 6. DeLano, W.L., Unraveling hot spots in binding interfaces: progress and challenges. Current Opinion in Structural Biology, 2002. 12(1): p. 14-20. 7. Thorn, K.S. and A.A. Bogan, ASEdb: a database of alanine mutations and their effects on the free energy of binding in protein interactions. Bioinformatics, 2001. 17(3): p. 284-285. 8. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Hot spots-A review of the protein-protein interface determinant amino-acid residues. Proteins-Structure Function and Bioinformatics, 2007. 68(4): p. 803-812. 9. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Unraveling the importance of protein-protein interaction: Application of a computational alanine-scanning mutagenesis to the study of the IgG1 streptococcal protein G (C2 fragment) complex. Journal of Physical Chemistry B, 2006. 110(22): p. 10962-10969. 10. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Detailed microscopic study of the full ZipA : FtsZ interface. Proteins-Structure Function and Bioinformatics, 2006. 63(4): p. 811-821. 11. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Computational alanine scanning mutagenesis - An improved methodological approach. Journal of Computational Chemistry, 2007. 28(3): p. 644-654. 12. Huo, S., I. Massova, and P.A. Kollman, Computational alanine scanning of the 1 : 1 human growth hormone-receptor complex. Journal of Computational Chemistry, 2002. 23(1): p. 15-27. 13. Massova, I. and P.A. Kollman, Computational alanine scanning to probe protein-protein interactions: A novel approach to evaluate binding free energies. Journal of the American Chemical Society, 1999. 121(36): p. 8133-8143. 14. Chakrabarti, P. and J. Janin, Dissecting protein-protein recognition sites. Proteins-Structure Function and Genetics, 2002. 47(3): p. 334-343. 15. Bahadur, R.P., et al., Dissecting subunit interfaces in homodimeric proteins. Proteins-Structure Function and Genetics, 2003. 53(3): p. 708-719. 16. Guharoy, M. and P. Chakrabarti, Conservation and relative importance of residues across protein-protein interfaces. Proceedings of the National Academy of Sciences of the United States of America, 2005. 102(43): p. 1544715452. 17. Bogan, A.A. and K.S. Thorn, Anatomy of hot spots in protein interfaces. J Mol Biol, 1998. 280(1): p. 1-9. 18. Li, J. and Q. Liu, 'Double water exclusion': a hypothesis refining the O-ring theory for the hot spots at protein interfaces. Bioinformatics, 2009. 25(6): p. 743-750. 66 FCUP Study of Protein-nucleic acid Complexes 19. Kosloff, M., et al., Integrating energy calculations with functional assays to decipher the specificity of G protein-RGS protein interactions. Nature Structural & Molecular Biology, 2011. 18(7): p. 846-U128. 20. Rajamani, D., et al., Anchor residues in protein-protein interactions. Proceedings of the National Academy of Sciences of the United States of America, 2004. 101(31): p. 11287-11292. 21. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Hot spot occlusion from bulk water: A comprehensive study of the complex between the lysozyme HEL and the antibody FVD1.3. Journal of Physical Chemistry B, 2007. 111(10): p. 26972706. 22. Ahmad, S., et al., Protein-DNA interactions: structural, thermodynamic and clustering patterns of conserved residues in DNA-binding proteins. Nucleic Acids Research, 2008. 36(18): p. 5922-5932. 23. Bradshaw, R.T., et al., Comparing experimental and computational alanine scanning techniques for probing a prototypical protein-protein interaction. Protein Eng Des Sel, 2011. 24(1-2): p. 197-207. 24. Zerbe, B.S., et al., Relationship between hot spot residues and ligand binding hot spots in protein-protein interfaces. J Chem Inf Model, 2012. 52(8): p. 223644. 25. Guharoy, M., et al., PRICE (PRotein Interface Conservation and Energetics): a server for the analysis of protein-protein interfaces. J Struct Funct Genomics, 2011. 12(1): p. 33-41. 26. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Unraveling the importance of protein-protein interaction: application of a computational alanine-scanning mutagenesis to the study of the IgG1 streptococcal protein G (C2 fragment) complex. J Phys Chem B, 2006. 110(22): p. 10962-9. 27. Law, M.J., et al., The role of positively charged amino acids and electrostatic interactions in the complex of U1A protein and U1 hairpin II RNA. Nucleic Acids Res, 2006. 34(1): p. 275-85. 28. Berman, H.M., et al., The Protein Data Bank. Nucleic Acids Res, 2000. 28(1): p. 235-42. 29. Kumar, M.D., et al., ProTherm and ProNIT: thermodynamic databases for proteins and protein-nucleic acid interactions. Nucleic Acids Res, 2006. 34(Database issue): p. D204-6. 30. Prabakaran, P., et al., Thermodynamic database for protein-nucleic acid interactions (ProNIT). Bioinformatics, 2001. 17(11): p. 1027-34. 31. Sarai, A., [Thermodynamic databases for proteins and interactions]. Tanpakushitsu Kakusan Koso, 2002. 47(8 Suppl): p. 1071-5. 32. Hoffman, M.M., et al., AANT: the Amino Acid-Nucleotide Interaction Database. Nucleic Acids Res, 2004. 32(Database issue): p. D174-81. 33. Luscombe, N.M., et al., An overview of the structures of protein-DNA complexes. Genome Biol, 2000. 1(1): p. REVIEWS001. 34. Norambuena, T. and F. Melo, The Protein-DNA Interface database. BMC Bioinformatics, 2010. 11: p. 262. 35. Kamada, K., et al., Structure of a replication-terminator protein complexed with DNA. Nature, 1996. 383(6601): p. 598-603. 36. Kaplan, D.L. and D. Bastia, Mechanisms of polar arrest of a replication fork. Mol Microbiol, 2009. 72(2): p. 279-85. 37. Bastia, D., et al., Replication termination mechanism as revealed by Tusmediated polar arrest of a sliding helicase. Proc Natl Acad Sci U S A, 2008. 105(35): p. 12831-6. 38. Neylon, C., et al., Interaction of the Escherichia coli replication terminator protein (Tus) with DNA: a model derived from DNA-binding studies of mutant FCUP Study of Protein-nucleic acid Complexes 67 proteins by surface plasmon resonance. Biochemistry, 2000. 39(39): p. 1198999. 39. Masse, J.E., et al., The S. cerevisiae architectural HMGB protein NHP6A complexed with DNA: DNA and protein conformational changes upon binding. J Mol Biol, 2002. 323(2): p. 263-84. 40. Allain, F.H., et al., Solution structure of the HMG protein NHP6A and its interaction with DNA reveals the structural determinants for non-sequencespecific binding. EMBO J, 1999. 18(9): p. 2563-79. 41. Coats, J.E., et al., Single-molecule FRET analysis of DNA binding and bending by yeast HMGB protein Nhp6A. Nucleic Acids Res, 2013. 41(2): p. 1372-81. 42. Zhang, J., et al., Basic N-terminus of yeast Nhp6A regulates the mechanism of its DNA flexibility enhancement. J Mol Biol, 2012. 416(1): p. 10-20. 43. McCauley, M., et al., Dual binding modes for an HMG domain from human HMGB2 on DNA. Biophys J, 2005. 89(1): p. 353-64. 44. Bochkarev, A., et al., Structure of the single-stranded-DNA-binding domain of replication protein A bound to DNA. Nature, 1997. 385(6612): p. 176-81. 45. Walther, A.P., et al., Replication protein A interactions with DNA. 1. Functions of the DNA-binding and zinc-finger domains of the 70-kDa subunit. Biochemistry, 1999. 38(13): p. 3963-73. 46. Brosey, C.A., et al., A new structural framework for integrating replication protein A into DNA processing machinery. Nucleic Acids Res, 2013. 47. Wyka, I.M., et al., Replication protein A interactions with DNA: differential binding of the core domains and analysis of the DNA interaction surface. Biochemistry, 2003. 42(44): p. 12909-18. 48. Ozawa, K., et al., Mapping of the 70 kDa, 34 kDa, and 11 kDa subunit genes of the human multimeric single-stranded DNA binding protein (hSSB/RPA) to chromosome bands 17p13, 1p35-p36.1, and 7p21-p22. Cell Struct Funct, 1993. 18(4): p. 221-30. 49. Wojciak, J.M., K.M. Connolly, and R.T. Clubb, NMR structure of the Tn916 integrase-DNA complex. Nat Struct Biol, 1999. 6(4): p. 366-73. 50. Yao, X.X., et al., Molecular dynamics study of DNA binding by INT-DBD under polarized force field. J Comput Chem, 2013. 51. Rajeev, L., K. Malanowska, and J.F. Gardner, Challenging a paradigm: the role of DNA homology in tyrosine recombinase reactions. Microbiol Mol Biol Rev, 2009. 73(2): p. 300-9. 52. Connolly, K.M., M. Iwahara, and R.T. Clubb, Xis protein binding to the left arm stimulates excision of conjugative transposon Tn916. J Bacteriol, 2002. 184(8): p. 2088-99. 53. Connolly, K.M., et al., Major groove recognition by three-stranded beta-sheets: affinity determinants and conserved structural features. J Mol Biol, 2000. 300(4): p. 841-56. 54. Gorfe, A.A., A. Caflisch, and I. Jelesarov, The role of flexibility and hydration on the sequence-specific DNA recognition by the Tn916 integrase protein: a molecular dynamics analysis. J Mol Recognit, 2004. 17(2): p. 120-31. 55. Katsamba, P.S., et al., Complex role of the beta 2-beta 3 loop in the interaction of U1A with U1 hairpin II RNA. J Biol Chem, 2002. 277(36): p. 33267-74. 56. Katsamba, P.S., D.G. Myszka, and I.A. Laird-Offringa, Two functionally distinct steps mediate high affinity binding of U1A protein to U1 hairpin II RNA. J Biol Chem, 2001. 276(24): p. 21476-81. 57. Showalter, S.A. and K.B. Hall, Altering the RNA-binding mode of the U1A RBD1 protein. J Mol Biol, 2004. 335(2): p. 465-80. 58. Law, M.J., et al., Kinetic analysis of the role of the tyrosine 13, phenylalanine 56 and glutamine 54 network in the U1A/U1 hairpin II interaction. Nucleic Acids Res, 2005. 33(9): p. 2917-28. 68 FCUP Study of Protein-nucleic acid Complexes 59. Oubridge, C., et al., Crystal structure at 1.92 A resolution of the RNA-binding domain of the U1A spliceosomal protein complexed with an RNA hairpin. Nature, 1994. 372(6505): p. 432-8. 60. Auweter, S.D., et al., Molecular basis of RNA recognition by the human alternative splicing factor Fox-1. EMBO J, 2006. 25(1): p. 163-73. 61. Kuroyanagi, H., Fox-1 family of RNA-binding proteins. Cell Mol Life Sci, 2009. 66(24): p. 3895-907. 62. Fogel, B.L., et al., RBFOX1 regulates both splicing and transcriptional networks in human neuronal development. Hum Mol Genet, 2012. 21(19): p. 4171-86. 63. Bitel, C.L., et al., Evidence that "brain-specific" FOX-1, FOX-2, and nPTB alternatively spliced isoforms are produced in the lens. Curr Eye Res, 2011. 36(4): p. 321-7. 64. Damianov, A. and D.L. Black, Autoregulation of Fox protein expression to produce dominant negative splicing factors. RNA, 2010. 16(2): p. 405-16. 65. Allinger, N.L., M.T. Tribble, and Y. Yuh, Androsterone. The structure by forcefield calculations. Steroids, 1975. 26(4): p. 398-406. 66. Stewart, E.L., et al., Molecular Mechanics (MM3) Calculations on OxygenContaining Phosphorus (Coordination IV) Compounds. J Org Chem, 1999. 64(15): p. 5350-5360. 67. Chen, K.H., et al., Molecular mechanics (MM4) study of amines. J Comput Chem, 2007. 28(15): p. 2391-412. 68. Brooks, B.R., et al., CHARMM: the biomolecular simulation program. J Comput Chem, 2009. 30(10): p. 1545-614. 69. Cheatham, T.E., 3rd, P. Cieplak, and P.A. Kollman, A modified version of the Cornell et al. force field with improved sugar pucker phases and helical repeat. J Biomol Struct Dyn, 1999. 16(4): p. 845-62. 70. Schyman, P. and W.L. Jorgensen, Exploring Adsorption of Water and Ions on Carbon Surfaces using a Polarizable Force Field. J Phys Chem Lett, 2013. 4(3): p. 468-474. 71. Pappalardo, M., et al., Free energy perturbation and molecular dynamics calculations of copper binding to azurin. J Comput Chem, 2003. 24(6): p. 77985. 72. Gruden-Pavlovic, M., S. Grubisic, and S.R. Niketic, Conformational analysis of octaand tetrabromo tetraphenylporphyrins and their Ni(II) and Tb(III) complexes. J Inorg Biochem, 2004. 98(8): p. 1293-302. 73. Kleinjung, J., et al., Implicit Solvation Parameters Derived from Explicit Water Forces in Large-Scale Molecular Dynamics Simulations. J Chem Theory Comput, 2012. 8(7): p. 2391-2403. 74. Wang, J., et al., Development and testing of a general amber force field. J Comput Chem, 2004. 25(9): p. 1157-74. 75. Case, D.A.P., D. A.; Caldwell, J. W.; Cheatham, T. E.; and J.R. Wang, W. S.; Simmerling, C.; Darden, T.; Merz, K. M.; Stanton, R. V.; Cheng, A.; Vincent, J. J.; Crowley, M.; Tsui, V.; Gohlke H.; Radmer, R.; Duan, Y.; Pitera J.; Massova, I.; Seibel, G. L.; Singh, U. C.; Weiner, P.; Kollman, P. , University of California: San Francisco, 2002. 76. Coleman, T.G., H.C. Mesick, and R.L. Darby, Numerical integration: a method for improving solution stability in models of the circulation. Ann Biomed Eng, 1977. 5(4): p. 322-8. 77. William L. Jorgensen, J.C., and Jeffry D. Madura Comparison of simple potential functions for simulating liquid water. AIP - The journal of chemical physics, 1983. 79. 78. Zou, H., et al., Molecular insight into the interaction between IFABP and PA by using MM-PBSA and alanine scanning methods. J Phys Chem B, 2007. 111(30): p. 9104-13. FCUP Study of Protein-nucleic acid Complexes 69 79. Huo, S., I. Massova, and P.A. Kollman, Computational alanine scanning of the 1:1 human growth hormone-receptor complex. J Comput Chem, 2002. 23(1): p. 15-27. 80. Yang, L., et al., New-generation amber united-atom force field. J Phys Chem B, 2006. 110(26): p. 13166-76. 81. Wang, W. and P.A. Kollman, Free energy calculations on dimer stability of the HIV protease using molecular dynamics and a continuum solvent model. J Mol Biol, 2000. 303(4): p. 567-82. 82. Bashford, D. and D.A. Case, Generalized born models of macromolecular solvation effects. Annu Rev Phys Chem, 2000. 51: p. 129-52. 83. Onufriev, A., D. Bashford, and D.A. Case, Exploring protein native states and large-scale conformational changes with a modified generalized born model. Proteins, 2004. 55(2): p. 383-94. 84. Luo, R., L. David, and M.K. Gilson, Accelerated Poisson-Boltzmann calculations for static and dynamic systems. J Comput Chem, 2002. 23(13): p. 1244-53. 85. Li, L., et al., DelPhi: a comprehensive suite for DelPhi software and associated resources. BMC Biophys, 2012. 5: p. 9. 86. Decherchi, S., et al., Between algorithm and model: different Molecular Surface definitions for the Poisson-Boltzmann based electrostatic characterization of biomolecules in solution. Commun Comput Phys, 2013. 13: p. 61-89. 87. Moreira, I.S., P.A. Fernandes, and M.J. Ramos, Computational alanine scanning mutagenesis--an improved methodological approach. J Comput Chem, 2007. 28(3): p. 644-54. 88. Hui Li, A.D.R., and Jan H. Jensen, Very Fast Empirical Prediction and Interpretation of Protein pKa Values. Proteins, 2005: p. 61, 704-721. . 89. Delphine C. Bas, D.M.R., and Jan H. Jensen, Very Fast Prediction and Rationalization of pKa Values for Protein-Ligand Complexes. Proteins, 2008: p. 73, 765-783. . 90. Rostkowski, M., et al., Graphical analysis of pH-dependent properties of proteins predicted using PROPKA. BMC Struct Biol, 2011. 11: p. 6. 91. Dolinsky, T.J., et al., PDB2PQR: an automated pipeline for the setup of Poisson-Boltzmann electrostatics calculations. Nucleic Acids Res, 2004. 32(Web Server issue): p. W665-7. 92. Zhang, X., et al., Application of new multi-resolution methods for the comparison of biomolecular electrostatic properties in the absence of global structural similarity. Multiscale Model Simul, 2006. 5(4): p. 1196-1213. 93. Price, D.J. and C.L. Brooks, 3rd, A modified TIP3P water potential for simulation with Ewald summation. J Chem Phys, 2004. 121(20): p. 10096-103. 94. Huggins, D.J., Correlations in liquid water for the TIP3P-Ewald, TIP4P-2005, TIP5P-Ewald, and SWM4-NDP models. J Chem Phys, 2012. 136(6): p. 064518. 95. Ribeiro JV, C.N., Moreira IS, Fernandes PA, Ramos MJ, Compasm: An ambervmd alanine scanning mutagenesis plug-in. Theoretical Chemistry Accounts - In press, 2012. 96. Ramos, R.M., L.F. Fernandes, and I.S. Moreira, Extending the applicability of the O-ring theory to protein-DNA complexes. Comput Biol Chem, 2013. 44: p. 31-9. 70 FCUP Study of Protein-nucleic acid Complexes Annexes RDF 1ECR ARG_194 GLN_246 r[Å] N H2O r[Å] N H2O 0,0500 0,0000 0,0500 0,0000 1,0500 0,0000 1,0500 0,0000 2,0500 0,0000 2,0500 0,0000 3,0500 0,2960 3,0500 0,5360 4,0500 0,9720 4,0500 1,7780 5,0500 1,8290 5,0500 3,7540 6,0500 3,3210 6,0500 7,0820 7,0500 7,0500 7,0500 12,1140 8,0500 14,2150 8,0500 19,7530 9,0500 25,5180 9,0500 30,0570 1J5N ARG_23 ARG_36 ARG_40 ASN_33 LYS_22 r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 3,0500 0,6060 3,0500 1,5470 3,0500 1,2050 3,0500 0,6520 3,0500 2,5600 4,0500 2,0660 4,0500 6,9890 4,0500 5,1090 4,0500 4,0310 4,0500 6,0830 5,0500 4,2630 5,0500 13,7540 5,0500 9,4510 5,0500 8,9340 5,0500 12,6620 6,0500 7,3160 6,0500 23,1130 6,0500 16,1530 6,0500 15,8120 6,0500 21,8940 7,0500 11,8290 7,0500 36,6490 7,0500 25,7220 7,0500 26,5360 7,0500 33,3100 8,0500 20,2540 8,0500 54,4670 8,0500 37,6770 8,0500 41,0400 8,0500 47,7640 9,0500 33,3780 9,0500 76,6920 9,0500 53,8290 9,0500 59,1880 9,0500 65,6580 1J5N LYS_53 LYS_54 LYS_58 LYS_60 LYS_67 r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 3,0500 2,4980 3,0500 2,2120 3,0500 2,4540 3,0500 2,5700 3,0500 2,7240 4,0500 5,3120 4,0500 5,9650 4,0500 5,3810 4,0500 5,8120 4,0500 4,7260 5,0500 10,4440 5,0500 11,5920 5,0500 10,2150 5,0500 12,2920 5,0500 8,3210 6,0500 18,3690 6,0500 19,7200 6,0500 17,1380 6,0500 21,8010 6,0500 14,0660 7,0500 29,1590 7,0500 30,3980 7,0500 26,1360 7,0500 34,4010 7,0500 21,2350 8,0500 44,3530 8,0500 44,4740 8,0500 38,0100 8,0500 50,8240 8,0500 31,0680 9,0500 64,7320 9,0500 63,0460 9,0500 53,4370 9,0500 71,3840 9,0500 45,2710 FCUP Study of Protein-nucleic acid Complexes 71 1J5N LYS_78 MET_29 PHE_48 TYR_28 TYR_81 r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 0,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 1,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 2,0500 0,0000 3,0500 2,1420 3,0500 0,0000 3,0500 0,0510 3,0500 0,0000 3,0500 0,0030 4,0500 5,6710 4,0500 0,0380 4,0500 2,6540 4,0500 0,0000 4,0500 0,8410 5,0500 11,8270 5,0500 0,4930 5,0500 6,6570 5,0500 0,0000 5,0500 3,2320 6,0500 19,9150 6,0500 1,7490 6,0500 12,7380 6,0500 0,0580 6,0500 6,8350 7,0500 30,2660 7,0500 4,0860 7,0500 21,7800 7,0500 0,8770 7,0500 12,4130 8,0500 42,6490 8,0500 8,8910 8,0500 33,4690 8,0500 3,9140 8,0500 21,9030 9,0500 57,3470 9,0500 17,0410 9,0500 49,5850 9,0500 10,4670 9,0500 33,9100 1J5N 1TN9 TYR_88 THR_13 r[Å] N H2O r[Å] N H2O 0,0500 0,0000 0,0500 0,0000 1,0500 0,0000 1,0500 0,0000 2,0500 0,0000 2,0500 0,0000 3,0500 0,0140 3,0500 1.4590 4,0500 1,2310 4,0500 4.3270 5,0500 3,9040 5,0500 8.9290 6,0500 7,9640 6,0500 16.2560 7,0500 14,7230 7,0500 25.7380 8,0500 23,6650 8,0500 37.7960 9,0500 37,2930 9,0500 52.8430 1JMC PHE_56 ARG_200 GLU_95 ARG_52 TRP_179 r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,05 0 0,05 0 0,05 0 0,05 0 0,05 0 1,05 0 1,05 0 1,05 0 1,05 0 1,05 0 2,05 0 2,05 0 2,05 0 2,05 0 2,05 0 3,05 0,025 3,05 0,804 3,05 0,696 3,05 0,677 3,05 0 4,05 1,406 4,05 2,124 4,05 1,556 4,05 2,9 4,05 0,001 5,05 3,472 5,05 3,743 5,05 3,212 5,05 4,631 5,05 0,123 6,05 6,262 6,05 6,656 6,05 3,916 6,05 7,136 6,05 0,956 7,05 9,142 7,05 11,712 7,05 6,476 7,05 8,648 7,05 2,375 8,05 13,34 8,05 18,762 8,05 10,26 8,05 12,542 8,05 5,291 9,05 19,526 9,05 28,982 9,05 14,921 9,05 16,908 9,05 11,45 72 FCUP Study of Protein-nucleic acid Complexes 1URn MET_50 GLN_53 PHE_55 r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,05 0 0,05 0 0,05 0 1,05 0 1,05 0 1,05 0 2,05 0 2,05 0 2,05 0 3,05 0,002 3,05 0 3,05 0 4,05 0,593 4,05 0,012 4,05 0 5,05 2,04 5,05 0,385 5,05 0,095 6,05 3,597 6,05 1,611 6,05 0,817 7,05 5,615 7,05 3,887 7,05 1,589 8,05 9,731 8,05 7,387 8,05 3,338 9,05 17,104 9,05 12,626 9,05 5,984 2ERR PHE_52 PHE_18 PHE_50 HID_12 r[Å] N H2O r[Å] N H2O r[Å] N H2O r[Å] N H2O 0,05 0 0,05 0 0,05 0 0,05 0 1,05 0 1,05 0 1,05 0 1,05 0 2,05 0 2,05 0 2,05 0 2,05 0 3,05 0,003 3,05 0,082 3,05 0,088 3,05 0,858 4,05 0,224 4,05 2,373 4,05 0,969 4,05 1,94 5,05 1,035 5,05 6,917 5,05 1,392 5,05 4,929 6,05 2,908 6,05 12,295 6,05 2,882 6,05 7,225 7,05 5,859 7,05 21,135 7,05 5,444 7,05 11,682 8,05 12,226 8,05 32,875 8,05 9,697 8,05 18,701 9,05 22,528 9,05 47,613 9,05 18,405 9,05 27,391 FCUP Study of Protein-nucleic acid Complexes 73 MM-PBSA ɛ1 ɛ2 ɛ3 AA Muted ∆∆Gexp / [kcal/mol] ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gabs ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gabs ∆∆Gpolar solv + ∆∆ɛele / [kcal/mol] ∆∆Gabs 1ECR R 1,2000 -8,08 9,28 0,31 0,89 3,02 1,82 Q 0,1100 -3,84 3,95 -1,89 2,00 -1,24 1,35 1J5N K 0,4300 3,33 2,90 5,11 4,68 5,69 5,26 R 0,6300 -8,25 8,88 1,51 0,88 4,58 3,95 Y 0,9000 -1,16 2,06 -0,56 1,46 -0,30 1,20 M 0,4600 -1,77 2,23 -0,88 1,34 -0,57 1,03 N 0,7400 -3,25 3,99 -1,32 2,06 -0,68 1,42 R 0,8400 2,23 1,39 4,86 4,02 5,72 4,88 R 0,7200 -5,43 6,15 1,09 0,37 3,21 2,49 F 0,4000 -4,26 4,66 -2,10 2,50 -1,37 1,77 K 0,0000 -0,11 0,11 2,21 2,21 3,00 3,00 K 0,0000 -0,32 0,32 1,93 1,93 2,70 2,70 K 0,4000 -3,01 3,41 1,62 1,22 3,15 2,75 K 0,2500 0,43 0,18 3,57 3,32 4,59 4,34 K 0,4300 -2,36 2,79 2,63 2,20 4,25 3,82 Y 0,2100 -0,86 1,07 -0,32 0,53 -0,16 0,37 Y 0,3600 -0,99 1,35 -0,40 0,76 -0,21 0,57 1JMC R 2,1500 7,14 4,99 6,51 4,36 5,95 3,80 F 0,1900 -2,70 2,89 -1,31 1,50 -0,84 1,03 E 1,3600 1,34 0,02 -0,26 1,62 -0,86 2,22 W 1,0900 -0,50 1,59 -0,10 1,19 0,05 1,04 R 1,9100 6,84 4,93 6,63 4,72 6,46 4,55 1TN9 T 0,0800 -3,96 4,04 -1,95 2,03 -1,28 1,36 1URN M 0,5400 -2,78 3,32 -1,19 1,73 -0,70 1,24 Q 4,8500 5,91 1,06 2,85 2,00 1,82 3,03 F 3,2300 -2,68 5,91 -1,16 4,39 -0,71 3,94 2ERR H 2,9800 -4,88 7,86 -2,30 5,28 -1,46 4,44 F 4,3100 -5,40 9,71 -2,60 6,91 -1,68 5,99 F 3,8700 -5,77 9,64 -2,69 6,56 -1,68 5,55 F 6,0800 -2,11 8,19 -1,01 7,09 -0,65 6,73 34 R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 Table 2 Description of the 112 residues that constitute our dataset, evidencing the respective system, PDB numeration, amino acid type and experimentally Gbinding. Protein ID #AA PDB #AA mutated Gbinding/ kcal mol−1 Reference 1MNM 16 K 0.54 Acton et al. (2000) 1MNM 17 E 0.24 1MNM 18 R 0.28 1MNM 20 K 0.64 1MNM 21 I 0.88 1MNM 22 E 0.20 1MNM 23 I 0.86 1MNM 24 K −0.13 1MNM 25 F 4.05 1MNM 26 I 4.05 1MNM 27 E 0.41 1MNM 28 N 1.75 1MNM 29 K −0.13 1MNM 32 R 2.76 1MNM 33 H −0.81 1MNM 34 V 0.58 1MNM 35 T 0.64 1MNM 36 F 4.05 1MNM 37 S 0.43 1MNM 38 K 4.05 1MNM 39 R 4.05 1MNM 40 K 2.83 1MNM 41 H −1.24 1MNM 43 I 3.24 1MNM 45 K 4.05 1MNM 46 K 4.05 1MNM 48 F −0.13 1MNM 49 E 2.27 1MNM 51 S 0 1MNM 52 V 0.11 1MNM 53 L 4.05 1MNM 66 T 0.64 1BDT 4 M 1.10 Brown et al. (1994) 1BDT 5 S 1.30 1BDT 6 K 1.30 1BDT 7 M 1.90 1BDT 9 Q 1.80 1BDT 11 N 2.00 1BDT 13 R 6.00 1BDT 23 R 0.20 1BDT 29 N −1.00 1BDT 31 R −0.30 1BDT 32 S −2.30 1BDT 33 V −2.30 1BDT 34 N 3.20 1BDT 35 S −0.20 1BDT 39 Q −0.40 1MSE 116 S 0.06 Oda et al. (1997, 1998, 1999) 1MSE 139 N 0.60 1MSE 141 E −0.10 1MSE 187 S 0.10 1B3T 469 R 3.41 Cruickshank et al. (2000) 1B3T 518 Y 2.62 1B3T 522 R 4.40 1QRV 9 L 0.02 Klass et al. (2003) 1QRV 13 M 1.20 1QRV 32 V −0.30 1J5N 18 P 0.00 Allain et al. (1999) 1J5N 22 K 0.43 1J5N 28 Y 0.90 1J5N 33 N 0.74 1J5N 36 R 0.84 1J5N 40 R 0.72 1J5N 53 K 0.49 1J5N 54 K 0.00 1J5N 58 K 0.00 1J5N 60 K 0.40 1J5N 67 K 0.25 1J5N 78 K 0.43 1J5N 81 Y 0.21 1J5N 85 K 0.21 1J5N 88 Y 0.36 1JMC 234 R 2.15 Walther et al. (1999) and Wyka et al. (2003) 1JMC 238 F 0.24 Table 2 (Continued) Protein ID #AA PDB #AA mutated Gbinding/ kcal mol−1 Reference 1JMC 263 K 1.01 1JMC 277 E 1.36 1JMC 382 R 1.91 1QZH 62 T 1.53 1QZH 64 D 1.46 1QZH 88 F 3.96 1QZH 91 Q 1.26 1QZH 115 Y 0.66 1QZH 122 L 1.00 1QZH 123 S −0.70 1TN9 5 R 0.78 Connolly et al. (2000) 1TN9 15 T 0.08 1TN9 18 S −0.14 1TN9 21 K 0.74 1TN9 24 R 1.25 1TN9 26 L −0.19 1TN9 42 W 0.48 1TN9 54 K 1.37 2A0I 3 S 4.15 Larkin et al. (2005) 2A0I 8 R 1.70 2A0I 19 D −0.30 2A0I 88 K 5.57 2A0I 147 D 0.30 2A0I 148 T 0.30 2A0I 149 S 3.41 2A0I 150 R 2.92 2A0I 153 E 1.70 2A0I 155 Q 2.20 2A0I 158 T 0.90 2A0I 187 E 2.10 2A0I 220 K 0.80 2A0I 221 H 2.82 2A0I 223 M 2.70 2A0I 237 R 4.39 2A0I 241 I 3.91 2A0I 242 R 1.20 2A0I 254 R 3.17 2A0I 265 K 1.80 3. Results The X-ray crystallographic structures of complexes available at the RCSB Protein Data Bank, as well as the tridimensional structures that are the basis of the initial knowledge for this work are static and do not show the conformational changes that occur in the system over time. The correct comprehension of the phenomena that occurs within the protein–DNA interface such as the structural adaptability and the right binding mode can benefit greatly from the use of computational methods capable of generating information based, not on a single structure, but on an ensemble of conformations generated by MD simulations. It also makes possible to understand the structural and functional role of water around HS and NS, since the great majority of biological processes occur in an aqueous medium. Our dataset consists of 10 protein–DNA complexes that were previously described in the methodological section and a total of 112 interfacial residues (Table 2). They have the following distribution by amino acid, group type and their hot and null-spot character: Glu (7%), Phe (4%), His (3%), Ile (4%), Lys (20%), Leu (4%), Met (4%), Asn (5%), Gln (4%), Arg (17%), Ser (10%), Thr (5%), Val (4%) and Tyr (4%); 50% are charged residues, 29% polar and 21% nonpolar; 28% are HS and 72% NS. We have a predominance of charged residues, especially positively charged amino acids such as Lys and Arg, which is consistent with previous studies within protein–DNA interfaces that characterize them as highly charged interfaces (Ahmad et al., 2008). The measure of the RDF profile of residues is a common procedure when dealing with a system formed with explicit solvent, since it allows the characterization of the interaction between the R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 35 Fig. 2. Representation of the RDF profile and average number of waters around the average HS (a, c) and the average NS (b, d). Table 3 Average number of water molecules at distances between 3 and 6˚ A of HS and NS, for each of the studied complexes. Complex 1MNM 1BDT 1MSE 1B3T 1QRV 1JN5 1JMC 1QZH 1TN9 2A0I Distance/Å HS NS HS NS HS NS HS NS HS NS HS NS HS NS HS NS HS NS HS NS 3 1.00 1.43 0.84 0.89 – 1.79 0.21 – – 0.01 – 1.50 0.68 0.29 0.00 0.43 – 0.70 0.53 1.89 4 2.97 4.21 2.23 3.12 – 4.23 1.87 – – 1.28 – 4.35 2.90 0.78 0.08 1.63 – 2.30 1.65 4.17 5 6.14 9.62 3.66 6.50 – 8.97 3.98 – – 3.24 – 9.07 4.63 1.45 0.98 3.70 – 4.69 3.75 8.63 6 10.76 17.15 5.90 11.22 – 14.76 5.93 – – 6.16 – 15.75 7.14 2.55 1.88 7.60 – 8.46 6.54 14.88 solute and the solvent molecules. Usually, a water RDF profile exhibits an oscillatory profile and a peak due to the presence of hydrogen bonds. We measured the RDFs for the 112 residues in our dataset. Fig. 2 shows two different RDF profiles that represent a HS and a NS. As it can be easily perceived, the exhibit behavior is quite different between the two. In Fig. 2b the peak located about 3˚ A is due to the strong interaction between the hydrogen atoms of water and the oxygen atoms of the carbonyl group of the NS. Some other peaks can be perceived in the remaining plot, which are less defined, due to the interaction between the water molecules and the atoms of the amino acid residue, with the exception of hydrogen. On the other hand, as can be seen in plot (a) of Fig. 2, the peaks are less defined when we are dealing with a putative HS. Plots (c) and (d) show the average number of water molecules around the HS and the NS, which are clearly distinct. We also measured the average number of water molecules around each residue, using distance cutoff values of 3, 4, 5 and 6˚ A. Table 3 summarizes the results for each of the complexes under study and Table 4 the global results in terms of HS and NS. These results were plotted in a graphic and shown in Fig. 3 for a simpler analysis. It is perceptible that the average number of water molecules around HS is notoriously lower when compared to NS. This difference for a cuttof value of 3˚ A is not significant, but as we increase the cutoff value for 4, 5 and 6˚ A the difference increases. As some individual systems do not possess both HS and NS (due to the difficulties in finding experimental binding free energy values upon alanine mutation Table 4 Average number of water molecules at distances between 3 and 6˚ A for the total of HS and NS for all the complexes. Distance/Å HS NS 3 0.70 1.13 4 2.23 3.42 5 4.63 7.24 6 7.92 12.64 for these protein–DNA complexes) we also present global results for HS and NS, which allows a better understanding of the phenomenon of occlusion. For a NS, and at a 5˚ A distance, there are an average of 7 water molecules around it; instead for an HS, the average number of water molecules decreases to 5. The average number of water molecules around a HS and a NS varies from 0.70 to 7.92 and 1.13 to 12.64, respectively. At shorter distances both HS and NS present an average number of water molecules relatively small, with less than 2 water molecules, which gradually increases with the distance. As we have previously stated, the interface is generally distinguished in a core and a rim, and the HS are usually at the core. With this in mind, and for every HS in our dataset, we selected the residues within 4˚ A and between 4 and 8˚ A distances of our HS and calculated the average number of water molecules for each of these coordination spheres. In theory this would give us an idea of Fig. 3. Representation of the global average results of water molecules around HS and NS. 36 R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 Fig. 4. Representation of the average number of waters around the HS (in brown), the 4˚ A solvation sphere (orange); and the 8˚ A solvation sphere (green). (For interpretation of the references to color in this figure legend and the text, the reader is referred to the web version of this article.) the environment surrounding each HS and act as a complement to the core and rim definition of the interface. The results obtained are shown in Fig. 4. This figure lists for all the 31 HS, the average number of waters for the HS itself (brown), for the residues inside the first coordination sphere (orange) and for the residues inside the second coordination sphere (green). It should be noted that in this figure the stacked vertical bars offer for each HS the given percentage over the three analyzed data (HS, 4˚ A sphere and between 4 and 8˚ A sphere). As expected, for the majority of the analyzed HS the average number of water molecules within the first coordination sphere is lower than in the second. The residues that closely surround the HS, and belong to the core, are more buried and have less accessibility to the solvent. On the other hand when we advance to the next coordination sphere, which can be related to the rim region, we find more residues with higher solvent accessibility that account for the difference obtained. Nevertheless, some of the residues within the two solvation spheres seem to present a different behavior. We cannot exclude that may be influenced by the existence of a HS within those spheres. As mentioned in the methodological part, the known HS were excluded but Gbinding values were not available for all the residues of the analyzed interfaces. The results we obtained, based on MD simulations of protein–DNA systems, clearly show that the average number of water molecules around a HS is much lower when compared to NS. This is in agreement with the O-ring theory and the fact that the HS are generally protected from the solvent. We also measured different SASA features for the complex and the monomers to complement the study as it allows a more profound characterization of the importance of the water molecules in the micro ambient of the HS and NS. SASA and relSASA, that were previously described, were also measured. The results we obtained are summarized in Table 5 and Fig. 5. Our results follow the tendency observed for PPI and indicate that a high value of SASA and relSASA is a necessary condition for a residue to be considered a HS. However, these features are insufficient to make a clear distinction between HS and NS, as some NS also present a high value of SASA. The type of amino acid plays a crucial role in the definition of the interface, so we have also opted to analyze the results according to the type of residue: charged (Asp, Glu, His, Lys e Arg), polar (Ser, Thr, Asn, Gln e Tyr), nonpolar (Val, Ile, Leu, Met, Phe e Trp) and aromatic (Phe, Trp, Tyr e His). Fig. 5 illustrates the different SASA and relSASA values achieved for HS and NS. The difference between the two sets of residues is notorious, and for all the 112 residues analyzed we Fig. 5. Schematic representation of the average SASA and relSASA values for each of the amino acid groups considered. have obtained values of 35.03 ± 4.31 ˚ A2(SASA) and 0.54 ± 0.05 (relSASA) for the HS and 26.57 ± 4.83 ˚ A2(SASA) and 0.30 ± 0.05 (relSASA) for NS. An exception can be found for the nonpolar residues, which contribution is higher for the NS. This statistical result is influenced by the small number of nonpolar residues that act as HS, when compared to the number of NS. However, it should be noted that the average SASA value for the nonpolar NS (31.08 ± 4.92 ˚ A2) is within the expected value of a typical HS, which indicates that the nonpolar residues have low accessibility to solvent. Charged, polar and aromatic residues show a clear difference between HS and NS, with HS having higher SASA and relSASA. We have to highlight that for charged, aromatic and polar residues the difference between HS and NS is significant, especially for relSASA in which the difference between HS and NS for charged and polar residues is more than the double. We took a similar approach as described to the RDF calculations. We have also measured the SASA characteristics of the residues inside the 4˚ A and between 4and 8˚ A spheres around the HS residues, and compared the results with the HS by themselves. The results were plotted in Fig. 6 (SASA) and Fig. 7 (relSASA). Both plots, especially Fig. 7, show a similar and expected trend. The HS have the higher value of relSASA (or SASA). As we advance for the first coordination sphere more residues are found with higher accessibility to solvent, and therefore there is a decrease in relSASA. It is more notorious when we considered the second coordination sphere that is formed almost by NS and residues with high R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 37 Table 5 Average and standard deviation values for SASA and relSASA, for each of the considered amino acid groups. Amino acid groups HS NS SASA/[Å2]All 35.03 ± 4.31 26.57 ± 4.83 ASP + GLU 13.08 ± 4.66 5.50 ± 2.06 LYS + ARG + HIS 41.37 ± 5.04 32.89 ± 7.13 Charged 38.04 ± 4.99 26.57 ± 5.96 Polar 37.54 ± 3.94 23.79 ± 3.09 Aromatic 27.80 ± 3.19 24.45 ± 4.91 Nonpolar 18.00 ± 2.07 31.08 ± 4.92 relSASA All 0.54 ±0.05 0.30 ± 0.05 ASP + GLU 0.22 ± 0.09 0.15 ± 0.05 LYS + ARG + HIS 0.57 ± 0.05 0.25 ± 0.05 Charged 0.53 ± 0.05 0.20 ± 0.05 Polar 0.84 ± 0.05 0.32 ± 0.04 Aromatic 0.35 ± 0.02 0.33 ± 0.04 Nonpolar 0.34 ± 0.05 0.44 ± 0.05 Fig. 6. Schematic representation of the average SASA considering the two coordination spheres analyzed for all HS. Fig. 7. Schematic representation of the average relSASA considering the two coordination spheres analyzed for all HS. solvent accessibility. Globally, we obtained on average 32.77 ± 4.04 ˚ A2, 13.32 ± 2.25 ˚ A2and 12.67 ± 2.00 ˚ A2for SASA on the HS itself, first coordination sphere and second coordination sphere, respectively; 0.54 ± 0.05, 0.28 ± 0.04 and 0.21 ± 0.03 for relSASA on the HS, first and second coordination spheres, respectively. Together with the RDF calculations, the SASA features seem to confirm that the O-ring theory is applicable to protein–DNA interfaces. Nevertheless, we have to stress out that our dataset is composed of ten complexes, with a known X-ray structure that fulfills the conditions mentioned in the methodological part. As long as new experimental data becomes available the study should be extended. Enough experimental data could also allow the differentiation between the different types of protein–DNA interface: enzyme, transcription factor and supporting proteins. 38 R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 4. Conclusions In the last years there has been an effort to understand the processes that govern the formation of biological complexes and the interactions that are behind them. The binding of proteins with other proteins, ligands or nucleic acids constitutes one of these processes that recently have been computational studied. For PPI it was proposed that only a small fraction of residues contribute significantly to the binding free energy – the hot-spots – which are protected from solvent molecules by null-spots. This theory later known as the O-ring theory also states that for a residue to be considered a HS it should have a low value of SASA. The O-ring theory has been the central aspect of a number of scientific papers, but the studies were mostly performed in PP complexes, and its study with other types of interfaces such as the PDI is missing. By measuring the average SASA features of the 112 residues of ten distinct protein–DNA complexes it was possible to obtain a clear perspective of the behavior of HS and NS. Radial distribution functions were also measured and helped to clearly distinguish between the two types of residues. Our results show that the HS tend to have fewer water molecules in their micro ambient and a higher value of SASA. So, they are occluded from the solvent by the NS, which ones have more water molecules in their neighborhood. In this work we were able to extend the applicability of the O-ring theory to protein–DNA complexes, since it was initially formulated in protein–protein complexes. We present evidence that the HS are indeed occluded from bulk solvent. Together with the previous works developed in protein–protein interfaces, it validates the O-ring theory. References Acton, T.B., Mead, J., Steiner, A.M., Vershon, A.K., 2000. Scanning mutagenesis of MCM1: residues required for DNA binding, DNA bending, and transcriptional activation by a Mads-Box protein. Molecular and Cellular Biology 20, 1–11. Ahmad, S., Keskin, O., Sarai, A., Nussinov, R., 2008. Protein–DNA interactions: structural, thermodynamic and clustering patterns of conserved residues in DNA-binding proteins. Nucleic Acids Research 36, 5922–5932. Allain, F.H.T., Yen, Y.M., Masse, J.E., Schultze, P., Dieckmann, T., Johnson, R.C., Feigon, J., 1999. Solution structure of the HMG protein NHP6A and its interaction with DNA reveals the structural determinants for non-sequence-specific binding. Embo Journal 18, 2563–2579. Bahadur, R.P., Chakrabarti, P., Rodier, F., Janin, J., 2003. Dissecting subunit interfaces in homodimeric proteins. Proteins-Structure Function and Genetics 53, 708–719. Bas, D.C., Rogers, D.M., Jensen, J.H., 2008. Very fast prediction and rationalization of Pk(a) values for protein–ligand complexes. Proteins-Structure Function and Bioinformatics 73, 765–783. Bochkarev, A., Pfuetzner, R.A., Edwards, A.M., Frappier, L., 1997. Structure of the single-stranded-DNA-binding domain of replication protein a bound to DNA. Nature 385, 176–181. Bochkarev, A., Bochkareva, E., Frappier, L., Edwards, A.M., 1998. The 2.2 Angstrom structure of a permanganate-sensitive DNA site bound by the Epstein-Barr virus origin binding protein, EBNA1. Journal of Molecular Biology 284, 1273–1278. Bogan, A.A., Thorn, K.S., 1998. Anatomy of hot spots in protein interfaces. Journal of Molecular Biology 280, 1–9. Brooks, B.R., Bruccoleri, R.E., Olafson, B.D., States, D.J., Swaminathan, S., Karplus, M., 1983. Charmm: a program for macromolecular energy, minimization, and dynamics calculations. Journal of Computational Chemistry 4, 187–217. Brown, B.M., Milla, M.E., Smith, T.L., Sauer, R.T., 1994. Scanning mutagenesis of the arc repressor as a functional probe of operator recognition. Nature Structural Biology 1, 164–168. Case, D.A., Darden, T.A., Cheatham, T.E., Simmerling, C.L., Wang, J., Duke, R.E., Luo, R., Merz, K.M., Pearlman, D.A., Crowley, M., Walker, R.C., Zhang, W., Wang, B., Hayik, S., Roitberg, A., Seabra, G., Wong, K.F., Paesani, F., Wu, X., Brozell, S., Tsui, V., Gohlke, H., Yang, L., Tan, C., Mongan, J., Hornak, V., Cui, G., Beroza, P., Mathews, D.H., Schafmeister, C., Ross, W.S., Kollman, P.A., 2006. Amber 9. University of California, San Francisco. Chakrabarti, P., Janin, J., 2002. Dissecting protein–protein recognition sites. ProteinsStructure Function and Genetics 47, 334–343. Chothia, C., Janin, J., 1975. Principles of protein–protein recognition. Nature 256, 705–708. Clackson, T., Ultsch, M.H., Wells, J.A., de Vos, A.M., 1998. Structural and functional analysis of the 1: 1 growth hormone: receptor complex reveals the molecular basis for receptor affinity. Journal of Molecular Biology 277, 1111–1128. Connolly, K.M., Ilangovan, U., Wojciak, J.M., Iwahara, M., Clubb, R.T., 2000. Major groove recognition by three-stranded beta-sheets: affinity determinants and conserved structural features. Journal of Molecular Biology 300, 841–856. Cruickshank, J., Shire, K., Davidson, A.R., Edwards, A.M., Frappier, L., 2000. Two domains of the Epstein-Barr virus origin DNA-binding protein, EBNA1, orchestrate sequence-specific DNA binding. Journal of Biological Chemistry 275, 22273–22277. Darden, T., York, D., Pedersen, L., 1993. Particle mesh Ewald –an N·Log(N) method for Ewald sums in large systems. Journal of Chemical Physics 98, 10089–10092. DeLano, W.L., 2002. Unraveling hot spots in binding interfaces: progress and challenges. Current Opinion in Structural Biology 12, 14–20. DeLano, W.L., Ultsch, M.H., de Vos, A.M., Wells, J.A., 2000. Convergent solutions to binding at a protein–protein interface. Science 287, 1279–1283. Guharoy, M., Chakrabarti, P., 2005. Conservation and relative importance of residues across protein–protein interfaces. Proceedings of the National Academy of Sciences of the United States of America 102, 15447–15452. Hornak, V., Abel, R., Okur, A., Strockbine, B., Roitberg, A., Simmerling, C., 2006. Comparison of multiple amber force fields and development of improved protein backbone parameters. Proteins: Structure, Function, and Bioinformatics 65, 712–725. Humphrey, W., Dalke, A., Schulten, K., 1996. VMD: visual molecular dynamics. Journal of Molecular Graphics 14, 33–38. Huo, S., Massova, I., Kollman, P.A., 2002. Computational alanine scanning of the 1: 1 human growth hormone-receptor complex. Journal of Computational Chemistry 23, 15–27. Izaguirre, J.A., Catarello, D.P., Wozniak, J.M., Skeel, R.D., 2001. Langevin stabilization of molecular dynamics. Journal of Chemical Physics 114, 2090–2098. Janin, J., 1995. Elusive affinities. Proteins-Structure Function and Genetics 21, 30–39. Jayaram, B., McConnell, K., Dixit, S.B., Das, A., Beveridge, D.L., 2002. Free-energy component analysis of 40 protein–DNA complexes: a consensus view on the thermodynamics of binding at the molecular level. Journal of Computational Chemistry 23, 1–14. Jones, S., Thornton, J.M., 1996. Principles of protein–protein interactions. Proceedings of the National Academy of Sciences of the United States of America 93, 13–20. Jorgensen, W.L., Chandrasekhar, J., Madura, J.D., Impey, R.W., Klein, M.L., 1983. Comparison of simple potential functions for simulating liquid water. Journal of Chemical Physics 79, 926–935. Klass, J., Murphy, F.V., Fouts, S., Serenil, M., Changela, A., Siple, J., Churchill, M.E.A., 2003. The role of intercalating residues in chromosomal high-mobility-group protein DNA binding, bending and specificity. Nucleic Acids Research 31, 2852–2864. Kosloff, M., Travis, A.M., Bosch, D.E., Siderovski, D.P., Arshavsky, V.Y., 2011. Integrating energy calculations with functional assays to decipher the specificity of G protein–RGS protein interactions. Nature Structural & Molecular Biology 18, 846-U128. Kumar, M.D.S., Bava, K.A., Gromiha, M.M., Prabakaran, P., Kitajima, K., Uedaira, H., Sarai, A., 2006. PROTHERM and PRONIT: thermodynamic databases for proteins and protein–nucleic acid interactions. Nucleic Acids Research 34, D204– D206. Larkin, C., Datta, S., Harley, M.J., Anderson, B.J., Ebie, A., Hargreaves, V., Schildbach, J.F., 2005. Interand intramolecular determinants, of the specificity of single-stranded DNA binding and cleavage by the F factor relaxase. Structure 13, 1533–1544. Lee, B., Richards, F.M., 1971. The interpretation of protein structures: estimation of static accessibility. Journal of Molecular Biology 55, 379–400. Lei, M., Podell, E.R., Baumann, P., Cech, T.R., 2003. DNA Self-recognition in the structure of POT1 bound to telomeric single-stranded DNA. Nature 426, 198–203. Li, J., Liu, Q., 2009. ‘Double Water Exclusion’: a hypothesis refining the O-ring theory for the hot spots at protein interfaces. Bioinformatics 25, 743–750. Li, H., Robertson, A.D., Jensen, J.H., 2005. Very fast empirical prediction and rationalization of protein Pk(a) values. Proteins-Structure Function and Bioinformatics 61, 704–721. Loncharich, R.J., Brooks, B.R., Pastor, R.W., 1992. Langevin dynamics of peptides – the frictional dependence of isomerization rates of N-acetylalanyl-N-methylamide. Biopolymers 32, 523–535. Masse, J.E., Wong, B., Yen, Y.M., Allain, F.H.T., Johnson, R.C., Feigon, J., 2002. The S. Cerevisiae architectural HMGB Protein NHP6A complexed with DNA: DNA and protein conformational changes upon binding. Journal of Molecular Biology 323, 263–284. Massova, I., Kollman, P.A., 1999. Computational alanine scanning to probe protein–protein interactions: a novel approach to evaluate binding free energies. Journal of the American Chemical Society 121, 8133–8143. Moreira, I.S., Fernandes, P.A., Ramos, M.J., 2006a. Detailed microscopic study of the full ZIPA:FTSZ interface. Proteins-Structure Function and Bioinformatics 63, 811–821. Moreira, I.S., Fernandes, P.A., Ramos, M.J., 2006b. Unraveling the importance of protein–protein interaction: application of a computational alanine-scanning mutagenesis to the study of the IGG1 streptococcal protein G (C2 fragment) complex. Journal of Physical Chemistry B 110, 10962–10969. Moreira, I.S., Fernandes, P.A., Ramos, M.J., 2007a. Computational alanine scanning mutagenesis – an improved methodological approach. Journal of Computational Chemistry 28, 644–654. Moreira, I.S., Fernandes, P.A., Ramos, M.J., 2007b. Hot spot occlusion from bulk water: a comprehensive study of the complex between the Lysozyme Hel and the Antibody FVD1.3. Journal of Physical Chemistry B 111, 2697–2706. R.M. Ramos et al. / Computational Biology and Chemistry 44 (2013) 31–39 39 Moreira, I.S., Fernandes, P.A., Ramos, M.J., 2007c. Hot spots – a review of the protein–protein interface determinant amino-acid residues. Proteins-Structure Function and Bioinformatics 68, 803–812. Murphy, F.V., Sweet, R.M., Churchill, M.E.A., 1999. The structure of a chromosomal high mobility group protein–DNA complex reveals sequence-neutral mechanisms important for non-sequence-specific DNA recognition. Embo Journal 18, 6610–6618. Oda, M., Furukawa, K., Ogata, K., Sarai, A., Ishii, S., Nishimura, Y., Nakamura, H., 1997. Identification of indispensable residues for specific DNA-binding in the imperfect tandem repeats of C-Myb R2r3. Protein Engineering 10, 1407–1414. Oda, M., Furukawa, K., Ogata, K., Sarai, A., Nakamura, H., 1998. Thermodynamics of specific and non-specific DNA binding by the C-Myb DNA-binding domain. Journal of Molecular Biology 276, 571–590. Oda, M., Furukawa, K., Sarai, A., Nakamura, H., 1999. Kinetic analysis of DNA binding by the C-Myb DNA-binding domain using surface plasmon resonance. FEBS Letters 454, 288–292. Ogata, K., Morikawa, S., Nakamura, H., Sekikawa, A., Inoue, T., Kanai, H., Sarai, A., Ishii, S., Nishimura, Y., 1994. Solution structure of a specific DNA complex of the Myb DNA-binding domain with cooperative recognition helices. Cell 79, 639– 648. Olsson, M.H.M., Sondergaard, C.R., Rostkowski, M., Jensen, J.H., 2011. PROPKA3: consistent treatment of internal and surface residues in empirical Pk(a) predictions. Journal of Chemical Theory and Computation 7, 525–537. Pérez, A., Luque, F.J., Orozco, M., 2011. Frontiers in molecular dynamics simulations of DNA. Accounts of Chemical Research 45, 196–205. Prabakaran, P., An, J., Gromiha, M.M., Selvaraj, S., Uedaira, H., Kono, H., Sarai, A., 2001. Thermodynamic database for protein–nucleic acid interactions (Pronit). Bioinformatics 17, 1027–1034. Rajamani, D., Thiel, S., Vajda, S., Camacho, C.J., 2004. Anchor residues in protein–protein interactions. Proceedings of the National Academy of Sciences of the United States of America 101, 11287–11292. Ryckaert, J.P., Ciccotti, G., Berendsen, H.J.C., 1977. Numerical-integration of cartesian equations of motion of a system with constraints – molecular-dynamics of Nalkanes. Journal of Computational Physics 23, 327–341. Schildbach, J.F., Karzai, A.W., Raumann, B.E., Sauer, R.T., 1999. Origins of DNA-binding specificity: role of protein contacts with the DNA backbone. Proceedings of the National Academy of Sciences of the United States of America 96, 811–817. Tan, S., Richmond, T.J., 1998. Crystal structure of the yeast mat alpha 2/MCM1/DNA ternary complex. Nature 391, 660–666. Thorn, K.S., Bogan, A.A., 2001. ASEDB: a database of alanine mutations and their effects on the free energy of binding in protein interactions. Bioinformatics 17, 284–285. Walther, A.P., Gomes, X.V., Lao, Y., Lee, C.G., Wold, M.S., 1999. Replication protein a interactions with DNA, 1. Functions of the DNA-binding and zinc-finger domains of the 70-Kda subunit. Biochemistry 38, 3963–3973. Wojciak, J.M., Connolly, K.M., Clubb, R.T., 1999. Nmr structure of the TN916 integrase-DNA complex. Nature Structural Biology 6, 366–373. Wyka, I.M., Dhar, K., Binz, S.K., Wold, M.S., 2003. Replication protein a interactions with DNA: differential binding of the core domains and analysis of the DNA interaction surface. Biochemistry 42, 12909–12918.