scieee AI-readable full text Open interactive document viewer

Computer-aided subsite mapping of alpha-amylases

Mótyán, János András; Gyémánt, Gyöngyi; Harangi, János; Bagossi, Péter

Full text

Elsevier Editorial System(tm) for Carbohydrate Research Manuscript Draft Manuscript Number: CAR-D-10-00518R1 Title: Computer-aided subsite mapping of α-amylases Article Type: Full Length Article Section/Category: Biochemistry and Enzymes Keywords: alpha-amylase; subsite mapping; binding energy; bond cleavage frequency; molecular modeling Corresponding Author: Dr. Péter Bagossi, Ph.D. Corresponding Author's Institution: University of Debrecen First Author: János A Mótyán Order of Authors: János A Mótyán; Gyöngyi Gyémánt, Ph.D.; János Harangi, Ph.D.; Péter Bagossi, Ph.D. Abstract: Subsite mapping is a crucial procedure in characterization of α-amylases (EC 3.2.1.1) which are extensively used in starch based industries and in diagnosis of pancreatic and salivary glands disorders. A computer-aided method has been developed for subsite mapping of α-amylases, which substitutes the difficult, expensive and time-consuming experimental determination of action patterns to crystal structures based energy calculations. Interaction energies between enzymes and carbohydrate substrates were calculated after short energy minimization by a molecular mechanics program. A training set of wild type and mutant amylases with known experimental action patterns of 13 enzymes of wide range of origin was used to set up the procedure. Calculations for training set resulted in good correlation in case of subsite binding energies (r2 = 0.827−0.929) and bond cleavage frequencies (r2 = 0.727−0.835). A set of eight novel barley amylase 1 mutants was used to test our model. Subsite binding energies were predicted with r2 = 0.502 correlation coefficient, while bond cleavage frequency prediction resulted in r2 = 0.538. Our computer-aided procedure may supplement the experimental subsite mapping methods to predict and understand characteristic features of αamylases. Reviewer #1: Recommendation: Minor Revision This manuscript reports the use of computation to determine bond cleavage frequencies and subsite binding energies from crystal structures of a number of α-amylases. This is an interesting choice of enzymes, as α-amylases are extensively used industrially and a great number of members of their family (GH13) have been sequenced and subjected to crystal structure determination. They are also endo-hydrolases, so different positions of starch chains can be cleaved. The argument is made that use of computation saves a great deal of time and money compared to obtaining the same data experimentally. This is undoubtedly true; certainly subsite mapping by experimental means is a long and hard process, as we can attest. On the other hand, computation must always take a back seat to experimentation if the latter can be used to obtain the data. After all, computation is subject to errors introduced by incomplete and inaccurate modeling of already existing experimental data. Comparison of experimental and computational data in this manuscript confirms the statement just above – the agreement between them is often rather poor. We completely agree with the reviewer therefore we emphasized throughout the manuscript that our method may complement experimental data but not substitute them. One of the reasons to choose relatively limited data sets was that we wanted to develop and apply our method to close relatives of enzymes where accuracy of modeling was usually higher than those of distant relatives. If the first try produce satisfactory results, we may extend the scope of the method later. Third, we also predicted the possible reasons where our method was failed and a special attention was needed during the work. Despite this, the manuscript is probably worth publishing as a first attempt to complement experimental subsite mapping. The English here is in general clear but is certainly the product of non-native English speakers. It will need some rewriting and correcting by a production editor. Following are specific comments as I read the manuscript: p. 3, line 11: Are plants not higher organisms? The sentence was corrected to the following: "They can be found in microorganisms, plants, animals and human where they play a dominant role in carbohydrate metabolism." p. 3, line 44: Is Domain C a carbohydrate-binding module? If so, what CBM families are found in the α-amylases studied here? *Response to Reviewers Domain C is not a carbohydrate-binding module (CBM) and enzymes of our data sets do not contain any well-defined separate CBM or starch-binding domain (SBD). α-Amylases belonging to GH-13 family are multidomain proteins that contain several characteristic domains such as the obligatory catalytic domain A and most of them possess a domain B and domain C. Some enzymes in GH 13 contain one or two additional all-β domains: domain D and/or domain E, at the C-terminal end, following the domain C (Janecek et al., 2003, Eur. J. Biochem. 270, 635–645). The function of domain D is unknown; however, domain E is referred to as SBD (the term of SBD still often used in the amylase research instead of the more "official" CBM) which is a distinct sequence-structural module that improves the efficiency of an amylolytic enzyme on raw starch. Fourand five-domain members of GH-13 can be referred to generally as the SBD-containing hydrolases and they can be classified into various CBM families. In contrast to the fact that none of the enzymes studied here contain domain D and domain E, they may contain secondary binding sites for carbohydrates in domains A, B or C. For example, the barley -amylase 1 (AMY1) provides two binding sites in addition to the catalytic cleft: a starch granule binding site within the catalytic domain A and a so-called sugar tongs within domain C (Nielsen et al., 2008, Biocatal. Biotransf., 26, 59-67). HSA has also surface binding sites located far from the active site (Ramasubbu et al., 2003, J. Mol. Biol., 325, 1061-1076). These enzymes are capable of binding and digesting raw starch without a specialized functional domain (e.g. SBD) in their sequence and structure (Janecek et al., 2003, Eur. J. Biochem. 270, 635–645). The manuscript was modified accordingly: "Besides domain A, B and C, the type and the number of extra domains such as domain D and/or domain E located at the C-terminus show wide variety within the α-amylase family. The function of domain D is unknown; however, domain E is referred to as carbohydrate-binding module (CBM) or starch-binding domain (SBD) which is a distinct sequence-structural module that improves the efficiency of an amylolytic enzyme on raw starch. It should be noted that none of the enzymes studied here contain domain D and/or domain E." p. 3, lines 53–56: Is it not true that α-amylases mainly act on α-(1,4)-linked glucans? Yes, it is true. Families 13, 70 and 77 of glycoside hydrolases contain structurally and functionally related enzymes catalyzing hydrolysis or transglycosylation of α-linked glucans, with retention of anomeric configuration. Alpha-amylases catalyze the endohydrolysis of (1→4)-α-D-glycosidic linkages in polysaccharides containing three or more (1→4)-α-linked D-glucose units and they act on starch, glycogen and related polysaccharides and oligosaccharides in a random manner. All these information can be found in the first two paragraphs of the manuscript. Nevertheless we rephrased the above mention sentence of the manuscript to the following: "Members of the α-amylase family act on several substrates (starch, glycogen, oligosaccharides) and the only shared structural feature of substrates is an α-(1-4)- linked glucose residue that should bind in the first aglycone subsite near the scissile bond." p. 5, line 35: It would be helpful to note that these PDB structures were of maltose and maltoheptaose crystallized with α-amylases. The sentence was modified to contain this information: "Structures of the maltooligosaccharide substrates were modeled based on the crystal structures of maltose (PDB code: 2GVY), maltoheptaose (PDB code: 1RP8) and a substrate-analogue inhibitor acarbose (PDB code: 1MFU, 1RPK, 1E3Z, 1OSE) complexed with various alpha-amylases." p. 5, line 42: What does “bumped” mean here? Crystal structures used for modeling of enzymes of our data sets contained shorter substrate or inhibitor than that of ours, therefore water molecules which occupied an "empty space" in the original x-ray structure may have a very close contact to the long substrate of the model (distance is much shorter than the sum of the van der Waals radius of the two atoms). Not surprisingly, it happened that oxygen atom of a hydroxyl group of our model substrate occupied approximately the same position as the oxygen atom of a water molecule in the crystal structure. In this case we removed that water molecule to avoid distortion caused by the huge initial energy during the minimization. p. 7, line 23: Except for the list of abbreviations, this is the first time that Etransf. has been mentioned. How is it calculated? The sentence was modified to explain the calculation: "After regression analysis, ESybyl data were linearly transformed to the scale of ESUMA data with the values of intercept and slope of the best fit line to get ETransf. (Fig. 2A and 2B) for each enzyme group." p. 8, line 31: A table or figure showing calculated BCF’s would be helpful. The table containing all BCF data would be too large here because BCFs of series of oligomeric substrates used to calculate a single set of binding energies. Furthermore, the BCF data will be published soon in a separate paper (Mori et al., in preparation). We may show an example for the correlation of the experimental and the calculated data, but these graph are still contained large number of points and they look a little bit messy. On the other hand, these graphs show basically the same qualitative message as the Fig. 2 but if the Reviewer suggests including any of them into the manuscript, we will do it. New figure? Correlation of experimental and calculated BCF values of wild-type and mutant AMY1 enzymes: a) training set (r2 = 0.801) and b) test set (r2 = 0.538). Majority of points located outside of the 95 % prediction interval range are belonging to mutant containing charged residue. To show this, the test set is divided into two classes: c) data of test set without those of mutant containing charged residues (r2 = 0.728) and d) data of charged mutant only. a) b) c) d) References: Journal titles should be abbreviated when they are of more than one word. References are reformatted. Table 1: What is being correlated here? Subsite binding energies were recalculated using our subsite models and this table shows the excellent correlation between the published energy values and our recalculated ones. This was explained only in the text therefore legend of Table 1 was modified to the following: "Original and modified subsite models of the studied enzymes and the correlation between the subsite binding energies of the original publications and the recalculated values based on subsite models of this work." Table 2: Individual subsite binding energies are very different between enzymes. Can a statement be made in the text linking this observation to the chain lengths of characteristic products of these enzymes? The following sentences were inserted into the Introduction: "The distribution of the high and low affinity binding subsites and the barrier subsites within the active site determines the hydrolytic efficiency on series of substrates with various length: high-affinity subsites next to the cleavage sites (−2 through +2) allow effective cleavage of shorter substrates, while distal high-affinity subsites control the action on longer substrates. Chain length and concentration of characteristic products can be determined from the BCF tables and energy contribution of each subsite to the total binding energy can be calculated." Table 3: The table legend mentions energies but not BCF’s. Legend of Table 3 was corrected to the following: "Parameters of the best fitted line of linear regression analysis of ESUMA and ESybyl values (middle panel) and the experimental and calculated values of BCFs (right panel) (m − slope, b − intercept, r2 − square of correlation coefficient)." Table 4: This table is not mentioned until the Conclusions. It should have been mentioned earlier. The tables were renumbered according to the order of appearance (earlier Table 4 now is Table 1) and the following text was inserted into the first paragraph of the Result and discussion: "Wild-type enzymes were chosen based on the existence of experimental BCF values as well as available crystal structure. In contrast to the various degree of sequence identity (11-86 %) between the wild-type enzymes (Table 1), they were all classified into the glycoside hydrolase family 13 and showed high degree of structural conservation." Figure 2A: A plot of kcal/mol vs. kJ/mol does not allow the reader to easily determine whether the slope of the regression line is unity. Fig. 2A shows only the raw data and it is not aimed to demonstrate that the slope is unity. Even if the kcal/mol is converted to kJ/mol, nobody can except that the absolute scale of the experimental and the calculated energy values would be the same. During the process of the calculation several assumptions were made, several factors were neglected, only partial minimizations were done, etc. This was the reason that the transformation step was needed and the unit conversion was also included in this step. Reviewer #2: Recommendation: Minor Revision The manuscript by Motyan et al. reports on an application of their SUMA algorithm to the problem of subsite mapping in alpha-amylases. The paper is relatively well written and addresses a moderately interesting problem in the field. The authors have collected a nice training and test set of data, and then proceed to apply their methodology to optimize empirical parameters to fit the data. This is a very common practice and methodologically, this paper does not present anything new. As such, the most interesting result in the paper is the model specifically trained for alpha-amylases. There are major weaknesses in the paper that should however be addressed before publishing. 1.) The size of the training and test sets are still small. Ideally, the authors should try to at least double the size of the training set to be able to perform more comprehensive crossvalidation. The fits are in fact not cross-validated, and the authors do not at all address whether they may in fact have overfit their model. This was our first attempt to study on in silico subsite maps of carbohydratemodifying enzymes and similar work had not been published in the literature, therefore we decided to limit our study to the relatively well characterized, structurally similar, catalytically uniform group of alpha-amylases. They still represented a wide range of enzymes of various species from bacteria to human: the sequence identity values were 11-86 % (Table 1 of the revised manuscript). We must limit our training set to the enzymes which experimentally determined BCF and crystal structure were both existed. We included all published experimental data in our training set except two: BCF and subsite map of rice amylase were published (Gyémánt et al., Eur. J. Biochem., 2002, 269, 5157-5162) but its crystal structure has not been determined yet and crystal structure of maltogenic α-amylase of Bacillus stearothermophilus was determined but BCF data has not been published yet. We admit that the size of our data set did not allow to handle all aspects in statistically satisfactory way but we think that our results and conclusions are correct and not go beyond the scope of this study. 2.) Why are the authors focusing only on alpha-amylases? Why not perform this analysis on all amylases, which would greatly increase the size of their training set and allow for much more 10-fold cross validation studies. The most known amylolytic enzymes are α-amylase (EC 3.2.1.1), β-amylase (EC 3.2.1.2) and glucoamylase (EC 3.2.1.3); however, they differ from each other in their primary and tertiary structures, their catalytic machineries and reaction mechanisms. They have therefore classified into different families: GH13 - α-amylases, GH14 - β-amylases, GH15 - glucoamylases. It may possible that our method can be applied not only to GH13 but further families too; however, it needs further studies and it is beyond the scope of this manuscript. SUMA program was used in several step of our method and this program was developed and validated only for α-amylases, as its full name in the Abbreviations and in the text indicates: SUbsite Mapping of α-Amylases. Please also see the answer for the first question of this reviewer. 3.) The authors are using the AMBER force field, yet there are far more accurate force-fields for sugars, e.g. glycam. The authors should evaluate how using different force-fields affect their results. Furthermore, the authors should also evaluate whether using explicit solvent in their simulations improves (or perhaps worsens) their fits to the experimental data. Several force fields available in Sybyl program: Kollman_All_Atom, Amber95 and Amber7_FF99 were tested, together with an in-house implementation of Glycam06 force field. Amber7_FF99 was chosen because its performance was the best in our hand. Furthermore, wide-variety of values of dielectric constant (1, 2, 3, 4 and 8), non-bonded cutoff (8, 10 and 12), number of iterations (20-100 by steps of 20 and 100-1000 by steps of 100) and the width of minimized shell (1-12 angstrom) of atoms around the substrate and the mutated residues in the initial energy minimization were also probed to find a good combination of parameters and it was turned out that these factors also significantly modified the results. Several, but not all possible combinations were tested because it would require large amount of computational power, therefore extensive discussion on the parameter choice was neglected. One of our aims was to build a fast procedure; therefore explicit molecules for bulk solvent and extensive minimization were avoided. We think that the readers of Carbohydrate Research are more interesting in the chemical/biochemical aspects of our method rather than the computational details of the parameter choice. The text of the manuscript was modified accordingly. 4.) The experimental data used to fit have been collected under a wide-variety of solution conditions. The authors need to discuss and assess how experimental error may in fact be limiting the resolution of their methodology. The wild-type enzymes were measured at various reaction conditions; however, they were measured at the optimal condition of each enzyme and this situation may be more appropriate for determination of bond cleavage frequencies and subsite binding energies than the choice of one reaction condition which is optimal for only one enzyme and suboptimal for the all others. Definitive conclusion of the effect of each reaction condition (e.g. pH, temperature or buffer composition) can be drawn only after extensive experimentation. Mutant enzymes were measured at the same reaction condition than their wild-type enzymes based on the assumption that their features were changed moderately. The same assumption was also necessary for the theoretical calculations. Furthermore, the effect of different reaction conditions could be eliminated in a large extent with our energy transformation procedure which was group-specific. Enzyme groups were based on each wild-type enzymes and mutants belonged to the corresponding wild-type group. Differences of the calculated subsite binding energies were below 0.4 kJ/mol in case of two independent experimental measurements of BCF, as we wrote in the Materials and methods. In a recent publication of us it was found that the same type of error was below 0.6 kJ/mol (Nielsen et al., 2009, Biochemistry 48, 7686–7697). 5a.) It is not clear how their results will be shared with the world. Can we access their parameters on a web server to test with other amylases? 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 5 Computer-aided analysis of X-ray crystallographic structures is an effective way to examine enzyme-substrate interactions and it could be supplemented by molecular mechanical/dynamical calculations to predict the interaction in a quantitative manner. In this work, we aimed to develop a computer-aided procedure for subsite mapping of α-amylases to predict and interpret functional features of wild type and mutant α-amylases. 2. Materials and methods 2.1. Calculation of subsite binding energies and BCFs by SUMA Experimentally determined BCFs were collected from the literature for each enzyme and based on them, number of subsites, position of the cleavage site and binding energy (ESUMA) of each subsite were calculated by SUMA. Apparent free energy values were optimized in a SUMA calculation by minimization of differences between experimental and calculated BCFs11. Differences of the calculated subsite binding energies were below 0.4 kJ/mol in case of two independent experimental measurements of action pattern of the same enzyme (data not shown). To standardize data evaluation, ESUMA and ESybyl data (see later) were calculated for the same number of subsites of wild-type and mutant enzymes. The temperature and the interpolation step were set to 37 °C and 0.01, respectively, in all calculation in contrast to the default values of the program. BCF prediction using calculated subsite binding energies was performed by SUMA according to the standard parameters (substrates of DP 3 to 11, temperature 37 °C). 2.2. Molecular modeling Crystal structures of wild-type enzymes were downloaded from the Protein Data Bank13. Structures of human salivary α-amylase (PDB code: 1SMD14), barley α-amylase AMY1 (PDB code: 1RP815) and AMY2 (PDB code: 1BG916), α-amylase of Bacillus amyloliquefaciens (PDB code: 3BH417), porcine pancreatic α-amylase isoenzyme II (PPA II) (PDB code: 1PIG18) and αamylase of Aspergillus oryzae (PDB code: 7TAA19) were used as templates for building of 3D structures of wild type and mutant enzymes using Sybyl program package (Tripos Inc., St. Louis, MO, USA). Structures of the maltooligosaccharide substrates were modeled based on the crystal structures of maltose (PDB code: 2GVY20), maltoheptaose (PDB code: 1RP815) and a substrate-analogue inhibitor acarbose (PDB code: 1MFU21, 1RPK15, 1E3Z22, 1OSE23) complexed with various alpha-amylases. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 6 Initial structure of an enzyme-substrate complex was generated by merging the substrate into the binding cleft of the enzyme and deleting the inhibitor if it existed in the original crystal structure. Water molecules bumped with substrate were also removed. In mutant proteins, wildtype residues were changed by Sybyl. To avoid bumping of mutant residues with unmodified ones, water molecules or substrate chain, the spatial position of the side chain of the modified residue was predicted by a short molecular dynamics run (length: 1000 step, step size: 1 fs, temperature: 300 K, dielectric constant: 1, AMBER7_FF99 force field). To refine the position of the substrate and the mutated residues, a short minimization procedure was performed by Sybyl (100 Powell iterations, dielectric constant 1, AMBER7_FF99 force field) on the 8 Å surround of the oligosaccharide chain and the mutated residues. The resulted complex was further energyminimized without any fixed atoms by Sybyl using the following parameters: AMBER7_FF99 force field, non-bonded cutoff 8 Å, 100 Powell iterations and dielectric constant 1. Interaction energy (ESybyl) between the enzyme and the carbohydrate substrate was calculated for each subsite. Calculations and visualization were performed on Silicon Graphics Fuel workstations (Silicon Graphics International, Fremont, CA, USA). 3. Results and discussion 3.1. Data sets The training set consisted of the subsite binding energies of all wild-type enzymes and some mutants of human salivary amylase (HSA) and barley amylase 1 (AMY1) enzymes. Wildtype enzymes were chosen based on the existence of experimental BCF values as well as available crystal structure. In contrast to the various degree of sequence identity (11-86 %) between the wild-type enzymes (Table 1), they were all classified into the glycoside hydrolase family 13 and showed high degree of structural conservation. Subsite binding energy values of HSA enzymes were recalculated by SUMA from the previously published action patterns of wild type HSA and its W58L24 and Y151M25 mutants. Recalculation of published subsite maps was necessary, because the published energies were calculated for 8 subsites (5 glycone and 3 aglycone subsites); however, reliable structural model for substrate can be built only for 7 subsites (4 glycone and 3 aglycone subsites). The published binding energy of subsite −5 was small (0.65 kJ/mol)24, therefore simplification of the subsite model should cause only marginal changes in the binding energies of other subsites. Indeed, the original and the recalculated binding energies were in excellent agreement, square of the correlation coefficients (r2) from the linear regression analysis were 0.986. Similarly, published subsite maps of all other wild-type 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 7 enzymes were recalculated and there were excellent correlations in all cases (Table 2). The training set consisted of total 13 wild-type and mutant enzymes and 97 subsite binding energy values (Table 3). A test set was created from subsite binding energies of novel mutants of AMY1: S48Y, V47A, V47D, V47F, V47I/S48I, V47K/S48G, V47G/S48D and V47L/S48A. These single and double mutant enzymes contain various mutations of V47 and/or S48 residues which were engineered previously with random mutagenesis for the investigation of the enzyme-substrate interactions26. Action patterns of these enzymes were determined experimentally using hydrolysis of CNP-MOSs of DP 3−11 and subsite maps were calculated by SUMA (Mori et al., in preparation). The subsite maps were calculated for 7 glycone and 4 aglycone binding sites by the same way as the AMY1 group of training set. The test set consisted of total 8 mutant enzymes and 73 subsite binding energy values (Table 3). 3.2. Steps of calculation Our working hypothesis was that the subsite binding energies could be predicted from enzyme-substrate interaction energies calculated by a molecular mechanical program. Homologous models of enzyme-substrate complexes of the training set were built and interaction energies were calculated between the protein and the carbohydrate residues of each subsite (ESybyl). These values were compared to those calculated based on the experimentally determined action patterns (ESUMA). To reach highest correlation between ESUMA and ESybyl data, parameters of the energy-minimization procedure were optimized. Several force fields available in Sybyl program: Kollman_All_Atom, Amber95 and Amber7_FF99 were tested, together with an in-house implementation of Glycam06 force field27. Amber7_FF99 was chosen because its performance was the best in our hand. Furthermore, wide-variety of values of dielectric constant (1, 2, 3, 4 and 8), non-bonded cutoff (8, 10 and 12), number of iterations (20-100 by steps of 20 and 100-1000 by steps of 100) and the width of minimized shell (1-12 angstrom) of atoms around the substrate and the mutated residues in the initial energy minimization were also probed to find a good combination of parameters. Our aim was to build a relatively fast procedure, therefore explicit molecules for bulk solvent and extensive minimization were avoided. Linear regression analysis was performed with the exception of energies of the −1 and +1 subsites because it is theoretically impossible to calculate them from experimental bond cleavage frequency data11. After regression analysis, ESybyl data were linearly transformed to the scale of ESUMA data with the values of intercept and slope of the best fit line to get ETransf. (Fig. 2A and 2B) for each enzyme group. Enzyme groups were based on the wild-type enzymes and 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 8 each mutant belonged to the group of the corresponding wild-type enzyme. Calculated ESybyl and ETransf. values of training enzymes (Table 3) showed good correlation with ESUMA values: r2 were in the range of 0.827−0.929 (Table 4). The procedure was applied for the test set to assess the predictive potential of the model. Calculated interaction energies (ESybyl) of the test set (containing only AMY1 mutants) were transformed to ETransf. values using equation of AMY1 group of the training set. Calculated ESybyl and ETransf. values of the test set (Table 3) showed correlation with ESUMA values (Fig. 2C): r2 was 0.502 (Table 4) which was lower than the corresponding value of 0.827 of AMY1 group of the training set, as expected. However, the error distribution was odd: all values out of 95% prediction interval were belonged to V47D, V47K/S48G and V47G/S48D mutants. All these mutants and only these mutants contained a substitution from neutral to charged residue. If all values of these mutants were omitted from the plot, the r2 value rose up to 0.638. Mutants of the test set contained single or double mutations in the position of V47 and/or S48 residues of AMY1. Crystallographic analysis showed that Val47 was involved in direct hydrogen bonds at the −7 subsite, while Ser48 formed indirect hydrogen bonds at the −1 and −2 subsites15. Docking studies by molecular modeling predicted that the influence of V47 was extended to the −6, −5, −4 subsites and S48 effected substrate binding at the −4, −3, −2 subsites28. Furthermore, V47 together with Y105, may be considered as the substrate “entrance” to the active site cleft where the substrate adopts a half-circle conformation centered on Val4715. Therefore adjacent V47 and S48 amino acids are surrounded by several glucose residues of bounded substrate chain and they may effect the dynamics of enzyme-substrate interaction in addition to influence of binding affinity. It seems that our simple procedure was only moderately able to model this combined effect as the correlation coefficient was dropped from 0.827 to 0.638 (neglecting mutants having substitutions to charged residues). Large changes in charge distribution of the mutated position turned out to be also problematic, as in cases of V47D, V47K/S48G and V47G/S48D mutants: the correlation coefficient was further dropped to 0.502. In contrast with the test set, enzymes of AMY1 group of training set contain the mutations of Y105 and/or T212, which amino acids located at −6 and +4 subsites and participated in substrate binding at only one or two outermost substrate binding site28 and these mutants did not harbor any substitution to charged residues. Some mutants containing amino acid changes in the secondary binding sites were also tested, but our procedure failed to correctly predict changes caused by these distant sites, as expected (data not shown). 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 9 3.3. Prediction of action pattern Computer program SUMA is capable not only to calculate ESUMA energies but also to recalculate BCFs from subsite maps11 thus it can be used for prediction of the primary experimental values. Bond cleavage frequencies were calculated using the ETransf. values as input data in SUMA and the predicted BCFs were correlated with the experimentally determined values. Linear regression analysis resulted in good correlation (r2 = 0.727−0.835) between predicted and published BCFs of enzymes of the training set (Table 4). In case of mutant AMY1 enzymes of the test set, the correlation coefficient was similarly low to that of binding energy analysis: r2 = 0.538 (Table 4). Not unexpectedly, the odd error distribution was also present here: if values of mutant enzymes having substitution to charged residues (V47D, V47K/S47G and V47G/S48D) were omitted, the value of r2 rose up to 0.728. 4. Conclusion The main structural feature of family 13 glycoside hydrolases is the presence of the characteristic (α)8-barrel domain with a varying number of extra domains. The substrate binding site is made from the residues of domains A and B, however, the architecture of the domain B varies between amylases and it is the major determinant in differences of substrate specificities. Present study was made with the aim to develop a computer-aided subsite mapping procedure to better understand the specificity of α-amylases. Our procedure was successfully adopted for α-amylases belonging to different kingdoms: BAA represents liquefying bacterial αamylases from Bacteria, further α-amylases were derived from Eukaryota: TAA represents amylases from fungi, AMY1 and AMY2 represent plant enzymes and PPA and HSA represent the mammalian enzymes. The structural conservation of these enzymes were much stronger than the sequence conservation (Table 4), which give the hope that our procedure can be successfully applied for wide range of α-amylases. The great clinical and industrial importances justify the studies of α-amylases. Described procedure could help to generate subsite maps of α-amylases as well as it could predict the primary experimental data (BCFs) via the use of computer program SUMA11. However, special care is needed when neutral residues was changed to charged one (and probably vice versa) in the substrate binding site of an α-amylase. The computer-aided subsite mapping may support the examination of the role of substrate binding residues and it can be used to supply protein engineers in design of modified enzymes with altered action pattern. Our method can 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 10 supplement the experimental subsite mapping procedure by prediction of subsite maps and action patterns in a fast and cheap way. Acknowledgements Erika Fazekas is acknowledged for help in determination of BCFs. References 1. Gupta, R.; Gigras, P.; Mohapatra, H.; Goswami, V.K.; Chauhan, B. Process Biochem. 2003, 38, 1599-1616. 2. Rohleder, N.; Nater, U.M. Psychoneuroendocrinology 2009, 34, 469-485. 3. Cantarel, B.L.; Coutinho, P.M.; Rancurel, C.; Bernard, T.; Lombard, V.; Henrissat, B. Nucleic Acids Res. 2009, 37, D233-D238. 4. MacGregor, E.A.; Janecek, S.; Svensson, B. Biochim. Biophys. Acta 2001, 1546, 1-20. 5. Janecek, S.; Svensson, B.; Henrissat, B. J. Mol. Evol. 1997, 45, 322-331. 6. Jespersen, H.M.; Macgregor, E.A.; Sierks, M.R.; Svensson, B. Biochem. J. 1991, 280, 51-55. 7. Janecek, S.; Svensson, B.; MacGregor, E.A. Eur. J. Biochem. 2003, 270, 635-645. 8. Kumar, V. Carbohydr. Res. 2010, 345, 893-898. 9. Davies, G.J.; Wilson, K.S.; Henrissat, B. Biochem. J. 1997, 321, 557-559. 10. Kandra, L.; Gyémánt, G.; Lipták, A. Biologia 2002, 57, 171-180. 11. Gyémánt, G.; Hovánszki, G.; Kandra, L. Eur. J. Biochem. 2002, 269, 5157-5162. 12. Kandra, L.; Abou Hachem, M.; Gyémánt, G.; Kramhoft, B.; Svensson, B. FEBS Lett. 2006, 580, 5049-5053. 13. Berman, H.M.; Westbrook, J.; Feng, Z.; Gilliland, G.; Bhat, T.N.; Weissig, H.; Shindyalov, I.N.; Bourne, P.E. Nucleic Acids Res. 2000, 28, 235-242. 14. Ramasubbu, N.; Paloth, V.; Luo, Y.G.; Brayer, G.D.; Levine, M.J. Acta Crystallogr. D Biol. Crystallogr. 1996, 52, 435-446. 15. Robert, X.; Haser, R.; Mori, H.; Svensson, B.; Aghajari, N. J. Biol. Chem. 2005, 280, 32968-32978. 16. Kadziola, A.; Sogaard, M.; Svensson, B.; Haser, R. J. Mol. Biol. 1998, 278, 205-217. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 11 17. Alikhajeh, J.; Khajeh, K.; Ranjbar, B.; Naderi-Manesh, H.; Lin, Y.H.; Liu, E.H.; Guan, H.H.; Hsieh, Y.C.; Chuankhayan, P.; Huang, Y.C.; Jeyaraman, J.; Liu, M.Y.; Chen, C.J. Acta Crystallogr. Sect. F Struct. Biol. Cryst. Commun. 2010, 66, 121-129. 18. Machius, M.; Vertesy, L.; Huber, R.; Wiegand, G. J. Mol. Biol. 1996, 260, 409-421. 19. Brzozowski, A.M.; Davies, G.J. Biochemistry 1997, 36, 10837-10845. 20. Vujicic-Zagar, A.; Dijkstra, B.W. Acta Crystallogr. Sect. F Struct. Biol. Cryst. Commun. 2006, 62, 716-721. 21. Ramasubbu, N.; Ragunath, C.; Mishra, P.J. J. Mol. Biol. 2003, 325, 1061-1076. 22. Brzozowski, A.M.; Lawson, D.M.; Turkenburg, J.P.; Bisgaard-Frantzen, H.; Svendsen, A.; Borchert, T.V.; Dauter, Z.; Wilson, K.S.; Davies, G.J. Biochemistry 2000, 39, 90999107. 23. Gilles, C.; Astier, J.P.; MarchisMouren, G.; Cambillau, C.; Payan, F. Eur. J. Biochem. 1996, 238, 561-569. 24. Ramasubbu, N.; Ragunath, C.; Mishra, P.J.; Thomas, L.M.; Gyémánt, G.; Kandra, L. Eur. J. Biochem. 2004, 271, 2517-2529. 25. Kandra, L.; Gyémánt, G.; Remenyik, J.; Ragunath, C.; Ramasubbu, N. FEBS Lett. 2003, 544, 194-198. 26. Svensson, B.; Jensen, M.T.; Mori, H.; Bak-Jensen, K.S.; Bonsager, B.; Nielsen, P.K.; Kramhoft, B.; Praetorius-Ibba, M.; Nohr, J.; Juge, N.; Greffe, L.; Williamson, G.; Driguez, H. Biologia 2002, 57, 5-19. 27. Kirschner, K.N.; Yongye, A.B.; Tschampel, S.M.; Daniels, C.R.; Foley, B.L.; Woods, R.J.J. J. Comput. Chem. 2008, 29, 622-655. 28. Bak-Jensen, K.S.; Andre, G.; Gottschalk, T.E.; Paes, G.; Tran, V.; Svensson, B. J. Biol. Chem. 2004, 279, 10093-10102. 29. Allen, J.D.; Thoma, J.A. Biochem. J. 1976, 159, 121-131. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 12 FIGURE LEGENDS Figure 1. Schematic representation of structure of AMY1 (bottom panel) together with the architecture of substrate binding site of studied wild-type enzymes (top panel). Green hexagons represent monosaccharide units of the substrate, large arrow shows the site of cleavage, subsites are labeled by negative and positive numbers for glycone and aglycone binding sites, respectively. Figure 2. Correlation between the ESUMA values of AMY1 group and the predicted energies (ESybyl and ETransf.). Black lines correspond to the best fitted regression line and green lines show the 95 % prediction interval range. A) AMY1 group of training set before linear transformation (ESybyl = 0.905 x ESUMA – 17.731) B) AMY1 group of training set after linear transformation (ETransf = 1.000 x ESUMA – 0.000) C) AMY1 group of test set after linear transformation. T212 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 13 Table 1. Sequence identities (%) of the studied enzymes based on structural alignment (UniProtKB codes of protein sequences used for sequence alignment: HSA – P04745, PPA – P00690, AMY1 – P00693, AMY2 – P04063, TAA – P0C1B3, BAA – P00692). HSA PPA AMY1 AMY2 TAA BAA HSA 100 -- -- -- -- -- PPA 86 100 -- -- -- -- AMY1 12 11 100 -- -- -- AMY2 12 11 74 100 -- -- TAA 18 19 14 13 100 -- BAA 16 18 15 15 20 100 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 14 Table 2. Original and modified subsite models of the studied enzymes and the correlation between the subsite binding energies of the original publications and the recalculated values based on subsite models of this work. Enzyme Subsite model square of correlation coefficient (r2) reference original work this work HSA −5 ... +3 −4 ... +3 0.986 24, 25 PPA −4 ... +3 −4 ... +3 0.942 11 AMY1 −8 ... +4 −7 ... +4 0.994 12 AMY2 −8 ... +4 −7 ... +4 0.999 12 TAA −3 ... +5 −3 ... +5 -- * 29 BAA −6 ... +4 −7 ... +3 0.996 29 * Energy values of subsites were not published in the original work. Figure2C Click here to download high resolution image