scieee AI-readable full text Open interactive document viewer

Machine learning techniques to discover genes with potential prognosis role in Alzheimer’s disease using different biological sources

Martínez Ballesteros, María del Mar; García Heredia, José Manuel; Nepomuceno Chamorro, Isabel de los Ángeles; Riquelme Santos, José Cristóbal

Abstract

Alzheimer’s disease is a complex progressive neurodegenerative brain disorder, being its prevalence ex pected to rise over the next decades. Unconventional strategies for elucidating the genetic mechanisms are necessary due to its polygenic nature. In this work, the input information sources are five: a public DNA microarray that measures expression levels of control and patient samples, repositories of known genes associated to Alzheimer’s disease, additional data, Gene Ontology and finally, a literature review or expert knowledge to validate the results. As methodology to identify genes highly related to this disease, we present the integration of three machine learning techniques: particularly, we have used decision trees, quantitative association rules and hierarchical cluster to analyze Alzheimer’s disease gene expres sion profiles to identify genes highly linked to this neurodegenerative disease, through changes in their expression levels between control and patient samples. We propose an ensemble of decision trees and quantitative association rules to find the most suitable configurations of the multi-objective evolutionary algorithm GarNet, in order to overcome the complex parametrization intrinsic to this type of algorithms. To fulfill this goal, GarNet has been executed using multiple configuration settings and the well-known C4.5 has been used to find the minimum accuracy to be satisfied. Then, GarNet is rerun to identify de pendencies between genes and their expression levels, so we are able to distinguish between healthy individuals and Alzheimer’s patients using the configurations that overcome the minimum threshold of accuracy defined by C4.5 algorithm. Finally, a hierarchical cluster analysis has been used to validate the obtained gene-Alzheimer’s Disease associations provided by GarNet. The results have shown that the ob tained rules were able to successfully characterize the underlying information, grouping relevant genes for Alzheimer Disease. The genes reported by our approach provided two well defined groups that per fectly divided the samples between healthy and Alzheimer’s Disease patients. To prove the relevance of the obtained results, a statistical test and gene expression fold-change were used. Furthermore, this rel evance has been summarized in a volcano plot, showing two clearly separated and significant groups of genes that are up or down-regulated in Alzheimer’s Disease patients. A biological knowledge integration phase was performed based on the information fusion of systematic literature review, enrichment Gene Ontology terms for the described genes found in the hippocampus of patients. Finally, a validation phase with additional data and a permutation test is carried out, being the results consistent with previous studies.

Full text

Machine learning techniques to discover genes with potential prognosis role in Alzheimer’s disease using different biological sources María Martínez-Ballesteros a, ∗, José M. García-Heredia b, Isabel A. Nepomuceno-Chamorro a, José C. Riquelme-Santos a a Dpto. Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Seville, Spain b Dpto. Bioquímica Vegetal y Biología Molecular, Universidad de Sevilla, Seville, Spain Keywords: Association rules Gene expression profiles Alzheimer’s disease Statistical significant genes Ensemble learning Biological knowledge integration a b s t r a c t Alzheimer’s disease is a complex progressive neurodegenerative brain disorder, being its prevalence expected to rise over the next decades. Unconventional strategies for elucidating the genetic mechanisms are necessary due to its polygenic nature. In this work, the input information sources are five: a public DNA microarray that measures expression levels of control and patient samples, repositories of known genes associated to Alzheimer’s disease, additional data, Gene Ontology and finally, a literature review or expert knowledge to validate the results. As methodology to identify genes highly related to this disease, we present the integration of three machine learning techniques: particularly, we have used decision trees, quantitative association rules and hierarchical cluster to analyze Alzheimer’s disease gene expression profiles to identify genes highly linked to this neurodegenerative disease, through changes in their expression levels between control and patient samples. We propose an ensemble of decision trees and quantitative association rules to find the most suitable configurations of the multi-objective evolutionary algorithm GarNet, in order to overcome the complex parametrization intrinsic to this type of algorithms. To fulfill this goal, GarNet has been executed using multiple configuration settings and the well-known C4.5 has been used to find the minimum accuracy to be satisfied. Then, GarNet is rerun to identify dependencies between genes and their expression levels, so we are able to distinguish between healthy individuals and Alzheimer’s patients using the configurations that overcome the minimum threshold of accuracy defined by C4.5 algorithm. Finally, a hierarchical cluster analysis has been used to validate the obtained gene-Alzheimer’s Disease associations provided by GarNet. The results have shown that the obtained rules were able to successfully characterize the underlying information, grouping relevant genes for Alzheimer Disease. The genes reported by our approach provided two well defined groups that perfectly divided the samples between healthy and Alzheimer’s Disease patients. To prove the relevance of the obtained results, a statistical test and gene expression fold-change were used. Furthermore, this relevance has been summarized in a volcano plot, showing two clearly separated and significant groups of genes that are up or down-regulated in Alzheimer’s Disease patients. A biological knowledge integration phase was performed based on the information fusion of systematic literature review, enrichment Gene Ontology terms for the described genes found in the hippocampus of patients. Finally, a validation phase with additional data and a permutation test is carried out, being the results consistent with previous studies. 1. Introduction Neurodegenerative diseases are complex syndromes characterized by a common feature: up to now, all of them progress ∗Corresponding author. E-mail addresses: [email protected] (M. Martínez-Ballesteros), jmgheredia@ inexorably. Although medical treatments can slow symptoms’ progression, there is no cure for any of them. One of the most common neurodegenerative diseases, constituting approximately 70% of all cases [1] , is Alzheimer’s Disease (AD). This multifactorial and heterogeneous disorder, characterized by a progressive loss of memory and a decline in cognitive function, affects to around 10% of people between 65–85 years, increasing its risk significantly with age, reaching percentages of up to 50–60% of people over 85 [2,3] . Although AD exhibits two markedly brain histological us.es (J.M. García-Heredia), [email protected] (I.A. Nepomuceno-Chamorro), [email protected] (J.C. Riquelme-Santos). characteristics, plaques and tangles accumulation [4,5] , the presence of one or both symptoms are not a proof of AD development. Indeed, around a 30% of normal aged people have similar levels of plaques than AD patients of similar ages [6,7] . The main cause of AD may be genetic. In fact, up to a 5% of AD cases are due to genetic inheritance, being responsible of the appearing of an early onset AD [8] . In these cases, mutations in three different genes ( APP, PSEN1 and PSEN2 ) have been connected to pathophysiology of the disease [8–11] . Assuming its genetic background, due to changes in gene regulation or mutations accumulated through lifetime, it is important to know which genes are usually altered in AD patients, mainly in its initial states, in order to design treatments that effectively slow down -or even stopAD progression. In this context, Data Mining and statistical techniques have been applied to find useful data patterns in Bioinformatics and Biomedicine fields to discover, for example, affected pathways in a specific syndrome or disease [12] . Specifically, several works have been published where such techniques have been focused on AD using gene expression profiles. These profiles are characterized by a low number of samples (patients) but a very high number, up to thousands, of features (gene expression) [13] . Example of data mining technique focused on AD are two types of non-supervised methods (principal component analysis and independent component analysis) that have been applied to extract and characterize the most relevant features from DNA microarray gene expression data of AD [14] . However, most of techniques used in the literature to increase AD knowledge present a low-dimensional solution. Thus, they are not sufficiently descriptive or the information provided is limited. Association Rules (AR) and particularly Quantitative Association Rules (QAR) [15] have emerged as a popular methodology to discover significant and apparently hidden relationships among attributes in a subspace of the dataset instance. The application of QAR in AD-related DNA microarray data analysis might provide a deeper knowledge into biological functions with higher relevance, since they can be used to identify gene expression patterns, helping to find which of them are associated with AD. In the context of this work, QAR are used fixing the consequent of the rule to the AD state (healthy or not) with the aim of addressing the problem as a rule-based classification. The purpose of this work is to present an ensemble of decision trees, rules and hierarchical cluster to analyze microarray gene expression data related to AD. Usually, ensembles of classifiers are used to obtain better predictive performance. Here, we proposed an ensemble of classifiers to obtain the best configuration settings of the evolutionary algorithm GarNet [16] to identify genes highly related to AD, through changes in their expression levels between healthy and AD samples. This algorithm requires the parametrization of the objectives to be optimized and the minimum threshold of some quality measures as many other algorithms do. With the aim of obtaining the most suitable parameters of GarNet, the well-known C4.5 algorithm has been applied to the gene expression data provided in [17] with the objective of setting the minimum accuracy threshold to be satisfied by the rules obtained by GarNet. These configurations have been used to rerun GarNet to find significant genes with potential prognosis role in AD. Furthermore, a hierarchical cluster analysis has been also applied to group healthy and AD patients using the genes obtained by GarNet. It can be noted that the use of prior knowledge can outperform our approach and increase the possibilities of correcting the spurious information existing in high-throughput technology data as gene expression data. Then, a phase of biological knowledge was performed including other biological sources such as systematic review of the literature in PubMed, Gene Ontology processes and protein-protein interaction networks. Thus, the systematic search in PubMed allowed us to detect altered functions specifically in neurons of patients affected by AD. These functions include genes that encode proteins connected to metabolism, cell signaling or protein-nucleic acid interaction, suggesting an effect on overall cell functions. In addition, enrichment analysis in the context of Gene Ontology and a mapping process to network protein-protein interactions allowed us to strengthen the relationship of the discovered genes with AD. Using this ensemble learning and fusion strategy, we have found more than 90 genes whose expression is modified during AD progression, affecting processes as diverse as lipid metabolism, transcriptional regulation or protein synthesis. In addition, some of the genes are connected to cardiovascular diseases or diabetes, disorders previously related to AD. The identified genes show the complexity of AD, and could be used to design preventive treatments in healthy people before the appearance of the first AD symptoms. The remainder of the paper is structured as follows. Section 2 presents a brief introduction about AD, summarizes the main concepts of AR and quality measures, overviews the most relevant concepts in classification and briefly presents the C4.5 algorithm. Section 3 thoroughly describes the six phases integrated in our approach to identify genes highly related to AD including the main features of GarNet algorithm. Section 4 presents and discusses the results obtained in each phase of the proposed analysis. Finally, Section 5 summarizes the conclusions drawn from the analysis conducted. 2. Preliminaries This section provides a brief description of AD, followed by main concepts of AR, QAR and quality measures, in addition to classification, decision trees and the well-known C4.5 algorithm. 2.1. Alzheimer disease As it has been stated in the introduction, plaques and tangles are common features in AD. The first are due to progressive extracellular deposition of amyloid β(A β) peptides, due to an inadequate clearance that produces the formation of plaques. In fact, plaques accumulation can be detected in non-affected individuals even more than twenty years before the appearance of AD symptoms [18–20] . Neurofibrillary tangles are formed by the accumulation of hyperphosphorylated tau protein inside neurons, being its levels up 2-fold higher that in normal brain, affecting its normal function [21] . This post-translational alteration affects tau protein normal function. In normal conditions, tau protein interacts with tubulin, promoting the microtubule assembly. However, under hyperphosphorylated state, tau protein cannot interact with tubulin, being capable to form tau helical filaments with no activity. Although some authors have considered that plaques and/or tangles may be considered the initiating event of AD [8,22] , it is not clear if they are the cause or just symptoms of this neurodegenerative disorder, remaining its main cause elusive. It can be considered as the junction of multiple imbalances, from mitochondrial dysfunction to oxidative stress, neuroinflammation and changes in gene regulation. The last one could be considered as the hard core of the disease, due to its role in other effects. In this way, formation of A βplaques and tau tangles are caused by the dysfunction of proteins associated to their normal processing. In addition, mitochondrial malfunction and oxidative damage can also be connected to gene regulation, due to a decrease in the synthesis of proteins from the oxidative stress pathway. This fact underlines the importance of developing tools to find common patterns in AD. 2.2. Association rule mining and quality measures The AR learning is a popular and well-known method in the Data Mining field used to discover interesting relationships among variables in large databases [23] . AR aim at identifying patterns that explain or summarize the data, instead of predicting the class of new data [24] . When the domain is continuous, the AR is known as QAR. In this context, let F = { F 1 , ..., F n } be a set of features or attributes, with values in R . Let A and C be two disjoint subsets of F , that is, A ⊂F, C ⊂F , and A ∩ C = ∅ . A QAR is a rule X ⇒ Y , in which features in A belong to the antecedent X , and features in B belong to the consequent Y , such that X and Y are formed by a conjunction of multiple boolean expressions of the form F i ∈ [ v 1 , v 2 ], (with v 1 , v 2 ∈ R ). Thus, in a QAR, the features or attributes of the antecedent are related with the features of the consequent, establishing a membership value interval for each attribute involved in the rule. For example, a QAR could be numerically expressed as EPHA 10 ∈ [2, 2.9] ∧ TOR 2 A ∈ [1.8, 2.6] ⇒ STRN 4 ∈ [3, 3.5] where EPHA 10 and TOR 2 A constitute the features appearing in the antecedent and STRN 4 the one in the consequent. There are several probability-based measures proposed in the literature to evaluate the generality and reliability of AR (and QAR) obtained in the mining process [25,26] . In this work, we have used the support, confidence, leverage, accuracy and gain measures to optimize and evaluate the quality of the QAR obtained by GarNet. The description and the mathematical definition can be found in [27] . Methods based on QAR have not been used to find gene associations in AD, although the technique has been used to analyze gene expression data [28] and other AD features [29] . In the first work, QAR has been used to analyze 23 genes known to be involved in arginine metabolism from yeast organism. In the second work, the authors combined a computer aided diagnosis with continuous attribute discretization and association rule mining for the early diagnosis of AD based on emission computed tomography images. Alternatively, Ponmary et al.used association rule mining to find AD patterns of amino-acid residues in the protein binding site of enzymes which has been described to have a role in AD [30] . It is noteworthy that the AD state (healthy or AD patient) has been fixed to appear in the consequent of the rules handled in this work. Thus, a rule is composed of a set of genes belonging to an interval (expression levels) in the antecedent and the AD state in the consequent. An example of the rules found is as follows: EPHA 10 ∈ [2.0, 2.9] and STRN 4 ∈ [3.0, 3.5] ⇒ AD state is 1 (not healthy). 2.3. Classification and ensembles In Machine Learning and the statistics, the classification goal is to identify to where a new observation belongs in a set of categories. It is important to consider a training set of data composed of observations or instances whose category membership is known. A wide range of classification algorithms can be found in the literature, for instance, Decision Trees, K-nearest Neighbor classifier, Bayesian Network, Neural Networks, Fuzzy Logic, Support Vector Machine, Boosting, etc. Decision Trees [31] are one of the most frequently used in the literature. There are many specific decision-tree algorithms, the most common are, among others, ID3 [32] , C4.5 [33] and CART [31] . The C4.5 algorithm [33] (and its predecessor ID3) is one of the most well-known and used decision trees. The model overfitting in the training dataset is considered as one of the main drawbacks of classification methods. Then, the use of additional techniques such as cross-validation, regularization or pruning, among others, is necessary. Specifically, we have applied a commonly cross-validation method named k-fold crossvalidation. The k-fold cross-validation is a frequently used model validation usually used to evaluate the results obtained by a Data Mining technique, specifically predictive techniques, with the aim at ensuring that the results can be generalized to an independent dataset. A detailed explanation of this common accuracy estimation method can be found in [34] . Ensembles of classifiers enhance the performance of simple classifiers combining the outputs of several others. In the literature, we can find many examples of ensemble systems such as [35,36] . The work presented in [37] proposes a supervised learning approach to the ensemble clustering of genes using known genegene interaction data to improve the results for already commonly used clustering techniques. In the context of the AD problem, an ensemble based in data fusion approach for early diagnosis of AD was proposed [38] . A two stage sequential ensemble is described in [39] to perform the classification of AD based on magnetic resonance imaging features. Most of existing ensembles of the literature are devoted to improve the predictive performance of a single classifier, whereas we use an ensemble of classifiers to obtain the best configuration settings of other classifier. 3. Methodology This section presents the main features of the techniques used to identify genes highly related to AD. The process conducted to detect genes with potential prognosis role in AD is mainly performed through the discovery of AR, in particular QAR, in six phases as can be observed in Fig. 1 . In general terms, a public microarray dataset obtained by laser capture microdissection from the entorhinal cortex was analyzed. The C4.5 algorithm has been applied to consider as minimum threshold the accuracy obtained in the model (first phase). The GarNet algorithm was applied using multiple configuration settings and a set of QAR with AD state fixed as consequent part was obtained for each experiment (second phase). The best experiments were selected in terms of accuracy measure in test sets using the accuracy obtained by C4.5 algorithm as a minimum threshold. GarNet algorithm was rerun using the selected configurations and the set of genes with potential prognosis role in AD were detected. Then, we obtained the top of frequent genes joining all gene-AD state pairs found in the rules (third phase). Gene groups were validated by hierarchical cluster analysis, in addition to statistical and biological significance validation techniques (fourth phase). A biological knowledge integration phase was performed (fifth phase). Finally, the results obtained were validated using additional data (sixth phase). The main features of each phase are detailed in the following Sections 3.1 –3.6 , respectively. 3.1. First phase - decision trees A well-known Machine Learning technique based on decision trees, frequently used to tackle the AD problem, was selected as a benchmark to filter the results obtained by GarNet. Decision trees, also known as classification trees or regression trees, are commonly used as predictive modelling approaches in statistics, Data Mining and Machine Learning. A wide range of decision-trees algorithms can be found in the literature. Specifically, the well-known C4.5 algorithm [33] has been selected to classify between healthy and AD samples due to the high performance usually presented [40] . The average percentage of instances correctly classified obtained by this technique was used as a minimum threshold to select the best configurations of GarNet to perform the second phase of the process. Fig. 1. Overview of general process. 3.2. Second phase - QAR mining process GarNet [16] is a multi-objective algorithm based on the NSGAII algorithm and it is able to find QAR in datasets with continuous attributes avoiding the discretization step. This algorithm aims at solving the main drawbacks caused by a fitness function based on a weighted objective scheme, trying to perform the best trade-off among all the measures optimized. GarNet was executed in this work to address the second phase of the analysis in order to obtain a gene subset highly related with AD. A thorough description of the algorithm can be found in the research work proposed in [16] . Additionally, the main features of GarNet are summarized below. GarNet uses adaptive intervals instead of fixed ranges to group samples whose features share certain sets of values in continuous domain. The search for the most appropriate intervals has been carried out by means of an evolutionary process. In this process, the intervals are adjusted to find QAR with high interpretability, generality, quality and precision. In the population, each individual constitutes a rule, being subjected to an evolutionary process, in which the genetic operators are applied. Furthermore, the instances already covered by the rules are penalized to emphasize the covering of instances not covered yet. Then, those samples covered by a few rules have a higher priority to be selected in order to generate the new population [41] . The evolutionary process ends when the number of generations is reached. The whole evolutionary process is repeated until the desired number of rules is achieved. Finally, GarNet returns the set of QAR found. The GarNet algorithm was executed modifying the minimum QAR confidence threshold obtained and second, optimizing different group of measures following the study proposed in [27,42] . Each minimum confidence threshold combined with each group of measures comprise an experiment. Support, confidence, leverage, gain and accuracy were the selected measures optimized and different groups of 3 measures were composed. Each experiment (each confidence minimum threshold in combination with each group of optimized measures) is executed several times using 5-fold cross-validation on the studied dataset to ensure that the performance of GarNet is stable and accurate. For each experiment, we have mined a model with a defined number of QAR in which the AD state has been fixed as consequent part of the rules and the average global accuracy (validation measure) of the model has been calculated to assess the percentage of correctly classified instances of the dataset [43] . Furthermore, we have computed the average percentage of instances incorrectly classified and the average percentage of instances not classified, that is, the instances not satisfied by any QAR of the model. 3.3. Third phase - selection process of genes with potential prognosis role in AD After completing the first and second phase, we selected those experiments (configuration settings) in which GarNet obtains a higher accuracy value than C4.5 algorithm in the test set using 5fold cross validation. Then, GarNet was executed again using such configuration settings in the original dataset (without using 5-foldcross-validation) in order to obtain a set of QAR that provide information of the entire dataset. Subsequently, all the gene-AD state associations were extracted for each corresponding set of rules. Finally, we selected the most frequent gene-AD state only considering the best rules-based models in order to find potential and relevant genes highly related with AD. This process aims at discovering the most frequent and strongest associations between genes and AD avoiding those relations that occur by chance. Furthermore, the most frequent gene-gene associations among the selected ADrelated genes have been extracted from the obtained rules. To conduct this process, the best rules set obtained for each configuration setting selected were splitted into sets of attribute pairs as follows: •First, genes belonging to the antecedent for each resulting rule were identified. Note that the AD was always fixed as a consequent of the rules. AD state can be 0 or 1, that is, 0 is used to define a healthy patient, while 1 denotes AD patients. •Afterwards, combinations between the genes of the antecedent and the consequent (AD stage) of each rule were performed obtaining gene-AD and gene-gene pairs. •For instance, let the following QAR be: EP HA 10 ∈ [2 , 2 . 9] and T OR 2 A ∈ [1 . 8 , 2 . 6] ⇒ AD stage is 1 The resulting gene-AD pairs (associations) that can be extracted from this rule are: EP HA 10 ⇒ AD stage is 1 T OR 2 A ⇒ AD stage is 1 The gene-gene associations found in this rule are: EP HA 10 ⇒ T OR 2 A After completing the inference process of GarNet for each configuration setting selected ( K number of configurations), the union among all the gene-AD pairs found was performed to find the most frequent gene-AD associations, hence, potential and relevant associations and, finally, gene-gene associations between these genes. Let K be the number of configuration settings used in the inference process. Let be the set of gene-AD pairs and gene-gene pairs obtained from the k th-configuration. The inference process output is defined as: = 1 ∪ 2 ∪ ... K (1) where k , k = 1.. K , is the set of gene-AD and gene-gene pairs obtained from the k th-configuration. The final step involved the selection of the most frequent geneAD associations from the union of the results obtained for the K configuration settings selected. Afterward, the most frequent genegene associations among these genes are found. Note that the configuration settings to perform the gene-AD extraction were selected taking into account a minimum threshold of accuracy using a benchmark method. Specifically, the well-known C.45 algorithm was applied. Finally, expression level changes between healthy and AD patients of the selected gene sets were determined through the intervals of the QAR in which the genes were involved. An example of how the decomposition process of a QAR set into gene-AD and gene-gene pairs is defined as follows. Let the following set of QAR obtained for 2 configuration settings be: 1. Gene-AD associations extraction •Configuration setting 1: –EPHA 10 ∈ [2, 2.9] and TOR 2 A ∈ [1.8, 2.6] ⇒ AD stage is 1 –FAM 158 A ∈ [2.7, 3.0] ⇒ AD stage is 1 –ALOX 12 B ∈ [2.1, 2.4] and VILL ∈ [2.1, 2.3] ⇒ AD stage is 1 –PRPF 40 b ∈ [2.7, 3.0] and DLGAP 2 ∈ [2.6, 2.8] and TAOK 2 ∈ [2.9, 3.3] ⇒ AD stage is 1 1 = { EPHA 10 - AD, TOR 2 A - AD, FAM 158 A - AD, ALOX 12 B - AD, VILL - AD, PRPF 40 B - AD, DLGAP 2 - AD, TAOK 2 - AD, EPHA 10 - TOR 2 A, ALOX 12 B - VILL, PRPF 40 b - DLGAP 2, PRPF 40 b - TAOK 2, DLGAP 2 - TAOK 2} •Configuration setting 2: –EPHA 10 ∈ [2.0, 2.9] and STRN 4 ∈ [3.0, 3.5] ⇒ AD stage is 1 –FAM 158 A ∈ [2.7, 3.0] and RPS 8 ∈ [9.4, 11.6] ⇒ AD stage is 1 –ALOX 12 B ∈ [2.1, 2.4] and PBX 4 ∈ [2.0, 2.4] ⇒ AD stage is 1 2 = { EPHA 10 - AD, STRN 4 - AD, FAM 158 A - AD, RPS 8 - AD, ALOX 12 B - AD, PBX 4 - AD, EPHA 10 - STRN 4, FAM 158 A - RPS 8, ALOX 12 B - PBX 4}  Output = 1 ∪ 2 = { EPHA 10 - AD, TOR 2 A - AD, FAM 158 A - AD, ALOX 12 B - AD, VILL - AD, PRPF 40 B - AD, DLGAP 2 - AD, TAOK 2 - AD, STRN 4 - AD, RPS 8 - AD, PBX 4 - AD EPHA 10 - TOR 2 A, ALOX 12 B - VILL, PRPF 40 b - DLGAP 2, PRPF 40 b - TAOK 2, DLGAP 2 - TAOK 2, EPHA 10 - STRN 4, FAM 158 A - RPS 8, ALOX 12 B - PBX 4} Expression levels: Expression level changes of AD patient genes were calculated considering the lower and upper Table 1 Example of how the gene expression levels are calculated. Gene symbol Avg. expression level QAR intervals Regulation Control AD Lower Upper bound bound EPHA10 3 .8 2 .3 2 .0 2 .9 Down-regulated TOR2A 3 .5 2 .6 1 .8 2 .6 Down-regulated FAM158A 4 .2 2 .8 2 .7 3 .0 Down-regulated ALOX12B 3 .1 2 .3 2 .1 2 .4 Down-regulated VILL 3 .1 2 .2 2 .1 2 .3 Down-regulated PRPF40B 3 .0 2 .9 2 .7 3 .0 Down-regulated DLGAP2 3 .1 2 .7 2 .6 2 .8 Down-regulated TAOK2 3 .8 3 .1 2 .9 3 .3 Down-regulated STRN4 5 .0 3 .4 3 .0 3 .5 Down-regulated RPS8 9 .0 10 .6 9 .4 11 .6 up-regulated PBX4 2 .5 2 .2 2 .0 2 .4 Down-regulated bounds of the QAR obtained by GarNet and the average values of control patients. If the interval bounds of the genes in the rules are below the average value of control patients, the gene is down-regulated in AD patients. Otherwise, the gene is up-regulated. The average values of control and AD patients, in addition to the lower and upper bound of each gene in the QAR detailed in the aforementioned example and the expression level obtained, are presented in Table 1 . 3.4. Fourth phase: validation by hierarchical cluster analysis and statistical tests The fourth phase is devoted to assess the quantitative significance of the genes related to AD obtained by GarNet from different a perspective. To fulfill this goal, hierarchical cluster analysis and statistical tests were used. 3.4.1. Hierarchical cluster analysis A hierarchical cluster analysis was conducted to cluster patients and genes using Spearman and Pearson correlation, respectively. This method is used to assess the capability of the genes obtained by GarNet to classify between healthy and AD patients according to changes in the expression levels (down-regulated or up-regulated). Alternatively, the hierarchical cluster analysis performed for the genes reported in the third phase is used to validate expression levels changes detected by the rules intervals obtained by GarNet. 3.4.2. Statistical and biological significance The Mann-Whitney U -test was applied to determine whether there is a statistically significant difference between expression levels of the reported genes in the previous phase in healthy samples and AD affected samples. Furthermore, the bioinformatics open source software named bioconductor , of the well-known R software environment was used to calculate the volcano plot that visualizes the changes (foldchange) versus the statistical significance ( p -value) of the expression levels of the genes selected by GarNet. The volcano plot has been plotted using the log-fold change and p -value of the topranked genes obtained from a linear model fit of the selected genes. The linear model fit was obtained by the lmFit function, and top-ranked genes were extracted by the topTable function, both included in the limma package of bioconductor (version 3.2). The p -value probability of a differentially expressed gene was computed through the Benjamin and Hochberg statistical tests [44] . The higher the negative log10 for each gene, the higher the probability that the gene is differentially expressed and not a false positive. The x-axis indicates the log2 value of fold-change between the two conditions. 3.5. Fifth phase: biological knowledge integration The fifth phase consisted on the integration of the biological knowledge using different biological sources. First, the altered functions in neurons affected by AD were detected by a systematic review of the literature using the well-known PubMed, a free web literature search service developed and maintained by the National Center for biotechnology Information (NCBI) [45] . Then, an enrichment analysis of the found genes was performed in the context of Gene Ontology, a structured, controlled vocabularies and classifications that cover several domains of molecular and cellular biology, using Fatigoo tool [46] which is integrated in Babelomics 5 analysis suite. The gene set reported by GarNet was also mapped onto the largest network of protein interactions related to Alzheimer’s referred to as AD PPI network . This network was reported by Soler et al. in [47] where 12 well known AD associated genes were selected from OMIM Database as seed. Through yeast two-hybrid matrix screen and two-hybrid library screen, authors generated interaction core set containing all the confirmed library and matrix interactions (200 interactions between 74 nodes including seedseed, seed-candidate and candidate-candidate interactions). This network was merged with direct interactions of seed AD genes extracted from several repositories (IntAct, DIP, MINT and HPRD). As a result, Soler et al. reported an AD network with 1704 nodes and 5881 interactions. In our analyses, the top genes were reported in official gene symbols and were converted to Uniprot accession numbers using DAVID tool [48] . Furthermore, we have analyzed which genes of those obtained by GarNet are associated to cerebral diseases using the well-known MalaCards database, in particular GeneCards [49] . This is an integrated database of all known predictable human genes, including information about diseases connected with each gene. 3.6. Sixth phase: validation using additional data The last phase is devoted to validate the strength of the results obtained by GarNet in the previous phases. To fulfill this goal, several AD-related datasets were used to check the significance of the subset of genes provided by GarNet. In particular, the gene regulation levels were calculated for six datasets collected from NCBI repository [45] . Then, the regulation level of each gene in the original dataset was compared with that presented for the six datasets. Additionally, permutation tests have been applied running the same analysis performed by GarNet multiple times. Specifically, two versions of the original dataset have been generated shuffling the class labels (AD and control patients) of the instances. The percentage distribution of instances of AD and control patients was the same as the original dataset. Then, the second and the third phases of our methodology ( Sections 3.2 and 3.3 ) were applied in the two random versions using the best configuration settings (#8, #13 and #14) described in Section 3.2 . Finally, the results obtained from the permutation tests were compared with those reported for the original dataset in order to check if the accuracy level and the genes selected are the same. 4. Results and discussion The dataset used to perform the analysis proposed is described in Section 4.1 . Furthermore, the results obtained for the six phases are presented and discussed in the following Sections 4.2 –4.7 , respectively. 4.1. Alzheimer dataset The dataset used as training data was retrieved from the gene expression data analysis described in [17] . The original dataset was provided by Dunckley et al. [50] in which 10 0 0 neurons were collected from each of the 33 samples by laser capture microdissection from entorhinal cortex. This single-cell gene expression data is formed by 33 samples and 35,722 probesets. The data were normalized by gcRMA [51] and the resulting probesets were mapped to genes by DAVID [52] . Specifically, the 13 normal controls correspond to the Braak stages 0–II and the average age of patients was 80.1 years. Regarding the AD affected samples, they belong to Braak stages III-IV considered as ‘incipient’ AD and the average age of patients was 84.7. We run our approach on a subset of 1663 genes obtained from this dataset. These 1663 genes were the result of preprocessing the data set of [50] as described in [17] . 4.2. First phase results - decision trees This section details the results obtained by the selected benchmark method. The decision tree built by C4.5 using 5-fold-crossvalidation over the studied dataset is shown below: NDRG 2 < = 2 . 52 : Healthy NDRG 2 < 2 . 52 : AD The first condition, NDRG 2 < = 2.52, corresponds to the control or healthy samples and the second condition, NDRG 2 > 2.52 determines if the samples presents AD. The 87.87% of samples was correctly classified. In particular, the 92.3% and 85% of control and AD samples, respectively, were correctly classified. Although the rate of samples correctly classified is close to 90%, the obtained model by C4.5 provides very poor information since that only one gene appears in the final model. The NDRG 2 gene was removed from the dataset, being the C4.5 algorithm rerun to obtain a different decision tree. In this case, the decision tree built by C4.5 removing this gene was as follows: SMARCD 2 < = 3 . 296 : Healthy SMARCD 2 > 3 . 296 : AD It can be observed that if we desire to achieve a model including a larger set of genes, it is necessary to remove the gene appearing in the decision tree and rerun the C4.5 algorithm. In contrast, AR-based methods can obtain models comprising a high number of genes that may provide useful and relevant knowledge for experts. In order to perform the second phase of the experimentation, the accuracy obtained by C4.5 algorithm was taken as minimum threshold to select the best configurations of GarNet. 4.3. Second phase results: QAR mining process Several experiments were carried out during the second phase to assess the performance of GarNet using multiple configuration settings with the aim at achieving the most optimal solutions for the problem at hand in this work. This section details the sensitivity study combining multiple confidence thresholds and measures of the results achieved by GarNet. Each minimum confidence threshold in combination with each group of optimized measures comprises an experiment. Each experiment was executed 100 times using 5-fold crossvalidation. The minimum thresholds for the confidence measure to be achieved by the rules were 0.8, 0.9 and 1, respectively. Alternatively, we selected the support, confidence, leverage, gain and accuracy measures and different groups of 3 measures have been used in the optimization of GarNet. These measures evaluate different features of QAR. For instance, support and confidence are Table 2 Configuration settings used by GarNet and results obtained applying 5-fold-cross validation. ID Min. conf. Groups of measures optimized Instances (%) Leverage Confidence Gain Support Accuracy Classified Misclassified #1 0 .8   80 .88 18 .98 #2 0 .8   80 .45 19 .26 #3 0 .8   87 .13 11 .26 #4 0 .8   85 .15 14 .37 #5 0 .8   77 .08 22 .62 #6 0 .9   82 .45 17 .34 #7 0 .9   85 .0014 .24 #8 0 .9   87 .80 10 .40 #9 0 .9   86 .92 11 .83 #10 0 .9   72 .82 20 .73 #11 1  85 .15 14 .18 #12 1  85 .01 12 .44 #13 1  88 .80 7 .80 #141  89 .03 8 .26 #15 1  60 .21 10 .35 devoted to assess the generality and the reliability of the rules, respectively. Note that confidence measure was always included in the group of measures optimized by GarNet. Table 2 shows the different configuration settings used to evaluate the performance of GarNet and the results obtained for each one. The configuration number of each experiment is identified in the first column. Second column details the different values for the minimum confidence threshold. Third column provides the group of quality measures optimized by GarNet in each experiment. For each configuration, the objectives optimized by GarNet are marked. Finally, the last column summarizes the results achieved by GarNet according to the percentage of instances classified correctly and percentage of instances not correctly classified, respectively, after performing the 100 executions using 5-fold-cross validation. The remaining percentage of instances until 100% were not classified instances. Note that only the average results obtained in the test set are presented in this table. The best results in terms of instances correctly classified are highlighted in bold. Specifically, we have selected those configuration settings of GarNet that achieve a percentage of correctly classified instances higher or equal than 87.8% according to the results obtained by C4.5. As can be observed, the group of measures composed of confidence, gain and accuracy optimized by GarNet is the best one, regardless the minimum confidence threshold used (configuration settings #3, #8 and #13). However, the best results in terms of instances correctly classified are those obtained by the group of measures leverage, confidence and accuracy when the minimum confidence threshold used is set to 1 (configuration setting #14). In contrast, the worst results werw obtained by the support and accuracy measures when the minimum confidence threshold used is also set to 1 (configuration setting #15). 4.4. Third phase results: genes with potential prognosis role in AD The results obtained by GarNet using the configuration settings #8, #13 and #14 ( Table 2 ), were selected to perform the third phase of the experimentation regarding the accuracy achieved by C4.5. Table 3 shows the top 100 frequent genes sorted by frequency and their expression levels appearing in the rules obtained by GarNet using the aforementioned configuration settings. The regulation of each gene was obtained according to the intervals of the rules discovered by GarNet. Green and purple colors are used to represent if the gene is downregulated or upregulated in AD, respectively. Gene Sym. column shows the Gene Symbol notation of each gene. Freq. column displays the normalized frequency of the presented top frequent genes. The value 1 represents the most frequent gene. Note that duplications were removed since some genes are frequent for both AD and control patient, thus, the final list is composed of 95 genes. Remarkably, 19 genes are up-regulated and 84 are downregulated. Furthermore, the group of genes obtained by C4.5, NDRG 2 (ID 21) and SMARCD 2 (ID 82) also appear in the set of genes obtained by GarNet, demonstrating, the robustness of the algorithm. The resulting network from the most frequent gene-gene interactions extracted using the selected set of AD-related genes is shown in Fig. 2 . Note that green and purple colors are used to identify the downregulated and upregulated genes, respectively, in AD. 4.5. Fourth phase results: validation by hierarchical cluster analysis and statistical tests This section presents and discusses the results obtained in the hierarchical cluster analysis and the statistical and biological validation. 4.5.1. Hierarchical cluster analysis Fig. 3 displays the heatmap of control and AD patients according to the top genes under study after applying the hierarchical cluster analysis. Note that data have been scaled and centered. The results are plotted as dendrograms. Spearman correlation and Pearson correlation have been used to cluster the columns (patients) and rows (genes), respectively. The resulting tree has been cut at specific height (1.5) with its corresponding clusters highlighted in the heatmap color bar. Four groups can be observed according to the patient type and gene expression levels. It is noteworthy that set of genes obtained by GarNet provided two well defined groups that perfectly divide the samples between control and AD patients, as can be observed in the abscissa axis. Note that the labels of patients were clustered by using Pearson correlation. The down-regulated genes are grouped at the top of the heatmap. The up-regulated genes are concentrated at the bottom of the heatmap. Note that the gene regulation levels determined by the rules found by GarNet (defined in Table 3 ) are also consistent with those provided in the heatmap using Pearson correlation. Additionally, the clusters obtained are displayed in Fig. 4 through the representation technique named clusplot [53] available in R software environment. It represents the principal component analysis of the two clusters of genes obtained in the heatmap using Pearson correlation. Red and blue ellipses indicate class boundaries Table 3 Top frequent genes extracted from the rules discovered by GarNet using the best configuration settings sorted by frequency. of the genes, that is, up-regulated and down-regulated, respectively. It can be observed that the two principal components obtained explain more than 95% of the point variability and all genes, except one, are divided into two independent groups or clusters. Thus, GarNet was able to discover a set of genes that successfully classified the samples between control and AD patients regarding their expression levels. Next section details the validation techniques applied to assess the statistical and biological relevance of the set of genes obtained by GarNet. 4.5.2. Statistical and biological significance The Mann-Whitney U -test was used to determine whether there is a statistically significant difference between expression levels in healthy samples and AD affected samples in the resulting subset of genes obtained by GarNet. This non-parametric hypothesis test assess if a particular population (AD samples) tends to have larger values than the other (healthy samples). The results of Mann-Whitney U -test are summarized in Table 4 . Control and AD columns shows the average of the expression levels obtained for each group of patients (healthy and Alzheimer’s patients respectively). Diff. column shows whether the distributions of control and AD samples differed significantly (  ). It can be noted that only 6 of the 95 genes do not significantly differed according to the MannWhitney U -test and 3 of them take the last positions in the ranking. Hence, GarNet is able to find rules containing genes highly related with AD. Alternatively, Fig. 5 shows the volcano plot that pictures the log-fold changes in base 2 on the x-axis versus the negative log of the p-value in base 10 on the y-axis. The most significantly differentially expressed genes, i.e., those that have a p -value lower than 0.05, are represented by a red dot. It can be observed that all the 95 genes presented a p -value lower than 0.05, therefore, all of them are statistically significant. A fold-change of 1.5 cut-off (log2 threshold of 0.58) was applied to identify significant genes due to the good performance presented in previous studies [54] . Genes significantly up-regulated in AD patients are located at right in the graph, and highlighted by a purple box (16 genes). These genes have a fold-change higher than 1.5, that is, a positive log2-fold change greater than 0.58. Genes significantly downregulated in AD patients are located at left in the graph and highlighted in a green box (61 genes). These genes have a fold-change lower than 0.66, that is, a negative log2-fold change lower than −0.58. Both significantly up-regulated and down-regulated genes are labelled with the corresponding Gene Symbol Identification. It can be noted that 77 genes (more than 80%) are biologically significant when taking into account an absolute fold-change higher than 1.5 and 100% of genes are statistically significant according to the minimum p-value considered. Again, we can state that GarNet was able to discover a set of genes highly significant both at statistical and biological level. 4.6. Fifth phase results: biological knowledge integration This phase details the achieved results after performing the biological knowledge integration process. Specifically, the altered functions in neurons affected by AD found in the literature, the Fig. 2. Network of gene-gene interactions among AD-related genes. Green color and purple color are used to identify the downregulated and upregulated genes in AD, respectively. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.) enrichment analysis, the mapping to known gene-disease associations related to AD, validation with additional datasets and permutation tests are presented in the following subsections. 4.6.1. Altered functions in neurons affected by AD As a result of applying rule-based strategy over a dataset of AD patients, we found more than 90 genes that were significantly altered when compared to healthy people. The identified genes covered a high range of functions, both neuron specific as pleiotropic functions. These results highlight the role of AD as a multifactorial syndrome, as it has extensively stated in the literature [5,55,56] . Some of the genes ( CDS1, KLF8, SPTBN1, DDX19A, TSC22D4, GPHN, NDRG2 ) found using GarNet had been previously linked to AD [17,57–62] . These genes cover different roles in neurons, including cell signaling ( CDS1, GPHN, NDRG2 ), cytoskeleton structure ( SPTBN1 ) or gene transcription ( KLF8, TSC22D4 ) [58,63] . Although some of the found genes have been identified in other AD gene array analysis ( CDS1, SPTBN1, DDX19A ) [17,59,60] , no functional information has been described yet. In addition, another additional gene, PIAS3 , encodes a protein that acts as an inhibitor of STAT3 , which has been connected to AD [64,65] . We also found a high number of proteins associated to neuronal processes. Amongst them, GFR α2 expression has been connected with the development or maintenance of cognitive abilities in mice [66] . In that way, GFR α2 altered expression in AD patients may be connected with the progressive loss of memory, one of the hallmarks of the disease. During AD, neurons experiment a loss of its metal homeostasis [67,68] . Regarding this, we found that some genes associated to zing finger proteins are down-regulated in neurons of AD patients (i.e., ZC3H3, ZNF202, ZNF583, Zc3h3 ). Interestingly, zinc levels are lower in the brain of AD patients [69] . However, two different metallothioneins related with zinc homeostasis (MT-1E and MT-1F) were also up-regulated. These proteins are related with heavy metals detoxification and are up-regulated in order to avoid zinc neurotoxicity [68] . This apparent contradiction may be due to differences in Braak staging of neurons used for gene analysis, showing the complexity of this neurodegenerative disease. In that way, there is also ample evidence of an increment of zinc levels in AD patients, as it has been reviewed by Watt et al. [70] . The apparent opposite results obtained by González-Domínguez et al. [69] may be due to a different progression stage of the disease. Also connected with zinc dyshomeostasis, diabetes and cardiac disease have been connected with AD [71–75] . Concerning this, we [30] D. Ponmary Pushpa Latha , D. Joseph Pushpa Raj , Measuring interesting amino acid patterns for alzheimer’s disease related studies targets on the binding site using association rule mining, J. Appl. Pharm. Sci. 3 (7) (2013) 25–30 . [31] L. Breiman , J. Friedman , C.J. Stone , R.A. Olshen , Classification and Regression Trees, CRC press, 1984 . [32] J.R. Quinlan , Induction of decision trees, Mach. Learn. 1 (1) (1986) 81–106 . [33] J.R. Quinlan , C4.5: programs for machine learning, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1993 . [34] R. Kohavi , A study of cross-validation and bootstrap for accuracy estimation and model selection, Morgan Kaufmann, 1995, pp. 1137–1143 . [35] J. Abellán , Ensembles of decision trees based on imprecise probabilities and uncertainty measures, Inform. Fusion 14 (4) (2013) 423–430 . [36] L.I. Kuncheva , Combining pattern classifiers: methods and algorithms, Wiley-Interscience, 2004 . [37] A.K. Rider , G. Siwo , S.J. Emrich , M.T. Ferdig , N.V. Chawla , A supervised learning approach to the ensemble clustering of genes, Int. J. Data Min. Bioinform. 9 (2) (2014) 199–219 . [38] R. Polikar , A. Topalis , D. Parikh , D. Green , J. Frymiare , J. Kounios , C.M. Clark ,An ensemble based data fusion approach for early diagnosis of alzheimer’s disease, Inf. Fusion 9 (1) (2008) 83–95 . [39] M. Termenon , M. Graña , A two stage sequential ensemble applied to the classification of alzheimer’s disease based on mri features, Neural Process. Lett. 35 (1) (2012) 1–12 . [40] S. Mestizo Gutiérrez , M. Herrera Rivero , N. Cruz Ramírez , E. Hernández , G.E. Aranda-Abreu , Decision trees for the analysis of genes involved in alzheimer’s disease pathology, J. Theor. Biol. 357 (0) (2014) 21–25 . [41] M. Martínez-Ballesteros , S. Salcedo-Sanz , J. Riquelme , C. Casanova-Mateo , J. Camacho , Evolutionary association rules for total ozone content modeling from satellite observations, Chemomet. Intell. Lab. Syst. 109 (2) (2011) 217–227 . [42] M. Martínez-Ballesteros , J. Riquelme , Analysis of measures of quantitative association rules, in: Proceedings of the International Conference on Hybrid Artificial Intelligent Systems, in: Lecture Notes in Computer Science, 6679, 2011, pp. 319–326 . [43] I.H. Witten , E. Frank , Data mining: practical machine learning tools and techniques, Morgan Kaufmann Series in Data Management Systems, 2nd, Morgan Kaufmann, 2005 . [44] Y. Benjamini , Y. Hochberg , Controlling the false discovery rate: a practical and powerful approach to multiple testing, J. R. Stat. Soc. Series B (Methodol.) 57 (1) (1995) 289–300 . [45] Pubmed resource, 2015, ( http://www.ncbi.nlm.nih.gov/pubmed/ ). [Online; accessed in October 2015]. [46] F. Al-Shahrour , P. Minguez , J. Tárraga , et al. , FatiGO: a functional profiling tool for genomic data. Integration of functional annotation, regulatory motifs and interaction data with microarray experiments, Nucleic Acids Res. 35 (2007) W91–W96 . [47] M. Soler-López , A. Zanzoni , R. Lluís , U. Stelzl , P. Aloy , Interactome mapping suggests new mechanistic details underlying Alzheimer’s disease, Genome Res. 21 (3) (2011) 364–376 . [48] David tools, 2016, ( https://david-d.ncifcrf.gov/ ). [Online; accessed in May 2016]. [49] N. Rappaport, N. Nativ, G. Stelzer, M. Twik, Y. Guan-Golan, T.I. Stein, I. Bahir, F. Belinky, C.P. Morrey, M. Safran, D. Lancet, Malacards: an integrated compendium for diseases and their annotation, Database 2013 (2013), doi: 10.1093/ database/bat018 . [50] T. Dunckley , T. Beach , K. Ramsey , A. Grover , D. Mastroeni , D. Walker , B. LaFleur , K. Coon , K. Brown , R. Caselli , W. Kukull , R. Higdon , D. McKeel , J. Morris , C. Hulette , D. Schmechel , E. Reiman , J. Rogers , D. Stephan , Gene expression correlates of neurofibrillary tangles in alzheimer’s disease., Neurobiol. Aging 27 (2006) 1359–1371 . [51] R. Irizarry , Z. Wu , H. Jaffee , Com parison of affymetrix genechip expression measures., Bioinformatics 22 (2006) 789–794 . [52] G. Dennis , B. Sherman , D. Hosack , J. Yang , W. Gao , H. Lane , R. Lempicki , David: database for annotation, visualization, and integrated discovery., Genome Biol. 4 (2003) P3 . [53] G. Pison , A. Struyf , P.J. Rousseeuw , Displaying a clustering with clusplot, Comput. Stat. Data Anal. 30 (4) (1999) 381–392 . [54] A .A . Margolin , S.-E. Ong , M. Schenone , R. Gould , S.L. Schreiber , S.A. Carr , T.R. Golub , Empirical bayes analysis of quantitative proteomics experiments, PLoS ONE 4 (10) (2009) e7454 . [55] K. Iqbal , I. Grundke-Iqbal , Alzheimer’s disease, a multifactorial disorder seeking multitherapies, Alzheimer’s Dementia 6 (5) (2010) 420–424 . [56] M. Storandt , D. Head , A.M. Fagan , D.M. Holtzman , J.C. Morris , Toward a multifactorial model of alzheimer disease, Neurobiol. Aging 33 (10) (2012) 2262–2271 . [57] L. Zhang , X. Ju , Y. Cheng , X. Guo , T. Wen , Identifying tmem59 related gene regulatory network of mouse neural stem cell from a compendium of expression profiles, BMC Syst. Biol. 5 (1) (2011) 1–12 . [58] R. Yi , B. Chen , J. Zhao , X. Zhan , L. Zhang , X. Liu , Q. Dong , Kruppel-like factor 8 ameliorates alzheimer’s disease by activating β-catenin, J. Mol. Neurosci. 52 (2) (2014) 231–241 . [59] M. Ray , W. Zhang , Analysis of alzheimer’s disease severity across brain regions by topological analysis of gene co-expression networks, BMC Syst. Biol. 4 (1) (2010) 1–11 . [60] M.-G. Hong , A. Alexeyenko , J.-C. Lambert , P. Amouyel , J.A. Prince , Genome-wide pathway analysis implicates intracellular transmembrane protein transport in alzheimer disease, J. Hum. Genet. 55 (10) (2010) 707–709 . [61] C.M. Hales , H. Rees , N.T. Seyfried , E.B. Dammer , D.M. Duong , M. Gearing , T.J. Montine , J.C. Troncoso , M. Thambisetty , A.I. Levey , J.J. Lah , T.S. Wingo ,Abnormal gephyrin immunoreactivity associated with alzheimer disease pathologic changes, J. Neuropathol. Exp. Neurol. 72 (11) (2013) 1009–1015 . [62] C. Mitchelmore , S. Buchmann-Moller , L. Rask , M.J. West , J.C. Troncoso , N.A. Jensen , Ndrg2: a novel alzheimer’s disease associated protein, Neurobiol. Dis. 16 (1) (2004) 48–58 . [63] A. Jones , K. Friedrich , M. Rohm , M. Schafer , C. Algire , P. Kulozik , O. Seibert , K. Muller-Decker , T. Sijmonsma , D. Strzoda , C. Sticht , N. Gretz , G.M. Dallinga-Thie , B. Leuchs , M. Kogl , W. Stremmel , M.B. Diaz , S. Herzig , Tsc22d4 is a molecular output of hepatic wasting metabolism, EMBO Mol. Med. 5 (2) (2013) 294–308 . [64] C. Chung , J. Liao , B. Liu , X. Rao , P. Jay , P. Berta , K. Shuai , Specific inhibition of stat3 signal transduction by pias3, Science (New York, N.Y.) 278 (5344) (1997) 1803–1805 . [65] C. Chung , J. Liao , B. Liu , X. Rao , P. Jay , P. Berta , K. Shuai , Tyk2/stat3 signaling mediates β-amyloid-induced neuronal cell death: implications in alzheimer’s disease, J. Neurosci. 30 (20) (2010) 6 873–6 881 . [66] V. Voikar , J. Rossi , H. Rauvala , M.S. Airaksinen , Impaired behavioural flexibility and memory in mice lacking gdnf family receptor α2, European. J. Neurosci. 20 (1) (2004) 308–312 . [67] M.C. Carreiras , E. Mendes , M.J. Perry , A.P. Francisco , J. Marco-Contelles , The multifactorial nature of alzheimer’s disease for developing potential therapeutics., Curr. Top. Med. Chem. 13 (15) (2013) 1745–1770 . [68] D. Mizuno , M. Kawahara ,The molecular mechanisms of zinc neurotoxicity and the pathogenesis of vascular type senile dementia, Int. J. Mol. Sci. 14 (11) (2013) 22067–22081 . [69] R. González-Domínguez , T. García-Barrera , J.L. Gómez-Ariza , Homeostasis of metals in the progression of alzheimer’s disease, BioMetals 27 (3) (2014) 539–549 . [70] I.J.W. Nicole T. Watt , N.M. Hooper , The role of zinc in alzheimer’s disease, Int. J. Alzheimers Dis. 2011 (2011) 1–10 . [71] C. Rosendorff, M. Beeri , J. Silverman , Cardiovascular risk factors for Alzheimer’s disease, Am. J. Geriatr. Cardiol. 16 (3) (2007) 143–149 . [72] R. Stewart , Cardiovascular factors in Alzheimer’s disease, J. Neurol., Neurosurg. Psychiatry 65 (2) (1998) 143–147 . [73] S. Craft , S. Watson , Insulin and neurodegenerative disease: shared and specific mechanisms, Lancet Neurol. 3 (3) (2004) 169–178 . [74] C. MacKnight , K. Rockwood , E. Awalt , I. McDowell , Diabetes mellitus and the risk of dementia, alzheimer’s disease and vascular cognitive impairment in the canadian study of health and aging, Dement Geriatr Cogn Disord. 14 (2002) 77–83 . [75] J. Janson , T. Laedtke , J. Parisi , P. O’Brien , Petersen , R.C. , P. Butler , Increased risk of type 2 diabetes in alzheimer disease, Diabetes 53 (2) (2004) 474–481 . [76] R. Graf , M. Munschauer , G. Mastrobuoni , F. Mayr , U. Heinemann , S. Kempa , N. Rajewsky , M. Landthaler , Identification of lin28b-bound mrnas reveals features of target recognition and regulation, RNA Biol. 10 (7) (2013) 1146–1159 . [77] E. Ahmady , S.A. Deeke , S. Rabaa , L. Kouri , L. Kenney , A.F.R. Stewart , P.G. Burgon , Identification of a novel muscle a-type lamin-interacting protein (mlip), J. Biol. Chem. 286 (22) (2011) 19702–19713 . [78] S. Sekar , J. McDonald , L. Cuyugan , J. Aldrich , A. Kurdoglu , J. Adkins , G. Serrano , T.G. Beach , D.W. Craig , J. Valla , E.M. Reiman , W.S. Liang , Alzheimer’s disease is associated with altered expression of genes involved in immune response and mitochondrial processes in astrocytes, Neurobiol. Aging 36 (2) (2015) 583–591 . [79] X.-F. Chen , Y.-w. Zhang , H. Xu , G. Bu , Transcriptional regulation and its misregulation in alzheimer’s disease, Mol. Brain 6 (1) (2013) 1–9 . [80] W.S. Liang , T. Dunckley , T.G. Beach , A. Grover , D. Mastroeni , K. Ramsey , R.J. Caselli , W.A. Kukull , D. McKeel , J.C. Morris , C.M. Hulette , D. Schmechel , E.M. Reiman , J. Rogers , D.A. Stephan , Altered neuronal gene expression in brain regions differentially affected by alzheimer’s disease: a reference data set, Physiol. Genomics 33 (2) (2008) 240–256 . [81] B.E.L. Lauffer , R. Mintzer , R. Fong , S. Mukund , C. Tam , I. Zilberleyb , B. Flicke , A. Ritscher , G. Fedorowicz , R. Vallero , D.F. Ortwine , J. Gunzner , Z. Modrusan , L. Neumann , C.M. Koth , P.J. Lupardus , J.S. Kaminker , C.E. Heise , P. Steiner , Histone deacetylase (hdac) inhibitor kinetic rate constants correlate with cellular histone acetylation but not transcription and cell viability, J. Biol. Chem. 288 (37) (2013) 26926–26943 .