Optimizing Dementia Diagnosis Through Distance-Correlation Feature Space and Dimensionality Reduction
Full text
This is a submitted version of the following article published by World Scientific Publishing: Zubasti, P., Patricio, M. A., Berlanga, A., & Molina, J. M. (2025). Optimizing Dementia Diagnosis Through Distance-Correlation Feature Space and Dimensionality Reduction. International Journal of Neural Systems, 35(09), 2550042. . DOI: https://doi.org/10.1142/S012906572550042X
December 17, 2025 17:49 output Optimizing Dementia Diagnosis through Distance-Correlation Feature Space and Dimensionality Reduction Pablo Zubasti Miguel A. Patricio Department of Computer Science and Engineering Department of Computer Science and Engineering Universidad Carlos III de Madrid Universidad Carlos III de Madrid Colmenarejo, 28270, Spain Colmenarejo, 28270, Spain E-mail: [email protected] E-mail: [email protected] ORCiD: 0009-0006-1906-118X ORCiD: 0000-0002-9304-826X Antonio Berlanga Jos´e M. Molina Department of Computer Science and Engineering Department of Computer Science and Engineering Universidad Carlos III de Madrid Universidad Carlos III de Madrid Colmenarejo, 28270, Spain Colmenarejo, 28270, Spain E-mail: ab[email protected] E-mail: [email protected] ORCiD: 0000-0002-5564-399X ORCiD: 0000-0002-7484-7357 The reduction of dimensionality in Machine Learning and Artificial Intelligence problems constitutes a pivotal element in the simplification of models, significantly enhancing both their performance and execution time. This process enables the generation of results more rapidly while also facilitating the scalability and optimization of systems that rely on such models. Two primary approaches are commonly employed to achieve dimensionality reduction: feature selection-based methods and those grounded in feature extraction. In the present paper, we propose a Distance-Correlation Feature Space, upon which we define a dimensionality reduction algorithm based on space transformations and Graph Embeddings. This methodology is applied in the context of dementia diagnosis through learning models, with the overarching objective of optimizing the diagnostic process. Keywords: Dimensionality reduction; Distance-Correlation; Graph Embeddings; Dementia Diagnosis 1. Introduction Dimensionality reduction represents one of the main areas of study within data preprocessing techniques that underpin subsequent Machine Learning methods. The increase in the volume of data1–3 and its dimensionality results in the generation of massive datasets, which are subsequently used by supervised and unsupervised models. The computational cost associated with Machine Learning and data analysis algorithms is considerably high, with most exceeding O(n2) and others easily reaching O(n4) in the worst cases,4and when combined with the large volume of data, this leads to significant temporal costs, rendering such models impractical for use in rapid response systems, such as real-time applications. Cognitive impairments diseases (such as Dementia or Alzheimer) represent a significant challenge in healthcare, affecting the daily functioning of individuals and serving as potential indicators of neurodegenerative diseases. Early and accurate classification of cognitive impairments is crucial for timely intervention and treatment planning. However, the complexity and heterogeneity of cognitive decline make traditional diagnostic approaches challenging. In the fields of medicine and neuroscience, Artificial Intelligence (AI) techniques are increasingly used in the development of tools capable of diagnosing,5–10 and even helping to explain symptoms that a patient 1
December 17, 2025 17:49 output may exhibit. A significant p ortion o f t he research focuses on the application of AI in the automated classification o f A lzheimer’s D isease ( AD). Several studies employ deep learning models, such as convolutional neural networks (CNN), to analyze medical imaging data, particularly magnetic resonance imaging (MRI).10, 11 These models demonstrate high accuracy in distinguishing between different cognitive states, such as cognitively normal (CN), early mild cognitive impairment (EMCI), and late mild cognitive impairment (LMCI). Some studies emphasize the importance of model explainability, employing techniques such as Grad-CAM to visualize critical regions in brain scans that contribute to classification d ecisions.11 Furthermore, ensemble machine learning models have been proposed as an alternative to deep learning, achieving comparable or even superior performance in AD classification while requiring less training data.10 Furthermore, an alternative line of research investigates the potential of electroencephalography (EEG) as a non-invasive and cost-effective t ool for Alzheimer’s diagnosis.12–16 These studies propose methodologies that use wavelet coherence and functional connectivity analysis to identify alterations in brain activity associated with neurodegeneration. In particular, the feasibility of using a reduced EEG montage with only four channels is explored, demonstrating accuracy comparable to traditional multichannel setups.13 This finding s uggests a promising avenue for portable and accessible early-stage diagnosis solutions. In addition, wavelet coherence models have been employed to analyze cortical connectivity with high temporal resolution, revealing significant differences i n c oherence p atterns b etween AD patients and healthy controls.12 Beyond machine learning applications, another set of studies reviews the role of multimodal imaging techniques and advanced classification algorithms in improving the precision of Alzheimer’s disease detection.17 These reviews highlight the importance of integrating different i maging m odalities, s uch as MRI and positron emission tomography (PET), to develop robust diagnostic biomarkers. More recent and powerful classification t echniques, s uch a s enhanced probabilistic neural networks, have been suggested to improve detection accuracy. Furthermore, research on graph-based complexity metrics has explored their applicability to detect neurodegenerative disorders by analyzing brain connectivity.18 These studies propose new complexity measures, such as Graph Index Complexity and Off-Diagonal Complexity, which have been successfully applied to both aging-related conditions and autism spectrum disorder. Overall, these studies underscore the critical role of AI-driven methodologies in cognitive impairment research. However, high-dimensional data often lead to overfitting and reduced interpretability.Our study addresses this limitation by proposing a dimensionality reduction framework that enhances classification accuracy while preserving the interpretability of key cognitive features. In the healthcare domain, the challenge of interpretability is of paramount importance, as decisions cannot be effectively made if they are not articulated in comprehensible terms. A doctor cannot rely on a decision that lacks a clear explanation, and similarly, a patient is unlikely to have confidence in a specialist who relies on results derived solely from computational methods.19,20 Explainable Artificial Intelligence (XAI)11,21–23 focuses on this last aspect, striving to create models that aid in the attribution of importance values to the features used. This approach facilitates an understanding of the significance and relevance of certain attributes in the context of the problem, as well as potential interactions between them and the target variable, among other factors. In this paper, a new method is proposed on reduction in dimensionality through the Distance-Correlation Feature Space (DCFS) and Graph Embedding, to correspond to the XAI field. Distance correlation provides a more informative measure of the relationships between features, taking into account both linear and nonlinear dependencies. This allows for a more meaningful feature space, where feature importance can be more accurately assessed. The use of graph embeddings translates feature relationships into a Euclidean space, making patterns and dependencies among variables more interpretable. The dimensionality reduction algorithm selectively merges highly related features while preserving their contribution to the model, simplifying the decision-making process, and improving the interpretability of the final predictions. In summary, in this paper we present a new method of dimensionality reduction in problems applied to the diagnosis of cognitive impairment dis-
December 17, 2025 17:49 output eases. Dimensionality reduction has the following advantages in the study of this type of disease: (1) Ability to eliminate variables that have no causal relationship with the objective. In this case, it may imply a reduction in the number of diagnostic tests. (2) It allows for the creation of less complex models with greater explanatory capacity. The paper analyzes the results of being applied to two wellknown public datasets: OASIS-224,25 dataset and the UDS (Uniform Data Set) dataset from the National Alzheimer’s Coordinating Center (NACC).26–28 2. Related work 2.1. Dimensionality reduction algorithms Dimensionality reduction techniques are divided into two main categories: feature selection and feature extraction.29 Techniques based on feature selection are often linked to explainability,30,31 as they involve preliminary analyses of attribute importance, which requires understanding of the relevance of each attribute to determine its inclusion or exclusion. On the other hand, dimensionality reduction methods based on feature extraction,32,33 such as Principal Component Analysis (PCA),34,35 produce sets of transformed variables that generate lowerdimensional spaces, often at the expense of the original meaning of the variables. This is because new variables are created during the transformation process, without directly selecting from the original set. Identifying a way to reduce dimensionality through transformations requires that such transformations be grounded in the relationships between the original attributes, which implicitly demands a prior understanding of both the interactions and importance of these variables. In this paper, we propose the definition of a novel feature space, termed Distance-Correlation Feature Space, upon which we introduce a dimensionality reduction algorithm based on transformations of the original data, leveraging Graph Embedding techniques. Previous studies have explored the potential of using Graph Embedding techniques36 to identify groups of interacting variables in medical applications, such as the detection of dementia through Machine Learning. In the present work, the authors define a flexible methodology that projects the attributes of a dataset - containing information on patients with and without dementia - onto a Euclidean space via a Graph Embedding algorithm, creating a vector representation that encodes the relationships between the variables as distances in this space. This vector representation in the embedding allows the application of classical data analysis and Machine Learning techniques, from which results are derived to identify groups of variables that exhibit similar behavior or impact predictive models in comparable ways. 2.2. Graph Embeddings Graph embedding techniques refer to a set of approaches that enable the representation of information stored in a graph within a δ-dimensional Euclidean space (typically of low dimension).37–39 The goal is to vectorize a graph such that each vertex is represented as a point (vector) in the embedding space. Classical data analysis and Artificial Intelligence techniques can then be applied to these embeddings. Among the most popular approaches for building embeddeds are those based on factorization, random walks, and neural networks. 3. Theory background The study of graph analysis has significantly advanced in the realm of Machine Learning, with graph embedding techniques playing a critical role in transforming graphs into a δ-dimensional space. This transformation simplifies the application of algorithms for tasks such as classification, regression, and clustering. Various methodologies have been proposed for the generation of graph embeddings, with the primary ones outlined in a review of graph embedding techniques.40 These methods are largely based on factorization, random walks, or deep learning. Our work specifically explores factorization-based approaches, with a focus on their potential applications in medicine and explainable AI (XAI). Factorization-based graph embedding algorithms first emerged in the 2000s, with the aim of representing each graph node as a vector and reducing dimensionality. One such algorithm is Local Linear Embedding (LLE),41 which utilized a quadratic objective function built from a neighborhood graph and solved it by calculating its eigenvectors.
December 17, 2025 17:49 output 3.1. Laplacian Eigenmaps The LLE principle was first introduced by the Laplacian Eigenmaps42 in 1971, where the idea behind the embedding problem was to optimize the following equation: min xJ(x) = |V| X i=1 |V| X j=1 ||xi−xj||2wij, s.t., |V| X i=1 ||xi||2= 1, |V| X i=1 xi= 0 (1) where xiand xjare vectors of the embedding space δ−dimensional and wij are the associated weights between the nodes iand j. The constraints defined for the optimization problem avoid making every xi= 0. The main idea behind equation (1) is that the large distances between pairs of vectors xiand xj in the embedding space correspond to small weigths wij in the original graph (associated with the edge eij that connects the nodes viand vj) and vice versa. This property is commonly referred to as the local topology preserving and mathematically is defined as follows: if wij ≥wpq →(xi−xj)2≤(xp−xq)2,∀i, j, p, q (2) This basically states that the larger (or smaller) a weight wij is in the original graph, the closer (or farther) xiand xjshould be in the embedding space. The usual approach required to solve the optimization problem presented in equation (1) is based primarily on the eigendecomposition of the Laplacian matrix (3) representing the graph. Mathematically: Lij ≜ deg(vi)if i =j −1if i =j and viis adjacent to vj 0otherwise L=D−W (3) where Dis the diagonal matrix that contains the degree of each corresponding vertex and Wis the adjacency matrix with the graph weights. The Laplacian Eigenmap embedding will be computed as follows: (1) Compute the laplacian matrix for the original graph: L=D−W (2) Normalize the laplacian matrix with the following: ˆ Lij ≜ 1if i =j and ki= 0 −1 √kikj if i =j and adjacent(ni, nj) 0otherwise ˆ L= (D+)1/2L(D+)1/2 where kirepresents the degree of the ith-node and D+represents the Moore-Penrose pseudoinverse. (3) Compute the eigenvalues and eigenvectors of the normalized Laplacian matrix ˆ L: Av =λv ⇒ˆ Lv=λDv (4) Select the first n-eigenvalues (and their corresponding n-eigenvectors) that are greater than zero. 3.2. Cauchy Graph Embeddings The Laplacian Eigenmaps42 were successful in treating large distances between pairs of vectors because the quadratic expression mentioned above emphasized more the distances that were significant rather than those that were smaller. This false notion of local topology preserving was explained by Luo et al., where they proposed a new version of the Laplacian Eigenmaps algorithm where the optimization function was modified so that short distances in the embedding space were also significant and local topology preserving was really accomplished. min xJ(x) = |V| X i=1 |V| X j=1 (xi−xj)2 (xi−xj)2+σ2wij, s.t., |V| X i=1 x2 i= 1, |V| X i=1 xi= 0 (4) The previous expression can be simplified with the following operations: min x |V| X i=1 |V| X j=1 (xi−xj)2 (xi−xj)2+σ2wij = min x 1− |V| X i=1 |V| X j=1 σ2 (xi−xj)2+σ2wij =
December 17, 2025 17:49 output max x |V| X i=1 |V| X j=1 σ2 (xi−xj)2+σ2wij = max x |V| X i=1 |V| X j=1 wij (xi−xj)2+σ2 Expressed in a general δ-dimensional embedding, we conclude: max RJ(R) = |V| X i=1 |V| X j=1 wij ||ri−rj||2+σ2 s.t., RRT=I, Re =¯ 0 (5) Optimizing Equation (5) allows us to compute the Cauchy Embedding of a given weighted graph. The authors of the original paper proposed an iterative algorithm that converges with the embedding solution after a few iterations. The embedding dimension is one of the key hyperparameters of the CGE algorithm. It is freely selectable, although typically representable dimensions (2D and 3D) are chosen, as they not only leverage embedding techniques but also allow for data visualization. The selection of the desired dimensions follows a process similar to Laplacian Eigenmaps, where the first n-eigenvalues and their corresponding n-eigenvectors greater than zero are selected. The desired embedding dimensions are then chosen in ascending order, akin to the approach used in PCA when selecting a subset of the principal components. 3.3. Distance Correlation The distance correlation43 is a metric that measures the statistical dependence between two vectors of arbitrary and not necessarily equal dimensions. It is established that if the correlation coefficient is equal to zero, then the two random vectors are independent. Based on these principles, distance correlation enables the detection of both linear and nonlinear associations. Compared to Pearson’s correlation, distance correlation offers a more flexible framework, as Pearson’s method can only capture linear dependencies between two random variables. To understand the fundamentals of distance correlation, it is necessary first to define distance covariance, which is essential for calculating the final metric. Consider two vectors whose values are the result of sampling two random variables, Xand Y. The first step in computing distance covariance involves calculating distance matrices (6): ajk =||Xj−Xk|| :j, k ∈ {1,2, . . . , n} bjk =||Yj−Yk|| :j, k ∈ {1,2, . . . , n}(6) where ||λ||, λ ∈Rndenotes the Euclidean norm. Then doubly-centered distances must be computed for each element of the distance matrices: Ajk ←ajk −¯aj·−¯a·k+ ¯a·· Bjk ←bjk −¯ bj·−¯ b·k+¯ b·· ¯aj·=1 n n X i=1 aji,¯a·k=1 n n X i=1 aik ¯a·· =1 n2 n X i=1 n X j=1 aij,¯ b·· =1 n2 n X i=1 n X j=1 bij (7) Thus, the distance covariance is computed as equation 8 states: dCov2 n(X, Y )≜1 n2 n X j=1 n X k=1 AjkBjk (8) Distance covariance can also be defined, as shown in (9). dCov2 n(X, Y ) = ZRp+q |φX,Y (s, t)−φX(s)φY(t)|2 cpcq|s|1+p p|t|1+q q dtds (9) where φX,Y (s, t), φX(s) and φY(t) are the characteristic functions of (X, Y ), Xand Y,pand qdenote the Euclidean dimension of Xand Y. On the other hand, cpand cqare constants. Having defined the distance covariance, we proceed to define the distance variance as shown in Equation 10. dV ar2 n(X) = dCov2 n(X, X) = 1 n2 n X j=1 n X k=1 A2 jk (10) With all components now computed, we proceed to define the formula for distance correlation, following the structural framework of Pearson’s correlation (11). dCorr2 (X, Y )=dCov2 n(X, Y ) pdV ar2 n(X)dV ar2 n(Y)(11)
December 17, 2025 17:49 output The authors of distance correlation outlined a set of relevant properties in their paper, among which we highlight the following. (1) 0 ≤dCorr(X, Y )≤1 (2) dCorr(X, Y ) = 0 iif Xand Yare independent (3) dCorr(X, Y ) = 1 implies that Y=A+bCX where Arepresents some vector, brepresents a scalar and Crepresents an orthonormal matrix. Recent studies explore newer properties of distance covariance (and correlation) and find similarities with Pearson’s correlation metric.44 In figure 1, we observe various data distributions (following linear and non-linear patterns), which illustrate the results obtained from Pearson’s correlation and distance correlation. As can be observed, distance correlation tends toward 1 when the correlation exhibits linear characteristics (property 3), while non-linear relationships are still captured, albeit with lower intensity. Compared to Pearson’s correlation, nonlinear relationships are better detected by distance correlation, as seen in figure 1, where Pearson produces a risky linear correlation value for the data distribution (which follows the polynomial P(x) = x3). 4. Distance-Correlation Feature Space In this paper, we propose what we refer to as the Distance-Correlation Feature Space (DCFS), which will be used below. The goal of DCFS is to represent the variables (attributes) of a problem in a Euclidean space, enabling the application of conventional Machine Learning algorithms. The DCFS must preserve the property described in Equation (2) once it has been represented as an embedding. The purpose of representing the variables as vectors in the space is that if the property (2) is maintained, the relationships selected to construct the embedding will be quantifiable in terms of distance. The procedure we propose to represent the variables as an embedding involves the following steps (shown in Figure 2): First, it is necessary to construct a weighted graph G= (V,E, W) where the set of vertices Vrepresents the variables in the problem, and the set of edges Erepresents the relationships between these variables. This requires building an adjacency matrix in which “all-to-all” relationships are studied, resulting in a fully connected graph of type Knwhere the edges are weighted (the weights are stored in the weighted adjacency matrix W). The relationships between variables (edges) must be numerical values that quantify the intensity (initially, both negative and positive considerations are taken into account) of those relationships. Originally, the methodology proposed in36 did not formally establish which metric should be used to assign weights to the edges of the graph, although Pearson and Spearman correlations were suggested as ways to measure the relationships between relevant attributes to quantify how important the variables are relative to the target variable, as well as how related the variables are to each other. However, in the current proposal, we unconditionally establish the use of distance correlation as the fundamental metric for constructing the weighted graph adjacency matrix A(12). A= dCorr(X1, X1). . . dCorr(X1, Xn) dCorr(X2, X1). . . dCorr(X2, Xn) . . ..... . . dCorr(Xn, X1). . . dCorr(Xn, Xn) (12) The properties outlined in the previous section of this document provide the rationale for using distance correlation as the fundamental metric. Specifically, the ability to capture both linear and nonlinear relationships surpasses the use of Pearson’s correlation as a metric. Additionally, the property that describes the range of values returned by distance correlation is highly significant. In Pearson’s correlation, the values are restricted to a bounded interval between -1 and +1, implying the possibility of obtaining negative correlations. In the context of graph construction, this results in edges with negative weights. If property (2) must be satisfied in all cases, a negative edge weight would naturally imply a lower value than a positive one, which would lead to the vertices being represented as widely separated in the embedding space. However, two variables that are highly negatively correlated are still highly relevant to the problem, as the correlation is significant despite showing an inverse relationship. Representing these vertices as far apart in the embedding space would imply the elimination of their relationship, since negative distances do not exist in such
December 17, 2025 17:49 output (a) (b) (c) (d) Figure 1. Examples of distance correlation: (a) linear, (b) quadratic, (c) cubic, (d) sinusoidal. Figure 2. Distance-Correlation Feature Space pipeline. spaces. Therefore, a significant quantitative separation would imply no relationship between attributes, which is inaccurate when considering the original graph information. The authors of36 propose squaring the individual values in the correlation matrix, which ostensibly resolves the issue of negative edges. This approach effectively treats a negative correlation c− ij as identical to a positive correlation c+ kl of the same magnitude |cij|=|ckl|, thus removing the original information provided by the correlation matrix. In contrast, distance correlation returns values bounded within the interval [0, 1], eliminating the need for external transformations of the adjacency matrices and thus completely avoiding the loss of original information. The relationship between the values returned by distance correlation and conditional dependence is particularly valuable in the context of the subsequent explainability to be applied to the DCFS embeddings. For example, groups of variables that are closely positioned in the embedding space not only exhibit some form of correlation, but may also suggest potential probabilistic dependence, which could offer insights into possible cause-and-effect explanations. The second and final step in constructing a DCFS involves applying the Cauchy Graph Embedding (CGE) algorithm to the adjacency matrix built using the distance correlation metric. There is the possibility of using alternatives to the CGE algo-
December 17, 2025 17:49 output rithm, such as variations of Laplacian Eigenmaps. The rationale behind selecting the CGE algorithm is that the authors of the original paper guarantee better preservation of distant relationships (edges with higher weights), a feature that Laplacian Eigenmaps were unable to achieve correctly. 5. Dimensionality reduction algorithm The second contribution of this paper consists of a dimensionality reduction algorithm that is based on the previously defined D CFS. T his a lgorithm provides a method for reducing the dimensionality of a given problem (both classification a nd regression) based on transformations of the original attributes, utilizing the information provided by the DCFS as “guidance” for the fusion process that will yield the new attributes. The algorithm comprises the following steps: (1) First, apply k-Means clustering and the CalinskiHarabasz index,45 which can be computed as follows: s=Tr(Bk) Tr(Wk)×nE−k k−1 Wk= k X q=1 X x∈Cq (x−cq)(x−cq)T Bk= k X q=1 nq(cq−cE)(cq−cE)T Where nErepresents the number of instances of a given set of data E, that has been clustered in k-groups. Cqrepresents the set of instances that form the qth-cluster, cqrepresents the centroid of the qth-cluster, cErepresents the centroid of the whole set of data Eand nqis the number of instances of the qth-cluster. The Tr(·) refers to the trace operator. This metric will be used to determine the optimal number of clusters (excluding the target variable during the clustering process). It is important to notice that the k-Means algorithm is only applied to the input variables, and doesn’t build groups with the target variable. (2) For each identified group, calculate the inverse of the Euclidean distance from each instance within the cluster to the target variable (a very small constant ϵshould be added to the denominator to prevent division by zero): Iq i=1 dq i+ϵ dq i=v u u t δ X j=1 (xq ij −tj)2 Where Iq irepresents the inverse of the distance, dq irepresents the euclidean distance between a point xq iand the target variable t. (3) Using the sum of the inverse distances for each group, a weighting factor is calculated for each instance, assigning a higher weight to those closer to the target variable and vice versa: fq i=Iq i Iq=Iq i Pi∈CqIq i (4) By scaling the original variables of the problem (outside the DCFS) into a [0, 1] range (using Min-Max scaling), a weighted linear combination is performed using the weights derived from the embedding information to create a new merged variable for each group of variables resulting from the clustering process: X(i) scaled =X(i)−X(i) min X(i) max −X(i) min X(q) new =X i∈Cq fq iX(i) scaled X(i)= x(i) 0 x(i) 1 . . . x(i) n−1 The underlying concept of the algorithm is to apply dimensionality reduction while sacrificing minimal predictive quality in the models used on the data. Representing the topology in the form of an embedding allows the relationships between variables to be mapped as distances in a Euclidean space (Figure 3), enabling the application of the k-means clustering algorithm, which supports the variable fusion phase described earlier. The information provided by the DCFS enables the calculation of the coefficients that will weight the fusion of the original variables.
December 17, 2025 17:49 output colors. The final result consists of 37 clusters. As observed in Figure 12, the majority of the 98 attributes selected for the problem are located in close proximity within a spatial region surrounding the target variable. This suggests the formation of compact clusters, which is expected to facilitate effective reduction in dimensionality. Figure 14. Resulting variables in the DCFS after the fusion process. The target variable is marked with a star. On the other hand, Figure 13 illustrates the clusters obtained after applying the clustering phase proposed by the dimensionality reduction algorithm. The resulting number of clusters is 37, which implies that the merging process will receive the 98 original variables and produce an output of 37 variables, achieving a dimensionality reduction of 62.24%. Figure 14 further presents a visualization of the merging process based on the clustering results, where the positions of the variables resulting from the fusion process are marked with ‘X’ in the embedding space. Analyzing the groups that have been merged into new variables, we observe that attributes such as “BILLS” (difficulty or need for assistance with writing checks, paying bills, or balancing a checkbook), “TAXES” (difficulty or need for assistance with assembling tax records, managing business affairs, or handling other documents), “GAMES” (difficulty or need for assistance with playing a game of skill, such as bridge or chess, or engaging in a hobby), “STOVE” (difficulty or need for assistance with heating water, making a cup of coffee, or turning off the stove), “EVENTS” (difficulty or need for assistance with keeping track of current events), and “PAYATTN” (difficulty or need for assistance with paying attention to and understanding a TV program, book, or magazine) are grouped together into a single new variable. This fusion has a clear medical rationale, as these assessments are reasonably interconnected in patients who exhibit some form of Mild Cognitive Impairment (MCI). Using the Random Forest and GBT models described in the experimentation with the OASIS-2 dataset, and the already mentioned 80-20 train-test split (respectively), the following results were obtained (without intensive hyperparameter optimization, as this is not the focus of the present paper): an accuracy of 77.25% for the Random Forest model and 76.49% for the GBT model (without dimensionality reduction). Repeating the training and evaluation process for the reduced dataset of 37 attributes, the accuracy achieved was 76.43% for the Random Forest and 75.66% for the GBT. This represents a minimal decrease of 0.82% and 0.83%, respectively. The accuracy results are reported as the weighted average for each of the four classes in the target variable. These results are comparable to those obtained by the authors in,46 but with a significant reduction in dimensionality, which provides several benefits, including: reduced computational time for preprocessing, learning, and especially hyperparameter optimization, as well as a more interpretable framework of interrelated variables (before merging). This could assist in determining whether certain tests are dispensable, alongside mitigating the wellknown “curse of dimensionality”, which refers to the vast number of instances required to solve a learning problem with many attributes. The reduction in dimensionality reduces the number of attributes, which in turn decreases the number of instances necessary for learning. In the medical context, this results in fewer subjects and tests per subject required to train models, directly impacting the time and cost associated with diagnostic studies and processes. 7. Performance optimizations for large-scale datasets Due to the computational limitations associated with the calculation of Distance Correlation, as previously discussed in this article, an optimization is
December 17, 2025 17:49 output proposed in terms of both time and space performance. Specifically, a p reviously p roposed accelerated version of the Distance Correlation algorithm47 reduces the computational time complexity from O(n2) to O(n log n). This computational improvement is achieved through the use of both AVL trees and the mergesort algorithm. Due to the improvement in asymptotic complexity, the computation of Distance Correlation can be applied to problems involving a large number of instances, thereby enabling quasi-linear scalability (see Figure 15). Figure 15. Evolution of the computation time during the application of the proposed dimensionality reduction algorithm for different problem sizes, exhibiting a quasilinear trend in relation to the growth in the number of dataset instances. In Figure 15, a quasi-linear trend in the growth of computational time can be observed, as expected given the theoretical O(nlog n) cost associated with the Fast Distance Correlation algorithm. This growth demonstrates that the system remains scalable even when the problem involves a large number of instances (such as 10,000), requiring an approximate computation time of 26 seconds, which is considered more than reasonable for a dataset with 98 attributes in its input space. 8. Conclusions The dual proposal of the Distance-Correlation Feature Space (DCFS) and the dimensionality reduction algorithm has proven to be effective in representing the attributes of a supervised learning problem, where variable importance can be calculated, leading to enhanced explainability. In addition, dimensionality reduction offers computational optimizations by simplifying the feature space in which predictive models operate. In the context of dementia and its diagnosis based on Machine Learning models, this study reinforces the findings of Zubasti et al., confirming that simple clinical tests are far more decisive than other attributes in the detection of dementia in patients. The space and algorithms proposed in this paper address the dementia problem by not merely selecting attributes but fusing those that exhibit strong relationships with each other and the target variable. In cases where the dataset contains many more instances and potentially new attributes not considered in the previous experiments, the dimensionality reduction algorithm becomes even more valuable. Accelerates training and testing times for the models used to diagnose dementia, while the DCFS serves as a foundation for identifying key relationships. This approach also helps explainability and assigns importance to the various medical tests performed in this domain. 9. Future works Since the proposed algorithm approaches supervised problems from a general standpoint (without being specific to any particular domain), it would be interesting to test it on problems from other fields, such as industrial and engineering applications. Comparing the proposed dimensionality reduction algorithm with advanced techniques such as Principal Component Analysis (PCA), which applies linear combination operations to represent attributes in a new space, would be highly valuable to identify potential improvements and limitations. This comparison would enable a more comprehensive statistical analysis of the results obtained within the field of dimensionality reduction. Using various predictive models for both classification and regression problems offers a compelling opportunity to validate whether the performance of the model degrades similarly between different techniques. A study based on selecting the least important attributes—when analyzed from their representation in the latent space (embedding)—for elimination would constitute a natural extension of the dimensionality reduction algorithm. This would further refine both the number and types of attributes used to address the present problem. Employing advanced clustering techniques as a substitute for the k-means algorithm can produce superior results in specific cases where the cluster distribution does not necessarily exhibit a spherical shape.
December 17, 2025 17:49 output Instead, when clusters present a more skewed and complex structure, density-based algorithms, such as DBSCAN or membership-degree-based approaches such as the Expectation Maximization (EM) algorithm, may provide a more suitable and accurate representation of the underlying data distribution. Acknowledgments This study was funded by public research projects of the Spanish Ministry of Science and Innovation PID2023151605OB-C22 and the project under the call PEICTI 2021-2023 with the identifier TED2021-131520B-C22. The NACC database is funded by NIA/NIH Grant U24 AG072122. NACC data are contributed by the NIA-funded ADRCs: P30 AG062429 (PI James Brewer, MD, PhD), P30 AG066468 (PI Oscar Lopez, MD), P30 AG062421 (PI Bradley Hyman, MD, PhD), P30 AG066509 (PI Thomas Grabowski, MD), P30 AG066514 (PI Mary Sano, PhD), P30 AG066530 (PI Helena Chui, MD), P30 AG066507 (PI Marilyn Albert, PhD), P30 AG066444 (PI David Holtzman, MD), P30 AG066518 (PI Lisa Silbert, MD, MCR), P30 AG066512 (PI Thomas Wisniewski, MD), P30 AG066462 (PI Scott Small, MD), P30 AG072979 (PI David Wolk, MD), P30 AG072972 (PI Charles DeCarli, MD), P30 AG072976 (PI Andrew Saykin, PsyD), P30 AG072975 (PI Julie A. Schneider, MD, MS), P30 AG072978 (PI Ann McKee, MD), P30 AG072977 (PI Robert Vassar, PhD), P30 AG066519 (PI Frank LaFerla, PhD), P30 AG062677 (PI Ronald Petersen, MD, PhD), P30 AG079280 (PI Jessica Langbaum, PhD), P30 AG062422 (PI Gil Rabinovici, MD), P30 AG066511 (PI Allan Levey, MD, PhD), P30 AG072946 (PI Linda Van Eldik, PhD), P30 AG062715 (PI Sanjay Asthana, MD, FRCP), P30 AG072973 (PI Russell Swerdlow, MD), P30 AG066506 (PI Glenn Smith, PhD, ABPP), P30 AG066508 (PI Stephen Strittmatter, MD, PhD), P30 AG066515 (PI Victor Henderson, MD, MS), P30 AG072947 (PI Suzanne Craft, PhD), P30 AG072931 (PI Henry Paulson, MD, PhD), P30 AG066546 (PI Sudha Seshadri, MD), P30 AG086401 (PI Erik Roberson, MD, PhD), P30 AG086404 (PI Gary Rosenberg, MD), P20 AG068082 (PI Angela Jefferson, PhD), P30 AG072958 (PI Heather Whitson, MD), P30 AG072959 (PI James Leverenz, MD). References 1. I. Lee, Big data: Dimensions, evolution, impacts, and challenges, Business Horizons 60(3) (2017). 2. D. Gupta and R. Rani, A study of big data evolution and research challenges, Journal of Information Science 45(3) (2019). 3. W. Jia, M. Sun, J. Lian and S. Hou, Feature dimensionality reduction: a review, Complex and Intelligent Systems 8(3) (2022). 4. S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms 2013. 5. A. Rajkomar, J. Dean and I. Kohane, Machine Learning in Medicine, New England Journal of Medicine 380(14) (2019). 6. C. J. Haug and J. M. Drazen, Artificial Intelligence and Machine Learning in Clinical Medicine, 2023, New England Journal of Medicine 388(13) (2023). 7. Q. Meng, Application of machine learning in medicine, Applied and Computational Engineering 33(1) (2024). 8. J. A. Sidey-Gibbons and C. J. Sidey-Gibbons, Machine learning in medicine: a practical introduction, BMC Medical Research Methodology 19(1) (2019). 9. G. M. d. M. Paix˜ao, B. C. Santos, R. M. de Araujo, M. H. Ribeiro, J. L. de Moraes and A. L. Ribeiro, Machine Learning in Medicine: Review and Applicability (2022). 10. N. Shaffi, K. Subramanian, V. Vimbi, F. Hajamohideen, A. Abdesselam and M. Mahmud, Performance Evaluation of Deep, Shallow and Ensemble Machine Learning Methods for the Automated Classification of Alzheimer’s Disease, International Journal of Neural Systems 34(7) (2024). 11. F. Mercaldo, M. Di Giammarco, F. Ravelli, F. Martinelli, A. Santone and M. Cesarelli, Alzheimer’s Disease Evaluation Through Visual Explainability by Means of Convolutional Neural Networks, International Journal of Neural Systems 34(2) (2024). 12. Z. Sankari, H. Adeli and A. Adeli, Wavelet Coherence Model for Diagnosis of Alzheimer Disease, Clinical EEG and Neuroscience 43(4) (2012) 268– 278. 13. E. Perez-Valero, C. Morillas, M. A. Lopez-Gordo and J. Minguillon, Supporting the Detection of Early Alzheimer’s Disease with a Four-Channel EEG Analysis, International Journal of Neural Systems 33(4) (2023). 14. J. P. Amezquita-Sanchez, A. Adeli and H. Adeli, A new methodology for automated diagnosis of mild cognitive impairment (MCI) using magnetoencephalography (MEG), Behavioural Brain Research 305 (2016). 15. J. P. Amezquita-Sanchez, N. Mammone, F. C. Morabito, S. Marino and H. Adeli, A novel methodology for automated differential diagnosis of mild cognitive impairment and the Alzheimer’s disease using EEG signals, Journal of Neuroscience Methods 322 (2019). 16. J. P. Amezquita-Sanchez, N. Mammone, F. C. Morabito and H. Adeli, A New dispersion entropy and fuzzy logic system methodology for automated classification of dementia stages using electroencephalograms, Clinical Neurology and Neurosurgery 201 (2021). 17. G. Mirzaei and H. Adeli, Machine learning techniques for diagnosis of alzheimer disease, mild cognitive disorder, and other types of dementia, Biomedical Signal Processing and Control 72 (2022) p. 103293. 18. M. Ahmadlou and H. Adeli, Complexity of weighted graph: A new technique to investigate structural complexity of brain activities with applications to aging and autism, Neuroscience Letters 650 (2017). 19. R. Porto, J. M. Molina, A. Berlanga and M. A. Patricio, Minimum relevant features to obtain explainable systems for predicting cardiovascular disease using the statlog data set, Applied Sciences (Switzerland) 11(3) (2021). 20. J. Adams, Defending explicability as a principle for the ethics of artificial intelligence in medicine, Medicine, Health Care and Philosophy 26(4) (2023). 21. G. Schwalbe and B. Finzel, A comprehensive taxonomy for explainable artificial intelligence: a systematic survey of surveys on methods and concepts,
December 17, 2025 17:49 output Data Mining and Knowledge Discovery (2023). 22. R. Dwivedi, D. Dave, H. Naik, S. Singhal, R. Omer, P. Patel, B. Qian, Z. Wen, T. Shah, G. Morgan and R. Ranjan, Explainable AI (XAI): Core Ideas, Techniques, and Solutions, ACM Computing Surveys 55(9) (2023). 23. A. Vassiliades, N. Bassiliades and T. Patkos, Argumentation and explainable artificial intelligence: A survey (2021). 24. D. S. Marcus, T. H. Wang, J. Parker, J. G. Csernansky, J. C. Morris and R. L. Buckner, Open Access Series of Imaging Studies (OASIS): Cross-sectional MRI data in young, middle aged, nondemented, and demented older adults, Journal of Cognitive Neuroscience 19(9) (2007). 25. D. S. Marcus, A. F. Fotenos, J. G. Csernansky, J. C. Morris and R. L. Buckner, Open access series of imaging studies: Longitudinal MRI data in nondemented and demented older adults, Journal of Cognitive Neuroscience 22(12) (2010). 26. M. Lin, P. Gong, T. Yang, J. Ye, R. L. Albin and H. H. Dodge, Big Data Analytical Approaches to the NACC Dataset: Aiding Preclinical Trial Enrichment, Alzheimer Disease and Associated Disorders 32(1) (2018). 27. J. Barnes, B. C. Dickerson, C. Frost, L. C. Jiskoot, D. Wolk and W. M. Van Der Flier, Alzheimer’s disease first symptoms are age dependent: Evidence from the NACC dataset, Alzheimer’s and Dementia 11(11) (2015). 28. F. Yi, H. Yang, D. Chen, Y. Qin, H. Han, J. Cui, W. Bai, Y. Ma, R. Zhang and H. Yu, XGBoostSHAP-based interpretable diagnostic framework for alzheimer’s disease, BMC Medical Informatics and Decision Making 23(1) (2023). 29. V. L. Chetana, S. S. Kolisetty and K. Amogh, A Short Survey of Dimensionality Reduction Techniques, Recent Advances in Computer Based Systems, Processes and Applications, 2020. 30. M. Vijayan, S. S. Sridhar and D. Vijayalakshmi, A Deep Learning Regression Model for Photonic Crystal Fiber Sensor With XAI Feature Selection and Analysis, IEEE Transactions on Nanobioscience 22(3) (2023). 31. C. van Zyl, X. Ye and R. Naidoo, Harnessing eXplainable artificial intelligence for feature selection in time series energy forecasting: A comparative analysis of Grad-CAM and SHAP, Applied Energy 353 (2024). 32. Haidar Khalid Malik and Nashaat Jasim Al-Anber, Comparison of Feature Selection and Feature Extraction Role in Dimensionality Reduction of Big Data, Journal of Techniques 5(1) (2023). 33. I. De-La-bandera, D. Palacios, J. Mendoza and R. Barco, Feature extraction for dimensionality reduction in cellular networks performance analysis (2020). 34. F. L. Gewers, G. R. Ferreira, H. F. De Arruda, F. N. Silva, C. H. Comin, D. R. Amancio and L. D. F. Costa, Principal component analysis: A natural approach to data exploration, ACM Computing Surveys 54(4) (2021). 35. J. P. Bharadiya, A Tutorial on Principal Component Analysis for Dimensionality Reduction in Machine Learning, International Journal of Innovative Research in Science Engineering and Technology 8(5) (2023). 36. P. Zubasti, A. Berlanga, M. A. Patricio and J. M. Molina, Assessing the Interplay of Attributes in Dementia Prediction Through the Integration of Graph Embeddings and Unsupervised Learning, Artificial Intelligence for Neuroscience and Emotional Systems, eds. J. M. Ferr´andez Vicente, M. Val Calvo and H. Adeli (Springer Nature Switzerland, Cham, 2024), pp. 371–380. 37. S. Ji, S. Pan, E. Cambria, P. Marttinen and P. S. Yu, A Survey on Knowledge Graphs: Representation, Acquisition, and Applications, IEEE Transactions on Neural Networks and Learning Systems 33(2) (2022). 38. M. Xu, Understanding Graph Embedding Methods and Their Applications, SIAM Review 63(4) (2021). 39. H. Cai, V. W. Zheng and K. C. C. Chang, A Comprehensive Survey of Graph Embedding: Problems, Techniques, and Applications, IEEE Transactions on Knowledge and Data Engineering 30(9) (2018). 40. P. Goyal and E. Ferrara, Graph embedding techniques, applications, and performance: A survey, Knowledge-Based Systems 151 (2018). 41. S. T. Roweis and L. K. Saul, Nonlinear dimensionality reduction by locally linear embedding, Science 290(5500) (2000). 42. Hall KM, An r-Dimensional Quadratic Placement Algorithm, Management Science 17(3) (1970). 43. G. J. Sz´ekely, M. L. Rizzo and N. K. Bakirov, Measuring and testing dependence by correlation of distances, Annals of Statistics 35(6) (2007). 44. D. Edelmann, T. F. M´ori and G. J. Sz´ekely, On relationships between the Pearson and the distance correlation coefficients, Statistics and Probability Letters 169 (2021). 45. T. Cali˜nski and J. Harabasz, A Dendrite Method For Cluster Analysis, Communications in Statistics 3(1) (1974). 46. N. An, H. Ding, J. Yang, R. Au and T. F. A. Ang, Deep ensemble learning for Alzheimer’s disease classification, Journal of Biomedical Informatics 105 (2020) p. 103411. 47. A. Chaudhuri and W. Hu, A fast algorithm for computing distance correlation, Computational Statistics and Data Analysis 135 (2019).