X-INCEPD: Enhancing Inception-based Model for Parkinson's Disease Prediction with Explainable AI
Abstract
This is the accepted version of the paper that will be published in the IEEE IISA 2025 Conference Proceedings.The final version will be available via IEEE Xplore once published.The DOI of the final published paper will be added after publication.
Full text
X-INCEPD: Enhancing Inception-based Model for Parkinson’s Disease Prediction with Explainable AI Nikos Tsolakis Christoniki Maga-Nteve Stefanos Vrochidis Information Technologies Information Technologies Information Technologies Institute Institute Institute Centre for Research & Centre for Research & Centre for Research & Technology Hellas Technology Hellas Technology Hellas Thessaloniki, Greece Thessaloniki, Greece Thessaloniki, Greece School of Informatics Aristotle University of [email protected] [email protected] Thessaloniki Thessaloniki, Greece [email protected] Abstract— Accurate and interpretable prediction of Parkinson’s Disease (PD) progression is essential for effective clinical decision-making and personalized care. Building on IncePD, a deep learning framework based on InceptionTime architecture, which demonstrated high predictive accuracy using wearable sensor data and clinical assessment score, we provide an explainable framework. In this paper, we present a significant advancement of that model by integrating explainable artificial intelligence (XAI) techniques into the prediction pipeline. Our enhanced framework, X-IncePD, preserves the temporal modelling strength of the original architecture, while offering interpretable insights into model predictions. We implement an enhanced LIME explainer to identify key features in model predictions. Experimental evaluations show that the model not only maintains great performance, but also provides transparent justifications combined with an easy-to-use prediction interface. Keywords—XAI, Deep Learning, Parkinson’s Disease, Healthcare, Explainability, AI I. INTRODUCTION Parkinson’s Disease (PD) is a progressive neurodegenerative disorder characterized by motor and non-motor symptoms that significantly affect patients’ quality of life [1]. The severity of PD is typically evaluated using two well-established clinical instruments: the Movement Disorder Society – Unified Parkinson’s Disease Rating Scale (MDS-UPDRS) and the Parkinson’s Disease Questionnaire-9 (PDQ-8) [2,3]. The MDS-UPDRS, an updated version of the original UPDRS, was developed to address limitations in the earlier scale. It is divided into four sections which collectively assess a wide spectrum of PD symptoms, inscluding non-motor experiences, motor abilities and overall impact on daily functioning and emotional health. In parallel, the PDQ-8, a concise alternative for PDQ-39, captures the patient’s self-reported quality of life through eight targeted questions covering mood, physical limitations, mental state and daily activities. Higher cumulative scores on the PDQ-8 indicate more severe impairment. Together, these tools offer robust, validated metrics and are widely utilized in both clinical evaluations and PD-related research. Early detection and accurate monitoring of symptom progression are critical for timely intervention and personalized treatment planning. Nick Bassiliades Georgios Meditskos School of Informatics School of Informatics Aristotle University of Thessaloniki Aristotle University of Thessaloniki Thessaloniki, Greece Thessaloniki, Greece [email protected]th.gr [email protected]th.gr
Recent advancements in Machnie Learning (ML) and wearable technologies have anabled the development of automated systems for PD prediction, offering promising avenues for remote monitoring and clinical support [4]. In our prior work, we introduced ncePD, a Deep Learning (DL) model based on the InceptionTime architecture, designed to predict PD severity using time series data from wearable sensors combined with clinical assessment scores such as MDSUPDRS and PDQ-8 [5,6]. While IncePD achieved great and competitive accurace in capturing disease progression patterns, its interpretability remain limited – a challenge common to most deep learning-based medical applications. In clinical settings, black-box models pose significant barriers to adoption, as healthcare professionals require transparent reasoning to trust and act on algorithmic decisions. This has led to a growing demand for explainable artificial intelligence (XAI) techniques tha can uncover the rationale behind predictions, ensuring models are not only accurate but also interpretable and cilinically aligned. In this study, we introduce X-IncePD, an enhanced version of our previous model that embeds explainability into its architecture and prediction pipeline. By integrating state-of-theart XAI methods such as LIME, X-IncePD offers insight into individualized decision pathways. These capabilities not only enhance trust and transparency, but also facilitate clinical validation of model behavior. We evaluate our approach on a multi-source dataset including wearable data and clinical scores, demonstrating that X-IncePD maintains high predictive performance while providing meaningful explanations that align with expert medical understanding. The rest of this paper is organized as follows. In Section 2, a literature review on previous related works on PD prediction and explainability frameworks is presented. In Section 3, the experimental setup and evaluation methods are described. In Section 5, the results and the explainability framework are discussed. Finally, Section 6 presents the conclusions and future directions. II. RELATED WORK Recent research efforts have increasingly focused on integrating machine learning and XAI to enhance the diagnosis and understanding of PD. These methods not only improve predictive performance but also offer critical insights into the underlying biomarkers and disease mechanisms. Bhandari et al.conducted an integrative gene expression analysis utilizng machine learning and XAI to identify significant genomic features associated with early PD diagnosis [7]. Aversano et al. [8] proposed an XAI-driven approach for early PD diagnosis that centers on motor symptoms. Their system utilized motor behavior data to predict disease onset and monitor patient progression, with an emphasis on model transparency to support clinical decision-making. An XGBoost with SHAP (Shapley Additive exPlanations) to enhance PD detection accuracy was introduced by Roy et al. [9]. Their work highlighted how feature attribution techniques could incover the most influential variables enabling both high accuracy and interpretability. In a broader context, Priyadarshini et al. developed a comprehensive framework for PD diagnosis that employs multiple machine learning techniques powered by XAI to analyze MRI imaging data [10]. Savaranan et al. developed a hybrid deep learning model combining VGG19 and GoogleNet for classifying spiral and wave drawings, achieving high accuracy in PD prediction [11]. Their integration of XAI techniques enabled local interpretability, offering valuable insights into how specific drawing features influenced model predictions. An AI-enabled system that leverages nocturnal breathing signals for detecting PD and tracking its progression proposed by Yang et al [12]. The non-invasive nature of the data and the model’s longitudinal prediction capacity mark a significant step toward passive monitoring solutions. Reddy et al. reviewed early diagnostic approaches using AI and ML, highlighting their capability to extract subtle biomarkers missed by conventional clinical methods [13]. Similarly, Dixit et al. conducted a comprehensive survey of AI models for PD diagnosis, emphasizing the superiority of these methods in handling complex, non-linear clinical data [14]. Finally, Velázquez et al. introduced an approach using X-vectors extracted from speech for automatic PD detection, showing high predictive accuracy with minimal feature engineering. This work supports the potential of voice as a viable biomarker [15].
III. METHODOLOGY Recent advancements in ML and DL have demonstrated significant promise in the prediction and diagnosis of various medical conditions, including PD. In this study, we leverage InceptionTime to develop models capable of assessing PD severity. A. Data Acquisition and Preprocessing This study employs data sourced from the mPower Public Research Portal [16], a largescale clinical observational initiative focused on PD. The Mpower study gathered a combination of sensor-derived data and self-reported survey responses from a broad participant base, all through a mobile application platform. The study was structured around seven distinct tasks, including physical activities (e.g., walking, tapping, voice, and memory) as well as surveys (demographics, MDS-UPDRS, and PDQ-8). Guided by medical expert recommendations, we selected four key components for our analysis: the walking task, the demographic survey, and the MDS-UPDRS and PDQ-8 questionnaires. The walking task comprises three segments—outbound walk, stationary rest, and return walk—during which the smartphone’s built-in accelerometer and gyroscope record three-dimensional motion and angular velocity. These motion signals are leveraged to detect PD-related motor impairments and to differentiate individuals with PD from healthy controls, while also estimating disease severity through predictive modeling of questionnaire scores. Extensive research has highlighted the importance of data preprocessing to improve the predictive accuracy of models estimating MDS-UPDRS and PDQ-8 scores [17, 18]. During this stage, issues such as missing entries, noise, and data inconsistencies are systematically addressed. The process began with identifying and isolating the most relevant subsets of the dataset aligned with the objectives of this study. Particular attention was given to the treatment of missing values, which often present a major challenge for Deep Learning systems. To enhance efficiency and model interpretability, feature selection was performed, helping reduce dimensionality and accelerate the training process. For this study, we focused on participants who had completed both relevant surveys and the walking task, as identified through guidance from the clinical collaborators involved. A critical aspect of preprocessing involved transforming raw sensor data into a structured format compatible with neural network inputs. In addition, harmonizing the dataset required establishing common indexing keys to handle overlapping and redundant entries across different data streams. Subsequently, time-series data was segmented into shorter windows using a sliding window technique—each segment spanning 5 seconds (500 samples at a 100 Hz sampling rate) with a 50% overlap between consecutive windows. Finally, the dataset was divided into training and testing subsets, with 80% allocated for training the models and the remaining 20% reserved for evaluation. B. Model Architecture The proposed Ince-PD in [5] model introduces several key architectural modifications to the original InceptionTime framework to enhance its effectiveness in predicting Parkinson’s Disease severity. The first major adjustment involves the removal of the Bottleneck layer within the inception modules. Empirical findings indicated that excluding this layer led to improved model efficiency without sacrificing accuracy. In contrast, residual connections—which link every third inception block—were retained from the original architecture, as they were shown to support better gradient flow and overall model optimization. Beyond changes to the inception modules themselves, the broader network architecture was refined. Additional layers such as Batch Normalization and Dropout (rate = 0.5) were incorporated to improve training stability and prevent overfitting, respectively. Furthermore, Rectified Linear Units (ReLU) were adopted as the activation function following the batch normalization layers to introduce nonlinearity and accelerate learning. Notably, instead of ending the network with a Softmax layer, the final output layer also utilizes a ReLU activation, which facilitated faster convergence and yielded improved performance in regression tasks.
To further enhance generalization, several hyperparameters were carefully tuned, including the convolutional kernel sizes, network depth, and number of filters. The sequence of operations—modified inception modules followed by normalization, activation, and pooling—was capped with a 1D Global Average Pooling layer, reducing feature dimensionality before the output stage. Ultimately, the trained models were used to predict aggregate scores for MDS-UPDRS Parts I & II and the PDQ-8 questionnaire, based on the preprocessed sensor data. A visual representation of the proposed architecture is presented in Figure 1. Fig. 1. IncePD architecture as presented in [5]. C. Model Performance Evaluation To assess the predictive performance of the proposed Inception-based model, we employed two standard evaluation metrics: Mean Absolute Error (MAE) and Mean Squared Error (MSE). These metrics were computed on both a per-window and per-patient basis to provide granular and overall insights into model accuracy. The mathematical definitions of these metrics are shown below, where (y_i ) , y_i are the predicted value and the actual value. MAE = 1 𝑛∑|𝑦𝑖− 𝑦𝑖| 𝑛 𝑖=1 (1) MSE = ∑(𝑦 𝑖−𝑦𝑖)2 𝑛 𝑛 𝑖=1 (2) To further validate our approach, we compared the performance of the proposed model against a range of existing methods reported in the literature for predicting MDS-UPDRS I & II and PDQ-8 scores. All models were evaluated using the same dataset and target labels to ensure consistency across comparisons. Details of the baseline architectures used for benchmarking are presented in the subsequent sections. D. Experimental Setup The Inception-based model introduced in this study was developed using Python 3.8 and executed on a workstation equipped with an Intel(R) Xeon(R) Silver 4210 CPU @ 2.20 GHz. The neural network was implemented with the TensorFlow framework [19], and training was conducted over 10 epochs, allowing the model sufficient opportunity to learn meaningful patterns from the data and improve performance. Model training was carried out using the fit function, applied to both input features and their corresponding labels. For optimization, we employed the Adam optimizer, a widely-used adaptive learning rate method known for its balance of efficiency and convergence speed. Hyperparameter tuning was performed through heuristic experimentation, supported by TensorFlow’s HParams library, allowing for dynamic adjustment of the model’s configuration. To evaluate and fine-tune the model's performance, we used MAE and MSE as primary evaluation metrics. These guided the optimization process by helping to identify the most effective model setup. The objective was to achieve the lowest possible error rates while maintaining a balance between computational efficiency and model complexity. Our proposed configuration was further benchmarked against optimized models trained on the same dataset to ensure robust comparison. IV. RESULTS
The performance of the initial Ince-PD model was first assessed by comparing it against several baseline architectures introduced in previous work. Thanks to the combination of Inception modules, ReLU activations, residual connections, and dropout regularization, Ince-PD offers improved learning efficiency and faster convergence while maintaining lower computational overhead. Despite careful tuning of the competing models, Ince-PD consistently outperformed them across both evaluation metrics—Mean Absolute Error (MAE) per window and per patient—for both the total MDS-UPDRS I & II and PDQ-8 scores. The model achieved a MAE of 1.97 (window) and 2.27 (patient) for MDS-UPDRS I & II, and 2.17 (window) and 2.96 (patient) for PDQ-8. These results reflect a significant improvement compared to standard CNN, LSTM, and hybrid architectures. IV. X-INCEPD EXPLAINABILITY FRAMEWORK The transparency and clinical interpretability in deep learning–based Parkinson’s Disease assessment is an important aspect in modern day clinical settings. We introduce X-IncePD— an explainability framework built around the IncePD model. X-IncePD integrates complementary techniques to provide meaningful insights into how predictions are made, supporting both scientific validation and potential clinical deployment. In this chapter, we present the key components of X-IncePD, including: • LIME-based local explanations with time-axis feature mapping, • Saliency maps highlighting input sensitivity across time and axes, • Combined visualizations merging attribution and gradient insights, • An interactive user interface for real-time exploration of predictions. A. Core Methodology And Results To enhance interpretability into our deep learning framework, we applied the Local Interpretable Model-Agnostic Explanations (LIME) methodology to explain the predictions made by the IncePD model [20]. LIME is a post-hoc explainability technique that approximates the local behavior of complex models by fitting an interpretable surrogate model—typically a linear regressor—around a specific prediction instance. Given that our model operates on multivariate time-series data with a shape of (500, 3), corresponding to 5-second windows of tri-axial accelerometer signals (at 100 Hz), the raw input was first reshaped into a 2D format compatible with LIME. Each instance was flattened into a single vector of 1500 features, enabling the use of the LimeTabularExplainer designed for tabular data. Α custom prediction function to reshape the flat input vectors back into the original three-dimensional form expected by the model was defined, allowing LIME to probe the model’s internal decision function without altering its architecture. The explainer was initialized using a representative subset of the flattened test dataset and configured in regression mode, appropriate for the continuous outputs of MDS-UPDRS I & II and PDQ-8 score predictions. For each prediction instance, LIME generated a set of perturbed samples around the data point of interest and recorded the model’s response to these variations. Using these observations, it trained a local linear model to estimate the importance of each feature. This provided feature attribution scores that indicated the relative contribution—positive or negative—of each timeseries component to the model’s final output. Notably, we visualized these results to identify which sensor channels and temporal segments influenced the prediction, providing clinically meaningful interpretations aligned with known symptom patterns in Parkinson’s Disease. Figure 2 presents the results of the explainer, providing meaningful explanations for the IncePD model.
Fig. 2. Feature importance in the IncePD model. Table 1 presents the feature importance summary for a single prediction instance, as generated by the LIME explainer. Each row corresponds to a feature (from the flattened time-series input), with its associated value used in the prediction. Features highlighted in blue indicate negative contributions, meaning they decreased the predicted severity score, while features in orange represent positive contributions, suggesting an upward influence on the output. For example, feature1165 and feature408 exhibit strong negative influence with values of −0.05 and −1.01 respectively, while feature583, with a relatively small positive value of 0.12, slightly increased the final prediction. TABLE I. FEATURES AND VALUES UTILIZING X-INCEPD Feature Value Feature1165 -0.05 Feature1203 -1.01 Feature492 -0.95 Feature408 -1.01 Feature583 0.12 B. Tracing Back Features in X-IncePD A core advantage of X-IncePD is its ability to bridge deep model reasoning with humaninterpretable movement patterns by tracing explanations back to the raw sensor signals captured during the walking tasks of the mPower study. Each model input is a 500×3 matrix, representing 500 time steps of tri-axial accelerometer data (x, y, z) collected from participants’ smartphones. For compatibility with local surrogate explainers, this matrix is flattened into a 1500dimensional vector. However, we maintain a precise and interpretable mapping from each flattened feature index back to its original sensor context. This mapping allows us to interpret any identified influential feature (e.g., feature1348) in terms of a specific moment in time and a specific axis of motion (e.g., time step 449, y-axis), thus reconnecting abstract model internals to real, physical movement. C. Saliency-based Enhancing of Explainability To further interpret the model’s internal behavior, we employed saliency map analysis, a gradient-based method that quantifies the importance of each input value with respect to the model’s prediction. By computing the gradient of the predicted output with respect to the input
features, we generated a 2D saliency heatmap that reveals the most influential time steps and sensor channels in a given input window. Figure 3 presents the saliency heatmap for a single sample (Sample 0) from the test set, covering a 5-second time window of tri-axial accelerometer signals (500 time steps at 100 Hz). The yaxis corresponds to the three sensor axes: 0 = X, 1 = Y, and 2 = Z, while the x-axis represents the temporal dimension. Each pixel's intensity reflects the absolute gradient magnitude (i.e., |∂y/∂x|), where brighter (yellow/white) values indicate higher sensitivity—i.e., input regions that had a greater impact on the model's prediction. The heatmap clearly shows discrete bursts of high saliency, particularly concentrated in the X and Y axes within the first 200 time steps. This suggests that early parts of the signal—likely representing the initial movement or gait transition in the walking task—play a pivotal role in the model’s decision. In contrast, the Zaxis generally contributes less across most of the sequence, which may indicate lower relevance of vertical motion in distinguishing Parkinsonian patterns. Fig. 3. Saliency heat map for a single sample. D. Feature Attribution Visualization To contextualize the most influential features in the prediction process, we visualized the raw accelerometer signal corresponding to a 5-second window (500 time steps) from Sample 0, highlighting the impactful features. As shown in Figure 4, the plot displays all three sensor axes: X (blue), Y (green), and Z (orange). These signals represent the participant’s linear motion as captured by the tri-axial accelerometer during the walking task. Key points of influence, as identified by X-IncePD, were overlaid as distinct scatter markers to indicate the top 10 features contributing most significantly to the model's output for this instance. Notably, these points align with distinct fluctuations or transitions in the sensor readings—often corresponding to gait events such as foot strikes, directional changes, or changes in movement intensity. Fig. 4. Raw tri-axial accelerometer signals for Sample 0, showing X (blue), Y (green), and Z (orange) axes across 500 time steps.. E. Unified Interpretability Visualization F. To offer a holistic view of model interpretability, we developed a combined visualization that overlays both LIME feature attributions and saliency-based relevance scores onto the original tri-axial accelerometer signals. This approach emphasizes the degree to which those values
affected the prediction. The visualization enables side-by-side validation of insights derived from local perturbation-based (LIME) and gradient-based (saliency) methods. As shown in Figure 5, the original sensor signals (X: red, Y: green, Z: blue) for a representative sample are plotted over 500 time steps (corresponding to a 5-second walking task). The background shading represents saliency magnitude: more saturated vertical bands indicate time-axis regions where the model is most sensitive to changes in input, calculated as the absolute gradient of the output with respect to each time–sensor value. Overlaid on this are the top 10 LIME features, depicted as scatter points. These points represent the most impactful input features identified through LIME’s locally linear surrogate model. Each feature is mapped back from the flattened input to its original time step and axis, allowing for direct visualization within the temporal signal context. Fig. 5. Combined interpretability view for Sample 0. Raw tri-axial accelerometer signals are shown in red (X), green (Y), and blue (Z). Saliency information is represented by background color bands, where darker hues indicate higher model sensitivity to input variation. G. Interactive Local Interpretability Tool To enhance model transparency and facilitate human-centered evaluation, a visualization tool was developed. Built upon the LIME framework, this tool allows users to explore how individual input features influence the predicted MDS-UPDRS scores generated by the IncePD model. As illustrated in Figure 6, users can select a specific test sample via an intuitive slider and receive an immediate prediction along with a corresponding local explanation. For the selected sample (e.g., index 17), the model predicted a UPDRS score of 12.23. The adjacent bar chart displays the most influential features affecting this prediction, separated into positive contributors (green bars) and negative contributors (red bars). These contributions are expressed as feature conditions (e.g., feature1383 > 0.92), along with their impact magnitude. The results of such explanations can be the input in developing evaluation frameworks as presented in [21].
Fig. 6. The X-IncePD Explainer interface showcasing a sample prediction and its corresponding explanation. The tool allows users to select a sample, view the predicted MDS-UPDRS score and interpret which input features (shown as conditional rules) most influenced the prediction outcome. Figure 7 illustrates the feature importance after identifying the features based on the methodology presented in section B of this chapter. In this example (Sample 42), we further enhanced interpretability by tracing each LIME-identified feature back to its corresponding time step and sensor axis, allowing us to present the explanation in a more meaningful format. Instead of abstract feature indices, the explanation is now expressed in terms of real signal components (e.g., "Time 426, Axis X"), helping to directly relate influential segments to physical movement patterns captured during the walking task. This mapping supports more intuitive understanding and facilitates clinical interpretation. Fig. 7. Interpretable Explanation with Time-Axis Mapping for Sample 42. V. CONCLUSION In this work, the IncePD model, a deep learning model based on the InceptionTime architecture, designed for the prediction of Parkinson’s Disease severity using tri-axial accelerometer signals was enhanced with an explainability framework. The initial model was trained to estimate clinically validated outcomes including total MDS-UPDRS Parts I & II and PDQ-8 scores and demonstrated great results. To address the critical need for transparency in medical AI systems, we developed X-IncePD, an explainability framework that surrounds the IncePD model with robust interpretability tools. X-IncePD leverages both LIME and gradient-based saliency maps to provide feature-level insights into the model’s predictions. These techniques expose which parts of the input signal (in time and sensor axis) were most influential, enabling a layer of posthoc interpretability that bridges the gap between black-box AI and clinical reasoning. In addition, X-IncePD includes a user-friendly, interactive visualization interface, allowing clinicians and researchers to explore individual predictions and their corresponding explanations. Built with Gradio, this tool presents local LIME explanations and predicted scores in real time, making model outputs accessible and actionable for healthcare professionals— even those without technical AI expertise. This ease of use makes X-IncePD a valuable asset in facilitating trust and adoption of AI tools in clinical practice. While the predictive performance of IncePD is well validated, the evaluation of the X-IncePD framework itself—in terms of clinical utility, explanation quality, and trustworthiness— remains a key direction for future work. Formal studies involving clinical experts, qualitative assessments of interpretability, and the incorporation of additional explanation metrics will be essential to assess and refine its effectiveness in real-world deployments. In conclusion, IncePD provides an accurate and efficient model for Parkinson’s Disease assessment, and X-IncePD enhances its transparency through actionable, human-centered explanations. Together, they represent a promising step toward interpretable and deployable AI tools in the neurological domain.