scieee AI-readable full text Open interactive document viewer

Machine learning for differential diagnosis of white matter lesions in Fabry Disease patients based on gait and cardiac characteristics

Braga, José António Fernandes

Abstract

Brain manifestations in FD include progressive white matter lesions (WMLs).This research aims to identify a set of gait and cardiac characteristics to discriminate FD patients with WMLs from FD patients without WMLs. For the gait study, 76 subjects walked through a predefined circuit using wearable sensors that continuously acquired different stride features. All strides of 16 gait variables were normalized using multiple regression models. The mean and the variability of each gait time series were calculated, resulting in 32 gait measures. Using the 32 gait measures, a feature selection algorithm were applied. Then, five different classifiers (LR, SVM Linear and RBF kernel, RF, and KNN) based on different selected set features were evaluated. CNN and LSTM algorithms were implemented using as input the gait time series. For FD patients with WMLs vs controls the highest accuracy of 71.50% was obtained using RF. For FD patients without WMLs vs controls, the best performance was observed using KNN with an accuracy of 86.67%. For FD patients with vs without WMLs the best models were obtained using the CNN algorithm and using LR algorithm with an accuracy of 81.43% and 80.76%, respectively. Regarding the cardiac data, was used the data from two exams: an electrocardiogram (ECG) and an echocardiogram. A total of 114 FD patients were evaluated with the ECG (61 have WMLs). For FD patients with vs without WMLs the highest accuracy of 79.72%. With the simultaneously use of both gait and ECG features, two models were evaluated on a test group of nine patients. The best result was 80% accuracy with LR. Finally, logistic regression analyses were also performed on the 22 features of the echocardiogram of ninety three FD patients (forty-nine with WMLs). The results confirmed that age is significantly associated with the presence of WMLs. These findings are the first step to demonstrate the potential of machine learning techniques based on gait and cardiac characteristics as a complementary tool to understand the role of WMLs in FD patients.

Full text

José António Fernandes Braga Machine learning for differential diagnosis of white matter lesions in Fabry Disease patients based on gait and cardiac characteristics Master’s Thesis Integrated Master’s in Industrial Electronics and Computers Engineering Work elaborated under supervision of: Prof. Dr. Estela Bicho (Supervisor) Dr. Flora Ferreira (Supervisor) June 2021 DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/ ii Acknowledgments I would like to thank the following people, without whom I would not have been able to complete this research, and without whom I would not have made it through my master’s degree! The Algoritmi Center at Universidade do Minho, especially to my supervisors Estela Bicho Erlhagen and Flora Ferreira, whose insight and knowledge into the subject matter steered me through this research. And special thanks to Carlos Fernandes, whose support as part of his research allowed my studies to go the extra mile. I would also like to thank Doctor Miguel Gago and Doctor Olga Azevedo for all the help regarding the medical issues that came through the development of this dissertation. And my biggest thanks to my parents and friends for all the support you have shown me through this research, the culmination of five years of learning, especially to my friend Coelho who has helped me across the years. And, last but not least, for my girlfriend Claudia, thanks for all your support, without which I would have stopped these studies a long time ago. I can only say thank you to everyone who has supported me and had to put up with my stresses and moans for the past five years of study and sorry for being even grumpier than normal whilst I wrote this thesis! iii STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. iv Resumo Machine learning para o diagnóstico de lesões na massa branca em pacientes com Doença de Fabry baseado na marcha e em características cardíacas As manifestações cerebrais da doença de Fabry (FD) incluem lesões na matéria branca (WMLs). O objetivo deste estudo é identificar quais as caracteristicas da marcha e cardíacas que permitem diferenciar os pacientes com FD com WMLS de pacientes com FD sem WMLs. Para o estudo da marcha, foram avaliados 76 sujeitos através de sensores vestíveis. Os valores da série temporal das 16 variáveis da marcha foram normalizados usando modelos de regressão múltipla. Usando as 32 medidas de marcha (média e variabilidade), foi aplicado um algoritmo de feature selection seguido de cinco classificadores diferentes (LR, SVM Linear e kernel RBF, RF e KNN). Os algoritmos CNN e LSTM foram implementados utilizando como input os conjuntos de séries temporais da marcha. Para pacientes com FD e com WMLs vs controlos, a maior exatidão de 71,50% foi obtida usando RF. Para pacientes com FD e sem WMLs vs controlos, o melhor desempenho foi observado usando KNN com uma exatidão de 86,67%. Para pacientes com FD com vs sem WMLs, os melhores modelos foram obtidos usando o algoritmo CNN e usando LR com base com uma exatidão de 81,43% e 80,76%, respetivamente. Em relação aos dados cardíacos, foram utilizados os dados de dois exames: o eletrocardiograma (ECG) e o ecocardiograma. Um total de 114 pacientes com FD (61 deste com WMLs) foram avaliados com o exame de ECG. Para pacientes com FD com vs sem WMLs, a maior exatidão foi de 79,72%. Com o uso simultâneo de marcha e ECG, dois modelos foram avaliados com um grupo de teste de nove pacientes. O melhor resultado foi a exatidão de 80% com o algoritmo LR. Finalmente, análises de regressão logística também foram realizadas nas 22 características do ecocardiograma de 93 pacientes com FD (49 com WMLs). Os resultados confirmaram que a idade está significativamente associada à presença de WMLs. Esses resultados demonstram o potencial das técnicas de machine learning baseadas na marcha e nas características cardíacas para entender o papel dos WMLs em pacientes com FD. Palavras-chave: Doença de Fabry, Machine Learning, Lesões da matéria Branca, Marcha, Exames cardíacos. v Abstract Machine learning for differential diagnosis of white matter lesions in Fabry Disease patients based on gait and cardiac characteristics Brain manifestations in FD include progressive white matter lesions (WMLs).This research aims to identify a set of gait and cardiac characteristics to discriminate FD patients with WMLs from FD patients without WMLs. For the gait study, 76 subjects walked through a predefined circuit using wearable sensors that continuously acquired different stride features. All strides of 16 gait variables were normalized using multiple regression models. The mean and the variability of each gait time series were calculated, resulting in 32 gait measures. Using the 32 gait measures, a feature selection algorithm were applied. Then, five different classifiers (LR, SVM Linear and RBF kernel, RF, and KNN) based on different selected set features were evaluated. CNN and LSTM algorithms were implemented using as input the gait time series. For FD patients with WMLs vs controls the highest accuracy of 71.50% was obtained using RF. For FD patients without WMLs vs controls, the best performance was observed using KNN with an accuracy of 86.67%. For FD patients with vs without WMLs the best models were obtained using the CNN algorithm and using LR algorithm with an accuracy of 81.43% and 80.76%, respectively. Regarding the cardiac data, was used the data from two exams: an electrocardiogram (ECG) and an echocardiogram. A total of 114 FD patients were evaluated with the ECG (61 have WMLs). For FD patients with vs without WMLs the highest accuracy of 79.72%. With the simultaneously use of both gait and ECG features, two models were evaluated on a test group of nine patients. The best result was 80% accuracy with LR. Finally, logistic regression analyses were also performed on the 22 features of the echocardiogram of ninety three FD patients (forty-nine with WMLs). The results confirmed that age is significantly associated with the presence of WMLs. These findings are the first step to demonstrate the potential of machine learning techniques based on gait and cardiac characteristics as a complementary tool to understand the role of WMLs in FD patients. Keywords: Fabry Disease, Machine Learning, White Matter Lesions, Gait, Cardiac exams. vi Contents Acknowledgments...................................... iii Resumo .......................................... v Abstract .......................................... vi List of Figures x List of Tables xiii Listofacronyms ...................................... xvi I : Introduction and State of the art 20 1 Introduction 21 1.1 Contextualization and Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 1.2 Objectives ...................................... 22 1.3 OverviewoftheThesis ................................ 23 1.4 Contributions of this thesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 2 State of Art 26 2.1 FabryDisease .................................... 26 2.2 Machine Learning for the diagnosis in healthcare . . . . . . . . . . . . . . . . . . . 29 vii CONTENTS II : Materials and methodology 31 3 Materials 32 3.1 Gaitassessment ................................... 32 3.2 Electrocardiogram .................................. 34 3.3 Echocardiogram ................................... 35 4 Methodology 38 4.1 Multiple Regression Normalization . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 4.2 Statisticaltests .................................... 39 4.3 MachineLearning................................... 40 4.4 Featureselection................................... 42 4.5 Modelstheoreticalbasis................................ 43 4.6 MachineLearningmetrics............................... 56 III : Results, Discussion, Conclusion, and Future Work 57 5 Gait assessment 58 5.1 Participants...................................... 58 5.2 GaitNormalization .................................. 59 5.3 FD patients with WMLs vs controls . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 5.4 FD patients without WMLs vs controls . . . . . . . . . . . . . . . . . . . . . . . . . 68 5.5 FD patients with vs without WMLs . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 5.6 Discussion ...................................... 82 6 Electrocardiogram 85 6.1 Participants and Holter data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 6.2 Feature selection with RFE . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 6.3 Electrocardiogram classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 6.4 Gait +Electrocardiogram analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 viii CONTENTS 6.5 Discussion ...................................... 96 7 Echocardiogram 98 7.1 Participants...................................... 98 7.2 Discussion ......................................103 8 Conclusion, limitations, and future work 104 8.1 Generalconclusions .................................104 8.2 Limitations......................................105 8.3 Futurework......................................106 IV : Appendix 107 ix LIST OF TABLES List of acronyms •α-gal A -α-galactosidase A •AIC - Akaike’s Information Criterion •CI - Confidence Interval •CNN - Convolutionary Neural Networks •ERT - Enzyme replacement therapy •ECG - Electrocardiogram •FD - Fabry Disease •Gb3 - Globotiaosylceramide •GNB - Gaussian Naive Bayes •KNN - K-Nearest Neighbor •LR - Logistic Regression •LSTM - Long-Short Term Memory •ML - Machine Learning •MR - Multiple Regression •NRR - Artificial Neural Network activated with ReLU function •RF - Random Forest •RFE - Recursive Feature Elimination •SD - Standard deviation xvi LIST OF TABLES •SVM - Support Vector Machine •VIF - Variance Inflation Factor •WMLs - White Matter Lesions xvii Part I : Introduction and State of the art 20 Chapter 1 Introduction 1.1 Contextualization and Motivation Fabry Disease (FD) is a rare disease that greatly affects the quality of life and may lead to premature death. This disease is an X-linked lysosomal storage disorder caused by the deficiency or absent activity of the enzyme α-Galactosidase A (α-Gal A). This disease also affects several organs, including the kidney, heart, and brain. Brain manifestations in FD include progressive white matter lesions (WMLs) (Buechner et al., 2008; Körver et al., 2018a). White matter is the brain region responsible for the transmission of nerve signals and for communication between different parts of the brain. Brain WMLs were an early manifestation, affecting 11.1% of males and 26.9% of females under 30 years of age, even without cerebrovascular risk factor outside FD (Azevedo et al., 2020). In (Körver et al., 2018a), WMLs were found in 46% of 1276 patients which tend to occur earlier in males and their prevalence revealed to increase with patients’ age. WMLs have been associated with gait impairment (Starr, 2003) and the risk of falls (Snir et al., 2019; Zheng et al., 2012). Gait abnormalities, such as slower gait and postural instability, have been reported in FD (Löhle et al., 2015). Furthermore, gait assessment has been recently revealed as a good complementary clinical tool to discriminate FD patients from healthy adults, however, not much is known regarding the impact of WMLs on gait performance in patients with FD. Specific location and distribution of WMLs suggest a specific underlying disease (Körver et al., 2018a) and may reflate in a 21 CHAPTER 1. INTRODUCTION different gait profile. Furthermore, heart abnormalities have been associated with the presence of WMLs (Forte et al., 2019). However, it remains unclear whether cardiac biomarkers are associated with WMLs in FD patients. It is then of extreme importance to explore different biomarkers that can assist in the differential diagnosis of WMLs in FD patients. Machine Learning (ML) is coming into its own, with a growing recognition that it can play a key role in a wide range of critical applications. ML provides potential solutions in several domains and is set to be a pillar of our future civilization. ML has been used in different industries, one of these being the healthcare industry. Machine learning is a growing field that has been extremely used because of the excellent results it can achieve when applied to different classification and pattern recognition tasks in the medical sector (see e.g. (Maity & Das, 2017)). The potential of ML methods based on gait, electrocardiogram, and echocardiogram has not yet been explored for the task of assisting in the differential diagnosis of WMLs in FD patients. This may be due to the fact that FD is a rare disease and data to develop different ML methods is not easily available. It happens that in the district of Guimarães, Portugal, there is an unusually high number of FD patients. This leads to the possibility of collecting a substantially high quantity of data about this disease and puts us in a unique position to study different patterns and biomarkers for the differential diagnosis of WMLs in FD patients using ML. 1.2 Objectives This dissertation aims to investigate different machine learning algorithms based on different datasets (gait, electrocardiogram, and echocardiogram) to develop a clinical decision support tool that enhances the differential diagnosis of WMLs in FD patients. More precisely, the objectives of this thesis are presented below: • To evaluate the effectiveness of machine learning methods to discriminate: –FD patients with WMLs from FD patients without WMLs, and these two groups from healthy controls based on gait data; 22 CHAPTER 1. INTRODUCTION –FD patients with WMLs from FD patients without WMLs based on electrocardiogram data; –FD patients with WMLs from FD patients without WMLs based on echocardiogram. • Investigate how the integration of more than one dataset (gait, electrocardiogram, or echocardiogram) can improve the predictive system for the differential diagnosis of WMLs in FD patients. 1.3 Overview of the Thesis Figure 1 presents an overview of the steps taken to achieve the stated objectives. Four different datasets were used: clinical and physical data, gait data, electrocardiogram data, and echocardiogram data. The first step consisted of analyzing each dataset to understand the influence of each feature: their significance (statistical tests) and their correlations. The effects of inter-subject variations in each dataset due to physical characteristics were analyzed. To reduce these effects normalization of gait data was performed, and electrocardiogram and echocardiogram datasets were subdivided. Then, different feature selection techniques were used to select the best subset of features of each dataset. Finally, with the subsets created by the feature selection from one or more datasets, different classification machine learning algorithms were evaluated. The results of the performance evaluation of the models provide prior knowledge to develop a supervised predictive system to work as a support for the early diagnostic of the presence or absence of WMLs in FD patients. This dissertation is composed of four parts and a total of eight chapters. The first part contains two chapters, introduction, and state of the art. Introduction talks about the general motivation and contextualization, the objectives, an overview of this thesis, and the scientific papers developed in the scope of this thesis. The last chapter of Part I is the state of art and it introduces the theoretical fundamentals of the concepts addressed in the dissertation and the work developed in the areas related to the dissertation. Part II describes the materials and methodology. The first chapter of Part II relates how the data gets collected and the inside of each dataset to be used further in the dissertation. Second chapter reports the methodology used across the development of this dissertation. Following is Part III with results, discussion, conclusion, and future work. The initial chapter of Part III contains the results of the data analysis on the 23 CHAPTER 1. INTRODUCTION Figure 1: Overview of the research steps of this thesis. information provided from the information in the third chapter plus a discussion for each of the different types of dataset. The next chapter concludes this thesis with a scope for future work. Finally, Part IV is the appendix, where is some extra information regarding the results. 1.4 Contributions of this thesis The following publications have been made based on the work developed in this dissertation: •José Braga, Flora Ferreira, Carlos Fernandes, Miguel F. Gago, Olga Azevedo, Nuno Sousa, Wolfram Erlhagen, Estela Bicho . “Gait characteristics and their discriminative ability in patients with fabry disease with and without white-matter lesions”. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in 24 CHAPTER 1. INTRODUCTION Bioinformatics). Vol. 12251 LNCS. Springer Science and Business Media Deutschland GmbH, July 2020, pp 415-428. ISBN: 978303058876. DOI: 10.1007/978-3-030-58808-330. • Flora Ferreira, Carlos Fernandes, José Braga, Estela Bicho, Wolfram Erlhagen, Miguel F. Gago.“The effect of levodopa medication on stride time variability in patients with Parkinsonism”. In: Journal of Statistics on Health Decision 2.2 (Oct. 2020), pp 63-64. DOI: 10.34624/jshd.v2i2.20109 25 Chapter 2 State of Art 2.1 Fabry Disease Fabry Disease (FD) is a rare genetic X-linked lysosomal disorder caused by the deficiency or absent activity of the enzyme α-galactosidase A (alpha-gal A), which is responsible for destroying a type of fat called globotiaosylceramide (Gb3 or GL-3). When these fat molecules start accumulating due to the lack of alpha-gal A, they start causing damage to the cells. FD has a wide variety of signs and symptoms, that goes from heart attacks and strokes to kidney diseases (Giugliani et al., 2016). With the proper care and treatment, patients with this disease can live longer and have a good life quality (Desnick et al., 2003). 2.1.1 Signs and Symptoms of FD As said before, this disease can be related to many signs and symptoms, that can vary from person to person. Several organs, including the kidney, heart, and brain may be affected (Giugliani et al., 2016). Brain manifestations in FD include progressive white matter lesions (WMLs) (Buechner et al., 2008; Körver et al., 2018a). In (Körver et al., 2018a), WMLs were found in 46% of 1276 patients which tend to occur earlier in males and their prevalence revealed to increase with patients’ age. White matter is the 26 CHAPTER 2. STATE OF ART brain region responsible for the transmission of nerve signals and for communication between different parts of the brain. WMLs have been associated with gait impairment (Starr, 2003) and the risk of falls (Snir et al., 2019; Zheng et al., 2012). Gait abnormalities, such as slower gait and postural instability, have been reported in FD (Löhle et al., 2015). However, not much is known regarding the impact of WMLs on gait performance in patients with FD. Specific location and distribution of WMLs suggest a specific underlying disease (Körver et al., 2018a) and may reflect in a different gait profile. 2.1.2 FD diagnosis “The timely diagnosis of FD is difficult” (Grünfeld, 2003). It is easy to understand that this disease has a hard diagnose, although since it is a genetic disease, analysing the family history can be very helpful. FD results in the accumulation of a fat molecule in cells, with the years the symptoms have a higher percentage to start appearing, although symptoms have been reported in children with two years old (Ramaswami et al., 2006). In the absence of a family member who has already received a diagnosis of the disorder, many cases are not diagnosed until adulthood (average age, 29 years) (Morgan et al., 1987), when the pathology of the disorder may already be advanced. Pain, skin rashes, heat, intolerance, stomach upsets, fatigue, lack of energy, and the inability to exercise are generally the first signs and symptoms to appear, but because these can be associated with other conditions, it may take many years for a diagnosis of Fabry disease to be made. In fact, up to 25% of patients are misdiagnosed (Atul Mehta, Michael Beck & Sunder-Plassmann, 2006). Kidney, heart, and brain problems tend to become noticeable between the ages of 30 to 45 and it is at this point that many individuals with FD are first diagnosed (MacDermot, KD and Holmes & Miners, 2001). What is particularly concerning about FD is that there is an average delay between the onset of symptoms and diagnosis of 12 years. This is the same for both sexes, although the onset of symptoms tends to occur about six years later in females than males (MacDermot, KD and Holmes & Miners, 2001). Since this disease has many symptoms equals to symptoms related to other most commons diseases, the diagnosis becomes harder. In figure 2, we can see the frequency of erroneous diagnoses in patients with FD enrolled in FOS – the Fabry Outcome Survey. 27 CHAPTER 3. MATERIALS 3.2 Electrocardiogram The data provided by the electrocardiogram (ECG) consists of 15 features that analyzed heart rate data and variability, QT (total time for the ventricles of the heart to depolarize and repolarize ( contract and relax)) episodes, and supraventricular ectopy. All features are described in Table 2. Table 2: ECG features. Heart rate data HR Min Minimum heart rate HR Mean Mean heart rate HR Max Maximum heart rate Heart rate variability ASDNN 5 Average standard deviation of all 5 mins normal R-R (heartbeat) intervals SDANN 5 Standard deviation of sequential 5 mins of normal R-R interval means SDNN Standard deviation of the normal R-R intervals RMSSD Root mean square differences of successive R-R intervals QT analysis QT Min Minimum value of the QT interval QT Mean Mean of the QT intervals QT Max Maximum value of the QT interval QTc Min Minimum value for the QT interval corrected for extreme heart rate QTc Mean Mean of the QT intervals corrected for extreme heart rate QTc Max Maximum value for the QT interval corrected for extreme heart rate QTc ≥450 QT interval percentage that is greater than 450 ms corrected for extreme heart rate Supraventricular ectopy Longest R-R Maximum difference between two R-R peaks These ECG features should now be analyzed in order to find the correlation between them and the 34 CHAPTER 3. MATERIALS presence of FD. With this study, it may be possible to find a predictor among these 15 features. In Table 3, are described the normal reference range of values used to control the outcome of the ECG. These values indicate the range of values in which a patient can be considered healthy. Table 3: ECG normal reference range of values (Rijnbeek et al., 2014; Umetani et al., 1998). Electrocardiogram features Normal reference range of values Male Female Heart Rate 65 to 74 66 to 72 QT 378 to 398 390 to 400 QTc 409 to 430 418 to 432 ASDNN 5 (ms) 43 to 88 38 to 66 SDNN 5 (ms) 117 to 182 114 to 147 SDANN (ms) 104 to 162 102 to 133 RMSSD (ms) 22 to 53 22 to 43 3.3 Echocardiogram The data provided by the echocardiogram consists of 22 features described in Table 4. The typical reference range of values for echocardiogram features is referred to in Table 5 and can be used to compare the actual result of a patient to check for abnormality in the outcome of the echocardiogram. 35 CHAPTER 3. MATERIALS Table 4: Echocardiogram features. Echocardiogram features MV E/A Ratio Mitral valve ratio between early diastole (E Wave) and atrial contraction (A Wave) MV A Vel Mitral valve A Wave blood flow velocity MV Dec T Mitral valve deceleration time MV E Vel Mitral valve E Wave blood flow velocity E’ Lateral E Wave using tissue doppler imaging at lateral mitral annulus position E’ Septal E Wave using tissue doppler imaging at septal mitral annulus position E/E’ Lateral Ratio between E Wave and E’ Wave measured at lateral mitral annulus position E/E’ Medial Ratio between E Wave and E’ Wave at medial mitral annulus position E/E’ Septal Ratio between E Wave and E’ Wave measured at septal mitral annulus position LVPWd Left ventricular posterior wall thickness ISVd Interventricular septum thickness at end-diastole LVIDd Left ventricular internal dimension at end-diastole LADiam/SC Left atrial diameter measured at subcostal position AoDiam Aorta diameter S’ Lateral Peak systolic velocity using tissue doppler tissue doppler imaging at lateral mitral annulus position LVdMassInd ASE Left ventricular mass at end-diastole indexed to body surface area according to the American Society of Echocardiography (ASE) LADiam Left Atrial diameter S’ Septal Peak systolic velocity using tissue doppler imaging at septal mitral annulus position A’ Septal A Wave using tissue doppler imaging at septal mitral annulus position A’ Lateral A Wave using tissue doppler imaging at lateral mitral annulus position LVDdMass ASE Left ventricular mass at end-diastole according to the ASE LVIDd/SC Left ventricular internal dimension at end-diastole at subcostal position 36 CHAPTER 3. MATERIALS Table 5: Echocardiogram normal reference range of values (Caballero et al., 2015; El Missiri et al., 2016). Echocardiogram features Normal reference range of values MV E/A Ratio 0.86 to 1.88 MV E Vel (cm/s) 0.61 to 0.93 MV A Vel (cm/s) 0.43 to 0.77 (must be smaller than MV E Vel) MV Dec T (ms) 138.6 to 237.4 E’ Lateral (cm/s) 9.5 to 17.5 E’ Septal (cm/s) 7.3 to 13.3 E/E’ Lateral 4.0 to 8.2 E/E’ Medial 4.6 to 8.6 E/E’ Septal 5.5 to 10.3 S’ Septal (cm/s) 6.7 to 9.5 A’ Septal (cm/s) 7.4 to 11.4 S’ Lateral (cm/s) 7.4 to 12.2 LVPWd (mm) 7.65 to 10.07 IVSd (mm) 7.78 to 10.06 LVIDd (mm) 43.56 to 52.16 LADiam/SC (mm/m2)Not defined AoDiam (mm) 23.22 to 27.88 LVdMassInd(ASE) (g/m2)43 to 115 LADiam (mm) 23.74 to 30.44 LVdMass(ASE) (g) 95.78 to 192.81 LVIDd/SC (mm/m2)Not defined 37 Chapter 4 Methodology 4.1 Multiple Regression Normalization Gait characteristics of a subject are affected by his demographics properties including height, weight, age, and sex, as well as by walking speed (Wahid et al., 2016; Wahid et al., 2015) or stride length (Alcock et al., 2018; Fernandes et al., 2019; Ferreira et al., 2019). To normalize the data regression models according to Wahid et al.’s method (Wahid et al., 2016; Wahid et al., 2015) were used. Comparing to other methods, such as dimensionless equations and detrending methods, MR normalization revealed better results on reducing the interference of subject-specific physical characteristics and gait variables (Mikos et al., 2018; Wahid et al., 2016; Wahid et al., 2015), thereby improving gait classification accuracy using machine learning methods (Fernandes et al., 2019; Wahid et al., 2015). First, to control the multicollinearity among predictor variables within this multiple regression, Variance Inflation Factor (VIF) was calculated (Thompson et al., 2017). This test measures the colinearity of the physical characteristics (age, weight, height, and sex), speed, and stride length, being a value of VIF greater than 5 an indicator of this strong correlation. Spatio-temporal gait variables and foot clearance variables were normalized as follows: ˆyi=β0+ p X j=1 βjxij +εi(4.1) 38 CHAPTER 4. METHODOLOGY where ˆyirepresents the prediction of the dependent variable for the ith observation; xij represents the jth physical property of the ith observation including age, weight, height, sex, speed or stride length, β0 represents the intercept term, βjrepresents the coefficient for the jth physical property and εirepresents the residual error for the ith observation. The model’s coefficients are estimated using the physical properties and the mean values of the gait variables of the healthy controls. Although at least 20 subjects per independent variable are recommended in multiple linear regression (Katz, 2011), based on similar studies (Mikos et al., 2018; Wahid et al., 2016) with higher sample sizes, MR models were computed for all combinations with 1, 2, and 3 independent variables using a bi-square weight function. For the models with all significant independent variables (p-value<0.05) Akaike’s information criterion (AIC) (Burnham & Anderson, 2004) and R-squared metrics were used to select the best-fitted model. Statistical assumptions of a linear regression including linearity, normality, and homoscedasticity were verified. In each subject group, the best fitted MR models are used to normalize each stride gait variable by dividing the original value yiby the predicted gait variable ˆyifrom (4.1), as follow: yn i=yi ˆyi (4.2) where yn irepresents the normalized value for the ith observation. After normalizing all strides of each of the 16 gait variables, the mean and the standard deviation (SD) of each variable (each gait time series) for all subjects were calculated. In this work, the SD value is used to measure the variability of each gait variable. 4.2 Statistical tests Independent T tests or non-parametric Mann–Whitney U tests (if the variable is not normally distributed) were used to compare differences between two independent groups. The normality of data was determined with Shapiro–Wilk tests. Fisher Exact T test was used to compare the differences between the groups when the data is categorical. The theory of statistical test of significance essentially involves the setting of a null hypothesis and 39 CHAPTER 4. METHODOLOGY competing for an alternative hypothesis. The null hypothesis is a hypothesis of “no difference among groups” or “ a sample came from a normally distributed population”. To decide if the null hypothesis is rejected or not, the p-value was calculated. The p-value means the probability of obtaining test results at least as extreme as the results actual observed, under the assumption that the null hypothesis is corrected. In practice, the significance level (α) is stated in advance to determine how small the p-value must be in order to reject the null hypothesis. Conventionally, it is considered statistically significant as p-value < α when α=0.05 and statistically highly significant as p-value < αwhen α=0.001 (Kirkwood & Sterne, 2003). 4.3 Machine Learning Machine Learning (ML) differs from typical programming because the goal here is to make computers learn from data, as said by Arthur Samuel (1959): “field of study that gives computers the ability to learn without being explicitly programmed.” Since this is an era where data can be accessed pretty easily, ML algorithms can revolutionize everything. Typically, to write a piece of code, we write all the rules necessary to make it logical as in Figure 3. Figure 3: Traditional approach to data analysis project (adapted from (Géron, 2019)). In a complex situation, the code can be pretty complicated and heavy. Another solution is to write a ML algorithm that can detect a pattern and learn by itself, making it much shorter, easier to maintain, and 40 CHAPTER 4. METHODOLOGY more accurate (Figure 4). Figure 4: Machine learning approach to data analysis project (adapted from (Géron, 2019)). There is also the advantage of finding patterns in a large amount of data, that human eyes cannot detect (data mining). The first step is to know exactly what is the objective of the ML algorithm. This is important because it will determine how to frame the problem, what algorithms to select and what performance measure should be used to evaluate the model. Second step consists of getting the data. Without data, ML is not possible because the computer needs “something” to learn from. Usually, it is provided a dataset (database table where every column of the table represents a specific variable (or feature), and each row corresponds to a sample) to the ML algorithm. Missing data, duplicate data, outliers, a large number of variables, or highly correlated variables are possible issues that originate at the source and that should be overcome to ensure good data. The performance of ML is greatly affected by the quality of the data. Different techniques such as missing data imputation, outlier detection, dimensionality reduction data transformations, are first applied to get data with quality for ML modeling (Gudivada et al., 2017). After the dataset is ready, it is time to create the test, train, and validation set: •Training Dataset: The sample of data used to fit the model. The actual dataset that we use to train the model. The model reads and learns from this data. 41 CHAPTER 4. METHODOLOGY •Validation Dataset: The sample of data used to provide an unbiased evaluation of a model fit on the training dataset while tuning model hyperparameters. This data is used to fine-tune the model hyperparameters hence the model never “learns” from it. •Test Dataset: The sample of data used to provide an unbiased evaluation of a final model fit on the training dataset. The test set provides the standard used to evaluate the model. Next, it is important to keep the relevant features so that the training is not clouded with irrelevant information because having many features will not help to have a better outcome (exactly the opposite). The model will only be capable of learning if the training data contains enough relevant features and not too many irrelevant ones. So it becomes critical to its success coming up with a good set of features to train on. This process is called feature engineering and involves: •Feature selection: selecting the most useful features to train on among existing features creating a subset of the original attributes (features or variables); •Feature extraction (or Feature projection): combining existing features to produce a more useful one, reducing this way the high-dimensional spaces into a space with fewer dimensions achieved by dimension reduction algorithms; • Creating new features by gathering new data. 4.4 Feature selection It is possible to state that feature selection is an inevitable part of a classifier design. For instance, the classifier-independent univariate filter methods have similar trends. Filter methods such as the T Test, Mann Whitney U Test or Fisher Exact T Test have better or similar performance with wrapper methods for harder problems. This improved performance is usually accompanied by significant peaking (Hua et al., 2009). 42 CHAPTER 4. METHODOLOGY In this study, a hybrid method (filter method followed by a wrapper method) was employed to select the most relevant gait and cardiac characteristics. First, a filter method based on Mann Whitney Utests and Spearman’s correlation between the variables was used to selected 10 gait characteristics. The selected 10 gait or cardiac characteristics are the ones that present a higher U-value and do not present a high correlation between them (|ρ|<0.90). Before applying a wrapper method, the selected 10 variables were scaled to have zero mean and unit variance. Based on previous work (Rehman et al., 2019), the wrapper was developed using the Recursive Feature Elimination (RFE) technique with three different ML classifiers: Logistic Regression (LR), SVM with Linear kernel, and RF. RFE has some advantages over other filter methods (Iguyon & Elisseeff, 2003). RFE is an iterative method where features are removed one by one without affecting the training error. The selection of the optimal number of features for each model is based on the evaluation metric F1 score (see definition in section 4.6) evaluated through 5-fold cross-validation, that consists in a procedure that divides a limited dataset into 5 non-overlapping folds, where each fold is at some point of the iteration used as a held-back test set, whilst all other folds collectively are used as a training dataset. The gait characteristics’ importance was quantified using the model itself (feature importance for LR and SVM with linear kernel and information gain for RF). The F1 score was used to assess the performance of the different gait characteristics combinations. 4.5 Models theoretical basis 4.5.1 Logistic Regression Some regressions can be used for classification. Logistic Regression is commonly used to estimate the probability that an instance belongs to a particular class. If the estimated probability is greater than 50%, then the model predicts that the belongs to that class, or else it predicts that it does not. Therefore, Logistic Regression is a binary classifier. To compute this, Logistic Regression computes a weighted sum of input features (plus a bias term), 43 CHAPTER 4. METHODOLOGY 2013). In R, the Minkowski metric can be employed: ||x0−xj||p= ( q X i=1 |(xi)0−(xi)j|p)1/p(4.18) which corresponds to the euclidean distance for p=2. For binary classification, KNN is defined as: fKNN (x0) =      1if Pi∈Nk(x0)yi≥0 −1if Pi∈Nk(x0)yi<0 (4.19) with neighborhood size Kand with the set of indices Nk(x)of the K-nearest patterns. The choice of Kdefines the locality of KNN. For a smaller Klittle neighborhoods arise in regions, where patterns from different classes are scattered. For bigger choices of K, patterns with labels in the minority are ignored. Therefore, the big question when building a KNN model is what Kto choose to achieve the best neighborhood and, consequently, the best performance. 4.5.5 Convolutional Neural Networks Convolutional Neural Networks (CNN’s) were developed back in 1989 by LeCun to perform automatic recognition of handwritten zip code digits by the U.S. Postal Service (LeCun et al., 1989). These networks are intended to process data of a grid-like structure nature. Time-series can be processed by this model if they are thought as a 1D grid of values with different time-steps. Also, imaging data can be thought of as 2D or a 3D grid, depending on the color channel is included (Papernot et al., 2016). CNN’s key idea is the convolution operations performed to input data. A CNN is fully filled with three processing layers: convolutional layers, pooling layers, and fully connected layers (Papernot et al., 2016). Convolutional layers This is the first layer of a CNN model and is based on the discrete convolution operation. This operation can be described as an input Drepresented by a 2D array of size n1∗n2convolved with a filter (also called kernel) of size a∗a. The filter slides across the input Daccording to the stride parameter. The output 50 CHAPTER 4. METHODOLOGY of a convolutional layer is a collection of feature maps of size q1∗q2, the number of generated feature maps is given by the number of filters i(Koushik, 2016). The ith feature map of the lth convolutional layer denoted C(l) iis computed as follows: C(l) i=B(l) i+ a(l−1) i X j=1 K(l−1) ij ∗C(l−1) j(4.20) where B(l) iis the bias, K(l−1) ij is the filter of size a∗athat connects the jth feature map in layer (l−1) with the ith feature map in layer land C(l−1) jrepresents the jth feature of layer l−1. The input of the first convolutional layer (l= 1) is the input data D, that is C(0) 1=D. Following the convolution layer, an activation function is applied to the feature maps and the output of the ith feature map is: Y(l) i=g(C(l) i)(4.21) where g(x)represents an activation function (ReLU for this work). Pooling layers Following the convolutional layer is always the pooling layer and its objective is to replace the feature maps’ output values at a certain location with a statistical summary of the nearby values reducing the dimensionality of the feature maps. This objective is achieved using a mask of size b∗bto perform a pooling operation on each of the feature maps. From the most common pooling operations available, maximum pooling operation was selected for this work and it outputs the maximum value within a neighborhood (Koushik, 2016; Papernot et al., 2016). This pooling operation can also increase the robustness of the model to noise and distortions since after this operation the representations of the feature maps are able to become approximately invariant to small translations of the input. Therefore, pooling layer is essential to improve the performance of the network for unseen data, but, no learning is done at the pooling layers. The only objective is to summarize the output responses over a neighborhood (Papernot et al., 2016). 51 CHAPTER 4. METHODOLOGY Fully connected layers Third and final layer of CNN is the fully connected layer. When convolutional and pooling layers are finished, the output of the last pooling layer is flattened and given as input to fully connected layers. A cost function is used to measure the discrepancy between the output of the network, the class labels, and the weights of the CNN are updated using backpropagation and Adam optimization method (Papernot et al., 2016). 4.5.6 Recurrent Neural Network Recurrent Neural Networks (RNNs) are a class of neural networks specialized in processing sequential data. They are built with a chain-like structure and can be used to process time series and predict their outcome (Papernot et al., 2016). The key idea is to share parameters across different parts of a model, sharing the same parameters across the different time steps (Graves, 2012; Papernot et al., 2016). A sequence that contains vectors x(t)with the time index tranging from 1 to τ, the value of the hidden node can be defined as (Papernot et al., 2016): h(t)=f(h(t−1), x(t);θ)(4.22) where h(t)represents the state of a hidden node and θrepresents the parameters of the network. the hidden state is then able to save a ”memory” of previous inputs and therefore influence the output of the network. The hidden nodes receive two input signals (Graves, 2012). An RNN with Iinput nodes, Hhidden nodes, and Koutput nodes, the forward pass can be computed as follows: u(t) h= I X i=1 wihx(t) i+ H X h=1 whhz(t−1) h(4.23) where wih represents the weights between the input vector xand the hidden node h,whh represents the weights between hidden node hat time step t−1and tand utrepresents the internal value of h 52 CHAPTER 4. METHODOLOGY at time step t. the output value z(t) hof the hidden node his computed by applying an activation function g(x): z(t) h=g(u(t) h)(4.24) The output value α(t) kat time step tis then computed as: α(t) k= H X h=1 whkz(t) h(4.25) z(t) k=g(a(t) k)(4.26) The loss of a RNN is the sum of the time step losses: J(θ) = τ X t=1 J(θ)(4.27) where the J(θ)represents the cost function. The training process of a RNN is made through backpropagation and this is how the gradients of the cost function respective to the parameters are calculated. In this specific case, the algorithm is backpropagation through time (BPTT). This algorithm unfolds an RNN into a Multilayer Perceptron allowing the application of standard back-propagation (Haykin, 2009). Using this, the gradients are easy to compute but very difficult to train due to vanishing or exploding gradient problems where the influence of the early time-steps either decays or blows up exponentially as the sequence is processed by the network (Allen-Zhu et al., 2018; Bengio et al., 1994; Sutskever et al., 2014). The problem is that when applying backpropagation from z(t) hto z(t−1) ha multiplication by WT hh is performed. This means that many factors of WT hh are involved in the gradient computation of z(0) h. The value of the gradients will either vanish or explode depending on the magnitude of WT hh. This can be described as (Sutskever et al., 2014): ∂J ∂Whh = τ X t=1 dz(t) kz(t−1)T h(4.28) 53 CHAPTER 4. METHODOLOGY where dz(t) k= ( τ Y t=1 WT hhg0(z(t) k))(4.29) If the values of WT hh are small then the contributions of the older inputs will rapidly tend to zero as t increases. The algorithm closest to solve this issue is Long Short-Term Memory (LSTM). Long Short-Term Memory Long Short-Term Memory (LSTM) is part of the RNNs algorithms. It is based on the idea of creating paths through time that derivatives that neither vanish nor explode. It was developed in 1997 by Hochreither and Schmidhuber and since then it has been further developed and applied to various fields with a high degree of success. It works by a cell state which has a linear self-loop controlled by different gates (forget and input gate) and the output is computed by the output gate. The operations of an LSTM (Papernot et al., 2016) can be described as: Forget Gate: f(t) i=bf i+ I X j=1 Uf i,jx(t) j+ K X j=1 Wi,jh(t−1) j(4.30) f(t) i=σ(f(t) i)(4.31) Input Gate: s(t) i=σ(bs i+ I X j=1 US i,jx(t) j+ K X j=1 Ws i,jh(t−1) j)˜ C(t) i(4.32) ˜ C(t) i=tanh(bc i+ I X j=1 Uc i,jx(t) j+ K X j=1 Wc i,jh(t−1) j)(4.33) Cell State: C(t) i= (f(t) iC(t−1) i) + s(t) i(4.34) Output Gate: o(t) i=σ(bo i+ I X j=1 Uo i,jx(t) j+ K X j=1 Wo i,jh(t−1) j)(4.35) 54 CHAPTER 4. METHODOLOGY h(t) i=o(t) itanh(C(t))(4.36) where is the symmetric tensor product and b, U, W are the bias, input weights, and recurrent weights of the LSTM gate respectively. σrepresents the sigmoid function, tanh represents the hyperbolic tangent function, f(t) irepresents the forget gate for time step tand cell i,s(t) irepresents the input gate for time step tand cell i,C(t) irepresents the cell state for time step tand cell i,o(t) irepresents the output gate for time step tand cell iand h(t) irepresents the hidden value of hidden node iat time step t. As RNNs, LSTM can be trained using backpropagation and gradient descent. The vanishing gradient problem is tackled because the LSTM cell state provides a way for the gradient to flow allowing the network to learn long term dependencies (Papernot et al., 2016). To go further into the analysis, CNN and LSTM approach were implemented with the group of features selected above from the recursive feature selection method to check if these two methods can improve the metrics. The features selected included its mean and standard deviation, but for CNN and LSTM the all stride values of the gait variables were used on the iteration. To perform these models, an algorithm iterating across different combinations of the number of neurons, learning rate, and kernel were developed to achieve the best model. The process for both CNN and LSTM is described in the flowchart in Figure 5. Figure 5: CNN and LSTM flowchart. From Figure 5, first, from the raw full stride series of the gait assessment the MR normalization described in Section 4.1. was performed. From this outcome, a feature selection algorithm using recursive feature elimination was developed to select which gait variables are used in CNN and LSTM models. 55 CHAPTER 4. METHODOLOGY 4.6 Machine Learning metrics All classifiers hyperparameters were tuned using randomized search and grid search method with 5fold cross-validation: LR regularization strength constraint and type of penalty according to regularization strength constraint; SVM regularization parameter, gamma, and different types of kernel (with different degrees); RF maximum depth, minimum samples leaf, minimum samples split, and the number of estimators; finally, KNN number of neighbors, weights, and metric. All classifiers were implemented in Python programming language using Scikit-learn library (Pedregosa et al., 2011). After tweaking your models for a while, eventually, the system performs sufficiently well. Now is the time to evaluate the final model on the test set. Often classification accuracy is used to measure the performance of the model, however, it is not enough to truly judge it. The most common metrics to evaluate a classification model are: •accuracy: the ratio of correct predictions; •precision: the ratio of correctly predicted positive observations to the total predicted positive observations; •sensitivity/recall: the ratio of positive instances that are correctly detected by the classifier; •specificity: the ratio of negative instances that are correctly detected by the classifier; •F1 Score: the harmonic mean between precision and recall whose range goes from 0 to 1 (Gu et al., 2009). It tells how precise the classifier is (how many instances it classifies correctly), as well as how robust it is (it does not miss a significant number of instances): F1Score =1 1 P recision +1 Recall (4.37) 56 Part III : Results, Discussion, Conclusion, and Future Work 57 Chapter 5 Gait assessment 5.1 Participants Gait assessment data from seventy-six subjects (28 males) were collected. Of this group, thirty-nine of the patients have FD (14 males) and thirty-seven are healthy subjects (14 males). Of the thirty-nine FD patients, twenty-five patients have White Matter Lesions (WMLs)(7 males). The subject demographics of the groups of FD with WMLs, FD without WMLs, and controls are summarized in Table 6. For all FD patients, the exclusion criteria were: less than eighteen years of age, the presence of resting tremor, moderate-severe dementia (CDR >2), depression, extensive intracranial lesions, or neurodegenerative disorders, musculoskeletal disease, and rheumatological disorders. Local hospital ethics committee approved the protocol of the study, submitted by ICVS/UM and Center Algoritmi/UM. Written consent was obtained from all subjects or their guardians. 58 CHAPTER 5. GAIT ASSESSMENT Table 6: Demographic variables for FD patients with and without WMLs and Controls. FD patients with WMLs (n = 24) FD patients without WMLs (n = 15) Controls (n = 37) Age (years) 59.29 (15.99) 36.93 (10.99) 52.76 (22.91) Male (%) 29% 47% 38% Weight (kg) 66.03 (10.47) 64.69 (6.67) 66.51 (9.12) Height (m) 1.59 (0.07) 1.66 (0.09) 1.62 (0.09) Data is presented as mean (standard deviation) 5.2 Gait Normalization The gait normalization was performed based on the physical properties and the mean values of the gait variables of the 37 healthy controls. Variance Inflation Factor (VIF) was employed to measure the colinearity of the physical characteristics (age, weight, height, and sex), speed, and stride length and the results are presented in Table 7. Table 7: Variance inflation factor for physical characteristics (age, weight, height, and sex), speed, and stride length. Age Weight Height Sex Speed Stride Length VIF 2.24 1.97 3.76 1.28 6.05 9.89 1.99 1.93 3.37 1.27 - 2.06 1.82 1.97 2.9 1.28 1.26 - Since when all variables are used both speed and stride length had a VIF value higher than 5 (an indicator of strong correlation), they can not be used simultaneously. When used separately no VIF value is higher than 5. Therefore, from the gait variables speed and stride length, only stride length is used on further investigation. All possible combinations for each feature were computerized by a developed algorithm, allowing to find the best possible model for each feature. The models created for each feature are summarized in Table 8 (right foot) and in Table 38 (left foot). The models obtained for both feet shown in Table 8 and Table 38 are very similar, therefore, all the 59 CHAPTER 5. GAIT ASSESSMENT Figure 10: F1 score vs number of features for selection of optimal numbers of gait characteristics (left) and feature importance results (right) obtained based on Logistic Regression, Support Vector Machine (SVM) Linear kernel and Random Forest in FD patients with WMLs vs controls. Recursive feature elimination was used through the 5-fold cross-validation (RFECV). Max.: maximum; Var.: variability; Min.: minimum. 5.3.3 Gait classification based on gait time series The groups selected were Top 3 (same as Top LR and Top 4): pushing, cycle duration, and foot flat; Top SVM: pushing, cycle duration, foot flat, maximum toe clearance 2 and maximum heel clearance; and 66 CHAPTER 5. GAIT ASSESSMENT Figure 11: Top 3 and Top 4 after performing RFE on gait features for FD patients with WMLs vs controls. Table 11: Classification accuracy on training and validation data for top common gait characteristics in FD patients with WMLs vs controls with LR, SVM (linear and RBF kernel), RF, and KNN. Top 3 Top 4 Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) LR 64.58 ±2.05 63.00 ±12.49 67.66 ±2.63 63.50 ±14.46 SVM (Linear Kernel) 67.63 ±5.24 61.50 ±15.46 65.00 ±6.09 61.50 ±16.70 SVM (RBF Kernel) 60.92 ±5.06 64.50 ±10.05 66.11 ±4.92 61.50 ±14.11 RF 77.61 ±2.04 71.50 ±10.91 70.18 ±3.67 64.50 ±4.58 KNN 71.04 ±7.43 60.57 ±13.08 74.38 ±3.64 60.12 ±15.30 LR: Logistic Regression, RBF: radial basis function, RF: Random Forest, KNN: K-Nearest Neighbour, SD: standard deviation Top RF: pushing, cycle duration, foot flat, maximum toe clearance 1, maximum heel clearance, swing, and maximum toe clearance 2. Table 12 shows the results for CNN and LSTM performed based on the gait time series of the FD patients with WML vs controls. Top RF achieved the best performance for both CNN and LSTM with a validation accuracy of 71.61% and 69.39%, respectively. Comparing CNN and LSTM algorithms, CNN achieved the highest percentage of validation accuracy for Top 5 and Top SVM, while Top 3 had the highest validation accuracy in LSTM method. Finally, comparing the results from Table 10, Table 11, and Table 12, Top 3 had a better performance 67 CHAPTER 5. GAIT ASSESSMENT Table 12: CNN and LSTM results on selected gait features from the gait time series for FD patients with WMLs vs controls. CNN LSTM Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Top 3 62.79 ±3.69 58.89 ±12.40 62.18 ±7.14 61.39 ±8.82 Top SVM 76.72 ±8.82 73.83 ±16.93 75.57 ±9.28 69.11 ±10.59 Top RF 83.11 ±6.71 71.61 ±17.35 75.03 ±12.76 69.39 ±15.54 CNN: Convolutional Neural Networks, LSTM: Long-Short Term Memory, SD: standard deviation using the first approach with RF algorithm with a validation accuracy of 71.50%. On the other hand, in Top SVM and Top RF both CNN and LSTM outperformed the rest of the algorithms, being CNN the algorithm with the best metrics achieving 73.83% and 71.61%, respectively. Regarding the standard deviation of the validation accuracy, there are no major differences between the results from all the models. 5.4 FD patients without WMLs vs controls To FD without WMLs vs Controls differentiation task was selected from the control group 15 subjects aged-matched with the group of FD patients without WMLs. The physical characteristics of these two groups are described in Table 13. Table 13: Demographic variables for FD patients without WMLs and Controls. FD patients without WMLs (n = 15) Controls (n = 15) p-value Age (years) 36.93 (10.99) 37.00 (13.12) 0.87 Male (n(%)) 47% 47% 1 Weight (kg) 64.69 (6.67) 69.73 (13.82) 0.325 Height (m) 1.66 (0.09) 1.71 (0.08) 0.246 Data is presented as mean (standard deviation). The comparison between the mean value of gait features in FD patients without WMLs vs controls for 68 CHAPTER 5. GAIT ASSESSMENT the raw gait data and the MR normalized gait is represented in Figure 12, where significant differences in gait features between FD patients without WMLs and controls are indicated with one asterisk (*p<0.001), and whiskers represent 95% confidence interval (CI) values. The data was scaled between 0 and 1 to fit onto the same plot. When using gait raw data: cycle duration, cadence, loading, foot flat, pushing, peak swing, strike angle, lift-off angle, maximum heel clearance, and minimum toe clearance are statistically significant between FD patients without WMLs and controls (for more details see Table A.37). After normalization, pushing is not statistically significant while stride length is now statistically significant. 69 CHAPTER 5. GAIT ASSESSMENT Figure 12: Comparison between the mean value of gait features in FD patients without WMLs vs controls. Data are shown for the raw gait data and the MR normalized gait. Max.: maximum; Var.: variability; Min.: minimum; H.C. : heel clearance; T.C. : toe clearance 70 CHAPTER 5. GAIT ASSESSMENT 5.4.1 Feature selection Recursive feature elimination algorithm was performed on the 10 remaining features from filter method: foot flat, peak swing, strike angle, minimum toe clearance mean and variability, stride length variability, loading mean and variability, lift-off angle variability, and maximum heel clearance variability. The heat map of these 10 features selected from the filter method algorithm is shown in Figure 13. Results are stated in Table 14 and Figure 14. Figure 13: Gait correlation heat map in FD patients without WMLs vs controls. Max.: maximum; Var.: variability; Min.: minimum. The results show that LR and SVM selected 5 features as the optimal number of gait characteristics with a F1 score of 72.31% and 74.90%, respectively, and RF selected 6 features with a F1 score of 79.81%. Table 14 presents the training and validation accuracies for the optimal models of each algorithm. LR 71 CHAPTER 5. GAIT ASSESSMENT presents the higher training and validation accuracies. Figure 14: F1 score vs number of features for selection of optimal numbers of gait characteristics (left) and feature importance results (right) obtained based on Logistic Regression, Support Vector Machine (SVM) Linear kernel and Random Forest in FD patients without WMLs vs controls. Recursive feature elimination was used through the 5-fold cross-validation (RFECV). Max.: maximum; Var.: variability; Min.: minimum. Taking into account the contribution of each gait characteristic in the classification model (Figure 14, right side), the common features were 3 (Top 3): loading mean and variability and lift-off angle variability. The common gait characteristics from LR and SVM were 5 (Top 5): foot flat, stride length variability, 72 CHAPTER 5. GAIT ASSESSMENT Table 14: Results for each of 3 machine learning algorithms (LR, SVM with linear kernel, and RF) with optimal gait characteristics’ from recursive feature elimination for FD patients without WML vs controls. Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) LR 88.33 ±3.12 76.67 ±8.16 SVM (Linear Kernel) 85.00 ±3.33 76.67 ±8.16 RF 94.17 ±4.25 70.00 ±12.47 SD: standard deviation loading mean and variability, and lift-off angle variability (see Figure 15). Figure 15: Top 3, Top 5, and Top 3 (LR & SVM) after performing RFE on gait features for FD patients without WMLs vs controls. 5.4.2 Gait classification based on gait measures These Top 3 and Top 5, as well as the Top 3 from LR and SVM (stride length variability, loading variability and lift-off angle variability), as shown in Figure 15, were evaluated with five classification models (LR, SVM Linear kernel, SVM RBF kernel, RF, and KNN) to identify the optimal combination of gait 73 CHAPTER 5. GAIT ASSESSMENT Table 15: Classification accuracy on training and validation data for top common gait characteristics in FD patients without WMLs vs controls with LR, SVM (linear and RBF kernel), RF, and KNN. Top 3 Top 5 Top 3 (LR & SVM) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) LR 83.33 ±4.56 80.00 ±12.47 88.33 ±3.12 76.67 ±8.16 80.83 ±2.04 80.00 ±12.47 SVM (Linear Kernel) 85.00 ±3.33 73.33 ±13.33 85.00 ±3.33 76.67 ±8.16 83.33 ±3.73 76.67 ±8.16 SVM (RBF Kernel) 80.00 ±4.08 80.00 ±19.44 87.50 ±4.56 80.00 ±6.67 78.33 ±4.86 70.00 ±12.47 RF 92.50 ±7.17 70.00 ±16.33 95.83 ±3.73 70.00 ±12.47 90.83 ±1.67 73.33 ±8.16 KNN 91.67 ±2.64 83.33 ±10.54 91.67 ±2.04 86.67 ±8.16 93.33 ±4.56 86.07 ±12.47 LR: Logistic Regression, RBF: radial basis function, RF: Random Forest, KNN: K-Nearest Neighbour, SD: standard deviation characteristics and the classification model with better performance. From Table 15, KNN achieved the highest validation accuracy in all three different top’s with a validation accuracy of 83.33% in Top 3, 86.67% in Top 5 from LR and SVM, and 86.07% in Top 3 from LR and SVM. The validation accuracy in KNN and SVM Linear Kernel increased by changing from Top 3 to Top 5 while the other models decreased (except SVM RBF Kernel who remained equal). Going from Top 5 to Top 3 from LR and SVM, produced a slight decrease in SVM RBF Kernel and increase in LR and RF (SVM Linear Kernel and KNN had no change). Overall, KNN showed higher mean validation accuracy followed by SVM RBF Kernel. 5.4.3 Gait classification based on gait time series The groups selected were Top 5 (same as Top LR and Top SVM): lift-off angle, loading, foot flat, and stride length; Top RF: loading, maximum heel clearance, minimum toe clearance, lift-off angle, and peak swing; Top 3: lift-off angle and loading; and Top 3 LR & SVM: stride length, loading, and lift-off angle. Table 16 shows the results for CNN and LSTM performed based on gait time series of the FD patients with WML vs controls. Top 5 achieved the best performance when using CNN algorithm with a validation accuracy of 80.48% while Top RF achieved the best performance for LSTM with a validation accuracy of 67.62%. Comparing CNN and LSTM algorithms, CNN outperformed LSTM in all group of features. Finally, comparing the results from Table 14, Table 15 and Table 16, KNN achieved the highest vali74 CHAPTER 5. GAIT ASSESSMENT Table 16: CNN and LSTM results on selected gait features from the gait time series for FD patients without WMLs vs controls. CNN LSTM Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Top RF 84.53 ±10.38 69.05 ±25.95 82.23 ±21.98 67.62 ±23.65 Top 3 85.37 ±14.89 65.24 ±18.53 59.73 ±19.75 48.57 ±22.08 Top 5 87.93 ±10.12 80.48 ±19.69 75.00 ±13.01 60.48 ±14.17 Top 3 (LR & SVM) 77.33 ±9.28 64.76 ±11.21 71.73 ±16.00 55.24 ±17.78 CNN: Convolutional Neural Networks, LSTM: Long-Short Term Memory, SD: standard deviation dation for Top 3, Top 5, and Top 3 (LR & SVM) with a validation accuracy of 83.33%, 86.67%, and 86.07%, respectively. For Top RF, using RF algorithm the validation accuracy was higher with 70%. Regarding the standard deviation of the validation accuracy, it showed higher for all the results in both CNN and LSTM. 5.5 FD patients with vs without WMLs For the analysis of FD patients with WMLs vs FD patients without WMLs all patients were included. The physical characteristics of these two groups are described in Table 17. Table 17: Demographic variables for FD patients with and without WMLs. FD patients with WMLs (n = 24) FD patients without WMLs (n = 15) p-value Age (years) 59.29 (15.99) 36.93 (10.99) <0.001 Male (%) 29% 47% 0.318 Weight (kg) 66.03 (10.47) 64.69 (6.67) 0.743 Height (m) 1.59 (0.07) 1.66 (0.09) 0.011 Data is presented as mean (standard deviation). The comparison between the mean value of gait features in FD patients with vs without WMLs for the 75 CHAPTER 5. GAIT ASSESSMENT Table 20: CNN and LSTM results on selected gait features from the gait time series for FD patients with vs without WMLs. CNN LSTM Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Top LR 88.53 ±11.87 78.93 ±16.34 87.22 ±8.81 68.93 ±10.98 Top SVM 89.19 ±9.00 71.07 ±16.88 79.54 ±8.39 66.43 ±17.40 Top RF 86.01 ±14.70 70.71 ±24.50 78.27 ±12.32 73.93 ±14.79 Top 3 95.63 ±8.75 81.43 ±14.49 75.08 ±6.41 71.43 ±13.27 Top 5 92.40 ±7.80 75.71 ±27.26 76.37 ±7.34 68.93 ±7.63 CNN: Convolutional Neural Networks, LSTM: Long-Short Term Memory, SD: standard deviation and Top 5, being only outperformed in Top RF. Finally, comparing the results from Table 18, Table 19, and Table 20, CNN achieved the highest validation for Top 3 with a validation accuracy of 81.43% and LR achieved the best performance for Top 5 with a validation accuracy of 80.76%. For Top LR, using LR algorithm the validation accuracy was higher with 80.76%. For Top RF, using LSTM achieved the best validation accuracy with 73.93%. Finally, for Top SVM, the better performance was obtained using SVM algorithm with a validation accuracy of 77.90%. Regarding the standard deviation of the validation accuracy, using CNN and LSTM the standard deviation values were slightly higher than with the other algorithms. 5.6 Discussion Based on previous literature (Aich et al., 2018; Fernandes et al., 2020; Mannini et al., 2016; Pradhan et al., 2015; Rehman et al., 2019; Wahid et al., 2015) different classification models were evaluated with different sets of gait characteristics selected using a filter method based on Mann Whitney Utests and Spearman’s correlation between the variables followed by a recursive feature elimination wrapper method with RF, SVM Linear kernel, and LR. Sixteen gait time series were obtained by two wearable sensors. All strides were normalized before developing any ML model according to previous studies (Fernandes 82 CHAPTER 5. GAIT ASSESSMENT et al., 2020; Mikos et al., 2018; Wahid et al., 2016). Then, for each gait time series the mean and the standard deviation (as a variability measure) were calculated obtaining 32 gait characteristics. From the feature selection analysis, foot flat, pushing, and maximum toe clearance 2 were identified as important characteristics to classify FD with WMLs. While stride length variability, loading variability, and lift-off angle variability, followed by loading, foot flat, and minimum toe clearance were identified as important gait characteristics to distinguish FD without WMLs from aged-matched healthy adults. Previous work (Fernandes et al., 2020) reveals that FD patients (with and without WMLs together) present lower percentages in foot flat and higher in pushing comparing with healthy adults. For FD patients with WMLs versus controls, validation accuracy of 62-72% and a similar training accuracy of 58-83% was achieved through the five selected classification models based on Top 3 gait characteristics, showing RF classifier the best performance with validation and training accuracy of 72% and 78%, respectively. With one more feature (foot flat variability) RF and LR revealed good performance with an accuracy of 65% and 64% for validation and 70% and 68% for training. These results corroborated the hypothesis that the gait characteristics can be used to distinguish FD patients with WMLs from controls. This goes in line with the premise that gait is a final outcome of WMLs (Snir et al., 2019; Starr, 2003; Zheng et al., 2012). Surprisingly, in the FD patients without WMLs versus controls classification higher training and validation accuracies of 60-93% and 49-86%, respectively, were obtained based on Top 3, Top 5 or Top 3 LR & SVM. KNN classifier displayed the best performance based on Top 3, with an accuracy of 83% for validation and 92% for training. By increasing the feature set for Top 5 (adding stride length variability and foot flat), overall validation accuracy of 61–87% was achieved, where for the SVM Linear Kernel, KNN, and CNN classifiers the accuracy slightly increased. Finally, reducing the number of features to 3 with Top 3 LR & SVM showed KNN as the best classifier with a top accuracy of 86% for validation and 93% for training. Also, reducing the number of features increased the performance of LR and RF. Similarly, in (Rehman et al., 2019) an increase in the model accuracy was observed with feature reduction. FD patients with vs without WMLs overall performance across all algorithms stood with 62-81% for validation and 61-96% for training. When using Top 5, LR produced the best performance with an 81% validation accuracy and 85% training. Reducing the number of features to 3 (Top 3) improved the accuracy of the best performance to 81% for validation and 96% for training. Further, feature selection (reduction) 83 CHAPTER 5. GAIT ASSESSMENT plays an important role to deal with the problem of model overfitting, reduces training time, enhancing the overall ML performance and implementation. These results suggest that selected gait characteristics could be used as clinical features for supporting diagnoses of FD patients even without WMLs from younger ages since the mean age of these patients is 36.93 ±10.99 years. Due to the number of subjects involved in this study, all dataset was used in the training and validation of the models and any independent (external) dataset was used for checking the model performances. To test the robustness of classification models based on the selected gait characteristics further research with independent datasets is needed. 84 Chapter 6 Electrocardiogram 6.1 Participants and Holter data One hundred and fourteen FD patients (44 males) were evaluated with a 24-hour ambulatory ECG (Holter) exam. From this group, 61 have WMLs (23 males). The group of patients with WMLs was significantly older than the group of patients without WMLs (Table 21). The age of about half of patients with WMLs ranged from 49 to 67 years and 91 years is the age of the oldest one while the age of 50% of FD patients without WMLs ranged from 34 to 55 years years and the oldest patient has 76 years old (Figure 20). Table 21: Demographic characteristics for 114 FD patients with and without WMLs. With WMLs (n = 61) Without WMLs (n = 53) p-value Age (years) 55.28 (16.52) 44.13 (14.90) <0.001 Male (%) 38% 40% 0.849 Data is presented as mean (standard deviation) plus p values from Mann Whitney U Test. Regarding the Holter data, Table 22 shows that there is a significant difference between the two groups on six ECG features: heart rate maximum, QT mean, minimum, and maximum, QT corrected mean, and QT corrected ≥450. 85 CHAPTER 6. ELECTROCARDIOGRAM Figure 20: Boxplot of age stratified for the presence of WMLs in FD patients. Table 22: ECG features for 114 FD patients with and without WMLs. With WMLs (n = 61) Without WMLs (n = 53) p-value HR Min 50.39 (5.88) 50.45 (6.84) 0.75 HR Mean 74.69 (7.32) 77.4 (8.43) 0.0197 HR Max 121.45 (12.5) 134.52 (16.91) <0.001 ASDNN 5 58.83 (22.58) 59.92 (19.18) 0.301 SDANN 5 119.17 (30.07) 120 (41.36) 1 SDNN 137.37 (32.71) 136.56 (44.73) 0.88 RMSSD 58.74 (52.9) 44.09 (29.6) 0.242 QT Min 309.08 (26.91) 292.5 (27.46) 0.005 QT Mean 403.12 (33.96) 384.48 (27.04) 0.003 QT Max 490.14 (76.93) 466.36 (67.49) 0.019 QTc Min 374.19 (34.3) 373.29 (32.07) 0.483 QTc Mean 443.73 (29.48) 430.26 (24.93) 0.007 QTc Max 556.33 (69.52) 545.43 (73.44) 0.197 QTc ≥450 35.33 (35.34) 19 (27.16) 0.012 Longest R-R 1.84 (0.95) 1.61 (0.4) 0.952 Data is presented as mean (standard deviation) plus p-values from T test or Mann Whitney U Test for normally or non-normally distributed continuous variables or Fisher Exact T Test for categorical variables. 86 CHAPTER 6. ELECTROCARDIOGRAM From the univariate logistic regression, heart rate maximum, QT minimum, QT mean, QT corrected mean, and QT corrected ≥450 present a significant association with the presence of WMLs. However, when adjusted by age no ECG features present a significant association with the presence of WMLs (Table 23). Table 23: Univariate and multivariate (adjusted for age) logistic regression model to predict the presence of WMLs. Univariate Logistic Regression Adjusted for age p-value p-value HR Min 0.739 0.467 HR Mean 0.228 0.613 HR Max <0.001 0.210 ASDNN 5 0.487 0.144 SDANN 5 0.827 0.331 SDNN 0.876 0.244 RMSSD 0.098 0.236 QT Min 0.005 0.061 QT Mean 0.003 0.361 QT Max 0.108 0.376 QTc Min 0.887 0.333 QTc Mean 0.004 0.367 QTc Max 0.304 0.362 QTc ≥450 0.004 0.720 Longest R-R 0.304 0.759 So, since age influences the presence of WMLs the group of patients was subgrouped into three age classes: age 19-39 years, age 40-59 years, and age > 59 years (Table 24). In each age class, there was no significant difference in age distribution between FD with WMLs and FD patients without WMLs. However, there is a big difference between the number of patients in age classes 19–39 and 60-91 years. While in subjects aged 19-39 years the number of patients without WMLs is higher in subjects aged 60-91 years the majority have WMLs. Due to this difference, these two age classes were not included in further study. Focusing only on patients between 40 and 59 years inclusive (age distribution is shown in Figure 21) same tests done earlier were repeated and the results are shown in Table 25 and Table 27. 87 CHAPTER 6. ELECTROCARDIOGRAM Table 24: Number of FD patients with and without WMLs per age classes With WMLs Without WMLs Total p-value [19,40[ 8 25 33 0.411 [40,60[ 23 25 48 0.331 [60,91] 30 3 33 0.683 P-values from Mann Whitney U Test. Figure 21: Boxplot of age stratified for the presence of WMLs in FD patients with ages between 40 and 59 years old, inclusive. Table 25: Demographic characteristics for patients between 40 and 59 years old, inclusive. With WMLs (n = 23) Without WMLs (n = 25) p-value Age (years) 50.61 (1.078) 49.00 (1.134) 0.331 Male (%) 35% 40% 0.552 Data is presented as mean (standard deviation) plus p-values from Mann Whitney U Test. In these subgroups of patients the age and sex differences are not statistically significant (Table 25). From the univariate logistic regression analysis, two ECG features were significantly associated with the presence of WMLs: heart rate variability SDANN5 and SDNN (Table 26). These two features were further analyzed using multivariate logistic regression (adjusted for age or sex), and both remained as 88 CHAPTER 6. ELECTROCARDIOGRAM Table 26: Univariate and multivariate (adjusted for age or sex) logistic regression model to predict the presence of WMLs for patients between 40 and 59 year old, inclusive. Univariate Logistic Regression Adjusted for age Adjusted for sex p-value p-value p-value HR Min 0.080 0.067 0.104 HR Mean 0.220 0.316 0.300 HR Max 0.437 0.721 0.589 ASDNN 5 0.149 0.091 0.190 SDANN 5 0.002 0.001 0.002 SDNN 0.004 0.002 0.004 RMSSD 0.151 0.122 0.163 QT Min 0.154 0.231 0.224 QT Mean 0.649 0.882 0.766 QT Max 0.301 0.166 0.281 QTc Min 0.423 0.380 0.393 QTc Mean 0.524 0.414 0.534 QTc Max 0.247 0.169 0.333 QTc ≥450 0.460 0.300 0.361 Longest R-R 0.329 0.401 0.289 independent predictors of the presence of WMLs in the adjusted models (see Table 27). 89 CHAPTER 6. ELECTROCARDIOGRAM Table 27: Association of heart rate variability SDANN5 and SDNN with WMLs in FD patients age ranged from 40 to 59 years. With WMLs (n = 23) Without WMLs (n = 25) Univariate Logistic Regression Adjusted for age Adjusted for sex p-value p-value p-value SDANN 5 133.18 (7.02) 106.11 (4.61) 0.002 group 0.001 group 0.002 age 0.105 sex 0.41 SDNN 145.43 (7.28) 119.51 (4.93) 0.004 group 0.002 group 0.004 age 0.052 sex 0.621 Data is presented as mean (standard deviation). 6.2 Feature selection with RFE Going further with the analysis of the electrocardiogram, the same feature selection technique used earlier with the gait assessment was employed again to develop a machine learning model that best fits the data. As done earlier, a filter method was performed to select the 10 most significant features (Mann Whitney UTest) not correlated (|ρ|<0.80). Recursive feature elimination algorithm was performed on the 10 remaining features from the filter method: heart rate minimum, mean and maximum, QT minimum, mean and maximum, QT corrected ≥ 450 and minimum, ASDNN 5, and SDANN 5. The heat map of these 10 features selected from the filter method algorithm is shown in Figure 22. Results are stated in Table 28 and Figure 23. LR selected 6 features as the optimal number of gait characteristics with a F1 Score of 64.94%, SVM selected 6 features with a F1 Score of 60.07% and RF selected 6 features with a F1 Score of 68.12%. Table 28 presents the training and validation accuracies for the optimal models of each algorithm. RF presents the higher training and validation accuracies. 90 CHAPTER 6. ELECTROCARDIOGRAM Figure 22: ECG correlation heat map in FD patients. Table 28: Results for each of 3 machine learning algorithms (LR, SVM with linear kernel and RF) with optimal ECG characteristics’ from recursive feature elimination for FD patients with vs without WMLs. Training accuracy% (Mean ±SD) Validation accuracy% (Mean ±SD) Logistic Regression 81.71 ±1.77 79.04 ±3.66 SVM (Linear Kernel) 78.76 ±1.12 77.80 ±2.98 Random Forest 84.75 ±0.09 79.08 ±3.42 SD: standard deviation 91 Chapter 7 Echocardiogram 7.1 Participants The dataset used in this section was the result of an echocardiogram on 93 FD patients (38 males). From this group, 49 of the patients have WMLs (25 males). As observed in the previous group of FD patients (see Section 6.1) the patients with WMLs were significantly older than the patients without WMLs (Table 32). Furthermore, sex difference was found between the two groups of patients (Table 32). By observing the boxplot (Figure 24), we find that the two groups of patients have a different distribution of age. While the age of about half of patients with WMLs ranged from 30 to 50 years, the age of 50% of patients with WMLs ranged from 50 to 70 years. Regarding the echocardiogram data, only two features, LVIDd/SC and A’Lateral, did not present a significant difference between the two groups of patients (Table 33). Table 32: Demographic characteristics for 93 FD patients with and without WMLs. With WMLs (n = 49) Without WMLs (n = 44) p-value Age (years) 57.24 (14.52) 39.02 (14.62) <0.001 Male (%) 51% 30% 0.029 Data is presented as mean (standard deviation) plus p values from Mann Whitney U Test. 98 CHAPTER 7. ECHOCARDIOGRAM Figure 24: Box plot of FD patients with vs without WMLs according to age. From the univariate logistic regression, also the same two features did not present statistical significance with the presence WMLs (Table 34). All the features were then adjusted by age and sex. While when adjusting by sex no changes are noticed, adjusting by age, except for LVIDd, all echocardiogram features become no significant predictor of the presence of WMLs. To decrease this influence of age, the group of patients was subgrouped into the same three age classes as in Section 6.1 (Table 34). There was no significant differences in age between the two group of patients in all three age classes. However, the number of patients in group of patients in age classes 19–39 and 60-91 are very different. Therefore, these two age classes were not included in further study. Focusing only on patients between 40 and 59 years inclusive (age distribution is shown in Figure 25) the univariate logistic regression analysis was performed again. The results are shown in Table 36. Only MV E/A Ratio, MV A Vel and S’ Lateral were significantly associated with the presence of WMLs and just the association of S’ Lateral with the presence of WMLs retained significance in the adjusted models for age. Furthermore, in the three models adjusted for age, the variable age was still a significant predictor of WMLs. 99 CHAPTER 7. ECHOCARDIOGRAM Table 33: Echocardiogram features for 93 FD patients with and without WMLs. With WMLs (n = 49) Without WMLs (n = 44) p-value MV E/A Ratio 1.01 (0.56) 1.48 (0.53) <0.001 MV A Vel 0.81 (0.19) 0.60 (0.18) <0.001 MV Dec T 245.58 (63.34) 202.08 (36.74) <0.001 MV E Vel 0.74 (0.17) 0.82 (0.13) 0.014 E’ Lateral 9.36 (4.32) 14.16 (5.21) <0.001 E’ Septal 7.37 (3.25) 10.70 (3.68) <0.001 E/E’ Lateral 9.27 (3.47) 6.53 (2.66) <0.001 E/E’ Medial 10.56 (3.96) 7.86 (2.98) <0.001 E/E’ Septal 11.72 (5.11) 8.73 (3.51) 0.001 LVPWd 11.19 (3.29) 9.04 (2.36) <0.001 ISVd 12.55 (4.18) 9.82 (3.32) 0.001 LVIDd 42.74 (4.56) 46.07 (3.57) 0.001 LADiam/SC 21.69 (3.55) 19.52 (2.92) 0.001 AoDiam 32.44 (3.69) 30.13 (4.23) 0.006 S’ Lateral 8.47 (2.67) 9.75 (2.82) 0.013 LVdMassInd ASE 108.67 (47.35) 86.16 (37.70) 0.020 LADiam 37.06 (5.76) 34.29 (5.51) 0.024 S’ Septal 7.06 (1.85) 7.80 (1.63) 0.027 A’ Septal 9.24 (2.08) 8.34 (2.11) 0.036 A’ Lateral 10.43 (2.48) 9.48 (2.87) 0.108 LVDdMass ASE 108.07 (82.25) 153.78 (71.95) 0.040 LVIDd/SC 25.11 (2.79) 26.06 (2.42) 0.063 Data is presented as mean (standard deviation) plus significance on T test, Mann Whitney U Test for normally or non-normally distributed continuous variables and Fisher Exact T Test for categorical variables. 100 CHAPTER 7. ECHOCARDIOGRAM Table 34: Univariate and multivariate (adjusted for age) logistic regression model to predict the presence of WMLs. Univariate Logistic Regression Adjusted for age Adjusted for sex p-value p-value p-value MV E/A Ratio <0.001 0.774 <0.001 MV A Vel <0.001 0.146 <0.001 MV Dec T <0.001 0.222 <0.001 MV E Vel 0.015 0.745 0.022 E’ Lateral <0.001 0.879 <0.001 E’ Septal <0.001 0.711 <0.001 E/E’ Lateral <0.001 0.663 <0.001 E/E’ Medial <0.001 0.868 0.001 E/E’ Septal 0.02 0.711 0.004 LVPWd 0.001 0.787 0.001 ISVd 0.001 0.861 0.003 LVIDd <0.001 0.016 <0.001 LADiam/SC 0.002 0.926 0.001 AoDiam 0.006 0.857 0.017 S’ Lateral 0.027 0.270 0.044 LVdMasInd ASE 0.014 0.366 0.033 LADiam 0.021 0.819 0.036 S’ Septal 0.044 0.293 0.081 A’ Septal 0.040 0.296 0.054 A’ Lateral 0.090 0.810 0.130 LVDdMass ASE 0.048 0.299 0.120 LVIDd/SC 0.084 0.241 0.113 Table 35: Division of patients according to respective ages plus age significance value for each age interval performed with Mann Whitney U Test. With WMLs Without WMLs Total p-value [17,40[ 5 22 27 0.416 [40,60[ 24 19 43 0.082 [60,91[ 20 3 23 0.404 101 CHAPTER 7. ECHOCARDIOGRAM Figure 25: Box plot of patients with vs without WMLs in the interval of ages between 40 and 59 years, inclusive. Table 36: Echocardiogram features presented as mean ±standard deviation, statistically significant after performing univariate logistic regression and adjusted logistic regression for age to predict the presence of WMLs for patients between 40 and 59 years, inclusive. With WMLs Without WMLs Univariate Logistic Regression Adjusted for age Adjusted for sex MV E\A Ratio 1.02 ±0.33 1.31 ±0.42 0.017 group 0.068 group 0.003 age 0.013 sex 0.091 MV A Vel 0.79 ±0.18 0.68 ±0.17 0.04 group 0.142 group <0.001 age 0.039 sex 0.219 S’ Lateral 9.25 ±2.88 8.42 ±2.35 0.027 group 0.039 group <0.001 age <0.001 sex 0.309 102 CHAPTER 7. ECHOCARDIOGRAM 7.2 Discussion In agreement with the results presented in the previous chapter, age was found to be significantly associated with the presence of WMLs. In the group of patients analyzed in this chapter, the patients with WMLs were 18 years older than the patients without WMLs on average. As mentioned before, with age, people tend to have their heart functions decreased (Strait & Lakatta, 2012) which also reflects on different values on the echocardiogram features. The results (Table 33) show that the difference in echocardiogram features between the two FD patients groups with WMLs and without WMLs are due to the age difference. Furthermore, after age stratification, in the group aged from 40 to 59 years, the results (Table 36) reveal that the values of the echocardiogram features are still affected by age. From this first analysis, the echocardiogram features not reveal to be good predictors of the presence of WMLs in FD patients. Then, it is not expected to get higher classification performance with the echocardiogram features comparing with the performance achieved previously with gait and ECG data. However, further research is needed to confirm this hypothesis. But, as the echocardiogram and ECG data are from two different groups of FD patients, this analysis is left for future work. 103 Chapter 8 Conclusion, limitations, and future work 8.1 General conclusions The major goals of this dissertation were to evaluate the effectiveness of machine learning methods for the differential diagnosis of WMLs in FD patients based on a dataset from one of three different sources: gait evaluation, electrocardiogram exam (Holter), or echocardiogram exam, and, to investigate how the integration of information from more than one source can improve the predictive system. To achieve these aims, several steps were performed including the evaluation of a MR normalization strategy for the de-correlation of physical properties, speed, stride length, and gait variables, the implementation of various feature selection methods using hybrid methods, the implementation and evaluation of different classification algorithms based on different datasets. To the best of our knowledge, this was the first study that explores gait characteristics and cardiac data and their discriminate power in FD patients with WMLs from FD patients without WMLs. Using gait data, for the discrimination of FD patients with WMLs from age-matched healthy adults the model with higher accuracy (72%) was achieved with RF classifier with the variables foot flat, pushing, and cycle duration as input. For the discrimination of FD patients without WMLs from age-matched healthy adults, the best-suited model was achieved using KNN based on the variables loading, foot flat, stride length variability, loading variability, and lift-off angle variability, with an accuracy of 87%. Finally, to 104 CHAPTER 8. CONCLUSION, LIMITATIONS, AND FUTURE WORK discriminate the presence of WMLs in FD patients using gait characteristics, the models that fit best were using CNN based on the time series of the variables minimum toe clearance, swing, and peak swing, with an accuracy of 82%, and using LR algorithm based on 5 gait measures lift-off angle variability, swing, minimum toe clearance variability, peak swing, and foot flat variability, with an accuracy of 81%. When using ECG features for the differential diagnosis of the presence of WMLs in FD patients the best-suited model was using RF with QT maximum, heart rate mean and maximum, and SDANN 5, with an accuracy of 80%. To analyze if the use of both gait and ECG features as input could improve the identification of the presence of WMLs, the best models described above: LR algorithm based on gait measures and RF algorithm based on ECG features were evaluated on a test group of nine patients. Six patients were well classified when the LR algorithm was performed based on gait or based on gait+ECG. However, with the RF algorithm based on ECG features 6 patients were well classification and this number increase to 7 when gait+ECG was used as input. This result shows some evidence that the use of both datasets could improve the performance of identification of WMLs in FD patients. The results from logistic regression analysis revealed that age was independently associated with the presence of WMLs, and the echocardiogram features that revealed some significant association with the presence of WMLs showed to be affected by age. Then, it is not expected to get higher classification performance with the echocardiogram features comparing with the performance achieved previously with gait and ECG data. However, further research is needed to confirm this hypothesis. 8.2 Limitations Several limitations were found across this study. Firstly, the developed MR models were based on a fairly small number of control subjects (n= 37), still comparable to other previous works (Mikos et al., 2018; Wahid et al., 2016; Wahid et al., 2015). A larger amount of subjects along with more independent variables might improve the robustness of the feature selection methods by reducing the correlation among variables and, therefore, improve the effectiveness of the classifiers. To validate the best set of gait 105 CHAPTER 8. CONCLUSION, LIMITATIONS, AND FUTURE WORK variables found by the feature selection methods further research with a higher number of subjects is required. For last, gait, ECG, and echocardiogram datasets were not collected for all the same patients, which made it impossible to develop a classification model with the three datasets as input. 8.3 Future work The findings reported in this thesis are the first step to demonstrate the potential of machine learning techniques based on gait and cardiac variables as a complementary tool to understand the role of WMLs in the gait impairment of FD. For future research, a larger sample size will be used to confirm and extend these findings. The implications of WMLs on gait compromise in FD or predictive value of each kinematic gait variable remain still elusive, warranting further investigation with a more enriched cohort. Due to the number of subjects involved in this study, in most of the cases, all subjects were used in the training and validation of the models, and only in Section 6.4, a small independent group test was used for checking the model performances. To test the robustness of classification models based on the selected gait and cardiac characteristics further research with a larger independent group test is needed. Since, in this study, the classification of FD was only evaluated based on gait, to go further with this study of the presence or absence of FD, is required to analyze: FD patients vs controls with ECG data, FD patients vs controls with echocardiogram data, and, finally, FD patients vs controls with gait + ECG + echocardiogram. 106 Part IV : Appendix 107 BIBLIOGRAPHY Koushik, J. (2016). Understanding Convolutional Neural Networks. Kramer, O. (2013). Dimensionality Reduction with Unsupervised Nearest Neighbors. Intelligent Systems Reference Library,51. https://doi.org/10.1007/978-3-642-38652-7 Kubota, K. J., Chen, J. A., & Little, M. A. (2016). Machine learning for large-scale wearable sensor data in Parkinson’s disease: Concepts, promises, pitfalls, and futures. Movement Disorders,31(9), 1314–1326. https://doi.org/10.1002/mds.26693 LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., & Jackel, L. D. (1989). Backpropagation Applied to Handwritten Zip Code Recognition. Neural Computation,1(4), 541–551. https://doi.org/10.1162/neco.1989.1.4.541 Löhle, M., Hughes, D., Milligan, A., Richfield, L., Reichmann, H., Mehta, A., & Schapira, A. H. (2015). Clinical prodromes of neurodegeneration in Anderson-Fabry disease. Neurology,84(14), 1454– 1464. https://doi.org/10.1212/WNL.0000000000001450 MacDermot, KD and Holmes, A., & Miners. (2001). Anderson-Fabry disease: clinical manifestations and impact of disease in a cohort of 98 hemizygous males. Journal of medical genetics,38, 769–775. Maity, N. G., & Das, S. (2017). Machine learning for improved diagnosis and prognosis in healthcare. IEEE Aerospace Conference Proceedings. https://doi.org/10.1109/AERO.2017.7943950 Mannini, A., Trojaniello, D., Cereatti, A., & Sabatini, A. (2016). A Machine Learning Framework for Gait Classification Using Inertial Sensors: Application to Elderly, Post-Stroke and Huntington’s Disease Patients. Sensors,16(1), 134. https://doi.org/10.3390/s16010134 Mauer, M., & Kopp, J. (2018). Fabry disease: Treatment. Mikos, V., Yen, S. C., Tay, A., Heng, C. H., Chung, C. L. H., Liew, S. H. X., Tan, D. M. L., & Au, W. L. (2018). Regression analysis of gait parameters and mobility measures in a healthy cohort for subject-specific normative values. PLoS ONE,13(6), e0199215. https://doi.org/10.1371/ journal.pone.0199215 Morgan, S. H., Cheshire, J. K., Wilson, T. M., MacDermot, K., & d.A. Crawfurd, M. (1987). Anderson-Fabry disease-family linkage studies using two polymorphic X-linked DNA probes. Pediatric Nephrology,1(3), 536–539. https://doi.org/10.1007/BF00849266 114 BIBLIOGRAPHY Myszczynska, M. A., Ojamies, P. N., Lacoste, A. M., Neil, D., Saffari, A., Mead, R., Hautbergue, G. M., Holbrook, J. D., & Ferraiuolo, L. (2020). Applications of machine learning to diagnosis and treatment of neurodegenerative diseases. https://doi.org/10.1038/s41582-020-0377-8 Papernot, N., Faghri, F., Carlini, N., Goodfellow, I., Feinman, R., Kurakin, A., Xie, C., Sharma, Y., Brown, T., Roy, A., Matyasko, A., Behzadan, V., Hambardzumyan, K., Zhang, Z., Juang, Y.-L., Li, Z., Sheatsley, R., Garg, A., Uesato, J., … McDaniel, P. (2016). Technical Report on the CleverHans v2.1.0 Adversarial Examples Library. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine Learning in Python Gaël Varoquaux Bertrand Thirion Vincent Dubourg Alexandre Passos PEDREGOSA, VAROQUAUX, GRAMFORT ET AL. Matthieu Perrot (tech. rep.). Pradhan, C., Wuehr, M., Akrami, F., Neuhaeusser, M., Huth, S., Brandt, T., Jahn, K., & Schniepp, R. (2015). Automated classification of neurological disorders of gait using spatio-temporal gait parameters. Journal of Electromyography and Kinesiology,25(2), 413–422. https://doi. org/10.1016/j.jelekin.2015.01.004 Prakash, C., Kumar, R., & Mittal, N. (2018). Recent developments in human gait research: parameters, approaches, applications, machine learning techniques, datasets and challenges. Artificial Intelligence Review,49(1), 1–40. https://doi.org/10.1007/s10462-016-9514-6 Prashant Gupta. (2017). Decision Trees in Machine Learning – Towards Data Science. Ramaswami, U., Whybra, C., Parini, R., Pintos-Morell, G., Mehta, A., Sunder-Plassmann, G., Widmer, U., & Beck, M. (2006). Clinical manifestations of Fabry disease in children: Data from the Fabry Outcome Survey. Acta Paediatrica, International Journal of Paediatrics,95(1), 86–92. https://doi.org/10.1080/08035250500275022 Rehman, R. Z. U., Del Din, S., Guan, Y., Yarnall, A. J., Shi, J. Q., & Rochester, L. (2019). Selecting Clinically Relevant Gait Characteristics for Classification of Early Parkinson’s Disease: A Comprehensive Machine Learning Approach. Scientific Reports,9(1), 1–12. https:// doi.org/ 10.1038 / s41598-019-53656-7 115 BIBLIOGRAPHY Rijnbeek, P. R., Van Herpen, G., Bots, M. L., Man, S., Verweij, N., Hofman, A., Hillege, H., Numans, M. E., Swenne, C. A., Witteman, J. C., & Kors, J. A. (2014). Normal values of the electrocardiogram for ages 16-90 years. Journal of Electrocardiology,47(6), 914–921. https://doi.org/10.1016/ j.jelectrocard.2014.07.022 Rost, N. S., Cloonan, L., Kanakis, A. S., Fitzpatrick, K. M., Azzariti, D. R., Clarke, V., Lourenco, C. M., Germain, D. P., Politei, J. M., Homola, G. A., Sommer, C., Üçeyler, N., & Sims, K. B. (2016). Determinants of white matter hyperintensity burden in patients with Fabry disease. Neurology, 86(20), 1880–1886. https://doi.org/10.1212/WNL.0000000000002673 Satriano, A., Afzal, Y., Sarim Afzal, M., Fatehi Hassanabad, A., Wu, C., Dykstra, S., Flewitt, J., Feuchter, P., Sandonato, R., Heydari, B., Merchant, N., Howarth, A. G., Lydell, C. P., Khan, A., Fine, N. M., Greiner, R., & White, J. A. (2020). Neural-Network-Based Diagnosis Using 3-Dimensional Myocardial Architecture and Deformation: Demonstration for the Differentiation of Hypertrophic Cardiomyopathy. Frontiers in Cardiovascular Medicine,7, 584727. https://doi.org/10.3389/fcvm. 2020.584727 Snir, J. A., Bartha, R., & Montero-Odasso, M. (2019). White matter integrity is associated with gait impairment and falls in mild cognitive impairment. Results from the gait and brain study. NeuroImage: Clinical,24. https://doi.org/10.1016/j.nicl.2019.101975 Starr, J. M. (2003). Brain white matter lesions detected by magnetic resosnance imaging are associated with balance and gait speed. Journal of Neurology, Neurosurgery & Psychiatry,74(1), 94–98. https://doi.org/10.1136/jnnp.74.1.94 Strait, J. B., & Lakatta, E. G. (2012). Aging-Associated Cardiovascular Changes and Their Relationship to Heart Failure. https://doi.org/10.1016/j.hfc.2011.08.011 Sunder-Plassmann, G., & Födinger, M. (2006). Diagnosis of Fabry disease: the role of screening and case-finding studies. Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. Advances in Neural Information Processing Systems,4(January), 3104–3112. Thompson, C. G., Kim, R. S., Aloe, A. M., & Becker, B. J. (2017). Extracting the Variance In flation Factor and Other Multicollinearity Diagnostics from Typical Regression Results. Basic and Applied Social Psychology,39(2), 81–90. https://doi.org/10.1080/01973533.2016.1277529 116 BIBLIOGRAPHY Tsai, C. H., Ma, H. P., Lin, Y. T., Hung, C. S., Hsieh, M. C., Chang, T. Y., Kuo, P. H., Lin, C., Lo, M. T., Hsu, H. H., Peng, C. K., & Lin, Y. H. (2019). Heart Rhythm Complexity Impairment in Patients with Pulmonary Hypertension. Scientific Reports,9(1). https://doi.org/10.1038/s41598019-47144-1 Umetani, K., Singer, D. H., McCraty, R., & Atkinson, M. (1998). Twenty-four hour time domain heart rate variability and heart rate: Relations to age and gender over nine decades. Journal of the American College of Cardiology,31(3), 593–601. https:// doi.org/ 10.1016/ S0735 - 1097(97)00554-8 Wahid, F., Begg, R., Lythgo, N., Hass, C. J., Halgamuge, S., & Ackland, D. C. (2016). A multiple regression approach to normalization of spatiotemporal gait features. Journal of Applied Biomechanics, 32(2), 128–139. https://doi.org/10.1123/jab.2015-0035 Wahid, F., Begg, R. K., Hass, C. J., Halgamuge, S., & Ackland, D. C. (2015). Classification of Parkinson’s disease gait using spatial-temporal gait features. IEEE Journal of Biomedical and Health Informatics,19(6), 1794–1802. https://doi.org/10.1109/JBHI.2015.2450232 Zheng, J. J., Lord, S. R., Close, J. C., Sachdev, P. S., Wen, W., Brodaty, H., & Delbaere, K. (2012). Brain white matter hyperintensities, executive dysfunction, instability, and falls in older people: A prospective cohort study. https://doi.org/10.1093/gerona/gls063 117