scieee AI-readable full text Open interactive document viewer

Depression Signs Detection through Smartphone Usage Data Analysis

Nino Rafael da Silva Rocha

Full text

Depression Signs Detection through Smartphone Usage Data Analysis Nino Rafael da Silva Rocha Mestrado integrado em Engenharia de Redes e Sistemas Informáticos Departamento de Ciência de Computadores 2014 Orientador Ana Vasconcelos (MSc), Fraunhofer Portugal Coorientador Rita P. Ribeiro (PhD), DCC - FCUP Todas as correções determinadas pelo júri, e só essas, foram efetuadas. O Presidente do Júri, Porto, ______/______/_________ Acknowledgements First of all, i would like to thank to my supervisors, Rita P. Ribeiro (PhD, DCC-FCUP) and Ana Vasconcelos (MSc, Fraunhofer Portugal) for all their kindness and dedication. To Dr. Ana Nobre, for the time spent with guidance in the field of psychology. To Fraunhofer Portugal for giving me the opportunity to develop this project. Next, thanks to my dear family, specially my father, mother, and sister, for believing in my skills, and for my education. All I am is because of them. To my dear friends, specially to André Francisco, Ana Raquel Azevedo, André Gaspar, Filipe Martins, Geno Pereira, Hugo Sousa, Inês Caeiro, José Pedro, and Tiago Melo. A special thanks to my best friend who always believes in me, God. 3 Resumo A depressão é uma das doenças mentais mais comuns entre a população geral. Os sintomas de depressão são diversos: insónias, perda de energia, tristeza e isolamento. A depressão torna-se mais comum em estágios tardios da vida, alimentada principalmente pela perda de familiares e cessação de actividade. Este trabalho propõem um estudo que utiliza informação colectada a partir de smartphones, como a localização, ou estatísticas de comunicação, entre outros, e conduz uma análise estatística e de machine learning para inferir conclusões sobre o estado depressivo de idosos. A análise estatística provou ser útil para determinar conclusões baseadas nos vários campos, enquanto que o machine learning deriva a suas conclusões a partir de hábitos populacionais. Um serviço web foi construído sobre esta implementação como uma ferramenta de visualização de dados. Este projeto é ainda um estudo preliminar que outros poderão utilizar como base para outros trabalhos. 4 Abstract Depression is one of the most common mental disorders among the general population. Depression symptoms are diverse: sleep disorder, energy loss, sadness, or isolation. Depression is most common in late stages of life, mostly fueled by the loss of relatives and retirement. This work proposes a study that uses data collected from smartphones, such as location sensors, communication statistics, among others, and uses statistical analysis and machine learning to infer conclusions on elders’ depression symptoms. The statistical analysis proved useful for determining field-wise conclusions from personal habits, while the machine learning process derives its conclusions from general population standards. A web service was constructed on top of this implementation as an useful data visualization tool. This is a preliminary study which might serve as ground basis for others to build upon. 5 Contents Acknowledgements 3 Resumo 4 Abstract 5 List of Figures 9 1 Introduction 12 1.1 Goals ...................................... 13 1.2 Thesis Structure ................................ 13 2 Seniors and Depression in Old-age 14 2.1 Age-related Changes ............................. 14 2.2 Depression ................................... 16 2.2.1 Aging and Depression ......................... 17 2.3 Successful Aging ................................ 19 3 Knowledge Discovery and Data Mining 23 3.1 Knowledge-discovery ............................. 23 3.1.1 Data Mining ............................... 24 3.1.2 Machine Learning ........................... 25 3.2 Data Mining Process: Cross-Industry Standard Process .......... 27 3.2.1 Phase 1: Business Understanding .................. 28 3.2.2 Phase 2: Data Understanding ..................... 28 3.2.3 Phase 3: Data Preparation ...................... 29 6 3.2.4 Phase 4: Modeling ........................... 29 3.2.5 Phase 5: Evaluation .......................... 29 3.2.6 Phase 6: Deployment ......................... 30 4 Market analysis and Existing Solutions 31 5 A CRISP-DM approach to DepSigns 34 5.1 Business Understanding ............................ 34 5.2 Data Understanding .............................. 35 5.2.1 Data Generation ............................ 35 5.2.2 Database Model ............................ 40 5.2.3 Statistical Analysis ........................... 43 5.2.4 Central tendency measures ...................... 43 5.3 Data Preparation ................................ 47 5.4 Modeling .................................... 52 5.5 Evaluation .................................... 55 5.6 Deployment ................................... 56 6 Technical Aspects and Implementation of DepSigns 57 6.1 Used Technologies ............................... 57 6.1.1 Scikit Learn ............................... 57 6.1.2 Web Services .............................. 58 6.1.3 Model-View-Controller ......................... 58 6.1.4 Database ................................ 59 6.1.5 Backend ................................. 59 6.1.6 Frontend ................................ 60 6.1.7 Chart Rendering ............................ 61 6.2 System Architecture .............................. 61 6.2.1 Structure ................................ 62 6.2.2 Database generation .......................... 63 6.2.3 Authentication and Authorization ................... 67 6.2.4 Views .................................. 68 7 6.2.4.1 Main Page .......................... 69 6.2.4.2 Logout ............................ 71 6.2.4.3 Feedback ........................... 72 6.2.4.4 Patient ............................ 72 6.2.4.5 Analysis ........................... 74 6.2.4.6 Patient List .......................... 74 6.2.5 Projection ................................ 75 6.2.6 Alerts .................................. 76 6.2.7 Decision Tree Predictions ....................... 78 7 Conclusion 81 7.1 Summary .................................... 82 7.2 Challenges ................................... 83 7.3 Future Work ................................... 83 A Acronyms 86 B Source code 88 References 90 8 List of Figures 3.1 Data Mining Taxonomy. ............................ 25 3.2 Process diagram showing the relationship between the different phases of CRISP-DM (reprinted from [54]). ...................... 28 5.1 Probability density function (reprinted from [3]). ............... 37 5.2 The database EER model ........................... 42 5.3 Sample projection of the data and statistical analysis. ........... 44 5.4 Example of an individual’s weekly time of activity, as well as its corresponding average ............................... 45 5.5 An example of a Learned Decision Tree, with leafs shown in yellow (reprinted from [52]). .................................... 53 5.6 Gini rule example(reprinted from [14]) .................... 54 5.7 Visualization of a Learned Decision Tree. .................. 55 6.1 DepSigns System Architecture ........................ 62 6.2 Login page ................................... 70 6.3 Registration form ................................ 70 6.4 Psychologist’s main page. ........................... 71 6.5 Caregiver’s main page. ............................. 71 6.6 Logout option. ................................. 72 6.7 Feedback psychologist option. ........................ 72 6.8 The patient page. ................................ 73 6.9 Visualization of the Learned Decision Tree. ................. 74 6.10 Page with the list of patients, and a search bar. ............... 75 6.11 Example of field-wise alert ........................... 77 9 CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 16 level, on close social relations. The negative events and particularly the death of important people in a close circle are some of the most frequent situations at this age, which leads almost inevitably to the realization that one "is old". There are other great challenges on a social level in the stage of old-age, such as the fear of eventual isolation and the integration in some institution, or the parting of one’s descendants from home. The role and status of an individual in society may also be a subject of many important changes for seniors - retirement and widowhood are two of the situations in which this matter is most evident. In the specific matter of retirement, seniors have the impression that one’s contribution on a working level is over, resulting in more free time which needs to be occupied somehow. In addition, elders go through countless changes in different levels, mostly as consequence of physiological, psychosocial and perceptual changes. Side by side with the decline seniors face on several levels, there is still an increase of social and familiar and even personal pressure on the senior to be able to deal with these factors, adjusting one-self to one’s environment and find a possible and desired balance [39]. 2.2 Depression Depression is one of the most common mental disorders among the general population and can be manifested from childhood to old age. This is a serious mental health problem that is even considered as the leading cause of disability related to illnesses and health problems. According to data from some studies [11], one in four people in the world suffers or will suffer from Depression. In Portugal, one in five Portuguese is depressed and 1200 deaths are a direct result from this illness [23]. This mental illness is essentially characterized by depressive symptoms that can be episodic, recurrent or chronic such as low mood, reduced energy and loss of interest or enjoyment. When these symptoms are detected, they are also usually associated with inactivity, physical pain, poor concentration, reduced self-confidence, pessimism, CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 17 disturbed sleep and altered appetite [7] leading a person to a substantial decrease in their ability to comply with their daily responsibilities. The causes associated with the manifestation of mental illness differ from person to person. In most cases, depressive episodes result from a variety of biological, psychosocial and family factors that increase the risk of a person developing a depressive disorder. Biological factors that result in increased risk of depression include: the females has predisposition to the illness, especially in adolescence;people who suffer from physical pain without explanation; chronic diseases such as hypertension, history of thrombosis, asthma, diabetes; or other diseases such as AIDS and Cancer. People with substance abuse such as alcohol or drugs are also prone to developing depressive symptoms [55]. At a family level, individuals who have a family history of depression, those who live with a family member carrying a serious or chronic disease and significant loss of a close relative are among the main factors for depression. The psychological and social factors have a significant role in the development of depression mainly on the individuals with stress-generating professions, people prone to anxiety and/or panic, loss of employment and loss of social relationships in their midst. 2.2.1 Aging and Depression Defined by the American Psychiatric Association’s in Diagnostics and Statistical Manual (DMS-IV), Depresion in the Elderly is the existence of a depressive syndrome in individuals with over 65 years of age [32]. Since the worldwide trend is that of an increasingly aged population, the importance of diagnosis and treatment of depression in this age group must be highlighted. According to the World Health Organization, in 2025 there will be 1.2 billion people over 60 years of age and, in Portugal, the National Statistics Institute predicts a considerable increase in population percentage over 65 years of age, rising from 17,6% in 2008 to 32.3% in 2060 [12]. CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 18 According to Melo and Ferruzzi in [34], Zimmerman (2000) stated that the probability of seniors suffering from this disease is greater than during youth or adulthood and this increase is justified by numerous losses and limitations associated with old age having consequences like low self-esteem. There are many risk factors associated with the onset of depression at this stage of life. At the psychosocial level, one of the most significant events is the end of an occupation. This can be negatively faced by the seniors since it leads to change routines and possible feelings of uselessness (FernandezBalesteros and Izal,1993 in [51]). The loss of loved ones and people in their social environment and the sometimes remoteness of family leads the seniors into a new cycle of life which can cause feelings of discouragement and loneliness (Figueiredo,2007 in [51]). Biological factors also pose a high risk of increased probability of depressive disorders. According to Nunes (2008) in [51] the cognitive decline associated with memory loss is one of the main complaints of people with depression. Associated with this are also functional insufficiency and physical illness that prevent seniors from participating in their usual daily activities. These factors combined can trigger feelings of inadequacy and social demotivation that are so often linked to Depression. According to Fernandes in [26], Marques and Col(1989) summarize the factors of depression in seniors in three major areas: environmental determinants, particularly the isolation and the lack of social interaction and job, death of the spouse, and social and occupational devaluation. The genetics area considers seniors as a group with a genetic predisposition to depression. Finally the organic area refers to the wide variety of organic illnesses that may present symptoms of depressive nature. It should also be noted the importance of concomitant diseases and their respective medications that may results in side effects such as depression. The branch of Medicine that focuses on the study of depressive symptoms in seniors is Geriatrics, aiming at the prevention and treatment of diseases at late stages of life. The monitoring and diagnosis are always done under the supervision of a doctor, but can be assessed through various evaluative scales available for patients. The most common scales cited in the literature are the Geriatric Depression Scale (GDS) and Center for Epidemiologic Studies Depression(CES-D). The GDS is a self-assessment questionnaire consisting of 30 questions with which the geriatric population can assess depressive symptoms. This same scale is used by professionals in monitoring and CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 19 comparing the depressed state of the patient [37]. The CES-D has a smaller version of 10 items, characterized by a high specificity and sensitivity in the diagnosis of major depression in the hospital context [25]. Other scales used for screenings are the Hamilton Depression Scale(HAM-D) and the Patient Health Questionnaire-9(PHQ-9). The HAM-D is a standard method used in clinical studies and is not specific to the geriatric population and includes some items related to somatic pain so it may cause some confusion with chronic diseases that may arise [37]. The PHQ-9 is a 9-item questionnaire derived from the DSM-IV for major depressive disorders and it is a instrument that aims to firstly, measure the severity of the depression and on the other hand is an auxiliary diagnostic tool for major depression [37]. 2.3 Successful Aging The concept of Successful Ageing was born in the 60’s and can be defined as a set of mechanisms of adaptation to the specific conditions of old age, looking to establish a balance between the capacity of the individual and the demands of the environment [20]. To study this concept, there is a need to overcome the stereotype commonly associated with this phase of life that this is a period in which the occurrence of disease is more frequent, the psychological and physical abilities decline and that there are obvious disabilities. To sustain a positive view of old age and the aging process, several studies were developed by The McArthur Foundation, such as "Study of McArthur Foundation"(1984) which is summarized in the book "Successful Ageing"(Rowe and Kahn,1998), which got a huge impact on the scientific community specialized in this area [50]. Rowe and Kahn(1998) argue that successful ageing, is based on "several factors that allow individuals to continue to work effectively, physically and mentally in old age [20]." However, there is no specific pattern of successful ageing as this is a complex construct. Baltes and Cartensen (1996) in [20] state that there is no theory or standard criteria to defining success in old age but there are two related processes. On the one hand, the CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 20 ability to adapt to losses that occur in old age, and on the other hand, the selection of certain lifestyles in order to maintain a good physical and mental activity. According to this authors, we can say that there is not a single path to reach the success in ageing, but a variety of factors that have a huge importance. Margoshes (1995) states that seniors should make the management of the available time in a more conscious and balanced way, transforming their lifestyle. In this conscious and balanced way are included a positive mental attitude, continuous challenges, cognitive exercises and preservation of healthy habits, thus ensuring the success of aging [20]. To Rowe and Kahn (1998), the low risk and disability related to diseases, mental and physical functioning/active engagement with life are the 3 components able to provide successful aging, concluding that successful ageing is dependent on choices and on the social behaviors that can be obtained through individual effort [20]. All theories of successful ageing see the individuals as pro-actives and able to regulate their quality of life by setting goals and work to achieve them. For seniors to maintain a successful ageing, they must adjust to age-related changes and involve themselves actively to preserve their well-being. Active ageing or successful ageing, corresponds to the adoption of appropriate strategies to deal with the inherent challenge of the ageing process. Considering that there is no single path to grow old, it is inevitable to state the absence of a similar way to all people, because the pathways of aging are all different and it is possible to reach satisfaction and success in life through different routes. Noting the complexity to outline a pattern in this personal success in aging, it is of most importance to evaluate several criteria: The Competence reveals itself of high importance because it is possible to predict the psychological state of individuals through this criterion. Each individual adapts in a dynamic way to the biological aging and all changes in social network and may say that aging is successful as greater is the adaption to changes. Paúl (2001), affirms that it is necessary to pay attention to the bio-psychosocial and behavioral complexity of seniors, properly valuing individual responses which are most appropriate to counter hypothetical losses of competence that jeopardize the autonomy of the subject. The integrity of autonomy in elderly life is implicit for successful aging, it is extremely necessary and required to have physical exercise and a constant existence of social relationships [20]. CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 21 Health Promotion is one of the aspects that have influence on successful ageing. According to Rowe and Kahn (1998), there should be prevention in adulthood emphasizing the influence of healthy lifestyle and health self-rated and how they reveal themselves in general well-being during aging [42]. Cognitive Activity is another criterion present to obtain a good aging. The adoption of compensatory measures to cope with an expected unfavorable evolution of biological variables such as sensory loss and decreased speed of information processing, emerges as an essential combat factor to the fatalist view that old age corresponds to the loss of capacity to understanding and learning, highlighting as preventive measures the physical exercise and cognitive training. Finally is it noted the importance of adopting strategy selections, compensation and optimization in relation to cognitive decline. Baltes and Cartensen (1999) in [20], describe this strategy in a distribution of available cognitive resources for the needs and goals to which the senior assigns more importance. Psychological Well-being is one of the central aspects of aging successfully. The competence, socioeconomic status and social integration emerge as the most important factors in measuring satisfaction and well-being. On the other hand, losses in areas such as retirement, widowhood and health issues, result in a negative impact on life satisfaction and psychological well-being [30]. The Context of Residence plays an equally important role to understanding the different patterns of aging and explain why some people achieve successful aging. This satisfaction is explained by the theory of person-environment fit (Kahana and Lawton,1992) in [20], which provides a satisfactory adaptation when in the transactions between the person and the environment, the individual characteristics of a person are congruent with the demands of the environment. The notion of aging-in-place is central to understanding the relationship between the context of residence and successful aging. Therefore, it is necessary to provide opportunities for older people to get a relationship with other people and find someone that they can trust, this being the best antidote against loneliness. The Quality of Life, is one of the criteria of the success formula at this age. The commonly indicators used to assess this criterion are focused on subjective well-being CHAPTER 2. SENIORS AND DEPRESSION IN OLD-AGE 22 (physical, material, social and emotional). Autonomy, activity, material indexes, economic resources, health, living conditions, intimacy, safety, place in community and personal relationships are also key indicators in assessing the quality of life of an elder. Fernandez-Ballesteros (1998) in [17], gives an insight into the measurement of quality of life by the context and circumstances in which the elderly lives, such as social status, age or gender, although corroborating the position of several authors, stating that the generalized standard patterns of quality of life can not be established. In sum, it seems generally agreeable that seniors quality of life is directly related to biological, psychological, social and behavioral factors, and only through mediation of all these factors we can think in standards patterns of quality of life for seniors. Chapter3 Knowledge Discovery and Data Mining Nowadays companies and institutions from various areas, such as science, business, or health care around the world have a huge amount of data stored with varied information, where the increase in the degree of complexity in their structures have a exponential tendency. The emergence of differentiated technologies for storage and retrieval of data, allow the user or analyst to obtain strategic information that can be transformed into useful knowledge from a collection of a logical structured data. 3.1 Knowledge-discovery Worldwide companies and organizations have big databases aggregated to their computer systems. The analysis of such voluminous amount of data is not humanly possible without the assistance of computational tools. In this context, it is important to understand all the terms and the differences between Knowledge Discovery (KD), Knowledge Discovery in Databases (KDD), Knowledge Discovery and Data Mining (KDDM) and the specific step of Data Mining (DM). Knowledge Discovery (KD) is a process through wich new knowledge or important information about the study domain is acquired. It involves many steps and each step 23 CHAPTER 3. KNOWLEDGE DISCOVERY AND DATA MINING 24 consists of a particular discovery task [28]. Knowledge Discovery in Databases (KDD) regards the application of the KD process to databases [28]. Fayyad et al. [16] defines it further as the non-trivial process of identifying valid, novel, potentially useful and ultimately understandable patterns in data. These result from the intersection of several areas, such as Machine Learning, Pattern Recognition, Databases, Statistics, Artificial Intelligence, Knowledge Acquisition for Expert Systems, Data Visualization and High-performance Computing. The Knowledge Discovery and Data Mining (KDDM) process regards to the KD process applied to any data source. It is considered the entire knowledge extraction process, which, according to [28], addresses issues like data storage and access, algorithms efficiency, results interpretation and visualization and, human-machine interaction. Data Mining and Machine Learning are part of all these processes. In the following subsections, we contextualize them in the knowledge discovery process. 3.1.1 Data Mining Data-Mining is a particular step of the KDD or KDDM processes. It consists on the application of specific algorithms and a variety of analysis tools to extract patterns and relationships from data that can be used to make valid predictions. It component of KDD currently strongly depends on known techniques such as Machine Learning, Pattern Recognition and Statistics to find patterns in data. However, it has become better known than the KDD and KDDM process themselves mostly due to it being the step where the search of knowledge techniques are applied. According to [28] and depending on the goal this knowledge discovery is distinguished into two types: Verification and Discovery.Verification only verifies the user’s hypothesis. Discovery is related to automatically finding new patterns. It can be subdivided into Prediction, where the found patterns are used to predict the future behavior of some entity; and Description, where the found patterns describe some entity in a way which is comprehensible to humans. The algorithms or methods in this task are supervised by an external source CHAPTER 3. KNOWLEDGE DISCOVERY AND DATA MINING 25 who knows the associated output values for each set of input attributes. According to the target variable type, numeric or categorical, we have Regression or Classification techniques. The algorithms used in Description assume the inexistence of a target variable, following here the system find patterns for presentation to a user in a way which is comprehensible to humans. The algorithms used in the this task ignore the output attribute following the unsupervised learning paradigm where is the Clustering and among others techniques. Figure 3.1 shows the taxonomy we have just described. However, we must stress that this a very elementary and over-simplistic view of Data Mining. Data Mining DIscovery Verification Supervised Learning Unsupervisioned Learning Classification Regression Clustering Figure 3.1: Data Mining Taxonomy. 3.1.2 Machine Learning Along with research fields, such as statistics and pattern recognition, Machine Learning is one of the research field on which Data Mining relies on. It concerns the design and development of learning algorithms that allow computers to automatically learn patterns CHAPTER 4. MARKET ANALYSIS AND EXISTING SOLUTIONS 32 Mobilyze application was developed by researchers at Northwestern University (USA) and the main goal is monitoring behavioral patterns and mood states by identifying states that trigger depression, thus preventing it. The strategy consists of a combination between the sensors of personal phone (such as GPS, accelerometer, and Wi-Fi), with the data provided by users (such as the state of mind and social context). The main advantage of this application focuses on the possibility of anticipating depressive moments of the individual being monitored. Still, there are some barriers that have not been exceeded, for example, the differentiation of a calm day to a sign of depressive disorder[33]. Xpression is a mood analyzer, developed by the British company Ei Technologies, that enables monitoring the state of mind of the individual using only the voice, more specifically, through the attribution of emotions during calls made by the user in their day-to-day, thus analysing the state of anxiety, stress and depression. This application is not available in the distribution market , although it is recognized it is importance in studies conducted in this field[46]. Menthal was developed by researchers at the University of Boon, in Germany, for the Android operating system. The main purpose of this application is to monitor the use of the personal phone, recording the time that the user spends on the phone, analysing which applications are used more often. This data is sent to an anonymous server where the information and statistics of each user are collected, and then analyzed by experts. Studies are being made using this application, with the aim of developing a component which allows for detecting depression[47]. Another type of existing solutions aim to guide or help the user with a tutorial to prevent depressive disorders, based on general risk factors of depression. In most of these applications, the user is required to answer the inquiry based on established scales for monitoring this disorder, so there are not any specificity of data on the user’s phone to monitor depression. We now list some solutions in the market following these guidelines: CBT Depression Self-Help Guide was developed by the Limited Liability Company, Excel at Life (USA). It is a guide support to detect depression, providing therapeutic articles, diary of cognitive thoughts to learn to challenge thoughts causing stress, providing CHAPTER 4. MARKET ANALYSIS AND EXISTING SOLUTIONS 33 positive thoughts, suggestions for helping track the concentration through motivation and screening tests with graphic support to monitor the severity of depression[15]. MyM3 is an application developed by collaborators at the University of Georgetown, Washington (USA) that, in collaboration with cognitive behavior therapists have created a list for self-evaluation of primary care to monitor potential symptoms of anxiety, indicating the relative risk of depressive symptoms or traumatic disorders. Such evaluation is made by the user and the data is then sent to a designated health care professional that analyzes the questionnaire and obtains relevant information helping the psychologist on getting a better diagnosis at the first visit with the patient[38]. Depression Calculator is an application developed by Egton Medical Information System Limited (UK), based on Patient Health Questionnaire Scale (PHQ-9), through which the user responds to this standard questionnaire trying to obtain a diagnosis of the depressive state. It is also provided a digital informative leaflet with important information and advices about Depression and antidepressants previously analyzed by experienced authorities in the field[40]. Chapter5 DepSigns: A CRISP-DM approach to Depression Signs Detection through Smartphone Usage Data Throughout this chapter are presented the six phases of the DepSigns project, following the CRISP-DM process. 5.1 Business Understanding Probably one of the most important phases of any Data Mining project is focusing on understanding the project’s objectives, transforming into a formal Data Mining definition, consequently developing a primary plan designed to achieve the project’s goals. It is necessary to determine the business’ objective, and the client’s goals and needs to be clear and describe all the resources available to reach these goals. Our project specification required the use Machine Learning to infer depressive symptoms in elders. Psychologists gave us the information that depression detection is mostly related to changes in habits. Therefore, we decided to also conduct a statistical analysis on elder’s data to cover this issue. However, the statistical analysis soon proved to have issues, such as when the elder is already depressed when we first 34 CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 35 start collecting data. The Machine Learning process is based upon the psychologist’s manual process, which is described in Section 5.2. It specifies which data should be analysed and how we decided on the structure, layout, and organization for our project, and chose an algorithm which would suit those requirements. We believe our method is closely related with the psychologist’s manual process and, therefore, results should be similar, but only further experiments will tell. Most of this process consisted of a set of requirements from our project specification and, therefore, a deep analysis on the business understanding for this project is outside the scope of this thesis. 5.2 Data Understanding This project handles sensitive data, which is personal to seniors. Therefore, there are some restrictions regarding how data can be disseminated. For this reason, we have created a data generator which simulates the physical and psychological states of abstract individuals. Next, we describe how this data is stored, organized, and how we have preprocessed it to build our dataset. 5.2.1 Data Generation At this stage, it is important to mention that no data source whatsoever was provided for this project. Originally, all the data was intended to be originated from the Smart Companion project of Fraunhofer Portugal, which is an android customization that was designed to address senior’s goals and needs, with the goal to support seniors in their daily activities[41]. However, the database did not fulfil the requirements for this project and there was not enough data collected to support a study of this kind. For these reasons, we decided to create a generator which simulates the necessary data. This has imposed a great challenge on the replication of real-life situations. Using computer CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 36 generated data makes it impossible to guarantee that the database represents actual real-life situations, as there is no way of matching this data to real-life everyday cases. As such, the generator was created under psychologist guidance, which have provided feedback on common cases on which data would be useful, as well as its relevance in this study. The generator uses a probabilistic approach to this problem. Instead of having data from real patients, we have generated a list of patients, as well as their behavioural patterns for the duration of several weeks. This was achieved by using a Beta distribution on the moods of the patients, from which the rest of the data is inferred. Beta distribution is a family of continuous probability distributions defined on the interval [0,1] configured by αand βshape parameters, which are the exponents of the random variable, and control the shape of the distribution. This distribution assumes the modulation of the behaviour of random variables limited to intervals of finite length. There are several areas of statistical description that use this distribution, such as genetic population or genetic heterogeneity in the probability of HIV transmission, among others. The standard Beta distribution gives the probability density of a value xon the interval [0,1], as shown in Equation 5.1. P(x) = (1 −x)β−1xα−1 B(α, β)(5.1) where: P(x)is the probability function α,βare the distribution parameters B(α,β) is the beta function Figure 5.1 shows a few examples of possible Beta distributions. CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 37 Figure 5.1: Probability density function (reprinted from [3]). From the beginning, this was intended to be a preliminary study. It is not known to us that any previous study of this kind exists, which limits to a large degree the basis we had as a starting point. Having no background on how to automatically detect depressive symptoms using our methodology, we chose our methods based on psychologist feedback and guidance, providing that our methods would closely relate to the methods used by the psychologist. For the psychologist analysis, it is important to know: •how did the elder perform, activity wise? •how did the elder perform, location wise? •how did the elder perform, communication wise? •and so on. The order in which these fields are evaluated is relevant. This means that the data concerned with elder’s activities is considered the most relevant by the psychologist, under which a failure would indicate immediate depressive symptoms. The order of relevance given by the psychologist for each field is the following: 1. Activities 2. Locations 3. Communications CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 38 1. Calls 2. Not Answered Calls 3. Messages 4. Moods 5. Ludic Activities The statistical study shows as a complement to the Machine Learning process in two distinct ways: •it complements the study by giving more relevance to personal habits; •it allows for a much better data visualization tool. We decided on using a generic central tendency metric which would allow future work to extend the set of available methods. On top of that, the application is also able to adapt to different methods used by different psychologists, provided that a suitable evaluation function exists and is implemented. However, we did focus on analysing the last two weeks of data, as previous studies exist that indicate that two weeks are sufficiently relevant to provide reliable results, as defined by the American Psychiatric Association [5]. One last thing to consider, related to the statistical process, consists in defining how the data is first evaluated. As an example, we consider the number of calls and not the time spent talking on the phone. The following lists how we evaluated each field as well as the reason that justifies our approach: •Activities - total time spent performing a given activity. The amount of time is considered relevant as it yields a better indicator of daily physical activity then a count of the number of activities performed. We consider nightly activities to be prejudicial to the individual, as it indicates poor sleeping habits. •Locations - total time spent away from home. Just leaving the house is not indicative of a positive activity, as there is no indication of the kind of activity. For instance, several activities could be performed, out of the house, like buying bread, which is not as relevant as leaving the household for several hours. Even so, we have no way of identifying the quality of the absence from home. That is, leaving the house, for how long as it might be, does not indicate whether the CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 39 visited place contributes positively to the elder’s state of mind, and there is no automatic process, that we are aware of, that does so. As such, we consider all locations to have equal relevance, and make no attempts on identifying the location whatsoever. •Communications - number of calls or messages sent and/or received. The amount of time spent using the phone was considered irrelevant by the psychologist, as the amount of time does not vouch for the quality of what is being talked; instead, we only count the number of actions taken using the phone. Take as an example a situation in which an elder spends several hours talking to the same person. This could either indicate a very active social life or a very strong need for attention. •Moods - count of the weekly mood self-evaluation. Seniors have four options to self-evaluate: Bad,Not Well,Fine, and Very Good. Each mood symbol is given a value by assigning to it a position in an array. Three is then subtracted from this position, making sure that bad moods yield negative values and good moods positive ones. The sum of these values is used as the the weight factor for each weekly mood. •Ludic Activities - amount of time spent performing the activity weighted by the achieved score. The reason for this relates to elderly cognitive changes, as detailed in Chapter 2, Section 2.1. For each name from a list of hand-given names the generator creates a patient entry. Each patient is then given a mood which we generate by assigning random integer values between 1 and 5 for the αand βparameters of the Beta distribution function, respectively. The reason these values are random is mainly because it is unknown to us what the correct distribution would be, thus allowing us to conduct several studies and getting to some final result which would seem reasonable. Again, the psychologist feedback was essential at this point, which has later confirmed reasonability after inspecting several cases. Once again, this is only justifiable from the absence of any real-life data source. As the probabilistic function generates real values between 0 and 1, we then assigned thresholds for mood values. According to psychologist advice, the moods can be CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 40 divided in four categories. Each category was given a 25 percentile of the range provided by the distribution function. As such, any mood value below 0.25 would be considered Bad, between 0.25 and 0.50 as Not Well, between 0.50 and 0.75 as Fine, and above 0.75 Very Good. These categories are then used as thresholds in the generation of all other data referring to the same patient, in the same week. One mood value was generated per week for the entire period of 12 weeks. Every other field is generated by using the same Beta distribution parameters; that is, the αand βvalues that were randomly generated for the mood. However, the mood plays an important role in the generation of these fields, as it is used as a threshold for whether the patient complies or not with some obligations. Because the Beta distribution generate good habits for bad moods, and vice-versa, the worst the patient’s mood the more we would require of him/her. That is, for values generated between 0 and 1, indicating how willingly the patient would comply with his/hers obligations.By obligations we mean any average person’s daily duties, like answering phone calls, going out, waking up in the morning, among others. We assume that a patient having a Bad mood would require a willingness threshold of 0.90; however, we would require just 0.70 for patients with a Not Well mood; 0.50 for Fine; and 0.25 for Very Good. Again, these thresholds are merely experimental and no sustainable data will ever be available unless provided from real-life situations. 5.2.2 Database Model The database was modeled to represent a set of constraints that were agreed with psychologist guidance. The goal, however, was to match, as closely as possible, the structure of the Smart Companion database, without compromising the physiologist’s requirements on elder data. To achieve this, the number of fields present in the database was kept at a minimum, allowing for improvement when trying to match it with the Smart Companion database. Several entities were listed by the psychologist as essential for this study (see Figure 5.2): CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 41 •Users are a presentation of a studied entity, which, in the context of the current application, means an elder. This field is central to the database and is used to relate all other fields by means of foreign keys. The data stored in this table encompasses the name and date of birth for each elder. •Activities represent daily physical activities separated by three different periods of the day: morning, afternoon, and night. The time at which the activity takes place indicates the period to which it belongs. These periods are later used to influence the relevance of the activity in the elder’s schedule. Each entry consists of a start and end timestamps indicating the period of the activity. •Locations indicate places that seniors visit when they leave home. The place itself is not stored, instead just the period in which the person was absent from their place of residence. This decision was justified by the computer’s inability to decide whether any place was a good or a bad influence on any one given person. As such, just the timestamps for when the elder leaves and later arrives at home were stored. •Phone calls are a registry of the elder’s voice communications over the telephone. Again, because the concrete context of the call cannot be determined as being positive or negative for the person under study, the only fields that were stored where the timestamps for the beginning and end of each phone call. This also allow us to compute the amount of time spent on the telephone. •Not Answered Phone Calls consist of a registry of missed phone calls, not answered by the peer under study. Because of our inability to know the relevance of the other peer for the elder’s mood, we only store timestamps. In this case, however, and because missed phone calls do not have a duration period, we only register one timestamp, indicating the time at which the call was missed. •Messages could be text messages or messages of other kinds of media. Again, the inability to infer whether the other peer is either a positive or negative influence to the person under study, forced us to discard any information about that peer, keeping the database structure minimal. For this field only one timestamp is stored, indicating when the message was first received. •Moods are supposed to be retrieved from elder feedback. That is, the senior him- CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 48 data, and even abstract the process. These queries have the following two purposes: 1. Data Reduction - The raw data is stored in a daily basis. However, for the purpose of this study, and under psychologist guidance, it was considered irrelevant to go as deep as analysing daily data. According to the psychologist’s opinion, the conclusions become more sustainable when analysing weekly periods, and therefore the first step consists in aggregating the data in that way. 2. Data Transformation - After being aggregated, we create weighted resumes of the data. How each resume is computed depends on the actual field being weighted. We next proceed to explain this process. As far as Data Transformation is concerned, the set of fields in the database can be divided in 5 different groups. •Compute three summations of the time spent performing an activity divided by periods of the day (morning, afternoon, and night, depending on the time of the day). Each summation is then multiplied by a constant weight factor which as been specifically assigned to its matching period. In this case we used a multiplier of 2.0 for the morning period, 1.0 for the afternoon, and -2.0 for the night period. These values, however, are not justified and should be considered future work. For this reason, we left the constants out of the query allowing for easy customization. Then all values are added up. This way we give more or less relevance to each period of the day, while still ending up with a single, descriptive, value. Only the field for Activities fit into this category. 1SELECT id, SUM(period *duration) as weight, week 2FROM( 3SELECT id, TRUNCATE(SEC_TO_TIME(SUM(TIME_TO_SEC(TIMEDIFF(end, start)))) /60,0) AS duration, 4(CASE 5WHEN TIME(start) BETWEEN ’06:00:00’ AND ’14:59:59’ THEN %f 6WHEN TIME(start) BETWEEN ’15:00:00’ AND ’23:59:59’ THEN %f 7WHEN TIME(start) BETWEEN ’00:00:00’ AND ’05:59:59’ THEN %f 8END)AS period, 9CONCAT(YEAR(start),’/’,WEEK(start)) AS week 10 FROM impl_activity 11 WHERE user_id = %d GROUP BY week,period ORDER BY start CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 49 12 )AS tmp 13 GROUP BY week Listing 5.2: Query used for Activities. The time periods are given as arguments to the query (shown as %f) yielding the weights used for each period. There are three daily periods (morning, between 6am and 3pm; afternoon, between 3pm and midnight; and night, between midnight and 6am). The week for each entry is represented in the format YEAR/WEEK. The user ID is also a parameter represented here as %d. •Compute the summation of time spent performing the activity on any given week multiplied by the performing scores. This gives not only the amount of time spent performing those activities, but also the weighted quality of the result. Only the field of Ludic Activities fit into this category. 1SELECT id, week, duration *score AS weight 2FROM ( 3SELECT 4id, score, 5TRUNCATE(SEC_TO_TIME(SUM(TIME_TO_SEC(TIMEDIFF(end,start)))) / 60, 0) AS duration, 6CONCAT(YEAR(start),’/’, WEEK(start)) AS week 7FROM impl_ludicactivity 8WHERE user_id = %d 9GROUP BY week 10 ORDER BY start 11 )AS tmp Listing 5.3: Query used for Ludic Activities. The score will be used as a weight factor to leverage the quality of the result. •Compute the summation of weighted mood values for each weekly period. Each mood is assigned a different weight, being -2 Bad, -1 Not Well, 1 Fine, and 2 Very Well. These values are not relevant per se, as long as bad moods are negative and good moods are positive, and equally relevant states of mood are equality distant from zero (that is, Bad is worth -2, while Very Well is worth it is absolute value). This way the summation of these values yields the overall quality of the elder’s mood state. Only Moods fit into this category. CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 50 1SELECT id, CONCAT(YEAR(day), ’/’, WEEK(day)) as week, SUM((FIND_IN_SET( moodvalue, ’B,N,?,F,V’) - 3)) AS weight 2FROM impl_mood 3WHERE user_id = %d 4GROUP BY week Listing 5.4: Query used for Moods Each mood symbol is given a value by assigning to it a position in an array. Three is then subtracted from this position, making sure that bad moods yield negative values and good moods positive ones. The sum of these values is used as the the weight factor for each weekly mood. •Compute the total amount of time spent performing a given obligation in a given week. This group applies to Locations. 1SELECT id, TRUNCATE(TIME_TO_SEC(SEC_TO_TIME(SUM(TIME_TO_SEC(TIMEDIFF( TIME(check_out), TIME(check_in)))))) / 60, 0) AS weight, CONCAT(YEAR (check_in), ’/’, WEEK(check_in)) as week 2FROM impl_location 3WHERE user_id = %d 4GROUP BY week Listing 5.5: Query used for Locations The time spent by the elder out of home is computed as the difference of the check-in and check-out timestampos, in total number of minutes. •Count how many times a given obligation was fulfilled in a weekly period. This group applies to Calls,Not Answered Calls, and Messages. 1SELECT id, CONCAT(YEAR(impl_notansweredcall.when), ’/’, WEEK( impl_notansweredcall.when)) as week, COUNT(*)as weight 2FROM impl_notansweredcall 3WHERE user_id = %d 4GROUP BY week Listing 5.6: Query used for Not Answered Calls, although Calls and Messages use a similar query. One of the consequences of this process of Data Transformation consists in reducing every weekly data set to a single entry, which consists of a numeric, weighted, repre- CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 51 sentation of the senior’s performance on any given week. A training sample consists of a matrix with Nrows (observations Oi,1≤i≤N) and Mvariables xj,1≤j≤M. There are exactly Kclasses, which must be known before hand, and which are used to classify the observations. In our case, there are only two classes, as we categorize elders as either having depressive symptoms or not. The number of rows Nis dictated by how many different elders the psychologist has given feedback about, and the number Mof variables is fixed and corresponds to twice the number of fields we are currently studying, as each field is analysed for the past two (2w) and four weeks (4w). Table 5.1 shows a sample of all fields discussed in Section 5.2. Each row is analysed by the CART algorithm to generate the Decision Tree, and the Classification column is used to classify the trained data. Observation O1O2... ON Activities (2 weeks) Activities (4 weeks) Locations (2 weeks) Locations (4 weeks) Calls (2 weeks) Calls (4 weeks) Not Answered Calls (2 weeks) Not Answered Calls (4 weeks) Messages (2 weeks) Messages (4 weeks) Mood (2 weeks) Mood (4 weeks) Ludic Activities (2 weeks) Ludic Activities (4 weeks) Classification Table 5.1: Example of a training sample. CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 52 5.4 Modeling We use a Decision Tree to classify depressive symptoms, mostly due to its high level of interpretability. To generate this tree, we allow the psychologist to provide feedback on a subset of elders, indicating whether depressive symptoms exist or not. We commit this feedback, in the form of a boolean flag, to the database associating it with the given elder instance. We can then generate a training set by correlating the elder’s history with the provided feedback. We generate this set in a format that the scikit-learn framework [9] understands and is capable of generating the Decision Tree from. We then save the tree to a file using pickle [22], enabling us to load it later, on demand, without having to generate it again. We also generate PDF and SVG files with the rendered tree, which is useful as visual support. We use the same format as the training set to classify other samples not used as training. Therefore, we generate a table like we did for the training sample, and use it to predict the elder’s state of mind according to previously trained rules. Contrary to the statistical analysis, the decision tree classifies symptoms according to global population standards. This study is useful in identifying elders with depressive symptoms, even if they were already depressed when data collection first started. Decision Tree Learning is one of the Machine Learning’s most used and practical methods for inductive inference, being the most popular among the inductive inference algorithms in large areas such as health, finances, learning to diagnose medical cases, or evaluate possible cases of financial risk. This method aims at approximating functions of discrete values with robustness on data with possibility of noise and learning disjunctive expressions. It is represented by a decision tree with possibility to associate if-then rules to improve human readability. Decision Tree learning classifies instances from the root of the tree to a leaf node that provides the class instance. Each node of the tree specifies the test of some attribute of the instance. An instance is classified starting at the root node of the tree and testing the attribute related to this node, following the branch that corresponds to the value of the attribute in the instance at matter. This process is then repeated for the subtree below until a leaf node is reached [35]. CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 53 The Learned Decision Tree shown in figure 5.5 has the objective of classifying the level of risk of heart failure, by checking if it is appropriate common symptoms. The instance < Systolic blood pressure ≤91, Age ≤62.5, Sinus tachycardia present =yes > would be sorted down to the leftmost branch of this decision tree and then will be classified as a negative instance. The tree predicts that Risk =High. In sum each path from the root of the tree that goes to a leaf corresponds to a conjunction of attributes tests and the tree itself a disjunction of these conjunctions. A Classification Tree is a classification method which uses training - or historical - data to construct a decision tree. The training set consists of a classified subset of data for a given sample. This learning sample is used to generate a decision tree which is able to classify untrained data. The decision tree divides the trained sample into subsets by asking boolean questions. At each iteration, these questions allow choosing one of the two subsets, according to the answer. That is, asking “Did the elder send over 15 text messages on a given week?” divides the sample data into two: those that did send over 15 text messages, and those that didn’t. The CART algorithm attempts to find the questions that provide the best division of the sample into parts as much homogeneous as possible. This process is repeated until all samples in the given subset are of the same class, which correspond to the tree’s leafs, shown in yellow in Figure 5.5. Figure 5.5: An example of a Learned Decision Tree, with leafs shown in yellow (reprinted from [52]). . CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 54 Maximum homogeneity is defined by an impurity function. As such, at each node the CART algorithm solves a maximisation problem. There are several impurity functions, but only two are generally used, the Gini splitting rule and Twoing splitting rule. Scikitlearn, our framework of choice, uses the Gini splitting rule. This rule looks for the largest class in the dataset and tries to isolate it from all the others. Figure 5.6 shows a decision tree generated from four classes, A, B, C, and D, with sizes 40, 30, 20, and 10, respectively. The Gini splitting rule first separates class A from the rest of the dataset, after that B, and so on. Such homogeneous splits are not always possible, in which case the dataset is split into less optimal groups 5.6. Figure 5.6: Gini rule example(reprinted from [14]) Figure 5.7 shows an example of a tree generated by the application. In this case, the field of activities for the past two weeks is evaluated first. In case the elder spent less than (approximately) 157 minutes practicing some form of activity, then the classification takes the right branch, which is a leaf, and indicates that there are depressive symptoms. Otherwise, the number of not answered calls for the past two weeks is analysed instead. In this case, the tree classifies the elder as having depressive symptoms if there are at least 16 not answered calls in the past two weeks. Next the field of not answered calls is analysed for the past two weeks, but in this case no conclusion is taken in either case, as neither child is a leaf. For the sake of brevity, we’ll analyse the right branch, which leads to the nearest leaf indicating the absence of depressive symptoms by analysing, again, the field of not answered calls for the past two weeks. In contrast with the previous analysis, less than 11 calls missed indicate that there are CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 55 no depressive symptoms. Figure 5.7: Visualization of a Learned Decision Tree. 5.5 Evaluation The evaluation step was probably one of the greatest impediments we found on this project. Not only because there’s no real data, but also because our solution remains to be tested in a real-life situation. Our tests consisted mainly of showing our results to a trained psychologist asking whether the results seemed reasonable, and whether our implementation would suit the standard depression-detection process. The feedback we got on that regard on the psychologist’s part was good, even though it was agreed that further testing was still necessary. As such, most of this process was delegated to future work, making our implementation not suitable for release. In order to finish evaluating our results, we would require three things: •access to the SmartCompanion database; CHAPTER 5. A CRISP-DM APPROACH TO DEPSIGNS 56 •further tests under psychologist guidance; •and real-life situation tests. Only then could a release version be considered. 5.6 Deployment The deployment phase is described in detail in chapter 6. Chapter6 Technical Aspects and Implementation of DepSigns In this chapter we present a descriptive view of how we structured and implemented this project. We start by exposing and explaining the technologies used and then proceed to correlate them with our own models. 6.1 Used Technologies Throughout this section, we describe used technologies to better understand the DepSigns implementation. 6.1.1 Scikit Learn Scikit Learn is an open source machine learning library for the Python programming language. In this library we have the possibility to use several Machine Learning algorithms such as Decision Tree Learning, Support Vector Machines, or Naive Bayes, among others. This library is designed to interoperate with the numerical and scientific Python libraries NumPy and Scipy [9]. 57 CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 64 represents. We start by explaining the generators for the two central modules in our design: users and moods. The user generator generates all permutations from two lists of names: one of given names and other of surnames. We also provide a year interval which will limit the minimum and maximum age of each generated individual, although this is quite irrelevant for our study, as all users are considered as being elders. For each permutation we generate a different entry, as we can see in Listing 6.1. 1class UserGenerator(Generator): 2 3def generate_names(self): 4for given in self.given_names: 5for sur in self.sur_names: 6yield given + ’’+ sur 7 8def generate(self): 9l = [] 10 for name in self.generate_names(): 11 user = User(name=name, birth=self.generate_date(self.min_year, self. max_year)) 12 l.append(user) 13 user.save() 14 return l Listing 6.1: The UserGenerator class We then proceed by defining the mood generator. As all the users have already been created, we now generate mood values using the process described in Section 5.2. In review, we create random alpha and beta parameters, between 1 and 5, for the beta distribution function, which generates values in the interval [0; 1], with a given distribution. Moods are then discretized into four categories: Bad,Not Well,Fine, and Very Well. Also, the start and end parameters indicate the period of data that will be generated. Between these two dates, we will generate an entry for each day. This also applies for every other generator from here on. 1class MoodGenerator(Generator): 2def generate(self): 3for user in self.users: 4user.alpha = random() *5+1 CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 65 5user.beta = random() *5+1 6start, l = self.start, [] 7 8while start != user.end: 9mood = betavariate(user.alpha, user.beta) 10 11 if mood < .25: mood = ’B’ 12 elif mood < .50: mood = ’N’ 13 elif mood < .75: mood = ’F’ 14 else: mood = ’V’ 15 16 mood = Mood(user=user, day=start, moodvalue=mood) 17 mood.save() 18 l.append(mood) 19 start += timedelta(days=1) 20 21 user.moods = l Listing 6.2: The MoodGenerator class The following code defines the generator for activities. Because every remaining entity generator uses pretty much the same method, we define the ActivityGenerator as a base class that implements the make(self, user, start, end) method (cf. Listing 6.3,line 31). This is overriden by subclasses to change the model instance that is created in the process. So, to generate the remaining entities, we use the same αand βparameters for the beta distribution function, as randomly generated when creating the moods. The moods are then, again, used as a threshold for obligation compliance, as previous described (See Chapter 5, Section 5.2). 1class ActivityGenerator(Generator): 2 3def generate(self): 4self.count_ok = { 0: 0, 1: 0, 2: 0 } 5self.count_not = { 0: 0, 1: 0, 2: 0 } 6for user in self.users: 7start, index = self.start, 0 8while start != self.end: 9mood, index = user.moods[index].moodvalue, index + 1 10 for iin range(2): 11 prob = betavariate(user.alpha, user.beta) CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 66 12 self.generate_period(i, prob, user, start, mood) 13 prob = 1 - betavariate(user.alpha, user.beta) 14 self.generate_period(2, prob, user, start, mood) 15 start += timedelta(days=1) 16 17 def generate_period(self, period, prob, user, day, mood): 18 if mood == ’B’: threshold = .90 19 elif mood == ’N’: threshold = .70 20 elif mood == ’F’: threshold = .50 21 else: threshold = .25 22 if prob < threshold: 23 return 24 start_hour, end_hour = self.time_according_to_period(period) 25 start, end = self.generate_time_2(start_hour, end_hour) 26 day = ("%s" % day).split(’ ’)[0] 27 start, end = "%s %s" % (day, start), "%s %s" % (day, end) 28 act = self.make(user, start, end) 29 act.save() 30 31 def make(self, user, start, end): 32 return Activity(user=user, start=start, end=end) Listing 6.3: The ActivityGenerator class All other generators simply override the make method to change the returned instance. The exception goes to the LudicActivityGenerator which also generates a score value. 1class LudicActivityGenerator(ActivityGenerator): 2def make(self, user, start, end): 3prob = betavariate(user.alpha, user.beta) *100 4return LudicActivity(user=user, start=start, end=end, score=prob) 5 6class LocationGenerator(ActivityGenerator): 7def make(self, user, start, end): 8return Location(user=user, check_in=start, check_out=end) 9 10 class NotAnsweredCallGenerator(ActivityGenerator): 11 def make(self, user, start, end): 12 return NotAnsweredCall(user=user, when=start) 13 14 def generate_period(self, period, prob, user, day, mood): CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 67 15 return super(NotAnsweredCallGenerator, self).generate_period(period, 1 - prob, user, day, mood) 16 17 class CallGenerator(ActivityGenerator): 18 def make(self, user, start, end): 19 return Call(user=user, start=start, end=end) 20 21 class MessageGenerator(ActivityGenerator): 22 def make(self, user, start, end): 23 return Message(user=user, when=start) Listing 6.4: The remaining generator classes The raw data is not in a convenient format for use in the application. On top of that, it was a well known fact from the start of this project that the data source could change in the future (that is, use of the Smart Companion database). As such, we created a Data Abstraction layer over the Database layer which allows to both abstract from the data source and convert the data into a more suitable format. 6.2.3 Authentication and Authorization Being multi-user in nature, this application required an authentication and authorization process. Authentication consists in identifying someone as actually being whom he/she claims to be. Each user is required to have a username and password. We use django as an underlying mechanism to implement this feature, as it already provides the necessary routines to do so. Authorization, on the other hand, consists in assigning each authenticated user a set of permissions. Our service only implements two kinds of users with different levels of permissions : psychologists and caregivers. The psychologist is allowed to: •access the main page, where main statistics, general alerts, a calendar, among others are present; CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 68 •search patients by name; •provide manual feedback to build the training dataset given to the decision tree learning algorithm; •see the learned decision tree, as generated by the CART algorithm; •see individual, personal, alerts for each elder; •see data history; •choose the centrality tendency metric used for each patient. The caregiver has a reduced set of privileges: •access to the personal page of only one elder; •see field-wise and population-infered alerts generated for that elder. By reducing the amount of information we allow the caregiver to see, we protect the seniors personal information. This data, however, is fundamental for the psychologist to derive conclusions on the seniors mental state. 6.2.4 Views This web service is composed of a set of pages which allow its users to explore the implemented functionality. Next we describe these pages, their URL mappings, authorization requirements, and objectives. The set of URL mappings constitute the service’s API. Each URL serves a different purpose, as described in table 6.1. CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 69 URL Purpose / The main page /login User authentication /logout User logout /register User registration /feedback Allows providing learning feedback, use to construct the learning decision tree /patient See patient history and personal alerts /analysis See a renderization of the learned decision tree, generated from psychologist feedback /patient_list See a list of available patients, as well as searching patients by name Table 6.1: Purpose of each URL mapping Some of these URLs return HTML, which consists of a visualization of some sort, others use AJAX to perform some distinct operation. We implement this by using django Views, which consists of classes that allow specifying how each URL behaves and what HTTP methods it supports. We now describe each view individually. 6.2.4.1 Main Page The main page has tree different views, depending if the user is: •not authenticated; •not registered; •authenticated. In case the user is not authenticated the main page shows a login form. The login form allows the user to identify him/herself.The figure 6.2 shows this page. CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 70 Figure 6.2: Login page If the user is not registered yet this page also offers a registration form, as shown in figure 6.3. Figure 6.3: Registration form If the user is already authenticated, the behavior for this page varies according to the type of user. In case the user is a caregiver, he/she will be redirected to the page describing the patient that has been assigned to him/her. If the user is registered as a psychologist the page will show a navigation menu, which allows navigating to other pages, a list of alerts identified by the Decision Tree, the set of predictions made by the Decision Tree, and other sections which demonstrate features to possibly implement in future versions of the service. Figure 6.4 shows a screenshot for the psychologist’s main page, where several features can be identified. The navigation menu (top left) has links for the list of patients and the learned decision tree projection; the lists of alerts (top center, in red) gives an overview of patients with infered alerts; Decision CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 71 Tree Predictions (the section with the same title) shows how the patients are classified with (Depressive Symptoms and No Symptoms); all other sections are future work and do not provide any functionality whatesoever. Figure 6.4: Psychologist’s main page. Figure 6.5: Caregiver’s main page. 6.2.4.2 Logout The logout view causes the user to be unauthenticated and redirects the page to the main page. CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 72 Figure 6.6: Logout option. 6.2.4.3 Feedback The feedback view uses AJAX to allow the psychologist to provide learning feedback to the service. The psychologist indicates whether, in his/her opinion, the patient belongs to the class that shows depressive symptoms or not. This feedback is then used to construct the training dataset provided to the decision tree learning algorithm (CART). No visualisation whatsoever exists for this URL mapping, being an AJAX call. This view can be called from the page showing patient history and personal alerts. This view also renders the visualization of the learned decision tree. By rendering the tree everytime feedback is provided, we are rendering it every time there are changes. As such, we need not render the same image several times, by caching this renderization on the server. This process is described in Chapter 5, Section 5.4. Figure 6.7: Feedback psychologist option. 6.2.4.4 Patient In this page we show a detailed history of the elder’s activities. Each type of registered activity may also issue an alert, indicating that the statistical analysis found deviances CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 73 regarding the centrality metric for the given field. Both the psychologist and caregiver can see these alerts, but only the psychologist can see the details for the seniors recent history. Also, this page allows choosing the generic evaluation function, which might change the field-wise alerts shown. Finally, two sections exist dedicated to the Machine Learning process. At the top of the page there’s an indication of the class inferred by the learned decision tree (that is, either the elder shows depressive symptoms or none at all). At the bottom of the page, there’s a section dedicated to psychologist feedback. Here, if the psychologist has not yet given feedback about this elder, he/she can do so. If the psychologist has already provided feedback, their conclusions are shown here, indicating what the feedback was at the time it was given. Note, however, that these two processes are independent from each other. The psychologist needs not to provide feedback for all elders in order to see Decision Tree inference for any given senior. Still, the psychologist can provide feedback to be later incorporated into the training data supplied to the decision tree learning algorithm. Figure 6.8 depicts the patient page, showing history for the seniors Activities, and alerts for each individual field. At the top left corner there’s a combo box that allows the psychologist to choose the evaluation function. At the center top there’s an indication of the Decision Tree classification; red indicates depressive symptoms and green the absence of symptoms. At the bottom there’s an indication that the psychologist has already indicated this patient as being depressed Figure 6.8: The patient page. CHAPTER 6. TECHNICAL ASPECTS AND IMPLEMENTATION OF DEPSIGNS 80 1def infer_depression_list(): 2return [usr for usr in User.objects.all() if infer_user_symptoms(usr)] Listing 6.8: Code for inference depression list 1var doughnutData = [ 2{ 3value: {{ count_depressed }}, 4color:"#E64C65" 5}, 6{ 7value : {{ count_not_depressed }}, 8color : "#11A8AB" 9}, 10 ]; 11 var myDoughnut = new Chart(document.getElementById("canvas").getContext("2d") ).Doughnut(doughnutData); Listing 6.9: Render Doughnut graphic Chapter7 Conclusion We finish by giving an overview of the conducted work in this thesis. We present our main contributions and challenges and discuss further work possibilities. We decided to incorporate in Depsigns data related to: activity level categorized by time of day (morning, afternoon, and night); locations (how often the user leaves the house); mood (a self-assessment made by the user); communications (calls and messages) and ludic activities (namely, cognitive exercises). We conduct a statistical analysis on the collected data and use it to infer depression symptoms from personal habits. The goal of this statistical analysis is to find deviations from the senior’s usual behaviour, by detecting standard patterns in the senior’s daily activities, and causing field-wise alerts to be shown to the psychologist or caregiver. To detect the mentioned outliers/deviances, we use an abstract evaluation function which can be configured by the psychologist using this software. The results achieved by conducting this study vary a lot according to the senior’s initial condition. The system will not generate an alert for senior’s that already showed depression signs when we first started sampling data, since a depressed state will correspond to their normal condition. This is why we also conduct a study using a Machine Learning algorithm. Decision Tree Learning is one of the most used and practical methods for inductive inference, being the most popular among the inductive inference algorithms in large areas such as health or financial, learning to diagnose medical cases or evaluate possible cases of financial risk [35]. We use this method to classify a set of symptoms as depressive/non-depressive condiction. For this purpose, we asked a psychologist to provide feedback on some 81 CHAPTER 7. CONCLUSION 82 seniors, indicating whether symptoms existed. We commit this feedback and generate a dataset by correlating the senior’s history with the provided feedback. A decision tree is trained using this dataset and the learned model is used to predict senior’s conditions on a new dataset. Contrary to the statistical analysis, the decision tree classifies symptoms according to global population standards[45]. 7.1 Summary As it should be clear by now, the psychologist’s feedback was essential in this project. Most of our choices towards depressive symptoms detection were based on this guidance - no other study towards that goal was ever conducted. The choice of algorithms therefore still lacks some justification and further analysis. As of this moment, our conclusions were driven from extreme cases, which the psychologist analysed superficially and indicated whether there were depressive symptoms or not, making it clear to us as to whether our inference process was correct. Even though we achieved a 100% success rate, we still lack the confidence to vouch for these results as it is clear to us that some considerings remain unmade, as we will see in Section 7.3. As such, we know not the actual success rate nor is it within our reach to find out, as some basic requirements would have to be met before: •real-life data; •tests performed by psychology experts; •fine tunning, as described in Section 7.3. We can conclude, however, that our process can at least identify extreme cases, when considering the generated data. Statistical analysis failed at identifying depressive symptoms when the elder is already admitted in a depressive state, while the Machine Learning process is not suitable for inferring conclusions from personal habits. Even so, psychologist feedback proved great as to how our process identifies depressive symptoms, just as much as it did with the quality of the result. CHAPTER 7. CONCLUSION 83 We consider that we have fulfilled the proposed objectives, perhaps with the exception of the evaluation phase, which lacks work towards validating our results, and other issues discusses next, in Section 7.3. 7.2 Challenges Gathering data and unprecedent related work were probably the greatest challenges we faced developing this project. Being personal in nature, the collected data is sensitive and imposes privacy issues. On top of that, the current database for the Smart Companion project did not fulfil the necessary requirements, having forced the fabrication of artificial data. This creates difficulties in achieving real-life representations of elder’s state of minds, if such is possible at all. As such, psychologist guidance and advice was essential and consisted of an important backbone for the development process, although multidisciplinary projects often impose other kinds of difficulties, such as achieving an understanding when the people involved stem from completely different backgrounds. Not having found any precedent work of the kind also imposed a great challenge. The choice of algorithms and methodology is not known to be optimal, having pushed much of the development process for future work. Once again, these choices were made under psychologist guidance, but such is not indicative of optimality, and we rather chose an approach which we consider to be good enough. We expect future work to be based upon our own, when results will possibly already be available. 7.3 Future Work It has been stated many times throughout this thesis that nothing can replace real data, and that our simulator is far from being validated. In order for this project to be continued this would have to be one of the first issues to be resolved. Access to the SmartCompanion database would be, therefore, a great contribute to further developing CHAPTER 7. CONCLUSION 84 this project. This form of automatic data gathering would be the next step towards making this a complete solution. However, other issues exist which still need further studying. A Fine tuning the parameters used to generate the tree is also something we still lack attention in. Decision trees are known to produce less-than-optimal results when upper and lower bound parameters are incorrectly setup or, as with in our case, nonexistent. That is, in order to ensure optimality, we would have to limit the tree’s depth to both minimum and maximum levels. However, these parameters are not known to us and only by solving the real-life data issue and further studying our results would it be possible to correctly identify and tune such parameters. On top of that, we cannot dismiss the possibility of using other Machine Learning methods. Our method should be tested, by comparing it with other Machine Learning algorithms. It is also not known which central tendency measure would produce the best results when conducting the statistical analysis. Instead, we allow the psychologist to choose such measure from a set of available functions, which can be further expanded. As such, we propose as future work that different types of measures are put to a test, even by different psychologists, and that the results are analysed and compared. Not only could this process identify the best suited method but also new central tendency measures which could further improve our solution. We do not limit our future work proposal to what we think our own work lacks attention in, but also to new features. We believe that this should be a complete tool, which a psychologist could use in the day-to-day work life. As such, we also propose the following: •prioritized alerts, in which the most critical situations would be shown first; •a calendar, in which the psychologist could associate events with elders, such as scheduling meetings, or take important notes; •better statistics of the general population. That is, a deeper statistical analysis of the underlying population, which contrasts with the current solution which only provides the percentage of individuals fitting in each class; •conduct an inquiry and further study other features which would be useful to both CHAPTER 7. CONCLUSION 85 the psychologist and caregiver. AppendixA Acronyms AIDS Acquired Immunodeficiency Syndrome DMS-IV Diagnostic and Statistical Manual of Mental Disorders GDS Geriatric Depression Scale CES-D Center for Epidemiologic Studies - Depression HAM-D Hamilton Scale - Depression PHQ-9 Patient Health Questionnaire KD Knowledge Discovery KDDM Knowledge Discovery and Data Mining KDD Knowledge Discovery in Databases DM Data Mining DepSigns Depression Signs Detection through Smartphone Usage Data Analysis URI Uniform Resource Identifier SOAP Simple Object Access Protocol WSDL Web Service Definition Language HTTP Hypertext Transfer Protoco URL Uniform Resource Locator 86 APPENDIX A. ACRONYMS 87 API Application Programming Language HTML HyperText Markup Language CSS Cascading Style Sheets AJAX Asynchronous JavaScript and XML XML eXtensible Markup Language HTTP Hypertext Transfer Protocol CRISP-DM CRoss-Industry Standard Process for Data Mining AppendixB Source code 1class User(models.Model): 2name = models.CharField(max_length=100, null=False) 3birth = models.DateField(null=False) 4class Activity(models.Model): 5user = models.ForeignKey(User) 6start = models.DateTimeField(null=False) 7end = models.DateTimeField(null=False) 8 9class LudicActivity(models.Model): 10 user = models.ForeignKey(User) 11 start = models.DateTimeField(null=False) 12 end = models.DateTimeField(null=False) 13 score = models.IntegerField(null=False) 14 15 class Mood(models.Model): 16 MOOD_CHOICES = ( 17 (’B’,’Bad’), 18 (’N’,’Not well’), 19 (’F’,’Fine’), 20 (’V’,’Very good’), 21 ) 22 user = models.ForeignKey(User) 23 day = models.DateField(null=False) 24 moodvalue = models.CharField(max_length=1,choices=MOOD_CHOICES) 25 26 class Location(models.Model): 88 APPENDIX B. SOURCE CODE 89 27 user = models.ForeignKey(User) 28 check_in = models.DateTimeField(null=False) 29 check_out = models.DateTimeField(null=False) 30 31 class NotAnsweredCall(models.Model): 32 user = models.ForeignKey(User) 33 when = models.DateTimeField(null=False) 34 35 class Call(models.Model): 36 user = models.ForeignKey(User) 37 start = models.DateTimeField(null=False) 38 end = models.DateTimeField(null=False) 39 40 class Message(models.Model): 41 user = models.ForeignKey(User) 42 when = models.DateTimeField(null=False) 43 44 class DepressionInference(models.Model): 45 user = models.ForeignKey(User) 46 depressed = models.BooleanField(null=False) Listing B.1: Django models used to create the database