scieee AI-readable full text Open interactive document viewer

Implementing electronic scales to support standardized phenotypic data collection - the case of the Scale for the Assessment and Rating of Ataxia (SARA)

Maarouf, Haitham

Abstract

The main objective of this doctoral thesis was to facilitate the integration of the semantics required to automatically interpret collections of standardized clinical data. In order to address the objective, we combined the best performances from clinical archetypes, guidelines and ontologies for developing an electronic prototype for the Scale of the Assessment and Rating of Ataxia (SARA), broadly used in neurology. A scaled-down version of the Human Phenotype Ontology was automatically extracted and used as backbone to normalize the content of the SARA through clinical archetypes. The knowledge required to exploit reasoning on the SARA data was modeled as separate information-processing units interconnected via the defined archetypes. Based on this approach, we implemented a prototype named SARA Management System, to be used for both the assessment of cerebellar syndrome and the production of a clinical synopsis. For validation purposes, we used recorded SARA data from 28 anonymous subjects affected by SCA36. Our results reveal a substantial degree of agreement between the results achieved by the prototype and human experts, confirming that the combination of archetypes, ontologies and guidelines is a good solution to automate the extraction of relevant phenotypic knowledge from plain scores of rating scales.

Full text

TESE DE DOUTORAMENTO Implementing electronic scales to support standardized phenotypic data collection - the case of the Scale for the Assessment and Rating of Ataxia (SARA) HAITHAM MAAROUF ESCOLA DE DOUTORAMENTO : Ciencias e Tecnoloxías da USC PROGRAMA DE DOUTORAMENTO EN : Investigación en Tecnoloxías da Información SANTIAGO DE COMPOSTELA 2018 DECLARACIÓN DO AUTOR DA TESE Implementing electronic scales to support standardized phenotypic data collection - the case of the Scale for the Assessment and Rating of Ataxia (SARA) D. Haitham Maarouf Presento miña tese, seguindo o procedemento adecuado ao Regulamento, e declaro que: 1) A tese abarca os resultados da elaboración do meu traballo. 2) No seu caso, na tese se fai referencia as colaboracións que tivo este traballo. 3) A tese é a versión definitiva presentada para a súa defensa e coincide ca versión enviada en formato electrónico. 4) Confirmo que a tese non incorre en ningún tipo de plaxio de outros autores nin de traballos presentados por min para a obtención de outros títulos. En Santiago De Compostela,18 de Decembro de 2017 Firmado. Haitham Maarouf AUTORIZACIÓN DO DIRECTOR / TITOR DA TESE Implementing electronic scales to support standardized phenotypic data collection - the case of the Scale for the Assessment and Rating of Ataxia (SARA) Dna. Prof. María Jesús Taboada Iglesias Dna. Dr. María Jesús Sobrido Gomez INFORMAN: Que a presente tese, correspóndese co traballo realizado por D. Haitham Maarouf, baixo a nosa dirección, e a utorizamos a súa presentación , considerando que reúne os r equisitos esixidos no R egulamento de Estudos de Doutoramento da USC, e que como director desta non incorre nas causas de abstención establecidas na Lei 40/2015. En Santiago de Compostela,18 de Decembro de 2017 Asdo. María Jesús Taboada Iglesias Asdo. María Jesús Sobrido Gomez I DEDICATION This thesis is dedicated to the memories of my father and uncles: Prof. Adel Maarouf, Abdul Rahim and Mahmoud. I also dedicate it to my mother, wife, brothers, sisters and Nabil for their endless love, support and encouragement. III ACKNOWLEDGMENT Writing an acknowledgment is the hardest part of writing a thesis. This work would not have been possible without the support and encouragement of several people. Please accept my sincere gratitude, irrespective of your name appearing in the acknowledgment. First and foremost, I would like to sincerely thank my supervisor, Prof. María Jesús Taboada Iglesias for her great support, untiring help, patience, commitment of time, guidance and being an exceptional supervisor through my long Ph.D. process. There are no words that can describe how much I am grateful to her. Thank you, Prof. Taboada, for giving me that much needed morale boost, when things seemed to be looking gray, and there is no light at the end of the tunnel. I am also indebted to her for expanding my knowledge, enhancing my research skills and helping me to publish a paper. I extend my thanks and gratitude to Dr. María Jesús Sobrido Gomez who provided me with the patient Data and all the required medical knowledge related to the Spinocerebellar Ataxia type 36 (SCA36) disease and the Scale for the Assessment and Rating of Ataxia (SARA). I also thank Dr. Manuel Arias and Dr. Ángel Sesar for participating in the validation process to test the validity of the developed tool. This Ph.D. thesis was supported by the National Institute of Health Carlos III [grant no. FIS2012-PI12/00373: OntoNeurophen], FEDER for national and European funding; PEACE II, Erasmus Mundus Lot 2 Project [grant no. 2013-2443/001-001-EMA2]. Special thanks go to the European Commission and Erasmus Mundus program. Furthermore, I express my deepest gratitude to Prof. Hatem Alamy, the Chairman of the Modern University for Business and Science (M.U.B.S), and to Dr. Bassem Kaissi for giving me the X the extraction of relevant phenotypic knowledge from plain scores of rating scales. KEYWORDS Rating scales, GDL, Human Phenotype Ontology, Clinical Archetypes, SARA, Spinocerebellar ataxia XI RESUMEN AMPLIADO A partir de la evidencia de que los pacientes con un mismo diagnóstico pueden presentar diferentes manifestaciones clínicas y, por tanto, reaccionar de forma distinta a la misma intervención, la medicina personalizada reconoce que cada paciente es único y por tanto debe ser tratado de forma individualizada. A partir de los noventa, con el impulso de la genómica y otras ciencias ómicas, se reconoce la importancia de estratificar a los pacientes, es decir, de clasificarlos en grupos similares biológicamente, con el objetivo de conseguir la respuesta óptima a las intervenciones planificadas en cada uno de los subgrupos. Adicionalmente, diversos estudios ya han demostrado que la identificación de estos subgrupos requiere analizar los datos ómicos junto con descripciones computacionales de calidad del fenotipo del paciente. En dominios clínicos, una anomalía fenotípica es una divergencia de la morfología, la fisiología o el comportamiento normal del paciente. Por tanto, el éxito en la estratificación de pacientes también dependerá, en gran medida, de los recursos computacionales disponibles para adquirir y representar el fenotipo de los pacientes, y para integrarlos adecuadamente con la información ómica y de imagen médica. Las ontologías, como artefactos informáticos del campo de la Inteligencia Artificial, facilitan la organización y armonización de la información compleja y heterogénea, proporcionando facilidades de consulta e inferencia lógica sobre los datos almacenados. En los últimos años, una de las ontologías que ha experimentado el avance más importante en su uso para estudiar el diagnóstico clínico en enfermedades con base genética es la Human Phenotype Ontology (HPO). A la vez, diferentes consorcios internacionales han estado desarrollado modelos de datos que promueven la estandarización en la adquisición de los datos de pacientes, tales como ISO 13606, HL7 CDA, NINDS CDE e Intermountain Healthcare. El uso de dichos modelos es crucial para comparar resultados entre diferentes estudios, XII integrar información entre diferentes aplicaciones, e implementar sistemas de ayuda a la decisión. Los esfuerzos de estos consorcios han dado lugar a especificaciones formales y computables del contenido clínico, que se conocen como arquetipos clínicos. Dichas especificaciones permiten representar, de forma consensuada, cualquier estructura de datos de la historia clínica del paciente, incluyendo tanto las definiciones (en forma de restricciones sobre las estructuras), como las interrelaciones entre dichas estructuras. Mientras que los arquetipos clínicos estandarizan la captura de los datos clínicos del paciente, las ontologías de fenotipos estandarizan su significado e interpretación. Hay que tener en cuenta que las descripciones fenotípicas de los pacientes (que aparecen, por ejemplo, en los informes clínicos textuales) están en un nivel de abstracción más elevado que los datos de paciente recopilados a través de cuestionarios o pruebas clínicas, lo que provoca impedance mismatch. Una posible forma de solucionar el desfase entre la estandarización de los datos clínicos y la de fenotipos es utilizar las facilidades del razonamiento basado en ontologías sobre los datos recopilados con arquetipos. Sin embargo, a día de hoy, esta opción es todavía un reto. Aunque las especificaciones de los arquetipos clínicos proporcionan formas de expresar alineamientos (mappings) de los ítems del arquetipo a los conceptos de las ontologías, no existen recursos que faciliten el razonamiento basado en las ontologías alineadas. Hasta el momento, se han propuesto varias alternativas que abarcan la conversión de arquetipos al lenguaje de ontologías OWL-DL (Ontology Web Language-Description Language) o la definición de alineamientos intensivos en conocimiento desde las fuentes de datos a los arquetipos clínicos. Sin embargo, estas propuestas siguen sin proporcionar una tecnología sencilla que facilite el razonamiento. Por otra parte, el alineamiento de los datos basados en arquetipos clínicos con las ontologías no es una tarea trivial, y prueba de ello es que la mayoría de los arquetipos públicos no contienen dichos alineamientos. Siguiendo la aproximación estándar de desarrollo de arquetipos clínicos, el alineamiento ontológico se suele realizar en las últimas etapas de modelado. Ello conlleva un esfuerzo extra, en parte debido al gran tamaño de las ontologías. Además, el diseño de arquetipos clínicos de XIII forma separada de las ontologías puede conllevar discrepancias muy elevadas en el significado de los ítems clínicos. Las escalas clínicas representan un recurso importante para la recopilación de datos estandarizados. Si bien las escalas clínicas se usan en todas las disciplinas médicas, son especialmente relevantes en especialidades que manejan variables fenotípicas complejas, como la neurología. Su uso incrementa la calidad de los datos, al reducir la subjetividad en las descripciones fenotípicas, y simplifica el diseño de los protocolos de recogida de datos en los estudios clínicos. Generalmente, las escalas clínicas valoran una o varias dimensiones clínicas mediante un conjunto de ítems y proporcionan una puntuación global. La hipótesis de partida de esta tesis doctoral es que reducir todo el contenido de la información recopilada a través de una escala de valoración a un único número (puntuación total) puede conllevar a la pérdida de información clínica relevante. El objetivo de esta tesis doctoral es demostrar que es posible realizar interpretaciones clínicas de forma automática sobre los datos recopilados por las escalas clínicas, de la misma manera que un experto clínico lo hace. Dichas interpretaciones automáticas pueden facilitar la evaluación médica, proporcionar ayuda para la escritura de informes de pacientes y la decisión médica. Para alcanzar el objetivo propuesto hemos desarrollado una aproximación novedosa orientada a modelar e implementar escalas clínicas electrónicas en el dominio de la neurología. Dicha aproximación busca representar computacionalmente tanto el contenido como la interpretación clínica de los datos recopilados. Para ello, se hace uso de los estándares de registros electrónicos de pacientes y de las tecnologías web semánticas. La principal innovación de nuestro trabajo ha sido el desarrollo de una aplicación que va más allá de una simple calculadora, con la incorporación del conocimiento clínico requerido para interpretar la información recopilada y generar automáticamente los correspondientes informes de pacientes. Los beneficios de nuestra solución innovadora son la provisión de estandarización clínica no sólo durante la recogida de los datos sino también durante la interpretación clínica de los hallazgos de pacientes, así como la producción de facilidades para la generación automática de informes, que liberan al médico de dicha tarea. XIV En este trabajo, optamos por abordar la Escala para la Evaluación y Clasificación de la Ataxia (SARA), un instrumento bien validado para evaluar la presencia y la gravedad de la ataxia cerebelosa. Esta escala tiene un uso muy extendido y ha sido aplicada por nuestro grupo para la evaluación de la ataxia espinocerebelosa tipo 36 (SCA 36). Para facilitar el alineamiento ontológico y evitar grandes discrepancias semánticas entre los arquetipos clínicos y las ontologías, se ha propuesto un método novedoso basado en la suposición de que el diseño de arquetipos debería ser soportado por ontologías. Por otro lado, la interpretación clínica de los datos recopilados por una escala de valoración requiere diferentes tipos de información para su automatización: datos para registrar (es decir, el contenido de la escala clínica), conocimiento sobre el significado de los términos en la escala (es decir, conocimiento terminológico), conocimiento de procedimientos para comprender el significado de los puntajes (que se pueden expresar fácilmente mediante guías clínicas) y conocimiento ontológico para deducir las anomalías fenotípicas de los pacientes. Elegimos utilizar una combinación de lenguajes de guías clínicas, arquetipos clínicos y ontologías para abordar los desafíos del modelado de la escala clínica. Las preguntas de investigación abordadas en este trabajo son: I) ¿La combinación de GDL (Guideline Definition Language), arquetipos clínicos y ontologías es adecuada para la descripción e interpretación de los datos colectados vía la escala SARA?, y II) ¿Es posible lograr la integración de estas herramientas computacionales para modelar e interpretar eficientemente la información clínica proporcionada por la escala clínica? Nuestro enfoque de modelado se basa en cuatro pasos principales: I) Creación de una versión reducida del HPO, mediante la extracción de los módulos de ontología relevantes para la escala SARA, II) anotación de las descripciones de texto libre de la escala clínica con los módulos de ontología, III) desarrollo de dos tipos de arquetipos clínicos (observación y evaluación), y IV) definición de unidades de procesamiento de información para expresar el sistema de apoyo a la interpretación clínica. El modelado de la escala SARA involucró un nivel de datos - representación de los ítems de SARA - y un nivel de conocimiento - referido a la estrategia para calcular el puntaje total y la XV interpretación del fenotipo. Los arquetipos se usaron para modelar el nivel de datos, mientras que GDL y OWL (Ontology Web Language) se usaron para modelar el nivel de conocimiento. Esta representación tenía restricciones, ya que los modelos openEHR son compatibles con GDL, pero no dan mucho soporte para OWL y el razonamiento relacionado. Para cerrar el gap entre los arquetipos clínicos y la ontología, se definieron alineamientos que facilitaron la traducción de las instancias de arquetipo al conjunto de datos OWL. Para la extracción del módulo de la ontología relevante a la escala, comenzamos revisando y recopilando documentos de texto que describían la escala. Luego anotamos las fuentes extraídas con los términos de la ontología HPO, utilizando el OBO Annotator, un sistema de anotación de conceptos fenotípicos desarrollado en el grupo. A continuación, mapeamos los ítems de la escala con las anomalías fenotípicas estándar del paciente descritos en su propia escala. Para el diseño del arquetipo, hicimos una diferencia entre los datos recopilados por la escala y que se calculan directamente utilizando la información proporcionada por la escala (ítems, dimensiones clínicas, puntajes individuales y puntaje total), y los datos que no están directamente disponibles en el escala. Teniendo en cuenta esta distinción, decidimos utilizar dos tipos de arquetipos clínicos: Observación para normalizar los datos proporcionados por la escala, y Evaluación, para estandarizar las interpretaciones sobre los datos descritos en la escala. El primero permite representar el contenido de la escala, mientras que el segundo registra las interpretaciones clínicas derivadas de la escala. Los dos tipos de arquetipos clínicos se desarrollaron utilizando el editor de arquetipos proporcionado por OpenEHR. Para modelar el arquetipo de observación, después de agregar los metadatos, estructuramos el contenido de la escala de acuerdo con este arquetipo, es decir, mediante una estructura de árbol con elementos para las tres dimensiones clínicas (el conjunto de elementos, la puntuación total y la fecha de las observaciones) y los valores para las puntuaciones individuales. Luego, definimos los elementos con tipos de datos adecuados, descripciones, comentarios, detalles, ocurrencias, restricciones y valores posibles. Asignamos cada elemento a la clase de ontología correspondiente de la versión reducida del HPO, utilizando las anotaciones obtenidas. Finalmente, estructuramos y organizamos los ítems de acuerdo con la XVI ontología. El arquetipo de evaluación se modeló para registrar tres tipos de interpretaciones: I) Interpretaciones deducidas directamente de elementos individuales o de la puntuación global, II) Interpretaciones deducidas directamente de valores específicos para un elemento individual, y III) Interpretaciones sobre las anomalías del paciente asociadas con las dimensiones clínicas de la escala. Con respecto a las unidades de procesamiento de información (o inferencias) que interpretan automáticamente los datos de la escala se distinguieron: • Las funciones de cálculo, que expresan las operaciones matemáticas para calcular las puntuaciones de los ítems individuales y la puntuación global de la escala. • La valoración del grado de deterioro que gradúa la intensidad de los elementos expresados en el arquetipo de observación, que generalmente está dado por niveles de corte. • La evaluación del síndrome cerebeloso que está basada en la puntuación global de la escala, y se calcula aplicando varias reglas heurísticas propuestas por el experto en neurología. • La generación de la sinopsis clínica que describe las características fenotípicas (signos) que acompañan al síndrome cerebeloso. Para la validación, utilizamos registros de datos de 28 sujetos anónimos con Ataxia Espinocerebelosa Tipo 36 (SCA36). Todos los pacientes fueron examinados siguiendo la SARA. La evaluación se llevó a cabo en tres pasos: I) Cálculo e interpretación de la puntuación total: completamos los datos de cada paciente para obtener la interpretación. El sistema dedujo automáticamente 1) la intensidad de cada ítem en la escala, 2) la gravedad del síndrome cerebeloso, 3) la gravedad de la ataxia troncal y XVII 4) la gravedad de la ataxia apendicular en los lados derecho e izquierdo. II) Interpretación por dos neurólogos independientes: se envió el mismo conjunto de datos a dos neurólogos, pero se agregaron las puntuaciones totales calculadas para cada paciente. Los neurólogos utilizaron su experiencia en ataxia para determinar la gravedad del síndrome cerebeloso a partir de la puntuación total y la gravedad de la ataxia troncal y apendicular, si está presente, de las puntuaciones individuales. III) Comparación de resultados entre el sistema y los expertos humanos: para validar el sistema, se realizaron los siguientes pasos: 1) creamos una hoja usando SPSS, 2) importamos las interpretaciones del sistema y de los neurólogos a SPSS, y 3) realizamos la prueba Weighted Kappa 12 veces para medir la fuerza del acuerdo entre el sistema y cada neurólogo, y entre los dos neurólogos. Para demostrar la funcionalidad de nuestro enfoque, desarrollamos un prototipo llamado "SMS" (Sistema de gestión SARA), que permite la gestión de los datos del paciente del síndrome cerebeloso. Usamos JAVA como lenguaje de programación y MySQL como un sistema de administración de bases de datos para almacenar toda la información necesaria. La arquitectura de "SMS" está estructurada en tres capas: "Persistencia", "Operación" e "Interfaz". La capa de persistencia utiliza un sistema de gestión de base de datos para almacenar toda la información requerida (los arquetipos clínicos, la ontología, los datos del paciente recogidos usando arquetipos clínicos e inferidos por el sistema, e información adicional). Para poblar la base de datos con la información del arquetipo, se usó el analizador ADL (Archetype Description Language) para generar un árbol de dependencias de términos y luego crear un archivo XML con las instancias de los arquetipos. La capa de operación incluyó todas las unidades de procesamiento de información que son responsables de ejecutar las funciones de cálculo, la evaluación del grado de deterioro, la evaluación del síndrome cerebeloso y la sinopsis clínica. La versión actual de GDL no proporciona una API de Java ni ningún mecanismo para manipular las reglas definidas con el editor. Por lo tanto, la única solución era volver a escribir las reglas en un motor de reglas (como Drools o Clips) o implementar directamente en Java (decidimos esta XVIII segunda solución). Sin embargo, la inferencia de fenotipos se ha podido implementar razonablemente utilizando la API de OWL y los alineamientos entre los ítems de los arquetipos y la ontología OWL. Los alineamientos son útiles para crear automáticamente individuos OWL de la clase definida a partir de las instancias de arquetipo almacenadas en la base de datos. La capa de interfaz consiste en varios formularios de entrada y salida. La forma principal de la herramienta contiene dos menús: el menú de ontología y el menú SARA. El primer menú permite a los usuarios modificar la estructura de los módulos de ontología HPO. Se pueden agregar, actualizar y eliminar términos en los módulos de ontología. El formulario proporciona la posibilidad de generar automáticamente un gráfico jerárquico que muestra todas las clases disponibles y sus relaciones jerárquicas. El menú SARA consta de dos submenús: 1) Observación, donde los neurólogos pueden seleccionar pacientes e ingresar los valores de los elementos evaluados definidos en la escala SARA y modelados en el arquetipo de observación, y 2) Evaluación, que proporciona tres características principales: • Una tabla que muestra todos los códigos de anomalías fenotípicas que tiene un paciente. El conjunto de estas anomalías se deduce automáticamente. • Un informe textual que resume el estado de un paciente según la recopilación de datos. El informe describe la gravedad del síndrome cerebeloso y una breve sinopsis de las anomalías fenotípicas. • Un gráfico de evaluación que visualiza todas las anormalidades del paciente Para facilitar el intercambio semántico interoperable de datos, se pueden generar tanto los datos XML recopilados por SARA (usando los esquemas que cumplen con el arquetipo de observación) como los datos inferidos por la aplicación (usando los esquemas que cumplen con el arquetipo de evaluación). La validación del sistema se basó en los valores de Kappa ponderado obtenidos de las 12 pruebas realizadas. Se obtuvieron unos valores Kappa en el rango entre 0.65 y 0.93. XIX En la actualidad, los trastornos atáxicos todavía no tienen una terapia farmacológica exitosa, y los pacientes sufren la inevitable progresión de la enfermedad degenerativa. El objetivo de las escalas clínicas es facilitar la comprensión de la historia natural de los trastornos atáxicos y evaluar adecuadamente la eficacia de los fármacos en los ensayos clínicos. En la presente tesis doctoral, nos centramos en proporcionar soporte automático para la interpretación clínica de los datos recopilados usando la escala SARA. El trabajo contribuye a una mejor comprensión de cómo los arquetipos clínicos, las guías clínicas y las ontologías se pueden combinar para modelar e implementar una escala de valoración en el dominio de la neurología. Hay varias contribuciones en esta investigación. Una contribución es un enfoque basado en ontologías para modelar los dos arquetipos clínicos propuestos, lo que reduce el esfuerzo necesario para crear alineamientos y evita grandes discrepancias semánticas entre los arquetipos modelados y los módulos de ontología. Otra contribución es la separación clara y explícita entre los componentes estándar de la escala relacionados con el contenido (es decir, ítems, dimensiones clínicas y puntajes), que han sido modelados usando un arquetipo de observación, y las interpretaciones clínicas de estos componentes, que han sido normalizadas por un arquetipo de evaluación para uso local. Finalmente, una contribución clave es la identificación clara de todos los diferentes tipos de conocimiento requeridos para interpretar los datos recopilados por la escala y su modelado como unidades de procesamiento de información que se comunican entre sí a través de los dos arquetipos definidos, proporcionando un mecanismo simple de combinación ontología y razonamiento basado en reglas. HAITHAM MAAROUF 2 manifestations. Examples of latent variables (or clinical dimensions) in neurological diseases include the quality and intensity of a tremor, the degree of gait imbalance or cognitive performance. These latent variables are assessed through a set of clinical questions (named statements or items) [7]. Each statement may have multiple ordered response options, for which an ordinal number (score) is assigned. The total score for the global clinical dimension is usually obtained adding up all individual scores for each statement. Well-known examples are the Mini-Mental State Examination (MMSE) [8], a 30-point survey used to measure cognitive impairment, or the Glasgow Coma Scale (GCS) [9], which is used to assess coma and impairment of consciousness. Using these instruments entails many advantages: improved data quality by reducing subjectivity during measurement, simplified design of data collection, and data harmonization across different clinical studies. Hence, computational implementation of rating scales offers a major chance for data quality improvement and harmonization across different clinical studies. Additionally, electronic rating scales are an important resource to support automated inference of patient phenotype from the data collection. Usually, rating scales grade several clinical dimensions, each of them assessed by different items. For instance, in addition to the movement disorder (i.e., disease state), the Unified Parkinson’s Disease Rating Scale (UPDRS) [10] assesses other clinical subdimensions (such as mental state, complications of treatment and activities of daily life) via 42 questions providing a total score that grades the progression of the disease. However, reducing all content of a rating scale to a unique number (score) inevitably causes the loss of some phenotype information implicitly collected by the scale. For example, in patients with the same total UPDRS score this number could be due to different clinical dimensions, therefore actually be quite different clinically. A more precise inference of the patient phenotype from the sub-scores would facilitate the automated codification of the clinical abnormalities for further analysis. Additionally, it would decrease subjectivity during the score interpretation and facilitate medical evaluation, report writing, and clinical decision-making. In this work, we chose to address the Scale 1 Introduction 3 for the Assessment and Rating of Ataxia (SARA)1 [11], a wellvalidated instrument to evaluate the presence and severity of cerebellar ataxia [12]. This scale is broadly used and it has been applied by our group in a research on the spinocerebellar ataxia type 36 (SCA 36). Formal description of rating scales using systematic clinical information models promotes computational data standardization comparison of results across studies [13], integration of information from different sources and medical records, and implementation of decision support systems. Different international projects and consortia have been developing standardized data models for clinical research and electronic health records, such as ISO 13606 [14], HL7 CDA [15], openEHR [16], NINDS CDE [3] and Intermountain Healthcare [17]. The commonality among these approaches is that they are focused on computable and formal specifications of clinical content in the form of information models known as clinical models or archetypes. These clinical models supply standardized data structures to represent the clinical statements included into rating scales. Additionally, mechanisms to link the clinical statements to classes of some standard terminology or ontology are provided. Hence, representing rating scales using clinical archetypes/models and terminologies/ontologies aims to get both clinical and computational harmonization of data collections. While both clinical archetypes and ontologies seek to structure the patient information, according to the needs of clinical research, however their perspectives are often dissimilar. In archetypes, the clinical statements that must be entered at the same time are aggregated together. Archetypes model the information to mirror patient records. For example, the items paraparesis and facial palsy were recorded together into the archetype Stroke Scale Neurological Assessment [18], which is available on the Clinical Knowledge Manager (CKM) [19] provided by the OpenEHR Foundation. Ontologies, on the other side, aim at representing the meaning of those clinical statements. Classes in the Human Phenotype Ontology [20] are arranged in a hierarchical structure of phenotypic abnormalities. For example, both paraparesis and facial palsy are represented as abnormalities of the nervous system 1 http://www.ataxia-study-group.net/html/about/ataxiascales/sara/SARA.pdf HAITHAM MAAROUF 4 in the HPO. However, the former is represented as an abnormality of the physiology, whereas the second one as an abnormality of the morphology. This ontological distinction cannot be reflected into the clinical archetype and however it is valuable to interpret the patient status. Thus, integrating ontologies with clinical archetypes would not only provide a static knowledge store, but also a dynamic resource to automatically infer patient phenotype and standardize data collection. Specifically, Braun et al. [21] developed four clinical archetypes (Timed 25-Foot Walk, Nine Hole Peg Test, Paced Auditory Serial Addition Test, and MSFC Score) [22-25] to represent a rating scale for the assessment of Multiple Sclerosis (MS) patients consisting of three neurological tests. They applied a standard archetype development approach, consisting of: analyzing the clinical domain and requirements, identifying the archetype contents and their organization from different sources (literature, record forms, etc.), selecting the archetype type, structuring the content according to the archetype type, and filling the parts of the archetype with the content. With this approach, terminology/ontology mapping is carried out during the later steps of archetype building, when the model is almost complete. At this stage, the effort required to create the mappings between archetype terms and ontology entities is substantial [26], due in part to the large size of the ontologies [27]. Furthermore, designing clinical archetypes separately from ontologies may lead to major discrepancies in the meaning of clinical statements. As a result, ontology mappings are not common in the openly accessible archetypes of the repositories. Nevertheless, mapping clinical archetypes and ontologies/terminologies is key to get semantic interoperability among different data sources. With the aim of facilitating the ontology mapping and preventing large semantic discrepancies between clinical archetypes and ontologies, we propose to reorganize the classical methodology. At the heart of our methodology is the assumption that the archetype design should be supported by ontologies in those clinical situations where it is expected that archetype contents be logically organized. Furthermore, automated assessment and evaluation upon the information represented by clinical archetypes is still an open research issue. Clinical archetypes/models aim to record standardized 1 Introduction 5 definitions of the clinical data in electronic medical records [28]. In the case of rating scales, this structure matches to its content (i.e., clinical dimensions, items and scores). Additionally, clinical archetypes/models offer the possibility to normalize clinical data by mapping them to formal ontologies. However, exploiting reasoning on this clinical knowledge is still limited and is another challenge. The Guideline Definition Language (GDL) [29] is a formal language recently authorized by openEHR for expressing decision support logic by a rule-based declarative strategy. Anani et al. [30] used GDL to implement knowledge on contraindications for using thrombolytic treatment in patients suffering acute stroke, and Lin et al. [31] to implement ten electronic clinical practice guidelines in the chronic kidney disease. As Anani et al. [30] have emphasized, GDL provides a rule authoring language aimed to represent declarative knowledge that can be shareable and standardized, as it is supported by OpenEHR clinical models. However, GDL does not yet provide much support for ontologies and related reasoning. Another alternative is to transform clinical archetypes/models into OWL-DL (Ontology Web Language – Description Language) [32-34]. Following this approach, ontology reasoners, such as Pellet2, Hermit 3 or Fact++4, can be used to both check the OWL-based archetype consistency [35], and use the ontology to draw any inferences on data collection. Additionally, having archetypes, ontologies and knowledge descriptive (inference rules) under the same syntactic structure provides support for interoperability of rule-based mechanisms [28, 33, 36]. However, having two separate, independent versions of the same standard model, one of them in the language of the model itself (ADL-Archetype Definition Language) [37] and the second one in OWL format, makes maintenance more difficult. Furthermore, procedural knowledge as the sum (counting or any complex mathematical calculation) of the scores in rating scales cannot be simply represented in OWL. An interesting alternative proposed by Mugzach et al. [36] to perform a particular counting (named k-of-N counting by the authors) in OWL was to develop a plugin meeting the specific requirements. However, different calculation functions would then require implementing specific plug-ins. Other 2 https://www.w3.org/2001/sw/wiki/Pellet 3 http://www.hermit-reasoner.com/ 4 http://owl.cs.manchester.ac.uk/tools/fact/ HAITHAM MAAROUF 6 researchers defined knowledge-intensive mappings from the data sources to openEHR archetypes [38, 39]. They distinguished between data-level and knowledge-level processing tasks. The former included calculation functions specified in the mappings and directly run on archetype data. The latter covered classification tasks defined using OWL classes with sufficient conditions. An integrated Personal Health Record is an alternative option proposed to simplify data integration and clinical decision-making [40]. 1.1 RESEARCH OBJECTIVE The main goal of our work was to develop an electronic rating scale in the clinical domain of the neurology representing both the content and the interpretation of the SARA, using the Electronic Health Record (EHR) standards and taking advantage of semantic web technologies to automatically interpret the phenotype from data collected by a rating scale. To that goal, specific objectives are:  To computationally represent the knowledge covered by the rating scales.  To define the phenotypes in a computationally accessible way.  To map the patients’ clinical data gathered from the rating scales to phenotypes.  To efficiently and computationally model the clinical information content provided by the rating scales using EHR standards.  To computationally define the interpretation of the rating scales using EHR standards. 1.2 RESEARCH QUESTIONS 1) Is the combination of GDL, openEHR clinical archetypes and ontologies suitable for the description of all knowledge and content covered by the SARA? 1 Introduction 7 2) Is it possible to achieve integration of these computational tools to efficiently model the clinical information provided by the rating scale? 1.3 RESEARCH CONTRIBUTION The contributions of this thesis are:  In order to facilitate ontology mapping and prevent large semantic discrepancies between clinical archetypes and ontologies. The classical methodology proposed by Braun et al. [21] was enhanced and we developed a novel method based on the assumption that archetype design should be supported by ontologies in those clinical situations where the archetype contents are logically organized.  The current HPO does not cover all the details of needed neurodegenerative phenotypes. New ontology modules relevant to the SARA were developed using OWL.  We demonstrated how the openEHR clinical archetypes Observation and Evaluation could be used to model the content of a rating scale and to record the clinical interpretations and phenotypic abnormalities.  GDL and OWL were effectively used to express all the required knowledge to understand the meaning and the scores of terms in the scale, and to deduce patient phenotypes.  We chose to use a combination of GDL, openEHR clinical archetypes and ontologies to address the challenges of modeling the rating scale. 1.4 STRUCTURE OF THE THESIS The rest of the thesis is structured as follows:  Chapter 2 describes the background, providing an overview of the main components of rating scales, as well as the models, languages and ontologies used in this work. HAITHAM MAAROUF 8  Chapter 3 describes the methodology proposed in this work to model electronic rating scales, with the specific example of an application to the SARA.  Chapter 4 presents the results of implementing and validating our method.  Chapter 5 highlights the implications of this work.  Finally, conclusions, limitations and future work are provided in Chapter 6. 9 2 BACKGROUND In this chapter, we present an overview about the components, tools and technologies used to develop this research work. The first part of this review is focused on the clinical domain, including information on rating scales and their components, the specific scale for the assessment and rating of ataxia (SARA) and the rare syndrome named spinocerebellar ataxia type 36 (also known as Ataxia da Costa da Morte). It also covers the technologies used in this thesis: the clinical data models or archetypes, OWL ontologies and the Human Phenotype Ontology, as well as the available openEHR formal language to implement computerized clinical decision support system. 2.1 CLINICAL RATING OR ASSESSMENT SCALES Most rating scales used in neurology are ordinal scales. They provide a set of items needed to quantify the severity of motor, sensitive, sensory, cognitive function or quality of life, whereby the rater has to assign a value, usually numeric, to the graded items. Thus, rating scales rank patients in degrees of disability according to certain external criteria. Some of them assay only one clinical attribute or item (single-item scale), such as the Modified Rankin Scale (MRS) [41]; whereas others consist of several items (multiple-item scale). In some cases, all items in the scale assess the same dimension (e.g., motor deficit), whereas in other cases, the scale consists of several multipleitem sub-scales, like UPDRS [10]. Usually, rating scales combine values of individual patient traits (items scores) into a total score, which measures the variable computed by the set of items. For example, the MMSE is a questionnaire with a total score of 30 that collects items related to different traits: orientation to time (5-score) and to place (5score), attention and calculation (5-score), language (2-score), etc. In the MMSE [8], the total score is derived by summation of all items, while in other cases the scoring may involve complex calculations. Different components can be distinguished in a rating scale (Fig. 2.1). HAITHAM MAAROUF 10 Fig. 2.1 Components of rating scales 2.1.1 Components of Rating Scales The following components can be distinguished in clinical rating scales: items of the scale, response options and item scores, subscales, total score, calculation function, and interpretation of the score. 2.1.1.1 Items Items of the scale are the different questions assessing a specific clinical dimension into the rating scale. For example, the UPDRS a scale that assesses disability due to Parkinson’s disease, contains 42 different questions grouped into four clinical dimensions: I) Cognitive, behavioral and mood (4 items), II) Activities of daily living (13 items), III) Motor performance (14 items), and IV) Complications of treatment (11 items). 2.1.1.2 Response options and scores Response options and scores are the possible values that raters can assign to items to quantify them. These values could belong to either Components of rating scales Subscales Items Response options Item Score Subscales score Total Score Calculation function Interpretation 2 Background 11 ordinal level, interval or ratio scales. For instance, the MRS ranges from 0 (No symptoms) to 6 (Death) to quantify the disability. 2.1.1.3 Subscales Subscales are the clusters of items, which together measure a particular clinical dimension. For example, in the MMSE, the question ‘What is the date?’, including questions on year, season, month, date and day of the week, let the rater assess the clinical dimension ‘orientation to time’. 2.1.1.4 Total Score Total score is the value assessing the global clinical dimension measured by the rating scale. It is calculated once the different items have been evaluated. 2.1.1.5 Calculation Function Calculation function is the set of mathematical operations performed on the item scores to calculate the total score. The sum of the item scores is the most usual approach for calculating the total score. However, alternative procedures, such as mean score or standardization to a reference population, are also frequent. 2.1.1.6 Interpretation of the total score Interpretation of the total score is the explanation of the measurement result, which is usually left open to the rater. In some cases, a simple standard procedure is attached to the rating scale. For example, in the MMSE, there are four criteria to qualify the degree of impairment: ’25-30 = questionably significant’, ’20-25 = mild’, ’10-20 = moderate’, and ’0-10 = Severe’. 2.2 CLINICAL ARCHETYPES OpenEHR developed a two-level approach to make a separation between the semantic of information and knowledge into two levels: HAITHAM MAAROUF 18 Fig. 2.5 displays the APGAR Observation archetype [50] that allows to record data after one minute, 2 minute, etc. Fig. 2.5 Modeling recorded data over time in APGAR archetype. 2.4 HUMAN PHENOTYPE ONTOLOGY Clinical archetypes use term mapping to standard terminologies/ontologies with the aim of normalizing the clinical data used in the model definition. Additionally, the use of a standard ontology provides the capability to automatically infer patient clinical phenotypes from data collected using the rating scale. The Human Phenotype Ontology (HPO) [20] delivers a structured and standardized vocabulary for phenotypic abnormalities encountered in human hereditary and other diseases. It is accessible at [5] and as of November 2017, it contains 13165 terms (classes) with 16794 is_a (is the same as subclassOf) relationships among those classes. Each class in the HPO describes an individual phenotypic abnormality. The is_a relationship describes the subclass-superclass relationships between the HPO classes. For example, dysarthria is_a neurological speech impairment. The is_a relationship is transitive, implying that if spastic dysarthria is_a dysarthria, which is_a neurological speech impairment, then spastic dysarthria is also a neurological speech impairment. As represented in OBO [51] format, each class can have up to 16 attributes 2 Background 19 (id, name, alternative ids, definition, synonym, references, is_a, etc.). Fig. 2.6 displays the HPO class Dysarthria and its attributes. Fig. 2.6 The HPO class Dysarthria and its attributes. 2.4.1 The sub-ontologies of the HPO The HPO has five sub-ontologies; Clinical Modifier, Mortality/Aging, Mode of Inheritance, Frequency, and Phenotypic Abnormality (Fig. 2.7) Fig. 2.7 The sub-ontologies of the Human Phenotype Ontology 2.4.1.1 Clinical Modifier The Clinical Modifier sub-ontology contains classes that describe typical modifiers of clinical symptoms. For example Severity, the Pace of progression, the Phenotypic variability or the Onset. It comprises terms such as Profound, Severe, Moderate, Mild, Profound, Childhood [Term] id: HP:0001260 name: Dysarthria alt_id: HP:0002327 def: "Dysarthric speech is a general description referring to a neurological speech disorder characterized by poor articulation. Depending on the involved neurological structures, dysarthria may be further classified as spastic, flaccid, ataxic, hyperkinetic and hypokinetic, or mixed." [HPO:curators] synonym: "Difficulty articulating speech" EXACT layperson [] synonym: "Dysarthric speech" EXACT [] xref: MSH:D004401 xref: SNOMEDCT_US:8011004 xref: UMLS:C0013362 is_a: HP:0002167 ! Neurological speech impairment All Clinical Modifier Mortality/Againg Mode of inheritance Frequency Phenotypic abnormality HAITHAM MAAROUF 20 onset, Variable progression rate, or Variable expressivity. Fig. 2.8 shows the subclasses of the class Severity which is a subclass of the Clinical modifier class. Fig. 2.8 The subclasses included into the class Severity 2.4.1.2 Mortality/Aging This sub-ontology describes Time of death and includes classes such as Death in early adulthood, Death in adolescence or Sudden death. 2.4.1.3 Mode of Inheritance This relatively small sub-ontology is intended to describe the mode of inheritance and contains terms such as Autosomal dominant inheritance, Gonosomal inheritance, Multifactorial inheritance, etc. 2.4.1.4 Frequency This sub-ontology defines the frequency with that patients do show a particular clinical feature. It comprises terms such as Very frequent, Very rare, Excluded, etc. 2.4.1.5 Phenotypic abnormality This is the core sub-ontology of the HPO and includes definitions of clinical abnormalities. It contains classes such as Abnormality of Clinical modifier Severity Borderline Mild Moderate Severe Profound 2 Background 21 blood and blood-forming tissues, Abnormality of the nervous system, Abnormality of the ear, etc. And it contains terms such as Intestinal carcinoid, Small intestine carcinoid, etc. 2.4.2 Terms Attributes The majority terms of the HPO belong to the Phenotypic Abnormality sub-ontology. Each class has a unique ID such as HP:0002503 and a name such as Spinocerebellar tract degeneration. Table 2.1 displays the name and description of each attribute that each class can have [51]. Table 2.1 The available attributes for each HPO class and their descriptions. Attribute Definition id The unique id of the current class. Cardinality: exactly one. is_anonymous To indicate if the current class has an anonymous id. Cardinality: zero or one. name The name of the current class. Each class may have only zero or one name defined. Cardinality: zero or one. alt_id Defines an alternate id for this class. Cardinality: any def The definition of the current class. Cardinality: zero or one comment A comment for this class. Cardinality: zero or one. subset It indicates a class subset to which this class belongs. Cardinality: any. synonym It gives a synonym for this class. Cardinality: any. xref It describes an analagous class in another vocabulary. It points to external disease databases such as Unified Medical Language System (UMLS)6 and Medical Subject Headings (MeSH)7. Cardinality: any. property_value It binds a property to a value in this instance. Cardinality: any. is_a It describes a subclassing relationship. Cardinality: any. created_by Name of the creator of the class. Cardinality: zero or one. creation_date The creation date of the class. Cardinality: zero or one. is_obsolete It indicates whether the current class is obsolete. The allowable values are "true" and "false". Cardinality: zero or one. replaced_by It specifies a class which replaces an obsolete class. Cardinality: any. consider It determines a class which is an appropriate substitute for an obsolete class. Cardinality: any. 6 https://www.nlm.nih.gov/research/umls/index.html 7 https://www.nlm.nih.gov/mesh/ HAITHAM MAAROUF 22 2.5 GUIDELINE DEFINITION LANGUAGE The Guideline Definition Language (GDL)8 is oriented to formally represent clinical procedural knowledge for computerized clinical decision support systems, using the format of knowledge rules [52]. GDL is designed to be natural language – and reference terminology - agnostic by leveraging the design of openEHR Archetype Model and openEHR Reference Model. GDL represents clear-cut clinical knowledge for singe-decision making. The importance of the GDL is:  It allows expressing the rules of CDS using archetypes both as input and output for the rule execution.  It is natural language-agnostic and support multiple language translations without changing the definitions of the rules  It is reference terminology-agnostic so various terminologies can be used.  It converts the CDS rules to main-stream general-purpose rule languages for execution  It facilitates the reusability of the CDS rules in different clinical contexts.  It allows grouping a set of related CDS rules in order to support complex decision making.  It is technology independent. 2.5.1 Components of GDL GDL has four main parts: Header, Definition, Rule and Ontology. 2.5.1.1 GDL Header The Header introduces the GDL and contains metadata such as authors, keywords and information about the purpose, etc. Fig. 2.9 displays the Header section of “CHA2DS2VASc” GDL guide document and it illustrates the main parts in this section. CHA2DS28 http://www.openehr.org/releases/CDS/latest/docs/GDL/GDL.html#_guideline_definition_language_gdl 2 Background 23 VASc is a clinical instrument for stroke risk stratification in atrial fibrillation [53]. Fig. 2.9 Extract of the Header section in "CHA2DS2VASc.gdl".It illustrates the current version of the guideline, authorship information, keywords, purpose and use of the guideline. (Taken from [29] ) 2.5.1.2 GDL Definition The Definition section contains all the elements used inside the guideline, the mappings to the archetypes and pre-conditions. Fig. 2.10 displays the Definition section of “CHA2DS2VASc” GDL guide document. It displays the archetype_binding within the (GUIDE) < gdl_version = <"0.1"> id = <"CHA2DS2VASc_Score_calculation.v1"> concept = <"gt0001"> language = (LANGUAGE) < original_language = <[ISO_639-1::en]> > description = (RESOURCE_DESCRIPTION) < details = < ["en"] = (RESOURCE_DESCRIPTION_ITEM) < keywords = <"atrial fibrillation", "stroke", "CHA2DS2-VASc"> purpose = <"Calculates stroke risk for patients with atrial fibrillation, possibly better than the CHADS2 score."> use = <"Calculates stroke risk for patients with atrial fibrillation, possibly better than the CHADS2 score."> > > original_author = < ["date"] = <"2012/12/03"> ["email"] = <"[email protected]"> ["name"] = <"Rong Chen"> ["organisation"] = <"Cambio Healthcare Systems"> > > > HAITHAM MAAROUF 24 guide_definition section, which binds data elements from the archetypes to variables used by GDL rules. It also illustrates that there is a condition (pre-condition) must be met before the rules inside the guide can be executed. For example, this guideline will not be executed unless the patient has atrial fibrillation. Fig. 2.10 Extract of the Definition section in "CHA2DS2VASc.gdl".(Taken from [29]) 2.5.1.3 GDL Rule The Rule section contains the condition and action parts of rules. Each rule consists of two parts: the first part contains the conditions needed for the rule to execute and it starts with the keyword ‘When’, and the second one comprises the actions that will be carried out once the rule is activated and it starts with the keyword ‘Then’. Fig. 2.11 shows rules that inspect different diagnoses relevant to CHA2DS2VASc score. definition = (GUIDE_DEFINITION) < archetype_bindings = < ["gt0002"] = (ARCHETYPE_BINDING) < archetype_id = <"openEHR-EHR-EVALUATION.problem-diagnosis.v1"> domain = <"EHR"> elements = < ["gt0015"] = (ELEMENT_BINDING) < path = <"/data[at0001]/items[at0002.1]"> > > > > pre_conditions = <"$gt0015!=null",...> > 2 Background 25 Fig. 2.11 Extract of the Rule section in "CHA2DS2VASc.gdl".(Taken from [29] ) 2.5.1.4 GDL Ontology In the Ontology section, all the terms are bond to user interface labels and description of the terms in supported natural languages. In addition, terms are bound to external terminologies. Fig. 2.12 illustrates how the atrial fibrillation term is bound to a specific code in ICD109. Fig. 2.12 Extract of the Ontology section in "CHA2DS2VASc.gdl".(Taken from [29]) The GDL editor10 enables users to create and run GDL files, and it is a multiplatform desktop application. 9 http://www.who.int/classifications/icd/ICD-10_2nd_ed_volume2.pdf 10 https://sourceforge.net/projects/gdl-editor/ rules = < ["gt0027"] = (RULE) < when = <"$gt0016!=null",...> then = <"$gt0007=1|local::at0028|Present|",...> priority = <11> > ["gt0028"] = (RULE) < when = <"$gt0017!=null",...> then = <"$gt0008=1|local::at0031|Present|",...> priority = <11> > > ontology = (GUIDE_ONTOLOGY) < term_bindings = < ["ICD10"] = (TERM_BINDING) < bindings = < ["gt0015"] = (BINDING) < codes = <[ICD10::I48],...> uri = <””> > > > >> HAITHAM MAAROUF 26 2.6 WEB ONTOLOGY LANGUAGE Ontologies are used to capture and model knowledge about some domain of interest. The Web Ontology Language (OWL) [54] is the most recent development in standard ontology languages. The OWL is a description-logic-based language used to formalize the concepts and relationships of a domain in terms of individuals, properties, classes, and data values. It represents rich and complex knowledge about concepts of a domain and relations between concepts. It has a richer set of operators – e.g. negation, union, and intersection. OWL depends on a different logical model which facilitates the definition and the description of concepts. Therefore, complex concepts can be built up. Additionally, the logical model provides the use of reasoners. Reasoners like Hermit are used to check if definitions and statements in ontologies are mutually consistent and also to identify which concepts fit under which definitions. Correctly maintaining the hierarchy is a critical issue, especially when dealing with cases where there are concepts with at least two parents. Reasoners are very helpful in preserving this hierarchy [55] . 2.6.1 Components of OWL Ontologies An OWL ontology consists of Individuals, Properties, and Classes [55]. 2.6.1.1 Individuals Individuals represent objects in the interested domain. They are also known as instances of classes. For example, Microsoft is an instance of a class called Company. 2.6.1.2 Properties Properties are binary relations between individuals – i.e. They link two individual together. For example, the property hasOwner might link the individual Microsoft to Bill Gates. Properties can have inverses. For example, isOWnedBy is the inverse of hasOwner. They can also be either symmetric or transitive. 2 Background 27 2.6.1.3 Classes OWL classes are sets that comprise individuals. They are described using formal descriptions. Classes’ descriptions specify precisely the requirements for memberships of the classes. For example, the class Person would contain all the individuals that are persons in the interested domain. Classes can be organized into a superclass-subclass hierarchy. Fig. 2.13 shows a representation of two classes Company and Person which are represented as circles, a representation of four individuals Microsoft, Facebook, Bill Gates, and Mark Zuckerberg which are represented as diamonds, and a representation of a property isOwenBy which is represented as a curved line. Fig. 2.13 Representation of Individuals, Properties and Classes. Individuals are represented by diamonds, properties by arrows and classes by circles. 2.7 THE SCALE FOR THE ASSESSMENT AND RATING OF ATAXIA (SARA) The SARA assesses severity of cerebellar dysfunction through the evaluation of eight items reflecting motor performance (gait, stance, sitting, speech disturbance, finger-chase test, nose-finger test, fast alternating hand movements and heel-shin test) [56] (See Appendix A for the full details of the items). For the last four items, upper and lower HAITHAM MAAROUF 34 the results of the scale. Our modeling approach is based on four main steps (Fig. 3.1): 1) Building of a reduced version of the HPO, through extraction of the ontology modules relevant to the SARA 2) Annotation of the free-text descriptions of the rating scale with the ontology modules 3) Development of two clinical archetypes (Observation and Evaluation) 4) Definition of the information-processing units to express the clinical interpretation support system. Each step involves several activities, described below and summarized in (Fig. 3.2). The modeling of SARA involved a data level - representation of the SARA itemsand a knowledge levelreferred to the strategy to compute the total score and interpretation of phenotype. Archetypes were used to model the data level, while GDL and OWL were used to model the knowledge level. This representation had restrictions, since openEHR models support GDL, but do not give much support for OWL and related reasoning. To bring the gap between the clinical archetypes and the ontology, mappings were defined. These mappings facilitated the translation of the archetype instances to the OWL dataset. 3 Methods 35 Fig. 3.1 The modeling approach. It includes: 1) Extraction of a reduced version of the HPO; 2) Annotation of the free-text descriptions of SARA items and scores; 3) Development of two archetypes (Observation and Evaluation); 4) Definition of the information-processing units in order to express the clinical interpretation of the SARA. Fig. 3.2 Summary of the main proposed activities for the modeling of electronic rating scales.They cover: extracting ontology modules, annotating the scale, building archetypes, and expressing interpretation. Textual Sources HPO ontology Ontology modules Extraction (1) Annotation (2) SARA items & Scores SARA with annotations Evaluation Archetype (+ Bindings) Observation Archetype (+ Bindings) Scoring System (GDL) Assessment of impairment (GDL) Diagnosis of cerebellar syndrome (GDL) Clinical synopsis (OWL reasoner) Development (3) Expression of the interpretation (4) Step3 BUILDING ARCHETYPE •Select archetype class •Name archetype •Select structure •Add data types •Add constraints •Add metadata •Copy scale annotations to the terminology bindings Step1 EXTRACTING ONTOLOGY MODULES •Review the domain •Select sources •Annotate sources with the HPO •Extract the initial ontology modules Step2 ANNOTATING THE SCALE •Match items partially •Revise candidate concepts •Propose new concepts and relationships •Reorganize structure Step 4 EXPRESSING INTERPRETATION •Express calculation •Express value interpretation •Express total score interpretation •Express phenotypic knowledge HAITHAM MAAROUF 36 3.2.1 Extracting the HPO ontology modules relevant to the SARA We reviewed and collected free-text sources describing the SARA, its component tests and application. For example, the following text describes the functional subdivisions of the cerebellum in order to explain the rationale of a coordination exam: “The cerebellum has 3 functional subdivisions, …. The first is the vestibulocerebellum. … Dysfunction of this system results in nystagmus, truncal instability (titubation), and truncal ataxia ... The following tests of the neuro exam can be divided according to which system of the cerebellum is being examined….”. We then annotated the sources to HPO terms, using the OBO Annotator [63], a phenotype concept recognition system (Fig. 3.3). In total, 12 HPO classes were annotated, which provided the set of seed terms required to extracting a self-contained portion of the HPO. Such reduced version of the HPO covered all the classes relevant to the SARA. In general, self-contained portions of an ontology are referred as ontology modules/segments [64], or slims in the context of the Gene Ontology. While these modules are subgroups of a base ontology, in our case the HPO, they are also equally valid on their own [65], but simpler and more manageable than the complete ontology. 3 Methods 37 Fig. 3.3 OBO Annotator interface to annotate a text to HPO classes.The text related to Gait Ataxia is annotated to HP:0002066 (Gait Ataxia) and HP:0002355 (Difficulties Walking). Annotating the rating scale SARA 3.2.2 Annotating the rating scale SARA The goal of this stage is to map the SARA items with standard patient phenotypes described into its own scale. Firstly, items were partially matched to ontology modules class names. Partial match happens when the item name is embedded inside some class name. An ontology engineering of our research group (Maria Taboada) programmed a specific method to run partial matches. Thus, one or more candidate classes for each item were obtained. Then the candidate mappings were revised by a neurologist (Maria Sobrido), who selected the most appropriated classes and proposed a minimal extension and reorganization of the the reduced version of the HPO, with the aid of the complete HPO, in order to cover all the details of the needed neurodegenerative phenotypes, but keeping it as close to the original HPO as possible. Six new classes, with the new relationships, were added to the ontology modules to precisely annotate four SARA items (stance, sitting, finger chase and heel-shin slide). Fig. 3.4 shows the HAITHAM MAAROUF 38 new classes and relationships, and the mappings between items in the SARA and classses in the ontology modules. The golden color is used to highlight the new classes and is_a relationships. On the other hand, each score in the SARA is accompanied by a textual description, describing a level of severity. In order to annotate the scores, we decided to reuse the general HPO classes describing the different levels of severity: borderline, mild, moderate, severe and profound. We introduced new subclasses to the eight HPO classes that were used to annotate the SARA items. The subclasses were defined based on the severity levels of their superclasses. For example, moderate_dysarthria was defined as a subclass of dysarthria with a moderate severity (Fig. 3.5). Adittionally, two scores were identified by the neurologist as two HPO phenotypes. The first one is the score 8 for the Gait item, which was bound to abasia and the second is the score 6 for the Speech Disturbance item, which was bound to anarthria. In addition, the neurologist added a new class (named astasia) to map the score 6 for the stance item. 3 Methods 39 Fig. 3.4 Excerpt from the set of mappings between the SARA and the ontology.Squares represent the SARA items. Gray and golden ellipses, respectively, are the original HPO classes and classes added to the ontology modules. Blue arrows are mappings between SARA items and HPO classes. Black and golden arrows, respectively, represent the original and the additional is_a relationships. Note that subsumptions in Electronic health record (Open Biological Ontologies) are represented using the is_a relationship, whereas in OWL using the subclassOf constructor. Gait and Balance Stance Sitting Gait Upper and Lower Coordination Nose Finger Test Fast alternating hand Finger Chase Heel-shin slide Speech Disturbance Truncal Ataxia (Midline Ataxia) Postural Instability Neurological Speech Impairment Gait Ataxia Abasia Mapping Ataxic Postural Instability Standing Instability is_a is_a is_a Astasia Sitting Instability is_a is_a is_a Mapping Dysarthria Anarthria is_a is_a Mapping Mapping New is_a Relationship Original is_a Relationship In HPO Upper Limb Dysmetria Dysmetria Limb Dysmetria Appendicular Ataxia Lower Limb Dysmetria Mapping Mapping Intention Tremor Dysdiadochokinesis Mapping is_a is_a is_a is_a is_a is_a Mapping is_a Mapping between SARA items and HPO classes New HPO class Original HPO class HAITHAM MAAROUF 40 Fig. 3.5. The five subclasses of dysarthria class. Once the ontology modules was modelled and the annotations were created, three main superclasses were identified taking the eight scale items into account: i) truncal ataxia (midline ataxia), which subsumed the classes gait_ataxia and ataxic_postural_instability (which subsumed standing_instability and sitting_instability); ii) dysarthria, annotating the fourth item; and iii) Appendicular Ataxia (limb ataxia which subsumed limb dysmetria, intention tremor and dysdiadochokinesis (Fig. 3.6). Thus, we structured and organized the rating scale items in accordance with the ontology, by inserting three CLUSTER nodes in a hierarchical structure: i) gait and balance, which is linked to truncal ataxia; ii) speech disturbance, which is relating to dysarthria; and iii) upper and lower limb coordination, which is associated with Appendicular Ataxia (limb ataxia). This new structure does not alter in any way the scale, as it continues to be based on an 8item performance. It simply arranges the items. This new organization is shown on the upper part of the Fig. 3.7 A. The new nodes can be viewed as clinical dimensions, but without any assigned score. Additionally, each item in the third node or clinical dimension was split into left part and right part, in order to capture the item on each side, as set out in the original rating scale. Once again, we organized them in a hierarchical structure. Finally, the SARA-specific, reduced version of the HPO was translated to Protégé11, and its properties were manually 11 http://protege.stanford.edu/products.php#desktop-protege 3 Methods 41 modeled. To check the consistency, the HermiT reasoner was used. Fig. 3.6. Excerpt from the domain ontology. It shows some classes of the reduced version of the HPO. HAITHAM MAAROUF 42 Fig. 3.7. The structure of the content of the SARA.(a) The new organization of the SARA, with three main patient’s phenotype components: gait and balance, speech disturbance, and upper and lower limb coordination; (b) The clinical observation archetype developed for the SARA. (a) (b) Gait and Balance Gait Stance Sitting Score 1 Score 2 Score 3 Speech Disturbance Speech Disturbance Upper and Lower Limb Coordination Score 4 Finger Chase Right Left Score 5 Nose Finger test Right Left Score 6 Fast alternating hand Right Left Score 7 Heel-shin Slide Right Left Score8 Total Score SARA Scale 3 Methods 43 3.2.3 Developing the archetypes for the SARA For the archetype design, we made a difference between data to be gathered by the scale or directly calculated using the information provided by the scale (items, clinical dimensions, individual scores and the total score), and data that are not directly available in the scale. Taking account of this distinction, we decided to use two types of clinical archetypes: Observation for normalizing data provided by the scale, and Evaluation, for data not directly available in the scale. The first one fits to capture the scale content, whereas the second one records clinically interpreted findings, such as the phenotypic abnormalities derived from the rating scale. The two clinical archetypes were developed using the archetype editor12 provided by OpenEHR. 3.2.3.1 Modeling the SARA Observation archetype To model the observation archetype, after adding the metadata (e.g. purpose, keywords, definition, author, etc.), we structured the content of the SARA according to this archetype, i.e., by means of a tree structure with elements for the three clinical dimensions (the set of items, the total score and the date of the observations), and values for the individual scores. Then, we defined the elements with proper data types, descriptions, comments, details, occurrences, constraints and possible values. We mapped each element to the corresponding ontology class of the reduced version of the HPO, using the achieved annotations (See Section 3.2.2). Finally, we structured and organized the items in accordance with the ontology, by inserting three CLUSTER nodes in a hierarchical structure: i) gait and balance, which was linked to truncal ataxia; ii) speech disturbance, which was linked to dysarthria; and iii) upper and lower limb coordination, which was associated with appendicular ataxia. This new structure did not alter the SARA, as it continued to be based on the same 8-items, but only arranged these items in a specific way (Fig. 3.7 A). The new nodes can be viewed as clinical dimensions, but without any assigned score. Fig. 3.7 B shows that the observation archetype meets the rating scale structure. As the archetype editor did not provide HPO in the list of 12 http://www.openehr.org/downloads/archetypeeditor/home HAITHAM MAAROUF 50 Clinical synopsis is inferred from the ontology modules partially based on the HPO. It requires using elements and values from the Evaluation archetype (such as, gait ataxia, sitting and standing instability) to infer other elements and values from the same archetype (such as, midline ataxia) using the OWL ontology. During this modeling phase, we only could simulate the ontology reasoning using the Protégé tool. From the values inferred by the GDL for the elements of the evaluation archetype, and taking into account the mappings of this archetype, we manually entered the individuals of the linked OWL classes and run the reasoner. 51 4 RESULTS With the aim of testing the appropriateness of the methods presented in the previous chapter, we implemented a prototype of electronic rating scale for the SARA, called SARA Management System. The prototype can be used both for the assessment of cerebellar syndrome and for the production of a clinical synopsis. Additionally, we validated the approach in a real clinical setting, by using the recorded SARA data from 28 anonymous subjects affected by Spinocerebellar Ataxia Type 36 (SCA36). This chapter details the implemented prototype and, dataset and validation of the method and the results of the validation. 4.1 THE SARA MANAGEMENT SYSTEM To demonstrate the functionality of our approach, we developed a framework entitled “SMS” (SARA Management System). SMS allows the management of patient data of cerebellar syndrome. We used the JAVA as a programming language, NetBeans as the integrated development environment, JAVA Swing as user interface toolkit, and MySQL as a database management system to store all the needed information. 4.1.1 SARA Management System Architecture The architecture of “SMS” was structured in three layers: “Persistence”, “Operation” and “Interface”. 4.1.1.1 Persistence Layer The persistence layer used MySQL to store the clinical archetypes, the patient input data, the data inferred by the system, and additional information. Fig. 4.1 shows the relational model of the database with three types of tables. The first type included the class, HAITHAM MAAROUF 52 subclass, scale, element and value tables. All of them modeled the attributes of the clinical archetypes (elements, values and mappings to ontology classes). To a greater extent, they were built based on the structure of the modeled archetypes. In order to populate the database with data from the archetypes, the ADL parser 13 provided by the OpenEHR Foundation was used to generate a dependency tree of terms and then create an XML file, which was aimed at producing archetype instances when needed. The second type of tables stored the patient data, and the last type (test and test_value tables) recorded the set of SARA tests. 4.1.1.2 Operation Layer The operation layer included all information-processing units (Table 3.2) that are responsible to run Calculation Functions, Assessment of the Degree of Impairment. Assessment of Cerebellar Syndrome, and Clinical Synopsis. The current version of GDL does not provide a Java API or any mechanism for manipulating the rules defined using the editor. Hence, the only solution was to rewrite the rules in a rule engine (such as Drools14 or Clips15) or to implement directly in Java (we decided this second solution). However deriving phenotype information was reasonably implemented using the OWL API [66] and the mappings linking archetype elements and values to OWL classes. The mappings were useful to automatically create OWL individuals from the archetype instances stored into the database. 13 https://github.com/openEHR/java-libs/tree/master/adl-parser 14 http://www.drools.org/ 15 http://www.clipsrules.net/ 4 Results 53 Fig. 4.1. Database Relational ModelThe table ‘patient’ recorded the patient identifier, the tables ‘test’ and ‘test_values’ stored the information about the tests covered by the SARA, and the rest of the tables recorded the information modeled in the clinical archetypes. HAITHAM MAAROUF 54 4.1.1.3 Interface Layer On the other hand, the interface layer consisted of several input and output forms. The main form of the tool contained two menus: the ontology menu and the SARA menu. The first menu (Fig. 4.2) allowed users to modify the structure of the ontology, by adding, updating or deleting classes in the ontology . It also provided an option to automatically generate a hierarchical graph that displayed all the available classes and their is_a relationships. The SARA menu consisted of two sub-menus: 1) Observation, where neurologists could select patients and enter the 12 values of the assessed elements defined in the SARA scale and modeled in the observation archetype (Fig. 4.3); and 2) Evaluation, which provided three main features (Fig. 4.4):  A table with all phenotypic abnormalities inferred from the collected patient data.  A textual report summarizing the patient status. The report included the assessment of cerebellar syndrome and a brief synopsis of the phenotypic abnormalities accompanying the syndrome.  An evaluation graph visualizing all the patient phenotypic abnormalities (Fig. 4.5). The green ellipses are the patient annotations based on the observation data and the yellow ones are the inferred abnormalities. To facilitate semantic interoperable data exchange, the approach was designed to deliver the XML data collected by the SARA (using the schemas compliant with the observation archetype) and data inferred by the application (using the schemas compliant with the evaluation archetype). Additionally, the system provided the facility to send the SARA results and the patient report by e-mail. The doctor could attach it to the patient medical record. 4 Results 55 Fig. 4.2. Screenshot of ontology update. The form can be used for checking, remove and update classes of the reduced version of the HPO used by the SMS. All the classes can be viewed in both graphical and tabular formats. Fig. 4.3. Observation form. It allows neurologists to enter the values of the SARA scale items. HAITHAM MAAROUF 56 Fig. 4.4. Evaluation form.It displays an example of phenotypic abnormalities derived from the data of a patient (on the left side) and a textural report summarizing the status of this patient (on the right side). Fig. 4.5. Screenshot of a graphical summary. It displays a graph visualizing the set of phenotypic abnormalities inferred by the SMS. The green ellipses represent the lower classes in the hierarchy and the yellow ellipses, the superclasses. Class id Class name Value Severity HP:002066 Gait Ataxia 3 Moderate new02 Standing Instability 2 Mild new08 Lower Limb Dysmetria 1.5 Mild Patient: P9 Observation: 33-2017-05-01 20:49:23.0 Score=6.5 Phenotypic abnormalities codes Patient status report Date:2017-05-01 20:49:23.0 ID patient:9 The patient has a total SARA score of 6.5. This corresponds to a MILD Cerebellar Syndrome, involving: - Moderate truncal ataxia ( Moderate gait impairment, Mild standing impairment, and Normal sitting instability). - Appendicular ataxia: o Upper Limb Dysmetria: Normal bilaterally o Intention tremor: Normal bilaterally o Dysdiadochokinesis: Normal bilaterally o Lower Limb Dysmetria: Mild on the right side and Moderate on the left side Display the Evaluation Graph 4 Results 57 4.2 DATASET AND VALIDATION OF THE METHOD For validation purpose, we used data records from 28 anonymous subjects with Spinocerebellar Ataxia Type 36 (SCA36) [67]. All patients were examined following the SARA. Only the set of scores for each item collected by the SARA was taken into account during this evaluation. The institutional research ethics committee approved the recruitment and study protocol, and all participants gave their written informed consent. Two independent neurologists validated the feasibility of our approach. The evaluation was carried out in three steps: 1) Total score calculation and interpretation: We filled out the score data for each patient to get the interpretation. The system automatically inferred:  The severity for each item in the scale  The severity of cerebellar syndrome,  The severity of truncal ataxia  The severity of appendicular ataxia on the right and left sides. 2) Interpretation by two independent neurologists: The same data set was sent to two neurologists, but adding the calculated total scores for each patient. The neurologists used their expertise in ataxia to determine the severity of the cerebellar syndrome from the total score, and the severity of truncal and appendicular ataxia, if present, from the individual scores. 3) Comparison of results between the system and the human experts (neurologists): The validation process was carried out based on the results-oriented validation perspective [68]. This perspective is based on comparing the performance of the developed tool with an expected performance provided by human experts, in order to assess whether the tool produces the required output correctly. There are many methods of assessing inter-rater agreement. Specifically, Cohen’s kappa [69] and Weighted kappa [70] have been widely applied in the HAITHAM MAAROUF 58 medical field [71]. Weighted Kappa is more compatible with ordinal scales [72-82], hence we decided to use it as the SARA is ordinal and the difference between severity levels is meaningful. To validate the system, the following steps were accomplished: 1) we created a sheet using SPSS16, 2) we imported the interpretations of the system and the neurologists into SPSS, and 3) we ran Weighted Kappa test 12 times to measure the strength of agreement between the system and each neurologists, and between the two neurologists themselves. 4.3 VALIDATION OF THE SYSTEM The validation of the system was based on the Weighted Kappa values obtained from the 12 tests conducted. Table 4.1 displays the results of measuring the strength of agreement between the system and the first neurologist, the system and the second neurologist and between the two neurologists themselves. Kappa values are between 0.65 and 0.93, as illustrated in Table 4.1. Table 4.1 Agreement between automated and manual ratings (Weighted Kappa) System vs. 1st Neurologist System vs. 2nd Neurologist 1st Neurologist vs. 2nd Neurologist Kappa value Kappa value Kappa value Cerebellar syndrome 0.86 0.84 0.85 Midline Ataxia 0.80 0.84 0.86 Appendicular Ataxia (Right side) 0.71 0.80 0.86 Appendicular Ataxia (Left side) 0.62 0.78 0.84 16 http://www.ibm.com/analytics/us/en/technology/spss/ 59 5 DISCUSSION In this doctoral thesis, a mixed method to support the development of the SARA has been presented. The method combined OpenEHR archetypes, guidelines, ontologies and reasoning. The innovation of our method rests on how these approaches were combined to get the full benefit of them. We distinguished between the modeling phase and the implementation phase. During the former, we addressed the calculation and assessment tasks by defining and executing GDL rules, and the clinical synopsis task by defining OWL classes and executing a reasoner. However, due to the lack of integration between these frameworks, we first ran the GDL framework, and then we manually entered the results in Protégé in order to infer the phenotypic abnormalities. During the implementation phase, we addressed the calculation and assessment tasks by rewriting the rules directly in Java, and the clinical synopsis task by integrating the OWL API into the system and using the mappings to create OWL individuals. We designed the approach as an archetype-based stand-alone application, providing a meaningful way for collecting and interpreting healthcare data. The application released the local EHR system of integrating the SARA, providing a standard way of delivering the collected and inferred data. Thus, the main role of this electronic rating scale was to collect the normalized data, execute the decision support logic and deliver both data and interpretations to the EHR system. Turning to the research questions in this research, a few conclusions can be drawn. With respect to question (1) - Is the combination of GDL, openEHR clinical archetypes and ontologies suitable for the description of all knowledge covered by the SARA?) - , we can conclude that a combination of OpenEHR, GDL and OWL offers a suitable framework for the purpose of describing the data and knowledge levels of the SARA. OpenEHR provides a formal specification at the data level, whereas GDL and ontologies offer formal specifications of different types of knowledge for data HAITHAM MAAROUF 66 highly satisfactory, we will develop a simple mobile application for the automatic transmission of the interpretation to the health information system. Table 5.1. Kappa interpretation rules-Landis and Koch (1977) Kappa Statistic Strength of agreement 0.00 Poor 0.00-0.20 Slight 0.21-0.40 Fair 0.41-0.60 Moderate 0.61-0.80 Substantial 0.81-1.00 Almost Perfect Table 5.2. Strength of agreement between automated and manual ratings. It follows the interpretation rules proposed by Landis and Koch. System vs 1st Neurologist System vs 2nd Neurologist 1st Neurologist vs 2nd Neurologist Cerebellar Syndrome Almost Perfect Almost Perfect Almost Perfect Truncal Ataxia (Midline Ataxia) Substantial Almost Perfect Almost Perfect Appendicular Ataxia (Right Side) Substantial Substantial Almost Perfect Appendicular Ataxia (Left Side) Substantial Substantial Almost Perfect Finally, although our approach was designed to implement a prototype for managing the SARA, it is rather generic and hence applicable to model other electronic rating scales, possibly in other 5 Discussion 67 clinical domains. To take an example, the approach could be applied to the domain of the autism spectrum disorders, which exhibit complex phenotypes affecting variables that are difficult to measure. As a consequence, standardized scales are often used to collect a large amount of phenotypic data. Recently, a phenotype ontology has been developed to identify behavioral features of importance [86]. The availability of this ontology and also the mappings to the rating scales would facilitate the implementation of prototypes like the one presented here. 69 6 CONCLUSIONS AND FUTURE WORK 6.1 CONCLUSIONS 1. Reducing all content of a rating scale to a unique number may lead to loss of useful clinical information about the dimensions implicitly collected by the scale. In this doctoral thesis, we developed a model to infer the full components of the patient’s phenotype from the clinical dimensions represented by the rating scores. This model provides automated support for medical evaluation, report writing, and clinical decision-making. The proposed approach has been shown for the Scale for the Assessment and Rating of Ataxia (SARA), a well-validated instrument to evaluate the presence and severity of cerebellar ataxia. 2. Integrating electronic rating scales with the electronic health records and related systems requires formally describing these scales using standard clinical information models, such as openEHR. In this doctoral thesis, a novel combination of the best performances from OpenEHR clinical archetypes, guidelines and ontologies has been proposed to be able to reason on clinical archetypes. We showed for the specific field of ataxias, how clinical information models can be mapped to standard terminologies or ontologies, which provide the required meaning of their concepts. 3. The integration of phenotype ontologies with clinical archetypes provides not just a static knowledge store, but also a dynamic resource that allows automatic inference of a patient´s medical status (phenotype) from systematized collection of clinical data. HAITHAM MAAROUF 70 6.2 SPECIFIC CONTRIBUTIONS This doctoral thesis work contributes to a better understanding of how clinical archetypes, guidelines and ontologies can be combined for modeling and implementing the SARA. There are several contributions in this research. 1. This research proposes an ontology-aware approach of clinical models, guidelines and terminologies to model electronic rating scales, where the ontology provides the backbone for normalizing the content of the scale through clinical archetypes. 2. The modeling approach distinctly clarifies the line of demarcation between the data level - representation of the scale itemsand a knowledge levelreferred to the strategy to compute the total score and the interpretation of patient phenotype. Archetypes facilitate the standard modeling of the data level, while GDL and OWL enable the standard modeling of the knowledge level. 3. The novel archetype development approach reduces the effort necessary for creating mappings, which is key to achieve semantic interoperability among different data sources. It also prevents large semantic discrepancies between the modeled archetypes and the ontology modules. 4. Additionally, a clear and explicit separation between the standard components of the scale related to the content (i.e., items, clinical dimensions and scores), and the clinical interpretations of these components are established. 5. Another key contribution was the clear identification of all different types of knowledge required to interpret the data collected by the scale. 6. The knowledge required to exploit reasoning on the scale data was modeled as separate information-processing units interconnected via the defined archetypes, providing a simple mechanism of combining ontology and rulebased reasoning. 6 Conclusions and Future Work 71 7. A prototype named SARA Management System was developed to demonstrate the validity of the modeling approach. The prototype can be used for both the assessment of cerebellar syndrome and the production of a clinical synopsis. 8. The prototype was validated using recorded data from 28 anonymous subjects affected by Spinocerebellar Ataxia Type 36 (SCA36). The results reveal a substantial degree of agreement between the results achieved by the ontology-aware system and the human experts. 6.3 LIMITATIONS OF THE WORK The innovation of our method rests on how clinical models, guidelines and terminologies were combined to get the full benefit of them. We have distinguished between the modeling phase and the implementation phase. During the former, we addressed the calculation and assessment tasks required by the scale by means of defining and executing GDL rules, and the clinical synopsis task by defining OWL classes and executing a reasoner. However, due to the lack of integration between GDL and OWL, we first ran the GDL framework, and then we manually entered the results in Protégé in order to infer the phenotypic abnormalities. This is clearly a limitation of the work, resulting from the current gaps in technology. Additionally, during the implementation phase, we addressed the calculation and assessment tasks by rewriting the rules directly in Java, and the clinical synopsis task by integrating the OWL API into the system and using the mappings to create OWL individuals. Once again, the inability of the current technology to automatically translate GDL rules to Drools or Clips rules to be integrated in a Java framework with the OWL API must be solved in the future work. From the results achieved in this doctoral thesis, we have concluded that a combination of OpenEHR, GDL and OWL offers a suitable framework for the purpose of describing the data and knowledge levels of the SARA. However, it should be emphasized that in our particular case, the knowledge level could be broken into separate information-processing units interconnected in a simple way through the two defined archetypes (one for observations and another HAITHAM MAAROUF 72 for evaluations). However, the interpretation of a rating scale may require more complex control mechanisms, demanding more interoperability between GDL and OWL. Furthermore, the current version of GDL uses archetype data as input and output variables for all the rules, but it provides no facility to define auxiliary variables. This type of variables is sometimes necessary to model procedural knowledge, such as the counting of the scores in rating scales. We showed that a full integration of the current technologies to model the rating scale is not possible at the moment. In the modeling stage, the use of GDL facilitated the development and interconnection of most processing units, without resorting to external resources and encouraging knowledge sharing. However, the current editor does not supply any facility for interoperability. For example, the generation of XML instances of the archetypes would be a remarkable advance to provide the option of combining the tool with other different inference engines, such as description logic reasoners. Finally, the interpretation of the results of our prototype reflects a very high degree of agreement between the prototype and the experts, confirming that the approach can be a good solution to develop electronic rating scales. Even so, these excellent results should also be viewed with much caution, as the validation was carried out only with 28 patient data, all of them affected by the same rare disease (SCA36). Additionally, although the two neurologists who carried out the assessment were independent, they work in the same hospital and one of them is in the same research group as the neurologist involved in the modeling process. It therefore has to be assumed that there exists consistency between the three neurologists. 6.4 FUTURE WORK Nowadays, the most ataxic disorders still have no successful pharmacological therapy, and patients suffer the unavoidable degenerative disease progression. The aim of well-validated rating scales is to understand better the natural history of ataxic disorders and evaluate properly drug efficacy in clinical trials. Rating scales facilitate clinical standardization of data collection, mainly in specialties with a richness of complex phenotypic variables, such as neurology. 6 Conclusions and Future Work 73 However, the current electronic approaches are simple calculators with no integration with the electronic health records and related systems. In this doctoral thesis, a new solution to work towards this goal is provided. Exploiting reasoning on clinical archetypes represents a challenge With the aim of achieving a full integration of the current technologies to model rating scales, we plan to evaluate the expressivity of the new major version of ADL. In particular, we plan to evaluate the specifications for defining explicit rules of invariant assertions. (i.e., expressions that should be satisfied by all instances of an archetype). If the definition of these rules provides the same functionality as GDL rules defined in our system, we will implement the facilities required to automatically execute these ADL rules and integrate with the OWL API. We will also evaluate Owlready2, a module for ontology-oriented programming in Python. We think that this module may provide the needed functionality for full integration. Furthermore, in our future work, we will evaluate the SARA application with a larger number of patient data that are affected by diverse cerebellar ataxias, and with the help of neurologists from different hospitals. This new evaluation will provide us a stronger validation of our approach. 75 REFERENCES 1. Robinson P N, Mungall C J, and Haendel M. Capturing phenotypes for precision medicine. Molecular Case Studies. 2015. 1(1): p. a000372. 2. Baynam G, Walters M, Claes P, Kung S, LeSouef P, Dawkins H, Bellgard M, Girdea M, Brudno M, and Robinson P. Phenotyping: targeting genotype's rich cousin for diagnosis. Journal of paediatrics and child health. 2015. 51(4): p. 381-386. 3. Grinnon S T, Miller K, Marler J R, Lu Y, Stout A, Odenkirchen J, and Kunitz S. National institute of neurological disorders and stroke common data element project–approach and methods. Clinical Trials. 2012. 9(3): p. 322-329. 4. Köhler S, Vasilevsky N A, Engelstad M, Foster E, McMurry J, Aymé S, Baynam G, Bello S M, Boerkoel C F, and Boycott K M. The human phenotype ontology in 2017. Nucleic acids research. 2017. 45(D1): p. D865-D876. 5. Human Phenotype Ontology. 2017. http://human-phenotypeontology.github.io. Accessed 13 April 2017. 6. Hobart J C, Cano S J, Zajicek J P, and Thompson A J. Rating scales as outcome measures for clinical trials in neurology: problems, solutions, and recommendations. The Lancet Neurology. 2007. 6(12): p. 1094-1105. 7. Martinez-Martin P. Composite rating scales. Journal of the Neurological Sciences. 2010. 289(1): p. 7-11. 8. Pangman V C, Sloan J, and Guse L. An examination of psychometric properties of the mini-mental state examination and the standardized mini-mental state examination: implications for clinical practice. Applied Nursing Research. 2000. 13(4): p. 209-213. 9. Teasdale G, Maas A, Lecky F, Manley G, Stocchetti N, and Murray G. The Glasgow Coma Scale at 40 years: standing the test of time. The Lancet Neurology. 2014. 13(8): p. 844-854. 10. Goetz C G, Tilley B C, Shaftman S R, Stebbins G T, Fahn S, Martinez‐Martin P, Poewe W, Sampaio C, Stern M B, and Dodel R. Movement Disorder Society‐sponsored revision of the Unified Parkinson's Disease Rating Scale (MDS‐UPDRS): 82 71. Viera A J and Garrett J M. Understanding interobserver agreement: the kappa statistic. Fam Med. 2005. 37(5): p. 360363. 72. Mielke Jr P W, Berry K J, and Johnston J E. Unweighted and weighted kappa as measures of agreement for multiple judges. International Journal of Management. 2009. 26(2): p. 213. 73. Cicchetti D V. Testing the normal approximation and minimal sample size requirements of weighted kappa when the number of categories is large. Applied Psychological Measurement. 1981. 5(1): p. 101-104. 74. Kramer M S and Feinstein A R. Clinical biostatistics: LIV. The biostatistics of concordance. Clinical Pharmacology & Therapeutics. 1981. 29(1): p. 111-123. 75. Banerjee M, Capozzoli M, McSweeney L, and Sinha D. Beyond kappa: A review of interrater agreement measures. Canadian journal of statistics. 1999. 27(1): p. 3-23. 76. Kingman A. Beyond weighted kappa when evaluating examiner agreement for ordinal responses. in Journal of Dental Research. 2002. INT AMER ASSOC DENTAL RESEARCHI ADR/AADR 1619 DUKE ST, ALEXANDRIA, VA 223143406 USA. 77. Ludbrook J. Statistical techniques for comparing measurers and methods of measurement: a critical review. Clinical and Experimental Pharmacology and Physiology. 2002. 29(7): p. 527-536. 78. Perkins S M and Becker M P. Assessing rater agreement using marginal association models. Statistics in medicine. 2002. 21(12): p. 1743-1760. 79. Fleiss J L, Levin B, and Paik M C. Poisson regression. Statistical Methods for Rates and Proportions, Third Edition. 2003: p. 340-372. 80. Kundel H L and Polansky M. Measurement of observer agreement 1. Radiology. 2003. 228(2): p. 303-308. 81. Schuster C. A note on the interpretation of weighted kappa and its relations to other rater agreement statistics for metric scales. Educational and Psychological Measurement. 2004. 64(2): p. 243-253. 83 82. Berry K J, Johnston J E, and Mielke Jr P W. Exact and resampling probability values for weighted kappa. Psychological reports. 2005. 96(2): p. 243-252. 83. Pathak J, Johnson T M, and Chute C G. Survey of modular ontology techniques and their applications in the biomedical domain. Integrated computer-aided engineering. 2009. 16(3): p. 225-242. 84. Smith B, Ashburner M, Rosse C, Bard J, Bug W, Ceusters W, Goldberg L J, Eilbeck K, Ireland A, and Mungall C J. The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nature biotechnology. 2007. 25(11): p. 1251-1255. 85. Landis J R and Koch G G. The measurement of observer agreement for categorical data. biometrics. 1977: p. 159-174. 86. McCray A T, Trevvett P, and Frost H R. Modeling the autism spectrum disorder phenotype. Neuroinformatics. 2014. 12(2): p. 291-305. 85 LIST OF FIGURES Fig. 2.1 Components of rating scales ................................................. 10 Fig. 2.2 Relationship of information types to the investigation process ............................................................................................................ 13 Fig. 2.3. Extract of the archetype GCS. ............................................. 16 Fig. 2.4 Modeling the overall score of Glasgow Coma Scale archetype. ........................................................................................... 17 Fig. 2.5 Modeling recorded data over time in APGAR archetype..... 18 Fig. 2.6 The HPO class Dysarthria and its attributes. ........................ 19 Fig. 2.7 The sub-ontologies of the Human Phenotype Ontology ...... 19 Fig. 2.8 The subclasses included into the class Severity ................... 20 Fig. 2.9 Extract of the Header section in "CHA2DS2VASc.gdl". ..... 23 Fig. 2.10 Extract of the Definition section in "CHA2DS2VASc.gdl". ............................................................................................................ 24 Fig. 2.11 Extract of the Rule section in "CHA2DS2VASc.gdl". ....... 25 Fig. 2.12 Extract of the Ontology section in "CHA2DS2VASc.gdl". 25 Fig. 2.13 Representation of Individuals, Properties and Classes. ...... 27 Fig. 2.14 The eight items included in the SARA scale for cerebellar ataxia. ................................................................................................. 28 Fig. 2.15 Cerebrum and Cerebellum. ................................................. 30 Fig. 3.1 The modeling approach. ....................................................... 35 Fig. 3.2 Summary of the main proposed activities for the modeling of electronic rating scales. ...................................................................... 35 Fig. 3.3 OBO Annotator interface to annotate a text to HPO classes.37 Fig. 3.4 Excerpt from the set of mappings between the SARA and the ontology. ............................................................................................ 39 Fig. 3.5. The five subclasses of dysarthria class. ............................... 40 Fig. 3.6. Excerpt from the domain ontology. ..................................... 41 Fig. 3.7. The structure of the content of the SARA. .......................... 42 Fig. 3.8 Excerpt of the modified "Terminology.xml" including HPO. ............................................................................................................ 44 Fig. 3.9. The Structure of the evaluation archetype. .......................... 45 Fig. 3.10. Example of a calculation function. .................................... 47 Fig. 3.11. GDL rule to assess the degree of sitting instability. .......... 47 Fig. 3.12. GDL rule for absence of cerebellar syndrome ................... 48 Fig. 4.1. Database Relational Model .................................................. 53 Fig. 4.2. Screenshot of ontology update. ........................................... 55 86 Fig. 4.3. Observation form. ................................................................ 55 Fig. 4.4. Evaluation form. .................................................................. 56 Fig. 4.5. Screenshot of a graphical summary. .................................... 56 Fig. 5.1 Auxiliary archetype. ............................................................. 60 Fig. 5.2 Entry form generated by GDL editor. .................................. 62 Fig. 5.3 The outcomes form generated by GDL. ............................... 62 Fig. 5.4 The mind map representation of the uploaded SARA Observation archetype . ..................................................................... 64 87 LIST OF TABLES Table 2.1 The available attributes for each HPO class and their descriptions. ....................................................................................... 21 Table 3.1 Example scenario for the SARA scale. .............................. 33 Table 3.2 Information-processing units. ............................................ 46 Table 3.3 SARA items severity levels. .............................................. 48 Table 4.1 Agreement between automated and manual ratings (Weighted Kappa) .............................................................................. 58 Table 5.1. Kappa interpretation rules-Landis and Koch (1977) ........ 66 Table 5.2. Strength of agreement between automated and manual ratings. ................................................................................................ 66 89 APPENDIX A: The eight items included in the SARA scale Gait item A patient is asked (1) to walk at a safe distance parallel to a wall including a halfturn (turn around to face the opposite direction of gait) and (2) to walk in tandem (heels to toes) without support. Value Description 0 Normal, no difficulties in walking, turning and walking tandem (up to one misstep allowed) 1 Slight difficulties, only visible when walking 10 consecutive steps in tandem 2 Clearly abnormal, tandem walking >10 steps not possible 3 Considerable staggering, difficulties in half-turn, but without support 4 Marked staggering, intermittent support of the wall required 5 Severe staggering, permanent support of one stick or light support by one arm required 6 Walking > 10 m only with strong support (two special sticks or stroller or accompanying person) 7 Walking < 10 m only with strong support (two special sticks or stroller or accompanying person) 8 Unable to walk, even supported Stance item A patient is asked to stand (1) in natural position, (2) with feet together in parallel (big toes touching each other) and (3) in tandem (both feet on one line, no space between heel and toe). Proband does not wear shoes, eyes are open. For each condition, three trials are allowed. Best trial is rated. Value Description 0 Normal, able to stand in tandem for > 10 s 1 Able to stand with feet together without sway, but not in tandem for > 10s 2 Able to stand with feet together for > 10 s, but only with sway 3 Able to stand for > 10 s without support in natural position, but not with feet together 4 Able to stand for >10 s in natural position only with intermittent support 5 Able to stand >10 s in natural position only with constant support of one arm 6 Unable to stand for >10 s even with constant support of one arm APPENDIX A 90 Sitting item A patient is asked to sit on an examination bed without support of feet, eyes open and arms outstretched to the front. Value Description 0 Normal, no difficulties sitting >10 sec 1 Slight difficulties, intermittent sway 2 Constant sway, but able to sit > 10 s without support 3 Able to sit for > 10 s only with intermittent support 4 Unable to sit for >10 s without continuous support Speech disturbance item Speech is assessed during normal conversation Value Description 0 Normal 1 Suggestion of speech disturbance 2 Impaired speech, but easy to understand 3 Occasional words difficult to understand 4 Many words difficult to understand 5 Only single words understandable 6 Speech unintelligible / Anarthria Finger chase item A patient sits comfortably. If necessary, support of feet and trunk is allowed. Examiner sits in front of proband and performs 5 consecutive sudden and fast pointing movements in unpredictable directions in a frontal plane, at about 50 % of proband´s reach. Movements have an amplitude of 30 cm and a frequency of 1 movement every 2 s. Proband is asked to follow the movements with his index finger, as fast and precisely as possible. Average performance of last 3 movements is rated. The right and left sides are rated independently, then the mean of both sides is calculated. Value Description 0 No dysmetria 1 Dysmetria, under/ overshooting target <5 cm 2 Dysmetria, under/ overshooting target < 15 cm 3 Dysmetria, under/ overshooting target > 15 cm 4 Unable to perform 5 pointing movements APPENDIX A 91 Nose-finger test sara item A patient sits comfortably. If necessary, support of feet and trunk is allowed. Proband is asked to point repeatedly with his index finger from his nose to examiner’s finger which is in front of the proband at about 90 % of proband’s reach. Movements are performed at moderate speed. Average performance of movements is rated according to the amplitude of the kinetic tremor. The right and left sides are rated independently, then the mean of both sides is calculated. Value Description 0 No tremor 1 Tremor with an amplitude < 2 cm 2 Tremor with an amplitude < 5 cm 3 Tremor with an amplitude > 5 cm 4 Unable to perform 5 pointing movements Fast alternating hand movements item A patient sits comfortably. If necessary, support of feet and trunk is allowed. Proband is asked to perform 10 cycles of repetitive alternation of proand supinations of the hand on his/her thigh as fast and as precise as possible. Movement is demonstrated by examiner at a speed of approx. 10 cycles within 7s. Exact times for movement execution have to be taken. The right and left sides are rated independently, then the mean of both sides is calculated. Value Description 0 Normal, no irregularities (performs <10s) 1 Slightly irregular (performs <10s) 2 Clearly irregular, single movements difficult to distinguish or relevant interruptions, but performs <10s 3 Very irregular, single movements difficult to distinguish or relevant interruptions, performs >10s 4 Unable to complete 10 cycles Heel-shin slide item A patient lies on examination bed, without sight of his legs. Proband is asked to lift one leg, point with the heel to the opposite knee, slide down along the shin to the ankle, and lay the leg back on the examination bed. The task is performed 3 times. Slide-down movements should be performed within 1 s. If proband slides down without contact to shin in all three trials, rate 4. The right and left sides are rated independently, then the mean of both sides is calculated. Value Description 0 Normal 1 Slightly abnormal, contact to shin maintained 2 Clearly abnormal, goes off shin up to 3 times during 3 cycles 3 Severely abnormal, goes off shin 4 or more times during 3 cycles 4 Unable to perform the task APPENDIX B 98 Set element "Gait severity valueMAGNITUDE" to 2 Set element Gait Ataxia to Mild Rule Gait moderate When (( Element Gait equals to Considerable staggering, difficulties in half-turn, but without support ) or ( Element Gait equals to Marked staggering, intermittent support of the wall required )) Then Set element "Gait severity valueMAGNITUDE" to 3 Set element Gait Ataxia to Moderate Rule Gait severe When (( Element Gait equals to Severe staggering, permanent support of one stick or light support by one arm required ) or ( (( Element Gait equals to Walking > 10 m only with strong support (two special sticks or stroller or accompanying person) ) or ( Element Gait equals to Walking < 10 m only with strong support (two special sticks or stroller or accompanying person) )) )) Then Set element Gait Ataxia to Severe Set element "Gait severity valueMAGNITUDE" to 4 Rule Gait profound When Element Gait equals to Unable to walk, even supported Then Set element Gait Ataxia to profound Set element "Gait severity valueMAGNITUDE" to 5 APPENDIX B 99 Rule Stance normal When Element Stance equals to Normal, able to stand in tandem for > 10 s Then Set element Stance to Stance Set element Standing Instability to Normal Rule Stance borderline When Element Stance equals to Able to stand with feet together without sway, but not in tandem for > 10s Then Set element "Stance severity valueMAGNITUDE" to 1 Set element Standing Instability to Borderline Rule Stance mild When (( Element Stance equals to Able to stand with feet together for > 10 s, but only with sway ) or ( Element Stance equals to Able to stand for > 10 s without support in natural position, but not with feet together )) Then Set element "Stance severity valueMAGNITUDE" to 2 Set element Standing Instability to Mild Rule Stance moderate When Element Stance equals to Able to stand for >10 s in natural position only with intermittent support Then Set element "Stance severity valueMAGNITUDE" to 3 Set element Standing Instability to Moderate Rule Stance severe When Element Stance equals to Able to stand >10 s in natural position only with constant support of one arm Then APPENDIX B 100 Set element "Stance severity valueMAGNITUDE" to 4 Set element Standing Instability to Severe Rule Stance profound When Element Stance equals to Unable to stand for >10 s even with constant support of one arm Then Set element Standing Instability to profound Set element "Stance severity valueMAGNITUDE" to 5 Rule Sitting normal When Element Sitting equals to Normal, no difficulties sitting >10 sec Then Set element Sitting Instability to Normal Set element Sitting to Sitting Rule Sitting borderline When Element Sitting equals to Slight difficulties, intermittent sway Then Set element "Sitting severity valueMAGNITUDE" to 1 Set element Sitting Instability to Borderline Rule Sitting mild When Element Sitting equals to Constant sway, but able to sit > 10 s without support Then Set element "Sitting severity valueMAGNITUDE" to 2 Set element Sitting Instability to Mild Rule Sitting moderate When Element Sitting equals to Able to sit for > 10 s only with intermittent support Then Set element "Sitting severity valueMAGNITUDE" to 3 Set element Sitting Instability to Moderate Rule Sitting severe When APPENDIX B 101 Element Sitting equals to Unable to sit for >10 s without continuous support Then Set element "Sitting severity valueMAGNITUDE" to 4 Set element Sitting Instability to Severe Rule speech disturbance normal When Element Speech Disturbance equals to Normal Then Set element Speech Disturbance to Speech Disturbance Set element Dysarthria to Normal Rule speech disturbance borderline When Element Speech Disturbance equals to Suggestion of speech disturbance Then Set element Dysarthria to Borderline Rule speech disturbance mild When Element Speech Disturbance equals to Impaired speech, but easy to understand Then Set element Dysarthria to Mild Rule speech disturbance moderate When (( Element Speech Disturbance equals to Occasional words difficult to understand ) or ( Element Speech Disturbance equals to Many words difficult to understand )) Then Set element Dysarthria to Moderate Rule speech disturbance severe When Element Speech Disturbance equals to Only single words understandable APPENDIX B 102 Then Set element Dysarthria to Severe Rule speech disturbance profound When Element Speech Disturbance equals to Speech unintelligible / anarthria Then Set element Dysarthria to Profound Rule Finger chase right normal When Element Finger chase-right hand equals to No dysmetria Then Set element Upper Limb Dysmetria Right to Normal Set element Finger chase-right hand to Finger chase-right hand Rule Finger chase right mild When Element Finger chase-right hand equals to Dysmetria, under/ overshooting target <5 cm Then Set element "Finger chase right severity valueMAGNITUDE" to 2 Set element Upper Limb Dysmetria Right to Mild Rule Finger chase right moderate When Element Finger chase-right hand equals to Dysmetria, under/ overshooting target < 15 cm Then Set element "Finger chase right severity valueMAGNITUDE" to 3 Set element Upper Limb Dysmetria Right to Moderate Rule Finger chase right severe When (( Element Finger chase-right hand equals to Dysmetria, under/ overshooting target > 15 cm ) or ( Element Finger chase-right hand equals to Unable to perform 5 pointing movements )) APPENDIX B 103 Then Set element "Finger chase right severity valueMAGNITUDE" to 4 Set element Upper Limb Dysmetria Right to Severe Rule Finger chase left normal When Element Finger chase-left hand equals to No dysmetria Then Set element Upper Limb Dysmetria Left to Normal Set element Finger chase-left hand to Finger chase-left hand Rule Finger chase left mild When Element Finger chase-left hand equals to Dysmetria, under/ overshooting target <5 cm Then Set element "Finger chase left severity value MAGNITUDE" to 2 Set element Upper Limb Dysmetria Left to Mild Rule Finger chase left moderate When Element Finger chase-left hand equals to Dysmetria, under/ overshooting target < 15 cm Then Set element "Finger chase left severity value MAGNITUDE" to 3 Set element Upper Limb Dysmetria Left to Moderate Rule Finger chase left severe When (( Element Finger chase-left hand equals to Dysmetria, under/ overshooting target > 15 cm ) or ( Element Finger chase-left hand equals to Unable to perform 5 pointing movements )) Then Set element "Finger chase left severity value MAGNITUDE" to 4 Set element Upper Limb Dysmetria Left to Severe Rule nose finger test right normal When Element Nose-finger test-right hand equals to No tremor APPENDIX B 104 Then Set element Intention Tremor Right to Normal Set element Nose-finger test-right hand to Nose-finger testright hand Rule nose finger test righ mild When Element Nose-finger test-right hand equals to Tremor with an amplitude < 2 cm Then Set element "Nose finger test right severity valueMAGNITUDE" to 2 Set element Intention Tremor Right to Mild Rule nose finger test righ moderate When Element Nose-finger test-right hand equals to Tremor with an amplitude < 5 cm Then Set element "Nose finger test right severity valueMAGNITUDE" to 3 Set element Intention Tremor Right to Moderate Rule nose finger test righ severe When (( Element Nose-finger test-right hand equals to Tremor with an amplitude > 5 cm ) or ( Element Nose-finger test-right hand equals to Unable to perform 5 pointing movements )) Then Set element "Nose finger test right severity valueMAGNITUDE" to 4 Set element Intention Tremor Right to Severe Rule nose finger test left normal When Element Nose-finger test-left hand equals to No tremor Then Set element Intention Tremor Left to Normal APPENDIX B 105 Set element Nose-finger test-left hand to Nose-finger test-left hand Rule nose finger test left mild When Element Nose-finger test-left hand equals to Tremor with an amplitude < 2 cm Then Set element "Nose finger test left severity valueMAGNITUDE" to 2 Set element Intention Tremor Left to Mild Rule nose finger test left moderate When Element Nose-finger test-left hand equals to Tremor with an amplitude < 5 cm Then Set element "Nose finger test left severity valueMAGNITUDE" to 3 Set element Intention Tremor Left to Moderate Rule nose finger test left severe When (( Element Nose-finger test-left hand equals to Tremor with an amplitude > 5 cm ) or ( Element Nose-finger test-left hand equals to Unable to perform 5 pointing movements )) Then Set element "Nose finger test left severity valueMAGNITUDE" to 4 Set element Intention Tremor Left to Severe Rule fast alternating hand movements right normal When Element Fast alternating hand movements-right hand equals to Normal, no irregularities (performs <10s) Then Set element Fast alternating hand movements-right hand to Fast alternating hand movements-right hand Set element Dysdiadochokinesis Right to Normal Rule fast alternating hand movements right mild When APPENDIX B 106 Element Fast alternating hand movements-right hand equals to Slightly irregular (performs <10s) Then Set element "Fast alternating hand movements right severity value MAGNITUDE" to 2 Set element Dysdiadochokinesis Right to Mild Rule fast alternating hand movements right moderate When Element Fast alternating hand movements-right hand equals to Clearly irregular, single movements difficult to distinguish or relevant interruptions, but performs <10s Then Set element Dysdiadochokinesis Right to Moderate Set element "Fast alternating hand movements right severity value MAGNITUDE" to 3 Rule fast alternating hand movements right severe When (( Element Fast alternating hand movements-right hand equals to Very irregular, single movements difficult to distinguish or relevant interruptions, performs >10s ) or ( Element Fast alternating hand movements-right hand equals to Unable to complete 10 cycles )) Then Set element "Fast alternating hand movements right severity value MAGNITUDE" to 4 Set element Dysdiadochokinesis Right to Severe Rule fast alternating hand movements left normal When Element Fast alternating hand movements-left hand equals to Normal, no irregularities (performs <10s) Then Set element Dysdiadochokinesis Left to Normal Set element Fast alternating hand movements-left hand to Fast alternating hand movements-left hand APPENDIX B 107 Rule fast alternating hand movements left mild When Element Fast alternating hand movements-left hand equals to Slightly irregular (performs <10s) Then Set element "Fast alternating hand movements left severity valueMAGNITUDE" to 2 Set element Dysdiadochokinesis Left to Mild Rule fast alternating hand movements left moderate When Element Fast alternating hand movements-left hand equals to Clearly irregular, single movements difficult to distinguish or relevant interruptions, but performs <10s Then Set element "Fast alternating hand movements left severity valueMAGNITUDE" to 3 Set element Dysdiadochokinesis Left to Moderate Rule fast alternating hand movements left severe When (( Element Fast alternating hand movements-left hand equals to Very irregular, single movements difficult to distinguish or relevant interruptions, performs >10s ) or ( Element Fast alternating hand movements-left hand equals to Unable to complete 10 cycles )) Then Set element "Fast alternating hand movements left severity valueMAGNITUDE" to 4 Set element Dysdiadochokinesis Left to Severe Rule heel shin slide right normal When Element Heel-shin slide-right hand equals to Normal Then Set element Heel-shin slide-right hand to Nose-finger testright hand Set element Lower Limb Dysmetria Right to Normal APPENDIX B 114 Element Sara Total Score is greater than or equals to 3 Element Appendicular right severity value equals to 2 Then Set element Appendicular Ataxia Right to Mild Rule Appendicular right moderate When Element Appendicular right severity value equals to 3 Element Sara Total Score is greater than or equals to 3 Then Set element Appendicular Ataxia Right to Moderate Rule Appendicular right severe When Element Sara Total Score is greater than or equals to 3 Element Appendicular right severity value equals to 4 Then Set element Appendicular Ataxia Right to Severe Rule Appendicular left and finger chase left When Element Appendicular left severity value is less than Finger chase left severity value Then Set element Appendicular left severity value to Finger chase left severity value Rule Appendicular left and nose finger test left When Element Appendicular left severity value is less than Nose finger test left severity value Then Set element Appendicular left severity value to Nose finger test left severity value Rule Appendicular left and fast alternating hand movements left When Element Appendicular left severity value is less than Fast alternating hand movements right severity value Then Set element Appendicular left severity value to Fast alternating hand movements right severity value APPENDIX B 115 Rule Appendicular left and heel shin slide left When Element Appendicular left severity value is less than Heel-shin slide left severity value Then Set element Appendicular left severity value to Heel-shin slide left severity value Rule Appendicular left normal When (( Element Sara Total Score is less than 3 ) or ( Element Appendicular left severity value equals to 0 )) Then Set element Appendicular Ataxia Left to Normal Rule Appendicular left mild When Element Sara Total Score is greater than or equals to 3 Element Appendicular left severity value equals to 2 Then Set element Appendicular Ataxia Left to Mild Rule Appendicular left moderate When Element Appendicular left severity value equals to 3 Element Sara Total Score is greater than or equals to 3 Then Set element Appendicular Ataxia Left to Moderate Rule Appendicular left severe When Element Sara Total Score is greater than or equals to 3 Element Appendicular left severity value equals to 4 Then Set element Appendicular Ataxia Left to Severe 117 ACRONYMS AND ABBREVIATIONS ADL Archetype Definition Language AM Archetype Model APGAR Appearance, Pulse, Grimace, Activity and Respiration CDA Clinical Document Architecture CDE Common Data Element CDS Clinical Decision Support CKM Clinical Knowledge Manager DL Description Logic ECG Electrocardiography EHR Electronic Health Record GCS Glasgow Comma Scale GDL Guideline Definition Language HL7 Health Level 7 HPO Human Phenotype Ontology ICD International Classification of Diseases MMSE Mini-Mental State Examination MRS Modified Rankin Scale MS Multiple Sclerosis NINDS National Institute of Neurological Disorders and Stroke OBO Open Biomedical Ontology OWL Web Ontology Language RM Reference Model RM Reference Model SARA Scale for the Assessment and Rating of Ataxia SCA36 Spinocerebellar Ataxia Type 36 SMS SARA Management System SNHL Slowly progressive sensorineural hearing loss UPDRS Unified Parkinson’s Disease Rating Scale