On diagnostic codes
Full text
On diagnostic codes Amalie Dahl Haue This thesis has been submitted to the Graduate School of Health and Medical Sciences University of Copenhagen August 19, 2021
Preamble The title of this thesis joins an integral part of clinical practice and a cornerstone of any computer language. While diagnosis codes represent a systematic answer to causes of disease and death; a diagnostic code may both refer to the content of a script and the dual nature of clinical practice. Since the earliest evidence of clinical practice, diagnostics has evolved from a practice that concluded manual inspection, palpation and functional assessment to standards evolving in a system that is heavily dependent on computers. This thesis explores the boundaries of diagnostics in a digitized reality. Diagnosis Patient Physician TRUE stands for CORRECT symbolizes refers to ADEQUATE Diagnostics and reality.∗ ∗Adapted from The Meaning of Meaning: A Study of the Influence of Language upon Thought and of the Science of Symbolism. Charles Kay Ogden &Ivor Armstrong Richards. Harcourt, Brace & World, Inc. New York (1925); p. 11.
Table of contents FRONT MATTER Preface i List of manuscripts ii Summary iii Summary in Danish v PART I SYNOPSIS 0 Objectives and overview 1 1 Introduction 3 1.1 Diagnostics in a web of multiple causes and consequences . . . . . . . . . . . . . . . 3 1.1.1 Ischemic heart disease is a common, chronic, multi-factorial disease . . . . . 3 1.1.2 Ischemic heart disease often co-exists with other common chronic diseases . 5 1.2 Classification systems harmonize healthcare data . . . . . . . . . . . . . . . . . . . . 6 1.2.1 Classification systems for medical causes and consequences . . . . . . . . . . 6 1.2.2 Classification systems that reflect clinical practice and examination . . . . . . 8 1.3 Strengths and limitations of medical classification systems in research . . . . . . . . 8 2 Materials 10 2.1 The Danish healthcare system as a research resource . . . . . . . . . . . . . . . . . . . 10 2.2 Data from millions of people over multiple decades . . . . . . . . . . . . . . . . . . . 11 2.2.1 The Danish National Patient Registry . . . . . . . . . . . . . . . . . . . . . . . 11 2.2.2 The Danish Registry for Causes of Death . . . . . . . . . . . . . . . . . . . . . 11 2.2.3 The Danish National Prescription Registry . . . . . . . . . . . . . . . . . . . . 12 2.3 Deeper data characterize patients with greater precision . . . . . . . . . . . . . . . . 12 2.3.1 The Eastern Danish Heart Registry . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.3.2 Electronic Health Records from Eastern Denmark . . . . . . . . . . . . . . . . 13 2.3.3 Copenhagen Hospital Biobank . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
3 Methods 14 3.1 Data, science, and scientific models evolve over time . . . . . . . . . . . . . . . . . . . 14 3.2 Comprehensive analysis of nationwide health registries . . . . . . . . . . . . . . . . . 16 3.2.1 Describing multi-morbidity using temporal disease trajectories . . . . . . . . 16 3.2.2 Targeting disease trajectories multi-morbidity in ischemic heart disease . . . 18 3.2.3 Prescription trajectories model multi-morbidity in the general population . . 19 3.3 Cluster analysis to define ischemic heart disease subgroups . . . . . . . . . . . . . . . 20 3.3.1 A model where multi-morbidity and text documents share properties . . . . 20 3.3.2 A mathematical graph representation of ischemic heart disease patients . . . 21 3.3.3 A cluster algorithm minimizes a cost function . . . . . . . . . . . . . . . . . . 22 3.3.4 Characterization of clusters using survival analysis . . . . . . . . . . . . . . . 23 3.4 Data-driven predictions of mortality in ischemic heart disease . . . . . . . . . . . . . 25 3.4.1 A conceptual introduction to artificial neural networks . . . . . . . . . . . . . 25 3.4.2 Ranking input features based on level of information . . . . . . . . . . . . . . 27 3.5 A probabilistic map of in-hospital drug dosage alterations . . . . . . . . . . . . . . . 28 3.5.1 Bayesian models can be decomposed into three distributions . . . . . . . . . . 28 3.5.2 Application of Bayes’ theorem to electronic healthcare data . . . . . . . . . . 30 4 Data protection and privacy 31 4.1 Informed consent and institutional permissions . . . . . . . . . . . . . . . . . . . . . . 31 4.2 Legal regulations and data management . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.3 Summary of authority approvals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 5 Results 34 5.1 A temporal analysis of multi-morbidity in ischemic heart disease . . . . . . . . . . . 34 5.2 A strategy to define more homogeneous ischemic heart disease patient subgroups . 36 5.3 Mortality in ischemic heart disease from a time-to-event machine learning algorithm 38 5.4 Evidence of drug-drug interactions in electronic health records . . . . . . . . . . . . 40 5.5 Establishing a map of longitudinal, nation-wide prescription patterns . . . . . . . . . 42 5.6 Internet applications display perspectives with data-driven research . . . . . . . . . 43 6 Discussion 46 6.1 Conclusion........................................... 46 6.2 Strengthsandlimitations................................... 47 6.3 Perspectives .......................................... 48 References 50 List of abbreviations 60
PART II FULL-LENGTH MANUSCRIPTS 63 7 Temporality in ischemic heart disease multi-morbidity 64 8 Risk stratification of 72,249 patients with ischemic heart disease: a retrospective study linking prior multi-morbidity, biochemical data, and genetics 89 9 PMHnet-alpha: a neural network-based discrete-time survival model for mortality prediction in ischemic heart disease 112 10 Polypharmacy and drug dosage modifications: a longitudinal analysis of 3.5 million electronic health records 135 11 Optimizing drug selection from a prescription trajectory of one patient 168 12 Disease trajectory browser for exploring temporal, population-wide disease progression patterns in 7.2 million Danish patients 199 PART III APPENDICES 210 A Appendix 211 B Appendix 214 C Appendix 237 D Appendix 253 E Appendix 269
Preface This thesis was written to complete the Graduate programme in Biostatistics and Bioinformatics at The Faculty of Health and Medical Sciences at University of Copenhagen and thereby fulfill the requirements for obtaining a PhD degree in agreement with the ministerial Order on the PhD Programme at the Universities and Certain Higher Artistic Educational Institutions (PhD Order). Financially, the work was carried out with generous support from The Department of Health and Medical Sciences, University of Copenhagen (CAG Precision Diagnostics in Cardiology), The Novo Nordisk Foundation (grant agreements NNF14CC0001 and NNF17OC0027594) and Innovationsfonden (project numbers 60833 and 3-3031-1731/1). For countless valuable discussions, timely insight, and firm trust, I thank family (immediate and extended), friends, colleagues, and the academic advisors. Individually and in concert you have expanded horizons, induced perspective; and encouraged me to find my own trajectory in this model of reality called science and the corner of reality known as clinic practice. Amalie Dahl Haue Copenhagen 2021 i
List of manuscripts 1. ∗Haue AD, Armentaros JJA, Holm PC, Moseley PL, Køber LV, Bundgaard H, Brunak S. Temporality in ischemic heart disease multi-morbidity. Manuscript in preparation. 2. ∗Haue AD, Holm PC, Banasik K, Lundgaard AT, Muse VP, Roder T, Westergaard D, Chmura P, Siggaard T, Christensen AH, Weeke PE, Sørensen E, Ostrowski SR, DF Gulbjartsson, Holm H, Iversen KK, Køber LV, Ullum H, Bundgaard H, Brunak S. Risk stratification of 72,249 patients with ischemic heart disease: a retrospective study linking prior multi-morbidity, biochemical data, and genetics. Manuscript submitted. 3. ∗Holm PC, Haue AD, Banasik K, Brunak S, Bundgaard H. PMHnet-alpha: a neural networkbased discrete-time survival model for mortality prediction in ischemic heart disease. Manuscript in preparation. 4. †Rodríguez CL, Mazonni G, Haue AD, Eriksson R, Biel JH, Cantwell L, Westergaard D, Belling KG, Brunak S. Polypharmacy and drug dosage modifications: a longitudinal analysis of 3.5 million electronic health records. Manuscript in revision. 5. †Aguayo-Orozco A, Haue AD, Jørgensen IF, Westergaard D, Moseley PL, Mortensen LH, Brunak S. Optimizing drug selection from a prescription trajectory of one patient. Manuscript in revision. 6. †Siggaard T, Reguant R, Jørgensen IF, Haue AD, Lademann M, Aguayo-Orozco A, Hjaltelin JX, Jensen AB, Banasik K, Brunak S. Disease trajectory browser for exploring temporal, population-wide disease progression patterns in 7.2 million Danish patients. Nat Commun. 2020 Oct 2;11(1):4952. DOI: https://doi.org/10.1038/s41467-020-18682-4. ∗: Main manuscripts †: Contributory manuscripts ii
Summary Introduction As the digitized reality is transforming modern medicine, data volumes and variety are growing considerably while the visions for healthcare are also being revised. Yet it remains an open question how to convert this wealth of data into clinically actionable evidence that is becoming even more needed in an ageing population. In this thesis, I present six original studies that investigate the potential of Danish healthcare data as a platform to perform data-driven research. The thesis has a strong focus on multi-morbidity in ischemic heart disease (IHD) and the applied methods range from classical regression models to machine learning methods. Multi-morbidity in IHD is both characterized with respect to the temporality between diagnoses and by analyzing more fine-grained healthcare data, such as electronic health records (EHRs) and genetic data. The practical consequences of multi-morbidity are also studied based on analyses of polypharmacy in an in-hospital setting and in the general population. Finally, the thesis presents an example sharing research results as a web tool made available to others without sharing person-sensitive data. Manuscript 1: Temporality in ischemic heart disease multi-morbidity. This study characterizes multi-morbidity in IHD using The Danish National Patient Registry (NPR) data from IHD patients in the period years 1994–2018. Multi-morbidity is characterized by application of a new version of the so-called disease trajectory program that establishes temporal associations between diagnoses. The study showcases differences between disease trajectories containing conditions that are solely risk factors, and conditions that can be both risk factors and complications. The study argues that disease trajectories can be used to systematically assess the temporality in multi-morbidity, which is important, yet often an underappreciated aspect, of multi-morbidity. Manuscript 2: Risk stratification of 72,249 patients with ischemic heart disease: a retrospective study linking prior multi-morbidity, biochemical data, and genetics. In this study, multimorbidity in IHD is addressed by characterizing distinct patient subgroups obtained by application of a cluster algorithm to diagnosis codes. In a cohort of roughly 75 000 patients, patient subgroups are identified based on similarities between patient-specific vectors that represent all diagnoses registered in NPR up to 24 years before IHD onset. Data from EHRs and Copenhagen Hospital biobank (CHB) are included to assess phenotypic differences in these patient subgroups obtained by application of the cluster algorithm. The patient subgroups associate with different progression rates as well as biochemical and genetic differences, indicating that the patterns captured by the cluster algorithm are biologically relevant. iii
Manuscript 3: PMHnet-alpha: a neural network-based discrete-time survival model for mortality prediction in ischemic heart disease. This study presents PMHnet-alpha, which is a neural network-based survival model for prediction of all-cause mortality in IHD patients. PMHnetalpha is based on a derivation cohort of approximately 40 000 IHD patients with verified coronary pathology according to The Eastern Danish Heart Registry (EDHR) data. Roughly 600 input features extracted from EDHR, NPR, EHRs, and CHB, including diagnoses, blood tests, and polygenic scores are the basis for development of PMHnet-alpha that outperforms the GRACE2.0 risk score. PMHnet-alpha is also subjected to explainability analysis where predictive valuesof input features are ranked and characterized as positive or negative depending on how they affect the prediction. Manuscript 4: Polypharmacy and drug dosage modifications: a longitudinal analysis of 3.5 million electronic health records In this study, polypharmacy is studied with reference to co-medication pairs that are defined by a combination of two distinct drugs where the likelihood of a dosage adjustment is more likely during concomitant treatment episodes than when administered as a monotherapy. By extracting information on drug regimens from EHRs from about one million inhospital patients in the period years 2008–2016, the study presents 3993 co-medication pairs that were cross-references with 15 publicly available drug-drug interaction databases. We argue that the analysis provides evidence for up to 600 previously undescribed drug-drug interactions which were identified as co-medication pairs and not present in at least one of the drug-drug interaction databases. Manuscript 5: Optimizing drug selection from a prescription trajectory of one patient. This study addresses polypharmacy by analyzing data from The Danish National Prescription Registry (DNPR) in the period years 1995–2019 comprehensively. The study presents the so-called prescription trajectories that map the prescription patterns from more than seven million people present in DNPR. Prescription trajectories contain up to seven distinct redeemed drugs that are more likely to be redeemed in one specific order. In the study, we argue that the analysis can support optimized treatment decisions by analyzing differences between people subjected to many changes in treatment regimens which can hopefully lead to more individualized care. Manuscript 6: Disease trajectory browser for exploring temporal, population-wide disease progression patterns in 7.2 million Danish patients. This study presents The Danish Disease Trajectory Browser, which is a web application that enables users to explore data from NPR in the period years 1994–2018 based on results from the disease trajectory program. Thus, the study exemplifies newer possibilities for sharing research results. In the browser, users can search for one or more diagnoses and search results will be displayed in the form of disease trajectories. The browser has an array of search filters that users can apply to refine searches as well as an interface that allows to navigate the search results and modify the visual representation. Search results can be exported in the form of graphical representations of disease trajectories and summary statistics. iv
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES Figure 1: Decline in cardiovascular diseases in relation to scientific advances. CABG: Coronary artery by-pass grafting. PCI: Percutanous coronary intervention. MI: Myocardial infarction. NHBPEP: National High Blood Pressure Educational Program. CASS: Coronary artery surgery study. TIMI: Thrombolysis in myocardial infarction. NCEP: The National Cholesterol Education Program. GISSI: Gruppo Italioano per la Sperimentazione della stretochinasi nell’INfarto Miocardico. ISIS-2: The Second Study of Infarct Survival. SAVE: Survival and Ventricular Enlargement Trial. ALLHAT: Antihypertensive and Lipid-Lowering Treatment to Prevent Heart Attack Trial. From [9]. To date, the Global Registry of Acute Coronary Events (GRACE) has provided the data foundation for one of the most widely used risk scoring schemes for IHD patients[11]. GRACE refers to a large multi-national registry, where patients with acute coronary syndrome were recruited from 14 different countries over the period from years 1999 to 2009. Data from GRACE comprise the foundation for a risk score of the same name that estimates shortand long-term mortality in patients with acute MI. The GRACE risk score has been updated several times and expresses the risk of six month mortality on a scale of 1to 372, where a score of 1is considered minimal risk[12]–[14]. Recent advances within genetic research have also paved the way for development of polygenic scores (PGSs) that have fueled perspectives within the precision medicine agenda, including more precise risk estimates[15]. There is already data supporting the hypothesis that polygenetic predictors can be used to identify individuals at increased risk of IHD and other multi-factorial diseases equivalent to that of genetic predictors of monogenic diseases. Further, it has been stipulated that the increasing interest in recycling population-wide healthcare data combined with continuous expansion ofcomputationalcapacity, present endlessopportunities of broadly integratinggeneticdata into clinical decision making[16], [17]. Ideally, principles that integrate conventional pathology obtained from CAGs and other clinical examinations combined with deep phenotypes harvested from 4
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES human data collections will increase healthcare quality for the individual[18]. However, the specific strategies that will translate these perspectives into clinical practice have not been fully established yet and accordingly much explorative work is being conducted within this field[19]. 1.1.2 Ischemic heart disease often co-exists with other common chronic diseases Up to 85% of IHD patients are co-morbid, broadly defined as a phenotype comprised of more than one chronic disease [9], [20], [21]. Co-morbidity and multi-morbidity are often used interchangeably. Technically, co-morbidity describes the phenotype with reference to an index disease, whereas multi-morbidity denotes the phenotype without reference to a particular index disease[22], [23]. In this thesis, I have chosen to use the term multi-morbidity in IHD to ensure consistency and stress that at any point in time several conditions may be equally important[24]. The importance of multi-morbidity has been articulated for more than 50 years[25]. Back then, there was already a focus on the consequences of failure to classify and analyze multi-morbidity owing to the potential causal role of multi-morbidity in prognosis and therapeutic effects. In randomized controlled trials (RCTs) and other types of clinical studies, unacknowledged multi-morbidity may confound or modify the results, making efficient methods for measuring multi-morbidity an unmet need[26]. Nevertheless, there is no consensus regarding multi-morbidity measures and the existing tools for measuring multi-morbidity vary greatly in validity, reliability and importantly, target population[26]. Historically, the most widely used co-morbidity measure is the Charlson co-morbidity index (CCI), which was derived from the correlation of 19 selected diseases with mortality among cancer patients. In a cohort of 685 breast cancer patients, the investigators found that only age and co-morbidity were risk factors of co-morbid death, i.e. death from other causes than metastatic disease[27]. Although there does not exist any specific multi-morbidity score for IHD patients, the evidence that multi-morbidity associates with the prognosis and burden of the individual patient is endless and in practice, CCI is widely used irrespective of main phenotype[28]. In acknowledging the prognostic interdependency of seemingly distinct diseases, it has been suggested that delivery of IHD healthcare should become more multidisciplinary rather than primarily designed to target one condition at a time[20], [29]. Examples of multidisciplinary approaches include integration of knowledge across different medical specialties and enhanced focus on patient preferences within a framework that is targeted treatment of IHD and other co-existing chronic conditions, i.e., multimorbidity[9]. A common therapeutic consequence of multi-morbidity is a condition known as polypharmacy, referring to the simultaneous usage of multiple drugs by a single patient[30]. And generally, there is a knowledge gap regarding benefits, risks, and intensity of treatment in cardiovascular patients, with an increasing prevalence as well as age[31]. Major concerns with polypharmacy are that it increases the risk of adverse reactions, reduces the likelihood of compliance, and there is very limited evidence regarding the long-term consequences of the drugs that are often prescribed life-long[30], [32]. In fact, the increasing complexity of medical therapy is gaining attention in so-called adaptive platform trials, where multiple therapeutic interventions may be compared in the same trials[33]. 5
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES 1.2 Classification systems harmonize healthcare data 1.2.1 Classification systems for medical causes and consequences Originally being the backbone of comparable mortality statistics dating back as long as to 1893, the International Statistical Classification of Diseases and Related Health Problems (ICD) has turned into an important resource in healthcare systems which are increasingly influenced by advances in information technology[34]. One inherent advantage of the ICD terminology is that it is generalizable across many different healthcare systems, as it is used in many countries[35]. Since 1994, the 10th version of the ICD system (ICD-10) has been the official diagnostic reporting system between Danish hospitals and the Danish health authorities[36]. The hierarchical structure of the classification system is comprised of 21 chapters that largely correspond to different functional-anatomical body systems. For example, chapter IX corresponds to the Diseases of the circulatory system. Each chapter can be divided into blocks that are further segmented into diagnosis codes at different resolutions, e.g. level 3 and 4 codes with level 4codes being the most specific. Level 4codes amount to a tabular listing of more than 155 000 distinct diagnosis codes (Figure 2). Figure 2: A search in the ICD-10 terminology for IHD where the level 3ICD-10 code chronic IHD was selected as indicated by the arrows in the panel on the left side. ICD-10 chapters, blocks, level 3and level 4codes are listed in the panel on the left. Descriptions of the selected level 3and 4 codes are on available on the right side. ICD-10: International Classification of Diseases and Related Health Problems, 10th version. IHD: Ischemic heart disease. Retrieved from https://icd.who.int/ browse10/2019/en#/I25. 6
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES The Anatomical Therapeutic Chemical (ATC) Classification System was originally developed as a tool to correctly interpret data on drug utilization[37]. Like ICD-10 codes, the ATC Classification System has a hierarchical structure, where groups at five different levels are used to classify each drug at increasing specificity. The groups range from the 1st level codes that consist of 14 anatomical groups specifying the primary organ system of action, to the 5th level codes, which specify the active chemical substance of a given drug[38]. In contrast to the ICD-10 terminology that has one unique code per diagnosis; the ATC Classification System may have more than one code for one drug. In cases where a single chemical substance has multipe routes of administration, this is standard. Similarly, although non-standard, chemical substances with more than one approved indication may also have more than one ATC code (Figure 3). Figure 3: Search results in the ATC Index for the 5th level ATC code B01AC06 which is acetylsalicylic acid corresponding to its indication as an antithrombotic agent. Green insert indicates 1st through 4th ATC Index levels. The description below explains the reason for the exemption that acetylsalicylic acid may also be classified in the chemical subgroups N02BA that contains salicyclic acid. In this case, the indication determines the ATC code. ATC: Anatomical Therapeutic Chemical. Retrieved from https://www.whocc.no/atc_ddd_index/. 7
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES 1.2.2 Classification systems that reflect clinical practice and examination The Nomenclature, Properties and Units in Laboratory Medicine (NPU) terminology, represented by NPU codes, index analyses within clinical laboratory sciences. NPU codes have been used since 1987 and are owned by the International Federation of Clinical Chemistry and Laboratory Medicine and International Union of Pure and Applied Chemistry. Since 1987 the number of available examinations from clinical laboratories has increased from a few hundreds to more than 30 000 different types of tests[39]. NPU codes were developed to convey information regarding the property of a studied object, i.e. the patient. The three essential elements of an NPU code are (i) the system, (ii) the component, and (iii) the kind-of-property. For example, an NPU code for troponin is NPU18583 which is defined by P—Troponin I, cardiac muscle; subst.c.= ? nmol/L. The component is "Troponin I, cardiac muscle". Anything before the component refers to the system (P in this case for plasma) and anything after refers to the kind-of-property, which in this case is the concentration measured in nmol/L. Importantly, NPU codes do not provide any information regarding the correctness of the measurement also meaning that reference intervals are defined locally at the clinical laboratory where the analyses are being performed[39]. Clinical examinations and procedures are yet another aspect of medical care that is being organized according to a terminology. For example, the Nordic Medico-Statistical Committee (NOMESCO) was established in the early 1980s to compare the frequencies of surgical activities in the Nordic countries. The first NOMESCO Classification of Surgical Procedures (NCSP) was published in 1996 and contains 20 chapters. Every NCSP code consists of three letters and three digits. The first letter of the NCSP code corresponds to the chapter. Similar to the case for ICD-10 and ATC codes, NCSP chapters are arranged according to functional-anatomic body systems. For example, the NCSP code for PCI is FNG05, where F indicates that it is surgery performed on the heart of major thoracic vessels and N specifies that it is performed on the coronary arteries[40]. In reporting from the Danish healthcare system, procedure codes are generally more reliable than diagnosis codes[41]. 1.3 Strengths and limitations of medical classification systems in research In many healthcare systems around the world, ICD-10 and ATC codes are an efficient means of comparing death statistics, drug usage, and handling billing purposes. Similarly, NPU and NCSP codes are instrumental administrative resources. Recently, the value of healthcare data derived from other sources than RCTs has been referred to as real-world evidence and is gaining attention as a resource to define other endpoints than all-cause mortality[42]. Potentially, this is highly valuable in an ageing population where factors such as quality of life and therapeutic efficacy are becoming increasingly important aspects of healthcare[31]. Although classification systems and clinical terminologies are necessary in this respect, they are not particularly designed to reflect differential diagnostic didactic nor to encompass the true complexity of multi-morbidity. In other words, there is no guarantee that a patient who has been assigned a diagnosis code also was subjected to adequate diagnostic examination. This also applies to ATC codes in the sense that if a person has redeemed a prescription for a drug, there is no evidence of compliance, i.e. whether or not the per8
CHAPTER 1. INTRODUCTION ON DIAGNOSTIC CODES son took the drug as prescribed. Similarly, an NCSP code as such has no information regarding the conclusion of the procedure that it describes. Essentially, in isolation no terminology provides any information regarding the true phenotype, where increasing prevalences of multi-morbidity and polypharmacy complicate matters even further. Moreover, even in the simplest cases, diagnostics remain an incomplete representation of reality and still automatized diagnostic systems are not used routinely[43], [44]. Yet, the classification systems presented in section 1.2 have become important tools to classify data points and harmonize healthcare data across countries. These efforts are particularly important in order to transform as much data as possible in to a foundation for improved medical treatment[42]. The adaption of ICD-10 and ATC world-wide has particularly been driven by the World Health Organization (WHO) leadership driving the ICD[45]. The systems are also being updated continuously reflecting a medical field that is constantly undergoing development. For example, ICD-11 in which the ATC codes are embedded, has been designed for use in a digital world and planned to come into effect in 2022[46]. The aphorism“garbagein, garbage out”hasresonatedbehind software ever sincethe earliest days of computer power[47]. And its implication remains the same. Thus, to make up for the fact that diagnostics in the classical sense was not performed over the course data-driven studies, including the onespresented in thisthesis; successfuladaptation ofmedicalclassificationsystemsis absolutely necessary to generate valuable data-driven models. In fact, it has been suggested quite recently to create more data-driven and mechanistically oriented disease classification systems that would match the aims of the precision medicine agenda better[48]. By discriminating between apparently similar phenotypes, using a wide range of data (e.g. registry, molecular, and genetic data) the diagnostic process may become much more individualized and in effect support better treatment decisions. A concrete attempt is the Human Phenotype Ontology (HPO) that now increasingly makes its way into healthcare systems[49]. The work in this thesis has a similar aim and perspective. 9
2|Materials This chapter gives an overview of the data sources that were the basis for exploring the potential of conducting data-driven research in the setting of Danish healthcare data. Section 2.1 introduces the basic structures within the Danish healthcare system of relevance to studies presented in this thesis, while section 2.2 describes the relevant nationwide registries and section 2.3 presents the sources of deeper phenotypic and genetic data. 2.1 The Danish healthcare system as a research resource Two of the essential aspects of the Danish healthcare system is that redistribution of taxes covers about 85% of the healthcare and the role of personal identification numbers[36]. Personal identification numbers are ten-digit numbers that since 1968 have identified every Danish citizen uniquely. The Danish Civil Registration System (CRS) administers all personal identification numbers[50]. Personal identification numbers function as unique identifiers of citizens across all public sectors including healthcare services[51]. Thus, in medical research based on Danish healthcare data, CRS facilitates linkage of different data sources given the study has been approved appropriately (cf. section 4). In addition to the linkage function, CRS contains information such as date of birth, sex and status (e.g. alive, dead, or emigrant) of the individual citizen. Figure 4 shows the data sources that were linked via CRS in this thesis prior to analysis. 1970 1980 1990 2000 2010 CRS† DAR∗NPR∗ DNPR∗EDHR•BTH• CHB‡ Figure 4: Overview of data foundation and resources ordered by year of first observation in resource. †: Foundation for Danish healthcare data infrastructure. ∗: Nationwide registries. •: Regional healthcare data with deeper phenotypic data. ‡: Genetic data. CRS: The Danish Civil Registration System. DAR: The Danish Registry of Causes of Death. NPR: The Danish National Patient Registry. EDHR: Eastern Danish Heart Registry. BTH: BigTempHealth (EHRs from Eastern Denmark). CHB: Copenhagen Hospital Biobank. EHRs: Electronic health records 10
CHAPTER 2. MATERIALS ON DIAGNOSTIC CODES 2.2 Data from millions of people over multiple decades 2.2.1 The Danish National Patient Registry The Danish National Patient Registry (NPR) was established in 1977 and is one of the oldest nationwide healthcare registries in the world[36]. It was established with the primary aim of continuous monitoring of hospital and healthcare service utilization for the Danish Health Authority. NPR contains data related to all discharges from Danish hospitals, which are reported in the form of timestamped patient contacts reflecting that activities within Danish hospitals are structured according to the so-called danske kontaktmodel[52]. A patient contact contains information such as duration and type of contact (e.g., in-hospital, out-hospital or emergency room visit), diagnosis codes assigned during the contact, type of diagnosis code, codes that document procedures performed during the contact and a myriad of other types of structured administrative information. One contact may span a few hours or several years, depending on the contact type and reason[53]. In NPR, registrations are reported in accordance with Sundhedsvæsenets Klassifikationssystem (SKS), which is organized in chapters corresponding to different aspects of healthcare.4For example, chapter D corresponds to the ICD-10 classification and chapter K contains NCSP codes. Yet other chapters contain Danish classification systems. One such example is chapter U that contains codes for non-surgical procedures such as radiologic examinations[40]. Diagnosis codes are unique to NPR in the sense that it is mandatory to report at least one diagnosis code for every single contact. In contrast, not all contacts have a relevant procedure code, which is obvious for contacts where not procedures were performed. However, it was not until year 2000 that it was required by the hospitals to report the performed procedures to the Danish health authorities[36]. Historically, the diagnosis codes have been used most extensively in research. They can either be assigned as a primary or non-primary diagnosis. Primary diagnosis codes refer to the diagnosis that best describes the reason for the contact and non-primary codes may compliment that description making them highly relevant in the context of multi-morbidity[36], [50]. 2.2.2 The Danish Registry for Causes of Death In Denmark, completion of death certificates has been mandatory by law since 1871. It was not until 1970 that deaths among citizens dying in Denmark were registered in individual electronic records in DAR. Until then the data from death certificates were archived on punched cards. Data in DAR originates from death certificates which can only by written by a physician. Today The Danish Registry for Causes of Death (DAR) is maintained by the Danish Health Authority. Death certificates that dates back before 1994 are archived in the Danish National Archives. Variables in DAR include manner (i.e. natural, accident, violence, suicide, and uncertain), time, place, and causes of 4The full SKS is available at https://medinfo.dk/sks/brows.php [Danish]. 11
CHAPTER 2. MATERIALS ON DIAGNOSTIC CODES death[54]. Underlying and contributory causes of death have been classified according to international classification systems since 1948 where they have been maintained by WHO. This means that causes of death are now being classified by ICD-10 codes. Before 1948 causes were registered using Danish and Scandinavian disease classifications. The registry was centrally validated until 2002 and since then, the registry has relied solely on information from the medical doctor who has verified the death[54]. There is some overlap between data in CRS and DAR as the time of death is a variable that is available in both CRS and DAR. The other variables related to death are unique to DAR[51], [54]. 2.2.3 The Danish National Prescription Registry The Danish National Prescription Registry (DNPR) contains data on all individual prescriptions dispensed at any Danish community pharmacy from 1994 and onwards. DNPR is a subregistry of Register of Medicinal Product Statistics, which is maintained by the Danish Medicines Agency. DNPR is administered by Statistics Denmark. The four main variable categories are variables related to the user (i.e. the citizen or patient), the prescriber (i.e. the treating physician), the drug and the pharmacy. Only redeemed prescriptions are entries in DNPR. Dispensed drugs are registered in accordance with the global ATC Classification System that is equivalent to chapter M of SKS. Overthe-counter (OTC) drugs are only entries in DNPR if they were dispensed as prescriptions[55]. 2.3 Deeper data characterize patients with greater precision A key element of this thesis is to explore the potential of the Danish healthcare data by analyzing it comprehensively. Deeper phenotypic data was therefore analyzed in addition to the nationwide registries presented in section 2.2. In broad terms, deep phenotypes are based on individual-level information that is more fine-grained than the information in for example healthcare registries[19], [56]. 2.3.1 The Eastern Danish Heart Registry The Eastern Danish Heart Registry (EDHR) is as subset of the Danish Heart Registry that was established in 2000. It contains data on all patients who have been subjected to CAG, PCI, coronary artery bypass grafting or heart valve surgery in Denmark. EDHR contains data from invasive cardiac procedures performed at a hospital in Eastern Denmark (equivalent to the Capital Region and Region Zealand) and patients younger than 14 at time of the procedure are excluded from the registry. Currently, EDHR contains data on about 100 000 unique patients many of whom have multiple entries in the registry. Each entry is described with up to 70 variables. The variables cover information on e.g., demographics, dispositions, prognostic factors, findings (e.g. coronary pathology), therapeutic consequences, and procedure-related complications[57]. 12
CHAPTER 2. MATERIALS ON DIAGNOSTIC CODES 2.3.2 Electronic Health Records from Eastern Denmark The EHR data used in the studies presented in this thesis originated from BigTempHealth (BTH), which denotes a research project based on EHRs from Eastern Denmark in the period years 2006 to 2016. The term EHR data is used to describe any routinely collected healthcare data, that are not necessarily archived in a health-registry, a biobank or a similar specialized institution[19]. Ideally, EHR data would contain continuous and for practically purposes complete monitoring of all patients during hospitalization[56], [58]. Biochemical tests, in-hospital prescriptions and free text originating from the raw clinical notes are the three components that make up BTH. The biochemical data were originally stored in the two databases Labka and BCC, covering The Capital Region and Region Zealand, respectively[59]. The biochemical tests were indexed using NPU codes or local systems and reference ranges were provided by the laboratory that performed the analyses. The prescription data available in BTH were originally archived in the three systems Electronic Patient Medication 1,3and OpusMedicine, covering The Capital Region and Region Zealand, respectively. Analogously to the prescription data in DNPR, the prescription data in BTH were indexed according to the ATC Classification. In contrast to the data in DNPR, the prescription data obtained from EHRs also contain information on the drug dosages of the individual dispensation. The clinical notes include unstructured data, where clinical variables such as blood pressure, height and weight can be extracted by text mining. Similarly, allergies and symptoms can in principle be identified in this part of the data. 2.3.3 Copenhagen Hospital Biobank CopenhagenHospitalBiobank(CHB)wasestablished in2009 tomakeuseof residual blood samples that were previously discarded. It was initiated at Rigshospitalet and later expanded to include samples from all public hospitals in the Capital Region of Denmark. CHB contains genetic data obtained from blood that was originally drawn for blood-type testing or red cell antibody screening[60]. The genetic analyses performed on these samples are currently being conducted in collaboration with deCODE genetics[61]. Since the biological material is leftover from routine tests, study participants should in principle be asked for informed consent prior to inclusion. However, the biobank has a dispensation from the general rule requiring consent. In effect, study participants are not asked for informed consent prior to inclusion but can opt out at any time. Samples frompatients younger than 18 years are excluded. Likewise, one patient can maximum be included once meaning that samples from patients who are already included in CHB are discarded. As of February 2020, genetic data were available on about 155 000 patients[60]. At the time of writing, this has now exceeded 300 000 patients. The analyses are performed only if studies are approved by the relevant authorities (cf. section 4). 13
CHAPTER 3. METHODS ON DIAGNOSTIC CODES redeemed over the study period, P12 indicate that prescription P1 was redeemed before prescription P2 and P02 indicate that only prescription P2 was redeemed in the study period (Figure 8). Significant prescription pairs were then defined by establishing a Poisson regression model for each prescription pair. The number of people who had redeemed a prescription for P2 was the dependent variable. The independent variables were age, sex, calendar year and whether a binary variable indicating whether the study participant had redeemed a prescription for P1 prior to redeeming P2. Like other regression models, a Poisson regression can be solved using maximum likelihood estimates and significance can be assessed using the Wald statistics[66], [76]. The interpretation of the RR was that it expresses the risk of redeeming a prescription for P2 after P1 relative to people who had not redeemed the prescription for P1. Longer prescription trajectories were pieced together from significant prescription pairs by application of the same principles as described in section 3.2.1. 3.3 Cluster analysis to define ischemic heart disease subgroups This section describes a strategy for studying multi-morbidity in IHD patients that is based on cluster analysis. In contrast to the disease trajectory approach described in section 3.2, the study of multi-morbidity based on cluster analysis only included multi-morbidities that were diagnosed before IHD onset. Section 3.3.1 describes the underlying concept for application of cluster analysis in the context of multi-morbidity used in this thesis. Next, section 3.3.2 describes the mathematical concepts that cluster analysis relies on. Section 3.3.3 describes how cluster analysis was applied to the entire spectrum of multi-morbidities present in an IHD cohort, that was identified in NPR and finally section 3.3 exemplifies how the physiological relevance of the clusters was assessed. 3.3.1 A model where multi-morbidity and text documents share properties The premise of this model of multi-morbidity was that the set of diagnosis codes assigned to a particular patient in a cohort can be considered a text document, where diagnosis codes correspond to the words in the document. Analogously, the cohort was considered a set of different text documents (i.e. text corpus), where the degree content overlap between the different documents differ. The idea is then to group patients based on the similarity in diagnosis codes, just like text documents can be sorted with respect to their theme based on the content of words. In this analogy, different manifestations of multi-morbidity among the patients in the cohort correspond to different text themes. This strategy for studying multi-morbidity originates from information retrieval and has was previously adopted in studying diabetes subgroups[77]. In broad terms, information retrieval is a discipline defined by the process of finding material with a specific piece of information from a larger set of information such as a data set. Classical information retrieval tasks refer to the task of finding specific documents from a larger collection which is usually stored on computers[78]. In the study of multi-morbidity that was based on these principles, the initial task was to identify patients (classically documents) in a cohort (classically a text corpus) that shared certain characteristics. The idea was that the resulting patient subgroups 20
CHAPTER 3. METHODS ON DIAGNOSTIC CODES could form the basis for a more precise segregation of highand low-risk patients. A certain risk group would then correspond to finding a specific piece of information in classical information retrieval tasks[78]. In classical information retrieval tasks a scaling factor is often introduced to reduce the effect of unspecific words. For example, the term "and" is an unspecific feature as it is relatively abundant and is not very appropriate for discriminating between documents. In the case of multi-morbidity in IHD patients, a diagnosis code for hypertension could correspond to "and" in being very frequent among IHD patients but not very discriminative, i.e. largely uninformative. The term frequency-inverse document frequency (tf-idf) is an example of a scaling factor that scales all features according to the number of occurrences (i.e., number of times it is assigned to a single patient) and the total number of occurrences in the entire data set. Thus, it dampens the effect of relatively uninformative terms[78]. Assuming that the principles from classification of text documents also apply to diagnosis codes in classification of patients, the tf-idf scaling factor was introduced to reduce the effect of relatively unspecific diagnosis codes. Mathematically, the tf-idf scaling factor is the product of the term frequency and the inverse document frequency and can be expressed by tf-idf =tft,d ·idft(3.2) where tft,d is the normalized frequency of a diagnosis code in the entire cohort and idftis the inverse of the ratio between the number of patients who are assigned that particular diagnosis code and the total number of patients in the cohort[77], [78]. Equations and further details for computation of tft,d and idftare available in equations A.2, A.3, and A.4 in appendix A. 3.3.2 A mathematical graph representation of ischemic heart disease patients A graph is a mathematical object that can be used to represent a class of entities[79]. It is defined by a triple consisting of a vertex set, an edge set and a relation that associates each edge with two vertices. It is not a requirement that the vertices are distinct as is the case in figure 9, where all nodes are connected by exactly two vertices. J A B C D E F G H I Figure 9: A graph with vertices (circles), edges (lines) and nnatural groups. 21
CHAPTER 3. METHODS ON DIAGNOSTIC CODES To model multi-morbidity in IHD, the Markov cluster (MCL) algorithm was applied to diagnosis codes from IHD patients, that conceptually corresponded to analyzing the words of different text documents in a text corpus. That is, in this model IHD patients corresponded to a class of entities. The overall aim was then to test for natural groups within this class, i.e, define multi-morbidity in a data-driven model by application of an unsupervised clustering algorithm to the entire set of diagnosis codes assigned prior to IHD onset. IHD patients were represented based on their diagnosis codes and using a graph where vertices corresponded to patients and edges represented patients belonging to the same subgroup. Thus, in the graph represented in figure 9, there is one cluster as all the nodes are related via a finite number of edges. Yet, it is also reasonable to argue that some of the nodes are more connected than others, illustrating that there may be natural groups within the graph. An example of such a group could be the nodes labelled I, J and H as it appears relatively detached from the other nodes (Figure 9). Many clustering algorithms have been developed, likely reflecting that the notion of a natural group is highly diverse; while also reflecting that cluster analysis is currently being applied in many different fields, including biology, communication, and computer science[80]. In sum and in broad terms, cluster analysis is concerned with finding ways to look at the data that enable a better understanding of mechanisms rather than finding the correct answer[81]. In the case of multi-morbiditity in IHD, cluster analysis is about finding ways to look at the data that, for example, allow to understand better why symptom burden and disease progression vary between patients. 3.3.3 A cluster algorithm minimizes a cost function The operational requirements for most clustering algorithms, including the MCL algorithm are (i) that the data are structured in a matrix where data points (observations) are related to data features (variables) and (ii) that there is a measure of similarity (qualitative features) or distance (quantitative features) between all data points[82]. When clustering algorithms are applied to healthcare data a matrix is usually having the individual patients (observations) as rows and different features (e.g. diagnosis codes) as columns [83], [84]. Before the matrix can be used as input to a clustering algorithm, it is typically optimized to better represent the data. In the study presented in manuscript 2, each patient was represented as a vector of the tf-idf scaled diagnosis codes. The second operational requirement for a clustering algorithm is that the similarity between observationsin the matrixisquantified either usinga similarityordistance function, suchas the cosine, Jaccard, and Hamming functions[80]. The cosine similarity was computed to quantify the similarity between the tf-idf scaled vectors. It is defined by an angle between two vectors which is equal to the dot product of the two vectors divided by the product of their lengths was used cos(θ) = xi•xj ||xi||·||xj|| (3.3) where θis the angle between the two vectors xiand xj. Finally, the MCL clustering algorithm was applied to the matrix that that represented the cosine similarities between the tf-idf scaled patient vectors. 22
CHAPTER 3. METHODS ON DIAGNOSTIC CODES Mathematically, the objective of the MCL algorithm is to partition a graph by minimizing a cost function, that is associated with the weight of the edges. The algorithm then works by identifying the graph partition that minimizes the cost function. Conceptually it works by stochastically simulating a sequence of random walks of length kin a graph, so similar regions in the graph will stay connected, while dissimilar regions will be partitioned in different clusters. The algorithm optimizes such that the length of kwill minimize the cost function. That is, the sequence of random walks will change the distribution of edges in a graph, such that vertices in dense regions (i.e. regions with many edges) will stay connected and vertices in sparse regions (i.e. regions with fewer edges) will not stay connected. The basic premise is that dense regions in the graph will contain relatively many paths of length k. Consequently, there will be relatively many walks in dense regions of the graph and fewer in sparse regions. Accordingly, following simulation and those regions will be considered similar[85]. Generally, additional data sources are necessary in the subsequent cluster characterization, as a comparison between subgroups obtained from cluster analysis comes with an extreme Type I error inflation rate, i.e. the risk of rejecting a true null-hypothesis is high because the seemingly difference is due to graph partition and thus a circular argument[86]. In effect, patient subgroups obtained by application of the MCL algorithm were characterized using biochemical and genetic data to evaluate the biological relevance of the patterns recognized by the algorithm. Obviously, patients who share disease patterns do not necessarily need to be genetically similar. That will depend on the disease in question, e.g. polygenic and multifactorial diseases. 3.3.4 Characterization of clusters using survival analysis The aim with cluster analysis in the setting of clinical research is often to identify distinct subgroups possibly characterized by different etiologies. In our cluster-based study of multi-morbidity in IHD, the physiological relevance of the clusters was assessed by different means. One of the strategies was to assess if there were differences in risk of disease progression between patients in the different clusters. Thus, clusters were used as covariates in a survival model where secondary ischaemic events and death from non-IHD causes were defined as two outcomes of interest. In broad terms, survival analysis refers to the type of analysis where the outcome of interest is time-to-event and is represented using survival curves, that describe the relation between number of survived subjects at different time-points (Figure 10). When there is more than one outcome of interest, the outcomes may be considered competing risks[87]. In the study presented in manuscript 2 the two outcomes were treated as competing risks as only one of the outcomes could occur in a single patient. In the study, clusters were also characterized using PGSs that were prepared from the raw genetic data in CHB by collaborators in the author group as well as the results from blood tests. Since I was not directly involved in these analyses, they are not described in this section. One of the strengths with survival analysis it that information from censored study participants is included in the model. Censoring refers to instances where the exact survival time of a study participant is unknown. Classical examples of censoring are that study participants are either diagnosed before the observation period began or are still alive when the observation period ended. Left-censoring applies to situations where the true survival time cannot be longer than the observed 23
CHAPTER 3. METHODS ON DIAGNOSTIC CODES Figure 10: Two survival curves and a risk table. Survival curves for patient subgroup A and B. X-axis: Time from start of follow-up to event, e.g. secondary ischemic events. Y-axis: Survival probability. Risk table: Number of patients in each subgroup at risk at time-points 0through 5. The number of patients at risk at a given timepoint is the number of patients where neither the outcome has occurred nor have they been censored at the timepoint that is being evaluated. Underlying data from study presented in manuscript 2. survival time. For example, there is left-censoring if the outcome of interest is time until disease onset and study participants are not screened continuously. In the case of NPR, most diagnoses are left-censored if time of diagnosis is used as a proxy for disease onset. Explicitly, a patient who has been assigned a diagnosis code for e.g. hypertension in NPR will most likely have developed hypertension before the point in time where it is registered in NPR. Conversely, right-censoring applies to cases where the true survival time is minimum as long as the observed survival time. Hence, if the outcome of interest is death, data will be right-censored unless all study subjects die during the follow-up period. Right-censoring is the most typical type of censoring in survival analysis[87]. The Cox proportional hazard (Cox PH) model is a semi-parameterized model and among the most widely used models for analysis of survival data[88]. This was also the model of choice in characterizing the patient subgroups obtained from application of the MCL algorithm. A central component of Cox PH is the baseline hazard function, also referred to as the conditional failure rate. Cox PH is semi-parameterized because the baseline hazard function is never specified[89]. Instead it is assumed that the baseline hazard is independent of the covariates also coined the proportional hazard assumption[87]. The Cox PH function is represented by the product of two terms, where 24
CHAPTER 3. METHODS ON DIAGNOSTIC CODES the first term is the baseline hazard function and the second term represents a linear combination of the explanatory variables. Mathematically it is given by the formula h(t, X) = h0(t)·e p P i=1 βiXi(3.4) where h0is the baseline hazard function and Xirepresents the explanatory variables. The baseline hazard function can be represented by a survival curve (Figure 10). The common interpretation is that the hazard function yields the risk which is quantified as the hazard in this analytical framework[87]. The hazard at a certain point in time (e.g. tn) is equal to the chance of survival at t0 given the study participant has survived the duration of time tn. Hazards are commonly expressed as hazards ratios (HRs) in comparing for instance a treatment and a placebo group[89]. 3.4 Data-driven predictions of mortality in ischemic heart disease This chapter explains the principles of ANNs, which are introduced in this thesis as a method for analyzing time-to-event data based on hundres of input features. Section 3.4.1 introduces the concept of ANNs and section 3.4.2 presents a model of explainability that responds to some general critique against the usage of ANNs and other machine learning methods in clinical research. 3.4.1 A conceptual introduction to artificial neural networks Originally, ANNs and the units they contain were developed to model information processing in the brain[90]. Accordingly, ANNs are built of neural units called artificial neurons. Neurons are typically described using the parameters weights and a bias term. In this context a bias is a threshold and determines how easy it is for the neuron to produce a certain output given its input[91]. Formally, the total, weighted input to a neuron in layer ixican be expressed as the weighted sum of incoming outputs from the previous layer xi=X j∈N−(i) wij ·yj+wi(3.5) where wij ·yjcorresponds to weighted sum of the inputs from the previous layer and wiis the bias term of a neuron in layer i. The weighted sum of inputs from the previous layer is then passed through a non-linear activation function, such as the sigmoid shaped logistic function, 1 1 + e−xi (3.6) to produce the output of a neuron in layter ithat serves as input to the next layer[92]. Multi-layer perceptrons are a class of ANNs that structurally are described by referring to three types of layers: the input layer that receives information to be processed, the output layer that delivers the computational result, and in between, one or more hidden layers (Figure 11). The hidden 25
CHAPTER 3. METHODS ON DIAGNOSTIC CODES and output layers are all comprised of a set of artificial neurons that process the data from the preceding layer. The model architecture of an ANN refers to e.g. the number of hidden layers in a ANN termed the depth of the ANN and also the relation between the neurons in the different layers[92]. Input Output Hidden layers Figure 11: Schematic representation of an ANN with two hidden layers. Artificial neurons are represented as boxes and colors indicate if the neurons belong to the input layer, output layer, or hidden layers. Input neurons are light blue (bottom). Output neurons are grey (top). ANN: Artificial neural network. Adapted from manuscript 3. A defining property of ANNs is that they learn from examples during development as weights and bias terms are tuned gradually in response to external stimuli (i.e. inputs). That is, ANNs of this kind learn from examples[92]. Before developing an ANN, a data set is typically split into a training set and a test set used for internal and external validation, respectively[92]. The former is used for internal validation to quantify and estimate the performance of the model. The test set is used to assess if the network works well on novel data or has been over-fitted on the training set (cf. section 3.4.2). One of the most widely applied learning algorithms is stochastic gradient descent (SGD), which is an algorithm that minimizes the cost function which is termed loss or objective function. In supervised learning, the cost function measures the difference between the actual output and the desired output[91], [93]. In principle, SGD solves the cost function by identifying the global minimum. In other words, the SGD works by minimizing the cost function. To avoid over-fitting, regularization of the data is often performed prior to model training. Regularization works by adding a penalty term to the cost function which ensures that it is not too complicated and hence prone to remember the data rather than learning some generalizable patterns in the data[94]. 26
CHAPTER 3. METHODS ON DIAGNOSTIC CODES Technically, SGD works by repeatedly applying the so-called update rule to weights and biases in a given network, which is exemplified for the weight wkhere wk→w0 k=wk−ηδC δwk (3.7) where C denotes the cost function and the change in the value of a weight wkto w0 kreflects testing if the cost function decreases when changing a weight. The update rule then tests if the cost function C decreases when changing a weight wkas indicated by wk→w0 k. The term ηrefers to the learning rate and is often explained as a "step-size". Step-size and ANN depth are yet other examples of socalled hyperparameters and ANN development is in part about optimizing the hyperparameters to reduce the cost function. 3.4.2 Ranking input features based on level of information A common critique against the usage of ANN in modern medicine is that they appear as black boxes in the sense that an ANN does not provide an explanation of why a certain input leads to a certain output, for example why the risk of death for a patient is estimated to be a certain value[95]. An obvious counterargument is that clinical decisions similarly appear as black boxes as they do not always generalize between cases. Another argument in favor of ANNs is that several methods for explaining the output of any machine learning model, including ANNs now exist[96]. This aspect of machine learning is called interpretability or explainability. An example of such method is the so-called SHapley Additive exPlanations value (SHAP value) named after the Nobel Prize-winning economist Lloyd Stowell Shapley[97]. The underlying idea of the SHAP value is borrowed from coorporative game theory. Here the Shapley value is a method that assigns payouts to players depending on their contribution to the total payout[98]. Analogously, the SHAP value can be used to measure the individual feature importance in an ANN, based on the informative value of that feature. The idea is that changing an informative feature will also change the output considerably; whereas the opposite is true for less informative features[99]. SHAP values refer to the results of differences between the output from the actual (baseline) model (i.e. an ANN), and the output of a new, explanatory model. The SHAP value for a given input feature measures how much that feature contributed to the model output in the context of its interactionwith theother features. So, the SHAPvaluefora given feature iscalculated by comparing the output of the actual model with the output of a model without that specific feature. The SHAP value is the difference between these two outputs. Thus, SHAP values are relative to the input data and can be used to interpret the model globally as well as interpretation of the individual predictions[99]. An important property of SHAP values is that they are additive, meaning that the importance of several input features can be pooled[100]. 27
CHAPTER 3. METHODS ON DIAGNOSTIC CODES Aside from this kind of explainability analysis, model performance quantified using calibration and discrimination are also relevant in evaluating a survival model based on an ANNs[101], [102]. Model calibration refers to the ability of the model to predict future outcomes compared to the observed outcomes, whereas model discrimination refers to a measure of the ability of the model to separate outcomes. That is, a model can be well-calibrated and at the same time have a poor discrimination if it predicts eight failures within a time period but the predicted subjects who fail do not correspond to the observed subjects who fail[103]. The classification performance, i.e. discrimination of the algorithm developed in manuscript 3, was evaluated using the time-dependent area under the receiver operating characteristic curve (tdAUC). The receiver operating characteristic curve describes the tradeoffs between benefits and cost, i.e. true positives and false positives, where the goal is to maximize the former and minimize the latter[101], [104]. 3.5 A probabilistic map of in-hospital drug dosage alterations Bayesian hierarchical logistic regression was applied to establish a model of polypharmacy, where estimated likelihoods of drug-dosage alterations were used as proxies for potential drug-drug interactions. Since the method represents a statistical paradigm that is fundamentally different from the frequentist paradigm centered on the null hypothesis, this chapter is dedicated to introducing the fundamentals of Bayesian statistics. Section 3.5.1 describes the components that comprise a Bayesian regression model and section 3.5.2 gives an overview of the application of such models in the context of EHRs. 3.5.1 Bayesian models can be decomposed into three distributions A defining attribute of Bayesian statistics is that a joint probability distribution is given to all observed andunobserved variables[105]. This contrasts to afrequentistparadigm, whereit is assumed that the effect of unobserved variables is constant. Consequently, Bayesian models are flexible and addition of data to the model may change the model parameters. Generally, Bayesian models are gaining attention due to an increasing amount of software that has been developed to perform the regression[106]. Conceptually, a Bayesian model quantifies a degree of (un-)certainty, which can be updated upon addition of new data, i.e. evidence[107]. Bayesian models can be decomposed into three parts that may be probability distributions[105]. These are a prior distribution, a likelihood distribution and a posterior distribution (Figure 12). Bayesian statistics is rooted in Bayes’ theorem from which a conditional probability can be derived without knowledge of the joint probability. Mathematically Bayes’ theorem is expressed by p(B|A) = p(A|B)p(A) p(B)(3.8) where p(B|A)expresses the probability of B, given A. The theorem states that this probability is equal to the product of the probability of A given B and the probability of A divided by the probability of B. In other words, Bayes’ theorem states that the probability of B given A is proportional to 28
CHAPTER 3. METHODS ON DIAGNOSTIC CODES Data Likelihood Priors Bayes’ theorem Posterior ∼ Figure 12: A schematic representation of a Bayesian logistic regression model. Upper row represents the three different distributions. Middle row represents a Bayesian logistic regression, that is the foundation for obtaining the joint distribution also coined the posterior distribution, which is represented in the bottom row. Adapted from [107]. the probability of A given B, i.e. p(B|A)∝p(A|B)p(A)(3.9) As indicated in figure 12, a Bayesian logistic regression model can be established by application of Bayes’ theorem to a data set, where the likelihood p(A|B)is the binomial (n, θ)and θis equal to the sum of the covariates. The prior reflects the level of information that is assumed to be in the data. It corresponds to p(A)in Bayes’ theorem and represents the prior beliefs about the relation between a set of variables and the outcome of interest. If the data is representative of the population that is being modelled, the prior will have a small variance. Conversely, the prior will have a large variance in cases where for example the data set is relatively small and thus not likely to be strongly related to the outcome of interest. Generally, priors are characterized based on their level of information ranging from informative (small variance) to diffuse or weakly informative (large variance). Importantly, the priors contain model subjectivity and are selected by the investigator although different priors can be compared[105], [108]. In a Bayesian logistic regression model, the goal is to model a posterior distribution that is as close to the true posterior as possible. The optimization can be performed using different algorithms, e.g., Markov Chain Monte Carlo sampling (MCMC)[109]. The algorithm estimates the posterior 29
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES Table 2: Summary statistics of selected disease trajectories. Population ICD-10 codes Mean age in years Counts Adj. P RR D1 D2 D3 D1 D2 D3 All E78 I21 60.8 62.9 76 334 <0.001 3.69 E11 I20 61.0 63.0 61 333 <0.001 1.56 E11 I50 65.6 70.4 42 023 <0.001 2.82 I20 I50 68.0 71.4 84 955 <0.001 1.94 E78 I20 I21 61.1 61.6 62.5 48 444 E11 I20 I50 64.2 65.6 69.3 23 431 I21 I21 I48 73.3 76.3 11 871 <0.001 1.09 I20, I21, I25 I48 I21 69.4 71.2 26 273 <0.001 1.09 From manuscript 1. 5.2 A strategy to define more homogeneous ischemic heart disease patient subgroups Traditionally, multi-morbidity has been quantified with reference to a few selected diseases and summarized in an index such as CCI. Some even argue that there does not exist any replicable profiles of multi-morbidity[27], [123]. To accommodate these limitations in future multi-morbidity classification systems, it has been argued that it is necessary to study it comprehensively, rather than purely assessing the presence of individual diseases in a binary fashion[9]. Studies that classify ischemic heart disease patients as well as heart failure patients based on application of cluster analyses to EHRs data have already been conducted[77], [124], [125]. Yet, cluster analysis in this domain is still relatively unexplored where the analyses are often limited to few diagnoses and thus only analyze a fraction of the available healthcare data volumes. Thus, there is a need to explore the potential of EHRs in this domain. To this end, the study presented in manuscript 2 addresses multi-morbidity in IHD based on a flexible clustering analysis of multi-morbidity by application of the MCL algorithm to the diagnosis codes of 74 249 IHD patients identified in NPR (Figure 15). The physiological relevance of the clusters (i.e. patient subgroups) obtained by application of the MCL algorithm was assessed based on a an enrichment analysis with respect to the diagnosis codes, a comparison of the subgroups that included estimates of risk for disease progression; and characterization derived from blood test results and PGS obtained from EHRs and CHB, respectively.6 Application of the MCL algorithm to the entire set of diagnosis codes assigned prior to IHD onset resulted in identification of highand low-risk patient subgroups. Patient subgroups were classified as highor low-risk based on survival analysis. For example, among three of the male subgroups 6For an in-depth description of the enrichment analyses as well as analyses of the biochemical and genetic data please refer to the full-length manuscript available in chapter 8. 36
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES Figure 15: Conceptual figure of the patient representation in the cluster analysis. The entire cohort is illustrated to the left and the vectors for five patients are illustrated on the right. A sample of diagnosis codes are listed horizontally in the upper right. Black circles indicate that diagnoses were assigned before IHD onset (included in cluster analysis). Hollow circles indicate diagnoses which were assigned after IHD onset (not included in cluster analysis). t0: IHD onset. IHD: Ischemic heart disease. Adapted from manuscript 2. (total of 31), mean age at IHD onset ranged between 54.9and 62.4years. These three patient subgroups were also characterized by different risks of disease progression, as the patients in one subgroup associated with reduced risk of progression (i.e. primary outcome) when compared to the the risk of the patients who were not in this subgroup. In contrast, the two other subgroups associated with increased risk of progression (Table 3). The results of the enrichment analysis showed that one of the subgroups characterized by increased risk of disease progression was enriched for diabetes whereas the other was enriched for more unspecific diagnoses such as diastasis of the muscle. Interestingly, the subgroup characterized by reduced risk of disease progression was enriched for acute MI. The distributions of PGS for total cholesterol (TC) and LDL cholestereol levels were significantly different in this subgroup with positive effect directions (Table 3). That is, patients in this subgroup generally had higher PGS for TC and LDL levels than the patients who were not in this subgroup. In sum, the study presented in manuscript 2 exemplified that application of the MCL algorithm to the entire multi-morbidity spectrum of IHD patients can be used to discriminate between highand low-risk IHD patients. Further, the study illustrates how, potentially, more homogeneous patient subgroups can be used to identify subgroups with genetic associations where PGS have positive or negative effect sizes (Table 3). In comparison to other cluster analyses of cardiovascular phenotypes, we developed a flexible approach that analyze input data in a data-driven manner based on the analogy to text documents[125]. Further, the characterization of the clusters involved addition of data sources and thus presents a pragmatic solution to reduce the impact of an inflated type I error rate[86], [124]. Potentially the methods presented in this study can be further optimized and outperform traditional multi-morbidity scores. In its current form, the discrimination of the subgroups is equally as good as CCI overall and in some subgroups, it is better. The results are presented more elaborately in the full-length manuscript, which is available in chapter 8. 37
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES Table 3: Summary statistics and characteristics for selected male clusters. Cluster Mean age HR, primary outcome Main enrichment PGS Years at index HR 95%CI Adj. P (effect direction) 4.M01 58.5 0.93 0.87 to 0.99 0.030 Acute myocardial TC(+), I25(-) infarction LDL(+) 4.M02 62.4 1.51 1.40 to 1.62 <0.001 Diabetes E11(+) 4.M25 54.9 1.38 1.15 to 1.65 0.030 Diastasis of muscle From manuscript 2. 5.3 Mortality in ischemic heart disease from a time-to-event machine learning algorithm Today, IHD patient prognostication is being carried out with risk prediction tools that largely have remained unchanged for decades[83]. While the volumes and variety of data are increasing in complexity, risk scoring algorithms are commonly extrapolated from classical regressions that model linear relationships between a limited set of well-characterized predictive features[11]. Machine learning methods, such as ANNs comprise a potential methodological solution to the fact that the increasingly complex healthcare data require analytical capabilities that often exceed that of the human mind[95]. There are already promising results of patient risk prediction models developed using machine learning methods within the cardiovascular domain[126], [127]. However, there are still no examples of risk predictions models in IHD where input features are being included based on availability rather than assumptions regarding their predictive value. Therefore, the study presented in manuscript 3 introduces PMHnet-alpha, an ANN-based survival model developed by application of a generic approach that was originally described by Gensheimer et al[128]. Based on a derivation cohort identified in EDHR, PMHnet-alpha was developedtopredict all-cause mortality in IHD patients from 596 inputfeatures (Figure16). The premise of the study was that the summarized predictive value of many features with varying predictive value could comprise a prediction model with higher performance than the summarized predictive value of fewer, highly predictive features. After model development, PMHnet-alpha achieved excellent performance and outperformed the GRACE 2.0 risk model in validation with a tdAUC of 0.869 at the one-year evaluation time point (Table 4). Using SHAP values, we were able to quantify the predictive importance of all 595 input features obtained from a combination of the different data types (Figure 16). The SHAP values enabled us to discriminate between the importance of different input features at a global and at a specific level. In a model based solely on diagnosis codes (ICD-10 codes), the diagnoses with the largest predictive values globally were diabetes (E11), acute myocardial infarction (I21), atrial fibrillation (I48), heart failure (I50) and chronic obstructive pulmonary disease (J44). All of them reduced the chances of survival. In contrast, the presence of angina pectoris (I20), disorders of lipoprotein metabolism and other lipidaemias (E78), open wound of wrist or hand (S61) and routine examination of specific systems (Z01) increased the chances of survival. In the model that was trained on all 595 input features, age was, not surprisingly, the most predictive variable. Interest38
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES Figure 16: Overview of the development of PMHnet-alpha. A: Derivation cohort where the input features of four patients are illustrated as indicated by the four dashed lines. B: The different data types that were used as input features. C: A schematic representation of the PMHnet-alpha architecture. D: PMHnet-alpha output here illustrated by survival curves for four individual patients. From manuscript 3. ingly, only troponins were among the most predictive features in evaluation of the performance at six months after index, whereas other biochemical tests taken before IHD onset were only among the most predictive features at the one-year and three-years evaluation timepoint. PMHnet-alpha is currently being validated at partner hospitals in Norway and Iceland, respectively. Table 4: Validation and benchmarking of PMHnet-alpha. Outcome Without imputation With imputation PMHnet-alpha GRACE 2.0 PMHnet-alpha GRACE 2.0 Six-months all-cause mortality 0.876 0.770 0.883 0.770 One-year all-cause mortality 0.869 0.782 0.872 0.756 Three-years all-cause mortality 0.843 0.738 0.831 0.727 From manuscript 3. Altogether, PMHnet-alpha exemplifies how data-driven studies can improve the performance of risk prediction both by analyzing non-linearly interacting features and also by including hundreds of features. Moreover, the study presents a solution to the common critique against machine 39
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES learning methods as PMHnet-alpha is an explainable model. For example, the SHAP analysis provided a framework for understanding why PMHnet-alpha based on diagnosis codes only had a similar performance to that of the GRACE 2.0 risk score. In contrast to other machine learning models within the cardiovascular domain, PMHnet-alpha was developed with essentially no a prior assumptions regarding potential predictive value[126]. We argue that this is an advantage in favour of PMHnet-alpha as the attribute makes it broadly applicable and presumably adaptable in to many different healthcare systems where data types and availability vary. This argument is presented more thoroughly in the full-length manuscript available in chapter 9, which is currently being prepared for submission. 5.4 Evidence of drug-drug interactions in electronic health records Multi-morbid patients, including IHD patients, are commonly subjected to polypharmacy which is often rooted in the simultaneous prescriptions of life-long treatment regimens[31]. Polypharmacy puts patients at risk of adverse drug reactions and reduces the likelihood of compliance. Traditionally, it has primarily been quantified with reference to the number of drugs only and drugs are rarely prescribed to patients who are diagnosed with only one chronic disease[30], [32]. In effect, evidence regarding the true therapeutic effect of drugs in multi-morbid patients is lacking as they are typically excluded from RCTs[9], [30]. This is particularly important in the context of IHD, where there are a wealth of different options for drugs used in primary and secondary treatment[129]. Figure 17: Conceptual figure illustrating concomitant treatment episodes during a hypothetical hospital admission. Three different drugs are depicted. Drug A (blue) is prescribed in two different regimens. Brug B (pink) is prescribed in three different regimens. Drug C (green) is prescribed one time. Three different modifications were considered: Drug addition, dosage adjustment, and drug discontinuation. ADD: Average daily dose. D/C: Discontinuation. From manuscript 4. 40
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES The objective of the study presented in manuscript 4 was to analyze polypharmacy with reference to drug dosage alterations based on in-hospital prescriptions from 1 069 873 hospitalized patients in the period years 2008 to 2016. Drug dosage alterations were analyzed systematically by identifying concomitant treatment episodes, defined by a point in time where a patient was prescribed two different drugs. Concomitant treatment episodes with a significant likelihood of dosage adjustments were defined as co-medication pairs comprised of an index-drug and a co-drug (Figure 17). Co-medication pairs were characterized specifically with reference to selected examples and also collectively by cross-referencing the results with data from 15 publicly available drug-drug interaction databases[130]. A total of 920 distinct drugs administered over a period of eight years gave rise to identification of 77 494 concomitant treatment episodes. In the data, there were a total of 3993 co-medication pairs. Index drugs were dominated by drugs belonging to the anatomical classes nervous system (ATC 1st level: N), cardiovascular system (ATC 1st C), and alimentary tract and metabolism system (ATC 1st A). Drugs classified as antithrombotic agents (ATC 2nd level: B01) and ACE inhibitors (ATC 2nd level: C09) were among the drugs that most frequently appeared as index drugs. They were represented in 316 (7.9%) and 155 (3.8%) of the co-medication pairs, respectively (Table 5). As expected, acetylsalicylic acid and clopidogrel comprised an example of a co-medication pair. Thus, when acetylsalicylic acid was co-administered with clopidogrel a dosage change in acetylsalicylic acid was 2.24 times more likely than in cases where acetylsalicyclic acid was administered as monotherapy. Similarly, clopidogrel was 1.52 times more likely to be dosages adjusted when coadministered with acetylsalicylic acid. Moreover, atorvastatin had an odds ratio of 1.31 for a dosage alteration when co-administered with clopidogrel. Thus the co-medication pairs represented expected as well as more surprising examples of dosage adjustments (Table 5). Table 5: Examples of co-medication pairs from highly represented therapeutic subgroups. Overview Selected co-medication pairs ATC Index Co-drug ATC is anatomical level of index drug. 2nd N N Index drug Co-drug Odds ratio 95% HDI A02 105 176 Pantoprazole Terlipressin 2.26 1.88 to 2.79 B01 316 314 Acetylsalicylic acid Clopidogrel 2.24 2.03 to 2.49 Clopidogrel Acetylsalicylic acid 1.52 1.39 to 1.67 Enoxaparin Liraglutide 1.78 1.23 to 2.61 C07 116 137 Metoprolol Rifampicin 1.45 1.09 to 1.94 C09 155 150 Losartan Quetiapine 1.32 1.10 to 1.57 C10 111 129 Atorvastain Clopidogrel 1.31 1.18 to 1.44 N07 56 52 Disulfiram Pantoprazole 1.20 1.05 to 1.37 From manuscript 4. Of the 3993 co-medication pairs, there were 3297 pairs that were already entries in a drug-drug interaction database. We argue that the remaining 696 drugs represent potential drug-drug interactions, given that the dosage adjustments are not rooted in clinical routines or guideline recommendations with other reationals than pure drug-drug interactions. Overall, the study presented in manuscript 4 exemplifies the potential importance of data-driven studies in this domain. The study adds to the literature by (i) presenting a model of actual in-hospital treatment patterns in 41
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES a population that is representative of patients hospitalized in Denmark and (ii) by capturing rarer therapeutic drug combinations that are not necessarily captured in smaller studies of polypharmacy. In this respect, cardiovascular drugs generally served as good positive controls during model interpretation both with respect to the quantity and quality, corresponding to the number content of the co-medications pairs. The results of the study are presented and discussed in greater detail in the full-length manuscript which is available in chapter 10. 5.5 Establishing a map of longitudinal, nation-wide prescription patterns Most of the drugs that are prescribed as primary or secondary prevention in treatment of IHD and relatedmulti-morbidityarebeing prescribed life-long[30]. Although the numberofpatientsneeded to treat to expect to prevent one adverse event are high; the process of deprescription is not standard in clinical care[131], [132]. Yet, there is only limited knowledge of the long-term consequences of polypharmacy, which is becoming even more prevalent in the ageing population[75]. Albeit these facts, very few studies have analyzed the long-term consequences of drugs used to treat multimorbidity, including IHD[30], [32]. Based on prescription trajectories established from 7 255 919 people in DNPR, the study presented in manuscript 5 mapped longitudinal prescription patterns in an entire Danish population over the years 1995 to 2019. This study transformed the individual level prescriptions in DNPR into 38 607 prescription trajectories of two distinct drugs that were most likely to be prescribed in a specific order. For example, the RR of redeeming digitalis glycoside after a beta blocking agent was 3.55 (95% CI: 2.30 to 5.47, P<0.001) compared to patients with no prior redemption of a beta-blocking agent. This observation is consistent with intensification of arrhythmia treatment and thus served as a positive control. We also found that redemption of direct factor Xa inhibitors associated with a reduced risk of subsequent redemption of other anti-dementia drugs (RR: 0.41,95% CI: 0.37 to 0.46, P <0.001). This association between anti-coagulative therapy exemplifies that data-driven studies can assist in identifying associations that are unlikely to be identified in purely hypothesis-driven studies. Potentially this specific prescription trajectory even points to a role of antithrombotic agents in primary prevention of stroke. Prescription trajectories that contained two prescriptions were also pieced together to form prescription trajectories of up to seven different prescriptions. We used this strategy to map the order of anti-hypertensive treatments for patients being prescribed up to seven different drugs within this drug group (i.e. anti-hypertension drugs). For example, only in very few cases did people who were prescribed an aldosterone antagonists (C03DA) redeem a prescription on another anti-hypertensive drug afterwards. And this trend seemed to be irrespective of the previous redeemed prescriptions as there were few prescription pairs on the right side of boxes labeled C03DA (Figure 18). Taken together, the prescription trajectories modelled polypharmacy in a novel manner and supported analyses of up to six different anti-hypertensive drugs prescribed to the same person over the course of 25 years (Figure 18). The results indicate that there is a considerable room for improvement when it comes to decisions regarding treatment regimens. It is of course important to 42
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES Figure 18: Alluvial plot illustrating the complexity of anti-hypertensive treatment. Each vertical bar represents an ATC level 4code containing anti-hypertensive drugs. Grey ribbons that connect the vertical bars represent a prescription trajectory comprised of two prescription, where the RR risk of redeeming the drug on the left first is higher than redeemed the drug on the right first. The height of the vertical bars corresponds to the number of people who follow all the prescription trajectories going from one vertical bar to another (left to right). ATC: Anatomical Therapeutic and Chemical Classification. RR: relative risk. From manuscript 5. interpret the results with caution. For example aldosterone antagonists are often considered an "add-on" drug in treatment of hypertension[133]. Also, any disease progression that could have been the rationale for changing treatment regimens were not included in the analysis. We argue that, given further model refinement, this concept of prescription trajectories can assist clinical decision making by identification of people where it unlikely that they will respond sufficiently to first-line therapy. The full-length manuscript is presented in chapter 11. 5.6 Internet applications display perspectives with data-driven research Most existing software developed to study healthcare data do either require users to upload their own data or limit users to query within a specific domain[134], [135]. It is paramount for datadriven research that the results of the analyses are being evaluated continuously as their usefulness often depend on proper examples [63]. The results described in sections above are presented with reference to relatively few examples that assist in interpretation of the model leaving out, potentially, 43
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES relevant information. To this end, we have developed The Danish Disease Trajectory Browser (DTB) which is an example of a software tool where users can decide which example to analyze. The study presented in manuscript 6 introduces DTB which is a web application that makes the disease trajectory program available to anybody, including people without coding experience, given internet access, and at the same time establishing a platform for model critique and suggestions. The browser was developed as a tool that enables users to characterize a phenotype of interest using disease trajectories. A set of functionalities were added to the browser, including search filters, that for example enable users to search for disease trajectories computed from either males or females. As the browser only contains summary statistics it does not provide any access to personal data. Figure 19: Screenshot from the DTB home page. The home page has three central elements: a search bar in the middle, an upper horizon ribbon and a vertical side panel to the right. In the search bar users can type in an ICD-10 code of interest, e.g. I25. Users can decide if the results of the query should be represented as a disease trajectory network or linear disease trajectories. The upper horizontal bar has a variety of different buttons that can be used to navigate the interactive phase that will replace the area with the search bar, when a query has been placed. The vertical side panel has the different search filters and options regarding layout and downloads. DTB: Danish Disease Trajectory Browser From manuscript 6. DTB is made available at http://dtb.cpr.ku.dk/ and lets users characterize a phenotype by providing access to nearly 25 years of data from more than seven million people. Users can query 44
CHAPTER 5. RESULTS ON DIAGNOSTIC CODES DTB by specifying a case population in the form of level 3ICD-10 codes, e.g. I25 for chronic IHD (Figure 19). The main result of the study was that the results of the disease trajectory program became publicly available with an interactive interface. Potentially, this may lead to future developments of the program that are based on input from a broad range of research communities. DTB has a user-friendly interface that was developed to support a graphical representation of temporal associations between diagnoses in NPR data. This is becoming more important than ever as large healthcare data are becoming increasingly difficult to manage and analyse because data volumes and variety are growing faster than ever. Search results are available for download as graphical representations or data tables with summary statistics. Moreover, a range of search filters can by applied to customize the query. Filters include number of diagnoses per trajectory, range for number of patients that should follow a trajectory for it to be included in the search, range for RRs and sex (males only, females only or both). The study exemplifies that the digital development and technologies provide new opportunities of sharing research results. This will obviously also make them more prone to external critique. We argue that in the long run, both researchers, clinicians, and patients will benefit from this level of transparency which hand-on usage and queries facilitates. To this end, DTB and other similar projects are only the beginning. The browser functionalities and interpretation of the disease trajectories are described in greater detail in the full-length manuscript, which is presented in chapter 12. 45
REFERENCES ON DIAGNOSTIC CODES [23] K. Nicholson, T. T. Makovski, L. E. Griffith, P. Raina, S. Stranges, and M. van den Akker, “Multimorbidity and comorbidity revisited: Refining the concepts for international health research,”Journal of clinical epidemiology,vol.105, pp. 142–146,2019.doi:10.1016/j.jclinepi. 2018.09.008. [24] S. P. Bell and A. A. Saraf, “Epidemiology of multimorbidity in older adults with cardiovascular disease,” Clinics in geriatric medicine, vol. 32, no. 2, p. 215, 2016. doi:10.1016/j.jacc. 2018.03.022. [25] A. R. Feinstein, “The pre-therapeutic classification of co-morbidity in chronic disease,” Journal of chronic diseases, vol. 23, no. 7, pp. 455–468, 1970. doi:10.1016/0021-9681(70)90054-8. [26] V. De Groot,H. Beckerman,G.J.Lankhorst, andL.M.Bouter, “Howtomeasure comorbidity: A critical review of available methods,” Journal of clinical epidemiology, vol. 56, no. 3, pp. 221– 229, 2003. doi:10.1016/s0895-4356(02)00585-1. [27] M. E. Charlson, P. Pompei, K. L. Ales, and C. R. MacKenzie, “A new method of classifying prognostic comorbidity in longitudinal studies: Development and validation,” Journal of chronic diseases, vol. 40, no. 5, pp. 373–383, 1987. doi:10.1016/0021-9681(87)90171-8. [28] F. Zhang, A. Bharadwaj, M. O. Mohamed, J. Ensor, G. Peat, and M. A. Mamas, “Impact of charlson co-morbidity index score on management and outcomes after acute coronary syndrome,” The American Journal of Cardiology, vol. 130, pp. 15–23, 2020. doi:10.1016/j. amjcard.2020.06.022. [29] R. Gupta and D. A. Wood, “Primary prevention of ischaemic heart disease: Populations, individuals, and health professionals,” The Lancet, vol. 394, no. 10199, pp. 685–696, 2019. doi:10.1016/S0140-6736(19)31893-8. [30] X. Rossello, S. J. Pocock, and D. G. Julian, “Long-term use of cardiovascular drugs: Challenges for research and for patient care,” Journal of the American College of Cardiology, vol. 66, no. 11, pp. 1273–1285, 2015. doi:10.1016/j.jacc.2015.07.018. [31] M. W. Rich, D. A. Chyun, A. H. Skolnick, K. P. Alexander, D. E. Forman, D. W. Kitzman, M. S. Maurer, J. B. McClurken, B. M. Resnick, W. K. Shen, et al., “Knowledge gaps in cardiovascular care of the older adult population: A scientific statement from the american heart association, american college of cardiology, and american geriatrics society,” Circulation, vol. 133, no. 21, pp. 2103–2122, 2016. doi:10.1161/CIR.0000000000000380. [32] S. M. Grundy, “Drug therapy of the metabolic syndrome: Minimizing the emerging crisis in polypharmacy,” Nature reviews Drug discovery, vol. 5, no. 4, pp. 295–309, 2006. doi:10.1038/ nrd2005. [33] D. C. Angus, B. M. Alexander, S. Berry, M. Buxton, R. Lewis, M. Paoloni, S. A. Webb, S. Arnold, A. Barker, D. A. Berry, et al., “Adaptive platform trials: Definition, design, conduct and reporting considerations.,” Nature Reviews Drug Discovery, vol. 18, no. 10, pp. 797–808, 2019. doi:10.1038/s41573-019-0034-3. [34] T. Boerma, J. Harrison, R. Jakob, C. Mathers, A. Schmider, and S. Weber, “Revising the icd: Explaining the who approach,” The Lancet, vol. 388, no. 10059, pp. 2476–2477, 2016. doi: 10.1016/S0140-6736(16)31851-7. 52
REFERENCES ON DIAGNOSTIC CODES [35] T. Siggaard, R. Reguant, I. F. Jørgensen, A. D. Haue, M. Lademann, A. Aguayo-Orozco, J. X. Hjaltelin, A. B. Jensen, K. Banasik, and S. Brunak, “Disease trajectory browser for exploring temporal, population-wide disease progression patterns in 7.2 million danish patients,” Nature communications, vol. 11, no. 1, pp. 1–10, 2020. doi:10.1038/s41467-020-18682-4. [36] M. Schmidt, S. A. J. Schmidt, J. L. Sandegaard, V. Ehrenstein, L. Pedersen, and H. T. Sørensen, “The danish national patient registry: A review of content, data quality, and research potential,” Clinical epidemiology, vol. 7, p. 449, 2015. doi:10.2147/CLEP.S91125. [37] W.H.Organizationet al.,“Whointernational workinggroupfordrug statisticsmethodology, who collaborating centre for drug statistics methodology, who collaborating centre for drug utilization research and clinical pharmacological services: Introduction to drug utilization research,” Introduction to drug utilization research. Geneva: World Health Organization, 2003. [38] Anatomical therapeutic chemical (atc) classification,https://www.who.int/tools/atc-dddtoolkit/atc-classification, (Accessed on 05/10/2021). [39] U. M. Petersen, R. Dybkær, and H. Olesen, “Properties and units in the clinical laboratory sciences. part xxiii. the npu terminology, principles, and implementation: A user’s guide (iupac technical report),” Pure and Applied Chemistry, vol. 84, no. 1, pp. 137–165, 2011. doi: 10.1515/cclm.2011.735. [40] N. NOMESCO, Nomesco classification of surgical procedures (ncsp), version 1.14,http://norden. diva-portal.org/smash/get/diva2:968721/FULLTEXT01.pdf, (Accessed on 05/17/2021), 2011. [41] K. Adelborg, J. Sundbøll, T. Munch, T. Frøslev, H. T. Sørensen, H. E. Bøtker, and M. Schmidt, “Positive predictive value of cardiac examination, procedure and surgery codes in the danish national patient registry: A population-based validation study,” BMJ open, vol. 6, no. 12, 2016. doi:10.1136/bmjopen-2016-012817. [42] J. C. Hong and A. J. Butte, “Assessing Clinical Outcomes in a Data-Rich World—A Reality Check on Real-World Data,” en, JAMA Network Open, vol. 4, no. 7, e2117826, Jul. 2021, issn: 2574-3805. doi:10.1001/jamanetworkopen.2021.17826. [Online]. Available: https: / / jamanetwork . com / journals / jamanetworkopen / fullarticle / 2782343 (visited on 07/30/2021). [43] E. S. Berner, G. D. Webster, A. A. Shugerman, J. R. Jackson, J. Algina, A. L. Baker, E. V. Ball, C. G. Cobbs, V. W. Dennis, E. P. Frenkel, et al., “Performance of four computer-based diagnostic systems,” New England Journal of Medicine, vol. 330, no. 25, pp. 1792–1796, 1994. doi:10.1056/NEJM199406233302506. [44] K.-H. Yu, A. L. Beam, and I. S. Kohane, “Artificial intelligence in healthcare,” Nature biomedical engineering, vol. 2, no. 10, pp. 719–731, 2018. doi:10.1038/s41551-018-0305-z. [45] P. R. Njølstad, O. A. Andreassen, S. Brunak, A. D. Børglum, J. Dillner, T. Esko, P. W. Franks, N. Freimer, L. Groop, H. Heimer, et al., “Roadmap for a precision-medicine initiative in the nordic region,” Nature genetics, vol. 51, no. 6, pp. 924–930, 2019. doi:10.1038/s41588-0190391-1. [46] S. W. Tu, O. Bodenreider, C. Çelik, C. G. Chute, S. Heard, R. Jakob, G. Jiang, S. Kim, E. Miller, M. M. Musen, et al., “A content model for the icd-11 revision,” Technical Report BMIR, pp. 2010–1405, 2010. 53
REFERENCES ON DIAGNOSTIC CODES [47] R. H. de Jong, “Computer crutches,” JAMA, vol. 240, no. 11, pp. 1176–1176, 1978. doi:10. 1001/jama.240.11.1176b. [48] M. A. Haendel, C. G. Chute, and P. N. Robinson, “Classification, ontology, and precision medicine,” New England Journal of Medicine, vol. 379, no. 15, pp. 1452–1462, 2018. doi:10. 1056/NEJMra1615014. [49] S. Köhler, N. A. Vasilevsky, M. Engelstad, E. Foster, J. McMurry, S. Aymé, G. Baynam, S. M. Bello, C. F. Boerkoel, K. M. Boycott, et al., “The human phenotype ontology in 2017,” Nucleic acids research, vol. 45, no. D1, pp. D865–D876, 2017. doi:10.1093/nar/gkw1039. [50] M. Schmidt, S. A. J. Schmidt, K. Adelborg, J. Sundbøll, K. Laugesen, V. Ehrenstein, and H. T. Sørensen, “The danish health care system and epidemiological research: From health care contacts to database records,” Clinical epidemiology, vol. 11, p. 563, 2019. doi:10.2147/CLEP. S179083. [51] M. Schmidt, L. Pedersen, and H. T. Sørensen, “The danish civil registration system as a tool in epidemiology,” European journal of epidemiology, vol. 29, no. 8, pp. 541–549, 2014. doi:10. 1007/s10654-014-9930-3. [52] Patientregistrering - Fællesindhold - Sundhedsdatastyrelsen,https://sundhedsdatastyrelsen. dk/da/rammer-og-retningslinjer/om-patientregistrering/patientregistreringfeallesindhod, (Accessed on 31/05/2021), Apr. 2021. [53] E. Lynge, J. L. Sandegaard, and M. Rebolj, “The danish national patient register,” Scandinavian journal of public health,vol.39,no. 7_suppl, pp.30–33,2011.doi:10.1177/1403494811401482. [54] K. Helweg-Larsen, “The danish register of causes of death,” Scandinavian journal of public health, vol. 39, no. 7_suppl, pp. 26–29, 2011. doi:10.1177/1403494811399958. [55] H. Wallach Kildemoes, H. Toft Sørensen, and J. Hallas, “The danish national prescription registry,” Scandinavian journal of public health, vol. 39, no. 7_suppl, pp. 38–41, 2011. doi:10. 1177/1403494810394717. [56] C. M. Delude, “Deep phenotyping: The details of disease,” Nature, vol. 527, no. 7576, S14– S15, 2015. doi:10.1038/527S14a. [57] C. Özcan, K. Juel, J. F. Lassen, L. M. von Kappelgaard, P. E. Mortensen, and G. Gislason, “The danish heart registry,” Clinical epidemiology, vol. 8, p. 503, 2016. doi:10.2147/CLEP.S99475. [58] P. B. Jensen, L. J. Jensen, and S. Brunak, “Mining electronic health records: Towards better research applications and clinical care,” Nature Reviews Genetics, vol. 13, no. 6, pp. 395–405, 2012. doi:10.1038/nrg3208. [59] A. F. Grann, R. Erichsen, A. G. Nielsen, T. Frøslev, and R. W. Thomsen, “Existing data sources forclinicalepidemiology:Theclinicallaboratoryinformationsystem(labka) researchdatabase at aarhus university, denmark,” Clinical epidemiology, vol. 3, p. 133, 2011. doi:10.2147/CLEP. S17901. [60] E. Sørensen, L. Christiansen, B. Wilkowski, M. H. Larsen, K. S. Burgdorf, L. W. Thørner, J. Nissen, O. B. Pedersen, K. Banasik, S. Brunak, et al., “Data resource profile: The copenhagen hospital biobank (chb),” International Journal of Epidemiology, 2020. doi:10 . 1093 / ije/dyaa157. 54
REFERENCES ON DIAGNOSTIC CODES [61] Decode genetics | a global leader in human genetics,https://www.decode.com/, (Accessed on 08/15/2021). [62] K. Popper, The logic of scientific discovery. Routledge, 2005. [63] S. Carrol and D. Goodstein, “Defining the scientific method,” Nat Methods, vol. 6, p. 237, 2009. doi:10.1038/nmeth0409-237. [64] D. A. Grimes and K. F. Schulz, “An overview of clinical research: The lay of the land,” The lancet, vol. 359, no. 9300, pp. 57–61, 2002. doi:10.1016/S0140-6736(02)07283-5. [65] A. C. Fanaroff, R. M. Califf, R. A. Harrington, C. B. Granger, J. J. McMurray, M. R. Patel, D. L. Bhatt, S. Windecker, A. F. Hernandez, C. M. Gibson, et al., “Randomized trials versus commonsense and clinicalobservation:Jaccreviewtopicof the week,”Journal of the American College of Cardiology, vol. 76, no. 5, pp. 580–589, 2020. doi:10.1016/j.jacc.2020.05.069. [66] R. Etzioni, Statistics for Health Data Science: An Organic Approach. Springer Nature, 2020. [67] K. J. Rothman, S. Greenland, and T. L. Lash, Modern epidemiology. Lippincott Williams & Wilkins, 2008. [68] S.Z.Abildstrom,S.Rasmussen, and M. Madsen,“Changesinhospitalizationrateandmortalityafteracutemyocardialinfarctionindenmark afterdiagnosticcriteriaandmethods changed,” European heart journal, vol. 26, no. 10, pp. 990–995, 2005. doi:10.1093/eurheartj/ehi039. [69] L. Østergaard, J. H. Butt, K. Kragholm, M. Schou, M. Phelps, R. Sørensen, M. Lamberts, G. Gislason, C. Torp-Pedersen, L. Køber, et al., “Incidence of acute coronary syndrome during national lock-down: Insights from nationwide data during the coronavirus disease 2019 (covid-19) pandemic,” American heart journal, vol. 232, pp. 146–153, 2020. doi:10.1016/j. ahj.2020.11.004. [70] P. N. Laursen, L. Holmvang, J. Lønborg, L. Køber, D. E. Høfsten, S. Helqvist, P. Clemmensen, H. Kelbæk, E. Jørgensen, J. F. Lassen, et al., “Comparison between patients included in randomizedcontrolled trials ofischemicheartdisease and real-worlddata.anationwidestudy,” American heart journal, vol. 204, pp. 128–138, 2018. doi:10.1016/j.ahj.2018.05.018. [71] A. B. Jensen, P. L. Moseley, T. I. Oprea, S. G. Ellesøe, R. Eriksson, H. Schmock, P. B. Jensen, L. J. Jensen, and S. Brunak, “Temporal disease trajectories condensed from population-wide registry data covering 6.2 million patients,” Nature communications, vol. 5, no. 1, pp. 1–10, 2014. doi:10.1038/srep36624. [72] M. K. Beck, A. B. Jensen, A. B. Nielsen, A. Perner, P. L. Moseley, and S. Brunak, “Diagnosis trajectories of prior multi-morbidity predict sepsis mortality,” Scientific reports, vol. 6, no. 1, pp. 1–9, 2016. doi:10.1038/srep36624. [73] J. X. Hu, M. Helleberg, A. B. Jensen, S. Brunak, and J. Lundgren, “A large-cohort, longitudinal study determines precancer disease routes across different cancer types,” Cancer research, vol. 79, no. 4, pp. 864–872, 2019. doi:10.1158/0008-5472.CAN-18-1677. [74] I. F. Jørgensen, A. Aguayo-Orozco, M. Lademann, and S. Brunak, “Age-stratified longitudinal study of alzheimer’s and vascular dementia patients,” Alzheimer’s & Dementia, vol. 16, no. 6, pp. 908–917, 2020. doi:10.1002/alz.12091. [75] C. López-Otín, M. A. Blasco, L. Partridge, M. Serrano, and G. Kroemer, “The hallmarks of aging,” Cell, vol. 153, no. 6, pp. 1194–1217, 2013. doi:10.1016/j.cell.2013.05.039. 55
REFERENCES ON DIAGNOSTIC CODES [76] E. L. Frome and H. Checkoway, “Use of poisson regression models in estimating incidence rates and ratios,” American journal of epidemiology, vol. 121, no. 2, pp. 309–323, 1985. doi: 10.1093/oxfordjournals.aje.a114001. [77] I. K. Kirk, C. Simon, K. Banasik, P. C. Holm, A. D. Haue, P. B. Jensen, L. J. Jensen, C. L. Rodríguez, M. K. Pedersen, R. Eriksson, et al., “Linking glycemic dysregulation in diabetes to symptoms, comorbidities, and genetics through ehr data mining,” Elife, vol. 8, e44941, 2019. doi:10.7554/eLife.44941. [78] H. Schütze, C. D. Manning, and P. Raghavan, Introduction to information retrieval. Cambridge University Press Cambridge, 2008, vol. 39. [79] D. B. West et al.,Introduction to graph theory. Prentice hall Upper Saddle River, 2001, vol. 2. [80] D.Xuand Y. Tian,“A comprehensivesurveyofclusteringalgorithms,” Annals of Data Science, vol. 2, no. 2, pp. 165–193, 2015. doi:10.1007/s40745-015-0040-1. [81] P. Spector. (2011). “Stat 133 class notes,” [Online]. Available: https://www.stat.berkeley. edu/~s133/all2011.pdf (visited on 03/18/2021). [82] L. Kaufman and P. J. Rousseeuw, Finding groups in data: an introduction to cluster analysis. John Wiley & Sons, 2009, vol. 344. [83] K. W. Johnson, J. Torres Soto, B. S. Glicksberg, K. Shameer, R. Miotto, M. Ali, E. Ashley, and J. T. Dudley, “Artificial intelligence in cardiology,” Journal of the American College of Cardiology, vol. 71, no. 23, pp. 2668–2679, 2018. doi:10.1016/j.jacc.2018.03.521. [84] I. S. Kohane, B. J. Aronow, P. Avillach, B. K. Beaulieu-Jones, R. Bellazzi, R. L. Bradford, G. A. Brat, M. Cannataro, J. J. Cimino, N. García-Barrio, et al., “What every reader should know about studies using electronic health record data but may be afraid to ask,” Journal of medical Internet research, vol. 23, no. 3, e22219, 2021. doi:10.2196/22219. [85] S. M. Van Dongen, “Graph clustering by flow simulation,” Ph.D. dissertation, 2000. [86] L. L. Gao, J. Bien, and D. Witten, Selective inference for hierarchical clustering, 2020. arXiv: 2012. 02936 [stat.ME]. [87] D. G. Kleinbaum and M. Klein, Survival analysis. Springer, 2012. [88] D. R. Cox, “Regression models and life-tables,” Journal of the Royal Statistical Society: Series B (Methodological), vol. 34, no. 2, pp. 187–202, 1972. [89] P. Royston and M. K. Parmar, “Restricted mean survival time: An alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome,” BMC medical research methodology, vol. 13, no. 1, pp. 1–15, 2013. doi:10.1186/1471-2288-13-152. [90] W. S. McCulloch and W. Pitts, “A logical calculus of the ideas immanent in nervous activity,” The bulletin of mathematical biophysics, vol. 5, no. 4, pp. 115–133, 1943. [91] M. A. Nielsen, Neural networks and deep learning. Determination press San Francisco, CA, 2015, vol. 25. [92] P. Baldi, S. Brunak, and F. Bach, Bioinformatics: the machine learning approach. MIT press, 2001. [93] P. Baldi, “Gradient descent learning algorithm overview: A general dynamical systems perspective,” IEEE Transactions on Neural Networks, vol. 6, no. 1, pp. 182–195, 1995. doi:10.1109/ 72.363438. 56
REFERENCES ON DIAGNOSTIC CODES [94] C. M. Bishop et al.,Neural networks for pattern recognition. Oxford university press, 1995. [95] E. J. Topol, “High-performance medicine: The convergence of human and artificial intelligence,” Nature medicine, vol. 25, no. 1, pp. 44–56, 2019. doi:10.1038/s41591-018-0300-7. [96] G. Tseng, Interpreting complex models with shap values,https://medium.com/@gabrieltseng/ interpreting-complex-models-with-shap-values-1c187db6ec83, (Accessed on 07/16/2021). [97] L. S. Shapley, H. Kuhn, and A. Tucker, “Contributions to the theory of games,” Annals of Mathematics studies, vol. 28, no. 2, pp. 307–317, 1953. [98] L. S. Shapley, “Stochastic games,” Proceedings of the national academy of sciences, vol. 39, no. 10, pp. 1095–1100, 1953. [99] S. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” arXiv preprint arXiv:1705.07874, 2017. [100] C. Molnar, Interpretable machine learning. Lulu. com, 2020. [101] N. R. Cook, “Useandmisuseof the receiver operatingcharacteristiccurve inriskprediction,” Circulation, vol. 115, no. 7, pp. 928–935, 2007. doi:10.1161/CIRCULATIONAHA.106.672402. [102] G. S. Collins, J. B. Reitsma, D. G. Altman, and K. G. Moons, “Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod) the tripod statement,” Circulation, vol. 131, no. 2, pp. 211–219, 2015. doi:10.1161/CIRCULATIONAHA. 114.014508. [103] F. E. Harrell Jr, Regression modeling strategies: with applications to linear models, logistic and ordinal regression, and survival analysis. Springer, 2015. [104] T. Fawcett, “An introduction to roc analysis,” Pattern recognition letters, vol. 27, no. 8, pp. 861– 874, 2006. doi:10.1016/j.patrec.2005.10.010. [105] R. van de Schoot, S. Depaoli, R. King, B. Kramer, K. Märtens, M. G. Tadesse, M. Vannucci, A. Gelman, D. Veen, J. Willemsen, et al., “Bayesian statistics and modelling,” Nature Reviews Methods Primers, vol. 1, no. 1, pp. 1–26, 2021. doi:10.1038/s43586-020-00001-2. [106] P.-C. Bürkner, “Brms: An r package for bayesian multilevel models using stan,” Journal of statistical software, vol. 80, no. 1, pp. 1–28, 2017. [107] T. M. Donovan and R. M. Mickey, Bayesian statistics for beginners: A step-by-step approach. Oxford University Press, USA, 2019. [108] A. Vehtari, A. Gelman, and J. Gabry, “Practical bayesian model evaluation using leave-oneout cross-validation and waic,” Statistics and computing, vol. 27, no. 5, pp. 1413–1432, 2017. doi:10.1007/s11222-016-9696-4. [109] M. Betancourt, “A conceptual introduction to hamiltonian monte carlo,” arXiv preprint arXiv:1701.02434, 2017. [110] J. Kruschke, “Doing bayesian data analysis: A tutorial with r, jags, and stan,” 2014. [111] J. K. Kruschke and T. M. Liddell, “Bayesian data analysis for newcomers,” Psychonomic bulletin & review, vol. 25, no. 1, pp. 155–177, 2018. 57
REFERENCES ON DIAGNOSTIC CODES [112] H. B. Bentzen and N. Høstmælingen, “Balancing protection and free movement of personal data:The neweuropeanunion generaldata protectionregulation,”Annals of internal medicine, vol. 170, no. 5, pp. 335–337, 2019. doi:10.7326/M18-2782. [113] H. K. Beecher, “Ethics and clinical research,” Bulletin of the World Health Organization, vol. 79, pp. 367–372, 2001. [114] What to notify? https://en.nvk.dk/howtonotify/whattonotify, (Accessed on 18/05/2021), May 2021. [115] C. J. Haug, “Turning the tables—the new european general data protection regulation,” New England Journal of Medicine, vol. 379, no. 3, pp. 207–209, 2018. doi:10.1056/NEJMp1806637. [116] M.D. Wilkinson,M.Dumontier,I. J.Aalbersberg,G.Appleton,M. Axton,A.Baak,N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al., “The fair guiding principles for scientific data management and stewardship,” Scientific data, vol. 3, no. 1, pp. 1–9, 2016. doi: 10.1038/sdata.2016.18. [117] K. K. Jensen, L. Whiteley, and P. Sandøe, RCR-A Danish textbook for courses in Responsible Conduct of Research. University of Copenhagen, Department of Food and Resource Economics, 2020. [118] F. Mölder, K. P. Jablonski, B. Letcher, M. B. Hall, C. H. Tomkins-Tinch, V. Sochat, J. Forster, S. Lee, S. O. Twardziok, A. Kanitz, et al., “Sustainable data analysis with snakemake,” F1000Research, vol. 10, p. 33, 2021. doi:10.12688/f1000research.29032.2. [119] Meet the gitlab team | gitlab,https://about.gitlab.com/company/team/, (Accessed on 05/17/2021). [120] M.Claussnitzer,J.H.Cho,R.Collins,N.J.Cox, E. T. Dermitzakis, M. E. Hurles,S.Kathiresan, E. E. Kenny, C. M. Lindgren, D. G. MacArthur, et al., “A brief history of human disease genetics,” Nature, vol. 577, no. 7789, pp. 179–189, 2020. doi:10.1038/s41586-019-1879-7. [121] M.Gheorghiade,G.Sopko, L.DeLuca, E. J.Velazquez,J.D.Parker,P. F. Binkley,Z. Sadowski, K. S. Golba, D. L. Prior, J. L. Rouleau, et al., “Navigating the crossroads of coronary artery disease and heart failure,” Circulation, vol. 114, no. 11, pp. 1202–1213, 2006. doi:10.1161/ CIRCULATIONAHA.106.623199. [122] R. R. Kalyani, “Glucose-lowering drugs to reduce cardiovascular risk in type 2 diabetes,” New England Journal of Medicine, vol. 384, no.13,pp. 1248–1260, 2021.doi:10.1056/NEJMcp2000280. [123] L. Busija, K. Lim, C. Szoeke, K. M. Sanders, and M. P. McCabe, “Do replicable profiles of multimorbidity exist? systematic review and synthesis,” European journal of epidemiology, vol. 34, no. 11, pp. 1025–1053, 2019. doi:doi.org/10.1007/s10654-019-00568-5. [124] T. Nagamine, B. Gillette, A. Pakhomov, J. Kahoun, H. Mayer, R. Burghaus, J. Lippert, and M. Saxena, “Multiscale classification of heart failure phenotypes by unsupervised clustering of unstructured electronic medical record data,” Scientific reports, vol. 10, no. 1, pp. 1–13, 2020. doi:10.1038/s41598-020-77286-6. [125] F. Crowe, D. T. Zemedikun, K. Okoth, N. J. Adderley, G. Rudge, M. Sheldon, K. Nirantharakumar, and T. Marshall, “Comorbidity phenotypes and risk of mortality in patients with ischaemic heart disease in the uk,” Heart, vol. 106, no. 11, pp. 810–816, 2020. doi:10. 1136/heartjnl-2019-316091. 58
REFERENCES ON DIAGNOSTIC CODES [126] F. D’Ascenzo, O. De Filippo, G. Gallone, G. Mittone, M. A. Deriu, M. Iannaccone, A. ArizaSolé, C. Liebetrau, S. Manzano-Fernández, G. Quadri, et al., “Machine learning-based predictionofadverse events following an acute coronary syndrome (praise):Amodellingstudy of pooled datasets,” The Lancet, vol. 397, no. 10270, pp. 199–207, 2021. doi:10.1016/S01406736(20)32519-8. [127] M. Motwani, D. Dey, D. S. Berman, G. Germano, S. Achenbach, M. H. Al-Mallah, D. Andreini, M. J. Budoff, F. Cademartiri, T. Q. Callister, et al., “Machine learning for prediction of all-cause mortality in patients with suspected coronary artery disease: A 5-year multicentre prospective registry analysis,” European heart journal, vol. 38, no. 7, pp. 500–507, 2017. doi: 10.1093/eurheartj/ehw188. [128] M. F. Gensheimer and B. Narasimhan, “A scalable discrete-time survival model for neural networks,” PeerJ, vol. 7, e6257, Jan. 2019, issn: 2167-8359. doi:10.7717/peerj.6257. [129] J.L.Mega,N. O. Stitziel,J. G.Smith,D.I. Chasman, M. J. Caulfield, J.J.Devlin,F.Nordio,C. L. Hyde, C. P. Cannon, F. M. Sacks, et al., “Genetic risk, coronary heart disease events, and the clinical benefit of statin therapy: An analysis of primary and secondary prevention trials,” The Lancet, vol. 385, no. 9984, pp. 2264–2271, 2015. doi:10.1016/S0140-6736(14)61730-X. [130] D. S. Wishart, C. Knox, A. C. Guo, S. Shrivastava, M. Hassanali, P. Stothard, Z. Chang, and J. Woolsey, “Drugbank: A comprehensive resource for in silico drug discovery and exploration,” Nucleic acids research, vol. 34, no. suppl_1, pp. D668–D672, 2006. doi:10.1093/nar/ gkj067. [131] A. Laupacis, D. L. Sackett, and R. S. Roberts, “An assessment of clinically useful measures of the consequences of treatment,” New England journal of medicine, vol. 318, no. 26, pp. 1728– 1733, 1988. doi:10.1056/NEJM198806303182605. [132] I. A. Scott, S. N. Hilmer, E. Reeve, K. Potter, D. Le Couteur, D. Rigby, D. Gnjidic, C. B. Del Mar, E. E. Roughead, A. Page, et al., “Reducing inappropriate polypharmacy: The process of deprescribing,” JAMA internal medicine, vol. 175, no. 5, pp. 827–834, 2015. doi:10.1001/ jamainternmed.2015.0324. [133] C. Rosendorff, D. T. Lackland, M. Allison, W. S. Aronow, H. R. Black, R. S. Blumenthal, C. P. Cannon,J.A.DeLemos, W. J. Elliott,L. Findeiss,et al., “Treatment of hypertensioninpatients with coronary artery disease: A scientific statement from the american heart association, american college of cardiology, and american society of hypertension,” Circulation, vol. 131, no. 19, e435–e470, 2015. doi:10.1016/j.jacc.2015.02.038. [134] M. Schmidt, L. V. Andersen, S. Friis, K. Juel, and G. Gislason, “Data resource profile: Danish heart statistics,” International journal of epidemiology, vol. 46, no. 5, 1368–1369g, 2017. doi:10. 1093/ije/dyx108. [135] A. Aguado, F. Moratalla-Navarro, F. López-Simarro, and V. Moreno, “Morbinet: Multimorbidity networks in adult general population. analysis of type 2 diabetes mellitus comorbidity,” Scientific reports, vol. 10, no. 1, pp. 1–12, 2020. doi:10.1038/s41598-020-59336-1. [136] J. Sundbøll, K. Adelborg, T. Munch, T. Frøslev, H. T. Sørensen, H. E. Bøtker, and M. Schmidt, “Positive predictivevalueofcardiovasculardiagnosesinthedanishnationalpatientregistry: A validation study,” BMJ open, vol. 6, no. 11, 2016. 59
REFERENCES ON DIAGNOSTIC CODES [137] D. S. Char, N. H. Shah, and D. Magnus, “Implementing machine learning in health care— addressing ethical challenges,” The New England journal of medicine, vol. 378, no. 11, p. 981, 2018. doi:10.1056/NEJMp1714229. [138] E. L. Fosbøl, J. H. Butt, L. Østergaard, C. Andersson, C. Selmer, K. Kragholm, M. Schou, M. Phelps, G. H. Gislason, T. A. Gerds, et al., “Association of angiotensin-converting enzyme inhibitor or angiotensin receptor blocker use with covid-19 diagnosis and mortality,” Jama, vol. 324, no. 2, pp. 168–177, 2020. 60
List of abbreviations ANN Artificial neural network ATC Anatomical Therapeutic Chemical BTH BigTempHealth CABG Coronary artery by-pass grafting CAG Coronary arteriography CCI Charlson co-morbidity index CHB Copenhagen Hospital Biobank Cox PH Cox proportional hazard CRS The Danish Civil Registration System DAR The Danish Registry for Causes of Death DNPR The Danish National Prescription Registry DTB The Danish Disease Trajectory Browser ECG Electrocardiogram EDHR The Eastern Danish Heart Registry GDPR General Data Protection Regulation GRACE Global Registry of Acute Coronary Events HDI Highest density interval HDL High-density lipoprotein HR Hazard ratio ICD International Classification of Diseases and Related Health Problems IHD Ischemic heart disease 61
, 3 Methods Data foundation and study population The main data source was the Danish National Patient Registry (NPR), where healthcare data from all encounters with a Danish hospital have been recorded since 1977. The data includes contact type (i.e. in-patient, out-patient and emergency room visits), date of contact start (e.g. admission), date of discharge, diagnosis codes and diagnosis type (e.g. primary codes that best describe the contact reason and secondary codes that complement the description of the contact)27. To obtain demographic data on patients such as date of birth, sex and status (dead or alive), data from NPR was linked to the Danish Civil Registration System (CRS) via the personal identification number28. Since 1994, diagnoses in NPR have been reported using the International Statistical Classification of Disease and Health Related Problems 10th Revision (ICD-10), which has a hierarchical structure comprising ICD-10 chapters, code blocks, level 3 and level 4 codes29. The NPR dataset used in this study covers the period 1994-2018 and contains data from 7 179 538 people corresponding to more than 142 million admissions. It contains 4 565 distinct level 3 ICD10 codes, which we here refer to as ICD-10 codes. Prior to analysis, level 4 codes were truncated to level 3 codes30. Patients that were deceased by the end of the study period were assigned the code Y99 and date of death was obtained from CRS28. To define the case population, all patients in NPR who had been assigned an ICD-10 code for angina pectoris (ICD-10 code: I20), acute myocardial infarction (ICD-10 code: I21) or chronic ischemic heart disease (ICD-10 code I25) in the period 1994-2018 were first identified. Next, patients who were assigned either of the diagnosis codes I20, I21, or I25 before the age of 18 years were excluded. Emigrants and tourists were also excluded, as their contacts with the Danish health-care system are likely to be sporadic and thus, not appropriate for modeling the temporality of IHD co-morbidities. All ICD-10 codes from chapters I-XIV assigned as a primary or secondary code (i.e., diagnosis types A, B or G) were included in the analysis. Codes assigned to less than 25 patients were excluded (Figure 1). Date of discharge in NPR was used to estimate age at first diagnosis (via linkage to CRS), identify pairs of diagnosis codes that were assigned to the same patient (i.e., diagnostic co-occurrences), define time-ordered directional diagnosis pairs and build disease trajectories. Experimental model and identification of diagnostic cooccurrences To study the temporal order of the entire spectrum of multi-morbidities in IHD patients, directional diagnosis pairs were computed and pieced together to form longer disease trajectories comprising three diagnoses. Disease trajectories were obtained by application of a modified version of the
, 4 disease trajectory program21. The two main steps in the program are (i) quantification of the overrepresentation of diagnostic co-occurrences between two diagnosis using relative risk (RR) and (ii) identification of directional diagnostic co-occurrences where the temporal order in which the diagnoses are assigned is statistically significant (i.e., directional diagnosis pairs). The first step (i) consists of a binomial test procedure that was developed in the original version of the disease trajectory program. Here all patients discharged with assignment of a certain diagnosis (e.g. D1) are considered exposed21. Following a pre-filtering step that uses the mean probability of assignment of any diagnosis from all discharges with diagnosis D1 (i.e., mean probability parameter), RRs for the diagnosis pairs (D1, D2) were obtained by considering each discharge as a Bernoulli sample. The algorithm then tests for statistical significance of the diagnosis pairs (D1, D2) being assigned to the same patient within five years compared to the mean probability parameter for D1. For each patient that were assigned diagnosis D1 at discharge (i.e., exposed), 10 000 comparison groups were formed by sampling from groups that were matched by sex, year of birth and week of discharge to correct for e.g., seasonal variation in diagnosis codes and changes in coding practices. RRs were then calculated as the ratio of exposed patients who were assigned diagnosis D2 compared to the ratio of unexposed patients who were assigned diagnosis D2. The level of significance was set to 0.001 to guard against false positives due to the binomial test procedure. The algorithm was run using R v. 3.4.0, Python v. 2.7, Python v. 3 and C++v. 1121,22. Defining directional diagnosis pairs and building disease trajectories The second step (ii) in the program establishes the directionality of diagnostic co-occurrences21. In contrast to previous versions of the disease trajectory program, a series of linear regression models (LRs) was introduced in this step. LRs were introduced to identify diagnostic co-occurrences (D1, D2) with a statistically significant difference between age at diagnoses D1 and D2, respectively. In cases where a patient was assigned the same ICD-10 code multiple times, only the earliest recorded diagnosis (with reference to end of contact) was included. In cases where a patient was assigned more than one diagnosis for the first time at the same contact, these instances would be included in the model. The dependent variable for the LRs was age at diagnosis and the independent variables were the diagnosis pair from step (i) (D1, D2), the type of diagnosis (A or B/G corresponding to primary or non-primary diagnosis), discharge date, type of patient (in-patient or out-patient), year of birth, and sex. These covariates were included to account for the possibility of differences in baseline
, 5 characteristics at diagnosis due to factors not related to the pathogenesis, e.g., changes in coding practice over the years. The P-value of the main effect for the diagnosis pair variable was used to determine the significance of difference in age between D1 and D2, where the predicted age at diagnoses D1 and D2 defined disease directionality. P-values were adjusted using the Bonferroni method. Significance level was set to 0.05. The LRs were applied to all diagnostic co-occurrences identified in step (i) and fitted using Statsmodels in Python 3.6.1031,32. The temporal order of the diagnostic co-occurrences (e.g., D1 and D2) with statistically significant age difference was determined by the age at diagnoses D1 and D2 that were computed from the LRs. Due to the number of covariates, it would be cumbersome to obtain a predicted age for each subgroup, e.g., primary diagnosis, females, and outpatients for each discharge year. Therefore, the predicted age was calculated using only the coefficients for diagnosis pairs and type of diagnosis, as we observed that the covariate for the diagnosis type generally had the highest impact (smallest Pvalues) on the age difference (Supplementary figure 1). The predicted age was calculated for D1 and D2 when assigned as primary diagnoses and the diagnosis that was predicted to be assigned at the youngest age was defined as the first diagnosis in the directional diagnosis (equivalent to length two trajectories) pair. In cases where the predicted age of D1 was less than that of D2 the length two trajectory would be D1 à D2. Patients who were assigned diagnoses D1 and D2 were characterized by following the trajectory and the order would be the one determined by the LR. To determine the directionality of disease trajectories containing three diagnoses, a similar approach to the directional diagnosis pairs was applied to sets of patients, who all were assigned diagnoses D1, D2 and D3. Diagnosis pairs with a significant directionality where diagnosis D2 of one pair was equal to diagnosis D1 of another pair were pieced together into a length three trajectory. The directionality of the disease trajectory was determined by extracting all patients with diagnoses D1, D2 and D3 and calculating the age of the three diagnoses. The age at diagnosis was calculated using the set of the three diagnoses (i.e., D1, D2 and D3) and type of diagnosis when assigned as primary diagnosis. The diagnoses were ordered by estimated age, from youngest to oldest age, e.g., the length three trajectory D1 à D2 à D3. As the directionality was recalculated for the disease trajectories, the directionality within a length two trajectory could change when combined with other diagnoses. In addition, length three trajectories could establish new directional associations, as their assemblage did not require the first and the last diagnoses to be a directional diagnosis pair (corresponding to diagnosis D1 and D3 in the example above). In the final analysis, only
, 6 trajectories computed based on sets of more than 50 patients were included; and patients with diagnoses D1 and D2 or D1, D2 and D3 were then said to follow the resulting trajectory. Using disease trajectories to characterize different IHD populations To assess if the temporal order of diagnoses was different between patients with different IHD manifestations, the cohort was split into seven groups defined by which IHD codes they were assigned. That is, one group was defined by having only I20, I21, or I25. Another group was defined by having two of the three codes, e.g. I20 and I21. A final group was defined by patients that were assigned I20, I21 and I25. These groups comprised a total of seven distinct index groups. Directional diagnoses pairs were now computed for each of these groups separately, following the same steps as described above (Figure 1). Results Cohort characterization and overview of the directional diagnosis pairs The case population was comprised of a total of 552 216 patients (57.2% males) diagnosed with IHD (ICD-10 codes I20, I21 or I25) in the 1994-2018 period. Mean age at first IHD diagnosis was 65.9 years for males and 70.8 years for females (Table 1). The size of the seven index groups ranged from 65 302 to 117 493 patients. Using the entire set of diagnosis codes assigned to patients in the case population, a total of 16 712 diagnostic co-occurrences (D1, D2) were identified. Among all the diagnostic co-occurrences, there were 1 459 pairs with a statistically significant difference in mean age at diagnosis D1 and D2, respectively. Those directional diagnosis pairs were followed by 546 704 patients (Figure 1). The directional diagnosis pairs that characterized the temporality of multi-morbidity in IHD were formed by a total of 408 distinct ICD-10 codes and generally, the temporal associations between IHD co-morbidities are highly diverse and dominated by other diseases of the cardiovascular system as well as metabolic diseases (Figure 2A and B). Among the 1 459 directional diagnosis pairs with a diagnosis code for IHD, 210 pairs contained a diagnosis code for either angina pectoris (ICD-10 code: I20), acute myocardial infarction (ICD-10 code: I21) or chronic IHD (ICD-10 code: I25). In 149 of these pairs (71.0%), the codes appeared as D2 consistent with the fact that IHD is a multi-factorial disease with a wide spectrum of etiologies and rarely the first manifestation of multimorbidity (Table 2).
, 7 Directional diagnosis pairs add information about the cohort As expected, essential hypertension (I10), non-insulin-dependent diabetes (E11), insulin-dependent diabetes (E10), atrial fibrillation (I48), heart failure (I50) and dyslipidemia (E78), were among the co-morbidities with the highest prevalence in the IHD cohort (Table 1). Not surprisingly, these diagnoses were also among the diagnoses that appeared in most distinct directional diagnosis pairs, indicating that they are diagnosed in a wide range of patient populations. Diagnoses that were prevalent in the cohort tended to be chronic conditions or manifestations of chronic disease, e.g., volume depletion (E86) (Table 2). A quantitative summary of the directional diagnosis pairs revealed additional characteristics regarding IHD co-morbidities that were not captured by the crude co-morbidity counts. For example, insulin-dependent diabetes (E10), obesity (E66), and disorders due to the use of alcohol (F10) were among the most frequently occurring diagnoses in the directional diagnosis pairs, although they were not among the diagnoses that were assigned to most patients in the population. Similarly, although rarely described in the context of IHD, diverticular disease of the intestine (K57) and osteoporosis without pathological fracture (M81) were among the co-morbidities that associated with most diagnoses in a directional manner. Conversely, a condition that is not chronic such as pneumonia (J18) was among the diagnoses assigned to most patients, yet it was not among the most frequently occurring diagnoses in the directional diagnosis pairs (Tables 1 and 2). Qualitive differences between risk factors, complications and diseases that can be both Generally, IHD risk factors appeared as D1 in most directional diagnosis pairs, e.g., hypertension (I10) and non-insulin dependent diabetes (E11) whereas conditions that may be related to both the etiology of IHD and represent a complication were equally likely to appear as either D1 or D2, e.g., heart failure (I50). Interestingly, atrial fibrillation (I48) most often appeared as diagnosis D1, indicating that in this population in this context, atrial fibrillation was often established at one of the earliest hospital encounters. However, the estimated age of a diagnosis of atrial fibrillation was not older than that of IHD, e.g. I20 à I48 (n = 80 560, P < 0.001). Patients following this directional diagnosis pair were also among the patients that were oldest when diagnosed with angina pectoris (ICD-10: I20). Among the ten most frequent directional diagnosis pairs containing I20, I21 or I25, there were considerable differences in the age distributions for age at IHD diagnosis. For example, the mean age at diagnosis for angina pectoris (I20) in the directional diagnosis pair E78 à I20 (n = 114 074, P < 0.001) was 6.6 years below the predicted age at angina pectoris for patients following the trajectory I20 à I48 (n = 80 560, P < 0.001) (Table 3).
, 8 Interestingly, we observed that the temporal relation between IHD and atrial fibrillation was not consistent across the different IHD subpopulations. For patients with a diagnosis code for acute myocardial infarction and no code for angina pectoris and chronic IHD I48 appeared as D1 in the directional diagnosis pair, I48 à I21 (n = 11 871, P < 0.001). In contrast, I21 à I48 appeared as a directional diagnosis pair when computed for the population who were indexed with diagnosis codes I20, I21 and I25 (n = 26 273, P < 0.001). This is consistent with the dual nature of atrial fibrillation that may be an age phenomenon as well as a disease complication. When the cohort was analyzed in its entirety, the diagnoses I21 and I48 did not comprise a directional diagnosis pair further pointing to the complexity of multi-morbidity in this domain (Table 4). In contrast to the differing trend between acute myocardial infarction and atrial fibrillation, the age of the diagnoses for heart failure (I50) and acute respiratory failure (J96) was similar in the two index groups where it appeared with a significant age difference between age at diagnosis of heart failure and respiratory failure. Among patients indexed with I20 and I25, estimated age at I50 was 70.5 years and 75.0 years for acute respiratory failure (n = 3 882, p < 0.0001). For patients indexed with I20, I21, and I25 it was 70.0 years and 74.8 years, respectively (n = 5 161, P < 0.001) (Table 4). Taken together, these observations suggest a more indisputable temporal association between I50 and J96, than between I21 and I48 among IHD patients. Comparison of length two and length three trajectories Next, the 1 459 directional diagnosis pairs were combined into 4 729 length three trajectories, i.e., disease trajectories comprised of three diagnoses. Selected length two and three trajectories with shared diagnoses were then compared. Generally, the predicted ages for IHD in trajectories containing IHD risk factors, e.g., dyslipidemia (E78) and hypertension (I10) were younger than trajectories containing a diagnosis code for IHD and no risk factors. In contrast, among diseased patients, the predicted age at IHD was older primarily reflecting that older people are more likely to be diseased. However, in trajectories that contained death (Y99) and a code for common IHD risk factors e.g., E78 or I10, age at death was generally young (Figure 3). Moreover, for some diagnoses the predicted age at one diagnosis varied considerably between trajectories. For example, the predicted age at angina pectoris (I20) was below 60.8 years for patients diagnosed with dyslipidemia and angina pectoris (i.e., E78 à I20). In contrast, the predicted age at angina pectoris was 68.0 for patients diagnosed with angina pectoris and heart failure (i.e., I20 à I50). When combined into a length three trajectory, i.e., E78 à I20 à I50 it was primarily the age at diagnosis
, 9 of heart failure that changed. The predicted age at diagnosis of heart failure among patients diagnosed with angina was 70.2 whereas it was 68.3 for patients diagnosed with dyslipidemia, angina pectoris and heart failure (Table 3). Interestingly, in some cases the length three trajectories revealed that the order of the diagnoses was not constant when comparing length two trajectories and length three trajectories, e.g., D1 à D2 and D3 àD2 à D1, respectively. For example, in the population of patients who had been assigned the diagnosis code for coxarthrosis (M16) and acute myocardial infarction (I21), the predicted age for coxarthrosis was younger than that of acute myocardial infarction (predicted ages 70.9 and 71.4 for M16 and I21, P < 0.001). However, for patients who had been assigned the diagnoses code for dyslipidemia (E78), coxarthrosis (M16) and acute myocardial infarction (I21) the order was reversed, i.e., the predicted age of I21 was younger than that of M16 (predicted ages 68.6 and 67.7 for M16 and I21). Again, this indicates the dual nature of coxarthrosis that might be a component of the metabolic syndrome or simply age-related degeneration (Table 3). In addition, the length three trajectories captured temporal trends in the cohort that length two trajectories did not identify. For example, albeit the length three trajectory I10 à I20 à I25 was followed by 105 652 patients, the diagnostic co-occurrence of I10 and I25 was not among the directional diagnosis pairs. The number of patients following the trajectory I10 à I20 à I25 corresponded to 66.1% of the patients following the trajectory I10 à I20 (n = 162 376, P < 0.001). A similar trend was observed when comparing the diagnoses E78 and I48 (no diagnostic co-occurrence) with patients following the trajectory E78 à I20 à I48 (n = 32 248), corresponding to 28.3% of patients who followed the trajectory E78 à I20 (n = 114 074) (Table 3). Discussion We have presented a strategy for analyzing the temporal order of IHD co-morbidities based on trends in nation-wide registry data of more than 500 000 IHD patients observed over a period of 24 years. By first establishing directional diagnosis pairs from diagnostic co-occurrences and then piecing together disease trajectories, we present a data-driven characterization of multi-morbidity based on nationwide health registry data. The directional diagnosis pairs captured temporal associations related to multi-morbidity that are usually omitted in a purely hypothesis-driven model. For example, the disease trajectory program identified osteoporosis as one of the co-morbidities that appeared in most directional diagnosis pairs. Similarly, there were cases where atrial fibrillation was diagnosed before IHD as well as after, consistent with the fact that it can represent an age phenomenon as well as a disease complication. Finally, the disease trajectories served as a tool to
, 10 identify cases where coxarthrosis was more likely to be a component of metabolic syndrome as opposed to an age-related degeneration in a single organ system (Table 2). The study has several strengths and limitations. Generally, the disease trajectory approach offers a novel strategy to identify temporal trends in the complex multi-morbidity landscape of IHD. We argue that this strategy can complement traditional studies, where multi-morbidity is assessed in a binary fashion, depending on their presence or absence11. An inherent limitation with NPR is that disease history of the individual patient is not complete, meaning that all data is conditioned on the fact that the patient went to the hospital. Similarly, there will of course be cases where the true age at first diagnosis was earlier than 1994 and hence will not be captured in the analysis. We assume that for patients with many contacts before 1994 this will apply to most diagnoses that are then likely to be registered at the same contact in the observation period. For patients with only few contacts before the start of the study that are close in time to 1994 this will only have limited impact on the results. Further, due to differences in year of birth for study participants combined with the broad inclusion criteria, differences in disease directionality may be confounded by factors not related to etiological differences. Ultimately, future studies will also include data from the Danish ICD-8 period, i.e., before 1994. In contrast to previous versions of the disease trajectory algorithm, we present a strategy where these confounders are adjusted for. In its current form the method can only define a direction in a predefined population in the sense that the direction is determined using the entire distribution for the population and not for the individual patient. Further studies are needed to determine the actual order of diagnoses in the individual patient and thereby disentangle the true mechanism in cases where one condition may both appear as a risk factor and a complication. In sum, we believe that the potential to conduct large-scale studies in nationwide registry data goes far beyond the more restrictive, classical usage. However, in its current state, the Danish disease trajectories are influenced by the inherent limitations within NPR and other health registries rooted in the universal healthcare in other Nordic regions, such as medical surveillance bias. So, additional studies that incorporate other data types are needed to identify the disease trajectories that reflect differences in disease progression patterns attributable to etiological difference. We believe that the value of studying nationwide health registry data comprehensively outweighs the limitations and calls for future collaborations between basic and physician scientists.
, 11 References 1. Timmis, A., Townsend, N., Gale, C. P., et al. European Society of Cardiology: Cardiovascular Disease Statistics 2019. Eur. Heart J. 41, 12–85 (2020) doi:10.1093/eurheartj/ehz859. 2. Forman, D. E., Maurer, M. S., Boyd, C., et al. Multimorbidity in Older Adults With Cardiovascular Disease. J. Am. Coll. Cardiol. 71, 2149–2161 (2018) doi:10.1016/j.jacc.2018.03.022. 3. Busija, L., Lim, K., Szoeke, C., Sanders, K. M. & McCabe, M. P. Do replicable profiles of multimorbidity exist? Systematic review and synthesis. Eur. J. Epidemiol. 34, 1025–1053 (2019) doi:10.1007/s10654-019-00568-5. 4. Gupta, R. & Wood, D. A. Primary prevention of ischaemic heart disease: populations, individuals, and health professionals. The Lancet 394, 685–696 (2019) doi:10.1016/S01406736(19)31893-8. 5. Ades, P. A. Cardiac Rehabilitation and Secondary Prevention of Coronary Heart Disease. N. Engl. J. Med. 345, 892–902 (2001) doi:10.1056/NEJMra001529. 6. Antman, E. M. & Braunwald, E. Managing Stable Ischemic Heart Disease. N. Engl. J. Med. 382, 1468–1470 (2020) doi:10.1056/NEJMe2000239. 7. Tran, J., Norton, R., Conrad, N., et al. Patterns and temporal trends of comorbidity among adult patients with incident cardiovascular disease in the UK between 2000 and 2014: A population-based cohort study. PLOS Med. 15, e1002513 (2018) doi:10.1371/journal.pmed.1002513. 8. Moseley, P. L. & Brunak, S. Identifying Sepsis Phenotypes. JAMA 322, 1416–1417 (2019) doi:10.1001/jama.2019.12591. 9. López-Otín, C., Blasco, M. A., Partridge, L., Serrano, M. & Kroemer, G. The Hallmarks of Aging. Cell 153, 1194–1217 (2013) doi:10.1016/j.cell.2013.05.039. 10. Barnett, K., Mercer, S. W., Norbury, M., et al. Epidemiology of multimorbidity and implications for health care, research, and medical education: a cross-sectional study. Lancet Lond. Engl. 380, 37–43 (2012) doi:10.1016/S0140-6736(12)60240-2. 11. Cosentino, F., Grant, P. J., Aboyans, V., et al. 2019 ESC Guidelines on diabetes, pre-diabetes, and cardiovascular diseases developed in collaboration with the EASD. Eur. Heart J. 41, 255–323 (2020) doi:10.1093/eurheartj/ehz486. 12. Køber, L., Swedberg, K., McMurray, J. J. V., et al. Previously known and newly diagnosed atrial fibrillation: a major risk indicator after a myocardial infarction complicated by heart
, 12 failure or left ventricular dysfunction. Eur. J. Heart Fail. 8, 591–598 (2006) doi:10.1016/j.ejheart.2005.11.007. 13. Benjamin, E. J., Levy, D., Vaziri, S. M., et al. Independent risk factors for atrial fibrillation in a population-based cohort. The Framingham Heart Study. JAMA 271, 840–844 (1994). 14. Gheorghiade, M., Sopko, G., De Luca, L., et al. Navigating the Crossroads of Coronary Artery Disease and Heart Failure. Circulation 114, 1202–1213 (2006) doi:10.1161/CIRCULATIONAHA.106.623199. 15. Adelborg, K., Szépligeti, S. K., Holland-Bill, L., et al. Migraine and risk of cardiovascular diseases: Danish population based matched cohort study. BMJ 360, k96 (2018) doi:10.1136/bmj.k96. 16. Kalyani, R. R. Glucose-Lowering Drugs to Reduce Cardiovascular Risk in Type 2 Diabetes. N. Engl. J. Med. 384, 1248–1260 (2021) doi:10.1056/NEJMcp2000280. 17. McMurray, J. J. V., Solomon, S. D., Inzucchi, S. E., et al. Dapagliflozin in Patients with Heart Failure and Reduced Ejection Fraction. N. Engl. J. Med. 381, 1995–2008 (2019) doi:10.1056/NEJMoa1911303. 18. Laursen, P. N., Holmvang, L., Lønborg, J., et al. Comparison between patients included in randomized controlled trials of ischemic heart disease and real-world data. A nationwide study. Am. Heart J. 204, 128–138 (2018) doi:10.1016/j.ahj.2018.05.018. 19. Rasmussen-Torvik Laura J., Shay Christina M., Abramson Judith G., et al. Ideal Cardiovascular Health Is Inversely Associated With Incident Cancer. Circulation 127, 1270– 1275 (2013) doi:10.1161/CIRCULATIONAHA.112.001183. 20. Fanaroff, A. C., Califf, R. M., Harrington, R. A., et al. Randomized Trials Versus Common Sense and Clinical Observation: JACC Review Topic of the Week. J. Am. Coll. Cardiol. 76, 580–589 (2020) doi:10.1016/j.jacc.2020.05.069. 21. Jensen, A. B., Moseley, P. L., Oprea, T. I., et al. Temporal disease trajectories condensed from population-wide registry data covering 6.2 million patients. Nat. Commun. 5, 4022 (2014) doi:10.1038/ncomms5022. 22. Siggaard, T., Reguant, R., Jørgensen, I. F., et al. Disease trajectory browser for exploring temporal, population-wide disease progression patterns in 7.2 million Danish patients. Nat. Commun. 11, 4952 (2020) doi:10.1038/s41467-020-18682-4. 23. Westergaard, D., Moseley, P., Sørup, F. K. H., Baldi, P. & Brunak, S. Population-wide analysis of differences in disease progression patterns in men and women. Nat. Commun. 10, 666 (2019) doi:10.1038/s41467-019-08475-9.
, 19 The Danish National Patient Registry in period 1994-2018 n patients = 7 179 538 Case population defined by ICD-10 codes I20, I21, I25 n patients = 552 216 ——– ——– ——– Index groups by ICD-10 codes (n patients): I20 (117 493), I21 (65 302), I25 (105 924), I20 I21 (14 346), I20 I25 (96 133), I21 I25 (66 293), I20 I21 I25 (86 725) ICD-10 codes excluded if not: ·assigned as an A, B or G code ·in chapter I-XIV ·assigned to >25 patients Patients following directional diagnosis pairs n patients = 546 704 1 Figure 1
, 20
, 21
, 22 IV VI III I II V VII Certain infectious and parasitic diseases Diseases of the blood and blood-forming organs and certain disorders involving the immune mechanism Neoplasms Endocrine, nutritional and metabolic diseases Diseases of the eye and adnexa Diseases of the nervous system Mental and behavioural disorders XI IX VIII X XII XIII XIV Diseases of the digestive system Diseases of the respiratory system Diseases of the circulatory system Diseases of the ear and mastoid process Diseases of the genitourinary system Diseases of the musculoskeletal system and connective tissue Diseases of the skin and subcutaneous tissue International Statistical Classification of Disease and Health Related Problems 10th Revision chapter I through XIV
, 23 ! !"##$%&%'()*+,-./"*%,0, Supplementary figure 1
8|Risk stratification of 72,249 patients with ischemic heart disease: a retrospective study linking prior multimorbidity, biochemical data, and genetics 89
, 1 Risk stratification of 72,249 patients with ischemic heart disease: a retrospective study linking prior multi-morbidity, biochemical data, and genetics Amalie D. Haue, MD1,2,*, Peter C. Holm, MSc1,*, Karina Banasik, PhD1, Agnete Troen Lundgaard, MSc1, Victorine P. Muse, MEng1, Timo Röder, MSc1, David Westergaard, PhD1, Piotr J. Chmura, MSc1, Troels Siggaard, MSc1, Alex H. Christensen, PhD2,3, Peter E. Weeke, PhD2, Erik Sørensen, PhD4, Sisse R. Ostrowski, DMSc4,5, Daníel F. Gudbjartsson, PhD6, Hilma Hólm, PhD6, Kasper K. Iversen, DMSc3, Lars V. Køber, DMSc2,5, Henrik Ullum, PhD7, Henning Bundgaard, DMSc2,5†, Søren Brunak, PhD1,8† 1Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Blegdamsvej 3B, DK-2200 Copenhagen, Denmark 2Department of Cardiology, The Heart Center, Rigshospitalet, Blegdamsvej 9, DK-2100 Copenhagen, Denmark 3Department of Cardiology, Copenhagen University Hospital, Herlev Hospital, Borgmester Ib Juuls Vej 1, DK-2730 Herlev 4Department of Clinical Immunology, Copenhagen University Hospital, Rigshospitalet, Blegdamsvej 9, DK-2100 Copenhagen, Denmark 5Department of Clinical Medicine, University of Copenhagen, Rigshospitalet, Blegdamsvej 3B, DK-2200 Copenhagen, Denmark 6deCODE Genetics-Amgen, Sturlugata 8, 101 Reykjavik, Iceland 7Statens Serum Institut, Artillerivej 5, DK-2300 Copenhagen, Denmark 8Copenhagen University Hospital, Rigshospitalet, Blegdamsvej 9, DK-2100 Copenhagen, Denmark *Denotes equal contribution. †To whom correspondence should be addressed: [email protected], Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, DK-2200 Copenhagen, Denmark Word count (excl. abstract, references, -tables, figures, and captions): 3,442
, 2 Key points Question: How can data on prior multi-morbidities be used for risk stratification and subgrouping of patients with ischemic heart disease (IHD)? Findings: Unsupervised clustering identified IHD patient subgroups defined by the entire multimorbidity spectrum could be stratified according to different risks of secondary ischemic events and death from non-IHD causes. Subgrouping was supported biologically by biochemical and genetic data. Meaning: Unsupervised clustering discriminated between patients at highand low risk of secondary ischemic events paving the way for differentiated management of patients with IHD, i.e., development of precision medicine in this domain.
, 3 Abstract Importance: The clinical use in risk stratification of prior multi-morbidity data spanning decades; and biochemical data and genetics needs to be realized. Objective: Explore how unsupervised clustering of multi-morbidity from electronic health records (EHRs) can be used to risk stratify patients with ischemic heart disease (IHD). Design: Retrospective study that used data from the Danish National Patient Registry, EHRs from Eastern Denmark and genetic data from Copenhagen Hospital Biobank. Unsupervised clustering for identification of IHD subgroups based on multi-morbidities in male and female patients was performed separately. Cox-proportional hazard models were used to compute hazard ratios (HRs) for primary and secondary outcomes in subgroups. For a subset of patients, subgroups were characterized by available biochemical data and polygenic scores (PGS) for 10 different traits associated with IHD. Discrimination by cluster-based risk stratification was tested against the Charlson and Elixhauser co-morbidity indices (CMIs) and evaluated using the C-index. Setting: A population-based study using nation-wide healthcare data. Multi-morbidity was assessed using spectrum-wide ICD-10 codes assigned prior to IHD. Participants: Patients diagnosed with IHD in the years 2004-2016. Exposures: None. Main Outcomes: Primary outcome was a composite outcome comprised of acute myocardial infarction, unstable angina pectoris, revascularization (not related to initial IHD diagnosis) and death caused by IHD. Secondary outcome was any death not caused by IHD. Results: In a cohort of 72,249 patients, we identified 31 male and 25 female subgroups. In total, 11 clusters had HR > 1 for the primary outcome only, and two clusters in each sex had HR > 1 for both the primary and secondary outcome. We also identified two subgroups that had HR < 1 for the primary outcome. Differences in biochemical data largely reproduced the subgrouping. The PGS for total cholesterol, LDL cholesterol, diabetes, atrial fibrillation had significant effect sizes in at least one subgroup. Cluster C-indices were generally higher for female than for male clusters and comparable to CMI C-indices. Conclusions and Relevance: Unsupervised clustering can be used to discriminate between highand low-risk subgroups based on prior multi-morbidities that are consistent with well-known phenotypes. Subgroup-specific biochemical and PGS characteristics were identified. Overall, the discrimination obtained by the clusters was comparable to CMIs. For patients who were not at increased risk it was better.
, 4 Introduction Worldwide, ischemic heart disease (IHD) affects more than 125 million people. Over the last decades, there has been a decline in mortality rates1,2. Yet, IHD patients remain at high risk of complications and have an age-adjusted mortality that is two to six times higher than in people without IHD3. Nearly 85% of IHD patients are multi-morbid making it a complex phenotype, with disease courses that vary between patients and between sexes4–6. Further, multi-morbidity is a crucial component for reliable outcome prediction, including susceptibility to disease progression4,7,8. It is currently being assessed using only few diagnoses and consequently often disregard the true complexity9. As biomedical research is being transformed by an exponential growth in patient profiling data and analytical capabilities, conditions to conduct studies that address the complexity of multi-morbidity are now conceivable8,10. Advanced analysis of routinely obtained biochemical data is expected to supplement data on multi-morbidities10,11. Polygenic scores (PGS) are becoming a means of patient-profiling although the true potential is still to be confirmed12. To date PGS are not incorporated in clinical care systematically and more studies using genetic data that model the risk of secondary ischemic events among IHD patients are needed13–17. This includes explorative studies that test how data from nation-wide health registries, EHRs (i.e. routinely collected data that capture most aspects of health care) and genotypes can be combined18–20. Aggregated and possibly weaker phenome-wide information that appear stronger in more homogenous subgroups may increase the statistical power of future trials21,22. In turn, this may refine risk stratification of IHD patients, allowing differentiation of treatment between patients10. Here, we present a study of 72,249 IHD patients demonstrating how unsupervised clustering based on multi-morbidities can discriminate between patients at highand low-risk of secondary ischemic events and death of non-IHD causes (i.e., highand low-risk patients). Among low-risk patients, the performance of our model is better than that of other existing co-morbidity indices. By integrating biochemical and genetic data, we assessed if the unsupervised clustering correlate with biology. In sum, the study presents a strategy for defining and using multi-morbidity, that advances precision medicine to practice.
, 11 Finally, the MCL algorithm also identified clusters where the secondary outcome was increased, while risks of the primary outcome were not significant. These included 4.M04, 4.M06, 4.M07, 4.M08, 4.M10, 4.M13, 4.M29, 4.F06, 4.F11, 4.F12 and 4.F21 (Figures 2 and 3); encompassing distinct clusters characterized by cancer (e.g. C44.3/C50.9), disorders due to the use of alcohol (e.g. F10.2), atrial fibrillation/heart failure (e.g. I48.9/I50.9) and chronic obstructive pulmonary diseases (COPD) (e.g. J44.9) (Table 2). In both sexes, the PGS for atrial fibrillation was significant in clusters enriched for atrial fibrillation/heart failure (4.M04 and 4.F11) (Table 2). In these clusters, C-indices for the CMIs were higher than the cluster C-indices. However, in some clusters without significantly altered risk for the primary or secondary outcome, a cluster-based risk stratification performed better than CMIs, e.g., 0.76, 0.73 and 0.72 for the secondary outcome for cluster 4.M14 (eTable 6 in the Supplement). Discussion The complexity and scale of multi-morbidity calls for development of risk stratification methods that acknowledge patient heterogeneity9. By application of the MCL algorithm to spectrum-wide multi-morbidities from 72,249 IHD patients, we identified patient subgroups at different risk of secondary ischemic events as well as death from non-IHD causes. For clusters at increased risk of secondary ischemic events, HRs ranged from 0.73 to 1.65 for males and 0.64 to 1.76 for females (Table 2). The study demonstrated that the MCL algorithm captured biologically relevant patterns within the ICD-10 terminology, which had support in biochemical and genetic data. The fact that only modest effects on HRs were observed when changing the inflation parameter suggests overall cluster stability (Figure 2). Importantly, discrimination of survival models fitted using clusters as co-variates was comparable the CMIs. In some subgroups a cluster-based estimation even outperformed CMIs (eTable 6 in the Supplement). In our analyses only one cluster (4.M01) at decreased risk of secondary ischemic events, associated with PGS distributions in comparison to that of the other clusters (Table 2). This adds to the growing body of literature regarding sex-specific manifestations of IHD5. In comparison to other clustering strategies, our model succeeded in including rarer phenotypes by application of a flexible algorithm that only depends on a limited set of a priori assumptions regarding explanatory variables6,43. Explicitly, we applied a data-driven strategy that discriminated between phenotypes that are at risk of both secondary ischemic events and death from non-IHD causes. For example, in both sexes, clusters
, 12 enriched for atrial fibrillation only associated with increased risk of death from non-IHD causes and correlated with the PGS for atrial fibrillation. The observation is consistent with a previous finding that anti-coagulative therapy should not be intensified in patients with atrial fibrillation and a recent percutaneous coronary intervention7. In addition we show that the association between IHD subgroups and PGS allows for adding genetics and possibly thereby enhancing the clinical utility of genetic data14,15. The study has several strengths, but also limitations. By performing an enrichment analysis based on past multi-morbidities and support it with biochemical and genetic data, we showcase how integration of complementary datatypes may lead to a more robust analysis. The study confirmed prior findings, while it potentially also leads to new discoveries. By studying patients aligned temporally according to IHD onset and their past spectrum-wide multi-morbidities, our work underscores the complexity of multi-morbidity in a model that in principle can be applied to any disease of interest. The complementary analysis of biochemical data and genetics provide evidence that this is indeed feasible. As our modelling strategy allows for analysis of cluster granularity (as a function of the inflation parameter), the complexity of multi-morbidity is an integral part of the model. Evidently, this is also a limitation. However, we argue that the framework adds value as this limitation is not unique to our model and often pre-fixed6,19. Yet, multi-morbidity is often studied binarily where neither the severity nor the order with which they were diagnosed is taken into account9. We assert that models allowing for such complexity will increase the likelihood of identifying relevant, yet weaker phenotypic and genetic information in subgroups. Information that potentially has the capacity to account for differences in disease risk and burden that are currently unknown. Further, the model requires data that are available in most health-care systems (i.e., ICD-10 codes). For successful translation of our findings into actionable risk stratification principles of value in clinical decisions making, further studies that apply the method at the N = 1 level are needed.
, 13 References 1. Nabel, E. G. & Braunwald, E. A tale of coronary artery disease and myocardial infarction. N. Engl. J. Med. 366, 54–63 (2012) doi:10.1056/NEJMra1112570. 2. Antithrombotic Trialists’ (ATT) Collaboration, Baigent, C., Blackwell, L., et al. Aspirin in the primary and secondary prevention of vascular disease: collaborative meta-analysis of individual participant data from randomised trials. Lancet 373, 1849–1860 (2009) doi:10.1016/S0140-6736(09)60503-1. 3. WHO | Prevention of Recurrences of Myocardial Infarction and Stroke Study. WHO https://www.who.int/cardiovascular_diseases/priorities/secondary_prevention/country/en/in dex1.html (2020). 4. Forman, D. E., Maurer, M. S., Boyd, C., et al. Multimorbidity in Older Adults With Cardiovascular Disease. Journal of the American College of Cardiology 71, 2149–2161 (2018) doi:10.1016/j.jacc.2018.03.022. 5. Pagidipati, N. J. & Peterson, E. D. Acute coronary syndromes in women and men. Nature Reviews Cardiology 13, 471–480 (2016) doi:10.1038/nrcardio.2016.89. 6. Crowe, F., Zemedikun, D. T., Okoth, K., et al. Comorbidity phenotypes and risk of mortality in patients with ischaemic heart disease in the UK. Heart 106, 810–816 (2020) doi:10.1136/heartjnl-2019-316091. 7. Lopes, R. D., Heizer, G., Aronson, R., et al. Antithrombotic Therapy after Acute Coronary Syndrome or PCI in Atrial Fibrillation. New England Journal of Medicine 380, 1509–1524 (2019) doi:10.1056/NEJMoa1817083. 8. Joshi, A., Rienks, M., Theofilatos, K. & Mayr, M. Systems biology in cardiovascular disease: a multiomics approach. Nat Rev Cardiol 18, 313–330 (2021) doi:10.1038/s41569020-00477-1. 9. Glynn, L. G. Multimorbidity: another key issue for cardiovascular medicine. The Lancet 374, 1421–1422 (2009) doi:10.1016/S0140-6736(09)61863-8. 10. Hemingway, H., Asselbergs, F. W., Danesh, J., et al. Big data from electronic health records for early and late translational cardiovascular research: challenges and potential. Eur. Heart J. 39, 1481–1495 (2018) doi:10.1093/eurheartj/ehx487. 11. Agniel, D., Kohane, I. S. & Weber, G. M. Biases in electronic health record data due to processes within the healthcare system: retrospective observational study. BMJ 361, k1479 (2018) doi:10.1136/bmj.k1479.
, 14 12. Mosley, J. D., Gupta, D. K., Tan, J., et al. Predictive Accuracy of a Polygenic Risk Score Compared With a Clinical Risk Score for Incident Coronary Heart Disease. JAMA 323, 627–635 (2020) doi:10.1001/jama.2019.21782. 13. Buysschaert, I., Carruthers, K. F., Dunbar, D. R., et al. A variant at chromosome 9p21 is associated with recurrent myocardial infarction and cardiac death after acute coronary syndrome: The GRACE Genetics Study. Eur Heart J 31, 1132–1141 (2010) doi:10.1093/eurheartj/ehq053. 14. Mega, J. L., Stitziel, N. O., Smith, J. G., et al. Genetic risk, coronary heart disease events, and the clinical benefit of statin therapy: an analysis of primary and secondary prevention trials. The Lancet 385, 2264–2271 (2015) doi:10.1016/S0140-6736(14)61730-X. 15. Nikpay, M. & Mohammadzadeh, S. Phenome-wide screening for traits causally associated with the risk of coronary artery disease. J Hum Genet 65, 371–380 (2020) doi:10.1038/s10038-019-0716-z. 16. Patel, R. S., Tragante, V., Schmidt, A. F., et al. Subsequent Event Risk in Individuals With Established Coronary Heart Disease. Circ Genom Precis Med 12, e002470 (2019) doi:10.1161/CIRCGEN.119.002470. 17. Vaara Satu, Tikkanen Emmi, Parkkonen Olavi, et al. Genetic Risk Scores Predict Recurrence of Acute Coronary Syndrome. Circulation: Cardiovascular Genetics 9, 172–178 (2016) doi:10.1161/CIRCGENETICS.115.001271. 18. Khera Amit V. & Kathiresan Sekar. Is Coronary Atherosclerosis One Disease or Many? Circulation 135, 1005–1007 (2017) doi:10.1161/CIRCULATIONAHA.116.026479. 19. Kirk, I. K., Simon, C., Banasik, K., et al. Linking glycemic dysregulation in diabetes to symptoms, comorbidities, and genetics through EHR data mining. eLife 8, e44941 (2019) doi:10.7554/eLife.44941. 20. Jensen, P. B., Jensen, L. J. & Brunak, S. Mining electronic health records: towards better research applications and clinical care. Nature Reviews Genetics 13, 395–405 (2012) doi:10.1038/nrg3208. 21. Mantel, N. & Haenszel, W. Statistical Aspects of the Analysis of Data From Retrospective Studies of Disease. J Natl Cancer Inst 22, 719–748 (1959) doi:10.1093/jnci/22.4.719. 22. Denny, J. C., Ritchie, M. D., Basford, M. A., et al. PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene–disease associations. Bioinformatics 26, 1205– 1210 (2010) doi:10.1093/bioinformatics/btq126.
, 15 23. Sørensen, E., Christiansen, L., Wilkowski, B., et al. Data Resource Profile: The Copenhagen Hospital Biobank (CHB). Int J Epidemiol dyaa157 (2020) doi:10.1093/ije/dyaa157. 24. Schmidt, M., Schmidt, S. A. J., Sandegaard, J. L., et al. The Danish National Patient Registry: a review of content, data quality, and research potential. Clin Epidemiol 7, 449– 490 (2015) doi:10.2147/CLEP.S91125. 25. Helweg-Larsen, K. The Danish Register of Causes of Death. Scand J Public Health 39, 26– 29 (2011) doi:10.1177/1403494811399958. 26. Nielsen, A. B., Thorsen-Meyer, H.-C., Belling, K., et al. Survival prediction in intensivecare units based on aggregation of long-term disease history and acute physiology: a retrospective study of the Danish National Patient Registry and electronic patient records. The Lancet Digital Health 1, e78–e89 (2019) doi:10.1016/S2589-7500(19)30024-X. 27. Schmidt, M., Pedersen, L. & Sørensen, H. T. The Danish Civil Registration System as a tool in epidemiology. Eur J Epidemiol 29, 541–549 (2014) doi:10.1007/s10654-014-9930-3. 28. Sundbøll, J., Adelborg, K., Munch, T., et al. Positive predictive value of cardiovascular diagnoses in the Danish National Patient Registry: a validation study. BMJ Open 6, e012832 (2016) doi:10.1136/bmjopen-2016-012832. 29. Adelborg, K., Sundbøll, J., Munch, T., et al. Positive predictive value of cardiac examination, procedure and surgery codes in the Danish National Patient Registry: a population-based validation study. BMJ Open 6, e012817 (2016) doi:10.1136/bmjopen2016-012817. 30. Enright, A. J., Van Dongen, S. & Ouzounis, C. A. An efficient algorithm for large-scale detection of protein families. Nucleic acids research 30, 1575–84 (2002). 31. Acosta-Herrera, M., Kerick, M., González-Serna, D., et al. Genome-wide meta-analysis reveals shared new loci in systemic seropositive rheumatic diseases. Annals of the Rheumatic Diseases 78, 311–319 (2019) doi:10.1136/annrheumdis-2018-214127. 32. Wuttke, M. A catalog of genetic loci associated with kidney function from analyses of a million individuals. Nature Genetics 51, 22 (2019). 33. Malik, R., Chauhan, G., Traylor, M., et al. Multiancestry genome-wide association study of 520,000 subjects identifies 32 loci associated with stroke and stroke subtypes. Nature Genetics 50, 524–537 (2018) doi:10.1038/s41588-018-0058-3. 34. Nielsen, J. B., Thorolfsdottir, R. B., Fritsche, L. G., et al. Biobank-driven genomic discovery yields new insight into atrial fibrillation biology. Nature Genetics 50, 1234–1239
, 16 (2018) doi:10.1038/s41588-018-0171-3. 35. Mahajan, A., Taliun, D., Thurner, M., et al. Fine-mapping type 2 diabetes loci to singlevariant resolution using high-density imputation and islet-specific epigenome maps. Nature Genetics 50, 1505–1513 (2018) doi:10.1038/s41588-018-0241-6. 36. van der Harst Pim & Verweij Niek. Identification of 64 Novel Genetic Loci Provides an Expanded View on the Genetic Architecture of Coronary Artery Disease. Circulation Research 122, 433–443 (2018) doi:10.1161/CIRCRESAHA.117.312086. 37. Surakka, I., Horikoshi, M., Mägi, R., et al. The impact of low-frequency and rare variants on lipid levels. Nature Genetics 47, 589–597 (2015) doi:10.1038/ng.3300. 38. Bland, J. M. & Altman, D. G. Multiple significance tests: the Bonferroni method. BMJ 310, 170 (1995) doi:10.1136/bmj.310.6973.170. 39. Violán, C., Roso-Llorach, A., Foguet-Boreu, Q., et al. Multimorbidity patterns with Kmeans nonhierarchical cluster analysis. BMC Family Practice 19, 108 (2018) doi:10.1186/s12875-018-0790-x. 40. R Core Team. R: A Language and Environment for Statistical Computing. (R Foundation for Statistical Computing, 2019). 41. Mölder, F., Jablonski, K. P., Letcher, B., et al. Sustainable data analysis with Snakemake. F1000Research 10, 33 (2021). 42. Editors, T. P. M. Observational Studies: Getting Clear about Transparency. PLOS Medicine 11, e1001711 (2014) doi:10.1371/journal.pmed.1001711. 43. MCL - a cluster algorithm for graphs. https://micans.org/mcl/index.html.
, 17 Tables Table 1: Cohort demographics, outcomes, and most prevalent non-IHD ICD-10 codes. Cohort demographics Total Males Females P-value1 Number of patients (%) 72,249 45,576 (63.1) 26,673 (36,1) Mean age at index (SD) 63.9 (11.9) 62.9 (11.6) 65.6 (12.1) < 0.001 Outcomes, number of cases Total Males Females P-value2 Secondary ischemic events (%) 11,476 (15.9) 7,891 (17.3) 3,585 (13.4) < 0.001 n Myocardial infarction 4,522 2,851 1,671 n Revascularization 4,910 3,660 1,250 n Death caused by IHD 1,943 1,379 664 Death from non-IHD causes (%) 6,177 (8.5) 3,902 (8.6) 2,275 (8.5) Censored (%) 54,596 (75.6) 33,783 (74.1) 20,813 (78.0) Outcomes, time to event Mean time to event in years (SD) P-value1 Count Males Females Secondary ischemic events 2.25 (1.91) 2.27 (1.92) 2.21 (1.89) 0.82 n Myocardial infarction 1.65 (1.41) 1.63 (1.42) 1.67 (1.40) 0.31 n Revascularization 1.48 (1.34) 1.49 (1.36) 1.45 (1.31) 0.33 n Death caused by IHD 1.14 (1.46) 1.18 (1.47) 1.05 (1.44) 0.060 Death from non-IHD causes 3.36 (1.80) 3.34 (1.81) 3.39 (1.80) 0.19 Censored 4.27 (1.13) 4.25 (1.15) 4.30 (1.11) 0.002 Total 3.72 (1.65) 3.67 (1.67) 3.81 (1.60) < 0.001 ICD-10 Description Males Females Total I10 Primary (essential) hypertension 14,519 10,316 24,835 E78.0 Hypercholesterolemia 7,843 4,939 12,782 E11.9 Type 2 diabetes mellitus: Without complications 4,895 2,667 7,562 I48.9 Atrial fibrillation and atrial flutter, unspecified 4,509 2,567 7,076 I50.9 Heart failure, unspecified 4,061 2,101 6,162 R07.9 Chest pain 3,450 2,423 5,873 H25.9 Senile cataract, unspecified 2,795 2.969 5,764 J18.9 Pneumonia, unspecified 3,329 2,265 5,504 E78.5 Hyperlipidaemia, unspecified 3,306 1,696 5,002 J44.9 Chronic obstructive pulmonary disease, unspecified 2,452 2,174 4,626 1 Students T-test, two-sided 2 χ2-test testing distribution of outcomes among males vs. females. Italics not included.
, 18 Table 2: Demographics of selected clusters. List of ICD-10 code descriptions in eTable 8. ID Mean age at index HR, primary outcome (95% CI) Adj. Pvalue Main enrichment Significant PGS. Trait (effect)5 Years (SD) P-value4 ICD-10 O/E-ratio 4.M01 58.5 (10.8) < 0.001 0.93 (0.87 to 0.99) 0.030 I21.3 2.46 TC (+) I25 (-) LDL (+) 4.M02 62.4 (10.8) 0.20 1.51 (1.40 to 1.62) < 0.001 E10.9 28.8 E11 (+) E11.9 11,7 H36.0 74.8 4.M19 67.1 (10.4) < 0.001 1.28 (1.11 to 1.48) 0.04 I69.4 33.7 None 4.M20 67.6 (9.3) < 0.001 1.65 (2.22 to 2.69) < 0.001 I70.8 58.4 I48 (+) 4.M25 54.9 (10.9) < 0.001 1.38 (1.15 to 1.65) 0.030 I30.9 2.6 None I73.0 4.4 M62.6 26.4 4.F01 62.7 (11.0) < 0.001 0.74 (0.66 to 0.84) < 0.001 M75.4 2.1 None I20.1 2.03 4.F02 64.0 (12.5) < 0.001 1.73 (1.57 to 1.91) < 0.001 E10.9 34.3 E11 (+) E11.9 11.9 H36.0 95.0 4.F04 63.7 (12.5) < 0.001 1.23 (1.10 to 1.38) 0.013 I21 2.52 None 4.F13 64.6 (11.6) 0.33 1.13 (0.93 to 1.38) > 0.99 I63.9 21.1 None 4.F16 70.0 (10.5) < 0.001 1.76 (1.48 to 2.10) < 0.001 I73.9 43.0 I25 (-) ID Mean age at index HR, secondary outcome (95% CI) Adj. Pvalue Main enrichment Significant PGS. Trait (effect) Years (SD) P-value4 ICD-10 O/E-ratio 4.M04 65.8 (10.6) < 0.001 1.51 (1.35 to 1.70) < 0.001 I48.2 11.1 I48 (+) E11 (-) I50.1 7.1 4.M08 66.2 (11.2) < 0.001 3.10 (2.81 to 3.4) < 0.001 J44.0 65.7 None 4.M10 57.9 (10.3) < 0.001 3.14 (2.72 to 3.62) < 0.001 F10.2 196.2 None 4.M23 71.4 (9.5) < 0.001 1.86 (1.57 to 2.20) < 0.001 C61 80.9 None 4.F06 67.4 (11.3) < 0.001 3.93 (2.93 to 3.69) < 0.001 J44.9 16.7 None 4.F11 70.1 (11.4) < 0.001 1.42 (1.19 to 1.70) < 0.001 I34.0 7.3 I48 (+) I48.9 12.7 I50.0 4.0 4.F12 69.9 (9.7) < 0.001 1.75 (1.47 to 2.08) < 0.001 C50.9 291.7 E11 (-) 4.F21 59.0 (11.2) < 0.001 3.13 (2.38 to 4.12) < 0.001 F10.2 104.7 None 4Two-sided Wilcoxon test for difference in mean age of patients in cluster against mean age of the other clusters. P-value Bonferroni adjusted. 5ICD-10 code except for LDL = LDL cholesterol and TC = total cholesterol. O/E ratio = Ratio of observed and expected term frequencies.
, 19 Figure captions Figure 1: Flowchart: Data sources, study population and outcomes. Gray: Identification. Blue: Screening. Red: Eligibility. Green: Inclusion and outcomes. NPR = The Danish National Patient Registry. IHD: ischemic heart disease (ICD-10 codes I20-I25 or block R94). CAG = coronary arteriography. CCTA: coronary computed tomography angiography. ICD10: International Statistical Classification of Diseases and Related Health Problems 10th Revision. SKS: Sundhedsvæsenets Klassifikationssystem (The Danish Health Authority Classification System). Figure 2: Clusters as a function of inflation parameter. Clusters plotted vertically at six different granularity levels displaying their branching in response to increased inflation going left to right. Number of male clusters (A) ranged from 24 to 35 and number of female clusters (B) ranged from 16 to 31 clusters across the six levels. Each black vertical bar corresponds to one granularity level (level 1 through 6 going left to right). Length of the bars is proportional to the number of patients in each cluster. Annotations inside black vertical bars denote a single, sex-specific cluster uniquely per granularity level. Granularity level indicated above black vertical bars. Color key according to the HR of a cluster. If a cluster is partly colored (i.e., different from grey) the entire cluster has the HR indicated by the color. Figure 3: Risk of secondary ischemic events and risk of death from non-IHD causes stratified by cluster. Forest plots where clusters at granularity level four are shown against HR for risk primary (left) and secondary (right) for males (A) and females (B), respectively. Table on the left with cluster, number of patients (size) mean age at index (age) and cluster characteristics based on O/E-ratios. “*”: No extreme enrichment in cluster, i.e., no terms in cluster with O/E-ratio > 10. Boxes correspond to HR and bars indicate 95% CI. A: Male clusters at granularity level four. B: Female clusters at granularity level four. X-axis: HR for a cluster relative to average HR of all other clusters. Y-axis: Clusters arranged by risk of secondary ischemic events, increasing from top to bottom. IHD: Ischemic heart disease. GERD: Gastroesophageal reflux disease. GI-tract: Gastro-intestinal tract. HR: Hazzard ratio. CI: Confidence interval. For a tabulated version of the results, see eTable 5 (survival analyses) and the Supplement with O/E-ratios.
, 20
, Feature selection and pre-processing Using exclusively data registered up until, and including, the index date, we extracted a total of 595 features from NPR, EDHR, BTH and CHB. Features included diagnosis codes, procedure codes (e.g. imaging examinations and surgical procedures), results from biochemical tests, clinical features such as blood pressure, height, weight, and smoking status, and coronary pathology at time zero. Moreover, a panel of ten different PGSs obtained from CHB were included and computed following a protocol described in previous work19. An overview of the features can be found in table 1. Diagnosis codes (ICD-10) and procedure codes (SKS/NOMESCO) were extracted from NPR and included in the model as one-hot encoded features.17 Diagnosis codes assigned more than 20 years prior to the index date and codes assigned to less than 1% of patients in the cohort were excluded. Biochemical test results and their reference ranges were originally stored in the databases Labka and BCC and in this study obtained from BTH.23 Tests were either annotated in accordance with the Nomenclature, Properties and Units (NPU) or various local coding systems.24 Biochemical tests drawn more than five years prior to the index date and tests performed on less than 5% of the cohort were excluded. Results of biochemical tests were included in the model as categorical variables. Each test could either take the value -1, 0, or 1 indicating if the value was below, within, or above the normal reference range. In cases where one patient had more than one value available the most recent was used. A total of 23 clinical features were included of which 8 were the same as those used in the GRACE Risk Score 2.0 (GRACE2.0) input features (“Clinical characteristics 1”, Table 1). Blood pressure and heart rate were obtained from the unstructured part of the EHRs using an in-house developed information extraction tool that recognizes features from Danish clinical text. The remaining clinical features were extracted from the structured data. Available continuous features were Z-score normalized prior to model development and missing values were then encoded with a value of zero. For the categorical features we used one-hot encoding with an additional category to designate when variables were missing. Model architecture and development To model time-to-event data and allow for censoring, we used the generic discrete-time survival model for neural networks described in Gensheimer 2019.25 In this model, follow-up time is divided into a fixed number of intervals and the model estimates a conditional hazard for each interval, i.e., the probability of failure given that no event has occurred before that particular interval. PMHnet-alpha uses 30 intervals separated in time such that event times in the training data were evenly distributed across all intervals. The implementation applied the PyTorch machine learning framework using the authors’ Keras implementation (Nnet-survival) as a reference.26 For benchmarking against GRACE2.0, the clinical features were divided into the two categories “clinical characteristics 1” that was comprised of the eight GRACE2.0 input features and “clinical characteristics 2” that contained the other 15 clinical features. We used a simple multilayer perceptron architecture with one to three hidden layers and 10-200 ReLUactivated neurons in each of the hidden layers. The final output layer was a fully connected sigmoid activated layer that outputs conditional hazards for each of the 30 different time points. We added dropout to each of the hidden layers to regularize the network and prevent over-fitting. Number of layers, number of neurons, and dropout rate for each layer were fine-tuned through hyperparameter optimization. In hyperparameter-optimization, we used five-fold cross-validation to obtain an optimism-corrected estimate of model performance on the training set. Networks were trained by application of stochastic gradient descent, with the negative log-likelihood as the loss function. Prior to modelling input features were grouped in the categories “diagnosis codes”,“procedure codes”,“clinical characteristics 1”,“clinical 4
, characteristics 2”,“clinical laboratory tests”, and “polygenic scores” (Table 1). PMHnet-alpha was developed by sequentially combining these six data type categories. Model performance PMHnet-alpha was used to predict survival curves on all individuals in the validation cohort. Model performance was evaluated through careful assessment of model discrimination and calibration based on these predictions. To assess calibration and discrimination graphically, the distribution of predicted survival probabilities at five years was used to construct five different risk strata. For each of the risk groups, the mean predicted survival curve was plotted together with a Kaplan-Meier estimate of the observed survival.27 For model calibration, patient outcomes were plotted against five-year mortality predictions and a locally estimated scatterplot smoother was used to obtain a non-parametric estimate of the calibration curve.28 The smoother was estimated using the loess function in R with the default bandwidth and the area 𝑎 between the regression line and the ideal line 𝑥 = 𝑦 was used as a measure of model calibration.15 Using bootstrap resampling, confidence intervals of the calibration score could be obtained.29 To complement the graphical methods, we also calculated time-dependent AUC scores as well as Harrel’s C-index. Time-dependent AUC (tdAUC) scores were calculated using the tdROC R-package.30 Harrel’s C-index was calculated using the rcorr.cens function from the Hmisc R-package.31 Time-dependent AUC scores were evaluated at six months, one year, three years, and five years after time zero. To further benchmark PMHnet-alpha we calculated the GRACE2.0 for all patients in the validation cohort.6We extracted the javascript source code for the GRACE2.0 webtool using Google Chrome’s Developer Tools. The javascript code was then manually converted to an R-package to enable automatised computation of GRACE2.0 on the entire cohort. Since GRACE2.0 does not allow for missing features, we imputed missing variables using the missForest R-package.32 Imputation was only performed to acquire a GRACE2.0 score for all patients; missing values were left missing in PMHnet-alpha. Explainability analysis was performed using SHapley Additive exPlanations (SHAP) values. The relative feature importance of the different input feature categories was computed by summarizing the absolute SHAP values for each feature category33. Ethics declaration Study design, methods, and results were reported in agreement with the TRIPOD statement.34,35 The scripts and source code are available upon reasonable request to the authors. The study was approved by The Danish National Committee on Health Research Ethics project IDs 60833 and 1708829. Results Baseline characteristics of derivation cohort A total of 39 746 patients (67·3% males) were included in the derivation cohort. It was randomly split into a training (n = 34 746) and a testing set (n = 5000) used for model development and validation, respectively (Figure 1). Table 2 summarizes the patient baseline characteristics and the data composition of the training set, testing set, and the derivation cohort as a whole. Mean age at index date was 66·1 5
, years and 19·5% of the cohort presented with a STEMI. The restricted mean follow up time was 1651 days (SE: 2·58) and the primary outcome (all-cause mortality) was observed in 21·3% of the subjects in the training set and 22·2% of those in the testing set (Figure S1). Performance and benchmarking of PMHnet-alpha PMHnet-alpha was developed by sequentially adding and combining the six different input feature types listed in Table 2, which led to six different intermediate models labelled 1 through 6. Through hyperparameter optimization, the best performing combination of hyperparameters were identified and then trained on the entire training data (n = 34746). The various PMHnet models then underwent validation using the unseen testing data (n = 5000). Figure 2 displays the tdAUC at six-months, one-year, and three-years evaluation time points for the different versions of PMHnet-alpha defined by their respective input feature combinations. The preliminary results have shown that the performance of PMHnet-alpha remains unchanged by simply adding the PGSs to model 6, indicating that the information from the PGSs may be included in the other data types. Serving as a reference for the observed performances, the GRACE2.0 score for six-months, one-year, and three-years was calculated on the entire training set. The eight variables used in GRACE2.0 were available for 51·4% of the cohort. Imputation was performed for patients with missing values (see Methods). The tdAUC of GRACE2.0 was 0·768, 0·768, and 0·729 at six-month, one-year, and three-year timepoints, respectively (figure S2). Thus, except for model 1, PMHnet-alpha outperformed GRACE2.0 in prediction all-cause mortality at all the evaluated timepoints (Figure 2). Generally, the performance of PMHnet-alpha improved as more feature categories were added. The version of PMHnet-alpha with diagnosis codes only, outperformed GRACE2.0 in prediction of three-year all-cause mortality but had poorer discrimination at six months and one year. Interestingly, the version of PMHnet-alpha trained using the feature category “clinical characteristics 1” (corresponding to the GRACE2.0 input features) outperformed GRACE2.0 at the six-months and three-years timepoints, likely because non-linear correlations are considered by PMHnet-alpha. At the one-year evaluation timepoint the performance of this version of PMHnet-alpha was almost identical to that of GRACE2.0 for the patients with missing values. However, the performance of GRACE2.0 on the patients without any missing values was better than that of PMHnet-alpha (Figure 2). In contrast to GRACE2.0, the predictive performance of PMHnet-alpha was the same for patients with and without missing values (Table 4). The best performing version of PMHnet-alpha was model 6 that based its predictions on all feature categories (Figure 2). The model has excellent model discrimination (Table 3) and a good calibration with a score, 1−𝛼, of 0·944 (CI: [0·925;0·964]) (Figure 3). In terms of calibration, the model provided too pessimistic predictions for patients in the 50% risk range but otherwise looked good (Figure S3). Explainability analysis of PMHnet-alpha Next, the predictive values (i.e., feature importance) of the different features was quantified using SHAP values for PMHnet-alpha model 6 (in the following referred to as the final version). Globally, age was the most predictive feature and its feature importance increased at increasing timepoints. That is, age contributed more to the prediction at the three-years than at the six-months evaluation timepoint. The sequential addition of input features types showed that diagnosis codes were an important addition to the model that only used GRACE2.0 features. However, in the final version of PMHnet-alpha, no diagnosis codes were among the top ranking features. This indicates that the performance of a model with many 6
, features with relatively little predictive value can be better than one with few highly predictive ones (Figure 4). In the final version of PMHnet-alpha, clinical features used to compute the GRACE2.0, e.g. Killip class, presentation with STEMI (yes/no), and SBP were among the most predictive ones. Aside from age, the two most predictive features (largest SHAP values) were coronary pathology and smoking status, which are not included in the GRACE2.0 score. Like age, coronary pathology was consistently important across the three evaluation timepoints. Yet, the importance decreased slightly over time suggesting that most patients with severe coronary pathology died relatively soon after index (i.e., earlier time points). The opposite was true for age, where the feature importance increased at increasing time points (Figure 4). For the feature coronary pathology, DA and 1VD were predictive features driving towards death whereas 2VD and 3VD were predictive features driving towards survival. Of the four possible categories for coronary pathology, 3VD had the largest impact on the prediction. The impact on model output for DA and 1VD was about the same. Unlike age and coronary pathology, the feature importance of smoking was relatively constant over the three different timepoints. SHAP values for smoking were negative, meaning that they contributed negatively to the outcome (Figure 4). The predictive importance of the feature categories was also assessed collectively. Although the biochemical tests were generally not among the highest ranked features, the input feature category “clinical laboratory tests” had the highest overall feature importance (Table 5) Explainability analysis of PMHnet-alpha for one patient, N = 1 Finally, the feature importance was evaluated for four individual cases with a predicted three-years survival chance between 0·75-1·00, 0·50-0·75, 0·25-0·50, and 0·00-0·25, respectively. Each case was represented by the linear sum of the SHAP values for the predictions at the three evaluation timepoints. Predictive features with the largest absolute SHAP values were annotated on figure 5. No SHAP values below 0·02 were annotated (Figure 5). Case 1 was characterized by very few SHAP values above the threshold for annotation indicating that the patient had few hospitalizations prior to index date. The ICD-10 code N81 (female genital prolapse) was one of few annotated values and contributed positively to the estimated chance of survival. In contrast, cases 2,3, and 4had overlapping top-ranked features. For example, the ICD-10 code J44 (chronic obstructive pulmonary disease) was among the top-ranked predictive features with the largest absolute SHAP value. In all three cases it contributed negatively to predictions (Figure 5). In contrast, 1VD contributed negatively to the prediction for case 3 and positively to the prediction for case 4. Similarly, age contributed negatively to the prediction for case 3 (82·7 years) and positively to the prediction for case 4(57·9 years). Age was not among the top-scoring predictive features with the largest absolute SHAP value for case 2. For case 2 the most predictive features were comprised of diagnosis codes, procedures codes, and two biochemical tests that were below reference range, where low arterial partial pressure of oxygen and low plasma glucose were among the most predictive biochemical features. In contrast, low plasma lactate was among the positive predictive features for case 3. Interestingly, missing information regarding familial IHD was among the most predictive features for case 4 (Figure 5). Generally, the feature attributions of PMHnet-alpha was consistent with physiological response mechanisms, e.g. low oxygen saturation and low plasma glucose vs. plasma lactate and the chances of survival (Figure 5). Thus, the SHAP analysis serves as a proof-of-principle of model development. In addition, the observations derived from the SHAP analysis demonstrate that the feature importance in PMHnet-alpha is context-dependent. For example, the analysis showed that in case 1 the age of 75·2 years contributed 7
, negatively, whereas the age of 82·7 contributed positively to the prediction for case 3; but that the overall survival probability was higher for case 1 than for case 3. Lastly, the SHAP analysis illustrates that PMHnet-alpha can integrate information of an administrative nature as well as information reflecting physiology, i.e., missing information regarding familiar IHD as well as biochemical test results. Discussion In this study we have developed PMNnet-alpha, a neural network-based model for prediction of all-cause mortality in IHD patients. PMHnet-alpha outperformed GRACE2.0 when the predictions were based on same eight input features as those used in GRACE2.0. By extending the volume and variety of input features using six different feature categories amounting to a total of 595 distinct input features, the tdAUC of PMHnet-alpha improved from 0·75 to 0·876 (Figure 2). It is a defining attribute of PMHnet-alpha that features were included with practically no a priori assumptions regarding potential predictive importance. Similarly, no assumptions regarding linear dependency between variables were made, which allowed us to include input features based on availability. This contrasts with most previously published prediction models in this domain, where model features are often limited to the few most predictive ones8,36. PMHnet-alpha is evidence that it is indeed an advantage to include more than the most predictive features in a risk scoring model by outperforming GRACE2.0. Using SHAP values, we were able to rank the level of information of the different input features and characterize the model systematically.33 Overall, the SHAP analysis served as a proof-of-principle and also exemplified yet another advantage with PMHnet-alpha over traditional risk scoring algorithms as PMHnet-alpha was able to estimate the survival rate for all patients in the cohort irrespective of data availability (Figure 5). Explicitly, GRACE2.0 could not provide a risk estimate of case 1 without feature imputation, whereas PMHnet-alpha successfully established a survival curve for this patient. Based on the results from the SHAP analysis, we could assess that the estimate seemed reasonable based on the predictive attribution to the available features (Figure 2). Further, PMHnet-alpha enabled us to identify predictive features that are not included in conventional models and the SHAP analysis allowed us to quantify their importance. For example, smoking, which is not included in GRACE2.0, was among the most informative features at all three evaluation timepoints. It was among the most predictive features at all three time-points with a nearly constant SHAP value (Figure 4). This adds to the literature regarding lack of improved survival in IHD patients, with a history of smoking.4Taken together, the flexibility of PMHnet-alpha with regards to input features is highly valuable in modern medicine, where data availability and access vary between countries. Thus, PMHnet-alpha can in principle be trained on a data set based on feature availability in another country or healthcare system, independent of the features that were available at the time and place of development. Ongoing work is centered on external validation and interpretation of the performance when PGSs are added. The actual realized disease history may contain the same information as the current PGSs could contribute. However, the lack of improvement could also be due to the current architecture of the neural network. As we are still setting up experiments targeting the analyses of the predictive performance of the PGSs, these results are not discussed thoroughly in this version of the manuscript. This study has several strengths but also some limitations. First, in this study PMHnet-alpha was only benchmarked against one of the extensively validated risk scores in this domain. However, generally other risk scores, e.g. TIMI, are not superior to GRACE2.0 and there is even evidence that GRACE2.0 has the best discrimination in some populations37. Further, PMHnet-alpha was developed in a secondary prevention setting and thus, the Framingham risk score is unsuitable for benchmarking.5However, the derivation cohort was different from that of GRACE2.0 as it included patients with acute manifestations 8
, of IHD as well as chronic IHD (Table 2). In contrast to the Framingham, TIMI and GRACE2.0 scores, PMHnet-alpha is not only based a limited set of risk factors only5. Overall we take our prediction model several steps further than conventional risk scores, by developing a model that is flexible with respect to volumes and variety of input data. We show that the performance of PMHnet-alpha is superior to conventional survival models based on the same features and that the performance of PMHnet-alpha increases gradually when the number of features is increased (Figure 2). Having developed and tested PMHnet-alpha, we argue that there is a need to rethink the usage of clinical data within medical research for modern medicine to take full advantage of the increase in data volumes, variety, and analytical capabilities. Another advantage of the strategy that we applied to risk assessment is that it may overcome the existing barrier between risk scoring algorithms and clinical practice.38 Explicitly, PMHnet-alpha could potentially be included as an application in EHRs for real-time risk assessment of IHD patients, as it computes a risk for all individual patients based on the available data (Figure 5). We have however not tested systematically at what point the missingness becomes a problem. Finally, PMHnet-alpha includes framework for model explainability, which is a common concern against the usage of ML models within the medicine9. In this study, we showcase that an explainable model is a crucial component of successful survival modeling based on ML, such as neural networks. We showcase that explainability analysis may both provide grounds for model confidence as well as discovery of important predictive features that have not been included in traditional survival models. 9
, References 1 James SL, Abate D, Abate KH, et al. Global, regional, and national incidence, prevalence, and years lived with disability for 354 diseases and injuries for 195 countries and territories, 1990– 2017: a systematic analysis for the Global Burden of Disease Study 2017. The Lancet 2018; 392: 1789–858. 2 Takahashi K, Serruys PW, Fuster V, et al. Redevelopment and validation of the SYNTAX score II to individualise decision making between percutaneous and surgical revascularisation in patients with complex coronary artery disease: secondary analysis of the multicentre randomised controlled SYNTAXES trial with external cohort validation. The Lancet 2020; 396: 1399–412. 3 Gupta R, Wood DA. Primary prevention of ischaemic heart disease: populations, individuals, and health professionals. The Lancet 2019; 394: 685–96. 4 Arbel Y, FitzGerald G, Yan AT, et al. Temporal trends in all-cause mortality according to smoking status: Insights from the Global Registry of Acute Coronary Events. International Journal of Cardiology 2016; 218: 291–7. 5 Wilson PWF, D’Agostino RB, Levy D, Belanger AM, Silbershatz H, Kannel WB. Prediction of Coronary Heart Disease Using Risk Factor Categories. Circulation 1998; 97: 1837–47. 6 Fox KAA, FitzGerald G, Puymirat E, et al. Should patients with acute coronary disease be stratified for management according to their risk? Derivation, external validation and outcomes using the updated GRACE risk score. BMJ Open 2014; 4: e004425. 7 Hung J, Roos A, Kadesjö E, et al. Performance of the GRACE 2.0 score in patients with type 1 and type 2 myocardial infarction. European Heart Journal 2021; 42: 2552–61. 8 Antman EM, Cohen M, Bernink PJLM, et al. The TIMI Risk Score for Unstable Angina/Non–ST Elevation MI. JAMA 2000; 284: 835. 9 Rajkomar A, Dean J, Kohane I. Machine Learning in Medicine. New England Journal of Medicine 2019; 380: 1347–58. 10 Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine 2019; 25: 44–56. 11 Brajer N, Cozzi B, Gao M, et al. Prospective and External Evaluation of a Machine Learning Model to Predict In-Hospital Mortality of Adults at Time of Admission. JAMA Network Open 2020; 3: e1920733. 12 Gibson WJ, Nafee T, Travis R, et al. Machine learning versus traditional risk stratification methods in acute coronary syndrome: a pooled randomized clinical trial analysis. Journal of Thrombosis and Thrombolysis 2019; 49: 1–9. 13 D’Ascenzo F, De Filippo O, Gallone G, et al. Machine learning-based prediction of adverse events following an acute coronary syndrome (PRAISE): a modelling study of pooled datasets. The Lancet 2021; 397: 199–207. 10
, 14 Motwani M, Dey D, Berman DS, et al. Machine learning for prediction of all-cause mortality in patients with suspected coronary artery disease: a 5-year multicentre prospective registry analysis. European Heart Journal 2016; : ehw188. 15 Steele AJ, Denaxas SC, Shah AD, Hemingway H, Luscombe NM. Machine learning models in electronic health records can outperform conventional survival models for predicting patient mortality in coronary artery disease. PLOS ONE 2018; 13: e0202344. 16 Özcan C, Juel K, Flensted Lassen J, von Kappelgaard LM, Mortensen PE, Gislason G. The Danish Heart Registry. Clin Epidemiol 2016; 8: 503–8. 17 Schmidt M, Schmidt SAJ, Sandegaard JL, Ehrenstein V, Pedersen L, Sørensen HT. The Danish National Patient Registry: a review of content, data quality, and research potential. Clinical Epidemiology 2015; : 449. 18 Nielsen AB, Thorsen-Meyer H-C, Belling K, et al. Survival prediction in intensive-care units based on aggregation of long-term disease history and acute physiology: a retrospective study of the Danish National Patient Registry and electronic patient records. The Lancet Digital Health 2019; 1: e78–89. 19 Sørensen E, Christiansen L, Wilkowski B, et al. Data Resource Profile: The Copenhagen Hospital Biobank (CHB). International Journal of Epidemiology 2021; 50: 719–720e. 20 Helweg-Larsen K. The Danish Register of Causes of Death. Scandinavian Journal of Public Health 2011; 39: 26–9. 21 Schmidt M, Pedersen L, Sørensen HT. The Danish Civil Registration System as a tool in epidemiology. European Journal of Epidemiology 2014; 29: 541–9. 22 Arlot S, Celisse A. A survey of cross-validation procedures for model selection. Statistics Surveys 2010; 4. DOI:10.1214/09-ss054. 23 Grann, Erichsen R, Nielsen, Frøslev, Thomsen R. Existing data sources for clinical epidemiology: The clinical laboratory information system (LABKA) research database at Aarhus University, Denmark. Clinical Epidemiology 2011; : 133. 24 Arendt JFH, Hansen AT, Ladefoged SA, Sørensen HT, Pedersen L, Adelborg K. Existing Data Sources in Clinical Epidemiology: Laboratory Information System Databases in Denmark. Clinical Epidemiology 2020; Volume 12: 469–75. 25 Gensheimer MF, Narasimhan B. A scalable discrete-time survival model for neural networks. PeerJ 2019; 7: e6257. 26 GitHub - MGensheimer/nnet-survival: Discrete-Time Survival Model for Neural Networks. GitHub. https://github.com/MGensheimer/nnet-survival (accessed Aug 18, 2021). 27 Royston P, Altman DG. External validation of a Cox prognostic model: principles and methods. BMC Medical Research Methodology 2013; 13: 33. 28 Harrell FE, Lee KL, Mark DB. Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Stat Med 1996; 15: 361–87. 11
, 29 Hesterberg T. Bootstrap. Wiley Interdisciplinary Reviews: Computational Statistics 2011; 3: 497–526. 30 Rodríguez-Álvarez MX, Meira-Machado L, Abu-Assi E, Raposeiras-Roubín S. Nonparametric estimation of time-dependent ROC curves conditional on a continuous covariate. Statistics in Medicine 2016; 35: 1090–102. 31 Hmisc: Harrell Miscellaneous. https://CRAN.R-project.org/package=Hmisc (accessed Aug 18, 2021). 32 Stekhoven DJ, Buhlmann P. MissForest--non-parametric missing value imputation for mixed-type data. Bioinformatics 2011; 28: 112–8. 33 Lundberg SM, Erion G, Chen H, et al. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence 2020; 2: 56–67. 34 Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD). Circulation 2015; 131: 211– 9. 35 Collins GS, Moons KGM. Reporting of artificial intelligence prediction models. The Lancet 2019; 393: 1577–9. 36 Boersma E, Pieper KS, Steyerberg EW, et al. Predictors of Outcome in Patients With Acute Coronary Syndromes Without Persistent ST-Segment Elevation. Circulation 2000; 101: 2557– 67. 37 Aragam KG, Tamhane UU, Kline-Rogers E, et al. Does Simplicity Compromise Accuracy in ACS Risk Prediction? A Retrospective Analysis of the TIMI and GRACE Risk Scores. PLoS ONE 2009; 4: e7947. 38 Manfrini O, Bugiardini R. Barriers to clinical risk scores adoption. European Heart Journal 2007; 28: 1045–6. 12
, Tables Table 1 Table 1: Overview of the different input feature categories. Category Features Hyperparameters Clinical characteristics 1 Age, pulse, systolic blood pressure, cardiac arrest at index, abnormal cardiac enzymes killip class, creatinine, st-segment deviation Clinical characteristics 2 Abnormal ECG, CCS class, diastolic blood pressure, coronary artery dominance, familiary IHD, patient height, icd-device or pm, ischemia test, LVEF, NYHA class, sex, smoking status, coronary pathology, patient weight Diagnosis codes 422 different level-4 ICD-10 diagnosis codes. Kept as-is or converted to level-3, block, or chapter during model development. • icd10 level • time cutoff Procedure codes 154 different NOMESCO procedure codes corresponding to various radiological examinations and surgical procedures • time cutoff Clinical laboratory tests 85 different lab tests with results categorised as below,within, or above the reference range • time cutoff Polygenic scores 10 different traits • winsorization 13
, Figure 5: Specific model evaluation for N = 1 of four distinct patients. Prediction survival rate represented as SHAP values for four distinct patients in the cohort represented as case 1,case 2,case 3, and case 4 reading left to right, top to bottom. Cases are from four different risk strata with a predicted risk of 15.1 (case 1), 34.6 (case 2), 72.8 (case 3), and 95.6 (case 4) at the three-years evaluation timepoint. X-axis: The three evaluation timepoints, i.e. six-months, one-year, and three-years. Y-axis: Predicted survival rate for the four different cases at the intersection of the two colors blue and red. Blue represents features that contribute positively to the prediction. Red represents features that contribute negatively to the prediction. Highest ranked features were annotated above a SHAP value of 0.02 were annotated. SHAP: SHapley Additive exPlanations. 20
, Figure S1: Kaplan-Meier estimates for the PMHnet derivation cohort. Kaplan-Meier estimates for the training (blue) and testing (red) set of 34749 and 5000 patients, respectively. X-axis: Follow-up time in years. Y-axis: Survival percentage. Figure S2: Time-dependent receiver operating characteristics for the GRACE2.0 on the PMHnet test set. Three tdROCs correspondong to the three evaluation six-months, one-years and three-years evaluatio timepoints goind left to right. Red: Patients in the test set where all GRACE2.0 input features were available. Blue: Patients where imputation was performed prior to computation of the GRACE2.0 due to missing features. 21
, Figure S3: Calibration curve of PMHnet-alpha model 6 Calibration curve showing the observed vs. predicted all-cause mortality at 5 years. The dotted line represents the ideal calibration line. The area between the non-parametric loess curve and the ideal line can be used as a measure for calibration. 22
10 |Polypharmacy and drug dosage modifications: a longitudinal analysis of 3.5 million electronic health records 135
, 1 Polypharmacy and drug dosage modifications: 1 A longitudinal analysis of 3.5 million electronic 2 health records 3 Cristina Leal Rodríguez1, Gianluca Mazzoni1, Amalie Dahl Haue1, Robert Eriksson1,2, Jorge 4 Hernansanz Biel1, Lisa Cantwell1, David Westergaard1,3, Kirstine G. Belling1, Søren 5 Brunak1,4* 6 1Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical 7 Sciences, University of Copenhagen, DK-2200 Copenhagen, Denmark 8 2Department of Pulmonary and Infectious Diseases, Nordsjællands Hospital, DK-3400 9 Hillerød, Denmark 10 3Department of Obstetrics and Gynaecology, Copenhagen University Hospital Hvidovre, 11 DK-2650 Hvidovre, Denmark 12 4Department of Health Technology, Technical University of Denmark, DK-2800 Kongens 13 Lyngby, Denmark 14 * To whom correspondence should be addressed: [email protected] 15
, 2 Abstract 16 Polypharmacy continues to grow in importance because of multimorbid, aging populations. 17 However, a comprehensive, drug-specific analysis of concomitant therapies and response by 18 means of dosage adjustments has not been performed. In this longitudinal population-wide 19 study of dosage adjustments, we used electronic health records from 3.5 million inpatient 20 admissions at Danish hospitals spanning the years 2008-2016 and identified associations 21 between concomitant drug pairs and dosages in close to 185 million treatment episodes. The 22 study suggests that dosage modifications are proxies of drug-drug interactions (DDIs) in 83% 23 of the cases. We found that drug pairs with no evidence of DDIs often had shared 24 metabolizing cytochrome enzyme and cellular transporters, and on average had a 25 significantly higher odds ratio of dosage modifications than drug pairs with established DDIs 26 (p-value < 0.001). Typically, these high odds ratio pairs were prescribed to significantly 27 fewer patients and described significantly less together in the literature. This may rationalize 28 the absence of evidence of DDIs among high odds ratio pairs as they have been harder to 29 identify in smaller-scale studies. 30
, 3 Introduction 31 In most studies, polypharmacy has been investigated using the number of drugs as a burden 32 indicator rather than the burden reflecting the actual combination of specific drugs. 33 Polypharmacy, broadly defined as the concomitant use of multiple medications by a patient, 34 makes prescription a non-trivial task as polypharmacy increases the risk of compromising 35 drug response through drug-drug interactions (DDIs) and by adverse drug reactions (ADRs)136 5. DDIs may influence the efficacy or the risk of toxicity of a drug6, 7, while the risk of ADRs 37 has been shown to increase by the number of medications6. The right drug at the right dose is 38 key in precision medicine to achieve optimal therapeutic effect. Although drugs are being 39 tested repeatedly in clinical trials, preand post-marketing, knowledge of the mechanisms 40 explaining the great variability in patients’ treatment response is limited8. Pharmacogenomics 41 is an emerging and cutting-edge field aiming to understand drug response variability9. The 42 polypharmacy-related dosage changes observed in real world settings are highly relevant in 43 this context as well10. 44 As life expectancy and prescription volumes are increasing, balancing patient safety among 45 the numerous treatment options is becoming a major challenge of modern medicine11-17. 46 Multiple, simultaneous treatments are potentially harmful and difficult to manage and have 47 been associated with in-hospital mortality and readmissions independently of age18-20. DDIs 48 are often studied in small cohorts and typically consist of standard evaluation of potential 49 pharmacokinetic inhibitors and inductors. Frequently, only major CYP450s and drug 50 transporters are tested in relation to likely interacting drugs that are known to be prescribed 51 concomitantly in selected target populations21-23. This is a problem particularly in the in52 hospital setting where the extent of polypharmacy is higher than in the general population24. 53 Large-scale clinical trials of polypharmacy characterising drug response are difficult to carry 54 out. In randomized controlled setups, the volume of data required to reproduce the full 55 spectrum of treatment regimens is often not considered and concomitant treatments are 56 typically addressed in the elderly or in selected phenotypes only25, 26. 57 As electronic health records (EHRs) are broadened into a research resource, studies 58 modelling prescription patterns, including identification of dose dependent ADRs, have been 59 conducted for entire populations and for selected patient groups27-35. However, the focus of 60 these studies has primarily been inappropriate prescription, improvement of patient 61 adherence, reduction of readmissions, ADR detection, and examination of the potential for 62
, 4 deprescription36-38. Yet, evidence for the relationship between polypharmacy and drug 63 response by means of dosage adjustments is sparse, and studies modelling polypharmacy are 64 typically limited to a smaller number of drugs30. Surrogate measures of polypharmacy like 65 the medication burden index39-41 and the comorbidity-polypharmacy score42, 43 are useful, but 66 still ignore many aspects of drug treatment contexts39. 67 In this study, we characterized the complexity of polypharmacy in an in-hospital setting by 68 investigating concomitant drug pairs subject to more frequent dosage adjustments. We 69 analysed in the order of 185 million concomitant treatment episodes derived from more than 70 24 million drug prescriptions. We aimed to provide insight into polypharmacy in a real-world 71 setting and decide whether identification of co-medication pairs subject to dosage 72 adjustments may improve the understanding of drug efficacy, toxicity, and known and 73 possibly unknown DDIs. Finally, we argue that this study can complement 74 pharmacogenomics studies for addressing response variability in patients subject to 75 polypharmacy. 76
, 5 Results 77 In-hospital drug use and polypharmacy 78 In this observational study, we used prescription and admission data from 1,069,873 79 inpatients covering approximately 50% of the Danish population in an eight-year period 80 (2008-2016). Table 1 presents a quantitative summary of the data. The median number of 81 uniquely prescribed drugs per patient was higher among hospitalized patients compared to 82 the general population38. Polypharmacy was prevalent in all age groups and the number of 83 different drugs correlated positively with age (Pearson ρ:0.40, 95% confidence interval 84 CI:0.39-0.40, p-value < 2.2x10-16, Supplementary figure 1). The degree of polypharmacy 85 varied across different admission reasons (primary diagnoses) with a median of drugs ranging 86 from three to eight (Supplementary table 1). Drugs were classified according to the 87 Anatomical Therapeutic Chemical (ATC) Classification System and we used the first and 88 second levels to refer to anatomical and therapeutic drug classes, respectively. Among the 89 24,379,285 inpatient drug prescriptions, 902 distinct drugs were prescribed to at least 50 90 patients; and these drugs were concomitantly administered with up to 857 other drugs. The 91 therapeutic classes that were prescribed to most patients were analgesics (N02), antibiotics 92 (J01), and antithrombotic agents (B01). These were also the ones with the highest number of 93 concomitant medications (ρ:0.93, 95% CI: 0.91-0.96, p-value = 1.70x10-37) (Fig. 1). 94 Condensation of millions of prescriptions into drug dosage patterns 95 In order to study drug dosage alterations comprehensively and provide a framework for the 96 discovery of potentially unknown DDIs, we characterized the 24 million drug prescriptions in 97 terms of treatment episodes, which then comprised the data foundation for a Bayesian 98 hierarchical logistic regression. For each patient, we identified both concomitant treatment 99 episodes (where two or more drugs were administered simultaneously) and monotherapy 100 episodes (where only one drug was administered). Concomitant treatment episodes were 101 structured as pairs of an index drug and a co-medication containing information regarding 102 drug addition, prescribed dosage, and discontinuation (Fig. 2). Only drugs with monotherapy 103 episodes, and drugs that appeared in concomitant treatment episodes of at least 50 patients, 104 were included in the Bayesian hierarchical logistic regression (Supplementary figure 2). 105 From the set of 902 unique drugs, 413 drugs were included in the regression model that 106 estimated likelihoods for dosage adjustments during concomitant treatment episodes. In total, 107
, 6 184,026,179 concomitant treatment episodes were identified among the 413 drugs that 108 combined into 77,494 different co-medication pairs (i.e. an index drug and a co-medication). 109 Each co-medication pair was then characterized by the likelihood of dosage adjustment for 110 the index drug using odd ratios (ORs) and the monotherapy episode of the corresponding 111 index drug as reference. For 309 index drugs. the dosage was more likely to be adjusted with 112 co-medications than during the monotherapy episodes. Among the 77,494 co-medication 113 pairs, 3,993 co-medication pairs had an OR >1 (Supplementary table 2). 114 Among the co-medication pairs, 59 ATC therapeutic subgroups were represented and 56 had 115 at least one drug with ORs >1. All index drugs belonging to antithrombotic agents (B01, e.g. 116 dalteparin, enoxaparin, aspirin, clopidogrel), drugs acting on the renin-angiotensin system 117 (C09, e.g. ramipril, losartan, enalapril, tandolapril), and lipid modifying agents (C10, e.g. 118 atorvastatin, simvastatin, rosuvastatin and ezetimibe) were subject to dosage adjustments 119 during concomitant treatment episodes. Yet, there was no clear pattern in classes generally 120 widely co-administered (e.g. N02, A02, J01, C03, N05) and the proportion of index drugs 121 subject to adjustments with co-medications. Three drug classes with low co-medication 122 burden had no adjustments with OR >1, i.e. vasoprotectives (C05), antibiotics and 123 chemotherapeutics for dermatological use (D06), and nasal preparations (R01) 124 (Supplementary figure 3). 125 The complexity of drug dosage adjustments in a real-world setting 126 Next, we characterized the co-medication pairs in relation to the therapeutic class of the index 127 drug and the anatomical class of the co-medications. While co-medications appeared 128 homogeneously distributed across therapeutic classes (Fig. 3a, left), co-medication pairs with 129 ORs >1 were dominated by nervous system (N), cardiovascular system (C), and alimentary 130 tract and metabolism system (A) co-medications (>60%) (Fig 3a, right). 131 Co-medication pairs where index drugs were classified in the therapeutic subgroups 132 psycholeptics (N05); psychoanaleptics (N06); antiepileptics (N03); other nervous system 133 drugs (N07); drugs for obstructive airway diseases (R03); corticosteroids, dermatological 134 preparations (D07); and antihypertensives (C02) had co-medications of the same anatomical 135 main group in >40% of the co-medication pairs with ORs >1. For the co-medications with 136 ORs >1 and an index drug classified as antiemetics and anti-nauseants (A04), the associated 137 co-medication classes were more evenly distributed. Dosages of immunosuppressant (L04) 138
, ! 7! Supplementary figure 6. Comparison between dosage adjustments and DDI evidence ! Boxplots indicating the odds ratio (OR) and the DDI evidence among co-medication pairs with ORs >1 across the different index therapeutic drug classes. Two-sample Mann-Whitney U-test, where *: p-value ≤ 0.05, **: p-value ≤ 0.01, ***: p-value ≤ 0.001, ****: p-value ≤ 0.0001. Therapeutic groups with less than five observations for each DDI class (DDI/Unknown DDI) are not shown. !
, ! 8! Supplementary figure 7. Comparison between patient volume and DDI evidence ! Boxplots indicating the prevalence (number of patients) and the DDI evidence among comedication pairs with ORs >1 across the different index therapeutic drug classes. Two-sample Mann-Whitney U-test, where *: p-value ≤ 0.05, **: p-value ≤ 0.01, ***: p-value ≤ 0.001, ****: p-value ≤ 0.0001. Therapeutic groups with less than five observations for each DDI class (DDI/Unknown DDI) are not shown. OR=Odds Ratio. ! !
, ! 9! Supplementary figure 8. Comparison between literature and DDI evidence ! Boxplots indicating the number of publications with co-mentioning and the DDI evidence among co-medication pairs with ORs >1 across the different index therapeutic drug classes. Two-sample Mann-Whitney U-test, where *: p-value ≤ 0.05, **: p-value ≤ 0.01, ***: p-value ≤ 0.001, ****: p-value ≤ 0.0001. Therapeutic groups with less than five observations for each DDI class (DDI/Unknown DDI) are not shown. OR=Odds Ratio. !
, ! 10! Supplementary figure 9. Average daily dose ! Calculation of prescribed dosage as the prescribed average daily dose (ADD). A prescription is a health-care program implemented by a physician or other qualified health care practitioner in the form of instructions that govern the plan of care for an individual patient.The term often refers to a health care provider’s written authorization for a patient to purchase a prescription drug from a pharmacist. Prescriptions can be classified into four different groups: (1) one-time prescriptions; (2) scheduled prescriptions; (3) Pro Necessitata (PN) or ‘as needed’ prescriptions; and (4) Variable dosage (VAO) prescriptions. One-time prescriptions consist on the administration of a single dose in a single day. As this type of prescriptions do not allow the possibility to study longitudinal dosage adjustments, we did not consider them for this study. PN consist of drugs only administered if needed and they are indicated with a maximum number of administrations/day that can be taken. An example of a PN prescription is the painkiller paracetamol or other analgesics indicated for the treatment of pain if the patient requires so after e.g. surgery. VAO drugs refer to drugs whose dosage is variable on a biochemical value or another physiological constant of the patient at the time of each administration. For example, this is the case for insulin, whose dosage depends on the blood glucose levels. For PN and VAO drug prescriptions, the ADD was coded categorically as ‘PN’ or ‘VAO’. Henceforth, quantitative ADD was calculated only for scheduled prescription types. Scheduled prescriptions consist of drug indications having an interval or cycle with a defined time pattern for each administration. Among this type, we can find many variations with different intervals and patterns. The figures a and b exemplify two different examples: a Dosing regimen for acyclovir (J05AB01), a drug used to treat infections caused by certain types of viruses (e.g. cold sores, shingles, chickenpox). The prescription regimen is indicated with a daily interval (1 day), a frequency of three administrations following the time pattern at 8 am, 2 pm and 10 pm. The prescribed dose
, ! 11! indicated at each administration is of 200 mg. The final calculated ADD for acyclovir is of 600 mg (see Equation 1). b A more complex regimen is presented with the medication Levothyroxine sodium (H03AA01), a thyroid hormone medication that is used to treat hypothyroidism and other hormonal conditions. The treatment plan for this drug follows a weekly interval (7 days) and different prescribed doses for every other day. The prescribed dose is 50 micrograms (mcg) for days 1,3,5 and 100 mcg for days 0,2,4,6. The final calculated ADD is 78.57 mcg, which is the sum of the two previous prescribed doses in the same weekly interval. !
, ! 12! Supplementary figure 10. Treatment episodes ! Construction of concomitant and monotherapy treatment episodes. Patients can have a drug prescribed multiple times during an admission. We constructed concomitant treatment episodes as the intervals of time when two different drug prescriptions were contemporaneously active. Monotherapy treatment episodes were used as the reference treatment episodes as the intervals of time when a drug was not concomitantly given with any other drug. !
, ! 13! Supplementary figure 11. Literature co-mentions diagram ! ! Diagram of the full-text article and abstract corpus generation for the text mining of drug-drug interactions. DTU DTM Corpus (~15 million articles), PubMed (~27 million articles) and MEDLINE were used for the collection of full-text articles and abstracts. !
, ! 14! Tables Supplementary table 1: Admission primary diagnosis ICD-10 chapter Chapter name No. unique drugs, median (IQR) 1 Certain infectious and parasitic diseases 8 (4-14) 2 Neoplasms 8 (5-12) 3 Diseases of the blood and blood forming organs and certain disorders involving the immune mechanism 7 (4-11) 4 Endocrine, nutritional and metabolic diseases 8 (4-12) 5 Mental and behavioural disorders 6 (4-9) 6 Diseases of the nervous system 5 (3-9) 7 Diseases of the eye and adnexa 4 (2-7) 8 Diseases of the ear and mastoid process 3 (2-5) 9 Diseases of the circulatory system 8 (4-12) 10 Diseases of the respiratory system 8 (4-13) 11 Diseases of the digestive system 6 (4-10) 12 Diseases of the skin and subcutaneous tissue 4 (2-9) 13 Diseases of the musculoskeletal system and connective tissue 8 (4-11) 14 Diseases of the genitourinary system 6 (3-10) 15 Pregnancy, childbirth and the puerperium 3 (2-5) 16 Certain conditions originating in the perinatal period 2 (1-5) 17 Congenital malformations, deformations and chromosomal abnormalities 4 (2-7) 18 Symptoms, signs and abnormal clinical and laboratory findings, not elsewhere classified 5 (2-9) 19 Injury, poisoning and certain other consequences of external causes 7 (3-12) 20 External causes of morbidity and mortality 5 (3-9) 21 Factors influencing heath status and contact with health services 5 (1-9) !
, ! 15! Supplementary table 2: Co-medication pairs with associated dosage adjustment List of co-medication pairs with ORs >1 for dosage adjustment. OR=Odds Ratio. Supplementary table 3: DDI evidence List of co-medication pairs with ORs >1 for dosage adjustment and their DDI evidence. OR=Odds Ratio. Supplementary table 4: Shared CYPs and transporters List of co-medication pairs with ORs >1 for dosage adjustment and their pharmacokinetic activities: shared metabolism and transport, and associated variants. OR=Odds Ratio. Supplementary table 5: Literature co-mentions List of co-medication pairs with ORs >1 for dosage adjustment, DDI evidence and comentioning in the literature. OR=Odds Ratio. Supplementary table 6: One-time drugs List of excluded index drugs for which the prescription type was ‘one-time’ or ‘singleadministration’ in more than 70% of the cases.
D|Appendix 253
, Supplementary figure 4 | Longitudinal trajectories formed by the same ATC group. a, trajectories formed by drugs from the same anatomical ATC group were included and then horizontally separated by average time between prescription redemptions. Each group of trajectories contains edges that separate the different prescription redemptions, whose height represent the number of trajectories that go from length 2 to length 3, to length N. Each group of trajectories is ordered from shortest to longest (top to bottom). *Refer to Figure 1 for ATC group legend. b, Number of individuals redeeming the subgroup indicated on the left side (node size) and how many of these patients redeem another drug used in hypertension (C02, C03, C07, C08, C09) posteriorly (ordered by chemical subgroup in ATC).
, Supplementary figure 5 | Individual grouping for Poisson modelling change over time. Each group of individuals, separated by sex and year of birth are dynamically allocated to different groups at each time point, depending on their prescription redemption at the time. That process is repeated for each stratum of the data (i.e., for each sex, year of birth and time of prescription), for each drug pair. As an example, women born in 1991 are used in the figure. They enter the study in 1995 (t1) for prescription redemption pair (P1P2). Some of these individuals redeemed both prescriptions at t1, so they are directly counted in P12 (P12 = has redeemed P1 and P2; P10 = has redeemed P1 but not P2; P02 = has redeemed P2 but not P1; P00 = has redeemed neither P1, nor P2). This individual is then censored from the rest of the years in the model. Other individuals in the same stratum at t1 will not have redeemed any of the drugs in the pairs under study, P1 or P2, hence they will be counted as P00. At each time, the position of the different individuals might change, as they might have redeemed one or both prescriptions, or they might have left the study (emigration or death).
, Supplementary figure 6 | Period of inclusion, 1995-2019, for Cox proportional hazard regression model cohort. Individuals who redeemed the prescription before 1995 are excluded from the cohort (left-truncated) and patients who leave the study, either due to end of window of study (2019), emigration or other are censored (right censored). The timeframe is the period for which prescription registry has information (1995-2019), leaving those patients who redeemed the prescription before the beginning of the registry truncated out of the model (Supplementary Fig. 2). Charlson Comorbidity Index (CCI) 20 was calculated using the ICD-10 coded disease data from the Danish National Patient Registry. Age, sex and CCI were added as covariates in the model and an interaction with them was included if the hazard proportionality assumption was violated – using a test based on Schoenfeld residuals with a p-value lower than 0.0001.
, Supplementary table 1 | Cohort characteristics. Male Female Number of patients (%) 3 586 032 (49 %) 3 669 887 (51%) Median age 1995, years 30 30 Mean age 1995, years 31·2 32·9 Range age 1995, years 0-106 0-110 Median age 2019, years 44 47 Mean age 2019, years 43·9 47·1 Range age 2019, years 0-110 0-110 Prescription redemption Median 11 17 Mean 13·9 19·1 SD 10·7 13·4 Range 1-164 1-129
, Supplementary table 2 | Prescription distribution over the Anatomical Therapeutic Chemical Classification System. Male Female Age 0-14 15-44 45-64 65-84 >85 0-14 15-44 45-64 65-84 >85 ATC mean (SD) A 1·3 (0·6) 1·7 (1·1) 2·1 (1·7) 2·5 (1·9) 2·3 (1·6) 1·3 (0·6) 2·0 (1·4) 2·3 (1·8) 2·9 (2·1) 2·7 (1·7) B 1·0 (0·3) 1·1 (0·4) 1·2 (0·5) 1·4 (0·7) 1·4 (0·6) 1·0 (0·2) 1·2 (0·5) 1·2 (0·5) 1·4 (0·7) 1·4 (0·6) C 1·3 (0·7) 1·5 (1·1) 2·8 (2·1) 3·2 (2·2) 2·2 (1·5) 1·3 (0·7) 1·5 (1·0) 2·6 (1·9) 3·2 (2·3) 2·4 (1·6) D 1·9 (1·3) 2·3 (1·7) 2·2 (1·6) 2·2 (1·6) 1·9 (1·3) 2·0 (1·3) 2·7 (2·0) 2·3 (1·6) 2·3 (1·6) 1·9 (1·3) G 1·0 (0·2) 1·1 (0·3) 1·2 (0·5) 1·4 (0·6) 1·2 (0·5) 1·1 (0·3) 2·1 (1·3) 1·7 (1·0) 1·4 (0·7) 1·2 (0·5) H 1·0 (0·1) 1·0 (0·2) 1·0 (0·2) 1·1 (0·3) 1·0 (0·2) 1·0 (0·2) 1·1 (0·4) 1·1 (0·4) 1·1 (0·4) 1·0 (0·3) J 1·0 (1·0) 2·2 (1·3) 2·1 (1·4) 2·5 (1·6) 2·2 (1·4) 1·9 (0·2) 3·3 (1·9) 2·6 (1·7) 2·7 (1·8) 2·4 (1·6) L 1·1 (0·3) 1·1 (0·3) 1·1 (0·3) 1·1 (0·3) 1·0 (1·8) 1·1 (0·2) 1·0 (0·2) 1·1 (0·3) 1·0 (0·2) 1·0 (0·2) M 1·1 (0·3) 1·5 (0·7) 1·7 (0·9) 1·7 (0·9) 1·4 (0·7) 1·1 (0·3) 1·6 (0·9) 1·8 (1·1) 1·9 (1·2) 1·5 (0·8) N 1·3 (0·8) 2·5 (2·3) 2·7 (2·3) 3·0 (2·3) 2·8 (1·9) 1·3 (0·8) 2·7 (2·4) 3·0 (2·5) 3·4 (2·5) 3·1 (2·1) P 1·0 (0·2) 1·2 (0·5) 1·1 (0·4) 1·1 (0·3) 1·0 (0·2) 1·0 (0·2) 1·3 (0·5) 1·2 (0·5) 1·1 (0·3) 1·0 (0·2) R 2·0 (1·3) 2·0 (1·3) 1·9 (1·3) 2·2 (1·7) 1·7 (1·1) 1·8 (1·2) 2·4 (1·7) 2·2 (1·7) 2·3 (1·8) 1·7 (1·5) S 1·5 (0·8) 1·6 (1·0) 1·7 (1·1) 2·1 (1·5) 1·7 (1·1) 1·4 (0·8) 1·7 (1·0) 1·8 (1·3) 2·2 (1·6) 1·9 (1·3) V 1·0 (0·1) 1·0 (0·1) 1·0 (0·1) 1·0 (0·1) 1·0 (0·2) 1·0 (0·1) 1·0 (0·1) 1·0 (0·1) 1·0 (0·1) 1·0 (0·1) Range A 1-13 1-23 1-21 1-21 1-14 1-18 1-23 1-22 1-21 1-18 B 1-6 1-7 1-7 1-7 1-6 1-6 1-8 1-8 1-8 1-6 C 1-9 1-19 1-20 1-21 1-15 1-9 1-18 1-20 1-21 1-15 D 1-16 1-21 1-21 1-19 1-15 1-18 1-21 1-19 1-18 1-15 G 1-4 1-7 1-11 1-7 1-5 1-5 1-14 1-11 1-9 1-7 H 1-4 1-5 1-5 1-5 1-4 1-5 1-7 1-6 1-5 1-4 J 1-17 1-20 1-16 1-14 1-11 1-17 1-21 1-16 1-15 1-13 L 1-4 1-5 1-6 1-5 1-3 1-5 1-5 1-5 1-5 1-4 M 1-6 1-9 1-11 1-10 1-7 1-6 1-10 1-11 1-11 1-9 N 1-19 1-33 1-34 1-27 1-18 1-14 1-36 1-30 1-28 1-19 P 1-5 1-6 1-7 1-5 1-3 1-5 1-7 1-6 1-5 1-4 R 1-14 1-17 1-17 1-17 1-13 1-14 1-19 1-19 1-18 1-13 S 1-13 1-16 1-16 1-17 1-13 1-13 1-17 1-16 1-15 1-14 V 1-2 1-3 1-3 1-3 1-2 1-2 1-3 1-3 1-2 1-2
, Supplementary table 3 | Relative risk for pairs in figure 2 ATC 1 ATC 12 Number patients RR 95% CI P-value N02AA N02AB 138 110 2.76 2.74 – 4.13 2.70e-08 R03CC R03BA 226 898 2.72 1.86 - 3.99 2.68e-07 C09AA C09BA 164 680 2.79 1.93 - 4.03 5.22e-08 C07AB B01AA 125 137 3.17 2.17 - 4.63 2.72e-09 C07AB C01AA 103 043 3.55 2.31 - 5.47 9.07e-09 A03FA N02AB 102 783 2.77 1.89 - 4.05 1.42e-07 N06AB N05AF 116 605 2.86 2.02 - 4.0 4.11e-09 B01AC C01AA 135 262 2.88 1.89 - 4.38 8.06e-07 N05BA N02AG 118 223 2.83 2.09 - 3.83 1.83e-11 C09AA C01DA 137 306 3.14 2.1 - 4.62 8.19e-09 J01CE D07BB 195 989 2.81 2.19 - 3.61 5.51e-16 M01AB N02AG 118 272 2.80 2.0 - 3.76 5.61e-12 M01AC M01AB 121 302 2.73 2.08 - 3.58 3.70e-13 G03AA J01EB 391 118 3.05 1.93 - 4.82 1.87e-06 J01EB G01AF 189 887 3.28 2.33 - 4.63 1.19e-11 A02BA N05BA 138 066 2.89 2.18 - 3.83 1.16e-13 M01AX C09AA 105 249 2.88 2.06 - 4.04 8.30e-10 C10AA C03CA 218 656 2.74 1.94 - 3.87 1.17e-08 G03AA G01AF 279 484 3.98 2.42 - 6.55 5.27e-08 A02BA A03FA 149 396 2.74 2.06 - 3.64 4.28e-12 D07BB J01FA 131 233 3.13 2.42 - 4.04 1.93e-18 C10AA A12BA 232 191 2.85 1.98 - 4.10 1.74e-08 S01GA R06AE 113 185 2.74 2.07 - 3.61 1.27e-12 G03AA D06BB 129 086 3.12 2.00 - 4.86 5.17e-07 M01AX B01AC 113 398 2.98 2.11 - 4.21 5.70e-10 S01GA S01GX 125 768 2.75 2.06 - 3.68 1.01e-11 C07AB C03DA 121 672 2.90 1.99 - 4.23 3.44e-08 D07BB D07AC 112 418 3.43 2.59 - 4.54 7.00e-18 A08AA M01AB 185 846 2.97 2.28 - 3.87 7.04e-16 A02BA R05FA 126 611 2.87 2.16 - 3.81 3.85e-13 G03AA J02AC 414 591 2.82 1.76 - 4.53 1.66e-05 D07BB S01AA 123 584 2.78 2.14 - 3.62 2.15e-14 M01AX C08CA 111 180 2.95 2.07 - 4.20 2.48e-09 D07BB D07AB 111 589 3.33 2.50 - 4.45 3.02e-16 D07BB D01AC 120 234 3.26 2.47 - 4.29 4.61e-17 M01AX C03CA 102 801 3.10 2.17 - 4.42 4.24e-10 C03AB R05CB 102 408 2.78 1.92 - 4.03 5.53e-08 D07BB A10BK 3 349 0.002 1.20e-03 - 3.15e-03 4.44e-142 D07CB A10BK 1 538 2.27e-03 1.24e-03 - 4.17e-03 2.46e-86 M03BA A10BK 1 264 2.32e-03 1.24e-03 - 4.35e-03 3.30e-80 S03BA B01AF 1 505 3.24e-03 1.80e-03 - 5.82e-03 8.93e-82 S03BA N05CH 1 789 3.25e-03 2.04e-03 - 5.19e-03 1.03e-127 G03CB B01AF 6 461 3.39e-03 1.74e-03 - 6.62e-03 2.50e-62 S03BA N05CH 1 789 3.25e-03 2.04e-03 - 5.19e-03 1.03e-127 G03CB B01AF 6 461 3.39e-03 1.74e-03 - 6.62e-03 2.50e-62 G03CB N05CH 6 023 3.71e-03 1.91e-03 - 7.18e-03 9.73e-62 C03CB A10BH 449 6.49e-01 2.74e-01 - 1.53 3.24e-01 G03CB R02AX 1 121 4.01e-03 2.02e-03 - 7.94e-03 2.24e-56 B01AF N06DX 1 344 4.15e-01 3.74e-01 - 4.60e-01 1.38e-62 A11EA G03FA 1 009 1.06e+01 8.51 – 13.1 6.12e-102 G01AG S01CA 1 696 1.06e+01 7.10 - 15.7 3.04e-31 B03AE A06AD 1 503 1.06e+01 6.51 – 17.1 1.40e-21
, S01KA M01AH 1 012 1.08e+01 7.57 – 15.4 1.49e-39 A11EA G04CA 1 052 1.09e+01 8.58 – 14.0 3.49e-82 A11EA G03CB 1 030 1.10e+01 9.47 - 12.8 8.01e-217 S01KA A02AA 1 081 1.12e+01 7.81 – 16.0 7.48e-40 S01KA A06AD 1 745 1.15e+01 7.85 – 16.7 1.40e-36 A10BB N02BE 97 371 1.60e+03 1.52e+03 - 1.68e+03 2.22e-308 A10BF N02BE 3 717 1.66e+03 1.61e+03 - 1.72e+03 2.22e-308 A10BB N02AX 68 483 1.67e+03 1.61e+03 - 1.75e+03 2.22e-308 A10BB N03AF 4 148 1.70e+03 1.65e+03 - 1.75e+03 2.22e-308 A10BB N02AJ 22 433 1.72e+03 1.67e+03 - 1.77e+03 2.22e-308 A10BB N03AA 1 403 1.74e+03 1.68e+03 - 1.80e+03 2.22e-308 A10BB N02AA 52 922 1.74e+03 1.68e+03 - 1.81e+03 2.22e-308 A10BB N03AG 2 891 1.75e+03 1.69e+03 - 1.80e+03 2.22e-308 A10BB N02AG 14 877 1.76e+03 1.67e+03 - 1.85e+03 2.22e-308 A10BB N02AE 12 015 1.79e+03 1.72e+03 - 1.87e+03 2.22e-308 A10BB N02AB 17 553 1.79e+03 1.71e+03 - 1.88e+03 2.22e-308 A10BF N03AX 1 178 1.86e+03 1.82e+03 - 1.90e+03 2.22e-308 A10BH A10BK 13 699 4.05e-01 3.24e-01 - 5.06e-01 2.34e-15 A10BD A10BK 12 727 4.06e-01 3.15e-01 - 5.24e-01 3.95e-12 A10BJ A10BK 13 156 4.19e-01 3.35e-01 - 5.23e-01 1.93e-14 A10BH A10BJ 13 537 4.55e-01 3.73e-01 - 5.56e-01 1.34e-14 N06DX N05AD 2 607 4.72e-01 3.90e-01 - 5.70e-01 8.22e-15 A10BH A10BD 8 676 4.80e-01 3.94e-01 - 5.85e-01 4.16e-13 A10BD A10BJ 11 035 5.03e-01 4.03e-01 - 6.28e-01 1.24e-09 B01AE B01AF 11 838 5.04e-01 4.23e-01 - 6.00e-01 1.41e-14 A10BG A10BH 1 436 5.25e-01 4.01e-01 - 68.6 2.31e-06 N06DX N02AB 4 151 5.30e-01 4.48e-01 - 62.7 1.10e-13 A11EA A01AB 2 882 9.20 6.88 - 12.3 1.93e-50 A11EA A11CA 1 065 9.44 6.95 - 12.8 9.89e-47 G01AG G03DA 1 351 9.57 6.64 – 13.8 8.72e-34 A11EA A10BA 1 331 9.73 6.71 - 14.1 2.88e-33 S01KA S01BC 1 116 9.89 6.56 – 14.91 6.24e-28 S01KA S01XA 1 175 10.28 7.71 - 13.7 9.65e-57 A11EA A12AX 1 857 10.5 7.97 – 13.85 2.62e-62 G01AG G02BA 2 112 10.70 6.62 – 17.5 8.48e-22 S01KA S01CA 2 092 10.8 7.11 – 16.4 8.13e-29 B03AE B03BA 1 103 9.87 6.78 – 14.37 6.95e-33
, Supplementary table 4 | Stratification of patients by prescription trajectories ACE treatment with no change ARB treatment with no change ACE treatment with change to ARB ARB treatment with change to ACE Number of patients 549 436 (49.3) 287 488 (25.8) 229 216 (20.6) 48 054 (4.3) Age, years Mean (SD) 63.3 (13·9) 61.7 (13·3) 60.7 (12·3) 61.0 (13·2) Median (IQR) 63.7 (20.0) 62.0 (18.9) 61.0 (17.4) 61.2 (18.9) Sex Male 257 099 (46.8) 153 527 (53.4) 122 549 (53.5) 24 484 (51.0) Female 292 337 (53.2) 133 961 (46.6) 106 667 (46.5) 23 570 (49.0) Charlson comorbidities Myocardial infarction 74 244 (13.5) 14 962 (5.2) 23 253 (10.1) 5 485 (11.4) Congestive heart failure 90 247(16.4) 13 696 (4.8) 26 115 (11.4) 6 668 (13.9) Peripheral vascular diseases 59 155 (10.7) 18 643 (6.5) 20 914 (9.1) 5 666 (11.8) Cerebrovascular disease 102 441 (18.6) 37 758 (13.1) 35 109 (15.3) 9 894 (20.6) Dementia 29 033 (4.9) 8 387 (2.9) 6 181 (2.7) 2 452 (5.1) Chronic obstructive pulmonary disease 77 099 (14.0) 31 317 (10.9) 28 892 (12.6) 7 029 (14.6) Rheumatoid disease 21 897 (3.9) 10 172 (3·5) 9 500 (4.1) 2 135 (4.4) Peptic ulcer disease 36 803 (6.7) 13 572 (4.7) 12 545 (5.4) 3 315 (6.9) Mild liver disease 12 827 (2.3) 5 527 (1.9) 4 739 (2.0) 1 175 (2.4) Diabetes without complications 90 186 (16.4) 27 143 (9.4) 37 033 (16.1) 8 391 (17·4) Diabetes with complications 31 608 (5·7) 7 560 (2.6) 12 894 (5.6) 2 806 (5.8) Hemiplegia or paraplegia 3 633 (0.7) 1 433 (0.5) 1 064 (0·4) 298 (0.6) Renal disease 29 369 (5.3) 8 112 (2.8) 11 624 (5.1) 3 236 (6.7) Cancer 102 249 (18.6) 44 873 (15.6) 39 402 (17.1) 9 362 (19.5) Moderate or severe liver disease 4 272 (0·7) 1 438 (0.5) 1 174 (0·5) 341 (0.7) Metastatic solid tumour 20 741 (3.7) 8 219 (2.8) 6 841 (2·9) 1 774 (3.7) AIDS/HIV 295 (0·05) 109 (0.04) 85 (0·03) 23 (0·04) Data are n (%) unless otherwise specified
, Supplementary table 5 | Population continent of origin Percentage Europe 92.50 Asia 2.93 Middle East 1.94 Africa 1.24 North America 0.81 South and Central America 0.39 Oceania 0.16 Not stated 0.03
E|Appendix 269