A section identification tool: Towards HL7 CDA/CCR standardization in Spanish discharge summaries
Abstract
We gratefully acknowledge the support of NVIDIA Corporation, United States with the donation of the Titan X Pascal GPU used for this research. This work was partially funded by the Spanish Ministry of Science and Innovation (DOTT-HEALTH/PAT-MED PID2019-106942RB-C31), the European Commission (FEDER), Spain, the Basque Government, Spain (IXA IT-1343-19), and the EU ERA-Net CHIST-ERA and the Spanish Research Agency (ANTIDOTE PCI2020-120717-2).
Full text
A Section Identification Tool: towards HL7 CDA/CCR Standardization in Spanish Discharge Summaries Iakes Goenagaa, Xabier Lahuertaa, Aitziber Atutxab, Koldo Gojenolab HiTZ Basque Center for Language Technology http: // www. hitz. eus University of the Basque Country (UPV/EHU), Spain aFaculty of Computer Science, PºManuel Lardizabal, 1 — 20018 Donostia-San Sebasti´an bSchool of Engineering, Paseo Rafael Moreno Pitxitxi, 3 — 48013 Bilbao Abstract Background. Nowadays, with the digitalization of healthcare systems, huge amounts of clinical narratives are available. However, despite the wealth of information contained in them, interoperability and extraction of relevant information from documents remains a challenge. Objective. This work presents an approach towards automatically standardizing Spanish Electronic Discharge Summaries (EDS) following the HL7 Clinical Document Architecture. We address the task of section annotation in EDSs written in Spanish, experimenting with three different approaches, with the aim of boosting interoperability across healthcare systems and hospitals. Methods. The paper presents three different methods, ranging from a knowledge-based solution by means of manually constructed rules to superEmail addresses: [email protected] (Iakes Goenaga), [email protected] (Xabier Lahuerta), [email protected] (Aitziber Atutxa), [email protected] (Koldo Gojenola) This is the accepted manuscript of the article that appeared in final form in Journal of Biomedical Informatics 121 : (2021) // Article ID 103875, which has been published in final form at https://doi.org/10.1016/j.jbi.2021.103875. © 2021 Elsevier under CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
vised Machine Learning approaches, using state of the art algorithms like the Perceptron and transfer learning-based Neural Networks. Results. The paper presents a detailed evaluation of the three approaches on two different hospitals. Overall, the best system obtains a 93.03% Fscore for section identification. It is worth mentioning that this result is not completely homogeneous over all section types and hospitals, showing that cross-hospital variability in certain sections is bigger than in others. Conclusions. As a main result, this work proves the feasibility of accurate automatic detection and standardization of section blocks in clinical narratives, opening the way to interoperability and secondary use of clinical data. Keywords: Section Identification, Interoperability, Electronic Discharge Summaries, HL7 Clinical Document Architecture 1. Introduction1 The outstanding advancement of Machine Learning (ML) technologies2 (e.g., Deep Learning) enable us to more efficiently harness the large amounts3 of data collected through healthcare processes such as clinical narratives in4 electronic health records (EHR) as well as electronic discharge summaries5 (EDS). EHRs contain a lifetime record of the patient’s complete medical6 history, diagnoses and treatment, medications, allergies and immunizations,7 as well as radiology images and laboratory results [1]. EDSs are an essential8 document to communicate patient journey and care planning regarding an9 hospitalization episode to the next practitioner [2]1. In 2016 the proportion10 of primary care practices using electronic clinical records was about 80% on11 1Some authors use these two terms interchangeably. 2
average across 15 EU countries [3], and in 2020 in the US the percentage12 is of 96% [4]. Digitalization of healthcare systems is contributing to the13 improvement of clinical and translational studies, and interoperability and14 information exchange between healthcare systems is more necessary than15 ever. For that reason, public policies and recommendations are pushing onto16 that way [5, 6, 7].17 There is an increasing interest for integrating heterogeneous health infor-18 mation for different reasons: to facilitate the cross-border interoperability of19 information among healthcare systems, federal states and countries to ensure20 that citizens can securely access and exchange their health data wherever they21 are, and also to make digital health information more usable to the bedside22 and beyond [5, 6]. Several standards as openEHR [8], HL7-FHIR [9], HL723 CDA/CCR [10] are examples of this standardization effort.24 However, despite the wealth of information contained in the clinical nar-25 ratives, interoperability and extraction of relevant information from docu-26 ments remains a challenge. Although the aforementioned standards exist, so27 far they have not been widely adopted, and even if so, the healthcare system28 at large still has a huge amount of untapped legacy clinical text.29 Healthcare systems provide guidelines for writing clinical documents,30 which for operative reasons typically follow some minimal principles to ensure31 the optimal interactions between health professionals and patients like SOAP32 (Subjective, Objective, Assessment, Plan), or APIE (Assessment, Plan, Im-33 plementation, and Evaluation) [11, 12]. Some systems assume that these34 principles are best reflected by using free text, due to flexibility to express35 anything that the health-care providers need to record. On the opposite36 3
extreme, some impose structured or semi-structured clinical documents in37 sections, where each section is a main block of information. In all cases,38 the automated processing of clinical texts is hampered by ambiguity, lexical39 variety, use of abbreviations, errors due to mistakes, redundancies, etc.40 Under this scenario, this work presents a first approach towards auto-41 matically standardizing Spanish EDSs following the HL7 Clinical Document42 Architecture (CDA) R2 template for Discharge Summaries [10] for both help-43 ing interoperability and secondary use of Electronic Discharge Summaries.44 The HL7 CDA R2 template contains a set of clinically relevant sections,45 and part of this standardization task is known as Section Identification. It is46 defined in [13] as detecting the boundaries of text sections and adding seman-47 tic annotations. They define a section as a text segment that groups together48 consecutive clauses, phrases or sentences that share the description of one49 dimension of a patient, patient’s interaction or clinical findings. A section50 can be marked explicitly, through structural demarcations (headings or sub-51 headings), or it can exist implicitly. The main assumption for making this52 identification is that unstructured texts have an explicit or implicit structure.53 Besides its relevance in terms of standarization and interoperability, sec-54 tion identification provides a deeper understanding of EDSs, for instance, by55 recognizing the section in which a medical entity is located. The same med-56 ical condition found in the “past personal medical history” or in the “family57 medical history” section might lead to different conclusions. Several works58 on secondary use of EHRs and EDSs have shown that section identification59 can be helpful for a variety of tasks [14] such as entity recognition [15], co-60 hort retrieval [16] and temporal relation extraction [17], and can help in most61 4
automatic medical processing tasks, as ICD-10 coding [18, 19, 20, 21]. This62 issue is rapidly becoming an important topic in both academia and industry.63 9027431 XX-XX-XXXX 66 años. VARON. MC: REFERENCIADO EN EL INFORME. INFORME AL ALTA : Paciente de 66 años. No alergias medicamentosas conocidas. A. PERSONALES: Enfermedad de Crohn diagnosticada en 1997 con afectación de íleon terminal (A3L1B2) por cuadros suboclusivos resueltos con enfermedad de íleon terminal asociada a mesenteritis fibrosa. Artrosis dorso-lumbar. Cirugía de hernia inguinal. Ci Tto: Dacortin 5: 1-0-0; Pariet 20: 1-0-0; Pentasa:1-1-1; Kilor 0-1-0, Clinutren: 2/día. E. ACTUAL: Acude a Urgencias por dolor abdominal generalizado con febrícula, sin tiritona, sin naúseas ni vómitos. Sin alteración del ritmo intestinal. Con pauta descendente de corticoides, después del último ingreso por cuadro suboclusivo. EXPL. FÍSICA: Paciente consciente, orientado, colaborador. Buena coloración de piel y mucosas. Cuello: no adenopatías cervicales. AC: rítmica sin soplos. No roncus ni crepitantes. Abdomen: distendido, timpánico. Peristaltismo ausente. Blumberg negativo. EEII: no edemas maleolares. PPP. RX ABDOMEN: Sugestivo de suboclusión intestinal. Se objetivan dos asas de delgado con niveles hidroaéreos incluso en la cámara gástrica. ANALÍTICA AL INGRESO. Urea, Creatinina, GPT, Amilasa dentro de límites normales. Leucocitos14.400. Segmentados 80 %. TP 100 %. Plaquetas 470.000. Hb 13.5. Hto 41.7 %. PCR 14.2. ANALÍTICA AL ALTA: GOT, GPT, Gamma GT, FA, Bilirrubina total, Amilasa, LDH, Colesterol total, Triglicéridos, Na, K dentro de límites normales. Alfa1 antitripsina 169, Albúmina 2.4. PCR 11.6. Fe 18. Transferrina 241. IS 5.3 %. Ferritina 24. Vit B12 247. A fólico 6.3. Hb 11.1. Hto 34.4 %. VCM 78.8. Plaquetas 367.000. Segmentados 75 %. VCM 7. EVOLUCIÓN Y PROCEDIMIENTOS: Se trata de paciente con enfermedad de Crohn, con afectación ileal y mesenteritis retractil que ingresa por cuadro suboclusivo, instaurándose tto. corticoideo siendo dado de alta con disminución progresiva de dicho tratamiento. Estando en tto. con 5 mg de Dacortin, ingresa de nuevo con cuadro suboclusivo. Se indica colocación de sonda nasogástrica para aspiración intermitente rechazando el paciente. Se inicia el tratamiento corticoideo iv a dosis plenas, mejorando la clínica del paciente. DIAGNÓSTICO: - CUADRO SUBOCLUSIVO INTESTINAL POR ENFERMEDAD DE CROHN CON AFECTACIÓN ILEAL (A3L1B2) Y MESENTERITIS RETRACTI TRATAMIENTO: - Dacortin 60: 1-0-0 durante 1 semana bajando 10 mg cada 10 días hasta 10 mg que mantendrá 15 días más y luego 5 mg 15 días más y suspender. H M H C I E E C E V T D Figure 1: Example EDS and its sections (H: Heading, MH: Medical History, CI: Current Illness, E: Exploration, EC: Complementary Exploration, EV: Evolution, D: Diagnosis, T: Treatment). 5
Given the difficulty in accurately extracting data from text, most non-64 research use of EHR and EDS data rely only on structured data. However,65 clinical notes contain highly valuable information not found in strictly struc-66 tured fields and, moreover, they give access to volumes of data that are67 orders of magnitude bigger and, consequently, improving retrieval accuracy68 from text would have great value.69 In this paper, we will explore the task of section annotation in EDSs70 written in Spanish (see Figure 1). We will experiment with three different71 approaches, ranging from a knowledge-based solution by means of manually72 constructed rules to supervised Machine Learning approaches, including the73 structured Perceptron algorithm and Deep Neural Networks. The paper will74 present a detailed evaluation of the three approaches and, as a main result,75 will prove the feasibility of automatically detecting section blocks in EDSs.76 The main contributions of this work are:77 •We describe an annotation format for EDSs that defines the section78 structure of a document. We have evaluated its feasibility annotating79 a dataset comprised of 300 documents and have measured a high inter-80 annotator agreement.81 •We implement three different approaches to automatic section identifi-82 cation, including a rule-based method, the Perceptron online learning83 algorithm and Neural Networks.84 •We conduct exhaustive experiments to explore the contribution of each85 method, also giving a detailed analysis of the strengths and weaknesses86 of the proposed approaches.87 6
The remainder of this paper is structured as follows. Section 2 discusses88 related work. The resources and corpus are presented in Section 3. Section 489 sketches the main results, while Section 5 provides an analysis of the results90 including a comparison of the different approaches as well as an estimation91 of the system’s ability to generalize across hospital settings and a qualitative92 evaluation of the encountered errors. Finally, Section 6 summarizes the main93 conclusions and future work.94 2. Background95 Pomares-Quimbaya et al. [13] reviewed several studies on clinical section96 identification, which varied on the kind of narrative, the type of section, and97 the application. The paper examines the characteristics of systems using a98 strategy for section identification, the methods used to identify implicit or99 explicit sections with different degrees of success, and the main application100 scenarios and contexts that have been used with good performance. From the101 technical point of view, the methods were classified into rule-based methods102 (59%), machine learning methods (22%) and a combination of both (19%).103 According to the authors, hybrid methods showed the best performance. 46%104 of the studies were able to identify explicit (using headings) and implicit105 sections. Regarding the language of application, most of the works (78%)106 were intended for English texts.107 Arnold et al. [22] present SECTOR, a model to segment documents into108 sections, under the hypothesis that topics, learned in an unsupervised way,109 characterize semantically coherent text segments (sections). Their deep neu-110 ral network architecture learns a latent topic embedding over a document, in111 7
order to classify local topics and to segment a document at topic shifts. They112 report a 56.7% F-score for segmentation and classification in the domain of113 diseases. Although the approach seems promising, its main inconvenient for114 our task is that, as topics are learned in an unsupervised manner, the topic115 clusters do not fit well with the nine HL7 section types of our documents,116 because topic clusters can be either finer or more coarse-grained.117 Choi et al. [23] claim that the structure underlying EHR data improves118 the performance of prediction tasks such as heart failure prediction. As most119 EHR data do not always contain complete structure information or is com-120 pletely unavailable, they experiment alternatives to the baseline consisting121 of treating EHR data as a flat-structured bag-of-features. The proposed122 model outperformed the baseline approach for various prediction tasks such123 as readmission and mortality prediction, indicating that the detection of EHR124 structure is beneficial for many tasks.125 Rosenthal et al. [24] developed a system to detect sections in EHRs, based126 on different architectures: an RNN based system and a transfer based system127 using BERT. To overcome the lack of annotated data they propose to use128 for training purposes sections learned from medical literature (journals, text-129 books, web content). They conclude that out of domain clinical literature130 is helpful when there is not enough EHR data, but its contribution is not131 significant with bigger sizes of the in-domain annotated dataset. Their system132 did not exploit the structure of the document, that is, the fact that sometimes133 sections are ordered in a canonical order (i.e., first the Chief complaint, then134 the Antecedents, ...), which we plan to use in our approach, as it can be135 helpful in deciding section types.136 8
Rush et al. [25] solve the section identification problem using a CRF137 classifier to mark each token as belonging to a section header, and then138 they apply a rule-based post-processing module to structure the annotated139 sections. Comparing to our work, they do not perform normalization and140 therefore the number of sections they identify is not fixed. In their system,141 similar section headers are considered different, while our aim is to normalize142 each section into a set of nine HL7 section types.143 Apart from the medical domain, other areas like legal decision-support144 systems also leverage the content structure of documents. For example,145 Branting et al. [26] exploit structural and semantic regularities in law case146 corpora to identify textual patterns that have both predictable relationships147 to case decisions and explanatory value for legal decision support and ex-148 plainable outcome prediction.149 To summarize, we can see that the identification of sections is currently a150 promising area of active research, specially for languages other than English.151 Historically, rule-based methods have been the most widely used approach,152 although the recent emergence of new ML and Deep Learning techniques that153 have revolutionized the state of the art on many tasks also presents avenues154 for new developments.155 3. Materials and Methods156 In this section we will explore all the corpora and tools we have used157 in order to carry out the experiments. In the first part (section 3.1), we158 present the annotated corpus, the defined annotation model and the inter-159 annotator agreement. Section 3.2 gives a description of the large unannotated160 9
Figure 3: Distribution of sections in both train and development splits. Section chief complaint, for instance, is present in 82% of the development split, while slightly less in the training split (79%). Documents Sections Tokens training 100 744 47,449 development 100 786 48,461 test 100 754 59,119 Table 3: Details of the annotated corpus. 16
3.1.3. Inter-Annotator Agreement246 The annotation process of the corpus was performed by annotation ex-247 perts. Annotators followed an iterative process of training until a high inter-248 annotator agreement was reached. The final agreement measure was cal-249 culated on a set of 25 EDSs that were doubly annotated by two different250 annotators, reaching a pairwise agreement of 93.47% Cohen’s Kappa, indi-251 cating that the agreement is very high. There were differences with respect252 to each section type, ranging from 86% for Diagnosis (lowest agreement) to253 100% for some section types, thus reaching a significant agreement for all254 section types.255 Our annotation strategy requires each section type to be matched exactly256 while taking into account its content, and additionally returns the first and257 last lines of each section. While this strategy might seem an overly stringent258 criteria, the task is well defined as evidenced by the high inter-annotator259 agreement.260 3.2. Textual Corpora261 Deep learning techniques usually require huge amounts of data. Although262 manually annotated data gives the best results, it is very expensive and time263 consuming. For that reason, the idea of acquiring useful information in an264 unsupervised manner is very attractive, and efficient and effective methods265 have been developed. Vectorial representations of words, also known as word266 embeddings [28, 29], that are learned from textual corpora, have proven useful267 as an information source for many Natural Language Processing tasks, such268 as Part-Of-Speech (POS) tagging, Named Entity Recognition or Machine269 Translation, due to their ability to acquire relevant generalizations. These270 17
embeddings are learned through solving an appropriate optimization objec-271 tive [28] under the assumption that similar words occur in similar contexts.272 As a result, vectors of similar words derived from such optimization tend273 to reside in the neighborhood in the vector space. There are two kinds of274 embeddings, static and contextual. Static embeddings capture in a vectorial275 representation information of a word form, while contextual embeddings are276 sensitive to context, representing both a word and its context.277 This way, an unsupervised system can utilize the information based on278 word similarity in a manner that associates unseen words with those already279 occurring in the annotated corpus, thereby allowing us to cover unseen and280 misspelled terms. For instance, infarct and stroke are similar terms but one281 of them may not be in the annotated data set. The resulting word vectors282 will be fed to the neural network as input during training (see Figure 4), thus283 providing a model of the language that can help obtain better generalizations284 and, consequently, increase the recall of the final tool.285 For this work we have employed heterogeneous embedding information286 both static and contextual in order to make the system sensitive to different287 granularity and domain specificity. Regarding the granularity, during the288 section identification training, adding a character embedding layer allows289 the system to learn at the character level. Besides the character embed-290 dings learned during the training, we incorporated pre-trained, character-291 based embeddings based on fastText [30] trained over the Spanish version292 of Wikipedia. Character-based embeddings are able to generalize over n-293 grams, enabling the system to take into account prefixes and suffixes as well294 as to capture information about the different n-gram variations on the sec-295 18
tion heading words. They also generalize over zero-shot words, words that296 do not appear in the training corpus as their building element are characters297 and not words.298 Table 4 presents the details of the different word embeddings we have299 used for the task. Static embeddings were obtained applying word2vec [30]300 to Electronic Discharge Summaries (50M words), together with pretrained301 embeddings that had been calculated with Wikipedia2Vec [31], representative302 of general domain. Additionally, we also used contextual string embeddings303 [32] we calculated from Electronic Discharge Summaries and Wikipedia .304 Technique Source Embedding Details text type word2vec EDS static window length = 1, dimensions = 300, algorithm = SkipNgram Wikipedia2Vec general window length = 5, domain dimensions = 300, algorithm = Skipgram FLAIR EDSs contextual layers=1, hidden size = 2,048, sequence length = 250, mini batch size = 32 general layers = 1, hidden size = 1,024, domain sequence length = 250, mini batch size = 100 Table 4: Overview of the different embedding types used in this work (static word embeddings and contextual character embeddings). 19
3.3. Approaches to Automatic Section Identification305 In this section we will explain the different approaches we have tried with306 the aim of automatically identifying sections in medical records. First of307 all, in subsection 3.3.1 we will specify the setup we used for the rule-based308 tool that we have developed. After that, in subsections 3.3.2, and 3.3.3 we309 will explore the ML algorithms we have employed, the Perceptron and Deep310 Learning, respectively. For both ML approaches, we have approached the311 task as a sequential learning process [33, 34], where the text is considered a312 sequence of tokens, and each token is associated with one tag indicating its313 corresponding section. We have used an IOB (Inside, Outside, Begin) tag314 model, where the beginning of each section is marked with a B tag (e.g., B-315 DIA for the token starting a diagnosis), the tokens inside a section are marked316 with an I tag (I-DIA will mark a token inside a diagnosis section), and using317 the O tag for elements that do not belong to any section (see Figure 5). This318 way, section identification can be viewed as the detection of extended and319 long entities. This approach has been successfully used in similar tasks as320 the identification of elementary discourse units (text segments consisting of321 one or several sentences) in Discourse processing [35] or topic segmentation322 [22]. Figure 4 presents an architecture of the system.323 20
027431 XX-XX-XXXX 66 años. VARON. MC: REFERENCIADO EN EL INFORME. INFORME AL ALTA : Paciente de 66 años. No alergias medicamentosas conocidas. A. PERSONALES: Enfermedad de Crohn diagnosticada en 1997 con afectación de íleon terminal (A3L1B2) por cuadros suboclusivos resueltos con enfermedad de íleon terminal asociada a mesenteritis fibrosa. Artrosis dorso-lumbar. Cirugía de hernia inguinal. Ci Tto: Dacortin 5: 1-0-0; Pariet 20: 1-00; Pentasa:1-1-1; Kilor 0-1-0, Clinutren: 2/día. E. ACTUAL: Acude a Urgencias por dolor abdominal generalizado con febrícula, sin tiritona, sin naúseas ni vómitos. Sin alteración del ritmo intestinal. Con pauta descendente de corticoides, después del último ingreso por cuadro suboclusivo. EXPL. FÍSICA: Paciente consciente, orientado, colaborador. Buena coloración de piel y mucosas. Cuello: no adenopatías cervicales. AC: rítmica sin soplos. No roncus ni crepitantes. Abdomen: distendido, timpánico. Peristaltismo ausente. Blumberg negativo. EEII: no edemas maleolares. PPP. RX ABDOMEN: Sugestivo de suboclusión intestinal. Knowledge-based system Perceptron Neural Network 027431 XX-XX-XXXX 66 años. VARON. MC: REFERENCIADO EN EL INFORME. INFORME AL ALTA : Paciente de 66 años. No alergias medicamentosas conocidas. A. PERSONALES: Enfermedad de Crohn diagnosticada en 1997 con afectación de íleon terminal (A3L1B2) por cuadros suboclusivos resueltos con enfermedad de íleon terminal asociada a mesenteritis fibrosa. Artrosis dorso-lumbar. Cirugía de hernia inguinal. Ci Tto: Dacortin 5: 1-0-0; Pariet 20: 1-00; Pentasa:1-1-1; Kilor 0-1-0, Clinutren: 2/día. E. ACTUAL: Acude a Urgencias por dolor abdominal generalizado con febrícula, sin tiritona, sin naúseas ni vómitos. Sin alteración del ritmo intestinal. Con pauta descendente de corticoides, después del último ingreso por cuadro suboclusivo. EXPL. FÍSICA: Paciente consciente, orientado, colaborador. Buena coloración de piel y mucosas. Cuello: no adenopatías cervicales. AC: rítmica sin soplos. No roncus ni crepitantes. Abdomen: distendido, timpánico. Peristaltismo ausente. Blumberg negativo. EEII: no edemas maleolares. PPP. RX ABDOMEN: Sugestivo de suboclusión intestinal. Regular expressions (EXPL | PRUEB).*COMPL.* ... BiLSTM layer B-CC I-CC CRF layer Character emb. layer <blank> p a …. <blank> s a <blank> El paciente ingresa por fuerte dolor B-CC I-CC …. I-CC Figure 4: Architecture of the system. Three different approaches have been used: regular expressions, the Perceptron algorithm and neural networks. 3.3.1. Rule-Based Approach324 Manually defined rules have been used since the early years of Artificial325 Intelligence, and are still a competitive method to achieve acceptable results.326 Their downside is the effort needed to include knowledge into the automatic327 system. Another drawback is their lack of generalization, because a change328 in the domain may imply a complete re-implementation of the rule system.329 Regarding the identification of sections in medical records, this approach330 21
has been used in many systems [13, 36], where acceptable results have been331 reported, although in several cases the approach has not been general, but332 rather limited to a reduced set of very specific sections or portions of text.333 Table 2 presents several examples of the beginning of different section types. The table shows how there is a high variability difficult to capture using rules, specially with implicit sections with no standard title, like in the Chief Complaint and Complementary Exploration. Examples (1) and (2) present two rules that try to capture the start of the Chief Complaint and the Current Illness sections, where the parentheses enclose optional elements. The objective was to cover the different options found in the training set. (1) MOTIVO(S) (DE(L)(A)) INGRESO|PETICION|334 EXP LORACION|ESTUDIO|CONSULT A (ACT UAL)335 (2) (E.|ENFERMEDAD|SIT UACI ´ ON|EP ISODIO|ESTADO)336 (A.|ACT UAL)|SINT OMAT OLOG´ IA337 3.3.2. Machine Learning: Perceptron338 For the application of ML to section identification, we modeled the prob-339 lem as a sequence to sequence problem. The task consists in learning to340 map from input word sequences w1...wm|wi∈Wto output tag sequences341 t1...tm|ti∈T.342 Although some approaches to section identification used sentence se-343 quences as input units [13], we preferred to model this problem using word344 sequences as input units to capture the fact that individual words in the345 22
right context are good signals for sections and also to reduce sparsity, be-346 cause sentence sequences are more sparse than word sequences. The problem347 is cast as the assignment of the correct tag to each token. Although the tag348 assignment is made token by token, the final evaluation will be done on the349 detection of complete sections.350 To do so, we employed the Averaged Structured Perceptron algorithm351 [37, 38] which combines the Perceptron algorithm for learning linear classi-352 fiers with an inference algorithm and converts a classification problem into a353 ranking problem. The objective of the algorithm is to find, for each sentence,354 the sequence of tags with the maximum score. This prediction decision pro-355 cess is divided into a sequence of smaller decisions made from left-to-right.356 Thus, at each step there is a word and its context, called the history, in357 which the local tagging decision is made, namely to predict the tag given the358 history. The history can be represented in several ways, using the prefixes of359 a given number of previous words, and/or the suffixes, or any other features360 that could be relevant for the task and then converted into a feature vector361 where each feature will get a weight through the learning process.362 Formally, the problem can be stated as follows. Given:363 •A sequence of input words w1...wm, for simplicity referred as w.364 •The sequence of tags t1...tmas t(this way, the set of possible tags is365 T).366 •In our case the context in which a tagging decision is made is repre-367 sented by the history tuple h:< t−2, t−1, w−2, w−1, w0, w+1, sx0, px0, cap,368 num, i >, where t−2and t−1are the previous two tags, w−2and w−1are369 23
the previous two words at a given position i(this way, Hcorresponds370 to the set of all possible histories).371 sx and px correspond to different sizes of word suffixes, (in this work,372 xvarying from 2 to 4) and prefixes of w0.373 cap and num correspond to two binary features to account for capital-374 ization and number status at the current word.375 The feature mapping function Φ : H×T→Rdmaps a history-tag376 pair to a d-dimensional feature vector we mentioned before. The Structured377 Perceptron models P(t|w) as P(t|h;α) where α∈Rdis a parameter vector378 representing the weight of each feature of Φ. P(t|h;α) is calculated as α·379 Φ(h, t) and the objective function is:380 ˆ t=argmaxtPd 1αi·Φi(h, t)381 Usually the Viterbi algorithm is applied when used on sequence data,382 in order to efficiently calculate the best tag sequence using dynamic pro-383 gramming. The algorithm is competitive to other options such as maximum-384 entropy taggers or CRFs [33].385 Figure 5: Simplified Architecture of the Structured Perceptron (the upper three rows exemplify the use of word features (first row), 3 letter prefixes and 3 letter suffixes (second and third rows). 24
We employed our own implementation of this tagger following [37]. We386 trained 100 iterations and selected the model corresponding to the iteration387 that achieved the best score on the development set. Although the algorithm388 achieves a competitive performance compared to state of the art methods,389 this approach requires a feature engineering effort to identify, select and390 properly encode relevant features.391 3.3.3. Machine Learning: Neural Networks392 In addition to a traditional neural network like Perceptron, we explored393 transfer learning methods. In this case, we used FLAIR [39], a bi-directionally394 trained Language Model (LM) using Recurrent Neural Networks (RNN),395 where the basic element is the character and not the word. Based on its char-396 acters, FLAIR generates pre-trained contextual embeddings for each word by397 concatenating the hidden state for the last character of the word in the for-398 ward neural network and the first character of the word in the backward399 neural network, as shown in Figure 6. As described in [39], formally, the ob-400 jective function of a character-based LM is to maximize the sum of the logs401 of P(xt|x0, ..., xt−1), that is to say, an estimate of the predictive distribution402 over the next character given past characters. FLAIR allows us to com-403 bine different types of embeddings by concatenating each embedding vector404 to form the final word vector. We employed a combination of embeddings405 as previously reported in section 3.2. One of the main advantages of these406 methods is that there is no need for feature engineering.407 25
Figure 8: Effect of training and testing on the same or a different hospital (H1: GaldakaoUsansolo hospital, H2: Basurto hospital), measured by F-score. The difference is significant in almost all types of sections. Specially503 relevant is the effect of the system trained on hospital H2 and tested on504 hospital H1 for the Heading section type (H column), where the F-score is505 very low. We examined the results and concluded that this happens mostly506 because the headings show a great variation, added to the fact that the507 data present in the headings is generated automatically most of the times,508 including record numbers or dates, and this implies that they can be different509 enough to confuse an automatic system. Surprisingly, this does not happen510 in the opposite direction, meaning that the data from hospital H1 shows511 more variability and is useful to account for the instance types of hospital512 H2. The difference is also significant for the second section type (Chief513 32
Complaint, CC), although less drastic. This was due to a cascade effect514 as a result of applying a sequence to sequence approach, as the errors in515 delimiting the first section of the document frequently are carried from one516 section to the next one. Finally, for some section types, like E(xploration),517 EV(olution) and D(iagnostic), we can see how applying a system trained on518 a different hospital can outperform the system based on data from the same519 hospital. This can be due to the fact that one hospital agrees more with the520 conventions of the other hospital for these section types.521 5.2. Error Analysis522 We looked at the errors given by the different systems. For simplicity, we523 will only examine the results of the best system based on neural networks.524 An examination of the divergences between the output of the system and the525 gold standard showed us the main causes of error:526 •Errors given by the inherent difficulty of spontaneously written section527 headings. Although explicit headings are an important clue to delimit528 sections, the variability of their writings together with the limited size529 of the training set (100 documents, which means that there are at most530 100 instances of each section type) is a source of errors.531 •Implicit sections. Some types of sections have a majority of instances532 without an explicit section heading, which means that the section must533 be detected using its content words (see Table 2).534 •Mixed sections. Although the annotators have decided the exact scope535 of each section with a high agreement, the use of unstructured and536 33
spontaneously written EDSs gives the writers freedom to describe any537 concept in different places. As an example, the section corresponding538 to the Medical History can contain passages related to past diagnoses,539 treatments and explorations, which can pose a challenge for an auto-540 matic system.541 •A special case of mixed sections can be the confusion between two542 related section types:543 –Chief Complaint and Current Illness. These two sections present544 the most diffuse definition [27], and are the cause of several errors.545 –Exploration and Complementary Exploration. Although the defi-546 nition of each section is precise, sometimes physicians mix them547 in the same block or paragraph.548 In Section 3.1.2, we mentioned that the ordering of section types shows549 a great variability. In order to measure its impact on the results, we split550 the test set in two subsets. The first subset corresponds to the documents551 that follow the canonical order (26% of the documents), while the rest of the552 documents conform the second subset (non-canonical order and/or missing553 sections, 74% of the documents). Since our sequence learning-based methods554 depend on the ordering for predicting the next token, this has an effect in the555 IOB-labeling prediction, with a F-score of 95.40 for the canonical documents556 and 89.81 for the non-canonical ones.557 Figure 9 presents the main types of mistakes made by the automatic558 tool. It shows how the errors are concentrated in some sections, like Chief559 Complaint (CC), Medical history (MH) and Diagnosis (D). Overall, the dis-560 34
tinction of different sections is reflected in the text by means of different clues,561 ranging from semantics (the content of each section) to syntax (e.g., use of562 section headings and separate paragraphs or text blocks for each section)563 but, in most of the errors, these conventions do not hold, and this causes the564 automatic tool to find an additional difficulty.565 Figure 9: Confusion matrix, where darker green means a higher frequency, of each instance (H: Heading, CC: Chief Complaint, MH: Medical History, CI: Current Illness, E: Exploration, EC: Complementary Exploration, EV: Evolution, D: Diagnosis, T: Treatment). 6. Conclusion566 We present a system for Section Identification in Discharge Summaries567 written in Spanish. We have adopted an annotation model based on H7 CDA568 R2 for Electronic Discharge Summaries (EDS) of the Spanish Health System,569 and we have applied it to manually annotate a corpus of 300 EDSs, obtaining570 a high inter-annotator agreement.571 We have evaluated the contribution of different rule-based and Machine572 Learning approaches and study the strengths and weaknesses of each option.573 Most previous works have used section identification as an auxiliary module574 35
for carrying on clinical processing, relying on a rule-based approach. How-575 ever, our results show that section identification is a task on its own, where576 simple methods do not obtain the best results. The Machine Learning sys-577 tems obtain results that are good enough for the application of the system578 in a production setting. Specifically, we show that Language Model tuning is579 a key factor, as a Language Model-based transfer learning provides the best580 performance. The paper has also studied the generalization ability of mod-581 els trained in different hospitals, showing that different section types have582 significant differences in some cases.583 The developed automatic annotation models and software are freely avail-584 able contacting the authors.585 Acknowledgements586 We gratefully acknowledge the support of NVIDIA Corporation with the587 donation of the Titan X Pascal GPU used for this research. This work was588 partially funded by the Spanish Ministry of Science and Innovation (DOTT-589 HEALTH/PAT-MED PID2019-106942RB-C31), the European Commission590 (FEDER), the Basque Government (IXA IT-1343-19), and the EU ERA-591 Net CHIST-ERA and the Spanish Research Agency (ANTIDOTE PCI2020-592 120717-2).593 References594 [1] C. Peterson, C. Hamilton, P. Hasvold, From innovation to implementa-595 tion – eHealth in the WHO European Region, World Health Organiza-596 tion, 2016.597 36
[2] M. Adnan, J. Warren, M. Orr, A. Ewens, J. Scott, S. Trubshaw, The598 quality of electronic discharge summaries for post-discharge care: Hos-599 pital panel assessment and it to support improvement, Health Care and600 Informatics Review Online 15 (2011).601 [3] Health at a Glance: Europe 2018 STATE OF HEALTH IN THE EU CY-602 CLE, https://ec.europa.eu/health/sites/default/files/state/603 docs/2018_healthatglance_rep_en.pdf, 2020. Last Online; accessed604 31-05-2021.605 [4] Health IT Data Summaries, https://dashboard.healthit.gov/apps/606 health-information-technology-data-summaries.php, 2021. Last607 Online; accessed 31-05-2021.608 [5] Connecting health and care for the nation: A shared nationwide in-609 teroperability roadmap. Office of the National Coordinator for Health610 Information Technology (ONC). Washington, DC: U.S. Department of611 Health and Human Services (HHS), 2015.612 [6] Recommendation on a European Electronic Health Record exchange613 format. European Commision, 2019.614 [7] State of Interoperability among U.S. Non-federal Acute Care Hospitals615 in 2018 , https://www.healthit.gov/sites/default/files/page/2020-616 03/State-of-Interoperability-among-US-Non-federal-Acute-Care-617 Hospitals-in-2018.pdf, 2020. Last Online; accessed 31-05-2021.618 [8] openehr, https://www.openehr.org, 2020. Last Online; accessed 31-619 05-2021.620 37
[9] Health Level Seven (HL7). FHIR, http://www.hl7.org, 2019. Last On-621 line; accessed 31-05-2021.622 [10] Health Level Seven (HL7). CDA, http://www.hl7.org, 2019. Last On-623 line; accessed 31-05-2021.624 [11] H. K., K. Saranto, N. P., Definition, structure, content, use and impacts625 of electronic health records: a review of the research literature, Int J626 Med Inform. 77 (5) (2008) 291–304.627 [12] W. LL., Medical records that guide and teach, N Engl J Med. 14) (1968)628 593–600.629 [13] A. Pomares-Quimbaya, M. Kreuzthaler, S. Schulz, Current approaches630 to identify sections within clinical narratives from electronic health631 records: a systematic review, BMC Medical Research Methodology 19632 (2019).633 [14] T. Edinger, D. Demner-Fushman, A. Cohen, S. Bedrick, H. W., Evalu-634 ation of Clinical Text Segmentation to Facilitate Cohort Retrieval, in:635 AMIA Annu Symp Proc., pp. 660–669.636 [15] J. Lei, B. Tang, X. Lu, K. Gao, M. Jiang, H. Xu, A comprehensive637 study of named entity recognition in chinese clinical text, Journal of the638 American Medical Informatics Association 21 (2014).639 [16] Y. Wang, L. Wang, M. Rastegar-Mojarad, S. Moon, F. Shen, N. Afzal,640 S. Liu, Y. Zeng, S. Mehrabi, S. Sohn, H. Liu, Clinical information extrac-641 tion applications: A literature review, Journal of Biomedical Informatics642 77 (2018) 34 – 49.643 38
[17] H.-J. Lee, Y. Zhang, M. Jiang, J. Xu, C. Tao, H. Xu, Identifying direct644 temporal relations between time and events from clinical notes, BMC645 Medical Informatics and Decision Making 18 (2018).646 [18] A. P´erez, K. Gojenola, A. Casillas, M. Oronoz, A. D´ıaz de Ilarraza,647 Computer aided classification of diagnostic terms in spanish, Expert648 Systems with Applications, 42(6), 2949–295 (2015).649 [19] A. Atutxa, A. D. de Ilarraza, K. Gojenola, M. Oronoz, O. P. de Vi˜naspre,650 Interpretable deep learning to map diagnostic texts to icd-10 codes, In-651 ternational Journal of Medical Informatics 129 (2019) 49 – 59.652 [20] K. Xu, M. Lam, J. Pang, X. Gao, C. Band, P. Mathur, F. Papay, A. K.653 Khanna, J. B. Cywinski, K. Maheshwari, P. Xie, E. P. Xing, Multimodal654 machine learning for automated icd coding, in: F. Doshi-Velez, J. Fack-655 ler, K. Jung, D. Kale, R. Ranganath, B. Wallace, J. Wiens (Eds.), Pro-656 ceedings of the 4th Machine Learning for Healthcare Conference, volume657 106 of Proceedings of Machine Learning Research, PMLR, Ann Arbor,658 Michigan, 2019, pp. 197–215.659 [21] A. Duque, H. Fabregat, L. Araujo, J. MartinezRomo, A keyphrasebased660 approach for interpretable ICD-10 code classification of Spanish medical661 reports, Artificial Intelligence in Medicine (2020).662 [22] S. Arnold, R. Schneider, P. Cudr´e-Mauroux, F. A. Gers, A. L¨oser, Sec-663 tor: A neural model for coherent topic segmentation and classification,664 Transactions of the Association for Computational Linguistics 7 (2019)665 169–184.666 39
[23] E. Choi, Z. Xu, Y. Li, M. W. Dusenberry, G. Flores, E. Xue, A. M.667 Dai, Learning the graphical structure of electronic health records with668 graph convolutional transformer, in: Association for the Advancement669 of Artificial Intelligence (AAAI).670 [24] S. Rosenthal, K. Barker, Z. Liang, Leveraging medical literature for671 section prediction in electronic health records, in: Proceedings of the672 2019 Conference on Empirical Methods in Natural Language Processing673 and the 9th International Joint Conference on Natural Language Pro-674 cessing (EMNLP-IJCNLP), Association for Computational Linguistics,675 Hong Kong, China, 2019, pp. 4864–4873.676 [25] E. Rush, I. Danciu, G. Ostrouchov, K. Cho, B. Mayer, Y.-L. Ho, J. Hon-677 erlaw, L. Costa, F. Linares, E. Begoli, Jsonize: A scalable machine678 learning pipeline to model medical notes as semi-structured documents,679 AMIA Joint Summits on Translational Science proceedings. AMIA Joint680 Summits on Translational Science 2020 (2020) 533–541.681 [26] L. K. Branting, C. Pfeifer, B. Brown, L. Ferro, J. Aberdeen, B. Weiss,682 M. Pfaff, B. Liao, Scalable and explainable legal prediction, Artificial683 Intelligence and Law (2020).684 [27] A. R. Terroba, Mejora de la calidad del informe cl´ınico de alta hospita-685 laria desde el punto de vista ling¨u´ıstico, PhD Thesis, University of La686 Rioja (2018).687 [28] T. Mikolov, K. Chen, G. S. Corrado, J. Dean, Efficient Estimation of688 Word Representations in Vector Space, CoRR abs/1301.3781 (2013).689 40
[29] J. Pennington, R. Socher, C. D. Manning, Glove: Global vectors for690 word representation, in: Empirical Methods in Natural Language Pro-691 cessing (EMNLP), pp. 1532–1543.692 [30] T. Mikolov, E. Grave, P. Bojanowski, C. Puhrsch, A. Joulin, Advances693 in pre-training distributed word representations, in: Proceedings of the694 International Conference on Language Resources and Evaluation (LREC695 2018).696 [31] I. Yamada, A. Asai, H. Shindo, H. Takeda, Y. Takefuji, Wikipedia2vec:697 An optimized tool for learning embeddings of words and entities from698 wikipedia, arXiv preprint 1812.06280 (2018).699 [32] A. Akbik, D. Blythe, R. Vollgraf, Contextual string embeddings for700 sequence labeling, in: Proceedings of the 27th International Conference701 on Computational Linguistics, pp. 1638–1649.702 [33] A. McCallum, W. Li, Early results for named entity recognition with703 conditional random fields, feature induction and web-enhanced lexicons,704 in: Proceedings of the Seventh Conference on Natural Language Learn-705 ing at HLT-NAACL 2003 - Volume 4, CONLL ’03, Association for Com-706 putational Linguistics, Stroudsburg, PA, USA, 2003, pp. 188–191.707 [34] A. Jagannatha, H. Yu, Bidirectional recurrent neural networks for med-708 ical event detection in electronic health records, CoRR abs/1606.07953709 (2016).710 [35] A. Atutxa, K. Bengoetxea, A. D. de Ilarraza, M. Iruskieta, Towards a711 top-down approach for an automatic discourse analysis for basque: Seg-712 41