Full text
GutBrainIE@CLEF25 Annotation Guidelines Vanessa Bonato, Giorgio Maria Di Nunzio, Nicola Ferro, Ornella Irrera, Stefano Marchesin, Marco Martinelli, Laura Menotti, Gianmaria Silvello, Federica Vezzani. Document Structure 1. Overview................................................................................................................. 2 2. Annotation Framework.......................................................................................... 4 Annotation process in practice.......................................................................... 5 3. Entity Annotation Guidelines................................................................................ 7 3.1. Entity Labels...............................................................................................7 3.2. General Entity Annotation Rules.............................................................. 14 4. Relation Annotation Guidelines..........................................................................21 4.1. Relation Labels.........................................................................................21 4.2. General Relation Annotation Rules.......................................................... 22 5. Bibliography......................................................................................................... 26 6. Appendix............................................................................................................... 28 Conceptual System......................................................................................... 28 GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 1
1. Overview The GutBrainIE@CLEF25 dataset aims to foster the development of Information Extraction (IE) systems that support experts by automatically extracting and linking knowledge from scientific literature, thereby enhancing the understanding of the gut-brain interplay and its role in neurological diseases. Recent evidence suggests a connection between neurological and gut disorders that might play a critical role in mental health-related conditions such as Multiple Sclerosis, Parkinson’s, and Alzheimer’s [7-9]. These guidelines aim to assist annotators in consistently labeling named entities and relationships within this dataset. Named entity labeling involves identifying and classifying specific text spans (entity mentions) into predefined categories, while relationship labeling determines if a particular relationship defined between two entity types holds or not [1]. In cases where multiple relationships are defined between two entity types, relationship labeling also determines which specific relation holds. The dataset will focus on biomedical titles and abstracts related to the gut microbiota and its effects on mental health [7], extracted from PubMed. The entities and relations of interest are defined by the conceptual system that can be found in the Appendix Section. Formally, the GutBrainIE task is divided into two subtasks: 1. Named Entity Recognition (NER): Participants are provided with PubMed abstracts discussing the gut-brain interplay and are asked to extract named entities related to genes, bacteria, intermediates, and disorders. 2. Relation Extraction (RE): Participants are tasked with identifying relations between pairs of extracted entities within a document (title + abstract). GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 2
The submitted results are evaluated using annotations created by a team of expert annotators and validated by a team of biomedical specialists. These guidelines build upon the methodologies and practices used in crafting EnzChemRED [1], BioRED [2], and BioASQ-QA [3], providing a common approach for annotating abstracts with minimal ambiguity and the highest possible inter-annotator agreement (IAA). GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 3
2. Annotation Framework The annotation process for the GutBrainIE dataset involves two tasks [2]: 1. Named Entity Recognition (NER): Identify and classify text spans into one of the defined entity categories. Formally speaking, the goal of NER is to identify mentions of the defined entity types in the text. The NER task can be seen as one of sequence labeling: text is represented as a sequence of tokens (in our case, words) where� = (�1,�2, ···, ��) denotes the length of the text, and the goal is to classify a sequence of tokens� with into a corresponding label .(��,��+1, ···, ��+�)�, (�+�)∈[0, �]��∈(�1,�2, ···, ��) The set is defined as the label set for the model, and each label (�1,�2, ···, ��)� represents a specific entity type that can be found in the texts. 2. Relation Extraction (RE): Identify relationships between entities, as explicitly or implicitly expressed in the text. This problem comprises two different but complementary tasks: ○Binary RE: Given a pair of entity mentions having labels(�1,�2) (�1,�2) assuming there is a relation defined from to or vice versa, where� ∈ � �1 �2 is the set of relations defined for the GutBrainIE dataset, the objective is to� state whether that relation holds between these two entities. The classification should be consistent with the predefined relation types � defined for the GutBrainIE dataset [1]. ○Ternary RE: Given a pair of entity mentions having labels , the (�1,�2) (�1,�2) objective is to determine if there is a relation (predicate) that links the �∈� ternary tuple or . The entity appearing before in the(�1,�,�2) (�2,�,�1)� ternary tuple is referred to as “head” or “subject”, while the one appearing after is referred to as “tail” or “object”. The relation must align with the� GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 4
context in which the entities appear and should fit the predefined relation types defined for the GutBrainIE dataset [1].� Normally, the process of dataset crafting also comprises the Named Entity Normalization (NEN) task, namely linking identified entity mentions to stable and unique database identifiers (e.g., MeSH,ChEBI) to standardize and contextualize entities [1]. At the moment we are not interested in this task. Annotation process in practice The documents are uploaded on the annotation platform MetaTron [4] with pre-annotations for entities performed by unsupervised algorithms to speed up the annotation process [1]. Before starting to annotate, annotators are required to read the guidelines carefully, paying particular attention to the defined entity and relation labels to ensure comprehensive understanding and consistency. We are aware that the defined annotation guidelines may impose certain restrictions that could result in the loss of potentially relevant or important information. However, these rules are crucial for minimizing uncertainties among annotators and, consequently, ensuring more uniform and consistent annotations. Annotators are required to follow a specific process when annotating each document to ensure consistency and accuracy: 1. Initial Reading: First, read carefully through the entire abstract without the pre-annotated mentions visualized. This helps in understanding the context without bias from pre-annotated labels [5]. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 5
2. NER Annotation: Once familiar with the abstract, enable the visualization of the pre-annotated labels and proceed with NER. When an entity instance that appears multiple times in the document is annotated for the first time, the ‘annotate all’ MetaTron feature may be used to uniformly annotate all subsequent occurrences of the entity, ensuring consistency throughout the document. 3. RE Annotation: After completing NER, proceed with RE to identify and classify relationships between the annotated entity mentions. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 6
3. Entity Annotation Guidelines 3.1. Entity Labels See the Appendix for a graphical representation of the schema (entities and relationships). GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 7 Entity Label URI Definition Reference URL Notes Examples Anatomical Location NCIT_C13717 Named locations of or within the body. The categories "anatomic site" of NCIT and "organism subdivision" of UBERON can be used as references for this label. link Instances of "... axis" (e.g., "gut-brain axis") and "... system" (e.g., "immune system”) should NOT be labeled as anatomical location entity mentions. The microbiota resides in various parts of the body, such as the oral cavity, nasal passages, lungs, gut, skin, bladder, and vagina. We found that PVD and cardiovascular disease were associated with lower microbiota diversity in the gut (i.e., α-diversity), while supplemental vitamin use was associated with higher α-diversity. Animal NCIT_C14182 A non-human living organism that has membranous cell walls, requires oxygen and organic foods, and is capable of voluntary movement, as distinguished from a plant or mineral. Human is a subordinate concept of Animal but for CLEF we distinguish only between humans and any other animal. link Instances of “... model”, such as “rodent model”, should NOT be annotated as animal entity mentions. Although approximately 30% mice are resilient to chronic social defeat stress (CSDS), the role of gut microbiota in this is unknown. We further demonstrated the role of Htr1a using AAV-shRNA to downregulate Htr1a in the mPFC of CUS mice. We compared the 16S ribosomal RNA (rRNA) gene sequences retrieved from fecal samples between control, CUMS-vulnerable, and CUMS-resilient mice.
GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 8 Entity Label URI Definition Reference URL Notes Examples Statistical Technique NCIT_C19044 A method of analyzing or representing statistical data; a procedure for calculating a statistic. link Instances of “randomized controlled trials” and “cohort studies” are too generic and should NOT be annotated. Pearson's correlation analysis was used to evaluate the association between bacterial taxa and psychotic symptoms. Linear Discriminant Analysis (LDA) revealed Ruminococcaceae as a discriminative feature. Biomedical Technique NCIT_C15188 Research concerned with the application of biological and physiological principles to clinical medicine. This category also includes the assay category, defined as a planned process with the objective of producing information about a material entity (the evaluant) by examining it. link Instances of “microbiota analysis” should NOT be annotated as biomedical technique entity mentions. The 16S rRNA amplicon sequencing method was performed to determine the fecal composition of fecal microbiota. The intestinal permeability biomarker zonulin was measured using enzyme-linked immunosorbent assays. Bacteria NCBITaxon_2 One of the three domains of life (the others being Eukarya and ARCHAEA), also called Eubacteria. They are unicellular prokaryotic microorganisms which generally possess rigid cell walls, multiply by cell division, and exhibit three principal forms: round or coccal, rodlike or bacillary, and spiral or spirochetal. link Microorganism that is unicellular and prokaryotic Here we demonstrated that stress resistance in mice was associated with more abundant Lactobacillus and Akkermansia in the gut, but less abundant Bacteroides, Alloprevotella, Helicobacter, Lachnoclostridium, Blautia, Roseburia, Colidextibacter and Lachnospiraceae NK4A136. The abundance of Akkermansia, Megamonas, Prevotellaceae NK3B31 group, and butyrate-producing bacteria, Lachnospira, Subdoligranulum, Blautia, and Dialister, and
GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 9 Entity Label URI Definition Reference URL Notes Examples acetate-producing bacteria, Streptococcus, in the gut microbiota of the MDD group was lower than that in the control (C) group. Chemical CHEBI_59999 A chemical substance is a portion of matter of constant composition, composed of molecular entities of the same type or of different types. This category also includes metabolites, which in biochemistry are the intermediate or end product of metabolism (more information here), and neurotransmitters, which are endogenous compounds used to transmit information across the synapses. link The list of chemicals reported in https://pubchem .ncbi.nlm.nih.go v/ can be used as a reference for this label. The administration of sodium butyrate and cryptotanshinone (CPT) led to… We observed differentially abundant microbial-derived neuroactive metabolites including multiple B-vitamins, kynurenic acid, gamma-aminobutyric acid and short-chain fatty acids. Microbial metabolites (short-chain fatty acids -SCFAs-, bile acids, amino acids, tryptophan -trpderivatives, and more), work as signaling pathways. The gut microbiota produces and modulates neurotransmitters such as GABA, serotonin, dopamine, glutamate, etc. Both Nef and Flu treatments induced significant increases in the levels of anti-depressant neurotransmitters, including dopamine (DA), serotonin (5-HT), and norepinephrine (NE).
●“Gut microbiota analysis are conducted to..”; ●“... plays a crucial role in stress reactivity over the life span”. ◆Annotate composite entities as a single entity if they belong to the same category. However, if entities belong to the same category but appear as a sequence, annotate them separately. Examples: ●"SMADs 1, 5, and 8" should be annotated as a gene; ●"breast or ovarian cancer" should be annotated as a disease; ●"b-sitosteryl and stigmasteryl linoleates" should be annotated as a chemical; ●In "Cytochrome P-450 genes (CYP1A1, CYP2A6, CYP2D6, and CYP2E1)," label each of "Cytochrome P-450 genes," "CYP1A1," "CYP2A6," "CYP2D6," and "CYP2E1" as separate gene entities. ➔Overlapping and ambiguous entities ◆DO NOT annotate overlapping entities. An entity cannot include words that are already part of another entity mention. ●Although we acknowledge that allowing overlapping annotations of entity mentions would be the best approach to better capture the complexity and nuances of biomedical texts, at the current stage, we do not permit overlapping annotations. This decision has been made to simplify the annotation process and reduce the ambiguity that could arise during manual annotation. ●In "anti-mouse IL-6 receptor antibody" annotators should label the entire text span as a chemical, while they DO NOT have to annotate "mouse" as an animal. ◆A text span can have only one entity label. If an entity could belong to multiple categories, use the context within the sentence to determine the appropriate label. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 16
●In "... blockade of interleukin-6 receptor in the periphery …" "interleukin-6" might be labeled as a gene, but since it is followed by the word "receptor" we label the entire "interleukin-6 receptor" as a neurotransmitter, therefore as a chemical. ●In some cases, it can be challenging to determine whether an entity mention should be annotated as a metabolite,gene,neurotransmitter, or chemical. This difficulty arises because these categories are all subsets of chemical, meaning that metabolites, genes, and neurotransmitters are also chemicals. When the appropriate label cannot be determined through annotation tools, context, or reasoning, the entity should be labeled with the most generic category chemical. Examples: ○In “the enzyme was inhibited by 3-hydroxybutyrate” it is not clear by the context if the entity mention is being used as a metabolite or a chemical. Therefore, it should be annotated as achemical. ○In “ASD patients shower increased dopamine compared to…” it is not clear by the context if dopamine is being used as a neurotransmitter or as a chemical. Therefore, it should be annotated as a chemical. ◆If an entity mention could be annotated in multiple ways, always annotate the longest version. Examples: ●“chronic sleep disorder” must be annotated in full, rather than just “chronic sleep disorder”; ●“major depressive disorder” must be annotated in full, rather than just “major depressive disorder”. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 17
◆When determining the span to annotate and deciding whether to keep the longest version, use the defined annotation tools to search if the longest version is recognized in a well-established knowledge source. ◆When determining the span to annotate, use the existence of reference acronyms in the literature as an indicator to keep the longest version of the entity mention. Examples: ●“major depressive disorder” is, in the literature, associated with the acronym MDD; ●“amyotrophic lateral sclerosis” should be annotated in full, rather than just “amyotrophic lateral sclerosis” or “amyotrophic lateral sclerosis”, since the acronym “ALS” is defined for the entire condition. ➔Special cases and abbreviations ◆DO NOT annotate terms that are identical to the entity labels. For example, the term “disease” by itself should not be annotated as a disease entity, while the term “Parkinson disease” should be annotated as a disease entity. ◆Annotate both the abbreviation and its long form separately, if possible. Example: ●In “Prostaglandin E2 (PGE2)” annotate both “Prostaglandin E2” and “PGE2” as separate chemical entities [2]. ◆If the boundary of an entity comprises both the full name name and its abbreviation, it should be annotated as a single entity. Example: ●"Deoxyguanosine kinase (dGK) deficiency" should be annotated as a single disease entity [2]. ◆DO NOT annotate words that are morphological variations of terms that would be entities. Examples: ●Do not annotate "hypertensive" as a disease, even though it is an adjective form of "hypertension." GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 18
●Do not annotate "VLCAD deficient" as a disease, even though it refers to "VLCAD deficiency." ◆Certain terms may appear to fall within one of the defined categories but should not be annotated because they do not represent actual instances of those entities. ●Instances of “xxx axis”, such as “gut-brain axis” or “hypotalamic-pituitary-adrenal axis”, even if they might be interpreted as anatomical locations, do not actually fit appropriately into any of the available entity labels and should be excluded from annotation. ●Instances of "xxx system," such as "immune system" or "endocrinal system," should be excluded from annotation. These systems are organizations of varying numbers and types of organs, arranged to perform complex functions for the body. Therefore, they do not represent a specific anatomical location but rather a set of anatomical locations linked together, and should not be annotated as anatomical location entity mentions. ●Instances of “xxx models”, such as “human models”, “animal models”, “rodent models”, etc.. should not be annotated. Although they contain terms like “human” or “animal”, these refer to experimental models rather than actual instances of human or animal entities. ●Instances of “xxx response” and “xxx mechanism”, such as “stress response” or “pathophysiological mechanism”, describe biological or physiological processes rather than specific DDF entity mentions and, therefore, should be excluded from annotation. ●Instances of "xxx transplant", "xxx therapy", "xxx treatment", "xxx intervention", and "xxx scale" are general medical procedures, GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 19
therapeutic approaches, or assessment frameworks and therefore should not be annotated as biomedical technique entity mentions. ●Instances of “microbiota analysis” are too generic and do not refer to specific biomedical techniques. Therefore, such mentions should not be annotated. ●Instances of “randomized controlled trials” and “cohort studies” are more aligned with research methods rather than specific biomedical or statistical techniques. Therefore, such mentions should not be annotated. ➔Use of full text and tools ◆Annotators can access the full text and use various tools detailed in the “Annotation Tools” section to clarify entity boundaries and labels [2]. If these tools are not sufficient to clarify their doubts, annotators are free to search the internet for more information. However, they must pay careful attention to the reliability of the websites being consulted, prioritizing reputable and authoritative sources. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 20
4. Relation Annotation Guidelines 4.1. Relation Labels Head Entity Tail Entity Predicate Anatomical Location Human / Animal located in Bacteria Bacteria / Chemical / Drug interact Bacteria Disease, Disorder, or Finding Influence Bacteria Gene change expression Bacteria Human / Animal located in Bacteria Microbiome part of Chemical Anatomical Loc. / Human / Animal located in Chemical Chemical interact / part of Chemical Microbiome impact / produced by Chemical / Dietary Supp. / Drug / Food Bacteria / Microbiome impact Chemical / Dietary Supp. / Food Disease, Disorder, or Finding influence Chemical / Dietary Supp. / Drug / Food Gene change expression Chemical / Dietary Supp. / Drug / Food Human / Animal administered Disease, Disorder, or Finding Anatomical Location strike Disease, Disorder, or Finding Bacteria / Microbiome change abundance Disease, Disorder, or Finding Chemical interact Disease, Disorder, or Finding Disease, Disorder, or Finding affect / is a Disease, Disorder, or Finding Human / Animal target Drug Chemical / Drug interact Drug Disease, Disorder, or Finding change effect Human / Animal / Microbiome Biomedical Technique used by Microbiome Anatomical Loc. / Human / Animal located in Microbiome Gene change expression Microbiome Disease, Disorder, or Finding is linked to Microbiome Microbiome compared to GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 21
4.2. General Relation Annotation Rules ➔Relation types to annotate ◆Annotate relations between entities only if they match the defined set of relations provided for the GutBrainIE dataset. ◆The head of a relation does not always precede the tail in the text; it may also come after the tail. Annotators should ensure that the correct entities are linked regardless of their order in the text. Examples: ●In “Depression has been historically treated through antidepressant medications. Nowadays, probiotics supplementation is being used to…” two change effect relations should be annotated, one from “antidepressant medications” (drug) and “depression” (DDF), and the other from “probiotics supplementation” (dietary supplement) to “depression” (DDF). ◆Annotate relations that are directly and explicitly mentioned in the text. Example: ●In “Firmicutes influences predisposition to major depressive disorder” an influences relation between “Firmicutes” (bacteria) and “major depressive disorder” (DDF) should be annotated ◆Annotate relations implied by the text, even if the specific term describing the relation is not used. Example: ●In “Firmicutes impact inflammation in the gut” an influences relation should be annotated between “Firmicutes” (bacteria) and “inflammation” (DDF), as it is implied by the verb “impact”. ◆Annotate relations that can be inferred from the context. Relations inferred by reasoning about the text are valid as long as there is clear contextual evidence in the text and no personal, previous, or external knowledge is used for that inference. Example: GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 22
●In “Firmicutes change the expression of gene ARID1B and, therefore, affect predisposition to Autism spectrum disorder (ASD) in young patients” it can be inferred that “Firmicutes” (bacteria) is related to “Autism spectrum disorder (ASD)” (DDF) since it plays a role in affecting the “gene ARID1B” (gene) related to its predisposition. Therefore, two relations should be annotated In this portion of text: the explicit one change expression between “Firmicutes” (bacteria) and “gene ARID1B” (gene), and the inferred one influence between “Firmicutes” (bacteria) and “Autism spectrum disorder (ASD)” (DDF). ●In “Bacteroides fragilis produces gamma-aminobutyric acid (GABA). GABA has been found to reduce anxiety in mice.” it can be inferred, although not explicitly stated, that “Bacteroides fragilis” (bacteria) has a role in reducing “anxiety” (DDF) through its production of “gamma-aminobutyric acid (GABA)” (chemical). Therefore, an influence relation can be inferred between "Bacteroides fragilis" (bacteria) and "anxiety" (DDF). ➔Contextual requirements for annotating relations ◆Co-occurrence is not required. The entities involved in a relation do not need to co-occur in the same sentence. Relations can be annotated even if the entities are in different parts of the abstract, as long as there is sufficient contextual information to support the relation. Example: ●In “The gut microbiome is known to produce short-chain fatty acids. [...] The latter populate the intestinal barrier and play a crucial role in maintaining its integrity.” an inferred produced by relation can be annotated between “short-chain fatty acids” (chemical) and “gut microbiome” (microbiome) from the first sentence, and a located in relation between “short-chain fatty acids” (chemical) and “intestinal GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 23
barrier” (anatomical location) can be annotated although they do not co-occur in the same sentence. ◆Avoid overgeneralization, namely, only annotate relations between specific entity instances if there is explicit, implicit, or inferred evidence of their connection within the abstract. Do not assume that all occurrences of the same entity pair have a relation unless it is explicitly or implicitly described. Example: ●If one mention of “gut microbiome” says “The gut microbiome is associated with reduced symptoms of Parkinson's disease” and another simply states “Parkinson's disease is prevalent among the elderly” only the former should be used to annotate the relation is linked to. ◆If there is a relation from a long version of an entity to another entity, the same relation should be annotated from the short/acronym version of the entity to the same target entity. Example: ●In “Selective serotonin reuptake inhibitors (SSRIs) have an important role in the pathogenesis of depression.” two relations change effect should be annotated, one from “Selective serotonin reuptake inhibitors” (drug) to “depression” (DDF), and a second one from “SSRIs” (drug) to “depression” (DDF). ➔Annotation context and scope ◆Annotators should determine if a relation between two entities holds by only considering the information presented in the abstract. Annotators are not allowed to access the full text of the article nor to use any external sources of information. If the abstract is not clear about the relationship, DO NOT annotate it. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 24
➔Doubtful cases ◆If annotators are uncertain whether a relation exists between two entities, they should annotate conservatively. Only annotate when there is sufficient evidence to support a direct, implied, or inferred relationship. GutBrainIE, Task 6 of the BioASQ Lab @ CLEF25 25