NERdME (Version 1.1): a Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
Abstract
NERdME is a scholarly information extraction dataset designed for named entity recognition (NER) from GitHub README files. It contains 200 manually annotated README documents covering ten entity types relevant to scholarly and technical domains, such as software, datasets, and evaluation metrics. The annotations capture entities embedded in free text written in Markdown format, enhancing the accessibility and reusability of research artifacts shared on GitHub.
Full text
NERdME - Annotation Guidelines Version 1.1 Major Changes in This Version •Added a general rule about description terms. •Added specifications to Ontology. •Removed two entities types: ACADDISC and DSOTHER. 1 Introduction This document outlines the annotation guidelines for the NERdME Named Entity Recognition (NER) task. The objective of this task is to detect entities mentioned within selected github README documents and classify them according to their respective types, as defined by the NFDI4DS Ontology 1. Clear definitions accompanied with examples for 10 distinct entity types are given to assist annotators in comprehending the task and performing annotations effectively. Note that we keep the contents as well as the markups in the README files to preserve the structured information in the markups for the following reasons. Markdown elements such as headers (#), links ([text](url)), and formatting (e.g., bold) provide valuable context for annotation. Structural features like headers can be leveraged to identify important entities, as they often highlight key topics. Inline styles (e.g., bold or italics) can indicate emphasis, suggesting that bolded entities are likely more significant. Links and code snippets also serve as useful metadata, potentially pointing to named entities such as repositories or external resources. Additionally, entities within the URLs themselves (e.g., repository names or topics embedded in links) should be annotated, as they may represent important references. Considering the complexity of the annotation task, in addition to the examples provided in the guidelines, we include a trial step where all annotators receive a specially prepared trial ReadME file. We emphasize that annotators 1https://ise-fizkarlsruhe.github.io/NFDI4DS-Ontology/ 1
should thoroughly read the entire guideline to fully understand the task and report any issues they encounter while annotating the trial files. All submitted trial annotations are reviewed, and annotators are provided with feedback to help refine their understanding of the task. This ensures consistency and accuracy in the subsequent annotations. Annotation Tool: In this task, the INCEpTION 34.5 2tool [1] will be used to perform both annotation tasks. Figure 1 illustrate general steps of NER Annotation with INCEpTION. 2 General Annotation Rules The following guidelines apply to the list of entities in section 3: Lexical Nature: All named entities must include a proper name 3 or a definite description having the status of a proper name. A proper name is a unique identifier to designate a specific entity, such as a dataset (e.g. BookSum), license, or project. Proper names distinguish a particular object, concept, or entity from others of the same category. In English and many other languages, proper names are associated with capitalization, but may appear in various formats depending on the style and structure of the text. For example, in the title “NDFIcore Ontology”, ontology is NOT part of the proper name NFDIcore. Descriptive Terms: Annotate the entire name but exclude generic descriptors if they are standalone. Examples: –“Dataset” in “BookSum Dataset” should NOT be annotated. –“License” in “MIT License” should be annotated. –See also the example in Figure 2. Determiners: Determiners, such as articles (the,a), demonstratives (this,that), quantifiers (all), and possessives (their) - are not included when annotating proper nouns, as in the example in Figure 2. Figure 2: Annotating a conference entity in the text “... the European Semantic Web Conference” However, when the determiner appears in publication titles, it is included in the span of the entity, as in “A Survey on...” shown in Figure 3. 2https://inception-project.github.io/ 3https://en.wikipedia.org/wiki/Proper_noun 2
1. Select a string. Spans can be overlapped. 2. Choose a class 3. Check if there are more pages 4. Finish a file 5. Go to next file The annotation guidelines. Figure 1: Screenshot of example using INCEpTION. 3
Figure 3: Annotating a publication entity in the text “title=A survey on intelligent transportation systems, author={Qureshi, Kashif Naseer and Abdullah, Abdul Hanan}, ...” Punctuation Marks: Punctuation marks are to be included only when they are part of named entities, such as abbreviations and publication titles. Examples: –An entity with (:) punctuation as shown in Figure 4 Figure 4: Annotating a publication entity in the text “Zhang, Yi, et al. ”PAV-SOD: A new task towards panoramic audiovisual saliency detection.” ... ” –An entity without the punctuations (*) as shown in Figure 5: Figure 5: Annotating an entity of type Dataset in the text “... is the ***FewEvent*** dataset for the paper accepted by ... ” Enumerations: Entity mentions that are part of a list should be separately annotated. An example is given in Figure 6. Figure 6: Annotating an entity of type Programming Language in the text “ Programming languages such as Python, PHP, and C++ ” Nested Entities: A nested entity occurs when one entity is fully contained within another, creating a hierarchical structure. In such cases, annotate both the shorter (nested) span and the longer span. An example is provided in Figure 7. Figure 7: An example showing nested entities 4
Discontinuous Entities : For this task, entity spans must be continuous. Discontinuous spans (e.g., parts of an entity separated by non-entity tokens) are not annotated or treated as valid entities. Acronyms: Entity mentions given as long form (acronym) or acronym (long form) should be annotated separately as shown in Figure 8. Figure 8: An example showing how to annotate acronyms Nested Codes: Readme files can contain nested code blocks or scripts that provide execution instructions or citations. All entities mentioned within these code blocks must be annotated as shown in Figure 9. Figure 9: An example showing how to annotate entities embedded in a code block. The code block starts and ends with ``` URLs: Entities enbedded in URLs should also be annotated, as in the example shown in Figure 10. Figure 10: An example showing how to annotate an entity embedded in the URL https://ai.meta.com/llama/ 5
Websites: Mentions of websites should NOT be annotated because they are not directly related to data science. Dates and Versions: When annotating entities of the following types: conference, workshop, publication, dataset, license, software, or programming language, always include date or version information directly associated with the entity mention. Examples are provided in Section 3. Single-Class Annotation: The sense of an entity should be determined by the context, e.g., python can be referred to as the python programming language or the python command (software). One entity mention should NOT be annotated by more than one class. 3 Entity Types The following classes of the NFDI4DS Ontology are to be used to annotate entity mentions. For each class/type, detailed information is provided, including: the URI (an identifier that directs you to the ontology where it is defined, should you wish to explore further), the tag (for use when annotating in INCEpTION), a definition, and some examples, are provided. 1. Conference •IRI: https://nfdi.fiz-karlsruhe.de/nfdi4dso/Conference •Tag: CONFERENCE •Definition: A conference event. •Examples: refer to Figure 11. Figure 11: Examples of entity mentions of type Conference 2. Dataset •IRI: https://nfdi.fiz-karlsruhe.de/ontology/Dataset •Tag: DATASET •Definition: A creative work that refers to a structured collection of data, organized typically for a specific goal such as analysis, research, or reference. Structured information about a resource provided by an organization or a person. •Rules: 6
–Data sources such as Wikidata, DBpedia, Reddit, and other possible explicit web or unstable, or NOT static sources, are not to be annotated as datasets as in the GSAP annotation guidelines4. –Generic datasets, i.e., mentions of datasets that refer to specific named entities but do not explicitly include the dataset’s official or proper name should NOT be annotated as a “Dataset”. Examples of mentions NOT to annotate: Generic mentions like “the DBpedia dataset”, “the Wikipedia datasets”, or “the Wikidata datasets”, which are descriptive references rather than explicit names. •Examples refer to Figure 12. Figure 12: Examples of entity mentions of type Dataset 3. Evaluation Metric •IRI: https://nfdi.fiz-karlsruhe.de/nfdi4dso/EvaluationMetric •Tag: EVALMETRIC •Definition: A quantitative measure used to assess the performance and effectiveness of a statistical or machine learning model. •Examples refer to Figure 13. Figure 13: Examples of entity mentions of type Evaluation Metric 4. License •IRI: https://nfdi.fiz-karlsruhe.de/ontology/License •Tag: LICENSE 4https://data.gesis.org/gsap/gsap-ner/documents/GSAP-NER-Annotation-Guideline-Version-1-2. pdf 7
•Definition An information content entity that refers to a legal instrument (usually by way of contract law, with or without printed material) governing the use or redistribution of the resource containing the licence. A document that provides legal guidelines for the use of a resource, e.g. a software. •Examples refer to Figure 14. Figure 14: Examples of entity mentions of type License 5. Ontology •IRI: https://nfdi.fiz-karlsruhe.de/ontology/Ontology •Tag: ONTOLOGY •Definition: Semantic expressivity that refers to a semantic framework that represents knowledge about a domain, including concepts, entities, properties, and relationships, in a structured and machine-understandable manner. A semantic framework that represents knowledge about a domain, including concepts, entities, properties, and relationships, in a structured and formalized manner, enabling machine-understandable reasoning and inference. •Rules: –Knowledge graphs should not be considered ontology. –Ontologies have no data instances. –In case of doubt, please look up the named entity on the internet. •Examples refer to Figure 15. Figure 15: Examples of entity mentions of type Ontology 6. Programming Language •IRI: https://nfdi.fiz-karlsruhe.de/ontology/ProgrammingLanguage •Tag: PROGLANG •Definition: A technological means that is used for implementing software. A formal language used for implementing a software. •Examples refer to Figure 6. 7. Project 8
•IRI: https://nfdi.fiz-karlsruhe.de/ontology/Project •Tag: PROJECT •Definition A planned process of a scientific or business endeavor that aims to conclude an investigation or to answer a research question. A scientific or business endeavor that aims to conclude an investigation or to answer a research question. •Examples: refer to Figure 16. Figure 16: Examples of entity mentions of type Project 8. Publication •IRI: https://nfdi.fiz-karlsruhe.de/ontology/Publication •Tag: PUBLICATION •Definition: Any creative work that is the output of a publishing process, except for datasets. A published scholarly work that reports on ongoing activity about or within a resource, e.g., proceedings of a conference, a journal article, or a preprint. •Rules: Only the title of the publication should be annotated. Exception: Conference names without mentioning ”Proceedings of, appear in booktitle or journal in BibTex citation blocks should NOT be annotated as PUBLICATION”. •Examples refer to Figure 17. Figure 17: Examples of entity mentions of type Publication 9. Software •IRI: https://nfdi.fiz-karlsruhe.de/ontology/Software •Tag: SOFTWARE •Definition A creative work that comprises a set of instructions, programs, or algorithms designed to perform specific tasks or functions on a computer or other electronic devices. A computer software provided by an organization or a person. 9