scieee AI-readable full text Open interactive document viewer

GREI Metadata Recommendations from DataCite schema version 4.6

Hahnel, Mark; Chodacki, John; Neumann, Stan; Nielsen, Lars Holm; Gautier, Julian; Hamelers, Audrey; Gueguen, Gretchen; Call, Mark; Scherer, David

Abstract

This document, version 2 of the GREI Metadata Recommendations, provides updated guidelines for generalist data repositories to standardize their metadata using the DataCite Metadata Schema 4.6 and based on community feedback received on version 1. The goal is to enhance the interoperability and discoverability of datasets, particularly those from NIH-funded projects. The recommendations strongly encourage repositories to collect a specific subset of metadata properties and use designated vocabularies and values to support common use cases for sharing, discovering, and tracking the impact of data. The core of the recommendation is a table that lists the DataCite metadata properties that the GREI repositories have agreed are most essential to FAIR data sharing, specifying recommended values and vocabularies for each field where applicable. Key recommendations include using DataCite-generated DOIs for identifiers, ORCID iDs for personal name identifiers, and ROR IDs for organizational identifiers and funder information. It also specifies using the value Dataset for resourceTypeGeneral, and Issued for the publication date. Additional subsections of the recommendation include: GREI Recommended Subset of Relation Types: This section outlines a subset of DataCite's relation types that are most useful for common interactions, providing both the DataCite relationType and a researcher-friendly label. Recommendations for Controlled Vocabulary Integration in Generalist Repositories: This section recommends using controlled vocabularies and lists quality criteria for selecting them, such as governance, coverage, and persistence. Proposed CRediT Role Subset for Generalist Data Repositories: This section proposes a subset of the official NISO CRediT roles that are most applicable to the lifecycle of dataset generation, curation, and publication. ORCID Recommendations: This section provides best practices for collecting, authenticating, and pushing ORCID iDs to improve attribution and metadata quality. Recommended Usage of Dates in GREI Metadata: This section provides guidance on using DataCite's Date and dateType fields, emphasizing that Issued must be used for the publication date and that other dates can provide additional context.

Full text

GREI Metadata Recommendations from DataCite schema version 4.6 Version 02: Last updated 2025-08-14 Overview 2 Recommendation 2 GREI Recommended Subset of Relation Types 4 Recommendations for Controlled Vocabulary Integration in Generalist Repositories 6 Controlled Vocabularies Quality Criteria 6 Implementing Controlled Vocabularies in the Repository Interface 6 Leveraging the Metadata Center and Community Efforts 6 MeSH for Subject Classification for NIH-Funded Research 6 Complementary Vocabularies by Domain 7 Proposed CRediT Role Subset for Generalist Data Repositories 7 Current Authoritative CRediT Roles (NISO) 7 Proposed Subset for Generalist Repositories 8 ORCID Recommendations 9 Collection of ORCID iDs 9 Authentication of ORCID iDs 9 Push to ORCID Registry 9 Recommended Usage of ‘Dates’ in GREI Metadata 9 Issued - Labeling the Publication Date 9 Recommended Additional Dates 10 Implementation Tips 10 What to Avoid 10 Example Metadata Block 11 Specific Use Case: When Metadata Are Public Before the Data 11 Step 1. Submitted 11 Step 2. Issued 11 Step 3. Available 12 Optional: Add an Embargo Note 12 Summary of ‘Dates’ Recommendations 12 Feedback on Recommendations 12 Next Steps 13 Overview One goal of GREI is to support interoperability and discovery of datasets across repositories by establishing common metadata standards for the generalist repositories. Having focused on an agreed standard, the DataCite Metadata Schema 4.6, the GREI Metadata working group has set its Year 4 goal for repositories to build on their existing work on metadata for research datasets. Version 1 of these metadata recommendations strongly encourage repositories to collect specific metadata when researchers publish their research outputs, including data from NIH-funded projects. It can be found online at https://doi.org/10.5281/zenodo.8101956. The GREI Metadata Task Group has continued to press for better interoperability as the project continues. As such, the working group has created this V2 set of recommendations to strongly encourage that each repository member collect the following metadata to support the generalist repository use cases for sharing, discovering and tracking the impact of data. We also hope this common metadata schema will be useful for data repositories beyond GREI to improve interoperability across data repositories and across the NIH data landscape. Recommendation The document lists strongly encouraged metadata to be collected by each GREI repository in alignment with the metadata collected by DataCite’s optional metadata properties. Where applicable, the values and vocabularies that repositories are encouraged to use have also been reviewed by the subcommittee and included in the recommendations. While these are the recommendations of the subcommittee to-date, the goal of the subcommittee is for the recommended common schema to evolve based on community feedback, new use cases, and emerging community initiatives to allow for updates and additions to the strongly recommended fields in the future. For DataCite’s property definitions, cardinality/quantity constraints, and other usage notes, see DataCite Metadata Schema 4.6 Page 2 of 14 DataCIte element ID DataCite-Property Subcommittee recommendations about property values 1 Identifier (Mandatory) DataCite generated DOI 2 Creator (Mandatory) 2.1.a nameType Use controlled list values from DataCite: “Personal” or “Organizational”. 2.2 givenName 2.3 familyName 2.4 nameIdentifier For names of people, use ORCID IDs. For names of organizations, use ROR IDs. 2.4.a nameIdentifierScheme For names of people, use “ORCID”. For names of organizations, use “ROR”. 2.4.b schemeURI For names of people, use ORCID’s URI, https://orcid.org/ For names of organizations, use ROR’s URI, https://ror.org/ 2.5 affiliation Use names of organizations from the Research Organization Registry (ROR). 2.5.a affiliationIdentifier Use IDs from ROR in URL format (e.g., https://ror.org/05h1kgg64). 2.5.b affiliationIdentifierScheme Use “ROR”. 2.5.c schemeURI Use ROR’s URI, https://ror.org/ 3 Title (Mandatory) 4 Publisher (Mandatory) 5 PublicationYear (Mandatory) 6 Subject The appropriate controlled vocabulary depends on the domain covered by the dataset; see the discussion below. 6.a subjectScheme 6.b schemeURI Page 3 of 14 6.c valueURI 8 Date 8.a dateType Use controlled list values from DataCite. All datasets should contain the date type Issued to indicate the publication date; other dates may provide additional context, such as Available or Collected. 8.b dateInformation 10 ResourceType 10.a resourceTypeGeneral Use controlled list values from DataCite; for datasets the appropriate value is “Dataset” 12 RelatedIdentifier 12.a relatedIdentifierType Use controlled list values from DataCite 12.b relationType Use controlled list values from DataCite; the recommended values and the recommended terminology to use in the user interface is described below 12.f resourceTypeGeneral Use controlled list values from DataCite. 15 Version 16 Rights 16.a rightsURI 16.b rightsIdentifier 16.c rightsIdentifierScheme 16.d schemeURI 17 Description 17.a descriptionType For general descriptions of the deposit, select the descriptionType “Abstract” from DataCite’s controlled list. 19 FundingReference 19.1 funderName 19.2 funderIdentifier Use IDs from ROR. Page 4 of 14 19.2.a funderIdentifierType Select “ROR” from DataCite’s controlled list. 19.2.b schemeURI Use https://ror.org/ 19.3 awardNumber 19.4 awardTitle GREI Recommended Subset of Relation Types The DataCite controlled list of types is very thorough and flexible; in practice a subset of this list is sufficient for most purposes. In addition, the names in the DataCite Controlled list are appropriate as an enumerated list for encoding metadata, but may not be appropriate for interactions with users. Accordingly, the table below provides the recommended types and terminology to use for most common interactions. Researcher-Friendly Label DataCite relationType Use Case Cited by this article IsCitedBy Dataset is cited in an article, that is not the original description of the dataset Cites this work Cites Dataset cites or references a paper or another dataset Supports this material IsSupplementTo Dataset is supplementary to other material (e.g. publication) Supported by these materials IsSupplementedBy Other materials (e.g. software) are supporting this dataset Part of a larger project IsPartOf Dataset belongs to a broader collection, grant, or study Includes other parts HasPart Dataset contains related datasets or components Updated version of IsNewVersionOf Dataset is a new version of a previously published dataset Older version of IsPreviousVersionOf Dataset is a previous version of a new published dataset Page 5 of 14 Has Version HasVersion As a canonical dataset (e.g. for an ongoing observational study or other research project) the Dataset may have versions of the data with version-specific DOIs Is a Version of IsVersionOf The dataset version is a specific version of a canonical dataset Duplicate of another record IsIdenticalTo Dataset exists in multiple repositories or records Collected using this instrument IsCollectedBy Dataset gathered by specific instrument These values are intuitive, align well with scholarly practices, and are supported by tools such as DataCite Event Data and initiatives such as Make Data Count. Recommendations for Controlled Vocabulary Integration in Generalist Repositories We recommend the use of controlled vocabularies wherever possible. The criteria listed below can help in the selection of a vocabulary. We further recommend the use of a broad, general vocabulary to facilitate comparisons across domains, such as the OECD Fields of Science and Technology. A more granular, domain-specific vocabulary, such as MeSH, can also be employed as necessary. Controlled Vocabularies Quality Criteria ● Governance: Vocabulary has clear governance, maintenance processes, and versioning ● Coverage: Provides adequate coverage of relevant domains ● Granularity: Offers appropriate level of specificity for the content being described ● Currency: Regularly updated to reflect evolving terminology ● Persistence: Maintains stable identifiers with a commitment to long-term support ● Machine-actionability: Available in machine-readable formats with semantic relations ● Documentation: Well-documented with clear definitions and usage guidelines Page 6 of 14 Implementing Controlled Vocabularies in the Repository Interface The technical implementation of controlled vocabularies significantly impacts their usability and effectiveness. Repositories should aim for implementations that: ● Provide Search and Browse Functionality: Allow users to search and browse datasets using controlled vocabulary terms. ● Offer Auto-completion or Suggestion Features: Assist depositors in selecting terms by providing suggestions from the controlled vocabularies as they type. ● Clearly Display Applied Terms: Make the controlled vocabulary terms used to index a dataset easily visible to users. Leveraging the Metadata Center and Community Efforts The CEDAR Metadata Center is a valuable resource for information on metadata standards and best practices, including those related to controlled vocabularies. GREI repositories should explore the resources and guidance available through the Metadata Center to inform their controlled vocabulary strategies. MeSH for Subject Classification for NIH-Funded Research Medical Subject Headings (MeSH) vocabulary, maintained by the National Library of Medicine. ● Recommended Vocabulary: MeSH (Medical Subject Headings) ● Scheme URI: https://meshb.nlm.nih.gov MeSH is widely adopted in biomedical literature and offers broad, hierarchical coverage of health and life sciences topics. Its use will support alignment with other NIH resources, including PubMed and clinicaltrials.gov. Importantly, we propose MeSH as a recommendation, not a requirement, to allow flexibility for datasets outside the biomedical domain or for repositories already using other structured vocabularies. Complementary Vocabularies by Domain Repositories should support additional vocabularies based on the domains they serve. For example: ● Gene Ontology (GO): For gene functions, cellular components, and biological processes ● NCBI Taxonomy: For organism classification ● Chemical Entities of Biological Interest (ChEBI): For chemical compounds ● Disease Ontology: For human disease concepts Page 7 of 14 For specialized needs, consider: ● HIVE (Helping Interdisciplinary Vocabulary Engineering): Provides access to multiple vocabularies through a single interface Proposed CRediT Role Subset for Generalist Data Repositories GREI repositories increasingly seek to recognize the varied contributions involved in dataset creation, curation, and sharing. The CRediT taxonomy provides a standardized framework for contributor roles, offering a path toward better attribution and metadata consistency. Some efforts have happened to customize CrediT to data use cases, such as the WSL Data Credit project. However, the authoritative list of 14 roles adopted by NISO do not specifically state how they map into our space. This document proposes a subset of official CRediT roles that are most relevant for generalist data repositories to implement. Current Authoritative CRediT Roles (NISO) 1. Conceptualization 2. Data curation 3. Formal analysis 4. Funding acquisition 5. Investigation 6. Methodology 7. Project administration 8. Resources 9. Software 10. Supervision 11. Validation 12. Visualization 13. Writing – original draft 14. Writing – review & editing Proposed Subset for Generalist Repositories We recommend supporting the following subset of roles, which most directly apply to the lifecycle of dataset generation, curation, and publication: Page 8 of 14 ● Conceptualization – Identifying the research goals behind the dataset. ● Data curation* – Managing, cleaning, and preparing data for sharing and reuse. ● Investigation* – Performing the experiments or data collection. ● Methodology – Designing the processes that guided data collection or analysis. ● Project administration – Overseeing project execution, including data management planning. ● Software* – Writing code or scripts used to generate or prepare the data. ● Supervision* – May refer to oversight of data collection or project execution, typically by senior researchers, PIs, or data managers. However, this role is often implicit or institutional rather than individually documented during dataset submission. ● Writing – original draft – Drafting associated documentation, such as README files or metadata. ● Writing – review & editing – Reviewing or improving documentation and descriptive metadata. The following roles may be relevant in specific contexts, depending on repository policies or contributor preferences: ● Formal analysis – May be useful for attributing those who derived or created secondary datasets. ● Funding acquisition – Often not applicable unless funders require attribution or the principal investigator’s role is explicitly recognized. ● Resources – May apply when someone acquires equipment, recruits subjects, or provisions instruments used in data generation. ● Validation* – May involve confirming the quality, accuracy, or reproducibility of data. While important, formal validation is not always part of the dataset deposit process. ● Visualization – May refer to generating visual representations of data (e.g., charts, maps, dashboards). These are often created outside the repository workflow or auto-generated, making attribution inconsistent. The ones denoted with * are those that map directly to the WSL Data Credit project. The first subset aims to strike a balance between completeness and feasibility, making it easier to implement structured attribution without overburdening depositors. However, the full list can also be used. ORCID Recommendations Page 9 of 14