Digital repository capabilities and characteristics mapping report
Abstract
This report documents the results of the desk-based mapping of repository capabilities and characteristics carried out in Task 5.2 of the FIDELIS project. Guided by the FIDELIS Transparent Trustworthy Repository Attributes Matrix (TTRAM), the report analyses domain-agnostic and domain-specific resources across five scientific communities (Agri-food, Biomedical Sciences, Climate Science, Linguistics, and Social Sciences). It identifies common practices, gaps, and opportunities for support, providing key insights to inform the development of the FIDELIS Network of Trustworthy Digital Repositories.
Full text
Project Title FIDELIS: Establishing A European Network of Trustworthy Digital Repositories Project Acronym FIDELIS Grant Agreement No. 101188078 Start Date of Project 2025-01-01 Duration of Project 36 months Project Website https://eden-fidelis.eu/ Digital repository capabilities and characteristics mapping report Work Package WP5 - Building common understanding of TDRs and a harmonised matrix of repository capabilities and characteristics Lead Author (Org) Philipp Conzett (UiT The Arctic University of Norway, 0000-0002-6754-7911) Contributing Author(s) (Org) Severine Duvaud (SIB Swiss Institute of Bioinformatics, 0000-0001-7892-9678) Thomas Jouneau (INRAE/Université de Lorraine, 0000-0001-5986-8128) Joel Kallio (Tampere University, Finnish Social Science Data Archive, 0009-0003-5076-9985) Terje Klemetsen (UiT The Arctic University of Norway, 0000-0002-2024-1798) Jonas Recker (GESIS - Leibniz Institute for the Social Sciences, 0000-0001-9562-3339) Pavel Straňák (Charles University, 0000-0002-6895-8536) Sanni Tujunen (Tampere University, Finnish Social Science Data Archive, 0009-0003-9076-8397) Mari Kleemola (Tampere University, Finnish Social Science Data Archive, 0000-0001-8855-5075) Henna Kaartinen (Tampere University, Finnish Social Science Data Archive, 0009-0001-2148-1822) Date 2025-10-29 Version V1.0 DOI https://doi.org/10.5281/zenodo.17471910 1 | Page
Dissemination Level X PU: Public PP: Restricted to other programme participants (including the Commission) RE: Restricted to a group specified by the consortium (including the Commission) CO: Confidential, only for members of the consortium (including the Commission) Versioning and contribution history Version Date Author Notes 0.1 2025.02.19 Philipp Conzett (UiT) First draft of spreadsheet for resource mapping 0.2 2025.06.25 Philipp Conzett (UiT), Severine Duvaud (SIB Swiss Institute of Bioinformatics), Terje Klemetsen (UiT), Sanni Tujunen (TAU-FSD), Joel Kallio (TAU-FSD), Jonas Recker (GESIS) First resources added to spreadsheet 0.3 2025.07.09 Philipp Conzett (UiT) Aligned spreadsheet with TTRAM v1.0 0.4 2025.05.28 Philipp Conzett (UiT) Outline and ToC of report 0.5 2025.10.17 Philipp Conzett (UiT), Severine Duvaud (SIB Swiss Institute of Bioinformatics), Terje Klemetsen (UiT), Sanni Tujunen (TAU-FSD), Joel Kallio (TAU-FSD), Jonas Recker (GESIS) First draft of Results section 0.6 2025.10.20 Philipp Conzett (UiT) Revised Results section and created first complete draft of report. 0.7 2025.10.26 Severine Duvaud (SIB), Terje Klemetsen (UiT), Sanni Tujunen (TAU-FSD), Joel Kallio (TAU-FSD), Jonas Recker (GESIS), Thomas Jouneau (INRAE) Reviewed complete draft. Revised Agri-Food sections, added methodology 2 | Page
0.8 2025.10.26 Philipp Conzett (UiT) Reviewed revised version and created final draft 0.9 2025.10.28 Mari Kleemola (TAU-FSD) WPL review 1.0 2025.10.29 Philipp Conzett (UiT) Final version Disclaimer FIDELIS has received funding from the European Commission’s Horizon Europe funding programme for research and innovation programme under the Grant Agreement no. 101188078. The content of this document does not represent the opinion of the European Commission, and the European Commission is not responsible for any use that might be made of such content. 3 | Page
Table of Contents Executive Summary.................................................................................................................................7 1. Introduction........................................................................................................................................ 8 2. Methodology.....................................................................................................................................10 3. Results...............................................................................................................................................14 3.1. Context.................................................................................................................................... 14 3.1.1. IDENTIFICATION & CONTACT (AF01).............................................................................. 14 3.1.2. MISSION & SCOPE (AF02)...............................................................................................16 3.2. Digital Object Management.....................................................................................................22 3.2.1. CONCEIVE, CREATE, COLLECT (AF03)..............................................................................22 3.2.2. DEPOSIT & APPRAISAL (AF04)........................................................................................25 3.2.3. CURATION, QUALITY & COMPLIANCE (AF05).................................................................31 3.2.4. DISCOVERY & IDENTIFICATION (AF06)........................................................................... 38 3.2.5. ACCESS (AF07)................................................................................................................50 3.2.6. REUSE (AF08)................................................................................................................. 57 3.2.7. WORKFLOWS (AF09)...................................................................................................... 63 3.2.8. PRESERVATION (AF10)....................................................................................................66 3.2.9. PROVENANCE & AUTHENTICITY (AF11)......................................................................... 75 3.2.10. SUPPORT (AF12)...........................................................................................................80 3.3. Organisational Infrastructure.................................................................................................. 84 3.3.1. GOVERNANCE (AF13).....................................................................................................84 3.3.2. POLICY & STANDARDS (AF14).........................................................................................87 3.3.3. RIGHTS (AF15)................................................................................................................92 3.3.4. RESOURCES (AF16).........................................................................................................97 3.3.5. PEOPLE & EXPERTISE (AF17)........................................................................................ 101 3.3.6. THIRD PARTY DEPENDENCIES (AF18)........................................................................... 104 3.3.7. CONTINUITY OF SERVICE (AF19).................................................................................. 107 3.3.8. EXTERNAL ENGAGEMENT (AF20).................................................................................111 3.3.9. RELEASE & PUBLISHING (AF21)....................................................................................115 3.3.10. INTEROPERABILITY (AF22)..........................................................................................117 3.3.11. LEGAL & ETHICAL (AF23)............................................................................................121 3.3.12. CRITERIA, ASSESSMENT, IMPROVEMENT (AF24)....................................................... 126 3.3.13. ANALYSIS & IMPACT (AF25)........................................................................................130 3.3.14. TRAINING (AF26)........................................................................................................132 3.3.15. RESEARCH & DEVELOPMENT (R&D) (AF27)............................................................... 135 3.4. Technology.............................................................................................................................137 4 | Page
3.4.1. STORAGE & INTEGRITY (AF28)..................................................................................... 137 3.4.2. TECHNICAL INFRASTRUCTURE (AF29)..........................................................................142 3.5. Security..................................................................................................................................146 3.5.1. SECURITY (AF30).......................................................................................................... 146 3.6. Other capabilities and characteristics................................................................................... 150 4. Analysis and Discussion...................................................................................................................150 4.1. Resource Characteristics........................................................................................................150 4.1.1. Network Beacon Communities.................................................................................... 150 4.1.2. Resource Types............................................................................................................ 151 4.2. Context.................................................................................................................................. 152 4.2.1. Identification & Contact (AF01)....................................................................................152 4.2.2. Mission & Scope (AF02)............................................................................................... 152 4.3. Digital Object Management...................................................................................................153 4.3.1. Commonalities Across Communities and Resource Types...........................................154 4.3.2. Specificities by Network Beacon Community.............................................................. 154 4.3.3. Resource Types............................................................................................................ 155 4.3.4. Discussion.................................................................................................................... 157 4.4. Organisational Infrastructure................................................................................................ 157 4.4.1. Commonalities Across Communities and Resource Types...........................................158 4.4.2. Specificities by Network Beacon Community.............................................................. 158 4.4.3. Resource Types............................................................................................................ 159 4.4.4. Discussion.................................................................................................................... 161 4.5. Technology.............................................................................................................................161 4.5.1. Storage & Integrity (AF28)............................................................................................161 4.5.2. Technical Infrastructure (AF29)....................................................................................162 4.5.3. Discussion.................................................................................................................... 163 4.6. Security..................................................................................................................................163 4.6.1. Commonalities Across Communities and Resource Types...........................................163 4.6.2. Specificities by Network Beacon Community.............................................................. 163 4.6.3. Resource Types............................................................................................................ 164 4.6.4. Discussion.................................................................................................................... 165 5. Summary and Main Implications for FIDELIS.................................................................................. 166 5.1.1. Key Findings................................................................................................................. 166 5.1.2. Implications for FIDELIS................................................................................................166 6. References.......................................................................................................................................168 5 | Page
TERMINOLOGY 6 | Page Terminology/Acronym Description AF Activity or Function drawn from the reference list of Activities and Functions in TTRAM API Application Programming Interface CARE Principles for Indigenous Data Governance: Collective benefit, Authority to control, Responsibility, Ethics CMDI Component MetaData Infrastructure – a metadata framework used in CLARIN CTS CoreTrustSeal – a certification for trustworthy data repositories DMP Data Management Plan EOSC European Open Science Cloud – a pan-European initiative to provide researchers with access to FAIR data and services FAIR Findable, Accessible, Interoperable, Reusable – guiding principles for scientific data management and stewardship GDPR General Data Protection Regulation LoRCaP Levels of Retention, Curation and Preservation OAIS Open Archival Information System – a reference model for digital preservation OAI-PMH Open Archives Initiative Protocol for Metadata Harvesting – a protocol for harvesting metadata from repositories PID Persistent Identifier – a long-lasting reference to a digital object, such as a DOI or Handle TDR Trustworthy Digital Repositories TRUST Transparency, Responsibility, User focus, Sustainability, Technology – principles for trustworthy data repositories TTRAM Transparent Trustworthy Repository Attributes Matrix – a reference model developed in FIDELIS to better understand repositories and align repository capabilities Zenodo An open-access repository for research outputs, often used for publishing deliverables in EU projects
Executive Summary This report presents one of three main outcomes of Task 5.2 (T5.2) “Preparing for Federation” within the FIDELIS project, which aims to establish a European Network of Trustworthy Digital Repositories (TDRs) aligned with the European Open Science Cloud (EOSC). The task addresses the challenge of supporting digital repositories in improving their practices and achieving greater interoperability and alignment with EOSC requirements. To this end, T5.2 conducted a comprehensive mapping of repository capabilities and characteristics across five scientific communities – Agri-food, Biomedical Sciences, Climate Science, Linguistics, and Social Sciences – through desk research and analysis of existing standards, best practices, solutions, landscape analyses, etc. The mapping was guided by the Transparent Trustworthy Repository Attributes Matrix (TTRAM), a reference model developed in Task 5.1, which defines 30 repository Activities and Functions (AFs) across thematic areas such as digital object management, organisational infrastructure, technology, and security. The report identifies commonalities and specificities in repository practices, highlights gaps in current capabilities, and provides insights into areas where FIDELIS can offer targeted support. These findings will inform the development of recommendations and tools to facilitate the federation of repositories and enhance their trustworthiness, discoverability, and alignment with EOSC. 7 | Page
1. Introduction This report presents a key outcome of Task 5.2 (T5.2): Preparing for Federation, within the FIDELIS project. The overarching goal of T5.2 is to chart the landscape of European digital repositories and gather insights on how FIDELIS can best support their efforts to improve, align, and federate with the European Open Science Cloud (EOSC) as Trustworthy Digital Repositories (TDRs). The work in T5.2 is structured into two main components: 1. Community Survey: Conducted jointly with Task 8.1, this survey collected input from repository stakeholders regarding their current practices, challenges, and needs. Topics included technical and organisational standards, existing federation models, and best practices for managing digital objects. The survey also explored expectations for how the FIDELIS network could provide support. The results are published on Zenodo (Jouneau et al., 2025). 2. Desk Research and Mapping: This present report documents a desk-based mapping of current landscape analyses, best practices, standards, catalogues, registries, and other solutions used and needed by repositories. These are collectively referred to as “repository capabilities and characteristics”. The mapping exercise is guided by the FIDELIS Transparent Trustworthy Repository Attributes Matrix (TTRAM), developed in Task 5.1. TTRAM serves as a reference model for alignment and cooperation across repositories in the FIDELIS network and EOSC. It enables a detailed examination of the artefacts that support repository practices and provide evidence for assessment, such as standards, semantic artefacts, policies, and procedures. The first version of TTRAM is available on Zenodo (L’Hours et al., 2025a, 2025b, 2025c). TTRAM defines 30 repository activities and functions (AFs), grouped into four main areas: ● Context o AF01 Identification & Contact o AF02 Mission & Scope ● Digital Object Management o AF03 Conceive, Create, Collect o AF04 Deposit & Appraisal o AF05 Curation, Quality & Compliance o AF06 Discovery & Identification o AF07 Access o AF08 Reuse o AF09 Workflows o AF10 Preservation o AF11 Provenance and Authenticity o AF12 Support ● Organisational Infrastructure 8 | Page
o AF13 Governance o AF14 Policy & Standards o AF15 Rights o AF16 Resources o AF17 People & Expertise o AF18 Third Party Dependencies o AF19 Continuity of Service o AF20 External Engagement o AF21 Release & Publishing o AF22 Interoperability o AF23 Legal & Ethical o AF24 Criteria, Assessment, Improvement o AF25 Analysis & Impact o AF26 Training o AF27 Research & Development (R&D) ● Technology o AF28 Storage & Integrity o AF29 Technical Infrastructure ● Security o AF30 Security This report is organised as follows: ● Section 2 outlines the methodology used for mapping repository capabilities and characteristics. ● Section 3 presents the results, structured according to the 30 TTRAM AFs and grouped by domain-agnostic and domain-specific resources. Domain-specific resources are further categorised by five research communities: Agri-food, Biomedical Sciences, Climate Science, Linguistics, and Social Sciences. ● Section 4 discusses key aspects of the reviewed and mapped repository capabilities and characteristics presented in Section 3. ● Section 5 summarises the main takeaways from the mapping. 9 | Page
Lin et al., 2024: The article emphasizes the importance of clear repository metadata such as the repository name, the existence of persistent identifiers, the organizational affiliation as well as contact information for support or inquiries. Climate Science The Activity/Function (A|F) Identification & Contact is not addressed by any of the reviewed resources from Climate Science. Linguistics For CLARIN B centres, the following requirements apply (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 2.c Visibility of connection to CLARIN. Requirement: Each centre needs to refer to CLARIN in a visible way on its website. ● 2.f Registration in the Centre Registry. Requirement: Each centre should be registered in the Centre Registry. [...] Centre Registry information is automatically transmitted to re3data.org – a registry of research repositories. Social Sciences The CESSDA Brief (Kleemola et al., 2025) recommends that data repositories should register in relevant registries and include repository identifiers on websites. Examples mentioned include re3data, OpenDOAR, Fairsharing, and Research Organization Registry (ROR). 3.1.2. MISSION & SCOPE (AF02) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function Mission & Scope follows: “Defining the purpose (or mission) of the organisation and the boundaries of activities for which the repository is responsible. Where activities involve partners or outsourced services, the organisation retains ultimate responsibility for the activities, functions and outcomes. Provides reference information and context for the rest of the activities and functions.” (L’Hours et al., 2025c, p. 11) and suggests the following items for transparent information: ● Organisation Mission Statement ● The Levels of Retention, Curation and Preservation (LoRCAP) offered by the repository ● Repository Type ● Designated Community ● Permitted/Restricted Object, Depositor and User Characteristics ● Geographic and linguistic coverage ● Whether a repository is FAIR-enabling or endorses the TRUST or CARE Principles The A|F Mission & Scope is covered by CoreTrustSeal requirement Mission & Scope (R01) requiring certified repositories to have an explicit mission to provide access to and preserve digital objects (CoreTrustSeal Standards and Certification Board, 2022, p. 11). 16 | Page
The nestor Seal addresses aspects of this A|F in “C2 Responsibility for preservation” and “C3 Designated communities”. C2 requires that the archive “responsibility for the long-term preservation of the information objects on the basis of legal requirements or its own objectives”. C3 specifies the need to define the archive’s Designated Community, including (nestor Certification Working Group, 2025). In the OAIS Reference Model, an Open Archival Information System (OAIS) is defined as an organization which has “accepted the responsibility to preserve information and make it available for a Designated Community [...]. The system meets a set of mandatory responsibilities that allows an OAIS Archive to be distinguished from other uses of the term ‘archive’” (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 1, p. 13). The mandatory responsibilities (Section 3.2) have additional implications for what the mission of an OAIS must entail. This includes ● defining the community for whose understanding and use the information objects will be preserved, ● ensuring that the preserved objects are “Independently Understandable” to this community, ● the implementation of “documented policies and procedures which ensure that the information is preserved against all reasonable contingencies, including the demise of the Archive”, and ● making the preserved information objects available to the defined community (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 3, p. 1). The EOSC Federation Handbook does not address the A|F Mission & Scope explicitly, but provides the following high-level guidelines for FAIR data repositories (EOSC Association, 2025, pp. 32–33): EOSC Node data repositories should align with the FAIR principles and aim to implement them as rigorously as possible. [...] Data repositories should implement the EOSC Guidelines for Research Data Sources. The repositories must be onboarded as a service in one of the EOSC Nodes [...]. [...] In the cases where the EOSC Node repositories are not community standards, they should link to community standard repositories if they exist, to increase their impact and reduce proliferation of data repositories. The Desirable Characteristics of Data Repositories for Federally Funded Research outlines its mission as to improve consistency across U.S. Federal departments and agencies in guiding researchers on selecting repositories for Federally funded research data, aiming to ensure data is Findable, Accessible, Interoperable, and Reusable (FAIR) while integrating privacy and security. Its scope defines repositories' responsibilities in managing and sharing data from Federally funded research, including specific considerations for human data, to enhance public access and the overall return on R&D investments (White House Office of Science and Technology Policy, 2022). The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term. Trusted repositories have specific provisions in place and offer 17 | Page
explicit information online about their policies, which define their services (e.g. acquisition, access, security of content, long-term sustainability of service including funding, etc) (European Commission, 2024). The Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements (Jahn et al., 2023) states that there is no single standardised system for organisations or individuals to codify and express commitment to and endorsement of specific repositories. For example, endorsement can be seen to happen through use, so that repositories popular among researchers can be considered endorsed by their research community. The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on repository coverage that the higher-level subject areas/disciplines the repository covers, as well as cross-disciplinary domains, such as the types of data, technology and study. The Research Data College Working Group has put forward exclusion criteria for selecting a trustworthy subject-specific repository for self-depositing data. This methodology was devised and applied to the French national list of recommended trust repositories: https://recherche.data.gouv.fr/en/repositories). When it comes to the mission and scope, the Research Data College Working Group rule out “subject-specific repositories that limit data deposits to certain scientific communities where only scientists affiliated to the institution hosting the repository are authorised to deposit” (Lamotte et al., 2024, p. 11). In addition the Research Data College Working Group recommends that subject-specific repositories provide transparent information about the following items (Lamotte et al., 2024, pp. 12–13): ● Disciplinary field. Where possible, the disciplinary field employs the nomenclature used by HAL, which proposes a corollary tailored to the human and social sciences, which is not always the case with other existing nomenclatures. Furthermore, it offers up to three different levels of granularity, which makes it possible to provide better descriptions of the selected repositories. ● Data accepted. Here it is a question of describing the type of data accepted by the repository, ensuring that the terminology specific to each discipline is employed in order to make the depositor’s choice easier (e.g. NMR spectra, 3D structures of biological molecules, TEI-encoded corpus, etc.). ● Volume limit. This information can prove important for researchers from disciplines that generate substantial volumes of data. It also helps to anticipate the expected cost if the repository requires a financial contribution above a certain volume of data. The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) implicitly recognizes organizational scope and mission by stressing that the relevance and priority of its indicators must be adjusted according to the specific requirements and practices of the disciplinary community. This community focus translates into explicit essential indicators requiring that both metadata and data comply with domain-relevant 18 | Page
community standards (R1.3), thereby setting the necessary technical boundaries for FAIRness within that defined scope. In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), all 14 respondents, among which 11 repositories and 3 organisations, had a mission statement. A closer examination of the statements revealed differences in the content and length, although there were similarities between them, particularly amongst the certified repositories. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) report presents a first iteration of guidelines that will help to expose relevant information at the organisational and object level to facilitate discovery, provide context, and support interoperability between repositories, registries, and other related stakeholders. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) discusses the mission of repository and how some are not in scope of CTS certification e.g. due to not having long-term preservation mission. In all cases, repositories need to be clear on their responsibilities and priorities. For repositories with a long-term preservation mission, CoreTrustSeal is the core certification. If, for example, data findability and accessibility are important, FAIR evaluation is useful. If security aspects are important, there are ISO standards (like ISO 2700141) for security, and if IT service management is essential, FitSM42 defines a baseline of IT service management effectiveness. The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) requires repositories to clearly define their operational boundaries by providing public documentation outlining the scope of accepted resources. Essential to the mission, the repository must also clearly indicate the organization responsible for management and governance, and maintain a publicly available policy detailing the fate of resources if operations cease. Domain-specific resources Agri-food No relevant mapping to the TTRAM Activity/Function AF02 Mission & Scope was found in the gathered Agri-Food resources. Biomedical Sciences The TTRAM Activity/Function AF02 Mission & Scope is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: Having a clear mission and scope is a defining feature of an ELIXIR Core Data Resource (CDR), as emphasized in the article. This is reflected in several key evaluation criteria: 19 | Page
● First, the scientific focus and quality of a resource is assessed by determining whether it functions primarily as a deposition database, accepting data submissions, or as a knowledge base, which adds value through integration, annotation, or interpretation. ● Second, the resource must provide a scope statement that outlines its scientific coverage and comprehensiveness. This includes specifying whether the resource addresses all species or a subset, particular families, or outputs from specific experimental methods. It also involves positioning the resource in relation to other similar data resources, highlighting its unique contributions or complementary role. ● Finally, the community and potential usage are considered by estimating the size of the global user community that could benefit from the resource. This helps gauge the resource’s relevance and impact within the life sciences domain. ● Together, these criteria ensure that each ELIXIR CDR has a well-defined mission and scope, providing essential context for its activities and supporting its long-term sustainability and strategic value. Lin et al., 2024: In the article, the mission of data repositories is to provide core services such as ingestion (intake) of data, data management, preservation, archival storage, administration, access. It emphasizes the importance of clearly articulating the repository’s mission, including the type of data it accepts, the target user community, and the goals of the repository (e.g., long-term preservation, open access, domain-specific reuse). Besides, the article supports transparency through the TRUST principles and the provision of information such as: ● Types of data repositories ○ Domain-specific repositories: These repositories store data of a specific type (e.g., protein structure, nucleotide sequence, clinical data) or discipline (e.g., cancer, neurology). They often form a nexus of resources for their research communities interested in these specialized data. ○ Generalist repositories: These repositories store data of multiple types and disciplines, accepting data regardless of its type, format, content, disciplinary focus, or research institution affiliation. NIH has established agreements with several generalist repositories under the NIH Generalist Repository Ecosystem Initiative (GREI)16. ○ Project-specific data repositories: These repositories store domain-specific data generated from a project or collaboration (e.g., NIH All of Us17) and enable data sharing and reuse by making the project-specific data available for reuse by other projects or researchers. This is not to be confused with a project data coordinating center (DCC), which facilitates the project collaboration, curation, and data analysis but does not serve as a repository as data are not widely available for reuse by other researchers. Note that a DCC may also later facilitate submission of the project data to a data repository. 20 | Page
○ Institutional repositories: These repositories store data primarily created by members of an institution or a group of institutions, such as principal investigators (PIs), postdocs, and students. This category addresses the needs of the institution’s staff and may serve to collect data from one-to-many projects, and, depending on the institutional mission, may function as a domain-specific or generalist repository. ● Data type: Categories of digital assets that are managed by a repository. Barrett et al., 2012: ● BioSample Database: Stores submitter-supplied metadata about biological materials (e.g., cell lines, tissue biopsies, organisms, environmental isolates). ● BioProject Database: Provides an organizational framework for accessing information about research projects. Karsch-Mizrachi et al., 2025: The article clearly states that INSDC captures, preserves, and presents globally comprehensive nucleic acid sequence information as part of the permanent scientific record and as a forum for sharing amongst the broad scientific community. Climate Science The Activity/Function (A|F) Research & Development (R&D) is not addressed by any of the reviewed resources from Climate Science. Linguistics For CLARIN B and E centres, the following requirements apply when it comes to the AF Mission & Scope (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3): ● Centres need to offer useful services to the CLARIN community ● Each centre needs to make clear statements about their policy of offering data and services ● Centres need to employ activities to relate their role in CLARIN to the research community in order to guarantee a research based status of the infrastructure and allow researchers to embed their services in their daily research work ● Centres that are offering infrastructure type of services (E) need to specify their services for CLARIN and the terms of giving service For CLARIN B centres, the following requirements apply (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 2.a Description of the repository's context, mission and scope. Requirement: The centre’s repository context, mission and scope needs to be clearly defined. ● 2.e Details about resources and services provided (updated). Requirement: Each centre needs to make explicit statements about CLARIN compliant resources and services available at the centre. Social Sciences 21 | Page
The CESSDA Data Management Expert Guide (CESSDA Training Team, 2022, Chapter 7. Discover) outlines the core mission of trusted social science domain repositories: integrating data into the research lifecycle to ensure it is published, shared, discovered, and reused in line with the FAIR principles. Moreover, they: ● archive and preserve data; ● offer and manage access to the data; ● provide complex services focused on data reuse for research, teaching and learning; ● check data quality and compliance; ● improve data interoperability, e.g. by accompanying data with rich standardised metadata; ● maintain data catalogues; ● seek to add new data to their collections; ● develop training for data producers and data users. 3.2. Digital Object Management Digital Object Management covers the lifecycle of digital objects from creation and deposit through appraisal, curation, access, reuse, and preservation, emphasizing transparency in practices and alignment with levels of retention, curation and preservation (LoRCAP). 3.2.1. CONCEIVE, CREATE, COLLECT (AF03) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Conceive, Create, Collect as follows: “The stages of the research data lifecycle, before data enters a repository. Working with data or metadata creators and owners can help improve quality and highlight the benefits of curation, preservation and reuse.” and suggests for transparent information “any information or guidance provided to researchers about how to manage their digital objects’ data and metadata during the pre-repository phase” (L’Hours et al., 2025c, p. 12) including: ● Funding Information ● Project Information ● Data Management Plan (DMP) The phase immediately preceding Ingest into the repository is covered in the PAIMAS (Consultative Committee for Space Data Systems (CCSDS), 2004) and PAIS (Consultative Committee of Space Data Systems (CCSDS), 2014) standards. PAIMAS defines roles and the different sub-phases of the pre-Ingest phase, beginning with the first contact between data producer and archive. PAIS provides guidance for preparing the information objects to be submitted to and ingested by the archive during the “formal definition phase” as defined in PAIMAS. The Desirable Characteristics of Data Repositories for Federally Funded Research requiring Federally funded researchers to develop Data Management Plans (DMPs) that specify how data will be managed and shared before it enters a repository. Federal agencies provide guidance to help 22 | Page
researchers select appropriate repositories (White House Office of Science and Technology Policy, 2022). The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) addresses the pre-repository phase by being designed for use during the development of Research Data Management Plans, allowing data producers and project managers to specify the expected level of FAIRness their resources should achieve before data and metadata are produced. In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), most respondents (10 out of 14) offered support for the initial conceptualisation of research projects, access to data, and the collection and/or creation of research (meta)data. The GREI Data citation best practices for repositories (Puebla et al., 2024) highlight that data citations provide credit for the data producer and are among the biggest motivators for researchers to publish their data. Within AF03 Conceive, Create, Collect, the Data Curation Network defines the following Curation Activity (Johnston et al., 2016): “Conversion (Analog): In effort to increase the usability of a data set, the information is transferred into digital file formats (e.g., analog data keyed into a database). Note: digital conversion is also used to convert “fixed” data (e.g., PDF formats) into machine-readable formats.” Domain-specific resources Agri-food The TTRAM Activity/Function AF03 Conceive, Create, Collect was only found to be addressed in the following Agri-Food resource: Harper et al., 2018: AgBioData recommends that repositories advise data creators on management plans, metadata completeness, and FAIR readiness before data submission. “Proper data management is a critical aspect of research and publication… Data are the lifeblood of research, and their value often do not end with the original study, as they can be reused for further investigation if properly handled.”. And also: “Provide training for researchers on responsible data management. Tutorials covering all aspects of data management, including file formats, the collection and publishing of high value metadata along with data, interacting with GGB databases, how to attach a license to your data, how to ensure your data stays with your publication and more will be useful training tools.” Biomedical Sciences The TTRAM Activity/Function AF03 Conceive, Create, Collect is addressed in the following Biomedical Sciences resources: Field et al., 2011: GSC emphasizes standards for describing and capturing rich contextual information, exemplified by the MIxS checklists, ensures the maximised usefulness, quality, and 23 | Page
quantity of data for public collections of genomes, metagenomes, and marker gene sequences before they enter a repository. Yilmaz et al., 2011: The MIxS standard forms mandatory guidelines for collecting and reporting comprehensive contextual data at the time of sample acquisition and sequencing, effectively serving as an "electronic laboratory notebook". Barrett et al., 2012: The BioProject and BioSample databases and Submission Portal represent a proactive approach by NCBI to organize and integrate data across interdisciplinary resources and to obtain a rich set of contextual metadata from data producers. Karsch-Mizrachi et al., 2025: The article describes how INSDC engages with data creators and submitters before data enters the repository by setting expectations for metadata quality, submission standards, and validation processes. Besides, INSDC has set up working groups to increase consistency and alignment with other standards organizations such as the Genomic Standards Consortium, the Public Health Alliance for Genomic Epidemiology, and the Global Alliance for Genomics and Health, across eleven categories of data and metadata. Rehm et al., 2021: The article emphasizes collaboration with data creators through its Driver Projects and Work Streams, which help shape standards and policies from real-world needs. GA4GH develops data models for clinical and phenotypic data (e.g., Phenopackets, Pedigree) and standardized file formats (e.g., SAM, BAM, CRAM, VCF/BCF, VRS, VA), for consistent quality, reusability, and effective data sharing before data formally enters a repository. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that data providers are responsible for an initial contact phase where they describe their project, data characteristics, and publication needs to the DKRZ Data Management (DM). Following this, they must undertake thorough data preparation, which includes ensuring files are in accepted open-source formats (like NetCDF with CF conventions), meet specific file header standards, adhere to file size and compression guidelines, and are consistently labeled and organized into appropriate Dataset directories. Linguistics The Activity/Function (A|F) Conceive, Create, Collect is not addressed by any of the reviewed resources from Linguistics. Social Sciences CESSDA’s Data Management Expert Guide supports researchers in the social sciences in caring for their data throughout the research data lifecycle in accordance with community standards (CESSDA Training Team, 2022). 24 | Page
The combination of informed consent, anonymization, and access control enables the sharing of personal data, so attention should be paid to these issues when collecting and processing data. (CESSDA Training Team, 2022, Chapter 5. Protect). Thorough and systematic documentation of research data is essential for ensuring it can be published, discovered, cited, and reused. Clear and comprehensive metadata enhances the overall quality and usability of the data. (CESSDA Training Team, 2022, Chapter 2. Organise and Document). Pre-ingest ensures data can be properly preserved, easing the workload for curators. While high value data should ideally be shared via trusted archives in a structured and well-documented format, researchers often face barriers like limited time, resources, or incentives. This leads to incomplete or unsuitable data submissions. Pre-ingest helps identify and resolve issues with data and metadata quality, completeness, and format early on. It also fosters collaboration with researchers, promoting a stronger culture of data sharing and good data management practices. (CESSDA Training Team, 2025a, Chapter 3.1 Pre-Ingest Basics). 3.2.2. DEPOSIT & APPRAISAL (AF04) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Deposit & Appraisal as follows: “Accepting custody of digital objects from depositors, transferring responsibility to the repository. It may also include appraising offered or requested deposits to ensure they meet established criteria for acceptance.”, and suggests for transparent information “[d]ocumentation of compliance criteria applied at the point of deposits, whether automated or manually applied and the degree to which they are required or optional” (L’Hours et al., 2025c, p. 13), including ● Collections Development and Appraisal Policy ● Deposit Compliance Criteria ● Deposit Procedures ● Acceptable/Preferred File Formats List ● ReAppraisal Plan, Data Management Plan ● Deposit licence (see also AF15 “Rights”). The suggested CoreTrustSeal Levels of Retention, Curation and Preservation (LoRCaP) address the deposit and appraisal aspect with level “D. Deposit Compliance”. On this level, “Data content and supporting metadata deposited are checked for compliance with defined criteria, e.g. data formats, metadata elements, and compliance with legal and ethical norms. Digital objects that do not meet these criteria may be rejected, or moved forwards to initial curation if provided by the repository” (CoreTrustSeal Standards and Certification Board, 2024) . Metadata required on objectand repository-level to demonstrate how the LoRCaP are implemented are proposed in L’Hours et al., 2024. The A|F Deposit & Appraisal is covered by CoreTrustSeal’s Deposit & Appraisal (R08), which requires certifying repositories to accept data and metadata based on defined criteria to ensure relevance and understandability for users. To document the fulfilment of R08, CoreTrustseal asks applicants to 25 | Page
● The standards that data, metadata and documentation must comply with to be acceptable for preservation and access. Whether these are general external standards, internally developed standards or specific to a community of practice. ● The quality control checks in place ensure the completeness and understandability of data and metadata. ● The approach to resolving issues e.g. whether the digital objects are returned to the depositor for rectification, fixed by the repository, noted by quality flags, and/or included in the accompanying metadata. ● The approach to managing changes to expected standards (e.g. new or updated data formats of metadata schemas) in response to changes in the technical environment or to changes in the needs of the Designated Community. ● Any links provided to other digital objects’ data and metadata e.g. related digital objects, publications, or the use of controlled vocabularies and ontologies. In the nestor Seal this A|F is addressed in the criteria C21-C25, which require that “the digital archive has issued specifications” for SIPs and AIPs as well as for the transformation of a SIP to an AIP and an AIP to a DIP. In addition, “C1 Selection of information objects and their representations” as well as “C20 Technical authority” are relevant here. (nestor Certification Working Group, 2025). In the OAIS Reference Model initial curation actions are performed in the Ingest functional entity. In particular, the “Generate AIP function” creates an Archival Information Package (AIP) from the SIP in accordance with archive standards. This may comprise the following actions: ● file format conversions, ● determining Transformational Information Properties1 (sometimes also referred to as significant properties), ● gathering adequate Representation Information, including from the Producer, and creating the Descriptive Information that completes the AIP, ● reorganizing the content information stored in the SIPs. (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 4/7). The EOSC Federation Handbook partially addresses the A|F Curation, Quality & Compliance, when requiring that “[d]ata repositories in the EOSC Federation should measure the FAIRness of their data through FAIR metrics and they should be able to demonstrate compliance to the FAIR principles by implementing at least the Findable and Accessible guidelines” (EOSC Association, 2025, p. 32). The Desirable Characteristics of Data Repositories for Federally Funded Research highlights that repositories should provide or facilitate expert curation and quality assurance to enhance the 1 “An Information Property the preservation of the value of which is regarded as being necessary but not sufficient to verify that any NonReversible Transformation has adequately preserved information content. This could be important as contributing to evidence about Authenticity. Such an Information Property is dependent upon specific Representation Information, including Semantic Representation Information, to denote how it is encoded and what it means.” (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 1/17). 32 | Page
accuracy and integrity of datasets and metadata, ultimately making Federally funded research data FAIR to the fullest extent possible (White House Office of Science and Technology Policy, 2022). The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates the Curation & Preservation Levels from Z. (Level Zero), D. (Deposit Compliance) C. (Initial Curation) B. (Logical-Technical Preservation) and A. (Conceptual preservation) for understanding and reuse. Data services, including repositories should specify all the levels of care they apply to objects within their collection, including through repository and digital object registry metadata. Different types of data services benefit from being transparent on their current level of storage, curation and preservation practice, as this increases trust by the user and funders alike. The EU Annotated Grant Agreement states repositories have mechanisms or provisions for expert curation and quality assurance for the accuracy and integrity of datasets and metadata, as well as procedures to liaise with depositors where issues are detected They ensure that contents are accompanied by metadata sufficiently detailed and of sufficiently high quality to enable discovery, reuse and citation and contain information about provenance and licensing. Their metadata is machine-actionable and standardized (e.g. Dublin Core, Data Cite, etc) preferably using common non-proprietary formats and following the standards of the respective community the repository serves, where applicable (European Commission, 2024). The Update of the Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements states that the lack of a public policy for preservation, curation and security of the contents is the most frequent reason, followed by not adhering to a specific metadata standard (Lazzeri, 2024). The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on data curation that review and annotation of the data performed by the repository (e.g. via a data submission tool that enforces some curation, or by its curation team). Does the repository curate its holdings? Findings from D1.3 Recommendations for a FAIR EOSC - White Paper of the FAIR-IMPACT Synchronisation Force (Grootveld, 2025) emphasize that building an inclusive and collaborative network of Trustworthy Digital Repositories (TDRs) is essential to share expertise and help repositories improve their trustworthiness, which includes achieving or preparing for certified status, often involving curation processes. Develop semantic interoperability through FAIR Semantic Artefacts, via adoption of policies, guidelines and mappings. Furthermore, achieving high data quality (fitness for purpose) is important for daily research and its value, even though quality is distinct from the assessment of FAIRness. The O'FAIRe makes you an offer: Metadata-based Automatic FAIRness Assessment for Ontologies and Semantic Resources (Amdouni et al., 2022) details a methodology and tool, O'FAIRe, for automatically assessing the level of FAIRness—a key standard of quality and compliance—for ontologies and semantic resources based largely on their metadata descriptions as managed by repositories. 33 | Page
The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) addresses curation, quality and compliance by setting Essential indicators for Reusability (R1), notably requiring that metadata and data comply with domain-relevant community standards (R1.3-01M, R1.3-01D) and that metadata is expressed in compliance with a machine-understandable community standard (R1.3-02M). Furthermore, achieving a high level of FAIRness relies on providing a plurality of accurate and relevant attributes (R1-01M) and including clear license information (R1.1-01M), both classified as Essential indicators necessary for enabling reuse. In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), most of the 14 respondents reported performing a combination of curation levels, with only five respondents selecting one level. The most common choice was enhanced curation by converting data to new formats and enhancing documentation, followed closely by basic curation by briefly checking the data and adding basic metadata. Six respondents reported providing data level curation by further editing deposited data for accuracy. A further two organisations indicated that they perform no curation for some data in their collection, alongside higher levels of curation for other parts of their collection. The D3.1—Report on Discipline Requirements and Needs (Andreassen et al., 2025) states in section 4.3.1, Contextual Data Quality, that it is strongly recommended to provide standardised, accessible, and comprehensive methods for applying documentation to datasets, intended to exist alongside standardized metadata schema. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that repositories express the levels of care offered and received by digital objects, which should detail the levels of curation and preservation in place and how these might change over time. Furthermore, standardizing the exposure of assessment results using the Data Quality Vocabulary (DQV) is recommended, which can embed FAIR assessment outcomes within the metadata of assessed datasets via DCAT, including crucial information such as the test date and the testing tool used. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) introduce several indicators that address aspects related to quality and standards, particularly under the Reusable (R) principle. Specifically, it is considered Essential (R1.3-01M, R1.3-01D) that both metadata and data comply with domain-relevant community standards, which directly supports the goal of ensuring defined levels of quality and compliance prior to making objects available for reuse. The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) emphasizes that the repository must undertake a lightweight review of basic metadata upon resource submission, and enhance it if necessary. Repositories must also provide documentation or a policy that outlines the curation processes applied to both the resources and their metadata to ensure defined criteria for access and reuse are met. 34 | Page
The GREI Data citation best practices for repositories (Puebla et al., 2024) recommend that repositories store citation links using specific DataCite metadata fields (such as relatedIdentifier and relationType) to establish a standardized relationship between the dataset and the citing scholarly output. Furthermore, to increase transparency and trust in the quality of the citation information, repositories must expose the provenance (source) for every asserted data citation listed on the dataset landing page. The Core Preservation Process CPP-019 Data Quality Assessment sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA evaluates and re-evaluates the data quality of Information Objects.” (EOSC EDEN T1.2 et al., 2025). Within AF05 Curation, Quality & Compliance, the Data Curation Network defines the following Curation Activities (Johnston et al., 2016): ● “Arrangement and Description: The re-organization of files (e.g., new folder directory structure) in a dataset that may also involve the creation of new file names, file descriptions, and the recording of technical metadata inherent to the files (e.g., date last modified).” ● “Code review: Run and validate computer code (e.g., look for missing files and/or errors) in order to find mistakes overlooked in the initial development phase, improving the overall quality of software.” ● “Data Cleaning: A process used to improve data quality by detecting and correcting (or removing) defects & errors in data.” ● “Deidentification: Redacting or removing personally identifiable or protected information (e.g., sensitive geographic locations) from a dataset prior to sharing with third-parties.” ● “File renaming: To rename files in a dataset, often to standardize and/or reflect important metadata.” ● “Indexing: Verify all metadata provided by the author and crosswalk to descriptive and administrative metadata compliant with a standard format for repository interoperability.” ● “Interoperability: Formatting the data using a disciplinary standard for better integration with other datasets and/or systems.” ● “Quality Assurance: Ensure that all documentation and metadata are comprehensive and complete. Example actions might include: open and run the data files; inspect the contents in order to validate, clean, and/or enhance data for future use; look for missing documentation about codes used, the significance of “null” and “blank” values, or unclear acronyms.” ● “Restructure: Organize and/or reformate poorly structured data files to clarify their meaning and importance.” Domain-specific resources Agri-food The TTRAM Activity/Function AF05 Curation, Quality & Compliance is addressed in the following Agri-Food resources: 35 | Page
Harper et al., 2018: AgBioData recommends that repositories define and publish data-quality and compliance policies to ensure community standards are met. “Developing data standards and practices to facilitate data curation, integration and reuse is essential. Standards will ensure data quality, facilitate interoperability and reduce redundant work.” Sen et al., 2020: WheatIS requires repositories to apply common quality criteria and ensure compliance with agreed data and metadata standards across distributed nodes. Top et al., 2022 : One of the cases in the fourth part of the article illustrates the need for a specific curation and data alignment process : “Selected data goes through a dedicated data curation procedure. This is a crucial step to enable use of the data beyond its original purpose of collection. In this procedure metadata is checked and completed using all information available in the data files, supporting documents, publications in data and scientific journals, etc. “ Biomedical Sciences The TTRAM Activity/Function AF05 Curation, Quality & Compliance is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: ● Compliance: FAIR is a set of guiding principles to make data Findable, Accessible, Interoperable, and Reusable. These indicators will be used to demonstrate that ELIXIR Core Data Resources (CDR) are compatible with the FAIR data principles. For instance, a CDR is expected to provide persistent and unique identifiers to enable usability and community-recognised standards for metadata and data to enable interoperability. ● Level of quality: scientific quality is one of the key indicators, especially for knowledge bases which add substantial value through expert curation, annotation of metadata. The curation effort and outputs linked to a resource are an important measure of the quality of ELIXIR CDRs. This is why the number of FTE corresponding to Curators position to support adherence to metadata requirements or support for extraction of information from the scientific literature is one of the indicators. Lin et al., 2024: In the article, curation, described as the process of employing various standards and best practices to transform data into meaningful organized, structured, and computable forms is one of the 6 common characteristics of data repositories. To ensure transparency, users of a data repository should understand the operational aspects of a repository, such as data validation and curation procedures. Field et al., 2011: The GSC developed the MIxS standard, including checklists (MIGS/MIMS/MIMARKS) that require core information upon data submission and publication to ensure compliance and richer entries. 36 | Page
Yilmaz et al., 2011: The MIxS standard is crucial in ensuring sequence data reaches a defined level of quality and standards compliance before it is made available for reuse by major data providers, including partners of the International Nucleotide Sequence Database Collaboration. Barrett et al., 2012: The article describes mechanisms that support metadata quality and standardization, but it does not fully implement the level of curation, quality assurance, and compliance enforcement envisioned by FIDELIS TTRAM. The responsibility for metadata quality largely remains with the submitters. Karsch-Mizrachi et al., 2025: The article explains how the INSDC ensures that nucleotide sequence data meets defined quality standards and compliance requirements before it is made available for reuse. This is achieved through the implementation of validation rules, metadata checklists, and minimal standards that are applied during the submission process. These efforts reflect a clear commitment to ensuring that digital objects are curated and meet quality and compliance benchmarks before being integrated into the repository and made accessible to the global scientific community. Climate Science The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) states that Curation, Quality and Compliance is actively managed through the CMIP5 Questionnaire, a metadata entry tool that enforces Controlled Vocabulary constraints, validates against the CIM XSD, and performs Schematron-based validation to check the deeper coherency of collected parameters, ensuring digital objects and their documentation adhere to defined standards before publication. The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) explains a multi-stage process involving a pre-publication quality control phase, which encompasses both Technical Quality Assurance (TQA) performed by WDCC and Scientific Quality Assurance (SQA) overseen by the data provider, verifying adherence to open source file formats, CF Metadata Conventions, and WDCC Minimal Standards for File Headers. Post-publication, curation activities maintain long-term usability and citability by enabling the addition of references, management of "open time series" for evolving datasets, and the creation of new versions or handling of errata with explanations. The NetCDF Climate and Forecast (CF) Metadata Conventions (Eaton et al., 2024) state that the CF conventions ensure requirements for self-describing metadata that is both human-readable and easily parsable by programs. This framework defines precise attributes for variables, units, coordinate systems, and data representation to facilitate unambiguous interpretation and interoperability. Linguistics The Activity/Function (A|F) CURATION, QUALITY & COMPLIANCE is not addressed by any of the reviewed resources from Linguistics. 37 | Page
Social Sciences (Kleemola et al., 2025) recommends that data repositories should: ● set standards for preservation and other repository operations by for example contributing to defining TDRs and by building consensus on curation and preservation levels. ● leverage domain-specific expertise in long-term data preservation, advocating for domain-specific services such as metadata creation and data curation and advocating the benefits of long-term data preservation for future research and society. A recommended best practice is for the repository to have a clearly defined data curation policy that outlines how data is maintained and how its value is enhanced to support reuse and long-term preservation. (CESSDA Training Team, 2025a, Chapter 4.4 Quality assurance of data and documentation material). The CESSDA Data Management Expert Guide addresses the limitations of the DDI metadata standard, which is recommended for social science research. It has limitations in “describing biases caused by data mining interfaces of social media platforms, in data availability and formats, explanations about code and scripts used in collection, cleaning and analysis etc.” These can be described only as free-text comments outside the structured standard. (CESSDA Training Team, 2022, Chapter 2. Organise and Document). The Guide also emphasizes the value of trusted CESSDA repositories as they ensure long-term access and data quality through expert guidance. Experts help improve metadata completeness, advise on suitable file formats for long-term preservation and review the quality of data. (CESSDA Training Team, 2022, Chapter 6. Archive&Publish). 3.2.4. DISCOVERY & IDENTIFICATION (AF06) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Discovery & Identification as follows: “Applying persistent identifiers and descriptive metadata to digital objects to support resource discovery. Providing discovery systems and making metadata available for harvesting.” and suggests for transparent information “[p]ersistent identifiers used (ideally chosen from a controlled vocabulary) for objects, organisations, researchers, software etc, and information about whether all objects are persistently identified. Information about identification at a more granular level (files within objects, questions within surveys, variables within statistics). Links to resource discovery systems that the repository provides or third parties that index metadata harvested from the repository collection.” (L’Hours et al., 2025c, p. 14) including ● Persistent Identifier system(s) ● PID Policy ● Discovery metadata and documentation 38 | Page
● Resource discovery catalogues ● Harvesting protocols (e.g. OAI-PMH) The A|F Discovery & Identification is covered by the CoreTrustSeal requirement with the identical name, requiring certifying repositories to enable users to discover the digital objects and refer to them in a persistent way through proper citation. The CoreTrustSeal extended guidance asks applicant to provide references to the following items (CoreTrustSeal Standards and Certification Board, 2022, pp. 22–23): ● The search facilities offered by the repository. ● The standards that a searchable metadata catalogue complies with. ● The approach to ensuring that identifiers are unique and persistent. ● Machine harvesting of the metadata. ● Repository, or repository data and metadata, inclusion in disciplinary or generic registries of resources. ● Recommended data citations. The nestor Seal criteria relevant to this A|F are “C16 Integrity: User Interface”, “C27 Identification”, and “C28 Descriptive Metadata”. C16 requires the archive to provide an interface “which allows users and the digital archive administration to check and maintain the integrity of the representations”. C27 deals with the use of unique and persistent internal and external identifiers. In addition, “C4 Access” mentions the need for “appropriate search possibilities” (nestor Certification Working Group, 2025). In the OAIS Reference Model the Data Management functional entity manages the metadata required for data discovery - i.e. the descriptive information generated during Ingest - and will return responses to queries from the Access functional entity (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2.3.5). The EOSC Federation Handbook addresses parts of the A|F Discovery & Identification in several of its sections (EOSC Association, 2025): ● 5.2.1 Research publications: The EOSC Federation will automatically reference publications through OpenAIRE and be searchable via the EU Node resource catalogue. ● 5.2.2 Research data sources: Linking between data repositories and databases is strongly encouraged and is fundamental to creating the Web of Data. ● 5.2.2.1 FAIR Data repositories: The repositories must be [...] be findable in the EOSC Federation catalogue and searchable through a custom search interface. Data repositories should implement a search function which is optimised to find data from the scientific domain they serve. [...] Data must be citable via a Persistent IDentifier (PID) according to the “Guidelines for creating a user tailored EOSC Compliant PID Policy” and the Persistent Identifier (PID) policy for the European Open Science Cloud (EOSC). [...] Data repositories must ensure that Metadata are harvestable via the standard protocols OAI-PMH (https://www.openarchives.org/pmh/). 39 | Page
The Desirable Characteristics of Data Repositories for Federally Funded Research emphasizing that repositories should assign unique, citable PIDs (like DOIs) to datasets and ensure they are accompanied by descriptive metadata using appropriate schemas to enable discovery, reuse, and citation (White House Office of Science and Technology Policy, 2022). The EU Annotated Grant Agreement states repositories assign persistent unique identifiers to contents (e.g. DOIs, handles, etc), such that the contents (publications, data and other research outputs) are unequivocally referenced and thus citable. They ensure that contents are accompanied by metadata sufficiently detailed and of sufficiently high quality to enable discovery, reuse and citation and contain information about provenance and licensing. Their metadata is machine-actionable and standardized (e.g. Dublin Core, Data Cite, etc) preferably using common non-proprietary formats and following the standards of the respective community the repository serves, where applicable. For an overview/comparison of EC GA metadata requirements, see EC GA Metadata Requirements (European Commission, 2024). The Update of the Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements show also in this case that the main reasons for not meeting the essential characteristics for trusted repositories are the lack of a licence field in the metadata, along with missing adherence to a specific standard for metadata. Furthermore, grant information is often not offered through separate metadata fields, resulting in a lack of machine-actionable, interoperable, and standardized metadata, as requested in the MGA (Lazzeri, 2024). The Recommendations for Services in a FAIR data ecosystem (Bangert et al., 2019) states that services supporting FAIR data should offer or make use of the following components: A. PID services for a wide range of objects, such as publications, researchers, data sets and organisations. Emerging PID types (e.g. for instruments) should be monitored and used when they are mature. B. Domain-specific ontologies, as domain-specific requirements have to be taken into account. C. Human and machine-readable standards to make datasets findable, reusable and interoperable (licences as one particular example of standards needed for machine readability). D. If applicable, metadata that complies with appropriate (domain) standards should be generated and captured automatically (for e.g by instruments) . The Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements (Jahn et al., 2023) states that for fulfilment of the Horizon Europe MGA metadata access and licensing criteria, we consider the repository to be required to provide metadata related to each digital object via a public domain dedication such as CC0, Public Domain or equivalent. However, as it is further shown in the Analysis section many repositories do not readily provide this information in a clear human or machine-readable manner. 40 | Page
The Practical Guide to the International Alignment of Research Data Management (Science Europe, 2021) states that: 1. Provision of Persistent and Unique Identifiers (PIDs): a. Allow data discovery and identification b. Enable searching, citing, and retrieval of data c. Provide support for data versioning. 2. 2. Metadata: a. Enable finding of data b. Enable referencing to related relevant information, such as other data and publications c. Provide information that is publicly available and maintained, even for non-published, protected, retracted, or deleted data d. Use metadata standards that are broadly accepted (by the scientific community) e. Ensure that metadata are machine-retrievable. The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on data and metadata standards that the community-defined standards the repository implements to enable the representation of data and/or metadata in a consistent, machine readable form (e.g. via models, formats, schemas, vocabularies, ontologies). These standards facilitate the discovery and interpretation of data and/or metadata. Which data and metadata standards (if any) has the repository implemented? Persistent Identifiers for Data. Globally unique and Persistent IDentifiers (PIDs). Does the repository assign PIDs to its holdings? If so, which PID schema has been implemented? Citation to related publications. A mechanism to link datasets to related articles or pre-prints. Does the repository enable data to article linking? At what stage of data deposition is article information required? The Current State and Future Directions for Open Repositories in Europe (Shearer et al., 2023) states that despite the fact that a repository supports certain metadata schemas and PIDs does not always equate to the collections having high quality metadata. While most repositories do support standardised and granular metadata schemas, they often rely on the author to fill in the metadata fields. Since authors may not be aware of the standards, this often leads to lower quality metadata records. While most repositories do undertake basic metadata curation and checking, this may not be sufficient to optimize discovery and reuse of repository resources. There are opportunities to improve the quality of metadata - either through data curation activities at the repository, or by introducing machine extraction of metadata information - but this may require greater commitment in terms of staff and technical resources at the repository. The exclusion criteria from the Research Data College Working Group rule out subject-specific repositories that do not assign long-term identifiers (Lamotte et al., 2024, p. 10). The O'FAIRe makes you an offer: Metadata-based Automatic FAIRness Assessment for Ontologies and Semantic Resource's (Amdouni et al., 2022) analysis of 149 semantic resources showed generally high scores for Findability (F), with an average normalized score of 70, confirming that most 41 | Page
self-describing through standardized, machine-parsable metadata. This design's use of persistent identifiers (e.g., DOIs for the document itself and LSIDs for data elements) directly supports automated discovery systems. Linguistics The Tromsø recommendations for citation of research data in linguistics (Andreassen et al., 2019) provide guidance on how to cite data from linguistics and language research and are further elaborated in the handbook chapter “Guidance for Citing Linguistic Data” (Conzett & De Smedt, 2022). The Tromsø recommendations suggest two templates for citation of datasets (Conzett & De Smedt, 2022, p. 146): 1. The template for a minimal bibliographic reference to a dataset has the following elements: Author, Date, Title, Publisher, Locator. 2. The template for an expanded bibliographic reference to a dataset including conditional elements (i.e., required in certain cases depending on resource characteristics) is as follows: Author, Other Attribution (Roles), Date, Title, Publisher, Locator, Version, Date accessed. Elements rendered in bold are part of the minimal template, in other words, they are always required, while elements rendered in italics are considered to be conditional. Conditional elements are elements whose presence is conditioned by either the characteristics of the resource (e.g., references to versioned datasets should include the version number), or on subfield-specific traditions (e.g., in language documentation, it is common to acknowledge the contributions of language consultants by name). Repositories and other resource providers are advised to provide metadata conformant to the following (Conzett & De Smedt, 2022, p. 153): ● At minimum, the metadata should include the elements in the minimal templates recommended herein. ● Metadata should preferably be structured according to a standard format (e.g., component MetaData Infrastructure, RIS) so that information from it can be extracted by programs. ● Metadata should be available freely, without cost or restrictions, even if the data themselves have restrictions. ● The metadata should allow persistent reference to the data set, and to the metadata themselves, to avoid link rot. This implies that PIDs should be assigned to the data and to the metadata. ● Data repositories and other resource providers should provide metadata in machine-readable as well as human-readable form and should preferably also generate ready-made citations, both in formats for export to reference managers and in textual format. The CLARIN Data Citation Guidelines version 1.0 (Matthiesen & Leonardič, 2025) differ from the Tromsø recommendations on the following items: 48 | Page
● “Labelling different versions is simply a matter of convention (rather than a necessary condition for identification) given that each new version of a dataset should be assigned a unique PID.” ● “For resources that are in CLARIN repositories, URLs other than those containing PIDs are not acceptable since they might break.” ● “Note that the use of PIDs makes mentioning “accessed at <date>” redundant. We therefore do not recommend it, in contrast to other recommendations.” For CLARIN B and E centres, the following requirements and recommendations apply when it comes to the AF DISCOVERY & IDENTIFICATION (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3): ● Centres need to offer component based metadata (CMDI) that make use of elements from accepted registries such as the CCR in accordance with the CLARIN agreements, i.e. metadata needs to be harvestable via OAI PMH. ● Centres need to associate PIDs records according to the CLARIN agreements with their objects and add them to the metadata record. ● Centres are advised to participate in the Federated Content Search with their collections by providing an SRU/CQL Endpoint. This content search is especially suitable for textual transcriptions and resources. CLARIN Type C centres are expected to serve metadata via the OAI-PMH protocol (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 2). For CLARIN B centres, the following requirements apply (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 3.c Licenses on data and metadata. Requirement: Data and metadata are licensed. ● 6.a Metadata harvestable via OAI-PMH to the VLO. Requirement: Computer access to the repository: Metadata harvesting to VLO should work - see https://vlo.clarin.eu/data/. ● 6.b CMDI metadata validation and use of CMDI profiles with CCR ConceptLinks. Requirement: The metadata should be CMDI-compliant (see https://www.clarin.eu/cmdi). ● 6.c Metadata PID and references to resources. Requirement: ○ 1) State if the harvested CMDI files contain a PID in the MdSelfLink header field; ○ 2) State if the harvested CMDI files refer to web-accessible files or a landing page with a ResourceProxy. ● 7. Persistent Identifiers. Requirement: Centres need to associate PIDs (handles or DOIs) with their metadata records. These PIDs should be suitable for both human and machine interpretation, taking into account the HTTP-accept header. Individual files (e.g. a text, zip or sound file) can be referred to with either the handle of the describing metadata record in combination with a part identifier or with another handle. ● 8. Federated Content Search (optional). Requirement: Centres can choose to participate in the Federated Content Search with their collections by providing an SRU/CQL Endpoint. 49 | Page
Social Sciences The CESSDA Data Access Policy (Bolton, 2022) emphasizes the importance of discovery and identification by requiring that metadata be openly available to anyone (Principle 1), aligned with FAIR principles (Principle 4), harvestable using controlled vocabularies (Principle 5), include persistent identifiers and citation guidance (Principle 6), and clearly state access conditions (Principle 7). The CESSDA Metadata Model (CMM) serves as a metadata schema for social science data. It contains metadata elements, their definitions and information on other requirements, such as repeatability. The model is intended to enhance the discoverability and clarity of data for users, as well as to support interoperability between CESSDA service providers. It is built from the viewpoint of quantitative (social science) data and is based on the DDI Lifecycle 3.2 metadata standard. (Akdeniz & Moilanen, 2023). (Kleemola et al., 2025) suggest that data repositories should extend the use of PIDs to cover datasets, studies, authors, data contributors, funders, and study components while standardising machine-readable metadata across repositories. The CESSDA Data Archiving Guide Chapter 2.4 (CESSDA Training Team, 2025a, Chapter 2.4 Data Access Policies) states that having a persistent identifier (PID) for data is a key element in enabling data access - and thereby supports discovery. Staff should ensure that metadata is created according to FAIR principles, making research data more accessible and reusable. Data descriptions follow international standards, such as those defined by the Data Documentation Initiative (DDI) and CESSDA. To support consistency and interoperability across datasets and catalogues, controlled vocabularies should be used for key metadata fields. Using the controlled vocabularies should be mandatory for all published datasets. (CESSDA Training Team, 2025a, Chapter 4.3.2 Vocabularies (CMM, DDI, etc.)). The article recommends that repositories should provide rich citation-related metadata in machine-actionable data format and citation suggestions that include at least the six core citation components, including persistent identifier (PID) and version information. PIDs should be available also to embargoed datasets. General metadata containers should describe relationships between the dataset and other entities, using machine-actionable PIDs (ORCID, ROR ID, DOI). The document recommends that citation requests could be included in metadata and repository entries. (Bornatici et al., 2025, pp. 6–8). 3.2.5. ACCESS (AF07) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Access as follows: “Determining the appropriate method of access—such as direct download or use within a secure remote environment—and facilitating user access accordingly. Access criteria are based on the characteristics of both the digital objects and the users, in alignment with rights management 50 | Page
policies (see AF15 "Rights").” (L’Hours et al., 2025c, p. 15) and suggests the following for transparent information: ● Access routes & methods ● Access Management Policy ● Documentation of access routes and methods, including machine harvesting ● Licence and terms of access (see also AF15 “Rights”) In the OAIS Reference Model access is handled by the Access functional entity. It maintains a search interface for the data users, e.g. a catalog, and creates Dissemination Information Packages (DIPs) to provide archive content to the users (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2.3.8). In the nestor Seal, criterion “C4 Access” is relevant to this A|F. It states: “The digital archive ensures that authorised users in the designated communities can access the representations. This includes appropriate search possibilities. The digital archive openly declares its conditions of use and any costs which may arise, listing these in a transparent manner”. In addition, C16 Integrity: User Interface has some relevance to this A|F (nestor Certification Working Group, 2025). The EOSC Federation Handbook addresses parts of the A|F Access in two of its sections (EOSC Association, 2025): ● 5.2.2.1 FAIR Data repositories: Data repositories must have a data policy that governs the access to the data which conforms with the general data policy of the EOSC Federation. ● 5.2.5 Research Services: All Research Services must be accessible either anonymously e.g. for open data repositories, or through an EOSC compliant AAI (cf. Chapter 4) in accordance with documented User Access Policies (see Chapter 6). The Desirable Characteristics of Data Repositories for Federally Funded Research policy emphasizes that repositories should provide broad, equitable, and maximally open access to datasets and metadata free of charge, ensuring clear documentation of access and use terms. For human data, repositories must implement specific procedures like tiered access, user credentialing, download control, and review of access requests to protect privacy and ensure fidelity (White House Office of Science and Technology Policy, 2022). The Practical Guide to the International Alignment of Research Data Management (Science Europe, 2021) states that: ● Data access and usage licences: ○ Enable access to data under well-specified conditions ○ Ensure data authenticity and integrity ○ Enable retrieval of data ○ Provide information about licensing and permissions (in ideally machine-readable form) ○ Ensure confidentiality and respect rights of data subjects and creators. 51 | Page
The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on data access conditions that data access mechanisms and terms to define access at repository and/or dataset level. Does the repository explicitly state data access conditions on dataset landing pages? The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states regarding powered by AI-ready infrastructure that adopting machine-readable metadata, linked data practices, and text/data mining capabilities to ensure repository resources remain accessible and valuable in an AI-driven research ecosystem. Findings from D1.3 Recommendations for a FAIR EOSC - White Paper of the FAIR-IMPACT Synchronisation Force (Grootveld, 2025) emphasize that facilitating secure data sharing and access to data requires defining a harmonized operational and legal framework, and recommend advocating for Creative Commons licenses as the default option to enhance the legal and organizational interoperability that governs access criteria. The Research Data College Working Group recommends that subject-specific repositories provide transparent information about whether “the repository offers the possibility of combining the deposit with an embargo” [...], as “[s]ome research teams may be keen to delay open-access publication of their datasets, specifying a specific embargo period” (Lamotte et al., 2024, p. 12). The O'FAIRe makes you an offer: Metadata-based Automatic FAIRness Assessment for Ontologies and Semantic Resources (Amdouni et al., 2022) found that the Accessibility (A) score for semantic resources in AgroPortal was high, with an average normalized score of 90, due to all ontologies being accessible via the open, free, and universal HTTP protocol (A1.1), and the repository (AgroPortal) systematically supporting versioning (A2) and authentication and authorization (A1.2) for access control. The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) addresses appropriate user access through numerous Essential indicators in the Accessible (A) area, requiring that both metadata and data be retrievable using standardized and free access protocols (A1.1-01M, A1-04M, A1-04D) and that manual access is possible (A1-02M, A1-02D). Although the most secure access methods (like those supporting authentication and authorization) are only classified as a Useful indicator (A1.2-01D), it is Important that the metadata contains the necessary information regarding access conditions and required actions for the user to retrieve the data (A1-01M). The FAIRsFAIR project D2.3 report (Behnke et al., 2020) recommends different access policies for different versions of the data. The D3.1—Report on Discipline Requirements and Needs (Andreassen et al., 2025) states in section 4.3.3, Metadata Quality, that standardized time recommendations for file restrictions, access control or embargoes are needed in order to strike a balance between FAIR requirements for data, publication requirements and researcher concerns surrounding open data sharing. A further need is 52 | Page
for repositories to ensure the existence of clear channels of communication to discuss restricted or embargoed data (e.g. author and/or repository contact information that is not time-sensitive). The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) propose that guidance should be incorporated directly into Data Management Planning (DMP) tools to support users in making informed choices about data deposit concerning access, along with storage, curation, and preservation, from the earliest stages of their research. Repositories are further advised to be transparent about the levels of care offered and received by digital objects, which should include details about the levels of retention, curation, and preservation in place and how they might change, impacting eventual access. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that access management is addressed through licensing, user support, and access portals. The report emphasizes the importance of transparency in access policies and the role of certification in improving access services. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) strongly aligns with appropriate method of access by making several access-related indicators essential, requiring that data and metadata identifiers resolve to the digital object or metadata record (RDA-A1-03M, RDA-A1-03D), that they are accessible through standardized protocols (RDA-A1-04M, RDA-A1-04D), and that metadata contains information necessary for the user to gain access (RDA-A1-01M, an Important indicator). Essential indicators also mandate that metadata is accessible through a free access protocol (RDA-A1.1-01M). In the D4.5 Report on Completed FAIR Data Standard Adoption and Certifications of Data Repositories in the Region (Alaterä et al., 2022), repositories are required to ensure that data can be retrieved and properly resolved via an open and free protocol, by testing the resolution of the data's GUID. Furthermore, any discovered GUID must support authentication and authorization mechanisms within its resolution protocol and comply with an explicit data access policy. Metadata must also include a clearly identified, machine-readable persistence policy to ensure long-term accessibility and reliability. The Core Preservation Process CPP-025 Enabling Access sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA gives access to its Information Objects to authorised internal users or end users.” (EOSC EDEN T1.2 et al., 2025). Within AF07 Access, the Data Curation Network defines the following Curation Activities (Johnston et al., 2016): ● “Embargo: To restrict or mediate access to a data set, usually for a set period of time. In some cases an embargo may be used to protect not only access, but any knowledge that the data exist.” ● “File download: Allow access to the data materials by authorized third parties.” 53 | Page
● “Restricted Access: In order to maintain the privacy of research subjects without losing integral components of the data, some data access will be protected and/or mediated to individuals that meet predefined criteria.” Domain-specific resources Agri-food The TTRAM Activity/Function AF07 Access is addressed in the following Agri-Food resources: Top et al., 2022: AgroDataCube provides a concrete repository model for transparent access documentation, combining DOI-based catalogue links, API-based retrieval, and license information. “The web page of AgroDataCube also provides information about the used license and the access token needed for data retrieval.” Drakos et al., 2015: agINFRA requires repositories to define user access methods through interoperable infrastructures and to validate access routes that enable collaboration. “Deploy tools (exchange standards, software, methodologies) for collaboration between European institutions in data-intensive research; validate the approach by enabling users to interact with each other and the data.” Sen et al., 2020: WheatIS directs repositories to establish standardized access methods enabling users to query distributed data sets securely and efficiently. Caracciolo et al., 2020: The Agrisemantics group recommends that repositories ensure access through machine-readable services and interoperable web standards. “Providing explicit and machine-readable description of data makes it possible to programmatically integrate and reuse data.” Biomedical Sciences The TTRAM Activity/Function AF07 Access is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: One of the key evaluation indicators for ELIXIR Core Data Resources (CDRs) is accessibility. ELIXIR CDRs must be freely and openly accessible to the scientific community, including direct access to data via downloads, APIs, or web interfaces. ELIXIR CDRs must support FAIR principles, among which accessibility.Data and metadata should be retrievable using standardized protocols. However, this methodology does not explicitly discuss secure remote environments or controlled access for sensitive data. It focuses primarily on open access to public data resources. Lin et al., 2024: The article showcases the various access models to data repositories depending on the nature of the data and user requirements. These include: ● Open access, where data is freely available without registration or fees. 54 | Page
● Registration-required access, which allows use after logging in, helping repositories track usage. ● Controlled access, typically for sensitive human data, requiring identity verification and research justification, often reviewed by a committee. ● Data enclaves, a subset of controlled access, where data cannot be downloaded and must be used within secure environments. ● Pay-to-access or donation-based models, which may support sustainability but can conflict with principles of equitable access. ● Closed access, where data is restricted, often for proprietary or commercial reasons. These models reflect how repositories balance accessibility, security, and rights management in alignment with ethical and regulatory frameworks. Sansone & Rocca-Serra, 2016: Addresses optimal interoperability access and use of data and other digital objects is completely automated, and accessible to both human and machine. Mentions the BioSharing/FairSharing (https://fairsharing.org) Barrett et al., 2012: The article describes multiple access mechanisms: ● Direct download of metadata and data via FTP in XML and tab-delimited formats. ● Programmatic access through Entrez Utilities. ● Interactive search and retrieval via the Entrez system and BioProject/BioSample portals. ● Submission of data requires authentication through NIH or NCBI accounts, ensuring controlled access for depositors. ● Clinical samples with privacy concerns are excluded from BioSample and handled via dbGaP, which supports controlled access. Karsch-Mizrachi et al., 2025: The article explains that INSDC provides free and unrestricted access to its nucleotide sequence data through publicly available databases and retrieval tools. The collaboration is built on the principle of open science, and all member organizations commit to making their data resources accessible without barriers. While the article does not go into detail about secure remote environments or differentiated access methods based on user characteristics, it does acknowledge that controlled-access human data is managed separately by each member and is not shared across the INSDC. Wilkinson et al., 2016: In the article, access is defined as a core principle (A) for scholarly digital objects, asserting that (meta)data must be retrievable by their globally unique and persistent identifiers using standardized, open protocols, which can include authentication and authorization procedures when required. This ensures that both humans and computational agents can effectively 55 | Page
access the data and metadata, with metadata remaining accessible even if the data itself is no longer available or is subject to restrictions, guided by clear usage licenses. Rehm et al., 2021: access is facilitated through technical standards and policy frameworks that enable secure and responsible sharing of genomic and health data, supporting both direct data retrieval and analysis within secure federated environments. This include standards like the Data Repository Service API, htsget, and Crypt4GH to manage data access, while using GA4GH Passports and the Data Use Ontology to enforce access criteria based on data sensitivity and user authorization, ensuring alignment with rights management policies including Machine Readable Consent Guidance. Climate Science The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) states that securing remote environments based on digital object and user characteristics, relies on the Earth System Grid Federation (ESGF) infrastructure for data publication and download in the climate modeling community. Crucially, the controlled vocabulary (CV) developed by the METAFOR project (Callaghan et al., 2010) provides standardized, high-level metadata essential for guiding end-users through data discovery, interpretation, and comparison by supporting tools like ESGF gateway interfaces and faceted browsing. The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC ensures data accessibility by offering DataCite DOI publication for archived items, which resolve to a WDCC-hosted landing page serving as an entry point for data access. The access criteria are primarily determined by the selected Creative Commons Licenses. Linguistics For CLARIN B centres, the following requirements apply (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 6.d User access to the repository. Requirement: Each centre should set up a repository (a web-accessible server that offers human access to language resources/services and their metadata). Specify how human access is enabled. ● 3.a The centre should offer data access/sharing for users from other CLARIN ERIC countries. Social Sciences The CESSDA Data Archiving Guide covers the A|F Access in chapter 2.4 Data Access Policies. It describes three levels of access that depend on the type of data and the agreements with data depositors. The levels are: ● Open access to data: Data holdings can be accessed by users at any time and by any means. ● Restricted data access: 56 | Page
1. Standard access: Data holdings are fully anonymized and available for scientific purposes, upon registration. 2. Special conditions access: User may gain access to data following the data depositor’s requirements (e.g. written approval). 3. Access under special licence: Data that require users to complete a detailed special licence form describing in detail the terms of using that data. ● Controlled data access: The archive reserves the right to permit access to data only through a physical or virtual secure environment or after special training. (CESSDA Training Team, 2025a, Chapter 2.4 Data Access Policies). 3.2.6. REUSE (AF08) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Reuse as follows: “Making and keeping digital objects usable and understandable depends on deposit, curation, and preservation activities. Facilitating reuse depends on understanding the needs and expectations of users, including—but not limited to—identified ‘designated communities’, through ongoing external engagement (see AF20 "External Engagement"). Repositories may mediate reuse through secure remote access systems, safe rooms, or other tools that keep the digital objects fully or partially under the service provider's control.”). Repositories may mediate reuse through secure remote access systems, safe rooms, or other tools that keep the digital objects fully or partially under the service provider's control.” and suggests for transparent information “References to artefacts including file formats, metadata and ontologies (potentially already mentioned under Deposit or Curation) that are specifically designed to enable reuse. Designated Community definitions, digital object models.” (L’Hours et al., 2025c, p. 15) including ● Terms and Conditions for Use ● Designated Community Consultation information CoreTrustSeal addresses the A/F Reuse through their requirement R13 with the same name. R13 requires repositories to “ensure that data and metadata continue to be understood and used effectively into the future despite changes in technology and the Designated Community’s knowledge base” and evaluates “the measures taken to ensure that data and metadata are reusable” by asking applicants to provide evidence including references to the following items(CoreTrustSeal Standards and Certification Board, 2022, pp. 23–24): ● The ways in which the repository engages with their Designated Community of users to identify their needs. ● The data formats, metadata schemas, controlled vocabularies and ontologies used to support reuse, and how these meet the community needs. ● The metadata and documentation provided at the point of access to support understandability and reuse appropriate to the Designated Community. This may include information specific to data type, e.g. manuals, calibration records, photos, protocols. 57 | Page
documentation related to how the repository approaches digital object management.”(L’Hours et al., 2025c, p. 16) The A|F Workflows is covered by R11 Workflows in CoreTrustSeal which requires the digital object management of repositories to take place “according to defined workflows from deposit to access” and elaborates on these workflows as follows (CoreTrustSeal Standards and Certification Board, 2022, p. 22): Workflows may be specified in a mixture of standard operating procedures, business process descriptions and diagrams that guide normal practice and provide mechanisms for handling exceptions. The response statement and evidence should include references to the following items: ● Workflows/business process descriptions covering the curation levels performed. ● How workflows are adjusted for different types of data and metadata. ● Decision handling within the workflows. ● Change management of workflows. ● Ability to track workflow execution, with mechanisms to handle exceptions. The functional model of the OAIS Reference Model is a generalized model of workflows relevant to performing digital preservation tasks in an OAIS. It can therefore be considered as a blueprint for specifying and documenting workflows in a given repository (CCSDS 2024, Chapter 4.2). The Desirable Characteristics of Data Repositories for Federally Funded Research (White House Office of Science and Technology Policy, 2022) state that the repository must support authentication of data submitters, has technical capabilities that facilitate associating submitter PIDs with those assigned to their deposited digital objects, such as datasets. It further states that the repository must provide broad, equitable, and maximally open access to datasets and their metadata free of charge in a timely manner after submission, consistent with legal and policy requirements related to maintaining privacy and confidentiality, Tribal and national data sovereignty, and protection of sensitive data. The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates FAIR-enabling practices to be undertaken by all data services should be defined. That FAIR-enabling practices undertaken by data services should be made transparent to users and funders to increase trust in services. Where responsibility is distributed, accountability should remain clear, including accountability for (meta)data loss or destruction. Digital object management outcomes, including preservation, should be integrated into a roles and responsibilities framework that integrates all actors and actions. The roles and responsibilities framework should be aligned with clear process models that meet the needs of different stakeholder communities. Different roles should use a living data management plan as a key artefact for periodic audit, review and revision. 64 | Page
The EU Annotated Grant Agreement states repositories have mechanisms or provisions for expert curation and quality assurance for the accuracy and integrity of datasets and metadata, as well as procedures to liaise with depositors where issues are detected (European Commission, 2024). The D1.3 Recommendations for a FAIR EOSC - White Paper of the FAIR-IMPACT Synchronisation Force (Grootveld, 2025) states that processes for managing, changing, and maintaining digital objects in line with legal, ethical, and policy standards, should be made more transparent in repositories to increase their overall trustworthiness. Furthermore, to develop semantic interoperability, transparent workflows and governance mechanisms are necessary to support the alignment and mapping of Semantic Artefacts (SAs). The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) approaches workflows primarily through its requirement for compliance with community standards across all stages, ensuring that the workflows producing and managing the data/metadata adhere to agreed-upon frameworks. Furthermore, the model requires transparency of information about the provenance and history of the data/metadata. In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), ten out of 14 respondents provided legal and ethical information, eight had metadata and data versioning policies in place, and seven provided descriptions of service processes or workflows. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories report (Kleemola et al., 2022) describes how repositories documented and improved their workflows during the certification process, including ingest, curation, and preservation workflows. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) itself is not a workflow description but acts as a tool to normalize assessment, which can be integrated into institutional workflows by specifying the expected level of FAIRness that resources must achieve within Research Data Management Plans (DMPs) before data is produced. Furthermore, the FDMM includes indicators that implicitly require established data management processes, such as the Essential indicator RDA-A2-01M, which verifies that metadata is guaranteed to remain available even after data is gone, an assurance based on the repository's documented lifecycle information. The Core Preservation Process CPP-013 Object Management Reporting sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA delivers reporting to enable the effective management of Objects. This must include a wide range of functional, operational, and statistical reports and analytics.” (EOSC EDEN T1.2 et al., 2025). Within AF09 Workflows, the Data Curation Network defines the following Curation Activity (Johnston et al., 2016): “Curation Log: A written record of any changes made to the data during the curation process and by whom. File is often preserved as part of the overall record.” 65 | Page
Domain-specific resources Agri-food The TTRAM Activity/Function AF09 Workflows was not found to be addressed in the gathered Agri-Food resources. Biomedical Sciences The TTRAM Activity/Function AF09 Workflows is addressed in the following Biomedical Sciences resources: Yilmaz et al., 2011: MIxS standard defines workflows and activities for the consistent acquisition, validation, and maintenance of metadata during submission to public databases. Barrett et al., 2012: The article outlines submission and metadata workflows that support structured data management and user interaction. However, it does not fully address the broader scope of repository-level workflows related to legal, ethical, and policy compliance Karsch-Mizrachi et al., 2025: The article outlines the processes and governance structures that manage the lifecycle of nucleotide sequence data within INSDC. These workflows include submission, validation, metadata enrichment, and public dissemination, all of which are guided by formalized policies and standards. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC manages digital objects and their associated data and metadata through a comprehensive workflow that includes initial contact, submission agreement, data preparation, metadata submission, data submission, quality control, DOI/PID assignment, and continuous curation. The NetCDF Climate and Forecast (CF) Metadata Conventions (Eaton et al., 2024) state that the CF framework includes an audit trail for modifications through the "history" attribute and a strong commitment to backward compatibility for sustained data utility and long-term maintenance. Linguistics The Activity/Function (A|F) WORKFLOWS is not addressed by any of the reviewed resources from Linguistics. Social Sciences The CESSDA Data Management Expert Guide and Data Archiving Guide give a broad overview of the workflows that need to be carried out and documented throughout the lifespan of research data (CESSDA Training Team, 2022, 2025a). 3.2.8. PRESERVATION (AF10) 66 | Page
Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Preservation as follows: “Monitoring both the technology landscape (see AF29 "Technical Infrastructure") and the user community (see AF20 "External Engagement") for any changes that could impact the usability or understanding (see AF08 "Reuse") of digital objects. When necessary, preservation actions are taken—such as migrating data formats, updating metadata (e.g., revising ontologies), or emulating the environment in which the digital object is used. These actions help ensure the long-term viability and continued accessibility of the digital object.” and suggests for transparent information “[p]reservation policies, strategies, plans and procedures that define how long artefacts will be preserved for, how potential preservation actions (emulation, file format updates, updates to semantic artefacts such as controlled vocabularies and ontologies) are approved and implemented. Preservation interacts with re-appraisal criteria set during Deposit & Appraisal, the changing Technical Infrastructure Landscape over time, and how the ReUse needs of the community are monitored.” (L’Hours et al., 2025c, pp. 16–17) The FIDELIS TTRAM also notes that “[s]pecialist repositories need specific skills in data formats, research methodologies, metadata, and ontologies related to disciplines or content types” and that “[u]ser characteristics influence community watch, object characteristics inform technology watch.” (L’Hours et al., 2025c, p. 17) The proposed CoreTrustSeal Levels of Retention, Curation and Preservation (LoRCaP) define requirements for Active Preservation as follows: “In addition to [D. Deposit Compliance and/or C. Initial Curation] the repository takes long-term responsibility for ensuring that the data and metadata can be understood and rendered as required by the designated community for reuse. The preservation actions can be aimed at logical-technical, semantic, or quality aspects of the (meta)data, for example, in response to the threat of technological obsolescence, to accommodate changing needs of the Designated Community, or in response to other considerations such as security or legal concerns” (L’Hours et al., 2024, p. 4). Metadata required on objectand repository-level to demonstrate how the LoRCaP are implemented are proposed in L’Hours et al., 2024. In CoreTrustSeal, the A|F Preservation is dealt with in requirement 9 (RO9), called Preservation plan, which requires repositories to “assume[...] responsibility for long-term preservation and manage[...] this function in a planned and documented way” and lists the following issues to be addressed by the repository (CoreTrustSeal Standards and Certification Board, 2022, pp. 19–20): ● The documented approach to preservation, including whether this involves format migration, emulation, etc. Ensuring bit level integrity is vital but not sufficient for preservation Ensuring bit level integrity is vital but not sufficient for preservation. ● File formats and metadata schemas for long term preservation. ● How the level of responsibility for the preservation of each item is defined. ● Plans related to future migrations or similar measures to address the threat of obsolescence. ● Actions relevant to preservation specified in documentation, including custody transfer, submission information criteria, and preservation information metadata. 67 | Page
● Measures to ensure these actions are taken. ● Any minimum stated retention and/or preservation periods. ● How often the digital objects are re-appraised and the possible outcomes of reappraisal. ● The repository approach to deleting/removing data and metadata from collection/holdings including the impact on persistent identifiers. In the nestor Seal criteria “C11 Preservation measures”, “C13 Significant properties” and “C24 Interpretability of the archival information” are relevant to this A|F. C11 states that the archive “should conduct strategic planning as a means of preserving the digital objects entrusted to it [...]. Long-term planning should be based on the monitoring of legal and social changes, the demands and expectations of the designated communities and all technical changes relevant for the sustained preservation and appropriate use of the information objects in the form of their representations”. C13 focuses on the properties “significant for preservation of the information objects” and C24 addresses “technical preservation measures [...] undertaken to ensure the interpretability of the archival information packages” (nestor Certification Working Group, 2025). In the OAIS Reference Model core preservation activities are performed by the Preservation Planning functional entity and the Administration functional entity. Preservation Planning (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2.3.7) includes a “Preservation Watch” function which compiles all information necessary to make informed preservation decisions. It receives input from a number of internal and external sources, including the “Monitor Designated Community” and the “Monitor Technology” functions. The former ensures that the archive continues to understand the needs and skills of the users who are the primary target group of the preserved digital objects while the latter guarantees that the archive has up to date information about developments in the technological landscape which may affect the renderability of the digital objects it holds. In addition, Preservation Planning is tasked with developing preservation strategies and standards, performing risk assessments, and developing preservation plans among other things. Administration in turn is tasked with actually performing any updates to be made to the archived objects (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2.3.6). Chapter 5 of the OAIS Reference Model addresses “Preservation Perspectives”, namely, “various practices that have been, or might be, used to preserve digital information and to preserve access services to digital information” (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 5/1), including different types of migration. The Desirable Characteristics of Data Repositories for Federally Funded Research (White House Office of Science and Technology Policy, 2022) address long-term organizational sustainability when the repository has a plan for long-term management of data, including maintaining integrity, authenticity, and availability of datasets; has contingency plans to ensure data are available and maintained during and after unforeseen events. The report FAIR Forever? Long Term Data Preservation Roles and Responsibilities (Currie & Kilbride, 2021) states that research repositories should urgently prioritise and adapt workplans to include quality improvement mechanisms where these do not already exist, including DPC Rapid Assessment 68 | Page
Model, establishing thereby a strategic framework to achieve baseline certification for primary preservation services, or identifying preservation pathways for data. For Research Repositories there is a medium priority, identifying costs of action versus inaction with respect to high value, critically endangered content. The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates that effective storage, including multi-copy redundancy and integrity measures, is necessary but not sufficient for preservation. Minimum criteria for acceptable storage practices in different scenarios should be defined as a foundation for all levels of retention, curation and preservation services. Different types of data services benefit from being transparent on their current level of storage, curation and preservation practice, as this increases trust by the user and funders alike. Data services, including repositories should specify all the levels of care they apply to objects within their collection, including through repository and digital object registry metadata. Digital objects should include metadata that specify their level of care and the timeframes or criteria for reappraisal of the level of care. Unique preservation functions and activities should be defined alongside functions and activities that apply to all (meta)data services. Preservation roles must include monitoring the changing needs of communities at the point of reuse. This community watch must be aware of the knowledge base, methodologies and technologies of the user communities. Preservation roles must include monitoring the changing nature of available technologies for the deposit, storage, curation, discovery, access and reuse of data and metadata. This technology watch must continue to meet the needs identified through community watch and be proactive as well as reactive. Managers must integrate preservation planning into operational management including staffing, funding, service development and procurement. The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term. They further facilitate midand long-term preservation of the deposited material (European Commission, 2024). The Update of the Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements states that the lack of a public policy for preservation, curation and security of the contents is the most frequent reason, followed by not adhering to a specific metadata standard (Lazzeri, 2024). The Practical Guide to the International Alignment of Research Data Management (Science Europe, 2021) states that: 1. Preservation: a. Ensure persistence of metadata and data b. Be transparent about mission, scope, preservation policies, and plans (including governance, financial sustainability, retention period, and continuity plan). 69 | Page
The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on data preservation policy that details how the preservation of the data is ensured. Does the repository provide information on its data preservation policies? The Research Data College Working Group recommends “prioritising maintained repositories with a data-preservation period of at least five years, following the practices put in place by Recherche Data Gouv” (Lamotte et al., 2024, p. 10). In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), respondents were asked whether their organisation employed any of the common digital preservation strategies e.g., format and schema migration and/or emulation. All three of the non-repository service providers selected ‘Can’t say or N/A’ which was expected, as they do not provide long-term digital preservation. Three of the repositories selected ‘Can’t say or N/A’; seven of the repositories selected ‘format and schema migration’ and one repository selected ‘both’. The D3.1—Report on Discipline Requirements and Needs (Andreassen et al., 2025) states in section 4.3.5, Preservation Objective/Designated Community Needs, that although all disciplines indicate that expert curators are utilized across all disciplines, variability in the level of curation requires that a standard level of acceptable curation that enables long-term digital preservation (LTDP) and digital object quality (DOQ) must be promoted. Across all disciplines, standardized and comprehensive strategic reappraisal and reassessment schedules and guidelines must be established and implemented. The existing tendency to follow long-standing practices or the simple statement that data will be preserved “forever” is inadequate and comprehensive strategies must be put in place. Regular, standard reappraisals and reassessment checks, that include methodologies for enhancement (e.g. file format, metadata enrichment), are required to ensure LDP and DOQ for both historical and legacy data, as well as relevant standards. The relevance of older data and standards is not sufficiently assessed across many disciplines, but can be highly relevant for state-of-the-art science that includes, for example, field recordings that capture linguistic variables, sampling of targeted organisms, or time-series climate data. Older data can be vital to understand temporal changes in many disciplines, and the rapid evolution of technology presents new challenges for both evaluating and supporting historical and legacy data plus standards. Concerning identifying needs in the designated community (community watch), developing a targeted, inclusive, and systematic approach to community engagement should be considered within the disciplines, preferably as a joint effort between repositories to reduce duplication of work. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that repositories express the levels of care offered and received by digital objects, detailing what levels of retention, curation, and preservation are in place and how these might change over time. Additionally, the guidelines conclude that digital objects not undergoing active preservation to address evolving technology (e.g., file format migration) or community changes (e.g., ontology updates) risk a deteriorating FAIR score over time. 70 | Page
The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that long-term preservation is a core requirement of TDRs. The report discusses OAIS-based responsibilities and the need for documented preservation policies. Some repositories lacked explicit preservation missions, which limited their eligibility for certification. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) emphasize persistence and preservation of metadata, notably through the Essential indicator RDA-A2-01M, which mandates that metadata is guaranteed to remain available even after the data itself is no longer accessible, a requirement evaluated based on the organization's documented life cycle information. Furthermore, the FDMM’s focus on Active preservation (A) within the Levels of Retention, Curation and Preservation (LoRCAP) variable directly supports this function by confirming that active preservation is defined as the long-term responsibility for ensuring that data and metadata can be understood and rendered for reuse by the designated community. The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) states that Preservation requires repositories to establish a digital preservation plan that outlines the duration of resource management, identifies responsible roles, and documents procedures for preserving different resource formats to ensure long-term usability. To enable necessary preservation actions, repositories must secure the rights to copy, transform, and store items from the depositor, and maintain a business continuity plan detailing procedures for managing disruptions like cyber-attacks or natural disasters. The GREI Data citation best practices for repositories (Puebla et al., 2024) require a standardized approach to tracking data usage, mandating the consistent use of DataCite metadata fields to store citation relationships. Furthermore, the recommendations ensure future access to this contextual information by requiring repositories to expose the provenance (source) of every asserted citation and by encouraging contribution to centralized resources like the Data Citation Corpus. The Core Preservation Process CPP-012 Risk Mitigation sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA enables the design, development and management of plans for mitigating identified preservation risks.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-014 File Migration sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA supports batch modifications of previously ingested Files to prevent a preservationor access-related risk.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-015 Emulation and Rendering Tools sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA enables the rendering of Objects via the application of emulation and/or other specialist tools.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-017 Disposal sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA enables the managed disposal of 71 | Page
Information Packages and permits the retention and maintenance of Metadata even when the content of the Information Package has been removed from the TDA.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-022 Significant Properties Definition sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA defines significant properties for sets of Information Objects (i.e., properties that it commits to preserve over the long term through preservation actions and rendition).” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-023 Risk Definition and Extraction sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA monitors technology evolution, identifies risks and defines detection methods for these risks.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-026 File Format Normalisation sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA performs operations on the data Objects prior to ingest in order to comply with its format requirements.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-028 Creation of Derivatives sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA generates derivative copies to address the specific needs of its Designated Community.” (EOSC EDEN T1.2 et al., 2025). Within AF10 Preservation, the Data Curation Network defines the following Curation Activities (Johnston et al., 2016): ● “File Format Transformations: Transform files into open, non-proprietary file formats that broaden the potential for long-term reuse and ensure that additional preservation actions might be taken in the future. Note: Retention of the original file formats may be necessary if data transfer is not perfect.” ● “File Inventory or Manifest: The data files are inspected periodically and the number, file types (extensions), and file sizes of the data are understood and documented. Any missing, duplicate, or corrupt (e.g., unable to open) files are discovered.” ● “Software Registry: Maintain copies of modern and obsolete versions of software (and any relevant code libraries) so that data may be opened/used overtime.” ● “Transcoding: With audio and video files, detect technical metadata (min resolution, audio/video codec) and encode files in ways that optimize reuse and long-term preservation actions. (E.g, Convert QuickTime files to MPEG4).” ● “Emulation: Provide legacy system configurations in modern equipment in order to ensure long-term usability of data. (E.g., arcade games emulated on modern web-browsers.” ● “Migration: Monitor and anticipate file format obsolescence and, as needed, transform obsolete file formats to new formats as standards and use dictate.” Domain-specific resources Agri-food 72 | Page
The TTRAM Activity/Function AF10 Preservation was not found to be addressed directly into the gathered Agri-Food resources. However, the following statements can be mentioned: Harper et al., 2018: AgBioData recommends that repositories develop long-term preservation and sustainability plans, including metadata and software curation, to prevent data loss. “Long-term sustainability and maintenance of databases is critical to ensure continued access to data and tools; mechanisms must be put in place to guarantee that curated data and metadata remain available to the community.” ; and also positions metadata as crucial to “long-term management of persistence and relevance, adaptation to changes in technologies, support of old data formats or conversion into new and continued integration with new data.” Caracciolo et al., 2020: The Agrisemantics group instructs repositories to maintain and update semantic artefacts and controlled vocabularies to preserve the interpretability of data. “Providing explicit and machine-readable description of data makes it possible to programmatically integrate and reuse data; keeping semantic resources updated ensures long-term interpretability.” Biomedical Sciences The TTRAM Activity/Function AF10 Preservation is addressed in the following Biomedical Sciences resources: Lin et al., 2024: In the article, Preservation, defined as “the extent to which the data repositories invest resources in archiving data for long-term use, including adapting to evolving user needs, changes in storage technology, and changes in media formats” is considered as one of the six major characteristics of the different types of data repositories. Besides, preservation is considered a mandatory responsibility as long as the project or institutional repository is relevant to its user base and mission. Therefore, the article encourages the adoption of principles such as TRUST and more specifically the Transparency and Sustainability practices. Field et al., 2011: GSC proactively monitors the evolving technology landscape, such as the democratization of sequencing access, and sociological changes within the scientific community. This vigilance ensures the continuous development and harmonization of shared standards. Karsch-Mizrachi et al., 2025: The Membership Arrangement described in the article states that INSDC Member institutions are independent government or non-profit organizations that manage nucleotide sequence databases that capture, preserve, and present comprehensive nucleotide sequence information and annotations to preserve the scientific record and enable broad sharing of such data. Additionally, INSDC has identified several areas of data management that are currently undergoing evolution or emerging as areas in which new development is needed to meet the needs of data submitters and users or to ensure the sustainability of sequence repositories. 73 | Page
pipeline, including Schematron-based validation, to guarantee internal coherency between the many descriptive parameters within the harvested metadata. The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that information regarding provenance is managed through versioning in WDCC. The NetCDF Climate and Forecast (CF) Metadata Conventions (Eaton et al., 2024) state that provenance is achieved through metadata attributes such as "history", which offers an audit trail for modifications to the original data, and "institution", "source", and "references", which detail the origin and production methods, thereby directly supporting the provenance and authenticity of the digital objects. Additionally, metadata for data quantization aims to provide provenance information so users can reproduce data transformations and understand how the data might differ from its original unquantized state. Linguistics The Activity/Function (A|F) Provenance and Authenticity is not addressed by any of the reviewed resources from Linguistics. Social Sciences The CESSDA Data Archiving Guide covers the A|F Provenance & Authenticity in Chapter 4.5 Updates and versioning. The guide strongly advises repositories to track the provenance and history of all files they receive and store. It is crucial to identify the current version of the data, and the guide presents structured versioning practices - such as changelog folders and versionor date-based file naming - to ensure that data is clear and reproducible. (CESSDA Training Team, 2025a, Chapter 4.5 Updates and versioning) Repositories are recommended to include data versioning and changelogs, and ensure access to metadata about all published versions of datasets, which contribute to documenting the history and authenticity of data. (Bornatici et al., 2025, p. 7). 3.2.10. SUPPORT (AF12) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Support is described as follows: “Providing guidance and responding to requests from depositors and users around deposit and appraisal, discovery, access and re-use of data and metadata.”, suggests for transparent information “[l]ink to primary location where the repository provides supporting information or offers supporting services to depositors, users and others.” and notes that “[s]pecialist repositories need specific skills in data formats, research methodologies, metadata, and ontologies related to disciplines or content types, to deliver relevant support.” (L’Hours et al., 2025c, pp. 17–18) 80 | Page
The A|F Support is not explicitly addressed by CoreTrustSeal, but has affinities with their requirement R06 Expertise and Guidance, which is discussed in section PEOPLE & EXPERTISE (AF17) below. In the OAIS Reference Model, interaction with producers is managed by the Administration functional entity (“Customer Service”, see Chapter 4, p. 14). The function “Coordinate Access Activities” in the Access functional entity “provides assistance to OAIS Consumers including providing status of orders and other Consumer support activities in response to an assistance request via the Deliver Response function.” (Chapter 4, p. 19). Beyond defining these functions, the OAIS Reference Model does not state any requirements for how such support functions should be implemented. PAIMAS provides detailed suggestions as to which information should be exchanged and which points clarified between the producer and the archive before Ingest (Consultative Committee for Space Data Systems (CCSDS), 2004, Chapter 4). The Desirable Characteristics of Data Repositories for Federally Funded Research (White House Office of Science and Technology Policy, 2022) address the topic of clear use guidance when the repository ensures datasets are accompanied by documentation describing terms of dataset access and use (e.g., reuse licenses and need for approval by a data use committee). Furthermore, broad and measured reuse is when the repository ensures datasets are accompanied by metadata that describe terms of reuse and provides the ability to measure attribution, citation, and reuse of data (e.g., through assignment of adequate and openly accessible metadata and unique PIDs). The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) states on user support that provides support to users during or after submission. Does the repository have a contact point (e.g. helpdesk email or contact form) to assist data depositors and data users? In the FAIRsFAIR project D2.3 report (Behnke et al., 2020), a strong emphasis is made on the APIs for which active support must be provided to users as documentation or training taught by the repository's staff. Besides, technical support for predefined file formats must be provided, as well. The D3.1—Report on Discipline Requirements and Needs (Andreassen et al., 2025) states in section 4.3.3, Metadata Quality, that direct communication between the repository and the depositor needs to be available, with increased communication about the implementation of, clarification/explanation of, and compliance with metadata standards. The communication or explanation needs can be expressed as both monitory, which includes repositories having one-on-one contact with depositors, and automated, with respect to improving workflows, standardizing metadata annotation and facilitating completeness. 4.3.5 Preservation Objective/Designated Community Needs: Researchers wish to deposit data efficiently and in line with good practice, and having access to an active support desk and/or a contact point inside the repository, where someone takes the time to answer, is strongly wanted. Interaction between repository personnel and depositor cannot be neglected, even when elaborated repository user guides exist. For the researcher, who often deposits data at a very irregular frequency, having someone with trustworthy digital repositories competencies who can give advice and answer 81 | Page
questions, is highlighted as very important for trust and user experience. For the repository, this human contact point aids monitoring the needs in the community. Common repository personnel-directed FAQs within disciplines may improve and facilitate live guidance, as well as reducing time spent formulating advice and answers. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) propose incorporating guidance directly into Data Management Planning (DMP) tools to help users make informed choices about data deposit concerning access, storage, curation, and preservation from the earliest stage of their research. Furthermore, transparency about the levels of care offered by repositories (retention, curation, and preservation) supports mutual trust, which is a critical precursor to trusted relationships between actors such as object depositors and users. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) provides one-on-one trust support work, webinars, and documentation templates. The repositories involved found the support process useful. Peer support and expert feedback were crucial for helping repositories navigate certification. Managing trust in the future entails peer collaboration through support programmes and networks, while ensuring solid resources for the upkeep of these collaborative efforts. Future endeavours to manage trust should make use of the existing and planned networks of trustworthy repositories that can share both expertise and responsibility, while recognising the need for more enduring sources of funding for managing trust sustainably. Some individual ERICs, such as CESSDA and CLARIN, provide targeted support for their members on issues related to trust and seeking compliance The GREI Data citation best practices for repositories (Puebla et al., 2024) conclude that repositories should provide guidance and information to users and contributors regarding the mechanisms used to manage data citations. Specifically, this guidance should be documented as part of the information pages for contributors and users, detailing how the repository collects and stores data citations, including the external sources it employs. Domain-specific resources Agri-food The TTRAM Activity/Function AF12 Support is addressed in the following Agri-Food resource: Harper et al., 2018: AgBioData explicitly recommends that repositories provide personal and responsive support to users and contributors to build trust and maintain communication. (See AF01 : Identification & Contact). Biomedical Sciences The TTRAM Activity/Function AF12 Support is addressed in the following Biomedical Sciences resources: 82 | Page
Durinx et al., 2017: A key component of quality of service for ELIXIR Core Data Resources is the provision of effective customer support and user engagement mechanisms. This includes operating a helpdesk to assist users with queries and technical issues, thereby ensuring accessibility and responsiveness. In addition, resources are expected to actively seek and incorporate user feedback into service design and improvement processes, fostering a user-driven approach to development. Finally, the provision of training activities—such as workshops, tutorials, or documentation—demonstrates a commitment to empowering users and enhancing their ability to effectively engage with the resource. Lin et al., 2024: In the article, User diversity is defined as a way to “accommodate a diverse audience, offering resources and support for users across multiple disciplines and skill levels” is considered as one of the six major characteristics of the different types of data repositories. Therefore, the article encourages the adoption of principles such as TRUST and more specifically the User focus practices. Field et al., 2011: GSC provides support for digital object management through its compliance working group to help depositors adhere to the MIxS standard. Additionally, its developer's working group provides technical assistance for implementing GSC standards in software and database projects. Barrett et al., 2012: Minimal support - BioProject and BioSample [...] submissions are supported by a web-based Submission Portal that guides users through a series of forms for input of rich metadata describing their projects and samples. Wilkinson et al., 2016: The FAIR principles provides intrinsic guidance and facilitates autonomous discovery, access, and reuse for both human and computational agents. Rehm et al., 2021: GA4GH provides guidance for data creation, deposit, discovery, access, and reuse across the genomic data lifecycle. Additionally, user support is achieved via the GA4GH Starter Kit and through community engagement initiatives such as the Genomics in Health Implementation Forum (GHIF) to foster understanding and share best practices among depositors and users. Karsch-Mizrachi et al., 2025: The article does not describe a formal helpdesk or user support system. However, it demonstrates that INSDC provides structured guidance and documentation to facilitate deposit, appraisal, discovery, access, and reuse of data and metadata. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC offers comprehensive support and guidance to depositors for the entire data publication process, encompassing initial data appraisal, preparation (e.g., format, structure), metadata submission, quality control, and assignment of persistent identifiers (DOIs/PIDs). Linguistics The Activity/Function (A|F) Support is not addressed by any of the reviewed resources from Linguistics. 83 | Page
Social Sciences When promoting data sharing, archives should highlight benefits such as: ● merits for data sharing (research visibility and increased recognition through higher citation rate); ● data from publicly funded research should be published and reused; ● long-term preservation of their own research data; ● opportunity to inspire further research and hands-on learning for students; ● contribution to quality, transparency and accountability of research; ● compliance to funders’, institutional, or journals’ data policies; ● address researchers’ fear of misinterpretation, misuse, and fear of losing control over data; ● address practical constraints that researchers perceive as hindrances to proper data management. Generalist repositories usually allow quicker data publication, which is why it may be beneficial to emphasize the benefits of using trusted domain-specific repositories in the social sciences, as they offer added value in terms of data quality, curation and preservation processes. This ensures data is more discoverable, reusable, and less prone to misinterpretation. (CESSDA Training Team, 2025a, Chapter 3.2.2 Advocate for data sharing). 3.3. Organisational Infrastructure Organisational Infrastructure encompasses governance, policy development, resource allocation, staffing, partnerships, and continuity planning, ensuring the repository’s operational sustainability and alignment with its mission. 3.3.1. GOVERNANCE (AF13) Domain-agnostic resources The FIDELIS TTRAM the repository Activity/Function (A|F) Governance is described as follows: “The organisational hierarchies and processes by which the entity's mission is managed and executed.” and suggests for transparent information “[m]anagement and Decision documentation, organograms, management roles, role of host organisation, funders etc.” (L’Hours et al., 2025c, p. 19) The A|F Governance is dealt with in CoreTrustSeal requirement R05 Governance & Resources, which requires repositories to have “adequate funding and sufficient numbers of staff managed through a clear system of governance to effectively carry out the mission” (CoreTrustSeal Standards and Certification Board, 2022, p. 16). For the Governance part of this requirement, CoreTrustSeal asks for applicants to provide evidence through “descriptions and diagrams of governance bodies, groups and hierarchies” (CoreTrustSeal Standards and Certification Board, 2022, p. 16). The nestor Seal states in C10 Organisation and processes: “The organisational structure should be appropriate for the objectives, tasks and processes of the digital archive. The structural and procedural organisation should be defined. The responsibilities should be established. The digital 84 | Page
archive is incorporated at the appropriate point in the schedule of responsibilities” (nestor Certification Working Group, 2025). The Functional Model of the OAIS Reference Model details the organization, processes and workflows of an Open Archival Information System (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2). The Desirable Characteristics of Data Repositories for Federally Funded Research policy (White House Office of Science and Technology Policy, 2022) emphasizes that a repository's "Organizational Infrastructure" must include Long-term Organizational Sustainability through comprehensive plans for data management, integrity, and authenticity, as well as Risk Management capabilities with documented safeguards for sensitive data. The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates that transparency about and analysis of current roles and responsibilities associated with data services functions and activities are necessary inputs into financial calculations related to salaries and funding streams. Identify and support staffing costs for preservation specific roles and responsibilities. Any decision to store, retain and then undertake initial curation will imply costs. Any additional costs needed to deliver active preservation are uniquely preservation costs. The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term (European Commission, 2024). The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states that aligned with institutional research policies. Integrate repository management into institutional data policies, funding allocation, and strategic planning. Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022) does not cover this AF – or at least not explicitly. However, the fact that most of the respondents had a business continuity plan as well as a preservation policy, procedure or plan in place, legal and ethical information, rights information and/or licences, metadata and data versioning policies, and descriptions of service processes or workflows suggest strong governance. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) conclude that, in the context of trust and transparency, repositories should be transparent about granular characteristics such as their governance structure. Furthermore, if a mechanism for information exchange, such as the one proposed by the guidelines, is used to align a network of Trustworthy Digital Repositories (TDRs), it will be the decision of the Network governance or decision-making body to set any expectations around the sharing or content of specific metadata fields. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that modern repositories are often developed through 85 | Page
partnerships and with a wide range of insource and outsource options. Trust between these different data service actors is essential for the data ecosystem, and future assessment and certification options and networks should cover data service providers and the issue of enabling FAIR data. Governance structures were evaluated as part of the certification process. Repositories were encouraged to clarify roles, responsibilities, and decision-making processes, especially in complex partnership models. The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) states that adhering to good practices regarding governance, a repository must clearly indicate the organization responsible for its management and the nature of its governance. Additionally, successful management and execution of the entity's mission require a publicly available policy stating what will happen to resources if operations cease. Domain-specific resources Agri-food The TTRAM Activity/Function AF13 Governance was not found to be addressed in the gathered Agri-Food resources. Biomedical Sciences The TTRAM Activity/Function AF13 Governance is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: A strong legal, funding, and governance infrastructure is essential for the long-term sustainability of ELIXIR Core Data Resources. One key element is the presence of an independent, international Scientific Advisory Board, which provides expert guidance, ensures scientific relevance, and supports strategic decision-making. In addition, resources must demonstrate sustainable support and funding, including evidence of past and ongoing commitments from host institutions and other funding bodies. These elements collectively ensure that the resource is not only scientifically robust but also institutionally and financially resilient. Field et al., 2011: GSC comprises a board, several standing committees, and working groups, with its mission managed and executed through face-to-face meetings, the formation of working groups, and the development of consensus products/standards. Yilmaz et al., 2011: MIxS standard is engaged by a network of experts from sequencing centers, major resource maintainers, and individual investigators. This governance includes structured processes such as community surveys for standard development, ensuring adoption by key data providers. Karsch-Mizrachi et al., 2025: The article describes the formal organizational structure of INSDC, which includes the establishment of two key committees: the Executive Committee (EC), responsible for strategic direction, and the Implementation Committee (IC), responsible for operational policies and procedures. This governance model was formalized through the signing of the Founders 86 | Page
Arrangement in 2023, which codifies the collaboration among the founding members and outlines their mutual responsibilities. The article also discusses the Membership Arrangement, which sets expectations for all members regarding data stewardship, access, and compliance. These structures and agreements demonstrate how the INSDC manages and executes its mission through defined hierarchies and processes. Rehm et al., 2021: GA4GH is a structured alliance with eight Work Streams (WSs) involving foundational groups for regulation, ethics, and data security, and technical groups that develop standards and policy frameworks driven by the needs of twenty four real-world genomic data initiatives (Driver Projects). Products are proposed to the Steering Committee, undergo review by specialized Work Streams and a Product Review Committee. Climate Science The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) states, for long-term management and execution of the CV as a standard, the governance structure includes the intent to establish an international governance committee under the auspices of the IS-ENES2 project to manage the CV's evolution and preservation. The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC is hosted by the DKRZ and maintained by its data management department, which is responsible for collecting, archiving, and disseminating Earth System data. Linguistics The Activity/Function (A|F) Governance is not addressed by any of the reviewed resources from Linguistics. Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “Organisation: Organisational Structure” that are relevant to the topic of Governance (https://www.cessda.eu/Resource-Directory?tree=17,22). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.3.2. POLICY & STANDARDS (AF14) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Policy & Standards is described as follows: “The adoption and development of policies, standards and other criteria to guide practice, aligned with governance structures (see AF13 "Governance") and operational workflows (see AF09 "Workflows"), and in compliance with the applicable legal and ethical framework (see AF23 "Legal 87 | Page
and Ethical").” and suggests for transparent information “[p]rocedures for adopting and developing policies and standards, lists of policies and standards.” (L’Hours et al., 2025c, pp. 19–20) In the OAIS Reference Model, it is a function of the Administration functional entity to “Establish Standards and Policies” based on input from other functions and functional entities (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 4/13). Among the standards and policies mentioned are “format standards, documentation standards and the procedures to be followed during the Ingest process” and “approved standards and Preservation Objectives [...] for example specifying goals relating to Transformational Information Properties and usability of the Content Information” (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 4/13). The Desirable Characteristics of Data Repositories for Federally Funded Research policy (White House Office of Science and Technology Policy, 2022) emphasizes that a repository's organizational infrastructure must include documented policies for data retention, provide clear use guidance detailing access and reuse terms, and implement risk management capabilities that comply with applicable confidentiality and monitoring requirements. The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) suggests that repositories adopt community-defined standards to ensure the consistent and machine-readable representation of data and metadata. These standards—encompassing models, formats, schemas, vocabularies, and ontologies—play a critical role in facilitating the discovery, accessibility, and interpretation of data and metadata. Considerations include: ● Data and Metadata Standards: What data and metadata standards, if any, has the repository implemented to enhance interoperability and usability? ● Persistent Identifiers (PIDs): Does the repository assign globally unique and persistent identifiers to its holdings? If so, which PID schema (e.g., DOI, Handle, ARK) is utilized to ensure long-term accessibility and traceability? ● Citation and Linking to Related Publications: Does the repository provide mechanisms to link datasets to related articles, preprints, or other publications? At what stage of the data deposition process is information about related articles required? ● By adhering to these standards and practices, the repository aims to support researchers in managing, sharing, and citing their data effectively, while fostering a robust ecosystem for data discovery and reuse. The O'FAIRe makes you an offer: Metadata-based Automatic FAIRness Assessment for Ontologies and Semantic Resources (Amdouni et al., 2022) focuses on technical compliance with the FAIR principles for semantic resources, it indirectly addresses the operationalization of standards by noting that high average scores for Findability (F) (70) and Accessibility (A) (90) reflect the success of repository policies in adopting standards like URIs and the HTTP protocol for access. The Global Community Guidelines for Documenting, Sharing, and Reusing Quality Information of Individual Digital Datasets (Peng et al., 2022) approaches policy and standards activities by making Essential the requirement that both metadata and data comply with domain-relevant community 88 | Page
standards (R1.3-01M, R1.3-01D) and that metadata adheres to a machine-understandable community standard (R1.3-02M), ensuring that policy outcomes are directly reflected in the technical quality and usability of the digital objects. These indicators establish a minimum compliance baseline, requiring repositories to adopt standards that facilitate the reuse of their resources by the designated community. In the FAIRsFAIR project D2.3 report (Behnke et al., 2020), explicit data policies (like versioning and dynamic data) and PID policies are recommended, together with global PID for each digital object or file. Besides, an explicit data deletion policy with the roles and responsibilities in human and machine interoperable ways must be provided to users. A tombstone procedure must be put in place, as well in case the repository ceases to operate. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that transparent information exposed by repositories should take account of, and map to, existing standards and criteria, such as DCAT, schema.org, and DataCite, to minimize divergence and maximize interoperability. Furthermore, specific information consumers (like F-UJI) can design their own requirements or standards regarding the content of information exposed, which can then be communicated to their audience. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories report (Kleemola et al., 2022) promotes the adoption of CoreTrustSeal, OAIS, ISO 16363, and TRUST principles. It also discusses the need for community-agreed definitions and minimum standards for data services. Feedback on the CoreTrustSeal certification pointed to the need to clarify its guidance and to define in even more detail the necessity to include evidence and documentation rather than spell out the procedures in the assessment. The CoreTrustSeal expectations are appropriate for SSH repositories seeking TDR status and no clashes between CoreTrustSeal and more specific SSH requirements were found. There is no apparent demand for anything more detailed and extensive than CoreTrustSeal at the moment. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) strongly emphasize compliance with standards, establishing several Essential indicators that require metadata and data to comply with domain-relevant community standards (RDA-R1.3-01M, RDA-R1.3-01D) and to use standardized protocols for access (RDA-A1-04M, RDA-A1-04D), thus providing specific criteria for assessing a repository's adherence to standards in practice. Furthermore, the FDMM itself serves as a standardizing framework, providing a common set of indicators and priorities intended to normalize assessment of FAIRness, which informs policy adoption and compliance efforts for data providers, publishers, project managers, and funding agencies. The GREI Data citation best practices for repositories (Puebla et al., 2024) mandate the use of DataCite metadata fields (relatedIdentifier, relationType, etc.) to store standardized citation relationships. This policy is aligned with broader community criteria, drawing upon the principles 89 | Page
Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC manages permissions, prohibitions, and obligations governing how digital objects are accessed and used by individuals and organizations primarily through licensing agreements. Data providers must decide under which license to publish their data. Linguistics CLARIN B and E centres are required “to make clear statements about their policy of offering data and services and their treatment of IPR issues” (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3). For CLARIN B centres, the following requirement applies (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 3.a Data offering & IPR. Requirement: Each centre needs to make clear statements about their policy of offering data and services and their treatment of IPR issues. CLARIN centres are recommended to offer their resources and services according to formal legal terms and conditions, which include (CLARIN Legal and Ethical Issues Committee, n.d.): 1) The general terms of service regulating access to CLARIN resources; 2) The end-user licenses further specifying the access conditions for each resource; 3) The deposition agreements that are signed by resource providers when depositing a resource to a CLARIN centre. These are formulated in a number of legal documents that have been made available by the Finnish CLARIN consortium (see links below). Terms of Service ● The general conditions under which end-users may access the resources distributed by CLARIN are stated in the CLARIN Terms of Service. ● The model document CLARIN-TOS-v1.0 includes the legal terms and conditions holding between a CLARIN centre and its end-users with regard to resource access. End-User Licenses ● All resources distributed by CLARIN are accompanied with licenses, which may include more specific conditions than those stated in the CLARIN Terms of Service. In fact, the same language resource can be distributed with more than one license, depending on the end user's role or intended use. ● In order to facilitate the management of language resources, CLARIN uses a classification system (see below), according to which licenses are divided into three main categories: ○ CLARIN PUB ('public use') 96 | Page
○ CLARIN ACA ('academic use') ○ CLARIN RES ('restricted use'). ● In addition, CLARIN uses a set of labels (aka 'laundry tags') that correspond to conditions of use that are frequently associated with the distribution of language resources. Deposition License Agreements ● When a resource provider (aka 'depositor') wants to include a new resource into the CLARIN domain, a specific deposition license agreement must be made between the depositor and the CLARIN centre that will host the resource. ● In the deposition agreement, the depositor also specifies the license(s) with which the resource will be distributed to end-users. Depending on the category under which the selected end-user license falls (see previous section), a minimal set of access conditions have to apply. These conditions are stated in the following three model agreements; additional or more specific conditions can also be agreed upon between the centre and the resource depositors. The minimal model agreements below can be used as checklists. ○ CLARIN-DELA-PUB-v1.0 (for public use resources) ○ CLARIN-DELA-ACA-v1.0 (for academic use resources) ○ CLARIN-DELA-RES-v1.0 (for restricted use resources). Social Sciences The CESSDA Data Access Policy (Bolton, 2022) requires that an agreement between the SP and the Data Owner covering data access arrangements must be in place for each data collection within each SP’s holdings (Principle 3). The CESSDA Data Archiving Guide addresses the A|F Rights in several of its chapters. First, it states that archiving processes in European archives are governed by both national laws and international regulations, and that the main legal concerns include copyright law and the sharing of personal data (i.e. GDPR). These frameworks impose obligations on how digital objects are managed within repositories. (CESSDA Training Team, 2025a, Chapter 1.9 What are relevant legislations concerning data archiving?). Next, the guide discusses data access. Data access and data usage permissions are managed through access conditions, and the guide outlines three levels of access, depending on the nature of the data and the agreements made. The levels are: (1) Open access to data, (2) Restricted data access and (3) Controlled data access. (CESSDA Training Team, 2025a, 2.4 Data Access Policies). Finally, the guide addresses data dissemination principles. Data dissemination principles reflect the archive’s position on safeguarding and sharing data. These principles can also refer to explicit policies and legal framework, such as GDPR compliance. (CESSDA Training Team, 2025a, 2.5 Data Dissemination Policies and Principles). 3.3.4. RESOURCES (AF16) 97 | Page
Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Resources is described as follows: “The organisational hierarchies and processes by which the organisational entity's human and financial resources are obtained and managed. This includes a knowledge of and an effective deployment of human, financial and energy assets.” and suggests for transparent information “[b]udget and finance strategies, plans, reports and projections.” (L’Hours et al., 2025c, p. 20) The A|F Resources is dealt with in CoreTrustSeal requirement R05 Governance & Resources, which requires repositories to have “adequate funding and sufficient numbers of staff managed through a clear system of governance to effectively carry out the mission” (CoreTrustSeal Standards and Certification Board, 2022, p. 16). For the Resources part of this requirement, CoreTrustSeal asks for applicants to provide evidence for the following items (CoreTrustSeal Standards and Certification Board, 2022, p. 16): ● Timescales for provision and renewal of funding for operational costs and recruitment; it is understood that permanent, ongoing funding cannot be perfectly quantified or guaranteed. ● Evidence that the repository is, or is hosted by, a recognized institution (supporting long-term stability and sustainability) appropriate to its Designated Community. ● Demonstrate that the repository can meet its obligations, including sufficient funding, staff resources, IT resources, and a budget for external engagement when necessary. The nestor Seal requirement C8 Funding corresponds to this A|F (nestor Certification Working Group, 2025). In the OAIS Reference Model the high-level governance of the OAIS is carried out by Management, which in the model’s logic is not part of the OAIS, but an element of the OAIS environment. “Management is often the primary source of funding for an OAIS and therefore should agree the budget for the Archive’s activities and may provide guidelines for resource utilization (personnel, equipment, facilities)” (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 2/11). The day-to-day utilization of these resources within the OAIS is managed within the Administration functional entity: “The Establish Standards and Policies function [...] receives approvals and budget information and policies such as the OAIS charter, scope, resource utilization guidelines, and pricing policies from Management” (Consultative Committee for Space Data Systems (CCSDS), 2024, p. 4/13). The Desirable Characteristics of Data Repositories for Federally Funded Research policy (White House Office of Science and Technology Policy, 2022) addresses the financial and management aspects of resources primarily through the requirement for Long-term Organizational Sustainability, which mandates that the repository have a plan for long-term management of data built on a stable technical infrastructure and funding plans. The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that 98 | Page
are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term (European Commission, 2024). The Repository Features to Help Researchers: An invitation to a dialogue (Cannon et al., 2021) includes the following questions on funding: Who funds the operation (organisations) of the repository and the type of funding (e.g. grants, donations, memberships). Does the repository provide information on its sustainability plans? The Current State and Future Directions for Open Repositories in Europe (Shearer et al., 2023) states that increased staffing at these repositories could help to address many of the challenges being experienced and ensure there is widespread adoption of good practices and next generation repository functionalities. Shared infrastructure models, which have already been adopted in a few countries, are another approach that offer economies of scale and could relieve some of the burden from individual institutions. Several respondents provided more information about their sustainability challenges, grouped into several categories: ● Time and resource requirements to properly curate metadata and content ● Replacement of repositories with CRIS systems, which do not fully support the needs for managing a variety of content types ● Complexities of regular software upgrades ● High cost of employing outside companies to support software upgrades and ongoing maintenance of the system ● Lack of expected functionalities of the repository platforms ● Understaffing In a survey, respondents perceptions of repository sustainability where 97% of respondents (351) felt their repository was either “very” or “somewhat” sustainable, with only 13 respondents (3%) indicating that it was “not sustainable”. The 3% of respondents that felt their repository was unsustainable came from different countries and repository types, so no geographic generalisations could be inferred. The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states regarding maintenance with sustainable funding and adequately staffed. Treat repositories as critical research infrastructure and ensure dedicated institutional budget lines and sustainable funding models. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that resource availability (staffing, funding) was a major factor in certification readiness. Smaller repositories often struggled with the resource demands of certification. Diversity of organisations combined with the fact that certification, particularly for the first time, demands time from the repositories points to the difficulty of finding certification solutions that are suitable for all organisations. The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) states that repositories must demonstrate effective resource 99 | Page
management by maintaining a long-term plan for managing and funding the repository, while also designating at least one staff member with the explicit responsibility of managing the services to support sustainability. Domain-specific resources Agri-food The TTRAM Activity/Function AF16 Resources was not found to be addressed in the gathered Agri-Food resources. Biomedical Sciences The TTRAM Activity/Function AF16 Resources is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: The article partially covers the “Resources” function, focusing on governance and financial sustainability. However, it does not delve into operational resource management or energy use, which may be outside the scope of ELIXIR evaluation framework. Field et al., 2011: Funding bodies for GSC include NERC International Opportunities Fund and National Science Foundation (NSF) grants. Yilmaz et al., 2011: The standards themselves are centrally maintained in a relational database system at the Max Planck Institute for Marine Microbiology Bremen on behalf of the GSC, utilizing established processes like a public issue tracking system and annual releases to manage and deploy these valuable knowledge assets to the scientific community. Rehm et al., 2021: GA4GH is an expert-driven structure involving over 1000 individuals from more than 90 countries. It has twenty four Driver Projects committing at least two full-time equivalents (FTEs) to standards development across its eight Work Streams. Financial resources are secured from a wide array of international funding bodies and institutions. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that financial resources come through external or third-party funding and associated grants from data projects during the metadata submission phase in MetaXA. Linguistics CLARIN B and E centres are required “to make explicit statements to the CLARIN boards about its technological and funding support state and its perspectives in these respects” (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3). For CLARIN B centres, the following requirement applies (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): 100 | Page
● 2.d Continuity of access and funding support. Requirement: Each centre needs to make explicit statements about perspectives of continuity of access and funding support to continue activities as an active CLARIN centre. Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “Organisation: Staffing, Management and Finances” that are relevant to the topic of Resources (https://www.cessda.eu/Resource-Directory?tree=17,20). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.3.5. PEOPLE & EXPERTISE (AF17) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) People & Expertise is described as follows: “The management of human resources to fulfil the mission. Includes ensuring that sufficient skills are available, internally or externally (see AF20 “External Engagement”).” and suggests for transparent information “Skills Statement, Skills Development Plan, Staff Roles & Skills List” (L’Hours et al., 2025c, p. 21). CoreTrustSeal address the A|F People & Expertise in their requirement R06 Expertise & Guidance, which requires that “[t]he repository adopts mechanisms to secure ongoing expertise, guidance and feedback-either in-house, or external” and for which evidence for the following items needs to be provided (CoreTrustSeal Standards and Certification Board, 2022, pp. 16–17): ● That guidance and expertise reflects the scientific scope of the repository, if relevant. ● The repository aligns internal recruitment and external engagement with the services it offers. ● The repository ensures that its staff have access to ongoing training and professional development. ● The range and depth of expertise of both the organisation and its staff, including any relevant affiliations (e.g. national or international bodies), is appropriate to the mission. ● In-house advisers, or external advisory committees that include technical, curation, data science, data security, and disciplinary experts. ● How the repository communicates with experts for advice. The nestor Seal criterion “C9 Personnel” requires that “[s]ufficient numbers of appropriately qualified staff are available. Updated job descriptions exist which set out the required qualifications of the digital archive personnel and contain an organisational chart and/or a staff development plan based on the tasks and objectives of the digital archive” add citation. 101 | Page
The Desirable Characteristics of Data Repositories for Federally Funded Research policy (White House Office of Science and Technology Policy, 2022) states that the repository provides or facilitates expert curation and quality assurance to improve the accuracy and integrity of datasets and metadata. The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates that all of the preservation-specific and supporting research data management roles across the data lifecycle require sustained training based on a rich knowledge base of preservation information. Clear responsibilities must be in place for developing standards and guidance, for communication and for training. The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term (European Commission, 2024). The Update of the Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements states that the lack of a public policy for preservation, curation and security of the contents is the most frequent reason, followed by not adhering to a specific metadata standard (Lazzeri, 2024). The Current State and Future Directions for Open Repositories in Europe (Shearer et al., 2023) states that as the needs of the user community expand and evolve with open science becoming mainstream, there will be an increasing strain on repository staff. Low staffing levels are due to the fact that repositories have not been a high priority service for universities and that they are also competing with the commercial sector for skilled technical staff. The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states regarding trained professionals ensure repository managers and librarians receive ongoing training in metadata curation, research data management, and emerging AI applications. Establish professional development programmes and participate in international capacity-building initiatives. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that repositories express information about Skills Statements, Skills Development Plans, and Staff Roles & Skills Lists to demonstrate the availability of sufficient expertise. Furthermore, these guidelines note that specialist repositories, in particular, require specific skills in data formats, research methodologies, metadata, and ontologies related to their disciplines or content types to execute their functions effectively. specific skills in data formats, research methodologies, metadata, and ontologies related to their disciplines or content types to execute their functions effectively. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that repository certification is a key part of building a trusted FAIR data ecosystem. It needs to be applied in a way that acknowledges the differences in 102 | Page
goals, practices and maturity of data repositories and other service providers. Interoperability, standards, automation and technology are all parts of the solution, but reusability of data and long term preservation of understandability is ultimately dependent on domain and disciplinary expertise. The report highlights the importance of staff expertise in data curation, preservation, and certification. Peer support and training helped build capacity in less mature repositories. Ensuring the sustainable management of trust is not solely dependent on assessment or certification, as trust goes beyond the technical aspects of repositories and also involves people. Domain-specific resources Agri-food The TTRAM Activity/Function AF17 People & Expertise is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData highlights the importance of identifying and crediting repository personnel and ensuring teams include the expertise necessary to maintain and support databases. See AF01, Identification & Contact. Sen et al., 2020: WheatIS points to the need for skilled data curators and technical staff within each contributing repository to maintain consistent standards and interoperability; as well as the need for a good range of different expertises to keep the network running. “It needed technical expertise to build and maintain a strong computational infrastructure and create data formats to make data sets readable; scientific expertise to understand different types of wheat data sets (including genetic, genomic, phenotypic, and metabolic); outreach capability to help build relationships to add new nodes with new data sets; and leaders who not only motivate and manage personnel, but also work with the Wheat Initiative and the broader wheat community to promote and support WheatIS. The need for dedicated and competent personnel with complimentary and overlapping expertise was crucial. For WheatIS, or for any scientific community for that matter, the critical question is the type of the expertise needed and how much time the experts can devote to a fledgling community.” Biomedical Sciences The TTRAM Activity/Function AF17 People & Expertise is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: Effective management of human resources is essential for fulfilling the mission of an ELIXIR Core Data Resource. This includes ensuring that the resource has access to the necessary expertise and skills. Resources support ELIXIR CDRs’ operations, maintain scientific quality, and respond to evolving community needs. This also involves strategic engagement with user communities, helping to ensure that the resource remains relevant and scientifically robust. Sansone & Rocca-Serra, 2016: Ideally, standards should be implemented by dedicated experts in tools, services, and infrastructure to make them "invisible" to regular users 103 | Page
Field et al., 2011: GSC develops standards through community-based activities. It achieves this by forming working groups and actively encouraging the wider scientific community to join and contribute their expertise. Yilmaz et al., 2011: MIxS is maintained and authored by representatives from genome sequencing centers, maintainers of major resources, and principal investigators of both largeand small-scale sequencing projects. Climate Science The Activity/Function (A|F) People & Expertise is not addressed by any of the reviewed resources from Climate Science. Linguistics The Activity/Function (A|F) People & Expertise is not addressed by any of the reviewed resources from Linguistics. Social Sciences The Activity/Function (A|F) People & Expertise is not addressed by any of the reviewed resources from Social Sciences. 3.3.6. THIRD PARTY DEPENDENCIES (AF18) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Third Party Dependencies is described as follows: “Identifying and managing third-party organisations, hosts, partners, and other actors that are essential to operations, including the delivery of data or metadata services.” and suggests for transparent information “[l]ists of partners, Partnership Statement, Partner Management Plan, Contracts/suppliers mapped to the functions and activities they provide for the organisation (e.g. a storage provider).” The A|F Third Party Dependencies is explicitly dealt with in item “(6) Cooperation and outsourcing to third parties, partners and host organisations” in the R0. Background Information & Context section of CoreTrustSeal. In cases where the applicant does not have have direct control of a repository function and/or supporting evidence, CoreTrustSeal requires repositories to list the relevant host organisation, partner or other third party, “[d]escribe the function or service they provide, the nature of the relationship or agreement (contractual, Service Level Agreement, Memorandum of Understanding, etc.) and whether they have any relevant certifications” . The OAIS Reference Model discusses this topic in chapter 6 “Archive Interoperability”. It distinguishes independent, cooperating, and federated archives as well as different “styles of resource sharing” ranging from “all in-house” to “distributed” (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 6). 104 | Page
The Recommendations Consultation. EOSC-A Long Term Data Preservation Task Force (Andreu et al., 2023) formulates that roles and responsibilities including for complex partnerships, third party relationships and outsourcing should be understood and transparent. Technical repository service providers’ (storage providers, ARCHIVER22 etc) portfolio of service offerings should be clear and comparable for client end-users. Where responsibility is distributed, accountability should remain clear, including accountability for (meta)data loss or destruction. In Repositories and Beyond: Analysis of Survey for SSHOC Organisations (Ala-Lahti et al., 2022), respondents were asked who provides their technical infrastructure. The responses to this question were more varied, with two out of the three non-repository service providers and four of the repositories indicating that their technical infrastructure was provided by ‘the organisation/the repository itself’. The remaining non-repository service provider and one repository indicated that their technical infrastructure was provided by ‘a third party (outsourced)’. Of the remaining six repositories, three selected that their technical infrastructure was provided by ‘the host organisation’ and the other three repositories indicated that they ‘shared responsibility’ with regards to their technical infrastructure. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) note that repositories depend on a wider partnership of metadata and data services, such as storage providers, to high-performance computing, and multiple registries, all of which play a critical role in research data infrastructure. Consequently, the guidelines specify that Security must extend to the boundary between the repository and external entities, including these third-party dependencies. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states in chapter 4.2 TDR Partnership Models and Outsourcing that repositories must document third-party relationships and ensure quality and security in outsourced services. Complex partnerships models and outsourcing pose challenges for certification. While outsourcing is not an obstacle to certification as such, it requires repositories outsourcing their functions to explicitly state these functions and related documentation. The GREI Data citation best practices for repositories (Puebla et al., 2024) rely on Third Party Dependencies for the delivery of data citation services, emphasizing that repositories harvest citations from and submit data to essential external actors like DataCite and Crossref. Additionally, the initiative relies on various external sources and third-party aggregators—including Dimensions, Europe PMC, NASA ADS, and the Chan Zuckerberg Initiative—to collect citation data and contribute to centralized resources like the Data Citation Corpus. Domain-specific resources Agri-food The TTRAM Activity/Function AF18 Third Party Dependencies was not found to be addressed in the gathered Agri-Food resources. 105 | Page
expressed in the following items that applicants are asked to address (CoreTrustSeal Standards and Certification Board, 2022, p. 17): ● The repository aligns internal recruitment and external engagement with the services it offers. ● The range and depth of expertise of both the organisation and its staff, including any relevant affiliations (e.g. national or international bodies), is appropriate to the mission. ● In-house advisers, or external advisory committees that include technical, curation, data science, data security, and disciplinary experts. How the repository communicates with experts for advice. Other aspects of the CoreTrustSeal requirement R06 Expertise & Guidance relate to the A|F People & Expertise (AF17).) The Desirable Characteristics of Data Repositories for Federally Funded Research policy (White House Office of Science and Technology Policy, 2022) emphasizes that repositories must facilitate Broad and Measured Reuse by monitoring the user community and having mechanisms to measure attribution, citation, and reuse of data, which requires ongoing engagement with external users and stakeholders. Additionally, the policy's overall purpose is to provide a consistent set of characteristics for repositories to inform external data repository developers and managers and improve coordination across agencies to enhance compliance and open science infrastructure. The Recommendations for Services in a FAIR data ecosystem presents four recommendations that stand out as being assigned at least medium priority by all, and top priority by two different groups. This is one: Foster global collaboration on FAIR implementation challenges and emerging solutions through organisations such as the Research Data Alliance (Bangert et al., 2019). The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states regarding connection to national and international repository networks that actively participating in repository networks and global initiatives enhance visibility, sharing of best practices, and align with evolving open science policies. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) emphasize that repositories depend on a wider partnership of metadata and data services, such as storage providers and multiple registries, which play a critical role in research data infrastructure. Consequently, the guidelines recommend that repositories must transparently share information related to their trustworthiness and services to foster mutual trust between human actors (e.g., researchers or funders) and machine agents, which is a critical precursor to trusted relationships with these external parties. The GREI Data citation best practices for repositories (Puebla et al., 2024) recommend repositories submit data citations to DataCite and harvest citations from third-party aggregators such as Crossref, Dimensions, Europe PMC, and NASA ADS repositories submit data citations to DataCite and harvest citations from third-party aggregators such as Crossref, Dimensions, Europe PMC, and NASA ADS, while also aligning their recommendations with external community work like the FORCE11 Data 112 | Page
Citation Principles. Furthermore, GREI encourages ongoing engagement by inviting repositories and interested parties to contribute data citations to the Data Citation Corpus and provide feedback during its development. The Core Preservation Process CPP-018 Community Watch sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA monitors its Designated Community in order to identify its evolving needs and knowledge.” (EOSC EDEN T1.2 et al., 2025). Domain-specific resources Agri-food The TTRAM Activity/Function AF20 External Engagement is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData explicitly recommends that repositories establish active communication with their user and depositor communities to strengthen engagement and accountability. “A strong collaborative community of database developers, curators, and users is critical to the success and sustainability of biological databases.” Drakos et al., 2015: agINFRA directs repositories to maintain continuous collaboration with partner institutions, funders, and data providers to ensure services meet the needs of external stakeholders. “Ensure stakeholder needs are met with regards to data management and sharing; involve as many agricultural data sources as possible to provide maximum value.” Sen et al., 2020: WheatIS demonstrates repository-level engagement by coordinating global partners and researchers to maintain interoperable data services and standards. “The Wheat Initiative tasked the WheatIS EWG’s to provide the international wheat research community with easy access to wheat genetics, phenotype with environmental information, genomic data and bioinformatics tools, and to support and promote the diverse wheat databases internationally.” Marrano et al., 2025: The authors underline that repositories must develop community engagement plans and communicate with funders to secure sustainable partnerships. “Engagement with funders, community stakeholders, and policy makers is essential for repositories to align objectives and ensure the long-term sustainability of services.” Šestak & Copot, 2023: External engagement is linked to repository resilience and policy impact, emphasizing that communication with policymakers and institutions strengthens sustainability. “A sustainable agri-data ecosystem depends on coordinated action among repositories, research institutions, and policymakers to ensure shared understanding and effective use of open data resources.” Ali & Dahlhaus, 2022: The paper identifies engagement with research networks and environmental agencies as necessary for repositories to maintain relevant and interoperable data services. “The 113 | Page
reliability of agricultural data infrastructures depends on strong engagement with research networks, water authorities, and other agencies that generate or maintain data.” Biomedical Sciences The TTRAM Activity/Function AF20 External Engagement is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: The article includes community-related indicators, such as size and diversity of the user base, engagement with scientific communities, and responsiveness to user needs. ELIXIR CDRs are often part of international consortia or collaborative networks, and the article highlights the importance of these relationships in sustaining and evolving the resource. The article discusses the role of host institutions and funding bodies, which are external entities essential to the resource’s sustainability. The presence of an independent, international Scientific Advisory Board is another form of external engagement, providing strategic input from outside experts. Field et al., 2011: GSC acts as an international body for external engagement with major data repositories like the International Nucleotide Sequence Database Collaboration (INSDC), as well as collaborating extensively with numerous partner organizations and the wider scientific community through community-led surveys and open calls for contributions to develop, maintain, and ensure the widespread adoption of its MIxS standard. Barrett et al., 2012: The article highlights NCBI’s role within the International Nucleotide Sequence Database Collaboration (INSDC), working alongside DDBJ (Japan) and EBI (Europe). This reflects sustained engagement with global partner organizations. NCBI collaborates with ATCC and Coriell to create Reference BioSample records for widely used biological materials, facilitating standardized reuse across the research community. The article references the Genomics Standards Consortium, whose MiXS checklists are integrated into BioSample workflows, demonstrating alignment with community-driven metadata standards. Karsch-Mizrachi et al., 2025: The article describes how INSDC interacts with a broad range of external parties, including current and prospective member organizations, standards bodies, and the wider scientific community. To reflect this commitment, INSDC Executive Committee will establish an ad hoc International Stakeholders Committee (ISC) which will include a diverse group of external experts who are expected to promote the principles and activities of INSDC and provide input on different stakeholders’ needs. Rehm et al., 2021: GA4GH collaborates with external standards development organizations like Health Level Seven (HL7), ISO, and Open Biological and Biomedical Ontology Foundry (OBO). GA4GH also engages with diverse data-hosting environments, to ensure global interoperability and uptake of its secure data sharing solutions. Climate Science 114 | Page
The METAFOR project: preserving data through metadata standards for climate models and simulations (Callaghan et al., 2010) advocates Community consultation and Intergovernmental Panel on Climate Change (IPCC) involvement. The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) describes the engagement in extensive External Engagement by conducting a wide consultation process with more than 35 climate modeling experts from 13 research centers representing 6 countries to define and establish scientific consensus on the Controlled Vocabulary (CV) content and granularity. This effort, which was supported by the EU 7th Framework Programme and involved collaboration with the US Earth System Curator project, included planning to establish an international governance committee under IS-ENES2 to manage the CV's future evolution and preservation. The NetCDF Climate and Forecast (CF) Metadata Conventions (Eaton et al., 2024) state that the CF framework promotes the processing and sharing of climate and forecast data across diverse sources and applications, and is designed to be backward compatible with other conventions like COARDS while also explicitly incorporating standards such as UGRID to enhance interoperability. Linguistics The Activity/Function (A|F) External Engagement is not addressed by any of the reviewed resources from Linguistics. Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “User Support and Communication” that are relevant to the topic of External Engagement (https://www.cessda.eu/Resource-Directory?tree=11). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.3.9. RELEASE & PUBLISHING (AF21) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Release & Publishing is described as follows: “The processes by which the organisation creates, manages, approves, releases, updates and withdraws information artefacts (not deposited digital objects) such as publications, policies, guidelines, and other content related to the repository activities and functions. It includes publishing information that provides evidence for external assessments (see AF24 "Criteria, Assessment, Improvement") and supports external engagement (see AF20 "External Engagement").” and suggests for transparent information “[r]ecords management policy, records retention schedule, publication approval process” (L’Hours et al., 2025c, p. 22). 115 | Page
The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) generally support the release and publishing activity by recommending that information exposed by repositories should include both internal information artefacts (like policies and procedures) and external information artefacts (like guidance and criteria), which are typically released through publishing processes. Moreover, the guidelines emphasize that repositories should expose information related to trustworthiness and certification status, which necessitates the organizational release of evidence for external assessment. Domain-specific resources Agri-food The TTRAM Activity/Function AF21 Release & Publishing is addressed in the following Agri-Food resources: Caracciolo et al., 2020: The Agrisemantics group directs repositories to publish semantic artefacts, metadata guidelines, and controlled vocabularies to promote transparency and reuse. Marrano et al., 2025: Marrano et al. highlight the importance of repositories releasing and maintaining formal records and documentation to evidence compliance and improvement processes. “Repositories are expected to publish and regularly update their internal policies, governance documents, and operational guidelines to provide transparency and facilitate external assessment.” Šestak & Copot, 2023: Sestak et al. connect release and publishing to open-science communication, advocating for repositories to publicly share reports, assessments, and guidance to ensure credibility and traceability. “Repositories contribute to a sustainable open-data environment by publishing documentation and assessment results that support transparency and accountability across the agri-data ecosystem.” Ali & Dahlhaus, 2022: The authors refer to the role of repositories in releasing documentation and technical specifications to support data integration within hydrological and agricultural systems. “Documentation and publication of technical standards and access mechanisms are essential for maintaining transparency and reproducibility across environmental and agricultural data infrastructures.” Biomedical Sciences The TTRAM Activity/Function AF21 Release & Publishing is addressed in the following Biomedical Sciences resources: Field et al., 2011: GSC release and publish in the Standards in Genomic Sciences journal. This journal functions as a formal voice for the GSC, supporting the publication of standardized genome, metagenome, and pan-genome reports, and other standards-supportive publications such as Standard Operating Procedures (SOPs) 116 | Page
Karsch-Mizrachi et al., 2025: The article discusses the publication of formal documents such as the Founders Arrangement, the Membership Arrangement, and the Membership Acceptance and Performance Guidelines, which are made available through the INSDC website. These documents provide transparency into the collaboration’s governance, policies, and operational expectations, supporting external engagement and enabling external assessment of the repository’s practices. The article also mentions updates to the INSDC website, including new sections that describe the mission, governance model, and plans for global participation, further demonstrating the repository’s commitment to releasing information artefacts that support its strategic and operational goals. Rehm et al., 2021: GA4GH releases its products through a multi-stage approval process involving its Work Streams, specialized committees, and the Steering Committee, with development work often conducted publicly on platforms like GitHub. Climate Science The Activity/Function (A|F) Release & Publishing is not addressed by any of the reviewed resources from Climate Science. Linguistics The Activity/Function (A|F) Release & Publishing is not addressed by any of the reviewed resources from Linguistics. Social Sciences The Activity/Function (A|F) Release & Publishing is not addressed by any of the reviewed resources from Social Sciences. 3.3.10. INTEROPERABILITY (AF22) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Interoperability is described as follows: “The protocols and processes by which people, processes, technologies and digital objects effectively interact with others, both within and across organisational entities.” and suggests for transparent information “Interoperability policy, plan, procedures. Links to and documentation of any machine accessible routes to data and/or metadata, e.g. APIs, OAI-PMH, aligned with Technical Infrastructure (see AF29 “Technical Infrastructure”). Crosswalks between relevant metadata standards and ontologies.” (L’Hours et al., 2025c, p. 23) The OAIS Reference Model discusses issues around interoperability in Chapter 6 “Archive Interoperability” (Consultative Committee for Space Data Systems (CCSDS), 2024). PAIS provides “a standard method for formally defining the digital information objects to be transferred by an information Producer to an Archive and for effectively packaging these objects in 117 | Page
the form of Submission Information Packages (SIPs)” and an XML schema for this purpose (Consultative Committee of Space Data Systems (CCSDS), 2014, Chapter 1, p. 1). PAIMAS defines the steps and phases by which producers and archives interact. It has the purpose “to identify, define and provide structure to the relationships and interactions between an information Producer and an Archive. This Recommendation defines the methodology for the structure of actions that are required from the initial time of contact between the Producer and the Archive until the objects of information are received and validated by the Archive” (Consultative Committee for Space Data Systems (CCSDS), 2004, Chapter 1, p. 2) . 1-2 The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states regarding interoperability with scholarly communication and research systems enable seamless data exchange with institutional research information systems (CRIS), funder databases, and scholarly publishing platforms. The D1.3 Recommendations for a FAIR EOSC - White Paper of the FAIR-IMPACT Synchronisation Force (Grootveld, 2025) emphasizes that enhancing these legal and organizational aspects requires scaling up knowledge via a central support program offering training and advice, particularly for Research Infrastructures that often face challenges due to fragmented setups. Furthermore, to advance legal and organizational interoperability, continued collaboration across Horizon Europe projects and the active advocacy of Creative Commons licenses as the default option are key recommendations The O'FAIRe makes you an offer: Metadata-based Automatic FAIRness Assessment for Ontologies and Semantic Resources (Amdouni et al., 2022) assessed by the O'FAIRe tool, the average normalized score for Interoperability (I) was notably low at 42, highlighting that the implementation of standardized protocols and processes for interaction is challenging. Interoperability is a strong focus of the FAIRsFAIR project D2.3 report (Behnke et al., 2020), covered by the mention of permanent identifiers as the manifestation of a data policy, the support to standard formats, the intensive use of metadata at the level of the repository itself but also for the digital objects, the granularity being under the responsibility of the repository. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that transparent information exposed by repositories should take account of, and map to, existing standards and criteria like DCAT and schema.org, and emphasize information exposed by repositories should take account of, and map to, existing standards and criteria like DCAT and schema.org, and emphasize that using DCAT (Data Catalog Vocabulary) is key for fostering interoperability between data catalogues on the web and maximizing machine-actionability for harvesters. The FAIR Data Maturity Model. Specification and Guidelines (FAIR Data Maturity Model Working Group, 2020) directly address Interoperability (I), defining several indicators as Important that relate to these organizational processes, such as requiring metadata (RDA-I1-01M) and data (RDA-I1-01D) to use knowledge representation expressed in standardized formats and for metadata (RDA-I2-01M) 118 | Page
to use FAIR-compliant vocabularies. Crucially, while the Interoperable area has 12 indicators, the FDMM notes that Essential indicators are completely absent from this FAIR area, which means a high level of FAIRness could be achieved even in the absence of capacity in these interoperability aspects. The GREI Data citation best practices for repositories (Puebla et al., 2024) mandate the standardized storage of citation relationships using DataCite metadata fields (e.g., relatedIdentifier and relationType). Furthermore, the protocols require repositories to actively interact with other entities by submitting collected data citations to DataCite and harvesting information from external sources such as Crossref, Dimensions, and Europe PMC. In the D4.5 Report on Completed FAIR Data Standard Adoption and Certifications of Data Repositories in the Region (Alaterä et al., 2022), it is recommended that data be structured using a formal language suitable for knowledge representation, along with a community-defined ontology or data template. Additionally, it is considered best practice for metadata to use terms that resolve to linked FAIR data, and for the metadata identifier itself to resolve to and utilize FAIR vocabularies, while also linking to third-party resources. Overall, datasets should be potentially representable as Linked Data and include references to other relevant metadata. Domain-specific resources Agri-food The TTRAM Activity/Function AF22 Interoperability is addressed in the following Agri-Food resources: Drakos et al., 2015: agINFRA defines interoperability as a core mission for agricultural repositories and provides explicit guidance for implementation, through the “interconnect [ion of] agricultural data repositories through extended metadata [and the] advanced implementation and adoption of European standards and specifications.” Sen et al., 2020: WheatIS demonstrates operational interoperability through federated metadata harvesting and standardized data formats. Caracciolo et al., 2020: The Agrisemantics group insists on semantic interoperability via shared vocabularies and metadata crosswalks. Harper et al., 2018: AgBioData identifies interoperability as one of its fundamental repository recommendations, promoting adoption of shared formats and APIs. Marrano et al., 2025: Marrano et al. stress that repositories must document and publish interoperability procedures and crosswalks between metadata schemas. “Repositories should maintain transparent documentation of interoperability protocols and crosswalks among metadata standards and ontologies to ensure consistency across federated infrastructures.” Šestak & Copot, 2023: The authors frame interoperability as a sustainability enabler, recommending coordinated data-exchange standards across Agri-food infrastructures. “A sustainable agri-data 119 | Page
ecosystem requires interoperable repositories built on open protocols and harmonized metadata frameworks to enable seamless information flow between research domains.” Ali & Dahlhaus, 2022: Interoperability is described as a key repository responsibility for connecting agricultural and hydrological data services. “Interoperability between repositories is essential for the integration of environmental and agricultural datasets; standardized metadata and API access ensure discoverability and reuse.” Biomedical Sciences The TTRAM Activity/Function AF22 Interoperability is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: ELIXIR Core Data Resources (CDRs) are evaluated on their use of community-agreed standards for data formats, metadata schemas, and ontologies. The CDRs are expected to provide programmatic access (e.g., APIs), enabling interaction with other tools, platforms, and services. The article discusses how ELIXIR CDRs often serve as foundational resources for other databases and services. This reach-through dependency highlights the importance of interoperability across organisational entities. Lastly, Interoperability is one of the four FAIR principles, which are central to the evaluation framework. CDRs must ensure that both data and metadata are machine-readable and semantically linked, enabling effective interaction. Lin et al., 2024: In the article, the authors emphasise the importance for project repositories to apply extra efforts to ensure data adherence to field-specific standards for increased interoperability and reusability. Besides, FAIR Principles are considered as one of the desirable characteristics for all repositories to serve as guidelines to enhance data’s discoverability, interoperability, and reusability. Sansone & Rocca-Serra, 2016: Describes a need to overcome the current fragmented and overlapping efforts due to a lack of central authority and coordination across different organizational types and domains in life science. Yilmaz et al., 2011: MIxS is compliant with many submission tools e.g., GenBank, EBI-ENA, SRA tools, MetaBar, QIIME, ISA, and designed to enable the seamless interaction and integration of sequence data and contextual metadata across different organizational entities and platforms. Barrett et al., 2012: BioProject and BioSample records are reciprocally linked to experimental data stored in multiple archival databases, including GenBank, SRA, GEO, and dbGaP, enabling seamless navigation and integration across NCBI resources. These databases are part of the International Nucleotide Sequence Database Collaboration (INSDC), which includes DDBJ (Japan) and EBI (Europe), with data exchanged regularly among partners—demonstrating robust interoperability across organizational boundaries. The BioSample database supports semantic interoperability through the use of controlled vocabularies and standardized checklists, such as the MiXS standards developed by the Genomics Standards Consortium. Additionally, collaboration with external providers like ATCC and Coriell to create standardized sample records further enhances interoperability by enabling consistent referencing and integration of datasets across institutions and platforms. 120 | Page
Karsch-Mizrachi et al., 2025: The article describes how INSDC operates as a coordinated collaboration among three major international organizations—NCBI, DDBJ, and EMBL-EBI—who exchange data daily to ensure each site maintains a complete and synchronized dataset. The article also discusses the alignment of INSDC standards with external bodies such as the Genomic Standards Consortium, PHA4GE, and GA4GH, which facilitates interoperability across different data systems and communities. Furthermore, the development of minimal standards and metadata checklists supports consistent interaction between data submitters, repository systems, and users, enabling effective integration and reuse of data across organizational boundaries. Wilkinson et al., 2016: The article addresses interoperability for scholarly digital objects (data, algorithms, tools, and workflows) and states they must utilize formal, shared, and broadly applicable languages and vocabularies, along with qualified references to other (meta)data, to enable effective interaction across diverse systems and entities. Rehm et al., 2021: The activity and function is not explicitly covered in the article. However, technical standards and policy frameworks, such as APIs and data formats are developed by GA4GH. These advocate seamless and responsible interaction among people, processes, technologies, and digital objects across diverse institutional and national boundaries, and facilitate federated data analysis and exchange. Climate Science The METAFOR project: preserving data through metadata standards for climate models and simulations (Callaghan et al., 2010) states CIM is designed for interoperability. The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) describes the established interoperability protocols and processes by adopting the CF convention for data formatting and developing a metadata information pipeline to bridge different systems. This pipeline automatically converts the human-readable Controlled Vocabulary (CV) into machine-readable formats like XML and OWL ontologies, enabling tools like the CMIP5 Questionnaire to interact effectively with distributed organizational entities such as the Earth System Grid Federation (ESGF) gateways via broadcast protocols like "atom feeds". Linguistics The Activity/Function (A|F) Interoperability is not addressed by any of the reviewed resources from Linguistics. Social Sciences It is recommended that repositories include metadata files or tags using a generic schema, as this improves interoperability and increases the likelihood that non-specialised infrastructures can reuse key citation elements and other metadata. (Bornatici et al., 2025, p. 8). 3.3.11. LEGAL & ETHICAL (AF23) 121 | Page
the scholarly ecosystem by focusing on the development of the Data Citation Corpus, a centralized resource intended to compile and make data citation information readily available and openly accessible to the community. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories report (Kleemola et al., 2022) states that CoreTrustSeal certification is used to assess and improve repository practices. Repositories benefited from self-assessment and peer support to enhance documentation, metadata quality, and compliance with standards. Within AF24 Criteria, Assessment, Improvement, the Data Curation Network defines the following Curation Activity (Johnston et al., 2016): “Repository Certification: The technical and administrative capacities of the repository undergo review through a transparent and well-documented process by a trusted third-party accreditation body (e.g., TRAC, or Data Seal of Approval).” Domain-specific resources Agri-food The TTRAM Activity/Function AF24 Criteria, Assessment, Improvement is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData calls for repositories to evaluate and refine their data management practices through community standards and feedback cycles. “AgBioData recommends that databases assess and update their practices regularly to ensure data quality, interoperability, and long-term usability.” Caracciolo et al., 2020: The Agrisemantics group encourages repositories to iteratively assess semantic quality and adopt improvement mechanisms aligned with FAIR. “We recommend systematic monitoring and evaluation of semantic artefacts and their use in Agri-food data repositories to ensure ongoing FAIR compliance.” Marrano et al., 2025: Marrano et al. emphasize repositories’ responsibility to document and publish assessment results and certification outcomes to demonstrate compliance and progress. “Repositories are expected to undertake regular self-assessment against international standards, document improvement actions, and publish results to support transparency.” Šestak & Copot, 2023: The authors link continuous improvement with sustainability, recommending that repositories monitor performance indicators and evolve with user and policy demands. “A sustainable agri-data ecosystem relies on regular evaluation of repository performance, ensuring adaptation to policy changes, user needs, and technological advances.” Ali & Dahlhaus, 2022: The paper points to assessment processes as necessary for repository credibility, particularly in multi-institutional data infrastructures. “Continuous assessment of data quality and repository practices is essential to maintain trust and usability across integrated environmental and agricultural data systems.” 128 | Page
Biomedical Sciences The TTRAM Activity/Function AF24 Criteria, Assessment, Improvement is addressed in the following Biomedical Sciences resources: Sansone & Rocca-Serra, 2016: Tracking the lifecycle of content standards is considered challenging, but initiatives like BioSharing use four specific indicators (Ready, Under Development, Uncertain, Obsolete) to assess a standard's readiness for implementation or use. Field et al., 2011: GSC fulfills its mission by fostering widespread adoption and collaboration across the scientific community. Yilmaz et al., 2011: MIxS standard is defined through GSC for core contextual data requirements and the use of specific ontologies for consistent reporting. Compliance is automatically assessed and validated by major databases like GenBank, EBI-ENA, and SRA via provided tools and templates. Karsch-Mizrachi et al., 2025: The article explains that INSDC has developed a Maturity Model that uses a series of objective criteria to increase levels of maturity in all of the categories in each of the five areas: ● Governance and Institutional Context ● Technical Infrastructure ● Data Operations ● Communications and Engagement ● Quality Rehm et al., 2021: GA4GH processes are piloted by its twenty four Driver Projects providing feedback. Interoperability testing is carried out through initiatives like FASP that identify areas for specification updates. Climate Science The Development and exploitation of a controlled vocabulary in support of climate modelling (Moine et al., 2014) describes established rigorous processes for assessing compliance by incorporating a mindmap validator and Schematron-based validation within the metadata information pipeline, ensuring that the collected climate model descriptions adhered to predefined encoding rules, CIM syntax, and parameter coherency. For continuous improvement and targeted enhancement of the metadata standards, the project planned to establish an international governance committee to manage the evolution and preservation of the controlled vocabulary, recognizing the necessity of reinvesting lessons learned from the complexity of the initial harvesting procedure into future projects. The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC's organizational infrastructures multi-step submission process includes Technical Quality Assurance by data management to verify metadata consistency and data integrity against defined criteria like CF Conventions. 129 | Page
Linguistics CLARIN B and E centres are required “to [...] participate in a quality assessment procedure as proposed by the CoreTrustSeal or the nestor Seal” (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3). CLARIN B centres “cannot be certified [...] until the CoreTrustSeal assessment has been successfully concluded (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023). Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “Organisation: Certification” that are relevant to the topic of Criteria, Assessment and Improvement (https://www.cessda.eu/Resource-Directory?tree=17,21). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.3.13. ANALYSIS & IMPACT (AF25) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Analysis & Impact is described as follows: “The analysis of internal data and external information to verify and demonstrate that the organisation is effectively fulfilling its mission (see AF02 "Mission & Scope").” and suggests for transparent information “[b]usiness information analyses, identification and validation of external impact. Dashboards or interfaces for metrics and usage information.” (L’Hours et al., 2025c, p. 24) In the OAIS Reference Model functions relevant to this A|F are part of the Administration functional entity, which issues report requests to other OAIS functional entities and receives reports from these in return (Consultative Committee for Space Data Systems (CCSDS), 2024, Chapter 4.2.3.6). The measurement of repository impact and digital object is also mentioned in the FAIRsFAIR project D2.3 report (Behnke et al., 2020), with the suggestion to use a publication tracker for associated datasets and allow citation of reuse of partial data or single elements of datasets. The GREI Data citation best practices for repositories (Puebla et al., 2024) center on the mandate to implement data metrics that enable reporting on the reach and impact of NIH-funded research data. Data citations are identified as a crucial component of these measures, signaling dataset usage and providing valuable evidence for research evaluation frameworks. Within AF25 Analysis & Impact, the Data Curation Network defines the following Curation Activity (Johnston et al., 2016): “Use Analytics: Monitor and record how often data are viewed, requested, and/or downloaded. Track and report reuse metrics, such as data citations and impact measures for the data over time.” 130 | Page
Domain-specific resources Agri-food The TTRAM Activity/Function AF25 Analysis & Impact is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData urges repositories to demonstrate their impact through documented community uptake and usability of their databases. “Metrics on data use, user engagement, and citation should be collected to demonstrate the value and impact of databases within the broader agricultural genomics community.” Drakos et al., 2015: agINFRA calls for monitoring repository outcomes to verify how interoperability and shared services contribute to community impact. Caracciolo et al., 2020: The Agrisemantics group links impact assessment with FAIR uptake, recommending that repositories measure how semantic resources improve findability and reuse. “We recommend systematic evaluation of the impact of semantic artefacts and FAIR implementation to ensure repositories contribute effectively to Agri-food data integration.” Marrano et al., 2025: Marrano et al. emphasize that repositories should analyze internal performance data and publish dashboards to demonstrate mission fulfillment. “Repositories are encouraged to monitor and publicly report indicators of usage, data quality, and service efficiency through dashboards or open reporting tools.” Ali & Dahlhaus, 2022: Repositories are positioned as measurable components in hydrological and agricultural data infrastructures, where impact is reflected in accessibility and integration. “The performance of agricultural and hydrological data repositories can be evaluated through their interoperability, accessibility, and contribution to evidence-based environmental management.” Biomedical Sciences The TTRAM Activity/Function AF25 Analysis & Impact is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: Understanding the impact of ELIXIR Core Data Resources involves examining their unique contributions to the scientific ecosystem. A useful approach is the counterfactual perspective—considering what would happen if the resource had never existed or were to disappear without replacement. Many ELIXIR CDRs are globally unique, and their absence would significantly disrupt dependent resources and scientific workflows. These resources also play a vital role in accelerating science, by setting standards, promoting data and software reuse, enhancing research efficiency, and enabling the extension of technical solutions across disciplines. To communicate their value effectively, resources may use translational figures—familiar metrics or examples that resonate with stakeholders and help illustrate the resource’s core function and broader relevance. 131 | Page
Lin et al., 2024: The article outlines a list of commonly collected repository metrics. According to the authors, these metrics are essential for monitoring the scientific impact of data usage, thereby supporting the continued operation and success of a repository. Moreover, they offer systematic parameters for evaluating the costs and benefits—essentially the return on investment—for various stakeholders, including managers, research institutions, funding agencies, and research communities. While data metrics are a key component of repository metrics, the two serve distinct purposes. Repository metrics are aggregate indicators that reflect the overall access, usage, and impact of the repository’s services across all hosted data. They provide a holistic view of the repository’s value and influence. In contrast, data metrics focus on individual datasets, offering granular insights into their reuse, value, and alignment with FAIR Principles over time. Yilmaz et al., 2011: MIxS is based on analysis of community needs and existing data gaps identified through surveys. The widespread adoption of these standards by major data providers and the INSDC serves as a demonstration of their success. Karsch-Mizrachi et al., 2025: The article describes how INSDB is fulfilling its mission by documenting its growth in data volume, its alignment with international standards, and its efforts to expand membership and representation. Rehm et al., 2021: GA4GH's Federated Analysis System Project tests implementations against real-world scenarios, with learnings feeding back to refine specifications and ensure interoperability and solutions. Community engagement through the Genomics in Health Implementation Forum (GHIF) helps verify that standards meet actual data sharing needs. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states that WDCC ensures it effectively fulfills its mission of long-term data accessibility and re-usability through Technical Quality Assurance (TQA). Linguistics The Activity/Function (A|F) Analysis & Impact is not addressed by any of the reviewed resources from Linguistics. Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “Organisation: Monitoring” that are relevant to the topic of Analysis & Impact (https://www.cessda.eu/Resource-Directory?tree=17,18). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.3.14. TRAINING (AF26) 132 | Page
Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Training is described as follows: “Leveraging internal expertise to train data and metadata producers, owners, depositors, and users. Training can cover the entire digital object management lifecycle, from the conception of research to the reuse of data and metadata. It also includes training for internal staff, as well as for peer organisations, partners, and third parties.” and suggests for transparent information “[l]ink to primary location where the organisation provides training materials and/or details of training programmes. (L’Hours et al., 2025c, p. 24)” The A|F Training is partially covered by the CoreTrustSeal requirement R06 Expertise & Guidance (“The repository adopts mechanisms to secure ongoing expertise, guidance and feedback - either in-house, or external.”), which among other things asks the repository to provide evidence that it “ensures that its staff have access to ongoing training and professional development” (CoreTrustSeal Standards and Certification Board, 2022, pp. 16–17). The EOSC Federation Handbook addresses the A|F Training in its chapter on research training (5.2.6), which stresses that “[t]raining material for research services and other services is essential for scientists to use them properly” and lists the following inclusion criteria for training resources to be registered in the EOSC Federation (EOSC Association, 2025, p. 36): ● Specify the learning outcomes, resource type (e.g. recorded lesson, textbook, activity plan, etc.), content resource type (e.g. video, slides, audio, etc.), and estimated duration (e.g. estimated work hours). ● Be in at least one of the European languages except from metadata information, which shall be available in English. ● Incorporate information about the expected level of training and expertise to be achieved (beginner, intermediate, advanced, all) and required qualifications to access the training resource. The EOSC Federation Handbook also encourages providers to use the Quality Assurance Certification Framework produced by Skills4EOSC. The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) recommend that guidance should be incorporated directly into Data Management Planning (DMP) tools. This guidance helps users make informed choices about data deposit concerning access, storage, curation, and preservation from the earliest stage of their research. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) explains that training was provided through webinars, documentation, and one-on-one support. Repositories gained skills in self-assessment and certification preparation. 133 | Page
Domain-specific resources Agri-food The TTRAM Activity/Function AF26 Training is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData directly calls for repositories to provide training and documentation to curators and users. “Training users and curators to adopt metadata standards and best practices is essential to ensure data quality and consistency across AgBioData repositories.” Drakos et al., 2015: agINFRA promotes training as part of repository collaboration, ensuring users and partner institutions can effectively use shared services. “Provide tools, training and support for collaboration between European institutions in data-intensive research, validating the approach by enabling users to interact with each other and the data.” Caracciolo et al., 2020: The Agrisemantics group links training with the adoption of semantic technologies, encouraging repositories to educate staff and users in metadata and ontology management. “Promote training in semantics and metadata to ensure repositories can effectively implement and maintain FAIR-compliant systems.” Marrano et al., 2025: Marrano et al. stress that repositories should document and publish training programmes and materials for transparency and replication. “Repositories should provide openly accessible training materials and descriptions of their training programmes for curators, depositors, and users.” Šestak & Copot, 2023: The authors connect training to repository sustainability, suggesting that education strengthens capacity for open data stewardship. “Sustainable agri-data infrastructures depend on continuous training and knowledge sharing among repository staff and the broader research community.” Ali & Dahlhaus, 2022: The paper highlights training in data management and interoperability as key to maintaining repository quality and coordination across environmental and agricultural data systems. “Training and documentation are necessary to ensure consistent data management practices and interoperability across distributed repositories.” Biomedical Sciences The TTRAM Activity/Function AF26 Training is addressed in the following Biomedical Sciences resources: Durinx et al., 2017: Under the Quality of Service indicators, the article asks whether the resource undertakes training, alongside helpdesk support and user feedback mechanisms. The article emphasizes engagement with the scientific community, which may include training as part of outreach and capacity building. However, there is no reference to lifecycle-based training (e.g., from 134 | Page
data creation to reuse), nor mention to internal staff training, or training for third-party partners or metadata producers. Sansone & Rocca-Serra, 2016: Need to foster collaboration beyond the pharmaceutical and biotech industries, including others key stakeholders such as publishers, librarians. Need for education, documentation, hackathons, training and courses materials (and events) targeting both producers and consumers of standards, and set to create a new career path. Climate Science The Activity/Function (A|F) Training is not addressed by any of the reviewed resources from Climate Science. Linguistics The Activity/Function (A|F) Training is not addressed by any of the reviewed resources from Linguistics. Social Sciences CESSDA Training Resources provides a wide range of training material focused on key aspects of research data, including data discovery, data management, data analysis, and data preservation. The platform offers content in various formats such as webinars, presentations, slides, and videos. (CESSDA Training Team, 2025b). The final CESSDA recommendation (8.) for data repositories regarding data citation encourages raising awareness and enhancing user guidance about data citation for data users and data producers. This can be done through, for example, general information, dataset documentation, teaching materials, and seminars. (Bornatici et al., 2025, pp. 8–9). 3.3.15. RESEARCH & DEVELOPMENT (R&D) (AF27) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Research & Development (R&D) is described as follows: “Projects and other activities outside current operations, routine maintenance and upgrades. This can include new approaches to data, metadata, business processes, technology, security, and research infrastructure.” and suggests for transparent information “[d]ocumentation of project, specification, development and delivery processes. Links to information about projects the organisation is involved with. How products move from ‘in development’ to ‘in production’.” (L’Hours et al., 2025c, p. 25) The GREI Data citation best practices for repositories (Puebla et al., 2024) focus on establishing standardized data citation practices and implementing data metrics to report on the reach and impact of NIH-funded research data across the ecosystem. This includes supporting the development of the Data Citation Corpus, which is a collaborative project by DataCite to create a centralized 135 | Page
resource that compiles data citations from various sources, addressing existing challenges in consistency and accessibility. Domain-specific resources Agri-food The TTRAM Activity/Function AF27 Research & Development (R&D) is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData positions repositories as innovation platforms, encouraging active participation in projects that develop new standards and tools for agricultural genomics. “AgBioData encourages collaboration and innovation among database developers to pilot new technologies and standards that improve interoperability and data reuse.” Caracciolo et al., 2020: The Agrisemantics group calls on repositories to engage in R&D that advances semantic technologies, ontologies, and FAIR data alignment. “We recommend repositories participate in semantic research and development to test and deploy new metadata standards and ontologies for the Agri-food domain.” Marrano et al., 2025: Marrano et al. underline the importance of documenting R&D processes and ensuring that prototype services are transitioned into production systematically. “Repositories should make R&D activities transparent, describing the development process, testing stages, and how prototypes move to production environments.” Šestak & Copot, 2023: The authors link R&D to sustainability, recommending that repositories continuously innovate to adapt to new technologies and policy contexts. “A sustainable agri-data ecosystem relies on ongoing research and development in repository technology, ensuring adaptation to emerging scientific and policy challenges.” Ali & Dahlhaus, 2022: R&D is highlighted as necessary for improving repository technologies and integration capabilities across environmental and agricultural systems. “Research and development are essential to enhance interoperability, develop new data services, and maintain repository relevance within multidisciplinary infrastructures.” Biomedical Sciences The TTRAM Activity/Function AF27 Research & Development (R&D) is addressed in the following Biomedical Sciences resources: Sansone & Rocca-Serra, 2016: New funding frameworks need to be created to provide catalytic support for activities necessary to: research new or apply existing methods to develop, extend, refine and harmonize interoperability standards, and also related tools and educational material. 136 | Page
Field et al., 2011: GSC actively leads the creation and evolution of new data and metadata standards, such as the extension Minimum Information About a Marker Gene Sequence (MIMARKS). R&D includes establishing new business processes and technology by collaborating with major public databases. Climate Science The Development and exploitation of a controlled vocabulary in support of climate modelling effort (Moine et al., 2014) established a sophisticated information pipeline and tool chain to automatically convert the newly developed controlled vocabulary (CV) from human-readable mindmaps into machine-readable formats like XML and OWL ontologies, ensuring data preservation, reuse, and extensibility. Linguistics The Activity/Function (A|F) Research & Development (R&D) is not addressed by any of the reviewed resources from Linguistics. Social Sciences The Activity/Function (A|F) Research & Development (R&D) is not addressed by any of the reviewed resources from Social Sciences. 3.4. Technology Technology addresses the repository’s technical infrastructure, including storage, integrity measures, and IT service management, to support reliable and scalable digital object handling and interoperability. 3.4.1. STORAGE & INTEGRITY (AF28) Domain-agnostic resources The FIDELIS TTRAM defines the repository Activity/Function (A|F) Storage & Integrity is described as follows: “The methods used to store and replicate digital objects’ data and metadata, including the validation of replicated copies and the restoration of data from backups in case of errors.” and suggests for transparent information “[s]torage Documentation, integrity measures used, approach to integrity checking” (L’Hours et al., 2025c, p. 26). The A|F Storage & Integrity has its equivalent in the CoreTrustSeal requirement R14 (CoreTrustSeal Standards and Certification Board, 2022, pp. 24–25), which requires repositories to apply “documented processes to ensure data and metadata storage and integrity”, which means that repositories “[i]n addition to maintaining ‘archival’ copies of digital objects [...] need to store data and metadata from the point of deposit, for curation and preservation, and for access by users”, and that “[f]or each storage location, measures should be in place to ensure that unintentional or unauthorised changes can be detected and correct versions of data and metadata recovered”. To 137 | Page
harvesting metadata and SWORD for mediated submission, while ensuring resources are stored in machine-readable, non-proprietary formats to facilitate reuse. The GREI Data citation best practices for repositories (Puebla et al., 2024) dictate the required technical infrastructure by mandating the consistent use of DataCite metadata fields to store citation relationships. This necessary technical capacity must support the submission of collected data citations to DataCite and enable the harvesting of citation data from various external sources and aggregators, such as Crossref and Dimensions. The Core Preservation Process CPP-006 AIP Batch Export sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “As part of an exit strategy, the TDA batch exports Information packages and all associated Metadata in a manageable format/structure for Ingest into another TDA.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-008 File Format Identification sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA identifies file formats to the appropriate level of precision, based on an existing registry (IANA MIME types, PRONOM, etc.).” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-009 Metadata Extraction sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA extracts characteristics (such as size, image dimensions, video codec, audio run time, creating application).” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-010 Format Validation sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA validates Files against File format specifications.” (EOSC EDEN T1.2 et al., 2025). The Core Preservation Process CPP-016 Metadata Ingest and Management sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “The TDA ingests and manages all required Metadata including Metadata appropriate for specific content types (e.g. geospatial, audio visual).” (EOSC EDEN T1.2 et al., 2025). Within AF29 Technical Infrastructure, the Data Curation Network defines the following Curation Activity (Johnston et al., 2016): “Technology Monitoring and Refresh: Formal, periodic review and assessment to ensure responsiveness to technological developments and evolving requirements of the digital infrastructure and hardware storing the data.” Domain-specific resources Agri-food The TTRAM Activity/Function AF29 Technical Infrastructure is addressed in the following Agri-Food resources: 144 | Page
Harper et al., 2018: AgBioData emphasizes that repositories need scalable, sustainable technical infrastructure to support growth and FAIR data management. “Databases should adopt modern, scalable technical architectures that allow long-term maintenance and community access, ensuring continued usability of data and tools.” Caracciolo et al., 2020: The Agrisemantics group recommends that repositories adopt technical infrastructures capable of supporting semantic interoperability and FAIR-aligned tools. “Repositories should implement and maintain infrastructures that support machine-actionable metadata, persistent identifiers, and semantic alignment with community standards.” Marrano et al., 2025: Marrano et al. stress transparent documentation of repositories’ technical infrastructure, including hardware, software, and service management frameworks. “Repositories are encouraged to publish their technical infrastructure plans, outlining software environments, IT service management, and updates to ensure long-term reliability.” Ali & Dahlhaus, 2022: The authors highlight that repositories supporting agricultural and environmental data need integrated, cloud-based infrastructure for stability and scalability. Biomedical Sciences The TTRAM Activity/Function AF29 Technical Infrastructure is addressed in the following Biomedical Sciences resource: Karsch-Mizrachi et al., 2025: The article does not provide detailed specifications of hardware or IT service management. However, the described infrastructure and coordination among members demonstrate a robust technical foundation that supports the repository’s operations and responsiveness to evolving scientific and technological demands. Climate Science The WDCC User Guide for Data Publications (Long Term Archive (LTA) group, 2024) states WDCC technical infrastructure includes hosting on systems like Levante and supporting open-source data formats like NetCDF, GRIB, and Zarr, along with the MetaXA graphical user interface for metadata submission. This is complemented by an IT service management, offering user guidance and automated technical quality assurance. Linguistics CLARIN B and E centres are required “to have a proper and clearly specified repository system” (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3). Social Sciences The CESSDA Resource Directory (CESSDA, 2025) contains links to resources on the topic of “Technical Infrastructure” that are relevant to this A|F (https://www.cessda.eu/Resource-Directory?tree=28). The CESSDA Resource Directory (RD) lists resources with the intention to “help to build sustainable 145 | Page
and mature data archives and support the development of new services and features within existing data archives. Information on relevant documents, training materials, tools and support services are collected, selected and reviewed, making the RD a curated inventory of existing resources” (CESSDA, 2025). 3.5. Security Security focuses on safeguarding digital assets and systems through comprehensive security policies, risk management, and compliance, ensuring trust in repository operations. 3.5.1. SECURITY (AF30) Domain-agnostic resources The FIDELIS TTRAM: defines the repository Activity/Function (A|F) Security is described as follows: “Ensuring security across the organisation’s infrastructure, digital object management, and systems. Security also extends to the boundary between the repository and external entities, including users, depositors and dependencies on third parties.” and suggests for transparent information “Information Security Statement, Information Security Plan, related certifications (e.g. ISO27001)” (L’Hours et al., 2025c, p. 27). The A|F Security has its equivalent in the CoreTrustSeal requirement R14 (CoreTrustSeal Standards and Certification Board, 2022, pp. 26–27), which requires repositories to “protect[...] the facility and its data, metadata, products, services, and users”, and asks applicants to address and provide evidence for the following items: ● The levels of security required for different data and metadata and environments, and how these are supported. ● The IT security system, employees with roles related to security (e.g. security officers), and any risk analysis approach in use. ● Measures in place to protect the facility. How the premises where digital objects are held are secured. ● Any security-specific standards the repository references or complies with. ● Any authentication and authorization procedures employed to securely manage access to systems in use. The nestor Seal addresses aspects of this A|F in “C34 Security”, which requires the organisation and the infrastructure protect the digital archive and its archived information objects and representations (nestor Certification Working Group, 2025). In the OAIS Reference Model this A|F is covered by the “Common Services” (chapter 4.2.3.2) which include Operating System services, Network services and Security. In addition, Physical Access Control is a function in the Administration functional entity (chapter 4.2.3.6). The Desirable Characteristics of Data Repositories for Federally Funded Research (White House Office of Science and Technology Policy, 2022) state that the repository has documented measures in 146 | Page
place to meet well established cybersecurity criteria for preventing unauthorized access to, modification of, or release of data, with levels of security that are appropriate to the sensitivity of data (e.g., the NIST Cybersecurity Framework: https://www.nist.gov/cyberframework). The EU Annotated Grant Agreement states repositories display specific characteristics of organisational, technical and procedural quality, such as services, mechanisms and/or provisions that are intended to secure the integrity and authenticity of their contents, thus facilitating their use and re-use in the shortand long-term. Trusted repositories have specific provisions in place and offer explicit information online about their policies, which define their services (e.g. acquisition, access, security of content, long-term sustainability of service including funding, etc). They meet generally accepted international and national criteria for security to prevent unauthorized access and release of content and have different levels of security, depending on the sensitivity of the data being deposited, to maintain privacy and confidentiality (European Commission, 2024). The Update of the Study on the readiness of research data and literature repositories to facilitate compliance with the Open Science Horizon Europe MGA requirements states the lack of a public policy for preservation, curation and security of the contents is the most frequent reason, followed by not adhering to a specific metadata standard (Lazzeri, 2024). The REPOSITORIES: Key Infrastructure For Μaintaining European Research Εxcellence (Shearer et al., 2025) states that resilient & secure against cyber threats and technological changes can be mitigated through ensuring regular backups, infrastructure monitoring, and compliance with best practices in long-term data integrity and protection. For depositing personal data that cannot be anonymised, the exclusion criteria from the Research Data College Working Group rule out subject-specific repositories that are “located outside the European Union [...], with the exception of Switzerland, Great Britain, Japan and Argentina, which are deemed to be GDPRcompliant” (Lamotte et al., 2024, p. 11). The M5.2—Guidelines for repositories and registries on exposing repository trustworthiness status and FAIR data assessments outcomes (Verburg et al., 2023) specify that repositories rely on a wider partnership of data and metadata services, such as storage providers and registries, and therefore the security provision must explicitly extend to the boundary between the repository and these external entities. Transparent information related to security, such as an Information Security Statement, an Information Security Plan, and related certifications like ISO27001, should be exposed by repositories. The D8.3 Trustworthy Digital Repository status update and certification solutions for SSHOC repositories (Kleemola et al., 2022) states that information security is another challenge when outsourcing, particularly for repositories that store sensitive digital objects and/or significant amounts of personal data (e.g., data concerning producers, repositories, researchers using the datasets etc.). A breach of information security (even when the repository is not responsible) can result in the repository sustaining significant reputational damage and possible legal consequences. 147 | Page
The COAR Community Framework for Good Practices in Repositories, Version 2 (Confederation of Open Access Repositories, 2022) states that ensuring security in repositories requires the application of security practices to prevent unauthorized manipulation of resources and the regular performance of integrity checks to detect unauthorized changes or accidental damage to digital objects. Security across the organization's infrastructure is also supported by maintaining a business continuity plan that details the response and procedures for handling cyber-attacks or natural disasters. The Core Preservation Process CPP-007 Virus Scanning sets the following good-practice baseline expectation for Trustworthy Digital Archives (TDA): “Information packages are virus checked, with appropriate facilities for quarantine.” (EOSC EDEN T1.2 et al., 2025). Domain-specific resources Agri-food The TTRAM Activity/Function AF30 Security is addressed in the following Agri-Food resources: Harper et al., 2018: AgBioData underlines the need for security policies that balance openness with data protection. “Repositories must define access control mechanisms and data security policies to ensure that sensitive information is protected while maintaining data availability for legitimate users.” Caracciolo et al., 2020: The Agrisemantics group links repository security to responsible stewardship of semantic and metadata resources. “Ensuring reliable and secure access to semantic resources is essential for maintaining their integrity and long-term usability within Agri-food repositories.” Marrano et al., 2025: Marrano et al. recommend formal documentation of information security procedures and external certification. “Repositories should maintain and publish an Information Security Statement, describing access control, risk management, and, where applicable, ISO 27001 or equivalent certification.” Šestak & Copot, 2023: The authors connect security with repository sustainability and resilience. “Secure data infrastructures are foundational to the resilience of agri-data ecosystems, protecting data integrity and ensuring continuity of service.” Ali & Dahlhaus, 2022: Security is discussed in the context of shared environmental–agricultural infrastructures. “Secure access and controlled sharing mechanisms are essential for maintaining trust between repositories and data providers in integrated agricultural systems.” Biomedical Sciences The TTRAM Activity/Function AF30 Security is addressed in the following Biomedical Sciences resources: 148 | Page
Lin et al., 2024, emphasize the critical importance of ensuring the security and integrity of data and metadata in repositories. They advocate for the use of documented and appropriate measures, considering these essential characteristics of trustworthy repositories. For repositories that manage human data, the authors highlight the need for a comprehensive breach response plan to address potential security incidents and unauthorized access. Furthermore, the article promotes the TRUST Principles, particularly the role of technology in delivering secure, persistent, and reliable services. Repositories are encouraged to adopt relevant and effective technologies and practices to mitigate security threats. Notably, technology infrastructure and security are identified as one of the three core assessment areas in evaluating repository trustworthiness and in certification standards. Climate Science The Activity/Function (A|F) Security is not addressed by any of the reviewed resources from Climate Science. Linguistics CLARIN B and E centres are required “to join the national identity federation where available and join the CLARIN service provider federation to support single identity and single sign-on operation based on SAML2.0 and trust declarations. In case all resources at a centre are open, setting up a Service Pro For CLARIN B and E centres, the following requirements apply when it comes to the AF Security (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023, p. 3): ● Centres need to adhere to the security guidelines, i.e. the servers need to have accepted certificates. ● Centres need to join the national identity federation where available and join the CLARIN service provider federation to support single identity and single sign-on operation based on SAML2.0 and trust declarations. In case all resources at a centre are open, setting up a Service Provider is optional. For CLARIN B centres, the following requirement applies (Wittenburg, Van Uytvanck, Zastrow, Straňák, et al., 2023): ● 4. Server Certificates. Requirement: Centres need to adhere to the security guidelines, i.e. the servers need to have accepted certificates. ● 5. Federated Identity Management. [Note: if a centre only provides and will provide fully open resources, this requirement is not applicable]. Requirement: Centres need to join the national identity federation where available and join the CLARIN service provider federation to support single identity and single sign-on operation based on SAML2.0 and trust declarations. ● 9. Attribute Checker (optional). Requirement: Centres can opt to configure the Shibboleth SP Attribute Checker which assists in case of failed logins. ● 10. Attribute Aggregator (optional). Requirement: Centres can opt to configure the Attribute Aggregator provided by Lindat which provides insights into failed login attempts. 149 | Page
Social Sciences The Activity/Function (A|F) Security is not addressed by any of the reviewed resources from Social Sciences. 3.6. Other capabilities and characteristics Some frameworks include other capabilities and characteristics than the ones in the FIDELIS TTRAM. These are briefly addressed in this section. The Data Curation Network includes peer-review among its curation activities and describes it as follows (Johnston et al., 2016): “The review of a data set by an expert with similar credentials and subject knowledge as the data creator for the purposes of validating the soundness and trustworthiness of the file contents.” Within the field of Linguistics, CLARIN B centres are required to have to be recognized by a CLARIN ERIC member or observer country or have to be established as a third party (Wittenburg, Van Uytvanck, Zastrow, & Offersgaard, 2023). The exclusion criteria from the Research Data College Working Group rule out subject-specific repositories that practice an excessive pricing policy, i.e., “repositories where every low-volume data deposit automatically incurs a fee [...]”, whereas “repositories that may require a financial contribution in return for depositing significant volumes of data (more than 50 GB) have not been excluded” (Lamotte et al., 2024, p. 11). 4. Analysis and Discussion This chapter discusses key aspects of the reviewed and mapped repository capabilities and characteristics presented in Chapter 3. In Section 4.1, we briefly describe the characteristics of the reviewed resources. Section 4.2 summarises identified commonalities across the reviewed resources, whereas Section 4.3 presents specificities that stand out across resources. 4.1. Resource Characteristics During the work on this report, a total of 85 resources were reviewed and added to the mapping spreadsheet. Of these, 75 were included in the present report. 4.1.1. Network Beacon Communities The resource mapping conducted in Task 5.2 of the FIDELIS project reviewed a total of 75 resources, which were categorized by their relevance to the six designated network beacon communities. These communities represent key scientific domains within the FIDELIS network and serve as focal points for aligning repository practices with the EOSC framework. The distribution of resources across the network beacon communities is as follows: ● Agri-food: 10 resources ● Biomedical Sciences: 4 resources 150 | Page
● Climate Science: 4 resources ● Linguistics: 7 resources ● Social Sciences: 8 resources ● Generic (domain-agnostic): 33 resources ● Cross-domain resources (relevant to multiple communities): o Biomedical Sciences, Climate Science, Linguistics, Social Sciences, Physical Sciences, Other: 1 resource o Biomedical Sciences, Generic: 1 resource o Linguistics, Social Sciences, Other: 2 resources o Other discipline, Biomedical Sciences: 4 resources o Social Sciences, Generic: 1 resource Notably, while the Physical Sciences community was not directly represented due to capacity limitations, it was indirectly covered through cross-domain resources. The Generic category, comprising 33 resources, reflects the strong presence of domain-agnostic standards, frameworks, and best practices applicable across all repository types. This distribution highlights the broad engagement of the FIDELIS project with both domain-specific and cross-cutting repository practices, ensuring that the resulting recommendations and frameworks (e.g., TTRAM) are grounded in a diverse and representative evidence base. 4.1.2. Resource Types The reviewed resources can be categorized into the following types, often with overlapping classifications: ● Standards: Formalized technical or procedural specifications (e.g., MIxS, CF Metadata Conventions, OAIS-RM) ● Recommendations: Community or project-based guidance (e.g., FAIRsFAIR D2.3, COAR Community Framework) ● Best Practices: Descriptive or prescriptive practices from domain-specific initiatives (e.g., AgBioData, WheatIS) ● Modes of Federation: Descriptions of federated repository models (e.g., CLARIN Centre Registry, INSDC) ● Solutions: Technical or organizational tools and frameworks (e.g., TTRAM, AgroDataCube, GA4GH APIs) ● Certification Bodies and Frameworks: CoreTrustSeal, nestor ● Guidelines and Handbooks: EOSC Federation Handbook, FAIR Data Maturity Model, EU Annotated Grant Agreement ● Surveys and Landscape Analyses: SSHOC, EOSC-Nordic, FAIRsFAIR, OpenAIRE/COAR studies These resource types were mapped against the 30 TTRAM Activities and Functions (AF01–AF30), providing a structured overview of how each resource related to repository capabilities and characteristics in the following five main areas: 151 | Page
● Context ● Digital Object Management ● Organisational Infrastructure ● Technology ● Security The remainder of Section 4 provides a comprehensive analysis and discussion of the mapping results, organised according to the five main TTRAM areas. 4.2. Context The Context area of the FIDELIS Transparent Trustworthy Repository Attributes Matrix (TTRAM) encompasses foundational elements that define a repository’s identity, mission, and scope. This includes how repositories present themselves (AF01: Identification & Contact) and articulate their purpose and responsibilities (AF02: Mission & Scope). The analysis of 75 resources across domain-agnostic and domain-specific contexts reveals both shared practices and community-specific nuances. 4.2.1. Identification & Contact (AF01) Commonalities: Across all Network Beacon Communities (NBCs), a strong emphasis is placed on transparency in repository identification. Domain-agnostic standards such as CoreTrustSeal (R0) and the FAIRsFAIR D2.3 report recommend the inclusion of persistent identifiers (e.g., DOIs, ROR IDs), clear contact information, and registration in trusted registries like re3data. These practices are foundational for discoverability and trust. Specificities by NBC: ● Agri-food repositories, as highlighted in Harper et al. (2018) and Drakos et al. (2015), stress the importance of naming responsible individuals and providing direct contact methods. Registration in agricultural aggregators like CIARD RING and AgriVIVO is also emphasized. ● Biomedical Sciences resources such as Durinx et al. (2017) and Lin et al. (2024) underscore the need for detailed repository metadata, including organizational affiliations and support contacts, especially for ELIXIR Core Data Resources. ● Linguistics repositories, particularly CLARIN centres, are required to visibly reference their affiliation with CLARIN and register in the Centre Registry, which feeds into re3data. ● Social Sciences repositories, as per CESSDA guidelines, are encouraged to register in multiple registries (e.g., OpenDOAR, FAIRsharing) and provide persistent identifiers. Resource Types: Standards and best practices dominate this AF, with CoreTrustSeal, FAIRsFAIR, and OAIS-RM providing comprehensive guidance. Domain-specific recommendations often take the form of best practices and community guidelines. 4.2.2. Mission & Scope (AF02) 152 | Page
Commonalities: Most repositories, regardless of domain, articulate a mission aligned with long-term preservation, open access, and FAIR principles. The CoreTrustSeal R01, nestor Seal, and OAIS Reference Model provide structured expectations for defining mission, designated communities, and scope of services. The EU Annotated Grant Agreement and FAIRsFAIR further emphasize the importance of transparency in repository responsibilities and alignment with community needs. Specificities by NBC: ● Agri-food repositories, such as those discussed by Drakos et al. (2015) and Harper et al. (2018), often tailor their missions to specific crops, data types, or stakeholder groups. The emphasis is on aligning repository goals with agricultural research and FAIR data sharing. ● Biomedical Sciences resources, notably Durinx et al. (2017) and Lin et al. (2024), distinguish between deposition databases and knowledge bases, with missions focused on supporting clinical and genomic research. Repositories are expected to define their scientific scope, user communities, and data types (e.g., nucleotide sequences, clinical data). ● Linguistics repositories, particularly CLARIN centres, must define their role within the infrastructure and specify the services and data they offer to the research community. ● Social Sciences repositories, as outlined in the CESSDA Data Management Expert Guide, focus on integrating data into the research lifecycle, supporting teaching, learning, and policy-making. Resource Types: This AF is supported by a wide range of resource types: ● Standards: OAIS-RM, CoreTrustSeal, nestor Seal ● Best Practices: AgBioData, WheatIS, CLARIN Requirements ● Recommendations: EU Annotated Grant Agreement, FAIRsFAIR, Research Data College Working Group Discussion: The Context area reveals a high degree of alignment across domains in terms of transparency and mission articulation. However, the specificities of each NBC – such as the need for domain-specific identifiers, community engagement strategies, and tailored mission statements –highlight the importance of contextualizing repository practices. The Generic resources provide a robust foundation, but domain-specific adaptations are essential for meaningful alignment with community expectations and EOSC integration. 4.3. Digital Object Management The Digital Object Management area of the TTRAM encompasses the full lifecycle of digital objects, from their conception and creation to deposit, curation, discovery, access, reuse, and preservation. This area is the most densely populated in the mapping, reflecting its centrality to repository operations. The analysis of 75 reviewed resources reveals a rich landscape of domain-agnostic standards and domain-specific practices, with varying levels of maturity and specialization across the Network Beacon Communities (NBCs). 153 | Page