FAIRCORE4EOSC Deliverable D5.1 - Report on Best Practices and Governance of Kernel Information Profiles
Abstract
Persistent identification and extensive metadata are essential for decision-making and processing digital objects. The Kernel Information, which is organised in Profiles and described by them, define the minimum requirements for metadata. This document describes current implementations and makes suggestions for future use.
Full text
FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. PAGE \* MERGEFORMAT 2 Project Number: 101057264 Start Date of Project: 01/06/2022 Duration: 36 months Deliverable 5.1 Report on Best Practices and Governance of Kernel Information Profiles Dissemination Level PU Due Date of Deliverable 31/03/2025, Month M32 Actual Submission Date 31/03/2025 Work Package WP 5, EOSC PID meta resolver and PID Kernel Information Profiles Task T5.3 Type R - Document, Report Version V1.0 Number of Pages p.1 – p.24 Deliverable Abstract Persistent identification and extensive metadata are essential for decision-making and processing digital objects. The Kernel Information, which is organised in Profiles and described by them, define the minimum requirements for metadata. This document describes current implementations and makes suggestions for future use. The information in this document reflects only the author’s views and the European Community is not liable for any use that may be made of the information contained therein. The information in this document is provided “as is” without guarantee or warranty of any kind, express or implied, including but not limited to the fitness of the information for a particular purpose. The user thereof uses the information at his/ her sole risk and liability.This deliverable is licensed under a Creative Commons Attribution 4.0 International License.
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 2 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. DELIVERY SLIP Contribution Name Partner/Activity Date Lead author(s) Sven Bingert GWDG 26/03/2025 Contributor(s) Wilko Steinhoff KNAW-DANS 26/03/2025 Wim Hugo KNAW-DANS 26/03/2025 Themis Zamani GRNET 26/03/2025 Richard Hallet DATACITE 26/03/2025 Morane Grunpeter INRIA 26/03/2025 Reviewer(s) Lassi Lager CSC 26/03/2025 Heinrich Widmann DKRZ 26/03/2025 DOCUMENT LOG Issue Date Comment Author/Editor/Reviewer v.0.1 14/03/2025 Sven Binger t v.0.2 20/03/2025 Internal review Lassi Lager, Heinrich Widmann v.0.3 26/03/2025 Incorporating reviewer suggestions All v.1 27/03/2025 Review by Project Lead v.1.1 27/03/2025 Incorporated internal reviewer comments Lassi Lager, Sven Bingert v.1.2 28/03/2025 New publication in Kernel Metadata included Lassi Lager, Sven Bingert v.1. 3 31/03/2025 Final version Elisa Halonen
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 3 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. Table of Contents 1Introduction ...................................................................................................................................................... 1 1.1 Strategic Research and Innovation Agenda ........................................................................................... 2 1.2 A Persistent Identifier (PID) Policy for the European Open Science Cloud .......................................... 2 1.3 General Workflow ..................................................................................................................................... 4 1.4 Relationship Between Custom and Kernel Metadata ............................................................................ 5 2Implementation Examples of Kernel Information Profiles ............................................................................ 6 2.1 DataCite DOI ............................................................................................................................................. 6 2.2 URN:NBN:FI .............................................................................................................................................. 8 2.3 SWHID - The Software Hash IDentifier ................................................................................................... 8 2.4 Signposting Approach ............................................................................................................................. 9 2.5 Fair Digital Object Approach ................................................................................................................. 12 2.6 Data Type Registry for PID Records...................................................................................................... 14 2.6.1 DTR Implementation examples .................................................................................................... 15 2.7 Comparison of Approaches to Kernel Information Provision ............................................................. 17 3Governance of Kernel Information Profile .................................................................................................... 18 3.1 Elements and tasks of the governance ................................................................................................ 18 4Best Practices and Conclusion ..................................................................................................................... 19 4.1 Governance-Related Recommendations ............................................................................................. 20 4.2 Optimal Metadata Management ........................................................................................................... 20 4.3 Implementation of Kernel Information Profiles ................................................................................... 20 4.4 Harmonisation of Features and Performance Measurement ............................................................. 21 5Summary ......................................................................................................................................................... 23 6References ...................................................................................................................................................... 23
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 4 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. List of Figures Figure 1: Actors and roles in the PID ecosystem. Taken from [R10] ..................................................................... 3 Figure 2: Visual representation of the workflow to retrieve the Kernel Information for making reasoned decisions. .................................................................................................................................................................. 5 Figure 3: Schematic overview of Signposting links that have the landing page, each of the metadata resources and each of the content resources as link origin. ................................................................................................ 11 Figure 4: Relation between the PID record and the FDO Profile for referencing FDOs, taken from [R7] ........... 13 Figure 5: Simple graphical illustration of the configuration type 4 taken from [R7]. .......................................... 14 Figure 6: PID of a Digital Specimen, using the DiSSCover Service (https://disscover.dissco.eu/search), and resolving the PID to the PID Record using the global resolver (http://hdl.handle.net) ...................................... 16 List of Tables Table 1: List of recommendations with the affected stakeholder and the problem approached. Table is taken from [R12] ............................................................................................................................................................... 23
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 1 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. Executive Summary Kernel Information Profiles are a tool for providing metadata that is necessary for the understanding and further use of the PID referenced information. Which information can be designated as necessary (kernel) depends heavily on the application and the domain. Earlier work has produced proposals and schemas. These are not applicable to all use cases. In addition, it is necessary to clarify how this kernel information can be accessed and how the profile can be interpreted. A generally valid and generic technical solution has not yet been developed. Rather, the current situation shows that a variety of PID systems are used, which in turn are applied in different ways. This document looks at the different aspects and describes how the kernel information concept is implemented in the different approaches. To begin with, the most important references to Kernel Information Profiles are reviewed to understand the requirements. In the second section, the examples of different implementations are discussed. These also include examples of how to use the Data Type Registry, which can be used to register the kernel information types and the kernel information profiles. The definition of kernel information and kernel information profiles should take place in coordinated processes. This is discussed in the chapter on governance. In the last chapter, best practices are derived from the collective experiences. 1 Introduction The term Kernel Information Profile (KIP) was mostly driven by the RDA working group on PID Kernel Information Profile Management. In their recommendation1 the “information contained in a PID record is represented by a PID Kernel Information profile”. It also states that the “PID Kernel Information profiles are registered schemas for interpreting PID Kernel Information records”. A record is the set of attributes that contain the kernel information. In more general terms, following the recommendation, the Kernel Information Profile describes the set of metadata to be expected when resolving a PID. As the process of resolving a PID depends on the PID system used, different approaches on how to implement Kernel Information Profiles have to be regarded. The term “Kernel” implies a mandatory set of attributes to be elements of a PID record. To define this mandatory set, detailed domain knowledge is required or a generic profile can be defined. The RDA WG on PID Kernel Information Profile Management provided a recommendation for a generic domain agnostic profile (see chapter 3 in [R2]). This profile includes, e.g., the attribute “etag” which should contain the checksum of object contents2, to which the PID refers. But KIPs can also be used to describe the records of PIDs pointing to other entities, like services, operations, and type definitions; in this case the concept of using a checksum might not be appropriate. An even more generic approach is required to guide a machine or a human in the decision process by providing metadata. The starting point of that process should be the profile. The minimum requirement of attributes, the kernel, is the profile itself and a set of attributes which will be discussed in later sections. 1 Weigel, T., Plale, B., Parsons, M., Zhou, G., Luo, Y., Schwardmann, U., Quick, R., Hellström, M., & Kurakawa, K. (2018). RDA Recommendation on PID Kernel Information. Research Data Alliance. [R2] 2 In the newest discussions the term digital object is used.
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 2 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 1.1 Strategic Research and Innovation Agenda In the Strategic Research and Innovation Agenda (SRIA) [R5] a gap analysis revealed, among others, the following two issues in the support for machine-actionable PIDs: ●PID Type Registries and accompanying Kernel profiles are not as yet standardised for different machine-readable data types and automated processing is largely missing or considered experimental. ●EOSC should consider the support of PID Type Registry services within EOSC, and develop services and use cases that exploit these services in automatic data analysis. This shows that machine-readability and machine-actionability is not yet achieved using PIDs. Machine readability refers to the capability of a computer system to analyse a given data or information without the need for human intervention. In order for (meta-)data to be machine-readable, it needs to be structured in a way that a computer can understand. Machine-actionability is one level higher and means that the (meta-)data is in a structured format that a computer can understand and act upon. Thus, by using machine-actionable (meta-)data software, a service can derive and decide on possible operations on the (meta-)data. The SRIA derived, among others, the following action item: ●Define specifications (schemata) for PID records / kernel information to support machine-actionable PIDs. Later on, in the “conclusion” section, we will evaluate whether this action item was successfully implemented. 1.2 A Persistent Identifier (PID) Policy for the European Open Science Cloud The PID Policy [R1] provides a set of generic definitions for the uniqueness, persistence, and resolvability of PIDs. PID resolution can serve at least three different purposes as resolving to the Digital Object (DO) itself, to a digital representation of the object (e.g. the landing page), or to information on how to access the object. Kernel Information is used to provide that information in a structured way. A Kernel Information Profile, registered in a public repository, serves as the machine-actionable implementation of the Kernel Information structure. Kernel Information can also be used for disambiguation when more than one identifier refers to the same object. Chapter 6 ‘PID Types’ of the EOSC PID Policy is relevant for the implementation of KIPs. It states: ●6.3: Classes of digital objects may need different attribute sets a PID is resolved to. It is the responsibility of a community of practice to define and document these attribute sets (PID Kernel Information Profiles). Defining a KIP requires a deep technical and domain understanding. In addition, there is a risk that the number of profiles to be specified will become very large, since there are a large number of digital object types or classes. Furthermore, it should be considered here that the same digital objects in different domains are probably described with different profiles, or even duplicates are created. The latter is not preferable for obvious reasons such as maintainability. Also, the definition focuses on digital objects and excludes services, operations, and other relevant information. In the section 2.5, the definition of the use of Kernel Information is extended to include precisely these elements. The EOSC PID policy also states the following in chapter 6:
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 3 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. ●6.4: The PID Kernel Information record is a non-authoritative source for metadata focused on facilitating automation of processes. If the information for an attribute duplicates metadata maintained elsewhere, the external source is the authority. Here two aspects are of importance: 1) the number of duplicate (non-authoritative) information in the PID record should be minimized and 2) automation of processes should be facilitated. The latter is in line with the requirement for machine actionable data described in the SRIA. Additionally, the reference to the external source with the authoritative metadata should be given in the kernel. It is important to take note of the definition of Actors in the PID Ecosystem, defined in the EOSC PID Policy, and elaborated further in the work done in collaboration between FAIRCORE4EOSC and FAIR-IMPACT [R10]. Figure 1summarises these actors and their relationship, and the actors and their roles are discussed in more detail below. Figure 1: Actors and roles in the PID ecosystem. Taken from [R10]
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 4 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 1. Many (but not all) PID Services are based on a published Scheme that is governed and maintained by an Authority, and is often linked to or based on one or more published Standards3. In some cases, the Authority is designated by the Standards Body to manage the Scheme. In other cases, an Authority may adopt a Scheme published by a Standards Body. 2. An Authority serves as the guarantor of identifier uniqueness, and often handles primary resolution duties and offers kernel metadata. Some Authorities provide PID registration, editing, and resolution services directly to the Owners and Users of the identifiers. 3. In some cases, the Authority directly enables a Provider to act as local registration and optionally resolution provider, and to optionally extend or manage a kernel metadata schema. Some PID Services, however, introduce an additional layer of resolution and/ or uniqueness management between Providers and the Authority (Multi-Provider Agencies or MPAs). Such MPAs may also extend the metadata kernel information defined in the Scheme. 4. Providers (frequently called ‘Services’) enable Managers to register PIDs on behalf of Owners, and Managers often make use of value-added services from the Provider. More often than not, Managers are repositories of physical and digital materials, or registries of concepts, and several other terms may be used to describe them - e.g. ‘Allocating Agents’. 5. The Provider maintains the relation between the referenced entity, the identifier, and the detailed kernel metadata for that object or concept where applicable - usually on behalf of the Owner of the object or concept. 6. When Users (humans or machines) encounter a PID, resolution is handled by the Provider in most cases, but if a PID was issued directly by an Authority, resolution is handled by the Authority. Sometimes, resolution is cascaded - redirected from a central registry at the Authority to a resolution mechanism at the Provider. 7. Resolution directs the user either directly to the entity (object) or to its proxy, which is often a (metadata) landing page maintained by the Manager on behalf of the Owner. In such cases, it is possible to obtain the entity either directly from the Owner, or from the Manager. This process may not be machine-actionable and may involve manual permission, retrieval, and sharing. 1.3 General Workflow The most important reason for implementing Kernel Information Profiles is to use this information to make informed decisions. It does not matter whether these decisions are made by machines or humans. The kernel information can help with making decisions in various areas. This information can be technical, contentrelated, legal or licensing-related. Based on this information, a decision can be made as to whether the information referenced by the PID is relevant and usable, or whether the necessary rights for (re-)use are available. Figure 2 shows the general process. The starting point is that the desired information/digital object is referenced with a PID. In step 1 of the workflow, the PID is resolved and the resolver is asked for the metadata. The specific call for how the metadata for a PID can be retrieved depends on the specific PID system. Examples are given in the sections below. If possible and available the resolver will return the metadata in step 3 In reality, many of the standards listed for PID stacks are not yet adopted by a Standards Body - they are often published RFCs that have not yet been adopted but serve as de facto standards [R11].
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 5 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 2. In order to understand the given metadata, a (technical) description of the profile is loaded in step 3. The attributes of the metadata become syntactically and semantically comprehensible. After the metadata has been processed and the decision made, the data object can then be requested in the last steps (step 4) and, if authentication and authorisation are successful, it can also be loaded in step 5. Figure 2: Visual representation of the workflow to retrieve the Kernel Information for making reasoned decisions. 1.4 Relationship Between Custom and Kernel Metadata In the workflow shown in Figure 2, the metadata service returns a metadata record, but the nature of this record is strongly dependent on the PID service. We define three types of metadata that could be returned. Primarily, a distinction is made between Kernel Metadata and Custom Metadata [R10]: In the EOSC PID Policy, specific responsibilities for maintenance and integrity of these metadata records are associated with specific actors in the ecosystem. a. Kernel Metadata: this is provided by the formalised system that offers the PID Stack, and standardised or based on well-publicised and managed schema in most (but not all) cases. Kernel metadata also functionally has two distinct parts, although these are sometimes made available in a single schema. b. Identifier Metadata:typically, information about the account that created the identifier, date of creation and/ or modification, and the original or current redirection URL(s) that are linked to the identifier. This metadata is maintained by the resolution service, usually operated by the Authority (which may be chained or federated in some cases). c. Resource Metadata:typically standardised and offered by the Provider, describes the resource in formal metadata terms, dealing with aspects of citation, coverage, subject and keywords, rights and licensing, and more. It is important to note that the resource metadata often only reflects the latest version, and that not all PID Stacks have standardised metadata schema. In this document, the record that is expected from a metadata service is assumed to be the kernel metadata record.
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 12 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. The content resource may also provide other Signposting links even though those will be redundant with the full set of links provided in the Link Set. The example HEAD response shows a "collection" link that points back to the landing page. The linkset16 will, amongst others, link to the landing page URI (“anchor”) and metadata (“describedby”) expressed in the given schema. This can be seen in the following json response: "linkset": [ { "anchor": "https://dataverse.nl/dataset.xhtml?persistentId=doi:10.34894/O1CHN7", "cite-as": [ { "href": "https://doi.org/10.5061/dryad.5d23f" } ], "describedby":[ { "href": "https://example.org/meta/7507/bibtex", "type": "application/x-bibtex" }, { "href": "https://doi.org/10.5061/dryad.5d23f", "type": "application/vnd.datacite.datacite+json" } ] (...) In the case of Signposting, the kernel information is not associated with a global or local PID registry as stated in the RDA PID Kernel Information Guiding Principle 3: “PID Kernel Information is stored directly at the resolving service and not referenced.” Most of the time the information will be available from the repository the object lives in. However, validating the existence of the minimum set is possible and accessible by a PID resolver or any other agents. FAIR Signposting therefore complies with 6 of the 7 RDA guiding principles and adheres to the primary purpose of PID Kernel Information stated by the RDA: “The primary purpose of PID Kernel Information is in support of smart machine actionable decisions that can be accomplished through inspection of the PID record alone.” Signposting therefore may be considered a “semi” Kernel Information Profile that adheres to most of the principles. In summary, it is good to see things in perspective here. The authority regarding the metadata for a scholarly object is the custodian of that object itself. After all, that is where the object is managed. The metadata in a PID registry is actually derived from the authoritative metadata of the custodian. This makes the Signposting approach a different one than a kernel Information profile. 2.5 Fair Digital Object Approach The Fair Digital Object Forum (FDO Forum or FDOF) published the core specifications17 to implement FDOs. The specification on implementation of Attributes, Types, Profiles and Registries [R4] requires the usage profiles to describe the attributes provided with the response. The FDO profile describes the structure of the (PID) record. This description lists the allowed and mandatory attributes for that specific record. Compared to a general KIP, an FDO profile requires additional information to reference digital objects as FDOs. 16 https://dataverse.nl/api/datasets/:persistentId/versions/1.0/linkset?persistentId=doi:10.34894/O1CHN7 17 https://fairdo.org/specifications/
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 13 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. According to the FDO Requirement Specifications 3.0 [R6] in the section 2.3 “FDO Layer Specification”: ●FDO-FDOR1: The content of each FDO record must be structured according to an FDO profile in accordance with an FDO defined schema registered in a recognized registry. ●FDO-FDOR2: The FDO record consists of a set of attribute-value18 pairs as defined by the FDO Profile and all used attributes need to be defined and registered according to the type specification schema. ●FDO-FDOR3: Each FDO record needs to contain the mandatory kernel attributes, as defined by the FDO profile, including the type of the FDO. Additionally, in [R4] it says: ● An FDO record contains as one needed attribute a key for the profile key and the value to this key is the specific FDO profile for the FDO. Thus, the reference, the PID, to the profile definition is an element of the profile itself. This allows machines to derive further processing steps by reading the profile definition before processing other elements of the PID record. Figure 4: Relation between the PID record and the FDO Profile for referencing FDOs, taken from [R7] As depicted in Figure 4 the FDO configuration defines the mandatory attributes within the FDO Profile. The FDO configurations are categorized in different types. Those types are initially discussed in [R7]. Although they are not yet part of the FDO specifications, their relevance is growing. Testbed implementation in the Project FDO One19, founded by the German Ministry for Digital and Transport, showed the importance of the configurations. To understand the FDO configuration types an example is provided here. We will examine configuration type 4 (c.f. Figure 5 and [R7]) to understand the process of ultimately using the bit sequence of a FDO. The process is as follows: 1. Resolve PID1 into the PID Record. This can be achieved, for example, by not redirecting when a URL is provided in a handle-based PID or by requesting a dedicated endpoint to retrieve the PID record. 18 For the reader: in other parts of the document the attribute is defined as a key-value pair. 19 https://fdo-one.org/
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 14 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 2. The PID Record will then contain: a. a reference to a Metadata object given as PID, b. a reference to a Bit-sequence given as PID, c. a reference to the profile definition, where the profile definition (configuration type 4) states that the mandatory attributes are i. the reference to the profile ii. a reference to a Metadata object given as PID iii. a reference to a Bit-sequence given as PID. The profile definition does not exclude other attributes and communities may want to derive a community specific configuration type 4 profile where community specific attributes are mandatory. The profile definitions should be registered in the open registries. 3. Resolve PID2 into the PID Record. 4. The PID2 Record contains a reference (URI) to the Metadata object. To follow the FDO idea the PID2 Record should also contain a reference to a profile definition (here configuration type 1) and may contain additional non-mandatory community specific attributes. 5. Resolve PID3 into the PID Record. 6. The PID3 Record contains a reference (URI) to the Bit-sequence. Same as for PID2, the PID3 record should have a profile and optional community specific attributes. 7. After examination of all PID Records (attributes) and the referenced Metadata an informed decision on processing the Bit-sequence can be done. Figure 5: Simple graphical illustration of the configuration type 4 taken from [R7]. In this example the minimum set of profiles is two, one for the main PID1 and one for reference objects via URI. These are registered in the FAIRCORE4EOSC Data Type Registry: ●Configuration Type 4: https://hdl.handle.net/21.T11969/16fae4022f0e5e75ca17 ●Configuration Type 1: https://hdl.handle.net/21.T11969/407f8043f1db837611ca In those profile definitions the attribute for the profile reference is called “FDO_Profile_Ref”. The definitions include the optional attributes, e.g. “FDO_Genre_Ref” or “FDO_Type_Ref” which are currently under discussion. The given profiles allow additional attributes. In order to make community specific attributes mandatory a new profile based on the given example needs to be defined, but always adding the profile reference as a mandatory attribute. 2.6 Data Type Registry for PID Records The Data Type Registry (DTR) enables the registration of PID-BasicInfoType and PID-InfoTypes based on it. Those types can then be used as building blocks for Profiles. In the initial work those profiles were called
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 15 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. KernelInformationProfiles. In the FAIRCORE4EOSC DTR instance20 created in WP4 during the course of the project the simpler term Profile was kept. A decision whether it is used for Kernel Information or extended information can be made later by the user. In the original DTR instance “dtr-test”21 a total of 58 KernelInformationProfiles are registered. In the new instance a total of 30 profiles are registered (as of end of February 2025). The DTR is based on Cordra22 and provides the possibilities to design schemas for types and register types based on those schemas. CORDRA uses JSON objects to describe type schemas and instances of types. Each schema and type is assigned with a persistent identifier. For the FAIRCORE4EOSC instance the ePIC namespace is used. The type definitions are human readable and contain metadata, e.g. about provenance, and the restrictions (e.g. enum, string). The DTR also creates and provides JSON schemas23 to allow validation of (meta-)data sets for specific type definitions. JSON objects are common and widely used technology which makes Cordra the best option for maintaining profiles. 2.6.1 DTR Implementation examples DiSSCo, the Distributed System of Scientific Collections is a new world-class Research Infrastructure for Natural Science Collections. DiSSCo aims to provide FAIR Data in the form of Digital Specimen, a particular type of digital object. The schema of the Digital Specimen is registered in the FAIRCORE4EOSC DTR: ●https://dtr-test.pidconsortium.net/#objects/21.T11148/d8de0819e144e4096645 In [R8] it is explained how to incorporate the results of the RDA working groups to DiSSCo infrastructure. They argue for the necessity of Kernel Information for efficient processing of large data sets. However, the profile does not specify which attributes are to be considered mandatory or optional. Thus, a clear distinction between Kernel Information and optional attributes is not possible. But it should be noted that the PID records of the PIDs on Digital Specimen actually have a profile specification (c.f. Figure 6) and thus follow the requirements of the FDO specifications. 20 https://typeregistry.lab.pidconsortium.net/ 21 All entries in the mentioned DTR services are persistent despite the domain name. 22 https://www.cordra.org/ 23 https://json-schema.org/
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 16 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. Figure 6: PID of a Digital Specimen, using the DiSSCover Service (https://disscover.dissco.eu/search), and resolving the PID to the PID Record using the global resolver (http://hdl.handle.net) B2Inst24 is a service to register and publish (scientific) instruments. The first instance25 is hosted at GWDG as one of the EUDAT26 services. The schema used in the service goes back to the works of the RDA Persistent Identification of Instruments Working Group27. The schema is registered and published in the FAIRCORE4EOSC DTR: ●https://typeregistry.lab.pidconsortium.net/#objects/21.T11969/e16dc712334588cd4831 As the service by default would not write into the PID Record (the Handle Record) an adapter is required to achieve PIDs conforming to the FDO specifications. The adapter would add the necessary attributes to the PID Record following a specific profile. As result of the FDO One project an adapter28 for the B2Share service was developed. As B2Inst is based on the same software stack as B2Share, the adapter could be reused for B2Inst. AVefi is an external community onboarded to FAIRCORE4EOSC. The developments are taking place in the DFG-funded project AV-efi29. The aim is to create a network system for film-holding institutions in Germany based on standardised film identifiers. The standardisation is achieved by the development of a schema using the LinkML framework and registration of the schema in the FAIRCORE4EOSC Data Type Registry. The special approach in this project is to store all metadata for a work30 in the PID record. A work has no references to its 24 The EUDAT service (https://eudat.eu) naming convention starts with B2 and the mentioned service focuses on Instruments. 25 https://b2inst.gwdg.de/ 26 https://eudat.eu/ 27 https://www.rd-alliance.org/groups/persistent-identification-instruments-wg/members/all-members/ 28 https://gitlab.com/fairdo/fdo-doip-b2share-adapter 29 https://www.av-efi.net/ 30“A moving image Work comprises both the intellectual or artistic content and the process of realisation in a cinematographic medium”, The FIAF Moving Image Cataloguing Manual,
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 17 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. manifestations, thus it is a very specific configuration type with respect to the FDO specifications. The project is not yet completed and a final decision on the configuration type has not yet been made. But the schema is defined in the Data Type Registry and actively used in the minting of new efi(s) (name of the PID based on Handles): ●https://typeregistry.lab.pidconsortium.net/#objects/21.T11969/e15fa59d1733320642b6 2.7 Comparison of Approaches to Kernel Information Provision In broad terms, there are two distinct approaches to implementation of Kernel Information Profiles, each with advantages and disadvantages31. 1. Detailed Metadata at the point of resolution (e.g. full FDO implementation): this will usually be the Authority or Multi-Primary Administrators (MPA) in the PID Stack. a. Advantages i. There are fewer redirects required to obtain the metadata if PID resolution is the primary means of access. However, a sizable proportion of redirects to scholarly content originates from web search and direct links to content or metadata, and for these, the benefit does not apply. ii. Obtaining the metadata is highly probable, if not guaranteed. iii. All the custom metadata records for a PID Service/ Stack can be found via a single endpoint, and the process/ workflow is standardised. b. Disadvantages i. It requires extra work for the Managers and Providers to maintain synchronisation between their versions and those of the redirection/ resolution endpoint. ii. The point of resolution will have to store a much larger volume of data, and while metadata records are small, some PID Stacks involve hundreds of millions (or even billions) of records and might encounter scalability issues. iii. Decentralised management and control of the metadata content and the means to reference it provides protection to the community in respect of monopolistic practices and abuse of power (provided minimum sustainability provisions are in place). iv. The authoritative version should be hosted where curation and quality assurance is performed (i.e. the Manager and/ or Provider). Authorities have little or no capacity or ability to curate content. v. It is highly unlikely that all Authorities and MPAs could be persuaded to host additional metadata, maintaining the diversity of workflows for developers and users of PIDs. vi. Custom metadata schemata vary considerably between Managers and Providers for sound reasons (format and domain diversity). Because of this diversity, representing such metadata in a single schema will be very challenging. https://www.fiafnet.org/images/tinyUpload/E-Resources/Commission-And-PIP-Resources/CDCresources/20160920%20Fiaf%20Manual-WEB.pdf 31 Adapted and extended from Hugo, W., Van de Sompel, H., & Hakala, J. (2025). The PID Landscape - a Technical View (1.1). Zenodo. https://doi.org/10.5281/zenodo.14881287
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 18 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. vii. The centralised resolution points are unlikely to add custom value to metadata (for example acting in domainor community-specific ways on machine-readable links in the metadata, or providing community-specific links to value-addition tools). 2. Custom Metadata managed at the target of resolution (e.g. Signposting): This will usually be a landing page, metadata record or resource curated by a Manager. a. Advantages i. Having Managers provide authoritative metadata has many advantages, including partial or full mitigation of the disadvantages listed above. ii. The relationship between owners (depositors) and Managers is important and trust (through quality assurance, curation, and guidance) is cemented by designating Managers as custodians of authoritative metadata. iii. Specialisation in terms of custom metadata and metadata guidance is very difficult to provide at scale. Managers usually deal with smaller and well-defined collections. b. Disadvantages i. Aggregation of metadata in the same domain, if spread over multiple Managers, requires harvesting to a central catalogue or federated search that is costly and complex, even though it is common practice. Best practice in our view would be to focus on the Manager and/ or Provider for authoritative metadata in all cases where this is possible and applicable, and extend Kernel Metadata for administrative content and machine-actionable support that is currently not available. This machine-actionable support includes, as a minimum, the automated discovery of a Kernel Information Profile, as presented in this report. 3Governance of Kernel Information Profile The governance of Kernel Information Profiles should be closely coupled to the governance of similar knowledge bases, e.g., the types in the DTR or the crosswalks in MSCR. All these definitions and entries in centralized registries will be created and maintained by domain experts. This includes experts in information design, metadata, or similar topics. 3.1 Elements and tasks of the governance In order to implement a governance model, the activities related to profiles need to be described. 3.1.1 The profile registry Based on the examples given in chapter 2, not all profiles or schemas are registered in a central registry. However, as required by the FDO specifications32 a registry should be used. Therefore, it is to be decided whether one (or small number of) registry should be used, or to establish a federation of registries. The federation would require technical and semantic interoperability to allow for harvesting, search and use of profiles registered across the distributed registries. As a result of the FAIRCORE4EOSC project a single data type registry is used, which serves the needs of the communities on-boarded or being part of the project. A federation of registries based on the same technology (Cordra) is under investigation. After the project lifetime a governance is required to steer to development of such registry services. 32 https://fairdo.org/specifications
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 19 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 3.1.2. The profiles The possible usage scenarios for profiles are very diverse. These can be very simple and generic profiles to enable machine processing, or detailed descriptions of referenced (digital) objects, services or operations. These profiles are discussed and created by different people or communities. This means that there must be established processes within the communities to reach a consensus. The governance of a profile registry should then be harmonized between different communities. Elements such as ownership and responsibility for maintaining Kernel Information Profiles need to be addressed. In more detail: ●Who can create and update profiles? ●Who can create and update the type schemas used for profiles? ●Who decides on the policy that profiles should not be deleted? ●Do we need to have prototypes and endorsed profiles? If so, who will endorse those profiles? ●Versioning as part of the profile registry. But who can create new versions? ●How should co-authors be referenced as creating a profile is a publication process? ●Who should decide on technology changes? These questions require the integration of the profile registry or registries into a larger infrastructure where an Authentication and Authorization Infrastructure (AAI) is available. Group Management and detailed rights management for the management of the profile would then be implemented in Identity and Access Management (IAM) system or realised by the profile registry directly. 3.1.3 Existing Structures Before setting up a new governance model or building new structures existing structures or initiatives should be evaluated. Examples of those are: ●Research Data Alliance: the RDA is focusing on creating output by working groups. Recommendations of those groups could be implemented as policies for the management of profiles. ●The FDO Forum is one example with a given governance structure to create policies and specifications for FAIR Digital Objects. Working groups within the FDO Forum can focus on different aspects, e.g., configuration types, profiles, operations, or services. Similar to the RDA the implementation of policies is not done within the FDO Forum activities, but rather with a member. ●EOSC: A registry service for profiles can benefit from an integration with the European Open Science Cloud. Especially the use of a specific Authentication and Authorization Infrastructure (AAI) would allow technical control of the creation and administration of profiles. Groups representing communities or domain experts could be created for this purpose. ●NFDI: The German National Research Data Infrastructure is an example for the implementation of national policies. Within the base service project PID4NFDI the use of PIDs, and therefore also Kernel Information, is promoted and service established. On the national level the governance of KIPs could be realised by a base service of the NFDI. 4 Best Practices and Conclusion The complexity of the scientific landscape is reflected in the use of PIDs. The different PID systems have very different properties regarding the metadata associated with a PID or the referenced object. As shown in the examples various approaches are already implemented.
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 20 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. 4.1 Governance-Related Recommendations Based on the existing services and communities following aspects would enhance the usage of Kernel Information and the implementation of profiles: ●Small number of profile registries: In order to maintain technical and organisational control of the development of profiles the number of registries should be minimized. These registries need to be promoted at all levels in the scientific domain. ●Interoperability between registries: Established or new profile registries should be technically interoperable. The use of profiles from several registries within one use-case should be possible. Interoperable registries also allow for cross-linking and would support a federated search. ●FDO specific and generic profiles are developed and specified, registered and published by the FDO Forum. These standards should be implemented by the profile registries. ●Open Access to the profile registries: Communities do have their own governance model and can create domain specific profiles. No barriers should exist to allow communities to register profiles. ●Standardise community profiles: Policies need to be defined to describe the requirements for domain specific profiles. This work could be done in the FDO Forum. 4.2 Optimal Metadata Management It is clear from the examples and discussion in the report that metadata could be managed and made available at several points in the ecosystem of actors involved in the provision of a PID service. To account for the variation in current implementation, and also promote uniformity and machine actionability, the following best practices are recommended: ●Custom metadata, available as the primary resolution target of the PID, and maintained by the Manager, is the authoritative metadata record. This record is available as a human readable and as a machine-readable version of the metadata through simple content negotiation (e.g. HTML and JSON versions). ●Custom metadata optionally offers Signposting information in its HEAD response, circumventing the need to retrieve the entire metadata record if only specific machine-actionable metadata is sought. ●The Signposting linkset must include a pointer to the Kernel Information Profile. ●Kernel metadata records must include a pointer to the Kernel Information Profile, and optionally can reference the Signposting Linkset. ●Similarly, in cases where an independent identifier metadata record is offered, it must include a pointer to the Kernel Information Profile, and optionally can reference the Signposting Linkset. 4.3 Implementation of Kernel Information Profiles A minimum set of elements should be defined for Kernel Information Profiles across all PID services, to which each PID service can add additional mandatory elements. These same mandatory elements should be implemented in Signposting Linksets. The use of profiles in the context of the FDO configuration types could be applied to combine approaches. The process can be well illustrated with the help of a signposting example. ●A Handle based PID could point to web resource which provides Signposting functionality ●The PID would therefore have an URL/URI attribute of the web resource ●Adding a profile to this PID would then add the following attributes the PID Record ○a profile reference specific for Signposting ○an attribute stating that the web resource referenced with the URL/URI will provide Signposting
D5.1 Report on Best Practices and Governance of Kernel Information Profiles 21 FAIRCORE4EOSC has received funding from the EU’s Horizon Europe research and innovation programme under Grant Agreement no. 101057264. ○an attribute stating the Signposting schema that is used by the web resource Another example could be to combine PIDs with RO-Crates to make them more findable. The same approach, as described for the Signposting example, could be taken to specify in the PID Record that the referenced object is a RO-Crate. Therefore, the configuration profiles provide a powerful tool to build machine actionable PIDs. 4.4 Harmonisation of Features and Performance Measurement Not all PID services exhibit a common, interoperable set of features that are required to assess the quality of service from an end user perspective. ●The most important of these involve the two critical characteristics of PIDs: persistence and resolvability. For neither of these, consistent and verified public evidence is available across the PID services. ● Availability of PID services (e.g. resolution endpoints) are also an important service metric. This metric is not consistently available from PID services. ●By implication, implementation or not of Kernel Information Profiles - following either an FDO-focused or Signposting approach - need to be transparently visible for each PID service. At present, there are no community benchmarks in respect of performance expected from the suppliers of persistent identifiers, and a mechanism to determine and maintain these is clearly required. The FAIRCORE4EOSC project, during implementation of the CAT, has proposed a set of default benchmarks that require community ownership and validation, but these need some form of governance and community involvement to be credible. Identifier Metadata: identifier metadata (i.e. always available through a request to the authoritative registry of a PID service) must include administrative metadata that is used inter alia to determine important metrics in respect of PID services. At a minimum, this information must include the date of first registration of a PID, and the original target for resolution. These are needed for measurement of resolvability and persistence. Actions: extension and harmonisation of PID service actions that can be requested from individual PIDs (e.g. consistent content negotiation implementation) and from registries (e.g. offering random samples of PIDs) are required. Poor practices to be avoided: There are several examples of these, including ●Resolution Variance: there are documented cases of resolution success depending on access methods (for example, authenticated users resolving a PID successfully but non-authenticated users do not). There may be cases where this variance is desirable, but for resolution to a metadata landing page, it is not good practice. ●Absence of Continuity Arrangements: In a few documented cases, PID services ceased to exist, or were interrupted for significant periods of time due to poor or nonexistent continuity arrangements. These persistent identifiers are either now not resolvable, or the future sustainability of the current services are in doubt. We will attempt to assess the extent of continuity planning in future releases of the Knowledge Base. More recommendations can be found in Table 1. Recommendation - What? Stakeholder affected - For whom? Problem statement - Why? Including sensitive metadata in the kernel met adata is to be Applies to the PID Manager and not to the PID Service Provider. Having no sensitive data in the kernel removes the need for