scieee AI-readable full text Open interactive document viewer

D6.3 - Data Management Plan

Seven Past Nine

Abstract

This deliverable report describes the creation of the BIO-SUSHY Research Output Management Plan (ROMP) as an extension of a Data Management Plan (DMP) and contains the first version as Annex I. With the success of the FAIR principles to foster making more and more data available for re-use, it became obvious that the high-quality knowledge management approaches should also be applied to other research outputs to increase their re-usability. To spearhead such a comprehensive approach, BIO-SUSHY has started to go beyond the DMP, mandatory for all projects of the Horizon Europe Framework Programme, by covering all research outcomes including but not limited to data, protocols, models, software, surveys and responses, reports and publications in the ROMP and specifying common management recommendations and guidelines across all research outputs as well as specific approaches and tooling for individual types. To be able to cover the requirements of all output types, BIO-SUSHY, in collaboration with the MACRAMÉ project (EU Horizon Europe research and innovation programme, grant agreement No. 101092686), performed an evaluation of existing DMP templates proposed in the DMPonline, DS Wizard and ARGOS tools. None of them is providing all functionality and flexibility needed by BIO-SUSHY, especially with respect to covering computational methods and research software. Thus, the BIO-SUSHY ROMP was newly designed by starting from the structure of the ARGOS tools, transferring sections of the other tools to strengthen specific aspects and adding additional sections based on the FAIR for Research Software (FAIR4RS), TRUST and CARE data guidance principles.

Full text

Funded by the European Union under the Grant Agreement 101091464. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Health and Digital Executive Agency (HaDEA). Neither the European Union nor the granting authority can be held responsible for them. Sustainable surface protection by glass-like hybrid and biomaterials coatings Deliverable D.6.3 Data Management Plan Deliverable Information Responsible partner: 7P9-SI Work package No and Title: WP6 - Data Management Plan Contributing partner(s): 7P9-DE Dissemination level1: PU Type: R Due date: 30/06/2023 Submission date: 30/06/2023 Version: V2 1 PU = PUBLIC fully open ((warning) automatically posted online on the Project Results platforms) SEN = Sensitive — limited under the conditions of the Grant Agreement EUCl = EU classified under Decision 2015/444 Ref. Ares(2023)4548311 - 30/06/2023 2 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan Project Profile Programme Horizon Europe Call HORIZON-CL4-2022-RESILIENCE-01 Topic HORIZON-CL4-2022-RESILIENCE-01-23: Safe and sustainable by design chemicals and materials (RIA) Number 101091464 Acronym BIO-SUSHY Name Sustainable surface protection by glass-like hybrid and biomaterials coatings Start Date 1 January 2023 Duration 48 months Type of action HORIZON Research and Innovation Actions Granting authority European Health and Digital Executive Agency Project Coordinator MATERIA NOVA Document History Version Date Entity Remarks V1 8 June 2023 7P9-SI, 7P9-DE First draft V2 26 June 2023 MANO, SIK, ZSI, AXIA, RESCOLL, CNR, PQSAR, WOODK+ Partner input 3 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan Publishable Summary This deliverable report describes the creation of the BIO-SUSHY Research Output Management Plan (ROMP) as an extension of a Data Management Plan (DMP) and contains the first version as Annex I. With the success of the FAIR principles to foster making more and more data available for re-use, it became obvious that the high-quality knowledge management approaches should also be applied to other research outputs to increase their re-usability. To spearhead such a comprehensive approach, BIO-SUSHY has started to go beyond the DMP, mandatory for all projects of the Horizon Europe Framework Programme, by covering all research outcomes including but not limited to data, protocols, models, software, surveys and responses, reports and publications in the ROMP and specifying common management recommendations and guidelines across all research outputs as well as specific approaches and tooling for individual types. To be able to cover the requirements of all output types, BIO-SUSHY, in collaboration with the MACRAMÉ project (EU Horizon Europe research and innovation programme, grant agreement No. 101092686), performed an evaluation of existing DMP templates proposed in the DMPonline, DS Wizard and ARGOS tools. None of them is providing all functionality and flexibility needed by BIO-SUSHY, especially with respect to covering computational methods and research software. Thus, the BIO-SUSHY ROMP was newly designed by starting from the structure of the ARGOS tools, transferring sections of the other tools to strengthen specific aspects and adding additional sections based on the FAIR for Research Software (FAIR4RS), TRUST and CARE data guidance principles. 4 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan Table of Contents Publishable Summary .................................................................................................................................. 3 List of Figures ................................................................................................................................................ 5 Table of Abbreviations ................................................................................................................................. 6 1. Objectives .............................................................................................................................................. 8 2. Creation of the first version of the Research Output Management Plan (ROMP) ........................... 8 3. BIO-SUSHY data management infrastructure ................................................................................. 10 4. Conclusions ........................................................................................................................................ 11 5. References .......................................................................................................................................... 12 6. Annexes .............................................................................................................................................. 13 Annex I: Initial Version of the BIO-SUSHY Research Output Management Plan (ROMP) ................ 13 5 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan List of Figures Figure 1: Data management concept in which all data from experiment, computational approaches and public resources is integrated to build the data lake used as input for data-driven modelling to guide the SSbD decisions for the development of the new coatings as well as to prepare upload to public databases) ...................................................................................................................................... 11 6 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan Table of Abbreviations Abbreviation Definition EU European Union EC European Commission H2020 Horizon 2020 research and innovation programme HE Horizon Europe research and innovation programme WP Work Package PFAS Perand polyfluorinated alkyl substance DMP Data Management Plan ROMP Research Output Management Plan RO Research Output TBD To Be Defined FAIR Findable, Accessible, Interoperable and Re-usable (data guidance principles) FAIR4RS FAIR for Research Software FIP FAIR Implementation Profile FER FAIR Enabling Resource FSR FAIR Supporting Resource TRUST Transparency, Responsibility, User focus, Sustainability, and Technology (data guidance principles) CARE Collective benefit, Authority to control, Responsibility, and Ethics (data guidance principles) GDPR EU’s General Data Protection Regulation DOI Digital Object Identifier ORCID Open Researcher and Contributor ID ROR Research Organisation Registry (identifier) ERM European Registry of Materials (identifier) CAS Chemical Abstract Service Registry Number IUPAC International Union of Pure and Applied Chemistry InChI IUPAC International Chemical Identifier ELIXIR European life sciences infrastructure 7 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan TeSS ELIXIR Training eSupport System SOP Standard Operating Procedure SSbD Safe-and-Sustainable-by-Design MODA MOdelling DAta reporting format CHADA CHaracterisation DAta reporting format QSAR Quantitative Structure Activity Relationship QMRF QSAR Model Reporting Format QPRF QSAR Prediction Reporting Format API Application Programming Interface EOSC European Open Science Cloud OECD Organisation for Economic Co-operation and Development VAMAS Versailles Project on Advanced Materials and Standards ISO International Organization for Standardization CEN European Committee for Standardization ECHA European Chemicals Agency 8 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 1. Objectives As a part of its open science philosophy, BIO-SUSHY is implementing high-quality knowledge and data management with the goals to facilitate, harmonise and accelerate information sharing within the project and, at the same time, prepare all BIO-SUSHY research output for open and FAIR (findable, accessible, interoperable, and re-usable) release. One important tool to organise data collection, documentation, preservation and preparation for sharing is the Data Management Plan (DMP). It is the central guidance document for the consortium specifying, first, the types of data to be generated within the project and what data management concepts, approaches and tools are recommended and, in subsequent versions, is providing more and more information on the data produced (data formats, size), applied (meta)data standards, annotations and solutions for long-term storage and indexing. Due to this central role, the DMP is now mandatory for projects funded under Horizon Europe. However, BIO-SUSHY is going even one step further. Already during the proposal writing, the consortium agreed that it will apply the high-quality management standards not only to data but to all research outputs and will document these in the Research Output Management Plan (ROMP), which integrates and extends the DMP. This was seen as necessary to be able to clearly document all steps in the development, production and safety and sustainability evaluation of the new coatings. Since many novel and non-standard methods are used in BIO-SUSHY, detailed and complete documentation of the design process, procedures (method descriptions, protocols, standard operating procedures), processing, analysis and extracted evidence and conclusions have to be provided as (meta)data to understand and evaluate all potential influences on the results. Additionally, the strong integration of physics-based and data-driven modelling and simulations extends the types of research outputs to be covered to computational models, workflows and software. To develop the best way to guide the consortium in the task of documenting all these research outputs and selecting appropriate approaches, BIO-SUSHY was collaborating with the MACRAMÉ project to create the structure of the ROMP. Besides the complete adoption of the open science and FAIR guidance principles1, addressing requirements of specific other research outputs and integrating aspects such as quality assurance, responsibility, trustworthiness and ethics were major objectives of the ROMP development. To achieve these, data management was completely aligned with BIO-SUSHY’s overall quality and risk management (see Deliverable D6.2) and the FAIR for Research Software (FAIR4RS)2, TRUST3 and CARE4 principles, besides the newest recommendations from the WorldFAIR project and the GO FAIR AvancedNano Implementation Network, were considered. 2. Creation of the first version of the Research Output Management Plan (ROMP) As introduced above, details of and guidance on the adopted management principles, FAIRification approaches, data security and ethics considerations as agreed on by the consortium are documented in the Research Output Management Plan (ROMP) as an extended version of a Data Management 9 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan Plan (DMP) covering not only data but also all other outputs including but not limited to protocols, models, software, surveys and responses, reports and publications. It is designed as a living document being updated constantly during the runtime of the project and is provided in its first version as Annex I to this report. Besides the initial version presented here, we anticipate that at least two additional stable versions of the plan (intermediate and final) will be compiled and publicly shared to accommodate reporting of the data generation at the midterm and end of the project but also to be able to adapt to the fast-changing field of Open and FAIR data in Europe. According to the EC recommendation to make data findable, accessible, interoperable and re-usable (FAIR), a Data Management Plan and, thus, also its extension, the ROMP, includes information on the following details: 1. Handling of research outputs (RO) during & after the end of the Project; 2. What RO are collected, processed and/or generated; 3. Which methodology and standards are applied; 4. Whether RO are shared/made Open Access; 5. How RO are curated and preserved (including after project end). The initial version of the BIO-SUSHY ROMP mainly documents decisions made by the consortium and guidelines to be implemented by the data providers and the data managers in collaboration with and under the supervision of the BIO-SUSHY data shepherd. This will subsequently be complemented throughout the project with specific information on the research outputs generated and resources re-used in the project as well as updates and refinements of the used FAIR Enabling and Supporting Resources (FERs and FSRs) and improvements in harmonisation and interoperability. Thus, the ROMP has, as already stated above, to be understood as a living document, in which changes are clearly tracked to show additions as well as decisions, which had to be changed, revised or replaced as a reaction to emerging new standards, tools and guidelines. In cases where the content of the document needs more extensive adaptations for a specific type of research output, a specific section for this type will be added or even a specific ROMP for this type will be created and referred to in this general ROMP covering the project globally. For creating the ROMP, three different online DMP tools were evaluated according to criteria that evaluate if they covered all necessary aspects and provided the appropriate structure needed to define the management approach and the generalisation to all research output in close collaboration with the MACRAMÉ project. The tools evaluated are: 1. DMPonline from the Digital Curation Centre, 2. DS Wizard and 3. ARGOS. Results of this evaluation have already been described in the public deliverable D3.1 of the MACRAMÉ project, which is currently still under review by the European Commission. Therefore, we will reproduce them here again and want to stress the importance covering computational models, workflows, and software as specific research outputs in the ROMP. This was a specific requirement of BIO-SUSHY because of the much more central role of physics-based and data-driven modelling and simulation approaches compared to MACRAMÉ. 16 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 1.1.2 Is it physical or digital? This ROMP covers the digital objects generated or re-used in BIO-SUSHY. If required, physical objects produced by the project and managed using specific centralised and standardised resources (e.g. reference material repositories, biobanks) will be addressed in a specific ROMP on physical objects aligned to the ROMP on digital objects presented here as much as possible. 1.1.3 Are you generating or re-using it? Digital research objects are generated as part of the BIO-SUSHY coating developments and SSbD evaluation. These are complemented by data and accompanying metadata mined and curated from public databases and scientific literature. Additionally, existing computational models, workflows and software will be re-used for simulating molecular interactions and predicting physicochemical characteristics, functionality, and human health and environmental safety. 1.1.4 What is the type of the described research output? As described above, this ROMP covers all research output from BIO-SUSHY. It can be grouped into the following general categories (more specific categories will be added during the continuous updating of this living document), which is used (including numbering) for the further structuring of the answers in this ROMP: 1. Study designs 2. Method specifications a. experimental b. computational 3. Protocols / SOPs 4. Computational models 5. Software and computational workflows 6. Data 7. Surveys and survey responses 8. Guidelines 9. Reports and knowledge collections 10. Training materials a. Manuals / tutorials b. Videos c. Handbooks / reference material 11. Publications a. Dissemination material b. White papers c. Peer-reviewed papers 1.1.5 What is its format? The variety of research output types listed above clearly shows that it is not possible to cover all of them in one exchange format. Additionally, different BIO-SUSHY partners entered the project with 17 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan very different knowledge management approaches applied in their internal processes and workflows, which have to be considered when creating the recommendations for project-wide data management and public sharing. Therefore, this initial ROMP lists all data formats currently in use and uses the (meta)data reported in these formats to develop recommendations for improvement of the data management practices at specific partners, select appropriate FAIRification tools for each setting, and develop (semi-)automated approaches to translate internal formats into harmonised and interoperable data documentation formats aligned with community and regulatory standards. 1. Study designs : instance maps 2. Method specifications : text documents, data templates (MODA, CHADA) 3. Protocols / SOPs : text documents, electronic lab notebooks 4. Computational models: data templates (QMRF, QPRF, MODA) 5. Software and workflows : data templates (QMRF, QPRF, MODA), notebooks (jupyter, colab), API definitions. 6. Data : customised spreadsheets, data templates (nanoFASE, eNanoMapper, MODA, CHADA), proprietary formats, data serialisation formats (json, yaml) 7. Surveys and survey responses: text documents, spreadsheets, Google forms 8. Guidelines : text documents 9. Reports : text documents 10. Training materials : text documents, slides, videos 11. Publications : text documents 1.1.6 What is its expected size? Limited information on expected size is available at the current state. Similar projects had relatively low requirements for data storage capacities. However, the distributed data storage approach adopted by BIO-SUSHY will also be able to cover much higher demands if necessary. This can be achieved by e.g. storing large raw data files only locally at the data providers or using data storing solutions specialised for a specific data type (e.g. omics) or for general storage of big data as e.g. provided by the European Open Science Cloud (EOSC). 1.1.7 Why are you collecting/generating or re-using it? Data is primarily generated and collected for re-use to satisfy the data needs of the BIO-SUSHY coating development in the three use cases including functionality and SSbD evaluation. The experimental and computational datasets will then be provided to guide defining SSbD criteria for coatings and, subsequently, as input for the standardisation and regulatory validation project initiated at the relevant bodies (e.g. OECD, VAMAS, ISO, CEN). 1.1.8 What is its origin / provenance? Each research output, both generated by BIO-SUSHY and collected from third parties, will specifically document its origin and provenance as part of the set of standardised metadata. 18 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 1.1.9 To whom might it be useful (“data utility”)? Primary users of the data will be the BIO-SUSHY consortium partners performing the development and optimisation as well as SSbD evaluation on the use case coatings. Secondary users will then be developers of SSbD criteria for bio-based and hybrid coating materials as well as labs (academia and industry) interested in performing Safe-and-Sustainable-by-Design (SSbD) material development. The data will then also be made available for starting the standardisation and regulatory validation projects for integrating specific SSbD criteria for coatings supporting registration of such new materials. 2 Links Between Outputs 2.1 Publications 2.1.1 Does the described output support any scientific publication? Due to the recent project start, BIO-SUSHY has not yet produced scientific publications supported by the newly generated research outputs. For data collected and mined from databases and literature, the corresponding references are stored as metadata providing clear provenance trails for the extracted information. 2.1.2 Is there a data availability statement provided along with the publication? Data availability statements will be provided in all future publications resulting from the BIO-SUSHY project. These will include unique, persistent identifiers (DOIs or data-specific identifiers), licensing information, clear provenance trails, and references to the data models used if applicable. 2.2 Qualified references to other objects 2.2.1 Does the described output support any other research output? BIO-SUSHY has adopted the approach of providing the different research output types as individual resources. Instead of creating complex (meta)data templates storing information on the methods, protocols and the generated raw and processed data, as e.g. implemented in the NIKC-NanoFASE templates or the MODA/CHADA system, all these components are managed using individual tools customised and optimised for each type (e.g. electronic lab notebooks, method-specific data formats and databases, version control systems). However, this results in the fact that many research outputs are needed to provide all information a specific piece of evidence or conclusion is based on. Managing all research output using the harmonised approach described in this ROMP has the advantage that all different types can be handled individually and, at the same time, outputs supporting each other can be linked together giving full access to all information relevant for e.g. a specific method or a BIO-SUSHY use case. For example, data include references to the methods and protocols used to generate it as metadata and methods can list all dataset used for testing and validation. 19 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 2.2.2 Is the research output integrated into a system of qualified references crosslinking outputs? The concept of individual resources has the advantage of access to optimal data management tools for each type. However, such a distributed system is putting more demands on keeping track of all the resources supporting each other and consistency of the resource references since losing links between the resources would destroy data completeness, understandability, interpretability, and, thus, trust and ultimately re-usability. BIO-SUSHY has established an advanced system for research output cross-linking composed of two services, the BIO-SUSHY Registry and the instance map tool. The first assigns a unique, even if only internal identifier to each research output. Additionally, basic metadata on data provenance and accessibility, references to use cases and project partners are provided. The second service is then providing ways to build further cross-links between the research outputs based on the unique identifiers, e.g. linking protocols to the generated data, and to visualise these as instance maps representing life-cycle stages of the materials and the experimental workflow of production and assessment. For public sharing and long-term storage, all relevant research, e.g. supporting a scientific publication, can be extracted from the internal management solution by: ● Replacing the internal identifiers automatically with global and persistent identifiers from authorities like DataCite/Zendo (DOIs), European Registry for Materials (ERM), ORCID, Research Organisation Registry (ROR) or public data/protocol repositories; ● Packing all resources specified in one instance map into a data package following the frictionless data or RO-Crate specifications either as (meta)data files or as links to other public resources; and ● Publicly sharing of the data packages in FAIR data storing solutions. 2.3 Pre-existing data 2.3.1 Are you using any pre-existing research output? Pre-existing research outputs will be re-used in the form of existing method descriptions, protocols/SOPs, computational models, workflows, and software as well as data mined and curated from public databases and scientific literature. 2.3.2 Is the pre-existing research output handled according to this ROMP including additional FAIRification if necessary? If not, provide links to DMP describing the treatment of these pre-existing research outputs. The pre-existing data will be handled in the same way as newly generated data as described in this ROMP. Additional data management steps will include integration into the BIO-SUSHY Registry, provision of data provenance trails, and harmonisation of data models and transfer formats. Additionally, FAIRification steps will be described here if they become necessary for specific preexisting resources. 20 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 3 Quality control, FAIR Practices and openness 3.1 Making research outputs findable, including provisions for metadata 3.1.1. What type(s) of persistent identifier(s) are used for the described research output? BIO-SUSHY is using an internal set of identifiers, which are assigned by the BIO-SUSHY Registry or, in specific cases, the project’s Google Shared Folder, are unique within the project and can be translated into globally unique and persistent identifiers at a later stage. In this way, the different (future) research outputs can be clearly identified starting from the planning phase on and global uniqueness and also indexing in the relevant identifier services is achieved for the final versions of the outputs meant for public sharing and long-term storage. Globally unique, persistent identifiers are currently available and integrated into community standards for: 1. Sampling plans : DOI (potentially) 2. Study designs : DOI (potentially) 3. Method specifications : DOI (Zenodo) 4. Protocols / SOPs : DOI (Zenodo, DataCite) 5. Computational models, software and workflows : GitHub 6. Data : DOI (Zenodo) a. Materials: ERM, NInChI b. Chemicals : CAS, InChI, (CAS) c. Providers : ORCID d. Institutions : ROR e. Projects : DOI (EU) 7. Surveys and survey responses: DOI (potentially) 8. Guidelines, reports, training materials : DOI, TeSS 9. Publications : DOI (from publisher) 3.1.1.1 Are components of the research output representing levels of granularity (software modules, experimental steps, materials) assigned distinct identifiers? Study designs, protocols, materials/chemicals, models, workflows, software as well as surveys and their results, reports and publications are referred to by individual identifiers in the internal system. These can thus also be mapped to individual distinct global identifiers later. The system also encourages splitting the experimental procedure into multiple protocols, e.g. for sample preparation, measurement, and processing, all with their own identifiers. Splitting into even smaller parts (individual protocol steps) is in principle also possible but currently not envisioned. This decision will be periodically reviewed. 21 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 3.1.1.2 Are different versions of the method descriptions, protocols, software assigned distinct identifiers? Development of all research outputs, which might exist in different versions, are managed in solutions with automatic version control (GitHub, Google documents). Stable versions used to generate results for specific studies will be specifically marked (named versions) and provided with distinct internal identifiers first and then global identifiers if they support specific public research outputs. 3.1.2 Will you provide metadata for the described research output? What metadata will be created? All research outputs will be accompanied by metadata, which will be standardised and made richer over the runtime of the project. At the current stage, the following high-level metadata fields are mandatory for all research output referred to as resources in the BIO-SUSHY Registry: 1. Unique internal identifier 2. Resource name 3. Type of resource (e.g. Material, Test method, protocol) 4. Status (e.g. Scheduled, in preparation, under internal review) 5. Resource link (access path) 6. Short description 7. Contributors and their roles 8. Licence 3.1.2.1 What disciplinary or general standards will be followed? The final metadata schema of BIO-SUSHY will be composed of archetypes, describing specific aspects like contributors, publications, (meta)data schema, and method-specific metadata. These archetypes will be constructed following existing standards. For example, contributors, institutions and publications will use the DataCite and Dublin Core specifications. For disciplinary components, standards are currently under development or revision (MODA/CHADA, eNanoMapper-based templates) and will be integrated when available. 3.1.2.2 In case metadata standards do not exist in your discipline, please outline what type of metadata will be created and how. The new production processes and test methods developed in BIO-SUSHY will require methodspecific metadata to fulfil minimum reporting requirements. These are currently under development based on the expertise from earlier projects and the first results provided by the BIO-SUSHY partners. They will be based on existing minimal reporting guidelines and regulatory requirements but will provide the flexibility to report metadata specific to the process/method. To guarantee interoperability to the highest possible extent, the metadata models/schemas implemented in the reporting guidelines will be provided as high-level metadata to the research output (data but also structured information from e.g. protocols, study designs). 22 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 3.1.3 Will search keywords be provided in the metadata to optimise the possibility for discovery and then potential re-use? Provision of structured (meta)data with well-defined (meta)data schemas can be used for data discovery within BIO-SUSHY and potential re-use using advanced searching/browsing features. These will heavily rely on the unique identifiers for materials, methods and protocols as well as semantic annotated metadata. Additional full text search will provide additional means to find relevant research outputs. If this does not satisfy all needed for information discovery, additional search keywords will be provided. 3.1.4 Will metadata be offered in such a way that it can be harvested and indexed? Metadata is available in a structured format in the BIO-SUSHY Registry and by providing the (meta)data schemas as high-level metadata. Options on how to perform harvesting and indexing will be provided as part of the metadata accessible via the application programming interfaces (APIs) of the Registry (project internal) and potentially of the long-term storage solutions (external). 3.2 Making research outputs accessible 3.2.1 Repository 3.2.1.1 In which repository will the dataset / output be deposited? The research output is currently managed and stored in the internal BIO-SUSHY Registry and different data storage solutions (mainly Google Shared Drive). Options for long-term storage are currently evaluated and selected based on their fitness for the specific output and their FAIRness considering general solutions like Zenodo, open institutional repositories, and domain-specific data warehouses. 3.2.1.2 Is the selected repository a trusted source? Long-term storage solutions will be selected based on community recommendations at the point of time. Since the research output is prepared according to the high BIO-SUSHY FAIR standards already for internal data sharing, it will be ready to be publicly released after semi-automatic transformation into the format requested by the storage solution. This allows a flexible selection of solutions, which will only consider trusted sources, according to the needs of the individual research output. 3.2.1.4 Are appropriate arrangements made with the repository(ies) where the described dataset will be deposited Arrangements will be made when the selection of the long-term storage solution is finalised. 3.2.1.5 Does the repository(ies) assign research outputs with persistent identifiers? Only solutions providing persistent identifiers will be considered in the selection process. 23 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 3.2.1.7 Does the repository support versioning? Only solutions providing versioning (when required by a specific research output like software) will be considered. 3.2.2 Data This section will be complemented whenever new research outputs become available since the answers need to be specific to these outputs. Currently, only general aspects of the internal data management system (BIO-SUSHY Registry and instance maps) are given. 3.2.2.1 What is the described research output title? “BIO-SUSHY (meta)data” (Additional and more specific titles will be added during the integration of specific research outputs.) 3.2.2.2 How is the research output shared? Specify reasons for the type of sharing selected (fully open, restricted, confidential) and embargo period, if applicable, separating legal and contractual reasons from intentional restrictions. Currently, all research outputs are only internally shared via the BIO-SUSHY Registry. This is the case since none of the outputs have already matured to their final version. When this status is reached for a specific output, sharing decisions are made on a case-by-case basis with preference for fully open sharing and licences allowing re-use in most situations as outlined in the consortium agreement. 3.2.2.4 Will the research output be accessible through a free and standardised access protocol? Final, long-term access will be provided using existing web-based FAIR data solutions with free and standardised access protocols (http, ftp and specific API endpoints). 3.2.2.5 Are there any methods or tools required to access the research output? All outputs will be provided without requiring any specific methods or tools other than standard applications (e.g. pdf). However, if relevant, raw data will be provided in proprietary formats since they offer re-use in advanced, data-type-specific analysis software. 3.2.2.8 Is the described research output supported by a data/research output access committee (e.g. to evaluate/approve access requests to personal/sensitive data)? The BIO-SUSHY General Assembly will function as the research output access committee guaranteeing that no confidential data is shared unauthorised. Since no sensitive personal data is meant to be collected in BIO-SUSHY, an additional external committee is deemed non-essential. 24 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan 3.2.2.9 Please specify how the research output will be accessed during and after the project ends especially if restrictions are in use. TBD 3.2.2.10 Please specify how long after the project has ended the research output will be made accessible for. TBD on a case-by-case basis. 3.2.2.11 How will the identity of the person accessing the data be ascertained? The BIO-SUSHY Registry, instance map tool and the Google shared drive used for temporal data storage are all protected by state-of-the-art authentication and authorisation management and accessible only after logging in using a personal account. In this way, ascertaining and protocolling the identity of persons accessing research outputs is assured. 3.2.3 Metadata 3.2.3.1. Will you provide metadata even if the described research output cannot be openly shared? Metadata for all research outputs are collected independent of their final usage and planned sharing options. This will also include information on how to get access to the output and the person responsible to handle all requests (see Section 3.1.2 above). 3.2.3.2. Under which licence will metadata be provided? Specify reasons for selecting the licence separating legal and contractual reasons from intentional restrictions. Metadata will be shared under the Creative Commons Attribution 4.0 License (CC-BY-4.0) if not prohibited for legal or contractual reasons. If such cases become relevant in the future, the reasons will be specified here. 3.2.3.3. Do metadata provide information about how to access the described research output? Information on the access routes and the person responsible for handling all requests will be provided as metadata to every research output. Currently, this is managed by the BIO-SUSHY registry providing links to the resources and responsible partners. This will be replaced with information on long-term storage solutions once the research output is publicly shared. 3.2.3.4. Do metadata provide a full description of the data model used for the research output (including input and output formats) or a link to such a description? It is anticipated that all research outputs provide the underlying (meta)data model as part of their high-level data documentation. This will be established as part of the additional FAIRification 25 of 31 Sustainable surface protection by glass-like hybrid and biomaterials coatings D.6.3: Data Management Plan following the currently ongoing evaluation of the reporting and management approaches at the individual BIO-SUSHY partners. 3.2.3.5. Will metadata remain available after the dataset/output is no longer available? Indexing of the metadata in standard FAIR Supporting Resources like Zenodo is planned guaranteeing availability even after the output is no longer available. 3.3 Making data and other outputs interoperable 3.3.1 Does your (meta)data use a controlled vocabulary? Controlled vocabularies will be used whenever available. On one hand, the DataModel Ontology and/or the Information Artefact Ontology will be used for the high-level data documentation including the description and semantic annotation of the (meta)data model. On the other hand, for the low-level annotation of method-specific metadata, different ontologies like the eNM ontology, the EMMO as well as multiple chemical and biological ontologies are available. However, it is expected that these will not cover all relevant aspects and BIO-SUSHY will collaborate with other projects to increase the ontological coverage. 3.3.2 If you created the vocabulary, where can it be found? Terminology resulting from BIO-SUSHY’s ontology work will be integrated into existing ontologies and will therefore be available from the services providing these ontologies (e.g. BioPortal). 3.3.3 Have you applied a standard schema for your (meta)data? As described above, harmonisation has been started with standardisation of high-level metadata reusing (parts of) existing standards (DataCite, Dublin Core). This will be complemented by schemas for the method-specific (meta)data documented as part of the low-level metadata. These are based on existing standards and minimum reporting guidelines and the enhanced versions created by BIOSUSHY (in intensive collaboration and alignment with other projects) will be proposed for standardisation. 3.3.5 What is the methodology followed? The following concepts, partly already described in previous sections, define the BIO-SUSHY methodology: • Individual management of research outputs according to their types allows optimal selection of tools for curation, documentation and sharing (e.g. electronic lab notebooks for protocols). • Separation of high- (biographical metadata, access options, licences, documentation of data model) and low-level (method-specific metadata) data documentation provides interoperability and computer-actionability for different applications (data discovery vs. data integration into computational workflows)