LiNES Project Data Management Plan
Abstract
LiNES Project Data Management Plan v1.1 april 2025 Životní cyklus nových zdrojů energie (LiNES)reg. n. CZ.02.01.01/00/23_020/0008508Horizon EuropeData Management Plan
Full text
Životní cyklus nových zdrojů energie (LiNES) Data Management Plan 1.1 29 May 2025 Data Management Plan created in Data Stewardship Wizard. Corresponds with Horizon Europe DMP template.
HISTORY OF CHANGES Version Publication date Changes 1.1 16 Apr 2025 update 2 / 13
Contributors The following contributors are related to the project of this DMP: Vladimír Církva Czech Academy of Sciences, Institute of Chemical Process Fundamentals (ÚCHP AV ČR) ORCID: 0000-0001-5351-4003 [email protected] Role: Data Steward Affiliation: Czech Academy of Sciences, Institute of Chemical Process Fundamentals (ÚCHP AV ČR) Martin Jakubec Czech Academy of Sciences, Institute of Chemical Process Fundamentals (ÚCHP AV ČR) ORCID: 0000-0001-7377-9359 [email protected] Role: Data Manager Affiliation: Czech Academy of Sciences, Institute of Chemical Process Fundamentals (ÚCHP AV ČR) • • 3 / 13
Acronym: LiNES Project Number: CZ.02.01.01/00/23_020/0008508 Start date: 2024-09-01 End date: 2028-12-31 Funding: Ministerstvo Školství, Mládeže a Tělovýchovy (Czechia): CZ. 02.01.01/00/23_020/0008508 (granted) Projects We will be working on the following project and for those are the data and work described in this DMP. Lifecycle of New Energy Sources, Životní cyklus nových zdrojů energie The project investigates new energy sources (NES) at several levels, from developing innovative technologies for energy storage and preventing risks associated with NSE operation to managing NSE as a source of secondary raw materials. With the expansion of new energy sources in the EU, a scientific problem arises in assessing safety, especially in the medium and long term. The project assesses the risks of NSE to human health and safety, optimizes using these sources, and examines the economic and environmental impacts. It focuses on research into new materials, battery diagnostics, and incident prevention, innovative environmentally friendly processing of NSE, and an overall assessment of the environmental impacts of NSE. 4 / 13
1. Data Summary Instrument datasets The following instrument datasets will be acquired in the project: CSV Data sets This dataset will be collected by experts in the project, with our own equipment. The equipment is very well described and known. Researchers working in other fields of research could be interested in using this data. We think that other researchers can use this data as follows: Repeating our experiments, verifying data, using collected data as a benchmark. Re-used datasets We have found the following reference datasets that we have considered for re-use: Crystallography Open Database (COD) It is available via: https://www.crystallography.net/cod/. It is used in the project. Owner of this dataset: The Crystallography Open Database (COD) is maintained by a collaborative team of scientists, with primary coordination and technical development led by Dr. Saulius Gražulis and Andrius Merkys at the Institute of Biotechnology, Vilnius University, Lithuania. For inquiries or feedback, you can contact the COD team via email at: [email protected].. The dataset can be used in the provided format without any conversion needed. We will use version " The latest revision number recorded in the COD is 291735." of this dataset. If a new version becomes available during the project, we will stay with the old version. We will keep a copy of the dataset and make it available with our results for the reproducibility. We will use the dataset as follows: Analyzing crystal structures of organic, inorganic, or metal-organic compounds. Studying polymorphism or phase transitions in crystalline solids. Screening structures for catalyst design, battery materials, or porous frameworks like MOFs. Reverse engineering crystallographic data into functional material properties. SciFinder-n It is available via: https://scifinder-n.cas.org. It is used in the project. Owner of this dataset: SciFinder-n is a subscription-based platform provided by the Chemical Abstracts Service (CAS), a division of the American Chemical • • • 5 / 13
Society. https://www.cas.org/solutions/cas-scifinder-discovery-platform/casscifinder. The dataset can be used in the provided format without any conversion needed. The provider keeps old versions around so the same reference data will be available to reproduce our results. We will use the dataset as follows: Substance records with full molecular details, reaction mechanisms and synthetic pathways, Journal articles, patents, and dissertations. NIST Chemistry WebBook It is available via: https://webbook.nist.gov/chemistry/? utm_source=chatgpt.com. It is used in the project. Owner of this dataset: The NIST Chemistry WebBook is maintained by the National Institute of Standards and Technology (NIST), specifically under the Standard Reference Data Program (https://webbook.nist.gov/chemistry/? utm_source=chatgpt.com). The dataset can be used in the provided format without any conversion needed. We will use version "While the database itself is identified as version 1.0, it undergoes continuous updates to incorporate new data and improvements. The most recent update to the data was in 2025 . " of this dataset. If a new version becomes available during the project, new analyses will be done with the new version. The provider keeps old versions around so the same reference data will be available to reproduce our results. We will use the dataset as follows: Thermochemical Analysis: 1) Determining heats of formation, heat capacities, phase change data, and reaction enthalpies., 2) Useful in reaction engineering, energy balance calculations, and environmental modeling. Spectroscopy: 1) Access to IR, mass, and UV/Vis spectra, 2) Valuable for compound identification, spectral comparison, and instrument calibration. . Agilent GC/MS Spectral Labraries It is available via: https://www.agilent.com/en/product/gas-chromatographymass-spectrometry-gc-ms/gc-ms-application-solutions/gc-ms-libraries? utm_source=chatgpt.com. It is used in the project. Owner of this dataset: The Agilent GC/MS Spectral Libraries are developed and maintained by Agilent Technologies, a global leader in life sciences, diagnostics, and applied chemical markets (https://www.agilent.com). . The dataset can be used in the provided format without any conversion needed. We will use version "Different versions of Agilent MassHunter or ChemStation software are compatible with specific library releases. Newer instruments may • • 6 / 13
require updated libraries. " of this dataset. If a new version becomes available during the project, we will stay with the old version. The provider keeps old versions around so the same reference data will be available to reproduce our results. We will use the dataset as follows: Compound Identification: 1) Match electron ionization (EI) spectra from your sample to those in the library, 2) Assign chemical names, structures, and confidence scores to unknown peaks. Targeted Screening: Use target libraries for fast, focused identification. We have found the following non-reference datasets that we have considered for reuse: nmrshiftdb2 (nmrshiftdb2) It is available via: https://fairsharing.org/FAIRsharing.nYaZ1N? utm_source=chatgpt.com. It is used in the project. Owner of this dataset: The nmrshiftdb2 project is an open-access NMR database for organic structures and their spectra, maintained by the Department of Chemistry at the University of Cologne (Universität zu Köln, https://nmrshiftdb.nmr.uni-koeln.de/nmrshiftdbhtml/credits.html? utm_source=chatgpt.com, [email protected]). The dataset can be used in the provided format without any conversion needed. We will use its online version without downloading it. It is a fixed dataset, changes will not influence reproducibility of our results. We will use the complete dataset. We will use the dataset as follows: NMR Spectrum Interpretation: 1) Compare experimental ¹H or ¹³C NMR data to predicted values from the database, 2) Use it as a supporting tool for structure elucidation. Spectral Prediction Tools: 1) Use nmrshiftdb2’s integrated prediction engine to estimate chemical shifts before running real NMR experiments, 2) Useful in planning syntheses or evaluating possible products. Open Science and Data Contribution: 1) Researchers can submit and share their own spectra, improving community resources, 2) Supports FAIR data principles. There is no need to harmonize different sources of existing data in our case. Data formats and types We will be using the following data formats and types: Rich Text Format (RTF) It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Document management -- Electronic document file format for long-term preservation -- Part 1: Use of PDF 1.4 (PDF/A-1) (ISO 19005-1:2005) • • • 7 / 13
It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Comma-separated Values (CSV) It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Tagged Image File Format (TIFF) It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Joint Photographic Experts Group Format (JPEG Format) It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Crystallographic Information Framework (CIF) It is a standardized format. This is a suitable format for long-term archiving. We will have only a small amount of data stored in this format. Audio Video Interleave (AVI) It is a standardized format. This is a suitable format for long-term archiving. We expect to have 500 GB of data in this format. 2. FAIR Data 2.1. Making data findable, including provisions for metadata Supporting information for publications (published) The dataset has the following identifiers: DOI: DOI provided by the repository We will distribute the dataset using: Domain-specific repository: ASEP - Repository of Czech Academy of Sciences Repozitář ASEP https://www.re3data.org/repository/ r3d100014032. We don't need to contact the repository because it is a routine for us. A persistent identifier will be assigned by the repository. The repository will make sure that the persistent identifier can be resolved to a digital object. The assigned persistent identifier is specified: To assign a DOI, we typically need to upload your dataset to a trusted repository. There won't be different versions of this data over time. We will be adding a reference to the published data to at least one data catalogue. • • • • • • ◦ ◦ 8 / 13
There are the following 'Minimal Metadata About ...' (MIA...) standards for our experiments: Nuclear Magnetic Resonance Markup Language (NMR-ML) We will use an electronic lab notebook to make sure that there is good provenance of the data analysis. We made a SOP (Standard Operating Procedure) for file naming. Everyone in the project will name files and folders according to the abbreviations of the institutions, and the folders will also be structured. personal codes will be used to identify the creator of the data, along with a code of the experiment. We will be keeping the relationships between data clear in the file names. 2.2. Making data accessible We will be working with the philosophy as open as possible for our data. All of our data can become completely open over time. Data that is not legally restrained will be released after a fixed time period (10 years), unconditionally. Metadata will be openly available without instructions how to get access to the data. Metadata will available in a form that can be harvested and indexed (managed by the used repository / repositories). We have a consortium agreement that arranges Intellectual Property. For the reference and non-reference data sets that we reuse, conditions are as follows: Crystallography Open Database (COD) It is freely available for any use (public domain or CC0). SciFinder-n It is available under specific restrictions, which we will follow in our project: Access typically requires an active institutional subscription and a registered personal account. . NIST Chemistry WebBook It is freely available for any use (public domain or CC0). Agilent GC/MS Spectral Labraries It is available under specific restrictions, which we will follow in our project: Access and usage of these libraries are governed by specific licensing agreements. This agreement outlines the permissible uses and restrictions associated with the software and data. The libraries are intended for use with Agilent's software platforms, such as MassHunter, and are designed to assist in the identification and analysis of compounds in gas chromatography/mass spectrometry (GC/MS) applications. nmrshiftdb2 (nmrshiftdb2) It is freely available for any use (public domain or CC0). • • • • • • 9 / 13