Full text
Deliverable D9.1a Update Open Data Management Plan 24th May 2024 DELIVERABLE NUMBER: D9.1a Update
Dissemination level - Public 2 General Information Disclaimer: © 2024 by FishEUTrust Funded by the European Union. Views and opinions expressed are, however, only those of the author(s) and do not necessarily reflect those of the European Union. The European Union cannot be held responsible for them. Call identifier: HORIZON-CL6-2021-FARM2FORK-01-10 GA Number: 101060712 Start date of project: 01/06/2022 Work Package: WP9 Type: Deliverable Number: 9.1a Update Version: 1 Due Date: 30/5/2024 Submission date: 29/5/2024 Responsible organisation: Jožef Stefan Institute (JSI) Prepared by: Barbara Koroušić Seljak Reviewed by: Javier de la Cueva Dissemination Level: Public Document Type PRO Technical/economic progress report (internal work package reports indicating work status) DEL Technical reports identified as deliverables in the Description of Work X MoM Minutes of Meeting MAN Procedures and user manuals WOR Working document, issued as preparatory documents to a Technical report INF Information and Notes Dissemination Level PU Public X PP Restricted to other programme participants (including the Commission Services) RE Restricted to a group specified by the Consortium (including the Commission Services) CO Confidential, only for members of the Consortium (including the Commission Services) CON Confidential, only for members of the Consortium
Dissemination level - Public 3 Table of Contents Executive Summary .................................................................................................................................................................................. 4 1. Data Summary .................................................................................................................................................................................... 4 2. FAIR data ................................................................................................................................................................................................. 7 2.1 Making data findable, including provisions for metadata .................................................................................. 7 2.2 Making data accessible ............................................................................................................................................................... 7 Repository ............................................................................................................................................................................................... 7 Data ............................................................................................................................................................................................................. 8 Metadata .................................................................................................................................................................................................. 9 2.3 Making data interoperable ....................................................................................................................................................... 9 2.4 Increase data re-use ................................................................................................................................................................... 10 3. Other research outputs ............................................................................................................................................................... 10 4. Allocation of resources .......................................................................................................................................................... 11 5. Data security ........................................................................................................................................................................................ 11 6. Ethics ................................................................................................................................................................................................. 12 7. Other issues .................................................................................................................................................................................. 13 Conclusions ................................................................................................................................................................................................... 15
Dissemination level - Public 4 Executive Summary FishEUTrust aims to collect, manage, and utilize data related to fish and shellfish. This data encompasses a wide range of information, including the quality of both wild and farmed fish and shellfish, regulatory and labelling standards, and consumer acceptance. Some of this data was sourced from previous projects and partners' activities, but the majority has been collected through ongoing FishEUTrust initiatives. The FishEUTrust Data Management Plan (DMP) is a core document addressing all open internal research data management activities in FishEUTrust. This is a working document that complies with the Horizon Europe template for DMP. It outlines data management strategies for satisfying the requirements for research data management in Horizon Europe. 1. Data Summary Will you re-use any existing data and what will you re-use it for? State the reasons if re-use of any existing data has been considered but discarded. What types and formats of data will the project generate or re-use? What is the purpose of the data generation or re-use and its relation to the objectives of the project? What is the expected size of the data that you intend to generate or re-use? What is the origin/provenance of the data, either generated or re-used? To whom might your data be useful ('data utility'), outside your project? At the beginning of the project, the FishEUTrust partners contributed with research and industrial (e.g., OXY) data on fish composition and other characteristics collected in previous projects (e.g., Realmed, SeeFoodTomorrow etc.). During the project, additional data on fish has been generated covering different aspects (i.e., chemical/environmental, economic, social etc.). The data collected so far is detailed in Table 1, which is dynamically updated on the project's NextCloud portal. The data are heterogenous, meaning that they are of different types and formats. An objective of the project is to develop various software solutions (such as web services, Machine learning pipelines and other workflows for data analysis) for re-using heterogeneous data (i.e., data of different types and formats) provided from different data sources. The purpose of linking and analysing such data using advanced computer methods (Machine learning, Natural language processing etc., described in detail in D4.3) is to find answers on open research questions. This might be useful for researchers as well as for other FishEUTrust stakeholders (like fish industry, especially SMEs, policy makers etc.). A comprehensive list of FishEUTrust stakeholders has been prepared by ABT (in Task 1.1 to be published in D1.2).
Table 1. Data summary - draft Partner Existing data Data types Data formats Utility for Objectives Data size Data origin Data utility outside the project JSI (1) Composition data for fish Isotope data, elemental composition data, fatty acids in fish and mussels Semi-structured (XML) Identification of fish origin and production type (wild / farmed) X data points for Y different fishes RealMed, Promedlife / FishEUTrust Identification of fish origin and production type (wild / farmed) IPMA (2) Composition data for fish Lipidomic data Semi-Structured (XML) Identification of production type (wild / farmed) X data points for Y different fishes FishEUTrust Identification of production type (wild / farmed) UNIBO (3) Results of household survey on consumers from Denmark, Italy, and Slovenia attitudes and choices related to fish products and aquaculture Survey data (at household level) SPSS file Identification of characteristics and behaviours of consumers in 3 countriers, Denmark, Italy, and Slovenia, representative of the Northern, Mediterranean, and Central regions of Europe 388 kb FishEUTrust Identification of consumers characteristics, attitudes, and choices related to aquaculture and aquaculture products EuroFISH (4) FishEUTrust UNIFI (5) - Duccio Composition data for fish, Metagenome data Semi-structured (XML) Identification of fish origin and production type (wild / farmed) X data points for Y different fishes ? / FishEUTrust Identification of fish origin and production type (wild / farmed) UNIFI (5) - Giovanna Data on toxins in mussels Sensor data UMF (6) Data on biological and chemical contaminants in fish Sensor data Semi-structured (XML) Identification of contaminants in fish ? / FishEUTrust Identification of contaminants in fish DTU (7) Dietary health impacts and environmental impacts for 18 impact categories for 5800 food items Environmental impacts data Semi-structured (Excel) Ecological footprint calculation X data points for Y different scenarios Pilot cases from Denmark and Portugal, and from literature / FishEUTrust Ecological footprint calculation BTU CS (9) Data on fish freshness Instrumental analysis data Semi-structured (XML) Quantification of the fish freshness, i.e., the time since the fish was caught FishEUTrust Quantification of the fish freshness, i.e., the time since the fish was caught
Dissemination level - Confidential 6 for particular storage conditions for particular storage conditions NORCE (10) FishEUTrust EuroFIR (11) Plant-based bioactive compounds Semi-structured Composition data on food waste streams Semi-structured (XML) REFRESH Food composition data Semi-structured (XML) EuroFIR UNIPD (12) DNA coding data ABT (13) Data on stakeholders Publicly availabe data + survey data Semi-structured (Excel) Stakeholder analysis + Survey analysis 500 FishEUTrust Business development REDINN (14) FishEUTrust BELIT (15) / FishEUTrust CETGA (18) FishEUTrust DIGITALSMART (18) IoT data on fish pH, T, DO, ORP, EC, GPS, QR code Structured data (JSON) Digital passport 5kb per measurement (measurement resolution is set to 15 minutes, could be adjusted) Pilot case from Montenegro / FishEUTrust Water quality monitoring, Cold storage transport monitoring, Batch tracebility BUGENVILA (20) Details about seafood recipes Seafood recipes Unstructured (textual data) Studying how consumers are able to distinguish between wild and farmed fish once prepared as dish Few tens Bugenvila / OXY (20) Composition on water matrix, fish and feed Quantitative content automatic measured with digital sensors. Manual registered data on growth, mortalities, medication etc. Semi-structured (XML) Digital passport through Cobália (production management software) Data stored and available in 'the cloud' FishEUTrust/fish farm & hatchery in DK Original data is owned by the fish farms - and as such not available for public use
2. FAIR data 2.1 Making data findable, including provisions for metadata Will data be identified by a persistent identifier? Will rich metadata be provided to allow discovery? What metadata will be created? What disciplinary or general standards will be followed? In case metadata standards do not exist in your discipline, please outline what type of metadata will be created and how. Will search keywords be provided in the metadata to optimize the possibility for discovery and then potential re-use? Will metadata be offered in such a way that it can be harvested and indexed? Some data provided by project partners is already identified with persistent identifiers, such as the Digital Object Identifier (DOI) and similar. For example, datasets provided through the the H2020 project Food Nutrition Security Cloud (FNS-Cloud)1 have assigned identifiers since they are stored in Zenodo. Data in a form of scientific papers (provided by JSI, UNIBO, UNIPD, UNIFI) is identified by DOI. Datasets collected by EuroFIR comply with the CEN/TC387 standard on food data, which specifies various metadata describing data quality. Moreover, data from these datasets have been classified with respect to global systems like LanguaL2 and FoodEx23. Fish microbiome data gathered by Unifi is already annotated with respect to the FNS-Harmony ontology4. This application ontology was developed in FNS-Cloud to describe concepts and entities from the food and nutrition domain. Isotopic data for fish that collected in the FishEUTrust project are represented using the ISO-FOOD ontology5 developed in the FP7 ERA Chair ISO-FOOD project6. The ISO-FOOD ontology has been linked with the FNS-Harmony ontology to enable semantic interoperability of FishEUTrust data with data from other food and nutrition domains (in T7.3). Regarding search keywords, an objective of the FishEUTrust project is to develop a platform where functionality of advanced searching data based on its metadata is being implemented (in WP7). Metadata has been designed and structured to allow its harvesting and indexing. Experiences from previous projects (e.g., ERA Chair ISO-FOOD, FNS-Cloud, COMFOCUS etc.) have been taken into consideration defining the FishEUTrust metadata model. 2.2 Making data accessible Repository 1 https://www.fns-cloud.eu/ 2 https://www.langual.org/default.asp 3 https://www.efsa.europa.eu/en/data/data-standardisation 4 https://github.com/panovp/FNS-Harmony 5 Eftimov, T., Ispirova, G., Potočnik, D., Ogrinc, N., Koroušić Seljak, B. (2019) ISO-FOOD ontology: A formal representation of the knowledge within the domain of isotopes for food science, Food Chemistry 277: 382-390, doi.org/10.1016/j.foodchem.2018.10.118. 6 http://isofood.eu
Dissemination level - Confidential 8 Will the data be deposited in a trusted repository? Have you explored appropriate arrangements with the identified repository where your data will be deposited? Does the repository ensure that the data is assigned an identifier? Will the repository resolve the identifier to a digital object? Each partner stores and archives data in a trusted repository suggested by its institution. However, the project openly accessible outcomes will be deposited in Zenodo7, where all meta data is openly available under CC0 licence and all content is openly accessible through open APIs. Zenodo helps researchers to receive credit by making the research results citable and through OpenAIRE integrates them into existing reporting lines to the European Commission and other funding agencies. Citation information is also passed to DataCite and onto the scholarly aggregators. Apropos appropriate arrangements between the Consortium and Zenodo, the Terms and Conditions made public by Zenodo8 do not make necessary to arrive to an arrangement as these Terms and Conditions are public and their acceptance is necessary to use their services. As stated before, Zenodo provides DOI to all uploaded data, therefore the repository ensures that the data is assigned an identifier and resolves the identifier to a digital object. Data Will all data be made openly available? If certain datasets cannot be shared (or need to be shared under restricted access conditions), explain why, clearly separating legal and contractual reasons from intentional restrictions. Note that in multi-beneficiary projects it is also possible for specific beneficiaries to keep their data closed if opening their data goes against their legitimate interests or other constraints as per the Grant Agreement. If an embargo is applied to give time to publish or seek protection of the intellectual property (e.g., patents), specify why and how long this will apply, bearing in mind that research data should be made available as soon as possible. Will the data be accessible through a free and standardized access protocol? If there are restrictions on use, how will access be provided to the data, both during and after the end of the project? How will the identity of the person accessing the data be ascertained? Is there a need for a data access committee (e.g., to evaluate/approve access requests to personal/sensitive data)? In this stage of the project, we are still discussing which data will be openly available. According to FishEuTrust Document of Action (DOA), no data embargo is previewed. All openly accessible data will be available through open APIs. For other data, a data access committee (DAC) will be established to evaluate/approve access requests to personal/sensitive data. The DAC members will be elected at the next FishEUTrust general assembly. No restriction on the use of openly accessible data exists, apart from giving the appropriate credit to the data creator. In case FishEUTRust partners need to exchange any sensitive (personal, industrial) data, encrypted tools and safe FTP-solutions provided by their IT-services will be used. 7 https://zenodo.org 8 https://about.zenodo.org/terms/
Dissemination level - Confidential 9 Metadata Will metadata be made openly available and licenced under a public domain dedication CC0, as per the Grant Agreement? If not, please clarify why. Will metadata contain information to enable the user to access the data? How long will the data remain available and findable? Will metadata be guaranteed to remain available after data is no longer available? Will documentation or reference about any software be needed to access or read the data be included? Will it be possible to include the relevant software (e.g., in open-source code)? All meta data deposited in Zenodo will be openly available under CC0 licence. It will contain enough information (including direct link) to enable a user to access the data. The data will remain available and findable for at least five years after the end of the project. All metadata is guaranteed to remain available also after data is no longer available. As the case may be, if the used software includes a CITATION.cff file9, it will be used as reference. Citation.Cff files have been developed by The Institute for Software Technology of the German Aerospace Center (DLR), The Netherlands eScience Center and The Software Sustainability Institute. It is supported by Github, Zenodo, Zotero. 2.3 Making data interoperable What data and metadata vocabularies, standards, formats or methodologies will you follow to make your data interoperable to allow data exchange and re-use within and across disciplines? Will you follow community-endorsed interoperability best practices? Which ones? In case it is unavoidable that you use uncommon or generate project specific ontologies or vocabularies, will you provide mappings to more commonly used ontologies? Will you openly publish the generated ontologies or vocabularies to allow reusing, refining or extending them? Will your data include qualified references to other data (e.g. other data from your project, or datasets from previous research)? FishEUTrust structured data are formatted using popular formats avoiding in all cases vendor lock in problems. To make FishEUTrust data of different types and formats interoperable, existing application ontologies, like FNS-Harmony and ISO-FOOD, have been linked and used to represent the FishEUTrust knowledge. Both ontologies are already openly published. In FishEUTrust we are further developing a web-based tool FoodViz10, which enables mapping of concepts and entities to different semantic resources (such as Hansard corpus, FoodOn, SnomedCT etc.). 9 https://citation-file-format.github.io/ 10 http://foodviz.env4health.finki.ukim.mk/#/cafeteria?recipeId=c-10048971