UNTWIST project: Data Management Plan
Full text
UNCOVER AND PROMOTE TOLERANCE TO TEMPERATURE AND WATER STRESS IN CAMELINA SATIVA THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 D6.3 Data Management Plan Ref. Ares(2022)6763777 - 30/09/2022
2 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 Table of Contents Document Summary ..................................................................................................................... 3 Abstract ........................................................................................................................................ 4 List of Abbreviations ..................................................................................................................... 5 1. Introduction ........................................................................................................................... 6 2. DMP Model ............................................................................................................................ 9 2.1. Data Summary ................................................................................................................................. 9 2.2. FAIR data ....................................................................................................................................... 10 Making data findable, including provisions for metadata ........................................................................... 10 Making data openly accessible .................................................................................................................... 11 Making data interoperable .......................................................................................................................... 13 Increase data reuse (through clarifying licences) ........................................................................................ 13 2.3. Allocation of resources .................................................................................................................. 14 2.4. Data security ................................................................................................................................. 14 2.5. Ethical aspects ............................................................................................................................... 15 2.6. Other issues................................................................................................................................... 15
3 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 Document Summary Deliverable number & title: D6.3 – Data Mangagement Plan Version & submission date: V2 – 30th September 2022 Lead Beneficiary: FZJ Related Work Package: WP6 Author(s): Björn Usadel (FZJ) Reviewer(s): Richard Haslam (RRes), Juan Rodrigues (RTDS), Filoklis Pileidis (RTDS), Stéphane Bernillon (INRAe), Maider Goñi (INI), Camino Fábregas (INI), Dominik Grosskinsky (AIT), Claudia Jonak (AIT) Datasets underlined: None Communication level: ☒ PU Public ☐ CO Confidential, only for members of the consortium (including the Commission Services) Approved by: ☒ AIT (COO) ☒ CCE ☒ INRAE ☒ RRes ☒ RTDS ☒ UNIBO ☒ FZJ ☒ INI Grant Agreement Number: 862524 Programme: SFS-30-2019 / Agri-Aqua Labs: Looking behind plant adaptation Start date of Project: 1st September 2020 Duration: 5 years Project coordinator: AIT
4 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 Abstract UNTWIST is part of the Open Data initiative of the EU. To best profit from open data it is necessary to not only store data but to make this findable, accessible, interoperable and reusable (FAIR). Open and FAIR data, however, considers the need to protect individual data sets. The aim of this document is to provide guidelines on: ● Principles guiding the data management in UNTWIST ● End point repositories for data deposition ● Answers to the EU questionnaire on DMP as a DMP document The Data Management Plan (DMP) details how data is to be handled during and after the project. The UNTWIST DMP is modelled around the H2020 Online Manual. It will be updated/its validity checked during the UNTWIST project several times. At the very least, this will happen at month 36 (D6.4) and at the project end (D6.6). Disclaimer: The opinions expressed on this deliverable are those of the author(s) only and should not be considered as representative of the European Commission’s official position. This public deliverable is a revised version of D6.3 Data Management Plan (submitted by M6), updated as requested by reviewers in the 1st project periodic assessment.
5 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 List of Abbreviations AIT AIT Austrian Institute of Technology GmbH CCE Camelina Company España S.L. CC Creative Commons DDBJ DNA Data Bank of Japan DEC dissemination exploitation and communication (for UNTWIST) DMP Data Management Plan DOI Digital Object Identifier EBI European Bioinformatics Institute EMPHASIS European Multi-Environment Plant Phenotyping And Simulation Infrastructure ENA European Nucleotide Archive EU European Union FAIR Findability, Accessibility, Interoperability, and Reusability FAME Fatty Acid Methyl Esters FZJ Forschungszentrum Jülich GmbH GA General Assembly (of UNTWIST) GDPR General data protection regulation (of the EU) GRA Grant Agreement INI Iniciativas Innovadoras, S.A.L. INRAe Institut national de recherche pour l’agriculture, l’alimentation et l’environnement IP Intellectual Property MIAMET Minimal Information about Metabolite experiment MIAPPE Minimal Information about Plant Phenotyping Experiment MinSEQe Minimum Information about a high-throughput Sequencing Experiment MS Mass Spectrometry NCBI National Center for Biotechnology Information NFDI National Research Data Infrastructure (of Germany) NGS Next Generation Sequencing PRIDE Proteomics Identification Database RNASeq RNA Sequencing SOP Standard Operating Procedures SRA Short Read Archive ONP Oxford Nanopore qRT PCR quantitative real time polymerase chain reaction RRes Rothamsted Research RTDS RTDS Association UNIBO Alma Mater Studiorum - Università di Bologna WP Work Package XC-MS any (x) chromatography coupled to mass spectrometry
6 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 1. Introduction UNTWIST follows an open access strategy. In addition to plant data, UNTWIST is dedicated to modelling a large set of multi-omics data, some of which are well defined and standardized in terms of raw output and easily comparable within this omics discipline (sequencing data), others are standardized for output (proteomics data) and the remainder have varying levels of standardisation for output (e.g. metabolite measurements or even plant phenology). Despite this variation, in all cases necessary standards and specialist repositories exist (e.g. ENA for transcript and genomics data). Hence UNTWIST will ultimately deposit large volumes of data in public repositories, besides keeping derived values and models in the UNTWIST Knowledge Hub. Standardized omics repositories have proven their reliability and perfect integration into the FAIR landscape. 2. Data Management Data Management and sample tracking is very important for UNTWIST. Hence UNTWIST has a separate work package for data management, sample tracking which are leading to data for the knowledge hub. The main two things that are being tracked: a) physical materials i.e. plants (plant lots) that are phenotyped for (whole) plant traits seeds which have been centrally produced in INRAE-IJPB and samples that are generated from plants which are grown. These might be shipped to other locations to be analysed (e.g. leaf samples from RRES to INRAE - Bordeaux for metabolite profiling) or are analysed on the spot (e.g. RRES leaf samples to be subjected to lipidomic profiling). Here samples can get a code as determined on the sampling workshop, however in addition samples are mostly identified by the primary growth location. In any case these samples are linked to the plants that they came from. After physical samples have been subjected to analysis (e.g. FAME profiles, carbon isotope discrimation or phenotypes such as plant height) these then result in b) measurements such as i) genomic information which is independent of the environment ii) measurements dependent on the environment (e.g. field conditions; greenhouse-imposed treatments such as drought by withholding water) such as • carbon isotope discrimination • total redox status • phenotypic data iii) “high throughput omics” measurements • FAME oil content • detailed lipidomic data • transcriptomic • metabolomic • epigenetic • proteomic
7 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 • enzyme activity datasets • detailed redox status iv) measurements and/or descriptions describing the environment. Here, the differentiation between ii and iii is often arbitrary but is reflected in the WP structure of UNTWIST. In any case, genomic data can thus be made a characteristic of a plant line for modelling in task WP4 and to be shared with the community (keeping in mind that heterozygous SNPs will get lost or fixed in individuals). All measurements which are dependent on the environment (i.e. ii and iii) need to be also clearly linked against the environment they were grown in. As samples are grown in several locations, measurement data is linked to a sample which is described by specific environmental information and can feature a “field map” which describes where plants where grown in a simple map in a greenhouse and/or a field (which is derived from coordinates describing precise location). These data are first checked for consistency and completeness by INRAE and FZJ and are afterwards made available in a database. There the values can be looked at, plotted and downloaded at convenience. Using the database ensures that once data is checked it is available and will not be lost. Nevertheless, underlying data tables are also kept on a sftp server and are backed up regularly. Raw data (such as sequencing run fastq data) that is to be uploaded into other repositories is also kept there where necessary. Bringing these data together in a database allows to link the different sets. For example, the sample metadata describing the sample ca be derived by determining the plant (line) and the environment the plant was grown in. By anchoring samples to the primary producer of the plant the samples were derived from it is also possible to check and recheck environments. Furthermore, the sample data base allows to search for plant and samples (See Figure 1) and allows plotting and comparing measurements within the database. Figure 1 search and autocomplete searching as of 8/2022 for C18 lipid measurements offers several data sets to select (in mol % as well as in µg /mg)
8 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 Figure 2 plotting a measurement set and showing the values for the plot in the database. The example shows C18:2 lipid measurements The location data allows to check for spatial effects but also allows to check easily if associations were correct and thus allows an additional way to check for data consistency. The database is thus conceptionally associating different plants, samples and environments as conceptualised in Figure 3. Figure 3 conceptional model of the UNTWIST database The actual flow of samples is determined by plant producers (lysimeter, glasshouse, controlled chambers) and by the partners analysing the samples depending on whether samples are measured on site or shipped for measurements. The location of a sample is an important characteristic allowing to trace its provenience. Hence
9 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 the location is kept not only in the short sample name, but also usually in freezer boxes etc. In addition, often sample tables are accompanying samples in physical form (i.e. a print out) as an additional measure to make sure that samples are not confused and that data in the database can thus be rectified in an updated version. This is particularly important for larger batches of shipments where only a simplified naming scheme is used on the actual sample vial to allow fast enough sampling in order to account for diurnal and other processes requiring fast sampling and based on the experience with “freezing resistant” stickers (please check revised version of 6.1). Hence, when entering data into the database the full sample code is being restored where applicable. Once all data is complete and correct it can then also be submitted as data sets to public domain specific databases such as the ones from the EBI. 3. DMP Model The sample tracking platform now relies on a first working version of the UNTWIST. It is password protected, so before any data can be obtained or samples generated an authentication needs to take place. (please check MS7) 3.1. Data Summary What is the purpose of the data collection/generation and its relation to the objectives of the project? For UNTWIST, data collection and integration is absolutely necessary as data is not only used to understand principles, but it is also used for developing models in WP6, which need to be informed about data provenance. It is therefore of importance that not only data is well generated, but also well annotated using open standards and metadata as it is laid out in the following section. As UNTWIST aims at uncovering new principles across diverse camelina lines, culture conditions (countries, field vs contained) and as it crosses multiple omics disciplines experiments, good documentation, data keeping and integration are necessary. What types and formats of data will the project generate/collect? We foresee that the following data will be collected and generated at the very least: Firstly, phenotypic data about camelina plants, secondly experimental descriptions and where necessary logging of the data, thirdly omics data including genomics, epigenomics, metabolomics, transcriptomics, lipidomic, enzyme activity profiles, and proteomics data sets stemming from different approaches such as next generation sequencing (NGS) and/or quantitative Real Time PCR-based approaches. In addition, derived data from the original raw data sets will also be collected. This is important, as different analytical pipelines might yield different results or include ad-hoc data analysis parts. Therefore, specific care needs to be taken to document and archive these resources (including the analytic pipelines) as well. Will you re-use any existing data and how? The project builds on existing data sets and relies on them. For instance, without a proper genomic reference it is very difficult to analyze NGS data sets, even though better genomes will be produced in the course of the project. It is also important to include existing data sets on the expression and metabolic behavior of camelina and the background knowledge of the partners. Genomic references can simply be gathered from reference databases for genomes/sequences like the National Center for Biotechnology Information: NCBI1 (US); 1 https://www.ncbi.nlm.nih.gov/
16 D6.3 Data Management Plan THIS PROJECT HAS RECEIVED FUNDING FROM THE EUROPEAN UNION’S HORIZON 2020 RESEARCH AND INNOVATION PROGRAMME UNDER GRANT AGREEMENT NO 862524 UNCOVER AND PROMOTE TOLERANCE TO TEMPERATURE AND WATER STRESS IN CAMELINA SATIVA UNTWIST PARTNERS