MetaTransfer: Harmonizing metadata of heterogeneous data resources in a common database framework
Abstract
Over the past decade, the increasing variety and number of deployed sensors, observation platforms and analytical devices, has resulted in a significant increase in the volume and complexity of scientific data in all disciplines of marine research. This trend is expected to continue with the development of new observation technologies and modelling approaches, calling for new strategies and approaches for data acquisition, processing, storage, analysis, and reuse. The widespread endeavor to manage such data according to F.A.I.R. principles led to the publication of vast amounts of datasets in a wide spectrum of databases and portals. However, while data publications steadily progress, the quality of the accompanying metadata does not yet keep up. High quality metadata, though, is essential to follow F.A.I.R. in its true intent to support and advance the reuse of published data. One of the obstacles to archiving high quality data and metadata is the lack of understanding of the added value of reusable data, as well as the requirements for reusability, and the often cumbersome and higheffort data archiving procedures and metadata standards. The current customary way requires scientists to provide metadata for administration and for publication in different locations such as modelling servers, doi servers, nucleotide sequence databases, or metadata catalogues. Meeting the requirements of all the used metadata standards of the different locations is timeintensive, requires expert knowledge, lacks incentivation, and might therefore prove too much effort, since for the scientists in possession of the required metainformation, their provision is only one out of many obligations and not the focus of their work. Thus two interwoven problems crystalize when publishing data including highquality metadata: Existing (extensive, but often not mandatory) metadata standards are not being used as intended or as comprehensively as necessary. Less extensive and easier to fulfil metadata standards may not facilitate data reuse. A solution needs to be a system that offers easy, onetime entry of metadata while at the same time ensuring that all relevant parameters for data reuse are collected. This may be achieved in two ways: interconnectedness of different data collection systems simplifies metadata submission, since entries can be transferred. Implementation of metadata standards beyond only their minimum mandatory fields to improve metadata quality for data reuse. In the case of the Leibniz Institute for Baltic Sea Research Warnemünde (IOW), several specialized metadata systems for different purposes already exist. These systems contain highquality, structured metadata: a Current Research Information System collecting all project information, involved scientists, publications, events, presentations, and collaborations, a Hyrax server for gridded datasets created from ocean model simulations, containing standardized metadata to provide access via the OPeNDAP protocol, a GeoNetwork metadata catalogue distributing metadata according to ISO standard for georeferenced metadata, a DOI server to publish datasets and software according to DataCite requirements, and a documentoriented database for DNA sequencing data and metadata in compliance with internationally accepted metadata standards (MIxS) and compatible with external data archives to facilitate data submission after project completion. Our aim is to connect all these systems in a «MetaTransfer Database» , enriching the metadata multifold. Ideally scientists need to input metadata once in an easytouse web application and only add partial information in case of using the other systems. This self registration system uses APIs to connect to the concerned databases and retrieve already existing metadata «per mouse click». Shared and specific metadata will be automatically transferred and may be exported in required formats. This presents one solution tailored to the specific needs and data infrastructure at IOW. It will provide an easier way to collect F.A.I.R. data and metadata and thereby offer incentives to publish better documented and reusable datasets.