Full text
From Interdisciplinary Research Data Management to Reproducible Research in Practice Steffen Strohm1, Mattis thor Straten1,2, Andrea Göhring3,4, Hartwig Bünning5, Peer Kröger1,4, Oliver Nakoinz4,6, Matthias Renz1,2,4, and Christoph Rinne4,6 https://doi.org/10.5281/zenodo.17361589 Abstract The proliferation of heterogeneous data in large-scale interdisciplinary projects, such as the Cluster of Excellence ROOTS, requires strategies ensuring FAIR research outcomes (data, code), and thus supporting scientific collaboration. This paper introduces the early conceptual framework and activities for the Data Management and Data Science Platform (DMSP), established to bridge Research Data Management (RDM) and Data Science in the archaeology-focused research project ROOTS. This work introduces the Modular Research Pipeline (MRP), which integrates the lifecycle perspective of RDM with the procedural stages of the Knowledge Discovery in Databases (KDD) pipeline as a core concept of the platform. MRP provides a flexible framework for structuring research into reusable procedures and products, accommodating the iterative nature of research and the need for structured human input at critical junctures. Additionally, levels of granularity and fields of activity in the DMSP perspective on interdisciplinary research are layed out. Key activities include fostering a trust-based culture, developing analytical methods that combine data science with domain knowledge through collaborative use-cases, and aligning the DMSP with the existing RDM ecosystem (recommendations, existing services and so on). The DMSP’s next steps of implementation at this early stage are organised across four strategic areas: people, research, infrastructure, and further concept development. By combining lifecycle-oriented stewardship with a broad data science perspective, the DMSP aims to transform interdisciplinary research practices within ROOTS into processes that are transparent, reproducible and yield sustainably FAIR outcomes. Keywords: Research Data Management, Data Science, Information Science, Interdisciplinary Research, Archaeology 1Department of Computer Science, Kiel University, 2NFDI4Objects – Research Data Infrastructure for the Material Remains of Human History, 3Leibniz Laboratory for Radiometric Dating and Stable Isotope Research, Kiel University, 4Cluster of Excellence ROOTS – Social, Environmental, and Cultural Connectivity in Past Societies, 5Research Group Ecosystem Research, Geoarchaeology and Polar Ecology, Kiel University, 6Institute of Prehistoric and Protohistoric Archaeology, Kiel University Correspondence [email protected]kiel.de 1
2 Steffen Strohm et al. 1. Introduction2 This work aims to introduce the conceptual framework as well as fields of activities planned3 for the Data Management and Data Science Platform (DMSP), an organisational unit within the4 DFG-funded Cluster of Excellence ROOTS1at Kiel University. The platform aims to utilise prac-5 tices from Data Science and Research Data Management (RDM) in order to support research6 within ROOTS and achieve research outcome-related goals which have become more and more7 common in the last years, e.g. fulfilling FAIR requirements (see Figure 1). With 25 Principal In-8 vestigators (PIs) involved in successfully submitting the ROOTS proposal for the second phase9 (2026-2032), the research project will continue to combine its members’ multidisciplinary ex-10 pertise to address transdisciplinary research questions, linking insights gained about the past to11 environmental and societal challenges of the present and near future.2From a DMSP standpoint,12 this comes with a highly dynamic research environment within ROOTS, multi-faceted research13 perspectives, and heterogeneous research data.14 Figure 1 – Introductory details about the Data Management and Data Science Platform (DMSP), incorporating the environment of Initiatives, Communities and Infrastructure as well as the core concept of Modular Research Pipeline and planned fields of activities. After introducing the project context and DMSP goals in this section, section 2will describe15 core concepts for DMSP activities, especially the Modular Research Pipeline and related charac-16 teristics of the DMSP perspective. Section 3looks at previous experiences and existing infras-17 tructure within the local and global ecosystem DMSP will be established in, illustrating links to18 relevant actors in the institutional environment at Kiel University, technical infrastructure on the19 web, and national or global initiatives in RDM and involved research communities. An overview20 of upcoming steps will be provided in section 4.21 1https://gepris.dfg.de/gepris/projekt/390870439?language=en 2https://www.uni-kiel.de/en/cluster-roots 2
Steffen Strohm et al. 3 1.1. Interdisciplinary Research Project ROOTS22 The Cluster of Excellence ROOTS at Kiel University is an interdisciplinary research project23 with a strong focus in archaeology, funded by the German Research Foundation (DFG). Estab-24 lished in 2019, the project investigates the dynamic interactions between human societies and25 their environments over time. It brings together researchers from archaeology, the natural sci-26 ences, the humanities, and the life sciences to explore how cultural practices, environmental27 factors, and social structures have been connected in past contexts. The research framework28 of ROOTS is organized into several subclusters, each addressing a central theme: environmen-29 tal hazards, dietary transformations, innovation and cognition, urbanization processes, social in-30 equality, as well as conflict and reconciliation. These thematic areas are complemented by archae-31 ological and historical laboratories, where empirical methods and theoretical perspectives are32 integrated. Besides the subclusters being thematic subdivisions with distinct research focuses,33 additional selected components shape the structure of ROOTS as a whole: the methods nucleus,34 the technical platform, as well as the data management and data science platform (DMSP).35 1.2. Goals of the Data Management and Data Science Platform36 The overarching goals of the Data Management and Data Science Platform (DMSP) are closely37 aligned with the guiding principles of sustainable, transparent, and interdisciplinary research38 within ROOTS. A primary objective is the implementation of the FAIR (Findable, Accessible, In-39 teroperable, Reusable) principles (Wilkinson et al., 2016), ensuring that data and other research40 outcomes produced within the Cluster are consistently retrievable for long-term usability. At the41 same time, DMSP is tailored to the specific needs of ROOTS, recognizing the project’s interdis-42 ciplinary character and the diverse formats, standards, and practices arising from archaeology43 and the other involved disciplines from the natural sciences and humanities. Building on existing44 work, the FAIR principles should be applied not only to research data in the narrow sense, but45 also to workflows, and consequently to research software (Barker et al., 2022). Beyond RDM,46 the platform seeks to foster collaborations and provide the resulting concrete outcomes in the47 form of data science tools, offering researchers accessible software solutions for analysis, visual-48 ization, and integration of heterogeneous datasets. To support interdisciplinary discovery, DMSP49 develops and promotes integrative workflows that combine machine learning techniques with50 established domain-specific methods, complemented by structured human input. Such feedback51 may stem from experts in particular disciplines, but also – where appropriate – from broader pub-52 lic engagement, aligning with open science values. Through this combination of FAIR-oriented53 infrastructure, tailored services, and transparent workflows, DMSP aims not only to counsel re-54 searchers in managing their data efficiently, but to transform it into a catalyst for new insights55 across ROOTS and a platform for collaborative action.56 2. Core Concepts and Approaches57 2.1. Modular Research Pipeline58 The Modular Research Pipeline (MRP) is proposed as a conceptual framework and practical59 perspective that combines two established and already closely related perspectives on research60 activity: the Research Data Lifecycle (UK Data Service, 2022) from an RDM point of view and61 the Knowledge Discovery in Databases (KDD) pipeline (Fayyad et al., 1996) from Data Science.62 While this work is the wrong place to add to the discussion about the scope, applicability, and63 3
4 Steffen Strohm et al. Figure 2 – Modular Research Pipeline (MRP) concept containing procedures (arrows) and products, and simplified example of an MRP instance view on a research use case (see Strohm et al., 2023c,d). shortcomings (Frické, 2009) of the Data, Information, Knowledge, Wisdom (DIKW) Hierarchy (Row-64 ley, 2007), it should be noted that the general idea is inherently embedded in both perspectives.65 The Research Data Lifecycle (RDLC) has become a central model for structuring the highly66 dynamic processes in research for systematic review in data management and curation, high-67 lighting stages through which research data (repeatedly) pass: planning, collection, processing,68 analysis, preservation, sharing, and eventual reuse (Cox and Tam, 2018)). The cyclic model under-69 scores the importance of sustainability, transparency, and reproducibility, while also emphasizing70 the iterative nature of research, where data and methods are often revisited, revised, adapted,71 and repurposed. The KDD pipeline, on the contrary, originates in the field of data science and72 describes the systematic process of extracting meaningful patterns from large datasets through73 stages such as data selection, preprocessing, transformation, mining, and interpretation (Fayyad74 et al., 1996). Clearly focusing on computational analysis, the KDD pipeline is often applied with75 less attention to long-term data stewardship, version management, and other related contextual76 challenges.77 The MRP seeks to bridge these two traditions, offering an approach tailored to identifying78 main subtasks in data analysis as well as in the interdisciplinary environment of the Cluster of79 Excellence ROOTS. By embedding lifecycle thinking into the more technically oriented KDD80 paradigm, MRP supports both methodological innovation and the responsible management of81 research data as a scholarly resource. One key feature of MRP is its recognition of iteration and82 modularity. With this in mind, a single, distinct research endeavour (individual research activity)83 which is both unique and somewhat separate from any other research endeavour is viewed as a84 modular process of productive procedures and products.85 Figure 2lays out the design of the Modular Research Pipeline, identifying five procedural86 steps which in itself may contain subtasks and will result in intermediary or final research out-87 comes, depending on the current research activity in scope. For better illustration any subtasks,88 repetition of steps (e.g. a complex computational analysis which can be separated as several89 analytical steps) and the overall cyclic nature of the pipeline are omitted from the visualisation.90 At first, general research preparation and data acquisition will result in raw data. If data have91 4
Steffen Strohm et al. 5 been collected using computer software or scripts, this first step may additionally have produced92 code, or notes in a manual data collection process. These notes or the code (assuming adequate93 comments provide the necessary minimum of documentation of intent) can be understood as94 descriptive metadata of the procedure and thus as produced outcome as well and should be han-95 dled as such. Having acquired raw data leads into the next step of pre-processing data through96 cleaning (handling missing values, ambiguities, outliers), transforming data structures and files97 into needed formats, fusing and synchronising several data sources – all of these usually going98 hand in hand with some descriptive statistics to better understand the underlying data. The re-99 sulting data set(s) is now a fitting input for computational analysis. In this third step of the MRP,100 analytical processing of the data takes place, e.g. data mining, machine learning, pattern recogni-101 tion or dimensionality reduction. The computational results will usually be evaluated regarding102 their meaning and explainability, contextualised and interpreted by experts of the corresponding103 field, which is the fourth step of the MRP model. The insights gained by experts are then the104 basis for publication which often consists of a textual description, e.g. in a paper or book chapter,105 necessary data and the code to reproduce or even replicate the presented methodical approach106 and results. In an ongoing research project over several years, preliminary research results might107 just be shared inside the project at first, before being published months or years later. The need108 for earlier internal sharing with collaborators will require an access regime and has implications109 for any practical implementation (see section 3).110 MRP allows cycles of refinement in which data are acquired, prepared, analysed, interpreted,111 and subsequently re-integrated into the research ecosystem. This iterative logic necessitates112 careful version control and documentation, ensuring that datasets, methods, and workflows re-113 main transparent, reproducible, and reusable over time.114 2.2. Levels of Application and Integrative Challenges115 The modular nature of MRP and its steps is intended to function on three levels of analysis116 and application within the research project: individual, subcluster, and cluster level. At the scale117 of individual research endeavours, it structures the workflow of single projects, making the transi-118 tions between data preparation, analysis, and interpretation explicit. Within ROOTS subclusters,119 MRP enables the coordinated integration of data and methods across related disciplines, facil-120 itating collaborative workflows and cross-validation of results. At the overarching cluster level,121 MRP provides a framework for harmonizing heterogeneous research activities, thereby creating122 a shared environment in which archaeological, historical, and natural science data can be brought123 into dialogue.124 Beyond the level of individual research endeavours, MRP instances can be understood as125 units, which are then connected in consecutive operations, essentially starting at any step and126 intermediary outcome of an individual MRP (see Figure 3). Connecting MRPs inevitably creates127 dependencies in research practice. Without claiming to be exhaustive, DMSP will focus on three128 integrative actions and related tasks required to guarantee relevant pre-conditions to be met.129 These three actions are: reuse of methods and reuse of data, which are closely related to the130 FAIR principles if methods are reduced to their artifacts when being implemented (e.g. code of131 a software, pseudo code specifying an algorithm), as well as data fusion, meaning the data-level132 combination of heterogeneous data sets for a more comprehensive view on a research object or133 phenomenon of interest.134 5
6 Steffen Strohm et al. Figure 3 – Simplified illustration of individual-level MRPs (1-3) and emerging integrative actions of reuse and interoperability on (sub)cluster level: (A) reuse of methods, (B) reuse of data , (C) data fusion. Reuse of data and methods within a subcluster or the whole cluster ROOTS comes with135 needs of standardised representation and generalisation in methodology, a central field where136 DMSP will be active to support researchers in finding feasible, effective, and at best efficient137 ways of collaboration across disciplines and research interests. Data fusion should be thought138 of as different from data integration, which often means the harmonisation of relatively simi-139 lar data sets. In a data fusion process, data from usually heterogeneous sources is combined140 in analyses to proliferate additional information with the goal of gaining deeper insights into141 (inter)dependencies, correlations and other characteristics hidden in the newly combined data142 corpus (see Zheng, 2015,2025).143 The implications of adopting the MRP in a larger interdisciplinary project are extensive. It pro-144 motes the creation of reusable data, methods and implementations. It aligns with the FAIR prin-145 ciples for research data management, and integrates structured human input at critical junctures.146 Expert feedback ensures that computational processes remain context-sensitive and epistemo-147 logically grounded. At the same time, MRP creates a framework to support workflow manage-148 ment, reproducibility, and scientific exchange (RDM, RSE), positioning itself as a flexible yet rigor-149 ous framework for interdisciplinary research. By combining lifecycle-oriented stewardship with150 a methodologically broad data science perspective, MRP aims to transform research practices151 within ROOTS into processes that are not only computationally powerful but also integrative,152 transparent and sustainable.153 6
Steffen Strohm et al. 7 3. Towards Practical Implementation154 Especially from a researcher’s point of view, it might be tempting to reinvent the wheel. How-155 ever, DMSP’s hybrid nature, which combines data science and research data management, re-156 quires a reasonable balance between new developments and the reuse of existing solutions.157 There are already numerous pre-existing concepts, recommendations, methodologies, technolo-158 gies as well as schemaand data-level standards which should be considered for deliberate159 decision-making. Partially, this already happens when aligning DMSP activities with preexisting160 RDM recommendations, current research, and when adapting to the conditions present within161 the ROOTS project environment, research communities, and infrastructural ecosystem. At this162 early stage in the project period, it is important to identify areas in which action is needed and163 to derive relevant fields of DMSP activity. This will form the basis for the continuous develop-164 ment of specifications as the project progresses and common understanding between the DMSP165 members and the researchers emerges and expands.166 3.1. Interdisciplinary Research Data Management167 Interdisciplinary Research Data Management (IRDM) summarises the effort of conceptualis-168 ing preconditions, challenges, design decisions and solutions to achieve set project goals while169 dealing with the complexity of research projects, the heterogeneity of people and data involved,170 and interdisciplinarity as a challenging characteristic itself (see Strohm et al., 2023a,b). The aim of171 IRDM is to support the project in four key areas: alignment with the FAIR principles as a core re-172 quirement by funding institutions, presentation and reporting of research results to funders, the173 research community and a broader audience, promotion of collaboration and early data exchange174 especially before publication of results, and provision of an integrated search space for explo-175 ration of all project data where applicable (see Strohm et al., 2024a,b). Practical implementation176 of roles, processes, and software require the application and adaptation of the IRDM concept,177 specifically for the ROOTS environment. This includes identifying and prioritising specific tasks,178 continuously reassessing and further developing the IRDM perspective, and providing support179 not only in handling research data, but also in reusing method implementations through work-180 flows based on the framework of the MRP concept. The notion of workflows comes with an181 understanding that certain tasks, steps or procedures can be automated. Finding such repeat-182 able parts in MRP instances within the ROOTS research practice will be essential for DMSP183 activity. It also provides the basis for generalisation of methods for cross-domain application.184 3.2. Alignment with the Existing Ecosystem185 The activities of the Data Management and Data Science Platform are embedded in a multi-186 layered ecosystem ranging from internal to local, national, and global initiatives. At the internal187 level of the ROOTS project, collaborations take place with individual researchers, subclusters,188 and cross-cutting units such as the technical platform or the methods nucleus.189 Locally at Kiel University, the opendata@uni-kiel repository, as the institutional home for190 FAIR research results, and the Centre for Interdisciplinary Data Science (CIDS)3play a central191 role; the latter promotes the exchange with subject-matter experts, advises researchers along192 3https://www.cids.uni-kiel.de 7
8 Steffen Strohm et al. the data life cycle and provides an expert network as basis for interdisciplinary method develop-193 ment. The ISAAK4group (a community of practice in the field of programming for archaeological194 analysis) is another relevant cornerstone to provide opportunities of practical insights for young195 (and old) researchers.196 At the national level, the DMSP follows the recommendations of the DFG (German Research197 Foundation) and seeks active exchange with NFDI4Objects (National Research Data Infrastruc-198 ture for Material Cultural Heritage in the Humanities and Natural Sciences). NFDI4Objects aims199 at FAIR-compliant, long-term archiving and standardised vocabularies and is developing its own200 knowledge graph for this purpose, for example. Other national partners include the IANUS5 201 research data repository, a DFG-funded data centre for archaeology and ancient studies in Ger-202 many, and re3data6, a global directory of research data repositories, which will be used as a203 source for domain-related active repositories if the need in contributing disciplines emerges.204 Globally, the DMSP connects as a part of ROOTS with CfAS (Coalition for Archaeological205 Synthesis)7, which is dedicated to boost archaeological data and knowledge synthesis, the gen-206 eral repository Zenodo8, the data publisher for earth and environmental sciences PANGAEA9.207 Services like Ariadne (a European research infrastructure for the integration of archaeological re-208 search data) and the European Open Science Cloud (EOSC), an EU initiative that aims to provide209 researchers with an open and trustworthy, multidisciplinary environment for publishing, find-210 ing and reusing research data, tools and services, are considered as relevant part of the DMSP211 ecosystem.212 All of these actors and potential partners shape the ecosystem in which DMSP will pursue its213 goals, and they come with constraining factors of differing intensity from, e.g. informational ma-214 terial to recommendations, fixed metadata requirements and standardised vocabularies. Finding215 a fitting, interoperable, or translatable representation of FAIR (and possibly open and CARE-216 compliant (Carroll et al., 2020)) research outcomes will be the result of prioritising and aggregat-217 ing such specifications.218 3.3. Scientific Method Development219 Methodological development focuses on specialising data science approaches, which usually220 follow the Knowledge Discovery in Databases (KDD) pipeline, for the specific requirements of221 the ROOTS cluster within the Modular Research Pipeline (see section 2.1) concept. In order to222 test and further develop the concepts of MRP and IRDM in practice, the DMSP actively partic-223 ipates in concrete use cases. These take the form of collaborations with researchers in ROOTS224 on questions of analytical research, which require adaptation of existing methods, development225 of new methods, and also allow the theoretical approaches to be directly challenged and vali-226 dated. Such use cases provide deeper insights in the modalities and contents of practically used227 research data as well as analytical goals. Creating a common understanding on the base level of228 research activity in ROOTS, will also enable researchers to start asking new questions and aiming229 for different, potentially more complicated research questions to achieve new research results.230 4https://isaakiel.github.io/ 5https://ianus-fdz.de/ 6https://www.re3data.org/ 7https://www.archsynth.org/ 8https://zenodo.org/ 9https://www.archsynth.org/ 8
Steffen Strohm et al. 9 However, this is an iterative process which can only succeed with adequate time investment and231 maintained interest on both sides of a collaboration.232 An important second step is the generalisation of methods in the transition from individual233 MRPs to the level of subclusters or the entire cluster, but also in collaborations beyond the234 project. This involves identifying patterns in method adaptation and application on a reflective235 basis with the potentials of automated workflows in mind, enabling a higher potential of reusing236 methods for cross-dataset and cross-domain application.237 3.4. Training and Counselling238 The training and counselling activities aim to strengthen the skills of ROOTS members and239 support knowledge transfer. A key partner in this regard is the Young Academy within ROOTS,240 for whose members specific courses are being developed, but also the Computing Centre at Kiel241 University and the Research Data Management division therein. Connecting experts in research242 data management (RDM), data science and computer science with ROOTS members from ar-243 chaeology and tightly related disciplines will help to enable researchers to meet their research244 goals and requirements of funding bodies through targeted support and advice. To this end, the245 content and formats of the training and counselling services are specifically defined and closely246 coordinated with the general activities and schedule of the ROOTS cluster. For practical compu-247 tation sessions and related courses, DMSP will be able to use the hardware capabilities of the248 CIDS Lab, including desktop clients, workstations, graphical interface hardware and more.249 4. Next Steps250 The successful establishment of the DMSP hinges on cultivating a people-centric approach,251 aiming for open exchange and building trust across all stakeholder groups. This involves identify-252 ing and staffing relevant roles in order to implement conceptual and procedural models used as a253 guideline for DMSP activity. ROOTS members, particularly Principal Investigators (PIs), must be254 actively engaged to collect and articulate specific data management needs and expectations in255 order to start an early synchronisation and development of a common understanding. In this way,256 they will participate in (re-)shaping DMSP concepts and prioritising activities. For ROOTS young257 researchers, expected to join the project in the second half of 2026, DMSP will prepare an ad-258 equate onboarding through counselling and courses covering computational methods, practical259 programming, and foundational aspects of research data management.260 The strategy for research focuses on fostering collaborative use-cases and integrating commu-261 nity feedback to ensure practical relevance. This includes the essential development of analyt-262 ical methods that effectively combine advanced data science techniques with domain-specific263 goals and knowledge. A parallel effort is dedicated to the adaptation of the Interdisciplinary Re-264 search Data Management (IRDM) and Modular Research Pipeline (MRP) concepts tailored to the265 ROOTS context. Ultimately, these activities are designed to increase common understanding of266 both the DMSP’s capabilities and the practical needs of researchers, while establishing vital links267 to broader communities in research data management, computational archaeology, data science,268 and computer science.269 Regarding infrastructure, the platform must actively explore, connect, and shape the ecosys-270 tem of data management tools and services. A key action involves investigating state-of-the-271 art techniques, focusing on data standards, integration, and fusion. Crucially, the platform will272 9