Full text
GRAPHIA Knowledge Graphs, AI Services and Next Generation Instrumentation for R&D in Social Sciences and Humanities Work Package WP1 Project Coordination and Management Deliverable D1.1 Data Management Plan Funding Instrument: Horizon Europe Call: HORIZON-INFRA-2024-TECH-01 Call Topic: R&D for the next generation of scientific instrumentation, tools, methods, solutions for RI upgrade Project Start: 2025-01-01 Project Duration: 36 months Document Identifier: 10.5281/zenodo.15690808 Funded by the European Union. Grant Agreement number 101188018. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
Deliverable Information WP Number: 1 WP Title: Project Coordination and Management Deliverable Number: D1.1 Deliverable Full Title: Data Management Plan Deliverable Short Title: DMP Document Identifier: GRAPHIA-D1.1-DMP-v1.0 Beneficiary in Charge: IBL PAN Report Version: v1.0 Report Submission Date: 2025-06-30 Dissemination Level: PU Nature: Report Lead Author(s): Mateusz Franczak/IBL-PAN, Nikodem Wołczuk/IBL-PAN Co-author(s): Reviewer(s): Luca De Santis/Net7, Marte De Leeuw/KU Keywords: data management, knowledge graphs, AI Status: Final 2 of 22
Change Log Date Version Author/Editor Summary of Changes made 2025/05/15 v0.1 Mateusz Franczak/IBL-PAN, Nikodem Wołczuk/IBL-PAN Initial draft 2025/06/10 v0.2 Luca De Santis/Net7 1st review 2025/06/12 v0.3 Marte De Leeuw/KU 2nd review 2025/06/17 v0.4 Mateusz Franczak/IBL-PAN Improvements suggested by reviewers 2025/06/23 v1.0 Nikodem Wołczuk/IBL-PAN Finalisation of the document 3 of 22
Table of Contents Deliverable Information 2 Change Log 3 List of Abbreviations 5 Executive Summary 6 1. Introduction 7 1.1 Purpose and Scope of the Document 7 1.2 Structure of the Document 8 2. Data summary 9 2.1 Overview of data types and purposes 9 2.2 Data formats 11 2.3 Estimated data volume 12 2.4 Data collection and handling 12 2.5 Access, storage, and sharing 13 3. FAIR data 15 4. Allocation of Resources 17 5. Data Security 19 6. Ethical Aspects 20 Appendix A. 22 4 of 22
List of Abbreviations AI Artificial Intelligence CNRS Centre national de la recherche scientifique DMP Data Management Plan DOI Digital Object Identifier EU European Union FAIR Findable, Accessible, Interoperable, Reusable GDPR General Data Protection Regulation JUST Judicious, Unbiased, Safe, Transparent KG Knowledge Graph LLM Large Language Model LLM4SSH Large Language Model for Social Sciences and Humanities LPG Labelled Property Graph RDF Resource Description Framework SME Subject-Matter Expert SSH Social Sciences and Humanities T Task WP Work Package 5 of 22
Executive Summary This Data Management Plan (DMP) defines the strategy for handling data generated and used throughout the lifecycle of the GRAPHIA project. The plan supports the project's overarching goals by ensuring the responsible, efficient and ethical management of various data types across all work packages. The DMP defines the types of data to be collected and describes their expected formats, sources and volumes. It provides a framework for storing, sharing and preserving data according to FAIR (Findable, Accessible, Interoperable, Reusable) principles and prioritises open access and long-term usability where possible. The document also defines a data management framework, specifying how responsibilities are distributed among project partners and how appropriate resources are allocated. In addition, the plan addresses data security by outlining general measures to protect information throughout its lifecycle. It defines expectations for secure storage, controlled access and responsible sharing of data, taking into account levels of sensitivity. Ethical issues are an integral part of the DMP. The plan establishes principles to ensure that all data involving people is collected and processed responsibly, in full compliance with legal obligations and institutional ethical standards. This DMP is a living document that will be revised as the project evolves, adapting to new types of data, emerging ethical issues and technological developments. It ensures that the data produced by GRAPHIA will make a significant contribution to future research, infrastructure development and open science. 6 of 22
1. Introduction 1.1 Purpose and Scope of the Document This Data Management Plan has been developed as an integral part of the GRAPHIA scientific infrastructure initiative. The purpose of this document is to outline how data generated, collected, processed, and stored throughout the project lifecycle will be handled to ensure its quality, accessibility, interoperability, and long-term preservation. Scientific infrastructure projects produce a wide array of datasets that are critical not only for the success of the project itself but also for the broader scientific community. Proper data management ensures that these valuable resources are made available in a FAIR (Findable, Accessible, Interoperable, Reusable) manner, in line with international best practices and funder requirements. All public datasets, documentation, and software outputs generated by GRAPHIA will acknowledge EU funding in accordance with Horizon Europe requirements and will include a reference to the Grant Agreement number. This ensures transparency and proper attribution of financial support. The outputs of the project—such as LLM4SSH, the SSH Knowledge Graph, and AI-based modules—are designed to be usable and accessible by external researchers, SMEs, and other projects. Open standards and APIs will be implemented to maximise interoperability and reuse beyond the project’s original scope. This plan also addresses issues related to data provenance and governance, ensuring the traceability, accountability, and stewardship of data assets throughout their lifecycle. These aspects are closely aligned with activities under Task 5.1 and the development of the Data Trust Framework, which provides a comprehensive policy and technical foundation for trustworthy data sharing. Intellectual Property generated during the project will be managed in compliance with the Consortium Agreement and relevant EU regulations. This includes the proper attribution, licensing, and exploitation of results to support both academic and commercial uptake. 7 of 22
Furthermore, the use of AI technologies in GRAPHIA will comply with emerging European regulatory frameworks, including the EU AI Act1. Ethical considerations, transparency, and risk management will be embedded in all AI-related processes to ensure responsible innovation. The plan will be continuously reviewed and updated as the project evolves, ensuring that it remains aligned with the project's objectives, technological developments, and institutional policies. 1.2 Structure of the Document This DMP addresses the following key areas: 1. A summary of the types of data that will be produced and used within the project (chapter 2). 2. Strategies to ensure that data adheres to FAIR principles (chapter 3). 3. Allocation of responsibilities and resources related to data management (chapter 4). 4. Measures to ensure data security and confidentiality (chapter 5). 5. Considerations related to ethical and legal compliance (chapter 6). 1 Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (AI Act). http://data.europa.eu/eli/reg/2024/1689/oj 8 of 22
2. Data summary The GRAPHIA project generates a wide range of data types reflecting its interdisciplinary, collaborative and infrastructure-focused nature. This section provides a structured overview of the nature, origin, purpose, and intended processing of the key datasets generated throughout the project lifecycle. Details about comprehensive description are available in Appendix A (GRAPHIA Data Identification Table). 2.1 Overview of data types and purposes GRAPHIA data can be grouped into the following categories: ● Administrative and project management data Includes deliverables, milestones, financial reports, meeting notes, mailing lists and internal communication materials. These datasets mainly come from WP1 (Tasks 1.1-1.3) and support project coordination and reporting. ● SSH Knowledge Graph data The SSH Knowledge Graph (SSH KG) comprises semantically structured datasets that represent research instruments, concepts, institutions, and communities in the social sciences and humanities (SSH). It is developed using a hybrid architecture that combines Resource Description Framework (RDF) and Labelled Property Graph (LPG) technologies, enabling both semantic reasoning and advanced graph analytics. The SSH KG integrates standard ontologies and interoperable data models to ensure reusability across platforms and domains. It supports complex use cases involving artificial intelligence and large language models (LLMs), including automated enrichment, reasoning, and entity disambiguation. Data used to populate the SSH KG originates from a range of curated sources, as outlined in the Grant Agreement. These include scientific publications, research project records, archival metadata, primary source documents, expert interviews, and scholarly articles. This reused data will be processed in accordance with applicable licenses and data governance policies, ensuring compliance and traceability. 9 of 22
Creative Commons licenses (e.g., CC BY) for public reports, deliverables, and structured knowledge resources like D2.1 (Ontology). This ensures transparency, reproducibility, and the possibility for third parties to build upon project outputs. Software developed within the project—including AI modules, data processing pipelines, and the LLM4SSH components—will be released under Open Source licenses, such as MIT, Apache 2.0, or GPL, depending on the component’s nature, compatibility, and intended level of openness. This includes codebases for enrichment pipelines, model training scripts, and reusable infrastructure tools. Any ethical or legal restrictions that apply to specific datasets will be clearly documented. In such cases, access conditions and reusability limitations will be transparently communicated to ensure responsible use while maximising the value of reusable outputs across the SSH and AI research communities. The FAIR strategy is applied dynamically across the project's lifecycle, ensuring that evolving data types and sources remain compliant with these foundational principles. 16 of 22
4. Allocation of Resources Data management responsibilities are distributed across WPs, with each Task Leader ensuring proper documentation, secure storage, and appropriate access control for the datasets under their responsibility. Each task includes allocated resources to support key data-related activities such as anonymisation, curation, metadata generation, and internal reporting. WP1, and specifically Tasks T1.1 to T1.3, serve a coordinating role by overseeing the implementation of data governance procedures. These tasks include setting up internal data repositories (e.g., Google Drive structure), ensuring compliance with FAIR principles, and providing guidance on ethical handling of personal data collected through interviews and co-design workshops. In addition, WP5 - and specifically Task 5.1 - develops legal and ethical guidelines that inform all aspects of data processing throughout the project. This includes the formulation of a Data Trust Framework covering privacy, access and data management. These efforts are aligned with relevant EU legislation, such as GDPR, European Data Governance Act3, Ethics Guidelines for Trustworthy AI4 and the EU AI Act, and provide a common reference to data sharing, consent and accountability. WP1 ensures integration of these principles with practical data workflows and resource planning. Resource allocation also includes: ● Technical infrastructure: Secure cloud storage (Google Drive) is provisioned and maintained centrally. ● Human resources: WP leaders and designated data stewards are responsible for metadata creation and DMP updates. ● Budget allocation: Specific line items in task budgets (e.g., in WP5 for interview transcription and survey analysis, in WP2 for ontology documentation and dissemination) cover data preparation and processing needs. 4 High‑Level Expert Group on AI. (2019). Ethics Guidelines for Trustworthy AI, European Commission, Brussels. https://digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai 3 Regulation (EU) 2022/868 of the European Parliament and of the Council of 30 May 2022 on European data governance (Data Governance Act). http://data.europa.eu/eli/reg/2022/868/oj 17 of 22
In addition, certain tasks such as T2.1 and T5.3 involve outputs intended for public sharing (e.g., ontology datasets, user evaluation results), and include budgetary and technical support for open access publishing and archiving (e.g., on Zenodo). WP1 ensures alignment of these activities with the overall data strategy and regulatory compliance, including periodic review and revision of this DMP. 18 of 22
5. Data Security Throughout the project, data will be securely stored using a dedicated cloud space (Google Drive) configured specifically for project-related work. Access to this space is strictly limited to authorised team members and controlled through Google Workspace's security mechanisms. Google Drive provides encryption during both transfer and storage, providing a secure environment for data. The platform also enables file versioning and change tracking to support data recovery when needed. In the rare cases where local storage is necessary, it is limited to non-sensitive data only, unless there is no viable alternative. In such situations, local storage must be on encrypted disks with up-to-date antivirus protection, and access must be protected by strong passwords. Sensitive or personal data - such as survey results - must be anonymised at the earliest possible stage of processing and stored only in password-protected, encrypted GDPR-compliant environments. All team members are instructed to follow best practices in data processing, including secure file transmission, regular password updates, and compliance with institutional and EU data protection regulations. Access rights are granted based on roles and responsibilities, periodically reviewed and revoked when no longer needed. Programming code developed within the project is stored in a GitLab repository. Only designated developers responsible for software development are granted editorial access. GitLab provides secure storage, user access control and detailed change tracking to ensure traceability and accountability. Regular backups of critical data are maintained via a cloud-based platform to minimise the risk of accidental loss. 19 of 22
6. Ethical Aspects Throughout the life of the GRAPHIA project, all partners involved are required to maintain the highest ethical standards in all aspects of data collection, processing, storage and dissemination. Ethical issues are integrated into the project's data management procedures from the very beginning and are constantly reviewed. Informed Consent All participants taking part in qualitative research, such as interviews and workshops, are fully informed of the nature and purpose of the project, the intended use of their data, and their rights as participants. Consent is obtained prior to participation, usually through the submission of appropriate statements to be accepted. Participants are informed of their right to withdraw at any stage without any negative consequences. Anonymisation To protect the privacy and confidentiality of individuals, all sensitive or personally identifiable information is removed from datasets before analysis, processing or dissemination. This includes direct identifiers (e.g., names, email addresses) and indirect identifiers that may lead to re-identification in combination with other data (e.g., job titles, locations, affiliations). Anonymisation is carried out in accordance with established best practices and in full compliance with the GDPR, using methods appropriate to the type of data. For textual data, this may include redaction or semantic masking; for audio or video recordings, voice distortion or transcription anonymisation may be used. Where full anonymisation is technically or contextually impossible - especially in qualitative research - pseudonymisation and strict access control mechanisms are implemented instead. All anonymisation and access control procedures are documented at the task level and periodically reviewed. Responsible Use of AI Tools In line with the project's use of artificial intelligence-based tools for data extraction, classification and enrichment, special attention has been paid to the ethical challenges associated with automated systems. These include the risk of factual inaccuracies (so-called “hallucinations”), propagation of social or algorithmic biases, and unintended reintegration of personal data. These concerns are addressed 20 of 22
through the ethical and legal framework developed in Task 5.1, which draws on the European Data Governance Act and the Ethics Guidelines for Trustworthy AI and applies the JUST principles (Judicious, Unbiased, Safe, Transparent) to guide responsible design and data handling. Output generated by artificial intelligence systems will be critically reviewed by human experts before dissemination or reuse. Where enrichment pipelines involve external or pre-trained models, metadata will clearly distinguish between humanand machine-generated content. Efforts are being made to trace the origin of models, document uncertainty and ensure compliance with relevant ethical principles, including fairness, non-discrimination and respect for privacy. Limited dissemination All datasets intended for dissemination undergo thorough ethical and legal review to ensure that they do not contain personally identifiable or otherwise sensitive information. Selected datasets - such as qualitative research results (e.g., transcriptions or recordings of interviews), stakeholder feedback and annotations generated by artificial intelligence - may remain restricted and will be shared only within the project consortium or under controlled access conditions. Access to such datasets may require specific agreements, such as signed declarations of data use or institutional consents, and is subject to the principles of proportionality, necessity and informed consent. For public dissemination, documentation and metadata will clearly indicate any restrictions or limitations on use. Ethical and legal justifications for non-disclosure will be recorded and periodically reviewed to ensure ongoing compliance with FAIR principles and data management obligations. 21 of 22
Appendix A. This appendix presents a preliminary data inventory table developed in the early stages of the GRAPHIA project as part of the preparation of this Data Management Plan. The table serves as a structured overview of the anticipated datasets across all work packages, including their origin, format, ownership, expected volume, level of access and sensitivity. The list was developed through a coordinated consultation process with work package leaders and task owners, who provided input on the types of data they expect to generate, collect or reuse. It reflects an initial understanding of the project's data landscape and builds on early planning and design documents. Given that the GRAPHIA project is still in its early stages, many of the items in this table are provisional or indicative. Specific details - such as exact data formats, licensing schemes and publication dates - will be refined as the project progresses and data collection activities develop. As such, the list should be viewed as a living, evolving component of the DMP, adapted to the dynamic nature of research workflows in large, interdisciplinary projects. The table will be regularly reviewed and updated in future versions of the DMP, ensuring that it remains accurate, complete and aligned with the project's technical and ethical obligations. A file to download (GRAPHIA_Data_identification.xlsx) will be made available in Zenodo together with this document. 22 of 22