scieee AI-readable full text Open interactive document viewer

956562 LeADS D7.5 Final Scientific ESRs Reports

Legatity Attentive Sata Scientists - LeADS - H2020 Project

Abstract

This report consists of a collection of scientific reports that the Early Stage Researcher (ESR) drafted to briefly illustrate the results of the research activities that they have carried out during the project.Each scientific report presents the state of the art of the topic analyzed by each ESR, summarizes the research findings, and discusses their broader implications.The report is organized as follows: Chapter 1 introduces the overarching objectives of the LeADS project;Chapter 2 explores the shifting domain boundaries between data protection, consumer protection and competition law on a conceptual as well as empirical level;Chapter 3 focuses on the development of concrete technical solutions for the implementation of privacy and security in various technologies;Chapter 4 delves into elements that enable data sharing and composing legal, technical and user-centred perspectives; andChapter 5 presents original views on the trustworthy development and deployment of AI.

Full text

Legality Attentive Data Scientists Grant Agreement No. 956562 D.7.5 - Final Scientific ESRs Reports The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 1 Deliverable Title D.7.5 - Final Scientific ESRs Reports Deliverable Lead: SSSA Partner(s) involved: All partners and all ESRs Related Work Package: WP7 Opening LeADS’s Crossroads Related Task/Subtask: 7.2 LeADS as an Innovator Enabler Main Author(s): Arianna Rossi (SSSA), Giovanni Comandé (SSSA) Other Author(s): Aizhan Abdrassulova (JU), Emmanouil Alexakis (UPRC), Tommaso Crepax (SSSA), Paul de Hert (VUB), Xengie Cheng Doan (UL), Sumeyra Dogan (JU), Armend Duzha (UPRC), Soumia Zohra El Mestari (UL), Afonso Ferreira (UT3), Mitisha Gaur (SSSA), Onntje Hinrichs (VUB), Bárbara da Rosa Lazarotto (VUB), Gabriele Lenzini (UL), Cristian Lepore (UT3-IRIT), Christos Magkos (UPRC), Elwira Macierzyńska-Franaszczyk (JU), Robert Lee Poe (SSSA), Imge Ozcan (VUB), Salvatore Rinzivillo (CNR), Yang Qifan (SSSA), Louis Sahi (UT3-IRIT), Maciej Zuziak (CNR) Dissemination Level: Public Due Delivery Date: 30.11.2024 Actual Delivery: 30.11.2024 Project ID 956562 Instrument: H2020-MSCA-ITN-2020 Start Date of Project: 01.01.2021 Duration: 48 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 2 02.10.2025 Version history table Vers. Date Modification reason Modifier(s) v0.1 23/10/2024 Initial draft of the document, no changes A. Rossi v0.2 28/10/2024 Collection of the majority of the ESRs Scientific Reports and Introduction, not harmonized A. Rossi v1.0 29/10/2024 Harmonization of the received ESRs Scientific Reports in one single document A. Rossi, Q. Yang, M. Zuziak v1.1 31/10/2024 Executive summary and conclusions added A. Rossi v1.2 12/11/2024 Revision of introduction and individual contributions and LateX errors solved G. Lenzini, G. Comandé v1.3 19/11/2024 Revision of introduction, conclusions and executive summary A. Rossi v1.4 28/11/2024 Removed errors, and general revision G. Lenzini v1.5 29/11/2024 General revision and conclusion G. Lenzini v2.0 01/10/2025 Final Version Reviewer’s comment addressed V. Virdis, A. Rossi Legal Disclaimer The information in this document is provided “as is”, and no guarantee or warranty is given that the information is fit for any particular purpose. The above-referenced consortium members shall have no liability for damages of any kind, including without limitation direct, special, indirect, or consequential damages that may result from the use of these materials subject to any liability which is mandatory due to applicable law. ©2021 by LeADS Consortium. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 3 Page intentionally left empty The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 4 CONTENTS CONTENTS Contents 1 Introduction 21 1.1 Interplay of data protection, consumer protection and competition law . . . . 23 1.2 Solving practical hurdles to privacy and security implementation . . . . . . . . 23 1.3 Legal, technical and user-centred enablers of data sharing . . . . . . . . . . . 25 1.4 Trustworthy development and deployment of AI . . . . . . . . . . . . . . . . . 27 2 Data Protection, Consumer Protection and Competition Law 29 2.1 Data as Property and Regulatory Subject Matter in European Consumer Law . . 29 2.1.1 Executivesummary............................ 29 2.1.2 Introduction ............................... 30 2.1.3 Research Findings and Analysis . . . . . . . . . . . . . . . . . . . . . . 31 2.1.4 Conclusion ................................ 43 2.2 Personal Data Protection and Market Competition . . . . . . . . . . . . . . . . 45 2.2.1 Executivesummary............................ 45 2.2.2 Introduction ............................... 45 2.2.3 Methodology............................... 46 2.2.4 Literaturereview ............................. 46 2.2.5 Research findings and analysis . . . . . . . . . . . . . . . . . . . . . . 48 2.2.6 Conclusions................................ 56 2.2.7 Futurework................................ 58 2.3 Unchaining Data Portability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 2.3.1 Executivesummary............................ 60 2.3.2 Introduction ............................... 60 2.3.3 Methodology............................... 75 2.3.4 ResearchFindings............................. 77 2.3.5 ResearchAnalysis............................. 83 2.3.6 Conclusions................................ 85 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 5 CONTENTS CONTENTS 2.3.7 Futurework................................ 86 3 Privacy and Security Implementations 87 3.1 Assessing Digital Identity Solutions . . . . . . . . . . . . . . . . . . . . . . . . 87 3.1.1 Executivesummary............................ 87 3.1.2 Introduction ............................... 88 3.1.3 Objectives and Scope. . . . . . . . . . . . . . . . . . . . . . . . . . . 88 3.1.4 StateoftheArt.............................. 89 3.1.5 Methodology............................... 89 3.1.6 A Model to Assess e-Identity Solutions . . . . . . . . . . . . . . . . . . 91 3.1.7 ResearchFindings............................. 95 3.1.8 ResearchAnalysis............................. 99 3.1.9 Recommendations ............................ 100 3.1.10 Conclusions................................ 102 3.2 Distributed reliability and blockchain-like technologies . . . . . . . . . . . . . 106 3.2.1 Executivesummary............................ 106 3.2.2 Introduction ............................... 106 3.2.3 Background and problem statement . . . . . . . . . . . . . . . . . . . 108 3.2.4 State of the Art: Towards collaborative data processing . . . . . . . . . 110 3.2.5 Objectives and Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . 111 3.2.6 Methodology............................... 113 3.2.7 ResearchAnalysis............................. 116 3.2.8 Conclusions and future work . . . . . . . . . . . . . . . . . . . . . . . 118 3.3 Data governance in distributed IoT systems and edge computing . . . . . . . . 123 3.3.1 Executivesummary............................ 123 3.3.2 Introduction ............................... 123 3.3.3 Methodology............................... 127 3.3.4 ResearchFindings............................. 128 3.3.5 ResearchAnalysis............................. 133 3.3.6 Conclusions................................ 135 3.3.7 Futurework................................ 136 3.4 User Empowerment in Information Management . . . . . . . . . . . . . . . . 137 3.4.1 Executivesummary............................ 137 3.4.2 Introduction ............................... 137 3.4.3 Background................................ 137 3.4.4 StateoftheArt.............................. 138 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 6 CONTENTS CONTENTS 3.4.5 1.3 Objective and scope . . . . . . . . . . . . . . . . . . . . . . . . . . 140 3.4.6 Methodology............................... 141 3.4.7 ResearchFindings............................. 141 3.4.8 RiskStratification............................. 141 3.4.9 Large Language Models . . . . . . . . . . . . . . . . . . . . . . . . . . 148 3.4.10 Personal data protection and PHIMS . . . . . . . . . . . . . . . . . . . 152 3.4.11 PHIMS and user empowerment . . . . . . . . . . . . . . . . . . . . . 153 3.4.12 Research Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154 3.4.13 Privacy and Security Framework . . . . . . . . . . . . . . . . . . . . . 154 3.4.14 Artificial Intelligence Integration . . . . . . . . . . . . . . . . . . . . . 154 3.4.15 Accessibility and User Empowerment . . . . . . . . . . . . . . . . . . 154 3.4.16 Technical Infrastructure and Standards . . . . . . . . . . . . . . . . . . 155 3.4.17 Healthcare Impact and Outcomes . . . . . . . . . . . . . . . . . . . . 155 3.4.18 Data Ethics and Governance . . . . . . . . . . . . . . . . . . . . . . . 155 3.4.19 Conclusions................................ 155 3.4.20Futurework................................ 156 3.5 PrivacyFriendlyAI................................. 158 3.5.1 Executivesummary............................ 158 3.5.2 Introduction ............................... 159 3.5.3 Methodology............................... 161 3.5.4 ResearchFindings............................. 162 3.5.5 ResearchAnalysis............................. 165 3.5.6 Conclusions................................ 166 3.5.7 Futurework................................ 166 4 Enablers of data processing 169 4.1 Sharing information for the public good . . . . . . . . . . . . . . . . . . . . . 169 4.1.1 Executivesummary............................ 169 4.1.2 Introduction ............................... 170 4.1.3 Research Findings and Analysis . . . . . . . . . . . . . . . . . . . . . . 172 4.1.4 Navigating Regulatory Complexities in the Digital Economy . . . . . . . 172 4.1.5 ResearchQuestions............................ 174 4.1.6 Methodology............................... 174 4.1.7 Conclusions................................ 175 4.1.8 Research Analysis and Conclusions . . . . . . . . . . . . . . . . . . . . 175 4.1.9 Futurework................................ 177 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 7 CONTENTS CONTENTS 4.2 Collective Consent Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 178 4.2.1 Executivesummary............................ 178 4.2.2 Introduction ............................... 178 4.2.3 ResearchFindings............................. 183 4.2.4 Research Analysis and future work . . . . . . . . . . . . . . . . . . . . 189 4.2.5 Conclusion ................................ 191 4.3 Balancing Bias Mitigation and Data Protection . . . . . . . . . . . . . . . . . . 192 4.3.1 Executivesummary............................ 192 4.3.2 Introduction ............................... 192 4.3.3 EHDS – Objectives and Key Components . . . . . . . . . . . . . . . . . 195 4.3.4 BiasinAISystems............................. 196 4.3.5 Data Handling under EHDS’s Secondary Use Framework . . . . . . . . 197 4.3.6 The AI Act and Bias Mitigation . . . . . . . . . . . . . . . . . . . . . . 199 4.3.7 Research Findings and Discussion . . . . . . . . . . . . . . . . . . . . 199 4.3.8 Conclusions................................ 200 4.3.9 Futurework................................ 201 4.4 Data Collaboratives - Theory and Practical Considerations . . . . . . . . . . . . 202 4.4.1 Introduction ............................... 202 4.4.2 State-of-the-Art: Decentralised Machine Learning . . . . . . . . . . . . 204 4.4.3 Data Collaboratives in Context of the European Data Governance . . . 219 4.4.4 Collaborative Contribution Function and Alpha-Amplification on the Example of Federated Learning . . . . . . . . . . . . . . . . . . . . . . . 231 4.4.5 Conclusions................................ 241 4.5 Boundaries of data ownership: from concept to practice . . . . . . . . . . . . 246 4.5.1 Abstract.................................. 246 4.5.2 Executivesummary............................ 247 4.5.3 Introduction ............................... 247 4.5.4 Methodology............................... 250 4.5.5 ResearchFindings............................. 250 4.5.6 ResearchAnalysis............................. 265 4.5.7 Conclusions................................ 272 5 Trustworthy development and deployment of AI 275 5.1 GuidingAdoptionofAI .............................. 275 5.1.1 Executivesummary............................ 275 5.1.2 Introduction ............................... 277 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 8 CONTENTS CONTENTS 5.1.3 Methodology............................... 301 5.1.4 ResearchFindings............................. 301 5.1.5 ResearchAnalysis............................. 308 5.1.6 Conclusion ................................ 308 5.1.7 Futurework................................ 309 5.2 The Complexity of Value Alignment . . . . . . . . . . . . . . . . . . . . . . . . 310 5.2.1 Background................................ 310 5.2.2 Methodology............................... 310 5.2.3 Beyond the Law or Back Again? . . . . . . . . . . . . . . . . . . . . . 312 5.2.4 A Transition: from Fair Machine Learning to Distributive Decisions . . . 335 5.2.5 Distributive Decisions . . . . . . . . . . . . . . . . . . . . . . . . . . . 337 5.2.6 Appendix: Code for Data Analysis for 5.2.3.5.1 . . . . . . . . . . . . . . 347 6 Conclusion 353 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 9 LIST OF FIGURES LIST OF FIGURES Acronyms Definitions EC European Commission ECHR European Convention on Human Rights EDPB European Data Protection Board EDPS European Data Protection Supervisor EHDS European Health Data Space EHR Electronic Health Record eIDAS electronic Identification, Authentication and trust Services EMR Electronic Medical Record EPOS Electronic Point Of Sale ESR Early Stage Researcher EU European Union FedAvg Federated Averaging FedOpt Federated Optimization FHIR Fast Healthcare Interoperability Resources FFNPD Free Flow of Non-Personal Data Regulation FL Federated Learning FLaaS Federated learning as a service GDPR General Data Protection Regulation GPC Global Privacy Control HCI Human-Computer Interaction HDAB Health Data Access Body HLEG High-Level Expert Group on AI HITL Human In The Loop HL7 Health Level Seven HODA Hoda case (Hoda v. Google) HTML HyperText Markup Language HOTL Human On The Loop HW/SW Hardware / Software IID Independent and Identically Distributed H2020-MSCA-ITN-2020 Marie Skłodowska-Curie Actions — Innovative Training Networks 2020 IDSA International Data Spaces Association IMI Internal Market Information system IoT Internet of Things IP Intellectual Property ISO International Organization for Standardization JSON JavaScript Object Notation LLM Large Language Model LOO Leave-One-Out (maschera / metodo) LoA Level of Assurance MDR Medical Device Regulation ML Machine Learning MNIST Modified National Institute of Standards and Technology The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 16 LIST OF FIGURES LIST OF FIGURES Acronyms Definitions MPC Multi-Party Computation MSCA Marie Skłodowska-Curie Actions N/A Non Applicable / Non Available NGO Non-Governmental Organization NN Neural Network NoSQL Not Only SQL OCR Optical Character Recognition OOP Object Oriented Programming OWL Web Ontology Language PDS Personal Data Store PETs Privacy Enhancing Technologies PHR Personal Health Record PHIMS Personal Health Information Management System PID Personal Identification Data PII Personally Identifiable Information PIMS Personal Information Management System PPML Privacy Preserving Machine Learning PSI Private Set Intersection PSU Private Set Union P2PFL Peer-to-Peer Federated Learning QoS Quality of Service QTSP Qualified Trust Service Provider RDF Resource Description Framework RE Requirements Engineering RF Random Forest (modello) RFC Random Forest Classifier RL Reinforcement Learning RNN Recurrent Neural Network SGD Stochastic Gradient Descent SP Service Provider SL Split Learning SMPC Secure Multi-Party Computation SME Small/Medium Enterprise SSI Self-Sovereign Identity SyRI Systeem Risico Indicatie TAM Technology Acceptance Model TEE / TEEs Trusted Execution Environment(s) TFL Transfer Federated Learning TILLS Technology and Law Innovation Labs UI/UX User Interface / User Experience UKVI UK Visa & Immigration vCard Virtual Card format The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 17 LIST OF FIGURES LIST OF FIGURES Acronyms Definitions VFL Vertical Federated Learning W3C World Wide Web Consortium XAdES XML Advanced Electronic Signatures XML eXtensible Markup Language XSD XML Schema Definition The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 18 LIST OF FIGURES LIST OF FIGURES Executive Summary This report consists of a collection of scientific reports that the Early Stage Researcher (ESR) drafted to briefly illustrate the results of the research activities that they have carried out during the project. Each scientific report presents the state of the art of the topic analyzed by each ESR, summarizes the research findings, and discusses their broader implications. The report is organized as follows: Chapter 1 introduces the overarching objectives of the LeADS project; Chapter 2 explores the shifting domain boundaries between data protection, consumer protection and competition law on a conceptual as well as empirical level; Chapter 3 focuses on the development of concrete technical solutions for the implementation of privacy and security in various technologies; Chapter 4 delves into elements that enable data sharing and composing legal, technical and user-centred perspectives; and Chapter 5 presents original views on the trustworthy development and deployment of AI. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 19 LIST OF FIGURES LIST OF FIGURES Page intentionally left empty The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 20 1. Introduction 1Introduction In an increasingly datafied world, data scientists become ever more indispensable to analyze, understand and anticipate human and natural phenomena, as well as to propose solutions to the wicked problems of our time. Because of the significance that the collection, processing and generation of data increasingly exert on individual and collective decision-making, with shortterm as well as long-term impacts, data scientists in the modern society are required to be able to grasp the ethical, legal and social implications (ELSI) of their activities. They are also asked to direct their professional activities towards the amelioration of the common welfare, while reducing risks deriving from fast-paced technological advancement. In the new digital decade imagined by the EU, a growingly complex interplay of norms attempts to regulate personal and non-personal data use, the development of artificial intelligence, the secure uptake of technologies in all sectors of society, and so on. This is why, even the new generation of jurists needs to be equipped with innovative skills: they are called to create and apply laws to an extremely wide set of emerging technologies, understand the intersections and the tensions between different regulatory instruments, and ensure legal certainty for fostering the economy of the internal market in a globalized world. Such a complete, enlightened understanding cannot derive from traditional training paths. It rather needs, and deserves, an ad hoc audacious, innovative training program that enables a safe-by-design, inclusive digital transformation of public and private services: LeADS. From the onset, LeADS has represented a highly ambitious program of training and research, that has set the mission of creating the next generation of researchers and practitioners who can easily cross the boundaries between the disciplines of law and data science. The LeADS project has carried out an unprecedented, intensive plan of training activities to prepare Early Stage Researcher (ESR) who are apt to positively contribute to society both within and outside academia. In the LeADS project, collaboration (as a soft skill) and interdisciplinarity (as a hard skill) have represented the guiding principles that have contributed to achieve groundbreaking results that would have been simply impossible otherwise. The four Crossroads reports (D7.1) are a tangible result of this approach: framed as living documents that were continuously re-elaborated for 3 years, they were drafted by small interdisciplinary working groups. Each member contributed with his/her specific disciplinary expertise to a collaborative dialogue that enabled an original co-creation of knowledge on intersectional topics such as Privacy and Intellectual Property (Crossroad 1), Trust in Data Processing and Algorithm Design (Crossroad 2), Data Ownership (Crossroad 3), and Empowering Individuals (Crossroad 4). Another significant occasion where the ESRs have further trained their ability to cross disciplinary boundaries and pioThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 21 1. Introduction neer new knowledge through collaborative discussions are the joint Working Papers (WOPAs) (D5.2). Some WOPAs have also been published in international venues, thereby amplifying the outreach of the project’s results. But engaging in interdisciplinary work does not simply occur naturally to individuals who have followed traditional, mostly mono-disciplinary academic curricula. This is exactly one of the challenges that the project intended to overcome by recurring to creative educational and training instruments that have fostered theoretical and applied research. The ESRs have participated in an intensive training program with a great breadth of topics at the forefront of research, policy-making and industrial progress spanning across three years. Moreover, they all have carried out research activities and actively continued their training during two secondment periods at the institutions of other Beneficiaries (i.e., other academic institutions) or project Partners (i.e., companies or independent authorities). The participation in the 4 Technology and Law Innovation Labs (TILLS), as well as the management of the Innovation Challenge and the design of boardgames, among the others, have all been occasions for learning to master newfound skills, including problem-solving, effective communication and successful presentation skills, and for applying their knowledge to relevant complex, real-world problems. This report, D.7.5 - Final Scientific ESRs Reports, illustrates the individual paths that the ESRs have followed to become Legality-Attentive Data Scientists: in addition to the training and research activities described above, all the ESRs have also pursued a personal research goal on an assigned topic, decided at the time of the drafting of the project proposal but that in several instances has evolved to adapt to changes in the regulatory framework, in technologies or in the individual ESRs’ interests. Although it was not a requirement in and for the LeADS’ project, in order to address this challenge, the ESRs have all been enrolled in a PhD program of their own institution and they are at date illustrating their research results in a PhD thesis. As the ESRs increasingly mastered their individual research topic and integrated their own original perspective, while society, regulations and technologies evolved, the assigned topics have also evolved to better address the more pressing needs of the contemporary era. This report contains the research findings of each ESR, in the form of a scientific report: each individual contribution contains the objectives and the scope of the research, the methodology, the research findings, a discussion, and future work. The scientific supervisors of the ESRs have reviewed and approved the included reports. The report is organized in four thematic sections. Chapter 2 explores the shifting domain boundaries between data protection, consumer protection and competition law on a conceptual as well as empirical level. Chapter 3 focuses on the development of concrete technical solutions for the implementation of privacy and security in various technologies. Chapter 4 delves into elements that enable data sharing composing legal, technical and user-centered perspectives. Chapter 5 presents original views on the trustworthy development and deployment of AI. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 22 1. Introduction 1.1. Interplay of data protection, consumer protection and competition law 1.1 Interplay of data protection, consumer protection and competition law The first contribution “Data as Property and Regulatory Subject Matter in European Consumer Law” by Onntje Hinrichs) constitutes the first comprehensive study that examines to what extent data has become regulatory subject matter in European consumer law. The regulation of data in consumer law might ultimately lead to legal fragmentation and thus risks undermining the objective of European data protection law to provide a fully harmonized legal framework that facilitates the free flow of data. Moreover, the ESR’s work has sought to understand how ‘data openness’ can be reconciled with privacy and economic interests in data, since existing intellectual property rights regimes do not confer property rights in data. The research evaluated the limitations of a property rights approach and highlighted how ultimately it was abandoned in the European Data strategy, in favor of a new focus on facilitating data access and data sharing. The study showcases how data is a difficult fit for traditional legal categories and the challenges it provides to policy-makers in creating a coherent legal framework across legal disciplines. The second contribution, titled “The Interplay Between Personal Data Protection and Online Platform Market Competition under the GDPR”, by Qifan Yang, explores the role of personal data in the business value creation of online platforms, specifically the social media market, and the impacts of the GDPR on the EU social media market concentration, based on data science, economics, and law. The findings reveal that between 2015 and 2020 the GDPR has resulted in a decrease in EU social media market concentration and a particular negative impact on large social media companies. The GDPR framework (due to transparency, data portability, and compliance costs) curbed the dominance of leading platforms and levelled the playing field for small competitors. In the long term, the impact of the GDPR diminished after 2020 as large companies adopt appropriate strategies to comply with the regulation. The study also identifies two factors that influence the impact of the GDPR: the scale of the Internet market and the level of technology. “Unchaining data portability” by Tommaso Crepax highlights the current ambiguity surrounding data portability. At present, there is no clear consensus on how much or what type of data should be transferable to meet legal standards. The study argues that, by focusing on data quality, and more specifically, on how well data serves its intended purpose (“fitness for purpose”), we can develop a clearer benchmark to guide existing and future regulatory reforms. This insight is far-reaching, with implications not only for the legal frameworks that govern our digital lives, but also for how data markets evolve, how digital businesses build their models, and how fundamental rights are protected in the process. 1.2 Solving practical hurdles to privacy and security implementation The second section of this report has a more practical nature since it proposes solutions for the effective implementation of privacy and security requirements into various technologies. “Assessing digital identity solutions” by Cristian Lepore revolves around the timely topic of selfsovereign identities (SSI) that aim to return control of personal data to citizens, thereby reducing The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 23 1. Introduction 1.2. Solving practical hurdles to privacy and security implementation the risk of misuse and building greater trust in digital services. However, the concept remains elusive, lacking a clear, uniform interpretation. Efforts to define SSI are mostly led by academic initiatives, but the concept often remains abstract and disconnected from practical business needs. As a result, many digital identity systems claim to follow SSI principles, but it is difficult to assess how closely they adhere to its core values. This work seeks to offer a formal definition of Self-Sovereign Identity and introduce a practical, structured model for evaluating digital identity solutions with the aim of providing both designers and citizens with a tool to help them choose the best platform for creating a digital identity. Louis Sahi, in “Distributed reliability and blockchain-like technologies”, describes his work on distributed reliability, blockchain-like technologies and trust in data processing. Starting from defining the main components of data processing, the study presents why collaborative data processing in Europe raises challenges of trust, reliability and legal compliance. Thanks to a systematic literature review and interviews with data experts, the author identified the relevant data quality criteria, and classified them by trust and reliability. Moreover, the thesis proposes a blockchain-based solution that implements some criteria and demonstrates how this solution can maintain data quality and ensure reliability and trust in decentralized data management. However, this solution is not taking the exhaustive list of data quality criteria. A future direction of the work will concern an automated data quality assessment that describes the use of algorithms, rules, or models to assess trust and reliability without the need for direct human intervention. “Data managementandanalyticsonedge computing andserverlessofferings” byArmendDuzha describes a new approach to protect against risks related to personal data exploitation, drawing a methodology for the implementation of data management and analytics in edge computing and serverless offerings in considering privacy properties to modulate prevention of risks and promotion of innovation. In addition, it establishes AI-driven processes to increase the users’ ability to define in a more accurate way their offerings in edge computing environments as well as the data management and analytics concerning privacy protection. The work also designs the architecture in terms of data governance and analytics linked with the resource management on such dynamic environments . “User empowerment in health through Personal Health Information Management Systems” by Christos Magkos examines Personal Health Information Management Systems (PHIMS) as transformative tools in healthcare data management, addressing three critical challenges: individual data control, insight generation, and privacy preservation. The investigation reveals significant potential in two key technological domains: risk stratification for predictive healthcare and Large Language Models (LLMs) for data processing and diagnostic support. The study identifies and addresses fundamental challenges including data quality issues (missing data, selection bias, non-stationarity), LLM limitations (explainability, bias, privacy), and regulatory compliance requirements (GDPR, AI Act, MDR). Key findings demonstrate that successful PHIMS implementation requires robust privacy-by-design approaches, standardized data integration protocols, and patient-centric interface development. The research emphasizes interoperability and security as foundational requirements, while highlighting the necessity of inclusive design to prevent demographic marginalization. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 24 1. Introduction 1.3. Legal, technical and user-centred enablers of data sharing In her thesis titled “Privacy friendly AI: benefiting from AI without giving away privacy”, Soumia el Mestari elaborated the fine lines between the guarantees required by the EU data protection directives and regulations and the actual threats and defense strategies with the technical guarantees they offer. A detailed threat model that considers stakeholders, deployment choices, and different phases of the pipeline before and after using PETs was built. Another finding was the limitations of the current legal framework against some complex machine learning (ML) settings, namely the case of transfer learning of models, where the basic aims of enabling data subjects to have a flow governance over their data may be bypassed and lost. As a further step, this work investigated the superficial mapping between the meaning of legal anonymization and the technical tools named as ’anonymization algorithms’, which leads to a narrow application of these algorithms, a shallow attempt to achieve lawful processing and a wrong usage of PETs. From another angle, the issues of privacy in the machine learning system may be tricky to spot before an actual leakage of massive data is revealed. One of these attacks is membership inference attacks in federated learning with a malicious aggregator and under transfer learning settings for large language models, where the findings suggest that this kind of attack is hard to detect and defend against, especially if one wants to preserve privacy and still maintain a good model performance. 1.3 Legal, technical and user-centred enablers of data sharing The topic of data sharing is at the centre of this third and last section of the report. Barbara da Rosa Lazarotto in “Sharing information for the public good” focuses primarily on the topic of governmental access to privately held data. The work examines not only the traditional legal framework of data ownership but also how control over data could be enhanced through technological and policy-based mechanisms, ensuring that the benefits of the digital economy were more equitably shared. The research further investigates the tensions between data protection and data sharing regulations within the European Union’s legal framework, highlighting the dual objectives of data protection, which seeks to safeguard the rights and control of data subjects, and the regulations promoting data sharing such as the Data Act and the Data Governance Act, which aim to enhance the digital economy through the reuse and exchange of data. Furthermore, the research identifies the potential for legal fragmentation resulting from these competing regulatory frameworks. The study offers a nuanced analysis of the protective and limiting aspects of data subject rights, exploring the complex interplay between individual control and the broader goals of economic innovation. “A framework for user-centered, legal-ethical collective consent models: genomic data sharing” by Xengie Doan is a theoretical and empirical study on collective consent processes for health data sharing. First, the study explores the need for collective consent from a privacy, biomedical, and legal-ethical lens and analyzed the transparency and user-relevancy of policies from notable Direct to Consumer (DTC) genetic testing companies to find gaps in the status quo. Then the author reports findings from the testing of methods in the framework for transparent, user-centered collective consent with employees from a welfare technology company to better understand pros and cons in future implementation. A survey of adults about their attitudes toThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 25 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law inventions in the case of patent laws, secret information that has commercial value in the case of trade secrets2, or databases that have been subject to substantial investment3. IP laws are thus concerned with the recognition and protection of limited exclusionary rights regarding different categories of subject matter. [DP18] Their justification from a utilitarian perspective4thereby aims at facilitating the creation of efficient markets for works that due to their intangibility would otherwise be difficult to commercialize. First, copyright protection is conferred on subject matter that can be regarded as an authorial work that constitutes the author’s own intellectual creation.5Copyright protection thus extends only to the author’s original treatment of subject matter but not the underlying ideas or information themselves. According to the idea/expression dichotomy, protected subject matter can only subsist of the creative expression of ideas. [Gin18] Copyright has therefore not created any new property rights in data as such.6Second, databases can be subject to two kinds of intellectual property regimes. On the one hand, databases ‘which, by reason of the selection or arrangement of their contents, constitute the author’s own intellectual creation shall be protected as such by copyright.’7Since many databases do not fulfil this originality criterion, the European legislator created a sui generis database right if ‘there has been qualitatively and/or quantitatively a substantial investment in either the obtaining, verification or presentation of the contents of that database.’8Contrary to copyright, the focus from the sui generis right lies thus not on the protection of creations of the human mind but on results of investment. Individual data, however, are not protected.9 Third, trade secrets protect information that (i) is secret (ii) has commercial value due to its secrecy (iii) and that has been subject to reasonable steps to be kept secret.10 Trade secrets can therefore often offer protection for information that cannot be protected by conventional intellectual property rights, for instance, in the case of data-generating patents. The exclusionary protective effect can, however, disappear once information is no longer secret. Trade secrets do therefore not protect data as such, but they rather protect information against breaches of a confidential relationship. [Ree21] Finally, according to the European Patent Conventions patents can be granted for inventions that are new, involve an inventive step and are suscepti2Directive 2016/943/EU on the protection of undisclosed know-how and business information (trade secrets) against their unlawful acquisition, use and disclosure 3Directive 96/9/EC on the legal protection of databases 4For an overview on other approaches to intellectual property see [Fis01] 5Case C-5/08 Infopaq International A/S v Danske Dagblades Forening [2009] ECR I-6569 and case C-393/9 Bezpečnostní softwarová asociace - Svaz softwarové ochrany v Ministerstvo kultury [2010]. 6Data can, however, be protected if part of a creative work. European copyright legislation has therefore created, for instance, exceptions for text and data mining purposes for scientific research, see Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market. 7Directive 96/9/EC on the legal protection of databases, article 3(1). 8Directive 96/9/EC of the European Parliament and of the Council of 11 March 1996 on the legal protection of databases, article 7(1). 9See however [Wes21] who fears that the case law of the CJEU might have created an exceedingly wide protection of databases subject to the sui generis right that in some instances might come close to protecting data as such. But see also the recent Data Act which explicitly specifies that data obtained from or generated by the use of a product or a related service is not protected by the sui generis right. 10Article 2(1) Directive 2016/943 on the protection of undisclosed know-how and business information (trade secrets) against their unlawful acquisition, use and disclosure. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 32 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law ble of industrial application.11 Presentations of information as such, however, are excluded from patentability.12 Data as such would therefore not qualify for a patent since only “inventions” that fulfil the above-mentioned criteria are an eligible subject matter. 2.1.3.2 Creating Property Rights in Data to Empower Individuals and Facilitate the Emergence of Data Markets? If data is not subject to a comprehensive protection through intellectual property law, should new property rights in data be created to empower both individuals and companies to be in control of ‘their’ data? Would such a property rights regime facilitate the emergence of competitive data markets and thereby ensure the availability of data to a wide variety of stakeholders? In other words, could data property rights be an instrument that achieves both the protection as well as the availability of data? We already intuitively speak about ‘our data’, which expresses how we feel about information that relates to us. Since the traces which we leave when we browse the internet oftentimes reveal private and intimate information about us, such as our sexual orientation or political views when we ‘like’ certain content on social media, it makes sense to intuitively claim ownership over such data and to prevent others from using them. ‘Our’ data should belong to us. Such language thus reflects how the notion of ownership can be a psychological concept that expresses a sense of belonging. Similarly in economics, property rights have been defined as an individual’s ability to consume a good directly or to consume it indirectly through exchange. [Bar] This economic definition of property rights is therefore not concerned with what people are legally entitled to do but rather what they believe they can do. Property in law, however, denotes something else than psychological attachment to or de facto control over objects. As stated by Van Erp, the essence of property law concerns legal relationships between people with regard to objects, from which rights ensue that may be invoked against more than a number of specific persons and which therefore involves the normative creation of objects that qualify as property. [VE19] Property law therefore creates things as normative concepts that are then assigned to persons, and property rights determine the extent of the granted exclusive power. [Rah10] Property law thus raises the distributive question of why we treat some things as objects of property and why we deny the same treatment to others.13 Counter-intuitively, the question if data should be ‘treated’ as property is not a new question but pre-dates the era of big data or the internet. Already in 1968, decades before the collection of data both online and offline has become ubiquitous, authors argued that the right we as individuals have over our personal information should be understood as a property right. [Wes68] Debates, however, intensified with the internet era where the collection of large amounts of consumer data was increasingly getting facilitated through technological means. Lessig proposed to link the protection of data with the incentives of the market by using the 11European Patent Convention Article 52(1). 12European Patent Convention Article 52(2). 13[Pen97] One explanation, for instance, is offered by Katharina Pistor [Pis19] who argues that law constitutes a ‘powerful social ordering technology’ and that legal coding (e. g. via contract or property rights) is key for the commodification of objects, their subsequent exploitation as well as the distribution of wealth. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 33 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law laws of property as a control mechanism. [Les99b] If data could only be transferred with the owner’s consent, individuals would be put in the position of valuing their privacy according to their personal preferences. [Rev99] Such a commodification of data, which would turn it into an object that can be bought and sold on the market, would thus enhance privacy since individuals would be in full control over their data. Such a property rights approach to data, however, has been criticized by different authors. Recognizing property rights in data might ultimately erode existing levels of privacy if individuals simply increasingly traded away their personal data. [Coh00] Furthermore, as the value of data is linked to uncertain future uses it would be almost impossible for individuals to ascribe value to information. Aggregated data and inferences that can be drawn from it further complicate this task since a ‘comprehensive collection of data about an individual is vastly more than the sum of its parts’. [Coh00] In such cases where trades are characterized by information asymmetries, property rules alone have thus been regarded as suboptimal to achieve efficient outcomes. Market solutions based on a property rights model would therefore not cure any of the problems related to control, but they would only legitimize them. [Lit00] A property rights approach has in particular been regarded as problematic in Europe where information privacy is seen as a fundamental right.14 Consequently, propertizing these rights would be ‘contrary to the theory, history, and practice of European information privacy and would require a concerted effort of dramatic proportions.’ [MS10] Other authors, however, remind us that although the EU data protection framework is conceived in fundamental rights terms, it equally functions to ensure the free flow of personal data. Similar to a property rights regime, EU data protection rules would provide a framework that legitimizes rather than discourages trade in personal data. [Lyn15] Furthermore, the status quo devoid of well-defined property rights in data would already result in a de facto assignment of economic property rights to the information industry, eroding data subject’s autonomy, privacy and informational self-determination. [Pur15] The debates on data property culminated when the European Commission proposed the creation of a ‘data producer’s right’ in non-personal data.15 Criticism on the creation of a new property right in data, however, prevailed. Since data would already be subject to transactions and businesses would have technical means to shield data against competitors and reach factual exclusivity, a contractual inter partes protection would suffice. [Dre17] Since Big Data analytics would rely on permanent access to real time data sources, individual ownership rights in data would present major challenges in terms of governance and enforcement of individual rights. [Dre17] Creating a new layer of rights in data would furthermore cause disruptive overlaps with both copyright and the sui generis databases right. [Hug18] The consequence would be multiple competing claims of ownership over the same content. Such a “tragedy of the anticommons” would hence risk resulting in in an underuse of data due to the high number and complex in14On the history of the fundamental right to data protection within the EU see e. g. [GF14] On the distinction between privacy and data protection see e. g. [DHG06b] For an overview how to conceptualize the link between privacy and data protection see [Lyn15] who distinguishes between s (i) data protection and privacy as separate but complementary rights (ii) data protection as a subset of the right to privacy (iii) data protection as an independent right that serves a multitude of purposes. 15European Commission, ‘Building a European Data Economy’ (2017) SWD(2017) 2 final. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 34 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law terrelations of overlapping property rights. [Ste20] Since data constitutes a non-rivalrous but at the same time excludable resource that might suffer from a contextual incentive problem for the production of certain types of data (raw versus inferred data), it would constitute a difficult object for traditional legal categories like property. [Ker17] 2.1.3.3 Abandoning the Idea of Property Rights and Moving Towards Data Access and Data Sharing Some years after the proposal of a “data producer’s right”, however, such an incentive-oriented approach has been abandoned. Neither the 2020 European Strategy for Data16 nor two of its key actions, the 2022 Data Governance Act (DGA)17 and the 2023 European Data Act18, contain any references to data ownership or sui generis rights in data. Whilst the goal of the European Strategy for Data remains to realize a ’genuine single market for data’, any reference to (legal) data ownership rights disappeared from the Commission’s communications and proposals. Commentators hence described the strategy as a ’paradigm shift’ in EU policy from data ownership towards facilitating data access and sharing. [MA21b] Since data would constitute an essential resource for companies to develop new products, services or train AI systems, a regulatory environment should enable stakeholders to have easy access to an almost infinite amount of data. Regulation should enable data to flow easily while at the same time ensure that European rules and values are fully respected. Key pillars of the European Strategy for Data thereby attempt to enable precisely the creation of such a regulatory environment. The DGA intends to facilitate access and encourage use of data held by public sector bodies. Whilst the Open Data Directive19 already established minimum rules governing re-use for data that is not subject to intellectual property rights or commercial confidentiality, the DGA harmonizes access conditions for re-use of protected categories of data. Moreover, the DGA puts in place a framework for ‘data intermediation services’ and ‘data altruism organizations’ which should play a key role in the data economy by supporting, facilitating, and promoting (voluntary) data sharing. Furthermore, the Commission intends to create Common European Data Spaces20 in strategic sectoral fields21 that bring together relevant data infrastructures and governance frameworks to facilitate data pooling and sharing. The new Data Spaces are strongly connected with the data intermediaries by the DGA which, as trustworthy, neutral data intermediaries, should facilitate the emergence of data spaces. Third, the Data Act constitutes together with the DGA the second horizontal legislative instrument of the European Data Strategy. Whilst the DGA aims at opening up data held by public sector bodies, the Data Act creates mandatory access rules to data held by private actors in B2C, B2B as well as B2G relations. The data access rights are thereby supposed to fulfil the twofold purpose of empowering individuals and companies over ‘their’ data while at the same time ensuring that data is available for other stakeholders. 16European Commission, ‘A European Strategy for Data’ (2020) COM(2020) 66 final. 17Regulation 2022/868 on European data governance (Data Governance Act). 18Regulation 2023/2854 on harmonised rules on fair access to and use of data (Data Act). 19Directive 2019/1024 on open data and the re-use of public sector information. 20European Commission, ‘Common European Data Spaces’ (2022) SWD(2022) 45 final. 21European Commission. 2022. Proposal for a Regulation on European Health Data Space. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 35 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law Finally, this paradigm shift towards fostering data access is not only visible in the European Strategy for Data but has also been observed in other new European digital regulations. For instance, provisions in the Digital Markets Act22 (Article 6(9)) have introduced obligations for gatekeepers to allow users to port data they provided or generated by their use of the platform’s services. Noto la Diega highlights that this would reflect the clear policy stance in favor of openness via portability and access and against proprietary considerations. [NLD23] With regard to the Digital Service Act (DSA)23 he remarks that it would leave out consumers by not doing much about private remedies. However, it would still provide incentives for opening up enclosed data by imposing obligations on intermediary service providers to, for instance, give access to data that are necessary to monitor and assess compliance with the DSA (article 31), or to provide users with information that they are being presented with advertisement, who is behind it, and which parameters were used to target them (article 24). The legislative initiatives by the EU of the past four years therefore reflect the EU’s attempt to open up data while at the same time finding a balance between new access and transparency obligations and competing interests in data. New access and/or portability obligations in the DMA, DSA, and the AI Act would therefore equally constitute part of the latest regulatory effort of the EU to end data enclosure while departing from a ‘single-minded rights-based and data ownership approach’. [NLD23] At the same time, however, access and data sharing are equally conflicting with privacy and economic interests of stakeholders. Creating data access obligations might be in conflict with, for instance, motivations by companies to keep their commercial information secret or their intellectual property rights protected. Similarly, consumer might have an interest to allow third parties access their data while simultaneously remain in control of it. This challenge of allowing access to data whilst at the same time protect the interests of ‘data holders’ (whether companies or individuals) further informed interdisciplinary research I conducted with two other ESRs from the LeADS project which resulted in a paper that was presented at the 2023 ACM Conference on Fairness, Accountability, and Transparency in Chicago. [Zuz+23] 2.1.3.4 Facilitating Data Access and Data Sharing Through Technology: An Interdisciplinary Perspective on Data Governance The starting point of our reflection was the evolution in EU data governance from property rights towards facilitating data access and data sharing. With a property-rights approach towards the regulation of data being discarded, we looked at literature on the Commons. Scholars on the commons have argued that an alternative approach to resource regulation, one that does not rely on property rights to prevent the tragedy of the commons24, is possible.". [Ost90] Since commons are based on the idea of access rather than exclusion, they would challenge the dom22Regulation (EU) 2022/1925 on Contestable and Fair Markets in the Digital Sector and Amending Directives (EU) 2019/1937 and (EU) 2020/1828 (Digital Markets Act). 23Regulation 2022/2065 on a Single Market For Digital Services (Digital Services Act). 24A famous example for an explanation of the necessity of property rights in society constitutes the ‘tragedy of the commons’ elaborated by Garrett Hardin, ‘The Tragedy of the Commons’ (1968) 162 Science In a society without property rights, scarce ressources would ultimately end up being overexploited and depleted: freedom in a commons would bring ruin to all. To avoid this tragedy, private property rights would enable efficient ressource allocation through the market. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 36 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law inance of private property [Mar17] or even constitute the opposite of property [MQ17]. The main function of commons would therefore be to institutionalize freedom to operate within symmetric constraints, free of the particular risk that any other can deny use of that resource set. [Ben14] The commons-framework therefore does not imply unmanaged access to a particular resource, but requires that groups engage in managed resource sharing. [Mad20] Whilst initially conceived for the management of natural resources whose characteristics are evidently very distinct from data, they would still be highly relevant with regard to data governance as they could provide insights for assessing the types of institutions needed to govern access and use of data. [MFS10] How data could be subject to regulation via commons governance institutions has been subsequently further analyzed within research on governing knowledge commons that have been defined as ”institutionalized community governance of the sharing and, in some cases, creation, of information, science, knowledge, data, and other types of intellectual and cultural resources”. [FMS14] Its key insight would be how information resources are governed as a shared resource via a collective and not in markets via intellectual property rights regimes or state intervention. [Mad20] New technologies thereby would have played a key role for the possible development of knowledge commons as they facilitated the ability to capture the previously uncapturable and to draw value from it by means of advanced data analytics. [Pur17] Studies therefore attempted to adapt Ostrom’s design principles for collective self-managing institutionsandtranslatethem toand extract usefulinsightsfordata governance[Ins21b] [Coy+20] [Pra19]. This further resulted in discussions and analyses [Ins21b] [WHB22], on various models of data stewardships, such as data trusts, data foundations, data cooperatives or data collaboratives [WHB22]. Data collaboratives thereby have been described as new emerging forms of partnerships where privately held data is made accessible for analysis. The collaboration between participants that is facilitated within these structures aims to result in new insights and innovation and to unlock the public good potential of previously siloed data. [Gov20] Building on that research, we analyzed how data collaboratives based on an application of decentralized learning could provide a substantial leap towards independent and self-sovereign data management inside trusted communities. By looking at both, the regulatory and the technological layer in data governance, we analyzed how legal obligations and policy objectives can be realized by technological solutions (decentralized learning). This interdisciplinary, i. e. legality attentive data scientists perspective, thus allowed us to better understand the importance of cooperating across academic fields to find legally compliant solutions to regulatory complexities. 2.1.3.5 A New Role for Consumer Law and Policy to Empower Consumers with Regard to Their Data The data access and sharing obligations created by the various legislative instruments over the past years by the EU, however, are not as comprehensive in scope as a property rights regime for data would have been. Whereas the Data Act, for instance, certainly creates new data access rights for consumers and might thus empower them to a certain extent with regard to ‘their’ data, it is in no way comparable to the degree of control new property rights would have conferred to them. [Hin22] A central motivation behind my LeADS research topic, however, related The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 37 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law to the question how to empower individuals with regard to their data – beyond the regulatory framework created by European Data protection law. For my PhD research I thus asked myself how the law, if not through property rights, can protect and potentially empower consumers with regard to their data. In my PhD thesis I therefore conduct the first comprehensive analysis to what extent data has become regulatory subject matter in European consumer law. Whereas data and consumer protection share the commonality that both intend to protect ‘weaker parties’ (the ‘consumer’ and the ‘data subject’) both were separate fields of law until 10 years ago. Contrary to data protection law, the consumer law acquis was firmly rooted in the analogue world and slow to adapt to legal challenges by digitisation [Hel+13b]. Data has traditionally not been subject matter that belonged to the regulatory ambit of consumer law. This, however, has changed gradually over the past decade. Today, the regulation of data is spread over various legal disciplines, with data protection law forming its core. [Str21] The EU legislator is thus confronted with the challenging task of constructing different dimensions of data law without infringing its core, i. e. to coordinate each dimension with the ‘boundary setting instrument’ of the GDPR. [DH23] Changes in the marketplace induced by the development and rapid diffusion of information and communication technology affected both the array of goods and services available and the way consumers buy these products. [OEC10] The datafication of consumer markets would go hand in hand with the digitalisation process. [Mak18] Whilst consumers benefited from these developments it equally gave rise to new complexities and challenges. Earlier accounts in EU policy documents acknowledged that electronic commerce would have the potential to most profoundly change the relationship between businesses and consumers and the nature of consumption itself.25 However, at first only the lack of harmonisation of national laws was identified as a principal problem. The resulting legal uncertainties would prevent businesses and consumers from fully taking advantage of the potential of e-commerce in the internal market26. This, however, changed with the growing awareness of the importance of data for economic growth. Scholars observed how the European legislator intends to expand the internal market’s four existing freedoms (free movement of goods, services, capital, and labour) with a fifth freedom: the free flow of data. [CM22] Both national27 as well as supranational institutions28 increasingly recognised the importance of data for an effective consumer protection online, and that data protection and privacy rules should be included in future consumer strategy programmes. Furthermore, with the advent of Big Data in EU policy and expectations attached to data-driven innovation, data was increasingly getting framed as a new crucial economic resource (“the new oil”, “the fuel of a new industrial revolution”, a “goldmine”) to be mined and 25European Commission, ‘Consumer Policy Action Plan 1999-2001’ (1998) COM(1998) 696 final 3. 26European Commission, ‘Green Paper on European Union Consumer Protection’ (2001) COM(2001) 531 final. 27The German Federal Ministry for Food, Agriculture, and Consumer Protection, ‘Charta on Consumer Sovereignty intheDigital World’ (2007) stressed growing theimportanceof data protection for consumerpolicyand that ensuring fair data processing would constitute a competitive advantage for companies. 28In 2008 the European Parliament adopted a resolution on the EU consumer policy strategy 2007-2013, pointing out that in the digital environment personal data would have become a “trade product as well as an ingredient of commercial methods, for example behavioural targeting” and that data protection and privacy rules should therefore be included in any consumer strategy, see European Parliament, ‘Resolution of 20 May 2008 on EU Consumer Policy Strategy 2007-2013’ 2009/C 279 E/04. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 38 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law exploited and to thus “turn this asset into gold”. [Rie18] Since consumers were at the origin of this valuable resource in the online environment, it was subsequently observed that data had become a new currency with which consumers are paying ‘free’ services. This, in turn, would engage fundamental rights of consumers and would ultimately raise questions on the relationship between constitutional and contract law [Mak13]. Furthermore, if conceived as a currency, it could potentially trigger the application of various consumer law instruments related to, for instance, unfair commercial practices (services advertised as “free”) or contract law (unfair contract terms). [Slo13] In a first attempt to address these new challenges and to facilitate the emergence of a vibrant digital single market29, the European Commission proposed a Regulation on a Common European Sales Law.30 To ensure that also ‘free’ digital content or services would fall under the scope of this instrument, it stipulated that it equally applied where consumers provided non-monetary consideration such as giving access to personal data. In addition to the growing number of scholars who recognised that the protection of personal data would constitute an integral part of consumer protection [Hel+13a], the European Data Protection Supervisor (EDPS) highlighted their complementarity. In 2014 the EDPS published a preliminary opinion on privacy and competitiveness in the age of big data where it highlighted the interplay between data protection, competition law and consumer protection in the digital economy. All three policy areas would converge around a twofold purpose: promote a thriving internal market while at the same time offering protection to the individual. An adequate regulatory response to technological change where ‘big personal data’ would have become a company asset and a currency for consumers to purchase ‘free’ services would have to rely on the three distinct policy areas. Consumer welfare would not only be determined by price but also other factors such as quality and consumer choice – both would constitute relevant concerns for data protection. At the same time, the 2014 EDPS opinion continued, consumer welfare might be at risk where control over one’s personal information is restricted. Consumers would be blinded by deceptive advertising offers that promote services as ‘free’ despite relying on the provision of personal data. The various consumer law measures that focus on the provision of clear information about the cost and value of services would mirror the importance of transparency in data protection law and to provide information on data processing in an intelligible way. Consumer law’s concern for product safety would complement the general risk-based approach of the GDPR, its impact assessment and in general the principle of accountability. Privacy and data protection were thus described as factors to evaluate the impact by business activities on consumer welfare – both would have become entangled as part of the economic interests of consumers. A lack of interaction between consumer and data protection policy in an economy where data represents a significant intangible asset might ultimately put consumer welfare at risk and fail to set incentives for the development of privacy-enhancing services. My PhD research thus concerned the central research question to what extent data has become regulatory subject matter in European consumer law which includes in particular the following 29European Commission, ‘A Digital Agenda for Europe’ (2010) COM(2010)245 final 7. 30‘Proposal for a Regulation on a Common European Sales Law’ (2011) COM/2011/0635 final. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 39 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law sub-questions: •How has data become regulatory subject matter in European consumer law? •How does the regulation of data under European consumer law complement and or/conflict with data protection law in reaching its twofold objective (i. e. the free flow of data and the fundamental right to data protection) •What does the regulation of data under European consumer law tell us about the general regulatory approach of the EU with regard to data and about the legal conceptualisation of data in EU law? These questions are answered through several chapters in my PhD which focus in particular on (i) how consumer law and policy developed a concern for the regulation of data over the past decades (ii) on legal anthropology and the difference between protecting ‘data subjects’ versus protecting ‘consumers’ (iii) the constitutionalization of the right to data protection and the principle of consumer protection in EU law (iv) the analysis how various EU consumer law instruments shape the regulation of data, such as the 2019 Digital Content and Services Directive , the 2005 Unfair Commercial Practices Directive , or the 1993 Unfair Contract Terms Directive and and (v) an analysis of the position of the consumer in the new data laws of the EU (such as the DA, DGA or DMA) . Whilst I am still in the writing process of my PhD thesis, some of my findings exemplify the regulatory complexity of the data economy – a core theme of the LeADS project. Whilst consumer law certainly has the potential of protecting consumers with regard to their data, it equally conflicts with the twofold objective of European data protection law: (i) it not only protects but equally risks polluting the privacy ecosystem [Hin23] (ii) it bears the risk of undermining the objective of European data protection law of creating a fully harmonised legal framework that facilitates the free flow of data [Hin24]. 2.1.3.6 Regulating Data Through Consumer Law: Protecting or Polluting the Privacy Ecosystem The latest EU consumer agenda identified the digital transformation as one of its key areas and states that “online” consumers should be protected on a comparable level as “offline” consumers.31 Consequently, the impact consumer law has on the regulation of data is likely going to increase. The perspective I offer in my research should thereby further the understanding how consumer policy interacts with data protection policy. By analysing the growing importance consumer law has for the regulation and protection of consumer data my research shows how consumer law may complement as well as conflict with data protection law. Whilst consumer policy is anchored in economic market integration, the regulation of data has traditionally been anchored in data protection law which is underlined by a fundamental right rationale. Whereas it has become clear that both policy areas no longer constitute separate monoliths and are often complementary in protecting consumer data, this does not imply that their pursued objectives might not conflict with each other at times. 31European Commission, ‘New Consumer Agenda: Strengthening Consumer Resilience for Sustainable Recovery’ (2020) COM(2020) 696 final. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 40 2. Data Protection, Consumer Protection and Competition Law 2.1. Data as Property and Regulatory Subject Matter in European Consumer Law In my research I disagree with authors who claimed that data and consumer protection might clash in their objective of protecting the weaker party, i.e. when consumer law offers stronger protection to consumer data than provided by the GDPR. [Sva18] Instead, it provides a new perspective by transposing tensions that exist between consumer and environmental law to the relationship between consumer and data protection law.32 Objectives in line with the economic interest of the consumer (low prices or access to goods and services) risk not only polluting the environment, but ultimately also the privacy ecosystem. Some protective dimensions of consumer policy thus clash with data protection policy when, for instance, consumer law tends to enable data (“counter-performance”) to be commodified in consumer contracts that package pieces of data into pieces of trade. Mirroring the concept of sustainable consumption which functions as a bridge between consumer and environmental policy, it exemplified that these tensions could at least in part be overcome by stronger integrating the concept of privacy by design and by default into the consumer law acquis and thereby strengthening the ties between both policy areas. 2.1.3.7 Consumer Law and the Free Flow of Data: Upsetting the Balance of the European Data Protection Framework My research further exemplifies how consumer law instruments, such as the 2019 Digital Content and Services Directive (DCDS) and the 1993 Unfair Contract Terms Directive (UCTD), interact with the free flow of both personal and non-personal data. In case studies on consumer directives, I demonstrate how consumer law potentially undermines the objective of data protection law to create a fully harmonized legal framework. Whilst both fields of law might often complement each other in the protection of consumer privacy, the growing intersections between consumer contract and data protection law bear the risk of leading to legal fragmentation. First, different legal traditions complicate the endeavor of harmonizing core areas of contract law, which was exemplified, for instance, by the failure of the proposal for a regulation that should create a common European Sales Law. Thus, when consumer law provisions that impact the free flow of personal data are interpreted (e. g. with regard to the unfairness assessments under the 1993 Unfair Contract Terms Directive or the Unfair Commercial Practices Directive) or implemented (e. g. with regard to the scope of the 2019 Digital Content Directive) differently on the national level, the results are heterogeneous levels of protection. Furthermore, they sometimes directly interfere with the interpretation of GDPR provisions when, for instance, the 2019 Digital Content Directive stipulates that national law should regulate contractual consequences in cases where consumers withdraw consent for the processing of personal data. Here, the result are again diverging solutions in national contract law in an area that belongs to the core of European data protection law (legal bases for the processing of personal data) and therefore should have been exhaustively harmonized on the European level through the GDPR. 32On the link and tensions between environmental and consumer policy see e. g. Klaus Tonner, ‘Consumer Protection and Environmental Protection: Contradictions and Suggested Steps Towards Integration’ (2000) 23 Journal of Consumer Policy 63; Ludwig Krämer, ‘On the Interrelation between Consumer and Environmental Policies in the European Community’ in Norbert Reich and Geoffrey Woodroffe (eds), European Consumer Policy after Maastricht (Springer 1994). The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 41 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition media market can exemplify the typical characteristics necessary for examining the interplay between personal data protection and market dominance, such as the multitude of platform providers competing for user attention and data, and the prevalence of user engagement and targeted advertising in this ecosystem [Oca;PD]. Data is collected from each EU member state and 32 other countries over the period 2009-2022 to explore the impact of the GDPR as an EU-level regulation on the market concentration of social media in the EU as well as its impact intensity. 2.2.5 Research findings and analysis 2.2.5.1 Personal Data as the Foundation of Online Platforms Online platforms derive revenue by providing content and products to the user side or by assisting third parties to reach the user side. When personal data is utilised for optimising the platform’s own services and all-sides users’ experience, it is generally considered as internal usage like personalizing content, improving algorithms, or enhancing user interactions [van+]. As more and more individual users join an online platform, the platform collects more personal data and analyses that data to improve the individual users’ experience, which in turn attracts more individual users, thus creating a positive loop [PS]. Internal data usage plays an important role in fostering network effects, a critical component of online platform success [BGL;Gaw; PNF;OECb;PPVA;Per;PS;Sch]. By contrast, external usage involves leveraging the collected data to provide services or insights to third parties, such as through targeted advertising or data analytics services [van+]. For instance, online platforms mainly based on the advertising revenue model need a larger user base and more refined data analytics capabilities, enabling the platform to offer more effective advertising solutions and attract more advertisers [PV]. Such online platforms capture user-generated data and de facto monetise it for their huge profits and dominant positions [Gaw]. Given the two (multiple) sides of the online platform, both uses exist and can occur at the same time in the online platform market. Fast et al. categorised the business value creation channels from user data into three domains, i.e., content and service personalisation, recommender systems, and targeted advertising. Content and service personalisation Content and service personalisation is to deliver “the right content to the right person at the right time” for maximising business opportunities [TH]. The abundance of personal data and better data quality can make online platforms easily build individual user profiles that approximate users’ interests and preferences [TMJ], and automatically tailor products for consumers’ needs and interests [KKR;SCS]. The large amounts of data can also be used to improve Artificial Intelligence, machine learning, deep learning, or Big Data Analytics [KD], and individual users’ data can be also analysed to explore patterns and trends of the market [de +;SL;SKa]. By some empirical studies, content and service personalisation can attract users to spend more time on their platform and even increase their willingness to pay [Ben]. The fulfilment of personalised needs will enhance perceived user experience [Lie+] and boost customer loyalty in different products or services markets [TS]. Content and service personalisation increases switching costs and improves customer loyalty, making it more difficult for competitors to capture customers The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 48 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition in online markets [AM]. Recommender systems A recommender system is a type of information filtering system that provides the most relevant suggestions from a potentially overwhelming number of items for a specific user [RRS]. Recommender systems can highlight the products or content that align with the preferences of targeted users and reduce search costs, based on users’ previous interactions [LGP]. Furthermore, recommendations are likely to boost demand and increase consumption by providing relevant products or complementary products when individual users browse an item or have already purchased an item [HL;Pat+]. The empirical studies from Adomavicius et al. also show that recommender systems can change individual users’ preferences and potentially increase their willingness to pay [Ado+a;Ado+b]. In a sense, the recommender system can switch costs and improve customer loyalty, and the platform will generate revenue by charging fees to thirdparty merchants or users based on the recommender system. Online Advertising The collection and analysis of personal data is almost the most critical factor separating the online advertising market from the traditional advertising market [GT]. Large amounts of low or no-cost personal data, location tracking, and other technological developments make it possible to send different personalized advertisements to different target groups. This behavioural targeting can help reduce companies’ costs and increase the probability of successful transactions, and related empirical studies confirm this value that behavioural targeting advertisements are 2.68 times more expensive than untargeted advertisements [ATb;Bea]. Farahat and Bailey estimate that the revenue generated by targeted advertisements was, on average, twice as high as that generated by untargeted advertisements [FB]. Goldfarb and Tucker revealed that with clickstream data showing the extent of a consumer’s product search, the performance of behavioural targeting can be improved, and they further show that the performance of behavioural targeting can be improved when combined with clickstream data that help to identify the consumer’s degree of product search [GT]. Targeting also lowers the cost of searching for consumers and improves the match between consumers and businesses [Cor]. The analysis and prediction of consumer choice and behaviour trends can assist companies in managing their supply chains more effectively and enhancing operational efficiency of companies [CV;GT]. Fast et al.(2023) categorise data-driven business value created by users’ data into three main categories: (1) improved customer retention, (2) increased revenue in the consumer market, and (3) increased revenue on other market sides [FSW]. 2.2.5.2 Impacts of the GDPR processing principles on data controllers under the static model Compared to principles relating to data quality in Directive 95/46, the main improvements in principles relating to the processing of personal data under the GDPR are in the following areas: (1) higher transparency and security requirements for the processing of personal data (Article 5(1)(a) and Article 5(1)(f)); (2) stricter limitations on the scope of personal data involved in data processing (Article 5(1)(b), Article 5(1)(c), Article 5(1)(d) and Article 5(1)(e)); (3) stronger accountability requirements for data controllers (Article 5(2)). To bridge the uncertainty created by the The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 49 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition Figure 2.1: Data-driven business value created by personal data principles, the GDPR also provides more detailed rules to clarify the obligations of controllers. During the pre-processing phase, the controller must plan and implement appropriate technical and organisational measures to ensure compliance with the GDPR (GDPR Article 24), consider data security from the outset (e.g. encryption, access control), and ensure that by default only necessary data is processed (GDPR Article 25). For those processes with a high risk to the rights and freedoms of natural persons, controllers must carry out data protection impact assessments or consult with the supervisory authority in order to sufficiently identify and mitigate security risks (GDPR Articles 35-36). Article 12 in the GDPR sets out the transparency rules of the information or communication: controllers should provide any information referred to the provision of information to data subjects (Articles 13-14), communications with data subjects concerning the exercise of their rights (Articles 15-22), and communications in relation to data breaches (Article 34) to the data subject in a concise, transparent, intelligible and easily accessible form, using clear and plain language, without undue delay and unreasonable fares. When processing personal data, controllers and processors must keep documentation of the processing activities (GDPR Article 30), implement appropriate technical and organisational measures (e.g. pseudonymisation and encryption) (GDPR Article 32), and promote the adoption of codes of conduct (GDPR Article 40) and authentication mechanisms (GDPR Article 42). Throughout this process, the controller must continue to monitor and assess the security measures to ensure the security level appropriate to the risk (GDPR Article 32), and notify the relevant supervisory authority and the data subject without undue delay in the case of a personal data breach (GDPR Article 33-34). And the DPO is responsible for overseeing compliance with The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 50 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition GDPR, including monitoring security measures, responding to data breaches, and advising the controller or processor on risk management (Article 39). Figure 2.2: Impacts of the GDPR processing principles on online platform providers For online platform companies as data controllers, the principles of transparency, security, accountability and their related obligations can bring about higher costs than before in theory, which may manifest as one-time costs of compliance (e.g. the cost of change to a new system or reshape the previous system) and ongoing costs of maintaining compliance (e.g. the cost of maintaining a competent system) [LYH]. According to the PricewaterhouseCoopers survey, 71 percent of survey respondents focus on information security, privacy policies, GDPR gap assessment and data discovery, and the cost of compliance is likely to exceed 1 million US dollars [Pri]. The strict limitation of the personal data scopes is most likely to have a direct negative impact on the amount of user data available to the data controller [GJS]. A pure reduction in data volume can deliver lower costs of collection, processing and storage, but the reduced data volume in the GDPR is embedded in the processing design of data minimisation, which may create some technical and administrative costs [LEC]. At the same time, the shrinking data volume may have weakened the company’s profitability somewhat [GA;Jan+]. Considering the online platform service based on personal data, the total provider revenue and total provider cost in this market are determined by two key factors: total personal data scale S The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 51 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition and technology level t, where t∈(0,1) represents the technological efficiency in utilizing personal data. In the baseline model, assume that total provider cost is driven by two components: fixed costs and variable costs that depend on the scale of personal data. The variable costs of processing personal data will rise as the personal data scale increases. Let the total provider cost function, C(S), be defined as: C(S)=C0+ (αS +f1(S))(1 −t) where C0represents fixed costs, αS represents the variable cost that grows linearly with the personal data scale S,α > 0is the cost coefficient in this linear function, and f1(S)reflects the variable cost of a nonlinear growth function with increasing marginal cost as personal data scale expands. In general, users on the online platform do not pay companies real money for their services, but instead provide their personal data. For companies, personal data is typically characterized by diminishing returns, i.e. the decrease in marginal benefit as the personal data scale is incrementally increased, like any other factor of production (Agrawal et al., 2019). The total revenue of the companies is given by: R(S)=tf2(S) where f2(S)is a nonlinear growth function with diminishing marginal benefit as personal data scale Sincreases. The marginal cost is given by: MC =dC dS = (1 −t)α+df1(S) dS  which indicates that the marginal cost is influenced by both the variable cost per unit of data (α) and the nonlinear growth function f1(S)of costs with the impact of technology. Similarly, the marginal revenue is derived from the revenue function: MR =dR dS =tdf2(S) dS which shows that marginal revenue depends on how much additional revenue the companies can extract from increasing their personal data scale, weighted by the technological efficiency in utilizing personal data (t). Theoptimal personaldata scale of productionS∗ pre isdeterminedbythe conditionwheremarginal cost equals marginal revenue: (1 −t)α+df1(S) dS =tdf2(S) dS After the GDPR, two types of compliance costs have been introduced into the cost structure: fixed compliance cost C1representing a one-time cost for GDPR compliance, and the variable compliance cost C2=βS +f3(S), where β > 0is the cost coefficient and f3(S)reflects The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 52 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition a nonlinear growth function with increasing marginal compliance cost as personal data scale grows. The new total cost function becomes: Cnew =C0+C1+βS +f3(S)+(αS +f1(S))(1 −t) And the new marginal cost function is: MCnew =dCnew dS =β+df3(S) dS +(1−t)α+df1(S) dS  The new optimal personal data scale of production S∗ post after the GDPR is given by: β+df3(S) dS +(1−t)α+df1(S) dS =tdf2(S) dS When comparing S∗ pre and S∗ post, the marginal benefit function tdf2(S) dS is consistent in both equations. However, the left-hand side of the equation for S∗ post includes an additional positive function β+df3(S) dS >0, indicating that the GDPR has introduced extra costs that increase marginal costs. Assuming that S∗ post ≥S∗ pre,df1(S) dS >0and d2f1(S) dS2>0, because f1(S)is a nonlinear growth function with increasing marginal cost. f1(S)is a convex function, so the line segment between any two distinct points on the function lies above or on the function curve, which implies: df1(S∗ post) dS ≥df1(S∗ pre) dS Since β+df3(S) dS >0, the following inequality can be obtained: β+df3(S∗ post) dS +(1−t)α+df1(S∗ post) dS >(1 −t)α+df1(S∗ pre) dS  df2(S∗ post) dS >df2(S∗ pre) dS The above inequality generates a contradiction with the inequality df2(S∗ post) dS ≤df2(S∗ pre) dS given by f2(S). Since f2(S)is a nonlinear growth function with diminishing marginal benefit, df2(S) dS >0 and d2f2(S) dS2<0. As f2(S)is a concave function, the line segment between any two distinct points on the function lies below or on the function curve. Given S∗ post ≥S∗ pre, we have: df2(S∗ post) dS ≤df2(S∗ pre) dS Thus, the increase in the compliance costs of the GDPR implies that the optimal data scale S∗ post must decrease, i.e. S∗ post < S∗ pre, which is also confirmed in the empirical studies in 2.2.5.3. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 53 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition 2.2.5.3 The Impact of the GDPR on Social Media Market Concentration in the EU The synthetic control method is employed to identify and quantify the causal effect of the GDPR reform on the social media market concentration at the EU level and EU average level. The results at the EU level and at the EU average level show the GDPR reform may have an ex-ante effect caused by the legislative process with voting and publication drafts, which is evidenced by a downward trend from 2015. The negative impact of the GDPR on the EU social media market reaches a maximum in 2017, when the CR8(social media) indexes decreased by 0.252 at the EU level and by 0.383 at the EU average level. Figure 2.3: CR8 trends in the social media market at the EU level and EU average This decline in social media market concentration may be attributed to the following reasons: (1) Transparency and accountability in the GDPR make it harder for social media platforms to misuse user data, which gives users more control over their data and force platforms to consider the scale of data collection based on costs. (2) Data portability provisions allow users to easily move data from one platform to another, further breaking down data exclusivity and market concentration. (3) A strong data protection regulation can impose compliance costs on dominant platforms, which may level the competitive landscape as these platforms cannot leverage their data collection advantages as freely. When dominant platforms are restricted in their data processing practices, new entrants can compete more effectively without the limitations of exclusive data to offer competitive services. After 2020, the negative impact of the GDPR on EU social media market concentration gradually diminishes. The latent reason is that leading companies have rebuilt their competitiveness through strategic adjustment to cope with the negative influence of the GDPR. Within the EU, The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 54 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition the results based on generalised DID estimator show that the GDPR significantly reduces market concentration within the social media sector across EU member states. Its impact on large companies in the EU social media market is more pronounced than on smaller companies. In this case, a 1 percent increase in the GDPR fines results in an approximate 0.216 reduction in market concentration among the three largest companies and a 0.0143 reduction for the eight largest companies. This difference may result from the heterogeneity in market share and technological levels, so this study further explores the moderating effect of the Internet market scale and high technologies between the GDPR and the social media market concentration in each EU member state. In social media markets, specifically, the Internet market scale to some extent enhances the impact of the GDPR because the compliance costs imposed by the GDPR clearly have a theoretical positive correlation with the scale and complexity of the market. With larger companies being hit more, SMEs may benefit from the reduced dominance of these larger firms. In this case, the large number of new entrants in the social media market may strongly shock the market concentration of the leading companies. Meanwhile, higher technological levels will likely mitigate the negative impact of the GDPR on smaller companies rather than the top companies. The technological advantage can help companies in the social media market to adapt to new regulatory systems faster and at a lower cost. Besides, the large companies with technological advantages in the social media market will likely come under strong attention from the DPAs, The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 55 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition thereby decreasing their market concentration more obviously than smaller companies. In-space placebo tests use randomly selected fake treatment units and in-time placebo tests employ fake treatment times. In the in-time placebo test, the adoption of GDPR (i.e., 2016) is shifted forward by one to four years (i.e., from 2015 to 2012) to create four false GDPR adoption periods. The impact of the GDPR on social media market concentration is then re-estimated. The estimated results are shown in 2.4 and 2.5.After including control variables, the in-time placebo test results show that the estimated coefficients are all nonsignificant at the 5 percent level, which indicates that the impact of GDPR on social media market concentrations in the EU is not attributable to potential effects generated by other policies. Figure 2.4: Results of the in-time placebo test (without control variables) The fake GDPR fines are randomly assigned to EU states, and regression is run to simulate random interference of the GDPR on social media market concentration. Iterating the above process 500 times yields a density plot of tin-space placebo estimated coefficients, as shown in 2.6. The figure shows that the true estimated coefficients are very far from the fake coefficients, which suggests that the effect of the GDPR on social media market concentration is not affected by unobservable random factors. 2.2.6 Conclusions The value creation from online platforms comes in the forms of communication, information, matching, and choice or diversity [Eur], and the network effects service as unique boosters for the process of their value creation. The existence of online platforms is finally to assist with economic decisions by participants through their services to match or integrate the information [Agy+;Gaw;Gen+;HMK;Hei+;Sem;SMJ]. As Sachs notes in in his book “The Ages of Globalization: Geography, Technology, and Institutions”, economic decisions are made by software The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 56 2. Data Protection, Consumer Protection and Competition Law 2.2. Personal Data Protection and Market Competition Figure 2.5: Results of the in-time placebo test (after including control variables) programs in the digital age, not by management in the modern industrial age [Acs+;Sac]. This shift underscores that the cornerstone and original source of online platforms even the digital economy is based on the generation and analysis of data that the online platform or subjects hold [Gaw], especially the personal data [ATa]. The value creation of online platforms is intrinsically linked to their ability to capture and utilise user data. By controlling user interactions and accumulating vast databases of user activities, online platforms enhance their services and create market dependencies that shape the dynamics of the digital economy. This work empirically studies the impact of the European Union’s General Data Protection Regulation (GDPR) on the social media market concentration in the EU, employing both the synthetic control method and the generalized difference-in-differences method. The findings reveal that the GDPR significantly reduced social media market concentration from 2015 to 2020, with a stronger impact on large companies. However, in the long term, the impact of the GDPR on EU social media market concentration is gradually fading, which has been very weak after 2020. Furthermore, the impact strength of the GDPR on the social media market concentration can be changed by Internet market scales and high technology levels. These insights contribute to a deeper understanding of how data protection policies shape the market dynamics of social media companies. This study along with other studies offers strong causal evidence of the interplay between personal data protection and digital market competition, and further proves that the impact of privacy protection is nuanced and context-specific [GT]. This means that the regulation of personal data protection is no longer the task of a single authority, but requires close cooperation across regulatory fields, such as market competition, consumer protection, and artificial intelligence. To ensure that the GDPR is effectively enforced, the DPAs should choose a regulatory model with long-term impact, and the regulatory practices should evolve with changes in technologies and market areas. A more cautious stance should be taken The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 57 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability the technical complexities of ensuring portability across heterogeneous systems. The research also considers the challenges of syntactic, semantic, and content heterogeneity. 2.3.2.2.1 Research’s Main Argument: The Indeterminacy of the Content of Data Portability’s Obligatory Claim Through the analysis of the three key dimensions, the ultimate objective of this work is to convince the reader that the greatest obstacle to realizing data portability is the non-determinability of the content of data portability’s obligatory claim. The subject matter of the legal obligation to data portability, according to the normative provisions, is incorporated in the data provided by the right holder. To say it differently, the right holder can request the simple delivery of the data they have provided. However, this determination is only seemingly straightforward, as it is fraught with complexities in its practical application. In concrete terms, the right holder has no certain way of knowing exactly how much data they have actually provided, nor how much has been processed and recorded by the service provider/obligated party–except through a complex procedure that does not guarantee the genuineness of the obligated party’s response. Further complicating the matter is that even if the obligated party acted in good faith to make the requested data available, they may still not provide a sufficient amount of information with respect to the purposes of the portability request. Discussions about reusability and data content can be found in [Pub21] This also means that not only it would not be possible for an appointed judge to assess the adequacy of the performance by the obligated party, but also for a portability right holder to know the extent of their right. Among the essential characteristics of an obligation is that it must be determined or determinable, meaning that, from the outset of the relationship, the content of the performance must be identified or, at the very least, all necessary elements must be present to define the performance when it is due to be executed. For any justice system to operate effectively on the basis of rights and obligations, these obligations must be either determined or determinable. This work will support and demonstrate the following statement: in the case of data portability, obligations are neither determined nor determinable. No metrics have been established to define the adequate content of a ported data set, preventing a data portability right holder from objecting that “necessary information is missing,” an obligor from asserting fulfillment of the obligation, or a judge from effectively ruling on the matter. 2.3.2.3 State of the Art The conversation around data portability sits at the intersection of law, market competition, and digital technologies with each field bringing its own perspective to how we understand and shape the concept. 2.3.2.3.1 The Metamorphosis of Data Portability in EU law An early insight supporting this research is the recognition that the meaning of data portability has significantly morphed as it has been carried across various regulations. What began as a consumer-focused right to transfer The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 64 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability data between service providers has gradually transformed into a market-driven obligation for providers to ensure data availability. Legally, data portability is primarily governed by Article 20 of the General Data Protection Regulation (GDPR), which gives individuals the right to transfer their personal data between service providers in a structured, commonly used, and machine-readable format. More recent regulations, like the Digital Markets Act (DMA) and the European Data Act (DA), have expanded on this right, introducing obligations for major digital players, or “gatekeepers,” to promote wider data-sharing practices. The literature has often focused on analyzing the concept of data portability within individual legislative texts, such as the GDPR, DMA, or DA. However, the research conducted adopts a holistic and cross-regulatory analysis of the legal texts to demonstrate that there is not just one unified concept of data portability. Instead, there are several concepts—often ontologically quite different—that are nonetheless referred to by the same name, despite having distinct natures, purposes, and implications depending on the regulatory framework in which they are situated. Through the deconstruction of the concept of portability in its basic building blocks that is, data movability, data transportability, and data ease of carry, different regulations envision data portability diversely. At its basic level, the concept of data portability should embed the characteristics of data mo(va)bility from one place and of transportability in a context dependent, sufficiently easy fashion. This notwithstanding, regulations such as the GDPR, Free Flow of NonPersonal Data Regulation (FFNPD) and DMA each directly define or indirectly intend portability as a set of data (a dataset, in fact) to be treated differently, depending on, for instance, the use of specific technical data format, the timing of service provision, and so on. The literature analysis reveals that data portability is not merely a matter of individual rights– which characterizes it, as such, as already complex,[Alb14] but also a regulatory tool aimed at fostering competition in the digital market. Relevant scholarship notes that portability serves as a bridge between data protection and competition law,[GHP18] enabling users to break free from the lock-in effects of dominant digital service providers, thereby promoting market dynamism.[DH+18;GHP18;Lun22;Lyn20] This duality–portability as a personal right and as a market regulator–[Cor16] presents a significant challenge for regulators, who must balance these sometimes conflicting goals.[Syr+21;Kra+23;TT24;SL13;VU17;Man21] In fact, contrary to the beliefs of the Article 29 Working Party,[Par17b] it is not possible to regulate data portability without impacting on data markets and data-based businesses and services. 2.3.2.3.2 Data Portability as Microcosm of Unresolved Conflicts Understanding the informational background of the right to data portability cannot be confined to its inter-legal conceptualization, but must extend to its role as a conceptual representation of unresolved conflicts that permeate the digital ecosystem today. In traditional, non-digital contexts, individuals maintain control over their private information by limiting disclosure and protecting it through various means. However, once that information crosses a “line of secrecy” and enters the public domain or is shared with third parties, it becomes vulnerable. A crucial characteristics of data governance is that once data exits the control of the individual, it is nearly impossible to “regain control over it.”[Koo14] In the physical world, The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 65 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability assets can be retrieved or secured through reactive measures, such as recovering a stolen car. In contrast, in the digital realm, data is non-rivalrous and non-excludable, meaning that once shared, it is difficult, if not impossible, to retract or fully control.[Spi+15] Counter-intuitively, the concept of data portability is built on the premise of reinforcing individual control through enhanced data sharing. Yet it has been argued that the corporate emphasis on privacy-as-control, on which data portability policies are based, serves rather the data extraction goals of companies than the true interests of users.[Wal21a] The pretence that more sharing equals more control is a facade, in reality allowing corporations to “justify their datagathering practices” while giving users the illusion of control. If anything, data portability tends to exacerbate the problem of data sharing rather than solve it. Once data is ported, it often ends up being further distributed, leading to further processing and exploiting by third parties and diminishing individuals’ control. Historically, critical privacy issues stemmed not from the initial data controller’s collection of data, but from the lack of mechanisms to control how the data could be shared with third parties.[Koo11] Thus, allowing users to port their data does not equate to giving them genuine control. Instead, it expands the number of recipients who can potentially misuse or exploit that data, leading to what Mitchell describes as a false sense of autonomy.[Mit23] As Koops rightly emphasizes,[Koo14] data protection regulations, as they are currently written, offer only an illusory form of control. Informational selfdetermination, the very principle on which data protection is built, becomes unenforceable in a world where data, once shared, can no longer be meaningfully controlled. This illusion of control provided by data portability, far from empowering users, aligns more closely with competition law than with privacy protection. Graef[GHP18] and Purtova and Newell[PN24] argue that data portability might be more suited for regulating competition and innovation, rather than being seen as a genuine privacy-enhancing tool. Thus, the concept of data portability, while seemingly empowering, ultimately contributes to the proliferation ofdata-sharing, andweakens individuals’grasp over theirpersonal data. Rather than providing a solution to privacy concerns, data portability might be better understood as a mechanism designed to stimulate competition and innovation in the digital market, but with limited efficacy in truly enhancing user control. Evidence suggests that portability can serveas thequintessence–theepitome, even amicrocosm– of the various conflicts and tensions that exist within the realm of data regulation, between control and re-use, between protection and free-flow. Some core tensions both inherent generally in data regulation and specifically in the study of portability are here summarized: •Privacy as confidentiality vs. privacy as control:[GA16] Data portability challenges the notion of privacy, shifting from a focus on secrecy to a dynamic of empowering users with control over their data. •Opacity vs. transparency:[DHG06a] data portability raises questions about the transparency of systems and whether users truly understand what happens to their personal data once shared and reused. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 66 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability •Protection vs. free flow: the right to data portability embodies the tension enshrined in article 1 of GDPR between safeguarding personal data and promoting the free flow of information across services. •Data minimization vs. reusability through multiplication: the GDPR principle of data minimization, requiring that only the personal data that is necessary, relevant, and limited to the specific purpose for which it is being processed should be collected and retained, which is central to personal data protection, can conflict with the European Data Strategy goals of data reusability and portability. •Privacy vs. data protection: while related, privacy–interpreted as the right to be left alone–and data protection–which allows third-party data processing for legitimate interests–can clash in the context of data portability, as seen in the balancing of individual rights and broader regulatory objectives. •Informational self-determination vs. informational self-determination: Portability both strengthens and complicates the notion of informational self-determination, offering control while introducing new risks of third-party use. •Private protection vs. sharing and reusing: Users are expected to protect their personal data, but at the same time, data portability encourages them to share and reuse it across platforms. •Business data protection vs. data market competition: Portability poses challenges to businesses that aim to protect proprietary data, as it could facilitate competition by enabling easier access to consumer data by competitors. This tension-filled landscape underscores the complex and multifaceted nature of data portability, touching on nearly all key areas of data law and policy. Some of such concepts are in the following–by no means exclusive–list: data governance, fundamental rights protection, data control, fairness, digital markets competition, personal and non-personal data protection, information system design, technological reference architectures, data modelling, artificial Intelligence, market power, economics of data, data security, and more. In this microcosm of tensions, it is also possible to observe an evolutionary parallel between, on the one hand, European policies on the management of the digital economy, markets, and digital services, both public and private, and on the other, the concept of data portability. It becomes evident how the focus has shifted from more protective models regarding data generated by individuals, businesses, and public administrations, supported by an individualistic approach to data, towards a model of sharing and reuse, supported by a pro-social approach. 2.3.2.3.3 Technological Landscape for Data Portability: Heterogeneity of Syntax, Structure, Semantics...and Content! After observing that data portability has undergone a conceptual transformation throughout its historical and legal evolution, and that the complexity of its regulation and implementation stems—if not primarily—from its role as an emblem of longstanding The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 67 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability conflicts within the digital society, the analysis of its informational background must inevitably shift focus towards the technical and technological aspects. The technical literature highlights the intrinsic complexity of data portability, particularly due to, on the one hand, the foundational difference between data and information, and, on the other, tothe heterogeneityof dataformats,[USM18] structures,[FI14;Flo11;Nic21;Bat+16;Gel22;Flo10; PN24] and metadata across different services.[Jan17] Differences Between Data and Information In the context of personal data portability, there is a crucial distinction between “data” and “information.” The GDPR equates data with information, defining personal data as “any information relating to an identifiable natural person.” However, this equation raises fundamental questions. The conceptual separation between data and information is often overlooked, with the two terms being used synonymously, leading to ambiguities in both legal interpretations and technical applications.[Byg15;Zin07;HDH16] Scholars such as Bygrave, Gellert, and Floridi have ventured into this complex distinction, yet the debate remains underexplored in both doctrine and case law.[Byg15] The challenge lies in their ontological differences: while data is often considered objective, syntactic, and quantitative, information is perceived as subjective, semantic, and qualitative.[Ros13] This creates a minefield for regulation, particularly when it comes to obligations like data portability, which require a precise understanding of what is being transferred and controlled–whether data or information. The literature reveals varied approaches to defining these terms, particularly within information theory.[CH03;Flo11] The legal field complicates matters by adopting inconsistent definitions across regulations.For instance, while the GDPR treats data as synonymous with information, the Data Act distinguishes between data as a digital representation and information as the interpretation of such data–as already noted in [Ros13] This distinction in legal texts has significant implications for the scope and application of data portability, as well as for broader digital regulations. There seem to exist varied regulatory logics within the GDPR itself. While data portability is treated as a right on data–therefore as meaning-agnostic right–many other GDPR provisions, such as rights to accuracy and rectification, are clearly meaning-driven, necessitating an understanding of the information’s semantic content.[PN24] This dual approach within the same regulation reflects a broader struggle to appropriately address both syntactic and semantic dimensions of data in a coherent legal framework. The importance of distinguishing between data and information is not a theoretical exercise, but a fundamental issue for effective legal regulation of the digital society. Without this clarity, “data laws” risk being blunt instruments that fail to address the nuanced challenges posed by the digital age. Data Syntax and Structure in Data Portability The technical scholarship examines the complexities of data formats, syntax, structure, and metadata in relation to data portability. It acknowledges the role of syntax in data handling and communication, drawing a parallel between grammatical rules in natural languages and the syntactic formalism of programming languages. Each programming language has its own syntax, which determines the structure and sequence of elements, as noted by Li,[LF21] who emphasizes that the correct organization and encoding of data within a format is essential for it to be readable and valid for machine processing . The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 68 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability Thestructural formalismof data is equally crucial. Differenttypes of data structures—structured, unstructured, and semi-structured—present varying challenges for data portability. Unstructured data, such as text documents or multimedia files, lacks a predefined schema, making it difficult to query or process. Structured data, exemplified by tabular formats like CSV, follows strict schemas that allow forefficient querying and processing. Between these two extremes lies semi-structured data, typified by formats such as XML and JSON, which contain tags or markers that provide a loose structure for hierarchical relationships. While these semi-structured formats are widely used in Application Programming Interfaces (APIs) and web services, they bring challenges due to their relative lack of strict constraints, as discussed by Wong and Henderson,[WH18] who note that such lax constraints often hinder seamless data portability. The literature also highlights the importance of schema heterogeneity in data portability. The integration of data from diverse sources, each with its own schema, presents significant challenges. Doan et al. illustrate that even when the same requirements are given, different developers will produce distinct schemas.[DHI12] Schema conflicts arise from differences in naming conventions (synonyms and homonyms) and the organizational structure of data, leading to logical mismatches between systems. This heterogeneity complicates data portability, as the importing system must often develop specific import procedures to accommodate the schemas of exporting providers. In the context of data portability, machine readability of data is paramount. Regulations such as the GDPR require data to be ported in a “machine-readable” format, which mandates that the data be organized in a way that machines can automatically process. Formats like CSV, JSON, and XML are frequently cited in the literature as examples of machine-readable formats. However, as Kranz et al. argue,[Kra+23] despite the technical compliance of many formats, true interoperability often falters when systems fail to adhere to standardized schemas. The challenge of maintaining machine readability and ensuring consistent data interpretation across systems remains a central concern in the literature. Metadata plays a critical role in data portability.[WH19a] Metadata–data about data–provides necessary context for understanding and processing dataset. Effective metadata allows systems to interpret the structure and content of data, but inadequate metadata can lead to significant obstacles in achieving true data portability. The lack of clear guidelines on the types and standards of metadata to accompany ported data exacerbates the challenges faced by service providers. Furthermore, metadata standards like the Dublin Core Metadata Terms (DCMT) and Data Catalog (DCAT) vocabulary are essential for achieving interoperability, as they provide a structured framework for cataloging and describing data, so improving interoperability. Yet, without sufficient documentation or standardized metadata, the receiving system may struggle to make use of the ported data. In summary, the literature emphasizes that the success of data portability hinges on a combination of factors, including adherence to syntactic rules, the resolution of schema heterogeneity, and the provision of adequate metadata. While semi-structured formats like XML and JSON provide flexibility, they also present challenges in achieving the level of structural consistency required for seamless data exchange. The role of metadata in ensuring both machinereadability and semantic interoperability is crucial but remains underdeveloped in current regThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 69 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability ulatory frameworks. As the literature demonstrates, addressing these technical issues is vital for realizing the goals of data portability as envisioned by regulations such as the GDPR, the DMA and the Data Act. Semantic Heterogeneity and Data Portability Semantic heterogeneity, the challenge of aligning meaning across different data systems, is a key issue in the context of data portability. It is important to understand how data retains or loses meaning in the transfer between systems. The relevant literature focuses, once again, on the distinction between data and information, the role of semantic heterogeneity, and approaches for resolving these issues through database integration methods. Wand and Weber outline two critical processes in information systems:[WW93] representation transformation, where a real-world system is modeled and populated with data, and interpretation transformation, where the system is used to infer insights.[Hil18] The disconnection between these two processes is particularly problematic in data portability because the data is generated and formatted in one system (the source) but interpreted in another (the recipient). Hildebrandt emphasizes the challenge of aligning these two transformations, as the meaning of data is often tied to its original context. The distinction between data and information is central to understanding semantic heterogeneity. Data refers to the raw symbols or signs, while information is the meaning derived from them. Veltman[Vel01] and Shannon[Sha48] argue that for effective communication, both syntax (structure) and semantics (meaning) must be aligned. Shannon’s entropy theory, which measures the uncertainty or randomness in a set of data and quantifies the amount of information needed to describe its possible outcomes, suggests that in communications, the goal is to reduce uncertainty so that the receiver correctly interprets the message. In data portability, this problem is exacerbated because the receiving system often has different expectations or requirements for data usage than the originating system, which leads to misalignment in the transfer of meaning. The literature shows that there have been plenty of approaches to addressing semantic heterogeneity, particularly from the field of database integration, which offers useful solutions for resolving these challenges. Rahm and Bernstein describe semantic schema mapping as a common method for addressing discrepancies between different systems.[RB01] Schema mapping involves aligning vocabularies or data structures across systems to ensure that data retains its intended meaning when transferred. Noy et al.[NDH05] and Doan[DHI12] extend this by focusing on creating unified semantic frameworks for integrating heterogeneous databases. AI and machine learning approaches are also emerging as effective tools for resolving semantic inconsistencies. Koutras highlights the ability of AI to analyze large datasets and recognize patterns, even when the data is unstructured or semi-structured.[Kou19] These techniques can help systems learn to reconcile differences between data sources, making it easier to integrate data without relying solely on predefined schemas. Ontology-based integration methods are another important approach. Gagnon [Gag07] and Allemang[AHG20] discuss the use of ontologies to define common vocabularies within a domain, enabling consistent interpretation of information across systems. Ontologies help establish shared meanings, which is critical for achieving semantic interoperability in large-scale data integration efforts, such as those required for the The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 70 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability Semantic Web.[Noy04] Data portability is oftentimes hindered by issues of heterogeneous data structures. The diverse nature of data in portability rights—ranging from personal data under the GDPR to business data under the DMA and Data Act—further complicates the issue of semantic heterogeneity. Strauch notes that NoSQL databases, which handle semi-structured data like JSON or XML,[Str11] are increasingly common in managing personal and business data. However, these databases differ from traditional relational databases, making integration more complex. Finally, Kranz et al. highlight the need for standardized formats and schemas to ensure successful data portability, particularly in cloud services.[Kra+23] Business data stored in cloud platforms like Amazon Web Services or Google Cloud often involves a variety of data formats and structures, requiring harmonization at the database level. This challenge is also present in the Data Act, where realtime access to raw data from connected devices must be accompanied by sufficient metadata to maintain data quality and meaning. In conclusion, semantic heterogeneity presents significant challenges in data portability, particularly when transferring data across systems with different structures and expectations. The literature provides valuable insights into resolving these issues through approaches like schema mapping, ontology-based integration, and AI-driven solutions. As data portability becomes more prominent in regulations, addressing the semantic challenges will be essential for ensuring meaningful data exchanges between systems. Integrating the solutions developed for database integration into data portability frameworks will help overcome these hurdles, ensuring that data retains its meaning across different contexts. Attempts at Addressing Structure and Semantic Heterogeneity Structural Harmonization and Semantic Artifacts in Data Portability. In the field of data portability, ensuring the exchange of meaningful and interoperable data across systems presents a complex challenge, particularly due to issues surrounding syntactic and semantic heterogeneity. Several strategies have emerged to address these issues, ranging from the development of syntactic data standards to the establishment of semantic frameworks that facilitate meaningful data exchange. Syntactic and Semantic Standardization. Syntactic standards refer to the formatting and structuring rules for data, which are often defined by widely used formats such as XML and JSON. These formats ensure that data can be stored and exchanged in a machine-readable manner. However, syntax alone does not guarantee interoperability because it lacks the ability to convey meaning. For example, XML Schema and JSON Schema define structural constraints, but the lack of semantic agreement may lead to confusion about the intended meaning of the data.[KSS20] Semantic standards, on the other hand, focus on ensuring that data is described consistently across systems by providing shared vocabularies and taxonomies. These standards allow systems to interpret data in a meaningful way by ensuring that terms, attributes, and relationships are uniformly defined. Controlled vocabularies and ontologies play an essential role in harmonizing semantics. Ontologies, such as those used in the Web Ontology Language (OWL), enable the definition of entities and their relationships in a way that is independent of the specific format in which data is represented. Merging Syntax and Semantics. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 71 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability Although syntactic and semantic standards have traditionally been developed separately, certain tools merge these two aspects to facilitate both structural and semantic interoperability. For instance, XML and JSON formats often allow developers to define tags and keys autonomously, but the development of controlled vocabularies or ontologies can introduce shared semantics within these syntactic structures. OWL, for example, provides a harmonized semantic framework that can describe entities independently of their syntactic representation, thus promoting interoperability. In specific sectors such as healthcare, comprehensive standards like HL7’s FHIR address both syntactic and semantic issues simultaneously.39 FHIR integrates standardized data models with semantic vocabularies to facilitate the exchange of clinical information on a global scale. This approach ensures that data exchanged within and between healthcare systems remains meaningful, even in highly complex environments. Coordinated Frameworks for Data Sharing. In competitive commercial settings, such as those governed by private data portability regulations,[SEE24] there is often little incentive for companies to share data among each other (Business-to-Business, or B2B), with consumers (Business-to-Consumer, or B2C) or the state (Business to Government, or B2G).However, solutions from non-competitivepublic-sectorframeworks offer useful insights.[SEE24] For example, the Open Data Directive40 and the Data Governance Act41 in Europe mandate the sharing and reuse of public data, often requiring standardized schemas for data exchange.42 In Italy, the National Guidelines for the Enhancement of Public Information Assets mandate that data schemas,such as XSD for XML and JSON Schema, be shared to ensure accurate data interpretation.[AGI23] Data Schemas and Structured Formats. The flexibility of syntactic standards like XML has contributed to their widespread adoption but has also introduced challenges for interoperability. XML allows for autonomous schema creation, which can result in syntactically valid data that lacks semantic consistency across systems. XML Schema (XSD) was developed to address this issue by specifying the structure, constraints, and data types required for valid XML documents.[Bik+13] Similarly, JSON Schema defines a structured approach to validate JSON data against predefined models. These schemas are critical for ensuring that data exchanged between systems is both syntactically correct and semantically meaningful. XSD and JSON Schemas are available, open source tools that could already be put to use to standardize data schemas and structures for purposes of portability. Semantic Artifacts and Ontologies. The need for semantic harmonization has led to the development of semantic artifacts such as controlled vocabularies, taxonomies, and ontologies. These tools provide a structured way to represent and share knowledge across systems. Ontologies, in particular, offer a high level of ex39See FHIR specification at http://hl7.org/fhir/index.html#0. 40Directive (EU) 2019/1024 of the European Parliament and of the Council of 20 June 2019 on open data and the re-use of public sector information, PE/28/2019/REV/1, OJ L 172, 26.6.2019. 41Regulation (EU) 2022/868 of the European Parliament and of the Council of 30 May 2022 on European data governance and amending Regulation (EU) 2018/1724 (Data Governance Act), PE/85/2021/REV/1, OJ L 152, 03/06/2022. 42See also the Regulation (EU) No 1024/2012 of the European Parliament and of the Council of 25 October 2012 on administrative cooperation through the Internal Market Information System and repealing Commission Decision 2008/49/EC (‘the IMI Regulation’), OJ L 316, 14.11.2012. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 72 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability pressiveness by defining concepts, relationships, and rules within a specific domain. Ontologies such as OWL, which build on the Resource Description Framework (RDF), enable a common understanding of data across different systems, promoting interoperability even in heterogeneous environments.[Gru93] In the public sector, for example, the European Union’s Core Vocabularies, developed under the Interoperable Europe initiative, provide simplified and reusable data models that capture the fundamental characteristics of entities in a context-neutral fashion. These vocabularies ensure that essential properties, such as those related to persons or organizations, are consistently shared across public-sector systems. The alignment of vocabularies across different namespaces, as seen in the mapping between Schema.org and the Core Vocabularies, further illustrates the importance of semantic consistency for effective data exchange. In conclusion, the literature on structural harmonization and semantic artifacts underscores the critical role of both syntactic and semantic standards in achieving effective data portability. Syntactic standards like XML and JSON ensure that data can be structurally validated, while semantic standards such as controlled vocabularies and ontologies guarantee that the data exchanged retains its intended meaning. As data portability regulations continue to evolve, the integration of these approaches—through tools like XML Schema, JSON Schema, and OWL—will be crucial for overcoming heterogeneity and ensuring meaningful data exchange across systems. Data Content Several studies have examined the practical application of data portability systems and their existing challenges. Syrmoudis et al.[Syr+21] and Wong and Henderson[WH19b] have identified key issues within these systems, notably data fragmentation and the variability in “data richness.” These challenges highlight how data formats and content differ across platforms and services, leading to inconsistencies in the quality and usefulness of transferred data. The issue of data format heterogeneity has been extensively covered in the literature, with a focus on standardizing data models through schemas like XML Schema (XSD) and JSON Schema, which define properties for specific data types. Object-oriented programming (OOP) concepts can help explain the structuring of data in these systems. For example, the “class” and “attribute” model in OOP, where high-level classes encapsulate general functionalities and lower-level classes capture specific details, has been applied to data types in services like Google Contacts. In the context of data portability, the concept of “data richness” becomes crucial. It refers to both the variety of data types a service collects and the amount of information embedded within each type. Google Takeout, for instance, offers a wide range of data types, such as contacts, emails, and photos–richness in types, with each type containing various attributes like name, email, and phone number–richness in information per data type. This richness in information per data type is defined by the number of data fields available within each data type, as illustrated by the case of Google Contacts, which includes 25 different data fields out of a potential 45 based on the vCard standard. This selective inclusion of data attributes highlights the arbitrary nature of data model design, where developers decide which fields are necessary or optional based on the functional needs of the system. The concept of data richness, when applied to the single data type (or class), refers to the amount of information, the attributes (fields) about that data type, therefore to how much inThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 73 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability tives, limiting third-party access and innovation. Furthermore, the choice of cloud storage services—restricted to options like Google Drive and Dropbox—raises additional concerns about limiting user choice and potentially exposing them to privacy risks. These examples emphasize the challenges in ensuring fair competition while respecting user rights to data portability, especially as new regulations like the Digital Markets Act come into force. The DMA, for instance, requires continuous, real-time data access, yet Google’s Takeout system currently only offers periodic access through a mediated link. This delay in access underscores the need for further adjustments to meet the DMA’s standards for seamless data portability. To address these issues effectively, future regulatory efforts should integrate both competition law and data protection, ensuring that commitments, like those made by Google, are assessed not only for their immediate impact but also for their long-term consequences on market fairness and user autonomy. 2.3.4.0.4 As Technological Design of Portability Shapes Policy Outcomes, Technological Specificity is Needed The research highlights the critical need to understand technology deeply, particularly in contexts like data portability, where the technical and legal realms are intricately connected. Regardless of one’s perspective—whether focused on the societal impacts of technology or on the regulatory challenges it presents—both sides agree on the essential role of technological literacy. Understanding the technical building blocks of data systems is foundational to any regulatory efforts aimed at governing complex systems like data portability. Concepts such as data formats, structures, syntax, and semantics, as well as advanced notions like ontologies and knowledge representation, are all crucial for creating a common technical language that bridges the gap between legal, technical, and economic discussions. The complexity of data protection as a legal field is a major theme. It is not only complex due to the varied nature of data itself, which includes information, knowledge, and its flow across systems, but also because it involves competing rights and societal values that must be protected simultaneously. As the literature points out, this inherent complexity is compounded by the risk-based nature of data protection regulation, which requires a balance between managing technological innovation and protecting individual rights. The interaction between technology and law is particularly important, as legal principles often emerge in response to technological innovations. In this regard, technology design processes may sometimes conflict with legal norms, especially when it comes to values like ethics, legitimacy, and democratic accountability. The research addresses the debate around technological neutrality versus specificity in regulation. While technology-neutral laws are generally preferred, especially for their adaptability to future technological developments, the discussion points out that certain areas, like data portability, require a more specific regulatory approach. This is because technology-specific laws can directly address unique challenges or moral concerns raised by specific technologies, ensuring that human rights and other fundamental principles are protected. Hildebrandt’s work is cited to underscore this point, highlighting that sometimes a technology-neutral approach fails to achieve its intended goals, and specific legislation becomes necessary to counter the risks introduced by technological design choices. In conclusion, regulatory success in fields like data portability and data sharing largely depends on the implementation of technical standards by industry actors. The effectiveness of laws governing these domains hinges not only on the legal frameworks themselves but also on how well they are integrated with the technological realities The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 80 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability of the systems they seek to regulate. The technical solutions used to implement data portability significantly impact the policy goals that are achieved. For instance, if dominant data holders control the structure and content of ported data, they may consolidate their market power by creating de facto standards. Conversely, more user-centric systems can empower individuals by giving them control over their data, while society at large benefits from increased data sharing and reusability. 2.3.4.0.5 Meaningful Data Portability is Information Portability The analysis of personal data portability within European regulations reveals a fundamental issue stemming from the conflation of “data” and “information.” The GDPR’s approach to data portability, which treats it as a meaning-agnostic problem, overlooks the inherent need for meaningful content in the transfer process. Scholars like Purtova and Newell criticize this oversight, arguing that the right to data portability should be understood as a problem of knowledge communication rather than mere data extraction and transmission. The crux of the issue is that data portability focuses on the transfer of data but fails to ensure that the information—data imbued with meaning—retains its significance in the new context of processing. Afunctional re-interpretation ofdataportabilitymusttreatitas a communication processwhere data is not only transferred but also remains meaningful to both the sender and the recipient system. As Floridi’s General Definition of Information emphasizes, data must be well-formed and meaningful to qualify as information. Thus, for data portability to be effective, data must be transferable in a way that maintains its semantic integrity, which is not currently achieved under the GDPR. Additionally, the legal framework must acknowledge the difference between data and information, as highlighted in the recent Interoperable Europe Act, which distinguishes between data, information, and knowledge. This distinction is crucial for a coherent legislative approach that aligns with the real-world application of data in digital environments. Ultimately, the goal of data portability should be the transfer of meaningful information rather than raw data. This requires a shift in how legislators and regulators conceptualize data portability, moving from a technical task of data migration to a process that ensures the preservation and usability of information across different systems. The failure to address this issue has led to the underachievement of data portability as a right, leaving individuals with the ability to transfer data but not necessarily the information they need. 2.3.4.0.6 Data Portability Needs Structural Harmonization and Semantic Artifacts The flexibility in designing data models and structures, while promoting innovation, has led to significant challenges in data portability due to structural heterogeneity. Kramer et al. (2020) highlight that the data models and formats used for representing, storing, and exchanging data are largely dependent on the type of data. This flexibility results in a lack of consistency across different systems, necessitating the development of standards like XML Schema (XSD), JSON Schema, and RDF Schema to standardize data structures. These schemas impose clear rules on properties within data models, ensuring consistency and interoperability across various platforms. For instance, Facebook’s Graph API, which uses nodes, edges, and fields, aligns with the structural concepts seen in these schema-based formats, facilitating data exchange between systems. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 81 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability Despite the apparent differences between data formats like XML, JSON, and RDF, the underlying structural similarities offer a pathway toward harmonization. Developed by W3C, many of these formats share a common foundation in how they model data objects and their properties. For example, XML and HTML, though designed for different purposes, have similar architectures based on tags, and both XML and RDF/XML serialize data in a similar manner. This allows for schema mappings across formats, meaning that data portability can be enhanced by translating the structures between different systems, reducing the risks associated with structural heterogeneity. Schemas also play a crucial role in restricting and defining the nature of data objects, their properties, and their relationships within a system. XSD and JSON Schema, for instance, allow developers to explicitly specify data types, set constraints, and determine mandatory and optional fields, which helps mitigate inconsistencies in data interpretation. In social media, for example, the Facebook Graph API’s use of nodes and edges to represent users and relationships demonstrates how these structured formats ensure that data remains actionable and interpretable, regardless of the platform receiving it. The research further suggests the need for unified conceptual blueprints across different industries to address data heterogeneity. Core schemas could be developed to define key objects and properties in fields like healthcare, finance, and social media, ensuring that data exchanged during portability requests is structured and meaningful. Similar efforts, such as the EU Core Vocabularies–in the inter-administrative sector–and Schema.org offer frameworks for standardizing these concepts. By mandating the use of shared schemas and namespaces in data portability laws, regulators can ensure that data exchanged is not only well-structured but also semantically consistent, enabling effective reuse of the data across different systems. This would uphold the essence of data portability rights by ensuring that transferred data retains its meaning and usability even in contexts of private data sharing.Private data sharing, in this context, means data sharing among private parties, such as businesses, organizations, professionals, or individuals.[SEE24][p. 3] 2.3.4.0.7 In Data Portability, Quantity Matters Research into data portability systems, particularly Google Takeout, reveals substantial fragmentation in both data format and content. For instance, Syrmoudis et al. found that after data export, there is a noticeable variation in “data richness,” which could either stem from the variety of data types used by the source system or the richness of content within each type. This finding underscores the complexity of ensuring that data retains its full utility when transferred between services, as companies often limit the amount of information available for export. Google, for example, limits the attributes included in the Contact data type, potentially withholding additional data fields that could be useful for the user in another context. Furthermore, data models are arbitrarily defined by developers, leading to variations in what is shared during portability. Google’s decision to include only 25 out of the possible 45 vCard attributes in its Contact data type, while still complying with the necessary requirements, exemplifies how companies use selective data modeling to control the flow of information. This practice creates challenges for users and platforms attempting to reuse the data meaningfully, The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 82 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability as the excluded data fields may contain valuable information. Additionally, even when data is transferred according to standard formats, the completeness of the information remains an issue, with developers having the discretion to include or exclude certain attributes. The findings also highlight the difficulty of regulating data portability in a heterogeneous digital ecosystem. The infinite variability of data types and properties across different services complicates the implementation of standardized data exchange practices. There is a need for more systematic approaches to simplify and categorize data types, reducing complexity and promoting meaningful reuse. By developing common guidelines and schemas for data exchange, similar to the EU Core Vocabularies, the effectiveness of data portability rights can be significantly enhanced, ensuring that transferred data retains its intended richness and utility across platforms. Finally, the research emphasizes the importance of moving beyond just data formats to consider the content being ported. Even if data is transferred in a compatible format, it loses value if it cannot be meaningfully reused by the receiving system. As the European Data Strategy emphasizes, “The value of data lies in its use and re-use.”43 This calls for a shift in focus towards specifying the quantity and quality of data that must be shared during portability requests, thereby creating a more predictable and enforceable framework for both data controllers and users. This approach would align the legal requirements of data portability with the practical realities of data reuse in a heterogeneous digital environment. 2.3.5 Research Analysis 2.3.5.0.1 Data Quality—meaning: fitness for use—is the Guiding Principle for Meaningful Data Portability The research demonstrates that data reusability is central to the European Data Strategy, emphasizing the need for data to remain accessible and valuable when transferred between systems. One significant finding is the application of the principle of “fitness for use” to data quality, showing that reusability is not solely about ensuring data can be technically transferred, but also about maintaining its usefulness in different contexts. For instance, the European Commission stresses that data must be provided in formats that allow potential reusers to understand and make use of it effectively. This highlights the responsibility of data holders to ensure that data is fit for purpose beyond its original system. Reusability involves ensuring both quantitative and qualitative aspects of data are addressed. The EC Data Quality Guidelines indicate that providing insufficient amounts of data reduces its usefulness, making reusability impossible. This issue aligns with the concept of “data quality as fitness for use,” where good quality data must be accurate, complete, and contextually relevant. An example of this principle in practice can be found in the ISO 8000-1:2022 standard, which reinforces the importance of data being fit for the specific task it is intended for. This finding stresses that data quality is essential for supporting meaningful data portability. Another key finding is that data quality is not a uniform concept but must be tailored to the needs of its intended users. The empirical approach highlighted in the literature, particularly by Wang and Strong, focuses on understanding data from the consumer’s perspective. This consumercentric view aligns with the idea that data portability must consider the expectations of the data 43European Commission, European Data Strategy, p. 6. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 83 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability recipient, whether that is the data subject or a third-party reuser. The findings suggest that by applying data quality principles, such as those found in ISO standards, the right balance can be struck between providing adequate content while ensuring the data remains reusable and fit for purpose. The research also identifies legal frameworks where data quality is explicitly referenced. In particular, the European Health Data Space and the Data Governance Act both highlight the role of data quality in ensuring effective data exchange. These legal instruments underline that data must not only be provided but must also retain its value across different uses, a principle that can be applied more broadly to other data portability systems. Ultimately, the findings suggest that data quality, particularly when framed as “fitness for use,” provides the only metric for ensuring that data portability obligations are both fulfilled and aligned with the broader goals of data governance and consumer protection. 2.3.5.0.2 Data Portability Really Means Data Of Data Portability One of the core issues identified is the lack of clarity regarding how much and what type of data needs to be ported to satisfy legal obligations. This problem of “data of the data”—specifically with regards to attributes—introduces a level of complexity that current regulations do not adequately address. By introducing data quality as a guiding principle, the report proposes a framework for determining what information must be included in the ported data to ensure it meets the needs of both users and regulatory bodies. 1. Proportionality-Based Metrics: The research introduces a proportionality-based approach to data portability, suggesting that the quantity and quality of ported data should vary depending on the regulatory context. In GDPR, where the focus is on informational self-determination, applying the principle of data quality to the right to portability would allow individuals to port as much personal data as they wish. In competition law, however, where the goal is to promote market functionality, the principle of data quality might allow that a lesser amount of data may be sufficient. This flexibility ensures that data remains fit for purpose across different legal frameworks. 2. Data Quality as Fitness for Purpose: Finally, data quality—interpreted as fitness for purpose—should become the guiding benchmark for data portability. By ensuring that ported data retains its value in subsequent services and applications, data quality serves not only to fulfill legal obligations but also to create a more functional digital ecosystem. The findings of this report reveal the inherent complexity of data portability, emphasizing that it cannot be fully understood or regulated within a single legal framework. Data portability is a microcosm of most, if not all, of the conflicts and tensions that are inherent in the European Data Strategy. These tensions, which span from data minimization versus data reusability, to privacy versus market openness, underscore the importance of an interdisciplinary approach to regulating data portability. 2.3.5.0.3 The Duality of Data Portability The dual purpose of data portability—to uphold informational self-determination and promote market competition—remains one of its most complex and unresolved aspects. As Graef et al. highlight,[GHP18] data portability has impliThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 84 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability cations for both privacy and competition, a view supported by Purtova,[Pur15] who stresses that while portability may enhance user control, it also introduces challenges related to market dynamics and power imbalances. The report extends this argument, showing that attempts to regulate portability solely through data protection law will inevitably have ripple effects on competition law and the design of information technologies. By focusing on the legality attentive data scientist approach, this research has identified key areas where the legal and technical aspects of data portability intersect. For instance, the determinability problem—the challenge of defining the scope of data to be ported—cannot be resolved without technical expertise regarding the attributes, data types, and metadata of the data. This is a point echoed by the scholarship,[Sha24] who cautions that dominant platforms may exploit the lack of clarity to cement their control over data formats and standards. 2.3.5.0.4 Data Quality and Fitness for Purpose The report introduces data quality, particularly its fitness for purpose, as a guiding principle that can bridge the gap between the legal and technical dimensions of data portability. This idea builds on Wang and Strong’s argument that the value of data is closely tied to its reusability and fitness for purpose,[WS96] particularly in a data-driven economy. Undeniably, data access and sharing foster competition and drive innovation, yet it is data quality that defines whether ported data remains useful across different services. By integrating data quality into the analysis of data portability, this report proposes a proportionality-based metric system that tailors the quantity and type of data to be ported based on the context. For example, in the context of GDPR, the emphasis is on informational self-determination, meaning that individuals should be able to port as much personal data as they wish. In competition law, the focus is on ensuring market functionality, so a more limited amount of data may suffice. This proportionality-based approach ensures that data portability remains meaningful across different regulatory frameworks, avoiding a one-size-fits-all solution that could undermine both privacy and competition goals. Graef et al. (2018) support this view, noting that data portability should not become a “mere formalistic exercise,” but rather a tool that balances the needs of individuals, companies, and society at large. 2.3.6 Conclusions The research on which this report is based makes significant contributions to the understanding of data portability within the European Union’s legal and regulatory framework. By employing the legality attentive data scientist approach, it reveals the interdisciplinary challenges inherent in regulating data portability and demonstrates that the concept cannot be adequately addressed without considering its dual purpose—enhancing informational self-determination and fostering market competition. The research on which this report is based advances the understanding of data portability through an array of contributions. 1. Data Portability as a Microcosm of Broader Conflicts: Data portability encapsulates many of the conflicts at the heart of the EU’s Data Strategy, including the tension between data minimization and reusability, as well as between privacy and competition. These tensions The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 85 2. Data Protection, Consumer Protection and Competition Law 2.3. Unchaining Data Portability cannot be resolved through regulation within a single domain, such as data protection or competition law, without affecting the other. 2. Technological Solutions and Policy Goals: The choice of technical solutions for implementing data portability has significant implications for policy outcomes. Systems that prioritize user control and data quality can empower individuals and promote transparency, while systems that allow dominant data holders to control the structure and content of ported data risk reinforcing existing market power. 3. Data Quality as a Central Principle: The introduction of data quality, particularly its fitness for purpose, as a guiding principle for data portability, represents a major contribution to this research. By ensuring that data retains its utility and completeness in subsequent services, this principle provides a framework for assessing the adequacy of ported data across different contexts.[Pur15] 4. Proportionality-Based Metrics: The report proposes a proportionality-based metric system that tailors data portability obligations based on the specific regulatory context, ensuring that the right remains flexible and adaptable to different legal frameworks. 2.3.7 Future work The future of data portability research lies in further exploring the interplay between data quality and existing regulatory frameworks. Taking the data quality fitness for purpose benchmark as a guiding principle of the digital world could be revolutionary. This principle has the potential to transform how data is shared and reused, affecting everything from the development of data markets to the functioning of public services, while achieving the overarching goal of the Data Strategy to give businesses and public authorities access to high-quality data. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 86 3. Privacy and Security Implementations 3Solving Practical Hurdles to Privacy and Security Implementation 3.1 Assessing Digital Identity Solutions Author: Cristian Lepore, Université Toulouse III 3.1.1 Executive summary The Internet, as it exists today, lacks robust and reliable mechanisms for trust and security. While this was not a significant concern in the early days of its development, the rapid growth of online platforms and increasing dependence on digital services have made trust an urgent issue. Digital identity offers a potential solution to many of the challenges surrounding trust and security online. For this reason, companies and governments are collaborating to provide citizens with secure and reliable electronic identification systems. For example, since 2014, the European Commission has implemented a framework for promoting equal access to public and private services throughout Europe. This initiative is known as eIDAS (Electronic Identification, Authentication, and Trust Services) and it is based on the proposed architecture of decentralization, such as the Self-Sovereign Identity (SSI). SSI aims to return control of personal data to citizens, reducing the risk of misuse and building greater trust in digital services. However, the concept remains elusive, lacking a clear, uniform interpretation. Efforts to define SSI are mostly led by academic initiatives, but the concept often remains abstract and disconnected from practical business needs. As a result, many digital identity systems claim to follow SSI principles, but it is difficult to assess how closely they adhere to its core values. For this reason, this work aims to close that gap. It seeks to offer a formal definition of Self-Sovereign Identity and introduce a practical, structured model for evaluating digital identity solutions. In the long term, this will provide both designers and citizens with a tool to help them choose the best platform for creating a digital identity. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 87 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions 3.1.2 Introduction 3.1.2.1 Background Digital services for businesses and citizens requiring an identity have become crucial for economic growth. However, every time an App or website asks us to create a new digital identity or easily log on to a platform, we have no idea what happens to our data. Nearly 72% of users wish to know more about how their data is processed when they open an account, and 63% of EU citizens want a trusted European identity.1The European Commission upstreams this challenge by providing a unique digital identity for over half a billion people. Since 2014, the European Commission has overseen eIDAS, a framework for the cross-border interplay of nationals e-identities with equal conditions for accessing public and private services [Eid]. The eIDAS aims to harmonize the use of electronic identification and trust services across Europe. It defines identification schemes that require users to authenticate to a service with a level of confidence that can be low, substantial, or high. Once authenticated, users have access to a range of services that allow them to sign documents (eSignature), ensure the origin of data (eSeal), and provide evidence (timestamp, delivery service, website authentication) [SAA22]. The provisioning of identity information starts with the issue of personal data (PIDs) for holders who store them in their wallet storage component. The holder then presents credentials to the Relying Parties [ASM21]. To meet the requirement for data flow, the Commission is working on a decentralized architecture that recalls the Self-Sovereign Identity (SSI) paradigm. Self-Sovereign Identity (SSI) is a theoretical concept from academia that promises users control of personal information [All]. The concept received great attention among practitioners and is today used as a reference guide for the design of e-identity solutions. However, SSI is still elusively defined as a set of principles, and past attempts to converge to an agreed-upon definition failed due to technical challenges [KP21], divergent visions from practitioners [Ruf18], interoperability issues [Ghi+23], and legal hurdles [SAA22]. Thus, a rigorous formalization of Self-Sovereign Identity would help design solutions and validate their completeness and correctness [SC22]. 3.1.3 Objectives and Scope. This report outlines the concept of Self-Sovereign Identity through a rigorous definition of principles. We outline concepts, relationships, and rules governing identity ecosystems’ entities and provide a formal specification of principles and a semantic taxonomy of SSI. We then delineate a model based on our implementable tweak of SSI principles to assess any worldwide digital identity system. We demonstrate our model value with an in-depth analysis of the major European digital identity framework. In the long run, we aim to enable future startups and governments 1Digital Identity for all Europeans; a personal digital wallet for EU citizens and residents. https://commission.europa.eu/strategy-and-policy/priorities-2019-2024/ europe-fit-digital-age/european-digital-identity_en The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 88 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions to rank solutions, spot weaknesses, and intervene accordingly. Our model also helps citizens to choose their best-fit e-identity system. 3.1.4 State of the Art SSI was firstly described as a set of principles in [All]. Subsequent works complemented these principles with social considerations [And16;Gil+20], along with a technical analysis [FCA19]. Principles were categorized in a three-way taxonomy by Tobin [TR16]. Andrieu provided a techfreedescriptionof SSI [And16], while otherpapersextendpreviousdefinitionstocoverblockchainbased e-identity systems [FCA19] and add security properties [Gil+20]. Sheldrake considers only essential principles of Self [She19]. Lastly, a business-oriented analysis of SSI results from [Blo]. Sometimes these definitions have been used to formalize the archetype for SSI systems. For example,[Sch+21] uses an empiricalapproachtosystematizee-identity offeringsinto eight archetypes and evaluates solutions. Satybaldy et al. [SNE20] propose a framework to describe, evaluate, and compare SSI systems. Tobin [Tob22] evaluates the eIDAS framework. However, those works need to define criteria for evaluation – otherwise, the contribution of their content analysis is limited. All works revealed the need for more effort to produce an agreed-upon evaluation model. Furthermore, all works require periodic adjustments anytime the identity system updates. 3.1.5 Methodology Our state of the art results from a systematic review study. A systematic literature review provides a coarse-grained overview of the research field through several steps: 1) defining the research questions, 2) conducting the search, 3) screening papers, 4) defining criteria for inclusion/exclusion of articles, and 5) extracting knowledge [PVK15]. First, we define broad research questions to cover the topic from interdisciplinary perspectives. From keywording research questions, we produced search strings and hit online databases to gather articles. We define classification criteria, filter out articles and categorize results. We plot results into a two-axe chart to help spot literature gaps. The steps are detailed as follows. 1. Defining the research questions. We aim to span the scope of the research field through broad search strings. We then use search strings to hit online databases. A scientific approach to comprehensively cover a topic of interest is to explore the intersection of legal, social, and technical fields (LST) [Cus+10;Bad+13]. First, we produce general research questions for each field of study (legal, social, technical). As a tool to guide us in designing research questions, we refer to the Five W’s (What, When, Where, Who, and Why) [Har96] and propose at least one or more research questions for each "W". Table 3.1 reports our research questions. Rows are the five W’s; columns are the fields of research. From keywording the research questions, we produced the following search strings. •Legal: ("governance" AND "digital identity") OR ("legal framework" AND "norms" AND "identity") OR ("rights" AND "regulation" AND "digital identity") The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 89 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions still need to notify an identity scheme under eIDAS.4 Outcome: 42%. Individuals can authenticate to a service and prove their existence under eIDAS in several ways, but a different definition of LoA among Member States generates entry-level barriers. •Persistence. Who can issue attributes? How many IdPs can provide the same attribute? Evaluation. Member States appoint special service providers as TSPs and QTSPs under the conditions specified by the Regulation. Although flexibility exists, the requirements are precisely defined by the European Union. Several trust service providers can issue the same trust services, thus producing overlaps. Examples exist in Italy where Postecom S.p.A. and InfoCert S.p.A. can issue the same attributes. By contrast, the regulation does not allow any other public or private body to issue trust services unless they become TSP. Outcome: 33%. Only trust service providers can issued an e-identity under eIDAS. Although, some attributes can be attested by many TSPs. •Protection. Who maintains the list of IdPs and SPs? Evaluation. Member States establish, maintain and publish the trusted list of TSPs and QTSPs (Article 22 (1) of the eIDAS regulation). The national supervisory bodies appointed by each Member State are responsible for the management and security of the list and notify the Commission (Article 22 (2)(3) of the eIDAS). The private sector, consortium, and other institutions cannot be a Trusted List maintainer if not appointed by the Member State. Outcome: 23%. Only authorities within the list of TSPs and QTSPs can issue e-identities. Table 3.4 (a) summarizes the evaluation for Individuals’ rights. σis 32%, with all principles under 50%. •Access. How users obtain information about their attributes? Can users access the list of IdPs? Evaluation. The Commission foresees a European Wallet to obtain, store, and issue credentials [Com23]. Notifications are periodic bulletins in the form of pop up that announce relevant information about the use of digital identity. The wallet functionalities listed in the ARF are precise, although the wallet implementation is ongoing, and no functionalities for a notification system are specified yet. The European Commission maintains a central map of trust service providers,5provides access to the Trusted Lists of all EU member states. Users can browse the list and retrieve information about each country’s trusted identity providers and service providers. Outcome: 60%. The Commission plans to introduce a European Wallet for obtaining, storing, and issuing credentials. To date, there is no specific implementation of the wallet. 4Overview of pre-notified and notified e-identity schemes under eIDAS.https://ec.europa.eu/ digital-building-blocks/wikis/display/EIDCOMMUNITY/Overview+of+pre-notified+ and+notified+eID+schemes+under+eIDAS 5EU/EEA Trusted List Browser. https://www.eid.as/tsp-map/#/ The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 96 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions •Control. Do users negotiate the release of attributes to SPs? Evaluation. Attributes exchange is service-dependent but generally possible with most service providers. PIDs are, instead, not negotiable as they are a unique set of personal information used to identify users and are difficult to negotiate with x509 certificates. Referring to service providers, users can freely choose the SP they wish to authenticate to, as long as it is in the list of trust service providers published by Member States. Outcome: 50%. The eIDAS gives only a limited possibility to negotiate attributes from a pre-set list of service providers. •Transparency. Are policies and rules to manage ecosystem members transparent and clearly stated? Evaluation. The eIDAS does not provide just guidelines but rules for electronic identification, authentication, and trust services. The regulation defines policies, procedures and promotes transparency through mutual recognition, which allows e-identities in a Member State to be recognized and accepted in other Member States. Outcome: 67%. The eIDAS establishes rules instead of general guidelines for the electronic identification. Table 3.4 (b) summarizes the evaluation for Trustworthiness. The σvalue is 59%, more that halfway from mere SSI, with no principle below 50%. •Consent. Does consent result adequately expressed and managed? Evaluation. Consent management is treated under several texts. The GDPR lays down the rules for processing personal data in the information society services in Europe, and (the GDPR) specifies conditions for consent to be valid (specific, informed, freely given, and unambiguous). However, the regulation leeways Member States with a certain degree of room to specify further norms. For example, ages specified within the GDPR do not apply to certain services (e.g., adult video content). Dynamic consent is not mandatory, while the service provider determines the best-fit consent management solution for its needs. So far, the GDPR and eIDAS do not foresee implementing other (post) consent management systems. Outcome: 22%. The eIDAS does not specify consent models other than the classical regulated under the GDPR. •Minimization. Does the service lawfully collect only the minimum amount of information? Do users employ techniques to limit data sharing? Evaluation. The service provider must collect adequate and relevant data for data processing. The digital certificates (e.g., the x509) allow sharing a subset of attributes but not the single attribute, or the associated information. Outcome: 23%. Data minimization is extremely compromised under eIDAS. Table 3.5 (a) summarizes the evaluation for Secrecy. σis stuck at 22.5% with all principles below 25%. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 97 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions •Cost. To what extent does the digital identity is profitable for stakeholders? Evaluation. The framework aims to decrease costs for public and private services and benefit citizens by fostering market competition [Dom20]. However, the lack of official studies and available data limits our assessment. Outcome: 50%. The framework aims to decrease costs for services in Europe. However, lack of evidences to prove it affect the score. •Interoperability. To what extent can IdPs attest attributes to SPs across different jurisdictions? Evaluation. The regulation enables cross-border transactions between public and private services. Several international transactions exist for exchanging goods or opening bank accounts between EU countries. Thus, the discriminant factor is the roomy coefficient; public and private services act multi-country within the European Union. Nevertheless, not all Member States declare an e-identity Scheme under eIDAS. Outcome: 75%. On paper, the eIDAS is an interoperability framework, but there is no evidence of concrete results. •Portability. To what extent can users transport the list of attributes on different ecosystems? Evaluation. Although transferring personal data across countries is an important aspect of data portability, the regulation does not focus on data portability, regulated by the GDPR and reinforced by further texts currently under discussion (e.g., the Data Act). It foresees the possibility for users to export data creating a copy of the wallet. No solution from the industry exists today. Outcome: 0%. The eIDAS does not foresees data portability. •Standard. Who can issue standards for e-identity systems? Evaluation. International organizations like ISO can promote standards for credentials exchange and electronic signature under eIDAS. For example, by implementing the XAdES protocol in eIDAS, trusted service providers, such as certificate authorities, can offer compliantsolutions for creatingand verifyingXML-based advancedelectronicsignatures [RAC21]. This standard enables organizations and citizens to use advanced electronic signatures securely and reliably for various purposes, such as signing contracts, submitting official documents, and more. Public agencies (e.g., European Committee for Standardization (CEN) and European Telecommunications Standards Institute (ETSI)) can play an important role in defining formal specifications for the eIDAS protocol. Outcome: 100%. The eIDAS employs a combination of standards from the industry, working groups, and academia. Table 3.5 (b) summarizes the evaluation for Sustainability. σscores is 56%. Three principles over 50%; one is at zero. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 98 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions 3.1.8 Research Analysis The eIDAS obtained a Global Score of 42,4%, almost halfway from a perfect SSI solution. The analysis of functional dependency reported <32; 59; 22; 56 >(values are in percent (%)). We plot results in a Kiviat chart to visually interpret relationships among categories (Figure 3.2). Our Kiviat chart consists of four axes radiating from a central point, each axis representing a category of our taxonomy. Outcomes from each category are along the corresponding axis, and a line connects points between adjacent axes. The thicker (blue) rumble is the eIDAS assessment. The dashed (outer) rumble represents the guideline for a fair SSI solution with the same values for all categories. The eIDAS performs well in two categories (Trustworthiness and Sustainability) with grades closer to 60%, with Trustworthiness marking the highest score. Two categories score less than 40%, with Secrecy stuck at 22,5%, indicating that work remains to enhance systems’ privacy and secrecy design. An overall analysis of principles shows that six principles score below 50% with four principles ≤25% and one at zero (Portability). Portability suffers from the lack of a standard for the EU wallet (the ARF provides only guidelines) and discrepancy of regulations in Member States. Five principles evaluate over 50% with Interoperability and Standard scoring significantly high (75% and 100%). Here the effort of the EU in pushing for a cross-border framework is worth it. Standard reflects the involvement of a manifold of players, although most likely rely on dated standards. We measure the variation or dispersion of results from principles from the average, namely the standard deviation [LIL15]. A low standard deviation implies that values cluster around the mean, whereas a large standard deviation indicates that values are more widespread. We calculate the standard deviation by taking the square root of the variance of the values subtracted from their average value [Wac09]. The Formula 3.1 computes the standard deviation. (x−µ) = v u u t 1 N N X i=1 (xi−µ)2(3.1) where xiis the ith principle value as resulting from step 3 of Algorithm 1, µis the mean of xvalues, Nis the number of principles (population size). The resulting standard deviation is 29.6, which suggests that values vary widely from the mean of 45%. This dispersion of results, particularly from the score of the Secrecy, affects the global evaluation. In the other three categories (except the Secrecy), about 60% of the principles score higher than the mean, and 80% of them above a score of 40%, confirming the effort of the Commission to foster several aspects of the framework in different disciplines. Only two principles: Protection and Portability, in three categories (except the Secrecy) score less than 25%. For both, the diversity of regulations from Member States lowers the result. The flaws of eIDAS are twofold. First, the Secrecy. Approaches like consent management and data minimization still refer to the old way of managing consent and sharing personal information. Consent management was dated back to the early Internet days when websites were different and mobile applications did not exist. Informed consent through pop-ups is tedious, and understanding what is "informed" is based on a case-by-case study. On the other hand, The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 99 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions data minimization is impossible through digital certificates in the form of x509, and policymakers should enforce legal value to coming standards for data exchange like Verifiable Credentials and Presentations. The second issue of eIDASspreads across (almost) all principles and concerns the European Union. Each Member State has a leeway to implement the framework through slightly different approaches that reflect, for example, the various implementations of the LoA or the design of the EU wallet. The Commission provides guidelines for the wallet functionalities through its Architecture Reference Framework (ARF). Four working groups provide use cases to test the wallet functionalities. Although the process may take longer, the final implementation will come from the collaboration with the Member States, which in the end, will play a crucial role in implementing the wallet. They will be the finalizer of the wallet functionalities. Yet, without Member States working together on common standards, the risk is to disperse the effort and obtain several wallet proposals, each with pros and cons. A better approach would be to find everybody working on one prototype for the wallet (maybe two) with strict constraints from the Commission. That would speed up the process and avoid discrepancies in actions. Our assessment is helpful in pinpointing weaknesses in the design of identity systems. For example, eIDAS proves weaknesses in consent management and data minimization. On the other hand, it strengthens the acceptance rate of solutions. Acceptance is users’ confidence in using an eidentity [Acc], which results in crucial system design. If people trust an e-identity system, they will be willing to adopt it. 3.1.9 Recommendations Our findings open the discussion on the quest for a perfect SSI solution. First, a mere SSI might only be worthy for some people. For example, having a system that exerts exclusive user control of the recovery key requires technical skills. If users lose the key, the wallet and its content will be burned. It is like being in a foreign country and losing the passport. Alternatively, someone (maybe a government) should have access to the wallet to back up its content periodically. On the other hand, a perfect SSI is also not feasible. Parameters are strongly tied, and increasing a score in one modifies the others. That means the eIDAS rumble of Figure 3.2 can be enlarged to a certain extent but not achieve a perfect diamond with corners at 100%. We believe a corner can be stretched as long as others are squeezed. Thus, an ideal SSI remains a utopistic idea. As a second point for discussion: corporations might have the power to stretch the diamond, but they might not be willing to invest in it, as opposed to startups. Suppose a mature solution ranks high in half of the parameters and zero in the others. It can receive a Global grade of 50%, still higher than a startup averaging 40% in each parameter. Apparently, the corporate solution is better than the other, but this analysis would not consider different aspect of the system. For this reason, a better definition of SSI would encode a combination of Global Score and σvalues, where the constraint for the Global Score Gwould be relaxed to justify the name "quasi" as long as the constraint for the functional dependency stays the same. The functional dependency study can be used to strengthen the constraints and force solutions to respect parameters in each development area. We then leverage a more pragmatic definition for SSI, as follows: Definition 1 (Quasi-SSI).Quasi-SSI is a lighter and pragmatic version of SSI that respects criteria 1 and 2, defined as follows: The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 100 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions Figure 3.2: A Kiviat chart reporting the Global Score and each category σvalue. The thicker (blue) rumble is the outcome of eIDAS. The inner thinner (green) rumble is our pragmatic definition of SSI. The dashed (outer) rumble represents the guideline for a fair SSI solution. 1. G≥40% 2. σi≥35% ∀σi∈Σ Where Gis the global score and σiresults from Step 5 of the functional dependency (Algorithm 1). The first criterion relaxes the constraints of SSI to a Global score above 40%. The second reinforces criterion 1 and pretends solutions will perform significantly well in every category of the taxonomy: all categories should have a score ≥35%. The overall result imposes constraints in the progression of development in each category (Individuals’ rights, Trustworthiness, Secrecy, and Sustainability). The two conditions must happen in conjunction. According to our new definition of quasi-SSI, eIDAS fails in the second criterion and cannot be considered a quasi-SSI solution. As an outcome of our analysis, we can refine the previously stated notion of acceptance in light of our new definition of quasi-SSI. Acceptance is the degree of confidence of users in using an e-identity system that, aftermath applying our model, respects the parameters defined in Definition 1. From the analysis of eIDAS through our model, we provide six legal and four technical recommendations for eIDAS from the twelve principles we analyzed. We specify between parenthesis the principles from which our recommendations come. From the legal advice: 1) the Commission should work to decrease the ambiguity of LoA, specifying parameters that are unique for Member States to follow. Today, Member States have tuned those parameters for LoA, generating inequality across e-identity schemes in the European Union. 2) Streamline the procedure for The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 101 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions service providers to become TSPs (Persistence). The process to become a TSP lasts from a few months up to a year (sometimes even more). Due to the importance of TSPs in the framework, the Commission should act to decrease the time to become TSP without relaxing the requirements. This would lift the number of TSPs from 14% in 2021 to the goal (of the Commission) of 80% by 2030. 3) Move the management of the list of service providers from Member States to a "super partes" entity of the European Union or an open community of contributors (Protection). Our objective is that no authority from a Member State shall control the list of IdPs/SPs, thus decreasing the risk of censorship from the Member States. 4) In terms of costs and benefits, once the revision of the regulation is applied, and the wallet adopted by the Member States free of charge for businesses and citizens (first quarter of 2025), analyzing eIDAS in terms of cost saving for citizens, public and private services (Cost). That would narrow the focus of the Commission to employ resources where needed. 5) Add a chapter in eIDAS that specifically addresses governance-related issues and portability to embrace and help quickly adopt coming standards from the technical field (Portability). Standardization will be key to interoperability between different platforms and networks, enabling the seamless use of identities, avatars, data, virtual assets, experiences or environments, and the associated rights across platforms and networks. This would also help to 6) embrace cutting-edge standards (e.g., the Verifiable Credentials) and speed up the process to give them legal value. Along with these recommendations, we highlight a few technical challenges. 1) The Commission should allow citizens to negotiate Personal Identification Data (PID), embracing coming standards different from digital certificates (Control). 2) Implement data minimization so that users should not disclose all personal information when they provide their identity (Minimization). 3) Enable consent preferences to travel along with claims (e.g., digital certificates, verifiable credentials, etc.) and provide a dashboard for users to manage consent preferences transparently (Consent). In general, more effort should be made by the Commission to extend the classical consent management proposal to include post-consent solutions. For example, controlling the flow of information and providing users with a dashboard to manage their preferences online. The final recommendation 4) the Commission should revise the use of middleware to foster interoperability (Interoperability). 3.1.10 Conclusions The Self-Sovereign Identity (SSI) ecosystem is young and immature, with a manifold of actors providing solutions that follow common design patterns. Yet, no tool exists to assess those offerings, thus allowing anyone to declare adherence to SSI. This work proposed a formalization of concept to study the convergence of systems to SSI. We tested our model on the major European identity framework (eIDAS) and discussed the results. The model promises to rank solutions, streamline the design pattern, and spot weaknesses. In the long run, it can be used to determine the historical progression of e-identity systems toward SSI, as opposed to periodrelated rankings. Our Comprehensive Rating System (CRS) is left purposely general to avoid overfitting it on specific solutions, and we also generalized some featuring aspects of eIDAS during the assessment. Many implementation-specific parts were left blank for practitioners to fill in and were not tested by our model. For example, we generalized privacy-related issues of The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 102 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions wallet authentication and trust services with LoA. However, the potential of our model is the process of moving from a list of theoretical principles to practical evaluation, and we are aware that in this migration process, we lose some information. Today, parameters and weights are based on the literature review, but often we cannot find proper guidelines to help us choose values for the weights and their ratio. In the future, we aim to strengthen the quality of our model, sharing parameters and weights with a recognized community of experts and conducting surveys to tune values. With this, we aim to achieve recognition from a global community of experts. Meanwhile, testing further industry-based initiatives will be a plus value in refining our parameters. Our model’s current state of the art led to a more pragmatic definition of SSI as quasi-SSI that promises to guide the future development of e-identity solutions. Quasi-SSI constrains our tweaked definition of SSI to distinguish practitioners doing well from solutions needing further improvement. Through a better model, we can also improve the definition of quasi-SSI to represent real-world values. The final breakthrough is to push policymakers and practitioners to strengthen the state of the art of digital identity to meet the promise of SSI and pave the road of the research field in evaluating e-identity offerings The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 103 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions Algorithm 1: Functional dependency (pseudo-code) Data: w1, . . . , wn∈Wsuch that w∈ {5,7,8,10,15,20};/*set of weights */ Input: mini-batches βi,βiis a category {w|w⊆Wweights of dimensions } βi={[(w1, c1),(w2, c2), . . .],[(wj, cj), . . .],[. . .]}where ∀i= 1, . . . , n, ci∈ {0,0.5,1} Output: Σ = {σ1→σ2→σ3→σ4};/*functional dependency */ Σ← ∅ In parallel for each mini-batch βido for each challenge Cdo Step 1: multiply weights by coefficients w′ j=wj×cjfor each dimension dj∈C Step 2: sum results Sc=Pdj∈Cw′ j Step 3: normalize weights ∥Sc∥=(Sc−wmin) [wmin,wmax];/*Min-max normalization */ end Step 4: compute weighted average σi=Avg P∥Sc∥;/*Avg of norms in mini-batch */ Step 5: collect results as a power of ten Σ∪ {σi×100};/*Cartesian product of 100 */ end Return: Σ← {σ1, σ2, σ3, σ4} The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 104 3. Privacy and Security Implementations 3.1. Assessing Digital Identity Solutions Table 3.4: Individuals’ rights an Trustworthiness assessment. Pts stands for points. Individuals’ rights (a) Principle Challenge Dimension Weight Coeff. Existence - What attributes can attest to an e-identity? - Assigned attributes (Username and Password) 5 pts • - Multi-Factor Authentication 7 pts • - Accumulated attributes (PIDs) 10 pts ◦ - Inherited attributes (Biometric data) 10 pts • - Digital certificates (e.g., x509) 10 pts ◦ - Other attributes (e.g, eSign, timestamp) 10 pts ◦ Persistence - Who can issue attributes? - Qualified Trust Service Providers (QTSPs) 10 pts • - Trust service providers (Non-Qualified) 10 pts • - Other public bodies (e.g., Gov agencies, Univ.) 10 pts ◦ - Other private bodies (e.g., Microsoft, Okta) 10 pts ◦ Protection - Who maintains the list of IdPs and SPs? - Private sector (e.g., banks, credit bureaus) 5 pts ◦ - Consortium of organizations (e.g., Kantara) 8 pts ◦ - Gov agencies (e.g., national identity authority) 10 pts • - Other gov institutions (e.g., EU Commission) 10 pts ◦ - Community of contributors (open) 10 pts ◦ Trustworthiness (b) Principle Challenge Dimension Weight Notify Access - How users obtain information about their attributes? - Can users access the list of IdPs? - Local agent (wallet) 10 pts ⊙ - Shared ledger of IdPs 10 pts • Control - Do users negotiate the release of attributes to SPs? - User negotiates attributes but PIDs 10 pts ⊙ - User negotiates PIDs 15 pts ◦ - Users can choose the service provider 15 pts • Transparency - Are policies and rules to manage ecosystem members transparent and clearly stated? - General guidelines only 5 pts ◦ - Transparent rules and procedures 10 pts • Table 3.5: Secrecy and Sustainability assessment. Pts stands for points. Secrecy (a) Principle Challenge Dimension Weight Coeff. Consent - Does consent result adequately expressed and managed? - Informed consent 10 pts • - Dynamic consent 15 pts ◦ - Post-consent 20 pts ◦ Minimization - Does the service lawfully collect only the minimum amount of information? - Do users employ techniques to limit data sharing? - Adequate and relevant data collection 10 pts • - Minimization of attributes but PIDs 10 pts ◦ - Minimization of PIDs 15 pts ◦ - Transfer only a subset of attributes 5 pts • - Transfer one attribute at a time 10 pts ◦ - Transfer the associated information only 15 pts ◦ Sustainability (b) Principle Challenge Dimension Weight Coeff. Cost - To what extent does the e-identity is profitable for stakeholders? - Profitable for public services 10 pts ⊙ - Profitable for private services 10 pts ⊙ - Profitable for citizens 10 pts ⊙ Interoperability - To what extent can IdPs attest attributes to SPs across different jurisdictions? - Among public services 10 pts ⊙ - Among private services 10 pts ⊙ Portability - To what extent can users transport the list of attributes on different ecosystems? - Between public authorities 10 pts ◦ - Between private authorities 10 pts ◦ - Full portability 15 pts ◦ Standard - Who can issue standards for e-identity systems? - Working/Community groups 10 pts • - Industry sector (e.g., Avast) 10 pts • - Public agencies (e.g., gov) 10 pts • The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 105 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies to regulations (RGPD, Data Act, Data Governance Act, AI Act, Open Data EU regulation...). There is a need for scientific research to develop a framework that aligns data quality criteria with the evolving regulatory landscape, particularly in the context of community-driven data spaces. 3.2.5.1 Ensure reliability of distributed systems Reliability is the likelihood that a system, including hardware and software materials, will successfully perform its intended tasks under specified conditions and over specified periods of time [AW13]. Distributed reliability extends this concept to distributed systems. Akimova et al. [AST18] define the reliability of distributed systems as the ability to maintain the defined criteria required to perform a given task under the influence of failures, breakdowns, hardware-based or human errors, etc. In the context of data processing, distributed system reliability means that each step is performed with fault tolerance by all the technologies involved, resulting in high quality outputs. The reliability of distributed systems depends on the systems, the data life cycle phases (data collection, data preparation, data analysis, etc [SPM21b]), the data transactions and, most importantly, the quality of the data outputs. It is therefore necessary to assess the reliability of all entities in distributed systems, i.e. the ability of each component to perform correctly and not degrade the quality of the data. Future research should focus on to create data quality contracts at each phases of the data life cycle based on appropriate data quality criteria. 3.2.5.2 Increase trust in decentralized governance Trust is a subjective concept characterizing a relationship among two or multiple entities with a common purpose [MFGL12]. In decentralized data governance, it is the belief that all participants in that data governance are willing and able to produce high quality data outputs. Indeed, having multiple entities managing data raises trust concerns, as the self-interest of one entity processing the data may conflict with the overall benefit of other entities [Yan22]. Assessing the trustworthiness of individual contributors and their commitment to producing high quality data outputs is essential. This need leads to the following questions: are data governance stakeholders able to make the right decisions to maintain data quality? What are the data quality criteria that can be used to assess trust in all data governance stakeholders based on their actions and decisions? What are the data quality criteria relevant to data governance? 3.2.5.3 Achieve compliance with regulations The concept of a Data Space implies a community-based approach for managing and exploiting data. Within such a community, commitments need to be formalized through contractual agreements, such as confidentiality or non-disclosure. This raises questions about how the underlying layers will ensure compliance and define a level of adherence to regulations. How can we formalize data quality through contracts, guarantee it, and establish a means of evaluation? Especially, regulatory requirements may change according to the type of data and its use. To the best of our knowledge, there is no existing work that categorizes data quality criteria according The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 112 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies to regulations (RGPD, Data Act, Data Governance Act, AI Act, Open Data EU regulation...). There is a need for scientific research to develop a framework that aligns data quality criteria with the evolving regulatory landscape, particularly in the context of community-driven data spaces. 3.2.6 Methodology To reach the resolution of the aforementioned challenges, we faced two main tasks: •Identify the relevant data quality criteria (DGC), •Classifying these criteria by trust, reliability or legal compliance. We managed these tasks by using respectively theoretical and practical methodologies which encompassed specific sub-steps. First, we conducted a systematic literature review to identify the DQCs. Then, we consolidated this list based on discussions and recommendations from data experts. Finally, we classified the resulting DQC. The remainder of this section describes all details of each steps of our methodology. 3.2.6.1 Systematic literature review This part describes the methodology to conduct our systematic literature survey. It explains how we selected and redefined the resulting DQC collected from the survey. Our methodology is similar to those of Shah et al. [SPM21c;SPM21b] and Bowling Ann et al. [Bow05]. The survey was structured in four main steps: 1. Formulation of the research questions. 2. Identification of the pertinent research works. 3. Selection of best articles through inclusion and exclusion criteria. 4. Analysis and verification. We realized a comprehensive analysis of the current surveys and research works on DQCs. We chose IEEE Xplore, Science Direct, ACM, Springer Link, Web of Science and Google Scholar. In each library, we formulate research queries to identify the pertinent DQC. The requests generated 3480 documents. Then, we picked the articles with headings related to "data" and "quality" or "information" and "quality". This resulted in 604 relevant scientific reports for our investigation. Moreover, we identified the top 30 research articles based on their number of citations. The last article on the list had 24, whereas the first had 2412. Next, we studied all the data quality criteria presented in these papers. These papers described a total of 270 DQC. However, several research articles proposed the same DQC. Thus, we regrouped the DQC having the same name and a similar definition to avoid redundancy. Then, we realized there were just 30 common DQC in the entire set of papers. Figure 3.4 summarizes this systematic literature review and the criteria are presented by figures 3.5, 3.6, 3.7 and 3.8. However, we proposed an exhaustive description of this systematic The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 113 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies Figure 3.4: Selection of relevant research articles The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 114 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies literature review in at the conference CSE 2023 untitled "Towards Reliable Collaborative Data Processing Ecosystems: Survey on Data Quality Criteria". 3.2.6.2 Semi-directive interviews with data experts This section describes our strategy for proposing a standardized framework of DQCs that is taking into account the overall context of collaborative data spaces in Europe and its corresponding challenges of reliability, trust and legality. It focuses on our list of 30 data quality criteria. As DCQs are essential in our mission, it was mandatory to validate the methodology of the systematic literature review and the resulting list of criteria in order to ensure the significance of our solutions. Then, we take advantage of the secondment program of LeADS Project to discuss our methodology with several professors in cybersecurity; business and innovation; and laws. Our objective is to validate and select the most relevant criteria by gathering feedback from several data experts. These interviews will allow us to evaluate the relevance of each criterion and understand how to apply them to real data sets. These exchange is a practical validation of the list of criteria to reduce the gap between academic methodology and industrial applicability. So, we had the opportunity to realise the aforementioned interviews with internal and partners having highest interest on quality of their data during the secondment at Intel Corporation (Belgium). The interviews at Intel Corporation aimed to gather the expertise of professionals and partners to help us to recognize ‘quality’ data sets. For each professional, they consisted in collecting firstly their opinions on the relevant criteria (their criteria, definitions, evaluation levels, etc.), then asking comments and feedback concerning our criteria, definitions and evaluation levels of DQCs. So, we selected nine data experts working in European companies and having important experiences on data management. They were coming from different backgrounds such as: cyber security (02) ; data privacy, protection and compliance (04) ; supply chain data management (01) ; statistics and performance (01) ; data documentation (01). Each interview was conducted online between 30 and 45 minutes. The questionnaire which led the discussions contained open and semi-directives questions such as: •Do you use datasets as part of your business? What type of data? Where does it come from (internal or external (linked to a contract or not)? What type of processing do you carry out on this data? •Have you encountered any problems with data quality during processing? •How do you define a dataset of good quality? And what are the quality criteria? •What definition do you propose? What are the evaluation levels for the criterion? Does this criterion enhance the trust, reliability or regulatory compliance of the data? We recorded all the interviews then extracted the relevant information of interest by taking into account the context of their activity and nature of their data. These discussions generated many information concerning data management and DQCs such as: The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 115 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies •31/32 criteria mentioned by the interviewees plus 4 others criteria •69 comments on the mentioned criteria. •These comments can be represented as full acceptation (06), recommendation on definitions or names and details on criteria and their evaluation methods (59), then the rejections (04). •The activities of the interviewees concerned the mandatory steps of the data life cycle and the governance task (collection, preparation, analysis, reuse and feedback, sharing, governance) 3.2.7 Research Analysis This part of the analysis concerns the solution we proposed to establishtrust, reliability and legal compliance in collaborative data processing. This solution will be based on blockchain which is a set of secure and distributed storage nodes. 3.2.7.1 Classified data quality criteria Performing the systematic literature review and the interviews with data experts allow us to gathered list of pertinent and consolidated DQCs which can be categorised in a data life cycle. This results is presented in the following figure "3.2". This picture describes the main contribution of our research. The methodology we used generate 4 categories of data quality criteria. Each categories contains criteria corresponding to a specific collaborative data processing challenges. Furthermore, this result proposes a step-bystep procedure to accomplish the aforementioned challenges. Security & infra and Data oriented DCQs: can be associated to reliability. In fact, distributed reliability consists in evaluating the ability of each component of the distributed system to perform correctly with fault tolerance and not degrade the quality of the data. The security & infra oriented DQCs focuses on the correct performance and fault tolerance part of the challenge. They ensure hardware and software of the overall ecosystem are aligned to perform successfully under specified conditions and over specified periods of time. The maintenance of these criteria is paramount because they prevent breakdowns, failures and hardware-based or human errors. On the other side data oriented DQC focuses on the quality of the content and structure of the data. In fact, as data is part of the information system it must be valuable and quality data, otherwise the overall system must not be considered as reliable. The content must be accurate and believable to generate valuable and safe results, when the structure must be built to facilitate the usage and reuse. Data oriented DCQs must be carefully guaranteed to justify the investment in system performance and keep the reliability of distributed system. These criteria need more attention during the data generation and collection. purpose and process oriented DQCs: can be considered as properties of trust. As mentioned above, trust characterizing a relationship among two or multiple entities with a common purpose [MFGL12]. Purpose oriented DQCs evaluate the commitment of data governance to proThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 116 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies ducing high quality data outputs. Also, they guarantee the interests of all entities managing data are not conflicting and whether they are aligned with regulations. They ensure the purpose of the data processing is valuable and aligned with data life cycle, and data reuse is compliant. They help to align data quality criteria with the evolving regulatory landscape. They help to determine the appropriate requirements according to the type of data and its use. They can be associated as indicator of compliance in collaborative data processing. Furthermore, process oriented DQCs assess the trustworthiness of all stakeholders in data management by tracking their actions and decisions on data. These criteria inform about who manage data, where, how, and when? They ensure transparency of data generation and lifecycle so that people will be able to trust the data governance. Finally, security & infra and data oriented will enforce reliability which is a sub-criterion of trust. When, purpose and process oriented DQCs will increase the trust in collaborative data processing by bringing respectively compliance and transparency. This overall analysis generated more meaningful information. It provided standardized definitions and names based on theoretical and industrial contribution which take compliance and legal boundaries into account. Furthermore, it gives the connection between all the criteria to understand how they support and impact each of them. 3.2.7.2 A blockchain based solution for data quality criteria traceability We implemented a solution to using blockchain and smart contracts to track datasets and DQC information. This platform aims to guarantee transparency, reliability and trust in data management. As centralized solutions lack of transparency, we aim to implement a storage platform using blockchain. In fact, centralized platforms such as unb.ca/cic/datasets and data.gouv.fr propose datasets and give details about their data generation without providing proof and traceability. For example unb.ca/cic/datasets, provides IoT, Dark web, operational technology and DNS datasets. Each dataset on the platform is described by a document called data paper which detailed the data lifecycle. The platform also gives the possibility to download the dataset. However, data users are not able to verify the data lifecycle and have clear information regarding DQC. To best of our knowledge, they can see information concerning the origin of the data but we are not able verify directly that origin, more, they can not reproduce these datasets. They have to trust the platform without real information. The trust relationship between data users and providers is completely focused on central party: the platform. At the opposite, blockchain gives high traceability and can not perform without transparency and awareness of all members in the consortium. It allows all members in the consensus to participate to the verification of new transactions. All action is transparent in the blockchain. To increase trust in data processing, we propose to use smart contracts to describe the protocol that manages information within the blockchain. A smart contract is a computer program that describes a transaction protocol to automatically execute, control or document events and actions according to the terms of a actions specified by the contract. A smart contract is a transparent secured stored procedure, as its execution and codified effects cannot be manipulated without modifying the blockchain itself. Therefore, smart contract limit trust assumptions. In our solution, any dataset creator can execute our smart contract to register his dataset as well as the metadata used to evaluate DQC. When this dataset depends on another dataset The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 117 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies (e.g. the produced dataset reuse data from another dataset), this relation of dependency can be specified too. As a consequence, the final user of any registered dataset can evaluate its quality by assessing the metadata of the dataset and the original datasets just by quering the blockchain transactions. Therefore, anybody can retrace the generation and reuse process of any datasets. This information is trustworthy because information inside the blockchain cannot be modified, and smart contracts available and executed by the nodes of the blockchain using a consensus algorithm. 3.2.8 Conclusions and future work This document describes our analysis of distributed reliability, blockchain-like technologies and trust in data processing . Starting from defining main component of data processing, we presented the way collaborative data processing in Europe raises challenges of trust, reliability and legal compliance. Knowing data quality is a central parameters to evaluate these needs we used a systematic literature review and interviews with data experts to identify the relevant data quality criteria before classifying them by trust and reliability. Moreover, we propose a blockchain-based solution to implement some criteria and demonstrate how this solution can maintain data quality and ensure reliability and trust in decentralized data management. However, this solutionis nottakingthe exhaustive listof dataquality criteria. A future direction ofthe work will concern an automated data quality assessment that describes the use of algorithms, rules, or models to assess trust and reliability without the need for direct human intervention. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 118 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 119 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 120 3. Privacy and Security Implementations 3.2. Distributed reliability and blockchain-like technologies The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 121 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing Figure 3.7: DDG Evolutionary Process pected tasks, (ii) the space needed to store data allowing its normal functioning, and (iii) the cost of communicating with other components to exchange requested data. When it was impossible to minimise all these three parameters at once due to conflict of conditions, the ‘best’ compromise has been selected. The criteria used to set the equilibrium entails giving the highest priority to the minimisation of the execution time requested to carry out the tasks, secondly to the minimisation of the communication overhead with other components, and finally to the minimisation of the amount of space required. The proposed approach is flexible and adaptable to the needs of the different sectors and stakeholders, and aims at increasing the value of data and minimize data processing related costs and risks. 3.3.4 Research Findings 3.3.4.1 Decentralised Data Governance This section presents the final iteration of the Decentralized Data Governance (DDG) model and describes the mechanisms for achieving and implementing the required provenance and reliability features while also respecting the privacy guidelines set by GDPR. 3.3.4.1.1 Reference architecture Figure 3.8 presents the high-level architecture of the Decentralised Data Governance (DDG) framework, which is an evolution of the architecture proposed in the first submission made to AIAI 2022 conference . It has been subsequently improved by integrating the Federated Learning as core component in the data processing layer. Moving outward, it stresses the importance of security, privacy and access control on one side, and ethical, legal and regulatory compliance on the other. These are combined with a blockchainenabled transactions tracking mechanism, ensuring transparency throughout the system. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 128 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing Figure 3.8: Decentralised Data Governance Framework The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 129 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing Figure 3.9: Overview of Privacy and Usage Control component The proposed framework consists of three layers: the data layers, the analysis layer, and the user layer. The data layer is responsible for the preparation of the data to be consumed by the processing layer. This includes cleaning and normalisation so that data is of good quality, as well as an anonymization process before any data activities commences. It follows a decentralised storage approach where each resource remains autonomous. The analysis layer consists of the AI models and analytics pipelines that process the data derived from the data layer to deliver valuable insights that support informed decisions. Finally, the user layer allows to share the processed results and new AI models parameters with data consumers (citizens, researchers and organisations). 3.3.4.1.2 Components, modules and support services Privacy and Usage Control This component (see Figure 3.9) is designed to handle the reconciliation and enforcement of machineprocessable policies, preferences, and conditions governing data usage. It facilitates a finegrained description of requirements, enabling both data providers and data consumers to define and enforce data-sharing policies tailored to their specific needs. For instance, policies can specify constraints such as "use data only for a particular purpose" or impose restrictions on data retention or third-party access. The model employs a structured and dynamic approach to ensure that these policies are both interpretable by machines and enforceable in real-time. To achieve this, it integrates a privacy module comprising a series of programmable filters or scripts that act as intermediaries between data resources and their consumers. These filters are configured to enforce agreed-upon conditions rigorously, ensuring compliance with the specified usage terms. This modular design not only improves flexibility but also simplifies updates to policy logic as requirements evolve. The privacy module operates as an active orchestrator, implementing key principles of privacyby-design and by-default. For example, the filters can enforce policies such as: 1. Purpose limitation: Ensuring data is accessed only for explicitly agreed-upon purposes. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 130 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing Figure 3.10: Overview of Federated Machine Learning component 2. Data Minimization: Limiting the scope of data shared to what is strictly necessary. 3. Retention controls: Automatically managing data expiration and deletion schedules. Incorporating these mechanisms enhances trustworthiness and accountability in data-sharing activities. Furthermore, the architecture supports auditing capabilities, enabling stakeholders to verify compliance with established policies and legal standards, such as the General Data Protection Regulation (GDPR) and the AI Act (AIA). Federated Machine Learning This component (see Figure 3.10) is dedicated to the deployment of federated ML algorithms that provide strong and quantifiable privacy guarantees. By extending computation to the edge (where raw data resides), it ensures that sensitive data is never exposed or transferred to centralized servers. This approach aligns with the principles of privacy-by-design and enhances compliance with data protection regulations. The component includes an ML orchestration framework that facilitates the collaborative development and training of highly accurate ML models on distributed local data stores. By enabling this decentralized learning paradigm, the system achieves several key benefits: 1. Data privacy: Local data remains securely stored on user devices or edge servers, minimizing the risk of breaches. 2. Data Minimization: Only model updates (e.g., parameters) are shared, significantly lowering bandwidth requirements. 3. Retention controls: Models can be tailored to local datasets, improving relevance and effectiveness while aggregating insights at a global level. To ensure robust privacy guarantees, the component integrates advanced privacy-preserving techniques such as: 1. Differential privacy: Adding carefully calibrated noise to model updates, ensuring that individual data points cannot be inferred from aggregated outputs. 2. Homomorphic Encryption: Enabling computations on encrypted data, ensuring that neither raw data nor intermediate results are exposed during training. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 131 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing Figure 3.11: Overview of Blockchain Transaction Tracking component The orchestration system also includes model evaluation and validation tools, ensuring that the collaboratively trained models meet high accuracy and fairness standards while addressing potential biases in the distributed datasets. Adaptive aggregation algorithms, such as weighted averaging, are employed to reconcile diverse data distributions across edge devices, improving model robustness and scalability. Blockchain-based Transaction Tracking This component (see Figure 3.11) leverages blockchain technology to log all data activities, including acquisition, sharing, access, and other interactions among the involved entities. By adopting a tamper-proof, decentralized ledger, the system ensures secure and immutable records of these activities, providing traceability, auditability, and non-repudiation. This transparency reinforces accountability while fostering trust among stakeholders in data-driven ecosystems. Moreover, the blockchain-basedlogging mechanismoperatesacrossmultiple layers -Data Level, Analysis Level, and User Level - creating a holistic view of all transactions. For instance, actions such as data submission, sharing, and access are securely recorded alongside metadata, including timestamps, user credentials, and purpose descriptions. These records are structured to adhere to confidentiality and data access permissions, ensuring that sensitive information remains protected even within a transparent framework. Security, Privacy and Access Management This component ensures the secure and efficient execution of data processing across all system layers, from edge devices to cloud-based data centres, while minimising the movement of data from its origin. It builds upon smart contracts deployed on private blockchains, enabling a fine-grained access control mechanism for datasets and data streams with a minimal overhead in the number of data transactions processed. This is particularly advantageous for environments with high transaction volumes, where scalability and efficiency are critical. To ensure secure data processing, data is temporarily brought into a secure execution environment through a distributed proxy mechanism. These execution environments – such as Trusted Execution Environments (TEEs) or confidential computing frameworks – isolate data and computation from unauthorized access, ensuring both data confidenThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 132 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing tiality and integrity during processing. Once the data processing tasks are completed, the proxy revokes access, ensuring that data remains confined to the provider’s secure repository and is not retained elsewhere. In addition, data provenance is maintained throughout the entire data lifecycle, documenting every stage of data management, from acquisition and manipulation to delivery. By leveraging blockchain’s immutability and transparency, the provenance records are securely logged, enabling end-to-end traceability of data manipulations. These records support auditing, compliance verification, and dispute resolution, providing a robust mechanism for accountability and trust. 3.3.5 Research Analysis Promoting privacy-by-default in DDG development is crucial for long-term success and widespread market adoption. To address evolving privacy challenges, the integration of advanced decentralized learning methodologies, such as federated learning, presents a robust approach. Federated learning enables collaborative model training across decentralized devices without sharing raw data, thereby minimizing privacy risks and ensuring compliance with data minimization principles. Combining FL with blockchain technologies enhances data integrity and security through transparent, immutable ledgers that manage data provenance and access control. Blockchain-based solutions can also facilitate decentralized identity management systems and smart contracts, providing automated governance and ensuring that data access aligns with user consent and regulatory mandates. Critically, these solutions must align with established legal and regulatory frameworks, including the General Data Protection Regulation (GDPR), AI Act (AIA), and Data Governance Act (DGA). Compliance with these standards not only ensures legal adherence but also builds public trust by addressing key concerns about accountability, transparency, and fairness in AI-enabled solutions. For instance, integrating explainability mechanisms and clear user consent protocols can further enhance transparency and trustworthiness. 3.3.5.1 Artificial Intelligence Integration For good or worse, AI has become a key player in the digital economy, driving innovation and shaping industries. Its capability to automate many data processing activities and extract valuable insights from vast amounts of raw data offers unprecedented opportunities for informed decision-making. However, this power comes with significant risks, as AI systems can be leveraged to subtly influence or manipulate user decisions. These risks are exacerbated by the presence of bias in AI models, particularly when users lack transparency regarding the underlying data sources and training methodologies. Such biases can perpetuate inequalities, distort perceptions, and unfairly sway public opinion. Addressing these challenges requires the integration of privacy-preserving and trust-enabling technologies. For instance, blockchain technology provides a decentralized, tamper-proof mechanism for recording data provenance, ensuring transparency in the data sources and algorithms used in AI systems. This enables users to verify the origin and integrity of the data, fostering greater accountability. Federated learning on the other side, enhances both privacy and fairness by enabling AI models to be trained collaboratively across decentralized datasets. This method ensures that sensitive data remains on local The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 133 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing devices, reducing the risk of data breaches and centralization-related biases. By maintaining data privacy while aggregating insights, federated learning can help mitigate the risks of overfitting and promote equitable model performance across diverse user groups. When combined, blockchain and federated learning offer a powerful synergy to ensure transparency, fairness, and reliability in AI outputs. These technologies empower users with greater control over their data and enable stakeholders to audit and verify the fairness of AI systems. 3.3.5.2 Data Usage Policy Enforcement By enforcing data usage restrictions, the DDG model can ensure the privacy and confidentiality of all involved entities, including users and organisation. Such safeguards are essential for building trust and fostering a sustainable digital ecosystem. However, to maximize the model’s acceptability and usability, additional measures must be taken to engage and inform data consumers effectively. First, communication about data usage policies must be tailored to the target audience. The language used should balance accessibility with precision, ensuring clarity without oversimplifying critical details. It is essential to align the terminology with the norms and standards of the relevant industry and technology, fostering familiarity and reducing misunderstandings. To enhance comprehension, the inclusion of real-world examples is recommended. These examples can illustrate complex concepts, showcasing their applicability in practical scenarios. Additionally, clear articulation of the expected benefits of the technology, along with specification of any potential risks, is crucial. For AI systems and applications, adhering to the risk-based approach outlined in the AI Act is imperative. This framework involves categorizing AI applications based on their risk levels and tailoring transparency and mitigation measures accordingly. For individuals with lower literacy levels or less familiarity with technical jargon, innovative approaches should be employed. These may include interactive tutorials, multimedia explanations, or knowledge assessments to ensure users have genuinely read and understood the policies rather than agreeing blindly. These assessments could take the form of brief, scenario-based questions that confirm comprehension while fostering informed consent. Finally, such efforts not only ensure compliance with ethical AI principles and legal requirements but also contribute to building user confidence and promoting widespread adoption. A thoughtful and inclusive approach to communication ensures that policies are not just regulatory obligations but tools for empowering users to make informed decisions in a privacy-centric digital ecosystem. 3.3.5.3 Regulatory and Ethical Framework The DDG is designed following the principles and guidelines established by key regulatory and ethical standards, including the GDPR, the DGA, and the AIA. Additionally, it adheres to the Ethics Guidelines for Trustworthy AI, as developed by the High-Level Expert Group on AI (HLEG). These frameworks collectively emphasize the importance of transparency, accountability, fairness, and respect for fundamental rights in data processing and AI-driven activities. In line with the GDPR and other legal mandates, the model emphasizes safeguarding individual rights, including the rights to data access, correction, erasure, and portability. To achieve this, DDG impleThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 134 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing ments mechanisms for meaningful user redress for individuals adversely impacted by AI technologies. These mechanisms include proactive measures such as risk assessment frameworks, real-time impact monitoring, and the provision of clear, accessible channels for users to challenge decisions or seek remedies. By integrating the ethical guidelines for trustworthy AI, the platform seeks to balance innovation with societal values. This entails respecting principles such as human-centricity, robustness, and sustainability, while ensuring compliance with evolving legal standards. Additionally, DDG incorporates tools to evaluate the societal and ethical impact of its AI systems, ensuring they align with the broader goals of fairness and inclusivity in digital services. Through this comprehensive approach, DDG not only strengthens its commitment to regulatory compliance but also fosters a privacy-centric and trustworthy digital ecosystem, contributing to long-term user confidence and societal trust in AI-enabled technologies. 3.3.6 Conclusions The decentralised data governance represents a critical step towards responsible and ethical utilization of data in an increasingly interconnected digital landscape. The proposed approach builds upon an extensive review of existing reference architectures, design styles, and patterns, coupled with insights drawn from lessons learned and best practices from past and ongoing initiatives. This synthesis of knowledge, informed by widely accepted models, has shaped a refined framework designed to address the multifaceted challenges of modern data governance. Central to the proposed framework is the emphasis on clear principles and policies supported by state-of-the-art privacy-preserving technologies. Techniques such as blockchain and federated learning play a pivotal role in safeguarding data integrity, enabling secure and transparent data usage while minimizing privacy risks. Blockchain ensures tamper-proof data records and facilitates traceable and accountable data transactions, whereas federated learning enables collaborative AI model training on decentralized datasets without exposing sensitive raw data. Together, these technologies empower individuals and organizations with granular control over how personal and organizational data is accessed, processed, and shared. The approach is deeply grounded in compliance with key legal and normative frameworks. For instance, adherence to the General Data Protection Regulation (GDPR) ensures that data governance practices respect fundamental rights such as data privacy and user consent. Similarly, alignment with the Artificial Intelligence Act (AIA) enforces a risk-based approach to AI systems, promoting accountability, transparency, and fairness throughout the data lifecycle. These legal foundations are further bolstered by principles outlined in the Ethics Guidelines for Trustworthy AI, fostering the development of systems that are not only legally compliant but also ethically sound. By putting control over data back into the hands of users, the proposed decentralized governance model offers a pathway to balancing innovation with privacy, security, and societal trust. This alignment of technological advancements with robust legal and ethical considerations ensures that data processing is conducted in a manner that respects individual autonomy, organizational accountability, and societal values. Such an approach is indispensable for creating a sustainable, inclusive, and trustworthy digital ecosystem. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 135 3. Privacy and Security Implementations 3.3. Data governance in distributed IoT systems and edge computing 3.3.7 Future work Looking ahead, the prospects that data governance can truly revolutionize the data economy is great. By facilitating secure and transparent data processing, decentralised data governance can drive innovations in different sectors, including healthcare, smart city and mobility, smart manufacturing, and finance. Its potential is huge, although some challenges still remain – respecting the ever-delicate balance between transparency and privacy and making sure that data processing does not jeopardise individual rights. The framework must continuously evolve in view of such problems, adapting to new types and sources of data, emerging technologies, or regulatory changes. By raising awareness of data value and governance principles, the benefits of the data governance can be realized while safeguarding privacy and data protection. This decentralised approach will help build trust and support for data-driven innovations, ultimately fostering a more responsible and secure data ecosystem. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 136 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management 3.4 User empowerment in health through Personal Health Information Management Systems Author: Christos Magkos, University of Piraeus Research Centre 3.4.1 Executive summary This research examines Personal Health Information Management Systems (PHIMS) as transformative tools in healthcare data management, addressing three critical challenges: individual data control, insight generation, and privacy preservation. Through comprehensive analysis spanning healthcare, ethics, law, and data science, the study establishes PHIMS as essential infrastructure for modern healthcare delivery. The investigation reveals significant potential in two key technological domains: risk stratification for predictive healthcare and Large Language Models (LLMs) for data processing and diagnostic support. The study identifies and addresses fundamental challenges including data quality issues (missing data, selection bias, non-stationarity), LLM limitations (explainability, bias, privacy), and regulatory compliance requirements (GDPR, AI Act, MDR). Key findings demonstrate that successful PHIMS implementation requires: robust privacy-bydesign approaches, standardizeddata integration protocols, and patient-centric interface development. The research emphasizes interoperability and security as foundational requirements, while highlighting the necessity of inclusive design to prevent demographic marginalization. Future development priorities include: technical specification refinement for recommendation systems, practical implementation frameworks for individualized health management, and strategies for achieving population-level impact through enhanced data management and patient engagement. This work provides a comprehensive framework for advancing personal health information management while maintaining essential balance between technological innovation and ethical considerations. 3.4.2 Introduction 3.4.2.1 Background The exponential growth in personal health data collection has created an urgent need for effective management systems that empower individuals while ensuring data security and privacy. Personal Health Information Management Systems (PHIMS) emerge as a critical solution to address the challenges of data control, accessibility, and insight generation in healthcare settings. 3.4.3 Background This research addresses the methodologies and limitations of Personal Health Information Management Systems (PHIMS) within the comprehensive framework of the LeADS project. PHIMS The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 137 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management reliable and justified results [DH08]. Considering the practical applications of PHIMS, a patient will need to rely upon a constantly updated and longitudinally evolving model, hence a metalearning approach could be used, where multiple algorithms for missingness are evaluated and used according to both the types of data sources and the assumptions of missingness. Selection bias While it is relatively simpler to define samples in deterministic studies, with predefined variables and specific outcomes, stochastic modeling of data that was not collected with specific intended research outcomes may present more complicated selection biases. We can identify two main sources of biases, at healthcare provider level and at research level. Healthcare providers can differ immensely in documentation practices and registering of data to EHRs and usually present with geographic intimacies; patients in two different areas will present different epidemiological characteristics [San+12]. Researchers performing analyses and creating models of disease can extract, process, and analyze data differently and interpret it in variable fashion, leading to diverse research outcomes. The demographic of healthcare providers is biased due to the very nature of healthcare services. Higher EHR completeness, for instance, may be associated with higher incidence of disease, as individuals who have more extensive EHRs have sought medical care more often [Con+24]. Certain conditions are associated with age, gender, ethnicity leading to confounding in extrapolations to the general population. Therefore, overrepresentation of certain patients may create strong statistical biases that are not representative of the total population, which is a likely target of PHIMS. Furthermore, the aging population of modern societies has shifted healthcare to address chronic disorders more often leading to an overrepresentation of older adults. An important factor to consider is willingness to share EHR data which also complicates the demographic [Jos+22]. Finally, geographic location can create sampling biases through differences in genetic composition or through local social determinants of health (SDOH), even at a neighborhood level [BEW11]. Depending on the area, socioeconomic differences can affect access, quality, and outcomes of health. Right censored vs. left censored Data censorship refers to unknown values or event times in the dataset. Data censorship can affect inclusion and exclusion criteria in the analysis which can further be complicated when patients switch health insurers due to potential interoperability and data sharing consent issues; this process is difficult to automate [Raz+15], hence having a central repository where personal health data is stored in the form of PHIMS and having access control regarding third parties in an interoperable format would greatly assist secondary data usage. Left censored data is below detectable threshold while right censored data is above the detection threshold but the exact value is unknown. Right censored data results in not knowing the exact label and is often excluded from modeling, while left censored may not suffice to derive features. In the latter case either further data collection is required, or imputation and synthetic data based on prior knowledge can be collected. Particular data limitations when using EMR records [PNMS13b] may be incomplete observations. Left censored patients may be patients that transferred to the study or provided their The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 144 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management data too late to provide complete profiles while right censored patients may have withdrawn or been discharged too early. When data is collected at discrete time points leading to unknown values or times of events, we observe interval censorship. This process can be addressed using PHIMS to an extent as additional data from IoT devices and log in could provide prior information and a central repository for data is conducive to complete longitudinal records. When it occurs, interval censoring may be addressed according to literature using methods such as expectation maximization algorithms and imputation-based approaches [Hsu+07;ZDS22]. PHIMS collected data can assist Bayesian methods of analysis through providing additional contextual information and priors or parameters to assist the analysis. Non stationarity In the process of constant medical data accumulation changes in technology, science, and incentives can change the practice of medicine as well as the clinical use of EHRs themselves, meaning the dataset itself can change. This process is called non-stationarity. When comparing a deep learning model to the status quo, slight improvements observed in the novel model compared to standard RFs may disappear with dataset shift or non-stationarity leading to overfitting the initial dataset [JS15]. This issue can occur for a variety of reasons: Time scale changes: Treatment, definition of the disease, and disease progression under treatment are constantly changing, affecting the predictive capabilities of the predictive model. Multivariate predictive nonlinear interactions: Unidentified interactions among variables can not only cause confounding. Subtle shifts among multiple variables that are undetected can lead to a cumulative sum change rendering the model less effective at prediction. Historical data overfitting: Certain models such as random forests can overfit specific data in the training set and be maladaptive to dataset shifts. So why is PHIMS a good solution to non-stationarity and what adaptations should be made when dealing with non-stationary data? Data The effect of non-stationarity on a predictive model is largely affected by data, including the volume, temporal coverage (discrete vs. longitudinal), and quality of data. Large and diverse datasets can improve generalizability, but they must assess potential data shifts to be effective. Longitudinal data spanning multiple time periods is crucial for identifying long-term trends and improving robustness, and high-quality data, derived from accurate data collection, storage, and validation to produce accurate features is essential to minimize the impact of noise and artifacts that can exacerbate non-stationarity. Model adaptations Preemptive design of models to deal with non-stationarity, particularly in the context of healthcare where data shifts are highly dynamic, is essential. Healthcare-derived time series can oftentimes be heteroskedastic and non-stationary, hence traditional time series assumptions can be inapplicable [Oga+10]. One example in literature concerning the case of transformers assessing longitudinal data in time series format where data distribution shifts across time, is to create two distinct modules [Liu+22a]. One module will provide stationarization, and provide integrated The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 145 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management statistics for inputs and heighten predictability. The second module provides a de-stationary attention mechanism to avoid overstationarization and recover the intrinsic non-stationary information into temporal dependencies by approximating differential attentions derived from the raw series data. Examining the case of non-stationarity due to intervention suggested by the very model is worth investigating as well, generating an intervention-tainted outcome. The ultimate goal of risk stratification is to inform and guide interventions. For example, a patient identified as low-risk might not be admitted to the hospital, but in certain cases, such interventions may have been the very reason certain patients, such as those with asthma or pneumonia, survived. Models can adjust for censoring effects of clinical interventions on patient outcomes, as seen in the TREWScore study (Targeted real-time early warning score) [Hen+15a] for septic shock, showing higher AUC in ROC when compared to MEWS (Modified Early Warning Score), commonly used in the clinic and Routine Screening. In this study, patients who did not develop septic shock after receiving treatment were considered right-censored after treatment; however, they were accounted for in the model development process. An inefficient updating of interventions can generate biases, as an additional factor is introduced in the causal pathway, hence causing predictive variables to be used in inferring their own effect [Lil+20]. An important consideration to mitigate nonstationarity hence arises, which is causal inference. The medical community has derived fairly accurate features for identification that may be in the causal pathway and can hence be efficient predictors of disease [Pea00]. Reducing the number of features and including primarily causal factors derived by expert opinion and literature such as the Cochrane guidelines [HG12] that directly infer the cause of disease rather than infer a correlation-derived probability risk can reduce uncertainty produced by nonstationarity. A Bayesian approach, for instance, using priors derived from within a causal framework could allow a research-informed adaptation to nonstationary new observations introduced [ZZT24]. This reduction of parameters can function two-fold in the context of PHIMS. Fewer variables improves interpretability and renders the analysis more explainable for both physicians and patients. Additionally, it is in accordance with the principle of data minimization [Sha+21a], which according to the EDPS (European Data Protection Supervisor) suggests “a data controller should limit the collection of personal information to what is directly relevant and necessary to accomplish a specified purpose”. In that case, the legal framework aids data analysts to shift towards more clinically actionable modeling of disease. 3.4.8.1 A framework for EHR based risk stratification analyses in the context of personal health Data selection to avoid biases Integration of External Data Sources: Linking EHR data with external datasets, present in the context of PHIMS such as population surveys or registries, can provide a more comprehensive view and help correct for selection biases [KAZ24]. External data can be provided through the PHIMS platform. Standardization of data Collection: Implementing uniform data entry protocols across healthThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 146 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management care providers and institutions can reduce variability and improve reliability [AS+24]. Use of Bias Adjustment Methods: Employ statistical techniques to adjust for known biases, enhancing the validity of prevalence estimates derived from EHR data [Con+24]. Generalization to population: 1. Raw frequency and exploratory data analysis 2. Data Censorship (a) Right Censored: Event occurs beyond the study period. (b) Left Censored: Values below the detectable threshold. (c) Interval Censored: Events occur within unobserved intervals. 3. Selection Bias evaluation (a) Sources: •Healthcare provider variability (documentation practices, geographic factors). •Researcher interpretation and model-building disparities. (b) Mitigation Strategies: •Consider geographic and socioeconomic disparities. •Address overrepresentation of specific demographics (e.g., older adults, frequent healthcare users). 4. Missingness evaluation (a) Classification: •Missing Completely at Random (MCAR): Data missing unrelated to variables. •Missing at Random (MAR): Missingness related to observed variables. •Missing Not at Random (MNAR): Missingness linked to unobserved factors, indicating underlying conditions related to observed variables. (b) Imputation Methods: •Multiple Imputation: Generates multiple datasets with varied imputations to improve accuracy and mitigate biases. •Sensitivity Analyses: Making use of external data, such as genetic risk scores, surveys to improve imputations and define uncertainty. (c) Data Provenance Analysis: Understanding the origin of data in order to confirm assumptions about missingness, supports granular classification, and reduces bias. 5. Post-stratification 6. Multiple Imputation and when possible modularisation using sub mechanisms for each variable of interest that presents with missingness patterns. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 147 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management 1. Causal factors and expert analysis to confirm feature selection 2. Adjust for non stationarity in the case of time series data (a) Causes: i. Temporal shifts in clinical practices and disease definitions. ii. Nonlinear interactions among variables. iii. Overfitting to historical data. (b) Model Adaptations: i. Incorporate stationarization techniques and dynamic modules for longitudinal data. ii. Adjust for intervention-induced shifts in outcomes. iii. Employ Bayesian approaches with causal priors to accommodate new data distributions. 3. Ethical and Legal Considerations (a) Data Minimization: Adhere to principles of collecting only relevant and necessary data for specific purposes. (b) Explainability: Simplify models to ensure transparency for both patients and clinicians. (c) Compliance: Align with data protection regulations such as the GDPR to ensure ethical data use. 3.4.9 Large Language Models LLMs, or Large Language Models refer to models that attempt to comprehend and generate speech or text using large text datasets [Zö+23]. LLMs use machine learning methods to generate the probability of the next speech fragment in a sequence effectively predicting the next token and simulating speech to an extent. Current LLMs such as GPT-4 (Generative Pre-Trained Transformer) can generate speech based on immense amounts of data and can sometimes be indistinguishable from human written content. In the context of PHIMS we can consider several scenarios where such models could prove useful. 1. Unstructured data summation 2. Chatbots for answering medical questions 3. Usage as a diagnostic tool The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 148 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management 3.4.9.1 LLMs for data extraction As LLMs are designed to analyze and produce text, they can be particularly useful in data processing. One example is named entity recognition (NER). LLMs can identify medical data and classify it, producing a structured format containing drug names, disorders and performed procedures from unstructured doctor’s notes or even from registered conversations [Gha+24]. The unstructured data can help with automation of registering data from multiple sources, as well as converting qualitative inputs into quantitative and categorical variables. Additionally it can be particularly useful in labeling large datasets for analysis as research suggests LLMs can be extremely efficient in the annotation of clinical data, with high accuracy and much lower human intervention. Furthermore, training LLMs in labeling using human expert knowledge significantly improves accuracy in the automated task [Goe+23]. When retrieving data from EHRs, having an additional LLM to evaluate the output greatly improved performance [Goe+23] with limited hallucinations and clinically acceptable accuracy. It is evident that LLMs can be particularly useful in the case of PHIMS especially when wishing to produce labeled structured data from patient inputs, diverse health IoT devices, physician notes and EHRs among others. 3.4.9.2 LLM performance in differential diagnosis (Ddx) LLMs offer the advantage that they can more readily analyze unstructured data such as patient histories, laboratory measurements etc. Google’s AMIE (Articular Medical Intelligence Explorer), whenconfronted withchallenging medicalcasesby theNewEngland Journalof Medicine, diagnosed patients correctly within its top 10 choices 59% of the time [Tu+24], outpacing doctors that took part in the study (34% accuracy). As standalone models, Chat GPT-4 and Claude produced the most accurate results, while a mixed method of doctors consulting with an ensemble of LLMs produced the most accurate diagnoses [Zö+24]. In a paper that assessed the accuracy of LLama 2, GPT-3.5, GPT-4 and Google GPT-4 displayed the best performance, although all LLMs were diagnostically accurate for common diseases. For rare disorder case files LLMs though displayed underperformance [San+24]. The LLMs could provide diagnosis support as well as suggest treatment and were suggested to be valuable clinical assistance tools. As these models have been trained on enormous amounts of data derived from both academic and layman resources, older information, derived from the internet, where rare diseases can be underrepresented in the most common bodies of text. Fine-tuning and pre-training can improve reliability to a great extent but can prove computationally intensive and can create massive costs in hardware [WZ24]. Equally important is correct prompting of the models, which can significantly improve performance in Ddx tasks [Whi+23]. Surprisingly, Chat GPT-4, despite not being fine tuned with medical data or given particular prompts, overperformed models specialized in medicine and trained on PubMed sourced data such as MedPalm [Nor+23]. Even more surprisingly GPT-4 passed all three stages of the USMLE (United States Medical Licensing Exam), indicating the capabilities of the model even without specialization. However, due to lacking accountability and misdiagnosing rare diseases along with hallucinaThe Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 149 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management tions, recommendations by LLMs in the context of PHIMS should be provided to the patient’s overseeing physician for evaluation. While a useful tool, LLMs should be used in compliance with a responsible trained professional. 3.4.9.3 LLM fairness and biases As mentioned above, commercially available LLMs such as GPT-4 have been trained on massive amounts of data, which means that biases derived by human predispositions are not derived from scientific inputs. It is therefore essential to evaluate whether human biases seep into analyses performed by LLMs for medical purposes. A particular case of interest is race based medicine. LLMs have been assessed in a Nature paper to promote in certain cases inaccurate, harmful, race-based medicine [Omi+23]. Since October 2023, when the paper was published, immense improvements have been made in LLMs and OpenAI, the founding company of GPT-4 have addressed to an extent certain fairness mishaps that were affecting outputs, suggesting GPT-4 presents low levels of bias [Elo+24] 3.4.9.4 Chat-GPT4 performance and biases assessment In order to assess the diagnostic capabilities of LLMs we performed a dry lab to investigate: 1. The performance of GPT-4 compared to Isabel (www.isabelhealthcare.com), a diagnostic tool, introduced into certain practices of the United Kingdom NHS (National Health System) 2. Potential racial biases introduced in the ddx process by Chat GPT Materials and Methods The dataset used was sourced by appendix A of the paper “DxGenerator: An Improved Differential Diagnosis Generator for Primary Care Based on MetaMap and Semantic Reasoning” [San+22], which also assesses a Ddx tool performance and compares it to the diagnostic system Isabel. The vignettes provided were deemed a ground truth as they were derived from credible sources of medical knowledge such as high impact factor journal case studies and medical textbooks. The symptoms and lab results for each case were fed into either Isabel or Chat GPT-3.5 and Chat GPT-4. GPT-4 outperformed in diagnostic accuracy GPT-3.5 which in turn outperformed in diagnostic accuracy Isabel for both top 1, 5 and 10 percent accuracy. We then attempted to assess the potential racial biases affecting diagnostic accuracy for GPT-4. Diagnoses that were unaffected by confounding racial factors were maintained and those associated with race were removed from the dataset. The data were then assigned a race and the diagnoses were reassessed for 4 different races: white, black, asian and hispanic. GPT-4 did not display statistical significant differences for the four races in terms of diagnostic accuracy. This unbiased performance represents a crucial advancement in AI-driven diagnostic support systems. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 150 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management 3.4.9.5 Explainability in LLMs for medical reasoning A paper on the journal of hepatology suggested LLMs offer limited interpretability and explainability of the generation process and a major limitation is that commercially available LLMs do not offer metrics of statistical uncertainty [VC24]. The black box mechanism behind these LLMs means that clinicians are unable to follow the reasoning behind the result; ultimately affecting trust in the system and hence reducing chances of implementation. Sequential prompt engineering, prompting step-by-step can steer LLMs into a specific direction and simulate a reasoning process, effectively improving accuracy in arithmetic, common sense, and symbolic reasoning tasks [Wei+22]. Prompting Chat-GPT into displaying chain-of-thought diagnostic reasoning offers interpretability for clinicians, as they can assess the “qualitative relationships between inputs and outputs” [Sav+24]. An interesting outlook was that diagnostic reasoning does not improve GPT-4’s accuracy, signifying it either does not derive the same benefits from reasoning processes that a human would or that a maximal accuracy that could not be surpassed by reasoning may have already been reached. Much more interestingly, the LLM could be generating the reasoning post hoc. In our opinion, while this would offer interpretability and ease of evaluation of the response, it is neither explainable nor necessarily accurate. While the authors of the study suggest ad hoc reasoning is unlike human reasoning, some neuroscience literature suggests humans may be reasoning in a similar fashion; the process of rationalization suggests providing seemingly logical reasons to justify behavior driven by unconscious impulses [JBL11], while the bounded reality theory suggests humans make decisions based on limited information to reach a good enough rather than optimal outcome, justifying this outcome ad hoc to suit their overall worldview and social context [GP24]. 3.4.9.6 LLMs for PHIMS LLMs offer a valuable tool in the field of personal health, especially for data extraction and summation, as well as an assistive tool for competent physicians. LLMs can function as a bridge for interoperability between providers, platforms, researchers requesting dataand doctors [Li+23b], being able to convert unstructured data into interoperable formats like FHIR. Second, inpersonalizedhealthcareapplications, LLMsdemonstrate significant potential through their fine-tuning capabilities using patient-specific data. This includes processing biometric and psychometric qualitative data, with demonstrated ability to detect signs of mental disorders and distress in speech patterns. This capability enables proactive healthcare intervention by alerting responsible physicians when concerning patterns are detected [Yan+24]. Third, regarding patient engagement, LLMs can facilitate direct user interaction in layman’s terms and provide customized patient education and health-related guidance. This includes the ability to respond to specific queries about personal health analytics, making complex health information more accessible and understandable to users [Fan+24]. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 151 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management 3.4.10 Personal data protection and PHIMS The protection of personal data in PHIMS presents multiple challenges across three critical dimensions: data handling risks, technical solutions, and regulatory compliance. Big Data and constant training means a higher risk to personal data protection. In particular, LLMs such as Chat GPT can store and process information such as conversation history and account information (https://openai.com/policies/eu-privacy-policy/), and hence sensitive information, such as medical data requires special handling and local models potentially in order to avoid data leaks in the future [VC24]. This challenge is amplified in the healthcare context, where ML models require extensive personal health data, yet sufficient safeguards are oftentimes overlooked due to limited awareness in the scientific community. 3.4.10.1 Technical Solutions and Their Limitations Several technical approaches have emerged to address these challenges. Decentralized learning techniques, particularly federated learning could provide a meaningful solution, and federated prompt and fine tuning for LLMs is currently being developed to address privacy leaks Similarly, blockchain technology where data is stored in distributed ledgers that are cryptographically safe, authenticated, and immutable can provide a secure, private and scalable environment for machine learning in healthcare. However, blockchain’s immutability poses challenges in complying with GDPR’s right to erasure and applied research solutions are required to mitigate these conflicts . Compliance with EU regulation such as the GDPR and the AI Act can mitigate to a certain extent damages to individuals. However, due to the massive rate of development of novel AI models, a proactive, privacy-by-design approach is essential, as misuse and potential privacy violations of the future may not be well defined by the regulations of today. The GDPR in particular presents ambiguous terminology in multiple articles and particularly with mobile health applications, which present multiple similarities with PHIMS, one can observe many violations in existing applications. Such are: incompleteness of privacy policy, inconsistency of data collection disclosure, and the insecurity of data transmission 3.4.10.2 Implementation Challenges In the case of machine learning, particularly LLMs, the black box nature of the models used in healthcare [VC24] is in direct contradiction with the transparency and explainability requirements of both the AI Act and GDPR. According to the Medical Device Regulation of the EU, a medical device is a tool used for “diagnosis, prevention, monitoring, prediction, prognosis, treatment or alleviation of disease,” hence qualifying any risk stratification or LLM described above as a medical device when used and applied in healthcare [VC24]. Medical devices require strict quality, security and risk control standards. Implementation in clinical workflows requires: •Local or decentralized learning approaches •Enhanced explainability mechanisms The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 152 3. Privacy and Security Implementations 3.4. User Empowerment in Information Management •Robust data management for transparency and quality •Compliance with GDPR, AI Act, and MDR requirements •Alignment with Digital Markets Act for data portability and sharing For future policy development, interdisciplinary teams, consisting of both scientists involved in state-of-the-art research with deep comprehension of healthcare analytics and legal/policy experts need to collaborate, to account for the rapid development of the field. 3.4.11 PHIMS and user empowerment Active patient participation into their own healthcare leads to better health outcomes and increased compliance and satisfaction PHIMS and personal health involvement can lead to a better understanding of personal health by the patients themselves and can ultimately inform policy and improve population health through patient education . Patient engagement can lead to change within healthcare organizations through mutual learning and more decentralized power systems. Patient engagement can have an immense impact on patient education, healthcare tool development, health organization planning, and novel policy development and implementation while enhancing existing service provisions and healthcare organization governance by affecting the care process or structural outcomes. Patient engagement requires further research as a domain considering the potential impact it can have on public health outcomes; however, it is under-defined and under-researched. Other than active participation in healthcare and a personalized approach to medicine, PHIMS can be used as a system for data sharing with researchers and a platform for personal data valorisation. The lack of control over personal data can cause users to withdraw from digital markets and assessment of personal data value is very difficult, as there is little involvement of users in the collection and processing of their data . PIMS can function as a personalized repository where personal data accessibility is managed and treated as an asset; personal data is of extreme value and users can scarcely decide to what extent they are involved in the data economy or receive sufficient compensation for their data. Personal data valorisation is underdefined and research has suggested that the main factor influencing individuals’ willingness to pay/be paid for data is the development of consciousness regarding data as a tradeable asset . The implementation of PHIMS must prioritize aspects that directly enhance patient engagement and improve healthcare outcomes. Active patient participation in their own healthcare leads to better health outcomes, increased compliance, and higher satisfaction levels. Through PHIMS implementation, patients gain better understanding of their personal health, which can ultimately inform policy and improve population health through enhanced patient education. The implementation strategy must address data valorisation aspects, recognizing that personal health data represents a valuable asset requiring careful management. PHIMS implementation should create a personalized repository where data accessibility is managed effectively, allowing users to participate meaningfully in the data economy while maintaining control over their information. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 153 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice Speaking about the economic potential of data, it should be noted that industrial data is of great value. “The value of the EU data economy was estimated at EUR 257 billion in 2014, or 1.85% of EU GDP. This increased to EUR 272 billion in 2015, or 1.87% of EU GDP (year-on-year growth of 5.6%). The same estimate predicts that, if policy and legal framework conditions for the data economy are put in place in time, its value will increase to EUR 643 billion by 2020, representing 3.17% of the overall EU GDP” 11. Consequently, with such a high cost of industrial data, the question of income distribution becomes acute. In “ensuring a fair allocation of the profits generated by analyzing the data...a clear property rule can provide the framework for a functioning data economy” [Zec15b] Today, industrial data is protected mainly by database rights and intellectual property rights. However, what if these data are not collected in a single database and do not represent the result of creative activity? The introduction of ownership rights to raw, machine-generated industrial data would go far beyond the main intellectual property regimes currently existing in Europe in the field of data and information, copyright and database rights.[Hug17] Continuing the topic of intellectual property, for instance, “copyright does not protect individual pieces of data. This claim is based on the idea/expression dichotomy, which lives at the core of copyright theory. Namely, ideas are not protected by copyright” [Hug17]. It is impossible not to agree with this point of view, since raw data, ideas and knowledge really cannot become the object of copyright. At the same time, data protection in the form of databases is possible with the help of copyright. At the same time, databases should be organized creatively and in an original way. Databases using technologies like IoT or AI may well fall under copyright protection. However, again, the raw data remains behind the legal board. In general, the introduction of the right to raw data would go far beyond the main intellectual property regimes currently existing in Europe in the field of data and information, copyright and database rights. Accordingly, there is a need for such a method of their protection is formally justified, and therefore proponents of the approach based on data ownership offer it. However, we are in no hurry to conclude that such an approach is the only possible and correct one. Thus, data ownership was proposed with the hope of successfully solving the problems that appeared with the development of the active phase of digitization and the circulation of large amounts of information. Thus, data property was supposed to firstly, simplify and fragment these very volumes of data, making it easier to manage them. The ability to organize and fragment data is needed for a reason, it was and it is an crucial task. Information flows are huge and incredibly confusing, they cannot be regulated by any single mechanism. Many problems that arise with data, such as data leakage, misuse of data, illegal disclosure of data, any unreasonable disposal of data, are difficult to control and find out all the relationships, reaching the true responsible party. It was assumed that they could be ordered using propertization. This is the so-called “erga omnes” effect, which entails that property rights apply to all persons, creating negative obligations for them without their consent. [Erp19] A certain ability to track relationships, the direction of data, protect data from misuse using the privileges of ownership rights is at the heart of this approach. The erga omnes effect could indeed, at least theoretically, ensure data security and 11European Commission, Staff Working Document on the free flow of data and emerging issues of the European data economy, Brussels, 10 January 2017 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 256 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice impose responsibility for unjustified actions with ”other’s” data. When talking about data structuring and fragmentation, we are talking not only about information security, but also about data control. To a greater extent, despite all attempts at legislation, “we have no real control over our personal data. And data property was declared a measure capable of regaining this control. Thus, if property rights are structured in a certain way, even after transfer of some control also by “selling” a fraction of rights an individual would always retain essential control over his personal data, e.g. allowing and defining the goals of data processing” [Pur07] Data control is important not only in relation to personal data, any data that is uncontrolled is a huge block of unresolved issues. And the data ownership was supposed to solve these issues, giving control over the data to individuals and organizations. Propertization, in addition to the above goals, of course, pursued the task of optimizing economic processes. Data ownership as a tool for regulating economically significant data was seen as an adequate response to the widespread processes of expanding firms’ work with data, making profits from data, selling and buying a variety of types of data. As Anastasios Dosis and Wilfried Sand-Zantma note in their study, “we define property rights as the ability to control the amount of data collected and to monetize it. In this setting, we highlight a central trade-off between two forms of contractual incompleteness: the first linked to the firm’s decision to process the data and the second related to the limit to data monetization”[AD20] . In addition to monetization and profit from data, the authors also mention data control, but in the context of data control by large firms with massive amounts of information. “If the firm owns the rights, it has full control over the collection and use of the data of all consumers”.[AD20]. Moreover, an economic forum in 2011 made the resounding conclusion that ownership of personal data already exists in the EU, if not expressly enshrined in law. that actually. “As EU data protection law evidently allows for the transfer of personal data between parties, we might conclude personal data is certainly a commodity, and thus can be considered property-like”12. . It is noteworthy that opponents of the approach under consideration have the opposite opinion that data protection law is incompatible with data ownership, which will be discussed in the next part. In addition, “the more data property rights are concentrated to large levels of subjects, the greater the value produced and the higher the efficiency” [Li22] The authors are confident that it is the ownership rights to data that can increase the profits of companies, and it is necessary to involve an increasing number of subjects in this structure. The propertization of data is supposed to fairly distribute the roles and will enable each subject to achieve higher efficiency. Summing up, propertization pursued several main tasks. First, information security, the possibility of being held accountable for illegal actions with data, systematization of the information space in such a way as to be able to follow the chain of actions with data and ”find” the responsible party. Secondly, data fragmentation was aimed at facilitating control over them, simplifying their management system. Also, data ownership implied the return of the possibility of personal management of their personal data. And of course, one of the most important goals is to expand the boundaries of the data economy and monetization of information. Like other objects of civil turnover, the data were supposed to be involved in civil transactions, bringing 12World Economic Forum, “Personal Data: The Emergence of a New Asset Class”, 2011 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 257 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice income and profit to their owners. 4.5.5.4 Argumentation of opponents of data ownership Opponents of the concept of ownership in relation to data appeal with many arguments and, first of all, the immaterial nature of data. “One primary reason for legal scholars to be critical of data ownership is that there are important differences between data and paradigm cases of property”[Zec12] First, unlike tangible entities, possession of data does not imply that one is the sole possessor and exclusive user. Data can be duplicated, and several people can use it at once. Second, it is difficult if not impossible to exclude third parties. Unless one manages to keep data secret, it can be duplicated and used by others. Due to being nonrival and non-excludable, data are public goods (as the term is used in economic theory, and contrasts with private goods, club goods, and common-pool resources). Moreover, data are non-depletable: they can be used more than once without losses in quality.[HP20] Some are convinced that the very idea of data exclusivity, necessary for the functioning of property rights, is questionable. And they refer this not only to data, but also to other goods. “Alienability of any object of some public significance, e.g. food, medication, children’s products, homes with minors as residents etc. – is heavily regulated by the state. Free alienability inherent in and necessary for the market function of property may be limited since property is never absolute”.[Pur07] Undoubtedly, data is of great public importance, but it is difficult to agree with the opinion that limited expropriation can affect the system of property rights, including in relation to data. From practical side there are explanations in the scientific community why the approach to data ownership is not able to solve specific practical problems. Many scientists are convinced that there is no practical need to own data, for example, there is a belief that even without data ownership, data would be produced, used, licensed, sold [HP20]. One of the issues of concern to policy makers is the issue of data availability. As far as we know, there are many ways to restrict access to data, there are doubts - will data ownership play a sharply negative role? This raises doubts “about the need to introduce a new ownership right to data due to the problem of data access. The access rights provided by the European Commission to facilitate the data market can indeed be implemented to a large extent by making changes to the existing rules. If the right of data ownership had been introduced, these access rights would have to take the form of restrictions on such a right”[FTF17] These doubts seem reasonable and justified, since it is known that one of the goals of the European Union is the availability of data. Data ownership implies very strict restrictions in this regard, which suggests that it is incompetent. The theory of great opportunities for data control using data ownership is being questioned, since such fragmentation can lead to even greater confusion. Many scientists are of the opinion that this approach could lead to the opposite effect and an even greater loss of control over the data. “We return to the unavoidability of death facto data control through both technological and legal means. When we consider these two aspects cumulatively, an additional layer The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 258 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice of property rights did not seem like a sufficiently attractive incentive which would dissipate de facto control, thereby ensuring greater circulation and use” [Gan] This point of view has been reflected in the works of many researchers. They express concerns that the creation of ownership rights, in particular, to personal data, is likely to create new regulatory problems. These reflections lead to the conclusion that “any benefits that could be obtained from the appropriation of personal data would be outweighed by the risk of potential harm” [P04] It is hard not to agree that property rights will potentially be the cause of many conflicts and litigation. Obviously, it will only increase the risk of unintentional violation of rights, thus complicating legal certainty. In the right of ownership, the key figure is the owner of the object of ownership. As mentioned above, it is not so easy to determine the original owner of the data. And in general, the number of participants involved the chain of information is often incredibly large, and the relationships between them are so intertwined that they are difficult to reflect in data protection measures. Speaking about data, it is difficult to say who each subject is in relation to specific information. And how to "assign" this or that status to subjects in relation to this information. All these difficulties are real obstacles to the data ownership approach. As mentioned above, some of the proponents of the data ownership approach appeal to the fact that this approach can make a more effective system of personal data protection. On the one hand, this sounds promising, however, won’t the powers conferred by the ownership right "overlap with the protection provided by the data protection act"? [Tho17] Since the data owner and the data subject will have the competence to mutually prohibit the use of data, this is quite a logical concern. Another issue is the dynamic nature of the development of the economic market. Data ownership does not seem to be a flexible enough approach to adapt to the fast pace and character of the economy. At least, in the continental system, property rights are already a “hardened” and well-established system that is not capable of rapid response. Reasoning more deeply about the inflexibility of the data ownership approach, such a characteristic of data as its dynamic nature comes to mind. It makes it "difficult, if not impossible, to determine a stable object of protection. The subject of law is simply too volatile" [P.B17] It is necessary to take into account the often rapid obsolescence of data, the constantly updated flow in databases, some instability in the right of the data producer. All these factors lead to the conclusion that creating a mechanism called “data ownership” is a challenging, if not impossible task. Not only economic, but also technical "trends" require an approach that can change and adapt to them, which is obviously not typical for data ownership. Thus, data ownership in the legal sense is not a suitable and working approach to data governance. Hugenholz put forward an idea against property rights based on the lack of legal certainty of the approach. He gives an example of data generated by machine processes. In his opinion, it is extremely difficult to imagine a right to data that is “sufficiently stable in terms of subject matter, scope and ownership rights”.[P.B17] As for the subject matter, if “the right extends to data generated by machine processes, what data will it protect? All the data that the machine outputs during a given period of time (for example, an hour, a minute or a second)”? [P.B17] It seems that the right of ownership The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 259 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice alone is not able to answer these questions. Among other things, the data ownership approach can become a threat to the confusion of overlapping rights with respect to data. Assuming data ownership as a working system, this "would seriously jeopardize the system of intellectual property law that currently exists in Europe. It would also contradict the fundamental freedoms enshrined in the European Convention on Human Rights and the EU Charter, distort the freedom of competition and the freedom to provide services in the EU, restrict scientific freedoms and generally undermine the prospects of big data for the European economy and society" [P.B17] It turns out that in addition to the "legal confusion", we are even talking about the violation of fundamental rights and freedoms, of course, according to the aforementioned author, but still this opinion takes place and seems not unfounded. We are of the opinion that data ownership is really capable of creating confusion and problems, for example, with intellectual property rights. In addition to intellectual property rights, enterprises and organizations also have a sufficient number of legal, factual and technical tools to protect their data. Speaking of the latter, when using technical protection measures, access to data is closed to any unauthorized persons and attempts of illegal interference and fraud that involve criminal liability. In addition, there is a database protection law that is able to effectively protect large databases. It is impossible not to mention the contract law, because in many cases it can "replace" the right of ownership, while taking into account the most detailed interests of the parties. We will return to contract law later dwell on it in more detail. These mechanisms offer a number of ways to ensure data protection, they are an established system, and despite the fact that each of the mechanisms has its weaknesses, "mixing" them with property rights is highly undesirable. Also, a legal inconsistency may arise due to a potential conflict with the ownership of the data carrier. “Since data can easily be copied, modified, and transferred to other storages, the entitlement to use and access the storage and data may diverge. A contractual agreement with a cloud operator usually does not mean that the operator receives legal ownership on the stored data assets. Instead, the user merely intends temporary retention and requests for exclusive access to the information. In this context, the agreement of the parties has to be cultivated” [Boe+18] The next potential conflict is data producer’s right. “The ‘data producer’s right’ would lead to extensive overlaps. As a consequence, the new right might give rise to multiple competing claims of ownership in the same content” [P.B17] It should take into account all existing mechanisms for the protection of copyright, related rights or the right to a database. The author gives an example that both copyright and database rights in the EU allow users to copy or extract data from databases for non-commercial research purposes. "The "right of the data producer" should include all the existing exceptions, and if we imagine the right of ownership, it will bring even greater difficulties. Thus, property rights cannot be considered as a separate independent category, without taking into account all related rights and the risk of conflicts with them. And these conflicts are inevitable, as illustrated above. Is there a need for such a complication and intersection of rights in the legal system? It seems obvious to us that there is a need to simplify it as much as possible, and apparently, the right of ownership is not able to cope with it. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 260 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice 4.5.5.5 Legal obstacles to data ownership approach To consider legal obstacles to the possession of information, first of all it is necessary to turn to the very concept of ownership. “The right of ownership is the most complete property right. According to its purpose, the right of ownership gives over a thing the full power compatible with the law. Its exclusivity means that no other person can have the same right to the same thing - property rights. The owner always excludes all others from ownership of this thing”.[V.P09] The theory of property rights focuses on absolute rights. The key question becomes how these absolute rights should be distributed in order to eliminate negative externalities, subsequently increasing efficiency. It follows from these that significant and socially useful resources must belong to someone, and this right must be guaranteed at the legislative level. The next thing we see from the definitions is that property rights must be distributed exclusively so that the owner can exclude others from the good. Otherwise, the distribution itself loses its meaning. And also an important point is the possibility of transferring of property rights. Speaking of the exclusion of third parties, it is necessary to mention the principle of “erga omnes”, almost “all systems of private law adhere to distinction between personal rights that apply between a chosen number persons and property rights that have consequences for third parties, sometimes even against the whole world. Property rights are special rights because they have effect against third parties, usually against everybody else”.[Akk15] However, the above definition suggests the impossibility of data ownership, as it faces the first obstacle on the way to this. By its very nature, a complete monopoly on information is difficult or completely impossible. Information "remains" with both the previous owner and the new one. This violates the very concept of property rights. The ability to exclude others from the field of ownership becomes a key issue in potential data propertization. The next principle to consider is the “numerus clausus” principle. It is interpreted differently in different legal systems, but the essence of the principle is that the content of the property rights is already predetermined, and the parties are limited to this already existing content and “thus are limited in the creation of new types of property rights”.[Akk15] This principle of property rights has both a substantive and procedural side [Erp], since this principle limits not only the number and content of such rights, but also how these rights are created, transferred and terminated. Following the principle of numerus clausus, next to full ownership, only property rights smaller than the property right (secondary or so-called "limited" property rights) can be created. Since we are talking about the classical understanding of property rights as an absolute right in legal doctrine, one can see from the concept discussed above that it is impossible to own, dispose and use information in the classical sense. And this is largely due to the contradictory nature of the information. It has much in common with material goods, while being an intangible object. Also, the right of ownership of information does not correspond to one of the leading principles of ownership - transparency. “The principle of transparency in a classical sense only had two aspects: the object as to which a property right was claimed had to be clearly described and delineated and it had to be made public”.[Erp] It is obvious that the principle of transparency did not imply the possession of information and is more directed towards material The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 261 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice objects. Thus, the data can be said to be inconsistent with the basic principles of property rights as such. “When looking at the characteristics of a potential data ownership right, most parameters are still very unclear. Academics have only just begun to think about the possible subject matter of protection (be it data on the syntactic level or information on the semantic level) and whether data or information should be protected as unitary asset (attributing rights irrespective of how others have obtained the subject matter of protection) or as a commodity asset (allowing others to use independently obtained data or information)” [FTF17] Speaking aboutdifferent levelsofinformation, thisissue was thoroughly studiedbyZech. [Zec16b] He distinguishes between three conceptual levels of data. The first is the semantic level or content level, which encapsulates human-understandable information that carries the essence and meaning. At this level, we can talk about the value of information as such. The second level is syntactic, where information is presented in the form of signs and symbols, also called the code layer. The semantic level is made up of the syntactic level, creating a third level, which is located on the physical storage medium. It is also called the physical level, which can be tangible by a person. This classification was a bright success in the scientific community, as it facilitated the understanding of the essence and nature of information. In addition, it is of great importance in relation to the question of this study, since “as it is crucial which layer will be the object of the property right. Granting the property right at different conceptual layers can have vastly different consequences for the law” [Ste23] Zech also formulated his definition of data ownership, which is based on the concept of property rights. “Building on the standard categories of property rights: use (usus), enjoying the benefits of the use (usus fructus), changing form and substance (abusus) and transfer of the property three basic categories of rights to information can be distinguished: possessing information, using information and destroying information”[Zec15b] This definition seems to be the closest to reality and accurate, since it reflects the capabilities of a potential data ownership system to the greatest extent. Continuing the question of the controversial legal nature of the information, the difficult question arises of how to determine their original owner? It can be said with certainty that “there is no consensus about the criteria on how to originally attribute the right to a right holder”.[FTF17] Since the data does not lose properties while simultaneously satisfying the needs of, say, hundreds of subjects, it is difficult to determine the original owner. However, it is quite certain and obvious that the familiar doctrine ceases to work. Many authors call today to “to avoid unnecessary and non-productive academic debate based on sterile dogmatic analysis and preconceived 19th century paradigms” [Erp] After all, the right of ownership is initially focused on the possession of physical objects that are fundamentally different from the object considered in this study. And when the object is something else, then there is no point in relying on the usual classical dogmas entirely. The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 262 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice 4.5.5.6 Legal analysis of GDPR provisions from the angle of the data ownership approach The General Data Protection Regulation (GDPR) is one of the most important legal framework that establishes guidelines for the collection and processing of personal data of individuals. It was approved in 2016 and entered into force in 2018. Its peculiarity lies in the fact that its power extends to all websites used by EU residents. In law, data ownership is mostly associated with the privacy of individuals. “With personal identifiable information being collected in an ever-increasing volume by large tech companies, this discipline aims at defining the actual owner of this data collection and the extent of control that remains with the data’s subjects. This legal perspective is particularly important as companies must be held accountable when it comes to data leakages or alienation...”[Leg21] The GDPR provisions are interesting to us from the point of view of the data ownership approach. What is personal data according to GDPR? According to the Article 4 personal data means “any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is a one who can be identified, directly or indirectly, in particular by reference to an identifier such a name, an identification number, location data...”13. The GDDR is, first of all, about the basic human rights and freedoms, their right to dispose of information about themselves. Is it possible to say that we "own" information about ourselves? According to Pearce, “data protection rights could prima facie be regarded as either personal or proprietary rights. Both kinds of rights share their orientation towards economic efficiency, civil liberties and avoidance of unjust interferences. But while proprietary rights do so on the basis of assigning transferable rights in external things, personality rights are internal, not directed at external things, relate to an individual’s personhood, character, and identity, and are inalienable”[H18] The important point is that the legislation does not directly address the issue of ownership, and firms and regulators circumvent many obstacles by avoiding direct references to ownership; instead, they focus on the potential control exercised by the data subject or the limitations of the data controller.[JM17] It should be noted that personal data in principle cannot be the subject of property protection. “The GAP in accordance with the European Charter of Fundamental Rights defines personal data as rights subject to special consideration, in respect of which no property rights can be exercised” [Ste23] Article 1(2) of the DR defines the purposes of the document, according to which it protects fundamental rights and freedoms and, in particular, the "right to personal data protection" 14., therefore providing a definition of the right separately from property. However, “This view is not necessarily shared in the US. Some privacy conscious academic advocate the propertization of privacy as the model for increasing privacy protection [Ste23] Different legal systems have different approaches, but we are focused on the EU approach. However, there is a directly opposite opinion that the European data protection system can be considered a regime similar to the property regime. This is explained by the legal powers vested in individuals, and the most striking arguments are the provisions of the Regulation on the right 13The General Data Protection Regulation, 2018, Article 4 14The General Data Protection Regulation, 2018, Article 1 (2) The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 263 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice to be forgotten and the right to data portability. Such control and management resembles the scope of property rights. “Significantly, some have suggested that despite the GDPR prima facie being premised on human rights rhetoric and the fact that it employs no explicit property terminology, its substantive provisions function remarkably similarly to regulated property regimes” [J] There are surprisingly a considerable number of adherents of this opinion, some of them even make such bold statements as “more specifically the GDPR appears to implicitly acknowledge that the concept of personal data itself has become akin to a commodity that is capable of changing hands and being traded” [J13] Another argument in favor of this point of view is the protection of personal rights, similar to the protection of property interests. Property rights protect the interests of an individual in relation to a particular object or resource, whether it is a tangible or intangible object. What is especially remarkable, “an interest will be protected from an unwanted use or taking by another party, under the assumption that property can only legitimately change hands if the consent of the right holder is given. The interest of rights holders in this context tend to be protected by courts and legislatures by way of fines, punitive damages or injunctions” [H18] Do we support this point of view? In part, yes, legal mechanisms do resemble the protection of property rights. And then we will try to show in more detail why. An important provision of the GDPR is the right to deletion (the “right to be forgotten”), which gives the data subject the right to receive from the controller the deletion of personal data concerning him without undue delay. The data subject also has the right to withdraw his consent to the processing of his personal data. The "right to be forgotten" is of interest from a legal point of view, which gives the data subject a special right to information about him. Thus, “the individual is granted a power of exclusive disposition concerning the processing of personal data that is—to some extent—comparable with the power of the owner over his property. In terms of property law, this could be understood as a negative dimension of an exclusive right, i.e., the power to exclude others from using one’s property”.[Boe+18] This point of view is quite justified, since it really resembles the regime of property rights and the exclusion of others from the possibility of owning an object. However, in the case of the right to be forgotten, it is not the exclusion of all others by the subject, but the deletion of the object itself. And if in the ownership right the owner continues to satisfy his ownership right over the object, then the same cannot be said here. Thus, although the "right to be forgotten" resembles the legal regime of property rights, they still cannot be correlated as identical. The followingprovisionsofthe Directive, inthe opinionof somescientists resembling theregime of property rights, is the right to data portability. The data subject must have the right to transfer his personal data “ to another controller without hindrance from the controller to which the personal data have been provided”.15. This provision gives the subject the right to dispose of his personal data almost in full, independently determining the place of their storage. This is a direct opportunity for the subject of personal data to make a decision on further processing and access, it goes far beyond consent to the processing of their data. “Therefore, it can be seen as another step towards a (privacy-based) concept of data ownership.” [Jü16] 15The General Data Protection Regulation, Art 20 (1) The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 264 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice 4.5.6 Research Analysis 4.5.6.1 Data as an object of civil law relations One of the key positions is the consideration of data as an object of civil rights. Civil law relations are most closely related to the daily needs of people and should cover at least the most important aspects of their lives related to civil law. This analysis provides an opportunity to get closer to the needs of individuals and to see gaps in the law regarding data. This is especially important, since today huge flows of economically valuable information circulate in civil turnover, and the question of its regulation arises. “In the big data era, data is an emerging object of civil rights” [Li22] It is impossible to deny the influence of data and their circulation in civil turnover as a new object of civil law, although not yet legally designated in the legislation of many countries. But in practice, data have long and firmly entered the system of civil legal relations, which is the need to consider them in more detail from this angle. “The civilian idea of ownership is an absolute dominion... over the relevant object”. This is indeed the case, and in the context of data means the most comprehensive property rights over the data. “From a comparative viewpoint, a main distinction can be drawn between the civil law and common law understanding of ownership” [Gra17] Data in the context of civil law is a different category, which needs to be delved into more deeply. Firstly, data as an object of civil rights is a very important category for the business sphere. As Teresa Scassa notes in her research “data ownership can play a role in commercializing data: It is common for companies and organizations to seek to control the data they collect through their activities in order to commercialize them. An ownership right can support various techniques for control, including contracts/licensing and technological protection measures”16. The economic value of data is growing day by day, and it is quite possible that in the future they will take a dominant place in the system of economic goods. Currently, the exchange of information and its storage and accumulation can already be called ubiquitous phenomena. However, when it comes to selling information as a commodity, it is a more unexplored and relatively new type of activity in the global economy. The data has unique characteristics as an inapplicable product. “In contrast to most other goods, data thereby are infinitely usable and are the source of increasing returns for companies” [JC19]. The non-susceptibility of data to physical wear is an important characteristic of them. At the same time, it is necessary to take into account the possible obsolescence of data, potential irrelevance or their ability to cease to be necessary as disclosure. All these are characteristics that distinguish data from other objects of civil law. It is necessary to understand the concept of the object of law as such. Indeed, while it is not difficult to determine the subject (it can be both an individual and a legal entity), some difficulties may arise with the object of law. Under the legal objects of civil law, most legal systems usually recognize tangible assets, such as land and movable property, and intangible assets, such as monetary claims, personal benefits, honor and dignity, etc. Regarding property rights “civil law 16CIGI Papers No. 187 — September 2018 Data Ownership Teresa Scassa, page 2 The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 265 4. Enablers of data processing 4.5. Boundaries of data ownership: from concept to practice nience for data chain participants. The expansion of the concepts of data user and data holder leads to a concept beyond the boundaries of data ownership. Such a concept is necessary today, when in the world of the Internet of Things there are a huge number of devices that interact with each other. Data should not only be collected and stored properly, but also transmitted and used by enterprises to improve the quality of their products and services. Moreover, the data itself becomes a commodity with its own value, and the Date Act reduces the obstacles to their transfer. Thus, without recognizing the data as property, the Date Act nevertheless opens up new possibilities comparable to ownership as such. It will not be possible to completely get away from the concept of ownership, because control, access and security are inextricably linked with it. 4.5.7 Conclusions The significance of the research findings in the context of the project objectives is crucial, as it can influence the development of data policy, especially personal data. The study focuses on the practice of working with data and empowering individuals. They are not always able to fully understand the processes that occur with their data, but they must be confident in transparency and the ability to get the data they need at any time. However, there is an opinion that “data protection is on the bottom of the list of priorities for the IT companies” [T07] Of course, this is unacceptable because only a synergy of legal and technical measures can ensure the most complete data security and protection. Data encryption, data anonymization, all these processes are closely related to computer science and legal science alone cannot overcome them. Reuniting the legal side with computer science is a powerful tool that can bring real benefits, including economic ones. When talking about privacy or intellectual property today, the introduction of technical measures and guarantees are timely responses. With respect to sensitive data, they are a direct necessity. Health data, genetic data should be processed in ways that are especially protected from external attacks. At the same time, the European data governance system should ensure the interests of other stakeholders. That is, each stakeholder should be able to control "their" data. In this way, the goal of increasing the value of the data and minimizing the associated costs can be achieved. The legal status of data ownership is difficult to determine, it is disputed by many scientists and has a different approach in different legal systems. This is due to the dynamic nature of the data, which is difficult to determine. The study provides a detailed analysis of the data as an object of civil rights, examines the positions of supporters and opponents of this approach with a legal justification. The legal obstacles to data ownership are also considered and the GDPR provisions regarding the concept of data ownership are analyzed. After conducting the study, the need for a new approach to data ownership that takes into account the interests of all parties, namely, a more flexible and modern one, is outlined. To a large extent, the research is aimed at the possibility of extracting benefits for data subjects from this concept and the possibility of its application in practice. Although ownership may not be an appropriate category for formulating and recognizing the requirements of data subjects, the proposals of owners to reform the data The Project has received funding from the European Union’s Horizon 2020 Innovative Training Networks. Grant Agreement ID: 956562 272 [Document text truncated for crawler view.]