scieee AI-readable full text Open interactive document viewer

A Green Paper on AI, Data Governance, and Metadata Policies for Europe's Music Ecosystem

Antal, Daniel

Abstract

This document (version 0.8) is an early-stage Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem: Practical Steps Towards a Decentralised and Open European Music Observatory. It is released for consultation and should not be considered a final work. It has been internally reviewed. Any comments are welcome for improvements in problem statements, omissions, recommendations. Prepared in line with Horizon Europe’s transparency rules and the Open Policy Analysis (OPA) framework, it is released early to enable consultation, incorporate stakeholder input, and ensure an auditable drafting process. All related deliverables, figures, datasets, and bibliographies are openly available via GitHub and Zenodo to support transparency and reuse. The Green Paper addresses three key reform layers: Fixing music data at the source (reducing redundancy, improving interoperability, reconciling attribution and privacy). Building a federated Open Music Observatory (as a European data-sharing space aligned with EIF, FAIR, EOSC, and ECCCH). Aligning AI with governance and value creation (supporting curative AI, shared utilities, and trustworthy frameworks that help small actors as well as large platforms). It serves as the basis for Deliverable D5.7 (Policy Brief) of the Open Music Europe consortium and will inform a subsequent White Paper to be discussed at LineCheck 2025 and the final policy forum in Brussels (December 2025). Transparency note: Always cite the latest versioned DOI available on Zenodo. Supporting documents and figures are accessible via our GitHub repository. Funding acknowledgement: This project has received funding from the European Union’s Horizon Europe programme under Grant Agreement No. 101095295. The views expressed are those of the authors only and do not necessarily reflect those of the European Commission or its agencies.

Full text

A Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem Practical Steps Towards a Decentralised and Open European Music Observatory Daniel Antal, CFA 2024-11-22 Table of contents Introduction 4 Glossary 9 Musicterms...................................... 9 Dataterms ...................................... 10 AI&SystemsTerms................................. 11 Dataprotectionterms ................................ 13 Data curation and collection terms . . . . . . . . . . . . . . . . . . . . . . . . . 13 Statisticalterms ................................... 14 Registers, authorities, standards and identifiers . . . . . . . . . . . . . . . . . . 15 Organisations..................................... 17 Otherabbreviations ................................. 18 1 Policy context and problem map 19 1.1 Three structural pressures . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 1.2 National and European pilots as anchors . . . . . . . . . . . . . . . . . . . 21 1.2.1 The Slovak Comprehensive Music Database (SKCMDb) . . . . . . 21 1.2.2 Unlabel ................................. 23 1.3 Questforefficiency............................... 23 1.4 Potentialsolutions ............................... 26 2 Fixing Music Data at the Source 28 2.1 Discussion.................................... 28 2.1.1 Structural fragmentation of data and value flows . . . . . . . . . . 28 2.1.2 Cost barriers in documentation and claims . . . . . . . . . . . . . . 31 2.1.3 Why one grand collection model will not work . . . . . . . . . . . . 31 2.1.4 Legacymetadata............................ 32 2.1.5 Named-entity resolution, attribution, and privacy . . . . . . . . . . 35 2.2 Policyproposals................................. 36 2.2.1 Reducing redundancy . . . . . . . . . . . . . . . . . . . . . . . . . 36 2.2.2 Reconciling attribution and privacy . . . . . . . . . . . . . . . . . . 37 2.2.3 Pragmatic metadata alignment . . . . . . . . . . . . . . . . . . . . 39 3 Open Music Observatory: Building a Shared Music Data Space 42 3.1 Discussion.................................... 43 3.1.1 Why centralisation is a futile model . . . . . . . . . . . . . . . . . . 43 3.1.2 Open Data Directive: right without means . . . . . . . . . . . . . . 45 3.1.3 Why voluntary workarounds do not scale . . . . . . . . . . . . . . . 46 2 3.1.4 Public infrastructures bypass music’s real data flows . . . . . . . . 47 3.1.5 Subsidiarity and infrastructures for scaling music data . . . . . . . 48 3.1.6 Economies of scale in metadata . . . . . . . . . . . . . . . . . . . . 50 3.2 PolicyProposals ................................ 50 3.2.1 Workflow playbooks and provenance trails . . . . . . . . . . . . . . 52 3.2.2 Federated infrastructure as a cost and governance solution . . . . . 53 3.2.3 Legal, standards, and funding levers . . . . . . . . . . . . . . . . . 54 3.2.4 Alignment with the European Open Science Cloud . . . . . . . . . 55 4 AI that Works for Music, Not Against It 56 4.1 Discussion.................................... 59 4.1.1 Structural problems for music businesses to apply AI . . . . . . . . 59 4.1.2 European regulation that misses the point . . . . . . . . . . . . . . 60 4.1.3 Policy issues at the intersection of AI, copyright, and GDPR . . . . 62 4.1.4 AI design without awareness of limits . . . . . . . . . . . . . . . . . 63 4.1.5 Unfreezing frozen assets . . . . . . . . . . . . . . . . . . . . . . . . 63 4.1.6 AI support for investment into new repertoire assets . . . . . . . . 65 4.2 Policy Proposals: Aligning AI with Governance and Value Creation . . . . 65 4.2.1 EU-Level Policy: Compass and Guardrails . . . . . . . . . . . . . . 66 4.2.2 Industry-Level Policy: Standards and Collaboration . . . . . . . . . 66 4.2.3 Organisational-Level Policy: Playbooks for CMOs, Publishers, Archives................................. 66 4.2.4 Curative AI and Reparative AI as a Remediation Solution . . . . . 67 4.2.5 Lowering Documentation Barriers . . . . . . . . . . . . . . . . . . . 68 4.2.6 Observatory: European = Open . . . . . . . . . . . . . . . . . . . . 68 4.2.7 The Open Music Observatory as a Collective Guardrail . . . . . . . 69 5 What Europe Should Do Next for Music Data & AI 71 Sources & Further Reading 73 3 Introduction There are musical works that are reinterpreted thousands of times across centuries. A symphony by Beethoven or a folk song from the Baltic coast can be heard again and again, each performance producing a new reading of something that never becomes “final.” The same is true of sound recordings. Some perennial recordings are rediscovered after sixty years, remastered, and brought into circulation for new audiences. Music assets, in other words, have an unusually long lifecycle. This is just as true of their documentation — the metadata that accompanies them from creation to archiving. Metadata does not freeze a work or recording in time. Instead, it evolves with it: from the moment of rights registration, through commercial distribution and playlisting, to preservation in a library or archive. Each new interpretation, remix, or reissue generates new metadata; and each new information system demands new connections and contexts. ĹWhy this Green Paper matters for music professionals? • Streaming has centralised power in platforms, but left rights-holders with microroyalties and huge admin burdens. • Metadata mistakes mean lost revenue — each unlinked ISRC or ISWC is money left on the table. • AI is already changing music — either it helps you fix documentation and get paid, or it floods the system with untracked works. • Europe needs federated, cooperative solutions so independents, CMOs, and archives can compete on fairer terms. 4 There is rarely a single moment when music metadata can be considered complete. Metadata, like music itself, is open to reinterpretation. A name can be reconciled with an identifier; a work can be linked to a new performance; a recording can be embedded in new file formats. Each act of documentation adds layers of meaning and makes the music informative in a new environment. This is not an invitation to reinvent the wheel. We can read Beethoven’s early prints as well as Iris Szeghy’s 21st-century scores because music notation — a standardised way of presenting the metadata of musical works — has remained remarkably stable for centuries. Notation shows that standardisation can endure, and that shared conventions make music legible across time, geography, and institutions. 5 The invention of the computer, and later the internet, introduced new ways to document and transmit music. These innovations brought powerful efficiencies: identifiers like the ISRC and ISWC, digital distribution pipelines, and networked catalogues have enabled the global circulation of music at unprecedented scale. But they also created new fragmentation. Standards proliferated, identifiers failed to interconnect, and workflows designed for one purpose often broke down in another. What was intended as progress sometimes left behind a mess of overlapping, incompatible, or incomplete metadata — a mess that now needs to be cleared up. ĹNote This Green Paper is an early-stage policy document, prepared in line with Open Policy Analysis and the Horizon Europe Data Management Guidelines. It has been released early to allow consultation, incorporate stakeholder input, and provide a transparent development process. This Green Paper extends the analysis developed in the first OpenMusE policy brief on music metadata mainstreaming and EU law (Deliverable D5.6), and its findings are condensed into the second policy brief (Deliverable D5.7), which incorporates wider stakeholder consultations.1 Transparency note: Following the principles of Open Policy Analysis, all related deliverables and technical documentation are publicly accessible to foster broad engagement and ensure a clear audit trail. Supporting documents for each chapter of this Green Paper are referenced in similar boxes. The current version (and future White Paper drafts) is available at https://zenodo.org/records/17075796. Standardised folders, figures, and bibliographies are available at https://github.com/dataobservatoryeu/open-music-data-white-paper. Funding acknowledgement: This project has received funding from the European Union’s Horizon Europe programme under Grant Agreement No. 101095295. The views expressed are those of the authors only and do not necessarily reflect those of the European Commission or its agencies.2 Citation note: When citing this Green Paper, please use the latest versioned DOI available on Zenodo, and include the date of access if referring to material hosted on our GitHub repository.3This is an early version (v0.84.) 6 Our document has been presented and discussed with industry specialists on the following forums: • Big Data Value Association, Gaia-X: Dataweek²�: Introducing a new European music dataspace4 • Echoes/ECCH: • Hungarian stakeholders interested in replication of the Slovak pilot versions 5 • CISAC: Protecting Creators’ Rights in the AI Era: OpenMusE at the European Committee Meeting, Vilnius, 29-30 April 6. • The Fair MusE - Prelude to a fairermusic industry Fair MusE project7 • IAMIC 8: The International Association of Music Information Centres and several key members of the organisation. • IAML: The International Association of Music Libraries, Archives and Documentation Centers and several national chapters and key members 9. 3The Policy Brief 1: Music Metadata Mainstreaming and EU Law (Senftleben et al. 2024) provides the legal and institutional framing for metadata mainstreaming in European copyright and data law. The present Green Paper builds on that foundation with a lifecycleand sovereignty-oriented conceptual framework, tested in pilots such as the Slovak Comprehensive Music Database. Its key recommendations are further condensed in OpenMusE Policy Brief 2: An Open, Scalable Data-to-Policy Pipeline for European Music Ecosystems (Deliverable D5.7, 2025) (Open Music Europe Consortium 2025), which integrates broader stakeholder consultations (CISAC, IAMIC, IAML, FairMusE, Music360, ECCCH forums, among others) and translates them into policy actions for EU institutions. 3This document has been prepared by Open Music Europe (OpenMusE) project partners as an account of work carried out within the framework of this contract. Any dissemination of results must indicate that it reflects only the author’s view and that the Commission Agency is not responsible for any use that may be made of the information it contains. Neither Project Coordinator, nor any signatory party of Open Music Europe (OpenMusE) Project Consortium Agreement, nor any person acting on behalf of any of them: (a) makes any warranty or representation whatsoever, express or implied, (i) with respect to the use of any information, apparatus, method, process, or similar item disclosed in this document, including merchantability and fitness for a particular purpose, or (ii) that such use does not infringe on or interfere with privately owned rights, including any party’s intellectual property, or (iii) that this document is suitable to any particular user’s circumstance; or (b) assumes responsibility for any damages or other liability whatsoever (including any consequential damages, even if advised of the possibility) resulting from your selection or use of this document or any information, apparatus, method, process, or similar item disclosed herein. 3Always use the latest versioned DOI when citing this Green Paper, available via Zenodo. If you rely on supporting material hosted in the GitHub repository, please add the date of access in your reference. The figures and charts can be found on FigShare and may be reused separately, citing their DOI and, for context, the Green Paper that contains them. 4Jun 5, 2024, Dataweek²�, Leuven, Belgium. 5Federation possibilities of the Slovak music data sharing space in Hungary (Antal 2024a) 6Protecting Creators’ Rights in the AI Era: OpenMusE at the European Committee Meeting, our presentation (Mikš 2025) 7We received useuful feedback for this Green Ppaer from the project and see further synergies in presenting our policy findings together. https://fairmuse.eu/about/ 8We presented and discussed these ideas at the International Association of Music Information Centres on the General Assembly and Annual Conference 2024 on November 21, 2024, at Music Austria, Vienna. See the presentation and its poster format (Antal 2024d). 9We presented and discussed these ideas at the International Association of Music Libraries, Archives and Documentation Centers on the General Assembly and Annual Conference 7th and 9th of July 2025 in Salzburg, Austria. See the presentation and its poster format (Antal 2025a, 2025b). 7 • Polifonia: In October 2023 Polifonia invited a few stakeholders - Podiumkunst.net, the Open Music Observatory, Uni Firenze, IC Fonseca School, Joséphine Simonnot/PRISM, Maria Luisa Onida/D’Istruzione Superiore Leonardo Da Vinci, Carnegie Hall Archive, Municipality of Bologna - for a work session, which gave us a great opportunity to strengthen the metadata framework of our policy recommendations and infrastructure planning. • Music Futures: the AHRC Creative Industries Cluster project MusicFutures in the United Kingdom. • Slovak national stakeholders interested in cultural data.10 • Wikimedia community and developers11. • European music industry stakeholders on LineCheck 2025 12 The CITF’s First Project Report (Ministry of Education and Culture, Finland, 2025) validates and extends the policy logic of OpenMusE’s Green Paper. While the CITF defines cross-sectoral requirements for trustworthy, machine-readable copyright infrastructures in the AI era, OpenMusE provides a concrete domain implementation within the European music ecosystem—demonstrating how interoperable identifiers, FAIR principles, and federated governance can function in practice13. 10Based on a memorandum of understanding with a broad range of public and private stakeholders, (Ministerstvo kultúry SR and Open Music Europe 2023) we developed a model for renewing statistical production for better cultural and music statistics (Antal 2023). 11Our work was presented in the Technology session of the Wikimedia CEE Meeting 2024 in Istanbul, and the Wikimedia CEE Meeting 2025 in Thessaloniki, and the Wikidata Conf 2025 online; we have built relationships with various national chapters and the Wikidata and Abstract Wikipedia teams, and joined the Wikidata Ontology Cleanup Task Force and the Wikidata Mereology Task Force to help the coordiantion of our open source technology, data curation and dissemination efforts. (Antal 2024b, 2025c; Antal, Pigozne, and Federico 2025). 12Open Access Music Dataspaces – Open Music Observatory presented on LineCheck 2025 (Mikš and Antal 2025) 13We are planning to give feedback to the (Partanen et al. 2025) on 19 November, and we asked the authors of the report to comment on our Green Paper, too. 8 Glossary Music terms audio recording: fixation of sounds (ISO 2019a) creator: in the context of this policy paper, we use the broad term for the arranger, author,composer,lyricist; for individual definitions see ISWC standard (ISO 2022) DSP or digital streaming platform: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. DSPs can provide access to music downloads, like Apple’s iTunes Store, or access to streaming music like Spotify, or even provide satellite-delivered content such as SiriusXM in the USA. expression: intellectual or artistic realisation of one and only one work Note: may take the form of a notation , sound, image, object, movement or text (ISO 2017b) manifestation: physical embodiment of an expression (ISO 2017b) movement: A principal division of a musical work. (ISO 2022) music video recording: fixation of sounds synchronized with pictures or moving pictures where (a) the fixed sounds are wholly or substantially a musical performance or (b) the recording is intended for viewing in association with a recording of a musical performance. This definition includes music videos and concert recordings, together with music-related interviews and documentaries, but does not extend to genera! audiovisual material, even if it includes music.(ISO 2019a) musical work: composed of a combination of sounds, with or without accompanying text (ISO 2022) original title: A title given to the work by its creator(s) shown in its original language. (ISO 2022) formal title: A standardized title in which the elements are arranged in a predetermined order, such as titles created for classical works. (ISO 2022) rights management (organisations): the function of managing the rights on behalf of rights owners. It can be companies whose sole purpose is to ensure that content that has been licensed has delivered royalties that are identified and accounted for. The role can be taken by collective management organisations or by private companies on behalf of songwriters, composers, performers, music publishers, or record labels. 9 within the global supply chain and to develop trade directories and similar services for the specialized market for music publications. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. DOI: The Digital Object Identifier is a standardised unique number given to many (but not all) articles, papers and books, by some publishers, to identify a particular publication. ORCID: the Open Researcher and Contributor ID is a unique, persistent identifier free of charge to researchers. URI: A Uniform Resource Identifier (URI) is a string of characters used to identify a resource on the internet. This resource can be either abstract or physical, such as a website, an email address, or a file. URIs are essential for enabling interactions with resources over a network using specific protocols. DDI: The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. Wikibase: Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, but it is available now for the creation of private, or public-private partnership knowledge graphs. It is developed by Wikimedia Deutschland. SDMX: Statistical Data and Metadata eXchange (SDMX), is an international initiative that aims at standardising and modernising (“industrialising”) the mechanisms and processes for the exchange of statistical data and metadata among international organisations and their member countries. CIDOC-CRM: The conceptual model of CIDOC, the standard conceptualisation of collection management systems in heritage organisations. RiC:Records in Context is a new conceptual model that replaces the four most important international archiving standards. DCTERMS or DCMI: the Dublin Core Metadata Terms is a vocabulary of metadata terms developed and maintained by the Dublin Core Metadata Initiative (DCMI). These terms are used to describe various aspects of digital resources, such as web pages, documents, and other online content. They provide a standardized way to assign metadata to resources, making them easier to discover, manage, and exchange. RDFS: the Resource Description Framework Schema is an extension of the Resource Description Framework (RDF) that provides a vocabulary for describing classes and properties of resources within an RDF graph. EDM: the Europeana Data Model is a framework for collecting, connecting, and enriching cultural heritage metadata. It’s designed to facilitate the sharing and reuse of cultural heritage information by providing a standardized way to represent and link data. 16 PROV-O: the Provenance ontology is a formal ontology developed by W3C to represent and interchange provenance information. MARC: MAchine-Readable Cataloging, is a standard digital format used by libraries to represent and exchange bibliographic information. DCAT: an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. Organisations AEPO-ARTIS: Organisation representing European artists-performers. Regroups most of the European CMO representing performers. ALOADED: is a company which distributes and exploits recordings. CISAC: The International Confederation of Societies of Authors and Composers is an international non-governmental, not-for-profit organisation that aims to protect the rights and promote the interests of creators worldwide. CNM (former CNV): the Centre National de la Musique is a public organisation managing a tax on concert tickets EMO: The European Music Observatory (EMO) is envisioned as a hub for collecting and analysing data on the music sector across Europe. Its primary aim is to address the current gaps and inconsistencies in music data collection, which have been a significant challenge for the sector. Europeana: a digital platform provided by the European Union that aggregates digitized cultural heritage from institutions across Europe. GESAC: The European Grouping of Societies of Authors and Composers (GESAC) comprises of 32 European authors’ societies in music, audiovisual, visual arts, literature and drama. IAML: International Association of Music Libraries, Archives and Documentation Centres IAMIC: International Association of Music Centres, an international network of organisations that collectively and collaboratively provides information and promotes the music of their countries or regions. ICMP: the global trade body representing the music publishing industry worldwide. SCAPR: International association for the development of the practical cooperation between performers’ collective management organisations (CMOs) SOZA: SOZA (Slovenský ochranný zväz autorský pre práva k hudobným dielam, Slovak Performing and Mechanical Rights Society) is a legal entity, non-profit civic association of authors and publishers of musical works, association of natural persons and legal entities. 17 Hudobné Centrum: Music Centre Slovakia is a music organisation with a mission to promote Slovak contemporaly music. Other abbreviations CEEMID: the Central European Music Industry Databases is a multi-country project that was a predecessor of Reprex’s Digital Music Observatory DSP: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. EIF: The European Interoperability Framework (EIF) is a set of recommendations and guidelines that aims to facilitate communication and collaboration between public administrations, businesses, and citizens within the European Union and across national borders. ECCCH: The European Collaborative Cloud for Cultural Heritage is a European Union initiative for a digital infrastructure that will connect cultural heritage institutions and professionals across the EU. EOSC: The European Open Science Cloud (EOSC) aims to create a trusted, open, and multidisciplinary environment for researchers and innovators in Europe. PPP: A Public-Private Partnership (PPP) is a collaborative arrangement between government entities and private sector companies aimed at financing, designing, implementing, and operating projects or services traditionally provided by the public sector. RDM: Research Data Management refers to the suite of practices, policies, and processes used to handle data throughout the lifecycle of a research project. W3C: The World Wide Web Consortium (W3C) is an international community that develops standards for the World Wide Web. Their mission is to lead the Web to its full potential by creating technical specifications and guidelines that are designed to be open and royaltyfree. These standards include HTML, CSS, and other web technologies, which ensure that web content is accessible across different browsers and devices. Our glossary is harmonised with relevant music-sector specific standards and with the • ISO Information technology Vocabulary (ISO 2023b); Cloud computing — Taxonomy based data handling for cloud services (ISO 2020); Cloud computing — Interoperability and portability (ISO 2017a); Metadata registries (MDR) — 1. Framework (ISO 2023a) standards and the Information and documentation — Foundation and vocabulary (ISO 2017b) standard. • ISO Information technology Artificial intelligence — Concepts and terminology (ISO/IEC 2022) and Artificial intelligence — Management system and (ISO/IEC 2023) standard’s vocabulary. 18 1 Policy context and problem map The European music ecosystem has undergone disruptive transformations in recent decades. In the 2010s, the arrival of agentic AI in streaming platforms radically reconfigured distribution and consumption. These systems centralised global sales, expanding the commercially available repertoire in a typical EU country from roughly 100,000 titles to over 100 million titles competing for attention. At the same time, the average transaction value collapsed from around €18 (in current prices) to less than €0.005. This shock hollowed out much of the traditional infrastructure — record stores, radios, and music television — and shifted value capture toward data-driven platforms able to control access through recommender algorithms. In the 2020s, the rise of generative AI further exacerbates this situation. Large-scale models can mass-produce new compositions and recordings, often imitating or plagiarising patterns of human creators. This inflates supply, undermines the position of professional authors and performers, and aggravates existing problems of remuneration and discoverability.1 EU-level studies and policy frameworks have recognised these dynamics and increasingly frame them as systemic challenges. The Feasibility Study for the Establishment of a European Music Observatory diagnosed the fragmented, scarce, and poorly harmonised nature of music data collection across Member States, calling it the fundamental reason for an EU-level observatory. The Music Ecosystem 2025 study reframes the sector as an interconnected ecosystem, where platformisation, market consolidation, and emerging technologies like AI interact with broader societal challenges such as precarity, gender inequality, and sustainability. The European Parliament, in its Resolution on cultural diversity and the conditions for authors in the European music streaming market, echoed these concerns with explicit calls for reform.2 1Music Ecosystem 2025: Study on the Music Ecosystem (Music Moves Europe 2024); it frames the sector as an adaptive, networked ecosystem, highlights AI’s ability to disrupt on pp. 6–7, and mentions it as an opportunity particularly on p. 23. Feasibility Study for the Establishment of a European Music Observatory (Commission et al. 2020); stresses the fragmented, scarce, and poorly harmonised nature of music data (pp. 9–10), the need for cooperation with rights organisations, statistical agencies, and industry stakeholders (p. 61), and introduces CEEMID as a best practice (pp. 147–148). CEEMID emerged from Budapest, Bratislava, and Zagreb as an early effort to address data poverty in Eastern EU Member States. 2European Parliament Resolution on cultural diversity and the conditions for authors in the European music streaming market (European Parliament 2024); it recognises streaming as the dominant global revenue source while leaving many authors with very low income (recitals F–H), stresses accurate metadata allocation at the time of creation using identifiers ISWC, ISRC, ISNI, IPI, and IPN (recital R, and 9.), highlights the lack of quality data to properly identify authors, performers, and rights holders (recital L), and warns that AI-generated tracks are flooding streaming platforms, aggravating discoverability and remuneration imbalances (recital O). 19 Our policy brief positions itself within this landscape. It aims to support and extend the Music Moves Europe framework by highlighting six crucial dimensions: 1. Practical solutions, grounded in dialogue between research and industry, and inspired by concrete experiences with open, federated data-sharing approaches. 2. Potential pitfalls where well-meaning initiatives may clash with legacy systems, existing business practices, or contradictions in legislation. 3. Legal and operational conflicts, such as the tension between GDPR’s data protection regime and the Berne Convention’s requirement of author attribution. 4. Cooperation and workflow sharing, recognising that no single actor can bear the full burden of metadata documentation. 5. Technology, including automation, entity recognition, reconciliation, and persistent identifiers. 6. AI adaptation and cooperative infrastructures, since most stakeholders cannot attract or retain scarce AI expertise. By foregrounding these issues, the brief complements the calls of the Music Ecosystem 2025 study and the European Music Observatory feasibility study, while remaining attentive to the practical challenges of implementation across Europe’s diverse music and cultural landscapes. 1.1 Three structural pressures Three structural pressures frame today’s metadata challenges: 1. Extreme efficiency pressure. Music is now monetised in micro-transactions worth a fraction of a cent. Each metadata mistake means lost royalties, while big-tech platforms enjoy economies of scale that self-releasing artists, small labels, and national CMOs cannot match. 2. AI-driven disruption. Agentic AI in streaming platforms has already displaced much of the traditional retail and promotion infrastructure. Generative AI risks flooding platforms with derivative works and further destabilising discoverability and revenues. Yet AI tools could also support documentation and reconciliation — if governance frameworks can enable them. 3. Governance and incentive conflicts. Identifiers such as ISWC, ISRC, ISNI, and IPN are essential for attribution and royalty distribution, but are maintained under costly, largely private regimes. Public policy increasingly demands more open metadata, but sustaining investment in these registers remains a challenge. These pressures mean that improving metadata is not only a matter of technical interoperability. It is also a question of economic sustainability, legal coherence, and cultural policy. 20 1.2 National and European pilots as anchors From the outset, we draw on concrete pilots that illustrate both the problems and possible solutions. Two of them — the Slovak Comprehensive Music Database (SKCMDb) and Unlabel — will recur throughout this paper as reference points. Together, they anchor the three thematic chapters: curation (Chapter 2), observatory (Chapter 3), and AI (Chapter 4). 1.2.1 The Slovak Comprehensive Music Database (SKCMDb) SKCMDb is our national pilot for federated metadata governance. It links together data from collective management (SOZA), national and city libraries, and archives, while ensuring that works can also be discovered in the digital environments where people actually listen: Spotify, YouTube, Apple Classical, and others. A further layer reconciles this metadata with the Slovak Statistical Office via a Satellite Business Register, so that cultural production is visible in official economic data. The SKCMDb is anchored in the Memorandum of Understanding signed between collective management organisations (SOZA, SLOVGRAM), cultural institutions (Hudobné centrum, Slovak National Library, Hudobný fond), and Reprex. This MoU formalises a federated governance model where: •Attribution (names of authors, performers, composers) is preserved as legally mandatory under copyright law. •Privacy is safeguarded by layered access: public data (names, works, identifiers) circulate broadly, while sensitive data (e.g., addresses, birth dates) remain restricted. •Interoperability is achieved by aligning with VIAF, ISNI, ISWC, ISRC, and Europeana. As such, the Memorandum provides the legal and institutional foundation for SKCMDb, turning a technical pilot into a national dataspace aligned with the EU Data Strategy. The SKCMDb in action The chart illustrates the biography and works of Slovak composer Iris Szeghy as an example: 21 Figure 1.1: A slide taken from: SKCMDb: Interoperability of Music Libraries and Archives with Public and Private Music Services (presentation at the IAML 2025 conference in Salzburg) <https://zenodo.org/records/16634558> •Left side: reconciliation of her works across SOZA, the Slovak National Library, the Bratislava City Library, and archives. •Right side: linking to listening platforms (Spotify, YouTube, Apple Classical). •Bottom: reconciliation with the Slovak Statistical Office via the Satellite Business Register. SKCMDb thus acts as a bridge between cultural memory institutions, rights management, digital distribution, and public policy. SKCMDb provides a pragmatic response to fragmentation and duplication. It anchors the discussion of preventive metadata strategies in Chapter 2. This challenge is not unique to Slovakia. A recent Horizon Europe policy brief has highlighted how inadequate metadata infrastructures and fragmented European initiatives risk leaving the field open to dominance by extra-European players (for example, the US Mechanical Licensing Collective).3 3See Policy Brief 1: Music Metadata Mainstreaming and EU Law (Senftleben et al. 2024) (Deliverable D5.6, OpenMusE project). That brief emphasises that without a European metadata infrastructure, EU repertoires may remain underexploited and culturally invisible, while foreign platforms consolidate hegemony. The present Green Paper extends on this line of argument by focusing on lifecycle-based interoperability and federated observatories as safeguards for European sovereignty.The Policy Brief 1 Annex references the Slovak Listen Local / SKCMDb project as a national pilot, underlining its relevance for EU-level policy design. The Green Paper complements this by situating the MoU as a replicable governance framework for federated metadata spaces. 22 1.2.2 Unlabel If SKCMDb focuses on building preventive infrastructures, Unlabel demonstrates how to repair the past. It is a collaborative pipeline connecting archives, libraries, collective rights organisations, and distributors to bring under-documented repertoires into the global digital supply chain. A striking example is the case of Hilda Griva, a bilingual Livonian–Estonian artist active in the interwar Finno-Ugric revival. Her recordings were rediscovered in the Latvian Archives of Folklore but lacked the metadata required for circulation. Through Unlabel, we translated and enriched her records, reconciled them with international authorities, and extended them with DDEX catalogue transfer metadata, enabling release via Spotify, YouTube, and Apple Music. ĹNote Infobox: Unlabel and Hilda Griva • Metadata repair began with archival records in the Latvian Archives of Folklore. • Records were translated, enriched, and reconciled with Wikidata, MusicBrainz, and VIAF. • DDEX-compliant catalogue transfer metadata enabled digital distribution. • The enriched catalogue allowed Hilda Griva’s recordings to be released and discovered globally. Unlabel demonstrates how public heritage institutions and private distributors can cooperate through shared standards. It anchors both the curative AI approaches in Chapter 4 and the observatory perspective in Chapter 3. 1.3 Quest for efficiency Technological progress, digitisation, automation, and now AI have transformed the music industry more dramatically than most sectors. After the collapse of the CD era under peerto-peer piracy, a newly configured recording industry emerged around global platforms. Traditional retail and wholesale jobs largely disappeared, replaced by streaming platforms such as YouTube, Apple Music, and Spotify. This shift coincided with a structural devaluation of music. The licensed streaming model never recovered the real revenues of the pre-collapse recording market, and from this diminished base, platforms take a significant share. Where a CD sale once brought around €10–18 in today’s terms, the unit of account in streaming is a fraction of a cent — typically $0.003–0.005 per play. To replace the economic weight of a single album sale, a rightsholder must now process and account for roughly 4,000 successful streams. This is not merely an economic shift, but an 23 administrative revolution. The documentation efficiency needed to handle millions of micro-transactions profitably is far higher than in the pre-streaming era. Streaming platforms are genuine big-data companies. Alphabet’s YouTube, Apple, and Spotify operate at a scale where billions of transactions and hundreds of millions of assets can be managed by autonomous agents and recommender engines. But the typical rightsholder — a self-releasing artist, an independent label, or even a national collective rights agency — works at a scale where each metadata mistake means lost royalties, and where IT or documentation specialists are often absent altogether. This asymmetry is so stark that even major CMOs rely on shared infrastructures like the digital services of “Mint” to manage repertoire at scale. Music, then, is now sold in extremely low-value transactions mediated by autonomous agents. This reality enforces a very strong pressure on the entire ecosystem to improve data interoperability and metadata quality. By contrast, in most industries administrative overhead is modest: • Retail/distribution: ~2–5% of net sales • Manufacturing: ~3–7% • Professional services: 10–15% (because administration blurs into the product) • OECD/EU cross-industry averages: 3–8% of turnover In “normal” industries, then, €50 of administrative cost is justified on €1000 of revenue. By comparison, in the recorded music industry, achieving that same 5% efficiency requires delivering faultlessly some 200,000 streaming transactions. This is a very tall order for a sector dominated by micro-enterprises and small independents without dedicated IT or metadata teams. The pressure for efficiency is not only present on the production side of the music business. In the non-profit sector, digitisation has profoundly transformed the workflows of archives, libraries, and heritage institutions as well. Streaming has reduced demand for physical collections, forcing libraries to reframe their role around digitisation, knowledge organisation, and community functions rather than lending CDs or scores. New spaces like creative studios and digital repositories are expected, but funding is limited, so efficiency is critical. At the same time, the vast amount of born-digital assets — and now the endless output of generative AI systems — creates a puzzle for archives that remains unsolved today.4 Metadata as provenance In today’s music ecosystem, almost every asset is born digital. A modern composer’s score is produced in notation software; a performer’s recording originates as a digital file; even 4See for example the Katona József Library’s adaptive strategies (Virág 2024). Archives, on the other hand, face a problem that instead of receiving records on paper, they are becoming gigantic data silos in the age of born-digital documents. They are being transformed into data through digitisation and born-digital records, face volumes too large for manual processing. This pressures traditional archival concepts such as provenance, original order, fixity, and authenticity (Colavizza et al. 2022). 24 printing, distribution, and promotion leave their own digital traces. From the very start, each musical work and each recording comes with a dense digital fingerprint. As these works move through their lifecycle — composition, registration, performance, recording, distribution, preservation — they accumulate provenance statements:“X composed this,” “Y registered that,” “Z archived this file.” Taken together, these traces form a chain of knowledge about the history of the work. Unlike in earlier centuries, this history is now almost continuously captured, though often fragmented or messy — the “shadows” that Karabinos has described. Figure 1.2: The PROV model helps us describe the lifecycle of music: who did what, when, and with what. A composer, performer, or software tool (agent) engages in an activity such as composing or recording, which results in a musical work or a sound recording (entity). Capturing these links over time makes provenance transparent, ensures correct attribution, and supports trustworthy data exchange across the music sector. Reuse: DOI: 10.6084/m9.figshare.30073210 Metadata is “data about data.” But in practice, what counts as data or metadata is relative: a duration may be descriptive for one actor, identifying for another, and algorithmic input for a third. This distributed record of provenance resembles a chain of statements, some verifiable, some contradictory, some lost in the shadows. The challenge is not to build a single immutable blockchain, but to make the distributed record reliable, reusable, and interoperable. As shown in Section 1.2, pilots like SKCMDb and Unlabel provide two complementary responses: preventive governance of metadata at creation (Chapter 2), and curative repair of legacy repertoires (Chapter 4; Chapter 3). 25 its clients release. None of these logics are wrong, but they are different. This is why attempts to force everything into one universal collection model have failed. In abstract terms, there is no single “conceptualisation” of the world that can fit a rights management organisation, a library, and a music archive equally well. On a very abstract level, the same lesson was drawn in mathematics and philosophy: Gödel showed that no formal system can capture all truths within itself, and Quine argued that reference is always relative to a conceptual scheme. In computer and information science, we know this as the impossibility of a universal ontology that could serve all databases.3These limits are well understood, but recognising them is not an excuse for inaction. It means we should work pragmatically: accept that multiple logics exist, and focus on making them interoperable where possible. ĹWhy collections differ in databases •Libraries collect under legal deposit rules: every book or score published in a country must be included, regardless of popularity. •Archives follow provenance: they keep what an organisation or individual produced, not necessarily what is “important.” •Collective management organisations (CMOs) must register only what their members submit — the collection reflects contracts and repertoire, not cultural completeness. •Distributors take what their clients release: the “collection” is shaped by market demand and contracts. Each of these logics is valid, but none can be reduced to the others. This is why a single “grand ontology” for all collections is not achievable. The pragmatic task is to connect them through lightweight, modular patterns that allow data to flow across boundaries while respecting institutional differences. 2.1.4 Legacy metadata The European Parliament has emphasised that accurate and standardised metadata is essential for ensuring fair remuneration and proper attribution in the music streaming 3As information science shows, a collection is not a mathematical set but a socially and institutionally constructed grouping, shaped by curatorial or organisational logics. Attempts to create one “gigaontology” for music metadata have consistently failed, because the sector is too heterogeneous — collective management organisations, libraries, archives, platforms, and distributors operate under different standards and governance models. At a more philosophical level, Quine reminds us that any ontology is relative to its conceptual scheme, and there is no absolute description of the world that can serve all purposes equally ((Quine 1968)). Gödel’s incompleteness results, likewise, show the inherent limits of formal systems, underscoring why computer science and database theory recognise that no single universal ontology can capture all possible cases. 32 market. It calls for identifiers such as ISWC, ISRC, ISNI, IPI, and IPN to be allocated at the moment of creation, and warns that the flood of AI-generated tracks will worsen discoverability and revenue imbalances if metadata remains incomplete or inconsistent.4 In practice, achieving this goal has proven very difficult. The registers that underpin music metadata are privately governed, require continuous investment, and cannot simply be rebuilt from scratch. Hundreds of millions of assets are already circulating, and billions of transactions are handled annually on the basis of this legacy infrastructure. Even the term metadata is ambiguous: in libraries and IT it means descriptive information (title, genre, provenance), but in the music industry it usually refers narrowly to administrative identifiers that drive royalty distribution. This gap in terminology adds to confusion and misplaced expectations. ĹForward-looking identifier pilots: PRS Nexus and Teosto ISNI Two recent initiatives show how the industry is moving towards better identifier coverage at source: •PRS for Music – Nexus. A new portal linking works (ISWC) and recordings (ISRC) at the moment of release. It already covers nearly 3 million works and offers APIs for rights-holders and DSPs (PRS for Music 2023; World Intellectual Property Organization (WIPO) 2023). By embedding ISWC allocation into distribution workflows, Nexus aims to accelerate royalty payments and reduce reconciliation delays that often last months or years. •Teosto – ISNI for authors. The Finnish CMO Teosto now assigns ISNIs to its members, giving authors and composers persistent identifiers that interlink with VIAF, ORCID, and Wikidata (Teosto 2024). This connects music rights data with library and research infrastructures and strengthens international interoperability. These projects simplify metadata at the point of creation and release, aligning with persistent identifier strategies in the research sector (Cruz and Tatum 2021). But they mainly address future repertoire. The much larger challenge lies in the hundreds of millions of legacy assets already circulating without complete identifier links — a problem that requires complementary solutions, discussed later in this chapter. Together, ISRC (recordings), ISWC (works), and ISMN (printed music) form the backbone of music identification. In theory they provide global coverage, but in practice they remain fragmented: many recordings never receive identifiers, links between identifiers are often missing, and uptake is uneven across registries. This fragility makes the European Parliament’s ambitions difficult to realise without new layers of interoperability, observability, and shared responsibility. The sheer growth in repertoire makes this gap impossible to close with manual workflows: by 2024, more music was released in a single day than in 4European Parliament resolution of 17 January 2024 on cultural diversity and the conditions for authors in the European music streaming market, recital 32 (European Parliament 2024). 33 the entire year of 1989 (Abing 2024)5. This scale of legacy under-documentation cannot realistically be resolved with manual workflows alone — it points directly to the need for curative AI approaches, which we return to in Section 4.2. Although metadata repair is indispensabl, metadata is never netural. Without corrected identifiers, reconciled names, and enriched annotations, works remain invisible in royalty and discovery systems. However, just as heritage data spaces show how repairing metadata can restore visibility while also reinforcing institutional logics, in music ecosystems the same repair practices can unexpectedly increase exposure to generative AI. By making works more legible to agentic applications, enriched metadata improves attribution but also sharpens the ability of AI systems to imitate and substitute. This paradox is most acute for small-scale repertoires and independent artists, whose economic position mirrors the epistemic vulnerability of minority heritage collections. ĹCase Study: Metadata Repair — Heritage and Repertoire Repairing heritage metadata - In the Finno-Ugric Data Sharing Space we worked with the Latvian Archive of Folklore and regional museums to repair and enrich metadata around Livonian, Latvian, and Hungarian folk music. - In Hungary, together with the House of Music, we began repairing the lost documentation of recordings suppressed under Communist censorship. Here, repair is not only a matter of accuracy but also of restitution: without corrected metadata, these works remain locked behind outdated copyright classifications long after the state label monopoly has ended. - Original records in both contexts were shallow, monolingual, and shaped by institutional or censored taxonomies. By reconciling names, places, languages, and cultural terms, we enabled works to be rediscovered across Wikidata and Wikipedia. - These are extreme cases of damaged metadata (through censorship, Soviet-type copyright, or minority language non-standardisation). Yet similar problems affect the long tail of European music heritage and today’s independent or self-releasing artists. - As our forthcoming academic paper shows, this is not a neutral “clean-up”: choices about vocabularies and identifiers determine what communities can see of themselves. Repair here means cultural repair — restoring epistemic visibility to communities, legal heirs, and cultural stewards. Repairing repertoire metadata - Through the Unlabel prototype, we apply similar practices to contemporary self5The International Standard Recording Code (ISRC) was introduced in 1986 as a 12-character identifier for recordings (ISO 3901) and is managed operationally by IFPI (ISO 2019a; International ISRC Registration Authority 2021). Persistent problems include retroactive assignment, inconsistent embedding, and weak interoperability with ISWC (Paskin 2006, p4). The International Standard Musical Work Code (ISWC) identifies compositions and lyrics (ISO 15707), managed by CISAC through the ISWC Agency (ISO 2022). Challenges include duplicate codes, mismatches with ISRC, and uneven adoption by CMOs (Paskin 2006, p7). The International Standard Music Number (ISMN, ISO 10957) identifies printed music publications (ISO 2021). It provides a bridge between bibliographic and rightsmanagement practices, but remains underused in digital workflows. 34 released music: enriching works with ISRC/ISWC codes, multilingual annotations, and library-standard metadata. - This makes previously “invisible” tracks legible to streaming platforms and collection societies, improving discoverability and royalty flows. - Again, repair is not neutral: the way identifiers and categories are assigned shapes how artists’ works are found, monetised, or sidelined. The paradox - These cases illustrate that metadata is never neutral. Repair empowers artists and communities, but it also encodes assumptions and makes works more legible to agentic applications. - In heritage, institutional schemas may flatten local epistemologies; in the market, generative AI may exploit enriched metadata to imitate and substitute — a problem we will discuss in Chapter 4. - In both contexts, metadata repair empowers and exposes — visibility and risk are two sides of the same process, which makes metadata governance a policy concern, not a purely technical one. Our approach - Our solution is to use decentralised systems like Wikidata and Wikibase together with strong ontological patterns. - Heavy-weight ontologies take up to a decade to develop, may introduce new biases through the non-neutral nature of metadata, and by the time they are created, they may not address new challenges — for example, providing guardrails against negative outcomes of agentic or generative AI. - As with the infrastructure in Chapter 3, we aim for decentralisation already at the metadata-definition level. An Open Music Observatory will allow metadata to be managed through flexible, open processes that create definitions and establish equivalences to existing standards. 2.1.5 Named-entity resolution, attribution, and privacy Attribution is not optional in music: the names of authors, performers, and producers are structurally necessary for copyright, royalties, and cultural record-keeping. Yet under GDPR, these names count as personal data, creating a contradiction at the very foundations of metadata curation. What is mandatory under copyright law becomes a liability under data protection law. In practice, private actors face repeated balancing tests, inconsistent interpretations, and the risk of complaints even when attribution is legally required. This contradiction drives up costs and discourages investment in better metadata. Small publishers and self-releasing artists already face disproportionately high OPEX (documentation, bookkeeping) and CAPEX (IT systems). Without affordable, legally secure ways to resolve named entities, their works perform badly on platforms and royalties are lost. Policy communities in Europe recognise these issues. The Big Data Value Association (BDVA) has long argued that trust frameworks and governance pillars are essential 35 for data sharing, while the Federation Working Group stresses that federation — not centralisation — is the only realistic model for connecting Europe’s fragmented data ecosystems (Big Data Value Association 2019; BDVA/DAIRO 2023; BDVA/DAIRO Federation Working Group 2023). These principles apply equally in music. But given the sector’s extreme fragmentation and micro-enterprise structure, implementing them here is especially difficult. How these structural problems can be addressed at systemic level is the subject of Section 3.1.3, where we show how data sharing spaces provide a way forward. 2.2 Policy proposals Figure 2.3: Explanation 2.2.1 Reducing redundancy The European Parliament has rightly highlighted that fragmented and unreliable metadata remains a major obstacle in the music sector. We agree with this diagnosis, but stress that the root cause lies partly in the need for backward compatibility with hundreds of millions of legacy assets, and in the costly redundancy of today’s practices: the same information must be repeatedly entered into separate systems such as ISNI, ISWC, ISRC, VIAF, or local authority files. This duplication creates errors, increases costs, and discourages accurate registration. 36 Our policy solution is to support redundancy-free registration by aligning the workflows of those who already maintain authoritative data. Instead of duplicating efforts, registration steps can be coordinated once and reused many times. We demonstrate this approach with our Open Music Registers pilot: a federated infrastructure that interconnects persistent identifiers (ISWC, ISRC, ISNI, VIAF) and, where relevant, links them to business and statistical identifiers (e.g. OpenCorporates, NACE, ISCO). This allows music creators and organisations to benefit from smoother workflows, while downstream users gain more reliable data for royalty distribution, cultural visibility, and AI-driven discovery. The Open Music Registers deliberately avoid centralisation. Each registrar — collective management organisations, libraries, archives, or statistical offices — retains ownership of its data but contributes to a shared semantic framework.6By connecting rather than merging registers, redundancy is reduced while subsidiarity, accountability, and trust are safeguarded across public and private actors. This distributed model directly answers European Parliament’s call for metadata systems that are reliable, inclusive, and supportive of creators.7 2.2.2 Reconciling attribution and privacy The problem of reconciling copyright attribution with GDPR obligations cannot be solved by ignoring either side: both are binding legal requirements. Our approach, tested in the Slovak Comprehensive Music Database (SkCMDb), shows that progress is possible through layered governance and careful balancing. Academic institutions and libraries, with their cultural and research mandates, can lawfully handle personal data under derogations for public-interest processing. Collective management organisations (CMOs) and private actors, by contrast, must rely on legitimate interest tests, supported by transparent documentation, notification to rightsholders, and opt-out mechanisms where possible. ĹInteroperability is a means, not a goal Our Slovak pilot, the Slovak Comprehensive Music Database (SKCMDb), links libraries, rights management, streaming services, and the statistical office. This is not “interoperability for its own sake.” Ontologies and crosswalks are valuable only insofar as they enable better services: 6Technically, this corresponds to a provenance-oriented modelling approach such as the W3C PROV-O standard (W3C 2013b, 2013a), which connects actors, activities, and entities in chains of attribution (“a composer authors a work, a performer interprets it, a producer records it…”). These chains can be expressed in the layered terms of the European Interoperability Framework (EIF), ensuring legal, organisational, semantic, and technical interoperability (Commission and Digital Services 2017). 7The Data Spaces Support Centre (DSSC) Blueprint v2.0 underlines that identifiers and rulebooks are the foundation of any common European data space (Data Spaces Support Centre 2025b). In the music sector, however, attribution identifiers themselves are caught in the GDPR contradiction (see Section 2.1.5), which underscores the importance of redundancy-free but legally robust registration practices. 37 •For audiences: making music findable and accessible across cultural and commercial platforms. •For rightsholders: ensuring that attribution, identifiers, and royalty flows are correct. •For policymakers: providing reliable data to support cultural policy and to measure the music economy. In short, interoperability at the data level is the condition for usable services at the societal level. The Slovak Memorandum of Understanding shows how attribution and data protection can be balanced in practice. -Names of authors, performers, and producers are treated as public-interest information necessary for copyright and royalty flows, justified under legitimate interest. -Sensitive fields (e.g., addresses, nationality, pseudonyms) are excluded from public layers and restricted to controlled-access tiers. -Governance is distributed across CMOs, libraries, and archives, ensuring subsidiarity and trust. This layered compliance model demonstrates that copyright attribution and GDPR obligations can coexist — and offers a template for other Member States and for the European-level Open Music Observatory. Balancing tests play a central role: each dataset is audited, divided into public and nonpublic categories, and then assessed again for personal vs. non-personal data. Public information such as names of authors, performers, and work titles—already widely available in catalogues and concert programmes—can justifiably be shared under legitimate interest, especially when linked to rights management purposes. Sensitive data (e.g. addresses, nationality, pseudonyms) require stricter access tiers and are only made available to selected stakeholders under contractual safeguards. This layered compliance model does not eliminate GDPR challenges, but it creates a robust defence: it demonstrates that the legitimate interest in accurate attribution and royalty distribution outweighs the minimal risks of publishing already public information. In practice, this means rights metadata can circulate across the ecosystem while privacysensitive data are contained. Building such workflows into federated observatories and data spaces allows the music sector to comply with data protection rules without undermining attribution, and provides a model for European-scale solutions. More broadly, these governance practices are supported by existing provisions in EU copyright and data legislation that already give metadata a central role. Rights Management Information (RMI) is explicitly protected under Article 7 of the InfoSoc Directive (2001/29/EC), making the removal or alteration of attribution data unlawful. The CRM Directive (2014/26/EU) obliges collective management organisations to maintain accurate and transparent repertoire and membership data. Under the CDSM Directive 38 (2019/790/EU), Article 17(4)(b) requires platforms to act expeditiously on notices where metadata enable rightholders to identify and claim their works, while Article 4(3) uses metadata as the operational basis for text and data mining opt-outs. Beyond copyright, the Data Governance Act (2022/868), the Data Act (2023/2854), and the Open Data Directive (2019/1024) provide the horizontal framework for treating music metadata as part of Europe’s emerging common data spaces.8 2.2.3 Pragmatic metadata alignment Attempts to build one comprehensive, harmonised schema for music metadata have repeatedly failed. The sector is too diverse: collective management organisations, libraries, archives, distributors, and platforms all operate with different standards and governance models. Trying to impose a single “grand schema” has proven brittle, costly, and unrealistic. A more workable solution is modular alignment. Instead of a single heavy ontology, small reusable building blocks can be combined to describe recurring patterns — for example, how people, works, recordings, and performances are related. This approach allows interoperability to grow step by step, without forcing any actor to abandon its systems.9 It also helps to separate two complementary tasks. On the one hand, we need conceptual scaffolding that lets different databases describe similar structures in comparable ways. On the other, we need identifier reconciliation to make sure that the same person, work, or recording can be linked across different registers. Neither of these tasks is sufficient on its own: they must work together if metadata is to remain reliable at scale.10 Other domains show how this can be done. Research infrastructures have reconciled ORCID with VIAF authority files, and libraries have mapped DataCite metadata to Dublin Core. Both examples show how two different standards can be aligned systematically while keeping their distinct scopes.11 8See Policy Brief 1: Music Metadata Mainstreaming and EU Law (Senftleben et al. 2024) (Deliverable D5.6, OpenMusE project). That brief analyses how these instruments can be mobilised to improve the reliability and circulation of music metadata. The present Green Paper complements this by showing how federated observatories and interoperability strategies can operationalise these obligations in practice. 9On ontology design patterns and modular approaches, see (Gangemi 2005; Blomqvist, Hammar, and Presutti 2016; Carriero et al. 2021). The Polifonia project applied these methods at European scale (Berardinis et al. 2023), aligning with MusicBrainz and the ChoCo knowledge graph (Albanese et al. 2023). While Polifonia did not focus on rights metadata, it provides a strong foundation for connecting musicological knowledge with industry identifiers. 10This distinction between ontology modelling and identifier reconciliation clarifies why both layers are necessary. Ontology patterns provide conceptual scaffolding (e.g. work–recording–performance), while identifier reconciliation ensures that an author in ISNI is the same as a VIAF authority record or a performer in MusicBrainz. 11For ORCID–VIAF reconciliation via OpenRefine, see (OpenRefine Community 2021; Jegan et al. 2023). For systematic mappings between DataCite and Dublin Core, see (DataCite 2021). 39 Figure 2.4: Pragmatic metadata alignment relies on modular patterns, not “giga-schemas.” The example shown here from our Wikibase pilot encodes roles, events, and provenance using reusable ontology design patterns. This allowed identifiers from rights management (ISWC, ISRC) to be reconciled with library authorities (ISNI, VIAF), proving that interoperability can be achieved incrementally without forcing any actor to abandon its systems. DOI: [10.6084/m9.figshare.30075379.v1](https://doi.org/10.6084/m9.figshare.30075379.v1) Music metadata needs the same periodic reconciliation. Rights identifiers such as ISRC, ISWC, and ISMN were designed separately and drift apart if not actively maintained. The same applies to personal and organisational identifiers such as ISNI, VIAF, and IPI. Without active cross-checking, records fragment, causing duplication and inconsistency.12 In our pilots, this modular alignment has already been tested. The Slovak Comprehensive Music Database reconciled rights identifiers with library authorities without schema unification. MusicBase used Wikibase to encode roles, events, and provenance in a way that let corrections propagate across systems. The Unlabel workflow streamlined metadata capture for self-releasing artists and libraries, allowing once-only documentation to be reused across distribution and preservation. These cases extend our proposal for Open Music Registers, which argued for federated, redundancy-free metadata workflows, into the broader governance framework of this Green Paper. Finally, this approach is consistent with work in the heritage sector. The Heritage Digital Twin Ontology (HDTO), developed within the European Cultural Heritage Cloud, uses the same principles of modularity and federation to describe tangible and intangible assets. Where HDTO provides a semantic framework for heritage “digital twins,” the Open Music 12On the divergence of identifiers if not maintained, see (Paskin 2006). 40 Observatory extends the same logic to music. Both models show how cultural and rights metadata can integrate with wider European data spaces while preserving subsidiarity and institutional diversity.13 13The ECHOES Heritage Digital Twin Ontology (HDTO) builds on CIDOC CRM extensions to model tangible and intangible heritage with space–time–cultural identity (ECHOES Ontology Task Force 2025). 41 dedicated workflow for music, and industry uptake remains minimal. As with ECCCH, music is underrepresented and rights-aware curation pathways are absent.9 The European Interoperability Framework (EIF) helps explain why these gaps persist. Interoperability depends not only on formats but also on legal, organisational, semantic, and technical alignment. Without shared governance and profiles, public and private systems diverge. The principle of subsidiarity adds another layer: stewardship over cultural data is distributed across national and regional authorities, as well as private actors. Centralisation is therefore both impractical and politically illegitimate. The challenge is not whether decentralisation should exist, but how to make decentralised contributions work together.10 This challenge directly motivates the Observatory’s bridging role with EOSC, Europeana, and ECCCH, elaborated in Section 3.211. 3.1.5 Subsidiarity and infrastructures for scaling music data The European principle of subsidiarity requires that decisions be taken as closely as possible to the citizens they affect. In cultural policy, this means that responsibilities are distributed across multiple levels: in some Member States, culture is managed regionally or provincially; in others, nationally. Beyond public administrations, many important datasets are held by private actors — collective management organisations, platforms, or archives. Any attempt to centralise music data governance would therefore risk losing both legitimacy and local relevance. Instead, subsidiarity must be built into the design of the Observatory. The European Interoperability Framework (EIF) provides a layered model — legal, organisational, semantic, and technical — for reconciling governance across institutions. The Data Governance Act (DGA) codifies the same principle: Member States retain stewardship over sensitive datasets, but EU-level standards ensure they can circulate securely and comparably across borders. The Data Space Support Centre (DSSC) extends this approach into practice, developing blueprints and building blocks that allow decentralised initiatives to scale. Together, these frameworks show how subsidiarity and federation are not barriers but design principles for data spaces.12 9EOSC provides federated access and persistence through Zenodo and OpenAIRE, but music workflows remain marginal. On EOSC’s role, see the European Strategy for Data (European Commission 2020). 10The EIF defines layered interoperability (legal, organisational, semantic, technical) (Commission and Digital Services 2017). The European Strategy for Data frames subsidiarity as compatible with federation (European Commission 2020). BDVA and the Federation Working Group emphasise that interoperability frameworks are needed to operationalise federation (BDVA/DAIRO 2023; BDVA/DAIRO Federation Working Group 2023). 11The EIF defines layered interoperability (legal, organisational, semantic, technical) (Commission and Digital Services 2017). The European Strategy for Data frames subsidiarity as compatible with federation (European Commission 2020). BDVA and the Federation Working Group emphasise that interoperability frameworks are needed to operationalise federation (BDVA/DAIRO 2023; BDVA/DAIRO Federation Working Group 2023). 12On subsidiarity and federation: the Data Governance Act (European Parliament and Council 2022) and the European Strategy for Data (European Commission 2020). On technical frameworks: DSSC’s 48 At the technical level, Wikidata and Wikibase provide a proven backbone for collaborative metadata management. They are already embedded in EU infrastructures such as the official EU Knowledge Graph and in national projects like MetaBelgica in Belgium. In Flanders, the performing arts field has gone further: since 2017, Kunstenpunt and meemoo have published decades of performing arts data on Wikidata, showing how enrichment happens automatically once data becomes part of a wider ecosystem. These pilots illustrate how subsidiarity and federation can work in practice, with decentralised actors maintaining control of their own data while contributing to a shared framework.13 The problem of scale makes such infrastructures essential. Large platforms and labels can manage millions of assets cheaply, but small actors cannot. Without shared systems, independent and community-based repertoires remain undocumented because the cost of proper registration exceeds likely revenue. Federated tools — strengthened by automation and AI — are the only realistic way to close this gap. ĹFinno-Ugric Data Sharing Space Our pilot with the Finno-Ugric Data Sharing Space illustrates subsidiarity in practice (see: https://finnougric.net/). By collaborating with regional NGOs and national archives, we curated and repaired datasets that would have remained invisible in a central repository. The project showed that decentralised actors are best placed to manage their own data, but that interoperability frameworks and shared observability layers can connect them effectively 14. International comparison confirms this. In the United States, the Mechanical Licensing Collective (MLC) was created in 2021 to administer a blanket mechanical license for streaming and downloads. It inherited more than $424 million in unmatched royalties and developed large-scale reconciliation systems to allocate them. By 2022, it had already distributed nearly $700 million. The MLC shows what can be achieved when identifiers such as ISWC and ISRC are used systematically and backed by law. But it also highlights the limits of centralisation: creators must still claim and maintain their records, education gaps persist, and disputes between platforms and rights bodies continue.15 ĹThe U.S. Mechanical Licensing Collective (MLC) The Mechanical Licensing Collective was created under the U.S. Music Modernization Act (2018) to administer a blanket mechanical license for streaming and downloads. It inherited more than $424 million in unmatched royalties from digital services and developed large-scale reconciliation systems to allocate them. By late 2022, it had blueprints (Data Spaces Support Centre 2025b, 2025a). On governance: BDVA (BDVA/DAIRO 2023) and the Federation Working Group (BDVA/DAIRO Federation Working Group 2023). 13On official adoption: EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021); SEMIC guidelines (SEMIC Support Centre 2023). On Belgian pilots: MetaBelgica (Stallmann et al. 2023) and Flemish performing arts enrichment (Magnus and Van D’huynslager 2021). 14See (Antal et al. 2025; Antal, Pigozne, and Federico 2025). 15On the MLC’s establishment and operations: (Mechanical Licensing Collective 2021); on contested governance and disputes with platforms: (Varghese 2024). 49 distributed nearly $700 million. The MLC shows what can be achieved when identifiers (ISWC, ISRC) are captured systematically and backed by legislation. But it also highlights the limits of centralisation: creators must still claim and maintain their records, education gaps persist, and disputes between platforms and rights bodies continue. For Europe, the lesson is clear: scaling metadata infrastructure is possible, but it must respect subsidiarity and federation rather than rely on a single central clearinghouse (Mechanical Licensing Collective 2021; Varghese 2024). 3.1.6 Economies of scale in metadata Large platforms and major labels can document millions of tracks at very low per-unit cost, because they manage everything in bulk. Smaller actors — independent labels, nonprofits, or community archives — face the opposite situation: the cost of registering and maintaining each track is often higher than the revenue it will ever generate. This imbalance explains why so many “frozen” assets remain unregistered and invisible in today’s digital ecosystem. Without a way to share infrastructure, small actors remain stuck. They cannot afford the per-track cost of full documentation, yet under-documentation ensures their work remains undiscovered. This is not just an accounting issue, but a structural barrier to diversity in music data flows. A federated approach, as outlined in Section 3.2.2, is essential to rebalance these inequalities and enable small actors to benefit from the same efficiencies as global players.16 3.2 Policy Proposals ÁEditing reminder • Open Music Observatory as the convening + conformance + observability layer (not a single database). • Workflow playbooks: rights→distribution→charting→preservation; changepropagation patterns; provenance trails that survive system boundaries. • Legal/standards/public investment inline: GDPR legal bases per flow; recom16Comparative research shows that costs per asset decrease sharply with catalogue size, creating scale advantages for majors and global platforms. Without shared infrastructures, small actors are disproportionately disadvantaged. The Feasibility Study for a European Music Observatory emphasised this imbalance as a structural barrier (Commission et al. 2020, p9), while the Music Ecosystem 2025 study highlighted how fragmentation and duplication reinforce these scale inequalities (Music Moves Europe 2024). 50 mended codes of conduct; lightweight policy for data fitness/quality; funding hooks (ECCCH pilots, national ministries). ĹPublic–private reconciliation in practice Reconciling public and private infrastructures: The ALOADED pilot in Latvia The Unlabel workflow was tested with Latvian archives and the distributor ALOADED, showing how public heritage metadata can be reconciled with private supply chains. • Archival recordings (Hilda Griva’s songs and Latvian/Latgalian midsummer songs) were located in the Latvian Archives of Folklore. • Metadata was translated, enriched, and aligned with international authority files. • ALOADED extended this metadata with DDEX-compliant catalogue transfer and ingested it into Spotify and other platforms. This demonstrated that reconciliation between public infrastructures (archives) and private infrastructures (distributors and platforms) is both technically and institutionally feasible, reconnecting suppressed or marginalised repertoires with contemporary audiences. See a more technical description of what we did here. Conformance and observability rules in the Open Music Observatory should be designed in line with the European Interoperability Framework (EIF) and the FAIR data principles. This ensures compatibility with wider European data space initiatives and reduces integration costs for institutions already adapting to these standards (Commission et al. 2020, p9). 51 Figure 3.2: The Open Music Observatory sits where open science, public sector information reuse, and music industry workflows overlap. By aligning with the European Interoperability Framework, it creates a shared space where libraries, rights managers, publishers, and researchers can collaborate. This positioning highlights OMO’s role as a bridge between cultural heritage, commercial distribution, and open knowledge. DOI: [10.6084/m9.figshare.30073267.v1](10.6084/m9.figshare.30073267.v1) 3.2.1 Workflow playbooks and provenance trails The Observatory should not only harmonise data formats but also document workflow playbooks that capture how metadata flows across the music lifecycle: • from rights registration, • to distribution and royalty attribution, • to charting and visibility, • to long-term preservation. Each step should include change-propagation rules: if a correction is made in one register, it should ripple through to others. Provenance trails must survive system boundaries, using standards such as PROV-O to show who did what, when, and under what authority. This makes corrections auditable, supports cross-border comparability, and prevents “data death” when an asset leaves its original system. 52 3.2.2 Federated infrastructure as a cost and governance solution The imbalance described in Section 3.1.6 makes one thing clear: small actors cannot compete on metadata without shared infrastructures. Federation, not centralisation, is the only viable way forward. Adata sharing space provides the framework. Instead of forcing everyone into a single metadata schema or legal agreement, it allows organisations to share and reuse data on an “as-needed” or “as-permitted” basis, while keeping full control of their own assets. For music — where rights, identifiers, and content are dispersed across hundreds of micro-actors and institutions — this model avoids both duplication and dependency. Crucially, it also avoids creating a new single gatekeeper: centralisation risks not only technical brittleness but also the emergence of a monopolistic intermediary able to close access or impose conditions on others17. Music is one of the most demanding test cases for European data governance. Attribution rules interact with privacy law, identifiers are used unevenly across the sector, and most music enterprises are too small to build their own compliance or documentation infrastructure. If a federated model can function in this environment, it can function anywhere. But decentralisation brings its own challenges: organisations with stronger infrastructures may prefer to protect competitive advantages by withholding data. Effective governance therefore has to make participation more attractive than isolation — through lower administrative costs, increased visibility, legal clarity, or shared compliance benefits. In practical terms, this means combining hard alignment (such as minimal metadata profiles and the use of basic identifiers) with soft alignment (such as mappings, crosswalks, and workflow playbooks). This mix allows different actors to negotiate interoperability without forcing anyone into a single model. For these reasons, the European Music Observatory cannot be designed as a single central database. It must act as a convening and observability layer — a place where decentralised contributions can be compared, connected, and reused. Such a structure reduces duplication, lowers costs for smaller actors, improves attribution, and provides a stable governance foundation for trustworthy AI and evidence-based cultural policy. Europe already has the policy and technical foundations for this. The European Strategy for Data, the Data Governance Act, and the Data Act all define data spaces as federated by design, supported by trust frameworks, rulebooks, and shared services.18 The Data Spaces Support Centre (DSSC) has translated these into practical blueprints that can be applied directly to the music sector.19 Other domains offer concrete precedents: the ISRC system 17Our definition here is an extended paraphrase of (Curry 2020) and reflects that a “data [sharing] space is an ecosystem of exchange, processing, sharing and provision of data between trusted partners, for a fee or not” from (EBU and Gaia-X 2022, p16). 18The European Strategy for Data (2020) defines Common European Data Spaces as federated ecosystems, while the Data Governance Act (2022) and Data Act (2023) supply the governance and access rules (European Commission 2020; European Parliament and Council 2022). 19The Data Spaces Support Centre (DSSC), led by KU Leuven with GAIA-X and BDVA, provides practical blueprints and building blocks for implementing federated data spaces in any domain (Data Spaces Support Centre 2025b, 2025a). 53 distributes responsibilities across national agencies; CISAC’s CIS-Net provides access to rights data without centralising ownership; and European statistical systems harmonise indicators through subsidiarity rather than through central repositories. The music sector can — and should — build upon the same logic: distributed stewardship, shared standards, and coordinated interoperability. Concerns about sovereignty make this design choice even more urgent. Without a European solution, metadata infrastructures risk drifting toward US-style centralisation, such as the Mechanical Licensing Collective (MLC), where market power and legislative mandates converge in a single hub. Deliverable D5.6 of the OpenMusE project explicitly warns against this outcome: unless Europe develops its own federated metadata infrastructures, it risks outsourcing control over visibility, attribution, and royalty data to foreign platforms.20 Our proposal for a federated Open Music Observatory therefore complements this legalinstitutional analysis. Where D5.6 highlights the provisions in EU law that can be mobilised, the present Green Paper demonstrates how a distributed dataspace model can translate them into practice. Taken together, they offer a dual strategy — one legalinstitutional, one cultural-sovereignty — for securing Europe’s music ecosystems. In practical terms, this means applying capture once, reuse many pipelines across the entire music lifecycle: from registration of works and recordings, through distribution and royalty attribution, to preservation and cultural statistics. To achieve this, the Observatory must be designed as a redundancy-free registration space, aligned with the European Interoperability Framework and provenance-oriented models such as PROV-O.21 Done well, this would rebalance the playing field: lowering costs for small actors, making datasets interoperable across institutions, and ensuring that Europe’s cultural and economic policies rest on reliable evidence rather than fragmented silos. 3.2.3 Legal, standards, and funding levers For these proposals to succeed, they must be backed by legal clarity, lightweight standards, and public investment: •Legal: GDPR legal bases should be specified for each data flow (legitimate interest for attribution; research and cultural heritage exemptions for archives). •Standards: Codes of conduct and minimum profiles should keep conformance achievable even for micro-enterprises. 20See Policy Brief 1: Music Metadata Mainstreaming and EU Law (Senftleben et al. 2024). The brief warns that US-style centralisation (e.g. the MLC) shows the risks of failing to establish European metadata infrastructures. This Green Paper complements this by presenting the federated, culture-led model of an Open Music Observatory as Europe’s sovereignty-preserving alternative. 21The EIF ensures interoperability across legal, organisational, semantic, and technical layers (Commission and Digital Services 2017). The W3C’s PROV model and PROV-O ontology offer a standard way to connect actors, activities, and entities in chains of attribution (W3C 2013b, 2013a). Applied together, they enable consistent tracking of economic and cultural flows without centralising databases. 54 •Funding: ECCCH pilots, national ministries, and EU programmes should explicitly support metadata fitness and data-quality improvements as public-interest infrastructure. Embedding these levers ensures that interoperability does not remain voluntary but becomes a supported and sustainable practice across Europe. 3.2.4 Alignment with the European Open Science Cloud Bridge cultural clouds and market workflows via a federated Music Data Sharing Space. Position the Open Music Observatory as the convening + conformance + observability layer that connects ECCCH/Europeana and GLAM authority files with industry pipelines. Concretely: 1. Capture once, reuse many across creation→registration→distribution→preservation. 2. Require minimal profiles that smaller actors can actually implement. 3. Prioritise identifier crosswalks (ISRC�ISWC�ISNI�VIAF/Wikidata) and changepropagation. 4. Use Wikibase/Wikidata as a low-friction backbone where appropriate. 5. Govern with EIF/FAIR-aligned rules, auditability, and PPP participation so rightsholders and memory institutions keep stewardship while interoperating. This reframes Europe’s investments from siloed repositories into a shared data space that lowers reconciliation costs, respects subsidiarity, and makes cultural metadata usable across public and commercial contexts — the practical foundation for any future European Music Observatory. 55 4 AI that Works for Music, Not Against It Most AI projects fail because they chase hype. MIT’s Project NANDA found that 95% of enterprise initiatives with generative AI delivered no measurable value. Budgets were spent on flashy pilots in sales or marketing, while the real potential — reducing back-office costs, prolonging the life of legacy systems, and avoiding constant IT churn — was overlooked. Our approach is different. We do not see AI as “for its own sake.” Instead, we treat it as a way to reduce IT churn, keep legacy systems alive longer, and cut both capital and operating expenses. Where once every new regulation, distributor change, or catalogue migration required costly upgrades, curative AI can patch outputs from existing software, extend the lifespan of old systems, and make them interoperable with new ones. Shared infrastructures make this practical for micro-enterprises, NGOs, and collective management organisations (CMOs), who could never maintain such capacity in-house. The European Parliament’s resolution on the music streaming market warns of the risks that AI-generated content poses for discoverability, attribution, and fair remuneration if metadata remains incomplete or unreliable. At the same time, the Music Ecosystem 2025 study highlights that AI will be both a disruption and an opportunity: while it can overwhelm systems with synthetic material, it also offers tools to automate documentation, reduce costs, and strengthen evidence-based policymaking (Music Moves Europe 2024, 23– 24). 56 Artificial intelligence is therefore central to the future of Europe’s music ecosystem. On one hand, it threatens to exacerbate existing inequalities by concentrating technological advantages in platforms and major rights holders. On the other, it can repair, enrich, and automate processes that are otherwise prohibitively costly for small actors. The challenge is not whether AI will be used, but whether its benefits will be distributed fairly across the ecosystem. European policy provides guidance for this balancing act. The Ethics Guidelines for Trustworthy AI underline that AI must be lawful, ethical, and robust throughout its lifecycle (Commission, Directorate-General for Communications Networks, and Technology 2019). The Getting the Future Right report by the Fundamental Rights Agency stresses the need to align AI with fundamental rights, especially where vulnerable groups and cultural participation are concerned (European Union Agency for Fundamental Rights 2020). Most recently, the AI Act enshrines a risk-based regulatory framework, defining obligations for providers and deployers of AI systems while reaffirming the principles of subsidiarity and proportionality in EU digital policy (European Parliament and Council 2024). Our own engagement with these issues began with the Listen Local feasibility study in 2020. By experimenting with the Spotify API, we discovered that Slovak users were rarely recommended Slovak music — not because Spotify was at fault, but because the data about local repertoire was sparse. Spotify’s open API was, in fact, uniquely transparent compared to competitors, and it enabled us to see a larger policy problem: without structured, machine-readable knowledge of diverse repertoires, algorithms cannot deliver fair outcomes. This lesson has guided our work ever since: improving metadata and interoperability is the first step to better AI governance. 57 The Unlabel pilot illustrates this problem: by treating catalogue transfers and documentation as high-cost, high-friction processes, valuable repertoires remain locked away. AIassisted metadata repair and DDEX-compliant catalogue transfer workflows provide a pathway to lower costs and bring neglected repertoires back into circulation. ĹNote Example: Old SQL Database in a Cultural Institution • A label or archive has a recording stored in a 20–30 year-old SQL database, built on a schema that was never fully documented. The system’s author is retired (or no longer alive). • The institution wants to re-release the recording, but to distribute it today, the metadata must be expressed in DDEX Catalogue Transfer messages — a completely different schema, designed decades later. Curative AI • Acts at the system level: it can “read” the old database structure, infer undocumented field meanings, and patch outputs so the legacy database can still talk to modern pipelines. • Instead of rebuilding or migrating the old database (expensive, risky), curative AI extends its lifespan by making its outputs usable. Reparative AI • Acts at the metadata/epistemic level: it can detect inconsistencies or missing fields (e.g., composer names stored in free-text notes, titles in mixed languages) and reformat or enrich them into structured DDEX-compliant fields. • This not only enables distribution but also restores visibility for works that might otherwise remain trapped in inaccessible formats. The policy point • Without curative/reparative AI, such recordings risk becoming “frozen assets”: legally owned but practically undistributable because the metadata cannot be transformed. • By investing in these AI uses, Europe can preserve access to cultural heritage, reduce IT churn, and ensure that both heritage archives and independent labels can connect to modern digital value chains. Unlike U.S.-style copyright, Europe’s author’s rights regime contains a moral component. Authors (and, for a period, their heirs) retain certain rights over how their works are used, even after economic rights expire. This recognises that works are part of a creator’s moral 64 and cultural heritage, not only economic assets. Various legal norms, for example, local content guidelines, also gave tool earlier to national or ethnic communities to provide some guardrails to the use of their shared heritage, even this means community stewardship and not inheritance in legal terms. Metadata repair and publication strengthen visibility, but also create risks that generative AI will use these works in ways that undermine moral rights, where heirs object to uses they see as distorting or trivialising an author’s legacy and community stewardship norms, where groups perceive their folk or minority heritage as being misappropriated, even when no legal infringement occurs. While we do not identify these challenges at this point as similarly actionable public policy challenges as the problems of GDRP and the creation of trustworthy music AI, regulators do face political risk if ethical expectations of communities around cultural stewardship are not addressed. Even if no author’s rights or other legal norms are breached, the ability to create “fake” Livonian, Latvian or Basque folk songs may strongly conflict with the expectation of communities on the ethical use of AI. 4.1.6 AI support for investment into new repertoire assets While generative AI that disregards human repertoires can undermine cultural value, AI also has constructive roles. Just as photographers benefit from embedded AI in tools like Photoshop or GIMP, musicians and producers can use AI to reduce the costs of composition, recording, and documentation. In practice, this means that creating new works and registering them with identifiers can become less burdensome and more accessible. This perspective aligns with the European Parliament’s call for “metadata from birth” (European Parliament 2024), but it goes further. AI can not only generate metadata automatically at the moment of creation, but also support sound recording, scoring, and archiving processes directly, ensuring that new assets enter circulation with complete, interoperable metadata. 4.2 Policy Proposals: Aligning AI with Governance and Value Creation Generative, agentic, and inference AI are now woven into the global creative economy. But value is not created by algorithms alone — it comes from governance, curated data, and institutions that ensure trust. Policy interventions are needed on three levels: EU,industry, and organisational. Our focus is the metadata and data needs of the music ecosystem — labels, distributors, publishers, managers, CMOs, archives — not the creative act of composing music itself. 65 4.2.1 EU-Level Policy: Compass and Guardrails •Embed cultural sectors in the EU AI Act & Data Spaces so music and cultural industries are not treated as “low risk.” •Subsidise shared AI utilities for identifier reconciliation, metadata repair, and fraud/plagiarism detection. •Adopt “metadata from birth” principles: embed ISNI/ISWC/ISRC identifiers at the point of creation. •Tax incentives for onboarding frozen assets, supporting digitisation and enrichment of under-documented catalogues. •Resolve attribution vs GDPR conflicts through legal clarification or jurisprudence, enabling fairness testing and copyright compliance. 4.2.2 Industry-Level Policy: Standards and Collaboration •Codes of conduct for AI in music, modelled on GDPR codes. •Identifier crosswalks across ISRC, ISWC, ISNI, VIAF, etc. •Federated AI services for claims, reconciliation, multilingual enrichment. •Training and reskilling to close the AI/data talent gap. •Working capital optimisation through AI-assisted claims and faster distributions. These principles do not stand in isolation: they echo and extend ongoing work such as the Responsible AI Music framework, ensuring that sector-specific practices in Europe are consistent with emerging international standards.5 4.2.3 Organisational-Level Policy: Playbooks for CMOs, Publishers, Archives •Embed AI in workflows so metadata is generated and validated during creation/distribution. 5The Responsible AI Music framework (RAIM) sets out principles for transparency, fairness, sustainability, and accountability in the use of AI in music (Herremans, Sturm, et al. 2025). Several of the codes of conduct proposed here — such as clarity around data provenance, safeguards for attribution, and limits on exploitative recommendation practices — align closely with RAIM’s recommendations. Where RAIM defines broad principles, this Green Paper provides concrete mechanisms for their operationalisation within European music data spaces and observatories. 66 •Capture once, reuse many times, reducing redundant re-entry. •Invest in knowledge capital, not IT churn (ontologies, vocabularies, multilingual enrichment). •Subscribe to shared AI utilities instead of bespoke in-house builds. •Develop internal AI governance — even small actors can appoint an “AI steward.” 4.2.4 Curative AI and Reparative AI as a Remediation Solution While data spaces establish rules for new data flows, they do not address the legacy backlog of poorly formatted or incomplete open data. Here, curative AI provides a complementary solution. AI-assisted services can detect duplicates, infer missing identifiers, reconcile heterogeneous formats, and enrich metadata with multilingual descriptions. In effect, they transform datasets that are legally open but practically unusable into resources that can circulate across the ecosystem. ĹNote Curative AI as regeneration, not replacement Figure 4.1: ���� (Ise Grand Shrine): a wooden sanctuary in continuous use for 1,600 years thanks to regeneration practices handed down through generations. The Ise Grand Shrine in Japan has been in continuous use for 1,600 years — not because its wooden beams never rotted, but because the knowledge of renewal was embedded and transmitted across generations. The true asset was the embedded know-how of regeneration, not any single plank of wood. 67 Curative AI can play the same role in the digital domain: - Extend the life of legacy systems by fixing patchy outputs from old ERPs, catalogues, or distributor software. - Preserve the methods of repair: how to reconcile corrupted records, reshape data for new systems, and upgrade databases while remaining compatible with older formats. - Transform investment logic: instead of constant capex for new IT systems, shared data infrastructures with curative AI reduce costs, smooth opex, and deliver futureproof and past-proof services. Our pilots — such as Unlabel and SKCMDb — show that new value can be created without additional IT investment or system upgrades by the participating companies, libraries, and rights management agencies. Thus, governance and remediation are two sides of the same coin: -Data sharing spaces ensure that new data is created in interoperable ways. -Curative AI repairs the inherited stock of legacy and low-quality datasets. Together, they close the gap between the right of reuse (granted by the Open Data Directive) and the means of reuse required for music, culture, and AI-driven innovation. 4.2.5 Lowering Documentation Barriers We propose to adapt Unlabel’s approach as a model for unfreezing frozen assets. By leveraging AI-assisted metadata repair and DDEX-compliant catalogue transfer workflows, documentation costs can be reduced enough to enable non-profits, small labels, and community archives to register and redistribute neglected repertoires. Public support should subsidise onboarding costs, create standardised pipelines, and incentivise low-friction reuse of metadata across systems. 4.2.6 Observatory: European = Open When we call for a European Music Observatory, the adjective “European” should not be read as a cultural filter that limits scope to European repertoires. Music is, and always has been, global. The task of the Observatory is not to create an insular archive of “European music,” but to build a governance and data architecture rooted in European values: •Data sovereignty — ensuring that creators, communities, and institutions have meaningful control over how their metadata and works are represented. •Subsidiarity — solutions should be built at the lowest effective level, allowing national archives, collective management organisations, and industry actors to contribute without being absorbed into a single monolith. 68 •Inclusiveness — minority repertoires, independent artists, and small markets must be equally visible alongside the global catalogues of multinational platforms. Our Finno-Ugric case studies show how fragile metadata can be repaired without erasing community perspectives — a model that must be embedded at Observatory scale6. This is why we chose the name Open Music Observatory (OMO). Even if the policy framework ultimately labels it the “European Music Observatory,” the essential principle must remain openness — of infrastructure, of governance, and of participation. The Observatory should be a federated, open knowledge space, not a centralised database. Europe has an opportunity to take a step that resonates beyond its borders. The U.S. Music Industry Licensing Collective (MILC) demonstrated how a single initiative could set standards and ripple globally. An Open Music Observatory, grounded in European governance but open to the world, could play a similar role — aligning sovereignty with interoperability, and showing how collective data architectures can provide guardrails for AI in a truly global music ecosystem. 4.2.7 The Open Music Observatory as a Collective Guardrail AI will only create sustainable value for music when governance, interoperability, and human capital are aligned. But building effective guardrails for agentic and generative AI cannot be done by individual firms or even national markets. - At the business level, companies lack the scale and incentives to police AI use of metadata. - At the industry level, cooperation is necessary but often fragmented by competing interests. This is where the European Union can play a decisive role: - Coordinating and aligning existing investments in Europeana, the European Collaborative Cloud for Cultural Heritage (ECCCH), and the new data sharing spaces. - Anchoring these initiatives in an Open Music Observatory (OMO) built around federated, Wikibase-compatible knowledge graphs. - Ensuring that metadata repair and publication feed into collective data architectures that double as guardrails — improving attribution and interoperability while reducing the risk of generative AI misuse. ĹWikidata Embedding Project: An Open Model for AI Guardrails In 2024–25, Wikimedia Deutschland, in collaboration with Jina.AI and DataStax, launched the Wikidata Embedding Project. - Its goal is to add vector-based semantic search to Wikidata, combining its 6We have created the second federated module of the Open Music Observatory with contemporary popular and authentic folk music of European Finno-Ugric minorities who do not have a nation state. (Antal et al. 2025) 69 multilingual knowledge graph with modern embedding models. - This enables context-aware retrieval for AI systems while anchoring results in a public, verifiable knowledge base. Why it matters for music policy - Shows that guardrails for generative AI can be built on open, communitymanaged graphs rather than proprietary black boxes. - Demonstrates how semantic search and retrieval-augmented generation can: - Reduce hallucinations by grounding outputs in human-verified data. - Combat misinformation with verifiable references. - Amplify underrepresented knowledge by balancing global visibility. Implication for the Open Music Observatory (OMO) - By adopting Wikibase-compatible knowledge graphs and existing ontological patterns, OMO can build similar guardrails for music. - This positions metadata repair and publication not just as technical fixes, but as part of a collective data architecture that keeps AI accountable. The OMO model would provide: -Compass and coordination at the EU level. -Standards and shared utilities through industry cooperation. -Flexible governance and playbooks for organisations. With this architecture, AI becomes an infrastructure for continuous renewal: prolonging legacy systems, unfreezing frozen assets, and supporting both heritage and new repertoires — while embedding guardrails against substitution and misappropriation into the very data fabric of Europe’s music ecosystem. 70 5 What Europe Should Do Next for Music Data & AI Europe’s music ecosystem is under pressure. Streaming pays in micro-royalties, metadata mistakes cost real money, and AI threatens to overwhelm platforms with untracked content. But solutions are within reach. This Green Paper sets out a path forward, built on three pillars: better metadata, shared data spaces, and AI that works for everyone. (See Chapter 1for the background and policy context.) The first step is to fix metadata at the source. Rights societies, platforms, labels, libraries, and archives all capture fragments of information about works and recordings. Today this is done in parallel, wasting effort and creating errors. Smarter pipelines, shared identifiers, and pragmatic exchange patterns can make documentation “capture once, reuse many.” This is not just a technical issue — it’s the foundation for fair royalties and cultural visibility. See Chapter 3for how shared infrastructures can make this possible. The second step is to build federated data sharing spaces. Instead of a single giant database, Europe should connect what already exists: collective management systems, heritage archives, and platform catalogues. Each actor stays in control of its own data but agrees to shared profiles, identifiers, and rules. This approach lowers costs, improves trust, and makes cross-border reuse realistic. The Open Music Observatory is our proposal for such a space: not a central repository, but a convening layer that makes decentralisation work. See Chapter 4for how artificial intelligence can be used to strengthen, not weaken, this foundation. The third step is to treat AI as a shared utility. Big platforms already use AI to document millions of tracks and to steer attention. Smaller players cannot compete unless Europe provides common tools: AI to reconcile identifiers, repair legacy datasets, enrich metadata in multiple languages, and help creators embed information “from birth.” If deployed in a federated way, AI reduces costs and unfreezes neglected repertoires — while respecting rights, attribution, and diversity. Taken together, these steps close the gap between the right of reuse granted by the Open Data Directive and the means of reuse that the music industry actually needs. Europe should: • Support metadata capture and cross-domain identifiers. 71 • Invest in federated data sharing spaces like the Open Music Observatory. • Provide pooled AI services that SMEs, CMOs, and archives can all use. This is how we make Europe’s music ecosystem more fair, efficient, and future-proof — for creators, for industry, and for audiences alike. 72 Sources & Further Reading Abing, Hannah. 2024. “Report Finds More Music Is Released in a Day in 2024 Than in All of 1989 Combined.” November 20. https://www.rareformaudio.com/blog/moremusic-released-in-a-day-2024-than-1989. Albanese, Davide, Matthias Balliauw, Alexis Macedo-Rouet, et al. 2023. “ChoCo: An Ontology and Knowledge Graph for Musical Chords.” Scientific Data 10 (1): 91. https: //doi.org/10.1038/s41597-023-02410-w. Antal, Daniel. 2019. Slovak Music Industry Report [Správa o slovenskom hudobnom priemysle]. 2019. https://doi.org/10.17605/OSF.IO/V3BE9. ———. 2020. Central And Eastern European Music Industry Report 2020. CEEMID, Consolidated Independent. https://doi.org/10.13140/RG.2.2.21450.31686. ———. 2021. An Empirical Analysis of Music Streaming Revenues and Their Distribution. Version 1.0. Zenodo. https://doi.org/10.5281/zenodo.5554089. ———. 2022. The Music Industry Value Chain. Figshare. https://doi.org/10.6084/m9. figshare.19174310.v1. ———. 2023. Pilot Program for Novel Music Industry Statistical Indicators in the Slovak Republic.https://doi.org/10.5281/zenodo.10372026. ———. 2024a. “A szlovák adatkicserélési tér magyarországi föderációjának lehetőségei.” In Az oktatás, a kutatás és a közgyűjtemények digitális transzformációja felsőfokon : NETWORKSHOP 2024 : 33. Országos Informatikai Konferencia : 2024. április 3–5. Eszterházy Károly Katolikus Egyetem, Eger, edited by József Tick, Károly Kokas, and András Holl, 192–98. Budapest: HUNGARNET Egyesület. https://doi.org/10.31915/ NWS.2024.25. ———. 2024b. Building a Music Data Sharing Space with Wikibase. Version 1.0. Open Music Observatory. https://doi.org/10.5281/zenodo.17078911. ———. 2024c. Open Music Observatory. Version 1.1. Digital Music Observatory. https: //doi.org/10.5281/zenodo.16539570. ———. 2024d. Trustworthy AI and Data-Sharing Spaces for the Slovak Music Centre. Poster Presentation at the IAMIC Conference 2024 on November 21, 2024, at Music 73 ness Analysis of the Media and Content Industries. The Music Industry. 25277 EN. Edited by Jean Paul Simon. Luxembourg: Publications Office of the European Union, 2012: Joint Research Centre Institute for Prospective Technological Studies (IPTS). http://ftp.jrc.es/EURdoc/JRC69816.pdf. Magnus, Bart, and Olivier Van D’huynslager. 2021. “Podiumkunstendata Op Wikidata: De Stap Naar Echte Linked Open Data.” Kunstenpunt / Flanders Arts Institute, February 11. https://www.kunsten.be/nu-in-de-kunsten/podiumkunstendata-op-wikidatade-stap-naar-echte-linked-open-data/. Mechanical Licensing Collective. 2021. 2021 Annual Report.https://www.themlc.com/ annual-report-2021. Mikš, Tomáš. 2025. OpenMusE: Towards a Sustainable Licensing Market for AI Use of Protected Works. April 29–30, 2025. Vilnius, Lithuania: Slovak Performing; Mechanical Rights Society (SOZA); OpenMusE Consortium; Presentation at the CISAC European Committee Meeting. https://www.openmuse.eu/wp-content/uploads/2025/ 05/20250427_CISAC_EC_2025_OpenMusE_final.pdf. Mikš, Tomáš, and Dániel Antal. 2025. Open Access Music Dataspaces – Open Music Observatory. Open Music Observatory. https://doi.org/10.5281/zenodo.17669739. Milosic, Klementina. 2015. “The Failure of the Global Repertoire Database (GRD).” Hypebot, August 2015. https://www.hypebot.com/hypebot/2015/08/the-failure-of-theglobal-repertoire-database-effort-draft.html. Ministerstvo kultúry SR, and Open Music Europe. 2023. Memorandum o porozumení o využití výsledkov analýz otvorených politík v kontexte slovenského kultúrneho a kreatívneho priemyslu a sektorových verejných politík v spolupráci s konzorciom pre výskum a inovácie s názvom OpenMuse. [Memorandum of Understanding on utilizing the Open Policy Analysis results of the OpenMuse Research and Innovation Consortium in the context of Slovak cultural and creative industries and sectors’ public policies]. https://www.crz.gov.sk/zmluva/7645338/. MIT Sloan School of Management. 2025. Project NANDA: Enterprise Generative AI Value Creation. Massachusetts Institute of Technology. https://mitsloan.mit.edu. Music Moves Europe. 2024. Music Ecosystem 2025: Study on the Music Ecosystem. Publications Office of the European Union. Luxembourg: European Commission, Directorate-General for Education, Youth, Sport; Culture. https://doi.org/10.2766/ 95340. Nakos, Basil, and Lazaros Tsoulos. 2022. “Web-Based Nautical Charts Automated Compilation from Open Hydrospatial Data.” Journal of Navigation 75 (6): 1–19. https: //doi.org/10.1017/S0373463322000489. 80 Open Music Europe Consortium. 2025. Policy Brief: An Open, Scalable Data-to-Policy Pipeline for European Music Ecosystems. EU Horizon Europe Deliverable D5.7. Open Music Europe Consortium. https://openmuse.eu/. OpenRefine Community. 2021. OpenRefine Reconciliation API Standard.https: //reconciliation-api.github.io/specs/latest/. Partanen, Niko, Philippe Rixhon, Karīna Bandere, Jānis Ziediņš, Pawan Kumar Dutt, Matīss Bolšteins, Matias Frosterus, et al. 2025. Interoperable, Trustworthy, and Machine-Readable Copyright Data in the AI Era: Report of the CITF First Project. Research report. Helsinki: Ministry of Education; Culture. https://urn.fi/URN:ISBN: 978-952-415-143-6. Paskin, Norman. 2006. “Identifier Interoperability: A Report on Two Recent ISO Activities.” D-Lib Magazine 12 (4): 1–20. https://doi.org/10.1045/april2006-paskin. Pomerantz, Jeffrey. 2015. Metadata. The MIT Press Essential Knowledge Series. Cambridge, MA, USA: MIT Press. Project, CEDAR. 2023. “A Hitchhiker’s Guide to High Value Datasets.” https://cedarheu-project.eu/articles/hitchikers-guide-high-value-datasets. PRS for Music. 2023. “PRS for Music Expands Pioneering Nexus Programme.” September 6. https://www.prsformusic.com/press/2023/prs-for-music-expands-pioneeringnexus-programme. PwC. 2023. Digital IQ 2023: Driving ROI on Digital Investments. PricewaterhouseCoopers International Limited. https://www.pwc.com/gx/en/industries/technology/ digital-iq-survey.html. ———. 2024. 27th Annual Global CEO Survey. PricewaterhouseCoopers International Limited. https://www.pwc.com/gx/en/ceo-agenda/ceosurvey/2024.html. Quine, Willard Van Orman. 1968. “Ontological Relativity.” The Journal of Philosophy 65 (7): 185–212. https://doi.org/10.2307/2024305. Sardo, Lucia, and Carlo Bianchini. 2022. “Wikidata: A New Perspective Towards Universal Bibliographic Control.” JLIS.it : Italian Journal of Library and Information Science 13 (1): 291–311. https://doi.org/10.4403/jlis.it-12725. Schnurr, Daniel. 2021. Open Government Data in Digital Markets: Effects on Innovation, Competition and Societal Benefits. SSRN Working Paper. https://papers.ssrn.com/ sol3/Delivery.cfm?abstractid=3743648. SEMIC Support Centre. 2023. Wikidata and Wikibase — SEMIC Support Centre. https://interoperable-europe.ec.europa.eu/collection/semic-support-centre/wikidata81 and-wikibase. Senftleben, Martin, Thomas Margoni, Joost Poort, Kacper Szkalej, and Etienne Valk. 2024. Policy Brief 1: Music Metadata Mainstreaming and EU Law. EU Horizon Europe Deliverable D5.6. OpenMusE Consortium. https://www.openmuse.eu/. Simoni, Marco U., Kristin A. Aasly, and Frode Schjøth. 2021. MINERAL Intelligence for Europe (Mintell4EU) – Case Study Overview. GeoERA. https://geoera.eu/wpcontent/uploads/2021/10/D4.1-Mintell4EU-Case-Study-Overview.pdf. Stallmann, Claudia, Koen Deneckere, Ruben Verborgh, et al. 2023. “MetaBelgica Project: A Linked Data Infrastructure Between Federal Scientific Institutes in Belgium.” Proceedings of the 19th Extended Semantic Web Conference (ESWC 2023) (Cham), 2023. https://doi.org/10.1007/978-3-031-33455-9_24. Teosto. 2024. “New ISNI Identifier Creates Better Opportunities for International Author Identification.” March 12. https://www.teosto.fi/en/new-isni-identifier-creates-betteropportunities-for-international-author-identification/. Varghese, Jacob. 2024. “Beyond the Metadata: How Can We Solve the Black Box Royalty Mystery?” Noctil, August 14. https://independentmusicinsider.com/editorial-articles/ 4820/. Virág, Barnabás. 2024. “Our Library’s Music Collection in the Era of Streaming Services.” Canadian Journal of Information and Library Science / La Revue Canadienne Des Sciences de l’information Et de Bibliothéconomie 47 (2): 175–87. https://doi.org/10. 5206/cjils-rcsib.v47i2.17436. W3C. 2013a. PROV-o: The PROV Ontology. Edited by Satya AND McGuinness Lebo Timothy AND Sahoo. W3C. https://www.w3.org/TR/prov-o/. ———. 2013b. PROV-Overview: An Overview of the PROV Family of Documents. Edited by Paolo Moreau Luc AND Missier. W3C. https://www.w3.org/TR/prov-overview/. World Intellectual Property Organization (WIPO). 2023. “Project Nexus: A Data Matching Project of PRS for Music.” April 19. https://www.wipo.int/edocs/mdocs/mdocs/ en/wipo_webinar_cr_2023_6/wipo_webinar_cr_2023_6_presentation.pdf. 82