Full text
Open Music Observatory Building an open data sharing space for the European music sector Daniel Antal, CFA Mester, Anna Márta 2025-11-30
Table of contents Open Music Observatory 7 DisclaimerofWarranties................................. 7 Glossary 9 Musicterms........................................ 9 Creatorsofmusicalworks ............................. 10 Datascienceterms.................................... 11 Dataprotectionterms .................................. 15 Data curation and collection terms . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 Statisticalterms ..................................... 16 Registers, authorities, standards and identifiers . . . . . . . . . . . . . . . . . . . . 17 Organisations....................................... 20 Otherabbreviations ................................... 21 Executive Summary 23 1 Introduction 26 2 Background & Concept 30 2.1 Why Europe Needs a Music Observatory . . . . . . . . . . . . . . . . . . . . . 31 2.2 Historical Precedent: CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2.3 Policy and Technological Evolution Enabling a New Observatory . . . . . . . 33 2.3.1 European Parliament and EU-level Mandates . . . . . . . . . . . . . . 33 2.3.2 Data (Sharing) Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 2.3.3 Preference for Open-Source and Open Standards in the EU . . . . . . 35 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures . . . 35 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy 36 2.3.6 European Interoperability Framework (EIF) . . . . . . . . . . . . . . . 36 2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds . . . 37 2.3.8 Summary .................................. 37 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model . . . . . . . 37 2.4.1 Lessons from CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 37 2.4.2 Requirements of the EU policy environment . . . . . . . . . . . . . . . 38 2.4.3 Requirements of the Grant Agreement . . . . . . . . . . . . . . . . . . 38 2.4.4 Technical rationale for decentralisation . . . . . . . . . . . . . . . . . . 39 2.4.5 Why decentralisation is essential for a European Music Observatory . 39 2.5 The Data-to-Policy Pipeline: How Open Music Europe Works . . . . . . . . . 39 2.5.1 1. Indicator and problem definition (WP1–WP3) . . . . . . . . . . . . 40 2.5.2 2. Data governance (WP1–WP3, WP6) . . . . . . . . . . . . . . . . . 40 2
2.5.3 3. Software for data collection (WP4) . . . . . . . . . . . . . . . . . . 40 2.5.4 4. Data acquisition (WP1, WP2, WP3) . . . . . . . . . . . . . . . . . 41 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) . . . . . . 41 2.5.6 6. Validation for analysis and dissemination (WP4, WP5) . . . . . . . 41 2.5.7 7. Analysis and modelling (WP1–WP3) . . . . . . . . . . . . . . . . . 42 2.5.8 8. Policy translation (WP5) . . . . . . . . . . . . . . . . . . . . . . . . 42 2.5.9 9. Dissemination and reuse (WP5) . . . . . . . . . . . . . . . . . . . . 42 2.5.10Summary .................................. 42 2.6 Stakeholder Engagement and Early Feedback . . . . . . . . . . . . . . . . . . 43 3 Core Services 44 3.1 Collect: Data Curation & Collection . . . . . . . . . . . . . . . . . . . . . . . 47 3.1.1 Microdata, Collections, Records . . . . . . . . . . . . . . . . . . . . . . 49 3.1.2 Primary data collection . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3.1.3 Metadata .................................. 51 3.1.4 Statistical indicators and datasets . . . . . . . . . . . . . . . . . . . . 52 3.2 Repair........................................ 52 3.3 Process ....................................... 52 3.3.1 Processing & re-processing microdata . . . . . . . . . . . . . . . . . . 53 3.3.2 Documentation............................... 54 3.4 Disseminate..................................... 55 3.4.1 Open Music Observatory . . . . . . . . . . . . . . . . . . . . . . . . . 55 3.4.2 EU Open Data Portal . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 3.4.3 Europeana Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 3.4.4 European Collaborative Cloud for Cultural Heritage . . . . . . . . . . 59 3.4.5 European Open Science Cloud . . . . . . . . . . . . . . . . . . . . . . 60 3.5 Metadata ...................................... 62 3.5.1 Wikibase & Wikidata . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 3.5.2 Music Observatory Website . . . . . . . . . . . . . . . . . . . . . . . . 63 3.5.3 APIEndpoint................................ 63 4 Architecture 64 4.1 Why Wikibase Is the Right Foundation for the Open Music Observatory . . . 64 4.1.1 Proven in real-world scenarios highly similar to music . . . . . . . . . 65 4.1.2 Already aligned with Europe’s digital knowledge infrastructure . . . . 66 4.1.3 Demonstrated support for required OMO functionality . . . . . . . . . 66 4.1.4 Fits EU policy preference for open-source and trustworthy AI . . . . . 66 4.1.5 The most widely used graph-editing interface in the world . . . . . . . 67 4.1.6 A hybrid model that fits real institutional workflows . . . . . . . . . . 67 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline . . . 67 4.2.1 Wikibase supports each stage of the pipeline . . . . . . . . . . . . . . 68 4.3 Summary ...................................... 69 5 Data coordination 70 5.0.1 The European Interoperability Framework (EIF) . . . . . . . . . . . . 71 5.0.2 Extending the EIF to public and private service coordination . . . . . 73 3
5.0.3 Datasharingspace............................. 74 5.1 Ontologies and Vocabularies in the Open Music Observatory . . . . . . . . . 75 5.1.1 Lightweight Ontology Patterns . . . . . . . . . . . . . . . . . . . . . . 77 5.2 Multiple roles, multiple workflows to support . . . . . . . . . . . . . . . . . . 78 5.2.1 Polyhierarchy................................ 80 5.2.2 Formalisation................................ 82 5.3 Future-Proofing................................... 83 5.3.1 Future-proofing through graph architecture . . . . . . . . . . . . . . . 83 5.3.2 Stabilising definitions through internationally defined standard vocabularies.................................. 84 5.3.3 Future-proofing through translatability and multiple serialisations . . 84 5.3.4 Future services through institutional interoperability . . . . . . . . . . 85 5.3.5 A concrete example: ALOADED, Livonian folk music, and DDEX . . 85 6 Federated Data Modules 88 6.1 Slovak Comprehensive Music Database (SKCMDb) . . . . . . . . . . . . . . . 89 6.1.1 Purposeandscope............................. 89 6.1.2 Datainputs................................. 90 6.1.3 Metadata and semantic alignment . . . . . . . . . . . . . . . . . . . . 91 6.1.4 Governance and legal basis . . . . . . . . . . . . . . . . . . . . . . . . 92 6.1.5 Interoperability and federation . . . . . . . . . . . . . . . . . . . . . . 93 6.1.6 Status and next steps . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 6.2 Hungarian Music Database (HUMDb) . . . . . . . . . . . . . . . . . . . . . . 93 6.2.1 Purposeandscope............................. 94 6.2.2 Datainputs................................. 95 6.2.3 Metadata and semantic alignment . . . . . . . . . . . . . . . . . . . . 95 6.2.4 Governance and legal basis . . . . . . . . . . . . . . . . . . . . . . . . 96 6.2.5 Interoperability and federation . . . . . . . . . . . . . . . . . . . . . . 96 6.2.6 Status and next steps . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 6.3 Finno-Ugric Data Sharing Space . . . . . . . . . . . . . . . . . . . . . . . . . 97 6.3.1 Purposeandscope............................. 97 6.3.2 Datainputs................................. 97 6.3.3 Metadata and semantic alignment . . . . . . . . . . . . . . . . . . . . 98 6.3.4 Governance and legal basis . . . . . . . . . . . . . . . . . . . . . . . . 98 6.3.5 Interoperability and federation . . . . . . . . . . . . . . . . . . . . . . 99 6.3.6 Status and next steps . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 6.4 Open Music Observatory Core Module . . . . . . . . . . . . . . . . . . . . . . 99 6.4.1 Purposeandscope............................. 99 6.4.2 Data inputs (WP1–WP4 contributions) . . . . . . . . . . . . . . . . . 100 6.4.3 Metadata and semantic alignment . . . . . . . . . . . . . . . . . . . . 100 6.4.4 Governance and legal basis . . . . . . . . . . . . . . . . . . . . . . . . 100 6.4.5 Interoperability and federation . . . . . . . . . . . . . . . . . . . . . . 101 6.4.6 Status and next steps . . . . . . . . . . . . . . . . . . . . . . . . . . . 101 6.5 Summary and Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101 4
7 Data Collection 102 7.1 Overview of the Data-Collection Framework . . . . . . . . . . . . . . . . . . . 102 7.2 Administrative and Register Data . . . . . . . . . . . . . . . . . . . . . . . . . 103 7.3 SurveyData.....................................103 7.4 Statistical and Economic Data . . . . . . . . . . . . . . . . . . . . . . . . . . 104 7.5 Platform and Streaming Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 104 7.6 Processing and Harmonisation . . . . . . . . . . . . . . . . . . . . . . . . . . . 104 7.7 Software Components Developed in WP4 . . . . . . . . . . . . . . . . . . . . 105 7.7.1 Data-Ingestion Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 7.7.2 Validation and Reconciliation Tools . . . . . . . . . . . . . . . . . . . 105 7.7.3 Harmonisation and Metadata Tools . . . . . . . . . . . . . . . . . . . 105 7.7.4 OMO Integration Tools . . . . . . . . . . . . . . . . . . . . . . . . . . 106 7.8 Position of Data Collection and Processing in the Pipeline . . . . . . . . . . . 106 7.9 Integration with the Open Music Observatory . . . . . . . . . . . . . . . . . . 106 8 Standardisation of Data & Terminology 107 8.1 Businessprocesses .................................107 8.2 Conceptual and information models . . . . . . . . . . . . . . . . . . . . . . . 108 8.3 Identification & Entity Linking . . . . . . . . . . . . . . . . . . . . . . . . . . 110 8.3.1 Registers & Authority Files . . . . . . . . . . . . . . . . . . . . . . . . 111 8.3.2 Open and persistent identifiers . . . . . . . . . . . . . . . . . . . . . . 112 8.3.3 Not open, music-industry specific identifiers . . . . . . . . . . . . . . . 113 8.3.4 Lyrics ....................................114 8.3.5 ISCC ....................................114 8.3.6 OMOIdentifiers ..............................115 9 Data Improvement & Innovation 117 9.1 Value-Added Data Services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 9.1.1 DataSharing................................117 9.1.2 Fix-the-data ................................118 9.1.3 DataLinking ................................118 9.1.4 Registration services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 9.2 UseCases......................................120 9.2.1 Data Health Services for Collective Management . . . . . . . . . . . . 121 9.2.2 Sustainability Reporting for Music Organisations . . . . . . . . . . . . 122 9.2.3 ListenLocal.................................124 9.2.4 Unlabel ...................................125 9.3 UseofAIsystems .................................126 10 Data Catalogue 129 10.1CollectionGuidelines................................131 10.2TopicalPillars ...................................132 10.2.1MusicEconomy...............................133 10.2.2MusicDiversity...............................135 10.2.3MusicSociety................................136 10.2.4Innovation..................................137 5
10.2.5Sustainability................................137 References 138 Appendices 144 Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network 144 ...............................................144 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 SKCMDb: Slovak Comprehensive Music Database 146 The Slovak Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 Slovak Comprehensive Music Database (public) . . . . . . . . . . . . . . . . . . . . 147 Slovak Comprehensive Music Database (private) . . . . . . . . . . . . . . . . . . . 148 Microdata......................................148 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 149 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 LīvMDb: Livonian Music Database 151 The Livonian Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152 Livonian Music Database (public) . . . . . . . . . . . . . . . . . . . . . . . . . . . 152 Livonian Music Database (private) . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 Microdata......................................153 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 154 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154 6
Open Music Observatory ¾Partly updated This document was presented as a planning document in 2023. It has been partly updated till 18 November 2025. Some text is still reflecting the planning phase. You can access all versions on https://zenodo.org/records/16539570 Disclaimer of Warranties This project has received funding from the European Union’s Horizon Europe, research and innovation programme, under Grant Agreement No. 101095295. This document has been prepared by Open Music Europe (OpenMusE) project partners as an account of work carried out within the framework of this contract. Any dissemination of results must indicate that it reflects only the author’s view and that the Commission Agency is not responsible for any use that may be made of the information it contains. Neither Project Coordinator, nor any signatory party of Open Music Europe (OpenMusE) Project Consortium Agreement, nor any person acting on behalf of any of them: (a) makes any warranty or representation whatsoever, express or implied, (i). with respect to the use of any information, apparatus, method, process, or similar item disclosed in this document, including merchantability and fitness for a particular purpose, or (ii). that such use does not infringe on or interfere with privately owned rights, including any party’s intellectual property, or 7
(iii). that this document is suitable to any particular user’s circumstance; or (b) assumes responsibility for any damages or other liability whatsoever (including any consequential damages, even if Project Coordinator or any representative of a signatory party of the Open Music Europe (OpenMusE) Project Consortium Agreement, has been advised of the possibility of such damages) resulting from your selection or use of this document or any information, apparatus, method, process, or similar item disclosed in this document. For the version history of this document, please refer to our open repository, where the change history can be reviewed with timestamps for every single file used to create the report: https://github.com/dataobservatory-eu/open-music-observatory This document is accompanied by the maturing A Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem. Practical Steps Towards a Decentralised and Open European Music Observatory document, which show’s our works datato-policy alignment. The latest development version of that document can be downloaded from the green paper’s website, and the latest identified version following the OPA protocol can be found on GitHub and on Zenodo. The temporary landing page of the OMO can be reviewed on https://dataobservatoryeu.github.io/omo-landing-page/. 8
Glossary Music terms audio recording: fixation of sounds (ISO 2019b) music video recording: fixation of sounds synchronized with pictures or moving pictures where (a) the fixed sounds are wholly or substantially a musical performance or (b) the recording is intended for viewing in association with a recording of a musical performance. This definition includes music videos and concert recordings, together with music-related interviews and documentaries, but does not extend to genera! audiovisual material, even if it includes music.(ISO 2019b) recording: result of a recording process independent of the type and number of audio or audiovisual carriers and technology used Note 1 to entry: The term “recording” applies to each recorded item which may be used as a separate unit regardless of whether it is issued as part of a larger recorded work (e.g. each separate track on an album of audio recordings). [SOURCE:ISO 3901:2001, definition 3.3] (ISO 2017b) work: distinct, abstract creation of the mind whose existence is revealed through one or more expressions (e.g. a performance) or manifestations (e.g. an object) (ISO 2022) musical work: composed of a combination of sounds, with or without accompanying text (ISO 2022) DSP or digital streaming platform: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. DSPs can provide access to music downloads, like Apple’s iTunes Store, or access to streaming music like Spotify, or even provide satellite-delivered content such as SiriusXM in the USA. rights management (organisations): the function of managing the rights on behalf of rights owners. It can be companies whose sole purpose is to ensure that content that has been licensed has delivered royalties that are identified and accounted for. The role can be taken by collective management organisations or by private companies on behalf of songwriters, composers, performers, music publishers, or record labels. duration: the elapsed playing time between the first and last recorded modulations of the recording. LP or Long Player: gramophone record usually on both sides comprising one or more sound recordings with a playing time of each side of normally round about 30 minutes and released and sold on its own (ISO 2017b) 9
anthology: document consisting of a collection of full documents or of extracts, usually of literary works (ISO 2017b) exhibition: curated display of objects on a clear concept and communicating a message [SOURCE:ISO 18461:2016, definition 2.4.6 modified] (ISO 2017b) curator: person responsible for overseeing a collection or exhibition (ISO 2017b) data curation: managed process, throughout the data lifecycle, by which data/data collections are cleansed, documented, standardized, formatted and interrelated (ISO 2017b) register: an official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country; a document, usually a volume, in which data are entered in a formal manner by a statutory authority Note 1 to entry: In modern usage, usually a database. (ISO 2017b) registration: act of giving an entity a unique identifier on its entry into a system (ISO 2017b) set of rules, operations, and procedures for inclusion of an item in a registry (ISO 2023a) registrant: organization or person that has either registered an authentication protocol or registered the adoption of an authentication protocol [SOURCE: ISO/IEC 24727-6:2010, definition 3.4] (ISO 2017b); an entity wishing to assign an ISRC to an applicable recording (ISO 2019b); aparty that requests an ISNI from the Registration Authority (ISNI 3.2 (ISO 2012, p15)) party: natural person or legal person, whether or not incorporated, or a group of either (ISO 2012) aggregation: acquisition of sensitive information by collecting and correlating information of lesser sensitivity (ISO 2023b) Statistical terms administrative records: data generated by a non‐statistical source, usually a public body, the main aim of which is not the provision of statistics. code list: predefined list from which some statistical coded concepts take their values (ISO 2013) data pipeline: a method in which raw data is ingested from various data sources and then ported to data store. FAIR or FAIR Guiding Principles for scientific data management and stewardship: guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital 16
assets, emphasising machine-actionability (i.e., the capacity of computational systems to find, access, interoperate, and reuse data with none or minimal human intervention.) indicator: the representation of statistical data for a specified time, place or any other relevant characteristic, corrected for at least one dimension (usually size) so as to allow for meaningful comparison. microdata: non‐aggregated observations or measurements of characteristics of individual units, without direct identifier. MVP or minimum viable product: a version of a work product with just enough features and requirements to satisfy early customers and/or provide feedback for future development [SOURCE:IEEE 2675-2021, 3.1] observation unit: an identifiable entity about which data can be obtained, it is also often called a statistical unit or data subject in case of a natural person. Open Policy Analysis Guidelines: a set of information management rules to make policy analysis more transparent. personal data: any information relating to an identified or identifiable natural person. pseudonymisation: processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information. survey: a systematic examination and record of a physical or social area and its features so as to construct a map, plan, or description. In social sciences it usually refers to a well-structured questionnaire and answers given to its items by a target population. statistics: quantitative and qualitative, aggregated and representative information characterising a collective phenomenon in a considered population. visualisations: schematic charts, drawings, photographs, and their collages will as still image files that help to explain the relationship between information carriers, data points, or processes. Registers, authorities, standards and identifiers IČO: The organisation identification number (IČO) is an identifier assigned to all types of legal entities, entrepreneurs and public authorities by the Statistical Office of the Slovak Republic. The Czech Republic’s organisation identifier is also called IČO. OpenCorporates: a public corporation database which sources data from national business registries. ISNI: an ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. VIAF: The Virtual International Authority File (VIAF) is an international service that consolidates multiple name authority files into a single database. Their primary goal is 17
to enhance the efficiency and usability of library authority files by linking and merging widely used authority records and making them accessible online. VIAF ID: The VIAF (Virtual International Authority File) combines multiple name authority files into a single OCLC-hosted name authority service. ISRC: The International Standard Recording Code (ISRC) is a standard identifying code that can be used to identify sound recordings and music video recordings so that each such recording can be referred to uniquely and unambiguously. ISWC: The purpose in creating an ISWC for musical works is to enable more efficient administration of rights to those works on a worldwide basis. The ISWC provides an efficient means of identifying musical works in computer databases and related documentation and for the exchange of information between rights societies, publishers, record companies and other interested parties on an international level. ISBN: the International Standard Book Number is an identification system for the publishing industry and its supply chains. ISMN: The International standard music number (ISMN) was developed by, and for, the music publishing sector as a separate system to complement the International standard book number (ISBN). The existence of the ISMN as a separate identifier system makes it possible to identify printed and notated music as a distinct category of publication within the global supply chain and to develop trade directories and similar services for the specialized market for music publications. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. DOI: The Digital Object Identifier is a standardised unique number given to many (but not all) articles, papers and books, by some publishers, to identify a particular publication. ORCID: the Open Researcher and Contributor ID is a unique, persistent identifier free of charge to researchers. URI: A Uniform Resource Identifier (URI) is a string of characters used to identify a resource on the internet. This resource can be either abstract or physical, such as a website, an email address, or a file. URIs are essential for enabling interactions with resources over a network using specific protocols. W3C: The World Wide Web Consortium (W3C) is an international community that develops standards for the World Wide Web. Their mission is to lead the Web to its full potential by creating technical specifications and guidelines that are designed to be open and royalty-free. These standards include HTML, CSS, and other web technologies, which ensure that web content is accessible across different browsers and devices. DDI: The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. Wikibase: Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, 18
but it is available now for the creation of private, or public-private partnership knowledge graphs. It is developed by Wikimedia Deutschland. GSBPM: The Generic Statistical Business Process Model is a international standard model that “describes and defines the set of business processes needed to produce official statistics.” | GSIM:Generic Statistical Information Model: a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation standards such as SDMX or DDI. SDMX: Statistical Data and Metadata eXchange (SDMX), is an international initiative that aims at standardising and modernising (“industrialising”) the mechanisms and processes for the exchange of statistical data and metadata among international organisations and their member countries. ESRS: The European Sustainability Reporting Standards (ESRS) are a set of guidelines developed by the European Financial Reporting Advisory Group (EFRAG) to standardise sustainability reporting across the European Union. These standards are designed to align with the Corporate Sustainability Reporting Directive (CSRD), which mandates detailed corporate reporting on environmental, social, and governance (ESG) issues for many companies operating within the EU. CIDOC-CRM: The conceptual model of CIDOC, the standard conceptualisation of collection management systems in heritage organisations. RiC:Records in Context is a new conceptual model that replaces the four most important international archiving standards. DCTERMS or DCMI: the Dublin Core Metadata Terms is a vocabulary of metadata terms developed and maintained by the Dublin Core Metadata Initiative (DCMI). These terms are used to describe various aspects of digital resources, such as web pages, documents, and other online content. They provide a standardized way to assign metadata to resources, making them easier to discover, manage, and exchange. RDFS: the Resource Description Framework Schema is an extension of the Resource Description Framework (RDF) that provides a vocabulary for describing classes and properties of resources within an RDF graph. EDM: the Europeana Data Model is a framework for collecting, connecting, and enriching cultural heritage metadata. It’s designed to facilitate the sharing and reuse of cultural heritage information by providing a standardized way to represent and link data. Europeana: a digital platform provided by the European Union that aggregates digitized cultural heritage from institutions across Europe. ESCO: the European Skills, Competences, Qualifications and Occupations classification is is a multilingual classification system developed by the European Commission to standardize the description of skills, competences, and qualifications relevant to the European labor market and education. 19
NACE: the European Union’s standard classification of economic activities for statistical purposes. The abbreviation stands for Nomenclature statistique des Activités économiques dans la Communauté européenne. ISCO: the International Standard Classification of Occupations (ISCO) is the International Labour Organization’s standardized system for classifying and organizing occupations according to jobs’ tasks and duties ISIC: the International Standard Industrial Classification of All Economic Activities (ISIC) is a standard classification system developed by the UN Statistics Division (UNSD) to categorize economic activities. PROV-O: the Provenance ontology is a formal ontology developed by W3C to represent and interchange provenance information. MARC: MAchine-Readable Cataloging, is a standard digital format used by libraries to represent and exchange bibliographic information. DCAT: an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. Organisations AEPO-ARTIS: Organisation representing European artists-performers. Regroups most of the European CMO representing performers. ALOADED: is a company which distributes and exploits recordings. CISAC: The International Confederation of Societies of Authors and Composers is an international non-governmental, not-for-profit organisation that aims to protect the rights and promote the interests of creators worldwide. CNM (former CNV): the Centre National de la Musique is a public organisation managing a tax on concert tickets EFRAG: The European Financial Reporting Advisory Group is a private association established in 2001 with the encouragement of the European Commission to serve the public interest. EFRAG extended its mission in 2022 following the new role assigned to EFRAG in the CSRD, providing Technical Advice to the European Commission in the form of fully prepared draft EU Sustainability Reporting Standards and/or draft amendments to these Standards. EMO: The European Music Observatory (EMO) is envisioned as a hub for collecting and analysing data on the music sector across Europe. Its primary aim is to address the current gaps and inconsistencies in music data collection, which have been a significant challenge for the sector. GESAC: GESAC comprises together 32 European authors’ societies in music, audiovisual, visual arts, literature and drama. 20
GESIS: Leibniz Institute for the Social Sciences. IAML: International Association of Music Libraries, Archives and Documentation Centres | IAMIC: International Association of Music Centres, an international network of organisations that collectively and collaboratively provides information and promotes the music of their countries or regions. ICMP: the global trade body representing the music publishing industry worldwide. SCAPR: International association for the development of the practical cooperation between performers’ collective management organisations (CMOs) SOZA: SOZA (Slovenský ochranný zväz autorský pre práva k hudobným dielam, Slovak Performing and Mechanical Rights Society) is a legal entity, non-profit civic association of authors and publishers of musical works, association of natural persons and legal entities. Hudobné Centrum: Music Centre Slovakia is a music organisation with a mission to promote Slovak contemporaly music. Other abbreviations CEEMID: the Central European Music Industry Databases is a multi-country project that was a predecessor of Reprex’s Digital Music Observatory CSRD: The Corporate Sustainability Reporting Directive (CSRD) is European Union (EU) legislation, effective from 5 January 2023, that requires EU businesses—including qualifying EU subsidiaries of non-EU companies—to disclose their environmental and social impacts, and how their environmental, social and governance (ESG) actions affect their business. DSP: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. EIF: The European Interoperability Framework (EIF) is a set of recommendations and guidelines that aims to facilitate communication and collaboration between public administrations, businesses, and citizens within the European Union and across national borders. ECCCH: The European Collaborative Cloud for Cultural Heritage is a European Union initiative for a digital infrastructure that will connect cultural heritage institutions and professionals across the EU. EOSC: The European Open Science Cloud (EOSC) aims to create a trusted, open, and multidisciplinary environment for researchers and innovators in Europe. PPP: A Public-Private Partnership (PPP) is a collaborative arrangement between government entities and private sector companies aimed at financing, designing, implementing, and operating projects or services traditionally provided by the public sector. RDM: Research Data Management refers to the suite of practices, policies, and processes used to handle data throughout the lifecycle of a research project. 21
Our glossary is harmonised with relevant music-sector specific standards (referred to in Chapter 8) and with the ISO Information technology — Vocabulary (ISO 2023b); Information technology — Cloud computing — Taxonomy based data handling for cloud services (ISO 2020); Information technology — Cloud computing — Interoperability and portability (ISO 2017a) and the Information and documentation — Foundation and vocabulary (ISO 2017b) and Information technology — Metadata registries (MDR) — 1. Framework (ISO 2023a) 22
Executive Summary Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. In the OpenMusE project, the development of the Open Music Observatory is coupled with a clear contractual expectation: the project must populate the Observatory’s four thematic pillars—Music Economy, Music Diversity, Music, Society & Sustainability, and Innovation & Future Trends —with initial, well-documented data and knowledge. This population process follows the project’s data-to-policy pipeline (see Chapter 1): each work package defines its indicators, establishes data governance and legal bases, collects or accesses relevant administrative, survey, statistical, and platform data, processes and harmonises them through WP4 tools, and finally activates them in reproducible analytical workflows in WP5. The result is that the Observatory is not only a technical prototype but a functional, databearing infrastructure: the first integrated demonstration of how Europe’s music data can be curated, linked, analysed, and made reusable across public, private, and civic actors. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries. An observatory has traditionally been a permanent location for observing terrestrial, marine, or celestial events. In the past 30 years, it has also been used for long-term digital data collection programs for markets, social sciences, and humanities. Our milestone requires the start of this observatory after a lengthy and intensive planning and prototyping phase. It can be seen as a modern reimagination of the data observatory model, or the observatory 2.0. We created a new observatory model that fully aligns with the European Interoperability 23
Framework but extends the governance of the digital services beyond public bodies, and allows the creation of a public-private partnership to manage the observatory. The European Interoperability Framework aims to create a four-layered approach to build digital research, marketing, rights management, collection management services for the sector. These layers are introduced in separate chapters of these documentation. 1. The technological alignment is introduced in Chapter 4; we decided to choose the technology of the world’s largest open knowledge graph, Wikidata, which already coordinates countless digital services in Europe’s cultural sector, and provides training and guardrails for many AI applications. 2. The semantic alignment is is introduced in Chapter 5. This chapter focuses on the semantic and organisational aspects of interoperability. 3. The organisational alignment means bridging actual data-driven and computer supported workflows with data semantics (what does a song’s title mean for a librarian, an ethnomusicologist, a collective rights management agency), and how they can work with various translated, alternative, historical, mistyped, and preferred titles in distributing royalties, loaning printed sheets, describing musical traditions. 4. The legal alignment creates a policy that lays out the rights, prohibitions and necessary permission processes to connect and use the data together. We give a concrete example in Section 5.3. We were informed and influenced by the creation of Europeana (which started out from a similar collaborative project, a cultural heritage oriented data sharing space) and the Commission’s new plans to extend their digital services into the European Collaborative Cloud for Cultural Heritage (ECCCH). We aimed for full interoperability with Europeana and we Reprex successfully sent data and concluded a Data Exchange Agreement. We were also aiming for interoperability with the ECCCH, which only published the first version of its Heritage Digital Twin conceptual model and ontology; we were the first to test them with music data. but we also bring a new element into their thinking. While they are mainly aggregating the work of public sector memory institutions, we are building a governance model that allows a more successful cooperation among the private sector and the public music sector. By the end of 2025, we aim to create an “observatory 3.0”, which already hosts many intelligent data improvement technologies and fuels innovative applications/services in line with our project’s exploitation pathways. These services are at different maturity levels, but they could not be brought to a testable MVP without building out the minimal digital infrastructure and governance model at this milestone. 24
ĹNote This document is licensed under the CC BY 4.0 LEGAL CODE Attribution 4.0 International license. You must refer to the document with the DOI 10.5281/zenodo.11385044. Canonical Licence URL:https://creativecommons.org/licenses/by/4.0/ Other formats:Plain Text;RDF/XMLlSee the deed 25
The study explicitly mentioned CEEMID as a promising bottom-up model for filling these gaps and providing a more modern, decentralised alternative to traditional observatory structures (Artisjus et al. 2014). The Feasibility Study also provided a clear definition of the stakeholders that a future European Music Observatory must serve. It distinguished three groups whose information needs and policy roles must be supported: •Industry: commercial organisations and agents involved in income-generating activities across performance, recording, distribution, and creation. •Civic: policymakers, NGOs, professional associations, and publicly funded intermediaries whose decisions shape the regulatory and support environment. •Public: consumers, cultural participants, education and training institutions, and third-sector organisations interested in the wider social and cultural roles of music. See (European Commission et al. 2020, p30) These stakeholder categories continue to structure the Observatory’s service model and interoperability requirements. 2.2 Historical Precedent: CEEMID The former CEEMID (originally: Central & Eastern European Music Industry Databases) collaboration began in 2014 as a voluntary, decentralised data initiative created by three collective management societies. Over time, it grew to include more than 60 stakeholders in 12 European countries. Its purpose was to fill the most pressing evidence gaps by combining: • voluntary data integration among partners, • open-data reprocessing, and • co-financed data collection. This work is documented in (Antal 2020a). CEEMID operated according to principles that would later become central to the European Union’s data (sharing) space strategy (formalised only years later). Its decentralised organisational model, distributed data stewardship, and emphasis on transparent, reusable methods demonstrated that a modern observatory in the digital era does not need to be a centralised institution. Instead, it can function as a federated ecosystem connecting statistical offices, cultural institutions, CMOs, and private actors. Long before the EU formalised its dataspace strategy, CEEMID also aligned its workflows with emerging statistical-system standards such as GSIM,DDI, and SDMX, anticipating later European requirements for interoperable, machine-readable statistical metadata. This early adoption provided a methodological bridge between cultural-sector data, administrative registers, and official statistics, and formed a direct precursor to the metadata foundations of the Open Music Observatory. 32
The Feasibility Study for the European Music Observatory explicitly recognised CEEMID as a potential building block for a new observatory model. Our proposal therefore sought to transform CEEMID’s prototype—referred to in the study as the Digital Music Observatory— into a scientifically robust and methodologically coherent system that could scale across Europe. This required grounding the work in state-of-the-art statistical science, data science, and computer science, and ensuring alignment with European interoperability and datagovernance frameworks. The prototype work that preceded Open Music Europe was shaped through two innovation environments: the Yes!Delft AI+Blockchain Lab, where product–market fit and technical feasibility were tested, and the JUMP Music Market Accelerator, where the first integrated prototype of a Digital Music Observatory was developed. These early iterations validated not only stakeholder demand but also the feasibility of a decentralised, standards-based architecture, and they informed the methodological and technical design choices taken forward in this project. 2.3 Policy and Technological Evolution Enabling a New Observatory Since the publication of the EMO Feasibility Study, the European Union has introduced a series of policy and infrastructure initiatives that strengthen the case for a decentralised, interoperable, and federated European Music Observatory. These developments span European Parliament mandates, Commission-funded research, cultural-heritage clouds, opendata regulation, and the EU’s overarching data-space strategy. Together, they establish the policy and technological foundations on which the Open Music Observatory is built. Our policy alignment is discussed in more detail in - Music Metadata Mainstreaming and EU Law -A Green Paper on AI, Data Governance, and Metadata–Policies for Europe’s Music Ecosystem2 2.3.1 European Parliament and EU-level Mandates The European Parliament, in its resolutions on the future of the music sector, explicitly called for: • the establishment of a European Music Observatory, • improved evidence for competitiveness, diversity, and fair remuneration, and • stronger coordination of public, private, and community data sources. These mandates update and reinforce both the Music Moves Europe framework and the findings of the EMO Feasibility Study. They frame the Observatory as an instrument that must serve industry, civic, and public actors through interoperable, reusable, cross-border data services. 2See (Senftleben et al. 2024); and (Antal 2025d), summarised in the internal document (Open Music Europe Consortium 2025). 33
The EU Music Ecosystem Study (2025) deepened this diagnosis, pointing to fragmentation across metadata, rights information, cultural statistics, and market data. It concluded that the sector requires a technical and governance model capable of linking these domains, rather than separate, siloed initiatives. The architecture of our dataspace responds directly to these recommendations. The European Parliament has rightly highlighted that fragmented and unreliable metadata remains a major obstacle in the music sector. European Parliament Resolution of 17 January 2024 on Cultural Diversity and the Conditions for Authors in the European Music Streaming Market 9. Emphasises that it is essential to improve the identification of anyone involved in the creation process, in particular authors and performers, on music streaming services, by ensuring the comprehensive and accurate allocation of metadata from the time of Directive 2014/26/EU of the European Parliament and of the Council of 26 February 2014 on collective management of copyright and related rights and multiterritorial licensing of rights in musical works for online use in the internal market (OJ L 84, 20.3.2014, p. 72). creation for any track uploaded to a music streaming service; encourages, in this regard, the use of all international identification codes (IPI, ISWC, ISRC, IPN, and ISNI); highlights that proper identification of creators plays a key role in the search for and discoverability of works, and enables proper remuneration for creators in the distribution of revenues. Our Observatory’s distributed model directly answers European Parliament’s call for metadata systems that are reliable, inclusive, and supportive of creators. Our policy alignment is explained in detail in our our policy paper, A Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem3. 2.3.2 Data (Sharing) Spaces The EU’s adoption of data (sharing) spaces provides the organisational and legal model for an Observatory that is not a centralised institution but a federated ecosystem. Curry defines dataspaces as: “an emerging approach to data management… Data is integrated on an ‘asneeded’ basis, with the labour-intensive aspects of data integration postponed until they are required.” (Curry 2020) The Design Principles for Data Spaces position paper further describes them as: “a federated data ecosystem within a certain application domain and based on shared policies and rules.” (Nagel and Lycklama 2021, p7) 3The Music ecosytem study: (Music Moves Europe 2024); the European Parliament’s resolution (European Parliament 2024) and our policy paper: (Antal 2025d). 34
These principles are fully consistent with CEEMID’s decentralised model and form the conceptual basis for the Open Music Dataspace (see Chapter 5). The CITF (2025) report arrives at the same architectural conclusion. It stresses that trustworthy copyright infrastructures in the AI era require federated governance, interoperable identifiers, and verifiable provenance chains rather than a single, centralised registry. Its three-layer model—foundational identifiers, shared semantics, and technical services—maps closely onto the Observatory’s dataspace design. Observatories created in the 1990s and early 2000s were built around centralised databases and slow-moving data-collection cycles. Since then, the rapid expansion of agentic AI in data collection, the widespread digitisation of live and recorded music, and the proliferation of large-scale, real-time data sources have made such centralised architectures obsolete. Modern evidence ecosystems require automated ingestion, continuous semantic enrichment, cross-domain reconciliation, and transparent provenance — all of which presuppose a federated, decentralised model rather than a single institutional database. The European Audiovisual Observatory (EAO), the European Market Observatory for Fisheries and Aquaculture Products (EUMOFA), and the European Observatory on Infringements of Intellectual Property Rights (EUIPO) provide valuable models of long-standing EU observatories. However, each operates within a centralised data-submission and aggregation framework appropriate to their legal mandates and sectoral data structures. The Feasibility Study acknowledged that the music sector lacks comparable legal obligations and contains far more fragmented, cross-domain, multilingual, and institutionally diverse datasets. Therefore, while these observatories offer important governance precedents, their centralised architectures cannot be replicated in the music ecosystem — strengthening the case for a federated dataspace model. 2.3.3 Preference for Open-Source and Open Standards in the EU Across the EU’s data and digital-transition strategies, there is a consistent preference for: • open-source software, • open standards, • open licensing, and • transparent, reproducible workflows. This aligns directly with the Observatory’s use of open-source R and Python pipelines, Wikibase for semantic interoperability, and FAIR-compliant metadata. 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures Europeana demonstrates how Europe manages distributed cultural-haritage collections at scale using: • persistent identifiers, • multilingual metadata, 35
• open licences (e.g. CC BY), • shared semantic standards (EDM, IIIF, rightsstatements.org), and • decentralised stewardship by libraries, archives, and museums. The Open Music Observatory follows the same principles. It uses: • semantic technologies, • PID-based cross-domain linking, and • open, reusable data models. This ensures interoperability with cultural-heritage collections, performing-arts archives, and national memory institutions, and aligns the music domain with the emerging European Collaborative Cloud for Cultural Heritage (ECCCH). 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy The EU Open Data Portal (data.europa.eu) establishes a common framework for: • open licences (e.g. CC BY 4.0), • machine-readable formats, • harmonised metadata (DCAT-AP), • and publication of public-sector information. The Open Music Observatory is designed so that: • public datasets can be harvested directly by the EU Open Data Portal, • indicators and derived datasets comply with open-data rules, and • metadata follow DCAT-AP and DataCite to support long-term reuse. This alignment ensures that the Observatory meets both Horizon Europe open-science requirements and broader EU open-data policy objectives. 2.3.6 European Interoperability Framework (EIF) The European Interoperability Framework (EIF) provides a four-layer model—legal, organisational, semantic, technical—for connecting: • public authorities, • cultural institutions, • rights-management organisations, • national statistical offices, and • private intermediaries. These are precisely the actors whose data must interoperate to support a European Music Observatory. By adopting the EIF, the Observatory can link diverse datasets into coherent, reusable services without centralising them. 36
2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds The European Open Science Cloud (EOSC) and the European Collaborative Cloud for Cultural Heritage (ECCCH) promote: • FAIR data, • open science workflows, • reproducible analysis, • transparent provenance, and • decentralised storage and processing. These principles inform the Observatory’s architecture through the use of: • open-source analytical pipelines, • SDMX and DataCite metadata, • persistent identifiers, and • federated linking across domains and institutions. 2.3.8 Summary Together, these EU policy instruments—the Parliament’s mandate, the EU Music Ecosystem Study, data-space strategy, Europeana, the EU Open Data Portal, the EIF, EOSC, and ECCCH—provide a unified rationale for an Observatory that is federated, decentralised, data-driven, and interoperable by design. They define the policy and technological environment in which the Open Music Observatory must operate and directly shape its architecture. The Chapter 4explains why we chose an architecture that is built around Wikibase and Wikiadta. 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model The Open Music Observatory adopts a decentralised, federated dataspace model because this is the only architecture that meets the needs identified by the EMO Feasibility Study, the EU Music Ecosystem Study, and the European Parliament’s resolutions, while also complying with the newer EU frameworks for interoperability, data governance, and cultural-heritage infrastructures. A centralised database model, common in observatories built in the 1990s or early 2000s, is no longer feasible or desirable for the music sector. 2.4.1 Lessons from CEEMID The CEEMID collaboration demonstrated that most music-sector data—repertoire, rights, cultural-heritage descriptions, business metadata, and statistical evidence—originate from many different institutions, each with its own mandates, legal obligations, and technical systems. Centralising such data is: 37
• legally constrained (e.g. GDPR, contractual confidentiality), • institutionally unrealistic (distributed ownership and stewardship), and • technically inefficient (rapidly evolving local systems). CEEMID showed that these data can nonetheless be made interoperable through: • shared identifiers and authority files, • open metadata standards, • reproducible R-based pipelines, and • rule-based, voluntary data sharing. These are the foundational principles of a data (sharing) space, which the EU has since elevated to a core strategic component of its digital-policy agenda. 2.4.2 Requirements of the EU policy environment As outlined in Section C, the EU now expects cultural and creative sectors to adopt: • federated data architectures, • FAIR and open data practices, • transparent governance models, • semantic interoperability, and • alignment with Europeana, EOSC, ECCCH, and data.europa.eu. This expectation reflects the broader transformation of European data governance, where sectors are encouraged to organise around data spaces rather than central repositories. A decentralised model also supports cultural and data sovereignty by allowing institutions to maintain control over their collections and data-processing rules. 2.4.3 Requirements of the Grant Agreement The Open Music Europe Grant Agreement defines the project explicitly as: “an open, scalable data-to-policy pipeline for European music ecosystems” and mandates the creation of: “a highly automated, decentralised intelligence hub that aggregates open data and creates dynamic, live policy documents.” To fulfil these contractual obligations, the Observatory must: • connect heterogeneous data sources without centralising them, • refresh indicators automatically as upstream data changes, • maintain legally sound provenance across many institutions, • support multilingual, cross-border metadata, and • integrate statistical, cultural-heritage, and industry systems. 38
These requirements can only be met in a federated dataspace, not in a single, centralised database. 2.4.4 Technical rationale for decentralisation The dataspace model makes it possible to: • keep sensitive or personal data (e.g. rights, royalties) within the institution that controls them, • link sources through semantic federation (Wikibase/Wikidata), • enable distributed curation by librarians, archivists, CMOs, and researchers, • integrate permanent identifier (PID) systems across domains (ISNI, VIAF, ROR, company registers), • use open standards (SDMX, DDI, DataCite, DCAT-AP), and • scale to new partners, genres, languages, and Member States. This structure mirrors the actual distribution of data in the music sector and the technical direction of the EU’s digital transition. 2.4.5 Why decentralisation is essential for a European Music Observatory For the European music ecosystem, decentralisation enables: • lower administrative and compliance burdens, • institutional autonomy and data sovereignty, • cross-border comparability without forced data transfer, • communityand expert-driven metadata improvement, • GDPR-compliant handling of personal data, and • sustainable expansion of the Observatory. A decentralised dataspace is therefore not an architectural choice but a necessary governance model for an Observatory that spans cultural heritage, rights management, statistical registers, community archives, and private-sector metadata across the EU. The Open Music Observatory is consequently designed as a federated, rule-based dataspace: an ecosystem where public, private, and civic stakeholders contribute knowledge, maintain authority records, and generate indicators while preserving full control over their own data. 2.5 The Data-to-Policy Pipeline: How Open Music Europe Works The Open Music Europe action is contractually defined as “an open, scalable data-topolicy pipeline for European music ecosystems” (see Grant Agreement). This is not a slogan: it is the methodological core of the project and the organising principle of all work packages (WP1–WP5). The pipeline connects indicator design, data governance, 39
data acquisition, semantic modelling, statistical analysis, and policy translation into a single reproducible workflow. This chapter introduces the logic of that pipeline and explains how it shapes the design of the Open Music Observatory. 2.5.1 1. Indicator and problem definition (WP1–WP3) Each thematic work package begins by identifying policy-relevant gaps and defining the indicators needed to address them. Deliverables D1.1, D2.1, and D3.1 specify: • the conceptual frameworks guiding each domain (economy, diversity, society), • the data requirements for measuring them, and • the procedures for ensuring comparability across countries and years. These definitions also appear in the Open Music Europe Data Management Plan (D6.3), which provides human-readable summaries and machine-readable metadata for all indicators. 2.5.2 2. Data governance (WP1–WP3, WP6) Before data can be collected or integrated, partners agree on: • sources, access rights, and sampling frames; • metadata standards (SDMX, DDI, DataCite); • ethical safeguards and GDPR-compliant procedures; • controlled vocabularies, authority files, and persistent identifiers. These agreements are formalised in D1.2, D2.2, D3.2, and the Data Management Plan (D6.3). They ensure compliance with FAIR, OPA, and EU data-governance principles. 2.5.3 3. Software for data collection (WP4) WP4 develops the open-source tools used to gather and ingest administrative data, survey data, platform usage data, and CMO records. These tools form the operational backbone of the pipeline. They include: • survey-data management scripts, • connectors for royalty and licensing accounts, • streaming API integration modules, • metadata templates for ingestion and harmonisation. All tools adhere to the reproducibility and interoperability requirements defined in Annex 1 and the DMP. 40
2.5.4 4. Data acquisition (WP1, WP2, WP3) Data are collected from: • collective management organisations (CMOs), • ministries and statistical offices, • cultural-heritage institutions, • surveys (enterprise and personal), • streaming-service APIs. Each domain follows its own protocol (e.g. WP1 T1.2 sampling frames; WP3 T3.1 participation and wellbeing indicators). Data collected are documented in the DMP and crossreferenced with OPA-compliant folders. 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) Raw inputs are processed using REPREX’s R-based openmusic-pipeline: • cleaning and pseudonymisation, • metadata harmonisation, • cross-linking with authority records, • structuring in SDMX/DataCite formats, • integration via persistent identifiers. This step transforms heterogeneous inputs into consistent, analysis-ready datasets. 2.5.6 6. Validation for analysis and dissemination (WP4, WP5) Before modelling can begin, WP4 and WP5 validate: • interoperability across sources, • semantic consistency, • statistical reproducibility, • versioning and provenance, • integration with the Open Music Observatory. This ensures that indicators can be reproduced from source data and that all transformations are transparent. 41
ĹNote Four types of data-collection principles have been identified as essential both by various branches of the music sector and also by policymakers at European, national and local levels: • The data-collection service provided by a European Music Observatory should help in mapping, understanding and analysing the main characteristics, trends and idiosyncrasies of the music sector in Europe; • The data collected should be neutral and available to decision-makers, music sector operators, and the public; • The data itself should cover the activities of the music sector across the entire European Union, be comparable between Member States, and rely on identified and stable indicators; • The data collection methods should be transparent and provide a strong degree of scientific accuracy. (European Commission et al. 2020, p28) In short, we collect data about music, as defined in the cultural statistics of any European Economic Area and EU candidate statistical office or by a representative European or international music organisation. In more detail, we a systematic data collection program requires a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. ĹNote Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. We work with a concept of a musical work, not with the entire work. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. Again, in an information system we obviously work with a concept of an author, and instances of authors represented by their unique data. The EMO feasibility study catalogues 45 data gaps that a future European music observatory should fill. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. We need agreed concepts of the composer,sound recording,work, to answer such questions. 48
The initial data collection guidelines of the Open Music Observatory are derived from the EMO Feasibility study. We see them as a starting point for further discussion with the Observatory Stakeholder Network. We introduce them with our data catalogue in Section 10.1. These guidelines are supported by our first conceptualisation, which is built on some widely used conceptualisations of creative works and statistics. This is the topic of Chapter 8. 3.1.1 Microdata, Collections, Records We treat “microdata” as a collection of structured data. Aregister is a document [in modern usage, usually a database], in which data are entered in a formal manner by a statutory authority (ISO 2017b). In statistical data collection ian official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country. Acollection is a group of objects, for example, musical works, sound recordings, printed scores, music enterprises, musician biographies, gathered together for some intellectual, artistic, or curatorial purpose. This is how radio playlists and charts, festival line-ups, local content guideline monitoring works; music labels and publisher select and musical works and their recordings or scores to place into commercial circulation. Such collections form the basis of census or sample surveys for statistical data collection. The documentation of collections relies on the work of registers. For example, music publishers can claim their revenues based on ISWC and ISMN identifiers provided to them by the collective management organisations that register works, or national libraries or other organisations that identify printed sheets. The maintenance of registers requires ongoing investment, and therefore registrars like the ISRC Authority or CISAC (the manager of the ISWC register) often restrict access to their data, or do not exchange data. In an increasingly globalised, automated music ecosystem where the number of identifiable works, recordings, scores, and related claims is growing exponentially, this situation puts the entire industry at a disadvantage, for example, against tech platforms. The Open Music Observatory is experimenting with innovative ways how registers can work together in some aspects of metadata standardisation, improvement and exchange in a way that keeps their core product intact and exclusive to them See: (Antal and Mester 2025). The Open Music Observatory works with metadata in a way that helps managing and improving registers, and it helps to create data about collections with authoritative data from registers. ĹNote Our first large database is the Slovak Comprehensive Music Database. Our aim is to publish a constantly refreshed database of every music composed or recorded in the territory of the current Slovak Republic, or composed and recorded by people from 49
Slovakia, or sung in the Slovak language. This database is partly based on registers, and partly on curated holdings of Slovak stakeholders. ⊠Our collections are always available on https://reprexbase.eu/skcmdb/�. Further details in Section 3.5.1. ⊠Whenever our collections fit in the collection and publication guidelines of Europeana, we make the collections available there, too. □We are investigating the possibility of synchronising our collections to the European Collaborative Cloud for Cultural Heritage. The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. The DDI plays a particularly important role in the creation of statistical surveys, particularly using questionnaires and question banks. The new Records in Context has replaced the international standards on archives in 2023. Its central concept is the record, which is a document according to DDI; a collection is a set of records. Our standardisation of microdata is explained in more detail in Chapter 8. 3.1.2 Primary data collection The Open Music Observatory is supporting high-quality primary data collection, and itself is carrying out such collection activities. The indicators derived from the processing of survey questionnaires will be comparable if the same concepts of interest (for example, concert visiting frequencies) are measured via the same questions and answering instructions. 50
Aconcert is a standard concept of a live performance of music. How many times in the previous [12 months] have you been to a concert? is a standard question accompanied by standardised answer options and processing in the Cultural Access and Participation surveys following the ICET model. Using standardised concepts and question banks, including question and instruction labels with standardised translations, is a cornerstone of ex-ante survey harmonisation. This process is a prerequisite for retrospective survey harmonisation and the subsequent creation of comparable statistical indicators, underscoring the importance of uniformity in data collection. ⊠We provide API and download access to harmonised, multi-language question banks. This allows music stakeholders to use the same question formulations and translations for comparability with European statistical and policy research programs. ⊠We provide tutorials to retroharmonize, a background open-source software of Reprex, which is an R library to retrospectively harmonise data from different survey programs that had asked the same questions. ⊠The Open Music Europe project will carry out some harmonised surveys to show and improve the methodology of harmonised data collection within the music sector of Europe. This data will be available as metadata (questionbank), as microdata (individual answers), and as processed statistical data. 3.1.3 Metadata The most common—and perhaps least useful—definition of metadata is that it is “data about data.” As catchy as this definition is, however, it is entirely ambiguous. First of all, what is data? And second, what does “about” mean? (Pomerantz 2015a, p19) The new ISO standard on Information technology — Metadata registries (MDR) defines metadata as data that define and describe other data. As Pomerantz eloquently argues, this is a definition that is not very helpful. We use his more functional (but not contradictory) definition. “Data is only potential information, raw and unprocessed, prior to anyone actually being informed by it. […] Data must be understood not as an abstract concept but as objects that are potentially informative. […] Metadata Is a Statement about a Potentially Informative Object.” (Pomerantz 2015a, p26) Following the metadata definition of “a statement about a potentially informative object,” we believe that any high-quality data can be used as metadata in certain circumstances. ĹNote Data or metadata? The data of birth can be seen as a metadata for disambiguation among authors with the exact same name in a copyright register. It can be seen as data for a curator of a 51
young author prize, or a music sociologist. Either way, the date of birth should be precise, and encoded in a way that makes it portable and interoperable. From a data management point of view, we do not distinguish between data and metadata. Of course, we acknowledge the fact that some types of data will always remain under the hood and will only serve the proper functioning of an information system. The music industry’s famous “metadata problems” usually arise when a music enterprise or institution wants to use metadata information from an authoritative source that is somehow corrupted. The Open Music Observatory can help with these metadata problems by disseminating proper, open authoritative data (as registers or collection) or by providing data improvement services that fix the metadata problems of a user. 3.1.4 Statistical indicators and datasets ÁWarning We will place our first statistical datasets to the EU Open Data Portal this week (pending their approvals) and will provide a screenshot and access conditions here. 3.2 Repair Throughout the project we realised that data and metadata repair is perhaps a more urgent challenge then data processing. While our team and our stakeholders gradually embraced the concept of a data sharing space, i.e., the idea that instead of starting new data collections from scratch it is more economic and useful utilise existing data, given the high level of music industry digitalisation and that almost all transactions leave a digital trail behind, we also realised that the music sector is “drowning in numbers”; it handles more data in various obsolete, undocumented, unstructured or ad hoc forms than it can utilise. In these cases, usually the data is already available somewhere, but in a format that prevents the data to be used to its potential. Most of our efforts therefore were concentrated on data and metadata repair instead of new collection. Metadata repair and data processing is usually hard-to-distinguish tasks that comprise of similar or same steps. They are various validation, normalisation procedures that allow that make the data informative. 3.3 Process We use the theory of metadata by Jeffrey Pomerantz, who defines Metadata as “a statement about a potentially informative object.” A dataset without such statements is not 52
findable, accessible, interoperable, and very hard to reuse. Pomerantz distinguishes among descriptive, administrative, structural, preservation, and use metadata. The Generic Statistical Information Model (GSIM) is a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation metadata standards such as SDMX or DDI. GSIM since its inception aims to bridge two important standards, SDMX and DDI. The Statistical Data and Metadata Exchange has been developed for decades and it is an ISO standard; it is more geared towards the aims of data sharing and preservation in RDM. DDI on the other hand is more focused on the documentation and quality control of primary data collection, or the reuse of often messy data sources, and supports the processes that make the data available for research. As DDI provides information about a much wider range of objects and processes, we are even more selective when we turn to this standard than SDMX; however, we cannot disregard DDI for microdata. 3.3.1 Processing & re-processing microdata ÁWarning We will place here an example that goes to the EU Open Data portal The EU Open Data Portal uses the following namespace definitions; these definitions refer to machine readable, explicit definitions (ontologies) of the way our datasets must be understood by a software agent. To demistify the process, we provide here an example of the metadata that we need to compile from the various steps of the data production pipeline. @prefix rdf:<http://www.w3.org/1999/02/22-rdf-syntax-ns#> . @prefix foaf:<http://xmlns.com/foaf/0.1/>. @prefix rdfs:<http://www.w3.org/2000/01/rdf-schema#> . @prefix xsd:<http://www.w3.org/2001/XMLSchema#> . @prefix owl:<http://www.w3.org/2002/07/owl#> . @prefix adms:<http://www.w3.org/ns/adms#> . @prefix dcat:<http://www.w3.org/ns/dcat#> . First we must translate the metadata of our datasets to any of the standard serialisations (file formats) of the World Wide Web Consortium’s Resource Description Framework definition, which allows the connection of data across the open internet. At the time of writing this report, the EU Open Data Portal was changing its backend, and for testing purposes, we worked with a dataset from the background of the Open Music Europe project (which had been earlier published by Reprex on Zenodo under the title *The turnover of the ration broadcasting industry in Europe*. ) The dataset itself cannot be downloaded from a data catalogue. It is an abstract intellectual work, similar to musical work or a literary work. A musical work is accessible in printed sheets or recordings, and a dataset in a distributed data file. 53
<https://doi.org/10.5281/zenodo.5652118> <a>"dcat:Dataset" ; <dcat:distribution><https://zenodo.org/records/5652118/files/codebook_trb.csv>,"https://zenodo.org/records/5652118/files/codebook_trb.csv" ; <dct:creator><https://orcid.org/0000-0001-7513-6760>; <dct:description>"\"The turnover of the ration broadcasting industry in Europe.\"@en" ; <dct:identifier><https://doi.org/10.5281/zenodo.5652118>; <dct:issued>"2022-06-03T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:modified>"2022-06-04T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:publisher><https://isni.org/isni/000000050973936X>; <dct:title>"A rádió szektor forgalma Európában\"@hu","\"Turnover of the Radio Broadcasting Industry in Europe\"@en" ; <edp:originalLanguage><rdf:resource><http://publications.europa.eu/resource/authority/language/ENG>. We can provide further provenance information about the dataset; in production, we will provide information on software agents (tools) used, researchers, data managers and curators and their organisations involved. As a bare minimum, we provide machine-readable information about the technical publisher of the dataset, Reprex B.V: <https://isni.org/isni/000000050973936X> <a>"foaf:Agent" . And then we point the user the downloadable files (distributions) of the dataset with the rights statements and licenses. We use the Creative Commons CC BY 4.0 license, similar to Eurostat on the EU Open Data Portal, and we state that the dataset is open for the public. <https://zenodo.org/records/5652118/files/codebook_trb.csv> <a>"dcat:Distribution" ; <dcat:accessURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:byteSize>"41672" ; <dcat:downloadURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:mediaType>"text/csv" ; <dct:license><http://publications.europa.eu/resource/authority/licence/CC_BY_4_0>; <dct:rights><http://publications.europa.eu/resource/authority/access-right/PUBLIC>; <owl:sameAs><https://zenodo.org/records/5652118>. 3.3.2 Documentation ÁWarning We will provide the link and screenshot of the documentation for each file that goes public. 54
3.4 Disseminate 3.4.1 Open Music Observatory The Open Music Observatory website provides 55
Figure 3.3: Our temporary landing page on <https://dataobservatory-eu.github.io/omolanding-page> The Open Music Observatory is a federated dataspace with joint services. The temporary data sharing space currently contains three, and soon four federeated spaces. • The Finno-Ugric Data Sharing Space is our experimental data sharing space that connects the music and broader immaterial and material heritage of fifteen small, mostly endangered ethnolinguist groups; including our Livonan Music Database (see Annex), Karelian Music Database, contemporary and heritage Mari, Udmurt, Székely, Csángó and other music. https://finnougric.net/en/ • The Slovak Comprehensive Music Database connects the services of the rights management agency SOZA, the Slovak Music Center, the Slovak National library and some public libraries of Slovakia. It is gradually filled up with music, and it shows where Slovak music (in the form of music recordings, videos, printed sheet, or other publications) can be found and listened to on streaming services, libraries, webshops. This federated unit is well-governed by Slovak national entities. https: //hudobnadatabaza.sk/en/ • The Open Music Observatory is mainly intended to give access to the Economy, Diversity, Society and Sustainability and Innovation data and publications of the Open Music Europe (OpenMusE) project as well as data from Poland, Latvia, Portugal and other countries. This federated unit is not yet finished by the OpenMusE consortium and is designated to be handed over to pan-European organisations of music http: //135.181.91.51:3008/en/ The Hungarian Comprehensive Music Database will follow soon. 56
Figure 3.4: Services for individuals, API access for industry and research partners, and planned trustworthy AI. The Observatory, unlike, for example, the European Audiovisual Observatory, has no permanent staff, and it is completely automated. It offers four layers of services. • Access to the Wikibase Suite GUI that gives very detailed description of datasets, publication, musical works and their textual or recorded, video manifestations. (See for example: fudss:Q4745) Manual access and ad hoc dumps for RDF serialisations of this information (See for example Q4745.ttl, or Q4745.jsonld.) • Access to the Sampo open-source semantic browsers that gives faceted search access for researchers, individuals, who cannot write API queries. See for example the https: //hudobnadatabaza.sk/en/ • API access to partners (requires password, authentication) that provides SPARQL queries with CSV download. • A planned extension for December-January is a trustworthy AI chatbot that connects AI agents via a strict Model Context Protocol to the knowledge base(s) of the observatory, and allows custom knowledge retrieval from the knowledge base without hallucations in natural language. 57
4 Architecture The Open Music Observatory (OMO) is designed as a decentralised, federated, semantically interoperable knowledge infrastructure. Its architecture is the direct consequence of the policy, governance, and methodological requirements described in the Background chapter (see Chapter 2, Section 2.3, Section 2.4, and Section 2.5). The architecture is not an aesthetic or purely technical choice: it is the only sustainable and legally compliant way to realise the “open, scalable data-to-policy pipeline” mandated in the Grant Agreement. The architectural choice is also reinforced by the findings of the CITF First Project Report (2025), which examined the requirements for trustworthy, lifecycle-aware copyright infrastructures in the AI era. CITF identifies the same structural needs—federated governance, interoperable identifiers, machine-readable rights metadata, and verifiable provenance—that underpin OMO’s design. This convergence shows that the Observatory’s dataspace architecture is not only technically justified but part of a broader European shift toward distributed, standards-based copyright and metadata infrastructures. Traditional observatories created in the 1990s and 2000s were built on centralised databases, static data submissions, and annual reporting cycles. Such architectures are no longer viable in an ecosystem where: • metadata is produced continuously across incompatible systems; • rights, repertoire, event, and business data change daily; • cultural-heritage and community archives hold essential non-market metadata; • statistical offices and ministries have different identifiers and legal mandates; • AI-assisted workflows require transparent provenance and versioning; • multilingual and cross-domain reconciliation is unavoidable. These conditions make a federated dataspace (see Section 2.4) the only feasible approach. Within such a dataspace, Wikibase/Wikidata provides the most suitable technological foundation. 4.1 Why Wikibase Is the Right Foundation for the Open Music Observatory The European music ecosystem produces data in many incompatible formats: • repertoire and rights databases (CMOs, publishers, labels), • cultural-heritage and performing-arts collections, • company registries and economic statistics, 64
• event metadata from festivals and venues, • community-maintained sources (Wikidata, folk archives, local heritage groups). These sources use different identifiers, legal definitions, languages, and metadata schemas. A knowledge graph is therefore indispensable: no relational or document database can reconcile these sources while maintaining provenance, multilinguality, and entity-level linking. ĹAlignment with CITF’s Copyright Infrastructure Model The CITF First Project Report defines three layers that a modern copyright infrastructure must satisfy: a foundational identifier layer (authoritative PIDs and registries), a semantic layer (shared meaning and mapping across domains), and a technical layer (APIs, services, resolution, and provenance). Wikibase operationalises this same structure in practice. Its support for persistent identifiers, ontology alignment, multilingual semantics, and transparent versioning makes it fully compatible with the CITF model and positions the Observatory as a concrete, domain-specific implementation of that broader European framework (Partanen et al. 2025). Wikibase is adopted because it has already proven its suitability in domains facing the same structural issues as the music ecosystem: fragmented identifiers, inconsistent authority control, multilingual metadata, cross-domain vocabularies, and parallel institutional workflows. 4.1.1 Proven in real-world scenarios highly similar to music Wikibase is widely used by libraries, archives, museums, national cultural bodies, and openscience infrastructures. These institutions face the same challenges OMO addresses: • reconciliation of people, works, events, organisations, and places; • multilingual labels and aliases; • authority file alignment (ISNI, VIAF, ORCID, GND, BNF, corporate registers); • provenance tracking and version history; • SPARQL-based validation and constraint checking. The GLAM-Wiki ecosystem, national knowledge graphs, and EU-funded linked-data projects have collectively demonstrated that Wikibase is an effective intermediary between: • authoritative PID systems; • domain ontologies (CIDOC-CRM, RiC-O, DDI, DCAT); • community-curated knowledge models. This track record gives OMO a mature, future-proof, and standards-aligned foundation. 65
4.1.2 Already aligned with Europe’s digital knowledge infrastructure Wikibase aligns with existing institutional practice across Europe. Many major knowledge centres and initiatives already use Wikibase/Wikidata: • national libraries and archives, • national cultural-heritage aggregators, • research infrastructures, • public-sector linked-data programmes, • the EU Knowledge Graph initiative. Adopting Wikibase ensures that the Open Music Observatory fits directly into the European interoperability ecosystem (see Section 2.3). Footnote: See Wikibase as an Infrastructure for Knowledge Graphs: the EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021) and (2020 2020). 4.1.3 Demonstrated support for required OMO functionality Everything OMO needs has already been demonstrated in production Wikibase environments: •authority control for creators, ensembles, organisations, venues; •multilingual and multiscript modelling for names, places, works; •cross-domain entity linking (work–recording–performance–rights–heritage); •event-based and entity-based models; •SPARQL validation, schema constraints, and automated reconciliation. This means OMO does not invent an untested paradigm: the consortium integrates proven practices from: • national registries, • performing-arts knowledge graphs, • the Slovak pilot and Finno-Ugric metadata federations developed inside the project. 4.1.4 Fits EU policy preference for open-source and trustworthy AI Wikibase is open-source, auditable, and non-proprietary. It aligns with: • the EU’s preference for open-source digital public infrastructure, • FAIR and CARE principles, • trustworthy AI requirements (provenance, transparency, versioning), • cross-border interoperability mandates, • decentralised data governance models. 66
This makes it compatible with Europeana, EOSC/ECCCH, DCAT-AP, and the European Interoperability Framework. It is used in EU organisations, too. 1 4.1.5 The most widely used graph-editing interface in the world Tens of thousands of data stewards, librarians, researchers, and citizen-scientists already know how to edit Wikibase/Wikidata. This provides OMO with: • an immediate user base, • a ready-made contributor community, • institutional familiarity across Europe, • workflows already adopted in GLAM and research sectors. No alternative open-source system has remotely this level of adoption. 4.1.6 A hybrid model that fits real institutional workflows Wikibase uniquely accommodates: • spreadsheet-based workflows (Excel, CSV), • relational database exports, • statistical microdata reference linking, • complex semantic modelling, • API-based ingestion, • R and Python pipelines. It is a practical compromise between triple stores, document databases, and relational systems — perfect for a music ecosystem where many partners still rely on basic tools. It has proven to be useful in music services2, and more generally on smalland large scale European knowledge institutions (national libraries, libraries, archives, museums3.) 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline The data-to-policy pipeline defined in the Grant Agreement and documented in the Background chapter (see Section 2.5) provides the methodological backbone of OME. Wikibase is the component that makes this pipeline operational. 1On official adoption: EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021); SEMIC guidelines (SEMIC Support Centre 2023). 2On Belgian pilots: MetaBelgica (Stallmann et al. 2023) and Flemish performing arts enrichment (Magnus and Van D’huynslager 2021). 3See for example: On Wikidata/Wikibase in heritage: (Bianchini, Bargioni, and Pellizzari di San Girolamo 2021; Sardo and Bianchini 2022). 67
4.2.1 Wikibase supports each stage of the pipeline 4.2.1.1 Data collection • imports from Excel, CSV, SQL, APIs, and legacy systems; • immediate linkage to persistent identifiers; • entity reconciliation as part of ingestion. 4.2.1.2 Validation and reconciliation • authority-control workflows for people, works, organisations, and places; • constraint-based quality checks; • SPARQL-driven validation; • alignment with external authority files. 4.2.1.3 Harmonisation and enrichment • multilingual labels and roles; • event-based and relationship-based modelling; • addition of contextual metadata by different institutions; • integration of domain vocabularies. 4.2.1.4 Activation for analysis • SPARQL endpoints for programmatic access; • JSON-LD, RDF dumps, and REST APIs; • R and Python pipelines use stable URIs for reproducibility. 4.2.1.5 Indicator construction • cross-domain indicators linking economic, cultural, rights, and heritage data; • entity-level referencing ensures indicators are traceable and verifiable. 4.2.1.6 Interpretation and contextualisation • experts review, annotate, and correct metadata through a human-readable interface; • provenance guarantees transparency. 68
4.2.1.7 Policy translation and observatory outputs • live, federated knowledge base powering the OMO front end; • entity profiles, metadata dashboards, and linked methodological documentation; • direct links from indicators to source entities and datasets. 4.2.1.8 Feedback loop • corrections flow back into the shared graph; • updated authority files update all downstream indicators; • the observatory improves over time. 4.3 Summary Wikibase is adopted because it best fulfils the policy requirements, semantic needs, technical constraints, and governance expectations described in the Background chapter: • It is proven in domains identical to ours. • It is aligned with the EU’s open-source, dataspace, and interoperability agenda. • It is widely adopted by the very institutions we must interoperate with. • It supports the entire Open Music Europe data-to-policy pipeline. • It allows the Observatory to operate as a decentralised, federated, evolving knowledge infrastructure. This architecture is therefore not an optional design preference but the only viable model for delivering the European Music Observatory described in the Grant Agreement. 69
5 Data coordination Cultural and music data in Europe is generated and maintained by a highly diverse set of organisations. National libraries, archives, museums, collective rights management organisations, academic research projects, streaming platforms, local cultural associations, and commercial distributors all describe the same creators, the same works, and the same recordings—but they do so for different purposes, using different conceptual models, and under different legal and organisational constraints. These systems have grown organically over decades, reflecting local traditions, professional cultures, and technical possibilities. As a result, the same piece of music may be catalogued differently in an archive, credited differently in a rights database, indexed differently in a library catalogue, and published differently on a digital music service. This fragmentation was already identified in the Feasibility Study for a European Music Observatory, which noted that music-sector metadata is “structurally scattered across private and public registers that seldom interconnect” (European Commission et al. 2020, 9–10). The Study explicitly highlighted the absence of a “functional link between ISWC, ISRC, ISNI and VIAF” as a core obstacle to evidence-building [p. 46–48]. The CITF First Project Report reaches the same conclusion in the textual domain, observing that rights metadata lacks “cross-registry alignment, lifecycle-based provenance, and interoperable identifiers” [Partanen et al. (2025), p. 20–22; p. 31]. These findings reinforce the need for a coordination model that can operate across independent registers without replacing them—precisely the purpose of a data sharing space. This fragmentation is not a sign of failure; it is the natural consequence of cultural stewardship being distributed across many institutions with distinct missions. However, fragmentation becomes a problem when people attempt to connect, reuse, or reinterpret information across domains. Without coordination, data cannot be trusted across systems: names do not match; contributor roles are incompatible; identifiers are missing or ambiguous; and important aspects of minority or community-based cultural heritage may be lost or become invisible. The need for coordination arises from this structural diversity. Institutions should not be forced to adopt a single metadata standard or a single software system, but they must be able to communicate meaningfully. Coordination provides the minimum shared semantic and technical foundations that enable such communication. Coordination also supports equity and cultural diversity. Smaller institutions—community archives, ethnographic collections, minority cultural groups, and local music organisations—often lack the resources to operate modern digital infrastructures. Without coordinated frameworks, their data cannot easily participate in larger European infrastructures or data spaces, which in turn reinforces cultural imbalances. 70
Coordination therefore serves not only technical efficiency but also cultural policy objectives related to inclusion, multilingualism, regional diversity, and the representation of endangered traditions. The European Parliament’s Resolution on Cultural Diversity and the Conditions for Authors in the European Music Streaming Market makes this a policy priority, stressing that fair remuneration and discoverability depend on “comprehensive and accurate allocation of metadata from the time of creation” and on adoption of all international identifiers (IPI, ISWC, ISRC, ISNI, IPN) (European Parliament 2024, recital 32). The EU Music Ecosystem Study similarly notes that fragmentation in metadata and identifier systems “prevents crossborder comparability and weakens the evidence base for policy” (Music Moves Europe 2024, 42–45). Coordination is therefore a legal and policy requirement, not simply a technical improvement. 5.0.1 The European Interoperability Framework (EIF) The European Interoperability Framework (EIF) provides guidance for how heterogeneous public-sector systems can work together without sacrificing their autonomy. It identifies four layers of interoperability—legal, organisational, semantic, and technical—and emphasises that meaningful interoperability requires all four to be addressed1. The EIF provides the most relevant cross-sectoral guidance for this problem. The European Interoperability Framework Implementation Strategy stresses that interoperability must balance institutional autonomy with cross-sector alignment, and that semantic mediation—not schema unification—is the correct method for heterogeneous domains (eif_2017?). This perspective is echoed in the Data Spaces Support Centre Blueprint v2.0, which states that common European data spaces must provide “shared semantics, identifier resolution, and rulebooks for interoperability” rather than imposing a single ontology (Data Spaces Support Centre 2025, 35–39). The EIF’s underlying philosophy is that institutions will always differ in their practices, constraints, and goals. Therefore, interoperability cannot be achieved by imposing uniformity. Instead, it must enable institutions to maintain their internal systems while still participating in a shared ecosystem. Although the EIF was not developed specifically for the cultural or music domains, but generally to public digital services, its principles apply directly, because most music services are digital, and they are often public in nature, too. 1The EIF defines layered interoperability (legal, organisational, semantic, technical) (European Commission 2017). The European Strategy for Data frames subsidiarity as compatible with federation (European Commission 2020). BDVA and the Federation Working Group emphasise that interoperability frameworks are needed to operationalise federation (BDVA/DAIRO 2023; BDVA/DAIRO Federation Working Group 2023). 71
Figure 5.1: The role of data interoperability is not to exchange or centralise data, but to create digital services in libraries, rights management, distribution, streaming platforms that work together well. For the service providers they should work with less data processing cost, and for the user they should give a higher value or better user experience. Cultural data infrastructures face the same challenges the EIF describes: • different legislation (e.g., copyright, orphan works, deposit laws), • different organisational workflows (cataloguing, digitisation, rights clearance), • and different semantic traditions (archival provenance, library authority control, rights metadata, music industry credits). The EIF also emphasises reusability, transparency, and proportionality—all central issues for cultural data governance, particularly when dealing with sensitive material, intangible heritage, and community-generated knowledge. Most importantly, the EIF recognises that interoperability is not a technical property alone. It is a sociotechnical negotiation across institutions, requiring trust, shared agreements, and mechanisms for maintaining alignment over time. This recognition makes the EIF a strong conceptual foundation for coordinating cultural and music data, even when many actors fall outside the classical public-sector domain. 72
5.0.2 Extending the EIF to public and private service coordination Cultural data ecosystems are unique because they span both public and private infrastructures. A library’s bibliographic record for a musical work interacts with rights information managed by collective management organisations, which in turn interacts with ISWC and ISRC registries maintained by industry actors. A museum’s digitised ethnographic collection may be reused by researchers, by minority communities seeking cultural revival, or by streaming services presenting curated heritage playlists. Public institutions depend on commercial metadata pipelines for up-to-date identifiers, contributor roles, and distribution metadata; commercial actors depend on public institutions for authoritative information, contextual enrichment, and long-term preservation. To coordinate across this mixed ecosystem, the EIF must be extended beyond its original scope. This extension is explicitly anticipated in the EU’s data governance framework. The Data Governance Act requires that cross-sector reuse of protected data be mediated through shared governance mechanisms and semantic alignment [Regulation (EU) 2022/868, Art. 8–12]. The Data Act reinforces that business-to-government and business-to-business reuse must rely on interoperable identifiers and machine-readable formats [Regulation (EU) 2023/2854, Art. 3–4]. The CITF report further argues that trustworthy AI requires alignment across private registries (publishers, collecting societies) and public infrastructures (national libraries, VIAF, ISNI), because “lifecycle-level provenance cannot be ensured in siloed systems” (Partanen et al. 2025, 101–2). Public–private coordination requires shared semantics, shared identifiers, and shared rules for linking, even when underlying workflows remain distinct. For example, DDEX roles must be interpretable by library and archival systems, just as archival provenance must be interpretable by music distributors when republishing historical recordings. Similarly, community-generated metadata (such as language revitalisation initiatives or fieldwork collections) must be able to participate in rights and distribution pipelines without being forced into commercial templates that erase cultural meaning. Extending the EIF in this way does not imply that private systems become part of public digital government infrastructures. It means that coordination mechanisms—identifier mapping, semantic mediation, minimal ontological alignment, provenance tracking, rights documentation—must operate across both sectors. This allows public and private services to interoperate responsibly while preserving their organisational autonomy and legal boundaries. The result is a more inclusive, resilient ecosystem, where cultural data can move across domains without losing meaning or trustworthiness. 73
the data sharing space preserves organisational autonomy while enabling semantic connections. The goal of the data sharing space is not data exchange for its own sake. The goal is to let heterogeneous systems work together across the lifecycles of cultural material, even when they treat the same agents and works very differently. Librarians, archivists, rights managers, ethnographers, cultural researchers, and music distributors all produce and need information about the same individuals—but they do so for different purposes, at different times, under different constraints, and using different data models. The data sharing space respects these differences rather than erasing them. Its contribution is to connect, not homogenise. By allowing a person to accumulate multiple, context-specific contributions—and by permitting these contributions to be classified differently in different organisational settings—the DSS creates an interoperative ecosystem without imposing a universal conceptual schema. This approach is completely aligned with the EIF: •Legal interoperability: each institution retains its own rights and contractual frameworks. •Organisational interoperability: each institution keeps its workflow and functional logic. •Semantic interoperability: meanings can be mapped without requiring conformity. •Technical interoperability: the systems exchange identifiers, links, and reference patterns. This is why a data sharing space succeeds where unified ontology engineering fails. It supports real institutional work, not an abstract harmonisation of concepts, enabling rights management systems, library catalogues, archival suites, and music distribution platforms to collaborate while remaining true to their respective missions. 5.2.1 Polyhierarchy The issue of polyhierarchy is not a minor edge case: it is at the heart of why large-scale, multidomain knowledge graphs—such as Wikidata, or our own data sharing spaces—struggle with representing cultural and musical reality accurately. In the Wikidata Ontology Cleanup Task Force and the Mereology Task Force, that the data professionals of the Open Music Observatory joined, we encounter the same structural tension: different communities use the “part of” relation in incompatible ways because they operate in different hierarchical systems that reflect different organisational needs. Wikidata currently contains probably over a hundred implicit or near-equivalent uses of the “part of” relation. Sometimes “part of” expresses a physical mereology (a page as part of a book), sometimes a conceptual hierarchy (a movement as part of a symphony), sometimes an organisational grouping (a track as part of an album), sometimes a legal– economic relationship (a track being part of a rights bundle), and sometimes an abstract 80
knowledge-organisation hierarchy (a concept being part of a broader field). These are not equivalent uses, and attempting to force them into a single hierarchy or a strict ontology rapidly leads to contradictions. This is not a flaw in Wikidata—it is an unavoidable consequence of operating a generalpurpose ontology used simultaneously by librarians, rights managers, musicologists, archivists, biologists, pharmaceutical researchers, geographers, and cultural heritage professionals. Each of these communities inherits its own modelling tradition. And crucially, those traditions reflect organisational workflows, not just semantic preferences. For example, in music: - A library catalogue does not care that a CD contains ten tracks composed by different authors; it treats the CD as a loanable unit. - A rights manager cares deeply about track-level authorship because royalties and licensing depend on it. - A musicologist may care about movements and sub-movements within a single work. - A digital distributor (DDEX) distinguishes between recordings, sound files, releases, and release bundles in ways foreign to both librarians and archivists. No single hierarchy can serve all of these purposes simultaneously. This is precisely the problem that the archival world has grappled with for decades. In traditional archival theory, the “unit of description” could be: a fonds, a series, box, folder in a fonds, a file, a letter or even a page of the letter. This is not a matter of ontological taste; it is an organisational reality shaped by the scale of collections, staff capacity, digitisation status, and cataloguing philosophy. Archives did not fail to formalise these differences because they lacked ontologists—they struggled because hierarchy itself is variable, contextual, and meaningful. In our efforts to create a pragmatic model for the Open Music Observatory we were informed and influenced by the outcomes of the ten-year effort to produce the Records in Contexts Conceptual Model (RiC-CM). RiC is essentially an attempt to design an ontology that could tolerate hierarchical variability between a deep archival system of a well-staffed national archive and a small community archive. RiC does not eliminate polyhierarchy; instead, it tries to provide a flexible, graph-oriented language for representing multiple hierarchical and non-hierarchical relationships at once. The result is powerful but also demanding. It is telling that, despite the conceptual elegance of RiC, few archives have rushed to replace ISAD(G). The conceptual shift is too large, the costs too high, and the implications for organisational workflow too complex. Music archive designers will have to support ISAD(G) and RiC for decades to come. (Needless to say, the designers of RiC were fully aware of this and the gradual transition is possibly.) Similarly, libraries never truly adopted FRBRoo, even though it theoretically offered a perfect hierarchy of Work → Expression → Manifestation → Item. Why? Because libraries operate under real organisational constraints, like existing cataloguing practices, legacy IT systems, patron-facing interfaces that users are loyal too, acquisition workflows, physical holdings and their storage,loan systems. FRBRoo’s clean theoretical hierarchy simply does not map onto the messy, practical realities of the often underfunded realities of a smaller music library’s operations. 81
This is exactly the lesson that Wikidata, and any data sharing space, must internalise: polyhierarchy is not a modelling flaw—it is a reflection of institutional diversity. Different institutions use different hierarchies not because they disagree, but because they serve different roles. A “part of” relation means something different in a conservation lab, a collective rights management organisation, a radio station’s playlist editor, a digitisation workflow, or an archival fonds description. The challenge is not to normalise all of these viewpoints into a single hierarchy. The challenge is to create a modelling environment where multiple hierarchies can coexist, where contradictions do not break the system, and where cross-domain linking is possible without forcing all contributors into one conceptual framework. This is precisely why Wikidata is cautious about enforcing strict subclass/instance models and the the a Mereology Task Force is needed just to understand the consequences of many near-equivalent meanings of the “part of” relationship. 5.2.2 Formalisation Advocates of data spaces sometimes describe them as “just connecting databases on an asneeded or as-permitted basis,” as if connection were a loose, informal network of API calls. But this is misleading. From a legal point of view, a database is either connected or not connected: if a system can dereference identifiers, look up metadata, or reuse statements, then the obligation to respect licenses, consent frameworks, provenance, and organisational constraints is triggered regardless of how lightly the connection is described. The same precision applies on the semantic level. Our data sharing space does not avoid formal modelling. We represent our conceptual model in the Wikibase ontology, and this can be—and in practice must be—compiled into RDF and OWL axioms so that machines can interpret the classes, properties, role patterns, qualifiers, and constraints. We use property semantics, subclass hierarchies, equivalence mappings, reified statements, and alignment patterns that are fully compatible with OWL 2 DL or OWL 2 RL, depending on the use case. Thus, “semantic mediation without semantic conformity” does not mean informality or the absence of a schema. It means something conceptually subtler: The data-sharing space maintains a formally specified ontology—but one that does not require all participating institutions to conform to a single, domainunifying conceptualisation. The formalism exists at the mediation layer, not as a global schema that all domains must adopt. A data space cannot function without a shared URI space, identity management, explicit class/property declarations, equivalence/near-equivalence mappings, domain/range expectations, inference patterns (even if limited), or consistent referential semantics. We just understand that we are making heavy trade offs for interoperability of systems, instead of serving one type of system’s internal workflows. 82
5.3 Future-Proofing Future-proofing in a Data Sharing Space means designing systems so that the knowledge we curate today will still make sense — technically, legally, and conceptually — ten, twenty, or fifty years from now. Because our audience spans librarians, archivists, music-industry professionals, rights managers, researchers, and IT staff with very different technical backgrounds, the future-proofing strategy of the Open Music Observatory must be both technically rigorous and described in familiar terms. To make this clearer, we describe future-proofing at four levels: the data model, the technology, the semantics, and the organisational workflows. 5.3.1 Future-proofing through graph architecture Many library and industry IT systems are still based on relational databases (MySQL, Oracle, SQL Server) or on simple tables (Excel, Access). Increasingly, people have heard the word “NoSQL,” even if they haven’t used such systems directly, and associate it with flexibility and modernity. A knowledge graph is a type of NoSQL system — but it is more than that. In a graph database: • the schema is not a separate document living in an IT department’s folder; • the schema is encoded in the data itself; • every relationship (composer of, recorded at, part of, published by) is explicitly stored alongside the entities. This means: • the structure of the data can evolve without breaking old records; • old data remains meaningful even after conceptual models change; • future systems can read the RDF graph and rebuild the entire database without needing to know the original software. Sadly, we often hear about cases when a developers passed away, or retired, and the latest database schema only existed in their heads. Graph databases connect the schema definition to every “table cell” forever. Often, our metadata repair work is re-discovering the not formalised schema of a legacy system. For underfunded library IT environments, this is crucial. A “TTL dump” or “JSON-LD export” is not just a backup — it is the guarantee of a future database, often with a path to transitioning to a cheaper open-source loan or archive management software with the explicit help of documenting the data first. In a music label, such future proofing of Excel repertoires helps to automate catalogue transfer to a more affordable or better quality distributor. 83
5.3.2 Stabilising definitions through internationally defined standard vocabularies When libraries, archives, or rights organisations use an SQL database, the meaning of a field often lives only inside that software. For example: • “CreatorName” inside an old PHP/MySQL-based webshop might mean a composer, or a performer, or someone who once clicked the upload button. • An “Author” field in a legacy library system may mix lyricists, arrangers, field collectors, annotators, and editors. But DCTERMS, RDFS/OWL, and DDEX give clear, internationally defined meanings to these concepts. This means: • once your data is mapped to these standard vocabularies, • future systems — even ones not invented yet — will know how to interpret it. This is why DDEX is so powerful: even if a label used a 20-year-old MySQL database, once mapped to DDEX, the meaning becomes future-proof. The same applies to: • ISWC (work identifiers), • ISRC (recordings), • VIAF/ISNI (persons), • RiC-O (archives), • MARC relator codes (libraries). Even if the technology changes, the semantics remain stable. 5.3.3 Future-proofing through translatability and multiple serialisations In legacy systems, a database backup is often useless outside its native software. For example: • An Access .mdb file from 2008. • A PHP webshop dumping csv files in an ad-hoc format. • A FileMaker Pro database from a defunct project. A graph database solves this because RDF data can be exported in many serialisations: • Turtle (.ttl) • JSON-LD • RDF/XML • N-Triples • JSON 84
Each of these files is both human-readable and machine-readable, and any future graph system can rebuild the knowledge base from them — exactly the same way software can be rebuilt from source code. If you take a look at the description of the Structural Business Statistics (OpenMuse) dataset (Q600), it can be read as om:Q600.ttl or om:Q600.json For cultural heritage preservation, this is vital: If we preserve the RDF, we preserve the meaning — not just the data. This is why OMO’s RDF exports act as both: • a functional database backup, and • a preservation format for future researchers, developers, and institutions. 5.3.4 Future services through institutional interoperability Future-proofing is not only technical. It is also about making sure that libraries, archives, rights organisations, and small labels can move into the future without needing to throw away their old systems. When a library aligns its catalogue with the SKCMDB (Slovak: hudobnadatabaza.sk) federated module of the Open Music Observatory: • the library gains a modern semantic layer, • the obsolete system gains a migration path, • future library software can import the mapped RDF as its starting point. This has already happened in Hungary and Slovakia: metadata that previously lived only inside ageing local catalogues has now been “semantic-lifted” into future-proof form. The same applies to: • music labels using old PHP/MySQL webshops, • rights organisations with 1990s-era SQL systems, • archives with hand-maintained Excel inventories, • community collections with no database at all. Once the data is mapped into the data sharing space, it becomes portable. 5.3.5 A concrete example: ALOADED, Livonian folk music, and DDEX The ALOADED proof-of-concept demonstrates future-proofing in action, because it shows how archival cultural-heritage metadata, contemporary music-industry workflows, and emerging data-space standards can all meet in a single semantic pipeline. 85
Figure 5.3: We explained this in a broader context in our Open Access Music Dataspaces – Open Music Observatory on LineCheck 2025. DOI: 10.5281/zenodo.17669739 Our workflow proceeded in the following steps: • We began with archival metadata about Livonian and Latvian folk music, including the tšitšōrlinki traditional dance song Q4169, with its associated field notes, handwritten scores, and contextual ethnographic information. • This material included a corresponding archival field recording Q4166, originally catalogued in a traditional archival system that follows a very different descriptive paradigm. • We then expressed this material using DDEX concepts Q4745, creating mappings between archival descriptions and DDEX’s contribution and release structures. This transformation enabled several outcomes: • the track could appear legally on Spotify https://open.spotify.com/track/184iSrExWt2K9z7S4FvJYd • the musical score and the archival recording became findable on Garamantas.lv and Europeana • creators and contributors could be credited transparently and correctly • rights metadata could be clearly documented • and the entire pipeline resulted in a FAIR, future-proof transformation of cultural heritage into a living, usable service We describe this workflow in more detail in A Feasibility Study With Two Datasets on the Application of the Early Stage HDTO Ontology in Practice (Antal 2025c). There, we 86
explore how early HDTO (Heritage Digital Twin Ontology) structures can support the future Europeana Data Space and the European Collaborative Cloud for Cultural Heritage, including creating digital-twin representations of cultural objects that: • have a non-commercial mechanical licence for scientific MIR processing • and a commercial streaming-safe version (e.g., Spotify) without download or AItraining permissions This demonstrates how one song can travel through different rights regimes and different infrastructures — all mediated through a coherent semantic layer. This is not only a proof-of-concept for technical interoperability. It also illustrates how a shared observatory infrastructure can enable the circulation of culturally vital but economically marginal repertoires. The Livonian ethnolinguistic community is critically endangered. Making their vocal music available in modern channels is essential for revitalisation, language learning, and community visibility. Without the shared services of the Open Music Observatory, distributing such high-cultural-value but low-market-value repertoires would be practically impossible. A listener can now: 1. legally stream a piece of folk music (Spotify, Apple Music Classical, Deezer); 2. then borrow or download the score from a library or from Garamantas.lv; 3. a researcher can analyse the non-commercially licensed version using MIR tools; 4. and anyone can explore the archival recording, contextual notes, and documentation connected to it. This is future-proof music librarianship — linking listening, learning, research, and cultural memory. One of our future development tasks is to model rights, permissions, obligations, and prohibitions directly inside the Observatory knowledge graph. We outline this in Wikibase as a Data Sharing Space: Connecting Rights, Communities, and GLAM through Federated Infrastructures (Wikidata Conference 2025) (Antal 2025e). 87
6 Federated Data Modules This chapter introduces the four federated data modules that currently constitute the backbone of the Open Music Observatory. Rather than treating “data sources” as a flat list, we follow the architectural logic established in Chapter 4and the data-to-policy pipeline in Chapter 2: data is contributed, curated, harmonised, and activated inside federated modules, each with its own governance, provenance, and semantic profile. The OpenMusE project was contractually expected to populate the Observatory’s four thematic pillars — Economy,Diversity,Society, and Innovation — through coordinated data collection, harmonisation, and processing. These contributions now materialise as a central Open Music Observatory module, complemented by three national or regional modules demonstrating how federation works in practice. Together, they show how the Observatory evolves as a network of interoperable, decentralised components rather than a single central repository. The four federated modules presented in this chapter are: • the Slovak Comprehensive Music Database (SKCMDb) — the first fully scaled national dataspace feeding the Observatory; • the Hungarian Music Database (HU-MDb) — a replication and enhancement of the Slovak model, demonstrating cross-border portability and semantic continuity; • the Finno-Ugric Data Sharing Space — an example of subsidiarity-based federation for regional, community, and low-depth archival collections; • the Open Music Observatory Core Module — the central, pan-European module containing data created by WP1, WP2, WP3, and WP4 of the OpenMusE project, and the point of semantic linkage between national and regional modules. Each module is described using the same structure: 1. Purpose and scope 2. Data inputs and contributing institutions 3. Metadata, identifiers, and semantic alignment 4. Governance and legal basis 5. Interoperability with OMO and other modules 6. Current status and next steps In the sections that follow, we provide concise summaries of the existing modules and describe how they federate with the Observatory. The temporary landing page of the OMO can be reviewed on https://dataobservatory-eu.github.io/omo-landing-page/. 88
Once the DMP is updated, all datasets described in these modules will be linked to the DMP’s Data Summaries, vocabularies, legal bases, and provenance statements, and ingestion workflows will proceed accordingly. 6.1 Slovak Comprehensive Music Database (SKCMDb) The Slovak Comprehensive Music Database (SKCMDb) is established to make Slovak music and music-related cultural assets more visible, discoverable, and interoperable across memory institutions, rights-management ecosystems, and public collections. It is available on https://hudobnadatabaza.sk/en/. 6.1.1 Purpose and scope • increase the precision of information about Slovak musical works, recordings, persons, and artefacts; • improve public accessibility to sheet music, sound recordings, books, and documents; • create a publicly accessible, linked database (“Slovak Summary Music Database”) as part of a broader Slovak Music Dataspace; • provide trusted, authoritative identifiers enabling cross-institutional coordination through VIAF, Wikidata, ISNI, library identifiers, and other persistent identifiers; • support digital curation, enrichment, and harmonisation of metadata across libraries, archives, collective management organisations, and publishers. This module embodies the subsidiarity principle: national institutions retain control over their data while contributing to a shared knowledge infrastructure compatible with the Open Music Observatory. The scope of the federated Slovak module covers: •musical works (compositions, arrangements) •recordings (audio carriers, CDs, digitised tapes, releases) •persons and organisations related to music (composers, performers, lyricists, pedagogues, musicologists, publishers, ensembles) •music-related documents (books, sheet music, archival holdings) •physical and digital artefacts in Slovak libraries and public collections •authoritative name control for Slovak creators and entities, with linkage to international authority files •cross-institutional identifiers, especially VIAF, ISNI, Wikidata Q-IDs, and internal catalogue identifiers 89
6.2.4 Governance and legal basis HuMDb governance is in formation. Current cooperation includes: •Hungarian Heritage House: stewardship of heritage and folklore collections; cooperation agreement based on the OpenMusE grant agreement between Reprex and the HHH. •House of Music Hungary: management of contemporary collections, studio workflows, and event metadata. ; cooperation agreement based on the OpenMusE grant agreement between Reprex and the HHH. •Reprex B.V.: semantic modelling, identifier strategy, mediation, and future-proofing •weCan: chatbot integration. • optional technical coordination with KeleSys developers. Each partner retains authority and control over its own datasets. The emerging model follows the Slovak Memorandum of Understanding structure but adapted to Hungarian institutions. 6.2.5 Interoperability and federation HuMDb is designed to interoperate with: • the Slovak module, using shared schemas for persons, works, and events • the Finno-Ugric Data Sharing Space, sharing heritage workflows and gazetteers • Wikidata and VIAF/ISNI ecosystems • open-source catalogue and archival systems including AtoM and Koha • industry platforms aligned with DDEX for distribution and rights workflows The Hungarian module introduces additional areas such as studio-file mediation and chatbotready knowledge-base services. 6.2.6 Status and next steps • heritage datasets have been semantically lifted following the feasibility work (Antal and Zagyva 2025) • the MZH cooperation document defines a two-month roadmap for studio files, events, and a public chatbot interface • ingestion templates and minimal metadata rules are being prepared • governance arrangements will be formalised after the first pilot integrations • next step: establishing a unified namespace and SPARQL endpoint for federation with the Open Music Observatory. 96
6.3 Finno-Ugric Data Sharing Space The Finno-Ugric Data Sharing Space (FUDSS) is the third federated module of the Open Music Observatory. The FUDSS can be accessed via https://finnougric.net/ It functions as a subsidiarity-based regional node, designed to support culturally endangered, minority, community, and heritage-rich datasets that lack the institutional structures available in larger countries. This module demonstrates how the OMO federated architecture can scale to low-resource, multicultural, multilingual, and distributed memory environments, and how open-source semantic technologies allow small organisations to participate in European data spaces. 6.3.1 Purpose and scope 1. Cultural and linguistic preservation Providing a sustainable semantic infrastructure for Finno-Ugric musical traditions, including Estonian, Finnish, Sámi, Mari, Udmurt, Komi, and Livonian materials. 2. Linking heritage and contemporary music ecosystems Connecting archival field collections to rights-aware DDEX-compatible distribution workflows and multilingual knowledge graphs. 3. Demonstrating regional federation Implementing the principles described in the CITF First Project Report and the Open Music Europe Green Paper in small and distributed heritage environments. The scope includes traditional songs, field recordings, contextual ethnographic metadata, contemporary reinterpretations, revival performances, notebook materials, archival finding aids, community-maintained materials, multilingual enrichment in Finno-Ugric languages and regional languages, and DDEX-ready metadata for selected recordings. It also includes digital twins for rights-aware distribution that separate non-commercial research uses from commercial streaming versions. 6.3.2 Data inputs Data inputs come from four primary sources. Latvian Archives of Folklore i.e., Garamantas.lv: - digitised field recordings and ethnographic notes - collector, performer, and informant metadata - archival structure (fonds, series, file, item) - settlement-level geographic data - Livonian and Latvian cross-border repertoires University research datasets: - ethnomusicology corpora - Finno-Ugric language and phonology datasets - structured vocabularies and contextual descriptions - annotations and transcriptions 97
Community organisations: - local archives of Sámi, Komi, Mari, Udmurt, Livonian, and other Finno-Ugric groups - recordings linked to cultural revitalisation - contextual information, translations, and performance metadata Reprex and Unlabel: - semantic models and reconciliation rules - enriched metadata mapped to VIAF, ISNI, Wikidata, and geographic registers - DDEX catalogue-transfer metadata - digital-twin transformations Selected distributors: - ALOADED proof-of-concept for turning archival metadata into releasable DDEX messages 6.3.3 Metadata and semantic alignment The Finno-Ugric module follows the same semantic principles used across the Open Music Observatory. Key elements: - multilingual labels and scripts - reified contribution roles for collectors, performers, informants, translators - preservation of local vocabularies instead of normalisation - settlement alignment with national and European gazetteers - archival provenance aligned with Records in Contexts (RiC-O) - alignment to CIDOC CRM, DCTERMS, EDM - equivalence mappings to OMO ontology and Wikibase schemas Industry identifiers are integrated through: - ISWC, ISRC, and DDEX fragments defined in the OMO namespace - release metadata compatible with commercial distribution pipelines - multilingual contributor roles This enables round-trip interoperability with archival systems, OMO, Wikidata, MusicBrainz, and commercial distributors where permitted. 6.3.4 Governance and legal basis Governance follows a distributed subsidiarity model. Each institution retains stewardship and decides which data to share. Rights metadata controls which versions are usable for research, public access, or distribution. Personal data follows GDPR balancing tests and cultural-heritage exemptions. Community organisations provide contextual data and approve sensitive uses. Reprex maintains the semantic layer and federated workflows. Legal bases include public-domain status, non-commercial research licences, community authorisation, and explicitly granted commercial licences. The module implements digitaltwin workflows separating versions for MIR research from versions licensed for streaming. 98
6.3.5 Interoperability and federation 6.3.6 Status and next steps The Finno-Ugric Data Sharing Space demonstrates how minority and low-resource cultural communities can participate in a European data sharing space using a lightweight, federated, rights-aware approach. It connects archival heritage, community knowledge, research datasets, and modern music-industry workflows, showing how the Open Music Observatory architecture supports both cultural preservation and contemporary reuse. 6.4 Open Music Observatory Core Module The Open Music Europe Core Module is the central, project-wide knowledge base that integrates the datasets, metadata structures, indicators, and workflows produced in WP1– WP4, in line with the Grant Agreement’s mandate to deliver a “360-degree intelligence” system for the European music ecosystem . Its function is not to act as a monolithic database, but to serve as the semantic, methodological, and interoperability backbone that connects the thematic, national, and regional modules of the Observatory. 6.4.1 Purpose and scope The Core Module provides the shared foundations required by the data-to-policy pipeline defined in Annex 1 of the Grant Agreement. It ensures that all datasets curated and produced within the project: • follow the project’s shared DMP, licensing, and data-protection rules • use common indicator definitions, harmonisation schemas, and metadata standards • expose reproducible processing workflows (survey ingestion, statistical pipelines, streaming sampling, register harmonisation) • feed policy analysis through persistent identifiers, versioning, and provenance that fulfil the project’s obligations for transparency and open policy analysis under Horizon Europe rules. It is the reference point against which national, regional, and thematic modules align their schemas, identifiers, and provenance models. 99
6.4.2 Data inputs (WP1–WP4 contributions) The Core Module ingests curated outputs from: • WP1 policy and landscape mapping (conceptual definitions, indicator families, regulatory mappings) • WP2 identification of data gaps, methodological foundations, and pan-European variable definitions • WP3 data collection instruments, including survey pipelines, national statistical harmonisations, platform-data sampling, and register-based contributions • WP4 data processing tools, ontologies, entity schemas, and automated ingestion scripts. These curated datasets correspond to the “backend datasets” referenced in the project summary: official statistics, survey participation data, rights-holder data, and streaming-service samples that power the OMO “living policy documents” 6.4.3 Metadata and semantic alignment The Core Module maintains the canonical Wikibase ontology layer for the project. It defines: • shared classes for persons, organisations, works, recordings, events, economic indicators • crosswalks between Eurostat, national statistical offices, cultural heritage vocabularies, DDEX categories, VIAF/ISNI/Wikidata identifiers • provenance and versioning rules required for reproducible scientific workflows • standardised definitions for contested domain concepts (composer, lyricist, average income, microlabel, participation index), in line with the project’s standardisation tasks. It acts as the semantic mediation layer between statistical, heritage, industry, and survey data—mirroring the federated architecture recommended in the CITF report. 6.4.4 Governance and legal basis The Core Module is governed by the consortium as a whole under Annex 1 of the Grant Agreement, with SINUS as coordinator and Reprex as technical steward for semantic modelling, ingestion, and interoperability. It is also complemented by the Consortium Agreement and the elements defined in these agreements: • The Data Management Plan of the consortium 100
• The OPA folders where the raw data and its contextual information, such as provenance information can be found. All partners contribute datasets or workflows according to their WP obligations. The Module operationalises Horizon Europe requirements on open science, open-source software, transparent methodology, and reproducibility, using the licensing, privacy, and provenance conditions defined in the DMP. 6.4.5 Interoperability and federation The Core Module is the linking hub for all other modules. It federates with: • the Slovak Comprehensive Music Database • the Hungarian Music Database (HHH + MZH tracks) • the Finno-Ugric Data Sharing Space • external infrastructures such as Wikidata, VIAF/ISNI, Europeana, the EU Open Data Portal, and the Cultural Heritage Cloud. It does so by providing the shared namespace, identifier strategy, ontological patterns, and schema alignment that allow national and regional nodes to retain autonomy while contributing to a unified European knowledge infrastructure. 6.4.6 Status and next steps The Core Module is operational and continuously expanded as WP3–WP4 pipelines mature. Next steps include: • finalising the unified indicator registry • completing automated ingestion for all WP datasets • publishing stable RDF/JSON-LD exports for long-term reproducibility • preparing the federation services that will onboard external stakeholders, as foreseen in the Grant Agreement and Amendment. 6.5 Summary and Integration This modular, federated structure reflects the principles described in Chapter 2and operationalises the architectural choices detailed in Chapter 4. The Observatory grows through interoperable contributions rather than central accumulation, aligning with the European data strategy, the Data Governance Act, and the European Interoperability Framework. 101
7 Data Collection This chapter summarises how data was collected, validated, harmonised, and processed within Open Music Europe. It connects the practical work of WP1–WP4 to the data-topolicy pipeline described in Section 2.5 and to the governing principles defined in the Data Management Plan. It also introduces the software components developed in WP4, which implement the project’s methodological requirements and prepare datasets for integration into the Open Music Observatory. Because the DMP has not yet been updated to reflect all datasets collected during implementation, this chapter includes placeholders indicating where the DMP manager must add or revise content. Once the DMP is updated, this chapter will be synchronised with it and will become the operational reference for data ingestion. After the DMP is updated, all data-collection workflows and summaries presented here will be aligned with the approved Data Summaries and metadata structures. ÁNot updated Because of the serious delay in WP6 with the updating of the Data Management Plan, this section should not be reviewed until harmonised with the DMP. 7.1 Overview of the Data-Collection Framework Data collection in Open Music Europe followed the pipeline logic defined in Section 2.5. Each thematic work package identified indicators and conceptual models (D1.1, D2.1, D3.1), which determined the data inputs required for analysis. These sources were then evaluated based on accessibility, legal compliance, interoperability, and their relevance to indicator construction. The data-collection effort covered four categories: • administrative and register data • survey data • statistical and economic data • platform and streaming data 102
The workflows for collecting, accessing, and harmonising these sources were originally documented in internal working files and the first version of the DMP. They must now be consolidated into the updated DMP. The DMP manager must add all final sources, access conditions, and descriptions to the Data Summaries section for traceability. 7.2 Administrative and Register Data Administrative datasets provided essential inputs to WP1 and WP3, where economic valuation, labour structures, and participation indicators required high-resolution administrative signals. These workflows included: • accessing royalty and licensing records from CMOs • collecting grant and programme data from ministries • extracting public business and organisation registers • harmonising venue, festival, and event registers • obtaining public-sector microdata in the Slovak pilot, including historical KULT survey data These datasets formed sampling frames, validation structures, and reference datasets for linking surveys and platform data to legal and organisational entities. The DMP must describe each administrative source, its controller, its legal basis, and its access pathway. 7.3 Survey Data Survey data in WP2 and WP3 filled gaps that administrative and platform data could not address. Survey work included: • enterprise surveys of MSMEs in the music sector • personal surveys on participation, wellbeing, and music behaviour • experimental modules on diversity, mobility, and cultural citizenship • harmonisation with national statistical surveys, especially the Slovak KULT survey Survey workflows required careful attention to: • sampling-frame construction • questionnaire metadata • informed-consent procedures • pseudonymisation and controlled processing • harmonisation using SDMX, DDI, and GSIM concepts The DMP must contain questionnaire versions, variable lists, consent protocols, and pseudonymisation details for each survey. 103
7.4 Statistical and Economic Data Statistical sources enabled European and national comparability across indicators. These included: • national accounts and satellite cultural accounts • labour-force data • business demography and structural-business statistics • external-trade and export data • cultural consumption and household budget surveys • price indices and cost-structure information These datasets were central to the economic modelling in WP1, circulation and diversity indicators in WP2, and societal-impact work in WP3. The DMP must include references, licences, and access conditions for all statistical datasets used. 7.5 Platform and Streaming Data Platform datasets (WP1, WP2, WP4) were collected via API-based sampling and automated scripts, including: • Spotify API samples (popularity, metadata, audio features) • YouTube API samples • playlist-localisation datasets • automated crawlers for repertoire discovery These data supported: - digital-market structure analysis - validation of economic indicators - repertoire and rights linkage - circulation and localisation studies Licensing and terms-of-service constraints mean these datasets have strict reuse limitations. The DMP must describe permitted uses and restrictions for each platform dataset. 7.6 Processing and Harmonisation After collection, datasets underwent several processing stages aligned with the architecture described in Chapter 4. These stages included: • cleaning and transformation • pseudonymisation of personal data 104
• cross-linking with authority files and identifiers • metadata enrichment • structural harmonisation (SDMX, DataCite, DDI) • conversion to formats suitable for ingestion into the OMO The openmusic-pipeline (WP4) implemented these processes in R, ensuring reproducibility and alignment with FAIR and OPA principles. The updated DMP must list controlled vocabularies, classifications, and harmonisation rules used. 7.7 Software Components Developed in WP4 WP4 created a suite of open-source tools that implement the pipeline and prepare data for the OMO. These are described here briefly, with technical detail deferred to annexes. 7.7.1 Data-Ingestion Tools These tools handle imports from Excel, CSV, SQL exports, APIs, and legacy systems. They include: - connectors for CMO and/or ministry datasets - survey-import scripts - API wrappers for streaming platforms 7.7.2 Validation and Reconciliation Tools These tools perform semantic alignment and quality checks: - authority-control reconciliation (ISNI, VIAF, ORCID, corporate registries) - SPARQL-based constraint checks - duplicate detection and entity merging tools 7.7.3 Harmonisation and Metadata Tools These implement the project’s semantic rules: • SDMX structure builders • DDI variable metadata generators • DataCite dataset metadata templates • vocabulary management utilities 105
Unfortunately, the music industry has long missed access to reliable, open registers. The reasons for this are beyond the scope of this report, but we highlight that the underlying reasons for closed and not interoperable registers are deeply rooted in the conflicts of interests among different sub-sectors of music and are unlikely to be solved in a short time. Therefore, music enterprises, researchers, professionals, and curators will need identification services and identity brokerage services for a long time. Creating and maintaining high-quality registers require significant professional and financial commitments, and they can form a vital service of a future European Music Observatory. Currently, we are experimenting with three service levels in the Open Music Observatory. • We create our own transparent and interoperable identifiers within the OMO for persons and their groups (ensembles, bands, orchestras, associations…), legal persons (music businesses, collective rights management agencies, …), events (recording, composing, performing events, festivals, conferences, …), musical works and their manifestation (books, works, recordings, sheets.) • We create integrity brokerage services and middle-term identification via Wikibase and Wikidata. Our identifiers are connected to middle-term Wikidata and Wikibase QIDs, which also serve as graph nodes to registry, library, collections, and industry-specific identifiers. • We are piloting data improvement services that can find erroneous identifiers or add correct identifiers to various datasets. 8.3.2 Open and persistent identifiers In line with the practice of the Netherlands, we prefer the use of the following identifiers: ISNI: preferred persistent identifier for names of people and groups. The use of ISNI is also preferred by Apple Music, Spotify, and as a pilot it was introduced by Teosto, the Finnish national collective management society; it is being considered in many use cases for adoption in all CISAC societies. ISNI is the ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. (Camp, Lieber, and IFLA 2022) For legal persons, we are discussing the terms to use the OpenCorporates ID, because many organisations at this point do not have an ISNI. ORCiD: preferred persistent identifiers for music researchers and scholars. This is in line with the Horizon Europe and the European Open Science Cloud recommendations; ORCiD itself only adds functionality to ISNI; i.e. each ORCiD ID is at the same time registered as an ISNI. VIAF: VIAF is the shared authority file of national libraries. It offers more services than ISNI and includes an ISNI for the author. DOI: we use the Digital Object Identifier for publicly released documents. ISBN: We issue ISBN identifiers for long-form publications of our partners. (ISO 2017c) 112
8.3.3 Not open, music-industry specific identifiers Book and music sheet publishing uses the ISBN and ISWN, professional and magazines and scholarly music journals use the ISSN, and the music rights management uses ISRC and ISWC. These standards usually resolve an identifier to some network location where metadata or the object itself can be found. There are many advantages and disadvantages of this model. For example, the ISWC identification of musical works is the backbone of copyright management, and it is a closed and consistent system developed over many decades by the member organisations of CISAC. The downside of this closed system is that the metadata about the works identified by ISWC is strictly available only to CISAC member societies. While CISAC offers an API for the individual lookup of ISWC for one example of a musical work, currently it does not allow bulk access to the registered data. We have already started a discussion with some music industry registers about connecting the Open Music Observatory to their systems. We are planning to present our proposals on the CISAC Good Governance seminar to be held in December 2024. Musical works ISWC: nternational Standard Musical Work Code is a unique identifier for musical works. It is adopted as international standard ISO 15707 (ISO 2022). OpenCollectons ID: Our ID for music works (only if we publish data about them.) Sound recordings ISRC: The International Standard Recording Code (ISRC) is the international identification system for sound recordings and music video recordings. (ISO 2019b; International ISRC Registration Authority 2021) OpenCollectons ID: Our ID for sound recordings (only if we publish data about them.) Music sheets ISWN: The International Standard Music Number currently identifies published music sheets (ISO 2022). ISBN-13: Before the introduction of ISWN, published sheets were identified by the ISBN book identifier. ISBN-10: The older format of the ISBN book identifier, which predates both the ISWN and the 13-digit ISBN used to identify music sheets. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. This is our preferred identifier for not published sheets. (ISO 2017c) For unpublished works, our preference is the use of the brand-new ISO-standard ISCC because it was designed precisely for the use case we were looking for. It is free to generate, generated from digital content (or its digital copy), and can connect various local or lesserused identifiers. Datasets DOI: DOIs are assigned to each distribution of a dataset. As datasets are often continuously filled, these datasets will have periodic versions with versioned DOIs (from Zenodo.) OpenCollectons ID: Our ID for unversioned (continous) datasets, pointing to the latest available version of the data. 113
Codebooks URI: Whenever possible, we use standard codebooks of SDMX or Eurostat, and provide a URI to the codebook, and provide dereferencing to the codebook definition. OpenCollectons ID: Our ID for our codebooks, regardless if they are same as the SDMX/Eurostat standards, or we create a non-standard coding for a novel dataset. Questionbank URI: Whenever possible, we use standard questionnaires, and provide a URI to the codebook, and provide dereferencing to the DDI questionnaire item definitions. OpenCollectons ID: Our ID for questionbank items. 8.3.4 Lyrics In many genres, lyrics are very important parts of a musical work, and there is a growing demand and need to provide or analyse the lyrics of the work. For example, in our X, we want to create location-aware music services and encourage the public performance of music made in Bratislava or music somehow specific to Bratislava within the public places or radio stations of Bratislava. One possible semantic connection to this environment is that a song is about Bratislava (Berlin, Paris, or Germany.) Access to the lyrics part of the music is not straightforward, mainly because the lyrics may be arranged from a literary work. We see lyrics identification and semantic analysis as the next immediate step to our location-aware application, for which we are looking for good industry solutions. In many cases, we will likely need to rely on the ISCC code as a temporary identifier for lyrics databases that were not available in a licensed format earlier. 8.3.5 ISCC The Open Music Observatory will start to implement the newest ISO-standard open identifier, the ISCC-CODE. ISCC is inverting the principle of a centralised register. It generates the ISCC code from the digital content object itself, therefore no third-party lookup is needed for finding the identifier of the object. ISCC registration becomes necessary when an ISCC code needs to be globally unique, publicly discoverable, resolvable, owned or authenticated. While these features inevitably require some kind of registry, not all of them require a centralised institutional registry. The ISCC specifies the necessary protocols to implement the aforementioned features in a decentralised, federated environment and across multiple public blockchains. Given a registered ISCC code, an application can unambiguously determine on what blockchain (if any), by which account, and at what time an ISCC has been registered. Registered ISCC codes refer to an authoritative public blockchain network. This indicator is part of the ISCC Code itself, such that codes registered on different networks cannot collide. This guarantees uniqueness of ISCC codes across multiple blockchains. Ownership of ISCC codes (not the identified content) is granted to the signatory of the first transaction for a given ISCC code on the corresponding blockchain. 114
As such the ISCC fulfils a distinct role and is not a replacement for established identifiers. Rather it is designed as an umbrella standard to augment established identifiers with enhanced algorithmic features. It can be used in the metadata of existing standards or support discoverability (reverse lookup). We will use for precisely this application: whenever we receive content that is not identified by a DOI,ISNI,ISWC,ISRC, or other standard identifier, we will assign an OMO identifier and enhance it with the ISCC features. This will help later linking to the preferred global, persistent identifiers. 8.3.6 OMO Identifiers We create our own identifiers for persons and things. We follow the practice of the Dutch national archives in the creation of PIDs, and we make them URIs following the W3C recommendation. music.dataobservatory.eu/{type}/{concept}/{reference} For {type} we utilise the following definitions: - {id}: an identifier {type} for dereferenced identifiers. •{doc}: a documentation {type} for the documentation of persons and objects. •{def}: a definition {type} for ontologies. For {concept} we utilise three categories: 115
•music.dataobservatory.eu/{id}/{person}/{reference} for persons, in order to synchronise with national and international name spaces. •music.dataobservatory.eu/{id}/{place}/{reference} for places, in order to synchronise with national and international name spaces. •music.dataobservatory.eu/{id}/{oc}/{reference} an other objects, such as musical work, a sound recording, a group a persons. 116
9 Data Improvement & Innovation The music sector was one of the early adopters of digitisation and is a highly data-driven sector of the economy. Because it relies on data, business and public policy problems often accompany data problems. In the previous section, we have shown how we aim to increase the data available for the sector. Now, we focus on improving the data’s quality and usability. ÁNot updated This section was created at an early planning stage, and had not yet been updated. This is no longer applicable, and should not be read or quoted. However, this part is clearly articulated in our green paper’s Fixing Music Data at the Source chapter (Antal 2025d). 9.1 Value-Added Data Services 9.1.1 Data Sharing “Data sharing” means securely sharing data among parties who do not want to expose their data to third parties or protect the personal data in the datasets. Agreeing and organising data sharing legally, semantically, syntactically, and technically can be challenging. This is the role of the Open Music Dataspace behind our observatory. Data sharing can reduce the redundancy in costly metadata collection, improvement, linking, updating activities, which are currently done often without coordination paralell among various authoritative database managers (for example, VIAF for libraries, ISNI, and ISRC or ISWC.) Data sharing can also greatly reduce the redundancy of parallel work at the level of collection managers, like individual libraries, publishers, labels, collective management organisations. The data sharing infrastructure behind our dataspace is the Reprexbase system, which is an extension of the Wikibase system using various open-source (and, in a few cases, non-opensource) software components to connect the Wikibase system with music sector databases and data sources. Given the sensitive nature of the data we handle, including business confidential and GDPRprotected data, we maintain strict segregation. Data batches from stakeholders are kept in separate instances and are integrated only after thorough review by the data protection officer and curatorial team, ensuring the highest level of data security. 117
In our Slovak prototype, we keep SOZA’s data in an insulated instance because the copyright management organisation must not release GDPR-protected and business-confidential information. Some of the data needed for NERD operations or the establishment of the Slovak Comprehensive Music Database is then sent to a joint instance concerning those data subjects (i.e. authors or their heirs) who agree with our data handling. This is where they meet public catalogue and database data from public libraries, open knowledge graphs, and the Slovak Music Center. The data that should be made public is then further exported to Wikibase Cloud, where it becomes public and available for all stakeholders. From Wikibase, it is also synchronised with Wikidata, the world’s largest open knowledge graph. 9.1.2 Fix-the-data “Fix-the-data” means improving the data quality by finding or imputing missing values or finding and replacing erroneous data entries. In terms of metadata, adding further machineactionable information to already existing datasets can improve their usability. The fix-the-data service can mean replacing missing our outdated metadata (such as a name change of a natural person or a corporate body), or recalculating aggregated accounting or statistical data after base change, or forecasting data that is not yet available. ĎTip Our fix-the-data services do not increase the size of the data available to our partners, but it increases the quality of their datasets or databases. 9.1.3 Data Linking “Data linking”, data fusion, or data matching means correctly joining data from different datasets (data sources.) Many fix-the-data problems initially arise from imperfect data linking, for example, mistakes in currency rates, units of measures, coding of geographical entities, misplaced decimal delimiters on the level of data, or misunderstandings of the meaning of “artist income” or “popularity score”, or other non-self evident variables. An even more subtle problem is joining data from two questionnaire surveys created with different sampling algorithms and different standard (measurement) errors. ĎTip Data linking or data fusion is a way to join many small databases into a large, federated dataset. This way, relatively small music organisatiosn can benefit from access to big data. 118
9.1.4 Registration services In Open Music Europe, other tasks deal with the policy problem plaguing the music industry: even though it needs access to an exceptionally high number of registers (due to the fragmentation of the copyright and several neighbouring rights), access to such registers is limited or impossible. Often, the registers carry legacy problems that make them less functional in trustworthy data and AI systems. Aregister is a document [in modern usage, usually a database], in which data are entered in a formal manner by a statutory authority (ISO 2017b). In statistical data collection a “register aims to be a complete list of the objects in a specific group of objects or population.” (Anders and Britt 2007). Statistical data collection and rights management are just two service areas whose workflows depend on well-functioning and accessible registers. The statistical business register is an essential tool for creating survey frames or sample frames, in other words, to organise statistical data collection. A copyright or neighbouring right register is necessary to organise royalty collection. ĹNote A statistical register is necessary to decide who should get a data request: • For a sample survey, the register is used to draw a lottery of population members who will be invited to provide data. • In a census-type survey, all registered members of the population, for example, all music labels, will receive an invitation to an interview or form. • In the case of a register-based survey, all members of the register, for example, all collective management societies in the territory, will be requested to send data directly from their databases. In other work packages of the Open Music Europe project, we are experimenting with statistical data coordination among the music sector and statistical authorities. Without recalling the details here, as digitisation exponentially increases the amount of structured data in the private sector, it is a growing trend in statistical innovation to rely on data held by the private sector to make more granular or timely official statistics. For consumer spending statistics, costly and imprecise surveys of randomly selected citizens putting their purchases in a diary, some statistical authorities directly process data from cash registers or credit card spending. We envision a similar statistical collaboration among statistical offices and collective rights management organisations because it is easier to report music royalty accounts than to ask musicians to talk about their complex income streams in interviews or on questionnaires. We see the role of the Open Music Observatory in providing a methodology and digital data infrastructure for such statistical collaboration. In other work package tasks, SOZA and Reprex will create so-called satellite business registers to harmonise the data collection of the observatory with the Slovak statistical authority. More about this work: (Antal 2023) 119
Such services, similar to data linking and some new services that will build on the data and the data API of the Open Music Observatory, rely on the provision of technical services for registration. The Open Music Observatory has its register, too. Registration is a costly data service with vast economies of scale, so providing more affordable registration services for the European music sector could be an important service. ⊠In our piloting phase, we rely on cooperation with the Slovak National Library to test the usability of the VIAF authority file system for identifying names. ⊠Our dataset distributions use the DOIs from Zenodo, which also provides our longterm archive. ⊠For certain assets, mainly photographs and scanned documents, we rely on the new ISCC registration, a long-term solution for some music industry applications. ⊠Reprex registered an imprint, the Digital Music Observatory, to place long-form publications as books into library systems. ⊠We are investigating the costs and benefits of finding an ISNI registrar partner or creating a roadmap for making the Open Music Observatory a registrar itself. 9.2 Use Cases In 2023 the Open Music Europe project applied for the Module A of the Horizon Results Booster (HRB) provided by Trust-IT Services�. The HRB aims to provide a tangible contribution to the dissemination of results and recommendations of research projects related to the European Commission Priority areas. 120
Figure 9.1: app 9.2.1 Data Health Services for Collective Management Entity linking and data linking are among the biggest technical problems in rights management. Because music authors, producers, and performers have three royalty streams and do not share an interoperable registry, the connection of musical works (compositions, ideally identified by an ISWC code), their sound recording manifestations (identified on all digital services with and ISRC code), and the various identifiers of performers require costly manual and technical identification. There are numerous projects underway in the music industry to resolve this problem going forward. In the United Kingdom, PRS’s Nexus programme� is developing a solution with the provisioning of preliminary ISWC registration to keep the recording and composition connected from the birth of a new recording. The Open Music Europe project, on the other hand, is pioneering a different route for already existing sound recordings, with the linking of public sector catalogues of heritage and library collections with rights management information; particularly with relying on the VIAF shared authority files. SOZA and Reprex are expected to present their MVP on the CISAC Good Governance seminar in December 2025. Modern registers typically assign a unique identifier, known as a URI, to their data subjects (our registered objects). A ‘Cool URI’, which resembles a URL, offers a practical advantage. When used as a URL, it generates a human-readable HTML file about the registered person or object. This can be particularly useful when processed by a graph application, as it 121
We do not consider that the system has wider risks or negative impacts. The algorithm is designed to cure sources of data biases that result in a late or missed payment for some rightsholders. 128
10 Data Catalogue The Open Music Observatory curates, maintains, and disseminates a data catalogue with the resources within the data catalogue: individual datasets and their series and API endpoints where the data can be queried in a custom format. In creating our data infrastructure, we considered the specifications of our dissemination nodes, which provide our data with a wide range of interoperability and easy access: the EU Open Data Portal, Europeana, Wikibase Cloud and Wikidata. From a thematic point of view, we relied on the definition of the EMO feasibility study, and created topical pillars (Section 10.2). The data curators of the Observatory Stakeholder Network (see Annex) and for the duration of the Open Music Europe project, the work packages (WP1-4 represent each “pillar”) can define and provide datasets or data series according to their topical collection guidelines (Section 10.1). Not updated ÁWarning This section was created at an early planning stage, and had not yet been updated. Unfortunately, due to the problems of the WP6 Data Management Plan task we cannot yet show how we will fill up the pillars of the observatory. A data catalogue formally is a metadata dataset: a dataset on information about our available datasets and their downloadable or queriable distributions. It follows the global World Wide Web DCAT standard. DCAT is an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. This document defines the schema and provides examples for its use (Albertoni et al. 2020). It is a global standard, which was further extended and specified for the release of statistical datasets (StatDCAT-AP) and for the needs of the EU Open Data Portal (DCAT-AP). (Sofou and Dragan 2019; Fragkou 2023) These extensions provide further metadata and organisations standards, but essentially they do not change the definition of the global standards. A data catalogue (dcat:Catalog) represents a catalogue, which is itself a dataset in which each individual item is a metadata record describing some resource: a description of a dataset, a data service, or other type of resource. dcat:Dataset represents a collection of data, published or curated by a single agent or identifiable community. We currently support two types of datasets: statistical datasets that conform to the datacube definition of SDMX, or collection datasets for microdata, which contain non-aggregated, structured data representing some unity criteria, for example, music works and recordings that have been present in the official radio charts of a given country. 129
The _dataset_, similar to a musical or literary work, is an abstract concept which can be used, downloaded, and stored in its manifestation. For a musical work, a manifestation may be a sound recording or music sheet; for a dataset, it is a distribution. A URI identifies a dataset; the URI does not allow the downloading of the dataset, because it refers to the abstract idea of the dataset; the URL for downloading the dataset belongs to the individual distributions. dcat:Distribution represents an accessible form of a dataset, such as a downloadable file. When the same dataset is distributed in different file formats (for example, CSV and SPSS files), each distribution is listed in the catalogue separately with a separate download link. Each distribution has its own URL where the dataset can be downloaded. Figure 10.1: The music.dataobservatory.eu/tag/music-economy/ URL lists the downloadable datasets on the Open Music Observatory website. They can be found on EU Open Data Portal, too. In the first days after launching our new service, around 1-10 June 2024, the datasets may be missing from the EU Open Data Portal, which is changing in these days its complete backend, and may have some backlog in accepting our datasets. dcat:DataService represents a collection of operations accessible through an interface (API) that provides access to one or more datasets or data processing functions. Our datasets are accessible on different platforms with their own datasets, and our internal data-sharing space also has its API. As data is added to the different platforms (EU Open Data Portal for statistical and microdata datasets, Europeana for collections dataset, Wikibase Cloud for further microdata, metadata and collections, and Reprexbase for confidential microdata and collections), we are updating the catalogue with the DataService entries. dcat:DatasetSeries is a dataset that represents a collection of datasets that are published separately but share some characteristics that group them; for example, a (play)list of sound recordings that were present in the weekly charts or the annual budget of an institution. A time series dataset is usually not defined as a data series, but the new time observations are added to an updated distribution of the time series dataset. Stakeholders who provide data to the Open Music Observatory can commit to making a data series; however, we only define a data series when we have at least two items available from the series. dcat:CatalogRecord represents a metadata record in the catalogue, primarily concerning the registration information, such as who added the record and when. 130
10.1 Collection Guidelines In short, we collect data about music. The initial data collection guidelines of the Open Music Observatory are derived from the EMO Feasibility study. We see them as a starting point for further discussion with the Observatory Stakeholder Network. ⊠Statistical data which is defined as cultural statistics of any European Economic Area and EU candidate statistical office or by a representative European or international music organisation. ⊠Statistical data (indicators and their datasets) defined by, or requested by members of the Observatory Stakeholder Network. ⊠Datasets about information gaps identified by the EMO feasibility study. ⊠Records of questionnaires, question banks, and any structured datasets used for the creation of the statistical datasets above. ⊠Collection datasets about musical works and their manifestations in sound recordings or musical sheets. ⊠Collection datasets about music events, including events of composition, recording, or live performance. ⊠Encyclopaedic, demographic, biographical data about music professionals and music enterprises. ⊠Collection datasets about books, publications, statutes and laws, standards related to music. The EMO feasibility study Curators are forming collections with the application of unity criteria which allow them to decide which musical work, sound recording, music enterprise or person is included in a collection list. The curators are responsibility for the comprehensive application of the unity criteria and ensuring that their collections are up-to-date (Wickett et al. 2013). Some examples of music data curation Hitlists use some kind of popularity metrics, and they follow rigorous rules which sound recordings are included every week, or year. Collective rights management organisations create comprehensive lists of works and sound recordings registered for rights protection and exploitation. Statistical agencies create business registers to carry out data collection. Who can curate our datasets? Any music professional or scholar can curate datasets in agreement with our Collection Guidelines. The quality review mechanisms will be set by the Observatory Stakeholder Network from a content point of view, and the Open Music Data Exchange from a technical point of view. 131
10.2 Topical Pillars Figure 10.2: The extended five pillars, with sustainability added. 132
ĹNote The suggested four-pillar model would categorise data-collection and analysis along the following lines: • Measure the contribution of music to the EU’s economic and legal environment, from a systemic perspective (Pillar 1). • Monitor the cross-border flows of repertoire, the mobility of artists and diversity (national, linguistic, genre-based) (Pillar 2). • Assess music’s impact on society and citizenship: how audiences access and consume music; how citizens participate in professional and not-for-profit music activities; the scale, value and quality of music education and training (Pillar 3). • Provide a framework to develop prospective research on the future of the music sector, supporting innovation and developing understanding of emerging practices from various perspectives (business, tech, policy) (Pillar 4)(European Commission et al. 2020, p30). 10.2.1 Music Economy ĹNote Main potential data-collection and research areas identified at this stage: � Macroeconomic patterns and trends (e.g. employment, revenue, competition) � Value chain mapping and analysis (e.g. characteristics of music organisations, copyright collection, collective management, remuneration of artists, spill-over effects) � Legal aspects (e.g. tax, labour laws, social security, contracts, case law) � Business regulations (e.g. live music regulations, consumer protection, licensing, anti-piracy rules) (European Commission et al. 2020, p114) According to the EMO feasibility study, one “of the key findings of the AB music working group report was a substantial appetite for cross-sectoral, neutral and comparable data on the music business at EU level. While recent studies (e.g. EY “Creating Growth” study) have attempted to measure the impact of music on the EU’s economy, systematic and comprehensive metrics do not exist at this stage.” (European Commission et al. 2020, p113) Deliverable D1.1 of Open Music Europe, Economy of Music in Europe: Methods and Indicators identifies critical research questions, data sources and gaps,and data collection methods regarding the economy of music in Europe . (Antal, Kmety Barteková, and Remeňová 2023) 133
The deliverable begins by reviewing definitions of “the music industry”, the categorisation of musical activities within the system of national accounts (SNA) and statistical classifications of economic activity (ISIC and NACE), and the three primary income streams within the music industry (the live music, author or publishing, and recording streams). It then turns to the topic of value, first identifying the types of value created by musical activity and then considering legal and economic dimensions of valuation. 134
Figure 10.3: Our website documents with visualisations each dataset, apart from providing links to the latest distribution downloads with visualisations (in zip) or the access points on the various dissemination nodes. The illustration is an experimental dataset from our background CEEMID catalogue. After introducing the concept of mixed enterprise and personal surveying as a means of improving insight on informal economic activity in the sector, the deliverable identifies data gaps relevant to national policy in our pilot study target country of Slovakia, critically reviews the data gaps relevant to EU-level policy first identified in the EMO feasibility study,and proposes data collection methods appropriate to filling specified data gaps. 10.2.2 Music Diversity According to the EMO feasibility study, “creating reliable tools to monitor what kind of repertoire circulates on digital platforms or via radio will require access to vast amounts of data from Digital Service providers (DSPs) or third party aggregators. The notion of European repertoire has to be clarified and very well defined; notion of language, of origin, of nationality, country of production, genres, and it should not be limited to the language sung in a given song. […] A European Music Observatory should also look into the possibility to collect regular data on the circulation of European repertoire at song and/or artist level, considering live performance/radio/ digital use, which will be available at a weekly/monthly/yearly basis to the music sector.” (European Commission et al. 2020, p34) The Open Music Europe project is developing two tools for capturing and turning the aforementioned data into informative indicators. In WP1, the project is developing big data statistical sampling algorithms to avoid the need for “access to vast amounts of data from Digital Service providers (DSPs)”. WP2 is working on a taxonomy and GDPR-conform representation of the “notion of language, of origin, of nationality, country of production, genres” based on our background (Antal 2020b). The result of this work will be the Slovak Comprehensive Music Database, which will create clear taxonomies and allow users and software applications to determine aspects of “Slovakness” for each sound recording. 135
The possibility “to collect regular data on the circulation of European repertoire at song and/or artist level” is an issue that must be addressed with strict adherence to GDPR. We are developing an opt-in, opt-out mechanism in Slovakia that will be replicable in all other jurisdictions in accordance with GDPR. (For potential non-European artists, we will apply GDPR, too.) ĹNote Main data-collection and research areas identified at this stage: � Cross-border circulation of works/repertoire (e.g. building common definition and indicators, mapping of cross-border access, sales and consumption flows � Cross-border mobility of artists and professionals (e.g. cross-border live performances, mobility of professionals, international music events) � Cultural diversity aspects (e.g. languages, genres, types of productions) � Legal aspects (freedom of movement, state aid, etc.) (European Commission et al. 2020, p115) An interesting opportunity here is that many European countries already collect information on subsidised music operators (e.g. associations or not-for-profit projects), not to mention the wealth of information available through Creative Europe supported initiatives, which could provide this Pillar with interesting data. (European Commission et al. 2020, p116) 10.2.3 Music Society Regarding music and society, D3.1 considers the reuse of various survey programs with retrospective survey harmonisation, and WP3 is planning to conduct surveys in 2025. Data will be added to the catalogue as it is becoming available. ĹNote Main data-collection and research areas identified at this stage: �Education, training, personal development �Audiences (music consumption, interaction, participation in music events, etc.) �Music and society (not-for-profit sector, associations, social inclusion, amateur music, heritage, participation in music) �Normative Aspects (broadcasting quota rules, diversity promotion schemes, freedom of speech rules) �Music and the environment (carbon footprint of venues, touring, festivals, merchandise manufacture, streaming services; issues around noise/neighbourhood impacts; good practice in these areas). (European Commission et al. 2020, p116) 136
10.2.4 Innovation The definition of the Innovation pillar in the EMO feasibility study is more a topic to be covered than a data need description. This pillar is less data-driven in that it will rely mostly on research conducted on topics relating to changes in the market place, new business models, disruptive technologies, etc. A European Music Observatory will have the latitude to pick certain topics based on priorities and input from sectoral stakeholders. An EMO should consider setting up an “innovation experts’ advisory committee,” constituted of respected professionals in their field who are known for their forward thinking views, to help identify key themes to be studied. (European Commission et al. 2020, p37) We will initiate an informal music innovation expert’s roundtable to discuss potential data needs in this pillar. 10.2.5 Sustainability In the EMO feasibility study the definition of sustainability was mentioned among the innovation topics. Because of the triple transition, introducing the Corporate Social Responsibility Directive and the European Sustainability Reporting Standards have increased the interest and need in sustainability data; we decided to create a separate topical pillar for environmental and social sustainability, or governance indicators (ESG.) We will publish datasets that will be used in the value-added service described in Section 9.2.2. 137
Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network Name of the stakeholder: □Corporate (institutional) name: This is mandatory for legal persons and groups; also, please provide a contact person. □given name: mandatory for natural persons □family name: mandatory for natural persons Legal status of the stakeholder: ⊠private person: your name will be made public among the members; we also ask to provide a persistent ID (ISNI, ORCiD or VIAF.) ⊠Legal person: your name and legal person ID will be made public among the members, but not the contact person. Please provide ISNI, or OpenCorporates ID. □other association or group without legal personality: your name and persistent ID will be made public among the members, but not that of the contact person. Please provide ISNI identifier. Logo or icon of the stakeholder: □not mandatory, but if provided, we will make it public Online contact details of the stakeholder: □Official website: we will make it public if provided □LinkedIn page: we will make it public if provided □Facebook page: we will make it public if provided □YouTube channel: we will make it public if provided □Instagram account: we will make it public if provided Permanent residence or seat of the stakeholder: ⊠Only the country and municipality will be made public □Please provide full postal address Official email address of the stakeholder: 144
□only used for invitations to stakeholder meetings, giving and revoking data handling; never made public. Contact telephone number of the stakeholder: □is only used to clarify potential problems and consents; it is never made public and is not used unless necessary. For the intake, we will also ask for a few lines of statement of interest in the Observatory Stakeholder Network and topic interests for data if they apply. We will make this information public, too. We are also happy to create an intake interview and publish it on our website to allow the members of the Observatory Stakeholder Network to get familiar with each others ideas and interests. ĺImportant We will provide in the final deliverable a link for the official intake to the stakeholder network. I would like to invite SOZA, Hudobné centrum, Aloaded, and HearDis! as first partners. Open Music Data Exchange ĺImportant This advisory body will not be open for invitations. The representative pan-European stakeholders can nominate here their technical providers to consult the technical aspects of data exchange. 145
SKCMDb: Slovak Comprehensive Music Database The Slovak Comprehensive Music Database (SKCMDb) is a national initiative aimed at making Slovak music more accessible, discoverable, and usable across libraries, archives, streaming services, and rights management organisations. It connects scores, recordings, and metadata using open standards and collaborative governance. As a functional module of the Open Music Observatory, the SKCMDb also serves as a testbed for developing shared data services, addressing the conceptual models, workflows, and governance rules required to link diverse music stakeholders. ‘ The SKCMDb is supported by a data-sharing space consisting of both shared and private databases. The data-sharing space currently comprises the following initial components: •Slovak Metadata Database: A database that facilitates connections between various Slovak stakeholders’ systems. •SKCMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 146
•SKCMDb HC-SOZA (private): A database used exclusively for rights management and library management, governed by an agreement between the Slovak Music Centre and SOZA. •SKCMDb HF (private): A technical dataset created to provide improvements, corrections, and enrichments for the Hudobný fond. The Slovak Metadata Database serves as a support layer that is partly public and partly private. Its metadata definitions and descriptive metadata are exported into the SKCMDb databases as needed and permitted. The Slovak Metadata Database The Slovak Metadata Database is developed in alignment with the metadata framework of the Open Music Observatory. •Ontological and thesauri patterns: Reuses standardized or widely adopted vocabularies. •Conceptualizations and definitions: Includes concept definitions, thesauri, and other elements developed specifically for the SKCMDb. •Public permanent identifiers: Uses identifiers that are public or can be made public. The metadata layer is generally licensed under CC0, though in some cases other licenses are used (for example, CC-BY). ĹNote Examples: • The definitions of musical work,printed sheet music, and the is score of relationship allow the description of connections between an abstract musical work—such as Bella’s Missa in C—and its actual printed manifestations. • The VIAF identifier 2737220 identifies Ján Levoslav Bella’s compositions across library systems. Slovak Comprehensive Music Database (public) The Slovak Comprehensive Music Database is a linked open database published by the Slovak Music Centre. It integrates elements from the Hudobné centrum’s own databases along with data made public by SOZA, Hudobný fond, and other organizations. The database is distributed under various Creative Commons licenses that allow both commercial and non-profit use. 147
The primary aim of the SKCMDb is to reduce the cost of maintaining public and private services that enhance the circulation, availability, visibility, and legally licensed use of Slovak music. Our licensing policies are designed to enable the widest possible use of the data while protecting the investments required to maintain registers, standards, and data integrity. Slovak Comprehensive Music Database (private) The private components of the SKCMDb consist of databases where the data is not intended for public sharing but is used to enhance rights management, music information services, library operations, or other specialized applications. These databases are maintained under agreements between the participating parties. Access to these private datasets serves specific, well-defined purposes and is governed by strict rules. Availability to third parties is determined solely at the discretion of the data owners and may vary depending on contractual or legal obligations. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category (e.g., a national musical works register). Collections typically rely on registers to identify works unambiguously and avoid duplication. Both are documented in structured datasets containing standard identifiers such as ISRC, ISWC, or ISMN codes. In the case of statistical data, microdata often refers to survey instruments and responses curated under defined methodological rules. •Metadata: Relevant elements from the Slovak Metadata Database that support the use of collections or registers. The SKCMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models 148
and data structures, enabling automatic aggregation. All datasets are available with RDF annotation and can be exported in all standard RDF serializations. Microdata is intended for institutional and professional use, not for the general public. It is annotated with standardized metadata suitable for applications such as music library cataloguing, distribution platforms, and rights management systems. ĹNote Example The musical work String Quartet in B-flat Major (composed by Ján Levoslav Bella) is described in a dataset that includes two publicly available printed scores and a publicly available sound recording. •Graphical view: Navigate and contribute to detailed entries on musical works, sound recordings, and related assets. � Explore on Wikibase •Semantic view: Export structured data in XML, JSON-LD, Turtle, or NTriples formats for reuse in research or digital projects. � Example Turtle file or download XML Recording of microdata follows defined rules. As a general principle, living natural persons may opt out of inclusion in the database. Statistical Data & Data Catalogue Our statistical data and catalogue consist of datasets aggregated using statistical methodologies. These comply with the SDMX standard and the W3C Data Cube vocabulary, making them compatible with spreadsheet software, statistical packages, and data science workflows in R, Python, or similar environments. Datasets are offered in multiple formats. In addition to RDF serializations, we provide standard CSV files and, when required, Excel or SPSS formats. We also publish data papers and related documentation that describe dataset usability and highlight key insights. Publications & Catalogue The SKCMDb’s most important publications are musical works, made available to end users as sound or video recordings and printed sheet music. These may be distributed in different sales formats, such as physical albums or books. Microdata and statistical datasets are treated as publications and are listed both in the general catalogue and a machine-readable data catalogue. The SKCMDb also includes methodological and musicological publications, as well as data papers explaining the use of datasets. The catalogue is designed for interoperability with libraries, archives, museums, and similar institutions. 149
In some cases, the SKCMDb may host the full publication, with the Open Music Observatory acting as publisher. In most cases, however, it provides catalogue entries with clear access points, such as webshops, public library lending systems, or repository links where legal copies of the documents can be obtained. 150
LīvMDb: Livonian Music Database The Livonian Music Database (LīvMDb) is a proof-of concept for working with music that has low documentation depth, weak institutions. The music of the Livonian people is scattered, and as native speakers of this small ethnic group died out, their heritage was dispersed, and largely not placed on modern digital platforms. The LīvMDb as a functional module of the Open Music Observatory, serves as a testbed for working with very low documentation, decolonisation, and other issues related to regional cultures, ethnic minorities. It offers insights into subsidiarity. Figure 4: The LīvMDb is federated with the Finno-Ugric Data Sharing Space and the Open Music Observatory. The first places Livonian music into a wider Finno-Ugric cultural context, the latter into a music context. The LīvMDb is supported by a data-sharing space consisting of both shared databases. The data-sharing space currently comprises the following initial components: •Finno-Ugric Metadata Database: A database that facilitates connections between various Finno-Ugric and Baltic stakeholders’ systems. •LīvMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 151
•LīvMDb (private): A technical dataset created to provide improvements, corrections, and enrichments. The Livonian Metadata Database serves as a support layer that is partly public and partly private. Its metadata definitions and descriptive metadata are exported into the LīvMDb databases as needed and permitted. The Livonian Metadata Database The Livonian Metadata Database is developed in alignment with the metadata framework of the Open Music Observatory. •Ontological and thesauri patterns: Reuses standardized or widely adopted vocabularies. •Conceptualizations and definitions: Includes concept definitions, thesauri, and other elements developed specifically for the LīvMDb. •Public permanent identifiers: Uses identifiers that are public or can be made public. The metadata layer is generally licensed under CC0, though in some cases other licenses are used (for example, CC-BY). ĹNote Examples: • The definitions of musical work,music recording and is recording of relationship allow the description of connections between an abstract musical work and its recording(s). • The ISRC code identifies recordings across all streaming platforms. Livonian Music Database (public) The Livonian Music Database is a linked open database published by the Reprex on behalf of the Open Music Observatory. The database is distributed under various Creative Commons licenses that allow both commercial and non-profit use. The primary aim of the LīvMDb is to provide a use case for very low documentation music ecosystems with very limited resources and challenging data curatorial scenarios. 152
Livonian Music Database (private) The private components of the LīvMDb are a staging area for data that has unclear provenance or legal status. Unlike in the case of LīvMDb, we hold minimal business confidential data (related to the royalty accounts of music that we published), but some data may have, for example, unclear GDPR status. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). Our work is centerred around the work of Hõimulõimed, a Finno-Ugric NGO, which curated many Finno-Ugric language collections, including the collection of Livonian-language songs available on Spotify. •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category. Our work buids on the register of Livonian placenames, because they offer the most straighforward curatorial help to find new Livonian (folk) music1. •Metadata: Relevant elements from the Livonian Metadata Database that support the use of collections or registers. The LīvMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models and data structures, enabling automatic aggregation. All datasets are available with RDF annotation and can be exported in all standard RDF serializations. Microdata is intended for institutional and professional use, not for the general public. It is annotated with standardized metadata suitable for applications such as music library cataloguing, distribution platforms, and rights management systems. 1Livonian place names: documentation, problems, and opportunities (Ernštreits 2020) and our gazetteer dataset: (Antal, Pigozne, and Mester 2025). 153