scieee AI-readable full text Open interactive document viewer

Open Music Observatory

Antal, Daniel; Lázár, Ádám; Mester, Anna Márta

Abstract

Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries.

Full text

Open Music Observatory Building an open data sharing space for the European music sector Daniel Antal, CFA Barát, Andor Kornél Lázár, Ádám Mester, Anna Márta 2025-11-18 Table of contents Open Music Observatory 7 DisclaimerofWarranties................................. 7 Glossary 9 Musicterms........................................ 9 Creatorsofmusicalworks ............................. 10 Datascienceterms.................................... 11 Dataprotectionterms .................................. 15 Data curation and collection terms . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 Statisticalterms ..................................... 16 Registers, authorities, standards and identifiers . . . . . . . . . . . . . . . . . . . . 17 Organisations....................................... 20 Otherabbreviations ................................... 21 Executive Summary 23 1 Introduction 26 2 Background & Concept 30 2.1 Why Europe Needs a Music Observatory . . . . . . . . . . . . . . . . . . . . . 31 2.2 Historical Precedent: CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 31 2.3 Policy and Technological Evolution Enabling a New Observatory . . . . . . . 32 2.3.1 European Parliament and EU-level Mandates . . . . . . . . . . . . . . 33 2.3.2 Data (Sharing) Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.3.3 Preference for Open-Source and Open Standards in the EU . . . . . . 34 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures . . . 34 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy 35 2.3.6 European Interoperability Framework (EIF) . . . . . . . . . . . . . . . 36 2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds . . . 36 2.3.8 Summary .................................. 37 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model . . . . . . . 37 2.4.1 Lessons from CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 37 2.4.2 Requirements of the EU policy environment . . . . . . . . . . . . . . . 38 2.4.3 Requirements of the Grant Agreement . . . . . . . . . . . . . . . . . . 38 2.4.4 Technical rationale for decentralisation . . . . . . . . . . . . . . . . . . 39 2.4.5 Why decentralisation is essential for a European Music Observatory . 39 2.5 The Data-to-Policy Pipeline: How Open Music Europe Works . . . . . . . . . 40 2.5.1 1. Indicator and problem definition (WP1–WP3) . . . . . . . . . . . . 40 2.5.2 2. Data governance (WP1–WP3, WP6) . . . . . . . . . . . . . . . . . 40 2 2.5.3 3. Software for data collection (WP4) . . . . . . . . . . . . . . . . . . 41 2.5.4 4. Data acquisition (WP1, WP2, WP3) . . . . . . . . . . . . . . . . . 41 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) . . . . . . 41 2.5.6 6. Validation for analysis and dissemination (WP4, WP5) . . . . . . . 42 2.5.7 7. Analysis and modelling (WP1–WP3) . . . . . . . . . . . . . . . . . 42 2.5.8 8. Policy translation (WP5) . . . . . . . . . . . . . . . . . . . . . . . . 42 2.5.9 9. Dissemination and reuse (WP5) . . . . . . . . . . . . . . . . . . . . 43 2.5.10Summary .................................. 43 2.6 Stakeholder Engagement and Early Feedback . . . . . . . . . . . . . . . . . . 43 3 Core Services 45 3.1 Collect: Data Curation & Collection . . . . . . . . . . . . . . . . . . . . . . . 47 3.1.1 Microdata, Collections, Records . . . . . . . . . . . . . . . . . . . . . . 48 3.1.2 Primary data collection . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3.1.3 Metadata .................................. 51 3.1.4 Statistical indicators and datasets . . . . . . . . . . . . . . . . . . . . 52 3.2 Repair........................................ 52 3.3 Process ....................................... 53 3.3.1 Processing & re-processing microdata . . . . . . . . . . . . . . . . . . 53 3.3.2 Documentation............................... 55 3.4 Disseminate..................................... 55 3.4.1 EU Open Data Portal . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 3.4.2 Europeana Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 3.4.3 European Collaborative Cloud for Cultural Heritage . . . . . . . . . . 57 3.4.4 European Open Science Cloud . . . . . . . . . . . . . . . . . . . . . . 58 3.5 Metadata ...................................... 60 3.5.1 Wikibase & Wikidata . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 3.5.2 Music Observatory Website . . . . . . . . . . . . . . . . . . . . . . . . 61 3.5.3 APIEndpoint................................ 61 4 Architecture 62 4.1 Why Wikibase Is the Right Foundation for the Open Music Observatory . . . 62 4.1.1 Proven in real-world scenarios highly similar to music . . . . . . . . . 63 4.1.2 Already aligned with Europe’s digital knowledge infrastructure . . . . 63 4.1.3 Demonstrated support for required OMO functionality . . . . . . . . . 64 4.1.4 Fits EU policy preference for open-source and trustworthy AI . . . . . 64 4.1.5 The most widely used graph-editing interface in the world . . . . . . . 64 4.1.6 A hybrid model that fits real institutional workflows . . . . . . . . . . 65 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline . . . 65 4.2.1 Wikibase supports each stage of the pipeline: . . . . . . . . . . . . . . 65 4.2.2 Datacollection ............................... 65 4.2.3 Validation and reconciliation . . . . . . . . . . . . . . . . . . . . . . . 65 4.2.4 Harmonisation and enrichment . . . . . . . . . . . . . . . . . . . . . . 65 4.2.5 Activation for analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 4.2.6 Indicator construction . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 4.2.7 Interpretation and contextualisation . . . . . . . . . . . . . . . . . . . 66 3 4.2.8 Policy translation and observatory outputs . . . . . . . . . . . . . . . 66 4.2.9 Feedbackloop................................ 66 4.3 Summary ...................................... 66 5 Data coordination 67 5.1 Ontologies and Vocabularies in the Open Music Observatory . . . . . . . . . 67 5.2 Registers....................................... 69 5.3 Global and Permanent Identifiers . . . . . . . . . . . . . . . . . . . . . . . . . 69 5.4 OpenCollections register services . . . . . . . . . . . . . . . . . . . . . . . . . 69 5.5 TBC......................................... 69 5.6 Adaptation to the music sector . . . . . . . . . . . . . . . . . . . . . . . . . . 70 6 Data Sources 71 6.1 How data sources were identified . . . . . . . . . . . . . . . . . . . . . . . . . 71 6.2 Administrative and register data . . . . . . . . . . . . . . . . . . . . . . . . . 72 6.3 Surveydata..................................... 73 6.4 Statistical and economic data . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 6.5 Platform and streaming data . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 6.6 Harmonisation tools, metadata structures, and vocabularies . . . . . . . . . . 75 6.7 Role of data sources in the data-to-policy pipeline . . . . . . . . . . . . . . . 76 6.8 Integration with the Open Music Observatory . . . . . . . . . . . . . . . . . . 76 7 Data Collection 78 7.1 Overview of the Data-Collection Framework . . . . . . . . . . . . . . . . . . . 78 7.2 Administrative and Register Data . . . . . . . . . . . . . . . . . . . . . . . . . 79 7.3 SurveyData..................................... 79 7.4 Statistical and Economic Data . . . . . . . . . . . . . . . . . . . . . . . . . . 80 7.5 Platform and Streaming Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 7.6 Processing and Harmonisation . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 7.7 Software Components Developed in WP4 . . . . . . . . . . . . . . . . . . . . 81 7.7.1 Data-Ingestion Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 7.7.2 Validation and Reconciliation Tools . . . . . . . . . . . . . . . . . . . 81 7.7.3 Harmonisation and Metadata Tools . . . . . . . . . . . . . . . . . . . 81 7.7.4 OMO Integration Tools . . . . . . . . . . . . . . . . . . . . . . . . . . 82 7.8 Position of Data Collection and Processing in the Pipeline . . . . . . . . . . . 82 7.9 Integration with the Open Music Observatory . . . . . . . . . . . . . . . . . . 82 8 Standardisation of Data & Terminology 83 8.1 Businessprocesses ................................. 83 8.2 Conceptual and information models . . . . . . . . . . . . . . . . . . . . . . . 84 8.3 Identification & Entity Linking . . . . . . . . . . . . . . . . . . . . . . . . . . 86 8.3.1 Registers & Authority Files . . . . . . . . . . . . . . . . . . . . . . . . 87 8.3.2 Open and persistent identifiers . . . . . . . . . . . . . . . . . . . . . . 88 8.3.3 Not open, music-industry specific identifiers . . . . . . . . . . . . . . . 88 8.3.4 Lyrics .................................... 90 8.3.5 ISCC .................................... 90 4 8.3.6 OMOIdentifiers .............................. 91 9 Data Improvement & Innovation 92 9.1 Value-Added Data Services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92 9.1.1 DataSharing................................ 92 9.1.2 Fix-the-data ................................ 93 9.1.3 DataLinking ................................ 93 9.1.4 Registration services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 9.2 UseCases...................................... 95 9.2.1 Data Health Services for Collective Management . . . . . . . . . . . . 96 9.2.2 Sustainability Reporting for Music Organisations . . . . . . . . . . . . 97 9.2.3 ListenLocal................................. 98 9.2.4 Unlabel ................................... 99 9.3 UseofAIsystems .................................100 10 Conclusions & Next Steps: Towards a European Music Observatory 103 10.1 Co-creating a governance model for the Observatory . . . . . . . . . . . . . . 104 10.1.1 Observatory Stakeholder Network . . . . . . . . . . . . . . . . . . . . 105 10.1.2 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . 106 10.2 Bottom-up expansion of the Observatory in a federation model . . . . . . . . 106 10.3 Practical next steps in the project . . . . . . . . . . . . . . . . . . . . . . . . 107 11 Data Catalogue 110 11.1CollectionGuidelines................................111 11.2TopicalPillars ...................................113 11.2.1MusicEconomy...............................114 11.2.2MusicDiversity...............................116 11.2.3MusicSociety................................117 11.2.4Innovation..................................118 11.2.5Sustainability................................118 References 119 Appendices 124 Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network 124 ...............................................124 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 Annex 2 - Data Curators Manuals 126 Inspiration ........................................127 Basic Data Organisation Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 Understanding the Data Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128 WorkingwiththeGUI..................................130 Sandboxenvironment ..................................130 Massimporting......................................131 5 Dataenrichment .....................................131 Quality Testing with SPARQL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131 Terminology 133 Mappingguidelines....................................133 Musicprofessionals.................................133 Musicalworks....................................134 Soundrecordings..................................134 Livepublicperformance..............................134 SKCMDb: Slovak Comprehensive Music Database 135 The Slovak Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 Slovak Comprehensive Music Database (public) . . . . . . . . . . . . . . . . . . . . 136 Slovak Comprehensive Music Database (private) . . . . . . . . . . . . . . . . . . . 137 Microdata......................................137 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 138 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 LīvMDb: Livonian Music Database 140 The Livonian Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 Livonian Music Database (public) . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 Livonian Music Database (private) . . . . . . . . . . . . . . . . . . . . . . . . . . . 142 Microdata......................................142 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 143 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 6 Open Music Observatory ¾Partly updated This document was presented as a planning document in 2023. It has been partly updated till 18 November 2025. Some text is still reflecting the planning phase. You can access all versions on https://zenodo.org/records/16539570 Disclaimer of Warranties This project has received funding from the European Union’s Horizon Europe, research and innovation programme, under Grant Agreement No. 101095295. This document has been prepared by Open Music Europe (OpenMusE) project partners as an account of work carried out within the framework of this contract. Any dissemination of results must indicate that it reflects only the author’s view and that the Commission Agency is not responsible for any use that may be made of the information it contains. Neither Project Coordinator, nor any signatory party of Open Music Europe (OpenMusE) Project Consortium Agreement, nor any person acting on behalf of any of them: (a) makes any warranty or representation whatsoever, express or implied, (i). with respect to the use of any information, apparatus, method, process, or similar item disclosed in this document, including merchantability and fitness for a particular purpose, or (ii). that such use does not infringe on or interfere with privately owned rights, including any party’s intellectual property, or 7 (iii). that this document is suitable to any particular user’s circumstance; or (b) assumes responsibility for any damages or other liability whatsoever (including any consequential damages, even if Project Coordinator or any representative of a signatory party of the Open Music Europe (OpenMusE) Project Consortium Agreement, has been advised of the possibility of such damages) resulting from your selection or use of this document or any information, apparatus, method, process, or similar item disclosed in this document. For the version history of this document, please refer to our open repository, where the change history can be reviewed with timestamps for every single file used to create the report: https://github.com/dataobservatory-eu/open-music-observatory 8 Glossary Music terms audio recording: fixation of sounds (ISO 2019b) music video recording: fixation of sounds synchronized with pictures or moving pictures where (a) the fixed sounds are wholly or substantially a musical performance or (b) the recording is intended for viewing in association with a recording of a musical performance. This definition includes music videos and concert recordings, together with music-related interviews and documentaries, but does not extend to genera! audiovisual material, even if it includes music.(ISO 2019b) recording: result of a recording process independent of the type and number of audio or audiovisual carriers and technology used Note 1 to entry: The term “recording” applies to each recorded item which may be used as a separate unit regardless of whether it is issued as part of a larger recorded work (e.g. each separate track on an album of audio recordings). [SOURCE:ISO 3901:2001, definition 3.3] (ISO 2017b) work: distinct, abstract creation of the mind whose existence is revealed through one or more expressions (e.g. a performance) or manifestations (e.g. an object) (ISO 2022) musical work: composed of a combination of sounds, with or without accompanying text (ISO 2022) DSP or digital streaming platform: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. DSPs can provide access to music downloads, like Apple’s iTunes Store, or access to streaming music like Spotify, or even provide satellite-delivered content such as SiriusXM in the USA. rights management (organisations): the function of managing the rights on behalf of rights owners. It can be companies whose sole purpose is to ensure that content that has been licensed has delivered royalties that are identified and accounted for. The role can be taken by collective management organisations or by private companies on behalf of songwriters, composers, performers, music publishers, or record labels. duration: the elapsed playing time between the first and last recorded modulations of the recording. LP or Long Player: gramophone record usually on both sides comprising one or more sound recordings with a playing time of each side of normally round about 30 minutes and released and sold on its own (ISO 2017b) 9 anthology: document consisting of a collection of full documents or of extracts, usually of literary works (ISO 2017b) exhibition: curated display of objects on a clear concept and communicating a message [SOURCE:ISO 18461:2016, definition 2.4.6 modified] (ISO 2017b) curator: person responsible for overseeing a collection or exhibition (ISO 2017b) data curation: managed process, throughout the data lifecycle, by which data/data collections are cleansed, documented, standardized, formatted and interrelated (ISO 2017b) register: an official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country; a document, usually a volume, in which data are entered in a formal manner by a statutory authority Note 1 to entry: In modern usage, usually a database. (ISO 2017b) registration: act of giving an entity a unique identifier on its entry into a system (ISO 2017b) set of rules, operations, and procedures for inclusion of an item in a registry (ISO 2023a) registrant: organization or person that has either registered an authentication protocol or registered the adoption of an authentication protocol [SOURCE: ISO/IEC 24727-6:2010, definition 3.4] (ISO 2017b); an entity wishing to assign an ISRC to an applicable recording (ISO 2019b); aparty that requests an ISNI from the Registration Authority (ISNI 3.2 (ISO 2012, p15)) party: natural person or legal person, whether or not incorporated, or a group of either (ISO 2012) aggregation: acquisition of sensitive information by collecting and correlating information of lesser sensitivity (ISO 2023b) Statistical terms administrative records: data generated by a non‐statistical source, usually a public body, the main aim of which is not the provision of statistics. code list: predefined list from which some statistical coded concepts take their values (ISO 2013) data pipeline: a method in which raw data is ingested from various data sources and then ported to data store. FAIR or FAIR Guiding Principles for scientific data management and stewardship: guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital 16 assets, emphasising machine-actionability (i.e., the capacity of computational systems to find, access, interoperate, and reuse data with none or minimal human intervention.) indicator: the representation of statistical data for a specified time, place or any other relevant characteristic, corrected for at least one dimension (usually size) so as to allow for meaningful comparison. microdata: non‐aggregated observations or measurements of characteristics of individual units, without direct identifier. MVP or minimum viable product: a version of a work product with just enough features and requirements to satisfy early customers and/or provide feedback for future development [SOURCE:IEEE 2675-2021, 3.1] observation unit: an identifiable entity about which data can be obtained, it is also often called a statistical unit or data subject in case of a natural person. Open Policy Analysis Guidelines: a set of information management rules to make policy analysis more transparent. personal data: any information relating to an identified or identifiable natural person. pseudonymisation: processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information. survey: a systematic examination and record of a physical or social area and its features so as to construct a map, plan, or description. In social sciences it usually refers to a well-structured questionnaire and answers given to its items by a target population. statistics: quantitative and qualitative, aggregated and representative information characterising a collective phenomenon in a considered population. visualisations: schematic charts, drawings, photographs, and their collages will as still image files that help to explain the relationship between information carriers, data points, or processes. Registers, authorities, standards and identifiers IČO: The organisation identification number (IČO) is an identifier assigned to all types of legal entities, entrepreneurs and public authorities by the Statistical Office of the Slovak Republic. The Czech Republic’s organisation identifier is also called IČO. OpenCorporates: a public corporation database which sources data from national business registries. ISNI: an ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. VIAF: The Virtual International Authority File (VIAF) is an international service that consolidates multiple name authority files into a single database. Their primary goal is 17 to enhance the efficiency and usability of library authority files by linking and merging widely used authority records and making them accessible online. VIAF ID: The VIAF (Virtual International Authority File) combines multiple name authority files into a single OCLC-hosted name authority service. ISRC: The International Standard Recording Code (ISRC) is a standard identifying code that can be used to identify sound recordings and music video recordings so that each such recording can be referred to uniquely and unambiguously. ISWC: The purpose in creating an ISWC for musical works is to enable more efficient administration of rights to those works on a worldwide basis. The ISWC provides an efficient means of identifying musical works in computer databases and related documentation and for the exchange of information between rights societies, publishers, record companies and other interested parties on an international level. ISBN: the International Standard Book Number is an identification system for the publishing industry and its supply chains. ISMN: The International standard music number (ISMN) was developed by, and for, the music publishing sector as a separate system to complement the International standard book number (ISBN). The existence of the ISMN as a separate identifier system makes it possible to identify printed and notated music as a distinct category of publication within the global supply chain and to develop trade directories and similar services for the specialized market for music publications. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. DOI: The Digital Object Identifier is a standardised unique number given to many (but not all) articles, papers and books, by some publishers, to identify a particular publication. ORCID: the Open Researcher and Contributor ID is a unique, persistent identifier free of charge to researchers. URI: A Uniform Resource Identifier (URI) is a string of characters used to identify a resource on the internet. This resource can be either abstract or physical, such as a website, an email address, or a file. URIs are essential for enabling interactions with resources over a network using specific protocols. W3C: The World Wide Web Consortium (W3C) is an international community that develops standards for the World Wide Web. Their mission is to lead the Web to its full potential by creating technical specifications and guidelines that are designed to be open and royalty-free. These standards include HTML, CSS, and other web technologies, which ensure that web content is accessible across different browsers and devices. DDI: The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. Wikibase: Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, 18 but it is available now for the creation of private, or public-private partnership knowledge graphs. It is developed by Wikimedia Deutschland. GSBPM: The Generic Statistical Business Process Model is a international standard model that “describes and defines the set of business processes needed to produce official statistics.” | GSIM:Generic Statistical Information Model: a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation standards such as SDMX or DDI. SDMX: Statistical Data and Metadata eXchange (SDMX), is an international initiative that aims at standardising and modernising (“industrialising”) the mechanisms and processes for the exchange of statistical data and metadata among international organisations and their member countries. ESRS: The European Sustainability Reporting Standards (ESRS) are a set of guidelines developed by the European Financial Reporting Advisory Group (EFRAG) to standardise sustainability reporting across the European Union. These standards are designed to align with the Corporate Sustainability Reporting Directive (CSRD), which mandates detailed corporate reporting on environmental, social, and governance (ESG) issues for many companies operating within the EU. CIDOC-CRM: The conceptual model of CIDOC, the standard conceptualisation of collection management systems in heritage organisations. RiC:Records in Context is a new conceptual model that replaces the four most important international archiving standards. DCTERMS or DCMI: the Dublin Core Metadata Terms is a vocabulary of metadata terms developed and maintained by the Dublin Core Metadata Initiative (DCMI). These terms are used to describe various aspects of digital resources, such as web pages, documents, and other online content. They provide a standardized way to assign metadata to resources, making them easier to discover, manage, and exchange. RDFS: the Resource Description Framework Schema is an extension of the Resource Description Framework (RDF) that provides a vocabulary for describing classes and properties of resources within an RDF graph. EDM: the Europeana Data Model is a framework for collecting, connecting, and enriching cultural heritage metadata. It’s designed to facilitate the sharing and reuse of cultural heritage information by providing a standardized way to represent and link data. Europeana: a digital platform provided by the European Union that aggregates digitized cultural heritage from institutions across Europe. ESCO: the European Skills, Competences, Qualifications and Occupations classification is is a multilingual classification system developed by the European Commission to standardize the description of skills, competences, and qualifications relevant to the European labor market and education. 19 NACE: the European Union’s standard classification of economic activities for statistical purposes. The abbreviation stands for Nomenclature statistique des Activités économiques dans la Communauté européenne. ISCO: the International Standard Classification of Occupations (ISCO) is the International Labour Organization’s standardized system for classifying and organizing occupations according to jobs’ tasks and duties ISIC: the International Standard Industrial Classification of All Economic Activities (ISIC) is a standard classification system developed by the UN Statistics Division (UNSD) to categorize economic activities. PROV-O: the Provenance ontology is a formal ontology developed by W3C to represent and interchange provenance information. MARC: MAchine-Readable Cataloging, is a standard digital format used by libraries to represent and exchange bibliographic information. DCAT: an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. Organisations AEPO-ARTIS: Organisation representing European artists-performers. Regroups most of the European CMO representing performers. ALOADED: is a company which distributes and exploits recordings. CISAC: The International Confederation of Societies of Authors and Composers is an international non-governmental, not-for-profit organisation that aims to protect the rights and promote the interests of creators worldwide. CNM (former CNV): the Centre National de la Musique is a public organisation managing a tax on concert tickets EFRAG: The European Financial Reporting Advisory Group is a private association established in 2001 with the encouragement of the European Commission to serve the public interest. EFRAG extended its mission in 2022 following the new role assigned to EFRAG in the CSRD, providing Technical Advice to the European Commission in the form of fully prepared draft EU Sustainability Reporting Standards and/or draft amendments to these Standards. EMO: The European Music Observatory (EMO) is envisioned as a hub for collecting and analysing data on the music sector across Europe. Its primary aim is to address the current gaps and inconsistencies in music data collection, which have been a significant challenge for the sector. GESAC: GESAC comprises together 32 European authors’ societies in music, audiovisual, visual arts, literature and drama. 20 GESIS: Leibniz Institute for the Social Sciences. IAML: International Association of Music Libraries, Archives and Documentation Centres | IAMIC: International Association of Music Centres, an international network of organisations that collectively and collaboratively provides information and promotes the music of their countries or regions. ICMP: the global trade body representing the music publishing industry worldwide. SCAPR: International association for the development of the practical cooperation between performers’ collective management organisations (CMOs) SOZA: SOZA (Slovenský ochranný zväz autorský pre práva k hudobným dielam, Slovak Performing and Mechanical Rights Society) is a legal entity, non-profit civic association of authors and publishers of musical works, association of natural persons and legal entities. Hudobné Centrum: Music Centre Slovakia is a music organisation with a mission to promote Slovak contemporaly music. Other abbreviations CEEMID: the Central European Music Industry Databases is a multi-country project that was a predecessor of Reprex’s Digital Music Observatory CSRD: The Corporate Sustainability Reporting Directive (CSRD) is European Union (EU) legislation, effective from 5 January 2023, that requires EU businesses—including qualifying EU subsidiaries of non-EU companies—to disclose their environmental and social impacts, and how their environmental, social and governance (ESG) actions affect their business. DSP: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. EIF: The European Interoperability Framework (EIF) is a set of recommendations and guidelines that aims to facilitate communication and collaboration between public administrations, businesses, and citizens within the European Union and across national borders. ECCCH: The European Collaborative Cloud for Cultural Heritage is a European Union initiative for a digital infrastructure that will connect cultural heritage institutions and professionals across the EU. EOSC: The European Open Science Cloud (EOSC) aims to create a trusted, open, and multidisciplinary environment for researchers and innovators in Europe. PPP: A Public-Private Partnership (PPP) is a collaborative arrangement between government entities and private sector companies aimed at financing, designing, implementing, and operating projects or services traditionally provided by the public sector. RDM: Research Data Management refers to the suite of practices, policies, and processes used to handle data throughout the lifecycle of a research project. 21 Our glossary is harmonised with relevant music-sector specific standards (referred to in Chapter 8) and with the ISO Information technology — Vocabulary (ISO 2023b); Information technology — Cloud computing — Taxonomy based data handling for cloud services (ISO 2020); Information technology — Cloud computing — Interoperability and portability (ISO 2017a) and the Information and documentation — Foundation and vocabulary (ISO 2017b) and Information technology — Metadata registries (MDR) — 1. Framework (ISO 2023a) 22 Executive Summary Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. In the OpenMusE project, the development of the Open Music Observatory is coupled with a clear contractual expectation: the project must populate the Observatory’s four thematic pillars—Music Economy, Music Diversity, Music, Society & Sustainability, and Innovation & Future Trends —with initial, well-documented data and knowledge. This population process follows the project’s data-to-policy pipeline (see Chapter 1): each work package defines its indicators, establishes data governance and legal bases, collects or accesses relevant administrative, survey, statistical, and platform data, processes and harmonises them through WP4 tools, and finally activates them in reproducible analytical workflows in WP5. The result is that the Observatory is not only a technical prototype but a functional, databearing infrastructure: the first integrated demonstration of how Europe’s music data can be curated, linked, analysed, and made reusable across public, private, and civic actors. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries. An observatory has traditionally been a permanent location for observing terrestrial, marine, or celestial events. In the past 30 years, it has also been used for long-term digital data collection programs for markets, social sciences, and humanities. Our milestone requires the start of this observatory after a lengthy and intensive planning and prototyping phase. It can be seen as a modern reimagination of the data observatory model, or the observatory 2.0. We created a new observatory model that fully aligns with the European Interoperability 23 Framework but extends the governance of the digital services beyond public bodies, and allows the creation of a public-private partnership to manage the observatory. The European Interoperability Framework aims to create a four-layered approach to build digital research, marketing, rights management, collection management services for the sector. These layers are introduced in separate chapters of these documentation. 1. The technological alignment is introduced in Chapter 4; we decided to choose the technology of the world’s largest open knowledge graph, Wikidata, which already coordinates countless digital services in Europe’s cultural sector, and provides training and guardrails for many AI applications. 2. The semantic alignment is is introduced in ?@sec-coordination. 3. The organisational alignment means bridging actual data-driven and computer supported workflows with data semantics (what does a song’s title mean for a librarian, an ethnomusicologist, a collective rights management agency), and how they can work with various translated, alternative, historical, mistyped, and preferred titles in distributing royalties, loaning printed sheets, describing musical traditions. 4. The legal alignment creates a policy that lays out the rights, prohibitions and necessary permission processes to connect and use the data together. We were informed and influenced by the creation of Europeana (which started out from a similar collaborative project, a cultural heritage oriented data sharing space) and the Commission’s new plans to extend their digital services into the European Collaborative Cloud for Cultural Heritage (ECCCH). We aimed for full interoperability with Europeana and we Reprex successfully sent data and concluded a Data Exchange Agreement. We were also aiming for interoperability with the ECCCH, which only published the first version of its Heritage Digital Twin conceptual model and ontology; we were the first to test them with music data. but we also bring a new element into their thinking. While they are mainly aggregating the work of public sector memory institutions, we are building a governance model that allows a more successful cooperation among the private sector and the public music sector. By the end of 2025, we aim to create an “observatory 3.0”, which already hosts many intelligent data improvement technologies and fuels innovative applications/services in line with our project’s exploitation pathways. These services are at different maturity levels, but they could not be brought to a testable MVP without building out the minimal digital infrastructure and governance model at this milestone. 24 ĹNote This document is licensed under the CC BY 4.0 LEGAL CODE Attribution 4.0 International license. You must refer to the document with the DOI 10.5281/zenodo.11385044. Canonical Licence URL:https://creativecommons.org/licenses/by/4.0/ Other formats:Plain Text;RDF/XMLlSee the deed 25 collective management societies. Over time, it grew to include more than 60 stakeholders in 12 European countries. Its purpose was to fill the most pressing evidence gaps by combining: • voluntary data integration among partners, • open-data reprocessing, and • co-financed data collection. This work is documented in (Antal 2020a). CEEMID operated according to principles that would later become central to the European Union’s data (sharing) space strategy (formalised only years later). Its decentralised organisational model, distributed data stewardship, and emphasis on transparent, reusable methods demonstrated that a modern observatory in the digital era does not need to be a centralised institution. Instead, it can function as a federated ecosystem connecting statistical offices, cultural institutions, CMOs, and private actors. Long before the EU formalised its dataspace strategy, CEEMID also aligned its workflows with emerging statistical-system standards such as GSIM,DDI, and SDMX, anticipating later European requirements for interoperable, machine-readable statistical metadata. This early adoption provided a methodological bridge between cultural-sector data, administrative registers, and official statistics, and formed a direct precursor to the metadata foundations of the Open Music Observatory. The Feasibility Study for the European Music Observatory explicitly recognised CEEMID as a potential building block for a new observatory model. Our proposal therefore sought to transform CEEMID’s prototype—referred to in the study as the Digital Music Observatory— into a scientifically robust and methodologically coherent system that could scale across Europe. This required grounding the work in state-of-the-art statistical science, data science, and computer science, and ensuring alignment with European interoperability and datagovernance frameworks. The prototype work that preceded Open Music Europe was shaped through two innovation environments: the Yes!Delft AI+Blockchain Lab, where product–market fit and technical feasibility were tested, and the JUMP Music Market Accelerator, where the first integrated prototype of a Digital Music Observatory was developed. These early iterations validated not only stakeholder demand but also the feasibility of a decentralised, standards-based architecture, and they informed the methodological and technical design choices taken forward in this project. 2.3 Policy and Technological Evolution Enabling a New Observatory Since the publication of the EMO Feasibility Study, the European Union has introduced a series of policy and infrastructure initiatives that strengthen the case for a decentralised, in32 teroperable, and federated European Music Observatory. These developments span European Parliament mandates, Commission-funded research, cultural-heritage clouds, opendata regulation, and the EU’s overarching data-space strategy. Together, they establish the policy and technological foundations on which the Open Music Observatory is built. Our policy alignment is discussed in more detail in - Music Metadata Mainstreaming and EU Law -A Green Paper on AI, Data Governance, and Metadata–Policies for Europe’s Music Ecosystem2 2.3.1 European Parliament and EU-level Mandates The European Parliament, in its resolutions on the future of the music sector, explicitly called for: • the establishment of a European Music Observatory, • improved evidence for competitiveness, diversity, and fair remuneration, and • stronger coordination of public, private, and community data sources. These mandates update and reinforce both the Music Moves Europe framework and the findings of the EMO Feasibility Study. They frame the Observatory as an instrument that must serve industry, civic, and public actors through interoperable, reusable, crossborder data services. The EU Music Ecosystem Study (2024) deepened this diagnosis, pointing to fragmentation across metadata, rights information, cultural statistics, and market data. It concluded that the sector requires a technical and governance model capable of linking these domains, rather than separate, siloed initiatives. The architecture of our dataspace responds directly to these recommendations3. 2.3.2 Data (Sharing) Spaces The EU’s adoption of data (sharing) spaces provides the organisational and legal model for an Observatory that is not a centralised institution but a federated ecosystem. Curry defines dataspaces as: “an emerging approach to data management… Data is integrated on an ‘asneeded’ basis, with the labour-intensive aspects of data integration postponed until they are required.” (Curry 2020) The Design Principles for Data Spaces position paper further describes them as: 2See (Senftleben et al. 2024); and (Antal 2025c), summarised in the internal document (Open Music Europe Consortium 2025). 3(eu_music_ecosystem_study?;ep_music_resolution?) 33 “a federated data ecosystem within a certain application domain and based on shared policies and rules.” (Nagel and Lycklama 2021, p7) These principles are fully consistent with CEEMID’s decentralised model and form the conceptual basis for the Open Music Dataspace (see Chapter 10). Observatories created in the 1990s and early 2000s were built around centralised databases and slow-moving data-collection cycles. Since then, the rapid expansion of agentic AI in data collection, the widespread digitisation of live and recorded music, and the proliferation of large-scale, real-time data sources have made such centralised architectures obsolete. Modern evidence ecosystems require automated ingestion, continuous semantic enrichment, cross-domain reconciliation, and transparent provenance — all of which presuppose a federated, decentralised model rather than a single institutional database. The European Audiovisual Observatory (EAO), the European Market Observatory for Fisheries and Aquaculture Products (EUMOFA), and the European Observatory on Infringements of Intellectual Property Rights (EUIPO) provide valuable models of long-standing EU observatories. However, each operates within a centralised data-submission and aggregation framework appropriate to their legal mandates and sectoral data structures. The Feasibility Study acknowledged that the music sector lacks comparable legal obligations and contains far more fragmented, cross-domain, multilingual, and institutionally diverse datasets. Therefore, while these observatories offer important governance precedents, their centralised architectures cannot be replicated in the music ecosystem — strengthening the case for a federated dataspace model. 2.3.3 Preference for Open-Source and Open Standards in the EU Across the EU’s data and digital-transition strategies, there is a consistent preference for: •open-source software, •open standards, •open licensing, and •transparent, reproducible workflows. This aligns directly with the Observatory’s use of open-source R and Python pipelines, Wikibase for semantic interoperability, and FAIR-compliant metadata. 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures Europeana demonstrates how Europe manages distributed cultural-haritage collections at scale using: 34 • persistent identifiers, • multilingual metadata, • open licences (e.g. CC BY), • shared semantic standards (EDM, IIIF, rightsstatements.org), and • decentralised stewardship by libraries, archives, and museums. The Open Music Observatory follows the same principles. It uses: •semantic technologies, •PID-based cross-domain linking, and •open, reusable data models. This ensures interoperability with cultural-heritage collections, performing-arts archives, and national memory institutions, and aligns the music domain with the emerging European Collaborative Cloud for Cultural Heritage (ECCCH). 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy The EU Open Data Portal (data.europa.eu) establishes a common framework for: • open licences (e.g. CC BY 4.0), • machine-readable formats, • harmonised metadata (DCAT-AP), • and publication of public-sector information. The Open Music Observatory is designed so that: • public datasets can be harvested directly by the EU Open Data Portal, • indicators and derived datasets comply with open-data rules, and • metadata follow DCAT-AP and DataCite to support long-term reuse. This alignment ensures that the Observatory meets both Horizon Europe open-science requirements and broader EU open-data policy objectives. 35 2.3.6 European Interoperability Framework (EIF) The European Interoperability Framework (EIF) provides a four-layer model—legal, organisational, semantic, technical—for connecting: • public authorities, • cultural institutions, • rights-management organisations, • national statistical offices, and • private intermediaries. These are precisely the actors whose data must interoperate to support a European Music Observatory. By adopting the EIF, the Observatory can link diverse datasets into coherent, reusable services without centralising them. 2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds The European Open Science Cloud (EOSC) and the European Collaborative Cloud for Cultural Heritage (ECCCH) promote: • FAIR data, • open science workflows, • reproducible analysis, • transparent provenance, and • decentralised storage and processing. These principles inform the Observatory’s architecture through the use of: • open-source analytical pipelines, • SDMX and DataCite metadata, • persistent identifiers, and • federated linking across domains and institutions. 36 2.3.8 Summary Together, these EU policy instruments—the Parliament’s mandate, the EU Music Ecosystem Study, data-space strategy, Europeana, the EU Open Data Portal, the EIF, EOSC, and ECCCH—provide a unified rationale for an Observatory that is federated, decentralised, data-driven, and interoperable by design. They define the policy and technological environment in which the Open Music Observatory must operate and directly shape its architecture. The Chapter 4explains why we chose an architecture that is built around Wikibase and Wikiadta. 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model The Open Music Observatory adopts a decentralised, federated dataspace model because this is the only architecture that meets the needs identified by the EMO Feasibility Study, the EU Music Ecosystem Study, and the European Parliament’s resolutions, while also complying with the newer EU frameworks for interoperability, data governance, and cultural-heritage infrastructures. A centralised database model, common in observatories built in the 1990s or early 2000s, is no longer feasible or desirable for the music sector. 2.4.1 Lessons from CEEMID The CEEMID collaboration demonstrated that most music-sector data—repertoire, rights, cultural-heritage descriptions, business metadata, and statistical evidence—originate from many different institutions, each with its own mandates, legal obligations, and technical systems. Centralising such data is: • legally constrained (e.g. GDPR, contractual confidentiality), • institutionally unrealistic (distributed ownership and stewardship), and • technically inefficient (rapidly evolving local systems). CEEMID showed that these data can nonetheless be made interoperable through: • shared identifiers and authority files, • open metadata standards, • reproducible R-based pipelines, and • rule-based, voluntary data sharing. These are the foundational principles of a data (sharing) space, which the EU has since elevated to a core strategic component of its digital-policy agenda. 37 2.4.2 Requirements of the EU policy environment As outlined in Section C, the EU now expects cultural and creative sectors to adopt: •federated data architectures, •FAIR and open data practices, •transparent governance models, •semantic interoperability, and •alignment with Europeana, EOSC, ECCCH, and data.europa.eu. This expectation reflects the broader transformation of European data governance, where sectors are encouraged to organise around data spaces rather than central repositories. A decentralised model also supports cultural and data sovereignty by allowing institutions to maintain control over their collections and data-processing rules. 2.4.3 Requirements of the Grant Agreement The Open Music Europe Grant Agreement defines the project explicitly as: “an open, scalable data-to-policy pipeline for European music ecosystems” and mandates the creation of: “a highly automated, decentralised intelligence hub that aggregates open data and creates dynamic, live policy documents.” To fulfil these contractual obligations, the Observatory must: • connect heterogeneous data sources without centralising them, • refresh indicators automatically as upstream data changes, • maintain legally sound provenance across many institutions, • support multilingual, cross-border metadata, and • integrate statistical, cultural-heritage, and industry systems. These requirements can only be met in a federated dataspace, not in a single, centralised database. 38 2.4.4 Technical rationale for decentralisation The dataspace model makes it possible to: • keep sensitive or personal data (e.g. rights, royalties) within the institution that controls them, • link sources through semantic federation (Wikibase/Wikidata), • enable distributed curation by librarians, archivists, CMOs, and researchers, • integrate permanent identifier (PID) systems across domains (ISNI, VIAF, ROR, company registers), • use open standards (SDMX, DDI, DataCite, DCAT-AP), and • scale to new partners, genres, languages, and Member States. This structure mirrors the actual distribution of data in the music sector and the technical direction of the EU’s digital transition. 2.4.5 Why decentralisation is essential for a European Music Observatory For the European music ecosystem, decentralisation enables: • lower administrative and compliance burdens, • institutional autonomy and data sovereignty, • cross-border comparability without forced data transfer, • communityand expert-driven metadata improvement, • GDPR-compliant handling of personal data, and • sustainable expansion of the Observatory. A decentralised dataspace is therefore not an architectural choice but a necessary governance model for an Observatory that spans cultural heritage, rights management, statistical registers, community archives, and private-sector metadata across the EU. The Open Music Observatory is consequently designed as a federated, rule-based dataspace: an ecosystem where public, private, and civic stakeholders contribute knowledge, maintain authority records, and generate indicators while preserving full control over their own data. 39 2.5 The Data-to-Policy Pipeline: How Open Music Europe Works The Open Music Europe action is contractually defined as “an open, scalable data-topolicy pipeline for European music ecosystems” (see Grant Agreement). This is not a slogan: it is the methodological core of the project and the organising principle of all work packages (WP1–WP5). The pipeline connects indicator design, data governance, data acquisition, semantic modelling, statistical analysis, and policy translation into a single reproducible workflow. This chapter introduces the logic of that pipeline and explains how it shapes the design of the Open Music Observatory. 2.5.1 1. Indicator and problem definition (WP1–WP3) Each thematic work package begins by identifying policy-relevant gaps and defining the indicators needed to address them. Deliverables D1.1, D2.1, and D3.1 specify: • the conceptual frameworks guiding each domain (economy, diversity, society), • the data requirements for measuring them, and • the procedures for ensuring comparability across countries and years. These definitions also appear in the Open Music Europe Data Management Plan (D6.3), which provides human-readable summaries and machine-readable metadata for all indicators. 2.5.2 2. Data governance (WP1–WP3, WP6) Before data can be collected or integrated, partners agree on: • sources, access rights, and sampling frames; • metadata standards (SDMX, DDI, DataCite); • ethical safeguards and GDPR-compliant procedures; • controlled vocabularies, authority files, and persistent identifiers. These agreements are formalised in D1.2, D2.2, D3.2, and the Data Management Plan (D6.3). They ensure compliance with FAIR, OPA, and EU data-governance principles. 40 2.5.3 3. Software for data collection (WP4) WP4 develops the open-source tools used to gather and ingest administrative data, survey data, platform usage data, and CMO records. These tools form the operational backbone of the pipeline. They include: • survey-data management scripts, • connectors for royalty and licensing accounts, • streaming API integration modules, • metadata templates for ingestion and harmonisation. All tools adhere to the reproducibility and interoperability requirements defined in Annex 1 and the DMP. 2.5.4 4. Data acquisition (WP1, WP2, WP3) Data are collected from: • collective management organisations (CMOs), • ministries and statistical offices, • cultural-heritage institutions, • surveys (enterprise and personal), • streaming-service APIs. Each domain follows its own protocol (e.g. WP1 T1.2 sampling frames; WP3 T3.1 participation and wellbeing indicators). Data collected are documented in the DMP and crossreferenced with OPA-compliant folders. 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) Raw inputs are processed using REPREX’s R-based openmusic-pipeline: • cleaning and pseudonymisation, • metadata harmonisation, • cross-linking with authority records, 41 of scientific accuracy. (European Commission et al. 2020, p28) In short, we collect data about music, as defined in the cultural statistics of any European Economic Area and EU candidate statistical office or by a representative European or international music organisation. In more detail, we a systematic data collection program requires a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. ĹNote Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. We work with a concept of a musical work, not with the entire work. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. Again, in an information system we obviously work with a concept of an author, and instances of authors represented by their unique data. The EMO feasibility study catalogues 45 data gaps that a future European music observatory should fill. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. We need agreed concepts of the composer,sound recording,work, to answer such questions. The initial data collection guidelines of the Open Music Observatory are derived from the EMO Feasibility study. We see them as a starting point for further discussion with the Observatory Stakeholder Network. We introduce them with our data catalogue in Section 11.1. These guidelines are supported by our first conceptualisation, which is built on some widely used conceptualisations of creative works and statistics. This is the topic of Chapter 8. 3.1.1 Microdata, Collections, Records We treat “microdata” as a collection of structured data. Aregister is a document [in modern usage, usually a database], in which data are entered in a formal manner by a statutory authority (ISO 2017b). In statistical data collection ian official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country. 48 Acollection is a group of objects, for example, musical works, sound recordings, printed scores, music enterprises, musician biographies, gathered together for some intellectual, artistic, or curatorial purpose. This is how radio playlists and charts, festival line-ups, local content guideline monitoring works; music labels and publisher select and musical works and their recordings or scores to place into commercial circulation. Such collections form the basis of census or sample surveys for statistical data collection. The documentation of collections relies on the work of registers. For example, music publishers can claim their revenues based on ISWC and ISMN identifiers provided to them by the collective management organisations that register works, or national libraries or other organisations that identify printed sheets. The maintenance of registers requires ongoing investment, and therefore registrars like the ISRC Authority or CISAC (the manager of the ISWC register) often restrict access to their data, or do not exchange data. In an increasingly globalised, automated music ecosystem where the number of identifiable works, recordings, scores, and related claims is growing exponentially, this situation puts the entire industry at a disadvantage, for example, against tech platforms. The Open Music Observatory is experimenting with innovative ways how registers can work together in some aspects of metadata standardisation, improvement and exchange in a way that keeps their core product intact and exclusive to them See: (Antal and Mester 2025). The Open Music Observatory works with metadata in a way that helps managing and improving registers, and it helps to create data about collections with authoritative data from registers. ĹNote Our first large database is the Slovak Comprehensive Music Database. Our aim is to publish a constantly refreshed database of every music composed or recorded in the territory of the current Slovak Republic, or composed and recorded by people from Slovakia, or sung in the Slovak language. This database is partly based on registers, and partly on curated holdings of Slovak stakeholders. 49 ⊠Our collections are always available on https://reprexbase.eu/skcmdb/�. Further details in Section 3.5.1. ⊠Whenever our collections fit in the collection and publication guidelines of Europeana, we make the collections available there, too. □We are investigating the possibility of synchronising our collections to the European Collaborative Cloud for Cultural Heritage. The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. The DDI plays a particularly important role in the creation of statistical surveys, particularly using questionnaires and question banks. The new Records in Context has replaced the international standards on archives in 2023. Its central concept is the record, which is a document according to DDI; a collection is a set of records. Our standardisation of microdata is explained in more detail in Chapter 8. 3.1.2 Primary data collection The Open Music Observatory is supporting high-quality primary data collection, and itself is carrying out such collection activities. The indicators derived from the processing of survey questionnaires will be comparable if the same concepts of interest (for example, concert visiting frequencies) are measured via the same questions and answering instructions. 50 ĹNote Aconcert is a standard concept of a live performance of music. How many times in the previous [12 months] have you been to a concert? is a standard question accompanied by standardised answer options and processing in the Cultural Access and Participation surveys following the ICET model. Using standardised concepts and question banks, including question and instruction labels with standardised translations, is a cornerstone of ex-ante survey harmonisation. This process is a prerequisite for retrospective survey harmonisation and the subsequent creation of comparable statistical indicators, underscoring the importance of uniformity in data collection. ⊠We provide API and download access to harmonised, multi-language question banks. This allows music stakeholders to use the same question formulations and translations for comparability with European statistical and policy research programs. ⊠We provide tutorials to retroharmonize, a background open-source software of Reprex, which is an R library to retrospectively harmonise data from different survey programs that had asked the same questions. ⊠The Open Music Europe project will carry out some harmonised surveys to show and improve the methodology of harmonised data collection within the music sector of Europe. This data will be available as metadata (questionbank), as microdata (individual answers), and as processed statistical data. 3.1.3 Metadata The most common—and perhaps least useful—definition of metadata is that it is “data about data.” As catchy as this definition is, however, it is entirely ambiguous. First of all, what is data? And second, what does “about” mean? (Pomerantz 2015a, p19) The new ISO standard on Information technology — Metadata registries (MDR) defines metadata as data that define and describe other data. As Pomerantz eloquently argues, this is a definition that is not very helpful. We use his more functional (but not contradictory) definition. “Data is only potential information, raw and unprocessed, prior to anyone actually being informed by it. […] Data must be understood not as an abstract concept but as objects that are potentially informative. […] Metadata Is a Statement about a Potentially Informative Object.” (Pomerantz 2015a, p26) Following the metadata definition of “a statement about a potentially informative object,” we believe that any high-quality data can be used as metadata in certain circumstances. 51 ĹNote Data or metadata? The data of birth can be seen as a metadata for disambiguation among authors with the exact same name in a copyright register. It can be seen as data for a curator of a young author prize, or a music sociologist. Either way, the date of birth should be precise, and encoded in a way that makes it portable and interoperable. From a data management point of view, we do not distinguish between data and metadata. Of course, we acknowledge the fact that some types of data will always remain under the hood and will only serve the proper functioning of an information system. The music industry’s famous “metadata problems” usually arise when a music enterprise or institution wants to use metadata information from an authoritative source that is somehow corrupted. The Open Music Observatory can help with these metadata problems by disseminating proper, open authoritative data (as registers or collection) or by providing data improvement services that fix the metadata problems of a user. 3.1.4 Statistical indicators and datasets ÁWarning We will place our first statistical datasets to the EU Open Data Portal this week (pending their approvals) and will provide a screenshot and access conditions here. 3.2 Repair Throughout the project we realised that data and metadata repair is perhaps a more urgent challenge then data processing. While our team and our stakeholders gradually embraced the concept of a data sharing space, i.e., the idea that instead of starting new data collections from scratch it is more economic and useful utilise existing data, given the high level of music industry digitalisation and that almost all transactions leave a digital trail behind, we also realised that the music sector is “drowning in numbers”; it handles more data in various obsolete, undocumented, unstructured or ad hoc forms than it can utilise. In these cases, usually the data is already available somewhere, but in a format that prevents the data to be used to its potential. Most of our efforts therefore were concentrated on data and metadata repair instead of new collection. Metadata repair and data processing is usually hard-to-distinguish tasks that comprise of similar or same steps. They are various validation, normalisation procedures that allow that make the data informative. 52 3.3 Process We use the theory of metadata by Jeffrey Pomerantz, who defines Metadata as “a statement about a potentially informative object.” A dataset without such statements is not findable, accessible, interoperable, and very hard to reuse. Pomerantz distinguishes among descriptive, administrative, structural, preservation, and use metadata. The Generic Statistical Information Model (GSIM) is a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation metadata standards such as SDMX or DDI. GSIM since its inception aims to bridge two important standards, SDMX and DDI. The Statistical Data and Metadata Exchange has been developed for decades and it is an ISO standard; it is more geared towards the aims of data sharing and preservation in RDM. DDI on the other hand is more focused on the documentation and quality control of primary data collection, or the reuse of often messy data sources, and supports the processes that make the data available for research. As DDI provides information about a much wider range of objects and processes, we are even more selective when we turn to this standard than SDMX; however, we cannot disregard DDI for microdata. 3.3.1 Processing & re-processing microdata ÁWarning We will place here an example that goes to the EU Open Data portal The EU Open Data Portal uses the following namespace definitions; these definitions refer to machine readable, explicit definitions (ontologies) of the way our datasets must be understood by a software agent. To demistify the process, we provide here an example of the metadata that we need to compile from the various steps of the data production pipeline. @prefix rdf:<http://www.w3.org/1999/02/22-rdf-syntax-ns#> . @prefix foaf:<http://xmlns.com/foaf/0.1/>. @prefix rdfs:<http://www.w3.org/2000/01/rdf-schema#> . @prefix xsd:<http://www.w3.org/2001/XMLSchema#> . @prefix owl:<http://www.w3.org/2002/07/owl#> . @prefix adms:<http://www.w3.org/ns/adms#> . @prefix dcat:<http://www.w3.org/ns/dcat#> . First we must translate the metadata of our datasets to any of the standard serialisations (file formats) of the World Wide Web Consortium’s Resource Description Framework definition, which allows the connection of data across the open internet. At the time of writing this report, the EU Open Data Portal was changing its backend, and for testing purposes, we worked with a dataset from the background of the Open Music Europe project (which had been earlier published by Reprex on Zenodo under the title *The turnover of the ration broadcasting industry in Europe*. ) 53 The dataset itself cannot be downloaded from a data catalogue. It is an abstract intellectual work, similar to musical work or a literary work. A musical work is accessible in printed sheets or recordings, and a dataset in a distributed data file. <https://doi.org/10.5281/zenodo.5652118> <a>"dcat:Dataset" ; <dcat:distribution><https://zenodo.org/records/5652118/files/codebook_trb.csv>,"https://zenodo.org/records/5652118/files/codebook_trb.csv" ; <dct:creator><https://orcid.org/0000-0001-7513-6760>; <dct:description>"\"The turnover of the ration broadcasting industry in Europe.\"@en" ; <dct:identifier><https://doi.org/10.5281/zenodo.5652118>; <dct:issued>"2022-06-03T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:modified>"2022-06-04T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:publisher><https://isni.org/isni/000000050973936X>; <dct:title>"A rádió szektor forgalma Európában\"@hu","\"Turnover of the Radio Broadcasting Industry in Europe\"@en" ; <edp:originalLanguage><rdf:resource><http://publications.europa.eu/resource/authority/language/ENG>. We can provide further provenance information about the dataset; in production, we will provide information on software agents (tools) used, researchers, data managers and curators and their organisations involved. As a bare minimum, we provide machine-readable information about the technical publisher of the dataset, Reprex B.V: <https://isni.org/isni/000000050973936X> <a>"foaf:Agent" . And then we point the user the downloadable files (distributions) of the dataset with the rights statements and licenses. We use the Creative Commons CC BY 4.0 license, similar to Eurostat on the EU Open Data Portal, and we state that the dataset is open for the public. <https://zenodo.org/records/5652118/files/codebook_trb.csv> <a>"dcat:Distribution" ; <dcat:accessURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:byteSize>"41672" ; <dcat:downloadURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:mediaType>"text/csv" ; <dct:license><http://publications.europa.eu/resource/authority/licence/CC_BY_4_0>; <dct:rights><http://publications.europa.eu/resource/authority/access-right/PUBLIC>; <owl:sameAs><https://zenodo.org/records/5652118>. 54 3.3.2 Documentation ÁWarning We will provide the link and screenshot of the documentation for each file that goes public. 3.4 Disseminate 55 3.4.1 EU Open Data Portal The portal is a central point of access to European open data from international, European Union, national, regional, local and geodata portals. It consolidates the former EU Open Data Portal and the European Data Portal. The portal is intended to: 1. give access and foster the reuse of European open data among citizens, business and organisations. 2. promote and support the release of more and better-quality metadata and data by the EU’s institutions, businesses, agencies and other bodies, and European countries, enhancing the transparency of European administrations. 3. educate citizens and organisations about the opportunities that arise from the availability of open data. It is funded by the EU and managed operationally by the Publications Office of the European Union in cooperation with the Directorate-General for Communications Networks, Content and Technology of the European Commission, responsible for EU open data policy. 56 We publish our data primarily on the EU open data portal for statistically processed datasets (datasets that contain the generalised characteristics of many data subjects without personal data that could identify them). 3.4.2 Europeana Integration We promised interoperability with Europeana, and after a lengthy process Reprex concluded aData Exchange Agreement with Europeana Sound (the music and sound aggregation center of Europeana hosted by the British Library’s Sound and Vision area) and sent pilot data from Latvia. 3.4.3 European Collaborative Cloud for Cultural Heritage Figure 3.2: Europeana is a bottom-up, decentralised data aggregator for the cultural heritage part of the music sector and the broader cultural and knowledge sector. Europeana is at the heart of the common European data space for cultural heritage, a flagship initiative of the European Union to support the digital transformation of the cultural heritage sector. Millions of cultural heritage items from over 3,500 data providers across Europe are available online via the Europeana website. We work to share and promote this heritage so that it can be used and enjoyed by educators and researchers, creatives and culture lovers across the world. While there is no agreed, cross-sectoral definition of “collections”, it is widely understood that in many cases, collections themselves are the entities that meet the information needs of music professionals or researchers (Wickett et al. 2013). The creation of collections is an important activity performed by music professionals and scholars as part of their work 57 4.1.3 Demonstrated support for required OMO functionality Everything OMO needs has already been demonstrated in production Wikibase environments: •authority control for creators, ensembles, organisations, venues; •multilingual and multiscript modelling for names, places, works; •cross-domain entity linking (work–recording–performance–rights–heritage); •event-based and entity-based models; •SPARQL validation, schema constraints, and automated reconciliation. This means OMO does not invent an untested paradigm: the consortium integrates proven practices from: • national registries, • performing-arts knowledge graphs, • the Slovak pilot and Finno-Ugric metadata federations developed inside the project. 4.1.4 Fits EU policy preference for open-source and trustworthy AI Wikibase is open-source, auditable, and non-proprietary. It aligns with: • the EU’s preference for open-source digital public infrastructure, • FAIR and CARE principles, • trustworthy AI requirements (provenance, transparency, versioning), • cross-border interoperability mandates, • decentralised data governance models. This makes it compatible with Europeana, EOSC/ECCCH, DCAT-AP, and the European Interoperability Framework. 4.1.5 The most widely used graph-editing interface in the world Tens of thousands of data stewards, librarians, researchers, and citizen-scientists already know how to edit Wikibase/Wikidata. This provides OMO with: • an immediate user base, • a ready-made contributor community, • institutional familiarity across Europe, • workflows already adopted in GLAM and research sectors. No alternative open-source system has remotely this level of adoption. 64 4.1.6 A hybrid model that fits real institutional workflows Wikibase uniquely accommodates: • spreadsheet-based workflows (Excel, CSV), • relational database exports, • statistical microdata reference linking, • complex semantic modelling, • API-based ingestion, • R and Python pipelines. It is a practical compromise between triple stores, document databases, and relational systems — perfect for a music ecosystem where many partners still rely on basic tools. 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline The data-to-policy pipeline defined in the Grant Agreement and documented in the Background chapter (see Section 2.5) provides the methodological backbone of OME. Wikibase is the component that makes this pipeline operational. 4.2.1 Wikibase supports each stage of the pipeline: 4.2.2 Data collection • imports from Excel, CSV, SQL, APIs, and legacy systems; • immediate linkage to persistent identifiers; • entity reconciliation as part of ingestion. 4.2.3 Validation and reconciliation • authority-control workflows for people, works, organisations, and places; • constraint-based quality checks; • SPARQL-driven validation; • alignment with external authority files. 4.2.4 Harmonisation and enrichment • multilingual labels and roles; • event-based and relationship-based modelling; • addition of contextual metadata by different institutions; • integration of domain vocabularies. 65 4.2.5 Activation for analysis • SPARQL endpoints for programmatic access; • JSON-LD, RDF dumps, and REST APIs; • R and Python pipelines use stable URIs for reproducibility. 4.2.6 Indicator construction • cross-domain indicators linking economic, cultural, rights, and heritage data; • entity-level referencing ensures indicators are traceable and verifiable. 4.2.7 Interpretation and contextualisation • experts review, annotate, and correct metadata through a human-readable interface; • provenance guarantees transparency. 4.2.8 Policy translation and observatory outputs • live, federated knowledge base powering the OMO front end; • entity profiles, metadata dashboards, and linked methodological documentation; • direct links from indicators to source entities and datasets. 4.2.9 Feedback loop • corrections flow back into the shared graph; • updated authority files update all downstream indicators; • the observatory improves over time. 4.3 Summary Wikibase is adopted because it best fulfils the policy requirements, semantic needs, technical constraints, and governance expectations described in the Background chapter: • It is proven in domains identical to ours. • It is aligned with the EU’s open-source, dataspace, and interoperability agenda. • It is widely adopted by the very institutions we must interoperate with. • It supports the entire Open Music Europe data-to-policy pipeline. • It allows the Observatory to operate as a decentralised, federated, evolving knowledge infrastructure. This architecture is therefore not an optional design preference but the only viable model for delivering the European Music Observatory described in the Grant Agreement. 66 5 Data coordination ĹNote The data-collection service provided by a European Music Observatory should help mapping, understanding and analysing the main characteristics, trends and idiosyncrasies of the music sector in Europe. Perhaps the most severe metadata problem in the music industry is the lack of interoperability among various musical work, recording, and rightsholder registers or effective identification services that would align interrelated objects (for example, an abstract composition and its recorded manifestation) across independent systems and applications. Areference model is an abstract framework or blueprint that provides a standard way of representing the components, relationships, and processes within a particular domain. It is a guide or template for designing and evaluating systems, ensuring consistency and interoperability. Statistical, record, and collection reference models help us find information about particular artists, their works, and their recorded manifestations. For example, a reference model helps us describe in a machine-readable way a sound recording of Beethoven’s 9th simphony. Areference object is a specific entity or instance used as a point of reference within a system or process. Registers as systems and authentic registration processes create such reference objects, such as the reference document of a newly registered musical work. Reference objects help us locate a specific recorded performance of the 9th Symphony in various music libraries, services, or catalogues. The concept of linked data requires access to identifiers for definitions in reference models (for example, an identifier of the definition “Collection”) and access to identifiers of authentic registers. 5.1 Ontologies and Vocabularies in the Open Music Observatory The Open Music Observatory (OMO) does not aim to create new ontologies. Instead, it integrates and reuses existing, well-established data models from the open data and cultural heritage ecosystems. In practice, this means that our data spaces and pilot projects (e.g. SKCMDb, Finno-Ugric Data Space, TextileBase) combine metadata drawn from: 67 •Open Government and Open Science: DCAT-AP — the European data portal metadata model. •Cultural Heritage and GLAM: EDM (Europeana Data Model), CIDOC CRM, and Records in Contexts (RiC). •Emerging Cultural Infrastructure: High-Definition & Collaborative Cultural Data Infrastructure (HDTO). •Common Metadata Principles: Dublin Core Terms (DCTERMS) — for general metadata interoperability. •Music Industry Standards: ISWC for works, ISRC for recordings, DDEX for metadata exchange and release categorisation. Our goal is not to add another layer of complexity, but to connect these existing models so that cultural heritage data, rights management systems, and research metadata can interoperate seamlessly. ĹOpen Music Observatory Ontology Approach • The Observatory uses a Wikibase-based ontology layer, derived from the Wikibase Ontology. This defines classes and properties that are compatible with the Wikidata ecosystem. •Equivalence links (owl:equivalentClass,owl:equivalentProperty) are added to map our entities to reference ontologies such as: –dcterms: (Dublin Core) –rico: (Records in Contexts) –crm: (CIDOC CRM) –edm: (Europeana Data Model) –dcat: (DCAT-AP) –and others used in European data spaces. • For industry standards (like ISWC, ISRC, and DDEX) that are not published in formal RDF/OWL form, the Observatory provides a minimal ontological scaffolding under the Open Music Ontology (omo) namespace. The source files of this minimal ontology are held at https://github.com/dataobservatory-eu/open-muisc-ontology This ensures that: • industry identifiers can be represented and queried alongside heritage data, 68 • and semantic equivalence can be maintained across research, rights, and cultural collections. In short, OMO acts as a semantic bridge, not as a new ontology. In summary: > The Open Music Observatory’s semantic stack connects ontologies — it doesn’t reinvent them. > It provides a thin, open, interoperable layer across public, research, and industry vocabularies to keep European music data FAIR and reusable. 5.2 Registers 5.3 Global and Permanent Identifiers 5.4 OpenCollections register services 5.5 TBC Building on the success of numerous data integration projects, we use Wikibase and its extensions, OpenCollections, for data coordination. This approach has many advantages. Wikidata is the largest open knowledge graph, with more than 1.5 billion semantic statements. We do not treat Wikidata as a primary data source because we work with highquality and well-maintained primary sources and generally work with data sources with higher data quality or trustworthiness than Wikidata. Where we see, like many other similar projects, is the reuse of many statements for starting and maintaining the structure of our knowledge graphs. Wikidata contains many statements that allow us to connect various graphs and access high-quality data sources. This functionality, often called “identity brokerage”, is valuable and perhaps indispensable in 2024. illustratio Wikidata is also an excellent platform for seeking consensus on terminology and definitions or aligning definitions used in different systems. We aim for compatibility with collections management, library and archive management systems, which use slightly different terminology to describe those persons as agents creating compositions or performing music. Wikidata, particularly our Wikibase instances, are good platforms to discuss such terminological alignment and immediately add it to the data model of our knowledge base. illustration 69 5.6 Adaptation to the music sector Rights management is critical in our @sec-music-economy-pillar because authors’ and neighbouring rights provide an essential part of most musicians’ income and the only income for composers. GLAM and rights management organisations use different data models for composing and performing music and for registering compositions or public performances of compositions and their recordings. Having reviewed the data models standardised by GLAM institutions (CIDOC, RiC, FBER), Wikidata, Music Ontology, and the Polifonia Ontology Network, we have not found an equally suitable formal, conceptual model for rights and collections management. Our initial data model results from more than a year’s work, but it is not set in stone. Such ambiguity is the most important reason why we avoid creating a central relational database system with a fixed schema and instead apply the dataspace model utilising a graph database. Wikibase provides an intuitive, easy-to-learn graphical interface for musicologists, rights management professionals, music economists and sociologists to agree on transparent definitions and to see the consequences of their choices. Our mapping guidelines in the Annex show how we deviated from the current Wikidata class definitions to retain compatibility with CIDOC, RiC, FBER and at the same time apply ISO standards of rights management. 70 6 Data Sources This chapter describes the data sources used in Open Music Europe and explains how they connect to the data-to-policy pipeline introduced in Section 2.5 and to the governance framework defined in the Data Management Plan. It is written in the same methodological structure as the Background and Architecture chapters: it explains where the data come from, how they were identified, how they will be governed, and how they will eventually be imported into the Open Music Observatory (OMO). The Data Management Plan (DMP) submitted at the beginning of the project has not been updated during the subsequent project cycles. As a result, many data sources that were collected, harmonised, or accessed during the project do not yet appear in the DMP’s Section “Data Summaries”,Section “Metadata and Standards”, or Section “Access, Retention, and Reuse”. For this reason, this chapter contains placeholders for the DMP manager indicating where updates are required before ingestion into the OMO can begin. Once the DMP’s Section Data Summaries and Section Metadata and Documentation are updated, the data listed in this chapter will be integrated into the Observatory and the ingestion workflow will resume. The vocabularies, classifications, licensing conditions, and provenance declarations listed here must be added to the DMP’s Section Standards and Vocabularies before ingestion. The purpose of this chapter is therefore twofold: 1) to describe the actual data sources used for Open Music Europe; 2) to prepare a DMP-aligned structure for importing them into the OMO once the DMP is updated. 6.1 How data sources were identified The identification of data sources followed the project’s data-to-policy pipeline described in Section 2.5. Work packages 1, 2, and 3 defined their indicator sets (D1.1, D2.1, D3.1), which in turn determined the necessary data inputs. These sources were documented in the internal methodological folders (see internal.qmd) and evaluated for: • accessibility and legal basis 71 • licensing and reuse conditions • data quality and completeness • compatibility with FAIR and GDPR principles • alignment with statistical and metadata standards (SDMX, DDI, GSIM) The resulting set of data sources falls into four functional groups: • administrative and register data • survey data • statistical and economic data • platform and streaming data These four categories reflect the cross-domain nature of the indicators: the music ecosystem cannot be understood solely through economic or cultural data, rights metadata alone, or platform data alone. The project therefore worked with a combined evidence base. When the DMP’s Data Summaries are updated, each source described here must be added as its own summary with variables, formats, provenance, and access conditions. 6.2 Administrative and register data Administrative sources form the backbone of WP1 and WP3, and they were essential for: • constructing sampling frames (WP1 T1.2) • validating enterprise and personal surveys • linking cultural, economic, and rights-related activities • generating indicators combining economic and cultural dimensions These sources include: • royalty and licensing records from CMOs • grant registers, project registers, and cultural funding databases • company registries and economic-activity registers 72 • non-profit and cultural-organisation registers • venue and festival registers • personal/professional registers where processing is permitted under GDPR Art. 6(1)(e) and Art. 89 In the Slovak pilot, these sources included historical microdata from the national KULT survey register, unpublished cultural-statistical microdata, and data held by Hudobné centrum and the Ministry of Culture. The DMP currently contains only partial references to administrative sources. To bring it into alignment, the DMP manager must update: •DMP Section: Data Summaries → add each administrative dataset with variables, formats, provenance, controller/processor roles •DMP Section: Legal and Ethical Compliance → include GDPR bases for each administrative source •DMP Section: Standards and Vocabularies → include controlled vocabularies used in registers (e.g. activity classifications, grant categories) These administrative data summaries will be added to the DMP before ingestion begins. 6.3 Survey data Survey data in the project covered: • enterprise surveys of MSMEs in the music sector • personal surveys on cultural participation, wellbeing, and music use • experimental diversity and circulation modules • harmonised components aligned with SDMX/DDI/GSIM Survey planning was delayed relative to the Grant Agreement timeline, but eventually harmonised with national cultural-statistics structures, including the Slovak KULT survey framework. For full DMP alignment, the following must be added: •DMP Section: Data Summaries → questionnaire versions, sample frames, variable lists 73 7.4 Statistical and Economic Data Statistical sources enabled European and national comparability across indicators. These included: • national accounts and satellite cultural accounts • labour-force data • business demography and structural-business statistics • external-trade and export data • cultural consumption and household budget surveys • price indices and cost-structure information These datasets were central to the economic modelling in WP1, circulation and diversity indicators in WP2, and societal-impact work in WP3. The DMP must include references, licences, and access conditions for all statistical datasets used. 7.5 Platform and Streaming Data Platform datasets (WP1, WP2, WP4) were collected via API-based sampling and automated scripts, including: • Spotify API samples (popularity, metadata, audio features) • YouTube API samples • playlist-localisation datasets • automated crawlers for repertoire discovery These data supported: - digital-market structure analysis - validation of economic indicators - repertoire and rights linkage - circulation and localisation studies Licensing and terms-of-service constraints mean these datasets have strict reuse limitations. The DMP must describe permitted uses and restrictions for each platform dataset. 7.6 Processing and Harmonisation After collection, datasets underwent several processing stages aligned with the architecture described in Chapter 4. These stages included: • cleaning and transformation • pseudonymisation of personal data 80 • cross-linking with authority files and identifiers • metadata enrichment • structural harmonisation (SDMX, DataCite, DDI) • conversion to formats suitable for ingestion into the OMO The openmusic-pipeline (WP4) implemented these processes in R, ensuring reproducibility and alignment with FAIR and OPA principles. The updated DMP must list controlled vocabularies, classifications, and harmonisation rules used. 7.7 Software Components Developed in WP4 WP4 created a suite of open-source tools that implement the pipeline and prepare data for the OMO. These are described here briefly, with technical detail deferred to annexes. 7.7.1 Data-Ingestion Tools These tools handle imports from Excel, CSV, SQL exports, APIs, and legacy systems. They include: - connectors for CMO and/or ministry datasets - survey-import scripts - API wrappers for streaming platforms 7.7.2 Validation and Reconciliation Tools These tools perform semantic alignment and quality checks: - authority-control reconciliation (ISNI, VIAF, ORCID, corporate registries) - SPARQL-based constraint checks - duplicate detection and entity merging tools 7.7.3 Harmonisation and Metadata Tools These implement the project’s semantic rules: • SDMX structure builders • DDI variable metadata generators • DataCite dataset metadata templates • vocabulary management utilities 81 7.7.4 OMO Integration Tools These tools prepare processed data for the Observatory: • Wikibase ingestion scripts • URI stabilisation and PID mapping utilities • JSON-LD and RDF exporters • R and Python client libraries for the OMO API The DMP must reference the software components used for metadata production and processing, as required by Horizon Europe guidelines. 7.8 Position of Data Collection and Processing in the Pipeline Data collection and processing bridge the conceptual work of WP1–WP3 and the semantic and technical infrastructure described in Chapter 4. They provide: • validated raw inputs • harmonised metadata • cross-domain entity linking • analysis-ready datasets These steps ensure the construction of traceable, reproducible indicators and policy outputs. 7.9 Integration with the Open Music Observatory Once the DMP is updated and the Data Summaries approved, all data collections and processing workflows described here will be integrated into the OMO: • datasets will be linked to persistent identifiers • metadata will be converted to humanand machine-readable forms • provenance will be documented in accordance with OPA and FAIR • ingestion routines will run automatically using WP4 tools Ingestion will begin once the DMP contains the complete, updated Data Summaries and metadata structures for all datasets. 82 8 Standardisation of Data & Terminology Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO Feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. Because standardisation is one of the key services of the envisioned European music observatory, we gave a lot of consideration to the standards to be applied, and the terminology negotiation process among the observatory’s stakeholders. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? 8.1 Business processes Since the Open Music Observatory is primarily a data dissemination hub, the definition of our services (Chapter 3) apply elements of the Generic Statistical Business Process Model (GSBPM), an international standard that describes and defines the set of business processes needed to produce official statistics. The GSBMP is accompanied by the General Statistical Information Model, which builds on the Data Documentation Initiative (DDI) and the Statistical Data and Metadata eXchange (SDMX) (Pellegrino and Grofils 2013). The DDI and SDMX are the foundations of working with social sciences archives, statistical microdata, and processed statistical data. Their key elements are described in the Resource Description Framework of the World Wide Web and can be used in Linked Data. Some elements of DDI are described with RDF: The DDI-RDF Discovery Vocabulary is a draft specification of the DDI Alliance. (Hartmann et al. 2024). Whenever possible, we rely in our observatory with this annotation; if that is not yet possible, we follow the DDI Lifecycle (3.3) Documentation (Data Documentation Initiative 2020). 83 8.2 Conceptual and information models Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? Numerous knowledge institutions store information about musical works, as well as natural persons (humans) who composed or performed these works and contributed to their recorded fixation. If we want to inquire about composers, we must know that a composer is always a human (animals or software agents with AI algorithms cannot be entitled to composer copyrights.) We also must know that a musical work is an abstract creation, manifesting as a notation (physical or digital sheets, MIDI files) or recording (analogue or digital-physical object, or a file.) If we want to validate the composer’s information connected to a recording of a particular musical work, we must access databases containing information about humans concerning some identifying properties of works or recordings. We imagine a future European Music Observatory that is not a specialised knowledge institution and is not a library, archive, museum, or statistical agency. Instead, it should be able to consolidate knowledge from all such institutions and find ways to bring together data from private enterprises and data collection programs to fill the information gaps of the European music sector stakeholders. Our services use the Wikidata Data Model as a data coordination and reconciliation model (Wikimedia Foundation n.d.). In this regard, we follow many successful EU and memberstate, (Alexiev et al. 2020; Diefenbach, De Wilde, and Alipio 2021; Rossenova, Duchesne, and Blümel 2022; Faraj and Micsik 2023) or music projects (Siler 2022). We particularly want to mention the excellent work of the University of Helsinki in creating WB CIDOC, a simple business process and data mapping between the Wikidata Data Model and the more complex CIDOC CRM used by extensive collection management systems (Kesäniemi, Koho, and Hyvönen 2022). 84 The StatDCAT-AP and the more general DCAT-AP definition of the EU Open Data Portal provide a bridge among library metadata systems, such as DCMI Metadata Terms (Dublin Core) for libraries, the World Wide Web DCAT standard for publishing datasets, and some core terms of the Statistical Data and Metadata eXchange. Figure 8.1: Our most important reference is the DCAT-AP 3.0 specification, and its extension to statistical data by the EU Open Data Portal. The Europeana Data Model (EDM) similarly provides a more straightforward connection tool among various library, museological or musical collections; it mainly builds on Dublin Core and offers equivalent classes for the more complex CIDOC CRM (Europeana 2017). We see no problem in connecting the EDM towards RiC. The CIDOC Conceptual Reference Model (CRM) provides an extensible ontology for concepts and information in cultural heritage and museum documentation (Bekiari et al. 2024). Last, we mention some novel standards and standard candidates related to documents, microdata, and metadata documentation, such as music survey questionnaires. The Records In Context (RiC) 1.0 CRM and ontology were adopted in November 2023 to replace four international archival standards with backward compatibility. The DDI-Discovery vocabulary is an evolving standard that aims to describe important DDI terms with the World Wide Web standard Resource Description Framework. To keep our systems future-proof, we adopt elements of RiC and DDI-Discovery to document our question bank and codebooks (International Council on Archives Expert Group on Archival Description 2023; Hartmann et al. 2024). 85 ĹNote A future European Music Observatory could help with coordinating European research activities in the music sector. An EMO could also develop tools to establish cooperation between various data collection bodies. The Observatory should, therefore, also be involved in setting standards and developing common EU wide definitions that are crucial for consistency. (European Commission et al. 2020, p80) Since the adaptation of the European Interoperability Framework and similar FAIR measures in open science, such terminological standardisation has taken place in the definition of formal ontologies, i.e., knowledge bases that software applications can use, too. The music observatory should have competent knowledge engineers and ontologists and should be involved in the discussions of sector-agnostic ontology, for example, on the possible improvements of CIDOC or EDM, for a better representation of music. There is also a need for the development of more usable and more widely accepted musicsector ontologies. In T5.1, we have reviewed the Polifonia Ontology Network and the Music Ontology, but we believe both have shortcomings for a full adaptation. 8.3 Identification & Entity Linking Entity linking, also referred to as named-entity linking (NEL), named-entity disambiguation (NED), named-entity recognition and disambiguation (NERD) or named-entity normalisation (NEN) is the task of assigning a unique identity to entities (such as famous individuals, locations, or companies) mentioned in a digital resource, such as a file. ĎTip The MusicBrainz free music database contains records of 20 artists named Paris (artists)�, and 15 locations using the same name Paris (locations)�, which all may enter a data-driven service as artists who must be credited for attribution or royalties, and as a place of an event, release, or publication. Connecting the word Paris to the correct person, group or location is the task of entity linking. Since the inception of the world wide web, data flows across organisations and countries, and the use of local identifiers is not a good solution. International organisations of music, heritage management, science, and national organisations are increasingly shifting to the use of persistent identifiers (or permanent Identifier or handle). ĎTip Apersistent identifier (or permanent Identifier or handle), is one that never changes, so that your bookmarks and links don’t break when a website or a database or an API service gets updated. 86 In 2024, there will be no European or international standard procedure for using PIDs, but several EU member states (Austria, Czechia, Germany, Netherlands) and other countries will have already adopted national PID strategies. Because Reprex is the current technical registrar of the Open Music Observatory, we losely follow the Dutch national strategy (Cruz and Tatum 2021) and the ID allocation practice of the Nationaal Archief, but this means no bias towards data partners in the Netherlands. The Dutch PID strategy does not use mandatory practices; it only recommends practices, and offers a thought-through consistent policy of using global identifiers that are not country-specific. The structure and management of global identifiers strongly correlates with the grade of achievable automation and the potential for innovation within and across different sectors of the media industries. Because of the prevailing problems of named entity linking, we are planning value added services to resolve named-entity recognition and disambiguation (NERD.) For this purpose, we are planning the use of AI (see Section 9.3). 8.3.1 Registers & Authority Files Registers record every data subject belonging to a category or class: every music publisher operating in a jurisdiction, music composer with copyright claims, or statistical dataset published. Registers are essential in identifying persons and objects (“things” in information science.) Authority files play a similar role in collections management: they provide identification information about persons or objects and tools for disambiguation. Authority files, for example, give the preferred name title for persons and musical works when available in different name or title formats, and they provide a language-independent, machine-readable identifier pointing to the correct name title. For two or more authors or performers with the same name, these identifiers help reference the proper person (or object.) Registers are valuable and indispensable for many digital workflows. They serve as the foundation of various processes, such as statistical sampling (determining who should receive a questionnaire) or copyright management (deciding who should receive the royalty payment). Their absence or inefficiency can significantly hamper these operations. Unfortunately, the music industry has long missed access to reliable, open registers. The reasons for this are beyond the scope of this report, but we highlight that the underlying reasons for closed and not interoperable registers are deeply rooted in the conflicts of interests among different sub-sectors of music and are unlikely to be solved in a short time. Therefore, music enterprises, researchers, professionals, and curators will need identification services and identity brokerage services for a long time. Creating and maintaining high-quality registers require significant professional and financial commitments, and they can form a vital service of a future European Music Observatory. Currently, we are experimenting with three service levels in the Open Music Observatory. 87 • We create our own transparent and interoperable identifiers within the OMO for persons and their groups (ensembles, bands, orchestras, associations…), legal persons (music businesses, collective rights management agencies, …), events (recording, composing, performing events, festivals, conferences, …), musical works and their manifestation (books, works, recordings, sheets.) • We create integrity brokerage services and middle-term identification via Wikibase and Wikidata. Our identifiers are connected to middle-term Wikidata and Wikibase QIDs, which also serve as graph nodes to registry, library, collections, and industry-specific identifiers. • We are piloting data improvement services that can find erroneous identifiers or add correct identifiers to various datasets. 8.3.2 Open and persistent identifiers In line with the practice of the Netherlands, we prefer the use of the following identifiers: ISNI: preferred persistent identifier for names of people and groups. The use of ISNI is also preferred by Apple Music, Spotify, and as a pilot it was introduced by Teosto, the Finnish national collective management society; it is being considered in many use cases for adoption in all CISAC societies. ISNI is the ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. (Camp, Lieber, and IFLA 2022) For legal persons, we are discussing the terms to use the OpenCorporates ID, because many organisations at this point do not have an ISNI. ORCiD: preferred persistent identifiers for music researchers and scholars. This is in line with the Horizon Europe and the European Open Science Cloud recommendations; ORCiD itself only adds functionality to ISNI; i.e. each ORCiD ID is at the same time registered as an ISNI. VIAF: VIAF is the shared authority file of national libraries. It offers more services than ISNI and includes an ISNI for the author. DOI: we use the Digital Object Identifier for publicly released documents. ISBN: We issue ISBN identifiers for long-form publications of our partners. (ISO 2017c) 8.3.3 Not open, music-industry specific identifiers Book and music sheet publishing uses the ISBN and ISWN, professional and magazines and scholarly music journals use the ISSN, and the music rights management uses ISRC and ISWC. These standards usually resolve an identifier to some network location where metadata or the object itself can be found. There are many advantages and disadvantages of this model. 88 For example, the ISWC identification of musical works is the backbone of copyright management, and it is a closed and consistent system developed over many decades by the member organisations of CISAC. The downside of this closed system is that the metadata about the works identified by ISWC is strictly available only to CISAC member societies. While CISAC offers an API for the individual lookup of ISWC for one example of a musical work, currently it does not allow bulk access to the registered data. We have already started a discussion with some music industry registers about connecting the Open Music Observatory to their systems. We are planning to present our proposals on the CISAC Good Governance seminar to be held in December 2024. Musical works ISWC: nternational Standard Musical Work Code is a unique identifier for musical works. It is adopted as international standard ISO 15707 (ISO 2022). OpenCollectons ID: Our ID for music works (only if we publish data about them.) Sound recordings ISRC: The International Standard Recording Code (ISRC) is the international identification system for sound recordings and music video recordings. (ISO 2019b; International ISRC Registration Authority 2021) OpenCollectons ID: Our ID for sound recordings (only if we publish data about them.) Music sheets ISWN: The International Standard Music Number currently identifies published music sheets (ISO 2022). ISBN-13: Before the introduction of ISWN, published sheets were identified by the ISBN book identifier. ISBN-10: The older format of the ISBN book identifier, which predates both the ISWN and the 13-digit ISBN used to identify music sheets. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. This is our preferred identifier for not published sheets. (ISO 2017c) For unpublished works, our preference is the use of the brand-new ISO-standard ISCC because it was designed precisely for the use case we were looking for. It is free to generate, generated from digital content (or its digital copy), and can connect various local or lesserused identifiers. Datasets DOI: DOIs are assigned to each distribution of a dataset. As datasets are often continuously filled, these datasets will have periodic versions with versioned DOIs (from Zenodo.) OpenCollectons ID: Our ID for unversioned (continous) datasets, pointing to the latest available version of the data. Codebooks URI: Whenever possible, we use standard codebooks of SDMX or Eurostat, and provide a URI to the codebook, and provide dereferencing to the codebook definition. OpenCollectons ID: Our ID for our codebooks, regardless if they are same as the SDMX/Eurostat standards, or we create a non-standard coding for a novel dataset. Questionbank URI: Whenever possible, we use standard questionnaires, and provide a URI to the codebook, and provide dereferencing to the DDI questionnaire item definitions. OpenCollectons ID: Our ID for questionbank items. 89 9.2.1 Data Health Services for Collective Management Entity linking and data linking are among the biggest technical problems in rights management. Because music authors, producers, and performers have three royalty streams and do not share an interoperable registry, the connection of musical works (compositions, ideally identified by an ISWC code), their sound recording manifestations (identified on all digital services with and ISRC code), and the various identifiers of performers require costly manual and technical identification. There are numerous projects underway in the music industry to resolve this problem going forward. In the United Kingdom, PRS’s Nexus programme� is developing a solution with the provisioning of preliminary ISWC registration to keep the recording and composition connected from the birth of a new recording. The Open Music Europe project, on the other hand, is pioneering a different route for already existing sound recordings, with the linking of public sector catalogues of heritage and library collections with rights management information; particularly with relying on the VIAF shared authority files. SOZA and Reprex are expected to present their MVP on the CISAC Good Governance seminar in December 2025. Modern registers typically assign a unique identifier, known as a URI, to their data subjects (our registered objects). A ‘Cool URI’, which resembles a URL, offers a practical advantage. When used as a URL, it generates a human-readable HTML file about the registered person or object. This can be particularly useful when processed by a graph application, as it provides crucial information about this person or object in a machine-readable (XML, JSON, TTL, or NQUAD) file. For example, the VIAF identifier number 89006617 can be placed into the http://viaf.org/ viaf/89006617 URL, which provides as access to the cataloging information of works created by, or written about the great etnomusicologist and modern composer, Béla Bartók. Modern platforms, such as Spotify, use similar identifiers. For example, the Spotify Artist ID 2fIUlieTjLTaNQUIKHX5B8 resolves to Celeste Buckingham’s available recordings on the platform via the URL https://open.spotify.com/artist/2fIUlieTjLTaNQUIKHX5B8. The problem is that music creators are often present on more than 200 digital platforms, each of which has its identifier policy and requires the repeated import of the artists’, works’, and recordings’ data. To consistently report such metadata is costly and complex, even for major labels and publishers with a dedicated IT system. No wonder we saw before our project in our own Feasibility study that more than 50% of artist data needed fixing on digital platforms. Relying on many local identifiers on otherwise interconnected computer systems will always create a costly and error-prone data exchange. Unfortunately, the music industry has never agreed to use genuinely open, high-quality registers. These changes were made during the period of our project. For example, large platforms like Apple, Spotify, and some collective rights management organisations started using the ISO-standard name identifier (ISNI) to avoid the high prevalence of multiple same-name persons and musical groups. This transition 96 is yet to begin, and it is incomplete, so the music sector will likely need to invest large IT resources into entity resolution in the next decade. 9.2.2 Sustainability Reporting for Music Organisations The Music Innovation Hub and Reprex will develop a CSRD-compliant sustainability reporting tool in 2024-2025. The reporting tool aims to provide an accurate and affordable ESG reporting facility that follows the European ESRS standards for music enterprises that create their financial reports according to the simplified reporting rules allowed by member states for microenterprises. More than 95% of European music enterprises (in some member states, this reaches 100%) apply simplified financial reporting. For such companies, there are no CSRD-compliant ESG reporting tools. We identify the reason for this market failure as follows: □The CSRD Directive imposes the responsibility of connected financial sustainability reporting on large and public companies and applies it to their entire value chain. The music industry lacks such large enterprises that would have taken a piloting role or played a pivotal role in establishing the standards. □Music enterprises and their trade associations do not act proactively because they believe they must follow the data provision instructions of the directly affected B2B buyers, financiers, or corporate sponsors. □The standardisation body EFRAG has de-prioritised the cultural and creative industries in setting industry-specific standards favouring sectors with a much higher adverse environmental impact. 97 □While small music businesses do not feel a compliance push, as they are not directly responsible for applying the ESRS, they also miss out on the opportunities provided by green financing and insurance. ⊠MiH and Reprex will pilot a service suitable for microenterprises, reducing compliance costs from 1500 euros to 500 euros per entity. ⊠This new application will rely on the Open Music Observatory’s Music Economy and Sustainability pillars and will derive its benchmarks, science-based targets and coefficients, and input-output tables. The MVP of this service was developed with a MusicAIRE microgrant, and it is the project’s background. A scale-up will be demonstrated with the use the Open Music Observatory’s open data API. 9.2.3 Listen Local Figure 9.2: The Feasibility Study On Promoting Slovak Music in Slovakia And Abroad is an important background of our project. In 2020, with a microgrant from the Slovak Arts Council, we created a Feasibility Study and a demo application called Listen Local (Antal 2020b). The study examined why the Spotify algorithm struggled to recommend Slovak music within Slovakia for Slovak people. We also created a demo application that modified the user’s Spotify recommendations to voluntarily comply with the local content guidelines applicable to local radio stations. The user could also listen to a lower or higher percentage of regional works. Our critical finding was the very pool data coverage and quality of the Slovak repertoire, which is mainly sent to distribution without the professional assistance of a commercial music label. Self-releasing artists and micro labels do not have the necessary metadata know-how, IT and data specialists to prepare their new releases for algorithmic curation by recommender engines of digital streaming platforms, radio stations, or large festivals. 98 Figure 9.3: Our conceptual demo application was able to make recommendations on voluntarily meeting the local content guidelines, but it was only supported by a relatively small Slovak Demo Music Database, and could only work with Spotify, which has the most transparent and open API of all streaming providers licensed to the territory of the Slovak Republic. We aim to develop applications to create a local content-aware public performance music stream. ⊠HearDis! aims to integrate such location-aware metadata into its background music playlisting service. ⊠We are planning Listen Local applications for radio stations to voluntarily review their current playlists for compliance with local content regulations and, if they fail to reach the statutory local content quotas, to recommend suitable recordings to their playlists. ⊠The OMO will disseminate the necessary data for these new services. □The data is not yet available, as the creation of the Slovak Comprehensive Music Database is a task of its own that will be ready by November 2025 in WP2. 9.2.4 Unlabel Unlabel is a planned service aimed at self-releasing artists and micro labels that need a functional data/IT department. Therefore, they are at a disadvantage compared to significant independent and major releases because they usually need to meet the high documentation standards necessary for a successful digital distribution strategy and engagement with algorithmic curation of streaming-, radio-, or festival playlists. Self-releasing artists and micro labels bring ill-documented new content to digital distributors like ALOADED. Digital distributors must maintain an arm’s length standard for all 99 labels, small or large, independent or major. ALOADED or other distributors cannot crossfinance the data problems of self-releasing and micro-label artists from the client revenues of more prominent labels. We identify the problem as a market failure and a technical failure: □In some developing markets, insufficient royalty revenues do not allow the presence or professionalisation of record labels with an IT and data management function because the payment of IT or data specialists or to keep external suppliers at least on a retainer cannot be financed from the label-artist revenue split. □Manual metadata provision without metadata specialists and tools leads to inferior data quality. Our feasibility study has shown that more than 50% of the releases have data shortcomings, and 17% have poor data representations that make algorithmic recommendations for these releases impossible. This creates a vicious circle because poor data quality translates into low visibility, low usage of such repertoire, and, therefore, low income. The cost of data improvement has no sustainable financial basis. ⊠In 2025, ALOADED, Reprex, Slovak Music Center and SOZA will conceptualise and plan a new public-private business model that aims at those rightsholders who do not have a technically proper label representation as a substitute for non-available market services. Our planned “Unlabel” service will provide documentation and metadata improvement services for self-releasing artists. This service, similar to current white-label services, will strictly address market failures and not compete with label services. We aim to provide a necessary level of data consolidation and improvement so that these artists can have equal opportunities in digital distribution services. The service will be connected to the Slovak Music Dataspace and its Slovak Comprehensive Music Database. We will provide a PPP business model for the onboarding and proper documentation of self-releasing artists on a large scale and the efficient, API-based provision of their digital distributor. Aloaded will provide the distribution services, Reprex will provide the data services, and SOZA and the Slovak Music Center will work out the details of minimal customer service for such labels. 9.3 Use of AI systems For the entity linking, related to our planned value added Section 9.2.1, we are planning to use in the future AI algorithms, particularly inference engines. The main goal of the system is to help matching correctly named entities, particularly rightsholders, musical works and recordings. The system is not yet in place. An adequate description will be provided for overview and will be brought to the attention of the Ethics Advisor during the upcoming meeting of the Ethics Board. 100 We cannot provide a full risk assessment because the service is not planned in detail yet. However, our preliminary risk assessment suggests low levels of risk, partly, because we plan to deploy AI in music/culture, which as a domain not seen as a high-risk area by the European regulation, and partly, because our system will not autonomous, will retain human-in-control, and will not influence the decisions or anyhow engage with end-users. We were conscious of the potential risk involved, and both the control structure and the data governance were planned over the course of 10 months. □Is the AI system designed to interact, guide or take decisions by human end-users that affect humans or society? No. The system will only help qualified persons in rights management to faster and more efficiently preview potentially unlinked entities. □Could the AI system affect human autonomy by interfering with the end-user’s decision-making process in any other unintended and undesirable way? No. The system in no way is considered as an end-user system. ⊠Please determine whether the AI system (choose as many as appropriate) overseen by a human: Is overseen by a Human-in-Command. ⊠Have the humans (human-in-the-loop, human-on-the-loop, human-in-command) been given specific training on how to exercise oversight? Yes. The system is not making autonomous decisions. ⊠Is your AI system being trained, or was it developed, by using or processing personal data (including special categories of personal data)? Yes. ⊠Did you put in place any of the following measures some of which are mandatory under the General Data Protection Regulation (GDPR), or a non-European equivalent? Yes. ⊠Data Protection Impact Assessment (DPIA) Yes. ⊠Designate a Data Protection Officer (DPO)24 and include them at an early state in the development, procurement or use phase of the AI system? Yes. ⊠Oversight mechanisms for data processing (including limiting access to qualified personnel, mechanisms for logging data access and making modifications)? Yes. □Measures to achieve privacy-by-design and default (e.g. encryption, pseudonymisation, aggregation, anonymisation)? Not applicable for NERD. The aim of the application is to detect errors in name attribution and to protect the moral and economic rights of the (named) rightsholders. ⊠Did you implement the right to withdraw consent, the right to object and the right to be forgotten into the development of the AI system? Yes. ⊠Did you consider the privacy and data protection implications of data collected, generated or processed over the course of the AI system’s life cycle? Yes. ⊠Did you consider the privacy and data protection implications of the AI system’s non-personal training-data or other processed non-personal data? Yes. 101 We do not consider that the system has wider risks or negative impacts. The algorithm is designed to cure sources of data biases that result in a late or missed payment for some rightsholders. 102 10 Conclusions & Next Steps: Towards a European Music Observatory In this report, we introduced the Open Music Observatory, a complex data curation and information management system with a novel observatory governance model. The system is technically ready and has been tested with real-life use cases and scenarios: it handles complex data extract-transform-load processes, GDRP and confidentiality problems, and many aspects needed to make a data (sharing) space work. Because it is based on opensource components and on open data standards, it is set on an optimal path to grow as a data ecosystem. The next development phase of this observatory will be critical: we must persuade music sector stakeholders to trust and use our system for data sharing well beyond the stakeholders of the Open Music Europe project. Once we reach out and onboard external partners, we will certainly have to modify business processes and permissions and provide further technical help, software components, and manuals. Besides these obvious iteration steps, we would also like to agree on a governance model that allows the extensions and the ownership transfer of this Open Music Observatory to a truly European, shared, publicprivate observatory that can serve the music sector for a long time. The European principle of subsidiarity requires that decisions be taken as closely as possible to the citizens they affect. In cultural policy, this means that responsibilities are distributed across multiple levels: in some Member States, culture is managed regionally or provincially; in others, nationally. Beyond public administrations, many important datasets are held by private actors — collective management organisations, platforms, or archives. Any attempt to centralise music data governance would therefore risk losing both legitimacy and local relevance. Subsidiarity must be built into the design of the Observatory. The European Interoperability Framework (EIF) provides a layered model — legal, organisational, semantic, and technical — for reconciling governance across institutions. The Data Governance Act (DGA) codifies the same principle: Member States retain stewardship over sensitive datasets, but EU-level standards ensure they can circulate securely and comparably across borders. The Data Space Support Centre (DSSC) extends this approach into practice, developing blueprints and building blocks that allow decentralised initiatives to scale. Together, these frameworks show how subsidiarity and federation are not barriers but design principles for data spaces.1 1On subsidiarity and federation: the Data Governance Act (regulation_dga_2022_868?) and the European Strategy for Data (European Commission 2020). On technical frameworks: DSSC’s blueprints (dssc_blueprint_intro_2025?;dssc_blueprint_interop_2025?). On 103 10.1 Co-creating a governance model for the Observatory Our current roadmap for the institutionalisation of the Observatory is based on the example of Europeana, a joint collection of European libraries and museums. Europeana 1.0 was built over 2.5 years in the European Digital Library Network project. It remained a decentralised network of three organisations: the Europeana Foundation, which is the technical operator of the digital collection; the Europeana Network Association, which is free to join for all interested parties; and the Aggregators’ Forum, which is a technical coordination body to help those technical data providers that send or exchange data with Europeana. To draw on this analogy, the Open Music Observatory is being created by the Open Music Europe project. As stated in our Grant Agreement and the connecting Consortium Agreement, we treat the duration of the Open Music Europe project as a prototyping and development phase when multiple institutionalisation forms are still possible. Based on our Agreements, we formed the Observatory Stakeholder Network as a stakeholder group to set priorities and express opinions on our work and the potential longer-term institutionalisation alternatives. We consulted several stakeholders throughout the project, and eventually, following their advice, we decided to form this advisory council when we already had a working data dissemination infrastructure and reviewable data in it to start their work. Not prejudicing any later presented organisational proposals, for the time being, we follow functionally the organisation of Europeana, because we need to deal with similar problems. governance: BDVA (bdva_discussion_paper_2023?) and the Federation Working Group (federation_wg_position_2023?). 104 Europeana started out as a project of a few national libraries and after 16 years of existence, it aggregates digital collections from more than 3000 organisations of Europe. Such a largescale but decentralised organisation requires a multi-tier, multi-functional governance structure, where large stakeholders can take ownership, smaller stakeholders can democratically participate and cooperate, and technical providers can work separately on technical-only problems. ⊠The Observatory Stakeholder Network is presented as a consulting body where future founders of the Observatory itself, and a Music Observatory Association of music data users, providers may find their future role. We hope that some of the initial members will become founders in the future European Music Observatory, and others will be forming a democratic Music Observatory Association for grassroot aggregation and dissemination. (Observatory Stakeholder Network�). ⊠Our Open Music Data Exchange Forum is a consultation body, similar to the Aggregators’ Forum in Europeana, for large data partners. ⊠As stated in our Grant Agreement, we will offer various institutionalisation proposals by the end of our project. If the representative stakeholders of the European music sector will choose not to create a European Music Observatory, we will continue to operate the Open Music Observatory with any interested party. (See project website: openmuse.eu, project data on CORDIS, (Open Music Europe 2023).) 10.1.1 Observatory Stakeholder Network The Observatory Stakeholder Network is a temporary organisation defined by the Open Music Europe project Consortium Agreement. It is a volunteer advisory body of the observatory. ĎTip The Europeana Network Association (ENA) is a strong and democratic community of experts working in the field of digital cultural heritage. We are united by a shared mission to expand and improve access to Europe’s digital cultural heritage. The Association is free to join and we encourage our members to get involved and benefit from all the ENA has to offer. We will invite every organisation that has promised data for a future European Music Observatory or has shown interest in the previous CEEMID or Digital Music Observatory collaborations to join our network. In the longer term, we envision some pan-European representative and umbrella organisations becoming founders or board members in the more permanent observatory organisation. For individuals, research groups, and micro-enterprises, we will suggest setting up the European Music Observatory Network Association (or similar entity), learning from the experience of Europeana. 105 ⊠Statistical data (indicators and their datasets) defined by, or requested by members of the Observatory Stakeholder Network. ⊠Datasets about information gaps identified by the EMO feasibility study. ⊠Records of questionnaires, question banks, and any structured datasets used for the creation of the statistical datasets above. ⊠Collection datasets about musical works and their manifestations in sound recordings or musical sheets. ⊠Collection datasets about music events, including events of composition, recording, or live performance. ⊠Encyclopaedic, demographic, biographical data about music professionals and music enterprises. ⊠Collection datasets about books, publications, statutes and laws, standards related to music. The EMO feasibility study Curators are forming collections with the application of unity criteria which allow them to decide which musical work, sound recording, music enterprise or person is included in a collection list. The curators are responsibility for the comprehensive application of the unity criteria and ensuring that their collections are up-to-date (Wickett et al. 2013). Some examples of music data curation Hitlists use some kind of popularity metrics, and they follow rigorous rules which sound recordings are included every week, or year. Collective rights management organisations create comprehensive lists of works and sound recordings registered for rights protection and exploitation. Statistical agencies create business registers to carry out data collection. Who can curate our datasets? Any music professional or scholar can curate datasets in agreement with our Collection Guidelines. The quality review mechanisms will be set by the Observatory Stakeholder Network from a content point of view, and the Open Music Data Exchange from a technical point of view. 112 11.2 Topical Pillars Figure 11.2: The extended five pillars, with sustainability added. 113 ĹNote The suggested four-pillar model would categorise data-collection and analysis along the following lines: • Measure the contribution of music to the EU’s economic and legal environment, from a systemic perspective (Pillar 1). • Monitor the cross-border flows of repertoire, the mobility of artists and diversity (national, linguistic, genre-based) (Pillar 2). • Assess music’s impact on society and citizenship: how audiences access and consume music; how citizens participate in professional and not-for-profit music activities; the scale, value and quality of music education and training (Pillar 3). • Provide a framework to develop prospective research on the future of the music sector, supporting innovation and developing understanding of emerging practices from various perspectives (business, tech, policy) (Pillar 4)(European Commission et al. 2020, p30). 11.2.1 Music Economy ĹNote Main potential data-collection and research areas identified at this stage: � Macroeconomic patterns and trends (e.g. employment, revenue, competition) � Value chain mapping and analysis (e.g. characteristics of music organisations, copyright collection, collective management, remuneration of artists, spill-over effects) � Legal aspects (e.g. tax, labour laws, social security, contracts, case law) � Business regulations (e.g. live music regulations, consumer protection, licensing, anti-piracy rules) (European Commission et al. 2020, p114) According to the EMO feasibility study, one “of the key findings of the AB music working group report was a substantial appetite for cross-sectoral, neutral and comparable data on the music business at EU level. While recent studies (e.g. EY “Creating Growth” study) have attempted to measure the impact of music on the EU’s economy, systematic and comprehensive metrics do not exist at this stage.” (European Commission et al. 2020, p113) Deliverable D1.1 of Open Music Europe, Economy of Music in Europe: Methods and Indicators identifies critical research questions, data sources and gaps,and data collection methods regarding the economy of music in Europe . (Antal, Kmety Barteková, and Remeňová 2023) 114 The deliverable begins by reviewing definitions of “the music industry”, the categorisation of musical activities within the system of national accounts (SNA) and statistical classifications of economic activity (ISIC and NACE), and the three primary income streams within the music industry (the live music, author or publishing, and recording streams). It then turns to the topic of value, first identifying the types of value created by musical activity and then considering legal and economic dimensions of valuation. 115 Figure 11.3: Our website documents with visualisations each dataset, apart from providing links to the latest distribution downloads with visualisations (in zip) or the access points on the various dissemination nodes. The illustration is an experimental dataset from our background CEEMID catalogue. After introducing the concept of mixed enterprise and personal surveying as a means of improving insight on informal economic activity in the sector, the deliverable identifies data gaps relevant to national policy in our pilot study target country of Slovakia, critically reviews the data gaps relevant to EU-level policy first identified in the EMO feasibility study,and proposes data collection methods appropriate to filling specified data gaps. 11.2.2 Music Diversity According to the EMO feasibility study, “creating reliable tools to monitor what kind of repertoire circulates on digital platforms or via radio will require access to vast amounts of data from Digital Service providers (DSPs) or third party aggregators. The notion of European repertoire has to be clarified and very well defined; notion of language, of origin, of nationality, country of production, genres, and it should not be limited to the language sung in a given song. […] A European Music Observatory should also look into the possibility to collect regular data on the circulation of European repertoire at song and/or artist level, considering live performance/radio/ digital use, which will be available at a weekly/monthly/yearly basis to the music sector.” (European Commission et al. 2020, p34) The Open Music Europe project is developing two tools for capturing and turning the aforementioned data into informative indicators. In WP1, the project is developing big data statistical sampling algorithms to avoid the need for “access to vast amounts of data from Digital Service providers (DSPs)”. WP2 is working on a taxonomy and GDPR-conform representation of the “notion of language, of origin, of nationality, country of production, genres” based on our background (Antal 2020b). The result of this work will be the Slovak Comprehensive Music Database, which will create clear taxonomies and allow users and software applications to determine aspects of “Slovakness” for each sound recording. 116 The possibility “to collect regular data on the circulation of European repertoire at song and/or artist level” is an issue that must be addressed with strict adherence to GDPR. We are developing an opt-in, opt-out mechanism in Slovakia that will be replicable in all other jurisdictions in accordance with GDPR. (For potential non-European artists, we will apply GDPR, too.) ĹNote Main data-collection and research areas identified at this stage: � Cross-border circulation of works/repertoire (e.g. building common definition and indicators, mapping of cross-border access, sales and consumption flows � Cross-border mobility of artists and professionals (e.g. cross-border live performances, mobility of professionals, international music events) � Cultural diversity aspects (e.g. languages, genres, types of productions) � Legal aspects (freedom of movement, state aid, etc.) (European Commission et al. 2020, p115) An interesting opportunity here is that many European countries already collect information on subsidised music operators (e.g. associations or not-for-profit projects), not to mention the wealth of information available through Creative Europe supported initiatives, which could provide this Pillar with interesting data. (European Commission et al. 2020, p116) 11.2.3 Music Society Regarding music and society, D3.1 considers the reuse of various survey programs with retrospective survey harmonisation, and WP3 is planning to conduct surveys in 2025. Data will be added to the catalogue as it is becoming available. ĹNote Main data-collection and research areas identified at this stage: �Education, training, personal development �Audiences (music consumption, interaction, participation in music events, etc.) �Music and society (not-for-profit sector, associations, social inclusion, amateur music, heritage, participation in music) �Normative Aspects (broadcasting quota rules, diversity promotion schemes, freedom of speech rules) �Music and the environment (carbon footprint of venues, touring, festivals, merchandise manufacture, streaming services; issues around noise/neighbourhood impacts; good practice in these areas). (European Commission et al. 2020, p116) 117 11.2.4 Innovation The definition of the Innovation pillar in the EMO feasibility study is more a topic to be covered than a data need description. This pillar is less data-driven in that it will rely mostly on research conducted on topics relating to changes in the market place, new business models, disruptive technologies, etc. A European Music Observatory will have the latitude to pick certain topics based on priorities and input from sectoral stakeholders. An EMO should consider setting up an “innovation experts’ advisory committee,” constituted of respected professionals in their field who are known for their forward thinking views, to help identify key themes to be studied. (European Commission et al. 2020, p37) We will initiate an informal music innovation expert’s roundtable to discuss potential data needs in this pillar. 11.2.5 Sustainability In the EMO feasibility study the definition of sustainability was mentioned among the innovation topics. Because of the triple transition, introducing the Corporate Social Responsibility Directive and the European Sustainability Reporting Standards have increased the interest and need in sustainability data; we decided to create a separate topical pillar for environmental and social sustainability, or governance indicators (ESG.) We will publish datasets that will be used in the value-added service described in Section 9.2.2. 118 References 2020, SEMIC. 2020. “Providing Sustainable Data Services Through Wikibase and Wikidata.” https://joinup.ec.europa.eu/sites/default/files/custom-page/attachment/202011/Parallel-track-4_B-Fischer_J-Thill_A-Angjeli%20final%20ppt.pdf. Albertoni, Riccardo, David Browning, Simon Cox, Alejandra Gonzalez Beltran, Andrea Perego, and Peter Winstanley, eds. 2020. “Data Catalog Vocabulary (DCAT) - Version 2.” W3C. https://www.w3.org/TR/2020/REC-vocab-dcat-2-20200204/. Alexiev, Vladimir, Plamen Tarkalanov, Nikola Georgiev, and Lilia Pavlova. 2020. “Bulgarian Icons in Wikidata and EDM.” Digital Presentation and Preservation of Cultural and Scientific Heritage 10: 45–63. https://doi.org/10.55630/dipp.2020.10.2. Anders, Wallgren, and Wallgren Britt. 2007. Register-Based Statistics. Administrative Data for Statistical Purposes. 1st ed. Chichester:United Kingdom: John Wiley & Sons Ltd. Antal, Daniel. 2020a. “Central And Eastern European Music Industry Report 2020.” CEEMID, Consolidated Independent. https://doi.org/10.13140/RG.2.2.21450.31686. ———. 2020b. “Feasibility Study on Promoting Slovak Music in Slovakia & Abroad.” https://doi.org/10.5281/zenodo.6427514. ———. 2023. “Pilot Program for Novel Music Industry Statistical Indicators in the Slovak Republic.” Zenodo. https://doi.org/10.5281/zenodo.8399254. ———. 2024a. “A szlovák adatkicserélési tér magyarországi föderációjának lehetőségei.” In Az oktatás, a kutatás és a közgyűjtemények digitális transzformációja felsőfokon : NETWORKSHOP 2024 : 33. Országos Informatikai Konferencia : 2024. április 3–5. Eszterházy Károly Katolikus Egyetem, Eger, edited by József Tick, Károly Kokas, and András Holl, 192–98. Budapest: HUNGARNET Egyesület. https://doi.org/10.31915/ NWS.2024.25. ———. 2024b. “Building a Music Data Sharing Space with Wikibase.” Open Music Observatory. https://doi.org/10.5281/zenodo.17078911. ———. 2024c. “Trustworthy AI and Data-Sharing Spaces for the Slovak Music Centre. Poster Presentation at the IAMIC Conference 2024 on November 21, 2024, at Music Austria, Vienna.” Open Music Observatory. https://doi.org/10.5281/zenodo.16540605. ———. 2025a. “SKCMDb. Interoperability of Music Libraries and Archives with Public and Private Music Services. Presentation at the IAML 2025 Conference in Salzburg, Austria Held on the 7th of July 2025.” Open Music Observatory. https://doi.org/10. 5281/zenodo.16634558. ———. 2025b. “Slovak Music Data Sharing Space. Poster Presentation at the IAML 2025 Conference in Salzburg, Austria Held on the 7th of July 2025.” Open Music Observatory. https://doi.org/10.5281/zenodo.15814286. ———. 2025c. “A Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem.” Open Music Observatory. https://doi.org/10.5281/zenodo. 119 17244314. Antal, Daniel, Mária Kmety Barteková, and Katarína Remeňová. 2023. “Economy of music in Europe: Novel data collection methods and indicators.” Zenodo. https://doi.org/10. 5281/zenodo.8334648. Antal, Daniel, and Anna Márta Mester. 2025. Open Music Registers.https://doi.org/10. 5281/zenodo.14767717. Antal, Daniel, Ieva Pigozne, and Anna Márta Mester. 2025. “Remapping the Livonian Coast: A Multilingual Gazetteer of the Settlements of Northern Kurzeme.” Finno-Ugric Data Sharing Space. https://doi.org/10.5281/zenodo.15668712. Artisjus, HDS, SOZA, and Candole Partners. 2014. “Measuring and Reporting Regional Economic Value Added, National Income and Employment by the Music Industry in a Creative Industries Perspective. Memorandum of Understanding to Create a Regional Music Database to Support Professional National Reporting, Economic Valuation and a Regional Music Study.” Bekiari, Chryssoula, George Bruseke, Erin Canning, Martin Doerr, Philippe Michon, Christian-Emil Ore, Stephen Stead, and Velios Athanasios, eds. 2024. “Definition of the CIDOC Conceptual Reference Model.” CIDOC CRM Special Interest Group. https://www.cidoc-crm.org/sites/default/files/cidoc_crm_version_7.2.4.pdf. Camp, Ann Van, Sven Lieber, and IFLA. 2022. “ISNI, a Top Tool for Quality Enhancement, Smooth Data Flows and Efficient Internal Processes.” Dublin: Ireland: International Federation of Library Associations; Institutions (IFLA). https://repository.ifla. org/handle/123456789/2008. Cruz, Maria, and Clifford Tatum. 2021. “NWO Persistent Identifier Strategy.” Zenodo. https://doi.org/10.5281/zenodo.4674513. Curry, Edward. 2020. “Dataspaces: Fundamentals, Principles, and Techniques.” In RealTime Linked Dataspaces: Enabling Data Ecosystems for Intelligent Systems, 45–62. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-296650_3. Data Documentation Initiative. 2020. “DDI Lifecycle (3.3) Documentation.” https://ddilifecycle-documentation.readthedocs.io/en/latest/index.html. Diefenbach, Dennis, Max De Wilde, and Samantha Alipio. 2021. “Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph.” In The Semantic Web – ISWC 2021, 12922:626–42. Lecture Notes in Computer Science. Cham: Springer. https://doi.org/10.1007/978-3-030-88361-4_37. Ernštreits, Valts. 2020. “Livonian Place Names: Documentation, Problems, and Opportunities.” Eesti Ja Soome-Ugri Keeleteaduse Ajakiri. Journal of Estonian and Finno-Ugric Linguistics 11 (1): 213–33. https://doi.org/10.12697/jeful.2020.11.1.09. European Commission. 2020. “A European Strategy for Data.” European Commission. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex:52020DC0066. ———. 2021a. “Brochure for Music Moves Europe Preparatory Action 2019.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/mme_2019_ brochure_final-web.pdf. ———. 2021b. “Music Moves Europe - First Dialogue Meeting. Final Report.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/mme-conferencereport-web.pdf. European Commission, Directorate-General for Education, Youth, Sport and Culture, M 120 Clarke, P Vroonhof, J Snijders, A Le Gall, B Jacquemet, et al. 2020. Feasibility Study for the Establishment of a European Music Observatory : Final Report. Publications Office of the European Union. https://doi.org/10.2766/9691. Europeana. 2017. “Definition of the Europeana Data Model V5.2.8.” Europeana. https://pro.europeana.eu/files/Europeana_Professional/Share_your_data/Technical_ requirements/EDM_Documentation//EDM_Definition_v5.2.8_102017.pdf. Faraj, Ghazal, and András Micsik. 2023. “Enriching Wikidata with Cultural Heritage Data from the COURAGE Project.” In, 407–18. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-36599-8_37. Fragkou, Pavlina. 2023. “DCAT-AP 3.0.” Edited by Makx Dekkers, Pavlina Fragkou, Natasa Sofou, and Bert Van Nuffelen. https://semiceu.github.io/DCAT-AP/releases/3. 0.0/. Hartmann, Thomas, Sarven Capadisli, Franck Cotton, Richard Cyganiak, Arofan Gregory, Benedikt Kämpgen, Olof Olsson, Heiko Paulheim, Joachim Wackerow, and Benjamin Zapilko. 2024. “DDI-RDF Discovery Vocabulary. A Vocabulary for Publishing Metadata about Data Sets (Research and Survey Data) into the Web of Linked Data.” Edited by Thomas Hartmann, Richard Cyganiak, Joachim Wackerow, and Benjamin Zapilko. W3C. https://rdf-vocabulary.ddialliance.org/discovery.html. International Council on Archives Expert Group on Archival Description. 2023. “Records in Contexts–Conceptual Model. Version 1.0.” International Council on Archives. https: //www.ica.org/app/uploads/2023/12/RiC-CM-1.0.pdf. International ISRC Registration Authority. 2021. “International Standard Recording Code (ISRC) Handbook. 4th Edition.” International ISRC Registration Authority. https: //www.ifpi.org/wp-content/uploads/2021/02/ISRC_Handbook.pdf. ISO. 2012. “International Standard Musical Work Code (ISNI). ISO 27729:2012.” International Organization for Standardization. https://www.iso.org/standard/44292.html. ———. 2013. “ISO 17369:2013(en) Statistical Data and Metadata Exchange (SDMX).” London:United Kingdom: International Organization for Standardization. https://www. iso.org/obp/ui/en/#iso:std:iso:17369:ed-1:v1:en. ———. 2017a. “ISO/IEC 19941:2017(en), Information Technology — Cloud Computing — Interoperability and Portability.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso-iec:19941:ed-1:v1:en. ———. 2017b. “ISO/IEC 5127:2017(en), Information and Documentation — Foundation and Vocabulary.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso:5127:ed-2:v1:en. ———. 2017c. “ISO 2108:2017 (En), Information and Documentation — International Standard Book Number (ISBN).” International Organization for Standardization. https: //www.iso.org/standard/65483.html. ———. 2019a. “ISO/IEC 20546:2019 Information Technology — Big Data — Overview and Vocabulary.” London:United Kingdom: International Standards Organisation. https: //www.iso.org/obp/ui/en/#iso:std:iso-iec:20546:ed-1:v1:en. ———. 2019b. “International Standard Recording Code (ISRC). ISO 3901:2019.” International Organization for Standardization. https://www.iso.org/standard/64817.html. ———. 2020. “ISO/IEC 22624:2020(en), Information Technology — Cloud Computing — Taxonomy Based Data Handling for Cloud Services.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso121 Figure 5: Dilinger is one of the best editors, and it is particularly suitable foor first-time markup users, as you immediately get visual feedback on how you mark up your text. Understanding the Data Model Wikidata is a collaboratively edited multilingual knowledge graph hosted by the Wikimedia Foundation. It is a common source of open data that Wikimedia projects, such as Wikipedia, and anyone else, is able to use under the CC0 public domain license. As of early 2023, Wikidata had 1.54 billion item statements, or small, verifiable, scientific statements about our world. It runs on Wikibase, the tool that we use for the data consolidation of the Open Music Observatory. Wikidata is a document-oriented database, focusing on items, which represent any kind of topic, concept, or object. 128 Figure 6: Wikidata is a document-oriented database. This document connects a lot of knowledge about the late English writer and humorist, Douglas Adams. Data curators are expected to understand the basics of the Wikibase Data Model and the idea of working with a document-oriented database. We could learn from many EU and member state projects in this regard because Wikibase is a tool that is very often used for similar tasks. Originally intended at the level of citizen scientists, it allows music domain experts, like musicologists, music economists, music librarians and other non-technical stewards of data, to work efficiently with a data coordination system that uses Wikibase. 129 Working with the GUI Figure 7: Identical to Wikidata: you must fill out at least the main Label of the item, and a description. We use English (en) as the master language for international cooperations. Sandbox environment Our manual is accompanied with a “sandbox” learning environment, where new data curators can try out various data manipuliation, editing, deleting, uploading actions without 130 endangering their data, or the shared database. Mass importing Mass importing data requires solid technical skills because, in almost all cases, the data arrives from a very differently structured database: spreadsheets or a relational database management system. These tasks are performed by Reprexbase , the software components developed by Reprex to connect music industry sources to the Wikibase Data Model. The guidelines provide information on how to prepare the data or what information should be given to us about the original schemas to start the mass import. Data enrichment The data enrichment are carried out with software components created by Reprex. It took about ten months to get enough data clearances in our data-sharing space to accumulate enough data for training enrichment algorithms. Results will be reported later in other tasks. Quality Testing with SPARQL SPARQL is the standard query language and protocol for Linked Open Data and RDF databases. Having been designed to query a great variety of data, it can efficiently extract information hidden in non-uniform data and store it in various formats and sources. SPARQL, pronounced ‘sparkle’, is the standard query language and protocol for Linked Open Data on the web or RDF triplestores. The SPARQL standard is designed and endorsed by the World Wide Web Consortium and helps users and developers focus on what they would like to know instead of how a database is organised. Our data curators must be able to run SPARQL queries and make elementary modifications to them. Because we often import very large datasets, it would be very difficult to manually control every record on the graphical user interface. We use pre-written SPARQL queries (the data curator is expected to run via a simple URL link, perhaps modifying a class’s QID or a property’s PID) that serve as so-called unit tests. These queries programmed by Reprex allow simple tests like these: ⊠If the curator gave us 5432 person records, we have 5432 persons in the Reprexbase instance; ⊠If the gender breakup of a person’s records is 2834:2598, the instance results in exactly the same persons of two genders (assuming that no third gender is used in the original data.) ⊠If we received data on Ján Levoslav Bella’s Symphony in B minor, the publication year is 1982. 131 # Composers: citizens of Slovakia SELECT ?item ?itemLabel ?givenNameLabel ?lastnameLabel ?birthdate ?deathdate ?nationalityLabel ?itemDescription WHERE { ?item wdt:P31 wd:Q5 . # instance of human ?item wdt:P106/wdt:P279*wd:Q36834. # occupation or subclass of occupation that is composer ?item wdt:P27 wd:Q214. # country of citizenship is Slovakia optional { ?item wdt:P735 ?lastname . } optional { ?item wdt:P734 ?givenName . } optional { ?item wdt:P569 ?birthdate . } optional { ?item wdt:P570 ?deathdate . } optional { ?item wdt:P27 ?nationality . } SERVICE wikibase:label { bd:serviceParam wikibase:language "en,sk,de,hu" } } order by ?itemLabel Try it out� 132 Terminology Mapping guidelines Music professionals Wikidata uses the human (Q5) class a subclass of person (Q215627) that was defined much later. This often makes the mapping to CIDOC-CRM, RiC and many other ontologies ambiguous, because many persons are lacking a statement about their personhood. In collections management E21 Person (collections) and RiC-E08 (archives) are the most likely anchors of persons who are composers or performers of music. For our use cases, the differentiation between living and deceased (not to mention imaginary) persons is necessary, so we encourage data curators to use the following mapping when importing their datasets: •deceased person: use this preferred label to dead human (Q18093576). Our description: human who is no longer living (equivalent with dead person on Wikidata.) •living person: use this preferred label to living human (Q18093573). Description: a human person who is alive, equivalent with living human on Wikidata. We will add for both deceased person and living person the human statement for compatibility with Wikidata. •imaginary person: use this preferred label to the imaginary character (Q115537581). Our description: character known only from narrations (fictional or in a factual manner) without a proof of existence; includes fictional, mythical, legendary or religious characters and similar; equivalent to the the Wikidata item imaginary character. Imaginary persons are not entitled to copyright. •deceased creator whose works are no longer protected by copyright: we create an inherited class of deceased person and creator. Creators who are living person or who do not belong to this class are assumed to have their works under copyright protection. Of course, in the case of multi-creator works, we cannot infere a public domain status from this class. 133 Musical works Sound recordings Live public performance 134 SKCMDb: Slovak Comprehensive Music Database The Slovak Comprehensive Music Database (SKCMDb) is a national initiative aimed at making Slovak music more accessible, discoverable, and usable across libraries, archives, streaming services, and rights management organisations. It connects scores, recordings, and metadata using open standards and collaborative governance. As a functional module of the Open Music Observatory, the SKCMDb also serves as a testbed for developing shared data services, addressing the conceptual models, workflows, and governance rules required to link diverse music stakeholders. ‘ The SKCMDb is supported by a data-sharing space consisting of both shared and private databases. The data-sharing space currently comprises the following initial components: •Slovak Metadata Database: A database that facilitates connections between various Slovak stakeholders’ systems. •SKCMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 135 •SKCMDb HC-SOZA (private): A database used exclusively for rights management and library management, governed by an agreement between the Slovak Music Centre and SOZA. •SKCMDb HF (private): A technical dataset created to provide improvements, corrections, and enrichments for the Hudobný fond. The Slovak Metadata Database serves as a support layer that is partly public and partly private. Its metadata definitions and descriptive metadata are exported into the SKCMDb databases as needed and permitted. The Slovak Metadata Database The Slovak Metadata Database is developed in alignment with the metadata framework of the Open Music Observatory. •Ontological and thesauri patterns: Reuses standardized or widely adopted vocabularies. •Conceptualizations and definitions: Includes concept definitions, thesauri, and other elements developed specifically for the SKCMDb. •Public permanent identifiers: Uses identifiers that are public or can be made public. The metadata layer is generally licensed under CC0, though in some cases other licenses are used (for example, CC-BY). ĹNote Examples: • The definitions of musical work,printed sheet music, and the is score of relationship allow the description of connections between an abstract musical work—such as Bella’s Missa in C—and its actual printed manifestations. • The VIAF identifier 2737220 identifies Ján Levoslav Bella’s compositions across library systems. Slovak Comprehensive Music Database (public) The Slovak Comprehensive Music Database is a linked open database published by the Slovak Music Centre. It integrates elements from the Hudobné centrum’s own databases along with data made public by SOZA, Hudobný fond, and other organizations. The database is distributed under various Creative Commons licenses that allow both commercial and non-profit use. 136 The primary aim of the SKCMDb is to reduce the cost of maintaining public and private services that enhance the circulation, availability, visibility, and legally licensed use of Slovak music. Our licensing policies are designed to enable the widest possible use of the data while protecting the investments required to maintain registers, standards, and data integrity. Slovak Comprehensive Music Database (private) The private components of the SKCMDb consist of databases where the data is not intended for public sharing but is used to enhance rights management, music information services, library operations, or other specialized applications. These databases are maintained under agreements between the participating parties. Access to these private datasets serves specific, well-defined purposes and is governed by strict rules. Availability to third parties is determined solely at the discretion of the data owners and may vary depending on contractual or legal obligations. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category (e.g., a national musical works register). Collections typically rely on registers to identify works unambiguously and avoid duplication. Both are documented in structured datasets containing standard identifiers such as ISRC, ISWC, or ISMN codes. In the case of statistical data, microdata often refers to survey instruments and responses curated under defined methodological rules. •Metadata: Relevant elements from the Slovak Metadata Database that support the use of collections or registers. The SKCMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models 137