scieee AI-readable full text Open interactive document viewer

Open Music Observatory

Antal, Daniel; Lázár, Ádám; Mester, Anna Márta

Abstract

Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries.

Full text

Open Music Observatory Building an open data sharing space for the European music sector Daniel Antal, CFA Barát, Andor Kornél Lázár, Ádám Mester, Anna Márta 2025-07-31 Table of contents Open Music Observatory 6 DisclaimerofWarranties................................. 6 Glossary 8 Musicterms........................................ 8 Creatorsofmusicalworks ............................. 9 Datascienceterms.................................... 10 Dataprotectionterms .................................. 14 Data curation and collection terms . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 Statisticalterms ..................................... 15 Registers, authorities, standards and identifiers . . . . . . . . . . . . . . . . . . . . 16 Organisations....................................... 19 Otherabbreviations ................................... 20 Executive Summary 22 1 Introduction 25 2 Background & Concept 27 2.1 History of the Music Observatory . . . . . . . . . . . . . . . . . . . . . . . . . 27 2.2 OpenMusicDataspace............................... 29 2.2.1 European Interoperability Framework . . . . . . . . . . . . . . . . . . 30 2.2.2 Embraceopen ............................... 31 2.2.3 Open Music Dataspace . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2.3 Prototyping in Slovakia . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2.4 Stakeholder presentations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.5 Otherobservatories................................. 34 2.5.1 Feasibilitystudy .............................. 34 2.5.2 Milk market observatory . . . . . . . . . . . . . . . . . . . . . . . . . . 35 2.5.3 European Audiovisual Observatory (EAO) . . . . . . . . . . . . . . . . 36 2.5.4 European Observatory on Infringements of Intellectual Property Rights(EUIPO) .............................. 36 2.5.5 European Market Observatory for Fisheries and Aquaculture Products(EUMOFA) .............................. 37 3 Core Services 39 3.1 Collect: Data Curation & Collection . . . . . . . . . . . . . . . . . . . . . . . 41 3.1.1 Microdata, Collections, Records . . . . . . . . . . . . . . . . . . . . . . 42 3.1.2 Primary data collection . . . . . . . . . . . . . . . . . . . . . . . . . . 44 2 3.1.3 Metadata .................................. 45 3.1.4 Statistical indicators and datasets . . . . . . . . . . . . . . . . . . . . 46 3.2 Process ....................................... 46 3.2.1 Processing & re-processing microdata . . . . . . . . . . . . . . . . . . 47 3.2.2 Documentation............................... 48 3.3 Disseminate..................................... 49 3.3.1 EU Open Data Portal . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 3.3.2 Europeana and the European Collaborative Cloud for Cultural Heritage 50 3.3.3 European Open Science Cloud . . . . . . . . . . . . . . . . . . . . . . 51 3.4 Metadata ...................................... 53 3.4.1 Wikibase & Wikidata . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 3.4.2 Music Observatory Website . . . . . . . . . . . . . . . . . . . . . . . . 55 3.4.3 APIEndpoint................................ 55 4 Data Catalogue 56 4.1 CollectionGuidelines................................ 57 4.2 TopicalPillars ................................... 59 4.2.1 MusicEconomy............................... 60 4.2.2 MusicDiversity............................... 62 4.2.3 MusicSociety................................ 63 4.2.4 Innovation.................................. 64 4.2.5 Sustainability................................ 64 5 Data Sources 65 5.1 Curation of reusable data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 5.1.1 Curation from data vendors . . . . . . . . . . . . . . . . . . . . . . . . 67 5.1.2 Noveldataassets.............................. 68 5.2 Dataproviders ................................... 68 6 Standardisation of Data & Terminology 70 6.1 Businessprocesses ................................. 70 6.2 Conceptual and information models . . . . . . . . . . . . . . . . . . . . . . . 71 6.3 Identification & Entity Linking . . . . . . . . . . . . . . . . . . . . . . . . . . 73 6.3.1 Registers & Authority Files . . . . . . . . . . . . . . . . . . . . . . . . 74 6.3.2 Open and persistent identifiers . . . . . . . . . . . . . . . . . . . . . . 75 6.3.3 Not open, music-industry specific identifiers . . . . . . . . . . . . . . . 75 6.3.4 Lyrics .................................... 77 6.3.5 ISCC .................................... 77 6.3.6 OMOIdentifiers .............................. 78 7 Data Improvement & Innovation 79 7.1 Value-Added Data Services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 7.1.1 DataSharing................................ 79 7.1.2 Fix-the-data ................................ 80 7.1.3 DataLinking ................................ 80 7.1.4 Registration services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 3 7.2 UseCases...................................... 82 7.2.1 Data Health Services for Collective Management . . . . . . . . . . . . 83 7.2.2 Sustainability Reporting for Music Organisations . . . . . . . . . . . . 84 7.2.3 ListenLocal................................. 85 7.2.4 Unlabel ................................... 86 7.3 UseofAIsystems ................................. 87 8 Conclusions & Next Steps: Towards a European Music Observatory 90 8.1 Co-creating a governance model for the Observatory . . . . . . . . . . . . . . 90 8.1.1 Observatory Stakeholder Network . . . . . . . . . . . . . . . . . . . . 92 8.1.2 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . 92 8.2 Bottom-up expansion of the Observatory in a federation model . . . . . . . . 93 8.3 Practical next steps in the project . . . . . . . . . . . . . . . . . . . . . . . . 94 References 97 Appendices 102 Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network 102 ...............................................102 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 Annex 2 - Data Curators Manuals 104 Inspiration ........................................105 Basic Data Organisation Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 Understanding the Data Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 WorkingwiththeGUI..................................108 Sandboxenvironment ..................................108 Massimporting......................................109 Dataenrichment .....................................109 Quality Testing with SPARQL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 Terminology 111 Mappingguidelines....................................111 Musicprofessionals.................................111 Musicalworks....................................112 Soundrecordings..................................112 Livepublicperformance ..............................112 SKCMDb: Slovak Comprehensive Music Database 113 The Slovak Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 Slovak Comprehensive Music Database (public) . . . . . . . . . . . . . . . . . . . . 114 Slovak Comprehensive Music Database (private) . . . . . . . . . . . . . . . . . . . 115 Microdata......................................115 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 116 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116 4 LīvMDb: Livonian Music Database 118 The Livonian Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 Livonian Music Database (public) . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 Livonian Music Database (private) . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 Microdata......................................120 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 121 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 5 Open Music Observatory Disclaimer of Warranties This project has received funding from the European Union’s Horizon Europe, research and innovation programme, under Grant Agreement No. 101095295. This document has been prepared by Open Music Europe (OpenMusE) project partners as an account of work carried out within the framework of this contract. Any dissemination of results must indicate that it reflects only the author’s view and that the Commission Agency is not responsible for any use that may be made of the information it contains. Neither Project Coordinator, nor any signatory party of Open Music Europe (OpenMusE) Project Consortium Agreement, nor any person acting on behalf of any of them: (a) makes any warranty or representation whatsoever, express or implied, (i). with respect to the use of any information, apparatus, method, process, or similar item disclosed in this document, including merchantability and fitness for a particular purpose, or (ii). that such use does not infringe on or interfere with privately owned rights, including any party’s intellectual property, or (iii). that this document is suitable to any particular user’s circumstance; or (b) assumes responsibility for any damages or other liability whatsoever (including any consequential damages, even if Project Coordinator or any representative of a signatory party of the Open Music Europe (OpenMusE) Project Consortium Agreement, has been advised of the possibility of such damages) resulting from your selection or use 6 of this document or any information, apparatus, method, process, or similar item disclosed in this document. For the version history of this document, please refer to our open repository, where the change history can be reviewed with timestamps for every single file used to create the report: https://github.com/dataobservatory-eu/open-music-observatory 7 Glossary Music terms audio recording: fixation of sounds (ISO 2019b) music video recording: fixation of sounds synchronized with pictures or moving pictures where (a) the fixed sounds are wholly or substantially a musical performance or (b) the recording is intended for viewing in association with a recording of a musical performance. This definition includes music videos and concert recordings, together with music-related interviews and documentaries, but does not extend to genera! audiovisual material, even if it includes music.(ISO 2019b) recording: result of a recording process independent of the type and number of audio or audiovisual carriers and technology used Note 1 to entry: The term “recording” applies to each recorded item which may be used as a separate unit regardless of whether it is issued as part of a larger recorded work (e.g. each separate track on an album of audio recordings). [SOURCE:ISO 3901:2001, definition 3.3] (ISO 2017b) work: distinct, abstract creation of the mind whose existence is revealed through one or more expressions (e.g. a performance) or manifestations (e.g. an object) (ISO 2022) musical work: composed of a combination of sounds, with or without accompanying text (ISO 2022) DSP or digital streaming platform: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. DSPs can provide access to music downloads, like Apple’s iTunes Store, or access to streaming music like Spotify, or even provide satellite-delivered content such as SiriusXM in the USA. rights management (organisations): the function of managing the rights on behalf of rights owners. It can be companies whose sole purpose is to ensure that content that has been licensed has delivered royalties that are identified and accounted for. The role can be taken by collective management organisations or by private companies on behalf of songwriters, composers, performers, music publishers, or record labels. duration: the elapsed playing time between the first and last recorded modulations of the recording. LP or Long Player: gramophone record usually on both sides comprising one or more sound recordings with a playing time of each side of normally round about 30 minutes and released and sold on its own (ISO 2017b) 8 single: gramophone record usually on both sides comprising one or two short sound recordings with a playing time of each side of normally no more than 7 minutes and released and sold on its own (ISO 2017b) track: single recording on a sound carrier (ISO 2017b) original version: The first established form of a work. (ISO 2022) medley: A continuous and sequential combination of existing works or excerpts. (ISO 2022) potpourri: A composite work with the addition of original material which have been combined to forma new work that has been published and printed. (ISO 2022) movement: A principal division of a musical work. (ISO 2022) original title: A title given to the work by its creator(s) shown in its original language. (ISO 2022) formal title: A standardized title in which the elements are arranged in a pre-determined order, such as titles created for classical works. (ISO 2022) expression: intellectual or artistic realisation of one and only one work Note: may take the form of a notation , sound, image, object, movement or text (ISO 2017b) manifestation: physical embodiment of an expression (ISO 2017b) Creators of musical works arranger: The author, or one of the authors, of an adapted text of a musical work. (ISO 2022) author: The creator, or one of the creators, of the text of a musical work. (ISO 2022) Entitled to authors’ right or copyright. composer: The creator, or one of the creators, of the musical elements of a musical work. (ISO 2022) Entitled to authors’ right or copyright. lyricist: The author of the text of a musical work, or a literary text that is arranged together with a musical work. Entitled to authors’ right or copyright. performer: The performer of a musical work; in case of a sound recording, the performer whose performance is fixed in the recording. They may be entitled to neighbouring or sound recording copyrights. producer: The person or legal entity that produces the recorded fixation of the sound recording. They are entitled to neighbouring or sound recording copyrights. 9 assets, emphasising machine-actionability (i.e., the capacity of computational systems to find, access, interoperate, and reuse data with none or minimal human intervention.) indicator: the representation of statistical data for a specified time, place or any other relevant characteristic, corrected for at least one dimension (usually size) so as to allow for meaningful comparison. microdata: non‐aggregated observations or measurements of characteristics of individual units, without direct identifier. MVP or minimum viable product: a version of a work product with just enough features and requirements to satisfy early customers and/or provide feedback for future development [SOURCE:IEEE 2675-2021, 3.1] observation unit: an identifiable entity about which data can be obtained, it is also often called a statistical unit or data subject in case of a natural person. Open Policy Analysis Guidelines: a set of information management rules to make policy analysis more transparent. personal data: any information relating to an identified or identifiable natural person. pseudonymisation: processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information. survey: a systematic examination and record of a physical or social area and its features so as to construct a map, plan, or description. In social sciences it usually refers to a well-structured questionnaire and answers given to its items by a target population. statistics: quantitative and qualitative, aggregated and representative information characterising a collective phenomenon in a considered population. visualisations: schematic charts, drawings, photographs, and their collages will as still image files that help to explain the relationship between information carriers, data points, or processes. Registers, authorities, standards and identifiers IČO: The organisation identification number (IČO) is an identifier assigned to all types of legal entities, entrepreneurs and public authorities by the Statistical Office of the Slovak Republic. The Czech Republic’s organisation identifier is also called IČO. OpenCorporates: a public corporation database which sources data from national business registries. ISNI: an ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. VIAF: The Virtual International Authority File (VIAF) is an international service that consolidates multiple name authority files into a single database. Their primary goal is 16 to enhance the efficiency and usability of library authority files by linking and merging widely used authority records and making them accessible online. VIAF ID: The VIAF (Virtual International Authority File) combines multiple name authority files into a single OCLC-hosted name authority service. ISRC: The International Standard Recording Code (ISRC) is a standard identifying code that can be used to identify sound recordings and music video recordings so that each such recording can be referred to uniquely and unambiguously. ISWC: The purpose in creating an ISWC for musical works is to enable more efficient administration of rights to those works on a worldwide basis. The ISWC provides an efficient means of identifying musical works in computer databases and related documentation and for the exchange of information between rights societies, publishers, record companies and other interested parties on an international level. ISBN: the International Standard Book Number is an identification system for the publishing industry and its supply chains. ISMN: The International standard music number (ISMN) was developed by, and for, the music publishing sector as a separate system to complement the International standard book number (ISBN). The existence of the ISMN as a separate identifier system makes it possible to identify printed and notated music as a distinct category of publication within the global supply chain and to develop trade directories and similar services for the specialized market for music publications. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. DOI: The Digital Object Identifier is a standardised unique number given to many (but not all) articles, papers and books, by some publishers, to identify a particular publication. ORCID: the Open Researcher and Contributor ID is a unique, persistent identifier free of charge to researchers. URI: A Uniform Resource Identifier (URI) is a string of characters used to identify a resource on the internet. This resource can be either abstract or physical, such as a website, an email address, or a file. URIs are essential for enabling interactions with resources over a network using specific protocols. W3C: The World Wide Web Consortium (W3C) is an international community that develops standards for the World Wide Web. Their mission is to lead the Web to its full potential by creating technical specifications and guidelines that are designed to be open and royalty-free. These standards include HTML, CSS, and other web technologies, which ensure that web content is accessible across different browsers and devices. DDI: The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. Wikibase: Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, 17 but it is available now for the creation of private, or public-private partnership knowledge graphs. It is developed by Wikimedia Deutschland. GSBPM: The Generic Statistical Business Process Model is a international standard model that “describes and defines the set of business processes needed to produce official statistics.” | GSIM:Generic Statistical Information Model: a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation standards such as SDMX or DDI. SDMX: Statistical Data and Metadata eXchange (SDMX), is an international initiative that aims at standardising and modernising (“industrialising”) the mechanisms and processes for the exchange of statistical data and metadata among international organisations and their member countries. ESRS: The European Sustainability Reporting Standards (ESRS) are a set of guidelines developed by the European Financial Reporting Advisory Group (EFRAG) to standardise sustainability reporting across the European Union. These standards are designed to align with the Corporate Sustainability Reporting Directive (CSRD), which mandates detailed corporate reporting on environmental, social, and governance (ESG) issues for many companies operating within the EU. CIDOC-CRM: The conceptual model of CIDOC, the standard conceptualisation of collection management systems in heritage organisations. RiC:Records in Context is a new conceptual model that replaces the four most important international archiving standards. DCTERMS or DCMI: the Dublin Core Metadata Terms is a vocabulary of metadata terms developed and maintained by the Dublin Core Metadata Initiative (DCMI). These terms are used to describe various aspects of digital resources, such as web pages, documents, and other online content. They provide a standardized way to assign metadata to resources, making them easier to discover, manage, and exchange. RDFS: the Resource Description Framework Schema is an extension of the Resource Description Framework (RDF) that provides a vocabulary for describing classes and properties of resources within an RDF graph. EDM: the Europeana Data Model is a framework for collecting, connecting, and enriching cultural heritage metadata. It’s designed to facilitate the sharing and reuse of cultural heritage information by providing a standardized way to represent and link data. Europeana: a digital platform provided by the European Union that aggregates digitized cultural heritage from institutions across Europe. ESCO: the European Skills, Competences, Qualifications and Occupations classification is is a multilingual classification system developed by the European Commission to standardize the description of skills, competences, and qualifications relevant to the European labor market and education. 18 NACE: the European Union’s standard classification of economic activities for statistical purposes. The abbreviation stands for Nomenclature statistique des Activités économiques dans la Communauté européenne. ISCO: the International Standard Classification of Occupations (ISCO) is the International Labour Organization’s standardized system for classifying and organizing occupations according to jobs’ tasks and duties ISIC: the International Standard Industrial Classification of All Economic Activities (ISIC) is a standard classification system developed by the UN Statistics Division (UNSD) to categorize economic activities. PROV-O: the Provenance ontology is a formal ontology developed by W3C to represent and interchange provenance information. MARC: MAchine-Readable Cataloging, is a standard digital format used by libraries to represent and exchange bibliographic information. DCAT: an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. Organisations AEPO-ARTIS: Organisation representing European artists-performers. Regroups most of the European CMO representing performers. ALOADED: is a company which distributes and exploits recordings. CISAC: The International Confederation of Societies of Authors and Composers is an international non-governmental, not-for-profit organisation that aims to protect the rights and promote the interests of creators worldwide. CNM (former CNV): the Centre National de la Musique is a public organisation managing a tax on concert tickets EFRAG: The European Financial Reporting Advisory Group is a private association established in 2001 with the encouragement of the European Commission to serve the public interest. EFRAG extended its mission in 2022 following the new role assigned to EFRAG in the CSRD, providing Technical Advice to the European Commission in the form of fully prepared draft EU Sustainability Reporting Standards and/or draft amendments to these Standards. EMO: The European Music Observatory (EMO) is envisioned as a hub for collecting and analysing data on the music sector across Europe. Its primary aim is to address the current gaps and inconsistencies in music data collection, which have been a significant challenge for the sector. GESAC: GESAC comprises together 32 European authors’ societies in music, audiovisual, visual arts, literature and drama. 19 GESIS: Leibniz Institute for the Social Sciences. IAML: International Association of Music Libraries, Archives and Documentation Centres | IAMIC: International Association of Music Centres, an international network of organisations that collectively and collaboratively provides information and promotes the music of their countries or regions. ICMP: the global trade body representing the music publishing industry worldwide. SCAPR: International association for the development of the practical cooperation between performers’ collective management organisations (CMOs) SOZA: SOZA (Slovenský ochranný zväz autorský pre práva k hudobným dielam, Slovak Performing and Mechanical Rights Society) is a legal entity, non-profit civic association of authors and publishers of musical works, association of natural persons and legal entities. Hudobné Centrum: Music Centre Slovakia is a music organisation with a mission to promote Slovak contemporaly music. Other abbreviations CEEMID: the Central European Music Industry Databases is a multi-country project that was a predecessor of Reprex’s Digital Music Observatory CSRD: The Corporate Sustainability Reporting Directive (CSRD) is European Union (EU) legislation, effective from 5 January 2023, that requires EU businesses—including qualifying EU subsidiaries of non-EU companies—to disclose their environmental and social impacts, and how their environmental, social and governance (ESG) actions affect their business. DSP: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. EIF: The European Interoperability Framework (EIF) is a set of recommendations and guidelines that aims to facilitate communication and collaboration between public administrations, businesses, and citizens within the European Union and across national borders. ECCCH: The European Collaborative Cloud for Cultural Heritage is a European Union initiative for a digital infrastructure that will connect cultural heritage institutions and professionals across the EU. EOSC: The European Open Science Cloud (EOSC) aims to create a trusted, open, and multidisciplinary environment for researchers and innovators in Europe. PPP: A Public-Private Partnership (PPP) is a collaborative arrangement between government entities and private sector companies aimed at financing, designing, implementing, and operating projects or services traditionally provided by the public sector. RDM: Research Data Management refers to the suite of practices, policies, and processes used to handle data throughout the lifecycle of a research project. 20 Our glossary is harmonised with relevant music-sector specific standards (referred to in Chapter 6) and with the ISO Information technology — Vocabulary (ISO 2023b); Information technology — Cloud computing — Taxonomy based data handling for cloud services (ISO 2020); Information technology — Cloud computing — Interoperability and portability (ISO 2017a) and the Information and documentation — Foundation and vocabulary (ISO 2017b) and Information technology — Metadata registries (MDR) — 1. Framework (ISO 2023a) 21 Executive Summary Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries. An observatory has traditionally been a permanent location for observing terrestrial, marine, or celestial events. In the past 30 years, it has also been used for long-term digital data collection programs for markets, social sciences, and humanities. Our milestone requires the start of this observatory after a lengthy and intensive planning and prototyping phase. It can be seen as a modern reimagination of the data observatory model, or the observatory 2.0. We created a new observatory model that fully aligns with the European Interoperability Framework but extends the governance of the digital services beyond public bodies, and allows the creation of a public-private partnership to manage the observatory. We were informed and influenced by the creation of Europeana (which started out from a similar collaborative project) and their new plans to extend their digital services into the European Collaborative Cloud for Cultural Heritage (ECCCH). We aim fully interoperability with Europeana and ECCCH, but we also bring a new element into their thinking. While they are mainly aggregating the work of public sector memory institutions, we are building a governance model that allows a more successful cooperation among the private sector and the public music sector. By the end of 2025, we aim to create an “observatory 3.0”, which already hosts many intelligent data improvement technologies and fuels innovative applications/services in line 22 with our project’s exploitation pathways. These services are at different maturity levels, but they could not be brought to a testable MVP without building out the minimal digital infrastructure and governance model at this milestone. Figure 1: we need a new version of this 23 ĹNote This document is licensed under the CC BY 4.0 LEGAL CODE Attribution 4.0 International license. You must refer to the document with the DOI 10.5281/zenodo.11385044. Canonical Licence URL:https://creativecommons.org/licenses/by/4.0/ Other formats:Plain Text;RDF/XMLlSee the deed 24 1 Introduction Task 5.1 Adding Economy, Diversity, Innovation and Society, Citizenship Pillars to the MVP of the Open Music Observatory and raising it to at least TLR Level 7 is a cornerstone of the OpenMusE project. It addresses the ultimate ambition of the initiative: creating and validating the technological, service, and governance components of an observatory that can serve as the foundation for a future European Music Observatory. Deliverable 5.1 Open Music Observatory marks the most significant milestone of this task, representing the public release of the Open Music Observatory on the World Wide Web. This report accompanies the deliverable. Its purpose is to explain the workings of the Observatory, outline the concepts behind its development, and describe our plans to evolve it into a platform that could underpin a future European Music Observatory. Following the Executive summary and this introduction (Chapter 1), the report is structured as follows: Chapter 2provides an overview of the historical plans for establishing a European music observatory and the rationale for a decentralized approach. It introduces CEEMID—an important precursor to the Open Music Europe project and a key building block identified in the EMO feasibility study. We then conceptualize the Open Music Observatory as a data-sharing space, a technical and legal innovation of the European Union. This section emphasizes compliance with the European Interoperability Framework and the adoption of open collaboration methods, open-source software, and open data as fundamental components. Finally, we review a selection of existing observatories, including those assessed in the EMO feasibility study, to compare their current services with our own. Chapter 3offers an overview of the current core services of the Observatory, which function as those of a modern statistical vendor and open data portal. In line with the EIF, we integrate our services into the EU Open Data Portal (as the primary gateway for disseminating statistical datasets), Europeana and the European Cultural Heritage Cloud (for music collections), and the Wikipedia ecosystem (Wikidata and Wikibase) to distribute data into open knowledge graphs. Chapter 4introduces our digital curation guidelines and the nucleus of our data offering. Before the release of Observatory 2.0, we could not ingest large, high-value datasets, so the current catalogue contains datasets derived solely from the background IP of the Open Music Europe project. These datasets are already available for testing, and we plan to present the first large-scale datasets and onboarding case studies at several international conferences in November 2025. One such large-scale case study is already underway and will be introduced later in Section 3.1.1. 25 2.2.3 Open Music Dataspace ĺImportant Open Music Dataspace: as the data management of the Open Music Observatory, it recognises that in a large-scale integration of music sector data involving potentially thousands of data sources, it is impossible to create a unifying schema across all sources. It is an exchange, processing, sharing and provision of data between trusted partners, for a fee or not. It is not necessarily copying or repatriating data centrally but ensures that each data holder has complete control over the conditions (e.g., who, when, and under what conditions) of access to their data. We create this data-sharing space view that certain national, genre-specific or other interest groups may want to apply a deeper level of integration and exchange and, through federation, extend the part of their data sharing to the entire Open Music Observatory. The application of automatic matching and mapping generation techniques, widely used conceptual models, and the concepts of the European Interoperability Framework allow the Observatory Stakeholder Network to integrate data on an ‘as-needed’ or as-permitted basis, keeping the data ready for integration whenever the data owners agree on such a need. 2.3 Prototyping in Slovakia Because a significant part of the Open Music Europe background was developed or tested in Slovakia, we decided to start prototyping the Open Music Observatory in this country. ĹNote Slovak Music Dataspace: a dataspace organised by SOZA, Reprex, and Hudobné centrum to share and exchange data on music related to the Slovak Republic. It provides a secure data exchange supported by trustworthy AI to harmonise and exchange data among representative Slovak music organisations and to create the public Slovak Music Register and the Slovak Comprehensive Music Database. In a significant stride towards our shared vision, we signed a Memorandum of Understanding in March 2023 (Ministerstvo kultúry SR and Open Music Europe 2023). This strategic alliance includes key stakeholders, such as the Ministry of Culture, and is aimed at establishing a robust public-private partnership. This partnership will play a pivotal role in sharing, exchanging, and improving data related to music in the Slovak Republic, thereby fostering a vibrant music ecosystem. Slovakia has a relatively advanced statistical system: one of the few EU member states with a satellite accounting system for the cultural and creative industry to augment the country’s national accounts. The methodological work related to coordinating governmental 32 and privately held data is the subject of other tasks; in making the Open Music Observatory, we only provide the infrastructure and the dissemination of knowledge for replication. Because we apply GSIM, DDI and SDMX, like the Slovak statistical system, our solutions are replicable in an EU/EEA member state or candidate country that has sufficiently aligned its statistical practices with the European Statistical System. 2.4 Stakeholder presentations We presented our paper at the Networkshop 2024 conference in Eger, Hungary in April 2024, because our first module in Slovakia overlaps most with Hungary’s music heritage due to the two countries many centuries of shared history [^background-3]. [^background-3]: Possiblities of a Hungarian data federation with the Slovak music data space (in Hungarian): (Daniel Antal 2024a). We presented our work and ideas on the International Association of Music Centres (IAMIC) general assembly and conference in November 2024 to find support, potential users and feedback.3. We presented our paper and poster on the International Association of Music Libraries, Archives and Documentation Centres (IAML) 4, again, with the aim to find support, potential users and feedback.5 3Poster: (Daniel Antal 2024b). 4SKCMDB: Interoperability of Music Libraries and Archives with Public and Private Music Services. Poster: (Daniel Antal 2025b), presentation: (Daniel Antal 2025a). 5Poster: (Daniel Antal 2024b). 33 2.5 Other observatories The EMO feasibility study presented several observatories to compare their services and organisational models. As more than five years have passed since the data collection for that study, we highlight here some of the considered comparators and add a few more. Figure 2.2: Various UN and OECD bodies, and particularly the European Union support or maintain more than 60 data observatories, or permanent data collection and dissemination points. 2.5.1 Feasibility study The Feasibility study directly mentions some observatories as possible good practice to be used wile working on the European Music Observatory. Chapter 1.5. is about the European Audiovisual Observatory (EAO) which has a great impact “on creating a consensual mapping environment for the audiovisual sector” (p. 25) and would be a good example to model the OMO after. The European Market Observatory for Fisheries and Aquaculture Products (EUMOFA) is known to operate under a service contract, which is awarded through a tender - a potential way the “to integrate the European Music Observatory setup directly within the competent services of the Commission” (p. 47) The basis of the Directorate-General for Agriculture and Rural Development (DG-AGRI) market observatories can be used to establish the OMO, but the downside is the lack of legal obligation for data transparency in the music sector, which is present in agriculture, thus making it difficult to replicate the legal basis. (p. 48) 34 The European Observatory on Infringements of Intellectual Property Rights (EUIPO) along with EAO is mentioned as possible partners to work with OMO and the tasks of OMO could be potentially integrated into within these partners. (p. 48) The details of this integration is discussed in chapter 3.8.6 (p. 89-91). 2.5.2 Milk market observatory The aim of the EU milk market observatory is to provide the EU dairy sector with more transparency by means of disseminating market data and short-term analysis in a timely manner. Figure 2.3: The European milk market observatory celebrates its 10th anniversary [link�](https://agriculture.ec.europa.eu/news/european-milk-marketobservatory-10th-anniversary-2024-04-16_en) • Publications: regular, in PDF document. • Statistical data access: Agri-food data portal • API: Agri-food data API • EU Open Data portal: yes, via Directorate-General for Agriculture and Rural Development • Newsletter: • License: The Commission’s reuse policy is implemented by the Commission Decision of 12 December 2011 on the reuse of Commission documents. Unless otherwise indicated 35 (e.g. in individual copyright notices), content owned by the EU on this website is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) licence. This means that reuse is allowed, provided appropriate credit is given and changes are indicated. 2.5.3 European Audiovisual Observatory (EAO) Figure 2.4: European Audiovisual Observatory • Link: https://www.obs.coe.int/en/web/observatoire/ • Publications: regular, in PDF and xlsx document • Statistical data access: • API: No. • EU Open Data portal: ? 2.5.4 European Observatory on Infringements of Intellectual Property Rights (EUIPO) “The European Observatory on Infringements of Intellectual Property Rights is a network of experts and specialist stakeholders that brings together representatives from EU bodies, authorities in EU countries, businesses and civil society. The aim of the observatory is to improve the fight against counterfeiting and piracy by sharing information and best practice, raising public awareness, strengthening cooperation, and developing better tools.” [https://single-marketeconomy.ec.europa.eu/industry/strategy/intellectual-property/enforcementintellectual-property-rights/european-observatory-infringements-intellectualproperty-rights_en] 36 The legal mandate of managing the observatory is in the Regulation No 386/2012. Figure 2.5: European Observatory on Infringements of Intellectual Property Rights • Link: https://www.euipo.europa.eu/en/about-us/observatory • Publications: regular, in PDF document • Statistical data access: Several, under “Services” Menu • API:The observatory does not appear to offer datasets. • EU Open Data portal: via Directorate-General for Communications Networks, Content and Technology some publication and datasets of the EUIPO are available. • License: The observatory does not appear to offer datasets. 2.5.5 European Market Observatory for Fisheries and Aquaculture Products (EUMOFA) The European Market Observatory for fisheries and aquaculture products (EUMOFA) is a market intelligence tool on the European Union fisheries and aquaculture sector, developed by the European Commission. It aims to increase market transparency and efficiency, analyses EU markets dynamics, and supports business decisions and policy-making. EUMOFA enables direct monitoring of volumes, values and prices of fisheries and aquaculture products, from the first sale to retail stage, including imports and exports. Data are collected from EU countries, the Faroe Islands, Iceland, Norway, the United Kingdom and from EU institutions, and are updated every day. 37 Figure 2.6: European Market Observatory for Fisheries and Aquaculture Products (EUMOFA) • Link: https://eumofa.eu • Publications: regular, in PDF documents. • Statistical data access: EUMOFA DATA • API: via the Directorate-General for Maritime Affairs and Fisheries API or via EU Open Data portal: yes, via Directorate-General for Maritime Affairs and Fisheries • License: European Commission Reuse and Copyright Notice (Decision of 12 December 2011). 38 3 Core Services This section introduces the core services of the Open Music Observatory. These services are described in the EMO Feasibility study, but we gave them a modern and subjective interpretation that relies more on the novel innovations of data regulation and data science. The Slovak Comprehensive Music Database (SKCMDb) is a national initiative to make Slovak music more accessible, discoverable, and usable across libraries, archives, streaming services, and rights organizations. It connects scores, recordings, and metadata using open standards and collaborative governance. As a functional module of the Open Music Observatory, SKCMDb also serves as a testbed for developing shared services. Our services fully embrace the EIF and EOSC (related) interoperability frameworks and reproducible research techniques that allow for a more timely, less costly, and higher quality data ingestion and processing than manual workflows. Almost all workflows are supported by open-source components, some of which came as the background of the Open Music Europe project, were developed by other projects, or are being developed in different tasks of this project. The introduction of the separate software components is not a subject of this report. Our main services are related to the collection, processing and dissemination of data. 39 The Generic Statistical Business Process Model (GSBPM) is an international standard model that “describes and defines the set of business processes needed to produce official statistics.” (UNECE 2019) A conceptual reference or information model accompanies this business process model, the General Statistical Information Model (GSIM)(UNECE 2014). We use the conceptualisation of GSIM so that our results will be similar in quality to official statistics; of course, similar processes allow us to create products that combine well with official statistical products. The Open Music Observatory is not only collecting and disseminating statistically processed data, but also collections datasets, i.e., structured microdata. The GSBPM and GSIM cover both because statistical business processes rely on collection-like datasets, such as registers, codebooks, and metadata thesauri. Figure 3.1: The two implementation standards, DDI and SDMX ensures that the Open Music Observatory can work together with statistical offices, or respectable social sciences data repositories. In this report of Task 5.1, we concentrate on the inner cycle of the business processes. The outer cycle, specification, design, and build are carried out in other work packages, as is most of the analysis phase. We provide further services, too, which could relate in GSBPM to the 9th Evaluate header, which allows for quality improvements. The two implementation standards, DDI and SDMX will ensure that the Open Music Observatory can work together with statistical offices, or respectable social sciences data repositories, because both our business processes and the way we organise data is following their standards. The Data Documentation Initiatve (DDI), will ensure that we will remain compatible with official statistical microdata and metadata services and other social sciences archives, like GESIS, the official data archive of all European Commission-mandated survey research dating back over 50 years (Vardigan, Heus, and Thomas 2008). The application of SDMX ensures that our microdata and statistically processed data will be interoperable with official statistics of the UN, OECD, Eurostat, and national statistical services (Stahl and Staab 2018). Our data improvements, go beyond improvements of statistical quality and application of GSBPM; we aim to fix and improve music industry datasets for rights management or digital curation. The data enrichment and improvement are innovative solutions that are not part of the services of an open data portal or an observatory. We aim to offer these value added services to create new value and therefore motivation for music industry data owners to work with the observatory. •“Fix-the-data” means improving the data quality by finding or imputing missing values or finding and replacing erroneous data entries. In terms of metadata, adding 40 further machine-actionable information to already existing datasets can improve their usability. •“Data linking”, data fusion, or data matching means correctly joining data from different datasets (data sources.) We ensure that data coming from sources can be meaningfully joined together; the variables have consistent meanings, the codebooks applied are harmonised; the timeframe or geographical frame is consistent. •Aggregation services: we turn your music-related datasets into statistical products or data publications. We clean, validate, and structure it to a format that it can be placed on the EU Open Data Portal, Europeana, or Wikibase for integration with Wikidata/Wikipedia. •Confidential data sharing: our data sharing space can be used for confidential data sharing and cross-pollination (for example, looking up missing ISWC/ISRC identifiers or misspelt names in each other’s datasets) without making the data public. These planned services will be discussed in Chapter 7. 3.1 Collect: Data Curation & Collection Data curation is the organisation and integration of data collected from various sources. It involves annotation, publication and presentation of the data so that the value of the data is maintained over time, and the data remains available for reuse and preservation. Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO feasibility study intuitively defines data gaps without an apparent reference to a data or conceptual model, but it recognises and stresses the need for terminological harmonisation. ĹNote Four types of data-collection principles have been identified as essential both by various branches of the music sector and also by policymakers at European, national and local levels: • The data-collection service provided by a European Music Observatory should help in mapping, understanding and analysing the main characteristics, trends and idiosyncrasies of the music sector in Europe; • The data collected should be neutral and available to decision-makers, music sector operators, and the public; • The data itself should cover the activities of the music sector across the entire European Union, be comparable between Member States, and rely on identified and stable indicators; • The data collection methods should be transparent and provide a strong degree 41 tors and their organisations involved. As a bare minimum, we provide machine-readable information about the technical publisher of the dataset, Reprex B.V: <https://isni.org/isni/000000050973936X> <a>"foaf:Agent" . And then we point the user the downloadable files (distributions) of the dataset with the rights statements and licenses. We use the Creative Commons CC BY 4.0 license, similar to Eurostat on the EU Open Data Portal, and we state that the dataset is open for the public. <https://zenodo.org/records/5652118/files/codebook_trb.csv> <a>"dcat:Distribution" ; <dcat:accessURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:byteSize>"41672" ; <dcat:downloadURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:mediaType>"text/csv" ; <dct:license><http://publications.europa.eu/resource/authority/licence/CC_BY_4_0>; <dct:rights><http://publications.europa.eu/resource/authority/access-right/PUBLIC>; <owl:sameAs><https://zenodo.org/records/5652118>. 3.2.2 Documentation ÁWarning We will provide the link and screenshot of the documentation for each file that goes public. 48 3.3 Disseminate 3.3.1 EU Open Data Portal The portal is a central point of access to European open data from international, European Union, national, regional, local and geodata portals. It consolidates the former EU Open Data Portal and the European Data Portal. The portal is intended to: 49 1. give access and foster the reuse of European open data among citizens, business and organisations. 2. promote and support the release of more and better-quality metadata and data by the EU’s institutions, businesses, agencies and other bodies, and European countries, enhancing the transparency of European administrations. 3. educate citizens and organisations about the opportunities that arise from the availability of open data. It is funded by the EU and managed operationally by the Publications Office of the European Union in cooperation with the Directorate-General for Communications Networks, Content and Technology of the European Commission, responsible for EU open data policy. We publish our data primarily on the EU open data portal for statistically processed datasets (datasets that contain the generalised characteristics of many data subjects without personal data that could identify them). 3.3.2 Europeana and the European Collaborative Cloud for Cultural Heritage Figure 3.2: Europeana is a bottom-up, decentralised data aggregator for the cultural heritage part of the music sector and the broader cultural and knowledge sector. 50 Europeana is at the heart of the common European data space for cultural heritage, a flagship initiative of the European Union to support the digital transformation of the cultural heritage sector. Millions of cultural heritage items from over 3,500 data providers across Europe are available online via the Europeana website. We work to share and promote this heritage so that it can be used and enjoyed by educators and researchers, creatives and culture lovers across the world. While there is no agreed, cross-sectoral definition of “collections”, it is widely understood that in many cases, collections themselves are the entities that meet the information needs of music professionals or researchers (Wickett et al. 2013). The creation of collections is an important activity performed by music professionals and scholars as part of their work process. For example, if we want to measure how many European or French works made it ever to an American hitlist, we have to contrast two collections (French works, and the collection of works that were ever on the particular hitlist) to calculate this indicator. The publishing policies of Europeana are restrictive, and therefore, there currently needs to be more musical works on this important open knowledge graph. After consultation with Europeana Sound, the British Library-based aggregator responsible for the music in the European collection, we decided to pursue two ways to make more extensive European music collections visible. We will publish collections with publicly viewable audiovisual material on Europeana via the Open Music Observatory’s collections. We will also start a discussion with Europeana’s new project, the European Collaborative Cloud for Cultural Heritage, which has less restrictive data licensing policies, about a more extensive dissemination point for our collection datasets. ::: callout-note More information about our services for curators of music collections and smaller independent repertoires can be found on our website�. 3.3.3 European Open Science Cloud The ambition of the European Open Science Cloud (EOSC) is to provide European researchers, innovators, companies and citizens with a federated and open multi-disciplinary 51 environment where they can publish, find and reuse data, tools and services for research, innovation and educational purposes. Naturally we want to ensure that users of the Open Music Observatory participate in these cloud services, either as research providers or as research users. The EOSC is recognised by the Council of the European Union among the 20 actions of the policy agenda 2022-2024 of the European Research Area (ERA) with the specific objective to deepen open science practices in Europe. It is also recognised as the “science, research and innovation data space” which will be fully articulated with the other sectoral data spaces. Given that the OMO is also following a data space architecture that is desigend to follow the European Interoperability Framework, the Open Music Observatory can be federated with, and can fully work with the EOSC. The Open Music Observatory is connecting to the EOSC via two key services of OpenAIRE. OpenAIRE itself is a Non-Profit Partnership of 50 organisations, established in 2018 as a legal entity, OpenAIRE A.M.K.E, to ensure a permanent open scholarly communication infrastructure to support European research, and it is a key implementer of the European Open Science Cloud. Or connection to OpenAIRE services guarantee our full compliance and use of EOSC. A key partner of OpenAIRE is CERN, which is manages the Zenodo open library and repository. We rely on the services of Zenodo for document identification, long-term archiving, and offering immediate access to our statistical datasets, visualisations, reports and other library-ready products. Similarly to the forming EU Open Research Repository, a Zenodocommunity dedicated to fostering open science and enhancing the visibility and accessibility of research outputs funded by the European Union, managed by CERN on behalf of the European Commission, we have created a similar Open Music Observatory Repository on the platform. Our repository is fully interoperable with the EU Open Research Repository (in pilot phase in June 2024) and with the entire EOSC. 52 Figure 3.3: For interoperability with EOSC and OpenAIRE and for long-term storage we store each distribution of the datasets among publications, documents on Zenodo, too, in the Open Music Observatory community. The semantic service of OpenAIRE, the OpenAIRE Graph is a collection of interlinked research objects that aggregates metadata records from more than 70K scholarly communication sources from all over the world for researchers, service providers, research managers and policy makers, by following a participatory approach. Figure 3.4: For interoperability with EOSC and OpenAIRE and for long-term storage we store each distribution of the datasets among publications, documents on Zenodo, too, in the Open Music Observatory community. 3.4 Metadata Europeana publishes the metadata in Turtle serialisation. The Open Music Observatory will provide access to the metadata in TTL and CSV distributions. 53 3.4.1 Wikibase & Wikidata Our main dissemination point for microdata are two Wikibase Suite installations, one for the Slovak Comprehensive Music Database, and one for other music. • For data related to Slovakia in the Slovak Comprehensive Music Database, we provide access at <https://reprexbase.eu/skcmdb/>. • For all other data we provide access at <https://reprexbase.eu/openmusic/>; this data has a less complex data management and governance. This access point provides access for individual datasets. We may use the openmusic.wikibase.cloud. Wikibase Cloud is an initiative of Wikipedia and Wikidata to bring more specialised data into the Wikimedia ecosystem. It may be a staging area for harmonisation with Wikidata. Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, but it is available now for the creation of private, or public-private partnership knowledge graphs. It was developed by Wikimedia Deutschland. Wikidata itself is a gigantic Wikibase instance. Their user interface is similar, but depending on what the administrator of your Wikibase instance allows you to do, you are likely to have more freedom to edit certain elements, like properties, than on Wikidata. Wikidata must protect the integrity of one of the world’s largest knowledge systems, and does not allow editing access to certain elements. Because of the success of Wikidata, several EU projects and institutions started to use Wikibase, the software that runs Wikidata. They aim to reuse the software to construct institutional or cross-institutional, domain-specific knowledge graphs. Several factors make Wikibase attractive: ⊠the fact that it is a well-maintained open-source software; ⊠there is a rich ecosystem of users and tools around it; ⊠Wikimedia Deutschland� (WMDE), the maintainer of Wikibase, has made considerable investments in optimising the software’s use outside of Wikidata or other Wikimedia projects; ⊠The EU Knowledge Graph� runs on Wikibase; ⊠The EU Academy and the EU Open Data Portal actively disseminate good practices and know-how on its implementation in cross-institutional data-sharing programs. Our main dissemination point for non-statistical data is the Wikibase Cloud. 54 3.4.2 Music Observatory Website ÁWarning We will completely revamp the website before submission and provide here a short overview with screenshots. 3.4.3 API Endpoint ÁWarning We will provide here a linked screenshot of our API endpoint before submission 55 4 Data Catalogue The Open Music Observatory curates, maintains, and disseminates a data catalogue with the resources within the data catalogue: individual datasets and their series and API endpoints where the data can be queried in a custom format. In creating our data infrastructure, we considered the specifications of our dissemination nodes, which provide our data with a wide range of interoperability and easy access: the EU Open Data Portal, Europeana, Wikibase Cloud and Wikidata. From a thematic point of view, we relied on the definition of the EMO feasibility study, and created topical pillars (Section 4.2). The data curators of the Observatory Stakeholder Network (see Section 8.1.1) and for the duration of the Open Music Europe project, the work packages (WP1-4 represent each “pillar”) can define and provide datasets or data series according to their topical collection guidelines (Section 4.1). A data catalogue formally is a metadata dataset: a dataset on information about our available datasets and their downloadable or queriable distributions. It follows the global World Wide Web DCAT standard. DCAT is an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. This document defines the schema and provides examples for its use (Albertoni et al. 2020). It is a global standard, which was further extended and specified for the release of statistical datasets (StatDCAT-AP) and for the needs of the EU Open Data Portal (DCAT-AP). (Sofou and Dragan 2019; Fragkou 2023) These extensions provide further metadata and organisations standards, but essentially they do not change the definition of the global standards. A data catalogue (dcat:Catalog) represents a catalogue, which is itself a dataset in which each individual item is a metadata record describing some resource: a description of a dataset, a data service, or other type of resource. dcat:Dataset represents a collection of data, published or curated by a single agent or identifiable community. We currently support two types of datasets: statistical datasets that conform to the datacube definition of SDMX, or collection datasets for microdata, which contain non-aggregated, structured data representing some unity criteria, for example, music works and recordings that have been present in the official radio charts of a given country. The _dataset_, similar to a musical or literary work, is an abstract concept which can be used, downloaded, and stored in its manifestation. For a musical work, a manifestation may be a sound recording or music sheet; for a dataset, it is a distribution. A URI identifies a dataset; the URI does not allow the downloading of the dataset, because it refers to the abstract idea of the dataset; the URL for downloading the dataset belongs to the individual distributions. 56 dcat:Distribution represents an accessible form of a dataset, such as a downloadable file. When the same dataset is distributed in different file formats (for example, CSV and SPSS files), each distribution is listed in the catalogue separately with a separate download link. Each distribution has its own URL where the dataset can be downloaded. Figure 4.1: The music.dataobservatory.eu/tag/music-economy/ URL lists the downloadable datasets on the Open Music Observatory website. They can be found on EU Open Data Portal, too. In the first days after launching our new service, around 1-10 June 2024, the datasets may be missing from the EU Open Data Portal, which is changing in these days its complete backend, and may have some backlog in accepting our datasets. dcat:DataService represents a collection of operations accessible through an interface (API) that provides access to one or more datasets or data processing functions. Our datasets are accessible on different platforms with their own datasets, and our internal data-sharing space also has its API. As data is added to the different platforms (EU Open Data Portal for statistical and microdata datasets, Europeana for collections dataset, Wikibase Cloud for further microdata, metadata and collections, and Reprexbase for confidential microdata and collections), we are updating the catalogue with the DataService entries. dcat:DatasetSeries is a dataset that represents a collection of datasets that are published separately but share some characteristics that group them; for example, a (play)list of sound recordings that were present in the weekly charts or the annual budget of an institution. A time series dataset is usually not defined as a data series, but the new time observations are added to an updated distribution of the time series dataset. Stakeholders who provide data to the Open Music Observatory can commit to making a data series; however, we only define a data series when we have at least two items available from the series. dcat:CatalogRecord represents a metadata record in the catalogue, primarily concerning the registration information, such as who added the record and when. 4.1 Collection Guidelines In short, we collect data about music. The initial data collection guidelines of the Open Music Observatory are derived from the EMO Feasibility study. We see them as a starting point for further discussion with the Observatory Stakeholder Network. ⊠Statistical data which is defined as cultural statistics of any European Economic Area and EU candidate statistical office or by a representative European or international music organisation. 57 4.2.4 Innovation The definition of the Innovation pillar in the EMO feasibility study is more a topic to be covered than a data need description. This pillar is less data-driven in that it will rely mostly on research conducted on topics relating to changes in the market place, new business models, disruptive technologies, etc. A European Music Observatory will have the latitude to pick certain topics based on priorities and input from sectoral stakeholders. An EMO should consider setting up an “innovation experts’ advisory committee,” constituted of respected professionals in their field who are known for their forward thinking views, to help identify key themes to be studied. (European Commission et al. 2020, p37) We will initiate an informal music innovation expert’s roundtable to discuss potential data needs in this pillar. 4.2.5 Sustainability In the EMO feasibility study the definition of sustainability was mentioned among the innovation topics. Because of the triple transition, introducing the Corporate Social Responsibility Directive and the European Sustainability Reporting Standards have increased the interest and need in sustainability data; we decided to create a separate topical pillar for environmental and social sustainability, or governance indicators (ESG.) We will publish datasets that will be used in the value-added service described in Section 7.2.2. 64 5 Data Sources Figure 5.1: sad In general terms, three main situations have arisen in the analysis of availability of data covering the music sector: 1. Data is available through stakeholders and would be supplied to the EMO at no cost or, if needed, at the cost of processing the data. The real cost would then be that of the human resources necessary to analyse and present the data. 2. Data is available through vendors whose business model is to sell or license data, research and analysis. The data would be made available to the EMO following a commercial and contractual negotiation, the terms of which were not available for this report. 3. Data is not available or not tailored for the needs of the EMO and therefore, access to such data would require EMO to establish the conditions for this data to exist, in partnership with stakeholders and data suppliers, at a cost that is difficult to determine without evaluating exactly the task at hand and the costbenefits of developing such data. (European Commission et al. 2020, p62–63) 65 5.1 Curation of reusable data ĹNote Data is available through stakeholders and would be supplied to the EMO at no cost or, if needed, at the cost of processing the data. The real cost lies in the human resources required to analyse and present the data. The EMO feasibility study gave limited attention to data curation and, in our view, presented an overly optimistic picture of data reuse. Based on our experience, the Open Music Observatory should ingest no data without reprocessing. Even pristine datasets from official statistical sources require proper provenance documentation to synchronize with future revisions or corrections from the original authority. For most music stakeholders, “data available through stakeholders” almost always requires significant curatorial investment. In working with around 100 music organisations, we have found that only a few—such as government-supported music information centres or wellfunded collective management agencies—employ trained staff dedicated to data curation. Our first use case in Slovakia was built around Music Center Slovakia because it has both a competent in-house library and an IT team. Establishing a curation workflow with trained librarians is straightforward when the data is well maintained. However, most music sector organizations are small, often with fewer than five staff, and lack dedicated information science or IT professionals. Figure 5.2: dsafsdf 66 To address these differences, our curation workflows are aligned with three stakeholder categories: 1. Data curation excellence centres – Institutions such as the EU Open Data Portal team, Europeana, Wikibase, Eurostat, national libraries, and official statistical offices. These have advanced competencies in linked open data, conceptual models, and data coordination. 2. Data expert organisations – Collective rights management agencies, digital distributors, larger publishers and labels, and music information centres with competent IT teams and experience reconciling local and external databases. 3. Industry specialists and music scholars – Often working with spreadsheet-based databases and limited IT capacity, they require either training or full curation support. Our technological choices are explained in Section 3.4.1. We have built our system around Wikibase—the software that underpins Wikipedia’s 329 language versions and coordinates extensive linked open data resources such as Wikidata, Wikispecies, and Wikiquote. Wikibase has proven effective in EU and member state projects where stakeholders lacked the resources of data curation excellence centres, bridging complex conceptual models with contributions from domain experts and citizen scientists. Wikibase’s strength lies in its flexible data model: based on the Resource Description Framework (RDF) like most international curation standards, yet simple enough to accommodate less formal data structures. This allows library and archive collection management systems to interface with Wikibase, importing or exporting curated music industry data into more complex models. We detail these procedures in our data improvement services (Chapter 7). 5.1.1 Curation from data vendors ĹNote Data is available through vendors whose business model is to sell or license data, research and analysis. The data would be made available to the EMO following a commercial and contractual negotiation, the terms of which were not available for this report. (European Commission et al. 2020, p39) At this stage, we have not considered acquisitions from commercial data vendors. 67 5.1.2 Novel data assets ĹNote Data is not available or not tailored for the needs of the EMO and therefore, access to such data would require EMO to establish the conditions for this data to exist, in partnership with stakeholders and data suppliers, at a cost that is difficult to determine without evaluating exactly the task at hand and the cost-benefits of developing such data. (European Commission et al. 2020, p39) All Work Packages of the Open Music Europe project will produce novel methodologies and experimental indicators. These will be added to the data catalogue by 31 December 2025. 5.2 Data providers Table 5.1: Organisations that indicated willingness to exchange data with a future European music observatory Stakeholder Description Access for observatory AEPO-ARTIS Organisation representing European artists-performers. Regroups most of the European CMO representing performers. Subject to nonfinancial partnership agreement. Would also require approval from SCAPR board. CEEMID Data collection and integration system based on open data, opensources and online surveys. Interested in contributing and working in partnership with a future EMO. CISAC Trade organisation regrouping rights societies in the world. Subject to nonfinancial partnership agreement. CNM (former CNV) Public organisation managing a tax on concert tickets Subject to nonfinancial partnership agreement. DDEX Standards-setting organisation regrouping all stakeholders in the digital value chain. Interested stakeholder in particular contributing to the Innovation & New Models pillar. GESAC European Grouping of authors societies Interested stakeholder. LIVEUROPE Initiative to support up-and-coming European artists through venues. Subject to partnership agreement. CEEMID is a pan European music data integration system based on open data, opensource software in open collaboration with the industry, statisticians and academia, best statistics, data science and AI practices. It uses many data 68 sources about the audience, the creators of music, music works and recordings, its circulation globally and its economy. Relevance: CEEMID can transfer thousands of indicators that are reproducible and verifiable, open-source software that creates them to a European Music Observatory. In particular, CEEMID provides a useful and interesting approach to harnessing the possibilities of open data in Europe in relation to the music sector, which should be further explored by the European Music Observatory in its start-up phase. (European Commission et al. 2020, p147) CEEMID was a predecessor of the Open Music Observatory. Although some of its data is dated, because it used reproducible research techniques in data collection and processing, some can still be updated and transferred to the Open Music Observatory. We have started this process and will continue to consult with partners about their needs or necessary approvals and permissions in the case of some datasets. 69 6 Standardisation of Data & Terminology Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO Feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. Because standardisation is one of the key services of the envisioned European music observatory, we gave a lot of consideration to the standards to be applied, and the terminology negotiation process among the observatory’s stakeholders. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? 6.1 Business processes Since the Open Music Observatory is primarily a data dissemination hub, the definition of our services (Chapter 3) apply elements of the Generic Statistical Business Process Model (GSBPM), an international standard that describes and defines the set of business processes needed to produce official statistics. The GSBMP is accompanied by the General Statistical Information Model, which builds on the Data Documentation Initiative (DDI) and the Statistical Data and Metadata eXchange (SDMX) (Pellegrino and Grofils 2013). The DDI and SDMX are the foundations of working with social sciences archives, statistical microdata, and processed statistical data. Their key elements are described in the Resource Description Framework of the World Wide Web and can be used in Linked Data. Some elements of DDI are described with RDF: The DDI-RDF Discovery Vocabulary is a draft specification of the DDI Alliance. (Hartmann et al. 2024). Whenever possible, we rely in our observatory with this annotation; if that is not yet possible, we follow the DDI Lifecycle (3.3) Documentation (Data Documentation Initiative 2020). 70 6.2 Conceptual and information models Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? Numerous knowledge institutions store information about musical works, as well as natural persons (humans) who composed or performed these works and contributed to their recorded fixation. If we want to inquire about composers, we must know that a composer is always a human (animals or software agents with AI algorithms cannot be entitled to composer copyrights.) We also must know that a musical work is an abstract creation, manifesting as a notation (physical or digital sheets, MIDI files) or recording (analogue or digital-physical object, or a file.) If we want to validate the composer’s information connected to a recording of a particular musical work, we must access databases containing information about humans concerning some identifying properties of works or recordings. We imagine a future European Music Observatory that is not a specialised knowledge institution and is not a library, archive, museum, or statistical agency. Instead, it should be able to consolidate knowledge from all such institutions and find ways to bring together data from private enterprises and data collection programs to fill the information gaps of the European music sector stakeholders. Our services use the Wikidata Data Model as a data coordination and reconciliation model (Wikimedia Foundation n.d.). In this regard, we follow many successful EU and memberstate, (Alexiev et al. 2020; Diefenbach, Wilde, and Alipio 2021; Rossenova, Duchesne, and Blümel 2022; Faraj and Micsik 2023) or music projects (Siler 2022). We particularly want to mention the excellent work of the University of Helsinki in creating WB CIDOC, a simple business process and data mapping between the Wikidata Data Model and the more complex CIDOC CRM used by extensive collection management systems (Kesäniemi, Koho, and Hyvönen 2022). 71 The StatDCAT-AP and the more general DCAT-AP definition of the EU Open Data Portal provide a bridge among library metadata systems, such as DCMI Metadata Terms (Dublin Core) for libraries, the World Wide Web DCAT standard for publishing datasets, and some core terms of the Statistical Data and Metadata eXchange. Figure 6.1: Our most important reference is the DCAT-AP 3.0 specification, and its extension to statistical data by the EU Open Data Portal. The Europeana Data Model (EDM) similarly provides a more straightforward connection tool among various library, museological or musical collections; it mainly builds on Dublin Core and offers equivalent classes for the more complex CIDOC CRM (Europeana 2017). We see no problem in connecting the EDM towards RiC. The CIDOC Conceptual Reference Model (CRM) provides an extensible ontology for concepts and information in cultural heritage and museum documentation (Bekiari et al. 2024). Last, we mention some novel standards and standard candidates related to documents, microdata, and metadata documentation, such as music survey questionnaires. The Records In Context (RiC) 1.0 CRM and ontology were adopted in November 2023 to replace four international archival standards with backward compatibility. The DDI-Discovery vocabulary is an evolving standard that aims to describe important DDI terms with the World Wide Web standard Resource Description Framework. To keep our systems future-proof, we adopt elements of RiC and DDI-Discovery to document our question bank and codebooks (International Council on Archives Expert Group on Archival Description 2023; Hartmann et al. 2024). 72 ĹNote A future European Music Observatory could help with coordinating European research activities in the music sector. An EMO could also develop tools to establish cooperation between various data collection bodies. The Observatory should, therefore, also be involved in setting standards and developing common EU wide definitions that are crucial for consistency. (European Commission et al. 2020, p80) Since the adaptation of the European Interoperability Framework and similar FAIR measures in open science, such terminological standardisation has taken place in the definition of formal ontologies, i.e., knowledge bases that software applications can use, too. The music observatory should have competent knowledge engineers and ontologists and should be involved in the discussions of sector-agnostic ontology, for example, on the possible improvements of CIDOC or EDM, for a better representation of music. There is also a need for the development of more usable and more widely accepted musicsector ontologies. In T5.1, we have reviewed the Polifonia Ontology Network and the Music Ontology, but we believe both have shortcomings for a full adaptation. 6.3 Identification & Entity Linking Entity linking, also referred to as named-entity linking (NEL), named-entity disambiguation (NED), named-entity recognition and disambiguation (NERD) or named-entity normalisation (NEN) is the task of assigning a unique identity to entities (such as famous individuals, locations, or companies) mentioned in a digital resource, such as a file. ĎTip The MusicBrainz free music database contains records of 20 artists named Paris (artists)�, and 15 locations using the same name Paris (locations)�, which all may enter a data-driven service as artists who must be credited for attribution or royalties, and as a place of an event, release, or publication. Connecting the word Paris to the correct person, group or location is the task of entity linking. Since the inception of the world wide web, data flows across organisations and countries, and the use of local identifiers is not a good solution. International organisations of music, heritage management, science, and national organisations are increasingly shifting to the use of persistent identifiers (or permanent Identifier or handle). ĎTip Apersistent identifier (or permanent Identifier or handle), is one that never changes, so that your bookmarks and links don’t break when a website or a database or an API service gets updated. 73 they meet public catalogue and database data from public libraries, open knowledge graphs, and the Slovak Music Center. The data that should be made public is then further exported to Wikibase Cloud, where it becomes public and available for all stakeholders. From Wikibase, it is also synchronised with Wikidata, the world’s largest open knowledge graph. 7.1.2 Fix-the-data “Fix-the-data” means improving the data quality by finding or imputing missing values or finding and replacing erroneous data entries. In terms of metadata, adding further machineactionable information to already existing datasets can improve their usability. The fix-the-data service can mean replacing missing our outdated metadata (such as a name change of a natural person or a corporate body), or recalculating aggregated accounting or statistical data after base change, or forecasting data that is not yet available. ĎTip Our fix-the-data services do not increase the size of the data available to our partners, but it increases the quality of their datasets or databases. 7.1.3 Data Linking “Data linking”, data fusion, or data matching means correctly joining data from different datasets (data sources.) Many fix-the-data problems initially arise from imperfect data linking, for example, mistakes in currency rates, units of measures, coding of geographical entities, misplaced decimal delimiters on the level of data, or misunderstandings of the meaning of “artist income” or “popularity score”, or other non-self evident variables. An even more subtle problem is joining data from two questionnaire surveys created with different sampling algorithms and different standard (measurement) errors. ĎTip Data linking or data fusion is a way to join many small databases into a large, federated dataset. This way, relatively small music organisatiosn can benefit from access to big data. 7.1.4 Registration services In Open Music Europe, other tasks deal with the policy problem plaguing the music industry: even though it needs access to an exceptionally high number of registers (due to the fragmentation of the copyright and several neighbouring rights), access to such registers 80 is limited or impossible. Often, the registers carry legacy problems that make them less functional in trustworthy data and AI systems. Aregister is a document [in modern usage, usually a database], in which data are entered in a formal manner by a statutory authority (ISO 2017b). In statistical data collection a “register aims to be a complete list of the objects in a specific group of objects or population.” (Anders and Britt 2007). Statistical data collection and rights management are just two service areas whose workflows depend on well-functioning and accessible registers. The statistical business register is an essential tool for creating survey frames or sample frames, in other words, to organise statistical data collection. A copyright or neighbouring right register is necessary to organise royalty collection. ĹNote A statistical register is necessary to decide who should get a data request: • For a sample survey, the register is used to draw a lottery of population members who will be invited to provide data. • In a census-type survey, all registered members of the population, for example, all music labels, will receive an invitation to an interview or form. • In the case of a register-based survey, all members of the register, for example, all collective management societies in the territory, will be requested to send data directly from their databases. In other work packages of the Open Music Europe project, we are experimenting with statistical data coordination among the music sector and statistical authorities. Without recalling the details here, as digitisation exponentially increases the amount of structured data in the private sector, it is a growing trend in statistical innovation to rely on data held by the private sector to make more granular or timely official statistics. For consumer spending statistics, costly and imprecise surveys of randomly selected citizens putting their purchases in a diary, some statistical authorities directly process data from cash registers or credit card spending. We envision a similar statistical collaboration among statistical offices and collective rights management organisations because it is easier to report music royalty accounts than to ask musicians to talk about their complex income streams in interviews or on questionnaires. We see the role of the Open Music Observatory in providing a methodology and digital data infrastructure for such statistical collaboration. In other work package tasks, SOZA and Reprex will create so-called satellite business registers to harmonise the data collection of the observatory with the Slovak statistical authority. More about this work: (Daniel Antal 2023) Such services, similar to data linking and some new services that will build on the data and the data API of the Open Music Observatory, rely on the provision of technical services for registration. The Open Music Observatory has its register, too. 81 Registration is a costly data service with vast economies of scale, so providing more affordable registration services for the European music sector could be an important service. ⊠In our piloting phase, we rely on cooperation with the Slovak National Library to test the usability of the VIAF authority file system for identifying names. ⊠Our dataset distributions use the DOIs from Zenodo, which also provides our longterm archive. ⊠For certain assets, mainly photographs and scanned documents, we rely on the new ISCC registration, a long-term solution for some music industry applications. ⊠Reprex registered an imprint, the Digital Music Observatory, to place long-form publications as books into library systems. ⊠We are investigating the costs and benefits of finding an ISNI registrar partner or creating a roadmap for making the Open Music Observatory a registrar itself. 7.2 Use Cases In 2023 the Open Music Europe project applied for the Module A of the Horizon Results Booster (HRB) provided by Trust-IT Services�. The HRB aims to provide a tangible contribution to the dissemination of results and recommendations of research projects related to the European Commission Priority areas. Figure 7.1: app 82 7.2.1 Data Health Services for Collective Management Entity linking and data linking are among the biggest technical problems in rights management. Because music authors, producers, and performers have three royalty streams and do not share an interoperable registry, the connection of musical works (compositions, ideally identified by an ISWC code), their sound recording manifestations (identified on all digital services with and ISRC code), and the various identifiers of performers require costly manual and technical identification. There are numerous projects underway in the music industry to resolve this problem going forward. In the United Kingdom, PRS’s Nexus programme� is developing a solution with the provisioning of preliminary ISWC registration to keep the recording and composition connected from the birth of a new recording. The Open Music Europe project, on the other hand, is pioneering a different route for already existing sound recordings, with the linking of public sector catalogues of heritage and library collections with rights management information; particularly with relying on the VIAF shared authority files. SOZA and Reprex are expected to present their MVP on the CISAC Good Governance seminar in December 2025. Modern registers typically assign a unique identifier, known as a URI, to their data subjects (our registered objects). A ‘Cool URI’, which resembles a URL, offers a practical advantage. When used as a URL, it generates a human-readable HTML file about the registered person or object. This can be particularly useful when processed by a graph application, as it provides crucial information about this person or object in a machine-readable (XML, JSON, TTL, or NQUAD) file. For example, the VIAF identifier number 89006617 can be placed into the http://viaf.org/ viaf/89006617 URL, which provides as access to the cataloging information of works created by, or written about the great etnomusicologist and modern composer, Béla Bartók. Modern platforms, such as Spotify, use similar identifiers. For example, the Spotify Artist ID 2fIUlieTjLTaNQUIKHX5B8 resolves to Celeste Buckingham’s available recordings on the platform via the URL https://open.spotify.com/artist/2fIUlieTjLTaNQUIKHX5B8. The problem is that music creators are often present on more than 200 digital platforms, each of which has its identifier policy and requires the repeated import of the artists’, works’, and recordings’ data. To consistently report such metadata is costly and complex, even for major labels and publishers with a dedicated IT system. No wonder we saw before our project in our own Feasibility study that more than 50% of artist data needed fixing on digital platforms. Relying on many local identifiers on otherwise interconnected computer systems will always create a costly and error-prone data exchange. Unfortunately, the music industry has never agreed to use genuinely open, high-quality registers. These changes were made during the period of our project. For example, large platforms like Apple, Spotify, and some collective rights management organisations started using the ISO-standard name identifier (ISNI) to avoid the high prevalence of multiple same-name persons and musical groups. This transition 83 is yet to begin, and it is incomplete, so the music sector will likely need to invest large IT resources into entity resolution in the next decade. 7.2.2 Sustainability Reporting for Music Organisations The Music Innovation Hub and Reprex will develop a CSRD-compliant sustainability reporting tool in 2024-2025. The reporting tool aims to provide an accurate and affordable ESG reporting facility that follows the European ESRS standards for music enterprises that create their financial reports according to the simplified reporting rules allowed by member states for microenterprises. More than 95% of European music enterprises (in some member states, this reaches 100%) apply simplified financial reporting. For such companies, there are no CSRD-compliant ESG reporting tools. We identify the reason for this market failure as follows: □The CSRD Directive imposes the responsibility of connected financial sustainability reporting on large and public companies and applies it to their entire value chain. The music industry lacks such large enterprises that would have taken a piloting role or played a pivotal role in establishing the standards. □Music enterprises and their trade associations do not act proactively because they believe they must follow the data provision instructions of the directly affected B2B buyers, financiers, or corporate sponsors. □The standardisation body EFRAG has de-prioritised the cultural and creative industries in setting industry-specific standards favouring sectors with a much higher adverse environmental impact. 84 □While small music businesses do not feel a compliance push, as they are not directly responsible for applying the ESRS, they also miss out on the opportunities provided by green financing and insurance. ⊠MiH and Reprex will pilot a service suitable for microenterprises, reducing compliance costs from 1500 euros to 500 euros per entity. ⊠This new application will rely on the Open Music Observatory’s Music Economy and Sustainability pillars and will derive its benchmarks, science-based targets and coefficients, and input-output tables. The MVP of this service was developed with a MusicAIRE microgrant, and it is the project’s background. A scale-up will be demonstrated with the use the Open Music Observatory’s open data API. 7.2.3 Listen Local Figure 7.2: The Feasibility Study On Promoting Slovak Music in Slovakia And Abroad is an important background of our project. In 2020, with a microgrant from the Slovak Arts Council, we created a Feasibility Study and a demo application called Listen Local (Daniel Antal 2020b). The study examined why the Spotify algorithm struggled to recommend Slovak music within Slovakia for Slovak people. We also created a demo application that modified the user’s Spotify recommendations to voluntarily comply with the local content guidelines applicable to local radio stations. The user could also listen to a lower or higher percentage of regional works. Our critical finding was the very pool data coverage and quality of the Slovak repertoire, which is mainly sent to distribution without the professional assistance of a commercial music label. Self-releasing artists and micro labels do not have the necessary metadata know-how, IT and data specialists to prepare their new releases for algorithmic curation by recommender engines of digital streaming platforms, radio stations, or large festivals. 85 Figure 7.3: Our conceptual demo application was able to make recommendations on voluntarily meeting the local content guidelines, but it was only supported by a relatively small Slovak Demo Music Database, and could only work with Spotify, which has the most transparent and open API of all streaming providers licensed to the territory of the Slovak Republic. We aim to develop applications to create a local content-aware public performance music stream. ⊠HearDis! aims to integrate such location-aware metadata into its background music playlisting service. ⊠We are planning Listen Local applications for radio stations to voluntarily review their current playlists for compliance with local content regulations and, if they fail to reach the statutory local content quotas, to recommend suitable recordings to their playlists. ⊠The OMO will disseminate the necessary data for these new services. □The data is not yet available, as the creation of the Slovak Comprehensive Music Database is a task of its own that will be ready by November 2025 in WP2. 7.2.4 Unlabel Unlabel is a planned service aimed at self-releasing artists and micro labels that need a functional data/IT department. Therefore, they are at a disadvantage compared to significant independent and major releases because they usually need to meet the high documentation standards necessary for a successful digital distribution strategy and engagement with algorithmic curation of streaming-, radio-, or festival playlists. Self-releasing artists and micro labels bring ill-documented new content to digital distributors like ALOADED. Digital distributors must maintain an arm’s length standard for all 86 labels, small or large, independent or major. ALOADED or other distributors cannot crossfinance the data problems of self-releasing and micro-label artists from the client revenues of more prominent labels. We identify the problem as a market failure and a technical failure: □In some developing markets, insufficient royalty revenues do not allow the presence or professionalisation of record labels with an IT and data management function because the payment of IT or data specialists or to keep external suppliers at least on a retainer cannot be financed from the label-artist revenue split. □Manual metadata provision without metadata specialists and tools leads to inferior data quality. Our feasibility study has shown that more than 50% of the releases have data shortcomings, and 17% have poor data representations that make algorithmic recommendations for these releases impossible. This creates a vicious circle because poor data quality translates into low visibility, low usage of such repertoire, and, therefore, low income. The cost of data improvement has no sustainable financial basis. ⊠In 2025, ALOADED, Reprex, Slovak Music Center and SOZA will conceptualise and plan a new public-private business model that aims at those rightsholders who do not have a technically proper label representation as a substitute for non-available market services. Our planned “Unlabel” service will provide documentation and metadata improvement services for self-releasing artists. This service, similar to current white-label services, will strictly address market failures and not compete with label services. We aim to provide a necessary level of data consolidation and improvement so that these artists can have equal opportunities in digital distribution services. The service will be connected to the Slovak Music Dataspace and its Slovak Comprehensive Music Database. We will provide a PPP business model for the onboarding and proper documentation of self-releasing artists on a large scale and the efficient, API-based provision of their digital distributor. Aloaded will provide the distribution services, Reprex will provide the data services, and SOZA and the Slovak Music Center will work out the details of minimal customer service for such labels. 7.3 Use of AI systems For the entity linking, related to our planned value added Section 7.2.1, we are planning to use in the future AI algorithms, particularly inference engines. The main goal of the system is to help matching correctly named entities, particularly rightsholders, musical works and recordings. The system is not yet in place. An adequate description will be provided for overview and will be brought to the attention of the Ethics Advisor during the upcoming meeting of the Ethics Board. 87 We cannot provide a full risk assessment because the service is not planned in detail yet. However, our preliminary risk assessment suggests low levels of risk, partly, because we plan to deploy AI in music/culture, which as a domain not seen as a high-risk area by the European regulation, and partly, because our system will not autonomous, will retain human-in-control, and will not influence the decisions or anyhow engage with end-users. We were conscious of the potential risk involved, and both the control structure and the data governance were planned over the course of 10 months. □Is the AI system designed to interact, guide or take decisions by human end-users that affect humans or society? No. The system will only help qualified persons in rights management to faster and more efficiently preview potentially unlinked entities. □Could the AI system affect human autonomy by interfering with the end-user’s decision-making process in any other unintended and undesirable way? No. The system in no way is considered as an end-user system. ⊠Please determine whether the AI system (choose as many as appropriate) overseen by a human: Is overseen by a Human-in-Command. ⊠Have the humans (human-in-the-loop, human-on-the-loop, human-in-command) been given specific training on how to exercise oversight? Yes. The system is not making autonomous decisions. ⊠Is your AI system being trained, or was it developed, by using or processing personal data (including special categories of personal data)? Yes. ⊠Did you put in place any of the following measures some of which are mandatory under the General Data Protection Regulation (GDPR), or a non-European equivalent? Yes. ⊠Data Protection Impact Assessment (DPIA) Yes. ⊠Designate a Data Protection Officer (DPO)24 and include them at an early state in the development, procurement or use phase of the AI system? Yes. ⊠Oversight mechanisms for data processing (including limiting access to qualified personnel, mechanisms for logging data access and making modifications)? Yes. □Measures to achieve privacy-by-design and default (e.g. encryption, pseudonymisation, aggregation, anonymisation)? Not applicable for NERD. The aim of the application is to detect errors in name attribution and to protect the moral and economic rights of the (named) rightsholders. ⊠Did you implement the right to withdraw consent, the right to object and the right to be forgotten into the development of the AI system? Yes. ⊠Did you consider the privacy and data protection implications of data collected, generated or processed over the course of the AI system’s life cycle? Yes. ⊠Did you consider the privacy and data protection implications of the AI system’s non-personal training-data or other processed non-personal data? Yes. 88 We do not consider that the system has wider risks or negative impacts. The algorithm is designed to cure sources of data biases that result in a late or missed payment for some rightsholders. 89 To reach our objectives related to the Open Music Observatory, we will take the following steps after this milestone: ⊠We introduce the system to those organisations that showed a willingness to share data with a future European observatory (See: Section 5.2), and ask for their feedback and at least test data samples; ⊠Based on the external data samples, in our supporting task, we will develop further extract-transform-load components to ensure a smooth data export/import according to the needs of the representative European music stakeholders; ⊠We will disseminate and improve with feedback our data sharing space model, and particularly the accompanying governance model, with experts in the field of data governance, data sharing, and various music industry data and metadata interest groups; ⊠We will finalise the creation of the Slovak Comprehensive Music Database, enabling the MVP demonstrations of several value-added services, as described in Section 7.1. We plan to start disseminating these use cases on large stakeholder forums in the second half of 2025 in the hope of finding additional replication partners and users. ⊠An adequate description of our planned use of AI will be provided for overview, and will be brought to the attention of the Ethics Advisor during the upcoming meeting of the Ethics Board (see Section 7.3.) ⊠We will ingest a large batch of survey data from other work packages in 2025 Q1. 96 References Albertoni, Riccardo, David Browning, Simon Cox, Alejandra Gonzalez Beltran, Andrea Perego, and Peter Winstanley, eds. 2020. “Data Catalog Vocabulary (DCAT) - Version 2.” W3C. https://www.w3.org/TR/2020/REC-vocab-dcat-2-20200204/. Alexiev, Vladimir, Plamen Tarkalanov, Nikola Georgiev, and Lilia Pavlova. 2020. “Bulgarian Icons in Wikidata and EDM.” Digital Presentation and Preservation of Cultural and Scientific Heritage 10: 45–63. https://doi.org/10.55630/dipp.2020.10.2. Anders, Wallgren, and Wallgren Britt. 2007. Register-Based Statistics. Administrative Data for Statistical Purposes. 1st ed. Chichester:United Kingdom: John Wiley & Sons Ltd. Antal, Daniel. 2020a. “Central And Eastern European Music Industry Report 2020.” CEEMID, Consolidated Independent. https://doi.org/10.13140/RG.2.2.21450.31686. ———. 2020b. “Feasibility Study on Promoting Slovak Music in Slovakia & Abroad.” https://doi.org/10.5281/zenodo.6427514. ———. 2023. “Pilot Program for Novel Music Industry Statistical Indicators in the Slovak Republic.” Zenodo. https://doi.org/10.5281/zenodo.8399254. ———. 2024a. “A Szlovák Adatkicserélési Tér Magyarországi Föderációjának Lehetőségei.” Eszterházy Károly Katolikus Egyetem. https://doi.org/10.31915/NWS.2024.25. ———. 2024b. “Trustworthy AI and Data-Sharing Spaces for the Slovak Music Centre.” Open Music Observatory. https://doi.org/10.5281/zenodo.16540605. ———. 2025a. “SKCMDb. Interoperability of Music Libraries and Archives with Public and Private Music Services.” Open Music Observatory. https://doi.org/10.5281/zenodo. 16634558. ———. 2025b. “Slovak Music Data Sharing Space.” Open Music Observatory. https: //doi.org/10.5281/zenodo.15814286. Antal, Daniel, Mester Anna Márta, Ieva Pigozne, and Mihaly Nagy. 2025. “Dataset of the Multilingual Gazetteer of the Settlements on the Livonian Coast of Northern Kurzeme.” Finno-Ugric Data Sharing Space. https://doi.org/10.5281/zenodo.15723051. Antal, Daniel, Mária Kmety Barteková, and Katarína Remeňová. 2023. “Economy of music in Europe: Novel data collection methods and indicators.” Zenodo. https://doi.org/10. 5281/zenodo.8334648. Antal, Daniel, Anna Márta Mester, and Ieva Pigozne. 2025. “Remapping the Livonian Coast: A Multilingual Gazetteer of the Settlements of Northern Kurzeme.” Finno-Ugric Data Sharing Space. https://doi.org/10.5281/zenodo.15668712. Antal, Dániel, and Anna Márta Mester. 2025. Open Music Registers.https://doi.org/10. 5281/zenodo.14767717. Artisjus, HDS, SOZA, and Candole Partners. 2014. “Measuring and Reporting Regional Economic Value Added, National Income and Employment by the Music Industry in a Creative Industries Perspective. Memorandum of Understanding to Create a Regional 97 Music Database to Support Professional National Reporting, Economic Valuation and a Regional Music Study.” Bekiari, Chryssoula, George Bruseke, Erin Canning, Martin Doerr, Philippe Michon, Christian-Emil Ore, Stephen Stead, and Velios Athanasios, eds. 2024. “Definition of the CIDOC Conceptual Reference Model.” CIDOC CRM Special Interest Group. https://www.cidoc-crm.org/sites/default/files/cidoc_crm_version_7.2.4.pdf. Camp, Ann Van, Sven Lieber, and IFLA. 2022. “ISNI, a Top Tool for Quality Enhancement, Smooth Data Flows and Efficient Internal Processes.” Dublin: Ireland: International Federation of Library Associations; Institutions (IFLA). https://repository.ifla. org/handle/123456789/2008. Cruz, Maria, and Clifford Tatum. 2021. “NWO Persistent Identifier Strategy.” Zenodo. https://doi.org/10.5281/zenodo.4674513. Curry, Edward. 2020. “Dataspaces: Fundamentals, Principles, and Techniques.” In RealTime Linked Dataspaces: Enabling Data Ecosystems for Intelligent Systems, 45–62. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-296650_3. Data Documentation Initiative. 2020. “DDI Lifecycle (3.3) Documentation.” https://ddilifecycle-documentation.readthedocs.io/en/latest/index.html. Diefenbach, Dennis, Max de Wilde, and Samantha Alipio. 2021. “Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph.” In ISWC 2021. Online, France. https://hal.science/hal-03353225. EBU, and Gaia-X. 2022. “Dataspace for Cultural and Creative Industries. Position Paper. v.2.0.” Gaia-X. https://gaia-x.eu/wp-content/uploads/2022/10/EBU_position-paper_ Media-Data-Space.pdf. Ernštreits, Valts. 2020. “Livonian Place Names: Documentation, Problems, and Opportunities.” Eesti Ja Soome-Ugri Keeleteaduse Ajakiri. Journal of Estonian and Finno-Ugric Linguistics 11 (1): 213–33. https://doi.org/10.12697/jeful.2020.11.1.09. European Commission. 2021a. “Brochure for Music Moves Europe Preparatory Action 2019.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/ mme_2019_brochure_final-web.pdf. ———. 2021b. “Music Moves Europe - First Dialogue Meeting. Final Report.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/mme-conferencereport-web.pdf. European Commission, Directorate-General for Communications Networks, Content, Technology, K Blind, S Pätsch, S Muto, M Böhm, T Schubert, P Grzegorzewska, and A Katz. 2021. The Impact of Open Source Software and Hardware on Technological Independence, Competitiveness and Innovation in the EU Economy : Final Study Report. Publications Office. https://doi.org/10.2759/430161. European Commission, Directorate-General for Education, Youth, Sport and Culture, M Clarke, P Vroonhof, J Snijders, A Le Gall, B Jacquemet, et al. 2020. Feasibility Study for the Establishment of a European Music Observatory : Final Report. Publications Office of the European Union. https://doi.org/10.2766/9691. Europeana. 2017. “Definition of the Europeana Data Model V5.2.8.” Europeana. https://pro.europeana.eu/files/Europeana_Professional/Share_your_data/Technical_ requirements/EDM_Documentation//EDM_Definition_v5.2.8_102017.pdf. Faraj, Ghazal, and András Micsik. 2023. “Enriching Wikidata with Cultural Heritage Data 98 from the COURAGE Project.” In, 407–18. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-36599-8_37. Fragkou, Pavlina. 2023. “DCAT-AP 3.0.” Edited by Makx Dekkers, Pavlina Fragkou, Natasa Sofou, and Bert Van Nuffelen. https://semiceu.github.io/DCAT-AP/releases/3. 0.0/. Hartmann, Thomas, Sarven Capadisli, Franck Cotton, Richard Cyganiak, Arofan Gregory, Benedikt Kämpgen, Olof Olsson, Heiko Paulheim, Joachim Wackerow, and Benjamin Zapilko. 2024. “DDI-RDF Discovery Vocabulary. A Vocabulary for Publishing Metadata about Data Sets (Research and Survey Data) into the Web of Linked Data.” Edited by Thomas Hartmann, Richard Cyganiak, Joachim Wackerow, and Benjamin Zapilko. W3C. https://rdf-vocabulary.ddialliance.org/discovery.html. Huyer, Esther, and Laura van Knippenberg. 2020. The Economic Impact of Open Data. Opportunities for Value Creation in Europe. Luxembourg: Publications Office of the European Union. https://doi.org/10.2830/63132. International Council on Archives Expert Group on Archival Description. 2023. “Records in Contexts–Conceptual Model. Version 1.0.” International Council on Archives. https: //www.ica.org/app/uploads/2023/12/RiC-CM-1.0.pdf. International ISRC Registration Authority. 2021. “International Standard Recording Code (ISRC) Handbook. 4th Edition.” International ISRC Registration Authority. https: //www.ifpi.org/wp-content/uploads/2021/02/ISRC_Handbook.pdf. ISO. 2012. “International Standard Musical Work Code (ISNI). ISO 27729:2012.” International Organization for Standardization. https://www.iso.org/standard/44292.html. ———. 2013. “ISO 17369:2013(en) Statistical Data and Metadata Exchange (SDMX).” London:United Kingdom: International Organization for Standardization. https://www. iso.org/obp/ui/en/#iso:std:iso:17369:ed-1:v1:en. ———. 2017a. “ISO/IEC 19941:2017(en), Information Technology — Cloud Computing — Interoperability and Portability.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso-iec:19941:ed-1:v1:en. ———. 2017b. “ISO/IEC 5127:2017(en), Information and Documentation — Foundation and Vocabulary.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso:5127:ed-2:v1:en. ———. 2017c. “ISO 2108:2017 (En), Information and Documentation — International Standard Book Number (ISBN).” International Organization for Standardization. https: //www.iso.org/standard/65483.html. ———. 2019a. “ISO/IEC 20546:2019 Information Technology — Big Data — Overview and Vocabulary.” London:United Kingdom: International Standards Organisation. https: //www.iso.org/obp/ui/en/#iso:std:iso-iec:20546:ed-1:v1:en. ———. 2019b. “International Standard Recording Code (ISRC). ISO 3901:2019.” International Organization for Standardization. https://www.iso.org/standard/64817.html. ———. 2020. “ISO/IEC 22624:2020(en), Information Technology — Cloud Computing — Taxonomy Based Data Handling for Cloud Services.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:isoiec:22624:ed-1:v1:en. ———. 2022. “International Standard Musical Work Code (ISWC). ISO 15707:2022.” International Organization for Standardization. https://www.iso.org/standard/83125. html. 99 ———. 2023a. “ISO/IEC 11179-1:2023(en), Information Technology — Metadata Registries (MDR) — Part 1: Framework.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso-iec:11179:-1: ed-4:v1:en. ———. 2023b. “ISO/IEC 2382:2015(en), Information Technology — Vocabulary.” London:United Kingdom: International Standards Organisation. https://www.iso.org/obp/ ui/en/#iso:std:iso-iec:2382:ed-1:v2:en. Kesäniemi, Joonas, Mikko Koho, and Eero Hyvönen. 2022. “Using Wikibase for Managing Cultural Heritage Linked Open Data Based on CIDOC CRM.” In New Trends in Database and Information Systems, edited by Silvia Chiusano, Tania Cerquitelli, Robert Wrembel, Kjetil Nørvåg, Barbara Catania, Genoveva Vargas-Solar, and Ester Zumpano, 542–49. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-03115743-1_49. Ministerstvo kultúry SR, and Open Music Europe. 2023. “Memorandum o porozumení o využití výsledkov analýz otvorených politík v kontexte slovenského kultúrneho a kreatívneho priemyslu a sektorových verejných politík v spolupráci s konzorciom pre výskum a inovácie s názvom OpenMuse. [Memorandum of Understanding on utilizing the Open Policy Analysis results of the OpenMuse Research and Innovation Consortium in the context of Slovak cultural and creative industries and sectors’ public policies].” https://www.crz.gov.sk/zmluva/7645338/. Nagel, Lars, and Douwe Lycklama, eds. 2021. “Design Principles for Data Spaces. Position Paper. Version 1.0.” Open DEI. https://doi.org/10.5281/zenodo.5244997. Open Music Europe. 2023. “Open Music Europe (OpenMusE) – An Open, Scalable, Data-to-Policy Pipeline for European Music Ecosystems.” https://doi.org/10.3030/ 101095295. Pellegrino, Marco, and Denis Grofils. 2013. “DDI-SDMX Integration and Implementation. Working Paper.” United Nations Economic Commission for Europe. https://unece.org/ fileadmin/DAM/stats/documents/ece/ces/ge.40/2013/WP5.pdf. Pomerantz, Jeffrey. 2015a. “Definitions.” In Metadata, 19–64. Cambridge, MA, USA: The MIT Press. http://www.jstor.org.proxy.uba.uva.nl/stable/j.ctt1pv8904.6. ———. 2015b. Metadata. The MIT Press Essential Knowledge Series. Cambridge, MA, USA: MIT Press. Rossenova, Lozana, Paul Duchesne, and Ina Blümel. 2022. “Wikidata and Wikibase as Complementary Research Data Management Services for Cultural Heritage Data.” In CEUR Workshop Proceedings.https://serwiss.bib.hs-hannover.de/frontdoor/deliver/ index/docId/2573/file/rossenova_etal2022-wikidata_research_data_mgmt.pdf. Siler, M. 2022. “Beyond the Fountain: Mapping a New Entry Point to the Society of Independent Artists.” Art Documentation 41 (2): 219–41. https://doi.org/10.1086/ 722172. Sofou, Natasa, and Adina Dragan. 2019. “StatDCAT-AP – DCAT Application Profile for Description of Statistical Datasets. Version 1.0.1.” European Commission. https://joinup.ec.europa.eu/collection/semantic-interoperability-communitysemic/solution/statdcat-application-profile-data-portals-europe/release/101. Stahl, Reinhold, and Patricia Staab. 2018. Measuring the Data Universe: Data Integration Using Statistical Data and Metadata Exchange. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-319-76989-9. 100 UNECE. 2014. “Generic Statistical Information Model. GSIM V2.0 Documents. UNECE Statswiki.” 2014. https://statswiki.unece.org/display/gsim/GSIM+v2.0+documents. ———. 2019. “Generic Statistical Business Process Model. GSBPM V5.2 Documents.” UNECE Statswiki. January 2019. https://statswiki.unece.org/display/GSBPM/GSBPM+ v5.1. Vardigan, Mary, Pascal Heus, and Wendy Thomas. 2008. “Data Documentation Initiative: Toward a Standard for the Social Sciences.” International Journal of Digital Curation 3 (1): 107–13. Wickett, Karen M., Antoine Isaac, Katrina S. Fenlon, Martin Doerr, Carlo Meghini, Carole L. Palmer, and Jacob Jett. 2013. “Modeling Cultural Collections for Digital Aggregation and Exchange Environments.” CIRSS Technical Report 201310-1, October. https://hdl. handle.net/2142/45860. Wikimedia Foundation. n.d. “Wikibase Data Model.” Wikimedia Foundation. Accessed May 26, 2024. https://www.mediawiki.org/wiki/Wikibase/DataModel. 101 Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network Name of the stakeholder: □Corporate (institutional) name: This is mandatory for legal persons and groups; also, please provide a contact person. □given name: mandatory for natural persons □family name: mandatory for natural persons Legal status of the stakeholder: ⊠private person: your name will be made public among the members; we also ask to provide a persistent ID (ISNI, ORCiD or VIAF.) ⊠Legal person: your name and legal person ID will be made public among the members, but not the contact person. Please provide ISNI, or OpenCorporates ID. □other association or group without legal personality: your name and persistent ID will be made public among the members, but not that of the contact person. Please provide ISNI identifier. Logo or icon of the stakeholder: □not mandatory, but if provided, we will make it public Online contact details of the stakeholder: □Official website: we will make it public if provided □LinkedIn page: we will make it public if provided □Facebook page: we will make it public if provided □YouTube channel: we will make it public if provided □Instagram account: we will make it public if provided Permanent residence or seat of the stakeholder: ⊠Only the country and municipality will be made public □Please provide full postal address Official email address of the stakeholder: 102 □only used for invitations to stakeholder meetings, giving and revoking data handling; never made public. Contact telephone number of the stakeholder: □is only used to clarify potential problems and consents; it is never made public and is not used unless necessary. For the intake, we will also ask for a few lines of statement of interest in the Observatory Stakeholder Network and topic interests for data if they apply. We will make this information public, too. We are also happy to create an intake interview and publish it on our website to allow the members of the Observatory Stakeholder Network to get familiar with each others ideas and interests. ĺImportant We will provide in the final deliverable a link for the official intake to the stakeholder network. I would like to invite SOZA, Hudobné centrum, Aloaded, and HearDis! as first partners. Open Music Data Exchange ĺImportant This advisory body will not be open for invitations. The representative pan-European stakeholders can nominate here their technical providers to consult the technical aspects of data exchange. 103 Annex 2 - Data Curators Manuals Our primary services involve data collection, processing, and dissemination. These services will not produce a high-quality data resource without competent data curators. Data curators, as professionals, are responsible for managing, maintaining, and enhancing the quality of an organisation’s data. Their work is instrumental in making data easily accessible, accurate, and relevant to the organisation’s needs. In large organisations, they collaborate closely with data engineers, analysts, scientists, and other stakeholders to establish a robust data ecosystem. ĹNote To work with the systems of the Open Music Observatory, we created two manuals. • In March 2023 we created (and subsequently updated with user feedbacks) a Contributors Guideline for the internal stakeholders of the Open Music Europe (OpenMusE) project. This manual is available on contributors.dataobservatory.eu • In 2024 we created a partly overlapping manual for the Observatory Stakeholder Network, i.e., for organisations that are providing and exchanging data with the Open Music Observatory. This manual is available on manual.opencollections.net. In the music sector, because of the dominance of microand small enterprise (institution) sizes, very few competent data curators and specialised data or knowledge engineers are present. Our approach to solving this problem is the following: ⊠We pool those music experts within the stakeholder network who have data curatorial skills (for example, music librarians) or, due to their job, have background skills or an affinity to data curation. ⊠We provide these data curators with robust tools that only require a little learning. ⊠We centralise all the knowledge and data engineering work in the centre of the datasharing space, i.e., at the Open Music Observatory. We provide small-group training and online manuals for the data curators, responsible for maintaining the quality of data ingested by the Open Music Observatory. For large data providers, for example, collective management organisations or music information centres, we train one in-house data curator because, due to data confidentiality issues, often only an 104 in-house person can review the totality of the data (and sort out the part of the data that can be shared within the dataspace with other observatory stakeholders.) Inspiration Very few music organisations employ data curators or employees with a library or information sciences background. We want to encourage music professionals working in various for-profit or social enterprises and research institutions to discover their “inner data curator.” We believe that a passion for music and the sector and deep knowledge and experience with music are more important than the technical skills needed to curate the data. We are encouraging the members of the Observatory Stakeholder Network (See: Section 8.1.1) to find professionals, researchers, or artists in their organisation who have deep subject-domain knowledge about the data we want to improve: they know a lot about organs in churches, about labels of a particular genre, sync licensing to films, or any other domain on which we collect data. Our ideal curators share a passion for data-driven evidence or visualisations, can learn tools that Wikipedia editors use, and have a robust and subjective idea about the data that would inform them in their work. Basic Data Organisation Concepts We are training and providing self-training material for two crucial but relatively simple concepts: tidy data and the text annotation and mark up. Figure 3: Following three rules makes a dataset tidy: variables are in columns, observations are in rows, and values are in cells. From R For Data Science - 12. Tidy Data Our documentation system works with MediaWiki, the mark up system developed by Wikipedia. 105 Musical works Sound recordings Live public performance 112 SKCMDb: Slovak Comprehensive Music Database The Slovak Comprehensive Music Database (SKCMDb) is a national initiative aimed at making Slovak music more accessible, discoverable, and usable across libraries, archives, streaming services, and rights management organisations. It connects scores, recordings, and metadata using open standards and collaborative governance. As a functional module of the Open Music Observatory, the SKCMDb also serves as a testbed for developing shared data services, addressing the conceptual models, workflows, and governance rules required to link diverse music stakeholders. ‘ The SKCMDb is supported by a data-sharing space consisting of both shared and private databases. The data-sharing space currently comprises the following initial components: •Slovak Metadata Database: A database that facilitates connections between various Slovak stakeholders’ systems. •SKCMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 113 •SKCMDb HC-SOZA (private): A database used exclusively for rights management and library management, governed by an agreement between the Slovak Music Centre and SOZA. •SKCMDb HF (private): A technical dataset created to provide improvements, corrections, and enrichments for the Hudobný fond. The Slovak Metadata Database serves as a support layer that is partly public and partly private. Its metadata definitions and descriptive metadata are exported into the SKCMDb databases as needed and permitted. The Slovak Metadata Database The Slovak Metadata Database is developed in alignment with the metadata framework of the Open Music Observatory. •Ontological and thesauri patterns: Reuses standardized or widely adopted vocabularies. •Conceptualizations and definitions: Includes concept definitions, thesauri, and other elements developed specifically for the SKCMDb. •Public permanent identifiers: Uses identifiers that are public or can be made public. The metadata layer is generally licensed under CC0, though in some cases other licenses are used (for example, CC-BY). ĹNote Examples: • The definitions of musical work,printed sheet music, and the is score of relationship allow the description of connections between an abstract musical work—such as Bella’s Missa in C—and its actual printed manifestations. • The VIAF identifier 2737220 identifies Ján Levoslav Bella’s compositions across library systems. Slovak Comprehensive Music Database (public) The Slovak Comprehensive Music Database is a linked open database published by the Slovak Music Centre. It integrates elements from the Hudobné centrum’s own databases along with data made public by SOZA, Hudobný fond, and other organizations. The database is distributed under various Creative Commons licenses that allow both commercial and non-profit use. 114 The primary aim of the SKCMDb is to reduce the cost of maintaining public and private services that enhance the circulation, availability, visibility, and legally licensed use of Slovak music. Our licensing policies are designed to enable the widest possible use of the data while protecting the investments required to maintain registers, standards, and data integrity. Slovak Comprehensive Music Database (private) The private components of the SKCMDb consist of databases where the data is not intended for public sharing but is used to enhance rights management, music information services, library operations, or other specialized applications. These databases are maintained under agreements between the participating parties. Access to these private datasets serves specific, well-defined purposes and is governed by strict rules. Availability to third parties is determined solely at the discretion of the data owners and may vary depending on contractual or legal obligations. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category (e.g., a national musical works register). Collections typically rely on registers to identify works unambiguously and avoid duplication. Both are documented in structured datasets containing standard identifiers such as ISRC, ISWC, or ISMN codes. In the case of statistical data, microdata often refers to survey instruments and responses curated under defined methodological rules. •Metadata: Relevant elements from the Slovak Metadata Database that support the use of collections or registers. The SKCMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models 115 and data structures, enabling automatic aggregation. All datasets are available with RDF annotation and can be exported in all standard RDF serializations. Microdata is intended for institutional and professional use, not for the general public. It is annotated with standardized metadata suitable for applications such as music library cataloguing, distribution platforms, and rights management systems. ĹNote Example The musical work String Quartet in B-flat Major (composed by Ján Levoslav Bella) is described in a dataset that includes two publicly available printed scores and a publicly available sound recording. •Graphical view: Navigate and contribute to detailed entries on musical works, sound recordings, and related assets. � Explore on Wikibase •Semantic view: Export structured data in XML, JSON-LD, Turtle, or NTriples formats for reuse in research or digital projects. � Example Turtle file or download XML Recording of microdata follows defined rules. As a general principle, living natural persons may opt out of inclusion in the database. Statistical Data & Data Catalogue Our statistical data and catalogue consist of datasets aggregated using statistical methodologies. These comply with the SDMX standard and the W3C Data Cube vocabulary, making them compatible with spreadsheet software, statistical packages, and data science workflows in R, Python, or similar environments. Datasets are offered in multiple formats. In addition to RDF serializations, we provide standard CSV files and, when required, Excel or SPSS formats. We also publish data papers and related documentation that describe dataset usability and highlight key insights. Publications & Catalogue The SKCMDb’s most important publications are musical works, made available to end users as sound or video recordings and printed sheet music. These may be distributed in different sales formats, such as physical albums or books. Microdata and statistical datasets are treated as publications and are listed both in the general catalogue and a machine-readable data catalogue. The SKCMDb also includes methodological and musicological publications, as well as data papers explaining the use of datasets. The catalogue is designed for interoperability with libraries, archives, museums, and similar institutions. 116 In some cases, the SKCMDb may host the full publication, with the Open Music Observatory acting as publisher. In most cases, however, it provides catalogue entries with clear access points, such as webshops, public library lending systems, or repository links where legal copies of the documents can be obtained. 117 LīvMDb: Livonian Music Database The Livonian Music Database (LīvMDb) is a proof-of concept for working with music that has low documentation depth, weak institutions. The music of the Livonian people is scattered, and as native speakers of this small ethnic group died out, their heritage was dispersed, and largely not placed on modern digital platforms. The LīvMDb as a functional module of the Open Music Observatory, serves as a testbed for working with very low documentation, decolonisation, and other issues related to regional cultures, ethnic minorities. It offers insights into subsidiarity. Figure 7: The LīvMDb is federated with the Finno-Ugric Data Sharing Space and the Open Music Observatory. The first places Livonian music into a wider Finno-Ugric cultural context, the latter into a music context. The LīvMDb is supported by a data-sharing space consisting of both shared databases. The data-sharing space currently comprises the following initial components: •Finno-Ugric Metadata Database: A database that facilitates connections between various Finno-Ugric and Baltic stakeholders’ systems. •LīvMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 118 •LīvMDb (private): A technical dataset created to provide improvements, corrections, and enrichments. The Livonian Metadata Database serves as a support layer that is partly public and partly private. Its metadata definitions and descriptive metadata are exported into the LīvMDb databases as needed and permitted. The Livonian Metadata Database The Livonian Metadata Database is developed in alignment with the metadata framework of the Open Music Observatory. •Ontological and thesauri patterns: Reuses standardized or widely adopted vocabularies. •Conceptualizations and definitions: Includes concept definitions, thesauri, and other elements developed specifically for the LīvMDb. •Public permanent identifiers: Uses identifiers that are public or can be made public. The metadata layer is generally licensed under CC0, though in some cases other licenses are used (for example, CC-BY). ĹNote Examples: • The definitions of musical work,music recording and is recording of relationship allow the description of connections between an abstract musical work and its recording(s). • The ISRC code identifies recordings across all streaming platforms. Livonian Music Database (public) The Livonian Music Database is a linked open database published by the Reprex on behalf of the Open Music Observatory. The database is distributed under various Creative Commons licenses that allow both commercial and non-profit use. The primary aim of the LīvMDb is to provide a use case for very low documentation music ecosystems with very limited resources and challenging data curatorial scenarios. 119 Livonian Music Database (private) The private components of the LīvMDb are a staging area for data that has unclear provenance or legal status. Unlike in the case of LīvMDb, we hold minimal business confidential data (related to the royalty accounts of music that we published), but some data may have, for example, unclear GDPR status. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). Our work is centerred around the work of Hõimulõimed, a Finno-Ugric NGO, which curated many Finno-Ugric language collections, including the collection of Livonian-language songs available on Spotify. •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category. Our work buids on the register of Livonian placenames, because they offer the most straighforward curatorial help to find new Livonian (folk) music1. •Metadata: Relevant elements from the Livonian Metadata Database that support the use of collections or registers. The LīvMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models and data structures, enabling automatic aggregation. All datasets are available with RDF annotation and can be exported in all standard RDF serializations. Microdata is intended for institutional and professional use, not for the general public. It is annotated with standardized metadata suitable for applications such as music library cataloguing, distribution platforms, and rights management systems. 1Livonian place names: documentation, problems, and opportunities (Ernštreits 2020) and our gazetteer dataset: (Daniel Antal, Mester, and Pigozne 2025). 120 ĹNote Example The village of Mazirbe (Livonian: Irē, German: Klein-Irben, Russian: �������) is the central location of the Livonian culture, which hosts the Livonian Community House. Folk songs collected and recordings made in Mazirbe may have a provenance of Irē, Mazirbe, Klein-Irben (or its Finnish and Estonian versions), or various Cyrillic transliterations. Mainitaing a clear metadata dataset on the geographical sources of Livonian music is necessary. •Graphical view: Navigate and contribute to detailed entries on musical works, sound recordings, and related assets. � Explore on Wikibase •Semantic view: Export structured data in XML, JSON-LD, Turtle, or NTriples formats for reuse in research or digital projects. � Example Turtle file or download XML We made available the dataset in standard RDFXML, JSON-LD, and TTL serialisations accompanies by a data paper explaining its use (Daniel Antal et al. 2025; Daniel Antal, Mester, and Pigozne 2025) Statistical Data & Data Catalogue Our statistical data and catalogue consist of datasets aggregated using statistical methodologies. These comply with the SDMX standard and the W3C Data Cube vocabulary, making them compatible with spreadsheet software, statistical packages, and data science workflows in R, Python, or similar environments. Datasets are offered in multiple formats. In addition to RDF serializations, we provide standard CSV files and, when required, Excel or SPSS formats. We also publish data papers and related documentation that describe dataset usability and highlight key insights. Publications & Catalogue The LīvMDb’s most important publications are sound recordings, which are made available as a archival recordings (not ready to be communicated to the public on commercial platforms) and publicly available sound recordings. We also provide access to some printed and and hand-written scores. Microdata and statistical datasets are treated as publications and are listed both in the general catalogue and a machine-readable data catalogue. The LīvMDb also includes methodological and musicological publications, as well as data papers explaining the use of datasets. The catalogue is designed for interoperability with libraries, archives, museums, and similar institutions. 121