scieee AI-readable full text Open interactive document viewer

Two Journeys: Insights on the Annotation of Large-Scale Optical Music Recognition Datasets

Torras, Pau; Dvořáková, Martina; Badal, Carles; Herzánová Vlková, Markéta; Asbert, Gerard; Mayer, Jiří; Fornés, Alicia; Hajič Jr., Jan

Abstract

The use of Optical Music Recognition (OMR) requires annotated music scores (namely, ground truth) for training or adaptation. Besides the level of structured encoding, a significant part of the existing OMR methods still needs ground truth also at image level, which is indeed still important for tasks adjacent to OMR. However, annotating images is a major part of OMR project costs: not just the work of the annotators themselves but also associated software development and process management. Unlike most other more prevalent image annotation tasks in industry, in the case of music notation, infrastructure is lacking, and many pitfalls exist. Thus, in this paper, we report on the combined experience of two ongoing OMR projects with major data collection and annotation efforts, and we formulate recommendations for managing music notation annotation projects, leading to efficient use of the inevitably limited resources.

Full text

Two Journeys: Insights on the Annotation of Large-Scale Optical Music Recognition Datasets Pau Torras1,2[0000→0003→0327→9046],MartinaDvo!áková 5[0009→0008→8539→2993], Carles Badal1,3[0000→0001→8948→9946],MarkétaHerzánová Vlková5[0009→0004→3409→9747],GerardAsbert 1[0009→0005→3316→5463],Ji!í Mayer4[0000→0001→6503→3442],AliciaFornés 1,2[0000→0002→9692→5336],andJan Haji"Jr.4[0000→0002→9207→567X] 1Centre de Visió per Computador 2Departament de Ciències de la Computació, Universitat Autònoma de Barcelona 3Departament d’Art i Musicologia, Universitat Autònoma de Barcelona 4Institute of Formal and Applied Linguistics, Charles University 5Moravian Library Abstract. The use of Optical Music Recognition (OMR) requires annotated music scores (ground truth) for training or adaptation. Besides the level of structured encoding, a significant part of the existing OMR methods still needs ground truth also at image level, which is indeed still important for tasks adjacent to OMR. However, annotating images is a major part of OMR project costs: not just the work of the annotators themselves but also associated software development and process management. Unlike most other more prevalent image annotation tasks in industry, in the case of music notation, infrastructure is lacking, and many pitfalls exist. Thus, in this paper, we report on the combined experience of two ongoing OMR projects with major data collection and annotation e!orts, and we formulate recommendations for managing music notation annotation projects, leading to e"cient use of the inevitably limited resources. Keywords: Optical Music Recognition ·Handwritten music scores · Annotation Software ·Ground-truthing ·Music Notation Formats 1Introduction Optical Music Recognition (OMR) consists of automatically reading images of musical scores into computer-processable formats [6]. One of its main applications is aiding musicologists in the long and demanding process of transcribing apiece.Nevertheless,despitetheadvancesinthefield,fullyautomaticrecognition of historical handwritten music scores is extremely di#cult due to the high variability in handwriting styles, music notation idiosyncrasies and severe supporting medium degradation. Existing State-of-the-Art OMR methods are based on powerful deep learning architectures, but at the cost of requiring significant amounts of labelled data (ground-truth) to train and to adapt to new music Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 608 2P.Torraset al. collections. In this context, it is important to note that it is not enough to provide a scanned music sheet together with the transcription in MusicXML/MEI, because the OMR system needs to learn the position and appearance of each music element in the image. During the last decades, the OMR community has devoted significant efforts in creating annotated datasets6for music recognition, sta$removal, writer identification, etc. But, not surprisingly, the ground-truth of handwritten music scores in Common Western Music Notation (CWMN) is barely available, becoming a bottleneck in the field. Therefore, the annotation of such scores is an essential step for advancing towards robust Optical Music Recognition systems, and to e$ectively apply them in practice with confidence to handwritten scores. However, ground-truthing is a resource-intensive process: not only in terms of the e$ort required to create the ground-truth, but also because of the overhead of managing human resources and the whole annotation process. In this paper, we describe the experience with two ongoing annotation processes7,8 of OMR image data: the first one focuses on coarse annotations, whereas the second one focuses on fine-grained annotations. We describe the associated project priorities, the resulting trade-o$s, the annotation tools and workflows, and the “user experience” from both the perspective of the annotators, annotation management, and project management. Although we do not make specific recommendations, from the discussion of the advantages and disadvantages of our approaches, we provide experience from extensive projects to help others with designing their own data collection e$orts. 2 Previous Work Very few projects focus on music annotation for OMR. A significant portion of modern-day OMR e$orts are focused on older music written in square and Mensural notation, which have been the focus of large-scale projects: SIMSSA [10] and PolifonIA [16]. Previous works on this notation system resulted in the incorporation of datasets such as SEILS [23], or CAPITAN [8]. For scores in Common Western Music Notation (CWMN), large-scale transcription e$orts have been less frequent. Most modern datasets are sourced from existing typeset transcriptions of music, which are usually adapted to a specific OMR-related task following specific design goals. The DeepScores dataset [28] and its V2 successor [29] are built for music object detection. They include symbol-level detection and segmentation annotations. The GrandSta$dataset [25] sources the KernScores 9repository to build a dataset of piano-form sheet music encoded in **kern format for performing page-level end-to-end recognition. The DoReMi Dataset [26] is a collection of classical works engraved using 6https://apacha.github.io/OMR-Datasets/ 7The Project on Music scores at the CVC: https://pages.cvc.uab.es/musicscores/ 8The OmniOMR project at Charles University and Moravian Library https://ufal. m!.cuni.cz/grants/omniomr. 9https://kern.ccarh.org/ Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 609 Two Journeys: Insights on the Annotation of Large-Scale OMR Datasets 3 Dorico that incorporates transcriptions and annotations for most of the widelyadopted formats. Other common sources of scores for OMR dataset creation include the OpenScore project [11], a corpus of classic public domain music works engraved by the MuseScore community, [15] produced by Hentschel et al., alargecorpusoftranscribedscoresformusicologicalanalysisandthePDMX dataset [17], the largest collection of MusicXML files to date. For handwritten scores, dataset construction projects almost always require the intervention of human annotators. The most widely known large-scale music annotation project produced the MUSCIMA++ Dataset [13], which was built atop the existing CVC-MUSCIMA writer identification and sta$removal dataset [9]. This dataset contains 140 pages with highly-detailed segmentation annotations and their relationships encoded using a custom graph-based format. Smaller datasets include the Pau Llinàs collection from Baró et al.[2], which contains only a few hundred measure-level samples. There are also a few symbol-level datasets such as HOMUS [7], which is primarily intended for Online music recognition but can be used for o%ine OMR by rasterising pen strokes. Most of these datasets are developed using di$erent conventions and requirements and it is not surprising that they are mostly incompatible. No encoding seems to suit all of the needs of the entirety of the community. Furthermore, experience in manually annotating large-scale OMR datasets is limited. 3OMRAnnotationProcessDesign Designing good OMR image annotation pipelines is a complex endeavour. Indeed, the requirements will vary depending on the exact needs of the final user, the ergonomic requirements for the transcribers, the OMR method (or methods) the dataset is designed to accommodate and the resources available for the production of the dataset. In particular, the OMR method to be applied has a direct implication on the level of annotation detail, the amount of data to be produced and the format in which the data will be presented. Moreover, the timeline and resources available for this ground-truth creation process condition whether or not the development of specific customized annotation software is possible. Format. Most current music notation formats are designed to interface with music engraving software, which tend to prioritise semantic encodings rather than literal representations of scores. This is a sensible choice to support the musician, as it will prevent the writing of unsound scores and use their knowledge of the language to their advantage to accelerate the engraving process. Nevertheless, this poses problems for the OMR practitioner, as most widespread formats – and, by extension, music engraving tools – represent music scores in a non-explicit, non-exhaustive manner [27]. In particular, MusicXML hides many details about the layout of the score, which are abstracted within music concepts e.g. Key signatures are usually represented by the number of fifths required to get to a specific Sharp/Flat configuration, not by modelling the position of each accidental separately. In some engravers, the position of each Accidental might change with the same key. The MEI format allows a more literal style Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 610 4P.Torraset al. of transcription, but we have found it to be significantly less well-supported on WYSIWYG transcription software necessary for making the detailed encodings. Moreover, the available conversion scripts from MEI to other formats seem to be unreliable. Therefore, we have found it not possible to rely on existing representations of music alone. Custom software development. One of the key requirements of modernday OMR is incorporating symbol-level annotations [6,30]. Mainstream music engraving software does not support this feature. Thus, it is necessary to design either a collection of supporting file formats that allow annotating this information in parallel to existing ones or creating one full music representation that powers all aspects of OMR-oriented music notation. Implementing a full music notation system requires having a solution to engrave, convert and edit these scores, which is a non-trivial engineering e$ort. On the other hand, while relying on existing tools has the advantage of not requiring as much new software, coordinating them can be complex, both for the implementers and for the annotators. Targeted OMR System. There exist di$erent types of OMR systems. End-to-end models can train using only image transcriptions extracted at the line level [2] or the page level [25]. On the other hand, pipeline or graph-based methods [3] usually require an object detection or segmentation step, for which fine-grained annotations for each symbol are additionally required for training. Level of annotation detail. Annotating music primitives amounts to hundreds or thousands of annotations per page, which is very costly. A possible solution is to use loosely annotated symbols in order to reduce this manual labor burden. The solutions presented in this work explore two options in this regard: one produces fewer, precisely annotated samples and the other produces more, yet less precisely annotated images. These are the most important decisions, but the list is far from exhaustive. Depending on the need, the expected OMR method to be applied, and available resources, one should opt for one of these choices. Next, we describe two di$erent annotation processes and discuss their advantages and disadvantages. 4Process1:CoarseAnnotations The first annotation process consists of a coarse annotation of the music scores. Concretely, the objective is to provide a bounding-box for each music element, together with its corresponding MusicXML category. The priority is to speed-up the process, labelling more pages with less human resources. 4.1 Annotation Pipeline 1 The coarse pipeline has two steps. The first step consists in the transcription of the music score using MuseScore [1], obtaining a MusicXML as output. The second step consists in the alignment of the MusicXML with the image, providing the bounding-boxes of each element. In this step, we have specifically designed Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 611 Two Journeys: Insights on the Annotation of Large-Scale OMR Datasets 5 an Android application for a Tablet with stylus pen. Figure 1 shows a summary of this first annotation pipeline. These steps are described next. Fig. 1. The coarse annotation pipeline. In green, steps that are assisted by human transcribers. In blue, automated steps. Transcription The process begins by assigning a score scan to a musicologist, who transcribes it using MuseScore 4.4 [1]. This transcription is performed as literally as possible – akin to an urtext edition. Each line of music is stored in a separate transcription file, with line being understood as any set of staves that are written to be played in unison or that belong to the same instrument,asseen in particellas, orchestral works or pianoform music. For instance, in orchestral works, all instruments that belong to the same line are written into the same file using di$erent parts. On the other hand, for single-instrument works, each of the lines on the same page are annotated in a di$erent file. These files are uniquely identified by the vertical position of the transcribed material. Key semantic information such as the part’s Clef, Time Signature and Key are always inserted at the beginning of each line in order to guarantee that the original full score can be reconstructed, but in those cases where they are not explicitly written in the score they are set to be hidden. Once the MuseScore files for all lines of a page are obtained, they are converted to MusicXML automatically using MuseScore’s batch conversion system. This line-based transcription process is chosen to guarantee that the layout of the original score is respected when processing the material using any MusicXML engraver – even those that do not fully support the format’s layout features – and to facilitate some downstream alignment steps. We have chosen to distribute MusicXML as the base format for our work because of better tooling support and because of its more extended use. However, in order to enable the future possibility of generating MEI or even MNX, we have chosen to also distribute the original MuseScore files with the dataset once it is ready. Given how the alignment pipeline works, we anticipate that adding such support alignment for MEI down the line would be quite straight-forward. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 612 6P.Torraset al. MusicXML-to-Image Alignment In order to align each musical symbol in the MusicXML with its representation in the image, a second intervention by the musicologist is required. For this task, we have designed an application that runs on an Android tablet with a stylus pen, since the labelling process is significantly faster and more accurate than using a computer with a mouse input. In order to perform the symbol-level alignment, the music objects need to be given some form of identification to assign them a bounding box. This is fairly trivial for those objects that MusicXML models explicitly and provides an id attribute for – e.g. Clefs or Barlines. We generate procedural IDs automatically for these objects after the transcription process. However, there are many objects that either do not have this id attribute or whose MusicXML definitions map to multiple-object compounds. We identify these objects using the id of their parent object and some form of disambiguation name for that specific element. Each symbol is related with their graphical representation in the score by leveraging the hierarchical primitive representation produced by the Verovio [24] engraving system. We have authored a fork 10 which inserts our generated identifiers in the output SVG, enabling the relationships between MusicXML and score representations. For those cases where Verovio’s primitive definitions do not match ours or the identifiers cannot be easily inserted within the engraving, a post-processing script inserts the missing information after-the-fact on the SVG directly. Once this decorated SVG file is produced, musicologists are tasked to label each of the desired symbols by drawing a rough circle around them using a stylus pen on an Android tablet. A screenshot of the main alignment screen of the Android app is shown in Figure 2. The application uses the rendered score’s SVG class information to select the symbols of interest, changes their colour and the user selects its matching symbol in the real image. Then, using the symbol’s ID stored in the SVG the final output alignment is generated. Only the final file containing the relationships between object bounding boxes and identifiers is necessary, which we have encoded as a slightly modified COCO format file (string-based IDs are used instead of integers). With this identifier, the XPath to the original MusicXML nodes can be constructed. 4.2 User Experience: Annotation Process 1 One senior musicologist has managed the workload of the eight musicologists that have worked in this coarse annotation process, obtaining more than 1.300 pages in less than one year. Concerning quality control, each page has been fully transcribed by one annotator and revised by another. The time required by an expert user to transcribe a handwritten music page into MuseScore in the style demanded by this method varies depending on several factors such as the number of staves per system or the variety of instruments involved. On average, transcription takes approximately 50 minutes per page, with a typical range between 30 and 90 minutes. In contrast, alignment usually 10 https://github.com/ptorras/verovio.git, ident branch. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 613 Two Journeys: Insights on the Annotation of Large-Scale OMR Datasets 7 Fig. 2. Screenshot of the Alignment Android application. At the top, the real handwritten document being annotated. At the bottom, its transcription. In green, objects that have already been annotated. In yellow, objects that have been skipped. In this example, the time signature is not present in the original document and thus it has not been annotated. takes less time, averaging around 30 minutes per page, although for pages containing large numbers of systems (ranging from 6 to 12 staves), alignment time can sometimes double. The alignment application in the Android tablet o$ers fast performance and responds promptly to input from the digital stylus pen, contributing to an efficient and precise user experience. Its interface is intuitive and user-friendly, facilitating navigation and task execution without requiring extensive prior training. The app facilitates alignment in multiple sessions and correcting previously aligned segments via discretionary auto-saving. Some downsides of the design of the application are that neither the image display nor the Verovio transcription automatically adjust to the position of the cursor. This shortcoming hinders precise task tracking, especially on pages with a high number of systems, where it is easy to lose visual reference of the current working point (hence the increase in alignment time for multi-sta$scores). This is a result of having a horizontally-optimised view for the workspace. A key limitation of this workflow is error correction. If transcription errors are identified during the alignment phase, the process must be paused, and the corresponding line must be revised in MuseScore before the alignment can continue. It is important to note that if a correction in MuseScore involves the addition or deletion of elements, misalignment will occur in the app from that point forward. This happens because Verovio identifiers shift to account for newly added elements or the removal of existing ones. Consequently, the alignment of the a$ected line must be temporarily halted. In contrast, when the correction consists only of a substitution (e.g., replacing one note with another), the identifiers Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 614 8P.Torraset al. remain unchanged. In such cases, the alignment can proceed as normal once the updated SVG file is loaded, since the alignment remains una$ected. Transcription errors are relatively frequent, with approximately 21% of lines containing at least one error on their first pass. Most errors are detected during the reviewing phase, but some can remain. This rate can be reduced through multiple rounds of revision, at the cost of more human resources. Nevertheless, it should be acknowledged that due to the double conversion of the original MuseScore (.mscz) files—first to MusicXML and then to SVG via Verovio—errors may still arise during the alignment stage due to format limitations. For this reason, it is essential that error correction remains possible in this final phase. 5 Process 2: Fine-grained Annotation The objective of this process is to obtain a dataset of 1000 pages with finegrained, very accurate annotation. This can then serve both as a comprehensive di#cult benchmark representative of the variety of scores found in a large library collection, and as training data for systems targeting such a domain. It is also intended to serve as a seed dataset for large-scale realistic music notation synthesis [22]: with a su#ciently broad collection of annotated scores, many di$erent outof-domain conditions for new collections can be simulated. The dataset should also enable comparing the performance of end-to-end and object detection based OMR pipelines on a real-world collection. For this reason, the dataset is intended to contain two layers: the Music Notation Graph (MuNG) that annotates the score image itself [13], and in parallel its MusicXML representation. 5.1 Annotation Tools 2 Obtaining the MusicXML layer of data was unproblematic: we simply use MuseScore [1]. We adapt transcription rules used by the OpenScore Lieder project [12]. Each page was transcribed by one annotator and revised by another. This process was straightforward, with average speeds under 2 hours per page, including the complete revision. The annotation toolchain for MuNG has proven much more troublesome. The MuNG format was originally designed and implemented in the MUSCIMA++ dataset [13] and had an annotation tool: MUSCIMarker [14]. However, due to lack of maintenance, MUSCIMarker has become unusable, as the features of the Python GUI Framework kivy it relies on have since been deprecated. Additionally, the all-Python implementation would have been di#cult to transfer into the online environment, which we identified as necessary for process management. Meanwhile, many commercial tools for image annotation have become available. Thus, faced with the choice between rewriting an old tool and using an existing interface with workflow management implemented, a Python SDK, and a su#ciently permissive academic licence, we decided to rather allocate software development resources to actual recognition rather than annotation tooling, and started using Labelbox for MuNG annotation. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 615 Two Journeys: Insights on the Annotation of Large-Scale OMR Datasets 9 Fig. 3. The workflow of fine-grained annotation process. Two separate paths for each ground truth type: MuNG (upper) and MusicXML (lower). Because the manual steps are done in one step separately for each ground truth layer, less automation was needed (though some symbol disambiguation is needed in the conversion from Labelbox outputs to MuNG format). Manual steps in green, automated in blue. Note that completing the MuNG path took approx. 5x longer than MusicXML path. Labelbox in principle did support this process. The only adjustment compared to MUSCIMA++ was that while the Labelbox annotation interface for images does support natively adding links between objects, this is only supported for polygons and polylines that have defined control points. We therefore defined the ontology to use polygons instead of pixel masks for all symbols except sta%ines and thin barlines, which were annotated as polylines. The complete ontology has 62 di$erent elements, of which 3 are edges (syntax, precedence, and sta$placement). Seven annotators worked on the process. All had significant musical experience: professionals or students, primarily from the early music community due to a greater familiarity with understanding musical manuscripts (a large portion of the annotated collection comes from the 17th and 18th centuries). Each image was fully annotated by one annotator, and a di$erent annotator would review it. Two annotation managers shared the workload of making sure the process ran smoothly. GitHub discussions were used to build up institutional knowledge, both on MuNG annotation policies and on optimal use of Labelbox. It took approx. 5 months to get to a point where the policies for both were largely complete and only rare situations needed discussion as they came up.11 Labelbox makes the annotations available via an API in its own JSON format. We convert these to MuNG format with a script. Compared to MUSCIMA++, we made several changes to the annotation policy to make the manual process easier: this mostly concerned merging several names into one and disambiguating them automatically afterwards. Most importantly, we only use one flag class instead of splitting it into flag8thUp,flag8thDown,etc.,because Labelbox had poor support for keyboard shortcuts to select an object class. 11 These keep coming up at roughly 1-2 per month. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 616