scieee AI-readable full text Open interactive document viewer

Pitch Spelling Jazz Lead Sheets and Solo Transcriptions

Bouquillard, Augustin; Jacquemard, Florent

Abstract

We present an algorithm for pitch spelling tailored for written jazz music. Receiving some input in a MIDI-like format, including information about note heights (expressed in semitones from a reference lowest note) and boundaries of bars (measures), it estimates appropriate note names, one global key signature, and one local scale for each bar. These related pieces of information are jointly assessed in two optimisation steps. In a first ""modal"" step, one likely scale is guessed for each bar, by minimising the number of accidentals that shall be printed in the engraved score, in a best-path search. Then, in a second ""tonal"" step, these local scales are used for estimating the key signature that would give the best note spelling on the whole piece. We report successful experiments on a set of lead sheets from the Real Book as well as transcriptions of Jazz solo recordings and basslines. Our procedure is originally designed for an application to music transcription, in particular the construction of digital collections of written jazz soli from audio recordings, in the context of musical analysis, teaching, and cultural heritage preservation. This method should also be useful in other tasks related to music notation processing. Moreover, we have defined for its purpose new distances between various common jazz scales, which might be of some interest in musicological studies.

Full text

Pitch Spelling Jazz Lead Sheets and Solo Transcriptions Augustin Bouquillard1[0009000303713196], and Florent Jacquemard2[0000000322697550] 1École polytechnique, Palaiseau, France [email protected] 2INRIA, CNAM/Cedric, Paris, France [email protected] Abstract. We present an algorithm for pitch spelling tailored for written jazz music. Receiving some input in a MIDI-like format, including information about note heights (expressed in semitones from a reference lowest note) and boundaries of bars (measures), it estimates appropriate note names, one global key signature, and one local scale for each bar. These related pieces of information are jointly assessed in two optimisation steps. In a first "modal" step, one likely scale is guessed for each bar, by minimising the number of accidentals that shall be printed in the engraved score, in a best-path search. Then, in a second "tonal" step, these local scales are used for estimating the key signature that would give the best note spelling on the whole piece. We report successful experiments on a set of lead sheets from the Real Book as well as transcriptions of Jazz solo recordings and basslines. Our procedure is originally designed for an application to music transcription, in particular the construction of digital collections of written jazz soli from audio recordings, in the context of musical analysis, teaching, and cultural heritage preservation. This method should also be useful in other tasks related to music notation processing. Moreover, we have defined for its purpose new distances between various common jazz scales, which might be of some interest in musicological studies. Keywords: Music information retrieval ·Music transcription ·Common music notation ·Pitch spelling. 1Introduction The preservation of outstanding jazz soli in sheet music has been a motivating goal of Automated Music Transcription (AMT) since the problem was first introduced. Pitch Spelling (PS) is an important subtask of AMT, whose purpose is to All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 334 A. Bouquillard and F. Jacquemard choose appropriate note names to denote, in Common Western Music Notation (CMN), pitch values initially expressed in an absolute number of semitones or equivalently, as key numbers on the piano keyboard (MIDI key). For a given note pitch, several possible notations exist in CMN, like for instance G\and AZ, and the choice of a particular note name is crucial for several reasons. Firstly, for the purpose of readability, by reducing the number of accidental symbols printed on the score: one global Key Signature (KS) is fixed for the whole piece (or a part of it), specifying 0 to 7 sharps or flats which are only printed at the very beginning of each line and apply by default to all bars of the line. Secondly, the presence of printed accidentals, not in the global KS, is a cue of the tonal function of the notes, and an indication of local modulations. The scope of an accidental is always one bar, which is not only a typical level of granularity for such tonal changes but also close to the harmonic rhythm, with generally one or two chords indicated per bar. Modulations are indeed tightly linked to harmonic context. Specific PS issues arise when transcribing jazz music, as correlation between accidentals and modulation is greatly reduced in improvised soli. Frequent modal incursions, chromatic movements and melodic lines rich in foreign notes as well as a varying degree of freedom from the underlying chord progression make the spelling task more difficult. State of the Art. Several PS algorithms have been proposed in the literature. Many of them have been designed according to musicological criteria, such as the analysis of voice-leading, interval relationships and local keys [6,16,19], or aprincipleofparsimony(minimisationofthenumberofaccidentals)[4,5],or to some relevant intermediate data structures, such as the Euler lattice [13] or weighted oriented graphs [22], in order to reduce PS to optimisation problems. Some other approaches to PS are based on training statistical models such as HMM [20] or RNN [10], with datasets of music scores. The last cited system, PKspell [10], has obtained state-of-the-art results on an iconic benchmark called Musedata, proposed by D. Meredith [16], made of 216 works by Baroque and Classical composers. In the former work [4], we proposed an algorithm for the purpose of PS and KS estimation in a strict tonal framework, which has obtained better results than PKspell [10] on the classical piano dataset ASAP [11] used for the training of the latter system (98.2% correct PS and 95.58% for KS estimation in [4], vs 96.50% for PS and 90.30% for KS estimation in [10]). To our knowledge, the applicability of the different approaches to jazz music, for example by trying to take into account the use of jazz modes, has not been studied so far. Contributions. In this paper, we present a PS algorithm aimed at processing jazz music, in particular transcriptions of improvised soli, and its evaluation on three jazz datasets of various kinds. Its principle is to guess, from the pitches, the accidentals that shall effectively be written in the engraved score according to the notational conventions, jointly with global key signature information as well as local tonality or rather local scale determination, since jazz musicians do not tend to stay inside a strict tonal framework. The somewhat naive starting idea is to first minimise the number of accidentals printed on the whole score, which Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 335 Pitch Spelling Jazz Lead Sheets and Solo Transcriptions leads to several possible key signatures. Then we determine the best spelling options by deriving for each bar the most plausible local scale. Depending on the user-specified version of the algorithm, the latter is either chosen only among tonalities, or not only tonalities but also from the seven diatonic modes as well as the minor and major Blues scales. Best-path search procedures are used at bar level to calculate the solution of minimal cost at each step. Our algorithm brings significant improvements to the former work [4] . Beyond extensive architectural and under-the-hood changes, the transition from a strictly tonal framework to handling jazz data required a substantial increase in the number of supported scales - from 30 to 165. To accommodate this expansion, we propose a generalisation of the Weber distance (a measure originally conceived for fully tonal contexts) extending it to encompass 11 modes across 165 distinct scales. This extension not only supports the broader melodic and harmonic vocabulary found in jazz but also constitutes a standalone contribution with potential relevance in studies oriented towards computational musicology. 2NamesandScales We first describe the problem studied, recalling basic notions and we then introduce a new distance between scales which is one key component of our method. 2.1 Problem Input Let us assume given in input a sequence of notes ⌫1,...,⌫ p,calledapart. It shall typically represent one staffin music notation, possibly including several voices and chords. Every note ⌫iin the input sequence is defined by: (inpi)a MIDI pitch value in 0..128, (insim)a boolean flag expressing whether ⌫iand the next note ⌫i+1 are played simultaneously (inbar)a boolean flag expressing whether ⌫ibelongs to the same bar as ⌫i+1. By convention, the two flags are set to false for the last note ⌫p. The MIDI pitch of a note ⌫corresponds to the distance in semitones from a reference lowest note. Its value modulo 12, called pitch class,isdenotedbypc(⌫).Wedonotassume the onset time nor duration of input notes to be given,however, we assume that they are enumerated by increasing onset. Moreover, we require information on the bar boundaries and note simultaneity. We call two notes simultaneous when they occur on the same onset, because they are involved in the same chord, or notes in different voices and starting simultaneously. However, a grace-note ⌫,be it single or involved in an ornament (appoggiatura, gruppetto, mordent, trill etc) is not considered simultaneous with the next note ⌫0,butprecedingit,although in a score, ⌫and ⌫0have theoretically the same onset. The above assumptions requires the time values to be quantised, i.e., expressible in a music score, in fractions of beats or bars. Our procedure is therefore especially relevant as a backend task in a music transcription framework. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 336 A. Bouquillard and F. Jacquemard 2.2 Problem Output Given input notes in the above form, we are expecting the following outcome: (outks)one estimated global key signature, (outspell)aspelling for each note, (outloc)one estimated local scale for each bar. The spelling of a note is made of: a name in A..G,asymbolofaccidental,amongst ^,Z,[,\,],andanoctave number in 2..9.Everynotenameisassociateda unique pitch class: 0 for Cup to 11 for B,andtheaccidentalsymbolactsasapitch class modifier: 2for [up to 2for ](the ^may be omitted). With the convention that the MIDI pitch 0has spelling C1,everynotespellingcanbeassociateda unique MIDI pitch value. pc 0D[CB \ 1DZC\B] 2E[DC ] 3F[EZD\ 4FZED ] 5G[FE \ 6GZF\E] 7A[GF ] 8AZG\ 9B[AG ] 10 C[BZA\ 11 CZBA ] Fig. 1. Enharmonic spellings for each pitch class. The opposite is not true: there exists several (2 or 3) alternative valid spellings for every MIDI value, summarised in Figure 1 for the 12 pitch classes. For instance, B\2,C^1 and D[1are alternative spellings for the MIDI pitch 0. A key signature (KS) is denoted by an integer kbetween37 and 7, which indicates that by default, |k|note names shall be altered by a \, when kis positive, or by a Z, when kis negative. The names of the notes altered are defined according to the order of fifths: F\,C\,G\,D\,A\,E\,B\for k>0,and BZ,EZ,AZ,DZ,GZ,CZ,FZ,for k<0. Amode is a sequence of intervals. In this work, we consider the following 9 heptatonic modes : ionian (or major mode), dorian, phrygian, lydian, mixolydian, aeolian (or natural minor), melodic minor, harmonic minor and locrian, as well as the major and minor blues modes. We call scale the pairing of a KS and a mode, functionally identical to the usual definition of a mode anchored to a tonic. In a tonal context, "scale" and "key" are synonymous concepts. Some scales can induce accidentals outside the KS, that we call here characteristic accidentals. It is for example the case of the leading tone at the seventh degree of harmonic minor scales. Intuitively, these accidentals, when printed, enable to recognise the scale at sight. 2.3 Distance between Scales Gottfried Weber defines in [21] a measure of distance between major and harmonic minor scales, which has been used for MIR tasks related to key estimation [9]. The Weber distance between two keys is the length of a shortest path between them in a 2D grid. Each node in the grid represents a specific key K with 4 neighbors that are considered close to K,eitherbecausetheydifferfrom Kby only one note (dominant key, subdominant key, relative key) or because they have the same tonic (homonym keys: same tonic but different mode). 3We do not consider KS 8or 8as they are very rarely found. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 337 Pitch Spelling Jazz Lead Sheets and Solo Transcriptions Cdo C`yAdo Fdo F`yDdo Bdo ZB`y ZGdo Blues Gae Gio Eae Cae Cio Aae Fae Fio Dae Dph Dmx Bph Gph Gmx Eph Cph Cmx Aph A`oAdo F\`o D`oDdo B`o G`oGdo E`o Fig. 2. Weber distance generalised to common jazz modes. We propose in this work an extension of Weber’s distance to the seven modern diatonic modes (ionian, dorian, etc), melodic and harmonic minor modes, and two blues modes, used for the estimation of local scales in the step described in Section 3.2 of the PS algorithm. We use a 3D structure represented in Figure 2. The modes are grouped by pairs, following the relationship between a major (i.e., ionian) scale and its minor (i.e., aeolian = natural minor) relative. It means in particular that a descending minor third always separates the first note of the "major-like" scale from the first note of its "minor-like" counterpart; for example, Flydian has Ddorian as its "modal relative". From every pair of "relative" modes we derive a new 2D grid similar to the original grid of [21]. All of the grids are aligned such that scales with a given nature of third (major or minor) and containing exactly the same notes are always placed on the same line; for example, the Flydian scale from the lydian-dorian grid, with a major third, is behind Cionian from the ionian-aeolian grid which is itself behind Gmixolydian from the mixolydian-phrygian grid. We add to this 3D structure the two other minor modes (melodic and harmonic), in the same places as their aeolian counterparts (with identical tonics), and finally the major and minor blues scales. The latter constitute a supplementary 2D grid placed between the lydian-dorian and ionian-aeolian grids, since the lydian distinctive augmented fourth is also present in the minor blues mode and the dorian mode is almost entirely contained in the major as well as in the minor blues mode. We assign a cost of 1 to any move in a straight line from one 2D grid to an immediate neighbouring 2D grid, and to any horizontal or vertical move within a single 2D grid. Finally, the distance between two scales is the minimum number of moves in the 3D grid to go from one to the other. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 338 A. Bouquillard and F. Jacquemard 3 Pitch Spelling Algorithm We present a method addressing the problem presented in Section 2, based on counting the accidentals that would be printed in a score following notational conventions recalled in 3.1 and using our extended distance (Section 2.3) between scales to refine its results. The possible choices are exhaustively explored through dynamic programming techniques. Our approach, guided by common principles of musical notation and writing, works in two steps: the first step (Section 3.2) evaluates one likely scale for each bar, using the characteristic accidentals in different spellings. The second step (Section 3.3) uses the estimated local scales in order to refine the selection of KS and spellings. The local scales are essentially a byproduct of our algorithm, used in an intermediate state to estimate the best spellings. 3.1 Conventional Spelling and Shortest Path For readability reasons, some accidentals are not printed in engraved scores. Following a principle of parsimony, the notational conventions [12] are roughly as follows: accidentals already in the key signature are omitted by default, and other accidentals need not be repeated in the same bar. There is an additional restriction to this rule, which we will treat as an option in the following: (optoct)An accidental applies only to the pitch at which it is written: each additional octave for the same pitch class requires a further accidental [12]. To ensure the principle described above, we consider a state, which is a mapping of note names (and octaves, in the case of optoct)intoaccidentalsymbols.For simplicity, we describe the case where (optoct)is disabled. Starting from a state , when processing a note ⌫with pitch class p,thereareuptothreepossible transitions to a new state 0,correspondingtothepossiblenameeand accidental afor pin Figure 1. If (e)=a,then0=,andtheaccidentalais not printed for ⌫, otherwise, 0(e)=a6=(e)(0is identical to for the other names), and ais printed. Given a scale S, each bar begins with an initial state 0,definedaccordingto one of the following options: (optunld)0contains exactly the accidentals defined by the KS of S, (optlead)0contains the accidentals of the KS of Sand the characteristic accidentals of S. All notes in the bar are then processed to determine which accidentals are printed. By assigning an integral cost value to each transition, depending on printed accidentals, we can compute an optimal path for each bar and scale using a Viterbi algorithm [14] tagging each state with the cumulated cost of the best path from 0into . The time complexity is linear in the number of states plus the number of transitions. In the worst case, the number of states can be exponential in the number of notes in the bar. However, we can prune Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 339 Pitch Spelling Jazz Lead Sheets and Solo Transcriptions unnecessary search branches by building the states and transitions on-the-fly, which keeps computation time reasonable in practice. For additional efficiency, we impose the following restriction: (ressim)two simultaneous notes (in the sense of insim)inthesamepitchclass must have the same name. It is ensured by adding to the states an additional mapping of {0,...,11}into note names. 3.2 Modal Step From now on, assume we are given pinput notes to be spelled, distributed across mbars, and npossible scales denoted S1,...,S n.Weconstructann⇥mtable G, called the grid, where each entry G[i, j]assigns an estimated scale to bar j, assuming the global starting scale is Si. This involves two substeps. Substep 1 is the construction of a n⇥mtable T, where T[i, j]is the cumulative cost of a best-path computed with the algorithm of Section 3.1 under the following conditions: the initial state 0is defined by (optlead), without the option (optoct),and,foreachtransition,acostvalueof0ifthecorrespondingaccidental ais not printed, 1if a2{ Z,^,\},or2if a2{ [,]}. Thus, T[i, j]is roughly the number of accidentals printed in bar junder global scale Si.Whendifferent best-cost paths arise, tie-breakers are used (details omitted for brevity). Substep 2 is the construction of the grid G, where G[i, j]gives the estimated local scale for bar j,assumingthestartingscaleisSi.For1inand 1jm,letrkT(i, j)be the rank of T[i, j]in the jth column of T,i.e.,the value ksuch that global scale Sigives the kth best cost in Ton bar j.Given i2[1..n]the index of a hypothetical starting scale, let us consider the vector: arg min i1,...,im2[1..n]0 @ m X j=1 rkT(ij,j)+ m X j=1 rkd(Si,S ij)+ m1 X j=1 rkd(Sij,S ij+1 )1 A where rkd(S, S0)is the rank of the distance d(S, S0),asdefinedinSection2.3, among hd(S, S1),...,d(S, Sn)i. This yields for any line ithe sequence hi1,...,i mi of local scale indices minimising a combination of spelling cost (first term in the sum), distance to the starting or "contextual" scale Si(second term), and modulation cost (third term). The sequence is computed using a shortest-path algorithm similar to that of Section 3.1. Finally, we set G[i, j]=Sij. 3.3 Tonal Step The second and final step estimates a global KS (outks) and a note spelling with this KS (outspell), using the local scales computed in grid G. To this end, we construct a new n⇥mtable P,similarlytoT,butthistime, with the option (optoct)enabled. Additionally, transition costs for computing P[i, j]now account for the number of accidentals not in the estimated local scale G[i, j]. Intuitively, the spelling in P[i, j]depends not only on the KS of Si, but also on how well it fits the local scale G[i, j]. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 340 A. Bouquillard and F. Jacquemard More precisely, let distance be the number of accidentals in the transition that are not in G[i, j]and let accid be the transition cost used for the construction of T[i, j](Section 3.2, Substep 1). The transition cost is defined as alexicographicallyorderedtriple:(i)accid +distance,(ii)distance,(iii)tie breakers as in Section 3.2. The estimated global KS (outks)isthekeysignatureKiof Siwhere row i of Pis the one with the smallest cumulative cost. The chosen spelling (outspell) corresponds to the minimal-cost path in that same row. 3.4 Rewriting Passing Notes After spelling is chosen, we apply local corrections by rewriting passing notes, following the rules from D. Meredith’s PS13 pitch-spelling algorithm [16], step 2. Each rule applies to a trigram of notes ⌫0,⌫ 1,⌫ 2,separatedby1or2semitones, in ascending, descending, or broderie patterns. If ⌫0and ⌫1(case `), or ⌫1and ⌫2(case r), share the same name, then ⌫1is rewritten. For example, in a broderie-up pattern, the rule CC \C!CD ZCapplies. These rules follow classical voice-leading principles, we will discuss their relevance in a jazz context in the next section. 4Evaluation 4.1 Implementation The algorithm of Section 3 was implemented4in C++20 (17k loc). This language was chosen for efficiency and integration into larger systems, in particular those designed for transcription, where quantised timings (especially bar boundaries) are computed before pitch spelling. A Python binding, based on pybind11,offers calls (in Python) to the C++ methods, and was used for evaluation. For evaluation, we used the Music21 toolkit [7], to parse the ground-truth MusicXML score files in the evaluation datasets, extract the required note information (Section 2.1), and compare the estimated spellings with those in the original scores. Evaluation feedback is provided as tables (one row per opus) as well as output XML scores annotated with color-coded spelling differences, original spellings written under the staffwhen errors occur, and the estimated local scales and global KS (in grey)5.ExamplefeedbackcanbefoundinFigure3. 4.2 Datasets We conducted an evaluation of our algorithm on three jazz datasets, based on the spelling in the original reference scores cited below (without annotations). 4See https://gitlab.inria.fr/pse/pse (branch beta) for the C++ code and the Python evaluation scripts. 5The complete outputs of our evaluations on the 3 datasets are available at https: //github.com/florento/PSjazzEval. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 341 Pitch Spelling Jazz Lead Sheets and Solo Transcriptions The first dataset comprises 200 lead sheets from the Real Book [1], in MusicXML format. Some were digitised by us, others were sourced from MuseScore, and have been manually curated to conform to the reference edition [1], in both spellings and chord symbols. Lead sheets are one page long on average, for a total of 6000 bars and 21000 notes in the whole dataset. In addition to the above lead sheets, we considered a digitised version of [8], in musicXML format, consisting of 50 transcriptions of complex tenor sax soli from the Charlie Parker Omnibook [2]6.Allscoresinthedatasetaretransposed for C instruments. Some specificities of the scores in this dataset, regarding in particular note spelling, are discussed below in section 4.3. The dataset contains atotalof3640barsand22700notes. Finally, as a third case study from the same repertoire, we consider the dataset FiloBass [17] made of 48 MusicXML verified transcriptions of basslines of jazz standards, whose backing tracks are digitised from the Aebersold series [3], for a total of 12500 bars and 53000 notes. 4.3 Evaluation Options and Ablation Tests We evaluated using the options defined in Section 3. Results are reported in Table 1. We vary the number of candidate scales for spelling (called S1,...,S n in Section 3): 30 refers to major and harmonic minor modes for all KS, whereas 165 covers all the above, plus 6 other diatonic modes, melodic minor mode, and minor and major blues (see Section 2.2). We evaluate with and without the passing note rewrite rules of Section 3.4 (post-processing). All datasets include one or several Chord Symbols (CS) per bar in standard jazz notation [15]. We offer 2 options regarding these CS during spelling. The first one, denoted by ⇥in Table 1, is simply to ignore them. The second option consists in extracting CS notes (via Music21 [7]) and constraining their names during spelling. In this "force" mode, at CS positions, only the transitions that match the CS note names are allowed in the shortest-path search (Section 3.1). This last option makes sense in contexts such as automatic transcription of jazz soli, when the lead sheet is a standard whose CS are known in advance. Editorial choices. In the Omnibook dataset [18], every opus uses a KS of 0, although the true global tonality is often not C major. We allow the option to force a global KS for outks at the step in Section 3.3, but we did not use this option in our evaluations. Moreover, this dataset does not contain any B\,CZ,E\, or FZ,noranydoubleaccidental[or ]. These can only be considered as editorial choices. Therefore, we offer the option to disable such spellings in our algorithm by removing the corresponding transitions (as in Figure 1) in the shortest-path search (Section 3.1). This editorial constraint is only applied to Omnibook as an option in our evaluations, as shown in the second results line for that dataset in Table 1. It is not applied to the other datasets, which do contain such spellings. 6We corrected a few spelling discrepancies between [8] and the original [2]. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 342