Proceedings of the 26th International Society for Music Information Retrieval Conference
Abstract
Full proceedings of ISMIR 2025, held in Daejeon, South Korea, September 21-25, 2025. Contains 99 peer-reviewed papers.
Full text
International
Society f or
Music Inf ormation
Retrie v al
Conf erence
ISMIR 2025
September 21–25, 2025
Daejeon, K orea
and Online
Pr oceedings
ISMIR 2025 was or ganized by the K orea Adv anced Institute of Science and T echnology (KAIST), Sogang Uni versity , the
K orean Society for Music Informatics, the International Society for Music Information Retriev al, and a di verse interna-
tional committee of or ganizers.
W ebsite: https://ismir2025.ismir.net
Confer ence theme: Harmon y of T radition and Modernity
Edited by:
Juhan Nam (KAIST , South K or ea)
Dasaem Jeong (Sogang Univer sity , South K or ea)
K eunwoo Choi (Gaudio Lab / KAIST , South K or ea)
Li Su (Academia Sinica, T aiwan)
Magdalena Fuentes (NYU , USA)
T omoyasu Nakano (AIST , J apan)
Xiao Hu (University of Arizona, USA)
Hao-W en (Herman) Dong (University of Michigan, USA)
ISBN: 978-1-7327299-5-7
T itle: Proceedings of the 26th International Society for Music Information Retrie val Conference, Daejeon, K orea, Septem-
ber 21–25, 2025.
Permission to make digital or hard copies of all or part of this w ork for personal or classroom use is granted without fee,
provided that copies are not made or distrib uted for profit or commercial advantage, and that copies bear this notice and
the full citation on the first page.
© 2025 International Society for Music Information Retrie val
Sponsor s
W e would like to e xpress our sincere gratitude to our generous sponsors whose support made ISMIR 2025 possible.
Platinum Sponsor s
Silver Sponsor s
Br onz e Sponsor s
WIMIR Sponsor s
W idening Inclusion in Music Information Retrieval
Patr on
Contrib utor
Supporter
Local Suppor t
W e gratefully acknowledge the local financial support of:
Or ganizing Institutions
The conference was supported by the Ministry of Education of the Republic of K orea and the National Research Founda-
tion of K orea (NRF-2024S1A5C3A03046168).
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Or ganizing Committee
General Chair s
Juhan Nam, KAIST , K orea
Dasaem Jeong, Sogang Uni versity , K orea
K eunwoo Choi, Gaudio Lab / KAIST , K orea
Scientific Pr ogram Chairs
Li Su, Academia Sinica, T aiwan
Xiao Hu, Uni versity of Arizona, USA
Magdalena Fuentes, NYU, USA
T omoyasu Nakano, AIST , Japan
T utorial Chair s
Hyung-Seok Choi, Ele venLabs, USA
Christof W eiß, Univ ersity of Würzbur g, Germany
Publication Chair
Hao-W en (Herman) Dong, Univ ersity of Michigan, USA
LBD Chair s
K osetsu Tsukuda, AIST , Japan
Y un-Ning (Amy) Hung, Moises AI, USA
Music Chair
Harin Lee, Max Planck Institute, Germany
Gabriel Meseguer Brocal, Deezer , France
Industry Chairs
Jaehun Kim, Pandora / SiriusXM, USA
Akira Maezaw a, Y amaha, Japan
DEI Chair s
K yung Myun Lee, KAIST , K orea
T aegyun Kwon, KAIST , K orea
iii
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Sponsor ship Chair
Minz W on, Suno, USA
Vir tual Chairs
Jeong Choi, LINE Music, Japan
Y u W ang, Spotify , USA
Chin-Y un Y u, Queen Mary Uni versity of London, UK
Grant Chair s
Seokjin Lee, K yungpook National Univ ersity , K orea
Alia Morsi, Uni versitat Pompeu F abra, Spain
Ne wcomer Initiative Chair
Seungheon Doh, KAIST , K orea
W eb Chair / Designer
Joonhyung Bae, KAIST , K orea
Local Or ganization Chairs
Jiyun Park, KAIST , K orea
Eunjin Choi, KAIST , K orea
Sein Lee, KAIST , K orea
Jongsoo Kim, KAIST , K orea
Hounsu Kim, KAIST , K orea
Hayeon Bang, KAIST , K orea
Danbinaerin Han, KAIST , K orea
V olunteer Chair s
Minsuk Choi, KAIST , K orea
Kirak Kim, KAIST , K orea
Social Media Chair
Jongmin Jung, Neutune / Sogang Uni versity , K orea
Unconference Chair
Geof froy Peeters, Télécom Paris / IP-P aris, France
i v
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
V olunteer s
Eekgyun Ahn
Luisa Lopes Carv alhaes
Hyeyoon Cho
Y erim Gim
Robie Gonzales
Dongyub Han
Jaesun Joung
Jaekwon Im
Jaeran Choi
Jaehyun On
Jiwoo Ryu
Jooeun Lim
Joohye Son
Kihong Kim
Daeyong Kw on
Y ujin Kim
Meilin L yu
Michaella Jung-Hyun Moon
Zheng Xun Ng
Beomjin Park
Christos Plachouras
Thiago Martin Poppe
Pedro Ramoneda
Baotong T ian
Michael Xie
Rui Y ang
Jingwei Zhao
T aehyeon Kim
Minsoo Kang
Hyunjae Kim
Seokbeom Park
Hyojin Kim
T aein Song
Mirinae Lee
Hoyeol Sohn
Minhee Lee
Sunjae W on
Y oonjeong P ark
Gyubin Lee
Carolina Carusi
Junwon Lee
K yung T aek Oh
Sangeun Cho
Hyerim Y un
Hannah Park
Sihun Lee
Dongmin Kim
Seonguk Ju
Seola Cho
Sojeong An
Gary Jiwon Ri
Minjun Kim
Hyeonseok Choi
Sungho Lee
K yungsu Kim
Saeyeon Hw ang
Eunsik Shin
T ae yeun Hwang
Jaeyoung Shin
Subeen Kim
Y eeun Shin
Y ideun (Eden) Park
Minji Kim
v
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Pr ogram Committee
Meta-Re viewer s
V inoo Alluri, IIIT - Hyderabad
V ipul Arora, IIT Kanpur
Claire Arthur , Geor gia Institute of T echnology
Andreas Arzt, Apple
Juan P . Bello, New Y ork Uni versity
Emmanouil Benetos, Queen Mary Uni versity of London
Rachel Bittner , Spotify
Dmitry Bogdanov , Univ ersitat Pompeu Fabra
Juan J. Bosch, Spotify
Nicholas J. Bryan, Adobe Research
John Ashley Bur goyne, Uni versity of Amsterdam
Jor ge Calvo-Zaragoza, Uni versity of Alicante
Carlos Eduardo Cancino-Chacón, Johannes K epler Univ ersity Linz
Mark Cartwright, Ne w Jersey Institute of T echnology
Kahyun Choi
Nathaniel Condit-Schultz, Geor gia Institute of T echnology
Simon Dixon, Queen Mary Uni versity of London
Chris Donahue, CMU
Hao-W en Dong, Univ ersity of Michigan
Stephen Do wnie, organization
Zhiyao Duan, Uni versity of Rochester
Sebastian Ewert, Spotify
Arthur Flex er , Johannes K epler Uni versity Linz
Ichiro Fujinaga, McGill Uni versity
Satoru Fukayama, National Institute of Adv anced Industrial Science and T echnology (AIST)
Masataka Goto, National Institute of Adv anced Industrial Science and T echnology (AIST)
Fabien Gouyon, P andora/SiriusXM
Dorien Herremans, Singapore Uni versity of T echnology and Design
Andre Holzapfel, KTH Royal Institute of T echnology in Stockholm
Y u-Fen Huang, Academia Sinica
Ozgur Izmirli, Connecticut College
Blair Kaneshiro, Stanford Uni versity
Jaehun Kim, Pandora / SiriusXM
Katherine M. Kinnaird, Smith College and USAF A
Katerina K osta, ByteDance
Audrey Laplante, Uni v ersité de Montréal
Stefan Lattner , Sony Computer Science Laboratories, P aris
Alexander Lerch, Geor gia Institute of T echnology
Florence Le ve, Uni versité de Picardie Jules V erne - Lab . MIS - Algomus
Cynthia C. S. Liem, Delft Uni versity of T echnology
Ethan Manilo w , Interacti ve Audio Lab, Northwestern Uni versity
Brian McFee, Ne w Y ork Uni versity
Cory McKay , Marianopolis Colle ge
Andre w McPherson, Imperial College London
Meinard Müller , International Audio Laboratories Erlangen
Hema A. Murthy , IIT Madras
Eita Nakamura, K yushu Univ ersity
Oriol Nieto, Adobe Research
Ser gio Oramas, Pandora
Bryan Pardo, Northwestern Uni v ersity
Johan Pauwels, Queen Mary Uni v ersity of London
vii
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Preface
Conference Theme: Harmon y of T radition and Modernity
ISMIR 2025 embraced the theme "Harmony of T radition and Modernity ," encouraging di verse perspecti ves on ho w MIR
can bridge past and present. W e welcomed research exploring the multifaceted intersections of tradition and inno va-
tion—from the preserv ation and analysis of traditional music forms to the study of contemporary music trends powered
by computational methods and data. The conference adv anced the understanding of music as a dynamic and e volving
cultural force, engaging with both its rich historical roots and its e ver -expanding horizons.
Confer ence Logo: The ISMIR 2025 logo draws inspiration from Ilwol-obongdo , the royal folding screen that traditionally
stood behind the K orean throne, depicting fiv e peaks with the sun and moon symbolizing cosmic balance. The design
integrates modern architectural elements from Daejeon—the Hanbit T o wer , Expo Bridge, and KAIST’ s signature blue
color—representing technological adv ancement. Musical innov ation is embodied by the pedal-up symbol, a musical
notation marking clear transition points. The color duality of blue (representing academic rigor) and red (representing
creati ve ener gy) reflects the conference’ s theme of harmonizing tradition with modernity . The logo was designed by
Joonhyung Bae, KAIST .
Messages fr om the General Chairs
It is our great pleasure to welcome you to the 26th Conference of the International Society for Music Information Retrie v al
(ISMIR 2025). The ISMIR conference is the world’ s leading forum for research on processing, searching, or ganizing,
and accessing music-related data. This year’ s edition takes place in Daejeon, K orea, from September 21 to 25, 2025, and
is jointly or ganized by the K orea Advanced Institute of Science and T echnology (KAIST), Sogang Uni versity , and the
K orean Society for Music Informatics.
W e are delighted to present the ISMIR 2025 program. This year , we recei ved 324 abstracts, from which 278 papers
were re viewed. Of these, 99 papers were accepted with an acceptance rate of 35.6% (35.84% in ISMIR 2024). As in
pre vious years, the revie w process was conducted under a double-blind, two-tier model, in volving 269 re vie wers and
74 meta-re viewers (up from 256 re vie wers and 70 meta-re viewers in 2024). Each paper recei ved at least three re vie ws,
including one from a meta-re viewer . W e are deeply grateful to all re vie wers and meta-revie wers for their time, expertise,
and dedication. The accepted papers, authored by 413 authors (328 unique authors), were presented in both oral and
poster sessions, with a mean of 4.17 authors per paper (median: 4; max: 10).
Guided by this year’ s special theme, Harmony of T radition and Modernity , the program features two ke ynote talks: a
legendary K-pop producer and a K orean traditional music expert. It also includes a special session introducing research on
Asian traditional music from musicological and anthropological perspecti ves, an industry session sho wcasing cutting-edge
music services from our sponsors, and a WIMIR session reporting ongoing ef forts to broaden di versity and inclusi veness in
the MIR community . Ev ening e vents include a music program that demonstrates creati ve applications of MIR technologies
in musical works, a concert of K orean traditional music, the e ver -popular jam session, and the RenCon challenge, which
e valuates systems capable of rendering e xpressi ve musical performances from symbolic scores. In addition, we prepare
K-Culture Night, a social e vent where participants can e xperience K orean traditional games, food, costumes, and music.
Continuing ISMIR’ s well-established traditions, the program also of fers a full day of tutorials, a half day of late-breaking/demo
and unconference sessions, and three satellite e vents before and after the main conference: the W orkshop on Human-
Centric Music Information Research (HCMIR25), the International Conference on Digital Libraries for Musicology
(DLfM), and the W orkshop on Large Language Models for Music & Audio (LLM4MA).
W e would like to e xpress our sincere gratitude to our 22 ISMIR sponsors and 3 WIMIR sponsors, raising $110,000 in
support. W e are particularly delighted to welcome 9 first-time ISMIR sponsors. Our sponsors represent a truly global
partnership: 10 from Asia, 7 from America, and 5 from Europe. Specifically , we thank: Adobe, Moises, Udio, Algo-
riddim, Google DeepMind, Spotify , Y amaha, Neutune, Steinber g, AlphaTheta, Suno, Uni versal Music Group, Cochl.,
AudibleMagic, BMA T , Deezer , Gaudio, MIPPIA, Roland, AudAI, Piascore, and Neutone. Their generous support makes
ISMIR 2025 possible. W e also gratefully ackno wledge the local financial support of the K orea T ourism Or ganization
and the Daejeon T ourism Organization. The conference was supported by the Ministry of Education of the Republic of
K orea and the National Research F oundation of K orea (NRF-2024S1A5C3A03046168). Finally , we thank the or ganizing
committee for their dedication, professionalism, and tireless work in bringing this conference to life.
xv
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
W e are not only organizers b ut also people who truly lov e ISMIR and its community . What we hav e learned, experienced,
and shared with so many colleagues at ISMIR has shaped us into the researchers we are today . It is a joy for us to return
the passion, inspiration, and kindness we hav e recei ved from ISMIR ov er the years. Every ISMIR has al ways remained in
our hearts as a joyful and inspiring time, and we ha ve prepared this year’ s e vent with the hope that it will be remembered
in the same way for you. W e warmly in vite you to enjoy the v ery first ISMIR in K orea to the fullest!
Conference Statistics
ISMIR 2025 brought together 596 participants: 489 full on-site attendees, 32 single-day on-site attendees, and 75 online
participants. W ith virtual attendees from 45 countries, supported by dedicated virtual volunteers and chairs, the confer-
ence fostered a truly global community . Ov er 600 members joined our Slack workspace for real-time communications,
and numerous vie wers followed the proceedings on Y ouT ube and engaged with the conference program through Mini-
Conf (ismir2025program.ismir .net). Interacti ve tutorial sessions were held via Zoom, demonstrating the conference’ s
far -reaching impact across continents.
Scientific Pr ogram
The scientific program recei ved 324 abstract submissions, from which 278 papers were revie wed. Of these, 99 papers
were accepted with an acceptance rate of 35.6% (35.84% in ISMIR 2024). The re view process w as conducted under
a double-blind, two-tier model, in volving 269 re vie wers and 74 meta-revie wers (up from 256 re viewers and 70 meta-
re viewers in 2024). Each paper recei ved at least three re vie ws, including one from a meta-revie wer . The 99 accepted
papers were authored by 413 authors (328 unique authors), with a mean of 4.17 authors per paper (median: 4; max: 10).
Accepted papers were presented in both oral and poster sessions throughout the conference.
Grants Pr ogram
The conference supported equitable access through a comprehensi ve grants program, distrib uting 99 registration wai v ers
(84 on-site, 15 virtual), 44 accommodation grants, and 5 tra vel grants. Grants were aw arded across four categories:
Paper Authors, WIMIR (including Accessibility and Childcare), Music Authors, and LBD/Satellite Events, enabling
participation from students, underrepresented groups, and researchers from lo w- or middle-income countries.
Late-Breaking/Demo Session
The Late-Breaking/Demo (LBD) session provided a platform for sho wcasing inno vati ve preliminary w ork in MIR. W ith
a capacity of 75 posters accepted on a first-come-first-serv ed basis, the session of fered an accessible entry point for new-
comers and early-career researchers to present prototypes, datasets, and initial concepts, fostering community engagement
and feedback.
Diver sity , Equity , and Inc lusion
ISMIR 2025 prioritized inclusi ve participation through multiple initiati ves. The conference maintained a Code of Con-
duct with clear reporting channels, provided accessibility accommodations including mobility and sensory support, and
implemented a photo consent policy using yello w lan yards to indicate "do not photograph me" preferences. These efforts,
coordinated across Registration, Local Or ganization, and V irtual teams, ensured a safe and welcoming en vironment for
all attendees.
xvi
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Ne w-to-ISMIR P aper Mentoring Pr ogram
The Ne w-to-ISMIR Paper Mentoring Session w as held to enhance accessibility and encourage participation from a
broader , more di verse community . T en mentees, ne w to ISMIR, recei ved pre-submission guidance on paper writing
and research direction from senior researchers who v olunteered their time as past Meta-Revie wers.
W e sincerely thank the following mentors for their v aluable service: Ale xander Lerch, Ajay Sriniv asamurthy , Brian
McFee, Cheng-i W ang, Chris Donahue, Cory McKay , Ethan Manilow , Jordan B. L. Smith, LEVE Florence, T aegyun
Kwon (Emer gency Mentor).
Ne wcomer Squad
The Ne wcomer Squad, coordinated by Seungheon Doh, matched 52 ne wcomers with 13 squad leaders, forming sub-
clusters based on research interests. An ISMIR onboarding session was held to help ne wcomers na vigate the conference
and connect with the community .
W e sincerely thank the following squad leaders for their dedication: Ajay Srini v asamurthy , V incent Lostanlen, Anja V olk,
Ser gio Oramas, J. Stephen Downie, P atricia Hu, Jin Ha Lee, Stefan Balke, Lele Liu, Bruno Di Giorgi, Jan Haji ˇ
c jr ., Jordan
B. L. Smith, Jingwei Zhao, and Gabriel Meseguer Brocal.
Evening Events and Special Pr ograms
The conference featured se veral e vening e v ents that complemented the scientific program. The Music Program demon-
strated creati ve applications of MIR technologies in musical works. A K orean T raditional Music Concert showcased
traditional K orean music heritage. The e ver -popular Jam Session brought together musicians from the community for
informal performances. K-Culture Night of fered participants an immersiv e experience of K orean traditional games, food,
costumes, and music. The Industry Session sho wcased cutting-edge music services from our sponsors, pro viding insights
into real-world applications of MIR technologies.
Juhan Nam, Dasaem Jeong, and K eunwoo Choi
General Chairs of ISMIR 2025
xvii
T able of Contents
K e y n o t e S p e a k e r s .................................................. 1
S p e c i a l S e s s i o n s ................................................... 2
T u t o r i a l s ....................................................... 5
S a t e l l i t e E v e n t s ................................................... 6
Papers – Session 1 9
GlobalMood: A Cross-Cultural Benchmark for Music Emotion Recognition
Harin Lee, Elif Celen, P eter Harrison, Manuel Anglada-T ort, P ol van Rijn, Minsu P ark, Mar c Sc hönwies-
ner , Nori J acoby ................................................ 1 1
RISE: Music Rearrangement for Realtime Intensity Synchronization W ith Ex ercise
Ale xander W ang, Chris Donahue, Dhruv J ain ................................ 2 0
Expanding the HAISP Dataset: AI’ s Impact on Songwriting Across T wo AI Song Contests
Lidia Morris, Michele Ne wman, Xinya T ang, Renee Singh, Mar cel Vélez Vásquez, Rebecca Le g er , Jin Ha
Lee ....................................................... 2 8
Quantifying Regularity in Music Structure Analysis
Brian McF ee ................................................. 3 6
On the De-Duplication of the Lakh MIDI Dataset
Eunjin Choi, Hyerin Kim, Jiwoo Ryu, J uhan Nam, Dasaem J eong ...................... 4 4
Conditional Dif fusion as Latent Constraints for Unconditional Symbolic Music Generation Models
Matteo P ettenò, Alessandr o Ilic Mezza, Alberto Bernar dini ......................... 5 2
Radif Corpus; Symbolic Dataset for Non-Metric Iranian Classical Music
Maziar Kanani, Seán O’Leary , J ames McDermott .............................. 6 0
Melodic and Metrical Elements of Expressi veness in Hindustani V ocal Music
Y ash Bhake , Ankit Anand, Pr eeti Rao ..................................... 6 8
Coloring Music: Bridging Music and Color Palettes for Graphic Design
T akayuki Nakatsuka, Masahir o Hamasaki, Masataka Goto ......................... 7 5
Exploring Network Adaptations for Minimum Latenc y Real-T ime Piano T ranscription
P atricia Hu, Silvan P eter , J an Schlüter , Gerhar d W idmer .......................... 8 3
A Systematic Ev aluation of Real-T ime Audio Score Follo wing for Piano Performance
Jiyun P ark, Carlos Eduar do Cancino-Chacón, Suhit Chiruthapudi, J uhan Nam .............. 9 1
Predicting Flutist Onset T iming in Duet Performance: A Multimodal Analysis of Gesture and Breath Cues
J aeran Choi, T ae gyun Kwon, J uhan Nam ................................... 1 0 0
xix
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
AI-Generated Song Detection via L yrics T ranscripts
Markus F r ohmann, Elena Epur e , Gabriel Mese guer Br ocal, Markus Schedl, Romain Hennequin . . . . . 107
Measuring Sensory Dissonance In Multi-T rack Music Recordings: A Case Study W ith W ind Quartets
Simon Schwär , Stefan Balke , Meinar d Müller ................................ 1 1 7
Papers – Session 2 125
Reformulating Soft Dynamic T ime W arping: Insights Into T arget Artifacts and Prediction Quality
J ohannes Zeitler , Meinar d Müller ...................................... 1 2 7
IT O-Master: Inference-T ime Optimization for Audio Ef fects Modeling of Music Mastering Processors
J unghyun K oo, Mar co Martinez-Ramir ez, W ei-Hsiang Liao, Gior gio F abbr o, Michele Mancusi, Y uki Mit-
sufuji ..................................................... 1 3 4
A Multidimensional Approach to Opera Analysis: Harmony , T empo, and Dramatic Interaction in W agner’ s
Siegfried Act III
P ascal Schmolenzky , Stephanie Klauk, Rainer Kleinertz, Christof W eiss, Meinar d Müller ......... 1 4 2
Exploring the Feasibility of LLMs for Automated Music Emotion Annotation
Meng Y ang, J on McCormac k, Maria T er esa Llano, W anchao Su ....................... 1 5 0
An Ev aluation Strategy for Local K ey Estimation: Exploiting Cross-V ersion Consistency
Y iwei Ding , Y annik V enohr , Christof W eiss .................................. 1 5 8
T uning Matters: Analyzing Musical T uning Bias in Neural V ocoders
Hans-Ulrich Ber endes, Ben Maman, Meinar d Müller ............................ 1 6 6
Aligning T e xt-to-Music Evaluation W ith Human Preferences
Y ic hen Huang, Zachary No vack, K oichi Saito, Jiatong Shi, Shinji W atanabe, Y uki Mitsufuji, John Thic k-
stun, Chris Donahue ............................................. 1 7 4
In v estigating Music T rack Liking in the Halo of Album Co vers
Ole g Lesota, Anna Hausber ger , Ivanna Pshenychna, Oleksandr Shvydanenko, Olha Y ehor ova, Markus
Schedl ..................................................... 1 8 2
Phylo-Analysis of F olk T raditions: A Methodology for the Hierarchical Musical Similarity Analysis
Hilda Romer o-V elo, Gilberto Bernar des, Susana Ladr a, José R. P aramá, F ernando Silva ......... 1 9 0
dPLP: A Dif ferentiable V ersion of Predominant Local Pulse Estimation
Ching-Y u Chiu, Sebastian Strahl, Meinar d Müller .............................. 1 9 8
PeakNetFP: Peak-Based Neural Audio Fingerprinting Rob ust to Extreme T ime Stretching
Guillem Cortès-Sebastià, Benjamin Martin, Emilio Molina, Xavier Serra, Romain Hennequin ....... 2 0 6
Generating Symbolic Music From Natural Language Prompts Using an LLM-Enhanced Dataset
W eihan Xu, J ulian McA uley , T aylor Ber g-Kirkpatrick, Shlomo Dubno v , Hao-W en Dong .......... 2 1 5
A Surve y on V ision-to-Music Generation: Methods, Datasets, Evaluation, and Challenges
Zhaokai W ang, Chenxi Bao, Le Zhuo, Jingrui Han, Y ang Y ue, Y ihong T ang, V ictor Shea-J ay Huang, Y ue
Liao ...................................................... 2 2 3
Emer gent Musical Properties of a T ransformer Under Contrastiv e Self-Supervised Learning
Y ue xuan K ONG, Gabriel Mese gues-Br ocal, V incent Lostanlen, Mathieu Lagr ange, Romain Hennequin . . 235
xx
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Papers – Session 3 245
Are Y ou Really Listening? Boosting Perceptual A wareness in Music-QA Benchmarks
Y ongyi Zang, Sean O’Brien, T aylor Ber g-Kirkpatrick, J ulian McA ule y , Zac hary Novac k .......... 2 4 7
GD-Retrie ver: Controllable Generati ve T e xt-Music Retrie val W ith Diffusion Models
J ulien Guinot, Elio Quinton, Györ gy F azekas ................................ 2 6 2
T ow ards Robust Automatic Music T ranscription By Measuring Cross-V ersion Consistency
Y annik V enohr , Y iwei Ding, Christof W eiss .................................. 2 7 1
Beyond Genre: Diagnosing Bias in Music Embeddings Using Concept Activ ation V ectors
Roman Gebhar dt, Arne K uhle , Eylül Bektur ................................. 2 7 9
LiLA C: A Lightweight Latent ControlNet for Musical Audio Generation
T om Baker , J avier Nistal ........................................... 2 8 7
What Song No w? Personalized Rhythm Guitar Learning in W estern Popular Music
Zakaria Hassein-Be y , Y ohann Abbou, Alexandr e d’Hooge , Mathieu Giraud, Gilles Guillemain, Aurélien
J eanneau ................................................... 2 9 6
Uni versal Music Representations? Ev aluating Foundation Models on W orld Music Corpora
Charilaos P apaioannou, Emmanouil Benetos, Alexandr os P otamianos ................... 3 0 3
A Theoretical Model of Musical Form
Martin Rohrmeier ............................................... 3 1 2
T ow ards Human-in-the-Loop Onset Detection: A T ransfer Learning Approach for Maracatu
António Pinto ................................................. 3 2 0
Instruct-MusicGen: Unlocking T ext-to-Music Editing for Music Language Models via Instruction T uning
Y ixiao Zhang , Y ukara Ikemiya, W oosung Choi, Naoki Murata, Mar co Martínez-Ramír ez, Liwei Lin, Gus
Xia, W ei-Hsiang Liao, Y uki Mitsufuji, Simon Dixon ............................. 3 2 8
T OMI: T ransforming and Organizing Music Ideas for Multi-T rack Compositions W ith Full-Song Structure
Qi He, Ziyu W ang , Gus Xia .......................................... 3 3 7
Automatic Melody Reduction via Shortest Path Finding
Ziyu W ang, Y uxuan W u, Rog er Dannenber g, Gus Xia ............................ 3 4 6
Expotion: Facial Expression and Motion Control for Multimodal Music Generation
F athinah Izzati, Xinyue Li, Gus Xia ...................................... 3 5 4
When V oices Interlea ve: T iming De viations in Six Performances of T elemann’ s Fantasias for Solo Flute
P atrice Thibaud, Mathieu Giraud, Y ann T e ytaut ............................... 3 6 3
Papers – Session 4 371
Audio Synthesizer In v ersion in Symmetric Parameter Spaces W ith Approximately Equi v ariant Flow Matching
Ben Hayes, Charalampos Saitis, Györ gy F azekas .............................. 3 7 3
SLAP: Siamese Language-Audio Pretraining W ithout Ne gativ e Samples for Music Understanding
J ulien Guinot, Alain Riou, Elio Quinton, Györ gy F azekas .......................... 3 8 2
PianoBind: A Multi-Modal Joint Embedding Model for Pop-Piano Music
Hayeon Bang, Eunjin Choi, Seungheon Doh, J uhan Nam .............. ............ 3 9 1
Enhancing Neural Audio Fingerprint Rob ustness to Audio Degradation for Music Identification
Recep Oguz Araz, Guillem Cortès-Sebastià, Emilio Molina, Joan Serr a, Xavier Serra, Y uhki Mitsufuji,
Dmitry Bogdano v ............................................... 3 9 9
xxi
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Beyond Notation: A Digital Platform for T ranscribing and Analyzing Oral Melodic T raditions
J onathan Myers, Dar d Neuman ........................................ 4 0 7
CMI-Bench: A Comprehensiv e Benchmark for Ev aluating Music Instruction Follo wing
Y inghao MA, Siyou Li, J untao Y u, Emmanouil Benetos, Akira Maezawa .................. 4 1 6
Lose the Frames: Exact Metrics for More Responsible Music Structure Analysis Evaluations
Qingyang Xi, Brian McF ee .......................................... 4 2 6
Unifying Continuous and Discrete Compressed Representations of Audio
Mar co P asini, Stefan Lattner , Györ gy F azekas ................................ 4 3 3
Improving BER T for Symbolic Music Understanding Using T oken Denoising and Pianoroll Prediction
J un-Y ou W ang, Li Su ............................................. 4 4 2
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Louis Bradshaw , Ale xander Spangher , Honglu F an, Stella Biderman, Simon Colton ............ 4 5 1
The Rhythm In An ything: Audio-Prompted Drums Generation W ith Masked Language Modeling
P atrick O’Reilly , J ulia Barnett, Hugo Flor es Gar cia, Annie Chu, Nathan Pruyne, Pr em Seetharaman,
Bryan P ar do .................................................. 4 6 0
Count the Notes: Histogram-Based Supervision for Automatic Music T ranscription
J onathan Y affe , Ben Maman, Meinar d Müller , Amit Bermano ........................ 4 6 9
Joint T ranscription of Acoustic Guitar Strumming Directions and Chords
Sebastian Mur gul, J ohannes Schimper , Michael Heizmann ......................... 4 7 7
Enabling Empirical Analysis of Piano Performance Rehearsal W ith the Rach3 MIDI Dataset
Alia Morsi, Suhit Chiruthapudi, Silvan P eter , Ivan Pilkov , Laura Bishop, Akira Maezawa, Xavier Serr a,
Carlos Eduar do Cancino-Chacón ...................................... 4 8 4
From Discord to Harmony: Consonance-Based Smoothing for Improv ed Audio Chord Estimation
Andr ea P oltr onieri, Xavier Serra, Martín Rocamor a ............................. 4 9 2
Papers – Session 5 501
K eyboard T emperament Estimation From Symbolic Data: A Case Study on Bach’ s W ell-T empered Cla vier
P eter V an Kr anenbur g, Gerben Bisschop ................................... 5 0 3
Refining Music Sample Identification W ith a Self-Supervised Graph Neural Network
Aditya Bhattacharjee , Ivan Mer esman Higgs, Mark Sandler , Emmanouil Benetos ............. 5 1 1
V ideo-Guided T e xt-to-Music Generation Using Public Domain Movie Collections
Haven Kim, Zachary No vack, W eihan Xu, J ulian McAule y , Hao-W en Dong ................. 5 1 8
PianoV AM: A Multimodal Piano Performance Dataset
Y onghyun Kim, J unhyung P ark, Joonhyung Bae , Kirak Kim, T ae gyun Kwon, Alexander Ler c h, J uhan Nam 528
LoopGen: T raining-Free Loopable Music Generation
Davide Marincione, Gior gio Strano, Donato Crisostomi, Roberto Rib uoli, Emanuele Rodolà ....... 5 3 6
Enhancing Music Recommender Systems W ith Multimedia Content: A Context-A w are Approach
Ole g Lesota, V er onica Clavijo, Attia Rizwani, Markus Sc hedl, Bruce F erwer da ............... 5 4 7
CultureMER T : Continual Pre-T raining for Cross-Cultural Music Representation Learning
Angelos-Nik olaos Kanatas, Charilaos P apaioannou, Ale xandr os P otamianos ................ 5 5 5
xxii
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
Adapti ve P ath of Prediction: An Unsupervised Method for Modeling Note-Lev el Informational Hierarchy of
Polyphony
Xiaoxuan W ang, Martin Rohrmeier ...................................... 5 6 5
V ersatile Music-for-Music Modeling via Function Alignment
J unyan Jiang, Daniel Chin, Xuanjie Liu, Liwei Lin, Gus Xia ........................ 5 7 3
Understanding Performance Limitations in Automatic Drum T ranscription
Philipp W e yers, Christian Uhle, Meinar d Müller , Matthias Lang ...................... 5 8 2
High-Resolution Sustain Pedal Depth Estimation From Piano Audio Across Room Acoustics
Hanwen Zhang, K un F ang, Ziyu W ang , Ichir o Fujinaga ........................... 5 8 9
In v estigating an Overfitting and De generation Phenomenon in Self-Supervised Multi-Pitch Estimation
F rank Cwitk owitz, Zhiyao Duan ....................................... 5 9 6
Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation
J uan Carlos Martinez-Se villa, J oan Cerveto-Serrano, Noelia Luna-Bar ahona, Gr e g Chapman, Cr aig
Sapp, David Rizo, J or ge Calvo-Zar agoza .................................. 6 0 4
Fx-Encoder++: Extracting Instrument-W ise Audio Ef fect Representations From Mixtures
Y en-T ung Y eh, J unghyun K oo, Mar co Martínez-Ramír ez, W ei-Hsiang Liao, Y i-Hsuan Y ang, Y uki Mitsufuji 612
Papers – Session 6 621
MIDI-V ALLE: Improving Expressi ve Piano Performance Synthesis Through Neural Codec Language Mod-
elling
Jingjing T ang, Xin W ang , Zhe Zhang, J unichi Y ama gish, Geraint W iggins, Györ gy F azekas ........ 6 2 3
Playability Prediction in Digital Guitar Learning Using Interpretable Student and Song Representations
Manuel Müllersc hön, Anssi Klapuri, Mar celo Rodriguez, Christian Car din ................. 6 3 1
Gregorian Melody , Modality , and Memory: Segmenting Chant W ith Bayesian Nonparametrics
V ojt ˇ ech Lanz, J an Haji ˇ c jr . .......................................... 6 3 8
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
Hitoshi Suda, J unya K oguc hi, Shunsuke Y oshida, T omohik o Nakamura, Satoru Fukayama, J un Ogata . . . 647
GO A T : A Lar ge Dataset of Paired Guitar Audio Recordings and T ablatures
J ac kson Loth, P edr o Sarmento, Saurjya Sarkar , Zixun Guo, Mathieu Barthet, Mark Sandler ........ 6 5 5
ST A GE: Stemmed Accompaniment Generation Through Prefix-Based Conditioning
Gior gio Str ano, Chiara Ballanti, Donato Crisostomi, Michele Mancusi, Luca Cosmo, Emanuele Rodolà . 663
Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?
Richa Namballa, Agnieszka Ro ginska, Magdalena Fuentes ......................... 6 7 1
Estimating Musical Surprisal From Audio in Autoregressi v e Diffusion Model Noise Spaces
Mathias Rose Bjar e, Stefan Lattner , Gerhar d W idmer ............................ 6 7 9
Improving Neural Pitch Estimation W ith SWIPE K ernels
David Marttila, J oshua D. Reiss ....................................... 6 8 8
Optical Music Recognition of Jazz Lead Sheets
J uan Carlos Martinez-Sevilla, F rancesco F oscarin, P atricia Gar cia-Iasci, David Rizo, Jor ge Calvo-Zara goza,
Gerhar d W idmer ............................................... 6 9 6
xxiii
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
T6: MIR for Health, Medicine, and W ell-being
Pr esenters: Anja V olk, Elaine Chew , Michael A. Casey
Abstract: This tutorial explores opportunities to emplo y MIR methods for music, health, medicine, and well-being. T op-
ics include MIR for music therapy , music heart theranostics, and neurology and music information in epilepsy research,
connecting MIR with interdisciplinary collaborations in healthcare.
Satellite Events
Satur da y , September 20, 2025
HCMIR25: 3rd W orkshop on Human-Centric Music Inf ormation Research
Wher e: Room #3229, P aik Nam June Hall N25 Building, KAIST
When: 14:00 – 18:00
Music and technology ha ve long intertwined, transforming how we create, share, and experience music. The 3rd edition of
the HCMIR workshop e xplores MIR’ s ethical, societal, and human-centred dimensions. Ho w can we ensure MIR systems
are inclusi ve, ethical, and aligned with human v alues? This workshop in vites researchers, practitioners, and artists to
engage in a multidisciplinary discussion.
K eynote: Understanding the Human Experience of Music to Shape Future T echnologies of MIR
Prof. Kyung Myun Lee, KAIST
F or further details: https://sites.google.com/view/hcmir25/home
Frida y , September 26, 2025
DLfM 2025: 12th International Confer ence on Digital Libraries for Musicology
Location: Gabriel Hall, Sogang Uni versity , Seoul, South K orea
Satellite event of: ISMIR 2025
The International Conference on Digital Libraries for Musicology (DLfM) presents a venue for those w orking on, and
with, digital library systems and content in the domain of music and musicology . DLfM welcomes contributions related to
any aspect of digital libraries and musicology , including musical archi ving and retriev al, cataloguing, musical databases,
music encodings, computational musicology , or the application of MIR to musicology .
Pr ogramme Chair: Elsa De Luca, CESEM, Uni versidade No v a de Lisboa
General Chair: David M. W eigl, mdw – Uni versity of Music and Performing Arts V ienna
Local Chair: Dasaem Jeong, Sogang Uni versity
F or further details: https://dlfm.web.ox.ac.uk/12th- international- conference- digital- libraries- musicology
F or further details: https://dlfm.web.ox.ac.uk/
Frida y , September 26, 2025
LLM4MA: Large Language Models f or Music & A udio
Location: Jung Geun Mo Conference Hall (5F), KAIST , Daejeon, K orea
T ime: 8:30 am – 5:00 pm
Online: https://zoom.us/j/99541677917
LLM4MA explores the rapidly e v olving intersection of large language models (LLMs) and music/audio understanding and
generation. The workshop pro vides a forum for discussing adv ances in tokenization, long-context modeling, multimodal
6
Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025
alignment, and controllability in music applications. It fosters early-stage research and community exchange on emer ging
methods, challenges, and ethical considerations in AI-dri ven music creation.
K eynote: Science of AI and AI for Science
Prof. Noah A. Smith, Univ ersity of W ashington & Allen Institute for AI
Organizing Committee: Chenghua Lin, SeungHeon Doh, Liumeng Xue, Ilaria Manco, Gus Xia, and others
7
P aper s – Session 1
GLOB ALMOOD: A CR OSS-CUL TURAL BENCHMARK FOR MUSIC
EMO TION RECOGNITION
Harin Lee 1 , 2 , 3 Elif Çelen 1 P eter Harrison 4 Manuel Anglada-T ort 5
P ol van Rijn 1 Minsu Park 6 Mar c Schönwiesner 3 Nori J acoby 1 , 7
1 MPI Empirical Aesthetics 2 MPI for Human Cogniti ve and Brain Sciences
3 Leipzig Uni versity 4 Uni v ersity of Cambridge 5 Goldsmiths, Univ ersity of London
6 Ne w Y ork Uni v ersity Abu Dhabi 7 Cornell Uni v ersity
ABSTRA CT
Human annotations of mood in music are essential for mu-
sic generation and recommender systems. Howe v er , ex-
isting datasets predominantly focus on W estern songs with
terms deri ved from English, which may limit generalizabil-
ity across di verse linguistic and cultural backgrounds. W e
introduce ‘GlobalMood’, a nov el cross-cultural benchmark
dataset comprising 1,180 songs sampled from 59 countries,
with lar ge-scale annotations collected from 2,519 indi vid-
uals across fi ve culturally and linguistically distinct loca-
tions: U.S., France, Mexico, S. K orea, and Egypt. Rather
than imposing predefined emotion and mood categories,
we implement a bottom-up, participant-dri ven approach to
or ganically elicit culturally specific music-related emotion
terms. W e then recruit another pool of human participants
to collect 988,925 ratings for these culture-specific de-
scriptors. Our analysis confirms the presence of a v alence-
arousal structure shared across cultures, yet also re veals
significant di ver gences in how certain emotion terms (de-
spite being dictionary equi valents) are percei v ed cross-
culturally . State-of-the-art multimodal models benefit sub-
stantially from fine-tuning on our cross-culturally balanced
dataset, particularly in non-English contexts. Broadly , our
findings inform the ongoing debate on the uni versality v er-
sus cultural specificity of emotional descriptors, and our
methodology can contrib ute to other multimodal and cross-
lingual research.
1. INTR ODUCTION
Music e vok es div erse emotional responses in listeners,
spanning a wide spectrum beyond basic emotional cate-
gories [1, 2]. A central challenge in Music Information
Retrie val (MIR) is designing algorithms that can replicate
this emotional sensiti vity . This is crucial for building rec-
ommendation systems that align with listeners’ mood and
© H. Lee, E. Çelen, P . Harrison, M. Anglada-T ort, P . v an
Rijn, M. Park, M. Schönwiesner , and N. Jacoby . Licensed under a Cre-
ati ve Commons Attribution 4.0 International License (CC BY 4.0). Attri-
bution: H. Lee, E. Çelen, P . Harrison, M. Anglada-T ort, P . van Rijn, M.
Park, M. Schönwiesner, and N. Jacoby , “GlobalMood: A cross-cultural
benchmark for music emotion recognition”, in Pr oc. of the 26th Int. So-
ciety for Music Information Retrieval Conf ., Daejeon, South Korea, 2025.
context [3–5], and for generating music that resonates with
indi vidual preferences [6]. More broadly , understanding
ho w music con ve ys emotion is a core question in the sci-
ence of music [7–9]. T o date, ho wev er , most algorithms
ha ve been trained on datasets deri ved from W estern listen-
ers and W estern music, using taxonomies primarily based
on English language (e.g., MIREX [10]).
A significant challenge is creating cross-cultural mod-
els capable of handling non-W estern music and emotion
v ocabularies be yond English. Addressing this challenge
is essential to de veloping algorithms that accurately reflect
global users’ preferences, including those whose musical
tastes extend be yond the limited range of styles currently
represented in training datasets. Moreov er , without cap-
turing culturally specific nuances of emotion, especially
those dif ficult to translate, key aspects of musical mean-
ing may be missed entirely . Direct dictionary translations
of English terms may be insuf ficient, as terms describing
emotions are deeply cultural and may lack exact equi v a-
lents [11–14].
T o address these issues, we introduce ‘GlobalMood’, 1
a ne w benchmark dataset designed to support culturally in-
clusi ve and linguistically di verse emotion and mood recog-
nition in music. Our contrib ution innov ates along three
ke y dimensions: (i) the div ersity of musical stimuli, drawn
from 59 countries; (ii) the div ersity of annotators, span-
ning fi ve distinct re gions (with plans to extens to ov er 20
languages and locations in future); (iii) a data-dri ven ap-
proach for collecting descriptors, generated org anically by
participants in their o wn language during the annotation
process.
Data were collected through two stages in volving a to-
tal of 2,519 participants and 1,180 songs balanced e venly
across 59 countries: In the first stage (Section 4.1; Fig-
ure 1), using a smaller subset of 200 songs, we employed
our recently de veloped iterati ve task that combines open-
ended elicitation with collectiv e refinement [13, 15, 16].
Rather than asking listeners to choose from a fix ed list of
pre-defined emotion terms, we asked them to describe the
percei ved emotion con veyed in the music using free-te xt
tags in their nati ve language, and at the same time, rate the
1 All code and data: https://github.com/harin- git/
GlobalMood
11
Figur e 1 . Elicitation and refinement of music emotion terms through it erati v e participant chains. (A) Schematic illustration
of the collaborati v e tagging process within a participant chain. P articipants contrib ute ne w emotion-related w ord tags for
each song, rate the rele v ance of e xisting tags, and can also flag irrele v ant content, creating a dynamic refinement system.
(B) T wenty most reliable emotion tags in each language, rank ed by their tag scores. Y -axis labels display tags in their
original language (left) and English translations (right).
tags pro vided by pre vious listeners. This approach w as k e y
to unco v ering emotion terms that w ould otherwise be o v er -
look ed by predefined, English-based taxonomies (such as
‘appeal/plead’ that appears in K orean only).
In the second stage (Sect ion 4.2; Figure 2), we selected
the top 20 elicited terms per language and cro wdsourced
ratings for each tag across the entire set of 1,180 songs.
This resulted in a total of 988,925 ratings, creating the most
comprehensi v e open-source cross-cultural emotion anno-
tation dataset in Music Emotion Recognition (MER) to
date.
W e le v eraged GlobalMood to test se v eral recent multi-
modal and multilingual models (Gemini, CLAP) by e v al-
uating their performance under zero-shot, fe w-shot, and
fine-tuned scenarios (Section 4.3; Figure 3). Models
trained only on English data performed poorly in some
cultural conte xts, b ut fine-tuning with our cross-cultural
data greatly impro v ed their performance in non-English
settings. This highlights the critical importance of cross-
cultural data in both training MER models and establishing
appropriate benchmarks for their e v aluation.
2. RELA TED W ORKS
2.1 Music Emotion and Mood Annotation Datasets
Se v eral datasets ha v e been de v eloped for MER systems
with v arying annotation approaches. 2 Early e xamples in-
2 Note that databases often e xtend the concept of emotion to include
related constructs such as mood or feeling . Here we adopt this broader
clude the widely used MIREX 2007 mood dataset [10]
with 240-250 W estern songs in fi v e mood clusters deri v ed
from AllMusic’ s English tags (e.g., ‘passionate–rousing’,
‘wistful–bittersweet’), and CAL500 [17] with 500 W estern
pop/rock songs annotated using 18 English mood terms by
U.S. under graduate listeners. Ov er time, lar ger datasets ap-
peared: the DEAM corpus (MediaEv al ‘Emotion in Music’
dataset [18]) containing 2,058 song e xcerpts with contin-
uous v alence/arousal annotations; mood tags mined from
lar ge corpora of Spotify music playlist s [19]; and the
MTG-Jamendo dataset [20], which pro vides mood/theme
tags for 18,486 songs. Notably , Jamendo’ s tags were freely
cro wdsourced (56 unique mood labels), which introduced
more label v ariety b ut still almost entirely in English.
A common limitation a cross these datasets is their re-
liance on predefined English descriptors, man y of which
stem from W estern music psychology (for an e xception,
see Strauss et al. [21]). F or instance, the Gene v a Emo-
tional Music Scale (GEMS) defines 45 emotion descrip-
tors (e.g., ‘jo yful acti v ation’) based on studies with Eu-
ropean listeners [ 1 ] , and this taxonomy has been used to
annotate datasets lik e Emotify [22]. Similarly , the mood
cate gories in MIREX and CAL500 were fix ed in adv ance
(dra wn from AllMusic or prior lit erature) and presented to
annotators as a closed set of options. Consequently , these
top-do wn approaches restrict annotators to the moods the
researchers en visioned, lea ving an y unlisted mood nuances
perspecti v e, while ackno wledging that subtle distinctions between them
do e xist.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
12
Figur e 2 . Association between emotion terms across languages. (A) MDS visualization of the emotion ter ms based
on mean ratings across the full song set. T erms positioned closer together e xhi bit similar rating patterns across songs,
suggesting similar interpretations across languages. (B) Comparison of terms with direct translation equi v alents across
languages. The area size indicates the de gree of semantic di v er gence despite apparent translation equi v alence.
uncaptured and undocumented.
Ackno wledging these limitations, recent research has
be gun e xploring MER be yond the W estern-centric scope.
Hu et al. [23] e xamined mood annotations of K-pop songs
pro vided by both K orean and American listeners. Their
approach in v olv ed translating the original MIREX mood
cate gories into K orean for local annotators. Although this
allo wed direct comparisons of mood classi fication between
K orean and American listeners, it inherently restricted K o-
rean annotations to terms originally defined within W estern
conte xts.
More recently , we compiled a balanced set of Ameri-
can, Brazilian, and K orean songs and g athered mood an-
notations across nine cate gories, where annotators rated
songs both from their o wn and the other tw o countries [12].
W e sho wed that certain mood terms lik e ‘ener getic’ and
‘sad’ are highly consistent across cultures, while more ab-
stract concepts lik e ‘lo v e’ and ‘dreamy’ di v er ge consider -
ably . Simila r findings ha v e been reported by other stud-
ies [13, 14], highlighting that when mood descriptors are
imposed from one language onto another , important mean-
ings can simply be ‘lost in translation’.
In summary , whil e e x i sting MER datasets and research
ha v e laid a solid groundw ork, the y remain limited by insuf-
ficient linguistic and cultural di v ersity . Because man y are
predominantly English-based and rely on top-do wn anno-
tation strate gies, the y may o v erlook ho w people in other
cultural conte xts percei v e emotion and mood in music.
2.2 A udio LLMs: the New Fr ontier in Music T agging
Recent adv ances in multimodal lar ge language models
(LLMs) ha v e opened promising a v enues for do wnstream
MIR tasks, including emotion recognition. These mod-
els combine the reasoning capabilities of LLMs with audio
perception systems (audio LLMs), enabling more fle xible
and nuanced music understanding than traditional classifi-
cation approaches [24–26].
Models lik e MuLan [27] and MER T [24] ha v e demon-
strated potential for zero-shot music emotion and mood
classification by embeddi n g audio and natural language de-
scriptions in a shared semantic space. Ho we v er , compre-
hensi v e benchmarks such as the MuChoMusic [25] high-
light a crucial limitation: these models rely hea vily on lan-
guage modality and do not attend suf ficiently to audi o, of-
ten f ailing with more nuanced audio e xamples for do wn-
stream MIR tasks. This limitation could be particularly
critical for non-W estern music and non-Engl ish emotion
descriptors, gi v en that their training data are lar gely from
W estern conte xts.
Similarly , closed-source models (e.g., Gemini) ha v e
sho wn promise in psychological te xtual analysis in multi-
lingual conte xts [28], while e v aluations in specialized do-
mains such as MIR remain scarce. The proprietary nature
of their training data complicates thorough assessment of
cross-lingual or cross-cultural performance. W e aim to ad-
dress these fundamental g aps by pro viding a lar ge set of di-
v erse, multilingual descriptors and annotations to support
broader cross-cultural generalizability of audio LLMs.
3. METHOD
3.1 P articipants
W e recrui ted tw o independent sets of participants across
the tw o stages of our data collection: Stage 1 for emo-
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
13
tion term elicitation (N = 778; see Section 4.1) and Stage
2 for subsequent ratings on top 20 terms (N = 1,741;
see Section 4.2). Participants had to be at least 18 years
old, reside in the target country , and speak the target lan-
guage as their primary language. Participants from the US
were recruited through Prolific, while participants from
the other four countries (France, Mexico, S. K orea, and
Egypt) were recruited through the CINT platform. All
participants provided informed consent under an appro ved
protocol (see Section 7). P articipants were instructed to
wear headphones and had to pass a headphone screening
task [29], and a language proficiency test [30] before be-
ing eligible for the main experimental task. Experiments
were conducted in each participant’ s nati ve language (En-
glish, French, Spanish, K orean, and Egyptian Arabic),
with instructions translated using GPT -4o. Code to repli-
cate the experiment through the PsyNet frame work [31]
and all data are a vailable at https://github.com/
harin- git/GlobalMood
3.2 Globally Representativ e Song Selection
T o create a globally representativ e music dataset, we used
weekly Y ouT ube top 100 music charts (year 2017-2023)
from 59 countries, spanning six continents. T o ensure each
country’ s charts reflected its distinct popular music, we
excluded an y track appearing in more than one country’ s
chart. This left us with a country-e xclusive pool of songs.
From this pool, we sampled 20 songs per country , yielding
1,180 songs in total. This di verse set is designed to capture
a wide range of musical traditions and serve as a rob ust
testbed for cross-cultural emotion recognition. Each 15-
second audio excerpt w as trimmed from a random starting
point in the full track, and normalized at -5dB loudness.
3.3 Model Evaluation
W e used the resulting GlobalMood dataset to ev aluate se v-
eral recent multimodal and multilingual models capable of
music understanding. Specifically , we assessed Google’ s
Gemini models ( 1.5 Flash , 2.0 Flash , and the latest 2.5
Pr o ), a family of multimodal lar ge language models capa-
ble of processing and reasoning across te xt and audio (b ut
also image and video). 3 W e compared zero-shot and fe w-
shot approaches, where the latter included 10 human-rated
emotion terms as examples.
Gi ven that Gemini is closed-source, we also included
CLAP (Contrasti ve Language-Audio Pretraining) [34] as
an alternati ve, open-source model that learns joint audio-
text embeddings. CLAP has demonstrated promise in MIR
applications [35] and serves as the foundation for music-
specific models like CLaMP [36]. Here, we conducted
zero-shot e valuations through: (1) extracting audio em-
beddings from CLAP , (2) computing cosine similarities
with text embeddings of emotion terms, and (3) compar-
ing these scores to human ratings.
3 Preliminary tests with other recent multimodal models sho wed per-
formance issues—Flamingo 2 [32] struggled with rating consistency and
GPT -4o [33] f ailed to generate musical descriptions or ratings from audio
alone—thus we excluded them from further analysis.
W e also fine-tuned CLAP on GlobalMood (train–test
split = 1,000:180) to assess potential performance im-
prov ements. T o preserve the continuous nature of our rat-
ings, we represented each term in proportion to its mean
rating (e.g., the term ‘calm’, with a mean rating of 3.0,
appeared three times in the text). This method retained
the nuanced information in our soft labels rather than re-
ducing them to binary categories. T o improv e generaliz-
ability , we created 10 augmented variations of each song
through pitch shifting (range of ± 3 semitones), loudness
adjustment (range of ± 15dB), and the addition of Gaus-
sian noise (amplitude of 0.005). Each augmented v ariant
randomly included one or two of these modifications.
4. RESUL TS
4.1 Bottom-up T erm Elicitation Across Languages
4.1.1 T agging pipeline
Many e xisting studies on music emotions rely on pre-
defined taxonomies or web-scraped data that of fer lim-
ited linguistic div ersity [10, 19]. T o ov ercome this limita-
tion, we employed a bottom-up, participant-dri ven tagging
method [13, 15, 16]. Specifically , we asked participants in
each country to complete independent ‘chains’ of iterati ve
annotations. A subsample of 200 songs from the 1,180
entire set was used as stimuli. This subsample consisted
of 180 balanced songs across countries, with an additional
20 local songs drawn from the participating country’ s pool.
This was to ensure that local participants encounter enough
music strongly tied to their background, allowing them to
elicit culturally specific emotion descriptors.
Figure 1A illustrates one such chain: (i) The first partic-
ipant annotates the song using single-word emotion tags in
their nati ve language; (ii) The second participant (from the
same country) rates the rele vance of these tags (1–5 scale),
flag irrele vant tags (e.g., genre- or lyrics-related rather than
emotion), add new tags as necessary; (iii) The third par-
ticipant sees all tags from earlier participants and repeats
these steps; (iv) This iterati ve process continues through
ten participants per chain, systematically refining and val-
idating emotion terms. In each country , we ran the entire
elicitation experiment twice and aggre gated the results to
increase the di versity of responses from a lar ger pool of
participants.
4.1.2 T op emer ging terms
Follo wing the remo val of tags flagged by more than tw o
participants in a chain, our STEP-T ag process yielded
an extensi v e, culturally specific lexicon of emotion terms
across languages (N unique terms: English = 644; French
= 528; Spanish = 870; K orean = 629; Arabic = 283). T o
identify the most salient terms in each language, we calcu-
lated a composite score for e very term by multiplying its
frequency of occurrences across chains by its mean rele-
v ance rating. Higher scores indicate terms frequently men-
tioned and consistently rated as highly rele vant.
W e consolidated closely related morphological vari-
ants (e.g., ‘happy’ and ‘happiness’ in English; gendered
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
14
forms such as ‘jo yeux’ and ‘jo yeuse’ in French) manu-
ally with nati v e speak ers. Figure 1B presents the resulting
20 highest-ranking tags per language, displaying both the
original w ord and English translations to f acilitate cross-
cultural comparisons.
Despite being e xplicitly ask ed to pro vide emotion
terms, participants often g a v e broader af fecti v e descrip-
tors lik e moods or feelings (e.g., ‘soft’ and ‘festi v e’). This
aligns with prior MIR literature that oft en includes both
emotion and mood, and gi v en its rele v ance for practical
use, we did not enforce a strict distinction.
4.2 Lar ge-scale Di v erse Human Ratings
4.2.1 Cr oss-cultur al r atings acr oss the entir e set
Ha ving identified the top 20 terms for each language, we
ne xt g athered e xhausti v e ratings for the entire 1,180 songs
of GlobalMood. W e recruited 1,741 ne w participants (see
Section 3.1) who listened to the 15-second e xcerpts and
rated ho w ef fecti v ely each e xcerpt con v e yed a gi v en emo-
tion and mood term (1–5 scale). F or each stimulus, partic-
ipants e v aluated se v en randomly selected terms from the
rele v ant language set. This systematic approach ensured
that, on a v erage, each song in each l anguage recei v ed 8.38
(SD = 2.40) unique participant ratings, resulting in an e x-
tensi v e collection of 988,925 ratings spanning across fi v e
languages.
4.2.2 Is ‘happy’ in my langua g e the same ‘happy’ in your
langua g e?
T o in v estig ate dif ferences in ho w each culture interprets
these terms, we constructed 100 rating v ectors (5 lan-
guages × 20 terms per language). Each v ector w as 1,180-
dimensional, capturing the mean rating per term across
the 1,180 song set. W e then performed non-metric mul-
tidimensional scaling (MDS) using correlation as the dis-
tance metric, projecting these v ectors in a tw o-dimensional
space. In this emotion ‘space, ’ terms that position close to
one another —e v en those from dif ferent languages—reflect
similar rating patterns across the musical e xamples, sug-
gesting comparable emotional interpretations across cul-
tures.
Figure 2A visualizes this emotion space. The terms
cluster into tw o main re gions: one re gion of high arousal
and high v alence (e.g., happy , ener g etic , and lively ; upper
re gion of the figure) and a second re gion of lo w arousal that
spans positi v e v alence (e.g., peaceful ; bottom left) to ne g-
ati v e (e.g., sad ; bottom right). Notably , man y transl ated
‘equi v alents’ appear close together , which might suggest
a general cross-cultural consensus on what music e v ok es
what emotions.
Ho we v er , e xamining s ix commonly shared terms that
ha v e direct translations in at least four of the fi v e lan-
guages ( fun , happy , rhythmic , lo ve , sad , and calm ) re-
v ealed v arying de grees of cross-cultural agreement (see
Figure 2B). F or each of these terms, between-country
agreement ( r between ) w as computed as the a v erage of pair -
wise corre lation coef ficients, while within-country agree-
ment ( r within ) w as calculated using split-half reliability with
Spearman-Bro wn formula. Ef fecti v ely , r within serv es as
measurement error to compare as baselines when e v alu-
ating r between .
The term calm sho wed the highest a v erage agreement
( r between = 0.52 [0.49, 0.55]; r within = 0.49 [0.38, 0.57]),
follo wed by fun ( r between = 0.46 [0.39, 0.53]; r within = 0.44
[0.27, 0.54]), lo ve ( r between = 0.44 [0.41, 0.47]; r within =
0.47 [0.33, 0.66]), sad ( r between = 0.41 [0.37, 0.45]; r within
= 0.43 [0.32, 0.58]), rhythmic ( r between = 0.38 [0.30, 0.45];
r within = 0.43 [0.25, 0.53]), and notably happy ( r between =
0.37 [0.28, 0.45]; r within = 0.45 [0.24, 0.59]).
Ov erall, considering within-country agreement (mean
r within = 0.43–0.48), most of these terms were compara-
ble in their between-country agreement (mean r between =
0.39–0.52). Ho we v er , despite being considered a basic
uni v ersal human emotion [37], happy e xhibited a consid-
erable g ap between between- and within-country agree-
ment. This emphasizes the necessity of incorporating di-
v erse cultural perspecti v es when modeling nuanced mu-
sical emotional responses. Reliance on either dictionary
translation or LLM-based translation alone could o v erlook
important, conte xt-specific nuances in emotion and mood
perception—particularly rele v ant when b uilding models
for global audiences.
Figur e 3 . Correlations between human ratings and multi-
modal model predictions. (A) Gemini models with zero-
shot prompting sho wing increase in performance with
ne wer models. (B) CLAP models in zero-shot and fine-
tuned scenarios sho wing ho w the use of multilingual an-
notations can substantially increase performance. Gray
dashed lines represent split-half reliability of human rat-
ings using the Spearman-Bro wn formula as ba seline refer -
ence of correlations achie v ed between humans. Error bars
indicate 95% CI of mean correlation across songs.
4.3 Human vs. Multimodal Models
Recent benchmarks ha v e e v aluated the capabilities of au-
dio LLMs across v arious do wnstream MIR tasks, b ut these
e v aluations ha v e also been restricted to English [25]. W e
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
15
P Y T H O N P R E - P R O C E S S I N G
I S O L A T E D D R U M T R A C K
S E C T I O N T Y P E
U s e r M u s i c
S o u r c e S e p a r a t i o n
U s e r E x e r c i s e P l a n
M u s i c S t r u c t u r e
A n a l y s i s
S E G M E N T S , B E A T T I M E S , A N D
I N T R A - S E G M E N T T R A N S I T I O N P O I N T S
R e a r r a n g e d M u s i c
U N I T Y R E A L T I M E A D A P T A T I O N
B e a t T r a c k i n g
A u d i o L o u d n e s s
M e t e r
I n t r a - S e g m e n t
C u t p o i n t s
B E A T T I M E S T A M P S
I N P U T
O U T P U T
H i g h - I n t e n s i t y
S e g m e n t s
Figur e 2 . System overview . RISE takes user music and ex ercise plans as input, preprocesses the music to identify intense
segments and intra-se gment cutpoints, and sends this information to Unity for real-time adaptation.
W e precompute a per-song loudness threshold τ , where
any se gment with a loudness abov e τ is considered high in-
tensity , and otherwise we consider it lo w intensity . Thresh-
old τ is computed relati ve to the loudest section in the
song: τ = max s n ∈ S L d ( s n ) − δ , where δ is a constant
that determines the relati ve threshold. In our implementa-
tion, we set δ = 5 decibels. If four or more consecutiv e
segments are labeled high-intensity , we reassign the seg-
ment with the lo west loudness as low-intensity and mer ge
consecuti ve se gments with the same intensity labels. Ul-
timately , our analysis induces a partition of the full track
into ≤ N segments and associated binary intensity la-
bels S ′ = { ( s ′
1 , i 1 ) , ( s ′
2 , i 2 ) , . . . } , where s ′
n are delineating
timestamps and i i ∈ { Lo w , High } .
3.2 Prepr ocessing - Estimating Intra-Segment
Cutpoints
T o better align intense segments of music with high-
intensity work out phases, we estimate a set of cutpoints
to facilitate seamless adaptation. A cutpoint is a pair of
timestamps in the music recording, consisting of a starting
timestamp and a destination timestamp. The objectiv e is to
estimate cutpoints that allo w smooth musical transitions,
such that if playback jumps from the start of a cutpoint to
its destination, users experience minimal disruption.
W e deriv e an initial set of cutpoints using the approach
proposed by Plachouras and Miron [25], which analyzes
recurrence matrices encoding the self-similarity of musi-
cal beats. Cutpoints are identified by detecting diagonals
in these matrices that correspond to repeated patterns, pin-
pointing transitions between musically coherent sections.
This results in an initial set of candidate cutpoints:
C = { ( c orig.
i , c dest.
i ) ∈ B × B }
where C is the set of all estimated cutpoint pairs, and each
cutpoint consists of a start time c orig.
i and an end time c dest.
i ,
both aligned to detected beat timestamps B .
In prior work, cutpoints hav e been applied to rear-
range music to fit external constraints, such as video du-
ration [29]. These approaches allo w cutpoints to cross sec-
tion boundaries, maximizing flexibility at the potential cost
of playback naturalness.
T o prioritize naturalness in our system, we enforce an
intra-se gment constraint, ensuring that cutpoints only jump
within a segment as opposed to across se gments. W e define
the filtered set of intra-segment cutpoints as:
C ′ : = { ( c orig.
i , c dest.
i ) ∈ C | ∃ s ′
n ∈ S ′ , s ′
n ≤ c orig.
i , c dest.
i < s ′
n +1 }
This guarantees that e very cutpoint’ s start and destination
timestamps fall within the same functional section s ′
n , pre-
serving the structural integrity of the music.
3.3 Adaptation With Cutpoints
In addition to music audio and ex ercise plan, our real-
time adaptation system takes as input the follo wing in-
formation estimated during pre-processing: (1) musical
segments and corresponding intensities, (2) seamless cut-
points, and (3) beat timestamps. T o adapt music in real
time, we define a state machine that gov erns playback be-
ha vior . The system operates in one of three possible states
(Figure 3):
• Loop State : If the system determines that a segment
should be extended, it selects a cutpoint c i where
c orig.
i > t current and c dest.
i < t current , looping pre vi-
ously played sections to increase the duration of the
current intensity segment.
• Skip State : If a segment duration needs to be short-
ened, the system selects a cutpoint c i where c orig.
i >
t current and c dest.
i > c orig.
i , skipping forward to reduce
the segment duration.
• Unmodified State : If no transition is required, the
music plays continuously without alteration.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
22
Original Music
Rearranged
Output
Original Music
Rearranged
Output
Looping
Skipping
Transition to earlier location
Transition to later location
Segment extended through repetition
Segment shortened through removal
A B C
A B C
B
A B C
A C
Figur e 3 . V isualization of adaptation modes. T op: Loop
mode extends a se gment by jumping back. Bottom: Skip
mode shortens a segment by jumping forw ard.
Playback state transitions are determined dynamically
based on the user’ s ex ercise plan, enabling the system to
adjust segment duration in real time.
3.3.1 F ilter-based tr ansitions
Sometimes, the system cannot immediately transition to
the tar get intensity because there are no a v ailable cutpoints
that matches the desired transition. In these cases, we use a
filter -based transition , gradually removing high-frequenc y
content before the transition and restoring it afterward,
similar to DJ fade techniques. These transitions are no-
ticeable and less seamless compared to cutpoint transi-
tions. W e currently apply filter -based transitions only in
unguided mode (section 3.4.1), when the ex ercise calls for
a timely transition to a high-intensity music state.
3.4 Usage Modes
W orkout habits may v ary greatly across different types of
ex ercises—weight training can require minutes of rest to
fully recov er and the actual timing can v ary greatly de-
pending on ho w exhausted the user is. Guided interv al
work outs, on the other hand, emphasize short bursts fol-
lo wed by short rest periods that are strictly timed to max-
imize time ef ficiency . Moti v ated by this observ ation, we
designed two usage modes for tw o scenarios. An unguided
User working when music is low intensity
User working when music is high intensity
User (started) resting when music is high intensity
User resting when music is low intensity
Unmodified
Skipping
Looping
Skipping
enter loop cycle
break from loop cycle
skip to high intensity
play normally until user starts working
1
2
3
4
Figur e 4 . Adaptation mode for different scenarios.
mode where the user is free to rest as long as the y need, and
a guided mode where the user follo ws a predefined ex ercise
plan. W e detail the system design of each mode belo w .
3.4.1 Unguided use
The unguided mode allo ws users to freely choose start
times and work/rest durations, b ut sometimes sacrifice
adaptation quality by using filter -based transitions when
no cutpoints are immediately a vailable. The system tak es
real-time work out state as input and adjusts the music ac-
cordingly (binary: work/rest). Currently , users manually
indicate state changes by pressing a b utton. As depicted
in Figure 4, RISE transitions between playback states de-
pending on the current work out status:
1. Starting exer cise during a low-intensity segment :
The system enters skip mode to quickly transition
to a high-intensity segment.
2. During exer cise in a high-intensity segment : The
system acti vates loop mode to sustain high-intensity
music until the user begins resting.
3. Starting r est in a high-intensity segment : The sys-
tem disables looping and switches to skip mode to
exit high-intensity se gments.
4. During r est in a low-intensity segment : The sys-
tem enters unmodified mode , allo wing the music to
play naturally until the user resumes ex ercising.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
23
3.4.2 Guided use
The guided mode provides seamless adaptations b ut re-
quires a precise ex ercise plan. The system takes two du-
rations (in seconds) for work and rest. W e iterate through
a vailable cutpoints and select the transition that results in
an adaptation closest to the desired duration. W e allow one
transition per segment to maximize naturalness. The re-
sults are close to the specified duration b ut rarely perfect
( e.g ., adjusting a 50-second segment to 32 seconds for a
30-second tar get). W e fade songs in at the beginning and
out at the end, adjusting the intro to match the rest duration.
4. QU ANTIT A TIVE EV ALU A TION
W e conducted an in-lab quantitativ e listening ev aluation to
assess the seamlessness of transitions.
4.1 Pr ocedure
W e collected 270 different songs, 30 songs each from nine
dif ferent official Spotify w orkout playlists with dif ferent
genre preferences. 13 songs were remov ed from the study
because they had similar intensity throughout the entire
song or had no a vailable intra-se gment cutpoints. For the
remaining 257 songs, we generated two 10-second audio
clips per song, one with a transition and one unmodified.
Each clip is selected from a random segment of the song.
W e randomized the transition timing to occur between tw o
and eight seconds within the 10-second clip.
This e valuation w as performed with six human listen-
ers. Each clip was randomly assigned to two listeners who
rated transition naturalness on a scale of 1-5 (1 = very jar -
ring, 5 very seamless/unnoticeable). W e compare the rat-
ings of modified and unmodified clips using a paired t-test.
4.2 Results
W e found no statistically significant differences between
the clips with transitions ( M = 4.3/5, SD = 1.1) and the
baseline clips ( M = 4.5/5, SD = 1.0), t (256) = -1.63, p =
.10, suggesting that our transitions are highly seamless and
comparable to the unmodified clips. W e observ ed a small
decrease in the a verage rating of transition clips (4.5 vs.
4.3) for two reasons. First, the beat detection algorithm is
not perfect. W e found instances where the transitions were
not perfectly aligned, causing a slight jump in rhythm. Sec-
ond, familiarity with a song influenced the detection of
transitions. Raters noted that, ev en when transitions were
completely natural, they percei ved dif ferences in e xpected
progression ( e.g ., altered lyrics) in songs they kne w well.
W e also found that unmodified clips did not receiv e perfect
scores. This was due to structural elements such as synco-
pated rhythms and abrupt breaks, intended to surprise the
listener , being perceiv ed as transitions by raters, despite
these elements being part of the original compositions.
5. USER STUD Y
W e conducted a user study to explore ho w users experience
RISE in both guided and unguided ex ercise settings.
5.1 Study Design
Participants e xercised to both unmodified music and our
adapti ve system across two blocks: guided interval training
and unguided weight training. W ithin each block, they e x-
perienced both adapti ve and non-adapti ve conditions, with
order fully counterbalanced. Each condition included a
brief tutorial, 8 minutes of ex ercise, and optional rest.
Participants selected tw o songs from a curated pool of 25
tracks, played identically across all conditions. The adap-
ti ve system modified the music in response to user acti vity ,
while the non-adapti ve v ersion left the music unchanged.
The full session lasted approximately 90 minutes.
Guided interv al training. Participants performed in-
terv al ex ercises of their choice ( e.g ., jumping jacks) ac-
cording to a 40s work / 30s rest schedule. For the non-
adapti ve system, these interv als are strict. For the adap-
ti ve system, the actual timer may v ary by seconds depend-
ing on the a vailable transitions in each section. Instruc-
tions were sho wn on a screen with countdown visuals and
sounds, modeled after popular work out timer videos [30].
Unguided weight training. Participants used dumb-
bells to perform any freeform weight e xercises ( e .g., bi-
cep curls). In adapti ve conditions, they v erbally indicated
when they were about to be gin or end a work segment,
allo wing the researcher to input music adaptation state
changes. No interface was sho wn.
Interview pr ocedure. After introducing the study and
obtaining participant demographic information and con-
sent, we ga ve them a verbal description of our system and
recorded their first impressions through a short intervie w .
After experiencing all conditions, we sho wed participants
ho w the system operated by replaying the music for them
along with visualizations of both segment intensity labels
and cutpoint transitions. This was done at the end of the
study to a void priming participants to focus on specific ma-
nipulations and to e v aluate whether the adaptations were
perceptible without guidance. The study concluded with a
semi-structured exit intervie w for v erbal feedback.
Participant and apparatus. W e recruited 12 partici-
pants (5 female, 7 male, age M = 24.8 years; SD = 2.1)
from a local uni versity . Participants re gularly ex ercise
(3: 1-2x/week, 7: 3-4x/week, 2: 5-6x/week) for consider -
able durations (1: 15-30 min, 4: 30-45 min, 1: 45-60 min,
5: 1-2 hours, 1: 2 hours+) in a v ariety of ex ercise types
(8: steady cardio, 3: interv al training, 10: strength training,
4: other). The study took place in a controlled lab en viron-
ment with music played through speakers for participant
safety . Participants recei ved $50 for their participation.
5.2 Findings
Intervie w transcripts were thematically analyzed using a
coding reliability approach [31] with the Recal2 tool [32].
The analysis resulted in an a verage raw accurac y of 86.1%
and an a verage Krippendorf f ’ s alpha of 0.72 between two
raters, indicating acceptable agreement ( α > 0.66). W e
present ke y findings below .
Participants wer e excited by the premise of their
music adapting to their work outs. After we e xplained
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
24
the study to participants b ut before they e xperienced our
system, we asked participants about their initial reactions
to the description of the study and our proposed system.
Some participants had already made manual ef forts to align
their work out with music (n = 8), such as coordinating
specific music sections with their work out (n = 4). All
participants expressed that music intensity alignment with
work outs could be helpful to them. Participants belie ved
alignment can increase moti vation, or help them feel more
ener gized and “pumped” during the workout (n = 11).
Estimated high-intensity segments aligned with their
intuitions. W e showed participants a visualization of our
system’ s estimated high-intensity music segments after the
study , and all participants found the results to be aligned
with their expectations (n = 12).
Participants f ound our cutpoint adaptation tech-
nique seamless, sub verting their expectations. Before
the study , we shared a high-le vel description of our work-
out adaptation system with participants. Many e xpressed
initial skepticism that the modifications might detract from
the naturalness of the music playback (n = 6). Out of the
participants who had concerns reg arding the naturalness
of music, all b ut one remarked that their concerns were
ov erturned after experiencing our system (n = 5). While
the filter -based modifications were regarded as more no-
ticeable, most participants either did not notice or barely
noticed any cutpoint-based modifications until the y were
sho wn a visualization at the end of the study (n = 10).
Participants appr eciated the extra motivation our
alignment appr oach provided. All but one participant
found the intensity alignment to enhance their workout e x-
perience (n = 11), gi ving them more motiv ation to push
harder in their work outs. P2 noted that when the music
intensity peaked just as the y were about to reach failure
in their work out, it helped them push through and get an
extra one or tw o reps. Similarly , P4 recalled an instance
where they e xpected the music to shift to a lo wer-intensity
segment in the middle of their w ork interval, but instead
the upbeat segment repeated, pro viding them the energy to
push through the remainder of their work out.
Participants v aried in their sensitivity to both music
modifications and alignment with exer cise. For instance,
P8 barely noticed music changes b ut felt their movements
were more consistent in the adapti ve condition, while the
nonadapti ve one felt chaotic. In contrast, P3, who regu-
larly coordinates music with work outs, immediately no-
ticed misalignments in the nonadapti ve condition and e ven
felt the ur ge to stop the music when high-energy se gments
played during their rest. Sensiti vity could also depend on
ex ercise. P6 mentioned that they treated music as back-
ground during simple ex ercises but v alued music align-
ment for moti vation during more challenging e xercises,
such as the bench press. This indicates the need to better
understand ho w different modes of music engagement in-
fluence tolerance and demand for system-dri ven changes.
Participants pr eferred adapti ve music but identified
issues with the unguided experience. Out of 12 partici-
pants, 11 preferred our adapti ve system o ver nonadapti ve
music: 6 preferred it in both scenarios, while 5 fav ored it
only in the guided setting. Those who preferred the system
only in the guided setting found two main issues with the
unguided experience. First, the input method—requiring
them to notify the researcher before ex ertion—was dis-
tracting and impractical (n = 7). Second, while cutpoint
transitions were deemed seamless and often unnoticeable,
filter transitions disrupted the natural flo w of the music.
These filter transitions were especially common during ex-
ercises with longer work durations and shorter rest peri-
ods, where excessi v e looping also occurred, leading to un-
natural adaptations of the music. As a result, participants
recommended improv ements in system customizability to
support dif ferent workout habits ( e .g., long work, short
rest) and music adaptation preferences (n = 8), manipu-
lation seamlessness (n = 4), and expanding the a v ailable
music pool or allo wing user input (n = 3).
6. LIMIT A TION AND FUTURE WORK
Our study re vealed se veral limitations, including users
finding manual input impractical and filter transitions dis-
rupting. W e propose future w ork below:
Fully automating the system with sensing technolo-
gies. T o eliminate the need for manual interaction during
work outs, we propose automating the system using real-
time sensing technologies. By integrating acti vity recogni-
tion, the system could automatically detect ex ercise phases
without manual input. The system could also model user
beha vior and predict upcoming exercise states based on
past patterns to prepare adaptation plans in adv ance.
Incorporating other types of modifications. While
cutpoints enable seamless transitions, they may be una v ail-
able when timely adaptation is required. Future work could
complement our segment-sensiti v e approach with existing
techniques, such as pace adjustment and song recommen-
dation, to compensate for minor time discrepancies. Addi-
tionally , exploring audio inpainting with generati ve mod-
els could enable smooth transitions between arbitrary seg-
ments, providing greater fle xibility and precision.
Determining high-intensity segments. RISE currently
aligns drum-prominent chorus and instrumental sections
with user ex ertion. While this approach worked well in
out study , future work can further explore ho w se gment-
le vel musical v ariations influence ex ercise and ho w these
ef fects may differ by genre or listener preference.
7. CONCLUSION
W e present RISE, a nov el system that adapts music to align
with ex ercise phases. A user study in v olving 12 partici-
pants re vealed that, despite initial skepticism, most users
appreciated the alignment and preferred it ov er nonadap-
ti ve music for their work outs. Our work represents a step
to wards expanding the design space of adapti v e music,
making tailored music experiences, once limited to video
games and precomposed soundtracks, applicable to real-
world scenarios lik e workouts.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
25
8. REFERENCES
[1] P . C. T erry , C. I. Karageorghis, M. L. Curran, O. V .
Martin, and R. L. Parsons-Smith, “Ef fects of music in
ex ercise and sport: A meta-analytic re view . ” Psycho-
logical b ulletin , vol. 146, no. 2, p. 91, 2020.
[2] S. Mitrof f, “Hitting the pa vement with
spotify running (hands-on) - cnet, ” https:
//www .cnet.com/tech/services- and- software/
hitting- the- pav ement- with- spotify- running- hands- on/,
2015.
[3] Apple, “ Apple fitness+ - apple, ” https://www .apple.
com/apple- fitness- plus/, 2024.
[4] P . Knees, M. Schedl, and M. Goto, “Intelligent user
interfaces for music disco very: The past 20 years and
what’ s to come. ” in Pr oceedings of the 20th Interna-
tional Society for Music Information Retrieval Confer -
ence, ISMIR 2019, Delft, The Netherlands, November
4-8, 2019 , 2019, pp. 44–53.
[5] M. Goto, “Grand challenges in music information re-
search, ” in Dagstuhl F ollow-Ups: Multimodal Music
Pr ocessing , M. Müller , M. Goto, and M. Schedl, Eds.
Dagstuhl Publishing, 2012, pp. 217–225.
[6] M. Schedl and A. Flex er , “Putting the user in the center
of music information retrie val. ” in Pr oceedings of the
13th International Society for Music Information Re-
trieval Confer ence, ISMIR 2012, Mosteir o S.Bento Da
V itória, P orto, P ortugal, October 8-12, 2012 , 2012, pp.
385–390.
[7] M. Kari, T . Grosse-Puppendahl, A. Jagaciak,
D. Bethge, R. Schütte, and C. Holz, “Sound-
sride: Af fordance-synchronized music mixing for
in-car audio augmented reality , ” in The 34th Annual
A CM Symposium on User Interface Softwar e and
T echnolo gy , 2021, pp. 118–133.
[8] L. Baltrunas, M. Kaminskas, B. Ludwig, O. Moling,
F . Ricci, A. A ydin, K.-H. Lüke, and R. Schwaiger , “In-
carmusic: Context-a ware music recommendations in
a car , ” in E-Commer ce and W eb T ec hnologies: 12th
International Confer ence, EC-W eb 2011, T oulouse,
F rance , A ugust 30-September 1, 2011. Pr oceedings 12 .
Springer , 2011, pp. 89–100.
[9] A. W ang, Y . F . Cheng, and D. Lindlbauer , “Maringba:
Music-adapti ve ringtones for blended audio notifica-
tion deli very , ” in Pr oceedings of the CHI Confer ence
on Human F actor s in Computing Systems, CHI 2024,
Honolulu, HI, USA, May 11-16, 2024 . New Y ork, NY ,
USA: Association for Computing Machinery , 2024.
[10] A. W ang, D. Lindlbauer , and C. Donahue, “T o wards
music-aw are virtual assistants, ” in Pr oceedings of the
37th Annual A CM Symposium on User Interface Soft-
war e and T ec hnology , UIST 2024, Pittsbur gh, P A, USA,
October 13-16, 2024 . Ne w Y ork, NY , USA: Associa-
tion for Computing Machinery , 2024.
[11] J. Shriram, M. T apaswi, and V . Alluri, “Sonus tex ere!
automated dense soundtrack construction for books us-
ing movie adaptations, ” in Pr oceedings of the 23r d
International Society for Music Information Retrieval
Confer ence, ISMIR 2022, Bengaluru, India, December
4-8, 2022 , 2022, pp. 535–542.
[12] S. Rubin, F . Berthouzoz, G. Mysore, W . Li, and
M. Agraw ala, “Underscore: musical underlays for au-
dio stories, ” in Pr oceedings of the 25th Annual A CM
Symposium on User Interface Softwar e and T ec hnol-
ogy , ser . UIST ’12. Ne w Y ork, NY , USA: Association
for Computing Machinery , 2012, p. 359–366.
[13] N. Masahiro, H. T akaesu, H. Demachi, M. Oono, and
H. Saito, “De velopment of an automatic music selec-
tion system based on runner’ s step frequency , ” in IS-
MIR 2008, 9th International Confer ence on Music In-
formation Retrieval, Dr exel Univer sity , Philadelphia,
P A, USA, September 14-18, 2008 , 2008, pp. 193–198.
[14] G. T . Elliott and B. T omlinson, “Personalsoundtrack:
context-a ware playlists that adapt to user pace, ” in
CHI’06 e xtended abstracts on Human factors in com-
puting systems , 2006, pp. 736–741.
[15] B. Moens, L. v an Noorden, and M. Leman, “D-jogger:
Syncing music with walking, ” in 7th Sound and music
computing Confer ence . Uni versidad Pompeu F abra,
2010, pp. 451–456.
[16] J. Hockman, M. M. W anderley , and I. Fujinaga, “Real-
time phase v ocoder manipulation by runner’ s pace. ” in
NIME , 2009, pp. 90–93.
[17] N. Oli ver and L. Kre ger-Stickles, “Papa: Physiology
and purpose-aw are automatic playlist generation. ” in
ISMIR 2006, 7th International Confer ence on Music
Information Retrieval , 2006, pp. 250–253.
[18] B. v an der Vlist, C. Bartneck, and S. Mäueler , “mobeat:
Using interacti ve music to guide and moti v ate users
during aerobic ex ercising, ” Applied psychophysiology
and biofeedbac k , vol. 36, pp. 135–145, 2011.
[19] Y . Chen, C.-C. Chen, L.-C. T ang, and W .-H. Chieng,
“Enhancing running ex ercise with iot, blockchain, and
heart rate adapti ve running music, ” IEEE Access , 2024.
[20] C. I. Karageorghis and D. Holland, “Music in the e xer -
cise domain: A revie w and synthesis (part ii), ” Interna-
tional Revie w of Sport and Exer cise Psycholo gy , v ol. 5,
no. 1, pp. 67–84, 2012.
[21] A. T urrell, A. R. Halpern, and A.-H. Jav adi, “When
tension is exciting: an electroencephalogram e xplo-
ration of excitement in music, ” bioRxiv , p. 637983,
2019.
[22] D.-L. Priest and C. I. Karageorghis, “ A qualitativ e in-
vestig ation into the characteristics and ef fects of music
accompanying e xercise, ” Eur opean physical education
r e view , v ol. 14, no. 3, pp. 347–366, 2008.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
26
[23] T . Kim and J. Nam, “ All-in-one metrical and func-
tional structure analysis with neighborhood attentions
on demixed audio, ” in IEEE W orkshop on Applications
of Signal Pr ocessing to A udio and Acoustics (W AS-
P AA) , 2023.
[24] R. Hennequin, A. Khlif, F . V oituret, and M. Moussal-
lam, “Spleeter: a fast and ef ficient music source sepa-
ration tool with pre-trained models, ” Journal of Open
Sour ce Softwar e , v ol. 5, no. 50, p. 2154, 2020.
[25] C. Plachouras and M. Miron, “Music rearrangement
using hierarchical segmentation, ” in ICASSP 2023-
2023 IEEE International Confer ence on Acoustics,
Speech and Signal Pr ocessing (ICASSP) . IEEE, 2023,
pp. 1–5.
[26] C.-W . Li and C.-G. Tsai, “The presence of drum and
bass modulates responses in the auditory dorsal path-
way and mirror -related regions to pop songs, ” Neur o-
science , v ol. 562, pp. 24–32, 2024.
[27] G. Madison, “Experiencing groove induced by music:
consistency and phenomenology , ” Music per ception ,
v ol. 24, no. 2, pp. 201–208, 2006.
[28] C. J. Steinmetz and J. D. Reiss, “pyloudnorm: A simple
yet flexible loudness meter in p ython, ” in 150th AES
Con vention , 2021.
[29] Adobe, “Remix in premiere pro, ” https:
//helpx.adobe.com/premiere- pro/using/
remix- audio- in- premiere- pro.html, 2024.
[30] W . M. W . T imer , “Interval timer with music | 40 sec
rounds 30 sec rest | mix 107, ” Y ouT ube video, 2021,
accessed: September 1, 2024. [Online]. A v ailable:
https://www .youtube.com/watch?v=lnBOQnc_p- E
[31] R. E. Boyatzis, T ransforming qualitative information:
Thematic analysis and code development . Sage, 1998.
[32] D. Freelon, “Recal2: Reliability for 2 coders, ” http://
dfreelon.or g/utils/recalfront/recal2/, 2010.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
27
EXP ANDING THE HAISP D A T ASET : AI’S IMP A CT ON SONGWRITING
A CR OSS TWO AI SONG CONTESTS
Lidia Morris 1 Michele Newman 1 Xinya T ang 1
Renee Singh 1 Mar cel Vélez Vásquez 2 Rebecca Leger 3 Jin Ha Lee 1
1 Information School, Uni versity of W ashington
2 Uni versity of Amsterdam
3 Fraunhofer Institute for Inte grated Circuits
[email protected] , [email protected]
ABSTRA CT
As artificial intelligence (AI) continues to shape creati ve
practices, understanding its role in human-AI songwriting
remains crucial. This paper expands the Human-AI Song-
writing Processes (HAISP) dataset by incorporating data
from the 2024 AI Song Contest, building upon the original
2023 dataset. By analyzing ne w submissions, we provide
further insights into AI’ s e v olving impact on songwriting
workflo ws, creati ve decision-making, and control. A com-
parati ve study of AI tool usage and participant strate gies
between the 2023 and 2024 contests re veals shifts in col-
laboration patterns and tool ef fectiv eness. Additionally , we
assess the dif ferences between general-purpose AI systems
and personalized, fine-tuned tools, highlighting their im-
pact on creati ve agenc y . Our findings offer k ey design im-
plications for AI-assisted songwriting tools, providing ac-
tionable insights for AI de velopers and music practitioners
seeking to enhance co-creati ve e xperiences.
1. INTR ODUCTION
Artificial Intelligence (AI) has rapidly become an integral
component of creati ve fields, reshaping artistic expression
across v arious domains. From visual arts to literature, AI-
po wered tools are being le veraged to augment human cre-
ati vity , raising ne w questions about authorship, originality ,
and the e volving nature of co-creation [1–3]. No where is
this transformation more e vident than in music composi-
tion, where AI systems are increasingly employed to gen-
erate melodies, harmonies, lyrics, and entire song struc-
tures [4, 5]. These adv ancements ha ve giv en rise to ne w
forms of collaboration between human musicians and AI,
necessitating a deeper understanding of the dynamics of
human-AI co-creation in songwriting.
The study of human-AI collaboration in music is par-
ticularly important due to the complex, often subjectiv e
© L. Morris, M. Newman, X. T ang, R. Singh, M.A. Vélez
Vásquez, R. Leger , and J.H Lee. Licensed under a Creati ve Commons
Attribution 4.0 International License (CC BY 4.0). Attribution: L. Mor -
ris, M. Ne wman, X. T ang, R. Singh, M.A. Vélez Vásquez, R. Leger , and
J.H Lee, “Expanding the HAISP Dataset: AI’ s Impact on Songwriting
Across T wo AI Song Contests”, in Pr oc. of the 26th Int. Society for
Music Information Retrieval Conf ., Daejeon, South K orea, 2025.
nature of the creati ve process. While AI can accelerate
composition workflo ws and generate nov el musical ideas,
its role in enhancing versus replacing human creati vity re-
mains a critical area of in vestig ation [6, 7]. Furthermore,
questions reg arding our changing definition of computa-
tional creati vity and the role AI can play in the creativ e
process as a tool or collaborator continue to loom large
ov er the field [8, 9]. Addressing these concerns requires
qualitati ve data that captures not just empirical informa-
tion on the use of generati ve AI, b ut the liv ed experiences
of creators working with AI in music production.
T o contribute to this gro wing field of study , the Hu-
man–AI Songwriting Processes (HAISP) dataset was in-
troduced in 2024 as a curated resource designed to ex-
plore the interaction between human musicians and AI sys-
tems [10]. The dataset was deri ved from submissions to
the AI Song Contest 2023, an annual competition that in-
vites teams of musicians, data scientists, and researchers
to explore the creati v e potential of AI in songwriting [11].
It comprises 34 coded entries documenting ho w teams
used AI tools in their songwriting processes. It provides a
structured frame work for analyzing v arious aspects of AI-
assisted music creation, including:
• The specific AI tools and models used in composi-
tion
• The songwriting methodologies employed by
human-AI teams
• Reflections on ethical considerations and challenges
related to AI in music
• T eams’ assessments of their collaborati ve e xperi-
ence with AI
The findings from the HAISP dataset highlighted the di-
verse w ays in which AI is integrated into songwriting, with
teams using AI for tasks ranging from melody generation
to performance synthesis. Howe v er , the dataset also under-
scored the limitations of AI tools, such as lack of creativ e
control, technical limitations, and concerns about origi-
nality . In addition, ethical concerns regarding the pro ve-
nance of data and the transparency of AI-generated content
emer ged as key themes.
28
Building on this foundation, the current study extends
the HAISP dataset by incorporating ne w data from the
2024 AI Song Contest, of fering a longitudinal perspecti ve
on the e volution of human-AI collaboration in songwriting.
By comparing data from 2023 and 2024, this e xpanded
dataset enables a deeper analysis of trends, emerging tech-
nologies, and shifting attitudes to ward AI in creati ve work.
Through this research, our goal is to provide v aluable in-
formation for musicians, AI de velopers, creativity schol-
ars, and beyond.
2. B A CKGR OUND
The intersection of AI and music composition represents
a rapidly e volving field that has a long history to e xplore.
AI-assisted music creation has progressed from early algo-
rithmic experiments and academic electronic music cen-
ters [12] to sophisticated machine learning models capa-
ble of composing complete musical pieces by the broader
public [13]. As these technologies become more accessi-
ble, they not only influence the w ay music is made but also
raise critical ethical, cultural, and artistic questions about
human-AI co-creation.
2.1 Evolution of AI in Music Composition
The application of computational techniques in music
composition can be traced back to the mid-twentieth cen-
tury , when early pioneers e xperimented with algorithmic
approaches to sound generation [13], such as the work
completed at the Columbia-Princeton Electronic Music
Center , which laid the groundwork for a v ariety of com-
posers careers and technological innov ations [12, 14]. By
the early 2000s, dev elopments in machine learning facili-
tated the creation of models that could autonomously gen-
erate melodies, harmonies, and song structures [15, 16].
The increasing sophistication of deep learning and gen-
erati ve AI in the past decade has further transformed the
landscape of music composition. Notable adv ances in-
clude OpenAI’ s MuseNet, Google’ s Magenta, and Meta’ s
MusicGen, all of which employ transformer -based archi-
tectures to produce di verse compositions [17–19]. These
tools enable musicians to collaborate with AI in v arious
ways, from generating musical ideas to assisting with ar-
rangement, and more [19]. The growing accessibility of
these technologies has been sho wcased in platforms such
as the AI Song Contest (AISC) [11].
2.2 AI Songwriting T ools and Methods
V arious AI-powered tools ha ve emer ged to facilitate
human-AI collaboration in songwriting. OpenAI’ s
MuseNet [20] is a deep neural network capable of com-
posing multi-instrumental pieces across multiple genres,
while Google’ s Magenta project provides open-source ap-
plications for AI-assisted melody generation, chord pro-
gression, and rhythm creation [18]. While these tools of fer
ne w creativ e possibilities, they also introduce challenges
related to artistic control, originality , and the implications
of AI as a co-creati ve entity , especially when the user is
an inexperienced music creator [2, 11, 21]. AI-generated
music is often constrained by its training data, leading
to concerns about predictability , stylistic homogenization,
and the potential for AI to reinforce e xisting musical con-
ventions rather than foster true innov ation [22, 23]. Fur-
thermore, the extent to which AI-generated compositions
can be considered “creati ve” in the same sense as human-
authored works is still being questioned, especially when
it comes to just ho w much of a role the AI plays in the
compositional process [24, 25].
2.3 Creativity Studies
The increasing adoption of AI in creati ve domains has
sparked debates about the nature of creati vity and the role
of machines in artistic e xpression [11, 26–28]. T raditional
vie ws of creativity emphasize human intuition, cultural
context, and emotional depth [29–31] - qualities that AI,
as a statistical modeling system, does not inherently pos-
sess. Can AI truly be considered a creativ e agent, or is it
merely an adv anced tool for pattern recognition and recom-
bination? Ethical concerns also e xtend to the implications
of AI’ s increasing role in the creati ve workforce [32, 33].
As AI-generated compositions become more sophisticated,
there is a potential for automation to displace human mu-
sicians in certain commercial contexts [34]. In response to
these challenges, scholars and industry professionals hav e
called for greater transparency in AI training data, ethical
guidelines for AI-assisted composition, and policies to en-
sure that human artists remain central to the creati ve pro-
cess [35].
3. D A T ASET EXP ANSIONS: METHODOLOGY
The HAISP Dataset is accessible as a .csv and .xlsx on the
Open Science Frame work (OSF) under a Creati ve Com-
mons Attrib ution-NonCommercial 4.0 International (CC
BY -NC) license, which allo ws for broad access and uti-
lization for research purposes [36].
W e generated the dataset via consensus coding [37].
One researcher coded a selection of the data entries, col-
lecting them into the dataset. A second coder then re-
vie wed the initial coding, validating the coding by ei-
ther marking agreement or disagreement with the cod-
ing choices within a comment on the code in the dataset,
adding what they felt the coder w as missing within their
codes from the data. In the case of disagreement, a third
researcher helped decide on the final code as a tie-breaker .
3.1 Data Collection
Similarly to the 2023 practice, each team had to fill in the
AI Song Contest 2024 Submission F orm via Google Forms
to participate in the contest [10]. The form consists of entry
fields that cov er all the basic information about teams and
songs:
• team (bio for the website, location, lev el of exper -
tise, moti vation to participate, how the y heard about
the AISC);
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
29
• song (title, length, link to music video/soundcloud/
blogpost, concept/idea, lyrics, li ve performance).
Each team was additionally ask ed to generate a process
document and sa ve it as a PDF file, which mainly includes
more detailed moti vation, songwriting workflo w and col-
laboration process, e valuation of co-creation, and ethical
considerations. In this way , each team had more space to
elaborate on their collaboration process than in pre vious
submissions.
Furthermore, all teams had to giv e their consent for their
responses to be published in a scientific paper . The an-
swers were collected into a Google Sheet with links to out-
standing PDF files. In total, 67 submissions were collected,
including 34 submissions that used text-to-music models
such as Udio and/or Suno in the human-AI collaboration
process and thus were considered disqualified due to the
judges inability to "assess the le vel and manner of the use
of AI in each entry" and an inability to obtain "a descrip-
tion of the data used to train the AI model." [38] After
these 34 submissions were excluded, 33 effecti v e partic-
ipating teams remain in the 2024 edition. The completed
questionnaires were then handed ov er to the research group
excluding personal data.
3.2 Methodology and V alidation
Due to the change in participant entry methods and sub-
mission format, the data dictionary from the initial HAISP
dataset was adapted by the coders to better e xtract the data
essential to the dataset from the written entries. The cate-
gories were e valuated one by one by the four researchers,
iterating three times, with testing of each new adaptation
to the dictionary before reaching final consensus. For each
iteration of the data dictionary , two coders tested it on two
sample entries to ensure that the categories were properly
defined and applicable for the ne w data.
3.3 Data Statistics
The HAISP dataset for the 2024 edition consists of data
from 33 teams, representing 22 dif ferent countries and re-
gions. The United States has the highest representation,
with 12 teams participating. The United Kingdom follo ws
with four teams, while Switzerland has three. Germany
and Spain are each represented by two teams.
Other represented countries include the Netherlands
(NLD), Colombia (COL), Japan (JPN), Italy (IT A), Brazil
(BRA), Thailand (THA), Canada (CAN), France (FRA),
T unisia (TUN), Denmark (DNK), Hungary (HUN), T urke y
(TUR), China (CHN), and Chile (CHL).
The type of af filiation of the HAISP dataset 2024 edi-
tion reflects a significant shift compared to the 2023 edi-
tion. The most notable change is a clear trend a way from
participants typically coming from academic backgrounds
to ward those who work in the creati ve industry , suggesting
a gro wing engagement of professional artists and creativ es
with AI-dri ven music composition, in artistic and commer -
cial sectors rather than academic research setting. In 2023,
58.8% of participants were af filiated with academia, mak-
ing it the dominant category . Ho we ver , in 2024, academic
af filiation dropped to just 19.57%, while the creati ve in-
dustry sur ged to 58.7%, making it the largest represented
category in this year’ s dataset.
The HAISP dataset for the 2024 edition of the AISC
sho wcases a div erse array of AI models and tools em-
ployed by participating teams. Compared to the 2023 edi-
tion, which saw the usage of 74 dif ferent AI tools, the 2024
dataset reflects an e ven broader spectrum of AI applica-
tions with 82 dif ferent AI tools, excluding the other tools
used by the disqualified participants. These tools include
AI-po wered music generation models, v oice cloning soft-
ware, AI-driv en mixing and mastering tools, AI-assisted
composition platforms, and more. Some teams utilized
publicly a vailable AI tools lik e Musicfy or Kits.ai, while
others employed custom-b uilt AI models tailored to their
specific creati ve needs, like Purr Data.
A significant portion (42%) of teams in 2024 contin-
ued to use ChatGPT and OpenAI’ s GPT -based models for
lyrics, structure, and creativ e assistance. Additionally , the
rise of Stability AI models, such as Stable-Audio-Open 1.0
and Stable Dif fusion XL, suggests a growing reliance on
AI for both music generation and visual content creation.
4. COMP ARA TIVE AN AL YSIS
The HAISP dataset re veals that while AI-assisted song-
writing can enhance creati vity and efficienc y , musicians
frequently encounter challenges related to control, trans-
parency , and process integration when w orking with AI
tools. Se veral recurring themes, as presented below ,
emer ge from the dataset that highlight the limitations of
current AI models.
4.1 Contr ol
T welve participating teams e xpressed frustration ov er the
lack of fine-grained control ov er AI-generated outputs, es-
pecially when it comes to using the more popular and
easily accessible AI systems. Users specifically choose
systems that allo w for greater control and flexibility o ver
the outputs, highlighting systems whose af fordances allow
them to control “...mechanisms such as text/audio prompt-
ing and loop generation” (T eam 63). W ith systems that
do not allo w for such user modifications, many feel that
the results are limited, and “...mostly based on seed luck
and good prompting” (T eam 28). One team in particular
noted that they chose to use an AI tool created by Ele v en-
Labs [39] not only because they felt it allowed for “greater
control, ” (T eam 52) but because of their ethical stance as
a team, noting that Ele venLabs was more ethically trans-
parent in the creation of its music database, using only li-
censed content from Shutterstock [40]. Decisions about
tools are not only based on control ov er the process or out-
put, but also on the team’ s control ov er ho w to accommo-
date or apply their ethical positions.
The 2024 HAISP dataset expansion reinforces man y of
the themes from the 2023 dataset, particularly in regards to
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
30
Figur e 1 . Bar Chart comparing the af filiation of AI Song P articipants in 2023 and 2024, sho wcasing an increase in creati v e
industry af filiation and decrease in academic af filiation.
maintaining control o v er their creation process. In 2023,
teams frequently encountered rigidity and limitations when
w orking with AI tools, as e x emplified by 2023’ s T eam 13’ s
e xperience w orking with Meta’ s MusicGen. Initially “f as-
cinated” by the outcomes that came from this tool, the y
soon realized the tool produced repetiti v e results, and fur -
ther attempts to refine the output resulted in “...an unwel-
come surplus of noise, leading to a sense of limitation. ”
This led them to switch to Google’ s Magenta so the y could
ha v e more control o v er the MIDI outputs. Other 2023
teams such as T eam 16 also noted that te xt-to-music mod-
els frequently needed more specific and detailed prompting
that included information on the k e y and tempo in order
to a v oid the “...incoherent and sometimes noisy outputs. ”
Ov erall, teams from both years sho w a stronger preference
for AI tools that allo w for real-time adjustments, iterati v e
prompting, and clear ethical stances so the y can acti v ely
and continuously mak e choices that gi v e them the control
the y desire o v er the creati v e process.
4.2 A pplications of AI
In the 2024 data, we noted that participants on a v erage
used 2-3 times more AI tools than the 2023 participants.
2024 AISC participants le v eraged AI for melody and har -
mon y generation, using models to produce initi al musical
ideas that were later refined through human music produc-
tion stages lik e mixing and mastering. V oice synthesis w as
another k e y application, with AI tools transforming v ocal
performances or generating synthetic v oices that could be
adjusted to fit the song’ s artistic vision. Unlik e 2023, the
majority of the tools used were openly a v ailable tools, and
not custom-b uilt and trained models. T ools lik e Stable Au-
dio, Ele v enLabs, and ChatGPT 3.5/4.0 were amongst the
most commonly used tools in the 2024 dataset.
In contrast, 2023’ s teams often emplo yed AI tools it-
erati v ely , using them to refine compositions and lyrics
throughout the process, which is described more as a recur -
si v e w orkflo w than step-by-step, wi th their AI tool “...pro-
viding creati v e suggestions and helping us iterate more
ef ficiently” (T eam 26). The iterati v e process means that
teams could listen to their AI-generated song elements and
add on to them with human elements as their submission
de v eloped, rather than ha ving the element be a generated
piece that cannot be recreated e xactly , e v en with the same
prompts. 2024’ s T eam 38 described their frustration with
this issue, writing that "The problem with all of this, and
what mak es this w orkflo w so granular , is the AI starts to
drift o v er time, losing sight of one aspect of the prompt in
f a v or of another; outputs be gin to dif fer in length, timbre,
language (for some reason) b ut most importantly tempo."
Additionally , man y of the tools used by the 2023 partici-
pants were either b uilt by the teams or were open source
models that were trained by the team. In 2023, lar ge-scale
AI systems lik e ChatGPT were mostly used to create sug-
gestions or generate ideas for song elements such as the
melody , which w as then played and recorded on real in-
struments by participants, or lyrics, which were then sung
by a separate AI tool or human team member .
4.3 Co-Cr eation vs. A utomation
The dataset indicates a strong preference for collaborati v e
AI tools o v er fully automated music generators. Man y
artists w ant AI to function as an assisti v e tool rather than
an autonomous composer . Users e xpress interest in AI
systems that respond dynamically to their inputs, rather
than generati ng static musical pieces that require e xtensi v e
manual re vision. As T eam 28 noted, it w as not just that
utilizing an AI tool that made the w ork co-creati v e, b ut the
combination of their o wn musical t raining and skills and
the w ork of the AI tools which allo wed for "... a seamless
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
31
T able 1 . Empirical mean ˜ ρ for uniformly sampled segment
durations d 1 ≤ d 2 in the range [4 , 60] seconds, for different
frame rates f (Hz) and tolerance windo w w (seconds).
w = 0 w = 0 . 25 w = 0 . 5
f = 40 0.009 0.349 0.477
f = 20 0.016 0.343 0.477
f = 10 0.029 0.261 0.477
f = 2 0.111 0.111 0.427
3.3 Pr operties
Before going into extensions and applications, it is worth
pausing to take note of a fe w properties of ρ, ˜ ρ , and R .
Boundedness Since gcd { d 1 , d 2 } ≤ min { d 1 , d 2 } , eqs. (1)
and (3) are bounded at 1, with equality when d 2 is an
integer multiple of d 1 (or vice v ersa). The minimal value
1 / min { d 1 , d 2 } is achie ved by relati vely prime ( d 1 , d 2 ) .
Scale-in variance F or any positi ve rational c such that c · d 1
and c · d 2 are integers, ρ ( c · d 1 , c · d 2 ) = ρ ( d 1 , d 2 ) .
Attainable values ρ ( d 1 , d 2 ) = 1 / N for some positi ve in-
teger N . This is because c = 1 / gcd { d 1 , d 2 } , which
satisfies the scale in v ariance property abov e, implies
ρ ( d 1 , d 2 ) = ρ ( c · d 1 , c · d 2 ) = 1 / min { c · d 1 , c · d 2 } .
Expected value T able 1 reports the empirical mean ˜ ρ for
uniformly sampled duration pairs ov er the range [4 , 60]
seconds. For the proposed default tolerance of w = 0 . 5 ,
the mean ˜ ρ ≈ 0 . 477 is stable for dif ferent frame rates f .
3.4 Extension 1: Section labels
Equation (2) a verages ov er all unordered pairs of distinct
segments. It often occurs that not all segments are rele-
v ant to include in this comparison: for example, introduc-
tory silences or cro wd noise may exist outside of musi-
cal time and therefore not participate meaningfully in reg-
ularity . Similarly , sections with significant deviations in
tempo from the remainder of the recording may result in
lo w scores under eq. (3), and a case could be made that
these should be treated separately .
More generally , one may consider a notion of restricted
regularity that only compares se gments with the same sec-
tion label ( e.g . , verse or c horus ). Under suitable label-
ing con v entions, this vie w encapsulates the examples listed
abov e, and provides a simple mechanism to e xclude seg-
ments with sporadically occurring labels. This idea can be
implemented with a straightforward modification to eq. (2)
where a collection of distinct segment pairs P ⊂ S × S is
provided rather than the entire se gmentation S :
R L ( P ) = 1
| P | X
( d 1 ,d 2 ) ∈ P
ρ ( d 1 , d 2 ) . (4)
2 The associati ve property of gcd and min also implies that the edge
case of a segmentation consisting of only one se gment should produce a
score of 1. This con v ention is adopted here.
3 δ is constrained to d + δ ≥ f so that eq. (3) is well-defined.
Label agreement is a simple way to generate the pair
set P , though the definition abov e supports other schemes,
e.g . automatic hierarchy expansion (for approximate agree-
ment) [12]. Relatedly , the temporal proximity observ ation
of Smith and Goto [9] can be implemented here by gener -
ating pairs of sequentially adjacent durations:
P = { ( d i , d i +1 ) | 0 ≤ i < | S | − 1 } .
3.5 Extension 2: Hierarchical r egularity
Equation (2) can be modified to e valuate the re gularity of
hierar chical se gmentations. Note that eq. (2) operates on
pairs of durations, but it does not require that the se gments
under comparison are disjoint in time or form a v alid seg-
mentation. If H = ( S 0 , S 1 , . . . ) denotes a multi-lev el seg-
mentation (with each S i denoting no w the collection of in-
terv als at the i th segmentation le vel), a pair set P can be
generated by matching each segment at le vel i to its max-
imally ov erlapping segment at each le vel j < i . The sim-
plified case of a two-le vel hierarchy H = ( S 0 , S 1 ) yields
P = ( | s | , | t | ) | t ∈ S 1 ∧ s = ar gmax s ∈ S 0 | s ∩ t | ,
where | s | denotes the duration of interv al s , and | s ∩ t | de-
notes the ov erlap duration between intervals s and t . Ev al-
uating ρ on each such pair captures ho w ev enly the hierar-
chy di vides segments from one le vel to the next.
3.6 Extension 3: Balance
Equation (1) captures a form of re gularity where durations
are related by simple ratios. This dif fers from previous no-
tions of regularity , which were designed to fa v or segments
of equal duration [5]. This notion can be recov ered by re-
placing the min normalization in eq. (1) by max :
β ( d 1 , d 2 ) := gcd { d 1 , d 2 }
max { d 1 , d 2 } . (5)
Equation (5) thus captures the balance of d 1 and d 2 : a
score of 1 is only achie ved when d 1 = d 2 , a score of 1 / 2 is
achie ved when the y are related by a factor of 2, and so on.
In general, β ( d 1 , d 2 ) ≤ ρ ( d 1 , d 2 ) , and it otherwise inherits
the boundedness, scale-in variance, and integer reciprocal
properties noted abov e. Repeating the calculations behind
table 1 for ˜
β results in an expected v alue of 0.216 for uni-
formly random durations and w = 0 . 5 .
As abov e, this also gi ves rise to an aggre gate pairwise
score B ( S ) , sampled versions ˜
β and ˜
B , and labeled and
hierarchical v ariations.
4. EXPERIMENTS
The proposed metrics are e valuated on the reference anno-
tations provided by a v ariety of commonly used structure
analysis datasets spanning multiple genres:
Beatles (TUT) 174 Beatles songs using the TUT segmen-
tations [13] and Isophonics beat annotations [14].
HarmonixSet 912 popular songs with segment and beat
annotations [15].
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
38
Figur e 2 . Segment durations for each dataset.
Jazz Structur e Dataset (JSD) 340 tracks [16]. For the
labeled metrics, the chorus and theme counter fields are
discarded from segment label strings, and segments la-
beled silence are treated as mutually distinct.
Jazz A udio-aligned Harmony (J AAH) 113 tracks with
labels deri ved from the parts annotations [17].
Real-world Computing (R WC) 211 tracks (100 popular ,
61 classical, 50 jazz) [18]. For labeled metrics, segments
labeled as "nothing" are treated as mutually distinct, and
labels are simplified by discarding parenthetical v aria-
tions ( e.g . , “chorus A (+1)” 7→ “c horus A” ).
SALAMI 1359 tracks from the publicly a vailable
dataset [10], consisting of 4486 annotations (2243 up-
per , 2243 lo wer). Sections labeled as "Z" or "silence"
are treated as mutually distinct for labeled metrics, and
v ariation markers are discarded ( e.g . , A 0 7→ A ).
Figure 2 illustrates the distrib ution of segment durations
for each dataset.
The e valuation seeks to e xplore the follo wing questions:
1. Ho w do the absolute time metrics ( ˜ ρ , ˜
β ) dif fer from
the musical time metrics ( ρ , β )?
2. Do structure annotations exhibit re gularity and/or
balance? Does this vary with genre?
3. Are multi-le vel segmentations re gular across lev els?
In service of the first question, we compared scores de-
ri ved from absolute time (using the approach described
in section 3.2) to the simpler forms deri ved from inte ger-
v alued durations measured in beats. This analysis is re-
stricted to the datasets with reference beat annotations:
Beatles, HarmonixSet, J AAH, and R WC. Each segment
boundary is mapped to its nearest beat, and segment du-
rations d are measured in beats between the start and end
boundaries. A preliminary study rev ealed sensiti vities to
rounding error in beat position identification, which were
resolved by including a maximization o ver { d − 1 , d, d + 1 } .
The (Pearson) correlation was then computed between the
musical-time and absolute-time metrics for each dataset.
For the second question, unlabeled and labeled forms
of the absolute time metrics were computed. As a point of
comparison, metrics were also computed under restriction
to adjacent segments [9], denoted here as R S , B S , etc .
T able 2 . Mean regularity and balance scores using musical
time, both unlabeled ( R, B ) and labeled ( R L , B L ).
R R L B B L
Beatles (TUT) 0.681 0.847 0.459 0.834
Harmonix 0.728 0.799 0.524 0.731
J AAH 0.741 0.878 0.488 0.869
R WC Classical 0.599 0.789 0.391 0.765
R WC Jazz 0.914 0.949 0.789 0.945
R WC Popular 0.820 0.958 0.587 0.945
T able 3 . Mean re gularity and balance scores using abso-
lute time, both unlabeled ( ˜
R, ˜
B ) and labeled ( ˜
R L , ˜
B L ).
˜
R ˜
R L ˜
B ˜
B L
Beatles (TUT) 0.704 0.820 0.394 0.805
Harmonix 0.730 0.789 0.498 0.719
JSD 0.732 0.606 0.344 0.591
J AAH 0.646 0.793 0.411 0.784
R WC Classical 0.506 0.709 0.298 0.673
R WC Jazz 0.818 0.856 0.720 0.838
R WC Popular 0.791 0.941 0.560 0.925
SALAMI (upper) 0.776 0.719 0.373 0.619
SALAMI (lower) 0.875 0.889 0.684 0.840
For the third question, we restrict attention to the
SALAMI dataset, and ev aluate hierarchical regularity and
balance using the paired upper - and lower -le vel annota-
tions for each track.
5. RESUL TS
5.1 Musical time vs. absolute time
T able 2 reports the av erage value for the re gularity and bal-
ance metrics on each of the datasets listed abov e for which
segment durations can be reliably measured in beats. As
should be expected, the labeled forms are generally sub-
stantially higher than the unlabeled forms. Each dataset
exhibits high labeled re gularity (significantly abov e 0.5),
as well as high labeled balance, indicating that similarly la-
beled segments do consistently span equi v alent durations.
T able 3 summarizes the absolute-time metrics across all
datasets, and Figure 3 illustrates the correlation between
these and the musical time data reported in table 2. The
correlations are generally high (abov e 0.6), with a few no-
table exceptions in the jazz and classical datasets. These
exceptions may be e xplained by the tempo distributions
of each dataset, illustrated in fig. 4. Recall that the ab-
solute time metric uses a tolerance windo w of 0.5 sec-
onds, equi valent to one beat at 120BPM. If a track is much
slo wer— e.g . , R WC Classical with median tempo of 87.1,
or R WC Jazz with median tempo of 89.4—the maximiza-
tion in eq. (3) will not cov er a full beat, so a larger windo w
may be warranted. Ho we ver , note that if the tempo is sta-
ble , this becomes less of an issue because absolute- and
musical-time are approximately proportional, which is ex-
ploited by the scale-in v ariance property of ρ .
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
39
R egularity (unlabeled)
R egularity (labeled)
Balance (unlabeled)
Balance (labeled)
Beatles (TUT)
Har monix
J A AH
R WC Classical
R WC Jazz
R WC P opular
0.73 0.80 0.85 0.81
0.91 0.93 0.95 0.95
0.59 0.66 0.85 0.70
0.55 0.70 0.59 0.74
0.48 0.38 0.82 0.36
0.74 0.80 0.86 0.84 1.0
0.5
0.0
0.5
1.0
Figur e 3 . Pearson correlation between musical- and
absolute-time metrics for each dataset.
Figur e 4 . T empo deri ved from reference beat annotations.
Each point corresponds to the mean tempo for one record-
ing. 120BPM is marked in red as a reference point.
Figure 5 illustrates the distrib utions of tempo stability ,
measured as the standard de viation of inter-beat-interv al.
Datasets with high tempo stability tend to exhibit high cor -
relation in fig. 3 e ven when the y contain many lo w-tempo
tracks ( e.g . , Beatles, Harmonix, and R WC Pop).
5.2 Unlabeled and labeled regularity
Figure 6 illustrates the relationship between labeled and
unlabeled regularity metrics. Consistent with the summary
in table 3, the unlabeled regularity scores are generally
quite dispersed, while the labeled scores ske w higher , con-
firming that segments belonging to dif ferently labeled sec-
tions may not conform to regular duration relationships.
T wo exceptions to this observ ation are JSD and SALAMI
(upper). In both cases, labeled regularity decreases from
the unlabeled scores. These cases may be explained by
the use of short silence segments, which di vide e venly into
most other segments, contrib uting many lar ge values to the
Figur e 5 . T empo stability for each dataset, as measured by
the standard de viation of local tempo deriv ed from inter-
beat interv als in the reference annotations. Each point rep-
resents the standard de viation of tempo for one recording.
Beatles (TUT)
Har monix
JSD
J A AH
R WC Classical
R WC Jazz
R WC P opular
S AL AMI (upper)
S AL AMI (lower)
Figur e 6 . Labeled vs. unlabeled re gularity metrics for
each annotation in each dataset.
a verage in eq. (2). In the labeled re gularity calculation,
each silence segment is treated as distinct, eliminating this
source of inflation. Segments of this nature are less pre v a-
lent in the other datasets ( e.g . , R WC or J AAH).
T able 4 summarizes the results of regularity and balance
when computed on adjacent segment pairs. While there are
clear regularity trends, confirming the prior w ork of Smith
and Goto, the ef fect is not generally as prev alent as the
label-agreement results reported in table 3.
5.3 Balance vs. Regularity
Figure 7 illustrates the distrib ution of the dif ference be-
tween labeled regularity and labeled balance in each
T able 4 . Sequential regularity and balance metrics in both
musical time ( R S , B S ) and absolute time ( ˜
R S , ˜
B S ).
R S ˜
R S B S ˜
B S
Beatles (TUT) 0.666 0.651 0.399 0.398
Harmonix 0.720 0.696 0.500 0.483
JSD — 0.729 — 0.479
J AAH 0.786 0.699 0.563 0.495
R WC Classical 0.604 0.521 0.380 0.311
R WC Jazz 0.935 0.872 0.818 0.780
R WC Popular 0.839 0.806 0.592 0.562
SALAMI (upper) — 0.746 — 0.420
SALAMI (lo wer) — 0.882 — 0.753
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
40
Figur e 7 . The distributions of dif ference between labeled
regularity and balance: ∆ = ˜
R L − ˜
B L for each dataset.
Figur e 8 . Hierarchical scores on SALAMI.
dataset. The balance scores cannot e xceed the regular -
ity scores—so the dif ference is non-negati ve—though each
dataset does exhibit v ery high correlation between reg-
ularity and balance: all correlation coef ficients exceed
0.95. While some datasets generally tend to match bal-
ance and regularity (Beatles, JSD, J AAH, R WC Jazz and
Pop), others div erge substantially (Harmonix, R WC Clas-
sical, SALAMI). This demonstrates that regularity and bal-
ance are indeed distinct qualities of segmentation.
5.4 Hierarchical r egularity
Figure 8 illustrates the distribution of hierarchical re gular-
ity and balance scores on the SALAMI dataset. As ex-
pected, the balance scores tend to be lo w due to the shorter
duration of segments in the lo wer le vel annotations.
Interestingly , the regularity scores are generally quite
high, with a median value of 0.969. This can be inter -
preted broadly as confirming that upper -lev el segments are
comprised of whole repetitions of lo wer-le vel se gment du-
rations. While this may be intuiti vely expected gi v en the
annotation rules, it is not an obvious conclusion from the
single-le vel analyses in the pre vious section. Figure 6 il-
lustrates that lo wer-le vel se gmentations tend to be highly
regular ( ˜
R L ≈ 0 . 889 ) and highly balanced ( ˜
B L ≈ 0 . 840 ),
while upper -lev el segmentations are slightly less re gular
( ˜
R L ≈ 0 . 719 ) and often less balanced ( ˜
B L ≈ 0 . 619 ).
6. DISCUSSION
From the findings abov e, we can draw some conclusions
about the role of regularity in music structure analysis.
First, because these analyses are conducted on refer -
ence annotations (not model outputs), the results reflect the
pbeha vior of human annotators, and not algorithms. The
distrib ution plots in fig. 6 indicate that although the mean
regularity scores are generally high across datasets, there
is considerable v ariability across individual tracks. While
these results deri ve from the absolute time metrics, the high
correlation with the musical time metrics suggests that this
is generally not explained by tempo v ariation, and rather
reflects widespread and meaningful structural irregularity
in many datasets. This suggests that regularity , if taken as
a design principle in segmentation algorithms, should be
treated with some care to allo w for irregular segmentations
when warranted by the track in question.
Second, the discrepancy between labeled and unlabeled
metrics can be quite large (Beatles, Harmonix, R WC Clas-
sical and Pop). This corresponds to non-tri vial interactions
between the regularity and repetition principles (as related
to segment label agreement), which had not been identi-
fied in pre vious studies. Modeling and fruitfully exploiting
these interactions would be an interesting direction for fu-
ture work in structure analysis algorithms.
Third, some datasets e xhibit significant discrepancies
between regularity and balance (Harmonix, SALAMI).
This demonstrates that segment durations in f act exhibit
more complex patterns than simple equi v alence.
7. LIMIT A TIONS
The proposed methods are applicable to quantitati ve e v al-
uation of segmentations, but the y do exhibit some limita-
tions. First, the absolute time definition does appear to ex-
hibit sensiti vity to tempo v ariation, in particular as it relates
to the choice of tolerance windo w . In situations where high
tempo v ariation may be expected, it may be preferable to
either apply the musical time formulation using estimated
beat positions (if they are reliable), or adapt the tolerance
windo w to fit the (estimated) tempo of the track.
Second, short segments may artificially inflate scores by
being easily di visible into long segments. This is partially
addressed by the labeled extension, as short se gments tend
to be sporadic and unrelated to the majority of a track, e.g . ,
a short silence segment at the be ginning or end.
Finally , the proposed metrics do not easily lend them-
selves to dif ferentiable formulations which may be in-
tegrated as learning objecti v es or penalties in current
gradient-based learning frame works. While it may be pos-
sible to do so, e.g. , by pre-computing a look-up table
of pairwise duration comparisons, other difficulties may
arise in adapting the ideas into practical segmentation al-
gorithms. Still, the proposed metrics may be more easily
integrated as post-processing steps, e .g. , to identify mean-
ingful le vels to include in a multi-le vel se gmentation, or to
select among a collection of proposed segmentations gen-
erated by an ensemble of methods.
8. A CKNO WLEDGMENTS
The author thanks Qingyang (T om) Xi and Meinard Müller
for helpful discussions and early feedback.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
41
9. REFERENCES
[1] G. Peeters, “Deriving musical structures from sig-
nal analysis for music audio summary generation:
"sequence" and "state" approach, ” in Computer
Music Modeling and Retrieval, International Sym-
posium, CMMR 2003,Montpellier , F rance , May
26-27, 2003, Re vised P apers , ser . Lecture Notes
in Computer Science, U. K. W iil, Ed., v ol. 2771.
Springer , 2003, pp. 143–166. [Online]. A v ailable:
https://doi.or g/10.1007/978- 3- 540- 39900- 1\_14
[2] J. Paulus, M. Müller , and A. Klapuri, “State
of the art report: Audio-based music structure
analysis. ” in Pr oceedings of the 11th International
Society for Music Information Retrieval Confer ence .
ISMIR, Aug. 2010, pp. 625–636. [Online]. A v ailable:
https://doi.or g/10.5281/zenodo.1417289
[3] O. Nieto, G. J. Mysore, C.-i. W ang, J. B. L. Smith,
J. Schlüter , T . Grill, and B. McFee, “ Audio-based mu-
sic structure analysis: Current trends, open challenges,
and applications, ” T ransactions of the International So-
ciety for Music Information Retrieval , Dec 2020.
[4] G. Sargent, F . Bimbot, and E. V incent, “ A
regularity-constrained viterbi algorithm and its ap-
plication to the structural segmentation of songs. ”
in Pr oceedings of the 12th International Society
for Music Information Retrieval Confer ence . IS-
MIR, Oct. 2011, pp. 483–488. [Online]. A v ailable:
https://doi.or g/10.5281/zenodo.1415950
[5] ——, “Estimating the structural se gmentation of
popular music pieces under regularity constraints, ”
IEEE/A CM T ransactions on A udio, Speech, and Lan-
guag e Pr ocessing , vol. 25, no. 2, pp. 344–358, 2017.
[6] A. Marmoret, J. E. Cohen, and F . Bimbot, “Barwise
music structure analysis with the correlation block-
matching segmentation algorithm, ” T r ansactions of the
International Society for Music Information Retrieval ,
Nov 2023.
[7] B. McFee and D. P . W . Ellis, “Learning to segment
songs with ordinal linear discriminant analysis, ” in
2014 IEEE International Confer ence on Acoustics,
Speech and Signal Pr ocessing (ICASSP) , 2014, pp.
5197–5201.
[8] A. Maezaw a, “Music boundary detection based on a
hybrid deep model of nov elty , homogeneity , repetition
and duration, ” in ICASSP 2019 - 2019 IEEE Inter-
national Confer ence on Acoustics, Speech and Signal
Pr ocessing (ICASSP) , 2019, pp. 206–210.
[9] J. B. L. Smith and M. Goto, “Using priors to improve
estimates of music structure. ” in Pr oceedings of the
17th International Society for Music Information
Retrieval Confer ence . ISMIR, Aug. 2016, pp.
554–560. [Online]. A v ailable: https://doi.org/10.5281/
zenodo.1416916
[10] J. B. L. Smith, J. A. Burgo yne, I. Fujinaga,
D. D. Roure, and J. S. Do wnie, “Design and
creation of a lar ge-scale database of structural
annotations. ” in Pr oceedings of the 12th International
Society for Music Information Retrieval Confer ence .
ISMIR, Oct. 2011, pp. 555–560. [Online]. A v ailable:
https://doi.or g/10.5281/zenodo.1416884
[11] C. Raf fel, B. McFee, E. J. Humphrey , J. Salamon,
O. Nieto, D. Liang, and D. P . W . Ellis, “mir_ev al: A
transparent implementation of common mir metrics. ”
in Pr oceedings of the 15th International Society for
Music Information Retrieval Confer ence . ISMIR,
Oct. 2014, pp. 367–372. [Online]. A v ailable: https:
//doi.or g/10.5281/zenodo.1416528
[12] B. McFee and K. Kinnaird, “Improving structure
e valuation through automatic hierarch y expansion, ”
in Pr oceedings of the 20th International Society for
Music Information Retrieval Confer ence . ISMIR,
Nov . 2019, pp. 152–158. [Online]. A v ailable: https:
//doi.or g/10.5281/zenodo.3527764
[13] J. Paulus, “Improving markov model based music
piece structure labelling with acoustic information. ”
in Pr oceedings of the 11th International Society for
Music Information Retrieval Confer ence . ISMIR,
Aug. 2010, pp. 303–308. [Online]. A v ailable: https:
//doi.or g/10.5281/zenodo.1416732
[14] C. Harte, “T o wards automatic e xtraction of harmony
information from music signals, ” Ph.D. dissertation,
Department of Electronic Engineering, Queen Mary ,
Uni versity of London, 2010.
[15] O. Nieto, M. McCallum, M. Davies, A. Robertson,
A. Stark, and E. Egozy , “The Harmonix Set: Beats,
do wnbeats, and functional segment annotations of
western popular music, ” in Pr oceedings of the 20th
International Society for Music Information Retrieval
Confer ence . ISMIR, No v . 2019, pp. 565–572.
[Online]. A v ailable: https://doi.org/10.5281/zenodo.
3527870
[16] S. Balke, J. Reck, C. W eiSS, J. AbeSSer , and
M. Müller , “JSD: A dataset for structure analysis in
jazz music, ” T ransactions of the International Society
for Music Information Retrieval , No v 2022.
[17] V . Eremenko, E. Demirel, B. Bozkurt, and X. Serra,
“ Audio-aligned jazz harmony dataset for automatic
chord transcription and corpus-based research, ” in
Pr oceedings of the 19th International Society for
Music Information Retrieval Confer ence . ISMIR,
Sep. 2018, pp. 483–490. [Online]. A v ailable: https:
//doi.or g/10.5281/zenodo.1492457
[18] M. Goto, H. Hashiguchi, T . Nishimura, and R. Oka,
“R WC music database: Popular , classical and
jazz music databases. ” in Pr oceedings of the 3r d
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
42
International Confer ence on Music Information Re-
trieval . ISMIR, Oct. 2002. [Online]. A v ailable:
https://doi.or g/10.5281/zenodo.1416474
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
43
ON THE DE-DUPLICA TION OF THE LAKH MIDI D A T ASET
Eunjin Choi 1 Hyerin Kim 2 Jiw oo Ryu 2
J uhan Nam 1 Dasaem Jeong 2
1 Graduate School of Culture T echnology , KAIST , South K orea
2 Department of Art & T echnology , Sogang Uni versity , South K orea
{jech,juhan.nam}@kaist.ac.kr, {kime0225, clayryu338}@gmail.com, {dasaemj}@sogang.ac.kr
ABSTRA CT
A lar ge-scale dataset is essential for training a well-
generalized deep-learning model. Most such datasets are
collected via scraping from v arious internet sources, in-
e vitably introducing duplicated data. In the symbolic mu-
sic domain, these duplicates often come from multiple user
arrangements and metadata changes after simple editing.
Ho wev er , despite critical issues such as unreliable train-
ing e valuation from data leakage during random splitting,
dataset duplication has not been extensi v ely addressed in
the MIR community . This study in v estigates the dataset
duplication issues reg arding Lakh MIDI Dataset (LMD),
one of the lar gest publicly a vailable sources in the sym-
bolic music domain. T o find and ev aluate the best retrie val
method for duplicated data, we employed the Clean MIDI
subset of the LMD as a benchmark test set, in which dif-
ferent versions of the same songs are grouped together . W e
first e valuated rule-based approaches and pre vious sym-
bolic music retrie val models for de-duplication and also in-
vestig ated with a contrastiv e learning-based BER T model
with v arious augmentations to find duplicate files. As a re-
sult, we propose three dif ferent versions of the filtered list
of LMD, which filters out at least 38,134 samples in the
most conserv ativ e settings among 178,561 files.
1. INTR ODUCTION
As data-dri ven approaches become mainstream and huge
generati ve neural networks become popular , the signifi-
cance of large-scale datasets is increasing. The tremen-
dous performance of generati ve models has been attrib uted
to massi ve dataset a vailability . For e xample, the av ailabil-
ity of billions of text-image pair data boosted the high-
quality image synthesis from text in the computer vision
domain [1]. Among se veral dataset collection strate gies,
web scraping [2], or synthesis [3] has emer ged as the most
practical method for assembling large-scale datasets, gi v en
the high costs associated with manual collection.
Ho wev er , collecting lar ge-scale datasets via web crawl-
ing can ine vitably cause problems, including priv ac y , copy-
© F . Author , S. Author , and T . Author . Licensed under a
Creati ve Commons Attribution 4.0 International License (CC BY 4.0).
Attribution: F . Author , S. Author, and T . Author , “On the de-duplication
of the Lakh MIDI dataset”, in Pr oc. of the 26th Int. Society for Music
Information Retrieval Conf ., Daejeon, South K orea, 2025.
right, and data duplication issues. In this paper , we focus
on data duplication, particularly highlighting issues in the
use of the Lakh MIDI Dataset (LMD) [2], which is widely
recognized as one of the lar ge-scale datasets in MIR. While
there are studies in other fields that discuss the do wnsides
of dataset duplication [4] and the benefits of de-duplication
[5], we found that such discussions are notably lacking in
the music information retrie val (MIR) community . One re-
lated work is the e xamination of dataset validity issues in
GTZAN [6, 7], and our work shares this intent by address-
ing the correct usage of LMD through de-duplication.
In particular , we ar gue that the current use of the LMD
in symbolic music generation impairs the v alidity of exist-
ing experiments. The issue arises from unreliable training
e valuations caused by data leakage during random splits.
Duplicates across training, v alidation, and test splits can
ske w e valuation metrics such as cross-entrop y loss, which
many pre vious studies used to assess their model perfor -
mance [8–13]. In the music generation domain, subjec-
ti ve e v aluation is costly , and objecti ve e valuation metrics
are often insuf ficient to fully capture the quality of gener-
ated music. Consequently , studies rely on v alidation loss
as a metric to claim non-ov erfitting and model ef fectiv e-
ness [13]. Duplicates can also bias listening-based e valua-
tions, particularly when the test split is used for condition-
ing. This is common practice in conditional music gener-
ation [10, 12–14]. If the duplication remains unaddressed,
random splits that lead to data leakage will continue to un-
dermine the reliability of e valuations.
Ho wev er , manually finding duplicates in a lar ge-scale
dataset such as LMD is virtually infeasible. Therefore, we
explored cleaning LMD using rule-based and neural ap-
proaches. Our contrib utions are as follows:
• How to clean the LMD?
W e ev aluated rule-based methods and e xisting symbolic
music retrie val models for duplicate detection. W e also
explored training an unsupervised contrasti v e BER T
model with augmentations to detect duplicates.
• How can we e valuate the de-duplication?
W e propose an ev aluation method to e v aluate duplicate
detection by utilizing the metadata of the Clean MIDI
subset of LMD, referred to as LMD-clean in this paper .
• How many duplicates exist in the dataset?
Using the duplicate detection method with the best per -
formance, we classify the duplicated files in the LMD.
44
Paper Dataset V ersion Split Evaluation
MuseGAN [15] LPD-5-matched N.A. .
MIDI-Sandwich2 [8] LPD-full N.A. NLL
LakhNES [9] LMD-full train, v alid PPL
PopMA G [10] LMD-matched random PPL, ˇ “ (
MMM [16] LMD-full N.A. .
PiRhDy [17] LMD-full N.A. .
MMD [18] LMD-matched N.A. .
Han et al. [19] LMD-full N.A. .
Musef ormer [11] LMD-full 8:1:1 PPL
MusicBER T [20] LMD-full N.A. .
FIGAR O [12] LMD-full 8:1:1 PPL, ˇ “ (
MIDI2V ec [21] LMD-matched 9:1 with CV .
YM2413-MDB [22] LMD-full 9:1:1 .
Sulun et al. [23] LM(P)D-full, matched N.A. .
Han et al. [24] LMD-full N.A. .
Anticipatory [13] LMD-full 87:6:6 PPL, ˇ “ (
text2midi [14] LMD-full (MidiCaps) N.A. ˇ “ (
T able 1 . Studies that used LMD for training. Studies men-
tioned the dataset duplication issue are bolded . LPD is the
piano roll version of LMD suggested by [15]. Split strate-
gies are N.A. when not described in the paper . CV means
cross-v alidation. The last column explains whether the au-
thors utilized the dataset during the e valuation. ˇ “ ( means
that the dataset is used in the listening test.
Even with the most conserv ati ve threshold of rejection,
we find 38,134 duplicated files to be filtered. Finally , we
present a filtering list of LMD from both our proposed
configuration and the most conserv ativ e threshold. 1
2. RELA TED WORKS
2.1 LMD and Related Studies
LMD [2] is a dataset released in 2016 with a method for ef-
ficiently matching lar ge-scale MIDI corpus collected from
the Internet to the Million Song Dataset (MSD) [25]. At the
time, the dataset was intended to be used for applications
such as content-based retrie val, corpus studies of music
structure and patterns, and transcription using paired au-
dio and MIDI. Di ver ging from its initially suggested appli-
cations, this dataset has become frequently used for train-
ing symbolic music generation models, primarily because
it is one of the lar gest av ailable symbolic music datasets.
Among the a vailable LMD v ersions, LMD-full, which con-
tains 178,561 files with unique MD5 hashes, is utilized as
the most popular for multi-instrumental pop music gener -
ation, and LMD-matched, which consists of 45,129 songs
matched with MSD is also used in se veral studies.
Since the release of LMD, larger -scale datasets based
on web scraping—such as MetaMIDI (MMD) [18] and
GigaMIDI [26]—ha ve emer ged. Notably , the recently re-
leased GigaMIDI dataset is a superset of LMD. Addition-
ally , the introduction of the MidiCaps dataset [27], which
incorporates LLM-generated text captions aligned with
LMD, has further expanded applications of LMD in te xt-
to-music generation [14] and music retrie val tasks [28, 29].
As sho wn in T able 1, most studies emplo yed the LMD-
full and split it for training without mentioning the split
1 All training and e valuation code is publicly a vailable: https://
github.com/jech2/LMD_Deduplication
strategies. Also, se v eral papers employed the NLL loss or
PPL v alues for their e v aluation. W e note that a fe w pa-
pers in T able 1 pointed out the duplication issues within
the dataset and remov ed the duplicated files using a rule-
based approach, such as MIDI encoding hash matching.
Ho wev er , we found that the rule-based approach is not suf-
ficient to remov e all of the duplicated files in the dataset,
which we will sho w in the ev aluation section. In this study ,
we used LMD-clean as our de-duplication study and test
dataset. This dataset contains 17,184 files or ganized by the
artist and song name in the directory and filename, which
we found to ha ve multiple MIDI files of the same song.
2.2 Dataset De-duplication and Related Issues
Recently , the need for dataset de-duplication has gained
attention across v arious domains. In computer vision, [4]
proposed a method to eliminate duplication by compress-
ing CLIP features using a contrasti ve feature compression
technique. In music information retrie val (MIR), [30] re-
cently introduced a method for detecting exact duplicates
in training data using audio-based music similarity met-
rics. Furthermore, in natural language processing, [5] in-
vestig ated the effects of dataset de-duplication by applying
exact substring matching and hash-based techniques. Their
findings sho wed that removing duplicates impro ves lan-
guage model performance, reduces training time, and lo w-
ers the rate of training data memorization without harm-
ing perplexity . In this work, we focus on the de-duplication
method for lar ge-scale symbolic music datasets.
2.3 Symbolic Music Understanding and Retrieval
W ith the adv ent of large-scale language understanding
models such as BER T [31] and B AR T [32] in the natu-
ral language domain, se veral counterparts ha ve been in-
troduced for symbolic music, including MidiBER T -Piano
[33], MusicBER T [20], and PianoB AR T [34]. Among
these, MusicBER T is a lar ge-scale pre-trained model
trained on LMD and a pri vate dataset. While earlier meth-
ods focused solely on the MIDI modality , recent ap-
proaches ha ve explored symbolic music retrie v al using
text, led by the introduction of CLaMP [35], which learns
joint embeddings of symbolic music and text through con-
trasti ve learning, using ABC notation as its input format.
CLaMP2 [28] extends this approach by enhancing the
symbolic encoder to support both ABC and MIDI formats
with multilingual support. CLaMP3 [29] further general-
izes the model to handle additional modalities, including
audio and images. W e used these models for our dataset
de-duplication task by le veraging their embeddings.
3. DUPLICA TION TYPES IN THE LMD
W e describe the types of duplicated MIDI files observed in
LMD-clean. In a broad sense, if we focus on music gen-
eration, all arrangements of the same song should be de-
fined as duplication (i.e., the same song cannot be in dif fer-
ent splits). Ho wev er , in tasks such as music arrangement,
dif ferent arrangements can be considered as dif ferent data
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
45
samples. Considering this aspect, we separated the dupli-
cation type into two cases: hard duplication and soft dupli-
cation 2 . This work focuses on detecting hard duplication.
3.1 Hard Duplication from Similar Arrangement
W e define hard duplication as files that share identical sec-
tions of arrangements with minor dif ferences. These dif-
ferences include instrument mapping or order , tempo, start
of fset, file length, missing tracks, or note-lev el alterations
such as pitch, duration, or velocity . Melodic variations may
also appear , such as added ornamentation or changes in the
number of chord tones played by a specific instrument. K ey
transpositions are considered hard duplication when other
musical elements remain nearly identical b ut are treated as
soft duplication if accompanied by significant stylistic or
structural changes. W e assume that some hard duplicates
were likely collected as users modified and re-uploaded e x-
isting files originally created by other arrangers.
3.2 Soft Duplication from Differ ent Arrangement
In soft duplication, MIDI files preserve essential musical
elements such as melody , harmon y , ornaments, and instru-
mentation b ut differ in arrangement style, reflecting the di-
verse styles of indi vidual arrangers. Such duplication in-
cludes cases where the core melody remains unchanged or
similar , b ut accompaniment styles (e.g., arpeggio, waltz),
pitch ranges, and ov erall structure v ary considerably . In
cases of extremely dif ferent arrangements, these v ariations
might lead a listener to percei ve them as distinct songs.
4. DUPLICA TE DETECTION: D A T ASET AND
CONVENTION AL APPRO A CHES
Along with the dataset for e valuation, we first e xplored
se veral approaches for identifying duplicates, including
simple rule-based methods and pre-trained symbolic mu-
sic retrie val models.
4.1 Dataset
W e used LMD-clean as our ev aluation benchmark to as-
sess ho w well each method detects duplicates within the
dataset. LMD-clean is org anized by artist folders and
song filenames, where duplicate instances of the same
song by the same artist are labeled with v ariations in
the filenames (e.g., Dancing Queen.mid , Dancing
queen.2.mid ). According to this metadata, 10,355 out
of 17,184 files in LMD-clean are considered duplicates.
4.2 Rule-based Appr oach
The follo wing rule-based methods serve as baselines for
identifying duplicates in the dataset. W e assumed that hard
duplicated samples share highly similar MIDI-le vel fea-
tures. Based on this assumption, we explored se v eral meth-
ods aimed at detecting and filtering out files with identical
or nearly identical features at the beat or pitch le vel.
2 W e sho w examples of duplication types in the companion website.
4.2.1 MIDI Encoding Hash
As discussed in Section 2.1, some studies [11, 20, 23] em-
ployed a hash-based approach to detect duplicated MIDI
files with dif ferent metadata. Here, we used the file de-
duplication code of MusicBER T [20]. The string versions
of Octuple representations are encoded according to the
MD5 hash v alue, and the hash values of all MIDI files from
LMD-clean are compared.
4.2.2 Beat P osition Entr opy
In our preliminary study , we found there are man y dupli-
cates that ha ve exact ly the same music b ut with different
instrument mapping or track order . T o detect these dupli-
cates, we applied a simple method that checks the distrib u-
tion of note position within a bar using the MIDI encoding
scheme of [36]. W e computed entropy v alues from note
position distrib utions at a 16th-note resolution. Files with
identical entropy v alues were identified as hard duplicates.
4.2.3 Chr oma-DTW
T o detect duplicates with similar pitch content, we mea-
sured the chroma-le vel distance between MIDI files. Piano
roll-based chromagrams were first generated and aligned
by transposing them with the highest pitch occurrence
across files. Dynamic T ime W arping (DTW) was then ap-
plied to measure the similarity between aligned chroma-
grams. T o reduce the computational cost of applying DTW
to the entire dataset, we first computed pitch histograms
for all files and measured the pairwise Kullback-Leibler
(KL) di ver gence. For each file, we selected the top 250
candidates with the lo west KL div ergence and then applied
chromagram-based DTW to these candidates. Although
this approach discards temporal information, it serves as
a rough prefiltering step.
4.3 Pre vious Symbolic Music Embedding Models
W e utilized pre-trained MusicBER T [20] and CLaMP
model series [28, 29, 35] since the y support multi-
instrumental MIDI. For MusicBER T , we used the pre-
trained MusicBER T -small and MusicBER T -base mod-
els for inference. For CLaMP , we used the pre-trained
CLaMP-512, CLaMP-1024, CLaMP2 and CLaMP3 mod-
els. Since the CLaMP-512 and CLaMP-1024 models use
XML for input files and ABC notation for their internal
data representation, the MIDI files are first con v erted using
Musescore batch processing and then con v erted with the
XML to ABC con v ersion algorithm. For CLaMP 2 and 3,
we con v erted MIDI to MTF , their MIDI encoding scheme.
5. DUPLICA TE DETECTION: A CONTRASTIVE
LEARNING-B ASED APPR O A CH
Pre vious pre-trained symbolic embedding models were
not originally trained for duplicate detection. Inspired by
[37, 38] that ev aluated the rob ustness of audio or music em-
bedding models against perturbations such as pitch shift,
we explore whether training with such perturbations w ould
improv e song identification despite v ariations. T o this end,
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
46
Lev el A ugmentation V alue
Note Onset Shift (-2, 2)
Note Duration Shift (-4, 4)
Note V elocity Shift (-3, 3)
T rack Pitch Octav e Shift (-24, 24)
T rack Inst Order Shuf fle
T rack Inst Mapping Except Drum
T rack Inst Drop Less than 50%
T rack Bar Drop 15%
Segment Bar Shift (1, 4)
Segment Note Drop 15%
Segment Pitch T ranspose (-6, 6)
T able 2 . List of augmentations for generating MIDI v aria-
tions. V alues in braces are in the token le vel (inclusi ve).
we de veloped a BER T -based model and applied v arious
augmentations to the positi ve samples in contrasti ve learn-
ing, which we refer to as CAugBER T .
5.1 Data Representation
W e used the LMD-full dataset, excluding all files present
in LMD-clean by matching MD5 hash v alues. The result-
ing dataset is referred to as LMD-filtered. W e randomly
split the LMD-filtered into 98:1:1 ratio and pre-processed
the MIDI files with Octuple encoding using MidiT ok [39].
Although we initially held out the remaining 1% for test-
ing, we did not use it in our final experiments, as we chose
to e valuate our model on LMD-clean instead.
5.2 Data A ugmentation
T o reflect the types of variations described in Section 3.1,
we constructed positi ve pairs for contrasti ve learning us-
ing dif ferent augmentations, referred to as MIDI variation
augmentation. Each augmentation of MIDI v ariation is in-
dependently and randomly applied during training. The de-
tails of these augmentations are provided in T able 2. In ad-
dition, follo wing [24], we used neighbor segments, which
are dif ferent parts from the same piece, as positiv e pairs
for training. Each piece was first se gmented into chunks
of 1024 tokens, and then a random se gment was selected.
MIDI v ariation augmentation was also applied to neighbor
segments to enhance rob ustness.
5.3 Model Description
The implementation of CAugBER T is based on the code
from [24], which applies contrasti ve learning to a BER T
architecture. T o align with the parameter settings of
MusicBER T -small, we used a 4-layer transformer with a
sequence length of 1024, hidden size and vocab ulary em-
bedding size of 512, and a feedforward dimension of 2048.
W e use a total batch size of 64 across two A6000 GPUs.
For mask ed language modeling (MLM), we adopted the
same element, compound, and bar -lev el masking strategies
used in MusicBER T . Contrasti v e learning was guided by
the NT -Xent loss [40]. The final loss was computed as
a weighted sum of the MLM and contrasti ve (NT -Xent)
losses, with weights of 0.3 and 1.0, respecti vely .
During training, each encoded MIDI segment w as aug-
mented using either MIDI v ariation or neighbor augmen-
tation described in Section 5.2, to maximize the di versity
of augmentations within each batch. For v alidation, fix ed
manual seed v alues were used to maintain consistent aug-
mentations across v alidation batches. T o e v aluate the ef-
fecti veness of the contrasti ve learning approach for dupli-
cation detection, we conducted an ablation study on the
contrasti ve loss, as presented in T able 3.
6. EV ALU A TION
W e ev aluated all approaches from two perspecti v es: (1)
Does the system rank duplicates as more similar than oth-
ers? (2) Ho w accurately does it identify true duplicates?
T o answer these questions, we utilized the metrics that are
commonly used for recommendation and retrie val systems.
6.1 Measuring Similarities
For MIDI Encoding Hash, the similarity between samples
was set as 1 when encoding hashes matched. Similarity of
beat position entropy w as computed by subtracting the ab-
solute dif ference in entropy from the maximum v alue of 1.
For Chroma-DTW , the similarity was calculated as 1 minus
the DTW distance. In the MusicBER T series, similarity
was measured using the cosine similarity of the a verage to-
ken embeddings from the T ransformer’ s final hidden layer .
For CAugBER T , we used the [CLS] token embedding from
the final hidden layer . All BER T -based models utilized
512-dimensional embedding. For CLaMP series, we used
a pre-trained 768-dimensional embedding where the last
hidden state was a verage pooled and passed through a pro-
jection layer , follo wing the code provided in [28, 29, 35].
6.2 Evaluation with Retriev al Metrics
T o ev aluate ho w well each method assigns higher simi-
larity scores to the duplicates, we adopt normalized Dis-
counted Cumulati ve Gain (nDCG) and Mean Reciprocal
Rank (MRR) as e valuation metrics. nDCG measures ho w
highly rele vant items are rank ed, assigning higher scores
when duplicates appear closer to the top of the retrie val
list. It is computed by normalizing Discounted Cumula-
ti ve Gain with the optimal ranking where all duplicates are
retrie ved at the highest possible ranks. F or each query in
LMD-clean, we assign rele v ance 1 to duplicates and 0 to
others when computing nDCG. MRR is defined as the a v-
erage of the in v erse ranks of the highest-ranked rele vant
item for the query . This corresponds to the a verage rank of
the highest similarity samples among the duplicates.
For the nDCG and MRR metrics, neural netw ork-based
approaches outperformed the rule-based methods. Among
them, the CLaMP model series consistently achie ved
higher scores than BER T -based models, with CLAMP3
sho wing the best ov erall performance. Since CLaMP mod-
els were specifically trained for retrie val tasks, the result is
consistent with its intended design.
Ho wev er , we observed that all approaches performed
belo w a certain upper bound. In particular , while neu-
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
47
(%) Contour Note Density Pitch Range Complexity
NM 63 . 36 76 . 63 46 . 70 48 . 67
P&L 37 . 44 10 . 11 0 . 41 50 . 59
LC-V AE-A 60 . 88 97 . 56 34 . 69 33 . 32
LC-V AE-SE 52 . 02 97 . 33 36 . 84 52 . 00
LC-Dif f 85 . 60 98 . 59 80 . 97 94 . 93
T able 1 : Pearson Correlation Coef ficient (PCC) between
tar get and decoded attributes.
condition on the noise le vel instead of the discrete dif fu-
sion step, which [7] later adopted for LDM-based symbolic
music generation. Thus, we are left with two continuous
conditioning signals to be passed onto the dif fusion model.
W e inject a and √ ¯ α t into ϵ θ through dedicated networks
(see Figure 1). First, we apply Sinusoidal Encoding (SE)
based on T ransformer positional embeddings [29]
Γ( u ) = sin ( ω i ( u )) , cos ( ω i ( u )) d/ 2
i =0 (7)
where ω i ( u ) = s u
b 2 i/d [7], with d ∈ N the (ev en) dimen-
sionality of the embedding, b ∈ R the base frequency ,
and s ∈ R a frequency scaling h yperparameter . The re-
sulting SE features are passed through a linear layer with
SiLU acti vations. Finally , we employ Feature-wise Linear
Modulation (FiLM) [30], where two fully-connected lay-
ers yield shifts and scales , respectiv ely , that modulate the
acti vations of the denoiser (see Figure 2). The two condi-
tioning branches run in parallel. This is equiv alent to learn-
ing a single af fine transformation, where scale and shift are
the sum of FiLM outputs from the attrib ute and noise lev el
conditioning networks.
T o enhance controllability over the generated samples,
we also apply Classifier -Free Guidance (CFG) [31] to the
noise prediction:
ˆ ϵ θ ( z t , ξ t , a ) = (1 + w ) ϵ θ ( z t , ξ t , a ) − w ϵ θ ( z t , ξ t ) , (8)
where ϵ θ ( z t , ξ t ) is the unconditional noise prediction and
w ∈ R ≥ 0 is the guidance scale. T o make CFG ef fecti ve,
the model must learn to predict noise both with and without
attrib ute conditioning. W e achiev e this through condition-
ing dr opout (depicted as 0 / 1 in Figure 1), i.e., setting the
outputs of the attrib ute conditioning network to zero with
a certain probability when e valuating (3).
3. EV ALU A TION
3.1 Dataset
The models are designed to learn pitch sequence represen-
tations from four -bar monophonic melodies. W e construct
a lar ge-scale dataset comprising melodies extracted from
176,581 MIDI files from the Lakh MIDI Dataset [32]. 1
First, we assess whether each MIDI file contains time
signature changes. If any are found, we segment the file
and retain only sections with a 4 / 4 time signature. Each
MIDI e vent is then quantized to the nearest sixteenth note.
A melody is defined as a sequence of pitches within
the standard 88-ke y piano range, played by an instrument
1 C. Raf fel, 2016, “The Lakh MIDI Dataset v0.1. ” [Online]. A v ailable:
https://colinraffel.com/projects/lmd
Contour Note Density Pitch Range Complexity
Uncond. V AE 41 . 44
NM 35 . 506 58 . 436 30 . 833 47 . 61
P&L 49 . 698 67 . 836 40 . 657 87 . 80
LC-V AE-A 30 . 197 29 . 450 30 . 257 32 . 435
LC-V AE-SE 29 . 161 30 . 124 31 . 274 30 . 166
LC-Dif f 19 . 299 20 . 559 31 . 695 17 . 51
T able 2 : Fréchet Music Distance [27].
mapped to a v alid MIDI program. A melody is considered
complete when a full measure of silence occurs. W e e xtract
only melodies spanning at least four bars and comprising at
least three distinct pitches. If multiple notes sound simul-
taneously , we follo w the approach proposed in [33] and
select only the highest-pitched note to ensure monophonic
sequences. Subsequently , four-bar se gments are extracted
using a stride of one bar .
For each melody thus e xtracted, we compute 13 musical
attrib utes, including those outlined in Section 3.2.
Melodies are encoded as sequences of N = 64 inte gers
in P = { 0 ,..., 129 } , where each element represents either
a MIDI note number ( 0 - 127 ) or one of two special tok ens:
note of f ( 128 ) and note hold ( 129 ). The dataset is divided
into training, validation, and test sets, with training data
augmented through transposition by a randomly selected
number of semitones within a range of ± 1 octa ve. The
final dataset, consisting of 10 , 126 , 676 unique melodies, is
publicly a v ailable. 2
3.2 Musical Attributes
As pre viously done in [25], we focus on four musical at-
trib utes: (i) Contour , which quantifies the melodic mov e-
ment in a sequence, measured by av eraging the pitch dif-
ferences between consecuti ve notes; (ii) Note Density ,
defined as the ratio between the number of notes in the
melody and the sequence length. It takes v alues in [0 , 1] ;
(iii) Pitch Range , defined as the dif ference between the
highest and lo west MIDI pitch v alues in the sequence, nor -
malized by the range of an 88 -ke y piano. It takes v al-
ues in 0 , 127
88 , where values abo ve one indicate a range
exceeding A0–C8; (iv) Rh ythm Complexity , ev aluated
using T oussaint’ s metrical comple xity measure [34], cor -
rected for the total number of notes in the sequence [26].
By definition, it takes on discrete v alues.
3.3 Unconditional Generative Model
As base unconditional model, we implement a β -V AE [35]
based on MusicV AE [33]. This model, pre viously used in
LDM-based symbolic music generation [7], also enables
direct comparison with existing attrib ute-regularized V AEs
(AR-V AEs) employing the same architecture [25, 26] (see
Section 3.5).
The encoder p ψ ( z | x ) consists of a two-layer bidirec-
tional LSTM network fed with four -bar pitch sequence rep-
resentations (see Section 3.1), followed by tw o linear lay-
2 M. Pettenò, Aug. 2024, “4 Bars Monophonic Melodies Dataset (Pitch
Sequence), ” Zenodo, doi: https://doi.org/10.5281/zenodo
.13369389
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
54
(a) NM (b) P&L (c) LC-V AE-A (d) LC-V AE-SE
Figur e 3 : Re gression plots comparing tar get and decoded Contour attrib utes across baseline methods.
Figur e 4 : Re gression plot comparing tar get and decoded
Contour attrib utes using LC-Dif f.
ers parameterizing the latent posterior . The hierarchical
decoder q ϕ ( x | z ) features tw o unidirectional LSTMs, with
the bottom-le v el netw ork autore gressi v ely estimating the
distrib ution o v er the sequence v alues via a softmax nonlin-
earity [33]. As such, each pitch sequence x ∈ P N is first
mapped onto a single latent code z ∈ R M , M = 256 , and
since decoding amounts to a ne xt tok en prediction task, the
standard β -V AE objecti v e [35]
L V AE = − E p ψ ( z | x ) [ l og q ϕ ( x | z ) ] + β D KL [ p ψ ( z | x ) ∥ p ( z ) ] ,
(9)
is implemented using cross-entrop y as reconstruction loss.
The unconditional model is trained for 40 , 000 iterations
on a single NVIDIA T itan R TX GPU with a batch size
of 512 . The objecti v e (9) is minimized using Adam, and
the learning rate is decreased e xponentially from 10 − 3 to
10 − 5 with a rate of 0 . 9999 . The h yperparameter β is an-
nealed e xponential ly from 0 to 10 − 3 , which encourages the
model to prioritize accurate sequence reconstruction dur -
ing the early part of the training. Similarly to [33], we
apply teacher forcing within the bottom-le v el decoder with
a probability follo wing a logistic schedule.
3.4 Conditional Diffusion Model
W ith latent codes being v ectors in R M , we implement a
DDIM model with a fully-connected denoiser netw ork. 3
Sho wn in Figure 2, the denoiser comprises an input layer
with 2048 linear units, f o l lo wed by three dense residual
blocks. Each residual block comprises tw o stacks of Lay-
erNorm, feature-wise modulation (responsible for joint at-
trib ute and time conditioning), SiLU, and a linear layer ,
plus a residual connection that shortcuts the input and out-
put of the block. Finally , the output is linearly projected
back onto R M .
3 Source code and audio e xamples are a v ailable at https://mpet
teno.github.io/controllable- latent- diffusion/
W e set the SE dimensionality to d = 128 . The attrib ute
and noise le v el conditioning netw orks ha v e 512 and 2048
units in the first linear layer and FiLM layers, respecti v ely .
In the forw ard process, β t follo ws a linear schedule
from 10 − 6 to 10 − 2 o v er T = 1000 steps. Con v ersely , the
number of sampling steps is set to T s = 100 .
W e train the model with an attrib ute conditioning
dropout probability of 20% . W e then apply CFG with a
guidance scale of w = 3 . 0 [31]. In our e xperiments, CFG
pro v ed fundamental to achie v e attrib ute re gularization.
The resulting denoiser netw ork has 43 . 1 million param-
eters, and con v er ges in just about 20 training epochs, half
the iterations required by the unconditional model.
3.5 AR-V AE Baseline Methods
F or comparison, we consider AR -V AEs [25, 26] with the
same architecture as the unconditional model described in
Section 3.3. AR-V AEs incorporate re gularization during
training by means of a supervised multi-task learning ap-
proach, with the goal of encoding the attrib ute a in the i -th
dimension z i of their latent spaces. This is achie v ed by
including an AR loss term in (9)
L AR-V AE = L V AE + γ L AR , (10)
where γ ≥ 0 is a tunable h yperparameter controlling the
strength of the re gularization.
Mezza et al. [26] propose the use of
L NM
AR = M AE ( z i , ˜ a ) , (11)
where M AE ( · , · ) denotes the mean absolute error , and ˜ a is
the z-score of a .
P ati and Lerch [25] introduced a re gularization term that
enforces a monotonic relationship between a and z i , i.e.,
L P&L
AR = M AE ( t an h ( δ D z ) , s i gn ( D a ) ) , (12)
where D z and D a are pairwise distance matrices between
z i and a of all samples in a batch, respecti v ely , and δ > 0
is a tunable h yperparameter . As in [25], we set γ = 1 and
δ = 10 . The remaining training details are the same as in
Section 3.3. F or bre vity , we will later refer to the former
AR method as “NM” and to the latter as “P&L. ”
3.6 LC-V AE Baseline Methods
Similarly to T ian and Engel [3], we implement LC through
a conditional V AE (cV AE) trained on the representations of
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
55
(a) NM (b) P&L (c) LC-V AE-A (d) LC-V AE-SE
Figur e 5 : Re gression plots comparing tar get and decoded Rh ythm Comple xity attrib utes across baseline methods.
Figur e 6 : Re gression plot comparing tar get and decoded
Rh ythm Comple xity attrib utes using LC-Dif f.
the base unconditional model (Section 3.3). The cV AE en-
coder consists of four linear layers with ReLU acti v ations,
follo wed by tw o Gating Mixing Layers (GML) that param-
eterize the innermost latent distrib ution. The decoder mir -
rors the encoder with four linear layers with ReLU acti v a-
tions, follo wed by an output GML. Except for using 2048
units in the fully-connected layers and M ′ = 128 latent
v ariables, the cV AE architecture is the same as in [3].
Let z ∈ R M be the latent representations of the uncon-
ditional model, z c ∈ R M ′ be the latent representations of
the cV AE, and a ∈ R the sequence attrib ute. The authors
of [3] considered binary labels and one-hot v ectors were
thus concatenated with z and z c . Instead, we deal with con-
tinuous attrib utes. W e implement tw o cV AE v ariants that
dif fer in ho w a is fed into the netw orks. In the first v ari ant,
later referred to as LC-V AE-A, we feed ˜
z = [ z T , a ] T to the
encoder , and ˜
z c = [ z T
c , a ] T to the decoder . In the second
v ariant, named LC-V AE-SE, we concatenate z and z c , re-
specti v ely , with the attrib ute SE, i.e., ˜
z = [ z T , Γ( a ) ] T and
˜
z c = [ z T
c , Γ( a ) ] T .
4. RESUL TS
4.1 Attrib ute-Contr olled Generation
T o e v aluate the controllability of the generati v e models un-
der scrutin y , we sample the tar get attrib utes uniformly in
the range of zero to the 99 th percentile of the attrib ute dis-
trib ution of the sequences in the test set. 4 These equally-
spaced v alues, which we refer to as tar g et attrib utes, are
fed to the respecti v e conditioning netw ork of LC-Dif f, suit-
ably transformed and plugged into the re gularized dimen-
4 Limiting the range to the 9 9 th percentile is meant to e xclude those
sequences with abnormally high attrib ute v alues. W e ar gue that these
sequences are spurious, and we att rib ute their e xistence to the choice,
borro wed from [33], of e xtracting melodies by naïv ely picking the highest
note at an y gi v en time.
sion z i of the AR-V AEs, and concatenated to the input v ec-
tor of the LC-V AE decoder netw orks.
T able 1 lists t he Pearson Correlation Coef ficients (PCC)
between the tar get attrib utes and those computed from the
generated sequences (the higher , the better). LC-Dif f con-
sistently outperforms the tw o AR-V AEs (NM and P&L)
and LC-V AEs (both with and without SE) for all attrib utes
considered. Notably , LC-Dif f is the only method among
those considered in the present study to yield correlation
scores higher than 80% across the board.
As for Contour , LC-Dif f achie v es a PCC of 85 . 60% ,
outperforming the ne xt-best model, NM, by o v er 22% . The
dif ference is less pronounced for Note Density , where LC-
Dif f ( 98 . 56% ) impro v es upon the second-best model by
just 1% . Nonetheless, LC-V AE-A and LC-V AE-SE al-
ready achie v e 97 . 56% and 97 . 33% , respecti v ely , suggest-
ing that constraining the generati v e model is v ery ef fecti v e
compared to AR methods when it comes to rendering the
desired number of notes. LC-Dif f also demonstrates sig-
nificant impro v ements in Pitch Range and Rh ythm Com-
ple xity . F or Pi tch Range, it achie v es a PCC of 80 . 97% ,
e xceeding NM ( 46 . 70% ) by 34 . 27% , while NM itself out-
performs LC-V AEs by approximately 10% . F or Rh ythm
Comple xity , LC-Dif f achie v es a remarkable 94 . 93% , sur -
passing LC-V AE-SE ( 52% ) by 42 . 93% .
Concerning AR models, while NM directly encodes the
(standardized) distrib ution onto the i th dimension of the
latent space, there is no a priori w ay to kno w the monotonic
relationship learned using the P&L re gularizati on in (12).
This e xplains the near -zero correlation observ ed for Pitch
Range, and, in general, the o v erall lo wer PCC.
Figures 3 through 6 sho w the re gression plots of Con-
tour and Rh ythm Comple xity . Figure 3 and Figure 4 illus-
trate the cas e of a continuous distrib ution, while Figure 5
and Figure 6 e x emplify a case where the attrib ute tak es on
inte ger v alues. Across both attrib utes, LC-Dif f is charac-
terized by a lo wer spread and a clear linear trend. In Fig-
ure 3, all baseline models sho w a tendenc y to produce e x-
cessi v ely high contour v alues, whereas LC-Dif f (Figure 4)
appears to mitig ate the issue. Lik e wise, Figure 5 re v eals
that all models b ut LC-Dif f (Figure 6) tend to f ail when the
tar get Comple xity v alues are lo w .
4.2 Data Fidelity
T o e v aluate the quality of the generated sequences, we use
the Fréchet Music Distance (FMD) [27], a metric that e x-
tends the f amily of Fréchet Inception Dist ance [36] and
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
56
(a) a = 1 . 0 → a g = 1 . 0 3 (b) a = 3 . 0 → a g = 2 . 9 5 (c) a = 6 . 0 → a g = 6 . 1 9
Figur e 7 : Examples of MIDI files generated by controlling the Contour attrib ute with LC-Dif f.
(a) a = 0 → a g = 0 (b) a = 1 5 → a g = 1 5 (c) a = 3 3 → a g = 3 3
Figur e 8 : Examples of MIDI files generated by controlling Rh ythm Comple xity with LC-Dif f. Quarter notes are indicated
by solid v ertical lines; odd pulses (strong) are indicated by dashed lines; e v en pulses (weak) are indicated by dotted lines.
Fréchet Audio Distance [37] to the symbolic music do-
main. FMD w as computed between 22 , 016 melodies from
the held-out test set and an equal number of generated se-
quences. T o pre v ent the FMD from measuring a spurious
di v er gence from the real att rib ute distrib ution, we condi-
tion the generation on the attrib utes of the reference se-
quences, rather than using e v enly-spaced control v alues as
in Section 4.1. By conditioning with attrib utes measured
from the test set, indeed, we aim to simultaneously com-
pare the fidelity of generated sequences and ho w well the y
conform to the desired attrib ute distrib ution.
T able 2 reports the results obtained using CLaMP 2
MIDI embeddings [38] (the lo wer , the better). F or com-
parison, we report the FMD between the reference test set
and the output of the unconditional V AE (see Section 3.3)
obtained by decoding 22 , 016 samples from N ( 0 , I ) .
The results presented in T able 2 demonstrate that the
proposed LC-Dif f model consistently achie v es the lo west
FMD v alues across most attrib utes, indicating superior per -
formance in generating samples that aligns more closely
with the statistical properties of real sequences. Notably ,
LC-Dif f outperforms all baselines in Contour ( 19 . 299 ),
Note Density ( 20 . 559 ), and Rh ythm Comple xity ( 17 . 51 ),
significantly impro ving o v er both AR-V AEs and LC-
V AEs. While LC-V AE-A achie v es the best Pitch Range
score ( 30 . 257 ), LC-Dif f remains competiti v e ( 31 . 695 ).
Ov erall, all LC methods outperform the unconditional
base model ( 41 . 44 ), sho wing that introducing post-hoc
control leads to more consistent and structured music gen-
eration, with better alignment to the desired attrib utes.
Finally , Figure 7 and Figure 8 illustrate the potential
di v ersity in the generated samples produced by LC-Dif f
when conditi o ne d on lo w , medium, and high v alues of
Contour and Rh ythm Comple xity , respecti v ely .
5. CONCLUSIONS
In this paper , we ha v e e xplored latent dif fusion through the
lens of Latent Constraints (LC), demonstrating the ef ficac y
of DDIM s as plug-and-play conditioning modules for sym-
bolic music generation. By k eeping the base generati v e
model fix ed, we trained dif fusion-based LC models (LC-
Dif f) capable of controlling a range of non-dif ferentiable
and continuous musical attrib utes, including contour , note
density , pitch ra ng e , and rh ythm comple xity . Our em-
pirical e v aluations re v eal that LC-Dif f significantly out-
performs attrib ute-re gularized V AEs and cV AE-based LC
methods in terms of both fidelity and controllability , with
absolute impro v ements of up to 12 . 65 in Fréchet Mu-
sic Distance and 43% in correlation between desired and
generated attrib utes. These results highlight the poten-
tial of denoising as a po werful tool for ad hoc f ader -lik e
control o v er mul tiple musical attrib utes along continuous
ax es, ef fecti v ely transforming a pre-trained unconditional
model into a controllable music generation system depend-
ing on the user’ s needs. Future w ork will focus on e x-
panding the library of LC-Dif f models to include a wider
range of musical attrib utes and e xploring the inte gration of
user interf aces for real-time control. Future e xperiments
could al so e xplore attrib ute-controlled input transforma-
tions by applying forw ard dif fusi on to encoded representa-
tions, rather than dra wing noise samples from the standard
normal prior . Furthermore, we aim to in v esti g a te the po-
tential for LC of other generati v e techniques, such as flo w
matching and consistenc y models.
6. REFERENCES
[1] J. Engel, M. Hof fman, and A. Roberts, “Latent con-
straints: Learning to g e nerate conditionally from un-
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
57
conditional generati ve models, ” in International Con-
fer ence on Learning Repr esentations , 2018.
[2] M. Dinculescu, J. Engel, and A. Roberts, “MidiMe:
Personalizing a MusicV AE model with user data, ” in
NeurIPS W orkshop on Machine Learning for Cr eativ-
ity and Design , 2019.
[3] Y . T ian and J. Engel, “Latent translation: Cross-
ing modalities by bridging generati ve models, ” arXiv
pr eprint arXiv:1902.08261 , 2019.
[4] R. Rombach, A. Blattmann, D. Lorenz, P . Esser ,
and B. Ommer , “High-resolution image synthesis
with latent dif fusion models, ” in Pr oceedings of the
IEEE/CVF confer ence on computer vision and pattern
r ecognition , 2022, pp. 10 684–10 695.
[5] S. W u and M. Sun, “Exploring the ef ficacy of pre-
trained checkpoints in text-to-music generation task, ”
in The AAAI-23 W orkshop on Cr eative AI Acr oss
Modalities , 2023.
[6] P . Jajoria and J. McDermott, “T e xt conditioned sym-
bolic drumbeat generation using latent dif fusion mod-
els, ” arXiv pr eprint arXiv:2408.02711 , 2024.
[7] G. Mittal, J. Engel, C. Hawthorne, and I. Simon, “Sym-
bolic music generation with dif fusion models, ” in Pr oc.
of the 22nd International Society for Music Informa-
tion Retrieval Confer ence (ISMIR) , 2021, pp. 468–475.
[8] M. Pasini, M. Grachten, and S. Lattner , “Bass accom-
paniment generation via latent dif fusion, ” in ICASSP
2024-2024 IEEE International Confer ence on Acous-
tics, Speech and Signal Pr ocessing (ICASSP) , 2024,
pp. 1166–1170.
[9] S. Li and Y . Sung, “MelodyDif fusion: Chord-
conditioned melody generation using a transformer-
based dif fusion model, ” Mathematics , vol. 11, no. 8,
2023.
[10] L. Min, J. Jiang, G. Xia, and J. Zhao, “Polyffusion: A
dif fusion model for polyphonic score generation with
internal and external controls, ” in Pr oc. of the 24th
International Society for Music Information Retrieval
Confer ence (ISMIR) , 2023, pp. 231–238.
[11] L. Kaw ai, P . Esling, and T . Harada, “ Attributes-a ware
deep music transformation. ” in Pr oc. of the 21st Inter-
national Society for Music Information Retrieval Con-
fer ence (ISMIR) , 2020, pp. 670–677.
[12] S.-L. W u and Y .-H. Y ang, “MuseMorphose: Full-song
and fine-grained piano music style transfer with one
transformer V AE, ” IEEE/A CM T ransactions on A udio,
Speech, and Languag e Pr ocessing , vol. 31, pp. 1953–
1967, 2023.
[13] M. E. Malandro, “Composer’ s Assistant 2: Interacti ve
multi-track MIDI infilling with fine-grained user con-
trol, ” in Pr oc. of the 25th International Society for Mu-
sic Information Retrieval Confer ence (ISMIR) , 2024,
pp. 438–445.
[14] G. Lample, N. Zeghidour , N. Usunier , A. Bordes,
L. Denoyer , and M. Ranzato, “Fader netw orks: Ma-
nipulating images by sliding attrib utes, ” in Advances
in Neural Information Pr ocessing Systems , 2017, pp.
5963–5972.
[15] H. H. T an and D. Herremans, “Music FaderNets: Con-
trollable music generation based on high-le vel features
via lo w-lev el feature modelling, ” in Pr oc. of the 21st
International Society for Music Information Retrieval
Confer ence (ISMIR) , 2020, pp. 109–116.
[16] J. Ho, A. Jain, and P . Abbeel, “Denoising dif fusion
probabilistic models, ” in Pr oc. of the 34th Interna-
tional Confer ence on Neur al Information Pr ocessing
Systems , 2020, pp. 1–12.
[17] J. Song, C. Meng, and S. Ermon, “Denoising dif fu-
sion implicit models, ” in International Confer ence on
Learning Repr esentations , 2021.
[18] J. Austin, D. D. Johnson, J. Ho, D. T arlow , and
R. v an den Berg, “Structured denoising dif fusion mod-
els in discrete state-spaces, ” in Advances in Neur al In-
formation Pr ocessing Systems , 2021, pp. 1–13.
[19] A. Lv , X. T an, P . Lu, W . Y e, S. Zhang, J. Bian, and
R. Y an, “GETMusic: Generating an y music tracks
with a unified representation and dif fusion framew ork, ”
arXiv pr eprint arXiv:2305.10841 , 2023.
[20] M. Plasser , S. Peter , and G. W idmer , “Discrete dif fu-
sion probabilistic models for symbolic music genera-
tion, ” in Pr oc. of the Thirty-Second International Joint
Confer ence on Artificial Intelligence , 2023.
[21] J. Zhang, G. Fazekas, and C. Saitis, “Composer style-
specific symbolic music generation using vector quan-
tized discrete dif fusion models, ” in 2024 IEEE 34th In-
ternational W orkshop on Machine Learning for Signal
Pr ocessing (MLSP) , 2024, pp. 1–6.
[22] ——, “Fast dif fusion GAN model for symbolic mu-
sic generation controlled by emotions, ” arXiv pr eprint
arXiv:2310.14040 , 2023.
[23] M. Zhang, L. J. Ferris, L. Y ue, and M. Xu, “Emotion-
ally guided symbolic music generation using dif fusion
models: The A GE-DM approach, ” in Pr oc. of the 6th
A CM International Confer ence on Multimedia in Asia ,
2024, pp. 1–5.
[24] Y . Huang, A. Ghatare, Y . Liu, Z. Hu, Q. Zhang, C. S.
Sastry , S. Gururani, S. Oore, and Y . Y ue, “Symbolic
music generation with non-dif ferentiable rule guided
dif fusion, ” arXiv pr eprint arXiv:2402.14285 , 2024.
[25] A. Pati and A. Lerch, “ Attrib ute-based regularization
of latent spaces for v ariational auto-encoders, ” Neural
Computing and Applications , v ol. 33, no. 9, pp. 4429–
4444, 2021.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
58
[26] A. I. Mezza, M. Zanoni, and A. Sarti, “ A latent rhythm
complexity model for attrib ute-controlled drum pattern
generation, ” EURASIP J ournal on A udio, Speech, and
Music Pr ocessing , vol. 2023, no. 1, 2023.
[27] J. Retko wski, J. Ste ¸ pniak, and M. Modrzejew-
ski, “Frechet music distance: A metric for gen-
erati ve symbolic music e v aluation, ” arXiv pr eprint
arXiv:2412.07948 , 2024.
[28] N. Chen, Y . Zhang, H. Zen, R. J. W eiss, M. Norouzi,
and W . Chan, “W a veGrad: Estimating gradients for
wa veform generation, ” in International Confer ence on
Learning Repr esentations , 2021.
[29] A. V aswani, N. Shazeer , N. Parmar , J. Uszkoreit,
L. Jones, A. N. Gomez, L. u. Kaiser , and I. Polosukhin,
“ Attention is all you need, ” in Advances in Neural In-
formation Pr ocessing Systems , vol. 30, 2017.
[30] E. Perez, F . Strub, H. de Vries, V . Dumoulin, and A. C.
Courville, “FiLM: V isual reasoning with a general con-
ditioning layer , ” in Pr oc. of the Thirty-Second AAAI
Confer ence on Artificial Intelligence , 2018, pp. 3942–
3951.
[31] J. Ho and T . Salimans, “Classifier -free diffusion guid-
ance, ” in NeurIPS 2021 W orkshop on Deep Gener ative
Models and Downstr eam Applications , 2021.
[32] C. Raf fel, “Learning-based methods for comparing se-
quences, with applications to audio-to-midi alignment
and matching, ” Ph.D. dissertation, Columbia Uni ver -
sity , 2016.
[33] A. Roberts, J. Engel, C. Raf fel, C. Hawthorne, and
D. Eck, “ A hierarchical latent v ector model for learn-
ing long-term structure in music, ” in Pr oc. of the 35th
International Confer ence on Mac hine Learning , 2018,
pp. 4364–4373.
[34] G. T oussaint, “ A mathematical analysis of African,
Brazilian, and Cuban cla ve rh ythms, ” in Bridges:
Mathematical Connections in Art, Music, and Science ,
2002, pp. 157–168.
[35] I. Higgins, L. Matthey , A. P al, C. P . Bur gess, X. Glo-
rot, M. M. Botvinick, S. Mohamed, and A. Lerchner ,
“beta-V AE: Learning basic visual concepts with a con-
strained v ariational framew ork. ” International Confer -
ence on Learning Repr esentations , v ol. 3, 2017.
[36] M. Heusel, H. Ramsauer , T . Unterthiner , B. Nessler ,
and S. Hochreiter , “GANs trained by a two time-scale
update rule con v erge to a local Nash equilibrium, ” in
Pr oc. of the 31st International Confer ence on Neural
Information Pr ocessing Systems , 2017, pp. 6629–6640.
[37] K. Kilgour , M. Zuluaga, D. Roblek, and M. Sharifi,
“Fréchet audio distance: A reference-free metric for
e valuating music enhancement algorithms, ” in Pr oc.
Interspeec h 2019 , 2019, pp. 2350–2354.
[38] S. W u, Y . W ang, R. Y uan, Z. Guo, X. T an, G. Zhang,
M. Zhou, J. Chen, X. Mu, Y . Gao, Y . Dong, J. Liu,
X. Li, F . Y u, and M. Sun, “CLaMP 2: Multi-
modal music information retrie val across 101 lan-
guages using lar ge language models, ” arXiv pr eprint
arXiv:2410.13267 , 2025.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
59
RADIF CORPUS: A SYMBOLIC D A T ASET FOR NON-METRIC IRANIAN
CLASSICAL MUSIC
Maziar Kanani
Uni versity of Gal way
[email protected]
Sean O’Leary
TU Dublin
[email protected]
J ames McDermott
Uni versity of Gal way
[email protected]
ABSTRA CT
Non-metric music forms the core of the repertoire in
Iranian classical music. Dastg ¯
ahi music serves as the un-
derlying theoretical system for both Iranian art music and
certain folk traditions. At the heart of Iranian classical mu-
sic lies the radif , a foundational repertoire that organizes
melodic material central to performance and pedagogy .
In this study , we introduce the first digital corpus rep-
resenting the complete non-metrical radif repertoire, co v-
ering all 13 existing components of this repertoire. W e
provide MIDI files (about 281 minutes in total) and data
spreadsheets describing notes, note durations, interv als,
and hierarchical structures for 228 pieces of music. W e
faithfully represent the tonality including quarter -tones,
and the non-metric aspect. Furthermore, we provide sup-
porting basic statistics, and measures of complexity and
similarity ov er the corpus.
Our corpus provides a platform for computational stud-
ies of Iranian classical music. Researchers might employ it
in studying melodic patterns, in vestigating impro visational
styles, or for other tasks in music information retrie v al,
music theory , and computational (ethno)musicology .
1. INTR ODUCTION
While ethnic music traditions from around the world ha ve
recently gained more attention in computational research,
many still lack the necessary datasets to support such stud-
ies. Iran, with its rich di versity of ethnic and folk musical
traditions, of fers great potential for computational analysis
that reflects its regional musical identity .
In this work, we take a step to ward addressing this g ap
by introducing a dataset specifically focused on Iranian
non-metric classical music, aiming to support and inspire
future studies in this area. W e begin by introducing Iranian
classical music and its core repertoire, the radif . After re-
vie wing previously published datasets, we present our o wn
dataset in detail. Finally , we provide a statistical and vi-
sual ov ervie w of the dataset, which can serve as a useful
reference for researchers and practitioners.
© M. Kanani, S. O’Leary , and J. McDermott. Licensed un-
der a Creati ve Commons Attribution 4.0 International License (CC BY
4.0). Attribution: M. Kanani, S. O’Leary , and J. McDermott, “Radif
Corpus: A Symbolic Dataset for Non-metric Iranian Classical Music”,
in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf .,
Daejeon, South K orea, 2025.
1.1 Foundations of Iranian Classical Music
Iranian classical music comes from a lar ger style of mu-
sic called dastg ¯
ahi music. This term describes the theo-
retical frame work underlying Iranian classical music and
certain styles of Iranian folk music, such as bakhti ¯
ari . The
core repertoire of Iranian classical music is radif (literally
"order"), a structured collection of melodies transmitted
across generations and foundational to performance and
pedagogy .
Radif is a collection of melodies org anized into a spe-
cific sequence, typically di vided into 12 subcategories (tra-
ditionally 13). Out of these, sev en are primary subcate-
gories kno wn as dastg ¯
ah , and fi ve (respecti vely six) are
secondary , referred to as ¯
av ¯
az , which can also be consid-
ered as smaller dastg ¯
ah and serve as subcate gories for the
primary se ven. Each of these subcategories is kno wn for its
distincti ve characteristics. They are typically recognized
based on their main mode (Introduced in the first g ¯
usheh ),
the functional roles of their tones within that mode, and the
specific sequence of g ¯
ushehs within them.
The dastg ¯
ahs are: shur , se g ¯
ah , nav ¯
a , hom ¯
ay ¯
un ,
chah ¯
ar g ¯
ah , m ¯
ah ¯
ur , and r ¯
astpanjg ¯
ah .
The ¯
av ¯
azes are: bay ¯
at-e-kor d , bay ¯
at-e-tork (also re-
ferred to as bay ¯
at-e-zand ), dasht ¯
ı , ab ¯
u’at ¯
a , afsh ¯
ar ¯
ı , and
bay ¯
at-e-esfah ¯
an .
Among the six ¯
av ¯
azes , bay ¯
at-e-esfah ¯
an is a subcategory
of the hom ¯
ay ¯
un , while the remaining are subcategories of
the shur . In many accounts, r adif is considered to hav e
5 ¯
av ¯
azes , as bay ¯
at-e-kor d is often omitted. The reason is
that most experts dispute the requirement of recognizing
it as a independent ¯
av ¯
az . In this study , we hav e included
bay ¯
at-e-kor d to ensure a complete representation.
Each of these subcategories comprises pieces called
g ¯
ushehs . These g ¯
ushehs can range from being as brief as a
single sentence to as extensi v e as a full composition, with
performances lasting se veral minutes.
G ¯
ushehs can be di vided into three types: modal,
melodic, and rhythmic. Modal g ¯
ushehs are played to in-
troduce a mode as a small frame work for improvisation.
Melodic g ¯
ushehs introduce a specific melody and its v aria-
tions, where that specific melody remains fixed in dif ferent
performance versions. Rhythmic g ¯
ushehs represent a spe-
cific rhythm and its v ariations. The same g ¯
usheh names
may appear in dif ferent dastg ¯
ahs or ¯
av ¯
azes . K er eshmeh is
a rhythmic g ¯
usheh that appears multiple times in the radif ,
sharing the same rhythmic pattern in each case. Another
60
example is haz ¯
ın , a melody that is performed in dif ferent
modes; it is classified as a melodic g ¯
usheh . Qarac heh is an
example of a modal g ¯
usheh that appears in more than one
dastg ¯
ah .
The term ¯
av ¯
az has three meanings: 1) broadly , it refers
to singing; 2) more generally , it refers to Iranian non-
metric music; and 3) more specifically , it signifies the seg-
ments of radif that are smaller than a dastg ¯
ah . This study
focuses on the third definition, though the other two mean-
ings are clarified where rele vant.
Non-metric music refers to musical or ganization that
lacks regular meter while potentially maintaining other
temporal structures [1]. The distinction between non-
metric music and free-rhythm music centers on the preser -
v ation of proportional durational relationships. [2] defines
free rhythm as “the rhythm of music without percei ved
periodic or ganization, ” encompassing music where tem-
poral or ganization serves non-rhythmic goals such as te xt
transmission or melodic exposition. W e consider that non-
metric music maintains relati ve proportional relationships
between note durations despite lacking metrical org aniza-
tion, while free-rh ythm music may abandon proportional
consistency entirely .
Tsuge discusses the concept of non-metric music and
emphasizes its greater importance in Iranian music com-
pared to other traditions [3]. He explains that the rhythmic
structure of ¯
av ¯
az (second definition) music is mainly based
on the poetic rhythm system, where a repeating pattern of
dif ferent number and size of syllables shapes its rhythm.
This structure is closely connected to the nature of the Per -
sian (Farsi) language and its classical poetry system which
plays a significant role in ho w the melody is formed and
percei ved. Kanani and Azadehfar [4] described the ke y
¯
av ¯
az (second definition) patterns commonly found in non-
metric traditional Iranian v ocal music.
The exact origins of the r adif system in Iranian music
are not clearly defined. Some sources, like Bruno Nettl, be-
lie ve it originated in the 17th century , while others suggest
the 18th century as the starting point [5–7]. What is clear ,
ho wev er , is that radif de veloped from the late Saf avid era
(1670s-1730s) through to the mid-Q ¯
aj ¯
ar period (1850s).
The lack of precise dating can be linked to the oral tradi-
tion of this music and the absence of recording technology
at the time.
It is belie ved that r adif was created to support the teach-
ing of musical modes and to enhance skills in improvi-
sation and modulation within Iranian art music [8]. The
same radif can be interpreted dif ferently by dif ferent mu-
sicians, and once a student becomes a master , they are able
to de velop their o wn version of the r adif . Over time, many
prominent music masters ha ve created their o wn interpre-
tations, leading to dif ferent versions of radif . These musi-
cians de veloped their r adif based on their personal under-
standing, experience, and e xpression of Iranian modes, ei-
ther for their o wn performances or to teach younger learn-
ers. T raditionally , radif was passed do wn orally from mas-
ter to student, preserving its legac y and technical details
through generations.
The version of r adif curated by musician and educator
M ¯
ırz ¯
a ’Abdoll ¯
ah has become the most widely used choice
in pri vate lessons, uni versity music education and conser -
v atories ov er the past century . Initially , it was mainly asso-
ciated with the t ¯
ar and set ¯
ar instruments, but today , it has
been adapted and performed on many k ey Iranian instru-
ments, including kamancheh , sant ¯
ur , ne y , q ¯
an ¯
un , o ¯
ud , and
qe ychak .
1.2 Exploring Datasets
In recent years, the creation and sharing of digital music
corpora has gained significant attention among researchers
in areas such as music information retrie val, computa-
tional musicology , and natural language processing. V ar -
ious studies ha ve demonstrated that well-curated datasets
can facilitate analysis of both symbolic and audio musi-
cal features, thereby promoting new insights into musical
traditions [9, 10], styles, and technologies.
Many music information retrie v al (MIR) corpora are
primarily audio-based, with annotations for pitch, timing,
structural information, etc., e.g. [11, 12], while others are
symbolic / score-based. Of the latter , some are based on
automated reading of paper scores [13]. Ours differs in
that we ha ve manually written the digital score.
Considering other musical traditions in the geographi-
cal region, there is no symbolic corpus a vailable for Ara-
bic Maq ¯
am music, whereas a symbolic corpus does exist
for T urkish Makams, known as SymbT r [14].
There are some audio datasets related to these musi-
cal traditions, such as the Dunya corpus, which includes
T urkish Makam [15], Carnatic, Hindustani [16], Beijing
Opera [17], and Arab-Andalusian music [18]. The Dunya
corpus is part of a lar ger project called CompMusic [19].
T o the best of our knowledge, there was no symbolic
corpus a vailable for Iranian music before our pre vious
work, in which we introduced the Shour Corpus [20]. This
corpus includes one section of the radif ( shur ) and was
used in our study on discov ering patterns and producing
meaningful v ariations through grammatical representation
(compression) in this musical style.
KUG Dastg ¯
ahi [21] and [22] are two audio datasets for
Iranian music. Na v a [23] is an audio dataset designed for
Iranian instrument recognition, while Ar-MGC is a dataset
for Arabic music genre classification [24].
1.3 Our Contribution
At the time of writing this paper , to the best of our kno wl-
edge, there is no symbolic dataset cov ering the entire radif .
This led us to create the Radif Corpus, which includes all
non-metric pieces from M ¯
ırz ¯
a ’Abdoll ¯
ah’ s radif . Out of the
se veral transcriptions of M ¯
ırz ¯
a ’Abdoll ¯
ah’ s radif , we ha ve
selected the edition titled “Radif Analysis - based on the
notation of M ¯
ırz ¯
a ’Abdoll ¯
ah’ s radif with annotated visual
description” by Dariush T alai [25]. This edition consists
of a recorded performance, together with a score deriv ed
from the performance, notated with hierarchical structure.
Although radif also includes some metric pieces, which
are usually performed at the end of each dastg ¯
ah / ¯
av ¯
az , our
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
61
corpus excludes them as we are focusing on non-metric
music. Figure 1 presents an example of a transcription
from the book.
Figur e 1 . T ranscription sample from Dariush T alai’ s
“Radif Analysis," illustrating the second g ¯
usheh structure
in the shur dastg ¯
ah . The three white boxes indicate the
main structural di visions, while the last white box contains
a further sub-di vision marked in gray .
In [19], Serra identified fiv e critical criteria for corpora
in the CompMusic project: Purpose, Cov erage, Complete-
ness, Quality , and Reusability . These are criteria that we
also considered in our work. Our purpose has already been
stated; Cov erage and Completeness are achiev ed by in-
cluding an entire radif collection. Regarding Quality , we
belie ve the transcription is accurate, as it has been double-
checked by one of the authors who is an e xpert in this mu-
sical style. Reusability is addressed by providing open data
in a documented format.
This corpus is a resource for MIR, computational musi-
cology , and ethnomusicology , enabling applications such
as melodic pattern recognition, automatic transcription,
and mode classification. It enables symbolic music gen-
eration and AI-assisted improvisation. The dataset also fa-
cilitates cross-cultural music studies, allo wing for compar-
ati ve analysis with T urkish Makams and Arabic Maqams,
as well as phrase-le vel e xaminations of melodic progres-
sion. Additionally , its structured format aids compu-
tational analysis of non-metric rhythm and hierarchical
structure understanding, making it a foundational dataset
for exploring Iranian classical music in both traditional and
computational domains.
2. RADIF CORPUS DESCRIPTION
Our corpus includes all these dastg ¯
ahs / ¯
av ¯
azes , featuring
228 non-metric g ¯
ushehs . The MIDI files contain a total
of 43,441 notes, with a total playback duration of approxi-
mately 16,825 seconds (about 281 minutes).
The dataset represents each musical piece as a sequence
of notes, where for each note we store microtonal pitch, du-
ration, pitch (quarter tones), interval, MIDI pitch number
and MIDI bend. Data formats are csv files and MIDI files,
which ha ve been manually transcribed from the book.
Additionally , we pro vide MusicXML files con verted
from the CSV data. These XML files preserve the mi-
crotonal pitch information using fractional <alter> v al-
ues follo wing MusicXML 4.0 standards. For non-metric
rhythm representation, we use fle xible time signatures that
accommodate the total duration of each piece. Howe v er ,
we note that some current music notation software imple-
mentations sho w limitations in both microtonal playback
and non-metric representation. W e observ ed that quarter-
tones are not played back correctly , and the software
tends to generate complex time signatures (e.g., 342/8) as
a workaround for representing non-metric music, which,
while functional, may not pro vide an aesthetically ideal
notation display . The MusicXML con version script is in-
cluded in the repository for researchers who wish to ex-
periment with dif ferent notation software or contrib ute to
improving microtonal MusicXML rendering capabilities.
These files don’ t represent hierarchical structures.
The dataset includes se veral figures that are e xplored
further in the continuation of this paper . Our digital ver -
sion exactly mimics the paper source, while to simplify the
dataset and a void additional comple xity , grace notes or or-
naments are not included in the dataset.
The accurate representation of Iranian classical music
in v olves dealing with two main issues: non-metric rhythm
and micro-tonal pitch. In the following subsections, we
describe our methods to address these challenges.
T onality . Notes are symbolized by
C, D , E , F , G, A, B , with accidental signs including
flat ( Z ), K or on ( k ), Sori ( s ), and sharp ( \ ). Here, “K oron”
and “Sori” indicate micro-tonal adjustments specific
to Iranian music - quarter tones lo wer and higher ,
respecti vely .
Chr omatic Scale. Although these interv als suggest a
24-quarter -tone chromatic scale per octa ve, which can be
seen in some contemporary compositions, Iranian instru-
ments traditionally employ only 18 specific notes: C, D Z ,
D k , D, E Z , E k , E, F , Fs, F \ , G k , G, A Z , A k , A, B Z , B k , B,
with corresponding quarter -tone interv als: 2, 1, 1, 2, 1, 1,
2, 1, 1, 1, 2, 1, 1, 2, 1, 1.
MIDI. Pitch bend is a commonly used method for rep-
resenting microtones in MIDI files in microtonal music
styles. T o encode K oron, we assign it a MIDI note number
one semitone lo wer than the natural note and a pitch bend,
i.e. increase of 2048 (a quarter tone); for Sori, the MIDI
note number is the same as the natural, with a pitch bend
increase of 2048.
Octa ves. W e consider the lowest note in the first g ¯
usheh
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
62
of each dastg ¯
ah / ¯
av ¯
az as the start note of the main octa ve.
For e xample in the first g ¯
usheh of our corpus the lo west
note is F3 so the main octa ve is from F3 to F4. The mu-
sic in the main octa ve is represented only by its symbols
and accidental signs (if any are present). F or notes in oc-
ta ves other than the main one, we use ‘+’ or ‘-’ follo wed
by a number to indicate the number of octa ves abov e or be-
lo w the main octav e. F or example, an A k from one octav e
higher than the main octa ve would be written as A k +1.
Interv als. The Intervals column represents the pitch
dif ference between consecutiv e notes, with "1" indicating
a quarter -tone step.
Durations. The term non-metric does not mean the
same as free rhythm. In this musical style, notes are related
to each other through proportional duration dif ferences,
with some being longer or shorter than others. These re-
lationships create the rhythmic structure, which is mostly
fixed and not altered by the performer . While the exact
durations are not strictly defined, they can be cate gorized
into four main types: very short, short, long, and very long.
These can be said to correspond to sixteenth, eighth, quar-
ter , and half notes [25] and are numerically represented as
1, 2, 4, and 8 in the corpus, where 1 rhythmic unit is equiv-
alent to 1 sixteenth note.
Gr eater Hierarchical Structur es. W e also docu-
mented the hierarchical structure of each piece, as provided
in the original printed source (see Figure 1). In our nota-
tion, brackets represent hierarchical relationships, forming
a tree structure. An open-bracket “[” in the datasheet marks
the beginning of a tree node, with following notes repre-
senting the contents of a section or subsection until the
matching close-bracket “]”. Each tune is enclosed within
brackets, representing the root node. Additional pairs of
brackets define child nodes, which can themselv es contain
further subsections, forming a nested hierarchy .
For e xample, in Figure 1, the abstract hierarchical struc-
ture can be represented as [[][][[]]] .
The outer brackets enclose the entire tune. The second pair
of brackets defines the first section, co vering the first three
lines in Figure 1. The third pair corresponds to the sec-
tion spanning lines three to six. The next open bracket is
follo wed by another open bracket, indicating the presence
of a subsection, which corresponds to lines eight and nine.
The subsection is highlighted in the last line.
3. ST A TISTICAL AND VISU AL O VER VIEW
In the dataset, for each g ¯
usheh , we provide both a pitch
histogram and an interv al histogram. Figure 2 provides an
example of the interv al histogram for a g ¯
usheh .
Each dastg ¯
ah or ¯
av ¯
az comes with a spreadsheet gi ving
information about its g ¯
ushehs , like the number of notes and
total duration. T able 1 presents the number of g ¯
ushehs in
each dastg ¯
ah or ¯
av ¯
az , along with the number of notes, du-
ration in both units and seconds, and pitch range.
M ¯
ah ¯
ur has the highest number of g ¯
ushehs with 34, fol-
lo wed by chah ¯
ar g ¯
ah with 31 and shur with 29. It also
contains the lar gest number of notes, with 6104 in m ¯
ah ¯
ur ,
5788 in chah ¯
ar g ¯
ah , and 4830 in shur . These three also
4
3
2
1
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
Interval
0
10
20
30
40
50
F r equency
07 - Gabri
Figur e 2 . Interval histogram of Gabri , a g ¯
usheh in ab ¯
u’at ¯
a
ha ve the longest durations, with v alues of 12480, 11606,
and 9839 rhythmic units, respecti vely .
Afsh ¯
ar ¯
ı has the fe west g ¯
ushehs with 4, follo wed by
bay ¯
at-e-esfah ¯
an with 5, and dasht ¯
ı and bay ¯
at-e-kor d with
6 each. The shortest dastg ¯
ah / ¯
av ¯
az in terms of duration is
dasht ¯
ı , with 1308 notes and a duration of 2685 units, fol-
lo wed by bay ¯
at-e-kor d with 1360 notes and a duration of
2432 units. afsh ¯
ar ¯
ı is the third shortest, with 1520 notes
and a duration of 2681 units.
Reg arding individual g ¯
ushehs , ker eshmeh in se g ¯
ah is
the shortest g ¯
usheh in the corpus, while bay ¯
at-e r ¯
aje’ va
for ¯
ud in bay ¯
at-e-esfah ¯
an is the longest.
dastg ¯
ah / ¯
av ¯
az g ¯
usheh
Count
Number
of
Notes
T otal
Duration
(unit)
MIDI Perfor-
mance Dura-
tion (second)
Pitch Range
Shur 29 4830 9839 1966 [F , A Z +2]
Bay ¯
at-e-kor d 6 1360 2432 486 [G-1, A Z +1]
Dasht ¯
ı 6 1308 2685 536 [F-2, G+1]
Bay ¯
at-e-tork 16 2544 4986 996 [F-1, G+1]
Abuata 7 2194 3959 791 [F , A Z +1]
Afsh ¯
ar ¯
ı 4 1520 2681 536 [F-1, A Z +1]
Se g ¯
ah 20 3283 6559 1310 [F , F+2]
Nav ¯
a 19 3252 5975 1194 [D-1, C+1]
Hom ¯
ay ¯
un 27 5323 9780 1954 [D, F+2]
Bay ¯
at-e-esfah ¯
an 5 1669 3264 652 [D-1, F \ +1]
Chah ¯
ar g ¯
ah 31 5788 11606 2319 [C-1, G+2]
M ¯
ah ¯
ur 34 6104 12480 2493 [C-1, G+2]
R ¯
astpanjg ¯
ah 24 4266 7958 1590 [D-1, C+2]
T able 1 . Summary of g ¯
usheh information for each dastg ¯
ah
and ¯
av ¯
az , including the number of g ¯
ushehs , total notes, du-
ration in units and seconds and pitch range.
3.1 Melodic Progr ession
One of the main objecti ves when a musician performs
a complete concatenated dastg ¯
ah is to follo w se yr , or
melodic mov ement [26]. In traditional Iranian music, se yr
refers to the progression of melodies within a piece, shap-
ing the ov erall pitch direction of the music through its in-
troduction, de velopment, climax, and resolution.
Se yr underlines the importance of transitional notes and
melodic phrases in establishing the identity and modal
character of the piece. These elements play a crucial role
in guiding the melodic flo w from one section to another ,
ensuring a coherent and expressi v e musical journey .
Each dastg ¯
ah / ¯
av ¯
az folder includes a pitch contour plot
to sho w its se yr . The pitch contour plot illustrates how the
melody and pitch e volv e across different g ¯
ushehs .
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
63
sponding to the consonant-v o wel (CV) transition.
The resulting alignments are check ed and manually cor -
rected for the occasional errors that arise mainly due to
the presence of long v o wels in singing, poor enunciation at
times, as well as the occurrence of significant pitch inflec-
tions. W e observ e the manually corrected onset locations-
i.e. the syllable onsets as realised by the artist- mark ed
at the top of t he w a v eform in the e xample of one full
tala c ycle of the bandish in Figure 2. T o obtain the beat
locations, we annotate the tabla strok e onsets using the
source-separated accompaniment file; we manually mark
the salient beats (do wnbeat ’x’ sam and 9 th beat ’o’ khali )
in 1 c ycle. Assuming a consistent local tempo, we di vide
each half c ycle (interv al between one do wnbeat and the ad-
jacent t th beat into 8 equal parts, thereby obtaining all the
estimated beat instants across the rendition.
Ne xt, we mark the canonical locations of the match-
ing syllables, positioning the syllables according to the
Bhatkhande notation (as in Figure 1). The position map-
ping of the realised syllable onsets to the corresponding
canonical syllable w as implemented follo wing the method
proposed in pre vious w ork [8]. W e observ e from Figure 2
ho w t h e realised onsets lag the canonical locations most of
the time.
Figur e 2 . Pitch contour (bottom) and sung syllable align-
ment (red) with canonical beat positions of the same sylla-
ble (black) for an e xcerpt of J a J a Re by ABD.
Finally the v oc al pitch is e xtracted at 10 ms interv als us-
ing an autocorrelation based method for fundamental fre-
quenc y and v oici ng [12]. Brief pauses and un v oiced re-
gions are linearly interpolated to obtain a continuous pitch
contour for each sung syllable re gion. The pitch contour is
con v erted to cents by normalisation with the kno wn perfor -
mance’ s tonic. Ev entually , we obtain for each performance
in our dataset, the se gmented audio of each sung line an-
notated at the syllable le v el with syllable name, boundaries
and the pitch (cents) at 10 ms interv als. W e use these lo w-
le v el features to define quantities that capture the singing
v ariations across repetitions of a bandish line within and
across singers. The reference for the comparison is the
syllable identity (i.e. its name and metrical location) as
defined in the canonical notation as presented in Figure 1.
Rag a Bhimpalasi Y aman
Bandish Ja Ja Re Y eri Aali
T ala T eentaal T eentaal
Sw ar S, R, g, m, P , D, n, S S, R, G, M, P , D, N, S
# Concerts 15 13
# Artists 15 12
# Repetitions (L1, L2, L3, L4) 167, 39, 47, 47 94, 32, 35, 23
Matra per min range 138–200 111–203
T able 1 . Summary of our dataset of concert recordings
across r a gas , bandish , and artists. Swar notation details
are in the supplementary .
4. MEASURING EXPRESSIVENESS
W e wish to quantify and compare the v ariability observ ed
in the acoustic realisation of a gi v en syllable, from a spe-
cific line of the bandish , across (i) repeated utterances
within an artist’ s performance, and (ii) utterances of the
same syllable across dif ferent performances/artists. The
acoustic parameters that we detect are: (i) the onset time
of the syllable, (ii) the syllable duration (as the time inter -
v al between the current syllable’ s onset and either the on-
set of the ne xt syllable or the start of the follo wing silence
se gment, whiche v er occurs first. and (iii) the pitch contour
shape across the syllable interv al. W e illustrate the process
by pro viding e xamples of the process ing and analyses of
the audio rendering of a chosen line by one artist.
4.1 T iming expr ession
The de viation of the detected onset from its reference as-
signed beat inde x in the canonical notation gi v es us an
estimate of the lag/lead of the singer for the syllable in
question. W e represent the de viation in terms of fraction
of the local beat interv al; this normalization f acilitates the
comparison across instances and concerts. W e can vie w
the thus measured timing of fsets as e vidence of e xpressi v e
timing, especially if this quantity sho ws v ariability across
repetitions of the syllable within the concert.
Figur e 3 . De viation of the sung syllable onsets from the
canonical locations measured in the units of beat duration
for J a J a Re Line 1 by ABD for multiple repetitions of the
line.
Figure 3 captures the onsets of syllables in bandish 1,
Line 1 as rendered by singer ABD. The syllable names
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
70
are sho wn with their canonical matr a locations at the bot-
tom. W e note that some syllables occup y 2 beats and oth-
ers 1 beat in the canonical form. W e observ e, for e xam-
ple, that "Jaa2" and "Man" (both 2-beat syllables) sho w
v ariations with a mean lag of about one beat. "Man" in-
stances ho we v er are much more dispersed. While a uni-
form of fset could potentially indicate a structural dif fer -
ence between the artist’ s v ersion of the bandish and that
of the Bhatkhande book, a high standard de viation (lik e
in "Man") points to the artist’ s in-the-moment e xpressi v e
v ariations. On the other hand, the 3 syllables preceding
"Man" sho w near -zero of fsets across repetitions.
4.2 Pitch expr ession
Analysing the rendered song for pitch-based e xpressi v e-
ness is a rather in v olv ed task. There are man y w ays in
which the artist injects e xpressi v enes s in their performance
via pi tch v ariation. As we w ant to analyse the dif ferent
w ays of realising the same line of a bandish , we are in-
terested in the v ariability in the pitch contour (PC) shape
of a syllable across the multiple repetitions of that line. A
greater di v ersity of PC shapes can then be interpreted as
higher e xpressi v eness in an artist’ s repetitions of the line.
Figur e 4 . PC quantisation to the nearest r a ga note for one
instance of J a J a Re bandish line 1 rendered by artist ABD
F or each syllable, the associated PC spans the duration
of that syllable; hence, this is not a fix ed-length time series.
W e represent this v ariable dimension PC for a gi v en sylla-
ble with a lo wer and fix ed-dimensional v ector . W e first im-
plement a piece-wise aggre g ate approximation (P AA) o v er
the syllable PCs. P AA is a time-series representation that
has been used widely in data-mining tasks [13]. This is
implemented as follo ws.
The syllable PC v alues are each first quantised to the
nearest r a ga swar (note), Figure 4 sho ws this process.
Ne xt, the quantised PC v alues for a syllable across repe-
titions, is aggre g ated by di viding the quantised PC into a
fix ed number of uniform interv als and assigning the mode
of the v alues to each interv al. The number of fix ed in-
terv als is set empirically to 10 interv als per beats allot-
ted to the syllable (treating the syllable e xtensions indi-
cated by ’-’ in Figure 1 to be a part of the pre vious syl-
lable, thereby adding to its allotted beats). The choice of
10 equal se gments per beat interv al is based on the tempo
range of our dataset (110-200 BPM or 300 ms to 545 ms
Figur e 5 . Three distinct renditions of the syllable "Jaa1"
by artist ABD, each represented by a fix ed number of uni-
form time interv als. Each interv al is mapped based on its
modal pitch to the nearest r a ga note, gi ving us the P AA
string representation for the s yllable’ s pitch shape. Fig-
ure 6 describes this process and the P AA strings for the
abo v e PCs.
per beat) and sampling period (10 ms/sample) of the pitch
contour . This results in se gments, each represented by a
short sequence of samples of the pitch contour . This bal-
ance allo ws capturing dynamic pitch fluctuations which
are pre v alent in Hindustani classical music (HCM), with-
out o v er -quantisation. No w , each P AA interv al within a
syllable is assigned a discrete symbol, where the symbols
are dra wn from a suitable alphabet which comprises notes
( swar ) across the rele v ant octa v e ranges. The resulting
string of note v al ues (one per P AA interv al) then represents
coarsely the realised pitch shape of the syllable. Figure 6
sho ws this process . The 3 string sequences in Figure 6 cor -
respond to the 3 PCs in Figure 5. This type of aggre g ation
presents a tradeof f of generality vs specificity .
Lik e in the case of syllable timing, we are interested
in the v ariation, if an y , in pitch shape of a gi v en syllable
across repetitions . W e achie v e this by computing the sim-
ilarity of the syllable PCs for pairs dra wn from the set of
repetitions in a single concert. The Le v enshtein edit dis-
tance [14] between the P AA strings pro vides us with the
number of note substitutions. W e e v aluate the Normalised
Le v enshtein Substitution Score (NLSS) for each pair as a
measure of the dissimilarity . Figure 7 sho ws a matrix rep-
resentation (heat map) of NLSS v alues for a chosen sylla-
ble as rendered by one artist across 14 repetitions.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
71
Figur e 6 . Processing pipeline for generating P AA string
representations for the pitch contour for dif ferent repeti-
tions of a syllable. The strings correspond to the PCs of
the syllable repetitions gi v en in Figure 5 spanning 2 beats.
Figur e 7 . Heat map sho wing NLSS for each pair of "Jaa1"
syllable PCs dra wn from the set of repeti tions of line 1 of
bandish J a J a Re rendered by artist ABD
5. OBSER V A TIONS AND DISCUSSION
W e are interested in the within-artist v ariation for a gi v en
bandish line and its syllables. This can f acilitate poten-
tially v aluable insights about (i) the preferred locations (in
terms of chosen syllables) for e xpressi v e gestures by indi-
vidual artists and across artists and (ii) the e xtent and na-
ture of e xpressi v e gestures for a gi v en artist. Due to space
limitations, we present the analysis results for line 1 of one
bandish , with the other bandish presented in the supple-
mentary .
Figure 8 presents for each artist and syllable, the stan-
dard de viation (s.d.) of the timing de viati on (as captured
in the e xample for artist ABD in Figure 3). The s.d. helps
us focus on the variability of of fsets rather than on ac-
tual of fset v alues (which might b e attrib uted to structural
dif ferences between the artist’ s v ersion and Bhatkhande’ s
v ersion of the bandish , rather than e xpression-related).
In Figure 8, we note the dominance of the first 3 syl-
lables for most artists. The full range of beha viours, ho w-
e v er , includes IN at one end with minimal v ariations to DG
and RK, who introduce ne w v ariations on practically all
syllables. That IN does not e x ercise an y fle xibility is not
Figur e 8 . Standard de viation of the distrib ution of the frac-
tional timing de viation for e v ery syllable o v er multiple rep-
etitions in one rendition, across artists.
Figur e 9 . Mean NLSS o v er all pairs of repetitions of each
syllable in line 1 of bandish J a J a Re , computed per artist.
serving as a dissi milarity measure across dif ferent repeti-
tions of a syllable by an artist.
surprising gi v en that his performances were e xplicitly cre-
ated to closely follo w the prescribed Bhatkhande notation,
as discussed here [15]. RK, on the other hand, is consid-
ered a virtuoso musician.
Aggre g ating across the ro ws, we obtain the per sylla-
ble beha viour across artists in Figure 10. The mean v alues
sho w that the first 3 syllables carry the most e xpressi v e tim-
ing, with "Jaa2" also sho wing the most spread across artists
(consistent with Figure 8). W e see, for e xample, that PT
sho ws a lar ge range in per syllable SD, ag ain agreeing with
Figure 8. In the case of the artist ABD, we can observ e that
the temporal de viation is spread relati v ely e v enly , while
peaking for a particular syllable "Man", which is the do wn-
beat. This can be easily appreciated in listening to the au-
dio, which can be accessed in the supplementary material.
T o assess pitch v ariability , we calculate the a v erage
number of pitch substitutions by taking the mean of the
NLSS for al l pairs of repetitions per artist and per syllable,
(normalised by the string length), calculated across all pos-
sible pairs of the gi v en syllable utterances within a concert.
Figure 9 sho ws the v alues per artist and per syllable of Line
1. W e can observ e in Figure 8 and Figure 9 that both the
pitch and temporal v ariation across repetitions and across
dif ferent artists is more prominent at the be ginning of the
line. Near the end of the line, the pitch v ariation increases,
which can be att rib uted to the emphasis on a semantically
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
72
Figur e 10 . Box-plot of s.d. of timing de viation param-
eter per syllable of J a J a Re Line 1 as aggre g ated across
concerts and artists. The tala -c ycle ends at the syllable
"Ne", and a ne w c ycle starts from do wnbeat at the syllable
"Man".
Figur e 11 . Box-plot of the mean of the a v eraged NLSS per
syllable of J a J a Re Line 1 as aggre g ated across concerts
and artists.
important w ord- "Mandira v a".
The syllables "P a" and "Ne" are near the tala -c ycle
boundary , and we observ e a minimum pitch and temporal
v ariation; this is consistent with the observ ations of [8]. As
the performer heads to w ards the tala -c ycle boundary , their
o v erall tendenc y of v ariation and impro visation decreases,
aiming rather to w ards resolving the melodic and temporal
e xpression the y ha v e come up with in the particular tala -
c ycle.
Aggre g ating across ro ws of Figure 9, we obtain the per
syllable beha viour o v er all artists in Figure 11. The dotted
line joining the means indica tes that the pitch v ariation de-
creases as one approaches the tala -c ycle boundary at the
syllable "Ne", while the the pitch e xpressi v eness is higher
in the start and end of the line, which f alls on the 7th and
5th beats of the tala -c ycle respecti v ely , which is f ar from
the c ycle boundary , hence pro viding more scope for timing
and pitch e xpressi v eness. W e see, for e xample, that A C
and AK sho w high pitch v ariation across the repetitions of
the same line.
Hierarchical clustering of all the pairs from the set of
repetitions of a syllable by an artist pro vides us with infor -
mation about the v ariation clusters. A threshold can be de-
fined t hat decides if a v ariation belongs to a cluster or not.
More di v erse v ariations w ould indicate more number of
clusters, indicating higher e xpressi v eness. Dendrograms
are e xcel lent for visualising such clusters. The realisation
of the pre vious syllable has an influence on which cluster
the follo wing syllable v ariation w ould belong to.
Comparing Figure 8 and Figure 9, we note that e xpres-
si v e gestures that utilise pitch are not necessarily at the
same locations that e xhibit timing de viation in terms of
the preferred syllable. It is rather interesting to look at the
least amount of e xpressi v eness in both pitch and timing lie
with the syllables that are near the tala -c ycle boundary . An
analysis at the indi vidual audio le v el w ould pro vide a more
accurate picture of the correlations, if an y , and is left to
future w ork.
Our observ ations, r eported here on the Line 1 of one
bandish , lar gely hol d with the second bandish . The under -
lying reasons for the choice of specific syllables e xhibiting
lar ger v ariability are similar to those discus sed by Mor -
ris [4]. These include lar ger v ariations at line or phrase
ending syllables due to the ef fect of pre vious and ne xt con-
te xts, and the choice of syllables belonging to more emo-
tionally loaded w ords in the lyrics.
6. CONCLUSION
In this paper , we articulated the problem of modeling e x-
pressi v e v ariations in the conte xt of performance of Hin-
dustani traditional compositions by established artists of
the genre. A well-kno wn bandish in the chosen r a ga is
al w ays sung at the be ginning of a concert with multiple
repetitions of the lines, mark ed by v ariations in the loca-
tion and type of the e xpressi v e gestures. Based on our
proposed methodology , we sho wed that it is possible to
arri v e at systematic patterns across artists by treating the
syllables of the lyrics as reference points for a study of the
range of v ari ation. This also helped us discuss interesting
correlations between the roles of melody and rh ythm in e x-
pressi v eness.
W e presented a dataset that w as annotated with a combi-
nation of manual and automatic tools to obtain a rich repos-
itory of distinct realizations of the lines of tw o popular
traditional compositions. While much further e xploration
remains possible, this w ork demonstrates the potential of
computational models for impro visation in the conte xt of
compositions in the Khayal genre. This w ork lays the
foundation for generati v e applications by capturing high-
le v el performance features that reflect an artist’ s distincti v e
style. A preliminary e xperiment w as pe rformed to generate
the temporal de viations discussed abo v e, where the distri-
b utions formed by all the de viations of a syllable from its
canonical location for an artist were used to sample out
ne w points for each syll able. A sine-tone based audio w as
synthesized from the generated pitch contour . The refer -
ence (Bhatkhande canonical form) and generated tracks are
a v ailable in the supplementary . This lets us create infinite
possibilities for rendering the same line, while at the same
time capturing some hint of the artist’ s style. Extending
this approach to other acoustic dimensions for e xpression
such as timbre and dynamics can enable the generation of
classical music that embodies the unique identit y of indi-
vidual artists.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
73
7. A CKNO WLEDGMENTS
W e take this opportunity to ackno wledge and express our
gratitude to all those who supported and guided us during
this research work. W e are thankful to Madhumitha S.,
whose thesis work w as a crucial base for our project. W e
also thank Mr . Himanshu Sati, and Mrs. Hemala Ranade
for their scholarly musicological insights, which ga ve di-
rection to our work.
8. REFERENCES
[1] B. W ade, “Music in India: The classical traditions, ”
Manohar Press, 2001.
[2] C. E. Cancino-Chacón, M. Grachten, W . Goebl, and
G. W idmer , “Computational models of expressi ve mu-
sic performance: A comprehensiv e and critical re-
vie w , ” F r ontier s in Digital Humanities , vol. 5, p. 25,
2018.
[3] P . N. Juslin, “Cue utilization in communication of emo-
tion in music performance: Relating performance to
perception. ” J ournal of Experimental Psycholo gy: Hu-
man per ception and performance , vol. 26, no. 6, p.
1797, 2000.
[4] A. Morris, T ransmission and performance of Khayal
compositions in the Gwalior gharana of Indian vocal
music. PhD thesis, Uni v . Of London, S.O.A.S., 2004.
[5] W . V an der Meer , “ Audience response and e xpressi ve
pitch inflections in a li ve recording of le gendary singer
kesar bai k erkar , ” in Expr essiveness in music perfor -
mance: Empirical appr oac hes acr oss styles and cul-
tur es , D. Fabian, R. T immers, and E. Schubert, Eds.
Oxford Uni versity Press (UK), 2014, pp. 170–184.
[6] S. Sankaran, P . V . K. Sekhar , and A. M. Hema, “ Au-
tomatic segmentation of composition in carnatic music
using time-frequency cfcc templates, ” in Pr oceedings
of 11th International Symposium on Computer Music
Multidisciplinary Resear c h (CMMR) , 2015.
[7] K. K. Ganguli and P . Rao, “ A study of variability in
raga motifs in performance conte xts, ” Journal of Ne w
Music Resear c h , vol. 50, no. 1, pp. 102–116, 2021.
[8] Y . Bhake and P . Rao, “Expressi ve timing in Hindus-
tani v ocal music, ” in Pr oc. of ICASSP 2025 W orkshop
on Indian Music Analysis and Generative Applications
(WIMA GA) , Hyderabad, India, 2025, accessed at:link.
[9] S. Rao and P . Rao, “ An o vervie w of Hindustani music
in the context of computational musicology . ” J ournal
of New Music Resear ch , v ol. 43, no. 1, 2014.
[10] V . Bhatkhande, Kramik Pustaka Malika . Sangeet
Karyalaya Hathras, India, 2013.
[11] D. Pov ey , A. Ghoshal, G. Boulianne, L. Bur get,
O. Glembek, N. Goel, M. Hannemann, P . Motlicek,
Y . Qian, P . Schw arz et al. , “The kaldi speech recog-
nition toolkit, ” in IEEE 2011 workshop on automatic
speech r ecognition and understanding . IEEE Signal
Processing Society , 2011.
[12] Y . Jadoul, B. Thompson, and B. De Boer , “Introducing
parselmouth: A python interface to praat, ” Journal of
Phonetics , v ol. 71, pp. 1–15, 2018.
[13] J. Lin, E. Keogh, L. W ei, and S. Lonardi, “Experienc-
ing SAX: a nov el symbolic representation of time se-
ries. ” Data Mining and knowledg e discovery , vol. 31,
no. B, pp. 107–144, April 2007.
[14] W . J. Heeringa, “Measuring dialect pronunciation dif-
ferences using Le venshtein distance, ” 2004.
[15] “2000 classic compositions from Bhatkhande on cd, ”
https://scroll.in/article/726180/, accessed: 2024-04-10.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
74
COLORING MUSIC: BRIDGING MUSIC AND COLOR P ALETTES
FOR GRAPHIC DESIGN
T akayuki Nakatsuka Masahiro Hamasaki Masataka Goto
National Institute of Adv anced Industrial Science and T echnology (AIST), Japan
{takayuki.nakatsuka, masahiro.hamasaki, m.goto}@aist.go.jp
ABSTRA CT
This paper explores the relationship between music and the
color palettes used for designing their corresponding mu-
sic cov er images, providing a comprehensi ve analysis that
bridges auditory and visual expression. Our findings re veal
a relationship between musical pieces and certain colors,
suggesting that the color palettes used in cov er image de-
sign are carefully selected to reflect the auditory e xperience.
Building on these findings, we propose a framew ork that
estimates appropriate color palettes for musical pieces to
support selecting colors for cov er image design. Using a
lar ge priv ate dataset of 582,894 pairs of a musical piece
and its corresponding co ver image from v arious music gen-
res, our frame work le verages deep learning techniques to
train our color palette estimator . W e demonstrate the effec-
ti veness of our proposed frame work in graphic design by
sho wcasing an application that generates cov er images us-
ing the estimated color palettes from gi ven musical pieces.
1. INTR ODUCTION
In multimodal music understanding, both music and their
corresponding music cov er images play a crucial role. For
instance, Oramas et al. successfully improv ed music genre
classification accuracy by incorporating image features in
addition to audio features [1]. In addition, L
¯
ıbeks and
T urnbull sho wed that cov er images in v olve distinct features
that can be used to predict music genre tags [2]. These
studies suggested that a cov er image embodies the essence
of its corresponding music content, thereby establishing
that analyzing these images yields a deeper understanding
of the music. This study focuses on the colors used in cov er
images and analyzes their relationship with music.
The colors used in cov er images tend to empirically re-
flect the characteristics of the corresponding music style. As
illustrated in Fig. 1, dif ferent music genres display distinc-
ti ve characteristics in the colors used in the co ver images.
As colors are closely linked to cultural conte xts [3], emo-
tions [4, 5], and the ability to attract visual attention [6],
cov er images contribute to the promotion of music content
© T . Nakatsuka, M. Hamasaki, and M. Goto. Licensed under
a Creati ve Commons Attribution 4.0 International License (CC BY 4.0).
Attribution: T . Nakatsuka, M. Hamasaki, and M. Goto, “Coloring Music:
Bridging Music and Color Palettes for Graphic Design”, in Pr oc. of the
26th Int. Society for Music Information Retrieval Conf., Daejeon, South
K orea, 2025.
Death m etal Country Ele ctronic
Figur e 1 . Example results of Google Search with the text
queries “{music genre} music alb um covers, ” where we
used the music genres ‘Death metal, ’ ‘Country , ’ and ‘Elec-
tronic. ’ Music cov er images for each genre are characterized
by the colors used in co ver image design: dark colors for
‘Death metal, ’ bro wnish colors for ‘Country , ’ and vivid col-
ors for ‘Electronic. ’
and enhance the ov erall music appreciation experience [7].
Therefore, this relationship between music and the colors
used in cov er images has been the subject of se veral stud-
ies [8
–
10]. Howe v er , these studies hav e mainly focused on
genres, not on musical pieces.
This paper first in v estigates the preferred colors for de-
signing cov er images across multiple genres in our prelimi-
nary study (Section 4) and further explores the relationship
between musical pieces and the colors used in their corre-
sponding cov er images based on our proposed frame work
(Section 5). In this study , we focus on not only a repre-
sentati ve color b ut also color palettes used in cov er images
because they play a crucial role in graphic design [11
–
13],
shedding light on the deliberate selection process of colors
that reflect the essence of the music content.
Based on our findings that a relationship exists between
musical pieces and the colors used in their corresponding
cov er images, we propose a framew ork to estimate appro-
priate color palettes for musical pieces. The ke y technical
75
aspects of our frame work are ho w to extract color palettes
from cov er images and ho w to estimate color palettes for
musical pieces. For a color palette e xtraction method, we
employ data-driven color manifolds [14], which are use-
ful in arranging the colors as a color palette. F or a color
palette estimator , we train a deep neural network to esti-
mate an appropriate color palette for each musical piece. In
this training, we lev erage a pretrained audio model ( con-
trastive langua ge-audio pr etrai ning (CLAP) [15] or A u-
dioT oken [16]) as an audio feature e xtractor to extract a
distincti ve feature from each musical piece. This frame-
work bridges musical pieces and their corresponding co ver
images using color palettes.
T o demonstrate the effecti v eness of our frame work, we
present an example application that generates co ver images
using the estimated color palettes from gi ven musical pieces
to support creating visually appealing cov er images.
2. RELA TED WORK
Se veral studies ha ve in vestigated the relationship between
music and color . W ells argued that there is a correlation
between music and color based on the principle of comple-
mentarity [17]. Furthermore, Pesek et al. suggested that
since music and emotions are closely related (e.g., [18, 19]),
as well as emotions and colors (e.g., [4, 5]), there exists
a relationship between music and color mediated by emo-
tions [20]. Ho wev er , these studies hav e only partially elu-
cidated the relationship between music and color , as they
analyzed this relationship using a limited number of colors.
Therefore, in this study , we use the colors used in music
cov er images that embody a musical essence [1, 2] as the
basis for our analysis.
In research exploring the colors used in co ver images,
pre vious studies hav e focused on specific genres (classi-
cal [8] and metal [9]). Seker [8] discov ered that the colors
used in cov er images for classical music predominantly fa-
v or neutral colors. Friconnet [9] found that cov er images
for metal music tend to use darker colors than those of
other genres, with a preference for black and orange [9].
Although these studies provide insights into the colors used
in cov er images of specific genres, no studies hav e explored
which color v alues are preferred for specific musical pieces
of v arious genres.
Additionally , color themes used in designing cov er im-
ages ha ve been studied [10]. Dorochowicz and K ostek [10]
analyzed cov er images across multiple genres with respect
to basic color analysis rules such as seasonal colors (e.g.,
spring (warm and bright), summer (cool and soft), autumn
(warm and soft), and winter (cool and bright)) and de grees
of brightness (e.g., light, medium, and dark). While their
findings provide v aluable insights into the color characteris-
tics of each genre, they focus on a limited number of color
palettes based on the basic color analysis rules.
In this paper , we in vestigate the relationship between
musical pieces and the color palettes used in their corre-
sponding cov er images and explore the application of this
relationship in cov er image design.
Input im ages
" -m eans
(cluste ring method)
Data -driven
color m anifolds
Figur e 2 . Comparison of color palette extraction methods.
Gi ven the input image (top ro w), the data-dri ven color man-
ifolds (bottom ro w) extract a color palette from the image
in consecuti ve color order , while
k
-means (middle ro w) ex-
tracts a color palette from the image in random color order .
3. COLOR EXTRA CTION
T o extract a representati ve color or color palettes from music
cov er images, we lev erage data-driven color manifolds [14],
a technique which aims to acquire color samples from im-
ages and learn a lo wer-dimensional manifold of the acquired
color samples. The learned manifold reflects the distribu-
tion of colors in cov er images, compressing areas of the
color space that are less commonly used and expanding
those that are more frequently utilized.
The technique in v olves se veral steps, starting with the ac-
quisition of color samples from cov er images. For success-
ful color manifold learning, a suf ficient number of samples
(ov er 10k) must be obtained from each image. Note that
we utilized all samples from
224 px × 224 px
-resized cov er
images, amounting to over 50k samples. These samples
are then used to estimate the density of each color in the
cov er images, with a focus on identifying and preserving the
most important colors. A self-organizing map [21], which
is used to reduce dimensionality , is then applied to deri ve
the one-dimensional or two-dimensional color manifolds.
W e utilize the one-dimensional color manifold to extract
color palettes from cov er images. In practice, we calculate
a discrete color manifold, which consists of
M ∈ N
colors,
to use the deri ved color manifold as a color palette. All
hyperparameter v alues related to density estimation and
dimensionality reduction were taken from [14], e xcept for
the smoothness parameter , which we set to r 0 = 1 .
The adv antage of this technique ov er clustering methods
such as
k
-means [22] is that the color palette extracted by
the data-dri ven color manifolds has a meaningful order -
ing, where the order of colors is determined by the deri ved
one-dimensional color manifold and thus results in consec-
uti veness, while the color palette e xtracted by a clustering
method has a random ordering (see Fig. 2). When using a
color palette consisting of multiple colors in graphic design,
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
76
Classica l Stage & S creen
Reggae L atin
Rock El ectronic
R
G
B
0
1
0
0
1
1
R
G
0
1 0
0
1
1
R
G
0
1
0
0
1
1
R
G
0
1
0
0
1
1
R
G
0
1 0
0
1
1
R
G
0
1
0
0
1
1
B B
B B
B
Low saturat ion
Warm colors
Wide variations
Figur e 3 . V isualization of a representativ e color used in
music cov er images by music genre. The larger the circle
in the visualization, the more frequently the circle’ s color
appears in the cov er images.
the color palette with a continuous color order based on the
data-dri ven color manifolds is intuiti ve and easy to use.
4. PRELIMINAR Y STUD Y
This section describes our preliminary study that aims to
analyze the preferred colors for designing music cov er im-
ages across multiple genres by le veraging color palettes
extracted from these images.
4.1 Experimental Setup
4.1.1 Dataset
W e randomly collected 3,887 cov er images (each image is
an RGB image) for the experiments. W e assigned genre tags
to each image based on the grouping of genres and styles
in Discogs
1
, in which v arious music is organized into 15
genres and styles (‘Blues, ’ ‘Brass & Military , ’ ‘Children’ s, ’
‘Classical, ’ ‘Electronic, ’ ‘Folk, W orld, & Country , ’ ‘Funk /
Soul, ’ ‘Hip-Hop, ’ ‘Jazz, ’ ‘Latin, ’ ‘Non-Music, ’ ‘Pop, ’ ‘Reg-
gae, ’ ‘Rock, ’ and ‘Stage & Screen’). A total of 5,150 genre
tags were assigned to 3,887 cov er images, which means an
a verage of 343.3 images per genre.
1
The grouping of genres and styles in Discogs is av ailable
at
https://support.discogs.com/hc/en- us/articles/
360005055213- Database- Guidelines- 9- Genres- Styles
.
T able 1 . List of representativ e colors most frequently used
in music cov er images for each music genre, excluding
grayscale colors.
Music genre RGB v alue Color
Blues (148,135,102)
Brass & Military (110,101,74)
Children’ s (139,178,241)
Classical (105,132,128)
Electronic (72,36,36)
Folk, W orld, & Country (108,101,68)
Funk / Soul (111,109,73)
Hip-Hop (73,36,36)
Jazz (181,145,109)
Latin (146,112,110)
Non-Music (165,127,156)
Pop (110,73,73)
Regg ae (168,132,68)
Rock (72,36,36)
Stage & Screen (174,172,106)
4.1.2 Implementation details
For representing colors in color manifolds, we utilized
an RGB color space, which is a widely used additiv e
color model. W e resized all of the co ver images into
224 px × 224 px
and normalized their RGB v alues to [0,
1]. Then, for the purpose of this preliminary study , we
simply extracted one representati v e color (i.e.,
M = 1
)
from the resized images using the data-dri ven color mani-
folds [14] as described in Section 3. Note that we extracted
more colors to form the color palettes in Section 5.4. W e
used all of the pixels in the resized images as color samples.
The constructed color manifold is di vided into eight bins,
and the samples are discretized into these bins, enabling
their visualization as a histogram.
4.2 Results
Fig. 3 represents three-dimensional histograms for the R,
G, and B v alues in the RGB color space. As sho wn in
Fig. 3, trends in the distrib ution of colors differ by genres:
for example, ‘Classical’ and ‘Stage & Screen’ music co ver
images ha ve lo w saturation, and ‘Reggae’ and ‘Latin’ music
cov er images tend to use warm colors such as red and
yello w . Additionally , it can also be observed that genres
such as ‘Pop’ and ‘Rock’ feature a wide v ariety of colors
in their cov er images. These genres hav e more di verse
subcategories and styles than other genres, resulting in such
color v ariation. Note that due to te xt on cov er images
and background colors, grayscale colors are prominently
displayed in the histogram. Therefore, T able 1 lists the
representati ve colors most frequently used in co ver images,
excluding grayscale colors. As sho wn in T able 1, the most
frequently used colors v ary by music genre. Our proposed
color palette estimation frame work is designed based on
this insight.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
77
5. COLOR P ALETTE ESTIMA TION FRAMEWORK
Using the close relationship between musical pieces and
colors used in their corresponding music cov er images, we
propose a frame work designed to estimate appropriate color
palettes for musical pieces. Fig. 4 shows an o vervie w of our
frame work. W e utilize an audio feature e xtractor and a color
palette estimator to estimate a color palette for each musical
piece. The audio feature extractor extracts audio features
from musical pieces. T o le v erage recent advancements in
audio models for do wnstream tasks, we use a pretrained
audio model as the audio feature e xtractor . Then, the color
palette estimator takes the e xtracted audio features as input
and estimates appropriate color palettes for musical pieces.
T o train the color palette estimator , we construct a large
pri vate dataset of musical pieces and their corresponding
cov er images, but we need the ground-truth color palettes to
be estimated. Therefore, we extract the ground-truth color
palette from each cov er image by le veraging the data-dri ven
color manifolds [14] described in Section 3 for our color
palette extraction method.
5.1 A udio F eature Extractor
Audio models trained on lar ge datasets hav e demonstrated
their capabilities in do wnstream tasks [15, 23]. F or exam-
ple, the outputs of the final layer of an audio model are
utilized in classification tasks, while audio embeddings are
used in generati ve tasks. In our approach, we le verage the
pretrained audio model
F
to extract audio features from
musical pieces.
Let
A = { a n ∈ R T } N
n =1
be a set of musical pieces,
where
T
is the length of each musical piece and
N
is the
number of musical pieces. Next, let
Z = { z n ∈ R d } N
n =1
be
a set of audio features, where
d
is the number of dimensions
of each audio feature. The audio feature
z n
can be extracted
from the musical piece
a n
by using the pretrained audio
model F as follo ws:
z n = F ( a n ) . (1)
W e utilize the pretrained audio model with all of its trainable
parameters fixed.
5.2 Color Palette Estimator
T o estimate color palettes from the extracted audio features
z n
, we propose the color palette estimator
G
, which consists
of three linear layers with a GELU function [24].
Let
C = { c n ∈ R 3 × M } N
n =1
be a set of color palettes,
where
M
is the number of colors in each color palette
and each color is represented by a set of three numerical
v alues (i.e., RGB values). The color palette
c n
can be
estimated from the audio feature
z n
by using the color
palette estimator G as follo ws:
c n = G ( z n )
= σ ( W 3 GELU( W 2 GELU( W 1 z n + b 1 ) + b 2 ) + b 3 ) ,
(2)
where
σ
is a sigmoid function. The parameters of the color
palette estimator
G
are defined by
W 1 ∈ R h × d , W 2 ∈
R h × h , W 3 ∈ R 3 × M × h , b 1 ∈ R h , b 2 ∈ R h
, and
b 3 ∈
Musical pi ece
Music cover
ima ge
Audio feature
extractor (fi xed)
Color palette
extraction
Color palette
estimator
! ! " #
Estim ated
color pa lette
Ground-truth
color pa lette
Trai ning procedure
Color pal ette esti mation fram ework
Color pal ette esti mation
Musical pi ece
Audio feature
extractor (fi xed)
Color palette
estimator (fixed)
Estim ated
color pa lette
Figur e 4 . Overvie w of our proposed color palette estima-
tion frame work. (T raining pr ocedure) W e start with an
original pair of a musical piece and its corresponding music
cov er image. The musical piece is processed by a fixed
audio model to extract its audio feature, and then a color
palette estimator is trained to estimate a color palette from
the audio feature. For this training, the ground-truth color
palette is extracted from the co ver image by using the color
palette extraction method. W e use the mean squared error
(MSE) loss function to optimize our color palette estimator .
(Color palette estimation) After training, the color palette
estimator can be used to estimate appropriate color palettes
for musical pieces.
R 3 × M
, where the dimension of the hidden layer
h
is set to
768. While training, a dropout with a probability of 0.2 is
applied to each output of the GELU functions.
5.3 Experimental Setup
5.3.1 Dataset
The lar ge pri vate dataset for training our color palette esti-
mator contains music audio excerpts (each e xcerpt is a 30 s
audio pre vie w for trial listening, with a 44.1 kHz sampling
rate) and their corresponding cov er images (each image is
an RGB image). The excerpts and their co ver images are
limited to single tracks, i.e., an original pair of an e xcerpt
and its corresponding cov er image is unique. The dataset
contains 582,894 pairs of an excerpt and its correspond-
ing cov er image by 115,113 artists. W e randomly split the
dataset into training, v alidation, and test sets with an eight-
one-one ratio (i.e., 466,316 pairs for the training set and
58,289 pairs for v alidation and test sets, respectiv ely) and
with no artists ov erlapping across these sets.
5.3.2 Implementation Details
As described in Section 4.1.2, we utilized the RGB color
space for color representation, resized the cov er images
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
78
into
224 px × 224 px
, and normalized their RGB v alues to
[0, 1]. W e used a single NVIDIA A6000 GPU to train the
color palette estimator . Our implementation was based on
PyT orch [25]. W e used the mini-batch size of 2,048. T o
train the color palette estimator , we used the Adam opti-
mizer [26] with a learning rate of
1 . 0 × 10 − 4
. W e calculated
the mean squared error (MSE) loss function
L M S E
between
the estimated and ground-truth color palettes to optimize
the parameters of the color palette estimator .
5.4 Experimental Settings
T o clarify which audio feature e xtractor would be most
ef fectiv e in estimating the color palette, we conducted com-
parati ve e xperiments to in vestigate the importance of selec-
tion. Additionally , as a reference, we used color palettes
composed of random colors.
5.4.1 Audio featur e e xtractor
T o extract audio features from musical pieces, we compared
two audio models: CLAP [15] and AudioT oken [16].
T o use CLAP [15], we set the parameters of pretrained
models a vailable at HuggingF ace’ s T ransformers [27] (i.e.,
“laion/clap-htsat-fused”). Each musical piece was con verted
to a mel spectrogram through a CLAP feature extractor ,
and the CLAP audio model used the spectrogram as input.
W e fixed all parameters of the model. By using the model,
we can obtain a 768-dimensional feature vector for each
musical piece.
W e also used AudioT ok en [16], which consists of a bidi-
r ectional encoder r epr esentation fr om audio transformers
(BEA Ts) [23] model and an embedder [16] model. W e
set the parameters of pretrained models a v ailable at offi-
cial GitHub repositories
2
. All parameters of the models
were fixed. By employing the models, we can obtain a
768-dimensional feature vector for each musical piece.
5.4.2 Number of Colors in Color P alette
W e used the number of colors
M = { 1 , 2 , 3 , 4 , 5 }
in the
color palettes for the experiments. T o e xtract the color
palettes from the cov er images, we used the data-driv en
color manifolds [14] as described in Section 3.
5.5 Evaluation Metric f or Comparative Experiments
In our comparati ve e xperiments, we used the minimum
color dif fer ence model (MCDM) [28], which is practically
designed to e v aluate the color difference between tw o color
palettes. The MCDM compares the two color palettes, each
consisting of
M
colors, to determine their av erage color dif-
ference. First, the colors in the color palettes are con verted
from RGB to CIELAB [29]. Then, the MCDM calculates a
CIELAB color dif ference between each color in one palette
and all colors in the other palette, identifying the closest
color match for each and recording the minimum dif fer-
ences. While multiple variants of CIELAB color dif ference
2
The pretrained models are av ailable at
https://github.com/
microsoft/unilm/tree/master/beats
for the BEA Ts model
and
https://github.com/guyyariv/AudioToken
for the em-
bedder model.
T able 2 . Results for the MCDM score on the test set of
our dataset. A lo wer MCDM score indicates a closer match
between the estimated color palettes and the ground-truth
color palettes.
Audio feature extractor M MCDM score
CLAP
1 28 . 37
2 25 . 66
3 22.68
4 21.72
5 21.17
AudioT oken
1 28.40
2 25.69
3 22 . 66
4 21 . 71
5 21 . 14
(Random)
1 69.29
2 68.09
3 66.35
4 65.90
5 65.50
exist, we here adopted the CIE1976 color dif ference [30]
for e valuation. This process is repeated for ev ery color
in the first palette, resulting in
M
color dif ference values,
which are then a veraged to obtain a mean v alue, denoted
as
m 1
. The same process is repeated for the second palette,
finding the closest matches in the first palette and a veraging
the
M
minimum dif ferences to obtain another mean value,
m 2
. Finally , the a verage of
m 1
and
m 2
gi ves the o verall
color dif ference between the two palettes. The lower the
MCDM score, the closer the two color palettes. W e le ver -
age this MCDM to compare a color palette estimated with
each experimental setting and a ground-truth color palette.
5.6 Results
T able 2 presents the results for the MCDM score under each
experimental setting. As sho wn in T able 2, our proposed
frame work achie ves a much lo wer (i.e., better) MCDM
score compared to random color palettes, with an improv e-
ment of ov er 40 points. This demonstrates the ef fectiv eness
of our frame work. Additionally , these results support that
there is a relationship between musical pieces and the color
palettes used for designing their corresponding cov er im-
ages because our frame work succeeds in training the color
palette estimator .
Reg arding the selection of each experimental setting,
there is no performance dif ference between the CLAP and
AudioT oken audio models, as sho wn in T able 2. This sug-
gests that either audio model can be selected based on the
intended application. As these results demonstrate, our
frame work ef fectiv ely estimates appropriate color palettes
for musical pieces.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
79
1
0
1
0
1
1024
0
1
160
2000 1500 1000 500 0 500 1000
0
1
Figur e 1 . Mobile-AMT centers STFT windows on the
time point to predict for , incurring a delay of 1024 samples.
W e shift the window to reduce the delay to 160 samples,
and change the windo w function to better use that limited
amount of future information.
combining weights and shift-tolerance ( T5 ).
For the ne xt experiments, we use binary targets with
weights. While the shift-tolerant loss performed fa v orably
here, our goal is to use a causal model, for which shift tol-
erance could result in systematically delayed predictions.
4.2 A udio Pr eprocessing
Mobile-AMT processes audio in spectrogram frames of
2048-sample windo ws centered on the time points to pre-
dict e vents for . Thus, e ven with a causal model that does
not process information from future frames, in real-time
inference, ev ery time a complete audio buf fer of 2048 sam-
ples is filled, inference will be triggered to predict ev ents
that are already 1024 samples (64 ms) in the past. Figure 1
illustrates this: the top ro w shows an audio w av eform, the
second ro w a typical STFT windo w centered at the onset.
Shorter filters reduce this delay , as the delay is fixed to half
the filter length, b ut this comes at the cost of lower fre-
quency resolution, which we should av oid: W ith a 2048-
sample STFT at 16 kHz, the bin width is 7.8 Hz, which is
already too coarse to achie ve semitone precision for the
lo west piano notes (A0 at 27.5 Hz, B Z 0 at 29.14 Hz).
W e can howe v er reduce this delay to a lo wer number
n s of samples (e.g., 160 samples or 10 ms) while k eeping
the windo w length and frequency resolution unchanged,
by shifting windo ws so they end n s samples after their
reference point instead of being centered. The third ro w
in Figure 1 sho ws the Hann window shifted to n s = 160
samples. The Hann window strongly attenuates the bound-
aries, the right one of which no w contains highly rele-
v ant information for the prediction. T o mitigate this un-
wanted attenuation, we can replace the Hann windo w with
an asymmetric windo w that tapers (2048 − n s ) samples
before and n s samples after the reference point. The last
ro w in Figure 1 illustrates this windowing function for a
1888/160 sample asymmetry . Note ho w we keep more in-
formation from the incoming samples in the gray shaded
area under the windo w function, albeit at the cost of in-
creasing spectral leakage (by about 20 dB for n s = 160 ).
T olerance 10 ms 20 ms 30 ms
H1 Hann 64 ms 17.43 ± 5.10 34.65 ± 7.40 39.86 ± 7.45
H2 Hann 10 ms 0.00 ± 0.01 0.00 ± 0.02 0.04 ± 0.10
T1 asym. 10 ms 22.25 ± 5.17 25.21 ± 5.20 25.84 ± 5.18
T2 asym. 20 ms 28.61 ± 7.09 33.76 ± 7.01 34.65 ± 6.89
T3 asym. 30 ms 27.61 ± 6.79 37.91 ± 7.54 39.43 ± 7.33
T4 asym. 40 ms 24.39 ± 6.50 37.51 ± 7.81 39.87 ± 7.60
T5 asym. 50 ms 20.99 ± 6.02 36.88 ± 7.75 40.47 ± 7.59
ST asym. 10 ms 0.41 ± 0.40 1.57 ± 0.82 11.86 ± 4.25
T able 2 . Note onset F1 scores on the MAESTR O v .3 vali-
dation set for dif ferent windowing functions. The last ro w
additionally uses a shift-tolerant training loss.
T o experiment with dif ferent windowing configurations
for reducing the delay in audio preprocessing, we mod-
ify our reference method to apply only causal processing,
as allo wing the model access to future frames would ren-
der our interventions meaningless. Specifically , we make
each con v olution causal, so the model’ s recepti ve field of
9 frames extends 8 frames into the past, rather than split-
ting 4 frames into the past and 4 frames into the future.
Additionally , we remov e the Squeeze-Excitation layers of
the MobileNet V3 blocks, which perform global av erage
pooling ov er both past and future frames in an excerpt.
T able 2 shows our results. The original centered Hann
windo w with our causal model ( H1 ) performs worse than
our non-causal starting point ( TP3 in T able 1). Shifting the
Hann windo w from a delay of 64 ms to a delay of 10 ms
( H2 ) seems to completely attenuate usable information in
the frames. Using an asymmetric windo w ( T1 ) improv es
performance, but still f alls behind the centered Hann win-
do w . Successiv ely increasing the delay up to 50 ms, we
see a strong improv ement ( T2 – T5 and Figure 2). For
a delay of 30 ms or more, we match performance of the
centered Hann windo w at an ev aluation onset tolerance of
30 ms. For stricter tolerances, shifted asymmetric windows
of 20 ms delay or more surpass the centered Hann windo w .
W e also take the chance to in vestig ate how a shift-
tolerant loss of ± 1 frame af fects results for the causal
model. The loss could allow the model to systematically
predict e vents one frame (10 ms) later than annotated. Sur-
prisingly , using an asymmetric window with 10 ms of de-
lay , we find that the shift-tolerant loss ( ST ) performs on
par with 30 ms delay ( T3 ) when admitting an e valuation
tolerance of 50 ms (not sho wn in table), but breaks do wn
with any stricter tolerance (as seen in the table).
For the third group of e xperiments, we keep the strictest
setting with asymmetric windo ws at a delay of 10 ms.
4.3 Model Architectur e
In our final group of e xperiments, we in vestig ate architec-
tural modifications to our reference model. The architec-
ture of Mobile-AMT consists of three acoustic stacks, each
consisting of recurrent con v olutional blocks. Each stack
learns a (onset, frame or velocity) tar get. For some tar gets,
the outputs of multiple stacks are concatenated to condi-
tion the final predictions. Compared to their reference of-
fline model [2], Mobile-AMT omits the acoustic stack for
the of fset target, reusing the stack of the frame tar get.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
86
10 20 30 50
Onset tolerance (ms)
0.200
0.225
0.250
0.275
0.300
0.325
0.350
0.375
0.400
Mean note onset F1 scor e
Asymmetric window with differ ent amounts of shif t
10ms
20ms
30ms
40ms
50ms
Figur e 2 . Note onset F1 scores (means only) for different
windo w delays and onset tolerance thresholds.
T olerance 10 ms 20 ms 30 ms
Note onset F1 mean ± std
A1 20.23 ± 4.29 23.03 ± 4.25 23.75 ± 4.24
A2 25.77 ± 4.65 30.66 ± 4.98 31.79 ± 4.93
A3 21.59 ± 4.79 24.48 ± 4.83 25.14 ± 4.80
A4 21.53 ± 4.31 24.51 ± 4.24 25.22 ± 4.19
A5 19.44 ± 4.20 22.39 ± 4.34 23.12 ± 4.30
A6 25.82 ± 4.48 30.52 ± 4.51 31.56 ± 4.44
Note onset and of fset F1 mean ± std
A1 3.25 ± 1.31 5.13 ± 2.54 7.20 ± 3.51
A2 3.56 ± 1.29 6.17 ± 2.80 8.39 ± 4.01
A3 5.84 ± 1.41 5.84 ± 2.48 7.71 ± 3.43
A4 5.94 ± 1.31 5.94 ± 2.48 7.89 ± 3.49
A5 4.84 ± 1.11 4.84 ± 2.36 6.61 ± 3.32
A6 6.79 ± 1.15 6.79 ± 2.53 8.89 ± 3.59
T able 3 . Note onset and onset-and-offset F1 scores on the
MAESTR O v .3 v alidation set across three onset tolerances
for architectural modifications and input representations.
Our experiments in volv e the follo wing adaptations,
each tested independently: First, in A1 we (re)introduce a
separate of fset acoustic stack to explore whether and ho w
it improv es of fset label prediction. In A2 we remov e the
velocity conditioning on the onsets. Next, we e xamine
whether further streamlining the architecture by sharing
a fourth ( A3 ), half ( A4 ) or all ( A5 ) of the con volutional
blocks in the model’ s acoustic stacks af fects performance.
Lastly , in A6 we examine the ef fect of training on the orig-
inal 10 seconds sequence length.
T able 3 presents the note onset and note onset-and-
of fset F1 scores of the model and data adaptations on the
MAESTR O v alidation set. Overall, it is e vident that the
combined impact of binary , hea vily imbalanced pointwise
tar gets, causal modeling, and shifted asymmetric window
results in a significantly harder learning problem, with
the same training duration (500 epochs) leading to signif-
icantly poorer scores than the base case ( TP1 in T able 1).
Ho wev er , across all experimental setups in Section 4.1
compared to the current one, all our causal modifications
demonstrate significantly stronger rob ustness to decreasing
tolerance thresholds, which is important to guarantee low
latency in predictions, and therefore appear promising for
further training.
Furthermore, when comparing all model architecture
modifications ( A1-5 ) on note onset and note-onset-and-
of fset F1 score, we observe tw o unexpected model be-
ha viours: First, adding a separate offset stack ( A1 ) does
not improv e of fset prediction. As our postprocessing de-
tects a note of fset as the earlier of either offset acti v ation
or frame inacti vation, we hypothesize that frame acti vity
is suf ficiently learned to compensate for the absence of an
of fset acoustic stack. Second, removing the v elocity con-
ditioning on onset prediction ( A2 ) results in a strong im-
prov ement in onset prediction. Furthermore, sharing the
acoustic stack across increasing proportions ( A3-5 ) does
not appear to hinder the model’ s ability to learn meaning-
ful representations. Finally , experiment A6 suggests that
the model benefits from the lar ger contextual windo w .
4.4 Final comparison
For our final comparison, we proceed with the follo wing
data and model configurations: we continue with the (160
samples) shifted asymmetric windo w for the STFT ( T1 in
Sec. 4.2), remov e the velocity conditioning ( A2 ) and share
all con v olutional layers in the acoustic stack across all tar-
gets ( A5 ). Mobile-AMT uses the original non-causal post-
processing described in Section 3.2, while our model use
the causal postprocessing introduced in Section 4.1.
T able 4 summarizes the results over dif ferent onset (and
of fset) thresholds for note onset and onset-and-offset met-
rics. As expected, Mobile-AMT outperforms our modified
causal model across all metrics, with a significant margin.
Upon re viewing all e xperiments conducted, we conclude
that the lar gest performance drops are attributed to the
shifted windo w function and the causal con volutions in our
model. When comparing Mobile-AMT and our adapted
model across v arious onset tolerance thresholds, we ob-
serve, similar to the pre vious experiment, that while our
modified causal model predicts fe wer targets with lo wer
accuracy o verall, it demonstrates higher precision and ro-
b ustness when ev aluated at stricter timing tolerances.
5. DISCUSSION AND OUTLOOK
In this work, we in v estigate whether and ho w the cur-
rent state of the art in real-time piano transcription can be
adapted to achie ve minimum-latenc y automatic piano tran-
scription suitable for real-time musical interaction.
What latency is suitable cannot be answered uni ver -
sally, so our choice of 10–30 ms is worthy of discus-
sion. While 10 ms is suggested in digital instrument de-
sign [8–10], thresholds for latency perception v ary depend-
ing on the musical situation, task, and instrument: for per -
cussi ve digital instruments, decreased ratings for v alues
of 20 ms and above were found [20], instrument-specific
thresholds between belo w 10 ms and 40 ms are reported
in a li ve monitoring setting [21], and about 30 ms were
found for gestural control [22]. Of fsets as low as 6 ms may
be percei ved in simple isochronously spaced stimuli [23],
while other researchers found just noticeable latency dif-
ferences at 27 ms and higher [24]. T ranscription-enabled
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
87
Note Onset Note Onset with Offset
Model T ol. (ms) Pr ecision Recall F1 Precision Recall F1
Causal-AMT 10 43.51 ± 6.87 25.60 ± 8.64 31.55 ± 7.76 6.86 ± 2.15 4.11 ± 2.03 5.03 ± 2.08
Mobile-AMT 10 22.70 ± 5.82 15.57 ± 2.96 18.26 ± 3.51 3.27 ± 1.19 2.32 ± 0.94 2.69 ± 1.00
Causal-AMT 20 50.86 ± 7.52 29.71 ± 9.11 36.70 ± 7.96 10.89 ± 3.91 6.34 ± 2.90 7.86 ± 3.16
Mobile-AMT 20 59.28 ± 7.41 41.87 ± 8.44 48.52 ± 6.96 11.03 ± 4.15 7.98 ± 3.73 9.16 ± 3.84
Causal-AMT 30 51.85 ± 7.49 30.24 ± 9.07 37.38 ± 7.85 14.49 ± 5.72 8.33 ± 3.93 10.37 ± 4.42
Mobile-AMT 30 81.17 ± 6.29 57.84 ± 12.79 66.80 ± 9.78 18.42 ± 6.89 13.27 ± 6.09 15.26 ± 6.30
T able 4 . Final comparison between our implementation of Mobile-AMT and our modified minimal-latency , strictly causally
adapated model.
real-time applications like interacti ve accompaniment or
generati ve impro visation are more akin to ensemble play-
ing than direct instrument control. In networked musical
contexts, researchers typically aim for 20–30 ms of latency
to meet performance conditions that mirror traditional in-
person ensembles [11, 12]. Ho we ver , studies also found
that musicians may be able to compensate for latencies up
to 50 ms [13, 25] or e ven 100 ms [26] for one piano piece, a
v alue that was deemed “neither musical nor interacti ve” in
another study [27]. A real-time transcription model should
not only be tolerable b ut enable fluent musical interaction,
so we took 30 ms as a minimal requirement, and 10 ms as
a goal for imperceptible latency .
W e in vestigate multiple adaptations to reduce la-
tency , including label encoding with causal postprocess-
ing, shifted asymmetric windo w functions during prepro-
cessing, and architectural modifications that enforce causal
processing within the model. Additionally , we reduce the
model size (from 320 to 160 GFLOPs for 3 seconds of in-
put) by sharing computations across core model compo-
nents for all tar gets.
In a first set of experiments, we assess the impact of
regression v ersus classification loss encodings for non-
causal models. The original regression tar gets only make
sense in conjunction with a lookahead as the targets be gin
to increase se veral frames before the actual onset which
is impossible for a causal model to predict. T o miti-
gate the cost in training stability and accurac y incurred by
localized, causal-ready tar gets, we experiment with loss
functions that weight the acti ve frames o ver the inacti ve
frame to combat label imbalance, and loss functions that
are tolerant to small temporal shifts. W e find that the
weighted classification losses approximate the baseline,
and the shift-tolerant losses reach the same le vel in the ab-
sence of tar gets requiring lookahead.
In a second experiment, we in vestigate the delay in-
curred by the computation of audio feature representations.
Specifically , we look at STFT windo ws and their cor-
responding centered tar gets. T ranscription requires high
frequency resolution for pitch estimation which requires
lar ge windows. Centering the targets results in an often
ov erlooked delay of half the windo w length, 64 ms in our
case. W e test configurations of shifted windo ws along with
asymmetric windo wing functions that do not attenuate the
most recent samples. W e find that aggressiv ely shifted
windo ws at 10 ms do deteriorate the transcription accuracy
by a lot, yet at 30 ms, we reach comparable performance to
an unshifted causal model. Here, a shift-tolerant loss does
not improv e performance. At the same time, configuring
the model architecture for strictly causal processing also
deteriorates performance with respect to the baseline with
more than 100 ms of lookahead.
In a third experiment, we assess dif ferent model archi-
tectures and their impact. W e observe that sharing the con-
v olutional components of the acoustic stack across differ -
ent tar get types proves beneficial. W e hypothesize that the
local acoustic features captured in the con v olutional layers
of the acoustic model can be ef fectiv ely learned indepen-
dently of sequential information, making them in v ariant to
the tar get type. Furthermore, removing the velocity con-
ditioning on the onsets strongly improv es the accuracy of
onset predictions.
Overall, we find that we can compensate well for algo-
rithmic issues: we can scale the model and use lookahead-
free tar gets without a major drop in performance. What
prov es dif ficult, howe ver , is to render the model strictly
causal and to ef fectiv ely process the incoming audio with-
out loss of rele vant information. For a latenc y of 10 ms, it
would be required that the model predicts pitches with at
most 10 ms of incoming audio samples. For onsets of the
lo west two octa ves on the piano, this means that there is not
e ven a full period of the fundamental frequency present in
the samples, and predictions may need to rely on harmonic
partials. Along with the transient phase and the conse-
quently blurry STFT frame, this leads to an increasingly
hard transcription task. W e hope that these findings and
pinpointed challenges will contrib ute to future research on
real-time, minimum latency automatic piano transcription.
While this study primarily focuses on the algorithmic
performance and rob ustness of a real-time transcription
model, we ackno wledge that a detailed analysis of pro-
cessing time—including both network inference and pre-
processing—across dif ferent hardware platforms remains
an important area for future work to allo w for the practi-
cal deployment of a real-time transcription system in real-
world scenarios. Like wise, we want to take a closer in-
spection into the design of the underlying windo w func-
tion and filter bank, in order to find an appropriate balance
between reducing prediction delay and increasing future
context, all while maintaining the desired STFT properties.
Lastly , across all experimental groups, our system adapta-
tions consistently outperformed the baseline at lo wer tim-
ing tolerance, which we consider a desirable property wor -
thy of further in vestigation.
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
88
6. A CKNO WLEDGEMENTS
This research ackno wledges support by the European Re-
search Council (ERC), under the European Union’ s Hori-
zon 2020 research and innov ation programme, grant agree-
ment No. 101019375 Whither Music? . The LIT AI Lab is
supported by the Federal State of Upper Austria.
7. REFERENCES
[1] C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-
Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and
D. Eck, “Enabling factorized piano music modeling
and generation with the MAESTR O dataset, ” in In-
ternational Confer ence on Learning Repr esentations ,
2019.
[2] Q. K ong, B. Li, X. Song, Y . W an, and Y . W ang, “High-
Resolution Piano T ranscription with Pedals by Re-
gressing Onset and Of fset T imes, ” IEEE/A CM T rans-
actions on A udio Speec h and Languag e Pr ocessing ,
v ol. 29, pp. 3707–3717, 2021.
[3] S. Sigtia, E. Benetos, and S. Dixon, “ An end-to-end
neural network for polyphonic piano music transcrip-
tion, ” IEEE/A CM T ransactions on A udio, Speech, and
Languag e Pr ocessing , v ol. 24, no. 5, pp. 927–939,
2016.
[4] R. K elz, S. Böck, and G. W idmer , “Deep polyphonic
adsr piano note transcription, ” in ICASSP 2019-2019
IEEE International Confer ence on Acoustics, Speec h
and Signal Pr ocessing (ICASSP) . IEEE, 2019, pp.
246–250.
[5] Y . K usaka and A. Maezawa, “Mobile-AMT : Real-
T ime Polyphonic Piano T ranscription for In-the-W ild
Recordings, ” in 2024 32nd Eur opean Signal Pr ocess-
ing Confer ence (EUSIPCO) . IEEE, 2024, pp. 36–40.
[6] A. Fernandez, “Onsets and V elocities: Af fordable
Real-T ime Piano T ranscription Using Con v olutional
Neural Networks, ” in 2023 31st Eur opean Signal Pr o-
cessing Confer ence (EUSIPCO) . IEEE, 2023, pp.
151–155.
[7] T . Kw on, D. Jeong, and J. Nam, “T o wards Ef ficient and
Real-T ime Piano T ranscription Using Neural Autore-
gressi ve Models, ” IEEE/A CM T ransactions on A udio,
Speech, and Langua ge Pr ocessing , 2024.
[8] D. W essel and M. Wright, “Problems and prospects for
intimate musical control of computers, ” Computer mu-
sic journal , v ol. 26, no. 3, pp. 11–22, 2002.
[9] A. P . McPherson, R. H. Jack, and G. Moro, “ Action-
sound latency: Are our tools f ast enough?” in
16th International Confer ence on Ne w Interfaces for
Musical Expr ession, NIME 2016, Griffith University ,
Brisbane, Austr alia, J uly 11-15, 2016 . nime.org,
2016, pp. 20–25. [Online]. A v ailable: https://doi.or g/
10.5281/zenodo.3964611
[10] F . Caspe, J. Shier , M. Sandler , C. Saitis, and
A. McPherson, “Designing neural synthesizers for lo w
latency interaction, ” arXiv pr eprint arXiv:2503.11562 ,
2025.
[11] L. T urchet and C. Rottondi, “On the relation between
the fields of network ed music performances, ubiqui-
tous music, and internet of musical things, ” P ersonal
and Ubiquitous Computing , vol. 27, no. 5, pp. 1783–
1792, 2023.
[12] E. Lakiotakis, C. Liaskos, and X. Dimitropoulos, “Im-
proving netw orked music performance systems us-
ing application-network collaboration, ” Concurr ency
and Computation: Practice and Experience , vol. 31,
no. 24, p. e4730, 2019.
[13] E. Che w , R. Zimmermann, A. A. Sawchuk, C. K yr-
iakakis, C. P apadopoulos, A. François, G. Kim,
A. Rizzo, and A. V olk, “Musical interaction at a dis-
tance: Distributed immersi ve performance, ” in Pr o-
ceedings of the MusicNetwork F ourth Open W orkshop
on Inte gr ation of Music in Multimedia Applications .
MusicNetwork Barcelona, 2004, pp. 15–16.
[14] T . Kw on, D. Jeong, and J. Nam, “Polyphonic Piano
T ranscription Using Autoregressi v e Multi-State Note
Model, ” in The 21th International Society for Music
Information Retrieval Confer ence (ISMIR) . Interna-
tional Society for Music Information Retrie val, 2020.
[15] D. Jeong and S. T elecom, “Real-time automatic piano
music transcription system, ” in Late Br eaking Demo.
International Society for Music Information Retrieval ,
2020, pp. 4–6.
[16] A. Ho ward, M. Sandler , G. Chu, L.-C. Chen, B. Chen,
M. T an, W . W ang, Y . Zhu, R. Pang, V . V asude van et al. ,
“Searching for MobileNetV3, ” in Pr oceedings of the
IEEE/CVF international confer ence on computer vi-
sion , 2019, pp. 1314–1324.
[17] C. Raf fel, B. McFee, E. J. Humphrey , J. Salamon,
O. Nieto, D. Liang, D. P . Ellis, and C. C. Raffel,
“MIR_EV AL: A T ransparent Implementation of Com-
mon MIR Metrics.” in ISMIR , v ol. 10, 2014, p. 2014.
[18] R. M. Bittner , J. J. Bosch, D. Rubinstein, G. Meseguer -
Brocal, and S. Ewert, “ A lightweight instrument-
agnostic model for polyphonic note transcription and
multipitch estimation, ” in Pr oceedings of the IEEE In-
ternational Confer ence on Acoustics, Speech, and Sig-
nal Pr ocessing (ICASSP) , Singapore, 2022.
[19] F . Foscarin, J. Schlüter , and G. W idmer , “Beat this!
Accurate beat tracking without DBN postprocessing, ”
arXiv pr eprint arXiv:2407.21658 , 2024.
[20] R. H. Jack, A. Mehrabi, T . Stockman, and A. McPher-
son, “ Action-sound latenc y and the perceiv ed quality of
digital musical instruments: Comparing professional
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
89
percussionists and amateur musicians, ” Music P er cep-
tion , vol. 36, no. 1, pp. 109–128, 09 2018. [Online].
A v ailable: https://doi.or g/10.1525/mp.2018.36.1.109
[21] M. Lester and J. Boley , “The effects of latenc y on li ve
sound monitoring, ” Journal of the A udio Engineering
Society , no. 7198, october 2007.
[22] T . Mäki-P atola and P . Hämäläinen, “Latenc y tolerance
for gesture controlled continuous sound instrument
without tactile feedback, ” in Pr oceedings of the 2004
International Computer Music Confer ence, ICMC
2004, Miami, Florida, USA, November 1-6, 2004 .
Michigan Publishing, 2004. [Online]. A vailable:
https://hdl.handle.net/2027/spo.bbp2372.2004.032
[23] A. Friberg and J. Sundber g, “T ime discrimination in
a monotonic, isochronous sequence, ” The Journal of
the Acoustical Society of America , vol. 98, no. 5, pp.
2524–2531, 1995.
[24] A. Schmid, M. Ambros, J. Bogon, and R. W immer ,
“Measuring the just noticeable dif ference for audio
latency , ” in Pr oceedings of the 19th International
A udio Mostly Confer ence: Explorations in Sonic
Cultur es, AM 2024, Milan, Italy , September 18-
20, 2024 , L. A. Ludovico and D. A. Mauro,
Eds. A CM, 2024, pp. 325–331. [Online]. A v ailable:
https://doi.or g/10.1145/3678299.3678331
[25] S. Dahl and R. Bresin, “Is the player more influenced
by the auditory than the tactile feedback from the in-
strument, ” in Pr oceedings of the Digital A udio Effects
Confer ence (D AFx) , 2001, pp. 6–9.
[26] A. A. Sawchuk, E. Chew , R. Zimmermann, C. Pa-
padopoulos, and C. Kyriakakis, “From remote media
immersion to distrib uted immersiv e performance, ”
in Pr oceedings of the 2003 A CM SIGMM W ork-
shop on Experiential T elepr esence , ser . ETP ’03.
Ne w Y ork, NY , USA: Association for Computing
Machinery , 2003, p. 110–120. [Online]. A v ailable:
https://doi.or g/10.1145/982484.982506
[27] C. Bartlette, D. Headlam, M. Bocko, and G. V elikic,
“Ef fect of network latency on interacti v e musical
performance, ” Music P er ception , vol. 24, no. 1,
pp. 49–62, 09 2006. [Online]. A v ailable: https:
//doi.or g/10.1525/mp.2006.24.1.49
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
90
MA TCHMAKER: AN OPEN-SOURCE LIBRAR Y FOR REAL-TIME PIANO
SCORE FOLLO WING AND SYSTEMA TIC EV ALU A TION
Jiyun Park 1 ∗ Carlos Cancino-Chacón 2 ∗
Suhit Chiruthapudi 2 J uhan Nam 1
1 Graduate School of Culture T echnology , KAIST , South K orea
2 Institute of Computational Perception, Johannes K epler Uni versity Linz, Austria
{june,juhan.nam}@kaist.ac.kr,
{carlos.cancino_chacon,suhit.chiruthapudi}@jku.at
ABSTRA CT
Real-time music alignment, also kno wn as scor e follow-
ing , is a fundamental MIR task with a long history and is
essential for many interacti v e applications. Despite its im-
portance, there has not been a unified open frame work for
comparing models, lar gely due to the inherent complex-
ity of real-time processing and the language- or system-
dependent implementations. In addition, low compatibil-
ity with the existing MIR en vironment has made it diffi-
cult to de velop benchmarks using lar ge datasets av ailable
in recent years. While ne w studies based on established
methods (e.g., dynamic programming, probabilistic mod-
els) ha ve emer ged, most e valuations compare models only
within the same family or on small sets of test data. This
paper introduces Matchmak er , an open-source Python li-
brary for real-time music alignment that is easy to use and
compatible with modern MIR libraries. Using this, we sys-
tematically compare methods along two dimensions: mu-
sic representations and alignment methods. W e e v aluated
our approach on a lar ge test set of solo piano music from
the (n)ASAP , Batik, and V ienna4x22 datasets with a com-
prehensi ve set of metrics to ensure rob ust assessment. Our
work aims to establish a benchmark frame work for score-
follo wing research while providing a practical tool that de-
velopers can easily inte grate into their applications.
1. INTR ODUCTION
Real-time music alignment, also kno wn as scor e follow-
ing , is the task of aligning performance data to the corre-
sponding position in the musical score in real-time. Ever
since it was first introduced independently by Roger Dan-
nenber g [1] and Barry V ercoe [2] ov er 40 years ago, music
alignment has become one of the fundamental MIR tasks.
* Equal contribution.
© J. Park, C. Cancino-Chacón, S. Chiruthapudi and J. Nam.
Licensed under a Creati ve Commons Attribution 4.0 International Li-
cense (CC BY 4.0). Attribution: J. Park, C. Cancino-Chacón, S.
Chiruthapudi and J. Nam, “Matchmaker: An Open-Source Library for
Real-T ime Piano Score Follo wing and Systematic Evaluation”, in Pr oc.
of the 26th Int. Society for Music Information Retrieval Conf., Daejeon,
South K orea, 2025.
Score follo wing is a necessary component of many inter -
acti ve applications (e.g., automatic accompaniment sys-
tems [3–6], automatic page turning [7, 8], lyrics align-
ment or tracking singing v oice [9–11], audiovisual/mul-
timodal [6, 12] and visualizations [13]. Music alignment
beg an as real-time score following [1, 2, 14–17] b ut, by the
mid-90s, had div erged into online and of fline methods (see,
e.g., early of fline work by Desain et al. [18]).
From its early use on monophonic sources like v oice
[17] and wind instruments, score following has gro wn to
support polyphonic instruments such as piano, ensemble,
and e ven full orchestral performances [17, 19–21]. Re-
search has also expanded across input modalities of the
performance, with systems operating on audio or MIDI,
and score representations including string format, sym-
bolic score, and sheet image [22].
The score follo wing challenge [23] in MIREX laid
the foundation to formalize the e valuation frame work, in-
troducing important metrics that include considerations
in real-time. Ho we ver , many subsequent studies ha ve
been de veloped in dif ferent en vironments—ranging from
system-dependent [24, 25] to language-dependent [26, 27]
implementations—often tailored to specific use cases and
without publicly shared source code. As a result, imple-
mentations became fragmented across platforms, making
it dif ficult to extend, reproduce, or compare methods in a
unified setting. This has hindered the de velopment of a
unified e valuation frame work and comparison o ver meth-
ods or features on shared datasets remain rare, limiting the
generalizability and reproducibility .
In this paper , we address these challenges by proposing
a unified, open frame work for the e valuation and bench-
marking of real-time audio-based score follo wing. Consid-
ering public datasets that of fer a range of dif ficulty le vels,
multiple renditions, and precise beat-le vel annotations, we
base our e valuation on three representati v e piano perfor-
mance datasets. W e implement this framew ork as an open-
source Python package called Matchmak er , 1 that allows
real-time ex ecution of representati ve baselines of score fol-
lo wing algorithms. In addition to benchmarking, it sup-
ports audio de vice input and has been validated in applica-
tion contexts through a standalone demo system.
1 https://github.com/pymatchmaker/matchmaker
91
2. A CONCEPTU AL FRAMEW ORK FOR SCORE
FOLLO WING
As a way to or ganize and compare the components of sys-
tems for score follo wing, we follow the structure proposed
by Müller [28]. This frame work consists of three core
components: (1) input music representations, (2) features,
and (3) online alignment algorithms.
2.1 Music Representation
Score follo wing aligns a fixed reference deri ved from mu-
sical scores with a time-e volving input from a perfor -
mance. The score can take v arious symbolic formats (e.g.,
MIDI, MusicXML) or sheet images, and is typically con-
verted into an intermediate representation such as synthe-
sized audio or e vent sequences. The performance input
may be gi ven as either audio or MIDI, each with distinct
representational and computational characteristics. Au-
dio input is continuous and latency-sensiti ve, while MIDI
is discrete and e vent-based. Instrumental factors also af-
fect alignment design: polyphonic or discrete-pitch instru-
ments (e.g., piano) dif fer from continuous-pitch sources
(e.g., violin, voice). Multi-instrument recordings pose fur -
ther challenges due to timbral ov erlap and source ambigu-
ity .
2.2 F eatures
Chroma features are the most commonly used in music
synchronization, with many v ariants for their computa-
tion [29–32]. Other works also use v arious spectral fea-
tures such as constant-Q transforms (CQT) [27, 33], non-
neg ativ e matrix factorization(NMF)-based [34] or spectral
template [35] for improv ed polyphonic alignment. Beyond
spectral representations, context-a ware features such as
onset-based feature [36] or beat-synchronous frames ha ve
been introduced to capture temporally salient e vents use-
ful for alignment. Later work e xplored learned features,
including feedforward mappings [27], semi-supervised
decompositions like NMF , and more recent neural ap-
proaches [37]. While these of fer richer contextual infor -
mation, they often rely on fix ed-length inputs and intro-
duce latency , making real-time usage more challenging.
2.3 Alignment Algorithms
T wo major families of alignment algorithms ha ve been
used in score follo wing: dynamic programming and prob-
abilistic models.
The dynamic programming approach, especially dy-
namic time warping (DTW), aligns tw o sequences by min-
imizing cumulati ve cost. Its online v ariant, On-Line T ime
W arping (OL TW) [38], enables causal alignment within
a fixed-size of windo w . V ariants include windo wed [39],
parallel [40], and constrained DTW [40, 41], as well as
tempo-aw are extensions [21, 42].
Probabilistic state-space models of fer an alternativ e by
treating alignment as latent state inference under uncer -
tainty [24, 29, 43]. HMM-based systems model each note
as a sequence of states (e.g., attack–steady–release), with
extensions including semi-Mark ov [44], hybrid [19], and
Bayesian v ariants [45]. Kalman filter models and switch-
ing state-space systems [46, 47] further incorporate tempo
dynamics, while particle filters [12, 29] handle multimodal
uncertainty in real time.
Other paradigms include early string-matching algo-
rithms [1] and reinforcement learning-based approaches
for multimodal or visual score alignment [48].
3. IMPLEMENT A TION
3.1 Python Package Structure
Matchmak er is an open source Python package that imple-
ments representati ve real-time music alignment algorithms
within a modular , e xtensible framew ork. Figure 1 illus-
trates the ov ervie w of the package and the whole pipeline.
The current version of Matc hmaker pro vides two types
of algorithms: 1) online time warping, with two v ariants:
OLTWDixon , based on the methods proposed in [38, 49],
and OLTWArzt , based on [21, 50]; and 2) an HMM-based
algorithm, similar to the one used in [3, 47]. A full descrip-
tion of the algorithms and their parameters can be found in
the supplementary Appendix. 2
Matchmak er supports two main usage scenarios: (1)
li ve streaming mode using the audio de vice and (2) sim-
ulation mode, which processes a performance file as in-
put. Figure 2 shows an e xample of running liv e streaming
mode with the default setting. The AudioStream object
handles the input stream by chunking the audio with o ver -
lapping windo ws to av oid padding artifacts. Both the syn-
thesized score audio and the performance audio are passed
to a Processor object that performs feature extraction.
The extracted features are pushed into a queue and con-
sumed by the OnlineAlignment object, which runs the
alignment methods in real time. Matchmaker tak es a mu-
sical score with all symbolic music formats (MusicXML,
MIDI, MEI, etc.) av ailable by partitura . 3 The returned
output is the current position in the score, represented in
beats as a musical unit according to the time signature in
the piece. More detailed description and API documenta-
tion of the package are av ailable here. 4
3.2 Design and Implementation Details
W e provide a simple and user -friendly interface to run
the score follo wing with minimal setup. As shown in Fig-
ure 2, users can instantiate a Matchmaker object with
a score file and ex ecute a run that iterates ov er the esti-
mated score position for each step. T o streamline real-time
processing, the AudioStream class is implemented as a
context manager that automatically handles stream initial-
ization and teardo wn. Furthermore, the alignment process
is designed as a generator , enabling users to recei ve score
positions concurrently while the alignment is in progress.
2 https://pymatchmaker.github.io/ismir2025_
supplementary_materials/
3 https://github.com/CPJKU/partitura
4 https://pymatchmaker.readthedocs.io/
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
92
Figur e 1 . Ov ervie w of the score follo wing package
1 from matchmaker import Matchmaker
2
3 mm = Matchmaker(
4 score_file= "path/to/score.musicxml" ,
5 input_type= "audio" ,
6 )
7 for current_position in mm.run():
8 print (current_position)
Figur e 2 . A code e xample for running the Matc hmak er in
a li v e streaming mode.
This design allo ws for ef ficient real-time inte gration with-
out requiring users to manage multiple threads, b uf fers, or
callbacks e xplicitly .
While the online mode uses a multi-threaded queue for
asynchronous audio b uf fering, the simulation mode pro-
cesses audio chunks in adv ance within a single-threaded
setup. By decoupling real-time I/O concerns from core
alignment e v aluation, it is intended to a v oid v ariability
from Python v ersion, OS-le v el threading, or queuing de-
lays, ensuring a consistent and reproducible benchmarking
en vironment. In addition, OL TW Arzt is implemented in
Cython [51] for ef ficienc y , a superset of Python designed
for C-lik e performance by incorporating C data types and
optimizing the e x ecution of Python code.
4. EXPERIMENTS
4.1 Datasets
W e use three public piano performance datasets: (n)ASAP
[52], Batik [53] and V ienna 4x22 [54], each of them of fer -
ing complementary characteristics for benchmarking score
follo wing. (n)ASAP , a subset of the MAESTR O dataset
including note-le v el score alignments, includes e xpressi v e
performances of technically demanding solo piano pieces,
of fering high dif ficulty and stylistic di v ersity . W e use only
the pieces i n the MAESTR O v2 test split. V ienna4x22
pro vides 22 distinct renditions for each of four relati v ely
easy pieces, which is suitable to test rob ustness to inter -
preti v e v ariation. Batik dataset contains recordings of 12
Mozart sonatas by a single pianist with the longest a v erage
piece duration among the three datasets, enabling e v alua-
tion across long-form classical repertoire.
W e use ground-truth beat-le v el annotations pro vided
with the (n)ASAP dataset, and e xtract equi v alent annota-
Dataset #Pieces #P erf #Beats #Notes Dur (h) Difficulty
(n)ASAP 43 59 26,329 100,958 2.65 6.53
Batik 30 30 18,789 102,421 2.85 5.67
V ienna 4 88 13,728 43,656 2.24 4.88
T otal 77 177 58,846 247,035 7.74 6.11
T able 1 . Datasets used in the e v aluation.
tions for Batik and V ienna4x22 from the .matc h files [55],
which contain note-wise score–performance alignments.
In addition, we incorporate the dif ficulty le v els of each
piece based on G. Henle Publishers, 5 which pro vides a
1-to-9 grading scale. The pieces used in our e xperiments
span le v els 4 through 9, representing a di v erse set of w orks
abo v e intermediate le v el. T able 1 pro vides the detailed
statistics of the datasets.
W e only included performances in the e xperiment that
recorded an MAE of less than 100 m s in the of fline test,
using the synctoolbox 6 with Chroma & DLNCO features.
The e v al uation w as conducted on 184 performances across
93 pieces, totaling o v er 58,000 beats and 247,000 notes,
with an o v erall duration of 7.74 hours of performances and
a piece-wise a v erage dif ficulty of 6.11.
4.2 Experiment Settings
W e conducted all e v aluations under simulation-based con-
ditions to ensure reproducibility . Li v e testing w as a v oided
due to v ariability introduced by room acoustics and hard-
w are setup, which complicates f air comparison across sys-
tems. The accurac y tests were carried out on an Intel i9-
9900K CPU (16 cores @ 3 . 6 G Hz ), Python 3.9, with a
sample rate of 44.1 kHz and a frame rate of 30, chosen to
balance latenc y and alignment accurac y . W e tested chro-
magram, mel-spectrogram, constant-Q transform (CQT),
mel-frequenc y cepstral coef ficients (MFCCs) [56] and a
simple STFT -based onset-sensiti v e representation similar
to the one used in Dixon [38], which we name log-spectral
ener gy (LSE). While results for all features were e v aluated,
we report detailed latenc y and accurac y metrics for the
best-performing configuration of each model. T o account
for hardw are v ariability , latenc y w as measured in multiple
setups: an Intel i9-9900K, an Apple M4 MacMini, and an
5 https://www.henle.de/Levels- of- Difficulty/
6 https://github.com/meinardmueller/synctoolbox
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
93
Figur e 3 . T w o e xamples of error calculation using the
mapping function. (a) sho ws a one-to-man y alignment at
the e v aluation point, while (b) illustrates a skipped align-
ment.
Apple M2 Pro MacBook, with the reported latenc y v alues
a v eraged across these de vices.
4.3 Pr epr ocessing
In the preprocessing step (see Fig. 1), the symbolic scores
are synthesized to audio using FluidSynth, pro vided by
partitur a . Since MusicXML often lacks tempo markings,
we set the synthesis tempo to each performance’ s a v er -
age—rounded to the nearest 20 BPM—assuming perform-
ers follo w approximate tempo indications.
T o generate beat annotations for the synthesized score
audio, we computed beat positions using the synthesis
tempo and the score’ s time signature. F or compound
meters (e.g., 6/8, 9/8, and 12/8), we adopted (n)ASAP’ s
beat annotation rules—counting them as tw o, three, and
four beats per measure, respecti v ely—across all datasets to
align score-side annotations with performance annotations.
Based on the synthesized audio, we then e xtract the feature
using the same Processor used in the online phase, b ut
precompute them of fline for the entire score sequence.
5. EV ALU A TION
Ev aluating score follo wing is challenging due to causality ,
timing precision, and output latenc y . Since the MIREX
challenge [23] pro vided foundational metrics, later studies
introduced alternati v e e v aluation strate gies including beat-
le v el e v aluations or asynchron y [3], reflecting the task’ s
frequent inte gration with automatic accompaniment sys-
tems.
In this w ork, we adopt tw o complementary e v aluation
perspecti v es. First, we e v aluate in the performance do-
main, where errors are measured in milliseconds based on
ground-truth annotations aligned to the audio. This ap-
proach is commonly used in audio-to-score alignment re-
search and enables precise, fram e-le v el e v aluation, since
the annotations directly reflect the actual timing of the per -
formance. Second, we also e v aluate in the score domain
measured in beat units as suggested in [29, 57], which bet-
ter reflects the nature of score follo wing as a task of pre-
dicting the corresponding score position at each moment
of the performance.
Figur e 4 . Defined delay types of the system. Only system
delay is considered in the e xperiment.
5.1 Ev aluation Metrics
W e select e v aluation metrics mostly adapted from score
follo wing MIREX benchmark [23] and audio-to-score
alignment (ASA) metrics [57]. W e use Alignment Rate
(AR) within a tolerance range of | θ e | , v arying from 50 ms
to 2000 ms. W e also compute Absolute Err ors (AE), both
in milliseconds and in beats, from which we deri v e the
A v erage Absolute Err or (AAE) and Median Absolute
Err or (MAE), along with the standard de viation σ e . T o
further characterize the distrib ution of errors, we report
kurtosis and sk ewness which capture the peak edness and
asymmetry of the non-absolute error distrib uti o n, respec-
ti v ely . In addition, we report the a v erage latency µ l a t , de-
fined as the system delay from the detection time to the end
of inference. Unlik e total latenc y , this e xcludes har dwar e
latency and is composed of tw o parts: (i) feature process-
ing and (ii) e x ecution of the online alignment algorithm
for each frame step (see Fig. 4). Errors e xceeding 2 sec-
onds (or 2 beats in the score domain) are e xcluded from
AE calculations, including both AAE and MAE, to a v oid
distortion from unbounded tracking f ailur es. W e report AR
in tw o w ays. The a v eraged piece-wise AR is a common
measure, while the total AR reflects the proportion of suc-
cessfully aligned beat e v ents across the entire dataset. The
latter a v oids o v errepresentation of shorter pieces and pro-
vides a more balanced vie w of o v erall performance.
T o e v aluate runtime latenc y under simulation, we mea-
sure tw o components: the a v erage duration for e xtracting
features from incoming audio frames, and the time tak en
by the alignment process to consume features and predict
score positions. Specifically , the latenc y w as computed
from the moment audio w as read to the time the score posi-
tion w as predicted—e xcluding hardw are I/O delays. This
tw o-step measurement allo ws for standardized latenc y re-
porting independent of the hardw are setup.
5.2 Alignment Mapping Function
Gi v en t h e alignment path, the alignment mapping function
is applied to transfer the beat positions on one axis (ei-
ther performance or score) to another axis to compute the
alignment error . Due to the local , stepwise nature of real-
time alignment, the resulting path is not necessarily mono-
tonic and may contain multiple correspondents or skipped
positions, depending on the impleme n t ation and purpose
of the methods. Un l ik e linear interpolation methods com-
monly used in of fline audio-to-score alignment, which as-
sume continuous mappings, our e v aluation relies only on
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
94
Dataset Method AAE(ms) ↓ ± σ MAE(ms) ↓ Skew . Kurt. Piece-wise AR (%) ↑ T otal AR ↑
( ≤ 2000ms, %)
≤ 50ms ≤ 100ms ≤ 500ms ≤ 1000ms ≤ 2000ms
(n)ASAP OL TWDixon 189.55 ± 281.55 97.09 3.20 17.97 40.3 58.5 82.5 88.3 92.0 89.4
OL TW Arzt 183.56 ± 263.95 91.18 0.75 11.79 44.1 58.3 84.8 92.0 95.1 92.8
HMM 487.73 ± 423.27 346.01 0.18 3.33 15.6 22.2 37.5 43.8 43.8 43.8
Batik OL TWDixon 186.97 ± 262.55 104.40 3.75 24.70 28.2 51.7 82.1 85.2 87.6 89.4
OL TW Arzt 193.36 ± 269.13 107.15 1.00 12.63 35.9 53.0 82.2 87.4 90.3 89.7
HMM 693.63 ± 376.58 641.77 0.11 0.98 4.5 10.8 34.0 46.2 64.2 61.9
V ienna4x22 OL TWDixon 285.43 ± 390.82 132.73 1.57 5.90 26.6 43.2 72.4 80.0 85.5 82.5
OL TW Arzt 300.41 ± 368.70 152.51 0.50 3.93 33.2 44.5 73.3 84.3 86.7 86.7
HMM 439.64 ± 427.02 319.13 0.15 3.79 23.5 33.3 51.1 57.1 63.0 75.9
T able 2 . Ev aluation results on three datasets using dif ferent score-following methods. The piece-wise alignment rate (AR)
is measured as the a verage ov er pieces, while the total AR indicates the global proportion of aligned beat ev ents across the
entire dataset. All tests were conducted with STFT -based Chroma as features.
predictions made prior to or at each e valuation time point.
T o reflect this, we define the mapping function as follows:
ˆ u k = min u i | ( u i , v i ) ∈ W , v i = max { v j | v j ≤ k } ,
where W = { ( u i , v i ) } is the warping path e xpressed in the
frame indices: u i is the score-rendered-audio frame index
and v i is the performance-audio frame index. The inner
max finds the latest performance frame v i not exceeding
the current frame k , and the outer min selects the smallest
score frame u i among those alignments. This mapping re-
lies solely on past or current frames to maintain causality .
It handles skipped or one-to-many mappings and a v oids
any interpolation methods that depend on future frames.
6. RESUL TS
T able 2 presents a comparison of alignment methods based
on performance-domain e valuation, measured in millisec-
onds. All methods exhibit positi v e ske wness in error
distrib ution, reflecting the expected lag of the beat esti-
mates in real-time alignment. The ov erall results sho w that
the OL TW -based method outperforms the HMM baseline
across all datasets in both alignment accuracy and co ver -
age. While OL TWDixon and OL TW Arzt sho w compa-
rable MAE depending on the dataset, OL TW Arzt consis-
tently achie ves higher co verage ( T otal AR ), suggesting that
it is more rob ust against o verall failures. The dif ference
likely stems from OL TWDixon skipping uncertain re gions,
while OL TW Arzt’ s “backward-forward” strate gy corrects
early misalignments and enhances cov erage. Despite ha v-
ing the lo west AR, the HMM sho ws the lowest sk ewness
and kurtosis primarily because significant errors (>2 s) are
excluded from the summary statistics and its “stick y” be-
ha vior to linger in the same state in local regions tends to
narro w the error distribution.
T able 3 presents an ev aluation comparison in beat units,
of fering a tempo-normalized perspectiv e. The overall
trend mirrors the performance-domain results in millisec-
ond, but these results are standardized across tempi. AAE
remains around 0.3 beats, with median v alues typically be-
lo w 0.2. T otal AR is consistently lo wer than the 2000 ms -
Dataset Method AAE ↓ (beats) ± σ MAE ↓ (beats) AR ↑ (%)
(n)ASAP OL TWDixon 0.22 ± 0.27 0.13 83.4
OL TW Arzt 0.27 ± 0.30 0.16 85.2
HMM 0.80 ± 0.54 0.66 76.9
Batik OL TWDixon 0.20 ± 0.27 0.11 88.9
OL TW Arzt 0.29 ± 0.34 0.18 88.8
HMM 0.80 ± 0.38 0.67 59.3
V ienna4x22 OL TWDixon 0.31 ± 0.33 0.19 78.3
OL TW Arzt 0.37 ± 0.38 0.24 84.0
HMM 0.76 ± 0.78 0.51 70.3
T able 3 . Beat-lev el e valuation results including total align-
ment rate (AR) (%).
F eature Process Online Alignment
T ype MAE (ms) Latency (ms) Method Latency (ms)
Chroma 265.50 3.05 OL TWDixon 1.22
mel 297.92 3.40 OL TW Arzt 0.07
CQT 341.25 42.58 HMM 3.59
LSE 241.85 0.91
MFCC 931.81 2.58
T able 4 . Comparison of feature types and alignment meth-
ods in terms of alignment error (MAE) and latenc y . LSE
is log-spectral ener gy feature that was adopted in [38]. La-
tency v alues are a veraged o ver the hardware setups e v alu-
ated in Section 4.
based metric, reflecting that most pieces ha ve tempi abo ve
60BPM, where two beats span less than tw o seconds.
In addition, a comparison of v arious feature types and
latencies of the alignment methods are reported in T able 4.
Among the features, log-spectral energy (LSE) sho ws the
lo west MAE ( 241 . 85 ms ) and delay ( 0 . 91 ms ), indicat-
ing strong performance with minimal ov erhead. In con-
trast, CQT and MFCC yield higher MAE, with CQT also
requiring considerable extraction time ( 42 . 58 ms ), which
limits its real-time suitability . For alignment methods,
OL TW Arzt achie ves the lo west latency ( 0 . 07 ms ), whereas
HMM sho ws noticeably higher delay ( 3 . 59 ms ) due to its
computational complexity . These results highlight a trade-
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
95
Figur e 2 : Br eath Cue and Onset Annotation Breath cue
is defined by the breath onset and of fset annotated on the
mel spectrogram. Six mark ers are annotated per trial. De-
tailed information can be found in Section 3.2
de grees in performance. Each flutist-pianist pair had not 173
pre viously performed together . 174
3.1.2 Placement and Equipment 175
The setup mimick ed a concert stage, with flutists posi- 176
tioned f acing a w ay from the pianists. The pianists could 177
observ e the flutists, while the flutists were instructed to 178
gi v e cues without turning or looking at the pianists. A 179
camera recorded flutists at 60 fps 2 , and microphones sep- 180
arately captured audio from each instrument at 44.1 kHz. 181
A neutral-colored screen behind flutists minimized back- 182
ground interference for accurate f acial mo v ement tracking. 183
3.1.3 Musical Piece 184
This dataset comprises simultaneously starting flute–piano 185
duets. P art 1 included a C major scale and P achelbel’ s 186
Canon performed at slo w (50 BPM), medium (100 BPM) 187
and f ast (150 BPM) tempos, each repeated twice as w arm- 188
up e x ercises (12 trials). P art 2 consisted of 18 classical 189
pieces simplified for piano and arranged for simultaneous 190
starts. Each piece w as assigned a specific tempo (50, 100, 191
or 150 BPM), and the entire set of 18 pieces w as repeated 192
three times, resulting in 54 trials. Each duet performed a 193
total of 66 trials, resulting in 1,320 trials o v erall. Sheet 194
music and sample audio were pro vided in adv ance. 195
3.1.4 Pr ocedur e 196
The recording procedure for each piece included: (1) An 197
e xperimenter’ s clap signaling start , follo wed by a measure 198
of clicks matching the gi v en tempo; (2) The flutist gi ving a 199
cue after clicks ended; (3) The duet be ginning in response 200
to the cue. 201
3.1.5 P ost-session Intervie w 202
The intervie ws collected the insights of the participants 203
on cue strate gies. P articipants reported pro viding cues ap- 204
proximately one beat (or half or tw o beats, depending on 205
the piece) ahead, using body mo v ements or breat h. Some 206
participants mentioned that in typical performance situa- 207
tions, the y adjust their cue timing based on the accompa- 208
n ying instrument and ensemble conte xt. 209
2 Some videos were recorded at 30 fps due to camera o v erheating
Figur e 3 : Gestur e Cue Example Gesture cue is defined
with the maximum(red) and minimum(black) peaks. The
interv al between these peaks, called ‘cue length’ ( x , red
line), and the duration between the maximum peak to the
flute onset, called ‘cue-onset length’ ( y , blue line).
3.2 Annotation and Pr epr ocessing 210
V ideo and audio data synchronization w as achie v ed 211
through an e xperimenter’ s clap at the s tart. Using the spec- 212
trogram vie wer in Adobe Audition, we m anually annotated 213
six mark ers per trial on the mel spectrogram (Figure 2): the 214
e xperimenter’ s clap ( Start ), breath sound onset and of fset 215
( Br eath Onset , Br eath Of fset ), initial note onsets of flute 216
and piano ( Flute Onset , Piano Onset ), and the flute’ s sec- 217
ond measure onset ( 2nd Measur e ). 218
4. METHODS 219
4.1 Gestur e Cue Detection 220
4.1.1 Motion Detection 221
T o detect gesture cues from flutists, we used MediaPipe’ s 222
f ace landmark detection 3 [19] to reliably identify f ace re- 223
gions, e v en when partially obscured by the flute. Subse- 224
quently , optical flo w methods [17, 18] were applied to track 225
f acial motion. A pilot study indicated optimal f ace land- 226
mark detection accurac y when the f ace occupied at least 227
50% of the video frame height; videos were accordingly 228
resized. 229
4.1.2 Motion F eatur e Extr action 230
W e analyzed f acial gestures by e xtracting position, v eloc- 231
ity , and acceler ation magnitude curv es from a v eraged y- 232
axis optical flo w v alues. Due to quantized pix el positions 233
causing discrete v elocity curv es, we applied zero-phase fil- 234
tering 4 to smooth the curv e while preserving peak posi- 235
tions. 236
4.1.3 Motion P eak Pic king 237
Figure 3 illustrates a typical gesture cue pattern. W ithin a 238
one-measure windo w preceding the flute onset (‘cue win- 239
do w’), we identified maximum and minimum peaks on 240
position, v elocity , and acceleration curv es using the find- 241
peaks 5 algorithm. This approach automatically det ected 242
3 a v ailable at: https://de v elopers.google.com/mediapipe
4 scip y .signal.filtfilt
5 scip y .signal.find_peaks
Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025
102
[Document text truncated for crawler view.]