scieee AI-readable full text Open interactive document viewer

Proceedings of the 26th International Society for Music Information Retrieval Conference

Nam, Juhan; Jeong, Dasaem; Choi, Keunwoo; Su, Li; Fuentes, Magdalena; Nakano, Tomoyasu; Hu, Xiao; Dong, Hao-Wen (Herman)

Abstract

Full proceedings of ISMIR 2025, held in Daejeon, South Korea, September 21-25, 2025. Contains 99 peer-reviewed papers.

Full text

International Society f or Music Inf ormation Retrie v al Conf erence ISMIR 2025 September 21–25, 2025 Daejeon, K orea and Online Pr oceedings ISMIR 2025 was or ganized by the K orea Adv anced Institute of Science and T echnology (KAIST), Sogang Uni versity , the K orean Society for Music Informatics, the International Society for Music Information Retriev al, and a di verse interna- tional committee of or ganizers. W ebsite: https://ismir2025.ismir.net Confer ence theme: Harmon y of T radition and Modernity Edited by: Juhan Nam (KAIST , South K or ea) Dasaem Jeong (Sogang Univer sity , South K or ea) K eunwoo Choi (Gaudio Lab / KAIST , South K or ea) Li Su (Academia Sinica, T aiwan) Magdalena Fuentes (NYU , USA) T omoyasu Nakano (AIST , J apan) Xiao Hu (University of Arizona, USA) Hao-W en (Herman) Dong (University of Michigan, USA) ISBN: 978-1-7327299-5-7 T itle: Proceedings of the 26th International Society for Music Information Retrie val Conference, Daejeon, K orea, Septem- ber 21–25, 2025. Permission to make digital or hard copies of all or part of this w ork for personal or classroom use is granted without fee, provided that copies are not made or distrib uted for profit or commercial advantage, and that copies bear this notice and the full citation on the first page. © 2025 International Society for Music Information Retrie val Sponsor s W e would like to e xpress our sincere gratitude to our generous sponsors whose support made ISMIR 2025 possible. Platinum Sponsor s Silver Sponsor s Br onz e Sponsor s WIMIR Sponsor s W idening Inclusion in Music Information Retrieval Patr on Contrib utor Supporter Local Suppor t W e gratefully acknowledge the local financial support of: Or ganizing Institutions The conference was supported by the Ministry of Education of the Republic of K orea and the National Research Founda- tion of K orea (NRF-2024S1A5C3A03046168). Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Or ganizing Committee General Chair s Juhan Nam, KAIST , K orea Dasaem Jeong, Sogang Uni versity , K orea K eunwoo Choi, Gaudio Lab / KAIST , K orea Scientific Pr ogram Chairs Li Su, Academia Sinica, T aiwan Xiao Hu, Uni versity of Arizona, USA Magdalena Fuentes, NYU, USA T omoyasu Nakano, AIST , Japan T utorial Chair s Hyung-Seok Choi, Ele venLabs, USA Christof W eiß, Univ ersity of Würzbur g, Germany Publication Chair Hao-W en (Herman) Dong, Univ ersity of Michigan, USA LBD Chair s K osetsu Tsukuda, AIST , Japan Y un-Ning (Amy) Hung, Moises AI, USA Music Chair Harin Lee, Max Planck Institute, Germany Gabriel Meseguer Brocal, Deezer , France Industry Chairs Jaehun Kim, Pandora / SiriusXM, USA Akira Maezaw a, Y amaha, Japan DEI Chair s K yung Myun Lee, KAIST , K orea T aegyun Kwon, KAIST , K orea iii Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Sponsor ship Chair Minz W on, Suno, USA Vir tual Chairs Jeong Choi, LINE Music, Japan Y u W ang, Spotify , USA Chin-Y un Y u, Queen Mary Uni versity of London, UK Grant Chair s Seokjin Lee, K yungpook National Univ ersity , K orea Alia Morsi, Uni versitat Pompeu F abra, Spain Ne wcomer Initiative Chair Seungheon Doh, KAIST , K orea W eb Chair / Designer Joonhyung Bae, KAIST , K orea Local Or ganization Chairs Jiyun Park, KAIST , K orea Eunjin Choi, KAIST , K orea Sein Lee, KAIST , K orea Jongsoo Kim, KAIST , K orea Hounsu Kim, KAIST , K orea Hayeon Bang, KAIST , K orea Danbinaerin Han, KAIST , K orea V olunteer Chair s Minsuk Choi, KAIST , K orea Kirak Kim, KAIST , K orea Social Media Chair Jongmin Jung, Neutune / Sogang Uni versity , K orea Unconference Chair Geof froy Peeters, Télécom Paris / IP-P aris, France i v Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 V olunteer s Eekgyun Ahn Luisa Lopes Carv alhaes Hyeyoon Cho Y erim Gim Robie Gonzales Dongyub Han Jaesun Joung Jaekwon Im Jaeran Choi Jaehyun On Jiwoo Ryu Jooeun Lim Joohye Son Kihong Kim Daeyong Kw on Y ujin Kim Meilin L yu Michaella Jung-Hyun Moon Zheng Xun Ng Beomjin Park Christos Plachouras Thiago Martin Poppe Pedro Ramoneda Baotong T ian Michael Xie Rui Y ang Jingwei Zhao T aehyeon Kim Minsoo Kang Hyunjae Kim Seokbeom Park Hyojin Kim T aein Song Mirinae Lee Hoyeol Sohn Minhee Lee Sunjae W on Y oonjeong P ark Gyubin Lee Carolina Carusi Junwon Lee K yung T aek Oh Sangeun Cho Hyerim Y un Hannah Park Sihun Lee Dongmin Kim Seonguk Ju Seola Cho Sojeong An Gary Jiwon Ri Minjun Kim Hyeonseok Choi Sungho Lee K yungsu Kim Saeyeon Hw ang Eunsik Shin T ae yeun Hwang Jaeyoung Shin Subeen Kim Y eeun Shin Y ideun (Eden) Park Minji Kim v Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Pr ogram Committee Meta-Re viewer s V inoo Alluri, IIIT - Hyderabad V ipul Arora, IIT Kanpur Claire Arthur , Geor gia Institute of T echnology Andreas Arzt, Apple Juan P . Bello, New Y ork Uni versity Emmanouil Benetos, Queen Mary Uni versity of London Rachel Bittner , Spotify Dmitry Bogdanov , Univ ersitat Pompeu Fabra Juan J. Bosch, Spotify Nicholas J. Bryan, Adobe Research John Ashley Bur goyne, Uni versity of Amsterdam Jor ge Calvo-Zaragoza, Uni versity of Alicante Carlos Eduardo Cancino-Chacón, Johannes K epler Univ ersity Linz Mark Cartwright, Ne w Jersey Institute of T echnology Kahyun Choi Nathaniel Condit-Schultz, Geor gia Institute of T echnology Simon Dixon, Queen Mary Uni versity of London Chris Donahue, CMU Hao-W en Dong, Univ ersity of Michigan Stephen Do wnie, organization Zhiyao Duan, Uni versity of Rochester Sebastian Ewert, Spotify Arthur Flex er , Johannes K epler Uni versity Linz Ichiro Fujinaga, McGill Uni versity Satoru Fukayama, National Institute of Adv anced Industrial Science and T echnology (AIST) Masataka Goto, National Institute of Adv anced Industrial Science and T echnology (AIST) Fabien Gouyon, P andora/SiriusXM Dorien Herremans, Singapore Uni versity of T echnology and Design Andre Holzapfel, KTH Royal Institute of T echnology in Stockholm Y u-Fen Huang, Academia Sinica Ozgur Izmirli, Connecticut College Blair Kaneshiro, Stanford Uni versity Jaehun Kim, Pandora / SiriusXM Katherine M. Kinnaird, Smith College and USAF A Katerina K osta, ByteDance Audrey Laplante, Uni v ersité de Montréal Stefan Lattner , Sony Computer Science Laboratories, P aris Alexander Lerch, Geor gia Institute of T echnology Florence Le ve, Uni versité de Picardie Jules V erne - Lab . MIS - Algomus Cynthia C. S. Liem, Delft Uni versity of T echnology Ethan Manilo w , Interacti ve Audio Lab, Northwestern Uni versity Brian McFee, Ne w Y ork Uni versity Cory McKay , Marianopolis Colle ge Andre w McPherson, Imperial College London Meinard Müller , International Audio Laboratories Erlangen Hema A. Murthy , IIT Madras Eita Nakamura, K yushu Univ ersity Oriol Nieto, Adobe Research Ser gio Oramas, Pandora Bryan Pardo, Northwestern Uni v ersity Johan Pauwels, Queen Mary Uni v ersity of London vii Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Preface Conference Theme: Harmon y of T radition and Modernity ISMIR 2025 embraced the theme "Harmony of T radition and Modernity ," encouraging di verse perspecti ves on ho w MIR can bridge past and present. W e welcomed research exploring the multifaceted intersections of tradition and inno va- tion—from the preserv ation and analysis of traditional music forms to the study of contemporary music trends powered by computational methods and data. The conference adv anced the understanding of music as a dynamic and e volving cultural force, engaging with both its rich historical roots and its e ver -expanding horizons. Confer ence Logo: The ISMIR 2025 logo draws inspiration from Ilwol-obongdo , the royal folding screen that traditionally stood behind the K orean throne, depicting fiv e peaks with the sun and moon symbolizing cosmic balance. The design integrates modern architectural elements from Daejeon—the Hanbit T o wer , Expo Bridge, and KAIST’ s signature blue color—representing technological adv ancement. Musical innov ation is embodied by the pedal-up symbol, a musical notation marking clear transition points. The color duality of blue (representing academic rigor) and red (representing creati ve ener gy) reflects the conference’ s theme of harmonizing tradition with modernity . The logo was designed by Joonhyung Bae, KAIST . Messages fr om the General Chairs It is our great pleasure to welcome you to the 26th Conference of the International Society for Music Information Retrie v al (ISMIR 2025). The ISMIR conference is the world’ s leading forum for research on processing, searching, or ganizing, and accessing music-related data. This year’ s edition takes place in Daejeon, K orea, from September 21 to 25, 2025, and is jointly or ganized by the K orea Advanced Institute of Science and T echnology (KAIST), Sogang Uni versity , and the K orean Society for Music Informatics. W e are delighted to present the ISMIR 2025 program. This year , we recei ved 324 abstracts, from which 278 papers were re viewed. Of these, 99 papers were accepted with an acceptance rate of 35.6% (35.84% in ISMIR 2024). As in pre vious years, the revie w process was conducted under a double-blind, two-tier model, in volving 269 re vie wers and 74 meta-re viewers (up from 256 re vie wers and 70 meta-re viewers in 2024). Each paper recei ved at least three re vie ws, including one from a meta-re viewer . W e are deeply grateful to all re vie wers and meta-revie wers for their time, expertise, and dedication. The accepted papers, authored by 413 authors (328 unique authors), were presented in both oral and poster sessions, with a mean of 4.17 authors per paper (median: 4; max: 10). Guided by this year’ s special theme, Harmony of T radition and Modernity , the program features two ke ynote talks: a legendary K-pop producer and a K orean traditional music expert. It also includes a special session introducing research on Asian traditional music from musicological and anthropological perspecti ves, an industry session sho wcasing cutting-edge music services from our sponsors, and a WIMIR session reporting ongoing ef forts to broaden di versity and inclusi veness in the MIR community . Ev ening e vents include a music program that demonstrates creati ve applications of MIR technologies in musical works, a concert of K orean traditional music, the e ver -popular jam session, and the RenCon challenge, which e valuates systems capable of rendering e xpressi ve musical performances from symbolic scores. In addition, we prepare K-Culture Night, a social e vent where participants can e xperience K orean traditional games, food, costumes, and music. Continuing ISMIR’ s well-established traditions, the program also of fers a full day of tutorials, a half day of late-breaking/demo and unconference sessions, and three satellite e vents before and after the main conference: the W orkshop on Human- Centric Music Information Research (HCMIR25), the International Conference on Digital Libraries for Musicology (DLfM), and the W orkshop on Large Language Models for Music & Audio (LLM4MA). W e would like to e xpress our sincere gratitude to our 22 ISMIR sponsors and 3 WIMIR sponsors, raising $110,000 in support. W e are particularly delighted to welcome 9 first-time ISMIR sponsors. Our sponsors represent a truly global partnership: 10 from Asia, 7 from America, and 5 from Europe. Specifically , we thank: Adobe, Moises, Udio, Algo- riddim, Google DeepMind, Spotify , Y amaha, Neutune, Steinber g, AlphaTheta, Suno, Uni versal Music Group, Cochl., AudibleMagic, BMA T , Deezer , Gaudio, MIPPIA, Roland, AudAI, Piascore, and Neutone. Their generous support makes ISMIR 2025 possible. W e also gratefully ackno wledge the local financial support of the K orea T ourism Or ganization and the Daejeon T ourism Organization. The conference was supported by the Ministry of Education of the Republic of K orea and the National Research F oundation of K orea (NRF-2024S1A5C3A03046168). Finally , we thank the or ganizing committee for their dedication, professionalism, and tireless work in bringing this conference to life. xv Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 W e are not only organizers b ut also people who truly lov e ISMIR and its community . What we hav e learned, experienced, and shared with so many colleagues at ISMIR has shaped us into the researchers we are today . It is a joy for us to return the passion, inspiration, and kindness we hav e recei ved from ISMIR ov er the years. Every ISMIR has al ways remained in our hearts as a joyful and inspiring time, and we ha ve prepared this year’ s e vent with the hope that it will be remembered in the same way for you. W e warmly in vite you to enjoy the v ery first ISMIR in K orea to the fullest! Conference Statistics ISMIR 2025 brought together 596 participants: 489 full on-site attendees, 32 single-day on-site attendees, and 75 online participants. W ith virtual attendees from 45 countries, supported by dedicated virtual volunteers and chairs, the confer- ence fostered a truly global community . Ov er 600 members joined our Slack workspace for real-time communications, and numerous vie wers followed the proceedings on Y ouT ube and engaged with the conference program through Mini- Conf (ismir2025program.ismir .net). Interacti ve tutorial sessions were held via Zoom, demonstrating the conference’ s far -reaching impact across continents. Scientific Pr ogram The scientific program recei ved 324 abstract submissions, from which 278 papers were revie wed. Of these, 99 papers were accepted with an acceptance rate of 35.6% (35.84% in ISMIR 2024). The re view process w as conducted under a double-blind, two-tier model, in volving 269 re vie wers and 74 meta-revie wers (up from 256 re viewers and 70 meta- re viewers in 2024). Each paper recei ved at least three re vie ws, including one from a meta-revie wer . The 99 accepted papers were authored by 413 authors (328 unique authors), with a mean of 4.17 authors per paper (median: 4; max: 10). Accepted papers were presented in both oral and poster sessions throughout the conference. Grants Pr ogram The conference supported equitable access through a comprehensi ve grants program, distrib uting 99 registration wai v ers (84 on-site, 15 virtual), 44 accommodation grants, and 5 tra vel grants. Grants were aw arded across four categories: Paper Authors, WIMIR (including Accessibility and Childcare), Music Authors, and LBD/Satellite Events, enabling participation from students, underrepresented groups, and researchers from lo w- or middle-income countries. Late-Breaking/Demo Session The Late-Breaking/Demo (LBD) session provided a platform for sho wcasing inno vati ve preliminary w ork in MIR. W ith a capacity of 75 posters accepted on a first-come-first-serv ed basis, the session of fered an accessible entry point for new- comers and early-career researchers to present prototypes, datasets, and initial concepts, fostering community engagement and feedback. Diver sity , Equity , and Inc lusion ISMIR 2025 prioritized inclusi ve participation through multiple initiati ves. The conference maintained a Code of Con- duct with clear reporting channels, provided accessibility accommodations including mobility and sensory support, and implemented a photo consent policy using yello w lan yards to indicate "do not photograph me" preferences. These efforts, coordinated across Registration, Local Or ganization, and V irtual teams, ensured a safe and welcoming en vironment for all attendees. xvi Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Ne w-to-ISMIR P aper Mentoring Pr ogram The Ne w-to-ISMIR Paper Mentoring Session w as held to enhance accessibility and encourage participation from a broader , more di verse community . T en mentees, ne w to ISMIR, recei ved pre-submission guidance on paper writing and research direction from senior researchers who v olunteered their time as past Meta-Revie wers. W e sincerely thank the following mentors for their v aluable service: Ale xander Lerch, Ajay Sriniv asamurthy , Brian McFee, Cheng-i W ang, Chris Donahue, Cory McKay , Ethan Manilow , Jordan B. L. Smith, LEVE Florence, T aegyun Kwon (Emer gency Mentor). Ne wcomer Squad The Ne wcomer Squad, coordinated by Seungheon Doh, matched 52 ne wcomers with 13 squad leaders, forming sub- clusters based on research interests. An ISMIR onboarding session was held to help ne wcomers na vigate the conference and connect with the community . W e sincerely thank the following squad leaders for their dedication: Ajay Srini v asamurthy , V incent Lostanlen, Anja V olk, Ser gio Oramas, J. Stephen Downie, P atricia Hu, Jin Ha Lee, Stefan Balke, Lele Liu, Bruno Di Giorgi, Jan Haji ˇ c jr ., Jordan B. L. Smith, Jingwei Zhao, and Gabriel Meseguer Brocal. Evening Events and Special Pr ograms The conference featured se veral e vening e v ents that complemented the scientific program. The Music Program demon- strated creati ve applications of MIR technologies in musical works. A K orean T raditional Music Concert showcased traditional K orean music heritage. The e ver -popular Jam Session brought together musicians from the community for informal performances. K-Culture Night of fered participants an immersiv e experience of K orean traditional games, food, costumes, and music. The Industry Session sho wcased cutting-edge music services from our sponsors, pro viding insights into real-world applications of MIR technologies. Juhan Nam, Dasaem Jeong, and K eunwoo Choi General Chairs of ISMIR 2025 xvii T able of Contents K e y n o t e S p e a k e r s .................................................. 1 S p e c i a l S e s s i o n s ................................................... 2 T u t o r i a l s ....................................................... 5 S a t e l l i t e E v e n t s ................................................... 6 Papers – Session 1 9 GlobalMood: A Cross-Cultural Benchmark for Music Emotion Recognition Harin Lee, Elif Celen, P eter Harrison, Manuel Anglada-T ort, P ol van Rijn, Minsu P ark, Mar c Sc hönwies- ner , Nori J acoby ................................................ 1 1 RISE: Music Rearrangement for Realtime Intensity Synchronization W ith Ex ercise Ale xander W ang, Chris Donahue, Dhruv J ain ................................ 2 0 Expanding the HAISP Dataset: AI’ s Impact on Songwriting Across T wo AI Song Contests Lidia Morris, Michele Ne wman, Xinya T ang, Renee Singh, Mar cel Vélez Vásquez, Rebecca Le g er , Jin Ha Lee ....................................................... 2 8 Quantifying Regularity in Music Structure Analysis Brian McF ee ................................................. 3 6 On the De-Duplication of the Lakh MIDI Dataset Eunjin Choi, Hyerin Kim, Jiwoo Ryu, J uhan Nam, Dasaem J eong ...................... 4 4 Conditional Dif fusion as Latent Constraints for Unconditional Symbolic Music Generation Models Matteo P ettenò, Alessandr o Ilic Mezza, Alberto Bernar dini ......................... 5 2 Radif Corpus; Symbolic Dataset for Non-Metric Iranian Classical Music Maziar Kanani, Seán O’Leary , J ames McDermott .............................. 6 0 Melodic and Metrical Elements of Expressi veness in Hindustani V ocal Music Y ash Bhake , Ankit Anand, Pr eeti Rao ..................................... 6 8 Coloring Music: Bridging Music and Color Palettes for Graphic Design T akayuki Nakatsuka, Masahir o Hamasaki, Masataka Goto ......................... 7 5 Exploring Network Adaptations for Minimum Latenc y Real-T ime Piano T ranscription P atricia Hu, Silvan P eter , J an Schlüter , Gerhar d W idmer .......................... 8 3 A Systematic Ev aluation of Real-T ime Audio Score Follo wing for Piano Performance Jiyun P ark, Carlos Eduar do Cancino-Chacón, Suhit Chiruthapudi, J uhan Nam .............. 9 1 Predicting Flutist Onset T iming in Duet Performance: A Multimodal Analysis of Gesture and Breath Cues J aeran Choi, T ae gyun Kwon, J uhan Nam ................................... 1 0 0 xix Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 AI-Generated Song Detection via L yrics T ranscripts Markus F r ohmann, Elena Epur e , Gabriel Mese guer Br ocal, Markus Schedl, Romain Hennequin . . . . . 107 Measuring Sensory Dissonance In Multi-T rack Music Recordings: A Case Study W ith W ind Quartets Simon Schwär , Stefan Balke , Meinar d Müller ................................ 1 1 7 Papers – Session 2 125 Reformulating Soft Dynamic T ime W arping: Insights Into T arget Artifacts and Prediction Quality J ohannes Zeitler , Meinar d Müller ...................................... 1 2 7 IT O-Master: Inference-T ime Optimization for Audio Ef fects Modeling of Music Mastering Processors J unghyun K oo, Mar co Martinez-Ramir ez, W ei-Hsiang Liao, Gior gio F abbr o, Michele Mancusi, Y uki Mit- sufuji ..................................................... 1 3 4 A Multidimensional Approach to Opera Analysis: Harmony , T empo, and Dramatic Interaction in W agner’ s Siegfried Act III P ascal Schmolenzky , Stephanie Klauk, Rainer Kleinertz, Christof W eiss, Meinar d Müller ......... 1 4 2 Exploring the Feasibility of LLMs for Automated Music Emotion Annotation Meng Y ang, J on McCormac k, Maria T er esa Llano, W anchao Su ....................... 1 5 0 An Ev aluation Strategy for Local K ey Estimation: Exploiting Cross-V ersion Consistency Y iwei Ding , Y annik V enohr , Christof W eiss .................................. 1 5 8 T uning Matters: Analyzing Musical T uning Bias in Neural V ocoders Hans-Ulrich Ber endes, Ben Maman, Meinar d Müller ............................ 1 6 6 Aligning T e xt-to-Music Evaluation W ith Human Preferences Y ic hen Huang, Zachary No vack, K oichi Saito, Jiatong Shi, Shinji W atanabe, Y uki Mitsufuji, John Thic k- stun, Chris Donahue ............................................. 1 7 4 In v estigating Music T rack Liking in the Halo of Album Co vers Ole g Lesota, Anna Hausber ger , Ivanna Pshenychna, Oleksandr Shvydanenko, Olha Y ehor ova, Markus Schedl ..................................................... 1 8 2 Phylo-Analysis of F olk T raditions: A Methodology for the Hierarchical Musical Similarity Analysis Hilda Romer o-V elo, Gilberto Bernar des, Susana Ladr a, José R. P aramá, F ernando Silva ......... 1 9 0 dPLP: A Dif ferentiable V ersion of Predominant Local Pulse Estimation Ching-Y u Chiu, Sebastian Strahl, Meinar d Müller .............................. 1 9 8 PeakNetFP: Peak-Based Neural Audio Fingerprinting Rob ust to Extreme T ime Stretching Guillem Cortès-Sebastià, Benjamin Martin, Emilio Molina, Xavier Serra, Romain Hennequin ....... 2 0 6 Generating Symbolic Music From Natural Language Prompts Using an LLM-Enhanced Dataset W eihan Xu, J ulian McA uley , T aylor Ber g-Kirkpatrick, Shlomo Dubno v , Hao-W en Dong .......... 2 1 5 A Surve y on V ision-to-Music Generation: Methods, Datasets, Evaluation, and Challenges Zhaokai W ang, Chenxi Bao, Le Zhuo, Jingrui Han, Y ang Y ue, Y ihong T ang, V ictor Shea-J ay Huang, Y ue Liao ...................................................... 2 2 3 Emer gent Musical Properties of a T ransformer Under Contrastiv e Self-Supervised Learning Y ue xuan K ONG, Gabriel Mese gues-Br ocal, V incent Lostanlen, Mathieu Lagr ange, Romain Hennequin . . 235 xx Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Papers – Session 3 245 Are Y ou Really Listening? Boosting Perceptual A wareness in Music-QA Benchmarks Y ongyi Zang, Sean O’Brien, T aylor Ber g-Kirkpatrick, J ulian McA ule y , Zac hary Novac k .......... 2 4 7 GD-Retrie ver: Controllable Generati ve T e xt-Music Retrie val W ith Diffusion Models J ulien Guinot, Elio Quinton, Györ gy F azekas ................................ 2 6 2 T ow ards Robust Automatic Music T ranscription By Measuring Cross-V ersion Consistency Y annik V enohr , Y iwei Ding, Christof W eiss .................................. 2 7 1 Beyond Genre: Diagnosing Bias in Music Embeddings Using Concept Activ ation V ectors Roman Gebhar dt, Arne K uhle , Eylül Bektur ................................. 2 7 9 LiLA C: A Lightweight Latent ControlNet for Musical Audio Generation T om Baker , J avier Nistal ........................................... 2 8 7 What Song No w? Personalized Rhythm Guitar Learning in W estern Popular Music Zakaria Hassein-Be y , Y ohann Abbou, Alexandr e d’Hooge , Mathieu Giraud, Gilles Guillemain, Aurélien J eanneau ................................................... 2 9 6 Uni versal Music Representations? Ev aluating Foundation Models on W orld Music Corpora Charilaos P apaioannou, Emmanouil Benetos, Alexandr os P otamianos ................... 3 0 3 A Theoretical Model of Musical Form Martin Rohrmeier ............................................... 3 1 2 T ow ards Human-in-the-Loop Onset Detection: A T ransfer Learning Approach for Maracatu António Pinto ................................................. 3 2 0 Instruct-MusicGen: Unlocking T ext-to-Music Editing for Music Language Models via Instruction T uning Y ixiao Zhang , Y ukara Ikemiya, W oosung Choi, Naoki Murata, Mar co Martínez-Ramír ez, Liwei Lin, Gus Xia, W ei-Hsiang Liao, Y uki Mitsufuji, Simon Dixon ............................. 3 2 8 T OMI: T ransforming and Organizing Music Ideas for Multi-T rack Compositions W ith Full-Song Structure Qi He, Ziyu W ang , Gus Xia .......................................... 3 3 7 Automatic Melody Reduction via Shortest Path Finding Ziyu W ang, Y uxuan W u, Rog er Dannenber g, Gus Xia ............................ 3 4 6 Expotion: Facial Expression and Motion Control for Multimodal Music Generation F athinah Izzati, Xinyue Li, Gus Xia ...................................... 3 5 4 When V oices Interlea ve: T iming De viations in Six Performances of T elemann’ s Fantasias for Solo Flute P atrice Thibaud, Mathieu Giraud, Y ann T e ytaut ............................... 3 6 3 Papers – Session 4 371 Audio Synthesizer In v ersion in Symmetric Parameter Spaces W ith Approximately Equi v ariant Flow Matching Ben Hayes, Charalampos Saitis, Györ gy F azekas .............................. 3 7 3 SLAP: Siamese Language-Audio Pretraining W ithout Ne gativ e Samples for Music Understanding J ulien Guinot, Alain Riou, Elio Quinton, Györ gy F azekas .......................... 3 8 2 PianoBind: A Multi-Modal Joint Embedding Model for Pop-Piano Music Hayeon Bang, Eunjin Choi, Seungheon Doh, J uhan Nam .............. ............ 3 9 1 Enhancing Neural Audio Fingerprint Rob ustness to Audio Degradation for Music Identification Recep Oguz Araz, Guillem Cortès-Sebastià, Emilio Molina, Joan Serr a, Xavier Serra, Y uhki Mitsufuji, Dmitry Bogdano v ............................................... 3 9 9 xxi Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Beyond Notation: A Digital Platform for T ranscribing and Analyzing Oral Melodic T raditions J onathan Myers, Dar d Neuman ........................................ 4 0 7 CMI-Bench: A Comprehensiv e Benchmark for Ev aluating Music Instruction Follo wing Y inghao MA, Siyou Li, J untao Y u, Emmanouil Benetos, Akira Maezawa .................. 4 1 6 Lose the Frames: Exact Metrics for More Responsible Music Structure Analysis Evaluations Qingyang Xi, Brian McF ee .......................................... 4 2 6 Unifying Continuous and Discrete Compressed Representations of Audio Mar co P asini, Stefan Lattner , Györ gy F azekas ................................ 4 3 3 Improving BER T for Symbolic Music Understanding Using T oken Denoising and Pianoroll Prediction J un-Y ou W ang, Li Su ............................................. 4 4 2 Scaling Self-Supervised Representation Learning for Symbolic Piano Performance Louis Bradshaw , Ale xander Spangher , Honglu F an, Stella Biderman, Simon Colton ............ 4 5 1 The Rhythm In An ything: Audio-Prompted Drums Generation W ith Masked Language Modeling P atrick O’Reilly , J ulia Barnett, Hugo Flor es Gar cia, Annie Chu, Nathan Pruyne, Pr em Seetharaman, Bryan P ar do .................................................. 4 6 0 Count the Notes: Histogram-Based Supervision for Automatic Music T ranscription J onathan Y affe , Ben Maman, Meinar d Müller , Amit Bermano ........................ 4 6 9 Joint T ranscription of Acoustic Guitar Strumming Directions and Chords Sebastian Mur gul, J ohannes Schimper , Michael Heizmann ......................... 4 7 7 Enabling Empirical Analysis of Piano Performance Rehearsal W ith the Rach3 MIDI Dataset Alia Morsi, Suhit Chiruthapudi, Silvan P eter , Ivan Pilkov , Laura Bishop, Akira Maezawa, Xavier Serr a, Carlos Eduar do Cancino-Chacón ...................................... 4 8 4 From Discord to Harmony: Consonance-Based Smoothing for Improv ed Audio Chord Estimation Andr ea P oltr onieri, Xavier Serra, Martín Rocamor a ............................. 4 9 2 Papers – Session 5 501 K eyboard T emperament Estimation From Symbolic Data: A Case Study on Bach’ s W ell-T empered Cla vier P eter V an Kr anenbur g, Gerben Bisschop ................................... 5 0 3 Refining Music Sample Identification W ith a Self-Supervised Graph Neural Network Aditya Bhattacharjee , Ivan Mer esman Higgs, Mark Sandler , Emmanouil Benetos ............. 5 1 1 V ideo-Guided T e xt-to-Music Generation Using Public Domain Movie Collections Haven Kim, Zachary No vack, W eihan Xu, J ulian McAule y , Hao-W en Dong ................. 5 1 8 PianoV AM: A Multimodal Piano Performance Dataset Y onghyun Kim, J unhyung P ark, Joonhyung Bae , Kirak Kim, T ae gyun Kwon, Alexander Ler c h, J uhan Nam 528 LoopGen: T raining-Free Loopable Music Generation Davide Marincione, Gior gio Strano, Donato Crisostomi, Roberto Rib uoli, Emanuele Rodolà ....... 5 3 6 Enhancing Music Recommender Systems W ith Multimedia Content: A Context-A w are Approach Ole g Lesota, V er onica Clavijo, Attia Rizwani, Markus Sc hedl, Bruce F erwer da ............... 5 4 7 CultureMER T : Continual Pre-T raining for Cross-Cultural Music Representation Learning Angelos-Nik olaos Kanatas, Charilaos P apaioannou, Ale xandr os P otamianos ................ 5 5 5 xxii Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 Adapti ve P ath of Prediction: An Unsupervised Method for Modeling Note-Lev el Informational Hierarchy of Polyphony Xiaoxuan W ang, Martin Rohrmeier ...................................... 5 6 5 V ersatile Music-for-Music Modeling via Function Alignment J unyan Jiang, Daniel Chin, Xuanjie Liu, Liwei Lin, Gus Xia ........................ 5 7 3 Understanding Performance Limitations in Automatic Drum T ranscription Philipp W e yers, Christian Uhle, Meinar d Müller , Matthias Lang ...................... 5 8 2 High-Resolution Sustain Pedal Depth Estimation From Piano Audio Across Room Acoustics Hanwen Zhang, K un F ang, Ziyu W ang , Ichir o Fujinaga ........................... 5 8 9 In v estigating an Overfitting and De generation Phenomenon in Self-Supervised Multi-Pitch Estimation F rank Cwitk owitz, Zhiyao Duan ....................................... 5 9 6 Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation J uan Carlos Martinez-Se villa, J oan Cerveto-Serrano, Noelia Luna-Bar ahona, Gr e g Chapman, Cr aig Sapp, David Rizo, J or ge Calvo-Zar agoza .................................. 6 0 4 Fx-Encoder++: Extracting Instrument-W ise Audio Ef fect Representations From Mixtures Y en-T ung Y eh, J unghyun K oo, Mar co Martínez-Ramír ez, W ei-Hsiang Liao, Y i-Hsuan Y ang, Y uki Mitsufuji 612 Papers – Session 6 621 MIDI-V ALLE: Improving Expressi ve Piano Performance Synthesis Through Neural Codec Language Mod- elling Jingjing T ang, Xin W ang , Zhe Zhang, J unichi Y ama gish, Geraint W iggins, Györ gy F azekas ........ 6 2 3 Playability Prediction in Digital Guitar Learning Using Interpretable Student and Song Representations Manuel Müllersc hön, Anssi Klapuri, Mar celo Rodriguez, Christian Car din ................. 6 3 1 Gregorian Melody , Modality , and Memory: Segmenting Chant W ith Bayesian Nonparametrics V ojt ˇ ech Lanz, J an Haji ˇ c jr . .......................................... 6 3 8 IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups Hitoshi Suda, J unya K oguc hi, Shunsuke Y oshida, T omohik o Nakamura, Satoru Fukayama, J un Ogata . . . 647 GO A T : A Lar ge Dataset of Paired Guitar Audio Recordings and T ablatures J ac kson Loth, P edr o Sarmento, Saurjya Sarkar , Zixun Guo, Mathieu Barthet, Mark Sandler ........ 6 5 5 ST A GE: Stemmed Accompaniment Generation Through Prefix-Based Conditioning Gior gio Str ano, Chiara Ballanti, Donato Crisostomi, Michele Mancusi, Luca Cosmo, Emanuele Rodolà . 663 Do Music Source Separation Models Preserve Spatial Information in Binaural Audio? Richa Namballa, Agnieszka Ro ginska, Magdalena Fuentes ......................... 6 7 1 Estimating Musical Surprisal From Audio in Autoregressi v e Diffusion Model Noise Spaces Mathias Rose Bjar e, Stefan Lattner , Gerhar d W idmer ............................ 6 7 9 Improving Neural Pitch Estimation W ith SWIPE K ernels David Marttila, J oshua D. Reiss ....................................... 6 8 8 Optical Music Recognition of Jazz Lead Sheets J uan Carlos Martinez-Sevilla, F rancesco F oscarin, P atricia Gar cia-Iasci, David Rizo, Jor ge Calvo-Zara goza, Gerhar d W idmer ............................................... 6 9 6 xxiii Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 T6: MIR for Health, Medicine, and W ell-being Pr esenters: Anja V olk, Elaine Chew , Michael A. Casey Abstract: This tutorial explores opportunities to emplo y MIR methods for music, health, medicine, and well-being. T op- ics include MIR for music therapy , music heart theranostics, and neurology and music information in epilepsy research, connecting MIR with interdisciplinary collaborations in healthcare. Satellite Events Satur da y , September 20, 2025 HCMIR25: 3rd W orkshop on Human-Centric Music Inf ormation Research Wher e: Room #3229, P aik Nam June Hall N25 Building, KAIST When: 14:00 – 18:00 Music and technology ha ve long intertwined, transforming how we create, share, and experience music. The 3rd edition of the HCMIR workshop e xplores MIR’ s ethical, societal, and human-centred dimensions. Ho w can we ensure MIR systems are inclusi ve, ethical, and aligned with human v alues? This workshop in vites researchers, practitioners, and artists to engage in a multidisciplinary discussion. K eynote: Understanding the Human Experience of Music to Shape Future T echnologies of MIR Prof. Kyung Myun Lee, KAIST F or further details: https://sites.google.com/view/hcmir25/home Frida y , September 26, 2025 DLfM 2025: 12th International Confer ence on Digital Libraries for Musicology Location: Gabriel Hall, Sogang Uni versity , Seoul, South K orea Satellite event of: ISMIR 2025 The International Conference on Digital Libraries for Musicology (DLfM) presents a venue for those w orking on, and with, digital library systems and content in the domain of music and musicology . DLfM welcomes contributions related to any aspect of digital libraries and musicology , including musical archi ving and retriev al, cataloguing, musical databases, music encodings, computational musicology , or the application of MIR to musicology . Pr ogramme Chair: Elsa De Luca, CESEM, Uni versidade No v a de Lisboa General Chair: David M. W eigl, mdw – Uni versity of Music and Performing Arts V ienna Local Chair: Dasaem Jeong, Sogang Uni versity F or further details: https://dlfm.web.ox.ac.uk/12th- international- conference- digital- libraries- musicology F or further details: https://dlfm.web.ox.ac.uk/ Frida y , September 26, 2025 LLM4MA: Large Language Models f or Music & A udio Location: Jung Geun Mo Conference Hall (5F), KAIST , Daejeon, K orea T ime: 8:30 am – 5:00 pm Online: https://zoom.us/j/99541677917 LLM4MA explores the rapidly e v olving intersection of large language models (LLMs) and music/audio understanding and generation. The workshop pro vides a forum for discussing adv ances in tokenization, long-context modeling, multimodal 6 Proceedings of the 26th ISMIR Conference, Daejeon, K orea, 2025 alignment, and controllability in music applications. It fosters early-stage research and community exchange on emer ging methods, challenges, and ethical considerations in AI-dri ven music creation. K eynote: Science of AI and AI for Science Prof. Noah A. Smith, Univ ersity of W ashington & Allen Institute for AI Organizing Committee: Chenghua Lin, SeungHeon Doh, Liumeng Xue, Ilaria Manco, Gus Xia, and others 7 P aper s – Session 1 GLOB ALMOOD: A CR OSS-CUL TURAL BENCHMARK FOR MUSIC EMO TION RECOGNITION Harin Lee 1 , 2 , 3 Elif Çelen 1 P eter Harrison 4 Manuel Anglada-T ort 5 P ol van Rijn 1 Minsu Park 6 Mar c Schönwiesner 3 Nori J acoby 1 , 7 1 MPI Empirical Aesthetics 2 MPI for Human Cogniti ve and Brain Sciences 3 Leipzig Uni versity 4 Uni v ersity of Cambridge 5 Goldsmiths, Univ ersity of London 6 Ne w Y ork Uni v ersity Abu Dhabi 7 Cornell Uni v ersity ABSTRA CT Human annotations of mood in music are essential for mu- sic generation and recommender systems. Howe v er , ex- isting datasets predominantly focus on W estern songs with terms deri ved from English, which may limit generalizabil- ity across di verse linguistic and cultural backgrounds. W e introduce ‘GlobalMood’, a nov el cross-cultural benchmark dataset comprising 1,180 songs sampled from 59 countries, with lar ge-scale annotations collected from 2,519 indi vid- uals across fi ve culturally and linguistically distinct loca- tions: U.S., France, Mexico, S. K orea, and Egypt. Rather than imposing predefined emotion and mood categories, we implement a bottom-up, participant-dri ven approach to or ganically elicit culturally specific music-related emotion terms. W e then recruit another pool of human participants to collect 988,925 ratings for these culture-specific de- scriptors. Our analysis confirms the presence of a v alence- arousal structure shared across cultures, yet also re veals significant di ver gences in how certain emotion terms (de- spite being dictionary equi valents) are percei v ed cross- culturally . State-of-the-art multimodal models benefit sub- stantially from fine-tuning on our cross-culturally balanced dataset, particularly in non-English contexts. Broadly , our findings inform the ongoing debate on the uni versality v er- sus cultural specificity of emotional descriptors, and our methodology can contrib ute to other multimodal and cross- lingual research. 1. INTR ODUCTION Music e vok es div erse emotional responses in listeners, spanning a wide spectrum beyond basic emotional cate- gories [1, 2]. A central challenge in Music Information Retrie val (MIR) is designing algorithms that can replicate this emotional sensiti vity . This is crucial for building rec- ommendation systems that align with listeners’ mood and © H. Lee, E. Çelen, P . Harrison, M. Anglada-T ort, P . v an Rijn, M. Park, M. Schönwiesner , and N. Jacoby . Licensed under a Cre- ati ve Commons Attribution 4.0 International License (CC BY 4.0). Attri- bution: H. Lee, E. Çelen, P . Harrison, M. Anglada-T ort, P . van Rijn, M. Park, M. Schönwiesner, and N. Jacoby , “GlobalMood: A cross-cultural benchmark for music emotion recognition”, in Pr oc. of the 26th Int. So- ciety for Music Information Retrieval Conf ., Daejeon, South Korea, 2025. context [3–5], and for generating music that resonates with indi vidual preferences [6]. More broadly , understanding ho w music con ve ys emotion is a core question in the sci- ence of music [7–9]. T o date, ho wev er , most algorithms ha ve been trained on datasets deri ved from W estern listen- ers and W estern music, using taxonomies primarily based on English language (e.g., MIREX [10]). A significant challenge is creating cross-cultural mod- els capable of handling non-W estern music and emotion v ocabularies be yond English. Addressing this challenge is essential to de veloping algorithms that accurately reflect global users’ preferences, including those whose musical tastes extend be yond the limited range of styles currently represented in training datasets. Moreov er , without cap- turing culturally specific nuances of emotion, especially those dif ficult to translate, key aspects of musical mean- ing may be missed entirely . Direct dictionary translations of English terms may be insuf ficient, as terms describing emotions are deeply cultural and may lack exact equi v a- lents [11–14]. T o address these issues, we introduce ‘GlobalMood’, 1 a ne w benchmark dataset designed to support culturally in- clusi ve and linguistically di verse emotion and mood recog- nition in music. Our contrib ution innov ates along three ke y dimensions: (i) the div ersity of musical stimuli, drawn from 59 countries; (ii) the div ersity of annotators, span- ning fi ve distinct re gions (with plans to extens to ov er 20 languages and locations in future); (iii) a data-dri ven ap- proach for collecting descriptors, generated org anically by participants in their o wn language during the annotation process. Data were collected through two stages in volving a to- tal of 2,519 participants and 1,180 songs balanced e venly across 59 countries: In the first stage (Section 4.1; Fig- ure 1), using a smaller subset of 200 songs, we employed our recently de veloped iterati ve task that combines open- ended elicitation with collectiv e refinement [13, 15, 16]. Rather than asking listeners to choose from a fix ed list of pre-defined emotion terms, we asked them to describe the percei ved emotion con veyed in the music using free-te xt tags in their nati ve language, and at the same time, rate the 1 All code and data: https://github.com/harin- git/ GlobalMood 11 Figur e 1 . Elicitation and refinement of music emotion terms through it erati v e participant chains. (A) Schematic illustration of the collaborati v e tagging process within a participant chain. P articipants contrib ute ne w emotion-related w ord tags for each song, rate the rele v ance of e xisting tags, and can also flag irrele v ant content, creating a dynamic refinement system. (B) T wenty most reliable emotion tags in each language, rank ed by their tag scores. Y -axis labels display tags in their original language (left) and English translations (right). tags pro vided by pre vious listeners. This approach w as k e y to unco v ering emotion terms that w ould otherwise be o v er - look ed by predefined, English-based taxonomies (such as ‘appeal/plead’ that appears in K orean only). In the second stage (Sect ion 4.2; Figure 2), we selected the top 20 elicited terms per language and cro wdsourced ratings for each tag across the entire set of 1,180 songs. This resulted in a total of 988,925 ratings, creating the most comprehensi v e open-source cross-cultural emotion anno- tation dataset in Music Emotion Recognition (MER) to date. W e le v eraged GlobalMood to test se v eral recent multi- modal and multilingual models (Gemini, CLAP) by e v al- uating their performance under zero-shot, fe w-shot, and fine-tuned scenarios (Section 4.3; Figure 3). Models trained only on English data performed poorly in some cultural conte xts, b ut fine-tuning with our cross-cultural data greatly impro v ed their performance in non-English settings. This highlights the critical importance of cross- cultural data in both training MER models and establishing appropriate benchmarks for their e v aluation. 2. RELA TED W ORKS 2.1 Music Emotion and Mood Annotation Datasets Se v eral datasets ha v e been de v eloped for MER systems with v arying annotation approaches. 2 Early e xamples in- 2 Note that databases often e xtend the concept of emotion to include related constructs such as mood or feeling . Here we adopt this broader clude the widely used MIREX 2007 mood dataset [10] with 240-250 W estern songs in fi v e mood clusters deri v ed from AllMusic’ s English tags (e.g., ‘passionate–rousing’, ‘wistful–bittersweet’), and CAL500 [17] with 500 W estern pop/rock songs annotated using 18 English mood terms by U.S. under graduate listeners. Ov er time, lar ger datasets ap- peared: the DEAM corpus (MediaEv al ‘Emotion in Music’ dataset [18]) containing 2,058 song e xcerpts with contin- uous v alence/arousal annotations; mood tags mined from lar ge corpora of Spotify music playlist s [19]; and the MTG-Jamendo dataset [20], which pro vides mood/theme tags for 18,486 songs. Notably , Jamendo’ s tags were freely cro wdsourced (56 unique mood labels), which introduced more label v ariety b ut still almost entirely in English. A common limitation a cross these datasets is their re- liance on predefined English descriptors, man y of which stem from W estern music psychology (for an e xception, see Strauss et al. [21]). F or instance, the Gene v a Emo- tional Music Scale (GEMS) defines 45 emotion descrip- tors (e.g., ‘jo yful acti v ation’) based on studies with Eu- ropean listeners [ 1 ] , and this taxonomy has been used to annotate datasets lik e Emotify [22]. Similarly , the mood cate gories in MIREX and CAL500 were fix ed in adv ance (dra wn from AllMusic or prior lit erature) and presented to annotators as a closed set of options. Consequently , these top-do wn approaches restrict annotators to the moods the researchers en visioned, lea ving an y unlisted mood nuances perspecti v e, while ackno wledging that subtle distinctions between them do e xist. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 12 Figur e 2 . Association between emotion terms across languages. (A) MDS visualization of the emotion ter ms based on mean ratings across the full song set. T erms positioned closer together e xhi bit similar rating patterns across songs, suggesting similar interpretations across languages. (B) Comparison of terms with direct translation equi v alents across languages. The area size indicates the de gree of semantic di v er gence despite apparent translation equi v alence. uncaptured and undocumented. Ackno wledging these limitations, recent research has be gun e xploring MER be yond the W estern-centric scope. Hu et al. [23] e xamined mood annotations of K-pop songs pro vided by both K orean and American listeners. Their approach in v olv ed translating the original MIREX mood cate gories into K orean for local annotators. Although this allo wed direct comparisons of mood classi fication between K orean and American listeners, it inherently restricted K o- rean annotations to terms originally defined within W estern conte xts. More recently , we compiled a balanced set of Ameri- can, Brazilian, and K orean songs and g athered mood an- notations across nine cate gories, where annotators rated songs both from their o wn and the other tw o countries [12]. W e sho wed that certain mood terms lik e ‘ener getic’ and ‘sad’ are highly consistent across cultures, while more ab- stract concepts lik e ‘lo v e’ and ‘dreamy’ di v er ge consider - ably . Simila r findings ha v e been reported by other stud- ies [13, 14], highlighting that when mood descriptors are imposed from one language onto another , important mean- ings can simply be ‘lost in translation’. In summary , whil e e x i sting MER datasets and research ha v e laid a solid groundw ork, the y remain limited by insuf- ficient linguistic and cultural di v ersity . Because man y are predominantly English-based and rely on top-do wn anno- tation strate gies, the y may o v erlook ho w people in other cultural conte xts percei v e emotion and mood in music. 2.2 A udio LLMs: the New Fr ontier in Music T agging Recent adv ances in multimodal lar ge language models (LLMs) ha v e opened promising a v enues for do wnstream MIR tasks, including emotion recognition. These mod- els combine the reasoning capabilities of LLMs with audio perception systems (audio LLMs), enabling more fle xible and nuanced music understanding than traditional classifi- cation approaches [24–26]. Models lik e MuLan [27] and MER T [24] ha v e demon- strated potential for zero-shot music emotion and mood classification by embeddi n g audio and natural language de- scriptions in a shared semantic space. Ho we v er , compre- hensi v e benchmarks such as the MuChoMusic [25] high- light a crucial limitation: these models rely hea vily on lan- guage modality and do not attend suf ficiently to audi o, of- ten f ailing with more nuanced audio e xamples for do wn- stream MIR tasks. This limitation could be particularly critical for non-W estern music and non-Engl ish emotion descriptors, gi v en that their training data are lar gely from W estern conte xts. Similarly , closed-source models (e.g., Gemini) ha v e sho wn promise in psychological te xtual analysis in multi- lingual conte xts [28], while e v aluations in specialized do- mains such as MIR remain scarce. The proprietary nature of their training data complicates thorough assessment of cross-lingual or cross-cultural performance. W e aim to ad- dress these fundamental g aps by pro viding a lar ge set of di- v erse, multilingual descriptors and annotations to support broader cross-cultural generalizability of audio LLMs. 3. METHOD 3.1 P articipants W e recrui ted tw o independent sets of participants across the tw o stages of our data collection: Stage 1 for emo- Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 13 tion term elicitation (N = 778; see Section 4.1) and Stage 2 for subsequent ratings on top 20 terms (N = 1,741; see Section 4.2). Participants had to be at least 18 years old, reside in the target country , and speak the target lan- guage as their primary language. Participants from the US were recruited through Prolific, while participants from the other four countries (France, Mexico, S. K orea, and Egypt) were recruited through the CINT platform. All participants provided informed consent under an appro ved protocol (see Section 7). P articipants were instructed to wear headphones and had to pass a headphone screening task [29], and a language proficiency test [30] before be- ing eligible for the main experimental task. Experiments were conducted in each participant’ s nati ve language (En- glish, French, Spanish, K orean, and Egyptian Arabic), with instructions translated using GPT -4o. Code to repli- cate the experiment through the PsyNet frame work [31] and all data are a vailable at https://github.com/ harin- git/GlobalMood 3.2 Globally Representativ e Song Selection T o create a globally representativ e music dataset, we used weekly Y ouT ube top 100 music charts (year 2017-2023) from 59 countries, spanning six continents. T o ensure each country’ s charts reflected its distinct popular music, we excluded an y track appearing in more than one country’ s chart. This left us with a country-e xclusive pool of songs. From this pool, we sampled 20 songs per country , yielding 1,180 songs in total. This di verse set is designed to capture a wide range of musical traditions and serve as a rob ust testbed for cross-cultural emotion recognition. Each 15- second audio excerpt w as trimmed from a random starting point in the full track, and normalized at -5dB loudness. 3.3 Model Evaluation W e used the resulting GlobalMood dataset to ev aluate se v- eral recent multimodal and multilingual models capable of music understanding. Specifically , we assessed Google’ s Gemini models ( 1.5 Flash , 2.0 Flash , and the latest 2.5 Pr o ), a family of multimodal lar ge language models capa- ble of processing and reasoning across te xt and audio (b ut also image and video). 3 W e compared zero-shot and fe w- shot approaches, where the latter included 10 human-rated emotion terms as examples. Gi ven that Gemini is closed-source, we also included CLAP (Contrasti ve Language-Audio Pretraining) [34] as an alternati ve, open-source model that learns joint audio- text embeddings. CLAP has demonstrated promise in MIR applications [35] and serves as the foundation for music- specific models like CLaMP [36]. Here, we conducted zero-shot e valuations through: (1) extracting audio em- beddings from CLAP , (2) computing cosine similarities with text embeddings of emotion terms, and (3) compar- ing these scores to human ratings. 3 Preliminary tests with other recent multimodal models sho wed per- formance issues—Flamingo 2 [32] struggled with rating consistency and GPT -4o [33] f ailed to generate musical descriptions or ratings from audio alone—thus we excluded them from further analysis. W e also fine-tuned CLAP on GlobalMood (train–test split = 1,000:180) to assess potential performance im- prov ements. T o preserve the continuous nature of our rat- ings, we represented each term in proportion to its mean rating (e.g., the term ‘calm’, with a mean rating of 3.0, appeared three times in the text). This method retained the nuanced information in our soft labels rather than re- ducing them to binary categories. T o improv e generaliz- ability , we created 10 augmented variations of each song through pitch shifting (range of ± 3 semitones), loudness adjustment (range of ± 15dB), and the addition of Gaus- sian noise (amplitude of 0.005). Each augmented v ariant randomly included one or two of these modifications. 4. RESUL TS 4.1 Bottom-up T erm Elicitation Across Languages 4.1.1 T agging pipeline Many e xisting studies on music emotions rely on pre- defined taxonomies or web-scraped data that of fer lim- ited linguistic div ersity [10, 19]. T o ov ercome this limita- tion, we employed a bottom-up, participant-dri ven tagging method [13, 15, 16]. Specifically , we asked participants in each country to complete independent ‘chains’ of iterati ve annotations. A subsample of 200 songs from the 1,180 entire set was used as stimuli. This subsample consisted of 180 balanced songs across countries, with an additional 20 local songs drawn from the participating country’ s pool. This was to ensure that local participants encounter enough music strongly tied to their background, allowing them to elicit culturally specific emotion descriptors. Figure 1A illustrates one such chain: (i) The first partic- ipant annotates the song using single-word emotion tags in their nati ve language; (ii) The second participant (from the same country) rates the rele vance of these tags (1–5 scale), flag irrele vant tags (e.g., genre- or lyrics-related rather than emotion), add new tags as necessary; (iii) The third par- ticipant sees all tags from earlier participants and repeats these steps; (iv) This iterati ve process continues through ten participants per chain, systematically refining and val- idating emotion terms. In each country , we ran the entire elicitation experiment twice and aggre gated the results to increase the di versity of responses from a lar ger pool of participants. 4.1.2 T op emer ging terms Follo wing the remo val of tags flagged by more than tw o participants in a chain, our STEP-T ag process yielded an extensi v e, culturally specific lexicon of emotion terms across languages (N unique terms: English = 644; French = 528; Spanish = 870; K orean = 629; Arabic = 283). T o identify the most salient terms in each language, we calcu- lated a composite score for e very term by multiplying its frequency of occurrences across chains by its mean rele- v ance rating. Higher scores indicate terms frequently men- tioned and consistently rated as highly rele vant. W e consolidated closely related morphological vari- ants (e.g., ‘happy’ and ‘happiness’ in English; gendered Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 14 forms such as ‘jo yeux’ and ‘jo yeuse’ in French) manu- ally with nati v e speak ers. Figure 1B presents the resulting 20 highest-ranking tags per language, displaying both the original w ord and English translations to f acilitate cross- cultural comparisons. Despite being e xplicitly ask ed to pro vide emotion terms, participants often g a v e broader af fecti v e descrip- tors lik e moods or feelings (e.g., ‘soft’ and ‘festi v e’). This aligns with prior MIR literature that oft en includes both emotion and mood, and gi v en its rele v ance for practical use, we did not enforce a strict distinction. 4.2 Lar ge-scale Di v erse Human Ratings 4.2.1 Cr oss-cultur al r atings acr oss the entir e set Ha ving identified the top 20 terms for each language, we ne xt g athered e xhausti v e ratings for the entire 1,180 songs of GlobalMood. W e recruited 1,741 ne w participants (see Section 3.1) who listened to the 15-second e xcerpts and rated ho w ef fecti v ely each e xcerpt con v e yed a gi v en emo- tion and mood term (1–5 scale). F or each stimulus, partic- ipants e v aluated se v en randomly selected terms from the rele v ant language set. This systematic approach ensured that, on a v erage, each song in each l anguage recei v ed 8.38 (SD = 2.40) unique participant ratings, resulting in an e x- tensi v e collection of 988,925 ratings spanning across fi v e languages. 4.2.2 Is ‘happy’ in my langua g e the same ‘happy’ in your langua g e? T o in v estig ate dif ferences in ho w each culture interprets these terms, we constructed 100 rating v ectors (5 lan- guages × 20 terms per language). Each v ector w as 1,180- dimensional, capturing the mean rating per term across the 1,180 song set. W e then performed non-metric mul- tidimensional scaling (MDS) using correlation as the dis- tance metric, projecting these v ectors in a tw o-dimensional space. In this emotion ‘space, ’ terms that position close to one another —e v en those from dif ferent languages—reflect similar rating patterns across the musical e xamples, sug- gesting comparable emotional interpretations across cul- tures. Figure 2A visualizes this emotion space. The terms cluster into tw o main re gions: one re gion of high arousal and high v alence (e.g., happy , ener g etic , and lively ; upper re gion of the figure) and a second re gion of lo w arousal that spans positi v e v alence (e.g., peaceful ; bottom left) to ne g- ati v e (e.g., sad ; bottom right). Notably , man y transl ated ‘equi v alents’ appear close together , which might suggest a general cross-cultural consensus on what music e v ok es what emotions. Ho we v er , e xamining s ix commonly shared terms that ha v e direct translations in at least four of the fi v e lan- guages ( fun , happy , rhythmic , lo ve , sad , and calm ) re- v ealed v arying de grees of cross-cultural agreement (see Figure 2B). F or each of these terms, between-country agreement ( r between ) w as computed as the a v erage of pair - wise corre lation coef ficients, while within-country agree- ment ( r within ) w as calculated using split-half reliability with Spearman-Bro wn formula. Ef fecti v ely , r within serv es as measurement error to compare as baselines when e v alu- ating r between . The term calm sho wed the highest a v erage agreement ( r between = 0.52 [0.49, 0.55]; r within = 0.49 [0.38, 0.57]), follo wed by fun ( r between = 0.46 [0.39, 0.53]; r within = 0.44 [0.27, 0.54]), lo ve ( r between = 0.44 [0.41, 0.47]; r within = 0.47 [0.33, 0.66]), sad ( r between = 0.41 [0.37, 0.45]; r within = 0.43 [0.32, 0.58]), rhythmic ( r between = 0.38 [0.30, 0.45]; r within = 0.43 [0.25, 0.53]), and notably happy ( r between = 0.37 [0.28, 0.45]; r within = 0.45 [0.24, 0.59]). Ov erall, considering within-country agreement (mean r within = 0.43–0.48), most of these terms were compara- ble in their between-country agreement (mean r between = 0.39–0.52). Ho we v er , despite being considered a basic uni v ersal human emotion [37], happy e xhibited a consid- erable g ap between between- and within-country agree- ment. This emphasizes the necessity of incorporating di- v erse cultural perspecti v es when modeling nuanced mu- sical emotional responses. Reliance on either dictionary translation or LLM-based translation alone could o v erlook important, conte xt-specific nuances in emotion and mood perception—particularly rele v ant when b uilding models for global audiences. Figur e 3 . Correlations between human ratings and multi- modal model predictions. (A) Gemini models with zero- shot prompting sho wing increase in performance with ne wer models. (B) CLAP models in zero-shot and fine- tuned scenarios sho wing ho w the use of multilingual an- notations can substantially increase performance. Gray dashed lines represent split-half reliability of human rat- ings using the Spearman-Bro wn formula as ba seline refer - ence of correlations achie v ed between humans. Error bars indicate 95% CI of mean correlation across songs. 4.3 Human vs. Multimodal Models Recent benchmarks ha v e e v aluated the capabilities of au- dio LLMs across v arious do wnstream MIR tasks, b ut these e v aluations ha v e also been restricted to English [25]. W e Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 15 P Y T H O N P R E - P R O C E S S I N G I S O L A T E D D R U M T R A C K S E C T I O N T Y P E U s e r M u s i c S o u r c e S e p a r a t i o n U s e r E x e r c i s e P l a n M u s i c S t r u c t u r e A n a l y s i s S E G M E N T S , B E A T T I M E S , A N D I N T R A - S E G M E N T T R A N S I T I O N P O I N T S R e a r r a n g e d M u s i c U N I T Y R E A L T I M E A D A P T A T I O N B e a t T r a c k i n g A u d i o L o u d n e s s M e t e r I n t r a - S e g m e n t C u t p o i n t s B E A T T I M E S T A M P S I N P U T O U T P U T H i g h - I n t e n s i t y S e g m e n t s Figur e 2 . System overview . RISE takes user music and ex ercise plans as input, preprocesses the music to identify intense segments and intra-se gment cutpoints, and sends this information to Unity for real-time adaptation. W e precompute a per-song loudness threshold τ , where any se gment with a loudness abov e τ is considered high in- tensity , and otherwise we consider it lo w intensity . Thresh- old τ is computed relati ve to the loudest section in the song: τ = max s n ∈ S L d ( s n ) − δ , where δ is a constant that determines the relati ve threshold. In our implementa- tion, we set δ = 5 decibels. If four or more consecutiv e segments are labeled high-intensity , we reassign the seg- ment with the lo west loudness as low-intensity and mer ge consecuti ve se gments with the same intensity labels. Ul- timately , our analysis induces a partition of the full track into ≤ N segments and associated binary intensity la- bels S ′ = { ( s ′ 1 , i 1 ) , ( s ′ 2 , i 2 ) , . . . } , where s ′ n are delineating timestamps and i i ∈ { Lo w , High } . 3.2 Prepr ocessing - Estimating Intra-Segment Cutpoints T o better align intense segments of music with high- intensity work out phases, we estimate a set of cutpoints to facilitate seamless adaptation. A cutpoint is a pair of timestamps in the music recording, consisting of a starting timestamp and a destination timestamp. The objectiv e is to estimate cutpoints that allo w smooth musical transitions, such that if playback jumps from the start of a cutpoint to its destination, users experience minimal disruption. W e deriv e an initial set of cutpoints using the approach proposed by Plachouras and Miron [25], which analyzes recurrence matrices encoding the self-similarity of musi- cal beats. Cutpoints are identified by detecting diagonals in these matrices that correspond to repeated patterns, pin- pointing transitions between musically coherent sections. This results in an initial set of candidate cutpoints: C = { ( c orig. i , c dest. i ) ∈ B × B } where C is the set of all estimated cutpoint pairs, and each cutpoint consists of a start time c orig. i and an end time c dest. i , both aligned to detected beat timestamps B . In prior work, cutpoints hav e been applied to rear- range music to fit external constraints, such as video du- ration [29]. These approaches allo w cutpoints to cross sec- tion boundaries, maximizing flexibility at the potential cost of playback naturalness. T o prioritize naturalness in our system, we enforce an intra-se gment constraint, ensuring that cutpoints only jump within a segment as opposed to across se gments. W e define the filtered set of intra-segment cutpoints as: C ′ : = { ( c orig. i , c dest. i ) ∈ C | ∃ s ′ n ∈ S ′ , s ′ n ≤ c orig. i , c dest. i < s ′ n +1 } This guarantees that e very cutpoint’ s start and destination timestamps fall within the same functional section s ′ n , pre- serving the structural integrity of the music. 3.3 Adaptation With Cutpoints In addition to music audio and ex ercise plan, our real- time adaptation system takes as input the follo wing in- formation estimated during pre-processing: (1) musical segments and corresponding intensities, (2) seamless cut- points, and (3) beat timestamps. T o adapt music in real time, we define a state machine that gov erns playback be- ha vior . The system operates in one of three possible states (Figure 3): • Loop State : If the system determines that a segment should be extended, it selects a cutpoint c i where c orig. i > t current and c dest. i < t current , looping pre vi- ously played sections to increase the duration of the current intensity segment. • Skip State : If a segment duration needs to be short- ened, the system selects a cutpoint c i where c orig. i > t current and c dest. i > c orig. i , skipping forward to reduce the segment duration. • Unmodified State : If no transition is required, the music plays continuously without alteration. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 22 Original Music Rearranged Output Original Music Rearranged Output Looping Skipping Transition to earlier location Transition to later location Segment extended through repetition Segment shortened through removal A B C A B C B A B C A C Figur e 3 . V isualization of adaptation modes. T op: Loop mode extends a se gment by jumping back. Bottom: Skip mode shortens a segment by jumping forw ard. Playback state transitions are determined dynamically based on the user’ s ex ercise plan, enabling the system to adjust segment duration in real time. 3.3.1 F ilter-based tr ansitions Sometimes, the system cannot immediately transition to the tar get intensity because there are no a v ailable cutpoints that matches the desired transition. In these cases, we use a filter -based transition , gradually removing high-frequenc y content before the transition and restoring it afterward, similar to DJ fade techniques. These transitions are no- ticeable and less seamless compared to cutpoint transi- tions. W e currently apply filter -based transitions only in unguided mode (section 3.4.1), when the ex ercise calls for a timely transition to a high-intensity music state. 3.4 Usage Modes W orkout habits may v ary greatly across different types of ex ercises—weight training can require minutes of rest to fully recov er and the actual timing can v ary greatly de- pending on ho w exhausted the user is. Guided interv al work outs, on the other hand, emphasize short bursts fol- lo wed by short rest periods that are strictly timed to max- imize time ef ficiency . Moti v ated by this observ ation, we designed two usage modes for tw o scenarios. An unguided User working when music is low intensity User working when music is high intensity User (started) resting when music is high intensity User resting when music is low intensity Unmodified Skipping Looping Skipping enter loop cycle break from loop cycle skip to high intensity play normally until user starts working 1 2 3 4 Figur e 4 . Adaptation mode for different scenarios. mode where the user is free to rest as long as the y need, and a guided mode where the user follo ws a predefined ex ercise plan. W e detail the system design of each mode belo w . 3.4.1 Unguided use The unguided mode allo ws users to freely choose start times and work/rest durations, b ut sometimes sacrifice adaptation quality by using filter -based transitions when no cutpoints are immediately a vailable. The system tak es real-time work out state as input and adjusts the music ac- cordingly (binary: work/rest). Currently , users manually indicate state changes by pressing a b utton. As depicted in Figure 4, RISE transitions between playback states de- pending on the current work out status: 1. Starting exer cise during a low-intensity segment : The system enters skip mode to quickly transition to a high-intensity segment. 2. During exer cise in a high-intensity segment : The system acti vates loop mode to sustain high-intensity music until the user begins resting. 3. Starting r est in a high-intensity segment : The sys- tem disables looping and switches to skip mode to exit high-intensity se gments. 4. During r est in a low-intensity segment : The sys- tem enters unmodified mode , allo wing the music to play naturally until the user resumes ex ercising. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 23 3.4.2 Guided use The guided mode provides seamless adaptations b ut re- quires a precise ex ercise plan. The system takes two du- rations (in seconds) for work and rest. W e iterate through a vailable cutpoints and select the transition that results in an adaptation closest to the desired duration. W e allow one transition per segment to maximize naturalness. The re- sults are close to the specified duration b ut rarely perfect ( e.g ., adjusting a 50-second segment to 32 seconds for a 30-second tar get). W e fade songs in at the beginning and out at the end, adjusting the intro to match the rest duration. 4. QU ANTIT A TIVE EV ALU A TION W e conducted an in-lab quantitativ e listening ev aluation to assess the seamlessness of transitions. 4.1 Pr ocedure W e collected 270 different songs, 30 songs each from nine dif ferent official Spotify w orkout playlists with dif ferent genre preferences. 13 songs were remov ed from the study because they had similar intensity throughout the entire song or had no a vailable intra-se gment cutpoints. For the remaining 257 songs, we generated two 10-second audio clips per song, one with a transition and one unmodified. Each clip is selected from a random segment of the song. W e randomized the transition timing to occur between tw o and eight seconds within the 10-second clip. This e valuation w as performed with six human listen- ers. Each clip was randomly assigned to two listeners who rated transition naturalness on a scale of 1-5 (1 = very jar - ring, 5 very seamless/unnoticeable). W e compare the rat- ings of modified and unmodified clips using a paired t-test. 4.2 Results W e found no statistically significant differences between the clips with transitions ( M = 4.3/5, SD = 1.1) and the baseline clips ( M = 4.5/5, SD = 1.0), t (256) = -1.63, p = .10, suggesting that our transitions are highly seamless and comparable to the unmodified clips. W e observ ed a small decrease in the a verage rating of transition clips (4.5 vs. 4.3) for two reasons. First, the beat detection algorithm is not perfect. W e found instances where the transitions were not perfectly aligned, causing a slight jump in rhythm. Sec- ond, familiarity with a song influenced the detection of transitions. Raters noted that, ev en when transitions were completely natural, they percei ved dif ferences in e xpected progression ( e.g ., altered lyrics) in songs they kne w well. W e also found that unmodified clips did not receiv e perfect scores. This was due to structural elements such as synco- pated rhythms and abrupt breaks, intended to surprise the listener , being perceiv ed as transitions by raters, despite these elements being part of the original compositions. 5. USER STUD Y W e conducted a user study to explore ho w users experience RISE in both guided and unguided ex ercise settings. 5.1 Study Design Participants e xercised to both unmodified music and our adapti ve system across two blocks: guided interval training and unguided weight training. W ithin each block, they e x- perienced both adapti ve and non-adapti ve conditions, with order fully counterbalanced. Each condition included a brief tutorial, 8 minutes of ex ercise, and optional rest. Participants selected tw o songs from a curated pool of 25 tracks, played identically across all conditions. The adap- ti ve system modified the music in response to user acti vity , while the non-adapti ve v ersion left the music unchanged. The full session lasted approximately 90 minutes. Guided interv al training. Participants performed in- terv al ex ercises of their choice ( e.g ., jumping jacks) ac- cording to a 40s work / 30s rest schedule. For the non- adapti ve system, these interv als are strict. For the adap- ti ve system, the actual timer may v ary by seconds depend- ing on the a vailable transitions in each section. Instruc- tions were sho wn on a screen with countdown visuals and sounds, modeled after popular work out timer videos [30]. Unguided weight training. Participants used dumb- bells to perform any freeform weight e xercises ( e .g., bi- cep curls). In adapti ve conditions, they v erbally indicated when they were about to be gin or end a work segment, allo wing the researcher to input music adaptation state changes. No interface was sho wn. Interview pr ocedure. After introducing the study and obtaining participant demographic information and con- sent, we ga ve them a verbal description of our system and recorded their first impressions through a short intervie w . After experiencing all conditions, we sho wed participants ho w the system operated by replaying the music for them along with visualizations of both segment intensity labels and cutpoint transitions. This was done at the end of the study to a void priming participants to focus on specific ma- nipulations and to e v aluate whether the adaptations were perceptible without guidance. The study concluded with a semi-structured exit intervie w for v erbal feedback. Participant and apparatus. W e recruited 12 partici- pants (5 female, 7 male, age M = 24.8 years; SD = 2.1) from a local uni versity . Participants re gularly ex ercise (3: 1-2x/week, 7: 3-4x/week, 2: 5-6x/week) for consider - able durations (1: 15-30 min, 4: 30-45 min, 1: 45-60 min, 5: 1-2 hours, 1: 2 hours+) in a v ariety of ex ercise types (8: steady cardio, 3: interv al training, 10: strength training, 4: other). The study took place in a controlled lab en viron- ment with music played through speakers for participant safety . Participants recei ved $50 for their participation. 5.2 Findings Intervie w transcripts were thematically analyzed using a coding reliability approach [31] with the Recal2 tool [32]. The analysis resulted in an a verage raw accurac y of 86.1% and an a verage Krippendorf f ’ s alpha of 0.72 between two raters, indicating acceptable agreement ( α > 0.66). W e present ke y findings below . Participants wer e excited by the premise of their music adapting to their work outs. After we e xplained Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 24 the study to participants b ut before they e xperienced our system, we asked participants about their initial reactions to the description of the study and our proposed system. Some participants had already made manual ef forts to align their work out with music (n = 8), such as coordinating specific music sections with their work out (n = 4). All participants expressed that music intensity alignment with work outs could be helpful to them. Participants belie ved alignment can increase moti vation, or help them feel more ener gized and “pumped” during the workout (n = 11). Estimated high-intensity segments aligned with their intuitions. W e showed participants a visualization of our system’ s estimated high-intensity music segments after the study , and all participants found the results to be aligned with their expectations (n = 12). Participants f ound our cutpoint adaptation tech- nique seamless, sub verting their expectations. Before the study , we shared a high-le vel description of our work- out adaptation system with participants. Many e xpressed initial skepticism that the modifications might detract from the naturalness of the music playback (n = 6). Out of the participants who had concerns reg arding the naturalness of music, all b ut one remarked that their concerns were ov erturned after experiencing our system (n = 5). While the filter -based modifications were regarded as more no- ticeable, most participants either did not notice or barely noticed any cutpoint-based modifications until the y were sho wn a visualization at the end of the study (n = 10). Participants appr eciated the extra motivation our alignment appr oach provided. All but one participant found the intensity alignment to enhance their workout e x- perience (n = 11), gi ving them more motiv ation to push harder in their work outs. P2 noted that when the music intensity peaked just as the y were about to reach failure in their work out, it helped them push through and get an extra one or tw o reps. Similarly , P4 recalled an instance where they e xpected the music to shift to a lo wer-intensity segment in the middle of their w ork interval, but instead the upbeat segment repeated, pro viding them the energy to push through the remainder of their work out. Participants v aried in their sensitivity to both music modifications and alignment with exer cise. For instance, P8 barely noticed music changes b ut felt their movements were more consistent in the adapti ve condition, while the nonadapti ve one felt chaotic. In contrast, P3, who regu- larly coordinates music with work outs, immediately no- ticed misalignments in the nonadapti ve condition and e ven felt the ur ge to stop the music when high-energy se gments played during their rest. Sensiti vity could also depend on ex ercise. P6 mentioned that they treated music as back- ground during simple ex ercises but v alued music align- ment for moti vation during more challenging e xercises, such as the bench press. This indicates the need to better understand ho w different modes of music engagement in- fluence tolerance and demand for system-dri ven changes. Participants pr eferred adapti ve music but identified issues with the unguided experience. Out of 12 partici- pants, 11 preferred our adapti ve system o ver nonadapti ve music: 6 preferred it in both scenarios, while 5 fav ored it only in the guided setting. Those who preferred the system only in the guided setting found two main issues with the unguided experience. First, the input method—requiring them to notify the researcher before ex ertion—was dis- tracting and impractical (n = 7). Second, while cutpoint transitions were deemed seamless and often unnoticeable, filter transitions disrupted the natural flo w of the music. These filter transitions were especially common during ex- ercises with longer work durations and shorter rest peri- ods, where excessi v e looping also occurred, leading to un- natural adaptations of the music. As a result, participants recommended improv ements in system customizability to support dif ferent workout habits ( e .g., long work, short rest) and music adaptation preferences (n = 8), manipu- lation seamlessness (n = 4), and expanding the a v ailable music pool or allo wing user input (n = 3). 6. LIMIT A TION AND FUTURE WORK Our study re vealed se veral limitations, including users finding manual input impractical and filter transitions dis- rupting. W e propose future w ork below: Fully automating the system with sensing technolo- gies. T o eliminate the need for manual interaction during work outs, we propose automating the system using real- time sensing technologies. By integrating acti vity recogni- tion, the system could automatically detect ex ercise phases without manual input. The system could also model user beha vior and predict upcoming exercise states based on past patterns to prepare adaptation plans in adv ance. Incorporating other types of modifications. While cutpoints enable seamless transitions, they may be una v ail- able when timely adaptation is required. Future work could complement our segment-sensiti v e approach with existing techniques, such as pace adjustment and song recommen- dation, to compensate for minor time discrepancies. Addi- tionally , exploring audio inpainting with generati ve mod- els could enable smooth transitions between arbitrary seg- ments, providing greater fle xibility and precision. Determining high-intensity segments. RISE currently aligns drum-prominent chorus and instrumental sections with user ex ertion. While this approach worked well in out study , future work can further explore ho w se gment- le vel musical v ariations influence ex ercise and ho w these ef fects may differ by genre or listener preference. 7. CONCLUSION W e present RISE, a nov el system that adapts music to align with ex ercise phases. A user study in v olving 12 partici- pants re vealed that, despite initial skepticism, most users appreciated the alignment and preferred it ov er nonadap- ti ve music for their work outs. Our work represents a step to wards expanding the design space of adapti v e music, making tailored music experiences, once limited to video games and precomposed soundtracks, applicable to real- world scenarios lik e workouts. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 25 8. REFERENCES [1] P . C. T erry , C. I. Karageorghis, M. L. Curran, O. V . Martin, and R. L. Parsons-Smith, “Ef fects of music in ex ercise and sport: A meta-analytic re view . ” Psycho- logical b ulletin , vol. 146, no. 2, p. 91, 2020. [2] S. Mitrof f, “Hitting the pa vement with spotify running (hands-on) - cnet, ” https: //www .cnet.com/tech/services- and- software/ hitting- the- pav ement- with- spotify- running- hands- on/, 2015. [3] Apple, “ Apple fitness+ - apple, ” https://www .apple. com/apple- fitness- plus/, 2024. [4] P . Knees, M. Schedl, and M. Goto, “Intelligent user interfaces for music disco very: The past 20 years and what’ s to come. ” in Pr oceedings of the 20th Interna- tional Society for Music Information Retrieval Confer - ence, ISMIR 2019, Delft, The Netherlands, November 4-8, 2019 , 2019, pp. 44–53. [5] M. Goto, “Grand challenges in music information re- search, ” in Dagstuhl F ollow-Ups: Multimodal Music Pr ocessing , M. Müller , M. Goto, and M. Schedl, Eds. Dagstuhl Publishing, 2012, pp. 217–225. [6] M. Schedl and A. Flex er , “Putting the user in the center of music information retrie val. ” in Pr oceedings of the 13th International Society for Music Information Re- trieval Confer ence, ISMIR 2012, Mosteir o S.Bento Da V itória, P orto, P ortugal, October 8-12, 2012 , 2012, pp. 385–390. [7] M. Kari, T . Grosse-Puppendahl, A. Jagaciak, D. Bethge, R. Schütte, and C. Holz, “Sound- sride: Af fordance-synchronized music mixing for in-car audio augmented reality , ” in The 34th Annual A CM Symposium on User Interface Softwar e and T echnolo gy , 2021, pp. 118–133. [8] L. Baltrunas, M. Kaminskas, B. Ludwig, O. Moling, F . Ricci, A. A ydin, K.-H. Lüke, and R. Schwaiger , “In- carmusic: Context-a ware music recommendations in a car , ” in E-Commer ce and W eb T ec hnologies: 12th International Confer ence, EC-W eb 2011, T oulouse, F rance , A ugust 30-September 1, 2011. Pr oceedings 12 . Springer , 2011, pp. 89–100. [9] A. W ang, Y . F . Cheng, and D. Lindlbauer , “Maringba: Music-adapti ve ringtones for blended audio notifica- tion deli very , ” in Pr oceedings of the CHI Confer ence on Human F actor s in Computing Systems, CHI 2024, Honolulu, HI, USA, May 11-16, 2024 . New Y ork, NY , USA: Association for Computing Machinery , 2024. [10] A. W ang, D. Lindlbauer , and C. Donahue, “T o wards music-aw are virtual assistants, ” in Pr oceedings of the 37th Annual A CM Symposium on User Interface Soft- war e and T ec hnology , UIST 2024, Pittsbur gh, P A, USA, October 13-16, 2024 . Ne w Y ork, NY , USA: Associa- tion for Computing Machinery , 2024. [11] J. Shriram, M. T apaswi, and V . Alluri, “Sonus tex ere! automated dense soundtrack construction for books us- ing movie adaptations, ” in Pr oceedings of the 23r d International Society for Music Information Retrieval Confer ence, ISMIR 2022, Bengaluru, India, December 4-8, 2022 , 2022, pp. 535–542. [12] S. Rubin, F . Berthouzoz, G. Mysore, W . Li, and M. Agraw ala, “Underscore: musical underlays for au- dio stories, ” in Pr oceedings of the 25th Annual A CM Symposium on User Interface Softwar e and T ec hnol- ogy , ser . UIST ’12. Ne w Y ork, NY , USA: Association for Computing Machinery , 2012, p. 359–366. [13] N. Masahiro, H. T akaesu, H. Demachi, M. Oono, and H. Saito, “De velopment of an automatic music selec- tion system based on runner’ s step frequency , ” in IS- MIR 2008, 9th International Confer ence on Music In- formation Retrieval, Dr exel Univer sity , Philadelphia, P A, USA, September 14-18, 2008 , 2008, pp. 193–198. [14] G. T . Elliott and B. T omlinson, “Personalsoundtrack: context-a ware playlists that adapt to user pace, ” in CHI’06 e xtended abstracts on Human factors in com- puting systems , 2006, pp. 736–741. [15] B. Moens, L. v an Noorden, and M. Leman, “D-jogger: Syncing music with walking, ” in 7th Sound and music computing Confer ence . Uni versidad Pompeu F abra, 2010, pp. 451–456. [16] J. Hockman, M. M. W anderley , and I. Fujinaga, “Real- time phase v ocoder manipulation by runner’ s pace. ” in NIME , 2009, pp. 90–93. [17] N. Oli ver and L. Kre ger-Stickles, “Papa: Physiology and purpose-aw are automatic playlist generation. ” in ISMIR 2006, 7th International Confer ence on Music Information Retrieval , 2006, pp. 250–253. [18] B. v an der Vlist, C. Bartneck, and S. Mäueler , “mobeat: Using interacti ve music to guide and moti v ate users during aerobic ex ercising, ” Applied psychophysiology and biofeedbac k , vol. 36, pp. 135–145, 2011. [19] Y . Chen, C.-C. Chen, L.-C. T ang, and W .-H. Chieng, “Enhancing running ex ercise with iot, blockchain, and heart rate adapti ve running music, ” IEEE Access , 2024. [20] C. I. Karageorghis and D. Holland, “Music in the e xer - cise domain: A revie w and synthesis (part ii), ” Interna- tional Revie w of Sport and Exer cise Psycholo gy , v ol. 5, no. 1, pp. 67–84, 2012. [21] A. T urrell, A. R. Halpern, and A.-H. Jav adi, “When tension is exciting: an electroencephalogram e xplo- ration of excitement in music, ” bioRxiv , p. 637983, 2019. [22] D.-L. Priest and C. I. Karageorghis, “ A qualitativ e in- vestig ation into the characteristics and ef fects of music accompanying e xercise, ” Eur opean physical education r e view , v ol. 14, no. 3, pp. 347–366, 2008. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 26 [23] T . Kim and J. Nam, “ All-in-one metrical and func- tional structure analysis with neighborhood attentions on demixed audio, ” in IEEE W orkshop on Applications of Signal Pr ocessing to A udio and Acoustics (W AS- P AA) , 2023. [24] R. Hennequin, A. Khlif, F . V oituret, and M. Moussal- lam, “Spleeter: a fast and ef ficient music source sepa- ration tool with pre-trained models, ” Journal of Open Sour ce Softwar e , v ol. 5, no. 50, p. 2154, 2020. [25] C. Plachouras and M. Miron, “Music rearrangement using hierarchical segmentation, ” in ICASSP 2023- 2023 IEEE International Confer ence on Acoustics, Speech and Signal Pr ocessing (ICASSP) . IEEE, 2023, pp. 1–5. [26] C.-W . Li and C.-G. Tsai, “The presence of drum and bass modulates responses in the auditory dorsal path- way and mirror -related regions to pop songs, ” Neur o- science , v ol. 562, pp. 24–32, 2024. [27] G. Madison, “Experiencing groove induced by music: consistency and phenomenology , ” Music per ception , v ol. 24, no. 2, pp. 201–208, 2006. [28] C. J. Steinmetz and J. D. Reiss, “pyloudnorm: A simple yet flexible loudness meter in p ython, ” in 150th AES Con vention , 2021. [29] Adobe, “Remix in premiere pro, ” https: //helpx.adobe.com/premiere- pro/using/ remix- audio- in- premiere- pro.html, 2024. [30] W . M. W . T imer , “Interval timer with music | 40 sec rounds 30 sec rest | mix 107, ” Y ouT ube video, 2021, accessed: September 1, 2024. [Online]. A v ailable: https://www .youtube.com/watch?v=lnBOQnc_p- E [31] R. E. Boyatzis, T ransforming qualitative information: Thematic analysis and code development . Sage, 1998. [32] D. Freelon, “Recal2: Reliability for 2 coders, ” http:// dfreelon.or g/utils/recalfront/recal2/, 2010. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 27 EXP ANDING THE HAISP D A T ASET : AI’S IMP A CT ON SONGWRITING A CR OSS TWO AI SONG CONTESTS Lidia Morris 1 Michele Newman 1 Xinya T ang 1 Renee Singh 1 Mar cel Vélez Vásquez 2 Rebecca Leger 3 Jin Ha Lee 1 1 Information School, Uni versity of W ashington 2 Uni versity of Amsterdam 3 Fraunhofer Institute for Inte grated Circuits [email protected] , [email protected] ABSTRA CT As artificial intelligence (AI) continues to shape creati ve practices, understanding its role in human-AI songwriting remains crucial. This paper expands the Human-AI Song- writing Processes (HAISP) dataset by incorporating data from the 2024 AI Song Contest, building upon the original 2023 dataset. By analyzing ne w submissions, we provide further insights into AI’ s e v olving impact on songwriting workflo ws, creati ve decision-making, and control. A com- parati ve study of AI tool usage and participant strate gies between the 2023 and 2024 contests re veals shifts in col- laboration patterns and tool ef fectiv eness. Additionally , we assess the dif ferences between general-purpose AI systems and personalized, fine-tuned tools, highlighting their im- pact on creati ve agenc y . Our findings offer k ey design im- plications for AI-assisted songwriting tools, providing ac- tionable insights for AI de velopers and music practitioners seeking to enhance co-creati ve e xperiences. 1. INTR ODUCTION Artificial Intelligence (AI) has rapidly become an integral component of creati ve fields, reshaping artistic expression across v arious domains. From visual arts to literature, AI- po wered tools are being le veraged to augment human cre- ati vity , raising ne w questions about authorship, originality , and the e volving nature of co-creation [1–3]. No where is this transformation more e vident than in music composi- tion, where AI systems are increasingly employed to gen- erate melodies, harmonies, lyrics, and entire song struc- tures [4, 5]. These adv ancements ha ve giv en rise to ne w forms of collaboration between human musicians and AI, necessitating a deeper understanding of the dynamics of human-AI co-creation in songwriting. The study of human-AI collaboration in music is par- ticularly important due to the complex, often subjectiv e © L. Morris, M. Newman, X. T ang, R. Singh, M.A. Vélez Vásquez, R. Leger , and J.H Lee. Licensed under a Creati ve Commons Attribution 4.0 International License (CC BY 4.0). Attribution: L. Mor - ris, M. Ne wman, X. T ang, R. Singh, M.A. Vélez Vásquez, R. Leger , and J.H Lee, “Expanding the HAISP Dataset: AI’ s Impact on Songwriting Across T wo AI Song Contests”, in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf ., Daejeon, South K orea, 2025. nature of the creati ve process. While AI can accelerate composition workflo ws and generate nov el musical ideas, its role in enhancing versus replacing human creati vity re- mains a critical area of in vestig ation [6, 7]. Furthermore, questions reg arding our changing definition of computa- tional creati vity and the role AI can play in the creativ e process as a tool or collaborator continue to loom large ov er the field [8, 9]. Addressing these concerns requires qualitati ve data that captures not just empirical informa- tion on the use of generati ve AI, b ut the liv ed experiences of creators working with AI in music production. T o contribute to this gro wing field of study , the Hu- man–AI Songwriting Processes (HAISP) dataset was in- troduced in 2024 as a curated resource designed to ex- plore the interaction between human musicians and AI sys- tems [10]. The dataset was deri ved from submissions to the AI Song Contest 2023, an annual competition that in- vites teams of musicians, data scientists, and researchers to explore the creati v e potential of AI in songwriting [11]. It comprises 34 coded entries documenting ho w teams used AI tools in their songwriting processes. It provides a structured frame work for analyzing v arious aspects of AI- assisted music creation, including: • The specific AI tools and models used in composi- tion • The songwriting methodologies employed by human-AI teams • Reflections on ethical considerations and challenges related to AI in music • T eams’ assessments of their collaborati ve e xperi- ence with AI The findings from the HAISP dataset highlighted the di- verse w ays in which AI is integrated into songwriting, with teams using AI for tasks ranging from melody generation to performance synthesis. Howe v er , the dataset also under- scored the limitations of AI tools, such as lack of creativ e control, technical limitations, and concerns about origi- nality . In addition, ethical concerns regarding the pro ve- nance of data and the transparency of AI-generated content emer ged as key themes. 28 Building on this foundation, the current study extends the HAISP dataset by incorporating ne w data from the 2024 AI Song Contest, of fering a longitudinal perspecti ve on the e volution of human-AI collaboration in songwriting. By comparing data from 2023 and 2024, this e xpanded dataset enables a deeper analysis of trends, emerging tech- nologies, and shifting attitudes to ward AI in creati ve work. Through this research, our goal is to provide v aluable in- formation for musicians, AI de velopers, creativity schol- ars, and beyond. 2. B A CKGR OUND The intersection of AI and music composition represents a rapidly e volving field that has a long history to e xplore. AI-assisted music creation has progressed from early algo- rithmic experiments and academic electronic music cen- ters [12] to sophisticated machine learning models capa- ble of composing complete musical pieces by the broader public [13]. As these technologies become more accessi- ble, they not only influence the w ay music is made but also raise critical ethical, cultural, and artistic questions about human-AI co-creation. 2.1 Evolution of AI in Music Composition The application of computational techniques in music composition can be traced back to the mid-twentieth cen- tury , when early pioneers e xperimented with algorithmic approaches to sound generation [13], such as the work completed at the Columbia-Princeton Electronic Music Center , which laid the groundwork for a v ariety of com- posers careers and technological innov ations [12, 14]. By the early 2000s, dev elopments in machine learning facili- tated the creation of models that could autonomously gen- erate melodies, harmonies, and song structures [15, 16]. The increasing sophistication of deep learning and gen- erati ve AI in the past decade has further transformed the landscape of music composition. Notable adv ances in- clude OpenAI’ s MuseNet, Google’ s Magenta, and Meta’ s MusicGen, all of which employ transformer -based archi- tectures to produce di verse compositions [17–19]. These tools enable musicians to collaborate with AI in v arious ways, from generating musical ideas to assisting with ar- rangement, and more [19]. The growing accessibility of these technologies has been sho wcased in platforms such as the AI Song Contest (AISC) [11]. 2.2 AI Songwriting T ools and Methods V arious AI-powered tools ha ve emer ged to facilitate human-AI collaboration in songwriting. OpenAI’ s MuseNet [20] is a deep neural network capable of com- posing multi-instrumental pieces across multiple genres, while Google’ s Magenta project provides open-source ap- plications for AI-assisted melody generation, chord pro- gression, and rhythm creation [18]. While these tools of fer ne w creativ e possibilities, they also introduce challenges related to artistic control, originality , and the implications of AI as a co-creati ve entity , especially when the user is an inexperienced music creator [2, 11, 21]. AI-generated music is often constrained by its training data, leading to concerns about predictability , stylistic homogenization, and the potential for AI to reinforce e xisting musical con- ventions rather than foster true innov ation [22, 23]. Fur- thermore, the extent to which AI-generated compositions can be considered “creati ve” in the same sense as human- authored works is still being questioned, especially when it comes to just ho w much of a role the AI plays in the compositional process [24, 25]. 2.3 Creativity Studies The increasing adoption of AI in creati ve domains has sparked debates about the nature of creati vity and the role of machines in artistic e xpression [11, 26–28]. T raditional vie ws of creativity emphasize human intuition, cultural context, and emotional depth [29–31] - qualities that AI, as a statistical modeling system, does not inherently pos- sess. Can AI truly be considered a creativ e agent, or is it merely an adv anced tool for pattern recognition and recom- bination? Ethical concerns also e xtend to the implications of AI’ s increasing role in the creati ve workforce [32, 33]. As AI-generated compositions become more sophisticated, there is a potential for automation to displace human mu- sicians in certain commercial contexts [34]. In response to these challenges, scholars and industry professionals hav e called for greater transparency in AI training data, ethical guidelines for AI-assisted composition, and policies to en- sure that human artists remain central to the creati ve pro- cess [35]. 3. D A T ASET EXP ANSIONS: METHODOLOGY The HAISP Dataset is accessible as a .csv and .xlsx on the Open Science Frame work (OSF) under a Creati ve Com- mons Attrib ution-NonCommercial 4.0 International (CC BY -NC) license, which allo ws for broad access and uti- lization for research purposes [36]. W e generated the dataset via consensus coding [37]. One researcher coded a selection of the data entries, col- lecting them into the dataset. A second coder then re- vie wed the initial coding, validating the coding by ei- ther marking agreement or disagreement with the cod- ing choices within a comment on the code in the dataset, adding what they felt the coder w as missing within their codes from the data. In the case of disagreement, a third researcher helped decide on the final code as a tie-breaker . 3.1 Data Collection Similarly to the 2023 practice, each team had to fill in the AI Song Contest 2024 Submission F orm via Google Forms to participate in the contest [10]. The form consists of entry fields that cov er all the basic information about teams and songs: • team (bio for the website, location, lev el of exper - tise, moti vation to participate, how the y heard about the AISC); Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 29 • song (title, length, link to music video/soundcloud/ blogpost, concept/idea, lyrics, li ve performance). Each team was additionally ask ed to generate a process document and sa ve it as a PDF file, which mainly includes more detailed moti vation, songwriting workflo w and col- laboration process, e valuation of co-creation, and ethical considerations. In this way , each team had more space to elaborate on their collaboration process than in pre vious submissions. Furthermore, all teams had to giv e their consent for their responses to be published in a scientific paper . The an- swers were collected into a Google Sheet with links to out- standing PDF files. In total, 67 submissions were collected, including 34 submissions that used text-to-music models such as Udio and/or Suno in the human-AI collaboration process and thus were considered disqualified due to the judges inability to "assess the le vel and manner of the use of AI in each entry" and an inability to obtain "a descrip- tion of the data used to train the AI model." [38] After these 34 submissions were excluded, 33 effecti v e partic- ipating teams remain in the 2024 edition. The completed questionnaires were then handed ov er to the research group excluding personal data. 3.2 Methodology and V alidation Due to the change in participant entry methods and sub- mission format, the data dictionary from the initial HAISP dataset was adapted by the coders to better e xtract the data essential to the dataset from the written entries. The cate- gories were e valuated one by one by the four researchers, iterating three times, with testing of each new adaptation to the dictionary before reaching final consensus. For each iteration of the data dictionary , two coders tested it on two sample entries to ensure that the categories were properly defined and applicable for the ne w data. 3.3 Data Statistics The HAISP dataset for the 2024 edition consists of data from 33 teams, representing 22 dif ferent countries and re- gions. The United States has the highest representation, with 12 teams participating. The United Kingdom follo ws with four teams, while Switzerland has three. Germany and Spain are each represented by two teams. Other represented countries include the Netherlands (NLD), Colombia (COL), Japan (JPN), Italy (IT A), Brazil (BRA), Thailand (THA), Canada (CAN), France (FRA), T unisia (TUN), Denmark (DNK), Hungary (HUN), T urke y (TUR), China (CHN), and Chile (CHL). The type of af filiation of the HAISP dataset 2024 edi- tion reflects a significant shift compared to the 2023 edi- tion. The most notable change is a clear trend a way from participants typically coming from academic backgrounds to ward those who work in the creati ve industry , suggesting a gro wing engagement of professional artists and creativ es with AI-dri ven music composition, in artistic and commer - cial sectors rather than academic research setting. In 2023, 58.8% of participants were af filiated with academia, mak- ing it the dominant category . Ho we ver , in 2024, academic af filiation dropped to just 19.57%, while the creati ve in- dustry sur ged to 58.7%, making it the largest represented category in this year’ s dataset. The HAISP dataset for the 2024 edition of the AISC sho wcases a div erse array of AI models and tools em- ployed by participating teams. Compared to the 2023 edi- tion, which saw the usage of 74 dif ferent AI tools, the 2024 dataset reflects an e ven broader spectrum of AI applica- tions with 82 dif ferent AI tools, excluding the other tools used by the disqualified participants. These tools include AI-po wered music generation models, v oice cloning soft- ware, AI-driv en mixing and mastering tools, AI-assisted composition platforms, and more. Some teams utilized publicly a vailable AI tools lik e Musicfy or Kits.ai, while others employed custom-b uilt AI models tailored to their specific creati ve needs, like Purr Data. A significant portion (42%) of teams in 2024 contin- ued to use ChatGPT and OpenAI’ s GPT -based models for lyrics, structure, and creativ e assistance. Additionally , the rise of Stability AI models, such as Stable-Audio-Open 1.0 and Stable Dif fusion XL, suggests a growing reliance on AI for both music generation and visual content creation. 4. COMP ARA TIVE AN AL YSIS The HAISP dataset re veals that while AI-assisted song- writing can enhance creati vity and efficienc y , musicians frequently encounter challenges related to control, trans- parency , and process integration when w orking with AI tools. Se veral recurring themes, as presented below , emer ge from the dataset that highlight the limitations of current AI models. 4.1 Contr ol T welve participating teams e xpressed frustration ov er the lack of fine-grained control ov er AI-generated outputs, es- pecially when it comes to using the more popular and easily accessible AI systems. Users specifically choose systems that allo w for greater control and flexibility o ver the outputs, highlighting systems whose af fordances allow them to control “...mechanisms such as text/audio prompt- ing and loop generation” (T eam 63). W ith systems that do not allo w for such user modifications, many feel that the results are limited, and “...mostly based on seed luck and good prompting” (T eam 28). One team in particular noted that they chose to use an AI tool created by Ele v en- Labs [39] not only because they felt it allowed for “greater control, ” (T eam 52) but because of their ethical stance as a team, noting that Ele venLabs was more ethically trans- parent in the creation of its music database, using only li- censed content from Shutterstock [40]. Decisions about tools are not only based on control ov er the process or out- put, but also on the team’ s control ov er ho w to accommo- date or apply their ethical positions. The 2024 HAISP dataset expansion reinforces man y of the themes from the 2023 dataset, particularly in regards to Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 30 Figur e 1 . Bar Chart comparing the af filiation of AI Song P articipants in 2023 and 2024, sho wcasing an increase in creati v e industry af filiation and decrease in academic af filiation. maintaining control o v er their creation process. In 2023, teams frequently encountered rigidity and limitations when w orking with AI tools, as e x emplified by 2023’ s T eam 13’ s e xperience w orking with Meta’ s MusicGen. Initially “f as- cinated” by the outcomes that came from this tool, the y soon realized the tool produced repetiti v e results, and fur - ther attempts to refine the output resulted in “...an unwel- come surplus of noise, leading to a sense of limitation. ” This led them to switch to Google’ s Magenta so the y could ha v e more control o v er the MIDI outputs. Other 2023 teams such as T eam 16 also noted that te xt-to-music mod- els frequently needed more specific and detailed prompting that included information on the k e y and tempo in order to a v oid the “...incoherent and sometimes noisy outputs. ” Ov erall, teams from both years sho w a stronger preference for AI tools that allo w for real-time adjustments, iterati v e prompting, and clear ethical stances so the y can acti v ely and continuously mak e choices that gi v e them the control the y desire o v er the creati v e process. 4.2 A pplications of AI In the 2024 data, we noted that participants on a v erage used 2-3 times more AI tools than the 2023 participants. 2024 AISC participants le v eraged AI for melody and har - mon y generation, using models to produce initi al musical ideas that were later refined through human music produc- tion stages lik e mixing and mastering. V oice synthesis w as another k e y application, with AI tools transforming v ocal performances or generating synthetic v oices that could be adjusted to fit the song’ s artistic vision. Unlik e 2023, the majority of the tools used were openly a v ailable tools, and not custom-b uilt and trained models. T ools lik e Stable Au- dio, Ele v enLabs, and ChatGPT 3.5/4.0 were amongst the most commonly used tools in the 2024 dataset. In contrast, 2023’ s teams often emplo yed AI tools it- erati v ely , using them to refine compositions and lyrics throughout the process, which is described more as a recur - si v e w orkflo w than step-by-step, wi th their AI tool “...pro- viding creati v e suggestions and helping us iterate more ef ficiently” (T eam 26). The iterati v e process means that teams could listen to their AI-generated song elements and add on to them with human elements as their submission de v eloped, rather than ha ving the element be a generated piece that cannot be recreated e xactly , e v en with the same prompts. 2024’ s T eam 38 described their frustration with this issue, writing that "The problem with all of this, and what mak es this w orkflo w so granular , is the AI starts to drift o v er time, losing sight of one aspect of the prompt in f a v or of another; outputs be gin to dif fer in length, timbre, language (for some reason) b ut most importantly tempo." Additionally , man y of the tools used by the 2023 partici- pants were either b uilt by the teams or were open source models that were trained by the team. In 2023, lar ge-scale AI systems lik e ChatGPT were mostly used to create sug- gestions or generate ideas for song elements such as the melody , which w as then played and recorded on real in- struments by participants, or lyrics, which were then sung by a separate AI tool or human team member . 4.3 Co-Cr eation vs. A utomation The dataset indicates a strong preference for collaborati v e AI tools o v er fully automated music generators. Man y artists w ant AI to function as an assisti v e tool rather than an autonomous composer . Users e xpress interest in AI systems that respond dynamically to their inputs, rather than generati ng static musical pieces that require e xtensi v e manual re vision. As T eam 28 noted, it w as not just that utilizing an AI tool that made the w ork co-creati v e, b ut the combination of their o wn musical t raining and skills and the w ork of the AI tools which allo wed for "... a seamless Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 31 T able 1 . Empirical mean ˜ ρ for uniformly sampled segment durations d 1 ≤ d 2 in the range [4 , 60] seconds, for different frame rates f (Hz) and tolerance windo w w (seconds). w = 0 w = 0 . 25 w = 0 . 5 f = 40 0.009 0.349 0.477 f = 20 0.016 0.343 0.477 f = 10 0.029 0.261 0.477 f = 2 0.111 0.111 0.427 3.3 Pr operties Before going into extensions and applications, it is worth pausing to take note of a fe w properties of ρ, ˜ ρ , and R . Boundedness Since gcd { d 1 , d 2 } ≤ min { d 1 , d 2 } , eqs. (1) and (3) are bounded at 1, with equality when d 2 is an integer multiple of d 1 (or vice v ersa). The minimal value 1 / min { d 1 , d 2 } is achie ved by relati vely prime ( d 1 , d 2 ) . Scale-in variance F or any positi ve rational c such that c · d 1 and c · d 2 are integers, ρ ( c · d 1 , c · d 2 ) = ρ ( d 1 , d 2 ) . Attainable values ρ ( d 1 , d 2 ) = 1 / N for some positi ve in- teger N . This is because c = 1 / gcd { d 1 , d 2 } , which satisfies the scale in v ariance property abov e, implies ρ ( d 1 , d 2 ) = ρ ( c · d 1 , c · d 2 ) = 1 / min { c · d 1 , c · d 2 } . Expected value T able 1 reports the empirical mean ˜ ρ for uniformly sampled duration pairs ov er the range [4 , 60] seconds. For the proposed default tolerance of w = 0 . 5 , the mean ˜ ρ ≈ 0 . 477 is stable for dif ferent frame rates f . 3.4 Extension 1: Section labels Equation (2) a verages ov er all unordered pairs of distinct segments. It often occurs that not all segments are rele- v ant to include in this comparison: for example, introduc- tory silences or cro wd noise may exist outside of musi- cal time and therefore not participate meaningfully in reg- ularity . Similarly , sections with significant deviations in tempo from the remainder of the recording may result in lo w scores under eq. (3), and a case could be made that these should be treated separately . More generally , one may consider a notion of restricted regularity that only compares se gments with the same sec- tion label ( e.g . , verse or c horus ). Under suitable label- ing con v entions, this vie w encapsulates the examples listed abov e, and provides a simple mechanism to e xclude seg- ments with sporadically occurring labels. This idea can be implemented with a straightforward modification to eq. (2) where a collection of distinct segment pairs P ⊂ S × S is provided rather than the entire se gmentation S : R L ( P ) = 1 | P | X ( d 1 ,d 2 ) ∈ P ρ ( d 1 , d 2 ) . (4) 2 The associati ve property of gcd and min also implies that the edge case of a segmentation consisting of only one se gment should produce a score of 1. This con v ention is adopted here. 3 δ is constrained to d + δ ≥ f so that eq. (3) is well-defined. Label agreement is a simple way to generate the pair set P , though the definition abov e supports other schemes, e.g . automatic hierarchy expansion (for approximate agree- ment) [12]. Relatedly , the temporal proximity observ ation of Smith and Goto [9] can be implemented here by gener - ating pairs of sequentially adjacent durations: P = { ( d i , d i +1 ) | 0 ≤ i < | S | − 1 } . 3.5 Extension 2: Hierarchical r egularity Equation (2) can be modified to e valuate the re gularity of hierar chical se gmentations. Note that eq. (2) operates on pairs of durations, but it does not require that the se gments under comparison are disjoint in time or form a v alid seg- mentation. If H = ( S 0 , S 1 , . . . ) denotes a multi-lev el seg- mentation (with each S i denoting no w the collection of in- terv als at the i th segmentation le vel), a pair set P can be generated by matching each segment at le vel i to its max- imally ov erlapping segment at each le vel j < i . The sim- plified case of a two-le vel hierarchy H = ( S 0 , S 1 ) yields P =  ( | s | , | t | ) | t ∈ S 1 ∧ s = ar gmax s ∈ S 0 | s ∩ t |  , where | s | denotes the duration of interv al s , and | s ∩ t | de- notes the ov erlap duration between intervals s and t . Ev al- uating ρ on each such pair captures ho w ev enly the hierar- chy di vides segments from one le vel to the next. 3.6 Extension 3: Balance Equation (1) captures a form of re gularity where durations are related by simple ratios. This dif fers from previous no- tions of regularity , which were designed to fa v or segments of equal duration [5]. This notion can be recov ered by re- placing the min normalization in eq. (1) by max : β ( d 1 , d 2 ) := gcd { d 1 , d 2 } max { d 1 , d 2 } . (5) Equation (5) thus captures the balance of d 1 and d 2 : a score of 1 is only achie ved when d 1 = d 2 , a score of 1 / 2 is achie ved when the y are related by a factor of 2, and so on. In general, β ( d 1 , d 2 ) ≤ ρ ( d 1 , d 2 ) , and it otherwise inherits the boundedness, scale-in variance, and integer reciprocal properties noted abov e. Repeating the calculations behind table 1 for ˜ β results in an expected v alue of 0.216 for uni- formly random durations and w = 0 . 5 . As abov e, this also gi ves rise to an aggre gate pairwise score B ( S ) , sampled versions ˜ β and ˜ B , and labeled and hierarchical v ariations. 4. EXPERIMENTS The proposed metrics are e valuated on the reference anno- tations provided by a v ariety of commonly used structure analysis datasets spanning multiple genres: Beatles (TUT) 174 Beatles songs using the TUT segmen- tations [13] and Isophonics beat annotations [14]. HarmonixSet 912 popular songs with segment and beat annotations [15]. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 38 Figur e 2 . Segment durations for each dataset. Jazz Structur e Dataset (JSD) 340 tracks [16]. For the labeled metrics, the chorus and theme counter fields are discarded from segment label strings, and segments la- beled silence are treated as mutually distinct. Jazz A udio-aligned Harmony (J AAH) 113 tracks with labels deri ved from the parts annotations [17]. Real-world Computing (R WC) 211 tracks (100 popular , 61 classical, 50 jazz) [18]. For labeled metrics, segments labeled as "nothing" are treated as mutually distinct, and labels are simplified by discarding parenthetical v aria- tions ( e.g . , “chorus A (+1)” 7→ “c horus A” ). SALAMI 1359 tracks from the publicly a vailable dataset [10], consisting of 4486 annotations (2243 up- per , 2243 lo wer). Sections labeled as "Z" or "silence" are treated as mutually distinct for labeled metrics, and v ariation markers are discarded ( e.g . , A 0 7→ A ). Figure 2 illustrates the distrib ution of segment durations for each dataset. The e valuation seeks to e xplore the follo wing questions: 1. Ho w do the absolute time metrics ( ˜ ρ , ˜ β ) dif fer from the musical time metrics ( ρ , β )? 2. Do structure annotations exhibit re gularity and/or balance? Does this vary with genre? 3. Are multi-le vel segmentations re gular across lev els? In service of the first question, we compared scores de- ri ved from absolute time (using the approach described in section 3.2) to the simpler forms deri ved from inte ger- v alued durations measured in beats. This analysis is re- stricted to the datasets with reference beat annotations: Beatles, HarmonixSet, J AAH, and R WC. Each segment boundary is mapped to its nearest beat, and segment du- rations d are measured in beats between the start and end boundaries. A preliminary study rev ealed sensiti vities to rounding error in beat position identification, which were resolved by including a maximization o ver { d − 1 , d, d + 1 } . The (Pearson) correlation was then computed between the musical-time and absolute-time metrics for each dataset. For the second question, unlabeled and labeled forms of the absolute time metrics were computed. As a point of comparison, metrics were also computed under restriction to adjacent segments [9], denoted here as R S , B S , etc . T able 2 . Mean regularity and balance scores using musical time, both unlabeled ( R, B ) and labeled ( R L , B L ). R R L B B L Beatles (TUT) 0.681 0.847 0.459 0.834 Harmonix 0.728 0.799 0.524 0.731 J AAH 0.741 0.878 0.488 0.869 R WC Classical 0.599 0.789 0.391 0.765 R WC Jazz 0.914 0.949 0.789 0.945 R WC Popular 0.820 0.958 0.587 0.945 T able 3 . Mean re gularity and balance scores using abso- lute time, both unlabeled ( ˜ R, ˜ B ) and labeled ( ˜ R L , ˜ B L ). ˜ R ˜ R L ˜ B ˜ B L Beatles (TUT) 0.704 0.820 0.394 0.805 Harmonix 0.730 0.789 0.498 0.719 JSD 0.732 0.606 0.344 0.591 J AAH 0.646 0.793 0.411 0.784 R WC Classical 0.506 0.709 0.298 0.673 R WC Jazz 0.818 0.856 0.720 0.838 R WC Popular 0.791 0.941 0.560 0.925 SALAMI (upper) 0.776 0.719 0.373 0.619 SALAMI (lower) 0.875 0.889 0.684 0.840 For the third question, we restrict attention to the SALAMI dataset, and ev aluate hierarchical regularity and balance using the paired upper - and lower -le vel annota- tions for each track. 5. RESUL TS 5.1 Musical time vs. absolute time T able 2 reports the av erage value for the re gularity and bal- ance metrics on each of the datasets listed abov e for which segment durations can be reliably measured in beats. As should be expected, the labeled forms are generally sub- stantially higher than the unlabeled forms. Each dataset exhibits high labeled re gularity (significantly abov e 0.5), as well as high labeled balance, indicating that similarly la- beled segments do consistently span equi v alent durations. T able 3 summarizes the absolute-time metrics across all datasets, and Figure 3 illustrates the correlation between these and the musical time data reported in table 2. The correlations are generally high (abov e 0.6), with a few no- table exceptions in the jazz and classical datasets. These exceptions may be e xplained by the tempo distributions of each dataset, illustrated in fig. 4. Recall that the ab- solute time metric uses a tolerance windo w of 0.5 sec- onds, equi valent to one beat at 120BPM. If a track is much slo wer— e.g . , R WC Classical with median tempo of 87.1, or R WC Jazz with median tempo of 89.4—the maximiza- tion in eq. (3) will not cov er a full beat, so a larger windo w may be warranted. Ho we ver , note that if the tempo is sta- ble , this becomes less of an issue because absolute- and musical-time are approximately proportional, which is ex- ploited by the scale-in v ariance property of ρ . Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 39 R egularity (unlabeled) R egularity (labeled) Balance (unlabeled) Balance (labeled) Beatles (TUT) Har monix J A AH R WC Classical R WC Jazz R WC P opular 0.73 0.80 0.85 0.81 0.91 0.93 0.95 0.95 0.59 0.66 0.85 0.70 0.55 0.70 0.59 0.74 0.48 0.38 0.82 0.36 0.74 0.80 0.86 0.84 1.0 0.5 0.0 0.5 1.0 Figur e 3 . Pearson correlation between musical- and absolute-time metrics for each dataset. Figur e 4 . T empo deri ved from reference beat annotations. Each point corresponds to the mean tempo for one record- ing. 120BPM is marked in red as a reference point. Figure 5 illustrates the distrib utions of tempo stability , measured as the standard de viation of inter-beat-interv al. Datasets with high tempo stability tend to exhibit high cor - relation in fig. 3 e ven when the y contain many lo w-tempo tracks ( e.g . , Beatles, Harmonix, and R WC Pop). 5.2 Unlabeled and labeled regularity Figure 6 illustrates the relationship between labeled and unlabeled regularity metrics. Consistent with the summary in table 3, the unlabeled regularity scores are generally quite dispersed, while the labeled scores ske w higher , con- firming that segments belonging to dif ferently labeled sec- tions may not conform to regular duration relationships. T wo exceptions to this observ ation are JSD and SALAMI (upper). In both cases, labeled regularity decreases from the unlabeled scores. These cases may be explained by the use of short silence segments, which di vide e venly into most other segments, contrib uting many lar ge values to the Figur e 5 . T empo stability for each dataset, as measured by the standard de viation of local tempo deriv ed from inter- beat interv als in the reference annotations. Each point rep- resents the standard de viation of tempo for one recording. Beatles (TUT) Har monix JSD J A AH R WC Classical R WC Jazz R WC P opular S AL AMI (upper) S AL AMI (lower) Figur e 6 . Labeled vs. unlabeled re gularity metrics for each annotation in each dataset. a verage in eq. (2). In the labeled re gularity calculation, each silence segment is treated as distinct, eliminating this source of inflation. Segments of this nature are less pre v a- lent in the other datasets ( e.g . , R WC or J AAH). T able 4 summarizes the results of regularity and balance when computed on adjacent segment pairs. While there are clear regularity trends, confirming the prior w ork of Smith and Goto, the ef fect is not generally as prev alent as the label-agreement results reported in table 3. 5.3 Balance vs. Regularity Figure 7 illustrates the distrib ution of the dif ference be- tween labeled regularity and labeled balance in each T able 4 . Sequential regularity and balance metrics in both musical time ( R S , B S ) and absolute time ( ˜ R S , ˜ B S ). R S ˜ R S B S ˜ B S Beatles (TUT) 0.666 0.651 0.399 0.398 Harmonix 0.720 0.696 0.500 0.483 JSD — 0.729 — 0.479 J AAH 0.786 0.699 0.563 0.495 R WC Classical 0.604 0.521 0.380 0.311 R WC Jazz 0.935 0.872 0.818 0.780 R WC Popular 0.839 0.806 0.592 0.562 SALAMI (upper) — 0.746 — 0.420 SALAMI (lo wer) — 0.882 — 0.753 Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 40 Figur e 7 . The distributions of dif ference between labeled regularity and balance: ∆ = ˜ R L − ˜ B L for each dataset. Figur e 8 . Hierarchical scores on SALAMI. dataset. The balance scores cannot e xceed the regular - ity scores—so the dif ference is non-negati ve—though each dataset does exhibit v ery high correlation between reg- ularity and balance: all correlation coef ficients exceed 0.95. While some datasets generally tend to match bal- ance and regularity (Beatles, JSD, J AAH, R WC Jazz and Pop), others div erge substantially (Harmonix, R WC Clas- sical, SALAMI). This demonstrates that regularity and bal- ance are indeed distinct qualities of segmentation. 5.4 Hierarchical r egularity Figure 8 illustrates the distribution of hierarchical re gular- ity and balance scores on the SALAMI dataset. As ex- pected, the balance scores tend to be lo w due to the shorter duration of segments in the lo wer le vel annotations. Interestingly , the regularity scores are generally quite high, with a median value of 0.969. This can be inter - preted broadly as confirming that upper -lev el segments are comprised of whole repetitions of lo wer-le vel se gment du- rations. While this may be intuiti vely expected gi v en the annotation rules, it is not an obvious conclusion from the single-le vel analyses in the pre vious section. Figure 6 il- lustrates that lo wer-le vel se gmentations tend to be highly regular ( ˜ R L ≈ 0 . 889 ) and highly balanced ( ˜ B L ≈ 0 . 840 ), while upper -lev el segmentations are slightly less re gular ( ˜ R L ≈ 0 . 719 ) and often less balanced ( ˜ B L ≈ 0 . 619 ). 6. DISCUSSION From the findings abov e, we can draw some conclusions about the role of regularity in music structure analysis. First, because these analyses are conducted on refer - ence annotations (not model outputs), the results reflect the pbeha vior of human annotators, and not algorithms. The distrib ution plots in fig. 6 indicate that although the mean regularity scores are generally high across datasets, there is considerable v ariability across individual tracks. While these results deri ve from the absolute time metrics, the high correlation with the musical time metrics suggests that this is generally not explained by tempo v ariation, and rather reflects widespread and meaningful structural irregularity in many datasets. This suggests that regularity , if taken as a design principle in segmentation algorithms, should be treated with some care to allo w for irregular segmentations when warranted by the track in question. Second, the discrepancy between labeled and unlabeled metrics can be quite large (Beatles, Harmonix, R WC Clas- sical and Pop). This corresponds to non-tri vial interactions between the regularity and repetition principles (as related to segment label agreement), which had not been identi- fied in pre vious studies. Modeling and fruitfully exploiting these interactions would be an interesting direction for fu- ture work in structure analysis algorithms. Third, some datasets e xhibit significant discrepancies between regularity and balance (Harmonix, SALAMI). This demonstrates that segment durations in f act exhibit more complex patterns than simple equi v alence. 7. LIMIT A TIONS The proposed methods are applicable to quantitati ve e v al- uation of segmentations, but the y do exhibit some limita- tions. First, the absolute time definition does appear to ex- hibit sensiti vity to tempo v ariation, in particular as it relates to the choice of tolerance windo w . In situations where high tempo v ariation may be expected, it may be preferable to either apply the musical time formulation using estimated beat positions (if they are reliable), or adapt the tolerance windo w to fit the (estimated) tempo of the track. Second, short segments may artificially inflate scores by being easily di visible into long segments. This is partially addressed by the labeled extension, as short se gments tend to be sporadic and unrelated to the majority of a track, e.g . , a short silence segment at the be ginning or end. Finally , the proposed metrics do not easily lend them- selves to dif ferentiable formulations which may be in- tegrated as learning objecti v es or penalties in current gradient-based learning frame works. While it may be pos- sible to do so, e.g. , by pre-computing a look-up table of pairwise duration comparisons, other difficulties may arise in adapting the ideas into practical segmentation al- gorithms. Still, the proposed metrics may be more easily integrated as post-processing steps, e .g. , to identify mean- ingful le vels to include in a multi-le vel se gmentation, or to select among a collection of proposed segmentations gen- erated by an ensemble of methods. 8. A CKNO WLEDGMENTS The author thanks Qingyang (T om) Xi and Meinard Müller for helpful discussions and early feedback. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 41 9. REFERENCES [1] G. Peeters, “Deriving musical structures from sig- nal analysis for music audio summary generation: "sequence" and "state" approach, ” in Computer Music Modeling and Retrieval, International Sym- posium, CMMR 2003,Montpellier , F rance , May 26-27, 2003, Re vised P apers , ser . Lecture Notes in Computer Science, U. K. W iil, Ed., v ol. 2771. Springer , 2003, pp. 143–166. [Online]. A v ailable: https://doi.or g/10.1007/978- 3- 540- 39900- 1\_14 [2] J. Paulus, M. Müller , and A. Klapuri, “State of the art report: Audio-based music structure analysis. ” in Pr oceedings of the 11th International Society for Music Information Retrieval Confer ence . ISMIR, Aug. 2010, pp. 625–636. [Online]. A v ailable: https://doi.or g/10.5281/zenodo.1417289 [3] O. Nieto, G. J. Mysore, C.-i. W ang, J. B. L. Smith, J. Schlüter , T . Grill, and B. McFee, “ Audio-based mu- sic structure analysis: Current trends, open challenges, and applications, ” T ransactions of the International So- ciety for Music Information Retrieval , Dec 2020. [4] G. Sargent, F . Bimbot, and E. V incent, “ A regularity-constrained viterbi algorithm and its ap- plication to the structural segmentation of songs. ” in Pr oceedings of the 12th International Society for Music Information Retrieval Confer ence . IS- MIR, Oct. 2011, pp. 483–488. [Online]. A v ailable: https://doi.or g/10.5281/zenodo.1415950 [5] ——, “Estimating the structural se gmentation of popular music pieces under regularity constraints, ” IEEE/A CM T ransactions on A udio, Speech, and Lan- guag e Pr ocessing , vol. 25, no. 2, pp. 344–358, 2017. [6] A. Marmoret, J. E. Cohen, and F . Bimbot, “Barwise music structure analysis with the correlation block- matching segmentation algorithm, ” T r ansactions of the International Society for Music Information Retrieval , Nov 2023. [7] B. McFee and D. P . W . Ellis, “Learning to segment songs with ordinal linear discriminant analysis, ” in 2014 IEEE International Confer ence on Acoustics, Speech and Signal Pr ocessing (ICASSP) , 2014, pp. 5197–5201. [8] A. Maezaw a, “Music boundary detection based on a hybrid deep model of nov elty , homogeneity , repetition and duration, ” in ICASSP 2019 - 2019 IEEE Inter- national Confer ence on Acoustics, Speech and Signal Pr ocessing (ICASSP) , 2019, pp. 206–210. [9] J. B. L. Smith and M. Goto, “Using priors to improve estimates of music structure. ” in Pr oceedings of the 17th International Society for Music Information Retrieval Confer ence . ISMIR, Aug. 2016, pp. 554–560. [Online]. A v ailable: https://doi.org/10.5281/ zenodo.1416916 [10] J. B. L. Smith, J. A. Burgo yne, I. Fujinaga, D. D. Roure, and J. S. Do wnie, “Design and creation of a lar ge-scale database of structural annotations. ” in Pr oceedings of the 12th International Society for Music Information Retrieval Confer ence . ISMIR, Oct. 2011, pp. 555–560. [Online]. A v ailable: https://doi.or g/10.5281/zenodo.1416884 [11] C. Raf fel, B. McFee, E. J. Humphrey , J. Salamon, O. Nieto, D. Liang, and D. P . W . Ellis, “mir_ev al: A transparent implementation of common mir metrics. ” in Pr oceedings of the 15th International Society for Music Information Retrieval Confer ence . ISMIR, Oct. 2014, pp. 367–372. [Online]. A v ailable: https: //doi.or g/10.5281/zenodo.1416528 [12] B. McFee and K. Kinnaird, “Improving structure e valuation through automatic hierarch y expansion, ” in Pr oceedings of the 20th International Society for Music Information Retrieval Confer ence . ISMIR, Nov . 2019, pp. 152–158. [Online]. A v ailable: https: //doi.or g/10.5281/zenodo.3527764 [13] J. Paulus, “Improving markov model based music piece structure labelling with acoustic information. ” in Pr oceedings of the 11th International Society for Music Information Retrieval Confer ence . ISMIR, Aug. 2010, pp. 303–308. [Online]. A v ailable: https: //doi.or g/10.5281/zenodo.1416732 [14] C. Harte, “T o wards automatic e xtraction of harmony information from music signals, ” Ph.D. dissertation, Department of Electronic Engineering, Queen Mary , Uni versity of London, 2010. [15] O. Nieto, M. McCallum, M. Davies, A. Robertson, A. Stark, and E. Egozy , “The Harmonix Set: Beats, do wnbeats, and functional segment annotations of western popular music, ” in Pr oceedings of the 20th International Society for Music Information Retrieval Confer ence . ISMIR, No v . 2019, pp. 565–572. [Online]. A v ailable: https://doi.org/10.5281/zenodo. 3527870 [16] S. Balke, J. Reck, C. W eiSS, J. AbeSSer , and M. Müller , “JSD: A dataset for structure analysis in jazz music, ” T ransactions of the International Society for Music Information Retrieval , No v 2022. [17] V . Eremenko, E. Demirel, B. Bozkurt, and X. Serra, “ Audio-aligned jazz harmony dataset for automatic chord transcription and corpus-based research, ” in Pr oceedings of the 19th International Society for Music Information Retrieval Confer ence . ISMIR, Sep. 2018, pp. 483–490. [Online]. A v ailable: https: //doi.or g/10.5281/zenodo.1492457 [18] M. Goto, H. Hashiguchi, T . Nishimura, and R. Oka, “R WC music database: Popular , classical and jazz music databases. ” in Pr oceedings of the 3r d Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 42 International Confer ence on Music Information Re- trieval . ISMIR, Oct. 2002. [Online]. A v ailable: https://doi.or g/10.5281/zenodo.1416474 Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 43 ON THE DE-DUPLICA TION OF THE LAKH MIDI D A T ASET Eunjin Choi 1 Hyerin Kim 2 Jiw oo Ryu 2 J uhan Nam 1 Dasaem Jeong 2 1 Graduate School of Culture T echnology , KAIST , South K orea 2 Department of Art & T echnology , Sogang Uni versity , South K orea {jech,juhan.nam}@kaist.ac.kr, {kime0225, clayryu338}@gmail.com, {dasaemj}@sogang.ac.kr ABSTRA CT A lar ge-scale dataset is essential for training a well- generalized deep-learning model. Most such datasets are collected via scraping from v arious internet sources, in- e vitably introducing duplicated data. In the symbolic mu- sic domain, these duplicates often come from multiple user arrangements and metadata changes after simple editing. Ho wev er , despite critical issues such as unreliable train- ing e valuation from data leakage during random splitting, dataset duplication has not been extensi v ely addressed in the MIR community . This study in v estigates the dataset duplication issues reg arding Lakh MIDI Dataset (LMD), one of the lar gest publicly a vailable sources in the sym- bolic music domain. T o find and ev aluate the best retrie val method for duplicated data, we employed the Clean MIDI subset of the LMD as a benchmark test set, in which dif- ferent versions of the same songs are grouped together . W e first e valuated rule-based approaches and pre vious sym- bolic music retrie val models for de-duplication and also in- vestig ated with a contrastiv e learning-based BER T model with v arious augmentations to find duplicate files. As a re- sult, we propose three dif ferent versions of the filtered list of LMD, which filters out at least 38,134 samples in the most conserv ativ e settings among 178,561 files. 1. INTR ODUCTION As data-dri ven approaches become mainstream and huge generati ve neural networks become popular , the signifi- cance of large-scale datasets is increasing. The tremen- dous performance of generati ve models has been attrib uted to massi ve dataset a vailability . For e xample, the av ailabil- ity of billions of text-image pair data boosted the high- quality image synthesis from text in the computer vision domain [1]. Among se veral dataset collection strate gies, web scraping [2], or synthesis [3] has emer ged as the most practical method for assembling large-scale datasets, gi v en the high costs associated with manual collection. Ho wev er , collecting lar ge-scale datasets via web crawl- ing can ine vitably cause problems, including priv ac y , copy- © F . Author , S. Author , and T . Author . Licensed under a Creati ve Commons Attribution 4.0 International License (CC BY 4.0). Attribution: F . Author , S. Author, and T . Author , “On the de-duplication of the Lakh MIDI dataset”, in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf ., Daejeon, South K orea, 2025. right, and data duplication issues. In this paper , we focus on data duplication, particularly highlighting issues in the use of the Lakh MIDI Dataset (LMD) [2], which is widely recognized as one of the lar ge-scale datasets in MIR. While there are studies in other fields that discuss the do wnsides of dataset duplication [4] and the benefits of de-duplication [5], we found that such discussions are notably lacking in the music information retrie val (MIR) community . One re- lated work is the e xamination of dataset validity issues in GTZAN [6, 7], and our work shares this intent by address- ing the correct usage of LMD through de-duplication. In particular , we ar gue that the current use of the LMD in symbolic music generation impairs the v alidity of exist- ing experiments. The issue arises from unreliable training e valuations caused by data leakage during random splits. Duplicates across training, v alidation, and test splits can ske w e valuation metrics such as cross-entrop y loss, which many pre vious studies used to assess their model perfor - mance [8–13]. In the music generation domain, subjec- ti ve e v aluation is costly , and objecti ve e valuation metrics are often insuf ficient to fully capture the quality of gener- ated music. Consequently , studies rely on v alidation loss as a metric to claim non-ov erfitting and model ef fectiv e- ness [13]. Duplicates can also bias listening-based e valua- tions, particularly when the test split is used for condition- ing. This is common practice in conditional music gener- ation [10, 12–14]. If the duplication remains unaddressed, random splits that lead to data leakage will continue to un- dermine the reliability of e valuations. Ho wev er , manually finding duplicates in a lar ge-scale dataset such as LMD is virtually infeasible. Therefore, we explored cleaning LMD using rule-based and neural ap- proaches. Our contrib utions are as follows: • How to clean the LMD? W e ev aluated rule-based methods and e xisting symbolic music retrie val models for duplicate detection. W e also explored training an unsupervised contrasti v e BER T model with augmentations to detect duplicates. • How can we e valuate the de-duplication? W e propose an ev aluation method to e v aluate duplicate detection by utilizing the metadata of the Clean MIDI subset of LMD, referred to as LMD-clean in this paper . • How many duplicates exist in the dataset? Using the duplicate detection method with the best per - formance, we classify the duplicated files in the LMD. 44 Paper Dataset V ersion Split Evaluation MuseGAN [15] LPD-5-matched N.A. . MIDI-Sandwich2 [8] LPD-full N.A. NLL LakhNES [9] LMD-full train, v alid PPL PopMA G [10] LMD-matched random PPL, ˇ “ ( MMM [16] LMD-full N.A. . PiRhDy [17] LMD-full N.A. . MMD [18] LMD-matched N.A. . Han et al. [19] LMD-full N.A. . Musef ormer [11] LMD-full 8:1:1 PPL MusicBER T [20] LMD-full N.A. . FIGAR O [12] LMD-full 8:1:1 PPL, ˇ “ ( MIDI2V ec [21] LMD-matched 9:1 with CV . YM2413-MDB [22] LMD-full 9:1:1 . Sulun et al. [23] LM(P)D-full, matched N.A. . Han et al. [24] LMD-full N.A. . Anticipatory [13] LMD-full 87:6:6 PPL, ˇ “ ( text2midi [14] LMD-full (MidiCaps) N.A. ˇ “ ( T able 1 . Studies that used LMD for training. Studies men- tioned the dataset duplication issue are bolded . LPD is the piano roll version of LMD suggested by [15]. Split strate- gies are N.A. when not described in the paper . CV means cross-v alidation. The last column explains whether the au- thors utilized the dataset during the e valuation. ˇ “ ( means that the dataset is used in the listening test. Even with the most conserv ati ve threshold of rejection, we find 38,134 duplicated files to be filtered. Finally , we present a filtering list of LMD from both our proposed configuration and the most conserv ativ e threshold. 1 2. RELA TED WORKS 2.1 LMD and Related Studies LMD [2] is a dataset released in 2016 with a method for ef- ficiently matching lar ge-scale MIDI corpus collected from the Internet to the Million Song Dataset (MSD) [25]. At the time, the dataset was intended to be used for applications such as content-based retrie val, corpus studies of music structure and patterns, and transcription using paired au- dio and MIDI. Di ver ging from its initially suggested appli- cations, this dataset has become frequently used for train- ing symbolic music generation models, primarily because it is one of the lar gest av ailable symbolic music datasets. Among the a vailable LMD v ersions, LMD-full, which con- tains 178,561 files with unique MD5 hashes, is utilized as the most popular for multi-instrumental pop music gener - ation, and LMD-matched, which consists of 45,129 songs matched with MSD is also used in se veral studies. Since the release of LMD, larger -scale datasets based on web scraping—such as MetaMIDI (MMD) [18] and GigaMIDI [26]—ha ve emer ged. Notably , the recently re- leased GigaMIDI dataset is a superset of LMD. Addition- ally , the introduction of the MidiCaps dataset [27], which incorporates LLM-generated text captions aligned with LMD, has further expanded applications of LMD in te xt- to-music generation [14] and music retrie val tasks [28, 29]. As sho wn in T able 1, most studies emplo yed the LMD- full and split it for training without mentioning the split 1 All training and e valuation code is publicly a vailable: https:// github.com/jech2/LMD_Deduplication strategies. Also, se v eral papers employed the NLL loss or PPL v alues for their e v aluation. W e note that a fe w pa- pers in T able 1 pointed out the duplication issues within the dataset and remov ed the duplicated files using a rule- based approach, such as MIDI encoding hash matching. Ho wev er , we found that the rule-based approach is not suf- ficient to remov e all of the duplicated files in the dataset, which we will sho w in the ev aluation section. In this study , we used LMD-clean as our de-duplication study and test dataset. This dataset contains 17,184 files or ganized by the artist and song name in the directory and filename, which we found to ha ve multiple MIDI files of the same song. 2.2 Dataset De-duplication and Related Issues Recently , the need for dataset de-duplication has gained attention across v arious domains. In computer vision, [4] proposed a method to eliminate duplication by compress- ing CLIP features using a contrasti ve feature compression technique. In music information retrie val (MIR), [30] re- cently introduced a method for detecting exact duplicates in training data using audio-based music similarity met- rics. Furthermore, in natural language processing, [5] in- vestig ated the effects of dataset de-duplication by applying exact substring matching and hash-based techniques. Their findings sho wed that removing duplicates impro ves lan- guage model performance, reduces training time, and lo w- ers the rate of training data memorization without harm- ing perplexity . In this work, we focus on the de-duplication method for lar ge-scale symbolic music datasets. 2.3 Symbolic Music Understanding and Retrieval W ith the adv ent of large-scale language understanding models such as BER T [31] and B AR T [32] in the natu- ral language domain, se veral counterparts ha ve been in- troduced for symbolic music, including MidiBER T -Piano [33], MusicBER T [20], and PianoB AR T [34]. Among these, MusicBER T is a lar ge-scale pre-trained model trained on LMD and a pri vate dataset. While earlier meth- ods focused solely on the MIDI modality , recent ap- proaches ha ve explored symbolic music retrie v al using text, led by the introduction of CLaMP [35], which learns joint embeddings of symbolic music and text through con- trasti ve learning, using ABC notation as its input format. CLaMP2 [28] extends this approach by enhancing the symbolic encoder to support both ABC and MIDI formats with multilingual support. CLaMP3 [29] further general- izes the model to handle additional modalities, including audio and images. W e used these models for our dataset de-duplication task by le veraging their embeddings. 3. DUPLICA TION TYPES IN THE LMD W e describe the types of duplicated MIDI files observed in LMD-clean. In a broad sense, if we focus on music gen- eration, all arrangements of the same song should be de- fined as duplication (i.e., the same song cannot be in dif fer- ent splits). Ho wev er , in tasks such as music arrangement, dif ferent arrangements can be considered as dif ferent data Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 45 samples. Considering this aspect, we separated the dupli- cation type into two cases: hard duplication and soft dupli- cation 2 . This work focuses on detecting hard duplication. 3.1 Hard Duplication from Similar Arrangement W e define hard duplication as files that share identical sec- tions of arrangements with minor dif ferences. These dif- ferences include instrument mapping or order , tempo, start of fset, file length, missing tracks, or note-lev el alterations such as pitch, duration, or velocity . Melodic variations may also appear , such as added ornamentation or changes in the number of chord tones played by a specific instrument. K ey transpositions are considered hard duplication when other musical elements remain nearly identical b ut are treated as soft duplication if accompanied by significant stylistic or structural changes. W e assume that some hard duplicates were likely collected as users modified and re-uploaded e x- isting files originally created by other arrangers. 3.2 Soft Duplication from Differ ent Arrangement In soft duplication, MIDI files preserve essential musical elements such as melody , harmon y , ornaments, and instru- mentation b ut differ in arrangement style, reflecting the di- verse styles of indi vidual arrangers. Such duplication in- cludes cases where the core melody remains unchanged or similar , b ut accompaniment styles (e.g., arpeggio, waltz), pitch ranges, and ov erall structure v ary considerably . In cases of extremely dif ferent arrangements, these v ariations might lead a listener to percei ve them as distinct songs. 4. DUPLICA TE DETECTION: D A T ASET AND CONVENTION AL APPRO A CHES Along with the dataset for e valuation, we first e xplored se veral approaches for identifying duplicates, including simple rule-based methods and pre-trained symbolic mu- sic retrie val models. 4.1 Dataset W e used LMD-clean as our ev aluation benchmark to as- sess ho w well each method detects duplicates within the dataset. LMD-clean is org anized by artist folders and song filenames, where duplicate instances of the same song by the same artist are labeled with v ariations in the filenames (e.g., Dancing Queen.mid , Dancing queen.2.mid ). According to this metadata, 10,355 out of 17,184 files in LMD-clean are considered duplicates. 4.2 Rule-based Appr oach The follo wing rule-based methods serve as baselines for identifying duplicates in the dataset. W e assumed that hard duplicated samples share highly similar MIDI-le vel fea- tures. Based on this assumption, we explored se v eral meth- ods aimed at detecting and filtering out files with identical or nearly identical features at the beat or pitch le vel. 2 W e sho w examples of duplication types in the companion website. 4.2.1 MIDI Encoding Hash As discussed in Section 2.1, some studies [11, 20, 23] em- ployed a hash-based approach to detect duplicated MIDI files with dif ferent metadata. Here, we used the file de- duplication code of MusicBER T [20]. The string versions of Octuple representations are encoded according to the MD5 hash v alue, and the hash values of all MIDI files from LMD-clean are compared. 4.2.2 Beat P osition Entr opy In our preliminary study , we found there are man y dupli- cates that ha ve exact ly the same music b ut with different instrument mapping or track order . T o detect these dupli- cates, we applied a simple method that checks the distrib u- tion of note position within a bar using the MIDI encoding scheme of [36]. W e computed entropy v alues from note position distrib utions at a 16th-note resolution. Files with identical entropy v alues were identified as hard duplicates. 4.2.3 Chr oma-DTW T o detect duplicates with similar pitch content, we mea- sured the chroma-le vel distance between MIDI files. Piano roll-based chromagrams were first generated and aligned by transposing them with the highest pitch occurrence across files. Dynamic T ime W arping (DTW) was then ap- plied to measure the similarity between aligned chroma- grams. T o reduce the computational cost of applying DTW to the entire dataset, we first computed pitch histograms for all files and measured the pairwise Kullback-Leibler (KL) di ver gence. For each file, we selected the top 250 candidates with the lo west KL div ergence and then applied chromagram-based DTW to these candidates. Although this approach discards temporal information, it serves as a rough prefiltering step. 4.3 Pre vious Symbolic Music Embedding Models W e utilized pre-trained MusicBER T [20] and CLaMP model series [28, 29, 35] since the y support multi- instrumental MIDI. For MusicBER T , we used the pre- trained MusicBER T -small and MusicBER T -base mod- els for inference. For CLaMP , we used the pre-trained CLaMP-512, CLaMP-1024, CLaMP2 and CLaMP3 mod- els. Since the CLaMP-512 and CLaMP-1024 models use XML for input files and ABC notation for their internal data representation, the MIDI files are first con v erted using Musescore batch processing and then con v erted with the XML to ABC con v ersion algorithm. For CLaMP 2 and 3, we con v erted MIDI to MTF , their MIDI encoding scheme. 5. DUPLICA TE DETECTION: A CONTRASTIVE LEARNING-B ASED APPR O A CH Pre vious pre-trained symbolic embedding models were not originally trained for duplicate detection. Inspired by [37, 38] that ev aluated the rob ustness of audio or music em- bedding models against perturbations such as pitch shift, we explore whether training with such perturbations w ould improv e song identification despite v ariations. T o this end, Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 46 Lev el A ugmentation V alue Note Onset Shift (-2, 2) Note Duration Shift (-4, 4) Note V elocity Shift (-3, 3) T rack Pitch Octav e Shift (-24, 24) T rack Inst Order Shuf fle T rack Inst Mapping Except Drum T rack Inst Drop Less than 50% T rack Bar Drop 15% Segment Bar Shift (1, 4) Segment Note Drop 15% Segment Pitch T ranspose (-6, 6) T able 2 . List of augmentations for generating MIDI v aria- tions. V alues in braces are in the token le vel (inclusi ve). we de veloped a BER T -based model and applied v arious augmentations to the positi ve samples in contrasti ve learn- ing, which we refer to as CAugBER T . 5.1 Data Representation W e used the LMD-full dataset, excluding all files present in LMD-clean by matching MD5 hash v alues. The result- ing dataset is referred to as LMD-filtered. W e randomly split the LMD-filtered into 98:1:1 ratio and pre-processed the MIDI files with Octuple encoding using MidiT ok [39]. Although we initially held out the remaining 1% for test- ing, we did not use it in our final experiments, as we chose to e valuate our model on LMD-clean instead. 5.2 Data A ugmentation T o reflect the types of variations described in Section 3.1, we constructed positi ve pairs for contrasti ve learning us- ing dif ferent augmentations, referred to as MIDI variation augmentation. Each augmentation of MIDI v ariation is in- dependently and randomly applied during training. The de- tails of these augmentations are provided in T able 2. In ad- dition, follo wing [24], we used neighbor segments, which are dif ferent parts from the same piece, as positiv e pairs for training. Each piece was first se gmented into chunks of 1024 tokens, and then a random se gment was selected. MIDI v ariation augmentation was also applied to neighbor segments to enhance rob ustness. 5.3 Model Description The implementation of CAugBER T is based on the code from [24], which applies contrasti ve learning to a BER T architecture. T o align with the parameter settings of MusicBER T -small, we used a 4-layer transformer with a sequence length of 1024, hidden size and vocab ulary em- bedding size of 512, and a feedforward dimension of 2048. W e use a total batch size of 64 across two A6000 GPUs. For mask ed language modeling (MLM), we adopted the same element, compound, and bar -lev el masking strategies used in MusicBER T . Contrasti v e learning was guided by the NT -Xent loss [40]. The final loss was computed as a weighted sum of the MLM and contrasti ve (NT -Xent) losses, with weights of 0.3 and 1.0, respecti vely . During training, each encoded MIDI segment w as aug- mented using either MIDI v ariation or neighbor augmen- tation described in Section 5.2, to maximize the di versity of augmentations within each batch. For v alidation, fix ed manual seed v alues were used to maintain consistent aug- mentations across v alidation batches. T o e v aluate the ef- fecti veness of the contrasti ve learning approach for dupli- cation detection, we conducted an ablation study on the contrasti ve loss, as presented in T able 3. 6. EV ALU A TION W e ev aluated all approaches from two perspecti v es: (1) Does the system rank duplicates as more similar than oth- ers? (2) Ho w accurately does it identify true duplicates? T o answer these questions, we utilized the metrics that are commonly used for recommendation and retrie val systems. 6.1 Measuring Similarities For MIDI Encoding Hash, the similarity between samples was set as 1 when encoding hashes matched. Similarity of beat position entropy w as computed by subtracting the ab- solute dif ference in entropy from the maximum v alue of 1. For Chroma-DTW , the similarity was calculated as 1 minus the DTW distance. In the MusicBER T series, similarity was measured using the cosine similarity of the a verage to- ken embeddings from the T ransformer’ s final hidden layer . For CAugBER T , we used the [CLS] token embedding from the final hidden layer . All BER T -based models utilized 512-dimensional embedding. For CLaMP series, we used a pre-trained 768-dimensional embedding where the last hidden state was a verage pooled and passed through a pro- jection layer , follo wing the code provided in [28, 29, 35]. 6.2 Evaluation with Retriev al Metrics T o ev aluate ho w well each method assigns higher simi- larity scores to the duplicates, we adopt normalized Dis- counted Cumulati ve Gain (nDCG) and Mean Reciprocal Rank (MRR) as e valuation metrics. nDCG measures ho w highly rele vant items are rank ed, assigning higher scores when duplicates appear closer to the top of the retrie val list. It is computed by normalizing Discounted Cumula- ti ve Gain with the optimal ranking where all duplicates are retrie ved at the highest possible ranks. F or each query in LMD-clean, we assign rele v ance 1 to duplicates and 0 to others when computing nDCG. MRR is defined as the a v- erage of the in v erse ranks of the highest-ranked rele vant item for the query . This corresponds to the a verage rank of the highest similarity samples among the duplicates. For the nDCG and MRR metrics, neural netw ork-based approaches outperformed the rule-based methods. Among them, the CLaMP model series consistently achie ved higher scores than BER T -based models, with CLAMP3 sho wing the best ov erall performance. Since CLaMP mod- els were specifically trained for retrie val tasks, the result is consistent with its intended design. Ho wev er , we observed that all approaches performed belo w a certain upper bound. In particular , while neu- Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 47 (%) Contour Note Density Pitch Range Complexity NM 63 . 36 76 . 63 46 . 70 48 . 67 P&L 37 . 44 10 . 11 0 . 41 50 . 59 LC-V AE-A 60 . 88 97 . 56 34 . 69 33 . 32 LC-V AE-SE 52 . 02 97 . 33 36 . 84 52 . 00 LC-Dif f 85 . 60 98 . 59 80 . 97 94 . 93 T able 1 : Pearson Correlation Coef ficient (PCC) between tar get and decoded attributes. condition on the noise le vel instead of the discrete dif fu- sion step, which [7] later adopted for LDM-based symbolic music generation. Thus, we are left with two continuous conditioning signals to be passed onto the dif fusion model. W e inject a and √ ¯ α t into ϵ θ through dedicated networks (see Figure 1). First, we apply Sinusoidal Encoding (SE) based on T ransformer positional embeddings [29] Γ( u ) =  sin ( ω i ( u )) , cos ( ω i ( u ))  d/ 2 i =0 (7) where ω i ( u ) = s u b 2 i/d [7], with d ∈ N the (ev en) dimen- sionality of the embedding, b ∈ R the base frequency , and s ∈ R a frequency scaling h yperparameter . The re- sulting SE features are passed through a linear layer with SiLU acti vations. Finally , we employ Feature-wise Linear Modulation (FiLM) [30], where two fully-connected lay- ers yield shifts and scales , respectiv ely , that modulate the acti vations of the denoiser (see Figure 2). The two condi- tioning branches run in parallel. This is equiv alent to learn- ing a single af fine transformation, where scale and shift are the sum of FiLM outputs from the attrib ute and noise lev el conditioning networks. T o enhance controllability over the generated samples, we also apply Classifier -Free Guidance (CFG) [31] to the noise prediction: ˆ ϵ θ ( z t , ξ t , a ) = (1 + w ) ϵ θ ( z t , ξ t , a ) − w ϵ θ ( z t , ξ t ) , (8) where ϵ θ ( z t , ξ t ) is the unconditional noise prediction and w ∈ R ≥ 0 is the guidance scale. T o make CFG ef fecti ve, the model must learn to predict noise both with and without attrib ute conditioning. W e achiev e this through condition- ing dr opout (depicted as 0 / 1 in Figure 1), i.e., setting the outputs of the attrib ute conditioning network to zero with a certain probability when e valuating (3). 3. EV ALU A TION 3.1 Dataset The models are designed to learn pitch sequence represen- tations from four -bar monophonic melodies. W e construct a lar ge-scale dataset comprising melodies extracted from 176,581 MIDI files from the Lakh MIDI Dataset [32]. 1 First, we assess whether each MIDI file contains time signature changes. If any are found, we segment the file and retain only sections with a 4 / 4 time signature. Each MIDI e vent is then quantized to the nearest sixteenth note. A melody is defined as a sequence of pitches within the standard 88-ke y piano range, played by an instrument 1 C. Raf fel, 2016, “The Lakh MIDI Dataset v0.1. ” [Online]. A v ailable: https://colinraffel.com/projects/lmd Contour Note Density Pitch Range Complexity Uncond. V AE 41 . 44 NM 35 . 506 58 . 436 30 . 833 47 . 61 P&L 49 . 698 67 . 836 40 . 657 87 . 80 LC-V AE-A 30 . 197 29 . 450 30 . 257 32 . 435 LC-V AE-SE 29 . 161 30 . 124 31 . 274 30 . 166 LC-Dif f 19 . 299 20 . 559 31 . 695 17 . 51 T able 2 : Fréchet Music Distance [27]. mapped to a v alid MIDI program. A melody is considered complete when a full measure of silence occurs. W e e xtract only melodies spanning at least four bars and comprising at least three distinct pitches. If multiple notes sound simul- taneously , we follo w the approach proposed in [33] and select only the highest-pitched note to ensure monophonic sequences. Subsequently , four-bar se gments are extracted using a stride of one bar . For each melody thus e xtracted, we compute 13 musical attrib utes, including those outlined in Section 3.2. Melodies are encoded as sequences of N = 64 inte gers in P = { 0 ,..., 129 } , where each element represents either a MIDI note number ( 0 - 127 ) or one of two special tok ens: note of f ( 128 ) and note hold ( 129 ). The dataset is divided into training, validation, and test sets, with training data augmented through transposition by a randomly selected number of semitones within a range of ± 1 octa ve. The final dataset, consisting of 10 , 126 , 676 unique melodies, is publicly a v ailable. 2 3.2 Musical Attributes As pre viously done in [25], we focus on four musical at- trib utes: (i) Contour , which quantifies the melodic mov e- ment in a sequence, measured by av eraging the pitch dif- ferences between consecuti ve notes; (ii) Note Density , defined as the ratio between the number of notes in the melody and the sequence length. It takes v alues in [0 , 1] ; (iii) Pitch Range , defined as the dif ference between the highest and lo west MIDI pitch v alues in the sequence, nor - malized by the range of an 88 -ke y piano. It takes v al- ues in  0 , 127 88  , where values abo ve one indicate a range exceeding A0–C8; (iv) Rh ythm Complexity , ev aluated using T oussaint’ s metrical comple xity measure [34], cor - rected for the total number of notes in the sequence [26]. By definition, it takes on discrete v alues. 3.3 Unconditional Generative Model As base unconditional model, we implement a β -V AE [35] based on MusicV AE [33]. This model, pre viously used in LDM-based symbolic music generation [7], also enables direct comparison with existing attrib ute-regularized V AEs (AR-V AEs) employing the same architecture [25, 26] (see Section 3.5). The encoder p ψ ( z | x ) consists of a two-layer bidirec- tional LSTM network fed with four -bar pitch sequence rep- resentations (see Section 3.1), followed by tw o linear lay- 2 M. Pettenò, Aug. 2024, “4 Bars Monophonic Melodies Dataset (Pitch Sequence), ” Zenodo, doi: https://doi.org/10.5281/zenodo .13369389 Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 54 (a) NM (b) P&L (c) LC-V AE-A (d) LC-V AE-SE Figur e 3 : Re gression plots comparing tar get and decoded Contour attrib utes across baseline methods. Figur e 4 : Re gression plot comparing tar get and decoded Contour attrib utes using LC-Dif f. ers parameterizing the latent posterior . The hierarchical decoder q ϕ ( x | z ) features tw o unidirectional LSTMs, with the bottom-le v el netw ork autore gressi v ely estimating the distrib ution o v er the sequence v alues via a softmax nonlin- earity [33]. As such, each pitch sequence x ∈ P N is first mapped onto a single latent code z ∈ R M , M = 256 , and since decoding amounts to a ne xt tok en prediction task, the standard β -V AE objecti v e [35] L V AE = − E p ψ ( z | x ) [ l og q ϕ ( x | z ) ] + β D KL [ p ψ ( z | x ) ∥ p ( z ) ] , (9) is implemented using cross-entrop y as reconstruction loss. The unconditional model is trained for 40 , 000 iterations on a single NVIDIA T itan R TX GPU with a batch size of 512 . The objecti v e (9) is minimized using Adam, and the learning rate is decreased e xponentially from 10 − 3 to 10 − 5 with a rate of 0 . 9999 . The h yperparameter β is an- nealed e xponential ly from 0 to 10 − 3 , which encourages the model to prioritize accurate sequence reconstruction dur - ing the early part of the training. Similarly to [33], we apply teacher forcing within the bottom-le v el decoder with a probability follo wing a logistic schedule. 3.4 Conditional Diffusion Model W ith latent codes being v ectors in R M , we implement a DDIM model with a fully-connected denoiser netw ork. 3 Sho wn in Figure 2, the denoiser comprises an input layer with 2048 linear units, f o l lo wed by three dense residual blocks. Each residual block comprises tw o stacks of Lay- erNorm, feature-wise modulation (responsible for joint at- trib ute and time conditioning), SiLU, and a linear layer , plus a residual connection that shortcuts the input and out- put of the block. Finally , the output is linearly projected back onto R M . 3 Source code and audio e xamples are a v ailable at https://mpet teno.github.io/controllable- latent- diffusion/ W e set the SE dimensionality to d = 128 . The attrib ute and noise le v el conditioning netw orks ha v e 512 and 2048 units in the first linear layer and FiLM layers, respecti v ely . In the forw ard process, β t follo ws a linear schedule from 10 − 6 to 10 − 2 o v er T = 1000 steps. Con v ersely , the number of sampling steps is set to T s = 100 . W e train the model with an attrib ute conditioning dropout probability of 20% . W e then apply CFG with a guidance scale of w = 3 . 0 [31]. In our e xperiments, CFG pro v ed fundamental to achie v e attrib ute re gularization. The resulting denoiser netw ork has 43 . 1 million param- eters, and con v er ges in just about 20 training epochs, half the iterations required by the unconditional model. 3.5 AR-V AE Baseline Methods F or comparison, we consider AR -V AEs [25, 26] with the same architecture as the unconditional model described in Section 3.3. AR-V AEs incorporate re gularization during training by means of a supervised multi-task learning ap- proach, with the goal of encoding the attrib ute a in the i -th dimension z i of their latent spaces. This is achie v ed by including an AR loss term in (9) L AR-V AE = L V AE + γ L AR , (10) where γ ≥ 0 is a tunable h yperparameter controlling the strength of the re gularization. Mezza et al. [26] propose the use of L NM AR = M AE ( z i , ˜ a ) , (11) where M AE ( · , · ) denotes the mean absolute error , and ˜ a is the z-score of a . P ati and Lerch [25] introduced a re gularization term that enforces a monotonic relationship between a and z i , i.e., L P&L AR = M AE ( t an h ( δ D z ) , s i gn ( D a ) ) , (12) where D z and D a are pairwise distance matrices between z i and a of all samples in a batch, respecti v ely , and δ > 0 is a tunable h yperparameter . As in [25], we set γ = 1 and δ = 10 . The remaining training details are the same as in Section 3.3. F or bre vity , we will later refer to the former AR method as “NM” and to the latter as “P&L. ” 3.6 LC-V AE Baseline Methods Similarly to T ian and Engel [3], we implement LC through a conditional V AE (cV AE) trained on the representations of Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 55 (a) NM (b) P&L (c) LC-V AE-A (d) LC-V AE-SE Figur e 5 : Re gression plots comparing tar get and decoded Rh ythm Comple xity attrib utes across baseline methods. Figur e 6 : Re gression plot comparing tar get and decoded Rh ythm Comple xity attrib utes using LC-Dif f. the base unconditional model (Section 3.3). The cV AE en- coder consists of four linear layers with ReLU acti v ations, follo wed by tw o Gating Mixing Layers (GML) that param- eterize the innermost latent distrib ution. The decoder mir - rors the encoder with four linear layers with ReLU acti v a- tions, follo wed by an output GML. Except for using 2048 units in the fully-connected layers and M ′ = 128 latent v ariables, the cV AE architecture is the same as in [3]. Let z ∈ R M be the latent representations of the uncon- ditional model, z c ∈ R M ′ be the latent representations of the cV AE, and a ∈ R the sequence attrib ute. The authors of [3] considered binary labels and one-hot v ectors were thus concatenated with z and z c . Instead, we deal with con- tinuous attrib utes. W e implement tw o cV AE v ariants that dif fer in ho w a is fed into the netw orks. In the first v ari ant, later referred to as LC-V AE-A, we feed ˜ z = [ z T , a ] T to the encoder , and ˜ z c = [ z T c , a ] T to the decoder . In the second v ariant, named LC-V AE-SE, we concatenate z and z c , re- specti v ely , with the attrib ute SE, i.e., ˜ z = [ z T , Γ( a ) ] T and ˜ z c = [ z T c , Γ( a ) ] T . 4. RESUL TS 4.1 Attrib ute-Contr olled Generation T o e v aluate the controllability of the generati v e models un- der scrutin y , we sample the tar get attrib utes uniformly in the range of zero to the 99 th percentile of the attrib ute dis- trib ution of the sequences in the test set. 4 These equally- spaced v alues, which we refer to as tar g et attrib utes, are fed to the respecti v e conditioning netw ork of LC-Dif f, suit- ably transformed and plugged into the re gularized dimen- 4 Limiting the range to the 9 9 th percentile is meant to e xclude those sequences with abnormally high attrib ute v alues. W e ar gue that these sequences are spurious, and we att rib ute their e xistence to the choice, borro wed from [33], of e xtracting melodies by naïv ely picking the highest note at an y gi v en time. sion z i of the AR-V AEs, and concatenated to the input v ec- tor of the LC-V AE decoder netw orks. T able 1 lists t he Pearson Correlation Coef ficients (PCC) between the tar get attrib utes and those computed from the generated sequences (the higher , the better). LC-Dif f con- sistently outperforms the tw o AR-V AEs (NM and P&L) and LC-V AEs (both with and without SE) for all attrib utes considered. Notably , LC-Dif f is the only method among those considered in the present study to yield correlation scores higher than 80% across the board. As for Contour , LC-Dif f achie v es a PCC of 85 . 60% , outperforming the ne xt-best model, NM, by o v er 22% . The dif ference is less pronounced for Note Density , where LC- Dif f ( 98 . 56% ) impro v es upon the second-best model by just 1% . Nonetheless, LC-V AE-A and LC-V AE-SE al- ready achie v e 97 . 56% and 97 . 33% , respecti v ely , suggest- ing that constraining the generati v e model is v ery ef fecti v e compared to AR methods when it comes to rendering the desired number of notes. LC-Dif f also demonstrates sig- nificant impro v ements in Pitch Range and Rh ythm Com- ple xity . F or Pi tch Range, it achie v es a PCC of 80 . 97% , e xceeding NM ( 46 . 70% ) by 34 . 27% , while NM itself out- performs LC-V AEs by approximately 10% . F or Rh ythm Comple xity , LC-Dif f achie v es a remarkable 94 . 93% , sur - passing LC-V AE-SE ( 52% ) by 42 . 93% . Concerning AR models, while NM directly encodes the (standardized) distrib ution onto the i th dimension of the latent space, there is no a priori w ay to kno w the monotonic relationship learned using the P&L re gularizati on in (12). This e xplains the near -zero correlation observ ed for Pitch Range, and, in general, the o v erall lo wer PCC. Figures 3 through 6 sho w the re gression plots of Con- tour and Rh ythm Comple xity . Figure 3 and Figure 4 illus- trate the cas e of a continuous distrib ution, while Figure 5 and Figure 6 e x emplify a case where the attrib ute tak es on inte ger v alues. Across both attrib utes, LC-Dif f is charac- terized by a lo wer spread and a clear linear trend. In Fig- ure 3, all baseline models sho w a tendenc y to produce e x- cessi v ely high contour v alues, whereas LC-Dif f (Figure 4) appears to mitig ate the issue. Lik e wise, Figure 5 re v eals that all models b ut LC-Dif f (Figure 6) tend to f ail when the tar get Comple xity v alues are lo w . 4.2 Data Fidelity T o e v aluate the quality of the generated sequences, we use the Fréchet Music Distance (FMD) [27], a metric that e x- tends the f amily of Fréchet Inception Dist ance [36] and Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 56 (a) a = 1 . 0 → a g = 1 . 0 3 (b) a = 3 . 0 → a g = 2 . 9 5 (c) a = 6 . 0 → a g = 6 . 1 9 Figur e 7 : Examples of MIDI files generated by controlling the Contour attrib ute with LC-Dif f. (a) a = 0 → a g = 0 (b) a = 1 5 → a g = 1 5 (c) a = 3 3 → a g = 3 3 Figur e 8 : Examples of MIDI files generated by controlling Rh ythm Comple xity with LC-Dif f. Quarter notes are indicated by solid v ertical lines; odd pulses (strong) are indicated by dashed lines; e v en pulses (weak) are indicated by dotted lines. Fréchet Audio Distance [37] to the symbolic music do- main. FMD w as computed between 22 , 016 melodies from the held-out test set and an equal number of generated se- quences. T o pre v ent the FMD from measuring a spurious di v er gence from the real att rib ute distrib ution, we condi- tion the generation on the attrib utes of the reference se- quences, rather than using e v enly-spaced control v alues as in Section 4.1. By conditioning with attrib utes measured from the test set, indeed, we aim to simultaneously com- pare the fidelity of generated sequences and ho w well the y conform to the desired attrib ute distrib ution. T able 2 reports the results obtained using CLaMP 2 MIDI embeddings [38] (the lo wer , the better). F or com- parison, we report the FMD between the reference test set and the output of the unconditional V AE (see Section 3.3) obtained by decoding 22 , 016 samples from N ( 0 , I ) . The results presented in T able 2 demonstrate that the proposed LC-Dif f model consistently achie v es the lo west FMD v alues across most attrib utes, indicating superior per - formance in generating samples that aligns more closely with the statistical properties of real sequences. Notably , LC-Dif f outperforms all baselines in Contour ( 19 . 299 ), Note Density ( 20 . 559 ), and Rh ythm Comple xity ( 17 . 51 ), significantly impro ving o v er both AR-V AEs and LC- V AEs. While LC-V AE-A achie v es the best Pitch Range score ( 30 . 257 ), LC-Dif f remains competiti v e ( 31 . 695 ). Ov erall, all LC methods outperform the unconditional base model ( 41 . 44 ), sho wing that introducing post-hoc control leads to more consistent and structured music gen- eration, with better alignment to the desired attrib utes. Finally , Figure 7 and Figure 8 illustrate the potential di v ersity in the generated samples produced by LC-Dif f when conditi o ne d on lo w , medium, and high v alues of Contour and Rh ythm Comple xity , respecti v ely . 5. CONCLUSIONS In this paper , we ha v e e xplored latent dif fusion through the lens of Latent Constraints (LC), demonstrating the ef ficac y of DDIM s as plug-and-play conditioning modules for sym- bolic music generation. By k eeping the base generati v e model fix ed, we trained dif fusion-based LC models (LC- Dif f) capable of controlling a range of non-dif ferentiable and continuous musical attrib utes, including contour , note density , pitch ra ng e , and rh ythm comple xity . Our em- pirical e v aluations re v eal that LC-Dif f significantly out- performs attrib ute-re gularized V AEs and cV AE-based LC methods in terms of both fidelity and controllability , with absolute impro v ements of up to 12 . 65 in Fréchet Mu- sic Distance and 43% in correlation between desired and generated attrib utes. These results highlight the poten- tial of denoising as a po werful tool for ad hoc f ader -lik e control o v er mul tiple musical attrib utes along continuous ax es, ef fecti v ely transforming a pre-trained unconditional model into a controllable music generation system depend- ing on the user’ s needs. Future w ork will focus on e x- panding the library of LC-Dif f models to include a wider range of musical attrib utes and e xploring the inte gration of user interf aces for real-time control. Future e xperiments could al so e xplore attrib ute-controlled input transforma- tions by applying forw ard dif fusi on to encoded representa- tions, rather than dra wing noise samples from the standard normal prior . Furthermore, we aim to in v esti g a te the po- tential for LC of other generati v e techniques, such as flo w matching and consistenc y models. 6. REFERENCES [1] J. Engel, M. Hof fman, and A. Roberts, “Latent con- straints: Learning to g e nerate conditionally from un- Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 57 conditional generati ve models, ” in International Con- fer ence on Learning Repr esentations , 2018. [2] M. Dinculescu, J. Engel, and A. Roberts, “MidiMe: Personalizing a MusicV AE model with user data, ” in NeurIPS W orkshop on Machine Learning for Cr eativ- ity and Design , 2019. [3] Y . T ian and J. Engel, “Latent translation: Cross- ing modalities by bridging generati ve models, ” arXiv pr eprint arXiv:1902.08261 , 2019. [4] R. Rombach, A. Blattmann, D. Lorenz, P . Esser , and B. Ommer , “High-resolution image synthesis with latent dif fusion models, ” in Pr oceedings of the IEEE/CVF confer ence on computer vision and pattern r ecognition , 2022, pp. 10 684–10 695. [5] S. W u and M. Sun, “Exploring the ef ficacy of pre- trained checkpoints in text-to-music generation task, ” in The AAAI-23 W orkshop on Cr eative AI Acr oss Modalities , 2023. [6] P . Jajoria and J. McDermott, “T e xt conditioned sym- bolic drumbeat generation using latent dif fusion mod- els, ” arXiv pr eprint arXiv:2408.02711 , 2024. [7] G. Mittal, J. Engel, C. Hawthorne, and I. Simon, “Sym- bolic music generation with dif fusion models, ” in Pr oc. of the 22nd International Society for Music Informa- tion Retrieval Confer ence (ISMIR) , 2021, pp. 468–475. [8] M. Pasini, M. Grachten, and S. Lattner , “Bass accom- paniment generation via latent dif fusion, ” in ICASSP 2024-2024 IEEE International Confer ence on Acous- tics, Speech and Signal Pr ocessing (ICASSP) , 2024, pp. 1166–1170. [9] S. Li and Y . Sung, “MelodyDif fusion: Chord- conditioned melody generation using a transformer- based dif fusion model, ” Mathematics , vol. 11, no. 8, 2023. [10] L. Min, J. Jiang, G. Xia, and J. Zhao, “Polyffusion: A dif fusion model for polyphonic score generation with internal and external controls, ” in Pr oc. of the 24th International Society for Music Information Retrieval Confer ence (ISMIR) , 2023, pp. 231–238. [11] L. Kaw ai, P . Esling, and T . Harada, “ Attributes-a ware deep music transformation. ” in Pr oc. of the 21st Inter- national Society for Music Information Retrieval Con- fer ence (ISMIR) , 2020, pp. 670–677. [12] S.-L. W u and Y .-H. Y ang, “MuseMorphose: Full-song and fine-grained piano music style transfer with one transformer V AE, ” IEEE/A CM T ransactions on A udio, Speech, and Languag e Pr ocessing , vol. 31, pp. 1953– 1967, 2023. [13] M. E. Malandro, “Composer’ s Assistant 2: Interacti ve multi-track MIDI infilling with fine-grained user con- trol, ” in Pr oc. of the 25th International Society for Mu- sic Information Retrieval Confer ence (ISMIR) , 2024, pp. 438–445. [14] G. Lample, N. Zeghidour , N. Usunier , A. Bordes, L. Denoyer , and M. Ranzato, “Fader netw orks: Ma- nipulating images by sliding attrib utes, ” in Advances in Neural Information Pr ocessing Systems , 2017, pp. 5963–5972. [15] H. H. T an and D. Herremans, “Music FaderNets: Con- trollable music generation based on high-le vel features via lo w-lev el feature modelling, ” in Pr oc. of the 21st International Society for Music Information Retrieval Confer ence (ISMIR) , 2020, pp. 109–116. [16] J. Ho, A. Jain, and P . Abbeel, “Denoising dif fusion probabilistic models, ” in Pr oc. of the 34th Interna- tional Confer ence on Neur al Information Pr ocessing Systems , 2020, pp. 1–12. [17] J. Song, C. Meng, and S. Ermon, “Denoising dif fu- sion implicit models, ” in International Confer ence on Learning Repr esentations , 2021. [18] J. Austin, D. D. Johnson, J. Ho, D. T arlow , and R. v an den Berg, “Structured denoising dif fusion mod- els in discrete state-spaces, ” in Advances in Neur al In- formation Pr ocessing Systems , 2021, pp. 1–13. [19] A. Lv , X. T an, P . Lu, W . Y e, S. Zhang, J. Bian, and R. Y an, “GETMusic: Generating an y music tracks with a unified representation and dif fusion framew ork, ” arXiv pr eprint arXiv:2305.10841 , 2023. [20] M. Plasser , S. Peter , and G. W idmer , “Discrete dif fu- sion probabilistic models for symbolic music genera- tion, ” in Pr oc. of the Thirty-Second International Joint Confer ence on Artificial Intelligence , 2023. [21] J. Zhang, G. Fazekas, and C. Saitis, “Composer style- specific symbolic music generation using vector quan- tized discrete dif fusion models, ” in 2024 IEEE 34th In- ternational W orkshop on Machine Learning for Signal Pr ocessing (MLSP) , 2024, pp. 1–6. [22] ——, “Fast dif fusion GAN model for symbolic mu- sic generation controlled by emotions, ” arXiv pr eprint arXiv:2310.14040 , 2023. [23] M. Zhang, L. J. Ferris, L. Y ue, and M. Xu, “Emotion- ally guided symbolic music generation using dif fusion models: The A GE-DM approach, ” in Pr oc. of the 6th A CM International Confer ence on Multimedia in Asia , 2024, pp. 1–5. [24] Y . Huang, A. Ghatare, Y . Liu, Z. Hu, Q. Zhang, C. S. Sastry , S. Gururani, S. Oore, and Y . Y ue, “Symbolic music generation with non-dif ferentiable rule guided dif fusion, ” arXiv pr eprint arXiv:2402.14285 , 2024. [25] A. Pati and A. Lerch, “ Attrib ute-based regularization of latent spaces for v ariational auto-encoders, ” Neural Computing and Applications , v ol. 33, no. 9, pp. 4429– 4444, 2021. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 58 [26] A. I. Mezza, M. Zanoni, and A. Sarti, “ A latent rhythm complexity model for attrib ute-controlled drum pattern generation, ” EURASIP J ournal on A udio, Speech, and Music Pr ocessing , vol. 2023, no. 1, 2023. [27] J. Retko wski, J. Ste ¸ pniak, and M. Modrzejew- ski, “Frechet music distance: A metric for gen- erati ve symbolic music e v aluation, ” arXiv pr eprint arXiv:2412.07948 , 2024. [28] N. Chen, Y . Zhang, H. Zen, R. J. W eiss, M. Norouzi, and W . Chan, “W a veGrad: Estimating gradients for wa veform generation, ” in International Confer ence on Learning Repr esentations , 2021. [29] A. V aswani, N. Shazeer , N. Parmar , J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser , and I. Polosukhin, “ Attention is all you need, ” in Advances in Neural In- formation Pr ocessing Systems , vol. 30, 2017. [30] E. Perez, F . Strub, H. de Vries, V . Dumoulin, and A. C. Courville, “FiLM: V isual reasoning with a general con- ditioning layer , ” in Pr oc. of the Thirty-Second AAAI Confer ence on Artificial Intelligence , 2018, pp. 3942– 3951. [31] J. Ho and T . Salimans, “Classifier -free diffusion guid- ance, ” in NeurIPS 2021 W orkshop on Deep Gener ative Models and Downstr eam Applications , 2021. [32] C. Raf fel, “Learning-based methods for comparing se- quences, with applications to audio-to-midi alignment and matching, ” Ph.D. dissertation, Columbia Uni ver - sity , 2016. [33] A. Roberts, J. Engel, C. Raf fel, C. Hawthorne, and D. Eck, “ A hierarchical latent v ector model for learn- ing long-term structure in music, ” in Pr oc. of the 35th International Confer ence on Mac hine Learning , 2018, pp. 4364–4373. [34] G. T oussaint, “ A mathematical analysis of African, Brazilian, and Cuban cla ve rh ythms, ” in Bridges: Mathematical Connections in Art, Music, and Science , 2002, pp. 157–168. [35] I. Higgins, L. Matthey , A. P al, C. P . Bur gess, X. Glo- rot, M. M. Botvinick, S. Mohamed, and A. Lerchner , “beta-V AE: Learning basic visual concepts with a con- strained v ariational framew ork. ” International Confer - ence on Learning Repr esentations , v ol. 3, 2017. [36] M. Heusel, H. Ramsauer , T . Unterthiner , B. Nessler , and S. Hochreiter , “GANs trained by a two time-scale update rule con v erge to a local Nash equilibrium, ” in Pr oc. of the 31st International Confer ence on Neural Information Pr ocessing Systems , 2017, pp. 6629–6640. [37] K. Kilgour , M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet audio distance: A reference-free metric for e valuating music enhancement algorithms, ” in Pr oc. Interspeec h 2019 , 2019, pp. 2350–2354. [38] S. W u, Y . W ang, R. Y uan, Z. Guo, X. T an, G. Zhang, M. Zhou, J. Chen, X. Mu, Y . Gao, Y . Dong, J. Liu, X. Li, F . Y u, and M. Sun, “CLaMP 2: Multi- modal music information retrie val across 101 lan- guages using lar ge language models, ” arXiv pr eprint arXiv:2410.13267 , 2025. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 59 RADIF CORPUS: A SYMBOLIC D A T ASET FOR NON-METRIC IRANIAN CLASSICAL MUSIC Maziar Kanani Uni versity of Gal way [email protected] Sean O’Leary TU Dublin [email protected] J ames McDermott Uni versity of Gal way [email protected] ABSTRA CT Non-metric music forms the core of the repertoire in Iranian classical music. Dastg ¯ ahi music serves as the un- derlying theoretical system for both Iranian art music and certain folk traditions. At the heart of Iranian classical mu- sic lies the radif , a foundational repertoire that organizes melodic material central to performance and pedagogy . In this study , we introduce the first digital corpus rep- resenting the complete non-metrical radif repertoire, co v- ering all 13 existing components of this repertoire. W e provide MIDI files (about 281 minutes in total) and data spreadsheets describing notes, note durations, interv als, and hierarchical structures for 228 pieces of music. W e faithfully represent the tonality including quarter -tones, and the non-metric aspect. Furthermore, we provide sup- porting basic statistics, and measures of complexity and similarity ov er the corpus. Our corpus provides a platform for computational stud- ies of Iranian classical music. Researchers might employ it in studying melodic patterns, in vestigating impro visational styles, or for other tasks in music information retrie v al, music theory , and computational (ethno)musicology . 1. INTR ODUCTION While ethnic music traditions from around the world ha ve recently gained more attention in computational research, many still lack the necessary datasets to support such stud- ies. Iran, with its rich di versity of ethnic and folk musical traditions, of fers great potential for computational analysis that reflects its regional musical identity . In this work, we take a step to ward addressing this g ap by introducing a dataset specifically focused on Iranian non-metric classical music, aiming to support and inspire future studies in this area. W e begin by introducing Iranian classical music and its core repertoire, the radif . After re- vie wing previously published datasets, we present our o wn dataset in detail. Finally , we provide a statistical and vi- sual ov ervie w of the dataset, which can serve as a useful reference for researchers and practitioners. © M. Kanani, S. O’Leary , and J. McDermott. Licensed un- der a Creati ve Commons Attribution 4.0 International License (CC BY 4.0). Attribution: M. Kanani, S. O’Leary , and J. McDermott, “Radif Corpus: A Symbolic Dataset for Non-metric Iranian Classical Music”, in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf ., Daejeon, South K orea, 2025. 1.1 Foundations of Iranian Classical Music Iranian classical music comes from a lar ger style of mu- sic called dastg ¯ ahi music. This term describes the theo- retical frame work underlying Iranian classical music and certain styles of Iranian folk music, such as bakhti ¯ ari . The core repertoire of Iranian classical music is radif (literally "order"), a structured collection of melodies transmitted across generations and foundational to performance and pedagogy . Radif is a collection of melodies org anized into a spe- cific sequence, typically di vided into 12 subcategories (tra- ditionally 13). Out of these, sev en are primary subcate- gories kno wn as dastg ¯ ah , and fi ve (respecti vely six) are secondary , referred to as ¯ av ¯ az , which can also be consid- ered as smaller dastg ¯ ah and serve as subcate gories for the primary se ven. Each of these subcategories is kno wn for its distincti ve characteristics. They are typically recognized based on their main mode (Introduced in the first g ¯ usheh ), the functional roles of their tones within that mode, and the specific sequence of g ¯ ushehs within them. The dastg ¯ ahs are: shur , se g ¯ ah , nav ¯ a , hom ¯ ay ¯ un , chah ¯ ar g ¯ ah , m ¯ ah ¯ ur , and r ¯ astpanjg ¯ ah . The ¯ av ¯ azes are: bay ¯ at-e-kor d , bay ¯ at-e-tork (also re- ferred to as bay ¯ at-e-zand ), dasht ¯ ı , ab ¯ u’at ¯ a , afsh ¯ ar ¯ ı , and bay ¯ at-e-esfah ¯ an . Among the six ¯ av ¯ azes , bay ¯ at-e-esfah ¯ an is a subcategory of the hom ¯ ay ¯ un , while the remaining are subcategories of the shur . In many accounts, r adif is considered to hav e 5 ¯ av ¯ azes , as bay ¯ at-e-kor d is often omitted. The reason is that most experts dispute the requirement of recognizing it as a independent ¯ av ¯ az . In this study , we hav e included bay ¯ at-e-kor d to ensure a complete representation. Each of these subcategories comprises pieces called g ¯ ushehs . These g ¯ ushehs can range from being as brief as a single sentence to as extensi v e as a full composition, with performances lasting se veral minutes. G ¯ ushehs can be di vided into three types: modal, melodic, and rhythmic. Modal g ¯ ushehs are played to in- troduce a mode as a small frame work for improvisation. Melodic g ¯ ushehs introduce a specific melody and its v aria- tions, where that specific melody remains fixed in dif ferent performance versions. Rhythmic g ¯ ushehs represent a spe- cific rhythm and its v ariations. The same g ¯ usheh names may appear in dif ferent dastg ¯ ahs or ¯ av ¯ azes . K er eshmeh is a rhythmic g ¯ usheh that appears multiple times in the radif , sharing the same rhythmic pattern in each case. Another 60 example is haz ¯ ın , a melody that is performed in dif ferent modes; it is classified as a melodic g ¯ usheh . Qarac heh is an example of a modal g ¯ usheh that appears in more than one dastg ¯ ah . The term ¯ av ¯ az has three meanings: 1) broadly , it refers to singing; 2) more generally , it refers to Iranian non- metric music; and 3) more specifically , it signifies the seg- ments of radif that are smaller than a dastg ¯ ah . This study focuses on the third definition, though the other two mean- ings are clarified where rele vant. Non-metric music refers to musical or ganization that lacks regular meter while potentially maintaining other temporal structures [1]. The distinction between non- metric music and free-rhythm music centers on the preser - v ation of proportional durational relationships. [2] defines free rhythm as “the rhythm of music without percei ved periodic or ganization, ” encompassing music where tem- poral or ganization serves non-rhythmic goals such as te xt transmission or melodic exposition. W e consider that non- metric music maintains relati ve proportional relationships between note durations despite lacking metrical org aniza- tion, while free-rh ythm music may abandon proportional consistency entirely . Tsuge discusses the concept of non-metric music and emphasizes its greater importance in Iranian music com- pared to other traditions [3]. He explains that the rhythmic structure of ¯ av ¯ az (second definition) music is mainly based on the poetic rhythm system, where a repeating pattern of dif ferent number and size of syllables shapes its rhythm. This structure is closely connected to the nature of the Per - sian (Farsi) language and its classical poetry system which plays a significant role in ho w the melody is formed and percei ved. Kanani and Azadehfar [4] described the ke y ¯ av ¯ az (second definition) patterns commonly found in non- metric traditional Iranian v ocal music. The exact origins of the r adif system in Iranian music are not clearly defined. Some sources, like Bruno Nettl, be- lie ve it originated in the 17th century , while others suggest the 18th century as the starting point [5–7]. What is clear , ho wev er , is that radif de veloped from the late Saf avid era (1670s-1730s) through to the mid-Q ¯ aj ¯ ar period (1850s). The lack of precise dating can be linked to the oral tradi- tion of this music and the absence of recording technology at the time. It is belie ved that r adif was created to support the teach- ing of musical modes and to enhance skills in improvi- sation and modulation within Iranian art music [8]. The same radif can be interpreted dif ferently by dif ferent mu- sicians, and once a student becomes a master , they are able to de velop their o wn version of the r adif . Over time, many prominent music masters ha ve created their o wn interpre- tations, leading to dif ferent versions of radif . These musi- cians de veloped their r adif based on their personal under- standing, experience, and e xpression of Iranian modes, ei- ther for their o wn performances or to teach younger learn- ers. T raditionally , radif was passed do wn orally from mas- ter to student, preserving its legac y and technical details through generations. The version of r adif curated by musician and educator M ¯ ırz ¯ a ’Abdoll ¯ ah has become the most widely used choice in pri vate lessons, uni versity music education and conser - v atories ov er the past century . Initially , it was mainly asso- ciated with the t ¯ ar and set ¯ ar instruments, but today , it has been adapted and performed on many k ey Iranian instru- ments, including kamancheh , sant ¯ ur , ne y , q ¯ an ¯ un , o ¯ ud , and qe ychak . 1.2 Exploring Datasets In recent years, the creation and sharing of digital music corpora has gained significant attention among researchers in areas such as music information retrie val, computa- tional musicology , and natural language processing. V ar - ious studies ha ve demonstrated that well-curated datasets can facilitate analysis of both symbolic and audio musi- cal features, thereby promoting new insights into musical traditions [9, 10], styles, and technologies. Many music information retrie v al (MIR) corpora are primarily audio-based, with annotations for pitch, timing, structural information, etc., e.g. [11, 12], while others are symbolic / score-based. Of the latter , some are based on automated reading of paper scores [13]. Ours differs in that we ha ve manually written the digital score. Considering other musical traditions in the geographi- cal region, there is no symbolic corpus a vailable for Ara- bic Maq ¯ am music, whereas a symbolic corpus does exist for T urkish Makams, known as SymbT r [14]. There are some audio datasets related to these musi- cal traditions, such as the Dunya corpus, which includes T urkish Makam [15], Carnatic, Hindustani [16], Beijing Opera [17], and Arab-Andalusian music [18]. The Dunya corpus is part of a lar ger project called CompMusic [19]. T o the best of our knowledge, there was no symbolic corpus a vailable for Iranian music before our pre vious work, in which we introduced the Shour Corpus [20]. This corpus includes one section of the radif ( shur ) and was used in our study on discov ering patterns and producing meaningful v ariations through grammatical representation (compression) in this musical style. KUG Dastg ¯ ahi [21] and [22] are two audio datasets for Iranian music. Na v a [23] is an audio dataset designed for Iranian instrument recognition, while Ar-MGC is a dataset for Arabic music genre classification [24]. 1.3 Our Contribution At the time of writing this paper , to the best of our kno wl- edge, there is no symbolic dataset cov ering the entire radif . This led us to create the Radif Corpus, which includes all non-metric pieces from M ¯ ırz ¯ a ’Abdoll ¯ ah’ s radif . Out of the se veral transcriptions of M ¯ ırz ¯ a ’Abdoll ¯ ah’ s radif , we ha ve selected the edition titled “Radif Analysis - based on the notation of M ¯ ırz ¯ a ’Abdoll ¯ ah’ s radif with annotated visual description” by Dariush T alai [25]. This edition consists of a recorded performance, together with a score deriv ed from the performance, notated with hierarchical structure. Although radif also includes some metric pieces, which are usually performed at the end of each dastg ¯ ah / ¯ av ¯ az , our Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 61 corpus excludes them as we are focusing on non-metric music. Figure 1 presents an example of a transcription from the book. Figur e 1 . T ranscription sample from Dariush T alai’ s “Radif Analysis," illustrating the second g ¯ usheh structure in the shur dastg ¯ ah . The three white boxes indicate the main structural di visions, while the last white box contains a further sub-di vision marked in gray . In [19], Serra identified fiv e critical criteria for corpora in the CompMusic project: Purpose, Cov erage, Complete- ness, Quality , and Reusability . These are criteria that we also considered in our work. Our purpose has already been stated; Cov erage and Completeness are achiev ed by in- cluding an entire radif collection. Regarding Quality , we belie ve the transcription is accurate, as it has been double- checked by one of the authors who is an e xpert in this mu- sical style. Reusability is addressed by providing open data in a documented format. This corpus is a resource for MIR, computational musi- cology , and ethnomusicology , enabling applications such as melodic pattern recognition, automatic transcription, and mode classification. It enables symbolic music gen- eration and AI-assisted improvisation. The dataset also fa- cilitates cross-cultural music studies, allo wing for compar- ati ve analysis with T urkish Makams and Arabic Maqams, as well as phrase-le vel e xaminations of melodic progres- sion. Additionally , its structured format aids compu- tational analysis of non-metric rhythm and hierarchical structure understanding, making it a foundational dataset for exploring Iranian classical music in both traditional and computational domains. 2. RADIF CORPUS DESCRIPTION Our corpus includes all these dastg ¯ ahs / ¯ av ¯ azes , featuring 228 non-metric g ¯ ushehs . The MIDI files contain a total of 43,441 notes, with a total playback duration of approxi- mately 16,825 seconds (about 281 minutes). The dataset represents each musical piece as a sequence of notes, where for each note we store microtonal pitch, du- ration, pitch (quarter tones), interval, MIDI pitch number and MIDI bend. Data formats are csv files and MIDI files, which ha ve been manually transcribed from the book. Additionally , we pro vide MusicXML files con verted from the CSV data. These XML files preserve the mi- crotonal pitch information using fractional <alter> v al- ues follo wing MusicXML 4.0 standards. For non-metric rhythm representation, we use fle xible time signatures that accommodate the total duration of each piece. Howe v er , we note that some current music notation software imple- mentations sho w limitations in both microtonal playback and non-metric representation. W e observ ed that quarter- tones are not played back correctly , and the software tends to generate complex time signatures (e.g., 342/8) as a workaround for representing non-metric music, which, while functional, may not pro vide an aesthetically ideal notation display . The MusicXML con version script is in- cluded in the repository for researchers who wish to ex- periment with dif ferent notation software or contrib ute to improving microtonal MusicXML rendering capabilities. These files don’ t represent hierarchical structures. The dataset includes se veral figures that are e xplored further in the continuation of this paper . Our digital ver - sion exactly mimics the paper source, while to simplify the dataset and a void additional comple xity , grace notes or or- naments are not included in the dataset. The accurate representation of Iranian classical music in v olves dealing with two main issues: non-metric rhythm and micro-tonal pitch. In the following subsections, we describe our methods to address these challenges. T onality . Notes are symbolized by C, D , E , F , G, A, B , with accidental signs including flat ( Z ), K or on ( k ), Sori ( s ), and sharp ( \ ). Here, “K oron” and “Sori” indicate micro-tonal adjustments specific to Iranian music - quarter tones lo wer and higher , respecti vely . Chr omatic Scale. Although these interv als suggest a 24-quarter -tone chromatic scale per octa ve, which can be seen in some contemporary compositions, Iranian instru- ments traditionally employ only 18 specific notes: C, D Z , D k , D, E Z , E k , E, F , Fs, F \ , G k , G, A Z , A k , A, B Z , B k , B, with corresponding quarter -tone interv als: 2, 1, 1, 2, 1, 1, 2, 1, 1, 1, 2, 1, 1, 2, 1, 1. MIDI. Pitch bend is a commonly used method for rep- resenting microtones in MIDI files in microtonal music styles. T o encode K oron, we assign it a MIDI note number one semitone lo wer than the natural note and a pitch bend, i.e. increase of 2048 (a quarter tone); for Sori, the MIDI note number is the same as the natural, with a pitch bend increase of 2048. Octa ves. W e consider the lowest note in the first g ¯ usheh Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 62 of each dastg ¯ ah / ¯ av ¯ az as the start note of the main octa ve. For e xample in the first g ¯ usheh of our corpus the lo west note is F3 so the main octa ve is from F3 to F4. The mu- sic in the main octa ve is represented only by its symbols and accidental signs (if any are present). F or notes in oc- ta ves other than the main one, we use ‘+’ or ‘-’ follo wed by a number to indicate the number of octa ves abov e or be- lo w the main octav e. F or example, an A k from one octav e higher than the main octa ve would be written as A k +1. Interv als. The Intervals column represents the pitch dif ference between consecutiv e notes, with "1" indicating a quarter -tone step. Durations. The term non-metric does not mean the same as free rhythm. In this musical style, notes are related to each other through proportional duration dif ferences, with some being longer or shorter than others. These re- lationships create the rhythmic structure, which is mostly fixed and not altered by the performer . While the exact durations are not strictly defined, they can be cate gorized into four main types: very short, short, long, and very long. These can be said to correspond to sixteenth, eighth, quar- ter , and half notes [25] and are numerically represented as 1, 2, 4, and 8 in the corpus, where 1 rhythmic unit is equiv- alent to 1 sixteenth note. Gr eater Hierarchical Structur es. W e also docu- mented the hierarchical structure of each piece, as provided in the original printed source (see Figure 1). In our nota- tion, brackets represent hierarchical relationships, forming a tree structure. An open-bracket “[” in the datasheet marks the beginning of a tree node, with following notes repre- senting the contents of a section or subsection until the matching close-bracket “]”. Each tune is enclosed within brackets, representing the root node. Additional pairs of brackets define child nodes, which can themselv es contain further subsections, forming a nested hierarchy . For e xample, in Figure 1, the abstract hierarchical struc- ture can be represented as [[][][[]]] . The outer brackets enclose the entire tune. The second pair of brackets defines the first section, co vering the first three lines in Figure 1. The third pair corresponds to the sec- tion spanning lines three to six. The next open bracket is follo wed by another open bracket, indicating the presence of a subsection, which corresponds to lines eight and nine. The subsection is highlighted in the last line. 3. ST A TISTICAL AND VISU AL O VER VIEW In the dataset, for each g ¯ usheh , we provide both a pitch histogram and an interv al histogram. Figure 2 provides an example of the interv al histogram for a g ¯ usheh . Each dastg ¯ ah or ¯ av ¯ az comes with a spreadsheet gi ving information about its g ¯ ushehs , like the number of notes and total duration. T able 1 presents the number of g ¯ ushehs in each dastg ¯ ah or ¯ av ¯ az , along with the number of notes, du- ration in both units and seconds, and pitch range. M ¯ ah ¯ ur has the highest number of g ¯ ushehs with 34, fol- lo wed by chah ¯ ar g ¯ ah with 31 and shur with 29. It also contains the lar gest number of notes, with 6104 in m ¯ ah ¯ ur , 5788 in chah ¯ ar g ¯ ah , and 4830 in shur . These three also 4 3 2 1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 Interval 0 10 20 30 40 50 F r equency 07 - Gabri Figur e 2 . Interval histogram of Gabri , a g ¯ usheh in ab ¯ u’at ¯ a ha ve the longest durations, with v alues of 12480, 11606, and 9839 rhythmic units, respecti vely . Afsh ¯ ar ¯ ı has the fe west g ¯ ushehs with 4, follo wed by bay ¯ at-e-esfah ¯ an with 5, and dasht ¯ ı and bay ¯ at-e-kor d with 6 each. The shortest dastg ¯ ah / ¯ av ¯ az in terms of duration is dasht ¯ ı , with 1308 notes and a duration of 2685 units, fol- lo wed by bay ¯ at-e-kor d with 1360 notes and a duration of 2432 units. afsh ¯ ar ¯ ı is the third shortest, with 1520 notes and a duration of 2681 units. Reg arding individual g ¯ ushehs , ker eshmeh in se g ¯ ah is the shortest g ¯ usheh in the corpus, while bay ¯ at-e r ¯ aje’ va for ¯ ud in bay ¯ at-e-esfah ¯ an is the longest. dastg ¯ ah / ¯ av ¯ az g ¯ usheh Count Number of Notes T otal Duration (unit) MIDI Perfor- mance Dura- tion (second) Pitch Range Shur 29 4830 9839 1966 [F , A Z +2] Bay ¯ at-e-kor d 6 1360 2432 486 [G-1, A Z +1] Dasht ¯ ı 6 1308 2685 536 [F-2, G+1] Bay ¯ at-e-tork 16 2544 4986 996 [F-1, G+1] Abuata 7 2194 3959 791 [F , A Z +1] Afsh ¯ ar ¯ ı 4 1520 2681 536 [F-1, A Z +1] Se g ¯ ah 20 3283 6559 1310 [F , F+2] Nav ¯ a 19 3252 5975 1194 [D-1, C+1] Hom ¯ ay ¯ un 27 5323 9780 1954 [D, F+2] Bay ¯ at-e-esfah ¯ an 5 1669 3264 652 [D-1, F \ +1] Chah ¯ ar g ¯ ah 31 5788 11606 2319 [C-1, G+2] M ¯ ah ¯ ur 34 6104 12480 2493 [C-1, G+2] R ¯ astpanjg ¯ ah 24 4266 7958 1590 [D-1, C+2] T able 1 . Summary of g ¯ usheh information for each dastg ¯ ah and ¯ av ¯ az , including the number of g ¯ ushehs , total notes, du- ration in units and seconds and pitch range. 3.1 Melodic Progr ession One of the main objecti ves when a musician performs a complete concatenated dastg ¯ ah is to follo w se yr , or melodic mov ement [26]. In traditional Iranian music, se yr refers to the progression of melodies within a piece, shap- ing the ov erall pitch direction of the music through its in- troduction, de velopment, climax, and resolution. Se yr underlines the importance of transitional notes and melodic phrases in establishing the identity and modal character of the piece. These elements play a crucial role in guiding the melodic flo w from one section to another , ensuring a coherent and expressi v e musical journey . Each dastg ¯ ah / ¯ av ¯ az folder includes a pitch contour plot to sho w its se yr . The pitch contour plot illustrates how the melody and pitch e volv e across different g ¯ ushehs . Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 63 sponding to the consonant-v o wel (CV) transition. The resulting alignments are check ed and manually cor - rected for the occasional errors that arise mainly due to the presence of long v o wels in singing, poor enunciation at times, as well as the occurrence of significant pitch inflec- tions. W e observ e the manually corrected onset locations- i.e. the syllable onsets as realised by the artist- mark ed at the top of t he w a v eform in the e xample of one full tala c ycle of the bandish in Figure 2. T o obtain the beat locations, we annotate the tabla strok e onsets using the source-separated accompaniment file; we manually mark the salient beats (do wnbeat ’x’ sam and 9 th beat ’o’ khali ) in 1 c ycle. Assuming a consistent local tempo, we di vide each half c ycle (interv al between one do wnbeat and the ad- jacent t th beat into 8 equal parts, thereby obtaining all the estimated beat instants across the rendition. Ne xt, we mark the canonical locations of the match- ing syllables, positioning the syllables according to the Bhatkhande notation (as in Figure 1). The position map- ping of the realised syllable onsets to the corresponding canonical syllable w as implemented follo wing the method proposed in pre vious w ork [8]. W e observ e from Figure 2 ho w t h e realised onsets lag the canonical locations most of the time. Figur e 2 . Pitch contour (bottom) and sung syllable align- ment (red) with canonical beat positions of the same sylla- ble (black) for an e xcerpt of J a J a Re by ABD. Finally the v oc al pitch is e xtracted at 10 ms interv als us- ing an autocorrelation based method for fundamental fre- quenc y and v oici ng [12]. Brief pauses and un v oiced re- gions are linearly interpolated to obtain a continuous pitch contour for each sung syllable re gion. The pitch contour is con v erted to cents by normalisation with the kno wn perfor - mance’ s tonic. Ev entually , we obtain for each performance in our dataset, the se gmented audio of each sung line an- notated at the syllable le v el with syllable name, boundaries and the pitch (cents) at 10 ms interv als. W e use these lo w- le v el features to define quantities that capture the singing v ariations across repetitions of a bandish line within and across singers. The reference for the comparison is the syllable identity (i.e. its name and metrical location) as defined in the canonical notation as presented in Figure 1. Rag a Bhimpalasi Y aman Bandish Ja Ja Re Y eri Aali T ala T eentaal T eentaal Sw ar S, R, g, m, P , D, n, S S, R, G, M, P , D, N, S # Concerts 15 13 # Artists 15 12 # Repetitions (L1, L2, L3, L4) 167, 39, 47, 47 94, 32, 35, 23 Matra per min range 138–200 111–203 T able 1 . Summary of our dataset of concert recordings across r a gas , bandish , and artists. Swar notation details are in the supplementary . 4. MEASURING EXPRESSIVENESS W e wish to quantify and compare the v ariability observ ed in the acoustic realisation of a gi v en syllable, from a spe- cific line of the bandish , across (i) repeated utterances within an artist’ s performance, and (ii) utterances of the same syllable across dif ferent performances/artists. The acoustic parameters that we detect are: (i) the onset time of the syllable, (ii) the syllable duration (as the time inter - v al between the current syllable’ s onset and either the on- set of the ne xt syllable or the start of the follo wing silence se gment, whiche v er occurs first. and (iii) the pitch contour shape across the syllable interv al. W e illustrate the process by pro viding e xamples of the process ing and analyses of the audio rendering of a chosen line by one artist. 4.1 T iming expr ession The de viation of the detected onset from its reference as- signed beat inde x in the canonical notation gi v es us an estimate of the lag/lead of the singer for the syllable in question. W e represent the de viation in terms of fraction of the local beat interv al; this normalization f acilitates the comparison across instances and concerts. W e can vie w the thus measured timing of fsets as e vidence of e xpressi v e timing, especially if this quantity sho ws v ariability across repetitions of the syllable within the concert. Figur e 3 . De viation of the sung syllable onsets from the canonical locations measured in the units of beat duration for J a J a Re Line 1 by ABD for multiple repetitions of the line. Figure 3 captures the onsets of syllables in bandish 1, Line 1 as rendered by singer ABD. The syllable names Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 70 are sho wn with their canonical matr a locations at the bot- tom. W e note that some syllables occup y 2 beats and oth- ers 1 beat in the canonical form. W e observ e, for e xam- ple, that "Jaa2" and "Man" (both 2-beat syllables) sho w v ariations with a mean lag of about one beat. "Man" in- stances ho we v er are much more dispersed. While a uni- form of fset could potentially indicate a structural dif fer - ence between the artist’ s v ersion of the bandish and that of the Bhatkhande book, a high standard de viation (lik e in "Man") points to the artist’ s in-the-moment e xpressi v e v ariations. On the other hand, the 3 syllables preceding "Man" sho w near -zero of fsets across repetitions. 4.2 Pitch expr ession Analysing the rendered song for pitch-based e xpressi v e- ness is a rather in v olv ed task. There are man y w ays in which the artist injects e xpressi v enes s in their performance via pi tch v ariation. As we w ant to analyse the dif ferent w ays of realising the same line of a bandish , we are in- terested in the v ariability in the pitch contour (PC) shape of a syllable across the multiple repetitions of that line. A greater di v ersity of PC shapes can then be interpreted as higher e xpressi v eness in an artist’ s repetitions of the line. Figur e 4 . PC quantisation to the nearest r a ga note for one instance of J a J a Re bandish line 1 rendered by artist ABD F or each syllable, the associated PC spans the duration of that syllable; hence, this is not a fix ed-length time series. W e represent this v ariable dimension PC for a gi v en sylla- ble with a lo wer and fix ed-dimensional v ector . W e first im- plement a piece-wise aggre g ate approximation (P AA) o v er the syllable PCs. P AA is a time-series representation that has been used widely in data-mining tasks [13]. This is implemented as follo ws. The syllable PC v alues are each first quantised to the nearest r a ga swar (note), Figure 4 sho ws this process. Ne xt, the quantised PC v alues for a syllable across repe- titions, is aggre g ated by di viding the quantised PC into a fix ed number of uniform interv als and assigning the mode of the v alues to each interv al. The number of fix ed in- terv als is set empirically to 10 interv als per beats allot- ted to the syllable (treating the syllable e xtensions indi- cated by ’-’ in Figure 1 to be a part of the pre vious syl- lable, thereby adding to its allotted beats). The choice of 10 equal se gments per beat interv al is based on the tempo range of our dataset (110-200 BPM or 300 ms to 545 ms Figur e 5 . Three distinct renditions of the syllable "Jaa1" by artist ABD, each represented by a fix ed number of uni- form time interv als. Each interv al is mapped based on its modal pitch to the nearest r a ga note, gi ving us the P AA string representation for the s yllable’ s pitch shape. Fig- ure 6 describes this process and the P AA strings for the abo v e PCs. per beat) and sampling period (10 ms/sample) of the pitch contour . This results in se gments, each represented by a short sequence of samples of the pitch contour . This bal- ance allo ws capturing dynamic pitch fluctuations which are pre v alent in Hindustani classical music (HCM), with- out o v er -quantisation. No w , each P AA interv al within a syllable is assigned a discrete symbol, where the symbols are dra wn from a suitable alphabet which comprises notes ( swar ) across the rele v ant octa v e ranges. The resulting string of note v al ues (one per P AA interv al) then represents coarsely the realised pitch shape of the syllable. Figure 6 sho ws this process . The 3 string sequences in Figure 6 cor - respond to the 3 PCs in Figure 5. This type of aggre g ation presents a tradeof f of generality vs specificity . Lik e in the case of syllable timing, we are interested in the v ariation, if an y , in pitch shape of a gi v en syllable across repetitions . W e achie v e this by computing the sim- ilarity of the syllable PCs for pairs dra wn from the set of repetitions in a single concert. The Le v enshtein edit dis- tance [14] between the P AA strings pro vides us with the number of note substitutions. W e e v aluate the Normalised Le v enshtein Substitution Score (NLSS) for each pair as a measure of the dissimilarity . Figure 7 sho ws a matrix rep- resentation (heat map) of NLSS v alues for a chosen sylla- ble as rendered by one artist across 14 repetitions. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 71 Figur e 6 . Processing pipeline for generating P AA string representations for the pitch contour for dif ferent repeti- tions of a syllable. The strings correspond to the PCs of the syllable repetitions gi v en in Figure 5 spanning 2 beats. Figur e 7 . Heat map sho wing NLSS for each pair of "Jaa1" syllable PCs dra wn from the set of repeti tions of line 1 of bandish J a J a Re rendered by artist ABD 5. OBSER V A TIONS AND DISCUSSION W e are interested in the within-artist v ariation for a gi v en bandish line and its syllables. This can f acilitate poten- tially v aluable insights about (i) the preferred locations (in terms of chosen syllables) for e xpressi v e gestures by indi- vidual artists and across artists and (ii) the e xtent and na- ture of e xpressi v e gestures for a gi v en artist. Due to space limitations, we present the analysis results for line 1 of one bandish , with the other bandish presented in the supple- mentary . Figure 8 presents for each artist and syllable, the stan- dard de viation (s.d.) of the timing de viati on (as captured in the e xample for artist ABD in Figure 3). The s.d. helps us focus on the variability of of fsets rather than on ac- tual of fset v alues (which might b e attrib uted to structural dif ferences between the artist’ s v ersion and Bhatkhande’ s v ersion of the bandish , rather than e xpression-related). In Figure 8, we note the dominance of the first 3 syl- lables for most artists. The full range of beha viours, ho w- e v er , includes IN at one end with minimal v ariations to DG and RK, who introduce ne w v ariations on practically all syllables. That IN does not e x ercise an y fle xibility is not Figur e 8 . Standard de viation of the distrib ution of the frac- tional timing de viation for e v ery syllable o v er multiple rep- etitions in one rendition, across artists. Figur e 9 . Mean NLSS o v er all pairs of repetitions of each syllable in line 1 of bandish J a J a Re , computed per artist. serving as a dissi milarity measure across dif ferent repeti- tions of a syllable by an artist. surprising gi v en that his performances were e xplicitly cre- ated to closely follo w the prescribed Bhatkhande notation, as discussed here [15]. RK, on the other hand, is consid- ered a virtuoso musician. Aggre g ating across the ro ws, we obtain the per sylla- ble beha viour across artists in Figure 10. The mean v alues sho w that the first 3 syllables carry the most e xpressi v e tim- ing, with "Jaa2" also sho wing the most spread across artists (consistent with Figure 8). W e see, for e xample, that PT sho ws a lar ge range in per syllable SD, ag ain agreeing with Figure 8. In the case of the artist ABD, we can observ e that the temporal de viation is spread relati v ely e v enly , while peaking for a particular syllable "Man", which is the do wn- beat. This can be easily appreciated in listening to the au- dio, which can be accessed in the supplementary material. T o assess pitch v ariability , we calculate the a v erage number of pitch substitutions by taking the mean of the NLSS for al l pairs of repetitions per artist and per syllable, (normalised by the string length), calculated across all pos- sible pairs of the gi v en syllable utterances within a concert. Figure 9 sho ws the v alues per artist and per syllable of Line 1. W e can observ e in Figure 8 and Figure 9 that both the pitch and temporal v ariation across repetitions and across dif ferent artists is more prominent at the be ginning of the line. Near the end of the line, the pitch v ariation increases, which can be att rib uted to the emphasis on a semantically Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 72 Figur e 10 . Box-plot of s.d. of timing de viation param- eter per syllable of J a J a Re Line 1 as aggre g ated across concerts and artists. The tala -c ycle ends at the syllable "Ne", and a ne w c ycle starts from do wnbeat at the syllable "Man". Figur e 11 . Box-plot of the mean of the a v eraged NLSS per syllable of J a J a Re Line 1 as aggre g ated across concerts and artists. important w ord- "Mandira v a". The syllables "P a" and "Ne" are near the tala -c ycle boundary , and we observ e a minimum pitch and temporal v ariation; this is consistent with the observ ations of [8]. As the performer heads to w ards the tala -c ycle boundary , their o v erall tendenc y of v ariation and impro visation decreases, aiming rather to w ards resolving the melodic and temporal e xpression the y ha v e come up with in the particular tala - c ycle. Aggre g ating across ro ws of Figure 9, we obtain the per syllable beha viour o v er all artists in Figure 11. The dotted line joining the means indica tes that the pitch v ariation de- creases as one approaches the tala -c ycle boundary at the syllable "Ne", while the the pitch e xpressi v eness is higher in the start and end of the line, which f alls on the 7th and 5th beats of the tala -c ycle respecti v ely , which is f ar from the c ycle boundary , hence pro viding more scope for timing and pitch e xpressi v eness. W e see, for e xample, that A C and AK sho w high pitch v ariation across the repetitions of the same line. Hierarchical clustering of all the pairs from the set of repetitions of a syllable by an artist pro vides us with infor - mation about the v ariation clusters. A threshold can be de- fined t hat decides if a v ariation belongs to a cluster or not. More di v erse v ariations w ould indicate more number of clusters, indicating higher e xpressi v eness. Dendrograms are e xcel lent for visualising such clusters. The realisation of the pre vious syllable has an influence on which cluster the follo wing syllable v ariation w ould belong to. Comparing Figure 8 and Figure 9, we note that e xpres- si v e gestures that utilise pitch are not necessarily at the same locations that e xhibit timing de viation in terms of the preferred syllable. It is rather interesting to look at the least amount of e xpressi v eness in both pitch and timing lie with the syllables that are near the tala -c ycle boundary . An analysis at the indi vidual audio le v el w ould pro vide a more accurate picture of the correlations, if an y , and is left to future w ork. Our observ ations, r eported here on the Line 1 of one bandish , lar gely hol d with the second bandish . The under - lying reasons for the choice of specific syllables e xhibiting lar ger v ariability are similar to those discus sed by Mor - ris [4]. These include lar ger v ariations at line or phrase ending syllables due to the ef fect of pre vious and ne xt con- te xts, and the choice of syllables belonging to more emo- tionally loaded w ords in the lyrics. 6. CONCLUSION In this paper , we articulated the problem of modeling e x- pressi v e v ariations in the conte xt of performance of Hin- dustani traditional compositions by established artists of the genre. A well-kno wn bandish in the chosen r a ga is al w ays sung at the be ginning of a concert with multiple repetitions of the lines, mark ed by v ariations in the loca- tion and type of the e xpressi v e gestures. Based on our proposed methodology , we sho wed that it is possible to arri v e at systematic patterns across artists by treating the syllables of the lyrics as reference points for a study of the range of v ari ation. This also helped us discuss interesting correlations between the roles of melody and rh ythm in e x- pressi v eness. W e presented a dataset that w as annotated with a combi- nation of manual and automatic tools to obtain a rich repos- itory of distinct realizations of the lines of tw o popular traditional compositions. While much further e xploration remains possible, this w ork demonstrates the potential of computational models for impro visation in the conte xt of compositions in the Khayal genre. This w ork lays the foundation for generati v e applications by capturing high- le v el performance features that reflect an artist’ s distincti v e style. A preliminary e xperiment w as pe rformed to generate the temporal de viations discussed abo v e, where the distri- b utions formed by all the de viations of a syllable from its canonical location for an artist were used to sample out ne w points for each syll able. A sine-tone based audio w as synthesized from the generated pitch contour . The refer - ence (Bhatkhande canonical form) and generated tracks are a v ailable in the supplementary . This lets us create infinite possibilities for rendering the same line, while at the same time capturing some hint of the artist’ s style. Extending this approach to other acoustic dimensions for e xpression such as timbre and dynamics can enable the generation of classical music that embodies the unique identit y of indi- vidual artists. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 73 7. A CKNO WLEDGMENTS W e take this opportunity to ackno wledge and express our gratitude to all those who supported and guided us during this research work. W e are thankful to Madhumitha S., whose thesis work w as a crucial base for our project. W e also thank Mr . Himanshu Sati, and Mrs. Hemala Ranade for their scholarly musicological insights, which ga ve di- rection to our work. 8. REFERENCES [1] B. W ade, “Music in India: The classical traditions, ” Manohar Press, 2001. [2] C. E. Cancino-Chacón, M. Grachten, W . Goebl, and G. W idmer , “Computational models of expressi ve mu- sic performance: A comprehensiv e and critical re- vie w , ” F r ontier s in Digital Humanities , vol. 5, p. 25, 2018. [3] P . N. Juslin, “Cue utilization in communication of emo- tion in music performance: Relating performance to perception. ” J ournal of Experimental Psycholo gy: Hu- man per ception and performance , vol. 26, no. 6, p. 1797, 2000. [4] A. Morris, T ransmission and performance of Khayal compositions in the Gwalior gharana of Indian vocal music. PhD thesis, Uni v . Of London, S.O.A.S., 2004. [5] W . V an der Meer , “ Audience response and e xpressi ve pitch inflections in a li ve recording of le gendary singer kesar bai k erkar , ” in Expr essiveness in music perfor - mance: Empirical appr oac hes acr oss styles and cul- tur es , D. Fabian, R. T immers, and E. Schubert, Eds. Oxford Uni versity Press (UK), 2014, pp. 170–184. [6] S. Sankaran, P . V . K. Sekhar , and A. M. Hema, “ Au- tomatic segmentation of composition in carnatic music using time-frequency cfcc templates, ” in Pr oceedings of 11th International Symposium on Computer Music Multidisciplinary Resear c h (CMMR) , 2015. [7] K. K. Ganguli and P . Rao, “ A study of variability in raga motifs in performance conte xts, ” Journal of Ne w Music Resear c h , vol. 50, no. 1, pp. 102–116, 2021. [8] Y . Bhake and P . Rao, “Expressi ve timing in Hindus- tani v ocal music, ” in Pr oc. of ICASSP 2025 W orkshop on Indian Music Analysis and Generative Applications (WIMA GA) , Hyderabad, India, 2025, accessed at:link. [9] S. Rao and P . Rao, “ An o vervie w of Hindustani music in the context of computational musicology . ” J ournal of New Music Resear ch , v ol. 43, no. 1, 2014. [10] V . Bhatkhande, Kramik Pustaka Malika . Sangeet Karyalaya Hathras, India, 2013. [11] D. Pov ey , A. Ghoshal, G. Boulianne, L. Bur get, O. Glembek, N. Goel, M. Hannemann, P . Motlicek, Y . Qian, P . Schw arz et al. , “The kaldi speech recog- nition toolkit, ” in IEEE 2011 workshop on automatic speech r ecognition and understanding . IEEE Signal Processing Society , 2011. [12] Y . Jadoul, B. Thompson, and B. De Boer , “Introducing parselmouth: A python interface to praat, ” Journal of Phonetics , v ol. 71, pp. 1–15, 2018. [13] J. Lin, E. Keogh, L. W ei, and S. Lonardi, “Experienc- ing SAX: a nov el symbolic representation of time se- ries. ” Data Mining and knowledg e discovery , vol. 31, no. B, pp. 107–144, April 2007. [14] W . J. Heeringa, “Measuring dialect pronunciation dif- ferences using Le venshtein distance, ” 2004. [15] “2000 classic compositions from Bhatkhande on cd, ” https://scroll.in/article/726180/, accessed: 2024-04-10. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 74 COLORING MUSIC: BRIDGING MUSIC AND COLOR P ALETTES FOR GRAPHIC DESIGN T akayuki Nakatsuka Masahiro Hamasaki Masataka Goto National Institute of Adv anced Industrial Science and T echnology (AIST), Japan {takayuki.nakatsuka, masahiro.hamasaki, m.goto}@aist.go.jp ABSTRA CT This paper explores the relationship between music and the color palettes used for designing their corresponding mu- sic cov er images, providing a comprehensi ve analysis that bridges auditory and visual expression. Our findings re veal a relationship between musical pieces and certain colors, suggesting that the color palettes used in cov er image de- sign are carefully selected to reflect the auditory e xperience. Building on these findings, we propose a framew ork that estimates appropriate color palettes for musical pieces to support selecting colors for cov er image design. Using a lar ge priv ate dataset of 582,894 pairs of a musical piece and its corresponding co ver image from v arious music gen- res, our frame work le verages deep learning techniques to train our color palette estimator . W e demonstrate the effec- ti veness of our proposed frame work in graphic design by sho wcasing an application that generates cov er images us- ing the estimated color palettes from gi ven musical pieces. 1. INTR ODUCTION In multimodal music understanding, both music and their corresponding music cov er images play a crucial role. For instance, Oramas et al. successfully improv ed music genre classification accuracy by incorporating image features in addition to audio features [1]. In addition, L ¯ ıbeks and T urnbull sho wed that cov er images in v olve distinct features that can be used to predict music genre tags [2]. These studies suggested that a cov er image embodies the essence of its corresponding music content, thereby establishing that analyzing these images yields a deeper understanding of the music. This study focuses on the colors used in cov er images and analyzes their relationship with music. The colors used in cov er images tend to empirically re- flect the characteristics of the corresponding music style. As illustrated in Fig. 1, dif ferent music genres display distinc- ti ve characteristics in the colors used in the co ver images. As colors are closely linked to cultural conte xts [3], emo- tions [4, 5], and the ability to attract visual attention [6], cov er images contribute to the promotion of music content © T . Nakatsuka, M. Hamasaki, and M. Goto. Licensed under a Creati ve Commons Attribution 4.0 International License (CC BY 4.0). Attribution: T . Nakatsuka, M. Hamasaki, and M. Goto, “Coloring Music: Bridging Music and Color Palettes for Graphic Design”, in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf., Daejeon, South K orea, 2025. Death m etal Country Ele ctronic Figur e 1 . Example results of Google Search with the text queries “{music genre} music alb um covers, ” where we used the music genres ‘Death metal, ’ ‘Country , ’ and ‘Elec- tronic. ’ Music cov er images for each genre are characterized by the colors used in co ver image design: dark colors for ‘Death metal, ’ bro wnish colors for ‘Country , ’ and vivid col- ors for ‘Electronic. ’ and enhance the ov erall music appreciation experience [7]. Therefore, this relationship between music and the colors used in cov er images has been the subject of se veral stud- ies [8 – 10]. Howe v er , these studies hav e mainly focused on genres, not on musical pieces. This paper first in v estigates the preferred colors for de- signing cov er images across multiple genres in our prelimi- nary study (Section 4) and further explores the relationship between musical pieces and the colors used in their corre- sponding cov er images based on our proposed frame work (Section 5). In this study , we focus on not only a repre- sentati ve color b ut also color palettes used in cov er images because they play a crucial role in graphic design [11 – 13], shedding light on the deliberate selection process of colors that reflect the essence of the music content. Based on our findings that a relationship exists between musical pieces and the colors used in their corresponding cov er images, we propose a framew ork to estimate appro- priate color palettes for musical pieces. The ke y technical 75 aspects of our frame work are ho w to extract color palettes from cov er images and ho w to estimate color palettes for musical pieces. For a color palette e xtraction method, we employ data-driven color manifolds [14], which are use- ful in arranging the colors as a color palette. F or a color palette estimator , we train a deep neural network to esti- mate an appropriate color palette for each musical piece. In this training, we lev erage a pretrained audio model ( con- trastive langua ge-audio pr etrai ning (CLAP) [15] or A u- dioT oken [16]) as an audio feature e xtractor to extract a distincti ve feature from each musical piece. This frame- work bridges musical pieces and their corresponding co ver images using color palettes. T o demonstrate the effecti v eness of our frame work, we present an example application that generates co ver images using the estimated color palettes from gi ven musical pieces to support creating visually appealing cov er images. 2. RELA TED WORK Se veral studies ha ve in vestigated the relationship between music and color . W ells argued that there is a correlation between music and color based on the principle of comple- mentarity [17]. Furthermore, Pesek et al. suggested that since music and emotions are closely related (e.g., [18, 19]), as well as emotions and colors (e.g., [4, 5]), there exists a relationship between music and color mediated by emo- tions [20]. Ho wev er , these studies hav e only partially elu- cidated the relationship between music and color , as they analyzed this relationship using a limited number of colors. Therefore, in this study , we use the colors used in music cov er images that embody a musical essence [1, 2] as the basis for our analysis. In research exploring the colors used in co ver images, pre vious studies hav e focused on specific genres (classi- cal [8] and metal [9]). Seker [8] discov ered that the colors used in cov er images for classical music predominantly fa- v or neutral colors. Friconnet [9] found that cov er images for metal music tend to use darker colors than those of other genres, with a preference for black and orange [9]. Although these studies provide insights into the colors used in cov er images of specific genres, no studies hav e explored which color v alues are preferred for specific musical pieces of v arious genres. Additionally , color themes used in designing cov er im- ages ha ve been studied [10]. Dorochowicz and K ostek [10] analyzed cov er images across multiple genres with respect to basic color analysis rules such as seasonal colors (e.g., spring (warm and bright), summer (cool and soft), autumn (warm and soft), and winter (cool and bright)) and de grees of brightness (e.g., light, medium, and dark). While their findings provide v aluable insights into the color characteris- tics of each genre, they focus on a limited number of color palettes based on the basic color analysis rules. In this paper , we in vestigate the relationship between musical pieces and the color palettes used in their corre- sponding cov er images and explore the application of this relationship in cov er image design. Input im ages " -m eans (cluste ring method) Data -driven color m anifolds Figur e 2 . Comparison of color palette extraction methods. Gi ven the input image (top ro w), the data-dri ven color man- ifolds (bottom ro w) extract a color palette from the image in consecuti ve color order , while k -means (middle ro w) ex- tracts a color palette from the image in random color order . 3. COLOR EXTRA CTION T o extract a representati ve color or color palettes from music cov er images, we lev erage data-driven color manifolds [14], a technique which aims to acquire color samples from im- ages and learn a lo wer-dimensional manifold of the acquired color samples. The learned manifold reflects the distribu- tion of colors in cov er images, compressing areas of the color space that are less commonly used and expanding those that are more frequently utilized. The technique in v olves se veral steps, starting with the ac- quisition of color samples from cov er images. For success- ful color manifold learning, a suf ficient number of samples (ov er 10k) must be obtained from each image. Note that we utilized all samples from 224 px × 224 px -resized cov er images, amounting to over 50k samples. These samples are then used to estimate the density of each color in the cov er images, with a focus on identifying and preserving the most important colors. A self-organizing map [21], which is used to reduce dimensionality , is then applied to deri ve the one-dimensional or two-dimensional color manifolds. W e utilize the one-dimensional color manifold to extract color palettes from cov er images. In practice, we calculate a discrete color manifold, which consists of M ∈ N colors, to use the deri ved color manifold as a color palette. All hyperparameter v alues related to density estimation and dimensionality reduction were taken from [14], e xcept for the smoothness parameter , which we set to r 0 = 1 . The adv antage of this technique ov er clustering methods such as k -means [22] is that the color palette extracted by the data-dri ven color manifolds has a meaningful order - ing, where the order of colors is determined by the deri ved one-dimensional color manifold and thus results in consec- uti veness, while the color palette e xtracted by a clustering method has a random ordering (see Fig. 2). When using a color palette consisting of multiple colors in graphic design, Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 76 Classica l Stage & S creen Reggae L atin Rock El ectronic R G B 0 1 0 0 1 1 R G 0 1 0 0 1 1 R G 0 1 0 0 1 1 R G 0 1 0 0 1 1 R G 0 1 0 0 1 1 R G 0 1 0 0 1 1 B B B B B Low saturat ion Warm colors Wide variations Figur e 3 . V isualization of a representativ e color used in music cov er images by music genre. The larger the circle in the visualization, the more frequently the circle’ s color appears in the cov er images. the color palette with a continuous color order based on the data-dri ven color manifolds is intuiti ve and easy to use. 4. PRELIMINAR Y STUD Y This section describes our preliminary study that aims to analyze the preferred colors for designing music cov er im- ages across multiple genres by le veraging color palettes extracted from these images. 4.1 Experimental Setup 4.1.1 Dataset W e randomly collected 3,887 cov er images (each image is an RGB image) for the experiments. W e assigned genre tags to each image based on the grouping of genres and styles in Discogs 1 , in which v arious music is organized into 15 genres and styles (‘Blues, ’ ‘Brass & Military , ’ ‘Children’ s, ’ ‘Classical, ’ ‘Electronic, ’ ‘Folk, W orld, & Country , ’ ‘Funk / Soul, ’ ‘Hip-Hop, ’ ‘Jazz, ’ ‘Latin, ’ ‘Non-Music, ’ ‘Pop, ’ ‘Reg- gae, ’ ‘Rock, ’ and ‘Stage & Screen’). A total of 5,150 genre tags were assigned to 3,887 cov er images, which means an a verage of 343.3 images per genre. 1 The grouping of genres and styles in Discogs is av ailable at https://support.discogs.com/hc/en- us/articles/ 360005055213- Database- Guidelines- 9- Genres- Styles . T able 1 . List of representativ e colors most frequently used in music cov er images for each music genre, excluding grayscale colors. Music genre RGB v alue Color Blues (148,135,102) Brass & Military (110,101,74) Children’ s (139,178,241) Classical (105,132,128) Electronic (72,36,36) Folk, W orld, & Country (108,101,68) Funk / Soul (111,109,73) Hip-Hop (73,36,36) Jazz (181,145,109) Latin (146,112,110) Non-Music (165,127,156) Pop (110,73,73) Regg ae (168,132,68) Rock (72,36,36) Stage & Screen (174,172,106) 4.1.2 Implementation details For representing colors in color manifolds, we utilized an RGB color space, which is a widely used additiv e color model. W e resized all of the co ver images into 224 px × 224 px and normalized their RGB v alues to [0, 1]. Then, for the purpose of this preliminary study , we simply extracted one representati v e color (i.e., M = 1 ) from the resized images using the data-dri ven color mani- folds [14] as described in Section 3. Note that we extracted more colors to form the color palettes in Section 5.4. W e used all of the pixels in the resized images as color samples. The constructed color manifold is di vided into eight bins, and the samples are discretized into these bins, enabling their visualization as a histogram. 4.2 Results Fig. 3 represents three-dimensional histograms for the R, G, and B v alues in the RGB color space. As sho wn in Fig. 3, trends in the distrib ution of colors differ by genres: for example, ‘Classical’ and ‘Stage & Screen’ music co ver images ha ve lo w saturation, and ‘Reggae’ and ‘Latin’ music cov er images tend to use warm colors such as red and yello w . Additionally , it can also be observed that genres such as ‘Pop’ and ‘Rock’ feature a wide v ariety of colors in their cov er images. These genres hav e more di verse subcategories and styles than other genres, resulting in such color v ariation. Note that due to te xt on cov er images and background colors, grayscale colors are prominently displayed in the histogram. Therefore, T able 1 lists the representati ve colors most frequently used in co ver images, excluding grayscale colors. As sho wn in T able 1, the most frequently used colors v ary by music genre. Our proposed color palette estimation frame work is designed based on this insight. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 77 5. COLOR P ALETTE ESTIMA TION FRAMEWORK Using the close relationship between musical pieces and colors used in their corresponding music cov er images, we propose a frame work designed to estimate appropriate color palettes for musical pieces. Fig. 4 shows an o vervie w of our frame work. W e utilize an audio feature e xtractor and a color palette estimator to estimate a color palette for each musical piece. The audio feature extractor extracts audio features from musical pieces. T o le v erage recent advancements in audio models for do wnstream tasks, we use a pretrained audio model as the audio feature e xtractor . Then, the color palette estimator takes the e xtracted audio features as input and estimates appropriate color palettes for musical pieces. T o train the color palette estimator , we construct a large pri vate dataset of musical pieces and their corresponding cov er images, but we need the ground-truth color palettes to be estimated. Therefore, we extract the ground-truth color palette from each cov er image by le veraging the data-dri ven color manifolds [14] described in Section 3 for our color palette extraction method. 5.1 A udio F eature Extractor Audio models trained on lar ge datasets hav e demonstrated their capabilities in do wnstream tasks [15, 23]. F or exam- ple, the outputs of the final layer of an audio model are utilized in classification tasks, while audio embeddings are used in generati ve tasks. In our approach, we le verage the pretrained audio model F to extract audio features from musical pieces. Let A = { a n ∈ R T } N n =1 be a set of musical pieces, where T is the length of each musical piece and N is the number of musical pieces. Next, let Z = { z n ∈ R d } N n =1 be a set of audio features, where d is the number of dimensions of each audio feature. The audio feature z n can be extracted from the musical piece a n by using the pretrained audio model F as follo ws: z n = F ( a n ) . (1) W e utilize the pretrained audio model with all of its trainable parameters fixed. 5.2 Color Palette Estimator T o estimate color palettes from the extracted audio features z n , we propose the color palette estimator G , which consists of three linear layers with a GELU function [24]. Let C = { c n ∈ R 3 × M } N n =1 be a set of color palettes, where M is the number of colors in each color palette and each color is represented by a set of three numerical v alues (i.e., RGB values). The color palette c n can be estimated from the audio feature z n by using the color palette estimator G as follo ws: c n = G ( z n ) = σ ( W 3 GELU( W 2 GELU( W 1 z n + b 1 ) + b 2 ) + b 3 ) , (2) where σ is a sigmoid function. The parameters of the color palette estimator G are defined by W 1 ∈ R h × d , W 2 ∈ R h × h , W 3 ∈ R 3 × M × h , b 1 ∈ R h , b 2 ∈ R h , and b 3 ∈ Musical pi ece Music cover ima ge Audio feature extractor (fi xed) Color palette extraction Color palette estimator ! ! " # Estim ated color pa lette Ground-truth color pa lette Trai ning procedure Color pal ette esti mation fram ework Color pal ette esti mation Musical pi ece Audio feature extractor (fi xed) Color palette estimator (fixed) Estim ated color pa lette Figur e 4 . Overvie w of our proposed color palette estima- tion frame work. (T raining pr ocedure) W e start with an original pair of a musical piece and its corresponding music cov er image. The musical piece is processed by a fixed audio model to extract its audio feature, and then a color palette estimator is trained to estimate a color palette from the audio feature. For this training, the ground-truth color palette is extracted from the co ver image by using the color palette extraction method. W e use the mean squared error (MSE) loss function to optimize our color palette estimator . (Color palette estimation) After training, the color palette estimator can be used to estimate appropriate color palettes for musical pieces. R 3 × M , where the dimension of the hidden layer h is set to 768. While training, a dropout with a probability of 0.2 is applied to each output of the GELU functions. 5.3 Experimental Setup 5.3.1 Dataset The lar ge pri vate dataset for training our color palette esti- mator contains music audio excerpts (each e xcerpt is a 30 s audio pre vie w for trial listening, with a 44.1 kHz sampling rate) and their corresponding cov er images (each image is an RGB image). The excerpts and their co ver images are limited to single tracks, i.e., an original pair of an e xcerpt and its corresponding cov er image is unique. The dataset contains 582,894 pairs of an excerpt and its correspond- ing cov er image by 115,113 artists. W e randomly split the dataset into training, v alidation, and test sets with an eight- one-one ratio (i.e., 466,316 pairs for the training set and 58,289 pairs for v alidation and test sets, respectiv ely) and with no artists ov erlapping across these sets. 5.3.2 Implementation Details As described in Section 4.1.2, we utilized the RGB color space for color representation, resized the cov er images Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 78 into 224 px × 224 px , and normalized their RGB v alues to [0, 1]. W e used a single NVIDIA A6000 GPU to train the color palette estimator . Our implementation was based on PyT orch [25]. W e used the mini-batch size of 2,048. T o train the color palette estimator , we used the Adam opti- mizer [26] with a learning rate of 1 . 0 × 10 − 4 . W e calculated the mean squared error (MSE) loss function L M S E between the estimated and ground-truth color palettes to optimize the parameters of the color palette estimator . 5.4 Experimental Settings T o clarify which audio feature e xtractor would be most ef fectiv e in estimating the color palette, we conducted com- parati ve e xperiments to in vestigate the importance of selec- tion. Additionally , as a reference, we used color palettes composed of random colors. 5.4.1 Audio featur e e xtractor T o extract audio features from musical pieces, we compared two audio models: CLAP [15] and AudioT oken [16]. T o use CLAP [15], we set the parameters of pretrained models a vailable at HuggingF ace’ s T ransformers [27] (i.e., “laion/clap-htsat-fused”). Each musical piece was con verted to a mel spectrogram through a CLAP feature extractor , and the CLAP audio model used the spectrogram as input. W e fixed all parameters of the model. By using the model, we can obtain a 768-dimensional feature vector for each musical piece. W e also used AudioT ok en [16], which consists of a bidi- r ectional encoder r epr esentation fr om audio transformers (BEA Ts) [23] model and an embedder [16] model. W e set the parameters of pretrained models a v ailable at offi- cial GitHub repositories 2 . All parameters of the models were fixed. By employing the models, we can obtain a 768-dimensional feature vector for each musical piece. 5.4.2 Number of Colors in Color P alette W e used the number of colors M = { 1 , 2 , 3 , 4 , 5 } in the color palettes for the experiments. T o e xtract the color palettes from the cov er images, we used the data-driv en color manifolds [14] as described in Section 3. 5.5 Evaluation Metric f or Comparative Experiments In our comparati ve e xperiments, we used the minimum color dif fer ence model (MCDM) [28], which is practically designed to e v aluate the color difference between tw o color palettes. The MCDM compares the two color palettes, each consisting of M colors, to determine their av erage color dif- ference. First, the colors in the color palettes are con verted from RGB to CIELAB [29]. Then, the MCDM calculates a CIELAB color dif ference between each color in one palette and all colors in the other palette, identifying the closest color match for each and recording the minimum dif fer- ences. While multiple variants of CIELAB color dif ference 2 The pretrained models are av ailable at https://github.com/ microsoft/unilm/tree/master/beats for the BEA Ts model and https://github.com/guyyariv/AudioToken for the em- bedder model. T able 2 . Results for the MCDM score on the test set of our dataset. A lo wer MCDM score indicates a closer match between the estimated color palettes and the ground-truth color palettes. Audio feature extractor M MCDM score CLAP 1 28 . 37 2 25 . 66 3 22.68 4 21.72 5 21.17 AudioT oken 1 28.40 2 25.69 3 22 . 66 4 21 . 71 5 21 . 14 (Random) 1 69.29 2 68.09 3 66.35 4 65.90 5 65.50 exist, we here adopted the CIE1976 color dif ference [30] for e valuation. This process is repeated for ev ery color in the first palette, resulting in M color dif ference values, which are then a veraged to obtain a mean v alue, denoted as m 1 . The same process is repeated for the second palette, finding the closest matches in the first palette and a veraging the M minimum dif ferences to obtain another mean value, m 2 . Finally , the a verage of m 1 and m 2 gi ves the o verall color dif ference between the two palettes. The lower the MCDM score, the closer the two color palettes. W e le ver - age this MCDM to compare a color palette estimated with each experimental setting and a ground-truth color palette. 5.6 Results T able 2 presents the results for the MCDM score under each experimental setting. As sho wn in T able 2, our proposed frame work achie ves a much lo wer (i.e., better) MCDM score compared to random color palettes, with an improv e- ment of ov er 40 points. This demonstrates the ef fectiv eness of our frame work. Additionally , these results support that there is a relationship between musical pieces and the color palettes used for designing their corresponding cov er im- ages because our frame work succeeds in training the color palette estimator . Reg arding the selection of each experimental setting, there is no performance dif ference between the CLAP and AudioT oken audio models, as sho wn in T able 2. This sug- gests that either audio model can be selected based on the intended application. As these results demonstrate, our frame work ef fectiv ely estimates appropriate color palettes for musical pieces. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 79 1 0 1 0 1 1024 0 1 160 2000 1500 1000 500 0 500 1000 0 1 Figur e 1 . Mobile-AMT centers STFT windows on the time point to predict for , incurring a delay of 1024 samples. W e shift the window to reduce the delay to 160 samples, and change the windo w function to better use that limited amount of future information. combining weights and shift-tolerance ( T5 ). For the ne xt experiments, we use binary targets with weights. While the shift-tolerant loss performed fa v orably here, our goal is to use a causal model, for which shift tol- erance could result in systematically delayed predictions. 4.2 A udio Pr eprocessing Mobile-AMT processes audio in spectrogram frames of 2048-sample windo ws centered on the time points to pre- dict e vents for . Thus, e ven with a causal model that does not process information from future frames, in real-time inference, ev ery time a complete audio buf fer of 2048 sam- ples is filled, inference will be triggered to predict ev ents that are already 1024 samples (64 ms) in the past. Figure 1 illustrates this: the top ro w shows an audio w av eform, the second ro w a typical STFT windo w centered at the onset. Shorter filters reduce this delay , as the delay is fixed to half the filter length, b ut this comes at the cost of lower fre- quency resolution, which we should av oid: W ith a 2048- sample STFT at 16 kHz, the bin width is 7.8 Hz, which is already too coarse to achie ve semitone precision for the lo west piano notes (A0 at 27.5 Hz, B Z 0 at 29.14 Hz). W e can howe v er reduce this delay to a lo wer number n s of samples (e.g., 160 samples or 10 ms) while k eeping the windo w length and frequency resolution unchanged, by shifting windo ws so they end n s samples after their reference point instead of being centered. The third ro w in Figure 1 sho ws the Hann window shifted to n s = 160 samples. The Hann window strongly attenuates the bound- aries, the right one of which no w contains highly rele- v ant information for the prediction. T o mitigate this un- wanted attenuation, we can replace the Hann windo w with an asymmetric windo w that tapers (2048 − n s ) samples before and n s samples after the reference point. The last ro w in Figure 1 illustrates this windowing function for a 1888/160 sample asymmetry . Note ho w we keep more in- formation from the incoming samples in the gray shaded area under the windo w function, albeit at the cost of in- creasing spectral leakage (by about 20 dB for n s = 160 ). T olerance 10 ms 20 ms 30 ms H1 Hann 64 ms 17.43 ± 5.10 34.65 ± 7.40 39.86 ± 7.45 H2 Hann 10 ms 0.00 ± 0.01 0.00 ± 0.02 0.04 ± 0.10 T1 asym. 10 ms 22.25 ± 5.17 25.21 ± 5.20 25.84 ± 5.18 T2 asym. 20 ms 28.61 ± 7.09 33.76 ± 7.01 34.65 ± 6.89 T3 asym. 30 ms 27.61 ± 6.79 37.91 ± 7.54 39.43 ± 7.33 T4 asym. 40 ms 24.39 ± 6.50 37.51 ± 7.81 39.87 ± 7.60 T5 asym. 50 ms 20.99 ± 6.02 36.88 ± 7.75 40.47 ± 7.59 ST asym. 10 ms 0.41 ± 0.40 1.57 ± 0.82 11.86 ± 4.25 T able 2 . Note onset F1 scores on the MAESTR O v .3 vali- dation set for dif ferent windowing functions. The last ro w additionally uses a shift-tolerant training loss. T o experiment with dif ferent windowing configurations for reducing the delay in audio preprocessing, we mod- ify our reference method to apply only causal processing, as allo wing the model access to future frames would ren- der our interventions meaningless. Specifically , we make each con v olution causal, so the model’ s recepti ve field of 9 frames extends 8 frames into the past, rather than split- ting 4 frames into the past and 4 frames into the future. Additionally , we remov e the Squeeze-Excitation layers of the MobileNet V3 blocks, which perform global av erage pooling ov er both past and future frames in an excerpt. T able 2 shows our results. The original centered Hann windo w with our causal model ( H1 ) performs worse than our non-causal starting point ( TP3 in T able 1). Shifting the Hann windo w from a delay of 64 ms to a delay of 10 ms ( H2 ) seems to completely attenuate usable information in the frames. Using an asymmetric windo w ( T1 ) improv es performance, but still f alls behind the centered Hann win- do w . Successiv ely increasing the delay up to 50 ms, we see a strong improv ement ( T2 – T5 and Figure 2). For a delay of 30 ms or more, we match performance of the centered Hann windo w at an ev aluation onset tolerance of 30 ms. For stricter tolerances, shifted asymmetric windows of 20 ms delay or more surpass the centered Hann windo w . W e also take the chance to in vestig ate how a shift- tolerant loss of ± 1 frame af fects results for the causal model. The loss could allow the model to systematically predict e vents one frame (10 ms) later than annotated. Sur- prisingly , using an asymmetric window with 10 ms of de- lay , we find that the shift-tolerant loss ( ST ) performs on par with 30 ms delay ( T3 ) when admitting an e valuation tolerance of 50 ms (not sho wn in table), but breaks do wn with any stricter tolerance (as seen in the table). For the third group of e xperiments, we keep the strictest setting with asymmetric windo ws at a delay of 10 ms. 4.3 Model Architectur e In our final group of e xperiments, we in vestig ate architec- tural modifications to our reference model. The architec- ture of Mobile-AMT consists of three acoustic stacks, each consisting of recurrent con v olutional blocks. Each stack learns a (onset, frame or velocity) tar get. For some tar gets, the outputs of multiple stacks are concatenated to condi- tion the final predictions. Compared to their reference of- fline model [2], Mobile-AMT omits the acoustic stack for the of fset target, reusing the stack of the frame tar get. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 86 10 20 30 50 Onset tolerance (ms) 0.200 0.225 0.250 0.275 0.300 0.325 0.350 0.375 0.400 Mean note onset F1 scor e Asymmetric window with differ ent amounts of shif t 10ms 20ms 30ms 40ms 50ms Figur e 2 . Note onset F1 scores (means only) for different windo w delays and onset tolerance thresholds. T olerance 10 ms 20 ms 30 ms Note onset F1 mean ± std A1 20.23 ± 4.29 23.03 ± 4.25 23.75 ± 4.24 A2 25.77 ± 4.65 30.66 ± 4.98 31.79 ± 4.93 A3 21.59 ± 4.79 24.48 ± 4.83 25.14 ± 4.80 A4 21.53 ± 4.31 24.51 ± 4.24 25.22 ± 4.19 A5 19.44 ± 4.20 22.39 ± 4.34 23.12 ± 4.30 A6 25.82 ± 4.48 30.52 ± 4.51 31.56 ± 4.44 Note onset and of fset F1 mean ± std A1 3.25 ± 1.31 5.13 ± 2.54 7.20 ± 3.51 A2 3.56 ± 1.29 6.17 ± 2.80 8.39 ± 4.01 A3 5.84 ± 1.41 5.84 ± 2.48 7.71 ± 3.43 A4 5.94 ± 1.31 5.94 ± 2.48 7.89 ± 3.49 A5 4.84 ± 1.11 4.84 ± 2.36 6.61 ± 3.32 A6 6.79 ± 1.15 6.79 ± 2.53 8.89 ± 3.59 T able 3 . Note onset and onset-and-offset F1 scores on the MAESTR O v .3 v alidation set across three onset tolerances for architectural modifications and input representations. Our experiments in volv e the follo wing adaptations, each tested independently: First, in A1 we (re)introduce a separate of fset acoustic stack to explore whether and ho w it improv es of fset label prediction. In A2 we remov e the velocity conditioning on the onsets. Next, we e xamine whether further streamlining the architecture by sharing a fourth ( A3 ), half ( A4 ) or all ( A5 ) of the con volutional blocks in the model’ s acoustic stacks af fects performance. Lastly , in A6 we examine the ef fect of training on the orig- inal 10 seconds sequence length. T able 3 presents the note onset and note onset-and- of fset F1 scores of the model and data adaptations on the MAESTR O v alidation set. Overall, it is e vident that the combined impact of binary , hea vily imbalanced pointwise tar gets, causal modeling, and shifted asymmetric window results in a significantly harder learning problem, with the same training duration (500 epochs) leading to signif- icantly poorer scores than the base case ( TP1 in T able 1). Ho wev er , across all experimental setups in Section 4.1 compared to the current one, all our causal modifications demonstrate significantly stronger rob ustness to decreasing tolerance thresholds, which is important to guarantee low latency in predictions, and therefore appear promising for further training. Furthermore, when comparing all model architecture modifications ( A1-5 ) on note onset and note-onset-and- of fset F1 score, we observe tw o unexpected model be- ha viours: First, adding a separate offset stack ( A1 ) does not improv e of fset prediction. As our postprocessing de- tects a note of fset as the earlier of either offset acti v ation or frame inacti vation, we hypothesize that frame acti vity is suf ficiently learned to compensate for the absence of an of fset acoustic stack. Second, removing the v elocity con- ditioning on onset prediction ( A2 ) results in a strong im- prov ement in onset prediction. Furthermore, sharing the acoustic stack across increasing proportions ( A3-5 ) does not appear to hinder the model’ s ability to learn meaning- ful representations. Finally , experiment A6 suggests that the model benefits from the lar ger contextual windo w . 4.4 Final comparison For our final comparison, we proceed with the follo wing data and model configurations: we continue with the (160 samples) shifted asymmetric windo w for the STFT ( T1 in Sec. 4.2), remov e the velocity conditioning ( A2 ) and share all con v olutional layers in the acoustic stack across all tar- gets ( A5 ). Mobile-AMT uses the original non-causal post- processing described in Section 3.2, while our model use the causal postprocessing introduced in Section 4.1. T able 4 summarizes the results over dif ferent onset (and of fset) thresholds for note onset and onset-and-offset met- rics. As expected, Mobile-AMT outperforms our modified causal model across all metrics, with a significant margin. Upon re viewing all e xperiments conducted, we conclude that the lar gest performance drops are attributed to the shifted windo w function and the causal con volutions in our model. When comparing Mobile-AMT and our adapted model across v arious onset tolerance thresholds, we ob- serve, similar to the pre vious experiment, that while our modified causal model predicts fe wer targets with lo wer accuracy o verall, it demonstrates higher precision and ro- b ustness when ev aluated at stricter timing tolerances. 5. DISCUSSION AND OUTLOOK In this work, we in v estigate whether and ho w the cur- rent state of the art in real-time piano transcription can be adapted to achie ve minimum-latenc y automatic piano tran- scription suitable for real-time musical interaction. What latency is suitable cannot be answered uni ver - sally, so our choice of 10–30 ms is worthy of discus- sion. While 10 ms is suggested in digital instrument de- sign [8–10], thresholds for latency perception v ary depend- ing on the musical situation, task, and instrument: for per - cussi ve digital instruments, decreased ratings for v alues of 20 ms and above were found [20], instrument-specific thresholds between belo w 10 ms and 40 ms are reported in a li ve monitoring setting [21], and about 30 ms were found for gestural control [22]. Of fsets as low as 6 ms may be percei ved in simple isochronously spaced stimuli [23], while other researchers found just noticeable latency dif- ferences at 27 ms and higher [24]. T ranscription-enabled Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 87 Note Onset Note Onset with Offset Model T ol. (ms) Pr ecision Recall F1 Precision Recall F1 Causal-AMT 10 43.51 ± 6.87 25.60 ± 8.64 31.55 ± 7.76 6.86 ± 2.15 4.11 ± 2.03 5.03 ± 2.08 Mobile-AMT 10 22.70 ± 5.82 15.57 ± 2.96 18.26 ± 3.51 3.27 ± 1.19 2.32 ± 0.94 2.69 ± 1.00 Causal-AMT 20 50.86 ± 7.52 29.71 ± 9.11 36.70 ± 7.96 10.89 ± 3.91 6.34 ± 2.90 7.86 ± 3.16 Mobile-AMT 20 59.28 ± 7.41 41.87 ± 8.44 48.52 ± 6.96 11.03 ± 4.15 7.98 ± 3.73 9.16 ± 3.84 Causal-AMT 30 51.85 ± 7.49 30.24 ± 9.07 37.38 ± 7.85 14.49 ± 5.72 8.33 ± 3.93 10.37 ± 4.42 Mobile-AMT 30 81.17 ± 6.29 57.84 ± 12.79 66.80 ± 9.78 18.42 ± 6.89 13.27 ± 6.09 15.26 ± 6.30 T able 4 . Final comparison between our implementation of Mobile-AMT and our modified minimal-latency , strictly causally adapated model. real-time applications like interacti ve accompaniment or generati ve impro visation are more akin to ensemble play- ing than direct instrument control. In networked musical contexts, researchers typically aim for 20–30 ms of latency to meet performance conditions that mirror traditional in- person ensembles [11, 12]. Ho we ver , studies also found that musicians may be able to compensate for latencies up to 50 ms [13, 25] or e ven 100 ms [26] for one piano piece, a v alue that was deemed “neither musical nor interacti ve” in another study [27]. A real-time transcription model should not only be tolerable b ut enable fluent musical interaction, so we took 30 ms as a minimal requirement, and 10 ms as a goal for imperceptible latency . W e in vestigate multiple adaptations to reduce la- tency , including label encoding with causal postprocess- ing, shifted asymmetric windo w functions during prepro- cessing, and architectural modifications that enforce causal processing within the model. Additionally , we reduce the model size (from 320 to 160 GFLOPs for 3 seconds of in- put) by sharing computations across core model compo- nents for all tar gets. In a first set of experiments, we assess the impact of regression v ersus classification loss encodings for non- causal models. The original regression tar gets only make sense in conjunction with a lookahead as the targets be gin to increase se veral frames before the actual onset which is impossible for a causal model to predict. T o miti- gate the cost in training stability and accurac y incurred by localized, causal-ready tar gets, we experiment with loss functions that weight the acti ve frames o ver the inacti ve frame to combat label imbalance, and loss functions that are tolerant to small temporal shifts. W e find that the weighted classification losses approximate the baseline, and the shift-tolerant losses reach the same le vel in the ab- sence of tar gets requiring lookahead. In a second experiment, we in vestigate the delay in- curred by the computation of audio feature representations. Specifically , we look at STFT windo ws and their cor- responding centered tar gets. T ranscription requires high frequency resolution for pitch estimation which requires lar ge windows. Centering the targets results in an often ov erlooked delay of half the windo w length, 64 ms in our case. W e test configurations of shifted windo ws along with asymmetric windo wing functions that do not attenuate the most recent samples. W e find that aggressiv ely shifted windo ws at 10 ms do deteriorate the transcription accuracy by a lot, yet at 30 ms, we reach comparable performance to an unshifted causal model. Here, a shift-tolerant loss does not improv e performance. At the same time, configuring the model architecture for strictly causal processing also deteriorates performance with respect to the baseline with more than 100 ms of lookahead. In a third experiment, we assess dif ferent model archi- tectures and their impact. W e observe that sharing the con- v olutional components of the acoustic stack across differ - ent tar get types proves beneficial. W e hypothesize that the local acoustic features captured in the con v olutional layers of the acoustic model can be ef fectiv ely learned indepen- dently of sequential information, making them in v ariant to the tar get type. Furthermore, removing the velocity con- ditioning on the onsets strongly improv es the accuracy of onset predictions. Overall, we find that we can compensate well for algo- rithmic issues: we can scale the model and use lookahead- free tar gets without a major drop in performance. What prov es dif ficult, howe ver , is to render the model strictly causal and to ef fectiv ely process the incoming audio with- out loss of rele vant information. For a latenc y of 10 ms, it would be required that the model predicts pitches with at most 10 ms of incoming audio samples. For onsets of the lo west two octa ves on the piano, this means that there is not e ven a full period of the fundamental frequency present in the samples, and predictions may need to rely on harmonic partials. Along with the transient phase and the conse- quently blurry STFT frame, this leads to an increasingly hard transcription task. W e hope that these findings and pinpointed challenges will contrib ute to future research on real-time, minimum latency automatic piano transcription. While this study primarily focuses on the algorithmic performance and rob ustness of a real-time transcription model, we ackno wledge that a detailed analysis of pro- cessing time—including both network inference and pre- processing—across dif ferent hardware platforms remains an important area for future work to allo w for the practi- cal deployment of a real-time transcription system in real- world scenarios. Like wise, we want to take a closer in- spection into the design of the underlying windo w func- tion and filter bank, in order to find an appropriate balance between reducing prediction delay and increasing future context, all while maintaining the desired STFT properties. Lastly , across all experimental groups, our system adapta- tions consistently outperformed the baseline at lo wer tim- ing tolerance, which we consider a desirable property wor - thy of further in vestigation. Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 88 6. A CKNO WLEDGEMENTS This research ackno wledges support by the European Re- search Council (ERC), under the European Union’ s Hori- zon 2020 research and innov ation programme, grant agree- ment No. 101019375 Whither Music? . The LIT AI Lab is supported by the Federal State of Upper Austria. 7. REFERENCES [1] C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.- Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTR O dataset, ” in In- ternational Confer ence on Learning Repr esentations , 2019. [2] Q. K ong, B. Li, X. Song, Y . W an, and Y . W ang, “High- Resolution Piano T ranscription with Pedals by Re- gressing Onset and Of fset T imes, ” IEEE/A CM T rans- actions on A udio Speec h and Languag e Pr ocessing , v ol. 29, pp. 3707–3717, 2021. [3] S. Sigtia, E. Benetos, and S. Dixon, “ An end-to-end neural network for polyphonic piano music transcrip- tion, ” IEEE/A CM T ransactions on A udio, Speech, and Languag e Pr ocessing , v ol. 24, no. 5, pp. 927–939, 2016. [4] R. K elz, S. Böck, and G. W idmer , “Deep polyphonic adsr piano note transcription, ” in ICASSP 2019-2019 IEEE International Confer ence on Acoustics, Speec h and Signal Pr ocessing (ICASSP) . IEEE, 2019, pp. 246–250. [5] Y . K usaka and A. Maezawa, “Mobile-AMT : Real- T ime Polyphonic Piano T ranscription for In-the-W ild Recordings, ” in 2024 32nd Eur opean Signal Pr ocess- ing Confer ence (EUSIPCO) . IEEE, 2024, pp. 36–40. [6] A. Fernandez, “Onsets and V elocities: Af fordable Real-T ime Piano T ranscription Using Con v olutional Neural Networks, ” in 2023 31st Eur opean Signal Pr o- cessing Confer ence (EUSIPCO) . IEEE, 2023, pp. 151–155. [7] T . Kw on, D. Jeong, and J. Nam, “T o wards Ef ficient and Real-T ime Piano T ranscription Using Neural Autore- gressi ve Models, ” IEEE/A CM T ransactions on A udio, Speech, and Langua ge Pr ocessing , 2024. [8] D. W essel and M. Wright, “Problems and prospects for intimate musical control of computers, ” Computer mu- sic journal , v ol. 26, no. 3, pp. 11–22, 2002. [9] A. P . McPherson, R. H. Jack, and G. Moro, “ Action- sound latency: Are our tools f ast enough?” in 16th International Confer ence on Ne w Interfaces for Musical Expr ession, NIME 2016, Griffith University , Brisbane, Austr alia, J uly 11-15, 2016 . nime.org, 2016, pp. 20–25. [Online]. A v ailable: https://doi.or g/ 10.5281/zenodo.3964611 [10] F . Caspe, J. Shier , M. Sandler , C. Saitis, and A. McPherson, “Designing neural synthesizers for lo w latency interaction, ” arXiv pr eprint arXiv:2503.11562 , 2025. [11] L. T urchet and C. Rottondi, “On the relation between the fields of network ed music performances, ubiqui- tous music, and internet of musical things, ” P ersonal and Ubiquitous Computing , vol. 27, no. 5, pp. 1783– 1792, 2023. [12] E. Lakiotakis, C. Liaskos, and X. Dimitropoulos, “Im- proving netw orked music performance systems us- ing application-network collaboration, ” Concurr ency and Computation: Practice and Experience , vol. 31, no. 24, p. e4730, 2019. [13] E. Che w , R. Zimmermann, A. A. Sawchuk, C. K yr- iakakis, C. P apadopoulos, A. François, G. Kim, A. Rizzo, and A. V olk, “Musical interaction at a dis- tance: Distributed immersi ve performance, ” in Pr o- ceedings of the MusicNetwork F ourth Open W orkshop on Inte gr ation of Music in Multimedia Applications . MusicNetwork Barcelona, 2004, pp. 15–16. [14] T . Kw on, D. Jeong, and J. Nam, “Polyphonic Piano T ranscription Using Autoregressi v e Multi-State Note Model, ” in The 21th International Society for Music Information Retrieval Confer ence (ISMIR) . Interna- tional Society for Music Information Retrie val, 2020. [15] D. Jeong and S. T elecom, “Real-time automatic piano music transcription system, ” in Late Br eaking Demo. International Society for Music Information Retrieval , 2020, pp. 4–6. [16] A. Ho ward, M. Sandler , G. Chu, L.-C. Chen, B. Chen, M. T an, W . W ang, Y . Zhu, R. Pang, V . V asude van et al. , “Searching for MobileNetV3, ” in Pr oceedings of the IEEE/CVF international confer ence on computer vi- sion , 2019, pp. 1314–1324. [17] C. Raf fel, B. McFee, E. J. Humphrey , J. Salamon, O. Nieto, D. Liang, D. P . Ellis, and C. C. Raffel, “MIR_EV AL: A T ransparent Implementation of Com- mon MIR Metrics.” in ISMIR , v ol. 10, 2014, p. 2014. [18] R. M. Bittner , J. J. Bosch, D. Rubinstein, G. Meseguer - Brocal, and S. Ewert, “ A lightweight instrument- agnostic model for polyphonic note transcription and multipitch estimation, ” in Pr oceedings of the IEEE In- ternational Confer ence on Acoustics, Speech, and Sig- nal Pr ocessing (ICASSP) , Singapore, 2022. [19] F . Foscarin, J. Schlüter , and G. W idmer , “Beat this! Accurate beat tracking without DBN postprocessing, ” arXiv pr eprint arXiv:2407.21658 , 2024. [20] R. H. Jack, A. Mehrabi, T . Stockman, and A. McPher- son, “ Action-sound latenc y and the perceiv ed quality of digital musical instruments: Comparing professional Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 89 percussionists and amateur musicians, ” Music P er cep- tion , vol. 36, no. 1, pp. 109–128, 09 2018. [Online]. A v ailable: https://doi.or g/10.1525/mp.2018.36.1.109 [21] M. Lester and J. Boley , “The effects of latenc y on li ve sound monitoring, ” Journal of the A udio Engineering Society , no. 7198, october 2007. [22] T . Mäki-P atola and P . Hämäläinen, “Latenc y tolerance for gesture controlled continuous sound instrument without tactile feedback, ” in Pr oceedings of the 2004 International Computer Music Confer ence, ICMC 2004, Miami, Florida, USA, November 1-6, 2004 . Michigan Publishing, 2004. [Online]. A vailable: https://hdl.handle.net/2027/spo.bbp2372.2004.032 [23] A. Friberg and J. Sundber g, “T ime discrimination in a monotonic, isochronous sequence, ” The Journal of the Acoustical Society of America , vol. 98, no. 5, pp. 2524–2531, 1995. [24] A. Schmid, M. Ambros, J. Bogon, and R. W immer , “Measuring the just noticeable dif ference for audio latency , ” in Pr oceedings of the 19th International A udio Mostly Confer ence: Explorations in Sonic Cultur es, AM 2024, Milan, Italy , September 18- 20, 2024 , L. A. Ludovico and D. A. Mauro, Eds. A CM, 2024, pp. 325–331. [Online]. A v ailable: https://doi.or g/10.1145/3678299.3678331 [25] S. Dahl and R. Bresin, “Is the player more influenced by the auditory than the tactile feedback from the in- strument, ” in Pr oceedings of the Digital A udio Effects Confer ence (D AFx) , 2001, pp. 6–9. [26] A. A. Sawchuk, E. Chew , R. Zimmermann, C. Pa- padopoulos, and C. Kyriakakis, “From remote media immersion to distrib uted immersiv e performance, ” in Pr oceedings of the 2003 A CM SIGMM W ork- shop on Experiential T elepr esence , ser . ETP ’03. Ne w Y ork, NY , USA: Association for Computing Machinery , 2003, p. 110–120. [Online]. A v ailable: https://doi.or g/10.1145/982484.982506 [27] C. Bartlette, D. Headlam, M. Bocko, and G. V elikic, “Ef fect of network latency on interacti v e musical performance, ” Music P er ception , vol. 24, no. 1, pp. 49–62, 09 2006. [Online]. A v ailable: https: //doi.or g/10.1525/mp.2006.24.1.49 Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 90 MA TCHMAKER: AN OPEN-SOURCE LIBRAR Y FOR REAL-TIME PIANO SCORE FOLLO WING AND SYSTEMA TIC EV ALU A TION Jiyun Park 1 ∗ Carlos Cancino-Chacón 2 ∗ Suhit Chiruthapudi 2 J uhan Nam 1 1 Graduate School of Culture T echnology , KAIST , South K orea 2 Institute of Computational Perception, Johannes K epler Uni versity Linz, Austria {june,juhan.nam}@kaist.ac.kr, {carlos.cancino_chacon,suhit.chiruthapudi}@jku.at ABSTRA CT Real-time music alignment, also kno wn as scor e follow- ing , is a fundamental MIR task with a long history and is essential for many interacti v e applications. Despite its im- portance, there has not been a unified open frame work for comparing models, lar gely due to the inherent complex- ity of real-time processing and the language- or system- dependent implementations. In addition, low compatibil- ity with the existing MIR en vironment has made it diffi- cult to de velop benchmarks using lar ge datasets av ailable in recent years. While ne w studies based on established methods (e.g., dynamic programming, probabilistic mod- els) ha ve emer ged, most e valuations compare models only within the same family or on small sets of test data. This paper introduces Matchmak er , an open-source Python li- brary for real-time music alignment that is easy to use and compatible with modern MIR libraries. Using this, we sys- tematically compare methods along two dimensions: mu- sic representations and alignment methods. W e e v aluated our approach on a lar ge test set of solo piano music from the (n)ASAP , Batik, and V ienna4x22 datasets with a com- prehensi ve set of metrics to ensure rob ust assessment. Our work aims to establish a benchmark frame work for score- follo wing research while providing a practical tool that de- velopers can easily inte grate into their applications. 1. INTR ODUCTION Real-time music alignment, also kno wn as scor e follow- ing , is the task of aligning performance data to the corre- sponding position in the musical score in real-time. Ever since it was first introduced independently by Roger Dan- nenber g [1] and Barry V ercoe [2] ov er 40 years ago, music alignment has become one of the fundamental MIR tasks. * Equal contribution. © J. Park, C. Cancino-Chacón, S. Chiruthapudi and J. Nam. Licensed under a Creati ve Commons Attribution 4.0 International Li- cense (CC BY 4.0). Attribution: J. Park, C. Cancino-Chacón, S. Chiruthapudi and J. Nam, “Matchmaker: An Open-Source Library for Real-T ime Piano Score Follo wing and Systematic Evaluation”, in Pr oc. of the 26th Int. Society for Music Information Retrieval Conf., Daejeon, South K orea, 2025. Score follo wing is a necessary component of many inter - acti ve applications (e.g., automatic accompaniment sys- tems [3–6], automatic page turning [7, 8], lyrics align- ment or tracking singing v oice [9–11], audiovisual/mul- timodal [6, 12] and visualizations [13]. Music alignment beg an as real-time score following [1, 2, 14–17] b ut, by the mid-90s, had div erged into online and of fline methods (see, e.g., early of fline work by Desain et al. [18]). From its early use on monophonic sources like v oice [17] and wind instruments, score following has gro wn to support polyphonic instruments such as piano, ensemble, and e ven full orchestral performances [17, 19–21]. Re- search has also expanded across input modalities of the performance, with systems operating on audio or MIDI, and score representations including string format, sym- bolic score, and sheet image [22]. The score follo wing challenge [23] in MIREX laid the foundation to formalize the e valuation frame work, in- troducing important metrics that include considerations in real-time. Ho we ver , many subsequent studies ha ve been de veloped in dif ferent en vironments—ranging from system-dependent [24, 25] to language-dependent [26, 27] implementations—often tailored to specific use cases and without publicly shared source code. As a result, imple- mentations became fragmented across platforms, making it dif ficult to extend, reproduce, or compare methods in a unified setting. This has hindered the de velopment of a unified e valuation frame work and comparison o ver meth- ods or features on shared datasets remain rare, limiting the generalizability and reproducibility . In this paper , we address these challenges by proposing a unified, open frame work for the e valuation and bench- marking of real-time audio-based score follo wing. Consid- ering public datasets that of fer a range of dif ficulty le vels, multiple renditions, and precise beat-le vel annotations, we base our e valuation on three representati v e piano perfor- mance datasets. W e implement this framew ork as an open- source Python package called Matchmak er , 1 that allows real-time ex ecution of representati ve baselines of score fol- lo wing algorithms. In addition to benchmarking, it sup- ports audio de vice input and has been validated in applica- tion contexts through a standalone demo system. 1 https://github.com/pymatchmaker/matchmaker 91 2. A CONCEPTU AL FRAMEW ORK FOR SCORE FOLLO WING As a way to or ganize and compare the components of sys- tems for score follo wing, we follow the structure proposed by Müller [28]. This frame work consists of three core components: (1) input music representations, (2) features, and (3) online alignment algorithms. 2.1 Music Representation Score follo wing aligns a fixed reference deri ved from mu- sical scores with a time-e volving input from a perfor - mance. The score can take v arious symbolic formats (e.g., MIDI, MusicXML) or sheet images, and is typically con- verted into an intermediate representation such as synthe- sized audio or e vent sequences. The performance input may be gi ven as either audio or MIDI, each with distinct representational and computational characteristics. Au- dio input is continuous and latency-sensiti ve, while MIDI is discrete and e vent-based. Instrumental factors also af- fect alignment design: polyphonic or discrete-pitch instru- ments (e.g., piano) dif fer from continuous-pitch sources (e.g., violin, voice). Multi-instrument recordings pose fur - ther challenges due to timbral ov erlap and source ambigu- ity . 2.2 F eatures Chroma features are the most commonly used in music synchronization, with many v ariants for their computa- tion [29–32]. Other works also use v arious spectral fea- tures such as constant-Q transforms (CQT) [27, 33], non- neg ativ e matrix factorization(NMF)-based [34] or spectral template [35] for improv ed polyphonic alignment. Beyond spectral representations, context-a ware features such as onset-based feature [36] or beat-synchronous frames ha ve been introduced to capture temporally salient e vents use- ful for alignment. Later work e xplored learned features, including feedforward mappings [27], semi-supervised decompositions like NMF , and more recent neural ap- proaches [37]. While these of fer richer contextual infor - mation, they often rely on fix ed-length inputs and intro- duce latency , making real-time usage more challenging. 2.3 Alignment Algorithms T wo major families of alignment algorithms ha ve been used in score follo wing: dynamic programming and prob- abilistic models. The dynamic programming approach, especially dy- namic time warping (DTW), aligns tw o sequences by min- imizing cumulati ve cost. Its online v ariant, On-Line T ime W arping (OL TW) [38], enables causal alignment within a fixed-size of windo w . V ariants include windo wed [39], parallel [40], and constrained DTW [40, 41], as well as tempo-aw are extensions [21, 42]. Probabilistic state-space models of fer an alternativ e by treating alignment as latent state inference under uncer - tainty [24, 29, 43]. HMM-based systems model each note as a sequence of states (e.g., attack–steady–release), with extensions including semi-Mark ov [44], hybrid [19], and Bayesian v ariants [45]. Kalman filter models and switch- ing state-space systems [46, 47] further incorporate tempo dynamics, while particle filters [12, 29] handle multimodal uncertainty in real time. Other paradigms include early string-matching algo- rithms [1] and reinforcement learning-based approaches for multimodal or visual score alignment [48]. 3. IMPLEMENT A TION 3.1 Python Package Structure Matchmak er is an open source Python package that imple- ments representati ve real-time music alignment algorithms within a modular , e xtensible framew ork. Figure 1 illus- trates the ov ervie w of the package and the whole pipeline. The current version of Matc hmaker pro vides two types of algorithms: 1) online time warping, with two v ariants: OLTWDixon , based on the methods proposed in [38, 49], and OLTWArzt , based on [21, 50]; and 2) an HMM-based algorithm, similar to the one used in [3, 47]. A full descrip- tion of the algorithms and their parameters can be found in the supplementary Appendix. 2 Matchmak er supports two main usage scenarios: (1) li ve streaming mode using the audio de vice and (2) sim- ulation mode, which processes a performance file as in- put. Figure 2 shows an e xample of running liv e streaming mode with the default setting. The AudioStream object handles the input stream by chunking the audio with o ver - lapping windo ws to av oid padding artifacts. Both the syn- thesized score audio and the performance audio are passed to a Processor object that performs feature extraction. The extracted features are pushed into a queue and con- sumed by the OnlineAlignment object, which runs the alignment methods in real time. Matchmaker tak es a mu- sical score with all symbolic music formats (MusicXML, MIDI, MEI, etc.) av ailable by partitura . 3 The returned output is the current position in the score, represented in beats as a musical unit according to the time signature in the piece. More detailed description and API documenta- tion of the package are av ailable here. 4 3.2 Design and Implementation Details W e provide a simple and user -friendly interface to run the score follo wing with minimal setup. As shown in Fig- ure 2, users can instantiate a Matchmaker object with a score file and ex ecute a run that iterates ov er the esti- mated score position for each step. T o streamline real-time processing, the AudioStream class is implemented as a context manager that automatically handles stream initial- ization and teardo wn. Furthermore, the alignment process is designed as a generator , enabling users to recei ve score positions concurrently while the alignment is in progress. 2 https://pymatchmaker.github.io/ismir2025_ supplementary_materials/ 3 https://github.com/CPJKU/partitura 4 https://pymatchmaker.readthedocs.io/ Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 92 Figur e 1 . Ov ervie w of the score follo wing package 1 from matchmaker import Matchmaker 2 3 mm = Matchmaker( 4 score_file= "path/to/score.musicxml" , 5 input_type= "audio" , 6 ) 7 for current_position in mm.run(): 8 print (current_position) Figur e 2 . A code e xample for running the Matc hmak er in a li v e streaming mode. This design allo ws for ef ficient real-time inte gration with- out requiring users to manage multiple threads, b uf fers, or callbacks e xplicitly . While the online mode uses a multi-threaded queue for asynchronous audio b uf fering, the simulation mode pro- cesses audio chunks in adv ance within a single-threaded setup. By decoupling real-time I/O concerns from core alignment e v aluation, it is intended to a v oid v ariability from Python v ersion, OS-le v el threading, or queuing de- lays, ensuring a consistent and reproducible benchmarking en vironment. In addition, OL TW Arzt is implemented in Cython [51] for ef ficienc y , a superset of Python designed for C-lik e performance by incorporating C data types and optimizing the e x ecution of Python code. 4. EXPERIMENTS 4.1 Datasets W e use three public piano performance datasets: (n)ASAP [52], Batik [53] and V ienna 4x22 [54], each of them of fer - ing complementary characteristics for benchmarking score follo wing. (n)ASAP , a subset of the MAESTR O dataset including note-le v el score alignments, includes e xpressi v e performances of technically demanding solo piano pieces, of fering high dif ficulty and stylistic di v ersity . W e use only the pieces i n the MAESTR O v2 test split. V ienna4x22 pro vides 22 distinct renditions for each of four relati v ely easy pieces, which is suitable to test rob ustness to inter - preti v e v ariation. Batik dataset contains recordings of 12 Mozart sonatas by a single pianist with the longest a v erage piece duration among the three datasets, enabling e v alua- tion across long-form classical repertoire. W e use ground-truth beat-le v el annotations pro vided with the (n)ASAP dataset, and e xtract equi v alent annota- Dataset #Pieces #P erf #Beats #Notes Dur (h) Difficulty (n)ASAP 43 59 26,329 100,958 2.65 6.53 Batik 30 30 18,789 102,421 2.85 5.67 V ienna 4 88 13,728 43,656 2.24 4.88 T otal 77 177 58,846 247,035 7.74 6.11 T able 1 . Datasets used in the e v aluation. tions for Batik and V ienna4x22 from the .matc h files [55], which contain note-wise score–performance alignments. In addition, we incorporate the dif ficulty le v els of each piece based on G. Henle Publishers, 5 which pro vides a 1-to-9 grading scale. The pieces used in our e xperiments span le v els 4 through 9, representing a di v erse set of w orks abo v e intermediate le v el. T able 1 pro vides the detailed statistics of the datasets. W e only included performances in the e xperiment that recorded an MAE of less than 100 m s in the of fline test, using the synctoolbox 6 with Chroma & DLNCO features. The e v al uation w as conducted on 184 performances across 93 pieces, totaling o v er 58,000 beats and 247,000 notes, with an o v erall duration of 7.74 hours of performances and a piece-wise a v erage dif ficulty of 6.11. 4.2 Experiment Settings W e conducted all e v aluations under simulation-based con- ditions to ensure reproducibility . Li v e testing w as a v oided due to v ariability introduced by room acoustics and hard- w are setup, which complicates f air comparison across sys- tems. The accurac y tests were carried out on an Intel i9- 9900K CPU (16 cores @ 3 . 6 G Hz ), Python 3.9, with a sample rate of 44.1 kHz and a frame rate of 30, chosen to balance latenc y and alignment accurac y . W e tested chro- magram, mel-spectrogram, constant-Q transform (CQT), mel-frequenc y cepstral coef ficients (MFCCs) [56] and a simple STFT -based onset-sensiti v e representation similar to the one used in Dixon [38], which we name log-spectral ener gy (LSE). While results for all features were e v aluated, we report detailed latenc y and accurac y metrics for the best-performing configuration of each model. T o account for hardw are v ariability , latenc y w as measured in multiple setups: an Intel i9-9900K, an Apple M4 MacMini, and an 5 https://www.henle.de/Levels- of- Difficulty/ 6 https://github.com/meinardmueller/synctoolbox Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 93 Figur e 3 . T w o e xamples of error calculation using the mapping function. (a) sho ws a one-to-man y alignment at the e v aluation point, while (b) illustrates a skipped align- ment. Apple M2 Pro MacBook, with the reported latenc y v alues a v eraged across these de vices. 4.3 Pr epr ocessing In the preprocessing step (see Fig. 1), the symbolic scores are synthesized to audio using FluidSynth, pro vided by partitur a . Since MusicXML often lacks tempo markings, we set the synthesis tempo to each performance’ s a v er - age—rounded to the nearest 20 BPM—assuming perform- ers follo w approximate tempo indications. T o generate beat annotations for the synthesized score audio, we computed beat positions using the synthesis tempo and the score’ s time signature. F or compound meters (e.g., 6/8, 9/8, and 12/8), we adopted (n)ASAP’ s beat annotation rules—counting them as tw o, three, and four beats per measure, respecti v ely—across all datasets to align score-side annotations with performance annotations. Based on the synthesized audio, we then e xtract the feature using the same Processor used in the online phase, b ut precompute them of fline for the entire score sequence. 5. EV ALU A TION Ev aluating score follo wing is challenging due to causality , timing precision, and output latenc y . Since the MIREX challenge [23] pro vided foundational metrics, later studies introduced alternati v e e v aluation strate gies including beat- le v el e v aluations or asynchron y [3], reflecting the task’ s frequent inte gration with automatic accompaniment sys- tems. In this w ork, we adopt tw o complementary e v aluation perspecti v es. First, we e v aluate in the performance do- main, where errors are measured in milliseconds based on ground-truth annotations aligned to the audio. This ap- proach is commonly used in audio-to-score alignment re- search and enables precise, fram e-le v el e v aluation, since the annotations directly reflect the actual timing of the per - formance. Second, we also e v aluate in the score domain measured in beat units as suggested in [29, 57], which bet- ter reflects the nature of score follo wing as a task of pre- dicting the corresponding score position at each moment of the performance. Figur e 4 . Defined delay types of the system. Only system delay is considered in the e xperiment. 5.1 Ev aluation Metrics W e select e v aluation metrics mostly adapted from score follo wing MIREX benchmark [23] and audio-to-score alignment (ASA) metrics [57]. W e use Alignment Rate (AR) within a tolerance range of | θ e | , v arying from 50 ms to 2000 ms. W e also compute Absolute Err ors (AE), both in milliseconds and in beats, from which we deri v e the A v erage Absolute Err or (AAE) and Median Absolute Err or (MAE), along with the standard de viation σ e . T o further characterize the distrib ution of errors, we report kurtosis and sk ewness which capture the peak edness and asymmetry of the non-absolute error distrib uti o n, respec- ti v ely . In addition, we report the a v erage latency µ l a t , de- fined as the system delay from the detection time to the end of inference. Unlik e total latenc y , this e xcludes har dwar e latency and is composed of tw o parts: (i) feature process- ing and (ii) e x ecution of the online alignment algorithm for each frame step (see Fig. 4). Errors e xceeding 2 sec- onds (or 2 beats in the score domain) are e xcluded from AE calculations, including both AAE and MAE, to a v oid distortion from unbounded tracking f ailur es. W e report AR in tw o w ays. The a v eraged piece-wise AR is a common measure, while the total AR reflects the proportion of suc- cessfully aligned beat e v ents across the entire dataset. The latter a v oids o v errepresentation of shorter pieces and pro- vides a more balanced vie w of o v erall performance. T o e v aluate runtime latenc y under simulation, we mea- sure tw o components: the a v erage duration for e xtracting features from incoming audio frames, and the time tak en by the alignment process to consume features and predict score positions. Specifically , the latenc y w as computed from the moment audio w as read to the time the score posi- tion w as predicted—e xcluding hardw are I/O delays. This tw o-step measurement allo ws for standardized latenc y re- porting independent of the hardw are setup. 5.2 Alignment Mapping Function Gi v en t h e alignment path, the alignment mapping function is applied to transfer the beat positions on one axis (ei- ther performance or score) to another axis to compute the alignment error . Due to the local , stepwise nature of real- time alignment, the resulting path is not necessarily mono- tonic and may contain multiple correspondents or skipped positions, depending on the impleme n t ation and purpose of the methods. Un l ik e linear interpolation methods com- monly used in of fline audio-to-score alignment, which as- sume continuous mappings, our e v aluation relies only on Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 94 Dataset Method AAE(ms) ↓ ± σ MAE(ms) ↓ Skew . Kurt. Piece-wise AR (%) ↑ T otal AR ↑ ( ≤ 2000ms, %) ≤ 50ms ≤ 100ms ≤ 500ms ≤ 1000ms ≤ 2000ms (n)ASAP OL TWDixon 189.55 ± 281.55 97.09 3.20 17.97 40.3 58.5 82.5 88.3 92.0 89.4 OL TW Arzt 183.56 ± 263.95 91.18 0.75 11.79 44.1 58.3 84.8 92.0 95.1 92.8 HMM 487.73 ± 423.27 346.01 0.18 3.33 15.6 22.2 37.5 43.8 43.8 43.8 Batik OL TWDixon 186.97 ± 262.55 104.40 3.75 24.70 28.2 51.7 82.1 85.2 87.6 89.4 OL TW Arzt 193.36 ± 269.13 107.15 1.00 12.63 35.9 53.0 82.2 87.4 90.3 89.7 HMM 693.63 ± 376.58 641.77 0.11 0.98 4.5 10.8 34.0 46.2 64.2 61.9 V ienna4x22 OL TWDixon 285.43 ± 390.82 132.73 1.57 5.90 26.6 43.2 72.4 80.0 85.5 82.5 OL TW Arzt 300.41 ± 368.70 152.51 0.50 3.93 33.2 44.5 73.3 84.3 86.7 86.7 HMM 439.64 ± 427.02 319.13 0.15 3.79 23.5 33.3 51.1 57.1 63.0 75.9 T able 2 . Ev aluation results on three datasets using dif ferent score-following methods. The piece-wise alignment rate (AR) is measured as the a verage ov er pieces, while the total AR indicates the global proportion of aligned beat ev ents across the entire dataset. All tests were conducted with STFT -based Chroma as features. predictions made prior to or at each e valuation time point. T o reflect this, we define the mapping function as follows: ˆ u k = min  u i | ( u i , v i ) ∈ W , v i = max { v j | v j ≤ k }  , where W = { ( u i , v i ) } is the warping path e xpressed in the frame indices: u i is the score-rendered-audio frame index and v i is the performance-audio frame index. The inner max finds the latest performance frame v i not exceeding the current frame k , and the outer min selects the smallest score frame u i among those alignments. This mapping re- lies solely on past or current frames to maintain causality . It handles skipped or one-to-many mappings and a v oids any interpolation methods that depend on future frames. 6. RESUL TS T able 2 presents a comparison of alignment methods based on performance-domain e valuation, measured in millisec- onds. All methods exhibit positi v e ske wness in error distrib ution, reflecting the expected lag of the beat esti- mates in real-time alignment. The ov erall results sho w that the OL TW -based method outperforms the HMM baseline across all datasets in both alignment accuracy and co ver - age. While OL TWDixon and OL TW Arzt sho w compa- rable MAE depending on the dataset, OL TW Arzt consis- tently achie ves higher co verage ( T otal AR ), suggesting that it is more rob ust against o verall failures. The dif ference likely stems from OL TWDixon skipping uncertain re gions, while OL TW Arzt’ s “backward-forward” strate gy corrects early misalignments and enhances cov erage. Despite ha v- ing the lo west AR, the HMM sho ws the lowest sk ewness and kurtosis primarily because significant errors (>2 s) are excluded from the summary statistics and its “stick y” be- ha vior to linger in the same state in local regions tends to narro w the error distribution. T able 3 presents an ev aluation comparison in beat units, of fering a tempo-normalized perspectiv e. The overall trend mirrors the performance-domain results in millisec- ond, but these results are standardized across tempi. AAE remains around 0.3 beats, with median v alues typically be- lo w 0.2. T otal AR is consistently lo wer than the 2000 ms - Dataset Method AAE ↓ (beats) ± σ MAE ↓ (beats) AR ↑ (%) (n)ASAP OL TWDixon 0.22 ± 0.27 0.13 83.4 OL TW Arzt 0.27 ± 0.30 0.16 85.2 HMM 0.80 ± 0.54 0.66 76.9 Batik OL TWDixon 0.20 ± 0.27 0.11 88.9 OL TW Arzt 0.29 ± 0.34 0.18 88.8 HMM 0.80 ± 0.38 0.67 59.3 V ienna4x22 OL TWDixon 0.31 ± 0.33 0.19 78.3 OL TW Arzt 0.37 ± 0.38 0.24 84.0 HMM 0.76 ± 0.78 0.51 70.3 T able 3 . Beat-lev el e valuation results including total align- ment rate (AR) (%). F eature Process Online Alignment T ype MAE (ms) Latency (ms) Method Latency (ms) Chroma 265.50 3.05 OL TWDixon 1.22 mel 297.92 3.40 OL TW Arzt 0.07 CQT 341.25 42.58 HMM 3.59 LSE 241.85 0.91 MFCC 931.81 2.58 T able 4 . Comparison of feature types and alignment meth- ods in terms of alignment error (MAE) and latenc y . LSE is log-spectral ener gy feature that was adopted in [38]. La- tency v alues are a veraged o ver the hardware setups e v alu- ated in Section 4. based metric, reflecting that most pieces ha ve tempi abo ve 60BPM, where two beats span less than tw o seconds. In addition, a comparison of v arious feature types and latencies of the alignment methods are reported in T able 4. Among the features, log-spectral energy (LSE) sho ws the lo west MAE ( 241 . 85 ms ) and delay ( 0 . 91 ms ), indicat- ing strong performance with minimal ov erhead. In con- trast, CQT and MFCC yield higher MAE, with CQT also requiring considerable extraction time ( 42 . 58 ms ), which limits its real-time suitability . For alignment methods, OL TW Arzt achie ves the lo west latency ( 0 . 07 ms ), whereas HMM sho ws noticeably higher delay ( 3 . 59 ms ) due to its computational complexity . These results highlight a trade- Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 95 Figur e 2 : Br eath Cue and Onset Annotation Breath cue is defined by the breath onset and of fset annotated on the mel spectrogram. Six mark ers are annotated per trial. De- tailed information can be found in Section 3.2 de grees in performance. Each flutist-pianist pair had not 173 pre viously performed together . 174 3.1.2 Placement and Equipment 175 The setup mimick ed a concert stage, with flutists posi- 176 tioned f acing a w ay from the pianists. The pianists could 177 observ e the flutists, while the flutists were instructed to 178 gi v e cues without turning or looking at the pianists. A 179 camera recorded flutists at 60 fps 2 , and microphones sep- 180 arately captured audio from each instrument at 44.1 kHz. 181 A neutral-colored screen behind flutists minimized back- 182 ground interference for accurate f acial mo v ement tracking. 183 3.1.3 Musical Piece 184 This dataset comprises simultaneously starting flute–piano 185 duets. P art 1 included a C major scale and P achelbel’ s 186 Canon performed at slo w (50 BPM), medium (100 BPM) 187 and f ast (150 BPM) tempos, each repeated twice as w arm- 188 up e x ercises (12 trials). P art 2 consisted of 18 classical 189 pieces simplified for piano and arranged for simultaneous 190 starts. Each piece w as assigned a specific tempo (50, 100, 191 or 150 BPM), and the entire set of 18 pieces w as repeated 192 three times, resulting in 54 trials. Each duet performed a 193 total of 66 trials, resulting in 1,320 trials o v erall. Sheet 194 music and sample audio were pro vided in adv ance. 195 3.1.4 Pr ocedur e 196 The recording procedure for each piece included: (1) An 197 e xperimenter’ s clap signaling start , follo wed by a measure 198 of clicks matching the gi v en tempo; (2) The flutist gi ving a 199 cue after clicks ended; (3) The duet be ginning in response 200 to the cue. 201 3.1.5 P ost-session Intervie w 202 The intervie ws collected the insights of the participants 203 on cue strate gies. P articipants reported pro viding cues ap- 204 proximately one beat (or half or tw o beats, depending on 205 the piece) ahead, using body mo v ements or breat h. Some 206 participants mentioned that in typical performance situa- 207 tions, the y adjust their cue timing based on the accompa- 208 n ying instrument and ensemble conte xt. 209 2 Some videos were recorded at 30 fps due to camera o v erheating Figur e 3 : Gestur e Cue Example Gesture cue is defined with the maximum(red) and minimum(black) peaks. The interv al between these peaks, called ‘cue length’ ( x , red line), and the duration between the maximum peak to the flute onset, called ‘cue-onset length’ ( y , blue line). 3.2 Annotation and Pr epr ocessing 210 V ideo and audio data synchronization w as achie v ed 211 through an e xperimenter’ s clap at the s tart. Using the spec- 212 trogram vie wer in Adobe Audition, we m anually annotated 213 six mark ers per trial on the mel spectrogram (Figure 2): the 214 e xperimenter’ s clap ( Start ), breath sound onset and of fset 215 ( Br eath Onset , Br eath Of fset ), initial note onsets of flute 216 and piano ( Flute Onset , Piano Onset ), and the flute’ s sec- 217 ond measure onset ( 2nd Measur e ). 218 4. METHODS 219 4.1 Gestur e Cue Detection 220 4.1.1 Motion Detection 221 T o detect gesture cues from flutists, we used MediaPipe’ s 222 f ace landmark detection 3 [19] to reliably identify f ace re- 223 gions, e v en when partially obscured by the flute. Subse- 224 quently , optical flo w methods [17, 18] were applied to track 225 f acial motion. A pilot study indicated optimal f ace land- 226 mark detection accurac y when the f ace occupied at least 227 50% of the video frame height; videos were accordingly 228 resized. 229 4.1.2 Motion F eatur e Extr action 230 W e analyzed f acial gestures by e xtracting position, v eloc- 231 ity , and acceler ation magnitude curv es from a v eraged y- 232 axis optical flo w v alues. Due to quantized pix el positions 233 causing discrete v elocity curv es, we applied zero-phase fil- 234 tering 4 to smooth the curv e while preserving peak posi- 235 tions. 236 4.1.3 Motion P eak Pic king 237 Figure 3 illustrates a typical gesture cue pattern. W ithin a 238 one-measure windo w preceding the flute onset (‘cue win- 239 do w’), we identified maximum and minimum peaks on 240 position, v elocity , and acceleration curv es using the find- 241 peaks 5 algorithm. This approach automatically det ected 242 3 a v ailable at: https://de v elopers.google.com/mediapipe 4 scip y .signal.filtfilt 5 scip y .signal.find_peaks Proceedings of the 26th ISMIR Conference, Daejeon, Korea, September 21-25, 2025 102 [Document text truncated for crawler view.]