The Essence Remains the Same: Generative Modeling of Expressive Percussion
Full text
Master thesis on Sound and Music Computing Universitat Pompeu Fabra “The Essence Remains the Same: Generative Modeling of Expressive Percussion” Anmol Mishra Supervisor: Martin Rocamora, Behzad Haki August 2025
Acknowledgments I would first like to express my deepest gratitude to my supervisors, Martin and Behzad, for their invaluable guidance and support throughout my master’s. I am equally grateful to Robin and Satya, with whom I had the pleasure of collaborating on several projects. I am indebted to my alumni mentor, Justin, whose advice gave me the clarity and conviction to move forward with the audio phase of this project. Special thanks to Hugo and Patrick, whose help was crucial in bringing this work to completion. At the MTG, I am thankful for the entire community, in particular: Pedro, for our countless conversations; Oguz, for his constant guidance; Hyon, for keeping my Korean skills intact; and Adithi, for being a trusted support. I am also grateful to Michael and Genís for their insightful discussions and steady encouragement, and to Esteban and David for being such wonderful friends. My thanks also go to Dmitry and Alastair for the technical tricks they shared during our supervision sessions, and to Lonce and Rafael for always being there when I needed support. Maria deserves a very special mention, not only for our amazing conversations but also for reintroducing me to my high school sweetheart, Physics, which eventually led me to explore diffusion models in depth. I will always cherish our discussions. I would also like to thank my Google Summer of Code supervisors, especially Jörg, for mentoring me on model serialization and export. I am sincerely grateful to Google for its generous TPU Research Cloud program and cloud credits, without which this work would not have been possible. Beyond academia, I am grateful to Mehul, Sahdev, Goka, and Bhavesh, as well as the entire Seoul Music Meetup community. A special thanks to Alexander for teaching me enough sheet music to thrive in this program. My heartfelt thanks also go to the members of the Spotlight DJ Group, who welcomed me wholeheartedly as the sole expat in the group. I especially want to thank President Jeong Jaeyoon for teaching me DJing, and Kevin for his constant support. Two people in particular deserve special thanks. First, Xavier, for everything you have done for me - from giving me the chance to pursue this master’s, to offering me an office at the MTG, and supporting me in various unconventional and deeply 2
meaningful ways. Second, Valerio, whose work truly changed the course of my life. Your Audio Signal Processing for Music course transformed what was a minor curiosity into my true calling. From attending the Generative Music Workshop twice to later serving as a TA, this sequence of events set everything into motion. Finally, I would like to thank my family. Over the years, as I have met people from different walks of life, I have come to realize how fortunate I am to have been raised in such a peaceful and supportive environment. My parents’ unwavering trust in my choices gave me the courage to quit my big-tech job and pursue my dreams. To my parents and my sister, thank you.
Contents List of Figures 6 List of Tables 7 1 Introduction 8 1.1 Motivation and Research Question . . . . . . . . . . . . . . . . . . . . . . 8 1.1.1 ResearchQuestion.............................. 8 1.1.2 Motivation .................................. 9 1.2 Objectives and Contributions . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.2.1 Objectives................................... 9 1.2.2 Contributions................................. 10 1.3 Structure of the Thesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 2 Learning Microrhythm in Uruguayan Candombe using Transformers 12 2.1 Abstract...................................... 12 2.2 Introduction.................................... 13 2.3 RelatedWork................................... 14 2.3.1 Analytical studies of microtiming . . . . . . . . . . . . . . . . . . . . . 14 2.3.2 Symbolic and audio-based representations . . . . . . . . . . . . . . . . 14 2.3.3 Computational modeling and generation . . . . . . . . . . . . . . . . 15 2.4 Dataset....................................... 15 2.4.1 IEMP Candombe Dataset . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.4.2 HVO representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.5 Method....................................... 17 2.6 Experiments.................................... 19 2.6.1 Distribution of Chico onsets in beat . . . . . . . . . . . . . . . . . . . 19 2.6.2 Distribution of microtiming in cycle . . . . . . . . . . . . . . . . . . . 20 2.7 Discussion..................................... 22 2.8 Conclusion and Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3 GotThatFlow: Flow-based Beatbox-to-Drum Generation 25 3.1 Abstract...................................... 25 3.2 Introduction.................................... 26 3.3 RelatedWork................................... 26 3.4 Dataset....................................... 27 3.4.1 Rhythm Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 3.4.2 TRIA Feature Extraction . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4
3.5 Method....................................... 30 3.5.1 Architecture ................................. 30 3.5.2 Conditioning Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.5.3 Pretrained Autoencoder for Audio Latents . . . . . . . . . . . . . . . 32 3.5.4 Diffusion Transformer Architecture for Latent Diffusion (Flow) Modeling ................................... 33 3.5.5 Inference.................................... 34 3.5.6 Training Objective and Conditioning . . . . . . . . . . . . . . . . . . . 35 3.6 Experiments and Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . 36 3.6.1 Rhythm Prompt Adherence . . . . . . . . . . . . . . . . . . . . . . . . 37 3.6.2 Timbre Prompt Adherence . . . . . . . . . . . . . . . . . . . . . . . . . 37 3.7 Discussion..................................... 38 3.8 Conclusion and Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . 38 4 Model Deployment for End Users 40 4.1 Introduction.................................... 40 4.2 Learning Microrhythm in Uruguayan Candombe using Transformers . . 40 4.2.1 Exporting via TorchScript . . . . . . . . . . . . . . . . . . . . . . . . . 40 4.2.2 Wrapping in NeuralMidiFX VST . . . . . . . . . . . . . . . . . . . . . 41 4.2.3 Real-World Deployment . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.3 GotThatFlow: Flow-based Beatbox-to-Rhythm Generation . . . . . . . 42 4.3.1 ONNXExport ................................ 42 4.3.2 Gradio Demo Deployment . . . . . . . . . . . . . . . . . . . . . . . . . 43 4.3.3 Toward a VST Plugin . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 4.4 Conclusion..................................... 44 5 Conclusion 45 5.1 Summary of Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 5.1.1 Learning Microrhythm in Uruguayan Candombe . . . . . . . . . . . . 45 5.1.2 GotThatFlow: Flow-based Beatbox-to-Rhythm Generation . . . . . 45 5.2 Practical Impact and Deployment . . . . . . . . . . . . . . . . . . . . . . . 46 5.3 Broader Perspectives and Future Directions . . . . . . . . . . . . . . . . . 46 5.4 ClosingRemarks ................................. 46 Bibliography 48
List of Figures 1 Modelarchitecture............................... 18 2 Chico pattern of the performance shown in Figure 3 in music notation (the lower line represents the hand, and the upper line represents the stick). The pattern is repeated for each of the four beats of the rhythmiccycle.................................. 20 3 Chico actual (top) vs predicted onsets (bottom) for all beats in one of the performances of the dataset. . . . . . . . . . . . . . . . . . . . . . 20 4 Predicted and actual velocities of chico onsets for all beats of the same performance of Figure 3. . . . . . . . . . . . . . . . . . . . . . . . . 21 5 Example of madera patterns actual vs predicted onsets . . . . . . . . . 22 6 The two madera patterns played by the repique drum in the performance of Figure 5, shown in music notation (×symbol indicates a madera hit). The performance starts with the top pattern and then transitions to the bottom pattern. . . . . . . . . . . . . . . . . . . . . . . 22 7 Architecture of the diffusion transformer (DiT). Cross attention includes text and rhythm conditioning. Prepend conditioning includes timbre conditioning and also the signal conditioning on the current timestep of the diffusion process. . . . . . . . . . . . . . . . . . . . . . . 31 8 Architecture of the autoencoder used for latent representation learning. The encoder maps raw audio waveforms to a compressed latent space, while the decoder reconstructs the audio from this latent representation. ................................... 33 9 Interface of the Candombe VST plugin. The plugin processes incoming MIDI events and applies microrhythmic transformations based on the learned Candombe grooves. . . . . . . . . . . . . . . . . . . . . . . . 41 10 Integration of the Candombe VST plugin within the MTG Toolbox environment. This setup allows users to easily install and utilize the plugin alongside other MTG Toolbox applications . . . . . . . . . . . . 42 11 Deployment of the GotThatFlow model as a Gradio web demo. Users can upload beatbox audio and receive generated drum patterns in real time. ....................................... 43 6
List of Tables 1 Input/output sequence representation for 2-bar beats in 4/4 with 16th note resolution for a total of 32 time steps (i), and 3 drum voices (j). 16 2 Chico actual and predicted mean, standard deviation and histogram intersection of offset distribution across beats computed for the entire dataset. ..................................... 21 3 Chico actual and predicted mean, standard deviation and histogram intersection of offset distribution across rhythmic cycles computed for theentiredataset. ............................... 23 4 Channelwise F1 scores for each model across CFG scales. . . . . . . . . 37 5 Mean timbre similarity for each model across CFG scales. . . . . . . . 38
Chapter 1 Introduction 1.1 Motivation and Research Question 1.1.1 Research Question This thesis focuses on using generative models for capturing the expressiveness in percussion music, the nuances that make human performances unique and engaging, also known as groove. Groove is a natural part of human musical performance. It encompasses the subtle timing deviations, dynamics, and articulations that make a performance feel alive and engaging. As a central aspect of music perception and appreciation, it is closely connected to the main functional uses of music; namely, dance, drill, and ritual. When seeking to find a relationship between music and the behavior that groove induces, synchronization and coordination, the temporal properties of the music signal are crucial to our understanding of groove. [Davies et al., 2013] Generative models have shown great promise in various creative domains, including music generation. They can learn from existing performances and generate new, expressive musical content. They can also learn to mimic specific styles or techniques, allowing for greater control and customization. Percussion music, in particular, offers a rich landscape for exploring expressivity. The nuances in timing, dynamics, and articulation are crucial for conveying the intended feel and groove of a piece. Percussion instruments, with their diverse timbres and playing techniques, present unique challenges and opportunities for generative modeling. The rich variety of sounds and rhythms in percussion music can be difficult to capture and reproduce accurately, but they also provide a wealth of material for training generative models. This diversity is what makes percussion music so compelling, and it is essential to develop methods that can effectively model and generate these intricate patterns. My research focuses on leveraging generative models to capture the nuances of per8
cussion and enhance its expressivity. 1.1.2 Motivation The motivation behind this research lies in the desire to bridge the gap between human expressivity and machine-generated music. By understanding and modeling the nuances of percussion performance, we can create systems that not only generate music but also enhance the creative process for musicians. This has the potential to revolutionize the way we compose, perform, and interact with music technology. While generative models have made significant strides in music generation quality, there is still a gap in the ways we can interact with these systems. Text input methods, while great for high-level descriptions, cannot be used to effectively describe the temporal evolution at a fine granularity. Talking about music often involves discussing its structure, harmony, and melody, but the subtleties of rhythm and timing are more challenging to convey through text alone. Talking about music also requires a shared understanding of its cultural and contextual nuances, which can be difficult to articulate. Music generation systems today are trained on large paired datasets of text and music, but these datasets often lack the depth needed to capture the intricacies of percussion performance. Moreover, the unique characteristics of percussion instruments, such as their diverse timbres and playing techniques, are often underrepresented in these datasets. This underrepresentation can lead to generative models that lack the ability to produce authentic and expressive percussion music. Another motivation is the desire to make these tools more accessible to a wider range of musicians and producers. A large number of these models are trained and deployed on cloud infrastructure, which can be a barrier for many users due to cost, latency, and privacy concerns. Music generation systems that can run efficiently on local hardware would enable more spontaneous and intimate music-making experiences. In addition, if the system can be adapted to the specific musical context and preferences of the user, it could lead to more personalized and meaningful musical experiences. 1.2 Objectives and Contributions 1.2.1 Objectives This research employs a multi-faceted approach to address the challenges identified in the previous sections. 1. Rhythms can be described to the model the way people naturally communicate about them, using a combination of verbal descriptions, visual notations, and audio examples. Verbal descriptions can include terms like "syncopation," "polyrhythm," and "groove," while visual notations can involve standard drum
2.4.1 IEMP Candombe Dataset The corpus consists of 12 performances: nine trios and three quartets. Trio performances feature three channels corresponding to the chico (C), piano (P), and repique 1 (R1) drums. Quartet recordings include an additional repique 2 (R2) channel. For modeling purposes, the quartets were reduced to three channels by discarding the R2 drum, allowing a uniform representation across all performances while preserving the essential rhythmic interactions of the ensemble. Each recording varies in duration, ranging approximately from 150 to 228 seconds per performance. The dataset provides rich temporal and dynamic annotations. Metre annotations are based on manually tapping the first beat of each cycle, establishing the location of downbeats and the overall cycle structure. Onset annotations are available at the level of sixteenth-note subdivisions, including both the precise onset time in seconds and peak amplitude for each instrument. Additional metadata specifies which repique channel is playing the clave versus the repique pattern, event density over the preceding two seconds, and performer identities. This structured information enables detailed analyses of both timing and dynamic interactions between ensemble members. To extract microtiming information, we first construct an isochronous grid for each rhythmic cycle, aligned to the manually tapped downbeats. For each annotated onset, the deviation from the corresponding sixteenth-note subdivision is computed, yielding the microtiming offsets that form the target for learning. Peak amplitude values provide a measure of expressive dynamics and are used in parallel with timing deviations to model the nuanced performance characteristics of each drum. By combining onset timing, dynamics, and ensemble configuration, this dataset allows the models to learn both the temporal and expressive structure of candombe drumming. 2.4.2 HVO representation Data Matrix Values Hits H32×3hij ∈{0,1} Velocities V32×3vij ∈[0,1] Offsets O32×3vij ∈[−0.5,0.5) Table 1: Input/output sequence representation for 2-bar beats in 4/4 with 16th note resolution for a total of 32 time steps (i), and 3 drum voices (j). Following prior work [Gillick et al., 2019, Haki et al., 2022], we represent the annotated candombe performances using the HVO (HitsVelocitiesOffsets) matrix representation, which has proven effective for capturing expressive percussion performance in machine learning models. This representation separates the key dimensions of performance: hits, indicating whether a drum is struck at a particular time step; velocities, encoding the relative strength or loudness of each hit; and offsets, representing the microtiming deviation of each onset from the underlying metrical grid.
By decoupling these aspects, the model can learn not only the rhythmic structure but also the subtle expressive variations that characterize human performances. Each performance is converted into three matrices of size T ×M, where T corresponds to the number of time steps and M to the number of drum channels. In our work, each time step corresponds to a sixteenth-note subdivision, and each matrix has three channels corresponding to chico, piano, and repique (R1, with R2 discarded for quartets). The hits matrix is obtained by quantizing the annotated onset times to the nearest sixteenth-note subdivision, assigning a value of 1 when a note is present and 0 otherwise. The offsets matrix captures the microtiming deviations, normalized to lie between -0.5 and 0.5 relative to a sixteenth-note duration, with negative values indicating anticipations and positive values delays. The velocities matrix encodes the strength of each hit, scaled linearly between 0 and 1 based on the peak amplitude of the onset annotations. Based on prior work [Gillick et al., 2019], we segment the performances into overlapping two-bar sequences, with a stride of one bar between consecutive segments. Given that each bar contains 16 sixteenth-note subdivisions, each HVO sequence has a length of T = 32 time steps. This overlapping segmentation ensures that the model can learn dependencies across bar boundaries while also increasing the total number of training samples. Across the 12 recorded performances, this procedure yields a total of 1,070 two-bar sequences, which form the ground truth dataset for training and evaluation. A summary of the HVO representation is provided in Table 1. 2.5 Method We approach the problem of modeling expressive drum performances as a sequenceto-sequence prediction task. Given a binary hits matrix indicating the presence or absence of hits across multiple drum channels over time, our goal is to predict both the velocity and micro-timing offset for each hit. By framing it this way, we can capture the temporal structure and inter-channel interactions that define a drum performance’s micro-rhythmic feel. To accomplish this, we employ a transformer encoder that takes the hits matrix as input and outputs velocity and offset matrices of the same shape. The architecture is based on the encoder of the transformer from [Vaswani et al., 2017], and is illustrated in Figure 1. We process drum patterns over T=32 time steps, encoding the hit patterns into expressive performance outputs. The encoder uses multi-head selfattention with 4 heads and a model dimension of 128. Feed-forward layers also have dimension 128, and the network consists of 11 stacked encoder blocks. We opt for atransformer encoder rather than a full encoder-decoder because the task involves predicting aligned sequences: each input hit corresponds directly to an output velocity and offset. The self-attention mechanism enables the model to capture both local and long-range temporal dependencies, as well as cross-channel interactions, which are crucial for modeling nuanced microtiming and groove.
Figure 1: Model architecture The model jointly predicts velocities (ˆ V) and timing offsets ( ˆ O). At each time step t, the output is split into two branches, both using tanh activations: one for velocity ˆvtand one for offset ˆot. Ground truth velocities and offsets (V, O) are scaled to (−1,1)to match the output range of tanh. A square error loss is computed at each time step tfor drum channel kas follows: Lt,k =(vt,k −ˆvt,k)2+(ot,k −ˆot,k)2. and mean is computed across all time steps and channels to obtain the final loss. To handle the sparsity of drum hits, we apply the hits matrix as a mask during loss computation, ensuring that only positions with actual hits contribute to the error. The network is trained end-to-end using teacher forcing, and parameters are updated with the Adam optimizer. We found that tanh activations provide better performance than sigmoid, as their range (−1,1)is zero-centered, unlike sigmoid’s (0,1), which helps optimization by
reducing bias in the activations. 2.6 Experiments For a qualitative evaluation of the microrhythms learned by our model, we focus on the musicological structure of candombe drumming. We select a single representative performance from the dataset, which includes multiple drums-Chico, Repique, and Piano-and extract its onset data at a resolution of sixteenth-note subdivisions. From these onsets, we infer velocities and timing offsets using our trained model. Candombe drumming exhibits repeating rhythmic structures at two distinct temporal levels: the beat and the full rhythmic cycle. We analyze the model’s inferences at both scales, evaluating its ability to reproduce the characteristic rhythmic patterns of the original performance. 2.6.1 Distribution of Chico onsets in beat In candombe, the chico assumes the role of the timekeeper, playing repeating patterns at the level of the beat throughout the entire performance (see Figure 2) [Fuentes et al., 2019]. To evaluate whether our model can capture its characteristic microtiming, we compare the distributions of chico onsets at the beat level between the actual performances and the model’s predictions. For this analysis, we employ the carat Python package [Jure and Rocamora, 2019], which allows precise examination of timing deviations across metric subdivisions. Figure 3 shows the distribution of chico onsets for both actual (green) and predicted (red) data at the beat level. Each beat is divided into four sixteenth-note subdivisions, labeled as .1, .2, .3, and .4. We compute the mean offsets at each subdivision as a percentage of the beat duration, providing an intuitive representation of the average microtiming deviations relative to the isochronous metric grid. The results indicate that the model successfully captures the characteristic timing deviations of the chico, reproducing the subtle micro-rhythmic articulations present in the original performance. Table 2 summarizes the mean, standard deviation, and histogram intersection values for each subdivision across the dataset, demonstrating that the predicted onset distributions closely match the ground truth. This confirms that the model effectively learns microtiming patterns unique to the chico, consistent with prior studies on candombe microtiming [Jure and Rocamora, 2016, 2019, Fuentes et al., 2019]. Accents play a central role in expressing groove [Danielsen et al., 2024], and since our model also predicts velocity, we compare actual and predicted velocity distributions at each beat subdivision. Figure 4 illustrates that the model captures the velocity trends observed in the ground truth. Notably, there is a discrepancy between the ground truth velocities and the theoretical pattern depicted in Figure 2, where an accent on the second subdivision is expected but not consistently reflected in the data. This highlights an interesting aspect of expressive performance that may
warrant further investigation. Figure 2: Chico pattern of the performance shown in Figure 3 in music notation (the lower line represents the hand, and the upper line represents the stick). The pattern is repeated for each of the four beats of the rhythmic cycle. Figure 3: Chico actual (top) vs predicted onsets (bottom) for all beats in one of the performances of the dataset. 2.6.2 Distribution of microtiming in cycle We extend our analysis to examine microtiming trends across the duration of the full rhythmic cycle. First, we compute the mean, standard deviation, and histogram intersection of the distributions captured by the model for the chico drum against the ground truth distributions for all subdivisions within the cycle. As shown in Table 3, the distributions repeat every four subdivisions (i.e., every beat), consistent with the beat-level analysis, confirming that the chico pattern is preserved across the cycle. Next, we focus on the madera pattern, which spans the full rhythmic cycle, as illus-
Figure 4: Predicted and actual velocities of chico onsets for all beats of the same performance of Figure 3. Sub Div Mean Std Hist Int Actual Pred. Actual Pred. .1 0.01 0.01 0.02 0.02 0.84 .2 0.25 0.26 0.03 0.03 0.94 .3 0.48 0.49 0.02 0.02 0.81 .4 0.72 0.73 0.02 0.02 0.84 Table 2: Chico actual and predicted mean, standard deviation and histogram intersection of offset distribution across beats computed for the entire dataset. trated in Figure 6. This pattern is initially played by all drums as an introduction and preparation for the rhythm, but during the main performance it is performed solely by the repique drum between phrases [Jure and Rocamora, 2016]. The IEMP candombe dataset provides annotations for sections containing the madera pattern. For this analysis, we consider the same performance used in Section 2.6.1, but focus on the cycles where the repique plays the madera pattern. We analyze 59 cycles of repique madera hits to evaluate whether the model captures cycle-level microtiming patterns. Figure 5 displays the distribution of repique onsets for both the ground truth and the model’s predictions. The model successfully reproduces the characteristic microtiming of the madera pattern. Notably, the onsets at the 4th subdivision of the first and fourth beats (1.4 and 4.4) occur slightly ahead of the isochronous grid, consistent with the expressive deviations observed in the original performance. This demonstrates that our model can learn and replicate not only beat-level timing but also subtler microtiming patterns that emerge across the full rhythmic cycle.
Figure 5: Example of madera patterns actual vs predicted onsets Figure 6: The two madera patterns played by the repique drum in the performance of Figure 5, shown in music notation (×symbol indicates a madera hit). The performance starts with the top pattern and then transitions to the bottom pattern. 2.7 Discussion In this work, we investigated the problem of learning microrhythmic characteristics in Uruguayan candombe drumming. We represented onset timing and strength data as hits, velocity, and offset (HVO) matrices and trained a transformer model on sequences of 2-bar length. After training, the model was used to infer velocities and timing offsets from hit information, allowing us to evaluate its performance at multiple temporal scales. Our qualitative analysis focused on two levels of rhythmic structure: the beat level, using the Chico drum as the timekeeper, and the cycle level, using the Madera pattern played by the Repique drum. At the beat level, we observed that the model accurately captured the microtiming deviations of the Chico drum, reproducing both
Sub Div Mean Std Hist Int Actual Pred. Actual Pred. 1.1 0.01 0.01 0.02 0.02 0.71 1.2 0.26 0.26 0.03 0.03 0.83 1.3 0.49 0.49 0.02 0.02 0.65 1.4 0.73 0.73 0.02 0.02 0.84 2.1 1.01 1.01 0.02 0.02 0.79 2.2 1.25 1.25 0.03 0.03 0.85 2.3 1.48 1.48 0.02 0.02 0.84 2.4 1.73 1.73 0.02 0.02 0.86 3.1 2.01 2.01 0.02 0.02 0.81 3.2 2.25 2.25 0.03 0.03 0.84 3.3 2.48 2.48 0.02 0.02 0.83 3.4 2.72 2.72 0.02 0.02 0.74 4.1 3.01 3.01 0.02 0.02 0.79 4.2 3.26 3.26 0.03 0.03 0.82 4.3 3.48 3.49 0.02 0.02 0.86 4.4 3.72 3.73 0.02 0.02 0.84 Table 3: Chico actual and predicted mean, standard deviation and histogram intersection of offset distribution across rhythmic cycles computed for the entire dataset. the mean and standard deviation of the original distributions. Additionally, the predicted velocity profiles closely matched those of the ground truth, demonstrating that the model is capable of learning not only the timing but also the dynamic accents that contribute to groove in candombe. At the cycle level, the Chico microtiming patterns were preserved across the full rhythmic cycle, reflecting the timekeeper behavior of this instrument. Furthermore, the model successfully reproduced the microtiming of the Madera pattern played by the Repique drum, capturing subtle deviations such as the anticipatory onsets on the 4th subdivision of specific beats. These results demonstrate that the model is capable of learning complex, hierarchical rhythmic structures and expressive timing across multiple temporal scales. 2.8 Conclusion and Future Work The transformer architecture used in this study proves effective for learning microtiming patterns from HVO representations, suggesting that the approach is generalizable to other datasets and musical genres. By capturing both timing and velocity distributions, our model provides a tool for analyzing, generating, and manipulating rhythmic microstructure in a data-driven manner. This has potential applications in algorithmic rhythm creation, music production, and the study of performance practice across different musical traditions.
In conclusion, our work demonstrates that deep learning models can successfully learn the microrhythmic structure of complex rhythmic traditions such as candombe. The model reproduces both timing deviations and dynamic accents at multiple temporal levels, highlighting the capability of data-driven approaches to capture expressive performance characteristics. Future work will focus on extending this methodology to other Latin American music genres, contributing to the development of more versatile tools for rhythm modeling and creative music applications.
Chapter 3 GotThatFlow: Flow-based Beatbox-to-Drum Generation In this chapter, I present the application of diffusion models for generating drum audio conditioned on rhythmic inputs such as beatboxing or tapping. This work has been carried out in collaboration with Patrick O’Reilly and Hugo Flores Garcia from the Interactive Audio Lab at Northwestern University. I have led the implementation of the code and the execution of the experiments, while benefiting greatly from the insightful guidance and suggestions provided by Patrick and Hugo in shaping the methodology and refining the experimental procedures. Throughout the chapter, the pronoun “we” is used to acknowledge this collaborative effort while reflecting my role in leading this work. 3.1 Abstract Musicians frequently rely on intuitive rhythmic gestures, such as tapping or beatboxing, to communicate drum patterns. While these vocal or percussive sketches convey hits and timing effectively, converting them into high quality drum tracks often requires substantial manual effort. To address this, we present GotThatFlow, a flow based generative model that maps raw rhythmic gestures to realistic drum recordings. GotThatFlow is developed over the Stable Audio Open Small model which is fine tuned following the Sketch2Sound paradigm. The model is conditioned using a combination of prepend conditioning and cross conditioning, enabling it to capture both timbral structure and rhythmic variation. Users can supply one audio input that encodes the rhythm (e.g., a beatboxing track) and another that specifies the target drumkit timbre. The result is a system capable of producing rhythmically coherent drum performances from unseen timbres in a zero shot manner. 25
(2) Timbre Conditioning via Latent Overwrite. Timbre control is achieved through a direct latent overwriting mechanism. Given a timbre prompt (a reference drum recording), we pass the audio through the SAO autoencoder to obtain latent embeddings that capture the target drumkit’s timbral characteristics. During training, a randomly selected prefix of the diffusion latent sequence (between 20 - 50%) is overwritten with the corresponding segment from the timbre latent. A binary mask is constructed to mark the overwritten positions, ensuring that no reconstruction loss is computed over these regions. This procedure constitutes a form of explicit latent conditioning, which circumvents the complexity of concatenation based methods, while robustly enforcing inheritance of timbral properties from the reference prompt. By combining these two conditioning pathways, TRIA-based additive rhythm control and latent overwrite timbre conditioning, our model learns to disentangle rhythmic structure from timbral style, enabling flexible generation of drum performances from simple sound gestures. Importantly, unlike transcription based approaches, this framework does not rely on symbolic intermediates, but rather leverages selfsupervised conditioning directly in the audio domain. 3.5.3 Pretrained Autoencoder for Audio Latents At the core of the latent generative framework lies a pretrained autoencoder, provided by SAO, which maps raw waveforms into a perceptually meaningful and temporally compressed representation. The encoder consists of five convolutional blocks, each performing strided convolutions for downsampling while simultaneously expanding the number of channels. Prior to each downsampling stage, the model applies a sequence of residual layers with dilated convolutions and Snake activations [Ziyin et al., 2020], which improve the representation of oscillatory signals and fine temporal structures. Figure 8 illustrates the architecture of the autoencoder. The bottleneck of the autoencoder is parameterized as a variational latent space with dimensionality 64, enabling stochastic sampling and regularization of the latent manifold. The decoder mirrors the encoder in structure, employing transposed strided convolutions to progressively upsample while reducing the channel dimension. This architecture yields a 64-channel latent representation operating at a temporal resolution of 21.5 Hz. Operating in this low rate latent space substantially reduces the computational burden of downstream generative modeling, while preserving high-quality reconstructions. The pretrained autoencoder comprises approximately 156 million parameters and is frozen during all subsequent training stages in our system.
Figure 8: Architecture of the autoencoder used for latent representation learning. The encoder maps raw audio waveforms to a compressed latent space, while the decoder reconstructs the audio from this latent representation. 3.5.4 Diffusion Transformer Architecture for Latent Diffusion (Flow) Modeling The generative backbone of the model is a Diffusion Transformer (DiT) [Evans et al., 2024], which extends the transformer paradigm to the diffusion setting. At its core, the DiT is composed of stacked transformer blocks, each containing serially connected self attention and gated multi layer perceptrons (MLPs), with residual skip connections around each sublayer. Bias-less layer normalization is applied before both the attention and MLP components, stabilizing training and improving generalization. Rotary positional embeddings [Su et al., 2023] are applied to half of the attention keys and queries, providing relative position encoding. To incorporate conditioning, each block includes a cross attention mechanism. Conditioning signals include text, timing, and diffusion timesteps. Text features are extracted via a pretrained T5-base encoder [Raffel et al., 2023], and timestep embeddings follow sinusoidal encodings [Ho et al., 2020]. Conditioning signals are introduced through a combination of cross attention and prepended embeddings, with text applied in cross attention, and timestep prepended to the input sequence.
Linear mappings are applied both at the input and output of the transformer to project between the autoencoder latent space and the transformer’s embedding dimension. For efficiency, block wise attention [Dao et al., 2022] and gradient checkpointing [Chen et al., 2016] are employed, reducing memory and compute requirements. The specific variant adopted here builds on the SAO framework [Evans et al., 2025], but with architectural modifications to improve efficiency while retaining generation quality. The DiT operates on the 64-channel latent representations produced by the pretrained autoencoder, and is conditioned on 109M-parameter T5 embeddings [Raffel et al., 2023]. Compared to the original 1.06B-parameter DiT, the model reduces the embedding dimension from 1536 to 1024, and the depth from 24 to 16 layers, while additionally incorporating QK-LayerNorm [Henry et al., 2020] and removing the “seconds start” embedding. These adjustments reduce the parameter count to 340M while maintaining synthesis quality and improving training stability. During inference, the base DiT is compiled using torch.compile, yielding further gains in runtime efficiency. 3.5.5 Inference At inference time, our system requires two inputs: a timbre prompt in the form of a short drum recording, and a rhythm prompt provided as a sound gesture (e.g., a tapping pattern, beatboxing, or another percussive input). The timbre prompt is first passed through the Stable Audio autoencoder to obtain its latent representation, which we use to initialize the prefix of the generation buffer. This prefix is left unmasked and remains fixed throughout the process, ensuring that the timbral identity of the final output faithfully matches the given recording. The remaining portion of the buffer, corresponding in length to the rhythm prompt, is fully masked and designated as the suffix to be generated. Rhythm features are then computed from the rhythm prompt and temporally aligned to this masked suffix, providing the model with explicit rhythmic conditioning. Generation proceeds through an iterative denoising procedure based on Euler sampling, carried out over 50 discrete steps. At each step, the model progressively refines the masked suffix latents, gradually transforming them from noise into coherent representations consistent with the provided rhythm. After every denoising update, the original timbre prefix is reinserted into the buffer before continuing to the next step. This replacement step is crucial: without it, the model could inadvertently alter the timbral prompt during sampling, drifting away from the intended drum identity. By continually restoring the prefix, we guarantee that the model is always conditioned on the correct timbre context while generating the suffix. The result is a sequence of latents in which the suffix is coherently “filled in” using both sources of conditioning: timbral characteristics from the prefix and rhythmic structure from the computed TRIA features. The tradeoff between strict adherence to conditioning and creative variability is controlled by classifier-free guidance (CFG). In our implementation, the CFG scale can be adjusted to emphasize either
timbre or rhythm conditioning, or to allow for looser interpretations that produce more diverse outputs. To make this process accessible to end users, we developed a Gradio interface in which the number of sampling steps and the CFG scale are exposed as adjustable parameters. This allows musicians and producers to experiment with different configurations, from faithful reconstructions to more imaginative generations. 3.5.6 Training Objective and Conditioning For training, we follow the flow matching framework [Lipman et al., 2023], where noise is applied to the encoded audio latents and the model is trained to denoise them. Let x0∈RF×Ddenote the encoded audio latents of dimension Dwith F=256 latent frames. We sample a timestep τ∼U(0,1)and construct noisy latents by convex combination with Gaussian noise: xτ=(1−τ)x0+τϵ,ϵ∼N(0,I).(3.1) The model is trained to predict x0from xτusing the rectified flow objective, i.e., LRF =Ex0,ϵ,τ [∥ˆ x0(xτ,τ,c)−x0∥2 2],(3.2) where ˆ x0(⋅)denotes the model prediction and care conditioning signals. Note that loss is computed only over the suffix portion of latents (outside the timbre prefix span). Training At each iteration, we begin by sampling a drum recording from the MusDB training set and extracting a random 11.89 second segment of the isolated drum track. This segment is passed through the pretrained Stable Audio autoencoder (SAO) to obtain a sequence of continuous latents, which serve as the reconstruction target. Flow matching noise is then applied to the latent sequence, and the model is trained using the rectified flow objective, with mean squared error (MSE) computed between predicted and target latents [Chang et al., 2022]. To condition generation, we incorporate both timbre and rhythm information in complementary ways. To avoid redundant information, rhythm features are zeroed out in the prefix region so that the model relies solely on timbre latents there, while in the suffix region the model must reconstruct the target latents by combining timbre information from the prefix with rhythmic cues from TRIA. To improve generalization to rhythm prompts from diverse sources such as tapping, beatboxing, or low quality recordings, we apply a set of augmentations to audio when computing rhythm features. These include additive Gaussian noise, pitch shifting, and high or low-pass filtering. Importantly, these augmentations are never applied to the audio used for encoding target latents, ensuring that the model always learns to predict clean latents while adapting to noisy or degraded rhythm conditioning. Finally, we employ classifier-free guidance (CFG) [Ho and Salimans, 2022] dropout
to control the degree of adherence to conditioning. During training, rhythm conditioning is independently disabled in 10% of iterations by replacing their embeddings with null prompts, enabling the model to learn both conditional and unconditional mappings. At inference time, we apply CFG with a scale of 2, interpolating between unconditional and conditional predictions to balance timbre prompt adherence with rhythm prompt adherence. 3.6 Experiments and Evaluation We developed the system through a series of experimental refinements: Initial model with no CFG control This was our first attempt at implementing the core architecture and training procedure. Even though the loss went down, the generated samples were of low quality and did not adhere well to either timbre or rhythm prompts. There was no mechanism to balance the two conditioning sources, leading to outputs that were often muddled or incoherent. Incorporation of classifier-free guidance (CFGScale) Adding CFG improved our ability to control the influence of timbre and rhythm prompts. At very high values of CFG scale (5-7), the model produced outputs that started to adhere to the rhythm prompt, but the timbre was often lost or distorted at high CFG scales. One notable observation was that the model struggled to maintain a consistent timbral identity when heavily conditioned on rhythm. Switching controls projection from latent space to DiT embedding space (DiTProj) Instead of adding our conditioning signals directly in the latent space, we introduced linear projections to map both the TRIA features and the timbre latents into the DiT’s embedding space. This change significantly improved the model’s ability to integrate conditioning information, leading to better adherence to both prompts. The outputs became more coherent, with clearer rhythmic structures and more faithful timbral characteristics, even at normal CFG scales. Data augmentations to enhance generalization to real world gestures (DataAug) To ensure the model could handle a variety of rhythm prompts, we applied augmentations such as noise addition, pitch shifting, and filtering when computing TRIA features. This step was crucial for improving robustness, as it exposed the model to a wider range of rhythmic inputs during training. The augmented training led to better performance on unseen rhythm prompts, particularly those derived from beatboxing or tapping. We evaluate the model from each stage on the MusDB test set to assess its performance. Evaluation focuses on two key aspects: Rhythm Prompt Adherence and Timbre Prompt Adherence.
3.6.1 Rhythm Prompt Adherence To assess how well our models preserve rhythmic information from the prompt, we adopt an automatic drum transcription based evaluation, on the held out test split of MusDB (50 tracks), which was not used for training. For each track, we consider the ground truth drum stem as a reference and generate corresponding drum outputs from our system. Both the ground truth stems and the model outputs are transcribed using the pretrained Frame-RNN drum transcription model [Zehren et al., 2023], yielding symbolic onset sequences for kick, snare, and hi-hat. We measure the correspondence between the two transcriptions using the onset F1 score at 30ms tolerance, as is standard practice in drum transcription evaluation [Vogl et al., 2017, Heydari et al., 2021]. Higher F1 values indicate closer alignment between the temporal placement of drum hits in the reference stem and the generated output. Model CFG F1 Kick F1 Snare F1 HiHats 1 0.00 0.00 0.00 CFGScale 2 0.00 0.00 0.00 3 0.05 0.01 0.02 1 0.25 0.30 0.00 DiTProj 2 0.43 0.46 0.16 3 0.43 0.52 0.24 1 0.59 0.43 0.15 DataAug 2 0.82 0.65 0.38 3 0.84 0.66 0.37 Table 4: Channelwise F1 scores for each model across CFG scales. 3.6.2 Timbre Prompt Adherence To evaluate timbre preservation, we again draw on the MusDB test set drum stems as source material. For each evaluation instance, we supply the model with a timbre prompt (a drum stem excerpt) and a rhythm prompt. The goal is to determine whether the generated output retains the spectral and textural qualities of the timbre prompt, while remaining unaffected by the rhythm prompt in terms of timbral content. We quantify timbre similarity using feature based measures computed on short time spectral representations. Specifically, we compare the 80 dimensional temporally averaged MFCC representations of the generated audio and timbre prompt via cosine similarity. A higher similarity indicates stronger timbral adherence.
Model CFG 1 CFG 2 CFG 3 DiTProj 0.94 0.95 0.96 DataAug 0.94 0.96 0.98 Table 5: Mean timbre similarity for each model across CFG scales. 3.7 Discussion Tables 4 and 5 summarize the results of our evaluations. Rhythm prompt adherence, as measured by onset F1 scores, shows a clear progression across model variants. The initial CFGScale model exhibits minimal rhythmic fidelity, with F1 scores near zero across all channels. The DiTProj variant demonstrates substantial improvements, particularly at higher CFG scales, indicating that projection space modifications significantly enhance the model’s ability to follow rhythmic cues. The DataAug model achieves the highest F1 scores, with values exceeding 0.8 for kick and 0.65 for snare at CFG2. Hi-hat F1 scores remain lower overall, reflecting the inherent difficulty of accurately capturing high frequency, rapid percussion elements. Timbre similarity remains consistently high across all models, with mean cosine similarities above 0.94. It must be noted that timbre evaluations here are restricted to drum sounds present in the MusDB dataset, in future work we aim to quantify timbre generalization to out of distribution drum sounds. Nevertheless, our listening tests also surfaced important limitations. While classifierfree guidance (CFG) and rhythm augmentations improved performance, timbre generalization is imperfect, particularly for rare drum sounds or heavily processed stems in the MusDB dataset. Similarly, transcription based evaluation introduces its own sources of error, as the accuracy of onset detection directly impacts rhythm adherence scores. Evaluation protocols that integrate perceptual listening tests or learned embeddings more closely aligned with timbre may yield more reliable assessments. Beyond the scope of quantitative metrics, qualitative experiments showed that the system can recombine rhythmic and timbral cues in musically compelling ways, often producing plausible and stylistically coherent drum tracks. This validates our broader vision of generative rhythm timbre transfer: enabling musicians to impose a desired groove structure onto a chosen drum sound, much like a producer layering the feel of one drummer with the kit of another. However, further work is needed to refine the model’s ability to handle extreme or out of distribution inputs, as well as to explore user interfaces that facilitate intuitive control over the generation process. 3.8 Conclusion and Future Work Several avenues emerge for future work. First, more robust timbre similarity measures, potentially based on pretrained audio embeddings, could complement MFCC cosine metrics and capture perceptual qualities beyond spectral energy. Second, user centric evaluation will be critical: perceptual studies with musicians and pro-
ducers could reveal how controllable and musically useful such systems truly are. Finally, extending this framework beyond drum synthesis toward other instrument classes may generalize the prefix suffix conditioning paradigm as a powerful tool for controllable audio generation.
Chapter 4 Model Deployment for End Users While the core of this thesis has focused on the design, training, and evaluation of machine learning models for modeling expressive rhythms, an equally important contribution lies in the exporting and deployment of these models into usable tools. Making research artifacts available to musicians and practitioners outside the lab is essential to bridging the gap between academic research and artistic practice. 4.1 Introduction In this chapter, we describe the process of exporting and deploying the models developed in the two independent works presented in this thesis: (1) Learning Microrhythm in Uruguayan Candombe using Transformers, and (2) GotThatFlow: Flow-based Beatbox-to-Rhythm Generation. For both projects, we detail the export formats used, the deployment contexts (plugins and demos), and the reception or current status of these deployments. 4.2 Learning Microrhythm in Uruguayan Candombe using Transformers 4.2.1 Exporting via TorchScript The trained Transformer model for microrhythm generation was exported using TorchScript, a serialization format provided by PyTorch that allows models to be run independently of Python. TorchScript enables integration into C++ environments and plugin frameworks, which is crucial for deployment in digital audio workstations (DAWs). Unlike ONNX or other intermediate representations, TorchScript provides a relatively seamless path for PyTorch-trained models to be used in C++ without introducing third-party runtime dependencies. This reduces deployment friction and 40
Figure 9: Interface of the Candombe VST plugin. The plugin processes incoming MIDI events and applies microrhythmic transformations based on the learned Candombe grooves. improves stability within DAW environments. The PyTorch model was traced using torch.jit.trace and subsequently saved as a TorchScript module. The exported module was validated against the Python version to ensure parity of results. 4.2.2 Wrapping in NeuralMidiFX VST To make the model accessible to musicians, the TorchScript model was integrated into the NeuralMidiFX [Haki et al., 2023b] VST wrapper. This wrapper provides a framework for embedding neural models inside VST plugins. Within this setup, the plugin accepts incoming MIDI input and outputs rhythmically transformed MIDI events, applying microrhythmic deviations learned from Candombe performances. The plugin interface is shown in Figure 9. Musicians can load the plugin in their preferred DAW (Ableton Live, Logic Pro, etc.) and directly enrich MIDI sequences with authentic Candombe grooves.
Bibliography Guillaume Alain, Maxime Chevalier-Boisvert, Frederic Osterrath, and Remi PicheTaillefer. DeepDrummer : Generating Drum Loops using Deep Learning and a Human in the Loop, August 2020. Jeffrey Adam Bilmes. Timing Is of the Essence : Perceptual and Computational Techniques for Representing, Learning, and Reproducing Expressive Timing in Percussive Rhythm. Thesis, Massachusetts Institute of Technology, 1993. Grigore Burloiu. Adaptive Drum Machine Microtiming with Transfer Learning and RNNs. In Extended Abstracts for the Late-Breaking Demo Session of the 21st Int. Society for Music Information Retrieval Conf., 2020. Antoine Caillon and Philippe Esling. RAVE: A variational autoencoder for fast and high-quality neural audio synthesis, December 2021. Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. MaskGIT: Masked Generative Image Transformer, February 2022. Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training Deep Nets with Sublinear Memory Cost, April 2016. Martin Clayton, Kelly Jakubowski, Tuomas Eerola, Peter E. Keller, Antonio Camurri, Gualtiero Volpe, and Paolo Alborno. Interpersonal Entrainment in Music Performance: Theory, Method, and Model. Music Perception, 38(2):136–194, November 2020. ISSN 0730-7829. doi: 10.1525/mp.2020.38.2.136. Anne Danielsen, Ragnhild Brøvig, Kjetil Klette Bøhler, Guilherme Schmidt Câmara, Mari Romarheim Haugen, Eirik Jacobsen, Mats S. Johansson, Olivier Lartillot, Kristian Nymoen, Kjell Andreas Oddekalv, Bjørnar Sandvik, George Sioros, and Justin London. There’s More to Timing than Time: Investigating Musical Microrhythm Across Disciplines and Cultures. Music Perception, 41(3):176–198, February 2024. ISSN 0730-7829. doi: 10.1525/mp.2024.41.3.176. Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, June 2022. Matthew Davies, Guy Madison, Pedro Silva, and Fabien Gouyon. The Effect of 48
Microtiming Deviations on the Perception of Groove in Short Rhythms. Music Perception, 30(5):497–510, June 2013. ISSN 0730-7829. doi: 10.1525/mp.2013.30. 5.497. Nils Demerlé, Philippe Esling, Guillaume Doras, and David Genova. Combining audio control and style transfer using latent diffusion. In Proceedings of the 25th Int. Society for Music Information Retrieval Conference. arXiv, July 2024. doi: 10.48550/arXiv.2408.00196. Jesse Engel, Lamtharn (Hanoi) Hantrakul, Chenjie Gu, and Adam Roberts. DDSP: Differentiable Digital Signal Processing. In International Conference on Learning Representations, September 2019. Zach Evans, Julian D. Parker, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Long-form music generation with latent diffusion, July 2024. Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable Audio Open. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, April 2025. doi: 10.1109/ICASSP49660.2025.10888461. Luis Ferreira. An Afrocentric Approach to Musical Performance in South Black Atlantic: The Candombe Drumming. Trans : Transcultural Music Review = Revista Transcultural de Música, ISSN 1697-0101, Nº. 11, 2007, January 2007. Anders Friberg and Andreas Sundström. Swing Ratios and Ensemble Timing in Jazz Performance: Evidence for a Common Rhythmic Pattern. Music Perception, 19(3):333–349, March 2002. ISSN 0730-7829. doi: 10.1525/mp.2002.19.3.333. Magdalena Fuentes, Lucas S. Maia, Martín Rocamora, L. Biscainho, H. Crayencour, S. Essid, and J. Bello. Tracking Beats and Microtiming in Afro-Latin American Music Using Conditional Random Fields and Deep Learning. In International Society for Music Information Retrieval Conference, 2019. Alf Gabrielsson. Interplay between Analysis and Synthesis in Studies of Music Performance and Music Experience. Music Perception, 3(1):59–86, October 1985. ISSN 0730-7829. doi: 10.2307/40285322. Hugo Flores García, Oriol Nieto, Justin Salamon, Bryan Pardo, and Prem Seetharaman. Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, April 2025. doi: 10.1109/ICASSP49660.2025.10888184. Jon Gillick, Adam Roberts, Jesse Engel, Douglas Eck, and David Bamman. Learning to Groove with Inverse Sequence Transformations. In Proceedings of the 36th International Conference on Machine Learning, pages 2269–2279. PMLR, May 2019.
Daniel Gómez-Marín, Sergi Jordà, and Perfecto Herrera. Network representations of drum sequences for classification and generation. Frontiers in Computer Science, 6, January 2025. ISSN 2624-9898. doi: 10.3389/fcomp.2024.1476996. Behzad Haki, Marina Nieto, Teresa Pelinski, and Sergi Jordà. Real-Time Drum Accompaniment Using Transformer Architecture. In Proceedings of the 3rd International Conference on on AI and Musical Creativity. AIMC, September 2022. doi: 10.5281/ZENODO.7088343. Behzad Haki, Cheuk Lun Isaac Lee, and Sergi Jordà. Taptamdrum: A Dataset for Dualized Drum Patterns. In Proceedings of the 24th International Society for Music Information Retrieval Conference, 2023a. Behzad Haki, Julian Lenz, and Sergi Jorda. NeuralMidiFx: A Wrapper Template for Deploying Neural Networks as VST3 Plugins. AIMC 2023, August 2023b. Holger Hennig, Ragnar Fleischmann, Anneke Fredebohm, York Hagmayer, Jan Nagler, Annette Witt, Fabian J. Theis, and Theo Geisel. The Nature and Perception of Fluctuations in Human Musical Rhythms. PLOS ONE, 6(10):e26457, October 2011. ISSN 1932-6203. doi: 10.1371/journal.pone.0026457. Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen. QueryKey Normalization for Transformers, October 2020. Mojtaba Heydari, Frank Cwitkowitz, and Zhiyao Duan. BeatNet: CRNN and Particle Filtering for Online Joint Beat Downbeat and Meter Tracking. In 22nd International Society for Music Information Retrieval (ISMIR) Conference. arXiv, August 2021. doi: 10.48550/arXiv.2108.03576. Jonathan Ho and Tim Salimans. Classifier-Free Diffusion Guidance, July 2022. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Diffusion Probabilistic Models, December 2020. Vijay Iyer. Embodied Mind, Situated Cognition, and Expressive Microtiming in African-American Music. Music Perception, 19(3):387–414, March 2002. ISSN 0730-7829. doi: 10.1525/mp.2002.19.3.387. Luis Jure and Martín Rocamora. Microtiming in the rhythmic structure of Candombe drumming patterns. In Fourth International Conference on Analytical Approaches to World Music, New York, USA, June 2016. Luis Jure and Martín Rocamora. Subir la llamada: Negotiating tempo and dynamics in Uruguayan Candombe drumming. In International Workshop on Folk Music Analysis, June 2018. Luis Jure and Martín Rocamora. Carat: A toolbox for computer–aided rhythm analysis. In Analytical Approaches to World Music Special Topics Symposium (AAWM 2019). Zenodo, July 2019. doi: 10.5281/ZENODO.10030090.
Olivier Lartillot and Fred Bruford. Bistate Reduction and Comparison of Drum Patterns. In Proceedings of the 21st International Society for Music Information Retrieval Conference, 2021. Stefan Lattner and Maarten Grachten. High-Level Control of Drum Track Generation Using Learned Patterns of Rhythmic Interaction. In IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA 2019). arXiv, August 2019. doi: 10.48550/arXiv.1908.00948. Antoine Lavault, Axel Roebel, and Matthieu Voiry. StyleWaveGAN: Style-based synthesis of drum sounds using generative adversarial networks for higher audio quality. In 30th European Signal Processing Conference (EUSIPCO 2022), Belgrade, Serbia, August 2022. Kyungyun Lee, Wonil Kim, and Juhan Nam. PocketVAE: A Two-step Model for Groove Generation and Control, July 2021. Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow Matching for Generative Modeling, February 2023. Anmol Mishra, Behzad Haki, Satyajeet Prabhu, and Martin Rocamora. Groove Transfer VST For Latin American Rhythms. In Extended Abstracts for the LateBreaking Demo Session of the 25th Int. Society for Music Information Retrieval Conf., San Francisco, November 2024. Anmol Mishra, Satyajeet Prabhu, Behzad Haki, and Martín Rocamora. Learning Microrhythm in Uruguayan Candombe using Transformers. In Proceedings of the International Computer Music Conference (ICMC), Boston, 2025. Luiz Naveda, Fabien Gouyon, Carlos Guedes, and Marc Leman. Microtiming Patterns and Interactions with Musical Properties in Samba Music. Journal of New Music Research, 40(3):225–238, September 2011. ISSN 0929-8215. doi: 10.1080/09298215.2011.603833. Javier Nistal Hurlé, Stefan Lattner, and Gael Richard. DrumGAN: Synthesis of drum sounds with timbral feature conditioning using Generative Adversarial Networks. In 21st International Society for Music Information Retrieval Conference (ISMIR), Toronto, Canada, August 2020. Zachary Novack, Zach Evans, Zack Zukowski, Josiah Taylor, C. J. Carr, Julian Parker, Adnan Al-Sinan, Gian Marco Iodice, Julian McAuley, Taylor BergKirkpatrick, and Jordi Pons. Fast Text-to-Audio Generation with Adversarial Post-Training, May 2025. Patrick O’Reilly, Hugo Flores Garcia, Prem Seetharaman, and Bryan Pardo. Masked Token Modeling for Zero-shot Anything-to-drums Conversion. In Extended Abstracts for the Late-Breaking Demo Session of the 25th Int. Society for Music Information Retrieval Conf.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, September 2023. Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. MUSDB18 - a corpus for music separation. December 2017. doi: 10.5281/zenodo.1117371. António Ramires, Rui Penha, and Matthew E. P. Davies. User Specific Adaptation in Automatic Transcription of Vocalised Percussion, November 2018. Martín Rocamora, Luis Jure, Bernardo Marenco, Magdalena Fuentes, Florencia Lanzaro, and Alvaro Gomez. An audio-visual database of Candombe performances for computational musicological studies. In Congreso Internacional de Ciencia y Tecnología Musical, September 2015. Vincent Rosinach and Traube Caroline. Measuring swing in Irish traditional fiddle music. In International Conference on Music Perception and Cognition, 2006. André C. Santos and F. Amilcar Cardoso. From Taps to Drums: Audio-to-audio Percussion Style Transfer. In Extended Abstracts for the Late-Breaking Demo Session of the 22nd Int. Society for Music Information Retrieval Conf, 2023. William Sethares. The geometry of musical rhythm: What makes a “good” rhythm good? Journal of Mathematics and the Arts, 8:135–137, December 2014. doi: 10.1080/17513472.2014.906116. Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced Transformer with Rotary Position Embedding, November 2023. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. Richard Vogl, Matthias Dorfer, and Peter Knees. Drum transcription from polyphonic music with recurrent neural networks. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 201–205, March 2017. doi: 10.1109/ICASSP.2017.7952146. I.-Chieh Wei, Chih-Wei Wu, and Li Su. Generating Structured Drum Pattern Using Variational Autoencoder and Self-similarity Matrix. In International Society for Music Information Retrieval Conference, 2019. Mickaël Zehren, Marco Alunno, and Paolo Bientinesi. High-Quality and Reproducible Automatic Drum Transcription from Crowdsourced Data. Signals, 4(4): 768–787, December 2023. ISSN 2624-6120. doi: 10.3390/signals4040042. Liu Ziyin, Tilman Hartwig, and Masahito Ueda. Neural Networks Fail to Learn Periodic Functions and How to Fix It, October 2020.