Non-Negative Matrix Factorization Deconvolution for Multi-Instrument Transcription of Solo Drum Performances
Abstract
Automatic drum transcription (ADT) is a key sub-problem of music information retrieval. While most ADT systems focus on a small set of drum instruments, such as kick, snare, and hi-hat, many practical applications require transcription of full drum kit performances. This paper presents a multi-instrument solo drum transcription method based on Non-negative Matrix Factorisation Deconvolution (NMFD). The system requires only minimal prior data collected through a soundcheck procedure. We compare various NMF-based methods and show that NMFD offers competitive performance with a low computational cost and high interpretability, making it suitable for real-world scenarios.
Full text
Non-Negative Matrix Factorization Deconvolution for Multi-Instrument Transcription of Solo Drum Performances Mikko Hakila1and Kjell Lemström2 1Master’s Programme in Computer Science, University of Helsinki, Finland [email protected] 2Department of Computer Science, University of Helsinki, Finland [email protected] Abstract. Automatic drum transcription (ADT) is a key sub-problem of music information retrieval. While most ADT systems focus on a small set of drum instruments, such as kick, snare, and hi-hat, many practical applications require transcription of full drum kit performances. This paper presents a multi-instrument solo drum transcription method based on Non-negative Matrix Factorisation Deconvolution (NMFD). The system requires only minimal prior data collected through a soundcheck procedure. We compare various NMF-based methods and show that NMFD offers competitive performance with a low computational cost and high interpretability, making it suitable for real-world scenarios. Keywords: Automatic Drum Transcription ·Non-negative Matrix Factorization ·Deconvolution ·Music Information Retrieval ·Audio Signal Processing 1Introduction Automatic drum transcription (ADT) is a central problem in music information retrieval (MIR), aiming to convert an audio recording of a drum performance into asymbolicrepresentationofnoteonsetsandinstrumentlabels.WhileearlyADT systems focused mainly on transcribing three core instruments—kick drum, snare drum, and hi-hat—many real-world applications require the transcription of full drum kit performances that include toms, cymbals, and auxiliary percussion. These applications span creative music production, performance analysis, music pedagogy, and interactive educational systems. Recent advances in machine learning have led to the development of powerful drum transcription models based on deep neural networks (DNNs), particularly those utilising convolutional and recurrent architectures. These methods have shown excellent performance in various benchmark datasets. However, they often rely on large amounts of annotated data, are computationally expensive, and offer limited interpretability. Their black-box nature can be a disadvantage in settings where transparency, user control, and customizability are essential. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1005
2 M. Hakila and K. Lemström In contrast, this work presents a transcription system based on Non-negative Matrix Factorisation Deconvolution (NMFD). This signal decomposition technique models audio as additive combinations of spectrotemporal templates and their temporal activations. Unlike recent drum transcription systems, which behave as high-performing but opaque black-box models, the proposed method offers full interpretability. Each component of the decomposition has a direct correspondence to musical events. This transparency enables users to understand, modify, and extend the system, making it suitable for research and real-world applications, including those with limited computational resources or training data. The system requires minimal prior data, collected through a short soundcheck procedure, and uses this to construct fixed spectral templates for each drum instrument. These templates are then used with NMFD to perform transcription on solo drum audio. The method is evaluated on standard benchmark datasets and real-world recordings, and its performance is compared with that of other NMF-based approaches. This paper is organised as follows: Section 2 first describes the Non-negative matrix factorisation process, then reviews related work on drum transcription, including NMFand deep learning-based approaches. Section 3 presents the proposed NMFD-based transcription system. Section 4 describes the experimental setup and evaluation datasets. Section 5 presents the results and analysis, and Section 6 concludes the paper with a discussion of potential applications and future work. 2Background 2.1 Non-Negative Matrix Factorisation Approaches Non-negative Matrix Factorization (NMF) presents the following problem: given the non-negative M⇥Nmatrix X,findthefactorizationtononnegative matrix factors Wand H,thatminimizesthecostfunctionDassociated with the factorization: min W,H D(X||WH),(1) where X2R0,M⇥N,W2R0,M⇥Rand H2R0,R⇥N, where the rank RM. The most commonly used cost functions are: Euclidean Distance (EU) (=2), Kullback-Leibler Divergence (KL) (=1)and (=0). All of these cost functions belong to a subset of Bregman divergences known as beta-divergences [4] that are defined as follows: D=8 > < > : x ylog x y1,=0 xlog x y+yx, =1 x+(1)yxy1 (1) ,2R\{0,1}, (2) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1006
Transcription of Solo Drum Performances 3 where xand ydenote individual elements of the input matrix Xand its approximation Y=WH, respectively. The divergence is evaluated element-wise and then summed over all matrix entries. In literature, there are several approaches to updating the mixing matrices Wand Hiteratively to find the local minimum divergence [2,3]. Lee and Seung [7] proposed Multiplicative Updates (MU) for EU and KL cost functions H H⌦W>(X/(WH)) W>1,(3) W W⌦(X/(WH))H> 1H>,(4) where ⌦denotes a Hadamard product (an element-wise multiplication), the division is also performed element-wise. The unity matrix 1is a matrix of ones in the appropriate dimensions. 2.2 Related Work Automatic drum transcription has evolved from simple rule-based systems to complex machine learning models. A significant shift occurred with the introduction of deep learning, where convolutional and recurrent neural networks (CNNs and RNNs) became the dominant architectures for modelling drum events in polyphonic music. Several studies have reported high transcription accuracy using deep networks trained on large annotated datasets [14,12,15]. Enhancements such as soft attention mechanisms [12] and joint beat-drum modelling [14] have further improved performance. Despite these advances, deep learning-based systems often behave as black boxes, providing limited insight into the decision-making process. They also typically require substantial amounts of labelled training data and computational resources. These limitations motivate the exploration of interpretable and lightweight methods, particularly in scenarios where transparency and adaptability are desired. Non-negative Matrix Factorisation (NMF) and its extensions offer an alternative, data-efficient approach. In NMF, an input spectrogram is decomposed into non-negative spectral templates and their corresponding activation functions. Variants such as Partially Fixed NMF (PFNMF) [16] and Semi-Adaptive NMF (SANMF) [5] improve performance by constraining or updating templates during optimisation. Non-negative Matrix Factor Deconvolution (NMFD) [11] extends NMF to incorporate temporal dynamics, making it especially well-suited for modelling percussive sounds with time-localised energy patterns. Recent variants such as Multi-layer NMFD [6] and Itakura-Saito-based formulations [10] further refine the model. In addition to pure NMF methods, hybrid systems have been proposed. These include neural models with NMF-inspired constraints [13] and frameworks that use NMF to improve the interpretability of neural classifiers [9]. Multi-pass Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1007
4 M. Hakila and K. Lemström strategies [1] and template adaptation [16] have also been explored for source separation and transcription tasks. While many of the above methods are designed for polyphonic music or mixtures involving multiple instruments, solo drum transcription, especially with a complete drum kit, is less frequently addressed. Moreover, few systems prioritise low-latency, real-time applicability with minimal training data. The method presented in this paper contributes to this space by offering a fully interpretable NMFD-based transcription system that operates on solo drum input using fixed templates collected via a soundcheck procedure. 3Method The proposed system for solo drum transcription combines fixed-template NMFDbased source separation with simple yet effective onset detection techniques. The approach is designed to be interpretable, computationally efficient, and usable with minimal prior data. An overview of the whole processing pipeline is presented in Fig. 1. The system begins with a short soundcheck session, during which isolated hits are recorded and used to construct spectral templates for each drum instrument. These templates are then used in the NMFD stage to separate instrument activations from complete performance recordings. Onset detection is performed directly on the resulting activation functions, followed by peak picking and formatting of the output. The following subsections provide a detailed description of each stage. 3.1 Preprocessing The source audio is transformed to time-frequency representation applying Short Time Fourier Transform (STFT) with a 211 sample frame length, 29sample hop length to a 44100 Hz sample rate source audio. The STFT is then smoothed with a one-frame-sized Hann window and half-wave rectified. The resulting 210 ⇥nSTFT is then filtered with a 48-band, 48-band Bark scale, non-overlapping rectangular filter bank to yield a 48 ⇥nSTFT, which is once more smoothed with a four-frame-long Hann window. 3.2 Onset Detection Two different Onset Detection Function (ODF)sareused.FirstaLogarithmic Spectral Difference ODF is used for the soundcheck step when the Onset Detection (OD) is performed on the soundcheck samples. LSD ODF with l1-norm is defined as ODFSF (n)= N 21 X k=N 2 Hlog filt(||X(n, k)||X(n1,k)||),(5) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1008
Transcription of Solo Drum Performances 5 Input audio soundcheck Prior templates Preprocessing NMFD with templates Activation matrix H Onset detection Transcription output Fig. 1. Flow diagram of the proposed NMFD-based drum transcription system using fixed templates and onset detection. where X(k,n)is frequency bin kin the frame n,Nthe window size and Hlog filt denotes half wave rectification of the logarithmic magnitude scaled STFT. The second ODF is used after the source separation step; it is simply the activations produced by the Non-negative Matrix Factorization Deconvolution (NMFD) algorithm, as described in Section 3.5. The activations are sharp peaks with a low noise floor, and therefore, they are suitable for peak picking. 3.3 Peak Picking Asimplepeak-pickingschemeisappliedtothe glsodf,selectingalllocalmaxima above a fixed threshold as onset candidates. Onset candidates are regarded as onsets if the ODF falls under the threshold between consecutive candidates, and the candidates are at least wframes apart. Empirical testing supported selecting w=3for this application. The threshold is set separately for each drum as described in section 3.4. 3.4 Prior Templates For NMF methods, a soundcheck is performed to create a set of prior templates W. These templates are evaluated from audio files recorded separately before analysing a drum performance. In the soundcheck, the drummer records nhits Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1009
6 M. Hakila and K. Lemström per drum. From these audio parts, the onsets are discovered by annealing the threshold until exactly nonsets are found. After this step, the templates are calculated for each drum by averaging the frequency response at the onset locations. For NMFD,eachtemplateisaveragedfromtheonsetlocationsandthe consecutive nine frames. Drum-specific thresholds drum for OD are determined by a simple hill-climbing algorithm that starts with a low and calculates the f-score for every iteration as increases. The search material is the NMFD separated soundcheck audio, where all separate soundcheck audio is concatenated to form a file containing nhits per drum. The optimal threshold range can be wide if the drum hits are separated well and the resulting Activations ODF has a low noise floor and sharp peaks. The optimal threshold range Ris the value range where the f-score is the highest. To set a good threshold that captures most hits but does not capture too many false positives, the threshold is moved up from the low end of the optimal range drum =↵⇤min(R)+1↵⇤max(R).(6) Based on conducted experiments, ↵=0.55 seems to be a good value. 3.5 Non-Negative Matrix Factorization Deconvolution NMFD was proposed by Paris Smaragdis [11]. The reasoning behind deconvolution was that a musical signal’s frequency response alters with time. One could achieve better approximations than regular NMF by considering also the temporal behaviour of frequency information of an event in a spectrogram. To achieve this, a time dimension is added to our prior basis matrix, and the basis matrix becomes a three-dimensional matrix W2R0,M⇥R⇥T, where Tis the number of frames stored in the templates. In other words, we extract a slice of time as our template, rather than a single moment on the timeline. The objective is once again to find Wand Hso that they approximate Xas well as possible. With the larger templates, the approximation becomes X⇡ T1 X t=0 Wt⇤t! H,(7) where X2R0,M⇥Nis the input we wish to decompose, and Wt2R0,M⇥R and H2R0,R⇥Nare the prior basis matrix at time tand the activations matrix. t! His a shift operation that shifts the columns of Htsteps to the direction of the arrow. The cost function for NMFD is according to Smaragdis [11] using the KL divergence: DKL = X⇤log X ⇤X+⇤ F ,(8) and divergence: DIS = X ⇤log X ⇤1 F ,(9) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1010
Transcription of Solo Drum Performances 7 where k.kFis the Frobenius norm, and ⇤=PT1 t=0 Wt⇤t! H.Tooptimisethe approximation, we use a similar approach as in NMF,butaddtheiteration over tand the shift operations to Lee & Seung MU for H H H⌦W> t t X ⇤ W> t1.(10) If a sparsity constraint µis used to control the trade-offbetween factorisation accuracy and sparseness [8], the update rule 10 becomes H H⌦W> t t X ⇤ W> t1+µ.(11) The proposed Automatic Drum Transcription (ADT) system follows the standard NMFD approach proposed by Smaragdis [11]. An is used as the cost function, and a stopping criterion stops processing when the cost decrease per iteration falls below a threshold. Basis matrix Wis fixed, and only His updated at every iteration. A sparsity constraint [10] is also used. A KL cost function performs equally well, but the divisions are already calculated in the MU step, so it seems convenient to use it. The update for the activations matrix His H H⌦PT tW> t t ˆ ⇤ PT tW> t1+µ,(12) where ˆ ⇤=X ⇤and µis the sparseness constraint. The cost function is DIS =X 2ˆ ⇤ ˆ ⇤log ˆ ⇤1,(13) and the stopping criterion error_Difference =|error[i]error[i1]| error[0] error[i].(14) We stop iterating if error_Difference falls below a preset threshold. There is also an upper limit on the number of iterations to control the maximum running time. The soundcheck audio is added to the analysis signal as a dummy target, ensuring that at least one onset of each drum present in the signal is detected. This prevents normalisation errors from occurring when zero onsets are in the signal for some particular drum. The onsets originating from the soundcheck audio are not included in the final transcription. 4 Experiments The experiments evaluate various NMF-based transcription methods using the IDMT-SMT-Drums (SMT) dataset [5]. This dataset contains 104 drum set recordings of kick, snare, and hi-hat instruments recorded in isolation and combination. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1011
8 M. Hakila and K. Lemström method precision recall f-score NMF tails 0.8906 0.8975 0.8793 NMF 0.8822 0.9083 0.8796 SANMF 0.8862 0.9113 0.8859 PFNMF 0.7993 0.7736 0.7590 NMFD tails 0.9342 0.9253 0.9205 NMFD 0.8932 0.9351 0.8991 SA-NMFD 0.9022 0.9316 0.9028 NMFD tails, full priors 0.9115 0.8400 0.8538 Table 1. Comparison of NMF methods. We compare the baseline NMF, Semi-Adaptive NMF (SANMF), Partially Fixed NMF (PFNMF), and Non-negative Matrix Factor Deconvolution (NMFD) variants with and without additional tail templates. All methods were evaluated using fixed prior templates derived from isolated hits. A ±20 ms tolerance window was used to match detected onsets with ground truth. Additional real-world recordings were conducted using a consumer laptop microphone and electronic and acoustic drum kits to further validate robustness. Although the RMBA13 and ENST datasets were considered, they were ultimately excluded from this study due to time constraints and annotation inconsistencies. Table 1 summarises averaged precision, recall, and F-score over 30 runs per method. 5Results The results presented in this section expand on the experimental setup described in Section 4. We evaluated each method over 30 runs, and the averaged precision, recall, and F-score are shown in Table 1. All the results align with Wu et al. [15], even though the correction of systematic error significantly improves the results of NMF methods. NMFD methods are superior to NMF. NMFD variants benefit from using tail templates, whereas NMF methods suffer from it. The semi-adaptive approach of [5] improves both NMF and NMFD performance slightly; PFNMF performs worse than standard NMF. PFNMF was not implemented as an NMFD extension due to time restrictions and preliminary evidence from NMF testing. Using all the training material drum hits for the basis matrix yielded considerably worse results, as expected, since the average frequency response of a drum would then include overlapping instrument frequencies. These findings support standard NMFD with a fixed basis matrix and additional tail basis vectors. This paper does not include results of Non-negative Vector Decomposition (NVD), as it differs significantly from the other tested methods. In addition to benchmark dataset evaluations, the system was tested using areal-worldusecasescenario.DrumpartswererecordedusingalaptopmicroProc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1012
Transcription of Solo Drum Performances 9 phone, which represents a low-budget or home environment. An electronic drum kit was used to create a MIDI annotation as ground truth for fine-tuning, while an acoustic drum kit was used for general testing. These recordings serve to validate the system’s performance under less controlled conditions. For the prior templates, xdrum hits per drum dwere recorded. For NMF, the prior is computed as the mean of the frequency response fat onset locations, Wd=1 xPx 0f0,x.ForNMFD,theframesfromanonsetonwardarestoredasa prior template vector, Wd(0...T)= 1 xPx 0f(0...T ),x.Anannotatedhitisregarded as a true positive if it is within ±25 ms of the ground truth. All actual use case audio was recorded with the laptop3microphone, without external signal processing. The recording level was set as low as possible to minimise distortion, but the microphone remained uncovered and within close range of the drums. The optimal number of frames for averaging the NMF templates varies with the recording conditions. For the test drum kit, a single frame proved optimal. Compared to the IDMT-SMT-Drums (SMT) dataset experiments, the NMF results degrade significantly as the kit size increases to nine instruments. Although the material is well separable due to the soundcheck procedure, the performance declines. Increasing the acceptable hit window to 50 ms improves the f-score to 0.65; a one-frame shift further enhances it to 0.63. However, this is more of an observation than a feasible improvement. The number of template frames in NMFD has a significant impact on performance. Too few frames fail to capture decay, while too many increase runtime and may introduce interference from tail components. A compromise of ten frames per template proved effective. While drum-specific template lengths could offer finer control, the longest template dictates the length for all instruments due to the fixed-size requirement in matrix multiplication. Preliminary tests showed that recall and runtime degrade with longer templates, although precision improves. The separate and detect approach with NMFD effectively isolates drum signals from drum-only audio. Even with a nine-piece drum kit, results exceeded expectations. A basic Python implementation transcribed a 90-second drum performance in 4 seconds on the test machine. Although the system is not intended for real-time use, the latency is low enough to make such applications conceivable. Reducing latency remains a goal for future research, especially in interactive or time-sensitive scenarios. 6Conclusions This study addressed automatic drum transcription (ADT), a crucial task in automatic music transcription, with a focus on accurately transcribing percussive instruments, particularly within Western music styles and complete drum kits. While most prior research has centred on three core instruments — bass drum, 32015 MacBook Pro with 2.2 GHz i7 four core processor, 16 GB of memory and an SSD. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 1013