Full text
A conditioned UNet for Music Source Separation Ken O’Hanlon1,2, Basil Woods2, Lin Wang1, and Mark Sandler1 1Centre For Digital Music, Queen Mary University of London 2AudioStrip Ltd. London Abstract. In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument vocabulary. In contrast, conditioned MSS networks accept an audio query related to a stem of interest alongside the signal from which that stem is to be extracted. Thus, a strict vocabulary is not required and this enables more realistic tasks in MSS. The potential of conditioned approaches for such tasks has been somewhat hidden due to a lack of suitable data, an issue recently addressed with the MoisesDb dataset. A recent method, Banquet, employs this dataset with promising results seen on larger vocabularies. Banquet uses Bandsplit RNN rather than a UNet and the authors state that UNets should not be suitable for conditioned MSS. We counter this argument and propose QSCNet, a novel conditioned UNet for MSS that integrates network conditioning elements in the Sparse Compressed Network for MSS. We find QSCNet to outperform Banquet by over 1dB SNR on a couple of MSS tasks, while using less than half the number of parameters. 1 Introduction Music Source Separation (MSS) attempts to separate a signal representing a song into several di↵erent signals containing the stems of individual instruments present in that song. MSS was previously considered a very difficult task with little success beyond specialised cases such as vocal separation [2]. The introduction of deep learning to MSS has resulted in consistent performance gains [27] [9] [7] [3] [8] [24] [19] [18] [28], particularly in the four-stem task in which the system attempts to separate vocals,bass and drum stems alongside a further catchall others category. Separation is often performed by a network with multiple output heads, one per instrument [7] [28], or by employing separately trained networks for each instrument [27] [3] [19]. Most of these networks employ UNetbased architectures [9] [7] [3] [8] [24] [28] that enhance encoder/decoder networks with skip connections between corresponding decoder and encoder modules. A notable exception are bandsplit networks [19] [18] in which disjoint spectrogram frequency bands are encoded, and similarly decoded, in a single block. Bandsplit RNN (BSRNN) [19] shares similarities with other recent high-performance networks, such as a dual-path approach in the neck between encoder and decoder [3] [28] [24], and the learning of complex masks that are applied to the input spectrogram to calculate stem approximations [28] [3]. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 322
One reason for the predominance of the four-stem task has been the existence of MusDb [22], a dataset that is well-formed for the task. Occasionally this task has been augmented with extra stem categories such as guitar [9], or guitar & piano [24]. However, these e↵orts required private data and still used fixed vocabularies while an ambiguity in some stem categories is also noted [24]. Alternatively, conditioned MSS networks [29] [1] [26] [17] [5] [15] [30] possess the potential to avoid such problems through the elimination of hard instrument categories in the network outputs. In such conditioned networks, the input signal is accompanied by an audio query related to the stem intended for separation. Typically, an embedding representation of the audio query is derived and presented to a separation network using a Feature-wise Linear Modulator (FiLM) [21] layer. The FilM layer is placed at some point in a network where it modulates channel activations to emphasise certain features, in this case related to the signal of a particular stem. Similar to the more general MSS problem, UNet architectures have primarily been used for conditioned MSS. Most of these conditioned UNets to date have employed MusDb [22] as the primary dataset [5] [15] [1] [29] [4] while the URMP dataset [16] has also been used with similar architectures [26] [17]. MusDb constrains the possibilities of conditioned MSS as it provides only a small set of four, mostly active, stem categories. URMP does provide a richer stem vocabulary, but the dataset is small and recorded in a homogenous fashion. A recent dataset, MoisesDb [20] consists of multi-track stems, with a stem hierarchy defined with 11 di↵erent stem categories each consisting of subcategories e.g. the guitar category consists of the sub-categories of {clean electric guitar, distorted electric guitar, acoustic guitar}. Having such a rich ontology, MoisesDb allows development of new tasks beyond four-stem separation. Banquet [30], a conditioned MSS network based upon BSRNN [19] is one of the first papers to exploit this new resource, with promising results seen on larger vocabularies. In Banquet, a state of the art music instrument identification network, PASST [14], is used to extract query embeddings. The queries are presented to a FiLM layer [21] placed just before the decoder, similar to [29]. This contrasts to a variety of locations seen in earlier networks; in every decoder block [5], in every encoder block [1] [26], at every convolutional layer [4] [13]. The adaptation of BSRNN in Banquet is based on a rationale that UNets are not suitable for conditioned MSS as they have problems with information flow [30] [31], which are averted with the frequency band split monolithic encoder and decoder layers of BSRNN. In this paper we consider that the assertion of poor information flow in UNets may not hold, as the skip connections in UNets should help information flow. Indeed, most multi-output MSS networks to date are UNets. Therefore, we propose the Query-SCNet (QSCNet), a conditioned variant of the Sparse Compressed Network (SCNet) [28], a UNet architecture that performs similar to BSRNN on the four stem task. We show superior MSS results on MoisesDb for some tasks outlined in the Banquet paper [30], particularly on the 6 stem problem where a very large improvement of 1.6dB SNR is seen. Meanwhile we observe that QSCNet requires only around 40% of the parameters of Banquet. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 323
2 Music Source Separation with UNet Music source separation seeks to separate a musical track into constituent stems that each contain a signal of an instrument or instrument class. Consider a musical signal, y2RC⇥N,withCchannels and Nsamples, and a set of instruments Ithat form the stem vocabulary, The MSS problem can then be defined as finding the set of sources {si2RC⇥N}such that: y⇡X i2I si. This is a difficult problem as the membership and cardinality of stems in Ifor a given piece may be unknown, and the stems typically outnumber the channels, of which there are 2 for stereo recordings. In most MSS e↵orts to date, the set of instruments employed is typically predefined e.g. I4={bass, vocals, drums, others}, and a network Nis applied to a signal leading directly to several corresponding outputs N(y)! {si}i2I(1) although in some cases one network is trained for each instrument Ni(y)! si.(2) The UNet [23] is a commonly used architecture in MSS problems, usually applied to the complex spectrogram: X2CF⇥T=STFT(y). UNets can be considered to comprise several common subnetworks, the encoder NEnc,decoder NDec, and the neck connecting these two subnetworks NNeck. While the original UNet was fully convolutional, variants employed in MSS typically cast the NNeck as an RNN-based[8] [28], or transformer-based [24] module. Similar to standard encoder / decoder architectures, the encoder in the UNet consists of several, L, sequential modules that reduce the feature dimension NEnc =(Nel)l2Lwhere L={1,..,L}while the decoder similarly consists of several modules NDec =(Ndl)l2Lthat increase the feature dimension. In timefrequency domain UNet MSS the encoder and decoder modules typically only alter the dimension in the frequency direction while the size of the temporal dimension remains constant through the network, although there are exceptions to this [3]. A defining feature of UNets is that the encoder and decoder modules with similar representation dimensions are joined by skip connections. Given that el=Nel(el1) describes the output of the lth encoding module, the operation at the (Ll)th decoding module is then described as dl=Ndl(dl1,eLl). The specific case of a mask based MSS UNet can be considered as the sequence of operations E=NEnc(X) N=NNeck(E) {Mi}=NDec(N,{el}l2L) {si}={ISTFT(Mi⌦X)}i2I (3) where Midenotes the complex mask related to the ith stem in Iand E=eL. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 324
2.1 Conditioned Music Source Separation An alternative approach to MSS does not require a fixed stem vocabulary (1) or an ensemble of stem-wise networks (2). Conditioned approaches to MSS employ one network for all possible instruments and supply either a representation of category [25] or an audio query, Q, as input alongside the input signal: N(y,Qi)! si.(4) Such an approach may also be referred to as query-based source separation. While Qcan be a one-hot vector with each dimension representing the activity of a given stem [25], it is more flexible to supply Qas a query embedding that represents a point, or area, in an audio feature space. The query embedding is usually output from a separate neural network such as one trained for musical instrument recognition, and typically derived from the activations in the penultimate layer of the network. A Feature-wise Linear Modulation (FiLM) [21] module is employed to condition the network activations P2RC⇥F0⇥T0at some point in the network based upon a query embedding, Q2Rq P NFilM(Q,P).(5) The FiLM module, NFiLM learns an affine transformation function with parameters o,othat are typically small neural modules that respond to the vector query input Q: =o(Q); =o(Q) (6) with the resultant vectors ,2RCapplied to the input representation Pc,f,t cPc,f,t +c.(7) As well as the query-based modulation, conditioned MSS di↵ers from the multi-stem based approach (3) at the mask formation and output stage. At this point, the set of masks has only one member, M, resulting in only one output signal, s. Nevertheless, multi-stem separation can still be performed in one batch by inputting several queries with one signal and performing post-FiLM processing as a batch. In this way, the computational load at inference time may be reduced as the signal encoding is performed only once. 3 Proposed Approach We propose a conditioned UNet for MSS called Query-SCNet (QSCNet). QSCNet is an adaptation of the Sparse Compressed Network [28] (SCNet) which we embellish with conditioning capabilities. Our rationale for adopting SCNet includes some perceived similarities to BSRNN, from which Banquet is adopted. BSRNN and SCNet are seen to perform similarly in the four-stem MSS task [28], and both possess features found in modern MSS networks such as dual-path RNN and complex mask learning. A further consideration in the selection of SCNet is its relatively small size compared to other state-of-the-art MSS networks [28]. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 325
3.1 SCNet The authors of the original paper [28] do not describe SCNet as a UNet. Rather, they opt to convey an additional fusion layer that accepts similar inputs as a UNet decoder module, (el,dLl) and outputs to the decoder. However, we state here that SCNet can be formulated as a UNet variant (3) simply by aggregating the fusion and decoder layers as they are defined in the paper [28]. In this light, the main point of di↵erence of the SCNet from other UNets is the novel proposed banded downsampling and upsampling modules employed in the encoder and decoder, and a novel dual-path RNN. We briefly outline these here as QSCNet naturally inherits these features. In SCNet the stereo complex spectrogram is first gathered into a tensor with 4 channels. At each time-frequency point in this tensor the 4 channels represent the corresponding complex STFT’s real and imaginary coefficients across the two stereo channels [28]. At each subsequent block of the encoder the inputs are subjected to a coarse banding, with fixed ratios, across the frequency dimension. This results in 3 separate banded tensors, to which di↵erent downsampling ratio and convolutional processing strategies are applied, before they are regathered at the output of the encoder block. Although fixed ratios are used in the banding, it is notable that this does not imply consistent frequency band splitting across the full decoder as the banding is only applied to a dimension with a linear frequency scale in the first downsampling layer. A corresponding recombination strategy is e↵ected in the decoder. The neck of SCNet consists of 6 dual-path bidirectional LSTMs that are processed in alternating fashion. Generally in MSS [3] [19] a dual-path RNN refers to first processing across the temporal dimension using a RNN before similar processing across the frequency dimension. A novel variation is present in the neck of SCNet, where alternating dual-path RNNs operate on di↵erent domains [28]. The first and subsequent odd numbered dual-path RNNs operate directly on the latent feature, while even numbered dual-path RNNs operate on the Fourier representation of the latent feature. Between the oddand evennumbered dual-path RNNs, the real FFT is applied to the feature map, and processing is performed on the latent feature in its frequency domain. After this processing is performed, between the even and odd numbered dual-path RNNs, the inverse RFFT is applied to bring the Fourier-based feature map back into the original latent feature domain [28]. 3.2 Conditioning Embedding-based queries are used for QSCNet. Similar to [30], we employ the PASST network [14] in order to extract embeddings, thereby maintaining similarities between QSCNet and Banquet. Specifically, we used the variant of PASST that was trained on the OpenMic [10] dataset for the instrument recognition task and is freely available from the authors [14]. This version of the network only has 20 outputs representing di↵erent instruments, or instrument categories Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 326
DownSample DownSample DownSample UpSample UpSample UpSample Separator Stereo Signal Query dependent single-instrument Signal Single-instrument audio query Audio embedding generation network ( PASST ) Modulation network ( FilM ) Fig. 1. A schematic diagram of QSCNet, or a generic UNet, with L=3, a PASST embedding generator network, and a FiLM modulator network integrated at end of encoder. and again similar to [30] we remove the final layer and employ the outputs as an embedding Q2R768 given an audio clip of up to 10s. We employed just one FiLM module, which we located at the end of the encoder and directly before the dual-path RNN as can be seen in Fig. 1. Otherwise put, in (5) we set P=E,whereEis the encoder output (3). Although we are aware that this is an unused location for FiLM in conditioned MSS, we consider that this position should be optimal for conditioning as it allows the instrument context to be defined before the sequential long-term processing. We recall that several works have used di↵erent conditioning locations, and have used several FilM modules at di↵erent locations. Perhaps of most interest here, in Banquet the conditioner is placed at the end of NNeck, just before start of the decoder. The FiLM module employed here uses similar networks for the parameters o,o(6) of the affine transformation (7). These networks consist of multi-layer perceptrons with two fully connected layers; the first layer has q= 768 inputs and c= 128 outputs, followed by an ELU activation [6]; the second layer has c= 128 inputs and outputs or , respectively (6). 4 Experiments We ran experiments to test the efficacy of the proposed QSCNet. We trained QSCNet on a 6 stem vocabulary I6={vocals, bass, drums, guitar, piano, others} similar to that employed in[24]. We also propose a six-stem variant of SCNet (SCNet6) which is also trained on the same vocabulary. I6contains the same Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 327
instrument stems as the Q:VDBGP vocabulary employed in [30], which allows a direct comparison with Banquet. However, I6also contains the others category, which allows QSCNet and SCNet6 to be compared more directly. As the conditioned network model can take any input query and return a corresponding output (4) there is a flexibility where, unlike in multi-stem networks, a model trained with one vocabulary can be tested on another. In this light, we also test on an extended variant of I6that includes some finer stems e.g. male vox and female vox rather than just the coarser vocals. 4.1 Data We used MoisesDb [20], which consists of 240 songs with multi-track audio with the training/validation/test splits proposed in [30] with 144/48/48 songs, respectively. MoisesDb uses a hierarchial structure with 11 separate stem categories, each of which may contain finer stem subcategories that number 30 in all. However, unlike 4-stem datasets like MusDb, the data distribution is skewed. [20]. Some stems such as bass, drums and vocals from I4are present in almost every song, while some instrument categories do not appear in all datasplits [30]. We first constructed the training dataset employing the vocabulary I6. For each stem in this category, all its finer subcategory tracks were added to form a single stem track. The others stems were then formed from all remaining tracks not assigned to one of the instrument categories. The validation and test sets were formed in a similar fashion. We employed a random data sampling strategy at training time in which training clips were generated by mixing randomly selected clips for the di↵erent stems. Such cacophonous mixtures of unrelated stems of music have been shown to be superior for training neural networks for MSS [11]. Audio clips of 10s were used to train QSCNet and the 6 stem SCNet. First, a set of candidate clips was assembled for each stem by selecting 10s audio segments spaced 1s apart in each audio track. Each candidate stem clip is accepted into the final pool of clips after an inspection for a simple silence detection that rejects segments in which more than 50% of samples are zero. At train time, an audio clip is randomly selected for each stem category. Some augmentation is applied to the stem clip data. This includes channel flipping in which the stereo channels for an instrument are swapped, and sign flipping where the sign of all elements of a signal are swapped. Both of these flipping augmentations are activated randomly with even chance. A gain augmentation is also applied in which each individual instrument stem segment is subjected to the application of a gain randomly selected between in the range (0.25,1.25). This is similar to the augmentation setup used in the Demucs networks [24] and in SCNet [28]. The training mixture is then formed by mixing the various augmented stem clips together. For the conditioned approach, a pool of queries for each instrument was also formed. Here, this is generated simply by further filtering of the audio clip pools used in training above. Specifically, only audio clips in which less than 20% of samples are zero are employed as queries. In the conditioned training, one Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 328
query instrument is selected from I6with an even chance of each instrument being selected. A random query from the pool is then selected, and input to the network alongside the augmented audio mixture. Similarly, at validation and test time a random query is selected for each instrument from its query pools in the validation or test set, respectively. A test set for I6was formed using the subcategory gathering strategy for stems, as above. An extra test set was formed from the same data using an extended vocabulary, referred to as I6E, that substitutes the vocals, guitar & piano categories with finer stems. Specifically the vocals are split into male & female vocals, the guitar is split into clean electric guitar, distorted electric guitar & acoustic guitar , while the piano is represented by the grand piano & electric piano subcategories. This results in a 10 stem category vocabulary, if the others category is included. 4.2 Parameters Spectrograms used for network inputs were produced from stereo audio files sampled at 44.1kHz using a window size of 4096 with 75% overlap. The root mean square energy (RMSE) was used as a cost function, as for SCNet [28]. Experiments were run on a single A100 chip with 80Gb of memory using a batch size of 8. Some initial experiments were run with batch sizes of 4 and 16 We found that training was sometimes poorer with a batch size of 4, while training was slower, requiring more epochs, when the batch size was 16. In each batch 32000 samples were taken resulting in 4000 mini-batches. The Adam [12] optimizer was employed for training with the learning rate set to 3 ⇥104. Each model was trained for 300 epochs, with validation performed after each epoch. Similar to [24][28] exponential moving averages were maintained after each epoch. The model, or averaged model, that performed best on the validation set was kept as the final trained model. 4.3 Metrics On each track, the signal-to-noise ratio for the ith instrument was calculated SNRi= 10 ⇥log10 kyik2 F kyisik2 F where yiand siare the known and approximated stem signals respectively. For each instrument the median SNR across all tracks of the same instrument is recorded. This is similar to the metric used in [30] which enables a direct comparison of results. 4.4 Results The results on I6are shown in Table 1, where QSCNet is compared with the proposed SCNet6 and its larger variant, SCNet6(L), which has the same architecture as SCNet6 but with double the number of channels in each layer. Here Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 329
Alg Bass Vocals Drums Guitar Piano Avg5 Others Banquet 11.0 8.0 9.5 3.3 2.5 6.9 - QSCNet 11.9 9.8 11.7 5.7 3.4 8.5 1.3 HTDemucs 10.9 8.9 11.6 2.4 1.7 7.1 - SCNet6 12.8 10.5 12.4 6.3 4.0 9.2 2.8 SCNet6(L) 13.5 12.2 13.4 7.0 4.6 10.1 3.4 Table 1. Results comparing 6 stem approaches HTDemucs6, SCNet6 & SCNet6(L), and two conditioned approaches, Banquet [30] and QSCNet for MSS on the MoisesDb test set. Instrumentwise results given in median SNR. they are also compared to state-of-the-art results given in [30] for the six stem HT-Demucs [24] and the conditioned Bandsplit network, Banquet [30] both of which are evaluated for 5 stems only, as the others category was not considered [30]. For each algorithm an average score over the 5 instrument stems (Avg5) is also recorded in order to a↵ord a simple summary comparison. In terms of the multi-output networks SCNet6 is seen to outperform HTDemucs by a large margin of 2.1dB, with improvements for all instruments. Some of these improvements are very large e.g. for the guitar there is an increase of 3.9dB. This is perhaps more impressive when it is considered that SCNet6 is trained only on 144 songs of MoisesDb while HT-Demucs was trained on a private dataset of 800 songs [24]. Further improvements are seen with SCNet6(L) with the Avg5 metric 3dB above the HTDemucs. Large improvements using SCNet6(L) relative to the standard SCNet6 are seen for the vocals and drums categories. Considering the conditioned networks, the proposed QSCNet is seen to be superior to Banquet by a large margin of 1.6dB on the Avg5 metric. Improvements are seen across all instruments when using QSCNet, notably on drums and guitar which both improve over Banquet by more than 2dB. The QSCNet is also seen to improve on the HTDemucs 6 stem, and reach a performance level around 0.7dB lower than that of SCNet6. It is notable that QSCNet uses many less parameters that SCNet6, as many parameters are employed in the formation of the individual masks. We observe that QSCNet uses 10.2M parameters; while SCNet4 and SCNet6 employ 20.4M and 26.6M parameters respectively. QSCNet is also significantly smaller than Banquet which uses 24.9M parameters. Vocals Guitar Piano Drums Bass Male Female Acoust. Clean Dist. Grand Elec. Avg9 Others Banquet 10.1 10.7 7.9 10.1 0.9 1.7 2.8 2.8 0.5 5.3 - QSCNet 11.8 11.6 8.5 11.8 1.3 3.6 4.0 3.2 0.7 6.3 1.1 Table 2. Results comparing Banquet[30] and the proposed QSCNet for MSS on the MoisesDb test set for the extended I6Evocabulary. Instrumentwise results given in median SNR. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 330