scieee AI-readable full text Open interactive document viewer

A Fine Tuning Strategy to Improve Musical Source Separation Quality for Indian Carnatic Music

Schweinitz, Serafin

Abstract

The computational analysis of Carnatic music from audio remains a field of high research interest due to the genre’s rich melodic and rhythmic complexity. However,despite the availability of large multitrack collections such as Saraga, the liverecorded nature of this repertoire leads to a scarcity of truly clean instrument andvocal stems, posing significant challenges for both musicological and technological studies. State-of-the-art music source separation (MSS) models perform poorly onCarnatic music due to a pronounced domain mismatch with their training data. This work proposes a fine-tuning strategy for improving separation of vocals, mridangam, and violin plus tanpura stems in Carnatic music. The approach uses aSparse Compression U-Net (SCNet) pretrained on MusDB18, extended with a curated training set combining clean Carnatic multitrack recordings and out-of-domaindata. To further reduce the domain gap, three data augmentations are introduced: (i) violin sampling augmentation, (ii) microphone-bleeding simulation, and (iii) room impulse response convolution. The proposed model achieves substantial SDR improvements over the baselines on a clean Carnatic benchmark derived from the Sanidha dataset, and a perceptualevaluation on Saraga confirms significant quality gains on all 3 separated sources. On the benchmark, the best configuration outperforms all baselines by a large marginin SDR, while training in under two days on a single 40GB GPU - making it considerably less resource-exhaustive than many similar deep learning-based MSS domain adaptation methods. All pretrained models, code, a cleaned version of the Saraga dataset, and the Sanidha benchmark are released alongside this work.

Full text

Master in Sound and Music Computing Universitat Pompeu Fabra A Fine Tuning Strategy to Improve Musical Source Separation Quality for Indian Carnatic Music Serafin Schweinitz Supervisor: Martín Rocamora Co-Supervisors: Adithi Shankar, Genís Plaja-Roglans July 2025 Contents 1 Introduction 1 1.1 Challenges in Carnatic Source Separation . . . . . . . . . . . . . . . . . 1 1.2 Motivation and Objectives . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2.1 CarnaticMusicology ............................ 2 1.2.2 Objectives.................................. 5 1.2.3 Carnatic Instrumentation . . . . . . . . . . . . . . . . . . . . . . . . . 6 1.3 Structure of the Report . . . . . . . . . . . . . . . . . . . . . . . . . . 8 2 State of the Art 10 2.1 Carnatic Music Source Separation . . . . . . . . . . . . . . . . . . . . . 10 2.2 MicrophoneBleeding............................ 11 2.2.1 Bleeding Aware Source Separation . . . . . . . . . . . . . . . . . . . . . 13 2.3 Bleeding Unaware Source Separation . . . . . . . . . . . . . . . . . . . 16 2.4 Data-Domain Adaptation Challenges in Music Source Separation . . . 17 2.5 U-Nets for Musical Source Separation . . . . . . . . . . . . . . . . . . . 18 2.5.1 (Hybrid) (Transformer) Demucs . . . . . . . . . . . . . . . . . . . . . . 20 2.5.2 SCNet.................................... 21 2.6 Datasets................................... 22 2.6.1 Saraga.................................... 22 2.6.2 MUSDB18.................................. 22 2.6.3 Sanidha ................................... 23 2.6.4 Carnatic Multi-stem Clean (CMC) . . . . . . . . . . . . . . . . . . . . 23 2.6.5 BachViolinDataset ............................ 24 2.6.6 Deep Noise Suppression Dataset (DNS) . . . . . . . . . . . . . . . . . . 24 3 A Finetuning Strategy - Methodology 25 3.1 Data Domain Investigation: Carnatic Music vs. Western Pop and Rock 26 3.1.1 Instrumentation Differences and Challenges in Separating Carnatic Music with Out-of-Domain MSS Models . . . . . . . . . . . . . . . . . . . 26 3.1.2 Genre-Specific Challenges in Carnatic MSS . . . . . . . . . . . . . . . . 27 3.1.3 Tailoring a Training Set and Benchmarks to Improve Carnatic MSS . . 28 3.2 DataAugmentations ............................ 30 3.2.1 Aligning SCNet Data Augmentations with the Carnatic Data Domain . 30 3.2.2 Violin Data Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . 31 3.2.3 Bleeding Augmentations . . . . . . . . . . . . . . . . . . . . . . . . . . 32 3.3 ExperimentalSetup............................. 34 3.3.1 FinetuningonCMC ............................ 34 3.3.2 Finetuning on CMC plus MUSDB18 . . . . . . . . . . . . . . . . . . . 34 3.3.3 Finetuning on CMC plus MUSDB18 with Violin Augmentation . . . . 34 3.3.4 Finetuning on CMC plus MUSDB18 with Violin and Bleeding Augmentations ................................. 35 3.4 Evaluation.................................. 36 3.4.1 SDR..................................... 36 3.4.2 Perceptual Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 3.4.3 Baselines................................... 37 4 Results 39 4.1 Improving the MSS quality for carnatic music . . . . . . . . . . . . . . 39 4.1.1 SDREvaluation............................... 39 4.1.2 On the Effectiveness of the Proposed Data Augmentations . . . . . . . 40 4.2 Perceptual Evaluation: SCNetc,m,v vs HTDemucsft ........... 41 4.2.1 Mridangam Separation . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.2.2 VocalSeparation .............................. 43 4.2.3 Violin and Tanpura Separation . . . . . . . . . . . . . . . . . . . . . . 44 4.3 Training Stability: Extending vs Replacing the Fine-tuning Dataset . . 46 5 Conclusion and Discussion 48 List of Figures 50 List of Tables 51 Bibliography 52 A Appendix 59 Abstract Abstract The computational analysis of Carnatic music from audio remains a field of high research interest due to the genre’s rich melodic and rhythmic complexity. However, despite the availability of large multitrack collections such as Saraga, the liverecorded nature of this repertoire leads to a scarcity of truly clean instrument and vocal stems, posing significant challenges for both musicological and technological studies. State-of-the-art music source separation (MSS) models perform poorly on Carnatic music due to a pronounced domain mismatch with their training data. This work proposes a fine-tuning strategy for improving separation of vocals, mridangam, and violin plus tanpura stems in Carnatic music. The approach uses a Sparse Compression U-Net (SCNet) pretrained on MusDB18, extended with a curated training set combining clean Carnatic multitrack recordings and out-of-domain data. To further reduce the domain gap, three data augmentations are introduced: (i) violin sampling augmentation, (ii) microphone-bleeding simulation, and (iii) room impulse response convolution. The proposed model achieves substantial SDR improvements over the baselines on a clean Carnatic benchmark derived from the Sanidha dataset, and a perceptual evaluation on Saraga confirms significant quality gains on all 3 separated sources. On the benchmark, the best configuration outperforms all baselines by a large margin in SDR, while training in under two days on a single 40 GB GPU - making it considerably less resource-exhaustive than many similar deep learning-based MSS domain adaptation methods. All pretrained models, code, a cleaned version of the Saraga dataset, and the Sanidha benchmark are released alongside this work. Keywords: Musical Source Separation, Carnatic music 6Chapter 1. Introduction Figure 2: A spectral analysis of the Mridangam strokes conducted by [15]. •Provide a bleeding-reduced version of the Saraga AV dataset to support future research in Carnatic musicology and audiovisual source separation. 1.2.3 Carnatic Instrumentation The remainder of this chapter provides a brief introduction to Carnatic instrumentation. Figure 6 depicts a Carnatic concert from the Saraga Audiovisual dataset. In the performance, a violin follows the improvised vocal line, while a mridangam player provides rhythmic accompaniment. Carnatic instruments can be broadly categorized into three groups: melodic instruments, rhythmic instruments, and drones. Melodic instruments include the violin and the Saraswati Veena. Drum instruments such as the Mridangam and Ghatam provide rhythm and percussive textures. Finally, drone instruments like the Tanpura create a wall-of-sound drone that typically remains in the background. Mridangam The mridangam is the primary percussion instrument, the main instrument used in Carnatic music concerts to keep the performance in a rhythmic pattern[4]. It is mentioned in historical manuscripts as far back as 200 B.C. and has gradually 1.2. Motivation and Objectives 7 developed into the most prominent percussion instrument played in South Indian classical music.[15]. The mridangam is a double-headed drum, see Figure 1.3(a). It is played using several different strokes on the treble (right) and bass (left) membrane. Different finger positions and methods of striking the drumheads produce a variety of tones, with some strokes producing harmonic sounds with a recognizable pitch, and some producing tonic-independent sounds[16]. These strokes, many of which got names over time, create a vocabulary of different sound-types. The authors of [15] categorize these strokes into 3 classes: •Ringing string-like tones played on the treble membrane –Dhin, Cha or Bheem •Flat, closed, crisp sounds. –Thi, Ta and Num •Resonant strokes played on the bass membrane –Thom Figure 2 shows spectrogram representations of the different stroke types introduced above. The presence of horizontal lines in the spectrograms indicates the harmonicrich nature of mridangam decays. In particular, the strokes Bheem,Cha, and Dhin - played on the treble side of the mridangam - exhibit long, sustained harmonic fade-outs. This stands in contrast to classic western drum kits, whose percussive elements typically lack discernible harmonics in their spectral representations. Tanpura Tanpura is a multi-stringed and fretless accompanying drone instrument extensively used in classical music in India. Instrumentalists create an underlying drone background sound by plucking the Tanpura by finger. Jitter, shimmer and complexity perturbations are found also in tanpura signals[17]. Figure 1.3(b) shows a photo of a Tanpura. 8Chapter 1. Introduction (a) Mridangam (b) Tanpura[18] (c) Violin Figure 3: The 3 carnatic instruments separated in this work Violin Violin-like string instruments have been used in indian classical music, before the western violin got introduced in india around the 18th century. Carnatic musicians play western violins with different posture and tuning - generally much lower and with different intervals. For a tonic Cthe four strings would be tuned to C3 - G3 - C4 - G4. The instrument furthermore is played in g¯ayaki style (see Section 1.2.1), usually follows the vocalists improvisation. In contrast to western violin, it follows a different scale with smaller intervals than the semitone. The violin also plays Gamakas - analog to vibrato or glissando in western music. Figure 1.3(c) illustrates that the Carnatic violin is build identical to its Western counterpart. 1.3 Structure of the Report This report is structured as follows: The state-of-the-art (Chapter 2) situates this work within prior research I accompanied at the MTG -Music Technology Group, Barcelona - on Carnatic source separation and bleeding-aware source separation. Furthermore, important developments in both areas are discussed, along with related works on source separation domain adaptation and a presentation of the neural networks used in this project. The methodology (Chapter 3) presents the proposed fine-tuning strategy. First, the data-domain mismatch between the Carnatic domain and Western pop/rock 1.3. Structure of the Report 9 music - represented by MUSDB18 - is analyzed. Next, I introduce a fine-tuning dataset and data augmentations tailored specifically to bridge these domain gaps. The experimental setup (Chapter 3.3) describes the four different configurations of the proposed models, the baselines and the evaluation procedure of the experiments. Finally, results, Chapter 4, presents the evaluation results followed by conclusion and discussion (Chapter 5) where they are summarized and critically discussed. Chapter 2 State of the Art This chapter provides an overview of the current state of research in Carnatic source separation by: •Surveying relevant literature and highlighting the link to research in the field of bleeding source separation due to the carnatic MSS datasets available •Introducing research works on similar source-separation domain adaptation problems •Presenting the architectural developments in source separation neural networks with a focus on models used within this work •Reviewing available datasets for Carnatic MSS and related collections that serve as the foundation for this thesis 2.1 Carnatic Music Source Separation Several studies have noted that publicly available pre-trained source separation models often exhibit poor generalization to Carnatic music. [19, 20, 21, 22]. Due to the limited availability of source-leakage free carnatic multistem audios, most literature as well as my own preceding research focuses on leveraging the bleeding-containing 10 2.2. Microphone Bleeding 11 Figure 4: Microphone bleeding[24] saraga stems to improve performance of carnatic source separation systems. Even if these systems didn’t yet outperform baselines like Hybrid Transformer Demucs [23], trained on clean out-of-domain data, significant improvements over baselines trained solely on the saraga bleeding stems have been made. These results suggest that source separation models are capable of deriving single-source information from training data that contains inter-source bleeding. This indicates that, in data domains where microphone leakage is common, bleeding-aware approaches are likely to outperform models that do not explicitly account for such interference in the future. The following chapter introduces microphone-bleeding and different bleeding-aware and unaware approaches for carnatic MSS. 2.2 Microphone Bleeding Microphone bleeding or Microphone leakage describes a phenomena of live multi-microphone recordings with multiple instruments/sound sources[24]. Figure 4 provides a simplified illustration of microphone bleeding. In a recording scenario with Nsound sources sn(k)in a reverberant environment, Mmicrophones capture signals denoted as xm(k). Let hmn(k)represent the room 12 Chapter 2. State of the Art impulse response modeling the acoustic path from source nto microphone m. Assuming that each microphone is primarily intended to capture a single target source, the signal at microphone mcan be expressed as: xm(k) = sm(k)∗hmm(k) + M X n=1 n=m sn(k)∗hmn(k) Then, the direct source, the convolution of the target source sm(k)with its direct path hmm(k)is defined via the first part of the equation: ˆsm(k) = sm(k)∗hmm(k) The second term, accounting for the contributions from all other sources due to room reflections and cross-talk, defines the bleeding component: um(k) = M X n=1 n=m sn(k)∗hmn(k) Neural network based MSS systems commonly minimize reconstruction losses like L1-distance or mean-square-error in a supervised training way requiring stem-data as groundtruths. If stems contain leakage, the model therefore learns to average a bleeding components ˆum(k), derived from all the bleeding components of the target sources in the dataset. Additionally, evaluation of the training success becomes increasingly challenging with the bleeding level of the dataset. Standard MSS evaluation metrics such as SDR (see Section 3.4.1) and ISR are designed to assess the overall signal quality and the level of interference between stems. However, these metrics can become misleading in bleeding scenarios. For example, if the target source is almost silent during a specific interval in the benchmark, the separation output may still contain a residual mixture signal due to learned bleeding patterns. As a result, metrics like SDR and ISR may yield highly negative scores, as they interpret the presence 2.2. Microphone Bleeding 13 of any residual energy as distortion or interference, despite the model reproducing the typical leakage observed during training. The following paragraphs summarizes developments in bleeding-aware approaches for carnatic MSS that rethink training objectives to tackle the aforementioned challenges. 2.2.1 Bleeding Aware Source Separation Recent work by Plaja et al. [22] introduces bleeding-aware techniques to improve source separation performance for Carnatic singing voice using diffusion based approaches. However, these approaches still face two key limitations: the overall separation quality does not yet match state-of-the-art results in the broader literature, and current efforts focus exclusively on separating the singing voice, neglecting other important sources. In a following paper Disentangling Overlapping Sources: Improving Vocal and Violin Source Separation in Carnatic Music [25], written by A. Shankar, myself, G. Plaja and M. Rocamora, we proposed a bleeding-aware two-stage training procedure targeting both voice and violin stems. This approach draws inspiration from advances in speech denoising and enhancement, where learned loss functions that approximate perceptual metrics - such as PESQ or STOI - have been shown to improve separation quality [26, 27]. For instance, Wuxuan et al. [28] train a dense convolutional neural network to predict PESQ and STOI scores, both of which are meant to correlate with the perceived cleanness of speech signals. During training, a random PESQ target value ypesq is sampled, and noise xnoise is scaled and added to a clean speech signal xspeech such that: ypesq =PESQ(xspeech, xspeech +s·xnoise) The model then learns to map the noisy input xspeech +s·xnoise to the target perceptual score ypesq. 14 Chapter 2. State of the Art In [25], we adapted this framework to the bleeding problem by simulating artificial bleed for Carnatic voice and violin stems. The leakage component is modeled following an algorithm introduced in the SDX Bleeding Challenge 2023 [29]. The provided benchmark and training dataset are generated via simulated inter-source leakage on the multi-stem sources from MUSDB18. The authors simulated bleeding following these assumptions: •The amount of bleeding in a recording is usually low –Authors define the bleeding level via source separation evaluation, aiming for a difference of 1db in SDR when models train on the bleeding containing data, compared to the same model, trained on clean MUSDB18 •Every single file contains bleeding •Every stem bleeds into every other stem in the same song •The bleeding component of one stem for another stem is obtained by: –Apply gain reduction, random between -7db and -12db –Filter: Choose Band or Lowpass filter with p=0.5 –Order: between 3and 10 –Lowpass: Cutoff frequency between 900 and 9000Hz –Bandpass: Low cutoff between 200 and 600Hz, high cutoff between 8 and 10 kHz All values are sampled from uniform distributions. The ranges were obtained empirically, compromising between bleeding realism and the desired goal of -1db in SDR[29]. Following this procedure, we implemented a PyTorch[30] dataloader that generates synthetic bleeding mixtures from clean Carnatic multi-stem studio recordings during each training step. The clean source material is sampled from the CMC dataset (see 2.2. Microphone Bleeding 15 Section 2.6.4). Before training, a target stem is selected - either vocal or violin - while the remaining stems (mridangam left,mridangam right, and tanpura) are treated as sources of leakage, following the approach described above. A bleeding level b∈[0,1] is randomly sampled for each example. Both the target stem and the combined bleeding stems are loudness-normalized. The final mixture is then computed as: xtrain = (1 −b)·xtarget +b·xbleed Note that xtarget is either xvocal or xviolin while xbleed is a filtered and gain reduced sum of all the remaining sources of the CMC dataset. A convolutional neural network, structurally identical to the discriminator used in MelGAN [31], is trained to learn the mapping from the input xtrain to the bleeding level b. We then pre-train a U-Net architecture [32] on the Saraga dataset for 300,000 iterations, minimizing L1 distance computed against bleeding-containing targets. Through this process, the model learns to suppress background instrumentation to match the average bleeding level present in the Saraga recordings. In the second stage, we fine-tune the pre-trained model by replacing the L1 loss with the learned bleeding estimator introduced earlier. During the finetuning, the model learns to lower the bleeding residuals significantly. The pre-training was performed on approximately 60 hours of Saraga data, while fine-tuning required only 300 iterations on a subset of 2 hours of Saraga. This suggests that the model already learns an internal representation of isolated, bleedingfree source components during pre-training. However, the discriminator loss introduced occasionally audible artifacts, steadily increasing during the finetuning stage. Due to the resulting limitation in effective duration of this second training phase the final separation quality does not yet match that of state-of-the-art systems trained 22 Chapter 2. State of the Art Figure 6: A Carnatic concert from the Saraga Audiovisual dataset. Instrumentation: Mridangam (left), Violin (right), Vocalist (center) 2.6 Datasets In the following section, all datasets used within this project are presented: 2.6.1 Saraga The Saraga-Audiovisual dataset [14] is the largest open-access, multistem collection of Carnatic music, comprising 64.8 hours of concert recordings. It consists exclusively of live Carnatic performances captured with multiple microphones on stage. Compared to the original Saraga dataset [47], the number of artists and r¯agas has approximately doubled, while the overall duration and number of recordings remain comparable [14]. Moreover, Saraga reflects the stylistic diversity, musical beauty, and nuanced performance practices of Carnatic music. This stands in contrast to other widely used open musical source separation datasets such as MUSDB18, which primarily feature studio-produced tracks with a commercial or advertisement-oriented sound. 2.6.2 MUSDB18 MUSDB18 [1] is the most widely used MSS dataset. The stems include: Vocal, Drums,Other and Bass. Other features many kinds of melodic instruments and electronic sounds. Bass includes baselines, played by bass-guitars and sub-frequency 2.6. Datasets 23 synthesizers. 2.6.3 Sanidha The Sanidha dataset [48] is a multimodal, audiovisual, multistem Carnatic music collection with little to no leakage between sources. It contains recordings of five Carnatic concerts performed at the School of Music, Georgia Institute of Technology (USA), featuring fifteen professional Carnatic musicians from Atlanta: three male vocalists, two female vocalists, four violinists, and six percussionists. The performances were captured in four soundproof rooms equipped with acoustic curtains, significantly reducing room reverberation. For synchronization purposes, each performer received a personalized monitor mix consisting of the other performers’ microphones, combined with an artificial tanpura drone. In the initial recording sessions, some performers required higher monitoring levels, which led to slight leakage from the headphones into their microphones. This issue was addressed in later sessions by switching to in-ear monitoring. As a result, concerts 2 and 3 are completely free from leakage, whereas the remaining three concerts exhibit minor bleed, especially on the vocalist tracks in the form of a metronome-like click sound. 2.6.4 Carnatic Multi-stem Clean (CMC) The Carnatic Multi-stem Clean (CMC) dataset is a private collection of 58 multistem Carnatic music recordings, totaling approximately 5hours of audio. Each recording includes one or two lead vocals, violin, mridangam, and tanpura, with some concerts also featuring t¯ala and veena. The dataset was provided by Shaale, a Bangalore-based company specializing in Indian art music education. It was originally recorded for educational purposes, enabling Carnatic musicians to practice alongside a complete ensemble. Notably, each instrument was captured in accoustic isolation, ensuring that no bleed occurs between tracks. Figure 1 compares the stems included in the four presented MSS datasets. 24 Chapter 2. State of the Art 2.6.5 Bach Violin Dataset The Bach Violin Dataset [49] contains high-quality public recordings of Johann Sebastian Bach’s sonatas and partitas for solo violin (BWV 1001–1006). It comprises 6.5 hours of performances by 17 professional violinists, recorded in a variety of sessions. In addition to the audio, the dataset provides reference scores and estimated alignments between the recordings and the corresponding scores, enabling score-informed processing and analysis. The dataset is bleeding and noise free. 2.6.6 Deep Noise Suppression Dataset (DNS) The DNS Challenge [50] is an annual competition that provides benchmarks and datasets for various speech enhancement tasks. It has been hosted at INTERSPEECH 2020,ICASSP 2021,INTERSPEECH 2021,ICASSP 2022, and ICASSP 2023. In the 2022 edition, the organizers released an additional curated dataset comprising 48 real and approximately 60,000 simulated room impulse responses (RIRs) to support research on joint speech denoising and dereverberation [51]. The RIRs were sourced from the OpenSLR26 and OpenSLR28 datasets [52] without further modification or specification. Chapter 3 A Finetuning Strategy - Methodology Previous research around Carnatic musical source separation shows that, at the time of this work, MSS models trained solely on 8 hours of the out-of-domain, leakagefree dataset MUSDB18 outperform models trained in a bleeding-aware manner on more than 60 hours of in-domain Carnatic data from the Saraga-AV dataset (see Section 2.2.1). Furthermore, related research on MSS domain adaptation highlights the superiority of fine-tuning pre-trained models with in-domain data over training from scratch using only in-domain data (see Section 2.4). Based on these findings, this work proposes a strategy for curating datasets and designing data augmentations to fine-tune a pre-trained SCNet using leakage-free data. First, the domain characteristics of Carnatic music are analysed, with emphasis on instrumentation, genre-specific differences compared to Western pop and rock, and recording/performance conditions. Next, I present a new dataset that draws from both out-of-domain and in-domain sources to represent the Carnatic music domain as comprehensively as possible. Additionally, multiple data augmentation techniques are introduced to further bridge the gap between data domains. Finally, a dedicated test set is proposed as a benchmark for evaluating separation quality 25 26 Chapter 3. A Finetuning Strategy - Methodology Table 1: Correlating stems per dataset. Dataset Vocals Drums Other Bass Unused Saraga Vocals Mridangam L/R Violin – Ghatam; Veena MUSDB18 Vocals Drums Other Bass – CMC Vocals Mridangam Tanpura; Violin – Taala; Veena Sanidha Vocals Vocals 1+2 Mridangam L/R Violin 1–2; Tanpura –Ghatam Far; Ghatam Close using standard MSS metrics as well as a perceptual test. 3.1 Data Domain Investigation: Carnatic Music vs. Western Pop and Rock To curate a fine-tuning dataset that accurately represents the Carnatic music domain and its specific challenges for MSS, this chapter examines the musicological characteristics of Carnatic music (see Section 1.2.1) in relation to the available datasets (see Section 2.6). The goal is to identify and analyse the data-domain gap between Carnatic music and the MUSDB18 dataset, with a particular focus on MSS performance in the Carnatic context. The following aspects are investigated: •Instrumentation characteristics and their domain-specific MSS challenges •Genre-specific differences and their MSS-related implications •MSS challenges arising from the live-recorded nature of the performances 3.1.1 Instrumentation Differences and Challenges in Separating Carnatic Music with Out-of-Domain MSS Models As described in Section 1.2.3, Carnatic music features a partly distinctive set of instruments, beyond the violin not typically found in Western popular music. In 3.1. Data Domain Investigation: Carnatic Music vs. Western Pop and Rock 27 preliminary experiments, I applied pre-trained MSS models - Hybrid Transformer Demucs (HTDemucs) and SCNet - both trained on MUSDB18, to separate Carnatic music recordings. MUSDB18 consists of four stems for all tracks: vocals,drums,bass, and other (see Section 2.6.2). When applied to Carnatic music, models trained on that data tend to produce the following mapping: •Vocals contain the lead vocal tracks. •Drums contain mridangam and ghatam. •Other contains taala,tanpura,veena, and violin. •Bass typically contains silence or low-frequency residual noise, but no isolated instrument stems. All separated stems contain residual bleed and, in some cases, omit portions of the target instruments. This work focuses on separating vocals,violin,mridangam, and tanpura, as these four instruments appear in every recording within the CMC, Sanidha, and Saraga datasets.1 As discussed in Section 1.2.3, the mridangam produces both harmonic and noisetransient components. The spectral analysis (Figure 2) shows the long harmonic decays of strokes such as Bheem,Cha, and Dhin. When separating the mridangam with a model trained on MUSDB18, the ringing, string-like strokes often lose substantial portions of their long decay envelopes. Figure 7 illustrates how the harmonic components of these strokes are largely removed by HTDemucsft, despite preserving the initial transients. 3.1.2 Genre-Specific Challenges in Carnatic MSS The g¯ayaki style in Carnatic music (see Section 1.2.1) creates a high melodic correlation and constant frequency overlap between the melodic instruments, particularly 1Note: In the Saraga dataset, the tanpura is not provided as an isolated stem, but present as leakage in every recording. 28 Chapter 3. A Finetuning Strategy - Methodology (a) Clean reference (b) HT_Demucs_ft separation Figure 7: Comparison of spectrograms: Mridangam audio separated from Sanidha concert 2 using an HTDemucsft model (right) and the clean reference (left). The HTDemucsft model preserves the transient components (vertical lines) but removes a significant portion of the harmonic decay (horizontal lines). the violin and the vocal stems. In contrast, Western pop and rock arrangements rarely feature instrumental parts that mimic the lead vocal so closely. State-of-theart MSS models such as SCNet and HTDemucsft operate entirely or partially in the frequency domain, using spectrograms as their primary input representation (see Sections 2.5.1 and 2.5.2). As a result, frequency-domain separation models trained primarily on Western data often struggle with Carnatic material, as the presence of violin in the same frequency bins as the vocal leads to significantly more residuals in the separated vocal stem and vice-versa. Furthermore, Carnatic music utilises intervals that are smaller than the semitone, the smallest interval in Western music[12]. This might further complicate separation tasks for models trained exclusively on Western datasets with bigger and overall more discrete pitch scales. 3.1.3 Tailoring a Training Set and Benchmarks to Improve Carnatic MSS Table 2 compares the four MSS datasets used in this work. Saraga, CMC, and Sanidha constitute the three Carnatic MSS datasets available at the time of writing. 3.1. Data Domain Investigation: Carnatic Music vs. Western Pop and Rock 29 Table 2: Overview of datasets used in this work. Dataset Length Genre Recording setup Bleeding Usage Saraga 64.8 h Carnatic Live on stage Yes Evaluation MUSDB18 10 h Pop and Rock Studio recorded No Training CMC 8 h Carnatic Studio recorded No Training Sanidha 4 h Carnatic Studio recorded Partial Evaluation Figure 8: Stem mapping between MUSDB18 and the CMC dataset. MUSDB18 is standardized to four stems for every track. In CMC, additional tracks (t¯ala,v¯ına) are present in some concerts but are not used in this work. Sanidha contains six concerts, of which concerts 2 and 3 are completely free of leakage and are therefore used as the primary evaluation benchmark for this work. MUSDB18 and CMC are leakage-free and thus suitable for training. I fine-tune models on CMC alone as well as on a combination of CMC and MUSDB18. In the latter case, Figure 8 shows the mapping between stems, as described in Section 3.1.1. Of the 58 tracks in the CMC dataset, only 24 include violin. Furthermore, all violincontaining tracks are recorded with the same Carnatic ensemble - featuring a single violinist and a single recording setup. Since models often struggle with violin–vocal separation, I oversample these 24 tracks by including them twice in the training set, applying the data augmentations described in Section 3.2.1 to increase variety. The Saraga dataset, with its stylistic diversity and real live-recorded performances, serves as the best available testing set for Carnatic music. Non-multistem Carnatic 30 Chapter 3. A Finetuning Strategy - Methodology recordings, such as publicly available concert recordings, are not suitable for the perceptual evaluation used in this work. State-of-the-art MSS models often remove parts of the target source during separation. To evaluate the completeness of the separated source, a comparative listening test with a reference is required. If the reference is the full mixture, assessing degradation becomes more difficult than when using the partially isolated, bleeding - containing stems of Saraga, where the target source is clearly in the foreground and perceptually distinct. Furthermore, this work aims to clean the Saraga AV dataset stems (see Chapter 1.2.2). 3.2 Data Augmentations 3.2.1 Aligning SCNet Data Augmentations with the Carnatic Data Domain The original SCNet training procedure uses an augmentation pipeline consisting of four operations, each applied sequentially with its own probability. While these augmentations are designed to improve generalization, their effectiveness depends heavily on the match between the augmentation strategy and the target data domain. In this work, I propose a modified pipeline that retains three of the original augmentations and introduces three new ones specifically tailored to the characteristics of Carnatic music. The authors of the SCNet propose the following augmentations: Flip Channels - Swaps the left and right stereo channels of each source with a probability of p= 0.5. Flip Sign - Inverts the phase of a source with a probability of p= 0.5. Scale - Multiplies each source by a random scaling factor sscale ∈[0.25,1.25] with a probability of p= 0.3. Remix - Randomly shuffles the same type of source across different mixtures within a batch. 3.2. Data Augmentations 31 The Flip Channels,Flip Sign, and Scale augmentations reduce overfitting and increase generalization capacities by varying model at every training step. Furthermore, the Scale augmentation is crucial for adapting to different mixdowns. For example, the CMC dataset contains only a single violin concert with a consistent loudness ratio between violin, vocal, and mridangam. Scaling these sources prepares the model for scenarios where the violin is mixed or played louder, or is more subdued in the background. The Remix augmentation, originally proposed in the Demucs paper [40] and shown to improve separation quality and generalization in MUSDB18-trained models, is excluded from the modified SCNet pipeline. Preliminary experiments indicated that Remix degraded vocal and violin separation quality in the Carnatic context by disrupting melodic correlation. When violin stems are shuffled across mixtures, the number of temporally aligned samples with strong frequency overlap to the vocal is reduced. This negatively impacts performance of carnatic violin and vocal separation. 3.2.2 Violin Data Augmentation This augmentation addresses the instrumentationand genre-specific MSS challenges in Carnatic music, particularly the g¯ayaki violin style (see Sections 3.1.2 and 3.1.1). Baseline models struggle with the strong frequency overlap between vocals and violin in Carnatic recordings, resulting in residual vocals in the violin stem and vice versa. Even after oversampling the violin-containing tracks in the CMC dataset (see Section 3.1.3), the models continue to have difficulty cleanly separating the two sources. This is exacerbated by the fact that CMC features only a single violinist, recorded in one session under identical conditions. To improve generalization in Carnatic violin separation, I introduce a violin data augmentation that adds more varied violin material to the training set. The augmentation is applied during training with a probability of p= 0.4. If the current sample does not already contain a violin, a random 11-second segment from the Bach Violin Dataset (See section 2.6.5) is selected and added to the other stem of 38 Chapter 3. A Finetuning Strategy - Methodology Experiment Training Data Finetuning Data Baselines HDemucsmmi MUSDB 800 songs HTDemucs MUSDB - HTDemucsft MUSDB + 800 songs MUSDB + 800 songs SCNet MUSDB - Proposed SCNetcMUSDB CMC SCNetc,m MUSDB MUSDB, CMC SCNetc,m,v MUSDB MUSDB,CMC Violin augment SCNetc,m,v,b MUSDB MUSDB, CMC Violin augment Bleeding augment Table 3: List of baselines and proposed models with their training and fine-tuning datasets. •Fine-tuned Hybrid Transformer Demucs (HT Demucsft, see Section 2.5.1) - included as the strongest baseline for Carnatic music. This model is trained on MUSDB18 plus 800 additional songs and averages the outputs of four different Hybrid-Demucs models during inference to maximize performance. Two additional Demucs configurations are also included: •Hybrid Demucs mmi (HDemucsmmi) - the preceding Demucs generation, which achieved state-of-the-art separation quality at the time of release. It is trained on MUSDB18 plus 800 additional songs and is included here due to its exceptional mridangam separation quality. •Hybrid Transformer Demucs (HT Demucs) - trained solely on MUSDB18, included for comparison with the fine-tuned variant. •Not-fine-tuned SCNet (SCNet see Section 2.5.2) - serves both as a baseline and as a direct point of comparison to evaluate the impact of fine-tuning. This configuration is trained exclusively on MUSDB18. Table 3.4.2 gives an overview on the training and finetuning dataset configurations of all baselines and proposed experiments. Chapter 4 Results 4.1 Improving the MSS quality for carnatic music The best checkpoint from each experiment is evaluated on the Sanidha and CMC benchmarks (see Section 3.4) and compared with the baselines (Section 3.4.3). Table 4.1 reports the complete SDR evaluation across all models for the Sanidha benchmark, Table A for the CMC testset. Audio samples are also provided for comparative listening1. 4.1.1 SDR Evaluation Overall, all four proposed models outperform all baselines across all three sources and on both test sets. Compared to the SCNet trained solely on MUSDB18, finetuning yields substantial improvements - up to +5.4 dB for vocals, +6.9 dB for violin + tanpura), and +5.5 dB for the mridangam source on the Sanidha benchmark. In particular, performance on the other stem (violin+tanpura) is noteworthy, with SDR values reaching nearly three times those of the best baseline model. The CMC evaluation delivers similar results - only the violin+tanpura source is more than 3db higher than the Sanidha evaluation, suggesting that the model might overfit on the CMC violin. It should be noted that the CMC benchmark contains 1Google Drive link with evaluation files 39 40 Chapter 4. Results Group Experiment Vocal Violin+Tanpura Mridangam Baselines HDemucsmmi 11.86 3.13 10.07 HTDemucs 12.27 3.52 9.74 HTDemucsft 12.63 3.92 8.65 SCNet 12.59 2.66 8.64 Proposed SCNetc14.76 7.13 13.77 SCNetc,m 16.67 8.08 13.06 SCNetc,m,v 18.02 9.65 14.16 SCNetc,m,v,b 16.30 8.52 14.18 Table 4: Signal-to-Distortion Ratio (dB) on the Sanidha benchmark. underline: strongest baseline, bold: strongest overall. only one song, performed by an ensemble that the proposed models encountered during training. 4.1.2 On the Effectiveness of the Proposed Data Augmentations Violin Augmentation The SCNetc,m,v configuration achieves the highest overall performance, exceeding the SDR of any other model by more than +1 dB for all sources except drums. A direct comparison with SCNetc,m - trained without the violin augmentation - demonstrates the effectiveness of the proposed violin augmentation. Figure 17 illustrates the impact. Both vocals,violin + tanpura and mridangam show substantial SDR gains compared to the preceding experiment, confirming the augmentation’s value for addressing the strong vocal-violin overlap in Carnatic music. Bleeding Augmentation In contrast, the bleeding augmentation applied in the SCNetc,m,v,b experiment slightly reduces SDR across all sources except drums. For drums, the improvement over SCNetc,m,v is marginal at +0.02 dB - well within the range of measurement noise and thus not statistically meaningful. 4.2. Perceptual Evaluation: SCNetc,m,v vs HTDemucsft 41 Vocal Other Drums 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 SDR (dB) 16.67 8.08 13.06 18.02 9.65 14.16 SCNet c , m SCNet c , m , v Sanidha Benchmark SDR With vs without violin augmentation Figure 11: Impact of the violin augmentation on the SDR evaluation on Sanidha dataset. Yet, a perceptual comparison between the models with bleeding augmentation and without, as described in Chapter 3.4.2, show the improved generalization capacities regarding the carnatic violin due to the bleeding augmentation. Figure 12 shows that the SCNetc,m,v,b yielded perceptually better separations in 42% and worse in only 18% of the samples. The bleeding augmentation led to 16% less artifacts/inferencecontaining samples and 10% less corrupted separations (See Figure 18). Yet, the perceptual quality decreased significantly for mridangam and vocal due to the increase of residuals of other sources in the separations. 4.2 Perceptual Evaluation: SCNetc,m,v vs HTDemucsft A second perceptual test, as described in Section 3.4.2, is conducted to evaluate whether the best proposed method, SCNetc,m,v,(1) generalizes to more realistic, non-studio-recorded data, and (2) whether perceptual quality correlates with the SDR measurement. The overall perceptual quality comparison with the strongest baseline, HTDemucsft, over 100 random Saraga AV samples is shown in Figure 15. Only 13% of the Demucs separations exhibit higher perceptual quality, while 27% 42 Chapter 4. Results 60% 4% 36% Vocals 66% 2% 32% Mridangam 18% 42% 40% Violin/Tanpura Saraga Perceptual Evaluation: Overall perceptual quality With vs without bleeding augmentation SCNetc , m , v better SCNetc , m , v , b better Same quality Figure 12: Impact of the bleeding augmentation on perceptual separation quality, computed on Saraga AV. show no perceivable difference. In 60% of the cases, the proposed method is rated as sounding perceivable better than the strongest baseline. Figure 13 compares the cleanliness and preservation quality between the two models. The proposed model removes substantially fewer frequencies than the baseline. Only 2% of the vocal separations and 4% of the mridangam samples are corrupted by SCNetc,m,v. The violin corruption rate is higher, at 26%, yet still significantly lower than the baseline’s 54%. 4.2.1 Mridangam Separation The mridangam emerges as the perceptually cleanest separated source. The best proposed model improves SDR by +4.11 dB over the strongest baseline (HDemucsmmi). Figure 14 illustrates how the proposed model preserves significantly more harmonic content (visible as horizontal lines in the spectrogram) from the resonant, tonal strokes - an area where all baselines struggled (compare Figure 7). Perceptually, SCNetc,m,v sounds almost as rich as the original source on nearly all test samples - there is as good as no source corruption (see Figure 13). The compared baseline, HTDemucsft, produces cleaner separations - 73% contain no perceivable instrument residuals or artifacts, compared to 61% for the proposed 4.2. Perceptual Evaluation: SCNetc,m,v vs HTDemucsft 43 0 20 40 60 80 100 Clean samples (%) 32% 61% 39% 21% 73% 35% Free of artifacts and interference SCNetc , m , v HTDemucsft Vocals Mridangam Violin/Tanpura 0 20 40 60 80 100 Not-corrupted samples (%) 98% 96% 74% 98% 25% 46% Source preserved Saraga Perceptual Evaluation: Cleanliness and preservation quality Strongest proposed model vs strongest baseline Figure 13: Amount of samples without artifacts/residuals and without corruption per source of HT Demucsft and SCNetc,m,v. model. However, this cleanliness is largely due to the heavy gating of the Demucs model, which results in separations containing only transient sounds. In contrast, the proposed model’s separations preserve the harmonic decays of the mridangam, although some resonant transients still contain residuals from string instruments or, more rarely, vocals performed at the same pitch as the drum. Overall, the proposed model is rated higher than the strongest baseline in 92% of the mridangam samples in the perceptual test (see Figure 15). 4.2.2 Vocal Separation In contrast to the mridangam, the vocal stem is generally well preserved by the strongest baseline, HTDemucsft (see Figure 13). The observed SDR improvement of +5.39 dB for SCNetc,m,v is primarily due to better removal of interfering sources, 44 Chapter 4. Results (a) SCN et (b) SCN etc,m,v Figure 14: Comparison of spectrograms: Mridangam audio separated from Sanidha concert 2 with the pretrained SCNet vs the finetuned SCNetc,m,v. The finetuned preserves the harmonics while the pretrained mainly preserves transients. Compare with Figure 7 for the reference and HT Demucsft separations. particularly the violin. Separation outputs from SCNetc,m,v contain significantly fewer violin residuals than those from the baselines. While 68% of the samples in the perceptual test contain instrument residuals and artifacts, these are considerably quieter than in the baseline, resulting in an overall substantially higher perceptual quality (see Figure 15). 4.2.3 Violin and Tanpura Separation The other stem, consisting of violin and tanpura, shows the largest relative improvement over the baselines. This is particularly important given the strong melodic correlation between violin and vocals in Carnatic music. The violin augmentation appears to play a key role here, as SCNetc,m,v achieves the highest SDR on both benchmarks for this stem. Perceptually, the violin and tanpura stem remains the weakest source of the proposed model on the Saraga dataset: frequency components are removed in 26% of the samples, and 61% contain artifacts and/or residuals. Loud, slightly distorted violin sounds in the Saraga dataset are sometimes not correctly identified by the proposed model. Nevertheless, even for this source, the perceptual test shows a significant 4.2. Perceptual Evaluation: SCNetc,m,v vs HTDemucsft 45 51% 13% 36% Vocals 67% 8% 25% Mridangam 62% 18% 20% Violin/Tanpura Saraga Perceptual Evaluation: Overall perceptual quality Best proposed model vs best baseline SCNetc , m , v best HTDemucsft best Both same quality Figure 15: Perceptual evaluation: Overall best separation quality on Saraga AV. improvement in overall separation quality relative to the strongest baseline (see Figure 15). 46 Chapter 4. Results 4.3 Training Stability: Extending vs Replacing the Fine-tuning Dataset Similar work on MSS domain adaptation has highlighted the advantages of finetuning pre-trained models, even when pre-training is on out-of-domain data (Section 2.4). This work further suggests that including the out-of-domain dataset in the fine-tuning process can be beneficial. Figure 16 compares SDR curves on a subset of Sanidha for SCNetc(CMC only) and SCNetc,m (CMC + MUSDB18). Including the pretraining dataset MUSDB18 enables longer, more stable fine-tuning and yields higher SDR results. The SDR values in this figure differ from Table 4.1 because only a subset of the Sanidha benchmark is used. There are several possible explanations for this behaviour. Training exclusively on CMC may lead to overfitting, given its smaller size and limited variability. The hyperparameters (e.g., learning rate) may not be optimally tuned for CMC. Furthermore, the SCNet checkpoints provided by the authors have been extensively trained on MUSDB18; removing this data from fine-tuning might destabilize training by moving the model too far from its pre-trained distribution. 4.3. Training Stability: Extending vs Replacing the Fine-tuning Dataset 47 10 12 14 16 18 SDR (dB) Vocal 4 6 8 10 SDR (dB) Violin + Tanpura 0 5 10 15 20 25 30 35 Epoch 8 10 12 14 SDR (dB) Mridangam SCNetc - Finetuning on CMC SCNetc , m - Finetuning on CMC and MusDB SDR development with finetuning epochs Figure 16: SDR development per fine-tuning epoch and source on a subset of the Sanidha benchmark. SCNetcshows a rapid quality boost after the first epoch followed by a decline. SCNetc,m declines only after 25 epochs. 54 BIBLIOGRAPHY [16] Chandramouli, K. & Sethares, W. Automatic transcription of drum strokes in carnatic music (2022). URL https://arxiv.org/abs/2211.15185.2211. 15185. [17] SENGUPTA, R., DEY, N., DATTA, A. K. & GHOSH, D. Assessment of musical quality of tanpura by fractal-dimensional analysis. Fractals 13, 245– 252 (2005). URL https://doi.org/10.1142/S0218348X05002891.https: //doi.org/10.1142/S0218348X05002891. [18] Salamon, J. Melody Extraction from Polyphonic Music Signals. Ph.D. thesis (2013). [19] Clayton, M., Rao, P., Shikarpur, N. N., Roychowdhury, S. & Li, J. Raga classification from vocal performances using multimodal analysis. In Proc. of the 23rd Int. Society for Music Information Retrieval (ISMIR), Bengaluru, India, 283–290 (2022). [20] Nuttall, T., Plaja-Roglans, G., Pearson, L. & Serra, X. The matrix profile for motif discovery in audio-an example application in carnatic music. In Int. Symposium on Computer Music Multidisciplinary Research, Tokyo, Japan, 228– 237 (2021). [21] Nuttall, T., Plaja-Roglans, G., Pearson, L. & Serra, X. In search of sañc¯aras: tradition-informed repeated melodic pattern recognition in carnatic music. In Proc. of the 23rd Int. Society for Music Information Retrieval Conf. (ISMIR), Bengaluru, India, 337–344 (2022). [22] Plaja-Roglans, G., Miron, M., Shankar, A. & Serra, X. Carnatic singing voice separation using cold diffusion on training data with bleeding. In 24th Int. Society for Music Information Retrieval Conf. (ISMIR), Milano, Italy (2023). [23] Rouard, S., Massa, F. & Défossez, A. Hybrid transformers for music source separation (2022). URL https://arxiv.org/abs/2211.08553.2211.08553. BIBLIOGRAPHY 55 [24] Kokkinis, E. K., Reiss, J. D. & Mourjopoulos, J. A wiener filter approach to microphone leakage reduction in close-microphone applications. IEEE Transactions on Audio, Speech, and Language Processing 20, 767–779 (2012). [25] Shankar, A., Schweinitz, S., Plaja-Roglans, G., Serra, X. & Rocamora, M. Disentangling overlapping sources: Improving vocal and violin source separation in carnatic music. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2025). [26] Fu, S.-W., Liao, C.-F. & Tsao, Y. Learning with learned loss function: Speech enhancement with quality-net to improve perceptual evaluation of speech quality. IEEE Signal Processing Letters 27, 26–30 (2020). URL http://dx.doi.org/10.1109/LSP.2019.2953810. [27] Xin Bai, H. Z., Xueliang Zhang & Huang, H. Perceptual loss function for speech enhancement based on generative adversarial learning. In 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (2022). [28] Wuxuan Gong, Y. L., Jing Wang & Yang, H. A no-reference speech quality assessment method based on neural network with densely connected convolutional architecture. In INTERSPEECH 2023, Dublin, Ireland (2023). [29] Fabbro, G. et al. The Sound Demixing Challenge 2023: Music Demixing Track (2023). URL http://arxiv.org/abs/2308.06979.2308.06979. [30] Paszke, A. et al. Pytorch: An imperative style, high-performance deep learning library (2019). URL https://arxiv.org/abs/1912.01703.1912.01703. [31] Kumar, K. et al. Melgan: Generative adversarial networks for conditional waveform synthesis (2019). [32] Kim, M., Choi, W., Chung, J., Lee, D. & Jung, S. KUIELab-MDX-Net: a two-stream neural network for music demixing (2021). URL http://arxiv. org/abs/2111.12203.2111.12203. 56 BIBLIOGRAPHY [33] Kang, D. & Hashimoto, T. B. Improved natural language generation via loss truncation. In Jurafsky, D., Chai, J., Schluter, N. & Tetreault, J. (eds.) Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 718–731 (Association for Computational Linguistics, Online, 2020). [34] Choi, W., Kim, M., Chung, J., Lee, D. & Jung, S. Investigating u-nets with various intermediate blocks for spectrogram-based singing voice separation (2020). URL https://arxiv.org/abs/1912.02591.1912.02591. [35] Sivasankar, A. S. Vocal source separation for carnatic music (2023). URL https://doi.org/10.5281/zenodo.8380379. [36] Hennequin, R., Khlif, A., Voituret, F. & Moussallam, M. Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software 1–4 (2020). [37] Chen, Z. Singing vocal source separation on jingju music (2024). URL https: //doi.org/10.5281/zenodo.13863042. [38] Stoller, D., Ewert, S. & Dixon, S. Wave-u-net: A multi-scale neural network for end-to-end audio source separation (2018). URL https://arxiv.org/abs/ 1806.03185.1806.03185. [39] Lin, K. W. E., T., B. B., Koh, E., Lui, S. & Herremans, D. Singing voice separation using a deep convolutional neural network trained by ideal binary mask and cross entropy (2018). URL https://arxiv.org/abs/1812.01278. 1812.01278. [40] Défossez, A., Usunier, N., Bottou, L. & Bach, F. Demucs: Deep extractor for music sources with extra unlabeled data remixed (2019). URL https: //arxiv.org/abs/1909.01174.1909.01174. [41] Tong, W. et al. Tfcnet: Time-frequency domain corrector for speech separation. In ICASSP, 1–5 (2023). URL https://doi.org/10.1109/ICASSP49357.2023. 10096785. BIBLIOGRAPHY 57 [42] Tong, W. et al. Scnet: Sparse compression network for music source separation (2024). URL https://arxiv.org/abs/2401.13276.2401.13276. [43] Su, J., Jin, Z. & Finkelstein, A. Hifi-gan: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks (2020). URL https://arxiv.org/abs/2006.05694.2006.05694. [44] Bahmaninezhad, F. et al. A comprehensive study of speech separation: spectrogram vs waveform separation (2019). URL https://arxiv.org/abs/1905. 07497.1905.07497. [45] Luo, Y. & Yu, J. Music source separation with band-split rnn. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31, 1893–1901 (2023). [46] Shazeer, N. Glu variants improve transformer (2020). URL https://arxiv. org/abs/2002.05202.2002.05202. [47] Srinivasamurthy, A., Gulati, S., Repetto, R. C. & Serra, X. Saraga: Open datasets for research on indian art music. Empirical Musicology Review 16, 85–98 (2021). [48] Krishnan, V. V., Alben, N., Nair, A. A. & Condit-Schultz, N. Sanidha: A studio quality multi-modal dataset for carnatic music, San Francisco, United States. In Proc. of the 25th Int. Society for Music Information Retrieval Conf. (2024). URL http://arxiv.org/abs/2501.06959. [49] Dong, H.-W., Zhou, C., Berg-Kirkpatrick, T. & McAuley, J. Deep performer: Score-to-audio music performance synthesis (2022). URL https://arxiv.org/ abs/2202.06034.2202.06034. [50] Dubey, H. et al. Icassp 2023 deep noise suppression challenge (2023). URL https://arxiv.org/abs/2303.11510.2303.11510. [51] Dubey, H. et al. Icassp 2022 deep noise suppression challenge (2022). URL https://arxiv.org/abs/2202.13288.2202.13288. 58 BIBLIOGRAPHY [52] Ko, T., Peddinti, V., Seltzer, M. & Khudanpur, S. A study on data augmentation of reverberant speech for robust speech recognition. 5220–5224 (2017). [53] Vincent, E., Gribonval, R. & Fevotte, C. Performance measurement in blind audio source separation. IEEE Transactions on Audio, Speech, and Language Processing 14, 1462–1469 (2006). Appendix A Appendix 59 60 Appendix A. Appendix Group Experiment Vocal Violin+Tanpura Mridangam Baselines HDemucsmmi 8.63 2.92 8.64 HTDemucs 8.98 3.17 7.71 HTDemucsft 9.89 4.06 7.66 SCNet 9.38 2.70 4.18 Proposed SCNetc12.93 7.97 10.37 SCNetc,m 15.68 10.46 13.04 SCNetc,m,v 17.74 12.76 13.65 SCNetc,m,v,b 15.95 10.76 12.60 Table 5: Signal-to-Distortion Ratio (dB) on the CMC benchmark. underline: strongest baseline, bold: strongest overall. The CMC benchmark only contains one track, not seen during training. Yet, the results are comparable to the Sanidha benchmark, see Figure 4.1. Vocal Other Drums 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 SDR (dB) 18.02 9.65 14.16 16.30 8.52 14.18 SCNet c , m , v SCNet c , m , v , b Sanidha Benchmark SDR With vs without bleeding augmentation Figure 17: Impact of the bleeding augmentation on the SDR evaluation, computed for the Sanidha benchmark. 61 0 20 40 60 80 100 Clean samples (%) 32% 64% 34% 26% 30% 50% Free of artifacts and interference SCNetc , m , v SCNetc , m , v , b Vocals Mridangam Violin/Tanpura 0 20 40 60 80 100 Not-corrupted samples (%) 98% 98% 72% 98% 94% 82% Source preserved Saraga Perceptual Evaluation: Cleanliness and preservation quality With vs without bleeding augmentation Figure 18: Impact of the bleeding augmentation on perceptual separation quality, computed on Saraga AV. Compared is the amount of samples without artifacts/residuals and without corruption per source