scieee AI-readable full text Open interactive document viewer

Conditioned Wave-U-Net for Acoustic Matching of Speech in Shared XR Environments

Luberadzka, Joanna; Gusó, Enric; Sayin Saraç, Umut

Abstract

Mismatch in acoustics between users is an important challenge for interaction in shared XR environments. It can be mitigated through acoustic matching, which traditionally involves dereverberation followed by convolution with a room impulse response (RIR) of the target space. However, the target RIR in such settings is usually unavailable. We propose to tackle this problem in an end-to-end manner using wave-u-net encoder-decoder network with potential for real-time operation. We use FiLM layers to condition this network on the embeddings extracted by a separate reverb encoder to match the acoustic properties between two arbitrarily chosen signals. We demonstrate that this approach outperforms two baseline methods and provides the flexibility to both dereverberate and rereverberate audio signals.

Full text

Conditioned Wave-U-Net for Acoustic Matching of Speech in Shared XR Environments Joanna Luberadzka1, Enric Gus´ o1,2, Umut Sayin1 1Eurecat, Centre Tecnol` ogic de Catalunya, Tecnologies Multim` edia, Barcelona 2Universitat Pompeu Fabra, Music Technology Group, Barcelona Abstract—Mismatch in acoustics between users is an important challenge for interaction in shared XR environments. It can be mitigated through acoustic matching, which traditionally involves dereverberation followed by convolution with a room impulse response (RIR) of the target space. However, the target RIR in such settings is usually unavailable. We propose to tackle this problem in an end-to-end manner using wave-u-net encoder-decoder network with potential for real-time operation. We use FiLM layers to condition this network on the embeddings extracted by a separate reverb encoder to match the acoustic properties between two arbitrarily chosen signals. We demonstrate that this approach outperforms two baseline methods and provides the flexibility to both dereverberate and rereverberate audio signals. 1. INTRODUCTION Extended Reality (XR) has attracted considerable attention in recent years. Beyond its diverse applications, it provides a platform for human communication, presenting opportunities for the integration of speech and audio technologies. The goal is to enable individuals from different remote locations to engage in a conversation that, from each user’s perspective, creates the illusion of being part of the real world. In addition to providing visual avatars for remote conversation partners, the speech captured by their respective microphones needs to be synthetically integrated into the user’s existing soundscape [1]. However, these augmented recordings may originate from spaces with very different acoustic properties than the user’s current location, leading to a perceived lack of acoustic coherence. To mitigate this issue, it is necessary to adjust the speech recorded in one space to match the acoustic attributes of the target space. Traditionally, this task involves: 1) dereverberating the signal, 2) estimating the room impulse response (RIR) characterizing the target acoustic space, and 3) convolving the dereverberated signal with the estimated RIR. The success of this pipeline relies on blind single-channel acoustic system identification, which, despite recent progress, remains a challenging problem [2], [3]. Advancements in deep learning continue to produce innovative solutions for traditional signal processing tasks, including dereverberation [2], [4], blind RIR estimation [5]–[7], room acoustic parameter estimations [8], [9], or artificial reverberation [10], [11]. At the same time, it is increasingly common to encounter methods that replace multi-stage analysis-synthesis frameworks with end-toend neural approaches. Following this trend, several architectures have been proposed for modifying the perceived sound environment, however, they either focus on artificial reverb effects in music production [12], [13], operate at low sampling rates [14], [15], facilitate transformation between only two domains [16], or require visual input [17]. The mentioned studies also do not comment on the realtime potential of the proposed systems. Thus, there are prospects for advancements in acoustic space transfer. In this work, we propose a time-domain end-to-end acoustic space transfer approach that matches the acoustic properties between two arbitrarily chosen reverberant speech signals. Our method (CWUNET) consists of three main components (See Figure 1): • a time-domain convolutional reverb encoder which extracts information about acoustic space from the reference speech, • a time-domain convolutional encoder-decoder which modifies the input signal so that it matches the target acoustic properties, • a conditioning mechanism, which allows to use the information about the target space in the transformation. We train the proposed approach with two different loss functions and compare them to two baselines: a combination of weighted prediction error (WPE) dereverberation [18] and the DNN-based blind single-channel RIR estimation [6], and a combination of DNN-based dereverberation [19] with the mentioned RIR estimation method. In addition to the baselines, we compare our models to a semi-oracle case, in which the RIR estimation is combined with anechoic signals. Using objective metrics and a listening test, we show that CWUNET transforms the reverberation of speech signals, both from more to less and from less to more reverberation and outperforms both baselines. Conv1d BatchNorm FiLM PReLU Downsample PReLU FiLM BatchNorm Conv1d Upsample Concat Concat Concat Reverb Encoder Reverb Encoder content prediction style Conv1d BatchNorm LeakyReLU Conv1d Tanh Reverb Encoder Transfer Network FiLM Conditioning Fig. 1: Conditional wave-u-net structure. 2. PROPOSED METHOD 2.1. Problem formulation Following the traditional nomenclature in neural style transfer [20], the goal of our approach is to modify the properties of the content signal s1r1 so that it sounds as if it originated from the acoustic space characterized by the style signal s2r2 , i.e. to estimate the reverberant target signal s1r2 , given two input signals s1r1 and s2r2 (See Figure 2). The signals are defined as follows: s1r1=s1∗r1;s2r2=s2∗r2;s1r2=s1∗r2, where s1 and s2 denote two distinct pseudo-anechoic speech samples, r1and r2denote two distinct RIRs, and ∗denotes convolution. * * * Fig. 2: Data generation and training. 2.2. Network structure Our approach consists of three main building blocks, detailed in the sections below (See Fig. 1). 2.2.1. Reverb Encoder: The purpose of this network is to learn the acoustically meaningful representation of reverberation. The network takes an audio waveform at the input and generates a fixed-length embedding at the output. Distinct speech samples originating from the same space should result in similar embeddings when passed through the network. We adopt the time-domain encoder originally proposed as one of the main components of the Filtered Noise Shaping (FINS) approach for blind RIR estimation [6]. FINS was shown to be highly effective at deriving RIR embeddings from reverberant speech signals. In our approach, the Reverb Encoder R(·) is used to generate embeddings zr1 and zr2 from reverberant speech samples s1r1 and s2r2 , i.e. R(s1r1) = zr1 and R(s2r2) = zr2 . Embeddings zr1 and zr2 serve as conditions for guiding the encoder and decoder of the Transfer Network, respectively. While the Reverb Encoder is not real-time tailored, the embeddings must be inferred only once for each environment pair. 2.2.2. Transfer Network: This building block has an encoderdecoder structure, and its goal is to transform the content sound s1r1 into the target sound s1r2 . Due to the significant correlation between the input and output data, we employ an encoder-decoder architecture with skip connections. We use wave-u-net [21]: a network topology broadly used in source separation [22] and speech enhancement [23]. It can be adapted to work in real-time [24], [25] and allows for manipulating the sound directly in the time domain. 2.2.3. FiLM Conditioning: The Transfer Network modifies the properties of the content signal s1r1 depending on the style sound s2r2 . Hence, it has to be flexible to perform the task of both removing and adding reverberation. We use Feature-wise Linear Modulation (FiLM) [26] as the conditioning mechanism, which has been successfully used in other U-Net-like architectures [27], [28]. 2.3. Dataset Our dataset contains speech convolved with synthetically generated RIRs. We use content s1r1 and style s2r2 at the input of CWUNET and target s1r2 as the ground truth (See Fig. 2). To obtain s1r1 , a pseudo-anechoic utterance s1 is convolved with RIR r1 . Other signals are created analogously. Speech samples are 2.73-second excerpts chosen from a pool of 93,046 speech audio files from VCTK [29] and PTBD [29] databases. The exact excerpt is randomly cut from the original audio file every epoch as a data augmentation strategy. RIRs are chosen from a pool of 10k RIRs synthetically generated using the Multichannel Acoustic Signal Processing Library (MASP) [30]. Each RIR comes from a different simulated shoebox room with a source placed at a random position inside the room. The room volume ranges from 13 m3 to 4000 m3 and reverberation time RT60 from 0.1s to 1.2s. The receiver is placed 10cm from the source to simulate the distance from the user’s mouth to the headset-microphone. Using the speech pool and the RIR pool, we create 300k random combinations of speech files and RIRs, which give 150k content-style pairs (113.75 hours of input data), divided into 80% train, 10% validation, and 10% test splits. The sampling rate is always 48kHz. 2.4. Loss and training details All network blocks are trained end-to-end from scratch in a traditional supervised manner (See Fig. 2). Given two input signals: content s1r1 and style s2r2 , the network outputs an estimate ds1r2 . The objective of the training is to minimize the reconstruction loss between the estimate and the target: LR(ds1r2, s1r2) . The reconstruction loss consists of a frequency-domain loss Lfand time-domain loss Lt: LR=λf· Lf+λt· Lt(1) We present results for two versions of the model, which differ in the employed frequency domain loss: 1) CWUNET-stft, which uses a multi-resolution STFT loss (with FFT sizes of 256, 512, 1024, 2048, and 4096, hop sizes of 25% of the respective FFT sizes, and Hann windows), and 2) CWUNET-mel, which uses a multi-resolution log-mel loss (with the same STFT parameters as the multi-resolution STFT loss and with 40 mel frequency bands distributed between 80 Hz and 7600 Hz). [31]. For the time-domain loss, both versions use the waveform shape loss [32], which computes the average L1 loss between max-pooled absolute waveforms across multiple window lengths, capturing shape differences at various time scales (we use window lengths of 300, 200, 100 samples). In both model versions λf= 0.8 and λt= 0.2 were set after performing a preliminary experiment. On top of the aforementioned reconstruction losses, we considered using an L2 style loss between the reverb embeddings of prediction ds1r2 and style s2r2 signals, but we found it to degrade performance. Each model was trained for 300 epochs with a batch size of 8, using the Adam optimizer and a learning rate of 1e-4. The Reverb Encoder network consisted of 12 blocks, leading to an embedding size D=512. The wave-u-net architecture included 12 encoder and decoder blocks, with one FiLM layer per each encoder and decoder block. 3. EVALUATION We evaluate the performance of our system in three ways: 1) by visualizing the information captured in the latent space of the reverb encoder, 2) using objective audio similarity metrics, and 3) using subjective MUSHRA-based evaluation. 3.1. Baselines We compare CWUNET-stft and CWUNET-mel with two baselines: WPE+FINS and DFNET+FINS, and with one semi-oracle scenario: oracle+FINS. In both baselines, we first dereverberate s1r1 into bs1 . Then, we estimate br2 from s2r2 , and finally we convolve these signals to get the estimate ds1r2=bs1∗br2 . For dereverberation, we use the weighted prediction error (WPE) technique [18] in WPE+FINS and a DNN-based denoising and dereverberation model (DFNET) [19] in DFNET+FINS. In both baselines, we use the FINS model [6] to estimate br2 . We train FINS with default parameters for 200 epochs, using our speech pool and our collection of synthetic RIRs. To provide an ablation of the FINS we report results on a semi-oracle case (oracle+FINS) where we directly take the pseudo-anechoic s1 and convolve it with br2 . To allow the comparison between the signals before and after acoustic matching, we also report the content condition, i.e., the similarity between unprocessed content signal s1r1 and the target s1r2. 3.2. Visualization of the latent space In the original work by [6], it was demonstrated how the latent space encodes information about the room acoustic properties. Here, we follow this practice. Figure 3 represents the reverb embeddings reduced to 2-dimensions using PCA [33]. The points in the latent space are arranged according to the acoustic parameters, which indicates that the network has learned how to encode information important for describing room acoustics. We also show the embeddings of 500 random pseudo-anechoic speech signals convolved with 5 different RIRs (100 samples per RIR). The embeddings of signals convolved with the same IR lie close to each other in the latent space. This confirms that the variability in the speech content has only a limited impact on the embedding (the reverberation of the signal influences the embedding, but the speech content does not). 40 30 20 10 0 10 20 30 40 20 0 20 colormap=RT60 a) 0.2 0.4 0.6 0.8 1.0 40 30 20 10 0 10 20 30 40 20 0 20 colormap=EDT 0.5 1.0 1.5 40 30 20 10 0 10 20 30 40 20 0 20 colormap=C50 10 20 30 30 20 10 0 10 20 30 40 20 10 0 10 20 colormap=RIR ID 1 2 3 4 5 b) c) d) Fig. 3: Visualization of the latent space: The reverb encoder network is used to extract the embeddings of reverberant speech samples, which we reduce to 2 dimensions using PCA. Each point represents one speech sample. In panels a)-c) the colors indicate the acoustic parameters of the RIR that the signal was convolved with: a) reverberation time, b) early decay time, and c) clarity. In panel d), the colors indicate which RIR was used to generate the signal. 3.3. Objective metrics The goal of the objective evaluation is to measure how close the reverb of the estimate ds1r2 is to the reverb of the style s2r2 . To our knowledge, there are currently no computational metrics that can evaluate the similarity of reverberation between two distinct speech signals. An alternative approach, which we employ in this work, is to assess the similarity between ds1r2 and the ground truth target s1r2 . The estimated ds1r2 should be more similar to the ground truth target s1r2 than the content signal at the input s1r1 . Hence, we treat an increase in the similarity as a measure of performance. Since the model must handle both deand rereverberation, it is important that the measures are independent of the direction of this transformation. Therefore, we use symmetric similarity metrics i.e., such that M(a, b) = M(b, a) . To achieve this, we either use metrics that are symmetric by default, enforce symmetry in non-symmetric metrics or compute the absolute difference in a metric before and after transformation. All objective metrics are specified in Table 1. We report our objective results in Table 2. In total, we compare our 2 model versions with 2 baselines and 1 semi-oracle case using 12 metrics. Overall, all evaluation metrics show that CWUNET-stft and CWUNET-mel outperform both baselines. Among the non-perceptual metrics (such as MR-MEL, MR-WAV, MR-STFT, L2-EMB, MCD), the CWUNET-stft model gives the best results. In contrast, the perceptual metrics, designed to reflect the judgements of human auditory system (like ∆PESQ,∆STOI,∆ni-PESQ,∆ni-STOI ), tend to favor the CWUNETmel model. On top of outperforming both baselines, CWUNET either reaches or surpasses the semi-oracle case (oracle dereverberation convolved with FINS-estimated RIR). Compared to the unprocessed signal (content), we see both CWUNET versions succeed in making the input signal more similar to the target. However, this is not always true for the baseline methods or even for the semi-oracle case. For example, in the DFNET+FINS baseline, the distance in MR-MEL between the processed sound and the target is 0.263, which is higher than the distance in MR-MEL between the original input and the target (0.170). This means that, according to MR-MEL the baseline actually made the sound less similar to the target. 3.4. Subjective evaluation Initial listening indicated that perceptual differences between models were not always aligned with objective metric scores. In general, methods that used FINS sounded more natural, but even the semi-oracle case introduced a specific coloration that made the output perceptually deviate from the target. In contrast, our models were able to closely match the spectral content of the target but introduced distortions that reduced perceived similarity. This was particularly noticeable for CWUNET-mel, which effectively replicated reverberation but introduced a strong metallic ringing artifact. Interestingly, this type of distortion was not captured by the perceptual metrics. To explore these observations in detail, an additional MUSHRA-based listening test [40] was conducted. The test included nine examples with clear differences in reverberation between the content and style speech 1 . In four examples, the transformation involved converting a highly reverberant signal to a less reverberant signal (dereverberation), while in the remaining five, it involved converting a slightly reverberant signal to a more reverberant one (rereverberation). Twelve participants, including seven expert listeners, participated in the study. In each trial, participants assessed seven audio signals played over headphones: five reverberation matching methods (our two model versions and three baselines), a hidden low anchor (content sound), and a hidden reference (target sound). They were asked to rate the similarity of the reverberation of each sample with the reverberation of a reference signal (target). All audio samples were normalized to a –17 dB LUFS. Pairwise comparisons between conditions were analyzed using the Wilcoxon signed-rank test [41] to account for the ordinal and nonGaussian nature of MUSHRA scores. Results of the MUSHRA test are depicted in Figure 4. Despite surprisingly low scores of oracle+FINS in the non-perceptual objective metrics (see Table 2 for comparison), this semi-oracle case obtained significantly higher human ratings than all other methods (pairwise comparison with CWUNET-stft: p= 5.0e-3 , with DFNET+FINS: p= 6.3e-9 , with CWUNET-mel: p= 9.2e-12 , with WPE+FINS: p= 3.1e-18 ). The second highest rated was our CWUNET-stft 1https://joaluba.github.io/CWUNET-demo/ Table 1: Evaluation metrics along with the methods applied to ensure symmetry. ds1r2 denotes the estimate, s1r2 the target signal, and s1 the clean reference pseudo-anechoic signal. Pearson correlation coefficients (r) quantify alignment with ratings from the MUSHRA listening test. Metric Name Direction Equation Symmetry Type r Multi-resolution log-mel loss [31] ↓MR-MEL([s1r2, s1r2) Symmetric by default -0.44 Multi-resolution waveform shape loss [32] ↓MR-WAVE([s1r2, s1r2)-0.25 L2 loss between embeddings ↓L2-EMB(R([s1r2),R(s1r2)) -0.45 Multi-resolution STFT loss [31] ↓MR-STFT = (MR-STFT([s1r2, s1r2) + MR-STFT(s1r2,[s1r2))/2 Enforced symmetry -0.32 Mel-Cepstral Distortion [34] ↓MCD = (MCD([s1r2, s1r2) + MCD(s1r2,[s1r2))/2-0.30 Frequency-weighted segmental SNR [35] ↑FWSNR = (FWSNR([s1r2, s1r2) + FWSNR(s1r2,[s1r2))/20.22 Intrusive PESQ [36] difference ↓∆PESQ = PESQ([s1r2, s1)−PESQ(s1r2, s1)  Absolute difference -0.52 Intrusive STOI [37] difference ↓∆STOI = STOI([s1r2, s1)−STOI(s1r2, s1) -0.27 Non-intrusive PESQ [38] difference ↓∆ni-PESQ = ni-PESQ([s1r2)−ni-PESQ(s1r2) -0.46 Non-intrusive STOI [38] difference ↓∆ni-STOI = ni-STOI([s1r2)−ni-STOI(s1r2) -0.70 Non-intrusive SISDR [38] difference ↓∆ni-SISDR = ni-SISDR([s1r2)−ni-SISDR(s1r2) -0.46 Non-intrusive SRMR [39] difference ↓∆ni-SRMR = ni-SRMR([s1r2)−ni-SRMR(s1r2) -0.25 Table 2: Comparison of different processing methods across multiple similarity metrics (see Sec. 3.3). Performance of each model can be assessed by comparing the metric value for the unprocessed signal (content) with the value for the estimated signal. MR-MEL ↓MR-WAVE ↓L2-EMB ↓MR-STFT ↓MCD ↓FWSNR ↑∆PESQ ↓∆STOI ↓∆ni-PESQ ↓∆ni-STOI ↓∆ni-SISDR ↓∆ni-SRMR ↓ content 0.170 0.051 24.777 1.091 3.192 11.606 0.378 0.024 0.054 0.854 5.439 1.407 WPE+FINS 0.231 0.070 24.452 1.368 5.135 8.507 0.394 0.104 0.133 1.031 9.792 1.956 DFNET+FINS 0.263 0.067 20.859 1.437 4.611 7.998 0.228 0.064 0.065 0.606 6.317 1.489 oracle+FINS 0.170 0.052 19.400 1.123 3.786 10.323 0.135 0.016 0.032 0.332 3.341 1.282 CWUNET-mel 0.136 0.042 18.803 1.009 2.471 12.696 0.182 0.022 0.035 0.553 4.120 1.345 CWUNET-stft 0.133 0.039 18.240 0.944 2.410 12.882 0.197 0.022 0.037 0.583 4.245 1.338 content WPE+FINS CWUNET-mel DFNET+FINS CWUNET-stft oracle+FINS target 0 20 40 60 80 100 MUSHRA score ref.signals our models baselines semi-oracle *** n.s. * Fig. 4: Results of the MUSHRA listening test: individual ratings overlaid with error bars. The mean of each condition is plotted as a point. The error bars represent the 95% confidence interval around the mean. Asterisks and n.s. indicate statistical significance (n.s.-not significant, * p < 0.05 , *** p < 0.001 ). All remaining pairwise comparisons between conditions had a highest significance level (***). model, achieving the best results of all non-oracle models. Next, DFNET+FINS showed slightly lower performance, though the difference was not statistically significant (p= 2.4e-1). When judged by humans, our CWUNET-mel model received significantly lower scores than DFNET+FINS ( p= 2.0e-4 ), which suggests that objective metrics may overestimate the performance of our models — either because they fail to capture certain distortions or because they overemphasize spectral differences, thus penalizing the baseline. This could also be attributed to the similarity between the evaluation metrics and the loss functions employed during training: our models may have the advantage of aligned optimization objectives with evaluation criteria. The WPE+FINS condition yields the lowest scores, which is consistent with objective scores and is perhaps caused by the limited effectiveness of WPE when applied as a single-channel dereverberation method [4], [12]. Notably, no objective metric clearly matches the trends shown by the listening test. The rightmost column of Table 1 shows the Pearson correlation between the mean human scores and the metric values obtained for the audio samples used in the listening test. The highest alignment with human scores is obtained by ∆ni-STOI ( r=−0.7 ), and the lowest by FWSNR ( r=−0.22 ). Future work is needed to investigate these differences in more depth. In summary, our method shows promising results outperforming baseline methods that relied on a two-step process of dereverberation followed by RIR estimation. However, there are still a few limitations to address in future work. First, the model introduces distortions that should be reduced, possibly through developing novel loss metrics focused on reverberation or through improved training strategies. Second, the current dataset lacks diversity in room types and source-receiver configurations. Expanding it to include more realistic scenarios will be essential to better assess its performance in practical applications. Finally, despite using a range of established audio metrics, our results show important discrepancies between perceptual ratings and objective scores. This highlights the need for new audio similarity measures that better capture reverberation characteristics. 4. CONCLUSION To address the problem of mismatched acoustics in remote mixed reality conversations, we propose a method for matching the reverberation between two speech segments with arbitrary reverberation. Our approach replaces traditional dereverberation and rereverberation steps with a time-domain wave-u-net conditioned on reverberation embeddings. It outperforms two baselines relying on a two-step approach and presents a novel use case for the successful reverb encoder from [6], extending its application to a broader task. REFERENCES [1] A. Neidhardt, C. Schneiderwind, and F. Klein, “Perceptual matching of room acoustics for auditory augmented reality in small rooms-literature review and theoretical framework,” Trends in Hearing, vol. 26, p. 23312165221092919, 2022. [2] P. Ochieng, “Deep neural network techniques for monaural speech enhancement and separation: state of the art analysis,” Artificial Intelligence Review, vol. 56, no. Suppl 3, pp. 3651–3703, 2023. [3] Y. Hua, “Blind methods of system identification,” Circuits, Systems and Signal Processing, vol. 21, pp. 91–108, 2002. [4] L. Zhao, W. Zhu, S. Li, H. Luo, X.-L. Zhang, and S. Rahardja, “Multi-resolution convolutional residual neural networks for monaural speech dereverberation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024. [5] I. Martin, F. Pastor, F. Fuentes-Hurtado, J. Belloch, L. Azpicueta-Ruiz, V. Naranjo, and G. Pi ˜ nero, “Predicting room impulse responses through encoder-decoder convolutional neural networks,” in 2023 IEEE 33rd International Workshop on Machine Learning for Signal Processing (MLSP). IEEE, 2023, pp. 1–6. [6] C. J. Steinmetz, V. K. Ithapu, and P. Calamia, “Filtered noise shaping for time domain room impulse response estimation from reverberant speech,” in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2021, pp. 221–225. [7] F. Llu ´ ıs and N. Meyer-Kahlen, “Blind spatial impulse response generation from separate room-and scene-specific information,” in ICASSP 20252025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5. [8] J. Eaton, N. D. Gaubitch, A. H. Moore, and P. A. Naylor, “Estimation of room acoustic parameters: The ace challenge,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, no. 10, pp. 1681– 1693, 2016. [9] N. J. Bryan, “Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 1–5. [10] S. Lee, H.-S. Choi, and K. Lee, “Differentiable artificial reverberation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 2541–2556, 2022. [11] G. Dal Santo, K. Prawda, S. Schlecht, and V. V ¨ alim ¨ aki, “Differentiable feedback delay network for colorless reverberation,” in International Conference on Digital Audio Effects. Aalborg University, 2023, pp. 244–251. [12] J. Koo, S. Paik, and K. Lee, “Reverb conversion of mixed vocal tracks using an end-to-end convolutional deep neural network,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 81–85. [13] A. Sarroff and R. Michaels, “Blind arbitrary reverb matching,” in Proceedings of the 23rd International Conference on Digital Audio Effects (DAFx-2020), vol. 2, 2020. [14] J. Im and J. Nam, “Diffrent: A diffusion model for recording environment transfer of speech,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 7425–7429. [15] J. Su, Z. Jin, and A. Finkelstein, “Acoustic matching by embedding impulse responses,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 426–430. [16] A. Mathur, A. Isopoussu, F. Kawsar, N. Berthouze, and N. D. Lane, “Mic2mic: using cycle-consistent generative adversarial networks to overcome microphone variability in speech systems,” in Proceedings of the 18th international conference on information processing in sensor networks, 2019, pp. 169–180. [17] C. Chen, R. Gao, P. Calamia, and K. Grauman, “Visual acoustic matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18 858–18 868. [18] T. Yoshioka, T. Nakatani, M. Miyoshi, and H. G. Okuno, “Blind separation and dereverberation of speech mixtures by joint optimization,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 1, pp. 69–84, 2010. [19] H. Schroter, A. N. Escalante-B, T. Rosenkranz, and A. Maier, “Deepfilternet: A low complexity speech enhancement framework for full-band audio based on deep filtering,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7407–7411. [20] L. A. Gatys, A. S. Ecker, and M. Bethge, “A neural algorithm of artistic style,” arXiv preprint arXiv:1508.06576, 2015. [21] D. Stoller, S. Ewert, and S. Dixon, “Wave-u-net: A multi-scale neural network for end-to-end audio source separation,” arXiv preprint arXiv:1806.03185, 2018. [22] A. D ´ efossez, “Hybrid spectrogram and waveform source separation,” arXiv preprint arXiv:2111.03600, 2021. [23] N. Babaev, K. Tamogashev, A. Saginbaev, I. Shchekotov, H. Bae, H. Sung, W. Lee, H.-Y. Cho, and P. Andreev, “Finally: fast and universal speech enhancement with studio-like quality,” Advances in Neural Information Processing Systems, vol. 37, pp. 934–965, 2024. [24] S. Nakaoka, L. Li, S. Inoue, and S. Makino, “Teacher-student learning for low-latency online speech enhancement using wave-u-net,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 661–665. [25] V. Kuzmin, F. Kravchenko, A. Sokolov, and J. Geng, “Real-time streaming wave-u-net with temporal convolutions for multichannel speech enhancement,” arXiv preprint arXiv:2104.01923, 2021. [26] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018. [27] A. Bitton, P. Esling, and A. Chemla-Romeu-Santos, “Modulated variational auto-encoders for many-to-many musical timbre transfer,” arXiv preprint arXiv:1810.00222, 2018. [28] Y. Gu, R. Zhang, L. Juvela, and Z. Wu, “Diff-ssl-g-comp: Towards a large-scale and diverse dataset for virtual analog modeling,” arXiv preprint arXiv:2504.04589, 2025. [29] J. Yamagishi, C. Veaux, K. MacDonald et al., “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),” University of Edinburgh. The Centre for Speech Technology Research (CSTR), pp. 271–350, 2019. [30] A. Perez-Lopez and A. Politis, “A python library for multichannel acoustic signal processing,” in Audio Engineering Society Convention 148. Audio Engineering Society, 2020. [31] R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 6199–6203. [32] Y.-C. Wu, I. D. Gebru, D. Markovi ´ c, and A. Richard, “Audiodec: An open-source streaming high-fidelity neural audio codec,” in ICASSP 20232023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5. [33] N. Kambhatla and T. K. Leen, “Dimension reduction by local principal component analysis,” Neural computation, vol. 9, no. 7, pp. 1493–1516, 1997. [34] R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE pacific rim conference on communications computers and signal processing, vol. 1. IEEE, 1993, pp. 125–128. [35] J. Ma, Y. Hu, and P. C. Loizou, “Objective measures for predicting speech intelligibility in noisy conditions based on new band-importance functions,” The Journal of the Acoustical Society of America, vol. 125, no. 5, pp. 3387–3405, 2009. [36] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 01CH37221), vol. 2. IEEE, 2001, pp. 749–752. [37] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on audio, speech, and language processing, vol. 19, no. 7, pp. 2125–2136, 2011. [38] A. Kumar, K. Tan, Z. Ni, P. Manocha, X. Zhang, E. Henderson, and B. Xu, “Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5. [39] T. H. Falk, C. Zheng, and W.-Y. Chan, “A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, no. 7, pp. 1766–1774, 2010. [40] B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union Radiocommunication Assembly, vol. 2, 2014. [41] R. F. Woolson, “Wilcoxon signed-rank test,” Encyclopedia of biostatistics, vol. 8, 2005.