Full text
Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack Yuxuan Liu, Rui Sang, Peihong Zhang, Zhixin Li, and Shengchen Li Xi’an Jiaotong-Liverpool University, Suzhou, 215123, China {yuxuan.liu2204, rui.sang22, peihong.zhang20, zhixin.li22}@student.xjtlu.edu.cn, [email protected] Abstract. Music Information Retrieval (MIR) systems are highly vulnerable to adversarial attacks that are often imperceptible to humans, primarily due to a misalignment between model feature spaces and human auditory perception. Existing defenses and perceptual metrics frequently fail to adequately capture these auditory nuances, a limitation supported by our initial listening tests showing low correlation between common metrics and human judgments. To bridge this gap, we introduce Perceptually-Aligned MERT Transformer (PAMT), a novel framework for learning robust, perceptually-aligned music representations. Our core innovation lies in the psychoacoustically-conditioned sequential contrastive transformer, a lightweight projection head built atop a frozen MERT encoder. PAMT achieves a Spearman correlation coe!cient of 0.65 with subjective scores, outperforming existing perceptual metrics. Our approach also achieves an average of 9.15% improvement in robust accuracy on challenging MIR tasks, including Cover Song Identification and Music Genre Classification, under diverse perceptual adversarial attacks. This work pioneers architecturally-integrated psychoacoustic conditioning, yielding representations significantly more aligned with human perception and robust against music adversarial attacks. 1Introduction Music Information Retrieval (MIR) systems play an essential role in numerous applications, including music recommendation [21,28], genre classification [8,14], and cover song identification [6,12]. These systems rely on machine learning models to extract and interpret intricate musical features. However, recent studies have shown that MIR models are vulnerable to adversarial attacks, where small perturbations, imperceptible to human listeners, can cause significant prediction errors in these systems [23,27]. Such vulnerabilities not only threaten the reliability of MIR systems but also expose critical weaknesses in their robustness [27]. Consequently, developing a perceptual model to quantify the auditory similarity All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 748
Y. Liu et al. between perturbed and original music is important for adversarial attacks and defenses. Traditional adversarial attack and defense strategies often rely on mathematical norms (e.g., Lpnorms) [3,24] or signal-level metrics such as Signal-to-Noise Ratio (SNR) or Log-Spectral Distance (LSD) [9] to evaluate the auditory similarity of adversarial examples. However, these metrics fail to account for the complexities of human auditory perception [7]. For instance, due to phenomena such as auditory masking [1, 32], the same perturbation magnitude may have varying perceptual e!ects depending on the spectral context. Such limitations highlight the disconnect between conventional perceptual metrics and human auditory perception. Recent e!orts have sought to incorporate human perception into adversarial attack evaluation. For example, psychoacoustic models have been integrated into adversarial attack design to better align perturbations with human auditory thresholds [29]. However, these models often rely on fixed psychoacoustic parameters, making them less adaptable to diverse audio contexts and listener variations [4]. Another line of work leverages human auditory ratings to evaluate adversarial perturbations [7], but the use of non-di!erentiable alignment methods, such as Dynamic Time Warping, hinders integration with gradient-based deep learning frameworks [23]. In music generation tasks [10,19,20], researchers often utilize PEMO-Q [13] and Fréchet Audio Distance (FAD) [15] to measure the perceptual distance between generated music and the original track. However, there is currently no experimental evidence to confirm the e!ectiveness of these metrics in adversarial perturbation tasks. To address these challenges, we conducted a human listening test to assess the e"cacy of commonly used auditory metrics, including SNR, LSD, FAD, and PEMO-Q, in evaluating adversarial perturbations. Our results reveal that these metrics exhibit low to moderate correlation with human perceptual judgments, underscoring their limitations in accurately capturing the auditory similarity of adversarial examples. As shown in Figure 1, pre-trained music representation models yield a Fréchet Audio Distance (FAD) more closely aligned with human judgments than traditional signal-based metrics. Inspired by recent advances in music representation learning, particularly the MERT (Acoustic Music Understanding Model with Large-Scale Self-supervised Training [17], we hypothesize that such pre-trained representations can be further refined to better align with human perception of adversarial perturbations. However, standard music representation models like MERT are not explicitly designed to be robust against imperceptible adversarial perturbations. To address this limitation, we propose Perceptually-Aligned MERT Transformer (PAMT), a novel approach that leverages MERT’s rich music representations and projects them into a perceptually-aligned feature space. Our method employs a Transformer-based projection head that is conditioned on psychoacoustic perturbation parameters through Feature-wise Linear Modulation (FiLM) [26] and trained with a sequential contrastive objective. This design allows the model to learn representations that remain consistent for perceptually similar inputs Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 749
PAMT Perception Model Table 1. The six adversarial perturbation types simulate common adversarial attacks [3, 7, 32]. Perturbations are applied with varying parameter ranges to capture di"erent attack strengths and characteristics. Category Perturbation Type (Parameter Range) Lp-based L2noise: Gaussian noise with →ω→2↑ε,ε↓[0.01,1.0] ↔ RMSsignal. L→noise: Uniform noise with →ω→→↑ϑ,ϑ↓[0.001,0.01] ↔ max(|signal|). Psychoacoustic-based Bark-band noise: Additive noise in a randomly selected Bark band, scaled by a factor of [0.1,0.5] relative to the original band energy. Semantic-based Pitch shift: n↓U(↗5,5) semitones. Speed change: s↓U(0.80,1.20) factor. Dynamics-based Dynamic range compression: Threshold [↗30,↗10] dBFS, ratio [2 : 1,8:1]. while di!erentiating between perceptually distinct samples, even when the mathematical di!erences are subtle. By explicitly conditioning on psychoacoustic perturbation characteristics, our model becomes sensitive to the context-dependent nature of human auditory perception. This context sensitivity can help detect subtle perturbations in adversarial examples. 2HumanListeningTestandMetricEvaluation To motivate our approach, we first conducted a human listening study to quantify the perceptual similarity between original and adversarially perturbed music. We then evaluated how well existing objective auditory metrics correlate with these human judgments, thereby demonstrating the need for a more perceptuallyaligned model. 2.1 Data Preparation for Subjective Evaluation We created a dataset of reference and perturbed audio pairs from 1,000 unique tracks sampled across the FMA [5], GTZAN [31], and MTG-Jamendo [2] repositories. All tracks were resampled to 16kHz and segmented into 10-second clips. We applied six distinct adversarial perturbation types to the reference clips, as detailed in Table 1, reflecting common attack methodologies from the literature [3, 7, 32]. By varying the parameters for each type, we generated 18,000 unique reference-perturbed pairs for subjective evaluation. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 750
Y. Liu et al. 2.2 Subjective Listening Test Design We conducted two listening experiments with 200 volunteers (50 with music backgrounds), who used their own high-quality headphones in quiet environments. MOS-style Perceptual Similarity Rating In our first task, participants rated the perceptual similarity of reference-perturbed pairs on a 5-point scale (1: identical, 5: dissimilar). Each of the 18,000 pairs was rated by at least five volunteers, yielding our "raw scores" dataset. Two-Alternative Forced Choice (2AFC) Test To mitigate rating biases [22, 33], we also used a 2AFC test. For each reference audio Sref,participantswere presented with two of its perturbed versions, Spi and Spj, and chose which sounded more similar to Sref.Foreachofthe3,000referencetracks,all15unique pairs of its perturbations were compared (!6 2"→3000 = 45,000 trials). We derived a"2AFCscore"foreachperturbedsamplebycountinghowmanytimesitwas chosen as more similar, forming our "2AFC scores" dataset. 2.3 Evaluation of Existing Objective Auditory Metrics We then evaluated a suite of objective metrics against our human perceptual data. The metrics included traditional, psychoacoustic, and deep learning-based models: –Signal-to-Noise Ratio (SNR): Ratio of signal to perturbation power. – Log-Spectral Distance (LSD) [9]: MSE between log-magnitude spectra. –PEMO-Q [13]: Psychoacoustic audio quality model. – Fréchet Audio Distance (FAD) [15]: Using embeddings from VGGish [11], PANNS [16], and MERT-v0 [17]. – Speech Quality/Perception Models: CDPAM [18], NOMAD [25], and SESQA [30] for cross-domain evaluation. We quantified the agreement by calculating the Spearman rank correlation (ω) between each metric and our human scores (both raw and 2AFC). The results are presented in Figure 1. As shown in Figure 1, existing metrics exhibited low to moderate correlation with human judgments. For raw MOS-style scores, correlations ranged from ω=0.18 (SNR) to ω=0.35 (FAD with MERT-v0). The more robust 2AFC scores yielded slightly better results, with FAD-MERT-v0 achieving the highest correlation (ω=0.44). Speech-domain models showed similarly limited performance. These results highlight a significant gap: even the best-performing metric, FAD with MERT-v0 embeddings, achieves only a modest correlation with robust human perceptual data. This limitation of existing metrics motivates our work on PAMT, a model that learns representations directly aligned with human auditory perception for enhanced robustness. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 751
PAMT Perception Model Fig. 1. Spearman Correlation Coe!cients (ϖ) between objective auditory metrics and human perceptual similarity ratings (mean ±std. dev.). Correlations are shown for both raw MOS-style scores and 2AFC-derived scores. Objective metrics include SNR, LSD, PEMO-Q, FAD with various embeddings (VGGish [11], PANNS-CNN1416k [16], PANNS-Wavegram-Logmel [16], MERT-v0 [17]), and speech-domain models (CDPAM [18], NOMAD [25], SESQA [30]). 3PAMT:LearningPerceptually-AlignedRepresentations The established limitations of existing objective metrics in capturing human perceptual similarity under adversarial conditions (Section 2) necessitate novel representation learning techniques that explicitly embed principles of auditory perception. To this end, we introduce the Perceptually-Aligned MERT Transformer (PAMT) framework. PAMT transforms features from the general-purpose music understanding model, MERT [17], into a new latent space that is demonstrably robust to imperceptible perturbations and closely aligned with human auditory perception. The core of the PAMT framework is our novel projection head: the Psychoacoustically-Conditioned Sequential Contrastive Transformer (PCSCT). 3.1 PAMT Framework Overview The PAMT framework, illustrated in Figure 2, uses a frozen MERT encoder for feature extraction. These features are then refined by our novel PCSCT projection head, which is trained via a sequential contrastive loss and conditioned on psychoacoustic perturbation parameters. The key components are: 1. Frozen MERT Encoder: A pre-trained, frozen MERT-v0 model [17] extracts a sequence of 768-dimensional embeddings, EMERT ↑RT→768,from the input audio (resampled to 24kHz). Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 752
Y. Liu et al. Fig. 2. Architectural Overview of the Proposed Perceptually-Aligned MERT Transformer (PAMT) Framework. Audio is processed by a frozen MERT encoder. For training, a psychoacoustic perturbation module generates perturbed audio and its parameters. A Perturbation Parameter Encoder (PPE) creates a conditioning vector cperturb from these parameters. The core PCSCT projection head (a Transformer) processes MERT’s sequential embeddings, conditioned by cperturb via FiLM layers, to produce robust, perceptually-aligned sequential embeddings ZPAMT . 2. Psychoacoustic Perturbation Module: During training, this module creates a perturbed audio version Apert from a reference Aorig using imperceptible alterations (see Table 1) and outputs the perturbation parameters Pparams. 3. Perturbation Parameter Encoder (PPE): A2-layerMLPencodesPparams into a compact 64-dimensional conditioning vector cperturb. 4. PCSCT Projection Head: This core component (detailed in Sec. 3.2) is aTransformerthatprocessestheMERTsequenceEMERT,conditionedon cperturb,toproducerobust128-dimembeddingsZPAMT . 5. Sequential Contrastive Learning: The PCSCT is trained to minimize the distance between representations of original and perturbed audio pairs (Zorig PAMT ,Zpert PAMT ) while maximizing their distance to other samples in the batch. 3.2 Psychoacoustically-Conditioned Sequential Contrastive Transformer (PCSCT) The PCSCT projection head transforms MERT features into a new space where imperceptible psychoacoustic variations are minimized. This is achieved by explicitly conditioning the transformation on the perturbation’s characteristics. Architecture PCSCT is a 4-layer Transformer encoder with 4 attention heads, a hidden dimension of 256, and a position-wise FFN with an inner dimension of 1024. We use GELU activations and a Pre-LN configuration [26]. It takes the 768-dim EMERT sequence as input and, after a final linear projection, outputs a 128-dim sequence ZPAMT . Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 753
PAMT Perception Model Psychoacoustic Conditioning with FiLM The key innovation is conditioning via Feature-wise Linear Modulation (FiLM) layers [26]. The 64-dim conditioning vector cperturb from the PPE is projected to generate modulation parameters (εl,ϑ l)↑R256 for each Transformer layer l. These modulate the output hl of the FFN sub-layer: FiLM(hl,c perturb)=εl(cperturb)↓hl+ϑl(cperturb)(1) where ↓is element-wise multiplication. This allows PCSCT to adapt its feature transformation based on the specific perturbation profile to which it must learn invariance. Sequential Contrastive Learning Objective PCSCT is trained with an InfoNCE loss. For each pair of reference Eorig MERT and perturbed Epert MERT embeddings, we generate conditioned outputs Zorig PAMT and Zpert PAMT . The loss is: LPCSCT =↔E#log exp $sim(Zorig PAMT ,Zpert PAMT ) ω% & k↑Batch k↓=orig exp 'sim(Zorig PAMT ,Zk PAMT ) ϖ( +exp'sim(Zorig PAMT ,Zpert PAMT ) ϖ( ) (2) The similarity function, sim(U, V ),isthecosinesimilarityofthetime-pooled mean vectors of sequences Uand V. The temperature ϖis set to 0.1. 3.3 Guiding Adversarial Defense Once trained, the PAMT embeddings ZPAMT (specifically, the mean-pooled vectors ¯zPAMT ) provide a robust feature space for downstream MIR tasks. For adversarial defense, we use adversarial training within this space. The perceptual distance dPAMT (A↔,A)is defined as the L2 norm between the PAMT embeddings of two audio clips A↔and A: dPAMT (A↔,A)=↗mean_pool(ZPAMT (A↔)) ↔mean_pool(ZPAMT (A))↗2.(3) Adversarial examples A↔are crafted to maximize the task loss while being constrained by dPAMT (A↔,A)↘ϱ. The downstream model fεis trained by solving: min εE(A,y)↗D*max dPAMT (A→,A)↘ϑLCE!fε(mean_pool(ZPAMT (A↔))),y "+,(4) where LCE is the cross-entropy loss. This strategy leverages the perceptual robustness of ZPAMT for stronger defense. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 754
Y. Liu et al. 4 Experiments This section validates our Perceptually-Aligned MERT Transformer (PAMT), evaluating its ability to learn perceptually-aligned representations and its e!ectiveness in enhancing adversarial defense. 4.1 Datasets and Evaluation Protocol We use the human listening test dataset from Section 2, annotated with both MOS and the more robust 2AFC perceptual similarity scores. For all evaluations, we use an 80/20 train/test split. Our primary evaluation metric is the Spearman rank correlation (ω)betweenamodel’spredictedsimilarityandthehumanscores on the test set. 4.2 Methods Compared We compare PAMT against two main categories of methods: 1. Objective Auditory Metrics: This includes the full suite of signal-based, psychoacoustic, and FAD-based metrics previously evaluated in Section 2.3, which serve as established benchmarks. 2. Learning-based Representation Models: Our primary comparison is against a strong baseline and our proposed model. – MERT+MLP-Contrastive: A strong baseline using the same frozen MERTv0 encoder and psychoacoustic augmentations as PAMT, but with a simpler MLP projection head trained with standard contrastive loss. This isolates the benefits of our PCSCT architecture. –PAMT (PCSCT - Ours): Our proposed framework, featuring the psychoacoustically conditioned, sequence-aware PCSCT projection head trained with a sequential contrastive loss. 4.3 Training Details for PAMT (PCSCT) We train the PCSCT head using the AdamW optimizer (LR = 1→10≃4,weight decay = 1→10≃5) with a batch size of 32. We use a cosine annealing schedule with a 10% warm-up and train for up to 100 epochs, with early stopping based on the validation set Spearman correlation (patience=10). The contrastive loss temperature ϖis 0.1. The PPE is a 2-layer ReLU MLP producing a 64-dim vector. Experiments were run on NVIDIA A100 GPUs using PyTorch. 4.4 Correlation with Human Perception Table 2 presents the Spearman correlation (ω)andaclassification-basedF1score against the 2AFC human judgments. The F1score evaluates the model’s utility in distinguishing noticeable from imperceptible perturbations. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 755
PAMT Perception Model Table 2. Evaluation of Perceptual Alignment on the Test Set. Spearman correlation (ϖ)andF1score (%) are reported against 2AFC human similarity judgments. Higher values are better. Method Spearman (ϖ) F1(%) Traditional & Speech Metrics SNR 0.261 56.1 LSD 0.315 57.0 PEMO-Q 0.307 58.1 CDPAM 0.22 53.5 NOMAD 0.21 53.1 SESQA 0.27 55.2 FAD with Pre-trained Embeddings FAD (VGGish) 0.33 57.9 FAD (PANNS-CNN14-16k) 0.37 59.5 FAD (PANNS-Wavegram-Logmel) 0.38 60.3 FAD (MERT-v0 raw) 0.44 63.1 MERT with Projection Heads MERT + MLP (Std. Contrastive) 0.55 68.5 PAMT (PCSCT - Ours) 0.65 81.2 The results clearly show PAMT’s superiority. It achieves a Spearman correlation of ω=0.65,substantiallyoutperformingallothermethods,includingFAD with raw MERT embeddings (ω=0.44) and the strong MERT+MLP contrastive baseline (ω=0.55). The high F1score of 81.2% further confirms its e!ectiveness. This significant improvement demonstrates that our psychoacoustic conditioning and sequence-aware training are crucial for learning representations that genuinely align with human auditory perception. 4.5 E!ectiveness in Adversarial Defense We test PAMT’s utility for adversarial defense on Cover Song Identification (CSI) and Music Genre Classification (MGC). We compare standard training (No Defense), adversarial training in the input audio space (Standard AT), and AT in the embedding spaces of our MERT+MLP baseline and our proposed PAMT.Robustnessismeasuredbytheworst-caseperformance("UnionRobust Acc.") against a set of strong perceptual attacks. As shown in Table 3, undefended models collapse under perceptual attacks. Critically, adversarial training in our PAMT embedding space yields the highest robust accuracy in both CSI (0.535 mAP) and MGC (0.465 Acc), significantly outperforming standard AT and AT in the baseline embedding space, all while maintaining high performance on clean data. This confirms that the superior perceptual alignment of PAMT embeddings translates directly into more e!ective and practical adversarial defenses. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 756