Full text
Master in Sound and Music Computing Universitat Pompeu Fabra Open-Domain Zero-Shot Audio Tagging: Evaluation via Semantic Embeddings Tolga Yapici Supervisor: Panagiota Anastasopoulou Co-Supervisor: Frederic Font July 2025
Contents 1Introduction 1 1.1 Motivation.................................. 1 1.2 KeyConcepts................................ 2 1.2.1 Folksonomy ................................. 2 1.2.2 Co-occurrence Based Tag Recommendation . . . . . . . . . . . . . . . 2 1.2.3 Zero-Shot Classification . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.2.4 Audio-Text Representation Models . . . . . . . . . . . . . . . . . . . . 3 1.2.5 LAION-CLAP................................ 4 1.3 Objectives.................................. 4 1.4 Scope .................................... 4 2Background 6 2.1 Freesound Tag Recommender (RankST) . . . . . . . . . . . . . . . . . 6 2.1.1 Co-occurrence Based Tag Recommendation . . . . . . . . . . . . . . . 6 2.1.2 Rank Aggregation and Adaptive Cutoff.................. 7 2.1.3 Limitations ................................. 8 2.2 Contrastive Language-Audio Pretraining . . . . . . . . . . . . . . . . . 8 2.2.1 CLAP Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . 8 2.2.2 Models&Benchmarks ........................... 9 2.2.3 PromptDesign ............................... 10 2.2.4 Challenges in Folksonomy-Based Tagging . . . . . . . . . . . . . . . . . 10 2.3 Semantic Evaluation of Tag Recommendations . . . . . . . . . . . . . . 11
2.3.1 Limitations of Exact Matching . . . . . . . . . . . . . . . . . . . . . . . 11 2.3.2 Embedding-Based Semantic Evaluation with SBERT . . . . . . . . . . 12 2.3.3 Application to Audio Tagging . . . . . . . . . . . . . . . . . . . . . . . 13 3Methods 14 3.1 Dataset and Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . 14 3.1.1 BSD10kDataset .............................. 14 3.1.2 Dataset Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 3.1.3 TagVocabulary............................... 16 3.2 LAION-CLAP Embeddings . . . . . . . . . . . . . . . . . . . . . . . . 16 3.2.1 AudioEmbeddings ............................. 16 3.2.2 TagEmbeddings .............................. 17 3.3 Tag Recommendation Systems . . . . . . . . . . . . . . . . . . . . . . . 17 3.3.1 RankST ................................... 17 3.3.2 Zero-ShotBaseline ............................. 18 3.3.3 Zero-Shot with DF Weighting . . . . . . . . . . . . . . . . . . . . . . . 18 3.4 Evaluation Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.4.1 EvaluationMetrics ............................. 19 4Results 21 4.1 Zero-Shot Performance . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 4.2 RankSTBenchmark ............................ 23 4.3 ResultsOverview.............................. 26 5Discussion 28 5.0.1 Zero-Shot Tagging Performance . . . . . . . . . . . . . . . . . . . . . . 28 5.0.2 Impact of Embedding-Based Semantic Evaluation . . . . . . . . . . . . 28 5.0.3 Tagset Quality and Noise . . . . . . . . . . . . . . . . . . . . . . . . . 29 5.0.4 Ground Truth Limitations . . . . . . . . . . . . . . . . . . . . . . . . . 29 5.0.5 System Performance Comparison . . . . . . . . . . . . . . . . . . . . . 30
6Conclusion 32 List of Figures 33 List of Tables 35 Bibliography 36 A First Appendix 39 BSecondAppendix 40
Dedication To my loving family, my dear friends, and the beautiful city of Barcelona.
Acknowledgement I thank my supervisors, Penny and Frederic, for their support and sincerity, and my friends in the Master’s program and the Music Technology Group who have supported my growth academically and beyond.
4Chapter 1. Introduction zero-shot audio classification by comparing similarity between audio and text embeddings, typically using cosine similarity [3][4]. 1.2.5 LAION-CLAP LAION-CLAP (Contrastive Language-Audio Pretraining) is an open-source audiotext representation model trained on over 630,000 audio–text pairs spanning diverse audio content. Its architecture (shown in Figure 4) uses separate encoders for audio and text, optimized via contrastive loss to produce aligned audio-text embeddings [4]. 1.3 Objectives This thesis aims to evaluate whether a zero-shot audio-text model can effectively recommend tags for heterogeneous audio content without supervised training on the target dataset. The specific objectives are: •Benchmark the performance of LAION-CLAP for tag recommendation in a zero-shot setting on Freesound. •Compare zero-shot predictions with Freesound’s co-occurrence-based predictions in tagging accuracy. •Evaluate whether semantic similarity metrics (via SBERT) capture meaningful tags missed by exact string matching. 1.4 Scope This thesis evaluates zero-shot audio tagging performance on a curated heterogenous subset of Freesound (BSD10k), representing diverse general-purpose audio content. Evaluation focuses on methods incorporating tag weighting via document frequency and semantic evaluation through sentence embeddings. Performance is measured using precision, recall, and F1-score computed under both exact string matching
1.4. Scope 5 and semantic matching criteria. Evaluations are performed in a zero-shot setting without fine-tuning or retraining of underlying models.
Chapter 2 Background 2.1 Freesound Tag Recommender (RankST) 2.1.1 Co-occurrence Based Tag Recommendation The RankST algorithm (Font et al. 2012) is the foundation of Freesound’s tag recommendation system. It estimates tag relevance from tag-tag co-occurrence statistics. Tag co-occurrence counts are used to compute the tag similarity matrix S=DD>, where each element sij indicates the number of sounds in which tags tiand tjappear together [2]. Given a set of input tags, RankST retrieves a list of candidate tags, taking the top-Nmost similar tags to each input tag. For example, for an input tag drums,thesystemmaysuggestrelatedtagssuchaspercussion or rhythm if these frequently co-occur in the annotations. This approach relies on the principle that if two tags frequently co-occur, they are likely to describe related sound content. 6
2.1. Freesound Tag Recommender (RankST) 7 Figure 3: Visualization of Tag Similarity Matrix S[2] 2.1.2 Rank Aggregation and Adaptive Cutoff RankST aggregates candidate tags using rank aggregation. Each input tag produces acandidatelistofCinputT agiof its top-Nco-occuring tags. For each candidate at position nin CinputT agi,rankvaluesareassignedas rank(neighborn)=N(n1) The closest tag to every input tag is assigned N,thesecondN1,andsoforth. Using this aggregation method, similarity is represented by these rank values, which are summed across the input tags. This gives greater weight to tags that consistently appear at the top of the co-occurrence lists. RankST then applies an adaptive cutoff based on the Anderson-Darling test to automatically determine how many tags to recommend, which improves F-measure performance compared to fixed-length outputs [2].
8Chapter 2. Background 2.1.3 Limitations RankST relies entirely on tag metadata and is constrained by the user-generated folksonomy and its inconsistencies. Its recommendations are limited to previously seen tags, and the quality of the output is heavily dependent on the seed tags, making the system sensitive to subjective or inconsistent annotations. In addition, noise in the folksonomy propagates into the recommendations. Spelling variations, user-specific tags, non-semantic terms, and misspellings often appear in the output providing no semantic utility. For instance, for the inputs drum and percussion,therecommendationsaredrums,percussive,percs,andusernametokens such as sandyrb (a username tagged across many uploads), seen in Figure 1. Further limitations of this system are the absence of audio content analysis and the seed requirement. As a result, this system cannot predict labels directly from the audio signal itself and cannot operate in a cold-start setting. 2.2 Contrastive Language-Audio Pretraining 2.2.1 CLAP Model Architecture CLAP models are trained on large corpora of paired audio–text data, enabling them to learn joint audio–text representations by training separate encoders for each modality [3][4]. In LAION-CLAP architecture, the audio encoder (e.g., PANN or HTSAT) processes mel-spectrograms, while the text encoder (e.g., BERT or RoBERTa) embeds natural language [4] (Figure 4). These embeddings are projected into a joint 512-dimensional latent space and optimized via contrastive learning, maximizing the similarity of semantically aligned audio-text pairs while minimizing it for unrelated pairs, a method adapted from CLIP [3][4][7]. With this framework these models are able to learn semantically aligned cross-modal representations making them suitable for downstream tasks such as text-to audio retrieval, zero-shot audio classification, supervised audio classifica-
2.2. Contrastive Language-Audio Pretraining 9 tion, and enabling them to generalize well to unseen categories without fine-tuning [4]. This framework enables inference through cosine similarity between an audio embedding and a set of candidate text embeddings, enabling audio classification in zero-shot settings [4]. Zero-shot performance of CLAP models have been validated on benchmark datasets such as ESC-50, UrbanSound8K and VGGSound, where they perform competitively with supervised models in classification accuracy [3][4]. Figure 4: Contrastive Language-Audio Pretraining Architecture [8] 2.2.2 Models & Benchmarks Several CLAP variants have been developed, each differing in model capacity, training data scale, and application domain. MS-CLAP is trained on 128k diverse audio–text pairs [3], while LAION-CLAP scales training to over 630k pairs [4]. On closed-domain benchmarks, MS-CLAP achieves 82.6% zero-shot classification
10 Chapter 2. Background accuracy on ESC-50 and 73.2% on UrbanSound8K. LAION-CLAP demonstrates improved performance with 89.1% accuracy on ESC-50 and 73.2% on UrbanSound8K [3][4]. LAION-CLAP’s larger training scale and data diversity suggest greater potential for generalization across diverse audio content. However, on VGGSound, a large opendomain dataset with over 310 sound classes, LAION-CLAP’s accuracy drops to 29.1%, highlighting the increased challenge of classification in open-domain settings [4]. Dataset Domain # Classes CLAP(MS) LAION-CLAP ESC-50 [9] Environmental 50 82.6% 89.1% US8K [10] Urban 10 73.2% 73.2% VGGSound [11] Open-domain 310+ N/A 29.1% Table 1: Zero-shot classification accuracy of CLAP models on benchmark datasets 2.2.3 Prompt Design LAION-CLAP is trained on audio–text pairs structured as natural language sentences, typically of the form This is a sound of [label] [4]. At inference time, the phrasing of candidate class tags significantly influences the model’s zero-shot classification accuracy. Prompts structured as natural language sentences (e.g., “This is a sound of a dog barking”) align more closely with the model’s training distribution, resulting in higher audio-text similarity scores and improved classification performance. The study by Olvera et al. (2024) demonstrates that prompting with complete sentences and detailed acoustic descriptions consistently outperforms using isolated labels [12]. This makes prompt formatting a critical design choice in zero-shot tagging and classification with CLAP-based systems. 2.2.4 Challenges in Folksonomy-Based Tagging Unlike standardized class labels typical of benchmark datasets, folksonomies consist of user-generated tags that are often noisy, inconsistent, and semantically overlapping [2]. As discussed in the Analysis of the Folksonomy of Freesound [13], the
2.3. Semantic Evaluation of Tag Recommendations 11 Freesound folksonomy is characterized by a continuously growing and largely uncontrolled vocabulary (Figure 5), where users label sounds without constraints leading to inconsistencies in describing audio content. In this setting, CLAP must score thousands of candidate tags with widely varying granularity and relevance, increasing the risk of inaccurate or irrelevant predictions. These conditions increase the complexity for adaptation of zero-shot models to heterogeneous vocabularies. Figure 5: Number of new tags introduced every month to Freesound (2005-2012) [13] 2.3 Semantic Evaluation of Tag Recommendations 2.3.1 Limitations of Exact Matching Conventional evaluation of classification systems relies on exact string matching, where a predicted class label must match the ground truth. This approach fails to account for semantic equivalence between lexically different tags, such as synonyms (e.g., birdsong vs. chirping), plural forms (e.g., drum vs. drums), or spelling variants (e.g., color vs. colour). This leads to an underestimation of system performance, particularly in tag recommendation tasks where the vocabulary is highly variable.
12 Chapter 2. Background 2.3.2 Embedding-Based Semantic Evaluation with SBERT Bidirectional Encoder Representations from Transformers (BERT) is a language model that produces token-level contextual embeddings, based on the multi-layer bidirectional Transformer encoder introduced by Vaswani et al. (2023) [14][15]. Sentence-BERT (SBERT) adapts BERT with siamese and triplet network training to generate sentence-level embeddings that capture semantic similarity beyond token-level representations [16]. Building on this, embedding-based approaches have been adopted in tasks such as audio caption quality assessment, measuring semantic similarity between predicted and reference captions [17], effectively addressing challenges of lexical variability. Figure 6: BERT Transformer encoder architecture based on Vaswani et al. 2017 [15][18]
2.3. Semantic Evaluation of Tag Recommendations 13 Figure 7: SBERT siamese adaptation of BERT architecture [16] Figure 8: Audio caption semantic similarity measurement framework [17] 2.3.3 Application to Audio Tagging Despite its adoption in captioning, semantic similarity has been underexplored as an evaluation method for tag recommendation systems. In this work, SBERT embeddings are used to evaluate tag recommendations, which are often multi-word concepts. Cosine similarity between embeddings of predicted and reference tags measures semantic relevance, effectively addressing the lexical variability of folksonomybased ground truth and capturing semantically relevant tags missed by exact string matching.
20 Chapter 3. Methods F1-score computes the harmonic mean of Precision and Recall: F1=2·P@10 ·R@10 P@10 + R@10 3.4.1.2 Semantic Matching Semantic similarity is computed using the all-MiniLM-L6-v2 SBERT model to encode predicted and ground-truth tags into 384-dimensional vectors. For each predicted tag tp, cosine similarity is calculated with all ground-truth tags tgt: sim(tp,t gt)=cos(embedding tp,embeddingtgt ) where embeddingtpand embeddingtgt denote SBERT embeddings. A predicted tag is a semantic match if max tgt2TGT sim(tp,t gt)⌧ where ⌧is a threshold parameter. Performance is evaluated using Precision@10, Recall@10, and F1 metrics using semantic matches. tptgt Cos. sim. Match (⌧=0.7) beat beat 1.000 X drums drum 0.867 X rhythmic rhythm 0.863 X vocal voice 0.824 X woman female 0.799 X metallic metal 0.884 X breaking break 0.920 X Figure 13: Semantic matching example for predicted tags tpagainst ground-truth tags tgt,collectedacrossmultipleaudioclips.
Chapter 4 Results Performance is reported using Precision@10, Recall@10, and F1 metrics as detailed in Section 3.4.1, under both exact string matching (Section 3.4.1.1) and semantic matching (Section 3.4.1.2) conditions. 4.1 Zero-Shot Performance 4.1.0.1 Exact Matching System Precision@10 Recall@10 F1 ZS Baseline 0.0051 ±0.0240 0.0093 ±0.0497 0.0062 ±0.0292 ZS DF-Weighted (↵=0.7)0.0305±0.0738 0.0515 ±0.1248 0.0367 ±0.0876 Table 3: Zero-shot tagging performance of CLAP-based systems under exact matching conditions. Reported as mean ±standard deviation across test clips. 21
22 Chapter 4. Results 4.1.0.2 Semantic Matching System Precision@10 Recall@10 F1 ZS Baseline 0.0130 ±0.0467 0.0230 ±0.0968 0.0157 ±0.0570 ZS DF-Weighted (↵=0.7)0.0488±0.1099 0.0837 ±0.1891 0.0590 ±0.1309 Table 4: Zero-shot tagging performance of CLAP-based systems under semantic matching conditions (⌧=0.7). Reported as mean ±standard deviation across test clips. 4.1.0.3 Performance Overview Figure 14: F1 performance of Zero-Shot systems (exact and semantic matching)
4.2. RankST Benchmark 23 Figure 15: Number of clips with 1hit for Zero-Shot systems (exact and semantic matching) 4.2 RankST Benchmark 4.2.0.1 Exact Matching System Precision@10 Recall@10 F1 RankST (k=1)0.0808±0.1285 0.1550 ±0.2445 0.0997 ±0.1513 RankST (k=2)0.1333±0.1441 0.2703 ±0.2875 0.1687 ±0.1740 RankST (k=3)0.1774±0.1598 0.3540 ±0.3070 0.2236 ±0.1890 Table 5: Tagging performance of RankST for k2{1,2,3}under exact matching conditions. Reported as mean ±standard deviation across test clips.
24 Chapter 4. Results 4.2.0.2 Semantic Matching System Precision@10 Recall@10 F1 RankST (k=1)0.1127±0.1609 0.2341 ±0.3585 0.1444 ±0.2056 RankST (k=2)0.1734±0.1747 0.3594 ±0.3756 0.2212 ±0.2157 RankST (k=3)0.2181±0.1826 0.4434 ±0.3754 0.2765 ±0.2197 Table 6: Tagging performance of RankST for k2{1,2,3}under semantic matching conditions (⌧=0.7). Reported as mean ±standard deviation across test clips. 4.2.0.3 Performance Overview Figure 16: F1 performance of RankST (k2{1,2,3})(exactandsemanticmatching)
4.2. RankST Benchmark 25 Figure 17: Number of clips with 1hit for RankST (k2{1,2,3})(exactand semantic matching)
26 Chapter 4. Results 4.3 Results Overview Exact Matching Semantic Matching System P R F1 Hits P R F1 Hits F1(%) Hits(%) ZS Baseline 0.005 0.009 0.006 72 0.013 0.023 0.016 136 +166.7 +88.9 ZS DF 0.031 0.052 0.037 293 0.049 0.084 0.059 369 +59.5 +25.9 RankST (k=1)0.0810.1550.100 651 0.1130.2340.144 739 +44.0 +13.5 RankST (k=2)0.1330.2700.169 983 0.1730.3590.2211053 +30.8 +7.1 RankST (k=3)0.1770.3540.22411450.2180.4430.2771200 +23.7 +4.8 Table 7: Performance across all systems. Reported values are mean Precision@10 (P), Recall@10 (R), F1 and clips with 1correct hit (Hits). F1(%) and Hits(%) values are exact vs semantic. Figure 18: F1 performance across all systems (exact and semantic matching)
4.3. Results Overview 27 Figure 19: Number of clips with 1hit across all systems (exact and semantic matching)
Chapter 5 Discussion 5.0.1 Zero-Shot Tagging Performance The baseline zero-shot system achieves an F1 score of 0.006 under exact matching, highlighting the inherent challenge of open-domain audio tagging without taskspecific training. Implementing normalized logarithmic document frequency (DF) weighting (Section 3.3.3) yields a substantial relative improvement, increasing the F1 score to 0.037, a 516.7% increase over the baseline. This approach effectively down-weights rare tags, which often correspond to uninformative labels, attenuating noise and improving discriminative performance in zero-shot tagging. 5.0.2 Impact of Embedding-Based Semantic Evaluation Zero-shot systems demonstrate substantial improvements under semantic evaluation compared to exact matching. The zero-shot baseline model achieves a 166.7% increase in F1 score and an 88.9% increase in clips with at least one correct prediction. The DF-weighted system similarly achieves proportional gains under semantic evaluation, with a 59.5% increase in F1 score and a 25.9% increase in clips with at least one correct prediction. Comparing the zero-shot baseline under exact matching (F1 = 0.006) to the DFweighted system under semantic matching (F1 = 0.059) reveals an 883% relative im28
29 provement in F1, demonstrating substantial gains in zero-shot performance through weighted tagging and semantic evaluation, which reveals latent performance missed by exact matching. In addition, semantic evaluation captures latent performance for RankST (up to a 44% increase in F1) highlighting the broader potential of embedding-based evaluation in MIR tasks. 5.0.3 Tagset Quality and Noise The zero-shot tag vocabulary, comprising 1,870 unique tags, provides broad semantic coverage for open-domain audio tagging. Compared to curated open-domain datasets such as FSD50K[22] (200 classes) and VGGSound[11] (310+ classes), the zero-shot vocabulary is substantially larger. However, the extensive tagset is not indicative of higher semantic utility; analysis reveals considerable variation in tag quality with a large fraction of the tagset comprising noise. These include overly specific and semantically uninformative tags, such as proper nouns, technical identifiers, subjective terms and fragments (Table 8). While DF weighting provides an immediate mechanism to attenuate noisy tags, it does not fully eliminate noise inherent to the folksonomy. This inflated vocabulary impairs CLAP generalization, compromising zero-shot tagging performance. Proper nouns barcelona, japan, nasa, sony, ableton Technical 16bit, bpm, midi, mono, vst, h4n Subjective nice, bad, cool, yes, no Fragments a, el, la, c3, fm, xy Table 8: Examples of noisy tags sampled from the zero-shot tag vocabulary 5.0.4 Ground Truth Limitations Ground truth comprises user-generated annotations exhibiting inconsistent descriptive coverage and high sparsity. This limits evaluation, as the system may produce relevant predictions not covered by the ground truth (Figure 20). Semantic
Bibliography [1] 2024 in numbers | The Freesound Blog. URL https://blog.freesound.org/ ?p=2141. [2] Font Corbera, F., Serrà Julià, J. & Serra, X. Folksonomy-based tag recommendation for online audio clip sharing (2012). URL http://hdl.handle.net/ 10230/22736. Publisher: International Society for Music Information Retrieval (ISMIR). [3] Elizalde, B., Deshmukh, S., Ismail, M. A. & Wang, H. CLAP Learning Audio Concepts from Natural Language Supervision. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP),1–5(2023). URLhttps://ieeexplore.ieee.org/abstract/ document/10095889.ISSN:2379-190X. [4] Wu, Y. et al. Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation (2024). URL http://arxiv. org/abs/2211.06687.ArXiv:2211.06687[cs]. [5] Folksonomy :: vanderwal.net. URL https://vanderwal.net/folksonomy. html. [6] Freesound. URL https://freesound.org/. [7] Radford, A. et al. Learning Transferable Visual Models From Natural Language Supervision (2021). URL http://arxiv.org/abs/2103.00020. ArXiv:2103.00020 [cs]. 36
BIBLIOGRAPHY 37 [8] LAION-AI/CLAP (2025). URL https://github.com/LAION-AI/CLAP. Original-date: 2022-03-06T20:12:49Z. [9] Piczak, K. J. ESC: Dataset for Environmental Sound Classification. In Proceedings of the 23rd ACM international conference on Multimedia, 1015–1018 (ACM, Brisbane Australia, 2015). URL https://dl.acm.org/doi/10.1145/ 2733373.2806390. [10] Salamon, J., Jacoby, C. & Bello, J. P. A Dataset and Taxonomy for Urban Sound Research. In Proceedings of the 22nd ACM international conference on Multimedia, 1041–1044 (ACM, Orlando Florida USA, 2014). URL https: //dl.acm.org/doi/10.1145/2647868.2655045. [11] Chen, H., Xie, W., Vedaldi, A. & Zisserman, A. VGGSound: A Largescale Audio-Visual Dataset (2020). URL http://arxiv.org/abs/2004.14368. ArXiv:2004.14368 [cs]. [12] Olvera, M., Stamatiadis, P. & Essid, S. A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification (2024). URL http://arxiv.org/abs/2409.13676.ArXiv:2409.13676 [cs]. [13] Font, F. & Serra, X. ANALYSIS OF THE FOLKSONOMY OF FREESOUND (2012). [14] Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019). URL http://arxiv.org/abs/1810.04805.ArXiv:1810.04805[cs]. [15] Vaswani, A. et al. Attention Is All You Need (2023). URL http://arxiv.org/ abs/1706.03762.ArXiv:1706.03762[cs]. [16] Reimers, N. & Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019). URL http://arxiv.org/abs/1908.10084. ArXiv:1908.10084 [cs].
38 BIBLIOGRAPHY [17] Mahfuz, R., Guo, Y. & Visser, E. Improving Audio Captioning Using Semantic Similarity Metrics (2023). URL http://arxiv.org/abs/2210.16470. ArXiv:2210.16470 [cs]. [18] Smith, B. A Complete Guide to BERT with Code (2024). URL https://towardsdatascience.com/ a-complete-guide-to-bert-with-code-9f87602e4a11/. [19] GitHub - allholy/BSD10k. URL https://github.com/allholy/BSD10k. [20] Anastasopoulou, P., Torrey, J., Serra, X. & Font, F. Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset (2024). URL http: //arxiv.org/abs/2410.00980.ArXiv:2410.00980[cs]version:1. [21] MTG/freesound (2025). URL https://github.com/MTG/freesound.Originaldate: 2012-11-07T18:25:03Z. [22] Fonseca, E., Favory, X., Pons, J., Font, F. & Serra, X. FSD50K: An Open Dataset of Human-Labeled Sound Events (2022). URL http://arxiv.org/ abs/2010.00475.ArXiv:2010.00475[cs].