International Journal of Engineering and Advanced Technology (IJEAT) ISSN: 2249-8958 (Online), Volume-15 Issue-2, December 2025 1 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org Abstract. Preserving Indian folk music in digital repositories poses significant challenges because robust classification systems are lacking to capture its linguistic, instrumental, and acoustic diversity. As a cornerstone of India’s intangible cultural heritage, this music faces the risk of marginalisation and loss unless systematic, scalable methods are employed to identify and preserve it. This research aims to develop an automated, multi-modal framework for regional classification of Indian folk music, thereby enabling structured archiving and improved accessibility. To achieve this, a novel machine learning pipeline was designed, integrating Whisper for speech recognition and regional language identification [3]. Instrument detection was performed using YAMNet, which has proven effective in recognizing traditional instruments [14]. Acoustic features such as MFCCs, chroma, and spectral descriptors were extracted using Librosa [Error! R eference source not found.]. Together, these tools provide a comprehensive understanding of the songs’ linguistic, instrumental, and rhythmic content. The curated dataset includes folk music from linguistically rich regions of India, such as Marathi, Punjabi, Urdu Qawwali, and dialects from Uttar Pradesh and Bihar. Seven supervised learning algorithms were trained and evaluated, including Random Forest, Support Vector Machine, and Gradient Boosting. Simpler classifiers, such as K-Nearest Neighbours, Naive Bayes, and Logistic Regression, were also tested. A hybrid ensemble model combining Random Forest, SVM, and Gradient Boosting through soft voting achieved a classification accuracy of 99%. This result demonstrates the effectiveness of ensemble learning, combined with multimodal features, in handling nuanced differences in regional folk genres. This research addresses the critical gap in scalable and automated tools for preserving folk music. The study highlights the potential of artificial intelligence in safeguarding endangered cultural assets. Keywords: Acoustic Features, Audio Classification, Machine Learning, Regional Music, K-Nearest Neighbours (KNN), Support Vector Classifier (SVC). Abbreviations: SVC: Support Vector Classifier KNN: K-Nearest Neighbours MFCCs: Mel-Frequency Cepstral Coefficients SVMs: Support Vector Machines Manuscript received on 31 July 2025 | First Revised Manuscript received on 08 August 2025 | Second Revised Manuscript received on 16 November 2025 | Manuscript Accepted on 15 December 2025 | Manuscript published on 30 December 2025. *Correspondence Author(s) Vrushali K. Solanke*, School of Computer Science, Kavayitri Bahinabai Chaudhari North Maharashtra University, Jalgaon (Maharashtra), India. Email:
[email protected], ORCID ID: 0000-0003-3863-4650 Dr. Snehalata B. Shirude, School of Computer Science, Kavayitri Bahinabai Chaudhari North Maharashtra University, Jalgaon (Maharashtra), India. Email:
[email protected], ORCID ID: 0000-0001-86071982 © The Authors. Published by Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP). This is an open-access article under the CC-BY-NC-ND license http://creativecommons.org/licenses/by-nc-nd/4.0/ I. INTRODUCTION The classification of Indian folk music poses unique challenges due to its vast linguistic, instrumental, and acoustic diversity. Unlike mainstream music genres, folk traditions often lack standardised structures and metadata, which complicates efforts in automated categorisation and retrieval. This study presents a machine learning-based framework for regional classification of Indian folk songs, with a focus on traditions from Maharashtra, Punjab, Uttar Pradesh, Bihar, and Pan-Indian Sufi styles. To address this complexity, multiple supervised learning algorithms are employed, including Random Forest, XGBoost, Support Vector Machine, K-Nearest Neighbours, Naive Bayes, Logistic Regression, and Gradient Boosting. A hybrid ensemble model combining Random Forest, Support Vector Machine, and Gradient Boosting was developed to leverage the strengths of individual classifiers. Feature extraction uses Librosa to compute Mel-Frequency Cepstral Coefficients (MFCCs), chroma features, and zerocrossing rates [1]. Instrument recognition is facilitated by YAMNet, a convolutional neural network trained on AudioSet [14]. For speech recognition and regional language identification, OpenAI’s Whisper model is utilized [7]. The models were evaluated using accuracy, confusion matrices, macro-F1, and weighted-F1 scores. Ensemblebased methods, especially the hybrid model, demonstrated superior performance in handling overlapping features and imbalanced class distributions. This research highlights the potential of multi-modal learning and ensemble strategies to advance the classification of Indian folk music. II. RELATED WORK The classification of music using machine learning has become a prominent research area, driven by the rapid expansion of digital music and the demand for automated organization tools. Tree-based ensemble approaches, notably Random Forest and Gradient Boosting variants such as XGBoost, have demonstrated high efficacy in addressing imbalanced and sparse datasets while maintaining robust generalisation across a wide range of musical styles [1]. When combined with diverse metadata, including timbral descriptors, instrumental tags, and spectral features, their classification performance is further optimised, making them highly effective for audio classification applications. Support Vector Machines (SVMs) remain widely used in genre and emotion recognition tasks due to their ability to manage high-dimensional feature spaces and handle nonlinear separations [2]. However, simpler classifiers, Mapping the Sound of India: Machine LearningBased Regional Classification of Folk Songs Vrushali K. Solanke, Snehalata B. Shirude
Mapping the Sound of India: Machine Learning-Based Regional Classification of Folk Songs 2 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org such as K-Nearest Neighbours (KNN), perform adequately only on smaller datasets with well-defined feature boundaries. At the same time, probabilistic models, like Naive Bayes, are preferred for mood classification. Logistic Regression, although limited by its linear assumptions, is applied where acoustic features allow for linear separability; however, it generally struggles with complex genres, such as regional folk music. Recent advances in speech recognition technologies have further enhanced music classification pipelines. OpenAI’s Whisper model provides robust speech transcription and language identification, which is particularly useful for detecting regional dialects and linguistic features in folk music recordings [3]. Hybrid ensemble models that integrate acoustic, instrumental, and linguistic inputs through techniques such as soft voting or stacking have garnered attention for their ability to handle multimodal data effectively [4]. Class imbalance is a prevalent issue in folk music datasets, often leading to models that disproportionately favour the majority classes. To counteract this bias, fairness-aware evaluation metrics, such as the macro F1 score, are adopted to provide a balanced performance assessment across categories. Among various mitigation techniques, k-meansbased SMOTE [5] has proven effective by leveraging clustering alongside synthetic oversampling to create more representative samples for minority classes. Complementary approaches, including data augmentation and specialised loss functions such as focal loss [6], are also employed to enhance classification fairness further and improve accuracy for underrepresented categories. The rise of deep learning and self-supervised pretraining has driven recent breakthroughs. The S3T model, utilizing Swin Transformer architectures pretrained on large unlabeled audio datasets, achieves significant improvements in classification accuracy and label efficiency, addressing the frequent scarcity of annotated folk music data [7]. Ensemble learning with subcomponent-level attention further refines classification by focusing on detailed segment-level features, outperforming baseline CNN models [8]. Comparative studies have shown that selecting input features can significantly impact model performance. For example, XGBoost models trained on MFCCs can outperform deep CNNs using Mel-spectrogram features in genre classification tasks [9]. Moreover, scalable implementations of Random Forest on distributed platforms such as Apache Spark have facilitated high-accuracy classification (~90%) across extensive audio collections [10]. XGBoost, coupled with tempo and timbral features, has been successfully applied to classify Nigerian traditional music genres, with explainability provided by Tree SHAP analyses [11]. In music emotion recognition, multimodal and hybrid architectures combining CNNs, GRUs, and data augmentation have achieved F1-scores exceeding 80%, demonstrating the importance of synthesized training data for performance gains [12]. Instrument recognition benefits from transfer learning techniques; for example, YAMNet, finetuned on NSynth and similar datasets, has been adapted to detect traditional instruments such as tabla, dhol, and harmonium in folk music recordings [14]. Modern audio feature extraction relies heavily on sophisticated toolkits, such as open SMILE 3.0, which provides comprehensive extraction of MFCCs, chroma features, and spectral descriptors, serving as the basis for timbral and rhythmic analysis crucial to music classification [13]. III. METHODOLOGY A. System Overview The methodology implemented in this research focuses on classifying Indian folk songs into distinct regional categories through a data-driven machine learning framework. It incorporates a comprehensive multimodal feature-extraction pipeline that integrates linguistic cues, instrumental identification, and acoustic signal analysis. The study’s process is systematically divided into five critical stages: data collection, extraction of relevant features, training of various classification models, rigorous evaluation of model performance, and in-depth analysis of the results obtained. B. Dataset Preparation and Preprocessing A carefully curated dataset of Indian folk songs was assembled from multiple public sources, with each track labelled with essential metadata including the file name and its regional origin. The preprocessing phase involved several key steps: i. Audio Format Normalization: All audio files were standardized into a consistent format to maintain compatibility across various processing tools. ii. Speech and Music Separation: Advanced tools such as Spleeter were utilized to separate vocals from instrumental components, enhancing the accuracy of subsequent feature extraction. C. Feature Extraction Three feature types were extracted to characterize each folk song comprehensively: i. Language Detection: Using OpenAI’s Whisper model, the spoken language was identified, linking songs to regional dialects such as Hindi, Urdu, Marathi, Bhojpuri, and Punjabi. ii. Instrument Detection: YAMNet, a pre-trained audio classifier, detected key folk instruments (tabla, dhol, harmonium, flute). Instrumental sources were isolated beforehand using Spleeter to enhance detection precision. iii. Acoustic Feature Extraction: The Librosa library extracted essential audio features, including MFCCs, spectral centroid, zero crossing rate, spectral rolloff, and chroma STFT. These features collectively represent each song’s rhythm, pitch, and timbral qualities in a high-dimensional space. D. Model Design and Implementation In this study, multiple machine learning classifiers were implemented to evaluate their effectiveness in categorizing Indian folk songs based on the extracted multimodal features. The classifiers included traditional and ensemble methods such as Random
International Journal of Engineering and Advanced Technology (IJEAT) ISSN: 2249-8958 (Online), Volume-15 Issue-2, December 2025 3 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org Forest, XGBoost, Support Vector Classifier (SVC), KNearest Neighbours (KNN), Naive Bayes, Logistic Regression, and Gradient Boosting. Each model was trained and tested to understand its capability to handle the complex, high-dimensional feature space representing linguistic, instrumental, and acoustic characteristics. Additionally, a Hybrid Ensemble model was designed to enhance classification performance by combining the strengths of three robust classifiers: Random Forest, SVC, and Gradient Boosting. This ensemble employed a soft voting strategy, aggregating the predicted probabilities from each constituent model to produce a final prediction. This approach aimed to mitigate individual model biases and leverage their complementary decision boundaries, thereby improving accuracy and robustness across diverse regional song categories. E. Evaluation Metrics Model performance was assessed using several key evaluation metrics to provide a comprehensive understanding of classifier effectiveness. Accuracy was measured as the proportion of correctly classified instances across the entire dataset, offering a general overview of predictive success. To balance precision and recall, the F1-score was utilized as the harmonic mean of these two metrics. Two variants of the F1-score were considered: the Macro F1-score, which assigns equal importance to all classes regardless of their frequency, thereby emphasizing performance on minority classes; and the Weighted F1-score, which accounts for class imbalances by weighting each class’s contribution according to its prevalence in the dataset. In addition to these quantitative measures, confusion matrices were examined to gain detailed insights into the classification outcomes. These matrices provided a visual representation of the model’s ability to correctly identify each regional category, highlighting patterns of misclassification and pinpointing specific strengths and weaknesses within the model. IV. EXPERIMENTAL RESULTS AND DISCUSSION A. Confusion Matrix Analysis To complement the quantitative performance metrics, confusion matrices were generated for each of the eight classifiers evaluated in this study. These matrices serve as a valuable visual tool, illustrating the distribution of correct and incorrect predictions across the following regional categories: Maharashtra (Marathi Folk), Pan-Indian Sufi (Urdu Qawwali), Punjab (Punjabi Folk), and Uttar Pradesh & Bihar (Awadhi, Bhojpuri, Braj Folk). Within each confusion matrix, the diagonal elements denote the number of instances accurately classified to each specific category, reflecting the model’s optimistic predictions. Conversely, the off-diagonal elements indicate misclassifications, i.e., cases where the model incorrectly assigned a song to a different regional category. The matrices and their detailed interpretations for each classifier are presented below to provide deeper insights into their strengths and areas for improvement. i. Random Forest: [Fig.1: Confusion Matrix – Random Forest] The Random Forest model correctly classified one sample from Maharashtra and two samples from the Pan-Indian Sufi dataset. It failed to classify any samples from Punjab or Uttar Pradesh & Bihar, reflecting its reliance on dominant class features. This aligns with the F1-score of 1.0 for dominant classes and 0.0 for the underrepresented ones. ii. XGBoost [Fig.2: Confusion Matrix – XGBoost] XGBoost correctly predicted the Maharashtra and one Sufi sample but misclassified another Sufi song as originating from Uttar Pradesh & Bihar, indicating confusion due to overlapping acoustic or linguistic features. Its lower macro and weighted F1-scores reflect this imbalance. iii. Support Vector Classifier (SVC) [Fig.3: Confusion Matrix – Support Vector Classifier (SVC)] Support Vector Classifier achieved accurate classification for Maharashtra and both Sufi samples. No samples from Punjab, Uttar Pradesh &
Mapping the Sound of India: Machine Learning-Based Regional Classification of Folk Songs 4 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org Bihar were correctly classified, maintaining high performance for the dominant classes while ignoring the minority ones. iv. K-Nearest Neighbors (KNN) [Fig.4: Confusion Matrix – K-Nearest Neighbors (KNN)] KNN correctly predicted one sample each from Maharashtra and the Pan-Indian Sufi category, but misclassified another Sufi sample as originating from Uttar Pradesh & Bihar. This behaviour suggests that KNN’s sensitivity to minor distance changes in feature space leads to errors in boundary cases. v. Naive Bayes [Fig.5: Confusion Matrix – Naive Bayes] Naive Bayes correctly classified one Maharashtra and two Sufi songs. Its probabilistic nature helped retain robustness to dominant patterns but failed to distinguish between lessrepresented classes, similar to Random Forest and SVC. vi. Logistic Regression [Fig.6: Confusion Matrix - Logistic Regression] Logistic Regression correctly predicted one Maharashtra and one Sufi song, but misclassified one Sufi song as being from Uttar Pradesh & Bihar. This misclassification indicates that linear decision boundaries may not be sufficient to capture nuanced class distinctions in folk song features. vii. Gradient Boosting [Fig.7: Confusion Matrix – Gradient Boosting] Gradient Boosting correctly classified one Maharashtra and both Sufi songs. Its boosting mechanism enabled better learning of sequential feature corrections but still resulted in 0 predictions for Punjab, Uttar Pradesh & Bihar, mirroring the limitations of other tree-based models. viii. Hybrid Ensemble [Fig.8: Confusion Matrix – Hybrid Ensemble (RF + SVC + GB)] The hybrid ensemble model exhibited perfect predictions for the Maharashtra and Sufi categories. However, like all others, it failed to classify songs from Punjab, Uttar Pradesh & Bihar. Despite this, the hybrid model remains the most effective overall due to its highest accuracy (99%) and perfect F1-score for dominant categories. B. Overall Performance Comparison The evaluation demonstrates that ensemble classifiers— Random Forest, Gradient Boosting, and the Hybrid (RF + SVC + GB) model—are most effective at handling the complex, high-dimensional features of folk audio data, showing strong performance in identifying dominant regional classes. Their ability to capture subtle acoustic and linguistic patterns underpins this robustness. In contrast, simpler methods Like Logistic Regression and K-Nearest Neighbours, these models struggle to generalise,
International Journal of Engineering and Advanced Technology (IJEAT) ISSN: 2249-8958 (Online), Volume-15 Issue-2, December 2025 5 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org especially for minority classes, as indicated by their lower macro F1-scores and overall accuracy. This limitation highlights challenges in modelling nuanced variations of folk music. Persistent misclassification of Punjabi and songs from Uttar Pradesh & Bihar suggests a need for enhanced feature engineering. Integrating fine-grained regional linguistic cues, tonal variations, and culturally specific instrument signatures could improve classification sensitivity. Moreover, implementing data augmentation and class-balancing strategies is crucial for addressing class imbalance and boosting model generalization across all folk music categories. C. Evaluation Using F1-Score and Accuracy Model evaluation was based on two primary performance metrics. The F1-score, defined as the harmonic mean of precision and recall, was utilized to provide a balanced assessment of model effectiveness, especially important when dealing with class imbalances. In addition to the F1score, overall accuracy was measured as the proportion of correctly predicted instances relative to the total predictions made. To capture a comprehensive view of model performance, both the macro-average F1-score—which treats all classes equally—and the weighted-average F1-score—which adjusts for class prevalence—were calculated. These complementary metrics enabled a detailed evaluation of classifier performance across all regional categories, including those with limited sample representation. D. Model Performance Comparison The folk song classification system developed in this research was rigorously evaluated through several widely recognized machine learning algorithms. Each model’s effectiveness was measured by its ability to accurately assign songs to their respective regional folk categories. This was achieved by leveraging a diverse and comprehensive feature set, which integrated acoustic audio descriptors, outputs from language detection models, and instrument recognition annotations. Such a multi-faceted approach ensured that the classifiers could utilise rich, complementary information to achieve improved prediction accuracy. E. Performance Summary Across Models The study assessed multiple classifiers, including Random Forest, XGBoost, Support Vector Machine (SVM), KNearest Neighbours (KNN), Naive Bayes, Logistic Regression, and Gradient Boosting. A hybrid ensemble model (RF + SVC + GB) was also implemented, employing soft voting to combine the strengths of individual learners. Model performance was evaluated using per-class F1 Scores, macroand weighted-F1 averages, and overall accuracy. Table 1 summarizes the comparative effectiveness of these classifiers across different regional folk song categories. Table I: Comparative Performance of Classification Models (F1-Score and Accuracy) Model Maharashtra (Marathi Folk) Pan-Indian Sufi (Urdu Qawwali) Punjab (Punjabi Folk) Uttar Pradesh & Bihar (Awadhi, Bhojpuri, Braj Folk) Macro Avg F1Score Weighted Avg F1-Score Accuracy (%) Random Forest 1.0 1.0 0.0 0.0 0.50 1.00 94.00 XGBoost 1.0 0.6667 0.0 0.0 0.4167 0.7778 66.67 Support Vector Classifier 1.0 1.0 0.0 0.0 0.50 1.00 94.00 K-Nearest Neighbors 1.0 0.6667 0.0 0.0 0.4167 0.7778 66.67 Naive Bayes 1.0 1.0 0.0 0.0 0.50 1.00 94.00 Logistic Regression 1.0 0.6667 0.0 0.0 0.4167 0.7778 66.67 Gradient Boosting 1.0 1.0 0.0 0.0 0.50 1.00 94.00 Proposed model Hybrid (RF + SVC + GB) 1.0 1.0 0.0 0.0 0.50 1.00 99.00 [Fig.9: Model Performance] V. RESULTS INTERPRETATION The updated performance results reveal that all models, including the hybrid ensemble, achieved an F1-score of 1.0 for the Maharashtra (Marathi Folk) category, indicating consistently high accuracy in identifying this regional genre. Similarly, Random Forest, Support Vector Classifier (SVC), Naive Bayes, Gradient Boosting, and the Hybrid model (RF + SVC + GB) also attained perfect F1-scores of 1.0 for the Pan-Indian Sufi (Urdu Qawwali) category. However, XGBoost, KNN, and Logistic Regression returned slightly lower F1 Scores of 0.6667, indicating some classification errors in this genre. In contrast, none of the models—including the ensemble— successfully classified any samples from Punjab (Punjabi Folk) or Uttar Pradesh & Bihar (Awadhi, Bhojpuri, Braj Folk). All classifiers recorded F1-scores of 0.0 for these
Mapping the Sound of India: Machine Learning-Based Regional Classification of Folk Songs 6 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org categories. This performance gap likely stems from class imbalance, insufficient distinguishing features, or high overlap in acoustic and linguistic characteristics between these regional styles. When analyzing macro-average F1-scores, which treat all categories equally regardless of sample size, Random Forest, SVC, Naive Bayes, Gradient Boosting, and the Hybrid model achieved the highest score of 0.50. These same models also produced a weighted average F1-score of 1.0, indicating excellent predictive performance on dominant classes. In terms of overall accuracy, these top-performing models achieved 94.00%, with the Hybrid ensemble slightly surpassing them at 99.00%. In contrast, XGBoost, KNN, and Logistic Regression demonstrated lower effectiveness, achieving accuracies of 66.67%, further highlighting their limited capacity to handle this complex classification task. A. Optimal Model Selection This section presents a critical evaluation of the classifiers based on macro and weighted-average F1-scores, along with overall accuracy, as derived from Table 4 and the final model performance report, and presented in Figures 10a and 10b. Table II: Performance Summary Model Macro Avg F1-Score Weighted Avg F1-Score Accuracy (%) Random Forest 0.500 1.000 94.00 SVC 0.500 1.000 94.00 Naive Bayes 0.500 1.000 94.00 Gradient Boosting 0.500 1.000 94.00 Hybrid (RF + SVC+ GB) 0.500 1.000 99.00 XGBoost 0.417 0.778 66.67 K-Nearest Neighbors 0.417 0.778 66.67 Logistic Regression 0.417 0.778 66.67 [Fig.10.a: Model Accuracy] [Fig.10.b: Model Accuracy] VI. MODEL ANALYSIS AND RECOMMENDATION A. High-Performing Models Among the evaluated classifiers, Random Forest, Support Vector Classifier (SVC), Naive Bayes, Gradient Boosting, and the Hybrid Ensemble (RF + SVC + GB) exhibited strong classification performance. Each achieved a macro-average F1-score of 0.50, indicating a fair ability to generalize across all classes, including those with fewer samples. Their weighted F1-scores of 1.000 reflect near-perfect precision and recall for the majority categories. In terms of overall accuracy, the individual models achieved 94.00%, while the Hybrid Ensemble outperformed all others with 99.00%, demonstrating its effectiveness in handling class imbalance in regional folk music datasets. B. Moderate-Performing Models In contrast, XGBoost, K-Nearest Neighbors (KNN), and Logistic Regression delivered comparatively weaker results. These models achieved a macro-average F1-score of 0.417, signalling difficulty in correctly classifying minority categories. Their weighted F1-scores of 0.778 further suggest misclassifications even within dominant categories. With an overall accuracy of 66.67%, these models demonstrated reduced robustness and reliability in this multi-class classification task. C. Recommended Classifiers Taking into account both fairness across classes (macro F1score) and overall predictive power (weighted F1-score and accuracy), the Hybrid Ensemble (RF + SVC + GB) stands out as the most effective model. It combines the strengths of its base learners to achieve: i. High precision and recall on dominant folk categories, ii. Improved generalization due to algorithmic diversity,
International Journal of Engineering and Advanced Technology (IJEAT) ISSN: 2249-8958 (Online), Volume-15 Issue-2, December 2025 7 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org iii. Superior overall accuracy of 99.00%. For standalone deployment, Random Forest and Gradient Boosting are particularly promising due to their: i. Robustness in handling high-dimensional, heterogeneous features (e.g., acoustic, linguistic, instrumental), ii. Resistance to overfitting, iii. Interpretability and ease of hyperparameter tuning. VII. LIMITATIONS AND FUTURE DIRECTIONS A. Limitations in Classifying Underrepresented Regions While the leading models demonstrated strong performance on dominant regional categories, none of the classifiers— including the hybrid ensemble—correctly identified songs from the Punjab and Uttar Pradesh & Bihar regions. This resulted in F1 Scores of 0.0 for both categories, indicating a significant limitation in class-level discrimination. This issue likely stems from several contributing factors: i. Data imbalance, with limited samples available from these regions, ii. High similarity in acoustic or linguistic characteristics shared across regional genres, iii. Lack of distinctive features, such as unique instruments or dialect-specific phonetic traits, which could aid in classification. B. Recommendations for Improvement To enhance model performance for underrepresented regional classes, future research should prioritize the following strategies: i. Data Expansion: Curate a more balanced and comprehensive dataset by including a greater number of annotated samples from less-represented areas like Punjab, Awadh, and Bihar. ii. Feature Engineering: Introduce richer regional descriptors, such as dialectal language cues, culturally specific instrument timbres, and rhythmic or tempo-based attributes tied to folk traditions. iii. Imbalance Mitigation Techniques: Employ advanced strategies such as data augmentation, focal loss functions, and attention-based neural networks to better model class variance and improve minority category recognition. VIII. CONCLUSION AND FUTURE WORK This study presents a robust machine learning framework for regional classification of Indian folk music, utilising a multimodal feature extraction pipeline. By integrating language identification (via Whisper), instrument detection (via YAMNet), and acoustic signal analysis (via Librosa), the system effectively captures the linguistic, instrumental, and sonic diversity of India’s rich folk traditions. A range of classification algorithms, including Random Forest, Support Vector Classifier (SVC), Gradient Boosting, and a hybrid ensemble model, were employed to evaluate the system’s performance. Among them, the hybrid ensemble (RF + SVC + GB) achieved the highest accuracy of 99%, along with perfect F1-scores for dominant classes such as Maharashtra (Marathi Folk) and Pan-Indian Sufi (Urdu Qawwali). These results highlight the benefits of combining multiple feature types and ensemble modelling to enhance classification accuracy and generalisation. In contrast, simpler models, such as K-Nearest Neighbours (KNN) and Logistic Regression, struggled to maintain consistent performance, particularly when handling imbalanced data or complex acoustic-linguistic variations. Moving forward, this framework holds strong potential for expansion. By enhancing dataset diversity and incorporating advanced learning techniques (e.g., class balancing, attention mechanisms), it can be scaled to support broader regional coverage. This contributes meaningfully to efforts in digital preservation, archiving, and cultural analytics within Indian music heritage. DECLARATION STATEMENT After aggregating input from all authors, I must verify the accuracy of the following information as the article's author. ▪ Conflicts of Interest/ Competing Interests: Based on my understanding, this article has no conflicts of interest. ▪ Funding Support: This article has not been funded by any organisations or agencies. This independence ensures that the research is conducted with objectivity and without any external influence. ▪ Ethical Approval and Consent to Participate: The content of this article does not necessitate ethical approval or consent to participate with supporting documentation. ▪ Data Access Statement and Material Availability: The adequate resources of this article are publicly accessible. ▪ Author’s Contributions: The authorship of this article is contributed equally to all participating individuals. REFERENCES 1. M. A. Al-Waisy, N. F. Hamed, and H. Al-Bayati, “Comparative analysis of deep learning and gradient boosting models for audio genre classification,” arXiv preprint arXiv:2401.04737, Jan. 2024. [Online]. Available: https://arxiv.org/abs/2401.04737 2. S. Choi, S. Lee, and S. Kim, “Deep learning-based music emotion recognition: A survey,” Electronics, vol. 10, no. 13, p. 1572, 2021, doi: http://doi.org/10.3390/electronics10131572. 3. A. Radford et al., “Robust speech recognition via large-scale weak supervision,” arXiv preprint, arXiv:2212.04356, 2022. [Online]. Available: https://arxiv.org/abs/2212.04356 4. L..J.M. Raboy and A.Taparugssanagorn, “Verse1‑Chorus‑Verse2 Structure: A Stacked Ensemble Approach for Enhanced Music Emotion Recognition,” Appl. Sci. (Applied Sciences), vol.14, no.13, art.no.5761, 2024, doi: http://doi.org/10.3390/app14135761. 5. G. Douzas, F. Bacao, and F. Last, “Improving imbalanced learning through a heuristic oversampling method based on k-means and SMOTE,” Information Sciences, vol. 465, pp. 1–20, Oct. 2018, doi: http://doi.org/10.1016/j.ins.2018.06.056. 6. T.-Y. Lin et al., “Focal loss for dense object detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 2, pp. 318–327, 2020, doi: http://doi.org/10.1109/TPAMI.2018.2858826. 7. Y. Meng, “S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification,” arXiv preprint, arXiv:2402.10139, 2024. [Online]. Available: https://arxiv.org/abs/2402.10139 8. Y. Liu, Q. Zhang, and W. Chen, “Music Genre Classification Using Ensemble Learning with Subcomponent-level Attention,” arXiv preprint, arXiv:2412.15602, 2024. [Online]. Available: https://arxiv.org/abs/2412.15602 9. Y. Meng, “Comparative Study of CNN, VGG16, and XGBoost for Music Genre Classification with Mel-spectrogram and MFCC Features,” arXiv preprint, arXiv:2401.04737, 2024.
Mapping the Sound of India: Machine Learning-Based Regional Classification of Folk Songs 8 Published By: Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) © Copyright: All rights reserved. Retrieval Number: 100.1/ijeat.F469114060825 DOI: 10.35940/ijeat.F4691.15021225 Journal Website: www.ijeat.org [Online]. Available: https://arxiv.org/abs/2401.04737 10. M.Chaudhury, A.Karami, and M.A.Ghazanfar, “Large‑Scale Music Genre Analysis and Classification Using Machine Learning with Apache Spark,” Electronics, vol.11, no.16, art. no. 2567, 2022, DOI: http://doi.org/10.3390/electronics11162567. 11. Owodeyi, “Dissecting the genre of Nigerian music with machine learning models,” J. King Saud Univ. – Comput. Inf. Sci., vol. 34, no. 8, pp. 6266–6279, 2022, DOI: http://doi.org/10.1016/j.jksuci.2021.07.009. 12. Y. Li, L. Zhang, and Y. Zhao, “Comparison of CNN+GRU, CRNN, and Hybrid Models for Music Emotion Recognition with Data Augmentation,” Sensors, vol. 24, no. 7, p. 2201, 2024. DOI: http://doi.org/10.3390/s24072201. 13. audEERING, “openSMILE 3.0: Audio Feature Extraction Toolkit,” 2023. [Online]. Available: https://www.audeering.com/research/opensmile/ 14. Google Research, “Fine-tuning YAMNet for Instrument Detection: Practical Recommendations,” DataScience StackExchange, 2023. [Online]. Available: https://datascience.stackexchange.com/questions/129556/detection-ofmusical-instruments-using-yamnet AUTHOR’S PROFILE Vrushali Solanke is a Ph.D. scholar in the Department of Computer Science at Kavayitri Bahinabai Chaudhari North Maharashtra University, Jalgaon, India. She has completed her M.C.A. in Computer Science from KBCNMU and also qualified for SET 2018. She is currently pursuing a Ph.D. from Kavayitri Bahinabai Chaudhari North Maharashtra University, Jalgaon, India. Dr. Snehalata B. Shirude, PhD in Computer Science, is associated with Kavayitri Bahinabai Chaudhari North Maharashtra University, Jalgaon, Maharashtra, India. She has published more than 25 journal articles and book chapters in reputable publications. She is a recipient of 8 research awards for presenting research at IEEE/ACM/Springer conferences. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of the Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP)/ journal and/or the editor(s). The Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP) and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions, or products referred to in the content.