Full text
Bridging Paintings and Music – Exploring Emotion-based Music Generation Through Paintings Tanisha Hisariya, Huan Zhang, and Jinhua Liang Queen Mary University of London, London, United Kingdom [email protected] {huan.zhang, jinhua.liang}@qmul.ac.uk Abstract. Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that resonates with the emotions depicted in visual arts, integrating emotion labeling, image captioning, and language models to transform visual inputs into musical compositions. Addressing the scarcity of aligned art and music data, we curated the Emotion Painting Music Dataset, pairing paintings with corresponding music for e!ective training and evaluation. Our dual-stage framework converts images to text descriptions of emotional content and then transforms these descriptions into music, facilitating e"cient learning with minimal data. Performance is evaluated using metrics such as Fréchet Audio Distance (FAD), Total Harmonic Distortion (THD), Inception Score (IS), and KL divergence, with audio-emotion text similarity confirmed by the pre-trained CLAP model to demonstrate high alignment between generated music and text. This synthesis tool bridges visual art and music, enhancing accessibility for the visually impaired and opening avenues in educational and therapeutic applications by providing enriched multisensory experiences. Keywords: Music generation ·Generative AI ·Transformers ·Images · Emotions 1Introduction “Art is not what you see but what you make others see.” - Edgar Degas. Visual art communicates information and emotions from the artist to the observer, encapsulating cultural and innovative influences from di!erent eras. Similarly, music as an art form evokes a broad spectrum of emotions through its composition and style, paralleling the expressive power of visual arts. This paper explores the innovative intersection of these two art forms by generating music that reflects the emotions perceived in visual artworks such as paintings. This approach not only aims to make art more accessible to the visually impaired by translating visual cues into auditory signals but also extends the research frontier in AI-driven generative models conditioned on images. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 686
2T.Hisariyaetal. Advancements in generative models, have shown remarkable achievements in mimicking human creativity and generating content that aligns with user expectations across various media, leveraging deep learning techniques [3, 25,21]. The models capture complex patterns from large datasets and generate outputs that often surpass traditional methods. Music generation is one significant application of AI, utilizing methods to create compositions previously thought unattainable by machines, with techniques categorized into symbolic representation and waveform generation. Symbolic generation focuses on note sequences and events, mainly used by musicians [37,42], while waveform generation produces continuous audio signals interpretable by a general audience, o!ering a broader application for everyday interactions [23,41]. Waveform generation’s e!ectiveness heavily relies on the sampling rate to fully capture the audio structure, requiring models that can train e!ectively on high-dimensional datasets [6]. The challenge of cross-modal generative AI, particularly converting images to music, lies in identifying relationships between these diverse modalities and the scarcity of paired data necessary for training. Despite these challenges, using transfer learning to apply a pre-trained music generation model fine-tuned to specific needs reduces the reliance on large datasets, allowing for the integration of di!erent architectures and modalities to create e"cient and seamless hybrid models. This paper aims to bridge the gap between visual and auditory arts, enhancing accessibility and merging distinct forms of artistic expression. The contribution of this work is summarized as follows: 1. We proposed a visual-guided music synthesis system that generates music by interpreting the emotions conveyed by images. 2. Our framework is decomposed into image-to-text and text-to-music tasks, facilitating e"cient learning with minimal data. We further enhance training e"ciency by exclusively training the decoder in the latent space. 3. We explore the influence of text descriptions by employing diverse textual conditions. To this end, we have curated for both training and evaluation purposes the Emotion Painting Music Dataset, which is publicly available at this URL 1. Our generated music is qualitatively measured across various metrics using the Fréchet Audio Distance (FAD) [15], Total Harmonic Distortion (THD), Inception Score (IS), and KL divergence. Audio-emotion text similarity has also been measured by pre-trained CLAP [10] model to demonstrate high alignment between generated music and text. This tool, bridging art and music, holds promise for enhancing learning experiences in educational environments or therapeutic contexts, providing unique multi-sensory engagements. 2RelatedWork This research covers three key areas in generative AI: image feature extraction and conversion, deep learning techniques for music generation, and multi-modality in audio and music generation. 1https://zenodo.org/records/13717256 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 687
Bridging Paintings and Music 3 2.1 Image Feature Extraction and Conversion Feature extraction from images can utilize either unimodal or multimodal approaches. The unimodal approach o!ers straightforward single-label classifications, while multimodal strategies employ language models like LLMs to process visual representations [17,18]. Deep neural networks (DNN)s, such as [11], laid the groundwork for the unimodal strategy by categorizing an image to a fixed set of labels. Subsequently, DNNs have been proven to yield the promising performance in other fields, such audio and graphs [20,8, 43]. The evolution of image captioning leveraged these CNN architectures alongside natural language processing, significantly advanced by the introduction of contrastive learning [31,18].Contrastive Language-Image Pretraining (CLIP) enhanced the alignment between textual and visual data using dual encoders to calculate cosine similarities between text and image vectors. Following CLIP, BLIP [17] further refined the model by integrating a flexible multimodal encoderdecoder framework that excelled in both understanding and generating visuallanguage tasks with high accuracy. The subsequent introduction of Multimodal large language model [1, 19,24] marked a major advancement, optimizing the handling of more complex contextual and visual information, setting new standards for image captioning capabilities. 2.2 Deep Learning Methods for Music Generation The use of Recurrent Neural Networks (RNNs) to generate musical melodies began with Todd [35], but their limited memory capacity led to the development of Long Short-Term Memory (LSTM) units [13], which improved the ability to retain musical sequences over time. The evolution continued with models incorporating Restricted Boltzmann Machines (RBM) for better melody generation [12], although challenges in long-term memory persisted until Google introduced advancements in RNNs for music with enhanced dependency handling. The introduction of generative architectures such as Convolutional Neural Networks (CNNs), Variational Auto Encoders (VAEs), and Generative Adversarial Networks (GANs) marked significant progress. MusicVAE [32] and MidiNet [39], a CNN-based GAN, expanded capabilities for creating coherent musical sequences, although GANs sometimes su!ered from mode collapse. MuseGAN [9] further advanced multi-track music generation, improving the coherence of generated compositions. The advent of WaveNet by Google in 2016 introduced a CNNbased model that employed an autoregressive approach for dynamic raw audio generation, heavily used in both text-to-speech and music generation tasks [36]. Significant advances were also seen with the introduction of Music Transformer [14], which utilized transformers and relative attention models to enhance long-term music generation, surpassing earlier RNN models in e"ciency but encountering issues with note redundancy. MuseNet in 2019 [28] and later developments like seqGAN [40] and models by Jacek and Teodara using Conditional VAEs and RNNs for emotion-driven music generation [27] highlighted the ongoing challenges of computational demands and complex architecture interpretability while pushing the boundaries of music generation technology. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 688
4T.Hisariyaetal. 2.3 Multi-Modality Audio and Music Generation In 2022, cMelGAN [30], a conditional GAN model based on Mel Spectrogram, was introduced to improve music generation e"ciency, though it faced challenges with GAN training instability. JukePix [38], a model for converting paintings to music using Convolutional GANs, showed potential in generating multi-track music but was limited by solo performance evaluations. Chen et al. developed MusicLDM [4], a complex model integrating CLAP, VAE, Hifi-GAN, and di!usion models, excelling in music generation but constrained by data sampling rates and computational resources. Further advancements include AudioLDM [23], which uses CLAP embeddings and a latent di!usion model to achieve high-quality audio generation. Building on this, AudioLDM 2 [22] leveraged GPT-2 to handle various input modalities, showing significant improvements in accuracy. Mousai [33] and MusicLM [2] further explored text-conditioned music generation, with MusicLM providing consistent output despite challenges in processing complex text structures. MeLoDy by Lam et al. [16], used a dual-path di!usion model combined with a language model to enhance semantic modeling and music generation. Despite its innovative approach, training data biases limited its diversity. She!er and Adi’s im2wav model [34], based on a transformer architecture and CLIP, aimed to generate highfidelity audio from image inputs but faced computational ine"ciencies. Recently, Chowdhury et al. introduced MelFusion [5], synthesizing music from images and text via advanced deep-learning di!usion models, setting new benchmarks in performance and enables further research in cross-modal generative AI. 3Method The method of generating music from images based on emotions consists of integrating deep learning models. Firstly, we will convert the images into their textual format, and then text along with its associated music will be used to fine-tune the MusicGen [7] model to generate the required music. By doing this, we are not only generating the music conditioned on emotions from images, but we are also exploring the e!ect of various textual descriptions during musical generation. In this research paper, we are exploring four models: an emotion labeling model, an image description model, a large language model, and a Music Generation model. The overview of our methodology can be seen in Figure 2. 3.1 Image Emotion Labelling Model To e!ectively perceive the emotions of images, we have introduced a classification model to label the emotions. The model will play an important part in determining the emotions during inference, helping to improve the consistency and relevance of generated music. Pre-trained ImageNet ResNet50 [11] has been chosen as the best model for this approach due to its ability to manage diverse and complex datasets with dense layers. Leveraging the technique of transfer learning, the model has been adapted to a given dataset with the additional two GRU layers along with Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 689
Bridging Paintings and Music 5 a multi-head Attention layer before the fully connected layer. The output layer of ResNet50 architecture has been flattening out to meet the output classes of the dataset. Further, to enhance the performance and to prevent overfitting, some dropping out of non-essential neurons have been incorporated before fully connected layers. Fig. 1. Architecture of Emotion Labelling model with pretrained ResNet50 and additional two bidirectional GRU and one Attention layer proceeding with dropout along fully connected layer. 3.2 Image Description Model The Image Captioning Model is very crucial as it is responsible for generating the captions of images reflecting emotions perceived by them. By doing this, we aimed to enhance the description of images by extracting more word tokens. We employed BLIP [11], a current state-of-the-art model for Image Captioning, due to its superior performance in generating diverse and descriptive texts aligning closely with visual information. The model is being conditioned on the emotion labels obtained from the emotion classifier enabling it to give better relevant emotional descriptions. The model is trained on a large amount of highly diversified data so we can directly incorporate the BLIP Large Captioning model to generate the caption. 3.3 LLM Model This model plays an integral role in the evolution of visual and musical information. It further enhances the description generated by the captioning model by incorporating some musical terms that reflect the mood, themes, and musical understanding terms that are very useful for generating the music. There is a need to enhance the description because the Image Captioning Model gives us descriptions based on image features such as objects, colors, and more, while the models for generating music need some music related component details in that to optimally perform. Providing this type of description leads to a better quality and resemblance of music. That’s why we are integrating the LLM model into our framework to fill the gap between the provided description and the expected inputs. The model not only works with the optimization of description but also ensures that the information about the visual image, especially the essence of emotion, is not lost. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 690
6T.Hisariyaetal. Fig. 2. Overview architecture of our model , which encompasses working flow of all the four models represented as di!erent colour giving the output as MS-G E (single label text from ResNet50+ MusicGen), MS-G N (image descritive text from BLIP+MusicGen), MS-G L (Enhanced description from Falcon+ MusicGen), MS-G O (Enhanced description from Falcon+enhanced finetuning method of MusicGen). The LLM model we have incorporated here is FalconRW-1B [29] because it is fully open-source, and able to perform well even with restricted computational power. It consists of language modelling decoder-based architecture that is only incorporated with many advanced techniques like Attention. The input requirement of this model is characterized into three parts as shown in Table 1: system message (intent to tell the behavior of the model), instructions (intent to give the proper input), and response (the response model is providing). Table 1. Description of input format given to falcon 1B model, emphasizing the system message and instruction, followed by a description provided. Role Content System message You are an enhanced description generator. You will be given with image description and you have to enhance those in musical terms. Instruction Generate a musical theme description for the following image description: “sad man in a sailor’s hat sitting at a table”. Include details like mood, genre, tempo, and melody in 2 lines. 3.4 Music Generation The MusicGen-small model, abbreviated as MG-S, was fine-tuned for generating music based on various textual inputs derived from corresponding image-totext models and audio files. Text and audio files were encoded into tokens using specialized encoder models, followed by a conditioning and fusion process involving an Attention mechanism. This was further processed by a transformer Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 691
Bridging Paintings and Music 7 model using masking techniques to generate relevant tokens for loss computation and parameter refinement. Our experimental process involved several iterations of the MG-S model to enhance music generation capabilities: MG-S Emotive: Utilized emotion descriptions from the Image Emotion Labeling model to adjust the model’s parameters. MG-S Narrative: Conditioned the music generation on image descriptions enriched with emotional cues, with comprehensive parameter tuning in the model. MG-S Lyrical:EnhancedthedescriptivecontentofimagesusingaLanguage Model that incorporates musical knowledge, thoroughly fine-tuning the model. MG-S Optimized: Introduced architectural improvements to optimize model performance. These improvements included pre-processing music and text files prior to training to stabilize and streamline the training process. The approach involved storing precomputed tensors, freezing initial layers to enhance stability, and modifying input handling for consistent learning. This version also adjusted how input descriptions were managed to ensure accurate performance evaluation on our dataset. These versions of MG-S represent a progression in the methodical enhancement of music generation, conditioned on textual and emotional cues, leading to a refined and e"cient training methodology. 4 Experiments 4.1 Datasets Data collection and preparation are elementary steps in developing any model. Due to the lack of an existing paired dataset encompassing both art and music, which share the same emotional attribute, we proposed to make our own bespoke paired dataset integrating two di!erent art forms, painting and music, while conveying the same emotion. For the image paintings dataset, we utilized WIKIART EMOTION DATASET [26], an openly available dataset of wiki art depicting various emotions. Wikiart is a collection of various paintings that evolved from di!erent eras of the past to present, symbolizing di!erent art forms and meanings. These paintings have been analyzed and further categorized into more than ten emotions in the emotion dataset. Based on these, we have manually analyzed datasets for five particular emotions - Happy, Angry, Sad, Fun, and Neutral and collected 1200 such paintings, which will serve as one part of our paired dataset. The manual selection process depicts the accurate representation of each dataset conveying the emotion. The next part contains the collection of music datasets that should convey the same emotion. For this purpose, we selected MIREX EMOTION DATASET, a dataset of 193 MIDI files of music depicting various emotions. Initially, we preprocess the musical dataset from its raw form of MIDI and convert it into .wav form at 32KHz making them compatible frequency for the MusicGen model. We further combined the emotional aspects of several parts to categorize it into five similar emotions as with paintings finally. Furthermore, as we are using a Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 692
8T.Hisariyaetal. fixed 30s audio segment chunk in our music generation model, the music has been trimmed into various 30s without any overlapping between two di!erent audios. By doing these preprocessing steps, we can generate more music samples that are uniquely identified from each other. Later, to form the final paired dataset, the music and paintings have been associated with each other randomly, depicting the same emotions. In this way, we can have 1200 di!erent pairs of paintings and music depicting five di!erent emotions. Furthermore, for the training purpose of the model, we took around 80% of data from each section of emotion, with the remaining 20% split evenly between evaluation and testing sets. 4.2 Evaluation Metrics To evaluate the quality and resemblance of generated music, we used a set of objective evaluation metrics to measure quality, smoothness, noise, and distortion in the generated music: Frechet Audio Distance (FAD). FAD, inspired by the Frechet Inception Distance, measures the similarity between the statistical distributions of generated and reference music sets using the VGGish model for feature extraction. A lower FAD score indicates greater similarity to the reference set. Contrastive Language Audio Pretraining (CLAP). CLAP score calculates the similarity between text descriptions and audio using embeddings from the pre-trained LAION CLAP model’s text and audio encoders, computing cosine similarity between them. A higher CLAP score indicates better alignment between text and audio. Total Harmonic Distortion (THD) score. THD measures the harmonic distortion present in the generated music, assessing the purity of the audio signal. Lower THD scores signify less distortion and higher audio quality. Inception score (ISc). IS reflects the variety and diversity of the generated audio group. A higher Isc indicates a diverser distribution of synthesise music. Kullback-Leibler (KL) divergence.KLdivergencequantifiesthedi!erence between the probability distributions of reference and generated features. A lower KL score indicates closer resemblance between two distributions, suggesting better generation fidelity. 4.3 Training All our experiments were performed on one NVIDIA A40 GPU with 48 GB of memory. This setup allowed us to perform training with a batch size of 16, using the small version of the MusicGen model over 40 epochs. We used an AdamW optimizer with early stopping and a learning rate of 1e-5, a cosine scheduler with warmup steps of 100. During inference, we took the topk as 250, selecting the top 250 most resembling audio tokens at a temperature of 1. 4.4 Results The architectural experiments with the four variants of the MusicGen model, as reflected in Table 2, o!er detailed insights into each model’s performance Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 693
Bridging Paintings and Music 9 Table 2. Objective comparison of test set data for emotion based image to music generation across all the models. Here the best results are made bold. Model FAD→CLAP↑KL→THD→ISc↑ MG-S Emotive 7.02 0.075 0.054 1.79 1.044 MG-S Narrative 5.22 0.096 0.045 1.73 1.032 MG-S Lyrical 5.06 0.11 0.046 1.92 1.031 MG-S Optimized 5.54 0.13 0.012 1.75 1.033 improvements and challenges. MG-S Emotive, our baseline model, uses ResNet50 and GRU to extract single-word emotion labels, demonstrating limitations with high FAD and KL scores, and low CLAP scores indicating poor text-to-music alignment and outputs with significant noise. MG-S Narrative advances this by using the BLIP model for richer emotional captioning, improving text-music alignment as seen in the higher CLAP scores and reduced noise, though it struggles with complex emotions like anger. MG-S Lyrical incorporates an LLM to enrich musical context in text descriptions, enhancing semantic appropriateness and improving FAD and CLAP scores but facing challenges in balancing complexity with fidelity as indicated by the KL and THD scores. Finally, MG-S Optimized integrates enriched contextual descriptions with a modified tuning pipeline, significantly reducing training times and achieving the highest CLAP scores while minimizing distortion and noise, demonstrating the most e!ective architecture in complex emotional contexts as evidenced by the spectrogram analysis. These progressive refinements highlight the enhanced capabilities of MusicGen in generating high-fidelity music aligned with complex textual descriptions. Fig. 3. CLAP analysis of generated song across model with their emotions. It is being meausred with providing “emotion song” as text during CLAP calculation. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 694