scieee AI-readable full text Open interactive document viewer

AI-Assisted Music Production: A User Study on Text-to-Music Models

Ronchini, Francesca; Comanducci, Luca; Marcucci, Simone; Antonacci, Fabio

Abstract

Text-to-music models have revolutionized the creative landscape, offering new possibilities for music creation. Yet their integration into musicians’ workflows remains underexplored. This paper presents a case study on how TTM models impact music production, based on a user study of their effect on producers’ creative workflows. Participants produce tracks using a custom tool combining TTM and source separation models. Semi-structured interviews and thematic analysis reveal key challenges, opportunities, and ethical considerations. The findings offer insights into the transformative potential of TTMs in music production, as well as challenges in their real-world integration.

Full text

AI-Assisted Music Production: A User Study on Text-to-Music Models Francesca Ronchini[0000→0001→6897→1645], Luca Comanducci[0000→0002→4167→5173], Simone Marcucci, and Fabio Antonacci[0000→0003→4545→0315] Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB), Politecnico di Milano Piazza Leonardo Da Vinci 32, 20133 Milan, Italy {name.surname}@polimi.it Abstract. Text-to-music models have revolutionized the creative landscape, o!ering new possibilities for music creation. Yet their integration into musicians’ workflows remains underexplored. This paper presents a case study on how TTM models impact music production, based on a user study of their e!ect on producers’ creative workflows. Participants produce tracks using a custom tool combining TTM and source separation models. Semi-structured interviews and thematic analysis reveal key challenges, opportunities, and ethical considerations. The findings o!er insights into the transformative potential of TTMs in music production, as well as challenges in their real-world integration. Keywords: Text-to-Music ·Generative Models ·Human-AI Interaction ·Human-computer co-creativity 1Introduction Deep learning is rapidly advancing music generation, particularly through Textto-Music (TTM) models that generate audio from text prompts [47,31]. These models are transforming music composition and production [16,22], enabling creative experimentation even for those without formal training, as seen with Suno [35] and Udio [38]. Despite advancements in AI music generation, TTM models still struggle to interpret musicians’ controls [44,30], underscoring the need for focused research on creator interaction and collaboration. Ronchini et al. [30] conducted a user experience study exploring how musicians interact with a TTM model enhanced by fine-tuning techniques to adapt to users’ specific preferences. However, the study was constrained by a homogeneous participant pool and a broad focus. Parallel to this paper, Fu et al. [12] conducted a similar investigation focused on music production, using a limited sample of university students with little to no production experience, where pressure due to deadlines and assignments may have influenced the results. All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 723 F. Ronchini et al. Fig.1: Workflow of the experimental procedure. The red lines mean optional actions, while the green lines mean the reiteration process. This paper overcomes previous limitations, proposing a study to better investigate the interpretation gap [44] by conducting a user case study on AI-assisted music production. The research involves 17 music producers with varying levels of experience, and evaluate how TTM models influence productivity and e!ectiveness, the usability of AI tools as creative collaborators, and the overall perspective on incorporating TTM models into creative workflow. Participants were invited to use their preferred setups to produce a musical track with the help of a custom-developed tool that integrates a TTM model and an AI-based source separation model, providing producers with greater control and flexibility over the generated composition [33]. Fig. 1 illustrates the workflow of the experiment. We believe these results o!er insightful perspectives on user experiences and ethical implications of TTM models. Supplementary material, including videos of interaction and code for replication, are available online. 2Background Text-to-music Models.Severaltext-basedgenerativemodels,basedondifferent approaches, have been proposed for the raw audio and music generation task [47]. The first model proposed is AudioLM [2], which was further developed in MusicLM [1]. Subsequently, transformer-based models have been proposed [17,7,49], and di!usion-based models have gained attention [43,19,10]. Recently, commercial solutions like Suno AI [35] and Udio [38] have been introduced. However, we opted not to consider them in order to adhere to the principles of open science, enabling fellow researchers to reproduce our experiments. Moreover, the use of closed-source models would have prevented us from integrating the source separation technique into a single interface. In recent research, control inputs such as tempo, rhythm, chords, and melody have been Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 724 AI-Assisted Music Production: A User Study on Text-to-Music Models integrated alongside text [41,18,37]. As text remains the primary control in these models, this study focuses on text-based approaches. For a more detailed survey, the reader is referred to [47,22]. Music Source Separation Models.Data-drivenmodelsforMusicSource Separation are categorized into spectrogram-based [36] and waveform-based models [9]. In [8], the authors propose hybrid approaches that combine both temporal and spectral domains. Spleeter [14] leverages pre-trained models. Recently, Hybrid Transformer Demucs (HT-Demucs) [32], inspired by Hybrid Demucs [8], and Banquet [39] have been introduced. Most models separate audio into four stems: vocals, drums, bass, and other (VDBO). Spleeter adds piano for five stems, HTDemucs includes guitar for six, and Banquet generates multiple stems [39]. Co-creativity and Human-AI interaction.SeveralstudiesexploreHumanAI interaction in music generative models [24,20], the perception of AI-generated music [4,34], and opportunities and challenges of AI for music [4,26]. In [21], the authors evaluate generative models and AI-steering interfaces in downstream creative tasks. Zhou et al. [48] explore ways to enhance user interaction by allowing direct control over a model’s sampling behavior, while Huang et al. [26] study human-AI co-creation, identifying challenges musicians face when composing with AI. Several interfaces have been proposed to improve interaction between generative models and users [3,20,29,46,48,42], AI-assisted music composition [28], along with creative supporting tools designed to enhance collaborative creativity [5,11,16]. However, research of TTM integration into music creation workflows remains limited, but crucial [30,12,44]. 3Studydesignandmethod Models and framework selection.WechoseMusicGen[7]formusicgeneration and HT Demucs [32] for source separation due to their performance tradeo!s and easy integration into a unified framework, ensuring reproducibility. MusicGen is a textor melody-conditioned autoregressive model available in three sizes, while HT Demucs [32] uses dual U-Nets and a Transformer to separate instruments. We use the htdemucs_6s model, which separates VDBO, guitar and piano. Combining music separation with generation enhances composition allows independent manipulation of individual sources. MusicGen-Stem [33], which integrates both functions, has been recently released. However, it was released during our study and supports only three stems. While recent studies provide additional controls for TTM models [41,37], text instructions continue to be the dominant approach. Given this and the study’s limited duration, we decided to focus the study on models that use text as input. Hence, we do not use the melody conditioning of MusicGen. Interface design.WedesignedtheinterfacebasedontheAI controllability exploration principles proposed in [40]. Inspired by [30,16], we implemented multiple output principles together with general and domain-specific controls, allowing users to explore a wide range of possibilities and better understand the model’s capabilities. To decide which control options to include, we analyzed Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 725 F. Ronchini et al. (a) (b) Fig.2: User interface: (a) music generation model; (b) source separation model. some interfaces already available on Hugging Face. To ensure ease of use and accommodate the experiment’s time constraints, we kept the interface simple with a limited set of controls. The interface, shown in Figure 2, presents one tab for music generation and one tab for source separation. Users can control the generation by selecting output duration, MusicGen size, and input prompt. The separation interface o!ers stems for drums, bass, guitar, piano,andother,with vocals excluded as MusicGen does not generate them. It was developed with Gradio for its intuitive design, rapid deployment, and focus on machine learning models. We built a custom interface for better experiment control; one author developed it, and two tested it. Experimental Procedure. The experiment was conducted live, either in person or online, using the same procedure. Participants scheduled a one-hour session with the experimenter and used their own laptops to use their preferred setup (e.g. DAW, headphones) to produce music. They accessed the interface through a web app, facilitated by ngrok [27], with secure remote access via SSH. The interface ran on a Quadro P6000 GPU on a private cluster. Each session was divided into three steps: 1. Pre-experiment step:participantswereintroducedtotheexperimentand then completed a survey on their music production experience, familiarity with AI tools, and expectations in the creative workflow. 2. Hands-on step:participantswereinvitedtocreateamusicexcerptusing the interface and their preferred setup. They prompted the model, using the available controls, then generated and downloaded the audio or used the separation model to obtain individual stems from the generated audio. They had full freedom to interact with the models, and restart the generation process as desired, with no music genre limitations. The session was recorded via screen and audio, with observational notes taken throughout. 3. Post-experiment step:attheendofthesession,participantsevaluatedtheir user experience, workflow creativity, productivity, TTM e!ectiveness, ethical considerations, and provided additional feedback in a semi-structured interview. The preand post-session questionnaires were adapted from [26,30,24] and the Goldsmiths Musical Sophistication Index [25]. Participants were recruited Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 726 AI-Assisted Music Production: A User Study on Text-to-Music Models 01234567891011121314 1 2 3 4 5 6 7 8 9 10 Number of Responses Q5 - Rate your expectations for AI tools to improve creativity Q4 - Rate your familiarity with MusicLM Q3 - Rate your familiarity with Udio Q2 - Rate your familiarity with MusicGEN Q1 - Rate your familiarity with SunoAI 12345 Fig.3: Diverging bar charts showing users’ familiarity with AI tools and users’ expectation for AI to improve creativity. through music producer communities and relevant mailing lists, with a minimum level of music production experience required. A total of 17 international participants from 7 di!erent countries participated in the study (8from Italy, 2from India, 2from the U.S.A, 2from the People’s Republic of China, and 1 each from the Republic of Korea, Greece, and Spain). 4participants attended in person and 13 remotely via Zoom platform; each participant was assigned a random ID. 4Results Section 4.1 and section 4.2 present Likert-scale and multiple/single-choice answers from the preand post-experiment questionnaire, respectively. Section 4.3 provides a thematic analysis of the open-ended questions of the post-questionnaire. Likert-scale answers are shown using diverging plots, where positive responses are on the right, negative on the left, and neutral responses are split between both sides. The bar length represents the number of responses, di!erent colors indicate specific scores. For the Likert-scale-based questions, we report the ratings on a scale from 1 to 5 due to space limitations. However, in the questionnaire, each question retained its appropriate responses according to the Likert-scale answer guide1. 4.1 Pre-experiment questionnaire: Musical background and AI tool familiarity Music Production Experience and Genre.Amongparticipants,10.5% have more than 5 years of experience, 42.1% have 3-5 years, 21.1% have 1-3 years, and 26.3% have less than 1 year of experience. Logic Pro is the most popular DAW, used by 35.7% of participants, followed by Ableton at 28.6% and Reaper at 14.3%. Other DAWs mentioned included Studio One, Cubase, Pro Tools, and Samplitude. 76.5% of participants had used AI tools for music creation before. 1www.extension.iastate.edu/documents/anr/likertscaleexamplesforsurveys.pdf Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 727 F. Ronchini et al. 012345678910111213141516 1 2 3 4 5 6 7 8 9 10 11 12 13 14 Number of Responses Q8 - After this experiment, how useful did you think AI could be in music production? Q7 - Before this experiment, how useful did you think AI could be in music production? Q6 - How effective did you find the source separation process in isolating and preserving individual elements of the music? Q5 - How satisfied are you with the tool’s responsiveness to your prompts? Q4 - Did the generated loops/instrumentals inspire your creativity? Q3 - How easy was it to integrate the generated loops/instrumentals into your workflow? Q2 - How much control did you feel you had over the AIgenerated music? Q1 - How hard was it to navigate the Gradio interface? 12345 Fig.4: Likert-scale answers from the post-experiment questionnaire. The main genre of music production is Pop,mentioned8times, then Classical, House / EDM / Dubstep and Rap / Trap mentioned 5times each, while Rock /Metaland R&B and Lo-Fi were mentioned 3times each. Carnatic Music, Urban, Latin, and Acoustic appeared once each. They could list multiple genres. TTM models familiarity and expectation. Fig. 3 shows users’ familiarity with key music generative models and their expectations for AI tools to enhance music creativity. Other mentioned AI tools include SynthPlant, NetEase, Tianyin AI, Watson Beats, ImprovNet, Stable Audio, Beatoven.ai, Soundraw, Mubert, iLoveSong.ai, and AIVA. Suno AI stands out, likely due to stronger brand recognition or better accessibility compared to others. While there is an optimistic outlook on AI’s role in creative processes, some participants remain cautious. To gain deeper insight into these expectations, we asked participants what they hoped to achieve using AI tools in their creative workflow. A brief thematic analysis of their answers revealed key themes: e"ciency and time savings, inspiration, customization and personalization, simplification for less experienced producers, skill development, and sound quality. 4.2 Post-experiment questionnaire: AI tools for music productivity Likert-scale questions. Figure 4 shows answers to 5-points Likert-scale questions. Participants responded positively to integrating AI-generated loops into their workflow, with many finding the loops inspiring creativity, positioning the tool as a potential creative collaborator. Perceptions of control over AI outputs and satisfaction with the tool’s responsiveness varied. Opinions on the source separation feature were similarly mixed. Single and multiple-choices questions. We asked participants where they would integrate the tool in the music production workflow. 94.12% of participants would use it “during the ideation phase”, while 47.06% consider it potential for “variations and experimentation”;35.29% identified composition and arrangement as a key stage, and 23.53% mentioned “remixing and reworking existing Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 728 AI-Assisted Music Production: A User Study on Text-to-Music Models tracks.” Other responses included “sound and track generation” and “digital audio generation”. We also asked if they would incorporate an AI-based source separation tool into their production workflow. 52.94% expressed a strong interest in using it for music production, while 41.18% indicated they would use it only in specific scenarios. 5.88% preferred manual mixing or alternative methods. We then explored which applications participants found most suitable for AI-generated music. 88.24% identified “background music for games, YouTube videos, or other media” and “music production assistance” as primary use cases, highlighting its potential to support the creative process. 64.71% pointed to “experimental production and sound design”,reflectinganinterestinexploringnew creative directions. 47.06% consider it useful for “film, TV, or commercial soundtracks,” while 35.29% mentioned “remixing and reinterpreting existing tracks”. 17.65% would use it for “direct listening”. 4.3 Thematic analysis An inductive thematic analysis was performed on the open-ended questions, resulting in 39 initial codes, which were iteratively merged into 14 themes, organized into 3sections using a"nity diagrams [16]. The sections are: productivity and e!ectiveness, usability and creative workflows, and ethical considerations. One author drafted the initial codebook, while two other authors iteratively refined the final themes, which are reported in bold in the following subsections. Productivity and e!ectiveness. Creative Misalignment.Discrepancybetween participants’ creative visions and the generated output emerged as a key theme. P12: “It didn’t go well with my initial ideas as planned”,P14:“I had something in mind but it proposed me some di!erent and unexpected things”, highlighting a mismatch between expectations and results. P15 mentioned, referring to the content generated: “sometimes it deviates from what I had in mind”, emphasizing the challenge of aligning with personal projects. P16: “I had some creative ideas about how the song I wanted should sound like, but I didn’t get the result that I had in mind”, while P2 stated: “The content generated didn’t reflect what I had in mind”.P11:“Some of the instrument sounds I asked for didn’t come out as expected”,furtherillustratingthisgap. Integration Challenges. Tempo, key, and beat alignment emerged as main issues for integrating generated content into projects. P12: “Generated output was not loop-based, and the tempo was not accurate. It was di"cult to slice them and beat-match to another track. About key, I wasn’t sure if they were correct”. P2 similarly pointed out (referring to the fact that the content generated did not reflect what they had in mind): “[..], especially in terms of BPM”.Source separation also posed di"culties. P17: “The guitar and piano separation tools did not pick up audible frequencies”,P8noted:“Noisy samples after the separation”. Participants P8 and P10 faced challenges with one-shot samples generation. From expectation to innovation. When outputs did not align with their original ideas, many participants changed their approach. P2: “I can’t use the Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 729 F. Ronchini et al. generated samples as main elements, but they’re useful and beautiful as background, not what I originally intended”.P7:“The AI tool redefined my project, inspiring new ways to create rather than helping me achieve my original idea”. Similarly, P12 added, referring to the fact that it did not go well with their initial idea: “However, by modifying the samples in GarageBand’s audio-flex functions, I was able to discover a very new direction of experimental music”.P14:“It transformed my idea into something new. I had something in mind, but it proposed di!erent and unexpected things”. These responses reflect how, despite results not aligning with initial ideas, the TTM model became a source of unexpected inspiration, encouraging participants to explore new creative directions. AI as an e!cient collaborator.Someparticipantsconsideredthetool as an e"cient collaborator and a creative partner. P4 highlighted the tool’s usefulness, even if “just for accelerating the generation process of something very precise I have in mind”.P11:“In some cases, it’s easier to generate what I want than it is to search for a loop”.P17:“It created a good foundation for work to continue but gave space for me to build something of my own, which is nice, not too prescriptive”. These insights suggest that, in specific scenarios, AI can speed-up the process and be an e"cient collaborator. Usability and creative workflows. Need for greater control. Concerns were raised about limitations in controlling the generation process. P15 recommended integrating “multi-modal prompts” for versatility. Many participants highlighted the need for control over tempo alignment, key selection, and loop duration. P17 proposed “More parameters like key signature or BPM for better alignment of the output”.P8emphasizedtheneedof“enhanced controllability of accurate BPM and key info”, while P6 wanted to “change/select tempo and measure of the loop”.P16recommended“fine-grained control”, while P1 suggested to allow users to choose the type of input prompt (text, tag, or reference audio). Lack of Editability. Participants expressed frustration with the inability to modify or refine the generated content. P6: “It was frustrating that I was not able to modify the tracks/ideas given".P7suggestedadding“a second prompt to a result, memorizing it and changing only a part of it” for more control over specific elements. P12 emphasized the need for an “iterative process of re-generation of samples” and proposed “music editing through a piano-roll interface on the generated audio samples” for precise manipulation. Partial usability of generated content. Concerns were raised about the limited usability of the generated content. P11: “After the clips were generated, some would be usable and others wouldn’t. Also, some of the instrument sounds I asked for didn’t come out as expected”,highlightingthatonlycertainpartsofthe output met expectations. P12 mentioned: “I expected a stem-wise output (e.g., by specifying jazz-marimba-solo), but it often comes with other instruments mixed in”,emphasizingthechallengeofisolatingspecificsounds. Ethical considerations. Copyright concerns and AI regulation. Copyright regulations were among the main considerations. P3: “a group of experts Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 730 AI-Assisted Music Production: A User Study on Text-to-Music Models should intervene to evaluate how much an AI-generated track takes from existing pieces”,suggestingtheneedforhumanoversight.P5emphasizedtheimportance of “strict regulations about copyright and AI training data”.P17calledfor a“broader overhaul of the copyright system”,advocatingforasystemicshiftin how copyright laws should adapt to AI-generated works. Artist Compensation.P17:“revision of the way artists are compensated for their work”.P16arguedthat“artists whose work has been used to train AI should be compensated”.P4stressedtheimportanceofensuringthat“artists are compensated each time their data is used”. Diversity and risk of homogenization in AI-generated music. The risk of homogenizing music was noted, stressing the need for diversity in music generation. P17: “AI training should span a wide range of music styles”;P3: “we should limit ourselves to get inspiration from the AI and to ease the sound generation process, not actually to make the AI create our music”.P6pointed out that in generating ethnic music genres, the model “fell short”, underscoring the need for better representation of diverse cultural sounds. 5Discussion While TTM models show promise in artistic applications, they primarily serve as aids for sketching and inspiration rather than as production-ready solutions. This reinforces the idea that AI serves as an innovative set of tools for creative exploration [15]. However, integrating TTM-generated audio samples into human creative workflows remains challenging due to control limitations. We recommend practical design improvements that address current limitations, such as improving tempo, key, and beat controls, editing capabilities, support for iterative re-generation, and, very importantly, DAW integration. These features would greatly strengthen the usability and creative potential of the tools. Additionally, interpretative gaps often arise between the AI’s output and the artist’s intended vision, particularly when producers have a clear idea of what they want. Some producers embrace these tools as a source of inspiration, adapting their creative process to incorporate the AI’s output. Others, however, report having to modify (sometimes completely changing) their original idea to accommodate the model’s unexpected outputs. This dynamic illuminates the role of AI as inspiration, prompting a question: does AI empower artists by enhancing creative exploration, or, on the flip side, does it constrain the process by steering creativity in unintended directions? Similar patterns are observed in [12]. Creative ideas are often discarded when they fail to integrate with other creative elements. Future models should consider additional input controls and allow for editable content to better align outputs with users’ intentions and empower them to refine specific sections. These challenges were already highlighted in earlier research, which led to the introduction of controls for tempo, rhythm, and chord progression [41,37,18,6]. However, many models are primarily evaluated based on technical performance, such as audio quality or prompt adherence, rather than through user-centered case studies. Misalignment between user inProc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 731