Full text
Master thesis on Sound and Music Computing Universitat Pompeu Fabra InScoreAI: Collaborative Score Inpainting with Anticipatory Transformers Manuel Lallana Babiloni Supervisor: Rafael Ramírez-Melendez July 2025
Master thesis on Sound and Music Computing Universitat Pompeu Fabra InScoreAI: Collaborative Score Inpainting with Anticipatory Transformers Manuel Lallana Babiloni Supervisor: Rafael Ramírez-Melendez July 2025
Contents 1 Introduction 1 1.1 Significance and Motivation . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Contributions................................ 3 1.3 Objectives.................................. 3 1.4 Structure of the Report . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2 State of the art 5 2.1 Introduction................................. 5 2.2 Human-human collaborative composition . . . . . . . . . . . . . . . . . 5 2.2.1 From co-creation to networks . . . . . . . . . . . . . . . . . . . . . . . 5 2.2.2 Collaborative Physical Interfaces . . . . . . . . . . . . . . . . . . . . . 6 2.2.3 Collaborative Digital Interfaces . . . . . . . . . . . . . . . . . . . . . . 6 2.2.4 Software Ecosystem Analysis . . . . . . . . . . . . . . . . . . . . . . . 8 2.3 Human-AI Music-Making . . . . . . . . . . . . . . . . . . . . . . . . . 8 2.3.1 AIMusicGeneration............................ 8 2.3.2 Human-AI collaboration . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.4 Bridging the gap: Human collaboration assisted by AI or AISCW . . . 13 3 Methods 17 3.1 Preliminarysurvey ............................. 17 3.2 InScoreAI Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.2.1 Webarchitecture .............................. 20 3.2.2 API ..................................... 25
3.2.3 Deployment................................. 30 3.3 Experiments................................. 30 3.3.1 EvaluationSurvey ............................. 31 3.3.2 Experiment 1. Individual composition without AI tools . . . . . . . . . 31 3.3.3 Experiment 2. Individual composition with AI tools . . . . . . . . . . . 32 3.3.4 Experiment 3. Collective composition with AI tools . . . . . . . . . . . 32 4 Results 33 4.1 Quantitative Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 4.1.1 Individual No-AI use findings . . . . . . . . . . . . . . . . . . . . . . . 33 4.1.2 AIusefindings ............................... 33 4.1.3 Collaboration findings . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 4.2 QualitativeAnalysis ............................ 34 4.2.1 Effect of Years of Experience . . . . . . . . . . . . . . . . . . . . . . . . 35 4.3 Tablesandgraphics............................. 36 5 Discussion 39 5.1 Discussion.................................. 39 5.1.1 Contributions................................ 40 5.1.2 Futuresteps................................. 41 5.1.3 Limitations ................................. 42 5.2 Conclusions................................. 43 Bibliography 45
Acknowledgement I would like to thank my supervisor, Rafael Ramírez-Melendez, for the effort and time he dedicated to this project. He believed in my ideas and proposals and always encouraged me to explore new solutions. I would like to thank Anna Xambó and Álvaro Barbosa for their guidance and Xavier Serra for putting us in contact. The conversations we had opened new paths and insights, and they generously dedicated their time and knowledge. I am thankful for my Master and Bachelor colleagues for all their support and feedback as well as for Valerio Velardo for providing us valuable knowledge on the most recent model architectures and sharing the Anticipatory Transformers with me. This work was made possible thanks to all the musicians and composers that filled the preliminary survey and participated in the experiments, it has been a pleasure to create something for you. Finally, I would like to thank my partner Laia, my parents and sister for all their love and care throughout an intense and fulfilling year.
Abstract Collaborative music composition is the process by which multiple individuals contribute to the creation of a musical work. In this project, an interface is created to help composers create collaboratively and reach consensus through flexible symbolic music generation using Anticipatory Transformers. This model generates MIDI inside a specific fragment of music taking into account the previous and following tokens. The central and most important part of the project are musicians and composers, they’ve been the focus from the start to the development and evaluation. A preliminary survey for experienced musicians and composers was developed to extract important features for the interface, the evaluation consisted of three tasks, one where the interface was used without the AI tools, one with the AI tools and the last one using the collaborative mode with AI tools. Results show that AI assistance significantly reduces the compositional effort and improves self-efficacy in individual workflows, but diminishes perceived ownership. In collaborative settings, the system effectively resolves interpersonal friction through AI-mediated "idea bridging," fostering consensus and mutual understanding, sense of ownership is higher compared to an individual AI setting though it is still a concern for some participants based on long responses. Professional composers showed concerns about AI making composers “lazy” and the works simple and predictable, while nonprofessionals outlined educational potentials for music collaborative settings. Keywords: Human computer interaction (HCI); Symbolic Music Generation; Transformer model; Collaborative music composition;
6Chapter 2. State of the art presented to give a bird’s-eye view of the topic. A well documented example is the collaboration between the experimental artists John Cage and David Tudor who worked with indeterminate scores. The performer had the freedom to define some musical parameters; Glover [1] considers this to be an example of collaborative musical practice. Later on, The League of Automatic Music Composers, founded by Jim Horton, put emphasis on connections between musicians and the use of computers to define new social contexts of music, while the collective known as The Hub (1980s) pioneered computer-mediated social interactions in music through real-time data exchange between performers, creating emergent structures from decentralized decisions [2]. Barbosa [3], who delivered the first systematic taxonomy of computer-supported collaborative music systems, analyzed this and other early manifestations of networked music and introduced the “Shared Sonic Environment” concept via the latencyadaptive Public Sound Objects (PSOs) platform. 2.2.2 Collaborative Physical Interfaces Tangible interfaces prioritized egalitarian participation through spatial interaction paradigms. Jam-O-Drum [4] used circular multitouch surfaces to enable non-hierarchical rhythmic collaboration, while the famous reacTable [5] mapped physical object positions to sound parameters, allowing simultaneous manipulation by multiple users. Multi-surface devices and smartphones have been used as a tool of music collaboration, showing that responsiveness and visual feedback is important for the user [6], and that real-time shaping of musical ideas fosters unexpected and experimental results [7]. These systems traded precision for accessibility, fostering “low-floor” engagement at the cost of limited expressive depth. 2.2.3 Collaborative Digital Interfaces Regarding collaboration using technological interfaces for music composition, a system for collaborative music composition over the web was introduced in [8]. This
2.2. Human-human collaborative composition 7 was one of the first collaborative tools for music and included logs of versions and voting. The development of shared virtual environments [9] and cloud-based file sharing [10] for music collaboration culminated in projects like DISSCO [11], a platform that integrates algorithmic composition, sound synthesis, and user input in real-time, enabling multiple participants—whether co-located or remote—to collaboratively create and modify musical compositions. Another important line of research has been the notion of mutual engagement in creative collaborations [12] [13]. This area explores how engagement in music collaboration in a digital tool can be measured. They found that adding instructions that emphasize collaboration can lead to more engagement, which was identified through several factors (like proximity of contributions, mutual modification, and acknowledgment, mirroring, or transformation of others’ inputs). Moreover, they found a critical challenge: participants reported a clash of ideas and a lack of communication channels. In this work, we will tackle this issue by introducing AI tools to create inpainting proposals in order to reach consensus. Digital interfaces for music collaboration have been also tackled by the field of music education, with platforms enabling “a powerful route in for compositional thinking” [14]. Studies of online composition [15] revealed that some asynchronous workflows improve organization and time management. Several studies have shown successful evaluation metrics for collaborative environments, especially the Creativity Support Index (CSI) [16], a psychometrically validated survey tool designed to evaluate how effectively a technology (e.g., software, digital tool) supports users in creative tasks, and a project by Burkhardt et al. [17], an evaluation of collaboration in tech-driven design (e.g., virtual environments) that emphasizes mixed-method frameworks blending quantitative metrics with qualitative insights.
8Chapter 2. State of the art 2.2.4 Software Ecosystem Analysis Sometimes, music software incorporates collaboration in the form of cloud syncing and file sharing, but it is rare to find real-time editing as being an essential part of it. The most notable music editing software, Musescore, Sibelius and Dorico, offer professional editing but lack real-time collaboration, although some plugins have been developed for this purpose [18]. Noteflight, Flat.io and Soundslice are web applications with collab support to some extent, but, as the already mentioned editing software, they lack AI tools for symbolic music generation. An exception of this is AIVA, an AI music generation assistant that invites the user to co-create tracks. Unfortunately, its input and visualization have the style of a DAW with piano roll, but it is not a score editing tool. In general, this tool seems to better fit music producers, not composers. Table 1 shows how current musical software supports diverse collaboration capabilities, but often lacks AI assistance and real-time collaboration, and when it does, it lacks score edition. Software Collab Type Notation Support AI Features Musescore Cloud Sync (MuseLab Plugin for collaboration) Professional None Dorico Async file sharing Professional None Sibelius Cloud sync Professional None Noteflight Real-time web Basic None Flat.io Real-time web Intermediate None Soundslice Sharing by links Basic None Staffpad Async cloud Handwriting None AIVA None DAW integration Full AI generation Table 1: Comparative Analysis of Collaborative Music Software 2.3 Human-AI Music-Making 2.3.1 AI Music Generation The taxonomy proposed by Zhu et al. [19] shows that AI-driven music generation systems can be grouped into three methodological families, each defined by the input representation and learning paradigm they employ.
2.3. Human-AI Music-Making 9 1. Parameter based models: These systems treat music as an ordered sequence of symbolic parameters (notes, durations, dynamics, etc.) and generate new sequences by modeling the statistical or heuristic relationships among them. •Markov chains (e.g. Markov Melody Generator): transition matrices store the probability of one musical event following another; generation proceeds by sampling from the learned matrix. •Rule based systems (e.g. MusicScore): expert defined harmonic and rhythmic constraints that delimit a search space where valid sequences are assembled. •Evolutionary algorithms (e.g. GenJam): candidate sequences go through mutation and crossover and a fitness function guides selection across multiple generations. •Neural networks (e.g. Magenta,Jukebox,MuseNet): recurrent or Transformer architectures that learn long-range dependencies in large MIDI or raw audio corpora and autoregressively predict the next token. 2. Prompt based (conditional) models: A natural language or audio prompt conditions the model, providing music that is, ideally, aligned with the description (e.g. Riffusion,MusicLM,MusicGen). 3. Visual based models: These approaches convert video or visual embeddings into an intermediate representation that conditions the music generator (e.g. CMT,Foley Music). They identified the next limitations in most of the AI music generation models: •Limited Control: Some models, such as MuseNet, may not adhere strictly to user-specified instruments, leading to unexpected results. •Data Dependency: Tools like Jukebox are limited by the diversity of their training datasets, affecting their ability to generate specific styles.
10 Chapter 2. State of the art •Computational Complexity: Models like Jukebox can be resource-intensive, making them less accessible for users without high-performance computing resources or for live applications. •Quality of Output: Generated music often requires manual fine-tuning, as the initial output may not meet user expectations. These limitations reveal the challenges of achieving fully autonomous and creatively rich AI music generation. AI Symbolic Music Generation Recent research on AI-driven symbolic music generation has highlighted the complexity of automating the composition process. Early deep learning approaches have excelled at capturing short-range dependencies to produce realistic local musical patterns, yet have often struggled with long-range structure, style alignment, and limited creative potential [20]. These limitations could also arise from the lack of unified methods for evaluating model performance, making it difficult to assess whether generated music meets professional standards of originality and coherence. Multiple studies advocate for collaborative approaches that foreground expressiveness and human–machine co-creation [21], this approaches often propose “music inpainting" solutions. Music inpainting Music inpainting addresses the completion of missing segments within a musical work, aligning closely with iterative human creative workflows. The task is to generate a musically coherent sequence that bridges the given past context with the future context. Early methodologies proposed probabilistic sampling. Gibbs sampling, a Markov Chain Monte Carlo (MCMC) technique, has been widely applied—notably in DeepBach for resampling notes/voices in Bach chorales. Similarly, Coconet (CNN-based)
2.3. Human-AI Music-Making 11 was developed to complete partial scores, framing generation as iterative “rewriting". These piano-roll models, however, were limited to fixed-length spans [21]. Extending this, Ippolito et al. [22] used Transformers for infilling missing MIDI performances via Gibbs sampling, enabling composers to iteratively rewrite or morph sections. Learning expressive latent spaces has also enabled greater diversity and user interaction. Pati et al. [23] introduced InpaintNet (VAE-based), generating connecting bars between clips via learned latent trajectories, allowing contextual manipulation. Chen et al. [24] proposed Music SketchNet, incorporating user-specified pitch/rhythm constraints ("sketches") to fill monophonic gaps. A key limitation of both bar-level VAEs was rigid alignment to measure boundaries. Recent work overcomes fixed-span limitations. Chang et al. [25] leveraged XLNet for variable-length completion, introducing relative bar encoding for nuanced positional awareness. Guo et al. [26] developed a Transformer framework enhancing stylistic fidelity to original material through control tokens governing key, track density, track polyphony, tensile strain, bar cloud diameter, and track occupation. This granular control facilitates interactive collaboration. Important technical advances in controlling symbolic outputs can be seen in works such as the Anticipatory Music Transformer, where this inpainting paradigm helps generate specific passages around user-defined anchors [27], and MIDI-GPT, which offers track-based conditional generation for comprehensive musical arrangements [28]. Finally, a recent model called NotaGen emphasizes training paradigms from large language models to improve musicality and user-controllable features [29], showing how domain-specific adaptation can refine structure and coherence in classical compositions. 2.3.2 Human-AI collaboration Zhu, et al. [19] mention the interdisciplinary potential of collaboration between humans and machines. This Human-AI collaboration is research that most of the time focuses on interactivity, steerability and control.
12 Chapter 2. State of the art Huang, et al. [30] emphasize the importance of developing AI tools with creative workflows in mind and incorporating existing musical practices. Creators see AI as a partner for ideation and exploration but emphasize maintaining control over the final product [31]. In the same lines, AI-steering tools were found to improve the sense of ownership in the users [32]. One of the most notable research efforts on AI-steering tools for Human-AI music creation is presented in [33]. The study introduces “Cococo”, a web-based music editor that integrates AI-steering tools to support iterative music composition and enhance collaboration. Their claim is that not as many studies have paid attention to the potential and limitations of co-creation with AI tools. Their approach is to provide a web interface where you can compose SATB chorales using a Deep Neural Network trained on Bach pieces. You can select what voice to generate and generated content appears as multiple alternatives you can audition and choose from, improving controllability over the final output. In their user study with music novices, the tools helped participants feel more empowered and connected to the compositions, while fostering a positive perception of AI’s role in creative tasks. This user study was evaluated by asking the users to compose a piece using a non-steerable AI tool like “Bach Doodle” and then composing a pice using “Cococo”. Then the users filled a survey with 7-point Likert scale items regarding efficacy, engagement, effort, controllability and other useful items of Human-AI collaboration. Based on the results extracted from analyzing and comparing those items from both tasks, we can observe that users benefited from: •Increased sense of control, trust, and ownership of compositions. •Improved ability to solve problematic areas, learn new musical structures, and explore AI’s limits. •Enhanced self-efficacy, creativity, and engagement in the co-creation process.
2.4. Bridging the gap: Human collaboration assisted by AI or AISCW 13 Users saw the AI as a collaborative partner when steering tools were available. Some participants envisioned using AI with different roles for different tasks. Therefore, future interfaces could let users define their creative goal, adapting the AI’s role accordingly. For fuzzy goals, the AI might automatically explore ideas; for clear goals, it could respond to specific requests. Steering tools thus offer not only control over direction but also the deliberate option to invite creativity and new possibilities by ceding control. 2.4 Bridging the gap: Human collaboration assisted by AI or AISCW Later, “Cococo” was tested in a collaborative human-human environment as a possible way to ease social friction during creative tasks [34]. Users were much more comfortable judging and debating AI proposals than human proposals, allowing them to establish common ground and progress efficiently. The study revealed that the non-human identity of the AI acted as a psychological safety net: participants freely rejected AI ideas without fear of offending a collaborator, reducing interpersonal tension. In addition, AI offered multiple alternatives during creative disagreements, allowing teams to move away from a dead end toward a compromise. The limitations of this study show that there is a lot of space for improvement: 1. Although “Cococo” facilitated collaboration, it was not originally designed for multi-human workflows. 2. Instead of making “Cococo” collaborative, they used the Google Remote Desktop extension, allowing one of the participants to access the screen of the other user. This makes them share the same view of the platform but lacks essential collaborative features like independent actions for each user and visual feedback of these actions. 3. Some of the participants indicated that the system turned humans from “composers” to “advisors” or “curators” of AI output.
14 Chapter 2. State of the art 4. Its Bach-inspired training data constrained stylistic diversity. Despite these limitations, the study uncovers the potential of AI as social glue: by sharing starting points, accelerating progress through auto-completion, and reframing criticism as collaborative “problem solving with a third party”. They evaluated the experiments by conducting semi-structured interviews from which high-level categories and themes were extracted. In our consideration, this approach could be benefited by the 7-point Likert scale items proposed by [33]. A recent study carried out by Fu et al. [35] used already developed AI audio tools (e.g. Suno and MusicGen) to carry collaborative compositions with novice creators. Their findings indicate that AI in collaborative environments can be perceived as: 1. As a "team member", with creative responsibilities (e.g., lyric generation, melody composition). 2. As a technical assistant with mixing, mastering, or genre-specific expertise where novices lacked proficiency. 3. As a mechanical output generator, criticized for producing content lacking emotional depth. Although these tools lowered entry barriers (enabling novices to create full songs in minutes), they introduced critical limitations: 1. AI outputs were often inflexible (e.g., uneditable WAV/MP3 files) and hard to refine and integrate with human-generated elements. 2. Participants emphasized over-reliance on AI risked creative identity. Some of the teams consciously limited the role of AI to preserve artistic agency. 3. Despite accelerating ideation, AI compressed the preparatory stage, causing “decision fatigue” during idea validation. Novices where overwhelmed by abundant low-quality options and they stick to “satisficing choices” due to limited domain knowledge.
2.4. Bridging the gap: Human collaboration assisted by AI or AISCW 15 The study identifies seven different stages of co-creation in their “Human-AI CoCreation Stage Model”: Problem identification, Preparation, Idea generation, Idea selection and validation, Collaging, refining and integration and Outcome. In the Collaging stage, participants manually merged fragmented AI outputs into cohesive works, a process that lacks controllability and was described by the participants as time consuming. For evaluating the study, they combined artifact analysis, semi-structured interviews, and embedded ethnography. Over a 10-week undergraduate capstone course, researchers systematically analyzed teams’ weekly process logs (e.g., ideation iterations, role assignments), creative outputs (draft lyrics, Spotify-published EPs), and collaborative platforms (Figma/Miro boards). In-depth interviews with 9 participants captured insights on workflow dynamics, perceived AI roles, and ownership challenges, song playback was used to stimulate reflection. All data were analyzed by thematic analysis, first coding emergent themes and then mapping the findings to the stage model, revealing critical deviations from it such as a compressed or inexistent preparation phase. This design captured behavioral dimensions (e.g., time spent editing AI outputs) and perceptual dimensions (e.g., self-efficacy shifts). To the best of our knowledge, these are the most notable and recent studies on how AI can affect collaborative musical composition practices. This thesis is an invitation to explore this gap that could be called “Humans-AI Collaboration” or “AI Supported Collaborative Work (AISCW)”. By joining the findings of both Human-Human and AI-Human music composition, we can answer some of the identified limitations of AI models (lack of control and creativity) as well as evolve the field of collaborative music composition by utilizing AI as and assistant of collaborative processes. Collaborative tools lack mechanisms to resolve the “clash of ideas” problem identified in [12], where participants struggle to reconcile divergent concepts without mediation. AI could act as a decentralized mediator through two key strategies: Idea Bridging: A Transformer model trained in past and future context (inpainting) samples could generate transitional material between conflicting contributions,
22 Chapter 3. Methods Figure 2: Web Architecture Figure 3: Web GUI tk.getPageCount() // Handles pagination tk.getTimeForElement() // Measure Timing tk.renderToMIDI() // Midi generation Measures were rendered as interactive SVG elements. Click handling leveraged
3.2. InScoreAI Architecture 23 Verovio’s xml:id to track selections (e.g., selectedMeasureIds). Start and end time of the elected measures to re-generate was extracted using Verovio’s time-mapping (tk.renderToTimemap()). Editing Workflow MEI DOM Manipulation was done by directly modifying MEI XML using DOMParser to insert/delete notes/rests based on user/MIDI input: const meiDoc =parser.parseFromString(originalMEI, "text/xml"); const noteElement =meiDoc.createElementNS(MEI_NS, 'note'); layerElement.appendChild(noteElement); Verovio validated edits against musical rules (e.g., measure duration checks using time signature data from scoreDef) so users could not insert durations over the maximum duration of each measure. The workflow is as follows: 1. User toggles edit mode with a button 2. User selects a measure to input new notes 3. User selects a note duration with the computer keys (1 for whole note, 2 for half note, 3 for quarter note, etc.) 4. User plays a note on their MIDI keyboard (received through the Web MIDI API) Playback System MIDI was generated via tk.renderToMIDI() and synchronized with Tone.js for audio playback. Visual feedback during playback used Verovio’s tk.getElementsAtTime() to highlight active notes. Verovio-generated MIDI was decoded and scheduled using Tone.Part, with timing derived from Verovio’s temporal data (playbackOffset). Users can play or stop playback of a selection of measures or the whole score.
24 Chapter 3. Methods Collaboration (PeerJS) As stated in the official documentation “PeerJS wraps the browser’s WebRTC implementation to provide a complete, configurable, and easy-to-use peer-to-peer connection API. A peer can have a data or media stream connection with a remote peer.” Verovio’s MEI state was serialized and synced in real-time between peers. Full-state broadcasts included MEI data and edit history. In our application User A connects to User B via Peer ID, full state synchronization is active and both users see real-time updates for: •Score changes •AI proposals (red highlight of the range generated) AI Proposal System AI-generated MEI fragments (received from the Flask backend) were integrated using Verovio’s measure-replacement functions (replaceMeasureInMEI()). The AI generation workflow is as follows: 1. User selects a group of measures 2. Clicks on harmonize, inpaint or change melody (harmonize will not change the soprano and change melody will not change alto, baritone and bass, inpainting will change all) 3. A “Generating proposal 1/2...” text gets displayed 4. Once the first proposal is generated it is displayed in the score, the notes are updated and the users can listen to it 5. A “Generating proposal 2/2...” text gets displayed, once the second proposal is generated a few buttons appear where the users can independently select accept, reject and listen to proposals
3.2. InScoreAI Architecture 25 6. Once one user accepts a proposal (or rejects), the score gets updated for both users The relationship with the API is as follows: 1. The web sends the MIDI of the selected region, the start_time and end_time of generation and the top_p(nucleus sampling parameter covered in the API section) 2. The API returns a new MusicXML file that gets rendered Challenges and Solutions Table 5 shows the different challenges and solutions faced during the development of the environment. Area Issue Solution SVG Re-rendering Loss of user state (e.g., selected measures) during SVG regeneration Persist measure IDs in selectedMeasureIds and re-apply CSS classes (.highlighted) after rendering Real-Time Collaboration Edit conflicts during concurrent operations Operational transforms using versioned history stacks (historyStack,redoStack) with MEI diffing before broadcast Performance Optimization Latency from full-score rerendering on edits Selective SVG updates using measure-level targeting and Verovio’s partial layout methods Table 5: Technical Challenges and Solutions in Web Implementation 3.2.2 API Symbolic Generation The first iteration of our AI symbolic generation used LLMs (Llama and GPT4o) in different environments (Autogen and CrewAI) to generate abc notation displayed on a web using abcjs. The quality of outputs varied considerably and the length
26 Chapter 3. Methods Figure 4: API Architecture of the generations was not consistent enough (even using pydantic models). It also lacked future context, that is why we went for an inpainting model recommended by the researcher Valerio Velardo, the Anticipatory Transformer. Anticipatory Transformers: The Anticipatory Transformer [27] leverages a decoderonly transformer with a context length of 1,024 tokens. We used the tokenization with best results shown in the original paper; the arrival-time encoding, which is done by representing events as (time, duration, note) triplets (vocab size: 27,512). This arrival-time encoding (3 tokens/event), enables context-free reordering for better sequence modeling. The model’s key innovation is its anticipation mechanism, which interleaves: •Events: Note onsets/offsets •Controls: User-defined constraints (e.g., future melody) During training, controls appear tseconds ahead of events (t= 5sin the implementation), enabling the model to learn dependencies between current events and future constraints. The model was trained on the Lakh MIDI dataset (178,561 files) with time quantized to 10ms resolution, pitch (0–127), and duration (0–10s). Three
3.2. InScoreAI Architecture 27 model sizes were evaluated: Small (128M parameters), Medium (360M), and Large (780M). During inference, nucleus sampling (top-p) is used for generation, sampling from tokens where cumulative probability ≤p(default: p= 0.95). Why using Anticipatory Transformers: Anticipatory transformers offer robust pre-trained models with high quality outputs. You can select the exact range to generate and the conditioning tokens. It also has a nucleus sampling top_phyperparameters users can interact with. Implementation Enhancements: Two critical improvements were made to the original implementation: •Compound Instrument Tokens: The initial code merged multi-track MIDIs (e.g., two piano tracks) into a single track. We introduced compound tokens combining (program, channel)using bit-shifting: compound_instr =(instr << 4)|channel This preserves instrument identity during tokenization and enables four-track separation during generation (e.g., piano, drums, bass, strings). •Metadata Recovery: Time signatures and tempo data (previously discarded) are now extracted during preprocessing and reinjected into generated MIDI files. This maintains rhythmic/metrical integrity: meta_track.append(mido.MetaMessage('set_tempo', tempo=original_tempo)) meta_track.append(mido.MetaMessage('time_signature',...)) Key utilities include: •midi_to_compound(): now preserves multi-track structure via compound tokens •events_to_midi(): Reconstructs MIDI with metadata recovery
28 Chapter 3. Methods •Velocity/note validation: Ensures MIDI compliance (velocity ∈[0,127], note ∈ [0,127]) The compound token fix turns the transformer from a single-track to a multi-track model, demonstrating that architectural adjustments can unlock new creative applications without retraining. Music Transformation Agents: Three agents were coded to take advantage of the transformer for user-driven transformations. All agents follow a unified workflow: 1. Preprocessing: •Convert input MIDI →tokenized events via midi_to_events() •Clip events to user-selected time range [start_time, end_time] 2. Generation: •Conditioned on history (events before start_time) and controls (constraints) •Generate tokens via nucleus sampling (top_p) 3. Postprocessing: •Merge generated/anticipated events →MIDI via events_to_midi() •Inject recovered metadata and clamp invalid values (e.g., MIDI pitches ∈[0,127]) Harmonization Agent: Generates accompaniment conditioned on a melody. Given input MIDI: 1. Extract melody (instrument 0) as controls 2. Generate accompaniment tokens: accompaniment =generate(model, controls=melody)
3.2. InScoreAI Architecture 29 3. Combine new accompaniment with original melody Infilling Agent: Inpaints missing segments using future context. 1. Extract events after end_time as anticipated controls 2. Generate infill tokens conditioned on future: infilling =generate(model, controls=anticipated) 3. Merge infill with original events before start_time and after end_time Melody-Changing Agent: Regenerates a melody based on (and preserving) accompaniment: 1. Treat the accompaniment as controls 2. Generate new melody conditioned on accompaniment: new_melody =generate(model, controls=accompaniment) 3. Combine new melody with original accompaniment API Workflow The Flask API exposes three endpoints: /upload,/uploadinfill, and /uploadchangemelody. Each endpoint: •Accepts a MIDI file and time range (start_time,end_time) •Routes to the corresponding agent (harmonizer, infiller,change_melody) •Converts output MIDI →MusicXML for web rendering: musicxml_str =midi_to_musicxml(temp_midi_path) return jsonify({"musicxml": musicxml_str})
30 Chapter 3. Methods 3.2.3 Deployment First we tried deploying the web and API on the Skynote server, but there were several space and memory issues. The API used the (resource heavy) Transformers and torch python libraries. We tried fixing it by: •Clearing pip caches •Using the lightest anticipatory model •Installing only the CPU version of torch library We solved this issue by: •Deploying the web part using Netlify •Deploying the API using Huggingface Spaces by making a docker container. The API is currently running on a 32 GB RAM server (0.03$ per hour) and can be upgraded to a 15 GB RAM GPU server (0.4$ per hour) The app can be accessed by entering https://inscoreai.netlify.app/ The generation time now is 7 seconds per score measure and all the experiments were conducted using this deployment. 3.3 Experiments Three experiments were designed to evaluate the collaborative AI-assisted composition app. 16 composers participated across 8 90-minute sessions, with sessions conducted via Google Meet. The author was present during the entire session, annotating feedback and issues, this represents approximately 24 hours of work time. Participants were paired according to availability (coordinated via Doodle), and feedback was collected through observation and a post-experiment survey (Google Forms). Musical material (starting and ending segments of 5 measures each) were
3.3. Experiments 31 distributed via WeTransfer. Participants without physical MIDI keyboards used virtual alternatives (MidiKeys for macOS and vmpk/loopMIDI for Windows). Technical problems arose when participants used VPNs, non-Chrome browsers, or mobile 4G network sharing, causing system instability. 3.3.1 Evaluation Survey The post-experiments survey employed a 7-point Likert scale (1: strongly disagree; 7: strongly agree) to measure user experience, drawing metrics from established frameworks (e. g. Creative expression metric: Rate 1-7 your agreement with (“I was able to express my creative goals in the composition made”). Core metrics for experiments 1,2 and 3 included: •Musical Coherence and Control from [32] •Creative Expression, Self Efficacy, Engagement, Completness, Uniqueness, Ownership, Collaboration with the system, Comprehensibility and Effort from [33] The specific metrics from established frameworks for evaluating collaboration in experiment 3 were: •Collaboration, Easiness of idea sharing and Idea tracking from [16] •Fluidity of collaboration, Mutual understanding and Consensus reaching from [17] 3.3.2 Experiment 1. Individual composition without AI tools Participants began by loading a starting 5 measure MusicXML file, adding four empty measures, and adding a 5 measure ending MusicXML file. The task was to manually fill the 4 empty measures via MIDI input without AI assistance. This baseline experiment evaluated individual composition without AI.
38 Chapter 4. Results Table 7: Raw p-values for Experimental Comparisons (Uncorrected) Q= 0.05 Exp1 vs Exp2 Exp2 vs Exp3 Metric p-value Sig. p-value Sig. Creative Expression 0.383 0.206 Self-Efficacy 0.007 0.774 Engagement 0.003 0.751 Completeness 0.708 0.034 Uniqueness 0.580 0.052 Ownership 0,001 0.227 Collaboration w/ System 0.718 0.126 Musical Coherence 0.652 0.055 Control 0.083 0.606 Comprehensibility 0.211 0.299 Effort 0.001 0.027 Table 8: Inferential Statistics (Benjamini-Hochberg Corrected) Q= 0.05 Exp1 vs Exp2 Exp2 vs Exp3 Metric p-value Sig. p-value Sig. Creative expression 0,602 0.378 Self-Efficacy 0,027 0,774 Engagement 0,089 0.826 Completeness 0,779 0,186 Uniqueness 0,798 0.190 Ownership 0,006 0,357 Collaboration w/ System 0,718 0.278 Musical Coherence 0,797 0.153 Control 0,181 0.740 Comprehensibility 0,386 0.412 Effort 0,004 0,293 Table 9: Collaboration Metrics in Exp 3 (N=16) Metric Mean Median Collaboration 6.19 7 Fluidity 6.13 6.5 Idea Sharing 6.31 7 Idea Tracking 5.25 5 Mutual Understanding 6.19 6 Consensus 6.38 6
Chapter 5 Discussion 5.1 Discussion Results confirm that AI integration distinctly reshapes creative processes across individual and collaborative environments, directly addressing our core research questions: 1. Measured metrics (Individual vs. Collaborative): •Individual Use: AI drastically reduced effort (µ= 2.56 vs. µ= 4.94, p=0.004) and boosted self-efficacy (µ= 5.31 vs. µ= 4.69, p=0.027), corroborating studies on AI-steering tools [32],[33]. However, ownership declined significantly (µ= 3.31 vs. µ= 5.62, p=0.006), echoing fears of "soulless" outputs and over-reliance [31],[35]. •Collaborative Use: Ownership remained low (µ= 3.88) and statistically indistinguishable from individual AI use (p=0.357), suggesting AI’s role as a generator rather than co-author dilutes ownership regardless of context. Yet, collaboration metrics (e.g., Consensus, Med=6) exceeded expectations, indicating AI mitigates creative bottlenecks without resolving ownership trade-offs. 2. Resolving Interpersonal Friction: 39
40 Chapter 5. Discussion High fluidity (Med=6.19) and consensus scores (Med=6.38) could suggest the efficacy of AI as "social glue" though qualitative critiques noted AI predictability could harness experimentation (P10, P11). Qualitative responses show that “We agreed fast and easy on which part to do” (P14). The author noted a varied set of approaches from the participants: dividing the composition in voices, in measures, using the AI to generate something while they were working on another part and re-harmonizing sections once they were finished. 3. Composer Requirements: Surveyed composers demanded real-time collaboration (90% voted >= 5), AI-aided technical tasks (e.g., harmonization: 65%), and "idea bridging" was mentioned in the qualitative responses. Our implementation addressed these via anticipatory transformers for multi-track inpainting and measure-level regeneration (features absent in state of the art software analyzed in Table 1). Our findings extend the "steerable AI" paradigm to collaborative settings, confirming that AI enhances co-creation efficacy but complicates ownership. Unlike Cococo [33], which elevated novice ownership via constrained AI roles, our professional cohort perceived AI as a substitutive force, urging designs that preserve human agency. 5.1.1 Contributions The contributions of this work are as follows: Anticipatory Transformers can now: •Generate multitrack MIDIs. •Be used in a new manner (Melody change, conditioned on accompaniment). •Be visualized in a web interface. •Be visualized collaboratively in real-time. A fully functional collaborative score editing + AI full-stack web app was designed and deployed. Composers now have a new tool to resolve:
5.1. Discussion 41 •Clash of ideas. (A problem identified in the state of th art of Human-Human Collaboration) •Lack of control with AI. (A problem identified in the state of the art of HumanAI Collaboration) Finally: insightful feedback from 20 professional composers on ideal features for future collaborative applications was extracted. Also, long answers from the final survey reveal some possible educational potentials of collaborative apps. 5.1.2 Future steps Web future steps As mentioned above, the preliminary survey revealed several features that can be implemented in the future (they were not implemented due to time constraints). These features are: a commenting/annotation system, contextual music theory guidance, version comparison analytics (or full version control) and contribution tracking metrics. Some of this would be aligned with the improvement of the Idea tracking metric, which, as we mentioned, was the lowest. Other features we suggest are user log-in and note-level generation inside the interface. Ideally a new open source score editing library for the web should be developed but, if this is not possible due to time and resources, the author suggests (for a production ready application) either contacting with the Noteflight team for the use of their notation editor or developing a plugin for MuseScore and keeping the web for visualization, AI, and version control. API future steps Future steps in the API could be focused on self-hosting the API on a powerful server. The anticipatory transformer could be fine-tuned to fit the desired genre of the composers for more rich details. One (not yet tested) implementation of LLMs into
42 Chapter 5. Discussion this application would be generating abc notation through natural language queries, turning this notation into MIDI and condition the anticipatory transformer on this MIDI. The researcher Dimitry Bogdanov proposed using embeddings from both composers corpus of works (these could be imported in the app, though privacy issues should be taken into account) so the inpainting would be more aligned to a mix of their style, resulting in a possible improvement of consensus. 5.1.3 Limitations Ideally, this application should be tested on more participants (now N=16), which was not possible due to the nature of the requirements set: •Professional composers •Proficient in score reading and editing •Access to a Midi keyboard (physical or digital) This N size is why some metric values should not be taken as a trend and generalization of the results cannot be taken for granted. The reason behind this was to focus on a specific population and perform long and exhaustive experiments where the author could be present for the whole time annotating feedback to extract developed insights and not short Yes/No answers. Thus, due to time constraints, this could not be done with a higher number of participants. Time constraints also lead to a discard of other quantitative analysis as analyzing the MIDI files generated by the users (by analyzing repetitiveness, complexity, etc.) and then comparing AI and no-AI files in a focus group to see if there are perceptual differences. The limited computational resources of the API also led to using the small version of the Anticipatory Transformer model on the experiments, limiting the quality of the output (which can even be higher).
5.2. Conclusions 43 5.2 Conclusions This thesis addressed a critical gap in the intersection of artificial intelligence and collaborative creative practices by designing, implementing, and evaluating InScoreAI, a web-based platform integrating anticipatory transformers for real-time, AI-assisted collaborative music composition. Our work bridges the domains of Computer Supported Collaborative Work (CSCW) and Human-AI co-creation, proposing a novel paradigm: AI-Supported Collaborative Work (AISCW) for music. Through empirical studies with professional composers, we address our core research questions and show the potential (and challenges) of AI mediation in creative workflows. AI significantly reduced compositional effort and enhanced self-efficacy, validating AI’s role as a powerful technical assistant that accelerates workflow and overcomes creative blocks. However, this comes at the cost of a significant reduction in perceived ownership, raising concerns about over-reliance and “soulless” outputs among professional composers. AI serves as an effective "social glue", demonstrably mitigating interpersonal friction during clashes of ideas. High median scores for Fluidity, Consensus, and Collaboration confirm that AI-generated "idea bridges" facilitate mutual understanding and efficient compromise. Although ownership remained low and comparable to individual AI use, the collaborative context fostered a marginally higher sense of shared ownership compared to solitary AI assistance, suggesting that collaboration partially mitigates the lack of sense of ownership. We successfully extended anticipatory transformers with multi-track capabilities and embedded it within a collaborative web editor. Key technical contributions include compound instrument tokens preserving multi-track structure during generation, metadata recovery (tempo, time signatures) ensuring rhythmic/metrical integrity, and three “steering” agents (Harmonization, Inpainting and Melody-Changing) offering flexible, user-controllable AI interventions conditioned on melody, accompaniment or future context. This provides a robust solution to the "lack of control" problem prevalent in autonomous AI music generators.
44 Chapter 5. Discussion InScoreAI effectively addresses the “lack of consensus” challenge identified in CSCW music research. By generating multiple transitional options (“idea bridges”) between conflicting human contributions and allowing real-time, independent audition/selection/rejection of proposals where AI acts as a neutral mediator. This transforms creative disagreements into collaborative problem-solving exercises focused on evaluating AI suggestions, thereby reducing interpersonal tension and accelerating progress toward shared goals. The system directly addresses the high-priority features identified by professional composers: real-time collaboration (90% high interest), AI assistance for technical tasks (Harmonization and Melodic Development, 65%), and the requested "idea bridging" functionality. However, qualitative feedback by professional composers raised concerns about AI potentially fostering creative laziness, simplifying musical output, and diminishing the perceived uniqueness and experimental nature of compositions. Qualitative insights highlight the significant educational value of AIassisted collaborative platforms, particularly for novice composers and students. Benefits include overcoming initial technical hurdles (e.g., harmonization), visualizing compositional possibilities, and motivating engagement by transforming exercises into fuller musical pieces within a supportive collaborative environment. This thesis demonstrates that anticipatory transformers, integrated as steerable partners within a real-time collaborative environment, offer a powerful paradigm (AISCW) for transforming music composition. InScoreAI successfully reduces effort, enhances self-efficacy, and resolves interpersonal friction, making collaborative composition more fluid and accessible. However, the core tension between AI’s efficiency gains and the dilution of perceived ownership/creative identity remains a fundamental challenge for the field. Future work must prioritize interface and AI designs that actively preserve human agency, foster stylistic diversity, and enhance transparency to address these concerns, particularly for professional users. The deployment of InScoreAI https://inscoreai.netlify.app/ and planned opensourcing of its code provide a foundation for further exploration of AI’s evolving role in human collaboration within music and other fields.
Bibliography [1] Glover, R. & Redhead, L. Collaborative and Distributed Processes in Contemporary Music-Making (Cambridge Scholars Publishing, 2020). [2] Gresham-Lancaster, S. The aesthetics and history of the hub: The effects of changing technology on network computer music. Leonardo Music Journal 8, 39–44 (1998). [3] Barbosa, A. Computer-suported cooperative work for music applications. Ph.D. thesis, Universitat Pompeu Fabra (2006). [4] Blaine, T. & Perkis, T. The Jam-O-Drum interactive music system: A study in interaction design. In Proceedings of the 3rd Conference on Designing Interactive Systems: Processes, Practices, Methods, and Techniques, 165–173 (ACM, New York City New York USA, 2000). [5] Kaltenbrunner, M., Jorda, S., Geiger, G. & Alonso, M. The reacTable*: A Collaborative Musical Instrument. In 15th IEEE International Workshops on Enabling Technologies: Infrastructure for Collaborative Enterprises (WETICE’06), 406–411 (IEEE, Manchester, UK, 2006). [6] Laney, R. et al. Issues and techniques for collaborative music making on multitouch surfaces. In 7th Sound and Music Computing Conference (Barcelona, 2010). [7] Xambó, A., Laney, R. & Dobbyn, C. TOUCHtr4ck: Democratic collaborative music. In TEI ’11 Fifth International Conference on Tangible, Embedded and Embodied Interaction (Funchal, Portugal, 2011). 45
46 BIBLIOGRAPHY [8] Jorda, S. & Wust, O. A system for collaborative music composition over the web. In 12th International Workshop on Database and Expert Systems Applications, 537–542 (2001). [9] Men, L. Exploring Collaborative Music Making Experience in Shared Virtual Environments. Thesis, Queen Mary University of London (2020). [10] Vlachakis, G., Kalaentzis, A. & Akoumianakis, D. Collaborative music composition as virtual work across boundaries. In 2014 International Conference on Telecommunications and Multimedia (TEMU), 202–207 (2014). [11] Tipei, S., Craig, A. B. & Rodriguez, P. F. Using High-Performance Computers to Enable Collaborative and Interactive Composition with DISSCO. Multimodal Technologies and Interaction 5, 24 (2021). [12] Bryan-Kinns, N., Healey, P. G. T. & Leach, J. Exploring mutual engagement in creative collaborations. In Proceedings of the 6th ACM SIGCHI Conference on Creativity & Cognition, C&C ’07, 223–232 (Association for Computing Machinery, New York, NY, USA, 2007). [13] Bryan-Kinns, N. Mutual engagement and collocation with shared representations. International Journal of Human-Computer Studies 71, 76–90 (2013). [14] Hart, A. & Williams, A. A Space for Making: Collaborative composition as social participation. Organised Sound 26, 240–254 (2021). [15] Biasutti, M. Assessing a collaborative online environment for music composition. Educational Technology & Society 18, 49–64 (2015). [16] Cherry, E. & Latulipe, C. The creativity support index. 4009–4014 (2009). [17] Burkhardt, J.-M. et al. An approach to assess the quality of collaboration in technology-mediated design situations. In European Conference on Cognitive Ergonomics: Designing beyond the Product — Understanding Activity and User Experience in Ubiquitous Environments, ECCE ’09 (VTT Technical Research Centre of Finland, FI-02044 VTT, FIN, 2009).
BIBLIOGRAPHY 47 [18] MuseScore. Muselab plugin. https://musescore.org/en/project/ muselab-real-time-collaboration [Accessed: 31-3-25]. [19] Zhu, Y., Baca, J., Rekabdar, B. & Rawassizadeh, R. A Survey of AI Music Generation Tools and Models (2023). 2308.12982. [20] Wang, L. et al. A review of intelligent music generation systems. Neural Computing and Applications 36, 6381–6401 (2024). [21] Ji, S., Yang, X. & Luo, J. A Survey on Deep Learning for Symbolic Music Generation: Representations, Algorithms, Evaluations, and Challenges. ACM Comput. Surv. 56, 7:1–7:39 (2023). [22] Ippolito, D., Huang, A., Hawthorne, C. & Eck, D. Infilling piano performances. In NIPS Workshop on Machine Learning for Creativity and Design, vol. 2 (2018). [23] Pati, A., Lerch, A. & Hadjeres, G. Learning to traverse latent spaces for musical score inpainting. CoRR abs/1907.01164 (2019). URL http://arxiv.org/ abs/1907.01164.1907.01164. [24] Chen, K., Wang, C.-i., Berg-Kirkpatrick, T. & Dubnov, S. Music sketchnet: Controllable music generation via factorized representations of pitch and rhythm. arXiv preprint arXiv:2008.01291 (2020). [25] Chang, C.-J., Lee, C.-Y. & Yang, Y.-H. Variable-length music score infilling via xlnet and musically specialized positional encoding. arXiv preprint arXiv:2108.05064 (2021). [26] Guo, R., Simpson, I., Kiefer, C., Magnusson, T. & Herremans, D. Musiac: An extensible generative framework for music infilling applications with multi-level control. In International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar), 341–356 (Springer, 2022). [27] Thickstun, J., Hall, D., Donahue, C. & Liang, P. Anticipatory Music Transformer (2024). 2306.08620.