scieee AI-readable full text Open interactive document viewer

Filling MIDI Velocity using U-Net Image Colorizer

He, Zhanhong; Cooper, David; Huang, Defeng; Togneri, Roberto

Abstract

Modern music producers commonly use MIDI (Musical Instrument Digital Interface) to store their musical compositions. However, MIDI files created with digital software may lack the expressive characteristics of human performances, essentially leaving the velocity parameter—a control for note loudness—undefined, which defaults to a flat value. The task of filling MIDI velocity is termed MIDI velocity prediction, which uses regression models to enhance music expressiveness by adjusting only this parameter. In this paper, we introduce the U-Net, a widely adopted architecture in image colorization, to this task. By conceptualizing MIDI data as images, we adopt window attention and develop a custom loss function to address the sparsity of MIDI-converted images. Current dataset availability restricts our experiments to piano data. Evaluated on the MAESTRO v3 and SMD datasets, our proposed method for filling MIDI velocity outperforms previous approaches in both quantitative metrics and qualitative listening tests.

Full text

Filling MIDI Velocity using U-Net Image Colorizer Zhanhong He1,2[0000000289408437], David Cooper2[0009000898058943], Defeng Huang1[0000000214318859],andRobertoTogneri 1[0000000237784633] 1University of Western Australia, Perth WA 6000, Australia 2Dolby Laboratories, Sydney NSW 2000, Australia [email protected], [email protected], {david.huang, roberto.togneri}@uwa.edu.au Abstract. Modern music producers commonly use MIDI (Musical Instrument Digital Interface) to store their musical compositions. However, MIDI files created with digital software may lack the expressive characteristics of human performances, essentially leaving the velocity parameter—a control for note loudness—undefined, which defaults to aflatvalue.ThetaskoffillingMIDIvelocityistermedMIDIvelocity prediction, which uses regression models to enhance music expressiveness by adjusting only this parameter. In this paper, we introduce the U-Net, awidelyadoptedarchitectureinimagecolorization,tothistask.By conceptualizing MIDI data as images, we adopt window attention and develop a custom loss function to address the sparsity of MIDI-converted images. Current dataset availability restricts our experiments to piano data. Evaluated on the MAESTRO v3 and SMD datasets, our proposed method for filling MIDI velocity outperforms previous approaches in both quantitative metrics and qualitative listening tests. Keywords: MIDI velocity prediction ·U-Net ·Image colorization ·Music expressiveness. 1Introduction MIDI (Musical Instrument Digital Interface), acting as digital sheet music playable by machines and software, is the dominant format in modern music production. A MIDI file resembling sheet music sounds mechanical due to the undefined velocity parameter, which defaults to a flat value. In contrast, as shown in Figure 1, MIDI files recorded from human performances capture performer skills through subtle timing and loudness variations, which infuse expressiveness [1]. Today, music producers are not always masterful in playing musical instruments [2,3], and low cost MIDI keyboards may lack sophisticated touch-sensitive sensors. This leads to a demand for automated systems designed to enhance the expressiveness of MIDI compositions. All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 949 Z. He et al. Enhancing the expressiveness of existing MIDI files is one important focus in music generation [4]. However, many of these systems modify multiple aspects of MIDI simultaneously [5 – 10], including note timing and loudness, and sometimes note quantity, which can introduce unwanted alterations. The loudness of each music note in a MIDI file is governed by a parameter called MIDI velocity. Studies [11] have shown that rendering only the MIDI velocity enhances expressiveness while preserving the original timing, making it a precise and controllable solution. This task of filling or rendering MIDI velocity has been treated as a sequential prediction problem by previous studies, which employed autoencoders [12] and sequential models [13]. Fig. 1. Comparison between MIDI notes with human performed velocity versus Music Software default velocity (64 if user not specified). The standard deviation of velocity (SDvelo)representsthedispersionofvelocitiesaroundtheirmeanacrossthepitches. Inspired by image colorization [14], we reframe MIDI velocity prediction as an image colorization problem, representing MIDI without velocity as a binary pianoroll and target velocity as a colored pianoroll. Image-based methods suit this case well, as they effectively capture the polyphonic structure of instruments such as the piano and guitar, which produce multiple simultaneous notes. While our work focuses on piano data, where well-annotated velocity datasets are most common, the universal nature of MIDI velocity makes cross-instrument generalization a promising direction for future research [15]. In this paper, we introduce the U-Net architecture to MIDI velocity prediction, leveraging its success in image colorization, and incorporate window attention to handle the sparsity of MIDI data. In addition, we design a custom loss function Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 950 Filling MIDI Velocity tailored to our task. The resulting model is evaluated through both objective quantitative metrics and qualitative assessments via a subjective listening test. 2RelatedWorks How to render a score to be more expressive (i.e., a performance-like MIDI) has been a long-standing topic in music research [4]. A central goal of this research, MIDI velocity prediction, has been to independently modify MIDI velocity. Early efforts to this problem involved linear basis models [16] and restricted Boltzmann machines [17]. More recently, Kuo et al. [12] implemented a convolutional autoencoder (ConvAE), while a Seq2Seq model [13] reported the best results by integrating Luong attention into a BiLSTM. While recent methods have treated MIDI as a sequence, those sequential models prioritize global features over local details [18], potentially affecting the scattered distribution of velocities. Inspired by image colorization, where precise grayscale images are overlaid with blurred color predictions [19], MIDI velocity prediction can be approached similarly by leveraging given MIDI notes. This suggests that U-Net, a widely used architecture in image colorization [20], can be effective for MIDI velocity prediction. U-Net also dominates image segmentation and is frequently combined with self-attention mechanisms [21,22]. Both U-Net and attention mechanisms have shown success in music information retrieval (MIR) research. While U-Net has been effective in automatic music transcription (AMT) [23,24], the attention mechanism has been used to refine velocity estimation from performance audio [25]. Since our task is MIDI-only, their audio-dependent approach is not applicable. 3Methods 3.1 Matrix Representation To process MIDI as images, we convert the MIDI into a three matrices with T⇥P dimension, where T is the number of time frames and P = 88 is the number of pitches. The three matrices include a binary onset roll O marking note starts, a binary frame roll F indicating note activation over time, and a velocity roll V acting as color intensity. For the velocity roll, integer values [0,127] are normalized to the range [0,1) to align with the model’s output activation layer. The final integer velocities will denormalize by scaling and rounding the model’s output. As shown in Figure 2, these matrices are highly sparse. 3.2 MIDI Segmentation The MIDI segment duration is a key consideration in our approach. Since a MIDI file often exceeds 3 minutes in length, we split it into short segments to Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 951 Z. He et al. manage computational load. The number of time frames T defines the temporal granularity of the MIDI-converted matrix, so the timestep resolution is given by: Resolution =Segment Duration T(1) Unlike tasks that require high timestep resolution for precise event detection, our task leverages known MIDI timings, making a computationally efficient timestep resolution feasible. With a fix input size T = 96 to keep size affordable, we experimented the segment durations of 5s, 10s, 15s, 20s and found that 10 seconds yielded the best results, as detailed in our hyperparameter search. 3 A 10-second segment likely provides superior semantic context by encapsulating a complete musical phrase (e.g., four measures at 120 BPM) compared to other durations we tested. Fig. 2. Proposed U-Net architecture. Model input ( F roll) comprises 88 pitch bins and 96 time frames. Attn block denotes the windowed scaled dot-product attention. The final velocity roll ( V roll) is generated during post-processing by extracting velocity at note positions, and then assigning each note the velocity at its onset. 3.3 Model Architecture The proposed architecture is shown in Figure 2. The U-Net extracts higher-level features through downsampling, with skip connections preserving details and global patterns. All convolution blocks use 3 ⇥ 3kernels, stride 1, and padding 1, followed by sigmoid activation and batch normalization. To scale features by a factor of 2, we perform downsampling with a standard 2 ⇥ 2MaxPool2D layer; for upsampling, we employ a transpose convolution with 4 ⇥ 4kernels and stride 2, a distinctive strategy introduced in [20]. 3 wandb report has concluded our experiment history of hyperparameter searching, available at: https://api.wandb.ai/links/zhanh-uwa/wpzvcb76 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 952 Filling MIDI Velocity To handle MIDI data sparsity, we integrate the windowed scaled dot-product attention (window attention) [26] into our U-Net. Traditional self-attention [27] operates on the full feature map X2RH⇥W , where H and W denote the height and width of the attention inputs, as depicted in Figure 2. Window attention partitions X into n⇥n non-overlapping windows and computes attention within each. This reduces computational complexity while enhancing feature aggregation [28]. The window size ( n )isatunablehyperparameterbalancinglocalandlongrange dependencies. We explored n2{ 1 , 2 , 4 , 8 } and found that a 2 ⇥ 2window attention yielded the best performance (see wandb report 3 ). This suggests that while window attention is effective, larger windows can over-compress information and reduce effectiveness. 3.4 Loss Function The proposed loss function combines binary cross-entropy (BCE) loss with cosine similarity (CosSim) introduced in [13], with ↵=0.2, defined as: LCombine =(1↵)LBCE +↵(1 CosSim) (2) where the BCE loss is used to optimize the prediction error; CosSim is computed for each pitch and then averaged, capturing the trending of velocity changes over time: CosSim =1 P P X p=1 PT t=1 yt,p ˆyt,p qPT t=1 y2 t,p qPT t=1 ˆy2 t,p (3) LBCE =1 TP T X t=1 P X p=1 lbce(yt,p,ˆyt,p)(4) here, CosSim and BCE are functions pre-built in PyTorch, with yt,p and ˆyt,p denoting the target and predicted velocity, respectively. The indices t and p represent the time and pitch dimensions. To deal with the sparsity, we apply amaskingoperation< m >usingtheonsetroll O , which ignores silent time steps and counts each note once at its onset. In addition, a weighting operation < w >isintroducedtoreducetheboundaryvelocitypredictionerrorsemphasized in [25,29,30]. Motivated by the Gaussian distribution of velocity observed in Figure 3, we design a V-shaped weighting centered on 64 (normalized to 0.5), with an empirical factor of 3 to enhance regions away from the midpoint. The updated BCE loss with masking and weighting is defined as: L<m,w> BCE =1 TP T X t=1 P X p=1 Ot,p ·1+3|Vt,p 0.5|·lbce(yt,p,ˆyt,p),(5) in which Ot,p is the onset-roll mask and (1 + 3 |Vt,p  0 . 5 | )is a weighting factor based on velocity roll V .Finally, LBCE in Eqn (2) is replaced with L<m,w> BCE to form L<m,w> Combine . The effectiveness of this loss was validated in our wandb report. 3 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 953 Z. He et al. Fig. 3. MIDI data distributions of the MAESTRO (blue) and SMD (orange) datasets, with density maps highlighting the similarity in their MIDI feature correlations. 4 Experiment 4.1 Dataset For training, we use the MAESTRO v3.0.0 dataset [31], recorded by skilled pianists on Yamaha Disklavier pianos during the International Piano-e-Competition. The default train/valid/test split is used. With 1,276 performances totaling over 200 hours, MAESTRO provides an ideal foundation for modeling MIDI velocity. For evaluation, we use the Saarland Music Data (SMD) dataset [32], which comprises 50 performances also recorded on a Yamaha Disklavier. We selected SMD for cross-dataset evaluation to assess model generalization, instead of the Piano-e-Competition dataset used in [12,13] which has significant performance overlap with MAESTRO. The suitability of SMD is confirmed in Figure 3. Furthermore, SMD is used for qualitative assessment through a subjective listening test of 8 selected performances, as listed in Table 1. As SMD only has composerstyle overlaps with MAESTRO (none of performance overlap), it allows us to evaluate our model on both seen and unseen compositional styles. Table 1. Selected SMD performances for the subjective listening test, with SMD composer statistics and their overlap with the MAESTRO train set. MAESTRO train set SMD dataset Composer Total Perf. Total Dur. Total Perf. Selected Perf. Chopin 145 19.9 h 13 Op010-04 Bach 114 11.2 h 8BWV849-02 Beethov. 110 20.5 h 7Op027No1-01 Liszt 93 16.0 h 3 - Schuman. 33 12.4 h 3 - Rachman. 29 4.2 h 3Op036-02 Haydn 29 3.7 h 4Hob017No4 Mozart 27 3.9 h 2 KV265 Scriabin 22 4.1 h 1 - Brahms 20 6.1 h 3 - Bartok 00h 3SZ080-03 Ravel 00h 2 JeuxDEau Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 954 Filling MIDI Velocity 4.2 Training Setup The models we trained include the proposed U-Net and a re-implemented ConvAE [12]. Following the training strategy of [33], we arranged continuous segments of a song into the same batch to preserve the musical structure, thereby aiding semantic learning. Both models were trained for 300 epochs on the MAESTRO train set with a learning rate of 1e-5, a batch size of 3, and the same loss function. Training took approximately 12 hours on an NVIDIA V100 32GiB GPU using the Ranger21 optimizer. The top three checkpoints of each model, selected based on the MAESTRO validation set performance, were tested on the MAESTRO test set and the SMD dataset for cross-dataset evaluation. 4.3 Evaluation Metrics All objective evaluation metrics are computed on the denormalized MIDI velocity, restored to the original scale of 0 to 127. We adopt the mean absolute error (MAE), mean square error (MSE), and standard deviation of velocity ( SDvelo ), which are standard metrics in MIDI velocity prediction [12,13]. We also incorporate the standard deviation of absolute error ( SDae )andRecall,bothprevalentin similar research [25,34]. The Recall uses a standard 10% error tolerance. The subjective listening test follows the mean opinion score (MOS) of MUSHRA [35] framework. Participants rated the expressiveness of MIDI-generated audio on a 100-point scale, mapped to values from 0 to 5, with five labeled intervals (from "bad" to "excellent") for ease of use. 5ResultsandDiscussion 5.1 Quantitative Results Tables 2 and 3 present the model performance, with all models trained exclusively on the MAESTRO train set. The Flat model assigns a fixed velocity of 64, representing default music software behavior. The Seq2Seq model uses pretrained weights from [13], while ConvAE [12] is re-implemented and trained in our framework. Both tables demonstrated that the proposed U-Net outperformed other models across all objective metrics. Table 2. Quantitative results on the SMD dataset, where " and # indicate whether higher or lower values are better. Model MAE #MSE #SDae #Recall "SDvelo " Flat (all velocities set to 64) 15.3 367.5 10.4 49.2% 0 Seq2Seq [13] 15.1 356.9 10.9 48.8% 8.5 ConvAE [12] (re-implemented) 12.5 258.1 9.6 58.5% 9.8 U-Net (proposed) 11.2 217.5 9.2 65.1% 11.1 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 955 Z. He et al. Table 3. Quantitative results on the MAESTRO test set, where " and # indicate whether higher or lower values are better. Model MAE #MSE #SDae #Recall "SDvelo " Flat (all velocities set to 64) 14.8 333.8 10.2 48.5% 0 Seq2Seq [13] 13.5 286.1 9.9 53.8% 6.2 ConvAE [12] (re-implemented) 12.3 250.7 9.7 59.4% 9.7 U-Net (proposed) 11.5 226.2 9.4 63.8% 10.7 The validity of evaluation metrics warrants further discussion. When visualizing the results, as demonstrated in Figure 4, we found: (1) Accuracy metrics (including MAE, MSE, SDae ,andRecall)usuallyshowsignificantdifferences, but not in certain cases (e.g., the Chopin Op10-04 segment), suggesting an artifact caused by a local optimum when mid-value predictions dominate. (2) SDvelo is a key differentiator that effectively reflects human-likeness. By capturing the dispersion of MIDI velocity around its mean, higher values reveal a greater clarity between the left and right hands, potentially enhancing expressiveness. Fig. 4. Comparison of human-performed velocity with ConvAE and U-Net predictions. The U-Net result is more human-like than ConvAE, but for Chopin Op. 10 No. 4, accuracy metrics fail to reflect this. 5.2 Qualitative Results through Listening Test AMUSHRA-likelisteningtestwasconductedtoevaluatethehuman-likenessof performances generated by our U-Net, ConvAE [12], and the Flat model. Participants. We recruited 11 expert listeners (aged  18)basedontheir substantial musical background and experience with critical listening studies. Participants completed survey anonymously through the Qualtrics platform [36]. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 956 Filling MIDI Velocity Stimuli. The test involved eight 10s MIDI segments, each selected from a performance listed in Table 1, with human-performed velocities removed for uniform model inputs. Outputs from Flat, ConvAE, and U-Net were rendered as stimuli, while the original MIDI (with human-performed velocity) served as a reference, all using the PianoTeq 8 plug-in with the Steinway Model D instrument. Calibration. The dataset used for the human-performed MIDI also contained corresponding recordings of the Yamaha Disklavier used in the performance. This is the audio that the pianist would have heard at the time of the performance. To most accurately render the MIDI, we calibrated the piano modeling software, PianoTeq 8 [37], to match those Disklavier recordings. This was achieved by adjusting the "Dynamics" control in PianoTeq 8 and comparing the Momentary Loudness via correlation coefficient and Loudness Range as described in BS.1770/EBU R128, against the reference. The results of this process presented in Table 4 indicated that a Dynamics setting of 60 dB most closely matched the audio of the original performance. This setting was used to process all files used for the test. Table 4. Metrics for comparing different PianoTeq configurations to the reference. Dynamics (dB) Loudness #Correlation Coefficient " 50 1.8 0.9583 60 0.3 0.9591 70 0.9 0.9561 80 1.9 0.9535 90 3.1 0.9479 Procedure. Participants were instructed to use headphones for accurate evaluation, as subtle velocity differences required higher playback volumes for clarity. They rated the similarity between audio rendered from model-predicted and human-performed MIDI, using a continuous slider from 0 to 5, and results were aggregated into MOS scores in Table 5. Table 5. Results of subjective listening test. MOS with 95% confidence interval are reported, where "most seen" includes {Chopin, Bach, Beethoven}, "less seen" includes {Rachmaninoff, Haydn, Mozart}, and "unseen" includes {Bartók, Ravel}. Model Most Seen MOS "Less Seen MOS "Unseen MOS "Overall MOS " Flat 1.58 ±0.36 1.64 ±0.44 1.18 ±0.42 1.50 ±0.37 ConvAE 1.93 ±0.34 2.40 ±0.39 1.80 ±0.44 2.08 ±0.33 U-Net 3.10 ±0.38 3.16 ±0.29 2.67 ±0.46 3.01 ±0.34 Results. Although a gap to human performance remains (no model achieved aperfectsimilarityscoreof5),ourU-Netmodelsignificantlyoutperformedother approaches. The "Flat" model performed worst, aligning with the survey in previous work [12]. A key limitation in this survey was distinguishing whether Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 957