Full text
Master Thesis on Sound and Music Computing Universitat Pompeu Fabra MuSA: A New TEL Platform for Enhancing Self-Reflection and Musical Understanding through Saliency Analysis of Performance Recordings Isabelle Oktay Supervisor: Rafael Ramirez-Melendez Co-Supervisor: Suvi Haeaerae July 2025
Contents 1 Introduction 1 1.1 Motivation.................................. 2 1.2 Objectives.................................. 5 1.3 CodeAccessibility ............................. 7 2 Background 8 2.1 Foundations of Musical Talent Development . . . . . . . . . . . . . . . 8 2.1.1 TheTADMusicModel........................... 9 2.2 Practice Strategies in Music Education . . . . . . . . . . . . . . . . . . 12 2.2.1 The Master-Apprentice Model . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.2 Constructivism and Dialogic Teaching . . . . . . . . . . . . . . . . . . 13 2.2.3 Scaffolding.................................. 14 2.2.4 Self-Regulated Learning . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.2.5 Self-Directed Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.3 Feedback................................... 19 2.3.1 CognitiveLoad ............................... 20 2.3.2 Modeling .................................. 22 2.3.3 SalientMoments .............................. 23 2.3.4 Synthesizing SRL, SDL, Feedback, and MuSA . . . . . . . . . . . . . . 23 2.4 Technological Pedagogical Content Knowledge . . . . . . . . . . . . . . 24 2.5 TEL Feedback Systems in Music Education . . . . . . . . . . . . . . . 25 2.5.1 OnlineResources .............................. 27
2.5.2 Non-Real-Time Recording Feedback . . . . . . . . . . . . . . . . . . . . 29 2.5.3 Real-Time Feedback . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 2.6 Analyzing Salient Performance Moments . . . . . . . . . . . . . . . . . 32 2.7 Background Research Summary . . . . . . . . . . . . . . . . . . . . . . 33 3 Designing and Implementing MuSA 34 3.1 InitialPrototype .............................. 34 3.2 Implementing MuSA-V2 . . . . . . . . . . . . . . . . . . . . . . . . . . 36 3.2.1 Selecting Musical Features for MuSA-V2 . . . . . . . . . . . . . . . . . 38 3.2.2 DesigningMuSA .............................. 42 3.3 MuSA System Architecture . . . . . . . . . . . . . . . . . . . . . . . . 43 4 MuSA Evaluation Study 49 4.1 Methods................................... 50 4.2 Results.................................... 53 4.2.1 Objective Performance Metrics . . . . . . . . . . . . . . . . . . . . . . 54 4.2.2 Perceived Performance Metrics . . . . . . . . . . . . . . . . . . . . . . 59 4.2.3 Self-Awareness Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 4.2.4 SentimentAnalysis............................. 65 5 Discussion 69 5.1 Discussion.................................. 69 5.1.1 Integrating Theory into Practice . . . . . . . . . . . . . . . . . . . . . . 69 5.1.2 Evaluating MuSA: Demand, Limitations, and Outcomes . . . . . . . . 71 5.1.3 Independent Practice and Learner Implications . . . . . . . . . . . . . 74 5.1.4 MuSA’s Limitations and Future Directions . . . . . . . . . . . . . . . . 75 5.2 Conclusions ................................. 80 List of Figures 82 List of Tables 86
Bibliography 88 A Appendix: MuSA Evaluation Study Results 114
Acknowledgement This project would not have been possible without the guidance and support of my supervisor, Rafael Ramirez-Melendez. He introduced me to SkyNote and allowed me to work on updating the SkyNote web application, an offshoot of the original TELMI [1] project. I enjoyed working on SkyNote so much that when Rafael proposed the project on saliency analysis and creating a non-real-time music performance feedback application, I was thrilled and immediately chose it as my thesis topic. Thank you, Rafael, for all your guidance and insight throughout this project—I have learned immensely, and I hope that someone will continue the next iteration of MuSA. I am also deeply grateful to my co-supervisor, Suvi Häärä, who studied alongside me during the Sound and Music Computing master’s program at Universitat Pompeu Fabra (2023–2024). Suvi has supported me on numerous projects, including this thesis, and has consistently been a positive influence, helping me stay focused and motivated. Thank you, Suvi, for all your encouragement and hard work—your support has been invaluable. Thank you to Anmol Mishra, who assisted me during a critical moment in deploying the application. Having never deployed anything before, I was quite lost, but Anmol patiently guided me through multiple calls, even while in India, to help solve the problem. Your help was essential—thank you, Anmol! I also owe immense gratitude to my mother, who was always there when I doubted myself or felt overwhelmed by this project. Her guidance, encouragement, and wisdom, particularly as a physician and academic, kept me motivated and grounded throughout this journey. Finally, thank you to my partner, Théo Fuhrmann, for unwavering support, for cooking meals during marathon coding sessions, and for helping me maintain perspective during long days of work. Your companionship, patience, and willingness to troubleshoot alongside me, sometimes as my rubber duck, kept my spirits high and made this project possible.
Abstract This thesis explores how a technology-enhanced learning (TEL) tool can improve music practice by addressing a critical, often-neglected component of skill development: the reflection phase. It focuses on the development and evaluation of MuSA (Musical Salience Analyzer), an application designed to provide a pedagogicallygrounded platform for analyzing recorded performances to make reflection more efficient and effective. MuSA’s design is informed by key educational theories, including the Talent-Development-in-Achievement-Domains (TAD) Music Model and learner-centered teaching (LCT) principles like scaffolding, self-regulated learning (SRL), and self-directed learning (SDL). Its central feature is saliency analysis, which algorithmically identifies key moments in a performance based on variability in musical features such as pitch, dynamics, and tempo. Unlike tools that offer prescriptive, "correct/incorrect" feedback, MuSA encourages a learner’s own interpretation. As an accessible, web-based platform, it allows users to upload or record audio for analysis independently of a teacher. To evaluate MuSA’s effectiveness, a mixed-methods, within-subjects study was conducted with 14 participants. While the study’s small sample size limited statistical power, the findings pointed to several exploratory trends. The data suggested a differential impact based on musical feature and experience level, with dynamics showing the most consistent trend toward objective and perceived improvement. The analysis also suggested a potential expertise reversal effect, where trends showed intermediate musicians gaining from the feedback while advanced musicians experienced neutral or slightly negative changes. Furthermore, the study’s self-awareness metrics indicated a general misalignment between participants’ self-ratings and objective performance, highlighting a core challenge in the self-reflection phase of independent practice. In conclusion, MuSA offers a potential contribution to TEL for music by leveraging computational analysis to provide targeted insights that can scaffold the reflective process. Although the quantitative results were inconclusive, positive qualitative feedback validates the demand for such a tool. This work provides a functional prototype and a research infrastructure for collecting labeled recording data, demonstrating the dual role of
6Chapter 1. Introduction of musical talent development and how students progress through them. Understanding these stages allows for a theoretical evaluation of how MuSA can support learners as they develop musical talent. 2. Compile a background of common practice strategies and learning frameworks used in both independent and traditional master-apprentice settings. Doing so informs how MuSA can be designed to integrate smoothly into existing practice routines and facilitate performance reflection. 3. Emphasize musical expression as a core dimension of musical talent to guide MuSA’s development toward providing flexible, personalized feedback (fostering expressive skills) rather than solely binary “correct”/“incorrect” judgments (training technical skills). 4. Evaluate existing TEL platforms for music practice and performance analysis, focusing on their capacity for personalized feedback and support for learning musical expression. This includes reviewing online learning strategies and assessing both real-time and non-real-time feedback applications (especially performance analysis tools) to identify gaps in TEL music platforms. To evaluate the pedagogical value of MuSA, a mixed-methods study was conducted. The study examined whether detecting salient moments and providing visual feedback affected musicians’ self-rated and objective performance (i.e., reflection) on features like pitch, dynamics, and tempo. The key contributions of this thesis include: 1. The design and implementation of MuSA, a TEL platform that analyzes salient moments of a music performance based on different musical aspects and can be used outside the traditional master-apprentice model to facilitate performance reflection. 2. Empirical findings from a mixed-methods user study that may inform future integrations of performance analysis technologies into independent music learning, specifically with regards to self-awareness during music performance. 3. A direction for future explorations that can further develop the MuSA
1.3. Code Accessibility 7 web application and explore how presenting salient moments to learners may not only support performance correction relative to a reference audio but also enhance musical understanding and reflection during practice through the visualization of performance recordings. The thesis is structured as follows: Chapter 1 introduces the research context and motivation; Chapter 2 reviews relevant literature and the state of the art; Chapter 3 details the design and development of MuSA; Chapter 4 outlines the evaluation study done on MuSA; and Chapter 5 discusses implications and future work. 1.3 Code Accessibility All code is accessible through Github6. Both MuSA7and the MuSA evaluation study8can be accessed online. 6https://github.com/isabelleoktay/audio-analyzer/ 7https://analyzer.appskynote.com/ 8https://analyzer.appskynote.com/testing/
Chapter 2 Background 2.1 Foundations of Musical Talent Development Musical talent is often perceived as innate, yet it more accurately emerges from the interaction between individual predispositions and cultivated skills [3]. Recognizing this interplay is essential for explaining why some musicians excel and for identifying how technology-enhanced learning (TEL) can best support their development. This perspective not only justifies the creation of MuSA but also clarifies the pedagogical principles that shape its core functionality. Some elements of musical talent are less easily influenced. Intelligence, for example, is largely determined by genetics [69] and comprises a range of cognitive abilities such as perceptual speed [70] and reasoning ability [71]. Individuals also differ in their auditory sensory discrimination even before receiving formal musical education, a capability known as musical aptitude [3]. Environmental factors can also predispose talent, as children whose abilities are recognized early by parents are more likely to receive enriched musical opportunities [72]. Other aspects are more directly shaped through intentional effort. Practice, for instance, is widely regarded as a pillar of talent development, as no one is born knowing how to read sheet music, play scales, or sight-sing. Early engagement with music further develops technical skills and fosters personality traits such as willpower and motivation, creating a 8
2.1. Foundations of Musical Talent Development 9 foundation for the continual accrual of musical knowledge [73, 74]. However, the relative impacts of these factors on developing musical talent remain debated, with practice being a focal point of contention. In this discussion, the term practice refers specifically to deliberate practice, which is understood as a set of activities typically designed by an external source, often a teacher, with the explicit goal of improving performance [6, 2, 7, 75]. This definition provides a narrower, more operational measure of practice, aligning with much of the literature against which competing influences (e.g., intelligence, musical aptitude) have been evaluated [4, 5]. Research shows that accumulated practice hours are the strongest predictor of technical skills like sight-reading [6, 7]. Yet meta-analyses estimate practice explains only 23% to 37% of the variance in overall music achievement [4, 5]. Critics of these findings cite methodological issues, including inconsistent definitions of deliberate practice [76]. Even so, genetic factors such as intelligence may surpass practice in predicting musical development—especially in beginners—and may predict sight-reading performance after controlling for practice [77, 78, 7]. Environmental influences also interact with genetic ones, as musically inclined parents may both foster enriched environments and pass on traits predisposing children to formal training. Such inherited tendencies to select and shape environments, known as gene–environment transactions [69], suggest that musical experiences may themselves be partly heritable. 2.1.1 The TAD Music Model While practice is vital for developing instrumental technique and personal skills, it operates alongside—and in constant interaction with—genetic, cognitive, and environmental influences. To situate these factors within a coherent developmental trajectory, this thesis adopts the Talent-Development-in-Achievement-Domains (TAD) Music Model [44, 45]. This empirically validated framework outlines distinct stages of talent acquisition and identifies psychological predictors that track progression over time, integrating both predispositions and learned skills. Using
10 Chapter 2. Background this model allows MuSA’s design to be tailored toward a clearly defined audience of learners. The TAD Music Model is not the first attempt to conceptualize musical talent development, but it is distinctive in its use of measurable predictors to link time and achievement [79]. Earlier models overemphasized genetic and environmental factors [80, 81, 82, 83], which are challenging to evaluate empirically due to the cost, complexity, and scarcity of longitudinal studies [84, 42, 85]. As a result, purely geneticor environment-based models have proven less reliable for predicting longterm musical achievement. Figure 1: The Talent-Development-in-Achievement-Domains Framework [44] The TAD Music Model (Figure 1) identifies four stages of musical talent development, each with characteristic predictors and learning needs: 1. Aptitude: innate mental and physical abilities, such as melodic, rhythmic, and tonal sensitivity, that predict a natural capacity for musical success [86]. 2. Competence: the acquisition of skills needed for independent practice, expressive improvement, and repertoire expansion. Key predictors include a
2.1. Foundations of Musical Talent Development 11 growth mindset [87, 88], self-motivation [79, 89], musical self-concept [89, 88], motor skills [90], and deliberate practice. 3. Expertise: a level of achievement characterized by peer recognition, creative problem-solving, and public engagement [91]. Predictors include psychological stability, openness, conscientiousness, resilience, and focused dedication [79]. 4. Transformational achievement: a level of expertise marked by significant, influential, and creative contributions. This stage is less predictable, shaped by factors such as chance, opportunity, creative potential [92], musical form [93, 94], and strategies for sustaining high-level output [95]. The TAD Music Model shows that musicians at different stages need different kinds of support. For example, in the Aptitude stage, a player needs basic technical training to turn natural ability into core technical skills. These include sight-reading, posture, ear-training, and the ability to play a piece with regards to musical aspects like timing, intensity, pitch, and timbre. On the other hand, in the Competence stage, the focus shifts to strengthening technique while also developing a personal style and expressive voice. Musical expression involves shaping musical aspects to convey a work’s emotional character and connect a performer’s interpretation with a listener’s experience [43, 96]. As Section 2.5 demonstrates, many TEL music tools focus on technical skills, providing binary feedback (correct/incorrect) tied to the musical score. This approach misses the interpretive dialogue essential to modern teaching frameworks and leaves a gap for musicians who need to develop their expressive abilities. Moreover, such tools are necessary to guide the reflective aspect of practice, which is guided by the learner’s interpretation of their performance after actively practicing. Therefore, in addition to technical skills, we develop MuSA in such a way as to facilitate musical expression in parallel, as both are essential for developing musical talent and are a core focus of TEL in music education.
12 Chapter 2. Background 2.2 Practice Strategies in Music Education Building on the factors shaping musical talent discussed in Section 2.1, this thesis focuses on practice, since it is the factor that learners can control most. To design TEL practice tools effectively, it is important to understand the different types of practice that influence talent development. Ericsson (2016) identified three types [97]: 1. Naive practice: Playing an instrument for enjoyment without specific goals or structured improvement. 2. Purposeful practice: Goal-oriented practice done independently, without continuous expert guidance or structured feedback. 3. Deliberate practice: Highly structured practice guided by a teacher with deep expertise. The difference in performance outcomes between purposeful and deliberate practice is substantial. In Ericsson (1993) seminal study of violinists at the Music Academy in West Berlin, the most accomplished musicians had accumulated the highest number of deliberate practice hours [6]. Later meta-analyses suggested that deliberate practice accounted for only 23% [4] to 37% [5] of performance variation. Ericsson et al. (2019) argued that these studies used incorrect definitions, and when deliberate practice is properly defined, its explanatory power rises to 61% [76], indicating a strong influence in musical talent development. TEL tools like MuSA support deliberate practice by strengthening the development of musical talent while making learning more effective and accessible through cognitive and pedagogical strategies. 2.2.1 The Master-Apprentice Model Given the central role of deliberate practice, it is unsurprising that traditional music education has long relied on the master–apprentice model, one popular learning strategy. In this approach, students receive in-person guidance from an experienced tutor, and the musical score serves as the core of instruction [98, 99, 100, 101, 10, 11]. Despite the rise of independent learning and online resources, this model
2.2. Practice Strategies in Music Education 13 remains dominant in conservatories and higher education worldwide [102], across both Western and non-Western traditions [103]. Most learners do not have constant access to detailed feedback available through conservatories and master-apprenticeship that can enable reflection, receiving only a single weekly lesson [104]. Left to practice alone for most of their time, these learners may develop incorrect techniques, lose direction, or form ineffective habits. Without consistent guidance, even motivated students can reduce deliberate practice to purposeful, or even naive, practice [105, 9, 8, 76]. To address this challenge, experts have developed teaching approaches that focus on the learner and their needs rather than just the tutor or musical score. The following sections examine constructivist strategies such as dialogic teaching, scaffolding, and modeling, which can promote learner autonomy, develop skills for more effective independent practice, and facilitate reflection. These frameworks guided MuSA’s design, providing a basis to support and enhance practice when expert guidance is limited. 2.2.2 Constructivism and Dialogic Teaching Constructivist learning principles position the learner as an active participant rather than a passive recipient of instruction. In constructivist theory, the teacher acts as a facilitator, designing activities that help learners build understanding by reorganizing and integrating prior knowledge [14, 106, 107, 15, 108]. This learnercentered teaching (LCT) , in contrast to the master-apprentice model’s emphasis on teacher and score authority, can provide the benefits of expert feedback while also fostering a collaborative environment that gives a student ownership over their learning [109]. Components of LCT include the use of interactive, personalized learning activities and the inclusion of technology to manage learning outside of the classroom. LCT has shown to lead to higher learning outcomes and motivation across a variety of fields including medicine [110], mathematics [111], and music [112, 113].
14 Chapter 2. Background A practical expression of LCT in music education is dialogic teaching, an approach in which sustained dialogue between teacher and student becomes the primary medium for learning. In dialogic teaching, rather than delivering one-way instructions, teachers engage learners through questioning, reflection, and collaborative exploration, enabling real-time adaptation of guidance to meet evolving needs [114, 115]. This open feedback channel encourages students to co-construct their own interpretations, which is why MuSA’s performance feedback avoids binary correct/incorrect labeling: MuSA is designed as a learner-centered tool for reflection rather than a score-centered tool for correction. 2.2.3 Scaffolding Another LCT approach is scaffolding, which involves tailoring support to a learner’s current ability and gradually transferring responsibility for performance and interpretation to the student. This process is rooted in the work of Vygotsky, a core contributor to constructivism. It leverages his concept of the zone of proximal development, which is the range of tasks a learner can perform with assistance but not independently, and provides tailored support that gradually decreases as the learner’s competence grows [106, 116]. The key principles of scaffolding (Figure 2) are: 1. Contingency: The teacher’s support must be continually adapted to the learner’s current skill level, preventing the learner from being overwhelmed or unchallenged [117, 118, 119]. 2. Fading: Support is gradually decreased as the learner’s competence grows. For example, a teacher may initially provide extensive modeling but reduce assistance as the learner becomes more proficient with a skill or piece [120, 118]. 3. Transfer of Responsibility: The goal of scaffolding is to shift learning responsibility to the learner, fostering autonomous competence so they can perform tasks independently [121]. 4. Cyclical Process: Scaffolding is a recurring cycle. Once a learner masters one sub-goal (e.g., a basic bow hold on a violin), the teacher introduces a new,
2.2. Practice Strategies in Music Education 15 slightly more challenging sub-goal. The scaffolding process then begins again for this new task [122]. Figure 2: Conceptual model of scaffolding [118]. MuSA can be contextualized within the scaffolding framework by providing a form of contingency and facilitating the transfer of responsibility by acting as an SRL tool. While scaffolding typically relies on a dedicated teacher to tailor support, MuSA’s feedback can act as an TEL intermediary. For example, if a learner is focused on a specific performance outcome, they can analyze a performance recording with MuSA. The platform then responds by highlighting interesting salient moments (see Section 2.3.3) during different musical aspects. This reflects how a teacher provides contingency by indicating salient moments in a student’s performance that may not be obvious to them. The goal is to initiate a transfer of responsibility for that specific task by helping the learner understand what to listen for, what performance outcomes to strive for, and what tools exist to help them practice.
22 Chapter 2. Background 2.3.2 Modeling One popular method for generating feedback in music education and potentially reducing cognitive load when learning a new task is modeling, in which a teacher demonstrates a desired sound or technique often based on a learner’s current performance state [139]. Modeling can take two forms: aural modeling, where learners listen to and emulate expert performances to build internal musical representations [51, 140, 141, 142], and visual modeling, which emphasizes physical aspects like posture and fingering with heavy reliance on written notation [143]. Aural versus visual modeling. While both visual and aural modeling have value, research shows that aural modeling has higher learning outcomes for both technical skills and musical expression since aural modeling focuses on external outcomes (how the instrument sounds) rather than just the internal action (how the body performs) [144, 145, 68, 146, 147, 148, 113]. This conclusion aligns with Meissner’s (2021) dialogic teaching framework, which combines cognitive load theory, dialogic teaching, and modeling-based feedback, suggesting that focusing on the overall outcome rather than becoming entangled in minutia can manage cognitive load [149]. Additionally, over-reliance on visual modeling may weaken a student’s ability to connect their physical actions with the resulting sound [150]. As a result, musicians may struggle to develop musical understanding, or the ability to recognize and manipulate patterns and relationships in music [151]. When modeling emphasizes symbolic information, it becomes score-centered rather than learnercentered, focusing on what to do instead of how and why, thus failing to work towards musical understanding. This is one reason MuSA relies on the sound of an individual’s performance instead of the sheet music behind it. Multi-Modal Modeling. Modeling is often paired with other learning methods in a combination known as multi-modal modeling to further enhance learning outcomes. For instance, while modeling alone can foster dependence, combining it with strategies like verbal feedback encourages students’ interpretive thinking and internal musical representations [149]. This is because verbal teaching methods,
2.3. Feedback 23 such as using metaphors or explicit explanations, are shown to improve musical understanding and strengthen expressive performance [152], which is in line with LCT dialogic teaching outcomes [112, 113]. 2.3.3 Salient Moments Often, when providing feedback using modeling, teachers naturally highlight important performance moments by exaggerating musical aspects like intonation, dynamics, or timing during model performances [67]. This thesis defines these highlighted moments as salient moments. A teacher may indicate salient moments as areas during a performance for technical correction or those where a learner can employ their own musical interpretation relative to an existing norm (like written notation). These may include pitch errors, timing issues, or expressive changes that deserve closer attention. Especially when focusing on musical expression, identifying a salient moment does not inherently imply assigning positive or negative value to it; rather, the identification of a salient moment can start conversation about expressive intent (dialogic teaching) or clarify if a learner’s technique or choices are actually leading to their intended expressive outcome. 2.3.4 Synthesizing SRL, SDL, Feedback, and MuSA Looking at the TAD Model, deliberate practice, scaffolding, and SRL, one theme stands out: reflection is essential at all stages of talent development for turning repetition into real progress, yet it is often neglected. Learners may plan and practice carefully, but without structured opportunities to evaluate their performance, practice can become mechanical and disconnected from long-term goals [153]. As Section 2.5.2 discusses, many music learners use self-recording to reflect on their playing, but this method is often inefficient: reviewing full sessions takes time [154], adds cognitive load, and can lead to passive listening that misses subtle mistakes [138]. As a result, reflection tends to be occasional rather than systematic. MuSA addresses this problem by automatically highlighting salient moments and focusing reflection tasks. By reducing the time and effort needed for review, a tool like MuSA
24 Chapter 2. Background can make reflection easier to do regularly. This approach helps integrate reflection into the practice cycle, improves planning through feedback on recurring patterns (e.g., “intonation drifts sharp in high positions”), and supports independent learning while hopefully developing critical listening and self-assessment skills. 2.4 Technological Pedagogical Content Knowledge Designing a new TEL tool to support the reflection process in music practice, such as MuSA, requires guidance beyond broad psychological or cognitive principles. Models such as Technological Pedagogical Content Knowledge (TPCK) and its expansion through Technology Mapping (TM) provide a structured way to connect these principles to the specific challenges of music education. The following section outlines how this model, which emphasizes the interplay of technology, pedagogy, content, and the affective domain, informs and justifies MuSA’s design. TPCK. Around 20 years ago, researchers introduced Technological Pedagogical Content Knowledge (TPCK) to guide the design of TEL tools by emphasizing how subject knowledge, pedagogy, and technology intersect in teaching [155, 17, 18]. Angeli and Valanides (2013) expanded this framework with an instructional design approach called Technology Mapping, which specifies how tools should be developed to use technology’s unique affordances for learning [19]. The principles of TPCK include: 1. Find high-impact content: Identify topics that are difficult for students to understand or for teachers to explain without the aid of technology. 2. Create unique representations: Develop new ways to present content that are only possible with technology, making it easier to grasp. 3. Leverage novel teaching methods: Identify and use teaching strategies that are difficult or impossible to implement with traditional methods alone. 4. Choose the right tools: Select technology with the appropriate features to support the teaching and learning goals. 5. Design learner-centered activities: Create learner-centered classroom ac-
2.5. TEL Feedback Systems in Music Education 25 tivities that effectively integrate technology. TM and the Affective Domain. Macrides and Angeli (2018) expanded Technology Mapping for music education by adding the affective domain (Figure 5) [156, 157]. This addition highlights that emotion, motivation, and personal expression are central to music learning, addressing a gap in the original framework [158], which only focused on pedagogical and cognitive aspects. Research shows that many students struggle with composition, improvisation, and listening because of limited skills or confidence [159, 160, 161]. By integrating the affective domain, the expanded model reframes these challenges as opportunities to connect technical learning with expression and emotion. Educators argue that musical experiences cannot be separated from their expressive character [162], and that emphasizing self-expression increases engagement and motivation [8, 151]. Using TPCK to Inform MuSA. MuSA’s design reflects this expanded model, which stresses that music learning involves both cognitive and emotional aspects [163]. MuSA provides flexible, non-corrective feedback that students can interpret in ways that suit their needs, echoing the view that focusing on expression can boost engagement [8]. By highlighting salient sections of a recording, MuSA directs attention to meaningful changes in dynamics, tempo, or articulation, features known to influence emotional responses [158, 164]. This approach lowers barriers for learners with different skill levels, since the focus is on recorded analysis rather than live performance [27]. At the same time, its saliency analysis supports SRL by helping students reflect on their playing and take more control of their progress, in line with TM’s emphasis on interactive, learner-centered use of technology [19, 163]. 2.5 TEL Feedback Systems in Music Education A wide range of TEL tools can support music learning, with different tools addressing different TPCK principles depending on their purpose, audience, and market focus. While the master–apprentice model still dominates music education [10, 11], technology is increasingly used both inside and outside the classroom by traditional
26 Chapter 2. Background Figure 5: A Technology Mapping model showing the interrelations of musical elements and concepts, technology, and affect [156].
2.5. TEL Feedback Systems in Music Education 27 and independent learners. These tools range from basic devices like metronomes and tuners to advanced systems that provide personalized feedback on performance [138]. This section examines the current state of TEL for music performance and practice to show how MuSA addresses a gap in this area. 2.5.1 Online Resources Internet Availability. The availability of TEL in music education would not be possible without the internet, which has been a powerful tool in the dissemination of information, particularly in regions where digital access is widespread. Internet connectivity has expanded across households in the United States and Europe, making these regions particularly relevant for examining technology’s role in music learning. A 2023 Pew Research Center survey reports that 95% of U.S. adults use the internet, 90% own a smartphone, and 80% have high-speed internet at home [165]. Similarly, the European Commission reported in 2021 that 96% of Europeans have access to a mobile phone, and 82% of households have internet access [166]. These statistics highlight the infrastructural foundation for digital learning tools in these regions. YouTube. A 2022 Pew Research Center survey found that YouTube4was the most widely used online video platform among American teenagers (96%) and European young adults (84%). YouTube serves as a primary hub for informal beginner learning via instructional mock video lesson demonstrating technique, exercises, and repertoire [138]. Advanced learners also use YouTube to access information addressing technical, philosophical, and career challenges. While the platform’s strengths lie in its accessibility, low cost, and self-pacing through video controls [138], its major limitation is the lack of personalized performance feedback. Music Education Websites. Music forums and music education websites like Dave Conservatoire5, Music Theory6, My Music Theory7, and SmartMusic8may fill 4https://youtube.com/ 5https://daveconservatoire.org/ 6https:/musictheory.net/ 7https://mymusictheory.com/ 8https://smartmusic.com/
28 Chapter 2. Background YouTube’s gap by generating conversation about technique and theory, but without personalized performance feedback, these platforms don’t address critical aspects of music talent development like posture, sound quality, or musical expression [138]. Young adult learners also rarely interact with these platforms and seem to gravitate towards unidirectional content [167], indicating that the content levied may be too tedious to learn without a dedicated instructor or a platform that engages better with the motivation or self-regulation principles of SDL. Online Learning. Personalized performance feedback can be delivered remotely via online learning, which includes synchronous, asynchronous, and hybrid formats. Online learning is suggested to be as effective as in-person instruction [168, 169] and may expand access to master-apprentice learning in domains like piano [170]. Critics of online learning mention that reliance on SRL and student motivation may impact student engagement [171, 172]. Regardless, this remains a key challenge for independent learning. Higher education institutions have adopted tools that replicate real-world musical workspaces such as LoLa9, JackTrip10, and UltraGrid11 to reduce location-based learning restrictions and promote student independence. Online Collaboration. Online collaboration is another way to give personalized performance feedback and cultivate motivation. Platforms like AudioMovers12, Source-Connect13, Sessionwire14, SonoBus15, and VST Connect16 offer products with professional-grade audio streaming for remote collaboration [173]. These tools prioritize high-quality sound, low-latency capabilities, and digital audio workstation integration, enabling remote music creation that may train creative and expressive skills, a key aspect of TPCK for effective technology implementation in music education [163]. Real-time "jamming" tools like MusicianLink17 and Artsmesh18 can 9https://lola.conts.it/ 10https://jacktrip.com/ 11https://ultragrid.cz/ 12https://audiomovers.com/ 13https://source-elements.com/products/source-connect/ 14https://sessionwire.com/ 15https://sonobus.net/ 16https://steinberg.net/vst-connect/ 17https://musicianlink.com/ 18https://artsmesh.com/
2.5. TEL Feedback Systems in Music Education 29 also facilitate collaboration without reliance on in-person sessions. However, these high-fidelity tools face challenges like latency, data loss, and high barrier-to-entry. 2.5.2 Non-Real-Time Recording Feedback Audio and Video Recording. Leveraging audio and video recordings is another way to receive personalized performance feedback and is proven to improve one’s own playing [61]. Many teachers already encourage students to make video record performances to review movements or gestures that are difficult to notice in real-time due to the high cognitive load of simultaneously performing and self-evaluating [138]. Separating performance and review enables more objective assessment as recordings, whether audio or video, allow musicians to analyze specific musical elements, compare recordings with those of teachers or professionals, and study physical technique in detail [138]. Reviewing recordings is also a tool for SRL and SDL by facilitating better musical understanding and setting practice goals [174]. However, despite the benefits of reviewing audio and video recordings, many learners fail to regularly review their recordings due to the time consuming nature of the process [154]. Studies show that using audio and video recordings to provide computer-assisted visual feedback can help musicians improve dynamic control and awareness by giving clear feedback on their expressive range [175]. Computer-Assisted Audio Recording Feedback. Audio recordings can be abstracted through the visualization of important musical aspects like frequency spectrum, fundamental frequency (pitch), intensity, and resonance areas (formants) using free audio analysis software like Praat19, Sonic Visualiser20, Audacity21, or LARA22 [176]. These tools also form the basis of purpose-designed visual feedback tools that display different audio analyses at different screen locations, sometimes aligned with reference audio to interpret results. For pitched audio, these include 19https://fon.hum.uva.nl/praat/ 20https://sonicvisualiser.org/ 21https:/audacityteam.org/ 22https://hslu.ch/en/lucerne-school-of-music/forschung/perfomance/lara/
30 Chapter 2. Background tools like Sing and See23, VocaVista24, and Tartini25, while for percussive audio, they may include audio analysis software plugins like BeatRoot or OnsetDS from Sonic Analyser [176]. However, many of these visual feedback tools have complicated interfaces with a high barrier to entry for the average learner, may require an expert with a sound and music computing background to interpret results, or are primarily oriented to real-time rather than non-real-time feedback. Computer-Assisted Video Recording Feedback. Video recordings can also be abstracted via visual feedback tools. Older systems relied on markers placed to track motion or mark clear contrast between an object and its background, like EyesWeb26, which analyzes expressive performance. Purdue University’s AIM project is developing an experimental tool called the Evaluator27 to assist string musicians with individual practice by analyzing a musician’s sound and video, comparing it to a digitized score to detect deviations in relevant musical aspects (intonation, rhythm, dynamics), and even offer posture correction suggestions. However, non-real-time video recording performance analysis software specialized to musical performances has not yet been deeply explored. General open-source video analysis tools like Kinovea28 are available for studying body mechanics and gestures. Onform29 is another tool, though specialized to sports performance, that allows instructors to give targeted feedback on video recordings using voice-overs and on-screen markup tools. 2.5.3 Real-Time Feedback Computer-Assisted Real-Time Audio Feedback. While available tools for computer-assisted audio recording feedback are either experimental tools or aimed at those with production, engineering, or sound and music computing backgrounds, many tools have been developed to provide learners real-time feedback based on performance audio. Some of these tools include those already mentioned such as 23https://singandsee.com/ 24https://vocevista.com/en/ 25https://cs.otago.ac.nz/graphics/Geoff/tartini/index.html 26https://casapaganini.unige.it/eyesweb_bp 27https://github.com/Purdue-Artificial-Intelligence-in-Music/Evaluator-code 28https://kinovea.org/ 29https://onform.com/
2.5. TEL Feedback Systems in Music Education 31 EyesWeb, Sing and See, VocaVista, and Tartini. General real-time feedback for ear training has been explored with EarMaster30. Real-time audio feedback for string instrument performance has been particularly explored with experimental opensource platforms Intonia31, which helps string players visualize intonation. More recently the SkyNote32 web application was developed to track a string player’s pitch in real-time and overlay their pitch curve with a score, allowing learners to see if their pitch and timing correctly align with given notation. Plectrus [177, 178] is another experimental application developed for training intonation in novce string players. Other commercial platforms leveraging real-time audio feedback include Yousician33, MakeMusic34, Uberchord35, and Tonestro36, and Riyaz37 which all compare user’s real-time audio to a selected score. Cortosia38 by KORG provides real-time feedback on features such as dynamics, pitch stability, timbre, and attack clarity, though receiving simultaneous aspects can make it difficult to independently assess each one. Computer-Assisted Real-Time Video Feedback. Real-time video feedback is well illustrated by the iMaestro project, which created an “augmented mirror” that uses visualization and sonification to analyze a string player’s gestures live. TELMI is another project for string players, offering real-time visual feedback on posture, gesture, and bowing, along with score-aligned features such as dynamics, pitch, and timbre [179]. The Evaluator instead compares student performances to pre-recorded examples with correct form. However, real-time video feedback systems remain less developed than audio ones, largely due to the technical demands of motion capture and the current limits of commercial AI motion-tracking software. 30https://earmaster.com/ 31https://intonia.com/index.shtml 32https://appskynote.com 33https://yousician.com/ 34https://makemusic.com/ 35https://uberchord.com/ 36https://tonestro.com/ 37https://riyazapp.com/ 38https://korg.com/us/products/software/cortosia/
38 Chapter 3. Designing and Implementing MuSA considerable overlap in key expressive features—such as intonation, vibrato, tempo, and dynamics—and also align with the primary research areas of my supervisor (violin) and co-supervisor (voice). 3.2.1 Selecting Musical Features for MuSA-V2 One of the key pieces of feedback I received from the violinists at the RCM based on MuSA-V1 was that the most important musical aspects to visualize and indicate moments of saliency were dynamics and tempo. Both of these aspects have been supported in the literature as indicators of saliency through variability [21, 25, 26]. Consequently, one of the main goals for the second version was to implement a reliable dynamics and tempo extraction algorithm that could also identify salient moments. This was a challenge in the original version: although an onset envelope was calculated using librosa.onset.onset_strength, smoothed with a median filter, and a tempogram was generated with librosa.feature.tempogram, estimating and visualizing global and local tempo using librosa.feature.rhythm.tempo proved unreliable, as the tempo curve could not be accurately inferred from the tempogram. Therefore, I needed to develop a new approach for visualizing tempo over time and select feature extraction algorithms that were robust. Additionally, I had to determine a set of musical aspects for MuSA-V2 that were interesting to implement for both violin and voice, while also narrowing the scope to allow for a complete redevelopment of the MuSA application. Prioritizing some musical features over others. Not all musical features from MuSA-V1 were carried into V2. While dynamics and pitch remained, timbre and articulation were omitted. Integrating these was possible, but my priority was developing a reliable dynamic and tempo visualization, and the RCM specifically emphasized dynamics and tempo. Exclusion of timbre. Timbre is generally particularly challenging to extract, as it depends on interactions among multiple lower-level features, including the energy spectrum, short-term transients, and correlations with loudness and pitch [182, 183],
3.2. Implementing MuSA-V2 39 Table 2: Feature extraction methods and criteria for salient moment selection (MuSA V2) with exact parameters. Musical Aspect Extraction Method Saliency Selection Dynamics Framewise RMS from librosa. feature.rms and smoothed RMS trace. Parameters: N_FFT = 2048, HOP_LENGTH = 512, smoothing: mean filter with window_percentage = 0.1. Sliding window: WINDOW_PERCENTAGE = 0.05, HOP_PERCENTAGE = 0.0125. High-variability sections computed using standard deviation over windows. Threshold = 50% of max variability. Returns frame indices and audio-time ranges. Pitch Framewise pitch in Hz with confidence, smoothed and filtered by RMS. Extracted using CREPE (EssentiaPitchCREPEmodel, CREPE_HOP_DURATION_SEC = 0.01) or Librosa pyin. Smoothing/segmentation: WINDOW_PERCENTAGE = 0.05, HOP_PERCENTAGE = 0.0125, RMS threshold = 0.01. High-variability sections computed using calculate_high_variability_sections with is_pitch=True. Sections with largest deviation from nearest equal-tempered piano note were highlighted. Tempo Dynamic tempo in BPM using Essentia BeatTrackerMultiFeature. Smoothed and interpolated to regular time axis. Sample rate = 44100 Hz. Smoothing: median filter. Window fraction: 0.1. No discrete highlighted sections; returns continuous interpolated tempo trace over time. Phonation Frame/window-level phonation class probabilities (breathy, flow, neutral, pressed) using VGGISH embeddings and PHONATION_MODEL (Keras). Window size = 2 frames, hop = 1 frame, frame duration = 0.1 s. No highlighted sections; returns per-class probability time series. Vibrato Per-frame vibrato oscillations around smoothed pitch. Detection uses stable-note identification and oscillation checks. Vibrato extent (amplitude) and rate (Hz) calculated per section. Thresholds: min_sustained_length = 5 frames, gap threshold = 5 frames. Extent threshold for highlighting = 1.9. CREPE_HOP_ DURATION_SEC = 0.01 used for time conversion. Highlighted sections are frames with extent > 1.9, grouped into consecutive segments. Returns both frame ranges and audio-time ranges.
40 Chapter 3. Designing and Implementing MuSA and so I chose to exclude it from MuSA-V2. Articulation was also left aside due to time constraints. Based on RCM feedback, tempo was added, along with vibrato and phonation mode for voice, guided by the work of my co-supervisor Suvi Häärä, with vibrato included because it is relevant for both voice and violin. The final set of musical aspects, their extracted audio features, and the method for selecting salient moments are summarized in Table 2. Implementing dynamic window and hop sizes. An important goal was to ensure that extracted feature curves maintained consistent smoothing and data representation regardless of audio length. For instance, a shorter audio segment using a fixed window size would yield fewer data points than a longer segment, which could lead to noisier curves if the window size were not adjusted. To address this, sliding window and hop sizes were defined as percentages of the audio length where applicable (e.g., for dynamics and pitch), allowing the feature curves to scale proportionally and remain comparable across recordings. Struggles with identifying variability thresholds. One key piece of feedback from MuSA-V1 was the need to identify multiple salient moments within each audio feature, rather than a single instance. This was addressed by selecting all moments exceeding a specific variability threshold. Alternative approaches could have included selecting the top nmost variable sections or applying a threshold to the top nsections, but the implemented method was chosen to avoid excluding potentially interesting moments. The thresholds were determined empirically during development rather than being drawn from existing literature, as few sources provided guidance on precise values. Establishing thresholds based on prior studies represents a clear avenue for improving the robustness and generalizability of the method in future work. Struggles with tempo extraction. Dynamic tempo tracking was a major challenge for this project. I tested several algorithms, including Essentia’s RhythmExtractor2013 and TensorflowPredictTempoCNN. Because voice and violin lack a strong beat, RhythmExtractor2013 returned few usable points, while TensorflowPredictTempoCNN
3.2. Implementing MuSA-V2 41 produced only 2–3 tempo values per segment, especially for short audio, limiting variability analysis. BeatTrackerMultiFeature provided more reliable estimates and a higher number of points, so it was selected for implementation. However, the density of points was still insufficient to confidently identify salient moments, so tempo variability was not fully integrated into MuSA-V2. Investigating optimal tempo extraction methods remains an interesting direction for future research, but it fell outside the scope of this thesis, which focused on designing and building the final MuSA-V2 web application. Implementing voice-specific features. Phonation mode was implemented as an experimental feature in MuSA-V2. Phonation mode describes the manner in which the vocal folds vibrate to produce sound, and changes in phonation throughout a performance can provide valuable insights for singers, particularly in refining expressive techniques and conveying emotion. Suvi Häärä developed an experimental phonation mode classification model using the Yesiler Phonation dataset, distinguishing four vocal modes: breathy (soft, airy), neutral (standard singing or speech), flow (resonant and clear), and pressed (tense, harsh). I incorporated visualizations of these phonation modes into MuSA with the intention of exploring salient moments based on mode transitions; however, due to time constraints, this functionality was not fully implemented. Nevertheless, all four phonation modes can be visualized in the application, representing a promising direction for future research. Feature extraction takeaways. Overall, feature extraction was a foundational component in developing MuSA-V2, enabling several successful salient moment visualizations. For example, dynamics produced clear highlighted sections, which may be supported by findings from the evaluation study (Section 4). Pitch extraction using CREPE resulted in accurate, smooth pitch curves, and the highlighted sections effectively captured moments of potential intonation—where frequencies deviated from their corresponding equal-tempered piano notes. Vibrato extraction was relatively straightforward, though improvements could be made to identify moments of variability more specifically, rather than relying solely on a fixed threshold. These
42 Chapter 3. Designing and Implementing MuSA and other limitations of salient moment identification, as well as potential alternative approaches, are discussed in Section 5.1.4. 3.2.2 Designing MuSA In designing MuSA-V2, I considered learner-centered teaching (LCT), self-regulated learning (SRL), and self-directed learning (SDL) principles (Section 2.3), alongside insights from Technological Pedagogical Content Knowledge (TPCK; Section 2.4). TPCK highlighted the importance of identifying topics that are difficult to teach or understand without technology, and of presenting content in ways only possible with technological tools. Accordingly, MuSA-V2 was designed to support both independent and classroom learning, using a web application to make salient moment identification and visual feedback on audio recordings more intuitive. Notably, algorithmically detecting salient moments based on variability is only feasible with such technological support. MuSA landing page. Based on critiques of MuSA-V1, I aimed to redesign the interface to minimize the number of steps required, reducing extraneous cognitive load for users. For example, users first select one of three instruments on the landing page—violin, voice, or polyphonic (Figure ??). The polyphonic option allows users to analyze recordings with multiple instruments or melodic lines, making the application more broadly accessible. Once an instrument is selected, the available analyses are automatically filtered to those relevant to that instrument. Additionally, audio can be recorded directly within the web application, eliminating the need to use separate platforms, although users can still upload pre-recorded audio if preferred (Figure ??). Feature-specific visualizations. Additionally, it was important to ensure that the musical aspect visualizations were properly contextualized against their own scales (in a musically relevant way, not just computationally), so users would be able to get meaning out of the results without dealing with the cognitive load associated with trying to interpret very numerical results. For example, pitch (Figure 8b)
3.3. MuSA System Architecture 43 and vibrato (Figure 8d) are displayed with a piano-roll background to contextualize frequency within musical notation. Dynamics (Figure 8a) are presented with simple amplitude or BPM scales. The different phonation mode classes (Figure 8c) are shown in separate windows to reflect it’s dimensional nature. Vibrato also has separate windows for extent and rate. Salient moment visualizations. Salient moments, when available for a specific feature, are highlighted and synchronized with the audio waveform. Users can click directly on a highlighted section to play the corresponding audio, providing an immediate connection between the audio and visual elements of the application. Zoom and pan functionality is available across all graphs, ensuring that highlighted sections and audio playback remain aligned at any level of detail. Other considerations. Users who are unfamiliar with MuSA-V2 require guidance on how to navigate the interface and identify the functions of each button. To address this, information tooltips are toggleable rather than always displayed. Since the application involves sending audio recordings to a server for analysis, users can also choose whether their data is collected via a toggleable consent button (default on). Additionally, feature calculations are cached to avoid redundant computations, improving the overall user experience (discussed further in the following section). 3.3 MuSA System Architecture Another central consideration in building MuSA-V2 was ensuring accessibility and scalability so that the application could meaningfully impact a wide range of users. To support deployment at scale and simultaneous multi-user access, the system was designed with an architecture comprising four main components. Figure 10 provides an overview of the system architecture. •Frontend (React4+ Tailwind5): Handles the interactive user interface and renders feature visualizations using D36. The component-based model of React 4https://react.dev/ 5https://tailwindcss.com/ 6https://d3js.org/
44 Chapter 3. Designing and Implementing MuSA Figure 7: MuSA landing page (top) and file selection or audio recording (bottom).
3.3. MuSA System Architecture 45 (a) Dynamics view. (b) Pitch view. (c) Phonation view. (d) Vibrato view. (e) Tempo view. Figure 8: Overview of MuSA analysis views.
46 Chapter 3. Designing and Implementing MuSA and the utility-first styling of Tailwind facilitate a flexible and maintainable design. •Node.js7API: Serves as the middle layer between the frontend and the MongoDB8database, managing session metadata and communication. •Python Service (Flask9+ Gunicorn10): Manages all audio-related tasks, including preprocessing, signal and feature extraction, saliency analysis, and caching. Redis11 is used for short-term caching to minimize redundant computation, and audio files are stored locally on the server. •Infrastructure (Ubuntu + Nginx12): Handles data management, deployment, and user authentication. Ubuntu is the server with Nginx configured as a reverse proxy. Redis manages caching, MongoDB stores user and session metadata, and JSON Web Tokens (JWT)13 secure user sessions. Explanation of architecture. A separate Python service was used instead of Essentia.js for several reasons. Essentia.js is optimized for short clips (under one minute), while this application targets full performance recordings. Its functionality is also more limited; for example, Essentia’s CREPE pitch-tracking implementation is not available in Essentia.js. Finally, server-side preprocessing—such as trimming silences with Librosa and resampling to the required sample rate—was necessary before analysis. Admittedly, a client-only approach using Essentia.js might have sufficed for a prototype on short audio segments, but this was a lesson learned in hindsight. To ensure user privacy, data is maintained in isolation through the use of Redis-based short-term caching and JWT authentication, ensuring that each user’s uploaded audio and corresponding analyses remain private. Example user interaction. To better understand how the archiecture works, a process diagram of an example user flow is shown in Figure 9. When a musician 7https://nodejs.org/ 8https://mongodb.com/ 9https://flask.palletsprojects.com/en/stable/ 10https://gunicorn.org/ 11https://redis.io/ 12https://nginx.org/ 13https://jwt.io/introduction#what-is-json-web-token
3.3. MuSA System Architecture 47 first logs into MuSA, their session is authenticated through a secure JWT connection managed by a Redis cache. Once inside, they select the instrument they want to analyze—for example, violin. The system then prompts them to choose which performance features to focus on, such as pitch stability, loudness, or articulation. After uploading their recording, the Python service processes the audio, extracts the chosen features, calculates saliency, and caches the results to speed up interactions. The user then explores these results in the React frontend, where they can play back their performance, zoom into specific passages, and see salient moments highlighted directly on the timeline. When they are finished, their analysis is stored automatically, with metadata saved in MongoDB and the audio preserved on the server so they can return to it in future practice sessions. 1. Authentication: JWT session with Redis cache 2. Instrument Selection: Violin, Voice, or Polyphonic 3. Feature Selection: Choose features to analyze 4. Feature Extraction: Python service analyzes audio, calculates saliency, caches results 5. Visualization: React frontend renders data, synchronizes playback, zoom, and highlighted salient moments 6. Storage: Metadata in MongoDB, audio on server Figure 9: Example user interaction with MuSA.
54 Chapter 4. MuSA Evaluation Study Figure 12: Distribution of the 14 participants retained for final analysis across musical experience levels: 1 Beginner, 7 Intermediate, 5 Advanced, and 1 Professional. When mapped onto the TAD Music Model (Aptitude, Competence, Expertise, Transformational Achievement), the sample reflects a concentration in the Competence (Intermediate) and Expertise (Advanced) stages, aligning with the sound and music computing student demographic at Universitat Pompeu Fabra. ation design allows for a more direct comparison between performance outcomes and underlying psychological and cognitive processes. Although familiar labels such as Intermediate and Advanced were used to describe participants, these categories correspond to the Competence and Expertise stages of the TAD framework. This mapping makes it possible to interpret differences in self-perception and performance not only in terms of musical skill but also through the lens of developmental processes in expertise acquisition. For instance, participants at the Competence stage may rely more heavily on external feedback to structure practice and refine their playing, whereas those at the Expertise stage may integrate feedback differently, drawing on internalized strategies and prior knowledge. 4.2.1 Objective Performance Metrics Objective performance was derived from computational analysis of recorded audio using the same feature extraction algorithms implemented in MuSA. Pitch was esti-
4.2. Results 55 mated using CREPE [189], dynamics were quantified via Root Mean Square (RMS) energy using the feature.rms function from librosa, and tempo was measured using Essentia’s BeatTrackerMultiFeature algorithm [190, 191]. Participant recordings under control and feedback conditions were compared with reference performances. To account for temporal variability, dynamic time warping (DTW) aligned performances for point-by-point comparison, a technique that has been proven to work effectively for aligning audio signals under a variety of conditions [192, 193]. Because pitch, dynamics, and tempo operate on different scales, analyses were conducted in three complementary ways: raw RMSE values for unprocessed deviations, baseline-standardized percentage improvements for withinsubject changes, and z-score standardization to allow cross-feature comparison and effect size evaluation. Effect sizes were calculated using Cohen’s d, which quantifies the standardized mean difference between groups, providing a measure of practical significance beyond statistical tests. This approach provides both feature-specific and integrative perspectives on the impact of MuSA feedback. Raw RMSE results. As shown in Figure 13, raw RMSE values before and after the intervention displayed considerable spread across all three musical features. Tempo was the only feature with a statistically significant effect, with the feedback group performing worse than the control group (p= 0.0096,d= 0.97; Appendix Table 4). Pitch and dynamics showed no significant condition effects (p > 0.32, |d|<0.27), though the wide variability illustrated in the boxplots suggests substantial individual differences that may have masked smaller effects. Taken together, these results suggest that tempo was most sensitive to disruption from feedback, while pitch and dynamics remained inconclusive at this level of analysis (see Appendix Table 5). Experience-level analyses further clarified these raw RMSE patterns. As shown in Figure 14, intermediate musicians showed moderate effects for pitch (∆ = −77.05, d=−0.37), small effects for dynamics (∆=0.0038,d= 0.15), and moderate effects for tempo (∆ = 5.71,d= 0.38), all corresponding to nonsignificant results
56 Chapter 4. MuSA Evaluation Study Figure 13: Raw RMSE performance across three musical features—pitch (Hz), dynamics (RMS), and tempo (BPM)—before and after the intervention for both control and feedback groups. Box plots display medians, quartiles, and outliers for each condition. Lower RMSE values indicate better performance. Sample sizes ranged from 9–14 participants per condition. Results show substantial spread across participants, with tempo most clearly disrupted by feedback, while pitch and dynamics show no clear condition-level effects. Figure 14: Change in RMSE for pitch, dynamics, and tempo across four musician experience levels (After – Before), with negative values indicating improved performance. No statistically significant effects were observed for condition or timing (all p > 0.05). Beginner and Professional groups had very small sample sizes (n=2), and although some large effect sizes were seen (e.g., Intermediate tempo d= 0.94, Beginner dynamics d= 7.58), these results are likely influenced by high variability. (p > 0.09). Advanced musicians exhibited small-to-moderate effects across features (pitch: ∆ = 14.65,d= 0.33; dynamics: ∆ = −0.0048,d=−0.16; tempo: ∆=9.54, d= 0.59), again nonsignificant (p > 0.32). Beginner and professional groups had extremely small sample sizes, producing highly variable and unstable effect sizes (e.g., Beginner dynamics Feedback vs Control: d= 7.58) that were also nonsignificant. Overall, although no comparisons reached statistical significance, the effect sizes and mean differences suggest that responses to feedback varied by experience, with
4.2. Results 57 intermediate musicians tending toward moderate improvements and beginners/professionals showing high variability and extreme but unreliable effects (Appendix Tables 6-8). Baseline-standardized improvements. To account for baseline performance differences, percent improvement scores were calculated for each feature (Appendix Table 9, Figure 15). Across all participants, feedback showed a modest relative advantage for dynamics (+14.1 pp compared to control), but no benefit for pitch (–10.1 pp) and a marked detriment for tempo (–110.3 pp). None of these contrasts reached statistical significance (p > 0.28), reflecting both limited sample size and high interindividual variability, but the pattern suggests that feedback selectively influenced temporal control while offering limited support for pitch or dynamic accuracy. Figure 15: Percent improvement from individual baselines across pitch, dynamics, and tempo. Violin plots with overlaid box plots show percentage change relative to pre-intervention baseline. Horizontal dashed line indicates no change. Statistical markers show effect sizes and p-values (all non-significant). When analyzed by musical experience level (Appendix Table 10, Figure 16), no statistically significant subgroup effects were observed. Descriptively, intermediate musicians showed the largest positive differences across features (dynamics +27.1%, tempo +116.7%, pitch +2.9%), while advanced musicians’ outcomes trended closer to zero or negative (pitch –28.5%, tempo –392.0%, dynamics +0.3%). Beginners and professionals contributed very limited data, with single cases per condition, making their results uninterpretable.
58 Chapter 4. MuSA Evaluation Study Figure 16: Percent improvement from individual baselines by musical experience level (Beginner, Intermediate, Advanced, Professional). Violin plots compare feedback and control groups within each category; all comparisons were non-significant. Dashed line shows no change from baseline. Z-score standardized comparisons. To account for cross-feature scale differences, RMSE values were z-score standardized against baseline distributions. As shown in Figure 17, feedback was associated with a small positive shift in dynamics (+0.41 SD, d= 0.41) and a negligible effect for pitch (–0.17 SD, d= 0.16), while tempo showed a small detrimental shift (–0.43 SD, d= 0.43). None of these differences reached statistical significance (p > 0.28). More summary statistics, including effect sizes and p-values, are provided in Appendix Tables 11–12. Figure 17: Z-score standardized improvement by feature (Pitch, Dynamics, Tempo) for control and feedback conditions. Violin plots show distributions, with overlaid box plots for median and quartiles. Positive values indicate above-baseline improvement. All comparisons were non-significant (p > 0.28). When broken down by musical experience (Figure 18), no subgroup effects reached statistical significance, though descriptive differences appeared. Intermediate musi-
4.2. Results 59 Figure 18: Z-score standardized improvement by experience level (Beginner, Intermediate, Advanced, Professional) for Pitch, Dynamics, and Tempo. Violin plots show distributions for control and feedback conditions. Statistical annotations indicate effect sizes (d) and p-values. Most differences were non-significant, highlighting heterogeneity in responses across experience levels. cians showed the largest positive shifts (dynamics +0.78 SD, tempo +0.46 SD, pitch +0.05 SD), while advanced musicians tended toward neutral or negative changes (pitch –0.46 SD, tempo –1.54 SD, dynamics +0.01 SD). Beginner and professional participants contributed only isolated data points, limiting interpretability. Overall, these standardized analyses suggest potential variation in feedback effects by experience level, but given small samples and non-significant tests, these trends should be considered exploratory. 4.2.2 Perceived Performance Metrics To evaluate participants’ subjective experience of improvement, self-ratings on pitch, dynamics, and tempo were collected before and after practice sessions under control (no feedback) and feedback conditions. Paired t-tests were conducted within each condition to assess changes over time, while between-condition comparisons (feedback vs. control) were performed on preand post-practice ratings. Effect sizes (Cohen’s d) were calculated to quantify the magnitude of observed differences. This approach allowed for examining both overall perceived performance changes and variations across musical experience levels, providing a complementary perspective to objective performance metrics. Overall, participants’ self-ratings varied by feature and condition (Figure 19). For
60 Chapter 4. MuSA Evaluation Study Figure 19: Comparison of self-rated performance before and after the intervention across different musical experience levels. Each subplot represents one musical feature (pitch, dynamics, tempo). Box colors distinguish ’Before’ and ’After’ ratings, highlighting changes in perceived performance across experience levels. pitch, perceived improvements were minimal, with feedback producing a small average increase (+0.15 points, p= 0.656,d=−0.122), indicating little impact of the feedback tool. Dynamics showed the largest perceived gains, increasing on average by +1.00 points with feedback, which was highly significant (p= 0.0009∗∗∗, d=−1.003). Tempo improvements were less consistent: feedback led to a moderate increase (+0.58 points, p= 0.089,d=−0.563), while practice alone produced a slightly larger improvement (+0.94 points, p= 0.0106∗), suggesting general practice contributed to tempo perception more than feedback (Appendix Table 13). Considering musical experience (Figure 21), differences emerged across features. For pitch, advanced participants showed the largest perceived improvement with feedback (+0.60), though this change was not statistically significant (p= 0.4766, d= 0.14), whereas beginners showed inconsistent changes. Dynamics benefited all experience levels, with beginners and intermediates showing the largest positive shifts (+2.00 and +1.29 points with feedback, respectively); only intermediates reached statistical significance (p= 0.0004∗∗∗,d= 0.92), while changes for beginners
4.2. Results 61 Figure 20: Self-reported performance ratings by musical experience across three features: (A) Pitch, (B) Dynamics, (C) Tempo. Box plots show ratings before (light purple) and after (medium purple) practice, combining control (no feedback) and experimental (feedback) conditions. Medians (thick line), interquartile range (box), and outliers (circles) are shown. Experience levels: Beg = Beginner, Int = Intermediate, Adv = Advanced, Pro = Professional. Figure 21: Mean perceived improvement (after minus before) by musical experience and intervention type for (A) Pitch, (B) Dynamics, (C) Tempo. Dark purple bars indicate feedback, light purple bars indicate control. Positive values reflect improvement, negative values reflect decline. Numbers above bars show exact values ≥0.01. Horizontal line at zero indicates no change. (+2.00) and advanced participants (+0.40) were not significant. Tempo improvements were most notable in the control condition for beginners (+3.00, p= 0.011∗), with smaller gains from feedback observed in other groups (intermediate: +1.05, p= 0.089; advanced: +0.00, p= 1.0). Overall, perception of dynamics performance improved with MuSA, particularly for intermediate participants, while perceived improvements in pitch were limited and tempo gains likely reflected general practice effects rather than MuSA’s reflective feedback (Appendix Table 14).
62 Chapter 4. MuSA Evaluation Study 4.2.3 Self-Awareness Metrics Self-awareness reflects the degree to which participants’ perceptions of their own performance align with objectively measured outcomes. To quantify this, perceived improvements were compared with actual performance changes using Pearson correlation coefficients. Correlations were calculated between self-reported ratings and RMSE values, between perceived improvement and percentage improvement, and between helpfulness ratings and actual performance gains, providing a primary index of self-assessment accuracy. For direct comparison across scales, objective improvements were rescaled to approximately match the Likert range by dividing by 7. Absolute self-awareness error was then computed as the absolute difference between subjective and scaled objective improvements, with lower values indicating greater accuracy. Condition-level differences were assessed using Mann–Whitney U tests, independent t-tests, and Cohen’s dto quantify effect sizes. Overall metrics. Comparing subjective self-ratings to objective measures across pitch, dynamics, and tempo revealed generally weak correlations, indicating poor self-awareness among participants. Overall, the Control group showed r=−0.027 (p= 0.874,n= 36) and the Feedback group r=−0.182 (p= 0.304,n= 34), with neither reaching statistical significance. Feature-specific analyses mirrored this pattern: for pitch, r= 0.089 (Control) and r=−0.160 (Feedback), p= 0.062; for dynamics, r=−0.267 (Control) and r= 0.101 (Feedback), p= 0.148; and for tempo, r= 0.317 (Control) and r=−0.363 (Feedback), p= 0.694. None of these correlations were statistically significant, suggesting that participants’ perceived improvements did not reliably align with objective performance across features or conditions (Appendix Table 15). To further characterize self-awareness, participants’ assessments were classified into four categories: Accurate,Overconfident,Underconfident, and Somewhat Inaccurate. Classification was based on the difference between subjective improvement and scaled objective improvement. A response was considered Accurate if this
4.2. Results 63 Figure 22: Correlation scores between self-ratings and objective performances for three musical features (pitch, dynamics, and tempo). Figure 23: Scatter plots examining the relationship between participants’ perceived improvement (subjective self-ratings) and actual performance improvement (objective measurements) across three musical features: pitch, dynamics, and tempo. Control participants are shown in orange, and feedback participants in blue. The red dashed line indicates perfect self-awareness. absolute difference was ≤1, accounting for typical perceptual noise. Overconfident responses reflected large positive subjective changes despite performance decline (subjective improvement >1and objective improvement <−5%), whereas Underconfident responses reflected negative self-ratings despite substantial improvement (subjective improvement <−1and objective improvement >10%). Remaining cases were labeled Somewhat Inaccurate, capturing moderate misalignments with-
70 Chapter 5. Discussion decisions. By highlighting salient performance sections and showing feature-level deviations in pitch, dynamics, or timing, MuSA enables learners to target these areas for particular outcomes based on their intended expression. In this way, the tool supports self-directed learning (SDL) and learner-centered teaching (LCT) principles, giving students ownership over how they analyze and refine their performances, and allowing expressive goals to emerge as a product of informed, feature-driven practice rather than prescriptive instruction. The concept of salience analysis was central to MuSA’s design. Rather than prescribing how a musician should express themselves, MuSA identifies moments in a performance that deviate from feature-specific expected norms. These norms are computed relative to the uploaded performance itself. For example, a fast dynamic change or high variability in dynamics compared to surrounding sections would be highlighted. For pitch, salient moments correspond to deviations from the expected note within a 12-semitone scale, and vibrato sections are detected and highlighted when they exceed a defined threshold. Importantly, these expected norms are inherently flexible and can be adapted based on a user’s goals, cultural context, or notation preferences. As a point of comparison, one of the goals of TELMI, the parent project of MuSA, was to provide real-time visual feedback on timbre, a feature that is inherently subjective. The success of TELMI demonstrates that meaningful feedback can be provided even when the underlying musical attribute is difficult to quantify, and that it is both acceptable and expected for TEL tools to interpret measurements or classifications of subjective data in order to offer targeted feedback to learners [1]. Similarly, by defining salient moments through a specific, feature-driven approach, MuSA helps narrow the learner’s focus during practice. Instead of having to consider all possible ways a moment could be interpreted as salient, the tool provides one concrete lens for attention, guiding practice in a manageable and structured way. This does not preclude the use of other techniques, such as human-guided feedback based explicitly on affect, which could be applied alongside MuSA to capture aspects of performance that extend beyond feature-based salient moment analysis.
5.1. Discussion 71 5.1.2 Evaluating MuSA: Demand, Limitations, and Outcomes The evaluation study was an essential step in validating the concepts behind MuSA. Despite its limitations—a small sample size with few usable data, the use of different microphones, and the remote, asynchronous nature of the study—it produced a number of findings. Subjective feedback from the post-study questionnaire indicated that participants perceived MuSA as a potentially helpful tool, particularly for improving awareness of dynamics. This aligns with the overall positive sentiment expressed in the sentiment analysis, where participants appreciated the clarity and usefulness of highlighted salient moments. Minor usability issues were noted, including lack of real-time responsiveness and occasional misalignment between visual feedback and actual performance, highlighting opportunities for refining the user experience. Objective performance metrics. The objective performance metrics revealed a more nuanced picture of MuSA’s impact across skill levels and musical features. Dynamics consistently benefited the most from feedback, whereas pitch improvements were minimal and tempo outcomes were inconsistent. Experience level further moderated these effects: intermediate musicians, corresponding to the Competence stage of the TAD model, often demonstrated measurable gains, while advanced musicians, corresponding to the Expertise stage, sometimes showed neutral or slightly negative changes, particularly for tempo. Although these trends were not statistically significant, they point to the expertise reversal effect, whereby feedback methods that benefit less experienced learners may disrupt the performance strategies of experts [194, 195, 196, 197, 198]. This also aligns with Colombo (2017), who suggested that young music students benefit from greater teacher support and monitoring, while more advanced students gain more from developing metacognitive skills such as reflecting on their learning strategies [199]. Perceived performance and self-awareness. Analysis of perceived performance revealed that participants reported the greatest improvement for dynamics, moderate improvement for tempo, and minimal change for pitch. Notably, some tempo
72 Chapter 5. Discussion improvements appeared to result from practice alone rather than tool-mediated feedback, emphasizing the importance of distinguishing intrinsic learning effects from the influence of MuSA. Self-assessment metrics demonstrated generally weak alignment with objective performance, suggesting limited self-awareness. Intermediate musicians showed modestly better self-assessment accuracy, whereas advanced musicians were often inversely aligned with their measured outcomes. These findings align with research on musician self-awareness, which suggests that learners’ self-assessments can be inaccurate [200]. Overall, poor self-assessment may directly reflect of a challenge in the self-reflection phase of the SRL cycle. They also highlight the complex interaction between experience, feature type, and self-evaluation in music practice, indicating that MuSA could be enhanced with metacognitive prompts to support more accurate performance judgments. Taken together, these results indicate that MuSA may selectively support musical practice as most effective for dynamics and for learners at intermediate skill levels, while advanced learners may experience neutral or slightly negative effects. Subjective impressions suggest that the tool successfully directs attention to salient performance moments, although future iterations could benefit from improvements in interactivity and alignment of visual feedback with audio. These insights underscore the importance of combining subjective and objective measures when evaluating digital feedback tools and point toward practical design considerations, including feature-specific guidance and adaptive feedback calibrated to experience level. Limitations. One of the main limitations of the MuSA evaluation study was that it tested the application as a whole, rather than isolating the effect of salient moment highlighting itself. Participants were asked to perform a task with the full analyzer and without it, meaning that only general statements about MuSA’s overall utility can be made, rather than conclusions specifically about the impact of indicating salient moments versus displaying raw feature curves alone. This distinction is important, particularly because tempo saliency analysis was not implemented due to algorithmic limitations. Interestingly, this provides a natural comparison: dynamics feedback showed a more consistently positive trend on performance, whereas tempo
5.1. Discussion 73 feedback—absent in the final version—might have had a markedly negative effect if implemented, illustrating a potential differential impact of feature type on practice outcomes. Another limitation is that the study did not specifically assess whether MuSA improves a musician’s expressive intent or ability to create expressive performance moments, and moreover, didn’t explicitly confirm its use as a reflective, non-corrective, post-performance tool. Testing this functionality and acquiring more qualitative data regarding its reflective potential would have required recruiting both highly skilled musicians capable of intentionally manipulating expressive parameters and less experienced participants, as well as facilitating extended, carefully structured practice sessions—constraints that exceeded the available thesis timeline, especially given the time-intensive process of developing the MuSA web application itself. Additionally, the choice of two simple pieces based on melodic segments from common children’s songs may not have been adequate to evaluate the use of MuSA across all stages of the TAD Music Model. For example, the simplicity of the reference audios may have been trivial for advanced participants, also potentially explaining the inverse relationship between expertise level and improvement. Ideally, the evaluation would be using different music pieces relevant to each participant, but this would have introduced an extra variable, perhaps overcomplicating the evaluation study based on the given thesis timeframe. Finally, the small usable sample size (14 out of 32 participants) limits the generalizability of the findings. Most observed effects were not statistically significant, making it difficult to draw definitive conclusions. Nevertheless, the study revealed intriguing patterns regarding feature-specific and experience-dependent effects of feedback, providing a valuable benchmark and setting the stage for future research and iterations of MuSA or similar non-real-time feedback applications. Implications. Nevertheless, the feedback from participants and the initial request from RCM musicians highlight a clear, unmet need for non-real-time feedback tools that engage with practice self-reflection. Existing tools, such as Sonic Analyser and
74 Chapter 5. Discussion Audacity, are powerful but lack pedagogical orientation and the ability to abstract raw audio into meaningful, musically salient moments. MuSA, therefore, represents a promising approach, but further research with a larger, more diverse participant sample is needed to validate its pedagogical impact and refine which types of salient moment analysis are most effective across skill levels. 5.1.3 Independent Practice and Learner Implications Developing MuSA was an important step in exploring how technology can enhance independent practice, which is often unstructured and lacks expert feedback. As discussed in Section 2.2.5, independent practice requires self-regulated learning (SRL) skills, such as self-reflection and goal-setting, which are challenging for learners without consistent teacher guidance. By highlighting salient moments, MuSA aims to scaffold the self-monitoring process, directing a learner’s attention to specific points for reflection and experimentation. This supports a cycle of reflective learning that is crucial for developing self-directed learning skills. This potential benefit is particularly important given the challenges learners face with metacognition, or the awareness of their own thinking and learning processes. The self-awareness metrics from this study suggest that metacognitive accuracy is a significant hurdle in independent practice, as participants across the board showed generally poor alignment between their subjective perceptions and objective performance outcomes (r=−0.122). This finding aligns with broader research indicating that learners often struggle to accurately assess their own performance [201, 202]. Interestingly, while the self-evaluation study suggested that intermediate (Competence) musicians were moderately accurate in judging their own performances after using MuSA, this pattern may partly reflect self-selection bias: individuals identifying as intermediate may approach self-assessment and reflection more cautiously and analytically, whereas those identifying as advanced (Expertise) may rely more on internalized, emotionally driven expectations of how they should perform at their level. Nevertheless, this overall lack of self-awareness underscores the importance of ex-
5.1. Discussion 75 ternal tools in supporting independent learning. Since metacognitive skills are not always innate, technology can provide a concrete foundation for reflection, whether through binary correct/incorrect feedback that highlights salient moments (as MuSA does) or by identifying performance characteristics such as affect. With respect to improving performance self-awareness (which was not MuSA’s primary goal), the present study did not reveal clear gains from using the analyzer tool, but it is difficult to draw definitive conclusions. Self-awareness typically develops gradually within scaffolding frameworks and dialogic learning contexts, where sustained external feedback such as guidance from a tutor supports its growth over time. It is possible that extended use of a tool like MuSA could contribute to shaping selfawareness, particularly as study participants were also managing the cognitive load of learning how to use the system, which may have influenced self-awareness results. 5.1.4 MuSA’s Limitations and Future Directions Expanding feature-driven saliency analyses. MuSA’s current approach to saliency analysis is primarily feature-driven, focusing on identifying moments where objective musical features—such as pitch, dynamics, or tempo—show significant variability relative to an established performance norm. During the implementation phase, however, several roadblocks emerged. For example, highlighting moments of high tempo variability proved difficult because the amount of tempo data extracted from a piece was much lower compared to pitch or dynamics. In addition, developing a signal processing algorithm that could consistently handle shorter audio segments within the given time constraints was not feasible. As a result, tempo variability could not be highlighted as reliably as other features. Future work should therefore prioritize implementing more robust algorithms for tempo analysis, as the MuSA framework already supports such functionality and only requires the right computational approach. Beyond tempo, other features present opportunities for further refinement. For instance, MuSA currently identifies vibrato moments using a fixed threshold. While this method works, it misses the expressive nuance conveyed through changes in
76 Chapter 5. Discussion vibrato rate and extent. Yang et al. (2013) demonstrated that variability in vibrato can signal expressive intent [25], suggesting that saliency detection could be improved by tracking vibrato variability across a performance rather than applying a hard cutoff. This would allow the tool to capture more abstract and musically meaningful performance moments. Similarly, phonation mode has been implemented for voice, but its role in saliency remains an open question. Is it the transition from one phonation mode to another that creates salience? Or is it sustained use of a particular mode in contrast to surrounding material? Investigating such questions would not only refine how salience is operationalized but also expand the interpretive power of MuSA in analyzing expressive vocal performance. Integrating more features. A missed opportunity in MuSA was also to detect salient moments through timbre, for example, as timbre is one of the primary mechanisms through which we detect emotion in music [43, 24]. This could, for example, involve detecting and displaying variable moments in timbre related low-level features like spectral centroid and flux to users and letting them interpret musical meaning and understanding from these highlighted areas based on their intention. Though the original prototype for MuSA included highlighting the most variable timbre section based on MFCCs, the final version did not include it for timing and scheduling reasons based on workload and the design of the application itself. Integrating affect-driven saliency analysis. Moreover, a complementary approach to the feature-based one could involve affect-driven saliency analysis, such as activity analysis [180]. In this case, a model could be trained to identify key moments across different instrumental performances—starting, for instance, with voice and violin—by detecting points where listeners consistently perceive a particular affective quality, such as sadness or happiness. Integrating activity analysis would also align more explicitly with the TDPK Music Framework (Section 2.4), which emphasizes that musical experiences are inseparable from their expressive character. Compared to feature-driven saliency analysis,
5.1. Discussion 77 this approach would provide more prescriptive feedback about potential affect and musical expression. At the same time, as the evaluation study suggests, such prescriptive feedback on expressive sections might be most effective for learners who are still developing these skills, while more advanced musicians may benefit less from this type of guidance (a potential reverse expert effect as previously mentioned). Need for musical understanding. This points to another potential limitation of the tool: effectively interpreting the feedback provided by MuSA likely requires a certain level of musical understanding, typically at the intermediate level or beyond. For beginners, an explicitly prescriptive, task-level feedback tool may be more effective, as suggested by the McPherson et al. feedback matrix [130]. At the same time, feedback must be appropriately tailored so as not to interfere with the selfregulatory abilities that intermediate and advanced learners have developed. In this sense, the fact that MuSA is not overly prescriptive and instead requires some interpretive background knowledge may, from the learner’s perspective, be considered a benefit rather than a drawback. Musa as a web application. One explicit benefit of MuSA is its fully online format. Users do not need to download software, create an account, or navigate complex setup processes. The range of available analyses is intentionally limited, which some might view as a drawback since users cannot access an extensive menu of options or review past uploads. However, this simplicity also reduces barriers to use: learners can either record directly within the application or upload existing audio files and receive visual feedback immediately. While tools such as Sonic Visualiser offer sophisticated analysis capabilities, they were not designed with a pedagogical audience in mind, but rather for computational research, and are therefore underutilized in performance review contexts. By contrast, MuSA’s streamlined design makes it more accessible for musicians, minimizing cognitive load compared to the steep learning curve of downloading, configuring, and interpreting results from specialized software like Sonic Visualiser.
78 Chapter 5. Discussion Current architectural limitations. Currently, MuSA functions as a minimum viable product. While it can handle a large number of simultaneous user interactions through its Redis-based setup, it is not yet optimized for large-scale deployment. Current constraints include shared server resources, redundancy from the dual Node.js and Flask services—which could be streamlined by consolidating database calls into the Python service—and reliance on local file storage, which would benefit from a cloud-based solution for improved scalability and resilience. All collected data is stored directly on the server, which has limited capacity. While this setup allows data to be retained, it also introduces the challenge of managing and maintaining growing storage demands. In practice, this means that someone must periodically transfer the collected audio files to an external drive, a process that can be both time-consuming and inconvenient. Additionally, the development stack uses Create React App, which has been deprecated; migrating to Vite would modernize the build process and improve efficiency. Addressing these redundancies and infrastructure limitations is a priority for future versions, but due to time constraints, a fully commercial-level tool was not feasible. Nevertheless, MuSA provides a functional foundation for delivering salience-based analysis to learners, enabling non-real-time performance feedback that can inform future iterations and related projects. Integrating written notation. In terms integration with notation, MuSA does not currently employ score alignment with performances, which some may view as a limitation. Score alignment, however, comes with its own drawback: it requires the performer to have a notated score for comparison, which may not always be available or relevant. MuSA deliberately avoids this feature because its focus is on the performance itself, privileging the performer’s expressive intent over adherence to notation, and distancing itself from the binary correct/incorrect prescriptions seen in real-time feedback tools like Yousician. That said, score alignment can also offer valuable insights into performance and expressive intent, and several studies highlight its usefulness. Future iterations of MuSA could therefore integrate score alignment as an optional feature to enhance salient-moment analysis, particularly
5.1. Discussion 79 when key score events are involved. Such integration could even allow users to set variability targets based on the score or receive feedback framed in more traditional musical terms. Importantly, incorporating score alignment may be more feasible in MuSA’s non–real-time setting, since real-time score following remains technically challenging to implement reliably. Performance-to-performance alignment. Another feature originally intended for MuSA but not implemented due to time and resource constraints was direct comparison to a target audio recording. This would have been valuable because the reference performance itself contains feature-driven salient moments that could be extracted and directly compared to those in the user’s recording. While comparative performance analysis is already used in the field, there remains a need for more sophisticated approaches beyond simple error detection [203]. In this sense, MuSA could eventually be expanded to compare salient moments between a student’s performance and a reference recording, offering feedback that is non-binary and more oriented toward expressive intent and musicality. However, implementing such a feature was beyond the scope of this iteration, especially given the technical and resource demands of building a large-scale, multi-user application. Data collection potential. While MuSA was primarily designed as a pedagogical application, it also functions as a valuable data collection tool. Since MuSA is deployed on a server hosted by the MTG at Universitat Pompeu Fabra, all uploaded recordings are automatically stored (with users explicitly consenting to data collection before using the tool). Additionally, each recording is labeled by instrument (currently voice and violin), allowing MuSA to serve as a passive, labeled dataset for these instruments. These recordings, accessible to those with server permissions, could be repurposed for future music performance analysis or music information retrieval tasks requiring labeled instrumental data.
List of Tables 1 Feature extraction and salient moment selection methods (MuSA V1). 37 2 Feature extraction methods and criteria for salient moment selection (MuSA V2) with exact parameters. . . . . . . . . . . . . . . . . . . . 39 3 Questionnaire items and corresponding labels. . . . . . . . . . . . . . 66 4 Two-way ANOVA results for Raw RMSE performance (Condition × Timing)...................................114 5 Group means and effect sizes (Cohen’s d) for RMSE performance. Significance: ** p < 0.01. . . . . . . . . . . . . . . . . . . . . . . . . 114 6 Pitch (Hz) RMSE comparisons by experience level. . . . . . . . . . . 115 7 Dynamics (RMS) RMSE comparisons by experience level. . . . . . . . 115 8 Tempo (BPM) RMSE comparisons by experience level. . . . . . . . . 115 9 Overall baseline-corrected RMSE improvements comparing feedback and control conditions. Cohen’s dand p-values indicate nonsignificant differences..................................116 10 Baseline-corrected RMSE improvements (%) for pitch, dynamics, and tempo by musician experience level. Positive values indicate an advantage for the feedback condition; all comparisons were nonsignificant.116 11 Overall Z-Score Standardized Improvement: Control vs Feedback . . 116 12 Z-Score Standardized Improvement by Experience Level: Control vs Feedback..................................116 13 Overall Perceived Performance Ratings by Condition. Significance: * p<0.05,***p<0.001..........................116 14 Perceived Performance Improvement (∆= After - Before) with tstatistics and p-values by Experience Level and Condition. . . . . . . 117 86
LIST OF TABLES 87 15 Correlation Analysis of Self-Assessment Accuracy by Feature . . . . . 117 16 Self-Awareness Comparison: Control vs Feedback Conditions . . . . . 117 17 Self-Awareness by Musical Experience Level (Correlation with ObjectivePerformance).............................117 18 Representative participant quotes by question and sentiment. . . . . . 118 19 Summary of sentiment analysis results, including overall statistics, per-question means and medians, and correlation between lexicon and VADERsentiment.............................119
Bibliography [1] Technology enhanced learning of musical instrument performance (telmi). European Commission Horizon 2020 Research and Innovation Programme, Grant Agreement No. 688269 (2016–2019). URL https://telmi.upf.edu/. Project duration: February 2016 – January 2019. [2] Ericsson, K. A. The Scientific Study of Expert Levels of Performance: general implications for optimal learning and creativity 1. High Ability Studies 9, 75– 100 (1998). URL https://doi.org/10.1080/1359813980090106. Publisher: Routledge. [3] McPherson, G. E., Blackwell, J. & Hallam, S. Musical Potential, Giftedness, and Talent Development. In McPherson, G. E. (ed.) The Oxford Handbook of Music Performance, Volume 1, 31–55 (Oxford University Press, 2022). URL https://doi.org/10.1093/oxfordhb/9780190056285.013.3. [4] Platz, F., Kopiez, R., Lehmann, A. C. & Wolf, A. The influence of deliberate practice on musical achievement: a meta-analysis. Frontiers in Psychology 5 (2014). [5] Macnamara, B. N., Hambrick, D. Z. & Oswald, F. L. Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis. Psychological Science 25, 1608–1618 (2014). URL https: //doi.org/10.1177/0956797614535810. Publisher: SAGE Publications Inc. 88
BIBLIOGRAPHY 89 [6] Ericsson, K. A., Krampe, R. T. & Tesch-Römer, C. The role of deliberate practice in the acquisition of expert performance. Psychological Review 100, 363–406 (1993). Place: US Publisher: American Psychological Association. [7] Meinz, E. J. & Hambrick, D. Z. Deliberate practice is necessary but not sufficient to explain individual differences in piano sight-reading skill: The role of working memory capacity. Psychological Science 21, 914–919 (2010). Place: US Publisher: Sage Publications. [8] Hallam, S. et al. The development of practising strategies in young people. Psychology of Music 40, 652–680 (2012). URL https://doi.org/10.1177/ 0305735612443868. Publisher: SAGE Publications Ltd. [9] McPherson, G. E., & Renwick, J. M. A Longitudinal Study of Self-regulation in Children’s Musical Practice. Music Education Research 3, 169–186 (2001). URL https://doi.org/10.1080/14613800120089232. Publisher: Routledge. [10] Gaunt, H., , C., Andrea, , L., Marion, & Hallam, S. Supporting conservatoire students towards professional integration: one-to-one tuition and the potential of mentoring. Music Education Research 14, 25–43 (2012). URL https: //doi.org/10.1080/14613808.2012.657166. Publisher: Routledge. [11] Gaunt, H. Apprenticeship and empowerment: the role of one-to-one lessons. In Rink, J., Gaunt, H. & Williamon, A. (eds.) Musicians in the Making: Pathways to Creative Performance, 0 (Oxford University Press, 2017). URL https: //doi.org/10.1093/acprof:oso/9780199346677.003.0003. [12] Phillips, J. A. Student self-assessment and reflection in a learner controlled environment (2016). URL https://arxiv.org/abs/1608.00313. [13] Dos Santos Silva, C. & Marinho, H. Self-regulated learning processes of advanced musicians: A prisma review. Musicae Scientiae 29, 89–108 (2025). URL https://journals.sagepub.com/doi/10.1177/10298649241275614.
90 BIBLIOGRAPHY [14] Magoon, A. J. Constructivist approaches in educational research. Review of Educational Research 47, 651–693 (1977). URL https://journals.sagepub. com/doi/10.3102/00346543047004651. [15] Lim, C. P. & Chan, B. C. microlessons in teacher education: Examining pre-service teachers’ pedagogical beliefs. Computers Education 48, 474– 494 (2007). URL https://www.sciencedirect.com/science/article/pii/ S0360131505000424. [16] Hattie, J. Visible Learning for Teachers (Routledge, 2012), 0 edn. URL https://www.taylorfrancis.com/books/9781136592331. [17] Angeli, C. & Valanides, N. Preservice elementary teachers as information and communication technology designers: An instructional systems design model based on an expanded view of pedagogical content knowledge. Journal of computer assisted learning 21, 292–302 (2005). [18] Angeli, C. & Valanides, N. Epistemological and methodological issues for the conceptualization, development, and assessment of ict–tpck: Advances in technological pedagogical content knowledge (tpck). Computers & education 52, 154–168 (2009). [19] Angeli, C. & Valanides, N. Technology mapping: An approach for developing technological pedagogical content knowledge. Journal of Educational Computing Research 48, 199–221 (2013). [20] Hattie, J. & Timperley, H. The power of feedback. Review of educational research 77, 81–112 (2007). [21] Sloboda, J. A. The role of musical structure in emotional response. Psychology of Music 24, 34–50 (1996). [22] Juslin, P. N. Cue utilization in communication of emotion in music performance: Relating expressive intention to acoustic features. Journal of Experimental Psychology: Human Perception and Performance 26, 1797–1813 (2000).
BIBLIOGRAPHY 91 [23] Gómez, E. & Herrera, P. An approach to the analysis of musical expressive timing in piano performance recordings. Journal of New Music Research 35, 3–17 (2006). [24] Eerola, T. & et al. An optimized factorial design for studying the emotion perception from music. In International Conference on Music Perception and Cognition (2012). [25] Yang, L. & et al. Vibrato performance style: a case study comparing erhu and violin. In 10th International Symposium on Computer Music Modeling and Retrieval (CMMR 2013), 109–116 (CMMR, 2013). [26] Li, P.-C., Su, L., Yang, Y.-H. & Su, A. W. Y. Analysis of expressive musical terms in violin using score-informed and expression-based audio features. In Proceedings of the 16th International Society for Music Information Retrieval Conference, 441–448 (ISMIR, 2015). [27] Bauer, W. I., Reese, S. & McAllister, P. A. Transforming music teaching via technology: The role of professional development. Journal of Research in Music Education 51, 289–301 (2003). URL https://journals.sagepub. com/doi/10.2307/3345656. [28] Savage, J. A survey of ict usage across english secondary schools. Music Education Research 12, 89–104 (2010). URL http://www.tandfonline.com/ doi/abs/10.1080/14613800903568288. [29] Greher, G. R. Music technology partnerships: A context for music teacher preparation. Arts Education Policy Review 112, 130–136 (2011). URL http: //www.tandfonline.com/doi/abs/10.1080/10632913.2011.566083. [30] Wise, S., Greenwood, J. & Davis, N. Teachers’ use of digital technology in secondary music education: illustrations of changing classrooms. British Journal of Music Education 28, 117–134 (2011). URL https://www.cambridge.org/ core/product/identifier/S0265051711000039/type/journal_article.
92 BIBLIOGRAPHY [31] Mroziak, J. & Bowman, J. Music tpack in higher education: Educating the educators. Handbook of technological pedagogical content knowledge (TPACK) for educators 285–296 (2016). [32] Miksza, P. Relationships among achievement goal motivation, impulsivity, and the music practice of collegiate brass and woodwind players. Psychology of Music 39, 50–67 (2011). URL https://doi.org/10.1177/0305735610361996. Publisher: SAGE Publications Ltd. [33] Miksza, P., Blackwell, J. & Roseth, N. E. Self-Regulated Music Practice. Journal of Research in Music Education 66, 295–319 (2018). URL https://www.jstor.org/stable/48588925. Publisher: [MENC: The National Association for Music Education, Sage Publications, Inc.]. [34] Acquilino, A. & Scavone, G. Current state and future directions of technologies for music instrument pedagogy. Frontiers in Psychology 13, 835609 (2022). URL https://www.frontiersin.org/articles/10.3389/fpsyg. 2022.835609/full. [35] Thorndike, E. Educational Psychology. Vol. I: The Original Nature of Man. Educational Psychology. Vol. I: The Original Nature of Man. (New York, Columbia Univ., 1913). Pages: Pp. xii+, 327. [36] Council, N. R. et al. How people learn: Brain, mind, experience, and school: Expanded edition, vol. 1 (National Academies Press, 2000). [37] Sweller, J. Cognitive load theory, learning difficulty, and instructional design. Learning and instruction 4, 295–312 (1994). [38] Sweller, J. Cognitive load theory. The psychology of learning and motivation: Cognition in education, Vol. 55 37–76 (2011). Place: San Diego, CA, US Publisher: Elsevier Academic Press. [39] Rego, C. & Montague, E. The impact of feedback modalities and the influence of cognitive load on interpersonal communication in nonclinical settings:
BIBLIOGRAPHY 93 Experimental study design. JMIR Human Factors 10, e49675 (2023). URL https://humanfactors.jmir.org/2023/1/e49675. [40] Lurie, N. H. & Swaminathan, J. M. Is timely information always better? the effect of feedback frequency on decision making. Organizational Behavior and Human Decision Processes 108, 315–329 (2009). URL https://www. sciencedirect.com/science/article/pii/S0749597808000745. [41] Karlsson, J. & Juslin, P. N. Musical expression: an observational study of instrumental teaching. Psychology of Music 36, 309–334 (2008). URL https: //journals.sagepub.com/doi/10.1177/0305735607086040. [42] McPherson, G. E., Davidson, J. W. & Faulkner, R. Music In Our Lives: Rethinking Musical Ability, Development, and Identity (Oxford University Press, 2012). URL https://doi.org/10.1093/acprof:oso/9780199579297.001. 0001. [43] Gabrielsson, A. Music Performance Research at the Millennium. Psychology of Music 31, 221–272 (2003). URL https://doi.org/10.1177/ 03057356030313002. Publisher: SAGE Publications Ltd. [44] Preckel, F. et al. Talent development in achievement domains: A psychological framework for withinand cross-domain research. Perspectives on Psychological Science 15, 691–722 (2020). URL https://doi.org/10. 1177/1745691619895030. PMID: 32196409, https://doi.org/10.1177/ 1745691619895030. [45] Müllensiefen, D. et al. 84talent development in music. In The Oxford Handbook of Music Performance, Volume 1 (Oxford University Press, 2022). URL https://doi.org/10.1093/oxfordhb/9780190056285.013.32.https: //academic.oup.com/book/0/chapter/357714233/chapter-ag-pdf/ 57022543/book_42624_section_357714233.ag.pdf. [46] Johnston, H. The use of video self-assessment, peer-assessment, and instructor feedback in evaluating conducting skills in music stu-
94 BIBLIOGRAPHY dent teachers. British Journal of Music Education 10, 57–63 (1993). URL https://www.cambridge.org/core/product/identifier/ S0265051700001431/type/journal_article. [47] Robinson, C. R. Singers’ self-assessment of choral performance: next-day recollections versus concert tape evaluation. Southeastern J. Music Educ 4, 224–233 (1993). [48] Daniel, R. Self-assessment in performance. British Journal of Music Education 18, 215–226 (2001). URL https://www.cambridge.org/core/product/ identifier/S0265051701000316/type/journal_article. [49] Hewitt, M. P. Self-evaluation tendencies of junior high instrumentalists. Journal of Research in Music Education 50, 215–226 (2002). URL https: //journals.sagepub.com/doi/10.2307/3345799. [50] Silveira, J. M. & Gavin, R. The effect of audio recording and playback on self-assessment among middle school instrumental music students. Psychology of Music 44, 880–892 (2016). URL https://doi.org/10.1177/ 0305735615596375. Publisher: SAGE Publications Ltd. [51] Woody, R. H. Learning expressivity in music performance: An exploratory study. Research Studies in Music Education 14, 14–23 (2000). URL https: //journals.sagepub.com/doi/10.1177/1321103X0001400102. [52] Woody, R. H. Learning from the experts: Applying research in expert performance to music education. Update: Applications of Research in Music Education 19, 9–14 (2001). URL https://journals.sagepub.com/doi/10. 1177/87551233010190020103. [53] Williamon, A. Musical ExcellenceStrategies and Techniques to Enhance Performance (Oxford University Press, 2004). URL http://www.oxfordscholarship.com/view/10.1093/acprof:oso/ 9780198525356.001.0001/acprof-9780198525356.
BIBLIOGRAPHY 95 [54] Juslin, P. N., Karlsson, J., Lindström, E., Friberg, A. & Schoonderwaldt, E. Play it again with feeling: Computer feedback in musical communication of emotions. Journal of Experimental Psychology: Applied 12, 79–95 (2006). URL https://doi.apa.org/doi/10.1037/1076-898X.12.2.79. [55] Timmers, R. & Sadakata, M. Training Expressive Performance by Means of Visual Feedback: Existing and Potential Applications of Performance Measurement Techniques, 304–327 (Oxford University PressOxford, 2014), 1 edn. URL https://academic.oup.com/book/27746/chapter/197952585. [56] Zimmerman, B. J. A social cognitive view of self-regulated academic learning. Journal of educational psychology 81, 329 (1989). [57] Zimmerman, B. J. Attaining self-regulation: A social cognitive perspective. In Handbook of self-regulation, 13–39 (Elsevier, 2000). [58] Zimmerman, B. J. & Moylan, A. R. Self-regulation: Where metacognition and motivation intersect. In Handbook of metacognition in education, 299–315 (Routledge, 2009). [59] Hatfield, J. L., Halvari, H. & Lemyre, P.-N. Instrumental practice in the contemporary music academy: A three-phase cycle of self-regulated learning in music students. Musicae Scientiae 21, 316–337 (2017). URL https:// journals.sagepub.com/doi/10.1177/1029864916658342. [60] Cleary, T. J. & Callan, G. L. Assessing Self-Regulated Learning Using Microanalytic Methods, 338–351 (Routledge, 2017), 2 edn. URL https://www.taylorfrancis.com/books/9781317448662/chapters/ 10.4324/9781315697048-22. [61] Nusseck, M., Wild, F., Sischka, C. & Spahn, C. Effects of audio feedback interventions with the disklavier on the performance of piano students. Frontiers in Psychology 16, 1568021 (2025). URL https://www.frontiersin. org/articles/10.3389/fpsyg.2025.1568021/full.
102 BIBLIOGRAPHY [110] Chung *, J. C. C. & Chow, S. M. K. Promoting student learning through a student-centred problem-based learning subject curriculum. Innovations in Education and Teaching International 41, 157–168 (2004). URL http: //www.tandfonline.com/doi/abs/10.1080/1470329042000208684. [111] Clark, K. The effects of the flipped model of instruction on student engagement and performance in the secondary mathematics classroom. The Journal of Educators Online 12 (2015). URL https://www.thejeo.com/archive/2015_ 12_1/clark. [112] Shively, J. Constructivism in music education. Arts Education Policy Review 116, 128–136 (2015). URL http://www.tandfonline.com/doi/full/ 10.1080/10632913.2015.1011815. [113] Meissner, H. & Timmers, R. Young musicians’ learning of expressive performance: The importance of dialogic teaching and modeling. Frontiers in Education 5, 11 (2020). URL https://www.frontiersin.org/article/10. 3389/feduc.2020.00011/full. [114] Alexander, R. J. Towards dialogic teaching: Rethinking classroom talk (2008). [115] Alexander, R. A dialogic teaching companion (Routledge, 2020). [116] Granott, N., Fischer, K. W. & Parziale, J. Bridging to the unknown: A transition mechanism in learning and development. Microdevelopment: Transition processes in development and learning 131–156 (2002). [117] Lajoie, S. P. Extending the scaffolding metaphor. Instructional science 33, 541–557 (2005). [118] Van de Pol, J., Volman, M. & Beishuizen, J. Scaffolding in teacher–student interaction: A decade of research. Educational psychology review 22, 271–296 (2010). [119] Van Der Linden, J., Johnson, R., Bird, J., Rogers, Y. & Schoonderwaldt, E. Buzzing to play: lessons learned from an in the wild study of real-time
BIBLIOGRAPHY 103 vibrotactile feedback. In Proceedings of the SIGCHI Conference on Human factors in Computing Systems, 533–542 (2011). [120] Pea, R. D. The social and technological dimensions of scaffolding and related theoretical concepts for learning, education, and human activity. Journal of the Learning Sciences 13, 423–451 (2004). URL http://www.tandfonline. com/doi/abs/10.1207/s15327809jls1303_6. [121] Reigosa, C. & Jiménez-Aleixandre, M.-P. Scaffolded problem-solving in the physics and chemistry laboratory: difficulties hindering students’ assumption of responsibility. International Journal of Science Education 29, 307–329 (2007). [122] Küpers, E., Van Dijk, M., McPherson, G. & Van Geert, P. A dynamic model that links skill acquisition with self-determination in instrumental music lessons. Musicae Scientiae 18, 17–34 (2014). URL https://journals. sagepub.com/doi/10.1177/1029864913499181. [123] Hallam, S. The development of metacognition in musicians: Implications for education. British journal of music education 18, 27–39 (2001). [124] Williamson, S. N. Development of a self-rating scale of self-directed learning. Nurse Researcher 14, 66–83 (2007). [125] Garrison, D. R. Self-Directed Learning: Toward a Comprehensive Model. Adult Education Quarterly 48, 18–33 (1997). URL https://doi.org/10. 1177/074171369704800103. Publisher: SAGE Publications Inc. [126] Parkes, K. A. Self-Directed Learning Strategies. In McPherson, G. E. (ed.) The Oxford Handbook of Music Performance, Volume 1, 0 (Oxford University Press, 2022). URL https://doi.org/10.1093/oxfordhb/9780190056285. 013.7. [127] Gibbons, M. The self-directed learning handbook: Challenging adolescent students to excel (San Francisco, 2002).
104 BIBLIOGRAPHY [128] Rashid, T. & Muhammad Asghar, H. Technology use, self-directed learning, student engagement and academic performance: Examining the interrelations. Computers in Human Behavior 63, 604–612 (2016). Place: Netherlands Publisher: Elsevier Science. [129] Sadler, D. R. Formative assessment and the design of instructional systems. Instructional Science 18, 119–144 (1989). URL http://link.springer.com/ 10.1007/BF00117714. [130] McPherson, G. E., Blackwell, J. & Hattie, J. Feedback in music performance teaching. Frontiers in Psychology 13, 891025 (2022). URL https://www. frontiersin.org/articles/10.3389/fpsyg.2022.891025/full. [131] Wisniewski, B., Zierer, K. & Hattie, J. The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in psychology 10, 487662 (2020). [132] Seufert, T., Hamm, V., Vogt, A. & Riemer, V. The interplay of cognitive load, learners’ resources and self-regulation. Educational Psychology Review 36, 50 (2024). URL https://link.springer.com/10.1007/s10648-024-09890-1. [133] Ng, K. C. et al. 3d augmented mirror: a multimodal interface for string instrument learning and teaching with gesture support. In Proceedings of the 9th international conference on Multimodal interfaces, 339–345 (2007). [134] Blanco, A. D., Tassani, S. & Ramirez, R. Real-time sound and motion feedback for violin bow technique learning: A controlled, randomized trial. Frontiers in Psychology 12, 648479 (2021). [135] Provenzale, C., Di Tommaso, F., Di Stefano, N., Formica, D. & Taffoni, F. Real-time visual feedback based on mimus technology reduces bowing errors in beginner violin students. Sensors 24, 3961 (2024). URL https://www. mdpi.com/1424-8220/24/12/3961. [136] Çorlu, M., Muller, C., Desmet, F. & Leman, M. The consequences of additional cognitive load on performing musicians. Psychology of Music
BIBLIOGRAPHY 105 43, 495–510 (2015). URL https://journals.sagepub.com/doi/10.1177/ 0305735613519841. [137] Michałko, A., Campo, A., Nijs, L., Leman, M. & Van Dyck, E. Toward a meaningful technology for instrumental music education: Teachers’ voice. Frontiers in Education 7, 1027042 (2022). URL https://www.frontiersin. org/articles/10.3389/feduc.2022.1027042/full. [138] Ramirez-Melendez, R. & Waddell, G. Technology-Enhanced Learning of Performance. In McPherson, G. E. (ed.) The Oxford Handbook of Music Performance, Volume 2, 0 (Oxford University Press, 2022). URL https: //doi.org/10.1093/oxfordhb/9780190058869.013.24. [139] Meissner, H. & Timmers, R. Teaching young musicians expressive performance: an experimental study. Music Education Research 21, 20–39 (2019). URL https://www.tandfonline.com/doi/full/10.1080/14613808.2018. 1465031. [140] Lindström, E., Juslin, P. N., Bresin, R. & Williamon, A. “expressivity comes from within your soul”: A questionnaire study of music students’ perspectives on expressivity. Research Studies in Music Education 20, 23–47 (2003). URL https://journals.sagepub.com/doi/10.1177/1321103X030200010201. [141] Sloboda, J. Exploring the musical mind: Cognition, emotion, ability, function (Oxford University Press, 2005). [142] Meissner, H. Instrumental teachers’ instructional strategies for facilitating children’s learning of expressive music performance: An exploratory study. International Journal of Music Education 35, 118–135 (2017). URL https: //journals.sagepub.com/doi/10.1177/0255761416643850. [143] Haston, W. Beginning wind instrument instruction: A comparison of aural and visual approaches. Contributions to Music Education 37, 9–28 (2010). URL http://www.jstor.org/stable/24127224.
106 BIBLIOGRAPHY [144] Wulf, G., McNevin, N. & Shea, C. H. The automaticity of complex motor skill learning as a function of attentional focus. The Quarterly Journal of Experimental Psychology Section A 54, 1143–1154 (2001). URL https:// doi.org/10.1080/713756012. Publisher: SAGE Publications. [145] Ebie, B. D. The effects of verbal, vocally modeled, kinesthetic, and audiovisual treatment conditions on male and female middle-school vocal music students’ abilities to expressively sing melodies. Psychology of Music 32, 405–417 (2004). URL https://journals.sagepub.com/doi/10.1177/ 0305735604046098. [146] Lisboa, T. Action and thought in cello playing: An investigation of children’s practice and performance. International journal of music education 26, 243– 267 (2008). [147] Duke, R. A., Cash, C. D. & Allen, S. E. Focus of Attention Affects Performance of Motor Skills in Music. Journal of Research in Music Education 59, 44– 55 (2011). URL https://doi.org/10.1177/0022429410396093. Publisher: SAGE Publications Inc. [148] Wulf, G. Attentional focus and motor learning: a review of 15 years. International Review of Sport and Exercise Psychology 6, 77–104 (2013). URL https://doi.org/10.1080/1750984X.2012.723728. Publisher: Routledge. [149] Meissner, H. Theoretical framework for facilitating young musicians’ learning of expressive performance. Frontiers in Psychology 11, 584171 (2021). URL https://www.frontiersin.org/articles/10.3389/fpsyg. 2020.584171/full. [150] Bruner, J. S., Olver, R. R., Greenfield, P. M. et al. Studies in cognitive growth. (1966). [151] Hallam, S. & Papageorgi, I. Conceptions of musical understanding. Research Studies in Music Education 38, 133–154 (2016).
BIBLIOGRAPHY 107 [152] Woody, R. H. Explaining expressive performance: Component cognitive skills in an aural modeling task. Journal of Research in Music Education 51, 51–63 (2003). URL https://journals.sagepub.com/doi/10.2307/3345648. [153] McPherson, G. E., Osborne, M. S., Evans, P. & Miksza, P. Applying selfregulated learning microanalysis to study musicians’ practice. Psychology of Music 47, 18–32 (2019). URL https://journals.sagepub.com/doi/10. 1177/0305735617731614. [154] Waddell, G., Perkins, R. & Williamon, A. The Evaluation Simulator: A New Approach to Training Music Performance Assessment. Frontiers in Psychology 10 (2019). URL https://www.frontiersin.org/journals/psychology/ articles/10.3389/fpsyg.2019.00557. [155] Shulman, L. S. Those who understand: Knowledge growth in teaching. Educational Researcher 15, 4–14 (1986). URL https://journals.sagepub.com/ doi/10.3102/0013189X015002004. [156] Macrides, E. & Angeli, C. Domain-specific aspects of technological pedagogical content knowledge: Music education and the importance of affect. TechTrends 62, 166–175 (2018). URL http://link.springer.com/10.1007/ s11528-017-0244-7. [157] Macrides, E. & Angeli, C. Investigating tpck through music focusing on affect. The International Journal of Information and Learning Technology 35, 181–198 (2018). URL http://www.emerald.com/ijilt/article/35/3/ 181-198/136592. [158] Juslin, P. N. & Sloboda, J. Introduction: aims, organization, and terminology. In Juslin, P. N. & Sloboda, J. (eds.) Handbook of Music and Emotion: Theory, Research, Applications, 3–12 (Oxford University Press, Oxford, 2011). [159] Burnard, P. & Younker, B. A. Problem-solving and creativity: Insights from students’ individual composing pathways 22, 59–76 (2004). URL https:// journals.sagepub.com/doi/10.1177/0255761404042375.
108 BIBLIOGRAPHY [160] Coulson, A. N. & Burke, B. M. Creativity in the elementary music classroom: A study of students’ perceptions 31, 428–441 (2013). URL https: //journals.sagepub.com/doi/10.1177/0255761413495760. [161] Todd, J. R. & Mishra, J. Making listening instruction meaningful: A literature review. Update: Applications of Research in Music Education 31, 4–10 (2013). URL https://journals.sagepub.com/doi/10.1177/8755123312473609. [162] Swanwick, K. & Cavalieri Franca, C. Composing, performing and audiencelistening as indicators of musical understanding. British Journal of Music Education 16, 5–19 (1999). URL https://www.cambridge.org/core/product/ identifier/S026505179900011X/type/journal_article. [163] Macrides, E. & Angeli, C. Music cognition and affect in the design of technology-enhanced music lessons. Frontiers in Education 5, 518209 (2020). URL https://www.frontiersin.org/article/10.3389/feduc. 2020.518209/full. [164] Gabrielsson, A. The Relationship between Musical Structure and Perceived Expression (Oxford University Press, 2016). URL https://academic.oup. com/edited-volume/34489/chapter/292614599. [165] Pew Research Center. Americans’ use of mobile technology and home broadband (2024). URL https://pewrsr.ch/49i9IFc. [166] European Commission. E-communications survey (2021). URL https: //europa.eu/eurobarometer/surveys/detail/2232. Fieldwork conducted November 2020 - December 2020. Requested by Communications Networks, Content and Technology. Reference: SP510. [167] Thorgersen, K. & Zandén, O. The internet as teacher. Journal of Music, Technology amp; Education 7, 233–244 (2014). URL https:// intellectdiscover.com/content/journals/10.1386/jmte.7.2.233_1. [168] Chisholm, M. A. et al. Influence of interactive videoconferencing on the performance of pharmacy students and instructors. American Journal of Pharma-
BIBLIOGRAPHY 109 ceutical Education 64, 152 (2000). URL https://www.learntechlib.org/ p/90385. [169] Nguyen, T. The effectiveness of online learning: Beyond no significant difference and future horizons. MERLOT Journal of online learning and teaching 11, 309–319 (2015). [170] Cao, Y. Online courses for piano teaching: comparing the effectiveness and impact of modern and traditional methods. Education and Information Technologies 29, 11407–11419 (2024). URL https://doi.org/10.1007/ s10639-023-12283-6. [171] Chiu, T. K. & Hew, T. K. Factors influencing peer learning and performance in MOOC asynchronous online discussion forum. Australasian Journal of Educational Technology 34 (2018). URL https://ajet.org.au/index.php/AJET/ article/view/3240. Section: Articles. [172] Chiu, T. K. F., Lin, T.-J. & Lonka, K. Motivating Online Learning: The Challenges of COVID-19 and Beyond. The Asia-Pacific Education Researcher 30, 187–190 (2021). URL https://doi.org/10.1007/s40299-021-00566-w. [173] Stickland, S., Scott, N. & Athauda, R. A framework for real-time online collaboration in music production. In Proceedings of the ACMC2018: Conference of the Australasian Computer Music Association, 79–86 (2018). [174] Volioti, G. & Williamon, A. Recordings as learning and practising resources for performance: Exploring attitudes and behaviours of music students and professionals. Musicae Scientiae 21, 499–523 (2017). URL https: //journals.sagepub.com/doi/10.1177/1029864916674048. [175] Hamond, L. F., Welch, G. & Himonides, E. The Pedagogical Use of Visual Feedback for Enhancing Dynamics in Higher Education Piano Learning and Performance. OPUS; v. 25, n. 3 (2019): Volume 25, número 3. Setembrodezembro 2019.DO - 10.20504/opus2019c2526 (2019). URL https://www. anppom.com.br/revista/index.php/opus/article/view/opus2019c2526.
110 BIBLIOGRAPHY [176] Expressiveness in music performance: Empirical approaches across styles and cultures (Oxford University PressOxford, 2014), 1 edn. URL https: //academic.oup.com/book/27746. [177] Tejada, J. & Fernández-Villar, M. Design and Validation of Software for the Training and Automatic Evaluation of Music Intonation on Non-Fixed Pitch Instruments for Novice Students. Education Sciences 13 (2023). [178] Tejada, J., Murillo Ribes, A., Morant Navasquillo, R., Bernabé Villodre, M. d. M. & Fernández-Vilar, M. Music teachers’ perceptions of a new software for the assessment of musical instrument intonation (2024). URL https: //hdl.handle.net/10550/108659. [179] Ramirez, R. et al. Enhancing music learning with smart technologies. In Proceedings of the 5th International Conference on Movement and Computing, MOCO ’18 (Association for Computing Machinery, New York, NY, USA, 2018). URL https://doi.org/10.1145/3212721.3212886. [180] Upham, F. & McAdams, S. Activity analysis and coordination in continuous responses to music. Music Perception 35, 253–294 (2018). URL https://online.ucpress.edu/mp/article/35/3/253/91976/ Activity-Analysis-and-Coordination-in-Continuous. [181] Sun, L., Hu, L., Ren, G. & Yang, Y. Musical tension associated with violations of hierarchical structure. Frontiers in Human Neuroscience 14, 578112 (2020). URL https://www.frontiersin.org/article/10.3389/ fnhum.2020.578112/full. [182] Schubert, E. & Wolfe, J. Does timbral brightness scale with frequency and spectral centroid. Acta Acustica United With Acustica 92, 820–825 (2006). URL https://api.semanticscholar.org/CorpusID:8061120. [183] Melara, R. D. & Marks, L. E. Interaction among auditory dimensions: Timbre, pitch, and loudness. Perception & Psychophysics 48, 169–178 (1990). Place: US Publisher: Psychonomic Society.
BIBLIOGRAPHY 111 [184] Williamon, A., Waddell, G. & Moghnieh, A. D2.5 – user evaluations of each technology in each scenario. Tech. Rep. TELMI-D-WP2-RCM-20181031-D2.5, Royal College of Music, TELMI Project (2019). URL https://telmi.upf. edu/. Horizon 2020 Research and Innovation Action, Grant Agreement No. 688269. [185] Raffel, C. & Ellis, D. P. W. Intuitive analysis, creation and manipulation of midi data with pretty_midi. In Proceedings of the 15th International Conference on Music Information Retrieval Late Breaking and Demo Papers (2014). [186] Bjørndalen, O. M. & Doursenaud, R. Mido: MIDI Objects for Python. https: //mido.readthedocs.io (2023). URL https://mido.readthedocs.io. Version 1.3.2. [187] Bianchini, S., Hanappe, P. & Lee, J. FluidSynth: Real-Time Software Synthesizer. https://www.fluidsynth.org/ (2025). URL https://www. fluidsynth.org/. Cross-platform software synthesizer based on the SoundFont 2 specification. [188] McFee, B. et al. librosa: Audio and Music Signal Analysis in Python. https: //librosa.org/ (2023). Version 0.10.2.post1. [189] Kim, J. W., Salamon, J., Li, P. & Bello, J. P. Crepe: A convolutional representation for pitch estimation. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 161–165 (2018). [190] Zapata, J., Holzapfel, A., Davies, M., Oliveira, J. & Gouyon, F. Assigning a confidence threshold on automatic beat annotation in large datasets. In International Society for Music Information Retrieval Conference (ISMIR), 157–162 (2012). [191] Zapata, J., Davies, M. & Gómez, E. Multi-feature beat tracker. IEEE/ACM Transactions on Audio, Speech, and Language Processing 22, 816–825 (2014).
118 Appendix A. Appendix: MuSA Evaluation Study Results Table 18: Representative participant quotes by question and sentiment. Question Sentiment Quote Experience Positive (n7oh8onc) It was good! It definitely was better than just plain practice. I liked the highlighted regions where it detected potential inconsistencies. Experience Positive (ndrqs3sg) I think that the feedback is useful for engagement and increase the desire of improvement. It could be useful, at some point that you have the graphics and try to adjust to them in real time. Experience Negative (chxrj0wq) The feedback tool was not in real-time so I could not make immediate corrections. It was not very easy to adjust my pitch after the analysis tool. The feedback tool was useful in understanding where I was correct and whe. . . Experience Negative (y6v3g63s) The tool looks great but seems to have some problem with alignments between what I sang and what was supposed to be sang. Improvement Positive (y6v3g63s) For sure! Improvement Positive (b068bafw) Yes, I focused more on nuances, the hardest one was tempo because I didn’t know if I was rushing or not. Improvement Negative (oplv7i5e) Not really because for me auditory cues are more helpful than visual cues. Improvement Negative (3d1ly0bg) It’s hard to say, I hope there’s a control group that is just practicing and listening to their own audio, because I feel like the practice inherently helps. Highlights Positive (w4m19emj) Yes they were helpful for showing me where I need to focus. Highlights Positive (xzdbn2c2) Yes definitely. The pitch curve and dynamics curve were quite good choices for the plot. Highlights Negative (3tzisfce) They were, but sometimes I wouldn’t see any highlighted area even though I still considered the performance not perfect at all. Highlights Negative (3d1ly0bg) Yeah they were. Payment Positive (etuygx3k) If I were recording music I would definitely pay for it because I think it is very useful for improving but at the moment I don’t have a need for it. Payment Positive (ndrqs3sg) If my work implied singing, yes, but I would compare it first to other similar tools if exists. If singing is a hobby I would not pay for it. Payment Negative (w4m19emj) I would pay a maximum of 5$ a month. Payment Negative (oplv7i5e) No.
119 Table 19: Summary of sentiment analysis results, including overall statistics, perquestion means and medians, and correlation between lexicon and VADER sentiment. Category Sentiment Mean Median N / Notes Overall Sentiment Statistics Overall Lexicon 0.223 0.200 56 Overall VADER 0.370 0.458 56 Sentiment by Question Experience Lexicon 0.308 0.306 14 Experience VADER 0.671 0.795 14 Improvement Lexicon 0.186 0.191 14 Improvement VADER 0.286 0.402 14 Highlights Lexicon 0.242 0.200 14 Highlights VADER 0.282 0.421 14 Payment Lexicon 0.156 0.000 14 Payment VADER 0.241 0.095 14 Correlation Lexicon vs VADER Pearson r0.549 — —