Full text
Feature-Based Modelling of Perceived Emotion in Film Music Ruby Olivia Nagano Crocker1and Gyorgy Fazekas1 Queen Mary University of London, Mile End Road, London E1 4NS {r.o.n.crocker, g.fazekas} @qmul.ac.uk https://www.aim.qmul.ac.uk/ Abstract. This study investigates how musical features in film scores relate to perceived emotional expression over time. Film music di!ers from general music, as it is designed to manipulate narrative, tension, and audience perception, providing context-specific emotional cues beyond melody or harmony. The study is based on the FME-24 dataset, a collection of 300 professionally composed film score excerpts (2002–2024) covering both contemporary and traditional practices, through which perceptual, rhythmic, and tonal features were analysed. Participants marked moments of perceived emotional change, described the emotion, and placed it in a valence–arousal space. The preliminary results showed that emotions were clearly perceived but often di"cult to verbalize. Analysis showed weak but significant correlations for certain features (e.g. ZCR and arousal), while chord types influenced arousal more strongly. Rhythmic and tonal features showed varied relationships with both dimensions. Arousal was generally perceived more consistently than valence. Isolating audio enables more precise mapping between musical features and perceived emotion, establishing a baseline for future audiovisual comparisons. Variability in participant reports highlights the subjectivity of film music perception and supports further feature-based modelling of emotional dynamics in cinematic scoring. Keywords: Music Emotion Recognition ·Music Information Retrieval ·Film Composition 1Introduction Music plays a vital role in film, shaping narrative and enhancing emotional impact. Understanding how film scores evoke emotion requires examining the relationship between musical features and listener perception over time. This study focuses on perceived emotion, capturing how listeners interpret musical expression through time-stamped annotations of emotional change. We analyze these responses using the FME-24 dataset, a curated collection of 300 film score excerpts spanning 2002–2024 [6]. In music emotion recognition (MER), emotions are studied as either induced; listeners’ internal a!ective states measured via physiological or behavioral responses, or perceived, the emotions listeners attribute to the music. This study focuses on perceived emotion, providing rich, Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 434
2R.O.N.CrockerandG.Fazekas temporally aligned data for dimensional and semantic analysis. Although film music is usually paired with visuals, this study isolates audio to examine its standalone emotional impact. Research shows that music alone can evoke stronger emotions than video-only or combined contexts [28]. Removing visuals highlights musical structure, harmony, and rhythm, the core drivers of emotional narrative. It allows for investigation of how film scores remain emotionally powerful without cinematic context, although it may increase interpretive variability. Combining music theory, psychology, and music information retrieval (MIR), this research uses timestamped annotations, feature extraction, and semantic clustering to study how composers shape emotional trajectories. Since emotion in narrative film composition is underexplored in MIR, this study fills that gap, o!ering data-driven insights relevant to both research and creative practice. 2RelatedWork Film music is a powerful tool for conveying emotion, with research focusing on its impact on audience emotion perceptions. Unlike standalone music datasets, which emphasize listening in isolation, film music is composed to serve narrative, manipulating emotion, guiding perception, and providing context-specific cues beyond text or visuals [27]. To capture this function, the FME-24 dataset was developed and used in this study [6]. Recent work highlights a shift from classical traditions toward experimental and electronic approaches, particularly in genres like horror, where “non-musical” and atmospheric elements challenge conventions [1]. This has led researchers to critique genre-based classifications as overly restrictive, since modern scores often blend styles and transcend traditional categories [12]. Enabled by digital tools, contemporary award-winning films increasingly feature hybrid and electronic scoring, yet studies show these diverse practices evoke emotions as e!ectively as classical forms [20], underscoring that emotional impact is not limited to traditional structures. However, despite these stylistic developments, systematic modelling of how film music evokes emotion remains underexplored. 2.1 Emotion Models Emotion research uses discrete models that categorize basic emotions (e.g., fear, happiness) and dimensional models that represent emotions along continuous axes such as arousal and valence [26]. Music emotion studies mainly adopt these frameworks, with arousal-valence favored for its empirical support and clarity. Musical emotions often di!er from everyday emotions in quality and expression, especially in film music, which evokes narrative-driven feelings like tension, anticipation, and hope, terms less common in typical music vocabularies [24]. The FME-24 dataset reflects this distinct semantic profile, highlighting the need for models that capture film score emotions more precisely. The Geneva Emotional Music Scale (GEMS) o!ers nine music-specific emotion categories for a nuanced perspective beyond valence and arousal [2], but Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 435
Feature-Based Modelling of Perceived Emotion in Film Music 3 lacks integration with dimensional models, limiting its ability to track continuous emotional changes. This study centers on perceived emotions within the arousal-valence framework, enhanced by qualitative descriptions and semantic clustering to better capture film music’s emotional complexity. 2.2 Features Significant research in music emotion recognition (MER) focuses on identifying audio features linked to emotion, especially within the valence-arousal model. However, existing models often miss aspects of musical expressivity and texture [21]. MER features fall into two main types: audio-based (low-level acoustic properties like timbre, pitch, rhythm, intensity) and symbolic (higher-level structural elements like chord progressions and key clarity) [29]. Combining both provides afullerrepresentationofmusicalstyleandemotion. Key acoustic features include MFCCs, spectral centroid, brightness, and inharmonicity, which influence qualities such as sharpness and energy [13]. Rhythmic features such as tempo and beat strength a!ect perceived intensity and expressiveness [9]. Perceptual features, articulation, pitch salience, danceability, and chord information, also strongly shape emotional perception [11,22]. Tools like Librosa [19], Essentia [4], MIRToolbox [17], and Chordino [18] enable extraction of these features, with Chordino providing chord timestamps crucial for analyzing harmonic emotion [10]. This multidimensional feature approach underpins modelling of perceived emotion, especially in complex settings like film music where acoustic texture and structure shape emotional narratives. 3Methods This study investigates the relationship between perceived emotional changes and musical features in film music, using data collected via an online annotation task, followed by audio feature extraction and analysis. As well as a preliminary subset participant experiment and classification of the soundtrack type. 3.1 Online Data collection Data was collected via a secure web interface where participants listened to 15second excerpts from the FME-24 dataset. They pressed a button upon perceiving emotional changes, adjusted a color-coded dot on a valence–arousal graph, added brief text descriptions, and rated familiarity (see Figure 1) [6]. A tutorial on the valence–arousal space preceded the task, with an optional help button o!ering emotion adjectives from established mood models [16]. Quality-control measures ensured reliable responses, and after screening for outliers, N = 93 participants’ annotations were included in the analysis. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 436
4R.O.N.CrockerandG.Fazekas Fig. 1. Online Interactive Survey for Emotion annotation collection 3.2 Feature Extraction Audio features were extracted around participant-annotated emotional change points using two temporal windowing strategies. Change-Based Windowing involved two 2-second windows before and after each change (e.g., 0–2s and 2–4s) to capture directional shifts in musical properties. Center-Aligned Windowing used a single 1-second window centered on the change (0.5–1.5s), suitable for perceptual and rhythmic features needing contextual symmetry. This dual approach accommodates how low-level features reflect directional changes, while higher-level perceptual features represent broader states, consistent with known perceptual bu!er times for emotion recognition [3,23]. Acoustic Features Extracted via Librosa [19] and related toolkits, these low-level spectral and temporal descriptors include MFCCs (mean and variance of 13 coe"cients), spectral centroid, bandwidth, rollo!, contrast, zero-crossing rate, and chroma-STFT (12 pitch bins). They relate closely to perceptual qualities such as timbre, pitch content, and rhythm, with MFCC variance indicating emotional consistency or irregularities [5,13]. Perceptually Motivated and Rhythmic Features: To capture musical expressivity beyond acoustics, features including BPM, beat counts, chord types and transitions, inharmonicity (dissonance/tension [23]), pitch salience, and danceability (rhythmic regularity and energy [15]) were extracted. These span multiple timescales critical for film music’s emotional narratives: chord progressions and harmonic shifts signal emotional change; inharmonicity induces unease; pitch salience highlights key moments; tempo modulates scene pacing and mood; and danceability influences physical engagement [15]. Combined with acoustic feaProc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 437
Feature-Based Modelling of Perceived Emotion in Film Music 5 tures, these provide a comprehensive, explainable framework for modelling musical emotion in film scoring. 3.3 Data Analysis To explore how audio feature changes relate to arousal and valence shifts, we used Spearman’s rank correlation to measure relationships and Cohen’s d to quantify change magnitude. We extracted 43 audio features using Change-Based and Center-Aligned Windowing. Non-normal data distributions (confirmed by Kolmogorov-Smirnov tests) led us to choose Spearman’s correlation and Cohen’s dfore!ectsizearoundemotionaltransitions. Emotion Space Segmentation The arousal-valence space was divided into four quadrants: Q1 (High Arousal, High Valence), Q2 (High Arousal, Low Valence), Q3 (Low Arousal, Low Valence), and Q4 (Low Arousal, High Valence), plus a neutral point at (0,0) for ambiguous segments. Transitions between quadrants grouped feature changes to evaluate emotional shifts, focusing on rhythmic, tonal, and perceptual features. This method combined with Change-Based Windowing to compare feature dynamics before and after transitions. 3.4 Preliminary Participant Experiment In the preliminary subset, most participants noted emotional changes in Songs 3–8, while Songs 1 and 2 elicited more varied, ambiguous reactions. Beyond common emotions like “happy” and “relaxed,” they used imagery, movement, and texture to describe feelings. Some felt uncertain or “emotionally illiterate,” often resorting to non-emotional descriptions, highlighting annotation challenges. Responses from CBT therapists, speech therapists, and musicians showed no clear consensus. 3.5 Soundtrack Type Classification Tracks were labeled as original (1), sourced (0), or mixed (0/1), with most annotations for original scores (1375), then sourced (253), and mixed (52). Emotional intensity was measured by each annotation’s distance from the neutral point (0,0) in arousal–valence space, categorised into neutral, moderate, or strong zones. The hypothesis is that original scores, designed to support narrative and evoke emotion, elicit stronger responses than sourced tracks, consistent with prior research [28]. Results appear in Section 4.7. 4Results This section presents findings on how musical features a!ect emotional perception. It covers participant agreement on emotion ratings, links between features and emotional changes. It includes analyses of spectral, harmonic, rhythmic, and perceptual features, chord extraction, participant insights, and annotation consistency. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 438
6R.O.N.CrockerandG.Fazekas 4.1 Participant Emotion Agreement Consistency in emotional annotation was evaluated using range, interquartile range (IQR), cosine similarity, and Euclidean distance metrics across participants. Arousal ratings demonstrated stronger alignment than valence, with lower variation and higher cosine similarity, indicating greater consensus on changes in arousal. These findings align with previous research [8, 14]. See Table 1 for detailed values. Range IQR Cosine Sim. Euc. Dist. Arousal 1.934 0.54775 0.3256 0.1958 Valence 1.941 0.66025 0.2495 0.1940 Table 1. Comparison of Range, IQR, Cosine Similarity, and Euclidean Distance for Arousal and Valence 4.2 Spectral & Temporal Features Zero-crossing rate (ZCR) showed a weak but significant association with arousal (r=0.05, p=0.03) and a notable e!ect size (Cohen’s d = –1.17), suggesting that smoother, more harmonic segments tend to accompany heightened arousal. Other spectral features had minimal or non-significant correlations with either arousal or valence 2. Feat. SC_A pSC_A SC_V pSC_V Cd_A Cd_V ZCR 0.052 0.033 -0.016 0.501 -1.166 -0.258 S.Cont 0.028 0.258 0.017 0.475 -0.584 -0.579 S.Cent 0.026 0.294 0.013 0.589 0.041 0.041 S.Band 0.020 0.415 0.033 0.170 -0.003 -0.003 S.Roll 0.022 0.365 0.012 0.629 0.054 0.054 Table 2. Top statistically significant features from Librosa with Spearman’s Coe"cient and Cohen’s d. 4.3 Harmonic Content Chroma bins 1–2 exhibited slight positive correlations with arousal, whereas bins 3–12 tended toward negative correlations. Valence e!ects were minimal, suggesting harmonic shifts may subtly influence perceived energy more than mood. This highlights a potential role for pitch class emphasis in modulating emotional intensity around transitions. 4.4 Chord Types Changes in arousal and valence were analysed for one-second segments containing asinglechord,classifiedasaugmented,diminished,major,orminor.Table3 shows average emotional shifts, where negative values indicate decreases. Augmented chords caused the largest decreases in both arousal (–0.31) and valence, suggesting darker, lower-energy emotions. Diminished chords slightly Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 439
Feature-Based Modelling of Perceived Emotion in Film Music 7 Fig. 2. Spearman’s Coe"cient of Chroma Features for Valence and Arousal lowered arousal but increased valence. Major chords had minimal e!ects, while minor chords slightly increased arousal with little change in valence, di!ering from previous findings linking minor chords to decreased valence [25]. ANOVA revealed a significant e!ect of chord type on arousal change (p = 0.0036) but not on valence (p = 0.66) (Table 4). Aweightedinfluencescorecombiningmeanchange,variability,andfrequency was used to rank chords by their emotional relevance. Rare chords like Fωm7ε5 scored highly, though common chords such as D major and G major also ranked prominently (Table 5). Here, ϑAro µrepresents the mean change in arousal, and ϑVal ϖthe change in valence variability. Type Mean Aro Change Mean Val Change Augmented -0.3055 -0.1236 Diminished -0.1087 0.0640 Major -0.0214 -0.0043 Minor 0.0335 -0.0001 Other 0.0547 0.0279 Table 3. Average Arousal and Valence Change by Chord Type Chord Analysis F-stat p-value Arousal Change 3.92 0.0036 Valence Change 0.60 0.6614 Table 4. ANOVA Results for Arousal and Valence Change by Chord Type 4.5 Perceptual & Rhythmic Features High-level perceptual and rhythmic features, such as BPM, beat count, inharmonicity, pitch salience, and danceability, were analysed using the Change-Based Windowing method described in Section 3.2. This involved comparing 2-second segments immediately before and after each annotated emotional shift to assess feature evolution across emotion quadrant transitions. Key findings are summarised in the following sections, supported by tables. Danceability Spearman correlations (Table 6) revealed a weak but significant relationship between arousal variance and danceability variance (r = 0.074, p Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 440
8R.O.N.CrockerandG.Fazekas Chord ωAro µωAro ε ωVal µωVal εFreq W-Score F#m7b5 0.613 0.600 0.507 0.507 61.319 D0.034 0.401 0.017 0.382 148 1.319 G-0.033 0.389 0.067 0.466 124 1.306 C-0.038 0.425 -0.022 0.442 131 1.291 Cm 0.085 0.319 -0.042 0.476 85 1.254 Dm 0.017 0.379 0.052 0.410 99 1.249 A-0.105 0.449 -0.011 0.454 81 1.226 Gm -0.034 0.430 0.001 0.450 88 1.197 Bdim 0.376 0.372 0.374 0.247 81.195 Em 0.066 0.372 -0.027 0.477 68 1.184 Table 5. Chord E!ect on Valence and Arousal < 0.005), indicating that emotional volatility is modestly linked with variability in danceability. No significant correlations were found with mean arousal or valence. Danceability, as a high-level feature, was analysed for its correlation with changes in valence and arousal. Analysis of quadrant transitions (Table 7) shows that movements into higharousal states (particularly Q1) tend to increase danceability. For instance, transitions from Q4 →Q1 showed more increases (29) than decreases (12), suggesting that energetic emotional states may enhance the perception of danceability. S_C S_Cp Aro_Var 0.0740 0.0046 Aro_Mean 0.0127 0.6277 Val_Var -0.0332 0.2047 Val_Mean -0.0228 0.3824 Table 6. Spearman Correlations with p-values for Arousal, Valence, Danceability Mean and Variance Transition Decrease Increase (0,0) →Q1 45 58 Q1 →Q1 155 144 Q1 →Q4 24 28 Q2 →Q2 85 108 Q2 →Q1 36 36 Q3 →Q4 812 Q4 →Q1 12 29 Table 7. Danceability Mean Transitions by Quadrant BPM & Beat Count Quadrant transitions exhibited subtle but telling changes in rhythmic patterns. Beat count decreased during (0,0) →Q1 transitions (Table 8), suggesting that emotionally intense, positive states may simplify rhythm, while Q1 →Q4 transitions (positive →negative) showed more decreases than increases. BPM followed a similar trend: transitions into high-arousal quadrants like Q1 generally led to stabilization or reduction, whereas Q3 (low arousal/negative valence) transitions showed greater BPM variability, reflecting emotional instability (Table 9). Inharmonicity Table 10 indicates increased inharmonicity during transitions with negative valence (e.g., Q4 →Q2), although statistical significance was lacking, likely due to small sample sizes or high variability. Nonetheless, the trend suggests inharmonicity may heighten in emotionally dissonant states. Pitch Salience Pitch salience tended to increase in high-arousal conditions (Q1 and Q2), as shown in Table 11. For example, Q1 →Q1 transitions showed more increases (160) than decreases (139). Interestingly, shifts toward more positive valence, even with decreasing arousal (e.g., Q1 →Q4), also led to increases in pitch salience, underscoring its role in marking uplifting emotional changes. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 441
Feature-Based Modelling of Perceived Emotion in Film Music 9 Transition Decr Incr N/c (0,0) →Q1 52 20 38 (0,0) →Q2 27 12 20 (0,0) →Q4 20 11 8 Q3 →Q2 615 8 Q1 →Q1 124 107 68 Q2 →Q2 72 70 51 Q3 →Q3 18 32 18 Q1 →Q4 21 16 15 Table 8. Quadrant Transitions for Beat Counts Transition Decr Incr N/c Q4 →Q1 23 9 9 Q2 →Q3 25 7 5 Q2 →Q2 108 41 44 Q1 →Q1 172 72 55 Q2 →Q1 38 13 21 (0,0) →Q1 60 17 33 Table 9. Quadrant Transitions for BPM Transition Decr Incr Q1 →Q2 27 18 Q2 →Q1 32 17 Q3 →Q2 12 7 Q3 →Q4 512 Q4 →Q2 511 Table 10. Quadrant Transitions for Inharmonicity Transition Decrease Increase Q1 →Q1 139 160 Q2 →Q2 82 111 Q1 →Q4 20 32 Q3 →Q4 911 Table 11. Significant Pitch Salience Changes Across Quadrant Transitions 4.6 Preliminary Experiment Results Most participants noted emotional changes in Songs 3–8, while Songs 1 and 2 drew varied, ambiguous responses. Beyond common emotion words like “happy” and “relaxed,” participants used imagery (e.g., “sunrise”), movement (e.g., “gliding”), and texture (e.g., “bright”) to express feelings. Some felt “emotionally illiterate” or uncertain, often relying on visual or non-emotional descriptions, highlighting annotation challenges. Responses from diverse professionals (CBT therapists, speech therapists, musicians) showed no clear consensus. 4.7 Score vs. Soundtrack Results To compare emotional responses to original film scores versus sourced commercial soundtracks, annotations were categorized into neutral, moderate, or strong emotional zones based on Euclidean distance from the emotional midpoint (0,0) in arousal–valence space: below 0.2 (neutral), 0.2–0.5 (moderate), and above 0.5 (strong). Tracks were labeled as original (0), sourced (1), or mixed (excluded for clarity). A chi-squared test showed a significant di!erence (p=0.0378)in emotional intensity distributions between soundtrack types. Original scores had Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 442