Investigating instrument identification in a virtual orchestral scene
Abstract
Preprint, supplementary materials, data sets, and scripts to the article "Investigating instrument identification in a virtual orchestral scene" re-submitted to the journal Music Perception in September 2025. Authors: Simon Jacobsen, Carl von Ossietzky Universität Oldenburg, Germany Félix Baril, McGill University, Canada Giso Grimm, Carl von Ossietzky Universität Oldenburg, Germany Kai Siedenburg, Graz University of Technology, Austria & Carl von Ossietzky Universität Oldenburg, Germany
Full text
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 1 Investigating instrument identification in a virtual orchestral scene Simon Jacobsen Carl von Ossietzky University of Oldenburg, Oldenburg, Germany. Félix Baril McGill University, Montréal, Canada. Giso Grimm Carl von Ossietzky University of Oldenburg, Oldenburg, Germany. Kai Siedenburg Graz University of Technology, Graz, Austria & Carl von Ossietzky University of Oldenburg, Oldenburg, Germany.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 2 Abstract Musical scene analysis (MSA) in an orchestral context is affected by the spatial positioning of instruments on stage, the reverberation of the concert hall, and the musical structure itself. Although studies have investigated the perceptual blending of instruments in musical mixtures, instrument identification has mostly been studied for sounds in isolation. How the acoustic and musical variables affect identification performance remains unclear. In the present study we used a cue-mixture identification paradigm to have participants identify eight sustained instruments in twelve short four-instrument mixtures taken from Wagner’s Tristan prelude. Excerpts were rendered with and without concert hall acoustics using a room-acoustic modeling software with instruments either spread out on stage or all in a single center location. Against initial hypotheses, reverberation of the virtual concert hall and spatial distribution of instruments did not affect identification performance, despite clear effects on sound quality ratings. On the other hand, instrument-dependent musical effects of voice, register, and acoustical content were observed. These results hint at instrument identification performance showing robustness to variations in acoustic conditions and highlight the complex nature of instrument perception in realistic sound mixtures. Keywords: instrument identification, musical mixture, room acoustics, virtual acoustics, timbre
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 3 Introduction In music listening, the perception of the mixture of sounds from instruments playing simultaneously is impacted by a host of influencing factors. Imagine a listener sitting in a spacious concert hall with a large orchestra on stage playing a symphony. While the listener closes their eyes they may – or may not – pick up on the different instruments that make up the sound mixture. This picture already gives us a hint at what these influencing factors might be. Firstly, the instruments – as separate sound sources – are distributed on stage and the sound waves they produce reach the listener from different directions. Secondly, not only the direct sound is perceived by the listener. The walls (and ceiling) of the concert hall create reflections of the instruments' sound waves of which early ones overlap with the direct sound depending on the position of the listener. Later reflections add to a sense of space by creating reverberation inside the hall. Critically, not only the acoustic environment shapes the listener's perception but also the music itself. Changes in orchestration and the structured interplay of instruments affect what is audible – of course based on the listener’s prior knowledge and experience with the music at hand. Many studies have focused on the individual aspects that shape this sonic experience. It is well understood that the acoustics of concert halls shape the listener's perception (e.g., Beranek, 2012; Lokki and Pätynen, 2020). Concerning instrument perception, Goad and Keefe (1992) have shown that variations in the acoustic properties of performance spaces affect timbre discrimination. Kato et al. (2014) reported variations in perceived timbral brightness of clarinet tones when convolved with two varying binaural room impulse responses. In a recent study, effects of room acoustic parameters were explored in the perception of musical blending (Thilakan et al., 2025). Sounds of two violins with varying levels of source level blending were rendered in different room acoustic environments that differed in their volumetric size, absorption behavior, and listener position. The authors found
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 4 significant effects for all variables, with source level blending contributing 60% and roomacoustic parameters 40% to the prediction of subjective blend ratings using a random forest regression model. Together, these studies demonstrate the influence room acoustics can have on the perception of instrument sounds and music. Another factor in our concert scenario is the placement of instruments on stage. Its influence on the parsing of individual sound sources can be described by auditory scene analysis (ASA; Bregman, 1990). In ASA, sound source separation is the key objective. This is governed by perceptual grouping mechanisms that are a described by Gestalt psychology. Sound events that are similar and follow specific organization principles tend to be perceived as one item. Important principles include proximity, closure, common fate, and continuity. Most of these grouping principles are results of primitive (bottom-up) processing in the auditory system. There also exists schema-driven (top-down) processing based on prior knowledge, memory, and attention. Spatial distribution as a principle of ASA allows for better sound source separation, at least for synthetic sounds (van Noorden, 1975). Together with other factors influencing the grouping of sounds, it might not be as strong of a cue (Bregman, 1990). However, other studies on the spatial release from masking (e.g., Litovsky et al., 2021) do attribute spatial distribution as a relevant cue to scene parsing. The third factor is the music itself. What becomes audible is governed by its orchestration and guided by ASA. In multi-source scenarios, simultaneous sounds from different instruments can be perceived as single fused entity (McAdams, 2019). This auditory fusion of instrument sounds is also called timbre blending and can lead to sounds with unique timbres (Sandell, 1995). Studies showed correlations of subjective blend ratings with acoustical descriptors such as energy maxima, spectral center of gravity, onset synchrony, but also crossing of voices and consonance (Reuter, 1996; Tardieu and McAdams, 2012; Lembke et al., 2017) – all of which are results of the individual instruments' sounds and their
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 5 orchestration. More recent studies investigated instrumental blends in complex orchestral excerpts and found several factors – score-based and acoustic – that influenced participants' blend and segregation ratings (Fischer et al., 2021; McAdams et al., 2025). These aspects of instrument timbre and orchestral scene perception are mostly measured using subjective ratings scales that can be prone to bias. Another way to determine what a listener hears out from an orchestral mixture can be realized through an instrument identification task as an objective measure. By asking whether an instrument was in the mixture or not – or by directly asking for the specific instrument – one obtains a binary decision, which can either be correct or incorrect. This yields a simple measure that does not depend on a subjective rating scale. Instrument identification as a perceptual listening task has already been studied extensively. In one of the first published studies, Eagleson and Eagleson (1947) claimed that identification of musical instruments is a rather “trivial” task. They had participants identify nine instruments directly and over a public address (PA) system. Identification accuracy was between 69% and 94% when heard directly and dropped by more than 10 percentage points when heard over the PA system (55% – 82%). Srinivasan et al. (2002) provided a baseline of identification performance by conservatory students using closed-sets of two, three, nine, and 27 instruments. Identification accuracies for the first three sets was very high with above 90% but dropped to around 56% for the large set of 27 instruments. There exist several other studies on instrument identification (Saldanha and Corso, 1964; Berger, 1964; Strong and Clark, 1967; Elliott, 1975), all of them reporting different overall identification accuracies for unaltered sounds in isolation. The selection and number of instruments to be identified and the varying expertise of participants (non-musicians, musicians) explain such differences. These early studies – but also recent ones – were concerned with the spectro-temporal ingredients that shape identification performance (Patil et al., 2012; Thoret et al., 2017;
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 6 Siedenburg, 2019; Siedenburg et al., 2019). Although stationary spectral envelopes of instruments' sounds are a major distinguishing factor, temporal features, such as, e.g. sounds' onsets also play a critical role in instrument identification. All of these studies, however, were constrained to instrument sounds in isolation. To our knowledge, there are no previous studies on instrument identification in sound mixtures. Hence, whether the found factors that affect identification in isolation are of similar importance in and generalize to instrument mixtures remains unclear. Furthermore, general aspects of room acoustics such as reverberation and spatial distribution have not yet been investigated in the light of instrument identification. Thus, in the present study, we investigate instrument identification in a virtual acoustic environment by using a novel cue-mixture identification task. As source material we used a realistic simulated score from Richard Wagner's prelude to the Opera Tristan und Isolde. Participants were instructed to hear out and identify a single instrument from a mixture by being provided a melodic line as a cue from the mixture. To consider varying acoustic conditions that might affect identification abilities, we present the instrument mixtures with and without reverberation and with instruments spread out or all centered on a stage in a virtual acoustic scene rendered via a loudspeaker array. Our focus on instrument identification was motivated not only by the need to extend the current state of research on identification to multi-source scenarios but also to complement findings on timbre blending. The extent to how well instruments blend is associated with their identifiability. Kendall and Carterette (1993) found that increasing blend correlated with decreasing identification in simultaneous orchestral wind instrument sounds. This inverse relationship of blend and identification was later formulated in more detail by Sandell (1995). Thus, an identification task could provide an objective way to capture blend ratings. Besides
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 7 the identification task, we also gathered sound quality ratings to assess the acoustic changes to the virtual scene qualitatively and provide a comparison for the audibility of these changes. Method This experiment consists of the main instrument identification task and a sound quality rating task to compare the quantitative identification results with qualitative measures. Participants The experiment had 21 participants. All participants were recruited through the online job board of the Carl von Ossietzky University of Oldenburg with the prerequisites of selfreported normal hearing and at least five years of continuous professional training on a musical instrument. The participants had a mean age of 25.0 years (SD = 3.3). In addition to the experiment, the participants had to complete two questionnaires on their musical perception and training abilities based on the Goldsmith Musical Sophistication Index (GoldMSI, Müllensiefen et al., 2014) The average scores were 4.69 out of 7 (SD = 1.84) for the perception questionnaire and 4.53 out of 7 (SD = 1.96) for the training questionnaire. Stimuli and task The stimuli consisted of short excerpts from the prelude to the opera Tristan und Isolde by Richard Wagner. A synthesized multitrack stereo mix from the OrchPlay Music Library (www.orchestraplayer.com) was used. It features hybrid instrument recordings of individual notes that were used to create a realistic mix with particular detail to intonation, dynamics, tempo, and blend between instruments. These features were hand-sculpted by a composer with a doctorate in music composition (co-author FB). The synthesized mix contained individual single-note instrument recordings from two sound libraries: the Berlin series from Orchestral Tools (OT) and the Vienna Symphonic Library (VSL). Both libraries do not
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 8 provide anechoic stimuli and there are differences in the level of dryness in the individual synthesized instrument tracks. Figure 1: (A) cue-mixture identification task. (B) Simplified overview of the dimensions of the shoebox-sized virtual concert hall in TASCAR and the orchestral seating configuration. Instruments used in the identification task are color-coded, additional full orchestration (sound quality rating task) in white. Connecting lines between the individual string sections visualize the stereo panning. The centered dashed circle corresponds to the positioning of instruments in the center condition.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 9 The identification task consisted of a cue and a mixture containing a target instrument, as visualized in Figure 1A. The sinusoidal cue provided tonal information repeated by the target instrument in the mixture. Participants were then asked Which instrument played the cued melody. As of our knowledge, no studies on instrument identification have used such a cuemixture paradigm. Previous studies on musical pattern or instrument detection have used a target-in-mixture approach (Bey and McAdams, 2002; Bürgel et al., 2021) where a target pattern or instrument had to be detected within a mixture of patterns or instruments – with the target being presented either before or after the mixture. This approach results in a binary Y/N decision. Although our identification task was not a binary decision because it involved more than two instruments, we extracted a binary scoring (correct, incorrect) for identification results. Other approaches of asking participants what instruments they heard in the mixture would not strictly result in such a binary scoring and would require a criterion for correct or incorrect labels of responses for mixtures with multiple instruments. Hence, we opted for the more streamlined and novel cue-mixture paradigm, essentially requiring two perceptual processes: the matching of the pitch information of the target in the mixture and the subsequent (or even parallel) identification of the instrument. The Toolbox for auditory scene creation and rendering (TASCAR; Grimm et al., 2019) was used for rendering the orchestral scene in a virtual acoustic environment. The toolbox has been used in many auditory listening studies (e.g., Gerken et al., 2014; Hendrikse et al., 2019). TASCAR uses a mirror source model for the early reflections and a feedback-delay network for diffuse late reverberation. The Sendesaal Bremen, Germany, was used as a concert hall model. It features individual angled wall elements while resembling the physical specifications a shoebox-shaped hall. Concert halls of this type are considered among the best concert halls in the world (Beranek, 2012) and comparatively straight-forward from the perspective of virtual acoustic simulation. The damping and reflection parameters of the
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 16 by-measure intercepts but no by-instrument intercept. As fixed effect, the model included the predictor of instrument. The model results confirmed that there was a main effect of target instrument (χ2 = 269.30, p < .001), that is, there were significant differences in identification accuracy between individual instruments. Where the violins (M = 0.65 [0.56, 0.73]), flute (M = 0.59 [0.50, 0.67]), and cellos (M = 0.54 [0.44, 0.63]) showed the highest accuracies, other woodwinds were below average with the lowest accuracy for the oboe (M = 0.18 [0.13, 0.23]), but still above chance. A full overview of identification accuracies across excerpt and instrument are included in the supplementary materials (see Table 3). Figure 3B depicts the confusion matrix for the identification task. Columns represent the true target instruments and rows represent the identified instruments. The matrix is normalized by column to indicate the proportion of identifications for each target instrument. The diagonal reflects the correct identifications and corresponds to the accuracy values in Figure 1A. Offdiagonal entries represent confusions between individual instruments and within instrument groups. Strong confusion occurred for most woodwinds (oboe, English horn, clarinet) with the oboe being misidentified as an English horn more often than correctly identified. Also, the bassoon showed moderate confusion with the oboe, English horn, and clarinet. Only the flute, belonging to the woodwinds, broke this pattern and was mostly confused with the sound of the French horn. The French horn was in turn confused with the bassoon. There was also confusion among the strings where cellos were strongly confused with the violins, but the violins only occasionally with the cellos. Overall, these results highlight strong variability in identification accuracy across instruments. Focus of this work were possible influences of spatial distribution of instruments on stage (center/spread) and the reverberation of the virtual concert hall (off/on) on the identification performance. Figure 3C depicts the identification accuracy across acoustic condition as the interactions of spatial distribution and reverberation. Surprisingly, there are no visible
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 17 differences between the acoustic conditions. The acoustic data were analyzed using a binomial GLMM with the predictors of spatial distribution (center, spread), reverberation (off, on), and their interaction term as fixed effects (R2cond = .20, R2marg = .00). The model results confirmed no main effect of spatial distribution and reverberation (χ2 = 5.39, p = .145) with the marginal means center/off: M = 0.38 [0.25, 0.52], center/on: M = 0.36 [0.24, 0.50], spread/off: M = 0.37 [0.25, 0.51], spread/on: M = 0.41 [0.28, 0.56]. There was, however, a marginal interaction effect (β = 0.06 [–0.03, 0.49], p = .078), reflected in the increased identification accuracy in the spread/on condition. These results show that – against initial hypotheses – instrument identification performance was robust to variability in the virtual acoustic conditions. Figure 4: Training effect on participants’ identification performance as identification accuracy across trial number on a logarithmic axis. Thin continuous lines depict the moving averages (10-trial window) including 95% confidence intervals as shaded areas. Thick dashed lines depict the fixed effects from the model output of a GLMM including 95% confidence intervals as shaded areas. We also investigated a possible training effect on the participants’ identification performance. Figure 4 depicts the mean identification accuracy across participants over the time course
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 18 (trial number) of the experiment on a logarithmic axis as moving averages (10-trial window). Over all trials (and conditions), identification accuracy increased with increasing trial number. Within the different acoustic conditions, the time course varied. The time course data were analyzed using a binomial GLMM with the predictors of trial number (scaled), spatial distribution (center, spread), reverberation (off, on) and their (three-way) interactions as fixed effects (R2cond = .22, R2marg = .02). The extracted fixed effects are overlayed in Figure 4 as dashed thick lines. The model results confirmed a main effect of trial number (β = 0.17 [0.10, 0.25], p < .001). That is, participants improved in the identification task over the course of the experiment. Although two-way interactions between trial number and spatial distribution or reverberation, respectively, showed no effect (β < –0.01, p > .058), the three-way interaction term was significant (β = –0.15 [–0.22, –0.07], p < .001). However, these results are difficult to interpret. It seems that the condition with no reverberation and instruments distributed across stage led to the strongest increase in identification accuracy over time. Likewise, its inverse with instruments at a center location and the presence of reverberation also showed a strong increase. Curiously, both conditions showed lower initial accuracies compared to the “realistic” acoustic condition, with instruments distributed and reverberation on, as of their moving averages and the model fit. The most “unrealistic” condition (instruments spread, reverb off) also showed low accuracies over the first 30 trials, but the fixed effect extracted from the model seems to flatline. Since each mixture consisted of four instruments, the excerpts were ordered by voice from highest (voice one) to lowest (voice four) – although not all voice configurations were clear due to the crossing of voices. Figure 3D shows the mean identification accuracy across voice from lowest to highest in terms of pitch height. Identification accuracy for outer voices (voice one: M = 0.44 [0.41, 0.46]; voice four: M = 0.38 [0.35, 0.41]) stood out from the inner voices (voice two: M = 0.33 [0.30, 0.36]; voice three: M = 0.29 [0.26, 0.31]). The voice data were
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 19 analyzed using a binomial GLMM. As fixed effects, the model included the predictor of voice as linear and quadratic terms (R2cond = .19, R2marg = .01). The model confirmed that there was a main effect of voice, both as linear (β = –0.90 [–1.35, –0.45], p < .001) and quadratic terms (β = 0.17 [0.08, 0.25], p < .001). That is, identification accuracy was affected considerably by the target instruments' note positions within the mixture. However, there was great variability in identification accuracy across target instruments – as already seen in the overall and instrument identification performance. Random effects variance in the model for the target instruments was 0.48, considerably larger than the variances for participants (0.18) and excerpts (0.11). This justifies a closer look at the orchestration for each measure.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 20 Figure 5: Mean identification accuracy for each instrument per excerpt as filled markers across voices within the excerpt. Error bars indicate bootstrapped 95% confidence intervals. Smaller empty markers indicate the top confused instrument if identification of the target instrument was below chance level (0.125). Transparent bars correspond to the acoustic conditions as interactions of spatial distribution and reverberation in the order center/off, center/on, spread/off, and spread/on. Figure 5 shows identification accuracies for mixture instruments in each excerpt across voice. Error bars indicate the 95% CIs and transparent bars correspond to the underlying accuracies for the interactions of spatial distribution and reverberation. Overall, there is substantial variability in accuracy per excerpt and instrument. Some excerpts exhibit the previously analyzed effect of the outer voices (excerpts M12, M41, M56, M74), some show a flat (M2, M10) or mixed trend, and others even display an inverse effect (M16, M34), where the
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 21 middle voices are more accurately identified. Furthermore, the plot shows the top confused instrument if its identification accuracy was above that of the correct target instrument. Depending on the voice within the excerpt, instruments were often confused with their lower counterparts, such as, e.g. the oboe in voice three with the English horn (M6, M10) or the cellos in the second voice with the violins (M72). Although spatial distribution of instruments and reverberation did not have an overall effect on the identification performance, individual trends can be seen from the transparent bars for each instrument (Fig. 5) reflecting the acoustic configuration instruments center/reverb off, center/on, spread/off, and spread/on. Cellos always benefited when being presented at their expected position on the right of the stage, opposite from the violins (M56, M72, M74). When both string instruments appeared in the same excerpt (M72) also the violins showed this improvement from spatial distribution. In the center condition, the accuracies were substantially lower. Other instruments were affected by reverberation as seen, e.g., for the French horn (M10, M72), where added reverberation led to an increase in identification accuracy. Other instruments showed mixed effects, and the flute seemed to be unaffected from the acoustic conditions. Again, these results highlight the strong variability individual instrument sounds impose on participants' identification performance. Additionally, a possible effect of instruments' F0-register was investigated since it can strongly affect an instrument's timbre. We argued that identification accuracy in instruments' middle register is higher as participants would be more familiar with more “common” sounds. Figure 3E shows the identification accuracy for low, middle, and high F0-registers. Mean identification accuracies indeed varied between the registers with the middle register showing the highest identification accuracy. Based on the given excerpt, not all instruments' registers were represented. Low register included oboe, clarinet, and French horn, middle register included all instruments except the clarinet, and high register included flute, oboe,
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 22 clarinet, bassoon, and cellos. A binomial GLMM was used to analyze the effect of instrument register on the identification accuracy. As a fixed effect, the model included the predictor of register as a factor with the levels low, middle, and high. The model confirmed that there was a significant effect of register (χ2 = 60.07, p < .001) even if the target instruments were included as random effects. Identification accuracy was significantly higher in the middle register at 0.47 [0.33, 0.61]. Marginal means for the low and high register were 0.23 [0.14, 0.36] and 0.30 [0.19, 0.44]. This result shows that instrument register did affect participants' instrument identification performance, likely due to a higher familiarity with sounds from instruments' middle registers. We also checked whether the register effect was a pseudo effect as a result of an underlying effect of pitch on a global level. For this, we computed a binomial GLMM with the predictor of pitch as linear and quadratic terms as well as their interaction. We did find an effect of pitch (χ2 = 162.15, p < .001) as linear (β = –0.39 [–0.59, –0.19], p < .001), quadratic (β = 0.12 [0.03, 0.21], p < .01), and interaction terms (β = 0.28 [0.23, 0.34], p < .001). It even outperformed the register model (χ2 = 102.08, p < .001), that is, pitch explained more variance in identification accuracy than register. However, when combining both register and pitch (linear, quadratic, and interaction terms) as fixed effects in a single model, in turn, it outperformed the pitch model (χ2 = 58.79, p < .001). Overall, although pitch (on a global level) is a better predictor for identification accuracy than register (on instrument level), the effect of register remains significant even after accounting for pitch. We investigated a potential relationship of identification and the acoustic similarity between instrument sounds. For each isolated instrument within a mixture of an excerpt, we extracted spectral centroid (SC) as a measure of the spectral center of gravity, spectral flux (SF) as a measure of spectral changes over time, and Mel-frequency cepstral coefficients (MFCCs, coefficients no. 2-13) as a measure of coarse spectral shape from 100 ms analysis windows.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 23 We conducted comparisons between the target and all three other instruments within each mixture by computing the average SC heights of target-instrument pairs, the Pearson correlation coefficient of their SF time series, and the mean of the Pearson correlation coefficients of their MFCC time series. For each target instrument, we computed two summary statistics per feature: the mean value across all three possible pairings, and the extreme value (minimum for average SC height, maximum for SF and MFCC correlations). That is, we compared the acoustic similarity of a specific instrument to either the average across all other mixture instruments (“average model”) or only the single most similar one within that mixture (“single model”). We then used two separate GLMMs with average SC height, SF correlation, and MFCC correlation as fixed effects predictors, as well as participant, excerpt, and target instrument as random effects to predict identification accuracy. Both models showed significant effects of all feature summary statistics (average model: χ2 = 128.71, p < .001; single model: χ2 = 101.99, p < .001), with the average model (R2cond = .50, R2marg = .20) describing more variance than the single model (R2cond = .44, R2marg = .20). There was a significant effect of average SC height in both models (average: β = 0.86 [0.61, 1.11], p < .001; single: β = 0.66 [0.46, 0.88], p < .001). That is, identification accuracy indeed increased with increasing average SC height. Also, SF showed significant effects in both models (average: β = –0.76 [–0.96, –0.55], p < .001; single: β = –0.75 [–0.97, –0.53], p < .001). That is, identification accuracy decreased with increasing SF correlation. MFCC correlation showed a significant effect in the average model (β = –0.31 [–0.45, –0.17], p < .001) but only a marginal effect in the single model (β = –0.21 [–0.38, –0.05], p = .011). That is, identification accuracy decreased with increasing MFCC correlation. Overall, these results show that acoustic similarity between the target and mixture instruments captured a substantial amount of variance seen in the identification data.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 24 Figure 6: Sound quality ratings across acoustic condition for the rating scales sonority, spatiality, and overall, with binned individual participant ratings and bootstrapped 95% confidence intervals as error bars. Sound quality Sound quality ratings of longer full-mixture excerpts were gathered using five 7-point Likert rating scales for room-acoustic and overall attributes of the sound. All rating scales were analyzed using individual LMMs. The conditional and marginal statistics for the five models were transparency: R2cond = .23, R2marg = .07; sonority: R2cond = .41, R2marg = .11; spatiality: R2cond = .35, R2marg = .21; concert hall: R2cond = .32, R2marg = .16; overall: R2cond = .26, R2marg = .12. Based on a Pearson correlation analysis on the residuals of each model the scales showed moderate to strong correlations between each other (Pearson's r(838) > .44, p < .001). Since the overall rating was strongly correlated with both the concert (r = .73) and transparency rating (r = .70), only the overall, sonority and spatiality ratings are analyzed in detail. Nonetheless, all five models showed main effects of reverberation (β > 0.21, p < .001) and spatial distribution (β > 0.35, p < .001). Also, their interaction effects were significant (β < – 0.16, p < .002).
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 25 Figure 6 shows the sound quality ratings across acoustic conditions for the sonority, spatiality, and overall rating scales. Results for the transparency and concert hall rating scales are included in the supplementary materials (see Figure 1). Mean sonority ratings (How thin/full was the sound?) in each condition were above neutral, indicating that the sound was generally perceived as full. A significant effect of both reverberation (β = 0.32 [0.29, 0.46], p < .001) and spatial distribution (β = 0.37 [0.24, 0.41], p < .001) individually lead to a strong increase in ratings (M = 5.20; M = 5.16) compared to the center and reverb off condition (M = 4.07). Their interaction, however, did not increase the ratings much further (spread/on: M = 5.35). For the spatiality ratings (How narrow/wide was the sound?) the picture is a bit different. Spatial distribution had the strongest effect (β = 0.73 [0.63, 0.83], p < .001) with a large increase in ratings from center/off (M = 3.42) to spread/off (M = 5.27) indicating the perception of the sound to be wider. However, also reverberation had a significant but smaller effect on the perceived width (β = 0.30 [0.20, 0.39], p < .001). Again, the interaction of spatial distribution and reverberation did not lead to an additional increase in ratings (M = 5.43). The overall ratings (Overall, how much did you like the sound?) showed similar effects of reverberation (β = 0.21 [0.12, 0.30], p < .001) and spatial distribution (β = 0.42 [0.33, 0.50], p < .001) leading to increases in ratings for the center/on (M = 4.61) and spread/off (M = 4.97) conditions compared to the center/off conditions (M = 3.75). Like for the other scales, there was no further increase in ratings for the spread/on condition (M = 4.97). Besides the main effects of reverberation and spatial distribution, musical effects of excerpt, dynamics, and orchestration were analyzed. Effects of excerpt were seen for all ratings scales (χ2 > 35.26, p < .001). Especially the sonority ratings differed from excerpt to excerpt (χ2 = 209.37, p < .001). This excerpt dependence can be attributed to changes in dynamic playing level. Namely, effects of dynamics were seen for the sonority ratings (χ2 = 21.80, p < .001),
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 32 but also music. Rendering of higher-order early reflections by geometrical considerations (e.g., use of mirror sound sources) produce physically accurate models of concert hall acoustics. However, this usually requires the use of anechoic recordings of musical instruments to accurately render reflections on the direct sound path. In our case, the multitrack recordings were not anechoic. This could lead to spectral alterations and changes in timbre, masking of onset transients, and smearing of the spatial cues. However, this might be less pronounced for sustained instrument sounds than for percussive sounds with sharper onsets. Furthermore, musical instruments have individual complex directivity patterns (Meyer, 2008). That is, the direct sound does not radiate outwards in a purely spherical pattern. Measured directivity patterns were found to be highly dependent on the dynamic playing level, the pitch produced by the instrument, the observed frequency band, and movements by the musicians (Ackermann et al., 2024). In a symphonic orchestral setting, woodwind and brass instruments are usually solo instruments, that is, they play individual parts. In contrast, the strings (e.g. violins, cellos) mostly play in unison as one section to create a tutti sound. The multi-track sounds used in our study did not include individual instrument sources for the string sections. In an ideal setting, each instrument in their section would be represented as an individual sound source. However, this also results in a loss of source-level blending between the individual instruments. To simulate the auditory width of an entire violin section, we took the left and right channels of the stereo sound and placed them separately on the virtual stage to create a panned impression. Another issue might be the unbalanced set of instruments to be identified, a consequence of using natural stimuli. Our selection of excerpts did not result in an equal distribution of instruments, that is, e.g. excerpts with French horn were overrepresented, those with cellos
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 33 underrepresented. Yet, this imbalance was accounted statistically for by including random byinstrument intercepts in our GLMMs. Although our manipulations to the acoustic scene were clearly audible – as reflected in the sound quality ratings – the extent to which these “realistic” acoustic conditions affected identification performance might be underexplored; more extreme manipulations to the spatial configuration and reverberation might provide a clearer picture and help consolidate our notion of robustness of instrument identification to acoustic variations. Such extreme manipulations could entail spatial separation of instruments of up to 180 degrees (or even more) and comparisons of anechoic source material with reverberation times of well over two seconds. Furthermore, our task design might have affected participants’ susceptibility to the acoustic variations. The identification paradigm was based on hearing out a musical cue from the mixture and subsequently naming the instrument that played said cue. Future studies could also incorporate a different task in which participants simply have to name all instruments they heard in the mixture. This approach would help disentangle possible conflicts of task design and acoustic cue relevance. Since our results showed strong differences in identification accuracy across instruments and random variance explained by target instrument was high, a follow-up study using controlled musical stimuli should be conducted. In this way, the hierarchy effect of certain instruments based on their voice within excerpts – as a direct result of orchestration practices – could be avoided. Furthermore, analyses using spectro-temporal features should be used to further investigate the present musical effects to ideally construct a computational model of instrument identification in complex musical mixtures.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 34 Conclusion A novel cue-mixture task was used as a tool to quantify instrument identification in realistic musical excerpts and to inform the blending of instruments within sound mixtures. Our results show that neither spatial distribution nor reverberation had a significant effect on identification performance in simulated room acoustics, whereas sound quality ratings were clearly affected. In contrast, musical factors – including an instrument’s register, voice, and pitch – together with timbral features reflecting acoustic similarity, significantly affected identification. Highlighting the complex nature of music perception in natural orchestral scenes, these results suggest that musical instrument identification remains robust even under varying room-acoustic conditions. In other words, the acoustical richness of a scene such as the availability of spatial cues may not be directly beneficial for the accurate inference of musical sound sources but may be interpreted primarily to affect auditory “comfort” and thus judgments of sound quality. It remains to be consolidated whether these two aspects of music perception, ASA and the perception of sound quality, are two sides of the same coin or constitute separate dimensions of musical experience.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 35 References Ackermann, D., Brinkmann, F., & Weinzierl, S. (2024). Musical instruments as dynamic sound sources. The Journal of the Acoustical Society of America, 155(4), 2302–2313. https://doi.org/10.1121/10.0025463. Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48, 2. https://doi.org/10.18637/jss.v067.i01. Beranek, L. (2012). Concert halls and opera houses: music, acoustics, and architecture. Springer, Heidelberg, Germany. Berger, K. W. (1964). Some factors in the recognition of timbre. The Journal of the Acoustical Society of America, 36(10), 1888–1891. https://doi.org/10.1121/1.1919287. Bey, C & McAdams, S. (2002). Schema-based processing in auditory scene analysis. Perception & Psychophysics, 64(5), 844–854. https://doi.org/10.3758/BF03194750. Bregman, A. S. (1990). Auditory Scene Analysis: The Perceptual Organization of Sound. MIT Press, Cambridge, MA. Bürgel, M., Mares, D., & Siedenburg, K. (2024). Enhanced salience of edge frequencies in auditory pattern recognition. Attention, Perception, & Psychophysics, 86, 2811–2820. https://doi.org/10.3758/s13414-024-02971-x. Listening in the mix: Lead vocals robustly attract auditory attention in popular music. Frontiers in Psychology, 12, 769663. https://doi.org/10.3389/fpsyg.2021.769663. Culling, J. F., Hodder, K. I., & Toh, C. Y. (2003). Effects of reverberation on perceptual segregation of competing voices. The Journal of the Acoustical Society of America, 114(5), 2871–2876. https://doi.org/10.1121/1.1616922.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 36 Eagleson, H. V. & Eagleson, O. W. (1947). Identification of musical instruments when heard directly and over a public address system. The Journal of the Acoustical Society of America, 19(2), 338–342. https://doi.org/10.1121/1.1902523. Elliott, C. A. (1975). Attacks and releases as factors in instrument identification. Journal of Research in Music Education, 23(1), 35–40. https://doi.org/10.2307/3345201. Fischer, M., Soden, K., Thoret, E., Montrey, M., & McAdams, S. (2021). Instrument timbre enhances perceptual segregation in orchestral music. Music Perception, 38(5), 473–498. https://doi.org/10.1525/mp.2021.38.5.473. Gerken, M., Hohmann, V., & Grimm, G. (2024). Comparison of 2d and 3d multichannel audio rendering methods for hearing research applications using technical and perceptual measures. Acta Acustica, 8, 17. https://doi.org/10.1051/aacus/2024009. Goad, P. J. & Keefe, D. H. (1992). Timbre discrimination of musical instruments in a concert hall. Music Perception, 10(1), 43–62. https://doi.org/10.2307/40285537. Grimm, G. & Herzke, T. (2024). TASCAR version 0.230,0. Available at https://github.com/gisogrimm/tascar. Grimm, G., Luberadzka, J., & Hohmann, V. (2019). A toolbox for rendering virtual acoustic environments in the context of audiology. Acta Acustica united with Acustica, 105(3), 566– 578. https://doi.org/10.3813/aaa.919337. Hake, R., Bürgel, M., Nguyen, N. K., Greasley, A., Müllensiefen, D., & Siedenburg, K. (2023). Development of an adaptive test of musical scene analysis abilities for normalhearing and hearing-impaired listeners. Behavior Research Methods, 56, 5456–5481. https://doi.org/10.3758/s13428-023-02279-y.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 37 Hendrikse, M. M., Llorach, G., Hohmann, V., & Grimm, G. (2019). Movement and gaze behavior in virtual audiovisual listening environments resembling everyday life. Trends in Hearing, 23, 2331216519872362. https://doi.org/10.1177/2331216519872362. Huron, D. (2001). Tone and voice: A derivation of the rules of voice-leading from perceptual principles. Music Perception, 19(1), 1–64. https://doi.org/10.1525/mp.2001.19.1. Jacobsen, S. & Siedenburg, K. (2024). Exploring the relation between fundamental frequency and spectral envelope in the perception of musical instrument sounds. Acta Acustica, 8, 48. https://doi.org/10.1051/aacus/2024038. Kato, K., Nagao, T., Yamanaka, T., Kawai, K., & Sakakibara, K.-I. (2014). Effect of room acoustics on timbral brightness of clarinet tones: Experimental investigation with two binaural room impulse responses. Acoustical Science and Technology, 35(6), 300–308. https://doi.org/10.1250/ast.35.300. Kendall, R. A., & Carterette, E. C. (1993). Identification and blend of timbres as a basis for orchestration. Contemporary Music Review, 9(1-2), 51–67. https://doi.org/10.1080/07494469300640341. Lembke, S.-A., Parker, K., Narmour, E., & McAdams, S. (2017). Acoustical correlates of perceptual blend in timbre dyads and triads. Musicae Scientiae, 23(2), 250–274. https://doi.org/10.1177/1029864917731806. Litovsky, R. Y., Goupell, M. J., Fay, R. R., & Popper, A. N. (2021). Binaural hearing. Springer. https://doi.org/10.1007/978-3-030-57100-9. Lokki, T. & Pätynen, J. (2020). Auditory spatial impression in concert halls. In J. Blauert & J. Braasch (Eds.), The Technology of Binaural Understanding, Modern Acoustics and Signal
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 38 Processing, (pp. 173–202). Springer Nature Switzerland AG. https://doi.org/10.1007/978-3030-00386-9_7. McAdams, S. (2019). Timbre as a structuring force in music. In K. Siedenburg, C. Saitis, S. McAdams, A. N. Popper, & R. R. Fay (Eds.), Timbre: Acoustics, Perception, and Cognition, (Vol. 69, pp. 23–57). Springer. https://doi.org/10.1007/978-3-030-14832-4_8. McAdams, S., Gianferrara, P. G., Korsmit, I. R., Goodchild, M., & Soden, K. (2025). Factors contributing to instrumental blends in orchestral excerpts. Music & Science, 8, 20592043251326391. https://doi.org/10.1177/20592043251326391. McAdams, S., Thoret, E., Wang, G., & Montrey, M. (2023). Timbral cues for learning to generalize musical instrument identity across pitch register. The Journal of the Acoustical Society of America, 153(2), 797–811. https://doi.org/10.1121/10.0017100. McFee, B., Matt McVicar, Daniel Faronbi, Iran Roman, Matan Gover, Stefan Balke, Scott Seyfarth, Ayoub Malek, Colin Raffel, Vincent Lostanlen, Benjamin van Niekirk, Dana Lee, Frank Cwitkowitz, Frank Zalkow, Oriol Nieto, Dan Ellis, Jack Mason, Kyungyun Lee, Bea Steers, … Waldir Pimenta. (2024). librosa/librosa: 0.10.2.post1 (0.10.2.post1). Zenodo. https://doi.org/10.5281/zenodo.11192913. Meyer, J. (2008). Musikalische Akustik. In S. Weinzierl (Ed.), Handbuch der Audiotechnik, (pp. 123–180). Springer, Heidelberg, Germany. https://doi.org/10.1007/978-3-540-343011_4. Müllensiefen, D., Gingras, B., Musil, J., & Stewart, L. (2014). The musicality of nonmusicians: an index for assessing musical sophistication in the general population. PLoS ONE, 9(2), e89642. https://doi.org/10.1371/journal.pone.0089642.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 39 Patil, K., Pressnitzer, D., Shamma, S. A., & Elhilali, M. (2012). Music in our ears: The biological bases of musical timbre perception. PLOS Computational Biology, 8(11), e1002759. https://doi.org/10.1371/journal.pcbi.1002759. Reuter, C. (1996). Die auditive Diskrimination von Orchesterinstrumenten. Peter Lang, Frankfurt am Main, Germany. Saldanha, E & Corso, J. F. (1964). Timbre cues and the identification of musical instruments. The Journal of the Acoustical Society of America, 36(11), 2021–2026. https://doi.org/10.1121/1.1919317. Sandell, G. J. (1995). Roles for spectral centroid and other factors in determining “blended” instrument pairings in orchestration. Music Perception, 13(2), 209–246. https://doi.org/10.2307/40285694. Siedenburg, K. (2019). Specifying the perceptual relevance of onset transients for musical instrument identification. The Journal of the Acoustical Society of America, 145(2), 1078– 1087. https://doi.org/10.1121/1.5091778. Siedenburg, K., Schädler, M. R. & Hülsmeier, D. (2019). Modeling the onset advantage in musical instrument recognition. The Journal of the Acoustical Society of America, 146(6), EL523–EL529. https://doi.org/10.1121/1.5141369. Siedenburg, K., Goldmann, K. & van de Par, S. (2021). Tracking musical voices in Bach’s The Art of the Fugue: Timbral heterogeneity differentially affects younger normal-hearing listeners and older hearing-aid users. Frontiers in Psychology, 12(608684), 1–9. https://doi.org/10.3389/fpsyg.2021.608684.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 40 Siedenburg, K., Röttges, S., Wagener, K. & Hohmann, V. (2020). Can you hear out the melody? testing musical scene perception in young normal-hearing and older hearingimpaired listeners. Trends in Hearing, 24, 1–15. https://doi.org/10.1177/2331216520945826. Srinivasan, A., Sullivan, D. & Fujinaga, I. (2002). Recognition of isolated instrument tones by conservatory students. In C. Stevens, D. Burnham, G. McPherson, E. Schubert, & E. Renwick (Eds.), Proceedings of the 7th International Conference on Music Perception and Cognition, (pp. 17–21), Sydney, Australia. Strong, W. & Clark, M. (1967). Synthesis of wind-instrument tones. The Journal of the Acoustical Society of America, 41(1), 39–52. https://doi.org/10.1121/1.1910327. Tardieu, D. & McAdams, S. (2012). Perception of dyads of impulsive and sustained instrument sounds. Music Perception, 30(2), 117–128. https://doi.org/10.1525/mp.2012.30.2.117. Thilakan, J., B T, B., Colella Gomes, O., Chen, J.-M. & Kob, M. (2025). Exploring the role of room acoustic environments in the perception of musical blending. The Journal of the Acoustical Society of America, 157(2), 738–754. https://doi.org/10.1121/10.0035563. Thoret, E., Depalle, P. & McAdams, S. (2017). Perceptually salient regions of the modulation power spectrum for musical instrument identification. Frontiers in Psychology, 8(587). https://doi.org/10.3389/fpsyg.2017.00587. van Noorden, L. P. A. S. (1975). Temporal coherence in the perception of tone sequences. Unpublished doctoral dissertation, Eindhoven University of Technology, Eindhoven, Netherlands.
INSTRUMENT IDENTIFICATION IN A VIRTUAL ORCHESTRAL SCENE 41 West, B. T., Welch, K. B. & Galecki, A. T. (2022). Linear Mixed Models: A Practical Guide Using Statistical Software. Chapman and Hall/CRC, Boca Raton, FL. https://doi.org/10.1201/9781003181064. Woods, K. J. & McDermott, J. H. (2015). Attentive tracking of sound sources. Current Biology, 25(17), 2238–2246. https://doi.org/10.1016/j.cub.2015.07.043.