Perceiving emotions in speech and music: modulations related to ageing, musical training, and Parkinson's disease
Full text
César F. Lima PERCEIVING EMOTIONS IN SPEECH AND MUSIC: MODULATIONS RELATED TO AGEING, MUSICAL TRAINING, AND PARKINSON’S DISEASE TESE DE DOUTORAMENTO PSICOLOGIA 2011
!
! ! PERCEIVING EMOTIONS IN SPEECH AND MUSIC Modulations Related to Ageing, Musical Training, and Parkinson’s Disease CÉSAR F. LIMA
! !
! ! PERCEIVING EMOTIONS IN SPEECH AND MUSIC Modulations Related to Ageing, Musical Training, and Parkinson’s Disease CÉSAR F. LIMA Thesis supervised by Professor São Luís Castro and presented at University of Porto for the Ph.D. degree in Psychology September 2011
! ! The studies presented in this thesis were supported by a doctoral grant from the Portuguese Foundation for Science and Technology (SFRH/BD/39306/2007) and by the Language Research Group of the Centre for Psychology at University of Porto (R&D Unit funded by FCT, 160/50; PTDC/PSI/66641/2006). They are also part of the project Emotional processing from language and music: Comparative neurocognitive and functional neuroimaging studies, funded by the Bial Foundation (N.º 29/08).
! 1 KEYWORDS Emotion recognition; speech prosody; music; ageing; musical training; Parkinson’s disease. American Psychological Association (PsycINFO! Content Classification Code System): 2300 Human Experimental Psychology 2326 Auditory & Speech Perception 2343 Learning & Memory 2500 Physiological Psychology & Neuroscience 2520 Neuropsychology & Neurology
! 2
! 3 RESUMO A fala e a música são dois poderosos meios de comunicação de emoções. A presente tese investiga como percebemos emoções em prosódia da fala – o tom de voz – e em música instrumental. Apresentamos uma série de estudos que visam três objectivos principais: determinar como o reconhecimento de emoções em música é modulado pela idade e pela formação musical (estudos 1 e 2); desenvolver uma base de estímulos de fala em Português para investigação em emoções prosódicas (estudo 3); e examinar a hipótese de que o reconhecimento de emoções em fala e em música depende de mecanismos neurocognitivos comuns (estudos 4 e 5). Utilizámos tarefas de escolha forçada e de julgamento de magnitude. Nos estudos 1 e 2, testámos adultos que variavam quanto à idade e formação musical no reconhecimento de alegria, serenidade, tristeza e medo em excertos musicais. Observámos que o aumento da idade está associado a uma diminuição na sensibilidade às emoções negativas, tristeza e medo, enquanto a sensibilidade às emoções positivas permanece estável. As mudanças na tristeza e medo foram significativas da meia-idade em diante. O número de anos de formação musical esteve correlacionado com uma maior sensibilidade às emoções musicais. Os efeitos de idade e formação foram independentes de diferenças em capacidades cognitivas gerais, o que sugere que aqueles efeitos têm uma origem primária. No estudo 3, duas falantes gravaram uma lista de frases e de pseudo-frases em Português variando a prosódia para comunicar neutralidade e seis emoções: alegria, fúria, medo, repulsa, surpresa e tristeza. Os perfis acústicos das diferentes emoções foram consistentes com descrições em outras línguas. As emoções foram reconhecidas com elevados níveis de exatidão, quer nas frases (75% correto), quer nas pseudo-frases (71%). Os estímulos foram incluídos numa base de dados (190 frases e 178 pseudofrases) que disponibilizámos para as comunidades de investigação e clínica. Nos estudos 4 e 5 determinámos em que medida o reconhecimento de emoções em fala e em música recruta mecanismos partilhados. Os resultados sugerem o envolvimento de uma combinação de mecanismos partilhados e específicos ao domínio. No estudo 4, investigámos se a formação num dos domínios, música, está associada a benefícios no outro domínio, fala. Comparámos músicos com participantes sem formação musical no reconhecimento de neutralidade e de seis emoções em fala. Encontrámos um efeito robusto de transferência entre domínios: os músicos foram mais exatos do que os participantes sem formação em todas as emoções. Este resultado indica que os mecanismos são pelo menos parcialmente partilhados. No estudo 5, implementámos um design comparativo entre fala e música na doença de Parkinson (DP). Será que a DP produz um perfil similar de défices em ambos os domínios? Encontrámos uma dissociação: os doentes tiveram défices no reconhecimento de emoções positivas em música, alegria e serenidade, mas foram normais na fala; tiveram défices no reconhecimento de tristeza em fala, mas foram normais na música. Os défices na música não podem ser explicados por dificuldades percetivas e cognitivas; na fala, o reconhecimento de emoções esteve fortemente associado à disfunção executiva dos doentes. Esta dissociação neuropsicológica é evidência de que fala e música podem recrutar mecanismos específicos ao domínio. No seu conjunto, os cinco estudos aqui apresentados contribuem para o avanço da nossa compreensão sobre diferenças
! 10 and structural analyses of speech and musical stimuli. Professors Anders Flykt and Arne Öhman kindly made available and gave permission for the utilization of the Karolinska Directed Emotional Faces database. Dr. Bruno Verschuere provided normative data on this database. Universidade Sénior das Antas and Universidade Sénior da Foz let us recruit their students as research participants and generously allowed us to collect data in their facilities. The studies presented here benefited very much from the critical comments and expert knowledge of many reviewers and journal editors. I am indebted to all research participants: students of different ages, musicians, and patients with Parkinson’s disease and their families. Without them this thesis really would not have been possible. They made the process of data collection a greatly personally rewarding experience. I am also in debt to the many people who helped in the difficult and timeconsuming process of finding and recruiting participants. I thank Angélica Relvas, Gaëtan Cousin, Manuela Cameirão and Paulo Santiago for their suggestions and comments on an earlier version of the text. My deepest acknowledgment goes to my family and to my friends – the old friends and the ones I was lucky enough to meet during this journey. They give personal meaning to my work. I specially thank my parents and my sister. Words cannot describe my gratitude for their unconditional love and support. Porto, summer 2011
! 11 ABBREVIATIONS ANCOVA Analysis of covariance ANOVA Analysis of variance AQ The Autism-Spectrum Quotient questionnaire dB Decibel e.g. Exempli gratia, for example ERP Event-related potential et al. Et alii, and others F0 Fundamental frequency fMRI Functional Magnetic Resonance Imaging Hu Unbiased hit rate i.e. Id est, that is MBEA Montreal Battery of Evaluation of Amusia MMSE Mini-Mental State Examination MOCA Montreal Cognitive Assessment ms Milliseconds ns Non-significant PD Parkinson’s disease PET Positron Emission Tomography RT Reaction time s Seconds SD Standard deviation TIPI Ten-Item Personality Inventory UPDRS Unified Parkinson’s Disease Rating Scale WAIS-III Wechsler Adult Intelligence Scale, version III
! 12
! 13 LIST OF PAPERS The work presented in this thesis is an expanded and updated version of five scientific articles which were submitted for publication in international peer-reviewed journals (same order as in the thesis): 1. Lima, C. F., & Castro, S. L. (2011). Emotion recognition in music changes across the adult life span. Cognition and Emotion, 25, 585-598. doi:10.1080/02699931.2010.502449 [Chapter VI, Study 1] 2. Castro, S. L., & Lima, C. F. (submitted). Emotion recognition in music is robust, and yet variable: The impact of age and musical expertise. [Chapter VI, Study 2] 3. Castro, S. L., & Lima, C. F. (2010). Recognizing emotions in spoken language: A validated set of Portuguese sentences and pseudo-sentences for research on emotional prosody. Behavior Research Methods, 42, 74-81. doi:10.3758/BRM.42.1.74 [Chapter VII] 4. Lima, C. F., & Castro, S. L. (2011). Speaking to the trained ear: Musical expertise enhances the recognition of emotions in speech prosody. Emotion, 11, 1021-1031. doi:10.1037/a002452 [Chapter VIII] 5. Lima, C. F., Garrett, C., & Castro, S. L. (submitted). Perceiving emotions in music and in speech prosody is dissociated in Parkinson’s disease. [Chapter IX]
! 14
Contents 15 CONTENTS Keywords 1 Resumo 3 Résumé 5 Abstract 7 Acknowledgments 9 Abbreviations 11 List of papers 13 Contents 15 List of Tables 19 List of Figures 22 List of Appendices 23 General introduction 27 PART 1 Sounds with emotional meaning: Emotion theories, emotional speech prosody, and musical emotions 31 Chapter I Emotion theories 33 1.1. Discrete emotion theories 36 1.2. Dimensional emotion theories 37 1.3. Emotion expression and recognition 39 Chapter II Speech prosody and emotion 41 2.1. Voice as a means of emotional expression 44
Contents ! 16 2.2. Acoustic cues of emotion in voice 45 2.3. Emotion communication through speech prosody 46 2.4. Neurocognitive basis of emotional prosody 52 2.5. Individual differences 57 2.6. When emotional prosody fails 58 Chapter III Music and emotion 61 3.1. Communication of emotions in music 64 3.2. Musical structure and expressive cues in performance 67 3.3. Emotion perception and experience 69 3.4. Which emotions are evoked by music? 71 3.5. Mechanisms underlying emotional responses to music 72 3.6. Neurocognitive basis of musical emotions 74 3.7. Individual differences 77 3.8. Clinical applications 79 Chapter IV Perceiving emotions in speech prosody and in music: How close? 81 4.1. Similar codes 84 4.2. Shared neurocognitive mechanisms? 87 Chapter V Goals of this thesis 91 PART 2 Empirical studies 99 Chapter VI Ageing and musical expertise modulate emotion recognition in music 101 6.1. Introduction 103 6.2. Study 1 110 6.2.1. Method 110 6.2.2. Results 113
Contents ! ! 17 6.2.3. Discussion 121 6.2. Study 2 123 6.2.1. Method 123 6.2.2. Results 126 6.2.3. Discussion 133 6.4. General discussion 135 6.5. Conclusion 141 Chapter VII Development and validation of a set of Portuguese stimuli for research on emotional prosody 143 7.1. Introduction 145 7.2. Method 149 7.2.1. Recording 149 7.2.2. Validation 151 7.3. Results and discussion 153 7.4. Conclusion 159 Chapter VIII Speaking to the trained ear: Musical expertise is associated with enhanced recognition of emotions in speech prosody 163 8.1. Introduction 165 8.2. Method 169 8.3. Results 171 8.4. Discussion 179 8.5. Conclusion 185 Chapter IX A dissociation between emotion recognition in music and in speech prosody in Parkinson’s disease 187 9.1. Introduction 189 9.2. Method 193 9.3. Results 200 9.4. Discussion 211
Contents ! 18 9.5. Conclusion 217 Chapter X General conclusions 219 10.1. Summarizing 222 10.2. Future directions 227 10.3. Concluding 229 References 231 Appendices 277 ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! !
Contents ! ! 19 LIST OF TABLES TABLE 1 Summary of empirical data on patterns of acoustic cues for anger, fear, sadness and joy in vocal expression. Adapted from Scherer (2003, p. 233). 49 TABLE 2 Patterns of acoustic cues for discrete emotions in vocal expression and in music performance. Reprinted from Juslin and Laukka (2003, p. 802). 86 TABLE 3 Demographic and background characteristics of the participants in each age group (Study 1). SDs are presented in parentheses. 111 TABLE 4 Distribution of responses (%) for each intended emotion as a function of age (Study 1). Diagonal cells in bold indicate accurate categorizations, i.e., match between the highest rating and the intended emotion. Ambivalent responses represent situations in which the highest rating was assigned to more than one emotion. Standard errors are presented in parentheses. 114 TABLE 5 Results from multiple regression analyses on listeners’ utilization of music compositional cues for subjective ratings, as a function of emotion and age (Study 1). Values correspond to beta weights and adjusted R2. 120 TABLE 6 Demographic and background characteristics of controls and musicians for each age group (Study 2). SDs are presented in parentheses. 124
! 26 ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! !
General Introduction ! 27 GENERAL INTRODUCTION Humans are hard-wired to connect with the sounds of voice and music. Words, tones of voice, laughter, sobs, screams, sighs, timbres, musical notes, melodies, rhythms, beats, all that, are building blocks of our auditory environment and mental life. Our senses are perpetually inundated by these sounds. Since ancient times in our species evolution, throughout our life span, across all societies, we are adept at making sense out of voice and music. We use them as ways to establish meaningful relations with the others and with ourselves. They are multifaceted and powerful communication tools, able to convey a multitude of kinds of information. Voice and music have also a startling capacity to move us. If you think of your own experience, it will be probably easy to remember of a situation in which you felt strongly aroused by a music that you particularly enjoyed, or deeply moved while listening to someone’s speech. This thesis is about a key common feature of voice and music – the communication of emotions. The focus is on how we perceive emotions communicated by speech prosody, the tone of voice, and by instrumental music. It comes naturally to us to tell whether someone is angry or scared by the sound of her or his voice, or whether a piece of music is happy or sad. The apparent ease of these operations, though, belies the complexity of the neurocognitive mechanisms that are involved in transforming streams of sound – pressure waves producing vibrations of the basilar membrane in the ear – into meaningful subjective representations of emotion. Despite recent advances in cognitive and neuroscientific research on this topic, many questions remain open. Herein, we set out investigate some of them in a series of studies using behavioural and neuropsychological approaches.
General Introduction ! 28 We pursue three main goals. The first one is to examine how the recognition of emotions in music changes with advancing age and as a function of musical training. The idea that ageing and musical training modulate many cognitive, neural and socioemotional functions is firmly entrenched (e.g., Habib & Besson, 2009; Hedden & Gabrieli, 2004; Salthouse, 2009; Samanez-Larkin & Carstensen, 2011), but their impact on musical emotions is still poorly understood. The second goal is to devise and validate a database of speech stimuli in Portuguese for research on vocal emotions. Analogous materials have been developed for different languages (Burkhardt, Paeschke, Rolfes, Sendlmeier, & Weiss, 2005; Makarova & Petrushin, 2002; Pell, 2002; Ross, Thompson, & Yenkosky, 1997; Staroniewicz & Majewski, 2009; Wu, Yang, Wu, & Li, 2006). They may be useful for studies with native speakers of the stimuli’s language, for cross-language research, as well as for the assessment of pragmatic skills in clinical settings. The third goal is to examine the extent to which emotion processing in speech prosody and in music engages shared neurocognitive mechanisms. This hypothesis has been highly debated in the last years (e.g., Juslin & Laukka, 2003; Juslin, Liljeström, Västfjäll, & Lundqvist, 2010; Juslin & Västfjäll, 2008; Nieminen, Istók, Brattico, Tervaniemi, & Huotilainen, 2011; Patel, 2008b; Peretz, 2010), but there is a dearth of empirical research on it. We investigate this issue in two ways. One consists of examining possible transfer effects from one domain to the other. Specifically, we determine whether expertise in music is associated with enhanced ability to recognize emotions in speech prosody. The other approach consists of using a comparative design between the two domains in a neuropsychological study. We determine whether Parkinson’s disease (PD), a neurodegenerative disorder involving damage in the basal ganglia, leads to a similar profile of impairments in emotion recognition in speech and music. The thesis is divided in two parts. Part 1 comprises a review of the literature. This review covers emotion theories (Chapter I), the acoustic and neurocognitive foundations of vocal and musical emotions (Chapters II and III), and a discussion on the parallels between both domains concerning emotion expression and processing (Chapter IV). In the end of this part the goals of the empirical studies are described (Chapter V). Part 2 comprises the empirical chapters. The first one (Chapter VI) presents two studies in
General Introduction ! 29 which listeners of different ages and varying in musical training are examined in the recognition of emotions in short instrumental music excerpts. The second empirical chapter (Chapter VII) describes the development and recording of a set of Portuguese speech stimuli expressing different emotions, as well as the acoustic and perceptual validation of these stimuli. The third empirical chapter (Chapter VIII) presents a study in which musicians are compared with musically untrained listeners in the recognition of emotions in speech prosody. The last empirical chapter (Chapter IX) corresponds to the neuropsychological study. PD patients are compared with healthy matched controls in the recognition of emotions in speech prosody and in music. Each empirical chapter follows the typical structure of a journal article, including the sections introduction, method, results, and discussion. A brief conclusion is also provided. In the end of Part II (Chapter X), which coincides with the end of the thesis, we summarize the findings and conclusions of the empirical studies and discuss avenues for future research.
! 30
! 31 PART 1 SOUNDS WITH EMOTIONAL MEANING: EMOTION THEORIES, EMOTIONAL SPEECH PROSODY, AND MUSICAL EMOTIONS
! 32
! 33 CHAPTER I Emotion Theories
! 34
Chapter I. Emotion Theories ! 35 It is common sense that emotions play a pivotal role in human existence. They accompany almost every significant event in our lives, and shape strongly the way we think, feel and act. What is an emotion? This question was the title of one of the most famous William James’ (1884) papers. Yet, more than one century afterwards a definitive answer is still lacking, even though the study of emotions is currently highly popular in cognitive neuroscience. They are notoriously hard to define and there is no general consensus regarding even what should count or not as an instance of emotion. However, experts do agree to a significant extent with respect to the general characteristics of emotions, as well as about the functions they serve (Izard, 2009). Emotions correspond to relatively brief and intense reactions to potentially important changes in our external or internal environment (e.g., stimuli or events representing subjective challenges or opportunities). They typically involve a number of subcomponents – subjective feeling, cognitive changes, physiological arousal, expressive behaviour, action tendencies, and regulation – that are more or less synchronized. They are coordinated reactions intended to promote adaptive behaviours in response to a changing environment (e.g., Davidson, Scherer, & Goldsmith, 2002; Ekman, 1999; Juslin & Sloboda, 2010; Levenson, 1999; Scherer, 2005; Smith & Lazarus, 1990). Although the mechanisms underlying emotional reactions may be of several kinds, in everyday life emotions are often the result of evaluating (cognitive appraisal) an event as relevant to subjective intentions, goals, motives and concerns (Scherer, 2005; Sloboda & Juslin, 2010). The two theoretical frameworks that have most strongly influenced research on emotions in the last decades are discrete categories and dimensional emotion theories, which we outline below.
! 42
Chapter II. Speech Prosody and Emotion ! 43 The human voice can be considered the most important sound of our auditory environment. First and foremost, it is the carrier of speech – we can speak and understand linguistic messages, and this ability defines us as humans. In parallel to speech, though, the voice expresses a wealth of socially relevant information. It communicates information about the speaker’s gender, identity, age, size, health, attractiveness, trustworthiness, dominance, feelings and mood, even when speech content is not available (for instance because it is a unknown language, or because the vocalization is nonverbal, such as a laughter or a sigh). Thus, the voice is very much like an “auditory face” (Belin, Bestelmeyer, Latinus, & Watson, in press). We are endowed with abilities to extract paralinguistic information from voices, in what reflects a more primitive and universal form of communication than language itself – vocalizations have been used by many species for millions of years before language emerged (Belin, Fecteau, & Bédard, 2004; Bruckert et al., 2010; Latinus & Belin, 2011). The expression of emotions through vocal cues has particular relevance for social interactions. The acoustic, psychological and neural mechanisms involved in the processing of vocal emotions are outlined in the following pages. The review focuses chiefly on the type of emotional vocal cue that we examine in the empirical part of this thesis, speech prosody. Briefly put, speech prosody refers to the musical aspects of speech, the tone of voice, which often convey information about the speakers’ emotions.
Chapter II. Speech Prosody and Emotion ! 44 2.1. VOICE AS A MEANS OF EMOTIONAL EXPRESSION The role of voice for emotional expression has long been recognized, as indicated by the treatises of the topic found in Greek and Roman manuals on rhetoric (Scherer, 2003), and by Darwin’s work on evolutionary biology in the 19th century (Darwin, 1872/2009). An interesting issue is that vocal expression may be the most phylogenetically continuous of all forms of emotional communication. A large number of non-human species use vocalizations to communicate motivational and emotional states. Vocalizations are especially important in social mammals, whose life is based on complex and cooperative interactions between individuals. There are similarities across many species of mammals, including humans, in the neural control of voice production, in voice production mechanisms, as well as in the acoustic characteristics of the expressions that signal different states (Fichtel, Hammerschmidt, & Jürgens, 2001; Grandjean, Bänziger, & Scherer, 2006; Juslin & Laukka, 2003; Scherer, 1995; Scherer, Johnstone, & Klasmeyer, 2003). Because the expression of vocal emotions serves a highly adaptive function for survival in social environments, it is likely that is was shaped by natural selection – vocalizations usually correspond to the communication of biologically important events (e.g., Hauser & McDermott, 2003). Examples of nonhuman vocal expressions are calls of warning, threat, submissive or affiliative states, mating, attention-getting, desire for social contact and companionship (Juslin & Laukka, 2003). The richness of the expressive repertoire varies across species depending on the complexity and differentiation of the sound-producing apparatus, with some species producing only a few innate vocal behaviours (e.g., frogs), and others exhibiting a varied repertoire of voluntarily controlled productions (e.g., non-human primates). In humans, the formalized and abstract systems of language acquired special prominence as a form of communication, namely for the expression of emotional meanings (through semantics). Notwithstanding, more ancient paralinguistic emotional expressions are also present in our vocal behaviours under different forms. We often produce purely emotive nonverbal vocalizations, such as laughter, sighs, sobs or screams. These are perceived rapidly (Sauter & Eimer, 2009), accurately (Sauter, Eisner, Calder, et al., 2010) and
Chapter II. Speech Prosody and Emotion ! 45 across cultures (Sauter, Eisner, Ekman, et al., 2010). In addition to nonverbal expressions, we routinely produce short interjections with affective connotation, such as “wow”, “Ah”, “yuck” or “oh” (Belin, Fillion-Bilodeau, & Gosselin, 2008; Scherer, 1995). Emotions are also expressed as an integrant part of the speech signal proper via speech prosody cues – suprasegmental vocal modulations that occur in the course of a spoken utterance. It is well established that emotional states are associated with changes in the way we speak: emotions involve psychophysiological changes, for instance in heart rate, blood flow and muscle tension, which affect respiration, vocal fold vibration and articulation, in such a way that all these influence the acoustic characteristics of certain voice features while we speak. For example, higher emotional arousal increases laryngeal tension and subglottal pressure, thereby increasing the intensity and changing the timbre of the voice. Thus, the suprasegmental signals that ride atop of speech provide diagnostic information concerning our states (Banse & Scherer, 1996; Scherer, 1986, 2003; Schirmer & Kotz, 2006). 2.2. ACOUSTIC CUES OF EMOTION IN VOICE Speech prosody comprises different acoustic cues: fundamental frequency (F0), intensity, tempo, rhythm and voice quality (e.g., Grandjean et al., 2006; Schirmer & Kotz, 2006). F0 reflects the frequency of vibration of the vocal folds during phonation and is subjectively perceived as voice pitch. Variations in the level, range and contour of F0 during the utterance are important cues of emotion. Intensity corresponds to the vocal energy and reflects the effort required to produce speech; it is subjectively perceived as loudness. Tempo corresponds to the number of phonemic segments per time unit, and rhythm to the structure of F0 accents, intensity peaks and distribution of pauses in the utterance. Finally, voice quality reflects the shape of the vocal tract, which can be modified, for instance, by the muscle tension in the larynx. Voice quality is
Chapter II. Speech Prosody and Emotion ! 46 acoustically characterized by the distribution of energy in the frequency spectrum, and is subjectively perceived as voice timbre (e.g., roughness and sharpness). The relative amount of energy in the high vs. low frequency region of the spectrum is highly informative to differentiate emotions (as the amount of high-frequency energy increases, the voice sounds more sharp and less soft). Modulations in these prosodic cues are associated with the communication of different emotional states, as we discuss below. An important notion is that changes in prosody can be spontaneous reflections of psychophysiological emotion-related states, but they can also be influenced by voluntary modulations. With regard to this issue, Scherer and colleagues have distinguished between push effects and pull effects in vocal expression (e.g., Banse & Scherer, 1996; Scherer, 1986; Scherer, 2003). Push effects refer to the direct impact of psychophysiological mechanisms involved in emotional episodes over the voice production system, which would be innately determined to a large extent. Pull effects, on the other hand, reflect the fact that prosody may incorporate sociocultural conventions and display rules, and can be strategically sculpted by our communicative intentions, independently of the presence of a real emotional state. During everyday life social interactions, the expression of emotions in prosody is often a joint product of push and pull effects. 2.3. EMOTION COMMUNICATION THROUGH SPEECH PROSODY As a communicative phenomenon, emotional prosody can be theoretically framed in the context of a modified version of the Brunswik’s functional lens model, which combines the analysis of expression and perception of affect signals (Brunswik, 1956; Grandjean et al., 2006; Juslin & Laukka, 2003; Scherer, 2003; Scherer et al., 2003). The process starts with the speaker expressing, or encoding, an emotional state through several acoustic cues, which can be objectively measured in the speech waveform. These
Chapter II. Speech Prosody and Emotion ! 47 objective parameters are called distal cues, as they are distant from the listener. Then they are transmitted, as part of the speech signal, and perceived by the auditory system of the listener, who forms a proximal percept. The perceived (subjective) cues are called proximal cues. The decoding process consists of combining and integrating the different internalized proximal cues in order to make a subjective inference regarding the speakers attitudes and emotions. Proximal cues are based on the distal ones, but they might suffer influences from the transmission channel (e.g., distance; noise) and from the characteristics of the perceptual system as well (e.g., selective enhancement of specific frequency bands). Within this framework, a functionally valid communication process occurs when the subjective judgments provided by the listeners correspond to the criteria for the state of the speaker (e.g., a specific emotional state, such as anger) with agreement rates above chance level. Over the last decades, a large number of empirical studies have been conducted to understand both expression and perception of emotional prosody. EXPRESSION | Studies that focus on expression have attempted to determine how differentiated patterns of acoustic voice cues are reliable indicators of specific emotional states. To that end, researchers collect vocal expressions corresponding to different emotional states and then measure them for acoustic attributes in order to examine associations between cue patterning and emotions. Classically measured acoustic cues are F0 (mean, variability and range), intensity and speech rate, although some studies include additional measures, such as jitter, high-frequency energy, formant frequency and F0 contours (e.g., Juslin & Laukka, 2001). An important methodological question concerns how to collect expressions for posterior analyses. Three major methods have been used, each of these with advantages and drawbacks (e.g., Scherer, 2003). One involves natural vocal expressions, that is, materials recorded during naturally occurring emotion episodes, for instance dangerous flight situations or journalists reporting emotionally charged events. This is a highly ecological method, but it has serious disadvantages: recordings are often very brief and obtained from a single speaker, they frequently suffer from bad recording quality, and experimental control is very limited. Another method consists of inducing emotional states experimentally, for example through films, and then recording speech samples while the
Chapter II. Speech Prosody and Emotion ! 48 speakers are under these states. This approach facilitates experimental control over the vocal materials produced, but it is difficult to induce strong and well-differentiated emotional states, such that the resulting recordings frequently express weak affect. Furthermore, one cannot assume that the same emotion elicitor procedure will produce the same emotional responses in all speakers. By far the most frequently employed method makes use of simulated actor portrayals, in which actors or untrained speakers are asked to pose expressions on the basis of emotion labels and/or scenarios. In a typical situation, they read the same verbal materials (e.g., sentences with neutral semantic content) while varying prosody to express different emotions. It has often been argued that this procedure may yield unnatural and stereotypical expressions, but acoustic convergence has been found between these expressions and the so-called natural ones (Juslin & Laukka, 2001). Acted portrayals contributed greatly for what is currently known about emotion communication through prosody (for a review of pros and cons of acted portrayals, Bänziger & Scherer, 2007). Acoustic analyses have revealed that there are different acoustic profiles for some emotion categories, thus confirming that voice inflections are informative about the speaker’s state (e.g., Banse & Scherer, 1996; Hammerschmidt & Jürgens, 2007; Juslin & Laukka, 2001; Paulmann, Pell, & Kotz, 2008b; Scherer, Banse, Wallbott, & Goldbeck, 1991). In a review of the available evidence, Scherer and colleagues (Banse & Scherer, 1996; Pittam & Scherer, 1993; Scherer, 2003) observed that anger is generally associated with increased mean F0 and voice intensity, although other features might also be found, such as increased variability and range of F0, increased highfrequency energy, faster rate of articulation and downward-directed F0 contours; fear is associated with increased mean and range of F0, increased high-frequency energy and faster rate of articulation; sadness is characterized by decreased mean and range of F0, decreased voice intensity and high-frequency energy, slower rate of articulation and downward-directed F0 contours; and joy/happiness appears to be associated with increased mean, range and variability of F0, increased voice intensity, and increased high-frequency energy and faster rate of articulation. These findings are summarized in Table 1.
Chapter II. Speech Prosody and Emotion ! 49 Table 1. Summary of empirical data on patterns of acoustic cues for anger, fear, sadness and joy in vocal expression. Adapted from Scherer (2003, p. 233). Acoustic cue Anger Fear Sadness Joy Intensity ! ! " ! F0 floor/mean ! ! " ! F0 variability ! " ! F0 range ! !(") " ! Sentence contours " " High-frequency energy ! ! " (!) Speech and articulation rate ! ! " (!) PERCEPTION | Research on the perception of emotional prosody, on the other hand, has examined the extent to which listeners are able to accurately infer the emotional state of a speaker from speech samples. In most studies, pre-recorded vocal stimuli – usually posed actor portrayals – expressing a number of different emotions are presented to a group of listeners. They frequently perform a forced-choice emotion recognition task, that is, they are asked to select the emotion that best describes the stimulus from a list of emotion labels. Then the percentage of stimuli correctly recognized per emotion is computed (percentage of matches between the selected emotion and the criterion for the stimulus’ emotion). It has been repeatedly shown that listeners are adept at perceiving emotions in prosody, with accuracy rates about four to five times higher than what would be expected if responses were given in a random manner (e.g., Juslin & Laukka, 2003; Pell, Paulmann, Dara, Alasseri, & Kotz, 2009; Scherer, 2003; Scherer et al., 2003). For example, Banse and Scherer (1996) found a global identification accuracy of 48% for 14 different emotion categories: hot anger, cold anger, panic fear, anxiety, desperation, sadness, elation, happiness, interest, boredom, shame, pride, disgust and contempt. Speech samples in this study consisted of pseudo-sentences composed of phonemes from several Indo-European languages, which were recorded by actors to express the intended emotions. Juslin and Laukka (2001) found an accuracy of 56% for stimuli consisting of sentences with neutral semantic content, which were recorded to express five emotions: anger, disgust, fear, happiness and sadness. Adolphs, Damasio and Tranel (2002) found a recognition accuracy of 81%, also for sentences with neutral semantic content expressing five emotions: anger, fear, happiness, sadness and surprise.
Chapter II. Speech Prosody and Emotion ! 50 Pell (2002) obtained an accuracy of 78% for pseudo-sentences that resembled English, which expressed six emotions: anger, disgust, happiness, pleasant surprise, sadness and neutrality. Altogether, these studies are strong evidence that we can perceive specific emotions from voice cues alone, i.e., even when concurrent semantic cues are not available or are emotionally neutral. Albeit recognition accuracy is generally high, significant differences are often detected between emotions. Sadness and anger are usually best recognized, followed by fear and happiness (Juslin & Laukka, 2003; Thompson & Balkwill, 2006). Disgust seems to be particularly difficult to recognize in prosody (Banse & Scherer, 1996; Scherer et al., 1991), in spite of being quite well detected in nonverbal vocal expressions (Belin et al., 2008; Sauter, Eisner, Calder, et al., 2010) and in faces (e.g., Goeleven, Raedt, Leyman, & Verschuere, 2008). The differential ease with which emotions and modalities are recognized might be related to differential evolutionary pressures (e.g., Scherer, 2003). An important issue is whether emotion recognition in prosody is governed by universal principles or determined by cultural and language-specific factors. Cross-cultural studies are pivotal to approach this question. Thompson and Balkwill (2006) examined how English-speaking listeners recognize joy, sadness, anger and fear as expressed in English, German, Chinese, Japanese and Tagalog speech samples (sentences with neutral semantic content). Recognition accuracy was highest for English stimuli, thus revealing an in-group advantage, but listeners were able to perceive emotions in the other languages well above the chance level, including in highly dissimilar non-Western languages. Consistently, Pell, Monetta, Paulmann, and Kotz (2009) found that Spanishspeaking listeners can recognize joy, sadness, anger, fear and disgust as expressed in Spanish, English, German and Arabic pseudo-sentences, even though they are better in their native language. Further evidence comes from a study by Scherer, Banse and Wallbott (2001), who compared how listeners from nine countries across Europe, America and Asia perceive anger, sadness, fear, joy, and neutrality as expressed by German actors (materials were pseudo-sentences). All listeners recognized emotions above the chance level, but accuracy rates varied across countries: generally, performance was higher when the native language of the listeners was closer to German (e.g., Dutch), and lower when it was highly dissimilar (e.g., Malay; see also Elfenbein
Chapter II. Speech Prosody and Emotion ! 51 & Ambady, 2002). Finding that prosodic emotions can be recognized cross-culturally suggests that the acoustic cues and the way in which listeners process prosody are universal to a significant extent. On the other hand, the in-group advantage indicates that cultural factors also exert a role. In other words, the recognition of emotional prosody appears to be determined by a combination of universal and culturally specific processes. Therefore, devising speech samples for different languages and analysing how speakers with different linguistic backgrounds process emotions is of paramount importance for better understanding emotion communication through prosody. FROM EXPRESSION TO PERCEPTION | Some studies have approached expression and perception of emotional prosody in a comprehensive manner in order to throw light on the nature of the inference rules used by the listeners during emotion recognition. Which acoustic cues (distal cues) are attended to in voice, and how do they lead to the recognition of a certain emotion? One strategy to answer this question consists of manipulating systematically specific acoustic cues (via synthesis or resynthesis) and examining how these manipulations affect emotion perception. For example, Breitenstein, Lancker and Daum (2001) manipulated prosodic stimuli for speech rate and F0 variability. They observed that speech rate was the most potent cue for subjective responses, with slow rate being reliably associated with sadness, and fast rate with anger, fear and neutrality; F0 variability was less potent but still determinant of responses, with reduced variability being associated with sad or neutrality, and large variability with fear, anger or happiness. Another strategy to explore the utilization of acoustic cues in emotion inferences involves measuring the acoustic properties of speech samples, and then correlating these with the subjective responses provided by the listeners. In regression analyses, Banse and Scherer (1996) showed that a significant proportion of variance in the listeners’ judgments was accounted for by acoustic attributes of the stimuli, such as mean and variability of F0, mean intensity, duration of voiced periods, relative proportion of high-frequency energy (cut-off 1000Hz) and energy drop-off in the spectrum. The proportion of explained variance was highest for hot anger, 36%, and for the other emotions it ranged from 7% to 25%. More recently, Juslin and Laukka (2001) found that subjective judgments for anger, disgust, fear, happiness and sadness could be significantly predicted by a set of nine acoustic cues in
Chapter II. Speech Prosody and Emotion ! 58 changes associated with adult development (e.g., Samanez-Larkin & Carstensen, 2011). Personality characteristics might modulate brain responses to prosody too. Individuals higher in social orientation, a trait reflecting interest in social exchange and care about others, have enhanced activity in the amygdala and in the orbitofrontal cortex in response to emotional vocal stimuli (Schirmer et al., 2008). Higher neuroticism is associated with enhanced activity in the right amygdala, left postcentral gyrus and medial frontal structures (Brück, Kreifelts, Kaza, Lotze, & Wildgruber, 2011). Trimmer and Cuddy (2008) uncovered that individuals with higher emotional intelligence are more sensitive to emotions in speech prosody. Training in music might be another factor explaining individual differences in prosody, but available evidence is inconclusive (Thompson, Schellenberg, & Husain, 2004; Trimmer & Cuddy, 2008). This hypothesis is empirically examined in Chapter VIII. 2.6. WHEN EMOTIONAL PROSODY FAILS Understanding the neurocognitive foundations of emotional prosody is important for the general advance of knowledge, and for applied clinical reasons as well. Difficulties in recognizing emotions in prosody are part of the phenotype of several developmental, neuropsychological and neuropsychiatric conditions. This is the case of autism spectrum disorders (Korpilahti et al., 2007; Lindner & Rosén, 2006), Williams syndrome (Pinheiro et al., 2011), paediatric and adult traumatic brain injury (Milders, Fuchs, & Crawford, 2003; Schmidt, Hanten, Li, Orsten, & Levin, 2010), acquired focal brain damage (Adolphs et al., 2002; Ross & Monnot, 2008), PD (Gray & Tickle-Degnen, 2010), Huntington’s disease (Speedie et al., 1990), Alzheimer’s disease (Taler, Baum, Chertkow, & Saumier, 2008), primary progressive aphasias (Rohrer, Sauter, Scott, Rossor, & Warren, in press), schizophrenia (Bach, Buxtorf, Grandjean, & Strik, 2009; Leitman et al., 2010), depression (Uekermann, Abdel-Hamid, Lehmkämper, Vollmoeller, & Daum, 2008), alcoholism (Monnot, Nixon, Lovallo, & Ross, 2001), posttraumatic stress disorder (Freeman, Hart, Kimbrell, & Ross, 2009) and psychopathy
Chapter II. Speech Prosody and Emotion ! 59 (Blair et al., 2002). Given that emotional prosody plays an important role for communication and social interactions, difficulties in this ability should be considered as a target in clinical settings, as they may compromise the patients’ psychosocial functioning (e.g., Trauner, Ballantyne, Friedland, & Chase, 1996). Higher accuracy in emotion recognition in prosody is associated with more positive interpersonal relationships and with less depression symptoms (Carton, Kessler, & Pape, 1999). It is also a predictor of social functioning, specifically of occupational performance (Hooker & Park, 2002). Even though our focus here is primarily on perception, it is important to note that deficits in the production of emotional prosody are also a feature of different clinical conditions, such as PD (e.g., Pell, Cheang, & Leonard, 2006) and autism spectrum disorders (Peppé, Cleland, Gibbon, O'Hare, & Castilla, 2011), with negative consequences for communication and social competence (e.g., Jaywant & Pell, 2010). After discussing how emotions are communicated in the music of speech, in the next chapter we turn to emotions in music itself.
! 60
! 61 CHAPTER III Music and Emotion
! 62
Chapter III. Music and Emotion ! 63 Like language, music is a uniquely human trait. Making and appreciating music are ubiquitous features of all human societies since ancient times. It is not possible to determine when exactly musical behaviours emerged during evolution – there are no fossil records of singing –, but archaeological findings date the first unequivocal evidence of musical instruments as back as circa 36,000 years, during the Upper Palaeolithic period (d'Errico et al., 2003). Musical behaviours and abilities also emerge very early in ontogenetic development, suggesting that they are guided by innate predispositions (Zentner & Eerola, 2010; Zentner & Kagan, 1996). As recently stated by Koelsch (2011), available evidence support the view that “musicality is a natural ability of the human brain” (p. 16). Indeed, there can be few individuals for whom music does not have a profound appeal, as indicated by the massive amounts of time that we spend on musical activities. A recent report found that music listening is placed among the most important and prevalent of leisure activities, above watching TV or reading books, for instance (Rentfrow & Gosling, 2003). Music is generally approached from a socio-cultural perspective, that is, as a cultural artefact and an exquisite form of art. The consideration of music as a biological faculty whose neurocognitive foundations can be examined is relatively recent. In the last years, though, we have witnessed a rapid expansion in biological, cognitive and neuroscientific research on music (e.g., Koelsch & Siebel, 2005; Patel, 2008a, 2010; Peretz, 2006). Remarkable advances have been made in many domains, such as in the understanding of the genetic basis of music cognition (e.g., Peretz, 2008), the development of musical abilities (e.g., Hannon & Trainor, 2007; Trehub & Hannon, 2006), the neurocognitive basis of music perception (e.g., Koelsch, 2011; Koelsch &
Chapter III. Music and Emotion ! 64 Siebel, 2005; Peretz & Coltheart, 2003), the disorders of music cognition (e.g., Liu, Patel, Fourcin, & Stewart, 2010; McDonald & Stewart, 2008; Peretz, Champod, & Hyde, 2003), the relations between music and language (e.g., Juslin & Laukka, 2003; Patel, 2008a; Schön & Besson, 2001; Schön et al., 2010), the effects of musical training on neurocognitive plasticity (e.g., Besson, Chobert, & Marie, 2011; Habib & Besson, 2009; Moreno et al., 2009; Pantev & Herholz, in press; Patel, 2011) and the processing of emotions in music (e.g., Juslin & Västfjäll, 2008; Koelsch, 2010). Emotions are perhaps the most important component of music. Ever since ancient Greece, there is an awareness of the ability of music to communicate emotions that we can perceive, enjoy and be moved by (e.g., Thompson, 2009). We use music to influence moods, to release emotions, to comfort and to enjoy ourselves, to match our emotional states and to alleviate stress (e.g., Juslin & Sloboda, 2010). The expressive power of music is also presumed in a number of applications in society, for example when we use it as a source of emotion in films (Cohen, 2010), marketing (North & Hargreaves, 2010) and therapy (Thaut, 2005; Thaut & Wheeler, 2010). How do the sounds of music express emotions? How does our neurocognitive system deal with emotions in music? In the following pages we review studies on emotion communication in music and on the neurocognitive foundations of musical emotions. 3.1. COMMUNICATION OF EMOTIONS IN MUSIC Theoretical treatises about musical emotions have a long history (e.g., Budd, 1985; Kivy, 2002; Meyer, 1956), but only recently have they become sidelined by empirical research in a systematic fashion. Musical emotions were considered too mysterious, elusive and personal to be examined rigorously in psychology and neuroscience. In the last two decades, however, this has changed and the topic of emotion processing in music is now expanding at a fast pace, paralleling the growing interest dedicated to emotion research in general (Sloboda & Juslin, 2010). A large number of studies focus
Chapter III. Music and Emotion ! 65 on how listeners perceive emotions communicated by music. In sharp contrast with the notion that this is a highly diffuse and variable experience, it has been repeatedly demonstrated that emotion recognition in music is remarkably consistent, accurate, immediate, precocious in ontogenetic development, and governed by universal principles to a significant extent. Bigand, Vieillard, Madurell, Marozeau, and Dacquet (2005) showed that emotional responses to music are highly consistent both within and between listeners. They presented 27 instrumental music excerpts from the Western classic repertoire to listeners with and without musical training, and asked them to group together the excerpts that conveyed similar emotions. The duration of the excerpts was manipulated, 30 s or 1 s, and listeners performed the task twice, with at least one week of interval (test-retest). Groupings were very stable across trained and untrained participants (between listeners), from the first to the second testing session (within listeners), and for both longer and shorter excerpts. There is also remarkable invariance across listeners in how they judge the emotion category expressed by a certain music excerpt. Agreement rates for emotion recognition in music tend to be high (e.g., Laukka & Juslin, 2007; Mohn, Argstatter, & Wilker, in press; Resnicow, Salovey, & Repp, 2004; Vieillard et al., 2008). For instance, Vieillard and colleagues (2008) obtained high accuracy rates for the recognition of four emotions expressed by short instrumental music excerpts: 99% for happiness, 84% for sadness, 82% for fear and 67% for peacefulness. Accordingly, in a meta-analytic review of 12 studies, Juslin and Laukka (2003) concluded that the recognition accuracy of five emotion categories in music performance (anger, fear, happiness, sadness and tenderness) is higher than what would be expected by chance alone. They found a global accuracy of 70%, which was as high as the one they found for vocal emotions. With respect to the temporal dynamics of emotion recognition, it has been shown that reliable judgments can be obtained after quite short segments of musical information, a fact that indicates that the underlying mechanisms are fast-acting and immediate. Peretz, Gagnon, and Bouchard (1998) observed that listeners need only 250 ms of music to distinguish happy from sad excerpts. In a related vein, Filipic, Tillmann, and Bigand (2010) demonstrated, using a gating task, that the distinction between moving and neutral music can be made on the
Chapter III. Music and Emotion ! 66 basis of 250 ms excerpts. Also using a gating task, Vieillard and colleagues (2008) determined the shortest duration that can trigger accurate emotion recognition in music: 483 ms for happiness, 1,446 ms for sadness, 1,737 ms for fear and 1,261 ms for peacefulness. From a developmental standpoint, available evidence indicates that the ability to perceive emotions in music emerges during early childhood. Cunningham and Sterling (1988) found that, by the age of four, children can already recognize happiness and fear in classical orchestral compositions with accuracy rates above chance, and by the age of six they are able to also identify reliably sadness and anger (see also Hunter, Schellenberg, & Stalinski, 2011). These early abilities are linked with sensitivity to specific musical features. Dalla Bella, Peretz, Rousseau, and Gosselin (2001) observed that six to eight-year-old children rely on both tempo (fast vs. low) and mode (major vs. minor) differences to distinguish happy from sad music, as adults do. Another attribute of emotion communication in music is that it is partly governed by universal principles. Some studies have shown that listeners are able to recognize emotions even in a nonfamiliar music style. Balkwill and Thompson (1999) found that Western listeners are sensitive to joy, sadness and anger in Hindustani raga excerpts. More recently, Fritz and colleagues (2009) compared how Western and Mafa listeners recognize happiness, sadness and fear in instrumental music excerpts composed according to the Western tonal tradition. The Mafa people are a native African ethnic group who had never been exposed to Western music. Western listeners had higher recognition accuracy that the Mafa, denoting a same-culture advantage, but the Mafa were able to recognize the three emotions with accuracy rates above chance level. Additionally, as Western listeners, they perceived consonant music as more pleasant than dissonant music. The fact that the recognition of musical emotions is accurate, fast, precocious and universal is evidence that the underlying mechanisms show some of the same properties of emotion stimuli important for social functioning, such as facial expressions and speech prosody (e.g., Doherty, Fitzsimons, Asenbauer, & Staunton, 1999; Elfenbein & Ambady, 2002; Grandjean et al., 2005; Pell, 2005; Pons, Harris, & Rosnay, 2004; Tracy & Robins, 2008).
Chapter III. Music and Emotion ! 67 3.2. MUSICAL STRUCTURE AND EXPRESSIVE CUES IN PERFORMANCE Which musical features contribute to the communication of a certain emotion? An important distinction is between features related to the music compositional structure, and those related to expressiveness in performance. Both are determinants of the perceived musical emotions (e.g., Thompson, 2009). Musical structure features correspond to those that are represented in conventional musical notation, such as tempo and dynamic markings, pitch, intervals, mode, melody, rhythm, harmony and formal properties (e.g., repetition; variation). These are manipulated by the composers to convey the intended emotional expressions (Gabrielsson & Lindström, 2010). Empirical research has shown that listeners are adept at perceiving emotions communicated by structural cues (Curtis & Bharucha, 2010; Gagnon & Peretz, 2003; Vieillard et al., 2008) since early childhood (Dalla Bella et al., 2001). Balkwill and Thompson (1999) observed that the structural properties tempo, rhythmic complexity, melodic complexity and pitch range are significant predictors of the listeners’ subjective emotion ratings. Several associations between specific structural cues and perceived emotion are well established. This is the case for tempo, which is one of the most important emotional cues. Fast tempo might lead to the perception of activity/excitement, happiness/joy/pleasantness, potency, surprise, anger, uneasiness and fear. Slow tempo, on the other hand, might lead to the perception of peacefulness, calmness/serenity, sadness, dignity/solemnity, tenderness, boredom and longing. Mode is also an important determinant of emotions, with major mode being usually associated with happiness/joy, serenity, gracefulness and solemnity, and minor mode with sadness, anger, tension, dignity and disgust. High pitch might be associated with expressions of happiness, serenity, excitement, potency, anger, fear and activity, and low pitch might be associated with sadness, dignity/solemnity and vigour. As for melody, a wide melodic range may suggest joy and fear, and a narrow melodic range may suggest sadness, dignity, delicacy, serenity and triumph. A simple and consonant harmony might convey happiness, dignity and serenity, whereas a complex and dissonant harmony might convey tension, anger, excitement, sadness and unpleasantness. A regular/smooth
Chapter III. Music and Emotion ! 74 might come to our mind while we listen to a certain melodic movement. The emotion experienced would result from the interaction between the image and the music. Episodic memory corresponds to a mechanism whereby we react emotionally to a piece of music because it evokes a memory of a specific event in our life. These emotions may be intense, possibly because the physiological responses and the experiential contents of the original events are stored along with each other in memory. Indeed, we often use music to remind ourselves of past valued events, a fact that indicates that music may play an important nostalgic function in everyday life. Finally, musical expectancy corresponds to a mechanism by which an emotion is induced because a feature of the music structure delays, violates or confirms our expectations concerning the continuation of the music. The expectations are built upon the previous experience with the music style. For instance, in the Western tonal system, the progression of E-F# prompts the expectation that the music will continue with G#, such that when this does not happen the listener may become surprised. These mechanisms would operate singularly or in a combined fashion while we listen to music. According to the authors (Juslin et al., 2010; Juslin & Västfjäll, 2008), each of them would correspond to distinct brain functions. This theoretical framework is preliminary, and so the precise characteristics of the mechanisms are to be examined. However, it is useful to derive specific hypotheses and to guide empirical research on musical emotions. 3.6. NEUROCOGNITIVE BASIS OF MUSICAL EMOTIONS In the last years, an increasing number of studies have examined the neural foundations of musical emotions (for reviews, Koelsch, 2010; Peretz, 2010). There is evidence that processing emotions in music involves subcortical brain systems, namely limbic structures and the striatum, as well as cortical systems, namely the temporal lobe,
Chapter III. Music and Emotion ! 75 orbitofrontal cortex, ventromedial prefrontal cortex, anterior cingulate cortex and the insula. Seminal evidence for the involvement of subcortical systems comes from a study by Blood and Zatorre (2001). They used positron emission tomography (PET) to measure regional cerebral blood flow changes while participants listened to 90 s of their own favourite music, to which they usually had highly pleasurable experiences of “shiversdown-the-spine” or “chills”. More intense chills correlated with cerebral blood flow decrease in the amygdala and left hippocampus, and with increase in the ventral striatum, dorsomedial midbrain and thalamus (chills also modulated cortical activity, with cerebral blood flow decreasing in the ventromedial prefrontal cortex, and increasing in the insula, orbitofrontal cortex and cingulate cortex; the functional role of these cortical regions is discussed below). The subcortical structures recruited by music are also recruited for reward responses to biologically relevant stimuli, which are important for survival, such as food and sex. Thus, as highlighted by the authors (ibd.), the results of this study indicate that the highly abstract sounds of music can engage ancient emotion and reward systems in the brain. Recently, Salimpoor, Benovoy, Larcher, Dagher, and Zatorre (2011) showed that the recruitment of the striatum by music-evoked chills reflects the modulation of dopaminergic activity. They found endogenous dopamine release in the striatum during music-induced peak emotional arousal, by using the neurochemical specificity of [11C]raclopride PET scanning, combined with psychophysiological measures of autonomic nervous system activity. The authors further observed, in a fMRI study with the same stimuli and listeners, a functional dissociation within the striatal system: the caudate was chiefly involved in the anticipation of reward, and the nucleus accumbens in the actual experience of intense emotional responses to music. Musical emotions recruit subcortical brain systems even when participants do not experience intense chills. For instance, Koelsch, Fritz, Cramon, Müller, and Friederici (2006) compared neural responses to consonant joyful instrumental dance-tunes (pleasant music) and to electronically manipulated dissonant music (unpleasant music) in an fMRI study. They detected subcortical activations for both types of music: pleasant consonant music activated the ventral striatum, and unpleasant dissonant music activated the amygdala, hippocampus and
Chapter III. Music and Emotion ! 76 parahippocampal gyrus. Similarly, Mitterschiffthaler, Fu, Dalton, Andrew, and Williams (2007) found that happy music activates the ventral and dorsal striatum, whereas sad music activates the hippocampus and amygdala. Thus, positively valenced responses to music appear to engage the striatum, through the modulation of dopaminergic activity, and relatively more negative emotional responses appear to engage medial temporal structures. Neuropsychological studies indicate that core subcortical neuroaffective systems are also critical for emotion recognition in music, not only for emotional experience. In three studies, Gosselin and colleagues (2005, 2007, 2011) found that patients with damage in the amygdala had impaired recognition of fear and sadness in instrumental music excerpts, while the recognition of positive emotions, happiness and peacefulness, was relatively preserved. This pattern could not be explained by defects in low-level music perception abilities. In another study, it was found that damage in the parahippocampal cortex correlated with abnormal perception of dissonance in music, such that patients rated dissonant music as slightly pleasant, whereas healthy controls rated it as unpleasant (Gosselin et al., 2006). The more extensive the patients’ damage was, the more indifferent they were to dissonance (see also Khalfa et al., 2008). Recently, Omar and colleagues (2011) examined emotion recognition in music (happiness, sadness, anger and fear) in patients with frontotemporal lobar degeneration, and explored associations between behavioural performance and grey matter losses. Patients’ performance was impaired and correlated with grey matter volumes in an extensive bilateral cerebral network that included subcortical structures, more precisely the parahippocampal gyrus, hippocampus, amygdala, nucleus accumbens and ventral tegmentum (the network included cortical structures as well, notably the insula, anterior cingulate, orbitofrontal cortex, medial prefrontal cortex, dorsal prefrontal, inferior frontal, anterior and superior temporal cortices, fusiform gyrus and posterior parietal cortices). Cortical systems comprise neural structures that are evolutionarily more recent. It has been shown that they play different functional roles in musical emotions. Auditory regions in the superior temporal cortex, particularly in the right hemisphere, support the
Chapter III. Music and Emotion ! 77 processing of low-level perceptual aspects of music, thus setting the stage for emotional responses and interpretative processes (Mitterschiffthaler et al., 2007; Peretz, Blood, Penhune, & Zatorre, 2001; Zatorre, Belin, & Penhune, 2002). There is solid evidence that the orbitofrontal cortex is a key region for the processing of emotions in music. It is recruited when participants experience intense chills (Blood & Zatorre, 2001), when they listen to music without feeling intense responses (Blood, Zatorre, Bermudez, & Evans, 1999; Menon & Levitin, 2005), and when they perform overt emotion recognition tasks (Omar et al., 2011). This region, which is connected with the amygdala and the striatum, appears to play a role in the analysis and integration of the reward value of emotional stimuli across sensory modalities, thereby guiding behavioural judgments (Blood & Zatorre, 2001; Menon & Levitin, 2005; Peretz, 2010). The ventromedial prefrontal cortex is also an important structure for musical emotions. It is active during chill experiences (Blood & Zatorre, 2001), during emotion recognition (Omar et al., 2011), and might be involved in triggering autonomic emotional responses to music. Johnsen, Tranel, Lutgendorf, and Adolphs (2009) found that patients with damage to the ventromedial prefrontal cortex are impaired in the generation of skin-conductance responses to music. Damasio’s somatic marker hypothesis focuses on this region (ibd.). Finally, the anterior cingulate and the insular cortex were also shown to be involved in music-evoked chills (Blood & Zatorre, 2001), music listening (Koelsch et al., 2006; Mitterschiffthaler et al., 2007) and emotion recognition (Omar et al., 2011). These structures seem to be implicated in coupling autonomic responses and subjective feeling states, as well as in motor-related functions and in the synchronization of biological subsystems during emotion episodes (such as physiological arousal, motor expression, motivation, monitoring processes and cognitive appraisal; Koelsch, 2010). 3.7. INDIVIDUAL DIFFERENCES Which factors explain differences across individuals in emotional responses to music? Personality traits might be important. For instance, Juslin and colleagues (2008)
Chapter III. Music and Emotion ! 78 observed that the prevalence of music-evoked emotions correlates with personality traits. Specifically, they found that higher levels of neuroticism are associated with more frequent occurrence of pleasure-enjoyment in response to music. They also observed that higher levels of extraversion are associated with higher overall prevalence of musical emotions, suggesting the extraverts and introverts use music differently. Barrett and colleagues (2010) found that individuals’ proneness to nostalgia predicts the intensity of music-evoked nostalgia, such that individuals with higher nostalgic tendencies are more likely to experience higher levels of nostalgia in response to music. Nostalgia proneness, in turn, correlated with neuroticism. Recently, Montag, Reuter, and Axmacher (in press) showed that persons who consider themselves as easily getting absorbed by arts and music (subscale “self-forgetfulness” of the character dimension “self-transcendence”) have decreased activity in the striatum while listening to their favourite songs. Individual factors might also modulate the ability to recognize emotional expressions in music. Mohn, Argstatter, and Wilker (in press) failed to detect significant associations with personality traits, but Vuoskoski and Eerola (2011) observed that the perception of sadness in film music excerpts correlates with higher neuroticism and with lower extraversion. Resnicow and colleagues (2004) found that individuals higher in emotional intelligence, as assessed by the Mayer-Salovey-Caruso Emotional Intelligence Test, are better in the recognition of emotions in music performance. We pointed out that ageing modulates emotion recognition in emotional prosody and in other emotional stimuli (e.g., facial expressions; body postures). Whether this is also the case for musical emotions remains poorly studied. In the only study that approached this question so far, Laukka and Juslin (2007) found that older adults are less accurate than younger ones in the recognition of negative emotions in music performance. Musical training might be another source of individual variability, but available evidence is inconclusive: while some studies reported suggestive evidence that musical training is associated with enhanced sensitivity to musical emotions (Bhatara et al., 2011; Livingstone, Muhlberger, Brown, & Thompson, 2010), other studies found no effects (Bigand et al., 2005; Resnicow et al., 2004). The impact of ageing and musical training
Chapter III. Music and Emotion ! 79 in emotion recognition in music is examined in the empirical part of this thesis (Chapter VI). 3.8. CLINICAL APPLICATIONS The therapeutic value of music in clinical settings has long been recognized. The scientific framework and rationales for music interventions have been altered on the basis of recent developments in the cognitive neuroscience of music. According to the model developed by the Center for Music Therapy Research in Heidelberg (Viktor Dulger Institute), there are five factors or ingredients by which music can improve the psychological and physiological health of children and adults: attention modulation, emotion modulation, cognition modulation, behaviour modulation and communication modulation (Hillecke, Nickel, & Bolay, 2005). These ingredients reflect the idea that music listening and music making engage a multitude of behaviours and brain structures related to cognitive, sensorimotor and emotional processing (e.g., Koelsch & Siebel, 2005). Thus, music could be used therapeutically to modulate activity at different levels and in different domains. For instance, it was shown that listening to relaxing music, as compared to silence, facilitates recovery from a psychologically stressful task; it leads to the reduction of salivary cortisol levels (Khalfa, Dalla Bella, Roy, Peretz, & Lupien, 2003). Recently, Särkämö and colleagues (2008) conducted a single-blind, randomized and controlled trial to determine the effects of everyday music listening in the recovery of cognitive and emotional functions after middle cerebral artery stroke. It was observed that two months of daily music listening (music was self-selected), as compared with listening to audio books or no listening material, led to significant improvements in verbal memory, focused attention, depression and confused mood. These effects are at least partly mediated by the emotional component of music. As noted by Koelsch (2010), because music modulates the activity of core emotion systems, such as the amygdala, hippocampus and striatum, it can in principle be used to
Chapter III. Music and Emotion ! 80 stimulate and regulate these systems when they are impaired (e.g., depression; PD; posttraumatic stress disorder). In an exploratory study, Koelsch, Offermanns, and Franzke (2010) observed that a music-therapeutic method involving music making in groups can produce positive changes in mood, as indexed by decreased depression/anxiety, decreased fatigue and increased vigour. These changes were accompanied by an increase in the experience of positive emotions (higher ratings for happiness) and by a decrease in negative emotions (lower ratings for anger, sadness and anxiety). Other studies suggest that musical-rhythmic stimuli can be used to regulate motor functions, namely gait and arm control, in patients with stroke and PD. The mechanisms behind these effects appear to be arousal and priming of the motor system, as well as timing and entrainment of the motor system (Thaut et al., 2009). Music therapeutic exercises can also be used effectively to improve the “feeling” experience of emotion, the identification of emotion, the expression of emotion, the understanding of emotional communication of others, and the synthesis, control and modulation of emotional behaviours (Thaut & Wheeler, 2010). Despite recent advances, the knowledge on the specific mechanisms that mediate the therapeutic effects of music is still limited. Research on the neurocognitive basis of music and on how music modulates brain plasticity will be essential to inform concepts and practices in music therapy. A more complete understanding of how music relates to other cognitive functions, namely speech and language, will also be illuminating for potential therapeutic applications. In the next chapter we discuss the relations between emotions in speech prosody and in music.
! 81 CHAPTER IV Perceiving Emotions in Speech Prosody and in Music: How Close?
! 82
Chapter IV. Perceiving Emotions in Speech Prosody and in Music: How Close? ! 83 The reason why music conveys specific emotions, induces mood changes and evokes emotional reactions remains puzzling. On one hand, it does not serve any obvious survival or reproduction functions. Possible adaptive roles of musical behaviours have been suggested, notably sexual selection, social cohesion and socio-cognitive development, but these are controversial (for reviews, Koelsch, 2011; Patel, 2008a). On the other hand, in everyday life, emotions are typically a response to events that are appraised as having the potential to interfere with our goal-directed actions, and music rarely either serves or blocks goals (Konecni, 2008; Scherer, 2004). In other words, while in most cases emotions are utilitarian (they have functions in the adaptation and adjustment of individuals), in music emotions are aesthetic (they are not triggered by appraisals concerning goal revelance, instead they derive from the intrinsic qualities of music; Scherer, 2004). One hypothesis is that musical emotions are an accidental by-product of capacities originally dedicated to other purposes. Music could have co-opted the neurocognitive machinery shaped by evolution for sounds of greater biological relevance, namely vocal emotions – speech prosody and nonverbal vocal expressions (Hauser & McDermott, 2003; Juslin & Laukka, 2003; Juslin et al., 2010; Nieminen et al., 2011; Patel, 2008b; Peretz, 2010; Pinker, 1997). Other theories suggest that music and language evolved as two different specializations of a common predecessor, which consisted of a vocalization system that subserved rudimentary referential-emotive communicative functions. This primitive system would have been highly adaptive for our ancestors and some of its features would still be present in the current forms of both music and
Chapter IV. Perceiving Emotions in Speech Prosody and in Music: How Close? ! 90 hypothesis awaits clarification. Another kind of study consists of focusing on individuals with a neuropathological condition that impairs the recognition of emotions in speech prosody, such as receptive aprosodia caused by stroke (e.g., Ross & Monnot, 2008) or PD (e.g., Pell & Leonard, 2003). If these individuals have a similar profile of impairments for emotion recognition in prosody and music, this would suggest that the damaged neural structures are shared by both domains. A differential profile of impairments, on the other hand, would be evidence that both domains are partly independent. Up to now, such a comparative design has never been implemented. According to Patel (2008b), this approach would be a powerful test to the resourcesharing hypothesis, provided that some methodological aspects are met: patients should be tested in other modalities, like facial expressions, to guarantee that the impairment is specific to the auditory domain; possible low-level defects in auditory processing should be excluded; and cognitive abilities (e.g., memory) should be controlled for. These studies will contribute to a better understanding of the neurocognitive relationship between emotions in speech and music. They can also be informative for applied psychological science. If both domains set into action common mechanisms, music therapy might be a useful device in language rehabilitation programs.
! 91 CHAPTER V Goals of this Thesis
! 92
Chapter V. Goals of this Thesis ! 93 In the preceding chapters, we have reviewed research focusing on emotion theories, emotional speech prosody, musical emotions, and on the parallels between emotion processing in speech prosody and in music. Several issues that we have referred to remain poorly understood, and they lie at the basis of the empirical part of this thesis. We highlighted that musical emotions show some of the same properties of emotion modalities of unequivocal biological and social importance, namely facial expressions and speech prosody. For instance, as for these modalities, the recognition of discrete emotions in music is accurate, quick, consistent, precocious in ontogenetic development and governed by universal principles to a significant extent. It is well established that ageing operates changes in emotion processing across different modalities (e.g., Ruffman et al., 2008), but so far it is unclear whether and how ageing affects also musical emotions. In other words, we do not know whether emotion processing in music undergoes developmental changes during the adult life, as emotion processing in other domains. It is also unclear whether training in music impacts on the processing of musical emotions, as it appears to impact on several aspects of brain and cognitive functioning (e.g., Besson et al., 2011; Habib & Besson, 2009). Concerning emotional speech prosody, we described studies showing that perceiving emotions in this modality is determined by a combination of universal and language-specific processes. Thus, it is of paramount importance to provide prosodic materials and perceptual studies in different languages. To the best of our knowledge, this has not been done for European Portuguese. Finally, we reviewed comparative research on emotional prosody and musical emotions, and concluded that there are striking cross-domain commonalities regarding the acoustic encoding of emotional qualities. We also pointed out that a currently debated hypothesis in cognitive neuroscience is that emotion processing in
Chapter V. Goals of this Thesis ! 94 speech prosody and in music recruits shared neurocognitive mechanisms. However, empirical research on this hypothesis is scarce. We set out to investigate these open questions in a series of empirical studies using behavioural and neuropsychological approaches. As we underlined in the review of the literature, emotion processing involves multiple components. It is therefore important to define at the outset what exactly we will assess. In all studies we focus in the overt recognition of discrete emotions, that is, in how individuals are able to identify and evaluate explicitly emotional qualities expressed by music excerpts and by speech samples. We resort to behavioural tasks that are widely used in emotion recognition literature, namely forced-choice and rating tasks. In conventional forced-choice tasks, participants are given a predefined list of emotions and select the one that best matches the emotion expressed by the stimulus. This task has been strongly defended by some authors (Ekman, 1994), though it may present some difficulties. When the number of options is small, results may reflect discrimination between alternatives rather than true emotion recognition (Banse & Scherer, 1996; Scherer, 2003; Scherer et al., 2003). Another criticism is that this task might inflate agreement rates because the response options are given beforehand to the participant (Sauter, 2006). This potential problem is somewhat mitigated in the studies presented here because we are chiefly interested in differences between groups, not so much in agreement rates per se. In rating tasks, participants use quantitative intensity scales to evaluate the degree to which different emotions are expressed by the stimulus. In a typical situation, participants are asked to rate how much each stimulus expresses every emotion included in the study (e.g., Adolphs & Tranel, 1999; Adolphs, Tranel, Damasio, & Damasio, 1995; Gosselin et al., 2005; Pell & Leonard, 2003; Trimmer & Cuddy, 2008). In the following paragraphs we briefly outline the empirical studies that were conducted.
Chapter V. Goals of this Thesis ! 95 Do ageing and musical expertise modulate emotion recognition in music? Little is known about individual differences in how we respond to music. One of the goals of this thesis is to examine the roles of ageing and musical training in emotion recognition in music. For modalities like facial expressions, voice and body postures, it has been demonstrated that advancing age is associated with decreased accuracy in the recognition of some emotions, particularly negative ones (e.g., Mill et al., 2009; Ruffman et al., 2008). Studies on facial expressions have also indicated that training may improve the ability to recognize emotions (Elfenbein, 2006; Matsumoto & Hwang, 2011; Silver, Goodman, Knoll, & Isakov, 2004). What about musical emotions? In two studies, presented in Chapter VI, we determined how emotion recognition in instrumental music excerpts changes along the adult life span, and whether individuals with musical training respond differently than those without training. We used an emotion rating task in both studies, and employed stimuli from the database developed and validated by Vieillard and colleagues (2008) for research on musical emotions. How are emotions recognized in emotional prosody in Portuguese? In the study presented in Chapter VII, we developed and validated a database of prosodic stimuli in Portuguese. The availability of this kind of resources is valuable for research with Portuguese-speaking participants and for cross-language studies. Analogous databases have been devised for different languages (e.g., Pell, 2002; Staroniewicz & Majewski, 2009; Wu et al., 2006). The verbal materials consisted of short sentences with emotionally neutral semantic content and of pseudo-sentences (utterances that include pseudo-words). Two speakers recorded these materials varying prosody in order to express neutrality and six emotion categories: anger, disgust, fear, happiness, sadness and surprise. In order to validate the stimuli, we performed acoustic analyses and conducted a perceptual experiment using a forced-choice emotion
Chapter V. Goals of this Thesis ! 96 recognition task. Reaction times and intensity judgments were also collected and analysed. Does emotion recognition in speech prosody and in music recruit shared neurocognitive mechanisms? As we emphasized in the review of the literature, we can use different research strategies to examine the extent to which emotion processing in speech and music engages common mechanisms. One consists of determining cross-domain transfer effects from musical training to emotional prosody. Available evidence on this issue is inconclusive (Thompson et al., 2004; Trimmer & Cuddy, 2008). In the study presented in Chapter VIII, we hypothesized that if neurocognitive mechanisms are shared across domains, then musicians should show enhanced emotion recognition in speech prosody. To investigate this hypothesis, we compared highly trained musicians with untrained participants in a forced-choice emotion recognition task using the prosodic stimuli previously validated. Another research strategy consists of determining whether a neuropathological condition that affects speech prosody also affects musical emotions in a similar manner. In the last empirical study, Chapter IX, we adopted this approach in a neuropsychological study on patients with PD. The recognition of emotions in speech prosody may be impaired in these patients, due to their dysfunction in the basal ganglia (Gray & Tickle-Degnen, 2010; Pell & Leonard, 2003). Whether the recognition of emotions in music also depends critically on the basal ganglia is still poorly understood (Omar et al., 2011), as is whether PD patients have difficulties in dealing with musical emotions (Tricht, Smeding, Speelman, & Schmand, 2010). We tested patients and matched healthy controls in an emotion rating task for music and speech prosody, and compared the effect of PD across domains. Participants were also assessed for emotion recognition in facial expressions, taken as a control measure to exclude general deficits in emotion recognition. Furthermore, patients and controls underwent an extensive
Chapter V. Goals of this Thesis ! 97 neuropsychological evaluation to ascertain that general cognitive or low-level perceptual problems would not account for putative impairments in emotion recognition.
! 98
! 99 PART 2 EMPIRICAL STUDIES
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 106 inhibition over the processing of negative emotional input (due to regulatory strategies), and uninhibited processing of positive input. The greater recruitment of frontal cortex by older adults in response to emotional stimuli is a consistent finding in the literature (St. Jacques, Bessette-Symons, & Cabeza, 2009; St. Jacques et al., 2008). It has also been shown that amygdala activations decrease with age for negative but not for positive information (Gunning-Dixon et al., 2003; see also, Kisley, Wood, & Burrows, 2007). An alternative explanation for ageing effects on emotion recognition stresses the role of neuropsychological deterioration in brain structures implicated in emotion processing. The recognition of some emotions would undergo changes and others would remain stable because rates of deterioration differ across neural systems of emotion (e.g., Raz et al., 2005). For instance, decline in the cingulate cortex and in the amygdala might contribute to explain changes in the recognition of fear and sadness in faces, and the relative preservation of basal ganglia structures might explain the stability observed for disgust (Calder et al., 2003; Ruffman et al., 2008). According to Cacioppo, Berntson, Bechara, Tranel, and Hawkley (2011), brain decline could also account for the positivity effect: age-related deterioration of the amygdala, which is important to monitor negative information, would lead to the reduction of negative affect. Moreover, it is possible that the increased frontal activations in older adults reflect compensation for brain decline, instead of motivation-related regulatory strategies (St. Jacques et al., 2009). Overall, the relative contribution of general cognitive decline, motivational changes, and brain deterioration, for age-related differences in emotion recognition is still a contentious issue. Because most research was conducted with visual stimuli, studies on other modalities are crucial to examine the generality of the effects and to provide insights into their causes. Does emotion recognition in music change with advancing age, undergoing similar developmental modifications as those of other emotional stimuli relevant for social functioning and communication? If so, are changes towards positivity, with responsiveness to negative emotions declining more than responsiveness to positive ones? Are the effects a consequence of general cognitive decline, or primary in origin? Here we are interested in these questions. It has been shown that musical emotions may exhibit some of the properties of other emotion domains (see Chapter III), but the
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 107 modulations related to ageing remain largely unexplored. In a first systematic approach to the topic, Laukka and Juslin (2007) compared younger and older adults, over 65 years of age, in the recognition of emotions in music and in speech prosody. Musical emotions were communicated only through expressive cues in performance: the same excerpt (the theme from W. A. Mozart’s Piano Sonata in A Major, K331) was played with differences in features like loudness, articulation, vibrato and phrasing, to express anger, fear, happiness, sadness and neutrality. Older adults had lower recognition rates than younger ones for the negative emotions of sadness and fear, both in music and in speech. No differences were observed for happiness, neutrality and anger. These results indicate that ageing may produce similar effects in emotion recognition in music and in speech prosody. The emotions which were hardest to recognize were not those that underwent age-related changes. Thus, general cognitive decline might not be the primary cause of the effects. However, several questions are to be examined. First, it is unknown how ageing impacts emotions expressed through music compositional structure, i.e., emotions encoded in the musical notation by the composer. Variations in structural features such as mode, tempo, pitch range and dissonance, determine the expressed emotion (e.g., Dalla Bella et al., 2001; Gabrielsson & Lindström, 2010; Hunter, Schellenberg, & Schimmack, 2008) and predict the listener’s subjective judgments (e.g., Balkwill & Thompson, 1999). Second, it remains to be determined whether effects for music are observed already in middle-aged, as described for other emotion modalities (e.g., Calder et al., 2003; Isaacowitz et al., 2007; Paulmann et al., 2008b). Third, possible associations between emotion recognition in music and general cognitive abilities need to be inspected to determine directly whether age differences are mediated by general, emotion-unspecific, factors. Furthermore, the study by Laukka and Juslin (2007), as most studies, included an unbalanced number of positive and negative emotions (only one positive emotion, as compared with three negative ones). This undermines conclusions regarding effects of valence. MUSICAL EXPERTISE AND EMOTION RECOGNITION IN MUSIC | Musical expertise might also be a source of individual differences in emotion recognition in music. Longitudinal and cross-sectional studies indicate that learning music changes brain morphology and cognitive processing (for a review, Habib & Besson, 2009). For
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 108 instance, Hyde and colleagues (2009) showed that 15 months of lessons in early childhood (mean age at start of study = 6.32 years) leads to behavioural improvements in melodic and rhythmic discrimination tasks, and to structural changes in auditory and motor brain areas. Adult professional musicians, as compared to amateur musicians and non-musicians, have increased grey matter volume in the primary motor and somatosensory areas, premotor areas, anterior superior parietal areas, inferior temporal gyrus, cerebellum, Heschl’s gyrus and inferior frontal gyrus (Gaser & Schlaug, 2003). Training impacts on non-musical cognitive abilities, such as discrimination of F0 in speech (e.g., Moreno et al., 2009), nonverbal reasoning (Forgeard, Winner, Norton, & Schlaug, 2008) and executive control (Bialystok & DePape, 2009). It is plausible that it also fine-tunes mechanisms underlying the processing of musical emotions because musical practice focuses extensively on the ability to deal with fine-grained modulations of acoustic features for expressive purposes, as well as on the relations between musical structure and the instantiation of emotions. This hypothesis appears intuitive, but delineating which specific musical abilities are enhanced in musicians is not a trivial endeavor, particularly in light of findings showing that untrained listeners perform like experts in some music processing tasks (e.g., learning new musical idioms) – they acquire sophisticated musical knowledge just through exposure (for a review, see Bigand & Poulin-Charronnat, 2006). Indeed, the study by Bigand and colleagues (2005) suggests that emotional responses to music might be largely similar in musically trained and untrained listeners. They asked participants to group music excerpts according to the emotion they evoked, and the number of groups produced was independent of training. The only effect of musical expertise was in enhancing the consistency of responses across testing sessions in a test-retest situation. However, this study assessed musical emotions at an implicit level, and explicit task conditions might be more sensitive to expertise effects (Bigand & Poulin-Charronnat, 2006). However, results for explicit tasks are scarce. In an exploratory analysis (trained participants, n = 9), Livingstone, Muhlberger, Brown, and Thompson (2010) found a significant correlation between the number of years of training and recognition accuracy for music excerpts that conveyed emotions through both structural and expressive cues. These results suggest that training does play a role, though it is difficult to ascertain whether the effect is driven by a better ability of musicians to process the expressive and/or the
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 109 structural emotion cues. Bhatara and colleagues (2011) found that musical training is associated with enhanced sensitivity to differences in expressive musical cues, namely timing and amplitude (n = 10 in one experiment + n = 10 in another experiment). Resnicow, Salovey, and Repp (2004), on the other hand, failed to find such an effect (number of trained participants unspecified). Hence, available results are mixed. To the best of our knowledge, no studies have addressed the impact of expertise for emotions expressed by structural features alone. Furthermore, prior research was not able to disentangle putative specific effects on musical emotions from domain-general cognitive benefits of training (e.g., Bialystok & DePape, 2009; Schellenberg, 2006) THE CURRENT STUDIES | The goal of the two following studies is to determine how ageing and musical expertise modulate emotion recognition in instrumental music. We also aim to establish whether these effects are independent of general cognitive differences. Stimuli were music excerpts previously validated to express two positive emotions, happiness and peacefulness, and two negative ones, sadness and fear/threat (Vieillard et al., 2008). These excerpts were unfamiliar to the participants – they were composed for experimental purposes. Emotions were communicated through compositional structural cues only. Participants rated how much each excerpt expressed each of the four emotions on intensity scales. This task does not involve a selection of a single category, and so it is less prone to the response biases that can affect conventional forced-choice tasks (e.g., Isaacowitz et al., 2007). Additionally, it mitigates the high decisional demands posed by forced choices, which might augment the impact of older adults’ general cognitive difficulties. Such a task was previously used in studies on emotion recognition in facial expressions (e.g., Adolphs, Schul, & Tranel, 1998), speech prosody (e.g., Péron et al., 2010), nonverbal vocal expressions (Sauter, Eisner, Calder, et al., 2010) and music (Gosselin et al., 2005, 2007). In Study 1, we focused primarily on age-related changes. We covered the full range of the adult life span, including younger, middle-aged and older participants. Based on the reviewed literature, we predicted that advancing age would produce decreased responsiveness to musical emotions, particularly to the negative ones (sadness and fear/threat). These changes would be significant already in middle-age, as previously observed for other modalities. Because a subset of participants had some degree of musical training, we
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 110 also undertook a preliminary examination of whether that training is associated with enhanced sensitivity to musical emotions. In Study 2, musical expertise and age were manipulated orthogonally. We compared expert musicians and musically untrained listeners (hereafter, controls) from two age cohorts, young and middle adulthood. In addition to emotion recognition, they were assessed for domain-general cognitive abilities, and completed brief control measures of personality and socio-communicative traits, because these may influence emotion processes (e.g., Hamann & Canli, 2004; Mill et al., 2009). We expected an advantage of musicians over controls for emotion recognition. If ageing and expertise effects on musical emotions are primary in origin, they should be independent of general cognitive and personality measures. In both Study 1 and Study 2, multiple regression analyses were conducted to explore how structural cues of music excerpts predict participants’ subjective emotion ratings, and whether age and expertise effects are linked with differences in weighting these cues. 6.2. STUDY 1 6.2.1. Method PARTICIPANTS | A total of 114 healthy adults (67 female) volunteered to take part in this study. They were aged between 17 and 84 years, and were categorized into three groups with 38 participants each: younger, middle-aged and older adults. Table 3 presents their demographic and background characteristics. The younger adults were undergraduate and graduate students from University of Porto; middle-age and older ones came from several local communities, including senior universities. The three groups were matched for education level, as assessed by the number of years attending school (p > .05). All participants had normal or corrected-to-normal vision, and none reported head trauma, substance abuse, cognitive or hearing difficulties, nor major psychiatric or neurological illnesses. For older adults, possible cognitive decline was evaluated by using the cognitive screening test Mini-Mental State Examination, MMSE
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 111 (Folstein, Folstein, & McHugh, 1975; Portuguese version, Guerreiro, Silva, Botelho, Leitão, & Garcia, 1994); they performed near ceiling, 29.3 on average (maximum 30), and none scored below 27. Older participants also completed a pure-tone audiometric screening to guarantee that they had acceptable hearing thresholds (at least 30 dB HL for frequencies of 500 Hz, 1000 Hz, 2000 Hz and 4000 Hz at left and right ears). Thirtythree participants reported having had some kind of formal musical training, including learning how to play an instrument; mean of 5.1 years of training (SD = 3.9; range = 1 – 14; age when training began, M = 11 years). Table 3. Demographic and background characteristics of the participants in each age group (Study 1). SDs are presented in parentheses. Age group Younger Middle-aged Older Age (years) 21.8 (3.5) 44.5 (6.2) 67.2 (6.2) Age range (years) 17 - 29 35 - 56 60 - 84 Gender 17F / 21M 21F / 17M 29F / 9M Education (years) 15.5 (2.2) 17.0 (3.7) 15.6 (3.1) Mini-Mental State Examination (/30) - - 29.3 (1.0) Musical training (years) 5 (4.1) 4.8 (4.2) 5.8 (2.8) Participants with musical training (n) 18 10 5 MUSIC STIMULI | The stimuli were 56 music excerpts validated for research on emotions by Vieillard and colleagues (2008). These excerpts were composed to express four emotions: happiness, peacefulness, sadness and fear/threat, 14 stimuli per category. They consist of a melody with accompaniment, follow the rules of the Western tonal system, and were produced in piano timbre by a digital synthesizer. The emotional information is expressed exclusively by music compositional structure, through variations in features such as mode, dissonance, pitch range, tone density, rhythmic regularity and tempo. Excerpts do not vary in performance-related expressive features like dynamics, vibrato, phrasing or articulation. The mean duration of the excerpts is 12.4 s. They were downloaded at the Internet site from the Isabelle Peretz Research Laboratory (www.brams.umontreal.ca/plab/publications/article/96). Perceptual
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 112 validation procedures including categorization, gating and dissimilarity judgments, confirmed that the intended emotions are perceived rapidly and with high agreement rates between listeners, high accuracy (Vieillard et al., 2008). Emotions can be recognized in these excerpts universally, that is, in the absence of prior exposition to the Western music (Fritz et al., 2009). The have also been used successfully in neuropsychological studies with patients with focal brain damage (Gosselin et al., 2005, 2007, 2011), neurodegenerative disorders (Drapeau, Gosselin, Gagnon, Peretz, & Lorrain, 2009), as well as with neurotypical children in developmental studies (Hunter et al., 2011). The stimuli were pseudo-randomized and divided into two blocks of 28 trials each; the presentation order of the blocks was counterbalanced across participants. Two additional stimuli were used as practice trials: 15 s from Debussy’s Golliwog’s Cakewalk, for happiness, and 18 s from Rachmaninoff’s 18th Variation from the Rhapsody on a Theme by Paganini, for peacefulness. PROCEDURE | Participants were tested individually in a single experimental session lasting about 30 minutes; exceptionally, some undergraduate students were tested in groups of two or three. They were told that they would listen to short excerpts of music expressing different emotional states, namely happy (alegre), peaceful (sereno), sad (triste), and/or fearful/threatening (assustador). All participants were well familiarized with the four emotion labels. They rated how much each stimulus expressed each of the four emotions on 10-point scales, from 0 (absent, ausente) to 9 (present, presente). It was stressed that participants should always respond to the four emotion scales, as they could perceive different emotions simultaneously, with similar or different degrees of intensity. For instance, even if they perceived a stimulus as happy, they should rate it not only with respect to happiness, but also with respect to the possible expression of sadness, fear and peacefulness. Each stimulus was presented only once and no feedback was given. The task started with practice trials, followed by the two blocks of stimuli. Stimulus presentation was controlled by SuperLab version 4.0 software (Abboud, Schultz, & Zeitlin, 2006), running on an Apple MacBook Pro computer. Excerpts were presented via high-quality loudspeakers, adjusted to a comfortable volume level for
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 113 each participant. Responses were collected via a pencil-and-paper questionnaire. Participants were encouraged to respond fast and spontaneously (they did not need to wait until the end of the excerpt to respond). The inter-stimulus interval was fixed for younger and middle-aged participants, 6 s, and contingent upon response for older ones, because some of them needed more than 6 s to provide the four ratings. In the same experimental session, participants completed a brief questionnaire exploring their demographic, background, and musical profile. Older adults also performed the audiometric screening and the MMSE before the experimental task. 6.2.2. Results To compare emotion recognition across age groups we derived a measure of accuracy from the raw ratings, based on the emotion that received the highest rating. For each subject and excerpt, when the highest of the four ratings matched the intended emotion, the response was coded as an accurate categorization. When the highest rating corresponded to a non-intended emotion, the response was considered inaccurate. When the highest rating corresponded to more than one emotion (e.g., giving 8 for sadness and also for peacefulness, and lower ratings for the other two categories), the response was considered ambivalent. Ambivalent responses indicate that, in a given excerpt, participants perceived more than one emotion with the same strength. This derived measure of accuracy was used previously in studies on emotion perception in music (Gosselin et al., 2005, 2007; Vieillard et al., 2008) and in prosody (e.g., Adolphs et al., 2002). It controls for the possibility that participants of different ages used the scales differently, and takes into account judgments both for the intended and nonintended emotions: a response is accurate only if the rating for the intended emotion is higher than the ratings for all the non-intended emotions.
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 114 Table 4 presents the percentage of accurate categorizations for each emotion, and the distribution of inaccurate and ambivalent responses, separately for younger, middleaged and older participants. Accuracy rates were always higher for the intended than for the non-intended categories, indicating that the highest ratings were assigned consistently to the intended emotions. Agreement rates ranged between 33% (fear in older participants) and 97% (happiness in younger ones), as can be seen in the diagonal lines of the Table 4 (bold cells). This confirms that the stimuli effectively communicated the four emotions. Table 4. Distribution of responses (%) for each intended emotion as a function of age (Study 1). Diagonal cells in bold indicate accurate categorizations, i.e., match between the highest rating and the intended emotion. Ambivalent responses represent situations in which the highest rating was assigned to more than one emotion. Standard errors are presented in parentheses. Age group / excerpt type Distribution of responses (%) Happy Peaceful Sad Scary Ambivalent Younger Happy 97 (1.1) 1 0 0 2 Peaceful 12 44 (4.9) 33 0 11 Sad 1 9 79 (3.6) 3 8 Scary 6 1 8 76 (2.9) 8 Middle-aged Happy 94 (1.9) 2 1 0 4 Peaceful 9 65 (4.6) 11 0 14 Sad 2 24 58 (4.8) 2 14 Scary 13 5 22 45 (5.5) 14 Older Happy 88 (3.1) 5 1 0 5 Peaceful 13 57 (4) 15 0 15 Sad 2 33 45 (4.3) 2 18 Scary 18 10 23 33 (4.1) 15 EFFECTS OF AGE ON ACCURACY | To examine how age modulates accuracy, correct categorizations per emotion were arcsine transformed and submitted to an ANCOVA, with emotion as repeated-measures factor (happy, peaceful, sad and scary), age as between subjects-factor (younger, middle-aged and older), and years of musical training as covariate. Musical training was partialled out here and in all the analyses on age
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 115 effects because the number of trained participants was unbalanced across age groups (see Table 3). Main effects and interactions were followed-up with post hoc Tukey HSD tests. Effect sizes are expressed as partial eta squared (!p2). The four emotions elicited significantly different accuracy rates: happiness reached the highest accuracy (93%, ps < .01), followed by sadness (61%) and peacefulness (55%), with similar rates (p > .1); fear had lower accuracy than happiness and sadness (52%, ps < .01), but not peacefulness [p > .5; main effect of emotion, F(3,330) = 90.16, p < .0001, !p2 = .45]. As predicted, age groups differed in emotion recognition, but not uniformly across categories [main effect of age, F(2,110) = 11.78, p < .0001, !p2 = .17; interaction Age x Emotion, F(6,330) = 12.71, p < .0001, !p2 = .19]. For the negative emotions, sadness and fear, middle-aged and older participants were less accurate than younger ones (ps < .001); middle-aged and older participants did not differ significantly (ps > .3). For the positive emotions, happiness and peacefulness, accuracy rates remained invariant across age groups (ps > .9), with the exception that for peacefulness middle-aged participants achieved better accuracy than younger ones (p < .01). We repeated this analysis with participants’ raw ratings on the intended emotion scales (raw ratings are displayed in Appendix 1). The pattern of age effects was globally replicated: older participants provided lower ratings than younger ones for sad and scary excerpts (ps < .01), and ratings for happy and peaceful ones did not change (ps > 1); for scary excerpts, differences were significant already in middle-age (p < .0001), but for sad ones they were not [p > .1; main effect of age, F(2,110) = 10.91, p < .0001, !p2 = .17; interaction Age x Emotion, F(6,330) = 15.28, p < .0001, !p2 = .22]. The age differences were also replicated with accuracy rates corrected for possible response bias (unbiased accuracy rates are displayed in Appendix 2)3: older participants had lower accuracy than younger !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 3 In forced-choice tasks, raw proportions of correct responses are not always the best measure of accuracy because they are vulnerable to response bias. Participants might be disproportionately likely to respond using some response categories rather than others. As a consequence, differences between groups might reflect this bias instead of true differences in the intended ability (for a review, see Isaacowitz et al., 2007). Our accuracy data are less prone to bias because they are derived from ratings on intensity scales and participants did not perform any choice between categories. However, to confirm that the pattern of age-related changes is not a by-product of possible response trends, we reanalyzed accuracy using unbiased hit rates, Hu (Wagner, 1993). Hu is a measure of accuracy which is insensitive to bias in responses. It represents the joint probability that a stimulus category is correctly recognized when it has been presented, and that a response is correct given that it has been used. It was calculated for each emotion and participant using the formula Hu = A2 / (B X C), where A corresponds to the number of stimuli correctly identified, B to the number of presented stimuli (14 per category), and C to the total number of responses provided for that category (i.e., including misclassifications). Values were arcsine transformed before being analyzed.
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 122 Because categorization accuracy for happiness was very high, 93% correct on average, it might be argued that the invariance observed across age groups reflects a ceiling effect. Notwithstanding, the observed rates still lie below the maximum score, and thus they would allow for variability in performance. Additionally, raw ratings for happiness were also highly stable across age groups (younger, 7.8; middle-aged, 7.7; older, 7.6; maximum 9), and this suggests true invariance, not lack of stimulus sensitivity to differences across listeners. Can decline in general cognitive functioning underlie the age-related differences in emotion recognition? This would predict changes for the emotions which are hardest to recognize, and we did not find that: there were changes for a relatively easy emotion, sadness, and stability for a relatively difficult one, peacefulness. Furthermore, older participants had a near ceiling performance on the MMSE, which indicates that they did not have prominent global impairments. Changes for emotion recognition might thus be primary in origin. Still, we cannot rule out the possible role of subtle cognitive changes, namely in the case of middle-aged participants, who were not inspected for domaingeneral abilities. In Study 2 we examine directly whether there are associations between emotion recognition and general cognitive performance. We also compare expert musicians with untrained listeners to follow-up the effect of musical training unveiled in Study 1. In Study 2 we include two age groups, younger and middle-aged adults; we selected middle-aged instead of older participants because developmental changes in the middle-years are frequently neglected in the literature. Additionally, we observed relative stability in the pattern of changes from middle-age onward (for similar results in other modalities, see Isaacowitz et al., 2007).
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 123 6.3. STUDY 2 6.3.1. Method PARTICIPANTS | Eighty participants took part in this study (none of them participated in Study 1). They were distributed into four groups according to musical expertise (musicians and controls) and age (younger and middle-aged), 20 per group (10 female). Table 6 presents their demographic and background characteristics. Younger participants were 23 years old on average (SD = 3.2), and middle-aged ones were 47.7 years old (SD = 4.4). Musicians were instrumentalists who played piano (n = 18), flute (n = 5), violin (n = 5), guitar (n = 3), double bass (n = 2), clarinet (n = 1), drums (n = 1), cello (n = 1), oboe (n = 1), viola (n = 1), trombone (n = 1) or accordion (n = 1); in addition to instrumental training, two of them had vocal training in classical singing. They had at least 8 years of formal musical training started in childhood, and practiced regularly their instruments at the moment of testing. Younger ones were advanced music students or professional musicians, and older ones were music teachers and/or orchestral performers. They were recruited from local music schools and orchestras, including Conservatório de Música do Porto, Escola Superior de Música e Artes do Espectáculo do Instituto Politécnico do Porto and Orquestra Sinfónica do Porto Casa da Música. Younger and older musicians were matched for the number of years of training, age of commencement and instrumental practice per week (Fs < 1; see Table 6). Controls have never had formal music lessons nor played any instruments. They were recruited from several local communities. The four groups were matched for education level, as assessed by the number of years attending school (Fs < 1). All participants had normal or corrected-to-normal vision and reported no head trauma, substance abuse nor major psychiatric or neurological illnesses. The groups rated themselves similarly for hearing acuity in a scale from 1, very good, to 6, very bad (M = 2.1; SD = .9; ps > .1). Participants were asked to self-rate their interest in music in a scale from 1, very high, to 6, very low; all reported that music is important to them, but musicians revealed stronger interest (1.03) than controls [2.18, F(1,76) = 51.94, p < .001, !p2 = .4]. Participants were financially compensated for their participation.
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 124 Table 6. Demographic and background characteristics of controls and musicians for each age group (Study 2). SDs are presented in parentheses. Controls Musicians Characteristics Younger Middle-aged Younger Middle-aged Age (years) 22.7 (2.8) 47 (4) 23.4 (3.6) 48.4 (4.8) Education (years) 15.7 (1.5) 17.1 (4) 15.4 (1.8) 16.5 (3.6) Musical training (years) - - 11.3 (3.1) 12.6 (3.2) Age of commencement (years) - - 9.2 (2.5) 8.4 (4.5) Instrumental practice (hours/week) - - 12.7 (11.6) 12.3 (11) Montreal Cognitive Assessment (/30) 28.2 (1.3) 26.8 (1.9) 28.1 (1.6) 27.8 (1.2) Raven’s APM (problems solved, /36) 20.5 (4.7) 13.9 (5.1) 19.7 (4.9) 16.8 (4.5) Stroop test1 Baseline (words/s) 2.09 (0.3) 2.14 (0.3) 2.37 (0.3) 2.39 (0.3) Conflict (colors/s) 1.04 (0.2) 0.87 (0.1) 1.08 (0.2) 1.04 (0.2) Ten-Item Personality Inventory Extraversion (/7) 4.8 (1.5) 4.6 (1.7) 4.8 (1) 4.9 (1.9) Agreeableness (/7) 5.3 (0.8) 5.5 (1) 5.2 (0.9) 5.5 (0.9) Conscientiousness (/7) 5 (1.3) 5.5 (1.3) 4.2 (1.4) 5.1 (1.5) Emotional stability (/7) 4.2 (1.4) 4.4 (2) 3.6 (1.2) 4.5 (1.5) Openness to experience (/7) 5.6 (0.8) 5.6 (1) 5.6 (1) 6.1 (0.9) Autism Spectrum Quotient2 (/50) 17.5 (5.4) 18 (4.8) 17.5 (5.7) 17.5 (5.5) 1 We computed the number of items named per s: number of correct responses/time taken to perform the task 2 Higher values indicate higher autistic-like socio-communicative traits (cut-off for clinically significant levels of autistic traits, 32) COGNITIVE AND PERSONALITY ASSESSMENT | The results of cognitive tests and personality scales are summarized in Table 6. Participants were screened for global cognitive dysfunction with the Montreal Cognitive Assessment (MOCA; www.mocatest.org; Portuguese version, Simões, Firmino, Vilar, & Martins, 2007), a brief instrument which inspects attention and concentration, executive functions, memory, language, visuoconstructional skills, conceptual thinking, calculations and orientation. Nonverbal general intelligence was assessed with a 20-minute timed version of the Raven’s Advanced Progressive Matrices; this version correlates strongly with the untimed one (Hamel & Schmittmann, 2006). To examine executive control, we used a Stroop task (Trenerry, Crosson, Deboe, & Leber, 1995). In the version used here participants performed speeded naming in two conditions: baseline, consisting of reading words denoting colour names (blue, azul; pink, rosa; grey, cinza; green, verde);
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 125 and an incongruous condition, consisting of naming the ink of written words that denoted an incongruent color name (e.g., the word "blue" printed in green ink; Portuguese version, Castro, Martins, & Cunha, 2003). Musicians and controls did not differ in the MOCA and Raven’s matrices (see Table 6; ps > .19). This indicates that they were similar for the general abilities assessed by these tests. In the Stroop task, musicians were faster than controls in both baseline [F(1,76) = 16.2, p < .001, !p2 = .18] and incongruous conditions [F(1,76) = 6.7, p < .05, !p2 = .08]. Their advantage in the incongruous condition only approached significance when baseline differences were controlled for [ANCOVA, F(1,75) = 3.7, p = .06, !p2 = .05]. This suggests that musicians’ enhanced executive control was partially due to an advantage in processing speed; effects of musical expertise on executive control have been found by Bialystok and DePape (2009). Middle-aged participants scored lower than younger ones in the MOCA [F(1,76) = 5.47, p < .05, !p2 = .07], Raven’s matrices [F(1,76) = 19.5, p < .01, !p2 = .2] and in the conflict condition of Stroop [F(1,76) = 7.7, p < .01, !p2 = .09]. These ageing effects were observed similarly for controls and for musicians (interactions Age x Expertise ns, ps > .05). The only measure that did not show age-related decline was the baseline condition of the Stroop test (F < 1; interaction Age x Expertise ns, F < 1). Because personality characteristics can influence emotion processing (e.g., Hamann & Canli, 2004; Matsumoto et al., 2000; Mill et al., 2009), participants completed the TenItem Personality Inventory (TIPI), a brief questionnaire that inspects the Big-5 personality domains: extraversion, agreeableness, conscientiousness, emotional stability and openness to experience (Gosling, Rentfrow, & Swann Jr., 2003; Portuguese version, Lima & Castro, 2009). They also completed The Autism-Spectrum Quotient (AQ), a questionnaire that assesses socio-communicative traits associated with the autistic spectrum in neurotypical adults (Baron-Cohen, Wheelwright, Skinner, Martin, & Clubley, 2001; Portuguese version, Castro & Lima, 2009). Autistic traits correlate with structural and functional differences in brain regions involved in emotion processing, even in neurotypicals (Hagen et al., 2011; Martino et al., 2009). Both age and expertise groups had similar personality and socio-communicative characteristics, as assessed by these measures (Fs < 1; see Table 6).
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 126 MUSIC STIMULI | As in Study 1, stimuli were taken from the set validated by Vieillard and colleagues (2008). Here were used 10 excerpts per emotion category, 40 in total. They were pseudo-randomized and divided into two blocks of 20 trials each. PROCEDURE | Participants were tested in an individual session lasting about 90 minutes. The instructions and structure of the experimental task were the same as in Study 1. Besides the experimental task and the cognitive and personality tests, participants also performed a forced-choice identification of seven emotional tones in speech prosody (these results are presented in Chapter VIII). 6.3.2. Results Comparisons across groups were based on derived accuracy rates, as in Study 1. Table 7 presents the percentage of accurate categorizations for each emotion, and the distribution of inaccurate and ambivalent responses, separately for age and expertise groups. Agreement rates ranged between 42% (peacefulness in younger controls) and 96% (happiness in younger controls). EFFECTS OF AGE AND MUSICAL EXPERTISE ON ACCURACY | Correct identifications per emotion were arcsine transformed and submitted to an ANOVA, with emotion as repeated-measures factor (happy, peaceful, sad and scary), and age (young and middleage) and expertise (controls and musicians) as between-subject factors. Happiness reached the highest accuracy, 94%, and peacefulness the lowest, 50%; sadness and fear elicited intermediate and similar accuracies, 67% [ps < .001; main effect of emotion, F(3,228) = 63.62, p < .001, !p2 = .46].
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 127 Table 7. Distribution of responses (%) for each intended emotion, as function of age and expertise (Study 2). Diagonal cells in bold indicate accurate categorizations. Standard errors are presented in parentheses. Distribution of responses (%) Controls Musicians Age group/ excerpt type Happy Peaceful Sad Scary Ambivalent Happy Peaceful Sad Scary Ambivalent Younger Happy 96 (1.3) 1 1 0 2 95 (2.8) 2 0 0 4 Peaceful 14 42 (7.3) 27 0 18 22 46 (7.5) 11 0 22 Sad 0 6 83 (3) 3 9 0 12 70 (6.6) 1 18 Scary 6 0 8 77 (4.4) 10 2 2 15 70 (7) 11 Middle-aged Happy 91 (3.8) 0 3 2 4 95 (2.3) 1 0 1 5 Peaceful 13 51 (5.4) 11 3 23 12 64 (5.7) 12 0 13 Sad 6 20 52 (7.1) 1 22 1 18 64 (5.6) 1 17 Scary 17 2 14 49 (5.8) 18 8 0 11 72 (5.5) 10
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 128 Concerning the impact of ageing, the pattern obtained in Study 1 was replicated: planned comparisons confirmed that middle-aged participants were less accurate than younger ones at categorizing the negative emotions, sadness (76% for younger vs. 58% for middle-aged, p < .01) and fear (73% for younger vs. 60% for middle-aged, p < .05). By contrast, the recognition of happiness remained stable (p > .4), and accuracy rates for peacefulness increased [44% for younger vs. 57% for middle-aged, p < .05; main effect of age ns, p > .1; interaction Age x Emotion, F(3,228) = 7.3, p < .001, !p2 = .09]. Note that the age-related changes for sadness and fear were observed in controls (ps < .01), but not in musicians [ps > .9; interaction Age x Expertise, F(1,76) = 5.79, p < .05, !p2 = .07]: younger and middle-aged musicians had similar accuracy rates. Our prediction that musical expertise would be associated with enhanced accuracy was confirmed, though for older musicians only: middle-aged musicians had higher accuracy (73%) than age-matched controls (61%, p < .05), but younger musicians (70%) and controls (73%) performed similarly [p > .3; main effect of expertise ns, p > .1; see above the interaction Age x Expertise]. This effect was independent of emotion category (interactions Expertise x Emotion, and Age x Expertise x Emotion ns, ps > .3)4. The pattern of age and expertise effects was globally replicated in ANOVAS on raw ratings (see raw ratings in Appendix 4) and on unbiased accuracy rates (see unbiased rate in Appendix 5), except that the age-related increase in responsiveness to peacefulness was not significant [raw ratings: middle-aged participants provided lower ratings than younger ones for sadness and fear, ps < .05, but for happiness and peacefulness ratings remained invariant, ps > .4, interaction Age x Emotion, F(3,228) = 7.55, p < .001, !p2 = .09; middle-aged musicians provided higher ratings on the intended emotions than controls, p < .05, interaction Age x Expertise, F(1,76) = 7.01, p < .01, !p2 = .08; unbiased rates: middle-aged participants were less accurate than younger ones for sadness and fear, ps < .05, but similar for happiness and peacefulness, ps > .05, interaction Age x Emotion, F(3,228) = 6.65, p < .001, !p2 = .08; middle-aged musicians performed better than age-matched controls, p < .05, interaction Age x Expertise, !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 4 We explored whether musicians’ performance varied according to the played instrument. Because they were widely distributed across many different instruments (see Methods), we first categorized them by type of instruments [keyboard, n = 18; strings, n = 12; woodwinds, n = 7; others, n = 3 (drums, accordion, and trombone)] and then computed an ANOVA with accuracy rates as dependent measure. Performance was similar across instrument types (p > .1).
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 129 F(1,76) = 4.35, p < .05, !p2 = .05]. Gender did not impact results (main effect of gender ns, F < 1; interactions with age, expertise and emotion ns, ps > .05). We also computed Pearson correlations between years of age and categorization accuracy for each emotion. These analyses were performed separately for controls and musicians because age affected differently both groups. For controls, recognition accuracy decreased with age for sadness (r = -.56, p < .0001) and fear (r = -.52, p < .001), but not for happiness (r = -.2, p > .1) and peacefulness (r = .09, p > .5). For musicians, recognition accuracy was not linked with age (ps > .05). The correlation between the number of years of musical training and recognition rates obtained in Study 1 was replicated in this sample of expert musicians. The longer they were trained, the more accurate they were at categorizing the excerpts (r = .33, p < .05), as illustrated in Figure 2. The correlation was marginally significant when it was conducted on all participants, musicians and controls, r = .21, p = .07. Figure 2. Scatterplot of emotion recognition accuracy (%) of musicians by years of training (Study 2).
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 130 DISTRIBUTION OF INACCURATE RESPONSES | An ANOVA was carried out for each emotion, with the non-intended response categories as repeated-measures factor (three non-intended emotions and ambivalent responses), and age and expertise as betweensubject factors. The pattern of responses broadly mirrored the one reported in Study 1. For happy music, misclassifications were rare and included chiefly ambivalent responses, 4% [0.75% for each of the other categories, ps < .001; main effect of category, F(3,228) = 8.74, p < .001, !p2 = .1]. For peaceful music, non-intended responses were similarly frequent for ambivalence (19%), sadness (15%) and happiness [15%, p > .3; 1% for fear, ps < .001; main effect of category, F(3,228) = 32.3, p < .001, !p2 = .3]. For sad music, non-intended responses were common for ambivalence (16%) and peacefulness [14%, p > .3; 2% for happiness and 1% for fear, ps < .01; main effect of category, F(3,228) = 54.98, p < .001, !p2 = .45]. Scary music was misclassified mainly as sad and ambivalent (both 12%), and less as happy [8%, p < .05; 1% for peacefulness, ps < .05; main effect of category, F(3,228) = 19.78, p < .001, !p2 = .2]. The pattern of non-intended responses for happy and scary music was similar across age and expertise groups (interactions Age x Category, Expertise x Category, and Age x Expertise x Category ns, ps > .05). Regarding peaceful music, younger participants gave more sad responses than middle-aged ones, and controls (younger and older) gave more ambivalent than happy responses; these differences were not observed in musicians [interaction Age x Expertise x Category, F(3,228) = 3.78, p < .05, !p2 = .05; interactions Emotion x Age and Emotion x Expertise ns, ps > .2]. As for sad music, younger participants gave less responses for peacefulness than middle-aged ones [interaction Age x Category, F(3,228) = 5.16, p < .001, !p2 = .06; the other interactions were ns, ps > .05]. CORRELATIONS BETWEEN GENERAL COGNITIVE ABILITIES, PERSONALITY AND EMOTION RECOGNITION | The four groups were matched for background variables such as education and sex. However, they differed for cognitive abilities (see Methods): middle-aged participants were worse than younger ones for global cognitive functioning (MOCA), nonverbal intelligence (Raven’s matrices) and executive control (Stroop test, conflict condition); and musicians had faster processing speed than controls and a trend for superior executive control (Stroop test, both conditions). Do these differences
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 131 account for performance in emotion recognition? We computed correlation analyses between global categorization accuracy (score averaged across the four emotions) and scores on the cognitive measures. Correlation coefficients and p values are displayed in Table 8. No significant associations were found. This suggests that domain-general cognitive abilities did not influence emotion recognition in music5. Another potential confound was musicians’ stronger interest in music, but this variable was also not related with recognition accuracy. The four groups had similar personality characteristics, as assessed by the TIPI, and socio-communicative traits, as assessed by the AQ. We explored possible associations between these measures and emotion recognition, but all analyses yielded nonsignificant results (see Table 8). Table 8. Correlations between cognitive and personality measures and global accuracy in emotion recognition in music. Emotion recognition accuracy Cognitive/personality measure r p Montreal Cognitive Assessment .1 .37 Raven’s Advanced Progressive Matrices .19 .1 Stroop test Baseline condition .15 .18 Conflict condition .07 .52 Ten-Item Personality Inventory Extraversion .03 .8 Agreeableness -.13 .27 Conscientiousness -.05 .67 Emotional stability -.17 .12 Openness to experience .06 .62 Autism Spectrum Quotient .1 .38 Interest in music -.16 .17 !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 5 We further confirmed that age and expertise effects in emotion recognition were not explained by general cognitive abilities by computing ANCOVAs on categorization accuracy with MOCA, Raven’s Matrices and Stroop test as covariates.
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 138 motivation to heighten attention towards positive information, and decrease attention towards negative information – positivity effect (Carstensen & Mikels, 2005; Charles & Carstensen, 2010; Mather & Carstensen, 2005; Samanez-Larkin & Carstensen, 2011). Consistently, we reported a pattern of decreased responsiveness to negative musical emotions, and stability for positive ones. This pattern fits coherently with studies showing that, with advancing age, brain responses decreases selectively for negative emotional stimuli (Gunning-Dixon et al., 2003; Kisley et al., 2007), and there is increased engagement of prefrontal systems which may implement regulatory effects over emotion processing (St. Jacques et al., 2009; St. Jacques et al., 2008; Williams et al., 2006). In a different vein, it has been hypothesized that age-related degradation in brain regions such as the amygdala and cingulate cortex contribute to explain the decline observed in the recognition of facial expressions of sadness and fear (Calder et al., 2003; Ruffman et al., 2008). Indeed, the volume of the amygdala was shown to decrease with advancing age (Curiati et al., 2009; Mu, Xie, Wen, Weng, & Shuyun, 1999; Walhovd et al., 2005). Neuropsychological research using the same stimuli and task as the present studies indicate that this structure is critical for the perception of fear and sadness in music (Gosselin et al., 2005, 2007). Thus, it is plausible that age-related changes for these emotions are linked with decline in the amygdala (see also, Cacioppo et al., 2011). On the other hand, the stability observed for the positive emotions might be related to the relative preservation of basal ganglia structures (e.g., Williams et al., 2006). Neuroimaging evidence showed that pleasant music recruits the striatum (e.g., Koelsch et al., 2006; Mitterschiffthaler et al., 2007). Note that some studies observed age-related shrinkage in the striatum, however (Raz et al., 2003). The present experiments are not ideally suited to specify the relative contribution of top-down regulatory mechanisms and neuropsychological deterioration for ageing effects in emotion recognition in music. Future studies combining behavioural measures with functional and structural brain imaging techniques will be critical to approach this question. These studies will also throw light on the differences that we observed in how younger and older participants relied on music structural cues to respond. For happy and peaceful responses, the utilization of the cues remained invariant with advancing age, but for sadness and fear older participants were less consistent than younger ones and weighted some cues differently. This might reflect deterioration of the neurocognitive
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 139 mechanisms that support the processing of these emotions, but can also be accounted for by top-down regulatory and motivational processes: older participants might implement an active and controlled disengagement of music cues when they signal negative emotional meanings. MUSICAL EXPERTISE AND EMOTION RECOGNITION IN MUSIC | Two findings indicate that musical training bolsters emotion recognition in music. First, in both studies, the length of musicians’ training correlated significantly with enhanced recognition accuracy. This suggests that learning music operates gradual fine-tuning on mechanisms underlying musical emotions. Second, in Study 2, middle-aged musicians were more accurate than age-matched controls at categorizing the four emotions. As observed for age, we showed for the first time that the effects of expertise cannot be explained by differences in socio-educational background, general cognitive abilities and personality characteristics. The correlation between training and accuracy extends previous results by Livingstone and colleagues (2010), who reported a similar correlation for musical emotions portrayed by both structural and expressive cues, and by Bhatara and colleagues (2011), who focused on the expressive features timing and amplitude. Here we established that the effect also occurs for emotions embodied solely on music structure. Further support for the notion that musical training influences emotion processing in music comes from a study by Dellacherie, Roy, Hugueville, Peretz, and Samson (2011). These authors found that musical experience affects both behavioural and psychophysiological responses to dissonance in music. Musical training was associated with more unpleasant feelings and stronger physiological responses to dissonance, as assessed by skin conductance and electromyographic signals. Resnicow and colleagues (2004) failed to find effects of training, possibly because their design had no power enough to attain statistical significance – they included only three stimuli per emotion; furthermore, it was not reported how many trained participants were included. It might be the case that the effect of expertise is small, a fact that would explain why it is not detected in some conditions (e.g., for younger musicians vs. controls in Study 2). Indeed, Bigand and colleagues (2005) observed that emotional responses to music in an
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 140 implicit task are highly consistent and similar across trained and untrained participants. Training only increased the consistency of responses across testing sessions. Moreover, in a review of studies on the effects of musical expertise, Bigand and Poulin-Charronnat (2006) concluded that the effects of expertise on music processing are usually small, particularly in light of the huge differences in training that typically exist between groups. According to the authors, this is evidence that musical abilities, namely emotion processing, are acquired through exposure, not requiring explicit training. Our data is also evidence that training is not a prerequisite to be adept at perceiving musical emotions: all participants, including musically untrained ones, identified the four intended emotions with high agreement rates. This consistency was obtained with stimuli that were short, unfamiliar, all played in the same genre and timbre, and expressing emotions exclusively through music structure cues. Additionally, participants were able to effectively use these cues to make emotion differentiations, as indicated by the multiple regression analyses. By claiming that learning music impacts emotion recognition, we intend to highlight the plasticity of the system, and not to argue that it needs training to operate robustly. How can musical training enhance the categorization of emotional qualities in music? Musicians, because of their extensive implicit and explicit practice with musical emotions, might develop more fine-grained, sharply defined and readily accessible categories to respond to emotion cues in music. Finding that they were more consistent than controls in using music structure cues to respond is in agreement with this hypothesis. Barrett and colleagues (Barrett, 2006b, 2009; Lindquist & Barrett, 2010) have suggested that individual differences in emotional responses and experience might reflect differences in the granularity of emotion concepts, which could be trained. According to this theoretical proposal, individuals higher in emotional granularity are more precise, specific and differentiated at categorizing affective states, while those with low granularity experience and describe emotions in a more global and undifferentiated fashion. Higher emotional granularity could be achieved through training and practice, just like wine and X-ray experts learn to perceive subtle differences that novices are unaware of. It is plausible that musical training contributes to increase the granularity of emotion concepts for music.
Chapter VI. Ageing and Musical Expertise Modulate Emotion Recognition in Music ! 141 6.5. CONCLUSION In two studies, we sought to determine the effects of ageing and musical expertise on emotion recognition in music. To our knowledge, this is the first examination of these effects that covers the full range of the adult life span, and that focus on emotions expressed through musical structure. We found that ageing modulates the recognition of musical emotions, with developmental shifts being observed early on in middle-age. Specifically, advancing age was associated with a gradual decrease in responsiveness to the negative emotions of sadness and fear in music, while the positive ones, happiness and peacefulness, remained invariant. This is evidence that emotion processing in music can follow the same developmental trends as stimuli important for biological survival and social functioning, such as facial expressions and speech prosody. We also showed that musical expertise is associated with more accurate categorization of musical emotions, thus contributing to the literature on the relation between musical experience and music perception. Both ageing and expertise effects were independent of domaingeneral cognitive or personality differences, and this suggests that they have a primary origin. Mechanisms supporting emotion recognition in music are robust, but also dynamic and variable: they respond plastically to the influence of experience.
! 142
! 143 CHAPTER VII Development and Validation of a Set of Portuguese Stimuli for Research on Emotional Prosody
! 144
Chapter VII. Set of Portuguese Stimuli for Research on Emotional Prosody ! 145 7.1. INTRODUCTION Well-devised and validated stimuli are important for research on emotional speech prosody (e.g., Burkhardt et al., 2005; Pell, 2002; Ross et al., 1997). Because prosodic information is overlaid on the speech signal, sets of stimuli need to be developed for different languages. The availability of materials in different languages is also crucial to shed light on the role of language-specific and universal factors in the processing of prosody (e.g., Pell, Monetta, et al., 2009; Thompson & Balkwill, 2006). In this study we develop and validate a database of prosodic stimuli in Portuguese6. In the literature on emotional speech prosody, typically researchers ask trained actors or untrained speakers to read some verbal materials aloud while portraying different emotional states, such as anger, disgust, fear, happiness, sadness and surprise (e.g., Adolphs et al., 2002; Banse & Scherer, 1996; Dara et al., 2008; Juslin & Laukka, 2001; Scherer et al., 1991). Distinct types of verbal materials have been used, including complete sentences (e.g., Adolphs & Tranel, 1999; de Gelder & Vroomen, 2000; Kotz et al., 2003; Mitchell, 2007), single words (e.g., Wiethoff et al., 2008) or monosyllabic utterances (e.g., Mitchell & Ross, 2008). It has often been argued that these posed portrayals may result in intense, prototypical expressions, which reflect conventionalized stereotypical norms, more than the natural psychophysiological changes on voice that occur in authentic emotion episodes. As stated by Scherer (2003), though, the so-called natural public expressions are also acted to some degree, since we self-regulate vocal expressions during everyday life, and sociocultural norms impose production constrains to a significant extent (pull effects). It is also well documented !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 6 Published in Castro, S. L., & Lima, C. F. (2010). Recognizing emotions in spoken language: A validated set of Portuguese sentences and pseudo-sentences for research on emotional prosody. Behavior Research Methods, 42, 74-81. doi:10.3758/BRM.42.1.74
Chapter VII. Set of Portuguese Stimuli for Research on Emotional Prosody ! 146 that listeners can identify reliably emotions in acted stimuli, indicating that they are at least partially modelled on the basis of natural expressions. Furthermore, although scarce, available evidence suggests that the acoustic patterns associated with specific emotions are qualitatively similar in portrayals and in natural expressions (Juslin & Laukka, 2001; Laukka et al., 2008). Another approach to devise prosodic stimuli involves inducing emotional states in the speakers (e.g., through films or pictures), and record expressions under these states. While this approach would be better suited to obtain spontaneous samples, it is difficult to evoke strong and differentiated emotional reactions in laboratory settings, so that the resulting voice stimuli are usually low intense and with relatively undifferentiated affect (Juslin & Laukka, 2001). Here we use acted portrayals and endorse the assumption that they are suitable for research on emotion communication through prosody (Bänziger & Scherer, 2007). A fundamental issue that has to be dealt with when creating prosodic materials is the interplay of prosody – the variable of interest – with the lexico-semantic content of the utterance. One strategy consists of removing semantic information through acoustic filtering procedures that suppress the segmental content. However, emotions are more difficult to recognize in these acoustically manipulated stimuli than in unaltered ones (Kotz et al., 2003). Alternatively, the influence of semantics can be controlled for by using standardized verbal materials, a procedure often called standard content paradigm (e.g., Juslin & Laukka, 2001, 2003). That is, the same verbal material is used to portray all emotions, and because content is similar across emotions, it is assumed that the effects observed in the listener’s judgments reflect solely the impact of prosodic cues (Scherer, 2003). This approach is used in most studies. The characteristics of the standardized content itself deserve also to be considered because they may modulate the neurocognitive processes that are set into action. For instance, neuropsychological and fMRI studies using the Aprosodia Battery show that stimuli with different degrees of verbal complexity recruit different neurocognitive processes (Mitchell & Ross, 2008; Ross & Monnot, 2008). The Aprosodia Battery (Ross et al., 1997) employs six intonations (angry, disinterested, happy, sad, surprised and neutral) in three types of materials: short sentences with emotionally neutral semantic content (“I am going to the other movies”), a type of stimuli used frequently in the literature (e.g., Adolphs, Tranel,
Chapter VII. Set of Portuguese Stimuli for Research on Emotional Prosody ! 147 & Damasio, 2001; Mitchell, 2007; Wildgruber et al., 2005); monosyllabic utterances (repeated ba); and asyllabic stimuli (prolonged aaaaahhhhh). It was observed that healthy listeners recognize emotions similarly well across conditions, about 70% correct on average, confirming that the processing of prosody can proceed independently of semantics. However, the relative contribution of the left and right hemispheres depended on the type of stimulus: lateral temporal lobes in both hemispheres were engaged by emotional prosody, but the contribution of the left hemisphere got smaller as linguistic complexity decreased (Mitchell & Ross, 2008; Ross & Monnot, 2008). The authors concluded that emotional prosody is right-lateralized to a significant extent, and that the contribution of the left hemisphere is related to the processing of verbal information. Accordingly, the dynamic dual-pathway model of language functions postulates a dissociation between semantic/syntactic processes versus prosodic ones: a left-lateralized temporo-frontal network supports syntactic and semantic processes, being the right hemisphere dominant for prosodic processes, depending on the linguistic demands of the stimuli or task – the higher these are, the greater is the co-involvement of the left hemisphere (Friederici & Alter, 2004). Hence, stimuli with different degrees of linguistic information can be useful to explore the processes that underlie emotional prosody. One can for example resort to meaningless speech and pseudo-sentences to create stimuli with reduced semantic content (e.g., Banse & Scherer, 1996; Pell, 2002; Péron et al., 2010). Pseudo-sentences are utterances that include pseudo-words. Their lexico-semantic content is substantially reduced, because pseudo-words have no meaning, yet they have the advantage of affording a good language-like quality because both phonetic-segmental and suprasegmental features of normal speech are present. Emotions can be recognized in pseudo-sentences with high accuracy (e.g., 78% correct in Pell, 2002), and this type of stimuli has been successfully used in neuropsychological (e.g., Dara et al., 2008; Péron et al., 2010), electrophysiological (e.g., Paulmann & Kotz, 2008b) and neuroimaging studies (Bach et al., 2008). Emotional stimuli like facial expressions (e.g., Ekman & Friesen, 1978), pictures (e.g., Lang, Bradley, & Cuthbert, 2008), music (Vieillard et al., 2008) and nonverbal vocal expressions (Belin et al., 2008; Sauter, Eisner, Calder et al., 2010; Sauter, Eisner, Ekman et al., 2010; Sauter & Scott, 2007) can be used independently of the linguistic
References ! 250 Juslin, P. N., Liljeström, S., Västfjäll, D., & Lundqvist, L. (2010). How does music evoke emotions? Exploring the underlying mechanisms. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of Music and Emotion: Theory, Research, Applications (pp. 605-642). New York: Oxford University Press. Juslin, P. N., & Sloboda, J. A. (2010). Introduction: Aims, organization, and terminology. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of music and emotion: Theory, research, applications (pp. 3-12). New York: Oxford University Press. Juslin, P. N., & Timmers, R. (2010). Expression and communication of emotion in music performance. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of Music and Emotion: Theory, Research, Applications (pp. 453-489). New York: Oxford University Press. Juslin, P. N., & Västfjäll, D. (2008). Emotional responses to music: The need to consider underlying mechanisms. Behavioral and Brain Sciences, 31, 559-621. doi:10.1017/S0140525X08005529 Kan, Y., Kawamura, M., Hasegawa, Y., Mochizuki, S., & Nakamura, K. (2002). Recognition of emotion from facial, prosodic and written verbal stimuli in Parkinson's disease. Cortex, 38, 623-630. doi:10.1016/S0010-9452(08)70026-1 Kehagia, A. A., Barker, R. A., & Robbins, T. W. (2010). Neuropsychological and clinical heterogeneity of cognitive impairment and dementia in patients with Parkinson's disease. Lancet Neurology, 9, 1200-1213. doi:10.1016/S14744422(10)70212-X Keltner, D., & Ekman, P. (2002). Introduction: Expression of emotion. In R. J. Davidson, K. R. Scherer & H. H. Goldsmith (Eds.), Handbook of Affective Sciences. Oxford: Oxford University Press. Khalfa, S., Della Bella, S., Roy, M., Peretz, I., & Lupien, S. J. (2003). Effects of relaxing music on salivary cortisol level after psychological stress. Annals of the New York Academy of Sciences, 999, 374-376. doi:10.1196/annals.1284.045
References ! 251 Khalfa, S., Guye, M., Peretz, I., Chapon, F., Girard, N., Chauvel, P., & LiégeoisChauvel, C. (2008). Evidence of lateralized anteromedial temporal strucures involvement in musical emotion processing. Neuropsychologia, 46, 2485-2493. doi:10.1016/j.neuropsychologia.2008.04.009 Kisley, M. A., Wood, S., & Burrows, C. L. (2007). Looking at the sunny side of life: Age-related change in an event-related potential measure of the negativity bias. Psychological Science, 18, 838-843. doi:10.1111/j.1467-9280.2007.01988.x Kivy, P. (1991). Music alone: Philisophical reflections on the purely musical experience. Ithaca New York: Cornell University Press. Kivy, P. (2002). Introduction to a philosophy of music. New York: Oxford University Press. Knösche, T. R., Neuhaus, C., Haueisen, J., Alter, K., Maess, B., Witte, O. W., & Friederici, A. D. (2005). Perception of phrase structure in music. Human Brain Mapping, 24, 259-273. doi:10.1002/hbm.20088 Koelsch, S. (2010). Towards a neural basis of music-evoked emotions. Trends in Cognitive Sciences, 14, 131-137. doi:10.1016/j.tics.2010.01.002 Koelsch, S. (2011). Toward a neural basis of music perception: A review and updated model. Frontiers in Psychology, 2, 1-20. doi:10.3389/fpsyg.2011.00110 Koelsch, S., Fritz, T., Cramon, D. Y. v., Muller, K., & Friederici, A. D. (2006). Investigating emotion with music: An fMRI study. Human Brain Mapping, 27, 239-250. doi:10.1002/hbm.20180 Koelsch, S., Gunter, T., Wittfoth, M., & Sammler, D. (2005). Interaction between syntax processing in language and in music: An ERP study. Journal of Cognitive Neuroscience, 17, 1565-1577. doi:10.1162/089892905774597290 Koelsch, S., Kasper, E., Sammler, D., Schulze, K., Gunter, T., & Friederici, A. D. (2004). Music, language and meaning: Brain signatures of semantic processing. Nature Neuroscience, 7, 302-307. Doi:10.1038/nn1197
References ! 252 Koelsch, S., Offermanns, K., & Franzke, P. (2010). Music in the treatment of affective disorders: An exploratory investigation of a new method for music-therapeutic research. Music Perception, 27, 307-316. doi:10.1525/MP.2010.27.4.307 Koelsch, S., & Siebel, W. A. (2005). Towards a neural basis of music perception. Trends in Cognitive Sciences, 9, 578-584. doi:10.1016/j.tics.2005.10.001 Kolinsky, R., Cuvelier, H., Goetry, V., Peretz, I., & Morais, J. (2009). Music training facilitates lexical stress processing. Music Perception, 26, 235-246. doi:10.1525/MP.2009.26.3.235 Konecni, V. J. (2008). Does music induce emotion? A theoretical and methodological analysis. Psychology of Aesthetics, Creativity, and the Arts, 2, 115-129. doi:10.1037/1931-3896.2.2.115 Korpilahti, P., Jansson-Verkasalo, E., Mattila, M.-L., Kuusikko, S., Suominen, K., Rytky, S., … Moilanen, I. (2007). Processing of affective speech prosody is impaired in Asperger Syndrome. Journal of Autism and Developmental Disorders, 37, 1539-1549. doi:10.1007/s10803-006-0271-2 Kotz, S. A., Meyer, M., Alter, K., Besson, M., Cramon, D. Y. v., & Friederici, A. D. (2003). On the lateralization of emotional prosody: An event-related functional MR investigation. Brain and Language, 86, 366-376. doi:10.1016/S0093934X(02)00532-1 Kotz, S. A., & Schwartze, M. (2010). Cortical speech processing unplugged: A timely subcortico-cortical framework. Trends in Cognitive Sciences, 14, 392-399. doi:10.1016/j.tics.2010.06.005 Kövari, E., Gold, G., Herrmann, F., Canuto, A., Holf, P., Bouras, C., & Giannakopoulos, P. (2003). Lewy body densities in the entorhinal and anterior cingulate cortex predict cognitive deficits in Parkinson's disease. Acta Neuropathologica, 106, 83-88. doi:10.1007/s00401-003-0705-2 Kramer, A. F., Bherer, L., Colcombe, S. J., Dong, W., & Greenough, W. T. (2004). Environmental influences on cognitive and brain plasticity during aging. Journal of Gerontology: Medical Sciences, 59A, 940-957. doi:10.1093/gerona/59.9.M940
References ! 253 Kraus, N., & Chandrasekaran, B. (2010). Music training for the development of auditory skills. Nature Reviews Neuroscience, 11, 599-605. doi:10.1038/nrn2882 Kwok, L. (2008). Sound Studio (Version 3.5.5). New York City: Felt Tip inc. Lang, P. J., Bradley, M. M., & Cuthbert, B. N. (2008). International affective picture system (IAPS): Affective ratings of pictures and instruction manual. Technical Report A-8. Gainesville, Florida: University of Florida. Latinus, M., & Belin, P. (2011). Human voice perception. Current Biology, 21, R143R145. doi:10.1016/j.cub.2010.12.033 Laukka, P. (2005). Categorical perception of vocal emotion expressions. Emotion, 5, 277-295. doi:10.1037/1528-3542.5.3.277 Laukka, P., & Juslin, P. N. (2007). Similar patterns of age-related differences in emotion recognition from speech and music. Motivation and Emotion, 31, 182191. doi:10.1007/s11031-007-9063-z Laukka, P., Linnman, C., Ahs, F., Pissiota, A., Frans, Ö., Faria, V., … Furmark, T. (2008). In a nervous voice: Acoustic analysis and perception of anxiety in social phobics' speech. Journal of Nonverbal Behavior, 32, 195-214. doi:10.1007/s10919-008-0055-9 Lawrence, A. D., Goerendt, I. K., & Brooks, D. J. (2007). Impaired recognition of facial expressions of anger in Parkinson's disease patients acutely withdrawn from dopamine replacement therapy. Neuropsychologia, 45, 65-74. doi:10.1016/j.neuropsychologia.2006.04.016 Leitman, D. I., Laukka, P., Juslin, P. N., Saccente, E., Butler, P., & Javitt, D. C. (2010). Getting the cue: Sensory contributions to auditory emotion recognition impairments in schizophrenia. Schizophrenia Bulletin, 36, 545-556. doi:10.1093/schbul/sbn115 Leitman, D. I., Wolf, D. H., Ragland, J. D., Laukka, P., Loughead, J., Valdez, J. N., … Gur, R. C. (2010). "It's not what you say, but how you say it": A reciprocal temporo-frontal network for affective prosody. Frontiers in Human Neuroscience, 4, 1-13. doi:10.3389/fnhum.2010.00019
References ! 254 Levenson, R. W. (1999). The intrapersonal functions of emotion. Cognition and Emotion, 13, 481-504. doi:10.1080/026999399379159 Levenson, R. W. (2011). Basic emotion questions. Emotion Review, 3, 1-8. doi:10.1177/1754073911410743 Lezak, M. D., Howieson, D. B., & Loring, D. W. (2004). Neuropsychological Assessment (4th ed.). New York: Oxford University Press. Lima, C. F., & Castro, S. L. (2009). TIPI - Inventário de Personalidade de 10 Itens, versão portuguesa [Ten-Item Personality Inventory, portuguese version]. Retrieved from http://homepage.psy.utexas.edu/homepage/faculty/gosling/scales_we.htm#Ten% 20Item%20Personality%20Measure%20(TIPI) Lima, C. F., & Castro, S. L. (2011). Emotion recognition in music changes across the adult life span. Cognition and Emotion, 25, 585-598. doi:10.1080/02699931.2010.502449 Lima, C. F., & Castro, S. L. (2011). Speaking to the trained ear: Musical expertise enhances the recognition of emotions in speech prosody. Emotion, 11, 10211031. doi:10.1037/a002452 Lima, C. F., Garrett, C., & Castro, S. L. (submitted). Perceiving emotions in music and in speech prosody is dissociated in Parkinson’s disease. Lindner, J., & Rosén, L. (2006). Decoding of emotion through facial expression, prosody and verbal content in children and adolescents with Asperger's syndrome. Journal of Autism and Developmental Disorders, 36, 769-777. doi:10.1007/s10803-006-0105-2 Lindquist, K. A., & Barrett, L. F. (2010). Emotional complexity. In M. Lewis, J. M. Haviland-Jones & L. F. Barrett (Eds.), Handbook of Emotions (3rd ed., pp. 513530). New York: The Guilford Press. Liu, F., Patel, A. D., Fourcin, A., & Stewart, L. (2010). Intonation processing in congenital amusia: Discrimination, identification and imitation. Brain, 133, 1682-1693. doi:10.1093/brain/awq089
References ! 255 Livingstone, S. R., Muhlberger, R., Brown, A. R., & Thompson, W. F. (2010). Changing musical emotion: A computational rule system for modifying score and performance. Computer Music Journal, 34, 41-64. Lundqvist, D., Flykt, A., & Öhman, A. (1998). The Karolinska Directed Emotional Faces – KDEF [CD]. Department of Clinical Neuroscience, Psychology section, Karolinska Institutet. ISBN 91-630-7164-9. Makarova, V., & Petrushin, V. A. (2002, September). RUSLANA: A database of Russian Emotional Utterances. Paper session presented at the 7th International Conference on Spoken Language Processing, Denver, Colorado, USA. Marques, C., Moreno, S., Castro, S. L., & Besson, M. (2007). Musicians detect pitch violations in a foreign language better than nonmusicians: Behavioral and electrophysiological evidence. Journal of Cognitive Neuroscience, 19, 14531463. doi:10.1162/jocn.2007.19.9.1453 Martino, A. D., Shehzad, Z., Kelly, C., Roy, A. K., Gee, D. G., Uddin, L. Q., … Milham, M. P. (2009). Relationship between cingulo-insular functional connectivity and autistic traits in neurotypical adults. American Journal of Psychiatry, 166(8), 891-899. doi:10.1176/appi.ajp.2009.08121894 Mather, M., & Carstensen, L. (2005). Aging and motivated cognition: The positivity effect in attention and memory. Trends in Cognitive Sciences, 9, 496-502. doi:10.1016/j.tics.2005.08.005 Mather, M., & Knight, M. (2005). Goal-directed memory: The role of cognitive control in older adults' emotional memory. Psychology and Aging, 20, 554-570. doi:10.1037/0882-7974.20.4.554 Matsumoto, D., & Hwang, H. S. (2011). Evidence for training the ability to read microexpressions of emotion. Motivation and Emotion, 35, 181-191. doi:10.1007/s11031-011-9212-2 Matsumoto, D., Keltner, D., Shiota, M. N., O'Sullivan, M., & Frank, M. (2010). Facial expressions of emotion. In M. Lewis, J. M. Haviland-Jones & L. F. Barrett (Eds.), Handbook of Emotions (3rd ed., pp. 211-234). New York: The Guilford Press.
References ! 256 Matsumoto, D., LeRoux, J., Wilson-Cohn, C., Raroque, J., Kooken, K., Ekman, P., … Goh, A. (2000). A new test to measure emotion recognition ability: Matsumoto and Ekman's japanese and caucasian brief affect recognition test (JACBART). Journal of Nonverbal Behavior, 24, 179-209. doi:10.1023/A:1006668120583 McDermott, J. (2008). The evolution of music. Nature, 453, 287-288. doi:10.1038/453287a McDermott, J. (2009). What can experiments reveal about the origins of music? Current Directions in Psychological Science, 18, 164-168. doi:10.1111/j.14678721.2009.01629.x McDonald, C., & Stewart, L. (2008). Uses and functions of music in congenital amusia. Music Perception, 25, 345-355. doi:10.1525/MP.2008.25.4.345 Menon, V., & Levitin, D. J. (2005). The rewards of music listening: Response and physiological connectivity of the mesolimbic system. NeuroImage, 25, 175-184. doi:10.1016/j.neuroimage.2005.05.053 Meyer, L. (1956). Emotion and meaning in music. Chicago: Chicago University Press. Meyer, M., Steinhauer, K., Alter, K., Friederici, A. D., & Cramon, D. Y. v. (2004). Brain activity varies with modulation of dynamic pitch variance in sentence melody. Brain and Language, 89, 277-289. doi:10.1016/S0093-934X(03)00350X Milders, M., Fuchs, S., & Crawford, J. (2003). Neuropsychological impairments and changes in emotional and social behaviour following severe traumatic brain injury. Journal of Clinical and Experimental Neuropsychology, 25, 157-172. doi:10.1076/jcen.25.2.157.13642 Mill, A., Allik, J., Realo, A., & Valk, R. (2009). Age-related differences in emotion recognition ability: A cross-sectional study. Emotion, 9, 619-630. doi:10.1037/a0016562 Mitchell, R. L. C. (2007). Age-related decline in the ability to decode emotional prosody: Primary or secondary phenomenon? Cognition and Emotion, 21, 14351454. doi:10.1080/02699930601133994
References ! 257 Mitchell, R. L. C., & Bouças, S. B. (2009). Decoding emotional prosody in Parkinson's disease and its potential neuropsychological basis. Journal of Clinical and Experimental Neuropsychology, 31, 553-564. doi:10.1080/13803390802360534 Mitchell, R. L. C., Kingston, R. A., & Bouças, S. B. (2011). The specificity of agerelated decline in interpretation of emotion cues from prosody. Psychology and Aging, 26, 406-414. doi:10.1037/a0021861 Mitchell, R. L. C., & Ross, E. D. (2008). fMRI evidence for the effect of verbal complexity on lateralization of the neural response associated with decoding prosodic emotion. Neuropsychologia, 46, 2880-2887. doi:10.1016/j.neuropsychologia.2008.05.024 Mithen, S. (2007). The singing Neanderthals: The origins of music, language, mind, and body. Cambridge: Harvard University Press. Mitterschiffthaler, M. T., Fu, C. H. Y., Dalton, J. A., Andrew, C. M., & Williams, S. C. R. (2007). A functional MRI study of happy and sad affective states induced by classical music. Human Brain Mapping, 28, 1150-1162. doi:10.1002/hbm.20337 Mohn, C., Argstatter, H., & Wilker, F. W. (in press). Perception of six basic emotions in music. Psychology of Music. Advance online publication. doi:10.1177/0305735610378183 Monnot, M., Nixon, S., Lovallo, W., & Ross, E. (2001). Altered emotional perception in alcoholics: Deficits in affective prosody comprehension. Alcoholism: Clinical and Experimental Research, 25, 362-369. doi:10.1111/j.15300277.2001.tb02222.x Montag, C., Reuter, M., & Axmacher, N. (in press). How one's favorite song activates the reward circuitry of the brain: Personality matters! Behavioural Brain Research. Advance online publication. doi:10.1016/j.bbr.2011.08.012 Moreno, S., Marques, C., Santos, A., Santos, M., Castro, S. L., & Besson, M. (2009). Musical training influences linguistic abilities in 8-year-old children: More evidence for brain plasticity. Cerebral Cortex, 19, 712-723. doi:10.1093/cercor/bhn120
References ! 258 Mu, Q., Xie, J., Wen, Z., Weng, Y., & Shuyun, Z. (1999). A quantitative MR study of the hippocampal formation, the amygdala, and the temporal horn of the lateral ventricle in healthy subjects 40 to 90 years of age. American Journal of Neuroradiology, 20, 207-2011. Musacchia, G., Sams, M., Skoe, E., & Kraus, N. (2007). Musicians have enhanced subcortical auditory and audiovisual processing of speech and music. Proceedings of the National Academy of Sciences, 104, 15894-15898. doi:10.1073/pnas.0701498104 Nambu, A. (2008). Seven problems on the basal ganglia. Current Opinion in Neurobiology, 18, 595-604. doi:10.1016/j.conb.2008.11.001 Nan, Y., Sun, Y., & Peretz, I. (2010). Congenital amusia in speakers of a tone language: Association with lexical tone agnosia. Brain, 133, 2635-2642. doi:10.1093/brain/awq178 Nieminen, S., Istók, E., Brattico, E., Tervaniemi, M., & Huotilainen, M. (2011). The development of aesthetic responses to music and their underlying neural and psychological mechanisms. Cortex, 47, 1138-1146. doi:10.1016/j.cortex.2011.05.008 Nolte, J. (2002). The human brain: An introduction to its functional anatomy (5th ed.). St. Louis: Mosby, Inc. North, A. C., & Hargreaves, D. J. (2010). Music and marketing. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of music and emotion: Theory, research, applications. New York: Oxford University Press. Obeso, J. A., Marin, C., Rodriguez-Oroz, C., Blesa, J., Benitez-Temiño, B., MenaSegovia, J., … Olanow, C. W. (2008). The basal ganglia in Parkinson's disease: Current concepts and unexplained observations. Annals of Neurology, 64, S30S46. doi:10.1002/ana.21481 Omar, R., Hailstone, J. C., Warren, J. E., Crutch, S. J., & Warren, J. D. (2010). The cognitive organization of music knowledge: A clinical analysis. Brain, 133, 1200-1213. doi:10.1093/brain/awp345
References ! 259 Omar, R., Henley, S., Bartlett, J., Hailstone, J., Gordon, E., Sauter, D., … Warren, J. D. (2011). The structural neuroanatomy of music emotion recognition: Evidence from frontotemporal lobar degeneration. NeuroImage, 56, 1814-1821. doi:10.1016/j.neuroimage.2011.03.002 Orbelo, D., Grim, M., Talbott, R., & Ross, E. (2005). Impaired comprehension of affective prosody in elderly subjects is not predicted by age-related hearing loss or age-related cognitive decline. Journal of Geriatric Psychiatry and Neurology, 18, 25-32. doi:10.1177/0891988704272214 Orgeta, V. (2010). Effects of age and task difficulty on recognition of facial affect. Journal of Gerontology: Psychological Sciences, 65B, 323-327. doi:10.1093/geronb/gbq007 Pantev, C., & Herholz, S. C. (in press). Plasticity of the human auditory cortex related to musical training. Neuroscience and Biobehavioral Reviews. Advance online publication. doi:10.1016/j.neubiorev.2011.06.010 Pantev, C., Oostenveld, R., Engelien, A., Ross, B., Roberts, L., & Hoke, M. (1998). Increased auditory cortical representation in musicians. Nature, 392, 811-814. doi:10.1038/33918 Parbery-Clark, A., Skoe, E., Lam, C., & Kraus, N. (2009). Musician enhancement for speech-in-noise. Ear & Hearing, 30, 653-661. doi:10.1097/AUD.0b013e3181b412e9 Patel, A. D. (2003). Language, music, syntax and the brain. Nature Neuroscience, 6, 674-681. doi:10.1038/nn1082 Patel, A. D. (2008a). Music, language, and the brain. New York: Oxford University Press. Patel, A. D. (2008b). A neurobiologial strategy for exploring links between emotion recognition in music and speech. Behavioral and Brain Sciences, 31, 589-590. doi:10.1017/S0140525X0800544X Patel, A. D. (2010). Music, biological evolution, and the brain. In M. Bailar (Ed.), Emerging disciplines (pp. 91-144). Texas: Rice University Press.
References ! 266 Samanez-Larkin, G. R., & Carstensen, L. (2011). Socioemotional functioning and the aging brain. In J. Decety & J. T. Cacioppo (Eds.), The handbook of social neuroscience. New York: Oxford University Press. Sander, D., Grandjean, D., Pourtois, G., Schwartz, S., Seghier, M. L., Scherer, K. R., & Vuilleumier, P. (2005). Emotion and attention interactions in social cognition: Brain regions involved in processing anger prosody. NeuroImage, 28, 848-858. doi:10.1016/j.neuroimage.2005.06.023 Särkämö, T., Tervaniemi, M., Laitinen, S., Forsblom, A., Soinila, S., Mikkonen, M., … Hietanen, M. (2008). Music listening enhances cognitive recovery and modd after middle cerebral artery stroke. Brain, 131, 866-876. doi: 10.1093/brain/awn013 Sauter, D. (2006). An investigation into vocal expressions of emotions: The roles of valence, culture, and acoustic factors (Unpublished doctoral dissertation). University College London, London. Sauter, D., & Eimer, M. (2009). Rapid detection of emotion from human vocalizations. Journal of Cognitive Neuroscience, 22, 474-481. doi:10.1162/jocn.2009.21215 Sauter, D., Guen, O. L., & Haun, D. B. M. (in press). Categorical perception of emotional expressions does not require lexical categories. Emotion. Sauter, D. A., Eisner, F., Calder, A. J., & Scott, S. K. (2010). Perceptual cues in nonverbal vocal expressions of emotion. The Quarterly Journal of Experimental Psychology, 63, 2251-2272. doi:10.1080/17470211003721642 Sauter, D. A., Eisner, F., Ekman, P., & Scott, S. K. (2010). Cross-cultural recognition of basic emotions through nonverbal emotional vocalizations. Proceedings of the National Academy of Sciences, 107, 2408-2412. doi:10.1073/pnas.0908239106 Sauter, D. A., & Scott, S. K. (2007). More than one kind of happiness: Can we recognize vocal expressions of different positive states? Motivation and Emotion, 31, 192-199. doi: 10.1007/s11031-007-9065-x Schellenberg, E. G. (2005). Music and cognitive abilities. Current Directions in Psychological Science, 14, 317-320. doi: 10.1111/j.0963-7214.2005.00389.x
References ! 267 Schellenberg, E. G. (2006). Long-term positive associations between music lessons and IQ. Journal of Educational Psychology, 98, 457-468. doi:10.1037/00220663.98.2.457 Schellenberg, E. G., & Moreno, S. (2009). Music lessons, pitch processing, and g. Psychology of Music, 38, 209-221. doi: 10.1177/0305735609339473 Scherer, K. R. (1986). Vocal affect expression: A review and a model for future research. Psychological Bulletin, 99, 143-165. doi: 10.1037/0033-2909.99.2.143 Scherer, K. R. (1995). Expression of emotion in voice and music. Journal of Voice, 9, 235-248. doi:10.1016/S0892-1997(05)80231-0 Scherer, K. R. (2003). Vocal communication of emotion: A review of research paradigms. Speech Communication, 40, 227-256. doi:10.1016/S01676393(02)00084-5 Scherer, K. R. (2004). Which emotions can be induced by music? What are the underlying mechanisms? And how can we measure them? Journal of New Music Research, 33, 239-251. doi:10.1080/0929821042000317822 Scherer, K. R. (2005). What are emotions? And how can they be measured? Social Science Information, 44, 695-729. doi: 10.1177/0539018405058216 Scherer, K. R., Banse, R., & Wallbott, H. G. (2001). Emotion inferences from vocal expression correlate across languages and cultures. Journal of Cross-Cultural Psychology, 32, 76-92. doi: 10.1177/0022022101032001009 Scherer, K. R., Banse, R., Wallbott, H. G., & Goldbeck, T. (1991). Vocal cues in emotion ecoding and decoding. Motivation and Emotion, 15, 123-148. doi:10.1007/BF00995674 Scherer, K. R., Johnstone, T., & Klasmeyer, G. (2003). Vocal expression of emotion. In R. Davidson, K. R. Scherer & H. Goldsmith (Eds.), Handbook of the affective sciences (pp. 433-456). New York: Oxford University Press.
References ! 268 Scherer, K. R., & Scherer, U. (in press). Assessing the ability to recognize facial and vocal expressions of emotion: Construction and validation of the emotion recognition index. Journal of Nonverbal Behavior. Advance online publication. doi:10.1007/s10919-011-0115-4 Schiess, M. C., Zheng, H., Soukup, V. M., Bonnen, J. G., & Nauta, H. J. W. (2000). Parkinson's disease subtypes: Clinical classification and ventricular cerebrospinal fluid analysis. Parkinsonism and Related Disorders, 6, 69-76. doi:10.1016/S1353-8020(99)00051-6 Schirmer, A., Escoffier, N., Zysset, S., Koester, D., Striano, T., & Friederici, A. D. (2008). When vocal processing gets emotional: On the role of social orientation in relevance detection by the human amygdala. NeuroImage, 40, 1402-1410. doi:10.1016/j.neuroimage.2008.01.018 Schirmer, A., & Kotz, S. A. (2006). Beyond the right hemisphere: Brain mechanisms mediating vocal emotional processing. Trends in Cognitive Sciences, 10, 24-30. doi:10.1016/j.tics.2005.11.009 Schirmer, A., Kotz, S. A., & Friederici, A. D. (2002). Sex differentiaties the role of emotional prosody during word processing. Cognitive Brain Research, 14, 228233. doi:10.1016/S0926-6410(02)00108-8 Schirmer, A., Kotz, S. A., & Friederici, A. D. (2005). On the role of attention for the processing of emotions in speech: Sex differences revisited. Cognitive Brain Research, 24, 442-452. doi:10.1016/S0926-6410(02)00108-8 Schirmer, A., Striano, T., & Friederici, A. D. (2005). Sex differences in the preattentive processing of vocal emotional expressions. NeuroReport, 16, 635-639. doi:10.1016/j.cogbrainres.2005.02.022 Schirmer, A., Zysset, S., Kotz, S. A., & Cramon, D. Y. v. (2004). Gender differences in the activation of inferior frontal cortex during emotional speech perception. NeuroImage, 21, 1114-1123. doi:10.1016/j.neuroimage.2003.10.048
References ! 269 Schmidt, A. T., Hanten, G. R., Li, X., Orsten, K. D., & Levin, H. S. (2010). Emotion recognition following pediatric traumatic brain injury: Longitudinal analysis of emotional prosody and facial emotion recognition. Neuropsychologia, 48, 28692877. doi:10.1016/j.neuropsychologia.2010.05.029 Schön, D., & Besson, M. (2001). Comparison between language and music. Annals of the New York Academy of Sciences, 930, 232-258. doi: 10.1111/j.17496632.2001.tb05736.x Schön, D., Gordon, R., Campagne, A., Magne, C., Astésamo, C., Anton, J. L., & Besson, M. (2010). Similar cerebral networks in language, music and song perception. NeuroImage, 51, 450-461. doi:10.1016/j.neuroimage.2010.02.023 Schön, D., Magne, C., & Besson, M. (2004). The music of speech: Music training facilitates pitch processing in both music and language. Psychophysiology, 41, 341-349. doi:10.1111/1469-8986.00172.x Schön, D., Ystad, S., Kronland-Martinet, R., & Besson, M. (2010). The evocative power of sounds: Conceptual priming between words and nonverbal sounds. Journal of Cognitive Neuroscience, 22, 1026-1035. doi:10.1162/jocn.2009.21302 Schott, B. H., Niehaus, L., Wittmann, B. C., Schutze, H., Seidenbecher, C. I., Heinze, H.-J., & Düzel, E. (2007). Ageing and early-stage Parkinson's disease affect separable neural mechanisms of mesolimbic reward processing. Brain, 130, 2412-2424. doi:10.1093/brain/awm147 Schröder, C., Nikolova, Z. T., & Dengler, R. (2010). Changes of emotional prosody in Parkinson's disease. Journal of the Neurological Sciences, 289, 32-35. doi:10.1016/j.jns.2009.08.038 Scott, S. K., Blank, C. C., Rosen, S., & Wise, R. J. S. (2000). Identification of a pathway for intelligible speech in the left temporal lobe. Brain, 123, 2400-2406. doi: 10.1093/brain/123.12.2400 Scott, S., Caird, F., & Williams, B. (1984). Evidence for an apparent sensory speech disorder in Parkinson's disease. Journal of Neurology, Neurosurgery, and Psychiatry, 47, 840-843. doi:10.1136/jnnp.47.8.840
References ! 270 Scott, S. K., Sauter, D., & McGettigan, C. (2010). Brain mechanisms for processing perceived emotional vocalizations in humans. In M. B. Stefan (Ed.), Handbook of Behavioral Neuroscience (Vol. 19, pp. 187-197). London: Academic Press. Silver, H., Goodman, C., Knoll, G., & Isakov, V. (2004). Brief emotion training improves recognition of facial emotions in chronic schizophrenia: A pilot study. Psychiatry Research, 128, 147-154. doi:10.1016/j.psychres.2004.06.002 Simões, M., Firmino, H., Vilar, M., & Martins, M. (2007). Montreal Cognitive Assessment (MOCA) - Versão experimental portuguesa [portuguese experimental version]. Retrieved from http://www.mocatest.org Sloboda, J. A., & Juslin, P. N. (2010). At the interface between the inner and outer world: Psychological perspectives. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of Music and Emotion: Theory, Research, Applications (pp. 73-97). New York: Oxford University Press. Smith, C. A., & Lazarus, R. S. (1990). Emotion and adaptation. In L. Pervin (Ed.), Handbook of personality: Theory and research (1st ed.). New York: Guilford Press. Snowdon, C. T., & Teie, D. (2010). Affective responses in tamarins elicited by speciesspecific music. Biology Letters, 6, 30-32. doi:10.1098/rsbl.2009.0593 Speedie, L., Brake, N., Folstein, S., Bowers, D., & Heilman, K. (1990). Comprehension of prosody in Huntington's disease. Journal of Neurology, Neurosurgery, and Psychiatry, 53, 607-610. doi:10.1136/jnnp.53.7.607 Spencer, H. (1857). The origin and function of music. Fraser's Magazine, 56, 396-408. St. Jacques, P., Bessette-Symons, B., & Cabeza, R. (2009). Functional neuroimaging studies of aging and emotion: Fronto-amygdalar differences during emotional perception and episodic memory. Journal of the International Neuropsychological Society, 15, 819-825. doi:10.1017/S1355617709990439 St. Jacques, P., Dolcos, F., & Cabeza, R. (2008). Effects of aging on functional connectivity of the amygdala during negative evaluation: A network analysis of fMRI data. Neurobiology of Aging, 31, 315-327. doi:10.1016/j.neurobiolaging.2008.03.012
References ! 271 Staroniewicz, P., & Majewski, W. (2009). Polish emotional speech database - recording and preliminary validation. Lectures Notes in Computer Science, 5641, 42-49. doi:10.1007/978-3-642-03320-9_5 Stewart, L., Kriegstein, K. v., Dalla-Bella, S., Warren, J. D., & Griffiths, T. D. (2008). Disorders of musical cognition. In S. Hallam, I. Cross & M. H. Thaut (Eds.), Oxford Handbook of Music Psychology (pp. 184-196). New York: Oxford University Press. Strait, D., Kraus, N., Skoe, E., & Ashley, R. (2009). Musical experience and neural efficiency - effects of training on subcortical processing of vocal expressions of emotion. European Journal of Neuroscience, 29, 661-668. doi:10.1111/j.14609568.2009.06617.x Sullivan, S., & Ruffman, T. (2004). Social understanding: How does it fare with advancing years? British Journal of Psychology, 95, 1-18. doi:10.1348/000712604322779424 Suzuki, A., Hoshino, T., Shigemasu, K., & Kawamura, M. (2006). Disgust-specific impairment of facial expression recognition in Parkinson's disease. Brain, 129, 707-717. doi: 10.1093/brain/awl011 Taler, V., Baum, S., Chertkow, H., & Saumier, D. (2008). Comprehension of grammatical and emotional prosody is impaired in Alzheimer's disease. Neuropsychology, 22, 188-195. doi:10.1037/0894-4105.22.2.188 Tavares, G. M. (2004). Poesia 1 [Poetry 1]. Lisboa: Relógio D’Água Editores. Thaut, M. H. (2005). The future of music in therapy and medicine. Annals of the New York Academy of Sciences, 1060, 303-308. doi:10.1196/annals.1360.023 Thaut, M. H., Gardiner, J. C., Holmberg, D., Horwitz, J., Kent, L., Andrews, G., … McIntosh, G. R. (2009). Neurologic music therapy improves executive function and emotional adjustment in traumatic brain injury rehabilitation. Annals of the New York Academy of Sciences 1169, 406-416. doi:10.1111/j.17496632.2009.04585.x
References ! 272 Thaut, M. H., & Wheeler, B. (2010). Music therapy. In P. N. Juslin & J. A. Sloboda (Eds.), Handbook of music and emotion: Theory, research, applications. New York: Oxford University Press. Thompson, W. F. (2009). Music, though, and feeling: Understanding the psychology of music. New York: Oxford University Press. Thompson, W. F., & Balkwill, L.-L. (2006). Decoding speech prosody in five languages. Semiotica, 158, 407-424. doi:10.1515/SEM.2006.017 Thompson, W. F., Schellenberg, E. G., & Husain, G. (2004). Decoding speech prosody: Do music lessons help? Emotion, 4, 46-64. doi:10.1037/1528-3542.4.1.46 Tracy, J. L., & Robins, R. W. (2008). The automaticity of emotion recognition. Emotion, 8, 81-95. doi:10.1037/1528-3542.8.1.81 Trauner, D., Ballantyne, A., Friedland, S., & Chase, C. (1996). Disorders of affective and linguistic prosody in children after early unilateral brain damage. Annals of Neurology, 39, 361-367. doi:10.1002/ana.410390313 Trehub, S. E., & Hannon, E. E. (2006). Infant music perception: Domain-general or domain-specific mechanisms? Cognition, 100, 73-99. doi:10.1016/j.cognition.2005.11.006 Trenerry, M. R., Crosson, B., Deboe, J., & Leber, W. R. (1995). The Stroop Neuropsychological Screening Test Manual. Tampa, Florida: Psychological Assessment Resources. Tricht, M. J. v., Smeding, H. M. M., Speelman, J. D., & Schmand, B. A. (2010). Impaired emotion recognition in music in Parkinson's disease. Brain and Cognition, 74, 58-65. doi:10.1016/j.bandc.2010.06.005 Trimmer, C. G., & Cuddy, L. L. (2008). Emotional intelligence, not music training, predicts recognition of emotional speech prosody. Emotion, 8, 838-849. doi:10.1037/a0014080
References ! 273 Uekermann, J., Abdel-Hamid, M., Lehmkämper, C., Vollmoeller, W., & Daum, I. (2008). Perception of affective prosody in major depression: A link to executive functions? Journal of the International Neuropsychological Society, 14, 552561. doi:10.10170S1355617708080740 Utter, A. A., & Basso, M. A. (2008). The basal ganglia: An overview of circuits and function. Neuroscience and Biobehavioral Reviews, 32, 333-342. doi:10.1016/j.neubiorev.2006.11.003 Vieillard, S., Peretz, I., Gosselin, N., Khalfa, S., Gagnon, L., & Bouchard, B. (2008). Happy, sad, scary and peaceful musical excerpts for research on emotions. Cognition and Emotion, 22, 720-752. doi:10.1080/02699930701503567 Vuoskoski, J. K., & Eerola, T. (2011). The role of mood and personality in the perception of emotions represented by music. Cortex, 47, 1099-1106. doi:10.1016/j.cortex.2011.04.011 Vytal, K., & Hamann, S. (2010). Neuroimaging support for discrete neural correlates of basic emotions: A voxel-based meta-analysis. Journal of Cognitive Neuroscience, 22, 2864-2885. doi:10.1162/jocn.2009.21366 Wagner, H. L. (1993). On measuring performance in category judgment studies of nonverbal behavior. Journal of Nonverbal Behavior, 17, 3-28. doi:10.1007/BF00987006 Walhovd, K. B., Fjell, A. M., Reinvang, I., Lundervold, A., Dale, A. M., Eilertsen, D. E., … Fischl, B. (2005). Effects of age on volumes of cortex, white matter and subcortical structures. Neurobiology of Aging, 26, 1261-1270. doi:10.1016/j.neurobiolaging.2005.05.020 Watson, G. S., & Leverenz, J. B. (2010). Profile of cognitive impairment in Parkinson's disease. Brain Pathology, 20, 640-645. doi:10.1111/j.1750-3639.2010.00373.x Wechsler, D. (2008). WAIS -III - Escala de Inteligência de Wechsler para Adultos - 3.ª edição [portuguese version of the Wechsler Adult Intelligence Scale - Third Edition, WAIS-III]. Lisbon: CEGOC-TEA. Whitman, W. (1855/2005). Leaves of grass. New York: Oxford University Press.
References ! 274 Wiethoff, S., Wildgruber, D., Kreifelts, B., Becker, H., Herbert, C., Grodd, W., & Ethofer, T. (2008). Cerebral processing of emotional prosody: Influence of acoustic parameters and arousal. NeuroImage, 39, 885-893. doi:10.1016/j.neuroimage.2007.09.028 Wildgruber, D., Ackermann, H., Kreifelts, B., & Ethofer, T. (2006). Cerebral processing of linguistic and emotional prosody: fMRI studies. Progress in Brain Research, 156, 249-268. doi:10.1016/S0079-6123(06)56013-3 Wildgruber, D., Riecker, A., Hertrich, I., Erb, M., Grodd, W., Ethofer, T., & Ackermann, H. (2005). Identification of emotional intonation evaluated by fMRI. NeuroImage, 24, 1233-1241. doi:10.1016/j.neuroimage.2004.10.034 Williams, L. M., Brown, K. J., Palmer, D., Liddell, B. J., Kemp, A. H., Olivieri, G., … Gordon, E. (2006). The mellow years?: Neural basis of improving emotional stability over age. The Journal of Neuroscience, 26, 6422-6430. doi:10.1523/ JNEUROSCI.0022-06.2006 Williams, L. M., Mathersul, D., Palmer, D. M., Gur, R. C., Gur, R. E., & Gordon, E. (2009). Explicit identification and implicit recognition of facial emotions: I. Age effects in males and females across 10 decades. Journal of Clinical and Experimental Neuropsychology, 31, 257-277. doi:10.1080/13803390802255635 Wittfoth, M., Schröder, C., Schardt, D. M., Dengler, R., Heinze, H.-J., & Kotz, S. A. (2010). On emotional conflict: Interference resolution of happy and angry prosody reveals valence-specific effects. Cerebral Cortex, 20, 383-392. doi:10.1093/cercor/bhp106 Wong, P., Skoe, E., Russo, N., Dees, T., & Kraus, N. (2007). Musical experience shapes human brainstem encoding of linguistic pitch patterns. Nature Neuroscience, 10, 420-422. doi:10.1038/nn1872 Wu, T., Yang, Y., Wu, Z., & Li, D. (2006, June). MASC: A speech corpus in Mandarin for emotion analysis and affective speaker recognition. Paper session presented at the IEEE Odyssey 2006 Workshop on speaker and language recognition, San Juan, Puerto Rico.
References ! 275 Yesavage, J. A., Brink, T. L., Rose, T. L., Lum, O., Huang, V., Adey, M., & Leirer, V. O. (1983). Development and validation of a geriatric depression screening scale: A preliminary report. Journal of Psychiatric Research, 17, 37-49. doi:10.1016/0022-3956(82)90033-4 Yip, J., Lee, T., Ho, S.-L., Tsang, K.-L., & Li, L. (2003). Emotion recognition in patients with idiopathic Parkinson's disease. Movement Disorders, 18, 11151122. doi:10.1002/mds.10497 Zatorre, R. J. (2001). Neural specializations for tonal processing. Annals of the New York Academy of Sciences, 930, 193-210. doi:10.1111/j.17496632.2001.tb05734.x Zatorre, R. J., Belin, P., & Penhune, V. B. (2002). Structure and function of auditory cortex: Music and speech. Trends in Cognitive Sciences, 6, 37-46. doi:10.1016/S1364-6613(00)01816-7 Zentner, M., & Eerola, T. (2010). Rhythmic engagement with music in infancy. Proceedings of the National Academy of Sciences, 107, 5768-5773. doi:10.1073/pnas.1000121107 Zentner, M., Grandjean, D., & Scherer, K. R. (2008). Emotions evoked by the sound of music: Characterization, classification, and measurement. Emotion, 8, 494-521. doi:10.1037/1528-3542.8.4.494 Zentner, M., & Kagan, J. (1996). Perception of music by infants. Nature, 383, 29. doi:10.1038/383029a0
Appendices ! 282
Appendices ! 283 3. Structural characteristics of the musical excerpts used in Study 1 and Study 2, Chapter VI Structural characteristics of the excerpts used in the Study 1 (all, n = 56) and Study 2 (coloured in grey, n = 40). Note: Tempo corresponds to crochet beats per minute; Mode, 1 = major, 2 = minor, 3 = indefinite; Pedal, 0 = without, 1 = with; Tone Dens. = Tone density (number of melodic events / total duration); M.P. Range = Melodic pitch range (number of semitones between the highest and the lowest tone of the melody); Diss. = Dissonance (minimum 1, maximum 5); Un. = Unexpected events (minimum 1, maximum 5); Rhy. Irr. = Rhythmic Irregularities (minimum 1, maximum 5). # Emotion Stimulus Tempo Mode Pedal Tone Dens. M.P. Range Diss. Un. Rhy. Irr. 1 Happy G01 112 1 0 3.4 17 1 1 1 2 Happy G02 143 1 0 4.4 21 1 1 1 3 Happy G03 91 1 0 3.7 18 1 1 1 4 Happy G04 180 1 0 5.0 20 1 1 1 5 Happy G05 100 1 0 5.5 22 1 1 1 6 Happy G06 195 1 0 5.2 17 1 1 1 7 Happy G07 180 1 0 3.2 14 1 1 1 8 Happy G08 143 1 0 3.7 21 1 1 1 9 Happy G09 143 1 0 2.5 10 1 1 1 10 Happy G10 160 1 0 3.8 12 1 1 1 11 Happy G11 140 1 0 4.5 19 1 1 1 12 Happy G12 107 1 0 2.6 12 1 1 1 13 Happy G13 120 1 0 3.3 19 1 1 1 14 Happy G14 112 1 0 5.7 19 1 1 1 16 Sad T01 40 2 1 0.8 5 1 1 1 17 Sad T02 40 2 1 0.7 14 1 1 1 18 Sad T03 40 2 1 0.4 12 1 1 1 19 Sad T04 54 2 1 0.5 9 1 1 1 20 Sad T05 60 2 1 0.9 5 1 1 1 21 Sad T06 54 2 1 0.6 8 1 1 1 22 Sad T07 42 2 1 0.8 10 1 1 1
Appendices ! 284 15 Sad T08 40 2 1 0.6 7 1 1 1 23 Sad T09 54 2 1 1.3 8 1 1 1 24 Sad T10 40 2 1 0.58 3 1 1 1 25 Sad T11 48 2 1 0.8 10 1 1 1 26 Sad T12 60 2 1 0.5 9 1 1 1 27 Sad T13 48 2 1 0.7 7 1 1 1 28 Sad T14 48 2 1 1.0 10 1 1 1 29 Scary P01 100 2 0 3.6 4 2 3 3 30 Scary P02 140 2 0 2.2 7 2 2 1 31 Scary P03 100 2 0 1.4 18 2 4 3 32 Scary P04 100 2 0 0.9 2 4 5 4 33 Scary P05 100 2 0 3.4 33 2 1 2 34 Scary P06 96 2 0 0.9 43 5 2 2 35 Scary P07 100 2 0 1.6 13 2 3 4 36 Scary P08 96 2 0 4.3 19 2 3 4 37 Scary P09 44 2 0 0.7 7 1 1 1 38 Scary P10 120 2 0 1.1 12 2 4 3 39 Scary P11 44 3 0 1.4 4 2 5 4 40 Scary P12 75 2 0 1.9 28 2 5 2 41 Scary P13 72 2 0 2.4 8 2 4 1 42 Scary P14 172 2 0 3.1 36 1 3 4 43 Peaceful A01 60 1 1 0.4 12 1 1 1 44 Peaceful A02 72 1 1 0.9 20 1 1 1 45 Peaceful A03 72 1 1 0.7 9 1 1 1 46 Peaceful A04 54 1 1 0.8 13 1 1 1 47 Peaceful A05 69 1 1 0.7 10 1 1 1 48 Peaceful A06 33 1 1 1.4 13 1 1 1 49 Peaceful A07 69 1 1 0.9 5 1 1 1 50 Peaceful A08 80 1 1 1.2 15 1 1 1 51 Peaceful A09 72 1 1 1.1 12 1 1 1 52 Peaceful A10 80 1 1 2.0 24 1 1 1 53 Peaceful A11 88 1 1 1.6 14 1 1 1 54 Peaceful A12 75 1 1 0.9 12 1 1 1 55 Peaceful A13 75 1 1 1.1 6 1 1 1 56 Peaceful A14 72 1 1 1.6 17 1 1 1
Appendices ! 285 4. Raw ratings for each musical emotion in Study 2, Chapter VI Raw ratings (minimum 0, maximum 9) for the intended and non-intended emotions, as a function of age and expertise groups. Diagonal cells in bold indicate ratings on the excerpts’ intended emotion. Standard errors are presented in parentheses. Age group/ excerpt type Ratings across response categories Controls Musicians Happy Peaceful Sad Scary Happy Peaceful Sad Scary Younger Happy 7.6 (0.2) 1.6 0.2 0.2 7.2 (0.4) 1.7 0.2 0.2 Peaceful 2.8 5.5 (0.4) 4.6 0.2 4.3 4.9 (0.5) 3 0 Sad 0.4 3.3 7.4 (0.2) 2.1 0.4 4.6 7 (0.3) 0.9 Scary 1 0.4 3.7 7.1 (0.3) 0.8 1.2 3.5 5.2 (0.5) Middle-aged Happy 6.7 (0.5) 2.1 0.6 0.4 7.4 (0.3) 1.6 0.6 0.6 Peaceful 3.9 5 (0.5) 2.7 0.4 2.9 5.2 (0.4) 3.2 0.4 Sad 1.5 3.9 5.5 (0.5) 0.5 0.7 4.4 6 (0.3) 1.3 Scary 1.8 0.6 2.9 3.7 (0.4) 1.5 0.8 2.8 5.3 (0.4) !
Appendices ! 286
Appendices ! 287 5. Unbiased accuracy rates for each musical emotion in Study 2, Chapter VI Hu for each musical emotion, as a function of age and expertise groups. Values vary between 0 and 1. Standard errors are presented in parentheses. Age group / excerpt type Hu scores Controls Musicians Younger Happy .8 (0.03) .8 (0.04) Peaceful .4 (0.07) .4 (0.07) Sad .6 (0.05) .5 (0.06) Scary .7 (0.04) .7 (0.07) Middle-aged Happy .7 (0.05) .8 (0.03) Peaceful .4 (0.05) .5 (0.05) Sad .4 (0.06) .5 (0.05) Scary .5 (0.06) .7 (0.05)
Appendices ! 288
Appendices ! 289 6. Acoustic and perceptual characteristics of the prosodic stimuli presented in Chapter VII Acoustic (duration, F0 mean, F0 SD) and perceptual (recognition accuracy, intensity) characteristics of each sentence and pseudo-sentence included in the database of prosodic stimuli. Sentences In Stimulus, s refers to sentence, p to pseudo-sentence, A to one speaker, and B to the other speaker. # Stimulus Content Duration (ms) F0 (Hz) F0 variability (SD, Hz) Accuracy (%) Intensity (1-7) 1 1sA_angry1 estaMesa 1278 339 99 75 4.5 2 2sA_angry2 oRadio 1245 323 81 80 5 3 3sA_angry3 aqueleLivro 1367 347 85 95 5.8 4 4sA_angry4 aTerra 1306 355 68 70 4.9 5 5sA_angry5 oCao 1166 299 99 80 3.6 6 6sA_angry6 eleChega 1066 370 95 75 5.3 7 7sA_angry7 estaRoupa 1224 349 85 80 4.4 8 8sA_angry8 osJardins 1415 333 75 70 3.5 9 9sA_angry9 asPessoas 1712 358 86 70 4.9 10 10sA_angry10 haArvores 1426 368 65 75 5.3 11 11sA_angry11 osTigres 1572 279 76 75 4.8 12 12sA_angry12 oQuadro 1301 371 78 80 5.9 13 13sA_angry13 alguemFechou 1455 336 93 95 5.7 14 14sA_angry14 osJovens 1477 338 80 95 5.7 15 15sA_angry15 oFutebol 1487 377 82 70 4.9 16 16sA_angry16 elaViajou 1323 329 93 85 5.2 17 17sB_angry17 estaMesa 1116 266 53 75 5.3 18 18sB_angry18 oRadio 1079 289 56 95 5.9 19 19sB_angry19 aqueleLivro 1186 284 51 55 4.6 20 20sB_angry20 aTerra 1068 288 53 80 5.6 21 21sB_angry21 oCao 1059 278 65 75 4.8 22 22sB_angry22 eleChega 897 252 41 70 5.2
Appendices ! 290 23 23sB_angry23 estaRoupa 1127 305 54 55 5.5 24 24sB_angry24 osJardins 1160 283 47 90 6.1 25 25sB_angry25 asPessoas 1378 324 47 70 5.5 26 26sB_angry26 haArvores 1202 293 61 85 6.1 27 27sB_angry27 osTigres 1387 281 56 55 5.2 28 28sB_angry28 oQuadro 1200 262 60 95 4.7 29 29sB_angry29 alguemFechou 1296 305 52 85 6.1 30 30sB_angry30 osJovens 1198 302 48 75 6.3 31 31sB_angry31 oFutebol 1215 264 24 65 4.6 32 32sB_angry32 elaViajou 1137 264 42 80 5.9 33 33sA_disgust1 oRadio 1661 270 76 60 5.3 34 34sA_disgust2 oCao 1538 263 69 45 4.2 35 35sA_disgust3 haArvores 1596 289 76 55 4.5 36 36sA_disgust4 osTigres 1908 279 77 45 5.2 37 37sA_disgust5 oQuadro 1800 286 78 60 4.3 38 38sA_disgust6 alguemFechou 1852 281 67 45 5.4 39 39sB_disgust7 oCao 1376 272 62 50 5.6 40 40sB_disgust8 estaRoupa 1317 266 64 45 4.9 41 41sB_disgust9 asPessoas 1875 289 75 60 5 42 42sB_disgust10 haArvores 1545 294 54 45 5.8 43 43sB_disgust11 oQuadro 1530 287 57 50 5.1 44 44sB_disgust12 oFutebol 1608 278 51 45 5 45 45sA_fear1 estaMesa 1687 276 28 55 5.4 46 46sA_fear2 oRadio 1501 336 45 75 4.5 47 47sA_fear3 aqueleLivro 1488 310 43 55 3.9 48 48sA_fear4 eleChega 1803 290 15 55 4.5 49 49sA_fear5 osJardins 2104 313 33 50 5.1 50 50sA_fear6 asPessoas 1772 310 45 65 4.4 51 51sA_fear7 haArvores 1954 291 38 80 5.2 52 52sA_fear8 osTigres 1659 290 35 55 4 53 53sA_fear9 oQuadro 2009 309 34 65 4.7 54 54sA_fear10 alguemFechou 1759 290 58 75 5.6 55 55sA_fear11 oFutebol 1645 255 60 50 3.4 56 56sA_fear12 elaViajou 1173 287 21 55 4.7 57 57sB_fear13 estaMesa 1086 300 28 85 5.6 58 58sB_fear14 oRadio 1129 281 34 80 5.2 59 59sB_fear15 aqueleLivro 1125 308 33 75 5.2 60 60sB_fear16 oCao 1063 323 54 55 5.7 61 61sB_fear17 eleChega 913 326 27 60 5.1 62 62sB_fear18 estaRoupa 1173 311 40 45 4.1 63 63sB_fear19 osJardins 1212 325 45 60 4.4 64 64sB_fear20 haArvores 1245 290 30 75 5.2 65 65sB_fear21 osTigres 1400 305 41 80 5.3 66 66sB_fear22 oQuadro 1200 302 27 60 4.6 67 67sB_fear23 alguemFechou 1320 311 37 65 5.2 68 68sB_fear24 oFutebol 1300 314 41 70 5.1 69 69sA_happy1 estaMesa 1451 360 83 85 4.8 70 70sA_happy2 oRadio 1436 386 94 75 4.3 71 71sA_happy3 aqueleLivro 1588 379 102 90 4.7 72 72sA_happy4 oCao 1368 377 98 75 5.5 73 73sA_happy5 eleChega 1207 385 68 100 6.1
Appendices ! 291 74 74sA_happy6 estaRoupa 1321 373 93 90 4.7 75 75sA_happy7 osJardins 1628 376 84 90 5.9 76 76sA_happy8 asPessoas 1772 391 85 90 5.7 77 77sA_happy9 haArvores 1574 360 76 75 4.9 78 78sA_happy10 osTigres 1700 385 102 55 4.3 79 79sA_happy11 oQuadro 1461 373 85 95 4.7 80 80sA_happy12 alguemFechou 1702 351 75 90 5.9 81 81sA_happy13 osJovens 1623 355 77 75 4.7 82 82sA_happy14 oFutebol 1647 371 83 90 5.7 83 83sA_happy15 elaViajou 1592 375 83 90 4.6 84 84sB_happy16 oRadio 1348 338 59 50 5.3 85 85sB_happy17 aqueleLivro 1370 311 69 60 5.3 86 86sB_happy18 aTerra 1271 356 81 65 5.6 87 87sB_happy19 oCao 1286 341 79 60 6.1 88 88sB_happy20 eleChega 1178 376 94 65 6.3 89 89sB_happy21 estaRoupa 1402 348 83 60 5.5 90 90sB_happy22 osJardins 1611 346 87 75 5.9 91 91sB_happy23 asPessoas 1660 320 79 75 5.2 92 92sB_happy24 haArvores 1470 345 93 70 5.8 93 93sB_happy25 oQuadro 1451 353 87 45 4.8 94 94sB_happy26 alguemFechou 1551 348 82 90 5.9 95 95sB_happy27 osJovens 1547 335 76 70 5.4 96 96sB_happy28 oFutebol 1368 332 80 60 5.2 97 97sB_happy29 elaViajou 1296 321 70 60 5.4 98 98sA_neutral1 estaMesa 1459 210 33 100 4.9 99 99sA_neutral2 oRadio 1280 213 54 95 4.7 100 100sA_neutral3 aqueleLivro 1536 208 51 100 4.4 101 101sA_neutral4 aTerra 1555 204 29 90 5.5 102 102sA_neutral5 oCao 1433 196 32 95 5.2 103 103sA_neutral6 eleChega 1088 203 25 90 4.9 104 104sA_neutral7 estaRoupa 1483 205 33 95 4.8 105 105sA_neutral8 osJardins 1575 239 79 80 4.6 106 106sA_neutral9 asPessoas 1792 217 56 95 5.3 107 107sA_neutral10 haArvores 1481 205 33 90 4.3 108 108sA_neutral11 osTigres 1842 196 28 100 4.6 109 109sA_neutral12 oQuadro 1438 194 26 95 4.7 110 110sA_neutral13 alguemFechou 1576 217 71 75 4.4 111 111sA_neutral14 osJovens 1795 208 70 100 5.5 112 112sA_neutral15 oFutebol 1838 223 84 90 5 113 113sA_neutral16 elaViajou 1358 191 21 95 5.2 114 114sB_neutral17 estaMesa 1609 215 24 100 5 115 115sB_neutral18 aqueleLivro 1282 221 40 85 4.5 116 116sB_neutral19 aTerra 1379 205 19 85 5.1 117 117sB_neutral20 oCao 1190 200 20 80 4.5 118 118sB_neutral21 eleChega 952 209 20 100 4.6 119 119sB_neutral22 estaRoupa 1294 218 23 90 4.7 120 120sB_neutral23 asPessoas 1671 215 20 70 5 121 121sB_neutral24 haArvores 1400 208 16 80 4.6 122 122sB_neutral25 osTigres 1682 209 54 80 4.1 123 123sB_neutral26 oQuadro 1361 221 53 70 3.3 124 124sB_neutral27 alguemFechou 1631 205 17 80 4.9
Appendices ! 298
Appendices ! 299 7. Unbiased accuracy rates for each prosodic emotion in Chapter VIII Hu for each prosodic emotion, as a function of age and expertise groups. Values vary between 0 and 1. Standard errors are presented in parentheses. Age group / emotion Hu scores Controls Musicians Younger Anger .6 (0.06) .8 (0.04) Disgust .5 (0.07) .6 (0.06) Fear .7 (0.06) .8 (0.04) Happy .5 (0.06) .6 (0.05) Sadness .7 (0.04) .8 (0.03) Surprise .6 (0.05) .6 (0.03) Neutrality .6 (0.05) .7 (0.04) Middle-aged Anger .4 (0.05) .6 (0.06) Disgust .2 (0.04) .3 (0.05) Fear .4 (0.06) .6 (0.07) Happy .3 (0.03) .4 (0.04) Sadness .7 (0.04) .6 (0.05) Surprise .4 (0.04) .5 (0.03) Neutrality .5 (0.05) .6 (0.04)
Appendices ! 300
Appendices ! 301 8. Intensity ratings for each prosodic emotion in Chapter VIII Intensity ratings (minimum 1, maximum 7) for each prosodic emotion, as a function of age and expertise groups. Standard errors are presented in parentheses. Age group / emotion Intensity ratings Controls Musicians Younger Anger 5.6 (0.2) 5.1 (0.2) Disgust 5.6 (0.2) 4.4 (0.3) Fear 5.8 (0.2) 4.9 (0.3) Happy 5.4 (0.3) 5.4 (0.2) Sadness 5.7 (0.3) 4.8 (0.4) Surprise 5.8 (0.3) 5.3 (0.2) Neutrality 5.6 (0.4) 5.2 (0.3) Middle-aged Anger 5.2 (0.3) 5.1 (0.3) Disgust 5.1 (0.3) 4.6 (0.3) Fear 5.2 (0.3) 5.1 (0.4) Happy 5.3 (0.2) 5.5 (0.2) Sadness 5.4 (0.3) 5.2 (0.3) Surprise 5.6 (0.2) 5.4 (0.2) Neutrality 5.2 (0.4) 4.9 (0.3)
Appendices ! 302
Appendices ! 303 9. RTs for each prosodic emotion in Chapter VIII RTs for each prosodic emotion, as a function of age and expertise groups. Standard errors are presented in parentheses. Age group / emotion RTs Controls Musicians Younger Anger 3,074 (106) 3,419 (185) Disgust 3,131 (117) 3,764 (189) Fear 3,111 (122) 4,509 (173) Happy 3,078 (74) 3,411 (187) Sadness 2,936 (100) 3,512 (229) Surprise 3,006 (107) 3,247 (171) Neutrality 2,960 (111) 3,390 (195) Middle-aged Anger 4,141 (218) 4,171 (251) Disgust 4,243 (273) 4,498 (328) Fear 4,171 (185) 4,197 (211) Happy 3,956 (277) 4,194 (262) Sadness 3,758 (280) 4,201 (221) Surprise 3,768 (182) 3,991 (147) Neutrality 3,963 (182) 3,861 (239)
Appendices ! 304
Appendices ! 305 10. Acoustic characteristics of the prosodic stimuli used in Chapter VIII Detailed acoustic characteristics of the prosodic stimuli. Note: Dur = Duration (ms); F0 mean = fundamental frequency mean (Hz); F0 var. = fundamental frequency variability (SD, Hz); Int. = mean intensity (dB); Int. var. = intensity variability (SD, dB); F0 min. = fundamental frequency minimum (Hz); F0 max. = fundamental frequency maximum (Hz); Jit. = Jitter (%); P.p. = Pause proportion (ratio between portions of silence and total duration); HFE = high-frequency energy (relative energy above vs. below 500 Hz). # Stimulus Dur F0 mean F0 var. Int. Int. var. F0 min. F0 max. Jit. P.p. HFE 1 4sA_angry4 1306 355 68 78 16 210 447 1.55 .31 4.50 2 5sA_angry5 1166 299 99 76 11 176 469 1.55 .28 2.17 3 7sA_angry7 1224 349 85 80 9 169 473 1.28 .11 0.47 4 14sA_angry14 1477 338 80 76 12 174 524 1.30 .35 1.33 5 16sA_angry16 1323 329 93 77 11 163 525 0.87 .21 2.43 6 18sB_angry18 1079 289 56 79 13 188 381 1.62 .21 4.33 7 24sB_angry24 1160 283 47 80 8 204 382 2.40 .10 1.24 8 26sB_angry26 1202 293 61 79 9 178 448 2.24 .30 3.38 9 28sB_angry28 1200 262 60 77 14 115 332 2.03 .37 4.00 10 29sB_angry29 1296 305 52 77 8 190 379 2.06 .05 4.00 11 33sA_disgust1 1661 270 76 76 8 171 469 2.10 .13 1.18 12 35sA_disgust3 1596 289 76 79 10 182 523 1.86 .10 3.20 13 36sA_disgust4 1908 279 77 76 8 191 507 1.83 .22 1.63 14 37sA_disgust5 1800 286 78 78 13 203 494 1.42 .26 4.13 15 38sA_disgust6 1852 281 67 78 9 195 495 0.92 .17 2.00 16 39sB_disgust7 1376 272 62 79 8 150 373 2.44 .14 0.92 17 40sB_disgust8 1317 266 64 78 11 90 333 2.12 .33 0.08 18 41sB_disgust9 1875 289 75 78 12 80 410 1.99 .20 0.26 19 43sB_disgust11 1530 287 57 76 11 151 399 3.46 .34 0.47 20 44sB_disgust12 1608 278 51 74 12 197 395 2.25 .61 0.25 21 46sA_fear2 1501 336 45 80 9 224 403 2.84 .13 0.22 22 50sA_fear6 1772 310 45 78 9 218 433 3.24 .12 0.06 23 51sA_fear7 1954 291 38 78 8 212 371 2.21 .10 0.17 24 53sA_fear9 2009 309 34 78 11 213 452 1.84 .20 0.12
Appendices ! 306 25 54sA_fear10 1759 290 58 79 7 77 491 2.34 .07 0.09 26 57sB_fear13 1086 300 28 78 10 224 328 1.38 .22 0.08 27 59sB_fear15 1125 308 33 77 12 201 325 3.21 .35 0.13 28 63sB_fear19 1212 325 45 79 10 234 387 1.79 .16 0.09 29 65sB_fear21 1400 305 41 79 9 213 359 1.94 .10 0.70 30 68sB_fear24 1300 314 41 77 11 269 482 4.28 .33 0.13 31 69sA_happy1 1451 360 83 80 7 189 435 1.08 .06 0.53 32 71sA_happy3 1588 379 102 77 9 188 491 0.92 .15 0.50 33 72sA_happy4 1368 377 98 80 8 187 499 0.96 .18 1.53 34 75sA_happy7 1628 376 84 79 8 192 495 1.55 .18 0.19 35 82sA_happy14 1647 371 83 76 13 199 508 1.41 .38 1.50 36 91sB_happy23 1660 320 79 78 13 181 481 1.74 .20 0.72 37 92sB_happy24 1470 345 93 79 10 81 525 1.92 .07 2.00 38 94sB_happy26 1551 348 82 77 9 176 464 1.16 .13 5.00 39 95sB_happy27 1547 335 76 77 10 178 489 1.58 .17 0.82 40 97sB_happy29 1296 321 70 78 9 170 479 1.69 .16 1.13 41 101sA_neutral4 1555 204 29 79 15 135 300 2.80 .29 0.36 42 103sA_neutral6 1088 203 25 80 5 153 284 1.93 .00 0.16 43 105sA_neutral8 1575 239 79 78 7 152 477 2.83 .09 0.23 44 108sA_neutral11 1842 196 28 78 7 145 262 3.87 .04 0.27 45 112sA_neutral15 1838 223 84 79 11 149 485 3.15 .22 0.58 46 119sB_neutral22 1294 218 23 79 10 175 262 2.98 .13 0.25 47 114sB_neutral17 1609 215 24 79 8 166 258 1.11 .05 0.37 48 115sB_neutral18 1282 221 40 80 9 175 494 1.30 .14 0.79 49 121sB_neutral24 1400 208 16 79 8 172 243 2.55 .22 1.18 50 124sB_neutral27 1631 205 17 78 7 175 233 1.54 .05 1.00 51 128sA_sad1 1756 197 22 76 8 166 243 2.09 .16 0.44 52 137sA_sad10 1996 178 26 75 6 149 251 3.36 .12 0.05 53 138sA_sad11 1599 184 25 76 12 148 253 1.93 .35 0.92 54 140sA_sad13 1760 181 20 77 8 149 233 1.81 .11 0.32 55 142sA_sad15 1555 171 17 77 7 147 238 2.03 .11 0.33 56 144sB_sad17 1294 218 55 77 10 170 509 2.39 .22 1.14 57 143sB_sad16 1367 216 30 79 9 178 272 1.37 .08 0.14 58 146sB_sad19 1278 205 24 78 12 169 247 2.68 .37 0.19 59 149sB_sad22 1101 208 34 78 10 179 265 2.98 .23 0.53 60 150sB_sad23 1491 244 74 77 10 182 499 2.34 .30 0.24 61 160sA_surprise2 1340 314 92 79 9 168 528 1.52 .09 0.86 62 163sA_surprise5 1360 291 66 78 11 204 464 1.24 .21 1.80 63 164sA_surprise6 1222 321 99 80 6 212 522 1.21 .10 0.72 64 168sA_surprise10 1586 295 59 77 11 218 516 2.43 .19 4.50 65 172sA_surprise14 1652 329 110 79 8 152 493 1.91 .11 0.32 66 175sB_surprise17 1243 319 44 78 9 239 406 1.85 .10 0.59 67 178sB_surprise20 1187 343 61 79 12 226 440 1.78 .21 0.50 68 182sB_surprise24 1458 360 64 78 8 236 464 1.53 .08 0.36 69 185sB_surprise27 1541 385 51 79 9 281 499 1.24 .14 1.41 70 187sB_surprise29 1432 393 65 79 7 268 486 1.07 .05 0.90
Appendices ! 307 11. RTs for prosody, music and facial expressions in Chapter IX RTs for emotion recognition in speech prosody, music and facial expressions, for PD patients and healthy controls. Standard errors are presented in parentheses. Modality / emotion RTs Patients Controls Music Happiness 4,482 (250) 3,706 (253) Peacefulness 5,972 (367) 5,374 (349) Sadness 6,169 (423) 5,505 (334) Fear 6,053 (407) 5,040 (359) Speech prosody Happiness 3,649 (357) 2,878 (184) Surprise 3,992 (331) 3,193 (202) Sadness 3,938 (350) 3,047 (225) Fear 4,332 (316) 3,574 (269) Facial expressions Happiness 2,694 (206) 2,125 (165) Surprise 3,437 (346) 3,094 (277) Sadness 3,654 (382) 2,822 (219) Fear 3,792 (368) 3,401 (313)