Full text
FACULTAD DE FILOLOGÍA DEPARTAMENTO DE LENGUA INGLESA THE EFFECT OF L1 DIALECT ON THE PERCEPTION OF PHONETIC VARIATION IN L2 MORPHOLOGICAL MARKERS TESIS DOCTORAL María Del Saz Caracuel Dirigida por: Profª. Dra. María Heliodora Cuenca Villarín Programa de Doctorado: Lengua y Lingüística Inglesas Sevilla, 2013
ii “Long is the way And hard, that out of Hell leads up to Light” (Paradise Lost, John Milton)
iii Abstract Second language (L2) speech perception depends on a number of factors related to the characteristics of the listener and to the characteristics of the stimuli employed in the experimental task. This doctoral dissertation explores the role that the dialect of the listeners (African American English and General American English) plays in the perception of two dialect variants of the same L2 (Western Andalusian Spanish and Castilian Spanish) in the morphological marker –s. Both General American English and Castilian Spanish use –s to mark the plurality of nouns and the person of verbs. African American English makes other uses of this marker, whereas Western Andalusian Spanish aspirates it. Our initial hypothesis proposes that the Castilian variant will be better identified than the Andalusian variant in general, although to a greater extent by General American English listeners. For this purpose, an identification task was designed with random sentences in which second-person and third-person verbs, as well as plural and singular nouns were embedded. Results corroborate our hypothesis and indicate that i) L2 proficiency level influences the perception of Andalusian aspiration; ii) the listener’s dialect influences the perception of Castilian sibilants, iii) the perception of both variants depends on the phonetic context of the stimuli. A subsequent acoustic analysis of the stimuli reveals that there are intrinsic characteristics in both L2 dialects that can explain these results, especially as far as fricatives and stops are concerned. As future investigation, attention to (inter)dental contexts is suggested, as they present the most acute results.
iv Resumen La percepción del habla de una segunda lengua (L2) depende de un número de factores relacionados con las características del oyente y con las características de los estímulos empleados en la prueba experimental. Esta tesis doctoral explora el papel que juega el dialecto de los oyentes (inglés afroamericano e inglés americano general) en la percepción de dos variedades dialectales de una misma L2 (español andaluz occidental y español castellano) en el marcador morfológico –s. Tanto el inglés americano general como el español castellano utilizan –s para marcar la pluralidad en los sustantivos y la persona de los verbos. El inglés afroamericano hace otros usos de este marcador, mientras que el español andaluz occidental lo aspira. Nuestra hipótesis de partida propone que la variante castellana será mejor identificada que la andaluza en general, aunque en mayor medida por parte de los oyentes de inglés americano general. Para ello se diseñó una prueba de identificación con frases aleatorias en la que se encontraban verbos de segunda y tercera persona así como nombres en plural y singular. Los resultados corroboran nuestra hipótesis e indican que i) el nivel de competencia en L2 influye la percepción de la aspiración andaluza, ii) el dialecto del oyente influye en la percepción de los sibilantes castellanos, iii) la percepción de ambas variantes depende del contexto fonético de los estímulos. Un posterior análisis acústico de los estímulos revela que existen características intrínsecas en los dos dialectos de L2 que pueden explicar estos resultados, especialmente en cuanto a fricativas y oclusivas se refiere. Como investigación futura, se sugiere prestar atención a los contextos (inter)dentales, ya que presentan los resultados más acusados.
v Acknowledgments I would like to thank the many people that have definitely helped me through this journey. First and foremost, my gratitude goes to my doctoral advisor, Dr. María Heliodora Cuenca Villarín, for her invaluable guidance, encouragement, and support through this ordeal. I would also like to thank Dr. J. Tamayo, Dr. J. Comesaña and Dr. F. Garrudo for their help in making this process continue in the right direction. I am also grateful to M. A. Lamprea, K. Goldsmith, J. Sánchez Torres, M. J. Daily, M. Buchner, Dr. J. A. Bravo-de-Rueda, Dr. G. Cowell, E. Cruz Jiménez, C. Carrasco Llopis, Dr. M. Tubino Blanco, Dr. A. Ascunce, Dr. T. Lorenz, Dr. C. Francom and Dr. I. Ortega-Santos for their help with participant recruitment, and to Dr. L.B. Schmidt for allowing me to know more about her work. Without them and without the participants and the speakers, this dissertation would have not been possible. I cannot fail to acknowledge the efficient and prompt work of P. Anaya, the computer technician who mounted the experimental task, as well as the invaluable help and expertise of M. Barrio in all matters related to Acoustic Phonetics. And last, but not least, I want to thank my family and friends for their patience, support, and understanding; especially to Dad and Nacho for those comments that gave me an outside and many times enlightening perspective on this project.
vi Table of Contents List of Tables ix List of Figures xi Introduction 1 Chapter 1. Description of L1 and L2 Dialects 5 1.1 L1 Dialect: African American English 5 1.1.1 The Phonology of African American English 9 1.2 L2 Dialect: Andalusian Spanish 20 1.2.1 The Phonology of Andalusian Spanish 21 1.3 The Morphological Marker –s 34 1.3.1 Description of [h] and [s] 36 1.4 Summary 41 Chapter 2. Acoustic Phonetics 42 2.1 Acoustic Cues 44 2.1.1 Stops 47 2.1.2 Fricatives 56 2.2 Summary 67 Chapter 3. L2 Speech Perception 68 3.1 Perception of L2 Dialect Variants 74 3.2 Models of L2 Speech Perception 79 3.2.1 Native Language Magnet Model 79 3.2.2 Speech Learning Model 81 3.2.3 Perceptual Assimilation Models 83 3.2.4 Automatic Selective Perception Model 87 3.3 Current Study 90 3.3.1 Predictions and Research Questions 91 3.4 Summary 93 Chapter 4. Method 95 4.1 Participants 95 4.1.1 Speakers 95 4.1.2 Listeners 96 4.2 Materials 98 4.2.1 Recording 101 4.3 Procedure 101 4.3.1 Language Background Questionnaire 105
vii 4.3.2 Spoken English Questionnaire 105 4.4 Data Coding 106 4.5 Statistical Analysis 108 4.6 Acoustic Analysis 109 4.6.1 Voiceless Stops 109 4.6.2 Fricatives 112 4.7 Rationale for the Identification Task 114 4.7.1 Pilot Study 115 Chapter 5. Results 117 5.1 Results of Perception 117 5.1.1 Pilot Study 117 5.1.2 Current Study 124 5.2 Results of Acoustic Analysis 160 5.2.1 Voiceless Stops 163 5.2.2 Fricatives 167 5.2.3 Speaker Gender 174 5.3 Summary 178 Chapter 6. Discussion and Conclusions 180 6.1 Perception of Aspiration and Sibilance 181 6.1.1 The Effect of L2 Proficiency and Instruction 183 6.1.2 The Effect of Syntactic Context 186 6.1.3 The Effect of Phonetic Context 188 6.1.4 The Effect of the Acoustic Characteristics of the Stimuli 191 6.2 Implications for L2 Speech Perception 195 6.3 Limitations and Future Directions 201 6.4 Conclusions 203 References 206 Appendix A. Grammatical Features of AAE 223 Appendix B. Grammatical Features of AS 232 Appendix C. List of Sentences 235 Appendix D. Informed Consent Form 239 Appendix E. Language Background Questionnaire 240 Appendix F. Spoken English Questionnaire 242
viii Appendix G. Waveforms and Spectrograms of WAS Aspirated Stops 244 Appendix H. Waveforms and Spectrograms of CS Unaspirated Stops 247 Appendix I. Waveforms and Spectrograms of WAS Fricatized Stops 250 Appendix J. Waveforms and Spectrograms of CS Sibilants Before Approximants 253 Appendix K. Spectral Slices of WAS Fricatives 256 Appendix L. Spectral Slices of CS Sibilants 258
ix List of Tables Table 1. Phonetic inventory of seseo, ceceo and distinction 28 Table 2. Classification of English and Spanish stops by VOT category 49 Table 3. Dialectal allophones of Spanish consonants in word-initial position after [s] and [h] 100 Table 4. Percentage of AAE usage reported by AAE listeners 106 Table 5. Mean accuracy percentages for all groups and variables 118 Table 6. Overall identification percentages of aspiration and sibilance in the three noise conditions 121 Table 7. Perception of aspiration and sibilance by years of instruction and listener group 134 Table 8. Identification of WAS and CS third-person verbs and singular nouns 140 Table 9. Reaction times (ms) for the identification of both conditions in the four syntactic contexts by group of listeners 140 Table 10. Wilcoxon test and statistical probability values 141 Table 11. Perception of aspiration in phonetic context by listener group 142 Table 12. Perception of sibilance in phonetic context by listener group 143 Table 13. Mean values for VOT (in ms) after aspiration and sibilance 163 Table 14. Mean values for closure duration (in ms) after aspiration and sibilance 163 Table 15. Vowel length with and without aspiration before voiceless stops 166 Table 16. Spectral moments of intervocalic [s] and [h] 167 Table 17. Spectral moments of [s] before voiceless stops 168 Table 18. Spectral moments of [z] before voiced approximants 169 Table 19. Spectral moments of WAS fricatized sounds 170 Table 20. Vowel length with and without aspiration before WAS fricatized sounds 171
3 We chose this type of task because it resembles what native speakers of a language do when decoding the acoustic signal they receive from continuous L1 speech, requiring the identification and categorization of phonetic segments according to their internalized language-specific categories to access meaning (Hawkins, 2011; Strange & Shafer, 2008). Studies have shown that experiments that address basic auditory capabilities and trigger language-general patterns of perception yield similar results for native, naïve, and experienced L2 listeners. As the cognitive demands of the task and the stimuli increase, language-specific patterns of perception are more likely to be reflected. We believe this may be especially relevant for elementary students, which have a more limited experience with the L2; thus, we also explored the proficiency level of the listeners when interpreting the results of the tests. The organization of the rest of this dissertation is as follows: Chapter 1 provides a detailed description of the phonology of AAE and WAS, and a specific description of [h] and [s] in Spanish and English. Chapter 2 is devoted to acoustic phonetics and, in particular, to the description of English and Spanish sounds in terms of their acoustic properties. Chapter 3 reviews the background literature pertinent to this study and poses predictions for the two groups of listeners as well as our research questions. Chapter 4 provides an account of the methods employed in the experiments, a description of the stimuli used, and a report of the participants who took part in this study. Furthermore, descriptions of the experimental task as well as the statistical and acoustic analyses carried out are also provided. In Chapter 5, we report the main findings obtained from the experiments; first, the results of the pilot test and, second, the results of the current experiment in terms of overall performance, identification patterns by individual group, and identification patterns across groups of listeners. Additionally, we analyze the results according to the syntactic and the
4 phonetic contexts in which stimuli are found, and we also provide an acoustic analysis of key stimuli to incorporate these findings to our discussion section. Finally, in Chapter 6, we discuss these findings, the implications of the results for L2 speech perception theory, and the limitations of this study with suggestions for future research.
5 CHAPTER 1 DESCRIPTION OF L1 AND L2 DIALECTS In order to understand how AAE and WAS differ from mainstream GAE and CS, respectively, this chapter provides a detailed description of the phonological systems of both dialects. Additionally, complementary information concerning grammar use can be found in Appendices A and B. 1.1 L1 Dialect: African American English The term AAE is generally employed to refer to the language varieties that African American people speak in the United States. However, as Baugh (2004b) points out, African Americans can fall into any of these three categories: a) GAE is their native language, b) GAE is not their native language, c) their native language is different than English. Most speakers of AAE belong to the second category, with the ability to switch from GAE to nonstandard English or AAE depending on the context and their interlocutors. We should also take into account that not all African Americans generally speak AAE nor are all speakers of AAE African Americans. As Green (2004a) describes it: African American English refers to a linguistic system of communication governed by well defined rules and used by some African Americans (though not all) across different geographical regions of the USA and across a full range of age groups. While AAE shares many features with mainstream varieties and other varieties of English, it also differs from them in systematic ways. (p. 77)
6 Ever since the early studies in the 1960s, AAE has seen a number of names applied to it, depending on the term employed to address African Americans at the time. “Negro dialect” and “Negro speech” were usual in the 1960s, “Black talk”, “Black dialect”, “Black English” and “Black Vernacular English” were popular in the 1970s, which then turned into “African American language” in the mid-1980s. “African American English” 3 became the preferred term in the 1990s, along with “African American Vernacular English” to design the nonstandard form of AAE. The term “Ebonics” (a blend of ebony and phonics) was initially employed to refer to the speech of those African Americans of West African descent but subsequently used as equivalent to those terms above, especially AAE. In any case, “the reality, however, is that most speakers of what is identified here as AAE do not have a name for their vernacular. Generally they say they speak English” (Mufwene, 2001, p.293). AAE has been the subject of much controversy, especially in education, since the Oakland School Board resolution in 1996 and subsequent resolution by the Linguistic Society of America in 1997, which recognized Ebonics as a systematic and rule-governed linguistic system, not related to English, and the primary language of African Americans. This resolution aimed at improving their proficiency in GAE and thus broadening their academic and professional future. However, there is still no full agreement as to whether it is considered a dialect of English or a separate language. On the one hand, apart from having its own distinctive characteristics, AAE shares the vast majority of its features and patterns with GAE, which would support the first position. One of the authors who support the term dialect is Dillard (1993), stating that it is “the first clearly discernible and reportable dialect of American English” (p. 60). On the other hand, AAE involves 3 There is still a difference of opinions among AAE speakers. Some prefer the term Black, “we have been here too long […] By now we have no African in us”. Others prefer to use African American to highlight “our origin and cultural identity.” (Smitherman, 1998, pp. 206-7)
7 sociological and ethnic connotations, as well as a unique background and development, which would account for the second position. This view is shared by authors such as Smitherman (1999), who states that “it is a language forged in the crucible of enslavement, US-style apartheid, and the struggle to survive and thrive in the face of domination” (p. 19). A third position invalidated by linguists and experts but still present in society, even among its own speakers, is that AAE is simply “bad English”. The origin of AAE is also a controversial issue, giving rise to three main views. The Anglicist hypothesis emerged in the mid-20 th century and defends that AAE originated from the various dialects of English that white immigrants from the British Isles spoke at the time. Later, towards 1970s, the Creolist hypothesis appeared in defense of the view that AAE may have started as a creole language such as Gullah or Jamaican Creole, with which it shares features, influenced by the languages of the slaves brought from other colonies. Therefore, contact with other dialects in the USA would have originated a slow process of decreolization, by which AAE is converging with other varieties of English. Finally, the Africanist hypothesis defends that AAE is similar to West African languages in structure and regards any similarity to English as only superficial. Even when it may have incorporated English features, the substrate influence of West African languages is still preserved. Other than these major views, there is a new one named Neo-Anglicist hypothesis that also believes that AAE originated from British dialects but has undergone a unique evolution that has made it diverge from GAE. We may never know how AAE exactly originated, given the scarce recordings and data of which linguists dispose. As Wolfram (2006) states:
8 Current evidence suggests more regional influence from English speakers than assumed under the Creolist Hypothesis and more durable effects from early language contact situations than assumed under the Anglicist positions, but the issue of regional accommodation and substrate influence continues to be debated. (p. 335) The term African American Vernacular English (AAVE) must be distinguished as the vernacular or nonstandard form of AAE which carries more stigmatized aspects, used mainly for everyday communication among its speakers. It is generally attributed to the working class, although the middle class can also use it depending on the context, e.g., informal situations, adding emphasis, expressing ethnic solidarity, etc. While it is true that it shares features with creole languages and Southern White Vernacular English (SWVE), AAVE is still a systematic and rulegoverned linguistic system with defined aspectual grammar, vocabulary of its own, distinctive phonology, and unique intonation which divert from GAE. At the same time, AAVE should not be regarded as mere slang, as slang refers to temporary vocabulary and expressions which grow out of fashion and are replaced by others with time. AAVE features are long-established and common throughout the country. Nowadays, research shows a tendency for different trajectories in the development of AAVE according to geographical and other sociological factors. There are instances of assimilation of the regional variety of English and reduction of AAVE characteristics, as well as instances in which AAVE characteristics are reinforced and resistance to the regional variety of English takes place. “Original settlement history, community size, local and extra-local social networks, and racial ideologies in American society must all be considered in understanding the course of change in African American speech” (Wolfram, 2006, p. 340).
9 1.1.1 The Phonology of African American English At first glance, the most distinctive features of AAE seem to lie in its morphology and syntax 4 . This has led to a great amount of research directed towards the origins of this variety and its implications for education. “In many ways phonology is the neglected stepchild of research on AAE. Even the most cursory review of the literature will show that morphology and syntax have long been the primary focus of work on AAE” (Bailey & Thomas, 1998, p. 85). The phonology of AAE presents different types of variables, the majority of which are systematic and context-dependent. This does not imply that all African Americans always use all of them; there is variation among these variables that are the most salient patterns in AAE. 1.1.1.1 Consonant clusters Consonant cluster reduction, especially when the second consonant is a stop, is well-known among AAE speakers. Even when this feature is common to other varieties of English, certain constraints in which it occurs seem to differ. The reduction is generally more likely to take place when the following word begins with a consonant sound (fast car) than when it begins with a vowel sound (cold air) or when the second consonant in the cluster is a morpheme, such as past tense –ed (talked). These phonological and grammatical constraints are found in all varieties of English. What is interesting is that phonological constraints seem to dictate cluster reduction in GAE, while AAE is more driven towards respecting grammatical constraints. In other words, AAE speakers are less likely to simplify the cluster when 4 For a description of the most salient grammatical features, see Appendix A.
10 it represents a morpheme, whether followed by a consonant or a vowel sound. However, there seems to be an exception to this rule and this is the case of irregular part tense verbs (kept), which are more likely to suffer reduction than regular past tense verbs, probably because the tense is additionally marked by a change in the vowel sound. Since utterances do not occur in isolation but within a context, it is possible for speakers to employ cluster reduction in past tense verbs when there are other clues of time in the sentence (Yesterday, she call me three times). The number one rule for consonant cluster reduction is that both consonant sounds must share voicing (fast, kind), but in spite of this rule we can also find an exception, and this is negative auxiliary verbs. It is common to hear can’t realized as [ˈkeɪn] and don’t uttered as [ˈdoʊn]. Wolfram and Thomas (2002, pp. 133-4) enumerate a list of constraints that affect the frequency of this type of reduction. First, simplification is less likely when both consonants are stops (pact) than when the first one is a sibilant (past). Second, it is less likely when the first one is a sibilant (past) than when the first one is /l/ (bold). Third, it is less likely when the first one is /l/ (bold) than when the first one is nasal (kind). And fourth, consonant cluster reduction is more commonly found in unstressed syllables than in stressed syllables. There seems to be opposing views about the origins of consonant cluster reduction. On the one hand, the reduction is believed to be a process that occurs according to the phonological context in which the cluster is given, as is the case in other varieties of English such as nonstandard British accents. On the other hand, this feature is attributed to the influence of West African languages, which do not allow final consonant clusters (Green, 2002b). In fact, there are speakers who actually do not seem to have a cognitive representation of the cluster; therefore, it is possible to
11 find the plural form of these reduced words as if indeed the cluster never existed, giving way to test > [ˈtes], the plural form of which would be [ˈtesəz], as in buses (Green, 2002a, 2002b; Mufwene, 2001). Furthermore, a more common plural form would be realized by lengthening the continuant as in tests > [ˈtes:] (Thomas, 2007). Additional features involving consonant clusters are metathesis and the backing of /str/ clusters. Metathesis consists of switching the position of the consonants in the cluster, whose main representative example is ask > [ˈæks]. Backing of /str/ cluster means that the cluster is realized as [skr], especially before high front vowels, as in street > [ˈskri:t]. 1.1.1.2 Fricatives A second feature attributed to AAE that is one of its most representative characteristics involves the absence of interdental fricatives /θ/ and /ð/, which are either labialized or stopped. The former phoneme is usually replaced by [t] in initial position (think > [ˈtɪŋk]) and final position (month > [ˈmʌnt]) or by [f] in final position (both > [ˈboʊf]), while the latter phoneme is replaced by [d] in initial position (this > [ˈdɪs]) and by [v] in medial position (mother > [ˈmʌvə]) and final position (bathe > [ˈbeɪv]). This feature is also found in other nonstandard varieties of English; nevertheless, it is much more commonly found in AAE and inversely correlated with social class and formality of speaking style. The constraints on this feature are somewhat unclear, as we can find the word with uttered in all four different ways: [ˈwɪt], [ˈwɪd], [ˈwɪf], and [ˈwɪv], mostly depending on the voicing of the following sound (Bailey & Thomas, 1998). One point should be made here: AAE speakers know how to realize interdental fricatives, as in thing > [ˈθɪŋ]. The alternative realizations are seen by Africanists as a West African
12 influence, whose languages do not include /θ/ or /ð/; by contrast, Anglicists claim that nonstandard dialects of British English also included initial [d] and final [f] for /θ/. Likewise, Creolists state that /θ/ realized as [t] or [d] is also a feature in creole and pidgin languages (Rickford & Rickford, 2000). Another case of fricative stopping takes place especially in medial position before nasal sounds. In this instance, it is the substitution of [b] for /v/, as in seven > [ˈsebm], and the substitution of [d] for /s/, as we can see in isn’t > [ˈidnt] (Rickford, 1999; Rickford & Rickford, 2000; Bailey, 2001). The first characteristic finds a similar phenomenon in creole languages, which realize /v/ as a bilabial continuant [β] (Lerer, 2007). 1.1.1.3 R-lessness and l-lessness The following features used to be shared by both AAE and SWVE at the beginning of the 20 th century; however, it is reversing for the latter while it seems to persevere with the former. It is the case of what is known as ‘r-lessness’ or nonrhoticity, i.e., the deletion or vocalization of constricted /r/ in any of the following phonetic environments: (1) postvocalic position (four), (2) word-medial position (carry), (3) unstressed syllable (mother), (4) and stressed syllable (work) –although deletion in this last case is mostly restricted to Southern AAE. (1) four > [ˈfoʊ], [ˈfoə], [ˈfo:] (2) carry > [ˈkæi] (3) mother > [ˈmʌvə] (4) work > [ˈwɜ: i k], [ˈwɜ:rk] Mufwene (2001) explains the frequency of occurrence of /r/ deletion or nonrhotic /r/ according to its position. The most frequent cases of deletion take place in
19 We have previously seen how the pronunciation of certain auxiliary verbs are affected in AAE and now we will pay attention to the pronunciation of two aspectual markers 5 of AAE, which share their forms with two GAE words but which differ in pronunciation (and, of course, in meaning). These are (1) remote been (written BIN due to its stressed pronunciation and not to be confused with GAE been) and (2) completive done (written dən due to its unstressed pronunciation and not to be confused with GAE done). The former is employed to indicate that something occurred a long time ago or that something has been happening for a long time to this day. The latter refers to the completion of an action with present results. (1) She BIN ate all the candy. GAE: She ate all the candy a long time ago. (2) I dən did my homework. GAE: I have done my homework. Rickford (1975) conducted an experiment in which one half of the informants were AAE speakers and the other half were GAE speakers. He presented them with different sentences in which these two aspectual markers were present and tested their understanding of their meanings. All AAE speakers obtained correct answers while only one GAE speaker answered all questions correctly. This example is to give an understanding of how this variety is ruled-governed and forms a well-designed system in some aspects different than GAE. As Rickford & Rickford (2000) put it, “these processes are highly systematic, and not the careless or haphazard pronunciations that observers often mistake them for” (p. 104). 5 See Appendix A
20 1.2 L2 Dialect: Andalusian Spanish Andalusian Spanish (AS) is the variety of the Spanish language spoken by people from and in the province of Andalusia, the southern region of Spain. It appeared as the result of the changes in the Medieval Castilian taken to the region with the Reconquest by Ferdinand III The Saint in the 13 th century. Record of these changes dates back to the 15 th and 16 th centuries and indicates that the variety was consolidated in the 18 th century. Whether it is a dialect or a variety of speech still remains undetermined; diachronically, it is a dialect which evolved from the historical Castilian brought to the region by settlers and colonizers around the 13 th century; synchronically, it is a linguistic variety of Spanish as other regional varieties are, which took form from elements of other dialects in the Iberian Peninsula and the influence of foreign languages. Additionally, there is a minority of researchers who defend that the origin of AS is not entirely Castilian but a pidgin language with a Castilian-based lexicon and morphosyntax combined Mozarabic features, a view to which Narbona, Cano, and Morillo (1998), Jiménez Fernández (1999), and Cano Aguilar and González Cantos (2000) oppose. Finally, there are a number of claims that AS is simply “bad Spanish”, even among its own speakers. In this study, we will abide by Alvar’s (2006) definition: Precisamente, diferencias e historia me hacen ver el andaluz como un dialecto y no aceptar que me digan que la «manera de hablar» una lengua es –así, sin más- “el sentido vulgar del término [dialecto], no el técnico” ... pues buen cuidado he tenido siempre en no confundir «la comprensión de un habla y el metalenguaje de una ciencia.» (p. 13) [Precisely, differences and history make me see Andalusian as a dialect and not accept to be told that “the way of speaking” a language is –just like that- “the vulgar sense of the term [dialect], not the technical” … since I have always been very careful not to mistake «the understanding of speech and the metalanguage of a science.»]
21 In Europe, unlike in America, it is usual to find dialects which are contemporary and even more ancient than the standard language of a country; therefore, its divergence from the norm must not be seen as simplifications of the standard language. In this case, the variation which presents more prestige and is used as the norm is northern Spanish, pushing southern Spanish to the background with a different social acceptance. The Castilian spoken in the Reign of Toledo rose as the standard variety of the language; thus termed Spanish, which was spread to Europe. The variety spoken in the Reign of Seville (Seville, Huelva, Cadiz), subsequently named Andalusian, was the norm transferred to the Canary Islands and South America. 1.2.1 The Phonology of Andalusian Spanish As we will see, none of the phonetic and phonological characteristics of AS is common to all speakers in the area nor are they all exclusive of Andalusia. AS also presents a faster and more varied rhythm than CS and some of its phonemes are realized in a more lax way, while others are uttered in a more tense way, than CS. In this section, we will see some characteristics which are spread all over the region, other characteristics that are less spread but still present, and some characteristics that are also found in other Spanish varieties but very commonly used in AS. 1.2.1.1 Andalusian /s/ One of the best-known features of AS is its /s/ realizations and the linguistic phenomena concerning this phoneme. Spanish /s/ is realized by placing the tip of the tongue in the alveolar region of the mouth with the tongue in a concave position. Andalusian /s/ has several realizations, usually with the tongue in a flat position –as
22 the [s] in Cordovaor in a convex position –as the [ṣ] in Seville. The manner of articulation is dental in these two cases, with the actual blade of the tongue and not the tip touching the teeth. In Andalusia, more than a third of the speakers make a distinction between /s/ and /θ/, about the same number merges these two sounds into a dental [s] (phenomenon called seseo), while less than a third merge both sounds into an interdental [θ] (phenomenon called ceceo), as in poso = pozo, casa = caza. The distinction between these two sounds is seeing a widespread tendency nowadays, particularly among young and educated speakers, partly favored by the media and the more accepted peninsular norm. It is mostly given in northern and eastern regions of Andalusia while, in the rest of the regions, there tends to be a coexistence of seseo and ceceo. Ceceo is considered as low status and is a minority phenomenon due to the presence of the distinction of sounds and seseo in urban areas and in the media. In speakers of low socioeconomic status, ceceo can become heheo, especially in rural areas. This means /s/ and /θ/ are uttered with a retracted position of the tongue in a relaxed and aspirated manner, such as [h]. We can also find ceseo and seceo, especially among those speakers who do not make a distinction between /s/ and /θ/. This means speakers realize those sounds as one or the other without a clear pattern, as in cerveza > [θerβésa] or [serβéθa]. 1.2.1.2 Aspiration In relation to the Andalusian /s/ realizations, the most characteristic feature of AS and the most widespread to other varieties of southern Spanish is the realizations of the syllable-final and word-final /s/ (implosive /s/). Being uttered with less articulatory force, as is the case of all Spanish final consonants, it can either be
23 maintained or derive into aspiration, assimilation to the following consonant, gemination, or deletion. This depends on the context, whether /s/ is placed before consonant, vowel or pause, and on geography and socioeconomic status. Aspiration of implosive /s/ occurs before any consonant sound in all geographical areas and social levels, and it is also characteristic of the varieties in the Canary Islands and South America. Reduplication or gemination of following consonants is the general tendency in informal and spontaneous situations. Before voiced stops /b, d, g/, educated speakers aspirate implosive /s/ without modification of the following consonant sound; however, in vernacular speech aspiration can be transferred to those consonant sounds, turning them into fricatives [f], [v], [θ], [ð] and [x]. Alvar (1996, p. 243) gives examples of the possible realizations and allophones derived from each phoneme: I) –s + b: lah brujah, lab bragah, lav viñah, lo brimbe, muncho fohqueh (= ‘las brujas, las bragas, las viñas, los brimbes, muchos bosques’) 6 . II) –s + d: loh dienteh, buenoð ðía, uno θeoh ( = ‘los dientes, buenos días, unos dedos,’). III) –s + g: lah gatah, log güebo, loj jabilane, la jraná ( = ‘las gatas, los huevos, los gavilanes, las granadas’). 7 Before voiceless stops /p, t, k/, aspiration of implosive /s/ also occurs without modification to the following consonants in careful speech. However, in vernacular speech aspiration derives into reduplication or gemination of following consonants. Speakers in Cordova and certain areas of Granada and Seville can infuse aspiration to /p, t, k/ (Gerfen, 2002; Torreira, 2007a, 2007b, 2012), and the latter sound can also be 6 Jiménez Fernández (1999) also adds labio-dentalization (resbalar > refvalá) to this list. 7 Jiménez Fernández (1999) also indicates complete assimilation (rasgo > ráho) as possible.
24 uttered with aspiration in the whole region, especially before –ié. This feature is omitted in formal situations. Alvar (1996, p. 243) exemplifies: I) loh pieh, doh toroh, lah casah; IIa) lo hp pieh, do ht toroh, la hk casah; IIb) lo p pieh, do t toroh, la k casah; III) lo pieh, do toroh, la casah; A more recent variant, which is the affricate palatal pronunciation of /t/, occurs when aspirated /s/ before this dental phoneme influences its articulation into [ts ]. This phenomenon is given in all areas where final /s/ is aspirated or dropped and it is more frequent among the young population and mid-class to high-class speakers (Moya Corral, 2007; Ruch, 2008, 2010, 2012). Jiménez Fernández (1999) mentions further contexts in which the aspiration of implosive /s/ is involved. Aspiration of /s/ before fricative consonants /f/, /s/ and /θ/ causes the gemination of these sounds with almost complete loss of aspiration. Before ch, ll, y, it is hardly maintained, with complete assimilation to these consonants. In the case of ch, aspiration can lead to fricatization. los llevo > [loɟéβo] más chico > [máʃíko] Before r and rr aspiration is lost, giving way to a complete assimilation, as in las ratas > [laráta]. Furthermore, in the case of l, there can be two solutions: aspiration + consonant gemination or no aspiration + consonant reduplication, as in muslo > [múl.lo] or [múhl.lo]. Implosive /s/ is aspirated before m, n, ñ, such as las niñas > [lahníɲa]. The nasal consonant can also be geminated in the presence of aspiration or, to a lesser extent, without it, as in mismo > [míhm.mo] or [mím.mo].
25 The aspiration of /x/ is found throughout the whole region and all social classes except in Jaen and some areas of Granada and Almeria, where we can find the full realization of velar /x/. We must distinguish between the voiceless [h], given in western Andalusia among educated speakers, and the voiced [ɧ], which is more relaxed and generally found in less educated speakers. A very weak aspiration is also possible but it is considered fairly vulgar. Additionally, there exists a halfway sound between velar /x/ and aspirated [h] which is found in eastern areas neighboring western Andalusia and educated speakers who intend to approach the standard. These two allophones can be represented as [h x ] or [x h ], depending on their approach to either the velar or the aspirated sound. The aspiration of Castilian /x/ can be traced back to the 16 th century and the evolution of sibilants during that period as a transition from Medieval to Modern language. The minimal pair /ʃ/-/ʒ/ merged into /ʃ/ after a process of devoicing, the result of which was forced to retract its place of articulation due to its similarity to /s/, giving way to the velar voiceless fricative /x/ that we know today. However, it was not until 1815 that the orthography reforms began and finally changed the spelling of /x/ from x to j. Nevertheless, in areas where aspiration was kept in words derived from f-initial Latin words, /ʃ/ did not result in /x/ but was confused with the aspiration of initial f, as in the case of Andalusia. The aspiration of h occurs when h is in initial position in words derived from Latin terms beginning with f, as facere > hacer. It is an archaic feature present in the whole western part of Andalusia, virtually in rural areas and uneducated speakers, which lacks in prestige among experts because it only applies to certain words. It is also linked to expressive and informal situations, and has been fixed in specific words and expressions, such as cante jondo (flamenco type of singing). The evolution of f- >
26 [h] > /ø/ is a much debated issue which has not seen a consensus to this day, although some historians link its geographical distribution in Andalusia to the Reconquest of this region. On the one hand, the reconquest and repopulation of Jaen was carried out by the Reign of Castile, which had lost the aspiration of f-initial words; thus this region does not preserve aspiration. On the other hand, the reconquest of Seville and Cordova took place after the unification of the Reigns of Castile and Leon, the latter of which still preserved the aspiration of f-initial words in the 16 th century. This feature spread throughout the whole western region of Andalusia and eastern Granada, the reconquest of which initiated in Seville. 1.2.1.3 Mergers Among the features that can be found in the region to a lesser extent is the merger of the alveolar vibrant /r/ and the alveolar lateral /l/ in syllable-final and wordfinal positions. In the western side of the region, both sounds tend to converge towards [r] and tend to be lost in word-final position. In the eastern part of the province, the sounds tend to converge towards [l], although it is following a decreasing tendency. In isolated areas /r/ can be aspirated as [h] and even dropped giving way to the gemination of the following consonants, usually /n/ and /l/. [r] soldado > [sordáo] [l] cuerpo > [kwélpo] [h] carne > [káhne] The consonant cluster rl can find diverse realizations: a) standard pronunciation, b) aspiration of /r/, c) gemination of /l/, d) complete assimilation, e) palatalization into [ʎ], given in areas where ll and y merge. The cluster rn usually undergoes aspiration of /r/ and germination of /n/, with or without aspiration, as in
27 carne > [káhne] > [káhn.ne] > [kán.ne]. In word-final position, as all final Spanish consonants, they tend to be lax and lose their phonemic opposition following the tendency to keep open syllables CVCV. In those cases where total deletion occurs, it can cause the opening of the previous vowel, as in ver > [bé] or [bé]. The first examples of this merger date back to the 12 th and 13 th centuries in Toledo and frequently given in Andalusian texts in 14 th -17 th centuries, although it seems it did not spread until the 16 th century, testimony of which is also found in America from this century onwards. Another type of merger is called yeísmo. This term refers to the merging of lateral palatal /ʎ/ and fricative palatal /ɟ / (pronunciation of ll and y). It is spread throughout the whole Spanish community except in some areas in Huelva and Seville, much rare in Cadiz, Malaga, and Almeria. However, in Jaen we can find an affricate pronunciation also given in Toledo and South America. 1.2.1.4 Fricatization of ch While Spanish ch is affricate /ʧ /, Andalusian ch can also be fricative /ʃ/ in expressive situations, with a variety of realizations that range from interdental or dental to palatal, being the pre-palatal version the most frequent. Although it is a feature decreasing in frequency, which speakers who use it tend to avoid in formal situations, it is still identified as a stereotyped characteristic outside the area. This feature is closely linked to the merger we have just described above, yeísmo, as they exemplify the Andalusian tendency to merge phonemes: [ʧ ] > [ʃ] and [ʎ]-[ɟ ] > [ɟ], giving way to the minimal pair of voiceless and voiced pre-palatal fricatives [ʃ]-[ɟ]. For Alvar (1990), these processes are irreversible and will lead to the establishment of the opposition mentioned above. What is clear for him is that ll will
28 never be lateral again and ch will never be affricate after fricatization is established. The second assumption in this statement seems overambitious, especially if we take into account Villena Ponsoda’s (2002) claims. Speakers in areas of seseo are relatively conservative and avoid the lenition of [ʧ ], while speakers in areas of ceceo are more innovative, allowing the lenition of [ʧ ] and [ɟ ]. As we have seen before, ceceo is a minority phenomenon and the least prestigious solution in the presence of seseo and distinction. Table 1 Phonetic inventory of seseo, ceceo and distinction. Based on Villena Ponsoda (2002, p. 199) seseo ceceo distinction labia l denta l palata l vela r labia l denta l palata l vela r labia l denta l palata l Vela r p t ʧ k p t k p t ʧ k b d ɟ g b t g b d ɟ g f s h f θ ʃ-ɟ h f θ-s h 1.2.1.5 Consonant deletion Other than the aforementioned /s/, /r/ and /l/, the rest of consonant phonemes tend to be uttered with a lax pronunciation in the whole southern region of Spain, with the possible appearance of total deletion of the final consonant. In syllable-final position within a word and followed by another consonant phoneme, the first consonant tends to disappear and the second one is geminated, as in obturar [otturár]. Before [h], /n/ can be dropped with the possible nasalization of the previous vowel sound, as in naranja [narãha]. In the case of word-final /n/, we can find two situations: a) stressed syllable and b) unstressed syllable, with the following solutions:
35 generally distinguished by the final /s/. The presence of masculine articles makes distinction easy for WAS speakers, as in El perro/Un perro (The dog/A dog) and Los perros/Unos perros (The dogs/Some dogs), because the form of the article changes. In the case of feminine nouns, speakers resort to aspiration, as in La mesa/Una mesa (The table/A table) and Las mesas/Unas mesas (The tables/Some tables). Connected to this aspect, O’Neill (2005) conducted a study on the production and perception of final /s/ by native speakers of WAS in second-person singular verbs and plural nouns. As the target words were in sentence-final position, /s/ tended to be uttered as a very weak aspiration. However, when compared to the production of third-person singular verbs and singular nouns, the author observed that there also seemed to be a very slight aspiration in these cases, giving way to the phonetic neutralization of this phonological contrast. These results were also reflected in his perception experiment, i.e., “in final position, in normal speech, there is no distinction between the sequence VS and V and therefore, the morphological distinctions which rely on this final sibilant element are lost in this position” (p. 159). In our current experiment, target words are embedded in initial or medial position precisely to avoid this neutralization. In AAE, the use of the morphological marker –s is different than in GAE, in the sense that it is usually omitted from third-person verbs and, to a lesser extent, from plural nouns in the presence of quantity markers, and also from genitive constructions. Nevertheless, it is employed to function as a narrative indicator or as indicative of habitual behavior in first-person verbs (Green, 2004a; Smitherman, 1999). Thus, the morphological marker –s exists in AAE, although its use and function differ from those in GAE.
36 In relation to this phenomenon, particularly interesting is the work by Johnson (2005) and de Villiers and Johnson (2007), who studied the comprehension of thirdperson singular /s/; in the first case, by AAE-speaking children and, in the second case, across dialects of American English (including AAE). Results from these studies indicate that AAE-speaking children do not understand /s/ as a number agreement marker. If it were part of their underlying system, in spite of its infrequent realization phonetically speaking, they would be sensitive to its perception: “If the speaker has available two competing grammars, then in comprehension, third /s/ would be understood as an agreement marker of singular subjects, but zero marking would be ambiguous between the two grammars” (Jonhson, 2005, p. 117). These findings have been recently supported by Beyer and Hudson Kam (2012). However, from these results we cannot determine if speakers of AAE may use third-person /s/ as a subject marker at a later age. We will revisit these studies when stating our objectives in Section 2.3.1. 1.3.1 Description of [h] and [s] Fricatives are sounds produced with an obstruction in the vocal tract that generates noise. “Frication noise is generated in two ways, either by blowing air against an object … or moving air through a narrow channel into a relatively more open space” (Hagiwara, 2009 10 ). The first description is pertinent to [s] while the second one refers to [h], our two sounds under study. The obstacle in the production of [s] is the teeth, at the front of the oral cavity, while [h] is produced at the back of the oral cavity without an obstacle against which air blows. 10 http://home.cc.umanitoba.ca/~robh/howto.html
37 Aspiration in Spanish can take place two ways: derived from the phoneme /x/ or from implosive /s/. In both cases, it has been traditionally transcribed as [h], without taking into account its position and its different realizations according to the surrounding context. As reported by Marrero (1990) in her study of the Spanish spoken in the Canary Islands, aspiration derives from /x/ tends to be pharyngeal [h], while aspiration of implosive /s/ is laryngeal and can be voiced [ɦ] in intervocalic position, similar to breathy speech, but velar [x] 11 before velar consonants. The laryngeal aspiration, she argues, is similar to the English /h/, with which Widdison (1993) agrees: “según Ladefoged, el sonido del murmullo corresponde, a groso modo [sic], a la pronunciación de la [h] intervocálica de las palabras inglesas ahead y behind” (p. 47) [“According to Ladefoged, the sound of the murmur corresponds, broadly speaking, to the pronunciation of intervocalic [h] of the English words ahead and behind”]. An example of WAS intervocalic [h] can be seen in Figure 1 below: Figure 1. Intervocalic [h] in “bebes agua” by a WAS female speaker 11 Although /x/ is used to label the CS phoneme, [x] is employed here to account for the velar place of articulation of aspiration. é e h á wa Time (s) 1.087 1.589 1.58918577bebes agua (WAS female)
38 In English, /h/ is defined as a voiceless glottal fricative sound by some authors (Collins & Mees, 2003; Ogden, 2009, Roach, 2010; among others), while other authors do not classify it as a consonant but rather as part of the vowel (Hagiwara, 2009; Jongman, Wayland, & Wong, 2000; Ladefoged, 1982). As Johnson (2012) points out, this sound is fricative “if we define the class as sounds produced with turbulent airflow; but, unlike other fricatives, they are nonconsonantal in the sense that they have … vowel-like spacing between formants” (p. 160) with higher amplitude in their higher formants than vowels. Lorenz (2012) explains that “the IPA chart lists it as glottal, but the constriction is rather somewhere in pharynx or larynx” (p. 30). In any case, in intervocalic position, /h/ also becomes voiced [ɦ]. The characteristics of [s] and [h] make [s] the strongest fricative sound while [h] the weakest fricative sound. In fact, the spectral peak in [s], with the highest frequency concentration of all fricative sounds, is near 8 kHz, with a minor peak around 4 kHz, whereas [h] has a much lower frequency. The pharyngeal fricative [h] peaks at 1.5 kHz, while the laryngeal fricative [ɦ] peaks at 2.56 kHz and the velar fricative [x] peaks at 3.45 kHz (Martínez Celdrán & Fernández Planas, 2007). In her study, Barreiro Bilbao (1994) conducted an acoustic cross-analysis of RP English and Spanish fricatives. As an isolated sound, she found that both English and Spanish [s] have a smaller range of frequencies than [f]. English [s] showed a concentration of energy around 13 349 Hz, while Spanish [s] had a concentration of energy around 10 915 Hz. In citation form, English [s] had a duration of 205.8 ms in initial position, 209.1 ms in medial position, and 299.7 in final position, while Spanish [s] presented a duration of 192.5 ms in initial position, 156.4 ms in medial position, and 197.5 ms in final position. The energy of English [h] stretched up to
39 8703 Hz and showed a duration of 91.9 ms in intial position and 156.8 ms in final position. Nevertheless, she did not analyze Spanish aspiration [h]. However, as seen in Section 1.2.1, aspiration of implosive /s/ is not only a matter of /s/ [h]. It is generally deleted in absolute position, while it is commonly realized as [h] or even [s] in word-final position followed by vowel. When followed by voiceless stops, we can observe preand/or post-aspiration; when followed by voiced stops, these tend to become fricatives (as we will see in the following chapter) and, when followed by nasals, lateral, and other fricatives, gemination is the most common solution (Romero Gallego, 1995). In the case of /s/, it is a voiceless alveolar fricative sound in CS, as well as in English, particularly in syllable-initial position and in syllable-final position when followed by a voiceless consonant (except /θ/ or /t/) or a pause, and in intervocalic position. When in syllable-final position and followed by a voiced consonant (except /d/), it becomes a voiced alveolar fricative sound. Additionally, when followed by /θ/ or /t/, it is a voiceless dental fricative sound, but if followed by /d/, it becomes a voiced dental fricative sound (Garrido Almiñana, Machuca Ayuso, & de la Mota Gorriz, 1998). Additionally, the following consonant not only affects the place of articulation and voice of /s/ but it can also affect its intensity, frequency, and duration. Nevertheless, implosive /s/ also has an effect on its surrounding context, e.g. it lengthens the preceding vowel (Widdison, 1993). As an example, Figure 2 below shows CS intervocalic [s]. Observe that, as opposed to Figure 1, the sibilant is clearly delimited between the vowels.
40 Figure 2. Intervocalic [s] in “bebes agua” by a CS female speaker In relation to this, Widdison (1993) conducted an experiment to explain the possible origins of Spanish aspiration in syllableand word-final position. A native Spanish speaker recorded words with and without implosive /s/, such as pasta and pata. The author then separated the vowel preceding /s/ and inserted it in the word without /s/, replacing its actual vowel. Upon doing this, he conducted an identification task with native speakers of Spanish, a great number of which identified the new word as containing /s/, even though it was not physically present. His conclusions were that the vowel alone already indicates the presence of the sibilant that the listeners associate with /s/ at a lexical level, whether it is /s/ that actually follows or aspiration, i.e. “[h] siempre está presente en la señal acústica de la vocal, pero sólo se percibe cuando los rasgos esenciales de [s] se reducen a un mínimo” (p. 55) [“[h] is always present in the acoustic signal of the vowel, but it is only perceived when the essential features of [s] are reduced to a minimum”]. é e s á wa Time (s) 0.7013 1.253 1.25287762bebes agua (PS female)
41 1.4 Summary In this chapter, we have seen how both AAE and WAS are dialects of English and Spanish that carry certain stigmatization and the origins of which are not completely defined. The two dialects have a set of phonetic features that makes them unique and sets them apart from the mainstream characteristics of English and Spanish. AAE is represented by an absence of interdental fricatives in syllable-initial or syllable-final position. Instead, we find alveolar stops and labiodental fricatives. It is also characterized by the absence of /r/ and /l/ in syllable-final position. WAS is mostly characterized by the aspiration of the sibilant /s/, which affects the following sounds: aspiration and post-aspiration in voiceless stops, fricatization of voiced stops, gemination of nasals and other consonants. Additionally, it displays a set of mergers and the fricatization of /tʃ/. What these two dialects have in common is the deletion of consonants in medial or final position and how their morphological marker for verb agreement and plurality is affected. AAE absence of the morphological marker –s from third-person verbs and plural nouns seems to come from internal grammatical rules while WAS aspiration of this marker in second-person verbs and plural nouns is derived from its phonetic characteristics. In the following chapter, we review the acoustic characteristics of the English and Spanish sounds that concern us in this study.
42 CHAPTER 2 ACOUSTIC PHONETICS When we talk about phonetics, we can do so from the point of view of production, transmission, and perception of speech sounds. Articulatory phonetics describes how sounds are formed by the vocal tract of the speaker; acoustic phonetics describes the characteristics of the sounds that reach our ears; and perceptual phonetics studies how these sounds are understood by the listener. The listener needs to actively participate in the process, extracting information from the signal in terms of intrinsic characteristics and context characteristics. The listener also uses information that is independent of the signal and in relation to their linguistics experience stored in memory. To decode a linguistic signal, listeners go through three stages (Marrero, 2001): i) audition, it is a passive and automatic mechanism by which the signal activates the fibers in the auditory nerve that allow us to distinguish sounds, ii) perception, when the nerve system converts the signal into linguistic units, and then segments, classifies and categorizes them, and iii) comprehension, which is the interpretation of the message in terms of grammatical and semantic meaning. Given that our articulatory system tends to produce sounds as similar as possible, and that our perceptive system needs sounds to be as distinguishable as possible, it seems that our perceptive system has played a key role in the evolution of language. In discrimination tasks, where listeners have to determine whether two sounds are similar or different, the mechanism activated is auditory, i.e., the characteristics of the sounds are essential. In identification or categorization tasks, when listeners have
43 to identify and label stimuli, they resort to their mental models that they have of such sounds to make a decision. One of the main differences between discriminating and identifying is that we can potentially detect minimal differences between sounds but our capacity to categorize and store them in memory is limited. Here, two processes of perception are at play, as we mentioned above. Auditory perception is a bottom-up process based on the physical characteristics of the sounds, while categorical perception is a top-down process that interprets sounds in terms of the pre-existing categories in memory. When we use this last process, we label sounds that share certain characteristics within the same category. Two sounds can considerable differ in parameters such as duration and frequency but still be assigned to the same category. This encompasses the variability that can be found in the signal, such as coarticulation and dialectal variation. On the contrary, other sounds may minimally differ in one property that is important enough to be categorized as two distinct sounds. As Martínez Celdrán and Fernández Planas (2007, p. 113) state “diferencias articulatorias pueden producir cambios acústicos muy destacables o, por el contrario, cambios mínimos” [“articulatory differences can produce very remarkable acoustic changes or, on the contrary, minimal changes”]. Therefore, categorical perception maintains the characteristics that distinguish sounds and minimizes irrelevant differences to compensate for the imperfect one-to-one correspondence between acoustic cues and phonetic features.
44 2.1 Acoustic Cues We make speech sounds audible when the air is pushed out of our lungs while producing a noise in our throat or mouth. By means of the actions of the tongue and the lips (articulators), we make changes in these basic noises. The speech sound is a spectrum of acoustic energy produced by the vibration of our vocal folds and then filtered by the articulators in our vocal tract. The mechanism of speech production involves four processes (Ladefoged & Johnson, 2010): i) the airstream process, that is, the ways in which we push air out of our lungs; ii) the phonation process, which involves the actions of our vocal folds. When they vibrate, they produce voiced sounds; when they do not vibrate, they produce voiceless sounds; iii) the oro-nasal process, by which we produce oral sounds when the air escapes through the oral cavity and nasal sounds when the air escapes through the nasal cavity); and iv) the articulatory process, by which our tongue and our lips interact with the roof of the mouth and the pharynx to articulate the sounds. Speech sounds can be divided into three categories (Hagiwara, 2009; Ladefoged & Johnson, 2010): i) periodic voicing, which is produced when the vocal folds vibrate; ii) devoicing, sounds produced without vibration of the vocal folds; and iii) aperiodic noise, which is when a turbulent airflow is produced in a random way. In the production of vowels, on the one hand, the vocal tract is relatively open and the air escapes without obstruction, which gives these sounds great loudness. The vocal folds vibrate and, thus, vowels are voiced. “The primary acoustic characteristics of vowels is the location of the formant frequencies, specifically, the first three formants (F1-F3)” (Reetz & Jognman, p. 182), which provide information about vowel quality. The rest of the formant frequencies above F3 provide more information
51 One study that deviates from the previous studies in terms of the type of stimuli used is the one conducted by Yao (2007) on the closure and VOT of voiceless stops in English connected speech. It seems that the factors taken into consideration proved to only account for 26% of the variability in closure and VOT values. Age and gender could only explain 1% of the variability (considering that their age range was from under 30 to over 40), speaking rate only accounted for 13% of variability in VOT and 4.5% in closure duration, place of articulation could only explain 2.2% of variability in VOT but a higher percentage of variability in closure duration (8.1%). Word frequency, on the contrary, was found to have an effect on both VOT a closure duration, i. e., “If some words occur extremely often, it is possible that they become the target of certain changes in production, for instance, acceleration, phone reduction and coarticulation” (p. 218). 2.1.1.1 Stops after WAS aspiration In Section 1.2.1 (Alvar, 1996; Gerfen, 2002; Torreira 2007a, 2007b, 2012), we saw a general description of how aspiration of implosive /s/ can affect the following sounds; more in particular, how it affects voiceless stops. As Parrell (2012) argues: the productions of /s/ in Western Andalusian Spanish are reported to be somewhat variable, ranging from a full sibilant to preaspiration to a breathy period at the end of the preceding vowel to post-aspiration … with the last being the most common. (p. 37) Observe Figures 4 and 5, where a clear difference in VOT duration can be detected between CS [t] and WAS [t h ]:
52 [t] vowel Figure 4. Word-initial [t] after sibilance [t h ] vowel Figure 5. Word-initial [t h ] after aspiration Time (s) 0 0.0898 -0.2503 0.393 0 Time (s) 0 0.0898 0 8000 Frequency (Hz) 0.0898036627 Time (s) 0 0.1269 -0.1644 0.1665 0 Time (s) 0 0.1269 0 8000 Frequency (Hz) 0.126916159
53 Torreira (2007a) was a pioneer in describing this phenomenon, although he acknowledges that two previous studies had already pointed in this direction (Maza, 1999; Vaux, 1998). His study first analyzed word-internal /st/ in laboratory-recorded speech of CS and WAS speakers, and subsequently in spontaneous speech of WAS and Eastern Andalusian Spanish (EAS) speakers. In both cases, he observed higher VOTs for the Andalusian stop after aspiration in comparison to the CS stop after sibilance. Despite variability found in the recordings, and factors such as speech rate, prosodic context, syllable stress, his findings were consistent with the premise that Andalusian aspiration induces longer VOTs. Another factor that was derived from aspiration is that both stop closure and the previous vowel were lengthened, as long as we consider the aspiration of /s/ as part of the vowel. Otherwise, as is the case with vowels before /s/, they were actually shorter. Subsequently, Torreira (2007b) compared the production of word-internal /st/ of WAS with the production of the same sequence by speakers of Porteño (Buenos Aires, Argentina) and Puerto Rican Spanish. What he found is that WAS displays shorter pre-aspiration and longer stop closure and post-aspiration period than the other two Spanish dialects. Under the Articulatory Phonology framework proposed by Browman and Goldstein (1989), in which articulatory gestures are seen as phonological units, the author seeks to provide an explanation for this phenomenon. This framework states that “gestures involved in syllable onsets tend to couple into an in-phase relationship, while gestures in coda position are left out of phase with respect to surrounding gestures” (Torreira, 2007b, pp. 118-119), i.e., at onsets, articulatory gestures tend to be simultaneous while at codas they tend to be more variable. His proposition is that of a gestural reorganization in which the glottal opening for the
54 aspirated /s/ and the supraglottal closure for the following stop overlap instead of being sequential, as is the case with dialects with pre-aspiration. Finally, in 2012, Torreira further investigated WAS aspiration before the three voiceless stops /p, t, k/ according to different speech rates and stress patterns. He found that, despite these two factors, VOT did not significantly vary in duration. Therefore, it seems that “the glottal and supraglottal gestures may be phased very closely even in conditions in which we would not expect much articulatory overlap, hence the lack of significant effects of speech rate and stress location on VOT” (p. 61). In reference to the variability found in WAS aspirated stops, Ruch (2008) researched the production of /st/ in Seville. What she found were nine possible realizations for this sequence: two with sibilants [st], [ s t]; four with aspiration [ h t], [ h t h ], [ s t h ], [t h ], one with assimilation [t:], one with complete deletion of /s/ [t], and finally, the new phenomenon that we mentioned in Chapter 1: the affricated [t s ]. The most common of these realizations was the post-aspirated stop [t h ] (49.1%), followed by the affricated stop [t s ] (22%). Additionally, Ruch (2012) conducted a sociophonetic study of the production of /t/ and /st/ in internal-word position with speakers from WAS (Seville) and EAS (Granada), taking into account their gender and their age. She concluded that young speakers produce post-aspiration significantly more frequently than older speakers not only in Seville, but also in Granada. They also produce less pre-aspiration, although this fact was only significant for speakers in Seville. Additionally, she found that female speakers showed greater differences in VOT values than male speakers. O’Neill (2009) also studied the sequence /st/ in WAS from Seville, narrowing down the effect of aspiration to two most frequent productions: [‘pa ɦ t h a] and [‘pat h a],
55 i.e. aspirated stops with or without pre-aspiration. What is interesting is that the author considers the second realization as part of a new set of phonemes in the dialect [p h , t h , k h ], working in opposition to their unaspirated counterparts. Instead of being the result of an overlap of gestures, as proposed by Torreira, “these pronunciations correspond to the phonetic realisation of a different sequence of phonemes” (p. 79), i.e., these set of sounds would be phonetic categories in itself and not the result of coarticulatory gestures. Parrell (2012) corroborates the claims by Torreira of a post-aspiration phenomenon in WAS, but he states that the question of “whether this reduction is an online phonetic process or a phonological one has not been thoroughly investigated” (p. 37). Finally, the most recent piece of work concerning post-aspirated voiceless stops of Seville is the study carried out by Horn (2013). She investigated the phenomenon in a sentence reading task from various perspectives. First, she studied whether the post-aspiration reported for /t/ also extended to /p/ and /k/. In this regard, she found that place of articulation “is the only robust predictor of the presence of significantly long postaspiration” (p. 81). Post-aspiration also extended to the velar sounds but not to the bilabial sound, opposite to the findings in Torreira (2012). Its duration was significantly shorter for /p/ than for the other two stops. Second, she aimed at analyzing the phenomenon from a social and linguistic perspective. The longest duration of post-aspiration was found when the preceding vowel was stressed and when in word-internal position, once more, in disagreement with Torreira’s claims (2007a, 2012). Although the social factor had no effect in these realizations, there was a tendency for younger women with college education level to reduce sibilance and produce longer post-aspiration. And third, she interpreted these results
56 under the Articulatory Phonology framework. Just as the previous studies, she concluded that there is a negative correlation between the presence of sibilance and post-aspiration. 2.1.2 Fricatives Fricative sounds are produced when the articulators constrict the passage through which the air escapes. These sounds are continuant, in the sense that “you can continue making them without interruption as long as you have enough air in your lungs” (Roach, 2000, p. 48). When the air passes through the articulators, it creates turbulence due to the size of the passage and the volume velocity of the airflow. Therefore, “the faster the air molecules move, the louder the sound … the narrower the channel, the louder the turbulent noise” (Johnson, 2012, p. 154). Nevertheless, most fricatives are produced when the air hits an obstacle in the passage, i.e., the teeth or the lips, increasing the amplitude of the turbulence. This turbulence noise is represented as a very dark area in the spectrogram. As it was the case with stops, fricatives can also be voiceless /s, f, θ, ʃ, x, h/ and voiced /z, v, ð, ʒ/ (the classification of /h/ is controversial, as we mentioned earlier). Fricatives can be described according to four characteristics: “spectral properties of the friction noise, amplitude of the noise, duration of the noise, and spectral properties of the transition into and out of the surrounding vowels” (Reetz & Jongman, 2009, p. 189). Sibilant fricatives have a more pronounced spectral shape because the air hits the teeth. Therefore, the alveolar sibilants typically present clear, distinct spectral shapes while labiodental and (inter)dental non-sibilant fricatives display a relatively flat spectrum. Velar fricatives present little energy at higher
57 frequencies since their greatest amount of energy concentrates at lower frequencies; particularly, in the area corresponding to the F2 of the adjacent vowel. Unlike their voiceless counterparts, voiced fricatives have two sources of energy: not only does it originate from the turbulent noise derived from the constriction of the air passage, but also from the vibration of the vocals folds, which generate low-frequency energy. The spectrograms of both types of fricatives are similar, with the exception that “they contain additional low-frequency energy corresponding to vocal fold vibration and slightly less intensity in the higher frequencies because part of the energy of the airstream serves to make vocal folds vibrate” (Reetz & Jongman, 2009, p. 192). In Spanish, the fricative sounds are /f, θ, s, x/, to which Quilis (1981) adds the allophones [h] and [ɦ]. In English, the fricative sounds are the voiceless /f, θ, s, ʃ, h/ and their voiced counterparts /v, ð, z, ʒ/. As we reported in Section 1.3.1, Barreiro Bilbao (1994) conducted a cross-sectional study of the acoustic characteristics of Spanish /f, θ, s, x/ and RP English /f, θ s, ʃ, h/. Among the characteristics measured, we find range of frequency, duration, and their spectral peaks. Concerning the range of frequency, she concluded that non-sibilant /f, θ/ present a great amount of dispersion of energy that extends between 1000 Hz and 15 400 Hz. Non-sibilant /x, h/ have a concentration of energy in the lowest area of the spectrum, from 0 Hz to 11 500 Hz. Sibilant /s, ʃ/ show a narrower band of frequency, from 1300 Hz to 14 800 Hz, although with higher intensity. With respect to this, their place of articulation has an effect on their respective frequency. /f, θ/ are articulated at the front of the oral cavity, /s, ʃ/ are articulated in the mid area of the oral cavity, while /x, h/ are articulated at the back of the oral cavity. The fricatives articulated at the back present lower frequency limits than the other sounds. Those articulated in the middle section
58 have low upper limits and higher lower limits. Finally, the fricatives articulated at the front of the oral cavity present higher lower limits and low upper limits. With respect to the duration of the fricatives, she found that both sets of fricatives had a similar duration according to their place of articulation. However, fricatives in word-internal position were shorter in Spanish, while fricatives in wordinitial position were shorter in English. Additionally, the velar sound /x/ had a similar duration to that of /f/, whereas English /h/ was very short. For the author, differences in spectral peaks are the key characteristic to distinguish these fricatives. This parameter is crucial to explain why the “trasvase de algunos de estos sonidos … de una lengua a otra conlleva una pronunciación errónea y, en otros casos … no supone cambios importantes a nivel perceptivo o articulatorio” (p. 477). [“transfer of some of these sounds … of a language to another leads to an erroneous pronunciation and, in other cases … it does not imply important changes on a perceptual or articulatory level”]. She divides them into three groups: i) Sibilants, which have formants with great amplitude due to the high-pass filter of the oral cavity. Spanish /s/ has a great concentration of energy in one formant from 3515 Hz to 6317 Hz, while English /s/ has this energy from 4336 Hz to 6619 Hz. “Cuanto más se retrae la punta de la lengua más baja es la frecuencia de dicho formante” (p. 467) [“The more retracted the tip of the tongue is, the lower the frequency of such formant”]. According to Quilis (1981), the closer the place of articulation is to the front of the oral cavity, i.e., the dental area, the less strident /s/ becomes. In other words, the length of the vocal tract from the point of constriction to the lips is inversely correlated to the frequency of the peak in the spectrum (Hughes & Halle, 1956).
59 ii) Labiodental non-sibilants have an almost flat spectrum. For Spanish /f/, the greatest information is contained in the first three formants, that is, below 6000 Hz. For English /f/, this information can also be found around 11 300 Hz. /θ/ presents more noise than /f/, without formants. Its information lies in both low and higher frequencies. For Spanish, it peaks up to 9000Hz, while for English it peaks up to 8000 Hz. iii) Velar and glottal non-sibilants present great energy in the lower area of the spectrum and have a marked coarticulation with the adjacent sounds. Spanish /x/ contains information in the first three formants below 3000Hz. Over 4000 Hz, it only presents noise without formants. English /h/ has five formants up to 8000 Hz. In fact, the spectrum of the fricatives articulated at the front of the oral cavity, in conjunction with neighboring vowels, see how their spectral peaks in the higher area of the spectrum increase their amplitude; those fricatives articulated in the middle section of the oral cavity suffer a decrease in their F1 and an increase in their F2; and the fricatives articulated at the back of the oral cavity suffer changes in amplitude and their formant frequencies. The energy of apical /s/ starts at 3500 Hz and reaches the highest point around the center of the spectrum (Martínez Celdrán, 2004; Martínez Celdrán & Fernández Planas, 2007). However, before dental stops /t/ and /d/, sibilance is said to suffer a process of dentalization, to which Quilis (1966) opposes, claiming that a dental allophone would be close to [θ]. Although there seem not to be great differences between apical /s/ and “dental” /s/, some differences in F1 seem to appear, as well as differences between intervocalic /s/ and “dental” /s/. Whether this is a question of an
60 assimilation process or a coarticulatory process, the authors point at a partial assimilation. As far as the rest of the Spanish fricatives are concerned, García Santos (2002, reported in Martínez Celdrán & Fernández Planas, 2007), state that their perception varies according to their duration. /f/ is perceived when longer than 90 ms; if its duration is shortened to 40-80 ms, it is then perceived as [v], while it is identified as the approximant [ß ] when its duration is less than 20 ms. Similarly, /θ/ is identified when its duration is longer than 85 ms, while it is perceived as the approximant [ð ] when its duration is shorter than 35 ms. Along these lines, Herrero Moreno and Supiot Ripoll (2002) investigated the characteristics than can distinguish voiceless fricatives /f, θ, x/ from the voiced approximants [ß , ð , ɣ ] of voiced stops /b, d, g/. In particular, they focused on voicing, noise, and duration as possible influential factors. They found that voicing and noise are not reliable factors to distinguish these sounds; on the contrary, duration counts as the key factor for distinction. In this case, the authors also equate duration and tension. On this aspect, Martínez Celdrán and Fernández Planas (2007) disagree with the notion of duration equated to tension, claiming that tension is not the product of duration but rather an increase in the tension is what leads to longer duration. Likewise, English /s/ also shows a large amount of energy at high frequencies (Ladefoged & Ferrari Disner, 2012), extending over 10 000Hz and with little energy below 3500 Hz. /ʃ/, in turn, concentrates energy around 3000 Hz, and thus is lower in pitch than /s/. On the contrary, /f/ and /θ/ show energy over a range of frequencies, i.e., greater dispersion, with higher concentration of energy around 3000-4000 Hz for the former and above 8000 Hz for the latter. Their voiced counterparts /z, ʒ, v, ð/, respectively, have less intensity because the movement of the vocal folds to produce
67 Figure 11. Spectral slice of [x] 2.2 Summary In this chapter, we have focused on one of the branches of phonetics: acoustic phonetics. We have seen how speech sounds can be described in terms of their acoustic properties, particularly as far as stops and fricatives are concerned. Additionally, we have reviewed several studies that investigated the nature of these sounds in WAS, as a result of the aspiration of sibilance given in this dialect. It seems that VOT is a good indicator of the presence of aspiration in WAS voiceless stops, while the spectral moments of fricative sounds have rendered diverse views until Jongman et al.’s (2000) work. In the following chapter, we cover the area of perceptual phonetics, specifically the area of L2 speech perception, and we explain how the acoustic cues of the sounds, along with the listeners’ characteristics, play a role in this process. Frequency (Hz) 0 2.205·10 4 Sound pressure level (dB/Hz) -20 0 20
68 CHAPTER 3 L2 SPEECH PERCEPTION Speech perception in general can be described as the decoding of the acoustic signal in speech into meaningful information for the listener. Native speakers, when processing continuous speech, ignore certain acoustic cues in favor of those that are relevant in their L1, despite age, gender, or rate of speech of the speaker, to “focus on the words being said, and not so much on exactly how they are pronounced” (Johnson, 2012, p. 100). The way we speak guides the way we interpret speech. This leads us to understand sounds according to the language-specific categories that we have learned to use in our L1. Thus, “we hear sounds that we are familiar with as talkers” (p. 107), and our perception is also guided by the linguistic knowledge that we have of our L1, i.e., the phonotactic rules of our native language. The perception of non-native sounds is said to depend on several factors related to the listener, such as L1, age of learning (AOL 13 ), and L2 experience. Initially, L1 listeners will have difficulty with L2 contrasts that are not phonetically contrastive in their L1. Contrasts that are given in the L2 but absent in the L1 may not be distinguished by the listeners. A classic example of this is the perception of English /r/ and /l/ by L1 Japanese listeners as one single L1 category (Miyawaki, Strange, Verbrugge, Liberman, Jenkins, & Fujimura, 1975; Best & Strange, 1992; Polka & Strange, 1985 among others). As this contrast is not given in their native language, L1 Japanese listeners are generally unable to distinguish these L2 sounds as separate phonemes. As also found by Flege, Bohn, and Jang (1997), L1 Spanish 13 This factor will be briefly addressed in Section 3.2.2
69 listeners in their study assimilated English /i:/ and /ɪ/ to Spanish /i/. Since this contrast is not present in their L1, they matched them to the only phoneme available in their native language. However, English /e/ was assimilated to Spanish /e/ and English /æ/ was assimilated to Spanish /a/, which are two distinct categories in Spanish. L1 Korean and Mandarin listeners also confused /i:/ and /ɪ/, as this contrast is not given in their L1 either. However, that was not the case for L1 German listeners, whose L1 does possess this contrast. Several studies have pointed out at the reliance on durational cues by L1 Spanish listeners in the perception of L2 English vowel contrasts, rather than on spectral cues inexistent in their L1 (Escudero & Boersma, 2004; Escudero, Benders, & Lipski, 2004). This serves as evidence that L1 experience may determine the way certain phonetic cues are used in L2 speech perception. Nevertheless, features that are shared by L1 and L2 on certain segments may not be transferred to new L2 sounds automatically. Consequently, the fact that L1 and L2 share the same features may not necessarily favor perception or learning. Another factor to be taken into account when examining L2 speech perception is the listeners’ experience in the L2, which may lead L2 learners to reorganize their phonetic systems as experience increases. Beginning L2 learners may find difficulties that can be overcome with increasing experience in the language. Bohn and Flege (1990) investigated the perception of English vowels /i:, ɪ, e, æ/ by experienced and inexperienced L1 German listeners. While experience was not an influential factor for the perception of vowels that had similar or identical counterparts in German (/i:, ɪ, e/), it proved to be crucial for the perception of /æ/, which was a new sound for the listeners. The inexperienced listeners performed significantly lower than the experienced listeners in the identification of this L2 sound and seemed to resort to
70 durational cues to distinguish it from /e/ (see also Flege & Liu, 2001; Flege, Takagi, & Mann, 1996 for further effects of experience). However, some studies have pointed out that experience may not render higher accuracy in some cases. Levy and Strange (2008) found that experience was influential in the perception of L2 French contrasts /u, œ/, /i-y/ and /y-œ/ for experienced and inexperienced L1 American English listeners. However, no differences were found between both groups of listeners for the perception of the contrast /u-y/. Levy (2009) also studied the perception of L2 French vowels by L1 American English listeners with no experience in French, with formal instruction in French, and with formal instruction and immersion in French. She concluded that higher accuracy was found for the most experienced listeners and in bilabial context. In this case, the acoustical similarities between French vowels were not sufficient to explain context-specific assimilation patterns. Instead “it is suggested that nativelanguage allophonic variation influences context-specific perceptual patterns in second-language learning” (p. 1138, see also Levy & Law II, 2010). To account for these contradictory results, two additional factors need to be taken into consideration in the perception of non-native sounds: the type of contrast under study and the type of acoustic cues of the L2 sounds (Barreiro Bilbao, 2002). Not all contrasts are similarly difficult; other than the L1 background and the L2 experience of the listeners, we should also look at the psychoacoustic salience of the sounds under study, that is, the sounds we perceive and experience as more salient in relation to our physiological capacity (auditory perception) and our phonetic knowledge (categorical perception). As pointed out by Strange and Shafer (2008):
71 “… in general, place-of-articulation contrasts in consonants, cued primarily by spectral differences of short duration, may be considered less salient than voicing contrasts, cued primarily by temporal parameters … Contrasts in manner of articulation (e.g., fricative vs. stop) may be considered very salient in that they are differentiated by differences in sound source characteristics.” (p. 175) In a series of studies (Hendrick & Carney, 1997; Hendrick & Younger, 2001; Hendrick & Younger, 2007), the role of relative amplitude and formant transitions of English stops and fricatives in speech perception was investigated. Although the nature of these studies was to investigate perception in hearing-impaired L1 listeners in relation to normal hearing listeners, insightful findings with respect to acoustic cues can be drawn. Studies have shown that in CV sequences manipulating a frequency region of the consonant in the syllable relative to the amplitude of the same frequency region in the following vowel (relative amplitude) influences the perception of place of articulation for fricatives and stop consonants. In this regard, Chen and Alwan (2003) studied the perception of English stops and fricatives by English L1 listeners in terms of place of articulation: labial /b, p, f, v/ and alveolar /d, t, s, z/, in three vowel contexts /a, i, u/. They found that “the perception of place for plosives and fricatives depends on whether the consonant is voiced or voiceless” (p. 1499), i.e., voiceless consonants were more robust than their voiced counterparts. Later, Alwan, Jiang, and Chen (2011) conducted a similar study in which they found that the identification of the distinction between labial and alveolar stops in noise depends on the manner of articulation and its interaction with voicing. In a cross-language perception study, Silbert, de Jong, and Park (2005) investigated the perception of English consonants by Korean listeners in terms of voicing, place of articulation (labial/coronal), manner of articulation (stop/fricative),
72 and position in the syllable (initial/final). The Korean language does not have nonsibilant fricatives produced at the front of the oral cavity and neutralizes voicing and manner of articulation in syllable-final position; thus, the identification task tested the effects of L1 specific phonological patterns in the perception of non-native features. The identification of voicing was rather good for labial and coronal stops and fricatives in syllable-initial position, although slightly worse for labial fricatives. In both cases, there was a bias towards voiceless classification. In syllable-final position, they exhibited a poor performance in the identification of voiced labial stops and coronal fricatives, with a tendency to identify the fricative sounds as voiceless sounds. In terms of poor performance, it seems that “being a fricative and being coronal both increase the likelihood that the listeners will call a segment voiceless” (p. 13), resulting in the perception of consonant noise of voiced and voiceless fricatives in a similar way. A study by Wagner, Ernestus, and Cutler (2006), focused on the role of L1 fricative inventory in the identification of L2 fricatives. They studied how listeners of German, Dutch, English, Spanish, and Polish identified spectrally similar fricatives /θ/ and /f/ in terms of formant transitions with and without manipulation. Since German and Dutch do not have spectrally similar fricatives, they were not affected by the changes in transitions, while listeners of the remaining three languages did. Their conclusion is that all listeners “may be sensitive to mismatching information at a low auditory level, but that they do not necessarily take full advantage of all available systematic acoustic variation when identifying phonemes” (p. 2267). In a similar study, Cutler, Cooke, Garcia Lecumberri, and Pasveer (2007) investigated the identification of GAE consonants in noise by native listeners, and Spanish and Dutch listeners. With respect to fricatives /f, θ/, due to the similarities of
73 their native inventory, Spanish and English listeners used the same cues, while Dutch listeners deviated more from native performance. Nevertheless, in the presence of noise, when transitional cues are difficult to distinguish, both English and Spanish listeners’ identification was affected negatively. In this case, the performance of Dutch listeners was not so affected because they did not rely on formant transition information in the first place, “but relied on the steady-state information in the fricative noise” (p. 1588). Similar results were found in Barreiro Bilbao (1999), who also researched the perception of fricatives by L2 listeners. In particular, she studied the effect of voicing and place of articulation in the categorization of two English contrasts that are not present in Spanish, that is, /s, z/ and /s, ʃ/. For the first pair of sounds, when the voice bar was removed, the results were random. Thus, Spanish listeners made use of voice to distinguish these two sounds. In the case of the second contrast, Spanish listeners relied on the frequency and amplitude of the fricatives, and not on the F2 transitions, just as Dutch listeners did in the study described above. Considering all the factors and the research mentioned above, there is still one more aspect of L2 perception that needs to be explored. Most research covers the perception of categorical sounds of mainstream languages; next section covers research concerning the perception, categorization, and identification of dialectal variations of a language.
74 3.1 Perception of L2 Dialect Variants Research indicates that categorization and discrimination varies across L2 contrasts and across L1s. L2 learners’ perception of L2 contrasts systematically depends on the phonotactic, allophonic, and coarticulatory patterns of their L1. Moreover, highly relevant for this dissertation is the assertion that not only does the L1 of the listener have an effect on the perception of a given L2 sound or contrast, but also “L1 and L2 dialect differences can both systematically affect perception of L2” (Best & Tyler, 2007, p. 19). This is why, when encountering an unfamiliar L1 dialect, perceptual learning may need to take place. Studies show that preference is given to unmarked dialects or mainstream varieties of a language. Clopper and Bradlow (2006, 2008) studied the intelligibility of dialect variation in noise. In favorable conditions, GAE and Southern English were better identified than Northern and Mid-Atlantic English. However, in unfavorable noise conditions, the intelligibility of GAE was greater than that of the other three dialects, suggesting that dialect information may be conveyed by aspects of the signal that are relatively vulnerable to perceptual disruption by noise. Sumner and Samuel (2009) also demonstrated this higher accuracy in identification of mainstream features of a language. Furthermore, they also found that being familiar with a dialect renders greater identification of its features. They researched word recognition in dialectal variation and the role of experience in perception and representation. With a series of tasks involving priming they targeted the perception and production of r-dropping in New York City (NYC) dialect, opposed to GAE full realization of –er > [ɚ]. Listeners in this study were i) speakers of NYC dialect, ii) speakers familiar with the dialect, and iii) speakers of GAE unfamiliar with the dialect. They came to the conclusion that dialect production is not
75 always representative of dialect perception and representation; listeners familiar with but not speakers of NYC dialect performed similarly to speakers of the dialect in perception tasks. Thus, experience seems to strongly affect a listener’s ability to recognize spoken words, although variants that are not regionally-marked are preferred overall. If we take this to the domain of L2 acquisition, differences in phoneme inventory between L1 and L2 pose a higher difficulty than L1 differences; L2 learners require exposure to, training in, and use of the L2 to attain the new features. One of the most recent works on L1 cross-dialectal perception is the study by Tuinman et al. (2011), which focused on the perception of British English intrusive [ɾ] by speakers of American English, who accurately perceived vowel-initial words despite intrusive [ɾ]. Nevertheless, these results are in contrast with the findings for the same materials presented to proficient L2 listeners (Tuinman et al., 2007), whose responses showed that they perceived intrusive [ɾ] as word-initial /r/. Although L1 dialect variation is not equivalent to L1-L2 differences, the results broadly showed a robust ability by L1 listeners to adjust to variation within the same language. A study by Cutler, Smits, and Cooper (2005) had also explored this dialect variation within the same L1 with the addition of subjects from an L2. They studied the identification of American English vowels in open and closed syllables by speakers of American English, Australian English, and Dutch. Both groups of English speakers clearly outperformed Dutch speakers; nevertheless, vowel tenseness judgment was more variable for Australian English speakers due to cross-dialectal differences. When speech input mismatches the native dialect, the difficulty is very much less than that which arises when speech input mismatches the native language in terms of the repertoire of phonemic categories available.
76 When we move towards L2 perception by listeners of different L1 dialects, we find works such as that by Chádková and Podlipský (2011), who studied the perception of Dutch /i:, ɪ/, characterized by spectral differences, by listeners of two dialects: Bohemian Czech and Moravian Czech, which have the same contrast. The first one is also based on spectral differences whereas the second one is based on durational differences. As predicted, Bohemian Czech speakers assimilated the Dutch contrast to two L1 categories while Moravian Czech speakers assimilated the L2 contrast to a single category, /ɪ/, supporting the claim that different L1 dialects can render different assimilation patterns of the same nonnative contrast. More recently, Escudero and Williams (2012) studied the perception of Dutch vowels by speakers of Peruvian Spanish (from Lima) and Peninsular Spanish (from Madrid), whose results indicate that acoustic differences in the native dialect can be more influential than proficiency in the L2. Peninsular Spanish speakers outperformed Peruvian Spanish speakers despite being less proficient in Dutch. Therefore, experience in this case does not seem to be most relevant for perception; results show that L1 dialect prevails. Moving towards our dialects under study, we found that research on AAE has been especially directed towards the description of the language in fields such as variation and change, grammar, phonology, lexicon and use, ethnic identity, education, origins and history, and recently hip hop culture (Alim, 2004; Baugh, 2000, 2004a; Billings, 2005; Fasold, 1972; Green, 2004b; Morgan, 2001; Mufwene, 2003; Poplack, 2000; Spears, 2001; Wolfram et al., 2001; Zeigler, 2001, among others). In the field of speech perception and production, research on AAE has traditionally focused on its implications for education, particularly for reading and writing among AAE-speaking children. In any case, research is mainly restricted to
83 bilingual’s representation is based on different features, or feature weights, than a monolingual’s. H7 The production of a sound eventually corresponds to the properties represented in its phonetic category representation. This mechanism of equivalence classification seen in H5 is a process by which an L2 sound can be perceived as identical, similar, or new with respect to an existing L1 sound. The L2 sound will be assimilated to the L1 sound if it is perceived as identical or similar, whereas a new category will be formed for the L2 sound if it is perceived as less similar or new (however, it is unclear what the terms ‘similar’ and ‘less similar’ exactly refer to.) Concerning the perception of non-native contrasts, SLM predicts that if two contrasting L2 sounds are perceived as similar to one L1 sound, then discrimination will be difficult. At the same time, if one of the L2 sounds is dissimilar to any L1 sound, then equivalence will not take place and a new phonetic category will be likely formed, so both perception and production can be carried out relatively accurately. Therefore, “the greater the perceived distance of an L2 sound from the closest L1 sound, the more likely it is that a separate category will be established for the L2 sound” (Flege, 1995, p. 264). 3.2.3 Perceptual Assimilation Models The Perceptual Assimilation Model (PAM), developed by Best (1995), focuses primarily on the perception of nonnative sounds by naïve listeners (i.e. monolingual) with no experience in the L2. This model presents a direct-realist view of speech perception based on gestural information which, unlike SLM, “is not built up from an analysis of simple acoustic features” (Best, 1995, p. 177) but detected from speech directly and actively by means of integrated perceptual systems. L2 sounds “tend to
84 be perceived according to their similarities to, and discrepancies from, the native segmental constellations that are in closest proximity to them in native phonological space” (Best, 1995, p. 193). Monolingual speakers can not only distinguish phonemes but also withincategory phonetic variations, rating them as good or poor exemplars of the category. This idea reflects the notion of warping that we have seen in NLM. According to PAM, assimilation of an L2 phone can follow any of these three patterns: i) the L2 phone can be assimilated to an L1 category as a good exemplar, an acceptable exemplar, or a deviant exemplar of that category; ii) the L2 phone can be classified as uncategorizable, i.e., recognized as speech but not an exemplar of any given L1 category; and iii) the L2 phone may not be assimilated to speech. Additionally, the model establishes six possible types of perceptual assimilation for nonnative contrastive sounds that differ in terms of difficulty: i) if the contrastive L2 sounds are assimilated to two different L1 categories (Two Category or TG type), then discrimination will be excellent; ii) if the contrastive L2 sounds are assimilated as equally acceptable or equally deviant exemplars of one single L1 category (Single Category or SG type), discrimination will be difficult (above chance level); iii) if the contrastive L2 sounds are assimilated to one single L1 category but their goodness to fit differs (Category Goodness or CG type), discrimination will be moderate to very good. Additionally, iv) when one of the L2 sounds is not perceived as similar to any L1 category (Uncategorized-Categorized or UC type), discrimination is expected to be very good.; v) if none of the L2 sounds are assimilated to any L1 category (Uncategorized-Uncategorized or UU type), discrimination will range from poor to very good; finally vi) if the L2 sounds are so different than any L1 sound that they are not perceived as speech at all (Non-Assimilable or NA type), discrimination
85 will range from good to very good (for a study in which a revision of the UC type is suggested, see Guion, Flege, Akane-Yamada, & Pruitt, 2000). 3.2.3.1 Perceptual Assimilation Model-L2 Furthermore, Best and Tyler (2007) developed the PAM-L2 to explain speech perception by late L2 learners and to additionally review SLM from PAM’s perspective. We must take into account that by the term L2 learner, they understand “people who are in the process of actively learning an L2 to achieve functional, communicative goals, that is, not merely in a classroom for satisfaction of educational requirements” (p. 16). On the one hand, the problem these authors see with a foreign language acquisition (FLA) environment is the L1-accented input that learners may receive along with the different dialectal varieties of the L2 language which can interfere with perception. In addition, a further limitation is the usual scenario of FLA being an educational requirement and not a process of active learning to achieve communicative and functional skills, as opposed to SLA learners. On the other hand, unlike naïve speakers, FLA learners are exposed to the L2; thus, the authors encourage research on perceptual adjustment to L2 contrasts in FLA settings as opposed to SLA contexts, which is what we did in this dissertation. Whereas its predictions of the perceptual assimilation of L2 contrasts by experienced listeners are similar to those posed about equivalence classification in the SLM and perceptual assimilation by naïve listeners in the PAM, the three models differ in one key aspect: PAM-L2 adds the phonological level of both L1 and L2 to judgments of L1-L2 similarity and dissimilarity; thus, perceptual assimilation can occur at the phonological, phonetic, or gestural/articulatory level.
86 This addition stems from the inclusion of L2 learners into this model who, unlike naïve listeners in PAM, have knowledge of the phonetic and phonological aspects of their L2. At the same time, this knowledge depends on their developmental stage and lexicon 14 acquired, making the phonological level a lexical-functional one where “listeners may identify L1 and L2 sounds as functionally equivalent (assimilated phonologically)”, which does not necessarily imply that “the associated phones are perceived as identical at the phonetic level” (Best & Tyler, 2007, p. 26). For example, such is the case of French /r/ > [ʀ], which American English learners of French assimilate to English /r/ > [ɾ] at a functional level. Late L2 learners, like naïve speakers, may also present difficulty in assimilating L2 contrasts which are not distinctive in their L1, especially if they have limited experience with the L2. However, as experience and familiarity with the L2 increases, so does the perception and production of the L2. PAM-L2 enumerates the following four possibilities for the perception of L2 contrastive sounds (Best & Tyler, 2007, pp. 28-30): 1. Only one L2 sound is perceptually assimilated to a given L1 phonological category, as in UC type. In this case, discrimination will have little difficulty. Alternatively, there exists the case in which the learner perceives an L2 sound as phonetically deviant from their L1 sound but yet phonologically and phonotactically similar on a lexical and functional level, and thus equates them phonologically. 2. Both L2 sounds are perceived as equivalent to the same L1 phonological category, but one is perceived as being more deviant than the other. This instance corresponds to the CG assimilation contrast. The good exemplar will be assimilated to the L1 category while it is estimated that, with L2 experience, the deviant exemplar 14 PAM-L2 considers that perceptual assimilation is more likely to succeed for listeners with limited L2 vocabulary; otherwise, incomplete perceptual learning before vocabulary expansion may give way to fossilization.
87 can move from a perceived phonetic variant of the good exemplar to a new phonological category. 3. Both L2 sounds are perceived as equivalent to the same L1 phonological category, but as equally good or equally poor examples of that category. In this case, it is an SC assimilation type, in which both L2 sounds will be assimilated to the L1 category and discrimination will be difficult. 4. No L1-L2 phonological assimilation. In this case, the L2 sounds will be uncategorized by the listener if they cannot be assimilated to any L1 phoneme but rather share characteristics of several L1 phonological categories. One limitation that the authors point out is that “some aspects of sensitivity to phonetic variation are related to similarities between nonnative stimuli and native speech patterns, but others reflect language-universal perceptual tendencies. The implications of these experience-tuned vs. universal phonetic sensitivities have not yet been fully resolved” (Best & Tyler, p. 18). We will see how Strange (2011) addresses this issue in the next section. 3.2.4 Automatic Selective Perception Model As a consequence of the models described in the previous sections, Strange (2006, 2011) developed the Automatic Selective Perception (ASP) working model to determine the mechanisms of speech processing that take place in the perception of L1 and L2, using neurobiological studies for the purpose. The focus is on adult naïve L1 listeners -category that also comprises beginning L2 learnersand on late L2 learners residing in a non-native country. Much like PAM, ASP is based on the direct-realist, ecological view of speech perception as “a purposeful, information-seeking activity whereby adult listeners
88 detect the most reliable acoustic parameters that specify phonetic segments and sequences in their native language (L1)” (Strange, 2011, p. 456). By this mechanism, adult L1 speakers resort to what she terms selective perception routines (SPRs) to detect relevant information for recognizing phonological sequences in their L1, which become automatic with the mastery of the language. In contrast, late L2 learners “must employ greater attentional resources in order to extract sufficient information to differentiate phonetic contrasts that do not occur in their native language” (p. 456). Therefore, L1 interference with L2 perception is seen as the attunement of L1 SPRs to the incorrect information in the L2 input. In this model, two modes of perception are described: the phonological mode and the phonetic mode. “These are “ways of perceiving” determined by an interaction of the listeners’ knowledge, purpose and intentions, the complexity of the stimulus materials, and the demands of the task to be accomplished” (Strange, 2011, p. 460). The phonological mode is employed by adult listeners to process continuous L1 speech, whether by speakers of the same variety or of dialects of the language familiar to the listener. The context-dependent phonetic variations are ignored in favor of the semantic message of the utterance, using automatic and robust SPRs even in nonoptimal conditions. The phonetic mode, on the other hand, is context-dependent and implies attentional focus to allophonic details and to those phonetic and phonotactic patterns necessary in their native dialect or language. It is also slower and may suffer in non-optimal conditions. Strange, Bohn, Trent, and Nishi (2004) and Strange, Bohn, Nishi and Trent (2005) studied the perceived similarity of German [u:] and [y] to American English vowels by naïve speakers of American English. Overall, the two vowel sounds were assimilated to their L1 [u:]; however, in citation-form /hVp/ contexts, [y] was classified as a poorer example of L1 [u:], while in sentence-
89 embedded /bVp/, /dVp/ and /gVk/ contexts, both German sounds were seen as good exemplars of L1 [u:], most likely because American English back rounded vowels are fronted in these contexts and become more similar to German front rounded vowels. Perception also depends on the design of experiment tasks: auditory salience 15 and perceptual salience 16 of the L2 sounds, memory and attention of listeners can all be targeted by the manipulation of the stimulus materials and the type of task employed in the experiment (see the Tetrahedral Model for Speech Perception Experiments by Strange, 1992). When the task and the stimuli are simple (citation words) and instructions direct listeners to pay attention to certain phonetic aspects, both naïve listeners and L2 learners can distinguish non-native L2 contrasts and determine similarities and dissimilarities between L1 and L2 sounds. However, as the complexity of the task increases, e.g. listeners must understand the semantic message of the utterance, so does the cognitive demand, and performance may suffer as listeners may resort to their L1 SPRs. Indeed, “as the complexity of the discrimination task increases, performance outcomes begin to reflect not only basic auditory sensory capabilities but increasingly the cognitive processes involved in categorization (including implicit labeling of presented stimuli)” (Strange & Shafer, 2008, p. 161). Even when listeners have enough experience in the L2 to have established L2 SPRs, these still may not be as automated as their L1 SPRs as “immersion experience alone may not be sufficient for L2 learners to develop and automate these SPRs” (Strange, 2011, p. 464). Instead, training for L2 learning is suggested as it can lead to 15 “The magnitude of the obligatory physiological response to a change from one to another contrasting lexical segment, tone, or sequence of segments in a normal hearing listener” (Strange, 2011, p. 458). 16 “Behavioral and physiological response “strength” that varies as a function of linguistic experience, as well as experimental manipulations of attentional focus” (Strange, 2011, p. 458).
90 the development of new SPRs to improve the detection of the most reliable cues in the L2. 3.3 Current Study What about the perception of two dialect allophones of the same phonological category? Initially, native speakers familiar with both L2 variants would assimilate both allophones to the same category (SG type) while native speakers unfamiliar with one of the L2 dialects would also assimilate both allophones to one single category but with differences in goodness-of-fit (CG type). As we saw at the beginning of this chapter, preference is given to the unmarked features of a language; thus, the marked allophone would be perceived as more deviant than the unmarked one. Native listeners may successfully discriminate two allophones “when the experimental task allows reliance on pre-existing mental representations of sounds” (Celata, 2007). Nevertheless, the perception of an allophonic contrast is generally less accurate than the perception of a phonemic contrast (Boomershine, Hall, Hume, & Johnson, 2008). The key point is that, in both types of assimilation, PAM and PAM-L2 consider the two contrastive sounds to be phonologically distinctive, but fail to determine how perception is carried out when the two sounds are allophones of one single category. The question pertinent to this study is how L1 listeners identify two dialect variants of the same L2 category, one of which is unfamiliar to them. The studies reviewed in this chapter suggest that the L1 dialect of the listener exerts a great influence on their discrimination and categorization of L2 segments. Thus, this study tested the perception of two dialect variants of implosive /s/ in Spanish, namely, aspiration [h] found in WAS and sibilance [s] characteristic of CS,
91 by native speakers of two American English dialects, GAE and AAE, whose L1 dialects differ in the use of final /s/ as a marker of plurality and verb agreement. 3.3.1 Predictions and Research Questions Listeners in this study are L2 learners of Spanish who, even at the elementary stage, are presumed to know that /s/ is phonologically distinctive in the L2 as it differentiates plural nouns from singular nouns and second-person verbs from thirdperson verbs in the present tense. What they ignore, especially when contact with an aspirating dialect has never occurred, is that [h] is a legitimate allophone of /s/ in certain Spanish dialects and it marks the same distinctions as [s]. A similar sound to the allophone under study [h] occurs in English as a contrastive sound in initial position but not as a legitimate variant of /s/ in implosive position, as is the case in WAS (and other varieties of Spanish). Even when aspirated /s/ and English [h] are acoustically and articulatorily similar to each other, listeners may not assimilate these two sounds, precisely due to the phonotactic biases of their L1. Can these listeners extract enough information from aspiration to identify it as functionally equivalent to [s]? They key may be in their experience with the L2 and their familiarity with the L2 dialectal feature. In this case, since we studied listeners of elementary and intermediate Spanish (levels 1 and 2) with no experience with aspirating dialects, the answer may be they cannot. It is in these levels where we can best determine if L1 dialect plays a role in perception. Thus, our first research question is as follows: Q1: Do AAE and GAE listeners differ in their ability to identify WAS aspiration of final /s/ in plural nouns and second-person verbs?
92 Contrastively, syllable-final /s/ is found as a legitimate sound in both GAE and AAE. Following the cross-language models reviewed, we predict that GAE listeners will assimilate CS [s] to GAE [s]. However, AAE speakers can regularly omit final /s/ from plural nouns and third-person verbs and, as we saw in the studies by Johnson (2005), and de Villiers and Johnson (2007), at least AAE children do not understand /s/ as an agreement marker, while GAE children do. Does this transfer to adulthood and to the perception of L2 features? Consequently, we pose our second research question: Q2: Do AAE and GAE listeners differ in their ability to identify CS sibilance in final /s/ in plural nouns and second-person verbs? Additionally, we have seen how context can affect the perception of stimuli and can render variation of results. For this reason we also wanted to explore how the syntactic and the phonetic contexts of the target variants can influence the perception of aspiration and sibilance. Our third research question is as follows: Q3: How do syntactic and phonetic contexts influence stimuli perception? Finally, we have seen that as experience with an L2 increases so does the identification and categorization of L2 sounds and contrasts. In Schmidt’s study (2011), there was no significant difference in [h] identification accuracy for level 1 and 2 listeners, it was not until level 3 that listeners began to identify [h] as a legitimate realization of implosive /s/, and not until level 5 that they performed similarly to native Spanish speakers. In this current study, we included listeners of elementary (level 1) and intermediate (level 2) Spanish of two different L1 dialects. In spite of not having enough experience with the target language, does proficiency level
99 (3PV) Tiene terreno en el campo Está tomando mucha verdura (He has land in the countryside) (He is eating a lot of vegetables) Figure 12 below shows the waveforms of WAS and CS /t/ after vowel: Figure 12. WAS and CS /t/ after vowel Target words ending in the morphological marker –s before /t/ were: (PN) Digo colas torpemente Digo amigos torpemente (I say tails awkwardly) (I say friends awkwardly) (2PV) Deberías tener más cuidado Necesitas tiempo para pensar (You should be more careful) (You need time to think) Table 13 displays the waveforms for /t/ after aspiration and sibilance: Figure 13. WAS aspiration and CS sibilance before /t/ From the 10 phonetic contexts that followed the target words, 6 of them have identical counterparts in English (/p, k, m, n, l, V/) in terms of place and manner of Time (s) 1.515 1.58 -0.2118 0.2131 0 Time (s) 1.014 1.087 -0.1486 0.1606 0 Time (s) 1.274 1.372 -0.08166 0.08466 0 Time (s) 0.8299 0.9571 -0.2421 0.218 0
100 articulation, while the remaining 4 (/b, d, g, t/) have similar but not identical counterparts in English. Since stimuli consisted of sentences, we need to consider a few allophonic variations that occur in Spanish due to the influence of the preceding sounds in connected speech. Voiced stops /b, d, g/ in word-initial position preceded by vowel or /s/ become voiced approximants [β, ð , ɣ ] (Garrido Almiñana, Machuca Ayuso, de la Mota Gorriz, 1998). This is true for CS after vowel and /s/ and for WAS after vowel. When WAS aspiration precedes these voiced stops, they become fricatives [v, ð, x]. Additionally, while /t/ is an alveolar stop in English, it is a dental stop in Spanish (see Table 3 below). The rest of the phonemes share place and manner of articulation with their English counterparts; however, WAS voiceless stops carry post-aspiration, while nasals and the lateral sound are geminated. Table 3 Dialectal allophones of Spanish consonants in word-initial position after [s] and [h] CS WAS Place Manner Place Manner b β bilabial approximant v labiodental fricative d ð dental approximant ð interdental fricative g ɣ velar approximant x velar fricative p p bilabial stop p h bilabial stop t t dental stop t h dental stop k k velar stop k h velar stop m m bilabial nasal h m.m bilabial nasal n n alveolar nasal h n.n alveolar nasal l l alveolar lateral h l.l alveolar lateral V V glottal open hV glottal open
101 4.2.1 Recording The 4 sets of sentences were recorded twice by four speakers of CS (two males, two females) and four speakers of WAS (two males, two females) at a 44.1 kHz and 16 bps sampling rate in a recording booth at the Phonetics Laboratory of the University of Seville (Spain), using a Marantz Professional PMD671 solid-state recorder and a Shure SM48 microphone, under the presence of the experimenter. Speakers were instructed to read as naturally as possible, as if they were talking to a friend at a normal conversational rate. Originally, this set of stimuli was added three levels of noise (30dB, 55dB, 65dB) with Akustyk for Praat (Plitcha, 2010), to be used in the pilot test only. With the pilot test we explored the extent to which aspiration and sibilance were subject to disruption by noise, as we will see in Section 5.1.1.4. As evidence suggested, at least the GAE listeners obtained native-like scores for sibilance in all noise conditions, and their identification of aspiration was generally less accurate than that of sibilance but increased with level of proficiency, generally despite noise condition. Thus, the effect of noise here may be confounded with proficiency level. Nevertheless, the evidence that was most interesting for our purposes came from the AAE listeners. Therefore, we eliminated the noise factor in our current experiment for this dissertation and focused on the performance of lower-level participants of both L1 dialects. 4.3 Procedure For this current experiment, we selected the four best exemplars out of the eight speakers from our corpus: one female and one male speaker per L2 dialect. CSM1 and WASM1 were discarded due to intonation and speech rate deviations in comparison with the rest of speakers. Subsequently, speakers CSF1 and WASF2 were
102 eliminated from this present study in order to have one exemplar speaker from each gender and L2 dialect. The resulting 320 sentences were converted to mp3 format and included in a test mounted in the experimenter’s university webpage supporting HTML, PHP, and MySQL. The experimental task was devised by a computer technician specifically for the purpose. This application developed in the Laboratory of Phonetics at the University of Seville is used to gather a great amount of data, which would not be possible otherwise. Participants must go through five sections to complete the test. The first section gathers general information for the sampling attributes of the experiment, such as age, gender, etc. The second section gathers linguistic information about the listeners’ L1 use in informal conversation, and aims at compiling data on L1 dialect use. The third section is a training exercise in which five samples of stimuli from the corpus appear, one at a time, with their corresponding solution. In the fourth and final section of the experiment, participants reproduce each individual stimulus twice before choosing an answer. The number of stimuli that appear in each test is fixed (60 sentences, in this case) but the order and type of stimulus is randomly presented by the application. Finally, in the last section the application asks the participant to confirm the submission of the results, and thanks the listener for their participation. The experimental task is programmed in a within a single webpage; therefore, during the completion of the task, the participant does not browse from one page to the next. This simple detail makes participants unable to use the browser to go back or go next, and lose the information provided up to that moment. Additionally, the Javascript functions that manipulate the webpage are invisible to the user, even to experienced programmers.
103 Thus, this web application was able to originate a different test for each of the participants in the study. Thanks to this randomness, we can have an unlimited number of stimuli in the corpus because they all have the same probability to be pooled by the program. Furthermore, the application records the listener’s reaction times to each stimulus. Figure 14 below shows a screenshot from the experimental task. Figure 14. Screenshot of the identification task in the current study Experiments were generally run in one of the following two settings: language classroom or at home. AAE listeners of both elementary and intermediate levels of Spanish took the test in language laboratories at a US university where the experimenter was present. These laboratories had 30 computer stations where students completed the identification task individually, using headphones. GAE listeners of both elementary and intermediate levels took the test at their US institutions, under the direction of their instructors. No monetary compensation was given to the participants but they were granted extra credit for their participation.
104 Participants received written instructions in English that they would listen to sentences in Spanish and would need to select the sentence that they heard from the two forced-choice written options given. When the target word was a plural noun or a second-person verb, the alternative option offered the same sentence with the same target word without the final –s, and vice versa. For example, if the sentence played was Nunca comes nada dulce (You never eat anything sweet), the two options given were Nunca come nada dulce (He/She never eats anything sweet) and Nunca comes nada dulce (You never eat anything sweet), so the correct option could not be inferred from reading the sentences alone. These instructions were presented in an informed consent document (see Appendix D) and repeated in the test itself. Listeners performed a self-paced sentence identification task in which each participant listened to a separate set of 60 sentences randomly chosen from the corpus, with no feedback provided. As a training method, the test played five sentences and showed the correct answer, so that they became familiar with the task and could adjust volume settings. Participants had to listen to each sentence twice in order to proceed to the next one, and were allowed unlimited time to complete the test, although a total duration of 15-20 minutes was estimated. Additionally, participants were required to sign the aforementioned informed consent form and fill out two initial questionnaires included in the test: a Language Background Questionnaire (see Appendix E) and a Spoken English Questionnaire (see Appendix F), both aimed at making a detailed profile of the listeners for classification and interpretation of the findings in this study.
105 4.3.1 Language Background Questionnaire The language background questionnaire gathered information about age and gender of the participants, birthplace of the participants and their parents or guardians, languages spoken at home and outside home, accent or dialect spoken by the participants, whether the participants had ever stayed in a Spanish-speaking country for over 3 months, dialect of Spanish currently exposed to, years of Spanish instruction, Spanish level (this question was later excluded; level was determined by section attending at the time of testing, as reported earlier), other languages in which participants were fluent, and hearing or speech disorders reported. 4.3.2 Spoken English Questionnaire The Spoken English Test listed 13 questions designed to test for dialectal features included among the most stable and rising in AAE speech (Wolfram, 2004). Specifically, the test looked into: copula absence + V-ing, habitual be + V-ing, thirdperson –s absence, copula absence + adjective, negative inversion, possessive they, existential they, noun plural absence, resultative be done, cluster reduction before vowel, regular past tense –ed deletion before vowel, and r-lessness before vowel. From the 122 AAE listeners, 49% of them reported using some of the features in this test. Table 4 below shows the percentages reported by these listeners for each of the elements tested.
106 Table 4 Percentage of AAE usage reported by AAE listeners Percentage of AAE listeners copula absence + v-ing 21 habitual be + v-ing 5 third person –s absence 12 copula absence + adjective 23 negative inversion 34 possessive they 18 existential they 32 noun plural absence 5 resultative be done 18 cluster reduction before vowel 39 -ed deletion before vowel 13 r-lessness before vowel 48 4.4 Data Coding Data were gathered by the program immediately after submission into an Excel sheet displaying all information submitted by each participant, i.e., their answers to the linguistic background and spoken English questionnaires and the 60 stimuli they listened to in order of appearance together with their score (1 = correct, 0 = incorrect). The following information was entered into a file using IBM SPSS Statistics 20: listener ID, age of listener, gender of listener (1 = female; 2 = male), other languages spoken at home (1 = yes; 2 = no), other languages spoken at school (1 = yes; 2 = no), dialect of listener (1 = GAE; 2 = AAE), stay in a Spanish-speaking country (1 = yes; 2 = no), years of instruction (1 = less than 1 year; 2 = 1-3 years; 3 = 3-5 years; 4 = more
107 than 5 years), level of instruction (1= elementary; 2 = intermediate; 3 = highintermediate; 4 = advanced; 5 = proficiency). At this point, participants who reported speech or hearing disorders were excluded so this variable was no longer present. We then included the following characteristics for each stimulus in order of presentation (1-60): speaker dialect (1 = WAS; 2 = CS), speaker gender (1= female; 2 = male), sentence type (1 = 3PV, 2 = 2PV, 3 = SN, 4 = PN), phonetic context ( 1 = [b]; 2 = [d]; 3 = [g]; 4 = [k]; 5 = [p]; 6 = [t]; 7 = [m]; 8 = [n]; 9 = [l]; 10 = [V]), score (1 = correct; 2 = incorrect), reactions times, and place of testing. On a separate SPSS file, we also included accuracy percentages (0-100%) per participant of [h] perception (aspiration), [s] perception (sibilance), [V] perception in WAS sentences and [V] perception in CS sentences, dialect and level to which they belonged (1 = AAE1, 2 = AAE2, 3 = GAE1, 4 = GAE2), and place where they took the test (1 = computer classroom; 2 = home). Additional SPSS files were created for the classification of stimuli according to their acoustic characteristics. For voiceless stops, we indicated gender of speaker (1 = female, 2 = male), L1 dialect of speaker (1 = WAS, 2 = CS), type of phonetic context (1 = sk, 2 = sp, 3 = st, 4 = k, 5 = p, 6 = t), duration of preceding vowel (ms), closure duration values (ms), and Voice Onset Time values (ms). For fricative sounds, we also indicated gender and L1 dialect of speaker as in the previous file, type of phonetic context (1 = sb, 2 = sd, 3 = sg, 4 = h, 5 = s), duration of previous vowel (ms), fricative intensity (dB), duration (ms), Center of Gravity (Hz), dispersion (Hz), kurtosis, skewness, and spectral peak (Hz). .
108 4.5 Statistical Analysis Accuracy results were obtained by dividing the number of correct answers per listener and variable by the number of stimuli they received from each variable. Thus, a participant that listened to 30 sentences where aspiration was present, and identified 20 of these sentences correctly, had an accuracy score of 66.67%. A general level of significance of p < .05 was assumed for all tests. However, when applicable, levels of significance were also expressed as p < .01, p < .005, and p < .001. We initially performed a Kolmogorov-Smirnov test to check for normal distribution. As not all groups showed normal distribution, we applied Spearman rank order correlations to determine the correlation between participant characteristics and variables tested. The initial characteristics we tested were i) stay in a Spanishspeaking country, ii) languages other than English at home, iii) languages other than English at school, iv) languages other than English spoken. Subsequently, participants who displayed influential factors were removed from the results. In the absence of normal distribution in most of the groups, we then proceeded to run non-parametric tests to analyze the results. For each group, we ran Wilcoxon tests (non-parametric equivalent to paired two-sample t-tests) comparing their intra-group performance in aspiration and sibilance first, and then between vowel identification in WAS and CS sentences. We then ran a Kruskal-Wallis test (non-parametric equivalent to ANOVA) to compare performance across all groups, with subsequent Mann-Whitney tests (nonparametric equivalent to unpaired two-sample t-tests) between pairs of groups. A third analysis was directed towards the syntactic context in which the target words were embedded and the phonetic contexts that followed [h] and [s], and finally, the years of instruction in Spanish that each group received. We explored the overall performance
115 4.7.1 Pilot Study As a preliminary study, we used 56 sentences from the set of stimuli that was added noise (8 speakers): seven sentences per speaker dialect and target word type. In this case, we only used /p, t, k/, /b, d, g/ and vowel as following sounds for L2 learners. Initially, the pilot test had 80 stimuli, as we also included /m, n, l/ as following sounds, but reduced the number of stimuli due to the duration of the test, which was discouraging for participants given the design of the platform in which it was mounted. Figure 17 shows a screenshot of the task. Figure 17. Screenshot of the identification task in the pilot study 4.7.1.1 Speakers Speakers in this pilot test were the four speakers in our current study with the addition of another set of four speakers (one male and one female per L2 dialect). The additional four speakers were a male speaker (CSM1) from Toledo (Castile), a female
116 speaker (CSF1) from northern Cordova (Northern Andalusia, at the border with Castile), who retained sibilance, a male speaker (WASM1) from Seville, and a female speaker (WASF2) from Seville (Western Andalusia). Three of them had higher-level education 18 (M age = 28.25), with the exception of speaker WASM1. 4.7.1.2 Listeners Twenty-four native Spanish listeners participated in this pilot identification task with the initial 80 stimuli, while 53 L2 learners of Spanish participated in the task with the final 56 sentences, either under the presence of the examiner or another trained instructor, or at home, during the spring semester of 2012. These listeners were classified according to their reported proficiency in the L2: Levels 1 and 2 were labeled under “low”, listeners of Levels 3 and 4 were named “mid” and listeners of Level 5 were termed “high”; and according to L1 dialect: AAE and GAE. Based on the findings from this pilot test, our current experiment only focused on elementary (Level 1) and intermediate (Level 2) L2 learners. We also examine these preliminary results in the following chapter. 18 Speakers CSM1 and WASF1 were pursuing their B.A. at the time when the recordings were carried out.
117 CHAPTER 5 RESULTS In this chapter, we present the results obtained from the pilot test and the subsequent identification task for this current study to answer the research questions stated at the end of Chapter 2. We first introduce the preliminary results for the native Spanish listeners and the L2 learners that participated in the pilot test, and then provide a review of the performance of the participants in the current identification task, in terms of accuracy identification of aspiration and sibilance in general, also according to the amount of instruction received by the listeners, and subsequently according to the syntactic and the phonetic contexts of the target words. Acoustic analyses are subsequently provided in search of an explanation for the results. 5.1 Results of Perception 5.1.1 Pilot Study 5.1.1.1 Native listeners Twenty-four native listeners (NL) of WAS participated in the pilot experiment. Their lowest accuracy score was for aspiration (M = 91.67, SD = 9.58), while their highest score was for sibilance (M = 100, SD = 0), with percentages of M = 98.51, 97.62; SD = 2.96, 5.45 for vowel in WAS sentences and vowel in CS sentences, respectively. Wilcoxon tests showed that the perception of sibilance was significantly higher than that of aspiration for this group of listeners [Z = -3.21, p = .001]. Taking into account that the stimuli was presented in noise, as we will see in Section 4.1.4,
118 aspiration seems to be vulnerable to disruption by noise, at least for these group of NL of Spanish. Precisely, it was at all levels of noise that NS presented differences (65dB: Z = -2.07, p < .05; 55dB: Z = -2.94, p < .005; 30dB: Z = -2.49; p < .05) between aspiration and sibilance. In fact, NL21, NL23 and NL24 showed remarkably lower scores in the identification of aspiration. This could be the main reason for such results. Finally, their identification of sentences ending in vowel was similar in both WAS and CS conditions. 5.1.1.2 L2 listeners Fifty-three L2 learners of Spanish, AAE low (n = 24), GAE low (n = 6), GAE mid (n = 19), GAE high (n = 4), took part in this initial test. Twenty-one had stayed in a Spanish-speaking country at the time of testing, none of which were AAE listeners, while 30 had not. Overall performance was as follows: accuracy in aspiration was markedly poorer (M = 32.28, SD = 28.47) than in all other conditions, followed by sibilance (M = 82.92, SD = 29.14), CS vowel (M = 83.42, SD = 27.76), and WAS vowel (M = 85.31, SD = 20.84). Table 5 shows the identification percentages obtained by each listener group per variable. Table 5 Mean accuracy percentages for all groups and variables aspiration sibilance WAS vowel CS vowel M SD M SD M SD M SD group AAE low 25.32 25.88 62.64 33.64 76.78 22.81 73.81 25.33 GAE low 21.53 18.15 100 0 83.93 23.60 96.43 4.12 GAE mid 32.98 24.65 99.56 1.91 92.50 11.23 95.36 8.44 GAE high 86.81 11.37 100 0 97.14 3.91 91.43 11.74
119 5.1.1.3 Preliminary analysis After applying a Kruskal-Wallis statistical test, we found significant differences across all L2 groups for aspiration [χ 2 (3) = 12, p < .01]; sibilance [χ 2 (3) = 28.7, p < .001]; WAS vowel [χ 2 (3) = 8.03, p < .05], and CS vowel [χ 2 (3) = 10.33, p < .05]. Wilcoxon statistical tests applied to each L2 group individually to extract intra-group performance revealed that the perception of sibilance was also significantly higher than the perception of aspiration for all L2 learner groups: [AAE low: Z = -3.64, p < .001; GAE low: Z = -3.08; p < .005; GAE mid: Z = -5.56, p < .001; GAE high: Z = -2.48, p < .05]. We then proceeded to analyze how AAE listeners compared with the three GAE groups in terms of aspiration and sibilance. For this purpose, we first considered whether AAE listeners who expressed overt AAE features in the Spoken Language Questionnaire (n = 14) and those who did not (n = 10) showed evidence of similar or different identification of aspiration and sibilance. In this case, there were no statistically significant differences (aspiration: U = 47, p = .19; sibilance: U = 39.5, p = .07). A Mann-Whitney test revealed that the perception of sibilance between AAE and GAE low-level groups was significantly higher for GAE listeners (U = 18, p < .005) but no differences were found for sibilance between these two groups (U = 71, p = .96). The same statistical test also revealed that the perception of sibilance between AAE listeners and GAE mid was also significantly higher for GAE listeners (U = 61, p < .001), but similar between both groups for aspiration accuracy (U = 175.5, p = .20). Upon comparison with GAE high listeners, accuracy for both aspiration (U = 1,
120 p < .001) and sibilance (U = 12, p < .05) was again found to be significantly higher for the GAE listeners. Subsequently, we analyzed the performance between the groups of GAE listeners. A comparison between GAE low and GAE mid showed that the perception of aspiration and sibilance was similar between both groups (U = 42.5, p = .36; U = 54, p = .88). When comparing GAE mid with GAE high, it was evident that the accuracy of aspiration identification (U = 2, p = .001) was higher for the most proficient learners but similar between the two groups for sibilance (U = 36, p = .91). Likewise, the performance between GAE low and GAE high listeners was also significantly favorable to the second for aspiration only (U = 0, p = .01), but identical between both for sibilance (U = 12, p = 1). We then compared the results of those who had stayed in a Spanish-speaking country and those who had not. A Mann-Whitney test showed that all differences were statistically significant, with higher accuracy for those who had stayed in a Spanish-speaking country before: [aspiration (U = 194.5, p = .01), sibilance (U = 147, p < .001), vowel WAS (U = 177, p < .005), and vowel CS (U = 176, p < .005). However, we have to consider that none of the AAE speakers (low-level) had ever stayed in a Spanish-speaking country while 13 out of the 39 GAE speakers did (at midand high-levels, but not at low-level). Finally, we compared the performance in aspiration and sibilance identification between L2 learners and NL. Mann-Whitney tests revealed that the identification of aspiration by NL was significantly better than that by the rest of the groups of L2 learners except for the GAE high group (AAE low: U = 3, p < .001; GAE low: U = 0, p < .001; GAE mid: U = 11.5, p < .001; GAE high: U = 31, p = .29). In the perception of sibilance, however, no significant differences were found
121 between GAE listeners and NL, but AAE low seemed to be significantly less accurate than NL (U = 79, p < .001). These preliminary results indicate that identification accuracy of aspiration for GAE listeners gradually increased with level of proficiency, rendering statistically significant differences between midand high-level learners. Native-like performance for GAE listeners was achieved at high-level of Spanish, with no differences in either aspiration or sibilance identification between these listeners and NL. Likewise, all L2 learners in this pilot study showed native-like performance in their identification of sibilance, but not in aspiration. While these results were predictable, a striking finding was the fact that significant differences in the identification of sibilance were found between GAE low and AAE low listeners in favor of the GAE listeners, suggesting that L1 dialect features may influence perception in this case. 5.1.1.4 The effect of noise As stated in the description of the stimuli employed for the pilot test, three levels of noise were added to the sentences in the task: 65dB, 55dB, and 30dB, the influence of which we analyze here. As we can see in Table 6 below, the L2 listeners’ overall identification of sibilance in the three conditions was similar [χ 2 (2) = .21, p = .90] while their identification of aspiration as a group was conditioned by the level of noise [χ 2 (2) = 9.7, p < .01]. Table 6 Overall identification percentages of aspiration and sibilance in the three noise conditions noise65dB noise55dB noise30dB M SD M SD M SD aspiration 24.53 33.43 36.16 36.06 36.14 30.31 sibilance 83.49 29.80 83.02 31.23 82.26 32.02
122 Figure 18 shows the performance of both low-level groups. At first sight, the figure already indicates what statistics can corroborate: GAE low listeners significantly outperformed AAE low listeners only in the identification of sibilance for the three levels of noise (65dB: U = 27, p < .05; 55dB: U = 30, p <.048; 30dB: U = 27, p <. 05). Figure 18. Identification percentages of aspiration and sibilance for AAE and GAE low-level groups according to noise level GAE mid (Figure 19) was the only group for which noise was an influential factor in their identification of aspiration (χ 2 (2) = 19, p < .001), which increased as level of noise decreased. Additionally, their performance was similar to that of GAE low listeners for both sibilance and aspiration at the three levels of noise, and significantly more accurate than that of AAE low listeners for sibilance (65dB: U = 92.5, p < .001; 55dB: U = 95, p < .001; 30dB: U = 85.5, p < .001) and for aspiration only at 30dB (U = 111.5, p < .005).
123 Figure 19. Identification percentages of aspiration and sibilance for GAEmid listeners according to noise level In comparison with GAE high listeners, both groups performed similarly for sibilance but GAE mid identified aspiration significantly more poorly than GAE high listeners at the three levels of noise (65dB: U = 8, p < .01; 55dB: U = 6.5, p < .01; 30dB: U = 2, p < .005). Figure 20. Identification percentages of aspiration and sibilance for AAEhigh and NL according to noise level There were no significant differences between GAE high and NL for any level of noise or target L2 feature, i.e., the performance of GAE high was similar to that of NL of Spanish (Figure 20). So far, these analyses confirm what was stated in the previous section. All GAE listeners and NL performed similarly in the identification of sibilance, in spite of
124 noise level. GAE listeners at lowand midlevel also performed similarly in the identification of aspiration, but it was not until high-level that GAE showed a significant improvement in the identification of aspiration, similar to that of NL, also regardless of noise level. AAE listeners, on the other hand, also identified aspiration in a similar manner to GAE low and GAE mid participants, with the exception that, at the lowest level of noise (30dB), GAE listeners of mid-level performed significantly better. The main difference here is that AAE listeners identified sibilance significantly less accurately than all groups of GAE listeners, including their low-level counterparts, in all three noise conditions. Our aim is to investigate L2 dialect speech perception, particularly in the lowest levels of learning without exposure to the target features, when languagespecific patterns of perception are more likely to be reflected. Therefore, we deemed it necessary to discard the use of noise in our following experiment given that its effect was irrelevant for these groups of L2 listeners in the pilot test. 5.1.2 Current Study As stated in Section 4.5, we initially ran statistical analyses to determine whether certain characteristics played a role in perception: i) stay in a Spanishspeaking country, ii) languages other than English at home, iii) languages other than English at school, iv) languages other than English spoken. The only characteristic that we found to be a significant factor was the stay in a Spanish-speaking country, which was inversely correlated with the perception of sibilance (r = -.165, p < .05); therefore, these participants were excluded from the study.
131 The perception of sibilance was well above 80% for all groups. Figure 27 below shows the accuracy percentages for the four groups of learners together. Figure 27. Perception of sibilance by all groups of listeners As with aspiration, we analyzed performance (a) by dialect group, (b) by proficiency level, and (c) across groups: (a) The identification of sibilance was again significantly higher for AAE2 listeners than for AAE1 listeners (U = 821.5, p < .05), and also higher for GAE2 than for GAE1 listeners (U = 620.5, p = .05) 19 . In this case, level of proficiency proved to be a significant factor in the perception of sibilance. (b) Among listeners within the same level of proficiency, we found that GAE1 listeners identified sibilance significantly better than AAE1 listeners (U = 879.5, p = .001) and that GAE2 listeners also performed significantly better than AAE2 listeners (U = 477.5, p < .05). For sibilance, GAE listeners seemed to have an advantage over AAE listeners at the two levels of Spanish. 19 Although a level of significance of < .05 was assumed, this is on the verge of statistical significance.
132 (c) In fact, AAE2 listeners’ identification of sibilance was similar to that of GAE1 listeners’, i.e., L1 dialect prevailed over L2 proficiency in this case. In light of these findings, the identification of aspiration [h] and sibilance [s] across listener groups in this current study can be summarized as follows: as expected, aspiration was significantly less accurately perceived than sibilance by each individual group of listeners. Although AAE listeners outperformed GAE listeners, this was not statistically significant (U = 4 798, p = .43). What was significant was that intermediate level listeners performed significantly better than elementary level listeners (U = 4 220.5, p = .05). The case of sibilance identification was different. It seemed to work to the advantage of GAE listeners (U = 2 879, p <. 001), and of intermediate level listeners (U = 3 011, p < .001). For both L2 variants, elementary level listeners obtained the lowest scores overall; GAE1 for aspiration and AAE1 for sibilance, which coincides with the findings from our previous pilot study. Order of Stimuli At this point, we analyzed how the order of presentation of the stimuli in the identification task affected the identification of the stimuli. In general, the order of stimuli did not have a correlation with the identification of the target stimuli. Nevertheless, although weak, some correlations were found. There was a positive correlation between the identification of aspiration in second-person verbs for GAE1 listeners (N = 271, r = .119, p = .05). Likewise, for GAE2 listeners, there was also a positive correlation between their identification of WAS third-person verbs and the order of the test. This implies that these listeners identified these tokens worse as the order of the test progressed. For AAE1 listeners, There was a negative correlation between their identification of WAS third-person verbs (N = 733, r = -.113, p < .005),
133 as well as CS third-person verbs (N = 758, r = -.097, p < .01) and CS singular nouns (N = 686, r = -.121, p < .005). This implies that, as the order of the stimuli progressed, their identification of these tokens improved. Finally, there was no significant correlation for the AAE2 group of listeners. 5.1.2.3 Results by years of instruction In this section, we explore the effect of number of years of instruction on perception, although we should take into account that the type of such instruction was not measured in this experiment. As stated in Section 3.4, years of instruction were classified into four groups (less than 1 year, 1 to 3 years, 3 to 5 years, and more than 5 years). Kruskal-Wallis tests revealed that the number of years of instruction was not significant for any of the groups of listeners individually (AAE1: χ 2 (3) = 3.25, p = .36; AAE2: χ 2 (2) = 4.10, p = .13; GAE1: χ 2 (2) = 1.9, p = .39; GAE2: χ 2 (3) = 2.14; p = .54). Nevertheless, as shown in Table 7, the perception of aspiration decreased with amount of instruction for AAE1 listeners and increased with years of instruction for GAE1 listeners. Additionally, the identification of sibilance progressively increased with instruction for both groups at elementary level of Spanish. At the elementary level of Spanish, it was AAE listeners with less than 1 year of instruction who identified aspiration best and those with more than 5 years of instruction identified it worst. At intermediate level, AAE listeners with 3-5 years of instruction obtained the highest scores for aspiration and GAE listeners with less than 1 year of instruction obtained the lowest scores (taking into account that there were no participants with less than 1 year of instruction in the AAE group).
134 Table 7 Perception of aspiration and sibilance by years of instruction and listener group aspiration sibilance N M SD M SD AAE1 years of instruction < 1 20 15.10 15.05 80.33 18.73 1 to 3 59 13.57 17.39 80.83 18.79 3 to 5 14 9.66 15.24 88.61 11.24 > 5 4 5.88 11.77 93.18 13.63 AAE2 years of instruction < 1 0 - - - - 1 to 3 9 15.71 26.05 90.08 15.67 3 to 5 11 23.51 21.49 91.93 17.11 > 5 5 21.52 11.39 85.31 12.31 GAE1 years of instruction < 1 14 7.01 10.64 92.10 9.34 1 to 3 8 6.49 7.17 92.73 7.93 3 to 5 8 12.79 11.46 94.97 8.86 > 5 0 - - - - GAE2 years of instruction < 1 2 3.13 4.42 97.22 3.93 1 to 3 27 14.37 18.63 93.53 15.92 3 to 5 17 21.70 31.62 96.71 5.48 > 5 8 17.11 12.80 93.12 11.65 For sibilance, GAE1 listeners with 3-5 years of instruction were the most accurate and AAE1 with less than 1 year of instruction were the least accurate. At intermediate level, GAE listeners with less than 1 year of instruction obtained the highest scores (although n = 2), followed by those with 3-5 years of instruction. AAE listeners with more than 5 years of instruction obtained the lowest scores. Upon analyzing identification performance across groups and years of instruction, we observed that no significant differences in perception were found between the two GAE groups or between AAE2 listeners and either of the GAE groups. Statistical differences worth mentioning were observed between AAE1 listeners and the other three groups.
135 Perhaps the most salient finding was that, with less than 1 year of instruction, AAE1 listeners significantly outperformed GAE1 participants in the identification of aspiration (U = 83.5, p < .05), and GAE1 listeners outperformed AAE1 participants in the perception of sibilance, which was on the verge of significance (U = 87, p = .06). With 1-3 years of instruction, the three groups perceived sibilance significantly better than AAE1 listeners (GAE1: U = 128.5, p < .05; GAE2: U = 308, p < .001; AAE2: U = 155, p < .05). GAE2 listeners also outperformed AAE1 listeners in the perception of sibilance with 3-5 years of instruction (U = 64.5, p < .05). Additionally, AAE2 listeners also perceived aspiration significantly better than AAE1 listeners with 3-5 years of instruction (U = 33.5, p < .05). Therefore, as to the number of years of instruction, it seemed to particularly affect AAE1 listeners in comparison with the rest of the groups. Once more, at elementary level with less than 1 year of instruction, in which we could regard listeners as truly “naïve” in the language, we observed that AAE listeners significantly outperformed GAE participants in the perception of aspiration, while GAE listeners identified sibilance significantly better than AAE participants. 5.1.2.4 Results by syntactic context The perception of aspiration and sibilance is described in this section in relation to the syntactic context in which target words were embedded: content sentences with second-person verbs in the present tense (2PV), in which the morphological marker –s determines verb person, and carrier sentences with plural nouns (PN), in which the morphological marker –s determines plurality. Examples of such sentences are as follows:
136 (2PV) Nunca comes nada dulce (You never eat anything sweet) (PN) Digo perros por la tarde (I say dogs in the afternoon) Aspiration Overall, aspiration was significantly better identified by all listeners when target words were second-person verbs (14.12%) than plural nouns (13.49%) (Z = - 11.43, p < .001). In spite of this general trend, AAE2 listeners (Figure 28) perceived aspiration in verbs significantly better than in nouns (Z = -5.63, p < .001). Figure 28. Perception of aspiration in syntactic context by listener group Analyses in terms of (a) listener dialect, (b) level of proficiency, and (c) across groups revealed the following: (a) The perception of aspiration in 2PV between the two AAE groups was statistically similar (U = 118 835, p =.63), whereas aspiration in PN was significantly higher for AAE2 (U = 55 755, p < .001). GAE2 listeners performed significantly better than GAE1 in both contexts (2PV: U = 19 440, p < .001; PN: U = 41 310, p < .001).
137 (b) At elementary level of proficiency, AAE listeners outperformed GAE participants in both contexts (2PV: U = 81 480, p < .001; PN: U = 93 120, p < .001). At intermediate level, GAE2 performed significantly better in 2PV sentences (U = 47 925, p < .001) while AAE2 listeners were more accurate in PN sentences (U = 41 850, p < .001). (c) Both intermediate groups outperformed elementary groups of the opposite dialect in both contexts. In general, all groups identified the morphological marker –s when realized as aspiration better in second-person verbs than in plural nouns, with the exception of AAE2 listeners, who identified nouns more accurately. In fact, they outperformed the rest of the groups in the identification of the plurality marker –s. Likewise, GAE2 were the most accurate for the identification of the second-person verb marker –s. As we have seen before, GAE1 listeners were less accurate than the rest of listeners at identifying the morphological marker –s when realized as aspiration in both verbs and nouns. Sibilance The identification of sibilance by all listeners as a group was, on the contrary, significantly higher in carrier sentences (90.79%) than in content sentences (83.88%) (Z = -21.32, p < .001). This pattern was followed by each individual group of listeners, as shown in Figure 29.
138 Figure 29. Perception of sibilance in syntactic context by listener group Upon analyzing performance in terms of (a) listener dialect, (b) level of proficiency, and (c) across groups, we found: (a) Within the AAE group, intermediate listeners significantly outperformed elementary listeners in both contexts (2PV: U = 19 400, p < .001; PN: U = 61 837.5, p < .001). For GAE listeners, elementary level participants outperformed intermediate level listeners in 2PV context (U = 66 420, p < .001) but both performed similarly in PN context (U = 80 190, p = .81). (b) At elementary level of Spanish, GAE listeners’ perception was significantly more accurate than that by AAE participants in both contexts (2PV: U = 46 560, p < .001; PN: U = 58 200, p < .001). At intermediate level, both GAE and AAE listeners performed similarly in 2PV context (U = 62 775, p < .112) while GAE listeners were more accurate in PN sentences (U = 58 050, p = .001). (c) Across groups of listeners, the GAE2 group outperformed AAE1 listeners in both syntactic contexts (2PV: U = 26 190, p < .001; PN: U = 57 618, p <
139 .001), while the GAE1 group performed significantly better than AAE2 listeners in PN sentences (U = 31 500, p < .001) but similarly in 2PV contexts (U = 34 875, p = .15). In this case, all groups identified the morphological marker –s in plural nouns better than in second-person verbs. In general, all groups performed similarly, with the exception of AAE1 listeners, who were the least accurate at identifying the morphological marker –s in both nouns and verbs. To check if these patterns are also true for singular nouns and third–person verbs, we also analyzed these two types of target words in the two L2 dialects. As Table 8 shows, all groups identified singular nouns in CS sentences significantly better than third-person verbs (AAE1: Z = -9.85, p < .001; AAE2: Z = -5, p < .001; GAE1: Z = -5.48, p < .001; GAE2: Z = -7.35, p < .001), just as they identified plural nouns significantly better than second-person verbs in CS sentences. However, in the case of WAS sentences, we found that both AAE1 and GAE1 listeners identified singular nouns significantly better than third-person verbs (AAE1: Z = -9.85, p < .001; GAE1: Z = -5.48, p < .001), as opposed to their identification of WAS aspiration in second-person verbs, which was significantly more accurate than in plural nouns. GAE2 listeners were consistent in the sense that they also identified WAS verbs ending in vowel significantly better than nouns (Z = -7.35, p < .001). AAE2 listeners, however, identified third-person verbs significantly better that plural nouns this time (Z = -5, p < .001).
140 Table 8 Identification of WAS and CS third-person verbs and singular nouns WAS CS 3PV SN 3PV SN M M M M AAE1 89.63 93.40 87.60 92.71 AAE2 95.24 93.01 95.61 97.30 GAE1 97.12 98.28 95.79 98.83 GAE2 95.32 93.20 91.78 93.43 Therefore, all groups consistently identified CS nouns more accurately than verbs, whether ending in [s] or vowel, while WAS sentences rendered several outcomes. Both elementary-level groups identified verbs ending in [h] better than nouns, and nouns ending in vowel better than verbs. AAE2 listeners were the opposite: aspiration was better identified in nouns than in verbs, while verbs ending in vowel were better identified than nouns. GAE2 listeners identified verbs in both conditions more accurately than nouns. Reaction Times We also measured the reaction times (RTs) of the L2 listeners, i. e., how long they took to choose an answer after listening to each stimulus. Table 9 shows their RTs in milliseconds (ms) for each type of sentence by L2 dialect. Table 9 Reaction times (ms) in both conditions in the four syntactic contexts by group of listeners GAE1 GAE2 AAE1 AAE2 M M M M 3PV WAS 1142 1319 1307 1108 CS 1061 980 911 826 2PV WAS 1447 1711 2088 1684 CS 839 986 977 936 SN WAS 988 1231 1227 1120 CS 834 939 805 764 PN WAS 1551 1400 1753 1307 CS 867 965 1065 1155
243 APPENDIX F continued SPOKEN ENGLISH QUESTIONNAIRE 9. I need to pick her up at 6pm • She be done finish by then • She’ll have finished by then 10. How would you normally say (not write) “cold air”? • cold air • col’ air 11. How would you normally say (not write) “She picked us up”? • She picked us up • She pick us up 12. How would you normally say (not write) “Stop for a minute”? • Stop for a minute • Stop fo’ a minute
244 APPENDIX G WAVEFORMS AND SPECTROGRAMS OF WAS ASPIRATED STOPS VOWEL CLOSURE VOT VOWEL Figure 43. Waveform and spectrogram of [p h ] Time (s) 0 0.253 -0.2346 0.4084 0 Time (s) 0 0.253 0 8000 Frequency (Hz) 0.253026992
245 APPENDIX G (continued) WAVEFORMS AND SPECTROGRAMS OF WAS ASPIRATED STOPS VOWEL CLOSURE VOT VOWEL Figure X. Waveform and spectrogram of [t h ] Time (s) 0 0.2227 -0.3068 0.5 0 Time (s) 0 0.2227 0 8000 Frequency (Hz) 0.222704719
246 APPENDIX G (continued) WAVEFORMS AND SPECTROGRAMS OF WAS ASPIRATED STOPS VOWEL CL VOT VOWEL Figure 45. Waveform and spectrogram of [k h ] 0 0.2964 -0.2758 0.4757 0 Time (s) 0 0.2964 0 8000 Frequency (Hz) 0.296440398
247 APPENDIX H WAVEFORMS AND SPECTROGRAMS OF CS UNASPIRATED STOPS VOWEL [s] CLOSURE VOT VOWEL Figure 46. Waveform and spectrogram of [sp] Time (s) 0 0.2615 -0.1461 0.2188 0 Time (s) 0 0.2615 0 8000 Frequency (Hz) 0.261521535
248 APPENDIX H (continued) WAVEFORMS AND SPECTROGRAMS OF CS UNASPIRATED STOPS VOWEL [s] CLOSURE VOT VOWEL Figure 47. Waveform and spectrogram of [st] 0 0.1974 -0.2755 0.5 0 Time (s) 0 0.1974 0 8000 Frequency (Hz) 0.19743904
249 APPENDIX H (continued) WAVEFORMS AND SPECTROGRAMS OF CS UNASPIRATED STOPS VOWEL [s] CLOSURE VOT VOWEL Figure 48. Waveform and spectrogram of [sk] Time (s) 0 0.2989 -0.3017 0.3324 0 Time (s) 0 0.2989 0 8000 Frequency (Hz) 0.298889925
250 APPENDIX I WAVEFORMS AND SPECTROGRAMS OF WAS FRICATIZED STOPS VOWEL [v] VOWEL Figure 49. Waveform and spectrogram of [v] Time (s) 0 0.2286 -0.2923 0.3782 0 Time (s) 0 0.2286 0 8000 Frequency (Hz) 0.228596378
251 APPENDIX I (continued) WAVEFORMS AND SPECTROGRAMS OF WAS FRICATIZED STOPS VOWEL [ð] VOWEL Figure 50. Waveform and spectrogram of [ð] 0 0.1944 -0.3012 0.3456 0 Time (s) 0 0.1944 0 8000 Frequency (Hz) 0.194424755
252 APPENDIX I (continued) WAVEFORMS AND SPECTROGRAMS OF WAS FRICATIZED STOPS VOWEL [x] VOWEL Figure 51. Waveform and spectrogram of [x] Time (s) 0 0.1897 -0.2671 0.4518 0 Time (s) 0 0.1897 0 8000 Frequency (Hz) 0.189711427
259 APPENDIX L (continued) SPECTRAL SLICES OF CS SIBILANTS Figure 61. Spectral slice of [z] before [ð ] Figure 62. Spectral slice of [s] before [ɣ ] Frequency (Hz) 0 2.205·10 4 Sound pressure level (dB/Hz) 0 20 40 Frequency (Hz) 0 2.205·10 4 Sound pressure level (dB/Hz) -20 0 20