scieee AI-readable full text Open interactive document viewer

Da corrente acústica à palavra: estádios do processamento da percepção da fala

Tânia Fernandes

Full text

UNIVERSIDADE DE LISBOA UNIVERSIDADE DE LISBOAUNIVERSIDADE DE LISBOA UNIVERSIDADE DE LISBOA FACULDADE DE PSICOLOGIA E DE CIÊNCIAS DA EDUCAÇÃO FACULDADE DE PSICOLOGIA E DE CIÊNCIAS DA EDUCAÇÃOFACULDADE DE PSICOLOGIA E DE CIÊNCIAS DA EDUCAÇÃO FACULDADE DE PSICOLOGIA E DE CIÊNCIAS DA EDUCAÇÃO DA CORRENTE ACÚSTICA À PALAVRA: ESTÁDIOS DO PROCESSAMENTO DA PERCEPÇÃO DA FALA Tânia Patrícia Gregório Fernandes DOUTORAMENTO DOUTORAMENTO DOUTORAMENTO DOUTORAMENTO EM PSICOLOGIA EM PSICOLOGIAEM PSICOLOGIA EM PSICOLOGIA GERAL GERAL GERAL GERAL Tese orientada pelo Prof. Doutor Paulo Ventura Fernandes da Rocha e co-orientada pela Prof. Doutora Régine Kolinsky 2007 20072007 2007 II The present work was supported by a doctoral grant of Fundação para a Ciência e a Tecnologia – Ministério da Ciência, Tecnologia e Ensino Superior, Portugal, ref SFRH / BD / 12290 / 2003. Part of this research was also conducted under the Project POCI / PSI / 56901 / 2004 - “Visual phonology and auditory orthography”, funded by Fundação para a Ciência e a Tecnologia – Ministério da Ciência, Tecnologia e Ensino Superior – and European Community FEDER funding. III Palavras-chave: Segmentação da Fala; Aprendizagem de Línguas Artificiais; Automaticidade; Condições de Audição; Pistas sublexicais. Key Words: Speech Segmentation; Artificial Language Learning; Automaticity; Listening Conditions; Sublexical cues. American Psychological Association (PsycINFO Classification Categories and Codes): 2300 Human Experimental Psychology 2326 Auditory and Speech Perception 2343 Learning and Memory 2346 Attention IV V ACKNOWLEDGEMENTS This was probably the most difficult section of this thesis to write. It is difficult to put into words (and even more difficult in just a few words) all the generosity I have received, both professionally and personally, throughout this scientific process. Since the beginning of my passion for Cognitive Psychology I had two mentors; my first words of gratitude will be for them. To Professor Carlos Brito#Mendes, who showed to me for the first time what the “ black box ” is capable of, who is a constant inspiration in all my work in science, and who I will never forget. To Paulo Ventura, my supervisor, who I esteem immensely as a scientist and as a person, for having shared with me these four years of work with friendship, and for all his help in this scientific process: theoretical reasoning, preparation of experimental materials, data analysis, discussion of results. In difficult times, when I most needed a friendly and encouraging comment, he was always there. I would also like to thank Régine Kolinsky, my co#supervisor, who I deeply admire, for all her generosity in accepting me as a PhD student, for sharing with me her objective and sharp reasoning, for many stimulating discussions, for all her support in my development as a research scientist, and for being a friend. Paulo and Régine have been a vital and constant source of reference and stimulation, while at the same time always allowing me to pursue an unconstrained research path in the directions that best suited my intellectual thirst and curiosity. To José Morais, I am sincerely grateful for all the stimulating, skeptical, and constructive questions that have enormously enriched this work, and for the conversations about Portugal and Literature. I would also like to express my grateful acknowledgement to: Sven Mattys, for sharing with me his theoretical insights and for having welcomed me at the University of Bristol; and Nicolas Dumay, the Language Group of the Psychology Department of the VI University of Bristol, the members of the Unité de Recherche en Neurosciences Cognitives (UNESCOG) and of the Laboratoire de Psychologie Expérimentale of the Université Libre de Bruxelles, for having discussed my work. The writing of this thesis was enormously benefited by the comments of Paulo, Régine, José, and Sven. I would also like to thank to Cristina Carvalho, Sandra Fernandes and Luís Querido for all the help in running some of the experiments. To them and to Isabel Leite I am also thankful for their being always available and willing to help me. This thesis was partly supported by a doctoral grant awarded by the Fundação para a Ciência e a Tecnologia, Ministério da Ciência, Tecnologia e Ensino Superior, Portugal. Its support is here gratefully acknowledge. I would also like to thank to Centro de Psicologia Clínica e Experimental – Desenvolvimento, Cognição e Personalidade, at Faculdade de Psicologia e de Ciências da Educaça}o da Universidade de Lisboa, Portugal, and to the Department of Psychology of Universidade de Évora, as well as to all the students from the Universities of Lisboa and Évora who participated in the experiments reported in this thesis. And last, but not least, I would like to thank to my closest ones. To my parents, who have supported (both economically and emotionally) all my development, who made me who I am, who accept me unconditionally. To Miguel, Lina, and Daniel for believing in my capabilities, even when I did not. To Pedro who honor me with genuine and enduring care, love, and words of support. To Beatriz and Filipa, for being true friends even through difficult times, and for all the philosophical conversations and interest in this work. To Carmen, Joaquina, and Carmencita who always share great moments with me in Brussels. This work is dedicated to my mother, who has always stimulated my curiosity, insisted on the importance of learning, and who supports me in everything since the very beginning # from whom I have received so much for so little. VII To Isilda To IsildaTo Isilda To Isilda “Céptico como os cépticos, crente como os crentes. A metade que avança é crente, a metade que confirma é céptica. Mas o cientista perfeito é também jardineiro: Acredita que a beleza é conhecimento. (A pessoa bela tem um segredo. Descobriu algo).” Gonçalo M. Tavares, In “Breves Notas sobre Ciência” VIII IX ABSTRACT Until now, the weighting of general domain and speech-specific cues in speech segmentation was left largely unspecified. In the present work, the impact of different (qualitative and quantitative) listening conditions on the weighting of sublexical sources of information: segmental (transitional probabilities - TPs), suprasegmental (universal prosody and lexical stress) and subsegmental (coarticulation), was investigated in artificial language learning (ALL) settings, on the grounds of Mattys and colleagues theoretical framework. In Experiments 1-3B we evaluated the impact of physical noise on these cues. In Experiment 1, with intact speech, coarticulation overruled TPs. However, its role was highly modulated by signal-quality, while TPs were very resilient to noise. In Experiment 2, universal prosody like TPs was also highly resilient to physical noise, with the former prevailing over the latter, and driving the segmentation process. The impact of stress was rather different than the one of universal prosody. Stress pattern effects only emerged in degraded conditions (Experiment 2 and Experiment 3B). The impact of cognitive noise was evaluated in Experiments 4-6 through attention-load. In Experiment 3, the impact of cognitive noise on the weighting of TPs and coarticulation was in sharp contrast to the pattern found with physical noise. While coarticulation processing was largely unaffected by a reduction of attentional resources, TPs computation was penalized. In Experiments 5 and 6, we evaluated in a more fine-grained manner the impact of cognitive noise on TP-based segmentation. TPs computation is attentional-resources’ dependent. Yet, it occurs even when the AL is not the focus of attention and scarce resources are available, suggesting some automaticity of statistically-driven segmentation processes. In Experiments 7-9 we combined ALL settings with conventional experimental techniques, demonstrating that adult listeners are able to use on-line statistical information. Listeners actually treated statistical learning outputs as potential words - the output of ALL exhibits a lexical competition signature revealed by the inhibitory priming effect of novel neighbors (e.g., cathedruke) on lexical decisions to real words (e.g., cathedral) - but only when segmentation cues were congruent (Experiment 8). In Experiment 9, no effect of lexicalization was observed with incongruent cues. Thus, speech segmentation is largely the product of the available cues and of the listening conditions. Furthermore, these results suggest that current models of spoken-word recognition need to incorporate the role that congruency between segmentation cues play in both speech segmentation and word learning. XVI XVII TABLE OF CONTENTS PART I – THEORETICAL BACKGROUND 1. General Introduction ………………………………………………………………........ 3 1.1. From the continuous stream into discrete units of meaning …………………………... 5 1.1.1. In the beginning there was the word ……………………………………………….. 5 1.1.2. Crossroads into the speech stream ………………………………………………….. 10 1.1.3. T_issues and issues …………………………………………………………………. 18 1.2. Conventional Experimental Paradigms ……………………………………………….... 23 1.2.1. Gating ……………………………………………………………………….. 23 1.2.2. Word-spotting ………………………………………………………………………. 24 1.2.3. Eye-tracking ………………………………………………………………….. 25 1.2.4. Cross-modal priming ………………………………………….………………. 29 1.3. The Artificial Language Learning Paradigm …………………………………………... 31 1.3.1. ALL and Implicit Learning ……………………………………………………….... 33 1.3.2. ALL and Speech Segmentation …………………………………………………….. 35 1.4. The hierarchical organization frame of Mattys and colleagues ……………………….. 37 1.4.1. Experimental evidences …………………………………………………………….. 38 1.4.2. Developmental and cross-linguistic implications …………………………………... 39 1.5. An overview of the present thesis ……………………………………….……………….. 42 1.5.1. The weighting of different sublexical sources of information in speech segmentation ……………………………………………….………………... 43 1.5.2. The listening conditions …………………………………………….………………. 46 1.5.3. An on-line evidence of statistical speech segmentation ……………………………. 51 1.5.4. The nature of speech segmentation outputs ………………………….……………... 54 PART II – EMPIRICAL CHAPTERS 2. Statistical and Coarticulatory Cues to word boundaries ……………….………….... 61 2.1. Introduction …………………………………………………………….………….. 61 2.1.1. Coarticulation – a subsegmental cue …………………………………….………….. 63 2.1.2. TPs – a segmental cue …………………………………………………….……….... 64 XVIII Experiment 1. TPs and Coarticulation in Physical Noise 2.2. Method ……………………………………………………………………………………. 68 2.2.1. Participants …………………………………………………………………………. 68 2.2.2. Material …………………………………………………………………………….. 68 2.2.3. Procedure ………………………………………………………………………….... 74 2.3. Results and Discussion ………………………………………………………………….... 75 2.4. General Discussion ……………………………………………………………………….. 80 3. Universal and Speech-specific Prosodic cues in ALL ……………………………….. 89 3.1. Introduction ………………………………………………………………………………. 89 3.1.1. The interaction between Prosody and TPs in speech segmentation ……….………... 91 3.1.2. An overview of the experiments ………………………………………….……….... 96 Experiment 2. Domain-general vs. Speech-specific Universal cues in segmentation 3.2. Method ……………………………………………………………………………………. 101 3.2.1. Participants …………………………………………………………………………. 101 3.2.2. Material …………………………………………………………………………….. 102 3.2.3. Procedure ………………………………………………………………………….... 105 3.3. Results and Discussion ………………………………………………………………….... 106 Experiment 3A. Duration is an acoustic correlate of word primary stress in EP 3.4. Method ……………………………………………………………………………………. 114 3.4.1. Participants …………………………………………………………………………. 114 3.4.2. Material …………………………………………………………………………….. 115 3.4.3. Procedure ………………………………………………………………………….... 116 3.5. Results and Discussion ………………………………………………………………….... 117 Experiment 3B. Domain-general vs. Language-specific cues in segmentation 3.6. Method ……………………………………………………………………………………. 122 3.6.1. Participants …………………………………………………………………………. 122 3.6.2. Material …………………………………………………………………………….. 123 3.6.3. Procedure ………………………………………………………………………….... 124 3.7. Results and Discussion ………………………………………………………………….... 125 3.8. General Discussion ……………………………………………………………………….. 131 4. Statistical and Coarticulatory Cues in Cognitive Noise …………………………….. 141 4.1. Introduction ………………………………………………………………………………. 141 XIX 4.1.1. Statistics and coarticulation: the impact of physical noise …………………………. 142 4.1.2. The impact of cognitive noise …………………………………………………….... 143 Experiment 4. The impact of cognitive noise on the relative weighting of speech segmentation cues 4.2. Method ……………………………………………………………………………………. 149 4.2.1. Participants …………………………………………………………………………. 149 4.2.2. Material …………………………………………………………………………….. 149 4.2.3. Procedure ………………………………………………………………………….... 152 4.3. Results …………………………………………………………………………………….. 154 4.3.1. Detection of Repetition in RSVP …………………………………………………... 154 4.3.2. ALL performance …………………………………………………………………... 156 4.4. Discussion …………………………………………………………………………………. 159 5. Attention in ALL ………………………………………………………………………. 167 5.1. Introduction ………………………………………………………………………………. 167 5.1.1. Statistically-driven speech segmentation and IL ………………………………….... 169 5.1.2. The role of attention in statistical learning …………………………………………. 170 5.1.3. An overview of the experiments ………………………………………………….... 174 Experiment 5. Exposition Time and Attentional resources in ALL 5.2. Method ……………………………………………………………………………………. 178 5.2.1. Participants …………………………………………………………………………. 178 5.2.2. Material …………………………………………………………………………….. 178 5.2.3. Procedure ………………………………………………………………………….... 179 5.3. Results and Discussion ………………………………………………………………….... 181 5.3.1. RSVP Task …………………………………………………………………………. 181 5.3.2. ALL ……………………………………………………………………………….... 182 Experiment 6. Intention to learn in ALL 5.4. Method ……………………………………………………………………………………. 191 5.4.1. Participants …………………………………………………………………………. 191 5.4.2. Material …………………………………………………………………………….. 191 5.4.3. Procedure ………………………………………………………………………….... 192 5.5. Results and Discussion ………………………………………………………………….... 193 5.5.1. RSVP Task …………………………………………………………………………. 193 5.5.2. ALL ……………………………………………………………………………….... 194 5.6. General Discussion ……………………………………………………………………….. 198 XX 6. On-line measures and lexicalization during ALL ………………………………….... 209 6.1. Introduction ………………………………………………………………………………. 210 6.1.1. The nature of statistical segmentation outputs ……………………………………... 213 6.1.2. An overview of the experiments ………………………………………………….... 216 Experiment 7. On-line statistically-driven speech segmentation 6.2. Method ……………………………………………………………………………………. 220 6.2.1. Participants …………………………………………………………………………. 220 6.2.2. Material …………………………………………………………………………….. 220 6.2.3. Procedure ………………………………………………………………………….... 222 6.3. Results and Discussion ………………………………………………………………….... 223 6.3.1. Spotting Responses …………………………………………………………………. 223 6.3.2. RT data ……………………………………………………………………………... 226 Experiment 8. The status of statistical segmentation outputs 6.4. Method ……………………………………………………………………………………. 232 6.4.1. Participants …………………………………………………………………………. 232 6.4.2. Material …………………………………………………………………………….. 232 6.4.3. Procedure ………………………………………………………………………….... 240 6.5. Results and Discussion ………………………………………………………………….... 240 6.5.1. ALL performance …………………………………………………………………... 240 6.5.2. Lexicalization test …………………………………………………………………... 242 Experiment 9. The status of the output of the segmentation procedure adopted by listeners 6.6. Method ……………………………………………………………………………………. 249 6.6.1. Participants …………………………………………………………………………. 249 6.6.2. Material …………………………………………………………………………….. 250 6.6.3. Procedure ………………………………………………………………………….... 251 6.7. Results and Discussion ………………………………………………………………….... 251 6.7.1. ALL performance …………………………………………………………………... 251 6.7.2. Lexicalization test …………………………………………………………………... 253 6.8. General Discussion ……………………………………………………………………….. 256 PART III – GENERAL DISCUSSION AND CONCLUSIONS 7. General Discussion …………………………………………………………………….. 273 7.1. Main findings of the present study …………………………………………………….... 274 7.1.1. Speech segmentation with physically degraded signal ………………………….…. 274 XXI 7.1.2. Speech segmentation in cognitive noise ……………………………………………. 280 7.1.3. Cognitive noise and statistically-driven speech segmentation ……………………... 282 7.1.4. The ALL paradigm revised ……………………………………………………….... 286 7.2. Theoretical framing of the present results …………………………………………….... 289 7.2.1. A theoretical proposal – the nature of the available sources of information ……….. 291 7.2.2. The nature of statistical segmentation output …………………………………..…... 299 7.3. Some unanswered questions ……………………………………………………………... 309 7.3.1. Compatibility between sublexical segmentation cues, prelexical units and lexical mechanisms ……………………………………………………... 309 7.3.2. Cross-species statistical learning ………………………………………………….... 316 7.4. The Artificial Language Learning Paradigm revisited ………………………………... 320 7.5. Thoughts for future research ……………………………………………………………. 322 7.6. In sum ……………………………………………………………………………………... 324 8. References …………………………………………………………………………….... 327 9. Appendices ……………………………………………………………………………... 363 9.1. Appendix I ……………………..……………………………………………….…. 365 9.2. Appendix II ……………………..…………………………………………………. 367 9.3. Appendix III ………………………………..……………………………………... 369 9.4. Appendix IV …………………………...…………………………………………... 373 9.5. Appendix V ……..…………………………..……………………………………... 375 9.6. Appendix VI ……....……………………………...………………………………... 377 XXII INDEX OF TABLES Table 1: Orthographic translation of a sample of the stream heard in the familiarization phase in the three cue conditions of Experiment 1. ……………………………………………………………...…............. 70 Table 2: Performance pattern (proportion of TP-Word responses, in percentage) according to cue condition (single cue; congruent cues; incongruent cues) and Signal Quality (intact speech; 22dB SNR; 10dB SNR), in Experiment 1, considering the TP gradient of the TP-words of the AL (High-TP; low-TP) …….…………………........ 76 Table 3: Orthographic translation of a sample of the stream heard in the familiarization phase in the four cues conditions and of the AL-stimuli (TP-words; Part-words 3#12; Part-words 23#1) of Experiment 2. ……..………........ 101 Table 4: Average proportions of TP-word responses (in percentage), separately for cues condition (single cue; incongruent cues - 1st syl; incongruent cues - 2nd Syl; and congruent cues) and signal quality condition (Intact vs. Mildly Degraded), considering Part-words type (3#12 vs. 23#1) in Experiment 2. ……………………...... 107 Table 5: Orthographic translation of a sample of the stream heard in the familiarization phase in the four conditions of stress location of Experiment 3B. …..……………………………………………….............. 124 Table 6: Average proportions of TP-Word responses (in percentage) in Experiment 2B, separately for each inputintelligibility condition (intact; mildly degraded; strongly degraded) and lexical stress location (TP-unstressed; 1st syllable; 2nd syllable; 3rd syllable), considering the TP-level of the AL-words (highvs. low-TP). ……........ 127 Table 7: Performance pattern (mean percentage of Correct Detections – CD – and false alarms – FA) for the Three AL-conditions, according to the number and congruence of segmentation cues available (single cue; congruent cues – congruent -; incongruent cues – incongruent), in the RSVP task, according to visual block (block1; block 2; block 3) and d’ scores in Experiment 4. ….…….…….………………………….. 155 Table 8: Average proportions of TP-Word responses (in percentage), separately for each cue condition (single cue; congruent cues; incongruent cues) and cognitive noise condition (low vs. high load), considering the TPgradient of the AL words (highvs. low-TP-words) in Experiment 4. ……...………………........ 157 Table 9: Average proportions of TP-Word responses (in percentage), separately for each familiarization time (7-; 14-; and 21-min) and attentional load (low vs. high load) condition, considering the TP level of the AL words (highvs. low-TP-words) in Experiment 5. …………….……………………………........ 183 XXIII Table 10: Average proportions of TP-Word responses (in percentage), separately for each learning condition (Incidental and Intentional) and attentional resources condition (low-attention vs. high-attention load), considering the TP-level of the AL words (highvs. low-TP-words) in Experiment 6. …....….…….……………........ 195 Table 11: Examples of materials used (in phonological form, according to IPA) in the AL familiarization phase (AL-Fam.), two-forced choice (ALL-test), and lexicalization test (Lexical Decision – LD on base-words) in Experiment 8, separately for the two artificial languages (ALs: ALA and ALB). …………………........... 231 Table 12: Base-Words characteristics by block (Blocks A and B; Experiment 8 and 9: mean log frequency; neighborhood density, mean transitional probability (TP between first and second syllable), Uniqueness Point (UP), and mean duration (in ms), as well as t-values (all ps > .5). ….….…………………........ 233 Table 13: Mean RTs (measured from target onset) for Base-Words A and B (BWd A; BWd B) and the Priming Effect (P.E., the difference in lexical decision latencies between primed and unprimed conditions) separately for each AL-condition (ALA and ALB) and moment of testing (Immediately and Post-1 week) in Experiment 8. …. 243 Table 14: Examples of materials used in AL-familiarization phase (AL-Fam.), ALL-test, and Lexicalization test (Lexical Decision – LD on Base-Words) on Experiment 9, regarding the two Artificial Languages (ALs). …..…………........ 248 Table 15: Mean RTs (measured from target onset) for Base-Words A and B (BWd A; BWd B) and the Priming Effect (P.E., the difference between primed and unprimed conditions) separately for each AL-condition (ALA2 and ALB2) and moment of testing (Immediately and Post-1 week) in Experiment 9. ……….…………........ 253 XXIV INDEX OF FIGURES Figure 1: Hierarchical organization of speech segmentation cues according to listening conditions, as proposed by Mattys and colleagues. The triangle represents the weighting of segmentation cues as well as their resilience to physical degradation of the signal (i.e., the base of the triangle corresponds to the strongest sensitivity to physical noise)………………………………….............................................. 38 Figure 2: Performance pattern (proportion of TP-Word responses, in percentage) according to cue condition (single cue; congruent cues; incongruent cues) and of the signal quality (intact speech; 22dB SNR; 10dB SNR) – Experiment 1. ………………………………………………………………….…........ 78 Figure 3: ALL performance pattern (proportion of AL-word responses, in percentage), in Experiment 2, broken-down by Part-word Type (3#12; 23#1), according to Input-Intelligibility (Intact; Degraded) and Cues available (single cue; incongruent cues first syllable; incongruent cues second syllable; congruent cues). .…...…………........ 108 Figure 4: Listeners’ proportion of responses on the three-alternative forced choice stress location task on AL-stimuli, broken-down by syllable detection responses (1st syl; 2nd syl; 3rd syl) in the two acoustic cue conditions (Unstressed vs. Stressed) in Experiment 3A. …………………..…………………………........ 117 Figure 5: Performance pattern (proportion of AL-word responses, in percentage) broken-down by TP-level (High-TP; Low-TP) of AL-words, according to Signal condition (Intact; mildly degraded - 22dB SNR; strongly degraded – 10dB SNR) and Lexical Stress location (TP – No stress; 1st Syllable; 2nd Syllable; 3rd Syllable) in Experiment 3B. Vertical bars denote standard error of the mean on each condition. …………………………...... 126 Figure 6: Performance pattern (proportion of AL word responses, in percentage) broken-down by TP level (High-TP-words; Low-TP-words), according to attentional load condition (High-attention Load; Low-attention Load) and familiarization time (7-min; 14-min; 21-min), in Experiment 5. …………….…………........ 184 Figure 7: Performance pattern (proportion of AL-word responses, in percentage) broken-down by TP level (High-TP-words; Low-TP-words) of AL words, according to attentional resources condition (Figure 7a: High-attention Load; Figure 7B: Low-attention Load) and learning condition (Incidental learning; Intentional learning) in Experiment 6. ………………………………………..….………………........ 196 Figure 8: Mean Adjusted RT (in ms; from AL-stimuli offset) for trisyllabic responses in spotting TP-words and part-words, separately for initial and final AL-stimuli-bearing sequences (in the Preceding vs. Following lists, respectively), in Experiment 7. ….….….……………………………………………........ 227 XXV Figure 9: Mean percentage of TP-word choices in the ALL-test (i.e., two-alternative forced choice test) for each participant, according to the moment of testing: immediately after the AL-familiarization phase (i.e., immediate) and one week after with no intervening AL-familiarization period in Experiment 8. ….…...……..…........ 241 Figure 10: Effects of AL-familiarization, according to the AL to which listeners were exposed (LA; LB), on lexical decision’s response to existing lexical items immediately after AL-familiarization: Mean RTs (in milliseconds) for the two blocks of Base-Words (Base-Words A; Base-Words B) in Experiment 8 …….….……........ 244 Figure 11: Inhibitory Effect (computed as the difference between lexical decision latencies to the primed block of Base-Words and the unprimed Block of Base-Words), according to the AL to which listeners were exposed (ALA; ALB) and the moment of testing (Immediately after and 1-week after) in Experiment 8. ………….... 246 Figure 12: Mean percentage of TP-word choices in the ALL-test for each participant, according to the moment of testing: immediately after the AL-familiarization phase (i.e., immediate) and one week after with no intervening AL-familiarization period, in Experiment 9. ….….….……………………………………........ 252 Figure 13: The theoretical proposal outlined in the present thesis allied with Mattys and colleagues proposal for language-specific cues (see also Figure 1 in section 1.4. of chapter one) …….…….….………........ 293 Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 7 nemes). Vroomen and de Gelder (1995) also demonstrated that, although lexical effects take time to develop, the direct competition between word candidates is an important component of word recognition and speech segmentation. In ambisyllabic nonsense contexts, in which the coda of the first syllable also belongs to the next syllable as onset (e.g., melkeum and melkaam), the cross-modal priming effect observed for a visual word-target (e.g., melk, in Dutch milk) was stronger when the second syllable of the disyllabic nonsense prime was the initial syllable of few words in Dutch (e.g, melkeum) in comparison to when it was the beginning of many words (e.g., melkaam). This patter of results is due to the fact that in the former case the first syllable of the prime, which corresponded to the word-target, received less inhibition from the (few) competitors initial overlapping with the second syllable of the prime than in the latter case. Tabossi, Burani, and Scott (1995) also demonstrated using cross-modal priming that participants presented with fragments of speech that could be segmented either as a long carrier word (e.g., visite – Italian for visits) or as a shorter embedded word (e.g., visi tediati – Italian for faces bored), were still considering the longer word as a potential candidate for recognition when they heard the first syllable of the second word (e.g., “tediati”). Gow and Gordon (1995) also demonstrated that both two-word (e.g., two lips) and one-word (e.g., tulips) utterances led to similar priming effects on the recognition of a semantic associate of the one-word utterance (e.g., flower). While hard to reconcile with the Cohort Model (Marslen-Wilson & Welsh, 1978), these evidences support models in which segmentation is a byproduct of lexical competition. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 8 Lexical Competition and Speech Segmentation Models in which lexical competition is conceived as occurring directly between word candidates compatible with the speech input, consider speech segmentation as a byproduct of lexical competition. As clearly defined by McClelland and Elman (1986) on TRACE model: “Word identification and segmentation emerge together as part and parcel of the process of word activation” (p. 61). “These remarkably simple mechanisms of activation and competition do a very good job of word segmentation, without the aid of any syllabification, stress, phonetic word boundary cues” (p. 64). Other models that incorporate lexical competition include the Neighborhood Activation Model (NAM; Luce & Pisoni, 1998; Luce, Goldinger, Auer, & Vitevich, 2000), the revised version of the Cohort Model (Marslen-Wilson, 1987; MarslenWilson, Moss, & van Halen, 1996), and the Shortlist Model (Norris, 1994). These models differ on the specific mechanism of lexical competition: either by direct inhibition as proposed by TRACE (McClelland & Elman, 1986) and Shortlist (Norris, 1994) models, and also by the Distributed Cohort Model (DCM, Gaskell & Marslen-Wilson, 1997; 1999; 2002), or by the indirect mediation of a decision stage as suggested by the Cohort Model (Marslen-Wilson, 1987; MarslenWilson et al., 1996) and the NAM (Luce & Pisoni, 1998; Luce et al., 2000). Nevertheless, all models agree that lexical activation will be reduced when a greater number of lexical candidates are compatible with the same speech stream, and lexical candidates with large phonological overlapping will act as strong competitors for recognition. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 9 Models such as TRACE (McClelland & Elman, 1986) and Shortlist (Norris, 1994) incorporate lexical competition between nonaligned candidates. For example, “cap” and all other words (even nonaligned candidates) in which the sequence /jzo/ is embedded (e.g., captain, captive) will directly compete for recognition through lateral inhibition. Thus, on both models, short words (e.g., cap) would only win the recognition process after its offset, when longer competitors are ruled out. TRACE vs. Shortlist Note however that although TRACE (McClelland & Elman, 1986) and Shortlist (Norris, 1994) suggest the same mechanism of lexical competition, expressed as lateral inhibition, between activated word candidates, they differ on three fundamental aspects. First, in TRACE (McClelland & Elman, 1986), all lexical nodes are potential candidates for recognition, continuously increasing or decreasing in activation as a function of their match with the incoming signal whereas in Shortlist (Norris, 1994) competition only takes place within a short list of word candidates that specifically match the input to some preset criterion. Second, TRACE (McClelland & Elman, 1986) is an expression of a highly interactive view of spoken word recognition in which there is a continuous two-way flow of information (i.e., both bottom-up and top-down); on the contrary, Shortlist (Norris, 1994) is an autonomous model, entirely bottom-up in its operation (see for a discussion between interactive vs. autonomous models on modularity grounds: Bowers, & Collin, 2004; McQueen, Norris, & Cutler, 2006). Importantly, while in TRACE speech segmentation is an exclusive byproduct of lexical competition, and thus no space remains for any effect of sublexical segmentation cues, Shortlist not only considers, Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 10 as TRACE, that segmentation can be a byproduct of lexical competition, but also incorporates explicit sublexical segmentation strategies (e.g., the Metrical Segmentation Strategy – MSS; Cutler & Norris, 1988), unified in the operation of the Possible Word Constraint (PWC; Norris, McQueen, Cutler & Butterfield, 1997). 1.1.2. Crossroads into the speech stream At a fundamental level, the speech segmentation problem faced by adults and young children is the same (Brent & Cartwright, 1996): adults sometimes encounter and learn novel words and children often recognize familiar words within continuous utterances. However, since adults are already familiar with a vastly larger proportion of words a lexical segmentation strategy can be, in adulthood, a very proficient mechanism (e.g., McClelland & Elman, 1986; Norris, 1994) but it has virtually no value in the very beginning of language onset. The conception of a solely lexical segmentation strategy, during language acquisition, would pose a chicken-and-egg problem (cf. Cairns et al., 1997): the identification of a particular stretch of speech as a meaningful unit would presuppose recognizing what that unit is, but that would only be possible once segmentation has been carried out. Since infant-directed speech does not reliably provide isolated words (Aslin, Woodward, LaMendola, & Bever, 1996), processing multi-word utterances is critical for word learning during language acquisition. On this view, children must consider as candidate words sound sequences that they have never heard in isolation. Thus, sublexical (bottom-up) sources of information could enable the bootstrapping of lexical boundary detection (e.g., Brent & Cartwright, 1996: Cairns, et al., 1997; Gleitman & Wanner, 1982), and hence would have a fundamental role in language Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 11 acquisition. Sublexical sources of information can be classified in three types of information: suprasegmental, subsegmental, and segmental information. Suprasegmental Information Suprasegmental information, i.e., prosody, was one of the first sublexical sources of information to be proposed as a speech segmentation cue. Gleitman and Wanner (1982) suggested that prosody (i.e., rhythm, intonation) might play an important role in the acquisition of syntax, acting as a “prosodic bootstrapping”, and it seems to be one of the more precocious segmentation cues. Newborns (as well as non-human primates: Ramus, Hauser, Miller, Mones, & Mehler, 2000; Tincoff, et al., 2005) are able to distinguish languages based on their rhythmic properties (e.g., Ramus, 2002; Ramus et al., 2000; Nazzi & Ramus, 2003). Two-month-olds are sensitive to Intonational Phrases1 (IPs; e.g., Dehaene-Lambertz & Houston, 1998) and at 6-month to preboundary length, pitch, and pause in clause segmentation (Seidl, 2007). After 7.5 month, infants are also sensitive to prosodic edges (Seidl & Johnson, 2006). Eight-month-olds English listeners are already able to use their native language metrical prosody for assisting speech segmentation (i.e., strong syllables are defined as possible word-beginnings; Curtin, Mintz, & Christiansen, 2005; Jusczyk, Houston, & Newsome, 1999a) and 9-month-olds treat strong-weak disyllables (i.e., trochaic units) rather than weak-strong ones (i.e., iambic units) as 1 A wide variety of terms are found in the literature, yet two levels above prosodic words are considered in prosodic hierarchy: Intonational Phrases (IPs), the highest prosodic unit, mostly corresponding to whole clauses or sentences and often marked by a pause at the end; and Phonological Phrases, below IPs, also referred as intermediate IPs or accentual groups (Grice, 2004; Werner & Keller, 1994). Both types of phrases are acoustically characterized by a final lengthening with a falling pitch contour (at their right edge; see Christophe, Peperkamp, Pallier, Block, Mehler, 2004; Shukla, Nespor, & Mehler , 2007). Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 12 cohesive units (Morgan & Saffran, 1995). Eight-month-old English infants are also able to use lexical stress in speech segmentation, since trisyllabic stimuli constituted by a sequence of strong-weak-strong syllables are only correctly extracted from the stream as long as the first syllable receives primary stress, but not when primary stress is located at the last syllable (Houston, Santelman, & Jusczyk, 2004). Twelve-month-olds French listeners (but not 8-month-olds) also use their native language’s prosodic units in segmentation (i.e., the syllable; Nazzi, Iakimova, Bertoncini, Frédonie, & Alcantara, 2006). Furthermore, 13-month-olds are also sensitive to prosodic boundaries, such as the ones of phonological phrases (Cristophe, Gout, Peperkamp, & Morgan, 2003). In fact, prosody is a general term associated with many kinds of suprasegmental information, such as rhythm, intonation and word primary stress (for a review see Cutler, Dahan, & van Donselaar, 1997), which do not necessarily have the same role in speech segmentation. For example, the role of IPs’ contour seems to be universal, probably based on physiological mechanisms such as breath groups (Grice, 2006; Shukla, Nespor, & Mehler, 2007; Werner & Keller, 1994). On the one hand, adult listeners are sensitive to those prosodic contours even in an unknown/foreign language (Shukla et al., 2007; Experiments 5 and 6). On the other hand, these prosodic correlates, such as the ones of prosodic right-edges (which characterize IPs in many different languages; Werner & Keller, 1994), are acoustic marks of the slowing down of the articulators within a breath group, which is reflected in the signal as final lengthening and low pitch (Grice, 2006, see also Cutler et al., 1997). In contrast to universal prosodic cues such as IPs, lexical (or word Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 13 primary) stress, although generally correlated with duration, F0, and amplitude (stressed syllables are lengthened, have higher pitch, and are louder), is not acoustically marked in the same manner in all languages. For example, while in Finish pitch seems to be the most important correlate of lexical stress (e.g., Iivonen, Niemi, & Paananen, 1998), in European Portuguese2 (EP; the native language of participants of the present work) lexical stress is marked by syllable lengthening (Delgado-Martins, 2002; d’Andrade & Lacks, 1996). The location of stress and whether it obeys to a fixedor free-pattern is also language-dependent (e.g., in Finish, always on the first syllable of a word: Iivonen et al., 1998; in EP lexical stress is by default on the penultimate syllable, although it can occur in any one of the three last syllables of a polysyllabic word, d’Andrade & Laks, 1996; Mateus & d’Andrade, 2000). Thus prosody is a broad category that comprises both universal (e.g., IPs) and language-specific (e.g., word primary stress) cues, which do not necessarily have the same role in speech segmentation. Subsegmental Information Subsegmental information corresponds to acoustic-phonetic cues (Davis, Marslen-Wilson, & Gaskell, 2002; Mattys, 2004). As a matter of fact, investigations of minimal pairs differing only in word boundary location (such has play taught and plate ought) identified a variety of acoustic (subsegmental) cues that are associated with word-boundaries (Lehiste, 1960). Allophonic differences in the articulation of segments in different contextual positions are observed in any natural language. For example, in EP, the allophones of [l] in mel (in EP, honey) and in lua (in EP, moon) 2 Although European Portuguese (EP) and Brazilian Portuguese (BP) share many properties, there are differences particularly as regards prosodic aspects (e.g., Frota & Vigário, 2001). For example, while vowel reduction is a prominent phenomenon in EP, it does not occur in BP (e.g., Abaurre & Galves, 1998). Thus in the present thesis we are exclusively considering EP. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 14 are differently produced according to their position in syllables and words (Mateus & d’Andrade, 2000). Word-initial segments are also longer in duration than equivalent segments that are not word initial (Lehiste, 1960). Durational differences are also observed between the same segments and syllables in short and long words (Klatt, 1980). Therefore, this acoustic-phonetic knowledge could provide listeners with clues about word boundaries location in fluent speech (Liberman & StuddertKennedy, 1978). As a matter of fact, two-month-old infants are already sensitive to allophonic differences, such as nigh rates and nitrates (Hohne & Jusczyk, 1994). However, only at 10.5 month, infants are able to use allophonic information in speech segmentation’s assistance (Jusczyk, Hohne, & Bauman, 1999b). The degree of coarticulation is also an important subsegmental cue (Fougeron & Keating, 1997) since an early phase in linguistic development (i.e., 8-month-olds; Johnson & Jusczyk, 2001). Coarticulation is usually defined as a change in the acoustic-phonetic content of a speech segment due to anticipation or preservation of adjacent segments (e.g., Kühnert & Nolan, 1999), and is the primary reason for the difficulty in specifying invariant acoustic properties of phonetic segments. Although all fluent speech is coarticulated, the extent of coarticulation between adjacent segments is influenced by the presence of a prosodic boundary. There is, generally, more coarticulation within words than between them (e.g., Byrd & Saltzman, 1998). On the one hand, the degree of overlap is greater for consonantal clusters at onset position than for adjacent consonants that belong to different words (Byrd, 1996). On the other hand, segment strength may convey information about the local coherence vs. disjuncture in connected speech (Fougeron & Keating, 1997; Keating, in press). Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 15 Segments at the beginning of prosodic domains are generally strengthened, while segments within those domains are not, and this domain-initial strengthening is a general phenomenon found in both stressed and unstressed syllables (Cho & McQueen, 2005). Segmental Information Another important source of information that can help listeners to locate word boundaries in any natural language is the statistical segmental information conveyed by the speech stream. The transitional probability (TP) between two linguist units (e.g., syllables or phonemes) corresponds to the conditional probability by which one first unit (X) is able to predict the immediately following one (Y): As a matter of fact, within any language, the TP from one unit of sound to the next will generally be highest when the two units follow one another within a word, whereas the TP spanning a word boundary will be relatively low (e.g., Perruchet & Peereman, 2004). This pattern of higher TPs within words than across them is also characteristic of natural infant-directed speech (Swingley, 2005). Human listeners, in different linguistic stages of development (i.e., infants, children and adults) and even nonhuman primates (Hauser, Newport, & Aslin, 2001) can use TPs between adjacent syllables to locate “word” boundaries within a continuous artificial language (AL) stream composed by nonsense syllables with no acoustic cues to “word” boundaries (e.g., Saffran et al., 1996a; 1996b). TP (Y|X) = Frequency (XY) Frequency (X) Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 16 Saffran and colleagues (e.g., Saffran et al., 1996a; 1996b) have implemented the AL learning (ALL) paradigm and proposed that a statistical learning mechanism is responsible for listeners’ ability to discover “words” embedded in an AL. In this paradigm, in the first, familiarization, phase, listeners are exposed to a continuous stream of nonsense syllables (i.e., the AL; e.g., bupadapatubi, Saffran et al., 1996b), with no other cues to word boundaries than the TPs between (adjacent or nonadjacent) syllables. Within the AL continuous stream, AL-syllables belonging to the same AL “word” have higher TPs (henceforth, TP-words; e.g., bupada, patubi) than AL-syllables that are part of different “words” (i.e., part-word stimuli; e.g., da#patu), straddling a TP-boundary (marked in the previous example by “#”). In the second phase of the ALL paradigm, participants are presented with isolated ALstimuli: AL-words vs. AL-nonwords3. With infant listeners, the Headturn Preference Procedure (Jusczyk & Aslin, 1995) enables the evaluation of infants’ sensitivity to the statistical information conveyed by the stream. Indeed, the discrimination between TP-words and partwords is not merely based on the raw frequency of occurrence of the stimuli, since TP sensitivity is observed for 8-month-olds even when part-words occurred as often as TP-words during the familiarization phase (Aslin, Saffran, & Newport, 1998). Furthermore, reliance on statistical information such as TPs can contribute to real-word acquisition. Since TP computation does not require any (even minimal) lexical knowledge, this conditional probability estimate might allow learners to segment speech from the discovery of troughs in the TPs distribution between 3 AL9nonwords are stimuli constituted by the syllables of the AL but with lower TPs than the TP9words, either because they did not occur in the familiarization phase (their TP being 0), or because, although having occurred, and even with the same raw frequency as TP9words (Aslin, Saffran, & Newport, 1998), their TP is lower. These last AL9nonwords are called Part9words, since their syllables occur in the stream as adjacent parts of different TP9words, hence straddling TP9words boundaries. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 23 1.2. Conventional Experimental Paradigms Most of the studies that demonstrated adult listeners’ sensitivity to sublexical segmentation cues used conventional experimental techniques such as gating, wordspotting, cross-modal (semantic or repetition) priming, and eye-tracking paradigms. Although all experimental tasks have pros and cons (Goldinger, 1999; Kolinsky, 1998), these conventional paradigms present the same broad limitation, since speech segmentation, as well as the role of sublexical sources of information on it, is indirectly inferred according to lexical activation of different word candidates. 1.2.1. Gating The dynamics of spoken word recognition have often been examined using gating tasks (for an overview see Grosjean, 1996). This paradigm has provided valuable information on the relationship between the acoustic signal (and the available acoustic-phonetic cues) and lexical access (e.g., Davis et al., 2002). The gating technique provides absolute control over the duration of the speech signal presented to participants. In this task, speech is presented in fragments (i.e., gates) of increasing duration and participants are asked to propose the word being presented and to give a confidence rating of their response. For a given word, the isolation point is defined as the duration of the gate at which participants correctly identified the stimulus-word and did not subsequently change their guess. Since listeners are encouraged to focus on accurate responses independently of the time taken to select them, gating provides a working measure of the steady state of the word recognition Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 24 system. It also provides measures reflecting the activation and identification of competing lexical hypotheses as acoustic information accumulates. However, gating is usually considered an off-line task that may include strategic effects, since it appeals to the sequential nature of the stimuli presented and possible bias can underlie listeners’ performance. Indeed, gating responses can be generated after the presentation of each gate, and most likely reflect lexical interpretation given the available sensory input after internal processing has reached an asymptotic stable state (e.g., Dahan & Gaskell, 2007). 1.2.2. Word-spotting According to McQueen (1996) the word-spotting task appears tailor-made for investigating lexical segmentation. In the word-spotting task, listeners are asked to detect any real word embedded in nonsense (pseudo-continuous) strings, which gives the task some “ecological validity” (McQueen, 1996). Indeed, since subjects are not told which word they are asked to spot, at first sight, this task seems to resemble the problem that listeners face in natural settings, with continuous speech. Kolinsky (1998) already suggested that tasks in which participant’s attention is not drawn to factors at study, reduce the use of late (i.e., metaphonological) representations involved in a post-recognition stage. However, despite the apparent resemblance between word-spotting task and recognition in connected speech, the slow reaction times (RTs) and high error rates suggest that this is a difficult task for listeners. Furthermore, it is not clear whether metaphonologic representations could be involved in listeners’ performance. In some metaphonological tasks, participants are asked to produce their response by deleting one linguistic unit (e.g., a syllable or Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 25 a phoneme) at a particular position of the stimuli presented. In some way, this seems to be similar to what participants are asked to do in word-spotting, particularly in conditions in which the location of the embedded word is given beforehand (whether it can occur in the beginning or in the end of the nonsense container string). Eventually, in order to correctly perform the word-spotting task, listeners are asked to delete one (or more) segment of the nonsense sequence and also evaluate whether this metaphonological operation results in any real word. This explicit, intentional analysis of recognition outputs would be far from the perceptual literacy-independent stage (see Kolinsky, 1998). Yet, this suspicion remains to be tested. It is important to note, however, that word-spotting has often been used to analyze the role of sublexical information in speech segmentation, such as metrical prosody (e.g., strong syllables, in English: Cutler & Norris, 1988; lexical stress, in Finish: Vroomen et al., 1998), and segmental information, like phonotactics (McQueen, 1998), the PWC (Norris et al., 1997), vowel harmony (Vroomen et al., 1998) and syllable onsets (Dumay, Frauenfelder, & Content, 2002). Crucially, the role of these sublexical segmentation cues demonstrated with word-spotting was also replicated with other paradigms (e.g., cross-modal priming, Mattys et al., 2005; the ALL: Saffran et al., 1996a; Vroomen et al., 1998). 1.2.3. Eye-tracking The eye movement paradigm (Allopena et al., 1998; for an overview see Tanenhaus & Spivey-Knowlton, 1996) has proved to be a valuable methodology for studying on-line lexical access in spoken-word recognition using continuous speech input. This fairly natural setting enables the exploration of subtle competitor effects Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 26 that possibly could not be observed with other conventional techniques. In the visual world task, participants are instructed to fixate a central cross then followed by a spoken instruction to move (using a computer mouse) one of four objects displayed on a computer screen (e.g., “click on the beaker with the computer mouse”; Allopena et al., 1998). Eye gaze is monitored as the spoken stimulus is heard and until listeners perform the instruction. Eye movements’ monitoring (i.e., location and latencies of eye fixations for each item presented in the visual display over time) allows the examination of spoken word recognition as the continuous speech stream unfolds, since eye movements occur concurrently with the spoken input in real time. Thus, fixations initiated at time t reflect the spoken word processing stage reached at that time, determining both the available input and internal processing dynamics. This measure provides a fine-grained estimate of target and competitor consideration over time, being extremely sensitive to the uptake of information during lexical access. Indeed, Allopena et al. (1998) demonstrated that the presence of a cohort competitor (e.g., beetle) in the display increased the latency of eye movements to the target (e.g., beaker) and induced participants’ frequent looks to that competitor. Thus, Allopena et al. (1998) directly demonstrated that two referents with phonologically similar names were in fact competing as the target word unfolds. Dahan, Magnuson, and Tanenhaus (2001) also demonstrated with eyetracking that frequency affects the earliest moments of lexical access. Indeed, when a referent picture (e.g., picture of a bench) was presented with two cohort members with highand low-frequency (e.g., bed and bell, respectively), participants fixated more often the high-frequency competitor than the low-frequency one. Thus, frequency effects occur early on as spoken word unfolds, suggesting that the locus of Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 27 frequency effect is not late-acting nor has a decision-bias locus. Furthermore, eyetracking data has been compared with simulations of connectionist models such as TRACE (e.g., Allopena et al., 1998; Dahan et al., 2001), providing an important tool for theoretical testing. Magnuson, Tanenhaus, Aslin, and Dahan (2003) have also used eye-tracking, after an object-label learning task, for examining the course of lexical activation and competition on the artificial lexicon in adulthood. Surprisingly, Magnuson et al. (2003) demonstrated signature processing effects of cohort and rhyme competitor modulated by target and neighbor frequency during “word” recognition in the artificial lexicon. Thus, artificial lexicons seem to exhibit the processing effects that characterize spoken word recognition with real words (Allopena et al., 1998; Dahan et al., 2001). These lexical signature effects for artificial lexical items were observed even immediately after the first session of object-label (i.e., new visual shape - new spoken word) learning. Recent eye-tracking studies (e.g., Ju & Luce, 2004; Salverda et al., 2003; Salverda et al., in press; Shatzman & McQueen, 2006) also demonstrated participants’ on-line sensitivity to signal derived cues, and their role on lexical competitors’ activation. Although eye-tracking provides fine-grained information of the temporal course of lexical activation, it suffers from two particular limitations: (i) it only measures activation for (the limited) displayed items; and (ii) only word-targets with clear referents can be evaluated using this paradigm. In a very recent study of Magnuson, Dixon, Tanenhaus, and Aslin (2007) a new version of the visual world paradigm that overcomes the first limitation pointed Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 28 out above was used. In this version the recognition of single monosyllabic words (varying in frequency and competitor’s environment) were displayed as a referent picture among three other unrelated distractors. Thus, competitors were never present in the same display as the target as it is usually done in eye-tracking experiments. In this new version was still possible to infer how the set of activated lexical candidates change as a word is heard, and hence to estimate the time course of lexical activation by measuring eye movements as target-words were presented. In fact, Magnuson et al.’s (2007) data demonstrated that while the competitor set changes dynamically, earlier on as the spoken word unfolds, stronger competition is observed between targets and cohorts (i.e., candidates sharing phonological onsets; cf. Marslen-Wilson, 1987). While in an early moment effects of word frequency and cohort density dominate, global neighborhood effects emerge later in the recognition process, and recognition occurs against a background of activated competitors that changes over time based on both fine-grained goodness-of-fit and competition dynamics. Notably, even studies in which participants are only asked to look at the display with no explicit instructions to search for particular targets (i.e., a passive condition: e.g., Huetting & McQueen, 2007) demonstrate the same effects observed in conventional conditions in which participants are asked to perform an over task on the visual display presented (e.g., Allopena et al., 2007). Recently, some studies (e.g., Huetting & McQueen, 2007; McQueen & Viebahn, 2007) started to overcome the second limitation referred above by using printed words instead of pictures in the visual display. In these studies, the same Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 29 phonological effects observed with visual referents (e.g., Allopena et al., 1998) were demonstrated. However, different patterns of results are found with words and pictures (see e.g., Huetting & McQueen, 2007), suggesting that specific aspects may differentiate the processes involved in the treatment of these two types of material. Thus further research is required on this matter before implementing this printed words version of the “visual-world” paradigm. 1.2.4. Cross-modal priming Both cross modal semantic (for an overview see Tabossi, 1996) and repetition priming allow the assessment of words activation in connected speech, and thus have also been used to examine the interaction of lexical and prelexical informations on speech segmentation (e.g., Gow & Gordon, 1995). In cross modal priming experiments participants usually perform a lexical decision task on visual targets after hearing an auditory prime. The inter-stimulus interval (ISI) between prime and target can also be manipulated. By comparing the RTs of lexical decision in the condition in which prime and target-word are related (e.g., semantic/associative relationship: tulips and flower) with the RTs in the condition in which no relation exists between prime and target (i.e., priming effect), it is possible to evaluate to which degree different lexical items are activated. In the semantic version, in critical trials, the visual target-words are semantically and/or associatively related either to the prime (e.g., flower and tulips; Gow & Gordon, 1995; Experiment 1) or to part of the prime (e.g., kiss and warm lips; Gow & Gordon, 1995; Experiment 1). In the repetition version, in critical trials, target-words can either correspond to the prime (in the repetition priming task: Davis Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 30 et al., 2002; Experiment 2) or share phonological information with part of the prime (e.g., diplenota and notable; Mattys, 2004). In the repetition version of cross-modal priming, the possible confounds produced by differences in semantic or associative relatedness are avoided, while retaining sensitivity to effects of lexical competition and mismatch. As a matter of fact, Norris, Cutler, McQueen and Butterfield (2006) have proposed that associative priming is not a direct consequence of automatic speech processing, since the conceptual interpretation of the speech input occurs only in a subsequent phase separated from the activation of phonological forms of the current lexical candidates. Importantly, the fragment version of cross-modal priming has been used to evaluate the weighting of different segmentation cues, incongruently available, and hence competing cues, in the speech input (i.e., the prime stimuli; Mattys, 2004; Mattys et al., 2005). However, one possible limitation of the cross-modal priming paradigm regards the fact that the non-observation of priming effects does not ensure that a lexical competitor that (partially or totally) matches with the input is not at least weakly activated. It is possible that a competitor might be weakly activated but not sufficiently so to be detected in this task, leading to no significant priming effects. As already suggested by Kolinsky (1998) the attribution of tasks to specific processing stages is not straightforward, and can easily become an oversimplistic approach. However, in all the conventional experimental tasks aforementioned, the role of sublexical cues in speech segmentation is indirectly inferred according to lexical activation of bottom-up matching lexical candidates. Thus, with those tasks it Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 31 is not possible to disentangle the contribution of sublexical sources of information per se from the role of high-level information in speech segmentation, since lexical activation and competition are also intrinsically involved on listeners’ performance. Notably, this shortcoming is not present on the ALL paradigm implemented by Saffran and colleagues (Saffran et al., 1996a, 1996b) and hence this technique could be a valuable tool for studying the role of sublexical information in speech segmentation, independently of the impact of high-level information in speech segmentation. This is particularly important in the context of a full mature speech perception system, since in this case high-level information per se has a major role in speech segmentation (e.g., McClelland & Elman, 1986; Norris, 1994). Thus, in order to disentangle the role of sublexical cues from the one of high-level information, it is important to adopt an experimental technique that enables the evaluation of those sublexical sources of information in the absence (or almost so) of available highlevel information. 1.3. The Artificial Language Learning Paradigm In fact, while in natural languages a rich set of correlated cues is available, which makes the isolation of each cues’ role in speech segmentation extremely difficult, ALs allow the systematic manipulation of particular sublexical sources of information, enabling the evaluation and preponderance of isolated cues over others. Additionally, in AL settings, the role of sublexical segmentation cues is not inferred according to high-level activation, as happens in conventional designs using the experimental techniques briefly reviewed in section 1.2. of the present chapter. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 32 In the ALL paradigm, participants are presented with an unknown language (i.e., the AL), and thus high-level information is absent (or at least largely reduced) and hence cannot be used in this speech segmentation processing. The rationale of this paradigm (Saffran et al., 1996a; 1996b) conceives that if listeners show learning effects (e.g., in a forced-choice task) after an AL familiarization period, it is because the information available in the AL corresponds to the one available in any natural language and used in speech segmentation. During the familiarization phase of an ALL study, as already referred in the previous section 1.1.2., listeners are exposed to a continuous AL stream constituted by a limited repertoire of concatenated syllables. Usually statistical information is the only available cue that can help listeners to locate word-boundaries, since the TP between adjacent syllables is higher within ALwords (i.e., TP-words) than between them. Thus, any preference for those TP-words over other AL-stimuli with lower TP (i.e., nonwords or part-words) in a subsequent ALL-test (e.g., two alternative forced-choice test) is a demonstration that listeners can parse TP-words from the continuous speech stream through the adoption of a statistical segmentation procedure. This test provides a measure of learning, and consequently allows evaluation of the reliability of the sublexical segmentation cues available in the speech stream, independently of high-level (lexical and post-lexical) information. Additionally, the ALL paradigm also provides the investigation of speech segmentation with a highly controlled situation, reducing the set of uncontrolled variables. This paradigm also permits the systematic manipulation of different sublexical cues simultaneously available in the AL stream, in conditions of both congruent (i.e., different available cues suggesting the same AL parsing) and Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 39 (2004; Mattys et al., 2005, Experiment 1A) demonstrated that the relative impact of (subsegmental) coarticulation and (suprasegmental) lexical stress information is modulated by signal quality. When these two cue types were put against each other, while with phonetically intact speech, the possible word boundaries indicated by coarticulation outweighed the ones indicated by lexical stress, when the signal was physically degraded by white-noise superimposition, the hierarchical organization of these cues was reversed (see Figure 1 for the illustration): listeners adopted a stressbased segmentation procedure, with the possible word boundaries indicated by stress outweighing the ones indicated by coarticulation. This metrical segmentation bias (cf. Mattys, 2004), dependent on physical degradation of the signal, was also observed when the incongruent information available was another sublexical cue represented at Tier II (i.e., phonotactics) and even when the incongruent cue available corresponded to high-level information, represented in Mattys et al.’s (2005) model at Tier I (i.e., lexical information and semantic-context: Mattys et al., 2005; Experiments 5 and 6A, respectively). Additionally, suprasegmental information was able to drive speech segmentation processing in strongly degraded conditions (i.e., white-noise superimposed at a signal to noise ratio – SNR – of -5dB) when multiple cues represented at each one of the other two above tiers were also available in the stream suggesting conflicting segmentation points (Mattys et al., 2005; Experiment 6B). 1.4.2. Developmental and Cross-linguistic implications Part of the differential weighting, and hence of the hierarchical organization according to the resilience of cues to noise, might be related to their structural grain Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 40 (Mattys et al., 2005). For example, the use of cues defined at a lower structural level, such as coarticulatory cues, depends on the ability to process fine-grained, low-level acoustic properties, which are more easily masked by noise than higher-level units. Indeed, the perceptual salience of coarticulation is much more affected by the presence of noise than other cue types (Mattys, 2004; Mattys et al., 2005). When only sublexical information is available in the stream, segmentation driven by coarticulation is only observed with phonetically intact speech (Mattys, 2004), and when information of all three tiers is available (high-level and sublexical information) coarticulation has only an effect in mild-noise conditions (i.e., 5dB SNR; Mattys et al., 2005, Experiment 6B). The representation of segmental and subsegmental information at the same middle tier by Mattys et al. (2005) is based on the fact that in any natural language segmental and subsegmental information tend to be intrinsically correlated, and hence they propose the same segmentation hypotheses. However, Mattys and colleagues also suggest that even cues represented at the same tier can be differentially weighted in particular languages, as long as the word-boundary predictability of those cues is not highly correlated in that language. Cross-linguistic differences would be particularly observed on the weighting of sources of information that are language-specific cues, i.e., specifically related with the particular language, such as lexical stress (e.g., in Finish always at the first syllable of a word, while in EP by default in the penultimate syllable). For example, while in Finish a lexical stress segmentation strategy could be a very reliable cue even with intact signal, in EP this is unlikely, since lexical location is neither fixed nor indicating a highly probable word-boundary. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 41 The hierarchical organization framework of Mattys and colleagues (Mattys, 2004; Mattys et al., 2005) is also underlain by a developmental conception of cues weighting, according to their role in different stages of linguistic development. Sublexical cues represented at the lowest tiers have a central role in development as guides to other segmentation cues, and hence they have a subordinated role, being lower-weighted, in the context of a full mature speech perception system. However, being a fundamental mechanism for language acquisition, though lower weighted in adulthood, confers it relative immunity to signal degradation. During linguistic development, these fundamental lowest-weighted sources of information were supplanted by other more reliable (and thus with high word-boundary predictability) and latter acquired segmentation cues. For example, the ability to track TPs appears to be precocious (e.g., Kirkham et al., 2002; Thiessen & Saffran, 2003), rapid (e.g., Saffran et al., 1996a), and with a possible phylogenetic basis (Hauser et al., 2001), enabling statistical information to act as a pivot mechanism (cf. Mattys et al., 2005) in the acquisition of not only words but also other word-boundary cues (Thiessen & Saffran, 2003). When TPs and lexical stress (i.e., language-specific cue) are both available in the stream suggesting different segmentation points, English 6.5-months-old infants consider TPs the more reliable cue (Thiessen & Saffran, 2003) whereas at 9-month-old, the weighting of these cues is reversed, with lexical stress being the more reliable one (Johnson & Jusczyk, 2001; Thiessen & Saffran, 2003). Furthermore, by 10.5 months, English infants are already able to integrate multiple sources of information, and lexical stress starts to loose its previous importance (Jusczyk et al., 1999a), remaining a lastresource segmentation heuristic in adulthood (Mattys et al., 2005). Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 42 Thus, the more resilient (but lower-weighted in normal listening conditions) cues in adult speech segmentation could correspond to the earliest acquired ones, which would also be the most critical cues at the onset of language development. Nevertheless, since Mattys’ (Mattys, 2004; Mattys et al., 2005) framework is quite recent, two fundamental aspects, studied in the work presented in this thesis, remain underspecified. First, it is not clear how different types of sublexical information, both represented at the same tier (e.g., segmental and subsegmental information) and at different tiers (e.g., segmental and suprasegmental information), interact with each other. Second, Mattys et al. (2005) already suggested that listening conditions go beyond physical quality of the signal. However, until the moment, all the studies addressing this aspect have only evaluated impoverished conditions by physical degradation of the signal. 1.5. An overview of the present thesis In the work described in this thesis we mainly adopted the ALL paradigm to evaluate two central aspects of the role of sublexical sources of information in speech segmentation on the grounds of Mattys et al.’s (2005) theoretical proposal. In particular, we evaluated the weighting of different sublexical cues and the role of (qualitatively) different listening conditions in speech segmentation in the context of a full, mature speech perception system. The ALL paradigm allows the simultaneous evaluation of the role of different sublexical sources of information in speech segmentation. Using this paradigm, it is Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 43 possible to manipulate systematically those cues in congruent and in incongruent conditions. In the congruent cues condition, the available cues would suggest the same segmentation outputs. This would enable the evaluation of whether the congruency of different available cues promotes an optimization of speech segmentation processing in comparison to a condition in which only one segmentation cue is available in the speech stream. In the incongruent cues condition, when different cues are available in the stream suggesting incongruent (and hence competing) segmentation hypotheses, it is possible to access which cue prevails over the other. In this incongruent cues condition, in the second phase of the ALL paradigm (i.e., forced-choice testing), participants are presented with the segmentation outputs of the cues previously available in the AL stream, and are asked to decide which units are more plausible “words” of the new language. Note that in this incongruent cues condition the available cues suggest different segmentation hypothesis. Thus, based on participants’ preference for the outputs of one segmentation procedure over another, we can evaluate the weighting of the available sources of information (i.e., which sublexical cues listeners consider more reliable). Since this incongruent cues condition can be presented in different listening conditions, we are also able to evaluate whether the weighting of the available cues (i.e., which cue prevails over the other) is modulated by the listening conditions as suggested by Mattys and colleagues (Mattys et al., 2005). 1.5.1. The weighting of different sublexical sources of information in speech segmentation Mattys et al. (2005) showed that when coarticulation (subsegmental) and Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 44 phonotactic (segmental) information (both at tier II) were available, indicating the same segmentation points, their segmentation hypotheses were the ones used by adult listeners, in two scenarios: (i) when lexical information was not available in the signal; (ii) when signal quality was sufficiently degraded as to reduce reliance on the semantic context, but, at the same time, allowing a relative availability of the acoustic-phonetic aspects. But, since in Mattys et al.’s study phonotactic and coarticulatory information were never put in conflict, we do not know whether one (and which one) of these two information types prevailed. In fact, Mattys and colleagues (Mattys, 2004; Mattys et al., 2005) posit that segmental and subsegmental information are represented at the same tier, and consequently their model confers them the same importance. Nevertheless, it is plausible that even sublexical types of information represented in the same tier have independent impacts on segmentation and differ by their degree of (in)dependence on signal quality. Indeed, as already posited by Mattys and colleagues, and mentioned in section 1.4.2. of the present chapter, at least in some languages, segmental and subsegmental information may be weighted differently according to their wordboundaries predictability. This proposal was investigated in Experiment 1 described in chapter 2 of the present thesis, as regards coarticulation (a subsegmental, acousticdependent, cue) and transitional probabilities, a segmental statistical cue related to a higher structural level (here, syllabic, i.e., the TP between adjacent syllables). In an ALL setting, we investigated, through three signal-quality conditions, the relative impact of these two types of sublexical information in speech segmentation: segmental statistical information – TPs between adjacent syllables, and subsegmental information – coarticulation. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 45 In addition, exclusively in physically degraded conditions for adult listeners, suprasegmental information seems to have a preponderant role over other sublexical sources of information such as statistical segmental cues (Mattys et al., 2005; Experiment 2). However these two types of information were only evaluated regarding word primary stress and phonotactics, and only with English listeners (Mattys et al., 2005). Indeed, Mattys et al. (2005) already suggested that the weighting of language-specific cues, such as lexical stress, probably depends on their word-boundary predictability in a specific language (cf. Mattys et al., 2005; see section 1.4.2. of the present chapter). Therefore, cross-linguistic differences would be particularly observed on the weighting of these language-specific cues. As a matter of fact, the majority of content words in stress-time languages like English have a metrically stressed syllable at word onset. Yet, that is far from true in EP, in which by default lexical stress occurs in the penultimate syllable of polysyllabic words. Thus, in EP, lexical stress does not generally indicate any word-boundary. Furthermore, as already mentioned above, prosody is a general term associated with many kinds of suprasegmental information, such as rhythm, intonation and word primary stress and no study until the present has evaluated whether these different types of suprasegmental information could act all as last-resource segmentation heuristics, only prevailing over (statistical) segmental information in physically impoverished conditions. In Experiments 2 and 3 presented in chapter 2 of the present thesis, we used two ALL settings to investigate the impact of physical noise on the weighting of prosodic information and statistical segmental information. In Experiment 2, we investigated the weighting of TPs computation and IPs right-edge Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 46 (i.e., acoustically marked by syllable lengthening and falling pitch). In order to evaluate whether the role of prosody in speech segmentation would depend on the particular type of suprasegmental information available in the speech stream, in Experiment 3B the suprasegmental information at study corresponded to a languagespecific prosodic cue: word primary stress. 1.5.2. The listening conditions One of the innovations and fundamental aspects of Mattys et al.’s (2005; Mattys, 2004) proposal regards the importance of listening conditions in the weighting of various speech segmentation cues. Until now this impact was only evaluated through physical degradation of the signal (i.e., physical noise), even though Mattys et al. (2005) already suggested that listening conditions go beyond this physical quality. Under this view, impoverishing listening conditions while maintaining intact the physical quality of the signal would also affect the kind of cues predominantly used in segmenting speech. A situation of cognitive noise could be achieved by reducing the available attentional resources necessary for extracting AL-words from the continuous stream. In Experiment 4 presented in Chapter 4 of the present thesis, we used an incidental ALL setting to investigate the impact of cognitive noise (i.e., attentional load) on the weighting of two types of sublexical cues to speech segmentation: statistical information – TPs between adjacent syllables – and subsegmental information – coarticulation. Using the same sublexical sources of information studied in Experiment 1 (see chapter 2), we were also able to evaluate whether the type of noise (physical in Experiment 1, cognitive in Experiment 3) would have a Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 47 differential impact on the hierarchical weighting of those cues, represented at the same middle tier in Mattys et al. (2005) proposal. Furthermore, we also evaluated whether attention-load (i.e., cognitive noise) has a graded rather than an all-or-none impact on speech segmentation, as already suggested for physical noise by Mattys (2004; Mattys et al., 2005). Until now, no study has ever evaluated the impact of cognitive noise on speech segmentation driven by coarticulatory information. However, for statistical information, two studies (i.e., Saffran et al., 1997; Toro, Sinett, & Soto-Faraco, 2005) evaluated the relationship between statistical learning (TP-based speech segmentation) and attentional resources. Statistical Learning and Attention The idea of a statistical segmentation procedure that is sensitive to temporal contingencies of the speech stream, operating largely independently of awareness and control (and hence in an automatic fashion) is attractive. This mechanism would allow the effortless acquisition of such a powerful cognitive ability as language, without the need to call on conscious (and controlled) strategies or processes. However, this does not imply that word segmentation based on TPs occurs automatically, regardless of the available attentional resources. Several studies suggest that statistically-driven segmentation can occur in the absence of attentional allocation to the AL stream. Human infants (Saffran et al., 1996a) and non-human primates (Hauser et al., 2001), who cannot be formally instructed to attend to the speech stream, are sensitive to the available TPs. In Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 48 addition, human adults are able to use the statistical information conveyed by the speech stream even in incidental passive listening (Toro et al., 2005) and when they have to perform a concurrent drawing task (Saffran et al., 1997). In an incidental ALL condition, in which participants were not informed about the subsequent performance of a task about their knowledge of the AL words, Saffran et al. (1997) reported that statistical speech segmentation was not affected by participants’ performance of an independent, primary visual color-drawing task, with the AL being presented as “background” nonsense auditory stream (Saffran et al., 1997). Since participants were told neither that the stream consisted of an AL, nor that they would be tested later in any way, the observation of an ALL effect in these conditions has been interpreted as evidence that statistical learning proceeds incidentally (Saffran et al., 1997) “as a by-product of mere exposure” (Saffran et al., 1999). Yet, incidental does not mean automatic; even if statistical learning occurs when it is not part of task requirement, the specific learning conditions on Saffran et al. (1997) do not permit the conclusion that this mechanism is automatic. Indeed, simply instructing participants to focus attention on another task is not sufficient to prevent attention-switching to the simultaneously presented speech stream. In situations of low perceptual load (cf. Lavie, 1995), such as the one used by Saffran et al. (1997), any attentional capacity not taken up by the processing of task-relevant stimuli (the color-drawing task) might involuntarily “spill over” to the perception of task-irrelevant information (e.g., Lavie, 1995, 2005). To ensure that attention is actually diverted from the AL stream, it is necessary to use another concurrent task of high processing load that engages all the available attentional resources (cf. Lavie, 2005). Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 55 in continuous speech streams. Whether this lexicalization effect would also hold true for the output of statistical segmentation in a continuous speech flow was investigated in Experiments 8 and 9 of the work presented in chapter 6 of this thesis. In both experiments two types of segmental information were available in the stream during the familiarization phase, namely the AL-statistical information (i.e., the TPs between adjacent AL syllables, which were always higher within than between TP-words) and the wordlikeness of trisyllabic AL-stimuli (i.e., according to both statistical segmental information on listeners native language and by the fact that AL-stimuli strongly overlapped from onset with real words of listeners’ native language – i.e., Base-Words). In the second phase of both Experiments 8 and 9, listeners performed the conventional forced-choice test and a Lexicalization test (i.e., lexical decision task on Base-Words, similar to the one used by Gaskell & Dumay, 2003a; 2003b). For clarifying whether any lexical engagement of the outputs of the segmentation procedure adopted by listeners was transient (observed only immediately after) or if it was stable and long-lasting, listeners also performed the ALL-test and the Lexicalization test post-one week, with no intervening ALfamiliarization period. The main difference between Experiments 8 and 9 was the congruency of those two available segmental cues during the AL familiarization phase. In Experiment 8, adult listeners were familiarized with one of two ALs constituted by trisyllabic TP-words that differed from the Base-Words in the last syllable, after the Base-Words’ uniqueness point (UP). Thus, the two types of segmental information (i.e., AL-TPs and the wordlikeness of TP-words) suggested the same segmentation outputs in a congruent cues condition. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 56 In Experiment 9, adult listeners were familiarized with two new ALs with the same phonological repertoire used in Experiment 8, and hence with the same two segmental sources of information available in the stream, but in an incongruent cues condition. In this experiment, the AL Part-words (even though these stimuli occurred adjacently in the stream during the familiarization phase, they had lower TPs than TP-words) were the AL-stimuli that differed from Base-Words after the UP. Thus, while AL-statistical information suggested that TP-words were the “units” of the new language, the wordlikeness of AL-stimuli suggested that part-words were the ones. Thus, we were able to evaluate whether listeners’ speech segmentation outputs would be always lexicalized, even when another incongruent cue was available in the stream, as in Experiment 9, or whether the lexicalization process is highly dependent of the congruency of the segmentation cues available in the speech stream, as in Experiment 8. In sum, in the work described in the present thesis we evaluated the weighting of different sublexical sources of information in speech segmentation, as regards the simultaneous availability of different types of sublexical cues (represented at the same and at different tiers in Mattys and colleagues theoretical proposal), as well as the role of listening condition in adulthood. In order to ensure that the differential weighting, or in Mattys et al.’s terms the hierarchical organization, of the available segmentation cues was not dependent on the activation of lexical (high-level) information, we mainly adopted the ALL paradigm through all work. The ALL paradigm has revealed itself ideally suited for studying the weighting of various sublexical cues in the absence (or almost so) of high-level Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 57 information, in the context of a full mature speech perception system. Importantly, the present work also directly evaluated the potential usefulness of ALL paradigm combined with other conventional experimental techniques, as well as the status of the sublexical segmentation outputs on a full mature consolidated mental lexicon. At the final end, in this series of experiments we had always as backdrop the natural settings’ perspective: a rich set of multiple cues available in the stream, assisting speech segmentation. Thus if we aim to obtain a realistic perspective of how these cues are weighted, it is important to study them in conjunction and also in different listening conditions. On the one hand, noise should no longer be confined to the physical quality of the signal. On the other hand, this enables the evaluation of speech segmentation’s optimization and the preponderance of cues over other incongruent ones. Chapter 1. General Introduction Chapter 1. General IntroductionChapter 1. General Introduction Chapter 1. General Introduction 58 59 PART II EMPIRICAL CHAPTERS Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 60 Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 61 Chapter 2. Statistical information and Coarticulatory cues to word boundaries: A matter of signal quality7. We investigated how statistical information  transitional probabilities, TPs  interacts with another sublexical cue  coarticulation  to word boundaries, and examined the impact of signal quality on the weighting of these cues, in this Experiment 1. In an artificial language learning setting, with phonetically intact speech, coarticulation overruled TPs, suggesting the prevalence of subsegmental, lowlevel information. However, while the role of coarticulation in segmentation was highly modulated by signalquality, TPs were very resilient to noise. When coarticulation was made unreliable by strongly degrading the input (10dB SNR), only statistical information drove segmentation. In a milder degraded (22dB SNR) condition, when some acoustic properties were still available, coarticulation was exploited, although with less reliability than in optimal conditions. These results can be interpreted according to a hierarchical approach (Mattys, White, & Melhorn, 2005) in which both the available segmentation cues and the listening conditions have an important role in speech segmentation. 2.1. Introduction As already presented in chapter one of this thesis, several signal-derived sources of information were pointed out as potential cues to word boundaries, and combining these cues to lexical-driven mechanisms may be helpful to word 7 The experiment presented in this chapter was published in Fernandes, Ventura, and Kolinky (2007). Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 62 segmentation (e.g., McQueen et al., 1994; Norris et al., 1995; see section 1.1.2. of chapter 1). However, since each of these signal-derived cues is only probabilistically associated with word boundaries, if we aim to obtain a realistic perspective of how listeners deal with multiple segmentation cues, it is essential to study them not only in isolation, but also in conjunction. Furthermore, according to Mattys and colleagues (Mattys, 2004; Mattys et al., 2005), the involvement of any segmentation cue, either lexicallyor signal-derived, is a graded rather than an all-or-none phenomenon. Mattys and colleagues defined a so-called hierarchical organization of both lexically-driven and signal-derived sources of information represented at three tiers, as already described in section 1.4. of the previous chapter. However, since Mattys’ (Mattys, 2004; Mattys et al., 2005) framework is quite recent, some of the possible interactions between the several information types available to the listeners remain underspecified. In particular, it is not clear how the various types of sublexical information interact with each other. Mattys et al. (2005) showed that when coarticulation (subsegmental) and phonotactic (segmental) information (both at tier II) were available, indicating the same segmentation points, their segmentation hypotheses were the ones used by adult listeners, in two scenarios: (i) when lexical information was not available in the signal; (ii) when signal quality was sufficiently degraded as to reduce reliance on the semantic context, but, at the same time, allowing a relative availability of the acoustic-phonetic aspects. But, since in Mattys et al.’s study phonotactic and coarticulatory information were never put in conflict, we do not know whether one (and which one) of these two information types prevailed. Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 63 In addition, Mattys and colleagues (Mattys, 2004; Mattys et al., 2005) posit that segmental and subsegmental information are represented at the same tier, and consequently their model confers them the same importance. Nevertheless, it is plausible that even different sublexical types of information represented in the same tier would have independent impacts on segmentation and differ by their degree of (in)dependence on signal quality. Indeed, as already posited by Mattys and colleagues, at least in some languages, segmental and subsegmental information may be weighted differently according to their word-boundaries predictability. This proposal was investigated in this chapter (Experiment 1) as regards coarticulation (a subsegmental, acoustic-dependent, cue) and TPs, a statistical cue related to a higher structural level (here, syllabic). 2.1.1. Coarticulation – a subsegmental cue Coarticulation is usually defined as a change in the acoustic-phonetic content of a speech segment due to anticipation or preservation of adjacent segments (e.g., Kühnert & Nolan, 1999; see section 1.1.2. of chapter one). Although all fluent speech is coarticulated, the extent of coarticulation between adjacent segments is influenced by the presence of a prosodic boundary. There is, generally, more coarticulation within words than between them (e.g., Byrd, 1996; Byrd & Saltzman, 1998), and segment strength may convey information about the local coherence vs. disjuncture in connected speech (Fougeron & Keating, 1997; Keating, in press). Indeed, domain-initial strengthening is a general phenomenon, found in both stressed and unstressed syllables (Cho & McQueen, 2005). Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 64 The importance of coarticulation in speech segmentation has been demonstrated, in optimal listening conditions, from an early age on (Johnson & Jusczyk, 2001) until adulthood (Mattys, 2004; Mattys et al., 2005). However, it seems to be strongly affected by the superimposition of noise to the speech input (Mattys, 2004). The main reason for this sensitivity to noise is possibly related to the nature of coarticulatory information (Mattys et al., 2005). Indeed, the use of a coarticulatory segmentation cue use depends on the ability to process fine-grained, low-level, acoustic properties. Noise probably masks such information easily, reducing the effectiveness of coarticulation in speech segmentation. 2.1.2. TPs – a segmental cue Another important source of information that can help locate word boundaries is the statistical information conveyed by the speech stream. The extraction of statistical information is a general mechanism (e.g., Saffran et al., 1999; Conway & Christiansen, 2005) that seems to be universal in nature. As a matter of fact, within any language, the TP from one unit of sound (e.g., a syllable) to the next will generally be the highest when the two units follow one another within a word, whereas the transitional probabilities spanning a word boundary will be relatively low (see section 1.1.2. of chapter one). Human listeners in different linguistic stages of development can use this statistical information to locate “word” boundaries (e.g., Saffran et al., 1996a, 1996b8). 8 While Saffran and colleagues considered statistical learning as based on statistical computations (Saffran et al., 1996a; Saffran et al., 1996b), another mechanism has been proposed: the formation of chunks (Perruchet & Vinter, 1998). It is difficult to decide between these interpretations on the grounds of their explanatory power (Perruchet & Pacton, 2006), but learning effects based on nonadjacent TPs (e.g., Kuhn & Dienes, 2005; Onnis et al., 2005) seem to challenge the chunking explanation. Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 71 The three natural versions of the AL differed in what regards the type of information available in the speech stream, as explained in the Introduction and illustrated in Table 1, in which an orthographic translation of a sample of the speech stream in each of the three versions of the AL is presented. Pretest 1: Natural vs. Synthesized Speech A pretest checked whether the use of naturally produced utterances, chosen here to achieve realistic coarticulation, had turned the task easier (Thiessen & Saffran, 2003) than the use of synthesized speech, which is more common in AL learning experiments. Synthetic stimuli (created using text-to-speech MBROLA software, cf. Dutoit, Pagel, Pierret, Bataille, & van der Vrecken, 1996) were presented to 21 independent volunteer undergraduates in the same AL experiment as the main one, but using only the single cue condition. Syllables were concatenated (with no other cues to word boundaries than TPs) using an EP female diphone database (available at http://tcts.fpms.ac.be/synthesis/mbrola.html) at 22.05 kHz and with a speech rate of approximately 270 syllables per minute. We evaluated two input-intelligibility conditions (intact speech and 22dB SNR), since a facilitation effect might appear for natural in comparison to synthetic speech in the more difficult (noisy) situation. This was not the case. The TP-word preference was similar in the synthesized and natural single cue conditions in the two signal intelligibility situations (with intact speech: t(19) = .73; with 22dB SNR: t(13) = .56; p > .10 in both cases). Thus, both in intact and noisy conditions, the natural speech stimuli used in the main experiment induced statistical learning based on TPs in the same way as the synthetic material more commonly used in AL learning studies does. Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 72 Pretest 2: Identification in Noise - degraded signal conditions In order to create the two degraded signal conditions, white noise was superimposed to each block of all three natural versions of the AL at 22dB SNR and at 10dB SNR. These SNR were selected on the basis of a pretest ran on an independent group of 23 volunteer undergraduate students. In this pretest, five between-subjects conditions were used: intact speech; 22dB SNR; 10dB SNR; 5dB SNR; and 0dB SNR. All the words and part-words of the AL were presented in randomized order, one at a time. They were played through headphones at 76dB SPL (which is approximately the level of conversational speech). Participants were informed that on each trial they would hear a pronounceable trisyllabic nonsense sequence and were required to write it down. This allowed us to choose the SNR that would reduce the stimuli intelligibility by approximately 50%, a value similar to the one used by Mattys (2004), operationalizing “intelligibility” as the total number of stimuli correctly identified. As expected for unfamiliarized listeners, correct responses did not significantly differ between the TP-words and part-words of the AL (t(22) = -1.534, p > .10). The best performance was observed for intact speech (91.7% correct identification, on the average). The 22dB SNR reduced performance by nearly 50%, leading to 47.9% correct identification, on the average, while the other SNR conditions induced much poorer performances (16.6% for 10dB SNR; 5% for 5dB SNR; 0% for 0dB SNR). In addition, the mean number of phonemes correctly identified (in their correct order) with 22dB SNR was 5.37 in sequences of 6 phonemes each. This indicates that the incorrect responses were not due to a large inability to identify any of the phonemes of the stimuli, as was the case, for example, Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 73 in the 0dB SNR condition. Thus, although the 22dB SNR impoverishes the signal, it does not affect phonetic information in a way that turns it inaudible. This makes the 22dB SNR a perfect option, since the signal is degraded, but noise is not so strong to disable the availability of many phonetic aspects of the speech input. Nevertheless, there are important differences between this pretest task, in which naïve participants were required to identify the TP-words and part-words of the AL presented one at a time, and the AL learning situation, in which participants are repeatedly exposed to these stimuli embedded in longer streams of speech. Thus, possibly, the intelligibility level obtained with a specific SNR in the identification in a noise task does not correspond to the degree of intelligibility obtained with the same SNR in an AL learning task. One should note that, usually, for words it is a 0dB SNR that reduces word intelligibility by about 50%. Although our AL material is rather different from real words, it is plausible that repeated exposition to the same material would lead to a higher level of intelligibility in the AL situation than the one used by Mattys (2004). Thus, another condition with superimposition of a higher level of white noise was also chosen, namely 10dB SNR. At least in the pretest, the 10dB SNR had a strong impact on the identification of the AL stimuli, reducing dramatically the correct identification not only of the nonwords but also of the phonemes that constitute them (3.9 out of 6). The forced-choice test included the six TP-words and six part-words. These part-words consisted of syllables of two different words that had appeared adjacently in the speech stream. Three part-words (Part-words 3#12 in Table 1) were formed by the last syllable of one word (e.g., /b5/ which is the last syllable of /luf5b5/) and the Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 74 first two syllables of the next word (e.g., /kil5/ which are the first two syllables of /kil5bu/). The others (Part-words 23#1 in Table 1) were constituted by the last two syllables of a word and the first syllable of the next word. For all conditions and groups, the stimuli used in the test phase were produced by concatenating the three syllables that constituted each word or part-word without white noise superimposed, thus avoiding responses based on acoustic matching between the stimuli of the familiarization and test phases. 2.2.3. Procedure Presentation in the familiarization phase was done with Windows Media Player with all the auditory stimuli presented at a comfortable level through Sennheiser HD 280 headphones. For the test phase, stimuli were also presented through headphones, with presentation, timing and data collection controlled by EPrime 1.1 (Schneider, Eschman, & Zuccolotto, 2002a; 2002b). All participants were instructed to listen to a new language that contained “words”, but no meaning or grammar. Their task was to find out what words constituted the new language. No information about the structure, phonology or length of the words was given. Participants were informed that the experiment consisted of three short listening blocks, followed by a test of their knowledge of the words that constituted the language. Participants in the degraded signal conditions were warned of the poor signal quality. After each of the 7-minutes blocks, a 5minutes break was provided. After the listening phase, participants were presented with a two-alternative forced-choice test. Each trial started with a warning tone, Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 75 followed by two trisyllabic strings, separated by 500 ms of silence, presented through headphones. One of these strings was a word from the AL, the other was a part-word. Each word was paired exhaustively with each part-word, rendering 36 trials. Immediately after participants gave their answer, another trial began. If no answer was registered after 10 seconds, the next trial began. Participants were told to always provide an answer, even if not totally sure about their decision. Nevertheless, accuracy was also emphasized. The test began with four practice trials, in which an animal and an environmental sound were presented and participants had to decide which of the two stimuli corresponded to the animal sound. Feedback was only provided for practice trials. Order of presentation of test trials was randomized for each participant, and order of presentation of the stimuli within trials was counterbalanced within each group. 2.3. Results and Discussion The percentage of TP-word choices was computed for each participant. In Table 2, the average AL learning performances are presented broken down by TP level (high vs. low level of TP of the TP-words). In all input intelligibility conditions, participants presented with the single cue chose the TP-word significantly more often than the part-word: with intact signal, 67% (SD= 6.4) (t(8) = 7.917, p < .0001); with milder degraded signal (22dB SNR), 62% (SD=8.2) (t(5) = 3.605, p < .025); with strongly degraded signal (10dB SNR), Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 76 58% (SD=7.1) (t(8) = 3.53, p < .01). A significant learning effect was also found in all intelligibility conditions for the participants presented with congruent cues: with intact signal, 84% (SD=9.4) (t(10) = 11.985, p < .001); with 22dB SNR, 72% (SD=12.2) (t(9) = 5.69, p < .001); with 10dB SNR, 61% (SD=4.7) (t(8) = 6.897, p < .001). Table 2: Performance pattern (proportion of TP-Word responses, in percentage) according to cue condition (single cue; congruent cues; incongruent cues) and Signal Quality (intact speech; 22dB SNR; 10dB SNR), in Experiment 1, considering the TP gradient of the TP-words of the AL (High-TP; low-TP). SIGNAL QUALITY Intact 22dB SNR 10dB SNR AL Condition High TP Low TP High TP Low TP High TP Low TP Single Cue 66.1 66.6 62.2 62.2 55.5 60.5 Congruent Cues 82.2 85.0 66.6 77.7 57.7 63.3 Incongruent Cues 39.4 41.6 50.0 48.1 55.5 59.4 Chance level corresponds to 50%. In sharp contrast with this pattern, with intact signal, participants presented with incongruent cues discarded the TP-words, with an average of only 41% TPword choices (SD=10.5). Thus, in this condition participants chose the part-words as the lexical units of the new language significantly more often than the TP-words (t(11) = -3.040, p = .01). With milder degraded signal, performance with incongruent cues did not differ from chance (t(11) = -.389, p = .704), with 50% TP-word choices (SD= 12.5). Thus, no learning effect was observed. It was only in the strongly degraded condition that participants presented with incongruent cues significantly preferred the TP-words over the part-words, reaching, on average, 57% TP-word choices [SD=7.6; t(8) = 2.794, p < .025]. Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 77 In order to evaluate directly both the weighting of the two studied cues and the impact of signal quality on this weighting, we ran an ANOVA including cue condition (single cue; congruent cues; incongruent cues) and signal quality (intact speech, 22dB SNR and 10dB SNR) as between-subject factors. In addition, we included the high vs. low level of TPs of the TP-words as a within-subject factor in order to check whether there was a TP gradient. The cue condition effect [F(2, 78) = 49.1, p < .001, MSe = 5.66] was significantly modulated by signal quality, as revealed by the interaction between these two factors [F(4, 78) = 11.551, p < .001]. No other main effect or interaction was significant [all Fs < 2, p > .1]. We further investigated the nature of the significant interaction through pairwise comparisons, using the Bonferroni-corrected alpha rate of .017. As clearly illustrated in the left-hand part of Figure 2, with intact signal a main effect of cue condition was found [F(2, 29) = 64.77, p < .001, MSe = 10.87], with congruent cues leading to better performance than both the single cue [F(1, 29) = 16.55, p < .0005] and the incongruent cues [F(1, 29) = 128.74, p < .0001] conditions. Not surprisingly, the incongruent cues condition, which led participants to prefer part-words over TP-words, also differed from the single cue condition [F(1, 29) = 43.47, p < .001]. As can be seen in Figure 2, with milder degraded (22dB SNR) signal, [main effect of cue condition: F(2, 25) = 11.24, p < .001, MSe = 17.33], incongruent cues led to a lower performance than congruent cues [F(1, 25) = 20.29, p < .001]. As displayed in the right-hand part of Figure 2, contrary to what had been observed in the former two input-intelligibility conditions, with a strongly degraded Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 78 (10dB SNR) signal, all three cue conditions led to a similar performance level [F(2, 24) = .73, p =.48, MSe = 2.815]. Figure 2: Performance pattern (proportion of TP-Word responses, in percentage) according to cue condition (single cue; congruent cues; incongruent cues) and of the signal quality (intact speech; 22dB SNR; 10dB SNR) – Experiment 1. Vertical bars denote 0.95 confidence intervals. Chance level corresponds to 50%. The data show that there was actually no major impact of signal quality in the single cue condition [F(2, 21)= 3.33, p >.05, MSe = 6.58]. Thus, the statistical learning mechanism operating on TPs between adjacent syllables seems very resilient to noise, allowing speech segmentation to be almost as efficient when the signal is quite distorted (i.e., 10dB SNR) as when it is highly intelligible [F(1,78) = 3.84, p > .05]. In contrast, signal quality significantly affected performance in the congruent cues condition [F(2, 27) = 14.94, p < .001, MSe = 11.46]. A significant linear trend Single Cue Congruent Cues Incongruent Cues Intact 22dB SNR 10dB SNR Speech Quality 0,00 10,00 20,00 30,00 40,00 50,00 60,00 70,00 80,00 90,00 100,00 % of TP-Word choices Single Cue Congruent Cues Incongruent Cues Intact 22dB SNR 10dB SNR Speech Quality 0,00 10,00 20,00 30,00 40,00 50,00 60,00 70,00 80,00 90,00 100,00 % of TP-Word choices Single Cue Congruent Cues Incongruent Cues Intact 22dB SNR 10dB SNR Speech Quality 0,00 10,00 20,00 30,00 40,00 50,00 60,00 70,00 80,00 90,00 100,00 % of TP-Word choices Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 79 [F(1, 78) = 30.05, p < .001] suggests that congruent cues to word boundaries are integrated in good listening conditions, allowing the optimization of the speech segmentation process. But integration was affected and hence the redundancy gain11 largely reduced as signal degradation increased. Indeed, the number of TP-words choices was significantly higher with intact speech than with noise superimposition [22dB SNR and 10dB SNR, F(1, 78) = 8.47 and = 30.05, respectively, p < .005]. In addition, performance was also better with milder (22dB SNR) than with strongly (10dB SNR) degraded speech [F(1, 78) = 6.72, p = .011]. This pattern of progressive reduction of the redundancy gain seems related to the strong sensitivity of coarticulatory cues to noise. Indeed, as reported above, the redundancy gain observed with congruent cues compared to the single cue condition was significant only with intact speech. It was still (numerically) present but not statistically significant anymore with mildly degraded speech (22dB SNR), and no longer observed at all with strongly degraded speech (10dB SNR). The sensitivity to noise of coarticulatory cues to physical noise is even more clearly revealed by the performance pattern observed with incongruent cues. As a matter of fact, with incongruent cues the effect of noise was significant [F(2, 30) = 6.16, p < .01, MSe = 14.53], and modulation of performance by signal quality is reflected by a significant linear trend [F(1, 78) = 15.74, p < .001]. As already reported, participants relied on coarticulation rather than on statistical information only when exposed to intact speech, a listening condition that differed significantly from strongly degraded (10dB SNR) speech [F(1, 78) = 15.74, p = .0001]. Indeed, with low noise intensity (22dB SNR), the inconsistency between the two types of 11 The term “redundancy gain” is used to emphasize the improvement in the segmentation process (revealed by the superior AL learning performance in the congruent cues condition with intact speech) as the outcome of the availability of different consistent segmentation cues in the stream. Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 80 information disrupted performance, but the influence of coarticulation totally vanished only with the most degraded input (10dB SNR), a situation in which listeners considered only statistical information, performing at the same level as listeners exposed to either a single cue or congruent cues. 2.4. General Discussion In Experiment 1, within an ALL setting, we investigated, through three signal-quality conditions, the relative impact of two types of sublexical information in speech segmentation: statistical information – TPs between adjacent syllables, and subsegmental information – coarticulation. The results presented in this chapter suggest, in accordance with Mattys and colleagues’ (2004; Mattys et al., 2005) proposal, that the segmentation process that listeners adopt varies as a function of both the types of cues available in the speech flow and signal quality. Indeed, with phonetically intact speech, coarticulation was a powerful segmentation cue, able to drive the segmentation process. Most importantly, coarticulation overruled statistical information when both cues were put in conflict in the speech stream, which is in line with Johnson and Jusczyk’s (2001) infant data. Mattys (2004; Mattys et al., 2005) had already demonstrated coarticulation reliability in adults, when this cue was in conflict with lexical stress. Our results thus add to this evidence. Thus, in good listening conditions, coarticulation is given priority to either lexical stress (Mattys, 2004; Mattys et al., 2005) or a general segmental statistical cue (TPs, in the present experiment). It is not clear however, if it is either the local coherence given by coarticulation or the perception of edges between concatenated Chapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word bounChapter 2. Statistical and Coarticulatory Cues to word boun Chapter 2. Statistical and Coarticulatory Cues to word boundaries dariesdaries daries 87 different patterns of coarticulation that often reflect the phonetic contrasts that are emphasized (see Manuel, 1999, for review), the importance of the languagespecificity vs. universality of the cues may be assessed by contrasting the role of TPs to the role of language-specific as opposed to language-general patterns of coarticulation in speech segmentation. Contrasting universal prosodic cues used by Shukla et al. (2007) to languagespecific prosodic cues like word stress patterns used by Mattys (2004) and Mattys et al. (2005) may also shed light on the relevance of this distinction. This was done in Experiments 2 and 3B of the work presented in this thesis (see Chapter 3). In the next chapter we have evaluated the weighting of both domain-general statistical mechanisms and speech-specific cues, such as the universal prosodic properties examined by Shukla et al. (2007) under physically degraded conditions (see Experiment 2, presented in chapter 3). Indeed, with intact speech Shukla et al. observed that phrasal prosodic cues seem to act as a filter, suppressing possible word-like sequences (trisyllabic sequences with high TPs) that straddle two prosodic constituents (this proposal is presented in more detail in Chapter 3). Whether this would hold true in noisy situations was tested in the experiments presented in the next chapter of this thesis. Importantly, although Mattys (2004; Mattys et al., 2005) already demonstrated that lexical stress acts as a last resource segmentation heuristic in speech segmentation for English adult listeners, we do not know whether this could also hold truth for EP listeners. As a matter of fact, this was evaluated in Experiment 3B of the next chapter. Thus, in Experiment 3 we were able to evaluate directly whether the nature of sublexical cues (i.e., their domain generality and their Chapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundariesChapter 2. Statistical and Coarticulatory Cues to word boundaries Chapter 2. Statistical and Coarticulatory Cues to word boundaries 88 role in particular languages) is in fact important on their weighting in different listening conditions. In summary, the AL learning patterns observed in Experiment 1 of the preent work have shown three important facts: (i) the modulation of coarticulation reliability by signal quality; (ii) the high resilience of statistical information based on TPs to noise superimposition; (3) the strong signal contingency of the weighting of the cues used to segment speech into words. This pattern of results can be well accommodated by Mattys’ (2004; Mattys et al., 2005) hierarchical proposal, and highlights the importance of studying speech segmentation in the context of multiple cues (e.g., Christiansen & Curtin, 2005) and in different listening conditions (Mattys, 2004; Mattys et al., 2005). Importantly, an integrated approach must be able to apprehend the role of sublexical information in speech segmentation, how sublexical cues interact with lexical and supra-lexical information, and how the weighting of several cue types is affected by listening conditions. Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 89 Chapter 3. Universal and Speech-specific Prosodic cues in Artificial Language Learning: The role of physical noise. Until now, the weighting of general domain and speechspecific cues in speech segmentation was left largely unspecified. In the present chapter, the impact of physical noise on the weighting of three sublexical cues: transitional probabilities (TPs), universal prosody and lexical stress, was investigated. In artificial language learning, while both universal prosody and TPs were highly resilient to signals’ physical degradation, with the former prevailing over the latter, being able to drive the segmentation process even when incongruent cues were available in the stream (Experiment 2), stress pattern effects only emerged in degraded conditions (Experiment 2 and Experiment 3B). Thus, speech segmentation does not seem to be the product of one preponderant cue acting as a filter of the outputs of another lower weighted cue. Instead, it mainly depends on the weighting of cues, according to their nature and the listening conditions . 3.1. Introduction The present chapter will focus on two types of sublexical cues that are available since a precocious phase in linguistic development, in particular: prosodic (suprasegmental) and statistical (segmental) sources of information. As already refer in section 1.1.2. of chapter one of the present thesis, prosody Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 90 was one of the first sublexical speech segmentation cues to be proposed (e.g., MMS, Cutler & Norris, 1988; Gleitman & Wanner; 1982), and listeners are sensitive to it from the very beginning of language onset. Adult listeners are also sensitive to several types of prosodic cues (for a review see Cutler, Dahan & van Donselaar, 1997), such as IPs’ edges (e.g., Shukla et al., 2007); phonological phrases boundaries (e.g., Christophe, Gout, Peperkamp, & Morgan, 2003), metrical units (e.g., Cutler & Norris, 1988), and lexical stress (e.g., Mattys, 2000), and these cues play some role in lexical activation (e.g., Christophe, Pepperkamp, Pallier, Block, & Mehler, 2004; McQueen et al., 1994; Norris et al., 1995; Salverda, Dahan, & McQueen, 2003). Another important speech segmentation cue is the statistical information conveyed by the speech stream. In particular, many studies have focused on TPs between adjacent syllables. Indeed, within any natural language, the TP from one syllable to the next is generally higher within a word than between words (e.g., Perruchet & Peereman, 2004; Swingley, 2005). The ability to track TPs appears to be precocious (e.g., Kirkham et al., 2002; Thiessen & Saffran, 2003), rapid (e.g., Saffran et al., 1996a), involuntary (e.g., Saffran et al., 1997; but see Toro et al., 2005) and age-independent (e.g., Saffran et al., 1997). Since TPs computation does not require any, even minimal, lexical knowledge, it might allow learners to segment speech from the discovery of troughs in the TPs distribution between adjacent syllables, acting as a pivot mechanism (cf. Mattys et al., 2005) in the acquisition of not only words but importantly also other word-boundary cues (Thiessen & Saffran, 2003). The integration of multiple available segmentation cues may provide Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 91 evidence about linguistic aspects that cannot be derived from any single source, promoting the optimization of speech segmentation (see Christiansen, Allen, & Seidenberg, 1998; Christiansen & Curtin, 2005). Therefore, the realistic understanding of natural speech segmentation and how cues are weighted on its assistance can only be achieved if the role of different cues, simultaneously available, is evaluated in both congruent and incongruent conditions, namely when these cues suggest either the same or incompatible segmentation hypotheses as regards the location of word boundaries, respectively. 3.1.1. The interaction between Prosody and TPs in speech segmentation Recently, a considerable amount of research has been devoted to the integrated study of statistical and suprasegmental cues. In optimal listening conditions (i.e., with intact speech), prosodic information (e.g., strong syllables, syllable lengthening, lexical stress) seems to be underestimated by adult listeners when other speech segmentation cues are available in the stream. Using wordspotting, McQueen (1998) found that phonotactic legality is a preponderant cue able to drive speech segmentation processes independently of metrical prosody. Within an artificial language learning (ALL) setting, Saffran et al. (1996b) also showed that statistical learning is not disrupted by the presence of an incongruent prosodic cue (i.e., initial syllable’s lengthening). As a matter of fact, in such a condition American listeners were as able to extract from the stream the AL words based on the high TPs between their syllables (henceforth, TP-words) as listeners only exposed to the Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 92 statistical information. Statistical learning immunity to incongruent prosodic cues is also observed with listeners from other linguistic backgrounds (e.g., French listeners: Vroomen et al., 1998). Using the cross-modal priming effect in a lexical decision task, Mattys (2004; Mattys et al., 2005) also demonstrated that listeners disregarded lexical stress when this suggested segmentation hypotheses that were incongruent with coarticulation, phonotactics, or lexical context. Even when prosodic cues are congruent with other sources of information, the expected benefit or redundancy gain (cf. Fernandes, Ventura, and Kolinsky 2007; see chapter two of the present thesis) is far from being consistently observed. On the one hand, Saffran et al. (1996b) reported that when TPs and prosodic information (final syllable lengthening) were congruent, listeners’ ALL performance was improved in comparison to a condition in which only statistical information was available in the stream. The same benefit was found for Finish and Dutch listeners (with initial stressed syllables marked by a F0 peak; Vroomen et al., 1998) and for French listeners (with lengthening and/or F0 peak of the last syllable; Bagou et al., 2002). But, on the other hand, there are numerous studies in which the congruency of prosody with other cues did not have any (positive or negative) impact. In ALL, Valian and Levitt (1996) only found a benefit promoted by “phrase prosody” (i.e., the rising pitch contour on the first two-word phrase of a sentence and a falling pitch on the other last two-word phrase) when listeners were unable to use other cues, since neither marker frequency nor a reference field were available. Mattys (2004; Mattys et al., 2005) also showed that when coarticulation, phonotactics, or lexical cues were available in the stream, the same pattern of results was found whatever the congruency of lexical stress with these cues. Toro-Soto, Rodríguez-Fornells and Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 93 Sebastián-Gallés (in press) evaluated whether Spanish listeners’ ALL performance would beneficiate from the availability of an increasing pitch in the “stressed” syllable of TP-words. Surprisingly, while listeners familiarized with the AL with TPwords stressed on the penultimate syllable (i.e., the default lexical stress pattern in Spanish) were not able to learn the TP-words, when stress occurred either in the first or in the last syllable of trisyllabic TP-words (which are also legal stress patterns in Spanish), listeners performed at a similar level than when only TPs were available in the stream. However, Toro-Soto et al.’s results need to be interpreted with cautious, since it is not clear that pitch is a strong correlate of lexical stress in Castilian Spanish. In fact, in Spanish, duration seems to be the strongest acoustic correlate of lexical stress, regardless of the presence of a pitch accent (Ortega-Llebaria, 2006). The impact of prosodic cues on speech segmentation is quite different when the speech signal is degraded by noise superimposition. In this case, lexical stress is able to override any other incongruent segmentation cue, either lexically-driven (e.g., the semantic context) or signal derived, like coarticulation and/or phonotactics (Mattys, 2004; Mattys et al., 2005). Thus, speech segmentation is largely the product of the differential weighting of the types of information available in the signal and of the listening conditions (Mattys, 2004; Mattys et al., 2005; see section 1.4. of chapter one). Shukla, Nespor, & Mehler (2007) In apparent contradiction with this model, Shukla et al. (2007) recently showed that IPs (i.e., one of the highest prosodic structures; Grice, 2004) act as important segmentation cues even when TPs between adjacent syllables (an Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 94 information represented at a higher tier in Mattys et al.’s 2005 model) are available in a phonetically intact signal. Using a variant of the ALL paradigm, when IPs contour was available (with left edge of IPs marked by a raising pitch and shorten duration of the beginning syllables, and right edge marked by a falling pitch and lengthen duration of final syllables), only TP-words within IPs (i.e., at middle positions) were correctly extracted from the stream. In contrast, TP-words straddling IPs (with TP-words’ first syllables at the right-edge of one IP and the last syllable at the left-edge of the next IP) were not selected by listeners as “words” of the new language, probably because, although statistically cohesive, these TP-words straddled a prosodic boundary. Importantly, when some TP-words were aligned with prosodic edges (and thus were cohesive units on both statistical and prosodic bases), while others were in IP’s middle position (not prosodically marked and hence only cohesive units on statistical grounds), only TP-words supported by the two types of information (i.e., TP-words at IP edges) were correctly extracted from the stream. Based on these results, Shukla et al. (2007) proposed that prosody could act to filter the output of statistical computation, with only TP-words compatible with it being selected. While Shukla et al.’s (2007) results may seem at odds with the unreliability of metrical prosody observed in previous studies with intact speech, the differential weighting of cues possibly depends also on two basic factors (cf. Fernandes et al., 2007; see chapter two), the first one being domain generality. The ability to track TPs involves a domain-general learning mechanism, since TPs are also extracted in non-linguistic materials (e.g., musical tones: Saffran et al., 1999; visual sequences: Fiser & Aslin, 2001; Kirkham et al., 2002; but see Conway & Christiansen, 2005, Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 95 2006) with a possible phylogenic origin (e.g., Hauser et al., 2001). In contrast, prosody is speech-specific. Within the speech domain, including prosodic information, we can also consider the further distinction between universal cues and language-specific cues. Universal cues are probably physiologically based (Grice, 2006; Shukla et al., 2007; Werner & Keller, 1994), used in any language from the onset of development as well as by adults confronted with unknown/foreign languages (Shukla et al., 2007). In contrast, language-specific cues are properties that depend on the particular language and thus latter acquired. For example the role of IPs’ contour seems to be universal, since some of its prosodic correlates such as the ones of prosodic right-edges (which characterize IPs in many different languages; Werner & Keller, 1994) are acoustic marks of the slowing down of the articulators within a breath group, which is reflected in the signal as final lengthening and low pitch (Grice, 2006, see also Cutler et al., 1997). On the opposite, lexical (or word primary) stress, although generally correlated with duration, F0 and amplitude (stressed syllables are lengthened, have higher pitch, and are louder), is not acoustically marked in the same manner in all languages. For example, while in Finish pitch seems to be the most important correlate of lexical stress (e.g., Iivonen et al., 1998), in EP lexical stress is marked by syllable lengthening (Delgado-Martins, 2002; d’Andrade & Lacks, 1996). The location of lexical stress and whether it obeys to a fixed or varied pattern is also language-dependent: while in Finish it is always on the first syllable of a word: Iivonen et al., 1998; in EP lexical stress is by default on the penultimate syllable, although it can occur in any one of the three last syllables of a polysyllabic word (d’Andrade & Laks, 1996; Mateus & d’Andrade, 2000). Thus, while both general-domain and universal speech-specific cues are used Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 96 from the very beginning of language onset, language-specific cues will only occupy a role latter on speech segmentation, which is possibly modulated by their ability to predict word boundaries in specific languages (cf. Mattys et al., 2005). In fact, when TPs and lexical stress (i.e., a language-specific cue) suggest different segmentation points in the speech stream, while 6.5-months-old infants seem to consider TPs as the more reliable cues (Thiessen & Saffran, 2003), at 9-month-old, the weighting of these cues is reversed (Johnson & Jusczyk, 2001; Thiessen & Saffran, 2003). At this age, infants are also sensitive to statistical segmentation cues of their native language, namely to phonotactics (the relative frequencies of segments and sequences of segments in syllables and words; cf. Mattys & Jusczyk, 2001), and by 10.5 months, infants are already able to integrate multiple sources of information, while language-specific prosodic cues start to loose their previous importance (Jusczyk et al., 1999a; 1999b). In line with Mattys et al.’s (2005) proposal, these language-specific sublexical segmentation cues, lower weighted in adulthood, such as lexical stress, seem to be early acquired, having a predominant critical role within a transitory phase in infancy. After this period, they gradually loose their reliance or are supplanted by other sublexical cues with higher word boundaries predictability. 3.1.2. An overview of the experiments presented in this chapter Until now, the question of how, in speech segmentation, universal cues, like domain-general TPs and speech-specific IPs interact with language-specific cues such as lexical stress has been left unsolved. To the best of our knowledge, no study on adult listeners has ever evaluated the relative power of these three types of cues in Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 103 words”, from 0.50 to 0.58, i.e., /buk5l5/, /b5buku/, /k5fubi/). This distributional gradient is probably similar to what happens in natural languages (Saffran et al., 1996b) and might allow a fine-grained evaluation of any effect of TPs (i.e., TPgradient) in ALL. Part-words. These AL-stimuli were constituted by syllables of two different TP-words that occurred adjacently in the speech stream during the familiarization phase, and corresponded to the part-words used in Experiment 1 (see chapter two). Three part-words (i.e., Part-words 3#12; e.g., /b5kil5/) consisted of the last syllable of one TP-word (e.g., /b5/ which is the last syllable of /luf5b5/) and the first two syllables of the next (e.g., /kil5/ which are the first two syllables of /kil5bu/). The other three part-words (i.e., Part-words 23#1; e.g., /fibulu/) consisted of the last two syllables of a TP-word (e.g., /fibu/ which are the last syllables of /fufibu/) and the first syllable of the next (e.g., /lu/ which is the first syllable of /luf5b5/). The ALstimuli are presented in Appendix I. The TPs of the AL-stimuli (TP-words and part-words) were computed by averaging the two TPs associated to each stimulus, with TPs between adjacent syllables always higher within than between TP-words (0. 68 and 0.38, respectively). Familiarization phase. Four synthesized versions of the AL were created. Each version included the same sequence of syllables, divided into three listening 7minutes blocks (rendering 21-minutes). Each block was created by concatenating 105 tokens of each TP-word (1890 syllables, 630 tokens of words) with the only criterion that two tokens of the same TP-word never occurred adjacently in the stream. The difference between the single cue version of the AL and the other three versions consisted on the number of segmentation cues available in the speech Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 104 stream during the familiarization phase. In the single cue version, only TPs between adjacent syllables could help listeners to locate word-boundaries, with no acoustic cues available, since the stream presented a flat 220Hz pitch and the average duration of all syllables was equivalent (i.e., 222 ms). In the other three versions, both statistical and prosodic informations were available. The prosodic information was added by acoustically marking one particular syllable of each TP-word through its lengthening on 150ms and linear decreasing its pitch (20Hz variation: from 220 Hz to 200 Hz), while the other syllables remained unchanged (i.e., on average a duration of 222 ms; flat 220Hz pitch). The three conditions in which prosody and TPs were available differed in the congruency of the cues (see Table 3 for an illustration). In the congruent cues condition both statistical and prosodic information suggested the same wordboundaries, since the last syllable of each TP-word was acoustically marked defining a prosodic right-edge. In the two incongruent cues conditions, statistical and prosodic information suggested different segmentation hypotheses. In the incongruent cues-1st syl condition, the first syllable of each TP-word was acoustically marked defining the prosodic right-edge. Thus, while TPs suggested that TP-words were the “lexical units” of the AL, they straddle a prosodic boundary, and prosodic information suggested that the part-words (straddling TP-boundaries) that end with the prosodically marked syllable were plausible words of the new language (in this cues condition, the part-words 23#1). In the incongruent cues-2nd syl condition, the prosodically marked syllable corresponded to the second syllable of each TP-word, and thus also here the two segmentation cues were put against each other (part-words 3#12 were the prosodic cohesive units). Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 105 The mildly degraded signal condition was made by superimposing white noise at a 22dB SNR to each 7-minutes block of all versions of the AL using the same method of Experiment 1 (see section 2.2.2. of Chapter Two). Forced-choice test phase. The three syllables that constituted each ALstimulus (TP-words and part-words) were synthesized with the same average duration and 220Hz flat pitch (with no prosodic acoustic correlates available), and concatenated without white noise superimposed, for avoiding that participants could respond on the basis of any acoustic matching between the stimuli of familiarization and test phases. The forced-choice test phase included 36 trials rendered by the exhaustive combination of all six TP-words and six part-words. In half of the trials, TP-words were confronted with part-words 3#12, and in the others they were confronted with part-words 23#1. 3.2.3. Procedure Identical to the one of Experiment 1 (see section 2.2.3. of the previous chapter). Participants were tested individually or in groups of two in a soundattenuated room, as done in both familiarization and test phased of Experiment 1. Importantly, as in the previous experiment, no information about the structure, phonology, length of the words or about the available prosodic cues was given. Participants in the mildly degraded signal conditions were warned of the poor signal quality. Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 106 3.3. Results and Discussion First, we evaluated whether any TP-gradient effect was found on listeners ALL performance. In the mixed ANOVA ran on raw TP-word choices with TP-level of AL-words (high vs. low) as within-subject factor, signal quality (intact vs. degraded) and cues (single cue; incongruent cues – 1st syl; incongruent cues – 2nd syl; congruent cues) as between-subject factors, only the main effect of cue condition was significant [F(3, 82) = 25.1, p < .0001, MSe = 12,54; Mp2 =.38]. No other significant main effects or interaction was found [Fs < 2, ps > .10]. Next, we evaluated whether listeners’ TP-word choices were influenced by the type of part-words to which TP-words were confronted with in the ALL test phase. In the mixed ANOVA with part-words type (part-words 3#12; part-words 23#1) as a within-subject factor, and signal quality and cue condition as betweensubject factors, the main effect of cue condition was significant [F(3, 82) = 25.1, p < .0001, MSe = 12,54; Mp2 = .48] and was modulated by part-word type [F(3, 82) = 10.0, p < .0001, MSe = 6,41; Mp2 = .27]. The three way interaction between all factors at study was also significant [F(3, 82) = 3.4, p < .05, MSe = 6,41; Mp2 = .12]. No other effect was found [Fs < 1]. Average ALL performances are presented in Table 4, as well as local onesample t-test comparisons with chance-level. In order to specifically analyze the weighting of prosodic and statistical information, we evaluated the effect of cue condition and part-words type separately in each input-intelligibility condition (see Figure 3). Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 107 With intact signal, the main effect of cue condition was significant [F(3, 41) = 15.0, p < .0001, MSe = 12,77, Mp2 = .52]. No main effect of part-word type was found [F < 1], but the interaction between, this factor and cue condition was significant [F(3, 43) = 8.0, p < .0005, MSe = 8,57, Mp2 = .37]. Table 4: Average proportions of TP-word responses (in percentage), separately for cues condition (single cue; incongruent cues 1st syl; incongruent cues 2nd Syl; and congruent cues) and signal quality condition (Intact vs. Degraded), considering Part-words type (3#12 vs. 23#1) in Experiment 2. Standard error of the mean for each condition is presented in parentheses as well as examples of the AL-material. SIGNAL QUALITY Intact Mildly Degraded Cues Condition Part-words 3#12 (ba#kila) Part-words 23#1 (fibu#lu) Part-words 3#12 (ba#kila) Part-words 23#1 (fibu#lu) Average Single Cue (fufibu#lufaba#kilabu) 64.4 (5.7)* 61.6 (4.7)* 58.6 (5.2)* 61.1 (3.9)* 61.4** Congruent cues (fufiBU#lufaBA#kilaBU) 88.4 (2.6)** 85.4 (3.8)** 83.3 (4.1)** 69.9 (4.3)** 81.7** Incongruent cues 1st syl (FUfibu#LUfaba#KIlabu) 58.8 (4.7) 37.5 (6.3)* 49.5 (3.8) 40.7 (2.8)** 46.6 Incongruent cues 2nd syl (fuFIbu#luFAba#kiLAbu) 47.8 (9.6) 74.4 (3.3)** 64.8 (7.0)* 69.4 (5.4)* 64.1** Local one-way t-tests (in comparison to chance level, i.e., 50%): *p < .05 ** at least p < .01 In both the single cue and the congruent cues condition, listeners were able to correctly choose TP-words as the lexical units of the new language, with overall above-chance performance levels (t(11) = 2.9, p = .01 and t(10) = 13.4, p < .0001, respectively). In both cases, performance was not affected by part-words type [both Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 108 Fs < 1]. Interestingly, the absence of a part-words type effect observed in the congruent cues condition suggests that listeners were not affected by primary stress. Had this been the case, we would have found worst performance on trials in which TP-words were confronted with part-words 23#1. With congruent cues, listeners performance was driven by the available statistical and prosodic cues, which enabled them to present the best performance in comparison to both the single cue [F(1, 43) = 16.6, p < .0005] and the incongruent cues conditions [vs. incongruent cues-1st syl condition: F(1, 43) = 43.6, p < .0001; vs. incongruent-cues 2nd syl condition: F(1, 43) = 17.6, p < .0001]. Figure 3: ALL performance pattern (proportion of AL-word responses, in percentage), in Experiment 2, brokendown by Part-word Type (3#12; 23#1), according to Input-Intelligibility (Intact; Degraded) and Cues available (single cue; incongruent cues first syllable; incongruent cues second syllable; congruent cues). Vertical bars denote standard error of the mean on each condition. Chance level corresponds to 50%. In clear opposition with the good performances observed in the single cue and Intact Signal Degraded Signal Single Cue 3#12 23#1 20 30 40 50 60 70 80 90 100 TP-word choices (%) Incongruent Cues 1st Syllable 3#12 23#1 Incongruent Cues 2nd Syllable 3#12 23#1 Congruent Cues 3#12 23#1 Intact Signal Degraded Signal Single Cue 3#12 23#1 20 30 40 50 60 70 80 90 100 TP-word choices (%) Incongruent Cues 1st Syllable 3#12 23#1 Incongruent Cues 2nd Syllable 3#12 23#1 Congruent Cues 3#12 23#1 Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 109 congruent cue conditions, with intact signal overall performance was at chance in the two incongruent cues conditions [1st syl condition: t(11) = -.5, p >.5; 2nd syl condition: t(9) = 1.8, p = .1]. As clearly illustrated in Figure 3, in both incongruent conditions TP-words choices were affected by part-words type [incongruent-cues 1st syl condition: F(1, 43) = 10.3, p < .005; incongruent-cues 2nd syl condition: F(1, 43) = 13.4, p < .01]. As a matter of fact, when the first syllable of TP-words was acoustically marked, performance was at chance when TP-words were confronted with partwords 3#12. Nevertheless, this level of performance did not differ from the one displayed in both the single cue [F(1, 43) = .5, p = .50] and incongruent cues-2nd syl [F(1, 43) = 1.7, p = .20] conditions. However, when TP-words were confronted with part-words 23#1, listeners considered these part-words as the lexical units of the AL, discarding the TP-words (see Table 4). This performance level was significant below the one found in any one of the other cue conditions at study [vs. single cue: F(1, 43) = 13.2, p < .001; vs. congruent cues: F(1, 43) = 50.1, p < .0001; vs. incongruent cues-2nd syl: F(1, 43) = 28.4, p < .0001]. Thus, when the statistical segmentation outputs (i.e., TP-words) were confronted with the lower-TP units that were consistent with the universal prosodic cue (part-words 23#1 like ka-la-#fu], where ] signals the prosodic edge, see Table 3), the latter were considered as the correct items of the new language, demonstrating the preponderance of the universal prosodic cue over general-domain statistics, with the former being able to drive segmentation even in the presence of an incongruent cue. In the other incongruent cues condition (when the second syllable of TPwords was acoustically marked), TP-words were chosen as being the AL units Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 110 significantly above chance (see Table 4) only when they were confronted to partwords that were not supported by any other available cue (i.e., to part-words 23#1 like ka-la-]#fu). This performance tended to be higher than the one in the single cue condition [F(1, 43) = 3.4, p = .07]. Importantly, when TP-words were confronted to AL stimuli that have lower TPs but form cohesive units according to IP edges (i.e., part-words 3#12 like ba-#ki-la]), listeners were no longer able to chose TP-words as the “lexical units” of the new language, performing at chance (see Table 4), because in this case the universal prosodic cue defined a right-edge at the last syllable of these part-words. With degraded signal (i.e., 22dB SNR), as already observed with intact speech, the main effect of cues condition [F(3, 41) = 11.1, p < .0001, MSe = 12,31; Mp2 = .45] and its interaction with part-words type [F(3, 41) = 3.3, p < .05, MSe = 4,24; Mp2 =.20] were significant, as was the case with intact speech. No main effect of part-words type was found [F(1, 41) =2.4, p > .15]. As can be seen in Figure 3, the performances of listeners exposed to the single cue and to the incongruent cues-1st syllable condition were similar with physically degraded signal as with intact signal. Indeed, in the single cue condition, overall performance was above chance (t(11) = 2.5, p < .05), with listeners choosing more often the TP-words than the part-words, independently of part-word type [F < 1]. In fact, statistical information was able to drive the segmentation process at a similar level with degraded and intact speech [F < 1], confirming its resilience to physical noise (cf. Fernandes et al., 2007). In the incongruent cues-1st syl condition, listeners had the worst performance, in comparison to both the single cue [F(1, 41) = 5.9, p = .01] and the congruent cues [F(1, 41) = 30.7, p <.0001] conditions, and even Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 111 to the other incongruent cues condition [F(1, 41) = 15.0, p < .0005]. They presented an overall performance almost below chance (t(11) = -2.0, p = .065), which was mainly due to listeners’ preference for part-words 23#1 over TP-words (see Table 4). Thus, confronted with TP-words and part-words that were prosodically cohesive according to IP edges (part-words 23#1), the latter units were considered as the plausible “words” of the AL. Thus, even when the signal is physically degraded, the universal prosodic cue is still able to drive segmentation processing, overruling the incongruent statistical cue at a similar level to the one already observed with phonetically intact speech [F < 1]. In opposition to what was observed with intact speech, when the signal was physically degraded performance with congruent cues was modulated by part-word type [F(1, 41) = 8.3, p < .01]. Indeed, listeners were more proficient on TP-words selection when TP-words were confronted to part-words 3#12 than when they were confronted to part-words 23#1. The proportion of TP-words chosen as the AL units over part-words 23#1 (which were acoustically marked on the penultimate syllable – here, second –, obeying to the default lexical stress pattern in EP) was lower than the one found for the same condition with intact speech [F(1, 41) = 5.9, p = .01], and did not differ from the one found in the single cue condition [F(1, 41) = 1.7, p > .10]. Thus, in this particular case, although the integration of statistical and universal prosodic cues was still able to drive the segmentation process, the confrontation between TP-words (outputs of this integration) and AL units with the default Portuguese stress pattern (part-words 23#1) induced a performance cost. This was not due to a general inability of congruent cues to promote a redundancy gain, since when confronted with part-words that were not supported by any cue (i.e., part-words Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 112 3#12, see Table 4), TP-words led to a well above-chance performance that was significantly better than in the single cue condition [F(1, 41) = 8.9, p < .005], being in fact equivalent with degraded and intact signal [F < 1]. In the incongruent cues-2nd syl condition, in which TP-words were acoustically marked on the penultimate syllable, which is incompatible with a universal prosodic edge but compatible with the default stress pattern in Portuguese, the performance pattern differed sharply from the one found with intact signal. Indeed, overall performance was above chance (t(11) = 2.9, p = .01), was not affected by part-words type [F < 1], and was similar to the one found in the single cue condition [F = 1.4, p > .10]. Furthermore, preference for TP-words over the partwords 3#12 (which were marked by a prosodic edge) was stronger here than when the signal was intact [F(1, 42) = 4.5, p < .05], revealing that lexical stress was possibly assisting TP segmentation when the signal was impoverished and hence that the conjunction of these two cues enabled paroxytone TP-words to be extracted from the stream even when confronted with stimuli supported by a universal prosodic cue (part-words 3#12). In agreement with Shukla et al.’s (2007) conclusions, the present results thus corroborate two facts. First, universal prosodic cues are more preponderant in speech segmentation than general-domain TPs. Second, listeners are sensitive to different (and even incompatible) segmentation byproducts, with their ALL performance being largely dependent on the type of test-stimuli presented to them (see Shukla et al., 2007; Experiment 4). This is an important result since it demonstrates that when different segmentation cues suggest incompatible “lexical units”, listeners do not simply adopt an all-or-none procedure, filtering out from the beginning of Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 119 quite able to correctly locate the stress of trisyllabic stimuli when the stressed syllable was lengthened in comparison to the unstressed syllables. Listeners who were presented with stimuli that included no acoustic marker of primary stress often considered that these stimuli were paroxytones, but not as often as those that were presented with lengthened stressed syllables, and their firstand second-syllable responses were at chance. Besides, the EP listeners’ tendency to consider acoustically neutral stimuli as paroxytones is by itself fascinating. We cannot ensure whether this effect corresponds to a decisional bias or to a perceptual “stress illusion” (see Dupoux, Kakehi, Hirose, Pallier, & Mehler, 1999), but it clearly illustrates the fact that suprasegmental properties of the native language affects processing of unfamiliar items (e.g., Dupoux, Pallier, Sebastian-Gallés, & Mehler, 1997). Porlex Psycholinguistic Database (Gomes & Castro, 2003) Inspection of EP Porlex database (Gomes & Castro, 2003) reveals that from the 29238 entries, about 98% (i.e., 28 859) are words with two or more syllables with 59% of them being paroxytones (i.e., stressed on the penultimate syllable). Considering only words with the same phonological structure as the AL-stimuli used in the present study (i.e., CV.CV.CV; which corresponds to about 14% of the entries), the predominance of words with stress on the penultimate syllable is even more clearer (i.e., 82%), while only 16% and 2% corresponds to proparoxytones and oxytones, respectively. Taking these statistical facts into account, stressed syllables in Portuguese likely correspond to the penultimate syllable of a word, which in the case of trisyllables correspond to the second syllable. Against what is observed in Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 120 English (e.g., Cutler et al., 1997; Mattys, 2004), in Portuguese the stressed syllable does not define a possible word onset (nor a word ending), which would seem to reduce the role of lexical stress in speech segmentation. However, the results found in Experiment 2 suggest that the lexical stress pattern helps speech segmentation in degraded listening conditions. This possibility was further evaluated in Experiment 3B, which used three input-intelligibility conditions, with a strongly degraded condition (10dB SNR; the strongly degraded condition of Experiment 1, see chapter two) in addition to the intact and mildly degraded (22dB SNR) conditions already used in Experiment 2. Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 121 Experiment 3B: Domain-general vs. Language-specific Cues in Segmentation. According to the results of Experiment 2 and to previous findings (e.g., Mattys et al., 2005; Valliant & Levitt, 1996), since lexical stress effects only emerge in degraded conditions, it is possible that when a resilient, general-domain, statistical cue like TPs, is still available in the stream, the impact of the former (which is a speech-specific cue) will be to restrain the units extracted from the stream to the ones supported by all available sources of information. Shukla et al. (2007) already proposed that the role of universal prosodic cues would be acting as a filter of the potential words set forth by general-domain statistical learning. If this were holding true for language-specific prosodic cues, it would be possible that lexical stress will also filter out the statistical units that are not compatible with it when the signal is physically degraded. However, on the basis of previous findings revealing that lexical stress is not able to drive the segmentation process when other cues are still available in the stream (Mattys et al., 2005; Valiant & Levitt, 1996), it is possible that lexical stress by itself is not able to drive speech segmentation, being unable to overrule the statistical cue. The present experiment evaluated this possibility, namely the relative weighting of TPs and primary word stress, through three input-intelligibility conditions (intact speech, and mildly and strongly degraded listening conditions), using four between-participants conditions according to the location of the stressed syllable. In the TP–unstressed condition, only statistical information was available in Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 122 the stream; it actually corresponded to the single cue version of Experiment 1; in the 1st syl-stressed condition, the first syllable of some exemplars of TP-words was lengthened by 100 ms while the other (unstressed) syllables of that TP-words were reduced by 50 ms each; in the 2nd syl-stressed condition, stress was located on the penultimate (second) syllable of TP-words (which corresponds to the Portuguese lexical stress default pattern); and in the 3rd syl-stressed condition, stress was located on the last syllable of the TP-words. If the role of primary stress was similar to the one of a universal prosodic cue we would expect to find a pattern of results similar to the one observed in Experiment 1, namely, a performance gain in the condition in which the prosodic and the statistical segmentation outputs are compatible. However, as already mentioned according to the results of Experiment 1 and to previous findings (e.g., Mattys et al., 2005; Valliant & Levitt, 1996), lexical stress effects are particularly observed in degraded conditions. Thus, we would find an impact of primary stress only in degraded conditions, with its effect being maximized in the strongly degraded condition at study. Most likely, it will be in that situation only that listeners will parse statistical cohesive units obeying to the default stress pattern in their native language, namely TP-words stressed on their 2nd syllable. 3.6. Method 3.6.1. Participants One hundred and twenty-eight undergraduate psychology students at the University of Lisbon participated in the experiment for a course credit. All were Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 123 monolingual EP speakers, with no reported history of speech or hearing disorders. Among them, 43 were randomly assigned to the intact speech condition (12 to the TP–unstressed, 11 to the 1st syl-stressed, 7 to 2nd syl-stressed, and 13 to the 3rd sylstressed condition), 47 were assigned to the mildly noise (22dB SNR) condition (9 to TP–unstressed, 11 to 1st syl-stressed, 14 to 2nd syl-stressed, and 13 to 3rd syl-stressed condition), and 38 to the strongly noise (10db SNR) condition (9 to TP–unstressed, 10 to 1st syl-stressed , 10 to 2nd syl-stressed , and 9 to 3rd syl-stressed condition). 3.6.2. Material All speech stimuli were synthesized using the same method of the previous experiments of the present chapter. The same AL of Experiment 1 was also used. For the familiarization phase, four synthesized versions of the AL were created. Each version included the same sequence of syllables, divided into three 7minutes listening blocks (rendering 21-minutes) as in Experiment 2. TP-unstressed of the AL version corresponded to the single cue version of Experiment 2. In the three stressed versions of the AL, the acoustic correlate of lexical stress that was manipulated was the duration of the AL syllables. Stressed syllables were lengthened by 100 ms while unstressed ones were reduced each by 50 ms. However, the manipulation of acoustic correlate of lexical stress (i.e., duration of the AL-syllables) here adopted was similar to the one with coarticulatory cues presented in Experiment 1 of Chapter Two (cf. Fernandes et al., 2007). Only some (approximately one third of the exemplars) of the TP-words within the AL stream presented a stress pattern. These stressed TP-words were located as close as possible in the three stressed conditions of the AL, the only difference between them being Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 124 the location of the stressed syllable: for the 1st syl-stressed condition, the lengthened syllable was the first; for the 2nd syl-stressed condition, it was the second; and for the 3rd syl-stressed condition, it was the last syllable. For illustration, an orthographic translation of a sample of the speech stream in each version is presented in Table 5. The two degraded signal conditions with white noise superimposed at 22dB SNR (i.e., mildly degraded condition) and 10dB SNR (i.e., strongly degraded condition) were created as in Experiment 1 and for the forced-choice test the same material of Experiment 2 was used (see previous section 3.3.2.). Table 5: Orthographic translation of a sample of the stream heard in the familiarization phase in the four conditions of stress location of Experiment 3B. The “#” defines word boundaries according to TPs; the “-” represents concatenation; TPwords with stressed syllables are underlined and the stress syllable is presented in capitalized letters. STRESS LOCATION CONDITIONS (available cues) SPEECH STREAM PARTWORDS 3#12 PARTWORDS 23#1 TP – unstressed ...#lu-fa-ba-#ki-la-bu-#ka-fu-bi-#ba-bu-ku- #bu-ka-la-#fu-fi-bu-#… ba-#ki-la ka-la-#fu 1st syl Stressed ...#lu-fa-ba-#KI-la-bu-#ka-fu-bi-#ba-bu-ku- #bu-ka-la-#FU-fi-bu-#… ba-#KI-la ka-la-#FU 2nd syl Stressed ...#lu-fa-ba-#ki-LA-bu-#ka-fu-bi-#ba-bu-ku- #bu-KA-la-# fu-fi-bu-#… ba-#ki-LA KA-la-#fu 3rd syl Stressed ...#lu-fa-BA-#ki-la-bu-#ka-fu-bi-#ba-bu-ku- #bu-ka-LA#fu-fi-bu-#… BA-#ki-la ka-LA#fu 3.6.3. Procedure Identical to the one of Experiment 2 (see previous section 3.2.3.). Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 125 3.7. Results and Discussion In the mixed ANOVA ran on participants’ TP-word choices with part-word type (part-words 3#12, part-words 23#1) as within-subject factor, signal quality (intact, 22db SNR, 10dB SNR) and lexical stress pattern (TP–unstressed, 1st sylstressed, 2nd syl-stressed, 3rd syl-stressed) as between-subjects factors, neither the main effect of lexical stress pattern nor its interaction with the other factors at study were significant [Fs < 2.1, p >.10]. The only significant effect found was the main effect of signal quality [F(2, 116) = 5.9, p < .005, MSe = 8,72; Mp2 = .09], with ALL performance declining with signal degradation [F(1, 116) = 11.7, p < .001; MSe = 8,72, Mp2 =.09]. Thus, contrary to what was found in contrast to what was found in Experiment 2, when the prosodic information was not a universal cue but a languagespecific one, as was the case in the present experiment, no impact on listeners’ TPwords choices was found. This is in accordance with previous findings on the unreliability of prosody in speech segmentation when other cues are still available (e.g., Vallian & Levitt, 1996). We next checked whether the ALL performance of listeners was affected by TP-level of TP-words (see Figure 5). According to the mixed ANOVA with TP-words type (high-; low-TP-words) as within-subject factor, and signal quality, as well as lexical stress pattern as between-subject factors, the main effect of signal-quality was significant [F(1, 116) = 6.1, p < .005, MSe = 8,74; Mp2 = .09], as was the main effect of TP-words type [F(1, 116) = 6.3, p < .01, MSe = 5,34; Mp2 = .05]. Indeed, overall listeners had better performances for highthan for low-TP words. However, this TP-gradient effect was Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 126 modulated by signal-quality [F(2, 116) = 3.5, p <.05, MSe = 5,34; Mp2 = .06]. Importantly, the three way interaction was also significant [F(6, 116) = 3.0, p < .01, MSe = 5,34, Mp2 = .14]. No other significant effect was found [Fs < 1]. Figure 5: Performance pattern (proportion of AL-word responses, in percentage) broken-down by TP-level (HighTP; Low-TP) of AL-words, according to Signal condition (Intact; mildly degraded - 22dB SNR; strongly degraded - 10dB SNR) and Lexical Stress location (TP – No stress; 1st Syllable; 2nd Syllable; 3rd Syllable) in Experiment 3B. Vertical bars denote standard error of the mean on each condition. Chance level corresponds to 50%. Average ALL performance broken down by TP-level is presented in Table 6. In order to specifically analyze the TP-gradient effect on each stress pattern, we evaluated separately each input-intelligibility condition at study (see Figure 5). With intact signal no significant effect was found [all Fs < 1.5, ps > .10]. All groups performed at a similar level, with no differences for lowand high-TP-words, independently of the absence/presence of an acoustic marker of primary stress and of Intact Mildly degraded (22dB SNR) Strongly degraded (10dB SNR) TP-no stress High_TP Low_TP 30 35 40 45 50 55 60 65 70 75 80 TP-word choices (%) 1st syllable High_TP Low_TP 2nd syllable High_TP Low_TP 3rd syllable High_TP Low_TP Intact Mildly degraded (22dB SNR) Strongly degraded (10dB SNR) TP-no stress High_TP Low_TP 30 35 40 45 50 55 60 65 70 75 80 TP-word choices (%) 1st syllable High_TP Low_TP 2nd syllable High_TP Low_TP 3rd syllable High_TP Low_TP Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 127 its location in TP-words (see Table 6 and Figure 5). All groups presented an abovechance overall performance [TP–unstressed: t(11) = 3.5, p < .005; 1st syl-stressed: t(10) = 2.3, p < .05; 2nd syl-stressed: t(6) = 2.0, p = .05; 3rd syl-stressed: t(12) = 3.9, p < .005], choosing significantly more often the TP-words than the part-words as the “lexical units” of the new language. Thus, with intact signal, listeners were able to use statistical information to correctly parse the TP-words from the speech stream, independently of the stress pattern of these parsed stimuli. Table 6: Average proportions of TP-Word responses (in percentage) in Experiment 2B, separately for each inputintelligibility condition (intact; mildly degraded; strongly degraded) and lexical stress location (TP-unstressed; 1st syllable; 2nd syllable; 3rd syllable), considering the TP-level of the AL-words (highvs. low-TP). Standard error of the mean for each condition is presented in parentheses. INPUT-INTELLIGIBILITY Intact Mildly degraded Strongly degraded Lexical stress location High-TP Low-TP High-TP Low-TP High-TP Low-TP Average TPunstressed 67.6 (4.17) 58.4 (4.31) 60.5 (4.82) 59.2 (4.98) 56.8 (4.82) 61.2 (4.98) 60.4 1st syllable 61.1 (4.36) 57.1 (4.51) 59.6 (4.36) 43.9 (4.51) 45.5 (4.58) 51.7 (4.73) 53.1 2nd syllable 59.5 (5.47) 66.7 (5.65) 66.3 (3.87) 51.6 (3.99) 59.4 (4.58) 41.7 (4.73) 57.5 3rd syllable 60.7 (4.01) 63.2 (4.15) 63.7 (4.01) 55.1 (4.15) 54.9 (4.82) 52.5 (4.98) 58.3 Average 62.2 61.3 62.5 52.5 53.9 51.7 In the mildly degraded (i.e., 22dB SNR) condition, the main effect of TPlevel of AL-words was significant [F(1, 43) = 12.8, p < .001, MSe = 5,83, Mp2 = .23]. No main effect of stress pattern [F(3, 43) = 1.3, p > .25] was found. The interaction Chapter 3. Universal and Speech Chapter 3. Universal and SpeechChapter 3. Universal and Speech Chapter 3. Universal and Speech- -- -specific cues in ALL specific cues in ALLspecific cues in ALL specific cues in ALL 128 between the two factors did not reached the conventional level of significance [F(3, 43) = 1.9, p = .10, MSe = 5,42, Mp2 = .12] but the size of this effect was moderated (Cohen, 1988). In fact, in line with the previous findings of Experiment 1(see chapter one; cf. Fernandes et al., 2007), and as illustrated in Figure 5, when only statistical information was available in the stream, listeners chose both highand low-TPwords above chance as the “words” of the new language [t(8) = 2.2, p < .05; t(8) = 2.1, p < .05, respectively], with no differences in the extraction of highand low-TPwords [F < 1]. In sharp contrast, for listeners exposed to the AL with the presence of an acoustic marker of primary stress, the TP-gradient effect was significant [F(1, 43) = 17.5, p < .005], reflecting the fact that for those listeners, only high-TP-words were extracted above chance [on average on 63% of the trials (SE = 2.0); t(37) = 6.0, p < .0005]. Performance for low-TP-words did not differ from chance [on average, 52% (SE = 1.9); t(37) = .3, p > .10], and was significantly poorer than in the TP– unstressed condition [F(1, 46) = 4.0, p = .05]. This pattern of results suggests that although listeners performance is not globally affected by the presence of lexical stress patterns (no main effect of stress pattern was found) lexical stress plays some role in speech segmentation when the signal is mildly degraded, possibly constraining the extraction of statistical segmentation outputs to the ones with higher support. When lexical stress was available in the stream, independently of the stress pattern (see Table 6), only high-TP-words were correctly selected above chance as the plausible words of the AL. This is not surprising, since although by default in Portuguese stress is located at the penultimate syllable, the other two stress patterns are also legal in that language. In the strongly degraded condition (i.e., 10dB SNR), the modulator impact of Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 231 but not to those of list B, and vice-versa for the other participants presented with we call the ALB. This ensures that any ALL effect (in the forced-choice test) across the two ALs could only be due to statistical learning and allowed estimating the speed of processing of each base-word with (primed condition) and without (unprimed condition) the influence of a latent novel competitor (the AL word). Table 11: Examples of materials used (in phonological form, according to IPA) in the AL familiarization phase (ALFam.), two-forced choice (ALL-test), and lexicalization test (Lexical Decision – LD on base-words) in Experiment 8, separately for the two artificial languages (ALs: ALA and ALB). The “#” defines word boundaries according to the statistical segmentation cues (AL-TPs and phonotactics); the “-” represents concatenation. Shared phonological onsets between TP-words and basewords are underlined. AL-Fam. phase ALL-test Lexicalization Test (21-min exposition) (forced choice test) (LD on base-words) ALs TP-word Part-word List A List B ALA Primed Unprimed f5-yn-m5 mu#f5-zo f5ynzu f5vet5 …#f5-zo-m5#pi-≤`- du#≤u-le-nu#fi-vD-ku#... ≤u-le-nu du#≤u-le ≤ulet5 ≤5vin5 ALB Unprimed Primed f5-ve-lu ku#f5-ve f5ynzu f5vet5 …#f5-ve-lu#≤5-viku#l5-t`-vu#fi-ni-m`#... ≤5-vi-ku lu#≤5-vi ≤ulet5 ≤5vin5 Since a pretest (see 6.4.2. section) had shown that the two lists of base-words led to similar lexical decision latencies and performances in participants that were not presented with any AL beforehand, in the present case any accuracy and/or latency difference between the two lists of base-words observed within each group of participants after familiarization to an AL would necessarily reflect a change in their processing due to the lexicalization of the outputs of the statistical segmentation Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 232 process (i.e., of the TP-words) that overlapped with the base-words. This inhibitory effect would be the “lexical footprint” of these novel items (Gaskell & Dumay, 2003a). Furthermore, listeners also performed the ALL test and the lexicalization test one week later, with no intervening new familiarization period. Previous work has shown that while a low rate of familiarization to novel items presented in isolation (e.g., 12 times each, Gaskell & Dumay, 2003a) leads to a slow lexicalization process, massive exposure to such novel items endorse immediate as well as long-lasting lexicalization effects (Gaskell & Dumay, 2003b). Whether the same holds true for massive exposure to items presented in a continuous stream was examined here, since each TP-word was present 189 times in the AL stream. 6.4. Method 6.4.1. Participants Thirty-two undergraduate psychology students at the University of Lisbon participated in the experiment for a course credit. None of them participated in the other experiments. Sixteen were familiarized with ALA and the other 16 with the ALB. 6.4.2. Material Participants were exposed to different materials in the three phases that constituted the first experimental session of this Experiment. All the AL-material was created based on real EP words (i.e., the Base-Words). Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 233 Base-Words Two lists of 10 trisyllabic (CV.CV.CV) real European Portuguese words constituted the base-words selected in order to create the two ALs and to evaluate the possible lexicalization of the TP-words. These base-words were selected from the Porlex database (Gomes & Castro, 2003). All of them were six phonemes long, with UP occurring at or earlier than the fourth phonemic position. Word frequency was based on Corlex (available at: http://www.clu.ul.pt) and log transformed for normalization of the distribution. The selected base-words were paroxytones and their second and third syllables onset were also possible word onsets in European Portuguese. Thus, neither lexical stress nor phonotactic violations could be used for segmentation. Table 12: Base-Words characteristics by block (Blocks A and B; Experiment 8 and 9: mean log frequency; neighborhood density, mean transitional probability (TP between first and second syllable), Uniqueness Point (UP), and mean duration (in ms), as well as t-values (all ps > .5). Standard error of the means is presented in parenthesis. Base-Words Block A Block B t-value Log Frequency 3.51 (0.38) 3.79 (0.36) -0.54 Neighborhood density 1.5 (0.17) 1.7 (0.26) -0.65 TP (1st and 2nd syl.) 5.0 (2.15) 5.0 (2.17) 0.00 UP 3.7 (0.15) 3.7 (0.15) 0.00 Duration (ms) 885.20 (16.10) 885.40 (13.04) -0.01 As illustrated in Table 12, the two lists of base-words (A and B; see Appendix V) were matched on word onset (and whenever possible on the first syllable), mean log frequency, neighborhood density (i.e., number of words that differed from them by 1-phoneme substitution, addition, or deletion, while Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 234 preserving relative position of the segments; e.g., Vitevitch & Luce; 1998; 1999; based on Porlex database, Gomes & Castro, 2003), mean TP between the first and the second syllable in European Portuguese, and mean duration. Lexical Decision Pretest To ensure that the base-words of the two lists were a priori equivalent, we checked whether they would lead to similar lexical decision performances when no prior familiarization phase to any AL is used. To this aim, we tested an independent group of 17 participants in lexical decision. No difference was found between the two lists on either accuracy [t(16) < 1, p > .10; with mean accuracy of 80 and 82% for lists A and B, respectively] or response latency [t(16) = 1, p > .10; average RT of 1166 and 1144 ms for lists A and B, respectively]. Thus, any accuracy and/or latency difference between the two lists of base-words observed after familiarization to the AL would necessarily reflect a change in their processing due to the lexicalization of the outputs of the statistical segmentation process (i.e., of the TP-words) that overlapped with them. AL Material Each one of the two ALs was based on one of the two lists of base-words, with ALA based on list A and ALB on list B. Each AL was constituted by 10 trisyllabic TP-words. These diverged from the base-words by their last CV syllable, after the fourth phoneme and thus after the UP. Each group of listeners was familiarized with one of the two ALs, in a between-subjects (ALA; ALB) counterbalanced design. For participants presented Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 235 with ALA, each TP-word (e.g., /fiudjt/; /f5ynl5.) was phonologically related to one base-word of list A (primed base-words, e.g., “fivela” /ehuDk5/, meaning “buckle”, and “gasoso” /f5ynyt/ meaning “gaseous”, respectively), but no TP-word was related to any base-word of the other list (unprimed base-words; e.g., “finito” /finitu/, meaning “finite”, and “gaveta” /f5uds5.ld`mhmfcq`vdq). Conversely, for those presented with ALB, each TP-word was phonologically related to one baseword of list B, but none was related to list A. Hence, across the two AL familiarization groups, each list of base-words occurred in each (primed vs. unprimed) condition. The third syllable of each TP-word was created guarantying that: (i) the resulting stimuli was not a real word in EP; (ii) the phonological sequence was phonotactically legal in EP words; and (iii) it occurred in EP words more frequently in final than in initial position [based on Porlex, Gomes & Castro, 2003; F (1,18) = 29.5, p < .001]. This last criterion was chosen due to material constraints. For having available two balanced types of segmental information (i.e., from AL and from listeners’ native language), we add this positional syllabic information, since in the AL, TP-words had not only higher TPs than part-words but also occurred three times more often. Thus, in the AL stream, two types of segmental information were available: the statistical information of the AL (i.e., TPs between adjacent syllables and raw frequency of the trisyllabic AL stimuli), and (according to the strong overlap from onset of TP-words and Base-Words and the relative position frequency of the last syllable of each TP-word in EP). For each AL, 10 trisyllabic part-words were selected, each consisting of the last syllable of one TP-word and the first two syllables of another (that occurred Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 236 adjacently in the stream during the familiarization phase). All AL stimuli (TP-words and part-words of the two ALs) are presented in Appendix VI. Porlex Data In order to guarantee that the two ALs were matched on statistical segmental information in listeners’ native language, we computed the average TP between the syllables that constituted each AL stimulus (TP-words and part-words) in European Portuguese words (based on Porlex database; Gomes & Castro, 2003), separately by stimulus type (TP-words and part-words) and AL (ALA and ALB). According to the mixed 2x2 ANOVA, neither the main effect of AL [F(1, 36) = 1.9, p >.10] nor the interaction between the two factors were significant [F < 1]. Importantly, the main effect of stimulus type was reliable [F(1, 36) = 11.0, p < .005], since TP-words had higher TPs in European Portuguese than part-words. Naming & Delayed Naming Pretest This was also confirmed by naming and delayed naming pretests on the AL stimuli (TP-words and part-words of the two ALs) ran on an independent group of 19 participants. They were presented, on each trial, with an auditory AL stimulus, which they should first name immediately and next name after an acoustical signal (occurring either at 1600, 1800, 2000, 2200, and 2400 ms after the offset of the stimulus). For the immediate naming condition, only the main effect of AL stimulus type (TP-words vs. part-words) was significant [F(1, 18) = 45.9, p < .0001]: TPwords were named faster than part-words, [ALs: F < 1; AL condition x AL stimulus type: F(1, 18) = 3.6, p = .10]. In delayed naming, no significant effects were found Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 237 [i.e., ALs: F(1, 18) = 1.8, p > .10; AL stimulus type: F(1, 18) = 1.3., p > .10); AL condition x AL stimulus type: F < 1]. AL stream The ALs (ALA; ALB) were synthesized using the same method as in Experiment 1. Each AL was constituted by 10 TP-words. In the AL stream, TP between adjacent syllables was always higher within (TP of 1) than between (TP of 0.33) TP-words and TP-words occurred three times more often than part-words. For the AL-familiarization phase, for each AL, three 7-minutes blocks (rendering 21-minutes) of the synthesized stream were created by concatenating, on each block, 63 tokens of each of the ten TP-words (1890 syllables, 630 tokens of TPwords per block), with the criterion that two tokens of the same TP-word never occurred adjacently in the stream. Only segmental informations (i.e., TPs between adjacent AL syllables and wordlikeness of AL-stimuli) were available in the stream, with no other acoustic or segmentation cues to word boundaries. For the two ALL-tests (i.e., two-alternative forced choice test), as regards each AL, the three syllables of each AL-stimulus (TP-words and part-words) were synthesized and concatenated. Lexical Decision Material The lexical decision task was constituted by 120 trials. Twenty were critical trials corresponding to the (10 x 2 lists of) base-words and the remaining 100 trials were trisyllabic fillers. There were 40 word fillers, with varied phonological structures and lexical stress in each possible position. The 60 nonwords used as Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 238 fillers were created by changing one phoneme (or syllable) of European Portuguese real words of varying phonological structures. This ensured a low proportion of basewords related with the AL to which participants were familiarized (i.e., 16.67%) and a proportion of 50% of the trials requiring a “word” response. In the lexical decision task, no AL stimulus was presented. A practice list of 10 items with the same characteristics of the experimental list was also generated for participants’ familiarization with the task. The stimuli used in lexical decision task were recorded in a soundproof booth (using M-Audio Fire Wire 410 and Adobe Audition 1.5 Program on a Carillon Audio Systems, Pentium IV PC) by an EP female native speaker using M-Audio Nova-Class A FET microphone and sampled at a rate of 22.05 kHz and 16-bit conversion. Editing of the digitized versions of the stimuli was made with Adobe Audition 1.5 and Praat 4.3.04 (available at http://fon.hum.uva.nl/praat/). 6.4.3. Procedure Participants were tested individually or in groups of two in a sound-attenuated room. All auditory stimuli were presented through headphones and the ALfamiliarization phase was similar to the one of Experiment 7. For the ALL test and lexical decision task, presentation, timing and data collection were controlled by EPrime 1.1 (Schneider et al., 2002a; 2002b). All participants were informed that there were two testing sessions, with the second session one week after the first. No other information about the second session was given. They were told that they were participating in an investigation Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 239 about word learning and that the first session was constituted by three phases (see Table 12). The AL familiarization phase was procedurally identical to the one of Experiment 7. After the AL familiarization phase, participants were presented with the ALL-test (two-alternative forced-choice test), according to the AL to which they were previously exposed. The structure of the two ALL-tests (for ALA and ALB) was identical, and formally similar to the one described in Experiment 7. Each trial started with a warning tone, followed by two trisyllabic strings, separated by 500 ms of silence, presented through headphones. One of these strings was a TP-word and the other a part-word. For avoiding needless repetitions of AL-stimuli, each TP-word was only paired with each one of the two part-words with which it shares syllables (20 trials) and with two other part-words with no relation with it (other 20 trials), rendering 40 trials. Accuracy was emphasized, although participants were told to always provide an answer. The test began with the four practice trials used in Experiment 7. Order of trials was randomized for each participant, and order of presentation of stimuli within trials was counterbalanced within each group and experimental session. All participants performed the same timed lexical decision task. No information about any possible relation between the first two phases and this one was given. Order of trials’ presentation was pseudo-randomized for each participant in each experimental session, obeying to three criteria: the first ten trials were always filler (warm-up) trials; critical trials (Base-Words) did not occur in consecutive order; and no more than three consecutive trials required the same (yes/no) response. Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 240 Participants were asked to perform a lexical decision, as quickly and accurately as possible, on each auditory stimulus presented, using their index-fingers to press one of two buttons of the (PST SRB 200A) button box: “yes” responses were given with the right index-finger. Response latencies were measured from stimulus onset, with a response deadline of 3000 ms and 1500 ms inter-trial interval. Before the experimental list, participants were familiarized with the task in a practice block, with feedback on correctness and speed provided only for practice trials. After the lexical decision task, participants were reminded about the second session and were dismissed. Overall, the first session lasted about 50 minutes. In the second experimental session, participants were informed that this was related to the previous first session with the only difference that they would not hear the AL again, but would perform the other two tasks (ALL-test and the lexicalization test). The second session lasted about 20 minutes. 6.5. Results and Discussion 6.5.1. ALL performance: two alternative forced-choice test Proportions of TP-word choices (as an index of ALL), computed for each participant, is presented in Figure 9, separately for the two moments of test. In order to evaluate if there was any difference between the two ALs (ALA and ALB) and/or between the two moments of testing (immediate; after 1 week) as regards the performance level reached in the two alternative forced-choice test, we ran a mixed 2x2 ANOVA on raw TP-choices, with AL as between-subjects factor Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 247 segmentation cues or uniquely by the wordlikeness of TP-words. That knowledge of phonotactic probabilities was used by the participants of the present experiment in ALL is suggested from their high performance level in the forced-choice test, which is in accordance both with former ALL studies manipulating this factor (Onnis et al., 2005; Perruchet et al., 2004) and with studies that have showed a redundancy gain with congruent segmentation cues (e.g., Fernandes et al., 2007; Vroomen et al., 1998). However, none of these studies evaluated lexical engagement. To what extend wordlikeness per se and/or cue integration may have favored the lexicalization of the new item configurations was evaluated in Experiment 9. Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 248 Experiment 9. The status of the output of the segmentation procedure adopted by listeners: When incongruent cues are available in the AL-stream. In the present experiment, the two types of segmental information studied in Experiment 8 (namely, TPs and wordlikeness of the AL stimuli) were also available in the AL stream, but were incongruent since they suggested incompatible segmentation outputs. As illustrated in Table 14, whereas the TPs between syllables suggested that TP-words were the “lexical items” of the new language, this time it was the part-words that were the AL stimuli that overlapped from onset with real words of the listeners’ native language. Table 14: Examples of materials used in AL-familiarization phase (AL-Fam.), ALL-test, and Lexicalization test (Lexical Decision – LD on Base-Words) on Experiment 9, regarding the two Artificial Languages (ALs). The “#” defines word boundaries according to the AL-TPs; the “-” represents concatenation. Shared phonological onsets between Part-words and Base-Words are underlined. All material is presented in phonological form, according to IPA. AL-Fam. ALL-test Lexicalization Test (21-mins exposition) (forced choice test) (LD on Base-Words) ALs TP-word Part-word Block A Block B ALA2 Primed Unprimed mu-f5-zo f5-yn#m5 f5ynzu f5vet5 …#mu-f5-zo# m5-Ri-R`#du-≤ule#d5-fi-vD#... du-≤u-le ≤u-le#nu ≤ulet5 ≤5vin5 ALB2 Unprimed Primed ku-f5-ve f5-ve#lu f5ynzu f5vet5 …#ku-f5-ve#lu-≤5-vi#d5-l5t`#du-fi-ni#... lu-≤5-vi ≤5-vi#ku ≤ulet5 ≤5vin5 On the basis of previous results obtained with incongruent cues in ALL settings, we hypothesized that listeners’ native language information would be Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 249 considered as more reliable than AL statistical information, at least with an intact speech signal, as was the case here. Indeed, using an AL in which both TPs and another acoustic-phonetic cue, namely coarticulation, were available, Fernandes et al. (2007) showed that, with intact speech signal, coarticulation was able to drive the segmentation process, even when TPs suggested other, incongruent segmentation hypotheses. The same pattern had been previously observed, with intact signal, in 8month-old infants (Johnson & Jusczyk, 2001). In the present experiment with incongruent cues, if listeners preferred native language information over AL statistical information, the output of their segmentation process would correspond to the stimuli that, although presenting lower TPs in the AL, strongly overlap from onset with real words of listeners’ native language. In other words, in this case, listeners would prefer the part-words as the probable “words” of the new language, discarding the TP-words, a preference that would be reflected by a significant below-chance performance in the two-forced choice test. The lexicalization test (also used in Experiment 8) allowed evaluating whether the output of the segmentation procedure adopted by listeners (according to their performance on the ALL test) would still be integrated in their lexical environment. 6.6. Method 6.6.1. Participants Twenty-four undergraduate psychology students (12 familiarized with new ALA2 and the other 12 with the new ALB2) at the University of Lisbon participated in Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 250 the experiment for a course credit. None of them participated in the previous experiments. 6.6.2. Material The base-words and AL stimuli were those used in Experiment 8, with the only difference that the TP-words of the present experiment corresponded phonologically to the part-words of Experiment 8, and vice-versa (see Table 14 for examples of the material used). Each one of the two ALs (ALA2 and ALB2) was constituted by ten TP-words (based on the AL-statistical information were the “lexical-units” of the new language). The ten part-words were trisyllabic stimuli constituted by the last two syllables of one TP-word and the first syllable of another, which strongly overlapped from onset until the fourth phoneme with real EP words (i.e., the Base-Words). Thus, part-words had a higher wordlikeness than TP-words (see Appendix VI). AL stream The two new ALs (ALA2 and ALB2) were synthesized using the same method of the previous experiment. Each AL was constituted by 10 TP-words (the Partwords of the previous experiment). The TP between adjacent AL syllables was always higher within (TP of 1.00) than between (TP of 0.33) TP-words, as well as their raw frequency (TP-words occurred three times more often than part-words: each TP-word occurred 189 times, while each part-word occurred 63 times). Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 251 For the AL-familiarization phase, for each new AL, three 7-minutes blocks of the synthesized stream were created as in Experiment 8. The only segmentation cues available were the incongruent segmental informations, with no other acoustic or segmentation cues to word boundaries. The ALL-tests were constructed as in Experiment 8, and the lexicalization test (i.e., lexical decision task) was the one used in the previous experiment. 6.6.2. Procedure Procedure and structure of the two experimental sessions (immediately and post-one week) were identical to the previous experiment. 6.7. Results and Discussion 6.7.1. ALL Performance: two alternative forced-choice test Average proportions of TP-word choices (as an index of ALL), computed for each participant, is presented in Figure 12, separately for the two moments of test. As in Experiment 8, we first evaluated if there was any difference between the two ALs (ALA2 vs. ALB2) and/or between the two moments of testing (immediate; after 1 week) as regards the performance level reached in the two alternative forced-choice test. This was not the case, since the mixed 2x2 ANOVA ran on raw TP-word choices, (with AL as between-subject factor and moment of testing as within subject-factor) showed no significant effect or interaction [all Fs ≤ 1]. Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 252 Figure 12: Mean percentage of TP-word choices in the ALL-test for each participant, according to the moment of testing: immediately after the AL-familiarization phase (i.e., immediate) and one week after with no intervening ALfamiliarization period, in Experiment 9. Chance level corresponds to 50%. Performance was below the chance level (i.e., below 50%), reaching 41.1% (SE = 3.6) immediately after familiarization and 43.5% (SE = 3.6) one week later [t(23) = -2.5, p < .01; and = -1.8, p < .05, respectively], with 13 out of the 24 participants presenting a performance clearly below chance level (≤ 45% of TPwords choices) in the two tests. Thus, listeners discarded the TP-words and considered more often the part-words as the “lexical units” of the new language. This is in line with former results suggesting that at least with intact speech (cf. Fernandes et al., 2007), listeners at different stages of their linguistic development consider native speech segmentation cues as more reliable than the statistical segmental information conveyed by an AL stream. Immediate 1 week after Moment of Testing 0 10 20 30 40 50 60 70 80 90 100 % of TP-word Choices Immediate 1 week after Moment of Testing 0 10 20 30 40 50 60 70 80 90 100 % of TP-word Choices Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 253 6.7.2. Lexicalization test: inhibitory priming in lexical decision As in Experiment 8, RTs were measured from stimulus onset to response onset, errors were analyzed separately and the same trimming procedure in RTs was applied (exclusion of less than 3.5% of the data). Mean RTs for the two lists of base-words (lists A and B) are presented separately for each AL and moment of testing in Table 15. Table 15: Mean RTs (measured from target onset) for Base-Words A and B (BWd A; BWd B) and the Priming Effect (P.E., the difference between primed and unprimed conditions) separately for each AL-condition (ALA2 and ALB2) and moment of testing (Immediately and Post-1 week) in Experiment 9. Standard errors of the mean in each condition are presented in parenthesis. Moment of Testing Immediately Post 1 Week AL BWd A BWd B BWd A BWd B ALA2 Primed Unprimed Primed Unprimed 943.4 (26.15) 916.5 (23.92) 26.85 895.9 (20.56) 852.8 (21.63) 43.12 ALB2 Unprimed Primed Unprimed Primed 1013.2 (26.15) 1020.2 (23.92) 6.95 933.73 (20.57) 933.7 (21.63) 20.80 Immediately after familiarization, the output of the segmentation procedure adopted by listeners was not lexicalized. Indeed, the mixed 2x2 ANOVA with AL as between-subject factor (ALA2; ALB2) and base-words list as within-subject factor (list A; list B) showed that neither the interaction between the two factors at study [F= 1; Mp2 = .04] nor the main effect of base-words list [F < 1; Mp2 = .01] were significant. We thus found no difference between the processing of primed and unprimed base- Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 254 words, namely between those that overlapped with the segmentation outputs (here, the part-words) and those that did not. Surprisingly, the main effect of AL was significant [F(1, 22) = 7.7., p = .01, MSe = 11687; Mp2 = .26], with listeners exposed to ALA2 responding faster than listeners exposed to ALB2. This result was unexpected, since the two ALs were well matched in all relevant parameters (see section 6.4.2. of the present chapter). Furthermore, in the previous experiment the same phonological repertoire was used in ALs and no global differences were found there. It seems to be due to the manipulation of the AL factor as a between-subjects variable. Much more important is the fact that there was no hint of an effect showing that the output of phonotactic segmentation had acquired some lexical status. Coherently with the non-significant effect of base-words list and its absence of interaction with AL, for both ALs the inhibitory effect (measured as the lexical decision latency difference between primed and unprimed basewords) did not differ from zero [ts(11) < 1, ps ≥ .10]. In the errors analysis, no significant effect was found [Fs < 2.5, ps > .10], showing that there was no speed/accuracy trade-off. For listeners familiarized with ALA2, errors reached on average 11% (SE = 3.0) for base-words A and 14% (SE = 4.0) for base-words B; for listeners exposed to ALB2, they reached 11% (SE = 3.0) and 12% (SE = 4.0), respectively. The same pattern found one week later. Listeners familiarized with ALA2 were faster than the ones familiarized with ALB2 [F(1, 22) = 5.9, p = .024, MSe = 2598, Mp2 = .21], but neither the main effect of base-word list [F(1, 22) = 1.5, p > .10, MSe = 2598; Mp2 = .06] nor the interaction between the two factors [F(1, 22) = 2.9, p Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 255 =.10, MSe = 2598; Mp2 = .11] were significant. The analysis on errors did not reveal any significant effect [AL condition: F(1, 22) = 1.32, p > .10; base-words list: F(1, 22) = 3.3, p = .10] or interaction between the factors [F < 1]. For listeners familiarized with ALA2, errors reached on average at 9% (SE = 3.0) for base-words A and 12% (SE = 4.0) for base-words B; and for listeners exposed to ALB2, they reached 7% (SE = 3.0) and 9% (SE = 4.0), respectively. Thus, although one may expected the incubation period, namely the delay in lexical engagement of speech segmentation outputs (cf. Dumay & Gaskell; 2003a; 2003b; 2007) to be particularly crucial in conditions of incongruent segmentation cues, one week after familiarization speech segmentation outputs were still not integrated into the mental lexicon as potential word-like units. Sub-analysis In order to guarantee that these null effects also hold true for participants whose segmentation outputs consistently corresponded to the part-words, we ran two further ANOVAs (according to moment of testing) with the same factors as in the previous analyses but considering only the participants that had displayed a clearly below-chance performance (≤ 45%) in the forced-choice test. Remarkably, both immediately after familiarization and post-one week, neither the base-word list main effect [immediate: F < 1; post-one week: F(1, 11) = 2.9, p > .10] nor the interaction of that factor with AL [immediate: F < 1; post-one week: F(1, 11) = 1.8, p > .10] were significant. The main effect of AL was also nonsignificant [F ≤ 1 for both immediate and post-one week tests). Thus, even for listeners that systematically Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 256 adopted a segmentation procedure based on the wordlikeness of part-words, no hint of a lexicalization effect of their segmentation outputs was found. The current results suggest that when different types of segmental information indicate incongruent word boundaries in the AL, listeners are still able to segment the stream, with knowledge of their native language being preponderant over AL statistical information, which is in line with previous findings (e.g., subsegmental information: Fernandes et al., 2007; Johnson & Jusczyk, 2001; prosodic contours: Shukla et al., 2007; Experiment 3 and 6) and reveals the differential weighting attributed to different types of sublexical cues in speech segmentation (Mattys et al., 2005). However, the AL units segmented on the basis of only wordlikeness and thus here against the statistical information conveyed by the AL streamdid not seem to have acquired any sort of lexical status, either immediately or one week after familiarization to the AL. Indeed, the speed of processing for existing lexical items that strongly overlapped with part-words was no different from the one found for other, unrelated, lexical items. 6.8. General Discussion The ALL paradigm has become a important tool in the study of language acquisition (for a review see Gómez, 2007) as well as in the context of a full mature speech perception system (e.g., Fernandes et al., 2007; Onnis et al., 2005; Saffran et al., 1996b; Vroomen et al., 1998). Indeed, it allows systematic manipulation of different segmentation cues, reducing the set of uncontrolled variables, and hence provides a powerful way to test various theoretical proposals (e.g., Fernandes et al., Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 263 inhibitory priming effect linked to onset overlap is highly debated, at least in immediate priming studies. For example, Goldinger (1999) argued that the inhibitory priming effect does not provide an accurate picture of lexical competition, because it co-occurs with evidence for response biases. More recently, Pitt and Shoaf (2002) argued that participants’ surprise when they encounter the first related trials might explain the inhibitory priming effect, which they showed to dramatically decrease between the beginning and the end of the experiment. Other authors have argued that long-term priming effects may also be contaminated by strategic processing. For example, Luce et al. (2000) consider as strategically-based the long-term priming effects observed by Monsell and Hirsh (1998) with lags of 1 to 5 min between primes and targets overlapping at their onset, among others because they used a high proportion of overlapping items in relatively small stimulus sets (Goldinger, 1998a; Goldinger et al., 1992), leading participants to notice the occurrence of overlapping items, as acknowledged by the authors (Monsell & Hirsh, 1998, p.1511). However, numerous studies support the notion that the inhibitory priming effect linked to onset overlap is largely lexical in nature and does not reflect purely strategic processes. Indeed, contrary to facilitatory priming, inhibition due to multiple-phoneme onset overlap is, in fact, stronger in conditions in which strategic confound is minimized, for example when the proportion of related prime-target pairs is low (Hamburger & Slowiaczek, 1996). Even surprise cannot fully explain the inhibitory priming effect, since such an effect is observed even when participants were presented with related pairs in the training session (Dufour, Frauenfelder, & Peereman, 2007). In addition, it is worth noting that the study of Pitt and Shoaf (2002) showed that the size of the inhibitory priming effect decreases as strategic Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 264 processes build-up over the course of the experiment (see also Hamburger & Slowiaczek, 1996). Thus, at least in immediate priming experiments, response biases are more likely to mask than to cause the inhibitory priming effect (Dufour et al., 2007). Second, the size of the inhibitory priming effect is influenced by lexical factors such as the lexicality of the primes (Slowiaczek & Hamburger, 1992), the relative frequency of the primes and targets (Radeau, Morais, & Segui, 1995), and the neighborhood density of the target words (e.g., Dufour & Peereman, 2003; Dufour et al., 2007; Luce & Large, 2001; Vitevitch & Luce, 1998; 1999; 2005). Third, the inhibitory effect observed in long-term priming studies is limited to word primes, supporting a lexical origin of the effect (Monsell & Hirsh, 1998; Sumner & Samuel, 2007). In Experiments 8 and 9, we tried to avoid response biases by never presenting the base-words to listeners before the lexical decision task, never giving them any information about their possible relation with the AL stimuli, and using a low proportion of trials (less than 17%) in which real words had a phonological relation with AL stimuli, as well as a long lag between presentation of AL isolated stimuli (in the forced-choice task) and presentation of these base-words (for lexical decision). However, one potential problem lies in the fact that we evaluated lexicalization only through lexical decision, a task that, although widely used, involves a decisional component and thus is also often considered as involving substantial post-lexical processing (e.g., Abrams & Balota, 1991; Goldinger, 1996a; Slowiaczek & Hamburger, 1992) and hence to be highly sensitive to strategic biases in priming studies (e.g., Holender, 1992; Norris, McQueen, & Cutler, 2002; Radeau, Morais, & Dewier, 1989; Slowiaczek, Soltano, Wieting, & Bishop, 2003). There is however Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 265 some evidence that lexicalization effects are not dependent on the specific lexicalization test chosen. As a matter of fact, Gaskell and Dumay (2003a) reported consistent lexicalization effects in both lexical decision and pause detection tasks. In addition, the latter task seems to be less prone to strategic biases because it offers a measure of the overall level of lexical activity in the absence of any explicit linguistic judgment, since listeners are merely asked to detect periods of silence in speech (Mattys & Clark, 2002). While future work should aim at checking whether the same holds true as regards lexicalization in ALL, it is worth noting that task idiosyncrasy cannot explain why an inhibitory priming effect was observed only in the context of the congruent cues used in Experiment 8, and not with the incongruent cues used in Experiment 9. Two factors can actually explain the inconsistent pattern of results of Experiments 8 and 9 as regards the lexicalization effect: the frequency of occurrence of the AL stimuli, and the (in)congruence of the available segmentation cues. Possibly, novel phonological sequences need to be sufficiently familiar before they are lexicalized. Indeed, it would be counter-productive if all phonological forms to which listeners are exposed would be stored as different categories in mental lexicon. In addition, it has been shown that novel item’s frequency has a role in lexicalization (Gaskell & Dumay, 2003b). In Experiments 8 and 9, due to material constraints, part-words occurred less often than TP-words during familiarization, as in most ALL studies, e.g., Saffran et al., 1996a). For sure, raw frequency of the AL stimuli is not able to explain all ALL patterns, since ALL has been reported even when the raw frequency of part-words and TP-words had been equated. (e.g., Aslin, Saffran & Newport, 1998; Saffran et al., 1996a). But this factor is well known to Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 266 affect language processing (e.g., Brent & Cartwright, 1996; Magnuson et al., 2003) and may account for at least some ALL patterns (e.g., Perruchet & Vinter, 1998). Thus, the absence of a significant lexicalization effect in Experiment 9 might be related to the lower frequency of part-words in comparison to TP-words. Nevertheless, previous data suggest that the raw frequency of part-words was sufficiently high in Experiment 9 to expect lexicalization to occur. Indeed, although using a different paradigm and presenting participants with already segmented stimuli, Gaskell and Dumay (2003a) reported a lexicalization effect in lexical decision with about 30 phoneme monitoring trials, and Gaskell and Dumay (2003b) even found an inhibitory effect post-one week for “low-frequency” novel items that had occurred only 12 times each during phoneme monitoring. In Experiment 9, partwords occurred in the AL stream 63 times each (i.e., twice more than in Gaskell and Dumay, 2003b, and five times more than Gaskell & Dumay’ 2003b low-frequency novel items), which was enough for listeners to consider them as the “words” of the AL (in the forced-choice test) but not to lexicalize them. It is of course possible that exposure to a continuous stream acts as some kind of noise to the lexicalization process, hence impeding massive exposure to AL stimuli to endorse an immediate lexicalization effect. However, the fact that we observed no lexicalization gain one week after familiarization suggests that frequency of occurrence is not the only factor at play. The pattern of results found in Experiments 8 and 9 might alternatively reflect the fact that the consistency of the segmentation hypotheses afforded by different sources of information is important not only for speech segmentation (e.g., Fernandes et al., 2007) but also for word learning, modulating the potential Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 267 lexicalization of the segmentation outputs. Thus, only phonological representations parsed from the continuous speech stream on the basis of the strongest evidence, i.e., with different segmentation cues suggesting the same parsing, would be lexicalized. Such a view would be in line with the suggestion that the integration of congruent segmentation cues provides evidence about aspects of linguistic structure that are not available from any single source of information, reducing the potential for making false generalizations (Christiansen, Allen, & Seidenberg, 1998; Christiansen & Curtin, 2005). More generally, further work should also aim at examining how close the present results may be related to the concept of ecological validity. In natural settings, listeners are exposed to a rich combination of different signal-derived and lexically driven segmentation cues. But most of the time these cues converge rather than diverge. Thus, in less ecological situations, e.g. when cues collide, as was the case in Experiment 9, listeners might be reluctant to accept the by-product of segmentation as “true” words. Whatever the outcome of this future research, our results clearly show that the lexicalization effect found in Experiment 8 is not only a matter of lexical similarity. Future work should also be aimed at examining which computational models may most adequately account for such word learning effects occurring in the context of speech segmentation and, in particular, for such rapid acquisition and integration of statistical information. While any discussion of this point remains highly speculative in the absence of precise simulations, it seems to us that the Adaptive Resonance Theory or ART (Carpenter & Grossberg, 2003; Grossberg, Boardman, & Cohen, 1997; Grossberg & Myers, 2000; see also Sumner & Samuel, 2007, and Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 268 Vitevitch & Luce, 1999) might be a good candidate to apprehend word recognition and word learning within a continuous flow, as well as to account for the present findings. In ART, the resonance created by the match between bottom-up information and top-down expectations (stored in Long Term Memory – LTM) is responsible for the creation of the conscious spoken percept (Grossberg et al., 1997; Grossberg, 2003) and also drives word learning (Carpenter & Grossberg, 2003). While the spoken input unfolds, it activates speech segments that in turn activate “items” represented at Working Memory (WM), where they are also unitized, into “list chunks” of variable length (e.g., phonemes, syllables or words) within the same masking (inhibitory) field. When novel words are presented in the speech stream, as happens during familiarization to an AL, novel information cannot form a good enough match with the expectations that are read-out by previously learned recognition categories. This will trigger a memory search, or hypothesis testing, that leads to selection and learning of a new recognition category that can better match the input. Since in the masking field, longer list chunks will promote stronger inhibition, a chunk that codes a longer novel list like /suset5/ (one of our TP-words in Experiment 8) will have apriori advantage over chunks that code shorter lists like the familiar nonsense syllables that constitute it. It is precisely the masking advantage of longer list chunks that enables novel words to successfully compete with “amplifier” familiar chunks of shorter lists. When top-down expectation achieves a good enough match with bottom-up data, this match process focus attention upon those feature clusters in the bottom-up input that are expected. If the expectation is close enough to the input Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization during ALL line measures and Lexicalization during ALLline measures and Lexicalization during ALL line measures and Lexicalization during ALL 269 pattern then a state of resonance develops as the attentional focus takes hold, and the novel word is learned. The resonant process will be interrupted or terminated by mismatch reset and hence a new speech segmentation point will be defined. The repeated exposure to specific spatial patterns (in the present case, the repeated occurrence of TP-words in the speech stream) permits learning by the LTM traces in the adaptive pathways between item nodes and the list nodes, reflecting the ALL and lexicalization effects found in Experiment 8. ART is also able to explain the pattern found in Experiment 9. When different informations are available in the stream, suggesting different competing chunks of the same length (e.g., TP-words vs. more real word-like part-words), resonance is actively reset by input mismatching even before reaching a stable state. The simultaneously activated competing list chunks promote only partial matching resonant loops and hence the new information cannot be stably stored in LTM. Consider for example a masking field that is tuned to expect two chunks like /kufune/ (a TP-words of the AL that based on bottom-up, on-line statistical information is an unifying category) and /funetu/ (a part-word of the same AL, but that according to its high wordlikeness and hence phonotactic information could easily constitute an unifying category). These chunks will strongly inhibit each other based on the shared items /fune/, and although both promote resonant loops, an equilibrated resonant state will not be achieved, and hence novel words will not be stored in LTM. Chapter 6. On Chapter 6. OnChapter 6. On Chapter 6. On- -- -line measures and Lexicalization line measures and Lexicalizationline measures and Lexicalization line measures and Lexicalization during ALL during ALL during ALL during ALL 270 PART III GENERAL DISCUSSION & CONCLUSIONS Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 272 Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 279 ensured that syllable duration is, in EP, a correlate of word primary stress (see also Delgado-Martins, 2002). In Experiment 3B the weighting of a statistical segmental cue (i.e., TPs between adjacent syllables) and word primary stress (i.e., language-specific prosodic information) was evaluated, within three input-intelligibility conditions. The role of language-specific prosodic cue evaluated in Experiment 3B was rather different than the one found in Experiment 2 with universal prosodic information. Indeed, in opposition to what was found in Experiment 2, listeners’ performance was not globally affected by the presence of lexical stress patterns. Yet, lexical stress has a role in speech segmentation when the signal is degraded, possibly constraining the extraction of statistical segmentation outputs to the ones with higher support. With phonetically intact speech, both highand low-TP-words were correctly extracted from the stream, both independently of the number of cues available in the stream and of the stress location. Notably, with mildly degraded signal, a lexical stress effect emerged, narrowing the extraction of AL-units: lowTP-words, which with intact speech were correctly parsed from the stream, were no longer selected as possible words of the new language, in the mildly degraded condition. Additionally, a strongly degraded signal (i.e., 10dB SNR) maximized the strength of primary stress effects (cf. Mattys et al., 2005). Consequently, in this condition, only the statistical outputs with the highest TPs that obeyed to the default stress pattern in EP were selected as plausible words. In stressed conditions diverging from the default one statistical learning was inhibited. The availability of a congruent lexical stress pattern did not promote any quantitative benefit in ALL. Instead, it narrowed the selection of statistical Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 280 segmentation outputs to the ones which had the strongest support from the conjunction of the two segmentation cues available in the stream, enabling the exclusive extraction of these units (i.e., high-TP-words obeying to the default stress pattern in EP). Remarkably this extraction was as efficient when the signal was quiet distorted as when it was intact. 7.1.2. Speech segmentation in cognitive noise The importance of the listening conditions in the weighting of various speech segmentation cues is an innovative and fundamental aspect of Mattys et al.’s (2005; Mattys, 2004) proposal. Until now this impact has been only evaluated through physical degradation of the signal (i.e., physical noise). Note, however, that Mattys et al. (2005) already suggested that the listening conditions go beyond this physical quality. Under this view, deteriorating the listening conditions while maintaining the signal physically intact would also affect the kind of cues predominantly used in speech segmentation. This proposal is based on naturalistic observations. Indeed, in natural settings, speech perception, and hence speech segmentation, often occur not only in conditions of physical noise, but also in conditions of attentional load, far from the optimal conditions of a laboratorial sound-proof room. The impact of cognitive noise (or attention-load) in speech segmentation was evaluated in Experiments 4, 5, and 6 of the present study. Experiment 4 Experiment 4 (see chapter four) explored if a hierarchical organization of Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 281 coarticulation and TPs, similar to the one found with physical noise (i.e., in Experiment 1 of the present work), would also be observed in cognitive noise conditions. The ALL patterns observed in Experiment 4 have shown that cognitive noise differently affects the use of these two cues. Notably, we have found a weighting change of these sublexical segmentation cues modulated by cognitive noise that tended to mirror the one observed with physical noise: The gradual effect of cognitive noise had a clear impact on statistical information but not on coarticulation. When strong cognitive noise prevented participants to optimally extract TP-units from the stream, listeners did rely on coarticulation when TPs were incongruent with it. This pattern of results is in sharp contrast to the one found with strong physical noise in Experiment 1. When coarticulation and TPs were congruently available in the AL stream, their integration was also achieved in strongly degraded attentional conditions. Indeed, with strong cognitive noise, a redundancy gain (better performance when both statistics and coarticulation suggested the same word-boundaries than when only TPs were available) was exclusively found for the AL words with low TPs. A soft reduction of the attentional resources available to speech segmentation through passive exposition to visual stimuli (i.e., weak cognitive noise condition, in Experiment 4) seemed to have the same (at least quantitative) impact as a mild physical degradation (i.e., mildly degraded condition in Experiment 1) of the auditory AL stream. In both cases, the product of the available incongruent cues was disrupted. Notably, the impoverishment of the statistically-driven segmentation Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 282 procedure by strong cognitive noise did not completely prevent statistical learning. Although listeners exposed to the AL with only statistical information available in the stream presented an impaired performance in the ALL-test (i.e., they were not able to select low-TP-words as the units of the new language), they still presented an above chance performance for AL units with the highest TPs. Thus, although statistically-driven speech segmentation is affected by cognitive noise (i.e., a severe reduction of attentional resources available) it still operates in these attentional adverse contexts. In order to specifically assess the impact of attention in statistically-driven speech segmentation, in Experiments 5 and 6 (see chapter five), we evaluated directly the role of amount of familiarization, and the role of the two facets of attention (i.e., as a selective mechanism, and as limited-capacity resources) in TPbased speech segmentation. 7.1.3. Cognitive noise and statistically-driven speech segmentation Experiment 5 In Experiment 5, we assessed the extent to which the amount of ALfamiliarization (i.e., between-participants familiarization phases of 7-min, 14-min, 21-min) would influence statistically-driven (i.e., TPs-based) speech segmentation in conditions of low- (i.e., weak cognitive noise) and high-attention load (i.e., strong cognitive noise). In the high-attention load condition, independently of the amount of Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 283 exposition to the AL, participants always succeeded in extracting the high-TPwords from the stream. However, it was only with 7-min of familiarization time that they were also able to extract the low-TP-words. With familiarization time, performance for high-TP-words actually increased, while performance for low-TPwords decreased, as though the cognitive system adapts itself to the reduction of attentional resources by focusing on the most salient “lexical” units. In contrast, in the low-attention load condition, the system seems to privilege the extraction of high-TP-words only when the familiarization time is drastically shortened (namely, with 7-min). Even in this shortest familiarization condition, performance was above chance level for both lowand high-TP-words, as it was in the other familiarization time conditions. In addition, with increasing familiarization times, the system is able to extract both highand low-TP-words at similar performance levels, although this procedure seems to be done at the expense of performance on high-TP-words. In Experiment 5 we were able to demonstrate clearly that attention-load affects statistical computation in a graded rather than in an all-or-none manner. This was already suggested in the strong cognitive noise condition when only statistical information was available (Experiment 4 presented in chapter four), Indeed, cognitive noise (or in other words, high-attention load) has a different impact on the extraction of highand low-TP-words. Reducing attentional resources and/or diverting attention from the AL do not completely prevent statistical learning. Instead, the extraction of the most salient units is slowed down, while the extraction of the less salient units is progressively inhibited. Indeed, it was in the longer familiarization time at study in Experiment 5 Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 284 (i.e., 21-min) that the impact of the reduction of the available resources was more severe. In this condition only high-TP-words were correctly extracted from the stream, leading to the observation of a TP gradient effect. Furthermore, this result also suggests that the TP gradient effect observed in Experiment 4 in the strong cognitive noise condition can not be due to a weak reduction of attentional load. Experiment 6 This experiment was aimed at disentangling the role of attention as a selective mechanism and as a central resource in TPs computation. In Experiment 6 we demonstrated that the efficient operation of the statistical learning mechanism is not a matter of selective attention, but one of attentional resources. In a condition of severe reduction of the available attentional resources through the requirement to perform another resource-consuming task, as in the highattention load condition (i.e., strong cognitive noise), statistically-driven speech segmentation operates in an impaired fashion. Additionally, even when participants were focusing their attention on the AL stream (i.e., intentional learning condition), statistical computations still suffered a dramatic impairment as long as the attentional resources were scarce. The TP gradient effect observed in high-attention load in the TP – single cue conditions (in Experiment 4, 5, and 6) is a deleterious effect of reduced available resources. In order to operate efficiently, the statistical learning mechanism requires at least some attentional resources. When this requisite is not fulfilled, TPs are still computed but in an impaired manner, leading to the inability of correctly extracting from the stream the AL-units with the lower TPs. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 285 Statistical learning occurs even in conditions in which processing the auditory stream is not required and is not beneficial (being presented as potentially harmful) to the performance of an independent task. In that sense, although statistically-driven speech segmentation is dependent of the attentional-resources available, and thus not a purely “stimulus-driven” mechanism (cf. Moors & De Houwer, 2006), it may nevertheless be conceived as automatic. This is in line with Tzelgov’s (1997) and Bargh’s (1992) proposals that an automatic process is one that is autonomous, running without monitoring (it can be incidental), and even when it is not part of participants’ task requirement. In fact, in the present study we have shown that cognitive noise has a differential impact on the available sublexical cues. It is important to note that the results observed in cognitive noise (high attention-load) conditions cannot be due to a general increase in task difficulty with attentional load. Indeed, in Experiments 1, 2 and 3B, we have demonstrated that sensory degradation (i.e., reducing the physical quality of the signal by white-noise superimposition), which also increased task difficulty, had a different impact on the relative weighting of the available speech segmentation cues. In other words, the observation of different patterns of ALL as a function of the nature of the degradation instigated (i.e., physical or cognitive) rules out general task difficulty as an alternative account for the effects observed in the present study. Both physical and cognitive noise increase general task difficulty but they clearly have independent and different impacts on statistically driven speech segmentation processes and in the weighing of different segmentation cues. The experiments presented in this thesis and reviewed thus far suggest that the ALL paradigm can be a useful tool. However, this paradigm suffers from one Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 286 particular limitation regarding the task by which the speech segmentation procedures adopted by listeners during the AL-familiarization phase are accessed. Indeed, the two-alternative forced-choice test provides only indirect evidence of speech segmentation processes. Furthermore, the status of statistical segmentation outputs in what regards listeners’ linguistic knowledge, in particular in the context of a full mature speech perception system, is far from clear. 7.1.4. The ALL paradigm revised In Experiments 7, 8, and 9 (see chapter six) of the present study, we adopted an ALL approach combined with conventional experimental techniques in order to obtain evidence of on-line statistical speech segmentation and to evaluate whether statistical segmentation outputs could acquire some lexical status in the adult speech system. Experiment 7 In Experiment 7, adopting a variant of the word-spotting task, we directly demonstrated that indeed listeners are able to use on-line statistical information on speech segmentation. In “AL-word-spotting”, listeners were both more accurate to spot and faster to recognize TP-words than part-words embedded in four-syllable nonsense strings. This is the first direct evidence that after an AL-familiarization period, in which listeners were never presented with previously segmented ALstimuli, they were able to correctly extract from a four-syllabic sequence the trisyllabic stimuli that corresponded to AL words. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 287 Experiment 8 In this experiment, TP-words based on the statistical information available in the AL stream also strongly overlapped from onset with real words of listeners’ native language. In this condition, TP-words were not only correctly parsed from the stream, as revealed by the highly accurate performance in the forced-choice test (used as an ALL-index), but these statistical segmentation outputs had also a lexical status, exhibiting a lexical competition signature. Indeed, both immediately and one week after the AL familiarization period, the speed of processing of real (already existing) words, members of the same cohort than the TP-words parsed from the AL stream, was penalized in comparison to the processing of unrelated real words. These results add to previous findings of word learning in the absence of semantic knowledge (e.g., Bowers et al., 2005; Gaskell & Dumay, 2003a; 2003) and demonstrate that novel lexical items can be acquired in the context of a continuous speech stream, which probably closely resembles natural conditions of word learning. This lexicalization effect found with adult listeners is also in line with infants’ data (Saffran, 2001; Swingley, 2007). Experiment 9 In the last experiment of this work, we evaluated whether wordlikeness of AL-stimuli by itself, against the statistical information available in the AL-stream, could drive the segmentation processing and whether its byproducts could also be lexicalized. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 288 In this condition, listeners adopted a segmentation procedure compatible with their native language’s specific information. This is in line with previous findings suggesting the preponderance of listeners’ native language information (e.g., coarticulation; prosody; phonotactics; vowel harmony) in AL speech segmentation (e.g., Fernandes et al., 2007; Onnis et al., 2005: Shukla et al., 2007; Vroomen et al., 1998). Nevertheless, the byproducts of this segmentation procedure adopted by listeners in the incongruent cues condition (Experiment 9) were not lexicalized. This is in contrast with what was found in the congruent cues condition (Experiment 8). These results were at first sight quite surprising. However, the difference found in Experiment 9 between the pattern of results on ALL-test and the lexicalization-test is in accordance with the proposal of Leach and Samuel (in press) on the distinction between “items configuration” (in the present study, the phonological representations parsed from the AL stream) and “lexical engagement” (their involvement in lexical dynamics). Indeed, the ALL-effect is the outcome of listeners’ ability to use the available segmentation cues on-line, as it was demonstrated in Experiment 7 of the present study, which corresponds to listeners knowledge of those AL-words’ configuration. However this information cannot predict either the potential lexical status or the lexical engagement of the segmentation outputs. In natural settings listeners are exposed to a rich combination of different signal-derived (and also lexically-driven) segmentation cues, and hence, it would be counter-productive if all phonological forms to which listeners are exposed would be stored as different categories in mental lexicon (cf. Gaskell & Dumay, 2003b). The pattern of results found in Experiments 8 and 9 is in line with the suggestion that the consistency of segmentation hypothesis proposed by different available sources of Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 295 prosody was to be used only as a “higher” filter of statistical outputs, allowing or suppressing the units extracted on a TP-basis, then when TPs and other sublexical cues suggested the same parsing (i.e., congruent cues condition; e.g., TPs and coarticulation in Experiments 1 and 4; TPs and universal prosodic cues in Experiment 2; TPs and wordlikeness in Experiment 8) no redundancy gain would be observed. In fact, in Experiment 2, a redundancy gain was observed both with intact and physically degraded signal, that was promoted by the congruent availability of TPs and a universal prosodic cue. In Experiment 1, a redundancy gain derived from the congruent availability of TPs and coarticulation was also observed in intact conditions. Additionally, universal prosody was also able to drive the segmentation process even when no statistical outputs were congruent with it and hence no TPwords were compatible with the universal prosodic segmentation procedure. The same pattern of results was observed in intact and strong cognitive noise conditions when, instead of universal prosody, coarticulation was the cue available in the AL stream. Moreover, three findings of Experiment 3B suggest that language-specific prosody does not act as a filter of statistical segmentation outputs. First, stress pattern had no impact em quê?(either negative or positive) in listeners’ ALL performance with intact speech. As a matter of fact, all listeners independently of the presence (/absence) and location of the primary stress’ (either in the first, second, or third syllable of TP-words) were able to correctly extract from the stream the statistical segmentation outputs, at the same level. Second, stress pattern effects only emerged in degraded signal, in line with Mattys and colleagues proposal (Mattys, 2004; Mattys et al., 2005) and were maximized in the strongly degraded (i.e., 10dB SNR) Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 296 condition. Third, when both statistical and stress cues were available in the stream and the strong physical impoverishment (i.e., 10dB SNR) enabled stress cues to operate as efficiently as the statistical learning mechanism, only the units parsed from the stream that were highly compatible with the conjunction (i.e., intersection) of the available cues were considered by listeners reliable outcomes of the segmentation process. Thus, speech segmentation does not seem to depend on one predominant cue acting as a “higher” filter of the outputs of another lower weighted cue. Instead, a more parsimonious account would suggest that the outcome of speech segmentation is largely the result of the conjunction of the available cues, both in accordance with their nature and their weighted reliability. Therefore, when lower weighted cues are available in the stream, only units strongly supported by the conjunction of those cues will be considered. In particular, as it was observed in the strong degraded condition in Experiment 3B, wthe available cues were lower weighted and thus weakly reliable in speech segmentation, as it is the case of lexical stress – the “last segmentation resource heuristic” (cf. Mattys et al., 2005; see also Valiant & Levitt, 1996) and general domain TPs computation, only the units extracted from the stream that are strongly supported by both cues in conjunction were considered likely correct. In that case, no redundancy gain should be observed, as it was found in Experiment 3B. As long as a reliable cue is available in the speech stream (e.g., in intact speech, coarticulation: Experiment 1; in any physical condition, universal prosodic cues: Experiment 2), its byproducts are considered highly reliable. If these units are also supported by a lower weighted cue, such as the general domain TPs, a Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 297 redundancy gain will probably be observed. Indeed, in Experiment 1 of the present study when both coarticulation and TPs were congruently available in the stream, with intact speech, an additive effect (for both highand low-TP-words) allowed the optimization of speech segmentation processing. Accordingly, in Experiment 2, when TP-words were acoustically marked by a universal prosodic edge, their correct parsing was also optimized. In Experiment 8, when the statistical segmentation outputs also strongly overlapped with existing words, listeners’ ALL performance reached a high level: note that 17 out of 32 participants chose TP-words as the “lexical units” of the new language in at least 90% of the forced-choice test trials. As already suggested by Christiansen and colleagues’ computational work (Christiansen et al., 1998; Christiansen & Curtin, 2005), a cue that insures a deeper encoding of structural regularities of the input also enables the reliance on more subtle aspects of the input for making correct predictions. Thus the integration of different cues does not necessarily promote a quantity gain (i.e., more units parsed from the stream) but it can promote a qualitative one (i.e., the correct parsing and deeper encoding of highly supported units). Consequently, the correct extraction of the byproducts of the conjunction (i.e., intersection) of the available sources of information can minimize errors and unwanted over-generalizations. This holds true, particularly, in conditions in which the available cues are lower-weighted. The fact that speech segmentation is a robust phenomenon emerges from the pattern of results presented in this study. In other words, as already suggested on Mattys and colleagues’ proposal (Mattys, 2004; Mattys et al., 2005) and supported by studies of anatomical and functional neural organization (for a review see Scott & Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 298 Johnsrude, 2003), speech perception (and hence speech segmentation) is very resilient to different listening conditions. The availability of multiple cues (both signal-derived and lexically-driven), differently weighted according to their nature (whether the tier at which they are represented, Mattys et al., 2005, or their distinction on domain-generality and their role in particular languages, as suggested in the present chapter) enables the listener to rapidly fulfill speech segmentation task in any listening conditions. In other words, the robustness of speech segmentation processing seems to be due to the fact that congruently available cues can act as complementary sources of information, being differential weighted according to the type of impoverishment (cognitive or physical) instigated on listening conditions. Therefore, an apparent paradox is observed in degraded conditions when different congruent sources of information (in the present study, sublexical cues) are available in the speech stream. Noise modulates the weighting of these segmentation cues, but at the same time, it does not obligatorily impair speech segmentation in a drastic way (note for example that in Experiments 1, 2, and 4 of the present study, listeners exposed to the AL with congruent cues available, had an ALL performance above chance in any one of the degraded listening conditions). This is in line with Mattys’ (2004, Mattys et al., 2005) proposal that speech segmentation is largely the product of listening conditions and the available segmentation cues. Therefore, at least at the levels of noise studied in the present experiments, speech segmentation was achieved even in degraded conditions due to the fact that lower weighted cues were called upon to drive speech segmentation in impoverished conditions and thus disabled any observation of a main impact of noise. The present study also demonstrated that TPs computation still play an on- Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 299 line role (see Experiment 7 presented in chapter six) in language processing in adults. For sure, TPs computation plays a less fundamental role in adulthood than in infancy. Nevertheless, adults are still sensitive to statistical information such as conditional probabilities, not only in speech processing (e.g., Saffran et al., 1996a; Vroomen et al., 1998), but also in other sequence learning tasks (e.g., Boyer et al. 2005; Conway & Christiansen, 2005; Fiser & Aslin, 2001; Perruchet & Pacton, 2006). Indeed, since the intensely dynamic environment that surrounds us mainly involves stimuli occurring in temporal and/or spatial sequences, statistical learning would enable adults to rapidly structure novel sequences of events into emergent units (i.e., unitized into sets of temporally or spatially contiguous events). More generally, the process of reduction of uncertainty (Gibson, 1991) drives learners to seek invariant structure in the stimuli array (Gómez, 2002). However, the nature of the phonological representations corresponding to statistical segmentation outputs was until the present largely obscured. In other words, at least for adults, whether the statistical segmentation outputs could actually be lexicalized was largely unknown. In order to shed light on this aspect in the context of a full mature speech perception system we designed the two last experiments of the study reported in this thesis. 7.2.2. The nature of statistical segmentation output In Experiments 8 and 9, adopting the approach of Gaskell and Dumay (2003a; 2003b) within an ALL setting, we demonstrated that, at least in some conditions, the byproducts of statistical segmentation can be rapidly integrated in listeners’ mental lexicon, exhibiting lexical competition signatures. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 300 The results found in Experiment 8 provide an important contribution to the understanding of the nature of statistical segmentation output, adding to Saffran’s (2001) and Estes et al.’s (2007) findings with infant listeners. The output of the statistical learning mechanism can have a lexical status, being rapidly integrated in mental lexicon, even in the context of a fully mature perception system. Notably, the lexical engagement of the output of statistical segmentation observed in Experiment 8 is long-lasting and it continued to be strengthening between the first (immediately) and second (post-one week) moments of testing. This pattern of results is in line with Dumay and Gaskell’s (2007) proposal. It is also in accordance with general evidences of sleep being implicated in memory consolidation (for a review see Walker & Stickgold, 2004) and, in particular, in word learning (Dumay & Gaskell, 2007; Fenn, Nusbaum, & Margoliash, 2003; see also Clay, Bowers, Davis & Hanley, 2007) and vocabulary growth (Gais, Lucas, & Born, 2006). Models of spoken word recognition must be able not only to apprehend how different sublexical sources of information are weighted (or in Mattys et al., 2005 theoretical terms, are “hierarchically organized”) in speech segmentation, but also what their role (and the one of their congruency) in word learning. Regarding the role of segmentation cues in word learning, it is important to consider their involvement in structural changes with consequences on how listeners represent speech at different (e.g., sublexical, lexical) levels (see e.g., Sumner & Samuel, 2007), as it was the case in Experiment 8 of the present work. Future work should be aimed at examining which computational models may most adequately account for such word learning effects occurring in the context of Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 301 speech segmentation and, in particular, for such rapid acquisition and integration of statistical information. While any discussion of this point remains highly speculative in the absence of precise simulations, it seems to us that the Adaptive Resonance Theory or ART (Carpenter & Grossberg, 2003; Grossberg, Boardman, & Cohen, 1997; Grossberg & Myers, 2000; see also Sumner & Samuel, 2007, and Vitevitch & Luce, 1999) might be a good candidate to apprehend word recognition and word learning within a continuous flow, as well as to account for the present findings. In particular the ARTWORD model (Grossberg & Myers, 2000) which extends the earlier ARTPHONE model (Grossberg et al., 1997), could integrate word recognition and word learning within a continuous speech flow (Carpenter & Grossberg, 2003). ART is also able to explain the pattern of results found in both the ALL-test and the lexicalization of the Experiments 8 and 9 of the present study. The ART For the ART model, the resonance created by the match between bottom-up information and top-down expectations (stored in LTM) is responsible for the creation of the conscious spoken percept (Grossberg et al., 1997; Grossberg, 2003). While the spoken input unfolds, it activates speech segments that in turn activate “item” nodes. “Items” processed through time, generate an evolving spatial pattern of activation across WM that represents item information (which items are stored) and temporal order (the sequence in which they are stored). This enables items to be grouped, or unitized, into categories or “list chunks”. These “list chunks” have their maximal length restrained by WM, but can represent items or larger groupings of variable length (e.g., phonemes, syllables or words), with all of them represented Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 302 within the same masking field. In ART, masking denotes inhibition between “list chunks” (of different lengths): the longer the list chunks, the stronger will be the inhibition (or masking) promoted. However, list chunks will only fire when enough bottom-up evidence is received, which occurs after all the items in the list had been activated. Thus list chunks begin sending top-down feedback (representing a learned expectation stored in LTM) to associated items that are currently stored in WM. Chunks whose topdown signals are best matched to the sequence of incoming data reinforce the WM items and, in turn, receive greater bottom-up signals in return, wining masking field competition and thereby selectively amplifying and focusing attention upon consistent WM items, while suppressing inconsistent ones. This resonance loop will result in an equilibrated resonant state, thereby creating an emergent conscious percept. In order to parse a continuous stream into multiple discrete words, the positive feedback loop of any resonance cannot continue indefinitely. ART explains it as an example of “resonant reset” (Grossberg, 2003). “Mismatch reset” occurs when new phonemic information arrives and is sufficiently different from the current activated WM pattern to warrant an arousal burst that rapidly resets activity in the masking field. The network is “reset” into a nonresonant state, so that the next resonance can be initiated. In ART, the resonant state also drives the learning process (Carpenter & Grossberg, 2003). When novel words are presented in the speech stream, as it happens during an AL-familiarization period, novel information cannot form a good enough match with the expectations that are read-out by previously learned Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 303 recognition categories. This will trigger a memory search, or hypothesis testing, that leads to selection and learning of a new recognition category that really matches the input. Since in the masking field, longer list chunks will promote stronger inhibition, a chunk that codes a longer novel list like any one of the trisyllabic TP-words of the ALs used in this study, will have an a priori advantage over chunks that code shorter lists, like the familiar nonsense syllables that constitute it. It is precisely the masking advantage of longer list chunks that enables novel words to successfully compete with “amplifier” familiar chunks of shorter lists. When a top-down expectation achieves a good enough match with bottom-up data, this matching process focus attention upon those feature clusters in the bottom-up input that are expected. If the expectation is close enough to the input pattern then a state of resonance develops as the attentional focus takes hold, and the novel word is learned. The resonant process will be interrupted or terminated by mismatch reset and hence a new speech segmentation point will be defined. The repeated exposure to specific spatial patterns (e.g., the repeated occurrence of TP-words in the speech stream) permits learning by the LTM traces in the adaptive pathways between item nodes and the list nodes, reflecting the ALL and lexicalization effects found in Experiment 8 of the present study. ART is also able to explain the unobservation of lexicalization effects when incongruent cues were available in the AL stream (see Experiment 9 presented at chapter 6 of this thesis). When different informations are available in the stream, suggesting different competing chunks of the same length (e.g., TP-words and Partwords, in conditions in which both are supported by the available incongruent cues), resonance is actively reset by input mismatching even before reaching a stable state. The simultaneously activated competing list chunks promote only partial matching Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 304 resonant loops and hence the new information cannot be stably stored in LTM. The competitive chunks will strongly inhibit each other based on shared items and although both promote resonant loops, an equilibrated resonant state will not be achieved, and consequently novel words will not be stored in LTM. ART model resembles the PARSER model (Perruchet & Vinter, 1998) in its focus on primitive chunks and on the importance of selective attention in chunks formation, although PARSER does not propose any mechanism of word learning. Perruchet and colleagues (Perruchet, 2005; Perruchet & Pacton, 2006; Perruchet & Vinter, 1998) proposed that simply relying on chunk formation through the repeated occurrence of stimuli is sufficient to explain statistical learning effects. This obviates the need to postulate that listeners are able to perform statistical computations such as the ones involved in TPs. Based on this fairly simple algorithm PARSER could represent an elegant proposal, challenging the statistical learning approach (e.g., Cleeremans et al., 1998; Saffran et al., 1996a; 1996b). However, the pattern of results presented in this thesis does not seem to be compatible with Perruchet and colleagues’ (Perruchet, 2005; Perruchet & Pacton, 2006; Perruchet & Vinter, 1998) proposal. The PARSER PARSER (Perruchet & Vinter, 1998) is a fragment-based model that proposes that the knowledge acquired in statistical learning tasks might result from simple learning algorithms, giving rise to little more than explicitly memorized short fragments or “chunks”. Thus sensitivity to statistical structure is conceived as a simple by-product of chunk formation (see for this discussion: Bonatti et al., 2006; Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 311 process of lexical access could be initiated. Therefore, segmentation cues often regard the discovery of where word boundaries are located, while perceptual units correspond to what intermediate representations possibly mediate lexical access. Therefore, the process of classification is obviously distinct from the process of segmentation, and segmentation units can differ from classificatory ones (Kolinsky, Morais & Cluytens, 1995). Segmentation (where) and classification (what) units Whereas classification inevitably entails segmentation, the reverse is not necessarily true. For example, a prosodic segmentation procedure such as the MSS (Cutler & Norris, 1988) is a segmentation device based on the rhythmic properties of listeners’ native language (e.g., in English and in Dutch, at the onset of any strong syllable; Cutler & Norris, 1988; Vroomen, van Zon, & de Gelder, 1996). Its role is to indicate where in the speech stream the process of lexical access could be initiated. Yet, this segmentation device does not provide any information about the prelexical representations (e.g., feet, syllables) of such classification. In this sense, segmentation procedures are compatible both with models of speech perception involving prelexical stages as well as with models involving no prelexical unit. Note however that word recognition would benefit if the speech code could be “broken” prelexically. The mediation of a prelexical level of representation (perhaps separated at sub-stages, McQueen, 2005) between the output of the auditory system and the lexicon could remove considerably redundancy that otherwise would have to exist at lexical level (but see Goldinger 1996; 1998). There is no general consensus, however, in what regards the size of these prelexical units. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 312 The size of prelexical units Among proposals of prelexical units, the most influential is probably the one of Mehler and colleagues (Mehler, Dommergues, Frauenfelder, & Segui, 1981). In a version of the monitoring paradigm (i.e., in the fragment detection task), response times of French listeners were longer when the target (e.g., pa) and the carrier word (e.g., pal#mier25) did not match on syllabic structure than when they did (e.g., pa in pa#lace; pal in pal#mier). This “syllabic effect” was taken as evidence for a syllabification procedure and hence for the existence of classification prelexical units (i.e., the syllable) that would operate during the process of word recognition: “(…) the syllable is probably the output of the segmentation device operating upon the acoustic signal. The syllable is then used to access the lexicon…” (p. 342, Mehler et al., 1981). However, Cutler, Mehler, Norris, & Segui (1983) demonstrated in a crosslinguistic study with French and English listeners that the syllabic effect was not (or at least it did not seem to be) universal. Indeed, English listeners did not present the syllabic effect observed in French listeners. This cross-linguistic difference was interpreted as a consequence of the specific structure of the two languages, probably based in their rhythmic properties. In French (a syllable-timed language), listeners segment the speech stream into syllables, since their language displays a clear and reduced variety of syllabic structures. Yet, in English (a stress-timed language), listeners would not use this syllabic strategy because it is not suited to a language presenting widespread ambisyllabicity and a large variety of syllabic structures. The 25 The “#” marks a syllabic boundary. Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 313 study of Cutler et al. (1983) introduced the shift on the focus of research from the universal perceptual building block to language-specific segmentation strategies (see also Kolinsky et al., 2000). The syllabic effect: compatible or incompatible with sublexical cues? Syllables have often been found to play a significant functional role in spoken word processing, not only in Romance Languages (e.g., in French: Dumay, Benraïss, Barriol, Colin, Radeau, & Besson, 2001; Dumay, Banel, Frauenfelder & Content, 1998; Kolinsky et al., 1995; Mehler et al., 1981; Pallier, Sébastian Gallés, Felguera, Christophe & Mehler, 1993), but also in other, non-Romance ones, like English (e.g., Finney, Protopapas & Eimas, 1996; Mattys & Melhorn, 2005) and Dutch (e.g., Zwitserlood, Schriefers, Lahiri, & van Donselaar, 1993). However, even in Romance Languages, the syllabic effect is not systematically observed in all experimental conditions nor with different materials (e.g., Content, Meunier, Kearns, & Frauenfelder, 2001; Sebastian-Gallés, Dupoux, Segui, & Mehler, 1992). Morais, Content, Cary, Mehler, & Segui (1989), demonstrated the important role of syllables in Portuguese speech processing, using a variant of the sequence monitoring paradigm. However, using the migration paradigm (for an overview see Kolinsky & Morais, 1996), Kolinsky et al. (1995; Kolinsky, 1998) reported that in Portuguese (for both EP and BP) syllables seems to have a less important role than other sub-syllabic units. Indeed, in Portuguese the initial consonant is the property that blends the most. Is it possible that listeners (at least Portuguese listeners) could be sensitive to different prelexical units? Kolinsky (1998) already suggested that all listeners are Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 314 able to use different types of routines (e.g., syllabic and subsyllabic) according to task requirements (for a striking example see Goldinger & Azuma, 2003) and language-specific strategies. However, in what regards the sequence monitoring paradigm, its use to study pre-lexical units can be seriously questioned, since it probably involves metaphonological processing (Kolinsky, 1998). Therefore, while results with the fragment monitoring task could reflect listeners’ strategies rather than underlying speech representations, the results found with tasks that do not require intentional retrieval of sublexical units (e.g., the migration paradigm) reduce the probability of the use of strategies and late (e.g., metaphonological) representations (see e.g., Kolinsky, 1998; Kolinsky et al., 2000). Interestingly, using the migration paradigm the role of the initial consonant was also found in EP preliterate children (Castro, Vicente, Morais, Kolinsky, & Cluytens, 1995) and illiterate adults (see Kolinsky, 1998; Kolinsky et al., 1995), suggesting the prelexical locus of this effect. Using the fragment-priming paradigm, Tabossi, Collina, Mazzetti, and Zopello (2000) showed that in Italian, words matching the syllable structure of the fragments (e.g., si#lenzio – silence - and si#l) were more strongly activated than words that mismatch the fragment (e.g., sil#vestre – sylvan – and sil#l), which is compatible with the “syllabic effect”. Nevertheless, small durational differences between the vowels in these fragments appear to have signaled the difference in syllabic structure. This could suggest that the “syllabic effect” is instead a product of subsegmental differences in the input (see also Content et al., 2001). In other words, Tabossi et al.’s (2000) syllabic effect could be due to sublexical segmentation procedures and not necessarily related with prelexical classification units. Indeed, the Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 315 importance of syllables in speech perception has often been related to the fact that, in natural speech, coarticulation is lower between than within syllables (Liberman, Cooper, Shankweiler, & Studdert-Kennedy, 1967). Furthermore, as demonstrated in the work presented in this thesis, coarticulation is a reliable cue in speech segmentation in adulthood (Experiment 1 and 4 of the present study). Therefore, sublexical segmentation cues such as the degree of coarticulation between segments could, by themselves, account for the observed syllabic effect. Indeed, using the migration paradigm, Kolinsky et al. (1995) also demonstrated the syllabic effect in French. At first sigh these results could suggest that the segmentation devices were the only procedures required for the observation of syllabic effects without calling upon any prelexical classificatory unit. However, if that was the case, we would always observe a syllabic effect mediated by sublexical segmentation devices, in every language. Note that the results observed in Portuguese with the migration paradigm (Kolinsky, 1998; Kolinsky et al., 1995) refute this argument. Based on the results found with coarticulatory cues in the study presented in this thesis (see Experiments 1 and 4, in chapters two and four, respectively), we already know that in EP coarticulation is a reliable segmentation cue. This means that the migration of a subsyllabic unit in EP (Castro et al., 1995; Kolinsky et al., 1995; Kolinsky, 1998) cannot be simply due to weaker coarticulatory influences in EP than in languages in which syllables are the unit that migrates the most (e.g., in French, Kolinsky et al., 1995; in English, Mattys & Melhorn, 2005). Thus, a variety of (both segmentation and classificatory) units could be involved in prelexical level, with their weight depending on their relevance in any particular language. An approach to the speech segmentation problem relying on several Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 316 sources of information extracted and integrated from infancy on, through all linguistic development, may provide a more powerful and realistic account of listeners’ accurateness in recognizing continuous speech (Christiansen & Curtin, 2005; Kolinsky et al., 2000; Mattys et al., 2005). Speech perception system could exploit different types of segmentation cues in order to constrain highly structured sublexical representations (Kolinsky et al., 1995). 7.3.2. Cross-species statistical learning The acquisition and processing of language is, beyond doubt, governed by universal constraints, many of which are derived from innate properties of the human brain (Christiansen & Ellefson, 2002). In Chomsky’s (1980; 1986) approach, the constraints on the acquisition and processing of language are represented in the form of a Universal Grammar (UG), i.e., a large biological endowment of linguistic knowledge, highly abstract, comprising a complex set of linguistic rules and principles that could not simply be acquired from exposure to language during development. Such characteristics are thus considered innate rather than learned and species-specific. Alternatively, the approach of Christiansen and Ellefson (2002) stresses the adaptation of linguistic structures to the biological substrate of the human brain rather than concentrating on biological changes to accommodate language. Therefore, many constraints on linguistic adaptation derive from non-linguistic limitations on learning and processing of hierarchically organized sequential structure. The underlying mechanisms existed prior to the appearance of language Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 317 but, presumably, also underwent changes after the emergence of language. Consequently, many language universals may reflect non-linguistic, cognitive constraints on learning and processing of sequential structure, rather than an innate UG. These aspects would correspond in Hauser, Chomsky, and Fitch’s (2002) terms to the Faculty of Language – broad sense (FLB). Although many aspects of FLB are shared with other species, the core aspects of Faculty of Language – narrow sense (FLN), such as the recursive aspect (and the notion of discrete infinity; Fitch, Hauser, & Chomsky, 2005), appear to lack any analog in non-human species and are thus unique in humans (for a discussion see Fitch et al., 2005; Jackendoff & Pinker, 2005; Pinker & Jackendoff, 2005). In this regard, Hauser et al. (2002) argue for a qualitative difference on the types of mechanism underlying language processing. Statistical learning based on TPs computation would be one example of a phylogenetic-based mechanism, or as in Hauser et al. (2002) terms, a FLB one. Indeed, Yang (2004) already suggested that general-domain statistical learning can be constrained by innate, domain-specific principles of linguistic structures, where UG could “instruct” the learner what cues (or what type of regularities) to attend. Note, however, that what is not so far clearily known are the sorts of statistical informations human listeners are able to use in natural linguistic settings. As Seidenberg, MacDonald, and Saffran (2002) pointed out: “Our understanding of the contribution of statistical learning is limited by incomplete knowledge of the kinds of statistics infants encode and whether these are the ones relevant to natural language. This view also leaves open the critical question of why only humans acquire language, as many other species are capable of simple forms of statistical learning.” (p. 554). Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 318 Therefore, although statistical learning is also observed in non-human primates (Hauser et al., 2001) and even in rats (Toro & Trobalón, 2005), the phylogenetic root of this mechanism does not imply that non-human species are able to extract the same regularities from a speech stream that human listeners do (e.g., Trout, 2001). Indeed, similar cross-species behaviors do not necessarily imply the same underlying computational abilities and units of analysis (Weiss & Newport, 2006). Cross-species differences in statistical computations Rats can segment an AL speech stream (Toro & Trobalón, 2005). However, they do so by using the overall frequency of co-occurrence among AL-items, a different kind of computation than that the one used by humans (who can rely on TPs: Aslin et al., 1998; Saffran et al., 1996a). On the contrary, cotton-top tamarins (like humans) are sensitive to TPs between adjacent syllables (Hauser et al., 2001) and also between non-adjacent ones (Newport, Hauser, Spaepen, & Aslin, 2004). These non-human primates are also sensitive to TPs between nonadjacent vowels but not between non-adjacent consonants (Newport et al., 2004). And here is where the distinction between statistical learning in humans and in other species probably lays. In opposition to what is observed in other species, human listeners are able to extract statistical regularities both between nonadjacent consonants and between nonadjacent vowels (Newport & Aslin, 2004). Bonatti, Pen}a, Nespor, & Mehler (2005) have demonstrated that human listeners find it easier to track TPs between nonadjacent consonant than between nonadjacent vowels, in opposite to what was Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 319 found with tamarins (Newport et al., 2004). Therefore, consonants seem to be more suitable than vowels to parse speech streams using statistical dependencies (see Bonatti, Pen}a, Nespor, & Mehler, 2007; Mehler, Pen}a, Nespor, & Bonatti, 2006). This could suggest a particular and species-specific role of consonants in language processing. Vowels and Consonants in human statistical learning An emergent bulk of research has suggested vowel-consonant independent (structural and functional) treatment (e.g., Boatman, Hall, Goldstein, Lesser & Gordon, 1997; Caramazza, Chialant, Capasso, & Miceli, 2000; Cutler, SebastiánGallés, Soler-Vilagelu, & van Ooijen, 2000; Nespor, Pen}a, & Mehler, 2003), since an early phase in linguistic development (Nazzi, 2005; Nazzi & New, 2007). In particular, Nespor et al. (2003) noted that vowels are the main carriers of prosodic information, providing information at the prosodic and syntactic levels and defining some basic linguistic distinctions at rhythmic properties of languages (e.g., Ramus, Nespor, & Mehler, 1999). Consonants have a privileged role at lexical level, outnumbering the number of vowels in many languages, and performing a primary role in lexical items’ distinction. In this particular demonstration of the different roles of vowels and consonants in speech processing, the ALL has revealed itself an important tool. The compatibility between particular types of statistical learning and natural languages can suggest that the structure of natural languages may be formed, at least in part, by the constraints and selectivity of what human learners find easy to acquire. In sum, although humans share a basic statistical learning mechanism with Chapter 7. General Discussion Chapter 7. General DiscussionChapter 7. General Discussion Chapter 7. General Discussion 320 other species (particularly other primates), it differs in what regards the more complex types of computations humans are able to perform. The study of the representations underlying statistical computations in humans can provide clues for the role of different linguistic units (e.g., consonant, vowels, syllables) in speech processing. ALL is therefore an additional, complementary paradigm for exploring and testing hypotheses about language evolution, language acquisition and sublexical segmentation procedures in adulthood. It also enables the study of speech processing of human listeners in different stages of linguistic development, as well as cross-species comparative studies (e.g., Hauser et al., 2001; Newport et al., 2004). 7.4. The ALL paradigm revisited The AL speech stream used in ALL experiments is clearly an extreme case: all words are usually of the same length, with no short word embeddings in ALwords, and with a reduced artificial “lexicon”. Thus AL obviously deviates from natural languages. However, this does not necessarily diminish its impact or usefulness to the study of language processing. An additional fact confirming the ALL importance is the fact that the ALL paradigm provides a highly controlled situation in which the available sources of information can be systematically manipulated. However, the generally used twoalternative forced choice task, which is an indirect (off-line) measure of ALL and hence of speech segmentation, is the Achilles’ heel of this paradigm. In the present study (experiments 7, 8 and 9; see chapter 6), we have demonstrated that the ALL REFERENCES References ReferencesReferences References 328 References References References References 329 Abaurre, M. B., & Galves, C. (1998). Rhythmic differences between European and Brazilian Portuguese: an optimalist and minimalist approach. D.E.L.T.A., 14 , 377-403. Abrams, R.A., & Balota, D.A., (1991). Mental chronometry: Beyond reaction time. Psychological Science, 2, 153-157. Allopena, P. D., Magnuson, J. S., & Tanenhaus, M. K. (1998). Tracking the time course of spoken word recognition using eye movements: Evidence for continuous mapping models. Journal of Memory and Language, 38, 419-439. Aslin, R. N., Saffran, J. R., & Newport, E. L. (1998). Computation of Conditional Probabilities by 8-Month-Old Infants. Psychological Science, 9, 321-324. Aslin, R. N., Woodward, J. Z., LaMendola, N. P., & Bever, T. G. (1996). Models of word segmentation in fluent maternal speech to infants. In J. L. Morgan and K. Demuth (Eds.), Signal to syntax: Bootstrapping from speech to grammar in early acquisition (pp. 117-134). Hillsdale, NJ: Erlbaum. Auer, E. T., & Luce, P. A. (2005). Probabilistic phonotactics in spoken word recognition. In D. B. Pisoni & R. E. Remez (Eds.), Handbook of speech perception (pp. 610-630). New York, NY: Blackwell Bagou, O., Fougeron, C., Fauenfelder, U. H. (2002). Contribution of prosody to the segmentation and storage of "words" in the acquisition of a new mini-language. In B. Bel & I. Marlien (Eds.), Proceedings of the Speech Prosody 2002 Conference (pp. 59–62). Aix-en-Provence: Laboratoire Parole et Langage. Bailey, T. M., & Hahn, U. (2001). Determinants of Wordlikeness: Phonotactics or Lexical Neighborhoods? Journal of Memory and Language, 44, 568-591. References ReferencesReferences References 330 Bargh, J. A. (1992). The ecology of automaticity : Toward establishing the conditions needed to produce automatic processing effects. American Journal of Psychology, 105, 181-199. Böcker, K. B. E., Bastiaansen, M. C. M., Vroomen, J., Brunia, C. H. M., & de Gelder, B. (1999). An ERP correlate of metrical stress in spoken word recognition. Psychophisiology, 36, 706-720. Boatman, D., Hall, C., Goldstein, M.H., Lesser, R., & Gordon, B. (1997). Neuroperceptual differences in consonant and vowel discrimination: As revealed by direct cortical electrical interference. Cortex, 33, 83-98. Bonatti, L. L., Pen}a, M., Nespor, M., & Mehler, J. (2005). Linguistic Constraints on Statistical Computations: The Role of Consonants and Vowels in Continuous Speech Processing. Psychological Science, 16, 451-459. Bonnati, L. L., Peña, M., Nespor, M. & Mehler, J. (2006). How to hit Scylla without avoiding Charybdis? Comment on Perruchet, Tyler, Galland, and Peereman (2004). Journal of Experimental Psychology: General, 135, 314-321. Bonatti, L. L., Pen}a, M., Nespor, M., & Mehler, J. (2007). On Consonants, Vowels, Chickens, and Eggs. Psychological Science, 18, 924-925. Bortfeld, H., Morgan, J. L., Golinkoff, R. M., and Rathbum, K. (2005). Mommy and Me: Familiar Names Help Launch Babies Into Speech-Stream Segmentation. Psychological Science, 16, 298-304. Boyer, M., Destrebecqz, A., & Cleeremans, A. (2005). Processing abstract sequence structure: learning without knowing, or knowing without learning? Psychological Research, 69, 383-398. References References References References 331 Bowers, J. S. (1999). Priming is not all bias: Commentary on Ratcliff and McKoon (1997). Psychological Review, 106, 582–596. Bowers, J. S. (2000). In defense of abstractionist theories of word identification and repetition priming. Psychonomic Bulletin & Review, 7, 83-99. Bowers, J., S. & Collin, C. J. (2004). Is speech perception modular or interactive? Trends in Cognitive Science, 8, 3-5. Bowers, J. S., Davis, C. J., & Hanley, D. A. (2005). Interfering neighbours: The impact of novel word learning on the identification of visually similar words. Cognition, 97, B45-B54. Brent, M. R., & Cartwright, T. A. (1996). Distributional regularity and phonotactic constraints are useful for segmentation. Cognition, 61, 93-125. Byrd, D. (1996). Influences on articulatory timing in consonant sequences. Journal of Phonetics, 24, 209-244. Byrd, D., & Saltzman, E. (1998). Intragestural dynamics of multiple prosodic boundaries. Journal of Phonetics, 26, 173-199. Cairns, P., Shillcock, R. C., Chater, N., Levy, J. (1997). Bootstrapping word boundaries: a bottom-up corpus-based approach to speech segmentation. Cognitive Psychology, 33, 111-153. Caramazza, A., Chialant, D., Capasso, R., & Miceli, G. (2000). Separable processing o fconsonants and vowels. Nature, 403, 428-430. Carpenter, G. A., & Grossberg, A. (2003). Adaptive Resonance Theory. In M.A. Arbid (Ed.), The Handbook of Brain Theory and Neural Networks (pp. 87-90). Cambridge, MA: MIT Press. References ReferencesReferences References 332 Castro, S. L., Vicente, S., Morais, J., Kolinsky, R., & Cluytens, M. (1995). Segmental representations of Portuguese in 5and 6-year-olds: Evidence from dichotic listening. In I. Hub Faria & J. Freitas (Eds), Studies on the acquisition of Portuguese, (pp. 116). Lisbon: Colibri. Cho, T., & McQueen, J. M. (2005). Prosodic influences on consonant production in Dutch: Effects of prosodic boundaries, phrasal accent and lexical stress. Journal of Phonetics, 33, 121-157. Chomsky, N. (1980). Rules and representations. New York: Columbia University Press. Chomsky, N. (1986). Knowledge of Language. New York: Praeger. Christiansen, M. H., & Allen, J. (1997). Coping with Variation in Speech Segmentation. In A. Sorace, C. Heycock, & R. Shillcock (Eds.), Language Acquisition: Knowledge Representation and Processing (pp. 327-332). University of Edinburgh Press. Christiansen, M. H., Allen, J., & Seidenberg, M. S. (1998). Learning to Segment Speech Using Multiple Cues: A Connectionist Model. Language and Cognitive Processes, 13, 221-268. Christiansen, M. H., & Curtin, S. (2005). Integrating multiple cues in language acquisition: A computational study of early infant speech segmentation. In G. Houghton (Ed.), Connectionist models in cognitive psychology (pp. 347-372). Hove, UK: Psychology Press. Christiansen, M. H., & Ellefson, M. R. (2002). Linguistic Adaptation without Linguistic Constraints: The Role of Sequential Learning in Language References References References References 333 Evolution. In A. Wray (Ed), The transition to Language (pp. 335-358). Oxford: Oxford University Press. Christophe, A., Gout, A., Peperkamp, S., & Morgan, J. (2003). Discovering words in the continuous speech stream: the role of prosody. Journal of Phonetics, 31, 585-598. Christophe, A., Peperkamp, S., Pallier, C., Block, E., & Mehler, J. (2004). Phonological phrase boundaries constrain lexical access: I. Adult data. Journal of Memory and Language, 51, 523-547 Clay, F., Bowers, J. F., Davis, C. J., & Hanley, D. A. (2007). Teaching Adults New Words: The Role of Practice and Consolidation. Journal of Experimental Psychology: Learning, Memory, and Cognition, 33, 970-976. Cleeremans, A., Destrebecqz, A., & Boyer, M. (1998). Implicit learning: news from the front. Trends in Cognitive Sicences, 2, 406-416. Cleeremans, A. & Jiménez, L. (2002). Implicit learning and consciousness: A graded, dynamic perspective. In R. M. French & A. Cleeremans (Eds.), Implicit Learning and Consciousness. (pp. 1-40).Hove, UK: Psychology Press. Cleeremans, A., & McClelland, J. L. (1991). Learning the structure of event sequences. Journal of Experimental Psychology: General, 120, 235-253. Coady, J. A., & Aslin, R. N. (2004). Young children’s sensitivity to probabilistic phonotactics in the developing lexicon. Journal of Experimental Child Psychology, 89, 183-213. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd edition). Hillsdale, NJ: Erlbaum. References ReferencesReferences References 334 Cohen, J. D., Aston-Jones, G., & Gilzenrat, M. S. (2004). A Systems-Level Perspective on Attention and Cognitive Control: Guided Activation, Adaptive Gating, Conflict Monitoring and Exploitation versus Exploration. In M. I. Posner (Ed.). Cognitive Neuroscience of Attention (chapter 6, pp. 61-90). Guilford Publications. Content, A., Meunier, C., Kearns, R.K., & Frauenfelder, U.H. (2001). Sequence detection in pseudowords in French: Where is the syllable effect? Language and Cognitive Processes, 16, 609-636. Conway, C. M., & Christiansen, M. H. (2005). Modality-Constrained Statistical Learning of Tactile, Visual, and Auditory Sequences. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31, 24-39. Conway, C. M., & Christiansen, M. H. (2006). Statistical Learning Within and Between Modalities: Pitting Abstract Against Stimulus-Specific Representations. Psychological Science, 17, 905-912. Creel, S. C., Tanenhaus, M. K., & Aslin, R. N. (2006). Consequences of Lexical Stress on Learning an Artificial Lexicon. Journal of Experimental Psychology: Learning, Memory, and Cognition, 32, 15-32. Cunillera, T., Toro, J. M., Sebastián-Gallés, N., Rodríguez-Fornells, A. (2006). The effects of stress and statistical cues on continuous speech segmentation: An event-related brain potential study. Brain Research, 1123, 168-178. Curtin, S., Mintz, T. H., & Christiansen, M. H. (2005). Stress changes the representational landscape: evidence from word segmentation. Cognition, 96, 233-262. References References References References 335 Cutler, A. (1986). Forbear is a homophone: Lexical prosody does not constraint lexical access. Language and Speech, 29, 201-220. Cutler, A., & Butterfield, S. (1992). Rhythmic cues to Speech Segmentation: Evidence from Juncture Misperception. Journal of Memory and Language, 31, 218-236. Cutler, A., Dahan, D., & van Donselaar, W. (1997). Prosody in the Comprehension of Spoken Language: A Literature Review. Language and Speech, 40, 141-201. Cutler, A., Mehler, J., Norris, D., & Segui, J. (1983). A language specific comprehension strategy. Nature, 304, 159-160. Cutler, A., & Norris, D. (1988). The role of Strong Syllables in Segmentation for Lexical Access. Journal of Experimental Psychology: Human Perception and Performance, 14, 113-121. Cutler, A., Sebastian-Gálles, N., Soler-Vilageliu, O., & van Ooijen, B. (2000). Constraints of vowels and consonants on lexical selection: Cross-linguistic comparisons. Memory and Cognition, 28, 746-755. Cutler, A., & van Donselaar, W. (2001). Vornaam is not (really) a Homophone: Lexical Prosody and Lexical Access. Language and Speech, 44, 171-195. d’Andrade, E., & Laks, B. (1996). Stress and Constituency: The Case of Portuguese. In J. Durant & B. Laks (Eds), Current Trends in Phonology: Models and Methods, volume I. (pp. 15-41). ESRI. Manchester: Universidade de Salford. Dahan, D. & Brent, M.R. (1999). On the discovery of novel wordlike units from utterances: an Artificial-language study with implications for native-language acquisition. Journal of Experimental Psychology: General, 128, 165-185. References ReferencesReferences References 336 Dahan, D. Gaskell, M. G. (2007). Temporal dynamics of ambiguity resolution: Evidence from spoken-word recognition. Journal of Memory and Language, 57, 483-501. Dahan, D., Magnuson, J. S., & Tanenhaus, M. K. (2001). Time course of frequency effects in spoken-word recognition: Evidence from eye movements. Cognitive Psychology, 42, 317-367. Davis, M. H., Marslen-Wilson, W. D., & Gaskell, M. G. (2002). Leading Up the Lexical Garden Path: Segmentation and Ambiguity in Spoken Word Recognition. Journal of Experimental Psychology: Human Perception and Performance, 28, 218-244. Dehaene-Lambertz, G., Hertz-Pannier, L., & Dubois, J. (2006). Nature and nurture in language acquisition: anatomical and functional brain-imaging studies in infants. Trends in Neurosciences, 29, 367-373. Dehaene-Lambertz, G., & Houston, D. (1998) Faster orientation latency toward native language in two-month-old infants. Language and Speech, 41, 21-43. Delgado-Martins, M. R. (2002). Fonética do Português: Trinta anos de investigação. Lisboa: Editorial Caminho. Dufour, S., Frauenfelder, U. H. & Peereman, R. (2007). Inhibitory priming in auditory word recognition: Is it really the product of response biases? Current Psychology Letters, Behaviour, Brain & Cognition, 22, Vol. 2. URL : http://cpl.revues.org/document2622.html. Dufour, S. & Peereman, R. (2003). Inhibitory priming effects in auditory word recognition: When the target’s competitors conflict with the prime word. Cognition, 88, B33-B44. References References References References 343 Huetting, F., & McQueen, J. M. (2007). The tug of war between phonological, semantic and shape information in language-mediated visual search. Journal of Memory and Language, 57, 460-482. Hunt, R. H., & Aslin, R. N. (2001). Statistical learning in a serial reaction time task: Simultaneous extraction of multiple statistics. Journal of Experimental Psychology: General, 130(4), 658-680. Iivonen, A., Niemi, T., & Paananen, M. (1998). Do F0 peaks coincide with lexical stresses? In S. Werner (Ed.), Nordic prosody: Proceedings of the VIIth conference, Joensuu 1996 (pp. 141–158). Frankfurt am Main: Peter Lang. Jackendoff, R., & Pinker, S. (2005). The nature of the language faculty and its implications for evolution of language (Reply to Fitch, Hauser, and Chomsky). Cognition, 97, 211-225. Jiang, Y., & Chun, M. M. (2001). Selective attention modulates implicit learning. Quarterly Journal of Experimental Psychology, 54A, 1105-1124. Jiménez, L., & Méndez, C. (1999). Which Attention is Needed for Implicit Sequence Learning? Journal of Experimental Psychology: Learning, Memory, and Cognition, 25, 236-259. Johnson, E. K., & Jusczyk, P. W. (2001). Word Segmentation by 8-Month-Olds: When Speech Cues Count More Than Statistics. Journal of Memory and Language, 44, 548-567. Johnson, E. K., Jusczyk, P. W., Cutler, A., & Norris, C. (2003). Lexical viability constraints on speech segmentation by infants. Cognitive Psychology, 46, 6597. References ReferencesReferences References 344 Ju, M., & Luce, P. A. (2004). Falling on Sensitive Ears: Constraints on Bilingual Lexical Activation. Psychological Science, 15, 314-318. Jusczyk, P. W. (1999). How infants begin to extract words from speech. Trends in Cognitive Science, 3, 323-328. Jusczyk, P. W., & Aslin, R. N. (1995). Infants' detection of sound patterns of words in fluent speech. Cognitive Psychology, 29, 1-23. Jusczyk, P.W., Hohne, E.A., & Bauman, A. (1999a). Infant's sensitivity to allophonic cues for word segmentation. Perception and Psychophysics, 61, 1465-1476. Jusczyk, P. W., Houston, D. M., & Newsome, M. (1999b). The beginning of word segmentation in English-learning infants. Cognitive Psychology, 39, 159–207. Jusczyk, P. W., Luce, P. A., & Charles-Luce, J. (1994). Infants’ sensitivity to phonotactic patterns in the native language. Journal of Memory and Language, 33, 630-645. Kahneman, D., & Chajczyk, D. (1983). Test of the automaticity of reading: Dilution of Stroop effects by color-irrelevant stimuli. Journal of Experimental Psychology: Human Perception and Performance, 9, 497-509. Keating, P. A. (in press). Phonetic Encoding of Prosodic Structure. In J. Harrington & M. Tabain (Eds.), Speech Production: Models, Phonetic Processes and Techniques. New York, USA: Psychology Press. Kirkham, N. Z., Slemmer, J. A., & Johnson, S. P. (2002). Visual statistical learning in infancy: evidence of a general learning mechanism. Cognition, 83, B35-B42. Klatt, D. H. (1980). Speech perception: A model of acousticphonetic analysis and lexical access. In R. A. Cole (Ed.), Perception and production of fluent speech (pp. 243-288). Hillsdale, N.J.: Erlbaum. References References References References 345 Kolinsky, R. (1998). Spoken Word Recognition: A Stage-processing Approach to Language Differences. European Journal of Cognitive Psychology, 10, 1-40. Kolinsky, R., Goetry, V., Radeau, M., & Morais, J. (2000). Human cognitive processes in speech segmentation and word recognition. In K. Jokinen, D. Heylen, & A. Nijholt (Eds), Proceedings of Centre for Evolutionary Language Engineering – University of Twelve Workshops on Natural Language Technology: Learning to behave / Internalizing Knowledge (pp. 149-169). Kolinsky, R., & Morais, J. (1996). Migrations in speech recognition. Language and Cognitive Processes, 11, 50-57 Kolinsky, R., Morais, J., & Cluytens, M. (1995). Intermediate Representations in Spoken Word Recognition: Evidence from Word Illusions. Journal of Memory and Language, 34, 19-40. Kuhn, G., & Dienes, Z. (2005). Implicit Learning of Nonlocal Musical Rules: Implicitly Learning more than chunks. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31, 1417-1432. Kühnert, B., & Nolan, F. (1999). The Origin of Coarticulation. In W. J. Hardcastle & N. Hewlett (Eds.), Coarticulation: Theoretical and Empirical Perspectives (pp. 61-75). Cambridge, UK: Cambridge University Press. Lavie, N. (1995). Perceptual Load as a Necessary Condition for Selective Attention. Journal of Experimental Psychology: Human Perception and Performance, 21, 451-468. Lavie, N. (2000). Selective attention and cognitive control: Dissociating attentional functions through different types of load. In S. Monsell & J. Driver (Eds.), References ReferencesReferences References 346 Control of Cognitive Processes: Attention and performance XVIII (pp. 175194). Cambridge, MA: MIT Press. Lavie, N. (2005). Distracted and confused? Selective attention under load. Trends in Cognitive Sciences, 9, 75-82. Leach, L., & Samuel, A. G. (in press). Lexical configuration and lexical engagement: When adults learn new words. Cognitive Psychology. Lehiste, I. (1960). An acoustic–phonetic study of internal open juncture. Phonetica, 5 (Suppl. 5), 5-54. Liberman, A. M., Cooper, F. S., Shankweiler, D. P., & Studdert-Kennedy, M. (1967). Perception of the speech code. Pshychological Review, 74, 431 - 461. Liberman, A. M. Studdert-Kennedy, M. (1978). Phonetic perception. In R. Held, H. Leibowitz H.-L. Teuber (Eds.), Handbook of sensory physiology: Perception (VIII, 143-178). Berlin: Springer-Verlag. Logan, D. (1988). Toward an Instance Theory of Automatization. Psychological Review, 95, 492-527. Logan, G. D., Taylor, S. E., & Etherton, J. L. (1999). Attention and automaticity: Toward a theoretical integration. Psychological Research, 62, 165-181. Luce, P. A. (1986). A computational analysis of uniqueness points in auditory word recognition. Perception & Psychophysics, 39, 155–158. Luce, P. A., Goldinger, S. D., Auer, E. T., & Vitevitch, M. S. (2000). Phonetic priming, neighborhood activation, and PARSYN. Perception & Psychophysics, 62, 615-625. Luce, P. A., & Large, N. (2001). Phonotactics, neighborhood density, and entropy in spoken word recognition. Language and Cognitive Processes, 16, 565–581. References References References References 347 Luce, P. A., & Pisoni, D. B. (1998). Recognizing Spoken Words: The Neighborhood Activation Model. Ear & Hearing, 19, 1-36. Magnuson, J. S., Dixon, J. A., Tanenhaus, M. K., & Aslin, R. N. (2007). The Dynamics of Lexical Competition During Spoken Word Recognition. Cognitive Science, 31, 133-156. Magnuson, J. S., Tanenhaus, M. K., Aslin, R. N., & Dahan, D. (2003). The Time Course of Spoken Word Learning and Recognition: Studies with Artificial Lexicons. Journal of Experimental Psychology: General, 132, 202-227. Manuel, S. (1999). Cross-language studies: relating language-particular coarticulation patterns to other language-particular facts. In W. J. Hardcastle & N. Hewlett (Eds.), Coarticulation: Theoretical and Empirical Perspectives (pp. 179-198). Cambridge, UK: Cambridge University Press. Marslen-Wilson, W. D. (1987). Functional Parallelism in spoken word recognition. Cognition, 25, 71-102. Marslen-Wilson, W. D., & Welsh, A. (1978). Processing interactions and lexical access during word recognition in continuous speech. Cognitive Psychology, 10, 29-63. Marslen-Wilson, W. D., Moss, H. E., & van Halen, S. (1996). Perceptual Distance and Competition in Lexical Access. Journal of Experimental Psychology: Human Perception and Performance, 22, 1376-1392. Mateus, M. H., & d’Andrade, E. (2000). The Phonology of Portuguese. Oxford, UK: Oxford University Press. Mattys, S. L. (2000). The perception of primary and secondary stress in English. Perception & Psychophysics, 62, 253-265. References ReferencesReferences References 348 Mattys, S. L. (2004). Stress Versus Coarticulation: Toward an Integrated Approach to Explicit Speech Segmentation. Journal of Experimental Psychology: Human Perception and Performance, 30, 397-408. Mattys, S. L., & Clark, J. H. (2002). Lexical activity in speech processing: evidence from pause detection. Journal of Memory and Language, 47, 343-359. Mattys, S. L., Jusczyck, P. W. (2001). Phonotactic cues for segmentation on fluent speech by infants. Cognition, 78, 91-121. Mattys, S. L., Jusczyk, P. W., Luce, P. A., & Morgan, J. L. (1999). Phonotactic and prosodic effects on word segmentation in infants. Cognitive Psychology, 38, 465-494. Mattys, S. L. & Melhorn, J. F. (2005). How do syllables contribute to the perception of spoken English? Evidence from the migration paradigm. Language and Speech, 48, 223-253. Mattys, S. L., & Melhorn, J., F. (2007). Sentential, lexical, and acoustic effects on the perception of word boundaries. Journal of the Acoustical Society of America, 122, 554-567. Mattys, S. L., & Melhorn, J., F., & White, L. (2007). Effects of syntactic expectations on speech segmentation. Journal of Experimental Psychology: Human Perception and Performance, 33, 960-977. Mattys, S. L., White, L., & Melhorn, J. F. (2005). Integration of multiple speech segmentation cues: A hierarchical framework. Journal of Experimental Psychology: General, 134, 477-500. McClelland, J. L., & Elman, J. L. (1986). The TRACE model of speech perception. Cognitive Psychology, 18, 1-86. [Document text truncated for crawler view.]