scieee AI-readable full text Open interactive document viewer

Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices

Kirkova-Naskova, Anastazija; Humanez-Berral, Pedro; Henderson, Alice

Abstract

These are the complete, compiled proceedings of the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. The document includes 10 chapters of original work by international scholars, an Introduction, Table of Contents, Acknowledgments, and Index.

Full text

Held at the University of Cantabria, Spain May 2024 Published 2025 Université Grenoble-Alpes https://doi.org/10.5281/zenodo.17319582 2025 Table of Contents Acknowledgements i Introduction: Looking back and moving forward for online EPIP proceedings Pedro Humánez-Berral, Alice Henderson and Anastazija Kirkova-Naskova iv The role of pronunciation in the assessment of listening skills in the Polish Matura exam Agnieszka Bryła-Cruz 1 EFL learners’ perceptions of naturalness, difficulty, and speech rate in native connected speech Ivana Duckinoska-Mihajlovska 15 Hungarian secondary school teachers’ pronunciation teaching practices: A questionnaire study Noémi Gyurka 28 Fluency in spontaneous description versus story-reading: Learner performance and views Xavier Martin-Rubió 43 A critical examination of research on the teaching of pronunciation in a second language Martha C. Pennington 55 Live or leave! Vowel quality and vowel length in the intelligibility of advanced Spanish-accented English Mateusz Pietrazek 71 Word-initial voicing in L2 English: Evaluations of /z/ productions before and after perceptual training Sidsel H. Rasmussen 84 Prepared to rescue Cinderella? Exploring Austrian EFL student teachers’ beliefs about pronunciation teaching and learning Karin Richter 97 Pronunciation features noticed by L1 Vietnamese MOOC users: In search of sociolinguistic salience Laura Rupp and Alice Henderson 109 The effects of explicit pronunciation instruction: A study of advanced Austrian EFL students Valerie White - Hautzinger, Marija Djenadic, Carolin Rumpler, and Miriam Fiala 118 I ndex 129 i Acknowledgements Working on the EPIP8 Proceedings has been a joyful experience for us as editors. With every published proceedings, we feel like we have achieved another milestone – this time adding another book to the EPIP online collection hosted on the HAL platform. We truly hope that the open access publication approach will make the authors more visible and their research more available. Exchanging ideas with other colleagues in the field, in such a friendly and supportive atmosphere, has been enriching and rewarding in many ways. We are ever so grateful to the brilliant reviewers we were fortunate to get on board. They generously dedicated their time and expertise to provide constructive feedback by giving insightful comments and suggestions, which have significantly enhanced the quality of the chapters. We are indebted to the following scholars (in alphabetical order): Gemma Archer University of Strathclyde, Scotland Alex Baratta University of Manchester, England Małgorzata Baran-Łucarz University of Wrocław, Poland Vincent Chanethom Princeton University, USA Tracey Derwing University of Alberta, Canada Jonás Fouz-González University of Murcia, Spain Joshua Gordon University of Northern Iowa, USA Nadine Herry-Bénit Paris Nanterre University, France Sophie Herment Aix-Marseille University, France Pekka Lintunen University of Turku, Finland John Levis Iowa State University, USA José Mompéan-Gonzalez University of Murcia, Spain Wayne Rimmer University of Manchester, England Arkadiusz Royczyk University of Silesia in Katowice, Poland Radek Skarnitzl Charles University in Prague, Czech Republic Sinem Sonsaat-Hegelheimer Iowa State University, USA Nick Travers Camosun College, Canada Gabor Turscan Aix-Marseille University, France Elina Vasu Tampere University, Finland Jan Volin Charles University in Prague, Czech Republic We also extend our heartfelt appreciation to the incredible authors who contributed their research for publication. Their professionalism and broad-mindedness made our cooperation a sheer delight. ii We are thankful to the team at the University of Grenoble-Alpes for lending support with the technical aspect of the publication process, in particular Sandrine Corvey-Biron and Anne-Christine Jacob at the university libraries. Their gentle guidance has been instrumental in overcoming various challenges and ensuring the accuracy of the texts when published on the platform. Special thanks are due to Aynur Kaso from Ss. Cyril and Methodius University in Skopje for helping us with the design of the book cover. Finally, we wish to express our deepest and sincerest gratitude to our loved ones, our pillars of patience, encouragement, and unwavering support. You have kept us focused and motivated during this creative journey. Alice Henderson, Anastazija Kirkova-Naskova and Pedro Humánez-Berral France, North Macedonia, and Spain October 2025 iii This chapter is the introduction to the proceedings of the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License, a copy of which can be found at http://crea tivecommons.org/licenses/by/4.0/ . Humánez-Berral, P., Henderson, A., & Kirkova-Naskova, A. (2025). Introduction: Looking back and moving forward for online EPIP proceedings. In A. Kirkova-Naskova, P. Humanez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. iv–xii). Université Grenoble-Alpes. https://doi.org/10.5281/zenodo.17273047 Introduction: Looking back and moving forward for online EPIP proceedings Pedro Humánez-Berral University of Cantabria Alice Henderson Grenoble-Alpes University Anastazija Kirkova-Naskova Ss. Cyril and Methodius University, Skopje These proceedings encompass extended accounts of selected oral presentations and plenaries from the conference English Pronunciation: Issues and Practices – EPIP8, held in 2024 at the University of Cantabria in Santander, Spain. EPIP is an international, bi-annual conference devoted to how English pronunciation is taught and learnt. Its main objective is to bring together and inspire connections between teachers and researchers from all over the world, be they preor in-service teachers, undergraduate or graduate students, early career or internationally renowned researchers. People from all career stages attend EPIP conferences, where they can network in person but also explore scientific, social and pedagogical issues related to English pronunciation. Some conference participants have the luxury of just attending, others invest a great deal of time and energy into preparing (and delivering) a talk – and then some of the latter submit proposals for the proceedings. A labour-intensive, three-stage process lies behind these compiled proceedings: first, conference presenters commit to shaping their work into a written text, which then undergoes a meticulous double-blind peer-review process, and finally the revised text is carefully edited by the co-editors – often involving many back-and-forth exchanges. Some contributions are rejected but, as EPIP prides itself on welcoming researchers into the field, reviewers are encouraged to provide constructive criticism and editors work hard to hone as many contributions as possible, respecting the principles of transparency and scientific rigour. To reach as many readers as possible and as quickly as possible, as editors we chose to make EPIP Proceedings available freely online starting in 2023 after the EPIP7 conference. Similar Humánez-Berral, Henderson & Kirkova-Naskova Looking back, moving forward v to the PSLLT archives1 at Iowa State University, the online EPIP proceedings2 are a reliable place to read innovative work being done now in our field, helping us all to update our knowledge even when we cannot attend conferences. We decided to use an Open Access digital archive and to procure Digital Object Identifiers (DOI) for each text. DOIs are important in the long term for the field, because documents remain accessible at a stable link. They are also crucial in the short term for many of our colleagues who live and work in countries where DOIs are (or are becoming) strategically important to their national institutions. Moreover, we sought to obtain the DOIs from a European institution (Zenodo3), because we support European Union aims of freedom, peace, social justice, scientific progress, and diversity in terms of culture and languages.4 Participating in the international open access and open data movement requires more than just good intentions. We also need an online space to ‘house’ the proceedings, and for this we chose HAL, the open science archive created by France’s ministry for research and higher education. Each time a reader opens or downloads a text via HAL, their actions are crossreferenced with search engines and interconnected with other services, such as ORCID. This is a win-win situation: the on-line profiles of authors are boosted, and so is our field’s representation in citation and abstract databases. HAL provides basic statistics on how readers use the archive, as Figure 1 shows. Figure 1 Downloads for 25 Texts in EPIP7 Proceedings5: August 1 and September 21, 2025 In August 2023 the EPI7 Proceedings went live online, yet the peak of interest came in 2024 – perhaps as colleagues began to prepare for EPIP8 which was held in May 2024. In the runup to EPIP96 in April 2026, at the University of Manchester, downloads have been increasing 1 https://apling.engl.iastate.edu/conferences/pronunciation-in-second-language-learning-and-teachingconference/psllt-archive/ 2 https://hal.univ-grenoble-alpes.fr/EPIP/ 3 https://about.zenodo.org/:“Zenodo is derived from Zenodotus, the first librarian of the Ancient Library of Alexandria and father of the first recorded use of metadata, a landmark in library history.” 4https://european-union.europa.eu/principles-countries-history/principles-and-values/aims-andvalues_en 5 https://hal.science/EPIP7/stat/dashboard. The number 25 includes the editors’ Introduction as well as the file compiling the Introduction and 23 contributions by authors. 6 https://epip9manchester.weebly.com/ Humánez-Berral, Henderson & Kirkova-Naskova Looking back, moving forward vi in autumn 2025 (e.g., up from approximately 1200 to 1800) and it will be interesting to continue tracking this activity via HAL. The EPIP7 proceedings include 23 finalised contributions, whereas EPIP8’s have only 10. To compare the research themes addressed at the two conferences, we thus turned to the books of abstracts: EPIP7 received 37 abstracts (word tokens 13,904: 2,737 word types) and EPIP8 46 abstracts (17,982 word tokens: 3,714 word types). As such, the analysis can reveal topics which were discussed at each conference, regardless of whether or not the presenter then submitted a text for publication in the Proceedings. Appendix 1 displays all the results of comparing the keywords7 as generated by AntConc.8 The keywords were generated by the AntConc software (Anthony, 2022), drawing on the conferences’ abstract booklets as corpora. One corpus was created for each conference, in which each abstract was a separate text, with references and the list of author-provided keywords removed. This helps to show how topics change over time, e.g., gestures and pointing appeared in the EPIP7 abstracts but do not appear for EPIP8. However, the number of occurrences of a term must be combined with its range is, i.e., in how many different abstracts it appears. A large number of occurrences may not necessarily be meaningful, e.g., the word <English> (EPIP7: 124 occurrences) and (EPIP8: 198 occurrences) occurs respectively across 86% and 82% of abstracts, merely confirming that it is a key topic of the conference. For each conference, the following paragraphs will draw on the selection of results presented in Table 1, in order to highlight which terms appeared frequently and in several abstracts, as well as what is missing or rare, what has increased or decreased, appeared or disappeared. A final paragraph focuses on which specific languages, language varieties or contexts were mentioned, referring to Table 2. Regarding terms used the least frequently, <CLIL>, <Foreign Language Anxiety> and <multilingual> do not appear in abstracts until EPIP8, where <foreign> (+ <accent> or + <accentedness>) is used in two abstracts, while <gesture> and <pointing> are lost. In both corpora, <shadowing>, <strategy> and <syllable> hold steady, but are mentioned in only one abstract. Once a term appears in at least 10 abstracts, it begins to constitute roughly a quarter of the total number of a conference’s abstracts, with a few developments worth highlighting. First, in the upward sense, <foreign language> has climbed from appearing in only 3 abstracts to 14 (30%) and <learners> bounds from 54% to 72%. The number of abstracts with the term <EFL> leaps from 27% to 39%, <instruction> rises from 32% to 43%, <feedback> rises from 13% to 22%, <acquisition> appears in 37% of EPIP8 abstracts (up from 19%). Minor increases include <intelligibility>, which moves up to 28% from 22%, while <comprehensibility> reaches 26% from 24%9. Reflecting a key objective of the conference – to bring together researchers and teachers – the following core topics are holding steady in EPIP abstracts, respectively the 7th and 8th conferences: <teaching> (54% and 48%), <learning> (43% and 48%), and <teachers> (32% and 37%). The sole notable inversion concerns <accent>, which has declined from 32% to 24% of abstracts, probably reflecting the notable sociolinguistic presence at the Grenoble EPIP. In relation to the ongoing terminological debate (Dewaele, 2018; Thomas & Osment, 2020), no occurrences of <LX> or <additional language> were found; 51% and 48% of abstracts continue to use the term <native>, respectively for EPIP7 and EPIP8, and the term <second language> appears in roughly 24% abstracts for both. 7 In contrast, for the Introduction to the EPIP7 Proceedings (Henderson & Kirkova-Naskova, 2023), we analysed the frequency of occurrence of author-provided keywords. 8 AntConc software developed by Anthony (2022). 9 See Appendix 1 for some terms occurring in between 3 and 10 abstracts. This chapter is based on the oral presentation given by the author at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons.org/licenses/by/4.0/ . Bryła-Cruz, A. (2025). The role of pronunciation in the assessment of listening skills in the Polish Matura exam. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 1–14). Université Grenoble-Alpes. https://doi.org/10.5281/zenodo.16020252 The role of pronunciation in the assessment of listening skills in the Polish Matura exam Agnieszka Bryła-Cruz Maria Curie-Skłodowska University Abstract In the second language (L2) classroom, pronunciation has been perceived as a natural component of speaking and has traditionally been assessed based on the learner’s speech, and yet pronunciation also plays a role in the assessment of L2 listening skills. While phonological competence is crucial for the activation of bottom-up processes in listening, which in turn are indispensable for the automatisation of word recognition and top-down processing (Buck, 2001; Field, 2008), the role of phonetics in receptive skills tends to be marginalised. The present study investigates the extent to which the role of pronunciation is recognised in listening comprehension tests in the Polish Matura, based on the exam papers from 2023. The study has shown that while the spoken texts used in the tests at the basic and the advanced levels are diversified in terms of length and genre, they hardly differ with respect to phonetic traits. The main difference between the two levels is the length of texts and the rate of speech, which is a bit faster for the advanced level, but still slower than in natural speech. Keywords: L2 pronunciation, pronunciation assessment, listening comprehension tests, scripted texts, spoken texts Bryła-Cruz Pronunciation & assessment of listening skills 2 1. Introduction In the second language (L2) classroom, pronunciation has been treated as a natural component of speaking skills and has typically been assessed in learners’ oral production. However, pronunciation also plays a role in the assessment of L2 listening ability. The role of phonetics in listening comprehension has been confirmed in empirical studies, as in Wagner and Toth (2017), where texts identical in lexis and grammar were understood differently by listeners, when the speaker’s pronunciation was modified. While phonological competence is crucial for the activation of bottom-up processes in listening, which in turn are indispensable for the automatisation of word recognition and top-down processing (Buck, 2001; Field, 2008), the role of phonetics in receptive skills tends to be marginalised. This study investigates the extent to which the role of pronunciation is recognised in listening comprehension tests in the Polish Matura, based on the English exam papers from 2023. The phonetic features in the audio materials used in the basic and advanced version of the exam were analysed in order to establish whether the phonetic difficulty of the spoken texts corresponds to the level of learners’ linguistic competence. 2. Theoretical framework The aim of a test is to assess learners’ language skills beyond the test context. Thus, L2 listening tests should mirror genuine spoken texts in terms of phonology, lexis, grammar, organisation and pragmatics (Buck, 2001; Wagner & Toth, 2017). Yet, many of the spoken texts used in L2 listening tests are not representative of natural speech. Rather than reflecting conversational, spoken language in the real-world, audio materials in listening comprehension tests are usually scripted, i.e., written down, edited, rehearsed and read aloud with an excessively clear enunciation (Wagner, 2013). This contradicts what is postulated by many researchers, namely that “given the inherent differences between oral and written discourse, it is imperative to investigate comprehension using listening passages that approximate natural speech” (Schmidt-Rinehart, 1994, p. 180). The differences between scripted and unscripted texts are manifold and have been extensively discussed in the literature (e.g., Cauldwell, 2014; Chafe, 1982; Gilmore, 2004; Henderson & Cauldwell, 2020; Ockey & Wagner, 2018). For instance, specific speech features impact the overall difficulty which a given text poses to L2 listeners, including: 1) hesitation phenomena (HP) present in unplanned speech; 2) speech rate; 3) phonological modifications characteristic of unplanned spoken texts; and 4) organisational and lexico-grammatical characteristics. The rest of this section discusses the first three categories of features, as the last lies beyond the scope of the present paper. Hesitative phenomena abound in spontaneous discourse and include filled and unfilled pauses, repetitions, false starts, redundancies, and hesitations. These reflect the fact that the spoken text is simultaneously composed and uttered. Studies into HP in L2 perception have rendered somewhat inconsistent results. On the one hand, HP may constitute a major perceptual problem for non-native speakers (Voss, 1979), because they may be confused with unstressed forms or parts of words and may be assigned semantic meaning (Griffiths, 1991; Voss, 1979). On the other hand, HP can help listeners, e.g., providing them extra time to process the message (Vandergrift & Goh, 2012). It has been suggested that such features make the text easier to listen to, by evoking natural oral language to a greater extent than planned speech (Shohamy & Inbar, 1991) It is widely recognised that the rate of speech influences L2 comprehension. The general conclusion from ample empirical research is that spoken texts delivered at a faster speech rate are more difficult for language learners to understand than texts spoken more slowly (Chiu & Bryła-Cruz Pronunciation & assessment of listening skills 3 Chen, 2023; Fujita, 2017; Griffiths, 1990). However, the reality is more nuanced. According to Wang and Narayanan (2007), “speech rate is primarily dependent on two factors: speaking style and the type of speech production (e.g., scripted or spontaneous)” (p. 2190). Typically, scripted texts are enunciated more slowly and more clearly. Bloomfield et al. (2011) note that “faster speech is usually less clear than slower speech although speech rate and auditory clarity are distinct properties” (p. 67). Moreover, while L2 listeners often report difficulty in comprehension due to faster delivery (Goh, 2000), they may in fact conflate problems with affective factors (anxiety). For example, the intrinsic nature of listening in many L2 language classrooms – namely, time pressure and the lack of control over the acoustic signal (Bloomfield et al., 2011) – is cause for anxiety in many learners. Furthermore, variations in speech rate are typical of natural, spontaneous speech; speakers have been found to speed up when delivering less relevant parts of utterances, such as parenthesis, digressions or repair sequences (Fon, 1999; Uhmann, 1992). The rate of speech also increases in chunks typical of spontaneous discourse, pre-fabricated expressions or formulaic sequences (McCarthy, 2010; Wood, 2006). Such clusters of words cause decoding difficulty due to their reduced prominence and they may vary in length from the two-word chunks (e.g., “I mean”) to as many as six (e.g., “do you know what I mean”) (Henderson & Cauldwell, 2020). Most real-world speaking events are characterised by connected speech processes (CSP). Alameen and Levis (2015), propose that CSPs can be divided into six main categories: linking, deletion, insertion, modification (palatalisation/yod-coalescence, assimilation, glottalisation, flapping), reduction, and other processes (lexical combinations, contracted forms). Brown and Kondo-Brown (2006) point out that although connected speech is more common in casual and informal contexts, it is a normal feature of all registers and speaking styles and should not be considered a marker of sloppy or careless speech. Bowen (1975) asserts that in informal speech the use of contractions is almost universal and that reduction is an ever-present feature of oral English. Yet, a speaker may wish to reduce CSPs if they attend consciously to their speech in pursuit of clearer enunciation and increased intelligibility (Ito, 2001). This is especially the case when the speaker reads a text aloud (Chafe, 1982). Against this backdrop it is therefore valid to check whether the two levels of the Matura exam acknowledge the role of phonetics in listening comprehension. This study focuses on test assessment rather than other, less formal types of evaluation, because “tests will often have an effect on classroom teaching” (Buck, 2001, p. 196). In other words, analysing the test provides insight into the teaching process with respect to its aims and resources. In the case of the highstakes Matura exam, the analysis of the listening component is expected to provide some indications as to whether the washback effect is beneficial or not. 3. Research methodology The instrument used in the present study consists of 18 recordings and their transcripts from the basic and the advanced level of the Matura exam from 2023 (see Table 1 for a detailed description of tasks). First, the extracts were analysed with reference to the main differences between scripted and unscripted texts, focusing on those related to pronunciation (HP, phonological modifications, and speech rate). The number of words and syllables in each of the recordings was counted with the use of Text Analyser.1 Speech rate was calculated as syllables per second (s/s), following common practice (Shriberg et al., 2000). Pauses and nonlinguistic utterances are treated as individual units, because “they might be indicators of conceptual planning and their existence might also contribute to rate perception” (Fon, 1999, 1 Online Utility – Free Online Software Utilities https://www.online-utility.org/ Bryła-Cruz Pronunciation & assessment of listening skills 4 p. 663). Phonological modifications (assimilations, elisions, linking devices) and HP (hesitative pauses, false starts, repetitions, redundancies) were traced auditorily and counted manually using both the transcript and the recording. The study explores the following research questions: RQ1: Is the phonetic difficulty of the spoken texts modified to correspond to specific levels of linguistic proficiency? RQ2: Is pronunciation aligned with the type of discourse? With reference to RQ1, it is expected that tests used for assessing listening comprehension skills should vary with respect to pronunciation depending on the test taker’s level. Regarding RQ2, the alignment of pronunciation with the type of discourse means that texts used in L2 listening comprehension tests should mirror spoken texts from the target context for language use. In real life, spoken texts differ in orality, meaning that some of them contain more features of spoken language as opposed to written language (Tannen, 1982), e.g., radio or TV news are scripted and thus less oral than spontaneous interactions or unplanned monologues. The passages used for L2 listening comprehension should be low or high in orality depending on their real-life counterparts. 4. Data analysis Table 1 presents the structure of the listening comprehension tests at both the basic and the advanced level. It shows that both versions of the exam consist of three tasks, which include Table 1 The Structure of the Listening Comprehension Tests Test type Listening section Text type Question type No of times the text is listened to No of questions Basic Task 1 1 dialogue true or false 2 5 Task 2 5 monologues matching a statement to the speaker 2 5 Task 3 2 monologues 1 dialogue multiple-choice 2 6 Advanced Task 1 2 monologues 1 mini interview (a question and an extended answer) multiple-choice 2 6 Task 2 5 monologues matching a statement to the speaker 2 5 Task 3 1 dialogue cloze test 2 4 Bryła-Cruz Pronunciation & assessment of listening skills 5 dialogues and monologues. Comprehension is checked by means of multiple-choice questions and matching a statement to the speaker. Additionally, true or false questions appear in the basic level and a cloze test in the advanced one. All texts are listened to twice. 5. Results The results of the analysis of the spoken texts used in the listening comprehension tests are presented in this section. The information pertaining to HP and the rate of speech is included in Tables 2 and 3, whereas CSPs are discussed separately in the text. In Tables 2 and 3 the last two columns juxtapose the degree of orality in the audio material with the degree of orality that its counterpart in real life is supposed to have. 5.1 Basic level The nine texts used in the basic-level exam vary with respect to length, style, and genre (see Table 2). Three texts could be classified as rather formal and the rest as semi-formal or informal. Irrespective of the type and length, however, all texts are scripted, which is justified only in four cases: two parts of a radio broadcast, a TV announcement, and partly in the tourist guide announcement where some pre-planning/rehearsing would be expected. The remaining texts are either personal narratives or conversations, which would be non-scripted yet specific to and/or typical of the targeted situation for language use. Only one recording (Text 3 in Task 3) contains some features of unplanned texts, such as one hesitative pause (Hmmm… his face is definitely triangular) and two silent pauses (I only need one pair…but hold on, Well…obviously you want the best pair for your face). All speakers enunciate clearly, with a subsequent restriction on the number of reductions which would normally occur in their speech. The most common feature of connected speech is linking /r/ (n = 11), found in examples (1)–(5), and intrusive /r/ (n = 2) in examples (6) and (7). (1) were a teenager /wər ə 'ti:neɪʤə/ (2) for only /fər 'əʊnlɪ/ (3) for as little as /fər əz 'lɪtl əz/ (4) bear at the same shop /'beər ət ðə 'seɪm 'ʃɒp/ (5) fear of people /'fɪər əv 'pi:pəl/ (6) I saw a few people / 'sɔ:r ə 'fju: 'pi:pəl/ (7) drawing /'drɔ:rɪŋ/ Elisions and assimilation are scarce. For instance, there is yod-coalescence in example (8), but it is not consistent, i.e., it does not appear in example (9) when the same speaker says a different phrase that meets the phonological conditions for assimilation. In Task 2, Text 2, there are some deletions and contractions, illustrated in (10) and (11) but not in all instances, as in sentence (12). There is no h-dropping, e.g., (13) and (14), or /t/-deletion as shown in (15)–(17). Plosive deletions were observed in (18)–(20), but not in (21). Similarly, instances of assimilation can be observed in (22) and (23), but not in (24) and (25). (8) would you /wʊʤʊ/ (9) what you need /wɒt jə 'ni:d/ Bryła-Cruz Pronunciation & assessment of listening skills 6 Table 2 Text Types (re Length, Style, Genre) Used at the Basic Level Listening task (parts) Text type Length HP Speech rate Orality Orality in real life Task 1 dialogue (an interview) 2 min. 330 words none 4.2 s/s low (scripted) high (unscripted) Task 2 Text 1 monologue (a personal film recommendation) 32 s 92 words none 3.1 s/s low (scripted) high (unscripted) Text 2 monologue (a tourist guide explanation) 31 s 93 words none 3.9 s/s low (scripted) rather low (preplanned) Text 3 monologue (a TV broadcast) 39 s 107 words none 3.7 s/s low (scripted) low (preplanned) Text 4 monologue (a radio programme) 38 s 102 words none 3.8 s/s low (scripted) rather low (semiscripted) Text 5 monologue (a personal narrative about a past event) 42 s 119 words none 3.7 s/s low (scripted) high (unscripted) Task 3 Text 1 monologue (a personal narrative about a past event) 46 s 139 words none 4 s/s low (scripted) high (unscripted) Text 2 monologue (a part of a TV or radio programme) 50 s 125 words none 3.5 s/s low (scripted) low (scripted) Text 3 dialogue (a conversation between a customer and a shop assistant) 1 min 19 s 214 words 1 hesit ative pause, 2 silent pauses 3.4 s/s low (scripted) high (unscripted) (10) why don’t you /waɪ dəʊn jə/ (11) you won’t regret it /jʊ wəʊn rɪ'gret ɪt/ (12) It is perfect for up to four guests /ɪt ɪz 'pɜ:fɪkt fər ʌp tə fɔ: 'gests/ (13) with him /wɪð hɪm/ (14) from him /frɒm hɪm/ (15) next week /nekst 'wi:k/ Bryła-Cruz Pronunciation & assessment of listening skills 7 (16) just five days /ʤʌst 'faɪv 'deɪz/ (17) just like Lego /ʤʌst 'laɪk 'legəʊ/ (18) just loved /ʤʌs 'lʌvd/ (19) second year /'sekən 'jɪə/ (20) first job /'fɜ:s 'ʤɒb/ (21) most of my time /məʊst əv maɪ 'taɪm/ (22) did you work /dɪʤʊ 'wɜ:k/ (23) did you ever find /dɪʤʊ evə 'faɪnd/ (24) about your work /ə'baʊt jɔ: 'wɜ:k/ (25) what are you doing /'wɒt jə 'du:ɪŋ/ Results show that there are a few contractions, such as I’d like to ask you, it’s great, I’m working, it’s in fact, it’s beautiful, but they are not always used consistently, for instance, because it is, which had just been hit, and it had been built. The speech rate ranges from 3.1–4.2 s/s, with the slowest tempo in the narrative and the fastest in the dialogue. While a faster rate for interaction is in line with what takes place in real life, it is still considerably slower than in casual conversations: 5–6 s/s (Deese, 1984) or 6.19 s/s (Pellegrino et al., 2011), and even slower than in radio broadcasts where the rate of speech is on average around 5.43 s/s (Kendall, 2009). Notably, due to the pre-fabricated character of the passages, there are no digressions, afterthoughts, repairs, or reformulations, which would introduce a variation in the speech rate. 5.2 Advanced level The texts used in the advanced version of the exam are diversified with respect to length and genre. The register of the texts is semi-formal or informal and all nine texts are scripted. This seems completely unjustified, because in real life none of these texts would be written. With respect to pronunciation, it is careful and clear with a fairly limited number of CSPs. Needless to say, some reductions are used as in English, irrespective of whether a text is scripted or non-scripted. These include contractions and weak forms of function words, but the latter have no consonant dropped, as shown in examples (26)–(29). Instances of assimilation are few but present, as in examples (30) and (31). The most frequent phenomenon is linking /r/, evident in the examples (32)–(36), and there is one case of intrusive /r/ in example (37). (26) must be white /mʌst 'bi: 'waɪt/ (27) life has revolved /'laɪf həz rɪ'vɒlvd/ (28) where I can purchase /'weər aɪ kən 'pɜ:ʧəs/ (29) I say that my friend /aɪ 'seɪ ðət maɪ 'frend/ (30) what could go wrong /'wɒt kəg 'gəʊ 'rɒŋ/ (31) does your /dʌʒjɔ:/ (32) producer of /prə'dʒu:sər əv/ (33) there any /ðeər 'enɪ/ (34) are a perfect /ɑ:r ə 'pɜ:fɪkt/ (35) for a turkey /fər ə 'tɜ:kɪ/ (36) or inside /ɔ:r ɪn'saɪd/ (37) pizza in the fridge /'pi:tsər ɪn ðə 'frɪdʒ/ Bryła-Cruz Pronunciation & assessment of listening skills 8 Table 3 Text Types (re Length, Style, Genre) Used at the Advanced Level Listening task (parts) Text type Length HP Speech rate Orality Orality in real life Task 1 Text 1 monologue (an anecdote) 57 s 173 words none 4.2 s/s low (scripted) high (unscripted) Text 2 dialogue (a radio interview - a question and an extended answer) 44 s 129 words none 4.4 s/s low (scripted) high (unscripted) Text 3 monologue (a personal narrative of a past event) 2 min. 2 s 385 words none 4.3 s/s low (scripted) high (unscripted) Task 2 Text 1 monologue (personal narrative relating to the speaker’s work experience) 43 s 141 words none 4.6 s/s low (scripted) high (unscripted) Text 2 monologue (personal narrative relating to the speaker’s work experience) 37 s 108 words none 4.5 s/s low (scripted) high (unscripted) Text 3 monologue (personal narrative relating to the speaker’s work experience) 31 s 106 words none 5.4 s/s low (scripted) high (unscripted) Text 4 monologue (personal narrative relating to the speaker’s work experience) 38 s 118 words none 4.3 s/s low (scripted) high (unscripted) Text 5 monologue (personal narrative relating to the speaker’s work experience) 34 s 95 words none 4.2 s/s low (scripted) high (unscripted) Task 3 dialogue (a TV interview with a food producer) 2 min. 10 s 365 words none 4.3 s/s low (scripted) rather high (semiscripted) Bryła-Cruz Pronunciation & assessment of listening skills 9 Contractions are used, but inconsistently within the same speakers. For instance, they are present in the examples we’ve all, you wouldn’t, I’m talking, it’s usually shown, can’t wait, they’re always, that’s another, aren’t, wouldn’t last, I won’t watch, I’ve been working, I wasn’t sure, and I’ll, but absent in they are a perfect substitute, you would get, potatoes have been freshly baked, and as it is about. The rate of speech ranges between 4.2–4.6 s/s. The slowest rate of speech appears in a narrative (anecdote; Task 2, Text 1) and the fastest also in a narrative (Task 1). The dialogues are not characterised by an increased rate of speech compared to the monologues, which is not typical in everyday speech. The overall rate of speech in all texts is a bit slower than in corresponding spontaneous texts in real life where it would be around 5.04 s/s (Grosz & Hirschberg, 1992). Because there are no phenomena reflecting conceptual planning, the rate of speech is devoid of any variation. 6. Discussion and key findings A summary of the findings is presented in Table 4. In both the basic and the advanced version of the exam, the same pair of speakers is employed. The texts are diversified with respect to genre, yet none of the exam versions exploits a casual conversation. A more varied selection of monologues is found in the basic version – TV and radio announcements, personal narratives and an extract delivered by a tourist guide. In the advanced version the monologues are of two types – a personal narrative and an anecdote. The basic version includes two texts which would also be scripted in real life, whereas they are not present in the advanced version. This may have been a deliberate choice on the part of the test designers, but it has no practical consequence, as even in the advanced version all texts are scripted and read out. Paradoxically, therefore, the texts in the basic version, not in the advanced one, better represent their real-life counterparts. Regarding RQ1, the phonetic difficulty of the spoken texts is not modified to correspond to levels of linguistic proficiency. This is partly due to the fact that both the basic and advanced version of the exam include scripted texts, which largely determines their pronunciation features. The main difference between the two versions of the exam is duration and the rate of speech. While the mean length of shorter passages is roughly the same for the two levels (39 sec. vs. 40 sec.), the longer extracts range from 1:19–2:00 minutes in the basic level and 2:02– 2:10 in the advanced level. Since a longer text usually means more information to process, the test designers may have deliberately introduced this difference to increase the difficulty of the advanced version. The number of phonological modifications is similar in both versions. There are contractions, weak forms and certain processes which appear in English irrespective of the type of text (written or spoken), but they are pronounced more clearly than in natural speech. In the basic version there are three instances of assimilation, specifically yod-coalescence, and five instances of elision; 11 cases of linking /r/ and two cases of intrusive /r/. The advanced version includes two instances of assimilation and two cases of linking /r/. The rate of speech in the two versions differs as well. Overall, the rate of speech in the exam texts is slower than naturally occurring monologues and dialogues. Both the range of speech and the mean values are lower for the basic version than the advanced version, respectively: 3.1–4 s/s (m = 3.6 s/s) and 4.2.–4.6 s/s (m = 4.3 s/s). As a faster speech rate usually requires more automatic processing, increasing it for the advanced level increases task difficulty. Furthermore, natural unplanned discourse is characterised by speech rate variation, yet variation is absent from the recordings; they are delivered at a steady rate of speech, without the normal digressions, afterthoughts, false starts, and redundancies which would occur at an accelerated tempo in real life. Bryła-Cruz Pronunciation & assessment of listening skills 10 Table 4 Comparison of the Main Features of Audio Materials Used in the Basic and Advanced Version of the Matura Exam Feature Basic Advanced Number of speakers 2 (male & female) 2 (male & female) Accent SSBE SSBE Type of texts scripted scripted Type of genre varied (monologues & dialogues) varied (monologues & dialogues) Rate of speech Range 3.1–4 s/s 4.2–4.6 s/s Mean 3.6 s/s 4.3 s/s Length in short extracts Range 31–50 s 31–57 s Mean 39 s 40 s Length in long extracts Range 1min. 19 s–2:00 2 min. 2 s–2:10 Mean 1min. 39 s 2 min. 6 s Hesitative pauses 1 filled pause (hmmm) 2 silent pauses none Repetitions False starts Reformulations Unfinished sentences none none Turn-openers 2 instances of well 1 instance of well Connected speech phenomena other than weak forms 3 instances of assimilation (yod-coalescence) 5 instances of elision 11 instances of linking /r/ 2 instances of intrusive /r/ 2 instances of assimilation 24 instances of linking /r/ Apart from duration and the rate of speech, the two versions do not differ considerably with respect to pronunciation. In other words, the phonetic competence required to perform the listening comprehension tasks appears not to have been taken into consideration. The test designers are not primarily focused on testing phonetic competence even though listening competence consists of, above all, dealing with the speech signal. The advanced version of the Matura exam is characterised by more advanced lexis and more complex grammatical structures, but in terms of pronunciation, it remains largely at the same level as the basic one. This is congruent with Buck’s (2001) observation that “unfortunately, many students progress Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 17 importance of exposing learners to diverse English varieties. For example, Miao et al. (2024) investigated the effects of integrating Global English (GE) varieties into classroom instruction and the impact on EFL learners’ listening comprehension and pronunciation. Their study compared an intervention group exposed to different accented English varieties with a control group that heard only American English. Data was collected through TOEFL-like listening exam questions and a speaking task, using a pre/post-test design. The results showed that listening comprehension of GE varieties of the intervention group improved, but only short term, and their pronunciation remained unaffected. Despite the lack of sustained long-term effects, the findings suggest that exposure to different English varieties does not negatively affect learners’ comprehension. Further research is needed on EFL learners’ perspectives on native connected speech to better inform both research and teaching practice. To that end, the current study examines Standard Scottish English (SSE). This variety was chosen because it remains underrepresented in course materials, where Standard Southern British English (SSBE) and General American (GA) continue to dominate as teaching models. Previous research has explored attitudes toward one variety of SSE, Glasgow Standard English, in terms of speaker competence and social attractiveness (McKenzie, 2008). The current study, however, focuses on the perceived naturalness of SSE, with particular attention to connected speech processes. 2.1 Research questions The literature reviewed above highlights the importance of exposing L2 learners to various language varieties, both native and non-native. Achieving intelligibility always involves a twoway process. Nevertheless, in NNS–NS interactions, CSPs place an additional cognitive demand on NNS speakers, as authentic speech can deviate from their expectations of spoken English (Alameen & Levis, 2015). With this in mind, the current study investigates the exposure to a less commonly represented native variety, Scottish English. It examines Macedonian EFL learners’ perceptions of the presence or absence of CSPs in native speech, by asking them which of two recordings sounds more natural, and then to explain their choice. Naturalness is not defined for participants but rather it is hoped that their explanations will mention CSPs. Thus, the present study attempts to answer the following research questions: RQ1: How do Macedonian EFL learners evaluate the naturalness of native speech in relation to: a) speech rate, and b) connected speech processes? RQ2: How do Macedonian EFL learners recognise and identify specific connected speech processes (vowel reduction, elision, intrusive sounds) in native speech? While Macedonian EFL learners may initially find speech with a less familiar accent more challenging, we assume that they will generally perceive both recordings as being easy to understand. However, we hypothesise that they will experience greater difficulty identifying specific connected speech processes, particularly in faster-paced speech. 3. Research methodology 3.1 Participants The participants in this study were 50 Macedonian EFL learners (Mage = 20.8, age range = 20– 25). There were 9 male and 40 female participants, with one participant identifying as nonbinary. All participants were second-year English majors at the University of Ss. Cyril and Methodius in Skopje. They were enrolled in the English Phonetics and Phonology course and Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 18 had completed one semester of instruction but had not yet studied connected speech processes. All had learned English through formal education, and 50% reported having basic knowledge of an additional foreign language. The class is obligatory but participation in the study was voluntary. A native female speaker of Scottish Standard English served as a speech model. She holds a Master of Research in Linguistics and works as a university tutor, teaching English pronunciation classes to international students. 3.2 Procedure The study consisted of two stages. First, the native speaker recorded a 295-word passage (adapted from The Guardian) twice. For the first recording (R1), the speaker was instructed to read the passage at a slower, more deliberate pace, aiming to produce more words in full forms. For the second recording (R2), the speaker read the passage at a faster pace, resulting in speech with more reduced forms. Both recordings, along with selected excerpts, were embedded in a questionnaire. In the second stage, the participants completed the questionnaire in a controlled environment using desktop computers and headphones and were supervised by a researcher.1 The task took approximately 40 minutes. 3.3 Data collection and analysis The main data collection instrument was a questionnaire comprising two parts: Part 1 gathered demographic information, whereas Part 2 focused on responses to the recordings (see Appendix). It comprised 14 questions: 8 open-ended, 2 Likert scale, and 4 multiple-choice questions, along with two recordings and five excerpts. The participants listened to Recordings 1 and 2 twice before answering the questions, while the excerpts could be played up to five times. These excerpts targeted three connected speech features: vowel reduction, elision, and intrusive sounds. Table 1 presents the selected features and transcriptions of their occurrence in a sentence. These five excerpts were chosen due to their representativeness and the prominence of the tested feature, thus facilitating recognition and distinction of their presence or absence in speech. In Excerpts 3, 4, and 5, the targeted feature was present in both recordings, further emphasising the fact that connected speech occurs naturally, irrespective of speech rate. Responses from the completed questionnaires were manually entered into Excel. The openended questions were analysed for the presence of recurring themes, while quantitative data was processed in SPSS 25 (IBM Corp., 2017). A Wilcoxon signed-rank test was used, appropriate for ranked data such as Likert-scale items. 4. Results The results, based on the analysis of both the qualitative and quantitative data from the questionnaire, are presented in two subsections. Quotes from individual participants are referenced with the participants’ codes P01–P50, where P stands for participant followed by the participant number. 1 The researcher is the author of this paper. Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 19 Table 1 Targeted Connected Speech Features in Excerpts Excerp t Target feature Phrase Transcription 1 Vowel reduction in grammatical words This was a problem. R1 /ðɪs wɒz ə ˈprɒbləm/ R2 /ðɪs wəz ə ˈprɒbləm/ 2 Vowel reduction in grammatical words …including politicians, presented as sports fans. R1 /ɪnˈkluːdɪŋ ˌpɒlɪˈtɪʃənz prɪˈzentɪd æz ˈspɔːts ˈfænz/ R2 /ɪnˈkluːdɪŋ ˌpɒlɪˈtɪʃənz prɪˈzentɪd əz ˈspɔːts ˈfænz/ 3 Elision But it can’t have remained more elusive in real life. R1 /kænt æv rɪˈmeɪnd/ R2 /kænt əv rɪˈmeɪnd/ 4 Intrusive /w/ …to its rapid layered rhythms… R1 & R2 /tu (w) ɪts ˈræpɪd ˈleɪəd ˈrɪðəmz/ 5 Intrusive /j/ … why all public figures, including politicians… R1 & R2 /waɪ(j) ɔːl ˈpʌblɪk ˈfɪɡəz ɪnˈkluːd ɪŋ ˌpɒlɪˈtɪʃənz/ 4.1 Recordings 1 and 2 The first set of questions in Part 2 of the questionnaire elicited participants’ views regarding their perceived ease of understanding of the speech in R1 and R2, the challenges faced while listening, and the differences between R1 and R2. For self-perceived ease of understanding, the participants rated each recording on a scale of 1–5 (1 = very difficult, 5 = very easy). The average score for R1 was M = 4.34, compared to M = 4.50 for R2, indicating that the participants did not find either recording particularly difficult to understand. To establish whether the difference in understanding was statistically significant, a Wilcoxon test was used. The results showed the difference between R1 and R2 was not statistically significant (p = .106). Regarding the challenges faced, the participants mentioned difficulties related to accent for R1 (n = 24, 48%), as illustrated in the following responses: “some words were difficult to understand because she was pronouncing them in an odd way” (P06), “it took me some time to adjust, but it wasn't difficult” (P12), or ”incomprehensible due to unfamiliar pronunciation” (P43). The majority found R2 less challenging than R1, although Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 20 a few (n = 10, 20%) reported difficulties with understanding R2 due to the heavy accent, fast pace of speech, pronunciation of some words, or the need for more processing time. When asked which recording sounded more natural, some participants (n = 8, 16%) selected R1, the majority (n = 33, 66%) selected R2, while the remaining (n = 9, 18%) said that both sounded natural to them. The main reasons for R1 not sounding natural included – again, in participants’ own words – its slow or robotic speech, and the perception that the speaker was non-native. In contrast, R2 was perceived as more natural due to its natural flow of speech, lack of pauses, the speaker sounding native despite the accent, and the overall tone being more conversational. The participants were also asked to identify any obvious differences between R1 and R2, choosing from a list of options; one served as a distractor (option d). Table 2 lists the main differences and the number of responses for each option. Table 2 Differences between Recording 1 and Recording 2 and the Number of Selected Answers Possible differences between Recording 1 and Recording 2 No of responses a) The speaker speaks faster. 38 b) Some sounds have been left out. 17 c) Some extra sounds have been added. 15 d) Some words have been left out. 5 e) Some words are pronounced shorter. 36 f) Sometimes there are no pauses between two words. 36 The results show that most participants noticed that the speaker in R2 spoke faster, resulting in some words being pronounced shorter or without pauses. Fewer participants noticed the omission or addition of sounds, while only five participants selected the distractor option that some words had been left out. 4.2 Excerpts The second set of questions elicited participants’ views regarding specific connected speech features present or absent in five different excerpts. Excerpts 1 and 2 targeted vowel reductions in grammatical words. In both excerpts, R1 contained a strong form of the grammatical words (i.e., was /wɒz/ and as /æz/), while R2 contained a weak form (i.e., was /wəz/ and as /əz/). When asked whether the grammatical word in Excerpt 1 was the same or different in both recordings, few participants (n = 8, 16%) said it was the same, while the majority (n = 42, 84%) noted that it was different. The participants who identified a difference cited factors such as Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 21 emphasis on the word, the presence of a strong versus a weak vowel, and the presence or absence of pauses before and after the word, as evident in the following remarks: “In R2 the tone is more conversational and the focus is not on <was>” (P04); “In recording one <was> is pronounced harder and it sound[s] more important, while in recording two <was> is pronounced more softly” (P14). 2 However, some responses referred to dialectal differences or provided reasoning not directly related to the target feature, such as: “The first one has an accent to it making it sound like <wos> and the second one is clearer and sounds like the [A]merican <was>” (P06), or “In the first recording, the <s> is pronounced as /z/, while in the second it’s /s/” (P42). The responses for Excerpt 2, targeting the word as, were more varied. Half of the participants (n = 25, 50%) reported that the pronunciation was the same in both recordings, many (n = 22, 44%) indicated that it was different, and few (n = 3, 6%) provided no answer. Those who perceived a difference attributed it to the presence of a schwa sound and hearing only the /s/ sound in R2, or noticing clearer pauses before and after the word in R1 compared to R2. Excerpt 3 involved instances of elision, specifically the omission of /h/, in both recordings. The only difference between R1 and R2 was the absence and presence of a weak form in the grammatical word have, respectively. When asked if the words in the verb phrase were fully pronounced, responses varied significantly. Twenty participants (40%) indicated that the words were either fully pronounced in both recordings, or not fully pronounced but their answers did not mention omission of /h/. Examples included comments such as the words sounding like one word in R1 but like two separate words in R2, or that words were fully pronounced only in R2. The remaining participants (n = 30, 60%) correctly noticed the elision of /h/ correctly, in at least one of the recordings, or identified the contracted form of have in the verb phrase, as illustrated in the following comments: “In Recording 2, the words <can’t> and <have> are connected with the neutral vowel in their pronunciation, resulting in the shortening of the word <have>, which is not the case in Recording 1” (P03) and “All of the words are pronounced in Recording 1[;] however, in Recording 2, the /h/ in <have> is left out and the word is connected to the previous one, <can’t> ” (P38). Concerning self-perceived ease of understanding of the phrase in Excerpt 3, the majority (n = 28, 56%) reported no difficulty. A smaller proportion found the speech easier in R1 than in R2 (n = 5, 10%), while others found it easier in R2 than in R1 (n = 6, 12%). The remaining participants (n = 10, 20%) said that it was not very easy, and they either had to rely on the context or listen to the phrase multiple times. One participant mentioned that familiarity with the accent helped. One participant did not answer the question. The last two excerpts, 4 and 5, focused on the intrusive sounds /w/ and /j/, respectively. For Excerpt 4, only one participant correctly identified the intrusive sound, stating: “It sounds like there is a /w/ sound in between them” (P11). The rest either did not notice any intrusive sound, or provided irrelevant information, such as hearing an extra /t/ or a prolonged /s/ sound. No participant identified the intrusive /j/ sound in Excerpt 5. However, one student reported difficulty understanding the phrase, stating: “It was partially understandable, since the first word of the phrase was incomprehensible.” (P34), suggesting that the presence of the intrusive sound may have affected comprehension. Regarding their self-perceived ease of understanding the phrase in Excerpt 4, the majority did not experience any difficulty (n = 29, 58%), while the rest (n = 20, 40%) noted that they had to listen to it multiple times, or that pauses between words aided comprehension. One participant did not answer the question. 2 In all original comments by participants, their inconsistent punctuation has been replaced: < > is used for graphemes and / / for phonemes, to maintain a text that is easy to read. Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 22 5. Discussion The first research question intended to investigate the participants’ perceptions of the naturalness of native speech and the difficulties they encountered understanding it. The fastpaced speech, which featured more instances of connected speech features, was generally perceived as more natural due to its absence of pauses, the rapid speech rate, and the speaker sounding native. These results suggest that fast speech is associated with fluency, despite the participants reporting difficulties in perceiving specific CSPs. This supports the view that CSPs contribute to the perceived regularity of English rhythm (Alameen & Levis, 2015). In terms of self-perceived ease of understanding, neither recording was perceived as especially difficult, and the difference was not statistically significant. This may be due to the participants’ overall familiarity with English through formal instruction and exposure to various native accents in the media. However, further research is needed to systematically investigate how such prior exposure affects the ability to process native speech and whether self-reported ease of understanding correlates with actual understanding, as authentic speech often remains challenging (Wong et al., 2021). When comparing the two recordings, the participants were able to identify more obvious distinctions, such as variations in speech pace and the absence of pauses between words. Nevertheless, more subtle CSPs, such as elision and intrusive sounds, proved more difficult. This was reflected in the responses to the speech excerpts explored in the second research question. For example, while more than half of the participants identified /h/ elision in Excerpt 3, this recognition was inconsistent across both recordings and mainly observed in R2. In contrast, the intrusive sounds /w/ and /j/ in Excerpts 4 and 5, respectively, were scarcely noticed. Only one participant identified the intrusive /w/, and another appeared to misinterpret the initial part of the phrase in Excerpt 5, potentially due to the presence of intrusive /j/. These findings align with previous research, reinforcing the notion that CSPs are difficult for EFL learners to perceive and decode accurately. Participants’ perception of vowel reduction in Excerpts 1 and 2 showed that they found it easier to detect the reduced sound in /wəz/ in R2 of Excerpt 1 than in /əz/ in Excerpt 2. The majority of participants (84%) identified the schwa sound in R2 of Excerpt 1, whereas nearly half as many (44%) noticed it or reported hearing only a /s/ sound in R2 of Excerpt 2. While these findings are based on a limited number of examples and should not be generalised, they are consistent with the results of a production study of vowel reduction among Macedonian EFL learners (Duckinoska, 2021). This study found that the L2 learners more readily produce the weak form of grammatical words containing the strong vowel /ɒ/ (e.g., was) than the weak form of words with the strong vowel /æ/ (e.g., am, can, shall, has). Nevertheless, further research is needed to establish a clearer link between learners’ perception and production of weak forms. Interestingly, the participants frequently referenced terms such as accent, native speaker, or specific dialect. Accent was often cited as a source of difficulty, while the delivery of the native speaker in R2 was perceived as more natural. These findings suggest that the participants’ perception of speech naturalness was primarily influenced by two factors: 1) the speaker being perceived as native, and 2) the absence of a noticeable accent. This insight carries significant pedagogical implications: learners should be made aware of the diversity of English accents – both native and non-native – and listen to a wide range of speech models. As exposure to Global Englishes does not hinder comprehension (Miao et al., 2024), such practice can broaden learners’ awareness of what constitutes intelligible and natural English. The results also suggest that certain CSPs are harder to detect than others. Thus, explicit instruction on how CSPs aid fluency and rhythm could be beneficial, as increased awareness improves decoding of rapid speech (Kennedy & Blanchet, 2014). For instance, teachers could Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 23 contrast sentences with and without intrusive sounds, highlighting how their presence helps maintain fluency. They could also design activities where students listen specifically for intrusive sounds in authentic speech and mark them in written transcripts. Moreover, since vowel reduction was easier to detect in was than in as, EFL instruction should explicitly highlight how vowel reduction varies across different grammatical words. Finally, the fact that only 60% of participants correctly noticed /h/ elision in have suggests that learners may not always be aware of common elisions in connected speech. Classroom activities should include targeted listening tasks with common examples of elision, followed by pronunciation practice to reinforce recognition. 6. Conclusion and limitations This study investigated Macedonian EFL learners’ perceptions of native speech with the presence and absence of connected speech features. Although the participants reported relatively high ease of understanding, many noted the need to listen to the excerpts multiple times, or rely on context, underscoring the fact that authentic speech remains challenging. Simultaneously, the perception of Recording 2 as more natural suggests that connected speech features contribute to the natural rhythm and perceived fluency of English. The current study has several limitations. The results are based on specific samples taken from a specific context; different excerpts from the recordings, or different contexts may have elicited different responses. In addition, using different speech models (e.g., Australian English) may influence participants’ responses. Further research could explore these variables and measure the effects of exposure to connected speech on perception and production. Despite these limitations, the findings contribute to a better understanding of how EFL learners perceive authentic spoken English and emphasise the importance of drawing learners’ attention to connected speech features as a step towards improved listening comprehension. References Ahmadian, M., & Matour, R. (2014). The effect of explicit instruction of connected speech features on Iranian EFL learners’ listening comprehension skills. International Journal of Applied Linguistics and English Literature, 3(2), 227–236. http://dx.doi.org/10.7575/aiac.ijalel.v.3n.2p.227 Alameen, G., and Levis, J. M. (2015). Connected speech. In M. Reed & J. M. Levis (Eds.), The handbook of English pronunciation (pp. 159–174). Wiley Blackwell. http://doi.org/10.1002/9781118346952.ch9 Brown, J.D., & Hilferty, A. (2006). The effectiveness of teaching reduced forms for listening comprehension. In J. D. Brown & K. Kondo-Brown, (Eds.), Perspectives on teaching connected speech to second language speakers (pp. 51–58). University of Hawai’i, National Foreign Language Resource Center. https://nflrc.hawaii.edu/PDFs/MG01TOC.pdf Derwing, T. M., & Munro, M. J. (1997). Accent, intelligibility, and comprehensibility: Evidence from four L1s. Studies in Second Language Acquisition, 19, 1–16. https://doi.org/10.1017/S0272263197001010 Duckinoska, I. (2021). Vowel reduction of English grammatical words by Macedonian EFL learners. In A. Kirkova-Naskova, A. Henderson, & J., Fouz-González (Eds.), English pronunciation instruction: Research-based insights (pp. 279–302). John Benjamins Publishing Company. https://doi.org/10.1075/aals.19.12duc IBM Corp. (2017). IBM SPSS Statistics for Windows, Version 25.0 [Computer software]. IBM Corp. Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 24 https://www.ibm.com/support/pages/downloading-ibm-spss-statistics-25 Ito, Y. (2006). The significance of reduced forms in L2 pedagogy. In J. D. Brown & K. KondoBrown, (Eds.), Perspectives on teaching connected speech to second language speakers (pp. 17–25). University of Hawai‘i, National Foreign Language Resource Center. https://nflrc.hawaii.edu/PDFs/MG01TOC.pdf Kang, O., Thomson, R., & Moran, M. (2020). Which features of accent affect understanding? Exploring the intelligibility threshold of diverse accent varieties. Applied Linguistics, 41(4), 453–480. https://doi.org/10.1093/applin/amy053 Kennedy, S., & Blanchet, J. (2014). Language awareness and perception of connected speech in second language. Language Awareness, 23, 92–106. https://doi.org/10.1080/09658416.2013.863904 McKenzie, R. M. (2008). Social factors and non‐native attitudes towards varieties of spoken English: A Japanese case study. International Journal of Applied Linguistics, 18(1), 63–88. https://doi.org/10.1111/j.1473-4192.2008.00179.x Miao, Y., Kang, O., & Meng, X. (2024). Incorporating global Englishes varieties into EFL classrooms: Development of listening comprehension and pronunciation. TESOL Quarterly. https://doi.org/10.1002/tesq.3328 Poesová, K, & Weingartová, L. (2018). Character of vowel reduction in Czech English. In J. Volín & R. Skarnitzl (Eds.), The pronunciation of English by speakers of other languages (pp. 96–116). Cambridge Scholars. Rahimi, M. & Chalak, A. (2017). The effect of connected speech teaching on listening comprehension of Iranian EFL learners. Journal of Applied Linguistics and Language Research, 4(8), 280–291. https://www.jallr.com/index.php/JALLR/article/view/722 Suzukida, Y., & Saito, K. (2019). Which segmental features matter for successful L2 comprehensibility? Revisiting and generalizing the pedagogical value of the functional load principle. Language Teaching Research, 25, 431–450. https://doi.org/10.1177%2F1362168819858246 Wong, S. W., Leung, V. W., Tsui, J. K., Dealey, J., & Cheung, A. (2021). Chinese ESL learners' perceptual errors of English connected speech: Insights into listening comprehension. System, 98, 1–10. https://doi.org/10.1016/j.system.2021.102480 Appendix Participants’ questionnaire Part 1 Please complete the questionnaire below. The information will only be accessible to and used by the researcher for research purposes and all the data will be anonymised. Name and surname Age Gender Male Female Other: What nationality are you? What is your first language/mother tongue? Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 25 Prior to starting university, did you study English outside school? a) private language courses Time period: ______ b) private lessons Time period: ______ Have you stayed in an English-speaking country? a) Yes Which one: Time period: b) No How would you rate your satisfaction with your English pronunciation on a scale from 1 to 5? 1 = not satisfied at all 5 = extremely satisfied 1 2 3 4 5 Do you think that formal training in English phonetics and phonology helps you improve your pronunciation? a) It is definitely helpful. b) It is helpful in some ways. c) It is not helpful at all. Do you speak any other foreign language(s)? If yes, please state which one(s). Part 2 In this part, you will listen to several recordings. We are interested in your thoughts about the recordings. You will not be tested on any phonetic/phonological aspect. Please follow the instructions as to what recording you need to listen to and how many times. You need to double click on the icon for the audio to start. Step 1: Listen to Recording 1 only once and then answer the questions below. Recording 1.mp3 How easy was it to understand recording 1? 1 = very difficult 5 = very easy 1 2 3 4 5 Did you find any pronunciation aspect particularly challenging to understand? Please describe it in your own words. Step 2: Listen to Recording 2 only once and then answer the questions below. Recording 2.mp3 Duckinoska-Mihajlovska Perceptions of naturalness, difficulty & speech rate 26 How easy was it to understand recording 2? 1 = very difficult 5 = very easy 1 2 3 4 5 Did you find any pronunciation aspect particularly challenging to understand? Please describe it in your own words. Step 3: Listen to Recording 1 and Recording 2 one more time, if necessary, and then compare Recording 1 and Recording 2. Please read the questions first before listening. Which of the two recordings sounds more natural to you? a) Recording 1 Please explain why: b) Recording 2 Please explain why: c) They both sound the same to me. Did you notice any obvious differences between the two recordings? Choose as many as applicable. a) The speaker speaks faster. b) Some sounds have been left out. c) Some extra sounds have been added. d) Some words have been left out. e) Some words are pronounced shorter. f) Sometimes there are no pauses between two words. Step 4: You will listen to excerpts from both recordings and answer the questions below. You can listen to each excerpt up to five times. Listen to Excerpt 1. Pay attention to the way ‘was’ is pronounced in both recordings. Excerpt 1.wav Is ‘was’ pronounced the same in both recordings? a) Yes b) No If not, what is different? Listen to Excerpt 2. Pay attention to the way ‘as’ is pronounced in both recordings. Excerpt 2.wav Is ‘as’ pronounced the same in both recordings? a) Yes b) No If not, what is different? Listen to Excerpt 3. Pay attention to the verb phrase in the clause in both recordings. Gyurka Pronunciation teaching practices in Hungarian schools 33 pronunciation specifically. These open-ended results are seemingly in line with the most commonly used pronunciation teaching technique the teachers reportedly used (see Table 2). Table 2 Descriptive Statistics of Pronunciation Teaching Techniques Pronunciation teaching technique M SD using listen-and-repeat 3.77 1.40 teaching an exercise from the coursebook 3.39 1.34 using pair work 3.06 1.50 marking pronunciation in drawing (e.g., underlining the stressed syllable) 2.91 1.47 using the minimal pair technique 2.89 1.40 teaching how to read the IPA symbols 2.39 1.33 using a game to practice pronunciation 2.30 1.23 giving metalinguistic explanations; pointing out the differences between English and Hungarian 2.20 1.12 teaching pronunciation by relying on touch (e.g., touching the neck to feel voiced consonants) 2.13 1.23 using videos 2.06 1.18 using online games for practice 2.03 1.25 teaching how to recognise IPA symbols 1.97 1.20 using Jazz Chants 1.89 1.10 using movement (e.g., clapping) 1.77 1.18 having the students record themselves 1.69 0.87 The participants stated that they most frequently employed the “listen and repeat” technique for pronunciation teaching, once every month. This suggests that they only used ‘listen and repeat’ when correcting errors, because vocabulary items are introduced more often than just every month. The second most commonly used technique was using exercises from a coursebook every two to three months. Interestingly, respondents reported very limited use of IPA symbols in the classroom. The results showed that IPA symbols were used for demonstration purposes less frequently than every two to three months, whereas teaching how to write and use IPA symbols was even rarer. In terms of pronunciation features (see Table 3), silent letters and the dental fricatives were taught the most often; both features were reported to come up around every month. Among the least frequently taught features were the stress-timed rhythm, the weak pronunciation of grammar words, and aspiration. These features were rarely addressed, less frequently than two to three months. Gyurka Pronunciation teaching practices in Hungarian schools 34 Table 3 Descriptive Statistics of Teaching Pronunciation Features Pronunciation feature M SD silent letters 4.05 1.03 the dental fricatives 3.87 1.24 word stress 3.84 1.35 the differences between SSBE (Standard Southern British English) and GA (General American) 3.55 1.25 the -ed suffix 3.53 1.28 the -s suffix 3.52 1.17 the difference between /w/ and /v/ 3.50 1.40 the /ŋ/ sound 3.25 1.32 vowel reduction in unstressed syllables 3.23 1.49 syllabic consonants 3.06 1.44 the difference between /əʊ/ and /ɔː/ 3.03 1.28 aspiration 2.83 1.28 the differences between native accents of English (other than SSBE and GA) 2.50 1.14 weak forms of grammar words 2.39 1.39 stressed-time rhythm 2.36 1.24 The overall results of these two scales suggest that Hungarian secondary school teachers include pronunciation features every two to three months (M=3.23; SD=0.91), and that they use some teaching techniques only occasionally (M=2.30; SD=0.71). 4.3 Variables influencing Hungarian teachers’ pronunciation teaching practices (RQ3) First, to determine which variables might impact on teachers’ pronunciation teaching practices, a series of regression analyses were run using the frequency of pronunciation teaching techniques, and then the frequency of teaching pronunciation features as a dependent variable. The 16 items measuring beliefs about pronunciation teaching were initially entered into the model as independent variables. Items that did not exert a statistically significant influence on the dependent variables were subsequently removed one at a time, with the analyses being rerun after each removal. The results (see Table 4) show that the explanatory power of the first model is somewhat stronger as it accounted for a greater percentage of variance in the case of the pronunciation features (R2 = 0.54) Notwithstanding, the second model also accounts for a considerable proportion of the variance (R2 = 0.49). The belief that pronunciation teaching is important at a secondary school level had a positive, albeit weak impact on the frequency of pronunciation features addressed in the lessons. Interestingly, the belief that pronunciation teaching can improve the learners’ confidence had a moderately positive effect on the various techniques employed during the lessons, but this did not have an impact on the pronunciation features Gyurka Pronunciation teaching practices in Hungarian schools 35 taught. The similarity between the two models is that the belief that pronunciation can improve on its own had a weak, but negative influence on how frequently teachers taught pronunciation features and used various pronunciation teaching techniques. Table 4 Linear Regression Predicting Frequency of Pronunciation Features and Teaching Techniques R2 Importance Confidence Improve Pronunciation features 0.54 0.36 n.s. -0.29 Pronunciation teaching techniques 0.49 n.s. 0.40 -0.23 Notes. (p<0.05). Importance = belief that teaching pronunciation is important at a secondary school level; Confidence = belief that pronunciation teaching improves the learners’ confidence; Improve = belief that the learners’ pronunciation can improve on its own. 4.4 The effect of training on pronunciation teaching practices (RQ4) To determine whether the training teachers received related to pronunciation teaching methodology played a significant role in terms of their practices, independent-samples t-tests were run on the construct of pronunciation teaching techniques and pronunciation features. As there was no significant difference between teachers with and without prior education in relation to either the frequency of teaching pronunciation features or the frequency of using pronunciation teaching techniques, the latter construct was examined for possible latent dimensions to see whether the different dimensions would show a significant difference. The 15 items of the scale were subjected to factor analysis. Table 5 Result of the Factor Analysis Pronunciation teaching techniques Factor loadings 1 2 reading IPA symbols .837 writing IPA symbols .888 teaching rules (metalinguistic explanation) .480 the minimal pair technique .692 listen and repeat .467 drawing .841 pair work .700 Notes. Extraction method: Maximum likelihood; Rotation method: Oblimin. Loadings smaller than 0.45 were removed. Gyurka Pronunciation teaching practices in Hungarian schools 36 Prior to this, the suitability of the data for factor analysis was verified.1 The results yielded a two-factor solution, which can be seen in Table 5. Factor 1 presents a dimension that subsumes teaching the IPA symbols, while factor 2 contains teaching techniques other than the use of IPA in the classroom. The independent-samples t-tests were conducted to compare teachers with prior education in pronunciation teaching methodology (n = 28) and teachers without prior education (n = 36). As the results in Table 6 indicate, there was a significant difference between the groups in terms of the frequency of pronunciation teaching when it comes to teaching IPA in the classroom, but not in the case of the other pronunciation teaching techniques. Table 6 The Results of the Independent-Samples t-tests M SD t p frequency of using other techniques with education 3.20 0.99 1.69 0.10 without education 2.78 0.97 frequency of teaching IPA with education 2.57 1.11 2.30 0.03 without education 1.90 0.84 5. Discussion The results revealed that Hungarian secondary school teachers indicate that they did not address pronunciation often during lessons, and whenever they did devote time to pronunciation teaching, it was to teach vocabulary items. Dealing with pronunciation only when teaching a new vocabulary item is not sufficient on its own; pronunciation should be considered as worthy of being taught, just like other language aspects, e.g., grammar. In terms of teaching techniques, “listen and repeat” and coursebook exercises were selected as among the most frequently used techniques, showing the same tendency as in other geographical contexts (see Jafari et al., 2021; Szyszka 2016). Furthermore, the extensive use of listen-and-repeat” is not surprising, as a considerable number of the participants considered vocabulary teaching or mere corrective feedback to constitute explicit pronunciation teaching. Respondents selected silent letters as the most commonly addressed pronunciation feature. This may be linked to pronunciation being included in the lessons whenever the pronunciation of a vocabulary item seemed problematic. While it is possible to teach the appearance of silent letters as a system, it is highly unlikely that teachers devoted time to demonstrate the pattern to their students. Based on the qualitative answers, teachers probably addressed this feature whenever it caused an issue related to a single vocabulary item. The other pronunciation feature with the highest frequency of teaching was the dental fricatives. This finding leads to an interesting discrepancy between the theoretical underpinnings of pronunciation teaching and 1 The Kaiser-Meyer-Olkin value was 0.73, which exceeded the 0.60 recommended minimal value (George & Mallery, 2019); the Bartlett’s Test of Sphericity proved to be statistically significant (Ⲭ2 = 138.70, p < 0.001), which further demonstrated that the data was sufficient for the analysis. Then, factor analysis was performed using maximal likelihood and Direct oblimin rotation as it was assumed that the results would correlate. Gyurka Pronunciation teaching practices in Hungarian schools 37 sactual practices, which indicates that the participants’ teaching practices might not have been informed by pronunciation teaching theory. As it has been established that pronouncing the dental fricatives does not influence intelligibility (Jenkins, 2000), it might be unnecessary to focus on teaching them. It may nevertheless be valuable to identify and target specific pronunciation features that present persistent challenges for Hungarian learners (e.g., word stress) and whose improvement could further enhance overall intelligibility. In relation to the different variables which might influence teachers’ pronunciation teaching practices, it was found that the belief that pronunciation improves on its own negatively impacts the frequency of teaching. Along these lines, the belief that pronunciation should be taught in secondary education and that it improves the students’ confidence had a positive impact on the frequency of inclusion. These findings undeniably highlight the need for well-designed pronunciation teaching methodology courses in teacher education, since only training can provide insight to teachers that can equip them with the necessary knowledge to inform their practices (Borg, 2011). The important role of teacher education in influencing teachers’ practices was also evident as there was a difference between teachers who had a course on pronunciation teaching and teachers who did not. As could be seen above (Table 6), there was a significant difference in terms of teaching IPA symbols, which was expected, since using the IPA symbols to adequately provide metalinguistic explanations might be difficult for untrained teachers. All of these results point to the same tendency observed in contexts outside Hungary: teacher education plays a pivotal role in later pronunciation inclusion (see Henderson et al., 2012; Shah et al., 2017). 6. Implications The findings of this study underscore several important pedagogical implications for pronunciation instruction in Hungarian secondary education. First, the limited nature of pronunciation teaching (typically tied to vocabulary instruction) suggests a need to move beyond this approach. Pronunciation should be taught systematically in the same manner as other language skills, and not merely addressed when individual lexical items are problematic or when introducing new vocabulary items. Second, the results of this study point to the importance of teacher education in pronunciation teaching methodology. More than half of the teachers indicated that they had not received formal instruction in this area. Therefore, there is a pressing need for pronunciation teaching methodology courses in pre-service and in-service teacher education in Hungary. If this came to pass, and teachers had adequate theoretical knowledge of why and how to integrate pronunciation effectively into their classroom practice, their beliefs might change and they might promote pronunciation in their classrooms just like any other language skill. Acknowledgements The research has been supported by the EKÖP-24 University Excellence Scholarship Program of the Ministry for Culture and Innovation from the Source of the National Research, Development and Innovation Fund. This study was also conducted in the frame of project no. PPKE-BTK-KUT-23-2, supported by the Faculty of Humanities and Social Sciences of Pázmány Péter Catholic University. Gyurka Pronunciation teaching practices in Hungarian schools 38 References Bai, B., & Yuan, R. (2019). EFL teachers’ beliefs and practices about pronunciation teaching. ELT Journal, 73(2), 134–143. https://doi.org/10.1093/elt/ccy040 Baker, A. (2011). ESL teachers and pronunciation pedagogy: Exploring the development of teachers’ cognitions and classroom practices. In J. Levis & K. LeVelle (Eds.), Proceedings of the 2nd pronunciation in second language learning and teaching conference (pp. 82–94). Iowa State University. https://apling.engl.iastate.edu/wpcontent/uploads/sites/221/2015/05/PSLLT_2nd_Proceedings_2010.pdf Baker, A. (2014). Exploring teachers’ knowledge of second language pronunciation techniques: Teacher cognitions, observed classroom practices, and student perceptions. TESOL Quarterly, 48(1), 136–163. https://doi.org/10.1002/tesq.99 Baran-Łucarz, M. (2017). FL pronunciation anxiety and motivation: Results of a mixedmethod study. In E. Piechurska-Kuciel, E. Szymańska-Czaplak & M. Szyszka (Eds.), At the Crossroads: Challenges of Foreign Language Learning (pp. 107–133). Springer Cham. https://doi.org/10.1007/978-3-319-55155-5_7 Borg, S. (2003). Teacher cognition in language teaching: A review of research on what language teachers think, know, believe, and do. Language Teaching, 36(2), 81–109. https://doi.org/10.1017/S0261444803001903 Borg, S. (2005). Teacher cognition on language teaching. In K. Johnson (Ed.), Expertise in second language learning and teaching (pp. 190–209). Palgrave Macmillan. https://doi.org/10.1057/9780230523470_10 Borg, S. (2011). The impact of in-service teacher education on language teachers’ beliefs. System, 39(3), 370–380. https://doi.org/10.1016/j.system.2011.07.009 Buss, L. (2016). Beliefs and practices of Brazilian EFL teachers regarding pronunciation. Language Teaching Research, 20(5), 619–637. https://doi.org/10.1177/1362168815574145 Dörnyei Z. (2007). Research methods in applied linguistics. Oxford University Press. Dörnyei, Z., & Csizér, K. (2012). How to design and analyze surveys in SLA research? In A. Mackey & S. Gass (Eds.), Research methods in second language acquisition: A practical guide (pp. 74–94). Wiley-Blackwell. https://doi.org/10.1002/9781444347340.ch5 Georgiou, G. P. (2019). EFL teachers’ cognitions about pronunciation in Cyprus. Journal of Multilingual and Multicultural Development, 40(6), 538–550. https://doi.org/10.1080/01434632.2018.1539090 Gilbert, J. (1995). Pronunciation practice as an aid to listening comprehension. In D. J. Mendelsohn & J. Rubin (Eds.), A guide for the teaching of second language listening (pp. 97–102). Dominie Press. Gyurka, N. (2022). Pronunciation activities in the EFL classroom: Important or useless? In K. Heller & I. Steiner (Eds.), Under the umbrella of applied linguistics (pp. 200–214). ELTE BTK Alkalmazott Nyelvészeti és Fonetikai Tanszék. [original title: Kiejtésfeladatok a középiskolai angol nyelvórán: fontos vagy felesleges? In K. Heller & I. Steiner (Eds.), Az alkalmazott nyelvészet esernyője alatt (pp. 200–214). ELTE BTK Alkalmazott Nyelvészeti és Fonetikai Tanszék.] https://www.eltereader.hu/media/2022/10/alknyelv-202122_jav.pdf Gyurka, N., & Piukovics, Á. (2023). Tailoring international activities for Hungarian learners of English. Research in Language, 21(3), 313–332. https://doi.org/10.18778/17317533.21.3.06 Henderson, A., Frost, D., Tergujeff, E., Kautzsch, A., Murphy, D., Kirakova-Naskova, A., Waniek-Klimczak, E., Levey, D., Cunningham, U., & Curnick, L. (2012). The English pronunciation teaching in Europe survey: Selected results. Research in Language, 10(1), 5– 27. https://doi.org/10.2478/v10015-011-0047-4 Gyurka Pronunciation teaching practices in Hungarian schools 39 Jafari, S., Karimi, M. R., & Jafari, S. (2021). Beliefs and practices of EFL instructors in teaching pronunciation. Journal for Language and Foreign Language Learning, 10(2), 147– 166. https://doi.org/10.21580/vjv11i110812 Jenkins, J. (2000). The phonology of English as an international language: New models, new norms, new goals. Oxford University Press. Kelly, L. G. (1969). 25 centuries of language teaching. Newbury House Publishers. Levis, J. M. (2005). Changing contexts and shifting paradigms in pronunciation teaching. TESOL Quarterly, 39(3), 369–377. https://doi.org/10.2307/3588485 Levis, J. M. (2018). Intelligibility, oral communication, and the teaching of pronunciation. Cambridge University Press. https://doi.org/10.1017/9781108241564 Murphy, J. M., & Baker, A. A. (2015). History of ESL pronunciation teaching. In M. Reed & J. M. Levis (Eds.), The handbook of English pronunciation (pp. 36–65). John Wiley & Sons. https://doi.org/10.1002/9781118346952.ch3 Nagle, C., Sachs, R., & Zárate-Sández, G. (2018). Exploring the intersection between teachers’ beliefs and research findings in pronunciation instruction. The Modern Language Journal, 102(3), 512–532. https://doi.org/10.1111/modl.12493 Nunan, D., & Miller, L. (Eds.). (1995). New ways in teaching listening (2nd ed.). TESOL. Pennington, M. C. (2021). Teaching pronunciation: The state of the art 2021. RELC Journal, 52(1), 3–21. https://doi.org/10.1177/00336882211002283 Piukovics, Á. (2016). The changing role of the IPA in the EFL classroom. In A. Á, Reményi, Cs. Sárdi, & Zs. Tóth (Eds.), Perspectives in contemporary Hungarian applied linguistics (pp. 155–164). Tinta Könyvkiadó. [original title: The changing role of the IPA in the EFL classroom. In A. Á, Reményi, Cs. Sárdi, & Zs. Tóth (Eds.), Távlatok a mai magyar alkalmazott nyelvészetben (pp. 155–164). Tinta Könyvkiadó.] Quoc, T. X., Thanh, V. Q., Dang, T. D. M., Mai, N. D. N., & Nguyen, P. N. K. (2021). Teachers’ perspectives and practices in teaching English pronunciation at Menglish Center. International Journal of TESOL & Education, 1(2), 158–175. https://ijte.org/index.php/journal/article/view/62 Shah, S. S. A., Othman, J., & Senom, F. (2017). The pronunciation component in ESL lessons: Teachers’ beliefs and practices. Indonesian Journal of Applied Linguistics, 6(2), 193–203. https://doi.org/10.17509/ijal.v6i2.4844 Sifakis, N. C., & Sougari, A. (2005). Pronunciation issues and EIL pedagogy in the periphery: A survey of Greek state school teachers’ beliefs. TESOL Quarterly, 39(3), 467–488. https://doi.org/10.2307/3588490 Szyszka, M. (2016). English pronunciation teaching at different educational levels: Insights into teachers’ perceptions and actions. Research in Language, 14(2), 165–180. https://doi.org/10.1515/rela-2016-0007 Zheng, H. (2015). Teacher beliefs as a complex system: English language teachers in China. Springer. https://doi.org/10.1007/978-3-319-23009-2 Gyurka Pronunciation teaching practices in Hungarian schools 40 Appendix Study questionnaire Part 1. The frequency of pronunciation teaching in general How often do you teach pronunciation in your lessons? 1 – never 2 – less frequently than every 2-3 months 3 – every 2-3 months 4 – weekly 5 – almost every single lesson Can you explain in a few sentences why? ________________________________________________________________________ Part 2. Beliefs about pronunciation teaching Below you will find statements about pronunciation teaching. Please indicate your level of agreement on a scale of 1 to 5. 1 – completely disagree; 5 – completely agree 1. Teaching pronunciation is important at secondary school level. 2. Pronunciation should be taught to beginners. 3. The language learner’s age influences the effectiveness of pronunciation teaching. 4. English spoken with a strong Hungarian accent leads to discrimination abroad. 5. The aim of pronunciation teaching at secondary school level is to eliminate the Hungarian accent. 6. The aim of pronunciation teaching at secondary school level is to make the learner’s accent intelligible. 7. The best teacher for teaching pronunciation is a native speaker. 8. When teaching pronunciation, teachers should point out the differences between the sound systems of English and Hungarian. 9. Pronunciation is best taught by explaining rules. 10. Students’ pronunciation improves spontaneously with language use. 11. Spending at least six months in the target language country helps improve the learner’s pronunciation. 12. Explicit pronunciation instruction helps improve pronunciation. 13. My students like it when I correct their pronunciation. 14. I think my students would like to change their pronunciation. 15. Teaching pronunciation gives students confidence when using the language. 16. Knowing the pronunciation of words improves the learners’ listening comprehension. Gyurka Pronunciation teaching practices in Hungarian schools 41 Part 3. Attitudes related to pronunciation teaching The following are statements related to your experience when teaching pronunciation. Please also indicate here how much you agree with them on a scale of 1 to 5. 1 – completely disagree; 5 – completely agree 1. I enjoy teaching pronunciation. 2. I am confident when teaching pronunciation. 3. Teaching pronunciation is difficult. 4. It is exhausting to teach pronunciation. Part 4. Pronunciation features How often do you teach the following features? Consider all your groups regardless of level and age. 1 – never, 2 – less frequently than every 2-3 months, 3 – every 2-3 months, 4 – monthly, 5 – weekly 1. “th” sounds, e.g., thanks, mother 2. the difference between wet and vet 3. the difference between boat and bought 4. in words ending in -ng, there is no pronounced ‘ɡ’, e.g., sing 5. the three-way pronunciation of the -ed suffix, e.g., loo‘kt’ – clea‘nd’ – nee‘id’ 6. the three-way pronunciation of the -s suffix, e.g., ca‘tsz’ – do‘ɡz’ – bus‘iz’2 7. word stress, e.g., poLICE, parTIcular 8. silent letters, e.g., debt, pneumonia, gnome, write 9. pronouncing a short ‘h’ after ‘p’, ‘t’, ‘k’, e.g., ‘ph’ie, ‘th’ap, ‘kh’ite 10. the pronunciation of -ism at the end of words, e.g., tourism pronounced ‘turizöm’ and not ‘turizmö’3 11. the pronunciation of grammar words with an ö-like sound, e.g., the sententence ‘There was a boy’ sounds like ‘dö vöz ö’ 12. unstressed syllables, e.g., pencil, lemon sound like ‘penszöl’, ‘lemön’, respectively 13. rhythm, e.g., pronouncing the sentence ‘Bob died’ takes the same amount of time as ‘Bob could have died’ 14. the difference between British and American accents 15. differences between other accents, e.g. Canadian, Australian, North of England, etc. Part 5. Pronunciation teaching techniques How often do you use the following techniques? Consider all your groups regardless of level and age. 1 – never, 2 – less frequently than every 2-3 months, 3 – every 2-3 months, 4 – monthly, 5 – weekly 2 Hungarian orthography was used to represent the endings (/s/, /z/, and /ɪz/) for easier understanding. 3 Hungarians often equate the sound /ø/ (represented with <ö> in spelling) with the schwa, which is why this item was formulated as is. Gyurka Pronunciation teaching practices in Hungarian schools 42 1. Completing the pronunciation tasks in the coursebook. 2. Teaching how to recognise phonetic symbols. 3. Learning how to write phonetic symbols. 4. Students listen to the teacher or a recording and then repeat the utterance/word, trying to imitate it as accurately as possible. 5. Practising the difference between words which only differ in one sound, e.g., bed – bad, bought – boat, wet – vet, etc. 6. Marking pronunciation in a text, e.g., marking stress, drawing intonation arrows, etc. 7. Teaching pronunciation by feeling, e.g., touching the neck with a hand while pronouncing voiced and voiceless sounds 8. Teaching pronunciation through movement, e.g., clapping, head bobbing, etc. 9. Using pronunciation teaching videos in class. 10. Using classroom games focusing on pronunciation. 11. Teaching pronunciation using online games. 12. Practising in pairs or groups with dialogue or role-play, with particular attention to pronunciation. 13. Explaining English pronunciation rules and comparing them with Hungarian. 14. Using Jazz Chants for rhythmic practice. 15. Learners record their pronunciation and listen to it. Part 6. The frequency of pronunciation teaching in general (again) I would like to ask you the very first question again. Please answer how often you usually teach pronunciation and if you think your opinion has changed, please describe why. 1 – never 2 – less frequently than every 2-3 months 3 – every 2-3 months 4 – weekly 5 – almost every single lesson Can you explain in a few sentences why? ________________________________________________________________________ About the author Noémi Gyurka is Assistant Lecturer at Pázmány Péter Catholic University (PPCU). She holds an MA in Teaching English as a Foreign Language and Teaching Hungarian Language and Literature as a First Language (PPCU, 2022). She is doing her PhD in Language Pedagogy and English Applied Linguistics at Eötvös Loránd University. Her research interests are pronunciation teaching and pronunciation integration in the Hungarian English as a foreign language classroom. She has been working at PPCU since 2022, as a full-time lecturer since September 2024. Email: [email protected] Martin-Rubió Fluency in 2 tasks: Learner performance & views 49 the story: the flatter lines represent those who paused more often, and the steeper lines represent those who read entire lines without pausing. 4.2 Fluency and prosody of story-reading The impact of fluency and prosody in story-reading can be illustrated by comparing two students with opposite approaches to the task. S12 and S22 had very similar MSRs for T1 (11.18 and 11.22 respectively), but their approach to T2 was different in many respects. In terms of the number of syllables, S22 produces the canonical 116 syllables, whereas S12 produced 119. S12 mistakenly said be aware (which adds one syllable) rather than beware, adds ah (one more syllable) before the second I’ll drop it that is not in the text, and does not use the contraction I’m the third time Stick Man tells us who he is, i.e., I AM STICK MAN, which has one more syllable than I’M STICK MAN. More detail for this comparison is provided in Table 2, where numbers in brackets indicate the rounded lengths (in milliseconds) of the pauses. S12 pauses 13 more times than S22, with a total length of pausing of 12.5 seconds for S12 and just 4.2 seconds for S22. S22 reads the first two lines without pausing, whereas S12 pauses three times. The same happens when the dog starts talking; S22 delivers the first two lines by the dog without pauses, whereas S12 pauses three times. Table 2 The Two Renderings of the First Four Pages by S12 and S22 S12 S22 Stick Man (0.5) lives in the family tree (0.8) with his Stick Lady Love (0.5) and their stick children three (1.2) one day (0.4) he wakes early and goes for a jog (0.8) Stick Man (0.4) oh Stick Man (1.0) be aware of the dog (1.0) a stick (0.9) barks the dog an excellent stick (0.6) the right kind of stick (0.3) for my favourite trick (0.7) I’ll fetch it (0.5) and drop it (0.4) and fetch it (0.3) and then ah I’ll drop it and fetch it and drop it again (0.8) I’m not a stick (0.3) why can’t you see I’m Stick Man (0.4) I’m Stick Man (0.3) I am Stick Man that’s me (0.4) and I want to go home to the family tree Stick Man lives in the family tree with his Stick Lady Love and their stick children three (0.7) one day he wakes early (0.3) and goes for a jog (0.4) Stick Man oh Stick Man beware of the dog (0.8) a stick barks the dog an excellent stick the right kind of stick for my favourite trick (0.3) I'll fetch it and drop it and fetch it and then (0.3) I'll drop it and fetch it and drop it again (1.1) I'm not a stick why can't you see I'm Stick Man I'm Stick Man I'm Stick Man that's me (0.3) and I want to go home to the family tree Note. Numbers in brackets indicate the rounded lengths (in milliseconds) of pauses. S12 and S22 also had to decide whether to individualise the voices of the narrator, the dog and Stick Man, and how to convey the emotions of these characters as represented by exclamation marks. S12 decided to vary the voices. For example, the voice of her dog had a Martin-Rubió Fluency in 2 tasks: Learner performance & views 50 different quality (lower pitch) from the other two voices, and she used a very wide pitch range (see Figure 2, blue line right). In contrast, S22 did not change voice quality and her pitch range remained unaltered (see Figure 2, blue line left). Figure 2 S22’s Dog (left) and S12’s Dog (right) Regarding productions for Stick Man, the contrast is even greater as shown in Figure 3. S12 adopted a low-pitched voice and decided to translate the italics and upper case into pitch variations. In the first rendering, S12 placed the tonic syllable on Stick and used a falling tone. For the text in italics she moved the stress to the word Man and used a fall-rise tone. For the upper-case version, she disregarded the contraction in the text and generated an extra syllable, and also used a fall-rise tone. Finally, in the last tone unit, she pronounced the tonic syllable me with a rise-fall tone (even though she wrote fall-rise (F/R) in her Praat image). S22’s version (see Figure 4) is a single tone unit in which the three times Stick Man shouts I’m Stick Man are read in exactly the same way, delivering all the syllables with no pauses for 0.25s or more (as shown in Table 2), and with the final falling tone on the tonic syllable me. Pitch variation is also minimal, compared to the productions of S12. The Praat images presented here are from students’ work, as they were taught how to use Praat to analyse the intonation of their own productions, using Tench’s (2011) tripartite notion of intonation as tonality, tonicity, and tone. The analysis for Activity 2 was submitted as a text document with Praat images on the page. The analysis for Activity 3 was submitted in a video, in which they shared their screen, and used a PowerPoint presentation including Praat images and corresponding audio files. Overall, they did well to master these basics of Praat, despite inevitable errors (as in Figure 3).3 3 See Herment (2018, 2023) on the difficulty of using Praat with students. Martin-Rubió Fluency in 2 tasks: Learner performance & views 51 Figure 3 S12’s Stick Man “I’m Stick Man.” “ I’m Stick Man. ” “I’M STICK MAN.” “that’s me.” Figure 4 S22’s Stick Man Martin-Rubió Fluency in 2 tasks: Learner performance & views 52 4.3 Student impressions of the tasks In order to understand how students experienced the tasks, a short questionnaire was used before the end of the course. It was answered by 20 of the 22 students analysed for this study. To the question: “Did you make voices for the dog, the girl? If not, did you consider doing it and decided not to in the end? Or you just thought there was no need?”, the answers were evenly split. Five tried to differentiate the characters somehow but considered what they had done did not amount to making voices as illustrated in the following comment: “I tried to change a little bit the intonation but I didn’t make voices” (S15). Ten students answered they had indeed made voices, and five gave a negative answer. Of these five responses, two students felt it was not necessary, S22 being one of them. They gave the following reasons: “it was necessary but I decided not to do it” (S02), “didn't think to do it” (S14), and “because I did not like how they were turning out” (S19). 5. Discussion This study presents the results of a teaching approach implemented in an undergraduate course. The approach focuses on using the differences between two tasks to make students reflect on the concepts of fluency, pronunciation accuracy, and prosody. The story-reading instructions allowed the students the freedom to choose whether they wanted to make voices or not, because not everyone would be comfortable making them; this freedom was suggested by a colleague who teaches optional, practical theatre workshops which these students can take for credit within the degree programme. Given the results, including the answers to the questionnaire, I would like to replicate this pilot study, encouraging students more explicitly to try and differentiate the voices more and to play with the possibilities at hand. Students’ questionnaire responses suggest that a good option may be to show students what people do when reading stories and to reflect on the impact these decisions have. For instance, some students make voices that sound very different from the narrator’s, with extreme pitch variation, others make enough minor variations so that the listener can identify a new character has appeared, and some do nothing at all. In the first task, it is just a matter of providing information in a clear way, and speech rate and pausing should correspond to their proficiency level, whereas the second task is an entirely different endeavour and students should be aware of this. Those who do nothing at all sound monotonous and would probably not capture children’s attention, and thus fail to fulfil the task’s entertainment function. In addition, being able to widen one’s pitch production and to realise that the intensity, pitch, and tonicity variations can impact on the audience’s perception is of great importance for improving learners’ overall speech performance. In other words, this story reading task could represent a golden opportunity for learners to experience success in playing with their voice, in changing, for example, their prosody and voice quality. This study presents insights derived from a pedagogical approach to teaching about phonemes, fluency, connected speech, and intonation. On the one hand, it confirmed my intuition that fluency measures would change in the two tasks because, apart from the student’s level, their approach to story-reading would translate into a different way of delivering syllables. On the other hand, it made me realise that some students simply see no need to vary pitch and differentiate characters, which might mean more work is needed to make them understand the relevance of pitch variation and character differentiation. Martin-Rubió Fluency in 2 tasks: Learner performance & views 53 6. Conclusion This pilot study presents an approach to teaching phonology in which students are asked to record themselves spontaneously describing images and reading a children’s book. For the spontaneous picture description task, the student’s proficiency level was a very important factor, something of which they were well aware. However, for the story-reading task, some participants seem to be unaware of what they need to do pronunciation-wise to make their speech more appealing. Providing learners with more opportunities to play with and grasp the importance of voice and pitch variations might raise their phonological awareness and further improve their L2 pronunciation. References Boersma, P., & Weenink, D. (2022). Praat: Doing phonetics by computer (Version 6.2.14) [Computer program]. www.praat.org Derwing, T. M., Rossiter, M. J., Munro, M.J. & Thomson, R.I. (2004). Second language fluency: Judgments on different tasks. Language Learning, 54(4), 655–679. https://doi.org/10.1111/j.14679922.2004.00282.x Donaldson, J. (2008). Stick Man. Alison Green Books. Ejzenberg, R. (2000). The juggling act of oral fluency: A psycho-sociolinguistic metaphor. In H. Riggenbach (Ed.), Perspectives on fluency (pp. 287–313). University of Michigan Press. Foster, R., & Skehan, P. (1996). The influence of planning and task type on second language performance. Studies in Second Language Acquisition, 18(3), 299–323. https://doi.org/10.1017/S0272263100015047 Ginther, A., Dimova, S., & Yang, R. (2010). Conceptual and empirical relationships between temporal measures of fluency and oral English proficiency with implications for automated scoring. Language Testing, 27(3), 379–399. https://doi.org/10.1177/0265532210364407 Herment, S. (2023). From research to teaching: The case of English rising contours. In A. Henderson & A. Kirkova-Naskova (Eds.), Proceedings of the 7th International Conference on English Pronunciation: Issues and Practices (pp. 83–96). Université Grenoble-Alpes. https://hal.science/hal-04168829 Herment, S. (2018). Apprentissage et enseignement de la prosodie : l’importance de la visualization. Revue française de linguistique appliquée, 23(1), 73-88. Read. K. (2014). Clues cue the smooze: Rhyme. pausing. and prediction help children learn new words from storybooks. Frontiers in Psychology, 5, 149. https://doi.org/10.3389/fpsyg.2014.00149 Starkweather, C. W. (1987). Fluency and stuttering. Prentice Hall. Tench, P. (2011). Transcribing the sound of English: A phonetics workbook for words and discourse. Cambridge University Press. Wennerstrom, A. (2000). The role of intonation in second language fluency. In H. Riggenbach (Ed.), Perspectives on fluency (pp. 102–127). University of Michigan Press. Martin-Rubió Fluency in 2 tasks: Learner performance & views 54 Appendix The two images for the spontaneous description task Image A Image B About the author Xavier Martin-Rubió holds bachelor degrees in English Philology and in Audiovisual Communication from Universitat de Lleida, and a master’s degree in European Studies from Maastricht Universiteit. He defended his doctoral dissertation in 2011. He worked as an English teacher in a high school, two official language academies and two private universities before moving back to Universitat de Lleida. He started working full-time at UdL in December 2014, where he is now an associate professor. He has been involved in several competitive projects as a member of CLA, and after a gap of several years, he is publishing again on the internationalisation of and EMI in higher education, beliefs and emotions in additional language learning, and fluency and accuracy measurements across tasks. Email: xavier.martin[email protected] This chapter is based on the plenary talk given by the author at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons .org/licenses/by/4.0/ . Pennington, M. (2025). A critical examination of research on the teaching of pronunciation in a second language. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 55–70). Université Grenoble-Alpes. https://doi.org/10.5281/zenodo.16695648 A critical examination of research on the teaching of pronunciation in a second language Martha C. Pennington University of Reading Birkbeck University of London Abstract Although there is a growing research base showing approaches to teaching pronunciation that have proven effective, both the types of approaches that have been systematically studied and the contexts in which those approaches have been applied are as yet relatively limited. There is still much to explore regarding what approaches work in what contexts and under what circumstances. I point out limitations in the research base and areas that have not yet received much attention from researchers. I then review recent studies that are moving pronunciation teaching and research away from its traditional narrow focus and artificial contexts to a broader conception and more authentic contexts. I encourage other researchers to replicate these studies and to adapt their methodologies for different contexts in order to systematically investigate these practices. I also encourage classroom teachers to draw on those methodologies to design action research for trying out new approaches in their classrooms that have proven effective in other contexts. Keywords: action research, applied research, language teaching, pronunciation Pennington Critical examination, L2 pronunciation teaching research 56 1. Introduction The increased interest in pronunciation teaching in the current era (Levis, 2019) has spurred a great expansion in the research base exploring effective practices (Levis, 2021a, 2022; Saito & Plonsky, 2019). The accumulation of research findings has made it possible to make generalisations about the positive effects of some approaches that have been researched, such as perception-oriented approaches (Sakai & Moorman, 2018), especially High Variability Phonetic Training – HVPT (Thomson, 2018), and approaches that address suprasegmental aspects of pronunciation (Lee et al, 2015). Notwithstanding these advances in our knowledge of approaches that work and the body of research supporting the value of teaching pronunciation (Derwing & Munro, 2015; Lee et al., 2015), the research base is still limited, and there is a vast territory for exploration and systematic study of the effectiveness of specific practices (Pennington, 2021). In what follows, I review the present state of research in the field and offer suggestions of areas that are ripe for further investigation by researchers as well as by teachers in their own classrooms, recommending action research as a way for the latter group to put research into practice. 2. Limitations of the existing body of pronunciation research A critical examination of the existing body of pronunciation research reveals a number of limitations, including:  restricted focus of the research;  relatively weak research design;  insufficient attention to methodological factors affecting research outcomes;  lack of systematic replication and comparison across different conditions;  limited attention to authentic language and language variety; and  limited attention to authentic communication. An examination of each of these limitations serves to delineate areas in need of attention in future research. 2.1 Restricted focus of the research The bulk of pronunciation teaching research centres on remediating segmental errors in the speech of students who have been learning English for some time (Lee et al., 2015; Levis, 2021b; Saito & Plonsky, 2019; Thomson & Derwing, 2015). A focus on segmentals tends to mean a focus on error and accuracy that the field is aiming to move away from and towards a focus on intelligibility (Levis, 2018). The focus on segmentals has meant comparatively less attention paid to suprasegmentals, discourse, and social attributes of pronunciation. Microlevel features of pronunciation have been attended to at the expense of macro-level features related to intonation, rhythm, stress, and the accompanying substantial modifications of phonemes that occur through linking and coarticulation during running speech. Because of the focus of research on relatively advanced learners (Saito & Plonsky, 2019), methodologies for teaching pronunciation to beginners have hardly been explored. In addition, the published research has centred mainly on learners of English, with limited attention paid to pronunciation teaching for learners of other languages (Levis, 2021b; Pennington, 2021). More research is needed on pronunciation in context, on younger and early-stage learners, and on learners of languages other than English. Pennington Critical examination, L2 pronunciation teaching research 57 2.2 Relatively weak research design Lee et al.’s (2015) discussion of their meta-analytical review points to weak research design and methodology as a problem in pronunciation research. Many studies showing positive findings need to be followed up by further studies exploring the same approaches with stronger designs in terms of methodology and research standards. Hardison (in press) lists all the features that characterise successful perception training for pronunciation suggesting the complexity and number of elements required to constitute a strong research design for successful auditory perception training: 1. compatible testing and training stimuli with verified intelligibility; 2. natural versus synthesised speech; 3. multiple exemplars of one or more target sounds in a range of phonetic environments to facilitate generalisation to new stimuli; 4. stimuli produced by multiple talkers (i.e., voices) to promote robust perceptual category development and generalisability; 5. an identification (vs. a discrimination) task; 6. relatively implicit training with feedback; 7. generalisation to improved perception of novel stimuli and unfamiliar voices; 8. transfer of skill to other tasks such as production in the absence of explicit production training; 9. retention of improved abilities. Only a small percentage of perception studies incorporate all of these features, HVPT being the focus of most of them. As explained in Thomson (2018), HVPT methodology for perceptual enhancement has been shown to help learners revise their mental representations of L2 sounds with potential transfer to production of the sounds whose perception has been trained. In addition to HVPT, many different ways of training perception have proven effective, as shown in the studies reviewed by Sakai and Moorman (2018) and also the perception-oriented studies reviewed by Lee et al. (2015) and Saito and Plonsky (2019), but most of the studies reporting positive findings do not use the same methodology. Pronunciation teaching would benefit by more replications and comparative studies of different methodologies to establish the reliability and generalisability of the effects of various approaches to training perception in different contexts. 2.3 Insufficient attention to methodological factors affecting research outcomes A problem with the advocacy of specific methods is that those methods may be implemented in different ways, with different effects. For example, a suprasegmental orientation for pronunciation can be as minimal as a focus on syllables rather than phonemes or as maximal as a focus on discourse intonation (Brazil, 1997; Brazil et al., 1980). In between these two extremes, it might focus on word stress, sentence stress, intonation, or phrase-level or sentencelevel linking between words. It might also be oriented to perception or production, or both. And the methods for teaching suprasegmentals within any of these orientations vary widely. It would be of value to design studies to specifically test the effects of different aspects of suprasegmentals, and especially to compare the effectiveness of more micro-level (e.g., syllableor word-level) vs. more macro-level (e.g., sentenceor discourse-level) approaches. Pennington Critical examination, L2 pronunciation teaching research 58 A recent computer-based study by Hirata (2024), for instance, found that in training pitch and durational contrasts for English learners of Japanese, sentence-level input was more effective than word-level. Another example of variation in implementation of pronunciation teaching methodology is the technique of shadowing in which the learner aims to repeat exactly what they hear with as little delay as possible, almost simultaneously. Shadowing, which is based on stretches of running speech, has been shown in some research to improve intelligibility and other global aspects of pronunciation (Foote & McDonough, 2017; Hamada, 2019; Murphey, 2001). “Pure” shadowing is done using continuous stretches of speech that the learner has not been exposed to before (in either written or spoken versions). The shadowing technique can however be performed in many different ways that affect the task and its outcomes (Hamada, 2019; Murphey, 2001), and most studies include moderating factors without systematic examination of how they affect results. In any given case, shadowing may or may not include explicit instruction, analysis of the stretches of spoken language that are shadowed, and use of written versions of what is heard. It may also involve manipulations such as a reduced rate of speaking and/or shadowing in short stretches with pauses in between each stretch rather than continuously as speech is produced in longer stretches. Shadowing may be done “live,” that is, on speech as it is occurring, or based on an audio recording, which is the common way in teaching contexts as it allows the shadowed material to be accessed multiple times, as well as via computer delivery. The technique may be done through repeating either out loud or silently, and with exact or complete repetition or only selective repetition of some words and phrases (Murphey, 2001). Comparative research on the different methodological options for shadowing can help to tease out which features, or combinations of features, are most effective. 2.4 Lack of systematic replication and comparison across different conditions A shortcoming related to the insufficient attention to methodological factors affecting research outcomes is the lack of systematic replication and comparison across different conditions, involving student populations, teaching contexts, pronunciation features focused on, and instructional or training methods. Nearly all studies have been carried out on just one class or with one group of learners (Lee et al., 2015; Saito & Plonsky, 2019), and only in a few areas of research ‒ notably, perception studies (Sakai & Moorman, 2018) ‒ have later studies been built on, and specifically compared to, the methodology of prior studies. HVPT is noteworthy in having carefully built a long line of related studies (Thomson, 2018) with some diversity across the studies in student populations, teaching contexts, and pronunciation features focused on ‒ though the HVPT research has tended to focus on certain populations and teaching contexts and on a few segmental contrasts, using essentially the same approach for the training. The research on pronunciation teaching effectiveness can be strengthened by increased attention to comparative research and replication in new contexts. 2.5 Limited attention to authentic language and language variety The limited attention paid to authentic language in teaching materials and methodologies is a major shortcoming of pronunciation practice and research (Saito & Plonsky, 2019). A longpractised tradition in pedagogy for teaching and learning to occur in a step-by-step sequence from simple and controlled to freer and more complex tasks can be seen in the mainstream approach to pronunciation teaching, particularly, in the form summarised in the Communicative Framework for Teaching Pronunciation (Celce-Murcia et al., 2010, pp. 44– 48), which is often cited as the instructional sequence followed in pronunciation teaching Pennington Critical examination, L2 pronunciation teaching research 65 teachers can have confidence that what they teach is supported by their own classroom research, as they establish their own linkages between research and practice. Action research is a type of applied research that has been recommended for language teachers and applied in second language contexts (Burns, 1999, 2005; Crookes, 1993; Edge, 2001; Nunan, 1990; Wallace, 1998). It involves trial-evaluation-reflection-adjustment cycles of: 1) trying out an intervention geared to a desired outcome or outcomes; 2) evaluating the effects and effectiveness of the intervention; 3) reflecting on the processes and results of the intervention, and then 4) adjusting actions to either continue on with the original intervention or to try out a different intervention aiming to achieve similar or better results. Teachers can pursue action research as a collaborative activity (Burns, 1999; Gordon, 2008), working together as research partners or groups to systematically investigate approaches for teaching pronunciation in their classes in a comparative way (see Pennington, forthcoming), by:  planning and preparing together,  deciding what pronunciation feature or features to focus on and how to assess them,  taking notes and comparing how things are going on a regular basis, if possible, observing each other’s classes to view the teaching-learning process, and  assessing the effects on students’ perception and production as well as on their motivation and attitudes to the approaches. I encourage those who conduct research on pronunciation teaching to expand their horizons to aid the field in creating a richer and more diverse teaching and research agenda. There is a need especially for more comparative studies—such as those examining the effects of teaching pronunciation at different stages of language learning, in different teaching contexts, or using contrasting pedagogical approaches or the same approach but with different constraints or affordances. Pronunciation teaching with technology also remains a vast territory awaiting further exploration. 5. Towards the future of pronunciation teaching and research In the current era, pronunciation teaching and research is transitioning beyond its traditional narrow and artificially constrained pedagogical focus to incorporate contextualised and authentic language and to enhance pedagogy through use of technology, media, and the internet. In addition to established research traditions, pronunciation teaching and research are rapidly developing new methodologies incorporating linguistic variety, discourse, and social context in innovative ways of teaching pronunciation with positive results. Future development of pronunciation teaching and research should attend to larger context features, including the complex of segmental and suprasegmental effects in stretches of speech and discourse of different kinds, the kinaesthetic accompaniments of speech, and the social dimensions of communication that shift teaching of pronunciation away from segmental errors and accent, and towards intelligibility, speaking style, and delivery. Researchers can assist in extending the research agenda by replicating and adapting the new methodologies in different contexts and for different student populations, as classroom teachers draw on the published research to try out new approaches in their classroom through action research. In this way, both researchers and teachers will be participating in building a future for pronunciation pedagogy that incorporates the richness of language and of the contexts in which it occurs. Pennington Critical examination, L2 pronunciation teaching research 66 References Bassetti, B. (2024). Orthographic effects in the phonetics and phonology of second language learners and users. In M. Amengual (Ed.), The Cambridge handbook of bilingual phonetics and phonology (pp. 699–720). Cambridge University Press. Block, D. (2003). The social turn in second language acquisition. Georgetown University Press. Brazil, D. (1997). The communicative value of intonation in English. Cambridge University Press. Brazil, D., Coulthard, M., & Johns, C. (1980). Discourse intonation and language teaching. Longman. Burns, A. (1999). Collaborative action research. Cambridge University Press. Burns, A. (2005). Action research. In E. Hinkel (Ed.), Handbook of research in second language teaching and learning (pp. 241–256). Erlbaum. Celce-Murcia, M., Brinton, D. M., Goodwin, J. M., & Griner, B. (2910). Teaching pronunciation: A course book and reference guide. Cambridge University Press. Crookes, G. (1993). Action research for second language teachers: Going beyond teacher research. Applied Linguistics, 14(2), 130–144. https://doi.org/10.1093/applin/14.2.130 Derwing, T. M., & Munro, M. J. (2015). Pronunciation fundamentals: Evidence-based perspectives for L2 teaching and research. John Benjamins. Ding, S., Liberatore, C., Sonsaat, S., Lučić, I., Silpachai, A., Zhao, G., Chukharev-Hudilainen, E., Levis, J., & Gutierrez-Osuna, R. (2019). Golden speaker builder – An interactive tool for pronunciation training. Speech Communication, 115, 51–66. https://psi.engr.tamu.edu/wpcontent/uploads/2019/11/1-s2.0-S0167639319302675-main.pdf Eckert, P., & Rickford, J. R. (Eds.). (2001). Style and sociolinguistic variation. Cambridge University Press. Edge, J. (2001). Action research. TESOL. Edwards, J. (2009). Language and identity: An introduction. Cambridge University Press. Evers, K., & Chen, S. (2022). Effects of an automatic speech recognition system with peer feedback on pronunciation instruction for adults. Computer Assisted Language Learning, 35(8), 1869– 1889. https://doi.org/10.1080/09588221.2020.1839504 Foote, J. A., & McDonough, K. (2017). Using shadowing with mobile technology to improve L2 pronunciation. Journal of Second Language Pronunciation, 3(1), 34–56. https://doi.org/10.1075/jslp.3.1.02foo Fouz-González, J. (2017). Pronunciation instruction through Twitter: The case of commonly mispronounced words. Computer Assisted Language Learning, 30(7), 631–663. https://doi.org/10.1080/09588221.2017.1340309 Fouz-González, J. (2019). Podcast-based pronunciation training: Enhancing FL learners’ perception and production of fossilised segmental features. ReCALL, 31(2), 150– 169. https://doi.org/10.1017/S0958344018000174 Fouz-Gonzáles, J. (2025). Teaching and learning pronunciation with technology: Current possibilities, pedagogical recommendations, and directions for the future. In J. M. Levis, M. Duris, S. SonsaatHegelheimer, & I. Na (Eds.), Proceedings of the 15th Pronunciation in Second Language Learning and Teaching Conference (pp. 1–15). Iowa State University. https://doi.org/10.31274/psllt.19027 Fouz-Gonzáles, J. (forthcoming). Technology-assisted pronunciation training: Bridging research and pedagogy. University of Toronto Press. Galimberti, V., Mora, J. C., & Gilabert, R. (2023). Audio-synchronized textual enhancement in foreign language pronunciation learning from videos. System, 116, 103078. https://doi.org/10.1016/j.system.2023.103078 Giles, H., & Robinson, W. P. (Eds.). (1990). Handbook of language and social psychology. John Wiley & Sons. Pennington Critical examination, L2 pronunciation teaching research 67 Gordon, J. (2021). Pronunciation and task-based instruction: Effects of a classroom intervention. RELC Journal, 52(1), 94–109. https://doi.org/10.1177/0033688220986919 Gordon, J., & Darcy, I. (2022). Teaching segmentals and suprasegmentals: Effects of explicit pronunciation instruction on comprehensibility, fluency, and accentedness. Journal of Second Language Pronunciation, 8(2), 168–195. https://doi.org/10.1075/jslp.21042.gor Gordon, S. P. (2008). Collaborative action research: Developing professional learning communities. Teachers College Press. Hamada, Y. (2019). Shadowing: What is it? How to use it? Where will it go? RELC Journal, 50(3), 386–393. https://doi.org/10.1177/0033688218771380 Hardison, D. M. (2004). Generalization of computer assisted prosody training: Quantitative and qualitative findings. Language Learning & Technology, 8(1), 34–52. http://dx.doi.org/10125/25228 Hardison, D. M. (2018). Visualizing the acoustic and gestural beats of emphasis in multimodal discourse: Theoretical and pedagogical implications. Journal of Second Language Pronunciation, 4(2), 231–258. Hardison, D. M. (in press). The multimodal context of phonological learning. University of Toronto Press. Hardison, D. M., & Pennington M. C. (2021). Multimodal second-language communication: Research findings and pedagogical recommendations. RELC Journal, 52(1), 62–76. https://doi.org/10.1075/jslp.17006.har Henderson, A., & Rojczyk, A. (2023). Foreign language accent imitation: Matching production with perception. In V. G. Sardegna & A. Jarosz (Eds.), English pronunciation teaching: Theory, practice and research findings (pp. 115–133). Multilingual Matters. Hirata, Y. (2024). Computer assisted pronunciation training for native English speakers learning Japanese pitch and durational contrasts. Computer Assisted Language Learning, 17(3-4), 357–376. https://doi.org/10.1080/0958822042000319629 Jenkins, J. (2000). The phonology of English as an international language. Oxford University Press. Jenkins, J. (2002). A sociolinguistically based, empirically researched pronunciation syllabus for English as an international language. Applied Linguistics, 23, 83–103. https://doi.org/10.1093/applin/23.1.83 Lantolf, J. P. (Ed.). (2000). Sociocultural theory and second language learning. Oxford University Press. Lantolf, J. P. (2011). The sociocultural approach to second language acquisition: Sociocultural theory, second language acquisition, and artificial L2 development. In D. Atkinson (Ed.), Alternative approaches to second language acquisition (pp. 24–47). Routledge. LaScotte, D. K., & Tarone, E. (2022). Channeling voices to improve L2 English intelligibility. Modern Language Journal, 106(4), 744 – 763. https://doi.org/10.1111/modl.12812 LaScotte, D., Meyers, C, & Tarone, E. (2021). Voice and mirroring in SLA: Top-down pedagogy for L2 pronunciation instruction. RELC Journal, 52(1), 3–21. https://doi.org/10.1177/0033688220953910 LaScotte, D., Meyers, C, & Tarone, E. (2023). Voice and mirroring in L2 pronunciation instruction. Equinox / University of Toronto Press. Lee, B., Jang, J., & Plonsky, L. (2015). The effectiveness of second language pronunciation instruction: A meta-analysis. Applied Linguistics, 36(3), 345–366. https://doi.org/10.1093/applin/amu040 Lee, B., Plonsky, L., & Saito, K. (2020). The effects of perceptionvs. production-based pronunciation instruction. System, 88, 1–13. https://doi.org/10.1016/j.system.2019.102185 Levis, J. M. (2018). Intelligibility, oral communication, and the teaching of pronunciation. Cambridge University Press. Pennington Critical examination, L2 pronunciation teaching research 68 Levis, J. M. (2019). Cinderella no more! Leaving victimhood behind. Speak Out! IATEFL Pronunciation SIG Journal, 60, 1–7. Levis, J. M. (2021a). Connecting the dots between pronunciation research and practice. In A. KirkovaNaskova, A. Henderson, & J. Fouz-Gonzáles (Eds.), Pronunciation instruction: Research-based insights (pp. 17–37). John Benjamins. Levis, J. M. (2021b). Editorial. L2 pronunciation research and teaching: The importance of many languages. Journal of Second Language Pronunciation, 7(2), 141–153. https://doi.org/10.1075/jslp.21037.lev Levis, J. M. (2022). Teaching pronunciation: Truths and lies. In C. Bardel, C. Hedman, K. Rejman, & E. Zetterholm (Eds.), Exploring language education: Global and local perspectives (pp. 39–72). Stockholm University Press. Levis, J. M. & Moyer, A. (2014). Social dynamics in second language accent. De Gruyter. Lippi-Green, R. (2012). English with an accent: Language, ideology, and discrimination in the United States (3rd ed.). Routledge. Long, M. (1991). Focus on form: A design feature in language teaching methodology. In K. de Bot, R. Ginsberg, & C. Kramsch (Eds.). Foreign language research in cross-cultural perspective (pp. 39– 52). John Benjamins. Long, M. (2015). Second language acquisition and task-based language teaching. Wiley-Blackwell. McCrocklin, S. (2019). ASR-based dictation practice for second language pronunciation improvement. Journal of Second Language Pronunciation, 5(1), 98–118. https://doi.org/10.1075/jslp.16034.mcc Mompean, J. A. & Fouz-Gonzáles, J. (2021). Phonetic symbols in contemporary pronunciation instruction. RELC Journal, 52(1), 155–168. https://doi.org/10.1177/0033688220943431 Mora, J. C., & Levkina, M. (2018). Training vowel perception through map tasks: The role of linguistic and cognitive complexity. In J. Levis (Ed.), Pronunciation in Second Language Learning and Teaching Proceedings (pp. 151-162). Iowa State University. https://www.iastatedigitalpress.com/psllt/article/id/15350/ Murphey, T. (2001). Exploring conversational shadowing. Language Teaching Research, 5(2), 128– 155. Namaziandost, E., Esfahani, F. R., & Hashemifarnia, A. (2018). The impact of using authentic videos on prosodic ability among foreign language learners. International Journal of Instruction, 11(4), 375–390. https://doi.org/10.12973/iji.2018.11424a Nunan, D. (1990). Action research in the language classroom. In J. C. Richards & D. Nunan (Eds.), Second language teacher education (pp. 62–81). Cambridge University Press. Pennington, M. C. (1989). Applications of computers in the development of speaking and listening proficiency. In M. C. Pennington (Ed.), Teaching languages with computers: The state of the art. Athelstan. Pennington, M. C. (1996). When input becomes intake: Tracing the sources of teachers’ attitude change. In D. Freeman & J. C. Richards, Teacher learning in language teaching (pp. 320–348). Cambridge University Press. Pennington, M. C. (2020). Pronunciation and international employability. In Bocanegra-Valle (Ed.), Applied linguistics and knowledge transfer: Employability, internationalisation and social challenges (pp. 225–244). Peter Lang. Pennington, M. C. (2021). Teaching pronunciation: The state of the art 2021. RELC Journal, 52(1), 3– 21. https://doi.org/10.1177/00336882211002283 Pennington, M. C. (forthcoming). The pronunciation book: A teacher’s guide. University of Toronto Press. Pennington, M. C., & Cheung, M. (1995). Factors shaping the introduction of process writing in Hong Kong. Language, Culture, and Curriculum, 8, 15–34. https://doi.org/10.1080/07908319509525185 Pennington Critical examination, L2 pronunciation teaching research 69 Pennington, M. C., & Ellis, N. C. (2000). Cantonese speakers’ memory for English sentences with prosodic cues. Modern Language Journal, 84(3), 372–389. https://doi.org/10.1111/00267902.00075 Pennington, M. C., Lau, L., & Sachdev, I. (2011). Diversity in adoption of linguistic features of London English by Chinese and Bangladeshi adolescents. Language Learning Journal, 39(2), 177–199. https://doi.org/10.1080/09571736.2011.573686 Pennington, M. C., & Richards, J. C. (1997). Re-orienting the teaching universe: The experience of five first-year English teachers in Hong Kong. Language Teaching Research, 1(2), 149–178. https://doi.org/10.1177/136216889700100204 Pennington, M. C., & Rogerson-Revell, P. (2019). English pronunciation teaching and research: Contemporary perspectives. Palgrave Macmillan. Robinson, P. (2011). Task-based language teaching: A review of issues. Language Learning, 61(1), 1– 36. https://doi.org/10.1111/j.1467-9922.2011.00641.x Robinson, W. P., & Giles, H. (Eds.). (2001). The new handbook of language and social psychology. John Wiley & Sons. Rogerson-Revell, P. M. (2021). Computer-assisted pronunciation training (CAPT): Current issues and future directions. RELC Journal, 52(1), 189–205. https://doi.org/10.1177/0033688220977406 Saito, K., & Plonsky, L. (2019). Effects of second language pronunciation teaching revisited: A proposed framework and meta-analysis. Language Learning, 69(3), 652–708. https://doi.org/10.1111/lang.12345 Sakai, M., & Moorman, C. (2018) Can perception training improve the production of second language phonemes? A meta-analytic review of 25 years of perception training research. Applied Psycholinguistics, 39(1), 187–224. https://doi.org/10.1017/S0142716417000418 Sánchez-Auñón, E., Férez-Mora, P. A. & Monroy-Hernández, F. (2023). The use of films in the teaching of English as a foreign language: A systematic literature review. Asian Journal of Second and Foreign Language Education 8(10), 1–17. https://doi.org/10.1186/s40862-022-00183-0 Sardegna, V. G., & Jarosz, A. (2023). Learning English word stress with technology. In R. I. Thomson, T. M. Derwing, J. M. Levis, & K. Hiebert (Eds.), Proceedings of the 13th Pronunciation in Second Language Learning and Teaching Conference (pp. 1-10). Iowa State University. https://doi.org/10.31274/psllt.16142 Sperti, S. (2017). Phonopragmatic dimensions of ELF in specialized immigration contexts. Working Papers del Centro di Recerca sulle Lingue Franchenella Comunicazione Interculturale e Multimediale, Dipartimento di Studi Umanistici, Università del Salento, 3. https://doi.org/10.1285/i24991449n3 Thornbury, S. (1996). Paying lip-service to CLT. EA [ELICOS Association] Journal, 14(1), 51–63. http://www.scottthornbury.com/articles.html Thomson, R. I. (2018). High variability [pronunciation] training (HVPT): A proven technique about which every language teacher and learner ought to know. Journal of Second Language Pronunciation, 4(2), 208–231. https://doi.org/10.1075/jslp.17038.tho Thomson, R. I. & Derwing, T. (2015). The effectiveness of L2 pronunciation instruction: A narrative review. Applied Linguistics, 36(3), 326–344. https://doi.org/10.1093/applin/amu076 Wallace, M. (1998). Action research for language teachers. Cambridge University Press. Wichmann, A., Dehé, N., & Barth-Weingarten, D. (2009). Where prosody meets pragmatics: Research at the interface. In A. Wichmann, N. Dehé, & D. Barth-Weingarten (Eds.), Where prosody meets pragmatics, Studies in Pragmatics 8 (pp. 1–20). Emerald. Yu, L. & Odlin, T. (2016). New perspectives of transfer in second language learning. Multlingual Matters. Pennington Critical examination, L2 pronunciation teaching research 70 Wisniewska, N., & Mora, J. (2020). Can captioned video benefit second language pronunciation? Studies in Second Language Acquisition, 42(3), 599–624. https://doi.org/10.1017/S0272263120000029 About the author Martha C. Pennington (PhD Linguistics, University of Pennsylvania) is a Research Fellow at Birkbeck University of London in the School of Creative Arts, Culture and Communication and is currently a Visiting Professor at the University of Reading in the Department of English Language and Applied Linguistics. Her books include Phonology in English Language Teaching: An International Approach (Longman, 1996), The Power of CALL (Athelstan, 1996), Phonology in Context (Palgrave Macmillan, 2007), English Pronunciation Teaching and Research: Contemporary Perspectives (Palgrave Macmillan, 2019, co-authored with Pamela Rogerson-Revell), and The Pronunciation Book: A Teacher’s Guide, to appear in her series Applied Phonology and Pronunciation Teaching, which has recently been acquired from Equinox by the University of Toronto Press. She has also published widely in literacy, including Why Reading Books Still Matters: The Power of Literature in Digital Times (Routledge, 2018, co-authored with Robert P. Waxler), and in language education, including Language Program Leadership in a Changing World: An Ecological Model (De Gruyter Brill, 2010, co-authored with Barbara J. Hoekje), volume 1 in her series Innovation and Leadership in English Language Teaching. Email: m.[email protected] This chapter is based on the oral presentation given by the author at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons.org/licenses/by/4.0/ . Pietraszek, M. (2025). Live or leave! Vowel quality and vowel length in the intelligibility of advanced Spanishaccented English. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 71–83). Université GrenobleAlpes. https://doi.org/10.5281/zenodo.16696002/ Live or leave! Vowel quality and vowel length in the intelligibility of advanced Spanish-accented English Mateusz Pietraszek Complutense University of Madrid Abstract Often characterised as short and long in both research and teaching practice (Cruttenden, 2008; Hancock, 2012, Walker, 2010), the KIT and FLEECE vowels are commonly confused by Spanish learners of English as their mother tongue lacks a similar distinction (RAE, 2011). Length is also traditionally insisted upon in instruction and even more recent intelligibility-based (Levis, 2005) approaches to international English (Jenkins, 2000; Walker, 2010) advocate its importance prioritising it over vowel quality. The goal of this paper is to examine the role of vowel length and quality in Spanish-accented English from a quantitative and qualitative exploratory perspective. Sixty advanced English speakers with L1 Spanish were recorded and the acoustic properties of their FLEECE and KIT vowels were analysed in the words live and leave. The overall intelligibility (Kang et al., 2018; Wang, 2007), comprehensibility, and foreign-accentedness (Derwing & Munro, 1997, 2015) of the participants’ speech were assessed online by an international sample of 330 listeners from 28 L1 backgrounds. No association was detected between the overall pronunciation scores and vowel length. However, some evidence suggested that distinguishing at least one formant might be associated with increased comprehensibility and decreased accentedness. Additionally, the analysis of phonological intelligibility breakdowns showed that these could only be explained with a reference to formant values but not length distinctions. Thus, it is argued that vowel length in KIT and FLEECE may be disregarded and the focus in receptive and productive instruction should be on quality distinctions. Keywords: accentedness, intelligibility, vowel length, vowel quality Pietraszek Vowel quality & length, Spanish-accented English 72 1. Introduction The English FLEECE /iː/ and KIT /ɪ/ vowels are most often characterised as long and short segments in both dictionaries and learning materials for L2 English students (Cruttenden, 2008; Hancock, 2012; Wells, 2008), especially – but not exclusively – those teaching British English. In educational contexts, this might lead to the conclusion that contrasting these vowels mainly based on length is sufficient in order to tease them apart, as isolated drills very often focus on length as the most integral feature to be trained when contrasting those vowels (see MompeánGonzález, 2001). However, this simplistic assumption poses several problems aggravated by the very fact that vowel length is a relative phenomenon. Firstly, the FLEECE and KIT vowels have different articulations regardless of their length, the latter being both more open and back than the former.1 Secondly, analyses of Standard Southern British English (Lindsey, 2019; Wells, 1982) and General American (Celce Murcia et al., 2010) tend to describe the FLEECE vowel as a diphthong [ɪj], which glides from a more open and centralised articulation towards a higher and fronter point of articulation. Thirdly, vowel length in English is context-dependent (Ladefoged, 2001; Odden, 2011) and it is commonly used as a cue regarding the voicing of the following consonant (Rojczyk, 2010) or even the preceding consonant (Viswanathan et al., 2020) by native speakers of English. As a result of pre-fortis clipping (Roach, 2009), long vowels may even be pronounced shorter than short vowels, e.g., when comparing a long vowel before a voiceless stop with a short vowel before a long one. Length also varies depending on the number of syllables in a word (Ladefoged, 2001), hence the first syllable of reading will be shorter than the word read in the same context. Finally, perceptual studies suggest length is insufficient when distinguishing vowel qualities among native speakers (Hillenbrand, 1995; Liu et al, 2014). It is therefore evident that vowels are never distinguished by length alone in English, and the descriptions of some varieties of English – including the most widely spoken, General American – make no reference to length as a relevant vowel feature (see Avery & Ehrlich, 1992). However, even relatively recent teaching recommendations – in all likelihood encouraged by canonical phonemic transcription in dictionaries (Wells, 2008) – insist on preserving length contrasts. Not only do they posit length as a feature necessary for internationally intelligible English, but they may even prioritise it over quality distinctions. This is the case, for example, of the Lingua Franca Core of features necessary for international communication in English as a global language (Jenkins, 2000, 2002; Walker, 2010). However, many languages do not have phonological length distinctions whereas all languages, including English and Spanish, distinguish between vowel qualities, which might be sufficient for basic vowel distinctions in English as an international language. In the particular case of Spanish, although vowel length differences exist in the language (Marín Gálvez, 1994), their phonological functions are not comparable to those in English (Fox et al., 1995). This feature, coupled with the fact that Spanish only has one high front vowel /i/, indeed poses pronunciation problems for L1-Spanish speakers of English (Finch & Ortiz Lira, 1982; Gómez-González & Sánchez-Roura 2016; Hancock, 2012; Walker, 2010). Previous studies have shown that, unlike English native speakers, Spanish speakers are largely insensitive to length differences (Fox et al., 1995; Viswanathan et al., 2020) or that their sensitivity to duration is lower than that of native speakers (Kondaurova & Francis, 2008). Other evidence suggests that instruction practices may lead Spanish speakers to rely on length even more than native speakers do in discrimination tasks (Mompeán-González, 2001). Finally, regarding international 1 See Durand (2005) for a comprehensive summary of the debate on the relevance of vowel length and quality in English phonology. Pietraszek Vowel quality & length, Spanish-accented English 73 intelligibility, the evidence in support of length distinctions to preserve intelligibility has so far been scarce and inconclusive (Jurado-Bravo, 2018). 1.1 Aims and research questions The main goal of this exploratory study is to examine the utility of insisting on length versus quality distinctions in teaching the FLEECE : KIT contrast in Peninsular Spanish, applying the intelligibility principle (Levis, 2005) in English as an international lingua franca (Jenkins, 2000). With this aim in mind, the following research questions (RQ) will be addressed: RQ1: Do Spanish speakers of English in the sample produce the FLEECE and KIT vowels with different length and/or quality? RQ2: Is a contrast between FLEECE and KIT in L2 English associated with improved general pronunciation performance scores? RQ3: What phonological intelligibility breakdowns may occur when FLEECE and KIT vowels are mispronounced? 2. Methodology 2.1 Participants, recordings, and scoring This mixed-methods (Creswell et al., 2008) exploratory study is part of a broader research project (Pietraszek, 2024a, Pietraszek, 2024b). The sample consisted of 60 university students (26 females and 34 males; aged 19–26, M = 21.2) majoring in disciplines unrelated to linguistics and having B2–C2 levels of English proficiency. The speakers were recorded using a set of semantically unpredictable sentences (SUS) (Benoît et al., 1996; Pietraszek, 2024a) and an elicitation paragraph (Pietraszek, 2024b). SUS were selected as the best instrument to test phonemic intelligibility, because most semantic information is eliminated (Benoît et al., 1996; Kang et al., 2018). Subsequently, 330 international listeners with 28 different L1s evaluated the speakers’ comprehensibility and foreign-accentedness, and transcribed their SUS for intelligibility measurements. The comprehensibility (COM) and foreign-accentedness (FA) (Derwing & Munro, 1997, 2015) scores are averaged ratings the speakers received from the listeners on a 1–6 semantic differential scale using excerpts from the paragraph (Pietraszek, 2024b). Intelligibility (INT) was operationalised as phonological utterance decoding at the segmental level (cf. Jenkins, 2000; Kang et al., 2018). Thus, the INT score was calculated as a mean score based on the number of words correctly interpreted by the listeners in the test following a specially designed protocol (Pietraszek, 2024a). Similarly, the qualitative part of the study, i.e., intelligibility breakdown analysis, is based on listener transcriptions of SUS in the online test (Pietraszek, 2024a). 2.2 Acoustic analyses The acoustic analyses are based on two words extracted from two syntactically comparable SUS: live for the KIT vowel and leave for the FLEECE vowel. The vowel length of the LIVE and LEAVE vowels was measured for each speaker using Praat (Boersma & Weenink, 2019). As “the lengths of segments depend on their position in the word, their position in the phrase and the whole utterance, where the stresses occur in the utterance, and many other factors” (Ladefoged, 2003, p. 102), in order for the measurements to be comparable, the instances were taken from SUS number 3 and 4 in series 3 for the minimal pair leave–live, i.e., Leave the sport and the thought and Live the sport and the fund. The fact that the vowels came from two Pietraszek Vowel quality & length, Spanish-accented English 74 consecutive sentences ensured a relatively stable pace of speech per speaker. Additionally, the syntactic position of both words was the same, in order to minimise contextual, phonotactic influence on vowel length. The exact procedure and the selection of starting and ending points in length measurements was particularly important owing to the phonological context of the targeted vowels. These were situated between a lateral approximant and a fricative, whose boundaries are more fluid than those of plosives, for example, as both laterals and fricatives are closer to vowels on the sonority scale (see Cruttenden, 2008). The boundaries were set using a spectrogram (see Figure 1) showing the intensity of vowel formants. The length was extracted using the accompanying waveform (see Figure 2) by selecting the first valley and the Figure 1 Spectrogram for /liːv/ Figure 2 Oscillogram for /liːv/ Pietraszek Vowel quality & length, Spanish-accented English 81 While this exploratory research cannot reliably answer the question of whether length plays a role in international English intelligibility, it provides some evidence that changing the quality of a vowel may lead to misunderstandings. Thus, dismissing vowel quality as a core international English feature while insisting on length seems unfounded and therefore risky. It also places an additional burden on the learners who are asked to develop a phonological vowel category foreign to their linguistic system, while their target language uses it mainly as a marker of final consonant voicing. Thus, it seems reasonable, as a pedagogical implication of the results presented in this study, to raise learners’ awareness of the varying duration due to prefortis clipping regardless of the exact vowel quality, rather than to treat duration as the most relevant distinctive feature of the FLEECE and KIT phonemes. References Avery, P., & Ehrlich, S. (1992). Teaching American English pronunciation. Oxford University Press. Benoît, C., Grice, M., & Hazan, V. (1996). The SUS test: A method for the assessment of text-to-speech synthesis intelligibility using Semantically Unpredictable Sentences. Speech Communication, 18(4), 381–392. http://dx.doi.org/10.1016/0167-6393(96)00026-X Boersma, P., & Weenink, D. (2025). Praat: Doing phonetics by computer [Computer program]. Version 6.4.36. https://praat.org Celce-Murcia, M., Brinton, D., Goodwin, J. M. & Griner, B. (2010). Teaching pronunciation: A course book and reference guide (2nd ed.). Cambridge University Press. Cohen, J. (1988). Statistical power analysis for the behavioral Sciences. Psychology Press. Creswell, J. W., Plano Clark, V. L., Gutmann, M. L., & Hanson, W. E. (2008). An expanded typology for classifying mixed methods research into designs. In V. L. Plano Clark, & J. W. Creswell (Eds.), The Mixed Methods Reader (pp. 161–196). SAGE Publications Inc. Cruttenden, A. (2008). Gimson's pronunciation of English. Hodder Education. Derwing, T. M., & Munro, M. J. (1997). Accent, intelligibility, and comprehensibility: Evidence from four L1s. Studies in Second Language Acquisition, 19(1), 1–16. https://doi.org/10.1017/S0272263197001010 Derwing, T.M., & Munro, M.J. (2015). Pronunciation fundamentals: Evidence-based perspectives for L2 teaching and research. John Benjamins. Durand, J. (2005). Tense/Lax, the vowel system of English and phonological theory. In P. Carr, J. Durand, & C. J. Ewen (Eds.), Headhood, elements, specification and contrastivity: Phonological papers in honour of John Anderson (pp. 77–97). John Benjamins. https://doi.org/10.1075/cilt.259.08dur Finch, D. F., & Ortiz Lira, H. (1982). A course in English phonetics for Spanish speakers. Heinemann Educational Books. Fox, R. A., Flege, J. E., & Munro, M. J. (1995). The perception of English and Spanish vowels by native English and Spanish listeners: A multidimensional scaling analysis. The Journal of the Acoustical Society of America, 97(4), 2540–2551. https://doi.org/10.1121/1.411974 Gómez González, M. d. l. A., & Sánchez Roura, T. (2016). English pronunciation for speakers of Spanish. De Gruyter. Hancock, M. (2012). English pronunciation in use. Cambridge University Press. Hillenbrand, J., Getty, L. A., Clark, M. J., & Wheeler, K. (1995). Acoustic characteristics of American English vowels. The Journal of the Acoustical Society of America, 97(5), 3099–3111. https://doi.org/10.1121/1.411872 Hisagi, M., Higby, E., Zandona, M., Acosta, A. P., Kent, J., & Tajima, K. (2024). Impact of speech rate on perception of vowel and consonant duration by bilinguals and monolinguals. JASA Express Letters, 4(5), 055201. https://doi.org/10.1121/10.0025862 Jenkins, J. (2000). The phonology of English as an international language. Oxford University Press. Jenkins, J. (2002). A sociolinguistically based, empirically researched pronunciation syllabus for English as an international language. Applied Linguistics, 23. 83–103. http://dx.doi.org/10.1093/applin/23.1.83 Pietraszek Vowel quality & length, Spanish-accented English 82 Jurado-Bravo, M. Á. (2018). Vowel quality and vowel length in English as a lingua franca in Spain. Miscelánea, 57(57), 13–34. https://doi.org/10.26754/ojs_misc/mj.20186310 Kang, O., Thomson, R. I., & Moran, M. (2018). Empirical approaches to measuring the intelligibility of different varieties of English in predicting listener comprehension. Language Learning, 68(1), 115–146. https://doi.org/10.1111/lang.12270 Kondaurova, M. V., & Francis, A. L. (2008). The relationship between native allophonic experience with vowel duration and perception of the English tense/lax vowel contrast by Spanish and Russian listeners. The Journal of the Acoustical Society of America, 124(6), 3959–3971. https://doi.org/10.1121/1.2999341 Ladefoged, P. (2001). Vowels and consonants. Blackwell. Ladefoged, P. (2003). Phonetic data analysis: An introduction to fieldwork and instrumental techniques. Blackwell. Levis, J. M. (2005). Changing contexts and shifting paradigms in pronunciation teaching. TESOL Quarterly, 39(3), 369–377. http://dx.doi.org/10.2307/3588485 Lindsey, G. (2019). English after RP. Oxford University Press. Liu, C., Jin, S., & Chen, C. (2014). Durations of American English vowels by native and non-native Speakers: Acoustic analyses and perceptual effects. Language and Speech, 57(2), 238–253. https://doi.org/10.1177/0023830913507692 Marín Gálvez, R. (1994). La duración vocálica en español. Estudios de lingüística, 10(10), 213–226. https://doi.org/10.14198/ELUA1994-1995.10.11 Mompeán-González, J. A. (2001). A comparison between English and Spanish subjects' typicality ratings in phoneme categories: A first report. International Journal of English Studies, 1(1), 115– 155. https://revistas.um.es/ijes/article/view/47641 Odden, D. (2011). The representation of vowel length. In Oostendorp, M. et al. (Eds.), Companions to linguistics: The Blackwell companion to phonology (pp. 465–490). John Wiley & Sons, Ltd. Peterson, G. E., & Lehiste, I. (1960). Duration of syllable nuclei in English. The Journal of the Acoustical Society of America, 32(6), 693–703. https://doi.org/10.1121/1.1908183 Pietraszek, M. (2024a). Associating speaker variables with English pronunciation ratings in Spanish tertiary education. Porta Linguarum An International Journal of Foreign Language Teaching and Learning, (41), 189–207. https://doi.org/10.30827/portalin.vi41.27083 Pietraszek, M. (2024b). ‘The sun was rising’: A new elicitation paragraph for English pronunciation research and assessment. In D. Martín González (Ed.), English linguistics meets the 21st century. Dykinson. RAE (Real Academia Española). (2011). Nueva gramática de la lengua española: fonética y fonología. Espasa Calpe. Roach, P. (2009). English phonetics and phonology (4th ed.). Cambridge University Press. Rojczyk, A. (2010). Preceding vowel duration as a cue to the consonant voicing contrast: Perception experiments with Polish-English bilinguals. In E. Waniek-Klimczak (Ed.), Issues in accents of English: Variability and norm (pp. 342–360). Cambridge Scholars. Viswanathan, N., Olmstead, A. J., & Aivar, M. P. (2020). The use of vowel length in making voicing judgments by native listeners of English and Spanish: Implications for rate normalization. Language and Speech, 63(2). https://doi.org/10.1177/0023830919851529 Walker, R. (2010). Teaching the pronunciation of English as a lingua franca. Oxford University Press. Wang, H. (2007). English as a lingua franca: Mutual intelligibility of Chinese, Dutch and American speakers of English. LOT. Wells, J. (1962). A study of the formants of the pure vowels of British English. [Master’s thesis]. University of London. Wells, J. C. (1982). Accents of English. Cambridge University Press Wells, J. C. (2008). Longman pronunciation dictionary. Pearson Longman. Pietraszek Vowel quality & length, Spanish-accented English 83 About the author Mateusz Pietraszek is an Assistant Professor at the Complutense University of Madrid. He holds a PhD in English Linguistics and an MA in Spanish Linguistics. He has been teaching English for Specific Purposes, General English, and English Phonetics and Phonology at the university level for 16 years. His research interests include English and Spanish phonetics (intelligibility, accent, and attitudes to pronunciation), English as a Lingua Franca (ELF), comparative morphosyntax, English-medium instruction (EMI), and higher education. A passionate language learner, he speaks seven languages (English, Spanish, Polish, French, Catalan, German, and Esperanto) and is a member of the International Hyperpolyglot Association (HYPIA). Email: mp[email protected] This chapter is based on the oral presentation given by the author at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons.org/licenses/by/4.0/ . Rasmussen, S. H. (2025). Word-initial voicing in L2 English: Evaluations of /z/ productions before and after perceptual training. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 84–96). Université Grenoble-Alpes. https://doi.org/10.5281/zenodo.16737662 Word-initial voicing in L2 English: Evaluations of /z/ productions before and after perceptual training Sidsel Holm Rasmussen Aarhus University Abstract This study examines the effect of perceptual training on the production of syllable onset L2 English /z/. L1 Danish participants’ productions of English /z/ and /s/ were evaluated in terms of intelligibility, defined as the ability of native English listeners to correctly identify the intended phone in a Two-Alternative Forced-Choice task. Overall, the training group was more intelligible than the control group. After training, trainees’ /z/ productions were identified with 61.3% accuracy (gain score of 12 percentage points), whereas the gain score for the control group averaged 4.5 percentage points (from 52.9% to 57.4%). Acoustic measurements were examined from the trainees whose /z/ tokens improved the most after training. Comparison between the two types of evaluations suggests that native listeners’ identification of the English voiced alveolar sibilant corresponds to a greater percentage (≥ 25%) of periodicity in the fricative. Methodological considerations concerning different types of production evaluations are discussed. Keywords: English sibilants, evaluation methods, foreign accent, fricative voicing, L2 phonetic training, perception and production Rasmussen Word-initial voicing of /z/, impact of training 85 1. Introduction A number of recent review papers delve into High Variability Phonetic Training (HVPT) as a widely used method for efficiently facilitating the acquisition of novel speech sounds (Barriuso & Hayes-Harb, 2018; Thomson, 2018). This line of research has a long tradition, with early studies demonstrating that exposure to several exemplars of the English /r – l/ contrast improved L1 Japanese listeners’ perception of the non-native speech sounds (Logan et al., 1991; Pisoni et al., 1994). Furthermore, the effects of perception training on L1 Japanese participants’ production of these English consonants were explored using different methods of evaluations which demonstrated that learning (for some trainees) led to improved productions (Bradlow et al., 1997). Meta-reviews have further scrutinised this potential for L2 perception training to lead to improvements in L2 production, and find that small-to-medium production gains are reported in the literature (Sakai & Moorman, 2018; Uchihara et al., 2024). Despite the wide range of results in studies examining the proposed link between perception and production in non-native speech learning, as well as the fact that the two modalities are now generally considered to be co-evolving in a bidirectional manner (Flege & Bohn, 2021), HVPT seems to offer a useful method for learners to receive targeted exposure to specific non-native contrasts. This paper provides results from some of the findings from an HVPT experiment, exploring the efficacy of perception training on L1 Danish participants’ production of English /z/. 1.1 L2 perception and production of English /s/ and /z/ The pair of English alveolar sibilants is known to cause problems for many learners whose L1 does not have voiced fricatives, and issues related to confusion between the two English phones are reported in both perception and production. For example, it was found that L1 Spanish speakers fail to produce a distinction (as measured acoustically) between word medial and word final English /z/ and /s/ (Mairano et al., 2021). With focused perceptual training, however, L1 Spanish students of English improved both identification and production of /z/ (FouzGonzález, 2019). Similarly, in a longitudinal transcription study, Kapranov (2022) found that Norwegian EFL learners consistently misheard /z/ as /s/ even after attending a course in English phonetics. Kapranov takes this result to be indicative of L1 Norwegian learners’ persistent (perceptual) difficulties with the similar sounding non-native /z/. For L1 Swedish speakers of L2 English, McAllister (2007) also reports difficulties related to the /s – z/ contrast in codas; while all productions of English /s/ were judged acceptable by a native English listener, the same was only true for about 22% of the Swedish participants’ /z/ productions, and half of the successful /z/ items were produced by the same two most successful participants. Evidently, many different EFL speakers encounter the same type of problems with the English voiced alveolar fricative. Acquisitional difficulties for the English /s – z/ contrast have also been explored for L1 Danish participants in previous studies. Examining the contrast in word-final position, Eger and Bohn (2015) found that L1 Danish speakers perceived the contrast less categorically than did native English listeners, yet performed more like native English participants in perceptually distinguishing word-final /s/ and /z/ than they did in production. These findings are in line with those reported by Trapp and Bohn (2002) who trained nine L1 Danish adolescents to perceptually identify coda /s/ and /z/ from minimal pairs. While the trainees’ mean identification scores had improved after training, there was no significant improvement in their measured production accuracy. In particular, “it was the production of the word-final /z/ rather than /s/ that created problems for the Danes” (p. 349). Moreover, acoustic analyses demonstrated that L1 Danish trainees’ vowel-to-fricative ratio was equally short before and Rasmussen Word-initial voicing of /z/, impact of training 86 after training, participants thus failing to employ the duration cue common in native English productions of coda /z/. Whereas these early studies were concerned with /s/ or /z/ in word final position, Horslund and Bohn (2022) examined how Danes perceive English consonants in initial position. First, they found that initial /s/ and /z/ assimilate to Danish /s/ as a Category Goodness (CG) assimilation type (Best & Tyler, 2007). The authors thus hypothesised that English /s/ is easily identifiable to L1 Danish listeners, while identification of /z/ depends on previous exposure to English. Interestingly, expectations from this hypothesis were not borne out in their following identification study – rather, Horslund and Bohn observed lower identification scores for /s/ than for /z/, and lower sibilant identification scores for the more experienced of two learner groups included. These unexpected results caused the authors to speculate about potential hypercorrection that may lead to such response biases favouring the less familiar phone in a CG pair such as /s/~/z/. The surprising results from Horslund and Bohn’s study (2022), coupled with previous research and reports of other L1 EFL learners’ issues, motivates this study of L1 Danish listeners’ perception and production of initial English /s/ and /z/ before and after training.1 The scope of the current paper is concerned only with the productions of the English sibilant contrast in onset position produced by L1 Danish speakers, as the aim of this paper is twofold: First, we examine the extent to which perceptual training of syllable initial English /s/ and /z/ has an effect on non-native production of the same speech sounds. Since Danish does not have any voiced fricatives, both English sibilants are assimilated to the native /s/ (Horslund & Bohn, 2022), and production of the unvoiced English /s/ is expected to be unproblematic (or sufficiently similar to be identified correctly by L1 English listeners), while production of the voiced /z/ is expected to be less nativelike. Secondly, this research also aims to explore the insights we may gain from different types of evaluations to non-native pronunciation. Specifically, this paper addresses the following research questions: RQ1: How accurately can L1 Danish speakers produce the syllable initial sibilant contrast /s – z/ (as judged by native English listeners) before training? RQ2: Do L1 Danish speakers’ production of /z/ (the more challenging of the sibilant pair) become more intelligible to native listeners after perceptual training? RQ3: Which production strategies are deployed by the speakers for whom correct /z/ identification by native evaluators is highest/most improved? 2. Method 2.1 Participants: Native Danish speakers Recordings from 45 (F = 34, M = 11) L1 Danish speakers were included in this study. Participants were randomly assigned to the experimental training group (TG, n = 33) or the control group (CG, n = 12). Their profile data are presented in Table 1. Their ages ranged from 20–76 (M = 44.6, SD = 21.9)2 and they had started learning English at the mean age of 10.2 (SD = 1.91).3 Participants completed a simple language background questionnaire from which 1 More information about the project and our group’s research on perceptual training is available on the website: https://cc.au.dk/en/phonetic-flexibility-in-old-age/our-research 2 Note that the participants were recruited to match one of our specified age groups (either 18–35 years or 60+) 3 Younger participants had on average started learning English earlier (8.7 yrs. old) than the older participants (11.7 yrs.) Rasmussen Word-initial voicing of /z/, impact of training 87 we confirmed that they had not spent more than 6 months in an English-speaking country. On a scale from 1–5 (poor–good) participants also reported their self-assessed level of “ability to understand spoken English” and their “ability to speak English”. Table 1 Background Information for L1 Danish Participants TG (n = 33) M (SD) CG (n = 12) M (SD) Age 44.6 (21.7) 44.7 (23.4) AOL English 10.2 (1.91) 10.3 (1.97) Score, understand 4.18 (0.81) 4.08 (1.08) Score, speaking 4.03 (0.95) 3.83 (1.53) 2.2 Production task Production of both initial and final /s/ and /z/ were elicited in a delayed repetition task (Flege et al., 1995). Participants first listened to a recording of a native speaker who produced 46 words4 separately, inserted in the carrier phrase: The next word is _____, after which they heard the same voice ask: What is the next word? They were then prompted to repeat the word in the carrier phrase: Let it be _____. Target words examined in the current study were 15 (near) minimal monosyllabic pairs of English words which differed in voicing in the initial fricative (e.g., sink  zink). Productions were elicited from participants before the perception tasks (Lab1), and the same production task was completed again after three weeks at a meeting in the lab for post-training test (Lab2). The focus of this paper is limited to the productions of initial /s/ and /z/. 2.3 Perception tests and perception training Perception of the English sibilant contrast was assessed in two separate Two-Alternative Forced-Choice (2AFC) identification tasks in which L1 Danish participants listened to CV and VC syllables, respectively. Token variability was introduced by two different talkers and by various vowel contexts: initial [si], [sa], [su], [zi], [za], [zu], and final [is], [as], [us], [iz], [az], [uz]. In each task participants were asked to determine which sibilant began or ended the syllable. Response buttons were labelled orthographically with “S” and “Z”, and no feedback was provided in the tests. Participants in the experimental group were then trained to perceive onset /s/ and /z/ in a similar identification task, using the same stimuli, but in which corrective feedback was provided after each trial. Trainees were asked to complete a total of 10 training sessions, each of 120 trials, in the course of 3–4 weeks after which the post-training test (identical to the first test) was administered in our lab. 4 15 monosyllabic word pairs beginning with /s/ or /z/; 8 word pairs ending with /s/ or /z/ Rasmussen Word-initial voicing of /z/, impact of training 88 2.4 Native listener evaluations Non-native productions of /s/ and /z/ were assessed auditorily by 40 native English judges who were recruited in or near Aarhus, Denmark. Target words were cut from the carrier phrase and presented to listeners who identified the initial sibilant orthographically as either “S” or “Z” in a 2AFC task. Each listening session included 30 English words recorded by three of the L1 Danish participants at Lab1 and Lab2 (30 words x 3 speakers x 2 recordings = 180 items) presented in randomised order, and half of the judges agreed to complete a second session on another day, on which they rated tokens from 3 other speakers. Productions from each of the Danish speakers were evaluated by 4 different native judges, which means that evaluations of each speaker’s items are based on the mean identification scores from four listeners. A calculated Intraclass Correlation Coefficient (ICC) of 0.84 indicated good reliability between native listener judges in the production evaluation experiment.5 Judges reported using English 50–100% in their daily lives (M = 88.5, SD = 15.3), and they had on average spent eight years in a non-English speaking country (SD = 8.15). 3. Results The native listeners’ evaluations of sibilant productions were computed in R studio (R Core Team, 2021) using the Tidyverse package (Wickham et al., 2019). Trials for which the target fricative was heard as intended received the score of 1 while trials in which the sibilant was not identified received the score of 0. Average evaluation scores for each participant were calculated as proportions. Sibilant productions from all 45 native Danish speakers were on average identified as intended in 65.8% of all instances when recorded before perceptual training. These evaluation results were based on 81.4% productions of intended /s/ and 50.3% productions of intended /z/ as summarised in Table 2. About half of the recorded tokens of intended /z/ produced before perception training were identified as /s/ by native listeners. Thirty-three participants trained identification of the /s – z/ contrast auditorily. In the remaining analysis and visualisations, data from the TG and CG are separated and compared. Table 2 Native English (NE) Listeners’ Identification of Intended /s/ and /z/ by all Danish (DA) Speakers before Training Fricative intended by DA speaker Identified by NE /s/ /z/ /s/ 81.4 % 49.7 % /z/ 18.6 % 50.3 % 5 The ICC estimates were calculated using R statistical software (R Core Team, 2021) with the lme4 package (Bates et al., 2015), based on a mean-rating (k = 4), absolute-agreement, one-way random effects model. Rasmussen Word-initial voicing of /z/, impact of training 89 While identification scores of the intended /s/ productions are consistently high for both groups at both times of recording (approximately 80% for trainees at Lab 1 and Lab 2; 85–86% for controls), the /z/ productions are less consistently heard as intended. Boxplots in Figure 1 furthermore reveal great variation among the evaluations of speakers’ recordings of initial /z/ for trainees and controls alike. A logistic mixed-effects model was run using the lme4 package (Bates et al., 2015) to examine the effect of training on the evaluation of sibilant production. Three factors were treatment coded, i.e., the experiment group (TG/CG) with TG as the reference, production session with Lab1 as the reference and intended fricative with “z” as the reference. From Lab1 to Lab2 the evaluated score for /z/ was significantly higher for trainees’ production, β = 0.93, z = 5.85, p < 0.001, and a post-hoc z-test comparison using emmeans (Lenth, 2025) with Tukey HSD corrections for multiple comparisons shows that that was not the case for the production of /z/ by the control group, z = -1.99, p = 0.49. The interaction between production sessions and the two groups was not statistically significant, β = -0.34, z = -1.02, p = 0.31, suggesting that it is uncertain if the trainees’ production actually does become more intelligible than those of the controls. Figure 1 Intelligibility of Sibilants as Judged by Native English listeners Note. Evaluations (EVAL) of intelligibility for the Training Group – TG (left) and Control Group – CG (right) 3.1 Individual variation Individual intelligibility scores for the 33 speakers in the TG were also calculated before and after training (see Appendix for overview). Figure 2 depicts the individual scores for each speaker, and it clearly demonstrates the varied nature of the data. Results from five selected speakers are highlighted (in yellow) to illustrate this point. For example, the /z/ productions of one speaker (S29) were very consistently identified at both times of testing (top yellow line), while /z/ tokens from another speaker (S02) were hardly ever heard as intended (bottom yellow line). Productions of three speakers improved by at least 35 percentage points after training as represented in the steeper slopes in Figure 2. Intelligibility scores from these five selected speakers are also summarised in Table 3. Rasmussen Word-initial voicing of /z/, impact of training 90 Figure 2 Proportion of Correctly Identified /z/ Items Note. Intelligibility scores for each individual participant in TG before and after training. The highlighted (yellow) lines represent results from five example speakers. Table 3 Intelligibility Scores percentage for /z/ Speaker number Lab1 score Lab2 score Gain score S29 98.3 100 1.67 S30 56.7 95 38.3 S08 45 80 35 S01 1.69 48.3 46.7 S02 1.67 1.67 0 3.2 Exploring the acoustic properties of /z/ To further probe the nature of successful /z/ productions, tokens from the five highlighted speakers in Figure 2 were analysed for duration of and periodicity in the fricative portion. The three speakers with the highest gain scores (S30; S08; S01) were those whose tokens were most improved in terms of target like pronunciation. Tokens from the two speakers with consistently highest (S29) and lowest (S02) intelligibility scores were also included. Altogether, acoustic measurements were obtained for 150 tokens (15 /z/ words x 2 times of recording x 5 talkers) by selecting the fricative portion and using the voice report in Praat (Boersma & Weenink, This chapter is based on the oral presentation given by the author at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons.org/licenses/by/4.0/ . Richter, K. (2025). Prepared to rescue Cinderella? Exploring Austrian EFL student teachers’ beliefs about pronunciation teaching and learning. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 97–108). Université Grenoble-Alpes. https://doi.org/10.5281/zenodo.16696158 Prepared to rescue Cinderella? Exploring Austrian EFL student teachers’ beliefs about pronunciation teaching and learning Karin Richter University of Vienna Abstract Over the past decade, there has been a growing interest in researching EFL pronunciation learning and teaching. Despite this increased scholarly attention, pronunciation remains a challenging area for many EFL teachers, largely due to the lack of specialised pedagogy courses in teacher education. This paper highlights findings from a study on the views of 78 EFL student teachers at the University of Vienna, Austria. Utilising a quantitative approach, the research explores student teachers’ beliefs about their perceived pronunciation proficiency and teaching competence. The results reveal that while most participants are satisfied with their own pronunciation skills, about half of them lack confidence in their ability to teach pronunciation effectively. Furthermore, there is a clear demand for a dedicated didactics course on pronunciation instruction, as the current curriculum prioritises other linguistic areas. Although an existing pronunciation class ostensibly helps improve their confidence in their own phonological competence, it does not adequately prepare them for teaching pronunciation. These findings confirm a need for teacher educators to offer pronunciation teaching courses within EFL training programs, ensuring future educators are equipped to address this often neglected area. Keywords: EFL teacher education, teacher beliefs, English pronunciation teaching, explicit pronunciation instruction Richter Austrian EFL student teachers’ pronunciation beliefs 98 1. Introduction Over the past decade, numerous researchers have advocated for a renewed emphasis on EFL pronunciation teaching, a field previously marginalised and metaphorically described as the “Cinderella” of language teaching (Celce-Murcia et al., 1996, p. 323) or “the neglected orphan” (Deng et al., 2009, p. 1; Gajewska, 2021, p. 20). This growing recognition of pronunciation as a legitimate research discipline has led some linguists to observe that Cinderella has seemingly transformed into the “Belle of the Ball” (Derwing, 2019, p. 27). However, the question arises whether intensified scientific attention translates into tangible improvements in pronunciation teaching practices in EFL classrooms. Indeed, this shift does not yet seem to have reached its happily-ever-after stage. As Levis (2021) points out, there remains a noticeable gap between research and practice, with many interesting studies lacking clear implications for teaching and practically-oriented publications showing insufficient grounding in research. In other words, pronunciation pedagogy continues to struggle to keep pace with advancements in research (Kirkova-Naskova, 2023). Recent studies have confirmed that pronunciation still remains a challenging area for numerous EFL teachers (e.g., Schaefer, 2023; Tegnered & Rentner, 2021). Commonly cited among the core reasons for this neglect is the scarcity of pronunciation pedagogy courses in teacher education programmes (e.g., Baran-Łucarz, 2022; Burri, 2015; Yokomoto, 2016). This is due to many universities still relying on general linguistics courses rather than offering targeted instruction on how to effectively teach pronunciation. 2. The pronunciation teaching paradox Pronunciation is without doubt a fundamental aspect of language competence, playing a pivotal role in both intelligibility and communicative effectiveness in EFL contexts (Celce-Murcia et al., 2010; Jenkins, 2000). Despite clear evidence supporting the benefits of explicit pronunciation teaching and its positive impact on student performance (e.g., Komar, 2019; Murphy & Baker 2019; Rao, 2019; Saito & Plonsky, 2019), many practitioners continue to bypass pronunciation in their lessons (Darcy, 2018; Yoshida, 2016), a phenomenon Darcy (2018) refers to as the “pronunciation teaching paradox” (p. 16). Research has consistently demonstrated that many EFL teachers confess to being reluctant to teach pronunciation (e.g., Henderson et al., 2012; Murphy, 2018; Schäfer, 2023). Even experienced teachers frequently express unease when correcting or evaluating their students’ pronunciation skills, with some avoiding pronunciation instruction altogether (Baker & Burri, 2016; Nangimah, 2020). This resistance is particularly strong among novice teachers, who, confronted with the complexity of diverse classrooms and evolving curricula, tend to view pronunciation as the least important aspect of language teaching (Baker & Burri, 2016). The literature often attributes this reluctance to the absence of emphasis in curricula (Jenkins, 2000), inappropriate teaching materials (Baker & Murphy, 2018), tight schedules and rigid syllabi (Levis & Sonsaat, 2019), as well as a lack of training and confidence (Henderson et al., 2012; Murphy, 2014; Murphy & Baker, 2015). Against this backdrop, studies by Baker (2011), Burri (2015), or Yokomoto (2016) have shown that EFL teachers’ personal experiences with pronunciation learning and teaching greatly impact their confidence in teaching pronunciation, as well as their beliefs about phonological instruction in general. Despite the crucial role these experiences play, many EFL teacher education programs fail to provide adequate training in this respect (Baran-Łucarz, 2022; Derwing, 2019; Murphy, 2014), often leaving teachers ill-prepared to address pronunciation issues in their classrooms. This deficiency in teacher preparation can perpetuate ineffective pronunciation teaching practices, which in turn, hampers learners’ progress in Richter Austrian EFL student teachers’ pronunciation beliefs 99 developing pronunciation skills. Conversely, a number of empirical studies have found that a dedicated pronunciation pedagogy course can in fact significantly boost EFL teachers’ confidence, leading them to address pronunciation more systematically in their classrooms, compared to those without targeted training (e.g., Baran-Łucarz, 2022; Nagle et al., 2018; Kochem & Levis, 2022). It is crucial to consider the views of future EFL teachers to tailor teacher education programmes to the needs of 21st language learners, and prevent pronunciation from being relegated to the sides or abandoned in English language classrooms. Given these concerns, it is imperative to explore the role of pronunciation in EFL teacher education more thoroughly. This research aims to shed light on the effects an explicit pronunciation course has on student teachers’ perceptions of pronunciation instruction. Unlike many European EFL teacher training programs, which often lack a focus on practical phonetics (Henderson et al., 2012), the Department of English and American Studies at the University of Vienna offers an explicit pronunciation course as part of its Bachelor of Education (BEd) programme. At present, no courses are available that specialise in pronunciation pedagogy. This prompts the question if a single language competence class is sufficient to build students’ confidence in teaching pronunciation while, at the same time, equipping them with a comprehensive repertoire of teaching tools and methods to effectively address pronunciation in their future classrooms. 2.1 Research questions Grounded in the premise that student teachers’ views and experiences can exert a significant impact on their prospective pedagogical practices, this study aims to investigate the beliefs advanced student teachers of English at the Department of English and American studies hold about pronunciation teaching and learning. To this end, the following three main research questions were formulated: RQ1: What do EFL student teachers think about their own pronunciation skills? RQ2: How confident do EFL student teachers feel about their pronunciation teaching skills? RQ3: What are EFL student teachers school-related experiences with explicit pronunciation instruction? 3. Research context and methodology 3.1 EFL teacher education at the University of Vienna As part of the Bachelor of Education programme, students enrolled in the teacher education programme at the Department of English and American Studies at the University of Vienna are required to complete an explicit pronunciation course. This course, Practical Phonetics and Oral Communication Skills (PPOCS), is designed to assist students in developing their own pronunciation skills (Richter, 2021). In addition to a weekly 90-minute interactive class with a lecturer, the course also includes a 90-minute compulsory language lab session every week for further practice. A closer look at the syllabus for PPOCS 1 reveals that each session focuses on a small number of specific segmental and suprasegmental features, which are practised productively and receptively. As far as pronunciation teaching pedagogy is concerned, pronunciation is addressed only marginally, if at all, in introductory EFL teaching courses. Richter Austrian EFL student teachers’ pronunciation beliefs 100 3.2 Participants The student teachers (N =78) who participated in this study were all enrolled in the Master of Education programme at the Department of English and American Studies at the University of Vienna. Data was collected from the winter semester 2021/22 to the winter semester 2022/23. The students were informed about the purpose of the study, and their participation was voluntary, ensuring adherence to ethical research guidelines. All the participants had taken PPOCS as part of their BA studies. At the time of data collection, 40% of the respondents were already teaching English at school or another educational institution. Three quarters (76%) of the respondents were female. This 3:1 ratio largely corresponds to the average gender distribution in the teacher education programme at the department. The average age of the participants was 24, and the level of English language competence the participants are expected to have reached at this stage is C1, according to the Common European Framework of Reference (Council of Europe, 2001). In order to ensure confidentiality, participants were coded as R1 through R78. 3.3 Instruments In order to elicit the student teachers’ views, a quantitative approach was adopted. A carefully constructed online survey with closed and open questions was administered using Google Forms. Participation was optional and anonymous. The online questionnaire encompassed items regarding students’ biographical data, their perception of their own pronunciation skills, and their views on and experiences with pronunciation teaching. The majority of the questions consisted of 4-point Likert-scale items. For some of these items, open questions asking the students to provide reasons for their choices were given. However, not all the respondents added verbal comments in the provided spaces. 4. Data analysis and results The collected data was analysed using Excel. Participant responses were analysed descriptively using mean scores, percentages, and frequencies. For the qualitative part, responses to openended items were grouped thematically through content analysis. The thematic analysis involved manually identifying, coding, and grouping responses into key categories based on recurring patterns and concepts. These categories, guided by the research questions, included participants’ positive and negative perceptions and experiences, challenges encountered, and suggestions for improvement. 4.1 Student teachers’ own pronunciation skills In order to address the first research question pertaining to the student teachers’ views about their own pronunciation skills, two items from the questionnaire were considered. Firstly, they were asked to indicate the extent to which they agreed with the statement that EFL teachers need to ensure that they have good pronunciation skills. Figure 1 shows that the majority of all the respondents either strongly agreed (52%) or slightly agreed (42%). The participants were also asked to rate their satisfaction with their own pronunciation on a 4-point scale with the options: 1 (very unhappy) to 4 (very happy). As Figure 2 shows, the participants’ overall satisfaction with their own pronunciation was very high. The majority of the respondents indicated that they were very happy (35.9%) or moderately happy (48.7%). Less than 16% chose the option ‘slightly unhappy’. Not a single respondent opted for the option ‘very unhappy’. Richter Austrian EFL student teachers’ pronunciation beliefs 101 Figure 1 Importance of Good Pronunciation Skills for Teachers Figure 2 Satisfaction with Own Pronunciation Skills 0,0% 10,0% 20,0% 30,0% 40,0% 50,0% 60,0% Strongly agree Slightly agree Slightly disagree Strongly disagree EFL teachers have to make sure they have good pronunciation skills. 35,9% 48,7% 15,4% 0,0% How happy are you with your pronunciation at the moment? very happy moderately happy slightly unhappy very unhappy Richter Austrian EFL student teachers’ pronunciation beliefs 102 To gain further insight into the data, the verbal comments some of the students provided were examined (n = 57). Among those who claim that they are happy with their pronunciation, most responses included references to the PPOCS course: “PPOCS really developed my pronunciation skills” (R9) or “I have learned during PPOCS how to improve my skills” (R24). At the same time, many respondents acknowledged room for improvement. as noted in the comments: “It could be better if I had more pronunciation classes and feedback” (R35) or “I wish I spent more time practising and getting better” (R72). 4.2 Student teachers’ confidence to teach pronunciation For the second research question about participants’ views regarding their abilities to teach pronunciation, responses to the item How do you feel about your pronunciation teaching skills were analysed. The participants chose their answers on a 4-point scale from 1 (very insecure) to 4 (very confident). As Figure 3 shows, the data displays a clear divide. Slightly more than half of the participants feel very confident (9%) or quite confident (46.1%), whereas 41% report feeling rather insecure, and 3.9% very insecure. Figure 3 Pronunciation Teaching Skills As far as their verbal comments are concerned, many respondents (n = 47) claim to have learned how to teach pronunciation from the PPOCS course: “[...] because of the teaching/didactic strategies I have acquired through PPOCS” (R9) or “[...] because I went through PPOCS” (R34). At the same time, many of the student teachers point out that they lack the necessary skills, as shown in the following response: “We never learnt how to teach it. Where should I start? Should I introduce the IPA alphabet at first or what and when do I teach pronunciation?” (R5). Another participant cited the lack of pronunciation teaching content in general: “This has never been the focus of any of the courses at uni” (R53). Interestingly, among those who express insecurity, quite a few assert that they themselves experienced problems in the PPOCS course: “I don’t have any experience with doing so and struggled with 9,0% 46,1% 41,0% 3,9% How confident do you feel about your pronunciation teaching skills? very confident quite confident rather insecure very insecure Richter Austrian EFL student teachers’ pronunciation beliefs 103 pronunciation myself” (R59). A further theme which emerged from the data relates to teachers’ views on what pronunciation instruction actually entails, as illustrated in the comments: “I feel confident that I can teach pronunciation when mistakes come up” (R33) or “Teaching pronunciation at school would mostly be done indirectly. Focusing on words that are difficult to pronounce, as they come up and focus on specific sounds that way” (R69). These remarks reveal a common misperception according to which pronunciation teaching solely consists of teachers correcting mispronounced words. 4.3 Student teachers’ experience with explicit school-based pronunciation teaching The third research question intended to gauge the participants’ experiences with explicit pronunciation teaching in a school context. To this end, three different items referring to their experience as learners at school, as observers during school practice, and as teachers were analysed. As can be seen in Figure 4, a staggering 80% of the student teachers report that they did not have any experiences with explicit pronunciation teaching when they were learners at school. Eighty-three percent assert that they had not observed any lessons with explicit pronunciation teaching elements delivered by experienced teachers in the course of their studies. Although 33% claim that they themselves have taught pronunciation explicitly (in the course of their school-based practica), the percentage of those who have not is still very high (almost 70%). Most of the comments related to this item refer to the lack of training, for instance: “No experience; no class taught me how to do that” (R53) and “I have the necessary background knowledge (theory) but am missing the didactic knowledge on how to transmit it in a fun way to the students” (R34). Figure 4 Experience with Pronunciation Teaching 0,0% 10,0% 20,0% 30,0% 40,0% 50,0% 60,0% 70,0% 80,0% 90,0% as a learner at school as an observer during teacher training as a teacher Experience with pronunciation teaching yes no can't remember Richter Austrian EFL student teachers’ pronunciation beliefs 104 In order to find out whether the student teachers feel the need for more support regarding pronunciation teaching, they were asked whether they would be interested in taking a course focusing on pronunciation pedagogy. Figure 5 shows that approximately 80% of the participants expressed an interest in the course (58% definitely, 28% yes, maybe). Some of the comments in favour of such a course provide further insight into the student teachers’ views, with many of them pointing at the crucial role of pronunciation, for instance: “It is an important skill to have. Pronunciation makes or breaks a language. Learning effective methods on teaching pronunciation could be useful for my future teaching career” (R10) or “I think I could benefit from activities or games that can be used in class to also motivate my students. I have also had the experience that especially younger learners find it fun to mimic the teacher and explore pronunciation (R37)”. Figure 5 Interest in a Pronunciation Teaching Course Among the 14% (11% I don’t think so, 3% certainly not) who claimed not to be interested in such a course, the opinion that pronunciation is not as important as other skills prevailed: “In my opinion, there are more important things than teaching pronunciation. You can survive without the perfect pronunciation” (R77). A preference for other language skills could also be noted in the responses, for instance: “It might be interesting, but in comparison to teaching other skills, such as vocabulary, speaking etc., I don’t think it’s as important” (R55). Another theme that emerged relates to how some student teachers question the necessity of teaching pronunciation, given the learners’ frequent exposure to online media, as illustrated in the following observation: “Because I don’t believe it’s really that necessary, most students speak with perfect pronunciation due to computer games, TV shows or podcasts anyways” (R62). 0,0% 10,0% 20,0% 30,0% 40,0% 50,0% 60,0% 70,0% yes, definitely Yes, maybe I don't think so Certainly not If the English department offered a pronunciation pedagogy course, would you be interested? Richter Austrian EFL student teachers’ pronunciation beliefs 105 5. Discussion of results and key findings This study set out to explore EFL student teachers’ perceptions of and experiences with pronunciation instruction. The findings show that the overwhelming majority of the respondents are satisfied with their own pronunciation abilities after completing a single explicit pronunciation course as part of their undergraduate programme. However, this contrasts sharply with the fact that almost half of them report lacking confidence in teaching pronunciation. Given that the respondents are nearing the completion of their academic education, this deficit is concerning, though not entirely unexpected. In fact, similar results were observed in a large-scale study conducted by Henderson et al. (2012), which surveyed 640 practising EFL teachers across eight different European countries. They also found that their participants rated their own pronunciation proficiency as favourable but their pronunciation teaching skills as poor. This trend seems to persist even more than a decade later, not only in Europe but also in other EFL contexts such as Uruguay (Couper, 2016) or Brazil (Buss, 2016) where teachers do not feel confident teaching pronunciation due to a lack of skills and knowledge. The ambivalence found in the students’ appreciation of the explicit pronunciation class was useful but additional training would have been even more beneficial, as observed in other research studies (e.g., Gordon & Barrantes-Elizondo, 2024). Another important observation from the data is the participants’ reported lack of experience with explicit school-based pronunciation instruction. The majority of the student teachers have had very little to no experience with pronunciation teaching as learners, observers, or teachers. About one third of them, however, reported some experience teaching pronunciation in an EFL classroom. The admittedly limited confidence to conduct such lessons without formal training might perhaps stem from the skills and the knowledge gained in a compulsory language competence course focusing on practical phonetics, or from a misconception of pronunciation teaching as merely correcting mispronounced words on the spot (Foote et al., 2016). This limitation to corrective feedback in the form of recasts has also been observed in a number of other studies (Foote et al., 2016; Murphy, 2011; Nguyen & Newton, 2020; Schäfer, 2023). On a more positive note, the vast majority of students expressed interest in a pronunciation pedagogy course should it be offered as part of their university studies. Other empirical investigations have also highlighted the need for specialised pronunciation training in teacher education programs to adequately prepare student teachers for effective pronunciation instruction in L2 classrooms. For instance, Burri (2015) demonstrated that targeted pronunciation instruction significantly impacts teachers’ confidence and effectiveness in the classroom, and thereby positively shapes their attitudes toward pronunciation teaching. Similar results have been found in other studies (Baker, 2014; Kochem, 2022) and even in professional development types of training (e.g., Nguyen & Newton, 2021). Taken together, these results corroborate earlier research demonstrating that pronunciation teaching still does not receive the attention it deserves in today’s EFL teacher education. Thus, these findings underscore the need for a systematic re-evaluation of teacher training programs to ensure that pronunciation is integrated as a core component. 6. Conclusion and implications The objective of the present study was to shed light on whether EFL student teachers enrolled at the Master of Education programme at the University of Vienna feel adequately prepared to rescue Cinderella and teach pronunciation in their (future) classrooms. The findings indicate that an explicit language competence course dedicated to enhancing the students’ own pronunciation skills can have a positive impact on their self-confidence regarding their own Richter Austrian EFL student teachers’ pronunciation beliefs 106 pronunciation skills. However, it does not result in a correspondingly high level of confidence in their ability to teach pronunciation. Though limited in their focus on a very specific group of student teachers, the insights gained in this project remain valuable in reaffirming a key trend: the persistent gap between research and practice. To ensure high-quality pronunciation instruction in EFL classrooms, student teachers need to be equipped with the methodological expertise and the necessary tools in order to help L2 learners to perceive and produce a new phonological system. Teacher educators have to realise that a lecture on phonetics and phonology or a single explicit pronunciation class does not inherently result in the seamless transformation of this knowledge into actual teaching practices. By making pronunciation instruction an obligatory component of EFL teacher training education, future generations of teachers will be better equipped to deliver effective pronunciation instruction, leading to improved learning outcomes and greater overall proficiency in English as a foreign language. Such a course for student teachers could then help lay the groundwork for future generations of EFL instructors to develop the confidence needed to teach a skill as vital to communicative language competence as pronunciation. References Baker, A. (2011). ESL teachers and pronunciation pedagogy: Exploring the development of teachers’ cognitions and classroom practices. In J. Levis & K. LeVelle (Eds.), Proceedings of the 2nd pronunciation in second language learning and teaching conference (pp. 82–94). Iowa State University. https://apling.engl.iastate.edu/wpcontent/uploads/sites/221/2015/05/PSLLT_2nd_Proceedings_2010.pdf Baker, A. (2014). Exploring teachers’ knowledge of second language pronunciation techniques: Teacher cognitions, observed classroom practices, and student perceptions. TESOL Quarterly, 48(1), 136– 163. https://doi.org/10.1002/tesq.99 Baker, A., & Burri, M. (2016). Feedback on second language pronunciation: A case study of EAP teachers’ beliefs and practices. Australian Journal of Teacher Education, 41(6), 1–19. https://doi.org/10.14221/ajte.2016v41n6.1 Baran-Łucarz, M. (2022). ‘The course should be obligatory!’: Attitudes of Polish future EFL teachers towards a course on pronunciation teaching. In J. Levis & A. Guskaroska (Eds.), Proceedings of the 12th Pronunciation in Second Language Learning and Teaching Conference (pp.1-10). Iowa State University. https://doi.org/10.31274/psllt.13260 Burri, M. (2015). Student teachers’ cognition about L2 pronunciation instruction: A case study. Australian Journal of Teacher Education, 40(10), 66–87. https://doi.org/10.14221/ajte.2015v40n10.5 Buss, L. (2016). Beliefs and practices of Brazilian EFL teachers regarding pronunciation. Language Teaching Research, 20, 619–637. https://doi.org/10.1177/1362168815574145 Celce-Murcia, M., Brinton, D., & Goodwin, J. (1996). Teaching pronunciation: A reference for teachers of English to speakers of other languages. Cambridge University Press. Celce-Murcia, M., Brinton, D. M., & Goodwin, J. M. (2010). Teaching pronunciation: A course book and reference guide (2nd ed.). Cambridge University Press. Council of Europe. (2001). Common European framework of reference for languages: Learning, teaching, assessment. Cambridge University Press. Couper, G. (2016). Teacher cognition of pronunciation teaching amongst English language teachers in Uruguay. Journal of Second Language Pronunciation, 2(1), 29–55. https://doi.org/10.1075/jslp.2.1.02cou Darcy, I. (2018). Powerful and effective pronunciation instruction: How can we achieve it? The CATESOL Journal, 30(1), 13–45. https://doi.org/10.5070/B5.35963 Deng, J., Holtby, A., Howden-Weaver, L., Nessim, L., Nicholas, B., Nickle, K., Pannekoek, C., Stephan, S., & Sun, M. (2009). English pronunciation research: The neglected orphan of second language Rupp & Henderson Pronunciation features noticed by Vietnamese MOOC users 113 English speaker who participated in the mentor programme of the MOOC.4 The coding was documented in a code book. 4. Data analysis and results In the 149 selected comments, Vietnamese MOOC users express challenges or concerns regarding their English pronunciation. Some of these comments (N = 109) were prompted. For example, in an exercise on long and short vowels, users are asked: Does your native language have long and short vowels?, or in an exercise on word endings: Do you omit word endings? Other comments were not prompted (N = 40), e.g., at the beginning of the MOOC, users are asked: Which English pronunciation features are difficult for you? and What are your needs regarding English pronunciation? Most of the unprompted comments (28/40) were about suprasegmental features. Table 1 provides an overview of the features which Vietnamese speakers commented upon: Table 1 Features Mentioned by Vietnamese EPGW Users Type of feature English feature # of mentions Total # of mentions Segmental vowel length; including in diphthongs and clipping 27 87 consonants, e.g., /f/, /ʃ/, /ʒ/, /d ʒ/, /θ, ð/, /p, b/ 21 vowels, e.g., NURSE 19 reduction of consonant clusters 17 rhoticity 3 Suprasegmental word stress 31 62 intonation 16 connected speech 15 Segmental and suprasegmental features are listed, followed by the number of mentions for each feature. The fact that both segmental and suprasegmental features were commented upon may be treated as evidence of salience, with segmentals being slightly more frequently mentioned. Across the two types of features, vowel length and word stress are mentioned a similar number of times (27 and 31 mentions, respectively) and thus seem more salient to these 4 For an outline of this programme, see Rupp and Henderson (forthcoming). Rupp & Henderson Pronunciation features noticed by Vietnamese MOOC users 114 EPGW users. Rhoticity, in contrast, is so rarely mentioned (3 times) that it does not seem salient to them. We analysed the comments to determine whether they constituted evidence of one or the other type of salience. For example, one Vietnamese enrolee is aware that their pronunciation difficulties affect how successfully they communicate: “I have difficulty in connected speech, speaking unnaturally and making it hard to understand what I am saying” (VMa23ad). They refer to a specific linguistic feature of English but they also assess their own speech as unnatural and difficult to decipher. Another enrolee expresses awareness of a norm, of true pronunciation: “I can speak sentences with true pronunciation but no intonation” (VMdb4d8). They claim their achievement – with ‘true’ presumably meaning being able to pronounce well at the segmental level or correctly in relation to a standard – but then they also state a perceived shortcoming (i.e., no intonation). Both types of comment seem to hint at language-external, i.e., sociolinguistic, salience: how their co-locutors potentially judge their speech and the imposition of a norm. Note in this relation that many of the features mentioned by the enrolees have been earmarked and researched from a contrastive analysis or a pedagogical perspective, hinting at sociolinguistic salience in educational practices. Studies of typical difficulties faced by Vietnamese English users have argued that vowel length is challenging for their intelligibility (Cunningham, 2009a, 2009b). The monophthongisation of diphthongs has also been described as a common feature (Swan & Smith, 2001). Single consonants, too, may be modified, as pointed out by Cunningham (2009a): “… final /f/ as in the word if is often pronounced as [ip]. This pronunciation is viewed as characteristic of Vietnamese accented English in Vietnam – teacher and thus learner awareness of this is high, and the feature is stigmatised” (p. 3). Consonants in final position can also be particularly difficult (Ngo, 2011; Nguyen, 2020) leading to the elision or the simplification of clusters, perhaps in part because Vietnamese does not allow consonant clusters in any position (Avery & Ehrlich, 1992). In terms of suprasegmentals, Vietnamese is a tonal language, where pitch carries lexical or grammatical meaning, i.e., distinguishes words or their inflections, whereas English can use word stress or the addition of another word such as this and that instead of a tone (Tang, 2007, 11). Enrolees’ comments on word stress (n = 31) constitute half of all comments on suprasegmental features. These comments seem to resonate with adult Vietnamese ESL speakers in a study by Zielinski (2006), whose non-standard word stress patterns made them less intelligible to three native Australian listeners. For example, as Cunningham (2009a) notes, “the use in English of the rising sac tone on syllables that have a voiceless stop in the coda … can result in a pitch prominence that may be interpreted as stress by listeners” (p. 2). Connected speech is another area where Vietnamese and English differ. English exhibits varied sandhi variation in all registers and speech rates, with widespread linking, assimilation and dissimilation, deletion, and epenthesis. In contrast, Vietnamese-accented English rarely exhibits any such features (Cunningham, 2009b). 5. Discussion of results and key findings In our search for evidence of sociolinguistic salience in Vietnamese EPGW enrolees, three key findings emerge: 1) Vietnamese EPGW enrolees post unprompted comments, so they do have some awareness of particular English pronunciation features; 2) Both segmental and suprasegmental features are commented upon; Rupp & Henderson Pronunciation features noticed by Vietnamese MOOC users 115 3) Given 1), a MOOC has potential as a suitable tool for enquiring into the salience of pronunciation features. The results of our study demonstrate that Vietnamese enrolees in EPGW show awareness of a range of both segmental and suprasegmental English pronunciation features, with a slight majority of the unprompted comments (28/40) being about suprasegmental features. They flag both as features that are either challenging for them or which they wish to develop in their English pronunciation. Does this awareness stem from linguistic or sociolinguistic salience? Two issues arise here. The issue with linguistic salience is that whilst relative distance in pronunciation may be quite straightforwardly measured for vowels, it is not straightforward for suprasegmental features such as stress and intonation. How does one measure the distance in stress between a monosyllabic language like Vietnamese and a stress-timed language like English? Or between a tonal language and a non-tonal language? An issue with sociolinguistic salience is that enrolees in EPGW are not asked about the reasons for flagging the English pronunciation features that they do. Is it because they were told about these features in their education? Our survey of the research literature on Vietnamese English speakers suggests that suprasegmental features are addressed to a lesser extent than segmental features. However, Ha Noi Open University provided us with the syllabi of two English Phonetics and Phonology courses that they run, which showed that they pay attention to segmental and suprasegmental features for an equal number of weeks. Perhaps the Vietnamese enrolees have had particular social experiences regarding the English pronunciation features that they mentioned in their daily lives; because the learning units in EPGW do not ask enrolees for the motivation of their comments, we cannot tell. Nonetheless, the fact that Vietnamese enrolees flagged the range of features observed reveals that these are important to them, regardless of whether the attached importance is due to linguistic or sociolinguistic salience. 6. Conclusion and implications/applicability The solution for probing sociolinguistic salience in EPGW would seem to be to explicitly ask enrolees to motivate their comments regarding English pronunciation features that they find challenging or would like to develop in their English pronunciation. This would work best in steps at the beginning of EPGW, when enrolees have not yet engaged with any steps that address particular pronunciation features (which may prompt them to post a comment about those features). More broadly, we would like to advocate an approach to English pronunciation teaching that takes account of and addresses English pronunciation features that may have social value for people developing their English pronunciation. Knowledge of language-internal factors (linguistic salience) constitutes a sound foundation for pronunciation teachers, yet they should also be knowledgeable about the language-external factors (sociolinguistic salience) which impact upon their learners. Teachers may prefer to not prioritise a feature if it contributes little to intelligibility, e.g., the English dental fricatives, yet learners may be particularly motivated to work on such features in order, for example, to avoid stigmatisation. Pronunciation teachers could take the opportunity to expertly balance their objectives with those of their learners. Following Boswijk and Coler (2020, p. 714), salience is not a static property of phonetic forms or an inherent attribute of units of this kind. The degree of salience of a form can differ between speakers. Rupp & Henderson Pronunciation features noticed by Vietnamese MOOC users 116 References Auer, P., Barden, B., & Grosskopf, B. (1998). Subjective and objective parameters determining ‘salience’ in long-term dialect accommodation. Journal of Sociolinguistics, 2(2), 163–187. https://doi.org/10.1111/1467-9481.00039 Avery, P., & Ehrlich, S. (1992). Teaching American English: Oxford handbook for language teachers. Oxford University Press. Boswijk, V., & M. Coler (2020). What is salience? Open Linguistics, 6(1), 713–22. https://doi.org/10.1515/opli-2020-0042 Cunningham, U. (2013). Teachability and learnability of English pronunciation features for Vietnamese-speaking learners. In E. Waniek-Klimczak & L. R. Shockey (Eds.), Teaching and researching English accents in native and non-native speakers (pp. 3–14). Springer. Cunningham, U. (2009a). Quality, quantity and intelligibility of vowels in Vietnamese accented English. In Ewa Waniek-Klimczak (Ed.), Issues in accents of English II: Variability and norm (pp. 3–20). Cambridge Scholars Publishing. Cunningham, U. (2009b). Phonetic correlates of unintelligibility in Vietnamese-accented English. FONETIK 2009 The XXIIth Swedish Phonetics Conference (pp. 108–111). Stockholm. https://urn.kb.se/resolve?urn=urn:nbn:se:du-4086 Ellis, N. (2018). Salience in usage-based SLA. In S. M. Gass, P. Spinner, & J. Behney (Eds.), Salience in second language acquisition (pp. 21–40). Routledge. Jansen, S. (2014). Salience effects in the Northwest of England. Linguistik Online, 64(4), 91–110. https://doi.org/10.13092/lo.66.1574 Kerswill, P., & Williams, A. (2002). ‘Salience’ as an explanatory factor in language change: Evidence from dialect levelling in urban England. In M. C. Jones & E. Esch (Eds.), Language change (pp. 81–110). De Gruyter. https://doi.org/10.1515/9783110892598.81 Levis, J. M. (2005). Changing contexts and shifting paradigms in pronunciation teaching. TESOL Quarterly, 39(3), 369–377. https://doi.org/10.2307/3588485 Llamas, C., Watt, D., & Johnson, D. E. (2009). Linguistic accommodation and the salience of national identity markers in a border town. Journal of Language and Social Psychology 28(4), 381–407. https://doi.org/10.1177/0261927X09341962 Ngo, P. A. (2011). L1 influence on Vietnamese accented English. Kajian Linguistik dan Sastra, 21(2),108–125. https://doaj.org/article/5f7dd263ef12446d9d25ba3335948b5a Nguyen, D. D. (2020). Common errors in English speaking lessons of second-year English major students at Haiphong Technology and Management University [Unpublished doctoral dissertation]. Haiphong Technology and Management University (Đại học Quản lý và Công nghệ Hải Phòng). Podesva, R. J. (2011). Salience and the social meaning of declarative contours: Three case studies of gay professionals. Journal of English Linguistics, 39(3), 233–64. https://doi.org/10.1177/0075424211405161 Rupp, L. (2019). English Pronunciation in a Global World. [MOOC]. FutureLearn. https://www.futurelearn.com/courses/english-pronunciation/ Rupp, L. & A. Henderson (forthcoming). Exploring the usefulness of MOOC English Pronunciation in a Global World for L2 learning, teaching and research. In E. Vasu & A. Kirkova-Naskova (Eds.), Achievements in second language pronunciation: Good practices for L2 teaching and learning. Cambridge University Press. Rupp, L., Das, A., Kamps A., & Acosta, E. (2023). English pronunciation in a Global World: A MOOC course for pronunciation teaching. In A. Henderson & A. Kirkova-Naskova (Eds.), Proceedings of the 7th International Conference on English Pronunciation: Issues and Practices (pp. 225–235). Université Grenoble-Alpes. hal-04178893 Siegel, J. (2010). Second dialect acquisition. Cambridge University Press. Swan, M., & Smith, B. (Eds.). (2001). Learner English: A teacher’s guide to interference and other problems (2nd ed). Cambridge University Press. Rupp & Henderson Pronunciation features noticed by Vietnamese MOOC users 117 Tang, G. (2007). Cross-linguistic analysis of Vietnamese and English with implications for Vietnamese language acquisition and maintenance in the United States. Journal of Southeast Asian American Education and Advancement, 2(3). https://doi.org/10.7771/2153-8999.1085 Trudgill, P. (1986). Dialects in contact. Blackwell. About the authors Laura Rupp is a Senior Lecturer in English Language and Linguistics and the director of the Centre for Global English at the Vrije Universiteit Amsterdam. She created the online course (MOOC) English Pronunciation in a Global World. Her research expertise is in grammatical language variation and change, and in English pronunciation learning and teaching. She has published internationally on both subjects. Her work has been funded by the British Academy and The Netherlands Initiative for Education Research. Her project Language Stories was granted patronage by the Dutch UNESCO Commission. Email: l.[email protected] Alice Henderson is a Full Professor at Grenoble – Alpes University, France, where she teaches English for Specific Purposes to STEM students and is currently the assistant director of the LIDILEM research group. She taught English phonetics and phonology to English majors for 24 years at another French university and has been involved in teacher training in France, Norway, Poland, and Spain. In 2009 she initiated the bi-annual international conference English Pronunciation: Issues & Practices (EPIP). She has published internationally on English pronunciation learning and teaching, the perception of foreign-accented speech, and English Medium Instruction (EMI). Email: [email protected] This chapter is based on the oral presentation given by the authors at the 8th International Conference English Pronunciation: Issues and Practices (EPIP 8) held May 8–10, 2024 at the University of Cantabria in Santander, Spain. It is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of the license, please go to: http://creativecommons.org/licenses/by/4.0/ . White-Hautzinger, V. et al. (2025). The role of pronunciation in the assessment of listening skills in the Polish Matura exam. In A. Kirkova-Naskova, P. Humánez-Berral, & A. Henderson (Eds.), Proceedings of the 8th International Conference on English Pronunciation: Issues and Practices (pp. 118–128). Université GrenobleAlpes. https://doi.org/10.5281/zenodo.16696495 The effects of explicit pronunciation instruction: A study of advanced Austrian EFL students Valerie White-Hautzinger, Marija Djenadic, Carolin Rumpler, and Miriam Fiala University of Vienna Abstract As English consolidates its role as a globalised language and a crucial medium for communication across nations, pronunciation emerges as a contributing factor to achieving international intelligibility. This, in turn, emphasises the importance of pronunciation teaching for English as a Foreign Language (EFL) learners. Students at the University of Vienna complete the obligatory explicit pronunciation instruction (EPI) course Practical Phonetics and Oral Communication Skills 1 (PPOCS1) and its accompanying language lab, focusing on one of the two offered target varieties. In this course, students receive EPI on both segmental and suprasegmental features. This study examines how such EPI affects 82 second-year EFL university students’ pronunciation, specifically analysing three segmental features (/ɒ – ɑː/, /əʊ – oʊ/, /r/) assessed by four trained raters using a rating scale. Respondents were recorded reading the same text at the beginning and the end of the 14-week course. The auditory analysis of the data reveals a significant shift towards their chosen target variety after EPI, most notably in the consistent use of /oʊ/ in General American English (GAE) and the diphthong /əʊ/ in Standard British English (SBE). Given these findings, this research aims to inspire curriculum designers to recognise the value of integrating EPI into their programmes. Keywords: auditory analysis, explicit pronunciation instruction, segmental features, tertiary EFL programmes White-Hautzinger et al. Explicit instruction & advanced Austrian EFL learners 119 1. Introduction As globalisation turns English into a key tool for communication, pronunciation emerges as a factor for personal and professional success. Research shows that explicit pronunciation instruction (EPI), involving direct structured teaching of pronunciation features, improves intelligibility and reduces listener effort (Darcy, 2018) across various language backgrounds, including Indonesian (Pardede, 2018), Japanese (Saito, 2011), Kurdish (Mahmood, 2023), and Tunisian Arabic (Bouchhioua, 2016). However, there is a research gap regarding how such instruction affects EFL learners from German backgrounds, particularly in approximating specific varieties of English. The existence of this gap is unsurprising as pronunciation instruction itself remains underrepresented in language teaching, particularly in tertiary language programmes (Darcy, 2018). This is especially concerning given that many university students of English aspire to become English teachers themselves, who function as role models for their students but often receive neither a sufficient amount of instruction to improve their own accents nor training to teach pronunciation (Henderson et al., 2015). For those pre-service teachers, clear and structured teaching of phonetic and phonological aspects (i.e., EPI) can be effective for modifying pronunciation. The mandatory Practical Phonetics and Oral Communication Skills 1 (PPOCS1) course at the University of Vienna provides such explicit training in both segmental and suprasegmental features, helping EFL students improve their oral proficiency in either Standard British English (SBE) or General American English (GAE) (Richter, 2021). This study examines how this EPI course impacts the development of three segmental features (/ɒ – ɑː/, /əʊ – oʊ/, and postvocalic /r/) in advanced EFL students who are mainly L1 German speakers (N = 82). Using preand post-recordings, this study provides insights into the effects of EPI on learners with a German language background across two different English varieties. 2. Previous research in the field Previous research highlights the positive impact of EPI on L2 learners’ pronunciation skills. For example, it enhances consistency at both the segmental and suprasegmental levels, along with improving word segmentation skills (Darcy, 2018). Beyond speaking skills, EPI has been found to strengthen auditory language perception (Linebaugh & Roche, 2015), spelling (Prator, 1971), and vowel perception (Ghorbani et al., 2016). These findings highlight the broad benefits of pronunciation instruction, though none focus on specific target varieties, such as SBE and GAE, or German EFL learners. These learners often present distinctive challenges, such as monophthongisation of diphthongs (Mende, 2009) and struggling with vowels rather than consonants (Tatzl, 2011), highlighting the need for targeted research on this learner group. Against this backdrop, various types of pronunciation instruction have been shown to improve advanced EFL learners’ pronunciation skills (Gordon & Darcy, 2022), with some approaches more effective than others. Atar (2018) emphasises the importance of EPI in meaningful contexts, focusing not only on perception but also on production to enhance fluency. To create such a meaningful context, Galante and Piccardo (2022) integrated controlled speech practice, along with weekly audio recordings, extensive feedback, and selfreflections, which proved highly successful. Studies comparing explicit and implicit approaches generally show that explicit instruction can have larger effects on students’ pronunciation. For instance, students who received eight weeks of EPI showed greater improvement in their English pronunciation compared to those who received implicit instruction (Yakut 2020), a finding also supported by Gordon et al. (2013), whose study used EPI in combination with a presentation-practice-production sequence over the course of three weeks. Both of these studies used preand post-recordings of students reading aloud. These White-Hautzinger et al. Explicit instruction & advanced Austrian EFL learners 120 findings suggest that the meaningful and controlled context of EPI could play a crucial role in improving advanced EFL learners’ pronunciation skills. To conclude, while existing research has established the positive effects of EPI across various learner populations and approaches, substantial gaps remain. In particular, learners from German language backgrounds have been largely overlooked, as have studies that focus on approximating specific EPI target varieties of English. 2.1 Research question and hypothesis In light of these gaps, the following research question was investigated: How does the pronunciation of EFL university students of selected segmental features – in particular, /ɒ – ɑː/, /əʊ – oʊ/, and postvocalic /r/ – develop over a semester of EPI? The following hypothesis was formulated: EFL university students show significant improvement in their pronunciation of the selected segmental features after a semester of EPI. 3. Methodology 3.1 Participants and setting The study was conducted at the Department of English and American Studies of the University of Vienna. Participants were fourth-semester bachelor students enrolled in PPOCS1, who are expected to have reached an upper intermediate B2 proficiency level (Council of Europe, 2020). PPOCS1 is a weekly two-hour class offering EPI. Prior to enrolment, students choose between one of the two offered varieties, SBE and GAE. PPOCS1 is complemented by a mandatory weekly two-hour language lab, focusing on the chosen variety. A total of 82 students participated in the study. Of these, 45 students attended the BE PPOCS1 course, while the remaining 37 took the GA PPOCS1 course. 84% (n = 73) of the participants reported German as their first language (L1), 6% (n = 5) were bilingual German speakers, and the remaining 10% (n = 9) had other L1s, including Albanian, Arabic, Bosnian/Croatian/Serbian, Czech, Dutch, Hungarian, Polish, Romanian, and Turkish. Despite varying language backgrounds, all participants spoke German. None of the participants identified English as their L1. 3.2 Pronunciation instruction Participants attended two EPI classes per week (4 hours total), receiving specific English variety instruction on segmental and suprasegmental features. PPOCS1 and the language lab combine theory with practical classroom activities, such as listening to audio guides and producing targeted features, which were shown to be effective practices (Galante & Piccardo, 2022). According to Richter (2021), the two courses aim at developing students’ theoretical and practical knowledge of segmental and suprasegmental features and help them approximate one of the two offered varieties of English. In their design, the courses offer not only hands-on opportunities to practise spoken language communicatively but also regular feedback on the production of segmentals and suprasegmentals (Richter, 2021). 3.3 Data collection Prior to data collection, the testing procedures were thoroughly explained to the students, and any questions were addressed to ensure clarity and comprehension. Participants provided their informed consent to participate in the study and agreed to the processing of the data. At the White-Hautzinger et al. Explicit instruction & advanced Austrian EFL learners 121 beginning of the semester, students completed a questionnaire to provide information on their language background and their chosen variety for PPOCS1. Speech samples were collected at the beginning of the course in week 1 (March 2023, pre-test) and at the end of the course in week 14 (June 2023, post-test) in the language lab. At both testing points, participants were required to record themselves individually reading a text passage aloud. While reading aloud is not an optimal measure of pronunciation change, it facilitates comparability across recordings and a more consistent analysis of the targeted features. Specifically, the text “The North Wind and the Sun” was employed due to its established use in previous research studies (e.g., Gass & Varonis, 1994) and its inclusion of all relevant features under investigation. This study focuses on three segmental features: two contrasts /ɒ – ɑː/ and /əʊ – oʊ/, and the postvocalic /r/, which are recognised as distinguishing characteristics between GAE and SBE and typical challenging features for German speakers (Schmitt, 2016). The text passage includes 30 words (see Table 1) featuring one or more of the targeted phonemes, which were embedded in the story to ensure that students read the words in context rather than in isolation. All speech stimuli were recorded using the software KALTURA1 and saved in MP4 format. A total of 164 recorded samples were analysed (82 at pre-test and 82 at post-test). Table 1 Distribution of 33 Selected Segmental Features Selected segmentals Total number of tokens Example words /ɒ – ɑː/ 6 stronger, off, stopped /əʊ – oʊ/ 7 cloak, closely, fold, so /r/ 20 North, were, stronger, traveller, dirty, first, considered, other, hard, more, warmly 4. Data analysis and results Data were analysed using R version 4.3.3 (R Core Team, 2024) and RStudio version 2024.04.2 (RStudio Team, 2024); specifically, the R packages “stats” (R Core Team, 2024), “rcompanion” (Mangiafico, 2024), and “vcd” (Meyer et al., 2023). 4.1 Raters and rating scale To investigate students’ development in pronunciation after receiving EPI, this study used an auditory analysis, as in prior research (e.g., Saito, 2011; Zhang & Yuan, 2020). During the 1 KALTURA https://corp.kaltura.com/ White-Hautzinger et al. Explicit instruction & advanced Austrian EFL learners 122 rating stage, four trained listeners rated the participants’ performances. These listeners were considered trained due to their roles as language lab tutors at the University of Vienna, where they are regularly exposed to a wide range of EFL university students’ speech and expected to provide feedback. Each listener was assigned 41 student recordings (i.e., four sets) per rating cycle, with two cycles in total to ensure inter-rater reliability. In the first cycle, each rater assessed one set, and in the second cycle, they assessed a different set, ensuring that each recording received two independent ratings. Subsequently, a balanced rating approach was adopted, which included assigning codes to each student recording. This coding system ensured participant anonymity and prevented raters from knowing whether they were evaluating preor post-instruction recordings. Additionally, each rater assessed recordings from both SBE and GAE varieties, which ensured a fully crossed design and further minimised potential rating effects. Weighted Cohen’s Kappa (κw)2 was calculated for each set of raters for all three target features. The range of κw for each feature is presented here to summarise the overall agreement levels across all four sets: For /ɒ – a:/, κw values ranged from .32 to .70, indicating a relatively broad level of agreement from fair to substantial; /əʊ – oʊ/ showed moderate to substantial agreement, with κw values between .50 and .74. Finally, regarding the postvocalic /r/, the agreement was consistently higher, with κw values ranging from .74 to .90, reflecting substantial to almost perfect agreement across the sets. Moreover, all raters utilised a standardised rating scale (see Appendix) developed specifically for this study to maintain consistency. Each segmental was assessed across a set number of tokens. For each token of a feature, raters determined whether the pronunciation approximated the target variety, assigning a binary score of 0 (not approximated) or 1 (approximated). The total number of correct tokens for each feature was calculated by summing the scores across all instances. This value was then converted into a percentage of accurate approximations relative to the total number of target instances for that feature. Based on the percentage of correct approximations, participants were assigned a score on a 5-point scale: not approximated to target variety (value 1), slightly approximated to target variety (value 2), moderately approximated to target variety (value 3), mostly approximated to target variety (value 4), and fully approximated to target variety (value 5). 4.2 Results To investigate the significance of changes from pre-test to post-test, a Wilcoxon signed-rank test was computed for each target feature. The analysis also included the effect sizes of significant differences; this study reports the rank biserial correlation coefficient (rrb), which is appropriate for non-parametric tests of difference (King et al., 2011). To interpret the magnitude of the rrb, the analysis followed field-specific benchmarks for second language research proposed by Plonsky and Oswald (2014), who classify standardised r-values of .25, .40, and .60 as small, moderate, and large effect sizes, respectively.3 To reduce the risk of alpha error accumulation, the significance threshold was set at .01. The Wilcoxon signed-rank test revealed significant improvements in students’ pronunciation from the pre-test to the post-test across all target features. Table 2 presents the 2 To interpret the magnitude of agreement, Landis and Koch’s (1977) benchmarks were employed. 3 While Plonsky and Oswald’s (2014) field-specific benchmarks for interpreting Pearson r were developed based on parametric data, the conceptual similarity between Pearson r and rank biserial r₍rb₎ (Rosenthal, 1991) allows for a tentative application of these thresholds. As argued by Kirby (2014), rrb is more appropriate for non-parametric tests as it takes into account rank sum differences. Therefore, rrb was expected to yield more nuanced results. Index accent vi, vii, xiii, 17, 19, 21, 22, 23, 40, 46, 60, 61, 84, 85, 112, 119, 125, 130 accentedness vi, vii, xi, 70, 73, 92 accents xii, 22, 34, 41, 60, 119, 125 accommodation 111 accuracy ix, x, 43, 44, 46, 47, 48, 52, 56, 61, 84, 85, 124 action research 55, 56, 64, 65 affective 3, 61 age 17, 31, 41, 42, 43, 60, 62, 73, 93, 100, 112 AI 64 American viii, 17, 34, 41, 72, 75, 77, 99, 100, 118, 119, 120 applied research 55, 65 Arabic viii, ix, 119, 120 ASR 64 assessment ix, x, 1, 2, 29, 59, 61, 92, 93 attitude(s) 17, 31, 32, 42, 43, 62, 66, 70, 110 auditory analysis 118, 122 Austrian viii, ix, 97,118, 125 authentic 11, 15, 16, 17, 22, 23, 24, 54, 56, 58, 59, 62, 63, 64, 65 authentic speech 11, 16, 17, 22, 23 awareness 15, 16, 22, 28, 53, 63, 81, 109, 110, 114, 115 British viii, 17, 29, 34, 41, 72, 75, 118, 119 Catalan viii, 43, 45, 63 chatbots 64 ChatGPT 63 Chinese viii, 16 CLT 59 clusters 3, 112, 113, 114 cognitive load 59 competence 1, 2, 10, 17, 30, 97, 98, 99, 100, 105, 106 computer-generated visuals 64 confidence x, 28, 30, 34, 35, 37, 40, 65, 97, 98, 99, 102, 105, 106 connected speech x, 3, 5, 10, 15, 16, 17, 18, 19, 20, 22, 23, 44, 52, 113, 114 consonants 33, 34, 85, 86, 112, 113, 114, 119, 124 Croatia viii, 120 CSP 15, 16 Czech viii, 120 Danish viii, ix, 84, 85, 86, 87, 88, 92, 93, 94 delivery 3, 16, 22, 58, 60, 61, 65 diphthongs 113, 114, 119, 124 discourse-level 57, 61 Dutch viii, 120 effectiveness x, 11, 40, 56, 57, 58, 59, 61, 65, 98, 105, 124 EFL learners 15, 16, 17, 22, 23, 85, 86, 119, 120, 124, 125 EFL students 118, 119, 124 elision x, 4, 5,9, 10, 15, 17, 18, 19, 21, 22, 23, 114 EPI 118, 119, 120, 124 error(s) xi, 16, 30, 33, 50, 56, 62, 65, 122 evaluation(s) 3, 65, 84, 85, 86, 88, 89, 91, 92, 93, 94, 105 explicit pronunciation teaching 30, 98, 103 facial expression 60 feedback vi, vii, xi, 30, 32, 36, 57, 62, 87, 102, 105, 119, 120, 122, 124 films 59, 64 fluency ix, xi, 22, 23, 43, 44, 46, 47, 48, 49, 52, 61, 64, 92, 119 French viii, 16, 60 fricative(s) x, 74, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93 dental 28, 33, 34, 36, 37, 62, 115 General American English 75, 118, 119 generalisability 57 German viii, 119, 120, 121, 124, 125 gesture(s) vi, vii, ix, xi, 60, 61 Global English 17, 22, 117 Golden Speaker 64 Haussa viii, ix hesitation 2 Hungarian viii, 28, 30, 31, 32, 33, 34, 36, 37, 40, 41, 42, 120 HVPT 56, 57, 58, 59, 85 identity 60, 61, 62, 111 Indian viii, ix 130 intelligibility vi, vii, ix, x, xi, 3, 16, 17, 29, 56, 57, 58, 61, 62, 65, 71, 73, 78, 80, 81, 84, 89, 90, 91, 93, 98, 111, 114, 118, 119 intervention x, 17, 65, 110, 123, 124 intonation xiii, 42, 44, 50, 52, 56, 57, 60, 61, 64, 113, 114, 115 intrusive 5, 7, 9, 10, 15, 17, 18, 19, 21, 22, 23 IPA 33, 35, 36, 37, 102 Italian viii Japanese viii, 58, 85, 119 L2 1, 2, 3, 4, 11, 16, 17, 44, 53, 57, 60, 61, 62, 63, 64, 72, 73, 84, 85, 92, 93, 94, 105, 106, 119 language acquisition 109, 110, 111 attitudes 110, 111 lab 99, 118, 120, 121, 122 listen and repeat 33, 35, 36 listener(s) vii, xiii, 2, 3, 30, 44, 46, 52, 62, 63, 73, 74, 78, 80, 84, 85, 86, 88, 89, 92, 93, 94, 115, 119, 122 listening comprehension xi, 1, 2, 3, 4, 5, 11, 16, 17, 24, 41 Macedonian viii, ix, 15, 17, 18, 23, 24 mentor 113 methodological factors 56, 57, 58 mirroring 60, 61, 65 mobile media 59 monosyllabic 87, 115 MOOC vii, ix, x, xi, xiii, 109, 110, 111, 113, 115 movement 33, 43, 61, 62 native speech x, 16, 17, 18, 22, 24, 85 naturalness xi, 15, 16, 17, 22, 23 non-native 2, 16, 17, 20, 23, 30, 85, 86, 88, 92, 93, 94 Norwegian ix, 85 open access i, v orthographic ix, 64, 88 outcomes x, 43, 56, 57, 58, 66, 107, 125 Pakistan ix pedagogy x, 30, 58, 59, 60, 63, 65, 97, 98, 99, 104, 105 perception and production x, 22, 23, 84, 85, 86 perceptions x, xi, 15, 16, 17, 22, 23, 44, 99, 100, 105 perceptual training x, 84, 85, 86, 88, 92, 94 Philippines viii phonetic training 56, 84 pitch 43, 44, 46, 50, 52, 53, 58, 61, 62, 68, 114 Polish viii, 1, 2, 120 practical phonetics 99, 105, 118, 119 proficiency xi, 4, 9, 44, 52, 53, 60, 62, 73, 93, 97, 105, 106, 119, 120 prominence 3, 18, 61, 109, 114 prosodic 44, 61, 62, 64 prosody vii, ix, x, 43, 44, 46, 49, 52, 59, 61, 62, 64 quantitative ix, 18, 32, 68, 71, 76, 97, 100 rating scale 118, 121, 122 reading aloud 46, 48, 119, 121 reliability 32, 57, 88, 122, 125 replication 56, 57, 58, 64 research base 55, 56, 64 research design ix, 56, 57 rhythm x, 16, 19, 22, 23, 33, 34, 41, 44, 56, 61 Romanian ix, 120 salience x, 109, 110, 111, 112, 113, 114, 115 Scottish 15, 17, 18 scripted texts 1, 2, 3, 9, 11 secondary school ix, 28, 29, 31, 34, 35 segmentals 56, 62, 113, 120, 124 shadowing vi, vii, ix, xi, 58, 59, 64 sibilants 84, 85, 86, 89, 93, 94 silent letters 28, 33, 34, 36, 41, 44 Sindhi viii social meaning 60, 61, 62, 111 social value 109, 115 Spain iv, viii, ix, 44 Spanish viii, 30, 63, 71, 72, 73, 78, 80, 85 speaking style 3, 61, 65 speech rate 2, 3, 6, 8, 9, 15, 16, 18, 22, 44, 46, 52, 76, 84, 114 spoken text(s) 1, 2, 4, 5, 9, 11 spontaneous texts 9 Standard British English 118, 119 131 stereotyped 62, 111 stigmatisation 115 story 42, 43, 44, 45, 46, 48, 49, 52, 53, 121 stress x, xi, 11, 31, 33, 34, 41, 42, 44, 45, 56, 57, 60, 61, 63, 64, 73, 112, 113, 114, 115 word 31, 34, 37, 41, 57, 112, 113, 114 student teachers ix, 97, 99, 100, 102, 103, 104, 105, 106 styling 60 suprasegmental(s) 10, 44, 56, 57, 60, 61, 62, 64, 65, 99, 109, 113, 114, 115, 118, 119, 120 Swedish viii, 85 target vii, x, xi, 4, 21, 30, 37, 40, 57, 81, 87, 88, 90, 92, 93, 110, 118, 119, 120, 122, 123, 124, 127 task-based 59 teacher education x, 29, 37, 97, 98, 99, 100, 105 technology 60, 63, 64, 65 television 59, 64 tertiary 118, 119, 124 thought groups 61 timing 61 tonal 114, 115 transfer 57, 60, 63 Ukrainian viii, ix Urdu viii variationist 109, 110 variety x, 17, 29, 30, 56, 59, 64, 65, 111, 118, 120, 121, 122, 127 video 33, 42, 50, 63, 64 Vietnamese viii, ix, 109, 110, 112, 113, 114, 115 vowel(s) xi, 71, 73, 74, 75, 76, 77, 78, 79, 113, 115, 119, 124 length 71, 72, 73, 74, 76, 113, 114 reduction 17, 18, 19, 20, 22, 23, 25, 34 voice quality 50, 52, 61, 71, 76, 78, 81 whole-body context 60 YouGlish 64 YouTube 64 130