The Story of the Learner Corpus LINDSEI_CZ
Abstract
22
Full text
The Story of the Learner Corpus LINDSEI_CZ Tomáš Gráf ABSTRACT: The article presents the recently completed Czech subcorpus of the multinational learner corpus of advanced spoken English LINDSEI and aims to draw attention to some of the methodological concerns the field of learner corpus linguistics faces. First, it describes the Louvain family of learner corpora, where this project originated, and provides adetailed description of LINDSEI, its history, design, structure, transcription system and metadata. It then outlines the nature of the Czech subcorpus LINDSEI_CZ, telling the story of its compilation and providing aquantitative description of the corpus size, task sizes and learner variables, as well as adescription of the transcription process. The core part of this text discusses methodological concerns affecting learner corpus design and construction and deals with such issues as task design, recording instructions, the matter of learner-participant proficiency, and transcription system employed. It concludes with aconsideration of various methodological suggestions and offers the possible view that, despite certain weaknesses, LINDSEI is an invaluable source of highly authentic learner data. The last section provides athematic categorisation of existing studies on LINDSEI and concludes with descriptions of some future projects. The article calls for athorough reconsideration of learner corpus design and practice and for the formulation of compilation and research standards which would lead to an increase in the reliability and exploitation potential of learner corpora. KEY WORDS: corpus methodology, learner corpora, learner corpus linguistics, LINDSEI, spoken corpora 1 INTRODUCTION Learner corpora— “electronic collections of spoken or written texts produced by foreign language learners” (Granger, 2004, p.124)— have long been awidely recognized part of the broader field of corpus linguistics. They offer direct evidence of the processes which are involved in the production of written and spoken texts in an L2,1 and through the wealth of the data they contain they enable researchers to explore areas as diverse as second language acquisition, psycholinguistics and natural language processing and to utilize their research findings within L2 pedagogy. Arecent survey of learner corpus research— The Cambridge Handbook of Learner Corpus Research (Granger— Gilquin— Meunier, 2015)— comprises no fewer than twenty-seven specialised chapters, which shows the enormous potential that learner data has for research. Somewhat sadly— if perhaps not surprisingly— the target language for most learner corpora is English. Other languages, however, are also appearing on the scene; of the 144 monolingual learner corpora listed in the survey of 1 See e.g. fluency studies by Götz (2013) and Gráf (2015; 2017) based on the LINDSEI data. OPEN ACCESS
TOMáš Gráf 23 “Learner Corpora around the World”,2 89 (62%) are for English, 12 (22%) for Spanish, 10 (18%) for German, 9 (16%), for French, 6 (11%) for Italian, 3 (6%) for Finnish, 2(4%) for Arabic, and 2 (4%) for Persian, and there are also single instances for such languages as Czech, Dutch, Estonian, Gaelic, Hungarian, Chinese, Korean, Norwegian, Russian, Slovene, and Swedish. Besides these, there also exist 14 multilingual corpora with diverse target languages including— besides English and some of the above-mentioned languages— also Catalan, Portuguese and Romanian. The field is thus becoming truly international in two aspects: in the increasing variety of L2s for which learner corpora are being produced, as well as in terms of the growing abundance of learner corpora featuring language produced by learners of English with different mother-tongue backgrounds.3 When Iwas approached by Ondřej Tichý in 2012 and asked whether Imight consider making acontribution to the international LINDSEI project by compiling its Czech subcorpus (i.e. learner English produced by speakers with Czech as their L1), Idid not hesitate. As an English teacher of twenty years’ standing and ateacher trainer at university level with astrong interest in learner language Ialready had high hopes that learner corpus research might have valid and perhaps even invaluable pedagogical implications. Furthermore, at that time there still existed no such corpus containing data produced by Czech learners of English. In this article, written five years after the launching of the LINDSEI_CZ project and two years after its completion, Iwould like to provide not only the story of that subcorpus but also an evaluation of the LINDSEI design. In so doing Iaim to provide an outline of how far we have come in our understanding of both pedagogical and learner-corpus-methodological implications. 1.1 THE LOUVAIN COLLECTION Of COrPOrA It was at the end of the 1980s that Sylviane Granger conceived the idea of the Centre for English Corpus Linguistics (CECL) at the Université catholique de Louvain, which set out to create learner and multinational corpora especially for pedagogical purposes. Since then the Centre has generated 14 corpora,4 some of which are amongst the largest of their kind and include contributions from several countries. This extensive work has led to the creation of anew learner corpus methodology: Contrastive Interlanguage Analysis (Granger, 1996). The truly pioneering project— in both the local and the international context— was the International Corpus of Learner English (henceforth ICLE). This was started 2 For acomprehensive list of learner corpora around the world see <https://uclouvain.be/ en/research-institutes/ilc/cecl/learner-corpora-around-the-world.html> (last accessed on 17 August 2017). 3 The term “mother-tongue background” is preferred over “speakers with different L1s” as the former includes speakers who share acommon L1 but come from different countries (e.g. the French of France and Belgium), thus making it possible for research to take into account not only linguistic but also educational and cultural factors. 4 See <https://uclouvain.be/en/research-institutes/ilc/cecl/corpora.html> for acomplete list. OPEN ACCESS
24 STUDIE ZAPLIKOVANÉ LINGVISTIKY 2/2017 in the early 1990s and its first version was launched in 2002, with asecond version appearing in 2009 and athird being planned for 2017. At the moment of writing it contains 16 national subcorpora5 and is accompanied by aparallel native corpus, LOCNESS (the Louvain Corpus of Native English Essays), with 324,304 words. ICLE consists of argumentative essays written by high-intermediate and advanced learners of English as L2 and contains 3.7 million words. As the years went by, CECL launched several other corpora. These include learner corpora such as FRIDA (the French Interlanguage Database), LINDSEI (the Louvain International Database of Spoken English Interlanguage), LONGDALE (the Longitudinal Database of Learner English) and VESPA (the Varieties of English for Specific Purposes Database); pedagogical corpora such as CoNNECT (the Corpus of Native and Non-Native EFL Classroom Teacher Talk) and TeMa (acorpus of textbook materials); translation corpora such as Label France, MUST (Multilingual Student Translation) and PLECI (Poitiers-Louvain Échange de Corpus Informatisé); and specialised corpora such as LOCRA (the Louvain Corpus of Research Articles), MULT-ED (the Multilingual Editorial Corpus) and NESSI (New Englishes Student Interviews).6 2 LINDSEI From the start, ICLE was to be accompanied by aspoken counterpart. This project was commenced in 1995 under the name The Louvain International Database of Spoken English Interlanguage (LINDSEI). It was started by Sylviane Granger, who was later joined by Gaëtanelle Gilquin as the project’s co-ordinator. The first version of LINDSEI was launched and released on CD-ROM in 2010; its second version is planned for release in 2018. LINDSEI is accompanied by aparallel native corpus, LOCNEC (the Louvain Corpus of Native English Conversation), containing almost 122,000 words in 50 interviews with native English students of linguistics, and following the same format as that of LINDSEI. LINDSEI is acorpus of spontaneous spoken English produced by high-intermediate and advanced learners of English with various mother-tongue backgrounds. Version One was made up of 11 national subcorpora7 which together contained approximately one million words, 554 interviews and 130 hours of recorded material. Version Two is to add afurther nine subcorpora8 to give atotal of 20 subcorpora, 1,000 interviews, approximately 250 hours of recordings and almost two million words. At the time of writing, the corpus is distributed with only the orthographic transcriptions and not the actual recordings. 5 Bulgarian, Chinese, Czech, Dutch, Finnish, French, German, Italian, Japanese, Norwegian, Polish, Russian, Spanish, Swedish, Tswana and Turkish. 6 Detailed information on these corpora may be found at <https://www.uclouvain.be/en258636.html>. 7 Bulgarian, Chinese, Dutch, French, German, Greek, Italian, Japanese, Polish, Spanish and Swedish. 8 Arabic, Brazilian, Basque, Czech, Finnish, Norwegian, Latvian, Turkish and Taiwanese. OPEN ACCESS
TOMáš Gráf 25 2.1 LINDSEI— COrPUS STrUCTUrE, TASK DESIGN, PrOfICIENCY, AND TrANSCrIPTION SYSTEM Each national subcorpus comprises aminimum of fifty transcriptions of approximately fifteen-minute recordings, all of which contain three tasks that are identical not only within the given national corpus but right across the whole LINDSEI project. The first is amonologue on one of achoice of three set topics (an experience which has affected you; ajourney which has affected you; or amemorable film or play), all of which invite the use of the past tense and the present perfect. The speaker is given two to three minutes to choose atopic and think about what to say on it; in an ideal situation, the task itself takes aminimum of three minutes. The second task is afree conversation with the interviewer, topics here typically including the student’s history of studying English, his or her experience of university life and studies, plans for the future, etc., and thus attempt to elicit avariety of grammatical tenses. The third task is areconstruction of astory based on aset of four pictures. This is aspontaneous, improvisatory task which poses demands on the student’s ability to construct coherent, logical text including linking devices and avariety of prepositions. Speakers are required and expected to be advanced learners of English. Proficiency is not, however, checked prior to the interview (i.e. the speakers are not expected to produce any proof of their proficiency), but is defined institutionally; the speakers are to be students in their 3rd or 4th year of study of English philology, therefore at or near the end of their BA studies, as it is assumed that such speakers ought to possess the requisite level of English.9 As Iwill later show, this selection method is far from ideal. The recordings are orthographically transcribed, and anonymized. There is no punctuation, and full stops are used to mark unfilled pauses. The transcription of filled pauses differentiates between short, long and nasalised pauses, and overlaps are marked using atag. Non-standard forms (such as cos, dunno, kinda etc.) are maintained, and so are contracted forms. Truncated words are transcribed using the equals sign (e.g. Ireme= remember). Acertain number of phonetic features are also recorded, including syllable or vowel lengthening (using colons, e.g. Iwent to: an interesting place) and stressed articles (e.g. a[ei] and the[i:]). Attention is also paid to prosodic features (whispering, laughing etc.) and non-verbal vocal sounds (e.g. coughing, lip smacking). The speakers’ turns are marked using the tags <A>, </A> for the interviewer and <B>, </B> for the learner. Example 1 shows ashort extract from one of the transcriptions. <B> (er) she: she was an economist but now she (er) takes care of afarm with horses </B> <A> that’s achange <overlap /> wow </A> <B> <overlap /> <starts laughing> yeah <stops laughing> </B> <A> and she married and Irishman <overlap /> didn’t she </A> 9 In the case of the Czech subcorpus, students completing their BA studies are expected to have attained C1 level of the Common European Framework of Reference for Languages. OPEN ACCESS
26 STUDIE ZAPLIKOVANÉ LINGVISTIKY 2/2017 <B> <overlap /> yeah . yeah </B> <B> so: . Iusually go to Dublin first and visit my friends over there and then . Imove to her place and spend . (er) really a. leisure time over there </B> example 1:Abrief extract from one of the LINDSEI_CZ transcriptions. 2.2 LINDSEI— METADATA Prior to the recording, the speakers are asked to sign an informed consent form and complete aquestionnaire. The purpose of the latter is to collect variables which are believed to play arole in the acquisition process; these include social and language-acquisition-related variables such as name, age, gender, nationality, language background (parents’ L1s, language(s) spoken at home, other languages spoken by the student), length of study of English at various levels of education, and lengths of stays in English-speaking countries. These are later transferred to adatabase along with information regarding choice of set topic, duration of the interview and length in tokens, as well as with basic information on the interviewer and his/her degree of familiarity with the student. 3 LINDSEI_CZ— THE CZECH SUBCORPUS OF LINDSEI The compilation of the Czech subcorpus of LINDSEI was started in 2012 and completed in 2015 at the Department of English Linguistics and ELT Methodology, the Faculty of Arts, Charles University, Prague. The project received financial support from the Institute of the Czech National Corpus, on whose web site it is currently hosted within the KonText interface.10 The requisite fifty interviewees were recruited from among 3rdand 4th-year students of English philology at the aforementioned English department; they were interviewed by two of their teachers,11 whose prime concern was to maintain as natural aflow of communication as possible by providing encouraging feedback and asking open questions whilst keeping their own oral input to aminimum. Most of the recordings were made at the recording studio of the same faculty’s Institute of Phonetics. 3.1 LINDSEI_CZ— DATA DESCrIPTION LINDSEI_CZ contains 123,761 tokens.12 These include filled pauses and truncated words. Contracted forms with an apostrophe are counted as one word. 77.5% of the total number of words are uttered by the students. 10 For more information see <http://wiki.korpus.cz/doku.php/cnk:lindsei_cz>. 11 Sarah Peters Gráfová and Tomáš Gráf. 12 The total count of positions including all special characters and tags is 135,366. OPEN ACCESS
TOMáš Gráf 27 A & B turns Mean B turns only Mean Min. (B turns) Max. (B turns) Length in tokens 123,761 2,475 (SD = 386) 95,904 1,918 (SD = 407) 920 3,045 Duration (hh:mm:ss) 12:52:25 15:27 (SD = 2:14) 10:37:42 00:12:45 (SD = 2:24) 0:06:51 0:17:21 Table 1:LINDSEI_CZ— corpus size. Length and duration of interviews. Table 1 provides basic descriptive data regarding the size of the corpus. As is clear from the standard deviations and the large ranges of the values, there is much variability in the data, which can be explained by the fact that the recorded students included both very reticent and rather talkative ones. The data, however, has anormal distribution. As was mentioned above, each interview is divided into three tasks. As is apparent from Table 2 and Chart 1, responses to Tasks One and Two are similar in terms of number of tokens and overall duration. Task One comprises 42% of the whole corpus, Task Two 47% and Task Three 11%. Task One Task Two Task Three Length in tokens 40,584 42,850 12,535 Mean token count 812 (SD = 329) 857 (SD = 284) 251 (SD = 85) Duration 4 hours 26 minutes 4 hours 38 minutes 1 hour 32 minutes Mean duration (mm:ss) 5:20 5:56 1:51 Table 2:LINDSEI_CZ— task sizes. CharT 1:LINDSEI_CZ— comparison of task sizes (Y-scale = number of tokens). OPEN ACCESS
28 STUDIE ZAPLIKOVANÉ LINGVISTIKY 2/2017 3.2 LINDSEI_CZ— LEArNEr VArIABLES Whilst we strove to achieve balance in the structure of the data, the fact that the majority of our students are women meant that we were unable to recruit the same number of men and women, the final ratio of females to males being 43:7. The average age of the speakers is 22.5 years (SD = 1.6), and prior to university studies they studied English for an average of 9.9 years (SD = 2.6). At the time of being interviewed they had spent 3.4 years (SD = 0.9) studying English at university and an average of 1.2 months in English-speaking countries. This shows that the students learnt their English largely in institutional settings. As for their knowledge of other foreign languages, 25 students mentioned German, 14 French, 7 Spanish, and 4 other languages including Russian, Italian and Dutch. In all of the cases, the home language was Czech. 3.3 LINDSEI_CZ— THE TrANSCrIPTION PrOCESS The laborious process of transcribing the recordings was performed by the speakers themselves as part of their credit requirements for courses in SLA and ELT methodology, both of which included seminars on the specifics of spoken language. Whilst the pedagogical dimension of this experiment proved to be meaningful and— as became clear in subsequent seminar discussions— led the students towards adeeper understanding of the complexity and structure of spoken discourse, the resulting transcriptions varied in quality. The coordinator of the corpus was obliged to spend an average of three hours checking each transcription to guarantee its accuracy and consistency. Features which were especially taxing from both the transcribers’ and the checker’s view were unfilled pauses (esp. their length), filled pauses (nasalisation and length) and the marking of overlaps. The students reported having needed on average of five hours to transcribe each recording. The total number of hours for transcribing and checking was thus about 400. It would perhaps have been more time-efficient to have had all the recordings transcribed by one person. The transcriptions followed the rules outlined in the Louvain transcription manual,13, which was made available to all LINDSEI coordinators, and some of whose details are described in section 2.1 above. 4 LINDSEI— METHODOLOGICAL CONCERNS AND CAVEATS The experience of compiling LINDSEI_CZ revealed anumber of methodological problem areas, some of which might— and perhaps rather ought to— have been considered prior to the project. Adolphs and Knight (2012) recommend that all phases of spoken corpus construction be considered at the design stage. It is perhaps due to the fact that LINDSEI was planned in the early days of learner corpus linguistics that some of these issues did not receive as much attention as they deserved. In the follow13 Its full version is available at <https://www.uclouvain.be/en-307849.html>. OPEN ACCESS
TOMáš Gráf 29 ing paragraphs Iwill attempt to explore some of these “weaknesses” whilst remaining constantly aware of the fact that LINDSEI was and is atruly pioneering project in the field of spoken learner corpora. 4.1 CONCErNS rEGArDING THE GENErAL PUrPOSE Of THE COrPUS AND THE DESIGN Of THE TASKS To this day any researcher interested in compiling aLINDSEI subcorpus receives what subsequent work on the given subcorpus reveals to be to be rather scanty information regarding the project. The idea stated is that LINDSEI is aspoken corpus of advanced learner English and aspoken counterpart to ICLE, but it is not specified whether any concrete research questions or interests are intended. How the interviews are actually carried out in terms of communicative content is thus left very much to the coordinator’s own experience or initiative. The positive side is that the resulting data is more spontaneous and authentic in being less controlled, but on the other hand anumber of questions arise and remain unanswered. Are the interlocutors to elicit particular grammatical forms and, if so, how is this to be achieved? What is the ideal proportion of the individual tasks in terms of duration? How conversationally active should the interlocutor become? These considerations can be illustrated by taking abrief look at task design. As described above, Task One consists in ashort speech on one of three set topics. These all seem to be designed to elicit statements in which the speakers need to contrast the present perfect and the past tense. But was this really the intention? Later analysis of the speech of those who chose the third topic (afilm or play which has impressed you) reveals that some speakers opt for the historical present. But might this be the case because it is actually the speakers’ avoidance strategy regarding the more complex tenses? As regards speech continuity, what is the interlocutor to do if the speaker is very brief? This was indeed the case with some of the more reticent speakers, and the interlocutor was obliged to intervene by asking supplementary questions and thus turning what had presumably been designed as amonologue into adialogue. Could not this situation have been easily avoided if the interlocutor had been specifically instructed to inform the student that he would be expected to talk without any interruptions for anumber of minutes and to say more than just afew sentences? As aresult, in some of the interviews Task One is more dialogical than in others and hence the formal (genre) distinction between Task One and Task Two, which could be significant for research, is negligible or even lost. The dialogue form may be limiting also in the sense that if the interlocutor asks specific (leading) questions his/her own use of grammar might prompt the speaker to produce grammatical constructions which the latter otherwise would not have used or might not even be capable of using. Task Two is intended to be adialogue spurred by the interlocutor’s questions. Here the instructions provide abrief list of possible topics (namely life at university, hobbies, and plans for after university) but can there be ahidden agenda here? Is the interlocutor to attempt to elicit different temporal references to the past, present and future? Indeed, this might have provided highly useful research data and would have OPEN ACCESS
30 STUDIE ZAPLIKOVANÉ LINGVISTIKY 2/2017 been quite easy to achieve if alist of suitable questions had been presented to the interlocutor with the explicit requirement that questions were to refer to the past, present and future. In this way more guidance would also have been provided as to the expected or ideal length of this part of the interview. Task Three is anarrative based on aset of four pictures. According to the instructions, the speakers are to make up ashort story, but no other guidance is given. To what extent is the interlocutor to intervene here, for example, if the speaker is very brief? Are questions allowed? Are the speakers to be prompted to produce more language, or is the task to be essentially monological? Could the instructions for the speakers regarding this part be more detailed and actually specify that, besides the retelling of the story, the pictures are also to be described? Ihave randomly checked this task in the various national subcorpora and found avariety of speaker solutions to this issue ranging from very short, one-sentence descriptions to lengthy dialogues about the pictures and even about the meaning of art. Whilst such data is authentic to the core and very valuable, comparisons across the subcorpora as well as of different speakers within one subcorpus may be problematic. 4.2 CONCErNS rEGArDING SPEAKEr PrOfICIENCY Carlsen (2012) describes proficiency in learner corpora as a“fuzzy variable”, claiming that in learner corpus research proficiency levels are often not adequately defined. As was stated above, LINDSEI uses an institutional definition of proficiency, which cannot guarantee comparability of results. Listening to the recordings indeed confirmed just how “fuzzy” this approach was; proficiency levels in them ranged from B2 to C2. Whilst LINDSEI sets out to be acorpus of advanced English, this proficiency definition problem would appear to make the whole corpus into amulti-level rather than exclusively advanced corpus. In 2016 an international team comprising Taiwanese researcher Lanfen Huang, co-ordinator of LINDSEI_TW, and the present author received aTaiwanese government grant14 for one of its joint projects: the carrying out of apost-hoc, perceptive proficiency rating in our two LINDSEI subcorpora (_TW, _CZ). To this end we engaged three professional IELTS examiners who had been previously trained to rate proficiency in accordance with CEFR levels. They gave ratings for lexical range, accuracy, fluency, phonological control, coherence and overall impression. As shown in Table3, the Czech subcorpus appears to be alevel ahead of the Taiwanese, with the majority of the Czech speakers rated as C1 and of the Taiwanese speakers as B2. In both corpora the figures for the respective highest levels (C2 for Czechs and C1 Taiwanese) are of low representativeness. It might also be argued that the high ratio within each corpus of speakers with lower proficiency somewhat undermines the corpus’s aims to be one of advanced English. 14 Ministry of Science and Technology, Taiwan, grant number MOST105-2628-H-158-001. OPEN ACCESS