Full text
KAROLINUM COMPILING AND ANNOTATING A LEARNER CORPUS FOR A MORPHOLOGICALLY RICH LANGUAGE CZESL, A CORPUS OF NON-NATIVE CZECH ALEXANDR ROSEN JIŘÍ HANA BARBORA HLADKÁ TOMÁŠ JELÍNEK SVATAVA ŠKODOVÁ BARBORA ŠTINDLOVÁ COMPILING AND ANNOTATING A LEARNER CORPUS CZESL, A CORPUS OF FOR A MORPHOLOGICALLY RICH LANGUAGE NON-NATIVE CZECH I loved the book and its versatility in presentation of learner corpus research in general and of CzeSL in particular. It takes up a variety of aspects connected to the theory and practice of development and use of learner corpora, their limitations and advantages, technical details and specifi cations. It is a complete and convenient package of research around L2 corpora, a reference book on L2 corpora construction and use, and CzeSL guidelines and manual – all in one. Elena Volodina Associate Professor of Linguistics, University of Gothenburg, Sweden The book goes much beyond systematically documenting a long-term learner-corpus project: On the one hand, it includes a conceptual discussion of the compilation and annotation of learner language in terms of linguistic properties and learner errors, including a discussion of other learner corpora providing a cross-language context. On the other, it systematically discusses the concrete, sustained eff ort of the CzeSL project building the Czech non-native language corpus, the reasoning behind the specifi c annotation schemes used, the annotation process, the soſt ware tools used to support this complex process, use cases for this corpus, as well as a compilation of lessons learned that will be directly relevant for the increasing number of research teams and languages for which such corpora are being built. The morphologically-rich aspect is an important factor for the annotation, off ering particular challenges for learner language analysis, and it nicely contrasts with the very dominant English focus, typical of learner corpus research, to point the way towards a more comprehensive refl ection of the diff erent linguistic domains that can and should be modeled to obtain comprehensive characterizations of learner language. The book fully confi rms why I had chosen an early version of the CzeSL corpus as one of the two exemplary case-studies included in my 2015 article in the Cambridge Handbook of Learner Corpus Research – even fi ve years later, it remains on the cutting-edge of learner corpus research. Detmar Meurers Professor of Computational Linguistics, University of Tübingen, Germany learner corpus_mont.indd 1 15/10/2020 11:27
Compiling and annotating a learner corpus for a morphologically rich language CzeSL, a corpus of non-native Czech Alexandr Rosen Jiří Hana Barbora Hladká Tomáš Jelínek Svatava Škodová Barbora Štindlová Reviewed by: Detmar Meurers, University of Tübingen, Germany Elena Volodina, University of Gothenburg, Sweden © Charles University, 2020 © Alexandr Rosen, Jiří Hana, Barbora Hladká, Tomáš Jelínek, Svatava Škodová, Barbora Štindlová, 2020 Published by Charles University Karolinum Press Ovocný trh 560/5, 116 36 Prague 1 Prague 2020 Typesed by Jiří Hana First edition ISBN 978-80-246-4759-3 ISBN 978-80-246-4765-4 (online: pdf)
Univerzita Karlova Nakladatelství Karolinum www.karolinum.cz [email protected]
Contents List of abbreviations 11 1 Introduction 13 1.1 About this book ............................. 13 1.2 Reasons to study non-native Czech ................... 14 1.3 Some properties of non-native Czech .................. 17 1.3.1 Morphology ............................ 18 1.3.2 Syntax ............................... 19 1.3.3 Word segmentation ........................ 21 1.4 Learner corpus .............................. 21 1.5 Roadmap ................................. 23 2 Learner corpora 25 2.1 Terminology ................................ 25 2.2 Various types of learner corpora ..................... 26 2.2.1 The choice of texts ........................ 26 2.2.2 Annotation ............................ 27 2.2.2.1 Textual annotation .................. 27 2.2.2.2 Linguistic annotation ................. 28 2.2.2.3 Error annotation – correction ............. 28 2.2.2.4 Error annotation – categorization .......... 29 2.2.2.5 Annotation scheme .................. 30 2.2.3 Data access ............................ 31 2.3 Some learner corpora ........................... 32 2.3.1 ASK ................................ 32 2.3.2 CLC ................................ 33 2.3.3 COPLE2 ............................. 34 2.3.4 CroLTeC ............................. 34 5
6CONTENTS 2.3.5 Falko ................................ 35 2.3.6 ICLE ............................... 36 2.3.7 MERLIN ............................. 36 2.3.8 RLC ................................ 37 2.3.9 SweLL ............................... 38 2.4 Relationships of CzeSL with other learner corpora .......... 39 3 Introducing the CzeSL project 41 3.1 Specifications of CzeSL .......................... 42 3.2 Intended usage .............................. 43 3.3 AKCES – the umbrella project ..................... 45 4 Procurement of texts 49 4.1 Text collection .............................. 49 4.2 Transcription ............................... 51 4.3 Anonymization .............................. 53 4.4 Metadata ................................. 54 5 Error annotation 59 5.1 Errors and learner language ....................... 59 5.2 More than one way to annotate errors in CzeSL ............ 64 5.3 A wishlist for error annotation ..................... 65 5.3.1 Interference and other types of explanation .......... 66 5.3.2 Interpretation in terms of TH .................. 66 5.3.3 Word order ............................ 67 5.3.4 Style ................................ 68 5.3.5 Communication goal ....................... 68 5.4 The two-tier annotation scheme ..................... 69 5.4.1 Annotation scheme as a compromise .............. 69 5.4.1.1 Why multiple tiers ................... 69 5.4.1.2 How many tiers .................... 71 5.4.1.3 Multiple tiers in a tabular format .......... 71 5.4.1.4 Content of the tiers .................. 72 5.4.1.5 A sample text with T1 vs. T2 corrections . . . . . 73 5.4.1.6 Links between tiers .................. 73 5.4.1.7 Error tags ....................... 76 5.4.1.8 Morphosyntactic references .............. 76 5.4.1.9 Follow-up corrections ................. 77 5.4.1.10 Alternative target hypotheses ............. 77
CONTENTS 7 5.4.2 Error tagset ............................ 78 5.4.2.1 Based on linguistic categories ............. 78 5.4.2.2 Grammar-based vs. formal errors ........... 80 5.4.2.3 Extent of the annotated unit ............. 81 5.4.3 Grammar-based tags ....................... 81 5.4.3.1 Errors at T1 ...................... 81 5.4.3.2 Errors at T2 ...................... 83 5.4.3.3 Coarse-grained ..................... 83 5.4.3.4 An example of complex annotation .......... 84 5.4.4 Evaluation of the manual tiered error annotation ....... 87 5.4.4.1 Inter-annotator agreement (IAA) ........... 88 5.4.4.2 A pilot annotation ................... 89 5.4.4.3 IAA on all doubly-annotated texts .......... 89 5.4.4.4 Error tags depend on target hypothesis ....... 93 5.4.4.5 Possible causes of the annotators’ disagreements . . 95 5.4.5 Formal tags ............................ 97 5.4.5.1 Automatic extension and modification of error annotation ......................... 97 5.4.5.2 Automatic detection of formal errors on T1 . . . . . 98 5.4.5.3 Formal orthographic errors .............. 99 5.4.5.4 Formal errors sometimes influencing pronunciation . 100 5.4.5.5 Formal errors influencing pronunciation . . . . . . . 101 5.4.5.6 Other types of errors .................103 5.4.5.7 Automatic classification of word-boundary errors . . 105 5.5 Implicit error annotation .........................105 5.6 Multi-dimensional error annotation (MD) ...............108 5.6.1 Focus on morphology ......................108 5.6.2 All annotation applied to the source text ...........109 5.6.3 Extent of the annotated unit ..................109 5.6.4 Alternative error domains ....................110 5.6.5 Source text, target hypothesis, annotated strings . . . . . . . 112 5.6.6 Domains and features ......................113 6 Linguistic annotation 119 6.1 Annotation with tools for Standard Czech ...............120 6.1.1 Annotation of target hypothesis .................120 6.1.2 Annotation of T1 ......................... 121 6.1.3 Annotation of source texts .................... 121 6.2 Annotation of interlanguage in UD ...................122
8CONTENTS 6.2.1 Tokenization ...........................124 6.2.2 Part-of-speech and morphology .................124 6.2.3 Lemmata .............................126 6.2.4 Syntactic Structure ........................128 6.2.5 Evaluation ............................130 7 Annotation process 131 7.1 Overview of the annotation process ................... 131 7.2 Transcription and anonymization of manuscripts ...........132 7.3 Tiered error annotation .........................133 7.3.1 Manual error annotation .....................134 7.3.2 Automatic annotation checking .................135 7.3.3 Data format for the tiered annotation scheme .........136 7.4 Automatic error tagging .........................136 7.5 Automatic correction ...........................139 7.6 Multi-dimensional error annotation ...................140 7.6.1 Morphemic analysis ....................... 141 7.6.2 Automatic error annotation ...................143 7.6.3 Experiments with automatic identification of errors in inflection144 7.6.4 Manual error annotation .....................147 7.6.5 Post-processing of manually annotated texts ..........149 7.7 Implicit annotation ............................150 7.8 Universal Dependencies .........................152 8 The CzeSL corpora 155 8.1 CzeSL-plain – without annotation and metadata ...........157 8.2 CzeSL-SGT – with automatic annotation ...............158 8.3 CzeSL-man – with manual annotation .................163 8.3.1 CzeSL-man v0 ..........................163 8.3.2 CzeSL-man v1 ..........................163 8.3.3 CzeSL-man v1 downloadable ...................164 8.3.4 CzeSL-man v1 searchable ....................165 8.3.5 CzeSL-man v2 ..........................167 8.4 CzeSL-TH .................................168 8.5 CzeSL-MD ................................168 8.6 CzeSL-UD .................................169 8.7 CzeSL-GEC and AKCES-GEC .....................169 8.8 CzeSL in TEITOK ............................170 8.9 Learner corpora of native Czech .....................170
CONTENTS 9 9 Tools 173 9.1 Annotation tools .............................173 9.1.1 feat ................................173 9.1.2 Speed ...............................174 9.1.3 brat ................................175 9.1.4 TrEd ................................175 9.1.5 Error annotation tools ......................175 9.1.5.1 Automatic error tagging in 2T ............175 9.1.5.2 Automatic error detection and tagging in MD . . . 176 9.1.6 Conversion tools .........................176 9.2 Search tools ................................177 9.2.1 SeLaQ ............................... 177 9.2.2 Sketch Engine and KonText ...................179 9.2.2.1 Token-based error annotation .............180 9.2.2.2 Error annotation using structures ..........183 9.2.3 TEITOK .............................189 10 Using the corpus 201 10.1 Learner corpora from the perspective of language teachers . . . . . . 202 10.2 The use of corpora in language research and teaching .........203 10.2.1 Benefits of learner corpora ....................204 10.2.2 Limitations of learner corpora ..................205 10.3 Corpus-based research and teaching of Czech as a foreign language . 207 10.3.1 The Czech National Corpus in the service of Czech as a foreign language ..............................207 10.3.2 Analyses based on learner corpus data .............209 10.4 Applications in natural language processing .............. 211 10.4.1 Text scoring ............................ 211 10.4.2 Text correction ..........................212 10.4.3 Natural language identification .................214 11 Lessons learned and perspectives 217 11.1 What we would do the same way again .................217 11.2 Blind alleys and second thoughts ....................220 11.3 Outlook ..................................225 12 Acknowledgements 227 A Notes about examples 231
16 CHAPTER 1. INTRODUCTION varieties have not received enough systematic attention.10 To summarize, the increasingly stronger position of Czech as L2, the merely intuitive understanding of L2 Czech based on the teacher’s experience and the persisting lack of modern didactic support for non-native speakers, including school children, make a strong case for a broadly conceived research, focused on those properties of non-native Czech which are significant, representative, identifiable, amenable to processing by formal tools, and comparable across learner texts of all types.11 Therefore, research in non-native Czech is important for both theoretical and practical reasons: 1. Like every IL, non-native Czech calls for identifying developmental patterns and orders of acquisition. In addition to the research on the sequence of acquisition of L1 structures, there are studies about the acquisition of specific aspects of L2, such as morphemes, pronouns and word order (e.g., Ellis and Barkhuizen 2005) but not for an inflectional language such as Czech. 2. Analyzing IL is essential for SLA research, which helps to reveal how language works in general. IL contributes to the understanding of linguistic universals in SLA (White 2003). Investigating SLA of Czech is important because of its typological specifics. 3. Non-native language also offers data for studying variability as a key indicator of how a situation affects the learners’ use of L2, either as free variations in the use of a language pattern which has not yet been completely acquired, or as systematic variations, determined by a linguistic, social or psychological context. Studying these phenomena helps to understand the development of learners’ IL while offering comparison with the acquisition of L1. The investigation of IL is worthwhile also in order to find out what types of errors learners make and what the errors say about their knowledge of target language and their ability to use it. This is important especially for didactic purposes: Czech should be described with regard to non-native speakers, and methods for teaching Czech to foreigners need to be elaborated and, i.a., translated into curricula (reflecting the relative difficulty of acquisition of individual phenomena). Analyzing IL is crucial also for language testing (individual features of IL need to be related to standard proficiency levels). the-follow-up-master-degree-programme/. 10For studies dealing with the presentation of specific linguistic phenomena, see §10. 11This applies especially to learners with a typologically distant L1, who are not acquainted with the European grammatical categories rooted in Latin.
1.3. SOME PROPERTIES OF NON-NATIVE CZECH 17 New methodologies, based on extensive data and computational tools, help to advance this line of research. Although unsupervised methods can be used, a formal model must be based on the research of relevant aspects of IL. Its absence has also practical consequences. Many NLP tools that are taken for granted (spell checkers with suggestions, Internet search supported by morphology, machine translation, etc.) perform much worse for non-native Czech or are simply unusable, because experts developing applications for the native language cannot rely on previous research. Moreover, the study of L2 also has an intrinsic value in itself, as a study of a cultural phenomenon: L2 is part of the non-native speakers’ identity. 1.3 Some properties of non-native Czech Non-native speakers deviate from the standard language in non-arbitrary ways (see, e.g., Ellis and Barkhuizen 2005); the deviations are to a large extent systematic and predictable, but also evolving as learners receive more input and revise their hypotheses about L2 (R. Ellis 2003, 33–35). They are influenced by: 1. The speaker’s native language (interlingual errors): •s matkoj →s matkou ‘with mother’ – Russian ending -oj instead of Czech -ou •jsem vietnamský →jsem Vietnamec ‘I am Vietnamese’ – adjective (as in English etc.) instead of a noun •žádný to ví →nikdo to neví ‘nobody knows it’; lit.: ‘none it not-knows’ – a single negative form žádný ‘none’ (as in English, German, etc.) instead of multiple negative forms nikdo ‘nobody’ and neví ‘not-knows’ 2. The general properties of the process of acquisition (intralingual errors): •v neděli spám dlouho →v neděli spím dlouho ‘I sleep late on Sundays’ – misuse of endings from another inflectional class: znát ‘to know’ – znám ‘I know’ vs. spát ‘to sleep’ – spím ‘I sleep’, a case of “false analogy” •tady jsou pět stoly→tady je pět stolů‘there are five tables here’, lit.: ‘there is five tables.GEN here’ – a case of “overgeneralization” from simpler quantifier-free patterns tady jsou stoly ‘there are tables here’ 3. The properties of the instructional process
18 CHAPTER 1. INTRODUCTION Many deviations of non-native Czech as compared with Standard Czech (SCz) belong to the domains of morphology and morphosyntax.12 This is why we focus on these levels and design a system of concepts capturing the deviations in a systematic, formal and linguistically motivated way. The concepts are supported by computational models. Various contrasts in the patterns of IL can thus be made explicit and linked to parameters such as stages of acquisition and differences due to linguistic backgrounds of the speakers. 1.3.1 Morphology In Czech, as an inflectional language, the syntactic functions of words are mostly expressed by their form, whereas word order is to a large extent constrained by information structure. Thus the domain of morphology plays a key role in Czech and due to its relative complexity represents the main source of deviations from the standard. It also deserves attention for a practical reason: many tools for computational language processing assume that methodologies and resources concerning morphology are available. To give an example, Czech nouns have seven cases, with distinct forms for singular and plural, which means that any noun may have up to 14 different forms, although in every declension paradigm some forms are identical due to syncretism in case and/or number. For nouns, there are 14 basic paradigms, and a larger number of paradigm subtypes. The paradigm žen|a‘woman’ (the most frequent for feminine nouns) has 10 forms, e.g., žen|ěis the form for dative and locative singular. Also adjectives, pronouns, numerals and verbs have many paradigms and inflected forms. It is nonetheless not necessary to master the entire Czech inflectional system in order to successfully communicate in Czech. It is enough to know how to use the most frequent cases and verbal forms for the common paradigms. For example, the understanding of the sentence in (1) is not disrupted by the error Úvalách → Úvalech.13 (1) Oslavil celebrated jsem aux Vánoce Christmas se with svými self’s příbuznými relatives v in jejich their domě house v in *Úvalách Úvaly.(loc) →Úvalech. Úvaly.loc 12For comparison, see an overview of native Czech grammar in Appendix B. 13A parenthesized morphosyntactic category in the gloss, such as .(loc), denotes an intended use of the category in an incorrect form. For a list of all conventions used in the examples, including identification of the source text, see Appendix A.
1.3. SOME PROPERTIES OF NON-NATIVE CZECH 19 ‘I celebrated Christmas with my relatives at their home in Úvaly.’ (KAR_MD_020 ru A2+) The error Úvalách →Úvalech is in the form of the locative plural of the name of a Czech town Úvaly, a plurale tantum: the case ending -ách used incorrectly instead of -ech is an existing ending, used to express the same morphosyntactic properties of nouns of another paradigm (Roztokách ‘Roztoky.loc’). As the incorrect form denotes the same nominal case of the noun, it may be noticed as unexpected by a native speaker, but it will not hinder the understanding of the whole sentence. Another type of inflection error, found in (2), can make the sentence slightly less understandable. (2) V in životě life dávám give.1sg přednost precedence *rodinu family.*acc →rodině. family.dat ‘In my life, I prefer family.’ (HRD_AS_221 ja A2) The form rodinu ‘family’ is a form for accusative singular of the noun rodina, in a construction where the dative form rodině is expected. The sentence is still understandable, as it is composed of only a few words, but the use of incorrect case makes it more challenging to be understood by a native speaker. When a completely random ending is used, unrelated to paradigm or case form, the understanding is even more disrupted. 1.3.2 Syntax Most syntactic deviations are found in morphosyntax, often when forms are lexically determined, e. g., by valency as in (3) or subject to a principle of grammar, e.g., agreement, as in (4). (3) S: myslím think.1sg *o about tobě you T: myslím think.1sg na on tebe you ‘I’m thinking about you’ (DGD_L5_143 ru A1) (4) S: *přišli came.*pl.*ma pět five studentů students.gen
20 CHAPTER 1. INTRODUCTION T: přišlo came.sg.neut pět five studentů students.gen ‘five students came’ Other deviations include non-standard word order due to an inappropriate topicfocus articulation (information structuring), or due to a misplaced clitic, such as in (5), where jsem ‘am’ and se – reflexive particle – are both 2nd position clitics and should follow the first constituent během studování na univerzitě ‘during university studies’ in that order. (5) S: během during studování studying na on univerzitě university *se refl seznámil met *jsem aux.1sg s with Evou Eva T: během during studování studying na on univerzitě university jsem aux.1sg se refl seznámil met s with Evou Eva ‘during my university studies I met Eva’ (BLAH_DZ_001 ky B2) Similarly, reflexive pronouns are often under-used, as in (6), where the possessive moji ‘my’ should be replaced by the reflexive possessive svoji. (6) S: miluji love.1sg *moji my práci work T: miluji love.1sg svoji self’s práci work ‘I love my work’ (TOD_P2_247 ru A2+) Various types of errors can occur within the same sentence, as in (7). (7) S: V in rodině family *byli be.*pl.*ma *jsme aux.1pl pět five – táta, dad *mátka, (mother) *dve (two) *sestři (sisters) a and já. I T: V in rodině family nás we.gen bylo be.3sg.neut pět five – táta, dad matka, mother dvě two sestry sisters a and já. I ‘We were five in my family – father, mother, two sisters and me.’ (HRD_1S_197 en A2) In (7), the three errors in morphology and morphonology (mátka,dve and sestři) are combined with an error in the agreement pattern involving quantified subject,
1.4. LEARNER CORPUS 21 shown in (4). The quantified subject NP, agreeing with a verb form in the 3rd person neuter singular (like přišlo ‘came.3sg.neut’ in (4)), includes the genitive form of the first person plural pronoun (nás – in the source sentence assumed to be nominative and thus pro-dropped). On the other hand, the 1st person plural past tense auxiliary jsme is dropped in the target sentence, because there is no auxiliary in the 3rd person past tense. 1.3.3 Word segmentation Inappropriate word segmentation is not a random phenomenon either. Words are often incorrectly split after prefixes homonymous with prepositions (do psat ‘in write’ instead of dopsat ‘finish writing’). Russian speakers influenced by their native language sometimes append reflexive pronoun to the verb (smějuse →směju se ‘I laugh’) or split verb and the negative particle (ne studuju →nestuduju ‘I don’t study’). Since clitics, such as some prepositions and short pronouns, form prosodic units with their host, speakers exposed primarily to spoken Czech might spell them incorrectly as one word, as in (8).14 (8) S: *opálímse suntan.1sg+refl ale but *musímít must.1sg+have opalovací krém sunscreen abych so-that-aux.1sg *semse aux.1sg+refl nespálil burned.neg moc too much T: opálím suntan.1sg se refl ale but musím must.1sg mít have opalovací krém sunscreen abych so-that-aux.1sg se refl nespálil burned.neg moc too much ‘I will get a sun tan but I must have a sunscreen so that I would not get sunburnt too much.’ (ss_dp_057_63 cs 11) 1.4 Learner corpus Investigating language acquisition by non-native learners helps to understand important linguistic issues and to develop teaching methods, better suited both to the specific target language and to specific groups of learners. These tasks can now be based on empirical evidence from learner corpora. A learner corpus consists of language produced by language learners, typically learners of a second or foreign language (L2). Such corpora may be equipped with 14Example (8) is from SKRIPT 2015, a corpus of young native Czech learners. The text ID is followed by the code for Czech (cs) and the age of the author.
22 CHAPTER 1. INTRODUCTION morphological and syntactic annotation, together with the detection, correction and categorization of non-standard linguistic phenomena. Learner corpora allow to compare non-native and native speakers’ language, or to compare interlanguage varieties, and can be studied on the background of standard reference corpora, which helps to track various deviations from standard usage in the language of non-native speakers, such as frequency patterns – cases of overuse or underuse – or foreign soundingness as compared with the language of native speakers. A range of studies have focused not only on the frequency of use of individual elements of language (e.g., Ringbom 1998), including phenomena such as negative and positive transfer, formulaic language, collocations (lexical patterns), prefabs and colligations (lexico-grammatical patterns, e. g., Nesselhauf 2005; Paquot and Granger 2012; N. C. Ellis 2017; Granger 2017; Vetchinnikova 2019), lexical analysis and phrasal use (e.g., Altenberg and Tapper 1998), but also developmental patterns, variability and the impact of the learning context (Meunier 2019; Granger, Gilquin, and Meunier 2015) and interlanguage complexity (e. g., Paquot 2019). An error-tagged corpus can be subjected to computer-aided error analysis (CEA), which is not restricted to errors seen as a deficiency, but understood as a means to explore the target language and to test hypotheses about the functioning of L2 grammar. CEA also helps to observe meaningful use of non-standard structures of IL. Such studies focus on lexical errors (e.g., Leńko-Szymańska 2004), wrong use of verbal tenses (e.g., Granger 1999) or phrasal verbs (e.g., Waibel 2008). The tasks of designing, compiling, annotating and presenting such corpora are often very much unlike those routinely applied to standard corpora. There may be no standard or obvious solutions: the approach to the tasks is often seen as an answer to a specific research goal rather than as a service to a wider community of researchers and practitioners. The difference between a standard and a learner corpus is mainly in their annotation. Texts in a learner corpus can be annotated in two independent ways: (i) by standard linguistic categories: morphosyntactic tags, base forms, syntactic structure and functions, and (ii) by error annotation: correct version of each ill-formed part of the source text, i. e., its target hypothesis (TH), and categories specifying the nature of errors. Reasonably reliable methodologies and tools are available for linguistic annotation (i) of many languages, as long as the text is produced by native speakers. The situation is different for non-standard language of non-native learners and for error annotation (ii), where manual annotation is quite common. However, with the growing volumes of learner corpora, the need for methods and tools simplifying such tasks is increasing. Yet the annotation of learner corpora remains a challenging task, even more so for a language such as Czech, with its rich inflection, derivation, agreement, and a largely information-structure-driven
1.5. ROADMAP 23 constituent order. 1.5 Roadmap Chapter 2:Learner corpora provides some context by listing several properties which make each learner corpus different from any other. The second half of the chapter presents an overview of nine learner corpora with features relevant for the CzeSL corpus. Chapter 3:Introducing the CzeSL project presents the foundations of the CzeSL project and outlines its main characteristics. Chapter 4:Procurement of texts deals with the initial tasks in the compilation of the CzeSL corpora. It is the first in the sequence of three chapters concerned with how the texts are treated and what kind of annotation they receive. These chapters do not focus on the actual pre-processing, which is the topic of Chapter 8, but rather on the description of the principles, categories and formats. Chapter 5:Error annotation presents the background and substance of several types of error annotation used in the CzeSL project. We focus on the original error annotation scheme, consisting of three parallel tiers for the source text and two tiers for its annotation. This type of error annotation is examined from several angles: we provide motivation behind this design, present the grammar-based and the “formal” error tagsets, complementing each other, and provide results of its evaluation in terms of inter-annotator agreement (IAA). The chapter follows by introducing two additional types of error annotation used in the CzeSL project more recently: annotation without explicit error tags, facilitating manual annotation, and a multidimensional scheme, complementing the original tiered system especially in the domain of morphonology. Chapter 6:Linguistic annotation examines the approaches adopted in CzeSL to the annotation of morphosyntactic categories, syntactic structure and functions. The chapter consists of three main parts: it starts with the methods analyzing the TH, proceeds to methods developed for standard language but used to annotate source learner texts, and concludes with the description of an approach to syntactic analysis designed specifically for learner Czech.
24 CHAPTER 1. INTRODUCTION Chapter 7:Annotation process looks at the transcription, anonymization and annotation from the perspective of a step-by-step procedure, including decisions about the share of manual tasks and suitability of automatic tools. The various types of annotation described in the preceding chapters are here described in terms of input, processing and output. Chapter 8:The CzeSL corpora provides an overview of searchable corpora or downloadable data sets containing the CzeSL texts. The various releases reflect the various approaches to the annotation, but they also differ in the choice of texts, availability of metadata and the data format, determining the search options, i.e., the choice of a suitable search tool. Chapter 9:Tools is an overview of tools used within the CzeSL project for processing and annotating texts on the one hand and for searching and viewing them on the other. Chapter 10:Using the corpus is concerned with how the CzeSL corpora are used in research and teaching of Czech as a foreign language, and also in NLP applications such as text scoring, text correction and natural language identification (NLI). Some of these types of use are closely related with the exploitation of standard reference corpora for the same purpose, which is why a section about the use of corpora of native Czech is also included. Chapter 11:Lessons learned and perspectives concludes the core chapters of the book by discussing positive and negative experience from implementing various solutions throughout the project and by an outlook into the future. Chapter 12:Acknowledgements should be seen as an important part of the book. There are many people and several funding agencies who deserve our credit for starting the project and for keeping the project alive throughout the years. Appendix A:Notes about examples briefly summarizes the presentation of examples. Appendix B:The Czech language presents an overview of Czech as a native language in its main features. This part may be useful especially for readers who do not speak or understand Czech.
Chapter 2 Learner corpora Since the release of the International Corpus of Learner English (Granger, Dagneaux, and Meunier 2002), learner corpora have become a well-established branch of corpus linguistics. Now they are an important source of data for foreign language teaching, second language acquisition and other related disciplines (McEnery 2018; McEnery et al. 2019). The growing number of learner corpora of various kinds, formats and access options have also led to efforts aimed at making the data, methods and tools used in various projects reusable by trying to achieve some degree of conceptual and structural interoperability (Chiarcos 2012; Stemle et al. 2019). 2.1 Terminology A learner corpus, also called interlanguage or L2 corpus, is a computerized textual database of language as produced by L2 learners (Leech 1998). A similar definition, where native language learners are excluded, is used by Granger (2008). Although our topic is Czech as L2, we would prefer to treat as a learner corpus each corpus concerned with language acquisition, no matter whether the language represented by the corpus is L1 or L2. The reasons enumerated by Granger (2008), related to the blurring of L1 and L2 in the context of various dialects of English around the world, apply also to Czech and its varieties, such as its Romani ethnolect or dialects used by communities of heritage Czech speakers abroad. Moreover, some corpora may intentionally include texts produced by both non-native learners and native speakers of a language. Once we agree that young native speakers are also learners, such corpora, including both L1 and L2, should also be called learner corpora. 25
32 CHAPTER 2. LEARNER CORPORA version. In some projects, handwritten texts or audio recordings are not disclosed to every corpus user to protect the learner’s identity. To some extent, the choice of query interface is determined by the data and annotation format. However, for many users the search tool is the only window to the corpus, so the options of querying and visualization offered by the search tool are crucial. For example, some users may prefer a tool such as TEITOK, which is able to display the text together with its annotation and properties of the handwritten source at the same time. Other users need an interface with a rich menu of statistical functions such as KonText or a tool combining the search and annotation environments (Korp and SVALA,TEITOK). 2.3 Some learner corpora This is a very partial overview of some currently available learner corpora. They were selected because some of their features are related in one way or another to the CzeSL project.7 2.3.1 ASK – Norsk andrespråkskorpus8 The corpus of Norwegian as second language, developed in 2006–2014, consists of transcripts of essays, hand-written by learners who had passed the higher level test in Norwegian for adult immigrants. The texts (1,936 items, 770 thousand words, 1,130 thousand tokens) were selected to achieve typological diversity in L1s: German, Dutch, English, Spanish, Russian, Polish, Bosnian-Croatian-Serbian, Albanian, Vietnamese and Somali. The corpus also contains texts from native Norwegians as control data. The texts and metadata are marked up in XML according to the TEI Guidelines.9The texts were typed in and validated using a standard XML editor (Oxygen). For error annotation, the TEI guidelines are extended by the attributes corr (corrections) and sic (errors), type (error category) and desc (subcategory). The 7For a more exhaustive overview of learner corpora see, e. g., Pravec (2002), Nesselhauf (2005), Štindlová (2011,2013), and Xiao (2008), or more up-to-date lists at https://www.uclouvain.be/ en-cecl-lcworld.html, and https://www.clarin.eu/resource-families/L2-corpora. 8https://clarino.uib.no/ask/, Tenfjord, Meurer, and Hofland (2006) and Tenfjord, Hagen, and Johansen (2009). The methodology and infrastructure of ASK were also used to build a pilot learner corpus for Slovene (PiKUST, Stritar 2009). 9In this respect, ASK preceded the learner corpora available in TEITOK (see §9.2.3), including CzeSL in TEITOK.
2.3. SOME LEARNER CORPORA 33 sic tags can be used recursively to mark up more than one error in a word or phrase. The error tagset is rather small in order to avoid inconsistencies in the error coding and redundancy due to the presence of POS tags.10 The tags are of the following seven types: lexical (lexeme, spelling, foreign word, word boundary, capitalization, derivation), morphology (category, paradigm), syntax (missing or redundant word or phrase, word order: inversion, adverbial), punctuation, uninterpretable, followup. Error annotation of the source texts is complemented by linguistic annotation using a tagger for standard Norwegian, with a facility for manual tag correction. The same tagger is applied also to the corrected texts. The source texts and their corrected versions are aligned as a parallel corpus. They can be searched and corresponding sentences displayed in parallel using Corpuscle, a corpus query engine and web-based corpus management system (Meurer 2012). The corpus is available under the CLARIN Res (Priv) license.11 2.3.2 CLC – Cambridge Learner Corpus12 CLC is an English learner corpus built and used by Cambridge University Press as a proprietary resource and a part of the Cambridge English Corpus.13 The texts are collected from learners taking one of the various types of Cambridge English Language Assessment exams in English: general, academic, business, legal, finance, or life skills. The whole corpus consists of 55 million words. The corpus is tagged and lemmatized, and about one third is error-annotated with a tagset of nearly 90 tags.14 Authorized users can search the corpus using Sketch Engine (see §9.2.2) with all its functionalities, including Word Sketches. Error annotation is implemented as pairs of XML structural elements err and corr, representing an incorrect form and its correction (see §9.2.2.2). E.g., to find all errors in incorrect verb tense associated with past participles the user should use the following query: 10Both reasons were also behind the decision to use a relatively small tagset in the manual tiered annotation of the CzeSL corpus. This CzeSL tagset consists of 26 tags. 11For details of the license see http://urn.fi/urn:nbn:fi:lb-2019071729. 12https://www.cambridge.org/sketch/help/; Nicholls (2003). 13The other part is the Cambridge Reference Corpus, consisting of 2 billion words of native English, both written and spoken. 14For a list of the CLC error tags see https://www.cambridge.org/sketch/error_codes_grouped. html.
34 CHAPTER 2. LEARNER CORPORA [tag=“VVN”] within <err type="#TV"/>.15 Alternatively, a simple error query interface can be used to search for the source word forms, error tags and corrections. The corpus metadata include L1, nationality, exam, CEFR level, year, educational level, age, years of English study, gender, pass or fail. A part of the corpus is accessible without error annotation as one of the Sketch Engine corpora under the name Open Cambridge Learner Corpus (Uncoded).16 This corpus consists of 11.5 thousand texts consisting of 3 mil. words from learners with 7 different L1s. Apart from its use in the publishing house for creating methodologies, textbooks and other English Language Teaching (ELT) materials, the corpus has also been used for creating the English Vocabulary Profile.17 2.3.3 COPLE2 – COrpus de Português Língua Estrangeira / Língua Segunda18 COPLE2 is a corpus of written and spoken texts produced by students of Portuguese as L2 and by applicants for exams in Portuguese, built since 2013. The corpus contains about 1,100 texts (230 thousand words) from learners with 15 different L1s and proficiency levels from A1 to C1, and covers different topics and tasks. The corpus is in the TEI format, built, maintained and searchable in the TEITOK environment (see §9.2.3). Together with CroLTeC, this corpus served as a model for CzeSL in TEITOK. The metadata include the L1, CEFR level, months of studying Portuguese, nationality, knowledge of other foreign languages, text genre, topic and text type. Manuscripts or oral productions are also available. The transcripts encode modifications by the student and the teacher. The corpus is annotated for POS, lemma, TH and error type. 2.3.4 CroLTeC – CROatian Learner TExt Corpus19 CroLTeC consists of essays written in weekly intervals as a part of a course in Croatian, collected since 2016 from 755 non-native learners of Croatian at all levels 15The error annotation based on the XML structural elements combined with the tabular (vertical) taken-based format has been adopted also in one of the CzeSL corpora (see §8.3.5). 16https://www.sketchengine.eu/cambridge-learner-corpus/ 17http://vocabulary.englishprofile.org 18http://teitok.clul.ul.pt/learnercorpus/; Mendes et al. (2016), Rio et al. (2016), and Rio and Mendes (2019). 19http://teitok.clul.ul.pt/croltec/; Preradović, Berać, and Boras (2015).
2.3. SOME LEARNER CORPORA 35 of proficiency with 36 different L1s. The size of the corpus is 1 million words. About 3.5 thousand texts were hand-written and transcribed, 1.2 thousand texts were digitally born. Like COPLE2,CroLTeC uses the TEITOK environment (see §9.2.3), which means that the corpus can be extended, modified, annotated and otherwise improved while being available for online searching at the same time. The transcripts encode corrections made by learners themselves (deletions, insertions and word order changes). The texts are POS tagged and lemmatized, hand-corrected and assigned error tags. Metadata include gender, age, nationality, mother tongue, bilingual and multilingual competence, parents’ language proficiency, required linguistic competence for the task, genre, scope, time limit, size limit and the task circumstances (homework, part of an exam, field work, etc.). 2.3.5 Falko – Ein fehlerannotiertes Lernerkorpus des Deutschen als Fremdsprache20 Falko, built since 2004, contains 641 texts (about 380 thousand words) written by non-native learners of German, complemented by 152 texts (about 92 thousand words) in its comparative native German section. The L2 part alone comprises several sub-corpora: text summaries, essays written by advanced learners, and a longitudinal corpus from learners with different proficiency levels. The comparative part includes texts for each of the non-native sub-corpora. All annotation is strictly stand-off, each type in a separate tier. The tiers for POS and lemmas are available for all texts. Error annotation, consisting of TH and error tags, is available only in some subcorpora. Additional tiers can be added at any time, which means that alternative THs are possible in addition to successive THs. For the essay subcorpus, alternative THs are available, tagged for POS and lemma: “minimal” – grammatically correct and “maximal” – approaching the standard native language. Error tags, annotating differences between a TH and the source text, are represented as separate tiers. Unlike the concept of parallel tiers in CzeSL (see §5.4), which allows for any reordering of words at the neighboring tiers while preserving the cross-tier links between corresponding (even non-contiguous sequences of) words, the tiers in Falko can be represented as rows in a table with columns standing for the cross-tier links. 20https://www.linguistik.hu-berlin.de/de/institut/professuren/korpuslinguistik/forschung/ falko; Reznicek et al. (2012).
36 CHAPTER 2. LEARNER CORPORA For any word order corrections, the cells for the word order region must be merged (horizontally) at the TH tier, which means that the cross-tier links between the individual words are lost. In an extreme case of a region afflicted by the need to correct an error in word order spanning an entire sentence, the whole sentence may end up as a single column. The tabular format is due mainly to the annotation tools21 rather than to the format or the search tool. The corpus is available under the CC BY 3.0 license and can be searched using the powerful ANNIS tool.22 2.3.6 ICLE –The International Corpus of Learner English23 The ICLE project, launched in 1990, includes essays written by university students of English mainly in their second or third year. In 2002 the corpus was released as a CD-ROM accompanied with a handbook. ICLE was the first academic learner corpus of a considerable size and is still seen as the paradigm of a methodologically mature approach to the design of the content of a learner corpus. ICLE v3, the latest, web-based and on-line searchable version, published in 2020, includes over 9 thousand essays (5 million words, the length of each between 500 and 1,000 words), written by learners from 26 mother tongue backgrounds. The corpus is balanced in terms of the share of various L1s. There are 14 metadata items about the learner and 7 items about the task. The texts are tagged and lemmatized, but they are without error annotation. Besides a trial version with some restrictions, the full version allows the download of entire texts.24 2.3.7 MERLIN – Multilingual Platform for European Reference Levels: Interlanguage Exploration in Context25 MERLIN, built in 2012–2014, consists of 2,286 texts (340 thousand words) from learners of three languages: German (1,033 texts), Italian (813 texts) and Czech (442 texts, 64.5 thousand words). The texts come from written exams of acknowledged test institutions, aiming to test knowledge across the CEFR levels A1–C1. The corpus is tagged, lemmatized, parsed and on-line searchable using a custom 21Falko add-in for Microsoft Excel or EXMARaLDA https://exmaralda.org/en/ 22https://korpling.german.hu-berlin.de/falko-suche/ 23https://corpora.uclouvain.be/cecl/icle/; Granger (1998b,2003b). 24The license is available for a fee or to institutions that are members of the eduGAIN interfederation (https://edugain.org/), using the Shibboleth log-in system. 25https://www.merlin-platform.eu; Wisniewski et al. (2014) and Boyd et al. (2014).
2.3. SOME LEARNER CORPORA 37 platform, based on ANNIS, with a detailed error taxonomy and the option of two target hypotheses (minimal and extended, similar to Falko). There is a specific motivation behind the project. MERLIN is meant to provide examples of authentic texts for the individual CEFR levels in order to highlight the distinctions on the basis of comprehensive empirical characteristics. Moreover, some of the error tags are designed to check whether the CEFR descriptors, concerning a language and a CEFR level, correspond to the way learners actually use the language. In addition to the CEFR specifications, the design of the error tagset is based on issues in SLA research, features reported by experts in teaching, analyses of textbooks, language tests and learner texts. In fact, each of the tags is labeled for its source. Throughout the corpus creation process, the strategy was to reuse existing methodologies, formats and tools, resulting in a combination of a number of tools, many of them adopted from the Falko project.26 The corpus is available under an open license (CC BY-SA 4.0). In addition, the project approach and computational architecture is designed to be adaptable to other languages for which CEFR level illustration is needed. 2.3.8 RLC – The Russian Learner Corpus27 As of 2016, RLC was a collection of 2,000 texts produced by learners of Russian as L2 and 1,500 texts by speakers of heritage Russian with various dominant languages, altogether 730 thousand tokens. The texts include academic writings, movie and picture descriptions, book summaries, expository essays and others. A part of the corpus are speech transcripts. Some texts constitute a longitudinal subcorpus of academic writing. The corpus is annotated by morphological tags and lemmas, and includes two tiers of error annotation, based on deviations from Standard Russian: formal corrections (spelling, case forms, gender/number agreement, tense and aspect) and lexical/constructional violations.28 There are 59 error tags29 for errors in spelling (6), morphology (6), syntax (2), constructions (1), lexicon (5), and 7 supplementary tags (combined with the above tags). There are 10 metadata items for each text. 26See Stemle et al. (2019) for an overview of the MERLIN strategy. 27http://web-corpora.net/RLC; Rakhilina et al. (2016). The corpus can be searched from http://web-corpora.net/RussianLearnerCorpus/search/. 28The two annotation tiers in RLC resemble the two-tiered error annotation scheme of CzeSL. However, the range of errors annotated at Tier 1 is larger in RLC. 29See http://www.web-corpora.net/RLC/help.
38 CHAPTER 2. LEARNER CORPORA 2.3.9 SweLL – research infrastructure for Swedish as a second language30 The aim of the SweLL project (2017–2020) is to provide methods and tools for processing learner texts and to build a corpus of L2 Swedish consisting of about 600 texts. Some of the texts are transcribed from manuscripts and some are digitally born. The results include a portal for data collection via file import and online exercises. Handling of sensitive data is a priority – all texts are anonymized or pseudonymized according to precise rules. Error annotation is done using SVALA, an annotation editor developed within the SweLL project (Volodina, Matsson, et al. 2019). In a way, the editor is similar to feat (see §9.1.1): there is a tier for the source text and another parallel tier for its corrected version (the TH) with links across the tiers connecting corresponding tokens. Error tags label links with a correction. The display with a sequence of vertical links connecting tokens on the two tiers is called spaghetti mode. In a text where the source and the target tokens correspond 1:1, the aligned spaghetti are straight, uncooked. Like in feat, there may be more than one or even no corresponding token on either of the two tiers, and a link may cross other links when the word order changes, resulting in a cooked spaghetto (curly and/or split). It is the task of the annotator to correct the TH tier, edit the alignment links and add error tags. The source and target substrings, corrected within a single form, are highlighted. There are altogether 36 error tags of five main types: orthographic (3), lexical (4), morphological (8), punctuation-related (4), and syntactic (11). The “Other” type (6) includes a tag for follow-up (“consistency”) and unidentified corrections, intelligible and foreign strings, and comments (internal and for the corpus user). Although the tagset is not too large, it specifies some error types with respect to a more detailed grammatical category: e. g., morphological errors include tags for errors in case, definiteness, gender and number. Annotating a single error by a combination of tags is allowed, e.g., for an error in orthography and morphology, lexicon and syntax, or morphology and syntax. In addition to POS and lemmas, linguistic annotation includes syntactic parse and word-sense disambiguation. The plans include experimental linguistic annotation of the source texts to obtain a parallel treebank. The corpus can be searched using the general CWB-based Korp tool.31 To see the annotation in the spaghetti mode, a click takes the user to the SVALA editor with the text including the concordance line. 30https://spraakbanken.gu.se/en/projects/swell; Volodina et al. (2016) and Volodina, Granstedt, et al. (2019). 31https://spraakbanken.gu.se/en/tools/korp
2.4. RELATIONSHIPS OF CZESL WITH OTHER LEARNER CORPORA 39 The corpus annotation will be available under the CLARIN RES (Priv) licence. 2.4 Relationships of CzeSL with other learner corpora Each of the learner corpora briefly described above was selected for a reason. Some of the projects, such as ICLE or Falko, were crucial by providing inspiration for the design and development of CzeSL, while other projects are noteworthy because CzeSL shares some important features with them. The concept of parallel annotation tiers in Falko, supporting alternative and successive THs and implemented in the stand-off fashion, was at the origin of the tiered annotation scheme of CzeSL (see §5.4). The main difference is in the flexibility of the cross-tier links: instead of the spreadsheet-like tabular format of the annotation editor used in Falko, the annotation editor used in CzeSL retains links between corresponding tokens at different tiers even though the tokens are moved to remedy incorrect word order. Another difference is that unlike Falko and like RLC,CzeSL allows only for two annotation tiers. Also, unlike Falko, we did not adopt ANNIS, a general-purpose search tool supporting stand-off annotation. We explored several other directions instead: a search tool built to fit the annotation scheme (see §9.2.1) and conversion startegies into several other formats: the standard token-based tabular format (see §8.3.4), the Sketch Engine format used in the CLC corpus (see §8.3.5) and the TEI XML format used in COPLE2 and CroLTeC (see §8.8). Together with CroLTeC and RLC,MERLIN is included as another learner corpus of a Slavic language, actually of Czech as one of its three languages. MERLIN is also interesting for its strategy to reuse existing tools and formats and for one of its goals: to discover language-specific features pointing to the individual proficiency levels. We find several meeting points with the two Scandinavian projects. Like one of the more recent CzeSL releases, ASK uses the TEI format, and like the CzeSL tiered annotation, ASK also uses a restricted error tagset to avoid inconsistency and redundancy in the presence of linguistic annotation (see §5.4.2). Probably a more common feature is the use of the same tagger for the source and the target text, as in an automatically annotated CzeSL release (see §8.2). From our perspective, the most interesting part of the SweLL project is SVALA, the annotation editor of two parallel texts: the source and the target, with links between the corresponding tokens, reminiscent of the annotation editor used in CzeSL for the tiered annotation (see §9.1.1). However, there are other interesting parts: the project’s policy con-
40 CHAPTER 2. LEARNER CORPORA cerning sensitive personal data, the search tool combining standard concordances with the parallel text view and the plan to turn the corpus into a parallel treebank. A more detailed picture of CzeSL will emerge from the following chapter.
Chapter 3 Introducing the project of Czech as a Second Language Czech as a Second Language is the name of a long-term project and its results – a series of learner corpora. After a historical note this chapter presents some context of the project and an overview of principles and properties embodied in the results. At first we list the main properties of the CzeSL corpora (see §3.1): the scope of L1s and CEFR levels, their size, annotation and available metadata. We follow by outlining the intended use (see §3.2) and an overview of AKCES, a larger project of which CzeSL is a part, which includes additional corpora of Czech as L1, spoken and written mostly by schoolchildren (see §3.3). In many ways, building a learner corpus of Czech as a second/foreign language has been a unique enterprise. To the best of our knowledge, CzeSL was one of the first learner corpus ever built for a highly inflectional language.1CzeSL texts have also been used in a number of studies related to FLT or SLA and in NLP applications (see §10). A case study of the CzeSL error annotation scheme appeared in The Cambridge Handbook of Learner Corpus Research (Meurers 2015). CzeSL has been advancing since 2009 in the volume and types of texts, in the extent and quality of annotation, and in the access options. Throughout the time, new methods and tools have been tested and implemented. CzeSL is still not a closed and finished project. It is extendable by additional annotation and more data, including longitudinal, spoken, comparative L1 texts. 1There was one learner corpus for a Slavic language available at the time CzeSL was released, namely PiKUST (Stritar 2009), including 35,000 words with error annotation adopted from the Norwegian project ASK (see §2.3.1) and one of the few using multi-layer annotation. 41
Chapter 4 Procurement of texts Each corpus starts with collecting its content and related tasks. As it happens with most steps related to learner corpora, there are more issues specific to learner texts than to standard texts produced by native speakers. Using CzeSL as an example, we show how learner texts can be obtained, transcribed, equipped with metadata and protected from potential infringement of personal rights. 4.1 Text collection The texts included in the CzeSL corpus in the first round (i.e., until 2012) were collected mainly from learners attending an educational institution in the Czech Republic. Most of the learners were adults (18 or older), but there were also some younger learners (15–17). A substantial share of responsibility was with the text collectors, often teachers of the class. In detailed guidelines and during extensive schooling, the collectors were instructed about the choice of learners and text topics, the handling of texts, the acquisition of metadata and the learner’s consent about the use of the text. However, the guidelines were not always fully observed and some texts eventually included in the corpus did not follow the rules. According to the rules, 3 to 4 texts were collected from a learner within a single time interval (a school term). The task specification, included as a part of the metadata of each text, were defined as follows: 1. A text on an assigned topic, depending on the proficiency level, such as “My Family”. 49
50 CHAPTER 4. PROCUREMENT OF TEXTS 2. A text on a topic selected by the learner from a list of up to 16 suggestions in the guidelines, adaptable to the learner’s age and proficiency level. 3. One or two texts on a free topic. However, this did not always mean a really free choice. It was often the case that a topic picked as appropriate by the teacher/collector was assigned. The texts were written in regular classes, as a part of final exams, but also as homework. The decision to include homeworks was mainly due to the fact that texts produced in class are rather short and not too many. This is because teachers prefer to use some of the classroom time to provide students with resources and backgrounds needed for the homework rather than spending the time on writing. Indeed, there is a risk that the use of various aids or applications in an environment beyond the teacher’s control may give a distorted picture about the learner’s vocabulary and grammar-related competence. However, students are allowed to use various aids even in the classroom and thus the difference between home and school does not matter that much. We also considered the opposite approach of collecting texts from test situations, when the use of aids is controlled. However, in testing learners tend demonstrate only a part of their real competence, choosing means they are sure of to avoid failing the test. A homework on a topic of their interest is a very different task. The learners are much more ready to experiment while trying to express even complex thoughts. The comparison with texts on the same topics included in the MERLIN corpus is striking: the MERLIN texts, written during the CEFR exams, are rather stereotypical and uncreative. The collectors also had an important role in the protection of personal rights. In addition to the task assignment, text handling, and the acquisition of metadata (see §4.4), the collectors had to negotiate the consent for making the anonymized learner texts public (see §4.3 about anonymization). In the days before GDPR, when the texts were collected, the legal demands were less strict, but the procedure was still taken seriously. For adult learners, the collector signed a solemn declaration that all participating learners agreed to the use of their texts in the project and in the corpus. For juvenile learners, the collector had to obtain the headmaster’s approval. For juvenile learners attending the preparatory language courses at Charles University in Prague, a person authorized to represent the parents had to agree. The authors of more recent additions are only adult learners and each of them signs a legally conformant statement of consent. For the previously obtained consents we assume the prohibition of legal retroactivity.
4.2. TRANSCRIPTION 51 4.2 Transcription Like most texts currently written by students in educational contexts, the materials we collected for CzeSL were mostly hand-written. This is usually the only available option, given that their most common source are language courses and exams.1 The avoidance of an electronic format is also due to the concern about the use of automatic text-editing tools by the students, which may significantly distort the authentic interlanguage. Therefore, many texts have to be transcribed.2 The manuscript properties are recorded in order to support the research of handwriting, especially of students with a different native writing system. Also captured are corrections made by the student (insertions, deletions, etc.), useful for investigating the process of language acquisition. While we strive to capture only the information present in the original hand-written text, often some interpretation is unavoidable. For example, the transcribers have to take into account specifics of hand-writing of particular groups of students and even of each individual student (the same glyph may be interpreted as iin the hand-writing of one student, eof another, and aof yet another). Parts of some texts may be completely illegible and are marked as such. Sometimes the text allows multiple interpretation, e. g., the case of initial letters or word boundaries are often unclear. When the transcriber is not able to provide a single interpretation, two or even more variants can be used. Their order is assumed to signify preference of the first variant as the most likely interpretation. Some of the downstream processing steps which do not accept variants take advantage of this order by accepting the first interpration and discarding the rest of them. While deciphering unclear handwriting, transcribers sometimes have to rely on context and their best guess. However, they are not instructed explicitly to apply the “principle of positive assumption” of Volodina, Granstedt, et al. (2019): “Whenever one of the alternatives involves better intelligibility or closer adherence to standard norms, that is the alternative which should be chosen.” Unlike this principle, the approach of encoding variant interpretations is more focused on details of the learner’s handwriting. In retrospect, a single interpretation guided by the principle would have prevented some processing issues downstream at a bearable cost. 1Electronic texts (BA, MA and Ph.D. theses) represent a minority. While these texts were not written in a class or with the aim to be included in a corpus, their final form may have been affected by an automatic spellchecker. More recently, learner texts typed in an electronic format have become more common additions to CzeSL. 2For transcription and anonymization from the perspective of annotation as process see §7.2.
52 CHAPTER 4. PROCUREMENT OF TEXTS Viktor je mladý pan z Polska Ruska. Studuje {češtinu}<in> ve škole, protože ne umí psat a čist spravně. Bydlí na koleje vedle školy, má jednu sestru Irenu, která se učí na univerzite u profesora Smutneveselého. Bohužel, Viktor není dobrý student, protože spí na lekci, ale jeho sestra {piše všechno -> všechno piše} a vyborně rozumí českeho profesora Smutneveselého {a brzo delá domací ukol}<in> . Večeře Irena jde na prohasku spolu z kamaradem, ale její bratr dělá nic. Jeho čeština je špatná, vím, že se vratit ve Polsko Rusk o u a tam budí studovat u pomalu myt podlahy. Kamarad Ireny je {A|a} meričan a chytry můž. On miluje Irenu a chce se vzít na ní. protože ona je hezká, taky chytra, rozumí ho a umí vyborně vařit. Viktor je mladý pan z <del>Polska</del><add>Ruska</add> . Studuje <add>češtinu</add> ve škole, protože ne umí psat a čist spravně. Bydlí na koleje vedle školy, má jednu sestru Irenu, která se učí na univerzite u profesora Smutneveselého. Bohužel, Viktor není dobrý student, protože spí na lekci, ale jeho sestra <subst><del>piše všechno</del><add>všechno piše</add></subst> a vyborně rozumí českeho profesora Smutneveseleho <add>a brzo delá domací ukol</add> . Večeře Irena jde na prohasku spolu z kamaradem, ale její bratr dělá nic. Jeho čeština je špatná, vím, že se vratit ve <del>Polsko</del><add>Rusk<del>o</del>u</add> a tam budí studovat u pomalu myt podlahy. Kamarad Ireny je Američan a chytry můž. On miluje Irenu a chce se vzít na ní. protože ona je hezká, taky chytra, rozumí ho a umí vyborně vařit. Figure 4.1: A sample hand-written document with its transcription in the plain and the XML-based format (NEM_GD_008 ru B2)
4.3. ANONYMIZATION 53 An example of a manuscript and its transcription can be seen in Figure 4.1.3 The text is transcribed in two different ways, which differ in how some relevant features of the handwriting, mostly self-corrections, are encoded. For example, the author replaced the form Polska ‘Poland’ by Ruska ‘Russia’, inserted a word češtinu ‘the Czech language.acc’ and changed the word order by moving the word všechno ‘everything’ leftwards. The codes are set on gray background. At first, the hand-written texts were transcribed using off-the-shelf editors supporting HTML (e.g., Microsoft Word or Open Office Writer). As in the first transcript, a set of codes is used to capture variants, illegible strings, self-corrections and emoticons.4Deletions are transcribed as strikeout text (Polska), insertions use transcription codes in angle brackets following a string in braces ({češtinu}<in>), word order changes are annotated using an infix arrow-like notation ({piše všechno -> všechno piše}). This format also supports alternative interpretations ({A|a}meričan), where the first option is the preferred reading. For example, the string … represents omission (…), &img; indicates the place, where there was a picture in the manuscript, &unclear; stands for an unrecognized word or passage, &rdot; is a string indicating the character with a dot above etc. Unreadable characters or words were transcribed as XXX. The original transcription method was prone to unchecked typos in the markup and was replaced later (in texts transcribed since 2018) by a different setup, based on an editor checking for inconsistencies in an XML-based format, including XML codes for the transcription markup.5 The second transcript uses such XML codes, e.g., <del>Polska</del> for deletion and <add>češtinu</add> for insertion. Alternatives are not supported in this format. The transcription and anonymization codes follow the TEI guidelines wherever possible.6 4.3 Anonymization In most cases, the author’s identity cannot be revealed in the text or through metadata. The hand-written texts are anonymized during the transcription: personal information is replaced either by generic names (e.g., for names of persons and 3This is the text presented with glosses and target hypotheses in Table 5.1 on page 74. 4For details, see Štindlová (2011, 106; in Czech), or an abbreviated transcription guide http: //utkl.ff.cuni.cz/~rosen/public/transcription-reference.pdf (in English). 5See http://utkl.ff.cuni.cz/~rosen/public/TranscriptionGuideXML-cs.pdf for a transcription guide and http://utkl.ff.cuni.cz/~rosen/public/TranscriptionMarkupXML-cs.pdf for a list of codes. 6See https://tei-c.org/release/doc/tei-p5-doc/en/html/CC.html.
54 CHAPTER 4. PROCUREMENT OF TEXTS towns) or special codes (e.g., for telephone numbers). The use of substitute names is sometimes called pseudonymization. In pseudonymization we strive to preserve agreement features (by matching the name’s gender, number and case) and some of the possible errors (e.g., capitalization and some errors in declension). We use different substitutes for declinable and non-declinable names, but we do not attempt to match the declension class – e.g., all declined female given names are substituted by an appropriate form of the name Eva, even if the original name (such as Lucie) has a different set of declension endings. Unsurprisingly, male given names are replaced with a form of Adam. According to the original guidelines, names of smaller places (streets, villages, small towns) and other potentially sensitive data were replaced by QQQ. Later, all such names, together with email addresses, phone numbers, zip codes and other information potentially revealing the author’s identity, were coded as &priv;, e.g., {ulice}<&priv;> for street (ulice). The XML-based anonymization codes use a single element anon with a type attribute, which identifies the type of the anonymized item, e. g., <anon type="female FirstName">Eva</anon>. The substitute names can be used according to similar rules also for place names and institutions, e. g., <anon type="street">Dlouhá </anon>. Substitutes need not be used where they cannot reflect any linguistically relevant irregularities in the original forms, e.g., <anon type="phone"/>.7 4.4 Metadata In a learner corpus, metadata about the author of the text are at least as important as all other types of annotation.8The same set of metadata items is available in most CzeSL corpora for nearly all texts. There are 15 items about the author of the text and 15 items about the text itself. The sociological and linguistic data about the learner include age, gender, first language, proficiency level in Czech according to CEFR, knowledge of other (nonnative) languages, bilingual competence, country of birth and residence, duration and conditions of the acquisition of Czech, including an indication of the institution, duration, or location (whether abroad or in the Czech Republic), textbooks used in learning Czech, and whether a family member has been a speaker of Czech. Specifications of the character of the text and circumstances of its production include 7For more details about both types of transcription and anonymization, including the codes, see the CzeSL site http://utkl.ff.cuni.cz/learncorp/ –CzeSL-man (for the HTML-based transcription), or CzeSL in TEITOK (for the XML-based transcription). 8The role of metadata has been emphasized by many authors (e. g., Granger 2003a,2008; Tono 2003).
4.4. METADATA 55 the availability of language reference tools, the extent and type of elicitation, and the temporal and size restrictions. The content of the individual items is listed in Table 4.1 and Table 4.2. Identifications of the items in the first column are used as XML attributes in the text headers of several downloadable and searchable CzeSL corpora.9 s_id Identification of the learner: e. g., TOU_H305 s_sex Sex: mor f s_age Age: e. g., 17 s_age_cat Age category: 6-11,12-15, or 16s_L1 First language: an ISO 639-1 code, e. g., sq (Albanian)10 s_L1_group Language group of the first language: IE (Indo-European non-Slavic), nIE (non-Indo-European), or S(Slavic) s_other_langs Knowledge of other languages: one or more ISO 639-1 codes s_cz_CEF Proficiency in Czech at the time of writing: A1,A1+,A2,A2+, B1,B2,C1, or C2 s_cz_in_family Knowledge of Czech in the family; one or more values: mother,father,partner,sibling,3(3 family members), other,nobody s_years_in_CzR Years in Czechia: -1,1,-2, or 2s_study_cz Past or present study; one or more values: 1to1 (individual tutoring), paid,TY (self-study), university,foreign, primary-secondary,other s_study_cz_months Months of studying Czech: -3,3-6,6-12,12-24,24-36, 36-48,48-60, or 60s_study_cz_hrs_week Hours of studying Czech per week: -3,5-15, or 15s_textbook Textbook used by the learner; one or more values: BC (Basic Czech), CC (Communicative Czech), CE (Čeština pro ekonomy), CMC (Chcete mluvit česky?), CpC (Čeština pro cizince), ECE (Easy Czech Elementary), NCSS (New Czech Step by Step), other s_bilingual Bilingual: yes or no Table 4.1: Metadata about the learner 9In the on-line searchable version of the CzeSL-SGT corpus the metadata items are identified by Czech labels. For a list of English and Czech metadata identifiers see http://utkl.ff.cuni.cz/ ~rosen/public/meta_attr_vals.html. 10See https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes. If necessary, the threecharacter code ISO 639-3 is used, e. g., xal (Kalmyk), see https://en.wikipedia.org/wiki/ISO_ 639-3.
56 CHAPTER 4. PROCUREMENT OF TEXTS t_id Identification of the text: e. g., TOU_H305_442 t_date Date of the text collection: YYYY-MM-DD t_medium Medium of the text: manuscript or pc t_limit_minutes Time limit in minutes: 10,15,20,30,40,45,60,other, or none t_aid Permitted resources; one or more values: yes,dictionary, textbook,other,none t_exam Was the text part of exam?; one or more values: yes,interim, final,n/a t_limit_words Assigned size limit in words: e. g., 150 t_title Title of the essay; one or more values: e. g., Událost, která změnila můj život t_topic_type Type of the topic: general or specific t_activity Activity before writing the text: exercise,discussion, visual,vocabulary,other, or none t_topic_assigned Assigned topic: multiple choice,specified,free, or other t_genre_assigned Assigned genre: free or specified t_genre_predominant Genre predominant in the resulting text: informative, descriptive,argumentative, or narrative t_words_count Actual number of words: integer t_words_range Range of the actual number of words: -50,50-99,100-149, 150-199, or 200Table 4.2: Metadata about the text In most texts, the metadata were specified by the text collector. For some items, such as the proficiency level, rather than applying a set of objectively defined criteria, collectors had to estimate the level by combining their impression of the learner with instructions received during training sessions and included in the collectors’ guidelines. As a result, this metadata item should be taken with a grain of salt. For more about the issue of inaccurate CEFR levels in the CzeSL corpus see §11, page 220. The representation of metadata is not the same in all CzeSL corpora. In the tiered format generated by the feat tool and used in CzeSL-man v1 downloadable (see §8.3.3), metadata concerning a specific text are stored in a separate file, together with files corresponding to the individual tiers and following the same naming convention. For example, a text identified as KAR_MI_005 has its metadata in a file named KAR_MI_005.meta.xml. The metadata items are represented as XML elements, see Figure 7.3, page 138. Releases of CzeSL corpora searchable in KonText have their metadata encoded
4.4. METADATA 57 as XML attributes in the text headers, see Figure 8.1 on page 162. This applies to CzeSL-SGT (including its downloadable version), CzeSL-man v1 searchable, and CzeSL-man v2 (including its downloadable version). Yet, the metadata for the searchable and downloadable releases of CzeSL-SGT also differ: the metadata items are named in Czech in the searchable version.11 In the CzeSL in TEITOK corpus, metadata are represented in a way conformant with the TEI guidelines, as long as they are available for the specific CzeSL items. They are part of the text XML header, for an example see Figure 8.3 on page 171. Metadata can be displayed in a user-friendly way and in a preferred language, depending on the available localization and the setting of the corpus tool for the specific corpus. Like the text itself, metadata can also be edited in the user interface by an authorized user. 11See http://utkl.ff.cuni.cz/~rosen/public/meta_attr_vals.html for a bilingual list of the attributes and values.
64 CHAPTER 5. ERROR ANNOTATION For a discussion about the practical issue whether TH and error tags are better annotated in one go or separately see §7.3.1. 5.2 More than one way to annotate errors in CzeSL Error annotation is a crucial component of most CzeSL corpora. Our approach to error annotation is one of the aspects of the corpus design reflecting the aim to serve many types of users. Rather than focusing on a narrow domain of learner language as the annotation target (such as spelling or lexical errors), the corpus is intended as open to as many research goals as possible. This is one of the main reasons why the target hypothesis is aimed at SCz and the error taxonomy is based on the standard linguistic concepts (spelling, morphology, syntax, semantics, agreement, valency), rather than on categories rooted in the concepts of interlanguage, communication strategy or specific research goals. Following the same approach of not aiming at any specific group of users, we have designed, tested and used two complementary error annotation schemes, and tested and used another one. Our first proposal, referred to as the two-tier annotation scheme – 2T, is based on parallel tiers, representing the source text and supporting successive corrections in two stages: corrections of spelling and all other corrections (see §5.4). The two annotation tiers were introduced as a compromise between several theoretically motivated levels and practical concerns about the process of annotation. They enable the annotators to register anomalies in isolated forms separately from the annotation of context-based phenomena but saves them from difficult theoretical dilemmas. To determine the target hypotheses and to apply the grammar-based error categorization (see §5.4.2) was the task of human annotators. Some of the texts were annotated independently by two annotators and evaluated (see §5.4.4). The tagset used by the annotators, slightly biased towards morphosyntax, and less detailed than most other error tagsets, was meant to be complemented by other annotation. The absence of POS distinctions in the error tags is a way to avoid redundancy in a corpus which is also annotated by POS tags (see §6). The lack of a detailed analysis of errors in spelling and morphonology is to some extent remedied by an automatically applied tagset identifying formal distinctions between the source forms and their corrections (see §5.4.5). Our second proposal, the multidimensional annotation scheme – MD, was developed to complement the 2T scheme by filling the gaps in the categorization of errors in spelling, morphonology and morphology, while allowing for alternative in-
5.3. A WISHLIST FOR ERROR ANNOTATION 65 terpretations of a single error, e. g., as an error which could be explained as an issue of spelling, morphonology, morphology or morphosyntax (see §5.6). The motivation for using implicit annotation (see §5.5), i.e., corrections without error tags, is twofold. Firstly, the full-fledged manual 2T or MD annotation requires a well-trained annotator and more time for the same amount of text. Eliminating the error tagging task makes the perspective to hand-annotate all currently available and new CzeSL texts realistic, while leaving open the option to assign error tags later. Secondly, corrections can be assigned to specific error interpretation levels, corresponding to the tiers in the 2T scheme, or to a more sophisticated system of linguistic domains, as in the MD scheme. Thus, the three error annotation schemes are compatible and complementary. In fact, the three schemes can be implemented in a single corpus. 5.3 A wishlist for error annotation Designing an error annotation scheme for non-native Czech is a challenging task. Czech, at least in comparison to most languages of the existing annotated learner corpora, has a more complex morphology and a less rigid word order, which opens annotation issues that had not been addressed before the error annotation of CzeSL started. As can be expected, the language of a learner of Czech may deviate from the standard in a number of aspects: spelling, morphology, morphosyntax, semantics, pragmatics or style. To cope with the multi-level options of erring in Czech and to satisfy the goals of the project, the annotation scheme should: 1. Properly handle Czech as an inflectional and free-word-order language, e. g., support successive corrections and annotation of errors in discontinuous expressions 2. Be detailed and informative but manageable for the annotators, e.g., preserve the original text alongside with its corrected version and represent syntactic relations for errors in agreement, valency, pronominal reference 3. Be open to future extensions, allowing for alternative/more detailed taxonomy to be added later 4. Provide solutions for issues of interference (see §5.3.1), interpretation (see §5.3.2), word order (see §5.3.3) and style (see §5.3.4) The resulting annotation scheme and the error typology is a compromise between the limitations of the annotation process and the demands of research into learner corpora.
66 CHAPTER 5. ERROR ANNOTATION 5.3.1 Interference and other types of explanation Interference figures prominantly among the candidates for the most relevant explanation of an error. Interference (also called positive or negative language transfer, or crosslinguistic influence) involves an inappropriate use of linguistic features from another language known to the learner, usually their native tongue, or the inappropriate avoidance of such features. A sentence such as Tokio je pěkný hrad ‘Tokio is a nice castle’ is grammatically correct, but its author, a native speaker of Russian, was misled by “false friends”, assuming hrad ‘castle’ as the Czech equivalent of Russian gorod ‘town, city’. Similarly in Je tam hodně sklepů ‘There are many cellars’. The formally correct sentence may strike the reader as implausible in the context. The interpretation becomes clear only with the knowledge that sklep in Polish means ‘shop’, not ‘cellar’ (i.e., sklep in Czech). However, to identify and correct the error without more or less thorough knowledge of the other language is impossible. In practical terms, the identification of all types of interference in a corpus with many L1s is very hard. Most of our annotators were no experts in Czech as a foreign language or in L2 learning and acquisition, and unaware of possible interferences between languages the learner knows. Thus they would have very likely failed to recognize an interferential error. Interference is just one of many types of error diagnostics which is different from grammar-based annotation or other relatively straightforward categorization. The perspective subsuming interference is concerned with the discovery of causes or explanations. Apart from the practical issue of annotating such properties without researching other resources, such as additional texts from the same learner, perhaps at different stages of the acquisition of L2, there is also a theoretical reason why the explanation of errors should be kept separate from the more down-to-earth types of linguistic annotation. Even though all annotation is interpretation, interpretation in terms of grammar-based categories or even stylistic appropriteness is governed by instructions, linguistic rules and/or relations to L1, while finding an explanation for an error can hardly be guided by guidelines. For such reasons, instead of its explicit annotation, interference and error explanations of other types are assumed to be identified by the corpus user in the process of interpreting the corpus data, other types of annotation and the metadata. 5.3.2 Interpretation in terms of TH For some types of errors, the problem is to define the limits of interpretation in terms of TH. Example (9) shows two possible interpretations (TH1 and TH2) of a
5.3. A WISHLIST FOR ERROR ANNOTATION 67 grammatically incorrect clause (S), corresponding to the concepts of “minimal” and “maximal” TH in the Falko corpus (see §2.3.5). The clause is roughly understandable as its TH1 version, but it can also be rewritten as TH2, which is further from the source clause. The TH1 version is less natural but closer to the original. However, to provide annotation in terms of TH2 the task of the annotator is interpretation rather than correction. (9) S: kdyby *citila na tebe *zlobna TH1: kdyby if se refl cítila felt na at tebe you rozzlobená angry ‘if she felt angry at you’ TH2: kdyby if se refl na at tebe you zlobila was-angry ‘if she was angry at you’ Without the option to provide both THs, as in Falko, it is difficult to provide clear guidelines. In the manual annotation of CzeSL, the TH is not supposed to aim at perfect Czech. Instead, the source text is corrected conservatively to arrive at a coherent and well-formed result, without any ambition to produce a stylistically optimal solution, refraining from too loose interpretation. In this sense, the annotator is instructed to minimize interpretation. In general, the ultimate TH in CzeSL is closer to Falko’s TH1 rather than TH2, unless the grammatically correct version is hard to understand or very unnatural.7Where a part of the input is not comprehensible, it is marked as such and left without correction. 5.3.3 Word order Czech constituent order reflects information structure (see §B.3) and it is sometimes difficult to decide (even in a context) whether an error is present.8The sentence rádio je taky na skříni ‘a radio is also on the wardrobe’ suggests that there are at least two radios in the room, although the more likely interpretation is that among other things which happen to sit on the wardrobe, there is also a radio. The latter interpretation requires a different word order: na skříni je taky rádio. In accordance with the preference of conservative target hypotheses (see §5.3.2), word order should be corrected only when it is perceived as ungrammatical. Misplaced 2nd position clitics are a typical example, as in rozhodli se jsme →rozhodli 7For a related discussion about what counts as an error in L2 see §5.1. 8See §5.1 for more about “covert errors”.
68 CHAPTER 5. ERROR ANNOTATION jsme se ‘we have decided’. However, in cases when word order (i) makes a difference in meaning, as in the switched order of the two NPs above (‘the radio’ and ‘the wardrobe’), and (ii) the context makes it clear which meaning is appropriate, word order should be corrected even though it is grammatical. In this sense, word order and lexical corrections share the same approach: correction is due whenever an item or pattern does not fit the meaning of the context. 5.3.4 Style The phenomenon of Czech diglossia (see Appendix §B) is reflected in the problem of annotating non-standard language, usually individual forms with colloquial morphological endings. Because learners may not be aware of the status of these forms and/or an appropriate context for their use, Colloquial Czech (CCz) is corrected under the rationale that the authors expect the register of their text to be perceived as unmarked. To give a prototypical example, one of the most frequent problems in learner texts is the absence of appropriate diacritics. At the same time, a missing acute accents on some verbal endings in written text is perceived as colloquial, because it is supposed to reflect the colloquial pronunciation of these forms: znam →znám ‘I know’, nosim →nosím ‘I wear’. There are nearly 2.5 thousand instances of 180 different apparently colloquial verbs in the 1st person singular in the 1 million CzeSL-SGT corpus. Cases like this are treated as errors (in spelling or morphonology), but they are also labeled as colloquial style, suggesting that the learner could have used a colloquial instead of an incorrect form.9It is up to the user of the corpus to interpret the annotation according to a wider context or the learner’s profile in the metadata. 5.3.5 Communication goal Other features of the learner language may also be considered as candidates for annotation, such as a measure estimating to what extent the learner’s communication goal is achieved. In fact, there is hardly anything that matters more in practice and could be reflected even at the level of individual utterances. 9The colloquial marker is in fact a category from the domain of linguistic rather than error annotation. However, it is used in the manual error annotation to make the point that some forms annotated as incorrect can also be interpreted as colloquial forms. The colloquial marker can be confirmed in the annotation provided by a tagger applied to the source text, although the automatic linguistic annotation of such forms may be less reliable then of their SCz counterparts.
5.4. THE TWO-TIER ANNOTATION SCHEME 69 On the other hand, not every aspect of the learner language must be explicitly annotated. It could even be a more proper move to leave some of the trickier phenomena such as interference or exhaustive interpretation for the corpus user while providing a reliable annotation of errors where a safer ground is available in linguistic theory, established categories and the annotators’ competence. Such annotation, provided ideally by the combination of the three annotation schemes, can help the user to interpret the search results or statistical findings in ways not previewed in the annotation. 5.4 The two-tier annotation scheme The two-tier scheme, including its error tagset, was designed to suit the specifics of learner Czech. In this respect, the scheme proved to be adequately expressive and practically useful.10 As the most sophisticated of the three schemes used in the CzeSL project, it deserves to be presented and examined from multiple angles, together with its merits and drawbacks. At first, we focus on foundations of the scheme, namely on why there are several parallel tiers, why exactly two tiers representing up to two successive THs, how the words represented at those tiers are related and how the errors can be tagged (see §5.4.1). The rest of the section is concerned with the error tagset (see §5.4.2), its evaluation (see §5.4.4), and a complementary “formal” tagset, used in rules comparing source forms and their corrections, without human intervention (see §5.4.5). 5.4.1 Annotation scheme as a compromise 5.4.1.1 Why multiple tiers After a careful examination of available options, we have arrived at a two-stage annotation design, consisting of three parallel tiers: a tier of the source text and two annotation tiers. Between the two opposite options of a flat inline annotation and a scheme consisting of more parallel tiers (see §2.2.2.5), we decided for a compromise solution. The choice of a multi-tier annotation scheme with a specific number of tiers calls for some justification. The optimal error annotation strategy is determined both by the goals and resources of the project and by the type of the language. A simple 10One of the two case studies in Meurers (2015) presents the scheme as a showcase example of a “state-of-the-art learner corpus annotation project integrating insights and tools from NLP”.
70 CHAPTER 5. ERROR ANNOTATION flat scheme with all annotation inline could be used for a specific narrowly defined purpose, such as investigation of morphological properties of the learner language, or for a language without an elaborate inflection system. Such a scheme can be appealing if corrections concern individual word forms or contiguous sequences of forms and successive or alternative corrections are not required. A scheme with a tier for the original text and a single parallel annotation tier would be appropriate if we were interested only in the original text and in the annotation at some specific level (fully emended sentences, or some intermediate stage, such as corrected word forms). This design could be used even if we insisted on registering some intermediate stages of the passage from the original to a fully emended text, and decided to store such information with the word-form nodes. However, such information might get lost in the case of significant changes involving deletions or additions. For example, in Czech as a pro-drop language, the annotator may decide that a misspelled personal pronoun in the subject position should be deleted. Then the information about the spelling error would disappear. Given the goals of the project and the properties of Czech, either of the two solutions – the inline annotation and a single annotation tier – was problematic. There were at least three reasons: 1. The corpus should be open to multiple research goals. Thus, it would not do to accommodate the analysis of a restricted set of linguistic phenomena within the inline annotation or a single tier. 2. Due to the fairly rich morphology and the relatively free word order of Czech, it is necessary to provide space for successive corrections. At the same time, it is important to maintain links between the original and the corrected forms even when the word order changes or when words are dropped or added. Otherwise it would be difficult to find the ultimate target hypothesis for a faulty expression or to find a corresponding expression in the original text given its target hypothesis. 3. Learner texts include word-boundary errors, i.e., incorrectly split or joined forms, or errors spanning multiple forms, even in discontinuous positions. The most natural way to annotate errors of this type with a target hypothesis and error label is in a multi-tier annotation scheme. Actually, the decision to use a multi-tier design was mainly due to our interest in annotating errors in single forms as well as those spanning (potentially discontinuous) strings of words.
5.4. THE TWO-TIER ANNOTATION SCHEME 71 5.4.1.2 How many tiers Once we have a scheme of multiple tiers available, we can provide them with theoretical significance and assign a linguistic interpretation to each of them. In a world of unlimited resources of annotators’ time and experience, this would be the optimal solution. Annotators could be free to use an arbitrary number of tiers to suit the needs of successive emendations. They could choose from a set of linguistically motivated tiers or introducing annotation tiers ad hoc. The first annotation tier would be concerned only with errors in graphemics, followed by tiers dedicated to morphemics, morphosyntax, syntax, lexical phenomena, semantics and pragmatics. More realistically, there could be a tier for errors in graphemics and morphemics, another for errors in morphosyntax (agreement, government) and one more for everything else, including word order and phraseology. On the other hand, annotators should not be burdened with theoretical dilemmas and the result should be as consistent as possible, which somewhat disqualifies a scheme using a flexible number of tiers. This is why we adopted a compromise solution with two tiers of annotation, distinguished by formal but linguistically founded criteria to make the annotator’s decisions easy. It is a compromise between an inline or single-tier annotation and an open multi-layer format, but a compromise preserving links between split, joined and re-ordered tokens, corrected in two stages, something not obviously supported in the multi-layered tabular format described below in §5.4.1.3. Each of the choices made in the design of the annotation scheme is a compromise between its feasibility in a practical large-scale annotation process and the requirement of a detailed and complex analysis. The restriction in the number of annotation tiers has proved its feasibility while still being useful and linguistically relevant. 5.4.1.3 Multiple tiers in a tabular format Many corpora use simple inline error annotation, denoting the scope, correction and categorization of an error. A few corpora such as Falko (see §2.3.5) adopt multi-tier annotation in a tabular format, with the option of specifying multiple corrections and several error types for single word tokens or strings thereof at several linguistically motivated tiers: orthography, morphology, syntax, lexicon, pragmatics, intelligibility. The tabular format is also used in MERLIN (see §2.3.7), one of the two currently available corpora including Czech. The format and the corresponding tools were considered also for the manual two-tier annotation of the CzeSL texts. Originally, the multi-tier tabular format and related tools were designed for an-
72 CHAPTER 5. ERROR ANNOTATION notating speech. The environment allows for an arbitrary segmentation of the input and multi-tier annotation of segments (Schmidt 2009). Typically, the annotator edits a table with columns corresponding to words and rows corresponding to tiers. A cell can be split or more cells merged horizontally to allow for annotating smaller or larger segments. This way, phenomena such as agreement or word order can be emended and tagged (Lüdeling et al. 2005). However, the tabular format is not quite suitable for languages with free word order and rich inflection, where a single form may be incorrect in several domains at once: typography, orthography, morphosyntax, lexicon, word order. In the tabular format, vertical correspondences between the original word form and its corrected equivalents or annotations at other tiers may be lost. It is difficult to keep track of links between forms merged into a single cell, spanning multiple columns, and the annotations of a form at other tiers (rows). This may be a problem for successive corrections involving a single form, starting from a typo up to an ungrammatical word order, but also for morphosyntactic tags assigned to forms, whenever a form is involved in a multi-word annotation, and its equivalent or tag is no longer present in the column of the source form. 5.4.1.4 Content of the tiers As a compromise between corpus users’ expected demands and limitations due to the annotators’ time and experience, the two-stage annotation design reflects the distinction roughly between errors in orthography and morphemics on the one hand and all other error types on the other. The scheme consists of three interconnected tiers – see Figure 5.1 for an annotated example glossed in (10).11 Annotation tiers are represented as a graph consisting of a set of interlinked parallel paths where a path is a sequence of word forms corresponding to a sentence at a given level. Each word in the input text is represented at every level, unless it is split, joined (as kdy by in Figure 5.1), deleted or added by the annotator. Whenever a word form is corrected, the type of error can label the link connecting the incorrect form with its corrected version (such as incorInfl or incorBase for morphological errors in inflectional endings and stems, stylColl as a stylistic marker, wbdOther as a word boundary error, and agr as an error in agreement). Tier 0 (T0) – Anonymized transcript of the hand-written original string of graphemes, with some properties of the manuscript preserved in the transcription mark-up (self-corrections, variants, illegible strings). 11Figure 5.1 is a screenshot of the annotation editor feat (see §9.1.1).
5.4. THE TWO-TIER ANNOTATION SCHEME 73 Tier 1 (T1) – The tier of orthographic and morphological normalization. As a rule of thumb, this is where forms incorrect in isolation are corrected. The result is a string consisting of correct Czech forms, even though the sentence may not be correct as a whole. The rule of “correct forms only” has a few exceptions: a faulty form is retained if no correct form could be used in the context, or if the annotator cannot decipher the author’s intention. On the other hand, a correct form may be replaced by another correct form if the author clearly misspelled the latter, creating an unintended homograph with another form. A formally correct form weak in a sentence such as I’ll see you in a weak would be corrected since the author clearly misspelled the form she intended to use, creating an unintended homograph. On the other hand, the form week in I’ll see you in two week is an error in morphosyntax and will be corrected at T2. Tier 2 (T2) – Handles all other deviations, resulting in a grammatically correct sentence. This includes errors in syntax (agreement, government), lexicon, word order, usage, style, reference, negation, or overuse/underuse. (10) T0: *Myslim think.(sg1) že that *kdy if *by would.*3rd byl was.masc se with *svim self’s *ditem child … T2: Myslím, think.sg1 že that kdybych if+would.1sg byl was.masc se with svým self’s dítětem, child … ‘I think that if I were with my child, …’ (KKOL_AV_007 ru B1) A more complex example is presented below in §5.4.1.5. 5.4.1.5 A sample text with T1 vs. T2 corrections To exemplify various types of deviations of L2 Czech from the standard, the sample text in Table 5.1 highlights errors according to the tier they are corrected. Forms wrong in any context due to an error in spelling or morphology, corrected at T1, are set in boldface, while forms wrong due to a morphosyntactic or lexical anomaly, corrected at T2, are underlined. Some forms may be faulty for both reasons; these are in bold and underlined. 5.4.1.6 Links between tiers While in the tabular format the correspondences between elements at various tiers are captured implicitly, in our annotation scheme these correspondences are explicitly encoded. The format supports the option of preserving correspondences across
80 CHAPTER 5. ERROR ANNOTATION 5.4.2.2 Grammar-based vs. formal errors For practical reasons we have abandoned the idea of alternative target hypotheses (see §5.4.1.10). However, this does not exclude the option to use alternative error categories for a single error with a single target hypothesis. For example, some errors in spelling can also be interpreted as errors in morphemics and morphosyntax. It is hard to decide which of the interpretations is correct without some more research into the individual learner’s competence in Czech. In the 2T scheme, the rule of thumb for choosing the appropriate interpretation is to prefer the “more sophisticated” error type, i.e., morphosyntax rather than spelling. As a result, the 2T scheme does not support alternatives in error categorization either. To compensate for this rather strict restriction, the grammar-based tags, for the most part manually assigned, are complemented by “formal” errors, assigned automatically. For the MD scheme, systematically supporting multiple interpretations, see §5.6. A single incorrect form is cross-classified as belonging to one or more types in each of the following two classes: • Grammar-based error types – a taxonomy classifying errors from a grammatical perspective: errors in spelling, morphology, word boundary, agreement, government, lexical issue, style, punctuation; these error types are similar to the “linguistic category classification” (James 1998, 104–113) or “linguistically based errors” (Lüdeling and Hirschmann 2015, 146) • Formal error types – error types capturing the formal nature of an error without referring to possible underlying grammatical reasons: diacritics, capitalization, metathesis, missing element; also called “target modification taxonomy” (James 1998, 104–113) or “edit-distance based errors” (Lüdeling and Hirschmann 2015, 146) Unlike the grammar-based types, the formal errors are detected by automatic tools by comparing the source forms with their corrections and the manually assigned tags. Yet the tools can also detect some of the grammar-based error types. Thus, errors can be identified in the following ways: • manually • automatically, by comparing the faulty and the corrected forms • automatically, by subdividing certain manually assigned error tags, often on the basis of the relevant word forms, their morphological tags or lemmas
5.4. THE TWO-TIER ANNOTATION SCHEME 81 5.4.2.3 Extent of the annotated unit In the 2T scheme, the minimal annotated text units are tokens. There are three exceptions: 1. For some error types the locus of the error is identified. To signal that a nonword is ill-formed, we manually distinguish an error in the stem (incorBase) from an error in the inflectional ending (incorInfl). 2. Word boundary errors are annotated by joining multiple tokens or splitting a single token. 3. The formal error tags are typically concerned with phenomena at the level of morphs or characters, but without specifying the exact location and only to the extent the error can be identified by an algorithm. The exact locus of the error within a word is identified in the MD annotation scheme (see §5.6). In the following, we first describe the grammar-based tags (§5.4.3). Then the grammar-based tagset is evaluated (§5.4.4). The formal error tags are discussed in §5.4.5. 5.4.3 Grammar-based tags Below, we describe the grammar-based tagsets for each tier separately. The tagset consists of 22 error tags, 8 for T1, 11 for T2, and 3 that can be used at both tiers. After a brief discussion of the granularity of the tagset, we show a commented example of a sentence annotated with the tagset. 5.4.3.1 Errors at T1 Errors in individual word forms, treated at T1, include misspellings (also diacritics and capitalization), misplaced word boundaries but also errors in inflectional and derivational morphology and unknown stems – fabricated or foreign words. Except for misspellings, all these errors are annotated manually. The result at T1 is the closest correct form, which can be further modified at T2 according to context, e.g., due to an error in agreement or semantic incompatibility of the lexeme. Table 5.2 lists the errors manually annotated at T1. Some error types (stylColl, stylOther and problem) are used also at T2. Some error categories, such as incor, have subtypes. While some of these subtypes are tagged manually (incorBase and
82 CHAPTER 5. ERROR ANNOTATION incorInfl), other tags are added automatically: in the absence of any other tag at T1, a corrected form satisfying some conditions is tagged as incorOther. The column with the heading “A” identifies whether the tag is assigned Manually or Automatically. The ↑column indicates the number of edges going upwards from the error label, located in between the tiers, in the source text direction, i.e., towards T0 from a T1 error or towards T1 from a T2 error. The ↓column indicates the number of edges going downwards from the error label towards the target hypothesis, i. e., towards T1 from a T1 error or towards T2 from a T2 error. The ↔column identifies how many outgoing morphosyntactic references are allowed for a given error type. Error type Error subtype Description Example A ↑↓↔ incor incorrect form incorInfl incorrect inflection pracovají→pracují v továrně; bydlím s matkoj→matkou M 1 1 0 incorBase incorrect word base lidé jsou moc mérní→mírní; musíš to posvětlit→vysvětlit M 1 1 0 incorOther other incorrect forms rád pivuju→piju pivo A 1-n1-n0 fw foreign, coined, unidentified word fwFab made-up, coined word pokud nechceš slyšet smášky→posměšky M 1 1 0 fwNc foreign word váza je na Tisch→stole; jsem ve truong→škole M 1 1 0 flex used with fw.. to mark inflection jdu do shopa→obchodu M 1 1 0 wbd wrong word boundary wbdPre prefix separated by a space or preposition without space musím to při pravit → připravit;veškole →ve škole M 1-2 1-2 0 wbdComp wrongly separated compound český anglický → česko-anglický slovník M 2-n1 0 wbdOther other word boundary error mocdobře →moc dobře;atak →a tak;kdy koli →kdykoli M 1-n1-n0 styl colloquial, bookish, dialect stylColl colloquial form dobrej→dobrý film M 0-n0-n0-n stylOther bookish, dialectal, slang, hyper-correct holka s hnědými očimi → hnědýma očima M 0-n0-n0-n problem problematic cases M 0-n0-n0-n Table 5.2: Grammar-based error tags at T1
5.4. THE TWO-TIER ANNOTATION SCHEME 83 5.4.3.2 Errors at T2 Corrections at T2 concern errors in agreement, valency, analytical forms, pronominal reference, negative concord, the choice of aspect, tense, lexical item or idiom, and also in word order. For the agreement, valency, analytical forms, pronominal reference and negative concord cases, there is usually a correct form, which determines some properties (morphological categories) of the faulty form. Table 5.3 shows a list of error types manually annotated at T2. The automatically identified errors include word order errors and subtypes of the analytical forms error vbx. 5.4.3.3 Coarse-grained In comparison to some other error tagsets, the taxonomy is relatively coarse-grained. There are several reasons: • We assume that error annotation can be combined with linguistic annotation, both in queries or in statistical analysis. For example, an error in agreement need not specify further that the incorrect form is an adjective because that information is available from the POS tag. Linguistic annotation for the source and corrected forms is provided by automatic tools. • Whenever the type of error can be determined from the way the incorrect form is corrected, the type is supplied by an automatic post-processing step (see §5.4.5). • Ever since the first design of the two-tier annotation scheme, we expected to provide a more detailed classification of errors or a classification of errors from other perspectives later. So far, two additional error annotation schemes – the multidimensional (MD) scheme (see §5.6) and the implicit annotation scheme (see §5.5) – and one linguistic annotation scheme, based on the Universal Dependencies guidelines (see §6.2), have been designed and tested on a smaller corpus sample. • Too detailed tags could be hard to apply consistently even by a single annotator. On the other hand, the usability of a tagset can be measured by the IAA score and depends on more factors than its granularity, such as well-defined denotation (see §6.2.5). Moreover, tags cross-classifying errors by using a clearly defined POS or grammar domain component can be easily applicable despite their large number in the tagset. Still, we see the relatively low number of error tags as both a practical and theoretical advantage.
84 CHAPTER 5. ERROR ANNOTATION Error type Error subtype Description Example A ↑↓↔ agr violated agreement rules to jsou hezké→hezcí chlapci;Jana čtu→čte M 1 1 0-n dep error in valency bojí se pes→psa;otázka čas→času M 0-1 0-1 0-n ref error in pronominal reference dal jsem to jemu i jejímu→jeho bratrovi M 1 1 0-1 vbx error in analytical verb form or compound predicate M 1-n1-n0-1 cvf analytical verb kluci jsou→∅ běhali M 1-n1-n0-1 mod modal or phase verb musíš přijdeš→přijít M 1-n1-n0-1 vpn compound predicates Petr má→je unavený M 1-n1-n0-1 rflx error in reflexive expression dívá ∅→se na televizi; Pavel si→se raduje M 0-n0-n0-n neg error in negation nikdo to ví→neví;půjdu ne →nepůjdu do školy M 1-n1-n0-1 odd redundant item Petr dělá→∅ čte M 1 0 0 miss missing item není ∅→to tak dávno M 0 1 0 wo wrong word order mají hezký velmi dům → mají velmi hezký dům A 1-n1-n0 lex error in lexicon or phraseology jsem ruská→Ruska; dopadlo to přírodně→přirozeně M 0-n0-n0-1 use error in the use of a grammar category pošta je nejvíc blízko → nejblíž M 1-n1-n0-1 sec secondary error (supplementary flag) stará se o našich rybičkách →naše rybičky M 1-n1-n0-n styl colloquial, bookish, dialect stylColl colloquial expression viděli jsme hezký→hezké holky M 0-n0-n0-n stylOther bookish, dialectal, slang, hyper-correct expression zvedl se mi kufr→žaludek M 0-n0-n0-n stylMark redundant discourse marker no;teda;jo → ∅ M 1 0 0 disr intelligible or disrupted construction kratka jakost vyborné ženy →? Mn n 0 problem supplementary label for problematic cases M 0-n0-n0-n Table 5.3: Grammar-based error tags at T2 5.4.3.4 An example of complex annotation Splitting, joining and reordering words, together with the morphosyntactic references, may result in a complex network of both labeled and unlabeled links, as in
5.4. THE TWO-TIER ANNOTATION SCHEME 85 Bojal jsme *feared aux incorInfl Bál jsme agr rflx Bál jsem se , I was afraid že ona se ne bude libila slavnou prahu , that she rflx not will *like famous Prague , wbdPre incorBase že ona se nebude líbila slavnou Prahu , dep vbx agr,sec dep že se jí nebude líbit slavná Praha , that she would not like the famous city of Prague, proto to bylo velmí vadí pro mně . therefore it was *very resent for me . incorBase proto to bylo velmi vadí pro mně . lex vbx dep protože to by mi velmi vadilo . protože to by mi velmi vadilo. 1 Bojal jsme *feared aux incorInfl Bál jsme agr rflx Bál jsem se , I was afraid že ona se ne bude libila slavnou prahu , that she rflx not will *like famous Prague , wbdPre incorBase že ona se nebude líbila slavnou Prahu , dep vbx agr,sec dep že se jí nebude líbit slavná Praha , that she would not like the famous city of Prague, proto to bylo velmí vadí pro mně . therefore it was *very resent for me . incorBase proto to bylo velmi vadí pro mně . lex vbx dep protože to by mi velmi vadilo . protože to by mi velmi vadilo. 1 Bojal jsme *feared aux incorInfl Bál jsme agr rflx Bál jsem se , I was afraid že ona se ne bude libila slavnou prahu , that she rflx not will *like famous Prague , wbdPre incorBase že ona se nebude líbila slavnou Prahu , dep vbx agr,sec dep že se jí nebude líbit slavná Praha , that she would not like the famous city of Prague, proto to bylo velmí vadí pro mně . therefore it was *very resent for me . incorBase proto to bylo velmi vadí pro mně . lex vbx dep protože to by mi velmi vadilo . because I would be very unhappy about it. 1 Figure 5.2: Two-level manual annotation of a sentence in CzeSL, the English glosses are added (VOB_KA_049 kk A1) Figure 5.2 (re-drawn from the feat representation for clarity and to save space). It is an authentic sentence, split in two parts for space reasons. • As in Figure 5.1, the three parallel strings of word forms represent the three tiers: the tier of transcribed input and the two annotation tiers. The tiers are parallel strings of word forms with links for corresponding forms. The asterisked forms in the English glosses below mark forms that are incorrect in any context, but they may be comprehensible – as is the case with all such forms in this example. • Correct words are linked directly with their copies at T1, for corrected words most links are labeled with an error type. The first line is T0, imported from the transcribed original. • T0 is followed by the level of orthographic and morphemic corrections (T1), where only forms incorrect in any context are treated. Errors at T1 are mainly non-word (OOV) errors while those at T2 are real-word and grammatical errors. However, a faulty form that happens to be spelled as a form which
86 CHAPTER 5. ERROR ANNOTATION would be correct in a different context, is still corrected at T1. Thus the result at T1 is a string consisting of correct Czech forms, even though the sentence may not be correct as a whole. • All other types of errors are corrected at T2, representing a grammatically correct, though stylistically not necessarily optimal target hypothesis. A syntactic error label may be linked by a pointer to a word token, specifying an agreement, valency or referential relation. • In the first (upper) part of the sentence, a form is labeled at T1 as an error in inflectional ending (incorInfl)bojal →bál, and another form as an error in the word stem (incorStem)libila →líbila. The rest of the T1 errors are purely orthographic: according to the rules of Czech spelling, the negative particle ne is joined with the verb using the error label wbdPre, and the initial character of the place name is capitalized (the error label is assigned automatically). • Staying with the first part of the sentence, at T2 another form is emended as an error in agreement (jsme →jsem ‘am’) with reference to a form exhibiting the correct morphological category of singular number (bál). • The missing reflexive particle se is inserted with reference to the inherently reflexive verb and the comma is inserted without any label, because this type of error is identified automatically. • The second reflexive particle se, a second position clitic, is misplaced in the source text and should be reordered (the wo label for a word-order error is assigned automatically). • The pronoun ona ‘she’ in the nominative case is governed by the form líbit se and should be assigned the dative case: jí ‘her’, with reference to the head verb. • The head verb has changed its finite form líbila into the infinitive, because it is now a part of the analytical future tense, identified by the error type vbx and a link to the future auxiliary. • The accusative case of Praha in the source is changed into nominative, again with a reference to the governing verb. The form of the adjective slavnou must be modified accordingly with an additional label sec as a secondary (follow-up) error.
5.4. THE TWO-TIER ANNOTATION SCHEME 87 • The result could still be improved by positioning Praha after the clitics and before the finite verb nebude, resulting in a word order more in line with the underlying information structure of the sentence, but our policy is to refrain from more subtle phenomena and produce a grammatical rather than a perfect result. • In the second (lower) part of the sentence, there is only one T1 incorBase error in diacritics (velmí →velmi ‘very’), but quite a few errors at T2. •Proto ‘therefore’ is changed to protože ‘because’ as a lexical error (lex). • The main issue in the second part of the sentence are the two finite verbs bylo ‘was’ and vadí ‘resents’. The most likely intention of the author is best expressed by the conditional mood. The two non-contiguous forms are replaced by the conditional auxiliary and the content verb participle in one step using a 2:2 relation. The intermediate node is labeled by vbx for complex verb forms. • The prepositional phrase (PP) pro mně ‘for me’ is another complex issue. Its proper form is pro mě (homonymous with pro mně, but with ‘me’ bearing accusative instead of dative), or pro mne. The accusative case is required by the preposition pro. However, the head verb requires that this complement bears bare dative – mi. Additionally, this form is a second position clitic, following the conditional auxiliary (also a clitic) in the clitic cluster. The change from PP to the bare dative pronoun and the reordering are both properly represented, including the pointer to the head verb. • What is missing is an explicit annotation of the faulty case of the prepositional complement, which is lost during the transition from T1 to T2. This is the price for a simpler annotation scheme with fewer levels. It might be possible to amend the PP at T1, but it would go against the rule that only forms wrong in isolation are corrected at T1.14 5.4.4 Evaluation of the manual tiered error annotation To evaluate the consistency of annotation of learner corpora, texts are annotated independently by two or more annotators and the results are compared. This is an approach commonly used for many other types of manual annotation. 14The implicit annotation scheme (see §5.5) can accommodate a case like this due to a more flexible system of successive corrections. In the setup of correction levels for CzeSL in TEITOK (see §8.8), mně →mě would be corrected as an orthographic correction (ort) while pro mě →mi as a morphosyntactic correction (gram).
88 CHAPTER 5. ERROR ANNOTATION However, this was not always the standard practice in learner corpus research. The issue of singly annotated learner texts, used as application training data, was raised for the first time by Tetreault and Chodorow (2008), who investigated nativespeakers’ classification of prepositions usage. They concluded that two native annotators performing the task of tagging errors in prepositions on the same text reach at best an agreement level on the border between moderate and substantial (their kappa value was κ= 0.63 – the metric is explained in §5.4.4.1 below). Rozovskaya and Roth (2010) also report low inter-annotator agreement (κ= 0.16−0.40) for the task of classifying sentences written by ESL learners. Meurers (2009) discusses the issue of verification of error annotation validity, viewing the lack of studies investigating inter-annotator agreement in the manual annotation of non-native speakers texts as a serious barrier for the development of annotation tools. 5.4.4.1 Inter-annotator agreement (IAA) The manual annotation of CzeSL was evaluated using the metric κ(kappa, Cohen 1960), the standard measure of inter-annotator agreement, especially for tagged corpora. It is calculated as: κ=P(A)−P(E) 1−P(E) where P(A)is the observed agreement among the annotators, and P(E)is the expected agreement, i.e., P(E)is the probability that the coders agree by chance. The values of κare within the interval [−1,1], where κ= 1 means perfect agreement, κ= 0 agreement equal to chance, and κ= 1 “perfect” disagreement. The problem is to determine which error tags in one annotation correspond to which error tags in the other and how their scopes align. T0, the source text, is shared by both annotations. However, annotators might use a different target hypothesis, and thus the higher tiers can differ. Moreover, they often differ not only in the shape of tokens but also in their number. Because of this, we project error tags to T0 tokens and then calculate differences relative to that tier. When there are multiple tokens on T0 corresponding to a token on the relevant tier, we project the tag on the first T0 token only.
5.4. THE TWO-TIER ANNOTATION SCHEME 89 5.4.4.2 A pilot annotation Early in the project, we calculated IAA on a pilot sample.15 It consisted of 67 texts totalling 9,848 tokens, most of them written by native speakers of Russian; the texts are classified according to the CEFR scale as A2 or B1. The sample was corrected and assigned error tags according to the error taxonomy presented above in §5.4.2 by 14 annotators. They were split into two groups. Each group annotated the whole sample independently. On average each annotator processed 1,475 words in 11 texts. Table 5.4 summarizes the distribution of selected error tags for the pilot sample and for all doubly annotated texts available at the time of the evaluation. The first column gives the error tag; some tags (marked with an asterisk) are used only in the evaluation as a more general error category.16 The column headed by ‘avg tags’ gives the number of times the tag was used by an average annotator (calculated simply as the total for the two annotators divided by two). For the comparison we only considered deviant (error-annotated) forms. Table 5.4 shows differences in error tags, ignoring potentially different THs. However, different THs are an obvious reason for disagreements in the error tags. For example, the THs were different in 54% of cases when the annotators did not agree on the use of the agr tag (κ= 0.54). The differences in THs were either on T1 (15%) or on T2 (39%). In a more extensive evaluation described below the impact of TH on the IAA in error tags was explored in more detail. 5.4.4.3 IAA on all doubly-annotated texts Using the feedback gained from the pilot experiment we modified the definition of some tags, refined the annotation guidelines and improved the training of annotators. In a few cases we also slightly modified the error taxonomy. A substantially larger subset of the transcribed texts was annotated by 31 annotators in three groups specializing on Slavic, non-Slavic and Roma learners. The evaluation was extended 15For complete results see Štindlová et al. (2012). Note that when there are multiple tokens on T0 corresponding to a token on the relevant tier, the tag was projected to all such tokens rather than to the first T0 token. Also note that the numbers for incorInfl and incorStem are switched by mistake in the reported results. 16As described in §5.4.2, the error taxonomy is hierarchical: the error types are partitioned into domains, which are further divided into more specific subcategories, tagged manually or automatically. For example, the domain of complex verb form errors on T2 can be further specified as errors in analytical verb forms (cvf), modal verbs (mod), verbo-nominal predicates, passive or resultative form (vnp).
96 CHAPTER 5. ERROR ANNOTATION tag total same emendations different emendations κavg. tags κavg. tags κavg. tags T1 incor* 0.88 14,380 0.95 12,376 0.48 2,004 incorBase 0.82 10,780 0.89 9,323 0.44 1,456 incorInfl 0.71 4,679 0.79 3,887 0.36 791 wbd* 0.56 840 0.71 525 0.33 315 wbdPre 0.75 484 0.90 336 0.40 148 wbdOther 0.69 842 0.90 479 0.41 363 wbdComp 0.22 58 0.38 23 0.12 34 fw* 0.36 423 0.45 235 0.24 187 fwNc 0.30 298 0.31 165 0.28 132 fwFab 0.09 125 0.13 70 0.04 55 stylColl 0.44 1,396 0.51 1,088 0.20 307 T2 agr 0.69 2,622 0.82 2,050 0.24 572 dep 0.58 3,064 0.71 2,303 0.19 760 rflx 0.42 141 0.58 98 0.05 43 lex 0.32 1,815 0.53 847 0.14 968 neg 0.23 48 0.62 16 0.03 32 ref 0.16 115 0.13 70 0.20 45 sec 0.26 415 0.43 224 0.06 191 stylColl 0.39 633 0.53 403 0.14 230 use 0.39 696 0.61 399 0.10 296 vbx 0.17 233 0.25 135 0.07 98 Table 5.7: IAA depends on emendation agreement 3. Different target hypotheses (see §5.4.4.4). Some annotations require a considerable amount of interpretation, while each annotator can have her/his own interpretation because of age, gender, education, etc. Moreover, in the case of multi-tier annotation, annotators can differ also on intermediate tiers, even though their target hypothesis might be identical. However, the annotation scheme of CzeSL, supporting corrections on both tiers, makes reasons for possible disagreements explicit.
5.4. THE TWO-TIER ANNOTATION SCHEME 97 5.4.5 Formal tags 5.4.5.1 Automatic extension and modification of error annotation From the first designs of the 2T CzeSL error annotation, it was assumed that, following the manual annotation, some types of errors would be annotated automatically. The automatic annotation would add a different point of view on the errors, based on a simple comparison of the source and corrected words. Manual annotation of errors on T1 is relatively simple, and the automatic annotation of errors can give the user much more detailed information about the error and sometimes even the most likely cause of the error. For example, the word form hřipku, used instead of the correct form chřipku ‘flu.sg.acc’, is probably not a case of a character omission, but an error in voicing (formVcd1) – the character his prototypically used for the voiced phoneme /H/, while the digraph ch for the voiceless phoneme /H/ ([h] and [x] form a voicing pair in Czech). The manual annotation of this error only assigns a simple distinction that the word contains an unspecified error in the stem (flective base of the word): incorBase. The automatic error annotation in CzeSL on T1 addresses errors of a formal nature, i. e., (broadly) orthographic errors, such as incorrect capitalization, wrong use of diacritics or wrong choice of the characters i↔y, and errors reflecting wrong pronunciation, such as voicing or (so called) hard and soft consonants, e. g., d↔ď: the character ď(dwith a caron) marks in Czech the phoneme /é/ (voiced palatal plosive). Apart from formal error detection, the automatic error annotation refines the error tagging of errors in word boundaries (incorrect division or joining of word forms). The T1-formal errors make no distinction whether the error occurs in the stem or in the inflectional affix; this distinction is assumed to come from the manual annotation: incorBase or incorInfl. The reason is that morpheme boundaries are often blurred in Czech (as in many other inflective languages) and it is not trivial to distinguish errors in inflection from errors in the stem automatically. Some T2 errors are annotated automatically to some extent – the annotators mark errors with a more general tag and an automatic procedure automatically assigns a detailed tag. The automatic annotation of T2 uses only a limited amount of information about morphological tags and lemmas. We do not attempt to identify the cause of the error. A form may be incorrect due to an incorrectly applied morphological paradigm or due to an incorrect syntax structure of the sentence. Instead, we only further specify or complement manually assigned tags. For example, an error in a periphrastic verb form is manually corrected and labeled with a general error tag,
98 CHAPTER 5. ERROR ANNOTATION the automatic annotation then adds a more detailed tag distinguishing past-tense errors and modal-verb errors. 5.4.5.2 Automatic detection of formal errors on T1 Automatic extension of the error annotation on T1 is performed for those T0 forms that are corrected on T1. It is based on a comparison of the source T0 form and the corrected T1 form. The result is the assignment of formal error tags: broadly orthographic errors, pronunciation-related errors. All such tags are prefixed by the form… string. Moreover, manually annotated errors in word boundaries (wbd) are specified as words incorrectly joined or split. For a list of all formal error tags on T1 see Table 5.8 and Table 5.9 (page 106–107). The algorithm identifies individual differences between the corresponding T0 and T1 forms (delamé →děláme ‘make.1pl’ contains three individual differences: e →ě,a→á,é→e) and assigns an error tag to each difference. The 2T CzeSL annotation scheme does not track the exact location of error, therefore error tags are assigned to the whole word. If there are multiple errors, the word is assigned multiple tags (delamé →děláme:formCaron0|formQuant0|formQuant1). A single difference can also be assigned multiple error tags (the difference i→ýin úteri → úterý ‘Tuesday’ is assigned formY0|formQuant0 to mark the confusion of i→yand a missing accent y→ý). For the formal tags, we use the following convention: they end with 0for incorrectly missing phenomena and 1for incorrectly realized phenomena. For example, incorrect spelling of words due to voicing assimilation can be marked either with formVcd0 or with formVcd1:pohátková →pohádková ‘fairytale.adj’ uses voiceless t instead of the correct voiced dand is therefore marked with formVcd0, while svadba →svatba ‘wedding’ uses voiced dinstead of the correct voiceless tand is therefore marked with formVcd1. For expository reasons, we classify automatically assigned formal tags into two groups: (broadly) orthographic errors, and formal errors affecting pronunciation. Orthographic errors concern misspellings that do not affect pronunciation of the misspelled form by native speakers. They are often the types of errors that even native speakers commonly make in their texts. Formal errors affecting pronunciation include errors in which there would be a noticeable difference in pronunciation between the original and the corrected form. The most common errors of this type include missing diacritics indicating the quantity of vocals or softening of consonants. We list only errors that actually occurred in student corpora; other errors are handled by the algorithm as well, but they are not present in the examined texts.
5.4. THE TWO-TIER ANNOTATION SCHEME 99 5.4.5.3 Formal orthographic errors Automatically identified orthographic errors include errors in capitalization. Czech uses capital letters to mark sentence beginnings and proper names. Rarely, they are also used for emphasis of certain words. The error tag formCap1 marks words that should be written with a lower-case letter, but are capitalized (ona Rozumí →rozumí trochu český ‘she Understands →understands little Czech’). Missing capitalization is labeled with formCap0 (Líbí se mi praha →Praha ‘I like prague →Prague’). Another group of orthographic errors are errors in the spelling of vowel groups with ě. In Czech, some phoneme sequences can be spelled in two ways: /bjE/ as bě or bje, /pjE/ as pě or pje, /vjE/ as vě or vje, /mñE/ as mě or mně (the spelling depends on the origin of the word, for example, bje,pje,vje are used across a morpheme boundary). We distinguish formal errors in the spelling of the phonemes /jE/ (rozbjehl →rozběhl ‘started to run’: formJe1) and errors in the spelling of /mE/ (vzpoměla →vzpomněla ‘remembered’ formMne0,mněla →měla ‘had’ formMne1). A similar phenomenon applies to the orthographic notation of phonetic groups with palatal plosives /c/ (ť), /é/ (ď) and palatal nasale /ñ/ (ň). These phonemes are spelled with a diacritic mark caron (wedge, hacek), but in a combination with the vowel e, they are spelled as tě,dě,ně and in a combination with the vowel i,í as ti,di,ni, and tí,dí,ní, respectively. The spelling with a caron on the consonant (kuchyňe →kuchyně ‘kitchen’) is wrong and is labeled with the error tag formDtn. Another error concerns the marking of length of the vowel u. In Czech, two variants with the same pronunciation are used: ú(uwith an acute diacritic) and ů(uwith a ring diacritic). Simplifying somewhat, úis used word and morpheme initially, and ůotherwise. Mistakes of this nature (dúm →dům ‘house’, ůkol → úkol ‘task’) are labeled with formDiaU. Non-native speakers sometimes make mistakes in spelling of the vowel ifollowed by a syllable boundary and another vowel. In original Czech words, the vowels are always separated by j, both in spelling and pronunciation. Borrowed words are often spelled without j, even though the glide is present in a correct pronunciation j(piano ‘piano’ is orthoepically pronounced as pijáno). Non-native speakers sometimes confuse this rule: we label with formEpentJ0 an incorrect omission of j(přiela →přijela ‘arrived’, žiou →žijou ‘live.3pl’, pieme →pijeme ‘drink.1pl’), and with formEpentJ1 a superfluous j(fotografijemi →fotografiemi ‘photos.inst’, studijum →studium ‘study’, dijamant →diamant ‘diamant’).
100 CHAPTER 5. ERROR ANNOTATION 5.4.5.4 Formal errors sometimes influencing pronunciation Several automatically identified types of formal errors are at the boundary between orthographic errors and errors affecting pronunciation. In some contexts, the error may affect pronunciation, in others it is purely orthographic, but we did not see any benefit in introducing another group of errors. The confusion of the homophonous letters i↔y, which are both used to spell the vowel [I], is labeled with formY0 when replacing ywith an incorrect i(kdiž →když ‘when’) or ýwith í(svími →svými ‘self’s’), or with formY1 when iis replaced with an incorrect y(hystorek →historek ‘stories.gen’), or íwith ý(ostatným →ostatním ‘others’). This error affects pronunciation in combination with the characters t,d and n, in the other cases it is a purely orthographic error. In the words of Czech origin, ynever follows č,ř,š,ž,j, and inever follows h,ch,k,r. After some characters (b,f,l,m,p,s,v,z), both iand yare commonly used in Czech; in some cases, the use of the character distinguishes the meaning of homophones (bil ‘beated’ vs. byl ‘was’). The pronunciation of the word-initial character jpreceding certain consonants (s,m,d) is optional (the prescribed pronunciation of jsem ‘am’ is [sEm]). In colloquial use, such words are often written without the initial j, even though it often distinguishes meaning (jsem ‘am’ vs. sem ‘here’). Such words in learners texts are given the error tag formProtJ0. Similarly, words with an incorrectly added initial j(jsi →si ‘refl.dat’) are labeled with formProtJ1. In these cases (significantly predominant in terms of frequency), this error is orthographic, however, the error tag is used for any missing or redundant jbefore any consonant at the beginning of a word. Therefore, it is also used for errors such as jšla →šla ‘went.fem’ where the change of pronunciation is possible. Errors at the border between orthographic errors and errors affecting pronunciation also include errors in voicing. In Czech, consonants in clusters assimilate in voicing, but their spelling preserves voicing based on their morphological and phonological structure (the word fotbalista ‘soccer player’ is pronounced as a [fodbalista] due to voicing assimilation of /t/ with the voiced /b/). Errors caused by “phonetic” spelling are labeled as formVcd0 (replacing a voiced character with its voiceless counterpart, e.g., skouší →zkouší ‘tries’) or formVcd1 (replacing a voiceless character with its voiced counterpart fodbalista →fotbalista ‘soccer player’). Czech voiced obstruents are devoiced word-finally, but they also preserve their voicing in spelling. Incorrect use of voiceless consonants in such cases is labeled with formVcdFin0 (kdyš →když ‘when’). Prepositions ending in a voiceless consonant optionally assimilate with the voiced consonant of the following word (přes hodinu ‘over an hour’ is pronounced as přezhodinu or přeshodinu), but their spelling does not change either.
5.4. THE TWO-TIER ANNOTATION SCHEME 101 The errors such as přez →přes ‘over’ are labeled as formVcdFin1. However, we use the formVcdFin1 tag for any incorrect use of a word-final voiced consonants, even for those that are not related to preposition voicing assimilation, including errors that a native speaker is unlikely to make (svěd →svět ‘world’). The remaining errors caused by an incorrect use of voiced consonants instead of voiceless ones and vice versa, which are not related to consonant cluster voicing assimilation and are not word final, are labeled as formVcd (přehod →přechod ‘crossing’, sůstala → zůstala ‘stayed.fem’). Sometimes, under the influence of the spelling rules in other European languages, jis incorrectly replaced with y(yá →já ‘I’, žiyu →žiju ‘live.1sg’, formYJ0), and kwith c(clientovi →klientovi ‘client.sg.dat’, culturu →kulturu ‘culture.sg.acc’ formCK0). The words would be pronounced incorrectly if we followed the rules of Czech pronunciation, but it is likely that the author pronounces it correctly and just used an incorrect spelling. In Czech, the letter yis always pronounced as the vowel [I], never as the glide [j]. The character cin Czech words (original Czech words and not recent borrowings) is always pronounced as /ţ/ (voiceless alveolar affricate). It is used for /k/ only in recently borrowed words. Another phenomenon that includes both orthographic and pronunciation changes involves double phonemes. In Czech, double consonants are sometimes pronounced as two phones and sometimes as one. Double pronunciation is used especially for vowels separated by a morpheme boundary (poloostrov ‘peninsula’, individuu ‘individual.dat’). Often they are separated by a glottal stop ([P]). Double pronunciation of consonants is also sometimes used to distinguish meaning (racci ‘seaguls’ vs. raci ‘crayfish’). Two identical consonants when one is in a prefix and another in a root are also often pronounced as two (oddálit ‘postpone’, dvojjazyčný ‘bilingual’), but not always (leccos ‘all sorts of things’). Double pronunciation across a root-suffix boundary is rather rare (vyšší ‘taller’, činnost ‘activity’, babiččin ‘grandma’s’). Errors in spelling of double and single letters are labeled with formGemin:formGemin0 is used for characters that should be doubled but are not (povinost →povinnost ‘duty’, polostrov →poloostrov ‘peninsula’), and formGemin1 is used for consonants that are doubled but should not be (sobbota →sobota ‘Saturday.acc’, rukoppis → rukopis ‘manuscript’). The designation of error is slightly misleading as Czech does not have a real gemination. 5.4.5.5 Formal errors influencing pronunciation One of the types of automatically identified formal errors affecting pronunciation (by native speakers) are errors caused by inappropriate writing of diacritics. Czech, has three diacritical marks: caron (wedge, hacek), acute accent and ring. Acute accent
102 CHAPTER 5. ERROR ANNOTATION and ring indicate a vowel is long (sila ‘silo.pl’ vs. síla ‘force’, pul ‘halve.imper’ vs. půl ‘half’). Caron indicates so-called softening of consonants, i.e., shifting the place of articulation of alveolar consonant backwards: to postalveolar as in z/z/→ž/Z/ or to palatal as in d/d/→ď/é/. The pronunciation of the ěcharacter depends on the previous consonant. Errors in vowel quantity are labeled with formQuant: missing diacritics with formQuant0 (libí →líbí ‘likes’), extra diacritics with formQuant1 (výprávěl →vyprávěl ‘narrated’). A missing caron is labeled with formCaron0 (pojd →pojď ‘come.imper’, neco →něco ‘something’), a superfluous caron is labeled with formCaron1 (kteřých →kterých ‘which’; věnkově →venkově ‘country’). The incorrect use of caron instead of acute or vice versa above the character eis labeled with formDiaE (možně →možné ‘possible’, obchodé →obchodě ‘shop’). Palatalization is a historically motivated consonant alternation. Velar and glottal consonants (k,h/H/, ch /x/, and gin borrowed words) change when followed by a morpheme originally containing the vowel yat, typically realized as e,ěor íin modern Czech: k→c/ţ/, h/H/→z,ch /x/→š/S/, g→z. Errors in palatalization occur most often in the declension of nouns whose stem ends in one of the above consonants. For example, the paradigm žena ‘woman’ has the ending -ě in dative and local singular. The noun řeka ‘river’ belonging to this paradigm has the form řece (stem-final kchanges to c, and the spelling of the ending -ě changes to -e in these cases). Missing palatalization is labeled with formPalat0 whether the author uses -ě (řekě →řece) or -e (řeke →řece). Non-native speakers sometimes make also the opposite error: applying palatalization in places where it should not be: koníčkem ‘hobby.inst’ has the ending -em that historically does not contain yat and thus there is no palatalization. All cases of unjustified palatalization are assigned the error tag formPalat1 (koníčcem →koníčkem ‘hobby.inst’, pracovnícem →pracovníkem ‘worker.inst’). The following formal error also has a historical connection. Yers, Proto-Slavic vowels, disappeared in Old Czech, some transforming into -eand some disappearing completely. This is the cause of alternating forms with and without -e- (pátek ‘Friday.nom’ – pátku ‘Friday.gen’). Synchronically, this is manifested as an epenthesised -emaking it easier to pronounce some words that would otherwise contain a consonant cluster both morpheme internally (kra ‘iceberg.nom’ – ker ‘icebergs.gen’ not kr) and across morpheme boundaries (roz +brát – rozebrat ‘take appart’). Forms with a missing -eare labeled with formEpentE0 error (odbereme →odebereme ‘remove.1pl’). However, the error is defined broadly: it applies to any missing -ebetween two consonants, and most occurrences of this error are thus only loosely related to the original -eepenthesis: odpoldne →odpoledne ‘afternoon’, přijla → přijela ‘arrived.fem.sg’, telvizi →televizi ‘TV.acc’. An extra -ebetween two consonants is labeled with formEpentE1. This is used both in cases where -eoccurs in
5.4. THE TWO-TIER ANNOTATION SCHEME 103 other forms of the paradigm (dáreky →dárky ‘presents’ cf. dárek ‘present’, páteku →pátku ‘Friday.gen’, cf. pátek ‘Friday.nom’), and in cases that are due to pronunciation difficulty of consonant clusters (jemenuje →jmenuje ‘is named’, čtvertek → čtvrtek ‘Thursday’). In spoken Czech, in Bohemia and Central Moravia, word-initial ois often preceded with a prothetic v. For example, some speakers pronounce the word okno ‘window’ as vokno and sometimes they even write it in that nonstandard way (but the phenomenon is probably currently declining). In the CzeSL project, we evaluate the written text against the rules of SCz, so we label occurrences of prothetic vas errors with the formProtV1 tag: vobčas →občas ‘sometimes’, vopravdu →opravdu ‘really’. During the evolution of Czech from Proto-Slavic, the original phoneme gchanged into h. However, this process did not occur in most other Slavic languages. This is a cause for another type of error when non-native speakers mostly of Slavic origin confuse the letters gand h. The incorrect use of the letter gin place of his labeled with formGH0:glavní →hlavní ‘main’, mnogo →mnoho ‘many’, gasič →hasič ‘firefighter.inst’. Czech uses gin newly borrowed words. An incorrect replacement of such gwith his labeled with formGH1:ciharetu →cigaretu ‘cigarette.acc’, prohramů →programů ‘programs.gen’, hrafička →grafička ‘graphic artist.fem’. Character metathesis is a relatively common type of error for both non-native and native speakers. We automatically identify two types of metathesis: swapping adjacent characters (sulnce →slunce ‘sun’, dobrodružtsví →dobrodružství ‘adventure’), and swapping two characters separated by another character (provůdce → průvodce ‘guide’, ojelů →olejů ‘oil.pl.gen’, zicích →cizích ‘foreign.pl.gen’). These errors are labeled with formMeta error tag. 5.4.5.6 Other types of errors The variability of errors in the texts of non-native speakers is too great, so it is not possible to systematically handle all cases. In this section we focus on automatically assigned tags that cannot be classified into any of the above categories. The tags attempt to provide at least some information about the nature of the difference in the original and corrected word. They almost always affect pronunciation. There are three error tags for labeling single-character mistakes that cannot be classified with any of the more descriptive tags above. The formSingCh tag indicates cases where one character is replaced by another: specifiské →specifické ‘specific.neut’, existije →existuje ‘exists’, ofjevit →objevit ‘appear.inf’. The formMissChar tag is assigned to cases where a single character is missing: učiteka →učitelka ‘teacher.fem’, výjmečnému →výjimečnému ‘exceptional.masc.dat’, zbudil
104 CHAPTER 5. ERROR ANNOTATION →vzbudil ‘woke up’. The formRedunChar tag is used in cases with an extra character: usmrdcení →usmrcení ‘killing’, privního →prvního ‘first’, kugličky →kuličky ‘marbles’. Another error tag marks mistakes due to a missing or extra prefix. Therefore, the error tag expresses that there are one or several characters missing or extra word-initially and that the characters are equal to one of the commonly used Czech prefixes. Errors where the prefix is missing in the original word are labeled with formPre0:hledu →pohledu ‘view.gen’, žaduje →vyžaduje ‘requires’, znamil → seznámil ‘introduced.masc.sg’. Words with extra prefixes in the original are labeled with formPre1:pojet →jet ‘drive.inf’, přezačít →začít ‘start.inf’, potrávíme → trávíme ‘spend.1pl’. In some cases, the tag is also assigned to incorrectly fused words that were manually annotated in a wrong way: the incorrectly fused word semnou, corrected to se mnou ‘with me’, should be labeled with the error tag wbdPreJoined and linked with each of the T1 words se and mnou. But because it was only linked to the word mnou, the automatic annotation incorrectly assigns the error tag formPre1. Word-initial errors where the difference cannot be classified as a common Czech prefix are labeled with formHead tags. The tag formHead0 is used for missing word-initial characters (busovou →autobusovou ‘bus.adj’), formHead is used for different word beginnings (prověděl →dozvěděl ‘learned’, chiny →Číny ‘China.gen’) and formHead1 for extra characters. However, the last situation is always a result of errors in manual annotation: incorrectly fused words were properly corrected (conejlíp →co nejlíp ‘as good as possible’) but instead of splitting the original T0 word into two (or more) T1 words, linked to the source and labeled together with the wbdOtherJoined tag, one of the T1 words was labeled as a correction of the original and the other word was inserted as a missing word. The opposite case, when the original and corrected words start in the same way but end differently, is labeled with formTail tags (formTail0,formTail1, formTail). The formTail0 tag is used when the original word is missing some characters at the end (t→tam ‘there’, ž→žít ‘live’; there are only few meaningful examples in the corpus), formTail is used for words with different ends (několiku → několika ‘several.gen’, šansu →šanci ‘chance.acc’), formTail1 for extra characters (no meaningful examples). Even less information is contained in the formLen tags, which are used for words that (1) differ significantly (but that still do not cross the threshold when we give up on marking differences), and (2) that they also differ in length. The tag formLen0 is used when the original word is shorter than the corrected word (ňákem →nějakém ‘some’, diš →když ‘when’), and formLen1 when the original word is longer (vidňanami →Vídeňany ‘Viennese’, recat →říct ‘say’).
5.5. IMPLICIT ERROR ANNOTATION 105 Cases when the original T0 word significantly differs from the corrected T1 word are labeled with the formUnspec error tag: omevy →umývá ‘washes’, choubů →hub ‘mushrooms.gen’, ěšče →ještě ‘still’. In retrospect, we should not have introduced the tags formLen,formTail and formHead – for words where partial differences cannot be easily automatically identified, it would be better to give up recognition completely and always use the formUnspec tag. 5.4.5.7 Automatic classification of word-boundary errors Word-boundary errors include words either incorrectly fused (semsi →sem si ‘aux.1sg refl.dat’) or split (ne chodila →nechodila ‘wasn’t going’). During correction, these errors are manually labeled with wbd tags: wbdPre is used for prepositions fused with the following words and separated prefixes, wbdOther is used for other word-boundary errors. The automatic procedure adds a tag to mark whether the forms were incorrectly separated (wbdPreSplit /wbdOtherSplit) or incorrectly joined (wbdPreJoined /wbdOtherJoined). 5.5 Implicit error annotation Either of the two components of error annotation (classification and correction) may be omitted. The decision to refrain from assigning error tags or from providing correct forms speeds up manual error annotation. However, such a decision can also be made due to theoretical reasons. Some authors intentionally avoid categorizing errors, advocating correction as sufficient error annotation. They see categorization as an interpretation model, influencing access to the data, while correction is viewed as an implicit explanation for the errors (Fitzpatrick and Seegmiller 2004; Mendes et al. 2016). Notwithstanding this theoretical argument, if correction is the only approach to error annotation, its advantage is the easier task of the annotator due to the absence of an error classification scheme (Fitzpatrick and Seegmiller 2001). The annotator does not need to learn any classification rules, which speeds up the annotation task and avoids misclassification. On the other hand, corrections without error labelling may not be sufficient to describe the error properly or substantiate the correction. The resulting annotation could then be too vague for specific queries or analysis by quantitative or statistical methods. As a compromise, corrections could be specified for specific annotation tiers, resulting in an implicit error classification, with an option to derive error tags automatically (Rio and Mendes 2019).
112 CHAPTER 5. ERROR ANNOTATION The form dovoleny ‘holiday’, where only the ending -y differs from the appropriate form dovolené, is apparently formed from a correct word base and an incorrect ending -y, which is a correct genitive singular ending for a different feminine paradigm, or a misspelled or mispronounced variant of the otherwise correct CCz genitive singular form dovolený. The error in dovoleny →dovolené is therefore undoubtedly an error in morphology, phonology or spelling, as the author of the text apparently fails to use the correct ending for the word she uses, but the case seems to be correct. The form šla ‘went.perf’, used instead of the correct form chodila ‘used to go.impf’, is a correctly formed Czech word. However, its perfective aspect is inappropriate in the context of the expression každý den ‘every day’. It is replaced by the imperfective form and classified as a lexical error.21 The omission of the diacritic on the vowel ain the word kamaradku →kamarádkou ‘friend’ can be classified as an error in spelling or phonology (non-native speakers of Czech often do not distinguish between short and long vowels). The second error in the word kamaradku ‘friend’, i. e., the inappropriate use of the ending -u (which is correct for accusative singular) instead of -ou (the ending for instrumental singular, required here after the preposition s‘with’), is either an error in morphology (the author of the text does not know what ending to choose to form the instrumental case), or an error in syntax (the author does not know which case should be used with the preposition s‘with’). The last incorrect form pláži ‘beach’, used instead of pláž, is a correct form of locative singular, a form that can be used after the preposition na ‘on’ (so that both na pláži and na pláž can be correct, depending on the context), but it is inappropriate with a verb of movement, such as jít/chodit ‘go’, so the error can be interpreted as an error in syntax. However, the ending -i is used in other feminine paradigms to form the accusative case, so we cannot exclude the possibility of an error in morphology (the appropriate accusative case is formed incorrectly). 5.6.5 Source text, target hypothesis, annotated strings • Error tags are assigned to those parts of the source text which are different from the single target hypothesis, corresponding to T2 in the 2T scheme. • A tag can annotate a part of a word, a whole word, or even multiple words. More than one tag can annotate a single text string. A tag may annotate a string which includes a shorter string annotated by a different tag, i. e., a tag 21The two forms are actually forms of two different verbs, because aspect in Czech is a lexical rather than morphological or syntactic category.
5.6. MULTI-DIMENSIONAL ERROR ANNOTATION (MD) 113 can be embedded in another tag. The spans of characters or words annotated by different tags may overlap. • As a rule, the shortest possible strings are annotated, a sequence of incorrect characters, sometimes just one character. Only lexical errors are annotated on the full word form. 5.6.6 Domains and features The MD scheme is based on five general categories of errors – domains – and a number of subcategories (4–13) – features – for each of the domains. For the full list of domains and features with examples see Tables 5.10–5.12 on pages 115–117.22 • Each error is assigned to at least one domain and to one feature appropriate for the domain. • Features are unique across the domains. Domains and detailed categories (domain-feature pairs) are thus identifiable by the feature tag. • Multiple domains assigned to an error are interpreted as alternative explanations of the error. The domain of orthography covers errors caused by ignoring the conventions of Czech writing, such as capitalization (praha →Praha ‘Prague’), conventions of transcription of some combinations of phonemes, e. g., ěrepresenting the phonemes jand ein vjec →věc ‘thing’, the use of diacritics (ďeti →děti ‘children’) etc. Many of such errors are fairly common even among native speakers. The domain of morphonology includes errors in phonology, e. g., the transcription of voiced and voiceless consonants (sůstala →zůstala ‘stayed’), or the distinction between the consonants rand l(sometimes ignored by native speakers of Chinese or Japanese, e. g., na kluku ‘on the boy’ vs. na krku ‘on the neck’, and incorrect forms of morphemes unrelated to inflection, e.g., učiteka →učitelka ‘teacher’. As errors in morphology we classify only errors related to nominal declension and verbal conjugation, including both non-words (na Erasmuse →Erasmu ‘on the Erasmus’; studovám →studuju ‘I study’) and existing forms of the given word, inappropriate in the given context (in this case, the error can be either morphological or syntactic). 22More details can be found in the (Czech) annotation manual (Škodová et al. 2019).
114 CHAPTER 5. ERROR ANNOTATION The domain of syntax covers errors caused by the incorrect use of word forms and function words (including prepositions) in a given context. Typically, this is where errors in valency, agreement, quantification and word order belong. Errors in the lexical domain concern cases when the original word is replaced in the correction by a different word with a different meaning and it is not the result of a random morphonological error. If necessary, two or more error domains can be used for the classification of any error. The alternative explanation of a single error by parallel annotation in multiple domains results in some regularities (see Table 5.3 on page 118) or frequent cooccurrence (see Figure 5.13 on page 118) of some error tags. In addition to some linguistic interest, relations of implication and predictable coincidence of some domains and features are used to alleviate the task of manual annotation. In addition to a partial segmentation into morphs and the assignment of automatically identifiable features, usually corresponding to the formal error tags with counterparts in the ORT and MPHON domains, preceding the manual annotation, some annotation is added in a post-processing step, based on the such relations. For more details about the annotation process concerning the MD scheme see 7.6 on page 140. 23except for voiced↔unvoiced
5.6. MULTI-DIMENSIONAL ERROR ANNOTATION (MD) 115 Feature Gloss Examples ORT Spelling only, not pronounced, except for some GEM errors IY i↔y analizovat→analyzovat, odpovýdají→odpovídají, mišlenka→myšlenka, babyčka→babyčka, viděl víli→víly ME mě↔mně etc. dítie→dítě, njekdo→někdo, tjišeji→tišeji, mněsíc→měsíc, jědí→jedí, pjet→pět rohlíků, obět→objet náměstí, konie→koně,ďítě→dítě, díťe→dítě, pro mně→mě, ťelo→tělo AT ú↔ů pújdu→půjdu, ůzký→úzký, domú→domů DIA diacritics (other) obˇjet→objet, řikat→říkat, unava→únava, můsel→musel GEM gemination denník→deník, pana→panna, rozlobit→rozzlobit, odálit→oddálit SUBST substitution (other) yako→jako, dal to Mariji→Marii CAP capitalization praha→Praha, Maminka→maminka PUN punctuation máma,a táta jsou doma →máma,a táta jsou doma;přišla aby se rozloučila. →přišla,aby se rozloučila SEG word boundary ne jsem→nejsem,byses→by ses,smaminkou→s maminkou MPHON Morphonology – errors altering non-native pronunciation VOC vocalization z→ze školy, v→ve škole, se→sMarií ASIM voiced↔unvoiced assimilating gdyž→když, noz→nos;ktyš→když, hrat→hrad, bes→beztebe, f→vkruhu, naschledanou→na shledanou NASIM voiced↔unvoiced non-assimilating grad→hrad, uglí→uhlí, výhodní→východní, roglík→rohlík, bod židlí→pod židlí;uchlí→uhlí, chlídám→hlídám, rochlík→rohlík, sjišťovat→zjišťovat SIB sibilants, affricates23 vajes→vajec, noc→nos, mucím→musím PAL palatalization g/k/ch, l/r ledničke→ledničce, článkech→článcích, ruke→ruce, páre→páře, Prahe→Praze, soši→sochy SOFT softening ě,d,t,n,r,s,c,z telo→tělo, tělefon→telefon, něbe→nebe, veda→věda QUANT vowel length Práha→Praha, učítelka→učitelka, počitač→počítač MET metathesis r/l/m/n pernamentka→permanentka, lefrektor→reflektor, verlyba→velryba, žlička→lžička, lorák→rolák EPENT epenthesis (including both iand e) volb→voleb, pesa→psa, ptáčeka→ptáčka, ližička→lžička PROT prosthetic vvokno→okno, vobjednat→objednat, von→on, vošklivej→ošklivej CNTR contraction děcký→dětský, bohactví→bohatství, czeský→český, morže→moře, bicze→biče, shok→šok ALT other alternations ve vůze→voze, koup→kup to CHAR additional or missing sounds večře→večeře, nesem→nejsem, cera→dcera, sedum→sedm Table 5.10: An overview of domains and features in the MD scheme, part 1 of 3
116 CHAPTER 5. ERROR ANNOTATION Feature Gloss Examples MORPH Morphology – errors due a wrong choice of affixes and stems NAFF affix incompatible with the stem spám→spím, neplavujeme→neplaveme, přečet→přečetl, Číné→Číny, Úvalách→Úvalech, rodičema→rodiči, dědečko→dědečka, Erasmusu→Erasmu FLEX inappropriately used affix in a paradigm uvidím dědečkovi→dědečka VBX compound verb forms zítra jsem→budu spát, jsem spát →spím,budu napsat → napíšu, musíš přijdeš→přijít,učit→učil ses RFL reflexives raduje si→se, má ráda její→své děti, směju →směju se PREP preposition bydlím na→vPraze, vystup v→na konečné, půjdu vles → půjdu do lesa SYN Syntax AGR agreement můj tatínek je už stará→starý DEP dependents pozdravuj Honza→Honzu, bojím se jí zavolám→zavolat SUBJ subject – missing or redundant pronoun dopoledne já čtu a já odpočívám →dopoledne čtu a odpočívám;já ne, ale to uděláš →já ne, ale ty to uděláš COMPL complement – missing (pronominal) object potřebuju tužku, maminka koupí →maminka ji koupí; přivedla muže, Jana neznala →kterého Jana neznala COP missing be, esp. copula země moc velká →země je moc velká;teď v Číně →teď je v Číně CONST other constituent přeju mu cestu →přeju mu šťastnou cestu CONJ connecting expression, including relative pronoun mám ráda, že→když prší;mám hlad, protože→proto se najím;mám rád Prahu, proto→protože je přátelská;chtěl, kdyby→abych přišel;Petr ale→aLucie se mají rádi; doporučuju každému, který→kdo má zájem WO word order mají hezký velmi dům →mají velmi hezký dům Table 5.11: An overview of domains and features in the MD scheme, part 2 of 3
5.6. MULTI-DIMENSIONAL ERROR ANNOTATION (MD) 117 Feature Gloss Examples LEX Lexicon CHOICE wrong lexeme jeli pěšky →šli;nudím se po domově →stýská se mi po domově ASP aspect celý den chytili→chytali ryby;denně vstanu→vstávám brzo MOD modality – verb, adverb, particle v pondělí může→musí jít do práce;hodně myslím, že byl hladový →byl určitě hladový NEG negation půjdu neráno →nepůjdu ráno;on ne→není velký;půjdu ne do školy →nepůjdu do školy;mám→nemám žádný čas; máma ani táta kouří→nekouří COIN coinage – innovative word formation slichtovaní názory (?); šťopinky špičurkatý (?); je to smíchovní→legrační FGN foreign or macaronic jdu do shopu;byla v hangu;to byl shock;hledám kleenexy USE suboptimal choice of (variant) forms, lexemes, collocations, independent categories říkám moje→svoje názory;dělám studium →studuji; ráno přišla rýma →ráno se mi spustila rýma;moucha chodila→lezla po stole;vidí jeho→ho v zrcadle;dívá se na ho→něj;na jaru→jaře všechno kvete;čte už dlouze→dlouho POS part of speech je to hezky→hezký muž;učím se český→česky/češtinu; moc rád pomoc→pomůže;jsem český→Čech;je to dobře→dobrá lekce;máma jméno Dana →máma se jmenuje Dana PHR construction Petr má rád lyžovat →Petr rád lyžuje;mám 17 let →je mi 17 let;já líbím Prahu →líbí se mi Praha;večer dostanu bolest hlavy →večer mne začne bolet hlava XDOM Cross-domain REG register koláč je dobrej→dobrý;přijdu s rodičema→rodiči;to je ale maglajz→zmatek SEC secondary (follow-up, subsequent) error jde na menzu→jde do menzy;Saná má úžasnýklimat → Saná má úžasnéklima PROBL problem Table 5.12: An overview of domains and features in the MD scheme, part 3 of 3
118 CHAPTER 5. ERROR ANNOTATION •MPHON:QUANT =⇒ORT:DIA An issue in the quantity of a vowel is also an issue in diacritics, i. e., missing or redundant acute accent or “ring”. •MORPH:NAFF =⇒no annotation in the SYN domain An incompatible affix results in a non-word, which excludes a syntactic error. •MORPH:FLEX ⇐⇒ SYN:AGR or SYN:DEP An inappropriate ending in a form, still correct within a paradigm, is always due to an issue in syntax, either in agreement or in case assignment or other requirements of a syntactic head. • if MORPH:NAFF or … (i. e., incompatible affix) if (SYN:DEP or SYN:AGR) and … (i. e., syntactic error) if no other MPHON or … (e. g., SOFT/QUANT/NASIM) if no ORT:IY/MNE/U/GEM/CAP (i. e., if it’s not spelling only) =⇒ MPHON:ALT or … (when replacing characters) MPHON:CHAR (when deleting or adding characters) Figure 5.3: Relations of implication and equivalence between features across error domains ORT MPHON MORPH SYN Example IY SOFT FLEX AGR všichniděti →všechnyděti IY FLEX AGR byli→byly DIA QUANT FLEX DEP dětí→děti,přátele→přátelé ALT FLEX AGR druhém→ruhým,jeden →jedno ALT FLEX DEP tradici→tradice CHAR FLEX AGR byl →bylo CHAR FLEX DEP noh →nohy,noc →noci Table 5.13: Co-occurrence of features across error domains
Chapter 6 Linguistic annotation Standard corpora documenting contemporary written native language are commonly annotated by POS tags, morphological categories, lemmas, sometimes syntactic structure and functions, or even by information about named entities or word senses. For many languages, tools and training data are available to perform these task with an error rate sufficient for many purposes. The result is a corpus more useful in a number of ways. A linguistically annotated corpus can be searched more efficiently. It may even be impossible to make some queries without such annotation. The same applies to statistical analyses. Also, linguistically annotated corpora are vital in the development of NLP tools, especially those based on machine learning, which require extensive training and testing. Learner corpora are no exception: together with error annotation, linguistic annotation helps to make them more useful. In fact, the error tagset used in the 2T scheme assumes that the texts are annotated at least by POS tags (see §5.4.3.3). However, linguistic annotation of learner corpora is not a straightforward task. This is due to several reasons: 1. Available tools are trained on standard language, mainly because it is difficult to obtain sufficiently large training data, comparable with the texts to be annotated in terms of text types and specifics of the learner language. Therefore, annotating learner texts by tools intended for standard language means that the reliability of automatic annotation could be lower than reported for native texts. The drop in success rate depends mainly on how far the texts diverge from the standard language. Even within a single learner corpus, texts authored by learners at different levels of proficiency and L1s can be annotated with various success. 119
120 CHAPTER 6. LINGUISTIC ANNOTATION 2. In addition to the higher error rate, adopting the standard language approach to learner language arouses a conceptual concern: categories and structures suited to the standard language might not suit learner language. 3. To avoid such problems, we can annotate the target hypothesis instead. This assumes that the source texts is reconstructed completely, to a grammatically correct version, including the correction of follow-up errors. However, some properties of the learner language may be lost in the annotation. As a possible solution, both the source text and the target hypothesis can be linguistically annotated. This chapter has two parts: 1. The first part (§6.1) focuses at automatic linguistic annotation performed with tools for Standard Czech. This is straightforward for target hypothesis. Exploiting the fact that the words in the 2T scheme are interlinked, the result is projected to the partial target hypothesis (T1). Applying existing tools on the source text is theoretically less sound, but for practical purposes, the results are still useful. 2. The second part (§6.2) describes manual syntactic annotation of a portion of CzeSL using the Universal Dependencies annotation scheme.1We argue that the more abstract syntactic categories are a more intuitive and less arbitrary alternative to morphological annotation of the source learner text. The annotated corpus is too small to train machine learning tools on it in the usual way, but it can be used for benchmarking such tools. 6.1 Annotation with tools for Standard Czech In order to make it easier for users to work with CzeSL, i.e., to enable a comfortable search for words according to base forms and grammatical categories, and to produce statistics based on linguistic categories, the words in the corpus were lemmatized and annotated with morphosyntactic tags. 6.1.1 Annotation of target hypothesis Because the target hypothesis at T2 is a native-like Czech sentence, we could apply a standard lemmatizer and tagger to assign all T2 words non-ambiguous annotation. We use the “Prague” positional tagset (Hajič 2004) in the version modified for the 1https://universaldependencies.org
6.1. ANNOTATION WITH TOOLS FOR STANDARD CZECH 121 Czech National Corpus. Each of the 16 positions corresponds to a morphosyntactic category, e.g., the first position stands for POS, the fifth position for case.2 Following lexical look-up, the disambiguation step proceeds in two stages: first a rule-based system removes most of the ambiguity, then a stochastic tagger resolves the remaining cases. For more details see Hnátková, Petkevič, and Skoumalová (2011). Moreover, the target hypothesis of CzeSL-man v1 searchable has a syntactic annotation according to the PDT standards, parsed automatically with TurboParser (Martins, Almeida, and Smith 2013). 6.1.2 Annotation of T1 The words in the 2T scheme are interlinked. We use this information to project lemmas and tags from T2 to T1, the intermediary target hypothesis, in the following way: 1. If the T1 word is identical to its T2 counterpart, it gets its lemma and tag. 2. Otherwise: a) If the T2 lemma is one of the possible T1 lemmas, we use that T2 lemma and the set of all tags associated with it which are consistent with the T1 form. For example, the homonym jí is either a form of the verb jíst ‘eat’ or dative singular of the personal pronoun ona ‘she’. Let us assume the ambiguous form on T1 was corrected as jedí ‘eat.3PL’ on T2. Because jedí is non-ambiguously a form of the verb jíst ‘eat’, jí on T1 is considered only as the form of this verb and it is assigned tags for the 3rd person plural and Common Czech 3rd person singular. b) Otherwise: T1 gets all possible lemmas and all possible tags. 6.1.3 Annotation of source texts In several releases of CzeSL (CzeSL-SGT,CzeSL-man v1,CzeSL-man v2,CzeSL in TEITOK), automatic lemmatization and morphosyntactic tagging is available for T0, the source version of the text. For this task, we used MorphoDiTa (Straková, Straka, and Hajič 2014), trained on standard native Czech data of the Prague Dependency Treebank (PDT, Hajič et al. 2018), rather than the hybrid tagger applied to native Czech texts of the Czech National Corpus and used also for several TH 2See https://wiki.korpus.cz/doku.php/en:pojmy:tag for a description of the tagset.
128 CHAPTER 6. LINGUISTIC ANNOTATION In that case, we try to be as conservative as possible and assume as little as possible: we use the form of the word as its lemma and mark it as unclear in the note field. The alternative is to use the correct lemma (Praha in (25) and večeře in (27)). Obviously, this would make the situation clearer and the annotation more reliable. However, the benefit would be minimal: error annotation already provides us with the correct forms so we can easily derive their lemmata using available approaches for standard native language. 6.2.4 Syntactic Structure In annotating syntactic structure, we again follow the rule of annotating the structure of interlanguage. For example, if the learner uses the phrase (29), the word místnost ‘room’ is annotated as a direct object (OBJ), even though a native speaker would use an adverbial (OBL)do místnosti ‘into room’ as in T2. (29) T0: vstoupit enter *místnost.OBJ room T2: vstoupit enter do into místnosti.OBL room ‘enter a/the room’ Examples (30) and (31) illustrate the difficulties we encountered during the annotation. Each example is followed by the corresponding sentence in standard native Czech. Missing že ‘that’ (30) T0: Myslím, think.1sg velmi very málokdo few-one dělá, does co what chce. wants T2: Myslím, think.1sg že that velmi very málokdo few-one dělá, does co what chce. wants ‘I think hardly anybody is doing what they want.’ (AA_IK_001 hu B1) •Annotation with interpretation. In the corresponding grammatically correct Czech sentence in T2, the connector že ‘that’ follows the verb myslím ‘I think’. This makes velmi málokdo dělá ‘hardly anybody is doing’ a subordinate complement (object) clause, and thus the verb dělá ‘wants’ would be annotated as ccomp.
6.2. ANNOTATION OF INTERLANGUAGE IN UD 129 •Annotation without interpretation. Without interpretation, we consider the clause velmi málokdo dělá, co chce to be coordinated with the previous Myslím, connected to it with conj. There is another possibility: the author uses a structure parallel to English I think very few people know … without that. Then the form dělá ‘wants’ would be annotated as ccomp as well. Using abych instead of že (31) T0: Rozhodla decided.f.sg jsem aux.1sg se, refl *abych conj+aux.1sg se refl naučila learned.f.sg nějaký some zajímavý interesting jazyk, language který which je is blízko close nás we.dat … … T2: Rozhodla decided.f.sg jsem aux.1sg se, refl že comp se refl naučím learn.1sg nějaký some zajímavý interesting jazyk, language který which je is blízko close nás we.dat … … ‘I decided to learn some interesting language that is close to us …’ (AA_IK_001 hu B1) •Annotation with interpretation. The verb rozhodla jsem se ‘I decided’ in the first clause requires a complement clause connected via the complementizer že ‘that’: že se naučím … jazyk ‘that I learn … language’, instead of an adverbial clause abych se naučila … jazyk ‘in order to learn … language’ Therefore the predicate naučím ‘learn’ of the subordinate clause would be annotated as ccomp. •Annotation without interpretation. The subordinate clause is considered a complement clause as well but marked with the conjunction aby ‘so-that’. This is parallel to (32). In this case, UD helps by not forcing us to make spurious distinctions. (32) Native Czech: Požádal asked.m.sg jsem aux.1sg ji, her aby comp+aux.3 se refl naučila.CCOMP learn.f.sg nějaký some jazyk. language ‘I asked her to learn some language.’
130 CHAPTER 6. LINGUISTIC ANNOTATION 6.2.5 Evaluation The manual annotation of CzeSL-UD was done by two annotators: an annotator with a philological background and a secondary-school student. They did not undergo any special training prior to the annotation, but instead relied on a secondary-school grammar training and the guidelines for Czech available at the UD project site.6When they were not sure about a particular construction, they referred to existing Czech and English UD corpora, compiling a shared guide and a cheat sheet7in the process. Technically, we used the TrEd editor with the ud extension to do the annotation.8The annotators corrected a default structure obtained by running UDPipe (Straka and Straková 2017) on target hypothesis (T2) text and projecting the output to learner text (T0). For a pilot annotation, we have randomly selected 100 sentences from CzeSLman shorter than 15 tokens. We measured their IAA using Cohen’s kappa (see §5.4.4.1) on part-of-speech labels, syntactic labels and unlabeled heads. The IAA scores 0.934, 0.89, 0.927, respectively, are good but not perfect. However, we believe that the most important result of the pilot UD annotation is not the actual annotation, but the guidelines that can be used as a basis for other non-native languages. The annotation is still a work in progress. Our goal is to eventually annotate the whole CzeSL-man corpus. So far more than 1600 sentences have been annotated. 6https://universaldependencies.org/guidelines.html 7https://bit.ly/UDCheat 8https://ufal.mff.cuni.cz/tred
Chapter 7 Annotation process In this chapter, we discuss technical, procedural and practical aspects of the compilation and annotation of the CzeSL corpus. For a discussion of the annotation scheme, see §5in case of error annotation, and §6in case of linguistic annotation. The supporting tools are covered in §9. 7.1 Overview of the annotation process The whole annotation process proceeds as follows. Most of the texts are processed in batches, and the individual steps are separated in time, space and people responsible for their correct execution. More recently, the texts can be processed as needed, also one by one. Thus the corpus can grow incrementally, with the texts ready for on-line searching soon after they become available as source documents (see §9.2.3). 1. Collection of texts: The original texts and their metadata are collected (see §3.1 and §4.4); handwritten manuscripts are scanned. For most texts, collection and procurement of metadata has been done in cooperation with the teacher or examiner. In some cases, especially in the assessment of the learner’s proficiency, the metadata item is based on the teacher’s estimate rather than on the performance of the learner in the text. This is why the information about the learner’s CEFR level in the corpus is not always completely reliable. 2. Transcription: Scanned manuscripts are manually transcribed and anonymized (see §4.3 and §4.4); each transcription is checked by a supervisor (see §7.2). 131
132 CHAPTER 7. ANNOTATION PROCESS 3. Error annotation and linguistic annotation: • Two-tier error annotation, including linguistic annotation of target hypothesis (2T; see §7.3.1) • Multi-dimensional annotation (MD; see §7.6) • Implicit error annotation (see §7.7) • Syntactic annotation of the source text in the Universal Dependencies framework (UD; see §7.8) Although these types of annotation are theoretically independent, there are some practical dependencies between them. The MD and UD annotation depend on the two-tier annotation: they use tokens derived from T1 and a default annotation based on T2. The implicit annotation is independent but compatible with the other annotation schemes. Implicit annotation can also be based on the annotation in one or more other annotation schemes. The annotation schemes can be integrated and used by a single corpus annotation, maintenance and search tool (see §9.2.3). 7.2 Transcription and anonymization of manuscripts To transcribe hand-written texts, at first we used off-the-shelf editors supporting HTML with simple transcription codes. Later, we switched to XML markup and an XML-aware editor. For details about the transcription formats see §4.2. The HTML-based format allows the transcribers to use a tool they are familiar with, which means that not much technical training is required. Some of the codes are supported via macros of the editor. This is how most of the hand-written texts in CzeSL were transcribed. The decision to use HTML produced by an off-the-shelf editor was made intentionally to minimize training time and not to limit the pool of potential transcribers – it is hard enough to find people who know the rules of handwriting of speakers of language X, it is even harder to find experts who are also able to transcribe into XML. However, in retrospect we feel this was not a correct decision, because the efforts needed to review the transcripts clearly outweigh the benefit of using a widely known tool. First, it is really important to minimize the occurrence of errors in transcription as they influence all the subsequent annotation steps. It is easier to enforce formal correctness in an XML editor such as XMLmind than in an HTML editor. Second, the ability to learn to use an XML editor is actually a good indication of other abilities that are important in the transcription process, for example the ability to follow the formal rules of a transcription manual.
7.3. TIERED ERROR ANNOTATION 133 Starting with a new batch of collected manuscripts, all transcripts have been done in the TEITOK tool in the XML-based format. Although the tool does not validate the content of the markup according to an XML schema, the use of transcription and anonymization codes is facilitated by pre-defined keyboard shortcuts and the XML syntax is checked on the fly. The texts previously transcribed and anonymized in the HTML-based format are converted into the XML-based format. 7.3 Tiered error annotation The tiered error annotation proceeds in the following steps: 1. Preprocessing: The transcript is converted into a format where T0 roughly corresponds to the tokenized transcript and T1 is set as equal to T0 by default. Both are encoded in PML, an XML-based format (see §7.3.3). The conversion includes basic checks for incorrect or suspicious transcription. 2. Manual error annotation: Errors in the text are manually corrected and tagged; each annotation is checked by a supervisor (see §7.3.1). Some texts are independently annotated twice (see §5.4.4). 3. Automatic annotation checks: Manually annotated texts pass through a series of automatic checks. Suspicious annotations are marked and manually reviewed. 4. Manual adjudication: Each doubly annotated text should be checked and adjudicated, resulting in a single annotated version. However, except for a small pilot, this has not been done yet, so a part of the corpus actually contains two independent annotations. 5. Linguistic (morphological) annotation: Target hypothesis is automatically annotated with lemmas and morphological tags, both full hypothesis on T2 (see §6.1.1) and individual words on T1 (see §6.1.2). 6. Automatic error annotation: Error information that can be inferred automatically is added by comparing original and emended words: type of spelling alternation, missing/redundant expressions, and inappropriate word order (see §5.4.5). Conversion to PML (see §7.3.3), annotation, supervision and adjudication are done with the help of feat, an annotation editor designed as a part of the project (see §9.1.1). The storage of the documents and their flow within this process is managed by Speed, a purpose-built text management system (see §9.1.2).
134 CHAPTER 7. ANNOTATION PROCESS 7.3.1 Manual error annotation Some of the transcribed texts are error-annotated manually according to the 2T scheme described in §5.4. The annotation was done in feat (see §9.1.1). The annotator corrected the text on appropriate tiers, modified relations between elements (by default all relations are 1:1) and annotated relations with error tags as needed. Figure 7.1 shows the annotation of a sample sentence as displayed by the tool. The top of the window shows the currently annotated part of the sentence, displaying the source text above the two annotation tiers. The context of the annotated text is shown both as a transcribed HTML document (bottom left of the window) and as a scan of the original document (bottom right). In the annotated part of the window, forms identified by the tool as non-words are underlined, corrections done by the annotators are in red. Unless the annotator decides otherwise, vertical links align the words across the tiers 1:1. When the error type cannot be identified automatically, the annotator is supposed to replace the Xlabel on the link between the incorrect form and its correction by one or more error tags. For some error tags, such as agr or dep, the annotator is instructed to provide a reference link to another word to substantiate the correction. It is usually the agreement source or the syntactic head of the corrected word. Each annotation was reviewed by a supervisor, who could approve it, modify it, or return it to the annotator with comments for revision. A subset of the texts annotated this way was independently annotated twice to assess the reliability of the annotation and the robustness of the tagset and the annotation scheme. After a pilot annotation, we used the result of the comparison to improve the annotation guidelines. See §5.4.4 for a detailed analysis of errors. The annotation guidelines do not make any strict requirement about the sequence of steps in the error annotation, or about the relation of normalization and categorization as the two error annotation tasks. They only make an assumption that the two tasks are done by the same annotator, typically while annotating the whole text in one go. The annotators tend to normalize and categorize errors at the same time anyway. The advantage of this approach is that error tags reflect THs (see §5.1, p. 62 about the relation beween TH and error categorization). On the other hand, separating annotation tasks in time and/or in the person of the annotator can result in better control of the annotator’s judgments and thus more robust annotation. We followed this idea in CzeSL-TH (see §8.4), a part of CzeSL, which is corrected at T1 a T2 according to the 2T scheme, without error categorization. Error tags can be assigned in a separate step at any time later. Given the 2T scheme, the annotators can also choose between annotating whole sentences, paragraphs or texts first on T1 and only then on T2, or annotating each
7.3. TIERED ERROR ANNOTATION 135 Figure 7.1: A sentence displayed in the feat annotation tool (see Table 5.1 on page 74 for the whole text) (NEM_GD_008 ru B2) text in parallel on both tiers. Some annotators prefer to annotate by paragraphs, first annotating the whole paragraph on T1 and then on T2, while others annotate by sentences, annotating a sentence on both tiers in parallel before moving to the next sentence. Despite the proofread status of the 2T annotation, additional checks by a different annotator as a part of the MD and implicit annotation have proved useful. This applies even to the doubly annotated part of CzeSL-man, due to the as yet unrealized plan of its adjudication. Manual categorization in the MD scheme is based on the TH made in the 2T scheme (more precisely, on its T2, see §5.6). The annotator can modify a TH which is clearly not correct. However, annotators are discouraged from substituting a more appropriate TH unless the existing TH is obviously wrong. In the implicit error annotation scheme (see §7.7), the annotators are free to use a TH suggested in 2T (if the text was annotated in 2T) or to use their own TH. 7.3.2 Automatic annotation checking The system designed for automatic error tagging is also used for evaluating the quality of manual annotation, checking the result for tags that are probably missing or incorrect. For example, if a T0 form is not known to the morphological analyzer,
136 CHAPTER 7. ANNOTATION PROCESS it is likely to be an incorrect word which should be corrected. Also, if a word was corrected and the change affects pronunciation, but no error tag was assigned, an incorBase or incorInfl error tag is probably missing. This approach cannot find all problems in error annotation, but provides a good approximate measure of the quality of annotation and draws the annotator’s attention to potential errors. 7.3.3 Data format for the tiered annotation scheme To encode the tiered annotation used in CzeSL-man (see §5.4), we have developed an annotation schema in the Prague Markup Language (PML).1PML is a generic XML-based data format, designed for the representation of rich linguistic annotation organized into tiers. Each of the higher tiers contains information about words on that tier, about the corrected errors and about relations to the tokens on the lower tiers. We had also considered using a TEI format.2However, at least for stand-off layered annotation, the support offered by PML was superior to that of TEI, mainly in the availability of tools and libraries. This concerns tasks such as validation, structural parsing, corpus management and searching. While some of those libraries do exist for TEI, many would have to be developed. More recently, we started using TEITOK §9.2.3 as the annotation and search tool, which explicitly supports some parts of the TEI standard. However, although it allows for stand-off annotation, its core uses the inline annotation format. T0 does not contain any relations, only links to the neighboring T1. In Figure 7.2, we show a portion (first two words and first two relations) of T1 of the sample sentence from Figure 5.2, encoded in the PML data format. To allow for data exchange, the feat editor supports import from several formats, including EXMARaLDA (Schmidt 2009; Schmidt et al. 2011); it also allows export limited to the features supported by the respective format. 7.4 Automatic error tagging After the manual error annotation in the 2T scheme the texts are automatically assigned tags identifying formal errors on T1 (see §5.4.5.2 for details). At the same time, some manually assigned error tags on T2 are automatically refined. The tool (see §9.1.5) compares the source and the T1 forms. Any difference is assigned a formal error tag following rules implemented as an algorithm. On T2, 1See Pajas and Štěpánek (2006) and https://ufal.mff.cuni.cz/pml. 2https://tei-c.org/
7.4. AUTOMATIC ERROR TAGGING 137 <?xml version="1.0" encoding="UTF-8"?> <adata xmlns="http://utkl.cuni.cz/czesl/"> <head> <schema href="adata_schema.xml" /> <references> <reffile id="w" name="wdata" href="r049.w.xml" /> </references> </head> <doc id="a-r049-d1" lowerdoc.rf="w#w-r049-d1"> ... <para id="a-r049-d1p2" lowerpara.rf="w#w-r049-d1p2"> ... <s id="a-r049-d1p2s5"> <w id="a-r049-d1p2w50"> <token>Bál</token> </w> <w id="a-r049-d1p2w51"> <token>jsme</token> </w> ... </s> ... <edge id="a-r049-d1p2e54"> <from>w#w-r049-d1p2w46</from> <to>a-r049-d1p2w50</to> <error> <tag>incorInfl</tag> </error> </edge> <edge id="a-r049-d1p2e55"> <from>w#w-r049-d1p2w47</from> <to>a-r049-d1p2w51</to> </edge> ... </para> ... </doc> </adata> Figure 7.2: A part of T1 of the sample sentence (Figure 5.2 on page 85) encoded in PML (VOB_KA_049 kk A1)
144 CHAPTER 7. ANNOTATION PROCESS identify several dozens of error subtypes (mainly in the domain of orthographic and morphonological errors). Some schemata are very simple (such as labeling errors in uppercase/lowercase letters, missing/inappropriate quantity), others are more complex, dependent on the phonetic environment, stem of the governing word, etc. (palatalization; decision whether to use error mark CHOICE etc.). 7.6.3 Experiments with automatic identification of errors in inflection We experimented also with automatic identification of errors in inflection, which (if reliable enough) would significantly reduce the workload on manual annotators. The experiments had promising results, but we decided not to implement this module before the manual annotation would provide enough data to test it automatically. The following text describes this experiment and its (partial) results. The experiment targets those words in the source text whose corrected form was identified as an inflectional word. Morphemic analysis, described in §7.6.1, was used to split both the source and the corrected word forms into a stem and an inflectional suffix (and sometimes prefix). For example, if the incorrect source form is stromom ‘tree’ and the TH is strom with an empty inflectional suffix, the system does not compare only the two last characters of both words (which are identical by chance), but compares the entire stems and determines that the suffix of the original word is -om (strom|om). Using the stems and inflectional affixes for both the original and the TH forms, a twodimensional comparison of the stems and the affixes is then performed. If the stems (original and TH) differ, two facts are checked: whether there are any minor errors in the stem (orthographic, phonological), and whether the source stem is an allomorph of the stem of the TH form, as in v Prahe →v Praze ‘in Prague’, where the original stem Prah (incompatible with the -e suffix) is used to form other (correct) forms of the same lemma, e.g., Prah|a, Prah|yetc. If the affixes differ, they are also checked for minor changes (orthography, e.g., diacritics). Another check tests whether the incorrect affix is used within the given paradigm for other morphosyntactic properties, or whether the affix is used with other paradigms to express the same morphosyntactic properties. The observed differences correspond roughly to the proposed error classification scheme: all errors in orthography and most of the errors in morphonology can be identified automatically. Incorrect affixes indicate an error in morphology; if the incorrect ending is an existing one, expressing the same morphosyntactic properties, it may be an error only in morphology, otherwise it has to be seen as a possible error in syntax as well. If the original word is correct and has the same morphosyntactic properties as the
7.6. MULTI-DIMENSIONAL ERROR ANNOTATION 145 A1 A2 B1 B2 C1 Total Number of tokens 6,961 42,252 39,987 28,182 5,522 122,904 Percentage of the data 5.66 34.38 32.54 22.93 4.49 100.00 Table 7.1: Data distribution by language proficiency A1 A2 B1 B2 C1 Total Correct 61.32 68.15 77.45 78.01 95.09 74.25 Incorrect ending 9.10 10.68 6.30 6.30 0.97 7.72 Incorrect stem 18.91 14.20 11.18 11.23 3.17 12.31 Incorrect whole 19.77 17.65 11.37 10.76 1.74 13.44 Total 100.00 100.00 100.00 100.00 100.00 100.00 Table 7.2: Proportion of correct and incorrect nouns by proficiency levels TH word, but the lemma is different, the error may belong to the lexical domain (except for function words). The relationship between the automatic classification and the classification into error domains is not straightforward. A manual test on a sample of 500 learner errors shows that the approach is reliable with more than 90% of categories determined correctly. As the system is rule-based, it can be fine-tuned by modifying the rules. We tested the rule-based system on nouns in the CzeSL-man corpus. Ill-formed nouns were identified as such using the disambiguated POS tags for corresponding corrected forms on T2. The texts were divided by language proficiency of the authors in terms of CEFR. The levels are not evenly distributed, as shown in Table 7.1. We performed two analyses of nouns in the CzeSL-man corpus: one more general, determining the proportion of incorrect nouns in the corpus, one detailed, focused only on errors in inflection suffixes of nouns. Table 7.2 shows the proportion of correct nouns, nouns with an incorrect suffix (jeskyne →jeskyně ‘cave’), with an incorrect stem and a correct suffix (Prahe →Praze ‘Prague.dat/loc’), with both stem and suffix incorrect (delki →délky ‘length.gen’), and impossible to split automatically (těmy →tématu ‘topic.gen’). The proportion of correct nouns increases with the proficiency level, but there is little change between B1 and B2. On the other hand, there is an unexpectedly large difference between B2 and C1 in the proportion of correct nouns. The highest proportion of incorrect suffixes is in the A2 level texts. We analyzed in more detail the errors in nominal suffixes: all nouns with either a correct stem, or with minor changes compared with the TH were examined. Two parameters were observed: whether the error in the suffix can be an error in orthog-
146 CHAPTER 7. ANNOTATION PROCESS A1 A2 B1 B2 C1 Total Other paradigm 12.19 14.90 14.16 18.97 16.20 15.23 Other paradigm & spelling 8.87 4.51 7.08 6.91 7.82 6.83 Paradigm 19.38 30.77 22.83 24.03 24.02 25.34 Paradigm & spelling 4.03 3.24 5.25 5.14 11.73 4.40 Spelling 7.65 2.94 3.65 7.00 7.82 4.06 Other 47.88 43.64 47.03 37.94 32.40 44.14 Total 100.00 100.00 100.00 100.00 100.00 100.00 Table 7.3: Proportion of types of errors in endings of nouns raphy, and whether the suffix is an existing Czech inflectional suffix used either to express the same case, number and gender in other paradigms, or is used in the same paradigm to express other morphosyntactic properties. Table 7.3 shows the analysis of errors in nouns in CzeSL. Six subtypes of nouns with incorrect suffix were registered: Other paradigms: a suffix used in other paradigms (Úvalách →Úvalech ‘Úvaly.loc’, a place name); likely syntactically correct Other paradigms & spelling: a suffix used in other paradigms and with a spelling error at the same time (Prázě →Praze ‘Prague.dat/loc’) Paradigm: a suffix used inside the paradigm for other morphosyntactic properties (na procházka.*nom →procházku.acc ‘on/for a walk’); this is probably an error in morphosyntax Paradigm & spelling: as above, with a spelling error at the same time (lidi → lidí ‘people’) Spelling: only a spelling error, none of the above (pracé →práce ‘work’) Other: all other instances We observe a steady decrease of “Other” errors, and an increase in the proportion of orthographic errors with language proficiency levels (the authors with a higher proficiency make less errors in general, but keep omitting diacritics). The system allows also for the analysis of individual suffixes: we observed, for example, that suffixes with high ambiguity such as -e,-i,-í are more prone to errors (already noted by Hudousková 2014, 220).
7.6. MULTI-DIMENSIONAL ERROR ANNOTATION 147 7.6.4 Manual error annotation The manual MD annotation is done in the brat annotation editor (see §9.1.3). The design and principles of the MD annotation scheme are described in §5.6 above.9 Figure 7.4 shows a text in brat, while Figure 7.5 shows the same text with a menu of error tags – labels of domains and features. The text has been preprocessed and manually annotated. As described above, pre-processing involves a partial morphemic analysis and automatic error annotation.10 Figure 7.4: A sample MD annotation in brat (AA_AO_002 pl B1) Each sentence in Figure 7.4 is displayed twice. The pre-processed source version, corresponding to T0, comes first. Inflective words are split into morphs by asterisks and errors are tagged by error labels, corresponding to the feature name. Nearly all errors are detected and partially categorized automatically in the pre-processing 9For the MD annotation manual (in Czech) see Škodová et al. (2019). 10A pilot annotation of 18 texts, based on the annotation manual, can be viewed and searched using brat at https://quest.ms.mff.cuni.cz/brat/czesl.err/index.xhtml#/anna_daniela/, or downloaded as a dataset in the brat format from https://bitbucket.org/czesl/czesl-md/.
148 CHAPTER 7. ANNOTATION PROCESS Figure 7.5: MD annotation in brat with the error tags menu (AA_AO_002 pl B1) step. The annotators proofread and modify the annotation by comparison with the target hypothesis of the sentence, shown below the source on the light grey background. The target hypothesis is adopted from the 2T annotation scheme. It is assumed to be correct, but can be modified when the annotator disagrees. For space reasons, the error domains are distinguished by different colors rather than by more verbose display of the domain tag with a gloss. Error labels assigned in pre-processing are denoted by the letter “a” preceding the feature tag, as in aDIA. Some of the other a-type tags may have been modified by the annotator, other tags were added manually. The span of the error, shown below the error tags, may be identified correctly by the morphemic analysis, but the annotator is free to modify it or annotate a new error with its specific error span. One of the main features of the MD annotation scheme is the option of multiple alternative interpretations of an error even in the case of a single target hypothesis. This might seem as an additional burden for the annotator. However, there are some regular patterns of co-occurring error tags, which are used in the pre-processing and post-processing steps. The patterns follow from the error taxonomy and most of them are easy to remember (cf. Figure 5.3 and Table 5.13 on page 118). For example, a feature tag in the MPHON domain is deducible from an automat-
7.6. MULTI-DIMENSIONAL ERROR ANNOTATION 149 ically assigned tag in the ORT domain. As a result, the annotator need not worry about the annotation in the MPHON domain when an ORT domain tag is in place. The same rule applies also in the opposite direction. On the other hand, tags in the LEX domain, except for CHOICE, and in the SYN domain must always be specified by hand. However, if the SYN domain tag is AGR or DEP, then the FLEX tag in the MORPH domain is always appropriate and need not be specified, and – if either ALT or CHAR are the correct tags in the MPHON domain – they are not needed either. To assist the annotator, the annotation editor is informed about possible combinations of tags and issues a warning whenever the annotator adds an incompatible tag for the same or overlapping span. The annotation spans can be of arbitrary length, from a single character to a sequence of words, where some words need not be completely included in the span. For some error types (defined in the brat annotation setup), even discontinuous sequences of words or characters are allowed. These error types include ORT:SEG, MORPH:VBX,SYN:WO or LEX:PHR. For multiple errors concerning different parts of a single word form, it is useful to specify spans for multiple substrings of a word, i.e., a morph or even a single character. On the other hand, some tags can only be used for entire words. This applies mainly to the LEX and SYN domains, except for SYN:AGR,SYN:DEP,LEX:NEG and LEX:USE. The MD annotation can be searched and viewed also in TEITOK (see §9.2.3). Figure 7.6 shows the same text again, this time in the TEITOK stand-off annotation view. The view shows only the MD annotation. Annotated words are underlined, details of the annotation are shown on mouse-over. In the right-hand column the error codes used in the text are listed at the top. A click on the tag shows all words annotated by that tag. All annotated forms with the spans highlighted are shown in the list of similarly clickable forms below the error tags. In TEITOK, the MD annotation can also be edited. Error tags can be deleted or added, and the span and error tag can be modified. 7.6.5 Post-processing of manually annotated texts The manual annotation in brat is followed by post-processing. This step adds error tags identifying phenomena implied by the manually assigned tags. For example, some morphonological tags entail orthographic errors: any voicing assimilation error, such as hodně chyp →hodně chyb ‘many errors’ is labeled as MPHON:ASIM (in this case, the spelling is incorrectly influenced by final devoicing), but it should be also labeled as ORT:SUBST (incorrect substitution of one letter by another). In cases when one error label implies another, the annotators are instructed to add only the former, the implied error is added by automatic post-processing.
150 CHAPTER 7. ANNOTATION PROCESS Figure 7.6: A sample of MD annotation in TEITOK (AA_AO_002 pl B1) The post-processing algorithm has not been fully implemented yet. It is waiting for a sufficient amount of annotated texts and more feedback from the annotators. As a more distant step, we consider designing an algorithm that would replace the manual annotation of other, possibly all, MD error tags. 7.7 Implicit annotation Manual error annotation can be easier when one of the two parts of the error annotation task is omitted (correction or categorization). In both of our two attempts to simplify error annotation this way we applied error correction, omitting catego-
7.7. IMPLICIT ANNOTATION 151 rization.11 The first approach is based on the previous experience with the 2T error annotation scheme. We used the same toolchain, including feat as the annotation editor, and the same annotation guidelines, including the distinction of T1 and T2 and the geometry of cross-tier links for splitting, joining and reordering tokens. However, no error tags were used. In 2017, 1,300 texts (180 thousand tokens) were manually corrected. The texts were selected from the pool of texts annotated only by automatic tools in CzeSLSGT to partially fill the under-represented groups of learners according to the combined L1 and CEFR specifications. The annotated texts are released as CzeSL-TH (see §8.4) and as a part of the AKCES-GEC dataset (see §8.7). We have also tested and adopted an approach based on a sequence of target hypotheses corresponding to linguistic notions such as spelling, morphology, syntax or lexicon with a radically simplified error categorization part, using TEITOK as the annotation tool.12 The first major application of this type of annotation was in a corpus of native learners of Czech – SKRIPT 2015 (see §8.9). The corpus is based on already existing transcripts. A part of the corpus overlaps with SKRIPT 2012, which was released without linguistic or error annotation. The texts were converted from transcripts using the original transcription markup into XML and, if necessary, manually anonymized. Then the texts were hand-corrected at four levels: (i) rectification of non-standard forms (resulting in a form that is still non-standard but spelled “correctly”), (ii) spelling and morphonology (correcting even “correctly spelled” nonstandard forms), (iii) morphosyntax and (iv) lexicon. Most of the levels were tagged and lemmatized, and annotated with the formal error tags (see §5.4.5). Due to a positive experience with this fairly large-scale manual annotation project, other CzeSL texts without manual annotation included in CzeSL-SGT are due to be annotated in the same way, while the already existing manually annotated parts of CzeSL will be integrated into the result – CzeSL in TEITOK. Importantly, the annotation in CzeSL in TEITOK is compatible with the 2T annotation scheme. Based on experience and options, the following data can be used in a corpus built in TEITOK: New texts (manuscripts or audio) can be transcribed and anonymized in TEITOK in the XML format 11See §5.5 for more about implicit error annotation. 12See §9.2.3 for more about the tool. Several learner and historical corpora are available in TEITOK, with annotation based mainly on corrections.
152 CHAPTER 7. ANNOTATION PROCESS Existing transcripts in the old format, possibly anonymized, can be converted into the TEITOK XML format using a conversion tool (see §9.1.6) 2T error-annotated texts – including texts without error tags, can be converted into the TEITOK XML format; some annotation can be expressed inline, other annotation (more complex cross-tier links) in a stand-off annotation format MD error-annotation can be added to the TEITOK XML in the stand-off format Once the data are included in a TEITOK corpus, they can be annotated in the following ways: Error annotation automatic: TH guessing (Korektor web service13), formal error tags manual: successive corrections, implicitly specifying the error type (corresponding to 2T tiers and some 2T error tags or to MD error domains) Linguistic annotation automatic: lemmas and tags for the source and/or any correction level (MorphoDiTa web service14), syntactic structure (UDPipe web service15) manual: checking and editing 7.8 Universal Dependencies A syntactically annotated learner corpus, CzeSL-UD, was built according to the framework of Universal Dependencies as described in §6.2. The annotation proceeded in three steps: 1. Preprocessing: An automatic parse of the target hypothesis (T2) is projected to the source text (T0) as the default syntactic structure. 2. Manual annotation: The default syntactic structure is corrected as necessary by annotators using TrEd. 13https://lindat.mff.cuni.cz/services/korektor/api-reference.php 14https://lindat.mff.cuni.cz/services/morphodita/api-reference.php 15https://lindat.mff.cuni.cz/services/udpipe/api-reference.php
7.8. UNIVERSAL DEPENDENCIES 153 3. Adjudication: A double-annotated subset of data had differences resolved. The manual annotation itself was done by two annotators: an annotator with a philological background and a secondary-school student. They did not undergo any special training prior to the annotation, but instead relied on a secondary-school grammar training and the guidelines for Czech available at the UD project site.16 When they were not sure with a particular construction, they referred to existing Czech and English UD corpora, compiling a shared guide and a cheat sheet17 in the process. Technically, we used the TrEd editor with the ud extension to do the annotation.18 As mentioned above, the annotation was not done from scratch, but the annotators corrected a default structure obtained by running UDPipe on target hypothesis (T2) text and projecting the output to learner text (T0). Obviously, using a default structure provides a certain bias, but we thought the bias to be minimal and the amount of manual work saved was quite large, so we decided it is worth the cost. Ideally, we would run a pilot study comparing the annotations resulting from annotations done from scratch and annotations based on correcting a default structure, but unfortunately this was not practically feasible. However, we did perform a pilot annotation to validate a general reliability of the annotation. We double annotated a sample of the sentences and compared the results. We used the analysis of the results to improve the guidelines. The pilot also showed that the differences between independent annotations were relatively small. See §6.2.5 for more details. 16https://universaldependencies.org/guidelines.html 17https://bit.ly/UDCheat 18https://ufal.mff.cuni.cz/tred
256 BIBLIOGRAPHY Lennon, Paul. 1991. “Error: Some Problems of Definition, Identification, and Distinction.” Applied Linguistics 12, no. 2 (June): 180–196. https://doi.org/10.1093/applin/12.2.180. Lüdeling, Anke. 2008. “Mehrdeutigkeiten und Kategorisierung: Probleme bei der Annotation von Lernerkorpora.” In Fortgeschrittene Lernervarietäten, edited by P. Grommes and M Walter, 119–140. Tübingen: Niemeyer. https://www.linguistik.huberlin.de/de/institut/professuren/korpuslinguistik/mitarbeiterinnen/anke/pdf/Luedeling_FLV-final.pdf. Lüdeling, Anke, and Hagen Hirschmann. 2015. “Error annotation systems.” In The Cambridge Handbook of Learner Corpus Research, edited by Sylviane Granger, Gaetanelle Gilquin, and Fanny Meunier, 135–158. Cambridge Handbooks in Language and Linguistics. Cambridge University Press. https://www.researchgate.net/publication/291835319. Lüdeling, Anke, Maik Walter, Emil Kroymann, and Peter Adolphs. 2005. “Multi-level error annotation in learner corpora.” In Proceedings of Corpus Linguistics 2005. Birmingham. https://www.researchgate.net/publication/228352566_MultiLevel_Error_Annotation_in_Learner_Corpora. Lukšija, Melita. 2009. “Korpus jako zdroj dat při prezentaci předložek do/na s místním směrovým významem ve výuce češtiny pro cizince.” Bachelor’s Thesis, Masarykova univerzita, Kabinet češtiny pro cizince. https://is.muni.cz/th/emun9/. . 2011. “Korpusy a česká deklinace ve výuce češtiny jako cizího jazyka.” Master’s thesis, Masarykova univerzita, Filozofická fakulta, Ústav českého jazyka. https://is.muni.cz/th/trima/. Machálek, Tomáš. 2017. “KonText – a modern, customizable corpus query interface.” Abstract of a talk presented at the conference Corpus Linguistics 2017, Birmingham. https://www.birmingham.ac.uk/Documents/collegeartslaw/corpus/conference-archives/2017/general/paper341.pdf. Marek, Michal, Pavel Pecina, and Miroslav Spousta. 2007. “Web Page Cleaning with Conditional Random Fields.” In Proceedings of the 3rd Web As a Corpus Workshop, Incorporating CLEANEVAL, 155–162. Louvain-la-Neuve, Belgium: UCL Pressess Universitaires de Louvain. https://ufal.mff.cuni.cz/~pecina/files/cleaneval-2007.pdf.
BIBLIOGRAPHY 257 Martins, André, Miguel Almeida, and Noah A. Smith. 2013. “Turning on the Turbo: Fast Third-Order Non-Projective Turbo Parsers.” In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 617–622. Sofia, Bulgaria: Association for Computational Linguistics. https://www.aclweb.org/anthology/P13-2109. McEnery, Tony. 2018. “Preface.” In Learner Corpus Research: New Perspectives and Applications, edited by Vaclav Brezina and Lynne Flowerdew, xiv–xvii. London: Bloomsbury Academic. http://www.bloomsburycollections.com/book/learner-corpus-research-newperspectives-and-applications/preface-tony-mcenery/. McEnery, Tony, Vaclav Brezina, Dana Gablasova, and Jayanti Banerjee. 2019. “Corpus Linguistics, Learner Corpora, and SLA: Employing Technology to Analyze Language Use.” Annual Review of Applied Linguistics 39:74–92. https://doi.org/10.1017/S0267190519000096. Mendes, Amália, Sandra Antunes, Maarten Janssen, and Anabela Gonçalves. 2016. “The COPLE2 corpus: a learner corpus for Portuguese.” In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), edited by Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, et al. Paris, France: European Language Resources Association (ELRA). http://www.lrec-conf.org/proceedings/lrec2016/summaries/439.html. Meunier, Fanny. 2019. “Tracking developmental patterns in learner corpora: Focus on longitudinal studies.” Selected papers on theoretical and applied linguistics 13:34–44. https://doi.org/10.26262/istal.v23i0.7319. Meurer, Paul. 2012. “Corpuscle – a new corpus management platform for annotated corpora.” In Exploring Newspaper Language: Using the Web to Create and Investigate a large corpus of modern Norwegian, edited by Gisle Andersen, 29–50. Studies in Corpus Linguistics 49. John Benjamins. https://doi.org/10.1075/scl.49.02meu. Meurers, Detmar. 2009. “On the Automatic Analysis of Learner Language: Introduction to the Special Issue.” CALICO Journal 26 (3): 469–473. http://www.sfs.uni-tuebingen.de/~dm/papers/meurers-09.pdf.
258 BIBLIOGRAPHY Meurers, Detmar. 2015. “Learner Corpora and Natural Language Processing.” In The Cambridge Handbook of Learner Corpus Research, edited by Gaëtanelle Gilquin Sylviane Granger and Fanny Meunier, 537–566. Cambridge University Press. http://www.sfs.uni-tuebingen.de/~dm/papers/meurers-15.pdf. Náplava, Jakub. 2017. “Natural Language Correction.” Master’s thesis, Charles University, Faculty of Mathematics and Physics. https://is.cuni.cz/webapps/zzp/detail/188260/. Náplava, Jakub, and Milan Straka. 2019. “Grammatical Error Correction in Low-Resource Scenarios.” In Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019), 346–356. Stroudsburg, PA, USA: Association for Computational Linguistics. https://www.aclweb.org/anthology/D19-5545/. Naughton, James. 2005. Czech: An Essential Grammar. Oxon, Great Britain and New York, NY, USA: Routledge. Nesselhauf, Nadja. 2005. Collocations in a Learner Corpus. Amsterdam: John Benjamins. Nicholls, Diane. 2003. “The Cambridge Learner Corpus: Error coding and analysis for lexicography and ELT.” In Proceedings of the Corpus Linguistics 2003 Conference, edited by Dawn Archer, Paul Rayson, Andrew Wilson, and Tony McEnery, 572–581. Lancaster, UK: Lancaster University: University Center for Computer Corpus Research on Language. http://ucrel.lancs.ac.uk/publications/cl2003/papers/nicholls.pdf. Novák, Michal, Jiří Mírovský, Kateřina Rysová, and Magdaléna Rysová. 2019. “Exploiting Large Unlabeled Data in Automatic Evaluation of Coherence in Czech.” In Text, Speech, and Dialogue, edited by Kamil Ekštein, 197–210. Cham: Springer International Publishing. Novák, Michal, Kateřina Rysová, Jiří Mírovský, Magdaléna Rysová, and Eva Hajičová. 2017. EVALD 2.0 for Foreigners. Data/software. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11234/1-2509.
BIBLIOGRAPHY 259 Novák, Michal, Kateřina Rysová, Magdaléna Rysová, and Jiří Mírovský. 2017. “Incorporating Coreference to Automatic Evaluation of Coherence in Essays.” In Statistical Language and Speech Processing, 58–69. Cham, Switzerland: Springer International Publishing. Osolsobě, Klára. 2010. “Jak se učit česky s korpusem.” In Přednášky a besedy z XLIII. běhu Letní škola slovanských (bohemistických) studií (LŠSS), 112–119. Brno: Masarykova univerzita. Kabinet češtiny pro cizince. Pajas, Petr. 2009. TrEd. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11858/00-097C-0000-0001-48F7-8. Pajas, Petr, and Jan Štěpánek. 2006. “XML-Based Representation of Multi-Layered Annotation in the PDT 2.0.” In Proceedings of the LREC Workshop on Merging and Layering Linguistic Information (LREC 2006), edited by Richard Erhard Hinrichs, Nancy Ide, Martha Palmer, and James Pustejovsky, 40–47. Genova, Italy. Pala, Karel, Pavel Rychlý, and Pavel Smrž. 2003. “Text Corpus with Errors.” In Text, Speech and Dialogue, edited by Václav Matoušek and Pavel Mautner, 2807:90–97. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. https://doi.org/10.1007/978-3-540-39398-6_13. Paquot, Magali. 2019. “The phraseological dimension in interlanguage complexity research.” Second Language Research 35 (1): 121–145. https://doi.org/10.1177/0267658317694221. Paquot, Magali, and Sylviane Granger. 2012. “Formulaic language in learner corpora.” Annual Review of Applied Linguistics 32:130–149. https://doi.org/10.1017/S0267190512000098. PDT – Prague Dependency Treebank 1.0. 2000. Prague. http://ufal.mff.cuni.cz/pdt/. Pečený, Pavel. 2017. “Užití spojovacích prostředků v textech nerodilých mluvčích češtiny.” PhD diss., Institute of Czech Language and Theory of Communication Faculty of Arts, Charles University. https://is.cuni.cz/webapps/zzp/detail/105300/. Petr, Jan, ed. 1987. Mluvnice češtiny. Praha: Academia.
260 BIBLIOGRAPHY Petrov, Slav, Dipanjan Das, and Ryan T. McDonald. 2011. “A Universal Part-of-Speech Tagset.” CoRR abs/1104.2086. arXiv: 1104.2086. http://arxiv.org/abs/1104.2086. Pravec, Norma A. 2002. “Survey of learner corpora.” ICAME Journal 26:81–114. http://icame.uib.no/ij26/pravec.pdf. Preradović, Nives Mikelić, Monika Berać, and Damir Boras. 2015. “Learner Corpus of Croatian as a Second and Foreign Language.” In Multidisciplinary Approaches to Multilingualism, edited by Kristina Cergol Kovačević and Sanda Lucia Udier, 107–126. Frankfurt am Main: Peter Lang. https://www.researchgate.net/publication/304622939_Learner_Corpus_of_ Croatian_as_a_Second_and_Foreign_Language#fullTextFileContent. Rakhilina, Ekaterina, Anastasia Vyrenkova, Elmira Mustakimova, Alina Ladygina, and Ivan Smirnov. 2016. “Building a learner corpus for Russian.” In Proceedings of the joint workshop on NLP for Computer Assisted Language Learning and NLP for Language Acquisition, 66–75. Umeå. http://www.aclweb.org/anthology/W16-6509. Ramasamy, Loganathan, Alexandr Rosen, and Pavel Straňák. 2015. “Improvements to Korektor: A case study with native and non-native Czech.” In ITAT 2015: Information technologies – Applications and Theory / SloNLP 2015, edited by Jakub Yaghob, 73–80. Prague: Charles University in Prague. http://ceur-ws.org/Vol-1422/73.pdf. Reznicek, Marc, Anke Lüdeling, Cedric Krummes, Franziska Schwantuschke, Maik Walter, Karin Schmidt, Hagen Hirschmann, and Torsten Andreas. 2012. Das Falko-Handbuch, Korpusaufbau und Annotationen, Version 2.01. Technical report. Humboldt-Universität zu Berlin, Institut für deutsche Sprache und Linguistik – Korpuslinguistik. https://www.linguistik.huberlin.de/de/institut/professuren/korpuslinguistik/forschung/falko. Richter, Michal. 2010. “An Advanced Spell Checker of Czech.” Master’s thesis, Faculty of Mathematics and Physics, Charles University. https://is.cuni.cz/webapps/zzp/detail/45334. . 2013. Korektor. Software. Charles University in Prague, UFAL. http://hdl.handle.net/11858/00-097C-0000-000D-F67C-5.
BIBLIOGRAPHY 261 Richter, Michal, Pavel Straňák, and Alexandr Rosen. 2012. “Korektor – A System for Contextual Spell-Checking and Diacritics Completion.” In Proceedings of COLING 2012: Posters, 1019–1028. Mumbai, India: The COLING 2012 Organizing Committee, December. http://www.aclweb.org/anthology/C12-2099. Ringbom, Håkan. 1998. “Vocabulary frequencies in advanced learner English: A cross-linguistic approach.” In Learner English on Computer, edited by Sylviane Granger, 41–52. Harlow: Longman. Rio, Iria del, Sandra Antunes, Amália Mendes, and Maarten Janssen. 2016. “Towards error annotation in a learner corpus of Portuguese.” In 5th NLP4CALL and 1st NLP4LA workshop in Sixth Swedish Language Technology Conference (SLTC). Umeå, Sweden: Umeå University. http://www.aclweb.org/anthology/W16-6502. Rio, Iria del, and Amália Mendes. 2019. “Error annotation in the COPLE2 corpus.” Revista da Associação Portuguesa de Linguística, no. 4, 225–239. https://ojs.apl.pt/index.php/RAPL/article/view/42/44. Rosen, Alexandr. 2001. “A constraint-based approach to dependency syntax applied to some issues of Czech word order.” PhD diss., Charles University. http://utkl.ff.cuni.cz/~rosen/public/THESIS/. . 2017. “Introducing a corpus of non-native Czech with automatic annotation.” In Language, Corpora and Cognition, edited by Piotr Pezik, Jacek Walinski, and Krzysztof Kosecki, 163–180. Frankfurt am Main, Bern, Bruxelles, New York, Oxford, Warszawa, Wien: Peter Lang. http://utkl.ff.cuni.cz/~rosen/public/2016_SGT_lodz.pdf. Rosen, Alexandr, Jirka Hana, Barbora Štindlová, and Anna Feldman. 2014. “Evaluating and automating the annotation of a learner corpus.” Language Resources and Evaluation – Special Issue: Resources for language learning 48, no. 1 (March): 65–92. https://www.researchgate.net/publication/234118426_ Evaluating_and_automating_the_annotation_of_a_learner_corpus. Rozovskaya, Alla, and Dan Roth. 2010. “Annotating ESL Errors: Challenges and Rewards.” In Proceedings of NAACL’10 Workshop on Innovative Use of NLP for Building Educational Applications. University of Illinois at Urbana–Champ. https://www.aclweb.org/anthology/W10-1004/.
262 BIBLIOGRAPHY Rychlý, Pavel. 2007. “Manatee/Bonito – A Modular Corpus Manager.” In 1st Workshop on Recent Advances in Slavonic Natural Language Processing, 65–70. Brno. https://www.sketchengine.eu/wp-content/uploads/ManateeBonito_2007.pdf. Schmidt, Thomas. 2009. “Creating and working with spoken language corpora in EXMARaLDA.” In LULCL II 2008: proceedings of the second colloquium on Lesser used languages and computer linguistics: Bozen-Bolzano, Italy, 13th–14th November, 2008, edited by Verena Lyding, 54:151–164. EURAC research. Bozen–Bolzano: Europäische Akademie. http://www.eurac.edu/ Org/LanguageLaw/Multilingualism/Projects/LULCL_II_proceedings.htm. Schmidt, Thomas, Kai Wörner, Hanna Hedeland, and Timm Lehmberg. 2011. “New and future developments in EXMARaLDA.” In Multilingual Resources and Multilingual Applications. Proceedings of GSCL Conference 2011 Hamburg. https://www.yumpu.com/en/document/read/8609912/new-andfuture-developments-in-exmaralda. Šebesta, Karel. 2010. “Korpusy češtiny a osvojování jazyka.” Studie z aplikované lingvistiky 1:11–34. Šebesta, Karel, Zuzanna Bedřichová, Kateřina Šormová, Barbora Štindlová, Milan Hrdlička, Tereza Hrdličková, Jirka Hana, et al. 2017. CzeSL Grammatical Error Correction Dataset (CzeSL-GEC). LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11234/1-2143. . 2019. AKCES-GEC Grammatical Error Correction Dataset for Czech. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11234/1-3057. Šebesta, Karel, Zuzanna Bedřichová, Barbora Štindlová, Milan Hrdlička, Tereza Hrdličková, Jirka Hana, Alexandr Rosen, et al. 2012. AKCES 4. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11858/00-097C-0000-000C-2293-0.
BIBLIOGRAPHY 263 Šebesta, Karel, Hana Goláňová, Tomáš Jelínek, Blanka Jelínková, Michal Křen, Jana Letafková, Pavel Procházka, and Hana Skoumalová. 2013. SKRIPT2012: akviziční korpus psané češtiny – přepisy písemných prací žáků základních a středních škol v ČR. Ústav Českého národního korpusu FF UK, Praha. http://www.korpus.cz. Šebesta, Karel, Hana Goláňová, Jana Letafková, and Jelínková Blanka. 2016. AKCES 1. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL). http://hdl.handle.net/11234/1-1741. Selinker, Larry. 1972. “Interlanguage.” International Review of Applied Linguistics in Language Teaching (IRAL) 10:209–231. https://doi.org/10.1515/iral.1972.10.1-4.209. . 1983. “Interlanguage.” In Second Language Learning: Contrastive analysis, error analysis, and related aspects, 173–196. Ann Arbor, MI: The University of Michigan Press. Sgall, Petr, and Jiří Hronek. 1992. Čeština bez příkras. Praha: H&H. Sgall, Petr, Jiří Hronek, Alexandr Stich, and Ján Horecký. 1992. Variation in Language: Code switching in Czech as a challenge for sociolinguistics. Amsterdam/Philadelphia: John Benjamins. Short, David. 1993. “Czech.” In The Slavonic Languages, edited by Bernard Comrie and Greville G. Corbett, 455–532. Routledge Language Family Descriptions. Routledge. Šindelářová, Jaromíra, and Svatava Škodová. 2013. “Práce s korpusy ve výuce žáků-cizinců.” https://clanky.rvp.cz/clanek/c/z/17481/PRACE-SKORPUSY-VE-VYUCE-ZAKU-CIZINCU.html/. Škodová, Svatava. 2017. “Realizace vybraných komunikačních funkcí v porovnání nerodilých a rodilých mluvčích češtiny.” Studie z aplikované lingvistiky 8:121–135. https://sites.ff.cuni.cz/studiezaplikovanelingvistiky/wpcontent/uploads/sites/19/2017/11/Svatava_Skodova_121-135.pdf. . 2018. “Sloveso JÍT v zrcadle užití nerodilými mluvčími češtiny.” In Čeština jako cizí jazyk v průsečíku pohledů, edited by Svatava Škodová and Milan Hrdlička. Praha: Filozofická fakulta UK v Praze. https://www.researchgate.net/publication/332698345_Sloveso_JIT_v_ zrcadle_uziti_nerodilymi_mluvcimi_cestiny.
264 BIBLIOGRAPHY Škodová, Svatava. 2020. “Sloveso JÍT jako reprezentant pohybové události v prostoru.” Studie z aplikované lingvistiky 11 (2). . n.d. “Genitivní a lokální vazby sloves v češtině nerodilých mluvčích.” In prep. Škodová, Svatava, Kateřina Rysová, and Magdaléna Rysová. 2019. “Comparison of Automatic and Human Evaluation of L2 Texts in Czech.” Journal of Slavic Languages (Seoul, Korea), 93–102. https://doi.org/10.30530/JSL.2019.04.24.1.93. Škodová, Svatava, Barbora Štindlová, Alexandr Rosen, Tomáš Jelínek, and Barbora Hladká. 2019. Příručka k morfologické anotaci češtiny nerodilých mluvčích. Univerzita Karlova, Praha. https://doi.org/10.13140/RG.2.2.34952.78080. Šotolová, Eva. 2008. Vzdělávání Romů. Praha: Karolinum. Spoustová, Drahomíra, Jan Hajič, Jan Votrubec, Pavel Krbec, and Pavel Květoň. 2007. “The Best of Two Worlds: Cooperation of Statistical and Rule-Based Taggers for Czech.” In Proceedings of the Workshop on Balto-Slavonic Natural Language Processing 2007, 67–74. Praha, Czechia: Association for Computational Linguistics. https://www.aclweb.org/anthology/W07-1709/. Stemle, Egon W., Adriane Boyd, Maarten Janssen, Therese Lindström Tiedemann, Nives Mikelić Preradović, Alexandr Rosen, Dan Rosén, and Elena Volodina. 2019. “Working together towards an ideal infrastructure for language learner corpora.” In Widening the Scope of Learner Corpus Research. Selected Papers from the Fourth Learner Corpus Research Conference, edited by Andrea Abel, Aivars Glaznieks, Verena Lyding, and Lionel Nicolas, 427–468. Corpora and Language in Use – Proceedings 5. Louvain-la-Neuve: Presses universitaires de Louvain. Stenetorp, Pontus, Sampo Pyysalo, Goran Topić, Tomoko Ohta, Sophia Ananiadou, and Jun’ichi Tsujii. 2012. “brat: a Web-based Tool for NLP-Assisted Text Annotation.” In Proceedings of the Demonstrations Session at EACL 2012. Avignon, France: Association for Computational Linguistics. https://www.aclweb.org/anthology/E12-2021/. Štícha, František. 2013. Akademická gramatika spisovné češtiny. Praha: Academia.
BIBLIOGRAPHY 265 Štindlová, Barbora. 2011. “Evaluace chybové anotace v žákovském korpusu češtiny.” PhD diss., Charles University, Faculty of Arts. https://is.cuni.cz/webapps/zzp/detail/25046/. . 2013. Žákovský korpus češtiny a evaluace jeho chybové anotace. Praha: Univerzita Karlova v Praze, Filozofická fakulta. Štindlová, Barbora, and Alexandr Rosen. 2012. “Návod k anotaci chybového korpusu.” https://doi.org/10.13140/RG.2.2.24106.64968. Štindlová, Barbora, Alexandr Rosen, Jirka Hana, and Svatava Škodová. 2012. “CzeSL – an error tagged corpus of Czech as a second language.” In Corpus Data across Languages and Disciplines, edited by Piotr Pęzik, 28:21–32. Łódź Studies in Language. Frankfurt am Main: Peter Lang. http://utkl.ff.cuni.cz/~rosen/public/2011-czesl-palc.pdf. Straka, Milan, and Michal Richter. 2015. Korektor 2. Software. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11234/1-1469. Straka, Milan, and Jana Straková. 2017. “Tokenizing, POS Tagging, Lemmatizing and Parsing UD 2.0 with UDPipe.” In Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, 88–99. Vancouver, Canada: Association for Computational Linguistics. http://www.aclweb.org/anthology/K/K17/K17-3009.pdf. Straková, Jana, Milan Straka, and Jan Hajič. 2014. “Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition.” In Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 13–18. Baltimore, Maryland: Association for Computational Linguistics. http://www.aclweb.org/anthology/P/P14/P14-5003.pdf. Stritar, Mojca. 2009. “Slovene as a Foreign Language: The Pilot Learner Corpus Perspective.” Slovenski jezik – Slovene Linguistic Studies 7:135–152. https://doi.org/10.17161/SLS.1808.5274. Tarone, Elaine. 2006. “Interlanguage.” In Encyclopedia of Language and Linguistics, edited by Keith Brown, 747–751. Boston: Elsevier.
272 BIBLIOGRAPHY Hardie, Andrew, 30 Harkins, William E., 233 Heid, Ulrich, 15 Hercíková, Barbora, 15 Hirschmann, Hagen, 15,59,60,62, 63,79,80 Hladká, Barbora, 169,214 Hnátková, Milena, 121 Hofland, Knut, 32 Holub, Martin, 214 Hronek, Jiří, 233 Hudousková, Andrea, 146,208,209, 219 Isahara, Hitoshi, 79 Israel, Ross, 15 Izumi, Emi, 79 James, Carl, 15,80 Janda, Laura A., 233 Janssen, Maarten, 190 Jarvis, Scott, 214 Jelínek, Tomáš, 194 Johansen, Hilde, 32 Johns, Tim, 203 Kaczmarska, Elżbieta, 217 Karlík, Petr, 233,242 Kilgarriff, Adam, 179,190 Kopřivová, Marie, 234 Křen, Michal, 204 Kříž, Vincent, 214 Lee, Sun-Hee, 15 Leech, Geoffrey, 25 Lennon, Paul, 61,62 Leńko-Szymańska, Agnieszka, 22,44 Lüdeling, Anke, 59,60,62,63,72, 79,80,95,107 Lukšija, Melita, 207,208 Machálek, Tomáš, 179 Marek, Michal, 212 Martins, André, 121 Matsson, Arild, 38 McDonald, Ryan T., 125 McEnery, Tony, 25 Mendes, Amália, 34,105,190 Meunier, Fanny, 22,25,27,201 Meurer, Paul, 30,32,33 Meurers, Detmar, 41,69,88,95,214 Mírovský, Jiří, 211 Náplava, Jakub, 139,140,169,213, 214,227 Naughton, James, 233 Nekula, Marek, 233,242 Nesselhauf, Nadja, 22,32 Neumann, Arne, 31 Ng, Hwee Tou, 169 Nicholls, Diane, 33,79 Novák, Michal, 211,219 Oliva, Karel, 244 Osolsobě, Klára, 207 Pajas, Petr, 136,175 Pala, Karel, 212 Paquot, Magali, 22 Pečený, Pavel, 209 Pecina, Pavel, 212 Petkevič, Vladimír, 121 Petr, Jan, 233 Petrov, Slav, 125 Pravec, Norma A., 32 Preradović, Nives Mikelić, 34,190 Ragheb, Marwa, 122 Rakhilina, Ekaterina, 37
BIBLIOGRAPHY 273 Ramasamy, Loganathan, 139,140, 213 Reznicek, Marc, 35 Richter, Michal, 139,192,212,213, 227 Ringbom, Håkan, 22 Rio, Iria del, 34,105,190 Rosen, Alexandr, 78,90,139,140, 158,192,212,213,227,244 Roth, Dan, 88 Roxendal, Johan, 30 Rozovskaya, Alla, 88 Rusínová, Zdenka, 233,242 Rychlý, Pavel, 179,190,212 Rysová, Kateřina, 211,212,219 Rysová, Magdaléna, 211,212,219 Schmidt, Thomas, 31,72,136 Šebesta, Karel, 45,46,169,172 Seegmiller, Steve, 27,105 Selinker, Larry, 15,26,60 Sgall, Petr, 233 Short, David, 233 Šindelářová, Jaromíra, 204 Škodová, Svatava, 113,147,204,206, 208–210,212 Skoumalová, Hana, 121 Smith, Noah A., 121 Smrž, Pavel, 212 Šotolová, Eva, 245 Spousta, Miroslav, 212 Spoustová, Drahomíra, 213 Stemle, Egon W., 25,37,59,189 Stenetorp, Pontus, 175 Štěpánek, Jan, 136 Štícha, František, 204 Štindlová, Barbora, 32,53,78,89, 140 Straka, Milan, 121,130,139,140, 169,194,212,214,227 Straková, Jana, 121,130,194,227 Straňák, Pavel, 139,140,192,212, 213,227 Stritar, Mojca, 32,41 Tapper, Marie, 22 Tarone, Elaine, 15 Tenfjord, Kari, 32 Tetreault, Joel, 88,214 Thomas, James, 204 Tono, Yukio, 54,201 Townsend, Charles E., 233 Tydlitátová, Ludmila, 214,215 Uchimoto, Kiyotaka, 79 Vališová, Pavlína, 204,207,208 Veselovská, Ludmila, 238 Vetchinnikova, Svetlana, 22 Vokáčová, Martina, 209 Volodina, Elena, 38,51,61 Votrubec, Jan, 227 Waclawičová, Martina, 234 Waibel, Birgit, 22,44 White, Lydia, 16 Wirén, Mats, 31 Wisniewski, Katrin, 36,111 Xiao, Richard, 32 Zasina, Adrian Jan, 204,208,209, 217 Zeldes, Amir, 31 Zinsmeister, Heike, 15 Zipser, Florian, 31 Znotina, Inga, 191
Index of Corpora A-GEC, see AKCES-GEC AKCES, 41,45,47,169,210,245 AKCES 1, 46,172 AKCES 2, 46 AKCES 3, 157 AKCES 4, 46,155,157,170,172 AKCES Grammatical Error Correction Dataset, see AKCES-GEC AKCES-GEC, 155,169,214 ASK, 32,39,41 C-GEC, see CzeSL-GEC Cambridge Learner Corpus, 33,39, 176,183,227 Cambridge Reference Corpus, 33 Chyby, 212 CLC, see Cambridge Learner Corpus CNC, 43,46,201,203,207,208,218 COPLE2, 34,35,39,190 Corpus de Português Língua Estrangeira / Língua Segunda, see COPLE2 Corpus of Arabic learners of Czech, 193 Croatian Learner Text Corpus, see CroLTeC CroLTeC, 34,39,190,217 Czech National Corpus, see CNC CzeFL-LONG, 46 CzeSL Grammatical Error Correction Dataset, see CzeSL-GEC CzeSL in TEITOK, 31,32,34,54, 57,87,121,151,155,156, 170,217,223 CzeSL-GEC, 155,169,213,214 CzeSL-LONG, 46 CzeSL-man, 54,108,130,135,136, 141,145,155,156,163, 167–170,191,212,213 CzeSL-man v.0, 156 CzeSL-man v.0, a2, 156 CzeSL-man v0, 31,163,165,169, 172,213,223,224 CzeSL-man v1, 30,121,163,164, 166,168,169,223,224 CzeSL-man v1 downloadable, 56, 164,165 CzeSL-man v1 searchable, 57,121, 165,180 CzeSL-man v2, 31,57,121,167,176, 223,224 CzeSL-MD, 31,140,155,168,170, 224 CzeSL-plain, 42,155,157,158,163, 275
276 BIBLIOGRAPHY 172,180,223,228 CzeSL-SGT, 30,55,57,68,121,139, 151,155,157,158,159–166, 168,170,180,181,213,215, 219,223,231 CzeSL-TH, 107,134,151,155,168, 170,229 CzeSL-UD, 124,130,152,155,169, 170,175,224 EARLYFAMILY 2018, 47 Ein fehlerannotiertes Lernerkorpus des Deutschen als Fremdsprache, see Falko Falko, 31,35,37,39,62,67,71 Frog Story Corpus, 47 Frog, where are you?, 47 Google Web1T, 139 ICLE, 36,39 International Corpus of Learner English, see ICLE MERLIN, 31,36,39,50,71,111, 203,209,211,218,221 Multilingual Platform for European Reference Levels, see MERLIN Norsk andrespråkskorpus, see ASK Open Cambridge Learner Corpus, 34 ORAL 2006, 234 PDT, 121,175 PiKUST, 32,41 Prague Dependency Treebank, see PDT RLC, 37,39,217 ROMi 1.0, 47 Russian Learner Corpus, see RLC SCHOLA 2010, 45,46 SKRIPT 2012, 46,151,155,170,172 SKRIPT 2015, 21,31,46,151,155, 170,172,190 SweLL, 38,39 SYN2015, 166 TLE, 169 WebColl, 212,213
Index of Tools ANNIS, 36,37,39,224 Bonito, 179 brat, 31,140,147–149,168,173,175, 176–178 CNC KonText, see KonText ConSpel, 139 Corpus Workbench, see CWB, 179 Corpuscle, 30,33 CQP, 190 CWB, 30,38,176,190 EVALD, 211 EXMARaLDA, 31,36,136,173 feat, 31,38,56,72,76,134–136,151, 164,173,174,178,194, 196,197,228 IMS Open Corpus Workbench, see CWB KonText, 30,32,43,46,56,157,158, 161,165,166,172,173,176, 178,179,180–184,186,189, 190,208,223,225,227 Korektor, 139,140,152,158,159, 181,192,212,213,215,227 Korp, 30,32,38 LINDAT, 46,47 Manatee, 179,190 Microsoft Excel, 36 MorphoDiTa, 121,152,194 NoSketch Engine, 179 Oxygen, 32 PAULA, 31 PML-TQ, 31 SeLaQ, 31,163,173,177,178,227 Sketch Engine, 30,33,34,39,173, 176,178,179,180,183, 186,187,189,190,227 Speed, 173,174 SVALA, 31,32,38,39 SyD, 208 TEITOK, 31,32,34,35,46,108, 133,136,139,149–152,172, 173,177,178,189,213, 223–225,227,229 TrEd, 130,152,153,173,175 TrEd-ud, 130,153,175 TurboParser, 121 UDPipe, 130,152,153,194,197 Word Sketch, 179 XMLmind, 132 277
Index alternatives categorization, 80,110,148 interpretation, 78 target hypothesis, see target hypothesis: alternative transcription, 77 annotation automatic, 97,120,121,136, 141,143,144,158,170,175, 219 editor, 173,175,189 error, 27,59,143 evaluation, 79,87,130 lemma, 126 linguistic, 28,119 manual, 43,134,147,152,163, 168–170,173–175,189 morphosyntax, 76,86,87,125, 158 morphs, 140,141 scheme, 30 syntax, 121,122,128,175 textual, 27 anonymization, 31,53 authenticity, 202 CCz, see Colloquial Czech CEFR, see Common European Framework of Reference clitic, 20,21,68,238,244 collection, 49 Colloquial Czech, 68,103,111,233 Common Czech, see Colloquial Czech Common European Framework of Reference, 15,37,42 computer-aided error analysis, 22 contrastive analysis, 59 corpus cross-sectional, 27 quasi-longitudinal, 27 longitudinal, 27,43,46 native Czech, 25,46,170 Corpus Query Language, 179,180, 186,198 crosslinguistic influence, see interference curriculum, 16 data-driven learning, 203,207 DDL, see data-driven learning developmental corpus, see corpus: longitudinal developmental patterns, 16 diglossia, 68 emendation, see target hypothesis 279
280 BIBLIOGRAPHY epenthesis, 102,109,110 error, 61 agreement, 19,21,76,77,83,86 aspect, 112 auxiliary, 83,86,87 case, 19,86,87,109,112,159, 210 categorization, 29,81,83,221 category, 29,76,93 clitic, 86,87 diacritics, 68,101 domain, 29,108,147 embedded, see error: successive follow-up, 77,86 formal, 80,97,136 grammar-based, 80,81,219 implicit categorization, 105,150 inflection, 19,20,81,86,109, 111,113,124,127,144,159, 209 interferential, 66 lexical, 83,87,91,92,114,209 location, 108 metathesis, 103 pronoun, 21,209 pronunciation, 97,100,101,111 real-word, 159 reflexive, 20,21,86 span, 61,81,109,112,149,175 spelling, 81,86,97,99,100,111, 113,159 stem, 86 successive, 62,70 taxonomy, 29,78,113 valency, 19,76,83,86,87 word boundary, 21,31,70,81, 105,185,197 word order, 20,67,83,86,87 error analysis, 59 error annotation, see error: annotation error tag, see error: category error tagset, see error: categorization false friends, 66 foreign language, see second language format, 223 conversion, 176 inline, 30,70,136,191 multi-tier, 31,35,39,69,136, 173,177 stand-off, 35,39,136,173,194 structure, 183 tabular, 30,36,71 vertical, 140,179 handwriting, 51,132,191 homonymy, 21,87,124,219 HTML, 53,132 IAA, see inter-annotator agreement IL, see interlanguage information structure, 67,244 intelligibility, 125 inter-annotator agreement, 87,88, 130 interference, 66 interlanguage, 15,16,22,26,43,44, 60,107,202 interlingual errors, 17 interpretation, 66,126,128 intralingual errors, 17 L2, see second language license, 31,43 manuscript, see handwriting metadata, 43,54,202,220 morphosyntax, 120 multi-word unit, 76
BIBLIOGRAPHY 281 natural language processing, 17,18, 211 NLP, see natural language processing normalization, see target hypothesis orders of acquisition, 16,18 overgeneralization, 17 palatalization, 102 PML, see Prague Markup Language Prague Markup Language, 133,136, 173 principle of positive assumption, 61 pro-drop, 70 protection of personal rights, 50 prothesis, 103 pseudonymization, 31 query interface, 32,177,179,189, 223 reconstruction, see target hypothesis Romani ethnolect, 25,42,43,46,245 SCz, see Standard Czech search interface, see query interface second language, 14,15,17,26 second language acquisition, 14–16, 37,59,201 second language teaching, 201 secondary error, see error: follow-up SLA, see second language acquisition stages of acquisition, see orders of acquisition Standard Czech, 18,64,233 style, 87 syncretism, 18 target hypothesis, 22,28,62,66,93, 120,168 alternative, 28,77 automatic, 139 successive, 29,35,87 TEI, see Text Encoding Initiative testing, 16 Text Encoding Initiative, 33,39,57, 136,189 TH, see target hypothesis token, 30,124,180,191,223 transcription, 28,51,132,222 transfer, see interference UD, see Universal Dependencies Universal Dependencies, 122,169, 175,194 word order, 67,244 XML, 39,53,132,136,191