High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS
Full text
252 Acta Linguistica Lithuanica LXXIII DAIVA ŠVEIKAUSKIENĖ Institute of Lithuanian Language Fields of research: computational linguistics, treebanking, layers of annotation. HIGH QUALITY SYNTACTIC ANNOTATED CORPUS OF LITHUANIAN – VILSINTEKS Aukštos kokybės sintaksiškai anotuotas lietuvių kalbos tekstynas VILSINTEKS ANNOTATION This paper presents a twofold annotation, which is used for the high quality annotation of the Lithuanian corpus. Comprehensive information about a sentence is given in a table and the syntactic structure of a sentence is presented in a picture. The experience of other languages is being used, and specific features of the Lithuanian language are taken into account. The insufficiency of the tree-representation for the syntactic structure of Lithuanian sentences is shown through the statistically annotated examples. The goal of the creation of the annotated corpus bearing exhaustive information is also clearly emphasized. The examples of the annotated sentences are given, which reflect the specific features of the Lithuanian language. KEYWORDS: syntactic annotated corpus; graph representation of the syntactic structure; layers of annotation in the Lithuanian corpus; insufficiency of tree-representation for Lithuanian sentences; goal of the syntactic annotated corpus. ANOTACIJA Straipsnyje aptariamas dviem lygmenimis atliekamas anotavimas, naudojamas aukštos kokybės anotuotam lietuvių kalbos tekstynui sukurti. Išsami informacija apie sakinį nurodoma lentelėje, o sintaksinė sakinio struktūra nubraižoma grafiškai. Naudojamasi kitų kalbų patirtimi, atsižvelgiant į specifinius lietuvių kalbos bruožus. Medžio nepakankamumas vaizduojant lietuvių kalbos sakinių sintaksinę struktūrą parodomas statistiniu
Straipsniai / Articles 253 High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS metodu anotuotų sakinių pavyzdžiais. Straipsnyje taip pat aiškiai pabrėžiamas anotuoto tekstyno, kuriame bus sukaupta išsami informacija, kūrimo tikslas. Pateikiami anotuotų sakinių pavyzdžiai, atspindintys specifinius lietuvių kalbos bruožus. ESMINIAI ŽODŽIAI: sintaksiškai anotuotas tekstynas, sakinio struktūros vaizdavimas grafu, tekstyno anotavimo lygmenys, medžio nepakankamumas vaizduojant lietuvių kalbos sakinius, tekstyno kūrimo tikslas. 1. INTRODUCTION The first annotation of corpora was a part of speech tagging POST [Church 1988: 136]. Such annotation is found in the Penn Treebank. Later, other sentences, carrying the syntactic information, were introduced. They were represented by using various schemes [Atwell et al. 2000: 12]. In 2012 Köhler mentioned the following: “...there is no general standard as to how corpora should be structured and notated” [Köhler 2012: 32]. Therefore, the Lithuanian corpus VILSINTEKS is annotated, bearing in mind the specific features of the Lithuanian language – a large amount of inflexion and a free word order in a sentence. Of the two leading types of syntactic structure representation, which are the phrase structure grammar and dependency grammar, the latter was chosen for Lithuanian because “...dependency grammar has appealed most to students of languages with relatively free word order…” [Kay, Gawron and Norvig 1994: 55]. The most common representation of the syntactic structure is a tree [Allen 1987: 41]. The name of the syntactically annotated corpus – treebank – is related to this term. The name of the Lithuanian treebank does not contain the word ‘tree’, because the syntactic structure of some Lithuanian sentences is represented by a graph with a cycle, that is, it does not meet the conditions established by the definition of a tree, as a tree is a connected acyclic graph [Swamy, Thulasiraman 1984: 33]. The acronym VILSINTEKS means VILniaus SINtaksinis TEKStynas – Vilnius syntactic corpus. 2. ANNOTATION OF THE LITHUANIAN CORPUS Recently, much has been said and written about the low degree of computerization of the Lithuanian language. According to the data of the META-NET
DAIVA ŠVEIKAUSKIENĖ 254 Acta Linguistica Lithuanica LXXIII project, Lithuanian belongs to the group of the least computerized European languages [Vaišnienė, Zabarskaitė 2012: 35]. The software created for other languages does not produce satisfactory results when it is applied to the Lithuanian language. It is high time to speak not about the low computerization of the Lithuanian language but about the quality of its computerization. Thus, it is worth looking back at the Lithuanian language and trying to create one’s own software for the computerization of the Lithuanian language so as to produce high quality computerization of the Lithuanian language. The first endeavors in the syntactic annotation of the Lithuanian corpus were made in 2013. The syntactic annotated corpus VILSINTEKS was created at the Institute of the Lithuanian Language in Vilnius. We aim at providing as much data as possible regarding a sentence, especially if they happen to be the data which are needed during the process of translation. Exhaustive information is given in the Prague Dependency Treebank, which has a three-level annotation [Hajič 2000: 103]. More levels of annotation in one picture could not FIGURE 1. Example of the syntactic structure of the Lithuanian sentence Jie gali būti kitokie – They can be different
Straipsniai / Articles 255 High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS FIGURE 2. Statistically parsed sentence Kas nerizikuoja, tas negeria šampano, bet graudžiai ir neverkia (Who does not risk, that does not drink champagne but does not cry tearfully either) [Kapočiūtė, Nivre, Krupavičius 2013: 15] FIGURE 3. Statistically parsed sentence Bet štai pro medį, kuriame sėdėjau, praslinko didelis šešėlis (But here through the tree in which I sat passed a small shadow) [Kapočiūtė, Nivre, Krupavičius 2013: 15] be achieved so we decided to divide the representation of information into two parts when we deal with Lithuanian sentences: a table, which contains the morphological, syntactic and semantic information, while the syntactic structure should be presented in a picture. A graph with cycles is used for Lithuanian sentences because a tree is not able to reflect the entire syntactic information which the Lithuanian sentence contains. The predicative attribute depends on two parts of the sentence: on the subject or the object, and on the predicate. These relationships are expressed formally and could not be ignored (for more details, see Šveikauskienė 2005: 412). Figure 1 shows the syntactic structure of a title sentence. It contains a predicative, which has formally expressed relationships with the subject and the predicate, and both relationships must be represented by annotating the corpus.
DAIVA ŠVEIKAUSKIENĖ 256 Acta Linguistica Lithuanica LXXIII The predicative must have the ending agreeing with the subject and at the same time it is adjoined to the predicate; that is, it is not an attribute of the subject. The first attempts to syntactically annotate the Lithuanian corpus demonstrated that the software crated for other languages did not produce good results in the Lithuanian language. Some 1500 sentences were annotated statistically at Kaunas University of Technology [Kapočiūtė, Nivre, Krupavičius 2013: 12]. The presented results showed the insufficiency of the information hidden in the FIGURE 4. VILSINTEKS parsed sentence Kas nerizikuoja, tas negeria šampano, bet graudžiai ir neverkia
Straipsniai / Articles 257 High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS FIGURE 5. VILSINTEKS parsed sentence Bet štai pro medį, kuriame sėdėjau, praslinko nedidelis šešėlis
DAIVA ŠVEIKAUSKIENĖ 258 Acta Linguistica Lithuanica LXXIII structure of sentences. The two figures illustrate two parsed sentences, both lacking very important information. The first sentence lacks the arrow between the predicate negeria and subject tas (Figure 2). In the second sentence the agreement relationship is not shown between two words medį and kuriame, which have agreeing endings (Figure 3) when the authors in the same article wrote: “an adjective modifying a noun has to agree in GENDER, NUMBER and CASE”. The words medį and kuriame have to agree in gender and number, and this information is absent in the structure of the sentence. It is difficult to agree that the relations between the words kuriame sėdėjau - in which I sat and pro medį - through the tree are of the same type. These sentences annotated according the method used in VILSINTEKS are shown in Figure 4 and Figure 5 respectively. 3. LAYERS OF ANNOTATION IN THE TABLE The information in the table has three types: • Information about the whole sentence, • Information about the words in the sentence, • Non-grammatical information about the words and the sentence. 3.1. Information about the Whole Sentence Information regarding the whole sentence consists of its code, that is, its position in the corpus, its type, and its features. The feature of the sentence is its characteristic taking into account its function in the text, that is, whether it is a title, an author, a subtitle or a text sentence. The type of the sentence indicates communicative information, that is whether it is declarative, imperative, interrogative, etc., and structural information: personal, impersonal, elliptical, simple, composite sentence, etc. [Ambrazas 1997: 573]. 3.2. Information about the Words in the Sentence Each word is provided with the data about its number in the sentence, morphological data (tense, case, gender, etc.), lemma, data on its lexical semantics,
Straipsniai / Articles 259 High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS syntactic function, direct syntactic relationships with other words in the sentence, and deep cases. Since the word order in the Lithuanian language usually does not have any syntactic information, the features of lexical semantics are very important for identifying its syntactic function. The noun in the accusative case with the feature of time is an adverbial modifier and without it – an object. If two nouns in the accusative case appear in the sentence the feature of the lexical semantics is sometimes the only criterion, which allows one to decide the syntactic function. 3.3. Non-grammatical Information about the Words and the Sentence The stylistic information about the word is given in the table. It helps to choose the right equivalent in the other language when translating the word. The antecedent of the pronoun is necessary. Mille, Wanner and Burga [2012: 5] describe coreferential structure, which links the pronoun with its antecedent in one sentence. The Lithuanian treebank provides the information about the antecedent of the pronoun if it is outside the sentence too. It is very important for translation because the gender of the noun (or pronoun accordingly) may differ in various languages. For example, the Lithuanian sentence Ji buvo graži has three translations into German. If it is “a girl”, the right translation is Es war schön. If it is “a cat” the right translation is Sie war schön, and if it is “a day”, the right translation is Er war schön. All three pronouns in Lithuanian are feminine because all three nouns are feminine, whereas in the case of the German language this pronoun has three different equivalents taking into consideration the noun it replaces. The table contains the missing words in elliptical sentences and the omitted subject, which is expressed by a personal pronoun. It is very often the case in the Lithuanian language. We can guess it from the ending of the verb. The copula of the predicate in the present tense is usually omitted too, so the table contains these missing words. Acronyms are represented in full words in the table. The numbers are additionally represented in the table by numerals and numerals by numbers because the lexical expression of numbers in various languages may differ.
DAIVA ŠVEIKAUSKIENĖ 260 Acta Linguistica Lithuanica LXXIII 4. TYPES OF INFORMATION IN THE SYNTACTIC STRUCTURE In the syntactic structure the sentence is divided at the first stage into a noun group, a verb group and the sentence-end character. The Lithuanian language has a free word order and sometimes the sentence-end character is the only means by which to determine the type of the sentence, i.e. whether it is a declarative or interrogative sentence, for example, Tu šiandien laimėjai prizą. – You have won a prize today. and Tu šiandien laimėjai prizą? – Have you won a prize today?. Figure 6 shows the syntactic structure of the interrogative sentence. The translation of the sentence is chosen according to the sentence-end character. Furthermore, the structure of the sentence is represented using dependency grammar. The head of the subject group is the subject and below are depicted the words that expand it, and the head of the predicate group is the predicate with the subordinated words located below. Other two fields, which are very important for annotating the Lithuanian corpus, are the type of the syntactic relations between the words and semantically irresolvable word groups. FIGURE 6. Syntactic structure of the interrogative sentence Tu šiandien laimėjai prizą? – Have you won a prize today?
Straipsniai / Articles 267 High Quality Syntactic Annotated Corpus of Lithuanian – VILSINTEKS Aukštos kokybės sintaksiškai anotuotas lietuvių kalbos tekstynas VILSINTEKS SANTRAUKA Pastaruoju metu labai daug kalbama ir rašoma apie tai, kad lietuvių kalba mažai kompiuterizuota. Perkama kitoms kalboms sukurta programinė įranga, kuri lietuvių kalbos atveju dažniausiai neduoda patenkinamų rezultatų. Jau atėjo metas, kai reikia pradėti aptarti lietuvių kalbos kompiuterizavimo kokybę, užuot kalbėjus apie menką jos kompiuterizavimą. Straipsnyje aprašomas aukštos kokybės sintaksiškai anotuotas lietuvių kalbos tekstynas, kuriame pateikta patikima informacija. Anotavimas atliekamas dviem lygmenimis: išsami informacija apie sakinį nurodoma lentelėje ir sintaksinė struktūra nubraižoma grafiškai. Anotuojant didelis dėmesys skiriamas sintaksinių ryšių vaizdavimui. Jie parodomi skirtingų spalvų ir tipų linijomis. Straipsnyje aprašomi neskaidomi žodžių junginiai – tai žodžių grupės, kurios tik kartu gali išplėsti kitą žodį, ir tik visą grupę gali pažymėti ją išplečiantis žodis. Kitų sakinio žodžių ryšys su vienu iš neskaidomo junginio dėmenų neturi prasmės. Struktūroje neskaidomi junginiai sudedami į vieną bloką parodant vidinius sintaksinius ryšius. Straipsnyje pateikiami sintaksiškai anotuotų sakinių pavyzdžiai. Sakinio struktūrai vaizduoti naudojamas grafas, nes medis, kuris sėkmingai taikomas anglų kalbos sakinių struktūrai, negali atspindėti visos sintaksinės informacijos, esančios lietuviškame sakinyje. Tai labai gerai matyti statistiniu metodu anotuotų sakinių pavyzdžiuose, kurie pateikiami šiame straipsnyje. Parodomos dviejų sakinių struktūros, kuriose trūksta labai svarbios informacijos: vienoje neparodytas sintaksinis ryšys tarp veiksnio ir tarinio, kitoje nėra ryšio tarp žodžių, kurie derinami skaičiumi ir gimine. Palyginimui straipsnyje pateikiami šie sakiniai, anotuoti ir VILSINTEKS naudojamu metodu. Pradėto kurti tekstyno tikslas – sukaupti išsamią ir patikimą informaciją apie lietuvių kalbos gramatiką, atliekant sakinio analizę be klaidų, t. y. kai kompiuterio darbo rezultatus dar peržiūri žmogus. Anotuotas tekstynas bus viešai prieinamas internete. Kaip dabar iš VDU tekstyno galima gauti pateikto žodžio pavartojimo pavyzdžius, taip VILSINTEKS tinklapyje bus galima gauti sintaksinių struktūrų, kurios turi tam tikrų požymių – pavyzdžiui, vientisinių sakinių, kuriuose veiksniu eina įvardis ir pan., – pavyzdžių. Anotuojant naudojamos Excel lentelės, nes jos leidžia vaizdžiai pateikti sakinio struktūrą ir lengvai pertvarkyti informaciją į XML formatą, kuris plačiai taikomas paieškai. Taigi anotuotą tekstyną bus galima panaudoti statistiniams lietuvių kalbos tyrimams, taip pat informacijai apie lietuvių kalbos gramatiką išgauti ir kt. Įteikta 2013 m. spalio 28 d. DAIVA ŠVEIKAUSKIENĖ Lietuvių kalbos institutas Petro Vileišio g. 5-216, LT-10308 Vilnius, Lietuva [email protected]
