scieee AI-readable full text Open interactive document viewer

Lietuvių kalbos morfologija atvirojo kodo „Hunspell“ platformoje

Virginijus Dadurkevičius

Full text

BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 VIRGINIJUS DADURKEVIČIUS University of Vilnius LITHUANIAN LANGUAGE MORPHOLOGY OPEN SOURCE HUNSPELL PLATFORMS 1 KEY WORDS: computer linguistics, Lithuanian language, grammar, morphology, morphological analysis, spell checking, text indexing, search, open source, Hunspell. INTRODUCTION More than 100 years ago, thanks to the efforts of Jonas Jablonskis (1901, 1918, 1922), the Lithuanian language began to be standardized and adapted to be the main means of unifying the state. Already in the first edition of Jono Jablonski’s Lithuanian Grammar (1901) the phonetics, morphology and syntax of the Lithuanian language were formalized: the basic concepts were defined, the rules of word substitution and sentence structure were presented (see Figure 1). 1 PAV. Excerpt from Jonas Jablonskis' Lithuanian Grammar 1901 2 The invaluable contribution of Jonas Jablonskis to the standardization of the Lithuanian language lexicon, explaining the semantic peculiarities of words. For a long time this was quite enough for pupils and teachers, writers and readers, officials and scientists. With the advent of the innovation that changes civilization – computers – classical grammar and vocabulary had to “put on a new garment” and become a full-fledged 1 The article was prepared based on the presentation Lithuanian language grammar in the open source world, read at the 24th international conference of Jon Jablonskis Digital language resources, their development directions and possibilities of use (2017 09 29); it was organized by the Lithuanian Language Institute General Language Research Center and Vilnius University Department of Lithuanian Language. 2 Paveikslas iš http://www.epaveldas.lt/vbspi/biRecord.do?biRecordId=25383. Virginijus Dadurkevičius. Lithuanian language morphology on the open source hunspell platform |2 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 technology. First of all, language subjects were needed in computers for encoding text (changing letters into numbers), automatically correcting text errors, spelling the text in columns of reciprocal alignment. Later – for information search, speech analysis and synthesis, machine translation from one language to another, automatic perception of content, “sensory” analysis of text, etc. Language subjects had to be formalized in detail and written in a different, computer-readable way. The most important of these subjects is morphology, as it is the basis for any more in-depth work in computational linguistics (syntax, semantic analysis, etc.). The first such work at the turn of the century was carried out by Vytautas Zinkevičius, programmer of the experimental factory “Bitas” (1996, 2000). For nearly two decades it was the indispensable and only tool for modern Lithuanian language research and computer applications. This system has been successfully used for detection of spelling errors in text editing programs, as well as for morphological analysis and synthesis in scientific and applied work at the Lithuanian Language Institute, Vytautas the Great University (where his work was named “Lemuoklis”) and Vilnius University. A variant of this system has even been developed, adapted for the analysis of the ancient Lithuanian language (Gelumbeckaitė et al. 2012). However, with the expansion of the needs of computer linguistics, the shortcomings of this system began to become apparent: non-standard, undescribed data structures; closed source; inadequate realization of proper nouns, illiative, plural, future tense, abbreviated and rare forms of primary verbs. In particular, the closed source of the system hindered its further development – the data and the algorithm could be changed practically only by the author himself. In order to avoid the previously described shortcomings and to create a new generation of computerized Lithuanian language morphology, the following requirements were raised: 1) use and develop only open source; 2) only the data, but not their form and interpretation, may have specific characteristics of the Lithuanian language; 3) all the program code must be universal, suitable for any other language; 4) to use solutions successfully adapted to other world languages. The Hunspell platform meets these requirements well. In addition, after describing the language in this way, it is possible to easily 3 check spelling errors in many (over 50) applications (OpenOffice, LibreOffice, Firefox, Chrome, Safari, InDesign, etc.), perform morphological analysis and synthesis of words, perform intelligent search of textual information, etc. 3 Prieiga internete: https://en.wikipedia.org/wiki/Hunspell. Virginijus Dadurkevičius. Morphology of Lithuanian Language on the Open Source Hunspell BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Platform |3 This article reviews a new implementation of Lithuanian language morphology based on the Hunspell platform. The related problems and their solutions are discussed, examples of successful testing and practical application are given. 1. FEATURES OF LITHUANIAN LANGUAGE DESCRIPTION IN HUNSPELL FORMAT 1.1. Assumptions and general principles of the Hunspell platform The open source Hunspell platform has evolved from simple spell checkers like Ispell, Aspell, MySpell, etc. Gradually making them more complex and adapting them to more and more diverse 4 languages (especially agglutinative languages – Finnish, Hungarian, Turkish, etc.), it became 5 possible to solve not only spelling checking but also morphological analysis and synthesis tasks. 6 The authors of this worldwide popular platform are mostly Hungarian scientists and engineers-programmers (Trón et al. 2005). It is not surprising that the first dot of its name "Hun-" is based on the name of Hungary in English. The morphology of the language is described in two text files: the.aff file contains the substitution rules, and the.dic file contains a dictionary of word forms with initial morphological information and references to one or more groups of substitution rules that can be applied additionally. A separate rule can be imagined as an instruction that specifies: 1) What can be changed to what in a word from the forms dictionary. For example, in the rule SFX 85 čias ty [^š]čias is:Masc_Sg_Voc, "SFX" indicates that the end of the word is being replaced, and it is possible to replace "–čias" with "–ty". "PFX" would indicate that the beginning of the word is being changed. No changes in the middle of the word are provided. 2) The number of the exchange rule group to which this rule belongs. A group can consist of one or more rules that are always applied together. In the previous example, the rule is assigned to a group numbered "85". 3) Conditions under which modification is possible (a certain formalism of regular expressions applies). In our example, the change is possible only if the word from the forms dictionary does not end in -ščias (e.g., „vaikiščias“). (4) Morphological information related to this modification. In our example, the morphological marker Masc_Sg_Voc (masculine, singular, vowel, e.g., „svety“) is associated with such a change in the word ending. 4 Prieiga internete: https://www.cs.hmc.edu/~geoff/ispell.html. 5 Internet access: http://aspell.net. 6 Available online: https://code.google.com/archive/a/apache-extras.org/p/ooo-myspell. Virginijus Dadurkevičius. More detailed information on the principles of Hunspell translation BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 rules and morphological information can be found on the website of this platform, and the possibilities of describing the morphology of various languages in this platform have been quite 7 well explored in the dissertation of Tommi A. Pirinen (2014). 1.2. Specific requirements for Lithuanian realization In order for such a formalization of language to be suitable for morphological analysis, synthesis, indexing of textual information and searching, the following two requirements must be met: 1) only lemas (heading forms) can be entered into the dictionary of word forms; 2) prepositional, prepositional and conjugation must be reflected in the ne.aff, o.dic file. The first requirement is not easy to implement in Lithuanian due to the great variety of verb conjugation (about 170 variants), but the Hunspell platform’s ability to call another conjugation rule helped to solve this problem. Although the "depth" of such calls can only be equal to one, this significantly reduces the size of the exchange rule file, from several million to several thousand lines. Satisfaction of the second requirement slightly increases the size of the.dic file, because, for example, the "take", "take", "take" and "not take" prefix-preposition-return derivatives of the same root will be written on separate lines. And this is not the only negative consequence of this requirement – it is easy to get confused and exclude some less commonly used verb or adjective derivative. True, this problem is solved in a fairly detailed way by selecting lemma's primarily not from published dictionaries, but from accumulated textbooks. Fragments of Lithuanian.dic and.aff files (examples) Several fragments of Lithuanian.dic file (dictionary) with emphasized equivalents of example words by Jonas Jablonskis are shown in Figure 2. 7 Online access: https://github.com/hunspell/hunspell. Virginijus Dadurkevičius. Morphology of Lithuanian language on the open source hunspell BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 platform |5 laukakmenis/118,122,132,137,141,145,157,159,9999 po:noun laukan po:adverb laukas/1,3,7,10,12,14,20,21,30,9999 po:noun Laukavičienė/738,740,745,9999 po:noun_family_name ... svečias/1467,1767,1775,1781,1784,1841,9999 po:adjective svečias/71,73,85,93,98,100,106,116,9999 po:noun svečiavimasis/1308,1310,1314,1315,1317,1319,9999 po:noun_reflexive Svečiulienė/738,740,745,9999 po:noun_family_name ... Gaidulys/118,122,132,141,145,157,9999 po:noun_family_name Gaidulytė/738,740,745,749,754,760,9999 po:noun_family_name gaidys/264,268,278,280,284,288,300,302,9999 po:noun Gaidys/264,268,278,284,288,300,302,9999 po:noun_family_name peilinis/2036,9999 po:adjective peilis/118,122,132,137,141,145,157,159,9999 po:noun peiliukas/1,3,8,10,12,14,20,21,30,9999 po:noun Peipus/444,446,452,9999 po:noun_geographic_name 2 PAV. Fragment ... of Word Forms.dic file. The numbers after the slash indicate references to the addressed groups of switching rules. Highlighted in red are the references to the rules groups shown in 3 EIA. Figure 3 shows some fragments of the Lithuanian translation rules.aff file: SFX 21 Y 1 SFX 21 as spring. is:Masc_Pl_Il ... SFX 98 Y 2 SFX 98 as he is. is:Masc_Pl_Nom SFX 98 as. is:Masc_Pl_Gen ... SFX 264 Y 6 SFX 264 is is. is:Masc_Sg_Nom SFX 264 ys ys. is:Masc_Sg_Nom SFX 264 is dead. is:Masc_Sg_Gen SFX 264 ys žio. is:Masc_Sg_Gen SFX 264 ai ai. is:Masc_Pl_Nom SFX 264 ai ų. is:Masc_Pl_Gen ... SFX 145 Y 12 SFX 145 is iams. is:Masc_Pl_Dat SFX 145 y iams. is:Masc_Pl_Dat SFX 145 is iam. is:Masc_Pl_Dat_short SFX 145 ys iam. is:Masc_Pl_Dat_short SFX 145 is ius. is:Masc_Pl_Acc SFX 145 ys ius. is:Masc_Pl_Acc SFX 145 is right. is:Masc_Pl_Inst SFX 145 ys iais. is:Masc_Pl_Inst SFX 145 is one of them. is:Masc_Pl_Loc SFX 145 and elsewhere. is:Masc_Pl_Loc SFX 145 is iuos. is:Masc_Pl_Loc_short SFX 145 and others. is:Masc_Pl_Loc_short 3 PAV. Change rules.aff file fragments. Several rules are highlighted, which are further used to illustrate morphological analysis. VIRGINIJUS DADURKEVIČIUS. Lietuvių kalbos morfologija atvirojo kodo hunspell platformoje |6 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Using the language description given in the examples and standard (language-independent) Hunspell analysis tools, we would get the following answer for the example words of Jonas Jablonskis (see. Figure 4. ). 4 PAV. Example of morphological analysis of words 2. PROCESS AND RESULTS Once a full-fledged language description in Hunspell format has been created (i.e. the corresponding.dic and.aff files are created), virtually unlimited possibilities for practical applications open up – from spelling or grammar checking to intelligent information retrieval systems. Due to the rather complex morphology of Lithuanian language (many rules and their groups) and the need to combine the list of lems with large (more than 1 billion words) texts reflecting the real state of the current written Lithuanian language, it took about 5 years from the first attempts to the creation of a smoothly operating morphology variant. The examples given earlier in this article are more illustrative, intended only for understanding the idea; in reality, the rules and their groups are much more complex, especially those related to verbs. In order to accelerate the creation of morphology and avoid errors in listing thousands of rules and assigning them to lemas, a preliminary, more abstract than the Hunspell specifications required morphology description was created (rules – in MS Excel table, lemas and their classification – in MS Access database). Software tools have also been developed to automatically convert these primary structures into.aff and.dic files or other forms defined by the application specificity, which are easier to read and understand by humans. The latest version of the Lithuanian.aff and.dic files is available on the author's GitHub site. 8 During the whole process of morphology formalization, the following main sources were used: 8 Online access: https://github.com/dadurka/hunspell_morphology_en. Virginijus Dadurkevičius. Morphology of Lithuanian language on the open source hunspell BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 platform |7 1) Grammar of contemporary Lithuanian language (DLKG 2006); 2) Vytauto Didžiojo universiteto (toliau – VDU) Dabartinės lietuvių kalbos tekstynas 9 – ~140 million words; 3) Vilnius University’s machine 10 translation textbook – ~800 million words, ~4 million phrases. unique; 4) Documents of the Seimas of the 11 Republic of Lithuania (hereinafter referred to as the LRS) – ~400 million words, ~1 million unique; 5) Dictionary of Contemporary Lithuanian Language (JZ 2006) – ~60 thousand lems; 6) Lithuanian language 12 dictionary – ~500 thousand articles; 7) Dictionary of International Words (VTŽŽe) – ~22 thousand lems; 8) Lithuanian Surname Dictionary (LPŽ) (~80 thousand names). In total, approximately 5,000 addressable rules groups (18,000 individual rules) were created. How the rules groups are distributed by language parts is shown in Figure 5. 5 PAV. Distribution of exchange rules groups by language parts.dic file (dictionary) was made up of 171 000 lemas. The distribution of the lemas according to the parts of the language is shown in Figure 6. 9 Online access: http://tekstynas.vdu.lt/tekstynas/. 10 Online access: https://www.versti.eu/. 11 Internet access: http://www.lrs.lt/. 12 Internet access: http://www.lkz.lt/. Nouns (370) Adjectives (238) Numbers (55) Pronouns (64) English verbs (4161) Adjectives (13) Virginijus Dadurkevičius. Morphology of Lithuanian language on the open source hunspell platform BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 |8 6 PAV. Distribution of lemas according to language parts of false, non-Lithuanian, unpresentable, Vykdant atranką buvo apsiribojama tiktai taisyklinga šiuolaikine lietuvių kalba ir vengiama offensive, obsolete words. Proper nouns were further divided into surnames, names, geographical names and other cases. Such selection facilitates the application of the created morphology to spelling check and allows the results of the analysis to be used in later stages of text analysis (syntax, recognition of named entities, indexing and search). The accuracy of the newly developed morphological analyzer is about 98%, i.e. on average 98 out of 100 words are interpreted correctly by the morphological analyzer (Kapočiūtė-Dzikienė et al. 2017). The remaining 2% of unrecognized words are typically spelling errors, words from other languages (first names, surnames, quotes from phrases, etc.), unpresentable, non-normative words and regular words that were not included in the list of lems. 3. PRACTICAL APPLICATION The newly developed morphology of the Lithuanian language on the Hunspell platform was adapted in the VDU Syntactic-Semantic Analysis System in 2011–2015 for spelling and grammar checking, morphological and syntactic analysis of texts (see Figure 7). The morphological analysis 13 was supplemented with mixed statistical-regular unambiguity (hidden Markov model, gold standard approximated trigram probabilities, Viterbi algorithm, several exceptions). More information on unidentification in morphological analysis can be found in Jurgita Kapočiūtė-Dzikienė, Erika Rimkutė and Loïc Boizou’s publication (2017). 13 Internet access: http://semantika.lt/SyntaticAndSemanticAnalysis/Analysis. Nouns (general) (41820) Nouns (verbal) (73697) Adjectives (14552) Numbers (153) Pronouns (53) English verbs (34662) Adjectives (3850) Other (2290) Virginijus Dadurkevičius. Morphology of Lithuanian language on the open source hunspell platform |9 7 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 PAV. Application of Hunspell Lithuanian language morphology in the VDU Syntactic-Semantic Analysis System 3.2. Search engine of the LRS register of legal acts In 2011–2013, parallel to the creation of the new Lithuanian language morphology, it was adapted to the search system of the Register of Legislative Acts in the LRS office. Usually, all words in the 14 documents are replaced with their lemma and the location of the lemma is noted. The text of the sistemose dar prieš vykdant pirmąją paiešką visi dokumentai yra indeksuojami lemuojant, t. y. search query is also determined in the same way before it is performed, so that not only the form of the word written in the query is found later, but also any other form of that word that is present in the searched documents. Since it was not possible to use the.dic,.aff data directly in this system, the original format required by the system (a text file, each line of which is [lemma] [part of language]  [list of forms]) was generated from the above-mentioned initial morphology description. Although this description can generate about 15 million theoretically possible forms of words, due to the internal limitations of the search engine, only about 1 million of the most common ones were used. This solution proved to be successful, the search functions quite smoothly and is used to this day (see Figure 8). 14 Online access: https://www.e-tar.l.t