scieee AI-readable full text Open interactive document viewer

Morfologiškai daugiareikšmių formų atsiradimo priežastys

Miglė Žemrietė

Abstract

Straipsnyje, remiantis duomenimis, gautais iš Lietuvių kalbos homoformų žodyno, aptariamos lietuvių kalbos morfologiškai daugiareikšmių formų, straipsnyje vadinamų homoformomis, atsiradimo priežastys. Homoformas pagal atsiradimo priežastis galima suklasifikuoti į tokias grupes: labiau teorinės homoformos, atsiradusios dėl kai kurių žodynų sudarymo specifikos; homoformos, atsiradusios dėl prozodinių elementų skirtumų; homoformos, atsiradusios dėl sutampančių skirtingų žodžių; homoformos, nulemtos lietuvių kalbos morfologijos sistemos; tos pačios leksinės reikšmės homoformos, atsiradusios dėl priskyrimo skirtingoms paradigmoms. Šios homoformų atsiradimo priežastys padeda suvokti, kodėl anotuojant morfologiniais anotatoriais  atsiranda vienokių ar kitokių morfologinio daugiareikšmiškumo atvejų, kodėl negalima gauti visiškai tikslių automatinio anotavimo rezultatų.

Full text

MiGle Earth. causes of morphologically polymorphic forms 107 MIGLĖ ŽEMRIETĖ Institute of Lithuanian Language oRciD id: 0009-0000-2174-3311 Research fields: linguistics of texts, computational linguistics, automatic morphological analysis, morphological polysemancy and its limitation. Doi: doi.org/10.35321/bkalba.2023.96.06 moRFoLogically Multivalued Causes of Forms KEY WORDS: morphological ambiguity, automatic morphological analysis, textbook, homoforms, dictionary. The article, based on data obtained from the Lithuanian language homoforms dictionary, discusses the causes of the emergence of morphologically polymorphic forms of the Lithuanian language, called homoforms in the article. Homoforms can be classified into the following groups according to the reasons for their emergence: More theoretical homoforms, which have arisen due to the specifics of compiling some dictionaries; homoforms, resulting from differences in prosodic elements; homoforms, resulting from the coincidence of different words; homoforms, determined by the morphological system of the Lithuanian language; homoforms of the same lexical meaning, resulting from attribution to different paradigms. These reasons for the occurrence of homoforms help to understand why annotation with morphological annotator causes some or other cases of morphological ambiguity, why it is not possible to obtain completely accurate results of automatic annotation. abstRact in the article, the reasons for the homoforms of the Lithuanian language are described on the basis of data obtained from the Dictionary of Homoforms of the Lithuanian Language. the author offers a classification of homoforms according to the reasons of their appearance, which is as follows: homoforms that are more theoretical in nature and appeared due to the specifics of creating dictionaries; homoforms resulting from differences in prosodic elements; homoforms resulting from different overlapping words; homoforms that arose due to the specifics of the Lithuanian morphology system; homoforms of the same meaning that appeared under different paradigms. these reasons help us to understand why annotation by morphological annotators often result in different cases of morphological ambiguity. 108 Common Language | 96 Questions Nowadays, the study of language is mostly done using various computer programs. they provide the possibility of processing language, solving linguistic problems, segmenting text, determining, classifying, clustering or correcting text errors, automatic translation. Computer programs can also perform semantic, morphological, syntactic analysis, create annotated or unannotated texts, audios, etc. The Lithuanian language is also being computerized. one of the stages of computerization of Lithuanian language – automatic morphological analysis. morphological ambiguity. Morphological ambiguity – only one small part of ambiguity, which includes: 1) variable and unvariable forms of words or words; 2) word forms or words of different and the same parts of speech; 3) words or their forms that differ in certain forms, prosodic elements and completely coincide (Rimkutė 2003a: 63). Morphologically polymorphic words or their forms are considered when two or more lemas are given for the same word form (e.g., the form fights can be given the lemas fights and fight) or two or more grammatical indications (e.g., night – vnk. noun or dgsk. final). According to Erika Rimkutė (2002: 86), morphological ambiguity, which emerges through automatic morphological analysis, is a very specific case, but people encounter ambiguous words on a daily basis. Most of the time, this does not cause communication problems, because they are able to guess quite easily which word or its form is used, and other variants do not even come to mind. Moreover, often morphologically polymorphic words in the sentence language are recognized by different intonation or place of accent. Although the problem of morphological polymorphism is basically solved (see Rimkutė, Daudaravičius 2007), even with improved morphological annotators of the Lithuanian language, it is not always possible to choose the most likely form from several morphologically polymorphic forms, and there are also frequent cases when the morphological annotator incorrectly chooses one form from several possible ones. The Lithuanian language can be analyzed with three morphological annotators that are described in detail and are publicly available – the Semantika.lt1 annotator available in the syntactic and semantic analysis system and the Lemuoklis 1 annotators See https://semantika.lt/analysis/textanalysis. MiGle Earth. Causes of morphologically polymorphic forms 109 and Morfuoklis2. The data presented in this article are based on Lemuoklis’s analysis, so this annotator will be presented in more detail. The aim of this article is to describe the causes of homoforms in Lithuanian language (this and other terms will be explained in Chapter 2). tasks – to review terms related to morphological ambiguity and, based on analyzed cases of morphological ambiguity, after describing and classifying them, to discuss the causes of homoforms. analyze the words in the Lithuanian Homoforms Dictionary (LkHŽ)3. quantitative analysis, textbook linguistics, descriptive methods applied to the study. LkHŽ provides unique examples of 1 million words of homoforms found in morphologically annotated text. from the dictionary it is possible to see which forms of Lithuanian words are morphologically polymorphic. readers can draw conclusions for themselves, which forms are more likely and what determines this. The dictionary contains about 36 thousand homoforms, which are classified into three categories: 1) homoforms of variable parts of speech; 2) homoforms of invariant parts of speech; 3) homoforms of variable and invariant parts of speech. homoforms of variable parts of speech, as well as homoforms of variable and non-variable parts of speech are further divided into 53 types (they will be mentioned below). e. Rimkutė, preparing a morphologically annotated textbook of the Lithuanian language, found that morphological polysemanticism includes both variable and non-variable parts of the language (Rimkutė 2006a: 136). According to her, most often the variable parts of the language coincide, most rarely – the variable with the non-variable ones. Although the dictionary does not provide statistical data4, according to which kinds of homoforms are assigned the most word forms, the following kinds of homoforms can be mentioned as common: (i)a conjugated nouns vnsk. dgsk of noun and (i)o conjugation nouns. ending coincidences (e.g.: considerable, acknowledged, third); (i)o and (i)a consonant nouns vnsk. final coincidences (e.g.: indifferent, eighth, sold); (i)o, ė of consonant nouns vnsk. nobleman and dgsk. coincidences of nominative and adjective (e.g.: abbeys, indifferent, most abstract) (for more details on the structure and preparation of the dictionary see the work of the author of this article: Piečytė 2023). 2 See https://sitti.vdu.lt/morfuoklis/. 3 this dictionary is being prepared for publication. not yet made public. authors – m. Žemrietė and e. Rimkutė. 4 Frequency estimation would require a separate study. but, after analyzing some of the homoforms in the texts, it becomes clear that some of them can be used thousands of times, others – only a few times. 110 common language | 96 The source of this article is a morphologically annotated textbook of 1 million words of Lithuanian language prepared in 2000–2005. it is annotated using the program Lemuoklis created by Vytautas Zinkevičius. in the morphologically annotated text almost half of the words are morphologically ambiguous. this phenomenon became apparent only when researchers began to process the annotated files of the automatic morphological analysis program, when they saw how language is treated without contextual knowledge, i.e. when each word form is analyzed as completely separate (Rimkutė 2006a: 136). Therefore, this article and its explanations should be useful for those who use Lithuanian morphological annotations. it should help to understand the reasons for any inaccuracies. LkHŽ is the first Lithuanian dictionary of its kind. it shows an unconventional approach to language, presented through the prism of computerized Lithuanian language analysis programs. This dictionary shows what language computer programs “understand”. Foreigners who are learning Lithuanian may also have a similar view of the Lithuanian language. Since the dictionary consists of a very large number of homoforms, it opens up more possibilities for researchers. Although, as mentioned, this dictionary is the first in Lithuania, similar dictionaries have already been published in other languages. For example, in 2001, a dictionary of Russian homoforms was created by zhanna grigoryevna anoshkina (Словаръ омонимичных словоформ)5. This dictionary includes both homonyms and homographs (words that are written the same but pronounced differently). The dictionary is arranged so that the word forms are on the left, and the lexicon on the right. parts of the speech are given in parentheses, plg.6: ехидна ехидной ехидна (с), ехидный (п) ехидною ехидны ехидно ехидно (н), ехидный (п) 5 The Russian dictionary of homoforms is available online at: http://cfrl.ruslang.ru/homoforms/index.htm. 6 This example shows the conjugation and parts of the word ехидна (their abbreviations are given in parentheses). In Russian, this word has more than one meaning, which depends on the context. The word ехидна in Lithuanian can mean a hedgehog, an Australian snake or can be used to describe an angry, cruel, treacherous person. MiGle Earth. 111 In 2008, a dictionary of English homophones7 was created by alfred aloisi (for more on the relationship between homophones and homoforms, see section 2 of this article). The description of the dictionary states that the dictionary is constantly updated, but there is a lack of information about where the words are collected, how many there are, etc. After contacting the author of the dictionary, it became clear that the first version of the dictionary was paper, but it was quickly understood that it was more practical to create an online version, because the paper version could quickly become obsolete and would require more than one edition. The dictionary contains more than 10,000 homophones, and they can be suggested by users of the dictionary. moreover, not only the lemas of homophones, their parts of speech, but also the definitions are presented here – this is not very common for dictionaries of this type. Differently spelled words with the same pronunciation, such as moo (lit. mū) and moue (lit. grimace), pg.: moo 1. :: verb-intransitive to emit the deep, bellowing sound made by a cow; low. 2. :: noun the lowing of a cow or a similar sound. moist 1.:: noun a small grimace; and handcuffs. English homophones are also included in paper dictionaries, such as Leslie Presson's Dictionary of Homophones, published in 1997. it is designed for learners of English, has more than 600 homophone pairs. In 2012, Reed published Reed's Homophones: A Comprehensive Book of Sound-alike Words, a comprehensive dictionary of American English homophones, with more than 1,000 pairs, and a definition, part of speech, and example of use for each word. In Lithuania, books, textbooks, scientific articles on similar topics have also been prepared. In 2004 jonas šukys published the book Similarly sounding words: norms and errors, which analyses paronyms in detail (the term will be explained in section 2). In the book, the scholar explains the normative meanings of the words compared, indicates and corrects incorrect usage, and also discusses aspects of the usage of paronyms and homographs. Some homomorphisms occur due to the random coincidence of words with different meanings. Rūta marcinkevičienė’s (2011) textbook The Meaning of the Word is dedicated to the analysis of such words. Dictionaries and textbooks, one section of which writes about polysemy, here homonyms are described in detail, their classification is given. 7 English homophone dictionary available online: https://www.homophone.com. 112 benDRinė kaLba | 96 e. Rimkutė published a particularly large number of works on morphological polysemanticism. We can mention her dissertation (Rimkutė 2006a), several scientific articles (Rimkutė 2002, 2003a, b), where various types of polysemanticism are described extensively, with particular attention paid to morphological polysemanticism and homoforms. This paper consists of an introduction (first part), two theoretical parts, two research parts and a summary. At the end of the thesis, a list of all cited literature sources is provided. The first morphologically annotated textbook of Lithuanian language is an electronic database created by Vida Žilinskienė and Laima Grumadienė. it was used specifically for the preparation of common dictionaries (Rimkutė 2006b: 34). v. Žilinskienė and L. Grumadienės Contemporary Written Lithuanian Language Common Dictionary electronic textbook (2002) consists of 1.2 million words, the texts contained in it are original (not translated) and belong to the journalistic, scientific, fictional or office style. These texts were annotated, and later dictionaries were created based on them – the Frequently Used Dictionary of the Present Written Lithuanian Language (1997) and the Frequently Used Dictionary of the Present Written Lithuanian Language (1998). Vytautas Magnus University Computer Linguistics Centre started creating the morphologically annotated textbook mentioned in the introduction in 20008. it was compiled semi-automatically, using the automatic morphological analysis program Lemuoklis. The first version was about 1 million words, but later the textbook was supplemented, inaccuracies were corrected9. The textbook is composed of 36 per cent journalistic texts, 24 per cent scientific literature texts, 19 per cent fiction texts, 2.8 per cent administrative style texts and 6.8 per cent the Seimas of the Republic of Lithuania stenograms. at the time of the compilation of this textbook, the morphological analysis program prepared by v. zinkevičius 8 this textbook is publicly available on the Internet: https://clarin.vdu.lt/xmlui/handle/20.500.11821/33. 9 The data for the Lithuanian language homoforms dictionary were taken from this first version of the textbook, because the aim was to record as many cases of morphological polysemy as possible. Later, statistical unambiguousness modules were installed in Lemuoklys (Rimkutė, Daudaravičius 2007), as a result of which this program began to work more accurately, providing much fewer morphologically unambiguous forms. MiGle Earth. 113 The lemma tool, presenting the lemmas and morphological indications, did not take into account the context, the program did not include information on semantics, word forms were usually determined by having a list of Lithuanian word roots, rather than using any lists of word forms with the grammatical indications of those forms (zinkevičius 2000: 30). This means that the program provided all possible grammatical indications and all possible lemas (secondary forms) for word forms (see examples in the introduction). It is also worth mentioning that Lemuoklis analyzed written forms without accents. For this reason, homoforms could also differ in prosodic elements (Rimkutė, grybinaitė 2004: 75–76). at present, the Semantika.lt annotator is used more frequently in language technology research, because it is more accurate; can be managed, updated and supplemented (bielinskienė et al. 2017: 2). but he also does not always annotate correctly – often does not recognize foreign words, jargon, abbreviations. not seeing the whole language or context as a human sees it, making annotation errors for morphologically polymorphic words (if a single word has several different grammatical meanings, morphological annotator sometimes chooses the wrong grammatical notation). very recently, at the end of 2023, a new word and text analysis program Morfuoklis based on the Semantika.lt annotator started operating. With this program it is also possible to synthesize, i.e. to generate the desired word or several according to the specified grammatical marks. The majority of annotation errors are due to morphological ambiguity: although programs are being improved, they still very often fail to recognize context (or recognize it incorrectly). Recently, considerable attention has been paid to solving this problem of automatic analysis of the Lithuanian language – various tools have been developed that can limit semantic, syntactic and morphological ambiguity, rules have been developed that help unambiguize texts, morphological ambiguity of the Lithuanian language has been analyzed and classified (Rimkutė, grigonytė 2006: 32). Homomorphisms are common not only in Lithuanian. A similar situation has also emerged when examining texts in other languages, for example, in Czech morphologically polymorphic forms account for 46%, in Slovenian – about 8.01%, in Romanian – 40%, in English – 38.65%, in Hungarian – 21.58% (Hajič 2004: 173, quoted from Rimkutė 2006a: 47). According to the data provided by Estonian scholars (Puolakainen 2012: 194), it can be said that almost half of the Estonian words are morphologically ambiguous. evalda jakaitienė (2009: 87) stated that there have been considerations as to whether such a phenomenon as word polysemanticism can exist at all – allegedly if one 114 common language | A sequence of 96 sounds has two different meanings, in which case it should be two separate words, not one polysemantic one. It is important to note that the researcher here speaks of lexical, not morphological, ambiguity. However, this opinion is not popular – according to aleksandravičiūtė (2008: 275), polysemanticism is a systematic language phenomenon, because it is related to the peculiarities of the Lithuanian language’s morphological system and syncretism of the vowels, and also to the economy and limitation of the language. 2. most important moRFoLogic In linguistics, morphological polysemancy is usually understood and referred to by several different terms. the following is a discussion of terms, i.e. how to name morphologically polymorphic forms: homoforms, homonyms, homographs or otherwise (terms related to other forms of polymorphism are not discussed, the main focus is on morphological polymorphism). The broadest term, which could be used in part in this article, is homonyms, they are usually defined as “words having the same pronunciation and completely different meanings” (jakaitienė 2009: 101), as “a word sounding the same as another word, belonging to the same part of the language, but having a different meaning” (gaivenis, keinys 1990) or as “words pronounced the same (there are also homonymous forms, morphemes)” (šukys 2004: 9). as mentioned earlier, morphologically polysemantic are called variable and unvariable, equally and differently pronounced forms. In analyzing morphological polysemancy, the lexical meaning is not important (this is usually emphasized when describing homonyms), so it is obvious that the term homonyms is too broad when talking about morphologically polysemantic forms. The term homoforms causes a lot of confusion among linguists. Some of them (barauskaitė 1979: 43–44; barauskaitė et al. 1995: 23) consider different forms of the same word or similarly sounding forms of different words as homoforms, e.g., aprašai – and the noun aprašas dgsk. noun, and the verb to describe the present tense of the second person form. Nikolaj Pavlovich Kolesnikov also calls homoforms paradigmatic homonyms (kolesnikov 1987: 7, quoted from nevzorova et al. 2005: 231), while Francis calls the overlapping forms of the same or different parts of the language syncretism (katamba 1994: 15). MiGle Earth. 115 For the purposes of this article, morphologically polymorphic forms are referred to as homoforms. As will be seen from the following overview of the usage of the term, the term homoforms is used in this article with a broader meaning than is usual in linguistic works. Therefore, without delving into all types of homonymy, the concept of homoforms in linguistic works is described below. e. jakaitienė (2009: 103) defines homoforms as separate pairs of completely non-homonymous words that coincide randomly in their phonological structure. as an example, the scientist gives the following sentence: Mama nuo vaiko muses vaiko, which coincides with the noun vaikas vnsk. the originator and the verb in the third person form of the present tense. According to the researcher, several grammatical forms of some words may also coincide, for example: lipti, lipa, lipo ("to move up or down") and lipti, limpa, lipo ("to hang, stick"). This is the case with the forms of the verbs and the forms made of them. It is problematic that others classify such cases not as homoforms but as partial homonyms (barauskaitė et al. 1995: 24). In the Kalbotyros terminų žodynas (ktŽ), homoforms are considered to be the same-pronounced and written forms of different words (gaivenis, keinys 1990). Thus, homoforms do not include cases of syncretism of vowels, when several vowels of the same word coincide. However, such a definition may create confusion in distinguishing between homoforms and partial homonyms. In the dictionary, the latter are defined as "homonyms whose morphemes do not completely coincide or belong to different parts of the language" (gaivenis, keinys 1990: 80). e. jakaitienė defines partial homonyms as such words, “which either have the entire grammatical forms of one of them coinciding with a part of the forms of another word, or have at least several forms of them coinciding” (jakaitienė 2009: 101–102). According to her, in such cases, one word is usually variable and the other is unchangeable, for example, gaila – it is both a noun, meaning regret, and an adverb (examples from morphologically annotated text). According to e. jakaitienė, in analytical languages, partial homonyms are words that may belong to different parts of the language, for example, the English word tender can be both a noun denoting a proposal and an adjective denoting gentle, sensitive, and a verb meaning “to offer, to present” (jakaitienė 2009: 102). other scholars (kaljanov 2023: 27) call such cases slightly differently – grammatical homonyms. this phonetically su- 122 common language | 96 In the International Dictionary of Words (IDW), the word bravas is defined as “in the 17th–18th centuries in Italy – a hired assassin, an adventurous man capable of using any kind of violence”, ata – “a hundredth of a Lao kip”, or – “1. covered wooden bicycle carriage used by central Asian peoples; where. in Afghanistan for the transport of heavy goods; 2. four-wheeled vehicle for the transport of grain; used in the Caucasus and P. Ukraine’. such words usually take time to distinguish from unlikely homoforms, because each of them has to be checked in different dictionaries and textbooks. It is also difficult to decide whether to include them in the LkHŽ, because it is no longer sufficient to rely on the data of other dictionaries. This reason for the occurrence of homoforms shows that there is a dilemma: the larger the lexicon of the language analysis program, the more words and their forms the program should recognize, but at the same time this means that unexpected coincidences – cases of morphological polysemanticism – can occur. therefore, in order to achieve the most accurate results, the principles of selection of words used in computer programs must be well considered. 4.2. Homoforms resulting from differences in prosodic elements Homoforms due to differences in prosodic elements (e.g., accent, preposition, vowel length) occur only in written language. Since Lithuanian morphological annotators analyze written forms of words without accents, homoforms often differ in prosodia. These homomorphic examples are mainly found among the overlaps of variable parts of the language, they are found in most of the smaller species. It is often the overlapping of two or more forms of the same word. There are 11 types of such examples in the dictionary: (i)o conjugated nouns vnsk. denominator coincidences with vnsk. as a noun and a pronoun (e.g.: agentūrà vs. agentudra, mamà vs. mãma); (i)o, ė of consonant nouns vnsk. nobleman and dgsk. denominator coincidences (e.g. absolute vs. absolute; adequate vs. adequate; square vs. square); ė amusing nouns vnsk. coincidences between the noun and the pronoun10 (e.g. katè vs. kãte, dumblè vs. duuble); dgsk of the past tense of the communicative and passive kinds of the senior g. participants coincidences of the denominator (e.g.: 10 note that the exclamation mark is only theoretically possible in many places. The semantics of nouns was not taken into account when compiling the homomorphic dictionary. MiGle Earth. Causes of morphologically ambiguous forms 123 apdrėbti vs. apdirbt; apeet vs. apeit); dgsk of the past tense senior g. participants of the passive and passive kind cognate coincidences (e.g. aplstų vs. aplytų; deñgtų vs. dengtų); (i)a verb noun vnsk. overlapping of placeholder and vowel (e.g. aidè vs. áide; garažè vs. waragē); (i)u and i inverted nouns vnsk. denominator and dgsk. final overlaps (e.g.: circumference vs. circumference; circle vs. circle); current time vnsk. coincidences of third and second person (e.g. neprigidi vs. neprigirdė; negãli vs. negalė); pronouns mot. g. vnsk. coincidences between the genitive and the adjective (e.g., tái vs. taang, šštai vs. šitaang; Of course, this can be a particle, but such coincidences in the dictionary fall into another category – the category of coincidences of variable and invariable parts of speech, because the dictionary, as mentioned, presents only grammatical notes suitable for a specific kind); (i)u conjugated nouns vnsk. back and dgsk. ancestral11 coincidences (e.g. ãals vs. ala; lieetų vs. lieted); the conjugation of verbs in the present tense in the second person (e.g. iñkšti vs. inkštas; klųsti vs. klystas). In addition, the differences between the two types of language can be explained by the differences between the different types of language. There are 7 types of homoforms of this type in the vocabulary: coincidences of nouns and verbs (e.g. : Adomo – corresponds to the proper noun Adõmas12 vnsk. of the progenitor and active species participant ãdomas vnsk. the form of a nobleman; apsaugos – coincides with noun protection vnsk. kilmininko, dgsk. future tense forms of the adjective and the verb protect, e.g.: apsaugõs vs. apsáugos); coincidence of nouns and adjectives (e.g. atlaidus – coincidence of noun atlaidus dgsk. ending and adjective atlaidus vnsk. nominative forms, cf. ãtlaidus vs. atlaidùs); verbs and adjectives coincidences (e.g. añtrinę, asmẽninę vs. antręnę, asmenęnę – coincidence of verbs seconded, personalized of the active kind of the past participle and adjectives secondary, -ė, asmeninis, -ė mot. g. vnsk. tailbone shape); overlapping of verbs and pronouns (e.g.: kitatu vs. kkattai; kùria vs. kurià – overlapping of the verbs kisti, kurti and the pronouns kitas, kuri); coincidences of nouns, adjectives and verbs (e.g.: balti – noun baltis vnsk. exclamation form baätti; adjective baltas dgsk. noun form baltė; communicative bálti); coincidences of nouns and numerals or pronouns (e.g. : Anos – coincides with the proper noun 11 times (as in the following cases) dgsk. originator possible only theoretically. 12 truth, the annotator is able to distinguish between uppercase and lowercase letters. However, morphological ambiguity could be created if the participant was capitalized, for example, at the beginning of the sentence. 124 common language | 96 Ana and pronoun form of Ana, plg. ânos vs. anõs; kelias – the noun kẽlias coincides with the pronoun keliàs); coincidences of nouns, verbs and pronouns or adjectives, or numerals (e.g. dvejas – coincidence of the numerals dvẽjos, the noun dveja vnsk. genitive form dvejõs and dgsk. nominative form dvẽjos, as well as the future tense verb dvejõs). Such coincidences were also observed between homoforms of variable and fixed parts of speech, they are 8 kinds in the dictionary: coincidences of verbs, adjectives and adverbs (e.g., blogai – coincidence of the verb blogti vnsk. second person past participle form blógai, adjective blogas, -a mot. g. vnsk. noun form blõgai and adverb blogaų); coincidences of adjectives and adverbs (e.g., aklai, efektyviai – coincidence of adjectives in the form of the noun ãklai, efektųviai and adverbs aklaaj, efektyviai); coincidence of nouns and adverbs (e.g. seniai – coincidence of the noun’s old dgsk. noun form sẽniai and the adverb seniaų); coincidences of nouns, adjectives or pronouns and adverbs (e.g.: aiškiai, aukštai – coincidence of nouns aiškis, aukštas dgsk. noun forms aahškiai, aũkštai, adjectives aiškus, -i, aukštas, -a mot. g. noun forms áiškiai, áukštai and adverbs áiškiai, aukštaų); coincidences of nouns, verbs, adjectives and adverbs (e.g., putih – coincidence of the noun buta dgsk. denominator and vnsk. vocative form bùkai, verb bukti past perfect tense vnsk. second person form bukakau, adjective bukas mot. g. vnsk. noun form bùkai, adverb bukaang); coincidences of nouns, prepositions and adverbs (e.g., pirma – coincidence of the first noun of the noun, -a mot. g. vnsk. noun and adverb, the genitive forms of the pronoun pirmà, pärma, and the preposition and adverb pärma); coincidences of verbs and adverbs (e.g. apmaudžiau – the verb apmaudyti in the past tense vnsk. first person form apmáudžiau and the adverb apmaudžiaũ of the higher degree), coincidences of nouns, pronouns and adverbs or particles (e.g. Ana – coincidence of the noun, conjunct and exclamation mark of the actual noun Ana vnsk., pronoun and particle forms ana, cf. Anà, ënna and anà). If Lemuoklis and other morphological annotators were able to analyze the interlaced texts13, morphological ambiguity would probably be significantly reduced. 13 it is not difficult to do this automatically, for example, the program Kirčiuoklis, it is available at https://kalbu.vdu.lt/mokymosi-priemones/kirciuoklis. On the other hand, this program is based on the results of an automatic morphological analysis program, so a vicious circle is formed: the quality of the spelling program depends on how well the morphological analyzer works. It should be noted that the pronunciation dictionary provides several variants of pronunciation for morphologically polymorphic forms, so this would not help to solve the problem of morphological polymorphism. MiGle Earth. The reasons for the emergence of morphologically polymorphic forms 125 and earlier were attempts to classify homoforms according to prosodic elements. e. Rimkutė (2002: 91–98) analyzed 10,000 grammatically annotated words used in the articles of Lietuvos rytas. 83% of the homoforms are phonetically identical, and 17% differ in the place of the consonant. Homoforms, which do not differ in the place of accentuation, the scholar further divided into two types: those that differ in the number of vowels (e.g., the verb mãno and the pronoun manye) and those that differ in the suffix (e.g., the verbs klaũsė (lema – to listen) and kláusė (lema – to ask)). The previous examples discussed both the place of conjugation and the conjugation-different homoforms. In the latter category, those cases that differ in the number of vowels are discussed together, but they are not separately distinguished. 4.3 Homoforms resulting from the coincidence of different words Part homoforms of variable parts of speech, homoforms of variable and non-variable parts of speech fall into this category. Theoretically, all cases of the homomorphic category of variable and invariant parts of speech are also words of different meanings, because they have different lemas. but this is not always true, because, for example, there is quite a lot of semantic commonality between adverbs and adjective overlaps, one form is often made of the other. Therefore, homoforms resulting from polysemantic words could be divided into two more smaller categories – homoforms of different lexical meaning and homoforms of partial lexical meaning. 4.3.1. This category includes homophones of different lexical meanings that are not related in adjective terms. Among them – 8 types of homoforms of variable and fixed parts of speech: homoforms of verbs, adjectives and adverbs, for example, griežtai – in this case, the participle griežtas, -a (lemma – strict) mot. g. vnsk. the form of the noun with the adverb having a completely different meaning stingrai and the adjective griežtas, -a mot. g. vnsk. the beneficiary form. Thus, in this case, the participle is linked to the adverb and the adjective by a different lexical meaning, and the adjective and the adverb are partially overlapping, because they are made of each other (see section 4.3.2 for more details); homomorphisms of nouns and adverbs or prepositions, conjunctions, particles, emoticons, adjectives (e.g.: dėlei – the forms of the noun dėlė vns. the noun and the preposition dėlei are the same; opa – the forms of the noun opa vns. the noun, exclamation mark, adjective and the adjective are the same); homoforms of pronouns and adverbs or particles, conjunctions 126 common language | 96 (e.g., who – coincidence of adverb and pronoun); homoforms of nouns, adverbs and adjectives or pronouns (e.g., mot. g. noun kandis, i.e. “a small moth, whose caterpillars chop clothes, eat grains, flour, plants” (LkŽ 2005), vnsk. the noun kandžiai and the senior noun kandis, i.e. „biting, biting, bitten place“ (LkŽ 2005), dgsk. noun kandžiai coincides with the adjective kandus, -i mot. g. vnsk. the noun and the adverb kandžiai (the lexical meaning of the adjective and the adverb coincides partially, while the noun kandis has a different lexical meaning); homomorphisms of nouns, verbs, adjectives, adverbs and prepositions (e.g., near – the forms of the noun arti vnsk. exclamation, the verb árti communicative, the adjective artus, the -i mot. g. vnsk. noun, the adverb and the preposition artaan); homomorphisms of verbs, nouns and adverbs (e.g., sometimes – the forms of the participle kártas dgsk. adverb, the noun katas dgsk. adverb and the adverb kartais coincide); homoforms of verbs and prepositions or particles, conjunctions, emoticons, abbreviations (e.g.: ak – the emoticon coincides with the noun akti in the third person singular singular) the form of the injunction; zavisti – the forms of the third person future tense of the verb zavisti and the adverb zavis coincide; vnsk – the graphic abbreviation of the word episkop coincides with the third-person imperative conjugation of the verb razvijati. forms); homomorphisms of nouns, pronouns and adverbs or particles (e.g. ana – coincidence of the proper noun Ana, the pronoun and the particle ana). homoforms of different lexical meanings could also exist between homoforms of variable parts of language. Homomorphisms of nouns, verbs and adjectives of different lexical meanings (e.g., anglių – coincidence of the noun anglys and anglė dgsk. pronoun forms); homoforms of nouns and verbs (e.g., adamą – the forms of the participle adamas vnsk. and the actual noun Adam vnsk. are the same); homomorphisms of verbs and pronouns or numerals (e.g. jos – the third-person future tense of the verb joti coincides with the pronoun ji dgsk. denominator and vnsk. the form of a nobleman; kitai – the second person of the past tense of the verb kisti and the pronoun kita vnsk. the form of the beneficiary; trejos – the third person form of the verb trejoti in the future tense and the form of the numerals treji, -os in the nominative); homomorphisms of nouns, adjectives and verbs (e.g.: kaltų – corresponds to the third person form of the verb kalti in the pronoun, the participle káltas, -a mot. and senior g. dgsk. adjective form, adjective kakt, -a mot. and senior g. dgsk. the nobleman's form and the misty Earth. Causes of morphologically polymorphic forms 127 noun káltas dgsk. the form of a nobleman; times – the card of the participant coincides with the senior g. dgsk. adjective times vnsk. noun and noun generation dgsk. tailbone shape); homoforms of adjectives, pronouns, numerals and nouns (e.g., viena – the same form of the proper noun Viena and the numerals, pronouns and adjectives viens, -a incl. noun, pronoun, exclamation); homoforms of nouns, verbs and pronouns, numerals (e.g.: dvejos – coincides with the noun dveja vnsk. kilmininko, dgsk. two forms of the noun, the verb dvejoti in the third person future tense and the numerals; maną – corresponds to the present tense participant of the active form made from the verb many. noun, pronoun manà mot. r vr. g. vnsk. of adjective and noun mãna vnsk. tailbone shape); homomorphisms of adjectives and pronouns (e.g., černým – coincidence of the pronoun it plural noun, suffix and adjective černý vnsk. nobleman's form). 4.3.2. Homoforms with partially overlapping lexical meaning such homoforms have different lemas, but they are semantically and operatively related, e.g.: absolutely – the forms of the adverb absolutely and the adjective absolute coincide; ilgai – the forms of the adverb ilgai and the adjective ilgas coincide. These are usually adverbs made up of adjectives, i.e. words of the same root, and only partially can they be considered words of different meaning (if viewed not formally morphologically, but semantically). Partial lexical homoforms in the dictionary consist of 8 kinds of homoforms of variable parts of language, namely: substantiva mobilia nouns, for example, homoforms of administrators (concurrence of the words administratoris and administratorė dgsk. pronoun forms), bėglių (concurrence of the words bėglys and bėglė dgsk. pronoun forms); homomorphisms of nouns and verbs (e.g. apkaustai – the present tense forms of the second person of the verb and the noun forms of the noun coincide; atsakai – corresponding noun response dgsk. noun and verb answer present tense on the first person form); homoforms of nouns and adjectives (e.g., absoliutus – coincides with the absolute dgsk of the noun. absolute vnsk of noun and adjective the form of the denominator; abstraktu – coincides with noun abstract vnsk. noun and adjective abstract, -i infinitive forms); homoforms of verbs and adjectives (e.g.: magnetinę – coincides with the participant’s magnetinįs past participle dgsk. noun and adjective magnetinis, -ė mot. g. vnsk. the shape of the tail; mass – coincides with the participant’s mass of the chief g. 128 common language | 96 dgsk substantive and adjective mass, -i mot. g. vnsk. tailbone shape); homomorphisms of nouns, adjectives and verbs (e.g., kultot – the kakat of the adjective and the kálta of the participle are the same) káltas of adjective and noun vnsk. of the final form, although in the latter case two different reasons intertwine. There is indeed a partial lexical overlap between the participle and the noun, since their meaning is related, whereas with the adjective kakttas participle and the noun káltas overlap for prosodic reasons); homomorphisms of adjectives and pronouns (e.g., firstly – coincidence of the first form of an adjective, -ia of a noun and the first form of a noun, adjective, pronoun and number); homomorphisms of nouns and numerals or pronouns (e.g. aštuonetą (a noun and a numerals), maniškė (a noun and a pronouns)); homomorphisms of nouns, verbs and pronouns or adjectives, numerals (e.g. niūri – coincidence of the noun niūris vnsk. exclamation, adjective niūrus, -i mot. g. vnsk. noun and the verbs niūrėti and niurti second and third person present tense forms). Examples of partial lexical polysemanticism can also be found among the coincidences of variable and invariable parts of speech – homoforms of verbs, adjectives and adverbs (e.g., aklinai – coincidence of the verb aklinti in the past participle of the second person, of the adverb aklinai and the adjective aklinas, of the -a mot. g. in the past participle of the noun), homoforms of adjectives and adverbs (indifferently, passionately, cyclically); homomorphisms of nouns and adverbs (e.g. dovanai, laukan, pusvokalsiu); homomorphisms of nouns, adverbs and adjectives or pronouns (e.g., cowards – coincides with the noun bailys dgsk. noun, adjective cowardly, -i mot. g. vnsk. forms of the noun and adverb bailiai, although the pronunciation of the noun bailiaų and the adjective and adverb baelliai differ. if we ignore the pronunciation, all three forms would have a partly overlapping lexical meaning. Also noteworthy are the homoforms of nouns, verbs and adverbs or prepositions (e.g., uždarai – coincide with the participle form made from the verb to close, uždaras, -a, vnsk. noun closed dgsk. noun and adjective closed, -a mot. g. vnsk. the closed forms of the noun and adverb; again, the accents of some words differ, e.g. ùždarai and uždaraaan), homoforms of verbs, nouns and adverbs (e.g., apsiaustai – the forms of the noun apsiaustas dgsk. nominative, participant apsiaustas, -a mot. g. vnsk. noun and adverb apsiaustai coincide); homoforms of numerals and adverbs (e.g.: keturiomis, pirmiau), migLė ŽemRietė. 129 homoforms of verbs and adverbs (e.g.: definitely, generalizedly – the forms of the participle definitely, generalizedly, -a, generalized, -a mot. g. vnsk. with the adverbs are the same), homoforms of nouns, pronouns and adverbs (e.g., nieko – the forms of the noun and pronoun niekas are the same as the adverb nieko), homoforms of adjectives, numerals and adverbs (e.g., first of all – this word, depending on the context, can be both an adjective, an adverb, and a numeral). 4.4 Homoforms determined by the morphological system of the Lithuanian language Some homoforms are determined by the morphological system of the Lithuanian language. the coincidences of nouns are caused by syncretism of certain nouns, coincidence of some nouns of different genus or number forms (e.g., new – both mot. and senior g. dgsk. pronoun form); person, number, time, etc. of verbs may coincide; the lexical meaning of such words is the same. the whole of this category consists of mere coincidences of variable parts of language. noun coincidences mostly, they can be divided into cases of number, gender and verbal syncretism. There are 7 types of co-ordinates: (i)o co-ordinate nouns vnsk. coincidences of noun, adjective, pronoun and adjective (e.g. nouns abbey, date, emulsion; adjectives actual, fatal, powerful); ė amusing nouns vnsk. coincidences of the noun and the pronoun (e.g., nouns beteise, tēde, bedante); (i)a, (i)u conjugated nouns dgsk. coincidences of the noun and the pronoun (e.g. nouns absolventai, akiniai, angelai); (i)a verb noun vnsk. coincidences of the placeholder and the pronoun (e.g.: albume, bate, šampione); coincidences of the noun and suffix of nouns and pronouns (e.g. pronouns judu, mudu, nouns abi, dvi (it is worth mentioning that in the General Lithuanian Dictionary (bLkŽ) the word abi is considered a pronoun); i nouns or participles, adjectives dgsk. coincidences of the nominative and the pronoun (e.g. stones, ducks, white, accused); coincidences of the denominator, suffix and suffix of numerals or pronouns (e.g. numerals eighteen, twelve, thirteen and the pronoun several). there are 7 kinds of gender coincidences in the dictionary: nouns masculine and feminine g. dgsk. coincidences of adjectives (e.g., participants of doubting, accentuating, generalizing, adjectives apynaujų, laisvas, blogas – coincide with 130 common language | 96 pronouns of the pronoun form v. g. vnsk. noun and (i)o conjugation nouns mot. g. dgsk. coincidences of the final (e.g. adjectives indifferent, broken, divine; participants differentiated, forced, received); (i)o and (i)a consonant nouns vnsk. ending coincidences (e.g. adjectives indifferent, blind, lazy; participants coloured, prohibited, received); of the active species participants coincidences of noun and adjective gender (e.g.: doubtful, eternal, dressed); coincidences of the nominative mot. and the genitive g. of numerals (e.g. three); coincidences of the final part of the noun with the principal part (e.g. tris); coincidences of the genitive and/or dative of numerals or pronouns (e.g.: dviem, judviem); nouns that coincide with the general genus (e.g., giminė, mazgotė – the general genus coincides with the feminine; ligonis – the general genus coincides with the masculine); as mentioned, the number of variable parts of the language may also coincide. There are 4 kinds of such cases in the dictionary: (i)o, ė consonant nouns vnsk. nobleman and dgsk. coincidences of nominative and pronoun (e.g., nouns doubts, alphabets, batteries; adjectives bare, loose, present; participants decentralized, degenerate, lined); (i)u and i of consonant nouns vnsk. denominator and dgsk. end matches (e.g. administrator, tribe, calendar); (i)u nouns with adjectives vnsk. back and dgsk. ancestral coincidences (e.g. aggressors, harvests, wise); (i) adjectives of the genitive form voice and dgsk. coincidences in the denominator (e.g. closer, louder, smaller). the masculine and feminine g.); (i)a often overlaps with different verb forms. such coincidences in the dictionary – 7 types. it: past tense of verbs of the communicative and passive kinds of participants. homoforms of the denominator (e.g. absorb, disassemble, varnish); dgsk of the third person verbs of the present tense and the past tense of the passive form. homoforms of the progenitor (e.g. adapted, produced, heard); homomorphisms of third and second person present tense verbs (e.g.: atsispindi, įsidėmi, kvėpteli); coincidences of participants with participants or semi-participants (e.g.: adorning, fearing – in this case the regressive participants coincide with the non-regressive present tense chief g. vnsk. the form of the participant of the denominator; the given semi-participle coincides with the present tense participants – both with mot. g. dgsk. both with the chief g. vnsk. denominator); communicative coincidences with verbs vnsk. in the second person (e.g., not to mistake, slip, splash); overlap of personified and non-personified forms of verbs (e.g.: dirbtinai, imtinai, maldautinai – mot. g. participants vnsk. the forms of the beneficiary coincide with the characteristics; returned, investigated – factors MiGle Earth. reasons for the emergence of morphologically polymorphic forms; 131 minors have present tense third person forms that coincide with the participles of the noun and the verb g – both with the noun and with the pronoun); homomorphisms of present and future participles (e.g.: interrogating, extending, fertilizing – the forms of future and present participles coincide). 4.5 Homoforms of the same meaning, resulting from the assignment to different paradigms Some nouns, adjectives, verbs, pronouns and adverbs with the same meaning may belong to two different paradigms, e.g. : Audrių could be both the form Audrius and the form Audrys vnsk. rear end; I could belong to both being, is, was paradigms, both being, is, was, and being, is, was; griuvus – could belong to the paradigms of both crash, crashes, crashes, and crash, crashes, crashes; the form of the spellcaster's lema could be both spellcaster and spellcaster; my – both the form of the pronoun with the lema my and the pronoun with the lema manas. a single species of LkHZ is allocated for such cases. 4.6. Causal relationships The five mentioned reasons for the emergence may be related to each other, and in addition, the homoforms of some words may be due to more than one reason, for example, abstrãkčiai vs. abstrakčiaė – here the polysemantic character is determined both by the differences in the prosodic elements (different accents) and by the partial coincidence of meaning (one word is made of another). another example is the coincidence of the adjective kaattas and the noun káltas. these words could fall both into the categories of prosodic elements and into the categories of homoforms with different lexical meanings. Furthermore, the different homoforms that make up one kind in the dictionary are not necessarily determined by the same reasons. For example, some overlaps between the participants of the pronoun and the passive kind (e.g. adaptuotų, brandintų) are due to differences in the prosodic elements (e.g. aplėtitų and aplytė), while others are due to the morphological system of Lithuanian language, i.e. simply overlap between two verb forms (e.g. automatizuotų, derėtų). The five reasons for homoforms help to understand even better why some or other cases of morphological ambiguity arise when annoting with Lemuoklis (or other morphological annotators). Addendum