„Dabartinės rašomosios lietuvių kalbos dažninis žodynas“ ir jo bazė
Full text
ACTA LINGUISTICA LITHUANICA XLVI (2002), 19-37 Frequently Asked Questions of the Lithuanian Language LAIMAGRUMADIENE Lietuviy kalbos institutas There are few electronic databases for Lithuanian. The largest are the Frequency Dictionary of Modern Written Lithuanian, compiled by V. Zilinskiené and L. Grumadiené, and the Lithuanian Language Corpus of Kaunas University. The former is presented in detail here: the principles of its compilation and its possible uses are discussed, and the hitherto published work in statistical linguistics based on it is briefly characterised. The main problems in connection with the databases for modern Lithuanian are: (1) their insufficient differentiation (they reflect only the written, not the spoken language), (2) their not being optimally used, as the Lithuanian scholarly community lacks the preparation needed for working with such tools, and (3) tne fact that the efforts of scholars compiling them receive too little public recognition and are not valued as original scholarly achievements. 1. ELECTRONIC AGEING OF DATABASE Right now Lithuanian linguistics is as if frozen in a tiger’s bench: the outstanding works that emerged after the impulse received mainly from structuralism have already been written, but the expectation of new, no less important works is still in the air. This requires not only a different approach to language phenomena, but also new, untested methods and working tools. In this age of information technology, the new tools of linguists are often connected to electronic databases, and the latter to linguistic statistical publications. It would seem obvious that they differ from each other, because electronic databases are dynamic, with them you can work independently, and publications are static, they are just a collection of data provided at the request of the authors. However, in the past, and sometimes even now in Lithuania, such databases and publications were or are identified. This is shown by the history of reverse (also called reverse) dictionaries. The first Lithuanian backward dictionary, prepared by the American David F. Robinson (1976) on the basis of the 1954 edition of DLKZ, covered about 45,000 words, it may not even have had its own electronic database, at least the author does not mention it. In addition, at that time such an electronic database would not have been available in Lithuania, because the dictionary itself was not widely used, since it was published in a small print run in America. Juozas Korsakas (1991) from Šiauliai prepared a Lithuanian inverse
20 | LAIMA GRUMADIENĖ dictionary, covering almost 22 000 words, its basis was also DLKŽ, but already the 1972 edition. The work was done using an electronic database, but the latter is not in public use (probably for technical reasons), so it is only available in the form of a printed book. Vidos Žilinskienės Atgalinis dabartinės lietuvių kalbos žodynas (1995) is not merely a copy of the DLKŽ, because the words are presented in reverse order, divided into parts of the language, and statistical analysis is added. It is now basically available in printed form as well, although an electronic version also exists. It was used relatively little (in the work group led by E. Jakaitienė writing the Lithuanian-Norwegian dictionary, in the articles of Daiva Murmulaitytė (1998, 1999 and others)). By the way, the latter dictionary and analysis were prepared using the computerized version of the DLKŽ (1993) and its statistical analysis, and not, as Rūta Marcinkevičienė (2000a: 16) points out in her review, the database of the Frequent Dictionary of the Current Written Lithuanian Language (Grumadienė, Žilinskienė 1997; 1998). The electronic version of DLKŽ (1993 edition) prepared by the joint efforts of scientists of the Institute of Mathematics and Informatics and the Institute of Lithuanian Language had a limited circulation (1995), because there was simply no right to distribute publications in electronic form. Now the Institute of Lithuanian Language has again taken up the DLKŽ, but already in 2000 edition, computerization (manager Stasys Keinys), has received financial support from the Lithuanian Science Foundation. As you know, the 1993 and 2000 editions of the DLKŽ differ little, so it is surprising that the decision was made not to use the previous electronic version simply because it was prepared in the already morally outdated DOS environment and FOXPRO base: again it was decided to scan the already available electronic text with a computer scanner, which requires only minor corrections, thus increasing the work costs and expenses. Perhaps even the statistical analysis will be redone? However, just like the computerized version of the 1993 edition of the DLKZ, many electronic databases are morally obsolete, so it is obvious that we need to be concerned about their maintenance. Doesn’t it happen similarly as with the electronic version of the LD50 (1993) that the commonly used electronic databases are more valued by the aforementioned institutes than the expensive and innovative electronic databases? Lithuanianists should agree on such electronic databases, which use specific characters, e.g., ancient scripts, transcribed dialect texts, because already now articles with such examples have to be re-selected if for one reason or another they are not submitted to, say, the Lithuanian Language Institute- ’s publications, but to one of the Vilnius University’s publications. Currently published dialect dictionaries and dialect texts, although in electronic form, are only a sign of the modernization of the printing press, because their texts cannot be used as an electronic database. Thus, in the absence of agreement among Lithuanians on a common encoding system (e.g., UNICODE or other), the public use of electronic data banks is greatly complicated, not to mention that the use of electronic technologies and programs for their analysis is very limited and sometimes impossible. Do we have so many of them that we are unable to cover them? The problem is that neither the linguists nor the IT specialists who assist them have yet properly understood that databases must be dynamic, that they must be able to work with a multitude of new programs and technologies that are constantly emerging, and that they must be able to be converted to a higher computer version. Until now, most of the electronic databases created in Lithuania are created as if they were the same sheets with
CONTEMPORARY LANGUAGE FREQUENCY DICTIONARY AND ITS BASE |21 language data, only collected by computer keyboard. Soon, these data may be comparable to dentistry carved in stone: on the screens of new generation computers, only “scratches” and “squares” will be visible, which can be deciphered by future linguists without comparing the reading of runes. The debates held on this issue so far have been fruitless, but there is obviously no need to lose hope, and it is worth calling for debate and consultation once again. By the way, attention should be paid to the terms: electronic form BASE, BANK, TEXT (CORPUS), ARCHIVE. The most fashionable of them —TEXTILE (BODY). It is defined differently, but common to all definitions is that TEXT is considered as a collection of texts in electronic form, so often almost everything that was previously called databases has been called textbooks (because the size of individual texts is not regulated). The term BANK is usually associated with terminology: term banks are usually not called textbooks. Now the term ARCHYVO is increasingly associated with the time dimension: an electronic collection of old data can be called an archive or a database. Thus, the most neutral term remains ELECTRONIC DATA BASE, which is why it is most often used here (cf. Kennedy 1998: 3 ff.). 2. ELECTRONIC CONTEMPORARY LITHUANIAN LANGUAGES There is no excess of current databases of the Lithuanian language in electronic form. Consequently, there are not many linguistic statistical publications, only linguistic statistical methods are used a little more frequently. This is the only phonetics and phonology based on its results that are distinguished here. Perhaps this happened because of the very nature of phonetics, i.e. because the frequent repetition of the same units — sounds — simply leads to conclusions based on statistics, or perhaps because of subjective factors, i.e. because of the emergence of researchers of innovative thinking: first of all, Aleksas Girdenis and his students, as well as Antanas Pakerys. Anyway, both these factors led to a leap in Lithuanian phonetics and phonology and even a kind of gap from other branches of linguistics. True, phonetics itself is an intermediate field between exact sciences and linguistics. Another area of linguistics that follows the frequency of repetitive units is morphology. Thus, it would be logical to assume that the need to operate with both linguistic statistical data and electronic databases should already exist in this field. So far, this has not happened, although such work and working instruments already exist. Vida Žilinskienė (1990) prepared the first linguistic- -statistical work on lexis and morphology in Lithuania using ESM. Her Lithuanian Frequently Used Words Dictionary is based on the data of one of the functional styles — publicistics — and contains an analysis of the uses of 300 000 words (i.e., both the same and different lexicons, and there were 18 776 such, after refusing to study further 11 956 verified words due to the generally accepted rules for compiling frequently used words dictionaries). All continuous texts, consisting of excerpts (sample) after 1000 words of use, were morphologically analyzed, interlaced, then perforated and inserted into the now morally obsolete and no longer used ESM “Minsk-22”. There is an electronic data bank, but it would be very difficult to use it. Unfortunately, neither the lexicological, nor the morphological, nor the accentological, nor the textual, nor the lingvostatistical analysis of one of the current variants of the Lithuanian language — publicistics — has so far been properly used for scientific or practical purposes, especially for writing textbooks and educational dictionaries. One of the electronically prepared (inter-
22 | LAIMA GRUMADIENĖ national CHILD program) and the synchronic morphology studies examined could be the doctoral dissertation of Ineta Savickienė (1999) on the morphology of children’s language. By the way, it should be mentioned that the works based on linguistic statistical data on other conjugations, stylistics and partly syntax, whose units have a much lower frequency, especially those by Audronė Bitinienė (1983, 1997, 1998, etc.), appeared already several decades ago. There is already another attempt based on linguistic statistics to study the current lexicon. Recently, such works are mostly done in the Vytautas Magnus University Contemporary Lithuanian Language Textbook (about 60 000 000 words) (e.g., Rūta Marcinkevičienė (2000a; 2000b) collation studies, Jurgita Mikelionienė (2000) doctoral dissertation on neologisms, etc.). Small electronic databases are used to prepare more and more lexical works, sometimes quite specific, for example, Astos Ryklienes (2000) doctoral dissertation on the vocabulary used by Internet browsers and its production, but in general there are few such works, and students are little involved in this process (except at least Vytautas Magnus University). The group led by Antanas Pakerius at Vilnius Pedagogical University is finishing the preparation of an electronic database for the analysis of lexicon, and then perhaps there will be more student works of this kind. This article is limited to a description of the tools needed for the study of the current written Lithuanian language, and therefore does not discuss the already proliferating series of studies of ancient scripts based on the electronic database of the Lithuanian Language Institute. The emergence of its own electronic databases will undoubtedly give a new impetus to Lithuanian dialectology: the Institute of Lithuanian Language is already preparing to establish a Centre of Dialects, where dialect records would be prepared in electronic form and transcribed texts deciphered from them. The same institute has started the work on creating an electronic database of onomastics and toponymy. The electronic database of the great Lithuanian language dictionary, which is already being prepared, should become a real turning point in Lithuanian lexicography, and much is expected from the electronic database of the General Lithuanian language dictionary. Vilnius University lecturers, led by Evalda Jakaitienė, are involved in one of the European Union projects and are preparing an electronic so-called universal list of Lithuanian words, on the basis of which a new series of bilingual dictionaries should emerge. To a certain extent, it is a firm competing with the General Lithuanian Language Dictionary edition. Mainly by the efforts of the State Lithuanian Language Commission, at the urging of the European Union, electronic so-called Terms or Foreign Words portfolios have been prepared, a sort of sample lists of the latest and most frequently used terms and words in various fields, distributed by the European Union to all its members and candidates, so that languages can adapt and prepare for the rising wave of translations, as well as machine translation. The electronic database of the current spoken language, like the Polish, is prepared on the basis of radio and television speech recordings. The start of the database has been made by a group of Vilnius University researchers (led by Meilutė Ramonienė). The same university is preparing to build an electronic database of Slavisms (managed by Vytautas Kardelis), which should contain both audio recordings and samples selected and transcribed from them. As experience shows, results can only be expected after learning to work with this type of data, after acquiring the skills to use new opportunities. For this purpose, laboratories are needed where this is taught, preferably from a young age. The very nature of electronic databases presupposes publicity, they should be accessible to all those who need it. The largest Lithuanian language data collections are the Lithuanian Language Institute
The Lithuanian Language Institute has a significant role in the development of the Lithuanian language, and therefore, in the course of its modernization, it should become the largest centre for the preparation of electronic databases of the Lithuanian language, a source for their public use, and a centre for the attraction and diffusion of works and ideas of this kind — perhaps this role will be played by the Centre for Lithuanian Studies, which is being established on the basis of the Lithuanian Language Institute? Those electronic databases that go on the Internet will become even more open. By the way, the place where you use the computer, when there is an opportunity to connect to the Internet, becomes irrelevant, only what is on the Internet is important. But these are plans for the future. Until now, in Lithuania, the publication of linguistic statistical works in electronic form has not been legalized, the issue of copyright has not been harmonized, therefore such works printed on paper are still relevant. However, there is a fundamental difference between electronic databases and linguistic statistical publications, the latter being only the tip of the iceberg. Next, a specific issue is discussed here: the relationship between electronic databases of the current written Lithuanian language and linguistic statistical dictionaries. The latter cannot exist without the database, but the possibilities offered by the database are far from limited to the publication of dictionaries. 3. ELECTRONIC CONTEMPORARY WRITTEN LITHUANIAN Electronic language databases, of course, are distinguished from other collections of language data, in particular, by their electronic form and the specificity of the use of computer technologies associated with it. The emergence of electronic databases has opened up completely new possibilities for studying language, but they are not only related to the emergence of modern electronic media. The fact that different tools produce essentially similar results, in particular in terms of the frequency of grammatical forms and lexical characteristics, shows that there is considerable commonality between modern electronic databases and their much older analogues. In France, the Larousse publishing house has been using the mass dictionaries (old generation electronic calculating machines, ESM) since 1956 for the preparation of dictionaries (Rozencvejg 1986: 82). The question of how to create modern electronic language databases arose, on the one hand, because with the emergence of new technologies it was already possible to deal with a large amount of information, and then the volume of text to illustrate a single word (i.e. its context) began to increase. This opened up new possibilities to study the presumptions of the word, and later of larger language units, and thus syntax and semantics. They began to admire increasingly larger text passages, eventually moving even to complete finished texts, until the cost of acquiring the copyrights of the creators restrained this. On the other hand, the impact of Gutenberg’s invention was probably matched by the emergence of computers, and most importantly — their spread at an unprecedented speed, so that with the geometric progression of the variety of human communications over distance, both with speakers of the same language and with non-speakers, there was an urgent social order to obtain a variety of statistical information related to language, in order to automate as many communication processes as possible: language teaching, translation, recognition, information compression, transmission, etc. Lithuanian linguistics was not left out of this general process, but, of course, it lags far behind in terms of time. Throughout the 20th century, especially in the second half, the world was already preparing for a new leap. First of all, these were common language dictionaries. The first appeared in the 19th century.
24 | LAIMA GRUMADIENĖ at the end of the 19th — beginning of the 20th century. The very first one was the German one, which was prepared by the stenographer E. Kaeding (1898) using a card library of 11,000,000 words. In 1929, Vander Beke published a dictionary of the French language. H. Eaton (1934) compiled a comparative dictionary of the first thousand most common words in four languages — English, French, German and Spanish. The first common dictionary of English was prepared by E. Thorndike and L. Lorge (1944), Spanish by V. Garcia Hozas (1953). The first common dictionary of Russian language in America was prepared by H. Josselson (1953), and N. P. Vachar (1966), also in America, published a common dictionary of Russian spoken language of the Soviet period, but essentially it is prepared only on the basis of the language of dramatic works. E. Steinfeldt (Štejnfel'dt 1963) published a common dictionary of Pushkin's language in Tallinn. Based on a textbook of 1,000,000 words, made up of more than half fiction and so-called informational prose (journalistics, stationery and scientific literature), the common Russian dictionary in Uppsala was prepared by L. Lonngren (1993). The first common dictionaries not only of Russian, but also of Romanian and Italian were also prepared in America — by A. Juillard, P. M. Edwards, I. Juillard (1965), A. Juillard, V. Traversa (1973) (by the way, A. Juillard, E. Ch. Rodrigues (1964) had already prepared another common dictionary of Spanish there, and later another one was prepared by J. R. Alameda and E. Cuetos (1995). A whole series of common dictionaries was based on the concept of functional language styles of the Prague Linguistic Club, as well as on Boris Larin’s concept of common language: the first, as it should be, was the Czech language, initially of individual functional styles, and later the general one, written by a collective of authors led by Marie Těšitelová (J. Jelinek, J. V. Bečka and others, see Těšitelova 1983). The Latvians, like the Czechs, were early in preparing separate functional styles, and later a two-part general dictionary — T. Jakubaite, D. and Kristovska, D. Gulevska, V. Ozola, R. Prūse, A. Rubine, N. Sika (1966-1976- ). The same methodology was used by V. A. Agrayev, V. V. Borodin, V. M. Muratova, E. V. Tisenko, under the direction of L. N. Zasorina (1977), in preparing a common dictionary of the modern Russian language based on 1,000,000 words from four functional styles. Only one functional style — fictional literature — is represented in the frequent dictionary of the Belarusian language prepared by N. S. Mazeika and A. J. Suprun (Mažejka, Suprun 1976). The Bulgarian common dictionary (Todorova, Pančovska 2001) is also in the publicistic style. The first Lithuanian-Latvian dictionary prepared by V. Žilinskienė (1990) was also of only one functional style — journalistic, using the same Czech methodology, with a lot of cooperation with Latvians. Over time, although Lithuanian cannot boast of this, the world has seen the appearance of a variety of common dictionaries: individual authors’ language, functional styles, regional language variants, written and spoken language, etc. Frequency dictionaries of the "general language" became non-functional, because they were needed to meet specific objectives, so the selection of words had to be made from texts in some field, rather than from their totality. On the other hand, the problem of compiling common dictionaries “came to another dimension” when American behaviorists and positivists, in search of examples of usage, found electronic textbooks: the rapid development of electronic technologies made it possible not only to store, but also to process huge text masses. Textbook linguistics began in 1961, when Nelson Francis and Henry Kučera announced and began to develop the famous electronic Brown textbook of written language (Brown university), John Sinclair —the LOB (Lancaster-Oslo/- Birmingen) textbook of the same format — and John
FREQUENCY DICTIONARY OF CONTEMPORARY LANGUAGE AND ITS BASE |25 tyną; As early as 1959 Randolph Ouirk announced the SEU (Survey of English Usage) written and spoken language corpus, which was followed by Jan Svartvik, who started the LLC (London-- Lund Corpus of Spoken English) spoken language corpus at Lund University. The volume of all these textbooks at that time was about 1-1.15 million words, and more similar textbooks appeared afterwards (see Aijmer, Altenberg 1991: 1tt.). The volume of the second generation textbooks is counted in tens of millions, the third — in hundreds and thousands. The world still faces the problem of compatibility of computer software and individual work programs: there are already accumulated enormous volumes of textbooks — hundreds of millions of words (that is basically the electronic form of the library), but how to use them comfortably, how to count the most elementary frequencies, how not to get lost in the ocean of words? After all, does quantity always determine quality? In the course of compiling various frequency dictionaries, it has been verified many times that half of the researched words are used only once. Therefore, every time you start work, you need to carefully consider whether the tasks really require you to work with an abundance of material, finding the rarest words. Sometimes, when some grammatical or word formation laws are studied, it is certainly irrational to increase the size of the study population, everything should be evaluated by statistical criteria, after all, they will show the probability when the sample is sufficient to make representative conclusions. Particular attention should be paid to which texts make up the study population, i.e., how the model of the study object is formed. In Lithuanian linguistics, perhaps the most important question in this regard is what should be considered as the common Lithuanian language, in what proportions should functional language styles be represented in it, or which layers of language should be distinguished according to other criteria. For example, the Vytautas Magnus University textbook was compiled according to one selection principle, while the Contemporary Lithuanian Written Language Common Dictionary was compiled according to another, which is why the results are different. Statistical data on, for example, individual parts of speech taken from different textbooks must be evaluated very carefully and objectively. It has been noted that in the works of this kind already written or being written, insufficient attention is paid to what texts have been studied and generalizations are made about the “contemporary Lithuanian language- ”. Working with numbers requires precision, otherwise misleading conclusions are obtained. The electronic textbook of the Contemporary Written Lithuanian Language Common Dictionary (1.2 million words) is dedicated to the first generation textbooks (Brown, LOB type) — over a million words in volume. It has another common feature with these famous texts — as rarely as any later text in the world, it is annotated. Computer technology has provided exceptional conditions for analytical languages: special, but quite widespread programs (especially well-known Oxford University — Oxford Concordance Program, as well as KWIC, KAYE, CLAN, OCE, WordCruncher, etc.) allow relatively easy preparation of lists of the most common words, presenting their concordances (Leech 1991: 10; Utka 2000). Such programs are usually used by Vytautas Magnus University researchers of contemporary Lithuanian language texts, but then it is possible to analyze only words, i.e. lexicon, because knowledge about the end of the word is presented, and it is, as we know, extremely important for Lithuanian language. This makes it more difficult for synthetic languages: after all, the programs described in this article prepare lists of word forms, not heading words. It takes quite a lot of additional work before word forms are subconjugated and become merely title words. Such lists of words, having a grammatical examination, are already called annotated, in the world there are not many of them, they are very appreciated, because they require considerable labor costs, therefore they are
26 | LAIMA GRUMADIENĖ are very expensive (V. Žilinskienė and I have calculated that in our case just the selection of texts, compilation on a computer keyboard, semi-automated morphological analysis (so-called LEMAVIMAS), correction and data insertion into a computer would require six years of work from one person). Calculations and subsequent corrections again take several years, because any inaccuracy in performing mathematical operations can turn into absurdity. In general, common dictionaries, especially with the advent of electronic textbooks, can now be created in a variety of ways — depending on what you are looking for. The most important reason why new common dictionaries are being prepared, especially those pretending to be dictionaries of the current n language, is that the basic condition of linguistic statistics, as well as of discrete statistics, is that the sample must be a representative model of the studying totality, or population. In short, there is constant doubt whether the sample really represents the intended totality. For example, the frequency dictionary of the contemporary Lithuanian language prepared by Vytautas Magnus University, which consists of as many as 60 million words, would be different from those already published by L. Grumadienė and V. Žilinskienė (1997, 1998), because its results would reflect the totality of texts, which are based on 70% periodical, 2696 non-periodical and 4% translated publications (Marcinkevičienė 2000a: 16). The texts of the published dictionaries have been selected on a completely different basis: they are original, non-translational texts of four functional language styles — journalistic, clerical, fictional and scientific. This was done because the common Lithuanian language, especially the written language, is still being formed, regulated and codified in the spirit of the ideas of the Prague Linguistic Club (by the way, it should be noted that the term STYLE itself is now defined quite differently in various schools of linguistics). The principles of other common dictionaries, especially Czech, Russian and Latvian, who also followed the previously described general language — the sum of functional language styles — were followed, their experience was applied, and therefore, like them, 300,000 words were selected from four functional styles, because it was believed that the principles of the Vytautas Didysis University textbook (which, by the way, appeared later than ours) were still too early for the written Lithuanian language, closer to the American concept of general (or standard, as it is usual there). The database of the current common dictionary of the written Lithuanian language sometimes lacks the domestic vocabulary, because it is forgotten that it was formed only on the basis of the public written language: it does not contain either private letters, diaries, wall notes or the like. However, even now the dictionary contains words from the spoken language, including not only foreign words and dialects, but also examples of pejorative vocabulary that are intolerable in common language, for which the authors of the dictionary are constantly rebuked. The question arises, what to consider fiction: these words were in fictional literature texts (original texts printed in Lithuania in more than 100 copies), and the postulate of creating a common dictionary is not to throw out any meaningful word, unless it is a true word, a number, an abbreviation or a proofreading error. By the way, the data show that after selecting samples of 300,000 words for each of the four functional styles, 284,687 meaningful words were found in the journalistic style (15,313 so-called “junk” words), 279,505 in the subjective style (20,495 of which were not included in the dictionary for the aforementioned reasons), 278,311 in the scientific style (21,689), and 290,561 in the fictional style (9,439). The heart did not allow to throw out the true word Lithuania, so it as an exception only true entered the dictionary (it turns out, 16th by frequency word, the first of the nouns, its frequency 5151 even 614 samples: here determined official documents,
27 names of institutions, etc., in which this word was frequently used in the noun vowel). The names of religious and other holidays (often used as generics), for example, Joninės, Kūčios, Žolinė, or the names of the countries of the world Rytai, Vakarai, Pietūs, Šiaurė, can also be considered as conditionally true. In accordance with the principle of representativeness, only those texts of each of the above-mentioned styles were selected that are appropriate for pure style: textbook texts, scientific publications, publications for a narrow circle of readers, drama, poetry, etc., which are considered to be paraibiniai, were abandoned. The texts were selected by Vida Žilinskienė from 19 central daily or weekly newspapers and widely read magazines in 1994-1995, representing five thematic groups: 1) politics, 2) industry and agriculture, 3) education and law, 4) art and 5) entertainment (sports, horoscopes, weather forecasts). Subjective style texts were selected by Pranas Kniūkšta and Danutė Vainauskienė from 1990-1995 publications and documents according to seven thematic groups: 1) laws and accompanying documents, commentaries, 2) legal documents, 3) economic and business references and documents, 4) political science publications, party documents, 5) standards and statistics publications, 6) information sources (bibliographies, catalogues, dictionaries), 7) individual social fields (national defence, social security, education, culture, etc. documents). Texts in scientific style were selected by Irena Andriukaitiene from scientific publications published in Lithuania in 1990-1995 or from works approved by scientific institutions according to eight thematic groups (direction of science): 1) agricultural sciences, 2) natural sciences, 3) humanities, 4) mathematics, 5) medicine, 6) social sciences, 7) technology, 8) theology. The fiction style texts were selected by Laima Grumadienė from prose works of fiction published in Lithuania in 1985-1990 (with small exceptions), 0 of which six “thematic” groups are divided according to the writers’ homelands: 1) Žemaičiai, 2) Eastern Highlanders, 3) Western Highlanders, 4) Southern Highlanders, 5) Urbanites, 6) Expatriates (the real homeland of the latter was not taken into account). It was decided that these texts could represent the present written Lithuanian language. From these, one section of 1000 words, the so-called sample, was selected in random order (pages and lines were selected according to a table of random numbers) and the computer keyboard was assembled (about two hours of work). After that, the samples were determined, i.e. morphologically analyzed by Vytautas Zinkevičius prepared MAN (Morphological Analysis and Normalization) program. The author adapted it specifically for this work from his own computer program for correcting the spelling of words, which he had prepared earlier. In 2000, on behalf of Vytautas Magnus University, he slightly improved it, moved it to a higher version and named it lemuoklis (Zinkevičius 2000). The work looked like this: you get a computer spreadsheet, from which you have to “sweep out” unnecesaš sary (underlined by me) lines. For example, text passages already processed by humans, not only by machines (...) for me this whole comedy (...) looks like this: 1 PAV. EXAMPLES OF TEXT DETERMINATION think me invr vns K think me vks dlv veik bk not inv vyr vns V
34 | The 34th Rank Frequency dkt vks bdv prv ívr 301-350 460-374 19646 10424 2420 2073 393 54,28% 28,80% 6,69% 5,73% 1,09% 351-400 373-313 18504 7550 3381 1347 654 53,04% 21,64% 9,69% 3,86% 1,87% 401-450 312-260 19156 7865 3976 522 304 57,20% 23,49% 11,87% 1,56% 0,91% 451-500 259-205 21718 9987 4806 2486 471 53,17% 24,45% 11,76% 6,08% 1,15% The minimum dictionary of the I-th qualification category of the State Lithuanian language (about 800 words) LITHUANIAN LANGUAGE COMPETENCY II QUALIFICATION LANGUAGE PARTS AND GENERAL WORDS FREQUENCY Rank Frequency dkt vks bdv prv jvr 501-550 204-155 22494 13153 7023 2481 525 46,97% 27,46% 14,66% 5,18% 1,10% 551-600 154-106 30834 17197 6221 4419 151 50,91% 28,39% 10,27% 7,29% 0,25% The minimum dictionary of the State Lithuanian Language II-th qualification category (about 1200 words) 3 TABLE: VALST. LITHUANIAN LANGUAGE COMPETENCY III QUALIFICATION LANGUAGE PARTS AND GENERAL WORDS FREQUENCY Rank Frequency dkt vks bdv prv pvr 601-650 105-56 45751 20328 12145 5822 433 52,31% 23,25% 13,88% 6,66% 0,50% The minimum dictionary of the third qualification category of the Lithuanian language (about 2500 words)
FREQUENT DICTIONARY OF CONTEMPORARY LANGUAGE IR JO BASE | 35 skt prl jng dll ist; jst Sum 426 - & 812 — 36194 1,1796 2,2496 3,02% - — 1407 2043 — 34886 4,03% 5,87% 291% 276 — 272 1118 — 33489 0,82% 0,81% 3,34% 2,79% 255 224 210 700 3 40857 0,6296 0,55% 0,51% 1,71% 3,40% words) overlaps 56,66% the text (sample iš 1,2 min. word uses). 56,66% CATEGORIES MINIMUM ZODYNO LINGUISTICAL STATISTICS ANALYZE: skt prl jng dll ist; jst Sum 530 342 177 795 371 47891 1,11% 0,71% 0,37% 1,66% 0,78% 3,99% 520 275 132 713 109 60571 0,86% 0,45% 0,22% 1,18% 0,18% 5,05% words) overlaps 65,70% the text (sample iš 1,2 2010-2011: 1.0 million word uses). 9,0496 Iš face: (=65,7096) LINGUISTICAL ANALYSIS OF THE MINIMUM VOCABULARY OF CATEGORY: skt prl jng dll ist; jst Sum 460 341 372 1697 102 87446 0,52% 0,40% 0,42% 1,94% 0,12% 7,29% words) overlaps 78,04% the text (sample iš 1,2 2010-2011: 1.0 million word uses). 7,2996 Iš face: (=78,04%)
AUMER, K., ALTENBERG, B. 1991: Introduction to the theory of the spatial space. Aijmer, K., Altenberg, B., eds., English Corpus Linguistics. Studies in Honour of Jan Svartvik, London-New York: Longman. ALAMEDA, J. R., CuEros, E 1995: Dictionary of frequencies of the linguistic units of Spanish, University of Oviedo. BITINIENE, A. 1983: Scientific style, Vilnius: Ministry of Higher and Special Education. BITINIENĖ, A. 1997: Functional Styles: Sentence Length and Structure, Vilnius: Vilnius Pedagogical University Press. BITINIENĖ, A. 1998: Administrative style and its sentence length. Kalbotyra 47 (1), 17-28. EATON, H. 1934: Comparative freguency list on the first thousand words in English, French, German and Spanish. Coleman, A., ed., Experiments and Studies in Modern Language Teaching, Chicago. Garcia Hoz, V. 1953: Vocabulario usual, comun y fundamental. Madrid: Consejo Superior de Investigaciones Científicas —Instituto „San José de Calasanz- ®. GRUMADIENE, L., ZILINSKIENE, V. 1997: The frequency order of the present written Lithuanian language, Vilnius: Institute of Mathematics and Informatics, Institute of Lithuanian Language. GRUMADIENĖ, L., ŽILINSKIENĖ, V. 1998: Frequently Asked Questions in Lithuanian (Alphabetical Order), Vilnius: Institute of Mathematics and Informatics, Institute of Lithuanian Language. HiLLERICH, R., www: eduplace. com: High-frequency words and vocabulary. JAKUBAITE, T. e.a. 1966-1976: Dictionary of the Latvian language, 1.4. edition, Riga: Zinatne. JosseLsoN, H. 1953: The Russian Word Count and Frequency Analysis of Grammatical Categories of Standard Literary Russian, Detroit: Periodical Service Co. JuiLLAND, A., RODRIGUES, E. Ch. 1964: Frequency Dictionary of Spanish words, The Hague: Mouton. JUILLAND, A., EDWARDS, P. M., JUILLAND, I. 1965: Frequency Dictionary of Roumanian Words, The Hague: Mouton. JUILLAND, A., TRAVERSA, V. 1973: Frequency Dictionary of Italian Words, The Hague: Mouton. KAEDING, E. 1898: Frequency Dictionary of the German Language, Steigliz bei Berlin (Selbstverlag). KEINYS, S., ed. 1993: Dictionary of Contemporary Lithuanian Language, Vilnius: Mokslo ir enciklopedijų leidykla. KEINYS, S., ed. 2000: Dictionary of Contemporary Lithuanian Language, Vilnius: Mokslo ir enciklopedijų leidykla. KENNEDY, G. 1998: An Introduction to Corpus Linguistics, London-New York: Longman. KORSAKAS, J. 1991: Inverse Dictionary of Lithuanian Language, Kaunas: Šviesa. KRUOPAS, J., ed. 1954: Dictionary of the Contemporary Lithuanian Language, Vilnius: State Political and Scientific Literature Publishing House. KRUOPAS, J., ed., 1972: Dictionary of Contemporary Lithuanian Language, Vilnius: Mintis. LEECH, G. 1991: The state of the art in corpus linguistics, Aijmer, K., Altenberg, B., eds., English Corpus Linguistics. Studies in Honour of Jan Svartvik, London-New York: Longman. LONNGREN, L. 1993: Yacmomnoiū crosape cospemennozo pycckozo aswika. Uppsala. (Acta Universitatis Upsaliensis, Studia Slavica Upsaliensia 32.) MARCINKEVICIENE, R. 2000a: Textual Linguistics (Theory and Practice). Works and Days 24, 7-64. MARCINKEVICIENE, R. 2000b: Patterns of word usage in corpus linguistics. Kalbotyra 49 (3), 71-80. MAZEJKA, N. S. (Max3iika, H. C.), SUPRUN, A. Ja. 1976: Yacmomnoi: cnoynix Genapyckaii Moet (macmaykas nposza), Minck: Beinasenrsa BJ1V ims V. Jlenina. MIKELIONIENE, J. 2000: New Lithuanian language lexicon (based on the 1991-1996 computerized periodical textbook- ). PhD dissertation, Kaunas: Vytautas Magnus University. MURMULAITYTĖ, D. 1998: Notation of word connections in the “Dictionary of Contemporary Lithuanian Language” (see abbreviations). Journal of Linguistics 39, 240-247. MURMULAITYTĖ, D. 1999: On the Computerization of the Common Lithuanian Language Dictionary. Lituanistica 2 (38), 92-98. RYKLIENĖ, A. 2000: Communication on the Internet: speaking while writing. Works and Days 24, 99-107.
ROBINSON, D. E. 1976: Lithuanian Reverse Dictionary, Michigan: Slavica Publishers Inc. ROZENCVEIG, V. Ju. 1976: Lithuanian Reverse Dictionary, Michigan: Slavica Publishers Inc. 1986: Onsir co3nanus HaUHOHANBLHBIX JIEKCHKOTpaOHUeCKHX CJIyxO 3a pybexom. I'm dead. H., pea., Mauunnor pono pycckozo asvika: udeu u cymcoenus, Mocksa: Hayka, 75-83. SAVICKIENE, I. 1999: Morphology of the Lithuanian child noun. PhD dissertation, Kaunas: Vytautas Magnus University. 1994: How to learn languages more rationally (Basic English), Vilnius: Žodynas. ŠTEJNFEL'- DT, E. A. 1963: Yacmomnbtū crosaps COBPEMEHHO20 PYCCKO20 iumepamypno2o A3bIKA, Tannun. TESITELOVA, M. e.a. 1983: Frekvenéni slovník češtiny věcného stylu, Prague: Academia- -Nakladatelstvi Československé Akademie Věd. THORNDIKE, E., LORGE, L. 1944: A Teacher's Word Book of the Twenty Thousand Words Found Most Frequently and Widely in General Reading for Children and Young People, New York: Appleton— Century Crofts. TopoRova, E., PANCOVSKA, R. 2001: Yecmomen peunux na 6wrzapckama nybauyucmuxa (1944— 1989), Codus: Ilencodr. UTkA, A. 2000: Language equipment and its possibilities. Works and Days 24, 275-285. VACAR, N. P. 1966: A Word Count of Spoken Russian. The Soviet Usage, Ohio: Ohio State University Press. VANDER BEKE, G. 1929: French Word Book, New York: Macmillan. ZASORINA, L. N., ed., 1977: Yacmomuwiit caoeape pycckozo aswika, Mocksa: UsnatensctBo «PycckHH A3BIKY. ZINKEVICIUS, V. 2000: Lemuoklis — morphological analysis. Works and Days 24, 246-273. ZILINSKIENE, V. 1990: Lithuanian Language Frequency Dictionary, Vilnius: Mokslas. 1995: A retrospective dictionary of contemporary Lithuanian language, Vilnius: Institute of Mathematics and Informatics. ŽILINSKIENĖ, V. 1998: Frequently Asked Questions in Lithuanian (first results of the research). Journal of Linguistics 39, 219-227. 1998: Lithuanian- -Russian minimum dictionary of the state language proficiency qualification categories (first qualification category), Kaunas: Šviesa. 1998: Lithuanian- -Russian minimum dictionary of the state language proficiency qualification categories (first qualification category), Kaunas: Šviesa. 1999a: Minimum Dictionary of Lithuanian- -Russian Languages of the State Language Proficiency Qualification Categories (Second Qualification Category), Vilnius: UAB „Nacionalinių tyrimų centro“ leidykla. 1999b: Minimum dictionary of Lithuanian-Russian languages (third qualification category) of the State language proficiency qualification categories, Vilnius: UAB "National Research Center" publishing house. ŽILINSKIENĖ, V. 2001: Development of the lexicon and morphology of Lithuanian and Latvian journalism from a statistical point of view. Linguistica Lettica 9, 155-168. ZILINSKIENE, V. Press. : Statistical characteristics of the morphology of the subjective style of Lithuanian language and their comparison with the corresponding characteristics of publicistics. Lituanistica. Laima Grumadiené Received 19 Apr 2002 Lithuanian Language Institute P. Vileišio g. 5, LT-2055 Vilnius, Lithuania laigruma(dtakas.lt
