scieee AI-readable full text Open interactive document viewer

Automatinis švietimo ir mokslo terminų nustatymas lingvistiniais metodais

Erika Rimkutė

Full text

54 Erika Rimkutė | Automatic translation of educational and scientific terms | nustatymaslingvistiniaismetodais Automatic determination of educational and scientific terms by linguistic methods ERIKA RIMKUTė Vytautas the Great University KEY WORDS: textbook, educational and scientific terms, linguistic methods, heading form, grammatical form INTRODUCTION Textbook Linguistics and Computer Linguistics in Lithuania are already counting their second decade – the beginning of this branch of science can be linked to the establishment of the Computer Linguistics Center of Vytautas Magnus University (see http://tekstynas.vdu.lt) in 1994. However, the principles of textbook linguistics are still scarcely applied in terminology and terminography. The 2010 article Possibilities of automatic terminology identification in Lithuanian language (Rimkutė 2010) reviewed the extent to which textbooks are used in various terminology works in Lithuania. Lithuanian language terms can be found in several international databases, term banks; A number of electronic terminology or databases have been developed2, but all Lithuanian electronic terminology is essentially based on terminology dictionaries. Many Lithuanian dictionaries were prepared based on relatively small volumes of texts and linguistic intuition, so it can be said that the prescriptive way of obtaining and describing terms still prevails in Lithuania. The article presents the results of the project3 dedicated to the identification and description of terms in education and science, related to the automatic terminology database. The comprehensive compendium of Lithuanian resources prepared until 2011 can be found at http://sruoga.vdu.lt/lituanistiniaiskaitmeniai-istekliai/lituanistiniu-istekliu-sarasas. 2 Referred to in that article. 3 This is a project funded by the Lithuanian Science Council Automatic Identification of Educational and Scientific Terms (HUNT 2) (contract no. LIT-2-44). The main objectives of this project are: 1) to develop and test the methodology of automatic term recognition in Lithuanian language texts as a new instrument of terminology management, 2) to create a special textbook for education and science, 3) to prepare an explanatory dictionary of education and science terms and the ontology of this field on the basis of the special textbook. 55Terminology | 2012 | 19 building, their description. The idea of the project for the automatic determination of terms arose from the observation that the existing normative sources4 are rather incomplete, as far as they do not orientuoti į pačius bendriausius švietimo politikos terminus ir beveik nė cover the field of science policy, and they do not allow for new terms arising from practical work; the standardisers of terms do not have time to standardize them. However, traditional norming is quite subjective, so usage and norms must be based on textbook linguistic methods. Therefore, this article will be written about a relatively new method of selection and analysis of terms in Lithuania, which could greatly facilitate the work of terminologists and give more objectivity. AUTOMATIC SETTING OF DEADLINES In order to automatically identify and define terms in a particular field, one field has been selected – science and education. A textbook of 4 million words on education and science (hereinafter referred to as E&S) was compiled, focusing on two main themes: 1) scientific research and 2) education. From the area of research, the main focus was on research policy and higher education policy. Post-secondary education is taken from the field of education, and the main focus is on higher education and continuing education. Vocational training is also included as an additional area. The aim of the textbook was to cover texts of various genres and types, e.g.: laws, other political documents of the highest authorities (decrees of the Seimas of the Republic of Lithuania, the Government of the Republic of Lithuania, the President of the Republic of Lithuania, statements, decisions of the Constitutional Court, etc.); strategic documents; implementing acts of laws; documents of institutions and projects implementing science and education policy; documents of research and study institutions; information publications, field surveys, assessments, feasibility studies; press releases, discussion material. Having a representative, well-balanced textbook of the research field5, it is possible to automatically identify terms. For this purpose, the most commonly used are4 The Lithuanian Republic Term Bank, L. Jovaišos Encyclopedia of Educational Terms (2007), TESE Thesaurus [accessed 2012-05-20], accessed via the Internet: http://eacea.ec.europa.eu/education/ eurydice/documents/tese/pdf/teselt_005_alphabetic.pdf. 5 Representativeness and size of textbooks are some of the most difficult issues in textbook linguistics. It is believed that the textbook of HM represents the chosen field quite well. The Bononia Legal Corpus (BOLC), for example, is approximately 18 million words in size (Rossini Favretti et al. 2001). 56 Erika Rimkutė | Automatic translation of educational and scientific terms | nustatymaslingvistiniaismetodais using three-fold methods: statistical, linguistic and hybrid models of statistical and linguistic methods. Several experiments were carried out to determine the most appropriate methods of determination of terms. The following statistical methods were tested: 1) WordSmith Tools function Keyword Clusters; 2) machine learning method KEA; 3) method of extraction of collocations using LICE tool (see Grigonytė et al. 2011). After these studies, it was concluded that linguistic methods are the most suitable for Lithuanian language. The following briefly explains the peculiarities of these methods. The main advantage of statistical term determination systems is that they do not require large databases with examples of terms given by people. Moreover, they are independent of language. The basic principle of statistical systems is that words that are often used together are related and should be treated as possible terms. Linguistic methods are usually based on morphological indications, word order. Morphologically annotated texts are required to be able to apply linguistic methods. In the case of a language, the grammatical forms of the terms are usually determined by linguistic methods of terminology. When determining grammatical forms, the peculiarities of languages must be taken into account. For example, for analytical languages, it is important to determine the part of the language of terms. The linguistic part of Lithuanian terms is also very important, but it is also necessary to know what other grammatical categories of information are required for the automatic determination of terms. It has been noted that the Lithuanian language still has numerous and verbal categories, and the category of gender, when automatically determining terms, is not particularly significant6. Researchers (e.g., Bowker 2002) have found that it is almost impossible to translate term definition tools developed for other languages, especially those that use linguistic methods. For example, a typical English term structure is bdv. + dkt., dkt. + dkt., while the patterns typical of French are dkt. + bdv., dkt. + prl. + dkt. Thus, in order to apply linguistic methods of recognition of terms, it is necessary to create a program that takes into account the grammatical system of a particular language. Further in this article will be presented attempts to automatically identify Lithuanian language terms using linguistic methods. 6 The genus becomes relevant in those cases where there is a syncretism of genders, e.g., researchers: the automatic morphological analysis program recognizes the form of the researcher as both male and female. The term should be given in the masculine form, i.e. researcher. nės daiktavardžio daugiskaitos kilmininką, todėl kaip antraštinę formą nurodo ir mokslodarbuotojas, ir 57Terminologija | 2012 | 19 Note that the program used to determine the possible terms (hereinafter GT)7 was developed by the Computer Linguistics Center of VDU. This program uses linguistic rules that specify which combinations of parts of speech or word forms are to be analyzed. As mentioned, using linguistic methods, morphologically annotated texts are required, therefore, firstly, the accumulated SH texts were morphologically annotated using the morphological annotator developed by the VDU Computer Linguistics Center8. Here is an example of automatically morphologically annotated text: <word=”PROJEKTŲ” lemma=”projektas” type=”dktv vyr.gim dgsk K”>9 <space> <word=”FINANCING” lemma=”financing” type=”dktv vyr.gim vnsk K”> <space> <word=’CONDITIONS’ lemma=’condition’ type=’dktv mot.gim dgsk K’> <space> <word=”DESCRIPTION” lemma=”description” type=”dktv vyr.gim vnsk V”> <p> <word=’I’ lemma=’I’ type=’rom skaič’> <sep=”.”> <space> <word=”COMMON” lemma=”common” type=”bdvr teig nelygin.l įvardž mot.gim vnsk K”> <space> <word=”NUOSTATOS” lemma=”nuostata” type=”dktv mot.gim vnsk K”>10 <p> <number=”1”> <sep=”.”> <space> <number=”2007”> <sep=”–”> 7 Automated term identification programs using various methods usually recognize possible terms, term candidates (see Zeller 2005), although among the automatically identified words or combinations there are often completely meaningless word samples or meaningful non-term word combinations, such as collocations. From these words and compounds, the actual terms are usually determined by an expert in a particular field. 8 For more details see http://tekstynas.vdu.lt/page.xhtml?id=morphological-annotator-how-to-use. 9 Word indicates the specific form of the word used in the text, lemma – the heading form (lemma), type – detailed morphological information. 10 It should be noted that morphological annotation is done automatically, so there are also errors, for example, in this example, due to syncretism of verbs, the incorrect form is indicated: instead of the feminine plural noun, the feminine singular pronoun is indicated. For more details on the inaccuracies caused by the specifics of the morphological annotator, see the section Problems with automatic terminology. 58 Erika Rimkutė | Automatic translation of educational and scientific terms | nustatymaslingvistiniaismetodais <number=”2013”> <space> <word=”m” lemma=”m” type=”sntrmp”> <sep=”.”> <space> The operating principle of the program that automatically detects GT can be described in a simple and brief way as follows: morphologically marked text is analyzed, searching for related word chains. The first criterion is to select continuous word combinations according to the marked text, sentence differences. Some sections of text are marked as punctuation, others can be marked as HTML tags, such as the <p> paragraph tag, the <space> word space tag, etc. Furthermore, some linguistic lo terms have to be developed and taisykles, kurios būtų pritaikytos programoje. Nustatant švietimo ir moksadapted to these basic rules11: 1. Only compounds containing at least one noun are analyzed. If the GT is a two-word or longer, then the other parts of speech in such a combination can be adjectives and participles. 2. Only those compounds were analyzed, the last word of which must necessarily be a noun12. 3. In the two-word GT, proper nouns are not analyzed, because they are usually included in the names of persons, companies, institutions, etc., which are not to be considered field terms (when establishing longer terms, proper nouns are also included in the lists of GT terms). 4. Numbers written in digits from compounds. 5. Words or combinations consisting of at least one word not recognized by the morphological annotator are not analyzed. These are mostly foreign language insertions, words with spelling errors, abbreviated word forms13. After applying such rules, GTs were identified, from which the expert14 selected SM terms. These terms are analysed in more detail, their structure is described in the section Structure of terms in education and science. 11 The rules are universal and generally suitable for the automatic determination of time limits in any field. 12 There are also other possible terms structures, e.g. noun with adjective, noun with prepositional construction. There are few such terms, so they are not analyzed. Term structures not covered by this article should also be analysed in future investigations. 13 When applying certain limitations to the selection of GTs, valuable information may be lost. This study does not count the amount of unrecognized real terms. When using automatic term recognition programs, you have to accept the fact that some information will be lost, but the automatically identified information will be easier to process, the resulting compounds will be grammatically more correct. 14 The expert of this project is Lithuanianist Giedrius Viliūnas, who has extensive pedagogical, administrative and managerial experience in the field of education and science. 59Terminologija | 2012 | 19 TERMINŲ ATRANKA By applying the method described in the previous chapter, the program and the rules used in it, about 11 thousand were identified. GTs consisting of one were iki penkiolikos žodžių15. Susidūrus su tokiu dideliu duomenų kiekiu, applied several selection criteria. The first is the length of terms: taking into account the general trends in the term structure, it was decided to analyze terms not longer than seven words. The second criterion is the frequency of use of terms. It was chosen to further analyze in detail the single words GT, which are used in the text at least 20 times; the two-word GT used at least 10 times; the three words GT, used at least 8 times; four words GT, used at least 6 times; five-word GT with a frequency of at least 4 times; six-word GTs with at least 3 occurrences and seven-word GTs with at least 2 occurrences. The automatic calculation of the frequency of appearance of GT was based on the heading (dictionary) forms of words, i.e. lemas. The terms identified by applying these selection criteria were examined in more detail (the structure is presented in the next section). After the expert review of the GTs determined according to the above-mentioned criteria, the real terms of the HM were selected. Their distribution by length is given Table 1. Table 1. Distribution of terms of education and science by length Number of terms, number words Number of terms, % Number of 1155 19,87 2474 60,77 3125 16,03 422 2,82 5 4 0,51 Total 780 100,00 15 Here are some of the longest terms found: in the cases specified in the description point, the repayment of the state-supported loan to the borrower by decision of the commission appointed by the director of the fund (15 words; figures omitted here, i.e. the point mentioned); at the address of the borrower’s place of residence specified in the publicly supported loan agreement concluded by the credit institution (13 words); the payment of university student tuition fees for the academic year is determined by the rector's order (11 words). It is obvious that these fragments of text cannot be considered as possible terms. 60 Erika Rimkutė | Automatic translation of educational and scientific terms | nustatymaslingvistiniaismetodais The table shows that about 97% of all ML terms consist of terms consisting of 1-3 words. There were no educational and scientific terms found among the six-word and seven-word GTs, so they were not further analyzed. This section describes in detail the structure of the automatically identified and expert-selected terms of the ESM, and presents their grammatical models. Terms are described from the shortest, i.e. one word, to the longest, i.e. five words. 1. Single-word terms. In total, 7,635 single words in GT were identified, of which 1,311 were used more than 20 times in the text. Only these words are analysed in more detail below. Of the 1311 words, 155 can be regarded as HM terms, i.e. about 20%. The most frequently used terms in the text are: studentas (420816), universitetas (3959), studijos (1562), mokslas (1508), sritis (1403). 2. ambiguous terms. 145,468 double-word GTs were found in the 4 million-word text. As mentioned earlier, only those two-word GTs with a frequency of at least 10 times were further analyzed – 4889 such terms were found. After reviewing the GT, only 474 biword terms remained in the final list of terms of the HM and they constitute the majority of all terms analysed – more than two thirds (see Table 1). Biword terms can be divided into three types, depending on their grammatical forms. 1) Mr. K. + Mr. D. the number of terms of type HM is 250, i.e. 52.7% of all biword terms. The most common terms of this structure in the text are study programme (1432), study level (306), study direction (303), researcher (289), quality of studies (253). 2) bdv. + dkt.17 compounds comprise 200 SM terms (of which 42.2% are biword terms). The most common terms used by the Ministry of Education are higher education (2160), scientific research (918), higher education (730), distance learning (453), vocational training (403). Bdv. + dkt. type compounds are considered GT in the case where the adjective with the noun is matched by gender, number and syllable. Otherwise, gau16 The parentheses indicate the frequency of the term in the 4 million-word text of the HM. 17 The verbs are not indicated for the two-word terms models, because adjectives and participles are combined with nouns, and nouns are usually the verb of the singular noun. 61Terminology | 2012 | 19 most important law, most important learning, best of the nami gramatiškai netaisyklingi18 junginiai, pvz., glaudesnisuniversitetų, year, or the combination is regular, although often incomplete, and these are not UM terms, e.g., largest in the world, most dynamic in the world, student-friendly. 3) dlv. + dkt. compounds as PM terms are only 24, i.e. 5 %. The most common terms in the MS: final thesis (162), elective subject (109), adult19 education (65), feedback (63), applied research (45). 3. Three-word terms. A total of 119,881 three-word GTs were identified. 2438 compounds with a frequency of at least 8 times in the text were analysed in more detail. There were 125 terms of HM, i.e. 16.03% of all terms. Most common terms in the Ministry of Education: higher education institution (251), vocational education institution (205), state-supported loan (149), general education school (147), vocational education institution (77). Three-word SM terms have seven grammatical patterns. Among the three-word terms of the LS, the most common combinations are bdv. K. + dkt. K. + dkt.20 type (they make up 48% of all three-word terms), e.g.: higher education institution (251), vocational training institution (205), general education school (147). Second among the three-word terms of the HM by frequency are dkt. K. + dkt. K. + dkt. type compounds (24.8 %), e.g.: study course regulation (66), study course description (60), state scientific institute (50). The third most frequent structural model of three-word SM terms is bdv. V. + bdv. V. + dkt. (12 %), e.g. national integrated programme (69 %), fundamental research (61 %), initial vocational training (52 %). Below, according to frequency, we can give just a few examples of three-word HM terms with structural patterns: dkt. K. + dlv. V. + dkt. (4 %) e.g. publicly supported loan (149), publicly recognised qualification (33), publicly funded student (21) Dr. K. + bdv. V. + dkt. (4 %), e.g.: foreign higher education institution (41), academic council of the college (32), doctoral research supervisor (20); dlv. V. / K. + dkt. K. + dkt. (3.2 %) e.g. final assessment of qualifications (36), free study places (13), applied research (13) bdv. V. + 18 Grammatically regular compounds in this work are considered such compounds, in which the dependent words are harmonized with the main, clear syntactic relationships between the elements of the word combination. 19 Participants include the following participants. 20 The vowel of the last noun is not indicated, because it is usually a noun. 62 Erika Rimkutė | Automatic dictionary of educational and scientific terms (or | nustatymaslingvistiniaismetodais K., if the participant is a noun) + dkt. (2.4 %) e.g. competitive subject (15), non-formal adult learning (13), research (12) dlv. V. + bdv. V. + dkt. (1.6 %), e.g. applied research (20), commissioned research (8). 4. Four-word terms. A total of 77,961 four-word GTs were found, 760 of which were used more than 7 times. Of these, the expert determined the terms of 22 SMEs (they constitute 2.82% of all terms). The most common combinations in this group of terms are research and study institution (568), science and technology park (62), national comprehensive programme (40), labour market vocational training (40). The four-word terms used in the text are of nine different grammatical patterns. Because there are few four-word terms used, each model is illustrated with only a few or one example. Here are presented all the structural models of four-word SM terms, grouped by frequency, and one example illustrating each model is given: bdv. V. + dkt. K. + dkt. K. + dkt. (22.7 % of all four-word terms): international scientific database (13); bdv. V. + bdv. K. + dkt. K. + dkt. (18.2 %): general subject of university education (8); bdv. K. + dkt. K. + dlv. V. + dkt. (13.6 %): student with special needs (8) Mr. K. + Mr. K. + Mr. K. + Mr. K. (13.6 %): research and study institution (568); bdv. K. + dkt. K. + bdv. V. + dkt. (9.1 %): High-level research (18) dlv. V. + bdv. V. + dkt. K. + dkt. (9.1 %): Recognised international database (11) bdv. V. + bdv. V. + bdv. V. + dkt. (4,5 %): national integrated framework (40); Mr. K. + Mrs. K. + Mr. K. + Mr. K. (4,5 %): public research institution (31); Mr. K. + Mr. K. + Mr. K. + Mr. K. (4.5 %): field of study of technology sciences (28) 5. Five-word terms. 49 768 five-word GTs were identified, more than 5 times used compounds were found in 601, of which 4 were terms of HM, i.e. 0.51% of all terms. It is difficult to make generalizations about their structure from such a small number of terms, so the terms are not structured. All five-word terms used in the text of the Ministry of Science and Technology are: research and experimental development (124), research and technological development (67), research and technological development (13), centre of excellence for scientific research (7). 69Terminologija | 2012 | 19 aUtOMatIC IDENtIFICatION OF SCIENCE aND EDUCatION tERMS USING LINGUIStIC MEtHODS The paper deals with possibilities and problems of automatic Lithuanian term extraction. Specifically, linguistic methods are discussed in the domain of Education and Science. Other researchers have shown that it is almost impossible to have language-independent term extraction tools; this is especially true for tools which are based on linguistic rules. Therefore a linguistic term extraction tool should incorporate methods that would deal with a language’s grammatical system. This paper presents a tool developed at the Centre of Computational Linguistics of Vytautas Magnus that employs linguistic rules for extracting domain-specific terminology. University In order to extract domain-specific terms automatically, some preparatory work should be completed: compilation of domain-specific corpus (a corpus of four million words has been compiled for this research), morphological annotation of the corpus, formulation of appropriate linguistic rules, and creation of methodology for filtering out irrelevant word combinations. The paper presents the linguistic rules that have been used for the extraction of Education and Science terms and the results of the extraction procedure. The identified terms are contrasted with approved terms in the Term Bank of the Republic of Lithuania. Some specific problems of automatic term extraction are discussed in the paper, e.g. number and case agreement of terms in a multi word term. Retrieved 2012-05-17. Erika Rimkutė Vytautas the Great University K. Donelaičio g. 52, LT-44248 Kaunas, Lithuania E-mail: [email protected]