Content interoperability as a prerequisite for re-using and re-purposing items of structured content as learning objects in eLearning – Seen under a standardisation perspective
Full text
6Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning Content interoperability as a prerequisite for re-using and re-purposing items of structured content as learning objects in eLearning – Seen under a standardisation perspective CHRISTIAN GALINSKI, BLANCA STELLA GIRALDO PÉREZ Infoterm KEYWORDS: structured content and unstructured content, verbal and non-verbal representations, appellations, morphology and morphemes, phraseology and phrasemes, collocations and multi-word terms, high complexity learning objects (HC-LO) and low complexity learning objects (LC-LO), language for general purposes (LGP) and language for special purposes (LSP), persons with disabilities (PwD) and assistive technologies, standards and standardization, certification, data models and data modelling, federated repositories, quality and interoperability requirements 1. STRUCTURED CONTENT “Content” here is seen as structured content at the level of lexical semantics comprising linguistic and non-linguistic representations of concepts (understood in science theory as a kind of “immaterial objects”). These representations can be designative (such as designations in terminology: comprising terms, symbols and appellations) or descriptive (such as various kinds of definitions, explanations or non-verbal representations), or hybrid. So far non-verbal designations and representations of concepts as well as appellations (i.e. proper names representing individual concepts) have been underrepresented in terminology theory and methodology. But they can be very important for designing learning objects (LO) in certain domains or fields of application – not to mention for language learning and translation. Equally important for designing multilingual, multimodal and multipurpose LO are a) similarities of LGP and LSP entries: - The designative representation of the entry (LGP lemma or LSP term) may be one word, a compound word or a multiword entity; - The verbal designative representation may be supplemented by a non-verbal designative representation (e.g. gesture, mimics, Blis-
7Terminologija | 2012 | 19 symbolics1 etc.), or – if required – even be replaced by it (e.g. graphical or other non-verbal symbol); - Each entry representing one concept has a unique entry identifier (which may for instance point to occurrences in text corpora); ° In LSP entries each designative representation should have its own item identifier, which among others would make it possible to trace conceptual change over time, ° In LGP entries each meaning of a designative representation should have its own item identifier, which makes multilingual entries possible, ° In analogy to termautonomy in terminology, representationautonomy applies to LSP and LGP as well as to non-verbal representations, ° Descriptive representations LGP (including explanations, contexts etc.) can be treated in analogy to those in LSP (explanations, contexts etc.), but exclude definitions in the strict sense, ° There may be additional fields for notes, examples, sources etc. - A verbal designative representation representing both a LGP as well as a LSP concept may have – depending on the degree of quasi-synonymy – to be covered by two entries, with proper cross-references. b) pronunciations of verbal items of entities of structured content, which should be foreseen in any entry at least potentially. c) non-verbal designative representations and non-verbal descriptive representations, which should be foreseen in any entry at least potentially, because they may be – depending on the domain or field of application – equally important to verbal ones and sometimes even preferred representations. d) components of entities of structured content, such as the morphemes of verbal designative representations or elements of a definition, explanation, context etc. They should be marked / tagged in order to facilitate cross-referencing or re-using as another LO (for instance prefixes and suffixes in medical nomenclature). e) larger entities of structured content, such as idioms, LGP collocations or LSP phrasemes, as well as metaphors, which can be formally treated similar to verbal designative representations in LGP (with proper cross-references to the relevant elements of the respective entity). At this 1 Blissymbolics is an ideographic writing system for cognitively impaired persons. Each of the several hundred basic symbols represents a concept, which can be composed together to generate new symbols that represent new concepts. Blissymbols do not correspond at all to the sounds of any spoken language.
8Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning level, the further refined differentiation of the above is necessary in corpus linguistics, but probably not for LO. It may also be of minor importance, whether the meaning is non-compositional or compositional. There are tools in corpus linguistics to identify and extract also such larger entities of structured content. f) the provision of placeholder fields or links to the respective designative representations for PwD, such as in Braille, sign language or Blissymbolics. Structured content resources at the level of lexical semantics so far were seen as mainly comprising lexicographical data, terminological data and other kinds of concept representations, including a few non-verbal ones, such as visual symbols (e.g. public symbols). But there may be also acoustic / audible symbols, haptic / tactile symbols, and others, which, in terminology management could occur as designations or even concept descriptions (such as non-verbal representations (ISO 10241-1:2011)). There may be further information – and the respective data categories – required for a systematic approach in managing structured content at large. In the light of the above, we can differentiate the kinds of structured content at the level of lexical semantics as follows: - Lexicographical data, such as: ° word entities (including compounds etc.), ° morphemes (morphology), ° collocations, metaphors; - Terminologies and similar kinds of language resources and other content resources, such as: ° nomenclatures, taxonomies, typologies, glossaries, vocabularies etc., ° terminological morphemes (morphology), ° terminological phrasemes (phraseology), ° proper names of all sorts as used for instance as items in different kinds of directories, ° graphical symbols and other non-verbal designations, ° (product) properties, characteristics, attributes etc.; - Thesauri, classification schemes (ISO/DIS 22274:2011), keywords and other kinds of documentation languages (or controlled vocabularies); - Encyclopaedic (knowledge) entries, covering among others: ° knowledge-enriched terminological entries,
9Terminologija | 2012 | 19 ° (explained) proper names and other kinds of data closely related to proper names; - Ontologies, topic maps and other kinds of knowledge-structuring systems. Like in terminology, any of the other kinds of structured content listed above may have - (one or more) phonetical representations, sign language representations etc.; - graphical and other non-verbal designative representations as well as non-verbal descriptive representations (some in addition to a verbal representation, others created as non-verbal representations independent from verbal ones). Non-verbal kinds of structured content are particularly meaningful in applications like eLearning and of vital importance in the communication: - with and among PwD (directly or through ICT devices functioning as assistivetechnologies), - between PwD and the devices they use, and - among these devices. Traffic sign designers are calling the elements of traffic signs “morphology” (in analogy to morphemes in linguistics). The same could be applied to other non-verbal designative representations. There have already been developed search methods to look for whole pictures / graphical representations by means of individual elements contained in them. The following figure shows an example for the “morphology” of graphical symbols (in this case a traffic sign): (from Schmitz’ Presentation at TSTT 20062; see Schmitz 2006) 2 3rd International Conference on Terminology, Standardization and Technology Transfer (TSTT 2006), Beijing, China, 25-26 August 2006
10 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning In certain domains (e.g. in biological nomenclatures) there are even examples for non-verbal descriptive representations replacing lengthy definitions or other concept descriptions, which could be useful also for eLearning purposes: (from Gonnissen and Mornie 1983, bird no 50) It only needs a comparatively short legend with explanations for the symbols (in several languages) and the graphical representation is easily understood by experts, students or hobby ornithologists alike. All of the above kinds of structured content are becoming more and more digitally accessible today (as eContent – increasingly also through mobile devices) and may - occur in digital texts (such as in technical documentation, scientific-technical writing etc.), - be combined with each other or embedded in each other, - have similar linguistic elements (letters, sounds, morphemes etc.) or different ones (such as non-verbal representations), - form complex content items, - need to be integrated or made interoperable with others in certain applications. At present, however, most of the existing repositories of structured content are not consistent within a repository and sometimes even contradictory between different repositories. Mostly, they are not based on proper metadata and data modelling methods, and therefore not integrable, not reliable and full of deficiencies. This is inacceptable for instance in applications, which support PwD, particularly in our aging societies,
11Terminologija | 2012 | 19 where more and more people have – sometimes multiple – impairments. It further does not allow their efficient use for eLearning purposes – which again is disadvantaging PwD. In contrast to structured content, a corpus can be described as unstructured content, namely “a body of naturally occurring language” (McEnery, Xiao, Tono 2006), thereby distinguishing a corpus from word lists, dictionaries, databases. Similarly a speech corpus (or spoken corpus) is a repository of speech audio files (and text transcriptions). These definitions are insufficient when it comes to non-linguistic elements in texts (such as in technical documentation) and with respect to elements of nonverbal communication (such as gestures, mimics etc.) necessarily accompanying spoken data in real life. If we look, where the items of structured content can be found, we recognize that most of them are not developed as a goal in itself, but are necessary elements of non-structured content, such as text corpora, speech corpora, etc. Therefore, the relation between structured content and corpora – especially with a view to making structured content occurring in non-structured content productive for instance for eLearning in the form of LO – should be further investigated. Today, the biggest corpus in the world is the Internet. However, in most cases it needs texts of assured quality for extracting LO – which is a challenge for classifying and quality rating of texts and tagging them appropriately. 2. STANDARDIZATION ISSUES REFERRING TO STRUCTURED CONTENT Standards documents today do not only comprise technical standards in the traditional sense, but also standards for terminology, testing, products, processes, services, interfaces, data etc. Some are basic standards that have a wide-ranging coverage or contain general provisions for one particular field, others are methodology standards. The standardized terminology of a given field can be considered as basic standard. The methodology standards concerning the principles and methods of terminology work and terminology standardization consequently are basic for all terminology work and terminology standardization / unification in standardization at large and beyond. Furthermore, the types of standards mentioned above only refer to some common types, and they are not mutually exclusive; for instance, a particular productstandard may also be
12 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning regarded as a testingstandard if it provides testmethods for characteristics of the product in question. As content today is necessary for nearly everything in science and technology, content related standards can fall under any of the above-mentioned types of standards. According to ISO/IEC Guide 2:2001 standardization is an activity for establishing, with regard to actualorpotentialproblems, provisions for commonandrepeateduse, aimed at the achievementoftheoptimum degreeoforderinagivencontext. In particular, the activity consists of the processes of formulating, issuing and implementing standards. Important benefits of standardization are improvement of the suitabilityofproducts, processesandservicesfortheirintendedpurposes, preventionofbarriers totrade and facilitationoftechnologicalcooperation. The preparation of standards is based on consensus, which is a generalagreement, characterized by the absenceofsustainedoppositiontosubstantialissues by any important part of the concernedinterests and by a processthatinvolves seekingtotakeintoaccounttheviewsofallparties (namely industry, research, public administration, consumers) concerned and to reconcile anyconflictingarguments. Standardization endeavours are governed by highly systemic approaches. In particular, methodology standards are aiming at generic solutions, which are appropriate in different applications, too. This also applies to methodology standards concerning structured content, such as standards for transcription and transliteration, data models, interchange formats, source identifiers (ISO 12615:2004) not to mention project management. It also applies to some kinds of standardized structured content itself, such as languages codes (ISO 639 series), script codes (ISO 15924:2004) and country codes (ISO 3166 series). In recent years aspects of interoperability of structured content have become a major issue in many fields. Therefore, standards concerning data quality, data administration, content management and workflows are increasingly becoming imperative. There are several technical committees dealing with various – more or less generic – aspects of content interoperability, for instance: - ISO/TC 37 Terminologyandotherlanguageandcontentresources - ISO/IEC-JTC 1/SC 32 Datamanagementandinterchange (especially its WG 2 MetaData); - ISO/IEC-JTC 1/SC 36 Informationtechnologyforlearning,education andtraining;
13Terminologija | 2012 | 19 - ISO/TC 184 Automationsystemsandintegration (especially its SC 4 Industrialdata); - ISO/TC 46 Informationanddocumentation. The standards developed by these committees do not yet take into account the specific requirements of eLearning with respect to content interoperability of entities of structured content at the level of lexical semantics. ISO/TC 37 was the first committee to take multilinguality fully into account (in fact as one of the basic principles of all its standardization efforts, which is concept-oriented, i.e. language-independent = multilingual from the outset) – other technical committees have followed suit, but often they are not sufficiently respecting this principle in practice. Today even lexicography has taken a turn towards concept-orientation (=multilinguality), as can be seen from the products of several dictionary publishers and large-scale online dictionaries accessible on the Internet. This should also apply to learning objects (LO) at the level of lexical semantics. Given the fact that content interoperability can only be achieved on the basis of commonly accepted rules, viz. international standards, there is a need for cooperation and coordination of the most important standardizing activities with respect to content interoperability. New standardizing activities for various aspects of content interoperability - are driven by technical developments in the direction of web-based cooperative / participatory content creation through ICT (in the form of social networks, cooperation platforms, mobile communication etc.) on the one hand; - will have a big impact on the various kinds of content and knowledge management (especially in eApplications, such as eLearning, eAccessibility&eInclusion, eHealth, multilingual product data management in eBusiness etc.) on the other hand. So far the standardization efforts in ISO/TC 37 Terminologyandother languageandcontentresources with respect to structured content focused on methodology standards related to - Data categories (not quite identical with metadata) used in the conceptual design of the entries of structured content; - Data models and data modelling methods;
14 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning - Metamodels to make competing data models interoperable; - Applications of the above; and to some extent on standardizing those kinds of content, which are of relevance to the TC. If ontologies in the meaning of knowledge representation are included, the above-mentioned metamodels need to be extended towards metaontologies and even a meta ontology language (ISO/CD 17347:2012), in order to provide the possibility to make ontologies interoperable. In this connection an ontology is seen as a “formal, explicit specification of a shared conceptualisation” (Gruber, June 1993), which represents a shared vocabulary and taxonomy that models a domain – that is, the definition of concepts and other information objects, as well as their properties and relations. An ontology language – different from mere knowledge representation – provides a metamodel for such formal, explicit specifications of a shared conceptualisation. Nearly all TCs in ISO and IEC standardize the most important terms and definitions of their respective domain or subject. Some also standardize other kinds of language and content resources. 3. VOLUMES OF STRUCTURED CONTENT: TERMINOLOGY AND OTHER LANGUAGE AND CONTENT RESOURCES According to ELRA (European Language Resource Association) language resources are: - text corpora, - speech corpora, - (lexicographical data and) terminologies. In a recent article (Galinski, Reineke 2011), an attempt was made to quantify the volumes of lexicographical and terminological entries in an ever increasing number of domains or subjects. The lexis of LGP in the highly developed languages may comprise up to 500,000 lexemes (including a considerable share of terminology). However, the total number of scientifictechnical concepts across all domains or subjects may well comprise ~100– 150 million. The number of identifiable chemical substances alone has passed the 60 million mark in 2011. In the light of these figures - Content interoperability is a prerequisite for avoiding a huge duplication of efforts; - The ISO/TC 37 approach is becoming more and more essential;
21Terminologija | 2012 | 19 In the figure above, the onyomi kyû (as L1 for K1 休) may carry any meaning of the kunyomi readings, but kyû can only be used for the kanji, if it occurs in kanji-combinations (or – which is an exception – if the kanji stands alone as an abbreviation). Kyûas LO would only make sense, if a learner wants to learn all kanji having the onyomi kyû for some reason or other. There are many kanji, where the onyomi represents one or more concepts and, therefore, each kanji+onyomi can be taken as a LO. l2 休む yasumu, L3 休まるyasumaru, L4 休めるyasumeru, L5 休み yasumi represent different – though semantically related – Japanese readings kunyomi of the kanji in question, which beyond their derived nature have an idiomatic meaning qualifying each of them as individual LC-LO of their own. However, if a LO of this kind has more than one meaning, any of these represents in principle a LC-LO. 休 in kunyomi readings only stands for the stem yasu. However, there are other stems read yasu with a different meaning written by different kanji, such as 安 in yasui 安い, which could lead the learner to other kanji LO having the Japanese reading yasu. l6 定休日 teikyûbi, L7 一休みする hitoyasumisuru and L8 休日 kyûjitsu represent the meaning of compounds of kanji (binoms, trinoms etc. having or not having endings) with onyomi, kunyomi or mixed reading. Thus, the core element of the HC-LO K1 (休 kyû) of this kanji flashcard can - be combined with two or more other kanji to form lexicographical LO, such as L6 teikyûbi, L7 hitoyasumisuru and L8 kyûjitsu. If any of them has more than one meaning, they would have to be taken as two or more LOs; - point to other kanji HC-O, such as K2 一, K3 定, K4 日; - be linked to look-alike kanji, such as K5 体 and K6 伏 or other kanji of historic or other relevance. Any LO of the kinds mentioned above should have foreseen a placeholder slot for: - pronunciation (or pronunciations, because there may be two or more; if there are homophones, pointers should lead to the respective LO); - sign language and other kinds of AAC (alternative and augmented communication) means. The highly complicated Japanese language and script has been chosen to illustrate the method of composing and decomposing LO at the level of lexical semantics. Taking Japanese, a complicated language with a com-
22 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning plicated script, as a basis for designing the data models of LO for a LO repository at the level of lexical semantics has advantages. It may make it easier to deal with phenomena in the world of non-linguistic symbols, such as in the figure below. In some countries there are even variants of the same traffic sign, depending where it is used in terms of position on the road or whether as a traditional traffic sign or in a variable message sign (VMS) (Galinski 2011). Most of Japanese kanji occur – often with irregular readings – in appellations, i.e. proper names of people, places, organizations, buildings and other objects. Especially in languages with non-Latin writing systems the correct writing and reading of proper names is often posing great difficulties to the learner. Not to mention the correct sorting of names in such writing systems. Flying over Siberia to China one can read on the flight monitor the Chinese name for Novosibirsk starting with 新xin, which means ‘new’, translating the name element novo into Chinese. The rest of the name sibirsk=Siberia is transcribed by Chinese characters. Thus, proper names may have to be transcribed or (partly or fully) “translated”. In any case proper names – equally to terminological and lexicographical entities – can be important LO and can be treated with the same approach outlined above. The approach of ISO/TC 37 can serve as a model for all kinds of structured content. But as already mentioned above, the DataCategoryRegistry forlanguageresources (DCR) (ISO 12620:2009) of ISO/TC 37 should and in fact can be extended towards other kinds of structured content beyond lexicographical and terminological data, as well as towards new eApplication needs, such as for eLearning and eAccessibility / eInclusion purposes. For those, who are frightened by the complexity, there may be a – not so novel, but in reality rarely applied – way out: federation of resources
23Terminologija | 2012 | 19 of structured content. Applying the philosophy of Dublin Core, a certain stratum of data categories (being the same for different kinds of structured content) should be standardized and implemented in any resource thus guaranteeing content interoperability. The full degree of complexity is not implemented in one super-resource, but distributed over several resources of different kinds of structured content. This way of ensuring content interoperability having the re-purposing of entities of structured content in mind still needs further investigation. 6. FUTURE STANDARDIZATION NEEDS In the light of the above, ISO/DCR may need additional data categories. ISO/DCR today contains the basic and many extended data categories for terminological and lexicographical data. The terminological data model as well as the lexicographical data model is based on them. The use of data categories (not quite the same as metadata or dataelements) seems to be highly appropriate for structured content in general and for LO at the level of lexical semantics in particular. Besides, data categories can be distinguished into primary data categories and secondary data categories (ISO 10241-1:2011) (or even tertiary data categories). Primary data categories refer to core data, such as term, definition etc. Secondary data categories refer among others to attributes, such as preferred, admitted or deprecated (term). Tertiary data categories refer to additional information (if they are not regarded as secondary data categories), such as those referring to the elements of the source reference of terminological data (applicable also to other kind of structured content) as outlined in ISO 12615:2004 Bibliographicreferencesandsource identifiersforterminologywork. Primary data categories according to ISO 10241-1 adapted to LO are (compared to Table 1 in the Annex 1): - entry number – unique for the LO database entry; - (metadata category:) designative representation of structured content → i.e. terms, covering also synonyms, homonyms etc. to be extended towards any kind of designative representation of structured content, each of which should have its own entity ID, thus extending term autonomy towards representation autonomy; - (metadata category:) descriptive representation of structured content → i.e. definitions, explanations, contexts etc. to be extended
24 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning towards any kind of descriptive representation of structured content including non-verbal descriptive representations; - (metadata category:) example → if necessary to be split up into different types of example; - (metadata category:) note → if necessary to be split up into different types of note; - (metadata category:) source → if necessary to be split up into different types of source and their respective conditions for use. If the entity of designative representation of structured content is verbal, the repeatability by language applies, as well as the repeatability within language in case of synonyms (each of which also should get its own entity ID). If the entity of designative representation of structured content is nonverbal, the repeatability by field of use / application applies. There may also be a kind of repeatability within the field of use / application (such as speed limit road signs depending on their location alongside the road or in the middle of a highway in some countries, or road signs of same meaning differing in various states of the USA). There are variations of sign language within the same sign language as well as of Blissymbols. The above shows that repeatability should be according to the “community of use” (viz. locale in localization) rather than according to language. Of course the communityofuse can also be a language community, but the repeatability by and within locales provides more flexibility to cover all the phenomena encountered in structured content. 7. CONCLUSIONS Over the last 10 years the limitations of semantic interoperability under a computer science perspective have become obvious. In addition to technical (i.e. hardwareand software-related) and organizational interoperability, semantic interoperability should comprise syntactic, conceptual andpragmaticinteroperability. Content interoperability provides an even broader and generic approach with respect to the communicative representations of information and knowledge – it also takes full-fledged re-usability and re-purposability of entities of different kinds of structured content into account (Galinski, Van Isacker 2010). Re-purposability may comprise for instance the adaptation of existing terminological entries
25Terminologija | 2012 | 19 - as learning objects (LO) in eLearning or - for being used by persons with disabilities (PwD) for general communication and of course also for learning purposes. International Standards for content interoperability are the prerequisite for: - avoiding a huge duplication of efforts, - developing methods (incl. certification) and devices to assure content quality, - introducing content interoperability into educational and training schemes, - enabling many eApplications to re-use and re-purpose existing structured content extensively. Increasingly, all kinds of structured content should be re-usable and re-purposable across system platforms. In addition, more and more webbased cooperative and distributed methods for content development should be implemented in cooperation with interrelated fields under an integrative approach. In this connection the intensified use of mobile technologies will have a huge impact: eContent is becoming mContent (mobile content). In the course of these developments, it would be worthwhile to develop stronger methodological and system design relations between content resource management and corpus linguistics in order to make better use of - existing and future corpora with improved features e.g. for extracting individual items of structured content in a systematic way (also taking the occurrence frequencies depending on domain, level / register and application into account so that the extraction of learning objects in context would be improved); - items of structured content in existing and future repositories for the sake of improving the processing of corpora for additional new purposes such as didactics; - existing and emerging methods as well as tools in fields, which so far show a low degree of methodological interrelation and integration, for the sake of mutual benefits; - systematically (and by participatory efforts) created and maintained learning objects in particular for CLIL (content and language integrated learning).
26 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning Efforts in standardization (and cooperation in standardization) should be stepped up. The proper integration of AAC requirements in the data modelling of structured content would benefit not only those in society, who need it most urgently, but ultimately everybody. This requires a stronger emphasis on pertinent assistive technologies in ICT-related education and training as early as possible. Participants at the ICCHP 20103 Conference confirmed that existing training and formal studies are not sufficient – even if certified under given certification / attestation systems – with respect to the skills and qualifications necessary for becoming familiar with the issues involved in global content interoperability and particular in eAccessibility&eInclusion. The “Recommendation on software and content development principles 2010” (see Annex 2) was formulated in a special workshop at ICCHP 2010 and thereafter endorsed by several technical committees in the field of standardization, and in 2012 by the Management Group (MoU/MG) of the ITU-ISO-IEC-UN/ECE Memorandum of Understanding concerning eBusiness standardization. REFERENCES Galinski Ch. 2011: A sign equals thousand words. Consistency of traffic / road signs and verbal messages. – Infrastructureandsafetyinacollaborativeworld.Roadtrafficsafety. Evangelos Bekiaris, Marion Wiethoff, Evangelia Gaitanidou eds., Heidelberg, Dordrecht, London, New York: Springer, 263–284. Galinski Ch., Reineke D. 2011: Vor uns die Terminologieflut. – eDITion 2/2011, 8–12. Galinski Ch., Van Isacker K. 2010: Standards-based Content Resources: A Prerequisite for Content Integration and Content Interoperability. – Computershelpingpeoplewithspecialneeds.12thInternationalConference, ICCHP2010,Vienna,Austria,July2010.Proceedings, Part I., K. Miesenberger e.a. (eds.), Berlin, Heidelberg, New York: Springer, 573–579. Gonnissen L., Mornie G. 1983: Bestimmenunderkennenleichtgemacht.Vögel.Diewichtigstenheimischen Arten [Birds of Europe], translated from French by Sieglinde Summerer and Gerda Kurz, Zürich/Köln: Benziger Verlag. Gruber T. R., June 1993: A translation approach to portable ontology specifications (PDF). KnowledgeAcquisition5 (2), 199–220. – http://tomgruber.org/writing/ontolingua-kaj-1993.pdf. Hodges M., Okazaki T. Japanesekanjiflashcards. Series 2, volume 1, Tokyo: White Rabbit Press, s.a. Kitazawa T., Windhab R., Galinski Ch. 2007: UntersuchungvonFlashcardsystemen(elektronischeLernkarteien)fürCALL(computergestütztenSpracherwerb)undCASPLL(computergestütztenFachsprachenerwerb) aufAnfängerniveau.MitdemSchwerpunktaufJapanisch-LernenzumErwerbeinerhohenlexikalischen Sprachkompetenz–ohnedaslangfristigeZieleinermehrsprachigenFlashcard-Lernplattformzuvernachlässigen. McEnery T., Xiao R., Tono Y. 2006: Corpus-basedLanguageStudies:AnAdvancedResourceBook, London/ New York: Routledge. Schmitz K.-D. 2006: Data modeling: From terminology to other multilingual structured content. – TSTT 2006.InternationalConferenceonTerminology,StandardizationandTechnologyTransfer.Proceedings.Beijing,China,2006-08-25/26. Yuli Wang, Yu Wang, Ye Tian (eds.), Beijing: Encyclopedia of China Publishing House. 3 12th International conference on computers helping people with special needs (ICCHP 2010), Vienna, Austria, July 2010
27Terminologija | 2012 | 19 REFERENCED StaNDaRDS ISO/IEC Guide 2:2004 Standardizationandrelatedactivities–Generalvocabulary. ISO 639 (series) Codesfortherepresentationofnamesoflanguages. ISO/IEC 2382-36:2008 Informationtechnology–Vocabulary – Part 36: Learning,educationandtraining. ISO 3166 (series) Codesfortherepresentationofnamesofcountriesandtheirsubdivisions. ISO/TS 8000-1:2011 Dataquality – Part 1: Overview. ISO 10241-1:2011 Terminologicalentriesinstandards – Part 1: Generalrequirementsandexamplesofpresentation. ISO 10241-2:2012 Terminologicalentriesinstandards – Part 2: Adoptionofstandardizedterminologicalentries. ISO 12615:2004 Bibliographicreferencesandsourceidentifiersforterminologywork. ISO 12620:2009 Terminologyandotherlanguageandcontentresources–Specificationofdatacategoriesand managementofaDataCategoryRegistryforlanguageresources. ISO 15924:2004 Informationanddocumentation–Codesfortherepresentationofnamesofscripts. ISO/CD 17347:2012 OntologyIntegrationandInteroperability(OntoIOp)– Part 1: TheDistributedOntology Language(DOL). ISO/DIS 22274:2011 Systemstomanageterminology,knowledgeandcontent–Concept-relatedaspectsfordevelopingandinternationalizingclassificationsystems. IEEE (Institute of Electrical and Electronics Engineers, Inc.). LOM standard 1484.12.1-2002 DraftStandardforLearningObjectMetadata. tURINIO SĄVEIKUMO BŪtINYBĖ NaUDOJaNt StRUKtŪRIZUOtO tURINIO VIENEtUS KaIP ELEKtRONINIO MOKYMOSI OBJEKtUS PaKaRtOtINaI aR PaGaL KItOKIĄ PaSKIRtĮ Daugėja internetinių turinio platformų, siūlančių vartotojams vieną ar daugelį ištek lių, bet vis dar trūksta teorinio ir metodologinio tokios veiklos pagrindimo, mažai atsižvelgiama į geriausias turinio sąveikumo užtikrinimo patirtis. Tokių platformų skaičius augs ir toliau, nes vis daugiau išteklių kuriama ir naudojama taikant nepakankamai veiksmingus metodus. Kad būtų užtikrinta turinio kokybė, o pirmiausiai patikimumas, reikalingas įvairių priemonių derinimas (standartai, informacinės ir komunikacinės technologijos, sertifikavimas ir kt.). Elektroniniam mokymuisi svarbus skirtumas tarp bendrosios kalbos (angl. LGP) ir specialiosios kalbos (angl. LSP). Tam, kad būtų galima sukurti daugiakalbius, daugiamodalinius, daugelio paskirčių mokymosi objektų duomenų modelius, leksinės semantikos lygmeniu būtina tobulinti ribotas dabartinių duomenų bazių galimybes. Reikėtų daugiau dėmesio skirti: a) duomenų bazėse pateikiamų bendrosios kalbos ir specialiosios kalbos įrašų panašumui; b) duomenų bazių įrašų žodinių elementų tarimo nuorodoms; c) pateikimui nežodinių žymenų ir atvaizdų, kurie, atsižvelgiant į taikymo sritį, gali būti tokie pat svarbūs kaip žodiniai žymenys ir net labiau už juos pageidautini; d) leksinių vienetų dėmenims, pavyzdžiui, morfologiniams elementams; e) didesniems leksiniams vienetams, pavyzdžiui, bendrosios kalbos kolokacijoms ar specialiosios kalbos frazemoms; f) neįgaliųjų komunikacinėms reikmėms. Straipsnyje pateikiami argumentai dėl žemiau išvardytų dalykų būtinumo: - standartų, kuriuose būtų pateikti pasaulinio turinio sąveikumo užtikrinimo reikalavimai ir gairės, įskaitant elektroniniam mokymuisi keliamus reikalavimus;
28 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning - koordinuotos strategijos, skatinančios tokių standartų taikymą (pavyzdžiui, pasitelkiant sertifikavimą); - priemonių, užtikrinančių standartų laikymąsi plėtojant sistemas. Šie standartai ir priemonės galėtų padėti suvaldyti jau dabar daugiau ar mažiau chaotišką turinio plėtrą (sukeliančią didžiulį veiklos dubliavimą) ar bent jau leistų aiškiai išskirti patikimo turinio saugyklas. Gauta 2012-03-20 Christian Galinski International Information Centre for Terminology (Infoterm) Gymnasiumstrasse 50, 1190 Vienna, Austria E-mail [email protected] Blanca Stella Giraldo Pérez International Information Centre for Terminology (Infoterm) Gymnasiumstrasse 50, 1190 Vienna, Austria E-mail [email protected]
29Terminologija | 2012 | 19 AnneX 1 table 1 — Overview of data categories of a standardized terminological entry in accordance with ISO 10241-1 Primary data categoriesaSecondary data categoriesb (including administrative data and usage information) Name Mandatory/optional; repeatable/nonrepeatable entry number … Mandatory; nonrepeatable — termc (or string of five halfhigh dots “·····” or other slot holder sign, …) in the order preferred term(s), admitted term(s), deprecated term(s) Mandatory; repeatable grammatical information in accordance with the rules of the standardizing body: e.g. gender, number, part of speech language code or script code, or both geographical use (e.g. country code) pronunciation normative status letter symbol … Optional (unless the letter symbol is internationally standardized); repeatable language code or script code, or both geographical use (e.g. country code) normative status graphical symbol … Optional (unless the graphical symbol is internationally standardized); repeatable geographical use (e.g. country code) normative status definition … Mandatory (unless a nonverbal repre sentation is used by convention within the respective domain or subject); non-repeatable domain or subject, if necessary
30 Christian Galinski, Contentinteroperabilityasaprerequisitefor Blanca Stella Giraldo Pérez re-usingandre-purposingitemsofstructured contentaslearningobjectsineLearning Primary data categoriesaSecondary data categoriesb (including administrative data and usage information) Name Mandatory/optional; repeatable/nonrepeatable non-verbal representation … Mandatory (if exists – complementing a definition or used instead of a definition by convention within the respective domain or subject); nonrepeatable — example … Optional; repeatable — note to entry (including note to term, letter symbol, graphical symbol, definition, context, non-verbal representation, example, a given language section of a multilingual terminological entry or the entire terminological entry) … Optional; repeatable — source of entire terminological entry (including source of term, letter symbol, graphical symbol, definition, context, non-verbal representation, example) or any language section of a multilingual terminological entry Optional (unless the terminolo gical entry or a language section of a multilingual termino logical entry is taken from an external authoritative source); repeatable additional information relating to the source, such as page number or clause number a All primary data categories except the data category “entry number” are repeatable by language and, therefore, apply to multilingual standards. Additional rules for terminological entries in a multilingual terminology standard are given in Clause 7. b All secondary data categories are optional except where they are crucial for disambiguation (…) and in cases where multilingual information is included in one terminological entry, in which case the primary data categories shall be complemented by a code for names of language in accordance with ISO 639, if necessary in combination with codes for names of countries in accordance with ISO 3166 or codes for names of scripts in accordance with ISO 15924. c For simplicity, only “term” is specified in Table 1 although other verbal designations, such as any existing synonymous terms, variants, full forms, abbreviated forms, homographs, antonyms, as well as equivalent terms in other languages are included under this data category name.
